Product R&D · Synthetic data
Test realistic workflows without copying source values.
Decoy is an in-development foundation for generating representative synthetic data inside the customer security boundary and testing exactly what is allowed to leave it.
Decoy is being built to run inside your security boundary as a single binary with read-only access to approved data sources. In the intended workflow, an authorized AI service analyzes schemas and relationships, then drafts provider rules for synthetic output. The customer defines what may cross the boundary and verifies the result against agreed fidelity and leakage tests.
Why realistic data is hard
Building enterprise applications requires realistic data with correct schemas, plausible distributions, proper relationships, and enough volume to test behavior. The source data may contain financial, healthcare, acquisition, or personnel information. AI-assisted development adds another boundary to govern. Decoy is designed to reduce source-record exposure and make the boundary testable; it does not replace privacy review, security controls, or re-identification risk analysis.
What Decoy Is Being Built to Do
Approved AI Analysis
An approved AI service in your environment can analyze schemas, field semantics, relationships, embedded structures, and data quirks. Provider, permissions, retention, and human review remain explicit operating constraints.
Provider Code Generation
AI drafts custom data-generation rules for approved fields, relationships, and distributions. Validation detects wrong types or broken constraints; the workflow revises the rules and retains evidence for accountable human review.
Full-Scale Synthetic Output
Configurable 1:1 volume targets. If your JSON has 1.2M lines, the synthetic version can be required to have 1.2M lines, with agreed field lengths, nested structures, and referential integrity. Workload tests then measure how closely the synthetic data represents the required behavior.
Documentation Package
The package includes schema maps, field profiles, relationship graphs, a data-quirks log, and a provider reference. It gives developers a durable guide to data they cannot inspect directly and supports onboarding and transition.
| Boundary | Contents |
|---|---|
| Proposed customer-side work | Approved data sources, Decoy binary, approved AI service, validation loop |
| Outputs reviewed before transfer | Synthetic dataset, documentation, provider code |
| Proposed development-side work | Application development, testing, and AI-assisted work using customer-approved synthetic outputs |
- Random values by type
- No cross-table relationships
- Uniform distributions
- Fixed row counts
- Breaks on first real query
- Approved analysis of field semantics
- Required referential integrity
- Agreed distribution tolerances
- Target production-scale volume
- Measured workflow representativeness
Target Data Sources
Target sources include S3 buckets, relational databases, data lakes, Microsoft Dataverse, JSON, YAML, Parquet, CSV, XML, and other structured formats. Configuration identifies approved connections and optional instructions for corner cases. The intended deployment grants read-only source permissions; the acceptance plan verifies that boundary.
Define and Test the Security Boundary
Define the approved environment, read-only sources, prohibited value leakage, schema and relationship fidelity, statistical tolerances, target scale, and audit evidence before delivery. The customer verifies those conditions inside its boundary. The configured capability, synthetic data, provider rules, schema documentation, accessibility and operating evidence, deployment instructions, and validation results form a portable handoff package.
Availability: Decoy is in development and currently offered only through a bounded Intelligrit engagement. Scope, controls, and pricing depend on data volume, sensitivity, and the approved environment. Let's Talk to discuss your needs.