DataDesigner: NeMo Data Designer: The Synthetic Data Compiler That Builds Datasets as Graphs
NVIDIA's open-source framework does not just prompt an LLM. It schedules dependencies, mixes samplers with templates, validates the result, and lets tools fill in the blanks.
- NeMo Data Designer treats synthetic rows as dependency graphs, so earlier fields can constrain later ones.
- Its real innovation is compilation, because templates, samplers, validators, and tools are scheduled as one pipeline.
- The column system makes generation composable, which is why the framework can mix free text, structured output, and statistical fields.
- The cleanest comparison is with prompt-only generators, which can imitate format but cannot reliably preserve relational structure.
The row is not flat. It is a graph.
NeMo Data Designer starts from a simple but useful correction: a synthetic row is rarely just one prompt. In practice, one field often depends on another. A category shapes a review. A seed value shapes a sampler. A validator decides whether the row survives at all.
That is the point of the project. The hard problem is not making text. It is making rows that agree with themselves, at scale, across fields that have to line up.
Why prompt-only synthetic data hits a ceiling
If you generate each column with a separate prompt, the output looks flexible right up until the fields have to cohere. The review sounds plausible, but the rating and tone do not match. The address looks real, but the city and postal code drift apart. The more relational the dataset becomes, the faster a flat prompt loop starts to fail.
NeMo Data Designer responds by treating dependencies as first-class. It does not ask an LLM to remember the whole world in one pass. It builds the world in steps, with earlier outputs feeding later ones.
We just open-sourced NVIDIA NeMo Data Designer — the framework behind some of the top open synthetic datasets on Hugging Face. Nemotron-Personas, Nemotron-PII, Synthetic Text-to-SQL, and large parts of the pre/post-training data for Nemotron all came from this stack.
How Data Designer compiles a dataset
The central move in the codebase is to turn configuration into an execution plan. The builder assembles a `DataDesignerConfig`, Pydantic validates the column definitions, and Jinja2 templates are scanned for dependencies. Once those references are known, the engine can order the work as a directed acyclic graph instead of a blind loop.
That matters because ordering is not a detail. If a review template references `product_category`, that field has to exist first. If a structured column depends on a sampled value, the sampler has to run before the model call. Data Designer turns those rules into a schedule.
The column system is the real primitive
The repo is not one generator with a few options bolted on. It is a family of column types. There are samplers for statistical distribution, text columns for open-ended language, structured columns for schema-constrained output, and a plugin path for custom generators that extend the type system without rewriting the core.
That composition is the quiet strength of the project. It lets a dataset mix different kinds of realism. A field can be deterministic, probabilistic, or language-driven, yet still live inside the same row-level graph.
Our team at NVIDIA just open sourced NeMo Data Designer, a general framework for generating high-quality synthetic data from scratch or based on your own seed data. Data Designer has been my primary focus since my time at Gretel before we joined NVIDIA.
Validation is part of the product
A synthetic pipeline only becomes useful when quality checks are built into it. Data Designer treats validation as a normal stage, not a post-processing chore. Python checks can inspect values directly. SQL can enforce relational constraints. LLM-as-a-judge can score higher-level quality signals when the rule is more semantic than syntactic.
That changes the contract. The tool is not just saying, "here is a row." It is saying, "here is a row that passed a set of explicit checks." For teams generating training data, that is the difference between a demo and a pipeline.
Built for agents, not just humans
The repo's shape makes the intended workflow obvious. The `.agents` directory, `CLAUDE.md`, and MCP support point to a project that expects coding agents to help extend it. The system is meant to be operated, adapted, and stitched into larger toolchains, not just clicked through by a person in a notebook.
What it replaces, and where it wins
| Approach | Dependency awareness | Mixed generation modes | Validation hooks | Tool use / MCP | Best fit |
|---|---|---|---|---|---|
| Prompt-only synthetic data | Weak. Each field is generated in isolation. | Usually text only. | Usually external or manual. | Rare. | Fast prototypes and simple augmentation. |
| General synthetic data platforms | Moderate to strong, depending on the product. | Often broad, but opinionated. | Often present, sometimes tied to UI flows. | Limited or indirect. | Teams that want a packaged workflow. |
| NeMo Data Designer | Strong. Dependencies are compiled into the plan. | Sampler, text, structured output, and custom columns. | Built in with Python, SQL, and LLM-as-a-judge. | Supported through MCP and extensible plugins. | Relational, validated synthetic datasets for LLM work. |
The comparison is blunt for a reason. Prompt-only systems can produce plausible rows, but they struggle to preserve relationships. Broader synthetic data platforms can cover more use cases, but NeMo Data Designer's advantage is the compiler mindset: make the dependencies explicit, execute them in order, then validate the result before it ships.