DataDesigner: NeMo Data Designer: The Synthetic Data Compiler That Builds Datasets as Graphs

NVIDIA's open-source framework does not just prompt an LLM. It schedules dependencies, mixes samplers with templates, validates the result, and lets tools fill in the blanks.

9 min read • View on GitHub • More from NVIDIA-NeMo

A drafting desk where a dataset is assembled like a blueprint. Blank cards and thin dependency lines form a branching structure, while a hand places the next piece only after the earlier pieces are aligned. It explains that synthetic data here is not a flat prompt loop, but a coordinated graph of fields.
NeMo Data Designer treats row generation like construction, not improvisation.
Key Takeaways

The row is not flat. It is a graph.

NeMo Data Designer starts from a simple but useful correction: a synthetic row is rarely just one prompt. In practice, one field often depends on another. A category shapes a review. A seed value shapes a sampler. A validator decides whether the row survives at all.

That is the point of the project. The hard problem is not making text. It is making rows that agree with themselves, at scale, across fields that have to line up.

Why prompt-only synthetic data hits a ceiling

If you generate each column with a separate prompt, the output looks flexible right up until the fields have to cohere. The review sounds plausible, but the rating and tone do not match. The address looks real, but the city and postal code drift apart. The more relational the dataset becomes, the faster a flat prompt loop starts to fail.

NeMo Data Designer responds by treating dependencies as first-class. It does not ask an LLM to remember the whole world in one pass. It builds the world in steps, with earlier outputs feeding later ones.

We just open-sourced NVIDIA NeMo Data Designer — the framework behind some of the top open synthetic datasets on Hugging Face. Nemotron-Personas, Nemotron-PII, Synthetic Text-to-SQL, and large parts of the pre/post-training data for Nemotron all came from this stack.

Alexander Watson · LinkedIn post
A split composition comparing two ways to generate data. One side shows a flat stream of identical prompts producing generic rows, while the other shows a branching build process where one field unlocks the next. It explains why dependency-aware generation preserves coherence better than one-prompt-per-column workflows.
The difference is not style. It is control over relationships.

How Data Designer compiles a dataset

The central move in the codebase is to turn configuration into an execution plan. The builder assembles a `DataDesignerConfig`, Pydantic validates the column definitions, and Jinja2 templates are scanned for dependencies. Once those references are known, the engine can order the work as a directed acyclic graph instead of a blind loop.

That matters because ordering is not a detail. If a review template references `product_category`, that field has to exist first. If a structured column depends on a sampled value, the sampler has to run before the model call. Data Designer turns those rules into a schedule.

The engine compiles field relationships into an ordered plan, then validates the row after generation.

A tight mechanical close-up of a template strip feeding into a watch-like scheduler. Small gears move the work into separate chambers for sampling, text generation, and validation. It explains the compile step, where configuration becomes ordered execution.
Templates are parsed first, then turned into a schedule of generation steps.

The column system is the real primitive

The repo is not one generator with a few options bolted on. It is a family of column types. There are samplers for statistical distribution, text columns for open-ended language, structured columns for schema-constrained output, and a plugin path for custom generators that extend the type system without rewriting the core.

That composition is the quiet strength of the project. It lets a dataset mix different kinds of realism. A field can be deterministic, probabilistic, or language-driven, yet still live inside the same row-level graph.

Our team at NVIDIA just open sourced NeMo Data Designer, a general framework for generating high-quality synthetic data from scratch or based on your own seed data. Data Designer has been my primary focus since my time at Gretel before we joined NVIDIA.

Johnny Greco, Research Scientist @ NVIDIA · LinkedIn post

Validation is part of the product

A synthetic pipeline only becomes useful when quality checks are built into it. Data Designer treats validation as a normal stage, not a post-processing chore. Python checks can inspect values directly. SQL can enforce relational constraints. LLM-as-a-judge can score higher-level quality signals when the rule is more semantic than syntactic.

That changes the contract. The tool is not just saying, "here is a row." It is saying, "here is a row that passed a set of explicit checks." For teams generating training data, that is the difference between a demo and a pipeline.

Built for agents, not just humans

The repo's shape makes the intended workflow obvious. The `.agents` directory, `CLAUDE.md`, and MCP support point to a project that expects coding agents to help extend it. The system is meant to be operated, adapted, and stitched into larger toolchains, not just clicked through by a person in a notebook.

A central registry of interchangeable modules with several plug-in branches entering from the sides. One branch reaches outward to an external tool box, showing how tool calls can be added without changing the core. It explains extensibility and agent-friendly integration.
The extension points are part of the design, not an afterthought.

What it replaces, and where it wins

ApproachDependency awarenessMixed generation modesValidation hooksTool use / MCPBest fit
Prompt-only synthetic dataWeak. Each field is generated in isolation.Usually text only.Usually external or manual.Rare.Fast prototypes and simple augmentation.
General synthetic data platformsModerate to strong, depending on the product.Often broad, but opinionated.Often present, sometimes tied to UI flows.Limited or indirect.Teams that want a packaged workflow.
NeMo Data DesignerStrong. Dependencies are compiled into the plan.Sampler, text, structured output, and custom columns.Built in with Python, SQL, and LLM-as-a-judge.Supported through MCP and extensible plugins.Relational, validated synthetic datasets for LLM work.

The comparison is blunt for a reason. Prompt-only systems can produce plausible rows, but they struggle to preserve relationships. Broader synthetic data platforms can cover more use cases, but NeMo Data Designer's advantage is the compiler mindset: make the dependencies explicit, execute them in order, then validate the result before it ships.