NodeSynth: The Taxonomy Engine for AI Safety Testing
A Google Research prototype that decomposes policy into a concept graph, then recombines it into grounded synthetic eval prompts.
- NodeSynth treats safety evaluation as a coverage problem, not a prompt-writing problem.
- Its core move is to decompose policy into a taxonomy and synthesize prompts from the relationships between concepts.
- The Streamlit app matters less as a product than as a readable prototype of a policy-to-query compiler.
- Compared with manual red teaming or generic generation, NodeSynth trades spontaneity for inspectable coverage.
The blind spot in safety testing
Most safety evals start too late. A team writes a few prompts, runs them through a model, and calls the gap analysis complete. That works for demos, not for policy.
NodeSynth attacks the earlier step. It asks what concepts a policy implies, how those concepts relate, and where the untested edges live. The interesting shift is not that it makes more text. It makes the missing space visible.
Why NodeSynth starts with policy, not prompts
That order is the point. Policy text is messy, but it still encodes a contract. NodeSynth decomposes that contract into a taxonomy, links the branches, and then synthesizes queries from the intersections. It behaves like a coverage compiler, not a text spinner.
This is why the repository feels different from a prompt library. The unit of work is not a prompt. It is a branch that policy implies but human reviewers have not yet made explicit.
Inside the pipeline
The pipeline is easy to describe once you separate the stages:
- Decompose the policy into smaller concepts so the evaluation space is not trapped in broad language.
- Build a taxonomy that makes those concepts visible as nodes rather than vague categories.
- Map relationships between nodes so the system knows which intersections deserve tests.
- Synthesize query candidates from those intersections instead of from a generic prompt pool.
- Validate the outputs so duplicate, thin, or ungrounded cases do not quietly pass as coverage.
That last step matters. A taxonomy can create the illusion of completeness if it is too neat. NodeSynth only earns its keep if the decomposition is honest enough to expose what the first prompt draft would have missed.
Who built it, and what the repository signals
The repo reads like a Google Research prototype, not a shipping product. The codebase is small and flat, with a single Streamlit app, a requirements file, and lightweight documentation. That shape tells you a lot: this is an instrument for exploring a method, not a polished application with broad surface area.
The stack reinforces that read. streamlit suggests fast research iteration, pandas points to tabular manipulation, and plotly hints that the taxonomy is meant to be inspected visually. In other words, the point is not just generation. It is legibility.
Our workflow centers on a curated taxonomy of programming knowledge derived from large-scale annotation of the Nemotron‑Pretraining‑Code‑{v1,v2} datasets. This taxonomy encodes thousands of programming concepts organized hierarchically, from fundamental constructs (e.g., strings, recursion) to advanced algorithmic and data-structure patterns.
How it differs from the old way
NodeSynth sits in the middle of two familiar extremes. Manual red teaming is flexible but incomplete. Generic synthetic generation is scalable but often blurry. NodeSynth tries to turn coverage into something you can inspect, discuss, and improve.
| Axis | NodeSynth | Manual red teaming | Generic synthetic generation |
|---|---|---|---|
| Input source | Policy decomposed into concepts and relationships | Ad hoc human prompts and instincts | A prompt or seed topic |
| Coverage model | Explicit, inspectable taxonomy | Implicit in the tester's intuition | Broad but often shallow |
| Strength | Finds blind spots systematically | Adapts to nuance fast | Scales quickly |
| Failure mode | Depends on taxonomy quality | Misses unseen corners | Produces noisy, poorly grounded cases |
| Best use | Safety eval coverage | High-touch audits and adversarial review | Volume generation and rough exploration |
The comparison also explains why this kind of tool is interesting now. As teams try to evaluate more policies across more model behaviors, intuition alone stops scaling. Structure becomes a feature, not bureaucracy.
A prototype with a strong idea
NodeSynth is still a prototype, and it looks like one. That is not a knock. It means the public value is the method itself: policy decomposition as a practical unit of evaluation. Once you see safety as a coverage problem, the repository makes immediate sense.
The bigger takeaway is simple. NodeSynth is not trying to invent better prompts. It is trying to make the space of forgotten prompts legible. That is a more durable idea than any single interface.