NodeSynth: The Taxonomy Engine for AI Safety Testing

A Google Research prototype that decomposes policy into a concept graph, then recombines it into grounded synthetic eval prompts.

8 min read • View on GitHub • More from google-research

A policy document feeds into a branching taxonomy tree, which then drops prompt cards on the far side. The image explains that NodeSynth treats coverage as a structured transformation, not a list of one-off examples.
NodeSynth turns policy language into a testable map of concepts, then uses that map to generate grounded prompts.
Key Takeaways

The blind spot in safety testing

Most safety evals start too late. A team writes a few prompts, runs them through a model, and calls the gap analysis complete. That works for demos, not for policy.

NodeSynth attacks the earlier step. It asks what concepts a policy implies, how those concepts relate, and where the untested edges live. The interesting shift is not that it makes more text. It makes the missing space visible.

Why NodeSynth starts with policy, not prompts

That order is the point. Policy text is messy, but it still encodes a contract. NodeSynth decomposes that contract into a taxonomy, links the branches, and then synthesizes queries from the intersections. It behaves like a coverage compiler, not a text spinner.

The central idea is a pipeline, not a prompt list: policy becomes taxonomy, then taxonomy becomes grounded evaluation coverage.

This is why the repository feels different from a prompt library. The unit of work is not a prompt. It is a branch that policy implies but human reviewers have not yet made explicit.

Inside the pipeline

A close-up of two taxonomy branches crossing, with a single prompt card stamped into the intersection. The image explains how NodeSynth uses relationships between concepts to produce grounded test cases.
The synthesis step happens at the intersection of concepts, where a relationship becomes a testable scenario.

The pipeline is easy to describe once you separate the stages:

  1. Decompose the policy into smaller concepts so the evaluation space is not trapped in broad language.
  2. Build a taxonomy that makes those concepts visible as nodes rather than vague categories.
  3. Map relationships between nodes so the system knows which intersections deserve tests.
  4. Synthesize query candidates from those intersections instead of from a generic prompt pool.
  5. Validate the outputs so duplicate, thin, or ungrounded cases do not quietly pass as coverage.

That last step matters. A taxonomy can create the illusion of completeness if it is too neat. NodeSynth only earns its keep if the decomposition is honest enough to expose what the first prompt draft would have missed.

Who built it, and what the repository signals

The repo reads like a Google Research prototype, not a shipping product. The codebase is small and flat, with a single Streamlit app, a requirements file, and lightweight documentation. That shape tells you a lot: this is an instrument for exploring a method, not a polished application with broad surface area.

The stack reinforces that read. streamlit suggests fast research iteration, pandas points to tabular manipulation, and plotly hints that the taxonomy is meant to be inspected visually. In other words, the point is not just generation. It is legibility.

Our workflow centers on a curated taxonomy of programming knowledge derived from large-scale annotation of the Nemotron‑Pretraining‑Code‑{v1,v2} datasets. This taxonomy encodes thousands of programming concepts organized hierarchically, from fundamental constructs (e.g., strings, recursion) to advanced algorithmic and data-structure patterns.

Joseph Jennings and Brandon Norick, Researchers at NVIDIA · NVIDIA Code Concepts

How it differs from the old way

NodeSynth sits in the middle of two familiar extremes. Manual red teaming is flexible but incomplete. Generic synthetic generation is scalable but often blurry. NodeSynth tries to turn coverage into something you can inspect, discuss, and improve.

AxisNodeSynthManual red teamingGeneric synthetic generation
Input sourcePolicy decomposed into concepts and relationshipsAd hoc human prompts and instinctsA prompt or seed topic
Coverage modelExplicit, inspectable taxonomyImplicit in the tester's intuitionBroad but often shallow
StrengthFinds blind spots systematicallyAdapts to nuance fastScales quickly
Failure modeDepends on taxonomy qualityMisses unseen cornersProduces noisy, poorly grounded cases
Best useSafety eval coverageHigh-touch audits and adversarial reviewVolume generation and rough exploration

The comparison also explains why this kind of tool is interesting now. As teams try to evaluate more policies across more model behaviors, intuition alone stops scaling. Structure becomes a feature, not bureaucracy.

A prototype with a strong idea

NodeSynth is still a prototype, and it looks like one. That is not a knock. It means the public value is the method itself: policy decomposition as a practical unit of evaluation. Once you see safety as a coverage problem, the repository makes immediate sense.

The bigger takeaway is simple. NodeSynth is not trying to invent better prompts. It is trying to make the space of forgotten prompts legible. That is a more durable idea than any single interface.