NVFlow: NVIDIA’s Recipe Engine for LLM Pipelines

A look at how NVFlow turns synthetic data, fine-tuning, and cluster execution into reusable workflows that speak NeMo and Slurm natively.

8 min read • View on GitHub • More from NVIDIA

A recipe book on a workbench opens into stacked workflow cards, YAML sheets, and a cluster job ticket, while a GPU rack hums in the background. The scene explains NVFlow’s core idea: workflow logic is packaged like a reusable recipe instead of scattered across scripts.
NVFlow treats a pipeline like a recipe. Prompts, stages, configs, and cluster settings travel together as one reusable unit.
Key Takeaways

NVFlow is interesting because it refuses to be a universal orchestrator. It narrows the problem to LLM pipelines inside the NeMo stack, then makes that narrow scope feel elegant: reusable recipes, declarative YAML, and Slurm jobs that chain themselves without a heavyweight control plane.

The Real Product Is the Recipe

That recipe metaphor is doing real work. In NVFlow, a workflow is not just a DAG of tasks. It is a bundled unit of prompts, stages, configs, and cluster settings that can be reused the way a training playbook can be reused.

Workflow orchestration framework for end-to-end synthetic data generation (SDG), training (SFT), and evaluation pipelines built on NVIDIA's NeMo ecosystem

NVIDIA, Project Owner · NVIDIA/nvflow

That changes the mental model. Instead of asking, "How do I schedule jobs?" you ask, "How do I package a repeatable method for synthetic data, fine-tuning, and evaluation?" NVFlow answers with a recipe layer that keeps the implementation close to the domain.

A rail holds a chain of job tickets, each stamped with a job ID and clipped to the next by visible dependency links. The image explains how NVFlow passes Slurm dependencies forward instead of relying on a central orchestrator heartbeat.
Slurm is not just the execution target. In NVFlow, it is the dependency engine.

Why It Is Not Trying to Be Airflow

Airflow, Prefect, Dagster, and Kubeflow solve broader orchestration problems. NVFlow does something narrower and more opinionated. It assumes you are already in the NeMo world, and it optimizes for the things that matter there: staged LLM workflows, Slurm execution, synthetic data generation, and reusable recipes.

ToolScopePrimary strengthScheduler modelBest fit
NVFlowNeMo LLM pipelinesRecipe-style reuse across SDG, SFT, and evaluationSlurm-native dependency chainingTeams already standardized on NVIDIA's stack
AirflowGeneral workflowsHuge ecosystem and mature schedulingCentral DAG schedulerBroad data platform coordination
PrefectGeneral workflowsDeveloper-friendly Python experienceCentral orchestration serviceFast-moving data teams
DagsterData assets and pipelinesStrong typing and asset lineageCentral orchestration serviceAnalytics and data engineering
KubeflowML platformKubernetes-native ML workflowsPlatform-managed orchestrationK8s-first ML infrastructure
NV-sflowDeclarative DAG executionGeneral YAML DAG running with pluggable backendsBackend-driven executionLightweight DAG execution outside NeMo

The point is not that NVFlow is more capable in every dimension. It is more fit for a specific job. When the domain is expensive, specialized, and already standardized around NVIDIA tooling, fit matters more than generality.

NVFlow’s registry turns short YAML labels into specific Python classes without name collisions.

# Conceptual shape of the registry pattern
@StageRegistry.register(recipe="finance", workflow="sdg", stage="generate_questions")
class GenerateQuestions(BaseStage):
    def execute(self, cluster, run_after=None):
        ...

# The same short stage name can exist elsewhere
@StageRegistry.register(recipe="finance", workflow="eval", stage="generate_questions")
class GenerateEvalQuestions(BaseStage):
    def execute(self, cluster, run_after=None):
        ...

The Registry That Makes Short Names Work

The extensibility trick is a hierarchy: Recipe -> Workflow -> Stage. NVFlow can keep YAML concise because the registry resolves a short stage name in context, not in isolation. The same name can exist in multiple workflows without collisions.

That matters because the YAML stays readable while the Python stays explicit. You get a human-friendly config surface and a deterministic runtime mapping to concrete stage classes.

Slurm Is the Orchestrator

This is the most distinctive part of the design. NVFlow passes Slurm job IDs through run_after, so one stage can submit the next stage with a dependency already attached. The scheduler enforces ordering, and NVFlow stays thin.

That removes the need for a central heartbeat and a persistent workflow brain. The orchestration logic becomes a chain of submissions, not a standing service that watches everything happen.

SEC filings enter a conveyor system and emerge as synthetic questions, grounded context, and filtered answers. The image explains the finance recipe as a transformation pipeline where each step makes the training data more trustworthy.
The finance recipe shows NVFlow doing more than scheduling. It turns raw filings into grounded training data.

YAML Inheritance Without the YAML Bloat

NVFlow uses _base_ inheritance to let one config inherit from another, then override only the parts that change. That keeps recipe files small even when the workflow branches across models, checkpoints, or execution modes.

The other smart move is dynamic stage expansion. Instead of manually writing a new evaluation stage for every checkpoint, the runner can expand a list into a set of repeated stages automatically. That is configuration reuse, but it is also guardrail design.

# conceptual pattern
_base_: sdg/base.yaml
model_path: /checkpoints/model-a
num_nodes: 8
checkpoints:
  - ckpt-1000
  - ckpt-2000
  - ckpt-3000

The Finance Recipe Is the Proof Case

The finance recipe is where the framework stops being abstract. It chains seed data creation, question generation, context mapping, and answer filtering into one pipeline. The important step is grounding: synthetic answers have to point back to the source text, or they are just polished noise.

That is why the example matters. NVFlow is not merely wiring tasks together. It is shaping the data quality loop that decides whether synthetic data becomes useful training data.

StageInputOutputWhy it matters
Create seed dataSource filingsSeed examplesStarts the chain with grounded material
Generate questionsSeed examplesCandidate questionsExpands coverage without manual authoring
Map questions to contextQuestions plus filingsGrounded context spansAnchors synthetic data to evidence
Filter answersContext plus candidate answersTraining-ready pairsRemoves weak or unsupported examples

What NVFlow Reveals About NeMo

NVFlow reads like infrastructure consolidation. It suggests NVIDIA wants one front door for synthetic data generation, fine-tuning, evaluation, and eventually adjacent tools like Data Designer and NeMo-Gym. That is a strategy signal, not just a repo choice.

The larger pattern is clear: a narrow orchestration layer can be more valuable than a broad one when the surrounding stack is already opinionated. NVFlow is not trying to win every orchestration contest. It is trying to make NeMo workflows feel native.

It exists to standardize how teams build, reproduce, and scale complex ML pipelines across domains with reusable stages and declarative workflows that run locally or on Slurm clusters.

NVIDIA, Project Owner · NVIDIA/nvflow