`nv-sflow`: The YAML Compiler That Makes Slurm Feel Declarative

Replica expansion, late-bound expressions, readiness probes, and GPU-aware placement turn static workflow files into a runtime graph for HPC and LLM inference.

10 min read · NVIDIA/nv-sflow

A YAML sheet feeds a Slurm control console while task slips split into a fan of concrete jobs. The image explains that nv-sflow treats the workflow file as input, then assembles a live execution graph after allocation and expansion.
The central trick is compilation, not scripting.
Key Takeaways

The workflow file is not the workflow

The most interesting thing about nv-sflow is that the file you write is not the thing that runs. It is closer to source code. The YAML gets validated, resolved, expanded, and only then assembled into a live task graph.

workflow:
  variables:
    model: llama3-70b
    replicas: 2
  tasks:
    prefill:
      operator: srun
      script: python serve.py --model ${{ variables.model }}
      probe:
        type: tcp
        port: 8000
    decode:
      depends_on: prefill
      replicas: ${{ variables.replicas }}
      script: python decode.py --upstream ${{ tasks.prefill.server[0].nodes[0].ip_address }}

That delay is the point. In Slurm, node names, IPs, and GPU visibility often do not exist when you author the workflow. nv-sflow waits until the cluster can answer those questions, then turns declarations into concrete runtime state.

What nv-sflow is actually compiling

The project behaves like an assembly layer. A static config object becomes a runtime state object, replica counts fan out into concrete task instances, and task references can point to values that are not known until allocation time.

A workflow is not born complete. It becomes complete in stages, as the assembly layer resolves, expands, and allocates it.

This is a very Kubernetes-like mental model applied to Slurm. You describe intent up front, then let the runtime fill in the missing pieces once the scheduler has handed you real resources.

Why Slurm needed a new orchestration layer

Slurm is a strong scheduler, but it is not a modern workflow runtime. Native job scripts can launch work, but they leave dependency management, parameter sweeps, health checks, and placement conventions to the user.

That gap matters most when the job itself is dynamic. Benchmarking multiple models across multiple replica counts, or waiting for one service to become healthy before the next starts, quickly turns a shell script into a maintenance problem.

The three tricks that make it feel modern

The modern feel comes from three features working together: late-bound expressions, probe-gated dependencies, and GPU-aware placement. Each one pulls a fragile piece of orchestration out of shell glue and into the system.

Late-bound expressions

Values can stay unresolved until the workflow reaches runtime. That means a task can point at an upstream node address, a resolved GPU count, or another task’s metadata without forcing the author to know those values in advance.

Probes as gates

Probes make dependencies feel real instead of ceremonial. A downstream task does not start because the previous command exited cleanly. It starts when the expected port opens, the expected HTTP endpoint answers, or the log line you care about actually appears.

A close-up gate blocks a downstream task until a small signal lamp turns green and a log-line strip clears. The image explains how readiness probes act as real execution gates instead of simple job dependencies.
Dependencies become observable, not assumed.

GPU-aware placement

Placement is where nv-sflow stops being generic. The assembly layer tracks GPU indices across nodes, then assigns non-overlapping CUDA_VISIBLE_DEVICES values so two tasks sharing a node do not trip over each other.

Two tasks sit on the same node while separate GPU lanes are assigned to each one. The image explains that nv-sflow tracks device-level placement, not just whole-node allocation.
GPU placement is treated as a first-class scheduling problem.

Why the plugin split matters

The architecture is clean because it separates where a task runs from how it runs. Backends handle placement. Operators handle execution. Probes handle readiness. That split keeps the system from collapsing into one giant Slurm wrapper.

It also makes the same workflow model portable. A task can run locally, under Slurm, or through a different operator without changing the basic shape of the workflow file.

Why NVIDIA built this for inference-heavy workflows

The repository reads like it was built for benchmarking pressure. Replica-heavy sweeps, disaggregated prefill and decode style runs, and repeated tuning loops all benefit from a system that can reassemble work after allocation instead of before it.

That explains the emphasis on developer experience too. The CLI can export skills for AI coding assistants, which is a practical acknowledgment that workflow YAML is often written with machine help now. If the assistant needs context, the project would rather provide a reliable one.

How it compares to the usual suspects

nv-sflow is not trying to be the universal workflow answer. It is narrower than that. The right comparison is against tools readers already know, because the distinction is about layer and intent more than raw feature count.

Hand-written Slurm scripts are pinned to a wall on the left, while a declarative workflow file feeds a runtime graph on the right. The image explains the shift from brittle shell stitching to compiled workflow assembly.
The old way is stitched by hand. The new way is assembled from intent.
Dimensionnv-sflowHand-written SlurmNextflowSnakemakePyxis
Workflow modelNative runtime graphManual job chainingNative DAGNative DAGNot a workflow layer
Runtime dataLate-boundShell variables and parsingEngine-managed paramsWildcards and paramsNo
Readiness gatesHTTP, TCP, and log probesManual pollingCustom codeCustom codeNo
GPU placementTopology-awareWhatever the script encodesResource requestsResource requestsNo
BackendsLocal and SlurmUsually one script per environmentMany executorsMany executorsSlurm container plugin
Best fitInference sweeps and HPC orchestrationSmall ad hoc clustersPortable scientific pipelinesRule-based pipelinesRunning containers under Slurm

The bigger bet

The bigger bet is cultural. HPC workflows are drifting toward declarative assembly, just as cloud systems did before them. nv-sflow shows what that looks like when the scheduler stays Slurm and the workflow file becomes the thing you compile.