`nv-sflow`: The YAML Compiler That Makes Slurm Feel Declarative
Replica expansion, late-bound expressions, readiness probes, and GPU-aware placement turn static workflow files into a runtime graph for HPC and LLM inference.
- `nv-sflow` treats YAML as source input and builds the real workflow only after Slurm allocation, resolution, and replica expansion.
- Late-bound expressions, readiness probes, and GPU-aware placement are the features that make the system feel modern without replacing Slurm.
- The plugin split keeps execution operators separate from backends, which lets the same authoring model stretch from local runs to cluster jobs.
- The project is narrow on purpose, aiming at reproducible HPC and inference workflows rather than becoming a general workflow platform.
The workflow file is not the workflow
The most interesting thing about nv-sflow is that the file you write is not the thing that runs. It is closer to source code. The YAML gets validated, resolved, expanded, and only then assembled into a live task graph.
workflow:
variables:
model: llama3-70b
replicas: 2
tasks:
prefill:
operator: srun
script: python serve.py --model ${{ variables.model }}
probe:
type: tcp
port: 8000
decode:
depends_on: prefill
replicas: ${{ variables.replicas }}
script: python decode.py --upstream ${{ tasks.prefill.server[0].nodes[0].ip_address }}
That delay is the point. In Slurm, node names, IPs, and GPU visibility often do not exist when you author the workflow. nv-sflow waits until the cluster can answer those questions, then turns declarations into concrete runtime state.
What nv-sflow is actually compiling
The project behaves like an assembly layer. A static config object becomes a runtime state object, replica counts fan out into concrete task instances, and task references can point to values that are not known until allocation time.
This is a very Kubernetes-like mental model applied to Slurm. You describe intent up front, then let the runtime fill in the missing pieces once the scheduler has handed you real resources.
Why Slurm needed a new orchestration layer
Slurm is a strong scheduler, but it is not a modern workflow runtime. Native job scripts can launch work, but they leave dependency management, parameter sweeps, health checks, and placement conventions to the user.
That gap matters most when the job itself is dynamic. Benchmarking multiple models across multiple replica counts, or waiting for one service to become healthy before the next starts, quickly turns a shell script into a maintenance problem.
The three tricks that make it feel modern
The modern feel comes from three features working together: late-bound expressions, probe-gated dependencies, and GPU-aware placement. Each one pulls a fragile piece of orchestration out of shell glue and into the system.
Late-bound expressions
Values can stay unresolved until the workflow reaches runtime. That means a task can point at an upstream node address, a resolved GPU count, or another task’s metadata without forcing the author to know those values in advance.
Probes as gates
Probes make dependencies feel real instead of ceremonial. A downstream task does not start because the previous command exited cleanly. It starts when the expected port opens, the expected HTTP endpoint answers, or the log line you care about actually appears.
GPU-aware placement
Placement is where nv-sflow stops being generic. The assembly layer tracks GPU indices across nodes, then assigns non-overlapping CUDA_VISIBLE_DEVICES values so two tasks sharing a node do not trip over each other.
Why the plugin split matters
The architecture is clean because it separates where a task runs from how it runs. Backends handle placement. Operators handle execution. Probes handle readiness. That split keeps the system from collapsing into one giant Slurm wrapper.
It also makes the same workflow model portable. A task can run locally, under Slurm, or through a different operator without changing the basic shape of the workflow file.
Why NVIDIA built this for inference-heavy workflows
The repository reads like it was built for benchmarking pressure. Replica-heavy sweeps, disaggregated prefill and decode style runs, and repeated tuning loops all benefit from a system that can reassemble work after allocation instead of before it.
That explains the emphasis on developer experience too. The CLI can export skills for AI coding assistants, which is a practical acknowledgment that workflow YAML is often written with machine help now. If the assistant needs context, the project would rather provide a reliable one.
How it compares to the usual suspects
nv-sflow is not trying to be the universal workflow answer. It is narrower than that. The right comparison is against tools readers already know, because the distinction is about layer and intent more than raw feature count.
| Dimension | nv-sflow | Hand-written Slurm | Nextflow | Snakemake | Pyxis |
|---|---|---|---|---|---|
| Workflow model | Native runtime graph | Manual job chaining | Native DAG | Native DAG | Not a workflow layer |
| Runtime data | Late-bound | Shell variables and parsing | Engine-managed params | Wildcards and params | No |
| Readiness gates | HTTP, TCP, and log probes | Manual polling | Custom code | Custom code | No |
| GPU placement | Topology-aware | Whatever the script encodes | Resource requests | Resource requests | No |
| Backends | Local and Slurm | Usually one script per environment | Many executors | Many executors | Slurm container plugin |
| Best fit | Inference sweeps and HPC orchestration | Small ad hoc clusters | Portable scientific pipelines | Rule-based pipelines | Running containers under Slurm |
The bigger bet
The bigger bet is cultural. HPC workflows are drifting toward declarative assembly, just as cloud systems did before them. nv-sflow shows what that looks like when the scheduler stays Slurm and the workflow file becomes the thing you compile.