robust-kbench: The benchmark that refuses to be gamed
Sakana AI’s robust-kbench turns CUDA evaluation into a robustness test, not a leaderboard trick.
- robust-kbench is built to catch CUDA kernels that only look good under narrow test conditions.
- Its filtering pipeline turns task selection into a statistical gate that rejects brittle benchmarks before they matter.
- The repo pairs evaluation with an agentic compile-run-profile loop, so verification and optimization happen in the same system.
- The real comparison is not custom CUDA versus Torch alone, it is custom CUDA versus modern compiler output under harder conditions.
Most CUDA benchmarks reward the wrong kind of clever. If a kernel only wins on one input shape, one weight scale, or one hand-picked initialization, the score is decoration, not evidence. robust-kbench exists to make that failure mode expensive.
Why Sakana AI built a benchmark that distrusts itself
The repo comes out of Sakana AI’s push into agentic systems, where models do not just write code, they revise, compile, profile, and try again. The paper frames the problem directly: existing kernel generation benchmarks have exploitable loopholes and too little diversity in testing conditions. In other words, a model can win by learning the test instead of learning the task.
A comprehensive benchmark suite designed to evaluate and validate CUDA kernels generated by Large Language Models (LLMs). This benchmark addresses the limitations of existing kernel benchmarks by implementing robust evaluation criteria that prevent LLMs from exploiting benchmark settings.
That framing changes the job of the benchmark. It is not just a scorekeeper. It is a filter for brittle problem setups, a harness for JIT-compiled kernels, and a guardrail against reward hacking.
The filter is the product
The most interesting code lives in robust_kbench/filter/forward.py. Before a task is allowed into the benchmark, the repository checks whether the task itself is robust enough to deserve attention. The pipeline looks for output ranges that are not trivial, output standard deviation that actually moves with the inputs, and variation across axes so a lazy kernel cannot ignore part of the tensor and still look correct.
There is also an LLM sanity pass. Instead of trusting reference code blindly, the project uses a model to look for redundancy or inefficiency in the PyTorch baseline. That is a sharp move. It treats the benchmark source material as something to audit, not something to canonize.
A tight loop from Python to CUDA and back
The architecture is split cleanly. tasks/ holds the problem definitions and reference forwards. robust_kbench/ contains the filtering logic, sandboxing, and execution primitives. Root scripts such as run_filter.py and run_kernel.py stitch the pipeline together. Python does the orchestration, CUDA is the target, and PyTorch’s JIT extension path keeps the feedback loop short.
robust-kbench/
├── tasks/
│ ├── mnist_linear/
│ │ ├── func_forward.py
│ │ └── config_forward.json
│ └── ...
├── robust_kbench/
│ ├── filter/
│ ├── primitives/
│ ├── sandbox/
│ └── parallel.py
├── run_filter.py
├── run_kernel.py
└── setup.py
The executor design matters as much as the filters. The repo uses process spawning, isolated temp directories, and GPU cleanup routines so multiple kernels can compile without stepping on each other. That sounds mundane until you have watched nvcc collide with itself in a parallel job.
Recent advances in large language models (LLMs) demonstrate their effectiveness in scaling test-time compute for software engineering tasks. However, these ap proaches often focus on high-level solutions, with limited attention to optimizing low-level CUDA kernel implementations. Additionally, existing kernel generation benchmarks suffer from exploitable loopholes and insufficient diversity in testing conditions, hindering true generalization assessment. To address these limitations, we introduce robust-kbench, a new benchmark for rigorous evaluation of kernel performance and correctness across varied scenarios.
Why this is more than another kernel benchmark
| Question | robust-kbench | KernelBench or manual profiling |
|---|---|---|
| What is being protected against? | Benchmark gaming through narrow shapes, single initializations, and brittle edge cases. | Mostly accidental overfitting to the test setup, or human bias during manual tuning. |
| What is the core mechanism? | Statistical filters, LLM-based sanity checks, and repeated evaluation across varied conditions. | Randomized tests, timing rules, or expert judgment applied after the fact. |
| What is the output? | A benchmark plus an agentic evaluation loop for compiling and profiling kernels. | Either a benchmark score or a faster kernel, but usually not both in one system. |
| What is the key comparison? | Custom CUDA versus PyTorch native and Torch compile. | Custom CUDA versus a hand-tuned baseline, often without the same robustness gate. |
That comparison is the point. robust-kbench does not just ask whether a hand-written kernel beats a vanilla PyTorch implementation. It asks whether the kernel still wins after the benchmark has been diversified enough to remove cheap shortcuts. If Torch compile already does the job, the bar gets higher, which is exactly what a serious benchmark should do.
The repo also signals a broader shift in systems work. The interesting unit is no longer just a kernel or a compiler pass. It is the loop that can generate code, catch its own failures, and re-run under stricter conditions until the result is trustworthy.
As always, thank you to the community –– including teams from Cognition AI, Meta, METR, Nvidia, Prime Intellect, SakanaAI, and many others whose work we may not yet be aware of –– for working on the benchmark and helping surface both its strengths and limitations. Your feedback directly informed many of the improvements in v0.1 and we’re excited to see what you’ll build with the new version.