openai/parameter-golf: The 16MB Weight Class for Language Models
A leaderboard repo where every byte, every second, and every optimizer choice changes the result.
- Parameter Golf treats storage and wall-clock time as the main objective, so model design becomes a budgeting problem instead of a scaling race.
- The repo is really a tournament archive, and its records folder matters because it preserves the exact recipes behind each submission.
- Muon, quantization, parameter tying, and hashing matter here because they increase useful capacity per byte, not because they are fashionable.
- The challenge is a sharp contrast to standard LLM work because it rewards compact systems that are fast to train and easy to reproduce.
Most model repos ask how big a model can get. openai/parameter-golf asks how much language ability survives the shrink ray. The rule set is severe enough to change the design problem itself: fit the artifact inside 16MB, finish training under a tight wall-clock cap, and still post strong bits-per-byte on FineWeb.
Train the smallest LM you can that fits in 16MB. Best model wins!
A leaderboard with a filesystem
The repo is not just code. It behaves like a living contest board, with baseline scripts at the root and a /records tree that preserves submission histories. The official entries live in track_10min_16mb, and the archive also keeps non-record attempts that overshoot the compute budget.
That structure matters because it shifts attention away from a single canonical model and toward the full recipe. Each submission is a self-contained artifact with its own train_gpt.py, metadata in submission.json, and logs that tell you how the run actually behaved.
How to spend bytes like they are the scarce resource
The challenge rewards a different kind of cleverness than mainstream scaling. Instead of chasing raw parameter count, successful entries lean on a small set of moves that increase useful capacity per byte.
- Parameter tying lets one set of weights do more work than a conventional stack would allow.
- Quantization compresses the artifact so the model can survive the 16MB ceiling.
- Hashing tricks shrink large embedding tables and vocab-related overhead.
- Geometry-aware optimizers such as Muon help narrow models keep learning signal alive.
We're excited to see how optimizing for a parameter-constrained setting pushes people toward unique architectures (test-time compute, aggressive parameter tying, depth recurrence, low-rank training, ...), compression schemes (low precision, QAT, bitnets, novel tokenizers, ...), and other creative submissions (test-time training, long context, megakernels ...).
That is why the interesting unit here is not the architecture by itself. It is the tradeoff. The winning idea is often less about inventing a new model family and more about packing familiar pieces into a tighter, better conditioned system.
How the baseline makes the constraints real
The baseline script does the unglamorous work of turning the contest rules into code. It uses environment-driven hyperparameters for quick sweeps, streams custom data shards instead of pretending memory is free, and enforces a hard stop so training cannot quietly drift past the budget.
That is the hidden design lesson of the repo. When the wall-clock limit is real, the training loop stops being a neutral implementation detail and becomes part of the model itself. Speed, data access, and checkpoint handling are no longer support functions. They are performance.
Muon is a better fit than default scaling habits
A striking pattern in the challenge ecosystem is the use of Muon, an optimizer that leans on orthogonalization instead of treating every update like a generic step downhill. In a tiny model, that kind of geometry-aware update can matter more than the usual defaults because each parameter has to carry more of the load.
The point is not that Muon is magical. The point is that the repo rewards methods that preserve useful directions when there are fewer of them to spare. In a bigger-model world, inefficiency can hide. In this contest, it shows up immediately.
| Project | What it optimizes | What that rewards |
|---|---|---|
| openai/parameter-golf | 16MB artifact, 10-minute training cap | Compression tricks, weight sharing, quantization, and fast training recipes |
| NanoGPT speedrunning | Shortest path to a target loss | Throughput and optimization efficiency under a time budget |
| Typical LLM training repo | Quality and scale with fewer explicit constraints | Infrastructure convenience and broad model capacity |
That comparison is the whole story. Parameter Golf is a counterweight to the usual scaling narrative, because it asks what happens when the most valuable skill is subtraction. Remove bytes. Remove latency. Remove redundant paths. Keep only the parts that still learn.