openai/parameter-golf: The 16MB Weight Class for Language Models

A leaderboard repo where every byte, every second, and every optimizer choice changes the result.

9 min read • View on GitHub • More from openai

A golf-themed machine shop where a tiny mechanical brain is being weighed against a brass 16MB weight while a stopwatch ticks overhead. The scene explains that this repository treats model size and training time as hard constraints, not afterthoughts.
Parameter Golf turns model design into budgeting under pressure. The score is not just loss, it is whether the whole system fits the rules.
Key Takeaways

Most model repos ask how big a model can get. openai/parameter-golf asks how much language ability survives the shrink ray. The rule set is severe enough to change the design problem itself: fit the artifact inside 16MB, finish training under a tight wall-clock cap, and still post strong bits-per-byte on FineWeb.

Train the smallest LM you can that fits in 16MB. Best model wins!

OpenAI Parameter Golf Repository, Project Documentation · parameter-golf README

A leaderboard with a filesystem

The repo is not just code. It behaves like a living contest board, with baseline scripts at the root and a /records tree that preserves submission histories. The official entries live in track_10min_16mb, and the archive also keeps non-record attempts that overshoot the compute budget.

That structure matters because it shifts attention away from a single canonical model and toward the full recipe. Each submission is a self-contained artifact with its own train_gpt.py, metadata in submission.json, and logs that tell you how the run actually behaved.

A close-up drawer filled with tightly packed submission cards, clamps, and stamped envelopes arranged by a careful archivist. The image explains that the repository preserves many competing recipes, not just one model checkpoint.
The records folder is the real memory of the project. It stores the design space, not just the winner.

How to spend bytes like they are the scarce resource

The challenge rewards a different kind of cleverness than mainstream scaling. Instead of chasing raw parameter count, successful entries lean on a small set of moves that increase useful capacity per byte.

We're excited to see how optimizing for a parameter-constrained setting pushes people toward unique architectures (test-time compute, aggressive parameter tying, depth recurrence, low-rank training, ...), compression schemes (low precision, QAT, bitnets, novel tokenizers, ...), and other creative submissions (test-time training, long context, megakernels ...).

OpenAI Parameter Golf Repository, Project Documentation · parameter-golf README

That is why the interesting unit here is not the architecture by itself. It is the tradeoff. The winning idea is often less about inventing a new model family and more about packing familiar pieces into a tighter, better conditioned system.

The repo is a pipeline with a scorecard at the end. Training, compression, and archiving are all part of the same contest.

How the baseline makes the constraints real

The baseline script does the unglamorous work of turning the contest rules into code. It uses environment-driven hyperparameters for quick sweeps, streams custom data shards instead of pretending memory is free, and enforces a hard stop so training cannot quietly drift past the budget.

That is the hidden design lesson of the repo. When the wall-clock limit is real, the training loop stops being a neutral implementation detail and becomes part of the model itself. Speed, data access, and checkpoint handling are no longer support functions. They are performance.

Muon is a better fit than default scaling habits

A striking pattern in the challenge ecosystem is the use of Muon, an optimizer that leans on orthogonalization instead of treating every update like a generic step downhill. In a tiny model, that kind of geometry-aware update can matter more than the usual defaults because each parameter has to carry more of the load.

The point is not that Muon is magical. The point is that the repo rewards methods that preserve useful directions when there are fewer of them to spare. In a bigger-model world, inefficiency can hide. In this contest, it shows up immediately.

ProjectWhat it optimizesWhat that rewards
openai/parameter-golf16MB artifact, 10-minute training capCompression tricks, weight sharing, quantization, and fast training recipes
NanoGPT speedrunningShortest path to a target lossThroughput and optimization efficiency under a time budget
Typical LLM training repoQuality and scale with fewer explicit constraintsInfrastructure convenience and broad model capacity

That comparison is the whole story. Parameter Golf is a counterweight to the usual scaling narrative, because it asks what happens when the most valuable skill is subtraction. Remove bytes. Remove latency. Remove redundant paths. Keep only the parts that still learn.