SakanaAI/google-code-golf-2025: The Repo That Turns ARC Reasoning Into Byte-Count Combat
A deep dive into the judge, minifier, and prompt loop that squeeze correct solutions down to the smallest possible Python.
- This repo treats ARC solving as a compression loop, where correctness matters only if it survives aggressive byte shaving.
- The judge, minifier, compare command, and prompt generator form a feedback system, not a static toolkit.
- ARC is the right substrate for this experiment because grid transformations reward structural insight, then punish every unnecessary character.
- Compared with ordinary ARC solvers or general code golf tools, this project optimizes for the smallest correct program, not just a correct one.
Most solver repos ask whether the answer is correct. This one asks a harder question: how many bytes of reasoning can you remove before correctness falls apart? SakanaAI/google-code-golf-2025 turns ARC into a compression test, with the judge, minifier, and prompt loop all tuned to one target, the smallest working Python.
Why ARC becomes a code-golf problem
ARC is a good fit because the tasks are structured, but slippery. You are not classifying images or filling in labels. You are inferring a transformation from a few examples, which means the winning move is often a small, reusable pattern rather than a sprawling algorithm.
That makes Python a natural host and a brutal judge. Python can express grid tricks in a tight footprint, but every shortcut has to preserve the exact transformation. In this repo, size is not a vanity metric. It is part of the problem definition.
The loop that shrinks solutions
This is the repo's real novelty. It does not just hold solutions. It runs a system that tests a candidate, strips it down, measures the result, and pushes the lesson back into the next prompt. The judge verifies correctness. The minifier attacks syntax. The compare step tells you whether the shrinkage was worth it. The prompt generator keeps the search moving.
uv run judge -n 8
uv run compare
uv run prompt task001
uv run submit
Inside the judge, minifier, and submission scripts
The judge is the heartbeat. It runs candidate solutions against ARC tasks, checks their outputs, and gives the repo a hard yes or no. That matters because code golf only works when the evaluator is unforgiving. If the answer is even slightly wrong, the shortest program in the world is still useless.
The minifier is where the repo starts to feel like a specialized compiler toolchain. A script such as minify_solutions.py can strip whitespace, comments, and other syntactic fat from working code, while the shell scripts handle the dull but essential pieces around it, from trimming spaces to packaging submissions. The result is a workflow that treats every byte as a design choice.
The surrounding tooling matters too. uv keeps the Python environment fast and reproducible, GitHub Actions can automate judging, and the Kaggle submission flow connects local experiments to competition infrastructure. This is not a notebook repo. It is an execution environment.
Why this repo is different from standard ARC solvers
| Approach | Goal | Output | Constraint | Success metric |
|---|---|---|---|---|
| SakanaAI/google-code-golf-2025 | Solve ARC with minimum bytes | Python solutions plus judge and minifier loop | Correctness under byte pressure | Passing the judge with shorter code |
| General ARC solver | Solve tasks robustly | Rule engine or model predictions | Accuracy | Task solve rate |
| Program-evolution framework | Search for better programs over time | Iterated candidate code and feedback | Open-ended improvement | Performance on target tasks |
| Standard code golf tooling | Minimize program length | Golfed snippets and format tricks | Byte count and language quirks | Shortest valid program |
That table is the key distinction. A normal solver is rewarded for being right. A golf tool is rewarded for being short. This repo forces both to matter at once, which is why the workflow feels closer to an optimization lab than a puzzle archive.
A framework that combines Large Language Models (LLMs) with evolutionary algorithms to drive scientific discovery.
That line from SakanaAI's broader work helps explain the logic behind this repo. The company is not treating code as a one-shot artifact. It is treating it as something you can search, mutate, validate, and compress in a loop until the shape of the solution changes.
What this says about AI-assisted programming
The important lesson is not that smaller code is automatically better. It is that compactness can be turned into a measurable objective, and once an objective is measurable, you can build tooling around it. That is where the judge, the minifier, and the prompt generator become more interesting than the individual solutions they produce.
This repo sketches a hybrid programming future. Humans find the structure. LLMs accelerate the search. Automated judges keep everyone honest. The end state is not just a correct program. It is a correct program that has been forced through repeated rounds of compression until only the necessary logic remains.