SakanaAI/google-code-golf-2025: The Repo That Turns ARC Reasoning Into Byte-Count Combat

A deep dive into the judge, minifier, and prompt loop that squeeze correct solutions down to the smallest possible Python.

8 min read • View on GitHub • More from SakanaAI

A massive mechanical press crushes a messy sheet of grid logic into a narrow strip of dense Python code. The scene explains the article's core idea: the repository treats reasoning as something that can be compressed without losing correctness.
The project is less a solver than a compression machine for solutions.
Key Takeaways

Most solver repos ask whether the answer is correct. This one asks a harder question: how many bytes of reasoning can you remove before correctness falls apart? SakanaAI/google-code-golf-2025 turns ARC into a compression test, with the judge, minifier, and prompt loop all tuned to one target, the smallest working Python.

Why ARC becomes a code-golf problem

ARC is a good fit because the tasks are structured, but slippery. You are not classifying images or filling in labels. You are inferring a transformation from a few examples, which means the winning move is often a small, reusable pattern rather than a sprawling algorithm.

That makes Python a natural host and a brutal judge. Python can express grid tricks in a tight footprint, but every shortcut has to preserve the exact transformation. In this repo, size is not a vanity metric. It is part of the problem definition.

A close-up of hands trimming a page of code with a razor blade and calipers, while tiny black shavings fall away. The scene translates the minification step into a tactile act of removing every unnecessary character.
Minification here is not cleanup. It is the core optimization step.

The loop that shrinks solutions

The repo is a loop, not a folder. Each pass through the system checks correctness, trims bytes, and feeds the result back into the next attempt.

This is the repo's real novelty. It does not just hold solutions. It runs a system that tests a candidate, strips it down, measures the result, and pushes the lesson back into the next prompt. The judge verifies correctness. The minifier attacks syntax. The compare step tells you whether the shrinkage was worth it. The prompt generator keeps the search moving.

uv run judge -n 8
uv run compare
uv run prompt task001
uv run submit

Inside the judge, minifier, and submission scripts

The judge is the heartbeat. It runs candidate solutions against ARC tasks, checks their outputs, and gives the repo a hard yes or no. That matters because code golf only works when the evaluator is unforgiving. If the answer is even slightly wrong, the shortest program in the world is still useless.

The minifier is where the repo starts to feel like a specialized compiler toolchain. A script such as minify_solutions.py can strip whitespace, comments, and other syntactic fat from working code, while the shell scripts handle the dull but essential pieces around it, from trimming spaces to packaging submissions. The result is a workflow that treats every byte as a design choice.

The surrounding tooling matters too. uv keeps the Python environment fast and reproducible, GitHub Actions can automate judging, and the Kaggle submission flow connects local experiments to competition infrastructure. This is not a notebook repo. It is an execution environment.

A judge console sits on a desk with an ARC grid on one side, a pass-fail light on the other, and a stack of punch-card-like solutions feeding into a slot from above. The image explains how the repository validates each attempt before it can move forward.
Every candidate passes through a judge before it earns another round of shrinkage.

Why this repo is different from standard ARC solvers

A split workspace shows two approaches side by side. The left side is spacious, with readable variables, comments, and a roomy solver notebook. The right side is cramped, with a one-line golfed program and a hovering byte counter. The image explains the trade-off between clarity and compression.
A normal solver wants clarity. This repo wants the shortest correct solution.
ApproachGoalOutputConstraintSuccess metric
SakanaAI/google-code-golf-2025Solve ARC with minimum bytesPython solutions plus judge and minifier loopCorrectness under byte pressurePassing the judge with shorter code
General ARC solverSolve tasks robustlyRule engine or model predictionsAccuracyTask solve rate
Program-evolution frameworkSearch for better programs over timeIterated candidate code and feedbackOpen-ended improvementPerformance on target tasks
Standard code golf toolingMinimize program lengthGolfed snippets and format tricksByte count and language quirksShortest valid program

That table is the key distinction. A normal solver is rewarded for being right. A golf tool is rewarded for being short. This repo forces both to matter at once, which is why the workflow feels closer to an optimization lab than a puzzle archive.

A framework that combines Large Language Models (LLMs) with evolutionary algorithms to drive scientific discovery.

SakanaAI, Author/Maintainer · SakanaAI/ShinkaEvolve

That line from SakanaAI's broader work helps explain the logic behind this repo. The company is not treating code as a one-shot artifact. It is treating it as something you can search, mutate, validate, and compress in a loop until the shape of the solution changes.

What this says about AI-assisted programming

The important lesson is not that smaller code is automatically better. It is that compactness can be turned into a measurable objective, and once an objective is measurable, you can build tooling around it. That is where the judge, the minifier, and the prompt generator become more interesting than the individual solutions they produce.

This repo sketches a hybrid programming future. Humans find the structure. LLMs accelerate the search. Automated judges keep everyone honest. The end state is not just a correct program. It is a correct program that has been forced through repeated rounds of compression until only the necessary logic remains.