The 800K-Parameter Reasoning Engine: Inside chenglou/sotaku

How a four-layer Transformer uses test-time compute scaling and 2D spatial embeddings to solve extreme Sudoku without learning the rules.

7 min read · chenglou/sotaku

An intricate mechanical hourglass where data tokens fall through a compact four-gear mechanism and assemble into a completed Sudoku grid.
Sotaku achieves 98.9% accuracy on extreme Sudoku puzzles by scaling compute at test time, continuously refining its predictions through a miniature four-layer Transformer.

Best result: **98.9%** puzzle accuracy at 1024 test-time iterations

chenglou, Author · chenglou/sotaku
Key Takeaways

The Test-Time Compute Microcosm

The AI industry obsesses over massive parameter counts, but true algorithmic reasoning often requires a different approach. The sotaku repository by Chenglou proves that a tiny 800K-parameter model can solve extreme algorithmic problems by simply being allowed to think longer. It is a miniature, open-source demonstration of the test-time compute scaling philosophy applied to Sudoku.

Portrait of Chenglou, author of sotaku.

This mind-bending statistic reveals the core phenomenon of iteration scaling. The model is trained on just 16 loops but tested on 1,024 loops. The iters/eval_more_iters.py script shows that the model's confidence and accuracy monotonically increase as it iterates. It continuously corrects its own mistakes the longer it runs.

Looping the Transformer

Instead of a deep 64-layer network, the flagship iters/exp_baseline_lr2e3.py model implements a four-layer Transformer block with shared weights across iterations. In each cycle, the model takes its previous softmax predictions, projects them back to the model dimension, and adds them to the initial puzzle encoding.

The iterative refinement loop feeds predictions back into the same neural layers, scaling accuracy through test-time compute.

This weight-sharing architecture acts as a strict regularizer. A structural ablation script, arch/exp_unrolled.py, implements an unrolled version with 16 different four-layer transformers. The shared-weight model with 800K parameters performs comparably to or better than the unrolled model with over 12 million parameters.

Injecting Geometric Intuition

Standard Transformers read 1D sequences, but Sudoku is a 2D grid. To bridge this gap, the project utilizes 2D Rotary Positional Embeddings (RoPE). Found in the pos_embedding/ directory, the model splits its 32-dimensional attention head into two 16-dimensional halves.

One half encodes row coordinates, and the other encodes column coordinates. This gives the sequence model an inherent geometric bias. It understands how rows and columns intersect without explicit programmed rules, using a very small base frequency tailored for a 9x9 grid.

The Reverse Curriculum

Standard machine learning often starts easy and gets harder. The script curriculum/train_curriculum_reverse.py details a counter-intuitive training strategy. The model is fed the absolute hardest Sudoku puzzles first.

By forcing the network to grapple with high-entropy features early on, it prevents the model from memorizing lazy patterns found in easy grids. Easy puzzles are introduced gradually only after the model has developed robust algorithmic reasoning skills.

Brute Force vs. Learned Algorithms

Traditional constraint satisfaction solvers like MiniSat rely on exhaustive search and back-propagation of rules. They guarantee a logical proof but cannot learn heuristics. Massive unrolled Transformers offer high accuracy with fast single-pass inference but suffer from extreme parameter bloat.

A massive factory floor with 64 identical stamping machines compared to a single polished machine with a looping conveyor belt.
Unrolled architectures require massive parameter bloat, whereas iterative shared-weight models recycle a single robust reasoning block.
ArchitectureParametersInference DepthAccuracy
MiniSat (Constraint Solver)ZeroVariable (Exhaustive)100%
Unrolled Transformer12M+Fixed (Single Pass)High
Sotaku (Iterative)~800KDynamic (Multi-Pass)98.9%

The Serverless Research Lab

The repository is a masterclass in modern solo research workflows. Heavy integration with Modal for serverless GPU acceleration allows the author to spin up massive H200 instances for short bursts of compute.

From-scratch experiments on iterative neural Sudoku solvers. See post

chenglou, Author · chenglou/sotaku

This workflow highlights how the barrier between hobbyist experimentation and industrial AI research has vanished. By treating model architecture as an engineering problem to be iterated upon rapidly, sotaku provides a blueprint for resource-efficient AI development.