The 800K-Parameter Reasoning Engine: Inside chenglou/sotaku
How a four-layer Transformer uses test-time compute scaling and 2D spatial embeddings to solve extreme Sudoku without learning the rules.

Best result: **98.9%** puzzle accuracy at 1024 test-time iterations
- Sotaku proves that test-time iteration scaling dramatically boosts accuracy without expanding parameter count.
- Shared-weight Transformers act as powerful regularizers for algorithmic reasoning tasks by forcing a single network to learn generalized refinement steps.
- Two-dimensional Rotary Positional Embeddings provide spatial intuition to sequence models without explicitly programming the rules of the grid.
The Test-Time Compute Microcosm
The AI industry obsesses over massive parameter counts, but true algorithmic reasoning often requires a different approach. The sotaku repository by Chenglou proves that a tiny 800K-parameter model can solve extreme algorithmic problems by simply being allowed to think longer. It is a miniature, open-source demonstration of the test-time compute scaling philosophy applied to Sudoku.
This mind-bending statistic reveals the core phenomenon of iteration scaling. The model is trained on just 16 loops but tested on 1,024 loops. The iters/eval_more_iters.py script shows that the model's confidence and accuracy monotonically increase as it iterates. It continuously corrects its own mistakes the longer it runs.
Looping the Transformer
Instead of a deep 64-layer network, the flagship iters/exp_baseline_lr2e3.py model implements a four-layer Transformer block with shared weights across iterations. In each cycle, the model takes its previous softmax predictions, projects them back to the model dimension, and adds them to the initial puzzle encoding.
This weight-sharing architecture acts as a strict regularizer. A structural ablation script, arch/exp_unrolled.py, implements an unrolled version with 16 different four-layer transformers. The shared-weight model with 800K parameters performs comparably to or better than the unrolled model with over 12 million parameters.
Injecting Geometric Intuition
Standard Transformers read 1D sequences, but Sudoku is a 2D grid. To bridge this gap, the project utilizes 2D Rotary Positional Embeddings (RoPE). Found in the pos_embedding/ directory, the model splits its 32-dimensional attention head into two 16-dimensional halves.
One half encodes row coordinates, and the other encodes column coordinates. This gives the sequence model an inherent geometric bias. It understands how rows and columns intersect without explicit programmed rules, using a very small base frequency tailored for a 9x9 grid.
The Reverse Curriculum
Standard machine learning often starts easy and gets harder. The script curriculum/train_curriculum_reverse.py details a counter-intuitive training strategy. The model is fed the absolute hardest Sudoku puzzles first.
By forcing the network to grapple with high-entropy features early on, it prevents the model from memorizing lazy patterns found in easy grids. Easy puzzles are introduced gradually only after the model has developed robust algorithmic reasoning skills.
Brute Force vs. Learned Algorithms
Traditional constraint satisfaction solvers like MiniSat rely on exhaustive search and back-propagation of rules. They guarantee a logical proof but cannot learn heuristics. Massive unrolled Transformers offer high accuracy with fast single-pass inference but suffer from extreme parameter bloat.
| Architecture | Parameters | Inference Depth | Accuracy |
|---|---|---|---|
| MiniSat (Constraint Solver) | Zero | Variable (Exhaustive) | 100% |
| Unrolled Transformer | 12M+ | Fixed (Single Pass) | High |
| Sotaku (Iterative) | ~800K | Dynamic (Multi-Pass) | 98.9% |
The Serverless Research Lab
The repository is a masterclass in modern solo research workflows. Heavy integration with Modal for serverless GPU acceleration allows the author to spin up massive H200 instances for short bursts of compute.

From-scratch experiments on iterative neural Sudoku solvers. See post
This workflow highlights how the barrier between hobbyist experimentation and industrial AI research has vanished. By treating model architecture as an engineering problem to be iterated upon rapidly, sotaku provides a blueprint for resource-efficient AI development.