Sakana AI
On a quest to create a new kind of foundation model based on nature-inspired intelligence.
GitHub
63 repos
4.0k followers
Explained projects
SakanaAI/google-code-golf-2025: The Repo That Turns ARC Reasoning Into Byte-Count Combat
A deep dive into the judge, minifier, and prompt loop that squeeze correct solutions down to the smallest possible Python.
8 min read · Apr 7, 2026
RePo teaches LLMs to rearrange their own context
Sakana AI’s open-source stack turns token order into a learnable layer. The payoff is strongest when the input is messy, the signal is uneven, and rigid sequence order gets in the way.
10 min read · Apr 7, 2026
SakanaAI/sparser-faster-llms: How TwELL Makes Sparse Transformers Run Like They Were Built for Hopper
A deep dive into the repo that couples activation sparsity with H100-tuned CUDA kernels, so the model’s zeros become real speed, lower memory use, and lower energy.
8 min read · Apr 7, 2026
SakanaAI/neuroevolution-for-ai: The README That Tries to Organize a Research Field
A curated Markdown map of labs, libraries, benchmarks, and learning resources shows how neuroevolution becomes usable when someone turns a field into a commons.
8 min read · Apr 7, 2026
TransEvalnia Makes Translation Scores Explain Themselves
Sakana AI's evaluation framework asks an LLM to critique, compare, and rank translations before it ever emits a final verdict.
9 min read · Apr 7, 2026
DiffusionBlocks: When Transformer Depth Becomes a Noise Schedule
This research repo treats training as denoising, splits ViT into independently trainable blocks, and uses diffusion math to rethink the memory wall.
8 min read · Apr 7, 2026
Doc-to-LoRA Turns Documents Into Temporary Model Weights
Sakana AI’s D2L skips the long prompt and generates LoRA adapters from the document itself.
10 min read · Apr 7, 2026
ike: Sakana AI’s Modular Shell Around DeepSpeed
A closer look at the framework that turns distributed LLM training into a plug-in architecture for data, loss functions, configs, and checkpoints.
8 min read · Apr 7, 2026
drq: When the benchmark fights back
Digital Red Queen uses an LLM to evolve Core War warriors in an adversarial loop, turning a retro programming game into a lab for open-ended adaptation.
10 min read · Apr 7, 2026
rl-razor-mnist: How a Tiny Task Flag Reveals Why RL Forgets Less
A minimal replication of Sakana AI’s RL’s Razor experiment, where a single extra input bit, a KL-minimal oracle, and group-relative RL turn catastrophic forgetting into something you can inspect, compare, and predict.
8 min read · Apr 6, 2026
petri-dish-nca: When Neural Cellular Automata Start Competing for Space
Sakana AI's PD-NCA turns growth into an ecosystem, where multiple learners fight, adapt, and survive inside one shared dish.
11 min read · Apr 6, 2026
DroPE: The LLM Context Trick That Works Better After You Drop the Map
Sakana AI’s repo extends pretrained models by removing positional embeddings, then recalibrating just enough to unlock longer context windows without brute-force long-context fine-tuning.
12 min read · Apr 6, 2026
Shachi: A Laboratory for LLM Societies
Sakana AI’s framework turns brittle agent-based modeling into repeatable experiments, with tools, memory, and a second parsing pass when models return messy output.
11 min read · Apr 6, 2026
fast-weight-product-key-memory: FwPKM: The Memory Layer That Writes Back
Sakana AI’s open-source experiment fuses product-key lookup with fast-weight updates, so an LLM can retrieve from a huge memory and revise what it remembers mid-stream.
10 min read · Apr 6, 2026
robust-kbench: The benchmark that refuses to be gamed
Sakana AI’s robust-kbench turns CUDA evaluation into a robustness test, not a leaderboard trick.
11 min read · Apr 6, 2026
IASC: The Compiler for Constructed Languages
Sakana AI's open-source pipeline turns phonology, grammar, lexicon, and orthography into a reproducible language build, not a one-shot prompt.
8 min read · Apr 6, 2026
ShinkaEvolve: The Survival of the Fittest Code
Sakana AI's framework for autonomous discovery uses "Islands of Intelligence" and a ruthless Novelty Judge to evolve programs that outperform human engineers.
· Mar 28, 2026
Yet to be explained
fugu
Shell
1.0k stars
Explain
pc-alm
PC-ALM
Python
238 stars
Explain
kame
Python
125 stars
Explain
digital-ecosystem
Interactive multi-agent NCA ecosystem simulation
JavaScript
91 stars
Explain
DreamCubed
Jupyter Notebook
77 stars
Explain
sheaf-admm
Sheaf-ADMM
Python
76 stars
Explain
CoffeeBench
[NeurIPS 2026] CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
Python
38 stars
Explain
kame_finetune
Python
35 stars
Explain
SearchCast
Ridge regression, with its preprocessing tuned, matches the deep models.
Python
9 stars
Explain
quanta
Repository for the paper "Neural Dynamics as the Composition of Quantized Units"
Python
2 stars
Explain
mimi2kana_sample
sample code for https://arxiv.org/abs/2509.20655
Python
0 stars
Explain
NumberNames
Number name ConLang grammars for testing frontier models
Python
0 stars
Explain
KamonBench
KamonBench: A Grammar-Based Dataset for Evaluating Compositional Factor Recovery in Vision-Language Models
Python
0 stars
Explain