RePo teaches LLMs to rearrange their own context

Sakana AI’s open-source stack turns token order into a learnable layer. The payoff is strongest when the input is messy, the signal is uneven, and rigid sequence order gets in the way.

10 min read • View on GitHub • More from SakanaAI

A chaotic stack of numbered index cards is being pulled into a tighter cluster by a drafting compass and a set of invisible rails. The scene explains RePo’s core idea: the model can reorganize context instead of accepting the input order as fixed.
RePo treats token order as a layout problem, not a law of nature.
Key Takeaways

Most LLMs inherit the same constraint: the first token is still the first token. RePo changes that assumption. Instead of asking attention to recover relevance from a rigid line, it learns a new layout for the line itself.

A model that sorts before it thinks

RePo comes from Sakana AI’s open-source repository, and it is built as a research system rather than a single clever module. The codebase wraps an OLMo-2 backbone, adds Hugging Face compatibility, ships evaluation tooling, and includes training scripts for large cluster runs. That matters, because the project is not just arguing for a new idea. It is showing how to wire that idea into an existing model stack.

Unlike standard approaches, RePo utilizes a differentiable module, $f_\phi$, to assign token positions that capture contextual dependencies, rather than replying on pre-defined integer range.

Sakana AI, Company/Project · Sakana AI RePo blog

The motivating frame is cognitive load. Standard positional encodings assume that physical order and semantic relevance should stay tightly coupled. RePo argues that this can waste the model’s finite working memory when the useful facts are scattered across noise, tables, or long prompts.

How the repositioning layer works

The core move is simple to describe and harder to implement. Instead of using only fixed integer positions, RePo learns continuous positions through a differentiable module, written in the repo as a positional re-mapping path inside the model. Because the positions are differentiable, the model can learn where information should sit relative to other information, not just whether it appeared earlier or later.

dynamic_pe_mode: 'swigluex_multihead'\ndynamic_pe_start_layer: 5\nsoftmax_auxiliary_loss: true\nauxiliary_loss_multiplier: 1e-5

That configuration says a lot. Early layers stay conventional, which gives the model a stable base representation before re-positioning begins. Later layers take over the more aggressive layout work, and the auxiliary loss keeps the system from collapsing into degenerate behavior where everything tries to occupy the same neighborhood.

The architecture is less about adding memory and more about changing the geometry of what the model already has.

That is also why the project’s configs include a small auxiliary loss multiplier instead of a loud new objective. RePo needs enough pressure to learn a useful layout, but not so much that it distorts the rest of training. The design looks like an engineer’s compromise, which is usually a good sign in research code.

What it buys you over fixed positions

QuestionFixed positional encodingsRePo
What does order mean?Order is assigned up front and mostly treated as given.Order becomes a learnable signal that can shift with content.
What happens to noise?Noise still occupies attention budget.Noise can be pushed away from relevant spans.
Where does flexibility come from?Mostly from longer context windows or retrieval.From a differentiable repositioning module inside the model.
Best fitClean sequences and conventional tasks.Noisy, structured, or long-context inputs where relevance is uneven.

RePo outperforms standard encodings on noisy contexts, structured data, and long-range dependencies while maintaining competitive general performance.

Sakana AI, Company/Project · Sakana AI RePo blog

The comparison that matters is not only against RoPE or other fixed encodings. RePo is also different from retrieval-augmented systems. Retrieval adds more information from outside the prompt. RePo tries to make the information already inside the prompt easier to use by changing how it is arranged.

The codebase looks like a real deployment path

The repo is structured like a working research pipeline. There is an OLMo core, a Hugging Face wrapper, an evaluation suite, modified inference libraries, and YAML configs for multiple model sizes and training stages. There are also SLURM scripts, BF16 mixed-precision settings, FSDP support, and custom CUDA and C++ pieces where the performance-sensitive parts live.

That is the difference between a paper demo and a platform. RePo is trying to prove an architectural claim, but it is doing it with the kind of scaffolding that lets other researchers train, evaluate, and adapt the idea without rebuilding everything from scratch.

It represents a step toward models that intelligently curate their own working memory rather than passively accepting input order.

Sakana AI, Company/Project · Sakana AI RePo blog

That final claim is the real thesis. RePo is not trying to make context longer for its own sake. It is trying to make context smarter, which is a much narrower and more interesting bet. If the project lands, the useful unit of progress may not be "more tokens," but better internal organization of the tokens you already have.