RePo teaches LLMs to rearrange their own context
Sakana AI’s open-source stack turns token order into a learnable layer. The payoff is strongest when the input is messy, the signal is uneven, and rigid sequence order gets in the way.
- RePo treats token order as a learnable layout problem, so the model can pull relevant spans together instead of living with a fixed sequence forever.
- The repository is a full research stack, not just a paper code dump, with OLMo integration, Hugging Face wrappers, evaluation tooling, and cluster training scripts.
- Its best case is not every language task, but noisy, structured, or long-context inputs where rigid positional encodings waste attention budget.
- The bigger idea is that working memory can be curated inside the model, not only enlarged from the outside.
Most LLMs inherit the same constraint: the first token is still the first token. RePo changes that assumption. Instead of asking attention to recover relevance from a rigid line, it learns a new layout for the line itself.
A model that sorts before it thinks
RePo comes from Sakana AI’s open-source repository, and it is built as a research system rather than a single clever module. The codebase wraps an OLMo-2 backbone, adds Hugging Face compatibility, ships evaluation tooling, and includes training scripts for large cluster runs. That matters, because the project is not just arguing for a new idea. It is showing how to wire that idea into an existing model stack.
Unlike standard approaches, RePo utilizes a differentiable module, $f_\phi$, to assign token positions that capture contextual dependencies, rather than replying on pre-defined integer range.
The motivating frame is cognitive load. Standard positional encodings assume that physical order and semantic relevance should stay tightly coupled. RePo argues that this can waste the model’s finite working memory when the useful facts are scattered across noise, tables, or long prompts.
How the repositioning layer works
The core move is simple to describe and harder to implement. Instead of using only fixed integer positions, RePo learns continuous positions through a differentiable module, written in the repo as a positional re-mapping path inside the model. Because the positions are differentiable, the model can learn where information should sit relative to other information, not just whether it appeared earlier or later.
dynamic_pe_mode: 'swigluex_multihead'\ndynamic_pe_start_layer: 5\nsoftmax_auxiliary_loss: true\nauxiliary_loss_multiplier: 1e-5
That configuration says a lot. Early layers stay conventional, which gives the model a stable base representation before re-positioning begins. Later layers take over the more aggressive layout work, and the auxiliary loss keeps the system from collapsing into degenerate behavior where everything tries to occupy the same neighborhood.
That is also why the project’s configs include a small auxiliary loss multiplier instead of a loud new objective. RePo needs enough pressure to learn a useful layout, but not so much that it distorts the rest of training. The design looks like an engineer’s compromise, which is usually a good sign in research code.
What it buys you over fixed positions
| Question | Fixed positional encodings | RePo |
|---|---|---|
| What does order mean? | Order is assigned up front and mostly treated as given. | Order becomes a learnable signal that can shift with content. |
| What happens to noise? | Noise still occupies attention budget. | Noise can be pushed away from relevant spans. |
| Where does flexibility come from? | Mostly from longer context windows or retrieval. | From a differentiable repositioning module inside the model. |
| Best fit | Clean sequences and conventional tasks. | Noisy, structured, or long-context inputs where relevance is uneven. |
RePo outperforms standard encodings on noisy contexts, structured data, and long-range dependencies while maintaining competitive general performance.
The comparison that matters is not only against RoPE or other fixed encodings. RePo is also different from retrieval-augmented systems. Retrieval adds more information from outside the prompt. RePo tries to make the information already inside the prompt easier to use by changing how it is arranged.
The codebase looks like a real deployment path
The repo is structured like a working research pipeline. There is an OLMo core, a Hugging Face wrapper, an evaluation suite, modified inference libraries, and YAML configs for multiple model sizes and training stages. There are also SLURM scripts, BF16 mixed-precision settings, FSDP support, and custom CUDA and C++ pieces where the performance-sensitive parts live.
That is the difference between a paper demo and a platform. RePo is trying to prove an architectural claim, but it is doing it with the kind of scaffolding that lets other researchers train, evaluate, and adapt the idea without rebuilding everything from scratch.
It represents a step toward models that intelligently curate their own working memory rather than passively accepting input order.
That final claim is the real thesis. RePo is not trying to make context longer for its own sake. It is trying to make context smarter, which is a much narrower and more interesting bet. If the project lands, the useful unit of progress may not be "more tokens," but better internal organization of the tokens you already have.