reasoning-bank: ReasoningBank: The Agent Memory System That Learns From Its Own Mistakes
A Google Research project that turns past reasoning traces, including failures, into reusable memory for web and coding agents.
- ReasoningBank treats failure as durable training signal, so the system learns from dead ends instead of discarding them.
- Its core unit is a reasoning trajectory, not a document, which makes the memory layer behavioral rather than purely factual.
- The project turns test-time scaling into a feedback loop, where more attempts create better memory and better memory improves the next attempt.
- Its real niche is not generic RAG or orchestration, but resilient reasoning across web and coding tasks under messy conditions.
The missing ingredient in agent memory: failure
Most agent memory systems are built to preserve wins. ReasoningBank flips that instinct. It stores reasoning trajectories from successful and failed runs, then turns them into reusable lessons for the next attempt.
That sounds small until you see the effect. A normal cache gives an agent more examples. ReasoningBank gives it a history of how it solved, stalled, recovered, or failed. For web agents and coding agents, that is often the difference between repeating a mistake and avoiding it.
To address the challenges posed by intermittent network connectivity, we introduce a reasoning-bank mechanism. This mechanism allows the system to store intermediate reasoning steps and resume from where it left off, ensuring that reasoning tasks can be completed even in the face of network disruptions.
What ReasoningBank actually stores
The project does not store raw documents as the main unit of learning. It stores traces: the agent’s internal reasoning and outward actions, typically represented as <think> and <action> outputs in the pipeline. That makes memory closer to a replayable execution history than to a search index.
In code terms, the repo separates writing from reading. `induce_memory.py` distills a run into a memory artifact. `memory_management.py` later ranks those artifacts by similarity and chooses what should come back into context. The memory bank is therefore not a log. It is an opinionated filter on past behavior.
# Simplified view of the loop
trajectory = load_run()
thoughts, actions = extract_think_and_action(trajectory)
if succeeded:
memory = induce_success_memory(thoughts, actions)
else:
memory = induce_failed_system_instruction(thoughts, actions)
save(memory)
query = embed(current_task)
chosen = select_memory(query, memory_bank, backend="gemini_or_qwen")
inject_into_next_attempt(chosen)
How memory gets written after a run
This is where the project stops feeling like a search layer and starts feeling like a learning loop. `induce_memory.py` reads trajectory files, extracts the reasoning trace, and asks the model to compress the run into a compact lesson. If the run succeeded, the output becomes a reusable pattern. If it failed, the output can become a warning, a pitfall, or a failed system instruction.
That distinction matters. Success memory says, effectively, “do this again.” Failure memory says, “do not repeat this move under these conditions.” Agents need both. The first helps them generalize. The second keeps them from looping into the same dead ends.
How the agent chooses what to remember
`memory_management.py` handles the other half of the loop. It embeds the current query, compares it against stored traces, and selects the closest matches with cosine similarity. The result is not a giant memory dump. It is a ranked shortlist of past reasoning that looks relevant to the task at hand.
The interesting part is backend flexibility. The code can swap between Gemini embeddings and Qwen embeddings, which suggests the authors care less about a single model brand and more about the retrieval behavior itself. That is a useful research instinct. If the memory layer works, the model underneath can change.
| Project | What it remembers | How it matches | What it is for |
|---|---|---|---|
| ReasoningBank | Reasoning trajectories and distilled lessons | Cosine similarity over embeddings | Reusable agent memory |
| RAG | Documents and chunks | Vector retrieval over text | Grounding with external knowledge |
| LangChain / LlamaIndex | Application state and tool flows | Developer-defined orchestration | Building LLM apps |
| vLLM | Model execution state | Serving and batching logic | Fast inference |
| Petals | Model layers across peers | Distributed layer routing | Decentralized inference |
That table is the key distinction. ReasoningBank is not trying to serve models faster, shard them across peers, or retrieve background documents. It is trying to remember how an agent got somewhere.
Why scaling matters more than just more parameters
The repo’s more ambitious idea is memory-aware test-time scaling. More parallel attempts create more trajectories. More trajectories create better memory. Better memory improves the next attempt. So scaling is not just about brute force search at inference time. It becomes a way to grow the agent’s experience bank.
This is a subtle but important shift. Traditional scaling asks a model to think harder in the moment. ReasoningBank asks the system to get wiser across attempts. The compute is doing double duty: solving the task now, and improving the memory that will help next time.
Why the project is built around WebArena and SWE-Bench
The repo is anchored in two noisy, realistic task families. WebArena forces an agent to survive browser state, UI churn, and multi-step interactions. SWE-Bench pushes it into software engineering tasks where local fixes can break other parts of the system.
Those environments are good tests because they punish shallow memory. A single fact is rarely enough. What matters is procedural memory: which steps were tried, where the run drifted, and which recovery tactic worked when the first approach failed.
| Benchmark | Why it matters | What memory helps with |
|---|---|---|
| WebArena | Browser tasks are stateful and fragile | Recovering from UI dead ends and partial progress |
| SWE-Bench | Code tasks demand multi-step debugging | Avoiding repeated mistakes and reusing repair patterns |
| Synthetic Q&A | Often too clean to be revealing | Usually less stress on procedural memory |
This is where ReasoningBank feels more concrete than the usual memory talk in agent papers. The system is not abstractly “remembering context.” It is remembering what happened in a real workflow that could go sideways.
How it compares to RAG, vLLM, Petals, and agent frameworks
The easiest way to misread ReasoningBank is to file it under familiar labels. It is not RAG, because it does not primarily retrieve source documents. It is not vLLM, because it is not an inference engine. It is not Petals, because it is not about decentralized model layers. And it is not LangChain, because it is not just orchestration.
| System | Primary job | Main unit | ReasoningBank difference |
|---|---|---|---|
| RAG | Answer from documents | Text chunks | ReasoningBank stores behavior, not just text |
| vLLM | Serve models efficiently | Tokens and batches | ReasoningBank improves agent learning, not throughput |
| Petals | Distribute model execution | Model layers | ReasoningBank distributes reasoning lessons, not layers |
| LangChain / LlamaIndex | Chain tools and data | Application logic | ReasoningBank defines a memory loop, not just a workflow |
That positioning is why the project is interesting. It sits in the gap between infrastructure and cognition. The code cares about network conditions and embeddings, but the intellectual move is about learning from traces.
What this suggests about the next generation of agents
ReasoningBank points toward a less glamorous but more powerful idea: agents may improve less by memorizing more facts and more by accumulating better habits. If the system can store its own reasoning history, then experience becomes a first-class asset.
That changes the unit of progress. The goal is no longer only a better prompt or a larger model. It is a tighter loop between action, failure, compression, retrieval, and retry. In that world, memory is not a sidebar. It is the mechanism that lets an agent learn in public from its own mistakes.