reasoning-bank: ReasoningBank: The Agent Memory System That Learns From Its Own Mistakes

A Google Research project that turns past reasoning traces, including failures, into reusable memory for web and coding agents.

9 min read • View on GitHub • More from google-research

A filing cabinet stands open with drawers labeled as successful runs, failed runs, detours, and recovered paths. A hand pulls a failed trajectory card from one drawer and feeds it toward a browser window and code editor, showing how mistakes become reusable guidance.
ReasoningBank treats a failed run as raw material, not waste. The surprising move is not storing answers, but storing the path that got there.
Key Takeaways

The missing ingredient in agent memory: failure

Most agent memory systems are built to preserve wins. ReasoningBank flips that instinct. It stores reasoning trajectories from successful and failed runs, then turns them into reusable lessons for the next attempt.

That sounds small until you see the effect. A normal cache gives an agent more examples. ReasoningBank gives it a history of how it solved, stalled, recovered, or failed. For web agents and coding agents, that is often the difference between repeating a mistake and avoiding it.

To address the challenges posed by intermittent network connectivity, we introduce a reasoning-bank mechanism. This mechanism allows the system to store intermediate reasoning steps and resume from where it left off, ensuring that reasoning tasks can be completed even in the face of network disruptions.

What ReasoningBank actually stores

The project does not store raw documents as the main unit of learning. It stores traces: the agent’s internal reasoning and outward actions, typically represented as <think> and <action> outputs in the pipeline. That makes memory closer to a replayable execution history than to a search index.

The loop is simple but unusual. A run becomes a lesson, and a lesson changes the next run.

In code terms, the repo separates writing from reading. `induce_memory.py` distills a run into a memory artifact. `memory_management.py` later ranks those artifacts by similarity and chooses what should come back into context. The memory bank is therefore not a log. It is an opinionated filter on past behavior.

# Simplified view of the loop
trajectory = load_run()
thoughts, actions = extract_think_and_action(trajectory)
if succeeded:
    memory = induce_success_memory(thoughts, actions)
else:
    memory = induce_failed_system_instruction(thoughts, actions)
save(memory)

query = embed(current_task)
chosen = select_memory(query, memory_bank, backend="gemini_or_qwen")
inject_into_next_attempt(chosen)

How memory gets written after a run

This is where the project stops feeling like a search layer and starts feeling like a learning loop. `induce_memory.py` reads trajectory files, extracts the reasoning trace, and asks the model to compress the run into a compact lesson. If the run succeeded, the output becomes a reusable pattern. If it failed, the output can become a warning, a pitfall, or a failed system instruction.

That distinction matters. Success memory says, effectively, “do this again.” Failure memory says, “do not repeat this move under these conditions.” Agents need both. The first helps them generalize. The second keeps them from looping into the same dead ends.

A workbench shows three parallel agent attempts on the left as tangled strips of thought, action, and response. On the right, those strips are compressed into a single lesson card that slots into a memory shelf, ready for reuse on a fresh task.
The point of induction is compression. Long, noisy trajectories become short memory artifacts that are cheap to retrieve and easy to reuse.

How the agent chooses what to remember

`memory_management.py` handles the other half of the loop. It embeds the current query, compares it against stored traces, and selects the closest matches with cosine similarity. The result is not a giant memory dump. It is a ranked shortlist of past reasoning that looks relevant to the task at hand.

The interesting part is backend flexibility. The code can swap between Gemini embeddings and Qwen embeddings, which suggests the authors care less about a single model brand and more about the retrieval behavior itself. That is a useful research instinct. If the memory layer works, the model underneath can change.

ProjectWhat it remembersHow it matchesWhat it is for
ReasoningBankReasoning trajectories and distilled lessonsCosine similarity over embeddingsReusable agent memory
RAGDocuments and chunksVector retrieval over textGrounding with external knowledge
LangChain / LlamaIndexApplication state and tool flowsDeveloper-defined orchestrationBuilding LLM apps
vLLMModel execution stateServing and batching logicFast inference
PetalsModel layers across peersDistributed layer routingDecentralized inference

That table is the key distinction. ReasoningBank is not trying to serve models faster, shard them across peers, or retrieve background documents. It is trying to remember how an agent got somewhere.

Why scaling matters more than just more parameters

The repo’s more ambitious idea is memory-aware test-time scaling. More parallel attempts create more trajectories. More trajectories create better memory. Better memory improves the next attempt. So scaling is not just about brute force search at inference time. It becomes a way to grow the agent’s experience bank.

This is a subtle but important shift. Traditional scaling asks a model to think harder in the moment. ReasoningBank asks the system to get wiser across attempts. The compute is doing double duty: solving the task now, and improving the memory that will help next time.


Why the project is built around WebArena and SWE-Bench

The repo is anchored in two noisy, realistic task families. WebArena forces an agent to survive browser state, UI churn, and multi-step interactions. SWE-Bench pushes it into software engineering tasks where local fixes can break other parts of the system.

Those environments are good tests because they punish shallow memory. A single fact is rarely enough. What matters is procedural memory: which steps were tried, where the run drifted, and which recovery tactic worked when the first approach failed.

BenchmarkWhy it mattersWhat memory helps with
WebArenaBrowser tasks are stateful and fragileRecovering from UI dead ends and partial progress
SWE-BenchCode tasks demand multi-step debuggingAvoiding repeated mistakes and reusing repair patterns
Synthetic Q&AOften too clean to be revealingUsually less stress on procedural memory

This is where ReasoningBank feels more concrete than the usual memory talk in agent papers. The system is not abstractly “remembering context.” It is remembering what happened in a real workflow that could go sideways.

How it compares to RAG, vLLM, Petals, and agent frameworks

The easiest way to misread ReasoningBank is to file it under familiar labels. It is not RAG, because it does not primarily retrieve source documents. It is not vLLM, because it is not an inference engine. It is not Petals, because it is not about decentralized model layers. And it is not LangChain, because it is not just orchestration.

SystemPrimary jobMain unitReasoningBank difference
RAGAnswer from documentsText chunksReasoningBank stores behavior, not just text
vLLMServe models efficientlyTokens and batchesReasoningBank improves agent learning, not throughput
PetalsDistribute model executionModel layersReasoningBank distributes reasoning lessons, not layers
LangChain / LlamaIndexChain tools and dataApplication logicReasoningBank defines a memory loop, not just a workflow

That positioning is why the project is interesting. It sits in the gap between infrastructure and cognition. The code cares about network conditions and embeddings, but the intellectual move is about learning from traces.

What this suggests about the next generation of agents

ReasoningBank points toward a less glamorous but more powerful idea: agents may improve less by memorizing more facts and more by accumulating better habits. If the system can store its own reasoning history, then experience becomes a first-class asset.

That changes the unit of progress. The goal is no longer only a better prompt or a larger model. It is a tighter loop between action, failure, compression, retrieval, and retry. In that world, memory is not a sidebar. It is the mechanism that lets an agent learn in public from its own mistakes.