rad-mem: The Memory Layer That Keeps Claude Code from Forgetting Itself
A persistent, token-aware system that captures agent observations, compresses them into durable context, and feeds back only what matters when the next session begins.
- rad-mem is less a transcript store than a memory compression layer that decides what agent history deserves to survive.
- Its core advantage is selective recall under token pressure, not maximal retention.
- The project combines capture, summarization, semantic retrieval, and recency filtering into one memory pipeline.
- That makes it feel like middleware for agent workflows, not just a Claude Code accessory.
The missing layer in agent workflows
Claude Code is strong inside a session and fragile across sessions. rad-mem is built to close that gap by sitting between the agent and its environment, capturing useful observations, compressing them, and feeding back only the parts that still matter. The point is not to preserve everything. The point is to preserve enough.
Rad-Mem seamlessly preserves context across sessions by automatically capturing tool usage observations, generating semantic summaries, and making them available to future sessions. This enables Claude to maintain continuity of knowledge about projects even after sessions end or reconnect.
What rad-mem actually stores
The core data flow starts with an observation. A tool runs, something meaningful happens, and rad-mem captures that result instead of dumping a raw terminal history into storage. Those observations are then summarized, stored in SQLite, and synced into Chroma so they can be searched semantically later.
POST /api/observations
// high-value tool output enters the queue
await observationQueue.enqueue({
projectId,
sessionId,
tool,
output,
timestamp: Date.now()
})
// worker extracts meaning, then persists and syncs it
await summarizeObservation(observation)
await saveToSQLite(summary)
await syncToChroma(summary)
Why vector-first changes the game
Plain logs answer the question, "What happened?" Vector search answers a better one: "What did we decide about this problem, even if we never used the exact same words?" That matters in agent workflows because the same design choice can surface under different prompts, different file names, and different task phrasing. rad-mem uses semantic retrieval to connect those dots, then adds a recency filter so stale decisions do not crowd out current work.
| Approach | What gets stored | How retrieval works | Token aware | Portable |
|---|---|---|---|---|
| Plain logs | Everything, raw and noisy | Exact text lookup | No | Usually yes |
| Naive vector search | Chunks of history | Semantic similarity only | Rarely | Usually yes |
| rad-mem | Observations, summaries, session context | Semantic search plus recency and budget gates | Yes | Yes, by design |
Progressive disclosure is the real product
This is the part that makes rad-mem feel opinionated. It does not try to flood the model with every available memory. It measures token cost, trims by freshness, and reveals context in layers. That is memory with restraint. In practice, that means the system is not optimizing for completeness. It is optimizing for relevance at the exact point where the next generation step needs it.
The repository describes this as progressive disclosure, and that phrase is doing real work. The system surfaces the minimum useful history first, then can reveal more if the task needs it. That is a better fit for agentic work than a giant dump of everything ever recorded.
The RAD Protocol idea
rad-mem is built around a broader claim than Claude Code alone. The structure points toward a platform-agnostic memory bus for agents, where the session, the platform, and the memory store are separate concerns. That matters because the same memory layer could serve multiple tools without rewriting the core idea each time.
| Layer | Claude Code only | Platform-agnostic memory bus |
|---|---|---|
| Scope | One tool chain | Multiple agents and interfaces |
| Memory model | Session-local convenience | Shared infrastructure with durable state |
| Retrieval goal | Help this session | Help any future session with the same work |
| Strategic value | Useful add-on | Potential middleware layer |
The project’s sharp edges
The repo is not pretending to be a generic, risk-free abstraction. It is licensed under AGPL 3.0, which is a deliberate choice about reuse and reciprocity. It also experiments with features like Endless Mode and branch switching, which makes the project feel like a working lab rather than a polished product brochure.
That edge is part of the appeal. rad-mem is built by someone who clearly needs the system to hold up under real agent work, not just demo traffic. The result is a repo that treats memory as infrastructure, with all the messiness that implies.
What it means for the next layer of AI tooling
The bigger takeaway is simple. Agent systems are moving from prompt engineering to context governance. Once work spans sessions, memory becomes part of the stack, not a bonus feature. rad-mem is small and specific, but it points at a durable pattern: store less, retrieve better, and respect the token budget.