rad-mem: The Memory Layer That Keeps Claude Code from Forgetting Itself

A persistent, token-aware system that captures agent observations, compresses them into durable context, and feeds back only what matters when the next session begins.

8 min read • View on GitHub • More from thedotmack

A terminal-shaped machine on a desk pulls loose scraps of command output, file names, and session notes into a central archive, then routes them into a new session window as ordered memory cards. The scene explains how rad-mem turns temporary agent activity into durable context for later reuse.
rad-mem treats agent output as something to compress, rank, and reuse, not just log and forget.
Key Takeaways

The missing layer in agent workflows

Claude Code is strong inside a session and fragile across sessions. rad-mem is built to close that gap by sitting between the agent and its environment, capturing useful observations, compressing them, and feeding back only the parts that still matter. The point is not to preserve everything. The point is to preserve enough.

Rad-Mem seamlessly preserves context across sessions by automatically capturing tool usage observations, generating semantic summaries, and making them available to future sessions. This enables Claude to maintain continuity of knowledge about projects even after sessions end or reconnect.

thedotmack, Author/Maintainer · Repository: thedotmack/rad-mem

What rad-mem actually stores

The core data flow starts with an observation. A tool runs, something meaningful happens, and rad-mem captures that result instead of dumping a raw terminal history into storage. Those observations are then summarized, stored in SQLite, and synced into Chroma so they can be searched semantically later.

The pipeline turns raw tool activity into a compact context packet for the next session.

POST /api/observations

// high-value tool output enters the queue
await observationQueue.enqueue({
  projectId,
  sessionId,
  tool,
  output,
  timestamp: Date.now()
})

// worker extracts meaning, then persists and syncs it
await summarizeObservation(observation)
await saveToSQLite(summary)
await syncToChroma(summary)

Why vector-first changes the game

Plain logs answer the question, "What happened?" Vector search answers a better one: "What did we decide about this problem, even if we never used the exact same words?" That matters in agent workflows because the same design choice can surface under different prompts, different file names, and different task phrasing. rad-mem uses semantic retrieval to connect those dots, then adds a recency filter so stale decisions do not crowd out current work.

A close-up of a memory gate with three stacked filters. Old observations pile up on one side, while only a small set of selected memory cards pass through the recency, relevance, and token budget layers into a compact context window. The image explains why rad-mem is selective rather than exhaustive.
More memory is not the goal. Better memory at the right moment is.
ApproachWhat gets storedHow retrieval worksToken awarePortable
Plain logsEverything, raw and noisyExact text lookupNoUsually yes
Naive vector searchChunks of historySemantic similarity onlyRarelyUsually yes
rad-memObservations, summaries, session contextSemantic search plus recency and budget gatesYesYes, by design

Progressive disclosure is the real product

This is the part that makes rad-mem feel opinionated. It does not try to flood the model with every available memory. It measures token cost, trims by freshness, and reveals context in layers. That is memory with restraint. In practice, that means the system is not optimizing for completeness. It is optimizing for relevance at the exact point where the next generation step needs it.

The repository describes this as progressive disclosure, and that phrase is doing real work. The system surfaces the minimum useful history first, then can reveal more if the task needs it. That is a better fit for agentic work than a giant dump of everything ever recorded.

The RAD Protocol idea

rad-mem is built around a broader claim than Claude Code alone. The structure points toward a platform-agnostic memory bus for agents, where the session, the platform, and the memory store are separate concerns. That matters because the same memory layer could serve multiple tools without rewriting the core idea each time.

LayerClaude Code onlyPlatform-agnostic memory bus
ScopeOne tool chainMultiple agents and interfaces
Memory modelSession-local convenienceShared infrastructure with durable state
Retrieval goalHelp this sessionHelp any future session with the same work
Strategic valueUseful add-onPotential middleware layer

The project’s sharp edges

The repo is not pretending to be a generic, risk-free abstraction. It is licensed under AGPL 3.0, which is a deliberate choice about reuse and reciprocity. It also experiments with features like Endless Mode and branch switching, which makes the project feel like a working lab rather than a polished product brochure.

That edge is part of the appeal. rad-mem is built by someone who clearly needs the system to hold up under real agent work, not just demo traffic. The result is a repo that treats memory as infrastructure, with all the messiness that implies.

What it means for the next layer of AI tooling

The bigger takeaway is simple. Agent systems are moving from prompt engineering to context governance. Once work spans sessions, memory becomes part of the stack, not a bonus feature. rad-mem is small and specific, but it points at a durable pattern: store less, retrieve better, and respect the token budget.