QMD: The Local Memory Engine That Teaches Agents the Missing Context

A hybrid search system for notes and docs that fuses lexical search, vectors, reranking, and hierarchical path metadata into one agent-ready memory layer.

9 to 11 min read • View on GitHub • More from tobi

A filing cabinet built like a living map, with drawers branching into nested note paths and project folders. Threaded cards flow from the drawers into a compact machine with two intake channels, showing how path context and hybrid retrieval merge into one ordered memory output.
QMD is not just searching text. It preserves where a note lives, then uses that context when it ranks what an agent should read next.
Key Takeaways

Why QMD is not just search

The obvious pitch is local search over notes. The better pitch is that QMD tries to stop LLMs from misunderstanding fragments by giving them the missing frame around those fragments. A chunk is useful, but a chunk plus its path-level meaning is much harder to misread.

That is the real trick behind context tree style retrieval. A file is not just text in a database. It is a meeting log for a project, a design note for a team, a code doc in a nested folder, and QMD can inject that hierarchy back into the prompt when it matters.

妈妈我的知识库有救了😭 Shopify 创始人Tobi 和 Karpathy 共同推荐! 完全本地,免费,使用简单,还帮你省一大笔 token! 配置提示词: 帮我装一下 qmd,把当前目录接进来做检索,我是中文笔记。 1.npm install -g @tobilu/qmd —— 装 qmd CLI 本体 https://t.co/GcDhyFS7Kc

诺鸭船长3, X user · @noahduck283 on X
A close-up of a Markdown file being cut into uneven blocks along semantic seams, with one code fence preserved as a sealed unit. Two ranking streams, one lexical and one vector-based, converge into a final stack of results, explaining how QMD combines smart chunking and hybrid retrieval.
QMD’s quality comes from the mechanics before the model sees anything: splitting on Markdown structure, preserving code fences, and merging lexical and semantic signals before reranking.

The retrieval stack under the hood

The core pipeline is disciplined rather than magical. QMD uses Markdown-aware chunking so it can split on semantic boundaries, not just character counts. That matters because a heading, a paragraph, and a code fence do not deserve the same treatment.

Once the text is chunked, QMD runs a hybrid retrieval loop. Lexical search gives you BM25-style relevance. Vector search gives you semantic similarity. Reciprocal rank fusion blends the two lists, and reranking pushes the most useful candidates toward the top.

QMD turns retrieval into a pipeline. The important part is not only ranking, but the extra context that survives all the way into the prompt payload.

The part that keeps reindexing honest

QMD also fingerprints its embedding setup. If the model or chunking strategy changes, it can detect stale vectors and trigger a reindex instead of pretending old embeddings still fit. That is a small detail with a big payoff: fewer silent failures, fewer weird results, less guesswork.

Why the storage layer matters

QMD leans on SQLite as a narrow, portable substrate, then pushes a surprising amount of intelligence into it. The project handles both Node and Bun, loads the vector extension carefully, and even works around platform quirks on macOS so the same search stack can run locally across runtimes.

// Simplified shape of the compatibility idea
const isBun = typeof Bun !== 'undefined';
const db = isBun ? new Bun.Database(path) : new Database(path);

loadSQLiteVecExtension(db);

export function openStore() {
  return {
    db,
    search: hybridSearch,
    reindexIfStale: ensureEmbeddingFingerprint,
  };
}

That choice matters because it keeps the system small enough to be trusted. The database is not an afterthought. It is where lexical search, vectors, and the indexing metadata all have to coexist without turning the app into a distributed system.

LayerWhat it gives QMDWhy it matters
Node or Bun + SQLiteA portable runtime and storage baseLets QMD stay local and shippable without a server
FTS5 + sqlite-vecLexical and semantic retrieval in one placeAvoids splitting the search stack across separate services
FingerprintingReindex safetyKeeps the index aligned with the model and chunking rules

Local models as plumbing, not a product

QMD’s `llm.ts` treats local models like infrastructure. The same local GGUF-backed layer can embed text, expand queries, and rerank results. That is a different posture from apps that call a remote model for every step and hope latency and privacy are someone else’s problem.

The strategic point is simple. Model work belongs inside the retrieval toolchain, not outside it. Once embeddings, generation, and reranking are all local tasks, the memory layer can behave like a durable system instead of a cloud dependency with a search box attached.

The agent interface is the product

QMD looks built for an agent before it looks built for a human. Its MCP support, `stdio` transport, and HTTP transport make it easy to plug into long-lived environments like Claude Desktop or Cursor. That is why the daemon posture matters more than a pretty UI.

This reframes the project. QMD is not trying to win note-taking. It is trying to become the memory substrate that other tools stand on. In that role, the best feature is not a dashboard. It is reliability under agent load.

What QMD is really competing with

QMD sits between three categories that usually stay separate. Cloud RAG platforms optimize for managed scale. Pure vector search tools optimize for semantic lookup. Desktop note search tools optimize for convenience. QMD tries to take one useful piece from each, then keep the whole stack local.

Product typeRetrieval strategyContext handlingWhere models runBest forMain limitation
Cloud RAG platformsUsually hybrid, often managedOften document-level onlyRemote APIsTeams that want turnkey infrastructureLess control over privacy, cost, and prompt plumbing
Pure vector search toolsSemantic similarity firstWeak or absent hierarchyOften remote or mixedFuzzy lookup across large corporaCan miss exact terms and structural cues
Desktop note search toolsMostly lexical searchFile and folder metadata onlyLocalFast human search over notesNot designed as agent memory
QMDHybrid lexical plus vector plus rerankingPath-aware context tree injectionLocal GGUF models and local SQLiteAgents that need durable local memoryMore moving parts than a simple note app

That is why QMD feels like a hybrid category rather than a single-feature tool. It is not chasing the fastest search UI. It is trying to be the most useful memory substrate for local agents.

The bet behind the project

The bet is bigger than search quality. QMD implies that the next useful layer for personal AI is a local memory engine with retrieval discipline, model plumbing, and context inheritance already built in. Once that exists, an agent does not just see more data. It sees data with shape.

That shape is the point. When a note can arrive with its parent path, its indexing discipline, and the right ranking signals already attached, an agent stops guessing what the fragment means. It starts working with memory that has a structure worth trusting.