QMD: The Local Memory Engine That Teaches Agents the Missing Context
A hybrid search system for notes and docs that fuses lexical search, vectors, reranking, and hierarchical path metadata into one agent-ready memory layer.
- QMD’s key move is not better search, but better context, because it carries folder meaning into retrieval instead of returning chunks in isolation.
- Its retrieval pipeline mixes Markdown-aware chunking, lexical search, vector search, fusion, and reranking so local notes behave more like a retrieval service than a file index.
- The storage and model layers stay deliberately narrow, using SQLite and local GGUF models to keep the whole system portable and offline.
- QMD is aimed at agents first, which makes its MCP interface and daemon-like posture more important than any standalone note UI.
Why QMD is not just search
The obvious pitch is local search over notes. The better pitch is that QMD tries to stop LLMs from misunderstanding fragments by giving them the missing frame around those fragments. A chunk is useful, but a chunk plus its path-level meaning is much harder to misread.
That is the real trick behind context tree style retrieval. A file is not just text in a database. It is a meeting log for a project, a design note for a team, a code doc in a nested folder, and QMD can inject that hierarchy back into the prompt when it matters.
妈妈我的知识库有救了😭 Shopify 创始人Tobi 和 Karpathy 共同推荐! 完全本地,免费,使用简单,还帮你省一大笔 token! 配置提示词: 帮我装一下 qmd,把当前目录接进来做检索,我是中文笔记。 1.npm install -g @tobilu/qmd —— 装 qmd CLI 本体 https://t.co/GcDhyFS7Kc
The retrieval stack under the hood
The core pipeline is disciplined rather than magical. QMD uses Markdown-aware chunking so it can split on semantic boundaries, not just character counts. That matters because a heading, a paragraph, and a code fence do not deserve the same treatment.
Once the text is chunked, QMD runs a hybrid retrieval loop. Lexical search gives you BM25-style relevance. Vector search gives you semantic similarity. Reciprocal rank fusion blends the two lists, and reranking pushes the most useful candidates toward the top.
The part that keeps reindexing honest
QMD also fingerprints its embedding setup. If the model or chunking strategy changes, it can detect stale vectors and trigger a reindex instead of pretending old embeddings still fit. That is a small detail with a big payoff: fewer silent failures, fewer weird results, less guesswork.
Why the storage layer matters
QMD leans on SQLite as a narrow, portable substrate, then pushes a surprising amount of intelligence into it. The project handles both Node and Bun, loads the vector extension carefully, and even works around platform quirks on macOS so the same search stack can run locally across runtimes.
// Simplified shape of the compatibility idea
const isBun = typeof Bun !== 'undefined';
const db = isBun ? new Bun.Database(path) : new Database(path);
loadSQLiteVecExtension(db);
export function openStore() {
return {
db,
search: hybridSearch,
reindexIfStale: ensureEmbeddingFingerprint,
};
}
That choice matters because it keeps the system small enough to be trusted. The database is not an afterthought. It is where lexical search, vectors, and the indexing metadata all have to coexist without turning the app into a distributed system.
| Layer | What it gives QMD | Why it matters |
|---|---|---|
| Node or Bun + SQLite | A portable runtime and storage base | Lets QMD stay local and shippable without a server |
| FTS5 + sqlite-vec | Lexical and semantic retrieval in one place | Avoids splitting the search stack across separate services |
| Fingerprinting | Reindex safety | Keeps the index aligned with the model and chunking rules |
Local models as plumbing, not a product
QMD’s `llm.ts` treats local models like infrastructure. The same local GGUF-backed layer can embed text, expand queries, and rerank results. That is a different posture from apps that call a remote model for every step and hope latency and privacy are someone else’s problem.
The strategic point is simple. Model work belongs inside the retrieval toolchain, not outside it. Once embeddings, generation, and reranking are all local tasks, the memory layer can behave like a durable system instead of a cloud dependency with a search box attached.
The agent interface is the product
QMD looks built for an agent before it looks built for a human. Its MCP support, `stdio` transport, and HTTP transport make it easy to plug into long-lived environments like Claude Desktop or Cursor. That is why the daemon posture matters more than a pretty UI.
This reframes the project. QMD is not trying to win note-taking. It is trying to become the memory substrate that other tools stand on. In that role, the best feature is not a dashboard. It is reliability under agent load.
What QMD is really competing with
QMD sits between three categories that usually stay separate. Cloud RAG platforms optimize for managed scale. Pure vector search tools optimize for semantic lookup. Desktop note search tools optimize for convenience. QMD tries to take one useful piece from each, then keep the whole stack local.
| Product type | Retrieval strategy | Context handling | Where models run | Best for | Main limitation |
|---|---|---|---|---|---|
| Cloud RAG platforms | Usually hybrid, often managed | Often document-level only | Remote APIs | Teams that want turnkey infrastructure | Less control over privacy, cost, and prompt plumbing |
| Pure vector search tools | Semantic similarity first | Weak or absent hierarchy | Often remote or mixed | Fuzzy lookup across large corpora | Can miss exact terms and structural cues |
| Desktop note search tools | Mostly lexical search | File and folder metadata only | Local | Fast human search over notes | Not designed as agent memory |
| QMD | Hybrid lexical plus vector plus reranking | Path-aware context tree injection | Local GGUF models and local SQLite | Agents that need durable local memory | More moving parts than a simple note app |
That is why QMD feels like a hybrid category rather than a single-feature tool. It is not chasing the fastest search UI. It is trying to be the most useful memory substrate for local agents.
The bet behind the project
The bet is bigger than search quality. QMD implies that the next useful layer for personal AI is a local memory engine with retrieval discipline, model plumbing, and context inheritance already built in. Once that exists, an agent does not just see more data. It sees data with shape.
That shape is the point. When a note can arrive with its parent path, its indexing discipline, and the right ranking signals already attached, an agent stops guessing what the fragment means. It starts working with memory that has a structure worth trusting.