Mole: The Research Agent That Refuses to Guess, Overspend, or Leak Your Data
A Go-based deep-research tool with a real budget ledger, verbatim quote enforcement, and a local-data boundary that keeps raw rows off the prompt.
- Mole turns research limits into hard infrastructure, so overspend, weak sourcing, and data leakage are blocked instead of merely discouraged.
- Its budget system uses a real ledger, which makes the spend cap behave like a gate rather than a warning.
- Its quote discipline forces claims to stay tethered to exact source text, which changes the epistemics of the report.
- Its local-data workflow keeps raw rows off the prompt and returns only aggregates, which makes privacy part of the execution model.
What Mole Actually Protects
Doing research with agents is fun until they blow way past budget, jumble the sources, and don't even give you the best possible answer, just sound confident.
Most research agents are optimized for output. Mole is optimized for refusal. It refuses to overspend, refuses to turn paraphrase into evidence, and refuses to ship raw local data into the prompt just because the model asked nicely.
That is the interesting move. Mole does not try to be a smarter chatbot. It tries to be a more trustworthy research appliance, where cost, sourcing, and privacy are enforced by the system itself.
Why a Ledger Beats a Token Budget
Mole stores spend as a hard ledger in micro-dollars, not as a fuzzy token estimate. The difference matters. A soft budget says "stop soon." A ledger says "this call cannot happen."
// Reserve first, then run.
if err := ledger.Reserve(ctx, costMicros); err != nil {
return err // hard stop before the model call
}
result, err := client.Call(ctx, prompt)
if err != nil {
ledger.Release(costMicros)
return err
}
return ledger.Settle(ctx, actualCostMicros)
That reserve-settle pattern makes the cap operational. It is not a dashboard number. It is a gate that can reject a request before the model sees a single token.
| Budgeting model | What it means | Failure mode |
|---|---|---|
| Soft token budget | Tracks usage and hopes the run stays inside bounds | The agent drifts past the limit before anyone notices |
| Ledger budget | Reserves spend before the call and settles after | The call is blocked if the ceiling would be breached |
| Post-hoc billing | Reports cost after the fact | You learn the problem only after the spend is gone |
The Verbatim Rule Changes Everything
Mole's sourcing discipline is stricter than ordinary retrieval-augmented generation. If the system cannot find the exact supporting quote, the claim is dropped. That is not a nicer prompt. It is a different standard of proof.
Every claim carries a source 3. Data stays local (give a CSV, it'll analyze it without the data ever leaving your machine)
This is the privacy story enterprise teams actually care about. Not "we support local models." Not "we are privacy-friendly." Mole keeps raw rows local by design, then limits the model to a narrow semantic interface.
Why the Daemon and the MCP Shim Are Split
Mole separates the long-lived daemon from the MCP shim. The daemon keeps state. The shim just pumps bytes between stdio and the socket. That split looks small until you think about editor restarts, session continuity, and what happens when a client process dies mid-research.
mole = long-lived research daemon
mole-mcp = thin stdio-to-socket bridge
Why it matters:
- session survives client restarts
- the daemon owns persistence
- integrations stay simple
It is a clean interoperability choice. Mole wants to be an engine other tools can plug into, not another monolith tied to one front end.
How the Planner Prevents Fan-Out Explosion
The planner is built to avoid curiosity spirals. Parameters like MaxDepth, MaxInitialLeads, and ReplanEvery shape how hard the system explores before it pauses to reassess.
That matters because research agents often fail by asking too many expensive questions too early. Mole asks enough to move the answer forward, then replans instead of spraying the web with leads.
What Mole Is Really Competing With
| Project | Strength | Where Mole differs |
|---|---|---|
| GPT Researcher | Broad report generation | Mole is stricter about budget enforcement and exact-source verification |
| STORM | Multi-perspective synthesis | Mole optimizes for terminal use, hard caps, and local-data privacy |
| OpenClaw | General agent harness | Mole is a specialist research engine that can plug into harnesses via MCP |
| Perplexity | Managed deep research service | Mole is local-first, self-hostable, and bound by your own ledger and keys |
The comparison is not about who has the prettiest report. It is about what is made non-negotiable. Mole optimizes for auditable behavior, not just fluent output.
Who This Is For
Mole fits developers, researchers, and teams that care about reproducibility, spend control, and sensitive inputs. It also suits people who want a tool they can inspect, not a service they have to trust on faith.
A test that passes with the fix removed proves nothing, and several of this project's own tests have been caught doing exactly that.
That line tells you a lot about the project. Mole is not built like a demo. It is built like software that expects its own assumptions to be attacked.
Sources: GitHub repository, Show HN announcement, and repository documentation including the README and contributing notes.