jevgrep: The CLI That Teaches Coding Agents Where to Look

A deep dive into the agent-first search layer that prunes repositories into high-signal context, using semantic judgment, AST-aware reading leads, and token-aware output.

9 min read • View on GitHub • More from dzhng

A wide editorial scene shows a coding agent at the edge of a sprawling repository map, with only a few illuminated drawers, folders, and declaration cards pulled into view. The concept is that the tool does not return everything it finds, but instead compresses a codebase into a small set of decision-ready leads.
jevgrep acts less like a search box and more like a context budgeter for agent workflows.

Introducing jevgrep - a research agent CLI powered by jev from @typesafeai that reduces your coding agent cost by 40% (verified on SWE-bench) Make sure to use the built in skill so your coding agent knows to use jg for context collection.

David Zhang, Creator / Aomni Founder · Introducing jevgrep
Key Takeaways

When a coding agent enters a repo it does not know, the expensive part is often not the answer. It is the wandering. jevgrep is built around that problem: it treats context retrieval as a software system of its own, with judgment, budgets, and output pruning baked in from the start.

What If the Search Tool Was Really a Context Budgeter?

The usual search story is simple. You type a string, you get matches. jevgrep is doing something stranger and more useful for agents: it decides which parts of a repository deserve attention before the primary model burns tokens reading the wrong files. That makes it closer to a context compiler than a grep replacement.

Jev is not a chat model. It takes unstructured state plus a typed question and returns a typed decision — a choice, a score, or a boolean, each with a confidence.

TypeSafe AI Documentation, Model Provider · dzhng/jevgrep README (via awesome-jev)

That framing matters because it changes the product’s unit of value. The output is not “all possible matches.” It is “the few pieces of repository evidence most likely to move the agent forward.” In other words, selection beats dumping.


Built for Agents, Not Just for Humans

The clearest sign of the design shift is the jg skill command. The tool does not only serve a human at the terminal. It also teaches Claude Code, Codex, and similar systems when to call back into jevgrep for discovery. The user is now partly the model itself.

That is a subtle but important product move. Instead of expecting an agent to discover a tool by chance, jevgrep ships instructions that make the tool part of the workflow contract. It is search infrastructure plus onboarding for the model.

A hedcut-style portrait of David Zhang derived from his GitHub avatar. The image identifies the creator behind the project and reinforces that the tool is tied to a specific workflow philosophy, not just a generic utility.

How the Retrieval Pipeline Chooses What Deserves Attention

The core engine lives in the retrieval flow. It walks the repository, asks what is worth following, and keeps tightening the scope until the result is compact enough to hand to the agent. The important detail is not that it searches. It is that it ranks, gates, and trims.

The retrieval funnel is a sequence of decisions, not a blunt scan. Each stage narrows the repository into a smaller, more useful context bundle.

This is where the internal vocabulary matters. The pipeline does not just produce files. It produces reading leads and source units. That gives the agent a path into the codebase, not just a pile of file names.

A close-up editorial illustration shows a repository branch splitting through a semantic gate into a declaration box, a call neighborhood, and a source excerpt. The image explains how structure becomes actionable retrieval units instead of a flat list of files.
The best outputs are not raw hits. They are structured leads that point the agent toward the next useful read.

Why Tree-Sitter Matters More Than Keyword Matching

Plain text search can tell you where a string appears. It cannot tell you whether the hit is a declaration, a call site, or the most useful neighborhood around a function. That is why the parser layer matters. jevgrep uses tree-sitter to shape retrieval around code structure, not just characters.

// Conceptual shape of the parser layer
// Not exact source, but the behavior is clear:

const lead = {
  file: "src/auth/session.ts",
  kind: "declaration",
  range: { start: 42, end: 118 },
  evidenceRange: { start: 55, end: 94 },
  neighbors: ["refreshSession", "expireSession"],
  score: 0.91
}

// The agent gets a lead, not a haystack.

That structure gives the model better footing. A function declaration plus its call neighborhood is far more useful than a page of line matches, especially when the question is behavioral: where does session expiry happen, or what code handles retries?

ToolBest forWhat it returnsStrengthWeaknessWhen to use it
rgKnown text or identifiersExact matchesFast and deterministicNeeds the right stringWhen you already know what to search for
jevgrepLikely relevant codeRanked leads, file locations, excerptsFinds useful structure for agentsDepends on semantic judgmentWhen you know the behavior, not the exact symbol
Agent-only explorationDeep reasoning across a codebaseNarrated investigationFlexible and adaptiveExpensive and context-hungryWhen the problem needs broader reasoning after the first pass

That comparison is the real middle layer. rg is still the right answer for exact lookup. Full agent exploration is still the right answer for broad reasoning. jevgrep sits between them, helping the model arrive at the right neighborhood before it starts thinking out loud.

The Hidden Product Is Output Pruning

The final trick is not discovery, but presentation. The renderer sorts by priority, respects a maximum output budget, and truncates intelligently so the agent stays oriented even when the entire result set would be too large. That is why the experience feels curated rather than verbose.

// Conceptual output behavior
// The goal is to preserve orientation under a byte budget.

renderResult({
  maxOutputBytes,
  sortBy: ["priority", "score"],
  keep: ["declaration locations", "reading leads"],
  prune: ["redundant source blocks", "low-value excerpts"]
})

This is a quiet but important shift in UX. The tool is not trying to be exhaustive. It is trying to be sufficient. That makes it feel less like a search result page and more like a briefing.

Where rg Stops and jevgrep Starts

The competitive frame is straightforward. ripgrep excels when the question already has a known string. jevgrep excels when the question is about behavior, ownership, or likely relevance. They are different tools for different parts of the discovery loop.

ToolBest forWhat it returnsStrengthWeaknessWhen to use it
rgKnown stringsLiteral matchesSpeedNo semantic judgmentWhen the query is exact
jevgrepBehavioral discoverySemantic leads and excerptsContext-aware selectionRequires a judge modelWhen the code surface is unknown
Agent-only explorationBroad reasoningNarrated investigationFlexibleContext expensiveWhen search is no longer enough

That is why the project should not be read as a replacement for grep. It is a different layer in the stack. Search finds. jevgrep selects.

The Economics of Asking the Right Model the Right Question

The cost story is the real thesis underneath the architecture. If a primary coding model spends less time reading irrelevant repository material, the whole workflow gets cheaper and faster. That is why the project’s benchmarks matter, even when they are not the whole accounting story.

dzhng/jevgrep checks the card catalog for your coding agent and pulls the right card, while Jev's coin meter charges for every drawer it opens. Tested commit baf2d1c4c719e339f10e7dfb20cf90f9fe402e01. Verdict. Adopt dzhng/jevgrep with care for finding code in repositories your coding agent does not know.

GitTested Reviews, Project Reviewer · Jevgrep Review: Find the Right Code by Asking What It Does

That tradeoff is clean. Spend a smaller model or decision layer to reduce waste in the larger model’s context window. The upside is better focus. The downside is that semantic judgment is never free, so the economics only work if the saved tokens and reduced wandering justify the extra step.