pi-grump: The AI Sidecar That Hunts Slop in Real Time

A grumpy terminal companion for the Pi coding agent, built to score risky changes, roast over-engineering, and keep AI code minimal enough to survive contact with reality.

8 min read · Evizero/pi-grump

A wide terminal courtroom where an AI agent stacks labeled blocks into an overbuilt tower while a grumpy wizard-like sidecar leans in from the margin, arms folded. The scene explains that this tool watches for AI-generated bloat and interrupts only when the code starts to smell wrong.
pi-grump does not act like a cheerful assistant. It behaves like a skeptical co-pilot with permission to object.
Key Takeaways

pi-grump is not trying to be a friend. It is trying to be a useful nuisance. That distinction matters, because the whole project is built around one sharp idea: if AI code starts drifting toward slop, the sidecar should notice early and say so.

The hook is timing. Most review tools wait until code exists. Most coding agents keep helping until the pile gets bigger. pi-grump sits closer to the moment of creation, watches the turn, and decides whether the situation is bad enough to interrupt.

Why Pi needed a grump at all

pi-grump makes sense because Pi itself is already a reaction against bloat. The host agent is minimal by design, and the extension follows that lead instead of fighting it. That gives the project a clean philosophical frame: keep the core small, then let the extension bring attitude and enforcement.

The README says it plainly enough. This is not a polite helper that smooths every edge. It is a character with a job, and the job is to push back when the code starts to look ceremonial, overbuilt, or just plain risky.

Pi is written by Mario Zechner and unlike Peter, who aims for “sci-fi with a touch of madness,” Mario is very grounded.

A personality system with actual teeth

The joke lands because the identity layer is not just flavor text. The grump has stats, rarities, and hardcoded legendary identities, which makes it feel like a game. But those choices also shape the prompt and the tone of the reaction, so personality becomes part of the operating logic.

That is the subtle win here. A lot of developer tools use personality as packaging. pi-grump uses it as a control surface. The persona can soften, sharpen, or redirect the critique without losing the underlying discipline of the system.

type GrumpStats = {
  grump: number
  wit: number
  yagni: number
  patience: number
}

const legendary = {
  pi: { grump: 9, wit: 8, yagni: 10, patience: 2 },
  gramps: { grump: 10, wit: 9, yagni: 9, patience: 1 },
}

How pi-grump decides to speak

The grump is not noisy. It is gated, scored, and fed by history, so it only reacts when the moment justifies the interruption.

A close-up evidence-board scene where separate drawers and folders labeled conversation, code glances, focus guidance, and anonymized text are assembled into a structured packet. The image explains that pi-grump feeds the model with tagged context instead of a loose blob of chat history.
The strongest technical idea in the repo is not the sarcasm. It is the structure around the sarcasm.

The trigger engine is the heart of the project. It looks for risky bash patterns, structural overbuilding, and repetition in reaction history, then turns those signals into a score. Only when that score clears the threshold does the grump wake up and speak.

That selectivity is the difference between character and noise. A bot that comments on everything becomes wallpaper. A bot that waits for a real smell, then fires once with a sharp read, starts to feel like a system with judgment.

What the model actually sees

The prompt pipeline is more disciplined than the joke suggests. Context is assembled with XML tags so the model can tell conversation history from code glances and focus guidance. That hierarchy matters because it prevents the reaction from becoming a soup of equally weighted tokens.

There is also an anonymization step, which is easy to miss and worth respecting. It means the project is not just trying to be clever with prompting. It is also trying to be careful about what gets exposed to the model in the first place.

This is where pi-grump stops looking like a mascot and starts looking like infrastructure. The system is opinionated, but the opinions are packaged with enough structure that the model can answer sharply without guessing at what matters most.

ToolWhen it actsPrimary jobContext stylePhilosophy
pi-grumpDuring generationInterrupt slop earlyTagged, selective, reactiveAnti-slop, minimalist
Octopus ReviewAfter code existsReview a PR for bugs and duplicationIndexed whole codebaseComprehensive review
Claude Code / CursorWhile generating codeHelp produce codeLarge interactive workspace contextHelpful and feature rich
LintersOn save or after the factEnforce rulesLocal file contextMechanical correctness

The comparison is not about which tool is smarter. It is about when the feedback lands. Linters complain after the shape is fixed. Review tools analyze after the code exists. pi-grump tries to catch the drift while the agent is still choosing the path.

Why this is not just another code review tool

That timing makes pi-grump feel more adversarial than a review bot and more restrained than a coding assistant. It is not trying to replace generation, and it is not trying to do a full audit. It is trying to stand beside the agent and object when the moment calls for it.

This is why the project is interesting beyond the gimmick. Personality is often treated as decoration in AI tooling. Here it behaves like policy. The grump does not just entertain the user. It shapes which mistakes become worth speaking about.

Pi, the shitty coding agent, is the same. It doesn’t do a whole lot out of the box... Simple core, massive customizability.

Hans Schnedlitz, Developer and blogger · Pi is Vim for Agentic Coding | Hans Schnedlitz

The deeper bet: personality as a control mechanism

The project’s real thesis is bigger than one extension. It suggests that persona can become an operational layer inside AI tools, not just a UI flourish. Tone, rarity, and trigger rules can all help determine when the system speaks and how hard it pushes back.

That is a useful direction for agentic tools in general. Not every assistant should be eager. Some should be selective. pi-grump shows how a tiny, opinionated sidecar can make restraint feel intentional instead of accidental.