Redline: the Claude Code plugin that makes your agent review itself

A Stop hook, a background reviewer, and Markdown commands turn code review into a live control loop inside the terminal.

9 min read • View on GitHub • More from alexanderatallah

A developer sits at a terminal desk and is stopped by a hard red barrier while a second figure reviews a printed diff under a lamp. The scene explains Redline's core idea: review is not a later stage, it is an interruption inside the same working session.
Redline turns a pause into a review trigger, so the agent can be interrupted while the context is still fresh.
Key Takeaways

The agent that stops itself

Most code review tools wait for a commit or a pull request. Redline does the opposite. It catches the moment a Claude Code session pauses, checks whether the agent just touched files, and can launch review before the work escapes the workspace.

That shift matters because it keeps the review in the same context that produced the change. Redline is not trying to be a smarter bot on a separate server. It is trying to become part of the agent's own control loop.

Why this only works inside Claude Code

Redline depends on Claude Code because the host can stop, explain why, and keep the reason inside the same conversation. That is the trick. A detached review service can comment on code, but it cannot hand that judgment back into the exact session that needs it.

WSJ-style hedcut portrait of Alexander Atallah. It shows a head-and-shoulders likeness derived from his GitHub avatar, reinforcing that Redline is a tightly scoped plugin built by a single maintainer.

Alexander Atallah built Redline as a plugin, which keeps the system small and legible. No framework, no heavy build layer, just hooks, Markdown commands, and a few scripts that wire the pieces together.

Markdown is the control plane

Redline's most interesting move is not the model call. It is that the behavior is encoded in plain text files: `review.md`, `adversarial.md`, `rescue.md`, and `check.md`. The repo reads less like an app and more like a small policy stack.

A close-up of labeled drawers and index cards for review.md, adversarial.md, rescue.md, and check.md. The image explains that Redline's behavior is governed by readable policy files, not buried in a framework.
Redline keeps policy in Markdown, so the review rules stay visible, editable, and easy to inspect.

The payoff is editability. A developer can tune what counts as a normal review, when to escalate into adversarial mode, and when to hand the task to a rescue path without hunting through a framework or recompiling anything.

How Redline routes a task to Codex

This diagram shows the whole trick: a pause becomes a routing event, and the review result comes back into the same conversation.

At the center is the Stop hook. When Claude pauses, Redline checks for uncommitted changes, decides whether the session deserves a review, and if so, tells the agent to spawn the background path.

The review itself runs through Codex, with OpenRouter supplying model access and the result flowing back into Claude's context. That last step is the whole trick. The reviewer is not a detached service. It is feedback the same session can actually use.

If the change is trivial, the plugin can avoid wasting attention on a full review cycle. If the diff looks risky, Redline does not wait for a later PR to discover it. It interrupts early, while the agent still has the full working set in memory.

What the adversarial mode changes

Adversarial mode changes the mood of the review. Instead of asking only whether the code works, `adversarial.md` pushes the reviewer to look for hidden assumptions, rollback risk, and failure modes that usually hide behind a passing test.

A split scene shows a slow pull request queue on one side and a tight terminal loop on the other. The contrast explains that Redline shifts review from a delayed handoff into an immediate interruption inside the coding session.
Redline changes timing, not just tooling. Review happens while the context is still hot.

That is a useful distinction. Plenty of tools can spot syntax errors or obvious bugs. Fewer are opinionated enough to ask what breaks when the change is deployed, reversed, or stressed.

Redline vs the usual review tools

ApproachWhen it runsWhat it is good at
PR botsAfter code is pushed or a pull request is openedCatch issues once context has already moved into review
Generic AI agentsDuring local or remote codingWrite or refactor code, but often lack a tight review gate
RedlineThe moment Claude Code stops with changes in flightInterrupts the agent and sends the diff back into the same conversation
Manual reviewWhenever a human has timeDeep judgment, but slow and expensive in attention

The difference is not that Redline is universally better. It solves a different timing problem. It makes review happen while the agent still has the full working set in memory, which is exactly when feedback is cheapest to absorb.

What this says about agentic coding

Redline hints at where useful agent tooling may go next. The best systems may not be louder copilots. They may be small governance layers that know when to interrupt, when to escalate, and when to let the worker continue.

The limitation is real, though. Redline is tightly coupled to Claude Code because it depends on that specific Stop hook behavior. That coupling is a feature, not a bug. The tool is less portable, but far more exact about the workflow it is trying to improve.