Loopwise: The Bash Framework That Turns Claude Code and Codex Into a Self-Checking AI Court

A shell-native review loop where one model writes, another challenges, and the system refuses to trust either one blindly.

8 min read • View on GitHub • More from haohappy

A wide courtroom scene built from Unix shell elements. A clerk-like creator model writes code on the left, while a judge-like reviewer model studies the output on the right with a magnifying glass over a diff. A narrow filing chute carries artifacts forward between them, explaining how the system separates generation from review.
Loopwise treats AI coding less like a chat session and more like a court with roles, evidence, and appeal paths.
Key Takeaways

Loopwise is not trying to be another all-purpose coding agent. It is trying to make AI coding less trusting, more inspectable, and harder to fool. The trick is simple to describe and uncommon in practice: one model writes, another model challenges, and the shell keeps score.

That matters because most AI coding tools still optimize for speed first. Loopwise optimizes for restraint. It treats uncertainty as a feature of the system, not a bug to hide behind a smoother interface.

AI Coding Needs an Enemy

The project’s core idea is separation of powers. Claude Code acts as the creator, Codex acts as the critic, and neither gets to decide the outcome alone. That gives Loopwise a sharper personality than a generic assistant loop: it assumes good output must survive skepticism.

It leverages two top-tier AI models (Claude Opus 4.6 + GPT-5.4) to cross-review each other, producing significantly higher quality output than either model alone.

Project README, Repository documentation · haohappy/loopwise README

The README frames the project as an automated review loop between Claude Code and Codex CLI. That description undersells the real story. This is less about automation than about building a controlled adversarial system around two non-deterministic tools.

A close-up control panel of shell switches, valves, and indicator lights. One switch is labeled parse JSON, another regex fallback, another degraded mode, and a red lamp signals the round limit. Ink threads move through the valves toward a locked bin labeled verified, dismissed, unverified_fix, showing how the loop survives messy model output.
The loop stays useful because it expects failure and routes around it instead of pretending the models will always behave.

Why Bash Is the Right Wrong Tool

The surprise is the implementation. Loopwise is not a Python agent framework or a cloud orchestration layer. It is Bash, `jq`, and two CLI models stitched together with the blunt discipline of Unix plumbing.

That choice changes how the tool feels. State is explicit. Files are inspectable. Every transition can be traced. Instead of hiding control flow inside a framework, Loopwise exposes it in scripts and local artifacts.

# Simplified shape of the loop
round=0
while [ "$round" -lt "$MAX_ROUNDS" ]; do
  run_claude_raw > "$artifact_file"
  run_codex_review "$artifact_file" > "$review_file"
  parse_review "$review_file"
  verify_findings
  if approved; then
    break
  fi
  apply_revisions
  round=$((round + 1))
done

That is the appeal. Bash is not elegant here. It is legible. The loop behaves like a control system because the control system is laid bare.

The Loop: Generate, Review, Parse, Verify, Revise

The loop is not just a chat chain. It is a state machine with checkpoints, dispositions, and a safe exit when the model output goes sideways.

The machine works in rounds. Claude produces an artifact. Codex reviews that artifact. A parser tries to turn the review into structured data. A verifier decides whether the findings are actionable. Then the system either revises, approves, or stops safely.

That sequence sounds ordinary until you look at the failure handling. Loopwise does not assume models will return clean JSON. It has a fallback chain, a degraded mode, and a round limit. Those are not accessories. They are the reason the loop can run at all.

DimensionNormal flowDegraded flow
ParsingDirect structured parse succeedsRegex extraction rescues fenced JSON
DecisionVerified findings route revisionsUnclear output is marked safely and does not spin forever
StoppingApproved or max rounds reachedDegraded termination prevents endless retries
StateAtomic writes keep artifacts inspectableState remains recoverable even after noisy output

What the Critic Actually Judges

The critic is not just looking for typos. The review model checks failure surfaces that matter in real software: correctness, maintainability, observability, concurrency, and evolution. That is a better test than generic praise, because it forces the system to care about code that survives contact with reality.

This is where the project’s adversarial design becomes visible. The reviewer is not a rubber stamp. It is a hostile peer with a checklist.

The point is not that one model always knows better. The point is that different failure modes can be checked by different roles.

How Loopwise Prevents AI From Arguing Forever

A lot of autonomous coding systems fail in a familiar way. They keep trying. Loopwise is more disciplined. It uses `MAX_ROUNDS`, fallback parsing, and disposition tracking so the system can say, in effect, enough.

The disposition model matters here. Findings are not just accepted or rejected. They can be `verified`, `dismissed`, or `unverified_fix`. That preserves nuance without letting ambiguity leak into the next round.

MechanismWhy it mattersWhat it prevents
Fallback parsingHandles messy model outputHard failures from malformed JSON
Disposition trackingSeparates verdict from noiseBrittle all-or-nothing review
Atomic rename writesKeeps state files coherentPartial writes and race conditions
Round limitsSets an upper bound on uncertaintyInfinite review loops

This is the most mature part of the repo. The brilliance is not in the models. It is in the refusal to let model instability become system instability.

The Gate Mode: Review Before Merge

Loopwise also has a one-shot gate mode. That is the practical version of the same philosophy, designed for staged changes, pre-commit checks, or a quick verdict in an ordinary workflow. It keeps the review discipline without requiring a long background session.

# .claude/commands/loopwise-gate.md
- collect staged or unstaged diffs
- ignore noisy paths like lockfiles and node_modules
- cap input size before context overflow
- ask Codex for a verdict: approve or needs_attention

That makes the project feel less like a demo and more like a reusable habit. The gate mode turns adversarial review into something you can actually put in front of a team.

How It Compares

ApproachModel separationFallback handlingState disciplineFailure posture
LoopwiseExplicit creator and critic rolesYes, with degraded modeArtifacts and atomic writesFails safely and stops
Generic autonomous loopOften one model doing everythingUsually minimalOften opaqueKeeps retrying
Single-model assistantNo real separationDepends on the UITransient chat stateTrusts the model too much

That comparison is the whole thesis. Loopwise is not trying to be the broadest tool in the category. It is trying to be the most skeptical one.

What Loopwise Is Really Pointing Toward

The bigger lesson is not about Bash or even about Claude and Codex. It is about governance for autonomous software. If AI coding is going to become more agentic, it will need systems that can inspect, challenge, and constrain themselves.

Loopwise is a small but sharp prototype of that future. It suggests that trust in AI tools will not come from asking models to be nicer. It will come from building architectures that expect disagreement, preserve evidence, and know when to stop.