Loopwise: The Bash Framework That Turns Claude Code and Codex Into a Self-Checking AI Court
A shell-native review loop where one model writes, another challenges, and the system refuses to trust either one blindly.
- Loopwise’s real innovation is governance, not automation, because it forces one model to generate while another is empowered to resist it.
- The project is surprisingly serious for a Bash repo, using shell-native plumbing, explicit state files, and atomic writes to make the loop inspectable.
- Its most important engineering work is failure handling, including fallback parsing, disposition tracking, and hard stop conditions that prevent endless model argument.
- Loopwise points toward a broader pattern for autonomous coding tools: systems that can challenge themselves without pretending they are certain.
Loopwise is not trying to be another all-purpose coding agent. It is trying to make AI coding less trusting, more inspectable, and harder to fool. The trick is simple to describe and uncommon in practice: one model writes, another model challenges, and the shell keeps score.
That matters because most AI coding tools still optimize for speed first. Loopwise optimizes for restraint. It treats uncertainty as a feature of the system, not a bug to hide behind a smoother interface.
AI Coding Needs an Enemy
The project’s core idea is separation of powers. Claude Code acts as the creator, Codex acts as the critic, and neither gets to decide the outcome alone. That gives Loopwise a sharper personality than a generic assistant loop: it assumes good output must survive skepticism.
It leverages two top-tier AI models (Claude Opus 4.6 + GPT-5.4) to cross-review each other, producing significantly higher quality output than either model alone.
The README frames the project as an automated review loop between Claude Code and Codex CLI. That description undersells the real story. This is less about automation than about building a controlled adversarial system around two non-deterministic tools.
Why Bash Is the Right Wrong Tool
The surprise is the implementation. Loopwise is not a Python agent framework or a cloud orchestration layer. It is Bash, `jq`, and two CLI models stitched together with the blunt discipline of Unix plumbing.
That choice changes how the tool feels. State is explicit. Files are inspectable. Every transition can be traced. Instead of hiding control flow inside a framework, Loopwise exposes it in scripts and local artifacts.
# Simplified shape of the loop
round=0
while [ "$round" -lt "$MAX_ROUNDS" ]; do
run_claude_raw > "$artifact_file"
run_codex_review "$artifact_file" > "$review_file"
parse_review "$review_file"
verify_findings
if approved; then
break
fi
apply_revisions
round=$((round + 1))
done
That is the appeal. Bash is not elegant here. It is legible. The loop behaves like a control system because the control system is laid bare.
The Loop: Generate, Review, Parse, Verify, Revise
The machine works in rounds. Claude produces an artifact. Codex reviews that artifact. A parser tries to turn the review into structured data. A verifier decides whether the findings are actionable. Then the system either revises, approves, or stops safely.
That sequence sounds ordinary until you look at the failure handling. Loopwise does not assume models will return clean JSON. It has a fallback chain, a degraded mode, and a round limit. Those are not accessories. They are the reason the loop can run at all.
| Dimension | Normal flow | Degraded flow |
|---|---|---|
| Parsing | Direct structured parse succeeds | Regex extraction rescues fenced JSON |
| Decision | Verified findings route revisions | Unclear output is marked safely and does not spin forever |
| Stopping | Approved or max rounds reached | Degraded termination prevents endless retries |
| State | Atomic writes keep artifacts inspectable | State remains recoverable even after noisy output |
What the Critic Actually Judges
The critic is not just looking for typos. The review model checks failure surfaces that matter in real software: correctness, maintainability, observability, concurrency, and evolution. That is a better test than generic praise, because it forces the system to care about code that survives contact with reality.
This is where the project’s adversarial design becomes visible. The reviewer is not a rubber stamp. It is a hostile peer with a checklist.
- Code quality and correctness
- Maintainability and evolution
- Observability and debugging
- Concurrency and race conditions
- Safety and rollback behavior
- Boundary conditions and error handling
- Operational fit for real workflows
The point is not that one model always knows better. The point is that different failure modes can be checked by different roles.
How Loopwise Prevents AI From Arguing Forever
A lot of autonomous coding systems fail in a familiar way. They keep trying. Loopwise is more disciplined. It uses `MAX_ROUNDS`, fallback parsing, and disposition tracking so the system can say, in effect, enough.
The disposition model matters here. Findings are not just accepted or rejected. They can be `verified`, `dismissed`, or `unverified_fix`. That preserves nuance without letting ambiguity leak into the next round.
| Mechanism | Why it matters | What it prevents |
|---|---|---|
| Fallback parsing | Handles messy model output | Hard failures from malformed JSON |
| Disposition tracking | Separates verdict from noise | Brittle all-or-nothing review |
| Atomic rename writes | Keeps state files coherent | Partial writes and race conditions |
| Round limits | Sets an upper bound on uncertainty | Infinite review loops |
This is the most mature part of the repo. The brilliance is not in the models. It is in the refusal to let model instability become system instability.
The Gate Mode: Review Before Merge
Loopwise also has a one-shot gate mode. That is the practical version of the same philosophy, designed for staged changes, pre-commit checks, or a quick verdict in an ordinary workflow. It keeps the review discipline without requiring a long background session.
# .claude/commands/loopwise-gate.md
- collect staged or unstaged diffs
- ignore noisy paths like lockfiles and node_modules
- cap input size before context overflow
- ask Codex for a verdict: approve or needs_attention
That makes the project feel less like a demo and more like a reusable habit. The gate mode turns adversarial review into something you can actually put in front of a team.
How It Compares
| Approach | Model separation | Fallback handling | State discipline | Failure posture |
|---|---|---|---|---|
| Loopwise | Explicit creator and critic roles | Yes, with degraded mode | Artifacts and atomic writes | Fails safely and stops |
| Generic autonomous loop | Often one model doing everything | Usually minimal | Often opaque | Keeps retrying |
| Single-model assistant | No real separation | Depends on the UI | Transient chat state | Trusts the model too much |
That comparison is the whole thesis. Loopwise is not trying to be the broadest tool in the category. It is trying to be the most skeptical one.
What Loopwise Is Really Pointing Toward
The bigger lesson is not about Bash or even about Claude and Codex. It is about governance for autonomous software. If AI coding is going to become more agentic, it will need systems that can inspect, challenge, and constrain themselves.
Loopwise is a small but sharp prototype of that future. It suggests that trust in AI tools will not come from asking models to be nicer. It will come from building architectures that expect disagreement, preserve evidence, and know when to stop.