`claw-code-parity`: Rewriting an AI agent without changing how it behaves
A Rust parity port that treats autonomy as a reproducible runtime problem, not a prompt trick.
- The repo treats the agent harness as the artifact worth preserving, not the model call itself.
- Recovery, compaction, and sandboxing are core product features because they keep the loop stable under load.
- The Green Contract turns autonomy into a gated system with explicit proof of readiness.
- Rust matters here because it tightens the runtime without changing the user-facing behavior.
The real product is the loop
Most agent repos try to get smarter. claw-code-parity tries to stay the same. The surprise is that this makes the project more ambitious, not less, because it treats an AI assistant like a runtime with behavioral guarantees.
The question is not whether the model can answer. It is whether the same loop, the same tools, the same recovery path, and the same stopping rules survive a rewrite. Once a system runs for hours instead of seconds, those details stop being plumbing and start being the product.
Why this fork exists
The README frames this branch as temporary parity work while the broader codebase migrates. That matters because the job is not to reinterpret the assistant in Rust. The job is to preserve the visible behavior of an already-tuned harness while replacing its execution engine.
The fastest repo in history to surpass 50K stars ⭐, reaching the milestone in just 2 hours after publication. Better Harness Tools, not merely storing the archive of leaked Claude Code but also make real things done. Now rewriting in Rust.
In other words, the fork is a compatibility layer for autonomy. If the agent used to recover from a tool failure in a certain way, the Rust port has to recover the same way. If a session used to compact at a certain point, that behavior becomes a contract.
At 4 AM on March 31, 2026, I woke up to my phone blowing up with notifications. The Claude Code source had been exposed, and the entire dev community was in a frenzy. My girlfriend in Korea was genuinely worried I might face legal action from Anthropic just for having the code on my machine — so I did what any engineer would do under pressure: I sat down, ported the core features to Python from scratch, and pushed it before the sun came up.
How the runtime keeps its promises
The codebase is split into crates so the responsibilities stay legible. The api layer handles provider differences and streaming, the runtime layer owns the conversation loop and persistence, the CLI sits at the edge, plugins extend the machine, and telemetry records what happened. That separation is what makes parity testable.
The mechanics are practical rather than mystical. Responses arrive over streaming transport, sessions persist across turns, long runs compact their history before the context budget breaks, and tool calls run through sandbox rules instead of raw shell access. The point is not to remove risk. The point is to make risk explicit.
Green Contract, red line
The Green Contract is the repo's trust model. Instead of asking whether the agent feels ready, it asks which level of readiness has been proven. A narrow fix may only need targeted tests. A broader change may need package, workspace, or merge-ready status before the agent is allowed to proceed.
- TargetedTests: validate a narrow change before the loop moves on.
- Package: prove one package is green under the contract.
- Workspace: keep the full workspace healthy, not just the local patch.
- MergeReady: clear the last gate before autonomous action.
That ladder is easy to miss, but it explains the philosophy of the whole project: autonomy is gated by evidence, not enthusiasm.
What it replaces, and what it refuses to be
| Approach | Optimizes for | What breaks first |
|---|---|---|
| Classic Python agent harness | Fast prototyping and quick edits | Behavior drifts, recovery stays ad hoc, and hardening takes effort |
| claw-code-parity Rust runtime | Behavior-preserving migration with runtime discipline | Change becomes more deliberate, which is the point |
| Generic agent framework | Broad flexibility and reusable abstractions | Parity guarantees and session-level rigor are usually weak |
claw-code-parity is its own category because it is optimizing for continuity, not novelty. Python makes experiments cheap. Generic frameworks make experiments broad. Parity work makes the runtime predictable enough to survive a rewrite without changing what the user sees.