recursive-mode: The Repo-Native Operating System for AI Engineering
A document-first workflow that moves agent state out of chat, into locked Markdown, and through a strict audit loop that humans can actually inspect.
- recursive-mode moves the source of truth for agent work out of chat and into locked repository files.
- Its real innovation is governance, not automation, because every phase can be audited and resumed from disk.
- The audit loop and content hashes make agent output behave more like controlled engineering records than disposable prompts.
- Its memory model turns one run into structured context for the next run without rewriting history.
Most agent tools optimize for a fast conversation. `recursive-mode` optimizes for a trustworthy record. That difference sounds small until a task stretches across hours, files, and model handoffs, where chat history turns into a messy approximation of state.
The Chat Box Is the Wrong Database
The project starts from a blunt premise: chat context is a poor medium for engineering state. Requirements, analysis, plans, test output, and verification all pile up in one long thread, and the thread becomes harder to trust as it grows. `recursive-mode` moves the important pieces into .recursive/ so the repo, not the sidebar, becomes the durable system of record.
chat context is a poor medium for engineering state.
That is why the workflow reads less like a prompt recipe and more like a governance layer. It treats the chat as a control surface for commands, not as a notebook for intent. Once you see it that way, the project stops looking like another agent wrapper and starts looking like an operating system for the work itself.
What Lives Inside `.recursive/`
Inside the repo, the workflow writes a ledger. Global files like RECURSIVE.md, STATE.md, and DECISIONS.md describe the rules, the current condition, and the reasoning trail. Run-specific folders under .recursive/run/<run-id>/ hold the live artifacts for one execution, including requirements, codebase analysis, implementation plans, and verification notes.
.recursive/
RECURSIVE.md
STATE.md
DECISIONS.md
memory/
incidents/
patterns/
run/
<run-id>/
00-requirements.md
01-as-is.md
02-to-be.md
03-implementation.md
04-verification.md
The point is not just organization. The point is that every artifact has a home, and every home has a job. A human can inspect the files without special tooling, while the agent can be forced to advance only through the accepted sequence.
The Audit Loop That Won't Let Go
This is the technical center of the project. Work advances through a strict phase machine: DRAFT -> AUDIT -> REPAIR -> RE-AUDIT -> PASS -> LOCK. The scripts recursive-status.py and recursive-lock.py do more than summarize progress. They enforce it.
`recursive-status.py` checks whether the run has actually earned the next state. `recursive-lock.py` calculates a content hash over the artifact, excluding the lock line itself, so any later drift invalidates the seal. That is a very different promise from a chat agent that says it is done because it sounds done.
recursive-mode addresses them by moving the important state into repository docs and advancing work through explicit recursive phases during planning, implementation, and verification.
Why Markdown Beats JSON Here
JSON would be machine-friendly, but it would be less legible to the people who need to audit a run at speed. Markdown keeps the records readable in plain editors, pull requests, and diffs. That legibility is not a cosmetic choice. It is part of the control system.
| Dimension | Markdown in recursive-mode | Typical JSON or DB-backed state |
|---|---|---|
| Human review | Readable in a normal editor and easy to diff | Often needs tooling or a UI to interpret |
| Audit trail | Lives in versioned files with history | Lives in records that may be hidden behind an app |
| Collaboration | Easy for humans and models to inspect together | Better for machines, less natural for review |
| Failure mode | A broken lock or mismatched file is obvious | State can drift quietly unless the app enforces it |
| Intent | Legibility is part of governance | Structure is optimized mainly for parsing |
The project is not using Markdown because it is quaint. It is using Markdown because a control system that humans cannot inspect is only half a control system.
Addenda, Not Rewrites
Locked history is treated as history. If the agent finds a mistake after a phase is sealed, it does not quietly mutate the old record. It appends an Addendum. That keeps the audit trail intact and makes the evolution of a decision visible instead of pretending the original mistake never happened.
| Action | Locked-history model | Rewrite-first model |
|---|---|---|
| Correcting an error | Add a new note that explains the correction | Edit the old file and blur the trail |
| Trust signal | Readers can see what changed and why | Readers must assume the latest version is truthful |
| Compliance style | Closer to engineering logs and legal records | Closer to disposable working notes |
| Cost | More disciplined, slightly more verbose | Shorter in the moment, weaker later |
That restraint matters because AI systems are good at producing confident revisions. A governed workflow needs the opposite instinct: preserve the record, then layer corrections on top of it.
Subagents Are Treated Like Untrusted Contractors
The controller does not accept a subagent's claim at face value. It checks the claimed file impact against the actual disk state before the phase can pass. That sounds bureaucratic until you realize it is the only sane response to multiple agents making changes in parallel.
This is where the project stops being merely tidy and starts feeling serious. If your orchestration layer cannot distinguish claimed work from actual work, it is just a narrative generator with a shell prompt.
How It Differs From Agents and Orchestrators
`recursive-mode` is not trying to compete with an executor or a fleet manager. It is the harness that constrains execution. That puts it in a different category from chat-driven coding agents and broader orchestration platforms.
| Primary role | Where state lives | How work advances | How trust is enforced | Best fit | Main limitation |
|---|---|---|---|---|---|
| `recursive-mode` | Repo docs in `.recursive/` | Strict phases and locks | File hashes, audits, and disk checks | Long-running engineering tasks that need custody | It is a harness, so it still depends on an executor |
| Typical coding agents | Chat history and local edits | Conversational prompting | Mostly model judgment and user review | Fast solo execution on smaller tasks | Context rot and weaker traceability |
| Agent orchestrators | Platform database or knowledge graph | Events, schedules, and integrations | Platform policies and dashboards | Coordinating many agents across systems | Less repo-native and usually more centralized |
recursive-mode is a workflow for agentic engineering that I have been using daily and refining over the past four months. It starts from a simple observation: chat context is a poor medium for engineering state. It solves that problem and opens up a few new possibilities for further coding performance improvements, by using generative recursion during planning, implementation and verification.
The clean distinction is that the tool does not promise to do all the coding itself. It promises to make the coding process governable.
Why This Matters Beyond One Repo
The larger idea is portable: agent work should be inspectable, resumable, and local-first. If a workflow can survive outside a chat transcript, it can survive handoffs, interruptions, and long-running tasks without collapsing into guesswork. That is the real payoff here, and it reaches well beyond one open-source repo.
The best systems do not just automate action. They preserve intent. `recursive-mode` is compelling because it treats that preservation as a first-class engineering problem.