eforge: the build system that makes AI review its own code blindfolded
A TypeScript repo that turns PRDs into isolated worktrees, separates builder from reviewer, and ships only what survives the pipeline.

I built it because I was tired of keeping the orchestration logic in my head - spawning a separate session for a blind review, switching back to the implementing session to evaluate results, deciding what to build next.
- eforge’s core move is to separate code generation from code approval so the same agent never grades its own work.
- The repo treats specifications as the unit of work, then routes them through planning, building, blind review, and validation.
- Git worktrees give the system real parallelism without turning the repository into a collision zone.
- The project argues that software delivery gets more reliable when AI behaves like a pipeline, not a conversation.
The real trick is the review boundary
Most AI coding tools optimize the front end of the job: produce code fast, keep the human in the loop, and hope the result is good enough. eforge makes a different bet. It assumes the hardest part is not writing code, but preventing the writer from becoming its own best critic.
Traditional build systems transform source code into artifacts. An agentic build system transforms *specifications* into source code - then verifies its own output.
From spec to queue
The input to eforge is not a file edit. It is an intent object: a prompt, a markdown note, a PRD, or a plan that gets normalized and dropped into a queue. A daemon claims the work, then drives it through a staged pipeline that can ask for clarification, plan the change, build it, review it, and validate it before merge.
# PRD: export issue data to CSV
- Add a command that turns issue records into a CSV file.
- Keep the change isolated.
- Validate with tests before merge.
Errand, Excursion, Expedition
eforge does not treat every task the same way. Small fixes stay light. Multi-file features get a more deliberate plan and a blind review. Larger refactors bring architecture docs, module decomposition, cohort-style validation, and parallel builds. That tiered model matters because most tools either under-process simple work or under-support hard work.
| Mode | What enters the pipeline | What gets added | Best fit |
|---|---|---|---|
| Errand | One clear change | Minimal orchestration | Single-file fixes and small edits |
| Excursion | A scoped feature or refactor | Plan, build, blind review | Multi-file work with moderate risk |
| Expedition | A broad change across modules | Architecture review, decomposition, parallel worktrees | Cross-cutting changes that need coordination |
Why worktrees matter
The technical spine is ordinary Git, used with unusual discipline. Each plan gets its own isolated worktree, so concurrent branches of work can run in parallel without trampling each other’s files. Then the system merges in dependency order, which turns Git into coordination infrastructure instead of a passive history log.
This is not Claude Code, and that is the point
Interactive assistants help a human write. eforge orchestrates the delivery system around that writing. The README describes it as a complement to tools like Claude Code and Pi, which is the right mental model: those tools are for planning and conversation, while eforge is for execution, review, and merge discipline.
| Layer | eforge | Interactive coding assistants | General multi-agent frameworks |
|---|---|---|---|
| Primary input | Specifications such as PRDs, markdown, and plans | Chat and live code context | High-level goals |
| Unit of work | A verified build pipeline | A human-guided coding session | A task or objective |
| Review model | Blind, staged, adversarial | Human in the loop | Depends on the framework |
| Concurrency model | Git worktrees and dependency order | Usually one active session | Often abstract orchestration |
| Best fit | Software delivery with quality gates | Thinking and editing with a human | Broad goal-seeking and experiments |
Built by someone who uses it on itself
The project feels lived in, not theoretical. The repository description and the public discussion both make clear that it is a young system moving fast, used daily, and already trusted to build itself. That kind of self-use matters because it turns the codebase into evidence, not just a pitch.
What eforge argues about the future of coding
The deeper argument here is not that AI should write more code. It is that software delivery becomes more trustworthy when the main artifact is a specification and the main system is a pipeline. eforge replaces the old question, "Can the model finish the file?" with a better one: "Can the build survive the review gates, the merge order, and the final validation?"