forge-framework: The AI coding pipeline that reviews the brief, the code, and the business goal
MatthiasMRC's Forge turns Claude Code into a three-gate production system, where tests, AI review, and vision alignment all have to agree before a PR ships.
- Forge is not really an AI coding helper, it is a judgment pipeline that forces a diff to survive tests, review, and brief-level scrutiny.
- Gate 3 is the project’s real invention because it asks whether the code still solves the original problem, not just whether it compiles.
- The /forge command behaves like a state machine, using a strict TDD loop and capped retries to keep the agent from wandering.
- Forge sits between prompt-only coding tools and standard CI, acting as a supervisory layer that protects product intent as much as code quality.
Forge is easy to mistake for another wrapper around an AI coding agent. It is not. The interesting move is that it treats generation as only one step in a longer editorial process, where a branch has to satisfy deterministic checks, survive AI review, and then prove it still matches the brief.
The real product is the review process
That is why Gate 3 matters. Most agent workflows stop when the code passes lint or tests. Forge keeps going until the system reopens the PRD and architecture docs, then asks a harder question: does this diff still solve the thing it was meant to solve?
Why MatthiasMRC built Forge
The repo has the feel of a personal operating system that was made public, not a broad platform product. MatthiasMRC's GitHub footprint points toward workflow-heavy work, including AI agent tooling and Flutter-oriented projects, so Forge reads like a bridge between planning and execution inside that larger stack.
There is not much public commentary around the repo, which makes the code itself the best source of intent. That helps here. Forge feels less like a framework in the abstract and more like a concrete answer to a very specific pain: how do you keep an AI agent from shipping something locally plausible but strategically wrong?
Three gates, three kinds of truth
Forge draws a sharp line between three kinds of validation. Gate 1 is mechanical, with linting, type checks, coverage, and pattern checks. Gate 2 is interpretive, where Claude reviews the diff for security, architecture, and completeness. Gate 3 is the most unusual, because it compares the change against the original product intent, not just the code around it.
That structure matters because it prevents the common failure mode of AI-assisted coding: a fast, coherent answer to the wrong question. Forge does not trust one review pass to carry both implementation quality and product judgment. It splits them apart, which is exactly how serious editorial systems work.
Inside /forge and the TDD loop
The `/forge` command is the engine room. It behaves like a state machine, finding the next story, forcing a red-green-refactor loop, and capping the number of attempts so the agent does not spin forever. That makes the workflow feel less like chatting with a model and more like running a tightly governed production checklist.
forge:
loop: RED -> GREEN -> REFACTOR
max_attempts_per_gate: 3
gate1:
lint: true
typecheck: true
coverage_threshold: 75
forbidden_patterns:
- print()
- TODO
gate2:
review_focus:
- security
- architecture
- completeness
gate3:
verify_against:
- prd
- architecture_docs
The template layer matters too. Forge is stack-aware, with checks that can be adapted for Python, JavaScript, or Flutter projects. That makes it feel like configuration as policy, not just a set of helper scripts.
What Forge is better than, and what it is not
Forge is not a model. It is not a build system in the conventional CI sense either. It sits on top of both, acting as a supervisory layer that decides when the machine is allowed to keep moving and when it has to go back and read the brief again.
| Workflow | What it optimizes | What it still leaves to chance |
|---|---|---|
| Prompt-only agent loop | Speed and local convenience | Whether the result still matches the original goal |
| Standard CI | Deterministic checks and repeatability | Product intent and architectural fit |
| Forge | A staged judgment pipeline | It still depends on good briefs and honest tests |
That comparison is the clearest way to place it. Prompt-only workflows are quick but porous. CI is reliable but narrow. Forge is trying to occupy the gap between the two, where AI-generated code becomes reviewable work instead of hopeful output.
The bigger argument
The real takeaway is bigger than Claude Code or any one repo. Forge suggests an AI-native SDLC where judgment is staged, not improvised. If that future arrives, the winning teams will not just generate code faster. They will get much better at deciding when code is actually ready to count.