forge-framework: The AI coding pipeline that reviews the brief, the code, and the business goal

MatthiasMRC's Forge turns Claude Code into a three-gate production system, where tests, AI review, and vision alignment all have to agree before a PR ships.

11 min read • View on GitHub • More from MatthiasMRC

A manuscript-like PR dossier moves across three editorial desks, each station marked by a different kind of approval. The scene shows that Forge treats shipping code as a sequence of judgments, not a single automated pass or fail.
Forge is not just code generation wrapped in a nicer interface. It is a staged review process that keeps asking whether the work still deserves to ship.
Key Takeaways

Forge is easy to mistake for another wrapper around an AI coding agent. It is not. The interesting move is that it treats generation as only one step in a longer editorial process, where a branch has to satisfy deterministic checks, survive AI review, and then prove it still matches the brief.

The real product is the review process

That is why Gate 3 matters. Most agent workflows stop when the code passes lint or tests. Forge keeps going until the system reopens the PRD and architecture docs, then asks a harder question: does this diff still solve the thing it was meant to solve?

Why MatthiasMRC built Forge

The repo has the feel of a personal operating system that was made public, not a broad platform product. MatthiasMRC's GitHub footprint points toward workflow-heavy work, including AI agent tooling and Flutter-oriented projects, so Forge reads like a bridge between planning and execution inside that larger stack.

A hedcut-style portrait of MatthiasMRC rendered in black ink on pure white. It frames the project as the work of a builder who thinks in systems and workflow rules rather than a glossy product team.

There is not much public commentary around the repo, which makes the code itself the best source of intent. That helps here. Forge feels less like a framework in the abstract and more like a concrete answer to a very specific pain: how do you keep an AI agent from shipping something locally plausible but strategically wrong?

Three gates, three kinds of truth

Forge draws a sharp line between three kinds of validation. Gate 1 is mechanical, with linting, type checks, coverage, and pattern checks. Gate 2 is interpretive, where Claude reviews the diff for security, architecture, and completeness. Gate 3 is the most unusual, because it compares the change against the original product intent, not just the code around it.

The gates separate mechanical correctness, code review, and product alignment. That separation is the whole point.

That structure matters because it prevents the common failure mode of AI-assisted coding: a fast, coherent answer to the wrong question. Forge does not trust one review pass to carry both implementation quality and product judgment. It splits them apart, which is exactly how serious editorial systems work.

Inside /forge and the TDD loop

The `/forge` command is the engine room. It behaves like a state machine, finding the next story, forcing a red-green-refactor loop, and capping the number of attempts so the agent does not spin forever. That makes the workflow feel less like chatting with a model and more like running a tightly governed production checklist.

forge:
  loop: RED -> GREEN -> REFACTOR
  max_attempts_per_gate: 3
  gate1:
    lint: true
    typecheck: true
    coverage_threshold: 75
    forbidden_patterns:
      - print()
      - TODO
  gate2:
    review_focus:
      - security
      - architecture
      - completeness
  gate3:
    verify_against:
      - prd
      - architecture_docs

The template layer matters too. Forge is stack-aware, with checks that can be adapted for Python, JavaScript, or Flutter projects. That makes it feel like configuration as policy, not just a set of helper scripts.

What Forge is better than, and what it is not

Forge is not a model. It is not a build system in the conventional CI sense either. It sits on top of both, acting as a supervisory layer that decides when the machine is allowed to keep moving and when it has to go back and read the brief again.

WorkflowWhat it optimizesWhat it still leaves to chance
Prompt-only agent loopSpeed and local convenienceWhether the result still matches the original goal
Standard CIDeterministic checks and repeatabilityProduct intent and architectural fit
ForgeA staged judgment pipelineIt still depends on good briefs and honest tests

That comparison is the clearest way to place it. Prompt-only workflows are quick but porous. CI is reliable but narrow. Forge is trying to occupy the gap between the two, where AI-generated code becomes reviewable work instead of hopeful output.

The bigger argument

The real takeaway is bigger than Claude Code or any one repo. Forge suggests an AI-native SDLC where judgment is staged, not improvised. If that future arrives, the winning teams will not just generate code faster. They will get much better at deciding when code is actually ready to count.