The Adversarial Developer: Inside xmkevinchen/agentic-engineering
How a Claude Code plugin uses operating system process isolation and late-stage adversarial agents to stop LLMs from hallucinating in echo chambers.

Current AI coding tools are too fragmented and single-threaded. I built Agentic Engineering to bring structure and collaboration to AI coding through a disciplined pipeline and specialized agents.
- Agentic Engineering weaponizes dissent by spawning late-stage adversarial agents to audit code decisions without access to prior chat history.
- The framework enforces an operating system model where agents communicate only with a central Team Lead, eliminating lateral token noise.
- Cross-model routing uses Gemini and OpenAI to bypass Claude's inherent training biases and provide independent validation.
- A strict state-machine pipeline prevents agents from writing code until a formal plan is explicitly marked as reviewed.
The AI Echo Chamber
Artificial intelligence coding assistants share a fundamental flaw: they are eager to please. When multiple Large Language Models (LLMs) are placed in a room and told to build software, they default to consensus. They reinforce each other's bad ideas and hallucinate complex solutions to simple problems.
This sycophancy loop means an AI will happily write 500 lines of boilerplate for an architecture that a senior engineer would reject in five minutes. The solution is not better prompting. The solution is automated, structural skepticism.
Enter the Red Team
Agentic Engineering is a plugin for Anthropic's Claude Code CLI. Rather than treating AI as an eager intern, it treats AI agents like untrusted operating system processes. Its core innovation is the Doodlestein Protocol, a suite of adversarial agents deployed late in the development cycle.
These agents (labeled strategic, adversarial, and regret) are intentionally blinded to the team's chat history. They are denied access to the sunk-cost debate that produced the current plan. Instead, they are handed the proposed architecture and strictly prompted to prove why the code will fail in six months. It is not just a code generator. It is an automated red team.
The Hub-and-Spoke Kernel
Under the hood, Agentic Engineering implements a Team Lead (TL) abstraction. In typical multi-agent frameworks, agents converse laterally, generating massive amounts of token noise. In this system, agents only communicate with the central Claude session.
This mimics an OS kernel managing isolated processes. It reduces context window pollution and ensures the human-facing session remains the single source of truth.
Bypassing the Blind Spots
To prevent Claude from grading its own homework, the architecture leverages Model Context Protocol (MCP) servers to route tasks to Gemini and OpenAI. Because different model families possess different training biases, using them as architectural judges provides genuine second opinions.
The execution phase is equally rigid. A 3-strike test-driven development circuit breaker prevents infinite looping. If an agent fails to fix a bug after three attempts, the process hard-stops and demands human intervention.
Process Over Play-Acting
While other multi-agent systems simulate a virtual company with a CEO and HR department, Agentic Engineering implements a ruthless CI/CD pipeline. The pipeline enforces a strict state machine: Analyze, Discuss, Plan, Work, Review.
You cannot trigger the implementation phase unless a plan is explicitly marked as reviewed. This approach prioritizes discipline over open-ended simulation.
| Feature | Virtual Companies (MetaGPT) | Agentic Engineering (ae) |
|---|---|---|
| Architecture Paradigm | Open-ended simulation | Strict state machine |
| Agent Communication | Lateral (Agent-to-Agent) | Hub-and-Spoke (Agent-to-TL) |
| Review Mechanism | Consensus building | Adversarial blind-spotting |
| State Persistence | Ephemeral chat history | Markdown files written to docs/ |