The Adversarial Developer: Inside xmkevinchen/agentic-engineering

How a Claude Code plugin uses operating system process isolation and late-stage adversarial agents to stop LLMs from hallucinating in echo chambers.

8 min read • View on GitHub • More from xmkevinchen

A single pristine gear sits in the center of a hall of mirrors, its reflections becoming increasingly distorted and tangled in the background. This represents the echo chamber problem of unconstrained multi-agent systems.
Without strict routing constraints, lateral agent-to-agent communication quickly degrades into an echo chamber of hallucinated complexity.

Current AI coding tools are too fragmented and single-threaded. I built Agentic Engineering to bring structure and collaboration to AI coding through a disciplined pipeline and specialized agents.

Kevin Chen, Creator · Introducing Agentic Engineering
Key Takeaways

The AI Echo Chamber

Artificial intelligence coding assistants share a fundamental flaw: they are eager to please. When multiple Large Language Models (LLMs) are placed in a room and told to build software, they default to consensus. They reinforce each other's bad ideas and hallucinate complex solutions to simple problems.

This sycophancy loop means an AI will happily write 500 lines of boilerplate for an architecture that a senior engineer would reject in five minutes. The solution is not better prompting. The solution is automated, structural skepticism.

Enter the Red Team

Agentic Engineering is a plugin for Anthropic's Claude Code CLI. Rather than treating AI as an eager intern, it treats AI agents like untrusted operating system processes. Its core innovation is the Doodlestein Protocol, a suite of adversarial agents deployed late in the development cycle.

A WSJ hedcut-style portrait of Kevin Chen.

These agents (labeled strategic, adversarial, and regret) are intentionally blinded to the team's chat history. They are denied access to the sunk-cost debate that produced the current plan. Instead, they are handed the proposed architecture and strictly prompted to prove why the code will fail in six months. It is not just a code generator. It is an automated red team.

An extreme close-up of a heavy jeweler's loupe resting on a perfectly machined metal cylinder. Inside the lens, a severe jagged hairline stress fracture is visible.
Adversarial agents are designed to find the hidden structural flaws that consensus-driven models ignore.

The Hub-and-Spoke Kernel

Under the hood, Agentic Engineering implements a Team Lead (TL) abstraction. In typical multi-agent frameworks, agents converse laterally, generating massive amounts of token noise. In this system, agents only communicate with the central Claude session.

This mimics an OS kernel managing isolated processes. It reduces context window pollution and ensures the human-facing session remains the single source of truth.

Unlike frameworks that allow agents to chat laterally, Agentic Engineering forces all communication through a central Team Lead.

Bypassing the Blind Spots

To prevent Claude from grading its own homework, the architecture leverages Model Context Protocol (MCP) servers to route tasks to Gemini and OpenAI. Because different model families possess different training biases, using them as architectural judges provides genuine second opinions.

The execution phase is equally rigid. A 3-strike test-driven development circuit breaker prevents infinite looping. If an agent fails to fix a bug after three attempts, the process hard-stops and demands human intervention.

Process Over Play-Acting

While other multi-agent systems simulate a virtual company with a CEO and HR department, Agentic Engineering implements a ruthless CI/CD pipeline. The pipeline enforces a strict state machine: Analyze, Discuss, Plan, Work, Review.

You cannot trigger the implementation phase unless a plan is explicitly marked as reviewed. This approach prioritizes discipline over open-ended simulation.

FeatureVirtual Companies (MetaGPT)Agentic Engineering (ae)
Architecture ParadigmOpen-ended simulationStrict state machine
Agent CommunicationLateral (Agent-to-Agent)Hub-and-Spoke (Agent-to-TL)
Review MechanismConsensus buildingAdversarial blind-spotting
State PersistenceEphemeral chat historyMarkdown files written to docs/