agent-relay: Agent Relay: The Stateless Execution Engine for Autonomous Engineering

Why the future of AI coding belongs to "Fresh-Instance" pipelines and adversarial review loops rather than long-form chat.

• View on GitHub • More from brooksjordan

A high-precision architectural blueprint that warps into surrealist, melting shapes as it scrolls down, while a mechanical hand tries to draw a straight line over the melting portion.
The "Context Drift" problem: as logical context builds up over a long session, the foundation degrades.

Key Takeaways

The Hallucination Spiral

The "Agentic Workflow" is suffering from a mid-life crisis. Early demonstrations of autonomous coding agents promised software that writes itself. In practice, developers quickly encounter the hallucination spiral: an AI agent works perfectly for the first ten minutes, but as the session stretches on, it loses the thread. It hallucinates APIs, forgets the architectural constraints it established earlier, and eventually starts trying to fix bugs it just created.

This degradation is known as context drift. When an agent operates in a continuous, long-running chat session, the accumulation of tokens acts like cognitive bloat. The model becomes overwhelmed by its own history. Agent Relay solves this not with more complex prompting, but with a brutalist, infrastructure-first approach: it kills the agent.

The Fresh-Instance Protocol

Instead of a monolithic session, Agent Relay treats AI agents like stateless serverless functions. Guided by a core script named loop.ps1, the system decomposes a high-level priority report into a discrete JSON manifest of tasks. For every single task, it spins up a completely fresh instance of the Claude Code CLI.

Once the task is complete, the instance is terminated. This "Fresh-Instance" protocol ensures a zero-hallucination baseline. There is no lingering context to corrupt the next operation. It is a DevOps approach to prompt engineering, favoring absolute state control over conversational continuity.

The 9-stage relay pipeline moves a feature from a priority report to a completed pull request through isolated execution environments.

PowerShell: The Unlikely Orchestrator

In an ecosystem dominated by Python, Agent Relay's choice of PowerShell stands out. The decision is rooted in the messy reality of orchestrating interactive CLI tools. Python often requires complex subprocess and pseudo-terminal boilerplate to manage tools that expect human interaction. PowerShell, particularly on Windows-heavy enterprise environments, handles these interactive streams and exit codes natively.

The framework's invoke-safeexpression wrapper shifts the definition of failure from "anything written to stderr" to strict non-zero exit codes. This prevents overzealous build tool warnings from crashing the automation loop. It also forces a "Human-in-the-Loop" visibility pattern by intentionally spawning visible terminal windows, ensuring developers can monitor the agent's real-time decisions.

Warning: Agent Relay runs AI agents autonomously. It consumes API credits and modifies code in your repository. Always run on a feature branch, never on main.

Project README, Repository documentation · brooksjordan/agent-relay README

Adversarial Review: Trust, but Verify

Generating code is only half the battle; validating it autonomously is the real bottleneck. Agent Relay implements an "Adversarial Review" phase. Rather than relying on the original agent to self-correct—which often leads to confirmation bias—the pipeline spins up a separate, isolated AI instance tasked exclusively with attacking the newly written code.

DimensionStandard Agent WorkflowsAgent Relay Architecture
Execution StateLong-running chat sessions (Volatile)Fresh-instance per task (Stateless)
Error CorrectionSelf-correction (High confirmation bias)Adversarial Review (Unbiased, isolated instances)
Context ManagementCumulative token bloatCompound Learning via extracted summaries
Failure HandlingOften halts or loops indefinitelyResumable state machine via JSON checkpoints

Synthesizing the Project Soul

If every agent is killed after its task, how does the system learn the quirks of a specific codebase? The answer lies in the "Post-mortem" stage. A script named gather-sessions.ps1 aggregates the raw transcripts, token usage, and git diffs from the freshly terminated agents.

A supervisor LLM analyzes this data to extract hard-won lessons—like "Auth is handled in /lib, not /api"—and appends them to a persistent CLAUDE.md file. This creates a long-term memory that survives the death of individual instances. The project gets smarter the more it fails, turning ephemeral agent activity into durable engineering capital.

Glowing, translucent cubes being dropped into a large, industrial mechanical sifter. Most of the cubes crumble and fall through as dust, but a few solid, dark-inked 'knowledge stones' remain on the mesh, being channeled into a leather-bound ledger.
Compound Learning extracts persistent architectural knowledge from the noise of individual agent sessions.