Babysitter: Putting a Hard-Coded Boss Inside the AI Agent

How a5c-ai uses event-sourcing and "Quality Convergence" to turn flaky LLM coding into a deterministic engineering pipeline.

a5c-ai/babysitter

A large, intricate clockwork machine with a glowing, organic spark trapped inside a glass sphere at the center. The spark tries to move freely, but the gears only turn when a physical 'Verification Key' is inserted.
Babysitter forces the probabilistic LLM (the spark) to wait for deterministic verification (the key) before proceeding.

We use Babysitter— an orchestration SDK that runs complex, multi-step AI workflows on top of Claude Code. It breaks down big tasks into quality-gated phases, delegates to specialized agents, and runs convergence loops until results meet a bar.

Eyal David, Maintainer/Developer at A5C AI (likely) · We Locked an AI in a Room With One Phone Line

Key Takeaways

The Hallucination of Progress

The dream of "vibe-coding"—giving an AI agent a prompt and watching it build an entire application—often collides with a frustrating reality. Agents are incredibly good at starting tasks but notoriously bad at finishing them correctly. They get stuck in hallucination loops, claiming a bug is fixed while the tests are still failing, or skipping crucial verification steps entirely.

Most agentic frameworks treat the Large Language Model (LLM) as the pilot, trusting it to navigate the execution flow. Babysitter, an orchestration framework by A5C AI, flips this dynamic entirely. Built primarily as a plugin for Anthropic's Claude Code, it treats the LLM as a talented but distractible intern. To manage this intern, Babysitter installs a deterministic, hard-coded "Manager" that literally cannot be lied to.

Babysitter enforces obedience to agentic workforces, enabling them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration.

The Two-Loop Architecture

The core innovation of Babysitter is its "Two-Loops" architecture. This design separates the authority that governs the workflow from the intelligence that executes the tasks.

The Outer Loop is symbolic. Written in JavaScript, it acts as the deterministic source of truth. It defines the mandatory phases of a project, sets the budget, and enforces quality gates. The Inner Loop is agentic. This is where the LLM (like Claude) actually writes code, runs commands, or analyzes logs.

Crucially, the Inner Loop cannot advance the overall state of the project. Only the Outer Loop can declare a phase complete, and it only does so when the agent provides cryptographic or evidence-based proof—like passing a specific test suite.

The Symbolic Outer Loop dictates the workflow, while the Agentic Inner Loop performs the work within strict constraints.

Portrait of Eyal David.

Event-Sourcing the Mind

To maintain this level of control, Babysitter employs event-sourcing. Instead of maintaining a mutable state database, it records every decision, tool call, and result as an immutable event in a journal (stored in .a5c/runs/).

This approach provides "Time-Travel Debugging" for AI agents. If an agent crashes or a human intervention is required, the system doesn't lose its place. It can deterministically replay the journal up to the point of failure, allowing developers to see exactly where the LLM's logic deviated from the plan.

A close-up of a hand-drawn ledger where each page is a punch-card. A mechanical pen is mid-stroke, recording a 'Decision Event.' Previous pages are chained together by literal iron links.
Every action the agent takes is immutably recorded, creating a verifiable chain of events.

Because LLMs frequently output malformed JSON or unexpected characters, Babysitter includes defensive engineering scripts like repair_json_text.py to sanitize responses before they hit the deterministic orchestrator, preventing the system from crashing due to a stray Unicode character.

Quality Convergence: The PID Controller for Code

Babysitter's most significant departure from standard agent frameworks is its concept of "Quality Convergence." Instead of running a prompt once and hoping for the best, it implements a feedback loop akin to a PID controller in engineering.

Using specific markdown templates (like score.md and decide.md) that act as structured prompts, the system forces the agent to evaluate its own work against strict criteria. The agent must iterate—submitting code, receiving a score, and refining the output—until it hits a predefined quality threshold (e.g., a score of 90+).

FeatureStandard 'Vibe-Coding' AgentsBabysitter Framework
Execution ModelLinear execution or probabilistic loopsDeterministic, event-sourced state machine
Completion LogicLLM declares 'I am done'System verifies evidence (e.g., passing tests)
DebuggingReading raw terminal logsDeterministic replay via the Journal
Human InterventionAd-hoc chat interruptionsStructured, hard-coded breakpoints

The Guardrail Layers

To prevent "runaway agents" from executing destructive commands or burning through API credits, Babysitter implements four distinct layers of guardrails:

Every action proposed by the agent must pass through four layers of validation before execution.

By treating the LLM as an unreliable component within a highly reliable system, Babysitter provides a glimpse into the future of autonomous engineering. It shifts the focus from making agents smarter to making them obedient, verifiable, and safe to deploy on complex, multi-step workflows.