Babysitter: Putting a Hard-Coded Boss Inside the AI Agent
How a5c-ai uses event-sourcing and "Quality Convergence" to turn flaky LLM coding into a deterministic engineering pipeline.

We use Babysitter— an orchestration SDK that runs complex, multi-step AI workflows on top of Claude Code. It breaks down big tasks into quality-gated phases, delegates to specialized agents, and runs convergence loops until results meet a bar.
- Babysitter uses a two-loop architecture that separates deterministic JavaScript authority from probabilistic LLM execution.
- The framework employs event-sourcing to create an immutable journal of every agent action for deterministic replay and debugging.
- Quality Convergence forces the agent to iterate on tasks until it meets evidence-based scoring thresholds rather than self-declaring completion.
- Four distinct layers of hard-coded guardrails prevent agents from exceeding budgets or violating sandbox security policies.
The Hallucination of Progress
The dream of "vibe-coding"—giving an AI agent a prompt and watching it build an entire application—often collides with a frustrating reality. Agents are incredibly good at starting tasks but notoriously bad at finishing them correctly. They get stuck in hallucination loops, claiming a bug is fixed while the tests are still failing, or skipping crucial verification steps entirely.
Most agentic frameworks treat the Large Language Model (LLM) as the pilot, trusting it to navigate the execution flow. Babysitter, an orchestration framework by A5C AI, flips this dynamic entirely. Built primarily as a plugin for Anthropic's Claude Code, it treats the LLM as a talented but distractible intern. To manage this intern, Babysitter installs a deterministic, hard-coded "Manager" that literally cannot be lied to.
Babysitter enforces obedience to agentic workforces, enabling them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration.
The Two-Loop Architecture
The core innovation of Babysitter is its "Two-Loops" architecture. This design separates the authority that governs the workflow from the intelligence that executes the tasks.
The Outer Loop is symbolic. Written in JavaScript, it acts as the deterministic source of truth. It defines the mandatory phases of a project, sets the budget, and enforces quality gates. The Inner Loop is agentic. This is where the LLM (like Claude) actually writes code, runs commands, or analyzes logs.
Crucially, the Inner Loop cannot advance the overall state of the project. Only the Outer Loop can declare a phase complete, and it only does so when the agent provides cryptographic or evidence-based proof—like passing a specific test suite.
Event-Sourcing the Mind
To maintain this level of control, Babysitter employs event-sourcing. Instead of maintaining a mutable state database, it records every decision, tool call, and result as an immutable event in a journal (stored in .a5c/runs/).
This approach provides "Time-Travel Debugging" for AI agents. If an agent crashes or a human intervention is required, the system doesn't lose its place. It can deterministically replay the journal up to the point of failure, allowing developers to see exactly where the LLM's logic deviated from the plan.
Because LLMs frequently output malformed JSON or unexpected characters, Babysitter includes defensive engineering scripts like repair_json_text.py to sanitize responses before they hit the deterministic orchestrator, preventing the system from crashing due to a stray Unicode character.
Quality Convergence: The PID Controller for Code
Babysitter's most significant departure from standard agent frameworks is its concept of "Quality Convergence." Instead of running a prompt once and hoping for the best, it implements a feedback loop akin to a PID controller in engineering.
Using specific markdown templates (like score.md and decide.md) that act as structured prompts, the system forces the agent to evaluate its own work against strict criteria. The agent must iterate—submitting code, receiving a score, and refining the output—until it hits a predefined quality threshold (e.g., a score of 90+).
| Feature | Standard 'Vibe-Coding' Agents | Babysitter Framework |
|---|---|---|
| Execution Model | Linear execution or probabilistic loops | Deterministic, event-sourced state machine |
| Completion Logic | LLM declares 'I am done' | System verifies evidence (e.g., passing tests) |
| Debugging | Reading raw terminal logs | Deterministic replay via the Journal |
| Human Intervention | Ad-hoc chat interruptions | Structured, hard-coded breakpoints |
The Guardrail Layers
To prevent "runaway agents" from executing destructive commands or burning through API credits, Babysitter implements four distinct layers of guardrails:
- Capability: Strict sandboxing, such as path restrictions that prevent the agent from accessing files outside the project directory.
- Budget: Hard limits on token spend and the maximum number of tool calls allowed per phase.
- Policy: Behavioral rules, such as "never exfiltrate secrets" or mandatory coding standards.
- Behavioral: Output formatting requirements, ensuring the agent always returns parsable structured data.
By treating the LLM as an unreliable component within a highly reliable system, Babysitter provides a glimpse into the future of autonomous engineering. It shifts the focus from making agents smarter to making them obedient, verifiable, and safe to deploy on complex, multi-step workflows.