babysitter-breakpoints: a5c-ai/babysitter: Putting AI Agents on a Leash
How a local-first middleware uses strict human breakpoints to turn unpredictable LLM execution into a deterministic state machine.
- Babysitter abandons the pursuit of total agent autonomy in favor of deterministic execution constrained by human breakpoints.
- The system uses a blocking loop in its CLI to physically pause an agent's process until a human reviewer provides an explicit release signal.
- A local Express.js web interface acts as a real-time pull request UI, allowing humans to inspect agent context safely before approving the next step.
- A background worker and SQLite state machine ensure that stale agent requests expire predictably if ignored by human reviewers.
The Illusion of Autonomy
The prevailing narrative in AI development is the pursuit of total autonomy. Frameworks push agents to plan, execute, and self-correct endlessly. In practice, this often results in expensive retry loops and unpredictable hallucinations. When an autonomous agent hits a snag, its instinct is to guess a fix rather than ask for help.
The a5c-ai/babysitter project proposes the exact opposite approach. It is a rebellion against probabilistic execution. By treating the LLM as an erratic worker that must be strictly managed, Babysitter proves that enterprise-grade AI requires less freedom and more human bottlenecks.
Babysitter enforces obedience to agentic workforces, enabling them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration.
The Agentic Pause
The core technical challenge is forcing an AI to stop. Most agents operate asynchronously, streaming tokens or chaining API calls. Babysitter solves this by providing a CLI that the agent is instructed to use. When the agent encounters a specific step or an error, it invokes the babysitter-breakpoints CLI with a context payload.
// The core blocking mechanism in bin/breakpoints.js
let status = "waiting";
while (status === "waiting") {
await new Promise(resolve => setTimeout(resolve, 2000));
const response = await fetch(`${API_URL}/breakpoints/${id}`);
const data = await response.json();
status = data.status;
}
This simple while loop is profoundly effective. It blocks the agent's process locally. The agent is frozen in time, unable to spend tokens or hallucinate code, until the API returns a changed status.
A Pull Request UI for LLMs
While the agent sits in its blocking loop, the human reviewer interacts with a local Express.js and Vanilla JS web interface. This UI flips the standard chat paradigm. Instead of a human prompting an AI, the AI has prompted the human.
The interface serves as a real-time pull request for agent actions. The API uses a resolveContextPath function to ensure that when a human views a file, the system only serves files within the defined repository root. This prevents path traversal vulnerabilities if an erratic agent attempts to attach sensitive system files to its context payload. Dynamic syntax highlighting turns the raw agent context into a readable review environment.
The Local-First State Machine
Babysitter relies on a local SQLite database to maintain state. This choice emphasizes its role as a local-first middleware designed to run alongside the agent on the same machine, rather than as a centralized corporate service.
To prevent agents from hanging indefinitely if a human ignores a request, a background worker queue manages Time-to-Live (TTL) expirations. Stale breakpoints are automatically marked as expired, allowing the agent to fail gracefully or retry according to its overarching instructions.
| Feature | General Autonomous Frameworks | a5c-ai/babysitter |
|---|---|---|
| Execution Model | Probabilistic Planning | Deterministic State Machine |
| Error Handling | LLM Self-Correction (Retry Loop) | Human Breakpoint (Pause) |
| State Storage | In-Memory or Vector DB | Local SQLite Journal |
| Primary Goal | Maximum Autonomy | Maximum Compliance |
Define your workflow in code - Babysitter enforces every step, ensures quality gates pass before progression, requires human approval at breakpoints, and records every decision in an immutable journal.
The Future of Agentic Skills
The repository includes a .codex/skills/ directory, which points to a broader shift in software development. We are moving from building libraries intended for humans to packaging tools explicitly built to be invoked by agents.
By providing a SKILL.md file, the repository teaches the agent how to call the CLI and what parameters to provide. Babysitter is not just a debugging tool. It is a blueprint for how developers will eventually structure all interactions between chaotic language models and rigid local infrastructure.