Citadel: The Markdown Operating System for AI Agent Fleets
How a repo-local brain, a work queue in plain text, and Git worktrees turn coding agents into something closer to an actual engineering system.
- Citadel’s core move is to make Markdown the control plane, so state stays legible, editable, and durable across sessions.
- The system is designed to defeat context amnesia by storing memory, handoffs, and work queues in repo-local artifacts instead of in a chat window.
- Routing and hooks matter as much as generation, because Citadel constrains intent before execution and gates tool use before it becomes chaos.
- Its real differentiator is not autonomous magic, but a Git-native operating loop for parallel agent labor.
Citadel’s most interesting trick is also its least glamorous one. It treats plain text as infrastructure. The repository’s own Markdown files hold memory, campaign state, task routing, and handoffs, which means both humans and agents can inspect and edit the same source of truth.
The Control Plane Is Just Markdown
That design choice changes the shape of the system. Instead of trusting a model to remember what happened three steps ago, Citadel writes the work down in repo-local artifacts. A task is not just a prompt. It is a record, a dependency, a status field, and a future recovery path.
| System | State model | Human editability | Coordination style | Safety model | Best fit |
|---|---|---|---|---|---|
| Citadel | Repo-local Markdown and Git state | High | Routed campaigns and work queues | Hooks and policy enforcement | Agent labor that must stay inspectable |
| Chat-first coding agent | Ephemeral conversation context | Low | Ad hoc prompts | Prompting and best effort | Quick, local edits |
| Generic multi-agent framework | Structured in framework objects | Medium | Agent graphs or role choreography | Framework-level guardrails | App builders who want abstraction |
| Managed assistant API | Vendor-hosted session state | Low to medium | Tool-enabled conversations | Platform controls | Product teams that want hosted infra |
Why Chat Fails, and Why Citadel Doesn’t Want Chat to Be the System
Most coding agents are strong at immediate execution and weak at continuity. They can edit files, call tools, and respond quickly, but they forget the reasoning behind the work as soon as the session shifts. Citadel is built around that failure mode.
The README frames the project as an agent orchestration harness for Claude Code, but the deeper idea is stricter: it wants to route any task through the right tool at the right scale. That means the system is not just a shell around a model. It is an attempt to make agent work persistent, steerable, and recoverable.
Citadel â Agent Orchestration Harness for Claude Code Run autonomous coding campaigns with Claude Code.
The Memory Layer: Blocks, Handoffs, and Verifiable Recall
Citadel’s memory system is not a vague vector store story. It is a set of explicit blocks in .planning/, organized so agents can inject prior decisions, failure patterns, and verification recipes into later work. That matters because engineering work is full of repeatable mistakes, and the project is trying to make those mistakes durable enough to learn from.
memory blocks:
- architecture-decisions
- failure-patterns
- verification-recipes
handoff:
- what changed
- why it changed
- what still needs verification
The important detail is the loop. Agents do not just write notes. They compile and lint their own operating memory, then hand that memory forward. In practice, that turns previous sessions into a reusable project asset instead of a pile of dead chat logs.
Routing Intent Before It Becomes Chaos
The `/do` command acts like a dispatcher. It classifies an intent, selects a skill or campaign, and sends the request through a defined protocol rather than letting the model wander. This is the difference between asking an assistant to help and asking a system to execute a workflow.
| Chat prompt | Routed command |
|---|---|
| "Can you fix this?" | Classify the task, select a skill, and enforce the protocol |
| Memory stays in the context window | Memory is written into .planning/ blocks |
| Tool use happens opportunistically | Tool use happens through a defined path |
| Success depends on the current conversation | Success depends on the stored operating loop |
Fleet Mode: Parallel Agents Without Merge-Conflict Theater
The fleet layer is where Citadel stops looking like a clever prompt wrapper. Work is serialized into a queue, dependencies are checked, and isolated Git worktrees keep parallel agents from colliding with one another. That is a real operating model, not a metaphor.
function isRunnable(task) {
return task.status === 'pending' && task.deps.every(dep => completed.has(dep));
}
// If runnable, assign the task to an isolated worktree
// then write the handoff back into .planning/
That choice is quietly decisive. Git worktrees give each agent a sandbox with shared history and isolated state. The queue decides who runs next. Dependencies decide what can run at all. Citadel is managing labor, not just generating code.
| Dimension | Single chat agent | Citadel fleet |
|---|---|---|
| Parallel work | Manual and brittle | Scheduled through a queue |
| State isolation | Shared conversation context | Separate Git worktrees |
| Dependencies | Implicit in prompting | Explicit in task metadata |
| Recovery | Ask the model again | Read the repo-local handoff |
Safety Is Not a Prompt. It’s a Hook.
Citadel’s hook system makes safety procedural. Before a tool action lands, the event is normalized and checked. If the policy says no, the action stops there. That is much stronger than asking a model to behave well in the abstract.
This is the right place for guardrails because it is the narrowest choke point in the loop. A planner can propose anything. A hook can still decide what may touch the filesystem, the shell, or the network.
The result is a system that treats agent behavior like software behavior. Inputs are normalized. Events are checked. Policy is enforced. Only then does execution proceed.
Where Citadel Fits in the Agent Landscape
Compared with LangGraph, AutoGen, CrewAI, AutoGPT, or managed assistant APIs, Citadel is more specific about where the truth lives. It is repo-native, Markdown-first, Git-native, and human-editable. That makes it less general than a framework, but more operational for teams that want the work itself to remain inspectable.
| System | Owns the work state? | Human-readable state? | Git-native? | Primary strength |
|---|---|---|---|---|
| Citadel | Yes | Yes | Yes | Durable agent operations inside a repo |
| LangGraph | Partly | Depends on implementation | No | Stateful graphs and orchestration |
| AutoGen | Partly | No | No | Multi-agent conversation patterns |
| CrewAI | Partly | No | No | Role-based collaboration |
| Managed assistant APIs | Vendor-owned | Limited | No | Hosted persistence and tooling |
Citadel’s thesis is simple enough to fit on one line. If the agent is going to do real work, the work needs memory, routing, scheduling, and safety that survive the chat window. Markdown just happens to be the most surprising way to make that practical.