Citadel: The Markdown Operating System for AI Agent Fleets

How a repo-local brain, a work queue in plain text, and Git worktrees turn coding agents into something closer to an actual engineering system.

8 to 10 min read • View on GitHub • More from SethGammon

A paper-like operations room where a large Markdown ledger sits at the center, with thin lines running out to several isolated agent workstations arranged like fenced plots. A human hand edits the ledger while the fleet keeps moving, showing that state lives in text the whole team can inspect.
Citadel treats Markdown as the control plane, so humans and agents work from the same durable state.
Key Takeaways

Citadel’s most interesting trick is also its least glamorous one. It treats plain text as infrastructure. The repository’s own Markdown files hold memory, campaign state, task routing, and handoffs, which means both humans and agents can inspect and edit the same source of truth.

The Control Plane Is Just Markdown

That design choice changes the shape of the system. Instead of trusting a model to remember what happened three steps ago, Citadel writes the work down in repo-local artifacts. A task is not just a prompt. It is a record, a dependency, a status field, and a future recovery path.

This diagram shows why Citadel feels less like a chatbot wrapper and more like a process manager for agent work.

SystemState modelHuman editabilityCoordination styleSafety modelBest fit
CitadelRepo-local Markdown and Git stateHighRouted campaigns and work queuesHooks and policy enforcementAgent labor that must stay inspectable
Chat-first coding agentEphemeral conversation contextLowAd hoc promptsPrompting and best effortQuick, local edits
Generic multi-agent frameworkStructured in framework objectsMediumAgent graphs or role choreographyFramework-level guardrailsApp builders who want abstraction
Managed assistant APIVendor-hosted session stateLow to mediumTool-enabled conversationsPlatform controlsProduct teams that want hosted infra

Why Chat Fails, and Why Citadel Doesn’t Want Chat to Be the System

Most coding agents are strong at immediate execution and weak at continuity. They can edit files, call tools, and respond quickly, but they forget the reasoning behind the work as soon as the session shifts. Citadel is built around that failure mode.

The README frames the project as an agent orchestration harness for Claude Code, but the deeper idea is stricter: it wants to route any task through the right tool at the right scale. That means the system is not just a shell around a model. It is an attempt to make agent work persistent, steerable, and recoverable.

Citadel — Agent Orchestration Harness for Claude Code Run autonomous coding campaigns with Claude Code.

Project README, Repository documentation · SethGammon/Citadel README
A close editorial scene of a notebook page with erased notes, a half-finished terminal window, and a rubber stamp reading reset. The image contrasts brittle short-lived memory with a system that preserves decisions in durable files.
Citadel is designed around the failure of ephemeral context, not around the fantasy that a model will simply remember.

The Memory Layer: Blocks, Handoffs, and Verifiable Recall

Citadel’s memory system is not a vague vector store story. It is a set of explicit blocks in .planning/, organized so agents can inject prior decisions, failure patterns, and verification recipes into later work. That matters because engineering work is full of repeatable mistakes, and the project is trying to make those mistakes durable enough to learn from.

memory blocks:
- architecture-decisions
- failure-patterns
- verification-recipes

handoff:
- what changed
- why it changed
- what still needs verification

The important detail is the loop. Agents do not just write notes. They compile and lint their own operating memory, then hand that memory forward. In practice, that turns previous sessions into a reusable project asset instead of a pile of dead chat logs.

Routing Intent Before It Becomes Chaos

The `/do` command acts like a dispatcher. It classifies an intent, selects a skill or campaign, and sends the request through a defined protocol rather than letting the model wander. This is the difference between asking an assistant to help and asking a system to execute a workflow.

Chat promptRouted command
"Can you fix this?"Classify the task, select a skill, and enforce the protocol
Memory stays in the context windowMemory is written into .planning/ blocks
Tool use happens opportunisticallyTool use happens through a defined path
Success depends on the current conversationSuccess depends on the stored operating loop

Fleet Mode: Parallel Agents Without Merge-Conflict Theater

The fleet layer is where Citadel stops looking like a clever prompt wrapper. Work is serialized into a queue, dependencies are checked, and isolated Git worktrees keep parallel agents from colliding with one another. That is a real operating model, not a metaphor.

function isRunnable(task) {
  return task.status === 'pending' && task.deps.every(dep => completed.has(dep));
}

// If runnable, assign the task to an isolated worktree
// then write the handoff back into .planning/

That choice is quietly decisive. Git worktrees give each agent a sandbox with shared history and isolated state. The queue decides who runs next. Dependencies decide what can run at all. Citadel is managing labor, not just generating code.

DimensionSingle chat agentCitadel fleet
Parallel workManual and brittleScheduled through a queue
State isolationShared conversation contextSeparate Git worktrees
DependenciesImplicit in promptingExplicit in task metadata
RecoveryAsk the model againRead the repo-local handoff

Safety Is Not a Prompt. It’s a Hook.

Citadel’s hook system makes safety procedural. Before a tool action lands, the event is normalized and checked. If the policy says no, the action stops there. That is much stronger than asking a model to behave well in the abstract.

This is the right place for guardrails because it is the narrowest choke point in the loop. A planner can propose anything. A hook can still decide what may touch the filesystem, the shell, or the network.

The result is a system that treats agent behavior like software behavior. Inputs are normalized. Events are checked. Policy is enforced. Only then does execution proceed.

Where Citadel Fits in the Agent Landscape

Compared with LangGraph, AutoGen, CrewAI, AutoGPT, or managed assistant APIs, Citadel is more specific about where the truth lives. It is repo-native, Markdown-first, Git-native, and human-editable. That makes it less general than a framework, but more operational for teams that want the work itself to remain inspectable.

SystemOwns the work state?Human-readable state?Git-native?Primary strength
CitadelYesYesYesDurable agent operations inside a repo
LangGraphPartlyDepends on implementationNoStateful graphs and orchestration
AutoGenPartlyNoNoMulti-agent conversation patterns
CrewAIPartlyNoNoRole-based collaboration
Managed assistant APIsVendor-ownedLimitedNoHosted persistence and tooling

Citadel’s thesis is simple enough to fit on one line. If the agent is going to do real work, the work needs memory, routing, scheduling, and safety that survive the chat window. Markdown just happens to be the most surprising way to make that practical.