codex-cli-hardening-cheatsheet: Building a Sandbox for the Future of Agency

Moving beyond "be careful" prompts to a zero-trust architecture for your terminal.

• View on GitHub • More from okdt

A massive, ornate vault door made of translucent glass. Inside, a small glowing spark (the AI) is trying to reach out, but its "hands" are stopped by physical gears labeled "TOML" and "Sandbox."
Physical constraints over verbal requests: The Agentic Firewall.

Key Takeaways

The Prompt Injection Illusion

For many developers, securing an AI coding agent means writing a sternly worded AGENTS.md file. We tell the model, "Don't delete my production database," and hope it listens. This is the fallacy of instruction-based safety.

The reality is that instructions are just prompts. If an LLM can be convinced to ignore its system prompt—a phenomenon known as prompt injection—your terminal is wide open. The okdt/codex-cli-hardening-cheatsheet repository argues that treating AI safety as a prompting problem is a fundamental mistake. It must be treated as a systems engineering problem.

The Four Pillars of the Agentic Firewall

The cheatsheet breaks down terminal security into four distinct architectural pillars, implemented via a config.toml file. This "Deny-by-Default" approach ensures that even a compromised model cannot escape its designated boundaries.

The Request Lifecycle: How a single command from the AI must pass through four distinct "filters" before hitting the OS.

By enforcing these constraints at the runtime level rather than the prompt level, the architecture shifts from asking the AI to be good to making it impossible for the AI to be bad.

Hardening for a Model That Doesn't Exist

Perhaps the most intriguing aspect of the configuration templates provided in the repository is the use of the placeholder model = "gpt-5.4". This signals a forward-looking philosophy.

The repository isn't just protecting against today's prompt injections; it's building a sandbox for a level of agency we haven't even seen yet. It acknowledges that as models become more capable, the risk of "Excessive Agency"—where the AI takes correct but unintended actions—grows exponentially.

A close-up of a "Human-in-the-Loop" gate. A mechanical arm (the Agent) is holding a "Delete" command, but the gate only opens when a physical human key is turned in a separate lock.
The `on-request` approval policy ensures human oversight for critical actions.

Lightweight vs. Fortress

The competitive landscape for AI agent security ranges from simple best practices to complex, enterprise-grade sandboxing. The okdt cheatsheet occupies a pragmatic middle ground.

Featureokdt/codex-cli-hardeningTrail of Bits Claude ConfigOfficial Anthropic Docs
ApproachConfig-level (TOML)OS-level (Seatbelt/Bubblewrap)Settings-level (JSON)
ComplexityLow (Drop-in Templates)High (Custom Hooks)Low (Native Settings)
Target AudienceIndie Devs / Small TeamsEnterprise Security TeamsGeneral Users

While projects like Trail of Bits offer more robust OS-level isolation, they require significant setup. The okdt approach provides immediate, copy-pasteable security wins for developers who need protection without the overhead of a full security audit.