codex-cli-hardening-cheatsheet: Building a Sandbox for the Future of Agency
Moving beyond "be careful" prompts to a zero-trust architecture for your terminal.
- Instruction-based safety is a fallacy because prompt injection can bypass any verbal command.
- The repository shifts security from the prompt layer to a system-level "Deny-by-Default" architecture.
- A central TOML configuration enforces physical constraints like filesystem isolation and network egress control.
- The framework provides lightweight, copy-pasteable security templates designed for immediate developer use.
The Prompt Injection Illusion
For many developers, securing an AI coding agent means writing a sternly worded AGENTS.md file. We tell the model, "Don't delete my production database," and hope it listens. This is the fallacy of instruction-based safety.
The reality is that instructions are just prompts. If an LLM can be convinced to ignore its system prompt—a phenomenon known as prompt injection—your terminal is wide open. The okdt/codex-cli-hardening-cheatsheet repository argues that treating AI safety as a prompting problem is a fundamental mistake. It must be treated as a systems engineering problem.
The Four Pillars of the Agentic Firewall
The cheatsheet breaks down terminal security into four distinct architectural pillars, implemented via a config.toml file. This "Deny-by-Default" approach ensures that even a compromised model cannot escape its designated boundaries.
- Sandbox Isolation: Utilizing OS-level primitives to restrict filesystem access to the current workspace.
- Approval Policies: Enforcing a "Human-in-the-Loop" threshold for destructive commands.
- Network Egress Control: Preventing the agent from exfiltrating data or reaching out to malicious domains.
- Ephemeral History: Mitigating the risk of local credential leakage through session logs.
By enforcing these constraints at the runtime level rather than the prompt level, the architecture shifts from asking the AI to be good to making it impossible for the AI to be bad.
Hardening for a Model That Doesn't Exist
Perhaps the most intriguing aspect of the configuration templates provided in the repository is the use of the placeholder model = "gpt-5.4". This signals a forward-looking philosophy.
The repository isn't just protecting against today's prompt injections; it's building a sandbox for a level of agency we haven't even seen yet. It acknowledges that as models become more capable, the risk of "Excessive Agency"—where the AI takes correct but unintended actions—grows exponentially.
Lightweight vs. Fortress
The competitive landscape for AI agent security ranges from simple best practices to complex, enterprise-grade sandboxing. The okdt cheatsheet occupies a pragmatic middle ground.
| Feature | okdt/codex-cli-hardening | Trail of Bits Claude Config | Official Anthropic Docs |
|---|---|---|---|
| Approach | Config-level (TOML) | OS-level (Seatbelt/Bubblewrap) | Settings-level (JSON) |
| Complexity | Low (Drop-in Templates) | High (Custom Hooks) | Low (Native Settings) |
| Target Audience | Indie Devs / Small Teams | Enterprise Security Teams | General Users |
While projects like Trail of Bits offer more robust OS-level isolation, they require significant setup. The okdt approach provides immediate, copy-pasteable security wins for developers who need protection without the overhead of a full security audit.