The Token Firewall: Disciplining Claude Code with claude-context-optimizer
How a silent middleware layer uses "displacement logic" and "waste metrics" to stop AI agents from drowning in their own context.
The average Claude Code session wastes 30-50% of tokens on files that are read but never actually used. Every `Read` call consumes context — whether the file was relevant or not.
- The optimizer intercepts tool calls to block redundant file reads that account for up to 50% of token waste.
- A displacement algorithm tracks when specific files have likely been pushed out of the model's active memory.
- Skeletal digests provide the agent with structural maps of files to prevent full ingestion of large boilerplate blocks.
- The system assigns efficiency grades and heatmaps to identify which files are consuming the most budget without being edited.
The High Cost of Curiosity
We treat Large Language Model context windows like magical, infinite buckets. The reality is much closer to a leaky, expensive pipe. When an AI agent assists with a codebase, it often exhibits "doom-scrolling" behavior. It reads configuration files, lockfiles, and boilerplate repeatedly without modifying a single line of code.
This passive reading incurs a massive penalty in API costs and degrades the model's ability to focus on relevant logic. The claude-context-optimizer project recognizes this architectural failure. It introduces a "Token ROI" metric that treats LLM attention as a finite engineering resource.
The Contextual Firewall
Most AI dev tools operate passively. They format prompts or pass data along. The optimizer takes an active, interventionist approach. It utilizes a PreToolUse hook within the Claude Code CLI to intercept the agent's intent before it hits the Anthropic API.
When the agent attempts to read a file, the context-shield.js script evaluates the historical return on investment for that specific file path. If the file is known to be a "high-waste" target, the tool actively blocks the command. It returns a cached response or a warning, effectively parenting the AI agent and preventing it from making expensive mistakes.
Managing the "Displacement" Window
Caching file reads is a standard optimization. The optimizer goes further by implementing a sophisticated staleness detection algorithm based on token displacement. It calculates how many tokens of other files have been ingested since the agent last viewed a specific file.
If 20,000 tokens have passed, the script assumes the original file has been pushed out of the LLM's active active window. This Least Recently Used (LRU) approach models the AI's memory decay accurately. It allows a re-read only when the agent has genuinely forgotten the context.
Skeletal Navigation
Blocking a read request completely can leave the AI blind. To solve this, the optimizer employs a file-digest.js module. Instead of feeding the agent a 1,000-line file, it uses regex landmarks to extract a structural map.
This skeletal map includes function names, exports, and critical framework tags. The agent learns that a specific update function exists on line 42 without needing to ingest the entire logic block. This reduces massive files to their bare navigational essentials.
| Metric | Standard Session | Optimized Session |
|---|---|---|
| Context Strategy | Raw file ingestion | Skeletal digests |
| Redundant Reads | Allowed unconditionally | Blocked via LRU cache |
| Token Waste | High | Minimized |
The Efficiency Grade
The optimizer tracks every file interaction to compute a usefulness score. Files that are edited gain points. Files that are passively read multiple times lose points. This data powers a session grading system.
Developers receive an efficiency grade from S to F at the end of a session. A visual heatmap identifies the exact files that burned the most tokens without contributing to the final output. This feedback loop trains developers to structure their projects and their prompts for maximum AI efficiency.