The Token Firewall: Disciplining Claude Code with claude-context-optimizer

How a silent middleware layer uses "displacement logic" and "waste metrics" to stop AI agents from drowning in their own context.

7 min read • View on GitHub • More from egorfedorov

A mechanical arm catching papers falling from a massive filing cabinet while a human hand holds a stopwatch and a coin purse, illustrating the finite nature of LLM context windows.
Context is a finite, billable resource. Treating an LLM's memory like an infinite bucket leads to massive API waste.
Portrait of Egor Fedorov

The average Claude Code session wastes 30-50% of tokens on files that are read but never actually used. Every `Read` call consumes context — whether the file was relevant or not.

Key Takeaways

The High Cost of Curiosity

We treat Large Language Model context windows like magical, infinite buckets. The reality is much closer to a leaky, expensive pipe. When an AI agent assists with a codebase, it often exhibits "doom-scrolling" behavior. It reads configuration files, lockfiles, and boilerplate repeatedly without modifying a single line of code.

This passive reading incurs a massive penalty in API costs and degrades the model's ability to focus on relevant logic. The claude-context-optimizer project recognizes this architectural failure. It introduces a "Token ROI" metric that treats LLM attention as a finite engineering resource.

The Contextual Firewall

Most AI dev tools operate passively. They format prompts or pass data along. The optimizer takes an active, interventionist approach. It utilizes a PreToolUse hook within the Claude Code CLI to intercept the agent's intent before it hits the Anthropic API.

When the agent attempts to read a file, the context-shield.js script evaluates the historical return on investment for that specific file path. If the file is known to be a "high-waste" target, the tool actively blocks the command. It returns a cached response or a warning, effectively parenting the AI agent and preventing it from making expensive mistakes.

A flow chart illustrating the lifecycle of a Read command intercepted by the optimizer. The process starts with a 'User' node pointing to a 'Claude Code' node. From there

Managing the "Displacement" Window

Caching file reads is a standard optimization. The optimizer goes further by implementing a sophisticated staleness detection algorithm based on token displacement. It calculates how many tokens of other files have been ingested since the agent last viewed a specific file.

If 20,000 tokens have passed, the script assumes the original file has been pushed out of the LLM's active active window. This Least Recently Used (LRU) approach models the AI's memory decay accurately. It allows a re-read only when the agent has genuinely forgotten the context.

A vertical stack of blocks representing files in an LLM context window. The interaction shows a 'Read New File' action dropping a new block on top. As this happens

Skeletal Navigation

Blocking a read request completely can leave the AI blind. To solve this, the optimizer employs a file-digest.js module. Instead of feeding the agent a 1,000-line file, it uses regex landmarks to extract a structural map.

This skeletal map includes function names, exports, and critical framework tags. The agent learns that a specific update function exists on line 42 without needing to ingest the entire logic block. This reduces massive files to their bare navigational essentials.

Metric Standard Session Optimized Session
Context Strategy Raw file ingestion Skeletal digests
Redundant Reads Allowed unconditionally Blocked via LRU cache
Token Waste High Minimized

The Efficiency Grade

The optimizer tracks every file interaction to compute a usefulness score. Files that are edited gain points. Files that are passively read multiple times lose points. This data powers a session grading system.

Developers receive an efficiency grade from S to F at the end of a session. A visual heatmap identifies the exact files that burned the most tokens without contributing to the final output. This feedback loop trains developers to structure their projects and their prompts for maximum AI efficiency.


Sources: