The Training-Free Context Engine: Inside JusperLee/BEAVER

How structure-aware "pages" and dual-path pooling are replacing token-level pruning to solve the LLM lost-in-the-middle problem.

7 min read • View on GitHub • More from JusperLee

A chaotic industrial paper shredder spitting out fragmented text puzzle pieces next to a precise mechanical librarian arm extracting whole pages from a ledger.
Token-level pruning shreds context, while BEAVER extracts intact structural pages.
Key Takeaways

Stop Mutilating the Grammar

Most prompt compression tools treat language like a bag of words. Methods like LLMLingua aggressively prune low-probability tokens until the prompt resembles a grammatically shattered telegram. This saves space but destroys discourse flow, causing the LLM to hallucinate or lose its train of thought when reasoning over long documents.

BEAVER introduces a fundamentally different paradigm. Instead of deleting isolated words, it chunks text into fixed-size "pages" using a SegmentPager. To ensure the model never sees a fragmented thought, BEAVER employs a "Sentence Smoother" (do_align) that uses sentence boundary IDs to guarantee cuts only happen at punctuation marks, never mid-sentence.

FeatureToken-Level Pruning (e.g., LLMLingua)BEAVER Page Selection
GranularityIndividual tokensFixed-size pages
GrammarFrequently brokenPreserved via Sentence Smoother
TrainingRequires specific fine-tuningZero-shot (uses hidden states)
ContextLoses local flowPreserved via Flow prior

The Anchor, Flow, and Flash Priors

The intelligence of BEAVER resides in QueryPlanner.py, which acts as a deterministic librarian. It uses three heuristic rules to decide which pages to keep from a massive document.

The three priors interact to protect structural integrity while extracting semantic relevance.

First, the Anchor prior mathematically protects the introduction and conclusion, acknowledging that titles and summaries are critical for context. Second, the Flash prior grabs the most semantically and lexically relevant pages based on the query. Finally, the Flow prior grabs the adjacent pages surrounding the Flash hits, ensuring the LLM isn't subjected to jarring logical leaps.

The Hidden State Hack

The elegance of BEAVER is that it requires absolutely zero fine-tuning. It achieves high fidelity by running a forward pass and using the LLM's own internal hidden states as a guide, implemented in PageEncoder.py.

It converts token-level hidden states into page-level embeddings using a "dual-path pooling" strategy—calculating both a weighted mean and a max-pooled representation. Furthermore, it utilizes In-Context Inverse Token Frequency (ITF) weighting. This naturally down-weights common stop-words within the page embedding without actually deleting them from the text.

Built for the Agentic Edge

BEAVER isn't just an academic proof of concept; it's built for production. The repository includes a HSPBlackBoxWrapper and a FastAPI server that automates the detection of special tokens depending on the specific model architecture (like Qwen or Llama).

Given any prefix-closed semantic constraint, BEAVER systematically explores the generation space using novel token trie and frontier data structures, maintaining provably sound bounds at every iteration.

Tarun Suresh, Nalin Wadhwa, et al., Authors · Paper page - BEAVER: An Efficient Deterministic LLM Verifier

By plugging directly into tools like the OpenClaw AI assistant, BEAVER allows smaller, faster models to act as intelligent filters for massive documentation folders, effectively granting an infinite context window limited only by the quality of the extraction.