The Death of Symmetric Memory: Inside TheTom/turboquant_plus

By proving that LLM Value caches can be aggressively compressed without lobotomizing the model, this experimental integration is rewriting the rules of local inference.

8 min read • View on GitHub • More from TheTom

Two bank vaults sitting side by side. The left vault is heavily armored with multiple thick steel locking mechanisms and laser sensors. The right vault is a simple, lightweight lockbox made of thin wire mesh. Both contain glowing geometric tokens. This illustrates the asymmetric KV cache concept.
While the Key cache requires intense protection and high precision, the Value cache can be stored in a lightweight 2-bit container with virtually no quality loss.
Key Takeaways

The "V is Free" Anomaly

For years, the local AI community has treated the Key and Value (KV) caches as a matched set. If you quantized the Keys to 4 bits, you did the exact same to the Values. This symmetric approach wasted gigabytes of memory.

TheTom/turboquant_plus challenges this core assumption. The experimental integration reveals a counterintuitive truth about how large language models pay attention. They are highly sensitive to Key degradation, but they are almost completely blind to Value degradation. By splitting the precision, you can crush the Value cache down to 2 bits.

V compression is free. Compressing the value cache (even down to 2 bits) has zero measurable effect on attention quality when key precision is maintained.

The Mathematical Blender

Achieving this extreme compression without destroying the model requires the TurboQuant algorithm, originally slated for ICLR 2026. The system uses a two-stage pipeline.

The first stage, PolarQuant, compresses the bulk of the data for minimum error. The second stage uses Quantized Johnson-Lindenstrauss (QJL) to eliminate residual bias. This split ensures that inner products required for attention are perfectly preserved.

The TurboQuant pipeline relies on a Fast Walsh-Hadamard Transform to prepare vectors before asymmetric compression.

Squeezing Silicon

Theoretical math is useless if it bottlenecks the GPU. The repository implements hardware-specific optimizations to translate Python research into usable Metal and CUDA kernels.

This includes a 4-mag look-up table detection mechanism specifically tuned for Apple M-series chips. The project also experiments with Sparse V dequantization, skipping low-weight Value positions entirely during inference. By ignoring 80 percent of the Value cache, the attention mechanism becomes a sparse operation, speeding up generation significantly.

A mechanical sorting chute dropping metallic spheres (data). A finely woven wire mesh sieve intercepts the flow. Most spheres fall straight through the gaps, while only a few heavy, polished spheres are caught and funneled into a brass processing tube.
Sparse V dequantization acts as a sieve, only processing the heavy, high-weight Value tokens and letting the rest fall away.
MetricSymmetric (q8_0 K + q8_0 V)Asymmetric (q8_0 K + turbo2 V)
VRAM Usage (128K Context)MassiveReduced by up to 5x
Attention QualityBaselineVirtually Identical
Dequantization SpeedStandardAccelerated via Sparse V

The Open Source Mutation

Because the project serves as a collaborative workspace for upstreaming to llama.cpp, community members are rapidly iterating on the foundational math. The open-source cycle is actively outpacing traditional academic publishing.

A prime example is the RotorQuant pull request, which swaps random orthogonal matrices for Clifford rotors. This single change drastically reduces the parameter overhead required to make the math work.

RotorQuant replaces TurboQuant's d×d random orthogonal matrix with Clifford rotors in Cl(3,0), achieving 88x fewer parameters with matching retrieval accuracy.

John D. Pope, Contributor · Pull Request #34 · TheTom/turboquant_plus