The Death of Symmetric Memory: Inside TheTom/turboquant_plus
By proving that LLM Value caches can be aggressively compressed without lobotomizing the model, this experimental integration is rewriting the rules of local inference.
- Asymmetric quantization proves that LLM attention mechanisms are highly sensitive to Key degradation but almost completely blind to Value degradation.
- TurboQuant+ compresses the Value cache down to 2 bits while keeping Keys at 8 bits, saving massive amounts of VRAM.
- The architecture relies on Fast Walsh-Hadamard Transforms and PolarQuant to minimize residual error without bottlenecking the GPU.
- Community contributions like RotorQuant are already iterating on the baseline math to reduce parameter counts by 88x.
The "V is Free" Anomaly
For years, the local AI community has treated the Key and Value (KV) caches as a matched set. If you quantized the Keys to 4 bits, you did the exact same to the Values. This symmetric approach wasted gigabytes of memory.
TheTom/turboquant_plus challenges this core assumption. The experimental integration reveals a counterintuitive truth about how large language models pay attention. They are highly sensitive to Key degradation, but they are almost completely blind to Value degradation. By splitting the precision, you can crush the Value cache down to 2 bits.
V compression is free. Compressing the value cache (even down to 2 bits) has zero measurable effect on attention quality when key precision is maintained.
The Mathematical Blender
Achieving this extreme compression without destroying the model requires the TurboQuant algorithm, originally slated for ICLR 2026. The system uses a two-stage pipeline.
The first stage, PolarQuant, compresses the bulk of the data for minimum error. The second stage uses Quantized Johnson-Lindenstrauss (QJL) to eliminate residual bias. This split ensures that inner products required for attention are perfectly preserved.
Squeezing Silicon
Theoretical math is useless if it bottlenecks the GPU. The repository implements hardware-specific optimizations to translate Python research into usable Metal and CUDA kernels.
This includes a 4-mag look-up table detection mechanism specifically tuned for Apple M-series chips. The project also experiments with Sparse V dequantization, skipping low-weight Value positions entirely during inference. By ignoring 80 percent of the Value cache, the attention mechanism becomes a sparse operation, speeding up generation significantly.
| Metric | Symmetric (q8_0 K + q8_0 V) | Asymmetric (q8_0 K + turbo2 V) |
|---|---|---|
| VRAM Usage (128K Context) | Massive | Reduced by up to 5x |
| Attention Quality | Baseline | Virtually Identical |
| Dequantization Speed | Standard | Accelerated via Sparse V |
The Open Source Mutation
Because the project serves as a collaborative workspace for upstreaming to llama.cpp, community members are rapidly iterating on the foundational math. The open-source cycle is actively outpacing traditional academic publishing.
A prime example is the RotorQuant pull request, which swaps random orthogonal matrices for Clifford rotors. This single change drastically reduces the parameter overhead required to make the math work.
RotorQuant replaces TurboQuant's d×d random orthogonal matrix with Clifford rotors in Cl(3,0), achieving 88x fewer parameters with matching retrieval accuracy.