First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.
81 stars
Python
Your first custom repo explanation is free. Reading existing public explainers always stays free.
This will take 10-20 minutes. You can close the tab and come back later.