tonbistudio/turboquant-pytorch
quality grade C, 52 out of 100From-scratch PyTorch implementation of Google's TurboQuant (ICLR 2026) for LLM KV cache compression. 5x compression at 3-bit with 99.5% attention fidelity.
- stars
- 1.0k
- stars gained this week
- —this week