← all repositories

SqueezeAILab/KVQuant

A quantization methodology for KV cache compression enabling 10M context length LLM inference on a single A100 GPU.

KVQuant
Not currently ranked — collecting fresh signals.
star history

KVQuant addresses the memory bottleneck in long-context LLM inference by quantizing the KV cache to low precision. It achieves high accuracy by exploiting consistent patterns in cached KV values, including per-channel pre-RoPE key quantization to handle outlier channels, non-uniform quantization for asymmetric activations, and dense-and-sparse quantization to mitigate numerical outliers. This enables serving LLaMA-7B with 1M context on a single A100-80GB or 10M context on an 8-GPU system.

Frequently asked

What is SqueezeAILab/KVQuant?
A quantization methodology for KV cache compression enabling 10M context length LLM inference on a single A100 GPU.
Is KVQuant open source?
Yes — SqueezeAILab/KVQuant is an open-source project tracked on heatdrop.
What language is KVQuant written in?
SqueezeAILab/KVQuant is primarily written in Python.
How popular is KVQuant?
SqueezeAILab/KVQuant has 429 stars on GitHub.
Where can I find KVQuant?
SqueezeAILab/KVQuant is on GitHub at https://github.com/SqueezeAILab/KVQuant.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.