Quantisation advanced 8 min read 7 flashcards

KV Cache Quantisation

Why the KV cache quantises along different axes for keys and values, how that asymmetry follows from where the outliers live, and why the cache is the one tensor whose quantisation error compounds over a generation.

Take Llama-3.1-70B: 80 layers, 8 key-value heads after grouped-query attention, head dimension 128. Each token stores \(2 \times 8 \times 128 = 2{,}048\) values per layer, so \(163{,}840\) values across the model, which at fp16 is 320 KiB per token. One 32k-token conversation holds 10 GiB of KV cache. The weights, at 4 bits, fit in 35 GiB. Past a few concurrent long sessions the cache, not the model, is what you ran out of.

Quantising it is the obvious move, and the cache turns out to be the tensor whose quantisation rules least resemble everything else's.

Keys and values want different axes

The standard choice for activations is per-token (per-row) quantisation: each token's vector gets its own scale. For values that is right. For keys it is wrong, and the reason is distributional. Key activations carry strong, persistent per-channel outliers, specific dimensions whose magnitude is consistently far above the rest, while value activations have no comparable structure. Quantising keys per token forces each token's scale to cover its own outlier channels, destroying precision in every other dimension of that token.

KIVI's contribution was to take this seriously and quantise along the channel dimension for keys and the token dimension for values, which makes 2-bit KV quantisation work with no fine-tuning at all. Reported results: almost unchanged quality on Llama, Falcon and Mistral with 2.6x lower peak memory, up to 4x larger batch size, and 2.35x to 3.47x throughput (Liu et al., 2024, KIVI, ICML 2024, arXiv:2402.02750). The asymmetry is now the default assumption, and the Hugging Face Transformers KV cache quantisation path derives from this work.

Per-channel key quantisation has an awkward consequence. Channels run across tokens, so a channel's scale cannot be finalised until the tokens in its group exist, while decoding appends one token at a time. The standard fix is a split cache: a quantised bulk plus a small fp16 residual window of recent tokens, flushed into the quantised store once a full group accumulates.

KVQuant adds two refinements worth knowing. Keys are quantised before RoPE is applied, because the rotation mixes channels and smears the outlier structure that per-channel scaling depends on, and a small set of numerical outliers is held in a sparse high-precision side table. With those, the authors report serving Llama-7B at 1 million tokens of context on a single A100-80GB and 10 million on an 8-GPU system (Hooper et al., 2024, KVQuant, NeurIPS 2024, arXiv:2401.18079).

The production version is cruder and mostly fine

What serving stacks ship is 8-bit floating point. vLLM exposes kv_cache_dtype with E5M2 and E4M3FN variants; E4M3 has an extra mantissa bit and a dynamic range of only \(\pm 240\), so it needs an fp32 scale per tensor. The documented benefit is roughly double the KV cache allocation for the same memory, the documented recommendation is calibrated scales rather than defaults, and the documented limitation is that only per-tensor scalar scales are supported (vLLM, Quantized KV Cache). Per-tensor is too coarse to express the key/value asymmetry above, which is the gap between the research results and the deployed ones.

Why this tensor is different from weights

Weight quantisation error is fixed at quantisation time and identical for every request. KV quantisation error is generated at inference and then read back by every subsequent token. A token quantised at position 500 is attended to at position 501 and at position 20,000, so its error enters thousands of downstream attention computations. Error in the cache accumulates along the sequence in a way weight error does not, which is why 4-bit weights are routine while 4-bit KV needs the per-axis care above.

The attention-sink phenomenon makes it worse at exactly one place. Models place disproportionate attention mass on the first few tokens regardless of their content, and dropping them collapses quality in streaming settings (Xiao et al., 2023, Efficient Streaming Language Models with Attention Sinks, ICLR 2024, arXiv:2309.17453). Those same positions carry unusually large activations, so an aggressive uniform scheme quantises the most heavily attended entries in the cache the most badly. Keeping the first handful of tokens in fp16 costs almost nothing and removes the failure.

When it breaks

Long context is the test, and short benchmarks will not show it. A 2-bit KV cache can match fp16 perplexity on 2k-token sequences and lose retrieval accuracy at 64k, because the error that matters is the error at distance. Evaluate on needle-style retrieval and multi-turn tasks at the context length you serve, not on perplexity.

The memory saving is smaller than the bit ratio. Group-wise scales, the fp16 residual window and any sparse outlier table all take space, so a nominal 4x from fp16 to 4-bit lands closer to 3x once metadata is counted.

Dequantisation sits on the critical path of every attention call. Decode attention is bandwidth bound, so trading bytes for arithmetic usually wins, but only with a kernel that fuses dequantisation into attention. Without one the cache is unpacked to fp16 before use, and you have paid the accuracy cost for a memory saving plus a latency penalty.

Quantised entries are not comparable across prefix-cache boundaries. If scales are computed per request, two requests sharing a prefix produce different quantised bytes for the same tokens, which breaks the byte-level reuse that prefix-aware routing and KV cache reuse depends on. Shared-prefix reuse and per-request dynamic scales are in direct tension.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Liu et al., 2024, KIVI, ICML 2024, arXiv:2402.02750 arxiv.org
  2. Hooper et al., 2024, KVQuant, NeurIPS 2024, arXiv:2401.18079 arxiv.org
  3. vLLM, Quantized KV Cache docs.vllm.ai
  4. Xiao et al., 2023, Efficient Streaming Language Models with Attention Sinks, ICLR 2024, arXiv:2309.17453 arxiv.org
Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track