Quantisation intermediate 7 min read 6 flashcards

When Quantisation Does Not Make Inference Faster

Weight-only quantisation buys bandwidth, not arithmetic, so its speedup is a function of batch size and phase; the arithmetic that predicts where the gain disappears, and why a kernel can be slower than fp16.

A team quantises a 70B model to 4-bit weights, measures single-request latency, and sees close to the 4x they hoped for. They deploy it behind a server handling 64 concurrent requests and the throughput gain is a few percent. Nothing is misconfigured. Weight-only quantisation buys one specific resource, and at batch 64 that resource was no longer the constraint.

The arithmetic that decides it

Decoding one token with batch size \(B\) reads every weight once and performs \(O(B)\) multiply-accumulates per weight. Arithmetic intensity therefore scales with \(B\), while the weight bytes read do not. Compare that against the machine: modern accelerators sit at a FLOP-to-byte ratio of roughly 100 to 200, so as long as fewer than about 25 to 50 tensor-core MACs are performed per 4-bit weight, the matmul is still bandwidth bound and shrinking the weight by 4x can still deliver close to a 4x speedup. This is the reasoning behind Marlin, a FP16xINT4 kernel that holds near-ideal 4x up to 16 to 32 tokens per batch where earlier kernels managed it only at 1 or 2, continues to accelerate at 64 to 128 with a gradually shrinking margin, and delivers up to 2.8x end to end inside vLLM (Frantar et al., 2024, MARLIN, PPoPP 2025, arXiv:2408.11743).

Read the same inequality from the other direction and it tells you where the gain goes. Past roughly 50 MACs per weight the matmul is compute bound, the weight read is amortised across the batch, and a 4-bit weight that must be dequantised to fp16 before entering the tensor cores does exactly as much arithmetic as the fp16 weight did, plus the unpacking. Prefill is the same story without the batch: processing a 4,000-token prompt performs thousands of MACs per weight on the first pass, so it is compute bound from the start and weight-only quantisation does approximately nothing for time to first token.

The empirical conclusion from the largest public sweep matches the arithmetic. Across more than 500,000 evaluations on the Llama-3.1 family, W4A16 is the most cost-efficient scheme for synchronous, latency-bound deployments, while W8A8 dominates under asynchronous continuous batching (Kurtić et al., 2024, "Give Me BF16 or Give Me Death"?, ACL 2025, arXiv:2411.02355). W8A8 wins at high concurrency not because 8 bits compresses better than 4, but because quantising activations too lets the matmul itself run on int8 or fp8 tensor cores, which is the resource that binds once arithmetic dominates.

What the scheme names are telling you

The WxAy notation is a deployment decision in disguise. W4A16 compresses weights and leaves activations in fp16, so it attacks memory traffic and capacity: maximum compression, useful in low-QPS serving, and supported on any GPU. W8A8-INT compresses both, needs channel-wise weights and dynamic per-token activations, and is the recommended path on NVIDIA GPUs below compute capability 8.9 where fp8 has no hardware support. W8A8-FP8 with dynamic per-token activation scales is described as the most performant option on hardware that has it and needs no calibration set at all (vLLM LLM Compressor, Compression Schemes).

So the selection rule is short. Interactive single-stream or memory-constrained: W4A16. High-concurrency server or offline batch: W8A8, fp8 on Hopper and later, int8 on Ampere and older.

When it breaks

No kernel, no speedup. A quantisation format is a promise about bytes; a kernel is what turns it into time. If the serving stack lacks a fused kernel for your exact combination of bit width, group size, packing layout and GPU architecture, it falls back to dequantising to fp16 and running the standard matmul. You then pay the memory saving's full accuracy cost, plus unpacking overhead, for a model that is slower than fp16. This is the single most common disappointment in quantisation work and it is a software-support question, not a numerics one.

Small group sizes erode the ratio. At group size 128 with an fp16 scale and a 4-bit zero point the metadata is about 0.16 bits per weight; at group size 32 it is about 0.6 bits, so a nominal 4-bit model stores 4.6 and the bandwidth saving falls short of the advertised 4x. See quantisation grids, scale and zero point.

The bottleneck may not be the weights at all. At long context the KV cache dominates both memory and the bandwidth budget, and compressing weights alone leaves it untouched. A 4-bit 70B model serving 32k conversations is limited by KV-cache memory and bandwidth, so the next useful quantisation is of the cache, not the weights.

Benchmarking at batch 1 flatters the result. A latency measurement on a single request is a measurement in the regime where weight-only quantisation looks best. Measure at the concurrency you will actually run, and measure goodput under your SLO rather than raw throughput, or the number you report will not be the number you get.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Frantar et al., 2024, MARLIN, PPoPP 2025, arXiv:2408.11743 arxiv.org
  2. Kurtić et al., 2024, "Give Me BF16 or Give Me Death"?, ACL 2025, arXiv:2411.02355 arxiv.org
  3. vLLM LLM Compressor, Compression Schemes docs.vllm.ai
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track