Quantisation intermediate 7 min read 6 flashcards

Calibration Data for Post-Training Quantisation

The few hundred kilobytes of text that decide how a 70B model rounds, why the default recipe of 128 random web sequences is a convention rather than a result, and what the evidence says about choosing better.

A 4-bit quantisation run on a 70B model touches every one of its 70 billion weights, and the only input that distinguishes a good run from a bad one is a small sample of text, usually a few hundred kilobytes. That sample is the calibration set. It is the least examined hyperparameter in the whole pipeline, it is rarely reported in model cards, and it moves downstream accuracy by more than the choice of algorithm does in some regimes.

What the calibration set is actually for

Weight-only post-training quantisation does not minimise the error in the weights. It minimises the error in the layer's output, which means it needs to know what inputs the layer sees. GPTQ formulates this as a per-layer reconstruction problem: given inputs \(X\), find quantised weights \(\hat{W}\) minimising \(\lVert XW - X\hat{W} \rVert_2^2\), solved greedily with error compensation using second-order information from the Hessian \(H = 2XX^{\top}\) (Frantar et al., 2022, GPTQ, ICLR 2023, arXiv:2210.17323). The Hessian is an empirical estimate built entirely from the calibration activations. Change the text and you change the Hessian, which changes which weights get the error budget.

AWQ uses the calibration set differently and more cheaply. It reads only the per-channel magnitude of activations, uses that to identify roughly 0.1 to 1 percent of salient weight channels, and solves for a per-channel scaling that protects them. Because there is no backpropagation and no reconstruction, the authors argue it cannot overfit the calibration distribution the way reconstruction-based methods can, and they report retained accuracy on coding and maths benchmarks and on multimodal models (Lin et al., 2023, AWQ, MLSys 2024 best paper, arXiv:2306.00978). The distinction is worth internalising: a statistics-only method is robust to calibration choice in a way a reconstruction method is not, and that is the main reason to prefer one over the other on a model whose training mixture you cannot sample.

Where the default recipe came from

The near-universal setting is 128 sequences of 2,048 tokens sampled from a web crawl, usually C4, which is about a quarter of a million tokens in total. It spread by inheritance from the GPTQ experiments rather than from any study showing it to be optimal, and later work describes it as exactly that: the practice followed from Frantar et al. rather than a tuned choice (Williams and Aletras, 2024, On the Impact of Calibration Data in Post-training Quantization and Pruning, ACL 2024).

Two sanity checks fall out of the arithmetic. 128 sequences of 2,048 tokens gives roughly \(2.6 \times 10^5\) activation rows per layer against a hidden dimension of 8,192 on a 70B model, so the Hessian is estimated from about 32 rows per dimension: enough to be well conditioned, nowhere near enough to be confident. And a 2,048-token sequence says nothing about how the layer behaves at position 100,000, which is why quantising a long-context model on short calibration text loses long-context behaviour while every short benchmark stays flat.

What the evidence says

The first systematic study across methods, datasets, tasks and models found that downstream task performance varies substantially with the calibration set, contradicting the earlier consensus that these methods were largely insensitive to it (Williams and Aletras, 2024). The parallel result on the pruning side is sharper still: at high sparsity the calibration data can matter more than which pruning strategy you picked, a modest quantity is sufficient, and data resembling the model's pre-training distribution produces better results. Since nobody outside the lab can sample a frontier model's pre-training mixture, the authors propose synthesising calibration text from the model itself, reporting gains of up to 2.68 percent over Wanda, DSnoT and OWL (Ji et al., 2024, Beware of Calibration Data for Pruning Large Language Models, ICLR 2025, arXiv:2410.17711). That result is measured on pruning, not quantisation, and both procedures consume calibration activations the same way, so it is suggestive here rather than conclusive.

The practical reading: match the calibration distribution to the deployment distribution where you know it, prefer self-generated text to an arbitrary web crawl where you do not, include long sequences if you serve long context, and ship the calibration set as a configuration item alongside the checkpoint.

When it breaks

Overfitting is invisible to the metric you are watching. A reconstruction-based run calibrated on C4 will report an excellent C4 perplexity, because that is approximately what it optimised. The loss shows up on a different domain, which is why a quantised model should be evaluated on the task class it will serve rather than on the corpus it was calibrated on. See evaluating quantised models.

Instruction-tuned models calibrated on raw web text lose formatting behaviour first. Chat templates, tool-call syntax and refusal behaviour are concentrated in a narrow slice of the activation distribution that a web crawl never exercises. Perplexity will not detect it; a few dozen formatted prompts will.

Too few sequences makes the result seed-dependent. With 32 or fewer sequences the Hessian estimate is noisy enough that two runs differing only in sample draw produce measurably different models. If a quantisation result cannot survive a change of random seed, it is not a result.

Calibration data is a leak surface. Production traffic is the most representative calibration set available and the most sensitive. A sample of real user prompts committed next to the checkpoint is a data-handling incident waiting to be noticed.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Frantar et al., 2022, GPTQ, ICLR 2023, arXiv:2210.17323 arxiv.org
  2. Lin et al., 2023, AWQ, MLSys 2024 best paper, arXiv:2306.00978 arxiv.org
  3. Williams and Aletras, 2024, On the Impact of Calibration Data in Post-training Quantization and Pruning, ACL 2024 aclanthology.org
  4. Ji et al., 2024, Beware of Calibration Data for Pruning Large Language Models, ICLR 2025, arXiv:2410.17711 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track