Quantisation and Reasoning Models
Reasoning models spend thousands of tokens per answer, which is exactly the regime where small per-token errors have room to compound; what the measurements show, where they disagree, and what to evaluate.
A standard chat model emits a few hundred tokens and the quality of the answer is roughly the quality of those tokens. A reasoning model emits several thousand, and the final answer depends on a chain in which an early arithmetic slip is never recovered. Quantisation introduces exactly the kind of small, unbiased, per-token perturbation that such a chain is built to amplify. The intuition says reasoning models should be the worst case for quantisation. The measurements say something more interesting.
What the measurements show
The first systematic study covered the distilled and RL-trained reasoning families across maths (AIME, MATH-500), science (GPQA) and code (LiveCodeBench), with weight-only quantisation via GPTQ and AWQ at 3 and 4 bits using asymmetric group-128 schemes, and KV cache quantisation via KVQuant and QuaRot at 3 and 4 bits. Its headline finding is that quantisation is not inherently destructive here: 8-bit weight-activation and 4-bit weight-only schemes stay close to lossless, and the dominant factors in how much is lost are model size, model origin and task difficulty rather than the quantisation algorithm (Liu et al., 2025, Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models, COLM 2025, arXiv:2504.04823).
Each of those three factors is worth taking separately.
Size. A larger model has more redundancy to absorb the same rounding error, so a 4-bit result on a 32B reasoning model says little about the same recipe on a 7B one. Scaling the model, or giving it more reasoning steps, recovers part of the loss.
Origin. Models distilled from a stronger reasoner and models trained to reason with reinforcement learning do not degrade identically, which matters because the two families are often treated as interchangeable deployment targets.
Difficulty. The loss concentrates in the hard problems. Accuracy on easy grade-school arithmetic can be untouched while competition-level maths drops several times further, which is the single most important fact for anyone choosing an evaluation set. Averaging over a benchmark that is mostly easy items hides the damage almost perfectly.
Where the field disagrees
The same study reports that quantised reasoning models do not produce longer outputs, which is a direct test of the compounding-error intuition and does not support it. Later work reports the opposite: under aggressive post-training quantisation, accuracy falls and chain-of-thought length rises, with up to 52 percent of failures being cases where the model reached the correct answer in an intermediate step and did not commit to it as the final answer (Quantized Reasoning Models Think They Need to Think Longer, but They Do Not, arXiv:2606.00206).
These are not obviously reconcilable, and the honest summary as of late 2026 is that output-length behaviour under quantisation depends on how aggressive the scheme is and on the model family, with the near-lossless regime showing no length inflation and the aggressive regime showing it. If the second result holds up, the cost of over-quantising a reasoning model is worse than an accuracy drop: it is an accuracy drop paid for with more tokens, which inverts the economics that motivated quantising in the first place.
What to evaluate, concretely
Perplexity is useless here and aggregate benchmark accuracy is close to it. Three measurements carry the signal:
- Accuracy stratified by difficulty, with the hard stratum reported separately. A single average across a mixed benchmark is the metric this failure mode hides behind.
- Mean and tail output length against the fp16 baseline at matched accuracy. A quantised model that matches accuracy using 30 percent more tokens has not matched the baseline on cost.
- Pass@1 with multiple samples per item, because reasoning benchmarks are small and high variance; AIME-style sets have tens of problems, so one or two items of movement is noise. See eval error bars and statistics.
When it breaks
The KV cache is the riskier half. A reasoning trace is a long self-attended sequence, so cache quantisation error at step 200 is read back at every one of the next several thousand steps. Weight-only 4-bit and KV 4-bit are not comparable risks on this workload even though both are called 4-bit. See KV cache quantisation.
Budget interactions make the comparison unfair. If a quantised model is allowed to generate until it stops and the baseline is capped, the comparison measures two different compute budgets. Fix the token budget, or report accuracy as a function of it, since inference-time compute and quantisation trade against each other directly (inference scaling laws and budgets).
Calibration on non-reasoning text is a mismatch. Reasoning traces have their own distribution: long, structurally repetitive, heavy in arithmetic and in the model's own reasoning delimiters. Calibrating on a web crawl collects activation statistics from text that looks nothing like what the model will generate. See calibration data for post-training quantisation.
A 4-bit frontier model may still beat an fp16 smaller one, and that is the real comparison. The decision is rarely "quantise or not" and usually "which model fits the budget". Evaluate the quantised large model against the unquantised small one on the same hardware and the same latency target, which is the comparison the deployment actually faces.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
7 flashcards for this concept
Click a card to reveal the answer.