Quantising Mixture-of-Experts Models
Why MoE models are the most attractive quantisation target and the most fragile one, how the router turns rounding error into a discrete change of computation path, and what belongs in high precision.
Count the parameters in Mixtral 8x7B. Each expert is a SwiGLU feed-forward block with three matrices of \(4096 \times 14336\), which is 176M parameters; eight experts across 32 layers is 45.1B. Attention contributes about 1.3B and the embeddings about 0.3B, summing to the published 46.7B, of which the paper notes only around 13B are used per token (Jiang et al., 2024, Mixtral of Experts, arXiv:2401.04088).
So roughly 97 percent of the weights are expert weights, and a mixture-of-experts model is, from a quantisation point of view, almost entirely the thing you are quantising. The memory pressure is also worst exactly here, because every expert must be resident even though each token touches two. No other architecture offers that much leverage. None is as easy to break.
The router makes the error discrete
Quantising a dense layer perturbs its output continuously: a small weight error produces a small activation error. Quantising a router does something categorically different. The router emits logits, takes a top-\(k\), and the output of that comparison is a set of experts. Perturb the logits slightly and most tokens keep their experts; the tokens whose \(k\)-th and \((k+1)\)-th scores are close switch, and for those tokens the computation path changes entirely.
This is now a named problem. Small quantisation-induced perturbations in router outputs can significantly alter top-\(k\) expert assignment and destabilise routing (Liu et al., 2025, EAQuant, arXiv:2506.13329), and the stability of a given token depends on the margin between the last selected and first rejected expert, which is what later alignment-based methods target directly (Value-and-Structure Alignment for Routing-Consistent Quantization, arXiv:2606.05688).
The practical consequence is a short list of tensors that stay in high precision. DeepSeek-V3, trained in FP8 throughout, exempts the embedding module, the output head, the MoE gating network, normalisation and the attention operators, and promotes FP8 GEMM partial sums to fp32 registers every 128 elements. Activations use 1x128 tiles and weights 128x128 blocks in E4M3, with optimiser states held in BF16 (DeepSeek-AI, 2024, DeepSeek-V3 Technical Report, arXiv:2412.19437). A lab willing to train a 671B-parameter model with 37B active per token in 8-bit still declined to put the gate in 8-bit. Follow that.
Experts are not equally sensitive, and not equally used
A uniform bit width across a sparse model ignores its structure twice over. Shared experts, present in architectures like DeepSeek's, process every token, while routed experts fire for a selective minority, so an error in a shared expert is an error on every forward pass. The benchmark work on MoE post-training quantisation makes this the central finding: precision should be allocated per structure rather than per model, with heuristics ranging from whole MoE layers down to individual linear layers, and different MoE structures genuinely want different bit widths (Li et al., 2024, QuantMoE-Bench, arXiv:2406.08155).
Expert activation frequency is the second axis and it is heavily skewed. A rarely selected expert contributes to few outputs, so its quantisation error is sampled rarely, and a frequently selected one is the opposite. The budget argument writes itself: more bits for shared and hot experts, fewer for cold ones. The budget argument also has a trap, below.
When it breaks
Per-expert calibration data is thin by construction. With top-2 of 8 routing, a calibration token reaches a quarter of the experts, so a 128-sequence calibration set gives each expert roughly a quarter of the activation statistics a dense layer would get. With 256 experts and top-8, as in the larger sparse models, each expert sees about three percent. Reconstruction-based methods are estimating a Hessian per expert from a sample that small, and cold experts get the worst estimates while also being the ones a mixed-precision scheme assigned the fewest bits. The two heuristics compound in the wrong direction.
Low-frequency experts are low-frequency on your calibration set. Routing is input-dependent. An expert that fires for three percent of web text can fire for forty percent of a particular customer's traffic, and a bit allocation derived from generic calibration data silently assigned it two bits. The mitigation is to derive activation frequencies from traffic resembling deployment, and to re-derive them when the traffic changes.
Routing drift does not show up in aggregate accuracy. Expert-selection consistency against the fp16 model is a separate measurement from benchmark score, and a quantised MoE can match the aggregate while routing a meaningful fraction of tokens differently. Causal work on route-mediated damage in quantised MoE reports partial recovery from repairing flips on some models and a null result on others, so the field does not yet have a settled fix (arXiv:2608.11212). Measure consistency; do not assume a benchmark covers it.
Mixed precision per expert costs you the grouped kernel. Expert matmuls are normally executed as one grouped GEMM over a batch of experts. Different bit widths per expert means different kernels or padding to the widest, and the dispatch overhead can eat the memory saving. See custom kernels for mixture of experts.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Jiang et al., 2024, Mixtral of Experts, arXiv:2401.04088 arxiv.org
- Liu et al., 2025, EAQuant, arXiv:2506.13329 arxiv.org
- Value-and-Structure Alignment for Routing-Consistent Quantization, arXiv:2606.05688 arxiv.org
- DeepSeek-AI, 2024, DeepSeek-V3 Technical Report, arXiv:2412.19437 arxiv.org
- Li et al., 2024, QuantMoE-Bench, arXiv:2406.08155 arxiv.org
- arXiv:2608.11212 arxiv.org
7 flashcards for this concept
Click a card to reveal the answer.