Knowledge Distillation intermediate 7 min read 6 flashcards

Multi-Teacher and Ensemble Distillation

Why compressing an ensemble was distillation's original purpose, how averaging teachers in probability space differs from averaging logits, and why teacher disagreement is the only part of the signal that is worth paying for.

An ensemble of five models beats any one of them and costs five forward passes. That trade is the problem distillation was invented to solve: the original proposal frames it as compressing the knowledge in an ensemble into a single model that is cheap to deploy (Hinton et al., 2015, Distilling the Knowledge in a Neural Network, arXiv:1503.02531). Single-teacher distillation, now the default framing, is the special case. Going back to the general case changes what the target actually contains.

Where the extra signal lives

With one teacher, the soft target carries the teacher's uncertainty, the dark knowledge covered in temperature and dark knowledge. With several teachers you additionally get their disagreement, and that is a different quantity. On examples where all teachers are confident and agree, the averaged distribution is close to one-hot and tells the student almost nothing a hard label would not. On examples where they disagree, the average encodes genuine epistemic uncertainty about the label, and that is the part a single teacher cannot supply at any temperature.

This has a practical consequence. Teachers trained from different seeds on the same data with the same recipe are highly correlated, so their disagreement is small and the ensemble adds little beyond variance reduction. Teachers that differ in architecture, data mixture, or objective disagree in structured ways, and that is where multi-teacher distillation earns its cost.

Averaging probabilities is not averaging logits

Two obvious combination rules give different students. Averaging in probability space,

\[\bar p(y) = \frac{1}{K}\sum_{k=1}^{K} p_k(y),\]

reproduces the ensemble's predictive distribution exactly, which is usually what you want, because the ensemble's calibration is part of what you are buying. Averaging logits before the softmax,

\[\bar\ell = \frac{1}{K}\sum_k \ell_k,\]

is a geometric rather than arithmetic mean after normalisation, so it is dominated by whichever teacher is most confident and it suppresses exactly the disagreement you were paying for. A third option, averaging the per-teacher KL losses rather than the targets, is equivalent to probability-space averaging in the gradient only when every teacher sees the same temperature.

Equal weights are also a choice, and usually a poor one. Weighting teachers per instance, rather than fixing one weight per teacher for the whole run, improves the student, because which teacher is trustworthy varies with the example and with the student's current capability (Yuan et al., AAAI 2021, Reinforced Multi-Teacher Selection for Knowledge Distillation, arXiv:2012.06048).

Intermediate layers cannot be averaged

Output distributions live in a shared space, so averaging is well defined. Hidden states do not: two teachers with different widths or depths have no correspondence between their intermediate representations. Multi-teacher methods that want intermediate supervision therefore use relational objectives instead of matching vectors directly, for example a triplet loss on relative dissimilarities so that the student reproduces the geometry each teacher induces rather than its coordinates (You et al., KDD 2017, Learning from Multiple Teacher Networks). This is the same reasoning behind single-teacher relational methods in feature and attention transfer distillation.

When it breaks

The averaged target may be unreachable. A mixture of two confident, disagreeing teachers is bimodal, and no single student of limited capacity can match a bimodal target while also matching the modes. This compounds the ordinary capacity gap: the student is now chasing a distribution that no model in the ensemble would have produced.

Correlated teachers cost K times as much for almost nothing. Measure pairwise disagreement on held-out data before committing to a second teacher. If mean pairwise KL between teachers is small relative to the student's own loss, the ensemble is not the bottleneck.

Per-instance weighting collapses. Learned weighting schemes frequently converge to putting nearly all weight on one teacher, which is a result worth checking for rather than assuming away; when it happens, the simpler single-teacher pipeline is the honest baseline.

Transfer-set cost scales with K. Every example needs K teacher forward passes unless you cache distributions, and caching a full vocabulary distribution per token is expensive for language models. In practice teams cache top-k probabilities, which quietly truncates the disagreement signal in the tail.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Hinton et al., 2015, Distilling the Knowledge in a Neural Network, arXiv:1503.02531 arxiv.org
  2. Yuan et al., AAAI 2021, Reinforced Multi-Teacher Selection for Knowledge Distillation, arXiv:2012.06048 arxiv.org
  3. You et al., KDD 2017, Learning from Multiple Teacher Networks dl.acm.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track