Knowledge Distillation advanced 8 min read 6 flashcards

Cross-Tokeniser Distillation

Why a KL divergence between a teacher and a student with different tokenisers is not even well defined, the two distinct mismatches involved, and how optimal-transport and alignment-based losses get around them.

Logit distillation assumes the teacher and student put probability on the same set of symbols at the same positions. Pick the best available teacher and a student from a different family, and neither assumption holds. The standard workaround is to give up on distributions and train the student on the teacher's sampled text, which throws away most of the signal. The methods below keep the distribution and pay for it differently.

Two mismatches, not one

The first is dimensional. The teacher's distribution is a vector over its vocabulary, the student's over a different vocabulary of different size, and a KL divergence between vectors indexed by different symbol sets has no meaning. There is no canonical permutation to apply, because the vocabularies are not two orderings of one set.

The second is positional, and it is the harder one. Tokenisers segment the same string differently, so position \(t\) is not the same prediction problem for both models. Take the string unbelievable. A teacher might emit un + believable, two steps, while a student emits un + bel + iev + able, four. After step one the two models are conditioned on identical text and are predicting different things: the teacher's step-two target spans the whole remainder, the student's spans three characters. Even with a perfect vocabulary correspondence, aligning step to step is a sequence-alignment problem, not a lookup.

Optimal transport over sorted probabilities

The approach that avoids both mismatches treats each distribution as a measure and asks for the cheapest way to move one onto the other. Universal Logit Distillation sorts both probability vectors, pads to a common length, and minimises a token-wise optimal-transport cost between them, which requires no correspondence between vocabularies at all (Boizard et al., 2024, Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs, arXiv:2402.12030, published in TMLR, 2025). What the loss then matches is the shape of the teacher's distribution, its concentration and tail mass, rather than which symbols carry it.

Later work extends transport beyond the single-token level, formulating the objective jointly over token and sequence levels so that the alignment is solved rather than assumed (Cui et al., AAAI 2025, Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation, arXiv:2412.14528).

The alternative family does the alignment explicitly: compute a dynamic-programming alignment between the two tokenisations of the same string, then aggregate teacher probabilities over the token spans that correspond to each student step. This preserves symbol identity, which transport discards, at the cost of needing the alignment to exist.

When it breaks

Sorted-probability transport is identity-blind. Matching the \(i\)-th largest teacher probability to the \(i\)-th largest student probability can pair unrelated tokens. The loss is then satisfied by a student whose distribution has the right entropy and the wrong support. The signal it reliably transfers is calibration, not content, which is useful and is less than it looks.

Alignment fails when tokenisations are not refinements of one another. Span aggregation works when one segmentation subdivides the other. Byte-level and unicode-normalising tokenisers can split a string at genuinely incompatible points, and a model with a different normalisation step may not even see the same character sequence.

Numbers, code and non-Latin scripts are the worst cases. These are precisely where tokenisers differ most, where segmentation boundaries carry meaning, and where token healing and boundary bias already cause trouble.

The teacher's full distribution is expensive. Transport costs need more than the top-k probabilities, so the cheap caching strategy that makes same-tokeniser distillation affordable is unavailable, and a 100k-wide float vector per token is a serious storage and bandwidth cost.

Sequence-level distillation remains the robust baseline. It is tokeniser-agnostic by construction and it is what the distribution-matching methods must beat. Compare against it honestly; see sequence-level vs token-level distillation.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Boizard et al., 2024, Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs, arXiv:2402.12030 arxiv.org
  2. Cui et al., AAAI 2025, Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation, arXiv:2412.14528 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track