Distilling Rerankers into Retrievers
How a cross-encoder's ranking signal is moved into a bi-encoder's parameters so it is paid at index time, why the target is a margin rather than a score, and why hard negatives decide whether any of it transfers.
A cross-encoder scores a query and a document jointly and is the most accurate ranker available. It is also unusable as a retriever: scoring a million documents means a million forward passes per query. A bi-encoder embeds documents once, offline, and compares with a dot product, which is cheap and markedly less accurate. Distillation is how the accuracy moves from the first architecture to the second, and the interesting part is that the usual distillation target does not work here.
The target is an order, not a distribution
Classification distillation matches a probability vector. Ranking has no such vector: relevance scores from different architectures live on different, arbitrary scales. A cross-encoder might emit logits in \([-8, 8]\) while a bi-encoder emits cosine similarities in \([-1, 1]\), and forcing the student to reproduce the teacher's absolute numbers wastes capacity on a calibration problem nobody asked about.
The fix is to distil differences. For a query \(q\) with a relevant passage \(d^{+}\) and a non-relevant \(d^{-}\), Margin-MSE matches the teacher's margin:
This is scale-free in the teacher's offset and transfers across architectures that score on different ranges, which is why it was introduced as a cross-architecture procedure, with an ensemble of three BERT cross-encoder teachers supervising DistilBERT-sized students on MS MARCO passage ranking (Hofstätter et al., 2020, Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation, arXiv:2010.02666).
The other standard target is a KL divergence over a candidate list: normalise the teacher's scores across the retrieved candidates for one query, do the same for the student, and match. ColBERTv2 does this with a 23M-parameter MiniLM cross-encoder as the teacher, combined with hard-negative mining, and couples it with residual compression that cuts the index footprint 6 to 10 times; it reports the best quality on 22 of 28 out-of-domain tests in its evaluation (Santhanam et al., 2022, ColBERTv2, arXiv:2112.01488).
Hard negatives are not an optimisation
The loss only teaches ordering among candidates you put in front of it. With random negatives, the teacher and the student already agree, the margin is large for both, and the gradient is near zero: you have spent teacher compute to learn nothing. The candidates must come from the retriever's own confusions, which is the same argument as in hard negative mining and contrastive embedding training, and it is why distillation and negative mining are reported together rather than as independent contributions.
When it breaks
The teacher is the ceiling, and its errors arrive with confidence. A cross-encoder that systematically prefers lexically overlapping passages teaches that preference as if it were relevance. Distillation narrows the gap to the teacher; it does not audit the teacher.
Distillation does not fix recall. The teacher only ever scores candidates the first-stage retriever returned. Documents the retriever failed to surface are invisible to the entire pipeline, so no amount of score distillation improves them. Recall is a candidate-generation problem.
Out-of-domain gains are fragile because the negatives are in-domain. The negative distribution is a property of the corpus the mining ran on. A student distilled on one corpus has learned to separate that corpus's confusions, which is not the same as learning relevance, and this is one mechanism behind the gap discussed in embedding benchmarks and the zero-shot problem.
Some of the gap is architectural and will not close. A bi-encoder must compress a document into a fixed vector before seeing the query, so query-document interaction is unrepresentable in principle. Distillation recovers what is representable. Recovering more needs a different architecture, which is the argument for late interaction.
Teacher inference dominates the training budget. Scoring enough query-candidate pairs to cover the corpus is often more compute than training the student, and it is the cost that decides whether a reranker-to-retriever pipeline is affordable at all.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
6 flashcards for this concept
Click a card to reveal the answer.