Transfer Sets and Data-Free Distillation
Distillation needs inputs to query the teacher on but not labels, which makes the choice of transfer set the dominant design decision, and makes distillation possible even when the original training data is gone.
A distillation loss compares two distributions over the same input. It never needs the ground-truth label. That single property, noted in the original paper, is why distillation can run on unlabeled data, and why the set of inputs you query the teacher on, the transfer set, is a first-class design decision rather than a detail (Hinton et al., 2015, arXiv:1503.02531). Teams routinely spend weeks on the loss function and minutes on the transfer set, which is the wrong allocation: a mismatched transfer set caps the student's deployment accuracy in a way no loss can recover.
What the transfer set has to cover
The student only learns the teacher's behaviour where it is asked about it. Formally, distillation minimises an expected divergence under the transfer distribution \(q\):
Regions where \(q\) puts no mass are unconstrained, and the student's behaviour there is whatever its inductive bias produces. So the requirement is not that \(q\) match the teacher's training distribution; it is that \(q\) cover the deployment distribution. Those differ whenever the student will be used on traffic the teacher was not trained for, which is the common case.
This reframes an apparent paradox. Distilling on data the teacher never saw often works well, because the teacher generalises and its outputs on that data are still informative. Distilling on a narrower set than deployment fails even when that set is exactly the teacher's training data.
When the data is gone
Often you have a teacher and no data: the corpus was licensed and the licence expired, it held personal data that cannot be retained, or it was never published. Data-free distillation synthesises a transfer set instead.
The first approach shipped activation statistics alongside the weights, as metadata, and reconstructed inputs that reproduce those statistics (Lopes et al., 2017, Data-Free Knowledge Distillation for Deep Neural Networks, arXiv:1710.07535). Removing even the metadata requirement, zero-shot distillation samples target output distributions from a Dirichlet prior fitted to the teacher's class structure and then optimises inputs, called data impressions, that elicit them. On MNIST this reached 98.77% with 24,000 data impressions, against 92.47% for the metadata-based approach and 86.70% for a few-real-examples baseline (Nayak et al., ICML 2019, Zero-Shot Knowledge Distillation in Deep Networks, arXiv:1905.08114).
For language models the synthetic-transfer-set move is now routine and usually goes by other names: generating prompts, then generating teacher responses to them. The engineering question is identical, and so is the failure mode.
When it breaks
Synthetic inputs are off-manifold, so you distil extrapolation. The teacher's output on an optimised, unnatural input is not knowledge; it is an artefact of the teacher's behaviour outside its training support. Data-free students inherit those artefacts and there is no held-out set that reveals it, because the held-out set is synthetic too.
Coverage is not measurable. You can measure the student's agreement with the teacher on the transfer set, which is the quantity you optimised, and it tells you nothing about the gap. The only honest check is agreement on a real sample of deployment traffic, which data-free settings by definition lack.
Published data-free results are mostly small-image benchmarks. MNIST and CIFAR-scale classification is a long way from a language model with a 100k-token vocabulary and open-ended outputs. Treat the technique as established for the former and exploratory for the latter.
The legal position is not the technical one. Querying a hosted teacher to build a transfer set is cheap and frequently prohibited. See distillation and terms-of-service constraints; the fact that a pipeline runs is not evidence that it is permitted.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
6 flashcards for this concept
Click a card to reveal the answer.