Borrowed Intelligence: How Knowledge Distillation Builds Small Language Models That Punch Above Their Weight
A 2-billion-parameter model that trades blows with one ten times its size is not an accident of architecture. It is the product of a teacher pouring its full probability distribution into a student, token by token.