Definition
Distillation transfers behavior from a teacher (large model) to a student (smaller model) by matching soft probabilities, hidden states, or generated outputs as well as hard labels.
It is widely used to deploy fast classifiers and increasingly to build smaller LLMs.
In simple terms
A master chef trains an apprentice, and the apprentice learns shortcuts and taste judgments along with the recipes.
Where you see it
- Mobile vision models distilled from large CNNs.
- Small chat models trained on teacher demonstrations.
How it works
1.Run the teacher
Collect soft labels or traces.
2.Train the student
Match teacher outputs plus task loss.
3.Deploy student
Cheaper latency and memory.
Why it matters
- Distillation bridges research-scale models and production constraints.
Often confused
The student always equals the teacher.
Students approximate teachers, expect some capability gap.