Knowledge Distillation
Training a small, cheap model to imitate a large one's outputs — the standard route to frontier-quality behavior on a narrow task at a fraction of the cost.
Last reviewed: July 25, 2026
What is distillation?
Distillation transfers capability from a large “teacher” model to a small “student”: generate high-quality outputs with the teacher, then fine-tune the student on them. The student never matches the teacher in general — but on a defined task, a distilled 8B model routinely matches a frontier model at 10–100× lower inference cost and latency.
Why it matters commercially
Distillation is the standard endgame of LLM cost optimization: prototype on the best model available, collect its outputs on your real traffic, distill into a small model, route the narrow high-volume task to it. It’s also how small open models got good — most competitive sub-10B models are trained heavily on outputs of larger ones (the R1-distill family made this explicit for reasoning).
Distillation vs quantization vs fine-tuning
Frequently confused, cleanly different: quantization shrinks the same model’s numeric precision; fine-tuning adjusts a model on your examples; distillation is fine-tuning where the training data is another model’s outputs. They stack — a distilled student, quantized to 4-bit, is the classic edge-deployment recipe.
The catch: licensing and drift
Many providers’ terms restrict using their outputs to train competing models — read them before building a business on distilled data. And the student inherits the teacher’s mistakes at collection time: teacher hallucinations become baked-in student beliefs, so filter the training set with evals or verifiable checks first.
What people get wrong
- Distilling before the task is stable. Every prompt or scope change means regenerating data and retraining; distill after product behavior settles.
- Expecting general capability. The student mimics the teacher on the training distribution; off-distribution it degrades much faster than the teacher.
- Skipping the quality filter. “Train on everything the teacher said” transfers errors as faithfully as skills.
Response-Based vs. Feature-Based Distillation
The simplest and most common form of knowledge distillation, response-based distillation, trains the smaller “student” model to match the larger “teacher” model’s final output — often using the teacher’s full probability distribution over possible next tokens (called “soft labels”) rather than just its single most likely answer, since the relative probabilities the teacher assigns to different options carry additional useful information about how confident or uncertain it was. Feature-based distillation goes further, training the student to match the teacher’s internal intermediate representations at various layers, not just its final output — a more involved technique that can transfer more of the teacher’s learned structure, at the cost of requiring the student and teacher architectures to be compatible enough for intermediate representations to be meaningfully compared.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.