AI / A CONCEPT NOTE
Distillation
compressing knowledge from a large teacher model into a smaller student model
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
A model compression technique where a smaller student model is trained on outputs, reasoning traces, or probability logits of a massive teacher model to match performance at lower latency.
02 / FOLLOW THE MECHANISM
How knowledge distillation works
Teacher Model (e.g. GPT-4)
generates high-quality responses and reasoning chains for dataset prompts.
Dataset Collector
filters and formats teacher completions into fine-tuning pairs.
Student Model (e.g. 8B)
is fine-tuned to minimize loss against teacher output targets.
Deployment
serves student model at 10x lower cost and fraction of latency.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
distillation loss logic
python -c print("KL-divergence loss matches student logits to teacher logits")05 / CHECK YOURSELF
Could you explain Distillation to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringLLMlarge language models that generate text by predicting the next token