AI / A CONCEPT NOTE

Distillation

compressing knowledge from a large teacher model into a smaller student model

~70 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

A model compression technique where a smaller student model is trained on outputs, reasoning traces, or probability logits of a massive teacher model to match performance at lower latency.

02 / FOLLOW THE MECHANISM

How knowledge distillation works

  1. Teacher Model (e.g. GPT-4)

    generates high-quality responses and reasoning chains for dataset prompts.

  2. Dataset Collector

    filters and formats teacher completions into fine-tuning pairs.

  3. Student Model (e.g. 8B)

    is fine-tuned to minimize loss against teacher output targets.

  4. Deployment

    serves student model at 10x lower cost and fraction of latency.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

distillation loss logic

python -c print("KL-divergence loss matches student logits to teacher logits")

Explore command anatomy in the CLI lab