AI / A CONCEPT NOTE
Inference vs Training
using vs building a model
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Training is the process of adjusting weights over vast corpora (requires massive clusters). Inference is using that trained model to generate predictions (requires low latency).
02 / FOLLOW THE MECHANISM
How compute paths separate
Training run
processes gigabytes of texts, updating model weights using backpropagation.
Weights save
freezes weights into static parameter files (checkpoints).
Inference deploy
loads static weights into server VRAM.
Generation query
receives user query, running forward pass to return next token predictions.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
execute model inference locally on your CPU/GPU
ollama run llama305 / CHECK YOURSELF
Could you explain Inference vs Training to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringGPUs & VRAMwhy AI needs special hardware