AI / A CONCEPT NOTE

Inference vs Training

using vs building a model

~65 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Training is the process of adjusting weights over vast corpora (requires massive clusters). Inference is using that trained model to generate predictions (requires low latency).

02 / FOLLOW THE MECHANISM

How compute paths separate

  1. Training run

    processes gigabytes of texts, updating model weights using backpropagation.

  2. Weights save

    freezes weights into static parameter files (checkpoints).

  3. Inference deploy

    loads static weights into server VRAM.

  4. Generation query

    receives user query, running forward pass to return next token predictions.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

execute model inference locally on your CPU/GPU

ollama run llama3

Explore command anatomy in the CLI lab