AI / A CONCEPT NOTE
Next-Token Prediction
how autoregressive models generate text word by word
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
LLMs do not generate full answers at once. Given prior context, the model computes a probability distribution over the vocabulary for the single next token, samples one, appends it to context, and repeats.
02 / FOLLOW THE MECHANISM
How autoregressive generation flows
Prompt
is tokenized into a sequence of integer IDs.
Model Forward Pass
processes visible sequence through transformer layers to output logits.
Sampling (Temp/Top-p)
selects the next token from the probability distribution.
Loop
appends selected token to context and repeats until EOS token or limit.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
stream tokens in real-time
curl -N http://localhost:11434/api/generate05 / CHECK YOURSELF
Could you explain Next-Token Prediction to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringContext Windowthe working memory limit of a single model request