AI / A CONCEPT NOTE

Next-Token Prediction

how autoregressive models generate text word by word

~60 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

LLMs do not generate full answers at once. Given prior context, the model computes a probability distribution over the vocabulary for the single next token, samples one, appends it to context, and repeats.

02 / FOLLOW THE MECHANISM

How autoregressive generation flows

  1. Prompt

    is tokenized into a sequence of integer IDs.

  2. Model Forward Pass

    processes visible sequence through transformer layers to output logits.

  3. Sampling (Temp/Top-p)

    selects the next token from the probability distribution.

  4. Loop

    appends selected token to context and repeats until EOS token or limit.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

stream tokens in real-time

curl -N http://localhost:11434/api/generate

Explore command anatomy in the CLI lab