AI / A CONCEPT NOTE

LLM

large language models that generate text by predicting the next token

~85 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

A Large Language Model is a neural network trained on vast text corpora. Given a prompt, it generates tokens (words or subwords) one at a time by predicting the most likely next token based on the preceding context. The 'magic' is pattern completion at massive scale — trillions of parameters in the largest models.

02 / FOLLOW THE MECHANISM

How an LLM generates a response

  1. Input prompt

    is tokenized — split into tokens and mapped to numeric IDs using the model's vocabulary.

  2. Transformer layers

    process the tokens through stacked self-attention and feed-forward layers, building a context-aware representation.

  3. Output head

    produces a probability distribution over the vocabulary for the next token.

  4. Sampling strategy

    selects the next token — greedy (highest probability), top-k, top-p (nucleus), or with temperature for creativity.

  5. Decode

    the selected token ID is mapped back to text, appended to the output, and fed back in for the next token.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

call an LLM API directly

curl -X POST -H "Content-Type: application/json" -d '{"model":"gpt-4","messages":[{"role":"user","content":"Hello"}]}' https://api.openai.com/v1/chat/completions

EXAMPLE 02 · REFERENCE

run an LLM locally with Ollama

ollama run llama3

Explore command anatomy in the CLI lab