AI / A CONCEPT NOTE
LLM
large language models that generate text by predicting the next token
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
A Large Language Model is a neural network trained on vast text corpora. Given a prompt, it generates tokens (words or subwords) one at a time by predicting the most likely next token based on the preceding context. The 'magic' is pattern completion at massive scale — trillions of parameters in the largest models.
02 / FOLLOW THE MECHANISM
How an LLM generates a response
Input prompt
is tokenized — split into tokens and mapped to numeric IDs using the model's vocabulary.
Transformer layers
process the tokens through stacked self-attention and feed-forward layers, building a context-aware representation.
Output head
produces a probability distribution over the vocabulary for the next token.
Sampling strategy
selects the next token — greedy (highest probability), top-k, top-p (nucleus), or with temperature for creativity.
Decode
the selected token ID is mapped back to text, appended to the output, and fed back in for the next token.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
call an LLM API directly
curl -X POST -H "Content-Type: application/json" -d '{"model":"gpt-4","messages":[{"role":"user","content":"Hello"}]}' https://api.openai.com/v1/chat/completionsrun an LLM locally with Ollama
ollama run llama305 / CHECK YOURSELF
Could you explain LLM to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringRAGgrounding LLM answers in your own data