AI / A CONCEPT NOTE
Tokens & Context
what the model can 'see'
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Tokens are text snippets (words or subwords). The context window is the maximum number of prompt and generation tokens a model can process in a single request.
02 / FOLLOW THE MECHANISM
How context limits operation
Prompt load
user sends a large context document. Tokenizer splits text into token IDs.
Limit check
compares prompt tokens against model limit (e.g. 200K tokens).
Processing
attention layers compute relationships across all tokens simultaneously.
Overflow
if prompt size exceeds limit, API throws a context-exhaustion exception.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
count tokens in prompt string
python -c "import tiktoken; enc = tiktoken.get_encoding('cl100k_base'); print(len(enc.encode('hello context')))"05 / CHECK YOURSELF
Could you explain Tokens & Context to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringInference vs Trainingusing vs building a model