AI / A CONCEPT NOTE
Context Window
the working memory limit of a single model request
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
The total number of tokens (system prompt + history + docs + response) the model can hold in memory during one call. Anything outside this window does not exist to the model.
02 / FOLLOW THE MECHANISM
How context window is allocated
System Prompt
establishes core instructions and constraints at the start.
Chat History
stores past messages, truncated when exceeding window limits.
Retrieved Context
injects RAG chunks relevant to the current query.
Completion Reserve
reserves max_tokens budget for model generation.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
count tokens in prompt
python -c import tiktoken; enc = tiktoken.get_encoding("cl100k_base"); print(len(enc.encode("test prompt")))05 / CHECK YOURSELF
Could you explain Context Window to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringAttentionhow models calculate context-sensitive word relationships