AI / A CONCEPT NOTE

Context Window

the working memory limit of a single model request

~60 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

The total number of tokens (system prompt + history + docs + response) the model can hold in memory during one call. Anything outside this window does not exist to the model.

02 / FOLLOW THE MECHANISM

How context window is allocated

  1. System Prompt

    establishes core instructions and constraints at the start.

  2. Chat History

    stores past messages, truncated when exceeding window limits.

  3. Retrieved Context

    injects RAG chunks relevant to the current query.

  4. Completion Reserve

    reserves max_tokens budget for model generation.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

count tokens in prompt

python -c import tiktoken; enc = tiktoken.get_encoding("cl100k_base"); print(len(enc.encode("test prompt")))

Explore command anatomy in the CLI lab