AI / A CONCEPT NOTE

Prompt Caching

optimizing LLM API latency and token reuse costs

~70 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Saving the context history. When sending massive prompt instructions (like full project codebases or reference docs) to an LLM, the provider caches the initial part of the prompt. Subsequent questions reuse that cached state, slashing costs and response wait times.

02 / FOLLOW THE MECHANISM

How the cache saves time

  1. First query

    submits a 20,000-word codebase. The AI provider processes it and caches the processed memory state.

  2. Second query

    sends a short question using the exact same codebase prefix.

  3. Cache match

    the provider finds the pre-processed codebase in memory, skipping recalculation.

  4. Fast output

    the AI responds instantly and charges you a discounted rate for the cached text.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · INCOMPLETE SKETCH

call API with explicit prompt caching

curl -X POST -d '{"model":"claude-3-5-sonnet","system":[{"type":"text","text":"...","cache_control":{"type":"ephemeral"}}]}' https://api.anthropic.com/v1/messages

The ellipsis omits required code or values. This sketch is not runnable as written.

Explore command anatomy in the CLI lab