AI / A CONCEPT NOTE
Prompt Caching
optimizing LLM API latency and token reuse costs
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Saving the context history. When sending massive prompt instructions (like full project codebases or reference docs) to an LLM, the provider caches the initial part of the prompt. Subsequent questions reuse that cached state, slashing costs and response wait times.
02 / FOLLOW THE MECHANISM
How the cache saves time
First query
submits a 20,000-word codebase. The AI provider processes it and caches the processed memory state.
Second query
sends a short question using the exact same codebase prefix.
Cache match
the provider finds the pre-processed codebase in memory, skipping recalculation.
Fast output
the AI responds instantly and charges you a discounted rate for the cached text.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
call API with explicit prompt caching
curl -X POST -d '{"model":"claude-3-5-sonnet","system":[{"type":"text","text":"...","cache_control":{"type":"ephemeral"}}]}' https://api.anthropic.com/v1/messagesThe ellipsis omits required code or values. This sketch is not runnable as written.
05 / CHECK YOURSELF
Could you explain Prompt Caching to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringStructured Outputsenforcement of JSON schemas on LLM response payloads