The mechanism
Every model call has a finite context window. User messages, tool schemas, tool results and prior responses compete for that space. Replaying an entire long conversation can increase latency and cost while burying the evidence needed for the current task. Context management decides what remains visible, what is summarized and what can be retrieved later.
Current Strands documentation exposes automatic and agentic context modes through an experimental strategy API. The automatic mode applies configured offloading and summarization behavior; the agentic mode gives the model more control over its working context. These are version-sensitive features, not an unconditional promise of lower cost or higher accuracy.
A worked context experiment
Experimental configuration; optional inference. This fragment shows the documented mode selection. Pin the SDK and check the current API before using it in a release.
from strands import Agent
agent = Agent(
model=model,
tools=[lookup_incident],
context_manager="auto",
callback_handler=None,
)
Do not claim that this setting always reduces token use by a particular percentage. Summarization can itself require model work, retrieval can add calls, and important details can be lost. The correct question is whether the complete task improves under a controlled comparison.
For ParcelOps, three facts must survive: the user forbids changes, the incident is INC-104 and the latest verified status is delayed at a specific revision. A summary that preserves the incident story but drops the no-write constraint is unacceptable even if it saves many tokens. Put critical authorization in application enforcement so it does not depend on summarization fidelity.
Practice: design a context stress case
Offline fixture construction. Create a conversation with an initial no-write instruction, several irrelevant synthetic records, a corrected incident revision and a final question about the current status. Mark which messages are necessary and which are distractors. Then write the ideal concise working summary yourself.
Expected observation: “latest” depends on revision and timestamp, not position alone. An older message may contain a still-active constraint; a newer message may contain untrusted instructions. A simple keep-the-last-N strategy can lose essential information, while a summary can introduce its own errors.
For an optional live comparison, run the same fixture under a baseline and one context strategy. Record quality, retained constraints, input and output tokens, number of retrievals, latency and failures. Repeat cases instead of presenting one favorable run as a benchmark. Label all planned values as unmeasured until runs exist.
Troubleshooting and trade-offs
If the agent forgets a constraint, first verify whether enforcement existed outside the prompt. Then inspect what context the strategy retained using synthetic data and permitted diagnostics. If the model repeatedly retrieves offloaded content, apparent savings may disappear. If summaries become progressively less precise, shorten the chain or retain authoritative references.
Do not confuse context management with long-term memory. Context management shapes what one active conversation sees; memory may carry selected knowledge across conversations. Session persistence stores the conversation so it can resume. The three mechanisms overlap operationally but answer different questions.
Interview practice
How would you evaluate a context compression strategy?
Use matched tasks with fixed model and tool configuration, repeated trials and explicit critical facts. Measure quality and constraint retention alongside tokens, latency, retrieval overhead and failures.
Which constraints should never depend only on a summary?
Authorization, tenant boundaries, approved destinations and side-effect permissions. Keep those in trusted application and destination controls, even when the prompt also states them.
Completion check
Identify the minimum evidence and constraints needed for the stress case. Explain why a smaller prompt is not automatically a better system. Mark the context strategy API as experimental in your environment card.
Sources and version notes
Checked 6 October 2026. Python examples target strands-agents==1.58.0 unless labelled otherwise. Live documentation can change; compare your installed version before adapting an example.
Make the understanding yours.
Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.
Self-assessed reading progress. This does not certify that a lab ran or a system is secure.