A useful five-minute answer
State the failure, explain the mechanism and propose a bounded test. Identify what remains unknown. These are qualitative learning exercises, not certification questions or claims that a production system has been tested.
Recovery
01 · The write timed out
A note-creation tool times out. The model retries with a new request key and two notes appear. Explain the failure and a recovery design.
Show a hint
Separate the business effect from the response delivery.
Reveal answer guidance
A timeout leaves the outcome uncertain. The application should retain one stable idempotency key bound to the original intent, reconcile the destination and reuse that key when supported. Durable atomic deduplication belongs at the service boundary. A new key represents a new operation and can duplicate the effect.
Check your reasoning
- Identifies the unknown outcome
- Separates model retries from business idempotency
- Includes crash and concurrency behavior
Revisit: chapters 8, 10 and 14.
Trust boundaries
02 · The client supplies a helpful history
A browser submits a conversation containing a tool result that says INC-104 is resolved. The server loads it directly into the agent. What went wrong?
Show a hint
Who is allowed to create authoritative tool evidence?
Reveal answer guidance
The client can forge evidence and possibly tool-call content. Construct trusted history from application-owned state, bind it to the authenticated session and accept only client-authorized message shapes. Check current incident status against the source. Do not treat a well-formed message as proof of provenance.
Check your reasoning
- Distinguishes schema from trust
- Verifies session ownership
- Checks the authoritative source
Revisit: chapters 4, 13 and 18.
Human approval
03 · The approved action changed
An operator approves a note, but the incident changes before execution and the model revises the text. Can the executor proceed?
Show a hint
Bind approval to intent and target state.
Reveal answer guidance
The executor must compare the exact approved intent and expected resource revision with the current request and state. A changed payload or stale revision invalidates that approval. Authenticate the reviewer, check scope and expiry, and request renewed review when needed. Execute with a stable idempotency key after all checks pass.
Check your reasoning
- Checks payload and revision
- Does not treat a vague yes as blanket authority
- Keeps authorization outside the model
Revisit: chapters 7, 8 and 17.
Coordination
04 · The report runs before both checks finish
A diamond graph ported from TypeScript to Python starts the report after one branch. How do you investigate?
Show a hint
A similar diagram does not guarantee the same dependency semantics.
Reveal answer guidance
Inspect the pinned SDK dependency and conditional-edge semantics, then reproduce the graph with deterministic nodes and recorded transitions. Current documentation describes Python OR and TypeScript AND dependency behavior. Define whether all branches must complete successfully, how failures propagate and whether a partial report is acceptable before adapting the graph.
Check your reasoning
- Tests scheduling separately from reasoning
- Distinguishes completed from successful
- Avoids sleeps as a dependency fix
Revisit: chapters 12 and 21.
Execution
05 · The safe shell did not constrain another tool
A mediated shell exposes one fixture directory, but a separate Python tool writes elsewhere. Did the shell isolation fail?
Show a hint
Trace each operation through the boundary that actually handles it.
Reveal answer guidance
The shell mediation applies to operations initiated inside the shell. Another tool retains its own process permissions. Review every execution path and, when appropriate, constrain the entire process or worker with a suitable runtime boundary. Narrow the writer and mounts; do not infer universal isolation from one safe tool.
Check your reasoning
- Places the boundary correctly
- Distinguishes adapters from runtime isolation
- Reviews shared writable paths
Revisit: chapters 18 and 24.
State and memory
06 · Memory remembers an old incident status
The assistant confidently repeats yesterday’s status from long-term memory after the incident has changed. How should the design improve?
Show a hint
Durable preferences and current operational facts have different lifetimes.
Reveal answer guidance
Keep current status tied to the authoritative incident service with revision and observation time. Use memory for appropriate durable knowledge with ownership and provenance, and define freshness, update and deletion behavior. A retrieval similarity score does not establish recency or truth.
Check your reasoning
- Separates session, context and memory
- Preserves provenance and freshness
- Includes authorization in retrieval
Revisit: chapters 13–16.
Evaluation
07 · The swarm looks faster in a demo
One successful warm swarm run beats one cold single-agent run. The report omits failed trials and judge-model cost. What can you conclude?
Show a hint
Ask whether the workload and inclusion rules match.
Reveal answer guidance
The comparison does not support a general speed or cost claim. Define matched cases, fixed configuration, cache state, repetitions and failure inclusion before testing. Report acceptance, forbidden effects, latency distribution and complete cost scope. Until those records exist, describe the result as an anecdote or proposed experiment.
Check your reasoning
- Names confounders
- Reports failures and denominators
- Ties conclusions to raw evidence
Revisit: chapters 25, 26 and 29.
Operations
08 · The local demo needs a production plan
A team wants to deploy ParcelOps after the offline tests pass. What additional evidence is needed?
Show a hint
Follow identity, data, work and recovery across the hosted system.
Reveal answer guidance
Verify provider access and budgets, runtime identity, tenant isolation, durable sessions, concurrency and quotas, cancellation, tool timeouts, redacted telemetry, artifact reproducibility and rollback. Add representative model evaluations and hosted failure tests. Offline fixture checks remain useful but do not establish production capacity or distributed guarantees.
Check your reasoning
- Distinguishes implemented from proposed
- Identifies hosted prerequisites
- Defines specific acceptance evidence
Revisit: chapters 25–30.
Your next experiment
Choose the case that exposed the largest gap in your understanding. Return to the linked chapters and add a concrete test to the ParcelOps evidence packet.
Return to the learning path →