The mechanism
Metrics show aggregate behavior, traces show the path of one request, and logs provide selected diagnostic events. A useful agent system needs all three at an appropriate level of detail. Counting successful HTTP responses alone misses unsupported answers, denied tool calls, budget exhaustion and work that continued after the caller disconnected.
ParcelOps correlates the incoming request, model calls, tool execution, policy decisions and final acceptance result. The correlation ID comes from trusted application code. It should be safe to share with support without exposing incident content or credentials. A trace is an operational record, not a reason to publish private model inputs.
A worked metrics extraction
Post-inference inspection. result is an actual AgentResult only after you run an approved model invocation. This code does not represent measured values in the handbook.
usage = result.metrics.accumulated_usage
record = {
"stop_reason": result.stop_reason,
"input_tokens": usage.get("inputTokens"),
"output_tokens": usage.get("outputTokens"),
"total_tokens": usage.get("totalTokens"),
"cycle_seconds": sum(result.metrics.cycle_durations),
"tool_names": list(result.metrics.tool_metrics.keys()),
}
print(record)
Missing values remain missing. Do not silently turn an unavailable usage field into zero, which would understate cost. Cycle duration is not necessarily identical to end-to-end request latency; queueing, transport, authentication and downstream work can add time outside the measured cycles. Capture an outer monotonic timer for the user-visible request as well.
Token usage also differs from billing. Cached inputs, auxiliary calls, provider-specific categories and external service charges may require separate accounting. Record currency and pricing date when estimating cost, and reconcile against actual billing before making financial claims.
Practice: diagnose three traces
Offline. Create three synthetic timelines: slow provider response, slow incident lookup and long retry backoff. All take the same total time. Label spans so an operator can distinguish them without reading raw incident text.
Expected observation: a single latency histogram cannot identify the cause. A trace can reveal where time was spent, while metrics show whether the pattern affects many requests. Add a final acceptance flag so a fast but unsupported answer does not look like a successful optimization.
Now insert a canary string into a synthetic prompt and tool argument. Define which telemetry fields may retain it and which must redact it. Verify redaction before data leaves the process. Restrict trace access and retention even when payload capture is disabled, because metadata can still reveal usage patterns.
Troubleshooting and trade-offs
High-cardinality labels such as full prompts or arbitrary incident IDs can overwhelm metrics systems and expose sensitive data. Keep aggregate labels bounded; use trace correlation for individual investigations. Missing spans may come from instrumentation or export configuration rather than absent work, so test the telemetry path with a known synthetic request.
Debug logging can reveal more than expected. Enable it only in a controlled environment and inspect the output before using real data. An observability dashboard should distinguish “not observed” from “did not happen”; dropped telemetry is an uncertainty that must remain visible.
Interview practice
Why measure both task acceptance and transport success?
A request can return HTTP success while producing an unsupported answer or stopping at a budget limit. Acceptance checks capture whether the intended task was completed under its constraints.
How would you investigate high latency without logging raw prompts?
Use request correlation, model and tool spans, retry counts, queue time, stop reasons and bounded metadata. Reproduce with synthetic fixtures if deeper payload inspection is needed.
Completion check
Design one metric, one trace span and one safe log event for ParcelOps. Explain missing usage fields and distinguish cycle duration from end-to-end latency. Keep all example measurements blank until observed.
Sources and version notes
Checked 6 October 2026. Python examples target strands-agents==1.58.0 unless labelled otherwise. Live documentation can change; compare your installed version before adapting an example.
Make the understanding yours.
Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.
Self-assessed reading progress. This does not certify that a lab ran or a system is secure.