The mechanism
A benchmark is a defined experiment, not a favorable anecdote. It specifies a workload, configuration, measurement procedure and interpretation. Agent systems vary across runs, model versions and external dependencies, so a useful comparison controls what it can and records what it cannot.
ParcelOps can compare a deterministic template, one agent and a specialist configuration. Another experiment compares context strategies. Change one major factor at a time so the result has an interpretable cause. Do not simultaneously change the prompt, model, tools and dataset and attribute the difference to the SDK.
A worked experiment plan
| Field | Proposed plan |
|---|---|
| Task | Explain synthetic incidents with supported evidence |
| Variants | Template baseline; one agent; specialist agent |
| Cases | Happy path, missing data, stale data, injection, outage |
| Primary outcome | Acceptance checks pass with no forbidden effects |
| Secondary outcomes | End-to-end latency, token categories, retries, tool calls |
| Repetition | Predeclare repeated trials and preserve failures |
| Status | Proposed; no benchmark results supplied |
Download the blank benchmark CSV and run record template. Empty measurements mean unmeasured, not zero. The templates intentionally contain no fabricated speedup, accuracy or savings values.
Include warmup and cache state in the procedure. A cached run and a cold run may exercise different paths. Record provider region, model identifier, SDK version, prompt and tool revisions, dataset hash and concurrency. Use an outer monotonic timer for end-to-end latency and preserve the provider usage categories available in the result.
Practice: detect a misleading comparison
Offline. Consider two reports. Report A averages only successful runs from the new configuration but includes all failures in the baseline. Report B uses different case sets and calls the lower average latency a speedup. Explain why both comparisons are invalid.
Expected observation: selection and workload differences can produce an apparent improvement without a better system. Define inclusion rules before running the experiment. Report failures and timeouts explicitly, and avoid replacing them with convenient latency values that hide their meaning.
Next, compute a simple acceptance rate from a fictional ten-case table that you label as teaching data. Then remove the table from any public results section unless the label remains unmistakable. A learning exercise with invented numbers must never become a claim that the handbook’s system achieved those numbers.
Interpreting real results
When real runs exist, report counts, distribution and uncertainty. Median latency is useful but can hide a poor tail; include a tail measure only with enough samples to interpret it responsibly. A small test set supports a narrow conclusion, not a universal ranking of frameworks.
Cost comparisons need a pricing date, currency and billing scope. Include judge-model and external-service costs if they are part of the experiment. Token reductions can be informative even before billing reconciliation, but label them as token measurements rather than verified monetary savings.
Troubleshooting and trade-offs
If variation overwhelms the difference, increase repetitions within an approved budget or narrow the claim. If the new configuration is faster but fails more negative cases, report the trade-off. If it performs better only on the tuning set, evaluate the held-out cases before recommending it.
Retain raw sanitized records and the analysis procedure so someone else can recompute the summary. Keep model text, private prompts and secrets out of public benchmark artifacts. A transparent limitations section makes the finding more useful, not less credible.
Interview practice
What must be fixed or recorded in an agent benchmark?
The task distribution, configurations, models, tool contracts, prompts, versions, cache state, concurrency, inclusion rules and measurement procedure. Record changing external dependencies and state the limits of reproducibility.
Why report failures beside latency?
A system can appear fast by failing early or excluding difficult cases. Quality, forbidden effects and completion rates are necessary context for interpreting performance measurements.
Completion check
Fill the experiment-plan fields while leaving results blank. Identify one confounder and one failure-inclusion rule. Explain what evidence would justify a narrow improvement claim and what would still remain uncertain.
Sources and version notes
Checked 6 October 2026. Python examples target strands-agents==1.58.0 unless labelled otherwise. Live documentation can change; compare your installed version before adapting an example.
Make the understanding yours.
Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.
Self-assessed reading progress. This does not certify that a lab ran or a system is secure.