Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 29 / 30 · Build with evidence

Benchmark methods before benchmark claims

Define comparable workloads, failure inclusion and measurement boundaries before testing.

4 min read + practiceWorked exerciseInterview practice

The mechanism

A benchmark answers a narrow question under specified conditions. “MCP is fast” is too broad. A useful question might compare median and tail latency for the same synthetic lookup over two transports using pinned implementations and equal concurrency.

Separate protocol overhead, business-handler time, authentication, network delay and model inference. A model-heavy workload can hide transport costs; a trivial in-memory handler can exaggerate their importance. Record warm and cold states, cache policy, payload size and client/server versions.

This chapter provides a proposed experiment, not measured results. No performance improvement, throughput figure or enterprise-scale claim is asserted. The downloadable CSV starts with headers only so the empty dataset cannot be mistaken for a completed benchmark.

Question
Controlled workload
All trials
Bounded conclusion

Worked example

This is the measurement schema for a future approved experiment. Keep failures in the dataset and explain timeouts as censored or failed observations rather than silently dropping them. Use a monotonic clock for elapsed time.

run_id,case_id,protocol,sdk_version,transport,concurrency,cache_state,payload_bytes,elapsed_ms,outcome,forbidden_effects,evidence_path

Practice: predict, inspect, explain

Offline exercise. Design five synthetic cases: small lookup, missing record, invalid input, denied operation and larger bounded output. Specify repetitions, warm-up handling, timeout budget and acceptance rules before measuring. Add a separate end-to-end model evaluation only if that is part of the question.

Expected observation: no numerical conclusion is available until data exists. Compare the same workload and report success rate, failure categories, latency distribution and resource scope. A faster run that omits authorization or returns less evidence is not an equivalent comparison.

Troubleshooting and trade-offs

If results vary wildly, inspect cache state, environment contention and background work. If p95 is reported from a tiny sample, show the sample size and uncertainty. If failed trials disappear, restore them before interpreting the result. Avoid combining optional model cost with transport cost without a clear accounting boundary.

Interview practice

Why include denied and invalid requests?

They exercise real policy and validation paths and reveal whether performance optimizations weaken controls or mishandle errors.

What invalidates a transport comparison?

Different workloads, versions, authorization work, cache states, concurrency or failure inclusion can confound the result. Match and disclose those conditions.

Completion check

Write a preregistered experiment plan and leave all result fields blank until evidence exists.

Sources and version notes

This edition targets MCP 2026-07-28, checked 6 October 2026. SDK examples are version-sensitive and labelled when not executed. Synthetic fixtures are learning material, not protocol conformance evidence.

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.