Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 29 / 30 · Evidence

Measure before making a performance claim

Design comparable trials and publish failures, spread and uncertainty alongside successful runs.

4 min readWorked exerciseInterview practice

What you will build

A benchmark protocol and blank results ledger. This edition contains no measured Docker Sandbox performance results. All numeric result fields remain unfilled until a real authorized experiment supplies evidence.

The mechanism at a glance
  1. Predeclare protocol
  2. Run comparable repetitions
  3. Retain raw successful and failed trials
  4. Analyze with explicit limits

Conceptual flow. Follow the lesson for prerequisites, exact commands and verification limits.

Read the mechanism

“Time to result” includes several stages: template preparation, VM startup, kit setup, agent reasoning, build/test execution, export and cleanup. A faster total can reflect a warmer cache or a shorter model answer rather than a faster isolation backend.

Separate deterministic workload benchmarks from agent-task evaluations. A fixed parser test suite can measure execution overhead. An agent fixing the parser introduces variable planning, tool calls, retries and model behavior. Both are interesting, but they answer different questions.

Control hardware, architecture, image identity, source revision, resource caps, policy and cache state. When conditions cannot be identical, report them and avoid a causal claim the experiment cannot support.

Worked example · predeclare a protocol

Choose three workspace modes—direct, clone and mountless—and one deterministic build/test workload. Before execution, state the repetition count, warm-up procedure, timeout, exclusion rules and how failures will be recorded.

Use the blank benchmark CSV or the JSON run template. A minimal result table begins like this:

ModeTrialsMedian task timeSpreadFailuresCost coverage
DirectNot run———Unknown
CloneNot run———Unknown
MountlessNot run———Unknown
Cloud, separate protocolNot run———Unknown

Use a monotonic clock for durations within one process. Define timestamps for ready, setup complete, task start/end, export end and confirmed cleanup. Do not subtract unsynchronized wall clocks from different machines as though they form one precise stopwatch.

Expected observations

After real trials, publish sample count and spread along with the central estimate. Include unsuccessful and timed-out runs in a separate visible category; silently dropping them can make an unreliable system look fast.

Cost can be known, estimated or unknown. Keep model charges separate from cloud compute and document the accounting window. A missing bill is not evidence of zero cost.

Useful findings can include: which phase dominates, whether warm state changes results, which failures recur, and which tasks cannot be compared because a backend lacks required behavior.

Troubleshooting

If timing starts after setup for one mode and before setup for another, fix the protocol before interpreting results. If one trial uses a different model or source revision, mark it non-comparable rather than quietly pooling it.

Do not claim statistical certainty from a tiny sample. Explain selection bias: a parser exercise does not represent all repositories, dependency graphs or long-running agent sessions.

Interview practice

Can one faster agent run establish that its sandbox backend is faster?

No. Model behavior and workload differ unless controlled. Separate deterministic execution measurements from agent-quality experiments.

Why report failed trials in a performance study?

Time and reliability jointly determine usefulness. Excluding failures changes the question to performance conditional on success and must be disclosed.

Completion check

Write a protocol another engineer could execute without guessing. Leave the findings section explicitly “not measured” until raw records exist.

Sources and version notes

Checked 6 October 2026; current baseline: sbx v0.46.0. sbx create · Compute sizes and limits · Docker Sandboxes v0.46.0 release

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.