What you will build
A benchmark protocol and blank results ledger. This edition contains no measured Docker Sandbox performance results. All numeric result fields remain unfilled until a real authorized experiment supplies evidence.
- Predeclare protocol
- Run comparable repetitions
- Retain raw successful and failed trials
- Analyze with explicit limits
Conceptual flow. Follow the lesson for prerequisites, exact commands and verification limits.
Read the mechanism
“Time to result” includes several stages: template preparation, VM startup, kit setup, agent reasoning, build/test execution, export and cleanup. A faster total can reflect a warmer cache or a shorter model answer rather than a faster isolation backend.
Separate deterministic workload benchmarks from agent-task evaluations. A fixed parser test suite can measure execution overhead. An agent fixing the parser introduces variable planning, tool calls, retries and model behavior. Both are interesting, but they answer different questions.
Control hardware, architecture, image identity, source revision, resource caps, policy and cache state. When conditions cannot be identical, report them and avoid a causal claim the experiment cannot support.
Worked example · predeclare a protocol
Choose three workspace modes—direct, clone and mountless—and one deterministic build/test workload. Before execution, state the repetition count, warm-up procedure, timeout, exclusion rules and how failures will be recorded.
Use the blank benchmark CSV or the JSON run template. A minimal result table begins like this:
| Mode | Trials | Median task time | Spread | Failures | Cost coverage |
|---|---|---|---|---|---|
| Direct | Not run | — | — | — | Unknown |
| Clone | Not run | — | — | — | Unknown |
| Mountless | Not run | — | — | — | Unknown |
| Cloud, separate protocol | Not run | — | — | — | Unknown |
Use a monotonic clock for durations within one process. Define timestamps for ready, setup complete, task start/end, export end and confirmed cleanup. Do not subtract unsynchronized wall clocks from different machines as though they form one precise stopwatch.
Expected observations
After real trials, publish sample count and spread along with the central estimate. Include unsuccessful and timed-out runs in a separate visible category; silently dropping them can make an unreliable system look fast.
Cost can be known, estimated or unknown. Keep model charges separate from cloud compute and document the accounting window. A missing bill is not evidence of zero cost.
Useful findings can include: which phase dominates, whether warm state changes results, which failures recur, and which tasks cannot be compared because a backend lacks required behavior.
Troubleshooting
If timing starts after setup for one mode and before setup for another, fix the protocol before interpreting results. If one trial uses a different model or source revision, mark it non-comparable rather than quietly pooling it.
Do not claim statistical certainty from a tiny sample. Explain selection bias: a parser exercise does not represent all repositories, dependency graphs or long-running agent sessions.
Interview practice
Can one faster agent run establish that its sandbox backend is faster?
No. Model behavior and workload differ unless controlled. Separate deterministic execution measurements from agent-quality experiments.
Why report failed trials in a performance study?
Time and reliability jointly determine usefulness. Excluding failures changes the question to performance conditional on success and must be disclosed.
Completion check
Write a protocol another engineer could execute without guessing. Leave the findings section explicitly “not measured” until raw records exist.
Sources and version notes
Checked 6 October 2026; current baseline: sbx v0.46.0. sbx create · Compute sizes and limits · Docker Sandboxes v0.46.0 release
Make the understanding yours.
Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.
Self-assessed reading progress. This does not certify that a lab ran or a system is secure.