What you will build
A Rithru Labs article draft and a claim audit. The draft below is deliberately framed as an evaluation plan because no handbook lab or benchmark has been executed for this publication.
- Choose a narrow claim
- Attach evidence to each claim
- Write findings and limits
- Review privacy and publish
Conceptual flow. Follow the lesson for prerequisites, exact commands and verification limits.
Read the mechanism
Enterprise readers need an operating decision: what problem a technique addresses, what authority it needs, what evidence supports it and where it stops being appropriate. Broad claims that an agent is now “safe” obscure the relevant boundaries.
Use four claim types: documented product behavior, measured observation, proposed application design and future work. Keep them visible while drafting. A proposed controller is not a Docker feature; a documented feature is not proof that your deployment has configured it correctly.
Every number needs a run record. Every version-sensitive product claim needs a current primary source. Every recommendation needs its assumptions. This turns writing into a final verification pass.
Worked draft · ready to complete after the POC
Evaluating isolated AI engineering with Docker Sandboxes
AI coding tools are most useful when they can run tests and build software. Those capabilities also create an authority problem: what files, services and credentials should an agent be allowed to use?
Docker Sandboxes offers a microVM-based execution boundary with an independent Docker Engine. Our proposed evaluation uses that boundary for a small parser repair, with explicit workspace sharing, constrained network access and a reviewed artifact export.
The evaluation separates three questions. First, can the agent produce the requested change? Second, do the configured boundaries behave as expected under harmless negative tests? Third, can the surrounding controller recover from interrupted execution or uncertain cleanup?
The proposed Rithru Sandbox Controller records task intent, resource identity, capability scope and lifecycle evidence. It treats the agent’s final message as a report to verify. A patch must satisfy the preserved acceptance tests, remain within scope and pass independent review before publication.
Results status: not measured. The planned comparison covers direct, clone and mountless local workspaces. Cloud evaluation is a separate stage because its credential, policy, port and lifecycle behavior differs. The experimental API is not treated as a stable production contract.
The key design constraint is explicit sharing. A direct workspace can be modified immediately. A read-only clone source can still disclose files. A host-run MCP tool carries host-side authority. Proxy-managed secrets reduce raw-value exposure while still allowing authenticated operations.
The next publication revision should replace this plan with sanitized evidence: input and output SHAs, version matrix, raw trial records, test results, negative-test outcomes, cleanup confirmation and the limits of the study. Until then, this is a reproducible evaluation approach, not a report of production adoption.
Expected observations
The finished draft should classify its claims clearly enough that a reviewer can tell a documented feature from a measured result or a proposed design. Before experiments, the correct result is an honest evaluation plan with no performance numbers. After experiments, every strengthened claim must point to the corresponding evidence packet.
Claim audit and evidence map
| Draft claim | Evidence required before stronger wording |
|---|---|
| Agent completed the parser fix | Reviewed patch plus independent test result |
| Policy blocked a request | Matching decision plus actual request evidence |
| Mode A was faster | Comparable repeated trials with spread and failures |
| Controller recovered correctly | Injected-failure trace and resource reconciliation |
| Enterprise-ready deployment | Defined requirements and verified controls, beyond this POC |
Link to the actual POC packet and benchmark dataset only after they exist. Remove private repository content, account details, student or family information and unredacted logs from public artifacts.
Troubleshooting
If a sentence cannot be traced to a source or run, narrow it or mark it as a proposal. If source docs conflict, describe the conflict and the version you validated. Do not use an attractive chart to conceal missing measurements.
Have a reviewer identify the strongest claim in the article and try to falsify it using the evidence packet. Their confusion is a signal to improve the article’s scope or supporting record.
Interview practice
What makes this article credible before measurements exist?
It clearly labels itself an evaluation plan, specifies the method and avoids invented results. Credibility comes from accurate scope, not the appearance of a completed benchmark.
What should change after real experiments?
Replace placeholders with traceable observations, retain failed trials and limitations, update the version matrix, and distinguish findings from recommendations.
Completion check
Audit every factual sentence and number. Then complete the scenario interview to connect the mechanisms across the whole handbook.
Sources and version notes
Checked 6 October 2026; current baseline: sbx v0.46.0. Docker Sandboxes · Docker Sandboxes API and SDK · Compare local and cloud sandboxes · AI Governance Audit Logs
Make the understanding yours.
Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.
Self-assessed reading progress. This does not certify that a lab ran or a system is secure.