What you will build
A negative-test report with explicit hypotheses. Use disposable fixtures and authorized endpoints; this is not an exploit exercise or a reason to probe real host secrets.
- Identify asset and threat
- Choose harmless probe
- Observe decision and side effect
- Record residual uncertainty
Conceptual flow. Follow the lesson for prerequisites, exact commands and verification limits.
Read the mechanism
A threat model connects assets, adversaries, entry points and controls. “The agent is untrusted” is a starting assumption, not a complete model. A package install script, repository instruction, external document or MCP result can steer the agent or execute code.
For each boundary, ask what would count as a failure. A canary file can test a specific read/write path. A controlled endpoint can test whether a request arrived. A mock tool’s counter can test side effects. The observation must be independent enough to distinguish a denied request from a broken test.
Passing a finite set of tests does not prove immunity to prompt injection, hypervisor flaws or every indirect execution path. Report the tested scope, version and limitations.
Worked lab · a negative-test matrix
| Boundary | Harmless test | Expected result | Required evidence |
|---|---|---|---|
| Direct workspace | Append a comment to a fixture | Host change is visible | Before/after diff |
| Clone source | Attempt a fixture write to source mount | Write rejected | Error plus unchanged host hash |
| Clone confidentiality | Inspect synthetic ignored marker | May be readable | Record that read-only is not secret isolation |
| Mountless input | Inspect only a deliberately copied file | Copy visible; no live synchronization | Transfer manifest |
| Egress deny | Bounded request to controlled destination | Denied | Rule plus request/log correlation |
| MCP write | Invoke simulated write | Counter unchanged under deny | Policy decision plus counter |
| Cleanup | Remove disposable resource | Identity no longer exists | Confirmed terminal state |
The positive direct-write row is intentional: it prevents the false expectation that direct mode protects host integrity. Not every safety test should expect denial.
For a clone-source write test, use only a named synthetic marker and a fresh disposable repository. Do not attempt kernel escapes, socket access, private-directory enumeration or destructive payloads.
Expected observations
Each test record should contain: hypothesis, prerequisites, exact fixture, version, expected result, actual result, relevant log, independent side-effect observation and remaining uncertainty.
For example, a denied HTTPS test might record “policy check denied; bounded curl failed; matching proxy decision references rule X; controlled endpoint recorded no request.” Even that result covers only the chosen destination, protocol and configuration.
An unsupported test should remain “not tested.” A missing command or DNS failure is not a successful security control.
Troubleshooting
If a probe behaves unexpectedly, verify the fixture and path before changing policy. If the destination does not resolve, fix the test design rather than celebrating a blocked connection. If a log contains no match, inspect the proxy path and correlation window.
Review outputs before executing them on the host. A malicious build script exported as an artifact is an indirect path that simple filesystem-read tests do not cover.
Interview practice
Why include an expected-allow case in a security test suite?
It verifies the test harness and documents intentional capability. A suite in which every probe fails may simply be broken.
What does a passing canary test leave uncertain?
Other paths, versions, protocols, integrations and undiscovered vulnerabilities. State those limits rather than turning a scoped result into a universal security claim.
Completion check
Complete at least one positive and one negative case with independent evidence. Explain why the outcome supports a narrow claim and does not certify the whole system.
Sources and version notes
Checked 6 October 2026; current baseline: sbx v0.46.0. Security model · Isolation layers · Local policy · MCP access policies
Make the understanding yours.
Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.
Self-assessed reading progress. This does not certify that a lab ran or a system is secure.