How to use this session
Spend five minutes on each case. Give your diagnosis, the first evidence you would inspect, the action you would avoid until clarified, and a recovery plan. These are qualitative practice prompts, not a certification exam.
01 · A sandboxed agent changed the host build
An agent completes a small bug fix in direct mode. The next host build unexpectedly runs a new script. A teammate says this proves the microVM failed. How do you investigate?
Reveal answer guidance
Identify the explicit read-write workspace first. Inspect the complete diff, dependency metadata, Makefile/package scripts, IDE tasks, agent instructions, untracked files and hooks without executing changed code. The agent may have modified a shared execution surface without crossing the microVM boundary. Preserve evidence, review from a clean checkout and narrow future sharing.
Listen for: Strong answers distinguish isolation from shared-file integrity and delayed execution. Weak answers assume a VM prevents every host effect.
Revisit: chapters 5 and 26.
02 · The private clone contains a secret
A repository ignores .env. The agent uses clone mode, yet can describe a synthetic secret marker placed in the host .env. Is that impossible?
Reveal answer guidance
No. Clone mode supplies a private writable clone while the host source is also mounted read-only at /run/sandbox/source. Ignored and untracked data can remain readable through that mount. Sanitize the source or use mountless transfer of approved files. Read-only is an integrity property, not a confidentiality guarantee.
Listen for: Strong answers identify the source mount and avoid equating Git tracking with visibility.
Revisit: chapters 6 and 7.
03 · Policy allowed, request failed
A host-and-port policy check says allowed. The HTTPS request fails. An engineer wants to reset all policy and try again. What evidence do you request?
Reveal answer guidance
Retain the active rules, source and version; correlate the exact request with the policy log. Destination checks do not validate HTTP method/path rules. Separate policy denial from DNS, TLS, upstream proxy and service failure. Do not reset global policy or broaden organization access to force success. Use a bounded request to a harmless controlled endpoint.
Listen for: Strong answers pair decision evidence with transport observations and know that reset changes persistent state.
Revisit: chapters 8 and 9.
04 · A read-only MCP tool writes to the host
The tool annotation says readOnly=true, but an approved local stdio server changes a host fixture. Why did sandbox isolation not prevent it?
Reveal answer guidance
The local server executes on the host outside the microVM. Its annotation is not a runtime proof. Review server identity and actual operations, constrain registration and use-time requests, test an independent side effect, and identify direct-client routes that bypass gateway governance. An OCI-packaged host server can also use host Docker.
Listen for: Strong answers place the server correctly and distinguish registration, loading, listing and invocation.
Revisit: chapters 18 and 19.
05 · A cloud create timed out
The API caller times out before receiving a resource name. The worker retries with a fresh request and creates two sandboxes. How should the controller recover?
Reveal answer guidance
Persist the exact request and its supported idempotency key before submission. A timeout is an unknown outcome. Recover the original request, inspect actual state and record any discovered identity before further work. Reconcile duplicates or orphan resources through an owned cleanup process. Keep stale-etag handling and bounded backoff explicit.
Listen for: Strong answers do not equate timeout with failure and keep durable ownership across worker crashes.
Revisit: chapters 23–25.
06 · The local demo becomes a cloud incident
A loopback-only local server is moved to cloud with the same assumptions. Its new endpoint is reachable publicly and the task leaves a running source sandbox. What was missing?
Reveal answer guidance
The migration preflight failed to compare port audience and lifecycle. Cloud publishing creates public HTTPS endpoints, and local policy/credential stores do not carry over as a promise of parity. A move copies filesystem state and leaves the source. Track both resources, require deliberate public exposure and verify TTL, deletion and budget controls.
Listen for: Strong answers include data destination, separate credentials, current cloud limitations and cleanup of both identities.
Revisit: chapters 22 and 27.
07 · Two patches pass alone and fail together
Parallel agents start from the same baseline and each passes tests. Their merged result fails. Does that mean the sandbox approach is ineffective?
Reveal answer guidance
No. Independent execution reduces some local races but does not prove semantic compatibility. Review both diffs, integrate serially and rerun the combined contract. Check shared external state, branch publication and resource contention too. The combined commit needs its own evidence.
Listen for: Strong answers distinguish isolation, concurrency and integration correctness.
Revisit: chapters 16 and 28.
08 · A persuasive but unsupported benchmark
An article says cloud is twice as fast based on one successful warm cloud run and one cold local run. Failed cloud attempts and model charges are omitted. How do you rewrite it?
Reveal answer guidance
Withdraw the general speed claim. Define comparable workloads and timing boundaries, report cache state, versions, sample count, spread, unsuccessful trials and cost coverage. Separate model behavior from deterministic runtime overhead. Until those records exist, publish a proposed evaluation method and state that performance is unmeasured.
Listen for: Strong answers connect every number to a raw run and describe what the evidence can actually establish.
Revisit: chapters 29 and 30.
A useful answer structure
- Asset: what data, identity, money or system state matters?
- Path: what crosses the boundary, and where does it execute?
- Control: which runtime or policy actually enforces the restriction?
- Evidence: what independent observation supports the diagnosis?
- Limit: what remains unknown after the proposed test?
The strongest answer is rarely “the sandbox makes it safe.” It explains a specific capability, the control around it and the evidence needed to trust the result.