Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 24 / 30 · Automation

Orchestrate create, exec, files and cleanup

Design for unknown outcomes, bounded retries and recoverable artifact transfer.

4 min readWorked exerciseInterview practice

What you will build

A fake-backend state machine that can be tested without Docker or a billed account. Its purpose is to make lifecycle bugs observable before connecting a real adapter.

The mechanism at a glance
  1. Durable job and request key
  2. Create / await / execute
  3. Validate and export output
  4. Delete / reconcile unknown state

Conceptual flow. Follow the lesson for prerequisites, exact commands and verification limits.

Read the mechanism

A network timeout is an observation about the caller’s wait, not proof that the server did nothing. A create request can succeed remotely while its response is lost. Blindly creating again can duplicate resources and cost.

Use one idempotency key for one exact request and preserve it during recovery where the API supports that mechanism. Resource mutation may also require the latest etag. A stale etag should trigger a fresh state read and reconciliation, not removal of the concurrency guard.

Process completion, output validation, export and cleanup are separate stages. If export fails after a successful task, deleting immediately may destroy the only result. If cleanup fails, the job must remain visible to an operator even when the artifact is already delivered.

Worked example · a controller contract

This TypeScript interface is application pseudocode, not the Docker SDK:

interface Backend {
  create(request: TaskRequest, key: string): Promise<Resource>;
  inspect(id: string): Promise<Resource>;
  execute(id: string, argv: string[], timeoutMs: number): Promise<ProcessResult>;
  readFile(id: string, absolutePath: string, maxBytes: number): Promise<Uint8Array>;
  delete(id: string, revision: string): Promise<void>;
}
type Phase =
  | 'planned' | 'creating' | 'ready' | 'running'
  | 'exporting' | 'cleanup-pending' | 'complete'
  | 'failed' | 'needs-reconciliation';

Persist the intended request and key before calling create. When an identity is returned, save it before execution. Export only allowed paths with byte limits, reject traversal and unexpected symlinks, and verify artifact hashes after transfer. A file’s existence does not establish its correctness.

Use a fake backend to inject these cases:

FailureRequired behavior
Create response lostRecover original request; no blind duplicate
Setup delayedBounded polling; distinguish VM from app readiness
Process exits nonzeroPreserve status and bounded stderr
Output too largeReject or truncate by explicit contract
Export interruptedRetain resource for retry within budget
Delete response lostInspect until confirmed or escalate unknown

Expected observations

The fake backend should record one logical resource after create-response recovery. A cleanup failure must leave a durable resource ID and an alertable reconciliation state. A process timeout must not silently become a successful exit.

Write down which behaviors the real API supports and which are controller responsibilities. Do not add guessed idempotencyKey fields to SDK calls without checking the pinned types and API reference.

Troubleshooting

Repeated retries can amplify a service outage. Use bounded attempts, backoff and a task deadline. Cancellation is also a transition: stop new actions, inspect in-flight work, preserve required evidence and reconcile cleanup.

Keep secrets out of durable job payloads. Store references to approved credential services rather than raw tokens.

Interview practice

What should happen after a lost create response?

Recover using the original exact request and its supported idempotency mechanism, then inspect current state. Do not assume failure and create a second resource.

Why keep cleanup pending after the user receives the output?

Delivery does not prove resource deletion. Remaining compute or authority needs a durable owner and reconciliation process.

Completion check

Demonstrate all six injected failures against a fake backend. Explain which transitions are safe to retry and which require inspection first.

Sources and version notes

Checked 6 October 2026; current baseline: sbx v0.46.0. Errors and retries · API concepts · Compute sizes and limits · SDK cookbook

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.