Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 27 / 30 · Evidence

Separate control and execution planes

Design durable jobs, bounded authority and observability for a proposed production service.

4 min readWorked exerciseInterview practice

What you will build

A production-readiness proposal and a table-top incident response. The architecture is a design exercise; it does not assert that the experimental API offers a particular SLA or that this system is already deployed.

The mechanism at a glance
  1. Authenticated control plane
  2. Queue and capability policy
  3. Isolated execution plane
  4. Artifact review and cleanup reconciler

Conceptual flow. Follow the lesson for prerequisites, exact commands and verification limits.

Read the mechanism

The control plane accepts and authorizes jobs, stores durable intent and controls lifecycle. The execution plane runs untrusted work in bounded environments. Credential handling, policy distribution and artifact review are distinct responsibilities that should remain explicit.

A queue provides backpressure, not unlimited capacity. Per-user concurrency, task deadlines, output limits and spend limits prevent one workload from exhausting shared resources. Capability profiles should be versioned so an old job can be interpreted after policy changes.

A reconciler compares desired state with observed resources. It catches jobs whose worker died after creating a sandbox, unknown cleanup outcomes and canceled tasks that still consume capacity. Durable state is essential because in-memory finally blocks do not survive process termination.

Worked example · incident table-top

Walk through these events without touching production:

IncidentImmediate stateRecovery evidence
Agent hangsTask deadline exceededProcess stop outcome and bounded logs
Worker crashes after createResource may existDurable request key and discovered ID
Credential expiresAuth failureNarrow provider error, no token value
Policy propagation lagsEffective policy uncertainRevision and active-rule inspection
Artifact export failsTask result retained temporarilyRetry budget and source identity
Delete times outCleanup unknownReconciler record and operator alert

Define a metric set before choosing dashboards:

sandbox_create_latency
kit_setup_latency
task_execution_duration
task_failure_count{bounded_reason}
artifact_export_bytes
policy_denial_count{policy_revision}
cleanup_latency
orphan_resource_count
known_compute_cost
unknown_cost_run_count

These are proposed metric names. Avoid high-cardinality secrets, full prompts, URLs with tokens or arbitrary user text in metric labels. Keep a correlation ID linking restricted detailed records.

Expected observations

Every incident should end in either a confirmed terminal state or a visible recoverable state with an owner and deadline. “Unknown” is acceptable evidence when it remains tracked; silently converting unknown into success is not.

Write service objectives only after measuring a representative workload and failure modes. A retry policy should specify when to stop, what remains active and who is notified.

Troubleshooting

If the queue grows, inspect arrival rate, capacity, job duration and stuck cleanup separately. If metrics show success while users report missing artifacts, your acceptance condition may end at process exit instead of validated delivery.

If the system requires broad long-lived credentials to operate, reduce the supported task surface before expanding access.

Interview practice

Why is an orphan resource an operational concern even after a task fails?

It may retain data, credentials, execution authority and cost. Failure of the task does not establish termination of its resources.

What belongs in the control plane rather than the agent prompt?

Identity, authorization, budgets, resource ownership, state transitions, retry limits, retention and cleanup reconciliation.

Completion check

Present the plane boundaries, six incident responses and a metric-to-action mapping. Clearly separate proposed controls from verified product features.

Sources and version notes

Checked 6 October 2026; current baseline: sbx v0.46.0. Compute sizes and limits · Errors and retries · Local audit logs · Monitoring policies

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.