Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

FINAL PRACTICE / EXPLAIN THE MECHANISM

The scenario interview.

Trace the failure, identify the boundary, choose a bounded check and state what remains unknown.

A useful five-minute answer

Try the case before opening the hint or answer. Explain your assumptions, propose an observable test and distinguish expected behavior from executed evidence. These are learning exercises, not certification questions.

Scheduling

01 · Plenty of capacity, but the Pod is Pending

The cluster has spare memory, yet no node accepts ParcelOps. How do you investigate?

Show a hint

Compute the intersection of requirements.

Reveal answer guidance

Inspect requests, node eligibility, affinity, taints and storage topology. Aggregate spare capacity is insufficient if no individual eligible node satisfies all constraints. Use scheduling events and a candidate-node table before relaxing a requirement.

Check your reasoning

  • Uses requests
  • Checks hard constraints
  • Avoids broad tolerations

Revisit: chapters 13 and 14.

Networking

02 · DNS resolves, but requests fail

The Service name resolves correctly, but the API cannot be reached. What is your next evidence?

Show a hint

Follow discovery to endpoints and the actual listener.

Reveal answer guidance

Inspect selectors and EndpointSlices, readiness, Service targetPort and the application listener, then source egress and destination ingress policy. DNS success only establishes name resolution. Distinguish refusal, timeout and application responses.

Check your reasoning

  • Checks endpoints
  • Traces both policy directions
  • Separates symptoms

Revisit: chapters 7–9 and 21.

Health

03 · A database slowdown restarts every replica

The liveness probe checks the database and all Pods restart during a shared outage. What should change?

Show a hint

Ask whether restarting can repair the condition.

Reveal answer guidance

Separate startup, readiness and liveness semantics. A shared dependency failure may justify unready status without restarting a healthy process. Use a liveness condition tied to recoverable process health and validate timing against workload behavior.

Check your reasoning

  • Separates probe actions
  • Avoids restart amplification
  • Requires application semantics

Revisit: chapters 5 and 15.

Network policy

04 · Default deny is present, but traffic is allowed

A namespace has default-deny and a second broad allow rule. Why might traffic still pass?

Show a hint

Policies are additive, and enforcement is external.

Reveal answer guidance

Verify CNI support and selection, then calculate the union of applicable allows for each direction. Default-deny does not subtract a broad allow. Test both permitted and prohibited flows in an approved disposable environment.

Check your reasoning

  • Checks enforcement
  • Explains additive rules
  • Tests both directions

Revisit: chapters 8, 21 and 22.

Availability

05 · Maintenance is blocked by the budget

Two replicas use minAvailable 1, but one is unready and eviction is blocked. Is the budget broken?

Show a hint

Count healthy replicas before the proposed disruption.

Reveal answer guidance

The remaining healthy replica may leave no disruption allowance. Investigate readiness and replacement capacity before maintenance. A PDB constrains supported voluntary disruptions; it does not prevent node failure or guarantee application availability.

Check your reasoning

  • Computes available budget
  • Preserves policy intent
  • Names limitations

Revisit: chapters 6, 15 and 16.

Identity

06 · A narrow Role still permits a sensitive action

An observer Role only lists Pods, but the user can read Secrets. How can that happen?

Show a hint

Review all grants and indirect capabilities.

Reveal answer guidance

RBAC permissions are additive. Another RoleBinding or ClusterRoleBinding may grant access. Inspect the exact identity, resource, verb and namespace across effective bindings, and review workload-creation or impersonation paths that can expose data indirectly.

Check your reasoning

  • Explains additive grants
  • Uses exact request attributes
  • Includes indirect authority

Revisit: chapters 11, 19 and 22.

Recovery

07 · The snapshot exists, but the application cannot recover

The team has an etcd backup but no verified database or volume restore. What is missing?

Show a hint

Map every authoritative data store.

Reveal answer guidance

Control-plane state is only part of recovery. Identify application databases, volumes, credentials, compatible infrastructure and restoration order. Run an approved separate restore exercise and validate business behavior before claiming recovery time or data-loss objectives.

Check your reasoning

  • Separates data stores
  • Requires restoration evidence
  • Avoids invented recovery claims

Revisit: chapters 12 and 27.

Evidence

08 · A local manifest test is called production proof

The synthetic workload lab passes and a report claims scheduling, isolation and performance are validated. How should it be corrected?

Show a hint

Match the claim to the executed mechanism.

Reveal answer guidance

Report only the selected offline invariants. API admission, scheduling, image startup, health, networking, identity and application outcomes remain unrun. Add a scoped integration plan and blank benchmark record, with prerequisites and acceptance criteria.

Check your reasoning

  • Bounds fixture coverage
  • Names missing runtime evidence
  • Keeps results blank

Revisit: chapters 28–30.

Your next experiment

Choose the case that exposed the largest gap. Return to the linked chapters and add a concrete test to your evidence packet.

Return to the learning path →