Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 24 / 30 · Protect and diagnose

Debug from the first failed transition

Localize failure before changing the system and preserve evidence before recovery.

4 min read + practiceWorked exerciseInterview practice

The mechanism

A useful debugging sequence follows the workload lifecycle. Did the API accept the object? Was the Pod scheduled? Could the image start? Did the process remain alive? Did readiness succeed? Did the Service select the correct endpoints? Did an end-to-end request work?

Start with observation, then form a narrow hypothesis and choose the least disruptive check that can distinguish it. Restarting or deleting a Pod first can erase evidence and temporarily hide a configuration defect. A successful restart is not proof that the root cause is understood.

ParcelOps has a synthetic failure: the Service targets port 8080 while the application listens on 9090. Both objects can look valid. The packet path exposes the mismatch more clearly than changing resource limits or scaling replicas.

First failed stage
Competing hypotheses
Bounded observation
Evidence-based correction

Worked example

This diagnostic decision table is a paper exercise. It identifies the next evidence source rather than prescribing a destructive repair. Use saved fixtures when an approved lab is unavailable.

Pending + unschedulable event -> requests, constraints, eligible capacity
Waiting + image error -> image reference and pull prerequisites
Repeated termination -> exit reason and previous container logs
Running + unready -> probe path, listener and application dependency
Ready + no Service traffic -> selector, EndpointSlices, target port, policy
HTTP response + wrong result -> application logic and data provenance

Practice: predict, inspect, explain

Offline exercise. Walk the port mismatch through the table. Write two competing hypotheses and the observation that distinguishes them. Then repeat for a Pod blocked by a missing ConfigMap key. Keep the expected result of each check in the incident note before proposing a change.

Expected observation: the earliest failed transition narrows the search. A precise hypothesis avoids unrelated modifications that make the incident harder to explain. Record proposed remediation separately from an executed correction and subsequent verification.

Troubleshooting and trade-offs

If several changes happen at once, you may lose the ability to identify which fixed the issue. If an error disappears, retain the evidence and note whether the cause is proven or only suspected. Debug containers and exec sessions can expose powerful runtime access; they require their own authorization and are not needed for this offline handbook path.

Interview practice

Why preserve evidence before restarting?

Restarts can remove transient state and logs, obscure the original failure and make a recurring defect harder to diagnose.

What makes a good next diagnostic step?

It distinguishes specific competing hypotheses with minimal impact and a clear expected observation.

Completion check

Diagnose the port mismatch and missing-key cases without changing a live workload.

Sources and version notes

Baseline checked 6 October 2026: the official release page lists Kubernetes 1.37.1. Verify your cluster and distribution prerequisites. All manifests are offline teaching examples; no cluster mutations or cloud resources are executed by this handbook.

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.