Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 23 / 30 · Protect and diagnose

Logs, events, metrics and useful evidence

Use complementary signals and preserve timestamps and object identity.

4 min read + practiceWorked exerciseInterview practice

The mechanism

Logs describe application or component events, metrics summarize measurements over time, and Kubernetes events report notable object-related observations. Each has different retention and completeness properties. No single signal is a perfect history of everything that happened.

Begin with the user-visible symptom and a time window. Correlate Deployment revision, Pod UID, node, restart count and request identifiers where available. A log from a replacement Pod can otherwise be mistaken for the process that failed earlier.

Kubernetes exposes useful mechanisms, but a durable logging and monitoring system requires additional components and policy. Do not assume that deleting or replacing a Pod preserves all diagnostic output. Retain only the evidence needed, with credentials and private payloads redacted.

Symptom + time
Object identity
Complementary signals
Testable hypothesis

Worked example

These are optional read-only commands for an existing approved disposable lab. They may reveal application data, so inspect synthetic workloads only and redact evidence before sharing. The handbook does not run them against a cluster.

kubectl --context handbook-lab -n handbook-lab get pods -o wide
kubectl --context handbook-lab -n handbook-lab get events
kubectl --context handbook-lab -n handbook-lab logs deploy/parcelops --tail=100
# Choose an actual lab Pod before requesting its previous container logs.

Practice: predict, inspect, explain

Offline exercise. Create a synthetic timeline with a rollout, a failed readiness probe and a spike in request errors. Assign each observation to a signal source and record its timestamp. Add a missing log interval and state what remains unknown rather than inventing a cause.

Expected observation: correlation suggests hypotheses but does not prove causation. A useful incident note preserves source identity, time, query scope and gaps. Record whether a value came from a synthetic fixture, a saved observation or a proposed command.

Troubleshooting and trade-offs

If logs are absent, inspect container selection, restart history and retention rather than assuming no failure occurred. If metrics are unavailable, identify the missing collection pipeline. If events have expired, say so. Avoid collecting every environment variable or full Secret object as a diagnostic shortcut; observability should not create a new disclosure surface.

Interview practice

Why combine signals?

Logs, events and metrics reveal different aspects and have different gaps. Correlation across them supports stronger hypotheses than any one source alone.

What makes evidence reproducible?

The exact query or command, target context, time window, object identity, source version and retained redacted output, plus known gaps.

Completion check

Build a synthetic incident timeline with one explicit uncertainty and no secret values.

Sources and version notes

Baseline checked 6 October 2026: the official release page lists Kubernetes 1.37.1. Verify your cluster and distribution prerequisites. All manifests are offline teaching examples; no cluster mutations or cloud resources are executed by this handbook.

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.