Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 15 / 30 · Operate the workload

Probes that measure the right condition

Separate startup, readiness and liveness so recovery does not amplify failure.

4 min read + practiceWorked exerciseInterview practice

The mechanism

A startup probe gives a slow-starting container time to initialize before liveness and readiness probes take over. Readiness controls whether the workload should receive ordinary Service traffic. Liveness can trigger a container restart when the process is considered unhealthy.

These questions differ. A temporarily unavailable dependency may make an application unready without making a restart helpful. If every replica restarts whenever a shared database is slow, the health check can amplify an outage instead of repairing it.

Design ParcelOps endpoints around application semantics. A liveness endpoint should answer whether restarting is a useful recovery action; readiness should reflect ability to serve the traffic routed to that replica. Probe timing values need workload evidence.

Startup complete
Readiness for traffic
Liveness for restart
Observed recovery

Worked example

This offline fragment uses illustrative paths and timings. The fictional application must implement these endpoints before the configuration has meaning. These values are not production defaults or measured recommendations.

startupProbe:
  httpGet: {path: /startup, port: 8080}
  periodSeconds: 5
  failureThreshold: 12
readinessProbe:
  httpGet: {path: /ready, port: 8080}
  periodSeconds: 5
livenessProbe:
  httpGet: {path: /live, port: 8080}
  periodSeconds: 10
  failureThreshold: 3

Practice: predict, inspect, explain

Offline exercise. Classify a slow initialization, deadlocked process and temporary downstream timeout. Decide which probe should fail and what recovery should follow. Estimate the startup allowance from the sample settings while acknowledging scheduling and execution timing are not exact clocks.

Expected observation: a failing readiness probe removes traffic eligibility without requiring a restart. A liveness failure should correspond to a condition that a restart can improve. Include an application-level request check in acceptance testing; passing health endpoints alone does not prove business correctness.

Troubleshooting and trade-offs

If containers restart during normal startup, inspect the startup allowance and initialization distribution. If all replicas become unready together, investigate shared dependencies and whether readiness is too broad. If probe endpoints are expensive, they can add load during an incident. Keep diagnostic detail separate from a minimal health response.

Interview practice

Why not use the same deep dependency check for every probe?

The recovery actions differ. Restarting a healthy process because a shared dependency is unavailable can worsen the outage.

What does a startup probe protect?

It prevents normal startup time from being misclassified by liveness or readiness checks until startup succeeds, within a bounded allowance.

Completion check

Map three failure cases to the correct probe and justify the resulting action.

Sources and version notes

Baseline checked 6 October 2026: the official release page lists Kubernetes 1.37.1. Verify your cluster and distribution prerequisites. All manifests are offline teaching examples; no cluster mutations or cloud resources are executed by this handbook.

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.