Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 16 / 30 · Operate the workload

Disruptions and availability budgets

Understand what a PodDisruptionBudget constrains and what it cannot prevent.

4 min read + practiceWorked exerciseInterview practice

The mechanism

A PodDisruptionBudget expresses how many selected replicas should remain available during supported voluntary disruption workflows. It is not a shield against hardware failure, every deletion path or application-level outage. Read the actual mechanism before treating it as an uptime guarantee.

Deployment rollout settings and disruption budgets address different operations. A rolling update has its own availability strategy. A node maintenance workflow using eviction interacts with the disruption budget. Both need enough healthy replicas and capacity to make progress.

For ParcelOps, two replicas with a minimum of one available can tolerate one eligible voluntary disruption under suitable conditions. If one replica is already unhealthy, the remaining budget may be exhausted. The operational response should investigate that state rather than bypassing the budget casually.

Healthy replicas
Disruption policy
Eviction decision
Replacement readiness

Worked example

This offline budget manifest selects ParcelOps Pods. It is not applied and no node is drained. The selector must match the intended workload; a valid object with the wrong selector protects the wrong set or nothing useful.

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: parcelops
  namespace: handbook-lab
spec:
  minAvailable: 1
  selector:
    matchLabels:
      app: parcelops

Practice: predict, inspect, explain

Offline exercise. Compute the conceptual disruption allowance with two healthy replicas, then one healthy replica, then zero. Explain how an ongoing outage changes a planned maintenance operation. Add a placement rule that keeps both replicas on one node and identify the remaining common failure risk.

Expected observation: replica count, health and failure-domain placement jointly affect availability. A budget cannot manufacture a healthy replacement. Record maintenance prerequisites and a stop condition instead of including destructive drain commands in the exercise.

Troubleshooting and trade-offs

If maintenance cannot evict a Pod, inspect budget status, selected replicas and readiness. If a workload is single-replica, a strict budget can intentionally block voluntary disruption until an availability plan exists. Do not delete the budget merely to unblock automation. Confirm whether the operation uses eviction and whether its behavior falls within the budget’s scope.

Interview practice

Does a PDB prevent node failure?

No. It constrains supported voluntary disruption paths. Involuntary failures and other application or infrastructure failures remain possible.

Why might a valid budget block maintenance?

Current healthy replicas may leave no allowed disruption. That can be the policy working as intended and indicates a capacity or health prerequisite.

Completion check

Explain the exhausted-budget case and name two availability risks outside PDB protection.

Sources and version notes

Baseline checked 6 October 2026: the official release page lists Kubernetes 1.37.1. Verify your cluster and distribution prerequisites. All manifests are offline teaching examples; no cluster mutations or cloud resources are executed by this handbook.

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.