The mechanism
Horizontal Pod Autoscaling adjusts a workload’s replica count using configured metrics and policy. It is a control loop with sampling, stabilization and readiness considerations, not an instantaneous reaction to every request. Resource-metric scaling requires the appropriate metrics pipeline.
CPU utilization targets are interpreted relative to resource requests. Missing or poorly chosen requests can make the signal unavailable or misleading. More replicas also need schedulable capacity and compatible shared dependencies; autoscaling cannot fix a database bottleneck simply by multiplying callers.
For ParcelOps, choose a metric related to demand and service quality. CPU may be useful for compute-bound work, while queue depth or request pressure may better explain another workload. Verify adapter and API prerequisites before claiming a custom metric works.
Worked example
This offline HPA fragment shows the relationship between minimum, maximum and a CPU target. It is not applied. Metrics availability, Deployment requests and cluster capacity remain unverified prerequisites.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: parcelops
namespace: handbook-lab
spec:
scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: parcelops}
minReplicas: 2
maxReplicas: 5
metrics:
- type: Resource
resource:
name: cpu
target: {type: Utilization, averageUtilization: 70}
Practice: predict, inspect, explain
Offline exercise. Explain what 70 percent means for a container requesting 100 millicores. Then consider a missing metrics API, a full cluster and a slow shared database. For each case, predict why increasing desired replicas may fail to improve user latency.
Expected observation: scaling is a system property involving signals, policy, capacity and application behavior. Write acceptance criteria for successful scale-out and safe scale-in, including observed request errors and latency. Keep all throughput numbers blank until measured.
Troubleshooting and trade-offs
If HPA shows unknown metrics, inspect metrics availability and resource requests. If desired replicas rise but Pods stay Pending, inspect scheduling capacity. If replicas oscillate, review stabilization and the workload signal. Do not raise the maximum indefinitely or create paid node capacity as a handbook troubleshooting step.
Interview practice
What is utilization relative to?
For resource-utilization targets such as CPU, it is relative to configured resource requests, not simply a percentage of an entire node.
Why can scaling worsen a bottleneck?
More callers can increase pressure on a shared dependency without increasing its capacity. Measure the whole request path before assuming replicas solve latency.
Completion check
Explain the metric, policy, capacity and dependency prerequisites for one proposed scaling decision.
Sources and version notes
Baseline checked 6 October 2026: the official release page lists Kubernetes 1.37.1. Verify your cluster and distribution prerequisites. All manifests are offline teaching examples; no cluster mutations or cloud resources are executed by this handbook.
- Official documentation: Horizontal pod autoscale
- Official documentation: Manage resources containers
- Official documentation: Assign pod node
Make the understanding yours.
Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.
Self-assessed reading progress. This does not certify that a lab ran or a system is secure.