The mechanism
Resource requests inform scheduling and resource allocation. Limits constrain runtime consumption through the container and operating-system mechanisms available for that resource. They are not interchangeable: increasing a limit does not reserve more node capacity, and lowering a request can make placement look easier while increasing contention.
CPU and memory behave differently. CPU consumption can be throttled; memory pressure can lead to termination rather than a smooth slowdown. Interpret observed throttling, out-of-memory events and application latency separately.
QoS classification depends on resource configuration and influences behavior under pressure, but it is not a complete availability guarantee. Size ParcelOps from representative measurements and failure behavior rather than copying a large limit from another service.
Worked example
This offline container fragment uses illustrative starting values. They are not measured recommendations. Two replicas request a combined 200 millicores and 256 MiB for these containers, before sidecars, Pod overhead or other workloads are considered.
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 256Mi
Practice: predict, inspect, explain
Offline exercise. Calculate aggregate requests for two, three and five replicas. Then compare a CPU-bound workload with a memory leak under the same limits. Define which measurements would justify changing requests or limits: working set, peak demand, throttling, latency and termination history.
Expected observation: a resource configuration is a hypothesis. A larger limit can postpone a failure without fixing the leak, and a low request can create contention during busy periods. Keep proposed values labelled illustrative until a representative workload validates them.
Troubleshooting and trade-offs
If a Pod is unschedulable, compare requests with eligible node allocatable capacity, not only current usage. If an application slows down, inspect CPU throttling and dependencies before adding replicas. If it is OOM-killed, inspect memory demand and leaks rather than increasing limits indefinitely. Include every container when calculating the Pod’s resource footprint.
Interview practice
Why does the scheduler use requests?
Requests express the resources the workload asks the scheduler to account for during placement. Current usage alone would not provide a stable capacity plan.
Are CPU and memory limits equivalent?
No. CPU may be throttled, while memory exhaustion can cause termination. Their failure symptoms and tuning evidence differ.
Completion check
Calculate aggregate requested resources and explain two different runtime failure modes.
Sources and version notes
Baseline checked 6 October 2026: the official release page lists Kubernetes 1.37.1. Verify your cluster and distribution prerequisites. All manifests are offline teaching examples; no cluster mutations or cloud resources are executed by this handbook.
- Official documentation: Manage resources containers
- Official documentation: Pod qos
- Official documentation: Assign pod node
Make the understanding yours.
Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.
Self-assessed reading progress. This does not certify that a lab ran or a system is secure.