Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 13 / 30 · Operate the workload

Requests, limits and resource evidence

Connect scheduling reservations to runtime enforcement without treating limits as capacity.

4 min read + practiceWorked exerciseInterview practice

The mechanism

Resource requests inform scheduling and resource allocation. Limits constrain runtime consumption through the container and operating-system mechanisms available for that resource. They are not interchangeable: increasing a limit does not reserve more node capacity, and lowering a request can make placement look easier while increasing contention.

CPU and memory behave differently. CPU consumption can be throttled; memory pressure can lead to termination rather than a smooth slowdown. Interpret observed throttling, out-of-memory events and application latency separately.

QoS classification depends on resource configuration and influences behavior under pressure, but it is not a complete availability guarantee. Size ParcelOps from representative measurements and failure behavior rather than copying a large limit from another service.

Requested resources
Scheduler placement
Runtime consumption
Measured pressure

Worked example

This offline container fragment uses illustrative starting values. They are not measured recommendations. Two replicas request a combined 200 millicores and 256 MiB for these containers, before sidecars, Pod overhead or other workloads are considered.

resources:
  requests:
    cpu: 100m
    memory: 128Mi
  limits:
    cpu: 500m
    memory: 256Mi

Practice: predict, inspect, explain

Offline exercise. Calculate aggregate requests for two, three and five replicas. Then compare a CPU-bound workload with a memory leak under the same limits. Define which measurements would justify changing requests or limits: working set, peak demand, throttling, latency and termination history.

Expected observation: a resource configuration is a hypothesis. A larger limit can postpone a failure without fixing the leak, and a low request can create contention during busy periods. Keep proposed values labelled illustrative until a representative workload validates them.

Troubleshooting and trade-offs

If a Pod is unschedulable, compare requests with eligible node allocatable capacity, not only current usage. If an application slows down, inspect CPU throttling and dependencies before adding replicas. If it is OOM-killed, inspect memory demand and leaks rather than increasing limits indefinitely. Include every container when calculating the Pod’s resource footprint.

Interview practice

Why does the scheduler use requests?

Requests express the resources the workload asks the scheduler to account for during placement. Current usage alone would not provide a stable capacity plan.

Are CPU and memory limits equivalent?

No. CPU may be throttled, while memory exhaustion can cause termination. Their failure symptoms and tuning evidence differ.

Completion check

Calculate aggregate requested resources and explain two different runtime failure modes.

Sources and version notes

Baseline checked 6 October 2026: the official release page lists Kubernetes 1.37.1. Verify your cluster and distribution prerequisites. All manifests are offline teaching examples; no cluster mutations or cloud resources are executed by this handbook.

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.