Skip to lesson
supraj.dev THE ENGINEERING HANDBOOKS
LEARN / BUILD / VERIFY2026 edition · checked 06 Oct

CHAPTER 01 / 30 · Understand the cluster

A cluster is a collection of control loops

Trace desired state from an API object to a running workload and observed status.

4 min read + practiceWorked exerciseInterview practice

The mechanism

Kubernetes is a declarative control system. You submit desired state to an API; controllers observe the world and work toward that state. The API server is the front door, etcd stores cluster data, the scheduler chooses nodes for unscheduled Pods, and node agents run workloads through a container runtime.

These responsibilities are separate. The scheduler does not start containers, and the API server does not directly repair application bugs. A controller may create a replacement Pod, but the application still needs correct behavior, data handling and health signals.

Our running example is a fictional ParcelOps API with two replicas. Every lab begins with synthetic objects or read-only inspection. This handbook creates no cluster, changes no kubeconfig and spends no cloud resources. Optional live observations require your own approved disposable environment.

Desired object
API + stored state
Controller + scheduler
Node + observed status

Worked example

This conceptual trace shows one replica disappearing. The names are illustrative. Notice that replacement means a new Pod identity; Kubernetes does not resurrect an identical process with all of its previous in-memory state.

Deployment wants 2 replicas
ReplicaSet observes 1 matching Pod
Controller creates a replacement Pod object
Scheduler selects an eligible node
Kubelet asks the runtime to start its containers
Readiness eventually determines whether it can receive Service traffic

Practice: predict, inspect, explain

Offline exercise. Draw the components and assign each sentence in the trace to its responsible component. Then imagine the API is reachable but no node has enough requested memory. Explain which object can exist and why the application is not yet running.

Expected observation: an accepted API object is not a completed deployment. Your diagram should include observed status flowing back to the API. Add one application responsibility that Kubernetes cannot infer, such as whether a shipment transaction was committed correctly.

Troubleshooting and trade-offs

When a workload fails, identify the stage before changing configuration: admission, scheduling, image startup, process execution or readiness. A single “cluster broken” label loses useful evidence. If a replacement Pod starts but data disappears, inspect where the application stored that data; replacement does not preserve arbitrary container files.

Interview practice

What does the scheduler do?

It selects a suitable node for an unscheduled Pod using constraints and available requested resources. The kubelet and runtime on that node perform container execution.

Does accepted desired state prove availability?

No. Controllers, scheduling, startup and readiness still need to converge. Inspect status and application behavior before concluding the service is available.

Completion check

Explain a pending replacement Pod by tracing responsibilities across the control plane and node.

Sources and version notes

Baseline checked 6 October 2026: the official release page lists Kubernetes 1.37.1. Verify your cluster and distribution prerequisites. All manifests are offline teaching examples; no cluster mutations or cloud resources are executed by this handbook.

YOUR NEXT STEP

Make the understanding yours.

Use the completion check above. Mark this chapter when you can explain the mechanism and its limits.

Self-assessed reading progress. This does not certify that a lab ran or a system is secure.