PRACTICE TRACK / 131 QUESTIONS

Kubernetes
Think it through.

Pods, scheduling, networking, and the control plane.

Choose a question, explain your approach, then reveal the supplied answer. Difficulty labels come from the existing question library.

131 questions

Answers stay closed until you choose to reveal them.

QUESTION 01KubernetesEasy

What is the difference between a Deployment and a StatefulSet?

#
Reveal answer guidance

A Deployment manages stateless, interchangeable replicas with random pod names and a shared scaling behavior. A StatefulSet gives each pod a stable, ordered identity (pod-0, pod-1…), stable network IDs, and stable persistent volumes — used for databases, Kafka, and anything that needs predictable identity or ordered rollout.

QUESTION 02KubernetesEasy

What is a Service and what are the main types?

#
Reveal answer guidance

A Service gives a stable virtual IP/DNS name in front of a changing set of pods (selected by labels). Types: ClusterIP (internal only, default), NodePort (exposes a port on every node), LoadBalancer (provisions a cloud LB), and ExternalName (maps to a DNS CNAME). Ingress sits above Services for HTTP routing.

QUESTION 03KubernetesMedium

How does the scheduler decide where to place a pod?

#
Reveal answer guidance

In two phases: filtering (predicates) removes nodes that can't run the pod — insufficient resources, taints without matching tolerations, node selectors/affinity, volume topology. Then scoring ranks the remaining nodes (spreading, least-requested, affinity weights) and the highest-scoring node wins. You influence it with requests, nodeAffinity, podAffinity/anti-affinity, taints/tolerations, and topology spread constraints.

QUESTION 04Control planeMedium

What happens, step by step, when you run kubectl apply?

#
Show a hint

Follow desired state from the API server to the node.

Reveal answer guidance

kubectl sends the manifest to the API server, which authenticates, authorizes (RBAC), runs admission controllers, validates, and persists the desired state in etcd. The relevant controller (e.g. Deployment controller) notices the new desired state and creates ReplicaSets/Pods. The scheduler assigns each pod to a node. The kubelet on that node pulls images and starts containers via the container runtime, then reports status back to the API server.

The sequence described in this answer
  1. kubectl sends the manifest
  2. API server authenticates, authorizes, admits and persists
  3. Controller creates ReplicaSets and Pods
  4. Scheduler assigns a node
  5. Kubelet starts containers and reports status

Check your reasoning

  • Explain the API server checks before persistence.
  • Distinguish the controller, scheduler and kubelet responsibilities.
QUESTION 05TroubleshootingHard

A pod is stuck in CrashLoopBackOff. How do you debug it?

#
Reveal answer guidance

Start with kubectl describe pod to see events and the last state/exit code. kubectl logs <pod> --previous shows the crashed container's output. Common causes: bad command/entrypoint, missing config/secret, failing readiness vs liveness probe killing it too early, OOMKilled (check exit code 137 and memory limits), or a dependency it can't reach. Fix probes/limits, or run kubectl debug / an ephemeral container to inspect the filesystem and environment.

Command excerpt from the existing answer; replace the pod placeholder · shell
kubectl logs <pod> --previous
QUESTION 06KubernetesHard

A pod is running but users see 503 errors. How do you debug it?

#
Reveal answer guidance

A running pod only means the container is alive — not that the application is healthy. Start by checking: (1) Service endpoints — kubectl get endpoints to verify the pod is an endpoint target. (2) Pod labels vs Service selectors — mismatched labels mean traffic routes nowhere. (3) Readiness probes — if failing, the pod is removed from the Service even though it's running. (4) Ingress configuration — verify routing rules, TLS, and that the backend Service name/port match. (5) Application logs — the app may be crashing on each request or failing its startup dependency.

QUESTION 07KubernetesHard

How does pod-to-pod networking work without NAT?

#
Reveal answer guidance

Kubernetes requires a flat network where every pod gets a unique IP and can reach every other pod directly, no NAT. A CNI plugin (Calico, Cilium, etc.) implements this — assigning pod IPs and programming routes or an overlay (VXLAN/eBPF). Services are virtual: kube-proxy (or eBPF in Cilium) load-balances the Service ClusterIP to backing pod IPs via iptables/IPVS rules. NetworkPolicies then restrict which pods may talk to which.

QUESTION 08KubernetesEasy

What is the difference between a DaemonSet and a Deployment?

#
Reveal answer guidance

A DaemonSet runs one pod on every eligible node, or a subset selected by node labels/affinity, and is used for node-level agents like log collectors, Node Exporter, CNI agents, and CSI node plugins. A Deployment manages a replica count for stateless application pods. Modern DaemonSet pods still go through the scheduler, so taints, tolerations, affinity, and resource requests matter. Rolling updates are node-oriented and controlled with maxUnavailable or maxSurge depending on cluster support and configuration.

QUESTION 09KubernetesEasy

What is a Namespace and why would you use multiple namespaces?

#
Reveal answer guidance

A Namespace partitions a cluster into virtual sub-clusters. They provide scope for resource names (no DNS conflicts), RBAC boundaries (restrict who can access what), resource quotas (per-namespace CPU/memory caps), and NetworkPolicy isolation. Common patterns: separate namespaces per team (team-a, team-b), per environment (dev, staging, prod), or per application. Many objects (nodes, PVs) are cluster-scoped and don't belong to any namespace. kubectl get namespaces lists all.

QUESTION 10KubernetesEasy

What is the purpose of a ResourceQuota, and what resource types can it limit?

#
Reveal answer guidance

A ResourceQuota sets aggregate hard limits on resource consumption within a namespace. It prevents a single team from starving others. Limits include: compute (requests.cpu, requests.memory, limits.cpu, limits.memory), storage (requests.storage, persistentvolumeclaims), object counts (pods, services, configmaps, secrets, replicationcontrollers), and extended resources (GPUs, FPGAs). Exceeding a quota causes the API server to reject the operation with a 403 Forbidden. View with kubectl describe resourcequota -n <ns>.

QUESTION 11KubernetesEasy

How do you restrict a container's resource usage, and what happens when it exceeds its limit?

#
Reveal answer guidance

Set resources.requests (guaranteed minimum) and resources.limits (hard cap) per container. CPU is throttled above the limit; memory exceeding the limit causes OOMKill (container terminates, kubelet restarts it per restartPolicy). requests help the scheduler pick a node with enough capacity. The QoS class is determined by the ratio: Guaranteed (limits == requests), Burstable (requests < limits, or requests set but limits omitted), BestEffort (neither set). Use kubectl describe pod to see QoS class and OOM score.

QUESTION 12KubernetesEasy

What is a ConfigMap and how does it differ from a Secret?

#
Reveal answer guidance

Both store configuration data as key-value pairs. ConfigMap stores plaintext data (max 1 MiB) intended for non-sensitive config like env vars, config files, or command-line arguments. Secret stores sensitive data (base64-encoded in YAML, automatically decrypted when injected) with optional encryption at rest if the cluster has a KMS provider configured. Secrets support additional access control via RBAC. Best practice: use ConfigMap for app config, Secrets for passwords/tokens/keys. Both can be mounted as volumes or env vars.

QUESTION 13KubernetesMedium

How does kube-proxy work and what are the differences between iptables, IPVS, and eBPF modes?

#
Reveal answer guidance

kube-proxy watches the API server for Service and EndpointSlice changes, then programs network rules so traffic to a Service's ClusterIP reaches backing pods. iptables mode: creates a chain of iptables rules per Service — O(n) linear traversal for large clusters, updates require rewriting all rules (slow for 5000+ Services). IPVS mode (default in many distributions): uses Linux kernel IP Virtual Server tables, O(1) lookup per packet, much faster for large clusters. Supports more scheduling algorithms (round-robin, least-connection, source-hash). eBPF mode (via Cilium): replaces kube-proxy entirely, attaches BPF programs to network interfaces at the kernel level, bypassing iptables/IPVS entirely. Offers the lowest latency, per-packet policy enforcement, and native NetworkPolicy without separate iptables rules. Cilium's eBPF kube-proxy replacement is the modern recommended approach for greenfield clusters.

QUESTION 14KubernetesMedium

How do you restrict pod-to-pod traffic using NetworkPolicy? Walk through an example that blocks all ingress except from a specific label.

#
Reveal answer guidance

NetworkPolicy is a Kubernetes resource that selects pods via podSelector and defines ingress/egress rules. Default is 'allow all' unless a policy selects a pod — once any policy matches a pod, only traffic conforming to the rules is allowed. Example blocking all ingress except from pods labeled app: frontend: kind: NetworkPolicy; metadata: {name: deny-all-except-frontend}; spec: {podSelector: {matchLabels: {app: api}}, ingress: [{from: [{podSelector: {matchLabels: {app: frontend}}}]}]}. The from can specify podSelector, namespaceSelector, ipBlock, or any combination. egress rules restrict outbound traffic similarly. CNI support varies: Calico, Cilium, and Weave Net support full enforcement; some CNIs (Flannel) don't support NetworkPolicy at all and require a separate network policy engine.

QUESTION 15KubernetesMedium

What are init containers, and what problems do they solve compared to regular containers?

#
Reveal answer guidance

Init containers run sequentially before the app containers in a pod start. Each init container must exit successfully (exit code 0) before the next one starts. If any init container fails, Kubernetes restarts the entire pod. Use cases: (1) waiting for a dependency to be ready (e.g., until pg_isready; do sleep 1; done); (2) downloading or preparing files into shared volumes before the app starts; (3) running privileged setup operations (e.g., setting sysctls, modifying kernel params) while the app container runs with reduced capabilities. Init containers can have different resource limits, images, and volume mounts than app containers.

QUESTION 16KubernetesMedium

How do you implement canary deployments in Kubernetes without a service mesh?

#
Reveal answer guidance

Without a service mesh, the simplest weighted canary is two Deployments behind one Service selector that matches both versions, for example app: api, while the pods carry different version labels. Traffic share is roughly proportional to ready pod count, so 9 stable pods and 1 canary pod is about 10% for new connections. This is coarse and connection-dependent. For explicit HTTP percentages, use an ingress controller feature such as NGINX canary annotations, or a progressive delivery controller like Argo Rollouts or Flagger with an ingress provider. Keep rollback simple: scale canary to zero or remove it from the Service selector.

QUESTION 17KubernetesMedium

What is the difference between a PersistentVolume and a PersistentVolumeClaim, and how does dynamic provisioning add storage on demand?

#
Reveal answer guidance

A PersistentVolume (PV) is a storage resource in the cluster — an abstraction over actual storage (EBS, GCE PD, NFS, etc.). It has a capacity, access mode (ReadWriteOnce, ReadOnlyMany, ReadWriteMany), and reclaim policy (Retain, Delete, Recycle). A PersistentVolumeClaim (PVC) is a request for storage by a user — it specifies desired size and access mode. Kubernetes binds a matching PVC to a PV (or waits if none match). Dynamic provisioning uses a StorageClass: when a PVC references a StorageClass with a provisioner, the CSI driver automatically creates the underlying storage volume and a PV object. This eliminates manual PV pre-creation. The reclaimPolicy on the StorageClass determines what happens to the volume when the PVC is deleted.

QUESTION 18KubernetesMedium

How does the HorizontalPodAutoscaler (HPA) work, and what are the key parameters for tuning it?

#
Reveal answer guidance

The HPA controller reads CPU, memory, custom, or external metrics through Kubernetes metrics APIs and computes desired replicas roughly as ceil(currentReplicas * currentMetric / targetMetric). It ignores small changes inside the tolerance band, commonly 10%, to avoid flapping. Important tuning lives under spec.behavior: scale-up policies limit how fast replicas can increase, scale-down policies and scaleDown.stabilizationWindowSeconds prevent rapid scale-down. The usual default scale-down stabilization is 300 seconds, while scale-up stabilization is commonly 0 unless configured. For custom metrics, you need a working adapter such as Prometheus Adapter or Datadog Cluster Agent, and requests must be set correctly for CPU utilization-based scaling.

QUESTION 19KubernetesMedium

What is a PodDisruptionBudget, and when should you use one for a stateful application?

#
Reveal answer guidance

A PodDisruptionBudget (PDB) limits the number of replicas of a set that can be voluntarily disrupted at a time. It does NOT protect against involuntary disruptions (hardware failure, kernel panic). PDBs are configured with minAvailable (e.g., minAvailable: 80%) or maxUnavailable (e.g., maxUnavailable: 1). Before an eviction (e.g., node drain via kubectl drain, cluster autoscaler scale-down), the disruption controller checks if evicting the pod would violate the PDB. If so, the eviction is blocked. For stateful applications (databases, message queues), set a PDB with maxUnavailable: 1 to ensure quorum during rolling updates or node maintenance. The PDB interacts with the cluster autoscaler: if a node needs to drain and the PDB blocks eviction, the autoscaler may fail to scale down.

QUESTION 20KubernetesMedium

How does DNS resolution work inside a Kubernetes cluster? Describe the full lookup path for curl myservice.myns.svc.cluster.local.

#
Reveal answer guidance

CoreDNS (default since K8s 1.13) runs as a cluster DNS Service (typically kube-dns or coredns). A pod's /etc/resolv.conf points nameserver to the CoreDNS ClusterIP and sets search domains: <pod-namespace>.svc.cluster.local, svc.cluster.local, cluster.local, <cloud-region>.compute.internal. For curl myservice.myns.svc.cluster.local, the OS resolver queries CoreDNS, which matches the cluster domain and returns the ClusterIP. CoreDNS uses Headless Service endpoints directly (A/AAAA records for each pod IP via endpoint slices). For ExternalName Services, CoreDNS returns a CNAME. The search domains allow short names: curl myservice resolves to myservice.<pod-ns>.svc.cluster.local. CoreDNS supports stub domains (custom upstreams), prometheus metrics, and autoscaling via cluster-proportional-autoscaler. Troubleshoot with kubectl run dnsutils --image=gcr.io/kubernetes-e2e-test-images/dnsutils:1.3 then nslookup kubernetes.default.

QUESTION 21KubernetesHard

How does etcd work under the hood — leader election, Raft consensus, quorum, and data storage?

#
Reveal answer guidance

etcd is a distributed key-value store based on the Raft consensus algorithm. Leader election: peers elect a leader via randomized timer-based voting. The leader handles all client writes and sends heartbeats (default 100ms) to followers. Quorum: a write succeeds only if a majority of nodes (N/2 + 1) acknowledge it. For a 3-node cluster, 2 nodes must confirm; writes are durable once committed to the majority. Reads: by default, reads are served from the leader's state machine (linearizable, consistent). With --quorum-backend-read, followers can serve reads but may return stale data. Data storage: etcd stores data in a persistent bbolt database on disk. Keys are organized in a B+ tree; writes go to a write-ahead log (WAL) for durability, then the in-memory tree, then snapshots. MVCC: each modification creates a new revision; the --compact flag reclaims old revisions. The Kubernetes API server stores all cluster state in etcd under /registry/. Snapshot size grows with cluster state; etcd defragmentation (etcdctl defrag) reclaims space. Best practices: 3-5 nodes, SSD storage, regular backups (etcdctl snapshot save), TLS client-to-server and peer-to-peer encryption.

QUESTION 22KubernetesHard

A node is in NotReady state. Walk through the full debugging process from kubectl to kernel diagnostics.

#
Reveal answer guidance

Step 1: kubectl describe node <node> and inspect Conditions, events, taints, and last heartbeat. Ready=Unknown usually means the control plane stopped hearing from kubelet; Ready=False means kubelet reported a problem. Also check DiskPressure, MemoryPressure, and PIDPressure. Step 2: on the node, check systemctl status kubelet, journalctl -u kubelet, container runtime health with crictl ps, and certificate or auth errors. Step 3: check resources and kernel signals: df -h, inode usage, free -h, CPU load, process count, CNI logs, routes, firewall rules, and DNS/API reachability. Step 4: verify control-plane health through component pods or managed-service status rather than relying on deprecated kubectl get cs. Recover by fixing the specific issue, cordoning/draining if needed, or replacing the node. Delete the Node object only when you understand whether the underlying machine should rejoin or be rebuilt.

QUESTION 23KubernetesHard

How would you design a Kubernetes cluster for multi-tenancy with strong workload isolation across tenants, including network, compute, and RBAC isolation?

#
Reveal answer guidance

There is no single multi-tenancy feature in K8s — it's a layered approach. (1) Namespaces per tenant: each tenant gets one or more namespaces with a ResourceQuota and LimitRange. (2) RBAC: create a ClusterRole with limited permissions (read-only on nodes, ability to create Deployments/Services in own namespace) and bind it via RoleBinding in each tenant namespace. Use ClusterRoleBinding only for cluster-wide access. (3) Network isolation: a default-deny NetworkPolicy in every namespace plus permissive ingress/egress policies per tenant. For cross-tenant isolation, use a CNI with native policy enforcement (Cilium, Calico). Cilium supports CiliumNetworkPolicy with layer 7 rules (HTTP methods, paths). (4) Compute isolation: use ResourceQuota per tenant namespace for hard caps. For CPU/memory guarantees, set requests == limits (Guaranteed QoS) for critical workloads. Use RuntimeClass (gVisor, Kata Containers) for untrusted tenant workloads — each tenant pod runs in a lightweight VM for kernel isolation. (5) Storage isolation: each tenant PV/PVC should be in their namespace; CSI drivers can enforce volume ownership. (6) vCluster: for true control-plane isolation, each tenant gets a virtual cluster (loft.sh/vcluster) — a full K8s API server backed by the host cluster. This adds overhead but provides near-complete isolation. No single approach is perfect; choose isolation level based on threat model and complexity budget.

QUESTION 24KubernetesHard

What happens when a node runs out of PID pressure or inode pressure? How does the kubelet detect and react to these conditions?

#
Reveal answer guidance

The kubelet monitors system-level pressure indicators via its node condition controller. By default, it checks nodefs.inodesFree (inode pressure) and pid.available (PID pressure) via the cgroup stats. PID pressure: if the node's PID limit (kernel kernel.pid_max, typically 32768) is reached, the kernel cannot fork new processes — containers crash with "resource temporarily unavailable". The kubelet sets NodeCondition.PIDPressure = True and begins evicting pods (lowest priority first, using the same eviction order as memory/disk). Inode pressure: depletion of inodes means no new files can be created (even if disk space is available). Causes: a rogue container writing millions of small log files or a Docker overlay2 graph driver accumulating orphan layers. Detection: nodefs.inodesFree based on kubelet's eviction signals. Reaction: kubelet sets DiskPressure = True and evicts pods. Monitoring: kubectl top node doesn't show PID/inode metrics — use node_exporter Prometheus metrics (node_namespaces_pid, node_filesystem_files_free). Prevention: set kubelet --eviction-hard=nodefs.inodesFree<5%,pid.available<10% or similar thresholds. Pod-level PID limiting via --pod-max-pids (kubelet flag) or cgroup v2 pids.max controller.

QUESTION 25KubernetesHard

How does the Kubernetes API server handle concurrent writes and what conflict resolution mechanism prevents lost updates?

#
Reveal answer guidance

The API server uses optimistic concurrency based on resourceVersion (a monotonically increasing integer from etcd's mod revision). When a client creates an object, it gets a resourceVersion. On update, the PUT request must include the resourceVersion of the object being updated. The API server compares the resourceVersion in the request against the current version stored in etcd — if they differ (meaning another client updated the object in the meantime), the API server returns a 409 Conflict error. The client must re-fetch the object and retry the operation. This prevents lost updates: without it, two concurrent clients could overwrite each other's changes silently. The kubectl apply command handles this internally via server-side apply (SSA) — it uses PATCH with application/apply-patch+yaml content type and the field manager pattern. SSA tracks which field manager owns each field, so overlapping patches merge correctly rather than overwrite. The resourceVersion check is also used for watches: the watch request includes resourceVersion and the API server streams events only from that version onward, ensuring no events are missed. The watch reconnects with the latest version on timeout. For high-contention resources, use kubectl replace (replace entire object) or strategic merge patch (surgical field update without needing full object). The fieldValidation flag (K8s 1.25+) rejects unknown fields at the API server level.

QUESTION 26KubernetesHard

What is the difference between a mutating admission webhook and a validating admission webhook? Walk through the full request lifecycle.

#
Reveal answer guidance

Admission webhooks intercept API requests after authentication/authorization but before persistence. MutatingAdmissionWebhook can modify the object in flight: it receives an AdmissionReview with the full object, can patch it, and returns the modified version. Use cases: injecting sidecars (Istio, Linkerd), setting default resource limits, adding node selectors, validating that namespaces have correct labels. ValidatingAdmissionWebhook cannot modify the object — it only accepts or rejects. Use cases: enforcing policies (OPA/Gatekeeper, Kyverno), checking compliance (e.g., forbid latest image tags, require specific labels). Full lifecycle: (1) Authentication/Authorization (RBAC); (2) MutatingAdmissionWebhooks (run in order, can modify object); (3) Object schema validation; (4) ValidatingAdmissionWebhooks; (5) ResourceQuota and LimitRanger admission; (6) Persist to etcd. Webhooks receive: AdmissionReview JSON with request.operation (CREATE, UPDATE, DELETE, CONNECT), request.object (new state), request.oldObject (for UPDATE/DELETE). Response: response.allowed = true/false, response.patch (base64-encoded JSON patch, mutating only), response.status.message (error message on rejection). Control with failurePolicy: Ignore (cluster survives if webhook is down) or FailurePolicy: Fail (cluster blocks requests if webhook unreachable). reinvocationPolicy: IfNeeded re-runs mutating webhooks after other objects are created/modified. Ordering: MutatingWebhookConfiguration rules have failurePolicy and matchPolicy; ordering within multiple mutating webhooks is not guaranteed in K8s < 1.15 (fixed with webhook ordering via names). TLS: webhooks require a TLS server certificate signed by a CA the API server trusts. Deploy webhooks as a Service backed by pods; the API server connects to the webhook Service endpoint. The kube-apiserver or aggregator can be configured with --admission-control-config-file for fine-grained webhook routing. Use objectSelector to restrict which objects trigger the webhook (label-based filtering) to reduce latency overhead. Watch out for deadlock: a mutating webhook that modifies a ConfigMap that the webhook itself reads can cause infinite loops. Prevent by setting reinvocationPolicy: Never.

QUESTION 27KubernetesHard

How do you migrate workloads from one node pool to another in a production cluster with zero downtime?

#
Reveal answer guidance

Zero-downtime migration requires draining nodes gradually while ensuring enough capacity and respecting PDBs. Step 1: cordon all source pool nodes (kubectl cordon <node>) — no new pods are scheduled. Step 2: drain nodes one at a time (kubectl drain <node> --ignore-daemonsets --delete-emptydir-data). The --ignore-daemonsets flag keeps DaemonSet pods running (they'll be rescheduled by the DaemonSet controller on the new pool). --delete-emptydir-data acknowledges loss of ephemeral data. Step 3: before draining each node, verify the target pool has enough allocatable capacity — use kubectl describe nodes on target pool nodes. Step 4: ensure PodDisruptionBudgets are configured so stateful workloads aren't evicted below quorum. Step 5: for stateful workloads with PVCs, ensure the new pool has the same storage class and topology zones — otherwise PVC binding fails. Use azcopy or Velero for PV migration if the new pool is in a different zone. Step 6: monitor rollout with kubectl get pods -o wide -w and dashboards. Step 7: after all pods have migrated, scale down the source pool (cloud provider) or delete nodes. Alternatives: (a) PDB-respecting eviction API — write a script that iterates pods and calls kubectl evict which respects PDBs. (b) For cluster autoscaler managed pools: taint the source pool with CriticalAddonsOnly=true:NoSchedule and set up a temporary node group on the target pool. (c) Use descheduler to proactively move pods from source pool based on label/taint-based eviction strategies.

QUESTION 28KubernetesHard

What is the role of the cloud-controller-manager (CCM), and how does it interact with nodes, Services, and routes?

#
Reveal answer guidance

The cloud-controller-manager is a Kubernetes controller that embeds cloud-specific control logic into the control plane. It runs as a set of loops: Node controller: watches nodes and annotates them with cloud provider info (instance type, region, zone, provider ID). When a node is unreachable, the CCM checks the cloud API to determine if the VM is healthy — if deleted, the CCM deletes the Kubernetes node object. This prevents orphaned node objects. Route controller: configures cloud networking routes so pods on different nodes can communicate without overlay (used with legacy cloud provider route-based networking, e.g., GCP routes, AWS VPC route tables). For modern clusters using CNI (Calico, Cilium), the route controller is unnecessary. Service controller: provisions cloud load balancers (AWS NLB/ALB, GCP TCP/HTTP LB, Azure LB) when a Service of type LoadBalancer is created. It also manages health checks, security group rules (AWS), and firewall rules (GCP). CCM communicates with the cloud API via IAM/IRSA credentials. It does NOT manage compute resources (VM creation) — that's the cluster autoscaler or cloud provider. In-tree cloud providers (gce, aws, azure) were removed upstream in K8s 1.27+; all clusters must now run out-of-tree CCMs. The CCM interacts with the API server via normal informers and workqueues. Failure of the CCM means LoadBalancer Services stop provisioning and node lifecycle events (instance termination) aren't handled.

QUESTION 29KubernetesHard

You deploy a ClusterIP Service, but pods can't reach it by name or ClusterIP. Walk through the full debugging flow.

#
Reveal answer guidance

Step 1: verify the Service exists and has endpoints: kubectl get svc my-svc and kubectl get endpoints my-svc. If endpoints are empty, the pod selector doesn't match — check kubectl describe svc my-svc for Selector. Step 2: verify pods are running and have the matching labels: kubectl get pods --show-labels. Step 3: inside a pod, test connectivity: kubectl exec <pod> -- curl -v <cluster-ip>:<port>. If curl hangs (no route), check the service chain. Step 4: check kube-proxy: kubectl logs -n kube-system kube-proxy-<hash>. If kube-proxy is not running, iptables/IPVS rules aren't updated. Step 5: on a node, check iptables rules: iptables-save | grep my-svc or ipvsadm -L -n. For IPVS: verify the virtual server exists with correct pod endpoints. Step 6: check kernel conntrack: conntrack -L | grep <cluster-ip>. Conntrack entries can be stale. Step 7: check net.ipv4.ip_forward=1 on all nodes. Step 8: check bridge-nf-call-iptables=1 (sysctl net.bridge.bridge-nf-call-iptables) — required for traffic through bridge interfaces to be filtered by iptables. Step 9: check CNI plugin health — kubectl get pods -n kube-system | grep cilium|calico|weave. A misconfigured CNI drops packets before they reach kube-proxy rules. Step 10: test Service DNS: kubectl exec <pod> -- nslookup my-svc. If DNS fails, check CoreDNS: kubectl logs -n kube-system -l k8s-app=kube-dns. Step 11: enable kube-proxy debug logging: edit the DaemonSet, add --v=4 to the kube-proxy args. This prints all iptables updates.

QUESTION 30KubernetesHard

How does Kubernetes garbage collection work for owner references? What happens when a parent object is deleted while child objects are still being created?

#
Reveal answer guidance

Owner references (metadata.ownerReferences) establish parent-child relationships between API objects. When a parent object is deleted, the garbage collector (GC) automatically deletes all child objects. The GC runs as a controller in the kube-controller-manager. Deletion modes: Foreground (default for some resources): GC first deletes all dependents (or adds orphanDependents finalizer to parent), then deletes the parent. The parent remains visible with a deletionTimestamp and a foregroundDeletion finalizer until all children are deleted. Background (default): GC propagates deletion concurrently — parent is deleted immediately, children are deleted by a background worker. Orphan: parent is deleted without deleting children (set deleteOptions.orphanDependents=true). When a parent is deleted while children are being created: if the child is created after the parent's deletionTimestamp, the GC still sees the owner reference (via the parent's UID) and deletes the child. Race condition: if a parent is deleted and a controller creates a child referencing the deleted parent, the child may appear briefly and then be garbage-collected. Use DeletionPropagation with foreground to avoid orphan races. The GC controller uses informers with caches; there's a small window between parent deletion and child deletion (bounded by GC sync period, default 1h for some resourced). Immediate deletion: set --controllers=garbagecollector --concurrent-gc-syncs=20 on the controller-manager. The GC uses UID references, not name references — so renaming a parent doesn't affect children. Shared owners: a single child can have multiple owner references (e.g., a custom resource owned by both a user and a system); all owners must exist for the child to survive. Cascade deletion of Jobs creates Pods: deleting a Job deletes its Pods via the owner reference (set by the Job controller). Prevent accidental cascade: use kubectl delete --cascade=orphan for specific cleanup scenarios.

QUESTION 31KubernetesHard

Explain the Kubernetes Control Plane high availability architecture. How does leader election work for the controller-manager and scheduler?

#
Reveal answer guidance

An HA control plane runs multiple replicas of the API server, controller-manager, and scheduler, typically across 3 or 5 availability zones. API servers are stateless — they front etcd which runs as a Raft cluster (3 or 5 nodes). API servers are load-balanced via a TCP load balancer or round-robin DNS. Controller-manager and scheduler are stateless but use leader election — only one instance actively operates at a time. Leader election mechanism: each replica creates a Lease object (or uses Endpoint in older versions) in the kube-system namespace. The Lease has a holderIdentity, leaseDurationSeconds (default 15s), renewTime, and acquireTime. Each candidate tries to update the Lease with its identity using an optimistic lock (resourceVersion). The first to succeed becomes leader; it renews the Lease before leaseDurationSeconds expires. If the leader fails, other candidates acquire the Lease after leaseDeadline (default 15s) + retryPeriod (default 2s). The scheduler and controller-manager each have separate Leases. Config: --leader-elect=true (default), --leader-elect-lease-duration=15s, --leader-elect-retry-period=3s, --leader-elect-renew-deadline=10s. The API server does NOT use leader election — all API server instances are active simultaneously. They share a single etcd cluster. The API server's cache (watch-based informer) ensures consistency across instances. etcd itself runs a separate Raft leader election with --election-timeout (default 1000ms, i.e., 1 second). The API server only talks to the etcd leader via the etcd client's built-in redirect. If the etcd leader fails, a new leader is elected in ~1 second, and API writes are briefly blocked during this window.

QUESTION 32KubernetesHard

Kubernetes incident: A Deployment rollout reaches 100% updated replicas but users get intermittent 503s. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check kubectl rollout status, ReplicaSet history, Service endpoints, readiness probe results, app metrics, and ingress/controller logs. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Use readiness probes that prove real serving health, set maxUnavailable: 0 and a controlled maxSurge for critical APIs, keep revisionHistoryLimit for rollback, and gate rollout on user-facing metrics. Do not treat updated replicas as proof of healthy traffic; verify endpoint membership and error rates.

QUESTION 33KubernetesHard

Kubernetes architecture scenario: Design a deployment strategy for stateless APIs that supports fast rollback and zero-downtime releases. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Use readiness probes that prove real serving health, set maxUnavailable: 0 and a controlled maxSurge for critical APIs, keep revisionHistoryLimit for rollback, and gate rollout on user-facing metrics. Do not treat updated replicas as proof of healthy traffic; verify endpoint membership and error rates.

QUESTION 34KubernetesMedium

Kubernetes security scenario: Developers can patch live Deployments directly and bypass the approved release path. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Use readiness probes that prove real serving health, set maxUnavailable: 0 and a controlled maxSurge for critical APIs, keep revisionHistoryLimit for rollback, and gate rollout on user-facing metrics. Do not treat updated replicas as proof of healthy traffic; verify endpoint membership and error rates.

QUESTION 35KubernetesHard

Kubernetes release scenario: Move a critical service from recreate deploys to rolling or canary deploys. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Use readiness probes that prove real serving health, set maxUnavailable: 0 and a controlled maxSurge for critical APIs, keep revisionHistoryLimit for rollback, and gate rollout on user-facing metrics. Do not treat updated replicas as proof of healthy traffic; verify endpoint membership and error rates.

QUESTION 36KubernetesMedium

Kubernetes reliability/cost scenario: Rollouts take too long because readiness gates and surge settings are conservative. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Use readiness probes that prove real serving health, set maxUnavailable: 0 and a controlled maxSurge for critical APIs, keep revisionHistoryLimit for rollback, and gate rollout on user-facing metrics. Do not treat updated replicas as proof of healthy traffic; verify endpoint membership and error rates.

QUESTION 37KubernetesHard

Kubernetes incident: Pods restart during startup because the liveness probe fails before the application is ready. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check kubectl describe pod, previous container logs, probe events, application startup timings, and endpoint readiness changes. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Use startupProbe to protect slow boot, readiness to control traffic, and liveness only for unrecoverable deadlocks. Health endpoints should be cheap, deterministic, and sanitized. Tune periodSeconds, timeoutSeconds, and failure thresholds from measured startup and dependency behavior.

QUESTION 38KubernetesHard

Kubernetes architecture scenario: Design health checks for an application with slow warmup and external dependencies. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Use startupProbe to protect slow boot, readiness to control traffic, and liveness only for unrecoverable deadlocks. Health endpoints should be cheap, deterministic, and sanitized. Tune periodSeconds, timeoutSeconds, and failure thresholds from measured startup and dependency behavior.

QUESTION 39KubernetesMedium

Kubernetes security scenario: Health endpoints expose dependency names, versions, or internal errors. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Use startupProbe to protect slow boot, readiness to control traffic, and liveness only for unrecoverable deadlocks. Health endpoints should be cheap, deterministic, and sanitized. Tune periodSeconds, timeoutSeconds, and failure thresholds from measured startup and dependency behavior.

QUESTION 40KubernetesHard

Kubernetes release scenario: Introduce startup, readiness, and liveness probes to an existing production service. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Use startupProbe to protect slow boot, readiness to control traffic, and liveness only for unrecoverable deadlocks. Health endpoints should be cheap, deterministic, and sanitized. Tune periodSeconds, timeoutSeconds, and failure thresholds from measured startup and dependency behavior.

QUESTION 41KubernetesMedium

Kubernetes reliability/cost scenario: Aggressive probes create restart loops and increase load during partial outages. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Use startupProbe to protect slow boot, readiness to control traffic, and liveness only for unrecoverable deadlocks. Health endpoints should be cheap, deterministic, and sanitized. Tune periodSeconds, timeoutSeconds, and failure thresholds from measured startup and dependency behavior.

QUESTION 42KubernetesHard

Kubernetes incident: Ingress returns 404 or 502 after a rule change, while the backend Service works inside the cluster. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check Ingress status, controller logs, generated load balancer config, backend Service ports, TLS secret, DNS records, and external health checks. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Separate DNS, certificate, routing, and backend health checks. Validate host/path matching and backend protocol, use staged DNS weights or parallel hostnames for migration, and restrict risky annotations with admission policy. Monitor both controller errors and external synthetic checks.

QUESTION 43KubernetesHard

Kubernetes architecture scenario: Design external HTTP routing with TLS, path routing, and safe multi-team ownership. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Separate DNS, certificate, routing, and backend health checks. Validate host/path matching and backend protocol, use staged DNS weights or parallel hostnames for migration, and restrict risky annotations with admission policy. Monitor both controller errors and external synthetic checks.

QUESTION 44KubernetesMedium

Kubernetes security scenario: Ingress annotations allow unsafe snippets or expose internal services publicly. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Separate DNS, certificate, routing, and backend health checks. Validate host/path matching and backend protocol, use staged DNS weights or parallel hostnames for migration, and restrict risky annotations with admission policy. Monitor both controller errors and external synthetic checks.

QUESTION 45KubernetesHard

Kubernetes release scenario: Migrate traffic from one ingress controller or gateway to another. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Separate DNS, certificate, routing, and backend health checks. Validate host/path matching and backend protocol, use staged DNS weights or parallel hostnames for migration, and restrict risky annotations with admission policy. Monitor both controller errors and external synthetic checks.

QUESTION 46KubernetesMedium

Kubernetes reliability/cost scenario: TLS renewal or load balancer changes cause avoidable downtime. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Separate DNS, certificate, routing, and backend health checks. Validate host/path matching and backend protocol, use staged DNS weights or parallel hostnames for migration, and restrict risky annotations with admission policy. Monitor both controller errors and external synthetic checks.

QUESTION 47KubernetesHard

Kubernetes incident: A newly applied default-deny NetworkPolicy breaks database access for multiple services. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check CNI policy logs, flow logs, kubectl describe networkpolicy, DNS tests, connection tests from affected pods, and namespace labels. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Start with namespace-level default-deny in lower environments, then add explicit ingress and egress for DNS, telemetry, and known dependencies. Use CNI flow visibility such as Cilium Hubble or Calico flow logs. Roll out in audit or narrow namespaces first and avoid policies that accidentally block CoreDNS.

QUESTION 48KubernetesHard

Kubernetes architecture scenario: Design namespace network isolation that still allows DNS, metrics, and required dependencies. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Start with namespace-level default-deny in lower environments, then add explicit ingress and egress for DNS, telemetry, and known dependencies. Use CNI flow visibility such as Cilium Hubble or Calico flow logs. Roll out in audit or narrow namespaces first and avoid policies that accidentally block CoreDNS.

QUESTION 49KubernetesMedium

Kubernetes security scenario: Pods can egress to the internet and call services outside their ownership boundary. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Start with namespace-level default-deny in lower environments, then add explicit ingress and egress for DNS, telemetry, and known dependencies. Use CNI flow visibility such as Cilium Hubble or Calico flow logs. Roll out in audit or narrow namespaces first and avoid policies that accidentally block CoreDNS.

QUESTION 50KubernetesHard

Kubernetes release scenario: Move from allow-all networking to default-deny policies. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Start with namespace-level default-deny in lower environments, then add explicit ingress and egress for DNS, telemetry, and known dependencies. Use CNI flow visibility such as Cilium Hubble or Calico flow logs. Roll out in audit or narrow namespaces first and avoid policies that accidentally block CoreDNS.

QUESTION 51KubernetesMedium

Kubernetes reliability/cost scenario: NetworkPolicy troubleshooting is slow because denied flows are not observable. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Start with namespace-level default-deny in lower environments, then add explicit ingress and egress for DNS, telemetry, and known dependencies. Use CNI flow visibility such as Cilium Hubble or Calico flow logs. Roll out in audit or narrow namespaces first and avoid policies that accidentally block CoreDNS.

QUESTION 52KubernetesHard

Kubernetes incident: Pods on different nodes cannot communicate after a node image or CNI upgrade. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check CNI DaemonSet health, node routes, eBPF maps or iptables rules, MTU, kubelet CNI errors, and node firewall settings. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Upgrade the CNI as critical infrastructure: one node pool or AZ at a time, with rollback manifests and node-level validation. Check MTU, route programming, policy enforcement, and kernel compatibility. Restrict write access to CNI resources and alert on agent restarts or datapath errors.

QUESTION 53KubernetesHard

Kubernetes architecture scenario: Choose and operate a CNI for a multi-zone production cluster. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Upgrade the CNI as critical infrastructure: one node pool or AZ at a time, with rollback manifests and node-level validation. Check MTU, route programming, policy enforcement, and kernel compatibility. Restrict write access to CNI resources and alert on agent restarts or datapath errors.

QUESTION 54KubernetesMedium

Kubernetes security scenario: The CNI grants broad host privileges and its DaemonSet can be modified by many users. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Upgrade the CNI as critical infrastructure: one node pool or AZ at a time, with rollback manifests and node-level validation. Check MTU, route programming, policy enforcement, and kernel compatibility. Restrict write access to CNI resources and alert on agent restarts or datapath errors.

QUESTION 55KubernetesHard

Kubernetes release scenario: Upgrade the CNI plugin without losing pod networking. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Upgrade the CNI as critical infrastructure: one node pool or AZ at a time, with rollback manifests and node-level validation. Check MTU, route programming, policy enforcement, and kernel compatibility. Restrict write access to CNI resources and alert on agent restarts or datapath errors.

QUESTION 56KubernetesMedium

Kubernetes reliability/cost scenario: Packet drops and conntrack pressure cause intermittent service failures. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Upgrade the CNI as critical infrastructure: one node pool or AZ at a time, with rollback manifests and node-level validation. Check MTU, route programming, policy enforcement, and kernel compatibility. Restrict write access to CNI resources and alert on agent restarts or datapath errors.

QUESTION 57KubernetesHard

Kubernetes incident: Pods intermittently fail DNS lookups during traffic spikes. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check CoreDNS logs and metrics, resolv.conf, query rate, cache hit ratio, upstream latency, NodeLocal DNSCache status, and failed lookup samples. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Scale CoreDNS by query rate, enable caching, consider NodeLocal DNSCache, and avoid excessive search-domain expansion from short names. Test stub domain changes in a canary CoreDNS deployment or single cluster first. Use network policy and DNS policy carefully because DNS is a shared dependency.

QUESTION 58KubernetesHard

Kubernetes architecture scenario: Design DNS resolution for high-query clusters and latency-sensitive services. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Scale CoreDNS by query rate, enable caching, consider NodeLocal DNSCache, and avoid excessive search-domain expansion from short names. Test stub domain changes in a canary CoreDNS deployment or single cluster first. Use network policy and DNS policy carefully because DNS is a shared dependency.

QUESTION 59KubernetesMedium

Kubernetes security scenario: Workloads can resolve or query internal domains they should not access. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Scale CoreDNS by query rate, enable caching, consider NodeLocal DNSCache, and avoid excessive search-domain expansion from short names. Test stub domain changes in a canary CoreDNS deployment or single cluster first. Use network policy and DNS policy carefully because DNS is a shared dependency.

QUESTION 60KubernetesHard

Kubernetes release scenario: Change CoreDNS configuration for stub domains or conditional forwarding. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Scale CoreDNS by query rate, enable caching, consider NodeLocal DNSCache, and avoid excessive search-domain expansion from short names. Test stub domain changes in a canary CoreDNS deployment or single cluster first. Use network policy and DNS policy carefully because DNS is a shared dependency.

QUESTION 61KubernetesMedium

Kubernetes reliability/cost scenario: CoreDNS CPU rises and application latency follows DNS timeout patterns. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Scale CoreDNS by query rate, enable caching, consider NodeLocal DNSCache, and avoid excessive search-domain expansion from short names. Test stub domain changes in a canary CoreDNS deployment or single cluster first. Use network policy and DNS policy carefully because DNS is a shared dependency.

QUESTION 62KubernetesHard

Kubernetes incident: HPA scales up too late during a traffic burst and scales down while requests are still queued. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check HPA status, metrics-server or adapter logs, Prometheus queries, queue depth, pod startup time, requests/limits, and workload latency. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Use the metric that matches demand: CPU for CPU-bound APIs, RPS or latency for services, and queue depth or lag for workers. Account for pod startup time with min replicas and scale-up policies. Use stabilization windows to prevent flapping, and protect metrics endpoints and labels from leaking sensitive data.

QUESTION 63KubernetesHard

Kubernetes architecture scenario: Design autoscaling for APIs, workers, and queue consumers. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Use the metric that matches demand: CPU for CPU-bound APIs, RPS or latency for services, and queue depth or lag for workers. Account for pod startup time with min replicas and scale-up policies. Use stabilization windows to prevent flapping, and protect metrics endpoints and labels from leaking sensitive data.

QUESTION 64KubernetesMedium

Kubernetes security scenario: Autoscaling metrics expose tenant or customer identifiers. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Use the metric that matches demand: CPU for CPU-bound APIs, RPS or latency for services, and queue depth or lag for workers. Account for pod startup time with min replicas and scale-up policies. Use stabilization windows to prevent flapping, and protect metrics endpoints and labels from leaking sensitive data.

QUESTION 65KubernetesHard

Kubernetes release scenario: Move from CPU-only HPA to custom or external metrics. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Use the metric that matches demand: CPU for CPU-bound APIs, RPS or latency for services, and queue depth or lag for workers. Account for pod startup time with min replicas and scale-up policies. Use stabilization windows to prevent flapping, and protect metrics endpoints and labels from leaking sensitive data.

QUESTION 66KubernetesMedium

Kubernetes reliability/cost scenario: Autoscaling increases cost but still misses latency SLOs. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Use the metric that matches demand: CPU for CPU-bound APIs, RPS or latency for services, and queue depth or lag for workers. Account for pod startup time with min replicas and scale-up policies. Use stabilization windows to prevent flapping, and protect metrics endpoints and labels from leaking sensitive data.

QUESTION 67KubernetesHard

Kubernetes incident: Pods remain Pending even though cluster autoscaler is enabled. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check Pending pod events, autoscaler logs, node taints, affinity, topology spread, PDBs, quotas, and cloud capacity errors. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Pending pods must be schedulable onto some allowed node shape. Check requests, taints/tolerations, affinity, topology constraints, max node group size, cloud quotas, and capacity availability. Use separate pools for critical, batch, GPU, and spot workloads, with quotas and PDBs so scaling decisions do not violate availability.

QUESTION 68KubernetesHard

Kubernetes architecture scenario: Design node pools for mixed workloads, GPUs, spot capacity, and critical services. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Pending pods must be schedulable onto some allowed node shape. Check requests, taints/tolerations, affinity, topology constraints, max node group size, cloud quotas, and capacity availability. Use separate pools for critical, batch, GPU, and spot workloads, with quotas and PDBs so scaling decisions do not violate availability.

QUESTION 69KubernetesMedium

Kubernetes security scenario: Workloads can force expensive node scale-ups by requesting large resources. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Pending pods must be schedulable onto some allowed node shape. Check requests, taints/tolerations, affinity, topology constraints, max node group size, cloud quotas, and capacity availability. Use separate pools for critical, batch, GPU, and spot workloads, with quotas and PDBs so scaling decisions do not violate availability.

QUESTION 70KubernetesHard

Kubernetes release scenario: Introduce Karpenter or change node pool instance families. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Pending pods must be schedulable onto some allowed node shape. Check requests, taints/tolerations, affinity, topology constraints, max node group size, cloud quotas, and capacity availability. Use separate pools for critical, batch, GPU, and spot workloads, with quotas and PDBs so scaling decisions do not violate availability.

QUESTION 71KubernetesMedium

Kubernetes reliability/cost scenario: Scale-down evicts useful pods or is blocked by disruption constraints. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Pending pods must be schedulable onto some allowed node shape. Check requests, taints/tolerations, affinity, topology constraints, max node group size, cloud quotas, and capacity availability. Use separate pools for critical, batch, GPU, and spot workloads, with quotas and PDBs so scaling decisions do not violate availability.

QUESTION 72KubernetesHard

Kubernetes incident: A node reports MemoryPressure or DiskPressure and evicts healthy pods. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check kubectl describe node, kubelet eviction events, cAdvisor metrics, container logs size, image GC state, and emptyDir usage. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Set realistic CPU, memory, and ephemeral-storage requests. Use limits for risky workloads, log rotation, image garbage collection, and namespace quotas. Evictions are based on node pressure and QoS/priority, so protect critical workloads with PriorityClass, Guaranteed or well-sized Burstable QoS, and enough node reserve.

QUESTION 73KubernetesHard

Kubernetes architecture scenario: Design node sizing and eviction thresholds for production workloads. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Set realistic CPU, memory, and ephemeral-storage requests. Use limits for risky workloads, log rotation, image garbage collection, and namespace quotas. Evictions are based on node pressure and QoS/priority, so protect critical workloads with PriorityClass, Guaranteed or well-sized Burstable QoS, and enough node reserve.

QUESTION 74KubernetesMedium

Kubernetes security scenario: Pods can consume node disk through logs or emptyDir and affect other tenants. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Set realistic CPU, memory, and ephemeral-storage requests. Use limits for risky workloads, log rotation, image garbage collection, and namespace quotas. Evictions are based on node pressure and QoS/priority, so protect critical workloads with PriorityClass, Guaranteed or well-sized Burstable QoS, and enough node reserve.

QUESTION 75KubernetesHard

Kubernetes release scenario: Introduce ephemeral-storage requests and limits across namespaces. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Set realistic CPU, memory, and ephemeral-storage requests. Use limits for risky workloads, log rotation, image garbage collection, and namespace quotas. Evictions are based on node pressure and QoS/priority, so protect critical workloads with PriorityClass, Guaranteed or well-sized Burstable QoS, and enough node reserve.

QUESTION 76KubernetesMedium

Kubernetes reliability/cost scenario: Overcommit improves utilization but causes noisy-neighbor evictions. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Set realistic CPU, memory, and ephemeral-storage requests. Use limits for risky workloads, log rotation, image garbage collection, and namespace quotas. Evictions are based on node pressure and QoS/priority, so protect critical workloads with PriorityClass, Guaranteed or well-sized Burstable QoS, and enough node reserve.

QUESTION 77KubernetesHard

Kubernetes incident: A StatefulSet rollout is stuck because one ordinal never becomes Ready. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check StatefulSet events, PVC/PV binding, pod ordinal logs, readiness gates, quorum state, topology spread, and storage backend health. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Respect ordered identity and quorum. Back up first, upgrade one ordinal at a time, validate replication health, and understand whether the application supports parallel rollout. Use pod anti-affinity or topology spread, zone-aware storage, scoped secrets, and tested restore procedures.

QUESTION 78KubernetesHard

Kubernetes architecture scenario: Design stateful workloads with stable identity, storage, backups, and failover. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Respect ordered identity and quorum. Back up first, upgrade one ordinal at a time, validate replication health, and understand whether the application supports parallel rollout. Use pod anti-affinity or topology spread, zone-aware storage, scoped secrets, and tested restore procedures.

QUESTION 79KubernetesMedium

Kubernetes security scenario: Database pods run with broad filesystem permissions and shared credentials. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Respect ordered identity and quorum. Back up first, upgrade one ordinal at a time, validate replication health, and understand whether the application supports parallel rollout. Use pod anti-affinity or topology spread, zone-aware storage, scoped secrets, and tested restore procedures.

QUESTION 80KubernetesHard

Kubernetes release scenario: Upgrade a database StatefulSet image or configuration safely. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Respect ordered identity and quorum. Back up first, upgrade one ordinal at a time, validate replication health, and understand whether the application supports parallel rollout. Use pod anti-affinity or topology spread, zone-aware storage, scoped secrets, and tested restore procedures.

QUESTION 81KubernetesMedium

Kubernetes reliability/cost scenario: Stateful pods concentrate in one zone and fail during zonal impairment. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Respect ordered identity and quorum. Back up first, upgrade one ordinal at a time, validate replication health, and understand whether the application supports parallel rollout. Use pod anti-affinity or topology spread, zone-aware storage, scoped secrets, and tested restore procedures.

QUESTION 82KubernetesHard

Kubernetes incident: Pods fail to start because PVCs are Pending or volumes cannot attach. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check PVC/PV status, StorageClass binding mode, CSI controller logs, VolumeAttachment objects, node zone labels, and cloud volume events. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Use WaitForFirstConsumer for zonal disks so scheduling and volume creation agree on topology. Check access modes, attach limits, reclaim policy, and CSI health. For migration, snapshot or replicate data, create new PVCs intentionally, validate application consistency, and avoid deleting retained PVs until restore is proven.

QUESTION 83KubernetesHard

Kubernetes architecture scenario: Design Kubernetes storage for RWO, RWX, snapshots, and disaster recovery. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Use WaitForFirstConsumer for zonal disks so scheduling and volume creation agree on topology. Check access modes, attach limits, reclaim policy, and CSI health. For migration, snapshot or replicate data, create new PVCs intentionally, validate application consistency, and avoid deleting retained PVs until restore is proven.

QUESTION 84KubernetesMedium

Kubernetes security scenario: Applications mount volumes with permissions broader than required. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Use WaitForFirstConsumer for zonal disks so scheduling and volume creation agree on topology. Check access modes, attach limits, reclaim policy, and CSI health. For migration, snapshot or replicate data, create new PVCs intentionally, validate application consistency, and avoid deleting retained PVs until restore is proven.

QUESTION 85KubernetesHard

Kubernetes release scenario: Migrate workloads to a new StorageClass or CSI driver. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Use WaitForFirstConsumer for zonal disks so scheduling and volume creation agree on topology. Check access modes, attach limits, reclaim policy, and CSI health. For migration, snapshot or replicate data, create new PVCs intentionally, validate application consistency, and avoid deleting retained PVs until restore is proven.

QUESTION 86KubernetesMedium

Kubernetes reliability/cost scenario: Volume attach limits or zone mismatch blocks scheduling during failover. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Use WaitForFirstConsumer for zonal disks so scheduling and volume creation agree on topology. Check access modes, attach limits, reclaim policy, and CSI health. For migration, snapshot or replicate data, create new PVCs intentionally, validate application consistency, and avoid deleting retained PVs until restore is proven.

QUESTION 87KubernetesHard

Kubernetes incident: A service account gets Forbidden in production but works with cluster-admin in dev. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check kubectl auth can-i, audit logs, Role/ClusterRole rules, RoleBinding subjects, service account tokens, and controller error logs. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Start from observed API verbs/resources and grant only what is needed. Prefer namespaced Roles when possible, reserve ClusterRoles for cluster-scoped resources, and test with kubectl auth can-i --as=system:serviceaccount:ns:name. Roll out RBAC separately from workload changes so authorization failures are easy to isolate.

QUESTION 88KubernetesHard

Kubernetes architecture scenario: Design least-privilege RBAC for controllers, apps, and platform teams. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Start from observed API verbs/resources and grant only what is needed. Prefer namespaced Roles when possible, reserve ClusterRoles for cluster-scoped resources, and test with kubectl auth can-i --as=system:serviceaccount:ns:name. Roll out RBAC separately from workload changes so authorization failures are easy to isolate.

QUESTION 89KubernetesMedium

Kubernetes security scenario: A namespace RoleBinding grants access to secrets or privileged resources unintentionally. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Start from observed API verbs/resources and grant only what is needed. Prefer namespaced Roles when possible, reserve ClusterRoles for cluster-scoped resources, and test with kubectl auth can-i --as=system:serviceaccount:ns:name. Roll out RBAC separately from workload changes so authorization failures are easy to isolate.

QUESTION 90KubernetesHard

Kubernetes release scenario: Replace cluster-admin permissions with scoped Roles. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Start from observed API verbs/resources and grant only what is needed. Prefer namespaced Roles when possible, reserve ClusterRoles for cluster-scoped resources, and test with kubectl auth can-i --as=system:serviceaccount:ns:name. Roll out RBAC separately from workload changes so authorization failures are easy to isolate.

QUESTION 91KubernetesMedium

Kubernetes reliability/cost scenario: RBAC changes break controllers during a production deploy. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Start from observed API verbs/resources and grant only what is needed. Prefer namespaced Roles when possible, reserve ClusterRoles for cluster-scoped resources, and test with kubectl auth can-i --as=system:serviceaccount:ns:name. Roll out RBAC separately from workload changes so authorization failures are easy to isolate.

QUESTION 92KubernetesHard

Kubernetes incident: A validating webhook outage blocks all pod creates in the cluster. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check admission webhook configuration, API server audit logs, webhook service endpoints, timeout settings, failurePolicy, and policy violation reports. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Admission webhooks must be highly available, fast, and scoped. Use failurePolicy: Fail only for controls that must block and have reliable webhook capacity; use namespace selectors and short timeouts. Prefer built-in Pod Security Admission for baseline restrictions and use policy engines for custom rules with staged enforcement.

QUESTION 93KubernetesHard

Kubernetes architecture scenario: Design admission controls for image policy, security contexts, and required labels. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Admission webhooks must be highly available, fast, and scoped. Use failurePolicy: Fail only for controls that must block and have reliable webhook capacity; use namespace selectors and short timeouts. Prefer built-in Pod Security Admission for baseline restrictions and use policy engines for custom rules with staged enforcement.

QUESTION 94KubernetesMedium

Kubernetes security scenario: Teams can deploy privileged containers, hostPath mounts, or images from untrusted registries. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Admission webhooks must be highly available, fast, and scoped. Use failurePolicy: Fail only for controls that must block and have reliable webhook capacity; use namespace selectors and short timeouts. Prefer built-in Pod Security Admission for baseline restrictions and use policy engines for custom rules with staged enforcement.

QUESTION 95KubernetesHard

Kubernetes release scenario: Move from PodSecurityPolicy-era controls to Pod Security Admission or policy engines. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Admission webhooks must be highly available, fast, and scoped. Use failurePolicy: Fail only for controls that must block and have reliable webhook capacity; use namespace selectors and short timeouts. Prefer built-in Pod Security Admission for baseline restrictions and use policy engines for custom rules with staged enforcement.

QUESTION 96KubernetesMedium

Kubernetes reliability/cost scenario: Admission latency slows deployments and causes API server request timeouts. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Admission webhooks must be highly available, fast, and scoped. Use failurePolicy: Fail only for controls that must block and have reliable webhook capacity; use namespace selectors and short timeouts. Prefer built-in Pod Security Admission for baseline restrictions and use policy engines for custom rules with staged enforcement.

QUESTION 97KubernetesHard

Kubernetes incident: A rotated Secret does not update running pods and the application keeps using old credentials. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check Secret version, pod env/volume mounts, app reload behavior, External Secrets status, RBAC access, and audit logs. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Environment variable secrets do not update until pod restart; mounted Secret volumes update eventually but applications must reload them. Use external secret managers, scoped service accounts, encryption at rest, and rollout automation for apps that cannot reload. Rotate with overlapping credentials where the backend supports it.

QUESTION 98KubernetesHard

Kubernetes architecture scenario: Design secret delivery and rotation for Kubernetes workloads. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Environment variable secrets do not update until pod restart; mounted Secret volumes update eventually but applications must reload them. Use external secret managers, scoped service accounts, encryption at rest, and rollout automation for apps that cannot reload. Rotate with overlapping credentials where the backend supports it.

QUESTION 99KubernetesMedium

Kubernetes security scenario: Secrets are exposed through environment variables, logs, or broad RBAC. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Environment variable secrets do not update until pod restart; mounted Secret volumes update eventually but applications must reload them. Use external secret managers, scoped service accounts, encryption at rest, and rollout automation for apps that cannot reload. Rotate with overlapping credentials where the backend supports it.

QUESTION 100KubernetesHard

Kubernetes release scenario: Move from static Kubernetes Secrets to External Secrets or CSI secret delivery. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Environment variable secrets do not update until pod restart; mounted Secret volumes update eventually but applications must reload them. Use external secret managers, scoped service accounts, encryption at rest, and rollout automation for apps that cannot reload. Rotate with overlapping credentials where the backend supports it.

QUESTION 101KubernetesMedium

Kubernetes reliability/cost scenario: Secret rotation causes connection storms and failed logins. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Environment variable secrets do not update until pod restart; mounted Secret volumes update eventually but applications must reload them. Use external secret managers, scoped service accounts, encryption at rest, and rollout automation for apps that cannot reload. Rotate with overlapping credentials where the backend supports it.

QUESTION 102KubernetesHard

Kubernetes incident: Controllers lag and kubectl calls intermittently time out during high churn. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check API server latency metrics, request rates by user-agent, audit logs, APF queues, watch counts, controller logs, and etcd latency. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Use informers and watches instead of polling, scope watches by namespace or label where possible, set client-side rate limits, and identify noisy user agents. Secure access with short-lived credentials and least privilege. For APF, classify critical controllers so noisy batch clients cannot starve control-plane operations.

QUESTION 103KubernetesHard

Kubernetes architecture scenario: Design controller and operator interactions that do not overload the API server. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Use informers and watches instead of polling, scope watches by namespace or label where possible, set client-side rate limits, and identify noisy user agents. Secure access with short-lived credentials and least privilege. For APF, classify critical controllers so noisy batch clients cannot starve control-plane operations.

QUESTION 104KubernetesMedium

Kubernetes security scenario: Clients use long-lived admin kubeconfigs outside controlled systems. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Use informers and watches instead of polling, scope watches by namespace or label where possible, set client-side rate limits, and identify noisy user agents. Secure access with short-lived credentials and least privilege. For APF, classify critical controllers so noisy batch clients cannot starve control-plane operations.

QUESTION 105KubernetesHard

Kubernetes release scenario: Introduce a new operator that watches many resources cluster-wide. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Use informers and watches instead of polling, scope watches by namespace or label where possible, set client-side rate limits, and identify noisy user agents. Secure access with short-lived credentials and least privilege. For APF, classify critical controllers so noisy batch clients cannot starve control-plane operations.

QUESTION 106KubernetesMedium

Kubernetes reliability/cost scenario: API Priority and Fairness throttles low-priority clients during incidents. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Use informers and watches instead of polling, scope watches by namespace or label where possible, set client-side rate limits, and identify noisy user agents. Secure access with short-lived credentials and least privilege. For APF, classify critical controllers so noisy batch clients cannot starve control-plane operations.

QUESTION 107KubernetesHard

Kubernetes incident: A CRD upgrade breaks an operator and existing custom resources fail validation. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check CRD versions, storedVersions, conversion webhook health, controller logs, reconcile rate, finalizers, and failed custom resources. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Treat CRDs as APIs. Add new versions before removing old ones, maintain conversion webhooks, validate schemas carefully, and migrate stored versions. Controllers should be idempotent, rate-limited, and scoped to the minimum resources they own. Never remove fields or finalizers without a migration plan.

QUESTION 108KubernetesHard

Kubernetes architecture scenario: Design CRDs and controllers with versioning, conversion, and backward compatibility. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Treat CRDs as APIs. Add new versions before removing old ones, maintain conversion webhooks, validate schemas carefully, and migrate stored versions. Controllers should be idempotent, rate-limited, and scoped to the minimum resources they own. Never remove fields or finalizers without a migration plan.

QUESTION 109KubernetesMedium

Kubernetes security scenario: A custom controller has broad permissions and can modify resources across all namespaces. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Treat CRDs as APIs. Add new versions before removing old ones, maintain conversion webhooks, validate schemas carefully, and migrate stored versions. Controllers should be idempotent, rate-limited, and scoped to the minimum resources they own. Never remove fields or finalizers without a migration plan.

QUESTION 110KubernetesHard

Kubernetes release scenario: Upgrade a CRD schema and controller without corrupting existing resources. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Treat CRDs as APIs. Add new versions before removing old ones, maintain conversion webhooks, validate schemas carefully, and migrate stored versions. Controllers should be idempotent, rate-limited, and scoped to the minimum resources they own. Never remove fields or finalizers without a migration plan.

QUESTION 111KubernetesMedium

Kubernetes reliability/cost scenario: A controller reconcile loop creates an event storm and overloads the API server. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Treat CRDs as APIs. Add new versions before removing old ones, maintain conversion webhooks, validate schemas carefully, and migrate stored versions. Controllers should be idempotent, rate-limited, and scoped to the minimum resources they own. Never remove fields or finalizers without a migration plan.

QUESTION 112KubernetesHard

Kubernetes incident: A Kubernetes minor version upgrade causes workloads using deprecated APIs to fail. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check deprecated API metrics, kubectl api-resources, admission warnings, add-on compatibility, node drain events, PDBs, and workload SLOs. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Run pre-upgrade scans for deprecated APIs and add-on compatibility, upgrade control plane first, then node pools gradually. Drain one failure domain at a time, respect PDBs, and keep rollback or surge capacity. Make upgrades routine so security patches do not become emergency migrations.

QUESTION 113KubernetesHard

Kubernetes architecture scenario: Design a cluster upgrade process for many teams and namespaces. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Run pre-upgrade scans for deprecated APIs and add-on compatibility, upgrade control plane first, then node pools gradually. Drain one failure domain at a time, respect PDBs, and keep rollback or surge capacity. Make upgrades routine so security patches do not become emergency migrations.

QUESTION 114KubernetesMedium

Kubernetes security scenario: Old cluster versions carry known vulnerabilities but upgrades are repeatedly delayed. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Run pre-upgrade scans for deprecated APIs and add-on compatibility, upgrade control plane first, then node pools gradually. Drain one failure domain at a time, respect PDBs, and keep rollback or surge capacity. Make upgrades routine so security patches do not become emergency migrations.

QUESTION 115KubernetesHard

Kubernetes release scenario: Upgrade control plane, node pools, and add-ons with minimal downtime. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Run pre-upgrade scans for deprecated APIs and add-on compatibility, upgrade control plane first, then node pools gradually. Drain one failure domain at a time, respect PDBs, and keep rollback or surge capacity. Make upgrades routine so security patches do not become emergency migrations.

QUESTION 116KubernetesMedium

Kubernetes reliability/cost scenario: Node upgrades evict too many pods or break workloads tied to node kernel behavior. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Run pre-upgrade scans for deprecated APIs and add-on compatibility, upgrade control plane first, then node pools gradually. Drain one failure domain at a time, respect PDBs, and keep rollback or surge capacity. Make upgrades routine so security patches do not become emergency migrations.

QUESTION 117KubernetesHard

Kubernetes incident: Critical pods are Pending because affinity, taints, and topology constraints conflict. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check pod scheduling events, scheduler logs, node labels, taints/tolerations, affinity rules, topology spread constraints, and resource requests. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Use hard constraints only for requirements and soft preferences for optimization. Protect platform nodes with taints and admission policy, and use topology spread with realistic maxSkew and fallback behavior. Validate changes with representative workloads before applying them cluster-wide.

QUESTION 118KubernetesHard

Kubernetes architecture scenario: Design scheduling rules for critical, batch, GPU, and zone-aware workloads. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Use hard constraints only for requirements and soft preferences for optimization. Protect platform nodes with taints and admission policy, and use topology spread with realistic maxSkew and fallback behavior. Validate changes with representative workloads before applying them cluster-wide.

QUESTION 119KubernetesMedium

Kubernetes security scenario: Untrusted workloads can schedule onto nodes meant for privileged platform components. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Use hard constraints only for requirements and soft preferences for optimization. Protect platform nodes with taints and admission policy, and use topology spread with realistic maxSkew and fallback behavior. Validate changes with representative workloads before applying them cluster-wide.

QUESTION 120KubernetesHard

Kubernetes release scenario: Introduce taints, tolerations, and topology spread constraints to an existing cluster. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Use hard constraints only for requirements and soft preferences for optimization. Protect platform nodes with taints and admission policy, and use topology spread with realistic maxSkew and fallback behavior. Validate changes with representative workloads before applying them cluster-wide.

QUESTION 121KubernetesMedium

Kubernetes reliability/cost scenario: Strict spreading improves HA but prevents scheduling during partial capacity loss. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Use hard constraints only for requirements and soft preferences for optimization. Protect platform nodes with taints and admission policy, and use topology spread with realistic maxSkew and fallback behavior. Validate changes with representative workloads before applying them cluster-wide.

QUESTION 122KubernetesHard

Kubernetes incident: Users report errors but cluster dashboards show all pods Running. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check service SLOs, ingress metrics, pod restarts, events, logs, traces, RED/USE metrics, and metric cardinality reports. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Running is not a health signal. Monitor request rate, errors, duration, saturation, dependency failures, and endpoint readiness. Add SLO burn-rate alerts tied to user impact, scrub sensitive fields from logs/traces, and control metric labels so cardinality does not break monitoring.

QUESTION 123KubernetesHard

Kubernetes architecture scenario: Design observability for Kubernetes services from node to user journey. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Running is not a health signal. Monitor request rate, errors, duration, saturation, dependency failures, and endpoint readiness. Add SLO burn-rate alerts tied to user impact, scrub sensitive fields from logs/traces, and control metric labels so cardinality does not break monitoring.

QUESTION 124KubernetesMedium

Kubernetes security scenario: Logs and traces contain secrets, tokens, or customer data. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Running is not a health signal. Monitor request rate, errors, duration, saturation, dependency failures, and endpoint readiness. Add SLO burn-rate alerts tied to user impact, scrub sensitive fields from logs/traces, and control metric labels so cardinality does not break monitoring.

QUESTION 125KubernetesHard

Kubernetes release scenario: Introduce SLO burn-rate alerts and service dashboards. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Running is not a health signal. Monitor request rate, errors, duration, saturation, dependency failures, and endpoint readiness. Add SLO burn-rate alerts tied to user impact, scrub sensitive fields from logs/traces, and control metric labels so cardinality does not break monitoring.

QUESTION 126KubernetesMedium

Kubernetes reliability/cost scenario: High-cardinality metrics make Prometheus expensive and slow. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Running is not a health signal. Monitor request rate, errors, duration, saturation, dependency failures, and endpoint readiness. Add SLO burn-rate alerts tied to user impact, scrub sensitive fields from logs/traces, and control metric labels so cardinality does not break monitoring.

QUESTION 127KubernetesHard

Kubernetes incident: A rollout fails because nodes cannot pull a private image. How do you investigate, recover service, and prevent the same failure?

#
Reveal answer guidance

Start by proving scope and the last change before editing resources. Check pod image pull events, registry auth, imagePullSecrets, node credential provider logs, admission policy, scan results, and image size history. Recover with the smallest reversible action, such as rollback, scaling, removing a bad endpoint, or temporarily relaxing a broken policy. The durable fix is: Use immutable image digests, private registries, short-lived pull credentials, vulnerability scanning, and signature verification with a staged exception process. Keep runtime images small and pre-pull critical images where needed. Do not rely on mutable tags for production rollbacks.

QUESTION 128KubernetesHard

Kubernetes architecture scenario: Design secure image build, signing, scanning, and promotion for Kubernetes. What design would you choose and what tradeoffs would you explain?

#
Reveal answer guidance

Design around failure domains, clear ownership, safe rollout, resource isolation, and observability. Prefer Kubernetes primitives first, then add controllers or platform tooling only when they solve an operational gap. For this scenario: Use immutable image digests, private registries, short-lived pull credentials, vulnerability scanning, and signature verification with a staged exception process. Keep runtime images small and pre-pull critical images where needed. Do not rely on mutable tags for production rollbacks.

QUESTION 129KubernetesMedium

Kubernetes security scenario: Clusters can run unsigned images or images from public registries. How do you harden it without breaking production?

#
Reveal answer guidance

Inventory current consumers, add audit or report-only checks first where possible, then roll out namespace by namespace. Keep a rollback path and alert on denied traffic, rejected admission requests, or auth failures. For this scenario: Use immutable image digests, private registries, short-lived pull credentials, vulnerability scanning, and signature verification with a staged exception process. Keep runtime images small and pre-pull critical images where needed. Do not rely on mutable tags for production rollbacks.

QUESTION 130KubernetesHard

Kubernetes release scenario: Enforce image registry and signature policy without blocking emergency fixes. How do you ship the change safely?

#
Reveal answer guidance

Use a staged rollout with preflight checks, small blast radius, health gates, and an explicit rollback. Watch workload readiness, endpoint count, error rate, latency, saturation, and events while the change rolls forward. For this scenario: Use immutable image digests, private registries, short-lived pull credentials, vulnerability scanning, and signature verification with a staged exception process. Keep runtime images small and pre-pull critical images where needed. Do not rely on mutable tags for production rollbacks.

QUESTION 131KubernetesMedium

Kubernetes reliability/cost scenario: Large images slow node scale-up and increase cold-start time. What signals do you inspect and what changes do you make?

#
Reveal answer guidance

Inspect requests, limits, throttling, memory pressure, pending pods, node utilization, disruption events, restart rate, and traffic shape. Tune the bottleneck instead of blindly adding nodes. For this scenario: Use immutable image digests, private registries, short-lived pull credentials, vulnerability scanning, and signature verification with a staged exception process. Keep runtime images small and pre-pull critical images where needed. Do not rely on mutable tags for production rollbacks.

CONTINUE PRACTICING

Try another perspective.