Why This Matters#
Kubernetes is not a deployment tool. It is a control loop over desired state: you write down what should exist, dozens of controllers keep comparing that to what does exist, and they act on the difference. Containers are the least interesting part. Most production incidents on Kubernetes come from the control plane and its defaults: a liveness probe that kills healthy pods under load, a pod with no CPU request that the scheduler packs onto a full node, an etcd that hits its 2 GiB quota and freezes every deploy in the company, an upgrade that silently drops a label some forgotten config depended on.
That is why "we'll run it on Kubernetes" is a sentence interviewers push on. The L5 candidate draws a box labelled "K8s" around the services and moves on. The L6 candidate says "each service is a Deployment with CPU and memory requests sized from load tests, a readiness probe that checks only the process itself, an HPA on requests-in-flight with the default 5-minute scale-down window, a PodDisruptionBudget so node drains can't take out quorum, and spread across 3 zones with topology constraints." The L7 candidate asks how many clusters the company should have, who runs them, and what it costs to keep 40 of them within the 14-month support window, because in year three the upgrade treadmill, not the YAML, is the bill.
The L5 → L6 gap is not knowing what a kubelet is. It is knowing that every Kubernetes object is a promise a controller will try to keep forever, so a wrong promise (a bad probe, a missing request, a too-tight budget) is enforced just as diligently as a right one.
The L5 → L6 → L7 Contrast#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Containerize the services and deploy to Kubernetes" | "What is the unit of failure and scaling? That decides Deployment vs StatefulSet vs Job, and which of these belong on Kubernetes at all." | "Should this team run a cluster, use the shared platform, or skip Kubernetes? That's a headcount and blast-radius decision, not a YAML one." |
| Health | "Add liveness and readiness probes" | Readiness gates traffic; liveness only detects a wedged process and never checks dependencies; startup probe covers slow boot | Writes the org probe standard and audits it, because one copy-pasted liveness template can restart a whole fleet during a database blip |
| Resources | "Set limits so pods don't use too much" | Requests sized from measured p95 usage; memory limit = request; CPU limit often omitted for latency services; QoS class chosen deliberately | Prices bin-packing: requests are what the company pays for; sets namespace quotas and chargeback by requested, not used, cores |
| Scaling | "HPA on CPU at 70%" | HPA on a demand signal (in-flight requests, queue lag), knows the 15s loop and 300s scale-down window, pre-provisions node headroom | Decides cluster autoscaler vs node-pool strategy fleet-wide and the capacity buffer the business pays for |
| Failure | "Kubernetes restarts failed pods" | Names what happens when etcd, the API server or a zone fails: running pods keep serving, changes stop | Designs cell boundaries so no single cluster or control plane is the whole company's blast radius |
| Ownership | "DevOps runs the cluster" | Platform owns nodes, control plane and add-ons; service teams own manifests, probes and requests | Defines the paved road: what the platform guarantees, what teams may customise, and who pays for upgrades |
Why "Health" separates levels
"Add liveness and readiness probes" is correct and incomplete in the most dangerous way. A liveness probe that calls /health, which in turn pings the database, turns a 20-second database slowdown into a fleet-wide restart: every pod fails liveness at once, every pod restarts, and every pod reconnects to the already-struggling database with a cold cache. The Kubernetes documentation itself warns that incorrect liveness probes can cause cascading failures. The Staff answer separates the questions: readiness asks "should I receive traffic right now?" and may consult dependencies; liveness asks "is this process wedged beyond self-recovery?" and must not. The Principal answer turns that into a standard with a linter, because a fleet of 300 services will otherwise copy whichever template was nearest.
Why "Resources" separates levels
The scheduler places pods by requests, not by actual usage and not by limits. A pod with no request is, to the scheduler, free, so it lands on whatever node has room on paper. Under load it competes with everything else and is the first candidate for eviction. A CPU limit, meanwhile, is enforced by the kernel's CFS quota in 100ms periods: a 4-thread service with a 0.5-CPU limit can burn its 50ms quota in 12.5ms of wall-clock time and then sit throttled for the remaining 87.5ms, adding up to ~90ms of latency to requests that looked cheap on average. The Senior answer sets limits; the Staff answer sets requests from data and treats CPU limits as a latency risk; the Principal answer recognises that the sum of requests is the compute bill.
The 60-Second Pitch#
"I'd run the stateless services on a managed Kubernetes cluster per region, spread over 3 zones. Each service is a Deployment with requests sized from load tests, a readiness probe on the process's own ability to serve, a startup probe for the 40-second JVM boot, and no dependency checks in liveness. Scaling is an HPA on in-flight requests per pod, with the cluster autoscaler adding nodes and ~10% placeholder headroom so a burst doesn't wait 2–3 minutes for a VM. Kubernetes gives us self-healing, bin-packing and a uniform deploy and rollback path for 40 services. It does not give us a database: the stateful core stays on a managed database, and nothing application-level lives in etcd. If the control plane goes down, running pods keep serving; only changes stop. I'd keep each cluster small enough that losing one is an incident, not an outage."
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Stateless service platform | Many services, frequent deploys, uniform ops | Deployments + HPA + Services; managed control plane; paved-road templates | Bad probe or missing requests causes fleet-wide restarts or noisy neighbours | Zero-downtime deploys; p99 unaffected by node loss |
| Batch and ML compute | Bin-packing, queueing, GPUs, preemption | Jobs/CronJobs, priority classes, a batch scheduler (Kueue, Volcano), large node pools | Thousands of pods hammer the API server; scheduler throughput becomes the bottleneck | Every job runs once to completion; fair share across teams |
| Stateful systems on Kubernetes | Stable identity, persistent volumes, ordered rollout | StatefulSets + operators + PVCs + PDBs; anti-affinity across zones | Zone loss strands volumes; operator bug or drain violates quorum | No data loss; quorum never broken by voluntary disruption |
🎯 Staff Move: "I'll design for the first intent: Kubernetes as the stateless service platform. The payments ledger stays on a managed database outside the cluster. Running stateful systems on Kubernetes is a legitimate second step, but only behind a mature operator and only once the team has rehearsed zone loss, so I won't make the core design depend on it."
The Staff Positions#
| Position | Rationale |
|---|---|
| Requests on every container, set from measurement | The scheduler, the HPA and eviction all reason from requests. No request means no placement guarantee and no CPU-based autoscaling. |
| Liveness never checks dependencies | A dependency outage should drain traffic (readiness), not restart the fleet (liveness). |
| Memory limit = memory request; CPU limit only where justified | Memory overcommit ends in OOM kills at the worst moment; CPU limits add throttling latency for little benefit on latency-sensitive services. |
| Running pods must not depend on the control plane | etcd, API server and scheduler are for changes. The data plane keeps serving the last state if they disappear. |
| Many medium clusters beat one huge one | A cluster is a blast radius and an upgrade unit; 5,000 nodes is a documented ceiling, not a target. |
| Application state does not go in etcd | etcd holds cluster metadata under a 2 GiB default quota; CRDs used as a database fill it and freeze the cluster. |
| Every workload has a PDB and zone spread | Node drains and upgrades are routine; without a budget they take out quorum or capacity. |
Architecture & Internals#
Only five internals change design decisions: the API server and etcd, controllers and the reconcile loop, the scheduler, the kubelet and probes, and Services/EndpointSlices.
The Control Plane: One Front Door, One Source of Truth#
Every object (Pod, Deployment, Service, ConfigMap, Secret, custom resource) lives in etcd and is read and written only through the API server. Controllers, the scheduler and every kubelet talk to the API server, mostly through long-lived watches. The API server keeps a watch cache so that thousands of clients don't watch etcd directly. The consensus mechanics, quotas and compaction of etcd are covered in etcd & ZooKeeper; here the point is what they mean for workloads.
Why it matters in design: requests from users never touch the control plane. If etcd loses quorum or the API server is down, existing pods keep running, kube-proxy keeps its programmed rules and traffic flows. What stops is everything that changes: deploys, scaling, rescheduling after a node failure, new endpoints. The Kubernetes etcd operations guide is blunt about it: if etcd is starved, no cluster state changes are possible and no new pods can be scheduled. That is the canonical control-plane/data-plane split, and it is the property you should draw when someone asks "what if the control plane dies?"
The Reconcile Loop#
Every controller runs the same loop: observe current state via a watch, diff against desired state, act to close the gap, repeat. Nothing is a one-shot command. A Deployment does not "deploy"; it declares that a ReplicaSet with template hash X should have 12 replicas, and the Deployment and ReplicaSet controllers keep making that true.
loop forever:
desired = spec from the API server (e.g. replicas: 12)
actual = observed status (e.g. 10 Running, 1 Pending, 1 Failed)
diff = desired - actual
if diff != 0:
act(diff) (create 2 pods, delete the failed one)
write status back
wait for next watch event or resync
Three consequences candidates miss:
- Level-triggered, not edge-triggered. A controller that misses an event still converges on the next one, because it compares states, not deltas. That is why Kubernetes survives controller restarts without a replay log.
- Controllers fight. An HPA setting
replicas: 18and a GitOps tool reapplyingreplicas: 6from Git will flap forever. One owner per field; omitreplicasfrom Git when an HPA owns it. - A wrong spec is enforced with the same diligence as a right one. If the spec says "restart when
/healthzfails" and/healthzchecks the database, the controller will faithfully restart every pod during a database brownout.
🎯 Staff Insight: "I treat every manifest as a standing order to a robot that never gets tired. So the review question isn't 'does this deploy?', it's 'what will the controllers do with this spec at 3am when a dependency is slow?'"
The Scheduler: Filter, Score, Bind#
The scheduler watches for pods with no node, then for each pod filters nodes that can't host it (insufficient requested CPU or memory, taints without tolerations, node selectors, affinity rules, volume zone constraints), scores the survivors (spread, resource balance, image locality, preferred affinity), and binds the pod to the winner. It never looks at actual utilisation.
| Placement control | What it does | Use for |
|---|---|---|
resources.requests | Reserves capacity for placement | Every container, always |
nodeSelector / node affinity | Restricts to labelled nodes | GPU pools, ARM pools, compliance pools |
| Taints + tolerations | Keeps pods off nodes unless tolerated | Dedicated pools (system, GPU, batch) |
| Pod anti-affinity | Keeps replicas apart | Quorum members on separate nodes |
topologySpreadConstraints | Bounded skew across zones or nodes | Spread 12 replicas 4/4/4 over 3 zones, maxSkew: 1 |
priorityClassName | Higher priority can preempt lower | System and critical services over batch |
The trap: affinity rules are evaluated per pod against every node. Heavy inter-pod affinity in a large cluster is the classic way to slow the scheduler from thousands of pods per minute to a crawl. Prefer topology spread constraints, which cover most of the same intent more cheaply.
Pod Lifecycle: From kubectl apply to Receiving Traffic#
Every arrow is a watch, so every step has propagation delay. In a healthy cluster a pod goes from bound to receiving traffic in seconds to tens of seconds, dominated by image pull and the startup and readiness probes. The interval between "pod removed from EndpointSlice" and "every proxy has stopped sending to it" is why graceful shutdown needs a preStop delay; the full mechanics are in Service Registry.
Kubelet, Probes and Node Health#
The kubelet is the node agent: it runs containers via the container runtime, executes probes, enforces resource limits through cgroups, evicts pods under node pressure, and reports node and pod status to the API server.
| Probe | Question it answers | Effect on failure | Who should define it |
|---|---|---|---|
| startupProbe | "Has the process finished booting?" | Liveness and readiness are held off until it passes; fails → restart | Service team; sized to p99 boot time |
| readinessProbe | "Should I receive traffic right now?" | Pod removed from EndpointSlices; no restart | Service team; may reflect dependency health or load |
| livenessProbe | "Is this process wedged beyond self-recovery?" | Container restarted | Service team, under a platform standard: no dependency checks |
Probe defaults: initialDelaySeconds 0, periodSeconds 10, timeoutSeconds 1, successThreshold 1, failureThreshold 3 (Kubernetes probe docs). With defaults, a pod is marked unready or restarted roughly 30 seconds after it starts failing, and a 1-second timeout is what turns a GC pause or a slow response under load into a probe failure.
Node failure timeline. When a node stops reporting, the node lifecycle controller marks it NotReady after the node-monitor grace period (tens of seconds), taints it, and pods on it are evicted only after the default 300-second toleration for not-ready and unreachable taints added by the DefaultTolerationSeconds admission plugin (well-known taints). So with defaults, pods on a dead node take about 5–6 minutes to be recreated elsewhere. Readiness removes them from traffic far sooner, which is why replicas across zones, not rescheduling, is the availability mechanism.
🎯 Staff Insight: "Kubernetes 'self-healing' takes about six minutes for a dead node with default tolerations. My availability comes from N+1 replicas spread across zones and readiness removing dead endpoints. Rescheduling restores capacity; it doesn't restore availability."
Services and EndpointSlices#
A Service is a stable virtual IP and DNS name in front of a label selector. The EndpointSlice controller watches pods matching the selector and writes their IPs into EndpointSlice objects, including only pods whose readiness probe passes. By default each slice holds up to 100 endpoints (configurable up to 1,000) to keep update fan-out small (EndpointSlices). kube-proxy (iptables, IPVS or nftables mode) or a mesh control plane watches those slices and programs every node or sidecar.
| Service type | What you get | Use for |
|---|---|---|
ClusterIP | In-cluster virtual IP, L4 load balancing per connection | Default service-to-service |
Headless (clusterIP: None) | DNS returns pod IPs directly | StatefulSets, client-side load balancing, gRPC |
NodePort / LoadBalancer | Exposes through node ports or a cloud L4 load balancer | Edge entry, non-HTTP protocols |
| Ingress / Gateway API | L7 routing, TLS, host and path rules | HTTP edge; see Envoy, Kong & NGINX |
The gRPC trap: ClusterIP balances connections, not requests. A gRPC client holding one long-lived HTTP/2 connection sends every request to one pod. Use a headless Service with client-side balancing, or a mesh that balances per request. The full discovery design, including mass-deregistration guards and panic thresholds, lives in Service Discovery.
Core Usage — "The Entire Game": The Workload Contract#
In Kafka the game is the partition key. In Kubernetes it is the workload contract: the five things every service team declares and the platform enforces. Get these right and the cluster mostly runs itself; get them wrong and the controllers will enforce the mistake at fleet scale.
The workload contract (per container):
1. requests -> where it can be placed and what it is guaranteed
2. limits -> when it is throttled (CPU) or killed (memory)
3. probes -> when it gets traffic and when it gets restarted
4. disruption -> how many replicas may be down during drains and upgrades
5. spread -> which failures (node, zone) it survives
Step 1: Requests and Limits — Size From Measurement#
The scheduler uses requests to decide placement; the kernel enforces limits. CPU over the limit is throttled; memory over the limit gets the container OOM-killed when the kernel detects memory pressure (resource management docs).
| QoS class | How you get it | Eviction order under node pressure | Use for |
|---|---|---|---|
| Guaranteed | Requests = limits for CPU and memory, every container | Last | Latency-critical, stateful, quorum members |
| Burstable | Requests set, limits higher or absent | Middle (by usage over request) | Most stateless services |
| BestEffort | No requests, no limits | First | Nothing in production |
Sizing recipe (a common production pattern):
memory.request = memory.limit = p99 working set under peak load test x 1.2
cpu.request = p95 CPU at target RPS per pod
cpu.limit = unset for latency services (or >= 2-4x request if policy demands one)
Example: checkout-api at 400 RPS per pod
load test: p95 CPU 0.7 cores, p99 working set 900 MiB
requests: cpu 700m, memory 1.1Gi
limits: memory 1.1Gi, no CPU limit
node: 16 vCPU / 64 GiB, ~1 vCPU and ~3 GiB reserved for system
pods/node: floor(15 / 0.7) = 21 by CPU, floor(61 / 1.1) = 55 by memory -> CPU-bound, 21
Why memory limit = request: memory is not compressible. If 21 pods each request 1.1 GiB but are allowed to burst to 3 GiB, the node is overcommitted ~2.7× on paper, and the first traffic spike triggers OOM kills across pods that did nothing wrong. Why no CPU limit: CPU is compressible. Without a limit, a pod borrows idle cycles and the request still guarantees its fair share under contention. With a limit, the CFS quota throttles it inside each 100ms period even when the node is idle. Watch container_cpu_cfs_throttled_periods_total; a ratio over ~5–10% of periods on a latency service is a p99 problem.
🎯 Staff Move: "I'll set requests from the load test, memory limit equal to request, and no CPU limit on the API pods. I'd rather manage noisy neighbours with requests and namespace quotas than pay 50–100ms of throttling at p99. Batch pools are different: there I'd set CPU limits so a runaway job can't starve its neighbours."
Step 2: Probes — Three Questions, Three Probes#
startupProbe: # p99 boot is ~40s; allow 60s before liveness takes over
httpGet: { path: /startup, port: 8080 }
periodSeconds: 5
failureThreshold: 12 # 12 x 5s = 60s budget
readinessProbe: # "send me traffic" - may reflect load and critical deps
httpGet: { path: /ready, port: 8080 }
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 2 # ~10s to drain a sick pod
livenessProbe: # "am I wedged?" - in-process only, never the database
httpGet: { path: /live, port: 8080 }
periodSeconds: 10
timeoutSeconds: 5 # generous: a GC pause is not a deadlock
failureThreshold: 6 # ~60s of consecutive failure before restart
The design rule: readiness is fast and sensitive, liveness is slow and conservative. Readiness false costs one pod's share of traffic for a few seconds. Liveness false costs a restart, a cold cache, a reconnect storm and, if correlated across the fleet, an outage.
| Probe check | Readiness | Liveness |
|---|---|---|
| Process can serve HTTP | Yes | Yes |
| Thread pool or event loop not deadlocked | Yes | Yes, this is the point |
| Database reachable | Maybe, if the pod is useless without it | Never |
| Downstream service healthy | Rarely; prefer circuit breakers | Never |
| Warm cache loaded | Yes | No |
| In-flight requests above shed threshold | Yes (load shedding) | No |
Step 3: Autoscaling — Three Loops With Different Clocks#
| Loop | What it scales | Default clock | Reaction time in practice |
|---|---|---|---|
| HPA | Replicas of a Deployment | Sync every 15s; scale-up stabilisation 0s; scale-down stabilisation 300s; tolerance 10% | 15–60s to decide, plus pod start |
| Cluster autoscaler / Karpenter | Nodes | Reacts to pods Pending for lack of capacity | VM boot + image pull: often 1–3 minutes |
| VPA | Requests per pod | Recommends from usage history | Hours to days; applies by restarting pods |
The HPA computes desiredReplicas = ceil(currentReplicas × currentMetric / targetMetric), ignores changes within the 10% tolerance, and calculates CPU utilisation as a percentage of the request, so a container with no CPU request cannot be autoscaled on CPU at all (HPA docs).
Example: checkout-api, 30 pods, target 200 in-flight requests per pod
Black Friday spike: observed 340 per pod
desired = ceil(30 x 340 / 200) = ceil(51.0) = 51 pods
51 x 0.7 CPU = 35.7 cores requested; existing headroom 12 cores
-> ~18 pods Pending -> cluster autoscaler adds 2 nodes -> ~2 minutes
-> placeholder pods (low priority, ~10% of cluster) are preempted instantly,
so ~17 of the 21 new pods start in seconds instead of minutes
The Staff point is that the three loops compound: HPA decides in ~15–30s, but if the new pods need new nodes, the user waits for a VM boot. Capacity headroom (low-priority placeholder pods that get preempted) is the cheap fix. The full capacity story, including asymmetric step caps and predictive scaling, is in Autoscaling & Capacity.
🎯 Staff Insight: "CPU is a lagging, indirect signal for an I/O-bound API. I'd scale on in-flight requests or queue lag via a custom metric, keep the 5-minute scale-down window, and hold 10% of the cluster as preemptible placeholder pods so a burst doesn't wait on a VM."
Step 4: Disruption Budgets and Rollouts#
A PodDisruptionBudget limits how many replicas voluntary disruptions (node drains, cluster upgrades, autoscaler scale-down) may take down at once. It does not prevent involuntary disruptions such as hardware failure, though those count against the budget (disruptions docs).
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 25% # extra pods during rollout (needs spare capacity)
maxUnavailable: 0 # never dip below desired replicas
minReadySeconds: 20 # a pod must stay Ready 20s before counting
---
kind: PodDisruptionBudget
spec:
maxUnavailable: 1 # for a 3-member quorum: never drain 2 at once
selector: { matchLabels: { app: ledger-consumer } }
Two traps: a PDB of minAvailable: 100% (or maxUnavailable: 0) blocks every node drain, so the cluster can never be upgraded; and a Deployment with 1 replica and any PDB is the same thing in disguise. The platform should reject both at admission.
Step 5: Spread Across Failure Domains#
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule # hard for zones
labelSelector: { matchLabels: { app: checkout-api } }
- maxSkew: 2
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway # soft for nodes
labelSelector: { matchLabels: { app: checkout-api } }
The availability arithmetic: to survive a zone loss at peak with 3 zones, you need 1.5× peak capacity spread evenly (each zone carries 50% of peak so two zones carry 100%). That 50% buffer is a business decision priced in nodes, and it is the most common place Kubernetes designs quietly assume capacity they haven't paid for.
The Tunable Tradeoff — Density vs Isolation#
Every Kubernetes platform decision moves along one axis: how tightly do we pack workloads together? Dense packing is cheap; isolation is predictable.
| Setting | Dense end | Isolated end | Who pays at the dense end |
|---|---|---|---|
| Requests vs actual usage | Requests below usage (overcommit) | Requests at p95–p99 usage | Service teams, as OOM kills and throttling |
| QoS | Burstable / BestEffort | Guaranteed | Whoever is evicted first |
| Pods per node | Near the 110 default | Fewer, larger pods | Everyone on the node when it dies or is noisy |
| Node pools | One shared pool | Dedicated pools by tier (system, critical, batch, GPU) | Platform, in utilisation; product teams, in noisy neighbours |
| Clusters | One per region for the company | Per domain or per cell | Every team during the one bad upgrade |
| Tenancy | Namespaces only (soft) | Separate clusters or sandboxed runtimes (hard) | Security, if a tenant is hostile |
Utilisation formula the finance team cares about:
cluster_efficiency = sum(actual usage) / sum(node capacity)
request_efficiency = sum(actual usage) / sum(requests)
allocation_ratio = sum(requests) / sum(allocatable)
Typical unmanaged fleet: request_efficiency 30-40%, allocation_ratio 70-80%
-> ~25-30% real utilisation: you pay for ~3-4x what you use.
Staff target for stateless tiers: request_efficiency 60-70% via VPA
recommendations + right-sizing reviews, allocation_ratio ~85% with headroom pods.
🎯 Staff Move: "I'll pack stateless services densely in a shared Burstable pool, keep quorum members and the payment path Guaranteed in a dedicated pool, and push batch onto a preemptible pool with lower priority. Density where failure is cheap, isolation where it isn't."
Who Pays for Each Choice#
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| No CPU limits | No throttling latency; idle cycles used | A runaway pod can consume a node's spare CPU | Neighbouring Burstable pods, mildly |
| CPU limits everywhere | Predictable per-pod ceilings | p99 latency from CFS throttling | Service team's SLO |
| Overcommitted memory | 20–40% more pods per node | OOM kills under correlated load | Service on-call at peak |
| Aggressive liveness probes | Wedged pods recovered fast | Fleet restarts on dependency blips | Every user during the cascade |
| Tight PDBs | Quorum never broken by drains | Upgrades block for days | Platform team's upgrade schedule |
| Single giant cluster | One place to operate, best packing | Company-wide blast radius per control-plane incident | Every team at once |
Anti-Patterns — What Kills Kubernetes Deployments#
1. Liveness Probes That Check Dependencies#
/health pings the database; the database has a 20-second slowdown; every pod fails liveness within ~30s and restarts; every restarted pod reconnects with a cold cache and connection-pool warm-up, amplifying load on the database that caused it. The official docs list exactly this: incorrect liveness probes cause restarts under high load, failed client requests and extra load on the remaining pods (probe concepts). Fix: liveness checks only in-process state; dependency health goes in readiness (carefully) or circuit breakers.
2. No CPU (or Memory) Requests#
To the scheduler, a pod without requests costs nothing, so a node can be packed with them until real usage hits 100%. They are BestEffort and evicted first; the HPA can't compute utilisation for them. Fix: an admission policy (LimitRange defaults plus a policy engine) that rejects production pods without requests.
3. CPU Limits on Latency-Critical, Multi-Threaded Services#
A JVM or Go service with 8 busy threads and a 1-CPU limit exhausts its 100ms quota in 12.5ms and stalls for 87.5ms. Average CPU looks like 40%; p99 is terrible. Fix: requests without CPU limits on latency tiers; watch throttled-period ratio.
4. A Single Cluster as the Whole Company's Blast Radius#
One cluster per region, every team on it. One bad admission webhook, CRD upgrade, CNI change or etcd quota breach freezes deploys or breaks networking for every service at once. Fix: multiple clusters by tier or cell, staged upgrades (canary cluster first), and the ability to shift traffic between clusters.
5. Kubernetes as a Database for Application State#
Custom resources used to store user sessions, job payloads or per-tenant records. Every object lands in etcd, under a 2 GiB default quota (8 GiB suggested maximum) and a 1.5 MiB request limit (etcd limits). When the quota trips, etcd goes read-only and nothing in the cluster can change. Fix: CRDs describe infrastructure intent (a desired database, a desired certificate), not business data. Business data goes in a database.
6. Fail-Closed Admission Webhooks Without an Escape Hatch#
A validating webhook with failurePolicy: Fail whose backing pods run in the same cluster: the webhook deployment crashes, no pod can be created, including the webhook's own replacement. Fix: exclude kube-system and the webhook's namespace, run 3+ replicas with a PDB, set short timeouts, and choose Ignore for anything that isn't a security control.
7. latest Tags and Mutable Images#
Pods on different nodes run different code under the same tag; rollback re-pulls the same broken image. Fix: immutable tags or digests; admission policy rejects :latest.
8. Ignoring the Upgrade Treadmill#
Minor releases arrive about 3 times a year, and each is supported for roughly 14 months (release cadence, patch support). A team that upgrades once a year is always rushing two versions at once, against deprecated APIs and add-ons nobody remembers installing. Fix: upgrades as a quarterly routine with an owner, a canary cluster and an API-deprecation scanner in CI.
The Technology Landscape — Head-to-Head Comparison#
| Dimension | Kubernetes (managed: GKE/EKS/AKS) | Self-managed Kubernetes | Serverless containers (Cloud Run, Fargate, ECS) | Nomad | Plain VMs + autoscaling groups |
|---|---|---|---|---|---|
| Model | Declarative reconcile over a rich API | Same, you run the control plane | Run a container, scale on requests | Single-binary scheduler, jobs and services | Machine images, instance groups |
| Scale ceiling | 5,000 nodes / 150,000 pods per cluster (documented) | Same | Per-service quotas; no cluster concept | Large; simpler model | Account and region quotas |
| Ops burden | Medium: node pools, add-ons, upgrades | High: etcd, certificates, control-plane upgrades | Low | Low–medium | Low–medium |
| Ecosystem | Largest: operators, Helm, service mesh, GitOps | Same | Thin | Moderate (HashiCorp stack) | Cloud-native tooling |
| Stateful workloads | Possible with operators and PVCs | Same | Mostly no | Possible | Natural |
| Cold start / scale-up | Seconds if nodes exist; minutes if not | Same | Seconds; scale to zero | Seconds | Minutes |
| Pick when | 20+ services, platform team, portability | Regulated or on-prem with deep expertise | Few services, small team, spiky traffic | Mixed workloads, simpler ops, HashiCorp shop | Few large services, VM-shaped workloads |
🎯 Staff Insight: "Managed Kubernetes removes etcd and the control plane from my pager, not the node pools, add-ons, probes, upgrades or cost. The real comparison for a small team is Kubernetes versus serverless containers, and below ~10 services the serverless option usually wins."
Patterns#
Pattern 1: Paved-Road Stateless Service#
Deployment + HPA + PDB + topology spread + ServiceAccount with least privilege + NetworkPolicy, generated from one template (Helm chart or a higher-level CRD like kind: WebService). Teams fill in image, port, requests and probe paths; the platform owns everything else. This is the default for 80%+ of services.
Pattern 2: Operators for Stateful Systems#
An operator is a custom controller that encodes the runbook of a stateful system (Postgres, Kafka, Elasticsearch): failover, backups, version upgrades, resizing, as a reconcile loop over a CRD. Use when the operator is mature, widely used and owned by someone who will patch it; avoid a home-grown operator for a database the team doesn't deeply understand. An operator automates expertise; it does not replace it.
Pattern 3: Batch Queues on Shared Clusters#
Jobs with priorityClassName: batch-low, preemptible node pools, a queueing layer (Kueue or Volcano) for quotas and gang scheduling, and API-server rate limits per tenant. The Job Scheduler design applies directly: Kubernetes Jobs are the execution layer, not the scheduler of record.
Pattern 4: Leader Election via Leases#
Controllers and singleton workers elect a leader with a coordination.k8s.io/Lease object (an etcd-backed lease with renew deadlines). Good for controllers already on the API server; for application-level locks with fencing, see Coordination Strategies: Leases, Leaders & Reconciliation and Distributed Consensus.
Pattern 5: Multi-Cluster Cells#
Each cell is a complete, independent copy of the stack serving a slice of tenants. A bad CRD, CNI upgrade or etcd incident takes out one cell (~33% of tenants here, ~5–10% in a mature fleet), not the company. The management cluster pushes configuration in waves and carries no user traffic, so its failure stops changes, not requests. State lives outside the clusters, which keeps every cluster disposable.
Pattern 6: GitOps as the Source of Desired State#
A controller (Argo CD, Flux) reconciles cluster state from Git. This extends the reconcile loop one level up: Git is desired state, the cluster is actual state, drift is corrected automatically. Rollback is git revert. The Reddit 2023 outage, below, is the cautionary tale of configuration that lived outside any such source.
Scaling#
The Numbers#
The Kubernetes project documents the supported envelope for a single cluster: no more than 110 pods per node, 5,000 nodes, 150,000 total pods and 300,000 total containers (Considerations for large clusters). These are tested limits, not targets.
| Resource | Documented limit or default | Design note |
|---|---|---|
| Nodes per cluster | 5,000 | Most fleets stay well under 1,000 per cluster for blast radius and upgrade time |
| Pods per node | 110 (kubelet maxPods default) | Also bounded by pod CIDR size and per-node IP limits on cloud CNIs |
| Total pods per cluster | 150,000 | API server memory and watch fan-out grow with object count |
| Total containers | 300,000 | Sidecars count; a mesh can double container count |
| etcd quota | 2 GiB default, 8 GiB suggested max | Events, oversized CRDs and Secrets are the usual culprits |
| etcd request size | 1.5 MiB | Bounds a single object; huge ConfigMaps fail here |
| Endpoints per EndpointSlice | 100 default, 1,000 max | A 5,000-pod Service is ~50 slices |
| HPA loop | 15s sync, 300s scale-down window, 10% tolerance | Scale-up is immediate by default |
| Pod eviction after node loss | 300s default toleration | Plus node-monitor grace period: ~5–6 min total |
| Graceful termination | 30s default terminationGracePeriodSeconds | Must exceed preStop + drain time |
| Minor releases | ~3 per year, ~14 months support each | 3 minor branches maintained at a time |
What Breaks First as a Cluster Grows#
~100 nodes: nothing interesting; defaults are fine
~500 nodes: API server LIST calls from badly written controllers spike memory;
DaemonSets x nodes = thousands of pods just for agents
~1,000 nodes: etcd DB size and write latency matter; split Events into their own etcd;
iptables-mode kube-proxy rule updates slow with 10K+ Services/endpoints
~2,500 nodes: scheduler throughput, CNI IP allocation, Prometheus scrape volume;
every node-level agent is now a load test against the API server
~5,000 nodes: documented ceiling; you are tuning API server flow control,
watch cache sizes and etcd hardware for a living
The documented guidance for large clusters includes storing Event objects in a separate etcd instance, running control-plane replicas in every failure zone and giving add-ons appropriate resource limits.
Scaling Moves in Order#
- Right-size requests (VPA recommendations, quarterly review). Often recovers 30–50% of nodes before you add any.
- Separate node pools by workload shape (system, general, memory-heavy, GPU, batch) so packing works.
- Tune the noisy clients: controllers that LIST instead of WATCH, agents that poll the API server, CI systems creating thousands of short-lived pods.
- Split etcd for Events; use API Priority and Fairness to stop one tenant's controller from starving the rest.
- Add clusters, not nodes, once a cluster is past a few hundred nodes or hosts more than one critical domain. This is where multi-cluster tooling, fleet GitOps and global load balancing become necessary.
🎯 Staff Move: "I'd cap clusters at a few hundred nodes and scale out by adding clusters. 5,000 nodes is what the project tests, not what I want to upgrade on a Tuesday. A cluster is my unit of blast radius and my unit of upgrade."
Failure Modes & Recovery#
1. Liveness-Probe Restart Cascade#
- Symptom: A dependency slows down; within a minute the restart count of every pod in the service jumps; error rate goes from 2% to 100%.
- Root cause: Liveness probe checks the dependency, or a 1-second probe timeout fails under load-induced latency.
- Detection:
kube_pod_container_status_restarts_totalrate across a Deployment; restarts correlated across many pods within one probe period;CrashLoopBackOffcount. - Fix: Patch the probe (remove dependency, raise
timeoutSecondsandfailureThreshold) and roll out; temporarily remove the liveness probe if needed. - Prevention: Platform probe standard enforced at admission; chaos test: slow the database by 5s and verify zero restarts.
2. etcd Quota Exhausted (NOSPACE)#
- Symptom:
kubectl applyfails withdatabase space exceeded; no pods can be created, deleted or rescheduled; running pods keep serving. - Root cause: Event storm, a controller writing a large CRD status every few seconds, history not compacted, or application data in CRDs.
- Detection:
etcd_mvcc_db_total_size_in_bytesover 80% of quota;etcd_server_quota_backend_bytes; object counts per resource type. - Fix: Delete the offending objects, compact, defragment one member at a time, disarm the alarm (see etcd & ZooKeeper).
- Prevention: Separate Events etcd, per-namespace object-count quotas, review of every CRD for size and write rate.
3. Node Pressure Evictions and OOM Kills#
- Symptom: Pods restart with
OOMKilledor areEvicted; latency spikes on unrelated services on the same nodes. - Root cause: Memory overcommit (limits far above requests), BestEffort pods, or a memory leak in a neighbour.
- Detection:
kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}; nodeMemoryPressurecondition; eviction events. - Fix: Raise requests to match working set; cordon and drain the worst nodes.
- Prevention: Memory limit = request; admission rejects pods without requests; VPA recommendations in review.
4. Upgrade Breaks the Data Plane#
- Symptom: Minutes into a control-plane or node upgrade, pod networking, DNS or ingress fails across the cluster.
- Root cause: A removed API or label, an incompatible CNI, CSI or webhook version, or configuration that lived outside version control.
- Detection: Synthetic probes per cluster (DNS, pod-to-pod, ingress);
apiserver_request_total5xx; CNI agent readiness. - Fix: Roll back the node pool or restore traffic to another cluster; restoring control plane from an etcd snapshot as the last resort.
- Prevention: Canary cluster that mirrors production add-ons; deprecated-API scanning; every add-on config in Git.
5. Control-Plane Overload From a Misbehaving Client#
- Symptom: API server latency climbs to seconds;
kubectltimes out; controllers fall behind; deploys stall. - Root cause: A controller or agent doing full LISTs on every resync across 1,000 nodes, a CI job creating 10,000 pods, or a watch on a huge resource.
- Detection:
apiserver_request_duration_secondsp99 by verb and client;apiserver_flowcontrol_rejected_requests_total; API server memory. - Fix: Identify the user agent and throttle or scale it down; raise API server replicas.
- Prevention: API Priority and Fairness with per-tenant flow schemas; informer-based clients only; load-test add-ons at target node count.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Liveness cascade | Correlated restarts across a Deployment | One service, then its callers | Patch probe, roll out | Service team; platform owns the standard |
| etcd NOSPACE | DB size > 80% quota | Every change in the cluster | Compact, defrag, delete offenders | Platform; offending team |
| OOM / eviction | OOMKilled reason, MemoryPressure | Pods on affected nodes | Raise requests, drain nodes | Service team |
| Upgrade break | Per-cluster synthetic probes | Whole cluster | Shift traffic to other cells, roll back | Platform |
| API server overload | Request latency by client | All controllers and deploys | Throttle client, APF | Platform; client's owner |
| Zone loss | Node NotReady by zone | One third of capacity | Pre-paid 1.5× capacity, spread | Platform + capacity owner |
| Webhook outage | Admission latency/errors | All pod creation | Fail-open non-security hooks, exempt system namespaces | Webhook's owning team |
When to Use vs. Alternatives#
| Need | Pick | Why |
|---|---|---|
| 20+ services, frequent deploys, a platform team | Managed Kubernetes | Uniform deploy, scale and recovery; huge ecosystem |
| A handful of stateless services, small team | Serverless containers (Cloud Run, Fargate, App Runner) | No nodes, no upgrades, scale to zero |
| Event-driven functions, spiky and short-lived | Functions (Lambda, Cloud Functions) | Pay per invocation |
| Primary database | Managed database outside the cluster | Durability and failover are the vendor's job |
| On-prem or regulated, deep in-house expertise | Self-managed Kubernetes | Control and portability, at real headcount cost |
| Few large, long-lived VM-shaped processes | VMs + autoscaling groups | Less machinery for no lost capability |
| Large batch and ML fleets | Kubernetes + a batch queue (Kueue/Volcano) or a dedicated scheduler | Bin-packing and preemption across teams |
When NOT to Use Kubernetes#
- A small team with simple workloads. Five engineers and four services get the cost of a platform (upgrades, node pools, ingress, observability add-ons) without the benefit of standardisation across many teams. Serverless containers or a PaaS deliver deploy, scale and rollback for zero platform headcount.
- No one owns the platform. A cluster is a product with a ~14-month version support clock. Without at least one engineer's sustained attention, it rots into an unpatched liability.
- The workload is the database. Running your primary datastore on Kubernetes is possible; doing it as your first Kubernetes workload is how teams learn PersistentVolume zone affinity during an outage.
- Hard multi-tenancy with hostile tenants. Namespaces are a soft boundary. Untrusted code needs sandboxed runtimes or separate clusters; see Online Code Runner.
- Latency-critical, hardware-tuned systems (exchange matching engines, kernel-bypass networking). The abstraction costs more than it saves.
Operational Concerns#
What the On-Call Actually Does#
- Watches per-cluster synthetic checks (DNS resolution, pod-to-pod, ingress,
kubectllatency). They catch data-plane breakage before service dashboards do. - Watches control-plane health: API server p99 latency and 5xx by client, etcd DB size against quota, etcd leader changes and fsync latency.
- Runs upgrades as a routine: control plane first, then node pools by surge, canary cluster a week ahead of production. A 300-node pool rolled at 10% surge with PDBs takes hours; that's normal.
- Drains nodes safely, and when a drain hangs, finds the PDB that allows zero disruptions and the team that owns it.
- Triages "my pod is Pending": insufficient requested resources, unsatisfiable affinity or spread, taints, PVC zone mismatch, quota exhausted.
kubectl describe podevents answer 90% of these. - Owns the admission policy set (required requests, probe standards, image provenance, no
:latest, PDB sanity) and the exceptions process.
Key Metrics & Alerts#
| Metric | Healthy | Alert |
|---|---|---|
apiserver_request_duration_seconds p99 (non-LIST) | < 1s | > 1s for 10m |
etcd_mvcc_db_total_size_in_bytes / quota | < 50% | > 80% (page) |
etcd_server_leader_changes_seen_total rate | ~0 | > 3 per hour |
kube_node_status_condition{condition="Ready",status="true"} | All nodes | > 5% of nodes NotReady |
Pods Pending > 5 minutes | 0 | > 0 for critical namespaces |
kube_pod_container_status_restarts_total rate per Deployment | ~0 | Correlated spike across > 20% of pods |
| CFS throttled periods ratio | < 5% | > 10% on latency tiers (ticket) |
| Cluster allocation ratio (requests / allocatable) | 70–85% | > 90% (no headroom) or < 50% (waste) |
| Days until the running minor version leaves support | > 120 | < 90 (ticket), < 30 (escalate) |
Upgrades as an Ongoing Cost#
Per cluster, per minor upgrade (a common production pattern):
prep: API deprecation scan, add-on compatibility matrix ~2-3 engineer-days
canary: upgrade canary cluster, soak 1 week ~1 day
rollout: control plane, then node pools by surge with PDBs ~0.5-1 day per cluster
~3 minors/year x N clusters -> at 30 clusters, upgrades alone are ~1-2 FTE
The cost scales with the number of clusters and the number of add-ons per cluster. Cell designs multiply clusters, so they only work with fleet automation and uniform add-on sets.
Interview Application — Staff-Level Plays#
Which Case Studies Use Kubernetes#
| Case Study | How Kubernetes Is Used | Key Pattern |
|---|---|---|
| Service Registry | Orchestrator as registry; readiness drives EndpointSlices | Platform registers from orchestrator truth; data plane fails static |
| Auto-Scaling & Capacity | HPA and cluster autoscaler as the reactive loop | 15s HPA loop, 300s scale-down window, placeholder headroom |
| Consensus Service | etcd under the control plane; Lease objects for leader election | Consensus for metadata, never per request |
| Distributed Job Scheduler | Jobs/CronJobs as the execution layer | Scheduler of record outside; Kubernetes runs the work |
| Online Code Runner | Sandboxed pods for untrusted submissions | Hard isolation needs sandboxed runtimes, not namespaces |
| API Gateway | Ingress and Gateway API in front of services | Edge routing as a platform-owned component |
| Load Balancing | Service VIPs, kube-proxy, cloud load balancers | L4 per-connection vs L7 per-request balancing |
| Metrics & Monitoring | Per-pod scraping, cardinality from pod labels | Pod churn multiplies time-series count |
| Multi-Region Active-Active | One cluster set per region, global traffic steering | Clusters are regional and disposable; state replicates separately |
| Video Streaming | Transcoding workers as batch Jobs on preemptible pools | Batch priority classes, queue-driven scaling |
| Deployment System | Rolling updates, readiness gates and the cluster as a deploy target | Waves and bake times live above the cluster, not inside a single Deployment |
| Multiplayer Game Backend | Dedicated game servers as pods via Agones fleets | Allocated servers are never drained mid-match; scale on a warm buffer |
Every System Design Question Has a Kubernetes Moment#
- URL shortener: "The redirect service is a stateless Deployment with an HPA on in-flight requests and a readiness probe that waits for the hot-key cache to warm. The database is outside the cluster."
- Chat: "WebSocket gateways are long-lived connections, so a rolling deploy must drain them: preStop marks the pod unready, the gateway tells clients to reconnect with jitter, and
terminationGracePeriodSecondsis raised to 5 minutes for that one Deployment." - Job scheduler: "Kubernetes Jobs execute the work; the scheduler of record with exactly-once dispatch lives in a database. I won't use CronJobs for anything that must run exactly once."
- Auth service: "Token validation pods are Guaranteed QoS in a dedicated pool with a PDB of maxUnavailable 1, because if auth is evicted under node pressure, every product is down. See Auth & Identity."
What Interviewers Probe#
| After You Say... | They Will Ask... | What They're Evaluating |
|---|---|---|
| "Kubernetes restarts failed pods" | "How long until a dead node's pods run elsewhere?" | Knowing the ~5-minute eviction default; availability from replicas |
| "Liveness and readiness probes" | "What does each check? What if the database is slow?" | The cascade failure mode |
| "HPA on CPU" | "What if the pods have no CPU request? How fast does it react?" | Requests-based utilisation; 15s loop; node provisioning delay |
| "Set resource limits" | "What does a CPU limit do to p99?" | CFS throttling |
| "One cluster for everything" | "What's the blast radius of a bad upgrade?" | Cells and multi-cluster |
| "Store the job state in a CRD" | "What's etcd's size limit?" | Kubernetes is not a database |
Common Interview Mistakes#
| What Candidates Say | What Interviewers Hear | What Staff Engineers Say |
|---|---|---|
| "Kubernetes gives us high availability" | Confuses rescheduling with redundancy | "Replicas across 3 zones give availability; rescheduling restores capacity in minutes." |
| "Liveness probe hits /health" | Fleet restart during a DB blip | "Liveness is in-process only; readiness may reflect dependencies." |
| "Set limits on everything" | Will discover CPU throttling at peak | "Requests from load tests; memory limit = request; no CPU limit on latency tiers." |
| "Autoscaling handles spikes" | Ignores node boot time | "HPA in 15–30s, nodes in minutes, so I hold 10% placeholder headroom." |
| "One big cluster is simpler" | Company-wide blast radius | "Clusters are cells; a few hundred nodes each, canary cluster first." |
| "Kubernetes for our 3 services" | Platform cost without platform benefit | "At this size, serverless containers; Kubernetes when we have 20+ services and an owner." |
L5 vs L6 vs L7 Responses#
| Scenario | L5 Answer | L6 / Staff Answer | L7 / Principal Answer |
|---|---|---|---|
| "Deploy 30 microservices" | Kubernetes cluster, Deployments, Services, HPA | Paved-road template: requests from load tests, probe standard, PDB, zone spread, HPA on demand signal, GitOps | Decides cluster topology (cells by tier), what the platform guarantees, and the upgrade budget in headcount |
| "Pods restart in a loop during an outage" | Increase probe timeout | Separate liveness from dependencies, startup probe for boot, add chaos test | Writes the probe standard, enforces it at admission, audits the fleet and tracks probe-caused incidents to zero |
| "Cluster costs doubled" | Smaller nodes | Right-size requests via VPA, bin-pack by pool, spot for batch, measure request efficiency | Chargeback on requested cores per team; capacity buffer priced and owned; retire under-used clusters |
| "Should the startup adopt Kubernetes?" | Yes, it's the industry standard | Not yet: 4 services fit serverless containers; revisit at ~20 services | Defines the trigger (service count, team count, compliance) and the migration path so the decision is re-made deliberately |
The Staff Kubernetes Checklist#
- Workload shape: "Stateless Deployment, StatefulSet behind an operator, or Job? And does this belong on Kubernetes at all?"
- Requests with math: "p95 CPU 0.7 cores at 400 RPS, so 700m; memory limit equals request at 1.1 GiB; no CPU limit."
- Probes: "Startup for boot, readiness for traffic, liveness in-process only with a generous timeout."
- Scaling chain: "HPA on in-flight requests, 15s loop, 5-minute scale-down window, 10% placeholder headroom for node lag."
- Failure domains: "3 zones, maxSkew 1, PDB maxUnavailable 1, 1.5× capacity to lose a zone at peak."
- Blast radius and ownership: "Clusters of a few hundred nodes as cells; platform owns control plane and add-ons; service teams own manifests and probes."
🎯 Staff Insight: Don't use Kubernetes as a database (etcd's 2 GiB default quota is for cluster metadata), as a scheduler of record for exactly-once jobs, or as a security boundary between hostile tenants. The strongest Kubernetes signal is saying what happens when the control plane is down: running pods keep serving and only changes stop.
Evaluation Rubric#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Workload contract | Deployments and Services | Requests, limits, probes, PDBs, spread set deliberately with numbers | Org standard enforced at admission, with exceptions process |
| Failure | "Pods restart" | Node eviction timing, zone loss capacity, control-plane-down behaviour | Cell boundaries, canary clusters, correlated-failure audit |
| Scaling | HPA on CPU | Demand signal, three-loop latency, headroom | Fleet capacity buffer and chargeback |
| Operations | "Monitor the cluster" | etcd quota, API server latency, upgrade routine | Upgrade cadence as funded work; cluster count as a cost driver |
| Adoption | Always Kubernetes | Knows when serverless or VMs fit better | Defines the trigger for adoption and the paved road |
Strong hire signals
| Signal | What It Sounds Like |
|---|---|
| Control-loop thinking | "Every manifest is a standing order; what will the controller do with it at 3am?" |
| Probe discipline | "Readiness drains, liveness restarts. Only one of those should ever see the database." |
| Knows the clocks | "15-second HPA loop, 5-minute scale-down, 5-minute eviction, minutes for a node." |
| Blast-radius sizing | "A cluster is a cell; I'd rather lose 10% of tenants than all of them." |
| Knows when not to | "Three services and five engineers: serverless containers." |
Lean no-hire signals
| Signal | Why It Misses the Bar |
|---|---|
| Kubernetes as the answer to availability | Rescheduling is minutes; redundancy is the mechanism |
| No requests or probes discussed | Will ship BestEffort pods and restart cascades |
| Application state in CRDs or etcd | Will freeze the cluster at the quota |
| One cluster for the company with no upgrade plan | Unpriced, uncontained blast radius |
Common false positives
- Fluent YAML ≠ judgment. Writing a Deployment from memory says nothing about whether the probe will cascade.
- Operator and CRD enthusiasm ≠ platform design. Ask who maintains the operator in year two.
- "We ran 2,000 nodes" ≠ knowing when not to. Ask what they'd recommend for four services.
Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
At Staff level Kubernetes is a runtime you configure correctly. At Principal level it is the company's operating model for compute: the platform team's product, the contract between that team and every service team, and a recurring cost that scales with cluster count, add-on count and the ~3 minor releases a year. The L7 question is not "Deployment or StatefulSet?" but "How many clusters should we have, what does the platform promise, what may teams customise, and what do we pay each year just to stay supported?"
🧭 Principal Move: "Before we talk about manifests, I want to agree on the paved road: what every service gets for free (deploy, scale, rollback, observability, a probe standard), what it must declare (requests, probes, owner), and what it may not do (state in etcd, cluster-admin, :latest). The YAML follows from that contract; the contract doesn't follow from the YAML."
The Org-Level Fault Line#
Shared multi-tenant clusters vs cluster-per-team.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| One shared cluster per region | Best bin-packing; one expert team; uniform policy | Company-wide blast radius; every upgrade needs every team's blessing; noisy neighbours | Every team, during the one bad afternoon |
| Shared clusters by tier or cell (critical, general, batch; or tenant cells) | Contained blast radius; tier-specific policies; staged upgrades | More clusters to upgrade; needs fleet automation | Platform headcount |
| Cluster per team | Maximum autonomy; clean cost attribution | 40 half-maintained clusters, drifting versions, duplicated add-ons, poor packing | Security and incident responders; finance through waste |
| No Kubernetes for some teams (serverless) | Zero platform cost for simple services | Two operating models; some tooling duplicated | Platform, to support both paths |
The Principal default: a small number of platform-owned clusters per region, split by tier and then by cell as the fleet grows, with namespaces as the team boundary, quotas and chargeback per namespace, and a serverless path for teams whose needs fit it. Cluster-per-team only for hard isolation requirements (regulated data, hostile code), and then from the same automated template.
🧭 Principal Insight: "The question isn't shared versus dedicated. It's who absorbs the upgrade treadmill. Shared clusters concentrate it in one team that gets good at it; cluster-per-team spreads it across 40 teams that each do it badly once a year."
Cost Model#
Assumptions: managed control plane ~$75/month per cluster; general node 16 vCPU / 64 GiB at ~$550/month on demand; observability and add-ons ~10–15% of node spend; loaded engineer $250K/year ($21K/month). Directional only.
| Scale | Clusters / nodes | Compute/month | Control planes + add-ons/month | Platform headcount | Total/month |
|---|---|---|---|---|---|
| Startup (8 services) | 2 / 12 | ~$6.6K | ~$1K | 0.5 FTE (~$10K) | ~$18K, versus ~$6–10K on serverless containers |
| Growth (120 services) | 8 / 250 | ~$140K | ~$18K | 3–4 FTE (~$75K) | ~$230K |
| Enterprise (1,500 services) | 60 / 3,500 | ~$1.9M | ~$250K | 15–20 FTE (~$370K) | ~$2.5M |
Two things stand out. At startup scale the platform engineer, not the compute, is the dominant cost, which is the quantitative case for not adopting Kubernetes early. At growth and enterprise scale, request efficiency is the lever: moving the fleet from ~35% to ~60% of requested CPU actually used removes roughly 40% of nodes, worth more than any other optimisation on this page.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse |
|---|---|---|
| Storing application state in CRDs/etcd | One-way once services depend on it | Data migration plus rewriting every controller that reads it |
| Pod and Service CIDR ranges, CNI with IP model | One-way for a running cluster | New cluster and migration |
| Cluster topology (shared vs per-team) | One-way-ish | Quarters of migration, workload by workload |
| Custom controllers and operators the business relies on | One-way-ish | Owning that code forever, or a migration |
| Managed provider (GKE vs EKS vs AKS) | Two-way, if manifests stay portable | Weeks per cluster; cloud-specific add-ons are the drag |
| Requests, probes, HPA targets | Two-way | Config change and rollout |
| Node instance types and pools | Two-way | Rolling node pool replacement |
| Service mesh adoption | Two-way-ish | Removing sidecars is easy; removing dependence on mesh features is not |
The Standard I'd Write#
RFC-PLAT-004: Workload Contract for Shared Kubernetes Clusters
Scope: Every workload in a platform-managed cluster, all environments above dev.
MUST
1. Declare CPU and memory requests on every container; memory limit = request.
2. Liveness probes check in-process health only; no network calls to
dependencies. Readiness may reflect dependencies. Startup probe if boot > 10s.
3. Run >= 2 replicas (>= 3 for critical tier) with a PodDisruptionBudget that
allows at least 1 disruption, and zone topology spread with maxSkew 1.
4. Use immutable image tags or digests; no :latest.
5. Store no business data in custom resources or ConfigMaps.
6. Declare an owning team label that routes pages.
SHOULD
7. Omit CPU limits on latency-critical services; set them on batch.
8. Scale on a demand signal (in-flight requests, queue lag) rather than CPU.
9. Handle SIGTERM: fail readiness, wait >= 5s, drain, exit before grace period.
Exceptions: Filed with the platform team; approved by platform lead and the
owning team's director; time-boxed to one quarter; listed on the exceptions dashboard.
Enforcement: Admission policy in warn mode for one quarter, then enforce.
Shadow -> warn -> canary cluster -> enforce everywhere.
Success metrics: zero probe-caused incidents per quarter; fleet request
efficiency >= 60%; every cluster within the supported version window;
median new-service time to production < 1 day.
What I'd Tell the VP#
"Kubernetes is how we run nearly all our services now, which means our clusters are a product with a yearly maintenance bill, not a one-time project. Today we have one large cluster per region, so a single bad upgrade can stop every team from deploying, and we use about a third of the computing capacity we pay for. I'm proposing three changes: split clusters so a failure hits a slice of customers rather than all of them, enforce a simple contract for every service so one bad health check can't restart a whole product, and charge teams for the capacity they reserve. That should cut our compute bill by roughly a third and turn platform outages into partial incidents. The cost is two more platform engineers and one quarter of migration work."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Treats the platform as a product | "The paved road is a contract: what we guarantee, what teams declare, what they may not do." |
| Prices the upgrade treadmill | "Three minors a year times 30 clusters is a standing team, so cluster count is a cost decision." |
| Sizes blast radius deliberately | "Cells of a few hundred nodes; a bad rollout hits 10% of tenants and the canary cluster first." |
| Prices requests, not usage | "We pay for what's requested. Chargeback on requests fixes right-sizing faster than any tool." |
| Knows when not to standardise | "Teams with three stateless services get serverless containers; forcing them onto clusters costs more than it saves." |
Staff answers that L7 interviewers find insufficient:
- "We'll run multiple clusters for isolation" without the upgrade cost, fleet tooling or how traffic shifts between them.
- "Enforce requests and probes" as advice, rather than an admission policy with a rollout path and an exceptions process.
- "Use managed Kubernetes to reduce ops" without saying which ops remain: node pools, add-ons, upgrades, policies and cost.
How Real Companies Built It#
Google Borg — The Ancestor#
Kubernetes descends from Borg, the cluster manager Google described in its EuroSys 2015 paper. Borg runs hundreds of thousands of jobs across clusters of up to tens of thousands of machines, and reaches high utilisation by combining admission control, efficient task packing, over-commitment and machine sharing with process-level performance isolation. The paper's lessons section discusses how Kubernetes revisited several Borg design choices, such as organising pods with labels rather than a fixed job-to-task grouping (Borg paper, EuroSys 2015).
Staff insight: Kubernetes inherited Borg's core bet: declare intent, let a scheduler bin-pack against requests, accept over-commitment for utilisation. When you discuss requests and QoS, you are discussing Borg's utilisation economics.
Pokémon GO on Google Kubernetes Engine — Launch at 50× Plan#
Niantic launched Pokémon GO in 2016 on Google Kubernetes Engine (then Google Container Engine). Google's account says the team planned for 1× traffic with a worst case of 5×, and actual traffic hit 50×. It describes the deployment as the largest Kubernetes Engine deployment at the time, adding over a thousand nodes, performing a live cluster version upgrade during the launch period, and moving to Google's HTTP/S load balancing. The Japan launch that followed is described as running without incident (Google Cloud blog).
Staff insight: The cluster scaled because nodes could be added and pods rescheduled, but the hard parts were the edges: load balancing and the version upgrade under load. In an interview, capacity buffer, the edge and the upgrade path are where a Kubernetes design gets tested.
OpenAI — 7,500-Node Clusters for ML#
OpenAI described scaling single Kubernetes clusters to 7,500 nodes to run research and training workloads including GPT-3, CLIP and DALL·E. Its workloads often place a single pod on an entire node to use all of its GPUs, and it describes moving off flannel to native pod networking for throughput, running API servers and etcd on dedicated nodes, and closely monitoring API server error rates (OpenAI engineering).
Staff insight: At the top of the envelope, the limits are the control plane and the network, not the pods. A batch or ML design on Kubernetes should name API server load, etcd placement and the CNI before it names the training framework.
Reddit — The Pi-Day Upgrade Outage#
On March 14, 2023, Reddit suffered an outage of roughly 314 minutes that began during an upgrade of a cluster from Kubernetes 1.23 to 1.24. Calico route reflectors were selected by the master node label, which 1.24 no longer applies in favour of control-plane, so no route reflectors ran and pod networking failed. The configuration had been set up years earlier outside version control (Reddit engineering post-mortem).
Staff insight: The upgrade itself was routine; the failure came from configuration no one knew existed. That is the case for a canary cluster with the same add-ons as production, every add-on's config in Git, and treating upgrades as a staged rollout instead of a maintenance task.
Practice Drill#
Prompt: "You run the platform for a 150-engineer company: 90 services on one managed Kubernetes cluster per region, 400 nodes each. Last month a database brownout turned into a 40-minute full outage, and the cluster is two minor versions behind. Redesign the platform."
Staff Answer
Start with the outage, because it is the cheaper fix. A 20-second database slowdown becoming 40 minutes of downtime is the signature of liveness probes that check dependencies: every pod failed liveness within ~30s, restarted, and reconnected with cold pools. I'd ship a probe standard immediately: liveness in-process only with timeoutSeconds: 5 and failureThreshold: 6; readiness may reflect the database; startup probes for the JVM services with 40–60s boot. Enforce it through an admission policy in warn mode for two weeks, then enforce, and add a chaos test that slows the database by 5s in staging and asserts zero restarts. Second, the workload contract: requests on every container from VPA recommendations, memory limit = request, no CPU limits on API tiers, PDBs that allow at least one disruption (several services currently have maxUnavailable: 0, which is also why upgrades are stuck), and zone spread with maxSkew: 1. Third, the upgrade debt: build a canary cluster with the same add-ons, scan for deprecated APIs, and upgrade one minor at a time each quarter until current, then keep the ~3-per-year cadence. Fourth, blast radius: split each region into a critical cluster (auth, checkout, payments; Guaranteed QoS, strict change control) and a general cluster, so a bad add-on upgrade can't take both. State stays in managed databases outside the clusters. Metrics: correlated restarts per Deployment, etcd size vs quota, API server p99, days to end of support, request efficiency.
Why this is L6:
- Diagnoses the cascade from the timeline (DB blip → fleet restart) rather than "add capacity".
- Connects the stuck upgrades to the PDB misconfiguration, one root cause with two symptoms.
- Rolls policy out shadow → warn → enforce with a chaos test proving the fix.
- Splits clusters by criticality to contain the next incident.
What L7 adds:
- Frames the platform as a product with a contract, an exceptions process and success metrics (zero probe-caused incidents, all clusters in support).
- Prices it: chargeback on requested cores, expected 25–35% node reduction from right-sizing, against two extra platform engineers.
- Sets the trigger for moving from tiered clusters to cells, and offers serverless containers for the long tail of small internal services.
Quick Reference Card#
Model: control loop over desired state; every object is a standing order
Source of truth: etcd via API server only; quota 2 GiB default, 8 GiB suggested max,
1.5 MiB max request; Events in a separate etcd for large clusters
Control plane down: running pods keep serving; deploys, scaling, rescheduling stop
Cluster limits: 110 pods/node, 5,000 nodes, 150,000 pods, 300,000 containers
Scheduler: places by REQUESTS (never usage); filter -> score -> bind
Limits: CPU over limit = throttled (CFS, 100ms periods)
memory over limit = OOM kill
QoS: Guaranteed > Burstable > BestEffort (evicted first)
Probe defaults: period 10s, timeout 1s, failureThreshold 3, success 1, delay 0
Probe rule: readiness drains traffic, liveness restarts; liveness never checks deps
Termination: terminationGracePeriodSeconds 30s default; preStop sleep before drain
Node loss: NotReady after grace period, eviction after 300s toleration (~5-6 min)
HPA: 15s sync, 10% tolerance, scale-up immediate, 300s scale-down window;
CPU utilisation is % of request (no request = no CPU scaling)
EndpointSlice: 100 endpoints per slice default, 1,000 max; only Ready pods
Releases: ~3 minors/year, ~14 months support each, 3 branches maintained
Zone loss math: 3 zones -> 1.5x peak capacity to survive one at peak
RED FLAGS
- Liveness probe that calls the database or a downstream service
- Containers without requests (BestEffort in production)
- CPU limits on multi-threaded latency-critical services
- One cluster as the whole company's blast radius
- Business data in CRDs or ConfigMaps
- PDB that allows zero disruptions (upgrades blocked)
- Fail-closed webhooks with no namespace exemptions
- Cluster more than one minor version behind with no upgrade owner