Hiring BarSupport

Kubernetes

Technology guide52 min read7 diagrams

Why This Matters#

Kubernetes is not a deployment tool. It is a control loop over desired state: you write down what should exist, dozens of controllers keep comparing that to what does exist, and they act on the difference. Containers are the least interesting part. Most production incidents on Kubernetes come from the control plane and its defaults: a liveness probe that kills healthy pods under load, a pod with no CPU request that the scheduler packs onto a full node, an etcd that hits its 2 GiB quota and freezes every deploy in the company, an upgrade that silently drops a label some forgotten config depended on.

That is why "we'll run it on Kubernetes" is a sentence interviewers push on. The L5 candidate draws a box labelled "K8s" around the services and moves on. The L6 candidate says "each service is a Deployment with CPU and memory requests sized from load tests, a readiness probe that checks only the process itself, an HPA on requests-in-flight with the default 5-minute scale-down window, a PodDisruptionBudget so node drains can't take out quorum, and spread across 3 zones with topology constraints." The L7 candidate asks how many clusters the company should have, who runs them, and what it costs to keep 40 of them within the 14-month support window, because in year three the upgrade treadmill, not the YAML, is the bill.

The L5 → L6 gap is not knowing what a kubelet is. It is knowing that every Kubernetes object is a promise a controller will try to keep forever, so a wrong promise (a bad probe, a missing request, a too-tight budget) is enforced just as diligently as a right one.

The L5 → L6 → L7 Contrast#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Containerize the services and deploy to Kubernetes""What is the unit of failure and scaling? That decides Deployment vs StatefulSet vs Job, and which of these belong on Kubernetes at all.""Should this team run a cluster, use the shared platform, or skip Kubernetes? That's a headcount and blast-radius decision, not a YAML one."
Health"Add liveness and readiness probes"Readiness gates traffic; liveness only detects a wedged process and never checks dependencies; startup probe covers slow bootWrites the org probe standard and audits it, because one copy-pasted liveness template can restart a whole fleet during a database blip
Resources"Set limits so pods don't use too much"Requests sized from measured p95 usage; memory limit = request; CPU limit often omitted for latency services; QoS class chosen deliberatelyPrices bin-packing: requests are what the company pays for; sets namespace quotas and chargeback by requested, not used, cores
Scaling"HPA on CPU at 70%"HPA on a demand signal (in-flight requests, queue lag), knows the 15s loop and 300s scale-down window, pre-provisions node headroomDecides cluster autoscaler vs node-pool strategy fleet-wide and the capacity buffer the business pays for
Failure"Kubernetes restarts failed pods"Names what happens when etcd, the API server or a zone fails: running pods keep serving, changes stopDesigns cell boundaries so no single cluster or control plane is the whole company's blast radius
Ownership"DevOps runs the cluster"Platform owns nodes, control plane and add-ons; service teams own manifests, probes and requestsDefines the paved road: what the platform guarantees, what teams may customise, and who pays for upgrades
Why "Health" separates levels

"Add liveness and readiness probes" is correct and incomplete in the most dangerous way. A liveness probe that calls /health, which in turn pings the database, turns a 20-second database slowdown into a fleet-wide restart: every pod fails liveness at once, every pod restarts, and every pod reconnects to the already-struggling database with a cold cache. The Kubernetes documentation itself warns that incorrect liveness probes can cause cascading failures. The Staff answer separates the questions: readiness asks "should I receive traffic right now?" and may consult dependencies; liveness asks "is this process wedged beyond self-recovery?" and must not. The Principal answer turns that into a standard with a linter, because a fleet of 300 services will otherwise copy whichever template was nearest.

Why "Resources" separates levels

The scheduler places pods by requests, not by actual usage and not by limits. A pod with no request is, to the scheduler, free, so it lands on whatever node has room on paper. Under load it competes with everything else and is the first candidate for eviction. A CPU limit, meanwhile, is enforced by the kernel's CFS quota in 100ms periods: a 4-thread service with a 0.5-CPU limit can burn its 50ms quota in 12.5ms of wall-clock time and then sit throttled for the remaining 87.5ms, adding up to ~90ms of latency to requests that looked cheap on average. The Senior answer sets limits; the Staff answer sets requests from data and treats CPU limits as a latency risk; the Principal answer recognises that the sum of requests is the compute bill.

The 60-Second Pitch#

"I'd run the stateless services on a managed Kubernetes cluster per region, spread over 3 zones. Each service is a Deployment with requests sized from load tests, a readiness probe on the process's own ability to serve, a startup probe for the 40-second JVM boot, and no dependency checks in liveness. Scaling is an HPA on in-flight requests per pod, with the cluster autoscaler adding nodes and ~10% placeholder headroom so a burst doesn't wait 2–3 minutes for a VM. Kubernetes gives us self-healing, bin-packing and a uniform deploy and rollback path for 40 services. It does not give us a database: the stateful core stays on a managed database, and nothing application-level lives in etcd. If the control plane goes down, running pods keep serving; only changes stop. I'd keep each cluster small enough that losing one is an incident, not an outage."

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Stateless service platformMany services, frequent deploys, uniform opsDeployments + HPA + Services; managed control plane; paved-road templatesBad probe or missing requests causes fleet-wide restarts or noisy neighboursZero-downtime deploys; p99 unaffected by node loss
Batch and ML computeBin-packing, queueing, GPUs, preemptionJobs/CronJobs, priority classes, a batch scheduler (Kueue, Volcano), large node poolsThousands of pods hammer the API server; scheduler throughput becomes the bottleneckEvery job runs once to completion; fair share across teams
Stateful systems on KubernetesStable identity, persistent volumes, ordered rolloutStatefulSets + operators + PVCs + PDBs; anti-affinity across zonesZone loss strands volumes; operator bug or drain violates quorumNo data loss; quorum never broken by voluntary disruption

🎯 Staff Move: "I'll design for the first intent: Kubernetes as the stateless service platform. The payments ledger stays on a managed database outside the cluster. Running stateful systems on Kubernetes is a legitimate second step, but only behind a mature operator and only once the team has rehearsed zone loss, so I won't make the core design depend on it."

The Staff Positions#

PositionRationale
Requests on every container, set from measurementThe scheduler, the HPA and eviction all reason from requests. No request means no placement guarantee and no CPU-based autoscaling.
Liveness never checks dependenciesA dependency outage should drain traffic (readiness), not restart the fleet (liveness).
Memory limit = memory request; CPU limit only where justifiedMemory overcommit ends in OOM kills at the worst moment; CPU limits add throttling latency for little benefit on latency-sensitive services.
Running pods must not depend on the control planeetcd, API server and scheduler are for changes. The data plane keeps serving the last state if they disappear.
Many medium clusters beat one huge oneA cluster is a blast radius and an upgrade unit; 5,000 nodes is a documented ceiling, not a target.
Application state does not go in etcdetcd holds cluster metadata under a 2 GiB default quota; CRDs used as a database fill it and freeze the cluster.
Every workload has a PDB and zone spreadNode drains and upgrades are routine; without a budget they take out quorum or capacity.

Architecture & Internals#

Only five internals change design decisions: the API server and etcd, controllers and the reconcile loop, the scheduler, the kubelet and probes, and Services/EndpointSlices.

The Control Plane: One Front Door, One Source of Truth#

Every object (Pod, Deployment, Service, ConfigMap, Secret, custom resource) lives in etcd and is read and written only through the API server. Controllers, the scheduler and every kubelet talk to the API server, mostly through long-lived watches. The API server keeps a watch cache so that thousands of clients don't watch etcd directly. The consensus mechanics, quotas and compaction of etcd are covered in etcd & ZooKeeper; here the point is what they mean for workloads.

Diagram: The Control Plane: One Front Door, One Source of Truth

Why it matters in design: requests from users never touch the control plane. If etcd loses quorum or the API server is down, existing pods keep running, kube-proxy keeps its programmed rules and traffic flows. What stops is everything that changes: deploys, scaling, rescheduling after a node failure, new endpoints. The Kubernetes etcd operations guide is blunt about it: if etcd is starved, no cluster state changes are possible and no new pods can be scheduled. That is the canonical control-plane/data-plane split, and it is the property you should draw when someone asks "what if the control plane dies?"

The Reconcile Loop#

Every controller runs the same loop: observe current state via a watch, diff against desired state, act to close the gap, repeat. Nothing is a one-shot command. A Deployment does not "deploy"; it declares that a ReplicaSet with template hash X should have 12 replicas, and the Deployment and ReplicaSet controllers keep making that true.

loop forever:
    desired = spec from the API server      (e.g. replicas: 12)
    actual  = observed status               (e.g. 10 Running, 1 Pending, 1 Failed)
    diff    = desired - actual
    if diff != 0:
        act(diff)                           (create 2 pods, delete the failed one)
    write status back
    wait for next watch event or resync

Three consequences candidates miss:

  1. Level-triggered, not edge-triggered. A controller that misses an event still converges on the next one, because it compares states, not deltas. That is why Kubernetes survives controller restarts without a replay log.
  2. Controllers fight. An HPA setting replicas: 18 and a GitOps tool reapplying replicas: 6 from Git will flap forever. One owner per field; omit replicas from Git when an HPA owns it.
  3. A wrong spec is enforced with the same diligence as a right one. If the spec says "restart when /healthz fails" and /healthz checks the database, the controller will faithfully restart every pod during a database brownout.

🎯 Staff Insight: "I treat every manifest as a standing order to a robot that never gets tired. So the review question isn't 'does this deploy?', it's 'what will the controllers do with this spec at 3am when a dependency is slow?'"

The Scheduler: Filter, Score, Bind#

The scheduler watches for pods with no node, then for each pod filters nodes that can't host it (insufficient requested CPU or memory, taints without tolerations, node selectors, affinity rules, volume zone constraints), scores the survivors (spread, resource balance, image locality, preferred affinity), and binds the pod to the winner. It never looks at actual utilisation.

Placement controlWhat it doesUse for
resources.requestsReserves capacity for placementEvery container, always
nodeSelector / node affinityRestricts to labelled nodesGPU pools, ARM pools, compliance pools
Taints + tolerationsKeeps pods off nodes unless toleratedDedicated pools (system, GPU, batch)
Pod anti-affinityKeeps replicas apartQuorum members on separate nodes
topologySpreadConstraintsBounded skew across zones or nodesSpread 12 replicas 4/4/4 over 3 zones, maxSkew: 1
priorityClassNameHigher priority can preempt lowerSystem and critical services over batch

The trap: affinity rules are evaluated per pod against every node. Heavy inter-pod affinity in a large cluster is the classic way to slow the scheduler from thousands of pods per minute to a crawl. Prefer topology spread constraints, which cover most of the same intent more cheaply.

Pod Lifecycle: From kubectl apply to Receiving Traffic#

Diagram: Pod Lifecycle: From kubectl apply to Receiving Traffic

Every arrow is a watch, so every step has propagation delay. In a healthy cluster a pod goes from bound to receiving traffic in seconds to tens of seconds, dominated by image pull and the startup and readiness probes. The interval between "pod removed from EndpointSlice" and "every proxy has stopped sending to it" is why graceful shutdown needs a preStop delay; the full mechanics are in Service Registry.

Kubelet, Probes and Node Health#

The kubelet is the node agent: it runs containers via the container runtime, executes probes, enforces resource limits through cgroups, evicts pods under node pressure, and reports node and pod status to the API server.

ProbeQuestion it answersEffect on failureWho should define it
startupProbe"Has the process finished booting?"Liveness and readiness are held off until it passes; fails → restartService team; sized to p99 boot time
readinessProbe"Should I receive traffic right now?"Pod removed from EndpointSlices; no restartService team; may reflect dependency health or load
livenessProbe"Is this process wedged beyond self-recovery?"Container restartedService team, under a platform standard: no dependency checks

Probe defaults: initialDelaySeconds 0, periodSeconds 10, timeoutSeconds 1, successThreshold 1, failureThreshold 3 (Kubernetes probe docs). With defaults, a pod is marked unready or restarted roughly 30 seconds after it starts failing, and a 1-second timeout is what turns a GC pause or a slow response under load into a probe failure.

Node failure timeline. When a node stops reporting, the node lifecycle controller marks it NotReady after the node-monitor grace period (tens of seconds), taints it, and pods on it are evicted only after the default 300-second toleration for not-ready and unreachable taints added by the DefaultTolerationSeconds admission plugin (well-known taints). So with defaults, pods on a dead node take about 5–6 minutes to be recreated elsewhere. Readiness removes them from traffic far sooner, which is why replicas across zones, not rescheduling, is the availability mechanism.

🎯 Staff Insight: "Kubernetes 'self-healing' takes about six minutes for a dead node with default tolerations. My availability comes from N+1 replicas spread across zones and readiness removing dead endpoints. Rescheduling restores capacity; it doesn't restore availability."

Services and EndpointSlices#

A Service is a stable virtual IP and DNS name in front of a label selector. The EndpointSlice controller watches pods matching the selector and writes their IPs into EndpointSlice objects, including only pods whose readiness probe passes. By default each slice holds up to 100 endpoints (configurable up to 1,000) to keep update fan-out small (EndpointSlices). kube-proxy (iptables, IPVS or nftables mode) or a mesh control plane watches those slices and programs every node or sidecar.

Service typeWhat you getUse for
ClusterIPIn-cluster virtual IP, L4 load balancing per connectionDefault service-to-service
Headless (clusterIP: None)DNS returns pod IPs directlyStatefulSets, client-side load balancing, gRPC
NodePort / LoadBalancerExposes through node ports or a cloud L4 load balancerEdge entry, non-HTTP protocols
Ingress / Gateway APIL7 routing, TLS, host and path rulesHTTP edge; see Envoy, Kong & NGINX

The gRPC trap: ClusterIP balances connections, not requests. A gRPC client holding one long-lived HTTP/2 connection sends every request to one pod. Use a headless Service with client-side balancing, or a mesh that balances per request. The full discovery design, including mass-deregistration guards and panic thresholds, lives in Service Discovery.


Core Usage — "The Entire Game": The Workload Contract#

In Kafka the game is the partition key. In Kubernetes it is the workload contract: the five things every service team declares and the platform enforces. Get these right and the cluster mostly runs itself; get them wrong and the controllers will enforce the mistake at fleet scale.

The workload contract (per container):
  1. requests    -> where it can be placed and what it is guaranteed
  2. limits      -> when it is throttled (CPU) or killed (memory)
  3. probes      -> when it gets traffic and when it gets restarted
  4. disruption  -> how many replicas may be down during drains and upgrades
  5. spread      -> which failures (node, zone) it survives

Step 1: Requests and Limits — Size From Measurement#

The scheduler uses requests to decide placement; the kernel enforces limits. CPU over the limit is throttled; memory over the limit gets the container OOM-killed when the kernel detects memory pressure (resource management docs).

QoS classHow you get itEviction order under node pressureUse for
GuaranteedRequests = limits for CPU and memory, every containerLastLatency-critical, stateful, quorum members
BurstableRequests set, limits higher or absentMiddle (by usage over request)Most stateless services
BestEffortNo requests, no limitsFirstNothing in production

Sizing recipe (a common production pattern):

memory.request = memory.limit = p99 working set under peak load test x 1.2
cpu.request    = p95 CPU at target RPS per pod
cpu.limit      = unset for latency services (or >= 2-4x request if policy demands one)

Example: checkout-api at 400 RPS per pod
  load test:  p95 CPU 0.7 cores, p99 working set 900 MiB
  requests:   cpu 700m, memory 1.1Gi
  limits:     memory 1.1Gi, no CPU limit
  node:       16 vCPU / 64 GiB, ~1 vCPU and ~3 GiB reserved for system
  pods/node:  floor(15 / 0.7) = 21 by CPU, floor(61 / 1.1) = 55 by memory -> CPU-bound, 21

Why memory limit = request: memory is not compressible. If 21 pods each request 1.1 GiB but are allowed to burst to 3 GiB, the node is overcommitted ~2.7× on paper, and the first traffic spike triggers OOM kills across pods that did nothing wrong. Why no CPU limit: CPU is compressible. Without a limit, a pod borrows idle cycles and the request still guarantees its fair share under contention. With a limit, the CFS quota throttles it inside each 100ms period even when the node is idle. Watch container_cpu_cfs_throttled_periods_total; a ratio over ~5–10% of periods on a latency service is a p99 problem.

🎯 Staff Move: "I'll set requests from the load test, memory limit equal to request, and no CPU limit on the API pods. I'd rather manage noisy neighbours with requests and namespace quotas than pay 50–100ms of throttling at p99. Batch pools are different: there I'd set CPU limits so a runaway job can't starve its neighbours."

Step 2: Probes — Three Questions, Three Probes#

startupProbe:            # p99 boot is ~40s; allow 60s before liveness takes over
  httpGet: { path: /startup, port: 8080 }
  periodSeconds: 5
  failureThreshold: 12   # 12 x 5s = 60s budget
readinessProbe:          # "send me traffic" - may reflect load and critical deps
  httpGet: { path: /ready, port: 8080 }
  periodSeconds: 5
  timeoutSeconds: 2
  failureThreshold: 2    # ~10s to drain a sick pod
livenessProbe:           # "am I wedged?" - in-process only, never the database
  httpGet: { path: /live, port: 8080 }
  periodSeconds: 10
  timeoutSeconds: 5      # generous: a GC pause is not a deadlock
  failureThreshold: 6    # ~60s of consecutive failure before restart

The design rule: readiness is fast and sensitive, liveness is slow and conservative. Readiness false costs one pod's share of traffic for a few seconds. Liveness false costs a restart, a cold cache, a reconnect storm and, if correlated across the fleet, an outage.

Diagram: Step 2: Probes — Three Questions, Three Probes
Probe checkReadinessLiveness
Process can serve HTTPYesYes
Thread pool or event loop not deadlockedYesYes, this is the point
Database reachableMaybe, if the pod is useless without itNever
Downstream service healthyRarely; prefer circuit breakersNever
Warm cache loadedYesNo
In-flight requests above shed thresholdYes (load shedding)No

Step 3: Autoscaling — Three Loops With Different Clocks#

LoopWhat it scalesDefault clockReaction time in practice
HPAReplicas of a DeploymentSync every 15s; scale-up stabilisation 0s; scale-down stabilisation 300s; tolerance 10%15–60s to decide, plus pod start
Cluster autoscaler / KarpenterNodesReacts to pods Pending for lack of capacityVM boot + image pull: often 1–3 minutes
VPARequests per podRecommends from usage historyHours to days; applies by restarting pods

The HPA computes desiredReplicas = ceil(currentReplicas × currentMetric / targetMetric), ignores changes within the 10% tolerance, and calculates CPU utilisation as a percentage of the request, so a container with no CPU request cannot be autoscaled on CPU at all (HPA docs).

Example: checkout-api, 30 pods, target 200 in-flight requests per pod
  Black Friday spike: observed 340 per pod
  desired = ceil(30 x 340 / 200) = ceil(51.0) = 51 pods
  51 x 0.7 CPU = 35.7 cores requested; existing headroom 12 cores
  -> ~18 pods Pending -> cluster autoscaler adds 2 nodes -> ~2 minutes
  -> placeholder pods (low priority, ~10% of cluster) are preempted instantly,
     so ~17 of the 21 new pods start in seconds instead of minutes

The Staff point is that the three loops compound: HPA decides in ~15–30s, but if the new pods need new nodes, the user waits for a VM boot. Capacity headroom (low-priority placeholder pods that get preempted) is the cheap fix. The full capacity story, including asymmetric step caps and predictive scaling, is in Autoscaling & Capacity.

🎯 Staff Insight: "CPU is a lagging, indirect signal for an I/O-bound API. I'd scale on in-flight requests or queue lag via a custom metric, keep the 5-minute scale-down window, and hold 10% of the cluster as preemptible placeholder pods so a burst doesn't wait on a VM."

Step 4: Disruption Budgets and Rollouts#

A PodDisruptionBudget limits how many replicas voluntary disruptions (node drains, cluster upgrades, autoscaler scale-down) may take down at once. It does not prevent involuntary disruptions such as hardware failure, though those count against the budget (disruptions docs).

strategy:
  type: RollingUpdate
  rollingUpdate:
    maxSurge: 25%          # extra pods during rollout (needs spare capacity)
    maxUnavailable: 0      # never dip below desired replicas
minReadySeconds: 20        # a pod must stay Ready 20s before counting
---
kind: PodDisruptionBudget
spec:
  maxUnavailable: 1        # for a 3-member quorum: never drain 2 at once
  selector: { matchLabels: { app: ledger-consumer } }

Two traps: a PDB of minAvailable: 100% (or maxUnavailable: 0) blocks every node drain, so the cluster can never be upgraded; and a Deployment with 1 replica and any PDB is the same thing in disguise. The platform should reject both at admission.

Step 5: Spread Across Failure Domains#

topologySpreadConstraints:
- maxSkew: 1
  topologyKey: topology.kubernetes.io/zone
  whenUnsatisfiable: DoNotSchedule      # hard for zones
  labelSelector: { matchLabels: { app: checkout-api } }
- maxSkew: 2
  topologyKey: kubernetes.io/hostname
  whenUnsatisfiable: ScheduleAnyway     # soft for nodes
  labelSelector: { matchLabels: { app: checkout-api } }

The availability arithmetic: to survive a zone loss at peak with 3 zones, you need 1.5× peak capacity spread evenly (each zone carries 50% of peak so two zones carry 100%). That 50% buffer is a business decision priced in nodes, and it is the most common place Kubernetes designs quietly assume capacity they haven't paid for.


The Tunable Tradeoff — Density vs Isolation#

Every Kubernetes platform decision moves along one axis: how tightly do we pack workloads together? Dense packing is cheap; isolation is predictable.

SettingDense endIsolated endWho pays at the dense end
Requests vs actual usageRequests below usage (overcommit)Requests at p95–p99 usageService teams, as OOM kills and throttling
QoSBurstable / BestEffortGuaranteedWhoever is evicted first
Pods per nodeNear the 110 defaultFewer, larger podsEveryone on the node when it dies or is noisy
Node poolsOne shared poolDedicated pools by tier (system, critical, batch, GPU)Platform, in utilisation; product teams, in noisy neighbours
ClustersOne per region for the companyPer domain or per cellEvery team during the one bad upgrade
TenancyNamespaces only (soft)Separate clusters or sandboxed runtimes (hard)Security, if a tenant is hostile
Utilisation formula the finance team cares about:
  cluster_efficiency = sum(actual usage) / sum(node capacity)
  request_efficiency = sum(actual usage) / sum(requests)
  allocation_ratio   = sum(requests)     / sum(allocatable)

Typical unmanaged fleet: request_efficiency 30-40%, allocation_ratio 70-80%
  -> ~25-30% real utilisation: you pay for ~3-4x what you use.
Staff target for stateless tiers: request_efficiency 60-70% via VPA
  recommendations + right-sizing reviews, allocation_ratio ~85% with headroom pods.

🎯 Staff Move: "I'll pack stateless services densely in a shared Burstable pool, keep quorum members and the payment path Guaranteed in a dedicated pool, and push batch onto a preemptible pool with lower priority. Density where failure is cheap, isolation where it isn't."

Who Pays for Each Choice#

ChoiceWhat WorksWhat BreaksWho Pays
No CPU limitsNo throttling latency; idle cycles usedA runaway pod can consume a node's spare CPUNeighbouring Burstable pods, mildly
CPU limits everywherePredictable per-pod ceilingsp99 latency from CFS throttlingService team's SLO
Overcommitted memory20–40% more pods per nodeOOM kills under correlated loadService on-call at peak
Aggressive liveness probesWedged pods recovered fastFleet restarts on dependency blipsEvery user during the cascade
Tight PDBsQuorum never broken by drainsUpgrades block for daysPlatform team's upgrade schedule
Single giant clusterOne place to operate, best packingCompany-wide blast radius per control-plane incidentEvery team at once

Anti-Patterns — What Kills Kubernetes Deployments#

1. Liveness Probes That Check Dependencies#

/health pings the database; the database has a 20-second slowdown; every pod fails liveness within ~30s and restarts; every restarted pod reconnects with a cold cache and connection-pool warm-up, amplifying load on the database that caused it. The official docs list exactly this: incorrect liveness probes cause restarts under high load, failed client requests and extra load on the remaining pods (probe concepts). Fix: liveness checks only in-process state; dependency health goes in readiness (carefully) or circuit breakers.

2. No CPU (or Memory) Requests#

To the scheduler, a pod without requests costs nothing, so a node can be packed with them until real usage hits 100%. They are BestEffort and evicted first; the HPA can't compute utilisation for them. Fix: an admission policy (LimitRange defaults plus a policy engine) that rejects production pods without requests.

3. CPU Limits on Latency-Critical, Multi-Threaded Services#

A JVM or Go service with 8 busy threads and a 1-CPU limit exhausts its 100ms quota in 12.5ms and stalls for 87.5ms. Average CPU looks like 40%; p99 is terrible. Fix: requests without CPU limits on latency tiers; watch throttled-period ratio.

4. A Single Cluster as the Whole Company's Blast Radius#

One cluster per region, every team on it. One bad admission webhook, CRD upgrade, CNI change or etcd quota breach freezes deploys or breaks networking for every service at once. Fix: multiple clusters by tier or cell, staged upgrades (canary cluster first), and the ability to shift traffic between clusters.

5. Kubernetes as a Database for Application State#

Custom resources used to store user sessions, job payloads or per-tenant records. Every object lands in etcd, under a 2 GiB default quota (8 GiB suggested maximum) and a 1.5 MiB request limit (etcd limits). When the quota trips, etcd goes read-only and nothing in the cluster can change. Fix: CRDs describe infrastructure intent (a desired database, a desired certificate), not business data. Business data goes in a database.

6. Fail-Closed Admission Webhooks Without an Escape Hatch#

A validating webhook with failurePolicy: Fail whose backing pods run in the same cluster: the webhook deployment crashes, no pod can be created, including the webhook's own replacement. Fix: exclude kube-system and the webhook's namespace, run 3+ replicas with a PDB, set short timeouts, and choose Ignore for anything that isn't a security control.

7. latest Tags and Mutable Images#

Pods on different nodes run different code under the same tag; rollback re-pulls the same broken image. Fix: immutable tags or digests; admission policy rejects :latest.

8. Ignoring the Upgrade Treadmill#

Minor releases arrive about 3 times a year, and each is supported for roughly 14 months (release cadence, patch support). A team that upgrades once a year is always rushing two versions at once, against deprecated APIs and add-ons nobody remembers installing. Fix: upgrades as a quarterly routine with an owner, a canary cluster and an API-deprecation scanner in CI.


The Technology Landscape — Head-to-Head Comparison#

DimensionKubernetes (managed: GKE/EKS/AKS)Self-managed KubernetesServerless containers (Cloud Run, Fargate, ECS)NomadPlain VMs + autoscaling groups
ModelDeclarative reconcile over a rich APISame, you run the control planeRun a container, scale on requestsSingle-binary scheduler, jobs and servicesMachine images, instance groups
Scale ceiling5,000 nodes / 150,000 pods per cluster (documented)SamePer-service quotas; no cluster conceptLarge; simpler modelAccount and region quotas
Ops burdenMedium: node pools, add-ons, upgradesHigh: etcd, certificates, control-plane upgradesLowLow–mediumLow–medium
EcosystemLargest: operators, Helm, service mesh, GitOpsSameThinModerate (HashiCorp stack)Cloud-native tooling
Stateful workloadsPossible with operators and PVCsSameMostly noPossibleNatural
Cold start / scale-upSeconds if nodes exist; minutes if notSameSeconds; scale to zeroSecondsMinutes
Pick when20+ services, platform team, portabilityRegulated or on-prem with deep expertiseFew services, small team, spiky trafficMixed workloads, simpler ops, HashiCorp shopFew large services, VM-shaped workloads

🎯 Staff Insight: "Managed Kubernetes removes etcd and the control plane from my pager, not the node pools, add-ons, probes, upgrades or cost. The real comparison for a small team is Kubernetes versus serverless containers, and below ~10 services the serverless option usually wins."


Patterns#

Pattern 1: Paved-Road Stateless Service#

Deployment + HPA + PDB + topology spread + ServiceAccount with least privilege + NetworkPolicy, generated from one template (Helm chart or a higher-level CRD like kind: WebService). Teams fill in image, port, requests and probe paths; the platform owns everything else. This is the default for 80%+ of services.

Pattern 2: Operators for Stateful Systems#

An operator is a custom controller that encodes the runbook of a stateful system (Postgres, Kafka, Elasticsearch): failover, backups, version upgrades, resizing, as a reconcile loop over a CRD. Use when the operator is mature, widely used and owned by someone who will patch it; avoid a home-grown operator for a database the team doesn't deeply understand. An operator automates expertise; it does not replace it.

Pattern 3: Batch Queues on Shared Clusters#

Jobs with priorityClassName: batch-low, preemptible node pools, a queueing layer (Kueue or Volcano) for quotas and gang scheduling, and API-server rate limits per tenant. The Job Scheduler design applies directly: Kubernetes Jobs are the execution layer, not the scheduler of record.

Pattern 4: Leader Election via Leases#

Controllers and singleton workers elect a leader with a coordination.k8s.io/Lease object (an etcd-backed lease with renew deadlines). Good for controllers already on the API server; for application-level locks with fencing, see Coordination Strategies: Leases, Leaders & Reconciliation and Distributed Consensus.

Pattern 5: Multi-Cluster Cells#

Diagram: Pattern 5: Multi-Cluster Cells

Each cell is a complete, independent copy of the stack serving a slice of tenants. A bad CRD, CNI upgrade or etcd incident takes out one cell (~33% of tenants here, ~5–10% in a mature fleet), not the company. The management cluster pushes configuration in waves and carries no user traffic, so its failure stops changes, not requests. State lives outside the clusters, which keeps every cluster disposable.

Pattern 6: GitOps as the Source of Desired State#

A controller (Argo CD, Flux) reconciles cluster state from Git. This extends the reconcile loop one level up: Git is desired state, the cluster is actual state, drift is corrected automatically. Rollback is git revert. The Reddit 2023 outage, below, is the cautionary tale of configuration that lived outside any such source.


Scaling#

The Numbers#

The Kubernetes project documents the supported envelope for a single cluster: no more than 110 pods per node, 5,000 nodes, 150,000 total pods and 300,000 total containers (Considerations for large clusters). These are tested limits, not targets.

ResourceDocumented limit or defaultDesign note
Nodes per cluster5,000Most fleets stay well under 1,000 per cluster for blast radius and upgrade time
Pods per node110 (kubelet maxPods default)Also bounded by pod CIDR size and per-node IP limits on cloud CNIs
Total pods per cluster150,000API server memory and watch fan-out grow with object count
Total containers300,000Sidecars count; a mesh can double container count
etcd quota2 GiB default, 8 GiB suggested maxEvents, oversized CRDs and Secrets are the usual culprits
etcd request size1.5 MiBBounds a single object; huge ConfigMaps fail here
Endpoints per EndpointSlice100 default, 1,000 maxA 5,000-pod Service is ~50 slices
HPA loop15s sync, 300s scale-down window, 10% toleranceScale-up is immediate by default
Pod eviction after node loss300s default tolerationPlus node-monitor grace period: ~5–6 min total
Graceful termination30s default terminationGracePeriodSecondsMust exceed preStop + drain time
Minor releases~3 per year, ~14 months support each3 minor branches maintained at a time

What Breaks First as a Cluster Grows#

~100 nodes:    nothing interesting; defaults are fine
~500 nodes:    API server LIST calls from badly written controllers spike memory;
               DaemonSets x nodes = thousands of pods just for agents
~1,000 nodes:  etcd DB size and write latency matter; split Events into their own etcd;
               iptables-mode kube-proxy rule updates slow with 10K+ Services/endpoints
~2,500 nodes:  scheduler throughput, CNI IP allocation, Prometheus scrape volume;
               every node-level agent is now a load test against the API server
~5,000 nodes:  documented ceiling; you are tuning API server flow control,
               watch cache sizes and etcd hardware for a living

The documented guidance for large clusters includes storing Event objects in a separate etcd instance, running control-plane replicas in every failure zone and giving add-ons appropriate resource limits.

Scaling Moves in Order#

  1. Right-size requests (VPA recommendations, quarterly review). Often recovers 30–50% of nodes before you add any.
  2. Separate node pools by workload shape (system, general, memory-heavy, GPU, batch) so packing works.
  3. Tune the noisy clients: controllers that LIST instead of WATCH, agents that poll the API server, CI systems creating thousands of short-lived pods.
  4. Split etcd for Events; use API Priority and Fairness to stop one tenant's controller from starving the rest.
  5. Add clusters, not nodes, once a cluster is past a few hundred nodes or hosts more than one critical domain. This is where multi-cluster tooling, fleet GitOps and global load balancing become necessary.

🎯 Staff Move: "I'd cap clusters at a few hundred nodes and scale out by adding clusters. 5,000 nodes is what the project tests, not what I want to upgrade on a Tuesday. A cluster is my unit of blast radius and my unit of upgrade."


Failure Modes & Recovery#

1. Liveness-Probe Restart Cascade#

  • Symptom: A dependency slows down; within a minute the restart count of every pod in the service jumps; error rate goes from 2% to 100%.
  • Root cause: Liveness probe checks the dependency, or a 1-second probe timeout fails under load-induced latency.
  • Detection: kube_pod_container_status_restarts_total rate across a Deployment; restarts correlated across many pods within one probe period; CrashLoopBackOff count.
  • Fix: Patch the probe (remove dependency, raise timeoutSeconds and failureThreshold) and roll out; temporarily remove the liveness probe if needed.
  • Prevention: Platform probe standard enforced at admission; chaos test: slow the database by 5s and verify zero restarts.

2. etcd Quota Exhausted (NOSPACE)#

  • Symptom: kubectl apply fails with database space exceeded; no pods can be created, deleted or rescheduled; running pods keep serving.
  • Root cause: Event storm, a controller writing a large CRD status every few seconds, history not compacted, or application data in CRDs.
  • Detection: etcd_mvcc_db_total_size_in_bytes over 80% of quota; etcd_server_quota_backend_bytes; object counts per resource type.
  • Fix: Delete the offending objects, compact, defragment one member at a time, disarm the alarm (see etcd & ZooKeeper).
  • Prevention: Separate Events etcd, per-namespace object-count quotas, review of every CRD for size and write rate.

3. Node Pressure Evictions and OOM Kills#

  • Symptom: Pods restart with OOMKilled or are Evicted; latency spikes on unrelated services on the same nodes.
  • Root cause: Memory overcommit (limits far above requests), BestEffort pods, or a memory leak in a neighbour.
  • Detection: kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}; node MemoryPressure condition; eviction events.
  • Fix: Raise requests to match working set; cordon and drain the worst nodes.
  • Prevention: Memory limit = request; admission rejects pods without requests; VPA recommendations in review.

4. Upgrade Breaks the Data Plane#

  • Symptom: Minutes into a control-plane or node upgrade, pod networking, DNS or ingress fails across the cluster.
  • Root cause: A removed API or label, an incompatible CNI, CSI or webhook version, or configuration that lived outside version control.
  • Detection: Synthetic probes per cluster (DNS, pod-to-pod, ingress); apiserver_request_total 5xx; CNI agent readiness.
  • Fix: Roll back the node pool or restore traffic to another cluster; restoring control plane from an etcd snapshot as the last resort.
  • Prevention: Canary cluster that mirrors production add-ons; deprecated-API scanning; every add-on config in Git.

5. Control-Plane Overload From a Misbehaving Client#

  • Symptom: API server latency climbs to seconds; kubectl times out; controllers fall behind; deploys stall.
  • Root cause: A controller or agent doing full LISTs on every resync across 1,000 nodes, a CI job creating 10,000 pods, or a watch on a huge resource.
  • Detection: apiserver_request_duration_seconds p99 by verb and client; apiserver_flowcontrol_rejected_requests_total; API server memory.
  • Fix: Identify the user agent and throttle or scale it down; raise API server replicas.
  • Prevention: API Priority and Fairness with per-tenant flow schemas; informer-based clients only; load-test add-ons at target node count.
Diagram: 5. Control-Plane Overload From a Misbehaving Client

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Liveness cascadeCorrelated restarts across a DeploymentOne service, then its callersPatch probe, roll outService team; platform owns the standard
etcd NOSPACEDB size > 80% quotaEvery change in the clusterCompact, defrag, delete offendersPlatform; offending team
OOM / evictionOOMKilled reason, MemoryPressurePods on affected nodesRaise requests, drain nodesService team
Upgrade breakPer-cluster synthetic probesWhole clusterShift traffic to other cells, roll backPlatform
API server overloadRequest latency by clientAll controllers and deploysThrottle client, APFPlatform; client's owner
Zone lossNode NotReady by zoneOne third of capacityPre-paid 1.5× capacity, spreadPlatform + capacity owner
Webhook outageAdmission latency/errorsAll pod creationFail-open non-security hooks, exempt system namespacesWebhook's owning team

When to Use vs. Alternatives#

NeedPickWhy
20+ services, frequent deploys, a platform teamManaged KubernetesUniform deploy, scale and recovery; huge ecosystem
A handful of stateless services, small teamServerless containers (Cloud Run, Fargate, App Runner)No nodes, no upgrades, scale to zero
Event-driven functions, spiky and short-livedFunctions (Lambda, Cloud Functions)Pay per invocation
Primary databaseManaged database outside the clusterDurability and failover are the vendor's job
On-prem or regulated, deep in-house expertiseSelf-managed KubernetesControl and portability, at real headcount cost
Few large, long-lived VM-shaped processesVMs + autoscaling groupsLess machinery for no lost capability
Large batch and ML fleetsKubernetes + a batch queue (Kueue/Volcano) or a dedicated schedulerBin-packing and preemption across teams

When NOT to Use Kubernetes#

  • A small team with simple workloads. Five engineers and four services get the cost of a platform (upgrades, node pools, ingress, observability add-ons) without the benefit of standardisation across many teams. Serverless containers or a PaaS deliver deploy, scale and rollback for zero platform headcount.
  • No one owns the platform. A cluster is a product with a ~14-month version support clock. Without at least one engineer's sustained attention, it rots into an unpatched liability.
  • The workload is the database. Running your primary datastore on Kubernetes is possible; doing it as your first Kubernetes workload is how teams learn PersistentVolume zone affinity during an outage.
  • Hard multi-tenancy with hostile tenants. Namespaces are a soft boundary. Untrusted code needs sandboxed runtimes or separate clusters; see Online Code Runner.
  • Latency-critical, hardware-tuned systems (exchange matching engines, kernel-bypass networking). The abstraction costs more than it saves.
Diagram: When NOT to Use Kubernetes

Operational Concerns#

What the On-Call Actually Does#

  1. Watches per-cluster synthetic checks (DNS resolution, pod-to-pod, ingress, kubectl latency). They catch data-plane breakage before service dashboards do.
  2. Watches control-plane health: API server p99 latency and 5xx by client, etcd DB size against quota, etcd leader changes and fsync latency.
  3. Runs upgrades as a routine: control plane first, then node pools by surge, canary cluster a week ahead of production. A 300-node pool rolled at 10% surge with PDBs takes hours; that's normal.
  4. Drains nodes safely, and when a drain hangs, finds the PDB that allows zero disruptions and the team that owns it.
  5. Triages "my pod is Pending": insufficient requested resources, unsatisfiable affinity or spread, taints, PVC zone mismatch, quota exhausted. kubectl describe pod events answer 90% of these.
  6. Owns the admission policy set (required requests, probe standards, image provenance, no :latest, PDB sanity) and the exceptions process.

Key Metrics & Alerts#

MetricHealthyAlert
apiserver_request_duration_seconds p99 (non-LIST)< 1s> 1s for 10m
etcd_mvcc_db_total_size_in_bytes / quota< 50%> 80% (page)
etcd_server_leader_changes_seen_total rate~0> 3 per hour
kube_node_status_condition{condition="Ready",status="true"}All nodes> 5% of nodes NotReady
Pods Pending > 5 minutes0> 0 for critical namespaces
kube_pod_container_status_restarts_total rate per Deployment~0Correlated spike across > 20% of pods
CFS throttled periods ratio< 5%> 10% on latency tiers (ticket)
Cluster allocation ratio (requests / allocatable)70–85%> 90% (no headroom) or < 50% (waste)
Days until the running minor version leaves support> 120< 90 (ticket), < 30 (escalate)

Upgrades as an Ongoing Cost#

Per cluster, per minor upgrade (a common production pattern):
  prep:       API deprecation scan, add-on compatibility matrix     ~2-3 engineer-days
  canary:     upgrade canary cluster, soak 1 week                   ~1 day
  rollout:    control plane, then node pools by surge with PDBs     ~0.5-1 day per cluster
  ~3 minors/year x N clusters -> at 30 clusters, upgrades alone are ~1-2 FTE

The cost scales with the number of clusters and the number of add-ons per cluster. Cell designs multiply clusters, so they only work with fleet automation and uniform add-on sets.


Interview Application — Staff-Level Plays#

Which Case Studies Use Kubernetes#

Case StudyHow Kubernetes Is UsedKey Pattern
Service RegistryOrchestrator as registry; readiness drives EndpointSlicesPlatform registers from orchestrator truth; data plane fails static
Auto-Scaling & CapacityHPA and cluster autoscaler as the reactive loop15s HPA loop, 300s scale-down window, placeholder headroom
Consensus Serviceetcd under the control plane; Lease objects for leader electionConsensus for metadata, never per request
Distributed Job SchedulerJobs/CronJobs as the execution layerScheduler of record outside; Kubernetes runs the work
Online Code RunnerSandboxed pods for untrusted submissionsHard isolation needs sandboxed runtimes, not namespaces
API GatewayIngress and Gateway API in front of servicesEdge routing as a platform-owned component
Load BalancingService VIPs, kube-proxy, cloud load balancersL4 per-connection vs L7 per-request balancing
Metrics & MonitoringPer-pod scraping, cardinality from pod labelsPod churn multiplies time-series count
Multi-Region Active-ActiveOne cluster set per region, global traffic steeringClusters are regional and disposable; state replicates separately
Video StreamingTranscoding workers as batch Jobs on preemptible poolsBatch priority classes, queue-driven scaling
Deployment SystemRolling updates, readiness gates and the cluster as a deploy targetWaves and bake times live above the cluster, not inside a single Deployment
Multiplayer Game BackendDedicated game servers as pods via Agones fleetsAllocated servers are never drained mid-match; scale on a warm buffer

Every System Design Question Has a Kubernetes Moment#

  • URL shortener: "The redirect service is a stateless Deployment with an HPA on in-flight requests and a readiness probe that waits for the hot-key cache to warm. The database is outside the cluster."
  • Chat: "WebSocket gateways are long-lived connections, so a rolling deploy must drain them: preStop marks the pod unready, the gateway tells clients to reconnect with jitter, and terminationGracePeriodSeconds is raised to 5 minutes for that one Deployment."
  • Job scheduler: "Kubernetes Jobs execute the work; the scheduler of record with exactly-once dispatch lives in a database. I won't use CronJobs for anything that must run exactly once."
  • Auth service: "Token validation pods are Guaranteed QoS in a dedicated pool with a PDB of maxUnavailable 1, because if auth is evicted under node pressure, every product is down. See Auth & Identity."

What Interviewers Probe#

After You Say...They Will Ask...What They're Evaluating
"Kubernetes restarts failed pods""How long until a dead node's pods run elsewhere?"Knowing the ~5-minute eviction default; availability from replicas
"Liveness and readiness probes""What does each check? What if the database is slow?"The cascade failure mode
"HPA on CPU""What if the pods have no CPU request? How fast does it react?"Requests-based utilisation; 15s loop; node provisioning delay
"Set resource limits""What does a CPU limit do to p99?"CFS throttling
"One cluster for everything""What's the blast radius of a bad upgrade?"Cells and multi-cluster
"Store the job state in a CRD""What's etcd's size limit?"Kubernetes is not a database

Common Interview Mistakes#

What Candidates SayWhat Interviewers HearWhat Staff Engineers Say
"Kubernetes gives us high availability"Confuses rescheduling with redundancy"Replicas across 3 zones give availability; rescheduling restores capacity in minutes."
"Liveness probe hits /health"Fleet restart during a DB blip"Liveness is in-process only; readiness may reflect dependencies."
"Set limits on everything"Will discover CPU throttling at peak"Requests from load tests; memory limit = request; no CPU limit on latency tiers."
"Autoscaling handles spikes"Ignores node boot time"HPA in 15–30s, nodes in minutes, so I hold 10% placeholder headroom."
"One big cluster is simpler"Company-wide blast radius"Clusters are cells; a few hundred nodes each, canary cluster first."
"Kubernetes for our 3 services"Platform cost without platform benefit"At this size, serverless containers; Kubernetes when we have 20+ services and an owner."

L5 vs L6 vs L7 Responses#

ScenarioL5 AnswerL6 / Staff AnswerL7 / Principal Answer
"Deploy 30 microservices"Kubernetes cluster, Deployments, Services, HPAPaved-road template: requests from load tests, probe standard, PDB, zone spread, HPA on demand signal, GitOpsDecides cluster topology (cells by tier), what the platform guarantees, and the upgrade budget in headcount
"Pods restart in a loop during an outage"Increase probe timeoutSeparate liveness from dependencies, startup probe for boot, add chaos testWrites the probe standard, enforces it at admission, audits the fleet and tracks probe-caused incidents to zero
"Cluster costs doubled"Smaller nodesRight-size requests via VPA, bin-pack by pool, spot for batch, measure request efficiencyChargeback on requested cores per team; capacity buffer priced and owned; retire under-used clusters
"Should the startup adopt Kubernetes?"Yes, it's the industry standardNot yet: 4 services fit serverless containers; revisit at ~20 servicesDefines the trigger (service count, team count, compliance) and the migration path so the decision is re-made deliberately

The Staff Kubernetes Checklist#

  1. Workload shape: "Stateless Deployment, StatefulSet behind an operator, or Job? And does this belong on Kubernetes at all?"
  2. Requests with math: "p95 CPU 0.7 cores at 400 RPS, so 700m; memory limit equals request at 1.1 GiB; no CPU limit."
  3. Probes: "Startup for boot, readiness for traffic, liveness in-process only with a generous timeout."
  4. Scaling chain: "HPA on in-flight requests, 15s loop, 5-minute scale-down window, 10% placeholder headroom for node lag."
  5. Failure domains: "3 zones, maxSkew 1, PDB maxUnavailable 1, 1.5× capacity to lose a zone at peak."
  6. Blast radius and ownership: "Clusters of a few hundred nodes as cells; platform owns control plane and add-ons; service teams own manifests and probes."

🎯 Staff Insight: Don't use Kubernetes as a database (etcd's 2 GiB default quota is for cluster metadata), as a scheduler of record for exactly-once jobs, or as a security boundary between hostile tenants. The strongest Kubernetes signal is saying what happens when the control plane is down: running pods keep serving and only changes stop.

Evaluation Rubric#

DimensionSenior (L5)Staff (L6)Principal (L7)
Workload contractDeployments and ServicesRequests, limits, probes, PDBs, spread set deliberately with numbersOrg standard enforced at admission, with exceptions process
Failure"Pods restart"Node eviction timing, zone loss capacity, control-plane-down behaviourCell boundaries, canary clusters, correlated-failure audit
ScalingHPA on CPUDemand signal, three-loop latency, headroomFleet capacity buffer and chargeback
Operations"Monitor the cluster"etcd quota, API server latency, upgrade routineUpgrade cadence as funded work; cluster count as a cost driver
AdoptionAlways KubernetesKnows when serverless or VMs fit betterDefines the trigger for adoption and the paved road

Strong hire signals

SignalWhat It Sounds Like
Control-loop thinking"Every manifest is a standing order; what will the controller do with it at 3am?"
Probe discipline"Readiness drains, liveness restarts. Only one of those should ever see the database."
Knows the clocks"15-second HPA loop, 5-minute scale-down, 5-minute eviction, minutes for a node."
Blast-radius sizing"A cluster is a cell; I'd rather lose 10% of tenants than all of them."
Knows when not to"Three services and five engineers: serverless containers."

Lean no-hire signals

SignalWhy It Misses the Bar
Kubernetes as the answer to availabilityRescheduling is minutes; redundancy is the mechanism
No requests or probes discussedWill ship BestEffort pods and restart cascades
Application state in CRDs or etcdWill freeze the cluster at the quota
One cluster for the company with no upgrade planUnpriced, uncontained blast radius

Common false positives

  • Fluent YAML ≠ judgment. Writing a Deployment from memory says nothing about whether the probe will cascade.
  • Operator and CRD enthusiasm ≠ platform design. Ask who maintains the operator in year two.
  • "We ran 2,000 nodes" ≠ knowing when not to. Ask what they'd recommend for four services.

Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

At Staff level Kubernetes is a runtime you configure correctly. At Principal level it is the company's operating model for compute: the platform team's product, the contract between that team and every service team, and a recurring cost that scales with cluster count, add-on count and the ~3 minor releases a year. The L7 question is not "Deployment or StatefulSet?" but "How many clusters should we have, what does the platform promise, what may teams customise, and what do we pay each year just to stay supported?"

🧭 Principal Move: "Before we talk about manifests, I want to agree on the paved road: what every service gets for free (deploy, scale, rollback, observability, a probe standard), what it must declare (requests, probes, owner), and what it may not do (state in etcd, cluster-admin, :latest). The YAML follows from that contract; the contract doesn't follow from the YAML."

The Org-Level Fault Line#

Shared multi-tenant clusters vs cluster-per-team.

OptionWhat WorksWhat BreaksWho Pays
One shared cluster per regionBest bin-packing; one expert team; uniform policyCompany-wide blast radius; every upgrade needs every team's blessing; noisy neighboursEvery team, during the one bad afternoon
Shared clusters by tier or cell (critical, general, batch; or tenant cells)Contained blast radius; tier-specific policies; staged upgradesMore clusters to upgrade; needs fleet automationPlatform headcount
Cluster per teamMaximum autonomy; clean cost attribution40 half-maintained clusters, drifting versions, duplicated add-ons, poor packingSecurity and incident responders; finance through waste
No Kubernetes for some teams (serverless)Zero platform cost for simple servicesTwo operating models; some tooling duplicatedPlatform, to support both paths

The Principal default: a small number of platform-owned clusters per region, split by tier and then by cell as the fleet grows, with namespaces as the team boundary, quotas and chargeback per namespace, and a serverless path for teams whose needs fit it. Cluster-per-team only for hard isolation requirements (regulated data, hostile code), and then from the same automated template.

🧭 Principal Insight: "The question isn't shared versus dedicated. It's who absorbs the upgrade treadmill. Shared clusters concentrate it in one team that gets good at it; cluster-per-team spreads it across 40 teams that each do it badly once a year."

Cost Model#

Assumptions: managed control plane ~$75/month per cluster; general node 16 vCPU / 64 GiB at ~$550/month on demand; observability and add-ons ~10–15% of node spend; loaded engineer $250K/year ($21K/month). Directional only.

ScaleClusters / nodesCompute/monthControl planes + add-ons/monthPlatform headcountTotal/month
Startup (8 services)2 / 12~$6.6K~$1K0.5 FTE (~$10K)~$18K, versus ~$6–10K on serverless containers
Growth (120 services)8 / 250~$140K~$18K3–4 FTE (~$75K)~$230K
Enterprise (1,500 services)60 / 3,500~$1.9M~$250K15–20 FTE (~$370K)~$2.5M

Two things stand out. At startup scale the platform engineer, not the compute, is the dominant cost, which is the quantitative case for not adopting Kubernetes early. At growth and enterprise scale, request efficiency is the lever: moving the fleet from ~35% to ~60% of requested CPU actually used removes roughly 40% of nodes, worth more than any other optimisation on this page.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to Reverse
Storing application state in CRDs/etcdOne-way once services depend on itData migration plus rewriting every controller that reads it
Pod and Service CIDR ranges, CNI with IP modelOne-way for a running clusterNew cluster and migration
Cluster topology (shared vs per-team)One-way-ishQuarters of migration, workload by workload
Custom controllers and operators the business relies onOne-way-ishOwning that code forever, or a migration
Managed provider (GKE vs EKS vs AKS)Two-way, if manifests stay portableWeeks per cluster; cloud-specific add-ons are the drag
Requests, probes, HPA targetsTwo-wayConfig change and rollout
Node instance types and poolsTwo-wayRolling node pool replacement
Service mesh adoptionTwo-way-ishRemoving sidecars is easy; removing dependence on mesh features is not

The Standard I'd Write#

RFC-PLAT-004: Workload Contract for Shared Kubernetes Clusters

Scope: Every workload in a platform-managed cluster, all environments above dev.

MUST
  1. Declare CPU and memory requests on every container; memory limit = request.
  2. Liveness probes check in-process health only; no network calls to
     dependencies. Readiness may reflect dependencies. Startup probe if boot > 10s.
  3. Run >= 2 replicas (>= 3 for critical tier) with a PodDisruptionBudget that
     allows at least 1 disruption, and zone topology spread with maxSkew 1.
  4. Use immutable image tags or digests; no :latest.
  5. Store no business data in custom resources or ConfigMaps.
  6. Declare an owning team label that routes pages.
SHOULD
  7. Omit CPU limits on latency-critical services; set them on batch.
  8. Scale on a demand signal (in-flight requests, queue lag) rather than CPU.
  9. Handle SIGTERM: fail readiness, wait >= 5s, drain, exit before grace period.

Exceptions: Filed with the platform team; approved by platform lead and the
owning team's director; time-boxed to one quarter; listed on the exceptions dashboard.

Enforcement: Admission policy in warn mode for one quarter, then enforce.
             Shadow -> warn -> canary cluster -> enforce everywhere.

Success metrics: zero probe-caused incidents per quarter; fleet request
efficiency >= 60%; every cluster within the supported version window;
median new-service time to production < 1 day.

What I'd Tell the VP#

"Kubernetes is how we run nearly all our services now, which means our clusters are a product with a yearly maintenance bill, not a one-time project. Today we have one large cluster per region, so a single bad upgrade can stop every team from deploying, and we use about a third of the computing capacity we pay for. I'm proposing three changes: split clusters so a failure hits a slice of customers rather than all of them, enforce a simple contract for every service so one bad health check can't restart a whole product, and charge teams for the capacity they reserve. That should cut our compute bill by roughly a third and turn platform outages into partial incidents. The cost is two more platform engineers and one quarter of migration work."

Principal Interview Signals#

SignalWhat It Sounds Like
Treats the platform as a product"The paved road is a contract: what we guarantee, what teams declare, what they may not do."
Prices the upgrade treadmill"Three minors a year times 30 clusters is a standing team, so cluster count is a cost decision."
Sizes blast radius deliberately"Cells of a few hundred nodes; a bad rollout hits 10% of tenants and the canary cluster first."
Prices requests, not usage"We pay for what's requested. Chargeback on requests fixes right-sizing faster than any tool."
Knows when not to standardise"Teams with three stateless services get serverless containers; forcing them onto clusters costs more than it saves."

Staff answers that L7 interviewers find insufficient:

  • "We'll run multiple clusters for isolation" without the upgrade cost, fleet tooling or how traffic shifts between them.
  • "Enforce requests and probes" as advice, rather than an admission policy with a rollout path and an exceptions process.
  • "Use managed Kubernetes to reduce ops" without saying which ops remain: node pools, add-ons, upgrades, policies and cost.

How Real Companies Built It#

Google Borg — The Ancestor#

Kubernetes descends from Borg, the cluster manager Google described in its EuroSys 2015 paper. Borg runs hundreds of thousands of jobs across clusters of up to tens of thousands of machines, and reaches high utilisation by combining admission control, efficient task packing, over-commitment and machine sharing with process-level performance isolation. The paper's lessons section discusses how Kubernetes revisited several Borg design choices, such as organising pods with labels rather than a fixed job-to-task grouping (Borg paper, EuroSys 2015).

Staff insight: Kubernetes inherited Borg's core bet: declare intent, let a scheduler bin-pack against requests, accept over-commitment for utilisation. When you discuss requests and QoS, you are discussing Borg's utilisation economics.

Pokémon GO on Google Kubernetes Engine — Launch at 50× Plan#

Niantic launched Pokémon GO in 2016 on Google Kubernetes Engine (then Google Container Engine). Google's account says the team planned for 1× traffic with a worst case of 5×, and actual traffic hit 50×. It describes the deployment as the largest Kubernetes Engine deployment at the time, adding over a thousand nodes, performing a live cluster version upgrade during the launch period, and moving to Google's HTTP/S load balancing. The Japan launch that followed is described as running without incident (Google Cloud blog).

Staff insight: The cluster scaled because nodes could be added and pods rescheduled, but the hard parts were the edges: load balancing and the version upgrade under load. In an interview, capacity buffer, the edge and the upgrade path are where a Kubernetes design gets tested.

OpenAI — 7,500-Node Clusters for ML#

OpenAI described scaling single Kubernetes clusters to 7,500 nodes to run research and training workloads including GPT-3, CLIP and DALL·E. Its workloads often place a single pod on an entire node to use all of its GPUs, and it describes moving off flannel to native pod networking for throughput, running API servers and etcd on dedicated nodes, and closely monitoring API server error rates (OpenAI engineering).

Staff insight: At the top of the envelope, the limits are the control plane and the network, not the pods. A batch or ML design on Kubernetes should name API server load, etcd placement and the CNI before it names the training framework.

Reddit — The Pi-Day Upgrade Outage#

On March 14, 2023, Reddit suffered an outage of roughly 314 minutes that began during an upgrade of a cluster from Kubernetes 1.23 to 1.24. Calico route reflectors were selected by the master node label, which 1.24 no longer applies in favour of control-plane, so no route reflectors ran and pod networking failed. The configuration had been set up years earlier outside version control (Reddit engineering post-mortem).

Staff insight: The upgrade itself was routine; the failure came from configuration no one knew existed. That is the case for a canary cluster with the same add-ons as production, every add-on's config in Git, and treating upgrades as a staged rollout instead of a maintenance task.


Practice Drill#

Prompt: "You run the platform for a 150-engineer company: 90 services on one managed Kubernetes cluster per region, 400 nodes each. Last month a database brownout turned into a 40-minute full outage, and the cluster is two minor versions behind. Redesign the platform."

Staff Answer

Start with the outage, because it is the cheaper fix. A 20-second database slowdown becoming 40 minutes of downtime is the signature of liveness probes that check dependencies: every pod failed liveness within ~30s, restarted, and reconnected with cold pools. I'd ship a probe standard immediately: liveness in-process only with timeoutSeconds: 5 and failureThreshold: 6; readiness may reflect the database; startup probes for the JVM services with 40–60s boot. Enforce it through an admission policy in warn mode for two weeks, then enforce, and add a chaos test that slows the database by 5s in staging and asserts zero restarts. Second, the workload contract: requests on every container from VPA recommendations, memory limit = request, no CPU limits on API tiers, PDBs that allow at least one disruption (several services currently have maxUnavailable: 0, which is also why upgrades are stuck), and zone spread with maxSkew: 1. Third, the upgrade debt: build a canary cluster with the same add-ons, scan for deprecated APIs, and upgrade one minor at a time each quarter until current, then keep the ~3-per-year cadence. Fourth, blast radius: split each region into a critical cluster (auth, checkout, payments; Guaranteed QoS, strict change control) and a general cluster, so a bad add-on upgrade can't take both. State stays in managed databases outside the clusters. Metrics: correlated restarts per Deployment, etcd size vs quota, API server p99, days to end of support, request efficiency.

Why this is L6:

  • Diagnoses the cascade from the timeline (DB blip → fleet restart) rather than "add capacity".
  • Connects the stuck upgrades to the PDB misconfiguration, one root cause with two symptoms.
  • Rolls policy out shadow → warn → enforce with a chaos test proving the fix.
  • Splits clusters by criticality to contain the next incident.

What L7 adds:

  • Frames the platform as a product with a contract, an exceptions process and success metrics (zero probe-caused incidents, all clusters in support).
  • Prices it: chargeback on requested cores, expected 25–35% node reduction from right-sizing, against two extra platform engineers.
  • Sets the trigger for moving from tiered clusters to cells, and offers serverless containers for the long tail of small internal services.

Quick Reference Card#

Model:             control loop over desired state; every object is a standing order
Source of truth:   etcd via API server only; quota 2 GiB default, 8 GiB suggested max,
                   1.5 MiB max request; Events in a separate etcd for large clusters
Control plane down: running pods keep serving; deploys, scaling, rescheduling stop
Cluster limits:    110 pods/node, 5,000 nodes, 150,000 pods, 300,000 containers
Scheduler:         places by REQUESTS (never usage); filter -> score -> bind
Limits:            CPU over limit = throttled (CFS, 100ms periods)
                   memory over limit = OOM kill
QoS:               Guaranteed > Burstable > BestEffort (evicted first)
Probe defaults:    period 10s, timeout 1s, failureThreshold 3, success 1, delay 0
Probe rule:        readiness drains traffic, liveness restarts; liveness never checks deps
Termination:       terminationGracePeriodSeconds 30s default; preStop sleep before drain
Node loss:         NotReady after grace period, eviction after 300s toleration (~5-6 min)
HPA:               15s sync, 10% tolerance, scale-up immediate, 300s scale-down window;
                   CPU utilisation is % of request (no request = no CPU scaling)
EndpointSlice:     100 endpoints per slice default, 1,000 max; only Ready pods
Releases:          ~3 minors/year, ~14 months support each, 3 branches maintained
Zone loss math:    3 zones -> 1.5x peak capacity to survive one at peak

RED FLAGS
  - Liveness probe that calls the database or a downstream service
  - Containers without requests (BestEffort in production)
  - CPU limits on multi-threaded latency-critical services
  - One cluster as the whole company's blast radius
  - Business data in CRDs or ConfigMaps
  - PDB that allows zero disruptions (upgrades blocked)
  - Fail-closed webhooks with no namespace exemptions
  - Cluster more than one minor version behind with no upgrade owner
  1. Loading the index…