Technologies referenced in this case study: API Gateways · ZooKeeper & etcd · Redis
Related: API Gateway · Service Discovery · CDN & Edge Caching · Circuit Breakers · Rate Limiting · Multi-Region · Real-time WebSockets · Degraded Mode · Consistent Hashing · Networking · Numbers to Know
How to Use This Case Study#
This case study is organized for the interview first and for reference second. Read it front-to-back once; then return to the fault lines and deep dives that match your weak spots. If you need the packet-level background — TCP, TLS, HTTP/2, anycast — start with Networking.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Deep Dives 1 and 4 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal Lens) and Appendix A (balancing algorithms in depth) |
What is a Load Balancer? — Why interviewers pick this topic
A load balancer accepts traffic addressed to one name or IP and spreads it across many backends, so that the service survives the loss of any backend and its capacity can grow by adding more. In practice it does more: terminates TLS, routes by host and path, retries failed requests, drains hosts during deploys, and decides — continuously, without a human — which backends are healthy enough to receive traffic.
That last job is the dangerous one. A load balancer is the component in the request path with the most authority and the least context. It can't see that a backend's database is slow; it sees that requests are slow. It can't see that a "healthy" backend returns errors in 2 ms; it sees a backend that finishes requests very quickly. Every load-balancing decision is made on proxies for health, and the proxies lie in predictable ways.
Before vs After — the "one bad host" scenario:
Without a designed balancing and health policy:
t=0: Backend b-17 loses its database connection pool. Every request now
returns 503 in 2 ms instead of 200 in 40 ms.
t=+1s: Least-connections LB sees b-17 always has 0 active requests.
Routes more and more traffic to it.
t=+10s: b-17 receives ~35% of all requests across a 40-host pool. 35% error rate.
t=+30s: Active health check (GET /health every 10s, needs 3 failures) still
passes: /health doesn't touch the database.
t=+8min: Someone notices the error graph, finds b-17, terminates it by hand.
With passive outlier detection and P2C least-request:
t=0: b-17 starts returning 503 in 2 ms.
t=+0.5s: 5 consecutive 5xx → b-17 ejected for 30s (cap: 10% of pool ejectable).
t=+0.5s: Error rate peak: ~0.6% for half a second. Traffic spread over 39 hosts.
t=+30s: b-17 re-admitted on probation, fails again, ejected for 60s.
t=+2min: Alert: host b-17 ejected 3 times → owner's auto-remediation replaces it.
Why interviewers reach for this question: Load balancing looks like a solved problem — "put an ALB in front" — which makes it an excellent test of depth. The candidate who stops at round-robin and health checks has described a working system on a good day. The interviewer wants to see whether you know how load balancers amplify failures: herding traffic onto broken hosts, ejecting healthy hosts during overload, breaking every connection on a config push, or turning a TLS certificate rotation into a CPU outage.
Mechanics Refresher: Layers and Algorithms
| Mechanism | How It Works | Pros | Cons |
|---|---|---|---|
| DNS round-robin / GSLB | Return different IPs per resolver; weight by region or health | Global steering; no data-path box | TTLs ignored by many clients; minutes to drain; no per-request control |
| Anycast | Same IP announced from many sites; BGP picks the nearest | Instant global spread; absorbs DDoS | Route changes can move flows mid-connection; coarse control |
| ECMP at the router | Router hashes 5-tuple across equal-cost next hops | Line-rate, no extra box | Next-hop set change rehashes flows unless backed by consistent hashing |
| L4 load balancer | Forwards packets/connections by 5-tuple; no payload inspection | Millions of packets/s per machine; protocol-agnostic | No HTTP awareness: can't route by path, retry requests or see 5xx |
| L7 proxy | Terminates the client connection, parses HTTP, opens/reuses backend connections | Per-request routing, retries, outlier detection, observability | TLS and parsing CPU; adds ~0.1–1 ms; must scale with requests, not packets |
| Client-side / sidecar balancing | Caller (or its local proxy) picks the backend from a discovered list | No central hop; per-request decisions with local latency data | Every client needs the logic; control plane fan-out to thousands of clients |
| Algorithm | How It Works | Pros | Cons |
|---|---|---|---|
| Round robin | Next backend in order | Simple, even on uniform requests | Ignores backend load and request cost variance |
| Least connections | Backend with fewest open connections | Adapts to slow backends | With HTTP/2 or keepalive, connections ≠ load; herds onto fast-failing hosts |
| Power of two choices (P2C) + least request | Pick 2 backends at random, send to the one with fewer in-flight requests | Near-optimal spread with O(1) work and stale data; avoids herding | Still fooled by fast-failing hosts without outlier detection |
| Consistent hashing (ring / Maglev / rendezvous) | Hash a key to a backend; minimal remapping on change | Affinity (caches, sessions); stable L4 flows | Uneven load on hot keys; needs bounded-load variants |
| Weighted / latency-aware (EWMA) | Score backends by latency × in-flight | Routes around slow hosts | Feedback oscillation if not damped |
For most production systems: anycast or DNS to a region, ECMP to an L4 tier using consistent hashing, an L7 proxy fleet terminating TLS, and P2C least-request with passive outlier detection toward backends. The algorithm is not the interview — health judgment and failure amplification are.
Executive Summary
If you only read one section, read this. Everything in this case study flows from the contrast below.
What This Interview Actually Tests#
A load balancer is not an algorithm question. Round-robin works fine on most days.
It is a failure-judgment question: a load balancer decides, many times per second and without a human, which machines deserve traffic — and the most common way it fails is by being confidently wrong at scale. It tests:
- Whether you know what each layer (DNS, anycast, L4, L7, client-side) can and can't see, and put each decision at the layer that has the information
- Whether you know where the latency actually goes — the proxy hop is sub-millisecond; the TLS handshake and connection setup are not
- Whether your health model distinguishes "dead" from "slow" from "fast-failing", and refuses to eject half the fleet during overload
- Whether you treat the load balancer's own config and control plane as the largest blast radius in the system
- Whether connection lifetime — keepalive, HTTP/2, WebSockets — is part of your balancing model
The key insight: The load balancer sits on every request, so its mistakes are correlated by construction. A backend bug hurts one host's share of traffic; a load-balancer bug — a bad health check, a herding algorithm, a config push — hurts all of it at once. Staff candidates design the LB to fail less confidently: bounded ejection, panic thresholds, staged config rollout, and draining that respects connection lifetime.
The L5 vs L6 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Draws clients → LB → backends, picks round-robin or least-connections | Asks "Internet edge or service-to-service? HTTP or raw TCP? Long-lived connections? What's the latency budget?" | Asks "How many balancing layers does the org already run, who owns each, and which one is the single config push that can take everything down?" |
| Latency | "LBs add latency, so we keep it to one hop" | "The L7 hop is ~0.1–1 ms. The cost is TLS: a full handshake is 1–2 RTTs plus CPU, so I'd maximize resumption and connection reuse" | Prices handshake CPU across the fleet — ECDSA certs and resumption can cut L7 fleet size by 30–50% — and sets the TLS policy org-wide |
| Algorithm | Least-connections | P2C least-request with outlier detection and slow start; consistent hashing only where affinity pays | Standardizes the algorithm in the shared proxy config so 200 teams don't each rediscover the fast-fail sinkhole |
| Health | Active /health check every 10s | Active checks for liveness + passive outlier detection for correctness; max 10–20% ejectable; panic mode below 50% healthy | Defines what /health may and may not check org-wide — shallow liveness only — so a shared dependency blip can't fail every service's checks at once |
| Failure | "Run two LBs, active-passive" | "N+1 stateless proxies behind ECMP; consistent hashing so a proxy loss moves only its flows; config changes staged by canary" | Designs the config-push failure posture: cell-by-cell rollout, automatic rollback on error-rate delta, a last-known-good config that can be restored in < 5 min globally |
| Ownership | Infra team owns the LB | Edge team owns the L4/L7 fleet and its config pipeline; service teams own their health endpoints, timeouts and retry budgets | Draws the contract: platform owns the data plane and policy guardrails, services own their routes — with validation that rejects unsafe routes before they ship |
Why "algorithm" separates levels
L5: "Least-connections, because it adapts to slow backends." This is a good instinct and correct in a single-LB, HTTP/1.1 world. It breaks in two common production situations. First, with many LB instances each seeing only its own connections, a newly added backend looks empty to all of them simultaneously, and they all herd onto it. Second, a backend that fails fast — returning 503 in 2 ms — always has the fewest active requests, so least-connections sends it the most traffic.
L6: "P2C least-request: each proxy picks two backends at random and sends to the one with fewer in-flight requests. Randomness breaks the herd across many proxies, and it needs no global state. On its own it still rewards fast failure, so I pair it with passive outlier detection — five consecutive 5xx ejects the host for 30 seconds — and slow start, so a new host ramps over 30–60 seconds instead of taking a full share cold."
L7: "The algorithm lives in one shared proxy config template, not in 200 service configs. Teams can choose consistent hashing when they need affinity, through a reviewed option — but the default is P2C, outlier detection on, max ejection 10%. Most balancing incidents I've seen were teams who overrode a good default without knowing why it existed."
Why "health" separates levels
L5: "Health check /health every 10 seconds; remove a host after 3 failures." The check usually tests that the process is up — and sometimes, overzealously, tests its database too. Neither tells you whether real requests are succeeding.
L6: "Two signals with two jobs. The active check is shallow — is the process alive and accepting connections — and is how new hosts join. Passive outlier detection watches real traffic: consecutive 5xx, error rate relative to peers, latency outliers. And there's a safety rail: if more than half the pool looks unhealthy, the LB stops trusting its health data and balances across everything, because it's far more likely the checks are wrong than that 50% of hosts died at once."
L7: "The worst health-check outage is the correlated one: every service's /health calls the same auth service, auth has a 20-second blip, and every LB in the company ejects every backend. I'd write the rule that health endpoints MUST NOT call shared dependencies, and enforce it with a lint on the health handler and a quarterly game day that kills a shared dependency."
Why "failure" separates levels
L5: "Two load balancers in active-passive with a floating IP." That protects against one machine dying. It doesn't address the more common outage: the LB fleet is fine, but a configuration change breaks routing on every instance simultaneously.
L6: "The data plane is a fleet of stateless proxies behind ECMP; losing one proxy moves ~1/N of flows, and consistent hashing at L4 means only those flows move. The bigger risk is config: route tables, TLS certs, health policies. Every config change goes canary → one AZ → one region → global, with automatic rollback if the 5xx rate rises more than 0.5 percentage points over baseline."
L7: "The LB config pipeline is the highest-blast-radius deploy system in the company. I'd give it the strictest change management we have: staged rollout enforced by tooling, not process; a global kill-switch to last-known-good; and edge config changes frozen during peak events with a VP-level exception path."
The Staff Positions#
| Position | Rationale |
|---|---|
| Two tiers: L4 (consistent hashing) in front of L7 (stateless proxies) | L4 scales packets cheaply and keeps flows stable; L7 does per-request intelligence. Neither can do the other's job well |
| P2C least-request as the default algorithm | Near-optimal spread, O(1), robust to many proxies with partial views; least-connections herds |
| Passive outlier detection with a max-ejection cap (10–20%) | Real traffic is the only honest health signal; the cap stops the LB from ejecting its way into an outage |
| Panic mode below ~50% healthy | When most hosts look unhealthy, the checks are wrong more often than the hosts are |
| Shallow active health checks | Liveness only; deep checks turn every shared-dependency blip into a fleet-wide ejection |
| TLS terminated at the edge with resumption and ECDSA certs | Full handshakes dominate proxy CPU and user latency; resumption makes repeat visits ~1 RTT cheaper |
| Config changes are deploys: staged, canaried, auto-rolled-back | LB config is the single change that can break every request at once |
| Retries at one layer only, with a retry budget (~10–20% of requests) | Retries at every layer multiply load 3–5× during an incident |
The Three Intents#
Three intents produce three different load balancers. Name them, then commit.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Internet edge (north-south) for a latency-sensitive HTTP service | p99 added latency < 5 ms; TLS for millions of clients; DDoS absorption | Anycast/DNS → L4 (ECMP + consistent hash) → L7 proxies terminating TLS → backends with P2C | TLS CPU exhaustion; config push outage; edge capacity loss in one region | Availability 99.99%; added p50 < 1 ms beyond TLS; no connection resets on deploys |
| Service-to-service (east-west) inside the data center | Sub-millisecond overhead; thousands of services; per-request retries and timeouts | Client-side balancing via sidecar or library, fed by service discovery | Control-plane fan-out; inconsistent client versions; retry storms | Zero extra hops; endpoint updates propagated < 5 s |
| L4 network load balancer for non-HTTP or very high packet rates | Millions of packets/s; preserve client IP; any TCP/UDP protocol | Kernel-bypass or XDP forwarding, consistent hashing, connection tracking, direct server return | Rehash on fleet change breaks flows; SYN floods; asymmetric routing | Existing connections survive LB fleet changes; line-rate forwarding |
🎯 Staff Move: "I'll design the internet edge for a latency-sensitive HTTP API — that's where TLS cost, L4/L7 layering and health judgment all matter. I'll sketch the L4 tier but not redesign packet forwarding, and I'll note where east-west balancing differs: it moves into the client, because a central hop for every internal call is a tax we don't need to pay."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Where to Terminate: L4 vs L7 | Per-request intelligence and TLS termination at L7, or line-rate, protocol-agnostic forwarding at L4 — and who pays the CPU? |
| 2 | Even Spread vs Affinity | P2C least-request spreads load best; consistent hashing keeps caches and sessions warm. You can't fully have both |
| 3 | Trusting Health Signals vs Ignoring Them | Eject fast to protect users from bad hosts, or eject cautiously so a bad signal can't empty the pool? |
| 4 | Central Proxy vs Client-Side Balancing | One fleet that's easy to operate, or balancing logic in every client that removes a hop but multiplies the control plane? |
| 5 | Connection Longevity vs Rebalancing | Long-lived connections save handshakes and latency, but pin load to old backends and make draining slow |
In the Wild: Real Production Systems#
Why this section belongs here: Naming how real edge networks settled these fault lines shows you've studied operations, not just cloud console options.
Google — Maglev#
Google's Maglev paper (NSDI 2016) describes a software L4 load balancer running on commodity servers. Routers spread packets across Maglev machines with ECMP; each machine uses Maglev hashing — a consistent hashing scheme built on a fixed-size lookup table populated by per-backend permutations — plus a local connection-tracking table, so packets of a flow reach the same backend even when the Maglev fleet or backend set changes. The fleet is active-active, scaling out instead of using active-passive pairs.
Staff insight: Maglev's design point is minimal disruption under change: when backends or balancers come and go, almost all existing flows keep their mapping. In an interview, "consistent hashing plus connection tracking at L4" answers the question "what happens to in-flight connections when you add a load balancer?"
Meta — Katran#
Meta open-sourced Katran in 2018, an L4 load balancer built on XDP and eBPF that processes packets in the kernel's earliest receive path. It sits behind ECMP, uses a Maglev-style consistent hash with a connection table, and encapsulates packets to backends, which reply directly to clients (direct server return) so return traffic bypasses the balancer.
Staff insight: Direct server return matters because responses are usually much larger than requests. If return traffic doesn't flow through the L4 tier, that tier is sized for inbound packets only — often a 5–10× smaller fleet. Mention it when sizing the L4 tier.
Lyft — Envoy#
Envoy was built at Lyft and open-sourced in 2016 as an L7 proxy used both at the edge and as a sidecar next to every service. Its load balancer offers round robin, least-request using power-of-two-choices by default, ring hash and Maglev for affinity, passive outlier detection, a panic threshold (by default, if fewer than 50% of hosts are healthy, it balances across all hosts), and dynamic configuration over the xDS APIs from a control plane.
Staff insight: Envoy's defaults encode the hard-won lessons in this case study — P2C over least-connections, outlier detection with a capped ejection percentage, panic mode. Citing why those defaults exist is far stronger than citing that the product exists.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Least connections" | "You have 50 LB instances and just added 10 backends. What happens in the next second?" | Herding with partial views; P2C, slow start |
| "Health checks every 10 seconds" | "A backend returns 503 in 2 ms on every request. When does it stop getting traffic?" | Passive outlier detection; fast-fail sinkhole |
| "We eject unhealthy hosts" | "The database is slow, and 70% of hosts now fail checks. What does the LB do?" | Max ejection, panic threshold |
| "We terminate TLS at the LB" | "Traffic just tripled from new users. What saturates first?" | Handshake CPU, resumption, cert type |
| "Active-passive pair" | "The config push had a typo. Both are running it." | Config as the real blast radius |
| "We use WebSockets / gRPC" | "You scaled from 20 to 40 backends. How long until the new ones get load?" | Connection longevity vs rebalancing |
| "We retry failed requests" | "Every layer retries 3 times. One backend tier is down. What's the multiplier?" | Retry budgets, single retry layer |
System Architecture Overview#
Quick-Reference: The 30-Second Cheat Sheet#
| Question | Staff Answer |
|---|---|
| Layers? | GeoDNS/anycast → L4 (ECMP + consistent hash) → L7 proxies → backends |
| What does the L7 hop cost? | ~0.1–1 ms in-region; TLS handshakes are the real cost (1–2 RTT + CPU) |
| Algorithm? | P2C least-request; consistent hashing (bounded load) only where affinity pays |
| Health? | Shallow active check for liveness; passive outlier detection for correctness |
| Over-ejection? | Cap ejection at 10–20% of pool; panic mode below 50% healthy |
| New host? | Slow start over 30–60 s |
| Deploy draining? | Stop new requests, finish in-flight, GOAWAY on HTTP/2, drain window ≈ p99 request time (seconds), longer for WebSockets |
| Retries? | One layer, idempotent requests only, budget ~10–20% of traffic |
| Config? | Validated, canaried per cell, auto-rollback, last-known-good restore < 5 min |
| East-west? | Client-side/sidecar balancing; no central hop |
Key Numbers Worth Memorizing#
| Number | Value | Context |
|---|---|---|
| L7 proxy added latency | ~0.1–1 ms p50 in-region | Parsing + routing + one local hop; tail grows under CPU pressure |
| L4 forwarding added latency | Tens of microseconds | Kernel-bypass/XDP forwarding |
| TLS 1.3 full handshake | 1 RTT before first request byte | TLS 1.2: 2 RTTs; 0-RTT resumption possible for replay-safe requests |
| Handshake CPU (server signature) | RSA-2048: ~1–3K signs/s/core; ECDSA P-256: ~10–40K/s/core | Why ECDSA certs shrink edge fleets |
| Cross-region RTT | ~60–150 ms | Why a user far from the edge pays 100+ ms per handshake RTT |
| Simple HTTP proxying | ~10–30K req/s per core | Varies with TLS, header size, filters |
| Concurrent connections per proxy | 100K+ idle keepalive connections | Memory-bound: ~10–50 KB per connection with TLS buffers |
| Envoy panic threshold default | 50% healthy | Below it, balance across all hosts |
| Outlier ejection defaults (Envoy) | 5 consecutive 5xx, 30 s base ejection, 10% max ejected | Base time multiplies on repeat ejections |
| ALB health check defaults | 30 s interval, 2 failures unhealthy, 5 successes healthy | Up to ~60 s to detect a dead host with active checks alone |
| ALB deregistration delay default | 300 s | Often far too long for short requests, too short for WebSockets |
| DNS TTL for steering | 30–60 s | Many clients and resolvers hold records longer; plan for minutes |
| P2C max load | ~log log n above average vs ~log n / log log n for random | Two choices capture most of the benefit of perfect information |
| Retry budget | ~10–20% of requests | Bounds load amplification during backend incidents |
Interview Walkthrough
The walkthrough below is a 45-minute script. The words in italics are meant to be said out loud.
Phase 1: Requirements & Framing (2–3 minutes)#
"Load balancer can mean three different things: the internet edge in front of a public service, balancing between internal services, or a pure L4 packet balancer. I'll design the internet edge for a latency-sensitive HTTP API, where TLS, L4/L7 layering and health judgment all matter. I'll mention how east-west differs at the end."
| Question | Assumed Answer | Why It Matters |
|---|---|---|
| Peak traffic? | 1M requests/s globally, ~400K/s in the largest region | Sizes the L7 fleet |
| New connections per second? | ~150K/s at peak (mobile clients reconnect often) | TLS handshake CPU is the binding constraint |
| Latency budget for the LB layers? | Added p99 < 5 ms in-region, excluding TLS RTTs | Rules out extra hops; pushes for connection reuse |
| Protocols? | HTTP/1.1, HTTP/2, some WebSockets, gRPC from partners | Long-lived connections change balancing and draining |
| Regions? | 3 regions, 3 AZs each | Global steering and zone-aware routing |
| Backends? | ~60 services behind the edge, 10–400 hosts each | Routing table size, per-service policy |
| Availability target? | 99.99% for the edge | ~4.3 min/month — config mistakes alone can burn it |
| Session affinity? | Not required by most services; one cache-heavy service benefits | Affinity is opt-in per route |
"99.99% gives the edge about four minutes a month. A single bad global config push can eat that in one go, so I'll treat config safety as a first-class part of the design, not an operations footnote."
Phase 2: Core Entities & API (1–2 minutes)#
| Entity | Key Fields | Notes |
|---|---|---|
| Listener | VIP, port, protocol, TLS policy, cert refs | One per public endpoint |
| Route | host, path prefix, headers → cluster, timeout, retry policy | Owned by service teams, validated by the platform |
| Cluster (pool) | name, endpoints, LB policy, health policy, outlier policy, circuit-breaker limits | Policy defaults from the platform template |
| Endpoint | IP:port, zone, weight, health status, drain state | Fed by service discovery |
| Config version | version, diff, rollout stage, owner | Every change is versioned and attributable |
# Control-plane API (service teams own routes; platform owns listeners and templates)
PUT /v1/routes/{service} { host, path_prefix, cluster, timeout_ms, retry: {...} }
PUT /v1/clusters/{service} { lb_policy: "p2c_least_request" | "ring_hash", hash_on?, ... }
POST /v1/endpoints/{service}/drain { endpoint, drain_seconds }
GET /v1/config/versions/{v}/status → { stage: "canary" | "zone" | "region" | "global", health_delta }
POST /v1/config/rollback { to_version } # platform on-call only
"Routes are owned by service teams; listeners, TLS policy and the defaults template are owned by the edge platform. That split matters later: it's what lets a team ship a route change without being able to change the timeout behavior for everyone else."
Phase 3: High-Level Architecture (≤5 minutes)#
"GeoDNS or anycast picks a region. In the region, routers ECMP across an L4 tier that uses Maglev-style consistent hashing with connection tracking, so flows stay stable when L4 nodes or L7 proxies change. The L7 fleet is stateless, terminates TLS, routes by host and path, and balances to backends with P2C least-request and outlier detection. The control plane pushes endpoints and config incrementally. I'll spend the depth on three places: where the latency goes, how health decisions fail, and how config changes fail."
Phase 4: Transition to Depth (1 minute)#
"The boxes are well known. Where this design succeeds or fails is the TLS cost at the edge, the health and balancing policy toward backends — because that's where load balancers amplify failures — and the config pipeline, because it's the one change that touches every request. Which would you like first?"
Phase 5: Deep Dives (25–30 minutes)#
Deep dive 1 — Where the latency goes (7 min).
"For a new client 80 ms away, two round trips — TCP and TLS 1.3 — cost 160 ms before the first byte of the request arrives. The proxy's own processing is well under a millisecond. So the levers are: terminate TLS as close to the user as possible, maximize connection reuse with HTTP/2 and long keepalives, and support session resumption so returning clients skip the certificate exchange. On the server side, a full handshake with RSA-2048 costs roughly ten times the CPU of ECDSA P-256; serving ECDSA certs to clients that support them is the biggest single lever on edge fleet size."
"Toward backends, the proxy keeps warm connection pools, so the backend hop is ~0.1–0.5 ms in-AZ with no handshake. If the backends are in another AZ, add ~1–2 ms round-trip — which is why I'd keep zone-aware routing on and only spill across zones when local capacity is short."
Deep dive 2 — Balancing and health (12 min).
"Balancing is P2C least-request: pick two healthy hosts at random, send to the one with fewer in-flight requests. With 50 proxies each choosing independently, randomness prevents them from herding onto the same host. New hosts get a slow-start weight that ramps over 60 seconds, so they warm caches and JIT before taking a full share."
"Health has two signals. Active checks every 5 seconds hit a shallow /healthz — the process is up and accepting — and decide admission. Passive outlier detection watches real responses: five consecutive 5xx or a success rate more than ~2 standard deviations below the pool ejects the host for 30 seconds, doubling on repeat. Two guards stop the LB from making things worse: no more than 10% of a pool can be ejected by outlier detection, and if fewer than 50% of hosts are healthy, we enter panic mode and spread across all of them. When half the fleet looks sick at once, the signal is wrong far more often than the fleet is."
Deep dive 3 — Config safety (6 min).
"Config changes — routes, clusters, TLS policies, certs — flow through one pipeline: schema validation, semantic checks (does every route point to an existing cluster? does any timeout exceed the listener's?), then a staged rollout: one canary proxy per region, then one AZ, then one region, then global, with a bake time of 5–10 minutes and automatic rollback if lb.upstream_5xx_rate or lb.added_latency_p99 worsens beyond a threshold. Proxies keep the last-known-good config and refuse a config that fails to load, rather than crash. The data plane must keep serving if the control plane is down."
Deep dive 4 — Draining and long-lived connections (3 min).
"On deploy, a backend is marked draining: no new requests, in-flight finish, HTTP/2 connections get GOAWAY so clients move. For short APIs the drain window is ~2× p99 request time — 10–30 seconds, not the 300-second default many cloud balancers ship with. WebSockets are different: they need an application-level 'reconnect' message and jittered reconnection across 5–10 minutes, or the drain becomes a reconnect storm."
Phase 6: Wrap-Up (2–3 minutes)#
"Summary: two tiers — L4 with consistent hashing for stable flows, L7 for TLS and per-request decisions. P2C least-request with outlier detection, capped ejection and panic mode, so the balancer can't amplify a bad signal into an outage. TLS cost managed with ECDSA and resumption. Config shipped like code with staged rollout and automatic rollback. Next I'd build: east-west balancing in a sidecar with the same policy template, and global steering that shifts load between regions on capacity, not just health."
Common Timing Mistakes#
| Mistake | Time Lost | Fix |
|---|---|---|
| Explaining TCP three-way handshake and OSI layers | 5 min | One sentence: "L4 sees connections, L7 sees requests" |
| Comparing 6 balancing algorithms in depth | 8 min | State P2C least-request and why not least-connections in one breath |
| Designing active-passive failover with VRRP | 5 min | "Stateless fleet behind ECMP; no pairs" |
| Never reaching config safety | Whole interview | Steer there in Phase 4 — it's the largest real-world blast radius |
| Treating the LB hop as the latency problem | 3 min | Quantify: < 1 ms vs 100+ ms for handshakes |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Every candidate has used a load balancer; very few have been paged because one made an incident worse. Interviewers use this question to separate candidates who think of the LB as plumbing from those who think of it as an automated decision-maker with fleet-wide authority. The Staff-level signal is recognizing the specific ways that authority fails: herding, over-ejection, sinkholes, correlated health checks, reconnect storms, and config pushes.
It's also a layering question. DNS, anycast, L4, L7 and client-side balancing each see different information and fail differently. Putting a decision at the wrong layer — trying to do request-level health at DNS, or connection-level stickiness at L7 for UDP — is a design error that no amount of tuning fixes.
1.2 The L5 vs L6 Contrast — Visual#
The Senior path builds a balancer that works when its inputs are true. The Staff path builds one that stays safe when its inputs are wrong.
1.3 The Staff Question That Cuts Through Everything#
"When this load balancer is wrong, how many requests does it get wrong — and how quickly does it stop?"
That one question surfaces the herd effect (wrong about which host is least loaded → all proxies pick it), over-ejection (wrong about health → removes capacity during overload), config pushes (wrong about routes → every request), and draining (wrong about when a host is gone → resets). Each part of the design either bounds the blast radius of a mistake or shortens its duration.
🎯 Staff Move: "I want every automated decision in the balancer to have a bound: ejection capped at 10% of the pool, panic mode below 50% healthy, retries capped at 20% of traffic, and config rolled out one cell at a time. The LB will be wrong sometimes. I'm designing for how wrong and for how long."
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Internet edge (north-south). Clients are many, far away and untrusted. The edge terminates TLS for millions of connections, absorbs abuse, routes by host and path to dozens of services, and is the place where a new connection's latency is mostly network round trips. The design centers on TLS economics, global steering, L4/L7 layering and config safety. This is the intent this case study commits to. For request-level policy on top of the edge — auth, quotas, transformations — see API Gateway.
Service-to-service (east-west). Clients are your own services, a millisecond or less away. Adding a central proxy hop doubles network traversals for every internal call. The design moves balancing into the caller — a library or a sidecar proxy — fed by service discovery. The hard problems shift to control-plane fan-out (pushing endpoint updates to 20,000 sidecars), consistent policy across languages, and retry budgets so a failing service isn't hit by every caller's retries at once.
L4 network load balancing. The unit is the packet or the connection, not the request. Use cases: non-HTTP protocols, UDP, extremely high packet rates, and preserving client IPs. The design centers on forwarding performance (kernel bypass, XDP), consistent hashing with connection tracking so flows survive changes, and direct server return so responses skip the balancer. It's the first tier of the edge design, but as a standalone problem it's a different interview.
| Dimension | Internet Edge | East-West | L4 Network LB |
|---|---|---|---|
| Unit of balancing | HTTP request | RPC | Connection / flow |
| Dominant latency cost | TLS + client RTTs | Extra hop if centralized | Microseconds |
| Health signal | Active + passive on responses | Passive per caller | Connection success, active checks |
| Biggest risk | Global config push | Retry storms, control-plane fan-out | Flow remapping on fleet change |
| Who owns it | Edge platform team | Mesh / RPC platform team | Network team |
2.2 When NOT to Use a (Dedicated) Load Balancer#
| Situation | Better Choice | Why |
|---|---|---|
| Internal RPC between services in one data center | Client-side or sidecar balancing | A central hop adds latency and a shared failure point for no benefit |
| Static assets and cacheable content | CDN — see CDN & Edge Caching | Serving from cache at the edge beats balancing to an origin |
| Stateful partitioned systems (databases, Kafka) | Client routing by partition map | The client must reach a specific node; balancing would be wrong |
| Very low traffic internal tools | DNS with 2–3 records or a single managed LB | Operating a fleet costs more than it protects |
| Batch/queue consumers | The queue itself distributes work | Pull-based consumers self-balance; see Message Queues |
🎯 Staff Move: "For east-west traffic I'd remove the load balancer, not scale it. Each internal call through a central proxy pays an extra hop and shares a failure domain with every other service."
2.3 What the Interviewer Leaves Underspecified#
| Unstated Assumption | Why It's Deliberately Vague | What to Say |
|---|---|---|
| Edge or internal | To see if you ask | "Is this the public edge or service-to-service? The designs diverge." |
| Connection lifetime | To see if you consider WebSockets/HTTP2 | "Are there long-lived connections? That changes draining and balancing." |
| New connection rate | To see if you know what saturates the edge | "How many new TLS connections per second? That sizes the fleet more than RPS." |
| Affinity needs | To see if you default to sticky sessions | "Does anything need affinity, or is state external? I'd default to none." |
| Health definition | To see if you question "healthy" | "What counts as healthy — process up, or serving correctly?" |
| Who changes config | To see if you think about ownership | "Who edits routes, and how fast do changes need to land?" |
2.4 Precise Terminology#
| Term | Precise Meaning | Common Misuse |
|---|---|---|
| L4 load balancing | Decisions per connection/packet using addresses and ports | "Any load balancer that's fast" |
| L7 load balancing | Decisions per request using HTTP semantics | Assumed to be "slow" — the hop is sub-ms |
| VIP | Virtual IP that clients target; backed by many machines | Confused with a single machine's IP |
| ECMP | Router spreads flows over equal-cost next hops by hashing headers | Assumed to keep flows stable when next hops change — it doesn't by itself |
| Direct server return (DSR) | Backend replies to the client directly, bypassing the LB | Assumed to work for L7 — it's an L4 technique |
| Outlier detection | Passive ejection based on observed responses | Confused with active health checking |
| Panic threshold | Healthy fraction below which the LB ignores health and uses all hosts | Seen as unsafe; it prevents self-inflicted outages |
| Slow start | Ramping a new host's weight over time | Confused with TCP slow start |
| Connection draining | Stop new work to a host, let in-flight finish | Assumed to cover long-lived connections automatically |
| Session affinity | Same client → same backend | Used to paper over state that should be external |
3. The Five Fault Lines#
3.1 Fault Line 1: Where to Terminate — L4 vs L7#
L4 is cheap per byte and blind to requests. L7 sees everything and pays for it in CPU, mostly at connection setup. The question isn't which one to use — it's which decisions go where, and who pays for the CPU.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| L4 only, TLS at backends | Cheapest edge; end-to-end encryption to the service; any protocol | No per-request routing, retries or outlier detection; every service manages certs and handshake CPU | Service teams (TLS ops, CPU); users (no request-level failover) |
| L7 only (anycast/DNS straight to proxies) | Full HTTP intelligence; one tier | Proxy changes remap flows without an L4 consistent-hash layer; DDoS lands on expensive L7 CPU | Edge team (fleet size, attack exposure) |
| L4 in front of L7 (two tiers) | L4 absorbs packet floods and keeps flows stable; L7 does request logic | Two fleets to run; one more hop (~tens of µs at L4) | Edge team (two systems) — cheap relative to the alternatives |
| L7 with TLS passthrough for some services | Services that need end-to-end TLS keep it | Those services lose L7 features at the edge | The services that opted in, knowingly |
The Staff default: two tiers. L4 with consistent hashing and DSR in front; L7 terminating TLS and re-encrypting to backends where policy requires (mTLS inside the network). Passthrough as a reviewed exception.
When to deviate: a single-service product at < ~20K RPS can use one managed L7 balancer and skip the L4 tier entirely — the cloud provider's L4 is effectively in front already. Non-HTTP protocols (game servers on UDP, MQTT) go L4 only.
🎯 Staff Move: "L4 for things that need to be fast and stable — flows and floods. L7 for things that need to be smart — routing, retries, health from real responses. And I'd re-encrypt to the backend rather than pass TLS through, unless a service has a compliance reason to own its keys."
Full reasoning: what TLS actually costs
A TLS 1.3 full handshake costs the server one asymmetric signature (the certificate proof) plus a key exchange (ECDHE, X25519 is cheap). The signature dominates: RSA-2048 signing manages roughly a thousand to a few thousand operations per second per core; ECDSA P-256 manages tens of thousands. At 150K new connections/s, RSA-only would need ~75–150 cores just for signatures; ECDSA needs under 10. Serving ECDSA certificates to clients that support them (nearly all modern clients), with RSA as fallback, is the single biggest lever on edge CPU.
Resumption is the second lever. With session tickets or PSK resumption, a returning client skips the certificate signature entirely. Resumption only works if every proxy that might receive the client's next connection can decrypt the ticket — so ticket-encryption keys must be shared across the fleet in a region and rotated on a schedule (e.g., every 12–24 h, keeping the previous key for decryption). A fleet where each proxy has its own ticket key gets near-zero resumption behind a load balancer, and nobody notices until a traffic spike.
The user-facing cost is round trips: TCP (1 RTT) + TLS 1.3 (1 RTT) before the first request byte; TLS 1.2 adds another. HTTP/3 over QUIC combines transport and TLS setup into 1 RTT, and 0-RTT resumption can send replay-safe requests (GETs) immediately — never non-idempotent ones, because 0-RTT data can be replayed.
3.2 Fault Line 2: Even Spread vs Affinity#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Round robin | Simple, stateless | Uneven when request cost varies 10–100×; ignores slow hosts | Users routed to the slow host |
| Least connections | Adapts to slow hosts on one LB | Herding across many LBs; fast-failing hosts attract traffic; meaningless with HTTP/2 multiplexing | Users during host failures; the new host that gets swamped |
| P2C least-request | Near-optimal spread with no coordination; robust to stale data | Still rewarded by fast failure without outlier detection | Nobody, if paired with ejection |
| Consistent hashing (ring/Maglev) | Cache/session affinity; stable mapping on change | Hot keys overload one host; weight changes move keys | The host owning the hot key; users on it |
| Consistent hashing with bounded load | Affinity until a host exceeds ~1.25× average, then spill to next | Some affinity lost under skew | Cache hit rate for hot keys |
| Sticky sessions (cookie) | Server-local session state works | Uneven load; failover loses sessions; drains take as long as sessions | Users whose host dies; ops during deploys |
The Staff default: P2C least-request everywhere; consistent hashing with bounded load only for routes where affinity measurably pays (a local cache with > 2× hit-rate improvement); no cookie stickiness — move session state to a shared store such as Redis.
When to deviate: a backend tier with large per-user in-memory state (real-time collaboration, game sessions) genuinely needs affinity; route by entity ID with consistent hashing and accept that rebalancing is a migration.
🎯 Staff Move: "Affinity is a cache-hit-rate optimization, so I'd only pay for it where I can measure the hit rate. Everywhere else, P2C — and if someone wants sticky sessions to hold session state, I'd rather fix where the state lives."
3.3 Fault Line 3: Trusting Health Signals vs Ignoring Them#
Health checking is a classifier with two error modes. False negatives keep a broken host in rotation — users see errors. False positives remove a healthy host — the remaining hosts get more load, which can cause more false positives. The second error mode compounds; the first doesn't.
| Policy | Detects Bad Host In | False-Positive Risk | What Breaks | Who Pays |
|---|---|---|---|---|
| Active shallow check only (process up) | 10–60 s for dead hosts; never for fast-failing ones | Low | Fast-failing sinkhole stays in rotation | Users routed to the broken host |
| Active deep check (checks DB, cache, deps) | 10–60 s | High and correlated: a shared dependency blip fails every host at once | Fleet-wide ejection during a dependency hiccup | Every user — total outage from a partial one |
| Passive outlier detection, uncapped | < 1 s | Medium; overload causes errors everywhere → eject everything | Cascading ejection under load | Every user |
| Passive + capped ejection (10–20%) + panic threshold (50%) | < 1 s for individual bad hosts | Bounded | Systemic problems aren't masked — they're visible as errors | Users during true systemic failure (unavoidable), not amplified |
The Staff default: shallow active checks (5 s interval, 2 failures to mark down, 2 successes to mark up) for liveness and admission; passive outlier detection (5 consecutive 5xx or success rate > 2σ below peers) for correctness; ejection capped at 10% of hosts; panic below 50% healthy.
When to deviate: for pools of fewer than ~5 hosts, a 10% cap means zero hosts ejectable — raise the cap to allow one host, and alert on any ejection.
🎯 Staff Move: "An LB that ejects aggressively is great when one host is bad and catastrophic when the signal is bad. I'll cap ejection at 10% and stop trusting health entirely below 50% healthy, because at that point the checks are much more likely to be wrong than half my fleet."
3.4 Fault Line 4: Central Proxy vs Client-Side Balancing#
| Model | Latency | Policy Consistency | What Breaks | Who Pays |
|---|---|---|---|---|
| Central L7 fleet for everything | +1 hop per call (~0.2–1 ms) | One place to change | Shared failure domain for all internal traffic; fleet scales with total RPC volume | Every service (latency); edge team (capacity) |
| Client library | No extra hop | Per-language implementations drift | Policy fixes need every service to redeploy; N languages × M versions | Service teams (upgrades); platform team (N libraries) |
| Sidecar proxy (mesh) | ~0.1–0.5 ms per side (local loopback hops) | One proxy implementation, central control plane | Control plane fan-out to every pod; sidecar CPU/memory per pod (~0.1–0.5 core, 50–200 MB) | Infra cost across the fleet; platform team (control plane) |
| Proxyless (gRPC xDS in client) | No hop | Central config, in-process | Only for supported languages | Teams outside the supported languages |
The Staff default: central L7 for north-south; sidecars or proxyless clients for east-west, all fed by the same control plane and the same policy template.
When to deviate: organizations with < ~30 services often do fine with a client library in one language and DNS-based discovery; a mesh is a platform commitment of several engineers.
🎯 Staff Move: "The edge is a central proxy because clients are untrusted and far away. Inside, I'd push balancing to the caller because it's the only place with per-call latency data and no extra hop — but with one policy template, so a timeout default means the same thing everywhere."
3.5 Fault Line 5: Connection Longevity vs Rebalancing#
Long-lived connections are good for latency (no handshakes) and bad for balance (load is pinned to whoever accepted the connection). HTTP/2 and gRPC multiplex thousands of requests on one connection; WebSockets live for hours.
| Approach | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Unlimited connection lifetime | Fewest handshakes | New backends get no load; old ones stay hot; draining takes hours | Hot backends (overload); deploys (slow) |
| Max connection age (e.g., 5–15 min) with jitter | Periodic rebalancing; bounded drain time | ~1 extra handshake per connection per age interval | Small CPU cost; negligible latency with resumption |
| Request-level balancing at L7 over pooled connections | Edge→backend balancing is per request regardless of client connection life | Client→edge connections still pinned to a proxy | Edge proxies (uneven across the fleet) |
| Application-level reconnect for WebSockets | Controlled migration with jitter | Requires client cooperation | Client teams (protocol support) |
The Staff default: client→edge connections capped at a max age (~10 min ± jitter) via GOAWAY; edge→backend balancing per request over pooled connections; WebSockets get an application "reconnect" frame and jittered reconnection spread over 5–10 minutes during drains.
When to deviate: for very latency-sensitive mobile clients on poor networks, lengthen the max age (30–60 min) — each reconnect costs them 2+ RTTs at 200 ms each.
🎯 Staff Move: "I'll balance per request, not per connection, wherever HTTP/2 is involved — otherwise I've built a load balancer that balances once per connection lifetime. And I'll cap connection age so a scale-out actually reaches new hosts within minutes."
4. Failure Modes & Operational Reality#
4.1 The Fast-Fail Sinkhole#
A backend breaks in a way that makes it fast. Load-aware balancing rewards speed, so the broken host attracts traffic.
t=0: Backend b-22 (of 40) loses its connection to the config service after a
cert expires. Handler returns 503 in 1.5 ms. Healthy p50: 35 ms.
t=+1s: Least-request: b-22 has ~0 in-flight at all times. Selected whenever it's
one of the two choices — and wins every comparison.
Share of traffic: 2.5% → ~5% (P2C limits the damage; least-conn would hit ~30%+).
t=+1s: Outlier detection: 5 consecutive 5xx on each proxy → b-22 ejected on most proxies.
t=+2s: Error rate: peak 4.8% for ~1 s, back to 0.05%.
t=+30s: b-22 re-admitted, fails again, ejected 60 s, then 90 s.
t=+5min: Alert: lb.host_ejections{host=b-22} > 3 in 5 min. Auto-replace triggers.
Detection: lb.upstream_5xx_rate by host, lb.host_ejections, latency distribution by host (a host that is suddenly much faster than peers is a signal).
Mitigation: outlier detection (consecutive 5xx + success-rate relative to peers); P2C limits a sinkhole's share to roughly 2/N even before ejection, vs a much larger share under least-connections.
Prevention: health endpoints that exercise the critical in-process dependency state (without calling remote shared dependencies); cert-expiry alerts at 30/14/7 days.
Owner: service team owns the bug and its health endpoint; edge platform owns the outlier policy defaults.
4.2 The Health-Check Cascade#
t=0: Shared auth service p99 rises from 20 ms to 3 s (GC pause storm).
t=+5s: 140 services' /health endpoints call auth to "verify dependencies."
Health checks time out (2 s timeout). Hosts start failing checks.
t=+15s: Pools drop below healthy thresholds. LBs without a panic threshold remove
hosts; remaining hosts take more load, fail checks faster.
t=+30s: 34 services at 0 healthy hosts → LB returns 503 "no healthy upstream"
for 100% of their traffic. Auth itself has recovered.
t=+90s: Hosts pass 3 consecutive checks again; pools refill gradually.
t=+4min: Full recovery. A 30-second auth blip became a 4-minute multi-service outage.
Detection: lb.healthy_hosts_pct per cluster; lb.no_healthy_upstream count; correlation of check failures across many services at the same moment (a correlated health drop is almost never real).
Mitigation: panic threshold — below 50% healthy, ignore health and use all hosts; if hosts are actually fine, traffic keeps flowing.
Prevention: health endpoints must not call remote shared dependencies; lint for it; dependency failures surface through passive outlier detection on real requests instead. Game day: inject latency in a shared dependency and verify no pool goes empty.
Owner: edge platform (panic threshold, policy); each service team (their /health); the platform that owns the health-endpoint standard.
4.3 The Global Config Push#
t=0: Route change merged: a new regex path matcher for a marketing page.
Pattern has catastrophic backtracking on certain paths.
t=+20s: Config pipeline validates syntax (passes) and pushes globally — no canary
stage for "route-only" changes.
t=+30s: Proxy CPU 30% → 100% across every region. Added latency p99 2 ms → 9 s.
t=+1min: Page: lb.added_latency_p99 > 500 ms, lb.cpu > 90% in all regions.
t=+6min: On-call identifies the change via config version on dashboards.
t=+9min: Rollback pushed. Proxies at 100% CPU take 2–3 min to apply it.
t=+13min: Recovered. 99.99% monthly budget (4.3 min) exceeded 3× in one incident.
A well-known public example of this class: in July 2019, Cloudflare deployed a WAF rule containing a regular expression that caused excessive backtracking; deployed globally at once, it drove CPU to exhaustion across its network for roughly half an hour.
Detection: lb.cpu_utilization, lb.added_latency_p99, config version annotations on every edge dashboard.
Mitigation: one-click rollback to last-known-good; proxies apply config on a separate thread pool or core reservation so a CPU-saturated data plane can still accept a rollback.
Prevention: every config change — including "just a route" — goes through canary → AZ → region → global with bake time; regex engines with linear-time guarantees (RE2-style) for any user-authored pattern; CPU-cost benchmarks of new config in CI.
Owner: edge platform team owns the pipeline and the guardrails; the route author owns the change.
4.4 The TLS Handshake Storm#
t=0: Scheduled ticket-key rotation runs. A bug distributes the new key to the
proxies without keeping the previous key for decryption.
t=+0s: Every resumption attempt fails → full handshakes.
Resumption rate: 65% → 0%. Full handshakes/s: 50K → 145K.
t=+40s: Proxy CPU 55% → 97%. Handshake latency p99 50 ms → 1.2 s.
Mobile clients time out at 10 s, retry → more handshakes.
t=+2min: Page: tls.full_handshakes_per_sec > 2× baseline, lb.cpu > 90%.
t=+5min: On-call scales the L7 fleet +40% (takes 4 min to warm).
t=+9min: Old key restored for decryption. Resumption back to 60%. CPU 60%.
Detection: tls.resumption_rate, tls.full_handshakes_per_sec, tls.handshake_latency_p99, proxy CPU.
Mitigation: restore previous ticket key; scale out; prefer ECDSA (if RSA was serving a share of clients, shifting cuts handshake CPU ~5–10×).
Prevention: ticket keys rotated with overlap (current + previous for decryption); resumption-rate alert at < 50% of baseline; capacity plan for zero resumption at peak, because a key mistake or a mass client reconnect produces exactly that.
Owner: edge platform (cert and key management).
4.5 The Long-Lived Connection Imbalance#
t=0: Partner gRPC traffic grows. Backend pool scales from 20 to 40 hosts.
t=+10min: The 20 new hosts serve ~3% of requests. The 20 old hosts at 85% CPU.
Clients hold HTTP/2 connections for hours; L4 balancing at the backend
tier balanced once, at connect time.
t=+30min: Old hosts p99 rises from 80 ms to 900 ms. Autoscaler adds more hosts —
which also get no traffic.
t=+40min: On-call restarts old hosts in batches → mass reconnect → spread.
Detection: request distribution skew across hosts (max/mean > 2), new hosts' RPS after scale-out, connection age distribution.
Mitigation: per-request balancing at an L7 hop in front of gRPC backends, or client-side balancing in the gRPC client; max connection age with GOAWAY.
Prevention: never use L4-only balancing for multiplexed protocols; alert on per-host RPS skew; scale-out tests that check new-host traffic share within 5 minutes.
Owner: edge platform (protocol-aware balancing); partner integration team (client config).
4.6 The L4 Rehash#
t=0: L4 tier scaled from 8 to 10 nodes for an event. Routers' ECMP next-hop
set changes. A naive mod-N hash would remap ~80% of flows to different
L4 nodes.
t=+0s: L4 nodes without shared consistent hashing forward remapped packets to
different L7 proxies than the ones holding the TCP state → RST.
t=+1s: ~800K established connections reset. Clients reconnect at once:
full TLS handshakes spike 10×.
t=+2min: L7 CPU saturates from the handshake storm (see 4.4).
Detection: tcp.resets_sent at L7, l4.flow_table_misses, sudden handshake spikes coinciding with L4 fleet changes.
Mitigation: consistent hashing (Maglev-style) on every L4 node using the same backend table, so any L4 node sends a given flow to the same L7 proxy; connection tracking for in-flight flows during backend set changes.
Prevention: L4 fleet changes during low-traffic windows; consistent hashing is mandatory; validate in staging that adding an L4 node resets < 1% of flows.
Owner: network / L4 team.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Fast-fail sinkhole | Per-host 5xx; host faster than peers | ~2/N of traffic with P2C until ejected | Outlier ejection | Service team (bug); edge (policy) |
| Health-check cascade | lb.healthy_hosts_pct drops across many clusters at once | Every service sharing the dependency | Panic threshold | Edge platform + service health owners |
| Global config push | lb.cpu, lb.added_latency_p99 in all regions | Everything behind the edge | Rollback to last-known-good | Edge platform |
| TLS handshake storm | tls.resumption_rate drop, handshake/s spike | A region or global | Restore ticket keys; scale out | Edge platform |
| Long-lived imbalance | Per-host RPS skew > 2× | One service | Per-request L7 balancing; max connection age | Edge + client teams |
| L4 rehash | TCP resets, flow-table misses | All connections through the changed tier | Consistent hashing + conn tracking | Network team |
| Retry amplification | lb.retry_rate > 20% of requests | Failing backend and its dependencies | Retry budget; single retry layer | Service team (policy), edge (budget enforcement) |
| Control-plane outage | xds.config_age_seconds rising | No config changes; data plane must continue | Last-known-good; static fallback | Edge platform |
| Zone imbalance | Per-zone RPS vs capacity | One zone overloaded | Zone-aware routing with spillover | Edge platform |
🎯 Staff Insight: Five of these nine failures are the load balancer amplifying a smaller problem — one bad host, a dependency blip, a key rotation, a fleet change. A good LB design is mostly a list of bounds on its own automated decisions.
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | Draws an LB in front of servers | Separates edge, east-west and L4; commits to one | Asks how many balancing layers exist and which are redundant |
| Latency | Minimizes hops | Locates cost in TLS and RTTs; resumption, ECDSA, reuse | Prices handshake CPU across the fleet; sets TLS policy org-wide |
| Algorithm | Least connections or round robin | P2C least-request; bounded-load hashing only where affinity pays | Standardizes defaults in a shared template; reviewed overrides |
| Health | Active checks | Shallow active + passive outlier detection; capped ejection; panic mode | Owns the health-endpoint standard to prevent correlated ejection |
| Failure | Active-passive pairs | Stateless fleets, consistent hashing, staged config with auto-rollback | Treats the config pipeline as the org's highest-blast-radius deploy system |
| Connections | Not considered | Max connection age, GOAWAY, per-request balancing for HTTP/2, WebSocket drain | Sets protocol policy (HTTP/3, 0-RTT scope) with security and client teams |
| Ownership | Infra owns the LB | Edge owns the fleet; services own routes, timeouts, health | Defines platform vs service contracts and the validation that enforces them |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Locates the real latency | "The proxy hop is under a millisecond. The two round trips for TCP and TLS are 160 ms for a user 80 ms away." |
| Knows how balancing herds | "Fifty proxies running least-connections all see the new host as empty at the same moment." |
| Bounds automated decisions | "Ejection capped at 10%, panic below 50% healthy, retries capped at 20%." |
| Separates liveness from correctness | "Active checks tell me the process is up. Real responses tell me it's working." |
| Treats config as a deploy | "Every route change goes canary, AZ, region, global — with automatic rollback." |
| Handles connection lifetime | "With HTTP/2 I balance per request, or I've balanced once per connection lifetime." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| Spends the interview comparing algorithms | Algorithm choice is rarely the failure; health and config are |
| Deep health checks that call the database | Creates correlated, fleet-wide ejection on a dependency blip |
| Active-passive pair as the HA story | Ignores that both members run the same bad config |
| Sticky sessions by default | Hides a state-placement problem and makes failover and deploys worse |
| Retries at every layer | Multiplies load during the exact incidents retries are meant to survive |
| "The LB adds latency" with no numbers | Misplaces the cost — and optimizes the wrong thing |
5.4 Common False Positives#
- Knowing Maglev's lookup-table math ≠ edge design. Impressive, but the interview is about bounding failures.
- Listing every cloud LB product ≠ judgment. Product names without policies (ejection caps, drain windows) are a catalog, not a design.
- Packet-level depth ≠ L7 understanding. Candidates strong on XDP and DSR sometimes miss request-level health entirely.
- "Use a service mesh" ≠ east-west design. A mesh is a mechanism; the policy (timeouts, retries, outlier rules) is the design.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–4 min | Commit to edge; capture RPS, new-connection rate, protocols, latency budget |
| Entities and API | 4–7 min | Listener, route, cluster, endpoint; who owns what |
| Architecture | 7–12 min | DNS/anycast → L4 → L7 → backends; control plane |
| Latency and TLS | 12–19 min | Where time goes; resumption, ECDSA, reuse |
| Balancing and health | 19–31 min | P2C, outlier detection, caps, panic, slow start |
| Config safety | 31–37 min | Validation, staged rollout, rollback, data plane independence |
| Connections and draining | 37–42 min | Max age, GOAWAY, WebSockets |
| Wrap-up | 42–45 min | Summary, east-west, global steering |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response |
|---|---|---|
| "Traffic triples overnight." | What saturates first | New-connection rate × handshake CPU; resumption rate; scale L7, not L4 |
| "An entire AZ goes dark." | Zone awareness, capacity | Remaining zones need N+1 headroom (~50% extra at 3 AZs); spillover policy; ECMP withdraws routes |
| "A backend returns errors fast." | Sinkhole awareness | Outlier detection; why least-connections makes it worse |
| "The database behind every service is slow." | Over-ejection | Panic threshold, capped ejection, shallow health checks |
| "We need sticky sessions." | Pushback on state | Externalize session; if truly needed, bounded-load consistent hashing |
| "Make it multi-region." | Global steering | GeoDNS/anycast, capacity-aware shifting, DNS TTL realities |
| "Design the east-west version." | Layer judgment | Remove the central hop; sidecar or proxyless with same policy template |
6.3 What to Deliberately Skip#
- OSI model recitation. "L4 sees connections, L7 sees requests" is enough.
- VRRP/keepalived active-passive details. Stateless fleets behind ECMP replace them.
- Exact Maglev table-population algorithm. Name it, state the property (minimal disruption), move on.
- Cipher suite lists. "TLS 1.3, ECDSA with RSA fallback, resumption" covers it.
- WAF rule design. It belongs to an edge-security interview.
6.4 Follow-Up Questions to Expect#
- "How does a new L7 proxy join without disrupting connections?" — L4 consistent hashing with connection tracking: existing flows stay pinned; only new flows map to the new proxy.
- "How do you drain a proxy for upgrade?" — Withdraw it from the L4 table for new flows, send GOAWAY /
Connection: close, wait for in-flight (≈ p99 request time), cap at a few minutes; WebSockets get reconnect frames. - "How do you pick the retry layer?" — Retry where you have the most context and the fewest multipliers: usually the edge for idempotent requests, with a budget; backends don't retry each other's retries.
- "How do you balance across zones?" — Prefer local zone; spill proportionally when local healthy capacity falls below demand; avoid permanent cross-zone traffic (latency + transfer cost).
- "What if the control plane is down?" — Proxies keep last-known-good config and endpoint lists; no new routes ship; endpoint churn is absorbed by outlier detection.
- "How do you protect the LB from DDoS?" — Anycast spreads volumetric attacks; L4 drops malformed packets and SYN floods (SYN cookies); L7 rate limits per client — see Rate Limiting.
- "How do you roll out HTTP/3?" — Advertise via Alt-Svc to a percentage of clients, compare latency and error rates, keep TCP fallback; UDP needs L4 support for QUIC connection IDs.
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design a load balancer."
Staff Answer
"Three different systems answer to that name: the internet edge in front of a public service, balancing between internal services, and a pure L4 packet balancer. I'll design the edge for a latency-sensitive HTTP API, sketch the L4 tier it sits on, and contrast east-west at the end.
Assumptions: 1M RPS globally, ~150K new TLS connections/s at peak, HTTP/1.1, HTTP/2 and some WebSockets, three regions with three AZs, 99.99% availability, added p99 under 5 ms excluding client RTTs. Plan: where latency goes → balancing and health, because that's where balancers amplify failures → config safety, because it's the largest blast radius → connection lifetime and draining."
Why this is L6:
- Distinguishes three intents with different designs and commits
- Asks for new-connection rate, not just RPS — the number that actually sizes the edge
- Orders the deep dives by where failures come from
What L7 adds:
- Asks which existing balancing layers this replaces or joins, and who owns them
- Frames the config pipeline as an org-level risk surface with its own change policy
❌ Common L5 Trap
"Two load balancers in active-passive behind a floating IP, least-connections to a pool of web servers, health check every 10 seconds, sticky sessions for logged-in users."
Why this misses: It's a working design for one data center on a good day. It herds on scale-out, sinkholes on fast failure, loses sessions on host death, and has no answer for a config mistake that both LBs share.
Drill 2: The Algorithm#
Prompt: "Why not least-connections?"
Staff Answer
"Three reasons. First, with many proxies each seeing only their own connections, they all see a newly added or recovered host as least-loaded at the same instant and herd onto it. Second, with HTTP/2 and keepalive, a connection carries anywhere from zero to hundreds of concurrent requests, so connection count isn't load. Third, a host that fails fast always has the fewest active requests, so it attracts the most traffic.
P2C least-request fixes the first two: each proxy samples two hosts at random and picks the one with fewer in-flight requests. Randomness spreads decisions across proxies, and the theory says two choices get you most of the benefit of perfect information — maximum load above average drops from roughly log n / log log n to log log n. The third needs outlier detection, because any load-aware algorithm rewards a fast failure."
Why this is L6:
- Names the specific failure modes, not just a preference
- Knows why P2C works with stale, partial information
- Admits P2C's remaining weakness and pairs it with ejection
What L7 adds:
- Puts P2C + outlier detection in the shared default template so teams can't drift into least-connections
- Ties algorithm choice to an incident class the org has actually had
Drill 3: The Health Check#
Prompt: "Design the health checking. A teammate wants
/healthto check the database, cache and auth service."
Staff Answer
"I'd push back. If /health checks shared dependencies, a 20-second blip in auth fails every host's check at the same moment, and the LB removes the whole fleet — turning a partial dependency problem into a total outage. Active checks should be shallow: the process is up, the listener accepts, in-process critical state is initialized. 5 s interval, 2 failures down, 2 successes up.
Correctness comes from real traffic: passive outlier detection ejects a host after 5 consecutive 5xx or when its success rate is > 2σ below peers, for 30 s doubling on repeats, capped at 10% of the pool. And if the healthy fraction falls below 50%, panic mode spreads traffic across all hosts. If the database is truly down, users will see errors either way — but the LB won't add 'no healthy upstream' on top for the minutes it takes hosts to pass checks again."
Why this is L6:
- Identifies correlated failure from deep checks
- Separates liveness (active) from correctness (passive)
- Bounds the LB's own ability to cause an outage
What L7 adds:
- Turns "health endpoints must not call remote shared dependencies" into an org standard with a lint and a game day
- Measures it: count of services whose pools went empty during the last shared-dependency incident
Drill 4: Make It Concrete — Sizing the Edge#
Prompt: "400K RPS in the largest region, 60K new TLS connections/s, average response 20 KB. Size the L7 and L4 tiers."
Staff Answer
"L7 has two loads: requests and handshakes. Requests: at ~15K simple proxied req/s per core including TLS record encryption, 400K RPS needs ~27 cores. Handshakes: assume 50% resumption and ECDSA for 95% of clients — 30K full handshakes/s; at ~20K ECDSA signatures/s/core that's ~1.5 cores, plus ~1.5K RSA handshakes/s at ~1.5K/s/core = ~1 core. Key exchange and record setup add a few more. Call it ~35 cores busy. Then plan for zero resumption — a key bug or mass reconnect — which roughly doubles handshake cost, and N+1 per AZ with 50% headroom so we survive losing an AZ: ~35 × 2 (headroom) × 1.5 (AZ loss) ≈ 105 cores → ~8 proxies of 16 cores, spread 3 per AZ, so 9.
L4: with DSR, the L4 tier sees only inbound packets. 400K RPS with a few packets per request plus ACKs is a few million packets/s — one or two XDP-based nodes can forward that; I'd run 4 across AZs for redundancy. Egress bandwidth, 400K × 20 KB ≈ 8 GB/s ≈ 64 Gbps, flows from L7 proxies to clients and never touches L4 — that's what makes the L7 NICs, not their CPUs, a likely limit: 9 proxies × 25 Gbps leaves headroom, while 9 × 10 Gbps doesn't survive losing an AZ (6 × 10 = 60 Gbps < 64)."
Why this is L6:
- Sizes L7 on both requests and handshakes and plans for zero resumption
- Includes AZ-loss headroom explicitly
- Notices DSR shrinks L4 and that NIC bandwidth may bind before CPU
What L7 adds:
- Converts the RSA fallback share into dollars: dropping it from 5% to 1% saves little; the ECDSA switch itself saved ~80% of handshake CPU
- Asks whether static content could move to a CDN, removing a large fraction of egress from the fleet entirely
Drill 5: The Dependency Goes Down#
Prompt: "A backend service's database slows down. 60% of its hosts start timing out. What does your LB do, minute by minute?"
Staff Answer
"First seconds: outlier detection starts ejecting hosts with consecutive timeouts, but hits the 10% ejection cap almost immediately — so most slow hosts stay in. That's intentional: the hosts aren't broken, their dependency is. P2C least-request steers away from hosts with long in-flight queues toward whichever are momentarily faster. Requests hit the route timeout — say 2 s — and the edge retries idempotent requests once, within a retry budget of 20% of traffic, so retries can't double load on a struggling service.
If the active checks were deep, we'd be in panic mode now; with shallow checks, hosts stay healthy and we serve what we can. Errors are visible to users and to the service's on-call — accurately attributed to the service, not to 'no healthy upstream'. The fix is in the service: shed load, fail fast with a circuit breaker on the database, serve degraded responses. The LB's job is to not make it worse."
Why this is L6:
- Walks through how each safety rail engages
- Uses a retry budget instead of unconditional retries
- Places the fix with the service and keeps the LB honest about error attribution
What L7 adds:
- Adds per-service load-shedding headers (e.g., backend signals overload; edge stops retrying) as a platform contract
- Defines which services get priority edge capacity when the fleet is constrained
Drill 6: The Hot Key#
Prompt: "A cache-heavy service uses consistent hashing on
user_idfor affinity. A celebrity account now generates 8% of its traffic."
Staff Answer
"Plain consistent hashing pins that 8% to one host, which in a 30-host pool is ~2.4× average load. I'd switch to consistent hashing with bounded load: a host may take at most ~1.25× the average in-flight load; above that, requests for its keys spill to the next host on the ring. The celebrity's requests spread over two or three hosts, which each warm their caches for that key — a small hit-rate cost on a single key in exchange for no overload.
Longer term: if one key is that hot, it should be cached closer — at the edge or in a per-proxy micro-cache with a 1–5 s TTL — so most of its requests never reach the pool."
Why this is L6:
- Quantifies the overload from the skew
- Uses bounded-load hashing to keep most affinity while capping any single host
- Moves the hottest content to a better layer
What L7 adds:
- Makes bounded load the default for every hashing policy in the template — unbounded hashing is an exception that needs review
- Connects to the CDN strategy: hot keys are an edge-caching problem first
Drill 7: Multi-Tenant Edge#
Prompt: "60 services share the edge fleet. One team's route change or traffic spike shouldn't hurt the others. How?"
Staff Answer
"Three kinds of isolation. Config: routes are owned per service, validated by the platform — a route can't change listener-level behavior, can't use non-linear-time regex, and can't set timeouts above the listener maximum. Changes roll out staged per cell, and a bad route rolls back automatically. Capacity: per-service circuit-breaker limits on the proxy — max connections, max pending requests, max concurrent retries to each cluster — so one slow service can't consume all proxy workers or connection pools. Traffic: per-service and per-client rate limits at the edge; an attack on one hostname is shed before it consumes shared CPU.
For the two or three services with very different needs — a large WebSocket service, for example — I'd give them a separate proxy pool behind the same L4 tier, so their connection counts don't dictate the shared fleet's memory profile."
Why this is L6:
- Covers config, capacity and traffic isolation separately
- Uses per-cluster limits on the proxy to stop resource exhaustion from one backend
- Recognizes when workloads should get separate pools
What L7 adds:
- Defines the platform/service contract: what service teams may configure, and what's reserved
- Uses per-service cost attribution (handshakes, egress) to make outlier tenants visible
Drill 8: Build vs Buy#
Prompt: "Cloud-managed load balancer, self-run open-source proxy, or build our own?"
Staff Answer
"Don't build a proxy — the data plane is a solved problem and owning its CVEs is a permanent tax. The real choice is managed vs self-run.
Managed L7 balancers are the default up to tens of thousands of RPS per service and when the team has no edge specialists: no fleet, built-in cert management, pay per use. Self-running an open-source proxy fleet (Envoy, HAProxy, NGINX) pays off when you need control the managed product doesn't give — custom outlier policies, per-request routing logic, multi-cloud consistency, east-west mesh with the same proxy — or when per-request pricing at hundreds of thousands of RPS exceeds the cost of a 3–5 person edge team. What I'd always build in-house is the config pipeline: validation, staged rollout and rollback — because that's where our outages come from, managed or not."
Why this is L6:
- Rejects building the data plane with a reason
- Gives triggers for each option
- Identifies the config pipeline as the part to own regardless
What L7 adds:
- Prices it: managed per-request and per-connection charges at 400K RPS vs ~$1M/yr for a small edge team
- Considers portability: self-run proxy config is portable across clouds; managed LB config isn't
Drill 9: Changing Routing Without an Outage#
Prompt: "Move 30% of traffic for
/checkoutto a new backend version. Then all of it. Safely."
Staff Answer
"Weighted routing on the route: checkout-v1: 99, checkout-v2: 1. Step through 1% → 5% → 30% → 100% with a bake of 10–30 minutes each, comparing per-version error rate and latency, auto-reverting if v2's 5xx rate exceeds v1's by 0.5 percentage points or p99 by 20%. The config change itself rolls out canary → AZ → region → global, so a mistake in the weight change is contained too.
Two subtleties. Sessions: if checkout carries state between requests, I'd hash on user ID for the split so a user doesn't bounce between versions mid-checkout. Draining v1 at the end: set weight to 0, wait for in-flight requests to finish — a drain window of 2× p99 — then remove the endpoints."
Why this is L6:
- Uses weighted routing with explicit, automatic success criteria
- Distinguishes the traffic shift from the config rollout of the shift
- Handles per-user consistency during the split
What L7 adds:
- Makes progressive delivery a platform capability with standard metrics, not a per-team script
- Ties rollout gates to the service's error budget
Drill 10: Multi-Region#
Prompt: "We're going from one region to three. How does global load balancing work, and what happens when a region fails?"
Staff Answer
"Two mechanisms. GeoDNS returns the nearest healthy region's VIPs with a TTL of 30–60 s; or anycast announces one VIP from all regions and BGP picks the nearest. Anycast reacts faster — withdrawing routes shifts traffic in seconds — while DNS failover takes minutes in practice because resolvers and clients hold records past the TTL.
Region failure is a capacity problem more than a routing problem. If each of three regions runs at 60% at peak, losing one sends ~30% extra to each survivor — 90%, uncomfortably close to saturation. So either each region runs at ≤ 50–60% with explicit failover headroom, or failover is capacity-aware: shift only what survivors can absorb and shed or degrade the rest. I'd also make failover a deliberate decision with criteria, and test it with quarterly region drains — the first real failover should not be the first time the survivors see that load."
Why this is L6:
- Compares DNS and anycast with realistic reaction times
- Recognizes failover as a capacity question with numbers
- Insists on regular drills
What L7 adds:
- Prices the headroom: running at 55% instead of 75% is ~35% more edge and backend capacity — a business decision with a dollar figure
- Coordinates with stateful tiers: shifting traffic to a region whose database is a read replica is a multi-region data question, not an LB question
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Incident — Launch Day Handshake Saturation#
Context: A marketing campaign drives a product launch. Requests per second rise 3×, but edge proxy CPU goes from 40% to 98% and handshake latency p99 reaches 2 seconds. Backends are at 50% CPU and fine. The on-call has started adding proxies; it takes 6 minutes per batch to warm up. You're pulled in.
Questions to Surface First:
- New connections per second vs baseline? Is this a request problem or a connection problem?
- What's the resumption rate? It should drop with new users — by how much?
- What fraction of handshakes use RSA vs ECDSA?
- Are clients retrying timed-out handshakes, compounding the load?
Typical L5 Approach: Scales the proxy fleet and waits. Possibly raises proxy worker counts. Recovers in 20–30 minutes once capacity arrives.
Staff Approach: Sees that new users mean near-zero resumption: full handshakes rose 8× while RPS rose 3×. Finds that the cert served is RSA-only because an ECDSA cert was never provisioned for the launch hostname. Deploys the ECDSA cert (canary first), cutting handshake CPU ~5×, while scaling out in parallel. Sets client retry backoff via config for the mobile app's next session.
Principal Approach: Makes "new-connection capacity at zero resumption" a launch-readiness criterion, and moves cert provisioning for new hostnames into the platform with dual ECDSA/RSA as the only supported option. Asks marketing to share campaign calendars with the edge team as part of the launch process.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Check tls.full_handshakes_per_sec (8× baseline) vs RPS (3×). Check tls.resumption_rate (65% → 12%). Confirm cert type on the launch hostname: RSA-2048 only. |
| Triage | Handshake CPU dominates: ~85% of proxy CPU in signature operations. Client retries on 10 s timeout add ~30% more handshakes. |
| Quick fix | Push dual ECDSA + RSA cert for the hostname through the cert pipeline, canary one proxy per AZ, then fleet. CPU 98% → 55% within 4 minutes. Continue scale-out for headroom. |
| Guardrails | Alert on tls.full_handshakes_per_sec > 2× baseline; capacity model includes zero-resumption scenario; cert lint rejects RSA-only certs for new hostnames. |
| Post-mortem | Why was RSA-only provisioned for a new hostname? Why didn't launch readiness include connection-rate load testing? |
Metrics to Watch: tls.full_handshakes_per_sec, tls.resumption_rate, tls.handshake_latency_p99, lb.cpu_utilization, tls.handshakes_by_sig_alg
Organizational Follow-up: edge platform owns cert provisioning (no hand-made certs); launch checklist includes a connection-rate test at 3× expected new users.
Ownership Question: "Who owns making sure the edge can handle a marketing launch?" Staff answer: The edge platform owns capacity and the readiness test. The launching team owns telling us — with the expected new-user count — two weeks ahead. If they don't, the readiness gate blocks the hostname from going live.
Key Takeaway: "Launch traffic is new-user traffic, and new users don't resume TLS sessions. Size the edge on new connections at zero resumption, not on requests."
What clears the Staff bar:
- Distinguishes connection growth from request growth
- Finds the cert type as a 5× CPU lever
- Fixes the cause while scaling, not instead of fixing
Deep Dive 2: Silent Failure — One Zone Running Hot for Weeks#
Context: A capacity review shows that one AZ's backends run at 75% CPU while the other two run at 35%. Latency p99 for users routed through that zone has been 40% worse for at least three weeks. No alert fired; overall SLOs were met. Cross-AZ transfer costs also rose 20% last month.
Questions to Surface First:
- Is zone-aware routing enabled? What's the spillover rule?
- Is the L7 fleet itself evenly spread across zones? Are the L4 tier's ECMP weights equal?
- Did a deploy change the number of backends per zone (e.g., one zone got fewer hosts after a capacity shortage)?
- Where does the cross-AZ transfer come from?
Typical L5 Approach: Adds hosts in the hot zone until CPU evens out.
Staff Approach: Finds the mechanism: an instance-type shortage left the hot zone with 12 backends vs 20 in the others, but zone-aware routing sends traffic in proportion to proxy placement, not backend capacity. The hot zone's proxies keep traffic local until hosts are "unhealthy", which they never quite become. Meanwhile, other services had spillover on, creating the cross-AZ transfer. Fixes routing to weight by healthy backend capacity per zone.
Principal Approach: Treats the absence of an alert as the main finding. Adds per-zone utilization skew as an SLO-adjacent signal across all services, and makes "capacity-weighted zone routing" the platform default so zone imbalance self-corrects.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Compare per-zone backend count, per-zone RPS and per-zone CPU. 12 vs 20 vs 20 hosts; RPS ~equal per zone. |
| Triage | Zone-aware routing keeps traffic local unless local healthy hosts fall below a threshold — capacity-unaware. Other services spill across zones at a fixed 30%, driving transfer costs. |
| Quick fix | Enable capacity-weighted zone routing: local share = local capacity / local demand, spill the remainder proportionally. Hot zone CPU 75% → 52%; cross-zone share for this service ~12%. |
| Guardrails | Alert on zone.utilization_skew = max/min > 1.5 for 1 h. Dashboard of cross-AZ bytes per service. |
| Post-mortem | Why could a zone run 3 weeks at 2× the others with no signal? Why was spillover a fixed percentage? |
Metrics to Watch: zone.utilization_skew, lb.cross_zone_request_pct, network.cross_az_bytes, per-zone latency_p99
Organizational Follow-up: capacity team reports per-zone host counts after any shortage; edge platform owns zone-routing defaults; service dashboards show per-zone breakdowns by default.
Ownership Question: "Who notices a zone imbalance?" Staff answer: The edge platform, through a fleet-wide skew alert. Individual service teams see aggregate SLOs that hide it; that's why it went three weeks.
Key Takeaway: "Averages hide zone skew. Route by capacity, not by location alone, and alert on the skew itself."
What clears the Staff bar:
- Finds the mismatch between proxy placement and backend capacity
- Fixes routing policy instead of throwing hosts at it
- Recognizes cross-AZ cost as a second symptom of the same policy
Deep Dive 3: Large-Customer Onboarding — 400K WebSockets From One NAT#
Context: An enterprise customer integrates a real-time dashboard product. They'll open ~400K WebSocket connections, all from a handful of corporate NAT egress IPs, and require their traffic to come from fixed IPs they can allowlist (for callbacks) and mTLS for their API calls. Go-live is in five weeks.
Questions to Surface First:
- Does any layer hash on source IP? If so, a few NAT IPs land on a few proxies.
- What's per-connection memory on the proxies with TLS? 400K × ~30 KB = ~12 GB across the fleet — fine if spread, a problem if concentrated.
- How will these connections drain during deploys — what's the reconnect behavior of their client?
- Is mTLS terminated at the edge with cert validation against their CA, or passed through?
Typical L5 Approach: Adds proxies to absorb the connections and enables sticky sessions by source IP for the WebSockets.
Staff Approach: Avoids source-IP hashing — with a handful of NAT IPs it puts thousands of connections on a few proxies. The L4 tier hashes on the full 5-tuple (source port varies), which spreads well. Places the WebSocket service on a dedicated proxy pool behind the same L4 tier so 400K long-lived connections don't shape the shared fleet. Agrees on a reconnect protocol: server sends a reconnect frame during drains; client reconnects with 0–300 s jitter. mTLS terminated at the edge with their CA in a per-tenant trust bundle; fixed egress IPs via a dedicated NAT for callbacks.
Principal Approach: Uses the deal to define an enterprise-edge tier — dedicated pools, tenant trust bundles, static egress, contractual reconnect behavior — priced into the enterprise SKU rather than built bespoke for one customer.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (week 1) | Load-test 400K connections from 4 source IPs against staging. Observe distribution with source-IP hashing (top proxy: 31% of connections) vs 5-tuple (top proxy: 4.1%). |
| Triage (design) | Dedicated WebSocket pool: 8 proxies, ~50K connections each, ~1.5 GB connection memory each. Max connection age 6 h with reconnect frame. |
| Quick fix (rollout) | Start at 10% of their users, then 50%, then 100% over two weeks. Drain test: deploy during business hours at 10% to verify reconnect jitter. |
| Guardrails | Alert on lb.connections_per_proxy skew > 1.5×; ws.reconnects_per_sec > 2× baseline. |
| Post-mortem (retrospective) | Document the enterprise-edge tier and its cost. |
Metrics to Watch: lb.active_connections per proxy, ws.reconnects_per_sec, tls.mtls_handshake_failures by tenant, proxy memory
Organizational Follow-up: edge platform owns the dedicated pool; the real-time product team owns the reconnect protocol; sales engineering knows that fixed egress and mTLS are enterprise-tier features.
Ownership Question: "Who decides when to drain the WebSocket pool?" Staff answer: The real-time product team schedules deploys; the edge platform enforces the drain protocol and rate. Neither can drain without the other's mechanism.
Key Takeaway: "Source-IP affinity fails exactly for enterprise customers, who arrive through a few NAT addresses. Long-lived connections need their own pool and a negotiated reconnect protocol."
What clears the Staff bar:
- Predicts the NAT hot spot and measures it before go-live
- Isolates long-lived connections from the shared fleet
- Designs draining as a protocol with the client, not a timeout
Deep Dive 4: Post-Mortem — 13 Minutes of Global Latency and 5xx From a Route Change#
Context: You lead the post-mortem for incident 4.3: a route change containing a backtracking regex was pushed globally and saturated every edge proxy for about 13 minutes. Initial root cause: "bad regex." Leadership asks how to make sure it can't happen again.
Questions to Surface First:
- Why did a route change skip canarying? Is there a class of change that's considered "safe"?
- Why could a service team author an arbitrary regex in edge config?
- Why did rollback take 3+ minutes to apply?
- Could the proxies have protected themselves — per-route CPU limits, linear-time regex?
Typical L5 Approach: Root cause: bad regex. Action items: review regexes more carefully; add a regex linter.
Staff Approach: The regex was the trigger. Root causes: (1) route changes bypassed staged rollout because they were classed "low risk"; (2) the regex engine allowed backtracking; (3) config application competed for CPU with the saturated data plane, slowing rollback. Fixes all three: no change class bypasses staging; linear-time regex engine only; config application reserved a core.
Principal Approach: Reframes the edge config pipeline as safety-critical infrastructure. Introduces an error-budget policy: when edge config changes consume more than 25% of the monthly budget, non-emergency edge config changes freeze until the pipeline gains a new safeguard. Reviews every system with global config push in the company — DNS, feature flags, WAF — for the same pattern.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Confirm all regions on last-known-good config. Freeze edge config changes pending review. |
| Triage | Pipeline history: "route-only" changes exempt from canary since a year ago to speed up marketing launches. Regex engine: backtracking. Rollback apply time: 3–4 min under 100% CPU. |
| Quick fix | Remove the exemption. Switch to a linear-time regex engine; reject constructs it can't support. Reserve one core per proxy for control-plane work. |
| Guardrails | CI benchmark: every config change measured for CPU per request on replayed traffic; > 10% regression blocks. Auto-rollback if lb.cpu rises > 20 points in a canary cell. |
| Post-mortem | Five whys ends at: speed of marketing launches was traded against edge safety without anyone deciding to make that trade. |
Metrics to Watch: config.rollout_stage_duration, config.auto_rollbacks, lb.cpu_utilization per cell, config.apply_latency_seconds
Organizational Follow-up: edge platform owns the pipeline with no exemption classes; marketing launches get a fast lane that still canaries (10 min instead of 60); the pipeline's error-budget consumption is reported monthly.
Ownership Question: "Who approved skipping canary for route changes?" Staff answer: Nobody explicitly — it was a pipeline change merged to fix a speed complaint. The edge platform owns that now: pipeline safety settings require the same review as the data plane.
Key Takeaway: "Any change that reaches every proxy at once is a global outage waiting for a trigger. There is no such thing as a safe-to-skip-canary config class."
What clears the Staff bar:
- Refuses "bad regex" as a root cause
- Fixes the pipeline, the engine and the rollback path
- Finds the organizational trade-off that created the exemption
Deep Dive 5: Multi-Region Expansion — Adding APAC#
Context: The company serves APAC users from a US region with ~180 ms RTT. Product wants a Singapore region in two quarters. Backends will be deployed there, but the primary database stays in the US for now. The edge team must decide how traffic reaches the new region and what happens on failure.
Questions to Surface First:
- Which requests can be served fully in-region (reads from replicas, static), and which must cross to the US (writes)?
- Anycast or GeoDNS? Do we have anycast address space and BGP presence in APAC?
- What happens to APAC users if the Singapore region fails — fail to US with +180 ms, or degrade?
- Capacity: can the US region absorb APAC traffic during a Singapore failure?
Typical L5 Approach: Adds a GeoDNS record for APAC pointing to Singapore and deploys the stack there.
Staff Approach: Starts with edge-only: terminate TLS in Singapore first, keep backends in the US, and send requests over warm, pooled connections across the backbone. That alone removes ~2 RTTs × 180 ms of handshake from every new APAC connection — most of the user-visible win — before any backend moves. Then moves read paths into the region, with writes routed to the US by path. Uses GeoDNS initially (no anycast space), with 60 s TTL and failover to the US edge; US capacity planned to absorb APAC at 30% headroom.
Principal Approach: Sequences the investment by user-visible latency per dollar: edge termination first (cheap, large win), read paths second, data residency and writes last (expensive, one-way doors). Ties the plan to the multi-region data strategy so the edge doesn't promise more locality than the data tier can deliver.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (planning) | Measure: APAC new-connection latency today ≈ 3 × 180 ms before first byte (TCP + TLS 1.2 in older clients). Target with local termination: ≈ 3 × 20 ms + one backbone round trip on a warm connection. |
| Triage (design) | Phase 1: Singapore edge (L4 + L7) with pooled connections to US backends. Phase 2: read-heavy routes served from Singapore backends + read replicas. Phase 3: write paths stay US until the data tier decides on regional writes. |
| Quick fix (rollout) | GeoDNS to Singapore for 5% of APAC resolvers, compare p50/p99 time-to-first-byte, ramp over 2 weeks. |
| Guardrails | Health-based DNS failover to the US edge; region.capacity_headroom alert if US can't absorb APAC; quarterly regional drain test. |
| Post-mortem (readiness) | Document per-route locality: which routes are served in-region, which forward to the US. |
Metrics to Watch: edge.ttfb_p50 by country, backbone.rtt_ms, dns.failover_events, region.capacity_headroom
Organizational Follow-up: edge team owns phase 1; service teams own deciding which routes are region-local; the data platform owns replica freshness guarantees that make read routing safe.
Ownership Question: "Who decides a route can be served from a read replica in Singapore?" Staff answer: The owning service team, against freshness requirements they sign off on. The edge only routes; it doesn't know whether stale data is acceptable.
Key Takeaway: "Terminating TLS near the user is the cheapest large latency win in a multi-region rollout. Move the handshake first, the reads second and the writes last."
What clears the Staff bar:
- Quantifies handshake RTT savings from edge termination alone
- Sequences the rollout by value and reversibility
- Leaves data-freshness decisions with the owning teams
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Distinguish edge, east-west and L4 load balancing and put each decision at the layer with the right information
- Locate latency correctly: sub-millisecond proxy hops vs round-trip-dominated TLS and TCP setup
- Size an edge fleet on both requests and new connections, including a zero-resumption scenario
- Explain why P2C least-request beats least-connections, and why it still needs outlier detection
- Design health checking that can't empty a pool: shallow active checks, passive ejection, caps and panic mode
- Treat config as a deploy with staged rollout, automatic rollback and data-plane independence from the control plane
- Handle long-lived connections with max age, GOAWAY, per-request balancing and negotiated reconnects
- Plan regional failover as a capacity question with explicit headroom
The Bar for This Question#
Mid-level (L4): Explains what a load balancer does, names round robin and least connections, and adds health checks and a redundant pair. Understands L4 vs L7 in textbook terms.
Senior (L5): Builds a sensible two-tier design with TLS termination, health checks, autoscaling and sticky sessions where needed. Knows the products. Misses how the balancer amplifies failures: herding, sinkholes, correlated health checks, config pushes.
Staff+ (L6): Starts from where latency and failure actually come from. Puts numbers on the TLS cost, chooses P2C with outlier detection and explains why, bounds every automated decision, ships config like code, and handles connection lifetime explicitly. Assigns ownership between the edge platform and service teams. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 The Load Balancer Hop Is Never Your Latency Problem#
| Claim | Evidence |
|---|---|
| A well-run L7 proxy adds ~0.1–1 ms | Request parsing and routing are microseconds; the in-AZ network hop is ~0.1–0.5 ms |
| Connection setup costs 1–3 RTTs | At 80–180 ms RTT, that's 160–540 ms for a new user |
| Removing a proxy hop rarely moves p99 | Tail latency comes from backends, queuing and retries |
The Staff position: Optimize connection reuse, TLS resumption and edge proximity before debating proxy hops.
Why this matters in interviews: Candidates who argue against an L7 tier "for latency" reveal they haven't measured where the time goes.
10.2 Least-Connections Should Be Retired as a Default#
| Claim | Evidence |
|---|---|
| It herds with many independent balancers | Every balancer sees a new host as least-loaded simultaneously |
| It rewards fast failure | A host returning instant 503s has the fewest active requests |
| Connection count isn't load under HTTP/2 | One connection may carry hundreds of concurrent streams |
The Staff position: P2C least-request plus outlier detection as the default; least-connections only for a single balancer with uniform, connection-per-request traffic.
Why this matters in interviews: It's the clearest single place where a Senior's reasonable answer has a known production failure mode.
10.3 Deep Health Checks Cause More Outages Than They Prevent#
| Claim | Evidence |
|---|---|
| Deep checks correlate failures | Every host calls the same dependency; a blip fails them all at once |
| Real traffic is a better correctness signal | Passive outlier detection sees the actual errors users see |
| The LB can't fix a dependency outage anyway | Ejecting every host replaces partial errors with total ones |
The Staff position: Health endpoints check the process, not its dependencies. Correctness comes from outlier detection; dependency problems surface as errors, honestly attributed.
Why this matters in interviews: Interviewers who've lived through a correlated ejection outage will recognize this immediately.
10.4 Sticky Sessions Are a State-Placement Bug#
| Claim | Evidence |
|---|---|
| Stickiness exists to keep server-local state | Which is lost when that server dies or deploys anyway |
| It worsens load balance | Long sessions pin load to old hosts after scale-out |
| It slows deploys | Draining waits for sessions to end, or breaks them |
The Staff position: Move session state to a shared store; use bounded-load consistent hashing only as a cache-affinity optimization, never for correctness.
Why this matters in interviews: Pushing back on a requirement and offering the real fix is a Staff behavior interviewers look for.
10.5 Your Edge's Biggest Risk Is Its Config Pipeline#
| Claim | Evidence |
|---|---|
| Data planes are mature and redundant | Stateless proxy fleets survive individual failures routinely |
| Config reaches every proxy | A bad change is correlated across the entire fleet by construction |
| Public post-mortems repeatedly feature global config pushes | Large edge and CDN outages are frequently triggered by a single config or rule change |
The Staff position: Invest in staged rollout, automatic rollback and CPU-cost testing of config before the next data-plane optimization.
Why this matters in interviews: It moves the conversation from boxes to change management — where real availability is won or lost.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
The Staff engineer designs a safe edge. The Principal engineer notices that a single request from a user's phone passes through five independently owned balancing decisions: GeoDNS run by the network team, a cloud L4 balancer, the edge proxy fleet, an API gateway owned by another platform team, and a sidecar mesh owned by a third. Each has its own health model, its own retry policy and its own config pipeline. A request can be retried at four layers; a backend can be ejected by three different health systems that disagree. The L7 problem is traffic-management governance: how many layers the org needs, which policies are set once, and who owns the end-to-end behavior of a request that no single team sees whole.
The Org-Level Fault Line#
One traffic platform vs layered teams vs per-service choice.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each layer owned independently | Teams move fast in their own layer | Retries multiply across layers; conflicting health models; five config pipelines to make safe | Users (amplified incidents); on-call (cross-team debugging) |
| One traffic platform owns edge, gateway and mesh | One policy model, one config pipeline, end-to-end visibility | Very large team; edge, gateway and mesh have different skills and cadences | Platform org (scope); services (one bottleneck for changes) |
| Separate data planes, one policy layer | Each team runs its data plane; timeouts, retries, outlier rules and config safety are defined once | Requires a shared policy schema and real enforcement | Platform architecture (the policy contract); teams (conformance) |
🧭 Principal Move: "I don't need one team running every proxy. I need one answer to 'how many times can a request be retried, by whom, and with what budget' — and one standard for how config reaches a proxy. The data planes can stay with the teams that know them; the policy and the change pipeline become shared."
Cost Model#
Assumptions: cloud-hosted, 3 AZs per region, fully loaded engineer ~$250K/year, average response 15 KB, ~40% of connections new at peak. Egress is billed separately and usually dwarfs balancer cost; it is excluded from infra but noted.
| Scale | Traffic | Infra ($/month) | Headcount | On-call Load | Dominant Cost Driver |
|---|---|---|---|---|---|
| Startup | ~2K RPS, 1 region | ~$500–2K (managed L7 balancer) | 0.25 eng (part of infra) | Shared rotation; balancer itself rarely pages | Managed per-hour + per-capacity-unit charges |
| Growth | ~50K RPS, 2 regions | ~$15–40K (managed or ~20 proxy instances + L4) | 2–4 eng edge team | Dedicated rotation, 2–4 pages/month, mostly config and certs | Handshake CPU at peak; cross-AZ traffic from poor zone routing |
| Enterprise | ~1M RPS, 3–5 regions, anycast | ~$150–400K (L4 + L7 fleets, dedicated pools, backbone) | 10–20 eng across edge, L4/network, config pipeline, certs | Follow-the-sun; per-region rotations | Egress and backbone; L7 fleet sized for zero-resumption peaks |
The pricing insight: at enterprise scale, the biggest controllable line items are not proxies — they're handshake headroom and cross-AZ transfer. Switching to ECDSA with high resumption can shrink the L7 fleet by 30–50%; capacity-weighted zone routing can cut cross-AZ transfer for the edge by half. Each is worth more than any proxy tuning, and each takes one or two engineers a quarter.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Anycast vs DNS-based global steering (and acquiring address space, BGP presence) | One-way-ish | Months: IP allocations, peering, client hardcoding of IPs |
| Public hostnames and IPs customers allowlist | One-way | Customer coordination; enterprise contracts list them |
| TLS termination point (edge vs passthrough to services) | One-way-ish | Cert ownership, compliance scope and service code all change |
| Session affinity promised to a product | One-way | Removing it requires moving state the product now depends on |
| Balancing algorithm default | Two-way | Config template change, staged |
| Outlier and panic thresholds | Two-way | Config, per cluster |
| Managed vs self-run L7 | Two-way (if config is abstracted) | 1–2 quarters of migration |
| Max connection age, drain windows | Two-way | Config |
The Standard I'd Write#
RFC-TRAFFIC-001: Traffic Management Standard
Status: Approved Owners: Edge Platform + Service Mesh Platform + SRE
Scope
Every balancing layer handling production requests: global steering,
L4, edge L7, API gateway, service mesh, and client libraries.
MUST
1. Default balancing is least-request with power-of-two choices. Hash-based
balancing MUST use bounded load (max 1.25x average).
2. Passive outlier detection is enabled with max ejection <= 20% of a pool;
panic threshold is 50% healthy.
3. Health endpoints used for load balancing MUST NOT call remote shared
dependencies.
4. Exactly one layer retries a given request. Retries apply only to
idempotent requests and are capped by a retry budget of 20% of traffic.
5. Every config change to any balancing layer rolls out in stages
(canary, zone, region, global) with automated rollback. No exempt classes.
6. Data planes continue serving on last-known-good config when the control
plane is unavailable.
7. Edge certificates are provisioned by the platform with ECDSA and RSA
fallback; ticket keys rotate with overlap.
SHOULD
1. Cap client connection age (10 min +/- jitter) via GOAWAY.
2. Use capacity-weighted zone-aware routing.
3. Use dedicated pools for long-lived connection workloads above 50K connections.
Exceptions
Filed with Edge Platform; reviewed within 5 business days; time-boxed to
2 quarters; exceptions to MUST 4 and MUST 5 need SRE sign-off.
Success metrics
- Incidents where a balancing layer amplified a smaller failure: trending to 0
- Config-change-caused edge minutes of impact: < 25% of edge error budget
- Edge-added latency p99 (excluding client RTT): < 5 ms
- TLS resumption rate at peak: > 50%
- Services with more than one retrying layer: 0
What I'd Tell the VP#
"Our biggest outages at the edge haven't come from machines failing — they've come from our own traffic systems overreacting: a config change that reached every server at once, health checks that removed healthy servers, retries piling up across layers. I'm proposing one shared standard for how all of our traffic layers behave and one safe path for changing them. It's about three engineers for two quarters. We expect it to remove our most common class of large edge incident, and the TLS and zone-routing changes that come with it should reduce edge infrastructure cost by roughly a third. Teams keep ownership of their own systems; what changes is that they follow the same rules."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Sees the whole request path | "This request crosses five balancing layers owned by four teams. Which of them is allowed to retry?" |
| Prices the levers | "ECDSA and resumption save a third of the fleet; that's worth more than any proxy we could tune." |
| Names one-way doors | "Customer-allowlisted IPs and anycast space are the decisions I'd slow down on." |
| Sets org failure posture | "No config class skips canary — and edge config changes freeze when they've eaten a quarter of the error budget." |
| Knows when not to centralize | "Mesh and edge keep separate data planes and teams; only the policy and pipeline are shared." |
Staff answers that L7 interviewers find insufficient:
- "I'd configure outlier detection and panic thresholds on the edge" — correct for one layer, but ignores the gateway and mesh making contradictory decisions on the same requests.
- "Retries happen at the edge with a budget" — doesn't ask whether the mesh and client libraries also retry, multiplying the budget.
- "We'll add a region for resilience" — no capacity headroom priced, no failover decision owner, no sequencing by cost.
Appendices
Appendix A: Mechanics in Depth#
A.1 Power of Two Choices with Least Request#
pick_host(cluster):
hosts = cluster.healthy_hosts()
if len(hosts) < cluster.panic_threshold * len(cluster.all_hosts()):
hosts = cluster.all_hosts() # panic: ignore health
a, b = random_sample(hosts, 2)
score(h) = (h.active_requests + 1) / h.effective_weight()
return a if score(a) <= score(b) else b
effective_weight(h):
w = h.configured_weight
if h.age_seconds < slow_start_window: # e.g. 60 s
w = w * max(0.1, h.age_seconds / slow_start_window)
return w
Why it's right: each proxy needs only its own in-flight counts — no coordination. Randomly sampling two hosts gives near-optimal balance even when counts are stale or partial, and different proxies rarely choose the same host at the same instant, so there's no herd. Why it can go wrong: a host that completes requests instantly (because it fails) always has low in-flight counts. Outlier detection is the necessary partner.
A.2 Outlier Detection#
on_response(host, status, latency):
if status >= 500 or timed_out:
host.consecutive_errors += 1
else:
host.consecutive_errors = 0
if host.consecutive_errors >= 5 and can_eject(cluster):
eject(host, duration = base_ejection * host.times_ejected) # 30 s, 60 s, 90 s...
every 10 s: # success-rate outliers, only if pool has >= 5 hosts with >= 100 requests
mean, stdev = success_rates(cluster)
for host where host.success_rate < mean - 1.9 * stdev:
if can_eject(cluster): eject(host)
can_eject(cluster):
return ejected_count(cluster) < max(1, 0.10 * len(cluster.all_hosts()))
The relative check (vs peers) is what makes outlier detection safe under systemic failure: if every host's success rate drops together, none is an outlier.
A.3 Consistent Hashing Variants#
| Variant | Mechanism | Disruption on Change | Balance | Use |
|---|---|---|---|---|
| Ring hash | Hosts placed at many virtual points on a ring; key goes to next point clockwise | ~1/N keys move | Needs ~100+ virtual nodes per host for evenness | General affinity |
| Maglev | Fixed lookup table (prime size, e.g. 65,537) filled via per-host permutations | Slightly more than minimal; very fast lookup | Very even | L4 flow hashing |
| Rendezvous (HRW) | Score every host with hash(key, host); pick max | Minimal | Even; O(N) per lookup | Small pools, simplicity |
| Bounded load | Any of the above, plus: skip hosts above c × average load | Some affinity lost under skew | Bounded by c (e.g. 1.25) | Hot keys |
See Consistent Hashing for the underlying theory.
Appendix B: Routing Keys and Identity#
| Key | Use | Pitfall |
|---|---|---|
| 5-tuple (src IP, src port, dst IP, dst port, proto) | L4 flow hashing | Changes on client reconnect; QUIC migrates connections across addresses — use QUIC connection IDs |
| Source IP | Legacy affinity | NAT concentrates thousands of users on one IP |
| Cookie (LB-issued) | Session affinity at L7 | Pins load; survives only while host lives |
Header (user_id, tenant_id) | Cache affinity, tenant pools | Hot tenants; requires trusted header injection |
| Path / host | Service routing | Regex cost; overlapping routes need explicit precedence |
Precedence rule: the most specific route wins (exact host > wildcard host; longest path prefix), and the config validator rejects ambiguous overlaps at commit time rather than letting proxies decide at runtime.
Appendix C: Coordination Mechanisms#
C.1 Endpoint Propagation#
Service discovery publishes endpoint sets; the control plane pushes incremental updates (deltas, not full snapshots) to proxies over a streaming API (e.g., xDS). Requirements: updates propagate in < 5 s at p99; proxies ACK/NACK each version so the control plane knows which proxies run which config; the control plane batches churn (a 400-host deploy shouldn't produce 400 pushes per proxy). See Service Discovery.
C.2 Quick Comparison#
| Mechanism | Protects Against | Doesn't Protect Against | Cost |
|---|---|---|---|
| Consistent hashing at L4 | Flow resets on L4/L7 fleet changes | Resets when the target proxy itself dies | Lookup table per node |
| Connection tracking | Flow remapping during backend set changes | Tracking-table loss on L4 node failure | Memory per flow |
| Outlier detection | Individual bad hosts | Systemic failures (by design) | Per-host counters |
| Panic threshold | Health signals that falsely empty a pool | Real total outages | None |
| Slow start | Cold hosts overloaded on join | Hosts that are slow forever | Longer scale-out ramp |
| Retry budget | Load amplification from retries | Non-idempotent retry hazards (handled by idempotency) | Some requests not retried |
| Staged config rollout | Global bad config | Bad config that only fails under peak load | Slower changes (10–60 min) |
Appendix D: API Contract & Client Behavior#
| Behavior | Rule |
|---|---|
| Keepalive | Clients reuse connections; edge idle timeout ~60–120 s; backend pools idle timeout shorter than backend's own |
| Max connection age | 10 min ± jitter via GOAWAY (HTTP/2) or Connection: close (HTTP/1.1) |
| Retries | Idempotent methods only (or requests with an idempotency key); one retry at the edge; retry budget 20% |
Retry-After / 503 | Edge honors backend overload signals and does not retry those requests |
| 0-RTT | Accepted only for safe methods (GET/HEAD) to replay-tolerant routes |
| WebSocket drain | Server sends reconnect message; client reconnects with 0–300 s jitter and exponential backoff on failure |
| Client timeouts | Mobile SDK: connect 10 s, request 30 s, exponential backoff with jitter on handshake failures |
Appendix E: Observability#
Core metrics:
| Metric | Why It Matters | Alert |
|---|---|---|
lb.added_latency_p99 | Edge's own contribution, excluding upstream | > 5 ms for 5 min |
lb.upstream_5xx_rate (per cluster, per host) | Backend health as users see it | Per-service SLO burn |
lb.no_healthy_upstream | Pool emptied — LB-caused outage | Any non-zero |
lb.hosts_ejected_pct | Ejection pressure | > 10% for 5 min |
lb.healthy_hosts_pct (fleet-wide correlation) | Correlated health failures | Many clusters dropping together |
tls.resumption_rate | Handshake cost driver | < 50% of baseline |
tls.full_handshakes_per_sec | Capacity driver | > 2× baseline |
lb.retry_rate | Amplification | > 20% of requests |
config.version_skew | Proxies on different config versions | > 1 version for > 15 min outside rollouts |
xds.config_age_seconds | Control-plane staleness | > 5 min |
Control plane vs data plane. The data plane must keep routing on last-known-good config when the control plane is down; endpoints going stale is tolerable for minutes because outlier detection ejects dead hosts from real traffic. Test it: kill the control plane in staging during a deploy.
Debugging the silent failure. Aggregate dashboards hide per-zone and per-host skew. Always keep: per-host RPS distribution (max/mean), per-zone utilization, and requests by config version. "Which config version served this request?" should be answerable from access logs.
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 5K RPS | One managed L7 balancer, DNS | Routing flexibility; per-service policies |
| 5K–100K RPS | Managed L4 + self-run L7 fleet with templated policy | Handshake CPU; config safety; zone imbalance |
| 100K–1M RPS | Own L4 (XDP, consistent hash, DSR), anycast, dedicated pools | Multi-region capacity; cross-layer retries |
| 1M+ RPS | Global steering on capacity, one policy layer across edge, gateway and mesh | Org coordination more than technology |
Multi-region path: GeoDNS with health-based failover → edge termination in more regions with backbone to central backends → regional backends for reads → anycast with capacity-aware steering.
What You Don't Build on Day One: your own L4 data plane; anycast; a custom proxy; per-request adaptive algorithms beyond P2C; cross-region active-active steering. Each is justified by a measured problem, not anticipated scale.
Appendix G: Multi-Tenancy, Fairness & Cost#
Per-service limits on shared proxies (prevent one backend from exhausting the proxy):
| Limit | Typical Default | Effect When Hit |
|---|---|---|
| Max connections to cluster | 1,024 per proxy | New requests queue or fail fast |
| Max pending requests | 1,024 | Fast 503 instead of unbounded queueing |
| Max concurrent requests | 1,024 (HTTP/2) | Same |
| Max concurrent retries | 3 per proxy, or budget-based | Retries dropped |
| Per-route timeout ceiling | ≤ listener max (e.g. 60 s) | Config rejected |
Cost attribution. Charge services for what they consume at the edge: full handshakes (CPU), requests, egress bytes and long-lived connection-hours. The first report usually shows one or two services dominating handshake cost — typically mobile clients with short keepalives — and fixing their client config pays for the attribution work.