Hiring BarSupport

Design a Load Balancer — Staff-Level Case Study

Case study86 min read6 diagrams

Technologies referenced in this case study: API Gateways · ZooKeeper & etcd · Redis

Related: API Gateway · Service Discovery · CDN & Edge Caching · Circuit Breakers · Rate Limiting · Multi-Region · Real-time WebSockets · Degraded Mode · Consistent Hashing · Networking · Numbers to Know

How to Use This Case Study#

This case study is organized for the interview first and for reference second. Read it front-to-back once; then return to the fault lines and deep dives that match your weak spots. If you need the packet-level background — TCP, TLS, HTTP/2, anycast — start with Networking.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Deep Dives 1 and 4
Deep Dive3+ hrsEverything, including Section 11 (Principal Lens) and Appendix A (balancing algorithms in depth)
What is a Load Balancer? — Why interviewers pick this topic

A load balancer accepts traffic addressed to one name or IP and spreads it across many backends, so that the service survives the loss of any backend and its capacity can grow by adding more. In practice it does more: terminates TLS, routes by host and path, retries failed requests, drains hosts during deploys, and decides — continuously, without a human — which backends are healthy enough to receive traffic.

That last job is the dangerous one. A load balancer is the component in the request path with the most authority and the least context. It can't see that a backend's database is slow; it sees that requests are slow. It can't see that a "healthy" backend returns errors in 2 ms; it sees a backend that finishes requests very quickly. Every load-balancing decision is made on proxies for health, and the proxies lie in predictable ways.

Before vs After — the "one bad host" scenario:

Without a designed balancing and health policy:
t=0:        Backend b-17 loses its database connection pool. Every request now
            returns 503 in 2 ms instead of 200 in 40 ms.
t=+1s:      Least-connections LB sees b-17 always has 0 active requests.
            Routes more and more traffic to it.
t=+10s:     b-17 receives ~35% of all requests across a 40-host pool. 35% error rate.
t=+30s:     Active health check (GET /health every 10s, needs 3 failures) still
            passes: /health doesn't touch the database.
t=+8min:    Someone notices the error graph, finds b-17, terminates it by hand.

With passive outlier detection and P2C least-request:
t=0:        b-17 starts returning 503 in 2 ms.
t=+0.5s:    5 consecutive 5xx → b-17 ejected for 30s (cap: 10% of pool ejectable).
t=+0.5s:    Error rate peak: ~0.6% for half a second. Traffic spread over 39 hosts.
t=+30s:     b-17 re-admitted on probation, fails again, ejected for 60s.
t=+2min:    Alert: host b-17 ejected 3 times → owner's auto-remediation replaces it.

Why interviewers reach for this question: Load balancing looks like a solved problem — "put an ALB in front" — which makes it an excellent test of depth. The candidate who stops at round-robin and health checks has described a working system on a good day. The interviewer wants to see whether you know how load balancers amplify failures: herding traffic onto broken hosts, ejecting healthy hosts during overload, breaking every connection on a config push, or turning a TLS certificate rotation into a CPU outage.

Mechanics Refresher: Layers and Algorithms
MechanismHow It WorksProsCons
DNS round-robin / GSLBReturn different IPs per resolver; weight by region or healthGlobal steering; no data-path boxTTLs ignored by many clients; minutes to drain; no per-request control
AnycastSame IP announced from many sites; BGP picks the nearestInstant global spread; absorbs DDoSRoute changes can move flows mid-connection; coarse control
ECMP at the routerRouter hashes 5-tuple across equal-cost next hopsLine-rate, no extra boxNext-hop set change rehashes flows unless backed by consistent hashing
L4 load balancerForwards packets/connections by 5-tuple; no payload inspectionMillions of packets/s per machine; protocol-agnosticNo HTTP awareness: can't route by path, retry requests or see 5xx
L7 proxyTerminates the client connection, parses HTTP, opens/reuses backend connectionsPer-request routing, retries, outlier detection, observabilityTLS and parsing CPU; adds ~0.1–1 ms; must scale with requests, not packets
Client-side / sidecar balancingCaller (or its local proxy) picks the backend from a discovered listNo central hop; per-request decisions with local latency dataEvery client needs the logic; control plane fan-out to thousands of clients
AlgorithmHow It WorksProsCons
Round robinNext backend in orderSimple, even on uniform requestsIgnores backend load and request cost variance
Least connectionsBackend with fewest open connectionsAdapts to slow backendsWith HTTP/2 or keepalive, connections ≠ load; herds onto fast-failing hosts
Power of two choices (P2C) + least requestPick 2 backends at random, send to the one with fewer in-flight requestsNear-optimal spread with O(1) work and stale data; avoids herdingStill fooled by fast-failing hosts without outlier detection
Consistent hashing (ring / Maglev / rendezvous)Hash a key to a backend; minimal remapping on changeAffinity (caches, sessions); stable L4 flowsUneven load on hot keys; needs bounded-load variants
Weighted / latency-aware (EWMA)Score backends by latency × in-flightRoutes around slow hostsFeedback oscillation if not damped

For most production systems: anycast or DNS to a region, ECMP to an L4 tier using consistent hashing, an L7 proxy fleet terminating TLS, and P2C least-request with passive outlier detection toward backends. The algorithm is not the interview — health judgment and failure amplification are.


Executive Summary

If you only read one section, read this. Everything in this case study flows from the contrast below.

What This Interview Actually Tests#

A load balancer is not an algorithm question. Round-robin works fine on most days.

It is a failure-judgment question: a load balancer decides, many times per second and without a human, which machines deserve traffic — and the most common way it fails is by being confidently wrong at scale. It tests:

  • Whether you know what each layer (DNS, anycast, L4, L7, client-side) can and can't see, and put each decision at the layer that has the information
  • Whether you know where the latency actually goes — the proxy hop is sub-millisecond; the TLS handshake and connection setup are not
  • Whether your health model distinguishes "dead" from "slow" from "fast-failing", and refuses to eject half the fleet during overload
  • Whether you treat the load balancer's own config and control plane as the largest blast radius in the system
  • Whether connection lifetime — keepalive, HTTP/2, WebSockets — is part of your balancing model

The key insight: The load balancer sits on every request, so its mistakes are correlated by construction. A backend bug hurts one host's share of traffic; a load-balancer bug — a bad health check, a herding algorithm, a config push — hurts all of it at once. Staff candidates design the LB to fail less confidently: bounded ejection, panic thresholds, staged config rollout, and draining that respects connection lifetime.

The L5 vs L6 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws clients → LB → backends, picks round-robin or least-connectionsAsks "Internet edge or service-to-service? HTTP or raw TCP? Long-lived connections? What's the latency budget?"Asks "How many balancing layers does the org already run, who owns each, and which one is the single config push that can take everything down?"
Latency"LBs add latency, so we keep it to one hop""The L7 hop is ~0.1–1 ms. The cost is TLS: a full handshake is 1–2 RTTs plus CPU, so I'd maximize resumption and connection reuse"Prices handshake CPU across the fleet — ECDSA certs and resumption can cut L7 fleet size by 30–50% — and sets the TLS policy org-wide
AlgorithmLeast-connectionsP2C least-request with outlier detection and slow start; consistent hashing only where affinity paysStandardizes the algorithm in the shared proxy config so 200 teams don't each rediscover the fast-fail sinkhole
HealthActive /health check every 10sActive checks for liveness + passive outlier detection for correctness; max 10–20% ejectable; panic mode below 50% healthyDefines what /health may and may not check org-wide — shallow liveness only — so a shared dependency blip can't fail every service's checks at once
Failure"Run two LBs, active-passive""N+1 stateless proxies behind ECMP; consistent hashing so a proxy loss moves only its flows; config changes staged by canary"Designs the config-push failure posture: cell-by-cell rollout, automatic rollback on error-rate delta, a last-known-good config that can be restored in < 5 min globally
OwnershipInfra team owns the LBEdge team owns the L4/L7 fleet and its config pipeline; service teams own their health endpoints, timeouts and retry budgetsDraws the contract: platform owns the data plane and policy guardrails, services own their routes — with validation that rejects unsafe routes before they ship
Why "algorithm" separates levels

L5: "Least-connections, because it adapts to slow backends." This is a good instinct and correct in a single-LB, HTTP/1.1 world. It breaks in two common production situations. First, with many LB instances each seeing only its own connections, a newly added backend looks empty to all of them simultaneously, and they all herd onto it. Second, a backend that fails fast — returning 503 in 2 ms — always has the fewest active requests, so least-connections sends it the most traffic.

L6: "P2C least-request: each proxy picks two backends at random and sends to the one with fewer in-flight requests. Randomness breaks the herd across many proxies, and it needs no global state. On its own it still rewards fast failure, so I pair it with passive outlier detection — five consecutive 5xx ejects the host for 30 seconds — and slow start, so a new host ramps over 30–60 seconds instead of taking a full share cold."

L7: "The algorithm lives in one shared proxy config template, not in 200 service configs. Teams can choose consistent hashing when they need affinity, through a reviewed option — but the default is P2C, outlier detection on, max ejection 10%. Most balancing incidents I've seen were teams who overrode a good default without knowing why it existed."

Why "health" separates levels

L5: "Health check /health every 10 seconds; remove a host after 3 failures." The check usually tests that the process is up — and sometimes, overzealously, tests its database too. Neither tells you whether real requests are succeeding.

L6: "Two signals with two jobs. The active check is shallow — is the process alive and accepting connections — and is how new hosts join. Passive outlier detection watches real traffic: consecutive 5xx, error rate relative to peers, latency outliers. And there's a safety rail: if more than half the pool looks unhealthy, the LB stops trusting its health data and balances across everything, because it's far more likely the checks are wrong than that 50% of hosts died at once."

L7: "The worst health-check outage is the correlated one: every service's /health calls the same auth service, auth has a 20-second blip, and every LB in the company ejects every backend. I'd write the rule that health endpoints MUST NOT call shared dependencies, and enforce it with a lint on the health handler and a quarterly game day that kills a shared dependency."

Why "failure" separates levels

L5: "Two load balancers in active-passive with a floating IP." That protects against one machine dying. It doesn't address the more common outage: the LB fleet is fine, but a configuration change breaks routing on every instance simultaneously.

L6: "The data plane is a fleet of stateless proxies behind ECMP; losing one proxy moves ~1/N of flows, and consistent hashing at L4 means only those flows move. The bigger risk is config: route tables, TLS certs, health policies. Every config change goes canary → one AZ → one region → global, with automatic rollback if the 5xx rate rises more than 0.5 percentage points over baseline."

L7: "The LB config pipeline is the highest-blast-radius deploy system in the company. I'd give it the strictest change management we have: staged rollout enforced by tooling, not process; a global kill-switch to last-known-good; and edge config changes frozen during peak events with a VP-level exception path."

The Staff Positions#

PositionRationale
Two tiers: L4 (consistent hashing) in front of L7 (stateless proxies)L4 scales packets cheaply and keeps flows stable; L7 does per-request intelligence. Neither can do the other's job well
P2C least-request as the default algorithmNear-optimal spread, O(1), robust to many proxies with partial views; least-connections herds
Passive outlier detection with a max-ejection cap (10–20%)Real traffic is the only honest health signal; the cap stops the LB from ejecting its way into an outage
Panic mode below ~50% healthyWhen most hosts look unhealthy, the checks are wrong more often than the hosts are
Shallow active health checksLiveness only; deep checks turn every shared-dependency blip into a fleet-wide ejection
TLS terminated at the edge with resumption and ECDSA certsFull handshakes dominate proxy CPU and user latency; resumption makes repeat visits ~1 RTT cheaper
Config changes are deploys: staged, canaried, auto-rolled-backLB config is the single change that can break every request at once
Retries at one layer only, with a retry budget (~10–20% of requests)Retries at every layer multiply load 3–5× during an incident

The Three Intents#

Three intents produce three different load balancers. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Internet edge (north-south) for a latency-sensitive HTTP servicep99 added latency < 5 ms; TLS for millions of clients; DDoS absorptionAnycast/DNS → L4 (ECMP + consistent hash) → L7 proxies terminating TLS → backends with P2CTLS CPU exhaustion; config push outage; edge capacity loss in one regionAvailability 99.99%; added p50 < 1 ms beyond TLS; no connection resets on deploys
Service-to-service (east-west) inside the data centerSub-millisecond overhead; thousands of services; per-request retries and timeoutsClient-side balancing via sidecar or library, fed by service discoveryControl-plane fan-out; inconsistent client versions; retry stormsZero extra hops; endpoint updates propagated < 5 s
L4 network load balancer for non-HTTP or very high packet ratesMillions of packets/s; preserve client IP; any TCP/UDP protocolKernel-bypass or XDP forwarding, consistent hashing, connection tracking, direct server returnRehash on fleet change breaks flows; SYN floods; asymmetric routingExisting connections survive LB fleet changes; line-rate forwarding

🎯 Staff Move: "I'll design the internet edge for a latency-sensitive HTTP API — that's where TLS cost, L4/L7 layering and health judgment all matter. I'll sketch the L4 tier but not redesign packet forwarding, and I'll note where east-west balancing differs: it moves into the client, because a central hop for every internal call is a tax we don't need to pay."

The Five Fault Lines#

#Fault LineThe Tension
1Where to Terminate: L4 vs L7Per-request intelligence and TLS termination at L7, or line-rate, protocol-agnostic forwarding at L4 — and who pays the CPU?
2Even Spread vs AffinityP2C least-request spreads load best; consistent hashing keeps caches and sessions warm. You can't fully have both
3Trusting Health Signals vs Ignoring ThemEject fast to protect users from bad hosts, or eject cautiously so a bad signal can't empty the pool?
4Central Proxy vs Client-Side BalancingOne fleet that's easy to operate, or balancing logic in every client that removes a hop but multiplies the control plane?
5Connection Longevity vs RebalancingLong-lived connections save handshakes and latency, but pin load to old backends and make draining slow

In the Wild: Real Production Systems#

Why this section belongs here: Naming how real edge networks settled these fault lines shows you've studied operations, not just cloud console options.

Google — Maglev#

Google's Maglev paper (NSDI 2016) describes a software L4 load balancer running on commodity servers. Routers spread packets across Maglev machines with ECMP; each machine uses Maglev hashing — a consistent hashing scheme built on a fixed-size lookup table populated by per-backend permutations — plus a local connection-tracking table, so packets of a flow reach the same backend even when the Maglev fleet or backend set changes. The fleet is active-active, scaling out instead of using active-passive pairs.

Staff insight: Maglev's design point is minimal disruption under change: when backends or balancers come and go, almost all existing flows keep their mapping. In an interview, "consistent hashing plus connection tracking at L4" answers the question "what happens to in-flight connections when you add a load balancer?"

Meta — Katran#

Meta open-sourced Katran in 2018, an L4 load balancer built on XDP and eBPF that processes packets in the kernel's earliest receive path. It sits behind ECMP, uses a Maglev-style consistent hash with a connection table, and encapsulates packets to backends, which reply directly to clients (direct server return) so return traffic bypasses the balancer.

Staff insight: Direct server return matters because responses are usually much larger than requests. If return traffic doesn't flow through the L4 tier, that tier is sized for inbound packets only — often a 5–10× smaller fleet. Mention it when sizing the L4 tier.

Lyft — Envoy#

Envoy was built at Lyft and open-sourced in 2016 as an L7 proxy used both at the edge and as a sidecar next to every service. Its load balancer offers round robin, least-request using power-of-two-choices by default, ring hash and Maglev for affinity, passive outlier detection, a panic threshold (by default, if fewer than 50% of hosts are healthy, it balances across all hosts), and dynamic configuration over the xDS APIs from a control plane.

Staff insight: Envoy's defaults encode the hard-won lessons in this case study — P2C over least-connections, outlier detection with a capped ejection percentage, panic mode. Citing why those defaults exist is far stronger than citing that the product exists.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Least connections""You have 50 LB instances and just added 10 backends. What happens in the next second?"Herding with partial views; P2C, slow start
"Health checks every 10 seconds""A backend returns 503 in 2 ms on every request. When does it stop getting traffic?"Passive outlier detection; fast-fail sinkhole
"We eject unhealthy hosts""The database is slow, and 70% of hosts now fail checks. What does the LB do?"Max ejection, panic threshold
"We terminate TLS at the LB""Traffic just tripled from new users. What saturates first?"Handshake CPU, resumption, cert type
"Active-passive pair""The config push had a typo. Both are running it."Config as the real blast radius
"We use WebSockets / gRPC""You scaled from 20 to 40 backends. How long until the new ones get load?"Connection longevity vs rebalancing
"We retry failed requests""Every layer retries 3 times. One backend tier is down. What's the multiplier?"Retry budgets, single retry layer

System Architecture Overview#

Diagram: System Architecture Overview

Quick-Reference: The 30-Second Cheat Sheet#

QuestionStaff Answer
Layers?GeoDNS/anycast → L4 (ECMP + consistent hash) → L7 proxies → backends
What does the L7 hop cost?~0.1–1 ms in-region; TLS handshakes are the real cost (1–2 RTT + CPU)
Algorithm?P2C least-request; consistent hashing (bounded load) only where affinity pays
Health?Shallow active check for liveness; passive outlier detection for correctness
Over-ejection?Cap ejection at 10–20% of pool; panic mode below 50% healthy
New host?Slow start over 30–60 s
Deploy draining?Stop new requests, finish in-flight, GOAWAY on HTTP/2, drain window ≈ p99 request time (seconds), longer for WebSockets
Retries?One layer, idempotent requests only, budget ~10–20% of traffic
Config?Validated, canaried per cell, auto-rollback, last-known-good restore < 5 min
East-west?Client-side/sidecar balancing; no central hop

Key Numbers Worth Memorizing#

NumberValueContext
L7 proxy added latency~0.1–1 ms p50 in-regionParsing + routing + one local hop; tail grows under CPU pressure
L4 forwarding added latencyTens of microsecondsKernel-bypass/XDP forwarding
TLS 1.3 full handshake1 RTT before first request byteTLS 1.2: 2 RTTs; 0-RTT resumption possible for replay-safe requests
Handshake CPU (server signature)RSA-2048: ~1–3K signs/s/core; ECDSA P-256: ~10–40K/s/coreWhy ECDSA certs shrink edge fleets
Cross-region RTT~60–150 msWhy a user far from the edge pays 100+ ms per handshake RTT
Simple HTTP proxying~10–30K req/s per coreVaries with TLS, header size, filters
Concurrent connections per proxy100K+ idle keepalive connectionsMemory-bound: ~10–50 KB per connection with TLS buffers
Envoy panic threshold default50% healthyBelow it, balance across all hosts
Outlier ejection defaults (Envoy)5 consecutive 5xx, 30 s base ejection, 10% max ejectedBase time multiplies on repeat ejections
ALB health check defaults30 s interval, 2 failures unhealthy, 5 successes healthyUp to ~60 s to detect a dead host with active checks alone
ALB deregistration delay default300 sOften far too long for short requests, too short for WebSockets
DNS TTL for steering30–60 sMany clients and resolvers hold records longer; plan for minutes
P2C max load~log log n above average vs ~log n / log log n for randomTwo choices capture most of the benefit of perfect information
Retry budget~10–20% of requestsBounds load amplification during backend incidents

Interview Walkthrough

The walkthrough below is a 45-minute script. The words in italics are meant to be said out loud.

Phase 1: Requirements & Framing (2–3 minutes)#

"Load balancer can mean three different things: the internet edge in front of a public service, balancing between internal services, or a pure L4 packet balancer. I'll design the internet edge for a latency-sensitive HTTP API, where TLS, L4/L7 layering and health judgment all matter. I'll mention how east-west differs at the end."

QuestionAssumed AnswerWhy It Matters
Peak traffic?1M requests/s globally, ~400K/s in the largest regionSizes the L7 fleet
New connections per second?~150K/s at peak (mobile clients reconnect often)TLS handshake CPU is the binding constraint
Latency budget for the LB layers?Added p99 < 5 ms in-region, excluding TLS RTTsRules out extra hops; pushes for connection reuse
Protocols?HTTP/1.1, HTTP/2, some WebSockets, gRPC from partnersLong-lived connections change balancing and draining
Regions?3 regions, 3 AZs eachGlobal steering and zone-aware routing
Backends?~60 services behind the edge, 10–400 hosts eachRouting table size, per-service policy
Availability target?99.99% for the edge~4.3 min/month — config mistakes alone can burn it
Session affinity?Not required by most services; one cache-heavy service benefitsAffinity is opt-in per route

"99.99% gives the edge about four minutes a month. A single bad global config push can eat that in one go, so I'll treat config safety as a first-class part of the design, not an operations footnote."

Phase 2: Core Entities & API (1–2 minutes)#

EntityKey FieldsNotes
ListenerVIP, port, protocol, TLS policy, cert refsOne per public endpoint
Routehost, path prefix, headers → cluster, timeout, retry policyOwned by service teams, validated by the platform
Cluster (pool)name, endpoints, LB policy, health policy, outlier policy, circuit-breaker limitsPolicy defaults from the platform template
EndpointIP:port, zone, weight, health status, drain stateFed by service discovery
Config versionversion, diff, rollout stage, ownerEvery change is versioned and attributable
# Control-plane API (service teams own routes; platform owns listeners and templates)
PUT  /v1/routes/{service}            { host, path_prefix, cluster, timeout_ms, retry: {...} }
PUT  /v1/clusters/{service}          { lb_policy: "p2c_least_request" | "ring_hash", hash_on?, ... }
POST /v1/endpoints/{service}/drain   { endpoint, drain_seconds }
GET  /v1/config/versions/{v}/status  → { stage: "canary" | "zone" | "region" | "global", health_delta }
POST /v1/config/rollback             { to_version }        # platform on-call only

"Routes are owned by service teams; listeners, TLS policy and the defaults template are owned by the edge platform. That split matters later: it's what lets a team ship a route change without being able to change the timeout behavior for everyone else."

Phase 3: High-Level Architecture (≤5 minutes)#

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

"GeoDNS or anycast picks a region. In the region, routers ECMP across an L4 tier that uses Maglev-style consistent hashing with connection tracking, so flows stay stable when L4 nodes or L7 proxies change. The L7 fleet is stateless, terminates TLS, routes by host and path, and balances to backends with P2C least-request and outlier detection. The control plane pushes endpoints and config incrementally. I'll spend the depth on three places: where the latency goes, how health decisions fail, and how config changes fail."

Phase 4: Transition to Depth (1 minute)#

"The boxes are well known. Where this design succeeds or fails is the TLS cost at the edge, the health and balancing policy toward backends — because that's where load balancers amplify failures — and the config pipeline, because it's the one change that touches every request. Which would you like first?"

Phase 5: Deep Dives (25–30 minutes)#

Deep dive 1 — Where the latency goes (7 min).

Diagram: Phase 5: Deep Dives (25–30 minutes)

"For a new client 80 ms away, two round trips — TCP and TLS 1.3 — cost 160 ms before the first byte of the request arrives. The proxy's own processing is well under a millisecond. So the levers are: terminate TLS as close to the user as possible, maximize connection reuse with HTTP/2 and long keepalives, and support session resumption so returning clients skip the certificate exchange. On the server side, a full handshake with RSA-2048 costs roughly ten times the CPU of ECDSA P-256; serving ECDSA certs to clients that support them is the biggest single lever on edge fleet size."

"Toward backends, the proxy keeps warm connection pools, so the backend hop is ~0.1–0.5 ms in-AZ with no handshake. If the backends are in another AZ, add ~1–2 ms round-trip — which is why I'd keep zone-aware routing on and only spill across zones when local capacity is short."

Deep dive 2 — Balancing and health (12 min).

Diagram: Phase 5: Deep Dives (25–30 minutes)

"Balancing is P2C least-request: pick two healthy hosts at random, send to the one with fewer in-flight requests. With 50 proxies each choosing independently, randomness prevents them from herding onto the same host. New hosts get a slow-start weight that ramps over 60 seconds, so they warm caches and JIT before taking a full share."

"Health has two signals. Active checks every 5 seconds hit a shallow /healthz — the process is up and accepting — and decide admission. Passive outlier detection watches real responses: five consecutive 5xx or a success rate more than ~2 standard deviations below the pool ejects the host for 30 seconds, doubling on repeat. Two guards stop the LB from making things worse: no more than 10% of a pool can be ejected by outlier detection, and if fewer than 50% of hosts are healthy, we enter panic mode and spread across all of them. When half the fleet looks sick at once, the signal is wrong far more often than the fleet is."

Deep dive 3 — Config safety (6 min).

"Config changes — routes, clusters, TLS policies, certs — flow through one pipeline: schema validation, semantic checks (does every route point to an existing cluster? does any timeout exceed the listener's?), then a staged rollout: one canary proxy per region, then one AZ, then one region, then global, with a bake time of 5–10 minutes and automatic rollback if lb.upstream_5xx_rate or lb.added_latency_p99 worsens beyond a threshold. Proxies keep the last-known-good config and refuse a config that fails to load, rather than crash. The data plane must keep serving if the control plane is down."

Deep dive 4 — Draining and long-lived connections (3 min).

"On deploy, a backend is marked draining: no new requests, in-flight finish, HTTP/2 connections get GOAWAY so clients move. For short APIs the drain window is ~2× p99 request time — 10–30 seconds, not the 300-second default many cloud balancers ship with. WebSockets are different: they need an application-level 'reconnect' message and jittered reconnection across 5–10 minutes, or the drain becomes a reconnect storm."

Phase 6: Wrap-Up (2–3 minutes)#

"Summary: two tiers — L4 with consistent hashing for stable flows, L7 for TLS and per-request decisions. P2C least-request with outlier detection, capped ejection and panic mode, so the balancer can't amplify a bad signal into an outage. TLS cost managed with ECDSA and resumption. Config shipped like code with staged rollout and automatic rollback. Next I'd build: east-west balancing in a sidecar with the same policy template, and global steering that shifts load between regions on capacity, not just health."

Common Timing Mistakes#

MistakeTime LostFix
Explaining TCP three-way handshake and OSI layers5 minOne sentence: "L4 sees connections, L7 sees requests"
Comparing 6 balancing algorithms in depth8 minState P2C least-request and why not least-connections in one breath
Designing active-passive failover with VRRP5 min"Stateless fleet behind ECMP; no pairs"
Never reaching config safetyWhole interviewSteer there in Phase 4 — it's the largest real-world blast radius
Treating the LB hop as the latency problem3 minQuantify: < 1 ms vs 100+ ms for handshakes

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Every candidate has used a load balancer; very few have been paged because one made an incident worse. Interviewers use this question to separate candidates who think of the LB as plumbing from those who think of it as an automated decision-maker with fleet-wide authority. The Staff-level signal is recognizing the specific ways that authority fails: herding, over-ejection, sinkholes, correlated health checks, reconnect storms, and config pushes.

It's also a layering question. DNS, anycast, L4, L7 and client-side balancing each see different information and fail differently. Putting a decision at the wrong layer — trying to do request-level health at DNS, or connection-level stickiness at L7 for UDP — is a design error that no amount of tuning fixes.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

The Senior path builds a balancer that works when its inputs are true. The Staff path builds one that stays safe when its inputs are wrong.

1.3 The Staff Question That Cuts Through Everything#

"When this load balancer is wrong, how many requests does it get wrong — and how quickly does it stop?"

That one question surfaces the herd effect (wrong about which host is least loaded → all proxies pick it), over-ejection (wrong about health → removes capacity during overload), config pushes (wrong about routes → every request), and draining (wrong about when a host is gone → resets). Each part of the design either bounds the blast radius of a mistake or shortens its duration.

🎯 Staff Move: "I want every automated decision in the balancer to have a bound: ejection capped at 10% of the pool, panic mode below 50% healthy, retries capped at 20% of traffic, and config rolled out one cell at a time. The LB will be wrong sometimes. I'm designing for how wrong and for how long."


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Internet edge (north-south). Clients are many, far away and untrusted. The edge terminates TLS for millions of connections, absorbs abuse, routes by host and path to dozens of services, and is the place where a new connection's latency is mostly network round trips. The design centers on TLS economics, global steering, L4/L7 layering and config safety. This is the intent this case study commits to. For request-level policy on top of the edge — auth, quotas, transformations — see API Gateway.

Service-to-service (east-west). Clients are your own services, a millisecond or less away. Adding a central proxy hop doubles network traversals for every internal call. The design moves balancing into the caller — a library or a sidecar proxy — fed by service discovery. The hard problems shift to control-plane fan-out (pushing endpoint updates to 20,000 sidecars), consistent policy across languages, and retry budgets so a failing service isn't hit by every caller's retries at once.

L4 network load balancing. The unit is the packet or the connection, not the request. Use cases: non-HTTP protocols, UDP, extremely high packet rates, and preserving client IPs. The design centers on forwarding performance (kernel bypass, XDP), consistent hashing with connection tracking so flows survive changes, and direct server return so responses skip the balancer. It's the first tier of the edge design, but as a standalone problem it's a different interview.

DimensionInternet EdgeEast-WestL4 Network LB
Unit of balancingHTTP requestRPCConnection / flow
Dominant latency costTLS + client RTTsExtra hop if centralizedMicroseconds
Health signalActive + passive on responsesPassive per callerConnection success, active checks
Biggest riskGlobal config pushRetry storms, control-plane fan-outFlow remapping on fleet change
Who owns itEdge platform teamMesh / RPC platform teamNetwork team

2.2 When NOT to Use a (Dedicated) Load Balancer#

SituationBetter ChoiceWhy
Internal RPC between services in one data centerClient-side or sidecar balancingA central hop adds latency and a shared failure point for no benefit
Static assets and cacheable contentCDN — see CDN & Edge CachingServing from cache at the edge beats balancing to an origin
Stateful partitioned systems (databases, Kafka)Client routing by partition mapThe client must reach a specific node; balancing would be wrong
Very low traffic internal toolsDNS with 2–3 records or a single managed LBOperating a fleet costs more than it protects
Batch/queue consumersThe queue itself distributes workPull-based consumers self-balance; see Message Queues

🎯 Staff Move: "For east-west traffic I'd remove the load balancer, not scale it. Each internal call through a central proxy pays an extra hop and shares a failure domain with every other service."

2.3 What the Interviewer Leaves Underspecified#

Unstated AssumptionWhy It's Deliberately VagueWhat to Say
Edge or internalTo see if you ask"Is this the public edge or service-to-service? The designs diverge."
Connection lifetimeTo see if you consider WebSockets/HTTP2"Are there long-lived connections? That changes draining and balancing."
New connection rateTo see if you know what saturates the edge"How many new TLS connections per second? That sizes the fleet more than RPS."
Affinity needsTo see if you default to sticky sessions"Does anything need affinity, or is state external? I'd default to none."
Health definitionTo see if you question "healthy""What counts as healthy — process up, or serving correctly?"
Who changes configTo see if you think about ownership"Who edits routes, and how fast do changes need to land?"

2.4 Precise Terminology#

TermPrecise MeaningCommon Misuse
L4 load balancingDecisions per connection/packet using addresses and ports"Any load balancer that's fast"
L7 load balancingDecisions per request using HTTP semanticsAssumed to be "slow" — the hop is sub-ms
VIPVirtual IP that clients target; backed by many machinesConfused with a single machine's IP
ECMPRouter spreads flows over equal-cost next hops by hashing headersAssumed to keep flows stable when next hops change — it doesn't by itself
Direct server return (DSR)Backend replies to the client directly, bypassing the LBAssumed to work for L7 — it's an L4 technique
Outlier detectionPassive ejection based on observed responsesConfused with active health checking
Panic thresholdHealthy fraction below which the LB ignores health and uses all hostsSeen as unsafe; it prevents self-inflicted outages
Slow startRamping a new host's weight over timeConfused with TCP slow start
Connection drainingStop new work to a host, let in-flight finishAssumed to cover long-lived connections automatically
Session affinitySame client → same backendUsed to paper over state that should be external

3. The Five Fault Lines#

3.1 Fault Line 1: Where to Terminate — L4 vs L7#

L4 is cheap per byte and blind to requests. L7 sees everything and pays for it in CPU, mostly at connection setup. The question isn't which one to use — it's which decisions go where, and who pays for the CPU.

OptionWhat WorksWhat BreaksWho Pays
L4 only, TLS at backendsCheapest edge; end-to-end encryption to the service; any protocolNo per-request routing, retries or outlier detection; every service manages certs and handshake CPUService teams (TLS ops, CPU); users (no request-level failover)
L7 only (anycast/DNS straight to proxies)Full HTTP intelligence; one tierProxy changes remap flows without an L4 consistent-hash layer; DDoS lands on expensive L7 CPUEdge team (fleet size, attack exposure)
L4 in front of L7 (two tiers)L4 absorbs packet floods and keeps flows stable; L7 does request logicTwo fleets to run; one more hop (~tens of µs at L4)Edge team (two systems) — cheap relative to the alternatives
L7 with TLS passthrough for some servicesServices that need end-to-end TLS keep itThose services lose L7 features at the edgeThe services that opted in, knowingly

The Staff default: two tiers. L4 with consistent hashing and DSR in front; L7 terminating TLS and re-encrypting to backends where policy requires (mTLS inside the network). Passthrough as a reviewed exception.

When to deviate: a single-service product at < ~20K RPS can use one managed L7 balancer and skip the L4 tier entirely — the cloud provider's L4 is effectively in front already. Non-HTTP protocols (game servers on UDP, MQTT) go L4 only.

🎯 Staff Move: "L4 for things that need to be fast and stable — flows and floods. L7 for things that need to be smart — routing, retries, health from real responses. And I'd re-encrypt to the backend rather than pass TLS through, unless a service has a compliance reason to own its keys."

Full reasoning: what TLS actually costs

A TLS 1.3 full handshake costs the server one asymmetric signature (the certificate proof) plus a key exchange (ECDHE, X25519 is cheap). The signature dominates: RSA-2048 signing manages roughly a thousand to a few thousand operations per second per core; ECDSA P-256 manages tens of thousands. At 150K new connections/s, RSA-only would need ~75–150 cores just for signatures; ECDSA needs under 10. Serving ECDSA certificates to clients that support them (nearly all modern clients), with RSA as fallback, is the single biggest lever on edge CPU.

Resumption is the second lever. With session tickets or PSK resumption, a returning client skips the certificate signature entirely. Resumption only works if every proxy that might receive the client's next connection can decrypt the ticket — so ticket-encryption keys must be shared across the fleet in a region and rotated on a schedule (e.g., every 12–24 h, keeping the previous key for decryption). A fleet where each proxy has its own ticket key gets near-zero resumption behind a load balancer, and nobody notices until a traffic spike.

The user-facing cost is round trips: TCP (1 RTT) + TLS 1.3 (1 RTT) before the first request byte; TLS 1.2 adds another. HTTP/3 over QUIC combines transport and TLS setup into 1 RTT, and 0-RTT resumption can send replay-safe requests (GETs) immediately — never non-idempotent ones, because 0-RTT data can be replayed.

3.2 Fault Line 2: Even Spread vs Affinity#

StrategyWhat WorksWhat BreaksWho Pays
Round robinSimple, statelessUneven when request cost varies 10–100×; ignores slow hostsUsers routed to the slow host
Least connectionsAdapts to slow hosts on one LBHerding across many LBs; fast-failing hosts attract traffic; meaningless with HTTP/2 multiplexingUsers during host failures; the new host that gets swamped
P2C least-requestNear-optimal spread with no coordination; robust to stale dataStill rewarded by fast failure without outlier detectionNobody, if paired with ejection
Consistent hashing (ring/Maglev)Cache/session affinity; stable mapping on changeHot keys overload one host; weight changes move keysThe host owning the hot key; users on it
Consistent hashing with bounded loadAffinity until a host exceeds ~1.25× average, then spill to nextSome affinity lost under skewCache hit rate for hot keys
Sticky sessions (cookie)Server-local session state worksUneven load; failover loses sessions; drains take as long as sessionsUsers whose host dies; ops during deploys

The Staff default: P2C least-request everywhere; consistent hashing with bounded load only for routes where affinity measurably pays (a local cache with > 2× hit-rate improvement); no cookie stickiness — move session state to a shared store such as Redis.

When to deviate: a backend tier with large per-user in-memory state (real-time collaboration, game sessions) genuinely needs affinity; route by entity ID with consistent hashing and accept that rebalancing is a migration.

🎯 Staff Move: "Affinity is a cache-hit-rate optimization, so I'd only pay for it where I can measure the hit rate. Everywhere else, P2C — and if someone wants sticky sessions to hold session state, I'd rather fix where the state lives."

3.3 Fault Line 3: Trusting Health Signals vs Ignoring Them#

Health checking is a classifier with two error modes. False negatives keep a broken host in rotation — users see errors. False positives remove a healthy host — the remaining hosts get more load, which can cause more false positives. The second error mode compounds; the first doesn't.

PolicyDetects Bad Host InFalse-Positive RiskWhat BreaksWho Pays
Active shallow check only (process up)10–60 s for dead hosts; never for fast-failing onesLowFast-failing sinkhole stays in rotationUsers routed to the broken host
Active deep check (checks DB, cache, deps)10–60 sHigh and correlated: a shared dependency blip fails every host at onceFleet-wide ejection during a dependency hiccupEvery user — total outage from a partial one
Passive outlier detection, uncapped< 1 sMedium; overload causes errors everywhere → eject everythingCascading ejection under loadEvery user
Passive + capped ejection (10–20%) + panic threshold (50%)< 1 s for individual bad hostsBoundedSystemic problems aren't masked — they're visible as errorsUsers during true systemic failure (unavoidable), not amplified

The Staff default: shallow active checks (5 s interval, 2 failures to mark down, 2 successes to mark up) for liveness and admission; passive outlier detection (5 consecutive 5xx or success rate > 2σ below peers) for correctness; ejection capped at 10% of hosts; panic below 50% healthy.

When to deviate: for pools of fewer than ~5 hosts, a 10% cap means zero hosts ejectable — raise the cap to allow one host, and alert on any ejection.

🎯 Staff Move: "An LB that ejects aggressively is great when one host is bad and catastrophic when the signal is bad. I'll cap ejection at 10% and stop trusting health entirely below 50% healthy, because at that point the checks are much more likely to be wrong than half my fleet."

3.4 Fault Line 4: Central Proxy vs Client-Side Balancing#

ModelLatencyPolicy ConsistencyWhat BreaksWho Pays
Central L7 fleet for everything+1 hop per call (~0.2–1 ms)One place to changeShared failure domain for all internal traffic; fleet scales with total RPC volumeEvery service (latency); edge team (capacity)
Client libraryNo extra hopPer-language implementations driftPolicy fixes need every service to redeploy; N languages × M versionsService teams (upgrades); platform team (N libraries)
Sidecar proxy (mesh)~0.1–0.5 ms per side (local loopback hops)One proxy implementation, central control planeControl plane fan-out to every pod; sidecar CPU/memory per pod (~0.1–0.5 core, 50–200 MB)Infra cost across the fleet; platform team (control plane)
Proxyless (gRPC xDS in client)No hopCentral config, in-processOnly for supported languagesTeams outside the supported languages

The Staff default: central L7 for north-south; sidecars or proxyless clients for east-west, all fed by the same control plane and the same policy template.

When to deviate: organizations with < ~30 services often do fine with a client library in one language and DNS-based discovery; a mesh is a platform commitment of several engineers.

🎯 Staff Move: "The edge is a central proxy because clients are untrusted and far away. Inside, I'd push balancing to the caller because it's the only place with per-call latency data and no extra hop — but with one policy template, so a timeout default means the same thing everywhere."

3.5 Fault Line 5: Connection Longevity vs Rebalancing#

Long-lived connections are good for latency (no handshakes) and bad for balance (load is pinned to whoever accepted the connection). HTTP/2 and gRPC multiplex thousands of requests on one connection; WebSockets live for hours.

ApproachWhat WorksWhat BreaksWho Pays
Unlimited connection lifetimeFewest handshakesNew backends get no load; old ones stay hot; draining takes hoursHot backends (overload); deploys (slow)
Max connection age (e.g., 5–15 min) with jitterPeriodic rebalancing; bounded drain time~1 extra handshake per connection per age intervalSmall CPU cost; negligible latency with resumption
Request-level balancing at L7 over pooled connectionsEdge→backend balancing is per request regardless of client connection lifeClient→edge connections still pinned to a proxyEdge proxies (uneven across the fleet)
Application-level reconnect for WebSocketsControlled migration with jitterRequires client cooperationClient teams (protocol support)

The Staff default: client→edge connections capped at a max age (~10 min ± jitter) via GOAWAY; edge→backend balancing per request over pooled connections; WebSockets get an application "reconnect" frame and jittered reconnection spread over 5–10 minutes during drains.

When to deviate: for very latency-sensitive mobile clients on poor networks, lengthen the max age (30–60 min) — each reconnect costs them 2+ RTTs at 200 ms each.

🎯 Staff Move: "I'll balance per request, not per connection, wherever HTTP/2 is involved — otherwise I've built a load balancer that balances once per connection lifetime. And I'll cap connection age so a scale-out actually reaches new hosts within minutes."


4. Failure Modes & Operational Reality#

4.1 The Fast-Fail Sinkhole#

A backend breaks in a way that makes it fast. Load-aware balancing rewards speed, so the broken host attracts traffic.

t=0:        Backend b-22 (of 40) loses its connection to the config service after a
            cert expires. Handler returns 503 in 1.5 ms. Healthy p50: 35 ms.
t=+1s:      Least-request: b-22 has ~0 in-flight at all times. Selected whenever it's
            one of the two choices — and wins every comparison.
            Share of traffic: 2.5% → ~5% (P2C limits the damage; least-conn would hit ~30%+).
t=+1s:      Outlier detection: 5 consecutive 5xx on each proxy → b-22 ejected on most proxies.
t=+2s:      Error rate: peak 4.8% for ~1 s, back to 0.05%.
t=+30s:     b-22 re-admitted, fails again, ejected 60 s, then 90 s.
t=+5min:    Alert: lb.host_ejections{host=b-22} > 3 in 5 min. Auto-replace triggers.

Detection: lb.upstream_5xx_rate by host, lb.host_ejections, latency distribution by host (a host that is suddenly much faster than peers is a signal).

Mitigation: outlier detection (consecutive 5xx + success-rate relative to peers); P2C limits a sinkhole's share to roughly 2/N even before ejection, vs a much larger share under least-connections.

Prevention: health endpoints that exercise the critical in-process dependency state (without calling remote shared dependencies); cert-expiry alerts at 30/14/7 days.

Owner: service team owns the bug and its health endpoint; edge platform owns the outlier policy defaults.

4.2 The Health-Check Cascade#

t=0:        Shared auth service p99 rises from 20 ms to 3 s (GC pause storm).
t=+5s:      140 services' /health endpoints call auth to "verify dependencies."
            Health checks time out (2 s timeout). Hosts start failing checks.
t=+15s:     Pools drop below healthy thresholds. LBs without a panic threshold remove
            hosts; remaining hosts take more load, fail checks faster.
t=+30s:     34 services at 0 healthy hosts → LB returns 503 "no healthy upstream"
            for 100% of their traffic. Auth itself has recovered.
t=+90s:     Hosts pass 3 consecutive checks again; pools refill gradually.
t=+4min:    Full recovery. A 30-second auth blip became a 4-minute multi-service outage.

Detection: lb.healthy_hosts_pct per cluster; lb.no_healthy_upstream count; correlation of check failures across many services at the same moment (a correlated health drop is almost never real).

Mitigation: panic threshold — below 50% healthy, ignore health and use all hosts; if hosts are actually fine, traffic keeps flowing.

Prevention: health endpoints must not call remote shared dependencies; lint for it; dependency failures surface through passive outlier detection on real requests instead. Game day: inject latency in a shared dependency and verify no pool goes empty.

Owner: edge platform (panic threshold, policy); each service team (their /health); the platform that owns the health-endpoint standard.

4.3 The Global Config Push#

t=0:        Route change merged: a new regex path matcher for a marketing page.
            Pattern has catastrophic backtracking on certain paths.
t=+20s:     Config pipeline validates syntax (passes) and pushes globally — no canary
            stage for "route-only" changes.
t=+30s:     Proxy CPU 30% → 100% across every region. Added latency p99 2 ms → 9 s.
t=+1min:    Page: lb.added_latency_p99 > 500 ms, lb.cpu > 90% in all regions.
t=+6min:    On-call identifies the change via config version on dashboards.
t=+9min:    Rollback pushed. Proxies at 100% CPU take 2–3 min to apply it.
t=+13min:   Recovered. 99.99% monthly budget (4.3 min) exceeded 3× in one incident.

A well-known public example of this class: in July 2019, Cloudflare deployed a WAF rule containing a regular expression that caused excessive backtracking; deployed globally at once, it drove CPU to exhaustion across its network for roughly half an hour.

Detection: lb.cpu_utilization, lb.added_latency_p99, config version annotations on every edge dashboard.

Mitigation: one-click rollback to last-known-good; proxies apply config on a separate thread pool or core reservation so a CPU-saturated data plane can still accept a rollback.

Prevention: every config change — including "just a route" — goes through canary → AZ → region → global with bake time; regex engines with linear-time guarantees (RE2-style) for any user-authored pattern; CPU-cost benchmarks of new config in CI.

Owner: edge platform team owns the pipeline and the guardrails; the route author owns the change.

4.4 The TLS Handshake Storm#

t=0:        Scheduled ticket-key rotation runs. A bug distributes the new key to the
            proxies without keeping the previous key for decryption.
t=+0s:      Every resumption attempt fails → full handshakes.
            Resumption rate: 65% → 0%. Full handshakes/s: 50K → 145K.
t=+40s:     Proxy CPU 55% → 97%. Handshake latency p99 50 ms → 1.2 s.
            Mobile clients time out at 10 s, retry → more handshakes.
t=+2min:    Page: tls.full_handshakes_per_sec > 2× baseline, lb.cpu > 90%.
t=+5min:    On-call scales the L7 fleet +40% (takes 4 min to warm).
t=+9min:    Old key restored for decryption. Resumption back to 60%. CPU 60%.

Detection: tls.resumption_rate, tls.full_handshakes_per_sec, tls.handshake_latency_p99, proxy CPU.

Mitigation: restore previous ticket key; scale out; prefer ECDSA (if RSA was serving a share of clients, shifting cuts handshake CPU ~5–10×).

Prevention: ticket keys rotated with overlap (current + previous for decryption); resumption-rate alert at < 50% of baseline; capacity plan for zero resumption at peak, because a key mistake or a mass client reconnect produces exactly that.

Owner: edge platform (cert and key management).

4.5 The Long-Lived Connection Imbalance#

t=0:        Partner gRPC traffic grows. Backend pool scales from 20 to 40 hosts.
t=+10min:   The 20 new hosts serve ~3% of requests. The 20 old hosts at 85% CPU.
            Clients hold HTTP/2 connections for hours; L4 balancing at the backend
            tier balanced once, at connect time.
t=+30min:   Old hosts p99 rises from 80 ms to 900 ms. Autoscaler adds more hosts —
            which also get no traffic.
t=+40min:   On-call restarts old hosts in batches → mass reconnect → spread.

Detection: request distribution skew across hosts (max/mean > 2), new hosts' RPS after scale-out, connection age distribution.

Mitigation: per-request balancing at an L7 hop in front of gRPC backends, or client-side balancing in the gRPC client; max connection age with GOAWAY.

Prevention: never use L4-only balancing for multiplexed protocols; alert on per-host RPS skew; scale-out tests that check new-host traffic share within 5 minutes.

Owner: edge platform (protocol-aware balancing); partner integration team (client config).

4.6 The L4 Rehash#

t=0:        L4 tier scaled from 8 to 10 nodes for an event. Routers' ECMP next-hop
            set changes. A naive mod-N hash would remap ~80% of flows to different
            L4 nodes.
t=+0s:      L4 nodes without shared consistent hashing forward remapped packets to
            different L7 proxies than the ones holding the TCP state → RST.
t=+1s:      ~800K established connections reset. Clients reconnect at once:
            full TLS handshakes spike 10×.
t=+2min:    L7 CPU saturates from the handshake storm (see 4.4).

Detection: tcp.resets_sent at L7, l4.flow_table_misses, sudden handshake spikes coinciding with L4 fleet changes.

Mitigation: consistent hashing (Maglev-style) on every L4 node using the same backend table, so any L4 node sends a given flow to the same L7 proxy; connection tracking for in-flight flows during backend set changes.

Prevention: L4 fleet changes during low-traffic windows; consistent hashing is mandatory; validate in staging that adding an L4 node resets < 1% of flows.

Owner: network / L4 team.

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Fast-fail sinkholePer-host 5xx; host faster than peers~2/N of traffic with P2C until ejectedOutlier ejectionService team (bug); edge (policy)
Health-check cascadelb.healthy_hosts_pct drops across many clusters at onceEvery service sharing the dependencyPanic thresholdEdge platform + service health owners
Global config pushlb.cpu, lb.added_latency_p99 in all regionsEverything behind the edgeRollback to last-known-goodEdge platform
TLS handshake stormtls.resumption_rate drop, handshake/s spikeA region or globalRestore ticket keys; scale outEdge platform
Long-lived imbalancePer-host RPS skew > 2×One servicePer-request L7 balancing; max connection ageEdge + client teams
L4 rehashTCP resets, flow-table missesAll connections through the changed tierConsistent hashing + conn trackingNetwork team
Retry amplificationlb.retry_rate > 20% of requestsFailing backend and its dependenciesRetry budget; single retry layerService team (policy), edge (budget enforcement)
Control-plane outagexds.config_age_seconds risingNo config changes; data plane must continueLast-known-good; static fallbackEdge platform
Zone imbalancePer-zone RPS vs capacityOne zone overloadedZone-aware routing with spilloverEdge platform

🎯 Staff Insight: Five of these nine failures are the load balancer amplifying a smaller problem — one bad host, a dependency blip, a key rotation, a fleet change. A good LB design is mostly a list of bounds on its own automated decisions.


5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
FramingDraws an LB in front of serversSeparates edge, east-west and L4; commits to oneAsks how many balancing layers exist and which are redundant
LatencyMinimizes hopsLocates cost in TLS and RTTs; resumption, ECDSA, reusePrices handshake CPU across the fleet; sets TLS policy org-wide
AlgorithmLeast connections or round robinP2C least-request; bounded-load hashing only where affinity paysStandardizes defaults in a shared template; reviewed overrides
HealthActive checksShallow active + passive outlier detection; capped ejection; panic modeOwns the health-endpoint standard to prevent correlated ejection
FailureActive-passive pairsStateless fleets, consistent hashing, staged config with auto-rollbackTreats the config pipeline as the org's highest-blast-radius deploy system
ConnectionsNot consideredMax connection age, GOAWAY, per-request balancing for HTTP/2, WebSocket drainSets protocol policy (HTTP/3, 0-RTT scope) with security and client teams
OwnershipInfra owns the LBEdge owns the fleet; services own routes, timeouts, healthDefines platform vs service contracts and the validation that enforces them

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Locates the real latency"The proxy hop is under a millisecond. The two round trips for TCP and TLS are 160 ms for a user 80 ms away."
Knows how balancing herds"Fifty proxies running least-connections all see the new host as empty at the same moment."
Bounds automated decisions"Ejection capped at 10%, panic below 50% healthy, retries capped at 20%."
Separates liveness from correctness"Active checks tell me the process is up. Real responses tell me it's working."
Treats config as a deploy"Every route change goes canary, AZ, region, global — with automatic rollback."
Handles connection lifetime"With HTTP/2 I balance per request, or I've balanced once per connection lifetime."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
Spends the interview comparing algorithmsAlgorithm choice is rarely the failure; health and config are
Deep health checks that call the databaseCreates correlated, fleet-wide ejection on a dependency blip
Active-passive pair as the HA storyIgnores that both members run the same bad config
Sticky sessions by defaultHides a state-placement problem and makes failover and deploys worse
Retries at every layerMultiplies load during the exact incidents retries are meant to survive
"The LB adds latency" with no numbersMisplaces the cost — and optimizes the wrong thing

5.4 Common False Positives#

  • Knowing Maglev's lookup-table math ≠ edge design. Impressive, but the interview is about bounding failures.
  • Listing every cloud LB product ≠ judgment. Product names without policies (ejection caps, drain windows) are a catalog, not a design.
  • Packet-level depth ≠ L7 understanding. Candidates strong on XDP and DSR sometimes miss request-level health entirely.
  • "Use a service mesh" ≠ east-west design. A mesh is a mechanism; the policy (timeouts, retries, outlier rules) is the design.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–4 minCommit to edge; capture RPS, new-connection rate, protocols, latency budget
Entities and API4–7 minListener, route, cluster, endpoint; who owns what
Architecture7–12 minDNS/anycast → L4 → L7 → backends; control plane
Latency and TLS12–19 minWhere time goes; resumption, ECDSA, reuse
Balancing and health19–31 minP2C, outlier detection, caps, panic, slow start
Config safety31–37 minValidation, staged rollout, rollback, data plane independence
Connections and draining37–42 minMax age, GOAWAY, WebSockets
Wrap-up42–45 minSummary, east-west, global steering

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response
"Traffic triples overnight."What saturates firstNew-connection rate × handshake CPU; resumption rate; scale L7, not L4
"An entire AZ goes dark."Zone awareness, capacityRemaining zones need N+1 headroom (~50% extra at 3 AZs); spillover policy; ECMP withdraws routes
"A backend returns errors fast."Sinkhole awarenessOutlier detection; why least-connections makes it worse
"The database behind every service is slow."Over-ejectionPanic threshold, capped ejection, shallow health checks
"We need sticky sessions."Pushback on stateExternalize session; if truly needed, bounded-load consistent hashing
"Make it multi-region."Global steeringGeoDNS/anycast, capacity-aware shifting, DNS TTL realities
"Design the east-west version."Layer judgmentRemove the central hop; sidecar or proxyless with same policy template

6.3 What to Deliberately Skip#

  • OSI model recitation. "L4 sees connections, L7 sees requests" is enough.
  • VRRP/keepalived active-passive details. Stateless fleets behind ECMP replace them.
  • Exact Maglev table-population algorithm. Name it, state the property (minimal disruption), move on.
  • Cipher suite lists. "TLS 1.3, ECDSA with RSA fallback, resumption" covers it.
  • WAF rule design. It belongs to an edge-security interview.

6.4 Follow-Up Questions to Expect#

  1. "How does a new L7 proxy join without disrupting connections?" — L4 consistent hashing with connection tracking: existing flows stay pinned; only new flows map to the new proxy.
  2. "How do you drain a proxy for upgrade?" — Withdraw it from the L4 table for new flows, send GOAWAY / Connection: close, wait for in-flight (≈ p99 request time), cap at a few minutes; WebSockets get reconnect frames.
  3. "How do you pick the retry layer?" — Retry where you have the most context and the fewest multipliers: usually the edge for idempotent requests, with a budget; backends don't retry each other's retries.
  4. "How do you balance across zones?" — Prefer local zone; spill proportionally when local healthy capacity falls below demand; avoid permanent cross-zone traffic (latency + transfer cost).
  5. "What if the control plane is down?" — Proxies keep last-known-good config and endpoint lists; no new routes ship; endpoint churn is absorbed by outlier detection.
  6. "How do you protect the LB from DDoS?" — Anycast spreads volumetric attacks; L4 drops malformed packets and SYN floods (SYN cookies); L7 rate limits per client — see Rate Limiting.
  7. "How do you roll out HTTP/3?" — Advertise via Alt-Svc to a percentage of clients, compare latency and error rates, keep TCP fallback; UDP needs L4 support for QUIC connection IDs.

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design a load balancer."

Staff Answer

"Three different systems answer to that name: the internet edge in front of a public service, balancing between internal services, and a pure L4 packet balancer. I'll design the edge for a latency-sensitive HTTP API, sketch the L4 tier it sits on, and contrast east-west at the end.

Assumptions: 1M RPS globally, ~150K new TLS connections/s at peak, HTTP/1.1, HTTP/2 and some WebSockets, three regions with three AZs, 99.99% availability, added p99 under 5 ms excluding client RTTs. Plan: where latency goes → balancing and health, because that's where balancers amplify failures → config safety, because it's the largest blast radius → connection lifetime and draining."

Why this is L6:

  • Distinguishes three intents with different designs and commits
  • Asks for new-connection rate, not just RPS — the number that actually sizes the edge
  • Orders the deep dives by where failures come from

What L7 adds:

  • Asks which existing balancing layers this replaces or joins, and who owns them
  • Frames the config pipeline as an org-level risk surface with its own change policy
❌ Common L5 Trap

"Two load balancers in active-passive behind a floating IP, least-connections to a pool of web servers, health check every 10 seconds, sticky sessions for logged-in users."

Why this misses: It's a working design for one data center on a good day. It herds on scale-out, sinkholes on fast failure, loses sessions on host death, and has no answer for a config mistake that both LBs share.


Drill 2: The Algorithm#

Prompt: "Why not least-connections?"

Staff Answer

"Three reasons. First, with many proxies each seeing only their own connections, they all see a newly added or recovered host as least-loaded at the same instant and herd onto it. Second, with HTTP/2 and keepalive, a connection carries anywhere from zero to hundreds of concurrent requests, so connection count isn't load. Third, a host that fails fast always has the fewest active requests, so it attracts the most traffic.

P2C least-request fixes the first two: each proxy samples two hosts at random and picks the one with fewer in-flight requests. Randomness spreads decisions across proxies, and the theory says two choices get you most of the benefit of perfect information — maximum load above average drops from roughly log n / log log n to log log n. The third needs outlier detection, because any load-aware algorithm rewards a fast failure."

Why this is L6:

  • Names the specific failure modes, not just a preference
  • Knows why P2C works with stale, partial information
  • Admits P2C's remaining weakness and pairs it with ejection

What L7 adds:

  • Puts P2C + outlier detection in the shared default template so teams can't drift into least-connections
  • Ties algorithm choice to an incident class the org has actually had

Drill 3: The Health Check#

Prompt: "Design the health checking. A teammate wants /health to check the database, cache and auth service."

Staff Answer

"I'd push back. If /health checks shared dependencies, a 20-second blip in auth fails every host's check at the same moment, and the LB removes the whole fleet — turning a partial dependency problem into a total outage. Active checks should be shallow: the process is up, the listener accepts, in-process critical state is initialized. 5 s interval, 2 failures down, 2 successes up.

Correctness comes from real traffic: passive outlier detection ejects a host after 5 consecutive 5xx or when its success rate is > 2σ below peers, for 30 s doubling on repeats, capped at 10% of the pool. And if the healthy fraction falls below 50%, panic mode spreads traffic across all hosts. If the database is truly down, users will see errors either way — but the LB won't add 'no healthy upstream' on top for the minutes it takes hosts to pass checks again."

Why this is L6:

  • Identifies correlated failure from deep checks
  • Separates liveness (active) from correctness (passive)
  • Bounds the LB's own ability to cause an outage

What L7 adds:

  • Turns "health endpoints must not call remote shared dependencies" into an org standard with a lint and a game day
  • Measures it: count of services whose pools went empty during the last shared-dependency incident

Drill 4: Make It Concrete — Sizing the Edge#

Prompt: "400K RPS in the largest region, 60K new TLS connections/s, average response 20 KB. Size the L7 and L4 tiers."

Staff Answer

"L7 has two loads: requests and handshakes. Requests: at ~15K simple proxied req/s per core including TLS record encryption, 400K RPS needs ~27 cores. Handshakes: assume 50% resumption and ECDSA for 95% of clients — 30K full handshakes/s; at ~20K ECDSA signatures/s/core that's ~1.5 cores, plus ~1.5K RSA handshakes/s at ~1.5K/s/core = ~1 core. Key exchange and record setup add a few more. Call it ~35 cores busy. Then plan for zero resumption — a key bug or mass reconnect — which roughly doubles handshake cost, and N+1 per AZ with 50% headroom so we survive losing an AZ: ~35 × 2 (headroom) × 1.5 (AZ loss) ≈ 105 cores → ~8 proxies of 16 cores, spread 3 per AZ, so 9.

L4: with DSR, the L4 tier sees only inbound packets. 400K RPS with a few packets per request plus ACKs is a few million packets/s — one or two XDP-based nodes can forward that; I'd run 4 across AZs for redundancy. Egress bandwidth, 400K × 20 KB ≈ 8 GB/s ≈ 64 Gbps, flows from L7 proxies to clients and never touches L4 — that's what makes the L7 NICs, not their CPUs, a likely limit: 9 proxies × 25 Gbps leaves headroom, while 9 × 10 Gbps doesn't survive losing an AZ (6 × 10 = 60 Gbps < 64)."

Why this is L6:

  • Sizes L7 on both requests and handshakes and plans for zero resumption
  • Includes AZ-loss headroom explicitly
  • Notices DSR shrinks L4 and that NIC bandwidth may bind before CPU

What L7 adds:

  • Converts the RSA fallback share into dollars: dropping it from 5% to 1% saves little; the ECDSA switch itself saved ~80% of handshake CPU
  • Asks whether static content could move to a CDN, removing a large fraction of egress from the fleet entirely

Drill 5: The Dependency Goes Down#

Prompt: "A backend service's database slows down. 60% of its hosts start timing out. What does your LB do, minute by minute?"

Staff Answer

"First seconds: outlier detection starts ejecting hosts with consecutive timeouts, but hits the 10% ejection cap almost immediately — so most slow hosts stay in. That's intentional: the hosts aren't broken, their dependency is. P2C least-request steers away from hosts with long in-flight queues toward whichever are momentarily faster. Requests hit the route timeout — say 2 s — and the edge retries idempotent requests once, within a retry budget of 20% of traffic, so retries can't double load on a struggling service.

If the active checks were deep, we'd be in panic mode now; with shallow checks, hosts stay healthy and we serve what we can. Errors are visible to users and to the service's on-call — accurately attributed to the service, not to 'no healthy upstream'. The fix is in the service: shed load, fail fast with a circuit breaker on the database, serve degraded responses. The LB's job is to not make it worse."

Why this is L6:

  • Walks through how each safety rail engages
  • Uses a retry budget instead of unconditional retries
  • Places the fix with the service and keeps the LB honest about error attribution

What L7 adds:

  • Adds per-service load-shedding headers (e.g., backend signals overload; edge stops retrying) as a platform contract
  • Defines which services get priority edge capacity when the fleet is constrained

Drill 6: The Hot Key#

Prompt: "A cache-heavy service uses consistent hashing on user_id for affinity. A celebrity account now generates 8% of its traffic."

Staff Answer

"Plain consistent hashing pins that 8% to one host, which in a 30-host pool is ~2.4× average load. I'd switch to consistent hashing with bounded load: a host may take at most ~1.25× the average in-flight load; above that, requests for its keys spill to the next host on the ring. The celebrity's requests spread over two or three hosts, which each warm their caches for that key — a small hit-rate cost on a single key in exchange for no overload.

Longer term: if one key is that hot, it should be cached closer — at the edge or in a per-proxy micro-cache with a 1–5 s TTL — so most of its requests never reach the pool."

Why this is L6:

  • Quantifies the overload from the skew
  • Uses bounded-load hashing to keep most affinity while capping any single host
  • Moves the hottest content to a better layer

What L7 adds:

  • Makes bounded load the default for every hashing policy in the template — unbounded hashing is an exception that needs review
  • Connects to the CDN strategy: hot keys are an edge-caching problem first

Drill 7: Multi-Tenant Edge#

Prompt: "60 services share the edge fleet. One team's route change or traffic spike shouldn't hurt the others. How?"

Staff Answer

"Three kinds of isolation. Config: routes are owned per service, validated by the platform — a route can't change listener-level behavior, can't use non-linear-time regex, and can't set timeouts above the listener maximum. Changes roll out staged per cell, and a bad route rolls back automatically. Capacity: per-service circuit-breaker limits on the proxy — max connections, max pending requests, max concurrent retries to each cluster — so one slow service can't consume all proxy workers or connection pools. Traffic: per-service and per-client rate limits at the edge; an attack on one hostname is shed before it consumes shared CPU.

For the two or three services with very different needs — a large WebSocket service, for example — I'd give them a separate proxy pool behind the same L4 tier, so their connection counts don't dictate the shared fleet's memory profile."

Why this is L6:

  • Covers config, capacity and traffic isolation separately
  • Uses per-cluster limits on the proxy to stop resource exhaustion from one backend
  • Recognizes when workloads should get separate pools

What L7 adds:

  • Defines the platform/service contract: what service teams may configure, and what's reserved
  • Uses per-service cost attribution (handshakes, egress) to make outlier tenants visible

Drill 8: Build vs Buy#

Prompt: "Cloud-managed load balancer, self-run open-source proxy, or build our own?"

Staff Answer

"Don't build a proxy — the data plane is a solved problem and owning its CVEs is a permanent tax. The real choice is managed vs self-run.

Managed L7 balancers are the default up to tens of thousands of RPS per service and when the team has no edge specialists: no fleet, built-in cert management, pay per use. Self-running an open-source proxy fleet (Envoy, HAProxy, NGINX) pays off when you need control the managed product doesn't give — custom outlier policies, per-request routing logic, multi-cloud consistency, east-west mesh with the same proxy — or when per-request pricing at hundreds of thousands of RPS exceeds the cost of a 3–5 person edge team. What I'd always build in-house is the config pipeline: validation, staged rollout and rollback — because that's where our outages come from, managed or not."

Why this is L6:

  • Rejects building the data plane with a reason
  • Gives triggers for each option
  • Identifies the config pipeline as the part to own regardless

What L7 adds:

  • Prices it: managed per-request and per-connection charges at 400K RPS vs ~$1M/yr for a small edge team
  • Considers portability: self-run proxy config is portable across clouds; managed LB config isn't

Drill 9: Changing Routing Without an Outage#

Prompt: "Move 30% of traffic for /checkout to a new backend version. Then all of it. Safely."

Staff Answer

"Weighted routing on the route: checkout-v1: 99, checkout-v2: 1. Step through 1% → 5% → 30% → 100% with a bake of 10–30 minutes each, comparing per-version error rate and latency, auto-reverting if v2's 5xx rate exceeds v1's by 0.5 percentage points or p99 by 20%. The config change itself rolls out canary → AZ → region → global, so a mistake in the weight change is contained too.

Two subtleties. Sessions: if checkout carries state between requests, I'd hash on user ID for the split so a user doesn't bounce between versions mid-checkout. Draining v1 at the end: set weight to 0, wait for in-flight requests to finish — a drain window of 2× p99 — then remove the endpoints."

Why this is L6:

  • Uses weighted routing with explicit, automatic success criteria
  • Distinguishes the traffic shift from the config rollout of the shift
  • Handles per-user consistency during the split

What L7 adds:

  • Makes progressive delivery a platform capability with standard metrics, not a per-team script
  • Ties rollout gates to the service's error budget

Drill 10: Multi-Region#

Prompt: "We're going from one region to three. How does global load balancing work, and what happens when a region fails?"

Staff Answer

"Two mechanisms. GeoDNS returns the nearest healthy region's VIPs with a TTL of 30–60 s; or anycast announces one VIP from all regions and BGP picks the nearest. Anycast reacts faster — withdrawing routes shifts traffic in seconds — while DNS failover takes minutes in practice because resolvers and clients hold records past the TTL.

Region failure is a capacity problem more than a routing problem. If each of three regions runs at 60% at peak, losing one sends ~30% extra to each survivor — 90%, uncomfortably close to saturation. So either each region runs at ≤ 50–60% with explicit failover headroom, or failover is capacity-aware: shift only what survivors can absorb and shed or degrade the rest. I'd also make failover a deliberate decision with criteria, and test it with quarterly region drains — the first real failover should not be the first time the survivors see that load."

Why this is L6:

  • Compares DNS and anycast with realistic reaction times
  • Recognizes failover as a capacity question with numbers
  • Insists on regular drills

What L7 adds:

  • Prices the headroom: running at 55% instead of 75% is ~35% more edge and backend capacity — a business decision with a dollar figure
  • Coordinates with stateful tiers: shifting traffic to a region whose database is a read replica is a multi-region data question, not an LB question

8. Deep Dive Scenarios#

Deep Dive 1: Peak-Traffic Incident — Launch Day Handshake Saturation#

Context: A marketing campaign drives a product launch. Requests per second rise 3×, but edge proxy CPU goes from 40% to 98% and handshake latency p99 reaches 2 seconds. Backends are at 50% CPU and fine. The on-call has started adding proxies; it takes 6 minutes per batch to warm up. You're pulled in.

Questions to Surface First:

  • New connections per second vs baseline? Is this a request problem or a connection problem?
  • What's the resumption rate? It should drop with new users — by how much?
  • What fraction of handshakes use RSA vs ECDSA?
  • Are clients retrying timed-out handshakes, compounding the load?

Typical L5 Approach: Scales the proxy fleet and waits. Possibly raises proxy worker counts. Recovers in 20–30 minutes once capacity arrives.

Staff Approach: Sees that new users mean near-zero resumption: full handshakes rose 8× while RPS rose 3×. Finds that the cert served is RSA-only because an ECDSA cert was never provisioned for the launch hostname. Deploys the ECDSA cert (canary first), cutting handshake CPU ~5×, while scaling out in parallel. Sets client retry backoff via config for the mobile app's next session.

Principal Approach: Makes "new-connection capacity at zero resumption" a launch-readiness criterion, and moves cert provisioning for new hostnames into the platform with dual ECDSA/RSA as the only supported option. Asks marketing to share campaign calendars with the edge team as part of the launch process.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Check tls.full_handshakes_per_sec (8× baseline) vs RPS (3×). Check tls.resumption_rate (65% → 12%). Confirm cert type on the launch hostname: RSA-2048 only.
TriageHandshake CPU dominates: ~85% of proxy CPU in signature operations. Client retries on 10 s timeout add ~30% more handshakes.
Quick fixPush dual ECDSA + RSA cert for the hostname through the cert pipeline, canary one proxy per AZ, then fleet. CPU 98% → 55% within 4 minutes. Continue scale-out for headroom.
GuardrailsAlert on tls.full_handshakes_per_sec > 2× baseline; capacity model includes zero-resumption scenario; cert lint rejects RSA-only certs for new hostnames.
Post-mortemWhy was RSA-only provisioned for a new hostname? Why didn't launch readiness include connection-rate load testing?

Metrics to Watch: tls.full_handshakes_per_sec, tls.resumption_rate, tls.handshake_latency_p99, lb.cpu_utilization, tls.handshakes_by_sig_alg

Organizational Follow-up: edge platform owns cert provisioning (no hand-made certs); launch checklist includes a connection-rate test at 3× expected new users.

Ownership Question: "Who owns making sure the edge can handle a marketing launch?" Staff answer: The edge platform owns capacity and the readiness test. The launching team owns telling us — with the expected new-user count — two weeks ahead. If they don't, the readiness gate blocks the hostname from going live.

Key Takeaway: "Launch traffic is new-user traffic, and new users don't resume TLS sessions. Size the edge on new connections at zero resumption, not on requests."

What clears the Staff bar:

  • Distinguishes connection growth from request growth
  • Finds the cert type as a 5× CPU lever
  • Fixes the cause while scaling, not instead of fixing

Deep Dive 2: Silent Failure — One Zone Running Hot for Weeks#

Context: A capacity review shows that one AZ's backends run at 75% CPU while the other two run at 35%. Latency p99 for users routed through that zone has been 40% worse for at least three weeks. No alert fired; overall SLOs were met. Cross-AZ transfer costs also rose 20% last month.

Questions to Surface First:

  • Is zone-aware routing enabled? What's the spillover rule?
  • Is the L7 fleet itself evenly spread across zones? Are the L4 tier's ECMP weights equal?
  • Did a deploy change the number of backends per zone (e.g., one zone got fewer hosts after a capacity shortage)?
  • Where does the cross-AZ transfer come from?

Typical L5 Approach: Adds hosts in the hot zone until CPU evens out.

Staff Approach: Finds the mechanism: an instance-type shortage left the hot zone with 12 backends vs 20 in the others, but zone-aware routing sends traffic in proportion to proxy placement, not backend capacity. The hot zone's proxies keep traffic local until hosts are "unhealthy", which they never quite become. Meanwhile, other services had spillover on, creating the cross-AZ transfer. Fixes routing to weight by healthy backend capacity per zone.

Principal Approach: Treats the absence of an alert as the main finding. Adds per-zone utilization skew as an SLO-adjacent signal across all services, and makes "capacity-weighted zone routing" the platform default so zone imbalance self-corrects.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Compare per-zone backend count, per-zone RPS and per-zone CPU. 12 vs 20 vs 20 hosts; RPS ~equal per zone.
TriageZone-aware routing keeps traffic local unless local healthy hosts fall below a threshold — capacity-unaware. Other services spill across zones at a fixed 30%, driving transfer costs.
Quick fixEnable capacity-weighted zone routing: local share = local capacity / local demand, spill the remainder proportionally. Hot zone CPU 75% → 52%; cross-zone share for this service ~12%.
GuardrailsAlert on zone.utilization_skew = max/min > 1.5 for 1 h. Dashboard of cross-AZ bytes per service.
Post-mortemWhy could a zone run 3 weeks at 2× the others with no signal? Why was spillover a fixed percentage?

Metrics to Watch: zone.utilization_skew, lb.cross_zone_request_pct, network.cross_az_bytes, per-zone latency_p99

Organizational Follow-up: capacity team reports per-zone host counts after any shortage; edge platform owns zone-routing defaults; service dashboards show per-zone breakdowns by default.

Ownership Question: "Who notices a zone imbalance?" Staff answer: The edge platform, through a fleet-wide skew alert. Individual service teams see aggregate SLOs that hide it; that's why it went three weeks.

Key Takeaway: "Averages hide zone skew. Route by capacity, not by location alone, and alert on the skew itself."

What clears the Staff bar:

  • Finds the mismatch between proxy placement and backend capacity
  • Fixes routing policy instead of throwing hosts at it
  • Recognizes cross-AZ cost as a second symptom of the same policy

Deep Dive 3: Large-Customer Onboarding — 400K WebSockets From One NAT#

Context: An enterprise customer integrates a real-time dashboard product. They'll open ~400K WebSocket connections, all from a handful of corporate NAT egress IPs, and require their traffic to come from fixed IPs they can allowlist (for callbacks) and mTLS for their API calls. Go-live is in five weeks.

Questions to Surface First:

  • Does any layer hash on source IP? If so, a few NAT IPs land on a few proxies.
  • What's per-connection memory on the proxies with TLS? 400K × ~30 KB = ~12 GB across the fleet — fine if spread, a problem if concentrated.
  • How will these connections drain during deploys — what's the reconnect behavior of their client?
  • Is mTLS terminated at the edge with cert validation against their CA, or passed through?

Typical L5 Approach: Adds proxies to absorb the connections and enables sticky sessions by source IP for the WebSockets.

Staff Approach: Avoids source-IP hashing — with a handful of NAT IPs it puts thousands of connections on a few proxies. The L4 tier hashes on the full 5-tuple (source port varies), which spreads well. Places the WebSocket service on a dedicated proxy pool behind the same L4 tier so 400K long-lived connections don't shape the shared fleet. Agrees on a reconnect protocol: server sends a reconnect frame during drains; client reconnects with 0–300 s jitter. mTLS terminated at the edge with their CA in a per-tenant trust bundle; fixed egress IPs via a dedicated NAT for callbacks.

Principal Approach: Uses the deal to define an enterprise-edge tier — dedicated pools, tenant trust bundles, static egress, contractual reconnect behavior — priced into the enterprise SKU rather than built bespoke for one customer.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (week 1)Load-test 400K connections from 4 source IPs against staging. Observe distribution with source-IP hashing (top proxy: 31% of connections) vs 5-tuple (top proxy: 4.1%).
Triage (design)Dedicated WebSocket pool: 8 proxies, ~50K connections each, ~1.5 GB connection memory each. Max connection age 6 h with reconnect frame.
Quick fix (rollout)Start at 10% of their users, then 50%, then 100% over two weeks. Drain test: deploy during business hours at 10% to verify reconnect jitter.
GuardrailsAlert on lb.connections_per_proxy skew > 1.5×; ws.reconnects_per_sec > 2× baseline.
Post-mortem (retrospective)Document the enterprise-edge tier and its cost.

Metrics to Watch: lb.active_connections per proxy, ws.reconnects_per_sec, tls.mtls_handshake_failures by tenant, proxy memory

Organizational Follow-up: edge platform owns the dedicated pool; the real-time product team owns the reconnect protocol; sales engineering knows that fixed egress and mTLS are enterprise-tier features.

Ownership Question: "Who decides when to drain the WebSocket pool?" Staff answer: The real-time product team schedules deploys; the edge platform enforces the drain protocol and rate. Neither can drain without the other's mechanism.

Key Takeaway: "Source-IP affinity fails exactly for enterprise customers, who arrive through a few NAT addresses. Long-lived connections need their own pool and a negotiated reconnect protocol."

What clears the Staff bar:

  • Predicts the NAT hot spot and measures it before go-live
  • Isolates long-lived connections from the shared fleet
  • Designs draining as a protocol with the client, not a timeout

Deep Dive 4: Post-Mortem — 13 Minutes of Global Latency and 5xx From a Route Change#

Context: You lead the post-mortem for incident 4.3: a route change containing a backtracking regex was pushed globally and saturated every edge proxy for about 13 minutes. Initial root cause: "bad regex." Leadership asks how to make sure it can't happen again.

Questions to Surface First:

  • Why did a route change skip canarying? Is there a class of change that's considered "safe"?
  • Why could a service team author an arbitrary regex in edge config?
  • Why did rollback take 3+ minutes to apply?
  • Could the proxies have protected themselves — per-route CPU limits, linear-time regex?

Typical L5 Approach: Root cause: bad regex. Action items: review regexes more carefully; add a regex linter.

Staff Approach: The regex was the trigger. Root causes: (1) route changes bypassed staged rollout because they were classed "low risk"; (2) the regex engine allowed backtracking; (3) config application competed for CPU with the saturated data plane, slowing rollback. Fixes all three: no change class bypasses staging; linear-time regex engine only; config application reserved a core.

Principal Approach: Reframes the edge config pipeline as safety-critical infrastructure. Introduces an error-budget policy: when edge config changes consume more than 25% of the monthly budget, non-emergency edge config changes freeze until the pipeline gains a new safeguard. Reviews every system with global config push in the company — DNS, feature flags, WAF — for the same pattern.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Confirm all regions on last-known-good config. Freeze edge config changes pending review.
TriagePipeline history: "route-only" changes exempt from canary since a year ago to speed up marketing launches. Regex engine: backtracking. Rollback apply time: 3–4 min under 100% CPU.
Quick fixRemove the exemption. Switch to a linear-time regex engine; reject constructs it can't support. Reserve one core per proxy for control-plane work.
GuardrailsCI benchmark: every config change measured for CPU per request on replayed traffic; > 10% regression blocks. Auto-rollback if lb.cpu rises > 20 points in a canary cell.
Post-mortemFive whys ends at: speed of marketing launches was traded against edge safety without anyone deciding to make that trade.

Metrics to Watch: config.rollout_stage_duration, config.auto_rollbacks, lb.cpu_utilization per cell, config.apply_latency_seconds

Organizational Follow-up: edge platform owns the pipeline with no exemption classes; marketing launches get a fast lane that still canaries (10 min instead of 60); the pipeline's error-budget consumption is reported monthly.

Ownership Question: "Who approved skipping canary for route changes?" Staff answer: Nobody explicitly — it was a pipeline change merged to fix a speed complaint. The edge platform owns that now: pipeline safety settings require the same review as the data plane.

Key Takeaway: "Any change that reaches every proxy at once is a global outage waiting for a trigger. There is no such thing as a safe-to-skip-canary config class."

What clears the Staff bar:

  • Refuses "bad regex" as a root cause
  • Fixes the pipeline, the engine and the rollback path
  • Finds the organizational trade-off that created the exemption

Deep Dive 5: Multi-Region Expansion — Adding APAC#

Context: The company serves APAC users from a US region with ~180 ms RTT. Product wants a Singapore region in two quarters. Backends will be deployed there, but the primary database stays in the US for now. The edge team must decide how traffic reaches the new region and what happens on failure.

Questions to Surface First:

  • Which requests can be served fully in-region (reads from replicas, static), and which must cross to the US (writes)?
  • Anycast or GeoDNS? Do we have anycast address space and BGP presence in APAC?
  • What happens to APAC users if the Singapore region fails — fail to US with +180 ms, or degrade?
  • Capacity: can the US region absorb APAC traffic during a Singapore failure?

Typical L5 Approach: Adds a GeoDNS record for APAC pointing to Singapore and deploys the stack there.

Staff Approach: Starts with edge-only: terminate TLS in Singapore first, keep backends in the US, and send requests over warm, pooled connections across the backbone. That alone removes ~2 RTTs × 180 ms of handshake from every new APAC connection — most of the user-visible win — before any backend moves. Then moves read paths into the region, with writes routed to the US by path. Uses GeoDNS initially (no anycast space), with 60 s TTL and failover to the US edge; US capacity planned to absorb APAC at 30% headroom.

Principal Approach: Sequences the investment by user-visible latency per dollar: edge termination first (cheap, large win), read paths second, data residency and writes last (expensive, one-way doors). Ties the plan to the multi-region data strategy so the edge doesn't promise more locality than the data tier can deliver.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (planning)Measure: APAC new-connection latency today ≈ 3 × 180 ms before first byte (TCP + TLS 1.2 in older clients). Target with local termination: ≈ 3 × 20 ms + one backbone round trip on a warm connection.
Triage (design)Phase 1: Singapore edge (L4 + L7) with pooled connections to US backends. Phase 2: read-heavy routes served from Singapore backends + read replicas. Phase 3: write paths stay US until the data tier decides on regional writes.
Quick fix (rollout)GeoDNS to Singapore for 5% of APAC resolvers, compare p50/p99 time-to-first-byte, ramp over 2 weeks.
GuardrailsHealth-based DNS failover to the US edge; region.capacity_headroom alert if US can't absorb APAC; quarterly regional drain test.
Post-mortem (readiness)Document per-route locality: which routes are served in-region, which forward to the US.

Metrics to Watch: edge.ttfb_p50 by country, backbone.rtt_ms, dns.failover_events, region.capacity_headroom

Organizational Follow-up: edge team owns phase 1; service teams own deciding which routes are region-local; the data platform owns replica freshness guarantees that make read routing safe.

Ownership Question: "Who decides a route can be served from a read replica in Singapore?" Staff answer: The owning service team, against freshness requirements they sign off on. The edge only routes; it doesn't know whether stale data is acceptable.

Key Takeaway: "Terminating TLS near the user is the cheapest large latency win in a multi-region rollout. Move the handshake first, the reads second and the writes last."

What clears the Staff bar:

  • Quantifies handshake RTT savings from edge termination alone
  • Sequences the rollout by value and reversibility
  • Leaves data-freshness decisions with the owning teams

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Distinguish edge, east-west and L4 load balancing and put each decision at the layer with the right information
  • Locate latency correctly: sub-millisecond proxy hops vs round-trip-dominated TLS and TCP setup
  • Size an edge fleet on both requests and new connections, including a zero-resumption scenario
  • Explain why P2C least-request beats least-connections, and why it still needs outlier detection
  • Design health checking that can't empty a pool: shallow active checks, passive ejection, caps and panic mode
  • Treat config as a deploy with staged rollout, automatic rollback and data-plane independence from the control plane
  • Handle long-lived connections with max age, GOAWAY, per-request balancing and negotiated reconnects
  • Plan regional failover as a capacity question with explicit headroom

The Bar for This Question#

Mid-level (L4): Explains what a load balancer does, names round robin and least connections, and adds health checks and a redundant pair. Understands L4 vs L7 in textbook terms.

Senior (L5): Builds a sensible two-tier design with TLS termination, health checks, autoscaling and sticky sessions where needed. Knows the products. Misses how the balancer amplifies failures: herding, sinkholes, correlated health checks, config pushes.

Staff+ (L6): Starts from where latency and failure actually come from. Puts numbers on the TLS cost, chooses P2C with outlier detection and explains why, bounds every automated decision, ships config like code, and handles connection lifetime explicitly. Assigns ownership between the edge platform and service teams. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 The Load Balancer Hop Is Never Your Latency Problem#

ClaimEvidence
A well-run L7 proxy adds ~0.1–1 msRequest parsing and routing are microseconds; the in-AZ network hop is ~0.1–0.5 ms
Connection setup costs 1–3 RTTsAt 80–180 ms RTT, that's 160–540 ms for a new user
Removing a proxy hop rarely moves p99Tail latency comes from backends, queuing and retries

The Staff position: Optimize connection reuse, TLS resumption and edge proximity before debating proxy hops.

Why this matters in interviews: Candidates who argue against an L7 tier "for latency" reveal they haven't measured where the time goes.

10.2 Least-Connections Should Be Retired as a Default#

ClaimEvidence
It herds with many independent balancersEvery balancer sees a new host as least-loaded simultaneously
It rewards fast failureA host returning instant 503s has the fewest active requests
Connection count isn't load under HTTP/2One connection may carry hundreds of concurrent streams

The Staff position: P2C least-request plus outlier detection as the default; least-connections only for a single balancer with uniform, connection-per-request traffic.

Why this matters in interviews: It's the clearest single place where a Senior's reasonable answer has a known production failure mode.

10.3 Deep Health Checks Cause More Outages Than They Prevent#

ClaimEvidence
Deep checks correlate failuresEvery host calls the same dependency; a blip fails them all at once
Real traffic is a better correctness signalPassive outlier detection sees the actual errors users see
The LB can't fix a dependency outage anywayEjecting every host replaces partial errors with total ones

The Staff position: Health endpoints check the process, not its dependencies. Correctness comes from outlier detection; dependency problems surface as errors, honestly attributed.

Why this matters in interviews: Interviewers who've lived through a correlated ejection outage will recognize this immediately.

10.4 Sticky Sessions Are a State-Placement Bug#

ClaimEvidence
Stickiness exists to keep server-local stateWhich is lost when that server dies or deploys anyway
It worsens load balanceLong sessions pin load to old hosts after scale-out
It slows deploysDraining waits for sessions to end, or breaks them

The Staff position: Move session state to a shared store; use bounded-load consistent hashing only as a cache-affinity optimization, never for correctness.

Why this matters in interviews: Pushing back on a requirement and offering the real fix is a Staff behavior interviewers look for.

10.5 Your Edge's Biggest Risk Is Its Config Pipeline#

ClaimEvidence
Data planes are mature and redundantStateless proxy fleets survive individual failures routinely
Config reaches every proxyA bad change is correlated across the entire fleet by construction
Public post-mortems repeatedly feature global config pushesLarge edge and CDN outages are frequently triggered by a single config or rule change

The Staff position: Invest in staged rollout, automatic rollback and CPU-cost testing of config before the next data-plane optimization.

Why this matters in interviews: It moves the conversation from boxes to change management — where real availability is won or lost.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

The Staff engineer designs a safe edge. The Principal engineer notices that a single request from a user's phone passes through five independently owned balancing decisions: GeoDNS run by the network team, a cloud L4 balancer, the edge proxy fleet, an API gateway owned by another platform team, and a sidecar mesh owned by a third. Each has its own health model, its own retry policy and its own config pipeline. A request can be retried at four layers; a backend can be ejected by three different health systems that disagree. The L7 problem is traffic-management governance: how many layers the org needs, which policies are set once, and who owns the end-to-end behavior of a request that no single team sees whole.

The Org-Level Fault Line#

One traffic platform vs layered teams vs per-service choice.

OptionWhat WorksWhat BreaksWho Pays
Each layer owned independentlyTeams move fast in their own layerRetries multiply across layers; conflicting health models; five config pipelines to make safeUsers (amplified incidents); on-call (cross-team debugging)
One traffic platform owns edge, gateway and meshOne policy model, one config pipeline, end-to-end visibilityVery large team; edge, gateway and mesh have different skills and cadencesPlatform org (scope); services (one bottleneck for changes)
Separate data planes, one policy layerEach team runs its data plane; timeouts, retries, outlier rules and config safety are defined onceRequires a shared policy schema and real enforcementPlatform architecture (the policy contract); teams (conformance)

🧭 Principal Move: "I don't need one team running every proxy. I need one answer to 'how many times can a request be retried, by whom, and with what budget' — and one standard for how config reaches a proxy. The data planes can stay with the teams that know them; the policy and the change pipeline become shared."

Cost Model#

Assumptions: cloud-hosted, 3 AZs per region, fully loaded engineer ~$250K/year, average response 15 KB, ~40% of connections new at peak. Egress is billed separately and usually dwarfs balancer cost; it is excluded from infra but noted.

ScaleTrafficInfra ($/month)HeadcountOn-call LoadDominant Cost Driver
Startup~2K RPS, 1 region~$500–2K (managed L7 balancer)0.25 eng (part of infra)Shared rotation; balancer itself rarely pagesManaged per-hour + per-capacity-unit charges
Growth~50K RPS, 2 regions~$15–40K (managed or ~20 proxy instances + L4)2–4 eng edge teamDedicated rotation, 2–4 pages/month, mostly config and certsHandshake CPU at peak; cross-AZ traffic from poor zone routing
Enterprise~1M RPS, 3–5 regions, anycast~$150–400K (L4 + L7 fleets, dedicated pools, backbone)10–20 eng across edge, L4/network, config pipeline, certsFollow-the-sun; per-region rotationsEgress and backbone; L7 fleet sized for zero-resumption peaks

The pricing insight: at enterprise scale, the biggest controllable line items are not proxies — they're handshake headroom and cross-AZ transfer. Switching to ECDSA with high resumption can shrink the L7 fleet by 30–50%; capacity-weighted zone routing can cut cross-AZ transfer for the edge by half. Each is worth more than any proxy tuning, and each takes one or two engineers a quarter.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Anycast vs DNS-based global steering (and acquiring address space, BGP presence)One-way-ishMonths: IP allocations, peering, client hardcoding of IPs
Public hostnames and IPs customers allowlistOne-wayCustomer coordination; enterprise contracts list them
TLS termination point (edge vs passthrough to services)One-way-ishCert ownership, compliance scope and service code all change
Session affinity promised to a productOne-wayRemoving it requires moving state the product now depends on
Balancing algorithm defaultTwo-wayConfig template change, staged
Outlier and panic thresholdsTwo-wayConfig, per cluster
Managed vs self-run L7Two-way (if config is abstracted)1–2 quarters of migration
Max connection age, drain windowsTwo-wayConfig

The Standard I'd Write#

RFC-TRAFFIC-001: Traffic Management Standard
Status: Approved   Owners: Edge Platform + Service Mesh Platform + SRE

Scope
  Every balancing layer handling production requests: global steering,
  L4, edge L7, API gateway, service mesh, and client libraries.

MUST
  1. Default balancing is least-request with power-of-two choices. Hash-based
     balancing MUST use bounded load (max 1.25x average).
  2. Passive outlier detection is enabled with max ejection <= 20% of a pool;
     panic threshold is 50% healthy.
  3. Health endpoints used for load balancing MUST NOT call remote shared
     dependencies.
  4. Exactly one layer retries a given request. Retries apply only to
     idempotent requests and are capped by a retry budget of 20% of traffic.
  5. Every config change to any balancing layer rolls out in stages
     (canary, zone, region, global) with automated rollback. No exempt classes.
  6. Data planes continue serving on last-known-good config when the control
     plane is unavailable.
  7. Edge certificates are provisioned by the platform with ECDSA and RSA
     fallback; ticket keys rotate with overlap.

SHOULD
  1. Cap client connection age (10 min +/- jitter) via GOAWAY.
  2. Use capacity-weighted zone-aware routing.
  3. Use dedicated pools for long-lived connection workloads above 50K connections.

Exceptions
  Filed with Edge Platform; reviewed within 5 business days; time-boxed to
  2 quarters; exceptions to MUST 4 and MUST 5 need SRE sign-off.

Success metrics
  - Incidents where a balancing layer amplified a smaller failure: trending to 0
  - Config-change-caused edge minutes of impact: < 25% of edge error budget
  - Edge-added latency p99 (excluding client RTT): < 5 ms
  - TLS resumption rate at peak: > 50%
  - Services with more than one retrying layer: 0

What I'd Tell the VP#

"Our biggest outages at the edge haven't come from machines failing — they've come from our own traffic systems overreacting: a config change that reached every server at once, health checks that removed healthy servers, retries piling up across layers. I'm proposing one shared standard for how all of our traffic layers behave and one safe path for changing them. It's about three engineers for two quarters. We expect it to remove our most common class of large edge incident, and the TLS and zone-routing changes that come with it should reduce edge infrastructure cost by roughly a third. Teams keep ownership of their own systems; what changes is that they follow the same rules."

Principal Interview Signals#

SignalWhat It Sounds Like
Sees the whole request path"This request crosses five balancing layers owned by four teams. Which of them is allowed to retry?"
Prices the levers"ECDSA and resumption save a third of the fleet; that's worth more than any proxy we could tune."
Names one-way doors"Customer-allowlisted IPs and anycast space are the decisions I'd slow down on."
Sets org failure posture"No config class skips canary — and edge config changes freeze when they've eaten a quarter of the error budget."
Knows when not to centralize"Mesh and edge keep separate data planes and teams; only the policy and pipeline are shared."

Staff answers that L7 interviewers find insufficient:

  • "I'd configure outlier detection and panic thresholds on the edge" — correct for one layer, but ignores the gateway and mesh making contradictory decisions on the same requests.
  • "Retries happen at the edge with a budget" — doesn't ask whether the mesh and client libraries also retry, multiplying the budget.
  • "We'll add a region for resilience" — no capacity headroom priced, no failover decision owner, no sequencing by cost.

Appendices

Appendix A: Mechanics in Depth#

A.1 Power of Two Choices with Least Request#

pick_host(cluster):
    hosts = cluster.healthy_hosts()
    if len(hosts) < cluster.panic_threshold * len(cluster.all_hosts()):
        hosts = cluster.all_hosts()                    # panic: ignore health
    a, b = random_sample(hosts, 2)
    score(h) = (h.active_requests + 1) / h.effective_weight()
    return a if score(a) <= score(b) else b

effective_weight(h):
    w = h.configured_weight
    if h.age_seconds < slow_start_window:              # e.g. 60 s
        w = w * max(0.1, h.age_seconds / slow_start_window)
    return w

Why it's right: each proxy needs only its own in-flight counts — no coordination. Randomly sampling two hosts gives near-optimal balance even when counts are stale or partial, and different proxies rarely choose the same host at the same instant, so there's no herd. Why it can go wrong: a host that completes requests instantly (because it fails) always has low in-flight counts. Outlier detection is the necessary partner.

A.2 Outlier Detection#

on_response(host, status, latency):
    if status >= 500 or timed_out:
        host.consecutive_errors += 1
    else:
        host.consecutive_errors = 0
    if host.consecutive_errors >= 5 and can_eject(cluster):
        eject(host, duration = base_ejection * host.times_ejected)   # 30 s, 60 s, 90 s...

every 10 s:  # success-rate outliers, only if pool has >= 5 hosts with >= 100 requests
    mean, stdev = success_rates(cluster)
    for host where host.success_rate < mean - 1.9 * stdev:
        if can_eject(cluster): eject(host)

can_eject(cluster):
    return ejected_count(cluster) < max(1, 0.10 * len(cluster.all_hosts()))

The relative check (vs peers) is what makes outlier detection safe under systemic failure: if every host's success rate drops together, none is an outlier.

A.3 Consistent Hashing Variants#

VariantMechanismDisruption on ChangeBalanceUse
Ring hashHosts placed at many virtual points on a ring; key goes to next point clockwise~1/N keys moveNeeds ~100+ virtual nodes per host for evennessGeneral affinity
MaglevFixed lookup table (prime size, e.g. 65,537) filled via per-host permutationsSlightly more than minimal; very fast lookupVery evenL4 flow hashing
Rendezvous (HRW)Score every host with hash(key, host); pick maxMinimalEven; O(N) per lookupSmall pools, simplicity
Bounded loadAny of the above, plus: skip hosts above c × average loadSome affinity lost under skewBounded by c (e.g. 1.25)Hot keys

See Consistent Hashing for the underlying theory.

Appendix B: Routing Keys and Identity#

KeyUsePitfall
5-tuple (src IP, src port, dst IP, dst port, proto)L4 flow hashingChanges on client reconnect; QUIC migrates connections across addresses — use QUIC connection IDs
Source IPLegacy affinityNAT concentrates thousands of users on one IP
Cookie (LB-issued)Session affinity at L7Pins load; survives only while host lives
Header (user_id, tenant_id)Cache affinity, tenant poolsHot tenants; requires trusted header injection
Path / hostService routingRegex cost; overlapping routes need explicit precedence

Precedence rule: the most specific route wins (exact host > wildcard host; longest path prefix), and the config validator rejects ambiguous overlaps at commit time rather than letting proxies decide at runtime.

Appendix C: Coordination Mechanisms#

C.1 Endpoint Propagation#

Service discovery publishes endpoint sets; the control plane pushes incremental updates (deltas, not full snapshots) to proxies over a streaming API (e.g., xDS). Requirements: updates propagate in < 5 s at p99; proxies ACK/NACK each version so the control plane knows which proxies run which config; the control plane batches churn (a 400-host deploy shouldn't produce 400 pushes per proxy). See Service Discovery.

C.2 Quick Comparison#

MechanismProtects AgainstDoesn't Protect AgainstCost
Consistent hashing at L4Flow resets on L4/L7 fleet changesResets when the target proxy itself diesLookup table per node
Connection trackingFlow remapping during backend set changesTracking-table loss on L4 node failureMemory per flow
Outlier detectionIndividual bad hostsSystemic failures (by design)Per-host counters
Panic thresholdHealth signals that falsely empty a poolReal total outagesNone
Slow startCold hosts overloaded on joinHosts that are slow foreverLonger scale-out ramp
Retry budgetLoad amplification from retriesNon-idempotent retry hazards (handled by idempotency)Some requests not retried
Staged config rolloutGlobal bad configBad config that only fails under peak loadSlower changes (10–60 min)

Appendix D: API Contract & Client Behavior#

BehaviorRule
KeepaliveClients reuse connections; edge idle timeout ~60–120 s; backend pools idle timeout shorter than backend's own
Max connection age10 min ± jitter via GOAWAY (HTTP/2) or Connection: close (HTTP/1.1)
RetriesIdempotent methods only (or requests with an idempotency key); one retry at the edge; retry budget 20%
Retry-After / 503Edge honors backend overload signals and does not retry those requests
0-RTTAccepted only for safe methods (GET/HEAD) to replay-tolerant routes
WebSocket drainServer sends reconnect message; client reconnects with 0–300 s jitter and exponential backoff on failure
Client timeoutsMobile SDK: connect 10 s, request 30 s, exponential backoff with jitter on handshake failures

Appendix E: Observability#

Core metrics:

MetricWhy It MattersAlert
lb.added_latency_p99Edge's own contribution, excluding upstream> 5 ms for 5 min
lb.upstream_5xx_rate (per cluster, per host)Backend health as users see itPer-service SLO burn
lb.no_healthy_upstreamPool emptied — LB-caused outageAny non-zero
lb.hosts_ejected_pctEjection pressure> 10% for 5 min
lb.healthy_hosts_pct (fleet-wide correlation)Correlated health failuresMany clusters dropping together
tls.resumption_rateHandshake cost driver< 50% of baseline
tls.full_handshakes_per_secCapacity driver> 2× baseline
lb.retry_rateAmplification> 20% of requests
config.version_skewProxies on different config versions> 1 version for > 15 min outside rollouts
xds.config_age_secondsControl-plane staleness> 5 min

Control plane vs data plane. The data plane must keep routing on last-known-good config when the control plane is down; endpoints going stale is tolerable for minutes because outlier detection ejects dead hosts from real traffic. Test it: kill the control plane in staging during a deploy.

Debugging the silent failure. Aggregate dashboards hide per-zone and per-host skew. Always keep: per-host RPS distribution (max/mean), per-zone utilization, and requests by config version. "Which config version served this request?" should be answerable from access logs.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 5K RPSOne managed L7 balancer, DNSRouting flexibility; per-service policies
5K–100K RPSManaged L4 + self-run L7 fleet with templated policyHandshake CPU; config safety; zone imbalance
100K–1M RPSOwn L4 (XDP, consistent hash, DSR), anycast, dedicated poolsMulti-region capacity; cross-layer retries
1M+ RPSGlobal steering on capacity, one policy layer across edge, gateway and meshOrg coordination more than technology

Multi-region path: GeoDNS with health-based failover → edge termination in more regions with backbone to central backends → regional backends for reads → anycast with capacity-aware steering.

What You Don't Build on Day One: your own L4 data plane; anycast; a custom proxy; per-request adaptive algorithms beyond P2C; cross-region active-active steering. Each is justified by a measured problem, not anticipated scale.

Appendix G: Multi-Tenancy, Fairness & Cost#

Per-service limits on shared proxies (prevent one backend from exhausting the proxy):

LimitTypical DefaultEffect When Hit
Max connections to cluster1,024 per proxyNew requests queue or fail fast
Max pending requests1,024Fast 503 instead of unbounded queueing
Max concurrent requests1,024 (HTTP/2)Same
Max concurrent retries3 per proxy, or budget-basedRetries dropped
Per-route timeout ceiling≤ listener max (e.g. 60 s)Config rejected

Cost attribution. Charge services for what they consume at the edge: full handshakes (CPU), requests, egress bytes and long-lived connection-hours. The first report usually shows one or two services dominating handshake cost — typically mobile clients with short keepalives — and fixing their client config pays for the attribution work.

  1. Loading the index…