Hiring BarSupport

Network Latency & Protocols

Foundation34 min read5 diagrams

Why This Matters#

Networking is not a protocol-trivia question. Nobody gets leveled Staff for reciting the TCP state machine. Networking is a latency budget question: every request has a fixed number of milliseconds before the user notices, and every hop, handshake, retry, and proxy spends some of them. The Staff question is who spends each millisecond, and who pays when the tail blows up.

Most candidates draw arrows between boxes as if arrows were free. They are not. An arrow inside a rack costs ~50µs. An arrow across availability zones costs ~1ms and $0.01 per GB in each direction. An arrow across an ocean costs ~80ms of physics that no amount of engineering will remove. A fresh HTTPS connection across that ocean costs three or four of those round trips before the first byte of payload moves. Candidates who know this draw fewer arrows, reuse connections, and put state close to the user. Candidates who don't design systems that are correct and slow.

The second reason networking shows up in Staff loops is tail latency. A service with a p99 of 10ms is fine in isolation. Put it behind a fan-out of 100 parallel calls and 63% of user requests now experience that p99. The p99 of a leaf becomes the median of the page. This is the single most under-appreciated number in system design, and it is where interviewers separate people who have operated fan-out services from people who have read about them.

If you can walk an interviewer from a latency budget ("the page has 200ms; 80ms is physics") to a protocol choice ("so we terminate TLS at the edge and hold warm HTTP/2 connections to origin") to a tail-latency mitigation ("hedge after p95, capped at 5% extra load") — and name who owns each piece — you are demonstrating the reasoning Staff interviews reward.

The 60-Second Version#

  • Physics sets the floor. Light in fiber travels ~200km per millisecond. US East ↔ US West is ~60–70ms RTT, transatlantic ~70–90ms, US ↔ Asia ~150–200ms. You cannot optimize below this; you can only avoid paying it more than once.
  • Handshakes multiply RTT. A cold HTTPS request costs TCP (1 RTT) + TLS 1.3 (1 RTT) + request (1 RTT) = 3 RTT; TLS 1.2 makes it 4. QUIC folds transport and crypto into 1 RTT, and 0-RTT resumption gets repeat visitors to 1 RTT total — at the cost of replay risk on non-idempotent requests.
  • Connection reuse is the single biggest latency win in most systems. A warm connection turns 3–4 RTT into 1 and skips TCP slow start (~14KB initial window, doubling per RTT).
  • HTTP/2 fixed application-level head-of-line blocking but kept TCP's. One lost packet stalls every multiplexed stream. HTTP/3 over QUIC removes transport HOL blocking — it matters most on lossy mobile networks (1–3% loss), barely at all inside a data center.
  • Long-lived connections break L4 load balancing. gRPC over HTTP/2 pins all requests on one connection to one backend. You need L7 (per-request) balancing, client-side balancing, or forced connection recycling (e.g., max connection age ~5–30 min).
  • Tail latency compounds under fan-out. P(slow page) = 1 − 0.99^N for N parallel calls. N=100 → 63%. Mitigations: hedged requests, tied requests, deadline propagation, and reducing N.
  • Retries multiply across layers. 3 attempts at each of 3 layers = up to 27× load on the bottom service during an incident. Retry once, at one layer, with a retry budget (~10% of traffic).

How Networking Works (for System Designers)#

The Latency Stack#

Every remote call pays some combination of these costs. The design question is which ones you pay per request versus once per connection versus never.

CostTypical ValuePaid WhenHow to Avoid It
DNS resolution0–100ms (cached: ~0)First request to a hostname, after TTL expiryLong-ish TTLs (60–300s), connection reuse, prefetch
TCP handshake1 RTTEvery new connectionKeepalive, connection pooling
TLS 1.3 handshake1 RTT (0 with resumption)Every new connectionSession resumption, TLS termination near the client
TLS 1.2 handshake2 RTTEvery new connectionUpgrade to 1.3
TCP slow startSeveral RTT for large responsesEvery new (or long-idle) connectionWarm connections, larger initcwnd at the edge
Serialization10µs–5msEvery requestProtobuf/binary formats on hot paths
Propagation~5µs/km one way (≈1ms RTT per 100km)Every requestMove compute/data closer; you cannot beat light
Queueing0 → unboundedEvery request under loadKeep utilization < ~70%; backpressure
Proxy hop (L7)0.2–2ms eachEvery requestFewer hops, sidecar-less where it matters

Reference Round-Trip Times#

PathRTTWhat It Means
Same host (loopback)~10–50µsSidecar hops are cheap but not free
Same rack / same AZ~0.1–0.5msChatty protocols are tolerable
Cross-AZ, same region~0.5–2msSynchronous replication across AZs is affordable
US East ↔ US West~60–70msOne synchronous cross-region call per request, at most
US East ↔ Europe~70–90msConsensus across the Atlantic costs this per write
US ↔ Asia-Pacific~150–200msSynchronous anything here is a product decision, not a technical one
Mobile last mile (4G)~30–100ms, 1–3% lossThe client link often dominates the whole budget

Where Protocols Live#

Application   HTTP/1.1 · HTTP/2 · HTTP/3 · gRPC · WebSocket · SSE
Security      TLS 1.2 / 1.3  (built into QUIC for HTTP/3)
Transport     TCP (reliable, ordered byte stream)  ·  UDP / QUIC (datagrams; QUIC adds streams + reliability)
Network       IP · anycast · BGP

L4 vs L7 is the distinction that matters most in design interviews:

LayerSeesBalancesCost per RequestTypical Tools
L4 (transport)IP, port, TCP/UDP flowsConnections~µs; millions of packets/s per nodeAWS NLB, Google Maglev, IPVS, Katran
L7 (application)HTTP method, path, headers, gRPC serviceRequests~0.2–2ms; tens of thousands of RPS per coreEnvoy, NGINX, HAProxy, AWS ALB

L4 is fast and dumb; L7 is slower and smart. L7 terminates TLS, can route by path or tenant, retry, rate limit, and balance per request. L4 forwards bytes and balances per connection — which is precisely why it mis-balances long-lived connections.

Core Strategies#

Strategy 1: Connection Reuse and Pooling#

The cheapest round trip is the one you never make. Keep connections warm and share them.

# Per-process pool to each upstream
pool = ConnectionPool(
    upstream       = "orders.internal:443",
    max_conns      = 64,          # per upstream host
    idle_timeout   = 55s,         # shorter than the LB's idle timeout (e.g., ALB default 60s)
    max_conn_age   = 10m,         # force periodic rebalancing across backends
    keepalive_ping = 30s,         # detect dead peers before a request finds them
)

call(req):
    conn = pool.acquire(deadline = req.deadline)   # never block past the caller's deadline
    return conn.send(req)

When to use: Always, for service-to-service traffic. The question is only how many connections and how long they live.

Failure mode: Idle-timeout mismatch. If the client keeps a connection idle for 90s and the load balancer drops it at 60s, the next request on that connection gets a reset. Symptom: a small, steady rate of connection reset by peer errors that correlates with low-traffic periods. Fix: client idle timeout strictly below every middlebox's idle timeout.

Strategy 2: Terminate Close, Travel Warm#

Put TLS termination at an edge PoP near the user (CDN or regional edge), then carry the request to origin over a pre-established, long-lived connection.

SetupUser in Sydney → Origin in Virginia (~200ms RTT)Time to First Byte
Direct, TLS 1.3, cold3 RTT × 200ms~600ms + server time
Edge in Sydney (~10ms RTT) + warm backbone to origin2 × 10ms handshake + 1 × 200ms~220ms + server time
Edge + cached response2 × 10ms + 0~20ms

When to use: Any global consumer product. This is why every CDN also sells "dynamic acceleration."

Failure mode: The edge becomes a hidden dependency. An edge configuration push that breaks TLS takes down every origin behind it — a correlated failure across teams who never talk to each other.

Strategy 3: Pick the Protocol for the Traffic Shape#

Traffic ShapeProtocolWhy
Public request/response APIHTTP/1.1 or HTTP/2 + JSONUniversal tooling, cacheable, debuggable with curl
Internal service-to-service RPCgRPC (HTTP/2 + protobuf)Schemas, codegen, deadlines, streaming, ~3–10× smaller payloads than JSON
Browser/mobile on lossy linksHTTP/3 (QUIC)No transport HOL blocking, connection migration across Wi-Fi ↔ cellular
Server → client push, one directionServer-Sent EventsPlain HTTP, auto-reconnect with Last-Event-ID, works through most proxies
Bidirectional, low-latency, interactiveWebSocketFull duplex over one TCP connection; you own reconnect, heartbeats, and fan-out
Infrequent updates, simplest possibleLong pollingWorks everywhere; ~1 request per update, wasteful at high rates
Loss-tolerant real-time mediaUDP / WebRTCLate data is worthless; retransmission would make it worse

Failure mode: Choosing WebSockets for a feature that pushes one update a minute. You now operate a stateful connection tier — sticky routing, reconnect storms on deploy, per-connection memory (~10–50KB each) — for traffic that SSE or polling would have handled statelessly.

Strategy 4: Deadlines, Not Timeouts#

A timeout is local ("I'll wait 2s"). A deadline is global ("this whole request must finish by 12:00:00.250"). Propagate the deadline and every hop can decide whether it is still worth doing work.

handle(req):
    remaining = req.deadline - now()
    if remaining < MIN_USEFUL_WORK:        # e.g., 5ms
        return DEADLINE_EXCEEDED           # fail fast; don't burn downstream capacity
    child_deadline = req.deadline - SAFETY_MARGIN     # leave time to respond
    a = call(inventory, deadline = child_deadline)
    b = call(pricing,   deadline = child_deadline)
    return merge(a, b)

When to use: Any call graph deeper than two hops. gRPC propagates deadlines natively; HTTP stacks need a header convention.

Failure mode: Uncoordinated static timeouts. The edge times out at 1s, the service at 2s, the database client at 5s. When the database slows down, the edge has already given up, but the service and database keep working for 4 more seconds on requests nobody is waiting for. That wasted work is what turns a slowdown into an outage.

Strategy 5: Retry Once, at One Layer, With a Budget#

retry_policy = {
    max_attempts:   2,                 # original + 1 retry
    retry_on:       [UNAVAILABLE, RESET, 503],   # never on 4xx or DEADLINE_EXCEEDED
    backoff:        exponential(base=25ms, cap=250ms, jitter=full),
    retry_budget:   0.10,              # retries ≤ 10% of requests over a 10s window
    idempotent_only: true,
}

Failure mode: Retry amplification. Three layers each retrying three times turns one failing request into 27 attempts at the bottom. The bottom service was slow because it was overloaded; now it receives 27× the load. Retry budgets cap the damage at ~1.1× instead of 27×.

Tail Latency: The Hard Sub-Problem#

Averages lie; tails compound. The single most important formula in this page:

P(request sees at least one slow leaf) = 1 − (1 − p_slow)^N

p_slow = 1% (i.e., each leaf's p99), N parallel leaves:
  N = 1     →  1.0%
  N = 10    →  9.6%
  N = 50    → 39.5%
  N = 100   → 63.4%
  N = 1000  → 99.99%

At N=100, the leaf's p99 is the page's median experience. This is why a search backend, a feed aggregator, or a scatter-gather query over 100 shards cannot be fixed by making the average faster. You must attack the tail directly.

Where Tails Come From#

SourceTypical MagnitudeSignature
GC pauses10–500ms (JVM), less with modern collectorsPeriodic spikes per host, uncorrelated across hosts
Queueing at high utilizationGrows as ~1/(1−ρ); at 90% utilization queue delay is ~10× service timep99 climbs sharply past ~70% CPU
TCP retransmissionMinimum RTO ~200ms on LinuxA cluster of requests at exactly +200ms
Delayed ACK + Nagle interaction~40msSmall writes stall at exactly ~40ms; fixed by TCP_NODELAY
Noisy neighbors2–10× service timeTail varies by host, not by request
Cold caches / cold connections1–3 extra RTTTail spikes right after deploys or scale-out
Background work (compaction, backups)10–100msTail correlates with maintenance windows

Mitigation Toolbox#

TechniqueHow It WorksExtra LoadWhen It Breaks
Hedged requestsSend a second copy to another replica if no response by the p95~5% (only the slowest 5% are hedged)Non-idempotent calls; correlated slowness (both replicas share the cause)
Tied requestsSend to two replicas at once; the first to start cancels the other~0–2% wasted workNeeds cancellation plumbing in the server
Reduce fan-outFewer, fatter shards; pre-aggregated indexesNegativeShard size limits, rebalancing cost
Partial resultsReturn what arrived by the deadline; mark the response incompleteNoneProduct must accept "good enough" answers
Utilization capKeep leaf CPU below ~60–70%Capacity cost (~30–40% headroom)Finance asks why the fleet is "idle"
Micro-partitioning10–100 partitions per server so load moves in small unitsMetadata overheadCoordination complexity

🎯 Staff Move: "This page fans out to about 40 services, so a 1% tail at each leaf means roughly a third of page loads see at least one slow call. I'd rather not fix that at each leaf. I'd propagate a 150ms deadline, hedge the three read-only calls at their p95, and render partial results for the recommendations module — product has to sign off that a page without recommendations is acceptable."

Full reasoning: why hedging at p95 costs only ~5%

If you send a hedge only when the first attempt has not returned by the p95 latency, then by definition only ~5% of requests trigger a hedge. The extra load is ~5%, not 100%. The payoff is large because the slow 5% of first attempts are usually slow for reasons local to that replica (GC, queueing, a noisy neighbor), so the hedge to a different replica usually completes near the median.

Two rules keep it safe:

  1. Only hedge idempotent reads. Hedging a payment capture is how you double-charge someone.
  2. Cancel the loser. Without cancellation, a hedged system under overload does 1.05× the work and the losing requests still occupy queue slots.

Hedging fails when slowness is correlated — for example, a slow shared dependency behind both replicas. Then both attempts are slow and you have added 5% load to an already-overloaded system. Watch hedge.win_rate: if hedges rarely win, turn them off.

Visual Guide#

A Cold HTTPS Request vs a Warm One#

Diagram: A Cold HTTPS Request vs a Warm One

Handshake Cost by Protocol#

ProtocolNew ConnectionResumed ConnectionCaveat
TCP + TLS 1.21 TCP + 2 TLS + 1 request = 4 RTT to response1 TCP + 1 TLS + 1 request = 3 RTTStill common on legacy clients
TCP + TLS 1.31 TCP + 1 TLS + 1 request = 3 RTT to response0-RTT data: 2 RTT (TCP still costs 1)0-RTT data can be replayed
QUIC (HTTP/3)1 combined handshake + 1 request = 2 RTT to response0-RTT: request in the first flight = 1 RTTUDP is blocked on a small share of networks; clients fall back to TCP

Choosing a Real-Time Transport#

Diagram: Choosing a Real-Time Transport

L4 vs L7 in a Typical Topology#

Diagram: L4 vs L7 in a Typical Topology

The pattern: L4 spreads connections across a stateless L7 tier cheaply; L7 spreads requests across services intelligently. Services talk to each other over gRPC with client-side or sidecar balancing so no single long-lived connection pins load to one backend.

Implementation Patterns#

HTTP/1.1 vs HTTP/2 vs HTTP/3#

DimensionHTTP/1.1HTTP/2HTTP/3
TransportTCPTCPQUIC over UDP
Concurrency1 request in flight per connection; browsers open ~6 per originMany multiplexed streams on 1 connection (default max ~100 concurrent streams)Many independent streams
Head-of-line blockingApplication-level (requests queue behind each other)Transport-level (one lost packet stalls all streams)None at transport level (loss stalls only the affected stream)
Header compressionNoneHPACKQPACK
HandshakeTCP + TLSTCP + TLS (ALPN negotiates h2)Combined 1 RTT; 0-RTT resumption
Connection migrationNoNoYes (connection IDs survive IP changes)
Where it winsSimplicity, debugging, legacyService-to-service in low-loss DCsMobile and lossy last mile

Staff default: HTTP/2 (gRPC) inside the data center; HTTP/3 at the edge for mobile and web clients, with TCP fallback. HTTP/3 inside the data center buys little — loss is near zero, and UDP processing is typically more CPU-expensive per byte than kernel-optimized TCP.

gRPC and the Long-Lived Connection Problem#

gRPC multiplexes every call over a small number of HTTP/2 connections. Behind an L4 load balancer, that means:

Client pods: 20, each with 1 connection
Server pods: 10 → 20 after autoscaling
Result: the 10 new servers receive 0 connections and 0 traffic
        until clients reconnect — which may be never.

Three fixes, in order of preference:

  1. Client-side load balancing — the client resolves all backend addresses (via DNS or a discovery service), keeps connections to each, and balances per call (round robin or least-outstanding-requests).
  2. L7 proxy / sidecar — Envoy balances per request across backends; the client holds one connection to its local proxy.
  3. Max connection age — servers send GOAWAY after 5–30 minutes so clients reconnect and redistribute. Cheap, crude, and a good backstop even when you do 1 or 2.

WebSockets and SSE at Scale#

ConcernNumberDesign Consequence
Memory per idle connection~10–50KB (kernel buffers + app state)1M connections ≈ 10–50GB across the fleet
Connections per well-tuned server100K–1M+ (file descriptors, ephemeral ports, memory)A gateway tier of tens of hosts for 10M users
Heartbeat interval20–60sMust be shorter than every NAT/LB idle timeout on the path
Reconnect storm on deployAll clients on a host reconnect within secondsDrain gradually; jittered reconnect backoff (1–30s) on the client

Separate the connection tier (holds sockets, stateless about business logic) from the business tier (stateless about sockets). A presence or routing service maps user_id → gateway_host. This is how chat and notification systems stay deployable.

Anycast and DNS-Based Steering#

  • Anycast: the same IP is announced from many PoPs via BGP; the network routes each user to the nearest. Failover is BGP convergence (seconds to a minute). Works best for short-lived or connectionless traffic (DNS, CDN).
  • DNS steering (GeoDNS / latency-based): resolvers get region-specific answers. Failover is bounded by DNS TTL plus resolvers that ignore it — plan for minutes, not seconds.

Failure Scenario: The Cross-Team Retry Storm#

t=0      Profile DB p99 rises from 15ms to 400ms (compaction + a hot partition).
t=+10s   Profile service (2s timeout, 3 attempts) starts retrying. DB load ×2.
t=+20s   Feed service (1s timeout, 3 attempts) times out on Profile and retries. Profile load ×3.
t=+30s   Edge gateway (3s timeout, 2 attempts) retries Feed. Total DB attempts per user request: up to 18.
t=+45s   DB connection pool exhausted. Every caller of Profile fails, including checkout.
t=+2min  Incident declared. Three teams on the bridge, each certain their retries are "reasonable".
t=+9min  Mitigation: gateway retries disabled via config; Profile shed to 50% for non-critical callers.
Diagram: Failure Scenario: The Cross-Team Retry Storm

Detection: rpc.client.retries ÷ rpc.client.requests above 10% on any edge; db.connections.in_use at pool max; rpc.server.deadline_exceeded rising at a service whose own dependencies look healthy. Blast radius: every service that shares the Profile DB, including checkout, which never retried at all. Mitigation: disable retries at the outermost layer first (it has the largest multiplier), shed non-critical callers, then let the DB recover. Prevention: one retrying layer, retry budgets in the shared RPC library, deadline propagation so abandoned requests stop consuming DB connections. Owner: the platform team that owns the RPC library — because no single product team can fix a behavior that emerges from three teams' defaults.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Retry stormrpc.client.retries ratio > 10%Every service sharing the slow dependencyDisable outer-layer retries; shed loadRPC platform team
Idle-timeout mismatchSteady connection_reset at low trafficOne client/LB pairClient idle timeout < LB idle timeoutCalling service
gRPC connection pinningPer-pod CPU skew > 3× after scale-outOne service's capacityMax connection age; client-side LBService owner + platform
Edge/CDN config push breaks TLSedge.tls_handshake_errors spike across all hostsEvery origin behind the edgeAutomated rollback; staged config rolloutEdge/traffic team
WebSocket reconnect storm on deployws.connects_per_sec 10–50× baselineConnection tier + auth serviceGradual drain; jittered client backoffReal-time platform team
Cross-region link degradationrpc.latency.p99 for cross-region calls; packet loss > 1%Features with synchronous cross-region callsFail over reads to local replicas; defer async replicationService owners; network team for the link

The Numbers in Context#

NumberValueWhat It Means for Your Design
Light in fiber~200km/ms one wayEvery 1,000km adds ~10ms RTT. Region placement is a latency decision.
Cross-AZ RTT~0.5–2msSynchronous replication across 3 AZs adds ~1–2ms per write — affordable.
Cross-region RTT60–200msSynchronous cross-region writes put this on every commit. Usually a no.
TCP initial window10 segments ≈ 14KBA 1MB response on a cold connection needs ~6–7 RTT of slow start. At 80ms RTT that is ~500ms.
Bandwidth-delay product1Gbps × 80ms = 10MBA single TCP flow needs a ~10MB window to fill a long fat pipe; default buffers often cap throughput far lower.
L7 proxy hop~0.2–2ms p50, more at p99A service mesh adds 2 hops per call (client + server sidecar). A 10-deep call chain pays 20 hops.
Linux min RTO~200msOne lost packet on a 1ms-RTT link turns a 2ms call into a 200ms call. This is a common source of p99.9.
Browser connections per origin (HTTP/1.1)~6The reason for domain sharding in the HTTP/1.1 era, and the reason HTTP/2 made it an anti-pattern.
Default HTTP/2 max concurrent streams~100 (server-configurable)A client exceeding this queues locally; heavy RPC clients need more connections.
Fan-out tailN=100 at p99 → 63% slowTail latency, not average latency, sets the user experience of any aggregator.
Retry amplification3 layers × 3 attempts = 27×One retrying layer, one retry, a 10% budget.
Cloud cross-AZ transfer~$0.01/GB each directionA chatty service moving 1PB/month across AZs pays ~$20K/month for the privilege.

How This Shows Up in Interviews#

Scenario 1: "Users in Asia say the app is slow. Our servers are in Virginia."#

The interviewer is testing whether you start from physics. Say: "Virginia to Singapore is roughly 200–230ms RTT. A cold HTTPS request is 3 RTT before any server time, so we are spending ~650ms on handshakes alone. Step one is an edge PoP in the region to terminate TLS at ~10ms RTT and carry traffic on warm connections; that cuts ~400ms without moving any data. Step two — only if reads dominate — is a regional read replica or cache so the remaining 200ms round trip disappears for reads. Writes still go to Virginia until someone signs off on a multi-region write story."

Scenario 2: "Design the API layer for a page that calls 30 backend services." (Full Walkthrough)#

Step 1 — Budget the latency. "The product SLO is p99 page render under 300ms. I'll reserve 50ms for the client link and rendering, which leaves 250ms server-side. With 30 parallel calls, if each leaf has a 1% chance of exceeding its p99, about 26% of pages hit at least one slow leaf. So I can't just budget each leaf's p99 at 250ms — I have to design for the tail."

Step 2 — Shape the call graph. "I'll split the 30 calls into critical (8: user, cart, pricing, inventory…) and decorative (22: recommendations, badges, promotions). Critical calls block the response. Decorative calls get a 120ms deadline and render as partial results if they miss — product signs off on that."

Step 3 — Propagate deadlines. "The aggregator sets a 250ms deadline and passes it down via gRPC metadata. Each service subtracts its own safety margin. Anything that would start with less than 10ms left fails fast instead of doing work nobody will read."

Step 4 — Attack the tail on reads. "For the four critical read-only calls with the worst p99/p50 ratio, I'll hedge at their p95. That costs ~5% extra load on those services. I'll put hedge.win_rate on a dashboard; if it drops below ~20%, the slowness is correlated and hedging is only adding load."

Step 5 — Control retries. "Retries happen only at the aggregator, once, for UNAVAILABLE, with a 10% retry budget. Leaf services never retry each other's calls. That caps worst-case amplification at 1.1× instead of compounding."

Step 6 — Name the owners. "The aggregator team owns the deadline and degradation policy. Each leaf team owns its p99 SLO. The platform team owns the RPC library that enforces deadline propagation and retry budgets — so the policy isn't re-implemented 30 times."

Why this is a Staff answer: It turns a vague performance question into a budget, uses the fan-out formula to justify the design, separates critical from decorative work with a product sign-off, and assigns ownership of the tail to the layer that can actually fix it.

Scenario 3: "We moved to gRPC and one backend pod is at 100% CPU while the rest idle."#

This tests the long-lived connection problem. The answer: "gRPC multiplexes everything over a few HTTP/2 connections, and our L4 load balancer balances connections, not requests. After the last scale-out, the old connections stayed pinned. Short-term, set a max connection age of ~10 minutes on the servers so clients reconnect. Long-term, move to client-side balancing with least-outstanding-requests, or put an L7 proxy in the path." Bonus points for noting that the problem gets worse with autoscaling — new pods get no traffic.

Scenario 4: "Should the live-scores feature use WebSockets?"#

This tests whether you pick the simplest transport that works. "Scores are server-to-client only and update every few seconds. That's SSE: plain HTTP, auto-reconnect with Last-Event-ID for resume, CDN-friendly fan-out. WebSocket buys bidirectionality we don't need and costs us a stateful connection tier. If we later add live chat under the scoreboard, that's the moment to reconsider."

Advanced Patterns#

PatternHow It WorksWhen to Use
Hedged requestsDuplicate a slow read to another replica after the p95; cancel the loserFan-out reads with uncorrelated tail causes
Deadline propagationAbsolute deadline travels with the request; each hop checks remaining budgetCall graphs deeper than 2 hops
Retry budgetsRetries capped at a % of successful traffic per clientEvery RPC client — prevents retry storms
Zone-aware routingPrefer same-AZ backends; spill cross-AZ only when local capacity is shortHigh-volume east-west traffic (cost + ~1ms latency)
Connection drainingOn deploy, stop accepting new connections, send GOAWAY, wait for in-flightEvery deploy of a connection-holding tier
Request collapsing at the edgeCDN coalesces concurrent misses for the same object into one origin fetchViral content, cache-miss storms
Outlier ejectionRemove backends whose error rate or latency is N× peers from the pool for 30s+Service mesh default; catches bad hosts before humans do
Adaptive concurrency limitsClient or server lowers in-flight limit when latency rises (TCP-Vegas-style)Protecting services without hand-tuned static limits

The Principal Lens#

Why L7 Sees This Problem Differently#

A Staff engineer tunes the network behavior of their service: pools, timeouts, protocol choice, hedging. A Principal engineer notices that every one of those settings is being chosen independently by 200 teams, and that the resulting emergent behavior — retry storms that cross team boundaries, timeout hierarchies nobody designed, cross-AZ traffic nobody budgeted — is the actual reliability and cost problem. At L7, networking stops being a protocol question and becomes a paved-road question: which behaviors do we bake into the one RPC stack everyone uses, so the organization's default is safe even when a team never thinks about it?

The Org-Level Fault Line#

Centralized RPC/mesh platform vs per-team client libraries.

OptionWhat WorksWhat BreaksWho Pays
Every team picks its own HTTP/gRPC client and settingsAutonomy; no platform bottleneckInconsistent retries and timeouts; incidents cascade across teams; no org-wide viewOn-call of the downstream team that absorbs other teams' retry storms
Shared RPC library per languageConsistent defaults; cheap per callUpgrading 200 services takes quarters; polyglot orgs maintain N librariesPlatform team maintaining 3–5 language ports
Service mesh (sidecar or proxyless)Policy changes without redeploying services; uniform telemetry and mTLS+2 proxy hops per call (~0.5–2ms each); sidecar CPU/memory tax; a new tier-0 dependencyEvery service pays the latency and compute tax; mesh team carries the pager for everything

The Principal position: a thin, mandatory RPC contract (deadlines, retry budgets, mTLS, standard telemetry) enforced by either a shared library or a mesh — and the choice between those two is driven by how polyglot the org is and how much latency the hottest paths can spare. Mandate the contract, not the implementation.

Cost Model#

Assumptions: AWS-style list pricing, cross-AZ transfer $0.01/GB each direction ($0.02/GB round trip), internet egress ~$0.05–0.09/GB, sidecar ~0.1–0.25 vCPU per pod at ~$30/vCPU-month, fully loaded engineer ~$25K/month.

ScaleEast-West TrafficCross-AZ Transfer $Mesh/Proxy Compute $Platform HeadcountOn-Call Load
Startup (30 services, 300 pods)~20TB/month, 60% cross-AZ~$250/month~$1–2K/month0.5 engineer (shared infra)Network issues ~1 page/month
Growth (300 services, 5K pods)~1PB/month, 60% cross-AZ~$12K/month~$20–40K/month3–4 engineers (~$90K/month)Dedicated mesh rotation, ~1 incident/week
Large (2,000 services, 50K pods)~30PB/month; zone-aware routing cuts cross-AZ to ~20%~$120K/month (vs ~$360K without zone-aware routing)~$200–400K/month10–15 engineersTier-0 rotation; mesh is on every incident bridge

The lever most orgs miss: zone-aware routing can pay for most of the networking platform team at large scale (~$240K/month saved in the Large row), because cross-AZ transfer is billed per GB and east-west traffic grows faster than user traffic.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversal Cost
Timeout, retry, and hedging valuesTwo-wayConfig push; minutes
HTTP/2 vs HTTP/3 at the edgeTwo-wayClients fall back automatically
Adopting a service meshMostly one-wayEvery service's traffic, identity, and telemetry now flows through it; exit is a multi-quarter migration
Public API protocol (REST/JSON vs gRPC vs GraphQL)One-wayExternal clients and SDKs pin you for years
Region placement for the primary data storeOne-wayData gravity; moving petabytes and re-pointing every dependency
Exposing WebSocket semantics to third-party clientsOne-wayEcosystem depends on connection behavior, message framing, and reconnect contract
Choice of edge/CDN vendorTwo-way in theory, one-way in practiceEdge logic (workers, rules, WAF) accretes vendor-specific code

The Standard I'd Write#

RFC: Service-to-Service Communication Standard (v1)

Scope: All synchronous internal RPC between production services. Excludes batch data transfer and external/public APIs.

MUST:

  1. Every request carries an absolute deadline; services MUST NOT start work when remaining budget is below their documented minimum.
  2. Retries MUST be limited to 1 additional attempt, only on idempotent methods, only for transport/UNAVAILABLE errors, and within a per-client retry budget of 10%.
  3. All traffic MUST use mTLS with workload identity; plaintext internal traffic is prohibited.
  4. Clients MUST emit standard metrics: rpc.client.latency (histogram), rpc.client.errors by code, rpc.client.retries, rpc.client.deadline_exceeded.
  5. Servers MUST support graceful drain (GOAWAY + in-flight completion within 30s).

SHOULD: prefer same-zone backends; set max connection age 5–30 minutes; hedge only reads whose p99/p50 > 5 and only after the p95.

Exceptions: Filed with the networking platform team; approved exceptions expire after 2 quarters.

Success metrics: zero cross-team retry-storm incidents per quarter; cross-AZ bytes ÷ total east-west bytes < 25%; ≥ 95% of services on the standard client within 3 quarters.

What I'd Tell the VP#

"Our services talk to each other millions of times a second, and today each team decides on its own how long to wait and how often to retry. That's why a slowdown in one service turned into a two-hour outage across checkout last quarter. We want to put one set of rules into the shared plumbing so every team gets safe behavior by default. That needs a platform team of about four engineers. Keeping traffic inside a zone also saves about $100K a year in data-transfer charges today, and that saving grows faster than our traffic. The risk is that the plumbing becomes critical infrastructure, so we'll roll it out gradually and staff it with a real on-call rotation."

Principal Interview Signals#

SignalWhat It Sounds Like
Treats network behavior as org policy"The problem isn't this service's timeout. It's that 200 teams each chose one, and nobody designed the hierarchy."
Prices the east-west graph"At 1PB a month of cross-AZ chatter we're paying ~$20K monthly to cross a line we could route around."
Separates contract from implementation"I'd mandate deadlines, retry budgets, and mTLS. Whether that's a library or a mesh depends on how polyglot we are."
Names correlated-failure risk"The mesh control plane and the edge config pipeline are now shared fate for every team — they need their own blast-radius cells and staged rollouts."
Knows when not to standardize"The trading path and the video path get an exemption — their latency budgets can't afford a sidecar."

Staff answers that L7 interviewers find insufficient:

  • "I'd set our client timeout to 500ms and retry twice." — Correct locally; ignores what every other caller in the graph is doing.
  • "We should adopt a service mesh." — Names a tool without the contract it enforces, the tax it charges, or the exit cost.
  • "Cross-AZ latency is only ~1ms, so it doesn't matter." — True for latency, blind to the transfer bill that grows faster than traffic.

In the Wild#

These are public, documented examples.

Google: "The Tail at Scale"#

Jeff Dean and Luiz André Barroso's 2013 Communications of the ACM paper formalized the fan-out tail problem using Google's own serving systems. It showed how a leaf-level 1-in-100 slow response becomes the common case for a root that fans out to 100 leaves, and introduced hedged requests and tied requests as the practical fixes. The paper reports that deferring a hedge until a request has been outstanding past its 95th-percentile latency sharply cuts the tail while adding only a few percent extra load.

Staff insight: Quote the formula, not the paper. "With 100 leaves, a 1% leaf tail is a 63% page tail" is the sentence that tells an interviewer you have reasoned about aggregators, not just built one.

Google QUIC → IETF HTTP/3#

Google designed and deployed QUIC in Chrome and its own services during the 2010s to remove TCP's head-of-line blocking and cut handshake round trips; the IETF standardized QUIC as RFC 9000 in 2021, and HTTP/3 (RFC 9114) runs on top of it. Major CDNs and browsers now support it, with automatic fallback to TCP when UDP is blocked.

Staff insight: The wins are concentrated where loss and RTT are high — mobile networks and long-haul links. Inside a data center, the benefit is small and the CPU cost of user-space UDP can be higher. Say where you'd use it, not just that it's newer.

Lyft: Envoy and the Service Mesh#

Lyft built Envoy as an L7 proxy to give every service consistent retries, timeouts, circuit breaking, outlier ejection, and observability without re-implementing them in every language. It was open-sourced in 2016 and became a CNCF graduated project and the data plane of several service meshes (e.g., Istio).

Staff insight: Envoy's origin story is an organizational one: polyglot services with inconsistent network behavior. That's the argument to make in an interview — a mesh is justified by the cost of inconsistency across teams, not by features any single service needs.


Staff Calibration#

What Staff Engineers Say (That Seniors Don't)#

ConceptSenior (L5)Staff (L6)Principal (L7)
Latency"We'll add a CDN to make it faster""It's 3 RTT at 200ms before server time. Edge TLS termination cuts ~400ms, then regional read replicas remove the last trip for reads""Region placement is a one-way door. I'd price a second write region against the revenue lost to latency in APAC before committing"
Protocol"gRPC is faster than REST""gRPC internally for schemas and deadlines; HTTP/JSON at the public edge for reach; we need client-side LB because of long-lived connections""The public protocol is a multi-year contract with external developers. Internally I'd standardize the RPC contract, not the library"
Timeouts"Set a 2-second timeout""Propagate a deadline from the edge; each hop fails fast below its minimum useful budget""Deadline propagation goes into the paved-road RPC stack so 200 teams get it without thinking"
Retries"Retry 3 times with backoff""One retry, one layer, idempotent only, 10% budget — otherwise 3 layers make 27× load""Retry storms are cross-team incidents; the retry budget is org policy with an exceptions process"
Tail latency"Our p99 is 50ms""We fan out to 40 leaves, so the leaf p99 is our page p67. Hedge reads at p95, partial results for decorative modules""Tail SLOs are set per tier and owned by the aggregator team; leaf teams are budgeted, not left to guess"
Real-time"Use WebSockets""SSE — one-way push, HTTP-native, resumable. WebSocket when we actually need bidirectional""A connection tier is a new stateful platform with its own on-call. I'd build one per company, not one per feature"
Why "Tail latency" separates levels

The Senior answer reports a leaf metric accurately. The Staff answer knows that the user experiences a different distribution — the maximum of N leaf samples — and redesigns the aggregator around it. The Principal answer recognizes that tail latency is a contract between teams: if the aggregator team owns the page SLO but 40 leaf teams each own their own p99, nobody owns the composition. L7 fixes the ownership model — tier-level budgets and a shared mechanism for hedging and partial results — so the math stops being rediscovered in every incident review.

Why "Retries" separates levels

"Retry 3 times with exponential backoff" is textbook-correct for a single client. Staff engineers compute what happens when every layer does it. Principal engineers notice that the retrying team and the team that gets paged are different teams, so the fix cannot live in either team's code review — it has to be a platform default with an exceptions process.

Common Interview Traps#

  • Drawing arrows as if they were free. Every cross-region arrow is 60–200ms. Say the number when you draw it.
  • Forgetting the handshake. "One request, 80ms" is wrong for a cold client. It's 240–320ms.
  • Putting gRPC behind an L4 load balancer and calling it balanced. Long-lived connections pin load; say how you rebalance.
  • Choosing WebSockets by default. A stateful connection tier is a platform. SSE or polling is often enough.
  • Retrying at every layer. Multiply the attempts; name the budget.
  • Reporting averages. Averages hide the tail, and the tail is what fan-out amplifies. Report p50/p99/p99.9.
  • Assuming DNS failover is instant. Resolvers and clients cache beyond TTL. Plan for minutes.
  • Using 0-RTT for writes. 0-RTT data can be replayed by an attacker; restrict it to idempotent requests.

Practice Drill#

Prompt: "Our checkout API has a p50 of 40ms but a p99 of 900ms. It calls 12 downstream services in parallel. Leadership wants the p99 under 250ms this quarter. What do you do?"

Staff Answer

First I'd confirm that the p99 is dominated by the fan-out and not by one leaf: with 12 parallel calls, even a clean 1% tail per leaf gives ~11% of checkouts at least one slow call, which is enough to drag the checkout p99 up to the worst leaf's p99.9. I'd pull per-leaf latency histograms and a trace sample of the slowest 1% of checkouts to see whether one or two leaves own most of the tail (usual) or it's spread evenly (queueing or GC across the fleet). Then, in order: (1) propagate a 250ms deadline from the gateway so we stop doing work after the user has given up; (2) split the 12 calls into must-have (cart, pricing, tax, payment-auth eligibility, inventory) and nice-to-have (loyalty points, recommendations, promo banners) — nice-to-haves get a 100ms deadline and degrade to empty; (3) hedge the idempotent reads among the must-haves at their p95, watching hedge.win_rate and capping extra load at ~5%; (4) check for the classic tail sources — 200ms TCP retransmit clusters, 40ms Nagle stalls, GC pauses, leaf CPU above 70% — and fix the ones present; (5) remove retries below the gateway and add a 10% retry budget at the gateway. Payment authorization is never hedged or blindly retried; it gets an idempotency key and a longer deadline carved out of the budget. Owners: checkout team owns the budget and degradation policy; each leaf team gets a p99 target derived from it; platform owns the RPC library changes.

Why this is L6:

  • Uses the fan-out formula to diagnose before prescribing.
  • Separates must-have from nice-to-have with an explicit degradation contract.
  • Refuses to hedge or retry non-idempotent calls, and says why.
  • Assigns ownership of each budget to a named team.

What L7 adds:

  • Turns the leaf-level p99 targets into an org-wide latency budget contract that is reviewed whenever a new dependency is added to checkout, so the problem doesn't regress next year.
  • Prices the fix: ~5% extra read load from hedging vs the conversion lift from a 650ms p99 reduction (conversion data from product), and presents the tradeoff to the business in dollars.
  • Pushes deadline propagation and retry budgets into the shared RPC stack so the next 50 services inherit the fix instead of rediscovering it.

Where This Appears#

  • Load Balancer — L4 vs L7 balancing, connection vs request distribution, health checks, and draining
  • API Gateway — Edge TLS termination, deadline propagation, and retry policy at the front door
  • CDN & Edge Caching — Anycast, edge termination, and why "terminate close, travel warm" dominates global latency
  • Real-Time Updates — WebSocket vs SSE vs long polling, connection tiers, and reconnect storms
  • Chat Messaging — Persistent connections at millions-of-users scale and presence routing
  • Circuit Breakers — Retry budgets, outlier ejection, and preventing cascading failure
  • Service Discovery — Client-side load balancing and how clients find backends for long-lived connections

Related Foundations & Patterns: Back-of-Envelope Estimation · API Design Patterns · Real-time Updates · Degraded Mode Framework

Related Technologies: API Gateways

  1. Loading the index…