Why This Matters#
Networking is not a protocol-trivia question. Nobody gets leveled Staff for reciting the TCP state machine. Networking is a latency budget question: every request has a fixed number of milliseconds before the user notices, and every hop, handshake, retry, and proxy spends some of them. The Staff question is who spends each millisecond, and who pays when the tail blows up.
Most candidates draw arrows between boxes as if arrows were free. They are not. An arrow inside a rack costs ~50µs. An arrow across availability zones costs ~1ms and $0.01 per GB in each direction. An arrow across an ocean costs ~80ms of physics that no amount of engineering will remove. A fresh HTTPS connection across that ocean costs three or four of those round trips before the first byte of payload moves. Candidates who know this draw fewer arrows, reuse connections, and put state close to the user. Candidates who don't design systems that are correct and slow.
The second reason networking shows up in Staff loops is tail latency. A service with a p99 of 10ms is fine in isolation. Put it behind a fan-out of 100 parallel calls and 63% of user requests now experience that p99. The p99 of a leaf becomes the median of the page. This is the single most under-appreciated number in system design, and it is where interviewers separate people who have operated fan-out services from people who have read about them.
If you can walk an interviewer from a latency budget ("the page has 200ms; 80ms is physics") to a protocol choice ("so we terminate TLS at the edge and hold warm HTTP/2 connections to origin") to a tail-latency mitigation ("hedge after p95, capped at 5% extra load") — and name who owns each piece — you are demonstrating the reasoning Staff interviews reward.
The 60-Second Version#
- Physics sets the floor. Light in fiber travels ~200km per millisecond. US East ↔ US West is ~60–70ms RTT, transatlantic ~70–90ms, US ↔ Asia ~150–200ms. You cannot optimize below this; you can only avoid paying it more than once.
- Handshakes multiply RTT. A cold HTTPS request costs TCP (1 RTT) + TLS 1.3 (1 RTT) + request (1 RTT) = 3 RTT; TLS 1.2 makes it 4. QUIC folds transport and crypto into 1 RTT, and 0-RTT resumption gets repeat visitors to 1 RTT total — at the cost of replay risk on non-idempotent requests.
- Connection reuse is the single biggest latency win in most systems. A warm connection turns 3–4 RTT into 1 and skips TCP slow start (~14KB initial window, doubling per RTT).
- HTTP/2 fixed application-level head-of-line blocking but kept TCP's. One lost packet stalls every multiplexed stream. HTTP/3 over QUIC removes transport HOL blocking — it matters most on lossy mobile networks (1–3% loss), barely at all inside a data center.
- Long-lived connections break L4 load balancing. gRPC over HTTP/2 pins all requests on one connection to one backend. You need L7 (per-request) balancing, client-side balancing, or forced connection recycling (e.g., max connection age ~5–30 min).
- Tail latency compounds under fan-out.
P(slow page) = 1 − 0.99^Nfor N parallel calls. N=100 → 63%. Mitigations: hedged requests, tied requests, deadline propagation, and reducing N. - Retries multiply across layers. 3 attempts at each of 3 layers = up to 27× load on the bottom service during an incident. Retry once, at one layer, with a retry budget (~10% of traffic).
How Networking Works (for System Designers)#
The Latency Stack#
Every remote call pays some combination of these costs. The design question is which ones you pay per request versus once per connection versus never.
| Cost | Typical Value | Paid When | How to Avoid It |
|---|---|---|---|
| DNS resolution | 0–100ms (cached: ~0) | First request to a hostname, after TTL expiry | Long-ish TTLs (60–300s), connection reuse, prefetch |
| TCP handshake | 1 RTT | Every new connection | Keepalive, connection pooling |
| TLS 1.3 handshake | 1 RTT (0 with resumption) | Every new connection | Session resumption, TLS termination near the client |
| TLS 1.2 handshake | 2 RTT | Every new connection | Upgrade to 1.3 |
| TCP slow start | Several RTT for large responses | Every new (or long-idle) connection | Warm connections, larger initcwnd at the edge |
| Serialization | 10µs–5ms | Every request | Protobuf/binary formats on hot paths |
| Propagation | ~5µs/km one way (≈1ms RTT per 100km) | Every request | Move compute/data closer; you cannot beat light |
| Queueing | 0 → unbounded | Every request under load | Keep utilization < ~70%; backpressure |
| Proxy hop (L7) | 0.2–2ms each | Every request | Fewer hops, sidecar-less where it matters |
Reference Round-Trip Times#
| Path | RTT | What It Means |
|---|---|---|
| Same host (loopback) | ~10–50µs | Sidecar hops are cheap but not free |
| Same rack / same AZ | ~0.1–0.5ms | Chatty protocols are tolerable |
| Cross-AZ, same region | ~0.5–2ms | Synchronous replication across AZs is affordable |
| US East ↔ US West | ~60–70ms | One synchronous cross-region call per request, at most |
| US East ↔ Europe | ~70–90ms | Consensus across the Atlantic costs this per write |
| US ↔ Asia-Pacific | ~150–200ms | Synchronous anything here is a product decision, not a technical one |
| Mobile last mile (4G) | ~30–100ms, 1–3% loss | The client link often dominates the whole budget |
Where Protocols Live#
Application HTTP/1.1 · HTTP/2 · HTTP/3 · gRPC · WebSocket · SSE
Security TLS 1.2 / 1.3 (built into QUIC for HTTP/3)
Transport TCP (reliable, ordered byte stream) · UDP / QUIC (datagrams; QUIC adds streams + reliability)
Network IP · anycast · BGP
L4 vs L7 is the distinction that matters most in design interviews:
| Layer | Sees | Balances | Cost per Request | Typical Tools |
|---|---|---|---|---|
| L4 (transport) | IP, port, TCP/UDP flows | Connections | ~µs; millions of packets/s per node | AWS NLB, Google Maglev, IPVS, Katran |
| L7 (application) | HTTP method, path, headers, gRPC service | Requests | ~0.2–2ms; tens of thousands of RPS per core | Envoy, NGINX, HAProxy, AWS ALB |
L4 is fast and dumb; L7 is slower and smart. L7 terminates TLS, can route by path or tenant, retry, rate limit, and balance per request. L4 forwards bytes and balances per connection — which is precisely why it mis-balances long-lived connections.
Core Strategies#
Strategy 1: Connection Reuse and Pooling#
The cheapest round trip is the one you never make. Keep connections warm and share them.
# Per-process pool to each upstream
pool = ConnectionPool(
upstream = "orders.internal:443",
max_conns = 64, # per upstream host
idle_timeout = 55s, # shorter than the LB's idle timeout (e.g., ALB default 60s)
max_conn_age = 10m, # force periodic rebalancing across backends
keepalive_ping = 30s, # detect dead peers before a request finds them
)
call(req):
conn = pool.acquire(deadline = req.deadline) # never block past the caller's deadline
return conn.send(req)
When to use: Always, for service-to-service traffic. The question is only how many connections and how long they live.
Failure mode: Idle-timeout mismatch. If the client keeps a connection idle for 90s and the load balancer drops it at 60s, the next request on that connection gets a reset. Symptom: a small, steady rate of connection reset by peer errors that correlates with low-traffic periods. Fix: client idle timeout strictly below every middlebox's idle timeout.
Strategy 2: Terminate Close, Travel Warm#
Put TLS termination at an edge PoP near the user (CDN or regional edge), then carry the request to origin over a pre-established, long-lived connection.
| Setup | User in Sydney → Origin in Virginia (~200ms RTT) | Time to First Byte |
|---|---|---|
| Direct, TLS 1.3, cold | 3 RTT × 200ms | ~600ms + server time |
| Edge in Sydney (~10ms RTT) + warm backbone to origin | 2 × 10ms handshake + 1 × 200ms | ~220ms + server time |
| Edge + cached response | 2 × 10ms + 0 | ~20ms |
When to use: Any global consumer product. This is why every CDN also sells "dynamic acceleration."
Failure mode: The edge becomes a hidden dependency. An edge configuration push that breaks TLS takes down every origin behind it — a correlated failure across teams who never talk to each other.
Strategy 3: Pick the Protocol for the Traffic Shape#
| Traffic Shape | Protocol | Why |
|---|---|---|
| Public request/response API | HTTP/1.1 or HTTP/2 + JSON | Universal tooling, cacheable, debuggable with curl |
| Internal service-to-service RPC | gRPC (HTTP/2 + protobuf) | Schemas, codegen, deadlines, streaming, ~3–10× smaller payloads than JSON |
| Browser/mobile on lossy links | HTTP/3 (QUIC) | No transport HOL blocking, connection migration across Wi-Fi ↔ cellular |
| Server → client push, one direction | Server-Sent Events | Plain HTTP, auto-reconnect with Last-Event-ID, works through most proxies |
| Bidirectional, low-latency, interactive | WebSocket | Full duplex over one TCP connection; you own reconnect, heartbeats, and fan-out |
| Infrequent updates, simplest possible | Long polling | Works everywhere; ~1 request per update, wasteful at high rates |
| Loss-tolerant real-time media | UDP / WebRTC | Late data is worthless; retransmission would make it worse |
Failure mode: Choosing WebSockets for a feature that pushes one update a minute. You now operate a stateful connection tier — sticky routing, reconnect storms on deploy, per-connection memory (~10–50KB each) — for traffic that SSE or polling would have handled statelessly.
Strategy 4: Deadlines, Not Timeouts#
A timeout is local ("I'll wait 2s"). A deadline is global ("this whole request must finish by 12:00:00.250"). Propagate the deadline and every hop can decide whether it is still worth doing work.
handle(req):
remaining = req.deadline - now()
if remaining < MIN_USEFUL_WORK: # e.g., 5ms
return DEADLINE_EXCEEDED # fail fast; don't burn downstream capacity
child_deadline = req.deadline - SAFETY_MARGIN # leave time to respond
a = call(inventory, deadline = child_deadline)
b = call(pricing, deadline = child_deadline)
return merge(a, b)
When to use: Any call graph deeper than two hops. gRPC propagates deadlines natively; HTTP stacks need a header convention.
Failure mode: Uncoordinated static timeouts. The edge times out at 1s, the service at 2s, the database client at 5s. When the database slows down, the edge has already given up, but the service and database keep working for 4 more seconds on requests nobody is waiting for. That wasted work is what turns a slowdown into an outage.
Strategy 5: Retry Once, at One Layer, With a Budget#
retry_policy = {
max_attempts: 2, # original + 1 retry
retry_on: [UNAVAILABLE, RESET, 503], # never on 4xx or DEADLINE_EXCEEDED
backoff: exponential(base=25ms, cap=250ms, jitter=full),
retry_budget: 0.10, # retries ≤ 10% of requests over a 10s window
idempotent_only: true,
}
Failure mode: Retry amplification. Three layers each retrying three times turns one failing request into 27 attempts at the bottom. The bottom service was slow because it was overloaded; now it receives 27× the load. Retry budgets cap the damage at ~1.1× instead of 27×.
Tail Latency: The Hard Sub-Problem#
Averages lie; tails compound. The single most important formula in this page:
P(request sees at least one slow leaf) = 1 − (1 − p_slow)^N
p_slow = 1% (i.e., each leaf's p99), N parallel leaves:
N = 1 → 1.0%
N = 10 → 9.6%
N = 50 → 39.5%
N = 100 → 63.4%
N = 1000 → 99.99%
At N=100, the leaf's p99 is the page's median experience. This is why a search backend, a feed aggregator, or a scatter-gather query over 100 shards cannot be fixed by making the average faster. You must attack the tail directly.
Where Tails Come From#
| Source | Typical Magnitude | Signature |
|---|---|---|
| GC pauses | 10–500ms (JVM), less with modern collectors | Periodic spikes per host, uncorrelated across hosts |
| Queueing at high utilization | Grows as ~1/(1−ρ); at 90% utilization queue delay is ~10× service time | p99 climbs sharply past ~70% CPU |
| TCP retransmission | Minimum RTO ~200ms on Linux | A cluster of requests at exactly +200ms |
| Delayed ACK + Nagle interaction | ~40ms | Small writes stall at exactly ~40ms; fixed by TCP_NODELAY |
| Noisy neighbors | 2–10× service time | Tail varies by host, not by request |
| Cold caches / cold connections | 1–3 extra RTT | Tail spikes right after deploys or scale-out |
| Background work (compaction, backups) | 10–100ms | Tail correlates with maintenance windows |
Mitigation Toolbox#
| Technique | How It Works | Extra Load | When It Breaks |
|---|---|---|---|
| Hedged requests | Send a second copy to another replica if no response by the p95 | ~5% (only the slowest 5% are hedged) | Non-idempotent calls; correlated slowness (both replicas share the cause) |
| Tied requests | Send to two replicas at once; the first to start cancels the other | ~0–2% wasted work | Needs cancellation plumbing in the server |
| Reduce fan-out | Fewer, fatter shards; pre-aggregated indexes | Negative | Shard size limits, rebalancing cost |
| Partial results | Return what arrived by the deadline; mark the response incomplete | None | Product must accept "good enough" answers |
| Utilization cap | Keep leaf CPU below ~60–70% | Capacity cost (~30–40% headroom) | Finance asks why the fleet is "idle" |
| Micro-partitioning | 10–100 partitions per server so load moves in small units | Metadata overhead | Coordination complexity |
🎯 Staff Move: "This page fans out to about 40 services, so a 1% tail at each leaf means roughly a third of page loads see at least one slow call. I'd rather not fix that at each leaf. I'd propagate a 150ms deadline, hedge the three read-only calls at their p95, and render partial results for the recommendations module — product has to sign off that a page without recommendations is acceptable."
Full reasoning: why hedging at p95 costs only ~5%
If you send a hedge only when the first attempt has not returned by the p95 latency, then by definition only ~5% of requests trigger a hedge. The extra load is ~5%, not 100%. The payoff is large because the slow 5% of first attempts are usually slow for reasons local to that replica (GC, queueing, a noisy neighbor), so the hedge to a different replica usually completes near the median.
Two rules keep it safe:
- Only hedge idempotent reads. Hedging a payment capture is how you double-charge someone.
- Cancel the loser. Without cancellation, a hedged system under overload does 1.05× the work and the losing requests still occupy queue slots.
Hedging fails when slowness is correlated — for example, a slow shared dependency behind both replicas. Then both attempts are slow and you have added 5% load to an already-overloaded system. Watch hedge.win_rate: if hedges rarely win, turn them off.
Visual Guide#
A Cold HTTPS Request vs a Warm One#
Handshake Cost by Protocol#
| Protocol | New Connection | Resumed Connection | Caveat |
|---|---|---|---|
| TCP + TLS 1.2 | 1 TCP + 2 TLS + 1 request = 4 RTT to response | 1 TCP + 1 TLS + 1 request = 3 RTT | Still common on legacy clients |
| TCP + TLS 1.3 | 1 TCP + 1 TLS + 1 request = 3 RTT to response | 0-RTT data: 2 RTT (TCP still costs 1) | 0-RTT data can be replayed |
| QUIC (HTTP/3) | 1 combined handshake + 1 request = 2 RTT to response | 0-RTT: request in the first flight = 1 RTT | UDP is blocked on a small share of networks; clients fall back to TCP |
Choosing a Real-Time Transport#
L4 vs L7 in a Typical Topology#
The pattern: L4 spreads connections across a stateless L7 tier cheaply; L7 spreads requests across services intelligently. Services talk to each other over gRPC with client-side or sidecar balancing so no single long-lived connection pins load to one backend.
Implementation Patterns#
HTTP/1.1 vs HTTP/2 vs HTTP/3#
| Dimension | HTTP/1.1 | HTTP/2 | HTTP/3 |
|---|---|---|---|
| Transport | TCP | TCP | QUIC over UDP |
| Concurrency | 1 request in flight per connection; browsers open ~6 per origin | Many multiplexed streams on 1 connection (default max ~100 concurrent streams) | Many independent streams |
| Head-of-line blocking | Application-level (requests queue behind each other) | Transport-level (one lost packet stalls all streams) | None at transport level (loss stalls only the affected stream) |
| Header compression | None | HPACK | QPACK |
| Handshake | TCP + TLS | TCP + TLS (ALPN negotiates h2) | Combined 1 RTT; 0-RTT resumption |
| Connection migration | No | No | Yes (connection IDs survive IP changes) |
| Where it wins | Simplicity, debugging, legacy | Service-to-service in low-loss DCs | Mobile and lossy last mile |
Staff default: HTTP/2 (gRPC) inside the data center; HTTP/3 at the edge for mobile and web clients, with TCP fallback. HTTP/3 inside the data center buys little — loss is near zero, and UDP processing is typically more CPU-expensive per byte than kernel-optimized TCP.
gRPC and the Long-Lived Connection Problem#
gRPC multiplexes every call over a small number of HTTP/2 connections. Behind an L4 load balancer, that means:
Client pods: 20, each with 1 connection
Server pods: 10 → 20 after autoscaling
Result: the 10 new servers receive 0 connections and 0 traffic
until clients reconnect — which may be never.
Three fixes, in order of preference:
- Client-side load balancing — the client resolves all backend addresses (via DNS or a discovery service), keeps connections to each, and balances per call (round robin or least-outstanding-requests).
- L7 proxy / sidecar — Envoy balances per request across backends; the client holds one connection to its local proxy.
- Max connection age — servers send
GOAWAYafter 5–30 minutes so clients reconnect and redistribute. Cheap, crude, and a good backstop even when you do 1 or 2.
WebSockets and SSE at Scale#
| Concern | Number | Design Consequence |
|---|---|---|
| Memory per idle connection | ~10–50KB (kernel buffers + app state) | 1M connections ≈ 10–50GB across the fleet |
| Connections per well-tuned server | 100K–1M+ (file descriptors, ephemeral ports, memory) | A gateway tier of tens of hosts for 10M users |
| Heartbeat interval | 20–60s | Must be shorter than every NAT/LB idle timeout on the path |
| Reconnect storm on deploy | All clients on a host reconnect within seconds | Drain gradually; jittered reconnect backoff (1–30s) on the client |
Separate the connection tier (holds sockets, stateless about business logic) from the business tier (stateless about sockets). A presence or routing service maps user_id → gateway_host. This is how chat and notification systems stay deployable.
Anycast and DNS-Based Steering#
- Anycast: the same IP is announced from many PoPs via BGP; the network routes each user to the nearest. Failover is BGP convergence (seconds to a minute). Works best for short-lived or connectionless traffic (DNS, CDN).
- DNS steering (GeoDNS / latency-based): resolvers get region-specific answers. Failover is bounded by DNS TTL plus resolvers that ignore it — plan for minutes, not seconds.
Failure Scenario: The Cross-Team Retry Storm#
t=0 Profile DB p99 rises from 15ms to 400ms (compaction + a hot partition).
t=+10s Profile service (2s timeout, 3 attempts) starts retrying. DB load ×2.
t=+20s Feed service (1s timeout, 3 attempts) times out on Profile and retries. Profile load ×3.
t=+30s Edge gateway (3s timeout, 2 attempts) retries Feed. Total DB attempts per user request: up to 18.
t=+45s DB connection pool exhausted. Every caller of Profile fails, including checkout.
t=+2min Incident declared. Three teams on the bridge, each certain their retries are "reasonable".
t=+9min Mitigation: gateway retries disabled via config; Profile shed to 50% for non-critical callers.
Detection: rpc.client.retries ÷ rpc.client.requests above 10% on any edge; db.connections.in_use at pool max; rpc.server.deadline_exceeded rising at a service whose own dependencies look healthy.
Blast radius: every service that shares the Profile DB, including checkout, which never retried at all.
Mitigation: disable retries at the outermost layer first (it has the largest multiplier), shed non-critical callers, then let the DB recover.
Prevention: one retrying layer, retry budgets in the shared RPC library, deadline propagation so abandoned requests stop consuming DB connections.
Owner: the platform team that owns the RPC library — because no single product team can fix a behavior that emerges from three teams' defaults.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Retry storm | rpc.client.retries ratio > 10% | Every service sharing the slow dependency | Disable outer-layer retries; shed load | RPC platform team |
| Idle-timeout mismatch | Steady connection_reset at low traffic | One client/LB pair | Client idle timeout < LB idle timeout | Calling service |
| gRPC connection pinning | Per-pod CPU skew > 3× after scale-out | One service's capacity | Max connection age; client-side LB | Service owner + platform |
| Edge/CDN config push breaks TLS | edge.tls_handshake_errors spike across all hosts | Every origin behind the edge | Automated rollback; staged config rollout | Edge/traffic team |
| WebSocket reconnect storm on deploy | ws.connects_per_sec 10–50× baseline | Connection tier + auth service | Gradual drain; jittered client backoff | Real-time platform team |
| Cross-region link degradation | rpc.latency.p99 for cross-region calls; packet loss > 1% | Features with synchronous cross-region calls | Fail over reads to local replicas; defer async replication | Service owners; network team for the link |
The Numbers in Context#
| Number | Value | What It Means for Your Design |
|---|---|---|
| Light in fiber | ~200km/ms one way | Every 1,000km adds ~10ms RTT. Region placement is a latency decision. |
| Cross-AZ RTT | ~0.5–2ms | Synchronous replication across 3 AZs adds ~1–2ms per write — affordable. |
| Cross-region RTT | 60–200ms | Synchronous cross-region writes put this on every commit. Usually a no. |
| TCP initial window | 10 segments ≈ 14KB | A 1MB response on a cold connection needs ~6–7 RTT of slow start. At 80ms RTT that is ~500ms. |
| Bandwidth-delay product | 1Gbps × 80ms = 10MB | A single TCP flow needs a ~10MB window to fill a long fat pipe; default buffers often cap throughput far lower. |
| L7 proxy hop | ~0.2–2ms p50, more at p99 | A service mesh adds 2 hops per call (client + server sidecar). A 10-deep call chain pays 20 hops. |
| Linux min RTO | ~200ms | One lost packet on a 1ms-RTT link turns a 2ms call into a 200ms call. This is a common source of p99.9. |
| Browser connections per origin (HTTP/1.1) | ~6 | The reason for domain sharding in the HTTP/1.1 era, and the reason HTTP/2 made it an anti-pattern. |
| Default HTTP/2 max concurrent streams | ~100 (server-configurable) | A client exceeding this queues locally; heavy RPC clients need more connections. |
| Fan-out tail | N=100 at p99 → 63% slow | Tail latency, not average latency, sets the user experience of any aggregator. |
| Retry amplification | 3 layers × 3 attempts = 27× | One retrying layer, one retry, a 10% budget. |
| Cloud cross-AZ transfer | ~$0.01/GB each direction | A chatty service moving 1PB/month across AZs pays ~$20K/month for the privilege. |
How This Shows Up in Interviews#
Scenario 1: "Users in Asia say the app is slow. Our servers are in Virginia."#
The interviewer is testing whether you start from physics. Say: "Virginia to Singapore is roughly 200–230ms RTT. A cold HTTPS request is 3 RTT before any server time, so we are spending ~650ms on handshakes alone. Step one is an edge PoP in the region to terminate TLS at ~10ms RTT and carry traffic on warm connections; that cuts ~400ms without moving any data. Step two — only if reads dominate — is a regional read replica or cache so the remaining 200ms round trip disappears for reads. Writes still go to Virginia until someone signs off on a multi-region write story."
Scenario 2: "Design the API layer for a page that calls 30 backend services." (Full Walkthrough)#
Step 1 — Budget the latency. "The product SLO is p99 page render under 300ms. I'll reserve 50ms for the client link and rendering, which leaves 250ms server-side. With 30 parallel calls, if each leaf has a 1% chance of exceeding its p99, about 26% of pages hit at least one slow leaf. So I can't just budget each leaf's p99 at 250ms — I have to design for the tail."
Step 2 — Shape the call graph. "I'll split the 30 calls into critical (8: user, cart, pricing, inventory…) and decorative (22: recommendations, badges, promotions). Critical calls block the response. Decorative calls get a 120ms deadline and render as partial results if they miss — product signs off on that."
Step 3 — Propagate deadlines. "The aggregator sets a 250ms deadline and passes it down via gRPC metadata. Each service subtracts its own safety margin. Anything that would start with less than 10ms left fails fast instead of doing work nobody will read."
Step 4 — Attack the tail on reads. "For the four critical read-only calls with the worst p99/p50 ratio, I'll hedge at their p95. That costs ~5% extra load on those services. I'll put hedge.win_rate on a dashboard; if it drops below ~20%, the slowness is correlated and hedging is only adding load."
Step 5 — Control retries. "Retries happen only at the aggregator, once, for UNAVAILABLE, with a 10% retry budget. Leaf services never retry each other's calls. That caps worst-case amplification at 1.1× instead of compounding."
Step 6 — Name the owners. "The aggregator team owns the deadline and degradation policy. Each leaf team owns its p99 SLO. The platform team owns the RPC library that enforces deadline propagation and retry budgets — so the policy isn't re-implemented 30 times."
Why this is a Staff answer: It turns a vague performance question into a budget, uses the fan-out formula to justify the design, separates critical from decorative work with a product sign-off, and assigns ownership of the tail to the layer that can actually fix it.
Scenario 3: "We moved to gRPC and one backend pod is at 100% CPU while the rest idle."#
This tests the long-lived connection problem. The answer: "gRPC multiplexes everything over a few HTTP/2 connections, and our L4 load balancer balances connections, not requests. After the last scale-out, the old connections stayed pinned. Short-term, set a max connection age of ~10 minutes on the servers so clients reconnect. Long-term, move to client-side balancing with least-outstanding-requests, or put an L7 proxy in the path." Bonus points for noting that the problem gets worse with autoscaling — new pods get no traffic.
Scenario 4: "Should the live-scores feature use WebSockets?"#
This tests whether you pick the simplest transport that works. "Scores are server-to-client only and update every few seconds. That's SSE: plain HTTP, auto-reconnect with Last-Event-ID for resume, CDN-friendly fan-out. WebSocket buys bidirectionality we don't need and costs us a stateful connection tier. If we later add live chat under the scoreboard, that's the moment to reconsider."
Advanced Patterns#
| Pattern | How It Works | When to Use |
|---|---|---|
| Hedged requests | Duplicate a slow read to another replica after the p95; cancel the loser | Fan-out reads with uncorrelated tail causes |
| Deadline propagation | Absolute deadline travels with the request; each hop checks remaining budget | Call graphs deeper than 2 hops |
| Retry budgets | Retries capped at a % of successful traffic per client | Every RPC client — prevents retry storms |
| Zone-aware routing | Prefer same-AZ backends; spill cross-AZ only when local capacity is short | High-volume east-west traffic (cost + ~1ms latency) |
| Connection draining | On deploy, stop accepting new connections, send GOAWAY, wait for in-flight | Every deploy of a connection-holding tier |
| Request collapsing at the edge | CDN coalesces concurrent misses for the same object into one origin fetch | Viral content, cache-miss storms |
| Outlier ejection | Remove backends whose error rate or latency is N× peers from the pool for 30s+ | Service mesh default; catches bad hosts before humans do |
| Adaptive concurrency limits | Client or server lowers in-flight limit when latency rises (TCP-Vegas-style) | Protecting services without hand-tuned static limits |
The Principal Lens#
Why L7 Sees This Problem Differently#
A Staff engineer tunes the network behavior of their service: pools, timeouts, protocol choice, hedging. A Principal engineer notices that every one of those settings is being chosen independently by 200 teams, and that the resulting emergent behavior — retry storms that cross team boundaries, timeout hierarchies nobody designed, cross-AZ traffic nobody budgeted — is the actual reliability and cost problem. At L7, networking stops being a protocol question and becomes a paved-road question: which behaviors do we bake into the one RPC stack everyone uses, so the organization's default is safe even when a team never thinks about it?
The Org-Level Fault Line#
Centralized RPC/mesh platform vs per-team client libraries.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Every team picks its own HTTP/gRPC client and settings | Autonomy; no platform bottleneck | Inconsistent retries and timeouts; incidents cascade across teams; no org-wide view | On-call of the downstream team that absorbs other teams' retry storms |
| Shared RPC library per language | Consistent defaults; cheap per call | Upgrading 200 services takes quarters; polyglot orgs maintain N libraries | Platform team maintaining 3–5 language ports |
| Service mesh (sidecar or proxyless) | Policy changes without redeploying services; uniform telemetry and mTLS | +2 proxy hops per call (~0.5–2ms each); sidecar CPU/memory tax; a new tier-0 dependency | Every service pays the latency and compute tax; mesh team carries the pager for everything |
The Principal position: a thin, mandatory RPC contract (deadlines, retry budgets, mTLS, standard telemetry) enforced by either a shared library or a mesh — and the choice between those two is driven by how polyglot the org is and how much latency the hottest paths can spare. Mandate the contract, not the implementation.
Cost Model#
Assumptions: AWS-style list pricing, cross-AZ transfer $0.01/GB each direction ($0.02/GB round trip), internet egress ~$0.05–0.09/GB, sidecar ~0.1–0.25 vCPU per pod at ~$30/vCPU-month, fully loaded engineer ~$25K/month.
| Scale | East-West Traffic | Cross-AZ Transfer $ | Mesh/Proxy Compute $ | Platform Headcount | On-Call Load |
|---|---|---|---|---|---|
| Startup (30 services, 300 pods) | ~20TB/month, 60% cross-AZ | ~$250/month | ~$1–2K/month | 0.5 engineer (shared infra) | Network issues ~1 page/month |
| Growth (300 services, 5K pods) | ~1PB/month, 60% cross-AZ | ~$12K/month | ~$20–40K/month | 3–4 engineers (~$90K/month) | Dedicated mesh rotation, ~1 incident/week |
| Large (2,000 services, 50K pods) | ~30PB/month; zone-aware routing cuts cross-AZ to ~20% | ~$120K/month (vs ~$360K without zone-aware routing) | ~$200–400K/month | 10–15 engineers | Tier-0 rotation; mesh is on every incident bridge |
The lever most orgs miss: zone-aware routing can pay for most of the networking platform team at large scale (~$240K/month saved in the Large row), because cross-AZ transfer is billed per GB and east-west traffic grows faster than user traffic.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversal Cost |
|---|---|---|
| Timeout, retry, and hedging values | Two-way | Config push; minutes |
| HTTP/2 vs HTTP/3 at the edge | Two-way | Clients fall back automatically |
| Adopting a service mesh | Mostly one-way | Every service's traffic, identity, and telemetry now flows through it; exit is a multi-quarter migration |
| Public API protocol (REST/JSON vs gRPC vs GraphQL) | One-way | External clients and SDKs pin you for years |
| Region placement for the primary data store | One-way | Data gravity; moving petabytes and re-pointing every dependency |
| Exposing WebSocket semantics to third-party clients | One-way | Ecosystem depends on connection behavior, message framing, and reconnect contract |
| Choice of edge/CDN vendor | Two-way in theory, one-way in practice | Edge logic (workers, rules, WAF) accretes vendor-specific code |
The Standard I'd Write#
RFC: Service-to-Service Communication Standard (v1)
Scope: All synchronous internal RPC between production services. Excludes batch data transfer and external/public APIs.
MUST:
- Every request carries an absolute deadline; services MUST NOT start work when remaining budget is below their documented minimum.
- Retries MUST be limited to 1 additional attempt, only on idempotent methods, only for transport/UNAVAILABLE errors, and within a per-client retry budget of 10%.
- All traffic MUST use mTLS with workload identity; plaintext internal traffic is prohibited.
- Clients MUST emit standard metrics:
rpc.client.latency(histogram),rpc.client.errorsby code,rpc.client.retries,rpc.client.deadline_exceeded.- Servers MUST support graceful drain (
GOAWAY+ in-flight completion within 30s).SHOULD: prefer same-zone backends; set max connection age 5–30 minutes; hedge only reads whose p99/p50 > 5 and only after the p95.
Exceptions: Filed with the networking platform team; approved exceptions expire after 2 quarters.
Success metrics: zero cross-team retry-storm incidents per quarter; cross-AZ bytes ÷ total east-west bytes < 25%; ≥ 95% of services on the standard client within 3 quarters.
What I'd Tell the VP#
"Our services talk to each other millions of times a second, and today each team decides on its own how long to wait and how often to retry. That's why a slowdown in one service turned into a two-hour outage across checkout last quarter. We want to put one set of rules into the shared plumbing so every team gets safe behavior by default. That needs a platform team of about four engineers. Keeping traffic inside a zone also saves about $100K a year in data-transfer charges today, and that saving grows faster than our traffic. The risk is that the plumbing becomes critical infrastructure, so we'll roll it out gradually and staff it with a real on-call rotation."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Treats network behavior as org policy | "The problem isn't this service's timeout. It's that 200 teams each chose one, and nobody designed the hierarchy." |
| Prices the east-west graph | "At 1PB a month of cross-AZ chatter we're paying ~$20K monthly to cross a line we could route around." |
| Separates contract from implementation | "I'd mandate deadlines, retry budgets, and mTLS. Whether that's a library or a mesh depends on how polyglot we are." |
| Names correlated-failure risk | "The mesh control plane and the edge config pipeline are now shared fate for every team — they need their own blast-radius cells and staged rollouts." |
| Knows when not to standardize | "The trading path and the video path get an exemption — their latency budgets can't afford a sidecar." |
Staff answers that L7 interviewers find insufficient:
- "I'd set our client timeout to 500ms and retry twice." — Correct locally; ignores what every other caller in the graph is doing.
- "We should adopt a service mesh." — Names a tool without the contract it enforces, the tax it charges, or the exit cost.
- "Cross-AZ latency is only ~1ms, so it doesn't matter." — True for latency, blind to the transfer bill that grows faster than traffic.
In the Wild#
These are public, documented examples.
Google: "The Tail at Scale"#
Jeff Dean and Luiz André Barroso's 2013 Communications of the ACM paper formalized the fan-out tail problem using Google's own serving systems. It showed how a leaf-level 1-in-100 slow response becomes the common case for a root that fans out to 100 leaves, and introduced hedged requests and tied requests as the practical fixes. The paper reports that deferring a hedge until a request has been outstanding past its 95th-percentile latency sharply cuts the tail while adding only a few percent extra load.
Staff insight: Quote the formula, not the paper. "With 100 leaves, a 1% leaf tail is a 63% page tail" is the sentence that tells an interviewer you have reasoned about aggregators, not just built one.
Google QUIC → IETF HTTP/3#
Google designed and deployed QUIC in Chrome and its own services during the 2010s to remove TCP's head-of-line blocking and cut handshake round trips; the IETF standardized QUIC as RFC 9000 in 2021, and HTTP/3 (RFC 9114) runs on top of it. Major CDNs and browsers now support it, with automatic fallback to TCP when UDP is blocked.
Staff insight: The wins are concentrated where loss and RTT are high — mobile networks and long-haul links. Inside a data center, the benefit is small and the CPU cost of user-space UDP can be higher. Say where you'd use it, not just that it's newer.
Lyft: Envoy and the Service Mesh#
Lyft built Envoy as an L7 proxy to give every service consistent retries, timeouts, circuit breaking, outlier ejection, and observability without re-implementing them in every language. It was open-sourced in 2016 and became a CNCF graduated project and the data plane of several service meshes (e.g., Istio).
Staff insight: Envoy's origin story is an organizational one: polyglot services with inconsistent network behavior. That's the argument to make in an interview — a mesh is justified by the cost of inconsistency across teams, not by features any single service needs.
Staff Calibration#
What Staff Engineers Say (That Seniors Don't)#
| Concept | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Latency | "We'll add a CDN to make it faster" | "It's 3 RTT at 200ms before server time. Edge TLS termination cuts ~400ms, then regional read replicas remove the last trip for reads" | "Region placement is a one-way door. I'd price a second write region against the revenue lost to latency in APAC before committing" |
| Protocol | "gRPC is faster than REST" | "gRPC internally for schemas and deadlines; HTTP/JSON at the public edge for reach; we need client-side LB because of long-lived connections" | "The public protocol is a multi-year contract with external developers. Internally I'd standardize the RPC contract, not the library" |
| Timeouts | "Set a 2-second timeout" | "Propagate a deadline from the edge; each hop fails fast below its minimum useful budget" | "Deadline propagation goes into the paved-road RPC stack so 200 teams get it without thinking" |
| Retries | "Retry 3 times with backoff" | "One retry, one layer, idempotent only, 10% budget — otherwise 3 layers make 27× load" | "Retry storms are cross-team incidents; the retry budget is org policy with an exceptions process" |
| Tail latency | "Our p99 is 50ms" | "We fan out to 40 leaves, so the leaf p99 is our page p67. Hedge reads at p95, partial results for decorative modules" | "Tail SLOs are set per tier and owned by the aggregator team; leaf teams are budgeted, not left to guess" |
| Real-time | "Use WebSockets" | "SSE — one-way push, HTTP-native, resumable. WebSocket when we actually need bidirectional" | "A connection tier is a new stateful platform with its own on-call. I'd build one per company, not one per feature" |
Why "Tail latency" separates levels
The Senior answer reports a leaf metric accurately. The Staff answer knows that the user experiences a different distribution — the maximum of N leaf samples — and redesigns the aggregator around it. The Principal answer recognizes that tail latency is a contract between teams: if the aggregator team owns the page SLO but 40 leaf teams each own their own p99, nobody owns the composition. L7 fixes the ownership model — tier-level budgets and a shared mechanism for hedging and partial results — so the math stops being rediscovered in every incident review.
Why "Retries" separates levels
"Retry 3 times with exponential backoff" is textbook-correct for a single client. Staff engineers compute what happens when every layer does it. Principal engineers notice that the retrying team and the team that gets paged are different teams, so the fix cannot live in either team's code review — it has to be a platform default with an exceptions process.
Common Interview Traps#
- Drawing arrows as if they were free. Every cross-region arrow is 60–200ms. Say the number when you draw it.
- Forgetting the handshake. "One request, 80ms" is wrong for a cold client. It's 240–320ms.
- Putting gRPC behind an L4 load balancer and calling it balanced. Long-lived connections pin load; say how you rebalance.
- Choosing WebSockets by default. A stateful connection tier is a platform. SSE or polling is often enough.
- Retrying at every layer. Multiply the attempts; name the budget.
- Reporting averages. Averages hide the tail, and the tail is what fan-out amplifies. Report p50/p99/p99.9.
- Assuming DNS failover is instant. Resolvers and clients cache beyond TTL. Plan for minutes.
- Using 0-RTT for writes. 0-RTT data can be replayed by an attacker; restrict it to idempotent requests.
Practice Drill#
Prompt: "Our checkout API has a p50 of 40ms but a p99 of 900ms. It calls 12 downstream services in parallel. Leadership wants the p99 under 250ms this quarter. What do you do?"
Staff Answer
First I'd confirm that the p99 is dominated by the fan-out and not by one leaf: with 12 parallel calls, even a clean 1% tail per leaf gives ~11% of checkouts at least one slow call, which is enough to drag the checkout p99 up to the worst leaf's p99.9. I'd pull per-leaf latency histograms and a trace sample of the slowest 1% of checkouts to see whether one or two leaves own most of the tail (usual) or it's spread evenly (queueing or GC across the fleet). Then, in order: (1) propagate a 250ms deadline from the gateway so we stop doing work after the user has given up; (2) split the 12 calls into must-have (cart, pricing, tax, payment-auth eligibility, inventory) and nice-to-have (loyalty points, recommendations, promo banners) — nice-to-haves get a 100ms deadline and degrade to empty; (3) hedge the idempotent reads among the must-haves at their p95, watching hedge.win_rate and capping extra load at ~5%; (4) check for the classic tail sources — 200ms TCP retransmit clusters, 40ms Nagle stalls, GC pauses, leaf CPU above 70% — and fix the ones present; (5) remove retries below the gateway and add a 10% retry budget at the gateway. Payment authorization is never hedged or blindly retried; it gets an idempotency key and a longer deadline carved out of the budget. Owners: checkout team owns the budget and degradation policy; each leaf team gets a p99 target derived from it; platform owns the RPC library changes.
Why this is L6:
- Uses the fan-out formula to diagnose before prescribing.
- Separates must-have from nice-to-have with an explicit degradation contract.
- Refuses to hedge or retry non-idempotent calls, and says why.
- Assigns ownership of each budget to a named team.
What L7 adds:
- Turns the leaf-level p99 targets into an org-wide latency budget contract that is reviewed whenever a new dependency is added to checkout, so the problem doesn't regress next year.
- Prices the fix: ~5% extra read load from hedging vs the conversion lift from a 650ms p99 reduction (conversion data from product), and presents the tradeoff to the business in dollars.
- Pushes deadline propagation and retry budgets into the shared RPC stack so the next 50 services inherit the fix instead of rediscovering it.
Where This Appears#
- Load Balancer — L4 vs L7 balancing, connection vs request distribution, health checks, and draining
- API Gateway — Edge TLS termination, deadline propagation, and retry policy at the front door
- CDN & Edge Caching — Anycast, edge termination, and why "terminate close, travel warm" dominates global latency
- Real-Time Updates — WebSocket vs SSE vs long polling, connection tiers, and reconnect storms
- Chat Messaging — Persistent connections at millions-of-users scale and presence routing
- Circuit Breakers — Retry budgets, outlier ejection, and preventing cascading failure
- Service Discovery — Client-side load balancing and how clients find backends for long-lived connections
Related Foundations & Patterns: Back-of-Envelope Estimation · API Design Patterns · Real-time Updates · Degraded Mode Framework
Related Technologies: API Gateways