Hiring BarSupport

Design Real-Time Updates (WebSockets) — Staff-Level Case Study

Case study65 min read7 diagrams

Technologies referenced in this case study: Redis · Apache Kafka · Cassandra · API Gateways · ZooKeeper & etcd

Start with the pattern, then come here for the system: Real-time Updates — Cross-Cutting Pattern covers SSE vs WebSocket, the connection registry and presence basics. This case study goes deeper on what the pattern page doesn't: the connection gateway fleet, fan-out at scale, reconnect storms, and delivery guarantees.

Related case studies: Chat Messaging · Notification System · Collaborative Editing · Load Balancer · Message Queues · Degraded Mode

How to Use This Case Study#

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 4, 5
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 → Section 4 (reconnect storm, hot channel) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including the Principal Lens and appendices
What is a Real-Time Update System? — Why interviewers pick this topic

A real-time update system keeps millions of long-lived client connections open and pushes events — new messages, likes, typing indicators, price ticks, order status — to the right subset of them within a few hundred milliseconds. Slack, Discord, Facebook, Uber's rider app, trading dashboards and every "live" feature run one.

Unlike request/response systems, the expensive resource isn't requests — it's connections. A connection is state that lives on one specific host, for hours, and every one of them must be re-established when that host goes away.

Before vs After — gateway deploy scenario (10M concurrent connections):

Without a Staff-level design:
t=0:        Routine deploy restarts 20% of gateway hosts at once.
t=+1s:      2M clients see the socket close. All reconnect immediately.
t=+2s:      Load balancer receives 2M TLS handshakes. Remaining hosts' CPU hits 100%.
t=+5s:      Handshakes time out. Clients retry with a fixed 1s delay.
t=+10s:     Healthy hosts fail health checks under handshake load. More clients disconnect.
t=+30s:     8M clients in a reconnect loop. Auth service sees 400K token validations/s.
t=+4min:    Auth service falls over. Full outage — self-inflicted by a deploy.
t=+45min:   Recovered by blocking traffic at the LB and admitting it region by region.

With a Staff-level design:
t=0:        Deploy drains 2% of hosts per wave, 10 minutes per wave.
t=+0s:      Draining host sends 'reconnect' frame with a random 0–120s delay per client.
t=+2min:    200K clients reconnected over 120s = ~1.7K/s. p99 handshake 40ms.
t=+2min:    Each client resumes from its last sequence number. Zero events lost.
t=+5h:      Deploy complete. Nobody noticed. Nobody was paged.

Why interviewers reach for this question: It tests whether you understand stateful infrastructure — capacity measured in connections not QPS, failure measured in reconnect storms not error rates, and correctness measured in "did the user eventually see everything" not "did this push succeed."

Mechanics Refresher: Transport Options
TransportHow It WorksProsCons
Short pollingClient asks every N secondsTrivial, statelessLatency = N/2; wasteful at scale (10M clients / 5s = 2M req/s of mostly empty responses)
Long pollingServer holds request until an event or ~30s timeoutWorks everywhere, HTTP semanticsReconnect per message; ordering and dedupe are your problem
Server-Sent Events (SSE)One-way HTTP stream, text/event-stream, built-in Last-Event-ID resumeSimple, HTTP/2 multiplexed, auto-reconnect in browsersServer→client only; some proxies buffer
WebSocketFull-duplex framed TCP after HTTP upgradeBidirectional, low overhead per message (2–14 byte frames)Stateful, no built-in resume, proxy/LB idle timeouts
Mobile push (APNs/FCM)OS-level push channelWorks when app is backgroundedSeconds of latency, no delivery guarantee, payload limits (~4 KB)

For most production systems: WebSocket (or SSE if the client never sends) for foreground apps, OS push for backgrounded mobile, and a durable sync API underneath all of them. The transport is the easy decision; see Real-time Updates for the comparison. This case study is about what happens behind it.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What This Interview Actually Tests#

Real-time updates is not a WebSocket question. Everyone knows how to open a socket.

It is a stateful-fleet and delivery-contract question that tests:

  • Whether you size the system in connections, memory and handshakes — not requests per second
  • Whether you know where each event's source of truth lives (hint: not the socket)
  • Whether you can design fan-out that survives one channel with 1M subscribers
  • Whether you treat reconnection — after deploys, crashes, network blips — as the dominant load event

The key insight: The WebSocket is a cache of a durable log, not the delivery mechanism of record. Push is a latency optimization; the guarantee comes from a per-user or per-channel sequence number and a sync API the client calls on every reconnect. Design it that way and the gateway can be lossy, restartable and boring — which is exactly what you want from a fleet holding 10M connections.

The L5 → L6 → L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"WebSocket servers behind a load balancer, Redis pub/sub between them"Asks "what's the delivery contract — can the user miss an event, and what's the source of truth?"Asks "how many product teams will push through this, and are we building a platform or a feature?"
CapacityQPS-based sizingSizes in connections (~100–200K/host), memory (~20–50 KB/conn), and reconnect handshakes/sPrices connections: $/million-concurrent/month, and the cost of blast radius per host size
Fan-outPublish to every server; each filtersSubscription-aware routing: servers subscribe only to channels they have members for; hot channels get tiered fan-outSets org limits: max channel size, per-team publish quotas, a topic registry
Delivery"WebSockets are reliable, TCP guarantees delivery"At-most-once push + sequence numbers + sync-on-reconnect = no gapsPublishes the delivery contract as an SLO ("99.9% of events visible within 2s, 100% within sync") that product teams build against
Failure"Clients reconnect automatically"Designs against reconnect storms: jittered backoff, drain with spread, admission control, resumable sessionsDesigns the org's failure posture: cells so a bad deploy hits 2% of users, not 100%; game days for mass reconnect
OwnershipEach team runs its own socket serverPlatform owns the gateway; product teams own topics and payload schemasWrites the standard; decides what never goes over the socket (large payloads, source of truth)
Why "delivery" separates levels

L5: "WebSocket runs over TCP, so delivery is reliable." TCP guarantees in-order bytes on one connection while it lives. Messages sent to a socket that's half-dead (mobile client entered a tunnel; NAT dropped the mapping) are acknowledged by the kernel's send buffer and silently lost when the connection eventually closes. There is no end-to-end ack.

L6: "I'll treat the push as best-effort. Every event gets a monotonic sequence number per user (or per channel). The client stores the last sequence it saw. On reconnect, it calls sync(since=seq) against the durable store. Push gives us 200ms latency; sync gives us the guarantee."

L7: "And I'd make that contract the platform's API. Product teams shouldn't each invent their own sequence scheme — the gateway team publishes the guarantee, and a product that needs stronger semantics (payments status, say) is told explicitly that real-time is the wrong channel."

Why "failure" separates levels

L5: Trusts automatic reconnection. Automatic reconnection is the failure: 2M clients reconnecting within the same second is a DDoS against your own load balancers, TLS terminators and auth service.

L6: Designs the reconnect path as the peak load path: exponential backoff with full jitter on clients, server-directed reconnect delays during drains, token-bucket admission on handshakes per host, cheap session resumption (TLS session tickets, a short-lived resume token so auth isn't re-hit), and host sizing chosen partly for blast radius.

L7: Recognizes correlated failure across the org: one gateway fleet serving every product means one bad config push disconnects the entire company's user base. Introduces cells (e.g., 50 cells of 200K users each) and a progressive deploy policy that's enforced by tooling, not by discipline.

Why "capacity" separates levels

L5: "Each server handles 10K requests/s; we need N servers." Idle connections cost almost no CPU, so QPS-based sizing wildly underestimates memory and wildly misses the handshake peak.

L6: Sizes three things independently: steady connections per host (memory-bound), message throughput (CPU-bound, driven by fan-out), and reconnect handshake rate (CPU-bound, driven by failure). The last one sets the headroom.

L7: Chooses host size as a blast-radius decision: 1M connections per host is achievable and cheap, but losing that host means 1M simultaneous reconnects. Prices the extra hosts for 200K/host against the cost of a storm.

The Staff Positions#

PositionRationale
The socket is not the source of truthEvery event is written durably first; push is a latency optimization over a sync API
Sequence numbers + sync-on-reconnect over server-side per-connection queuesConnection-scoped buffers die with the host; a cursor into a durable log doesn't
Gateways are dumb and stateless beyond the socketSubscription and routing state lives in a separate layer so any gateway can be killed
Size hosts for blast radius, not max density100–200K connections/host, not 1M — a lost host must be a tolerable reconnect wave
Reconnect is the peak load, design for itFull-jitter backoff, server-directed drain delays, handshake admission control
Subscription-aware fan-out; special-case hot channelsBroadcasting every event to every gateway is O(events × gateways) and dies at ~100 hosts
WebSocket for foreground, OS push for backgroundDon't hold sockets for backgrounded mobile apps; the OS will kill them anyway

The Four Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Activity notifications (likes, badges, feed refresh hints)Huge scale, low value per eventBest-effort push, coalescing, drop under pressureMissed hint → stale badge until next syncEventually visible; loss acceptable
Conversational delivery (chat, comments, DMs)Ordered, no gaps, ~200msSequence per conversation/user, durable log, sync on reconnectGaps or duplicates visible to usersEvery event visible exactly once in order after sync
Collaborative sessions (docs, whiteboards, multiplayer)Low-latency bidirectional, per-room stateRoom affinity: route all members of a room to one session serverSession server loss = room state rebuildConvergence (OT/CRDT), not just delivery
Live broadcast (scores, live-stream comments, auctions)One publisher, 100K–10M subscribersTiered fan-out, sampling, edge relaysHot channel melts one tierLatest value wins; sampling acceptable

🎯 Staff Move: "I'll design a general real-time delivery platform for conversational delivery — chat-like semantics where users must not miss events — because it has the strictest contract. Notifications are the same pipeline with dropping allowed. Collaborative sessions need room affinity, which I'll call out as a different routing mode. Live broadcast I'll handle as a hot-channel special case."

The Five Fault Lines#

#Fault LineThe Tension
1Stateful Gateway vs Stateless GatewayKeep subscriptions and buffers on the host holding the socket (simple, fast) or externalize them (restartable, more hops)?
2Broadcast vs Subscription-Routed Fan-OutSend every event to every gateway (simple) or route by who's subscribed where (scales, but needs a routing layer)?
3Push Guarantees vs Pull SyncMake the socket reliable (acks, per-connection queues) or keep it best-effort and guarantee via cursor sync?
4Reconnect Speed vs Fleet ProtectionLet clients reconnect instantly (best UX for one) or throttle and spread (survives mass events)?
5Platform Gateway vs Product-Owned SocketsOne shared fleet for every team (efficiency, one blast radius) or per-product sockets (autonomy, duplication)?

In the Wild: Real Production Systems#

Why this section belongs here: Real-time fleets at scale converge on the same shape. Citing them shows you've seen the operational reality, not just the protocol.

Slack — Gateway Servers, Channel Servers, and an Edge Cache#

Slack has publicly described its real-time architecture: stateful Gateway Servers hold client WebSockets and subscribe on behalf of their clients; Channel Servers own channels via consistent hashing and fan messages out to subscribed gateways; separate services handle presence; and an edge cache ("Flannel") serves the boot/state payload clients need on connect, so reconnects don't hammer the primary databases.

Staff insight: The separation of "who holds the socket" (gateway) from "who owns the channel" (channel server) is the Fault Line 1 and 2 answer in production. And Flannel exists because reconnect is the expensive path — the boot payload, not the socket, is what melts backends.

Discord — Per-Guild Processes on the BEAM#

Discord runs its real-time gateway on Elixir/Erlang, with a process per guild (server) that fans out to member sessions. Their engineering blog has described scaling to millions of concurrent users and the special engineering needed for very large guilds — where a single guild's fan-out becomes the bottleneck — as well as work to compress gateway traffic to cut bandwidth.

Staff insight: Discord's big-guild problem is the hot channel fault line. The general design works until one channel has hundreds of thousands of members; then fan-out for that one channel needs its own architecture (relays, lazy member lists, sampled presence).

Netflix — Zuul Push#

Netflix open-sourced Zuul Push, a push notification service that holds persistent WebSocket/SSE connections for millions of devices. Their public talks describe the operational lessons: connections that live too long pin load to old hosts and make deploys dangerous, so they cap connection lifetime with randomization to continuously rebalance, and they emphasize that the herd of reconnecting clients is the main scaling risk.

Staff insight: Randomized connection lifetimes turn a rare, catastrophic mass reconnect into a constant, gentle trickle. That's a Staff move: convert a tail-risk event into routine background load you're always sized for.

(Also worth knowing: WhatsApp publicly reported handling around 2M concurrent connections on a single Erlang server in 2012. It proves density is achievable; it doesn't mean it's the right blast radius for you.)

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"WebSocket servers + Redis pub/sub""One channel has 2M subscribers. What happens?"Hot-channel fan-out
"Clients reconnect automatically""You deploy and 2M clients reconnect in the same second."Reconnect storm awareness
"TCP guarantees delivery""The phone went into a tunnel for 40s. Which messages did it miss, and how does it know?"Delivery contract
"We store connection → server in Redis""That's 10M entries churning on every blip. What's the write rate during a storm?"Registry cost under failure
"We'll scale to 1M connections per host""What happens when that host dies?"Blast-radius sizing
"Each team can publish events""A team ships a bug that publishes 50K events/s to every user."Platform guardrails

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Writes go to the durable path first — the Write API assigns a per-channel sequence number and persists before anything is pushed. The log feeds the subscription router, which knows which gateways have members of each channel and sends each event only there. Gateways are deliberately thin: they hold sockets, check resume tokens, and forward frames. On any reconnect, the client calls the Sync API with its last sequence number and fills gaps from the store. Backgrounded mobile clients are served by OS push from the same log.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Delivery"TCP is reliable""Push is best-effort. Per-channel sequence numbers + sync on reconnect make it gap-free."
Fan-out"Redis pub/sub to all servers""Subscription-routed: events go only to gateways with subscribers. Hot channels get a relay tier."
Capacity"N servers at 10K QPS""~200K connections/host at ~30 KB each; headroom set by reconnect handshakes/s, not steady state."
Reconnect"Auto-reconnect""Full-jitter backoff, server-directed drain delays, handshake admission control, resume tokens."
Deploys"Rolling restart""Drain 2% per wave with spread reconnects; randomized max connection lifetime keeps load rebalanced."
Ownership"Each service opens sockets""One gateway platform; product teams own topics and schemas under quotas."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Memory per idle WebSocket (tuned runtime)~10–50 KBSets connections per host: 200K × 30 KB ≈ 6 GB
Connections per host (practical)100K–500K; ~2M demonstrated (WhatsApp 2012)Density is possible; blast radius is the constraint
Heartbeat interval20–30sMust beat LB/NAT idle timeouts (commonly 60s; some mobile NATs shorter)
Full TLS handshake CPU (ECDSA)~1–2 ms CPU → ~1–3K/s per coreReconnect storms are CPU-bound on handshakes
TLS resumption~5–10× cheaper than full handshakeSession tickets are storm insurance
Reconnect spread for a drained host60–120s200K conns / 120s ≈ 1.7K/s — absorbable
Target push latencyp99 < 500ms publish → on-screenBeyond ~1s users perceive "not real-time"
WebSocket frame overhead2–14 bytesWhy small events are cheap on the wire
Mobile push payload limit~4 KBOS push carries hints, not data
Registry churn during a 1M-conn storm~1M writes + deletes in minutesThe registry must survive the storm it's tracking
Fan-out cost of naive broadcastevents/s × gateways50K events/s × 200 hosts = 10M deliveries/s of mostly-dropped messages

Interview Walkthrough

The most common mistake: Candidates spend 15 minutes on WebSocket vs SSE vs long polling. The interviewer assumes you know. Pick one in 30 seconds and spend the time on the fleet, fan-out, reconnects and the delivery contract.


Phase 1: Requirements & Framing (2–3 minutes)#

Functional scope in one breath:

"Clients keep a live connection and receive events for the channels, conversations and objects they care about within about half a second. Clients can also send small events — typing, acks, presence — upstream."

Then the non-functionals that drive the design:

"The key question is the delivery contract. Can a user miss an event? For chat-like delivery, no — so I'll make the socket a best-effort accelerator over a durable, sequenced log, and guarantee completeness through sync on reconnect. Second, this is a stateful fleet: I'll size it in connections and reconnect handshakes, not QPS."

Commit to numbers:

"Assume 50M DAU, 10M peak concurrent connections, 1M published events/s, average fan-out 10 — so ~10M deliveries/s. Push latency p99 under 500ms. And one hot-channel case: a live event with 2M subscribers."

🎯 Staff Move: "The transport choice is the least interesting part of this problem. I'll use WebSockets for foreground clients and OS push for background, and spend our time on what breaks: fan-out, reconnects and the delivery guarantee."


Phase 2: Core Entities & API (1–2 minutes)#

  • Connection (conn_id, user_id, device_id, gateway_host, connected_at, resume_token)
  • Subscription (conn_id, channel_id) — derived from membership, not stored per connection durably
  • Event (channel_id, seq, type, payload_ref or small payload, created_at)
  • Cursor (client-side: last_seq per channel, or a per-user inbox seq)

API surface:

WS  /connect?resume_token=…           → HELLO { conn_id, heartbeat_ms: 25000 }
WS  ← EVENT { channel, seq, type, body }      (server → client)
WS  → ACK { channel, seq } (optional, for metrics only)
WS  ← RECONNECT { after_ms: 0–120000 }         (server-directed drain)
GET /sync?cursors=ch1:1042,ch2:77&limit=500   → { events[], more: bool }
POST /channels/{id}/events { … }               → 201 { seq }   (write path, HTTP not socket)

🎯 Staff Move: "Writes go over HTTP to the write API, not over the socket. The socket is for delivery. That keeps the gateway free of business logic and lets me restart it without losing writes."


Phase 3: High-Level Architecture (≤5 minutes)#

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk it in 60 seconds:

  1. A producer writes an event through the write API, which assigns the next per-channel seq and persists it.
  2. The event is appended to the log; the subscription router consumes it.
  3. The router looks up which gateways have subscribers for that channel and forwards only to them.
  4. Each gateway writes the frame to its local sockets for that channel.
  5. If a client was disconnected, on reconnect it calls sync with its cursors and gets anything it missed.

🎯 Staff Move: "That's the whole data path. It's correct on a quiet day. It breaks in three ways at scale: one channel with 2M subscribers, 2M clients reconnecting at once, and gaps nobody notices. Let me take those."


Phase 4: Transition to Depth (1 minute)#

"Three deep dives: fan-out routing including hot channels, the reconnect storm — which is the real peak load of this system — and the delivery contract with sequence numbers. I'd start with reconnects, because that's the failure that turns a routine deploy into a company-wide outage. Preference?"


Phase 5: Deep Dives (25–30 minutes)#

Deep dive 1: Reconnect storms (8–10 min)

"Steady state is easy — idle sockets cost memory, not CPU. The CPU peak is reconnection: a full TLS handshake costs ~1–2ms of CPU, and re-establishing session state costs an auth check plus a boot payload. If 2M clients reconnect in 2 seconds, that's 1M handshakes/s — orders of magnitude over what the fleet is sized for."

Five controls, in order of leverage:

ControlWhat It DoesNumber
Client full-jitter backoffsleep(random(0, min(cap, base × 2^n)))base 1s, cap 60s
Server-directed drainDraining host sends RECONNECT{after_ms} with a random delay per clientspread over 60–120s
Handshake admission controlPer-host token bucket on new connections; excess gets 503 + Retry-After faste.g., 2K new conns/s/host
Cheap resumptionTLS session tickets + a signed short-lived resume token (no auth service call)5–10× cheaper
Randomized max connection lifetimeEvery connection closes after 2–4h ± jittercontinuous rebalancing

"The last one is what Netflix described for Zuul Push: if connections recycle constantly, a mass reconnect is just a faster version of normal, and new hosts pick up load without a thundering herd."

Deep dive 2: Fan-out (8–10 min)

"Naive design: every gateway subscribes to one Redis pub/sub bus and filters. That's O(events × gateways): 1M events/s × 50 hosts = 50M messages/s, 95% discarded. Instead, the subscription router tracks channel → set of gateways with at least one subscriber, maintained by gateways as clients join/leave. An event goes to ~k gateways where k is the number of hosts with members — for a 10-person chat, usually 1–10."

"Hot channels flip the math: a live event with 2M subscribers spread across all 50 hosts. Per-event cost is 50 router sends and 2M socket writes — fine for one event, fatal for 1K events/s (2B writes/s). So hot channels get: server-side coalescing (send the latest score every 500ms, not every change), sampling for comments ('show 20/s of 10K/s'), and a relay tier so the router sends once per region."

Deep dive 3: Delivery contract (6–8 min)

"Every event gets a monotonic seq per channel at write time. The client tracks last_seq per channel. Three cases: if a pushed event is last_seq + 1, apply it. If it's higher, there's a gap — call sync(since=last_seq). If it's lower or equal, it's a duplicate — drop it. On every reconnect, sync all channels the client cares about (batched, one request). This gives effectively-once, in-order visibility without any per-connection durable queue on the gateway."


Phase 6: Wrap-Up (2–3 minutes)#

"The design principle is that the socket is a cache of a durable log. That makes gateways stateless enough to kill, which makes deploys and failures routine; the durable path gives the guarantee; and the expensive event — reconnection — is designed for explicitly with jitter, drain spreading, admission control and cheap resumption."

Organizational closer:

"The long-term risk is that every product team pushes through this fleet. I'd run it as a platform with per-topic quotas and a topic registry, and cells so a bad config affects 2% of users. The most important thing I'd publish is the delivery contract — so teams don't put payments status on a best-effort channel and call it real-time."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Transport debate10 min on WebSocket vs SSE vs long pollPicks in 30s, links the tradeoff, moves on
QPS sizing"10K QPS per server"Connections/host, memory/conn, handshakes/s
Reconnect ignored"Clients auto-reconnect"Treats reconnect as the peak and designs for it
Delivery hand-waved"TCP is reliable"Sequence numbers + sync + gap detection
One fan-out designPub/sub to all serversSubscription routing + hot-channel special case
No ownership storyStops at architecturePlatform gateway, topic quotas, cells, delivery contract

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Most system design prompts are stateless request/response problems where scaling means "add more boxes." Real-time delivery is the canonical stateful prompt: each connection is pinned to a host for hours, so every failure, deploy and scaling action has a reconnection cost. Interviewers use it to find candidates who have carried a pager for a stateful fleet — the ones who know that the dangerous moment isn't peak traffic, it's the deploy at 2pm on a Tuesday.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"A user's phone loses signal for 40 seconds in the middle of a busy group chat. Walk me through, message by message, how their client ends up showing exactly the right conversation — no gaps, no duplicates, correct order."

This single question exposes whether the candidate knows that TCP send buffers lie, that the server can't tell a dead connection from a quiet one until heartbeats fail, and that the only robust answer is sequence numbers plus a sync path.


2. Problem Framing & Intent#

2.1 The Four Intents — Explained#

Activity notifications → best-effort, coalesced

  • Constraint: very high volume, low value per event (like counts, "someone's typing")
  • Strategy: push without durability for ephemeral events; coalesce per client (at most 1 badge update per 2s); drop under backpressure
  • Failure mode: stale badge until the next page load
  • Who pays for imperfection: nobody meaningfully — product signs off that ephemeral events are lossy

Conversational delivery → gap-free, ordered

  • Constraint: user-visible messages must all appear, in order, within ~500ms when online
  • Strategy: durable write + per-channel seq + push + sync on reconnect
  • Failure mode: gap if the client doesn't detect missing seq; duplicates if it doesn't dedupe
  • Who pays for imperfection: users (missed messages), support volume, trust

Collaborative sessions → affinity, convergence

  • Constraint: bidirectional, 50–100ms, shared mutable state per room
  • Strategy: route all participants of a document/room to the same session server (consistent hash on room_id); the session server is authoritative for ordering; see Collaborative Editing
  • Failure mode: session server loss forces room reload from snapshot + op log
  • Who pays: users in that room see a 1–3s freeze

Live broadcast → latest-value, sampled

  • Constraint: 100K–10M subscribers per channel, one or few publishers
  • Strategy: coalescing, sampling, relay tiers, sometimes CDN-delivered SSE or HTTP polling for the long tail
  • Failure mode: hot channel saturates router or gateway CPU
  • Who pays: every other channel sharing those gateways, unless isolated

🎯 Staff Move: "These four share a gateway fleet but not a delivery contract. I'd tag every topic with its class — ephemeral, durable, session, broadcast — and the platform enforces different rules per class: durable topics must have sequence numbers and a sync endpoint; broadcast topics must declare a coalescing interval."

2.2 When NOT to Use WebSockets#

  • Updates every 30s or slower are fine. Poll. 10M clients polling every 60s is ~170K req/s of cacheable GETs — often cheaper and far simpler than a stateful fleet.
  • Server→client only, web clients. SSE gives you resume (Last-Event-ID) and HTTP semantics for free.
  • The app is backgrounded. Mobile OSes kill background sockets; use APNs/FCM with a hint payload and sync on open.
  • Correctness-critical state (payment status, order confirmation). Deliver a hint over the socket if you like, but the client must read the authoritative state from an API. Never make the socket the system of record.
  • Very large payloads (files, images). Send a reference; fetch via CDN. See Handling Large Blobs.

2.3 What the Interviewer Leaves Underspecified#

  • Delivery semantics. Can users miss events? Must they be ordered? Per channel or globally?
  • Multi-device. Does a user have 3 devices each needing every event, with read-state synced?
  • Channel size distribution. Median 8 members, but the tail (the company-wide channel, the live event) decides the design.
  • Upstream traffic. Typing indicators and presence can exceed message volume 10×.
  • Offline behavior. Mobile push? How long do we retain events for sync — 7 days? 30?
  • Geography. Single region or global users connecting to the nearest region?

2.4 Precise Terminology#

TermWhat It MeansWhy the Precision Matters
GatewayHost terminating client socketsIts only job is holding connections; keep it thin
Subscription routerMaps channel → gateways with subscribersDecides fan-out cost
Connection registryMaps user/conn → gatewayNeeded for user-targeted sends; churns during storms
Sequence number (seq)Monotonic per channel (or per user inbox)Enables gap detection and idempotent apply
SyncPull API returning events after a cursorThe actual delivery guarantee
DrainGracefully moving a host's connections elsewhereDeploys are drains, not restarts
Reconnect stormMany clients reconnecting in a short windowThe system's true peak load
Hot channelA channel whose subscriber count or event rate dominates a tierNeeds its own fan-out strategy
CellAn independent slice of the fleet serving a subset of usersLimits blast radius

3. The Five Fault Lines#

3.1 Fault Line 1: Stateful Gateway vs Stateless Gateway#

The tension: The host holding a socket is the cheapest place to keep per-connection state — subscriptions, unacked buffers, session context. It's also the one place guaranteed to disappear.

StrategyWhat WorksWhat BreaksWho Pays
Fat gateway (subscriptions, buffers, business logic on the socket host)Fewest hops, lowest latency, simplest codeEvery deploy/crash loses state; gateways can't be killed freely; product logic in the gateway means every team deploys itOn-call (risky deploys), users (lost buffered events)
Thin gateway (sockets + auth + frame forwarding; subscriptions derived and re-registered on connect)Hosts are cattle; deploys are drains; one platform codebaseExtra hop to router; re-registration cost on reconnectPlatform team maintains router; ~1–5ms extra latency
Session-affine gateway (for collaborative rooms)Authoritative per-room ordering in memoryRoom reload on host lossUsers in that room (1–3s freeze)

Staff default: Thin gateways for delivery. The only state on the gateway is the socket and an in-memory index of which local connections subscribe to which channels — rebuildable in seconds from membership when the client reconnects. Collaborative rooms use a separate session tier with explicit affinity, not the delivery gateway.

When to deviate: Small scale (< 100K connections, one team) — a fat gateway with a single deploy unit is fine and faster to build.

🎯 Staff Move: "I want to be able to kill any gateway at any time with no data loss. That's the test for what's allowed to live on it: the socket, yes; anything that isn't reconstructible on reconnect, no."


3.2 Fault Line 2: Broadcast vs Subscription-Routed Fan-Out#

The tension: Broadcasting every event to every gateway is trivially correct and scales as O(events × gateways). Routing only to gateways with subscribers scales with actual interest but needs a routing layer whose state churns with every connect/disconnect.

Back-of-envelope:

Events/s published:         1,000,000
Gateways:                   50 (at 200K conns each for 10M)
Broadcast deliveries/s:     50,000,000 gateway messages — ~95% dropped on arrival
Routed (avg 3 gateways/ev): 3,000,000 gateway messages
Router state:               channels with ≥1 online member × gateways ≈ 20M entries
Router churn (steady):      ~10M conns / 3h avg lifetime ≈ 1K conn events/s × ~20 channels = 20K updates/s
Router churn (storm):       1M reconnects in 120s ≈ 8K conns/s × 20 = 160K updates/s
StrategyWhat WorksWhat BreaksWho Pays
Global pub/sub broadcastSimple; works to ~10–20 gatewaysNetwork and CPU explode linearly with fleet sizeInfra bill; eventually latency
Pub/sub topic per channel (gateways subscribe to channel topics)Routing is the pub/sub system's jobMillions of topics; subscribe/unsubscribe storms during reconnectsPub/sub cluster on-call
Channel-owner servers (consistent-hashed owners hold channel → gateways)Scales horizontally; owner can do per-channel logic (ordering, coalescing)Owner failover must rebuild membership; hot channel pins one ownerPlatform team
Hot-channel relay tier2M-subscriber channels become 1 send per relay, relays fan to gatewaysExtra tier, extra latency (~10–50ms)Platform team; only for tagged channels

Staff default: Channel-owner servers (Slack's channel-server shape) for normal channels, sharded by consistent hash on channel_id — see Consistent Hashing. Channels above a threshold (e.g., > 10K online subscribers or > 100 events/s) are promoted to broadcast class: coalesced, sampled, and relayed.

Diagram: 3.2 Fault Line 2: Broadcast vs Subscription-Routed Fan-Out

🎯 Staff Move: "Fan-out is designed for the median channel and special-cased for the tail. The median chat has 8 members on 3 gateways. The tail has 2 million on all of them — and for that one I stop promising every event and start promising the latest state."


3.3 Fault Line 3: Push Guarantees vs Pull Sync#

The tension: You can make the socket reliable (server-side per-connection queues, client acks, retransmit) or keep it best-effort and put the guarantee in a pull path.

StrategyWhat WorksWhat BreaksWho Pays
Fire-and-forget push onlySimplest, fastestSilent loss on half-open connections, deploys, crashesUsers (missed events), support
Per-connection durable queue + acksReliable while the host livesQueue dies with the host; must replicate or persist per connection → heavy; multi-device duplicatesPlatform (complexity), storage bill
Best-effort push + per-channel seq + sync on reconnect/gapGuarantee lives in the durable store; gateways stay statelessNeeds a sync API that can handle storm load; clients must implement gap logicClient teams (logic), store (sync load)
Per-user inbox seq (one ordered stream per user)One cursor per client; simple syncWrite amplification: one event × N recipients inbox writesStorage; write path cost

Staff default: Best-effort push + sequence numbers + sync. Use per-channel seq when channels are small-to-medium and the client tracks < ~1K channels; use a per-user inbox seq for notification-style feeds where one cursor per client is valuable and fan-out is bounded.

Diagram: 3.3 Fault Line 3: Push Guarantees vs Pull Sync

The subtle part — the race between sync and live push: subscribe first, then sync. If you sync first, events written between the sync read and the subscription land in neither. Subscribe, buffer live events, sync, then apply buffered events deduped by seq.

🎯 Staff Move: "I don't make the socket reliable. I make it unimportant. If every push can be lost and the user still sees everything after sync, then gateways can crash, deploy and drain freely — and that operational freedom is worth more than any ack protocol."


3.4 Fault Line 4: Reconnect Speed vs Fleet Protection#

The tension: The individual user wants to reconnect in 100ms. The fleet wants reconnects spread over minutes. Both can't win during a mass event.

StrategyWhat WorksWhat BreaksWho Pays
Immediate reconnectBest single-user UXSynchronized herds; self-DDoSEveryone during mass events
Fixed-delay retrySimpleHerds stay synchronized — every client retries at t+1s, t+2sEveryone
Exponential backoff with full jitterDesynchronizes herdsUnlucky users wait up to the cap (60s)Some users wait longer
Server-directed reconnect delayServer knows the load; spreads preciselyOnly works for graceful drains, not crashesUsers on drained hosts (up to 120s of degraded, push-less state — sync covers it)
Admission control on handshakesProtects healthy hosts from collapseRejected clients see errorsClients beyond the admission rate

Staff default: All three together — full-jitter backoff (base 1s, cap 60s) in every client SDK, server-directed RECONNECT{after_ms} during drains (spread 0–120s), and per-host admission control (token bucket, e.g., 2K new connections/s/host, fast 503 + Retry-After). Plus cheap resume: TLS session tickets and a signed resume token (HMAC, 10-minute TTL) so reconnection doesn't hit the auth service.

While a client is waiting to reconnect, it isn't blind: it can poll the sync API every 10–30s. The UX degrades from 200ms to ~15s, not to "broken."

🎯 Staff Move: "I'm deliberately making some users wait up to two minutes to reconnect after a deploy, because the alternative is that everyone waits forty-five minutes after the fleet falls over. And because sync covers the gap, 'waiting' means a slightly delayed message, not a missing one."


3.5 Fault Line 5: Platform Gateway vs Product-Owned Sockets#

The tension: Every product wants real-time. If each builds its own socket server, the client holds 4 connections and the company runs 4 fleets with 4 reconnect-storm postures. If one platform owns it, one bad config push disconnects everything.

ModelWhat WorksWhat BreaksWho Pays
Per-product socket serversAutonomy, independent deploysMultiple sockets per client (battery), duplicated storm engineering, inconsistent guaranteesMobile battery; N teams on-call
One platform gateway, sharedOne socket per client; one team masters stormsShared blast radius; platform becomes a bottleneck for new event typesPlatform team; everyone during incidents
Platform gateway, cellular, with topic registry and quotasOne socket, bounded blast radius, self-serve topicsNeeds a registry, schema governance, quota toolingPlatform team (upfront investment)

Staff default: One multiplexed platform gateway, deployed in cells (e.g., 50 cells × 200K users), with a topic registry: each topic declares owner team, delivery class, max event rate, payload size cap (e.g., 16 KB), and schema. Publish quotas are enforced at the router. Product teams own their topics; the platform owns the fleet and the contract.

🎯 Staff Move: "The platform owns the socket and the guarantee; product teams own the topic and the payload. A team that wants to push 50K events/s to every user files a registry change and gets a quota review — not a surprise at 3am."


4. Failure Modes & Operational Reality#

4.1 Reconnect Storm After a Gateway Crash — Full Timeline#

Scenario: A memory leak in a new gateway build crashes 10 hosts (2M connections) within 30 seconds of each other at peak.

t=0:       10 gateways OOM. 2M sockets close (clients get RST or nothing).
t=+0–5s:   Clients that saw the close start jittered backoff (0–1s first attempt).
t=+1s:     ~1M reconnect attempts in the first second window (jitter spreads the rest).
t=+1s:     LB distributes to 40 healthy hosts → 25K handshakes/s/host vs 2K admission rate.
t=+1s:     Admission control rejects excess fast with 503 + Retry-After: 5–30s (randomized).
t=+2s:     Healthy hosts' CPU at 70% — handshake work bounded by admission rate.
t=+5s:     Clients with dead TCP (no RST) detect it only via missed heartbeat: +25–50s.
t=+30s:    Autoscaler adds hosts. Resume tokens mean auth service sees < 5K validations/s.
t=+90s:    1.6M reconnected. Sync API p99 = 300ms at 20K req/s (edge-cached boot payload).
t=+3min:   All reconnected. Zero events lost (sync filled gaps). Nobody else disconnected.

Detection: gateway.connections_active (sudden drop by host), gateway.new_connections_per_sec, gateway.admission_rejected_total, sync.requests_per_sec, auth.validations_per_sec.

Mitigation: Admission control and backoff do the work automatically; on-call rolls back the build and ensures the autoscaler isn't adding hosts running the bad build.

Prevention: Canary new gateway builds on one cell for 24h (leaks show up over hours); memory ceiling alerts per host; cap per-host connections so any single crash is ≤ 1% of users.

Owner: Gateway platform on-call. Sync API owner must be looped in — sync is the second system that sees the storm.


4.2 Hot Channel Meltdown#

Scenario: A celebrity joins a public live Q&A channel. 1.5M users subscribe in 10 minutes. The chat rate reaches 20K messages/s.

t=0:       Channel has 50K subscribers; owner server handles it fine.
t=+5min:   800K subscribers across all 50 gateways; 8K msgs/s.
           Owner server: 8K × 50 gateway sends = 400K sends/s → CPU 100%.
t=+6min:   Every other channel hashed to that owner (~2% of all channels) sees 5s delays.
t=+7min:   Gateways: 8K msgs/s × 16K local subscribers each = 128M socket writes/s fleet-wide.
t=+8min:   Gateway CPU saturates; heartbeats delayed; clients time out and reconnect → storm.

Detection: router.channel_events_per_sec top-K, router.owner_cpu, gateway.socket_writes_per_sec, push_to_visible_ms_p99 for unrelated channels (collateral damage signal).

Mitigation: Promote the channel to broadcast class at runtime: migrate it off the shared owner to a dedicated relay, coalesce to at most 20 displayed messages/s with sampling, and disable per-message push in favor of batched frames every 500ms.

Prevention: Automatic promotion thresholds (> 10K online subscribers or > 100 events/s); isolate broadcast-class channels on their own owner pool; per-channel rate caps in the topic registry.

Owner: Platform (automatic promotion); product (decides sampling policy — "users see a sample of messages in huge channels" is a product decision).


4.3 Silent Gaps — The Client That Never Syncs#

Scenario: A client release has a bug: on reconnect it re-subscribes but skips the sync call when the app resumes from background. Users report "missing messages" days later.

Detection: Server-side: sync.requests_per_reconnect ratio by client version (should be ~1.0; buggy version shows 0.3). Client-side: client.gap_detected_total (events arriving with seq > last_seq + 1) and client.gap_filled_total — the difference is the silent loss.

Mitigation: Force-upgrade or server-side flag that sends a SYNC_REQUIRED frame after every resume.

Prevention: The client SDK owns the seq/sync logic — product client code can't skip it; contract tests in CI that kill the socket and assert no gaps.

Owner: Client platform team (SDK). This is why the delivery contract must be implemented once, in a shared SDK.


4.4 Half-Open Connections and the Heartbeat Trap#

Scenario: A mobile carrier's NAT drops idle mappings after 30s. Your heartbeat is 45s. Connections look alive on the server but packets go nowhere.

Detection: gateway.heartbeat_timeouts_total by ASN/carrier; push_ack_ratio (if clients ack optionally) dropping for specific networks; server-side send-buffer growth per connection.

Mitigation: Lower heartbeat to 20–25s (costs 10M / 25s = 400K tiny frames/s fleet-wide — cheap); detect dead connections within 2 missed heartbeats.

Prevention: Heartbeat interval configurable per client platform; per-connection send buffer cap (e.g., 256 KB) — exceeding it closes the connection rather than letting memory grow.

Owner: Gateway platform team.


4.5 Connection Registry Overload During a Storm#

Scenario: User-targeted sends look up user → gateways in a Redis registry. During a 2M-connection storm, every connect/disconnect is a write.

Detection: registry.write_latency_p99, registry.ops_per_sec, router.lookup_miss_total.

Mitigation: Batch registry updates per gateway (every 1s, not per connection); tolerate stale entries (a send to a gateway without the user is a cheap no-op); rate-limit registry writes during storms.

Prevention: Prefer channel-owner routing (gateway registers interest in channels, batched) over per-user registry for most traffic; shard the registry; TTL entries with heartbeat refresh (every 60s) so crashes self-clean without delete storms.

Owner: Platform team. Key lesson: the system tracking the storm must survive the storm.


4.6 Slow Consumers and Head-of-Line Blocking#

Scenario: A client on a 2G connection can't drain its socket fast enough; the gateway's write buffer for it grows, and if writes are synchronous per channel, other clients wait.

Detection: gateway.conn_send_buffer_bytes p99, gateway.write_queue_depth, per-host event-loop lag.

Mitigation: Non-blocking writes with per-connection bounded buffers; on overflow drop ephemeral events first, then close the connection (client will sync on reconnect).

Prevention: Class-aware buffers: ephemeral events are droppable, durable events are recoverable by sync — so the gateway can always drop.

Owner: Platform team.


4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Gateway crashconnections_active drop per hostConnections on that host (≤ 1% by design)Admission control, jittered reconnect, syncGateway platform
Bad deploy across fleetError/crash rate on new buildUp to one cell per waveCell-by-cell rollout, auto-rollbackGateway platform
Hot channelTop-K channel event rate, owner CPUChannels sharing the owner, gateways with membersPromote to broadcast class, coalesce, samplePlatform + product
Silent client gapssync_per_reconnect by client versionUsers on buggy versionServer-forced sync frame, SDK fixClient platform
Half-open connectionsHeartbeat timeouts by ASNUsers on affected networksShorter heartbeat, buffer capsGateway platform
Registry overloadRegistry p99, ops/sUser-targeted sends delayedBatch updates, tolerate stalePlatform
Sync API overload after stormsync.p99, sync.rpsAll reconnecting clientsEdge cache boot payload, rate-limit sync per clientSync API owner
Publish flood from one teamrouter.publish_rate{topic}Everyone sharing routerPer-topic quotas, circuit-break topicTopic owner team

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Delivery contract"TCP/WebSocket is reliable"Best-effort push + seq + sync; subscribe-then-sync race handledPublishes the contract as a platform SLO; forbids system-of-record data on the socket
Capacity modelQPS per serverConnections/host, memory/conn, handshakes/s as the headroom driverHost size as a priced blast-radius decision; $/million concurrent
Fan-outPub/sub broadcastSubscription-routed with channel owners; hot-channel promotionOrg limits on channel size and publish rates; registry governance
FailureAuto-reconnectJitter, drain spreading, admission control, resume tokens, randomized lifetimesCells, progressive rollout enforced by tooling, mass-reconnect game days
OwnershipTeam-owned socket serverPlatform gateway; product-owned topicsDecides centralization boundary; funds a client SDK as the contract carrier
EvolutionAdd serversMulti-region with regional gateways and sync3-year path from single fleet to cellular multi-product platform

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Separates delivery from durability"The socket is a cache of the log. Push is latency; sync is the guarantee."
Sizes for failure, not steady state"Idle sockets cost memory. The CPU peak is reconnection, so I size for handshakes/s."
Treats deploys as the main risk"A deploy is a drain. I'd spread reconnects over 120 seconds and do 2% per wave."
Special-cases the tail"Median channel has 8 members. The 2M-member channel gets coalescing and a relay tier."
Handles the race"Subscribe first, then sync, then apply buffered events deduped by seq."
Names ownership"Platform owns the socket and the guarantee; product teams own topics under quotas."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
10+ minutes on WebSocket vs SSE vs pollingSpends time on the least differentiating decision
"TCP guarantees delivery"Misunderstands half-open connections and send buffers
Reconnection never discussedMisses the dominant failure mode of a stateful fleet
Every event broadcast to every serverDoesn't scale past ~20 hosts; no awareness of fan-out cost
Per-connection queues as the guaranteeState dies with the host; doesn't handle multi-device
Sockets used for writes with business logicCouples gateway deploys to every product change

5.4 Common False Positives#

  • Knowing the WebSocket frame format ≠ designing a fleet. Opcode trivia doesn't predict operational judgment.
  • "We'll use a managed service" ≠ design. Fine as a build-vs-buy answer, but the candidate still needs to state the delivery contract and reconnect behavior.
  • Redis pub/sub fluency ≠ fan-out design. The question is routing cost at 50+ hosts and hot channels, not the API.
  • Mentioning Kafka ≠ durability story. The log is only useful if the client has a cursor into it and a sync path.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing + intent0–4 minDelivery contract, numbers, transport in 30s
Entities + API4–6 minSeq per channel, sync API, writes over HTTP
High-level design6–11 minGateway, router, log, store, sync
Deep dive: reconnect storms11–20 minJitter, drain, admission, resume
Deep dive: fan-out + hot channels20–30 minChannel owners, promotion, coalescing
Deep dive: delivery30–40 minSeq, gaps, dedupe, subscribe-then-sync
Wrap-up40–45 minPlatform, cells, contract

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Direction
"Add presence (online/offline)."Upstream write amplificationHeartbeat-driven TTL, coarse status, fan-out presence only to open conversations; see the pattern page
"Add typing indicators."Ephemeral classNon-durable, coalesced to 1 per 3s per user, droppable
"A user has 5 devices."Multi-device deliveryPer-device cursors; send to all device connections; read-state as its own synced event
"Now go multi-region."Stateful geo designRegional gateways; cross-region event replication; sync from nearest replica with seq
"What if Kafka is down?"Durable path dependencyWrites fail (fail-closed for durable topics); ephemeral topics bypass the log
"Mobile app in background?"Transport boundariesOS push hint + sync on foreground; close the socket

6.3 What to Deliberately Skip#

  • The WebSocket handshake details and frame format. One sentence.
  • Message storage schema in depth. (channel_id, seq) clustered key in a wide-column store is enough; link to Chat Messaging.
  • Presence internals. Covered in Real-time Updates; mention TTL-based presence and move on.
  • Authentication flows. Token at connect, resume token after; that's it.

6.4 Follow-Up Questions to Expect#

  1. "How do you know a connection is dead?" — (Heartbeats every 25s, dead after 2 misses; send-buffer cap closes stuck connections.)
  2. "Where does ordering come from?" — (Per-channel seq assigned by the write path — a single sequencer per channel, e.g., the channel owner or a DB counter; no global order.)
  3. "How big can one host get?" — (Hundreds of thousands to ~1M+ is technically feasible; I'd cap at ~200K for blast radius.)
  4. "How do you deploy the gateway?" — (Drain cell by cell, 2% of hosts per wave, server-directed reconnect with 0–120s spread.)
  5. "What does the load balancer do?" — (L4 pass-through or TLS termination with session tickets; least-connections balancing — round robin skews toward new hosts badly after storms.)
  6. "What if the sync API is down?" — (Clients keep receiving live pushes; gaps are detected and retried with backoff; alert on sync.error_rate.)
  7. "How do you rebalance after scaling up?" — (Randomized max connection lifetime of 2–4h moves load gradually; never mass-disconnect to rebalance.)

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design a system to push real-time updates to our users."

Staff Answer

"First, the delivery contract: can users miss events? For chat-like updates, no; for likes and typing, yes. I'll design for the strict case and let ephemeral events opt out of durability. The transport is WebSocket for foreground clients and OS push for background — I won't spend time there. The key principle is that the socket is a cache of a durable log: every durable event gets a per-channel sequence number when written, push is best-effort, and clients sync from their last sequence on every reconnect. Numbers: 10M concurrent, 1M events/s, fan-out 10, p99 push latency under 500ms. I'll cover the gateway fleet, fan-out routing including hot channels, and reconnect storms — which are the real peak load."

Why this is L6:

  • States the delivery contract before drawing anything
  • Dismisses the transport question quickly and correctly
  • Names reconnect storms as the peak before being asked

What L7 adds:

  • "I'd also ask how many product teams will publish through this — if more than one, it's a platform with a topic registry and quotas."
  • Frames host sizing as a blast-radius and cost decision
❌ Common L5 Trap

"We'll use WebSockets because they're bidirectional. Let me compare them with SSE and long polling... Each server subscribes to Redis pub/sub, and when a message comes in we send it to the connected users."

Why this misses: A correct small-scale design; no delivery contract, no reconnect story, fan-out that's O(events × servers). The interviewer now has to pull every Staff-relevant topic out one question at a time.


Drill 2: Capacity Math#

Prompt: "How many gateway hosts do you need for 10M concurrent connections?"

Staff Answer

"Three independent constraints. Memory: ~30 KB per connection with a tuned runtime → 200K connections ≈ 6 GB, fine on a 32 GB host. Steady CPU: 10M deliveries/s across the fleet — at ~200K frame writes/s per host that's ~50 hosts. Reconnect headroom: if the worst tolerated event is losing 5% of the fleet at once (500K clients) and we spread over 60s, that's ~8K handshakes/s, which the rest of the fleet absorbs easily with ECDSA and session tickets. So ~50 hosts at 200K each, plus 20–30% headroom → ~65 hosts. I could run 10 hosts at 1M each, but then one crash is 10% of users reconnecting — I'd rather pay for 55 extra machines."

Why this is L6:

  • Separates memory, throughput and reconnect constraints
  • Chooses density with blast radius in mind and prices it

What L7 adds:

  • Expresses it as cost per million concurrent per month and compares with a managed service
  • Aligns host count with cell design (e.g., 13 cells × 5 hosts)

Drill 3: Make the Delivery Guarantee Concrete#

Prompt: "How do you guarantee a user sees every message?"

Staff Answer

"The write path assigns a monotonic seq per channel and persists before pushing. The client SDK keeps last_seq per channel. On push: if seq = last+1 apply; if greater, gap → sync since last; if less or equal, drop as duplicate. On connect: subscribe first, buffer live events, call sync with all cursors in one batched request, apply, then drain the buffer deduped by seq. That closes the race where an event lands between sync and subscribe. Retention for sync is 30 days; beyond that the client does a full reload. Result: effectively-once, in-order per channel, with no per-connection state on the gateway."

Why this is L6:

  • Covers gap, duplicate and race cases explicitly
  • Puts the guarantee in the durable path, not the socket

What L7 adds:

  • Ships the logic in a single client SDK so no product team re-implements it
  • Defines a measurable SLO: gap_detected − gap_filled = 0 and 99.9% of events visible < 2s

Drill 4: The Deploy#

Prompt: "You need to deploy a new gateway version to 65 hosts holding 10M connections. Go."

Staff Answer

"Canary one cell (~5 hosts) for 24h — memory leaks in connection servers show up over hours, not minutes. Then waves of ~2% of hosts. For each host: remove from LB for new connections, send RECONNECT{after_ms} with a uniform random 0–120s delay to each client, wait for connections to drain to near zero (hard kill after 5 minutes), restart, re-add. 200K connections over 120s is ~1.7K/s arriving across the remaining fleet — trivial. Clients resume with a resume token so the auth service sees nothing. Total: ~50 waves, a few hours. Auto-halt if connections_active fleet-wide drops more than 3% unexpectedly or push_to_visible_ms_p99 exceeds 1s."

Why this is L6:

  • Treats a deploy as a drain with spread, not a restart
  • Includes canary duration matched to the failure mode (leaks)
  • Has an automatic halt condition

What L7 adds:

  • Makes progressive, cell-by-cell deploys enforced by the deploy tooling for every stateful fleet in the org
  • Adds randomized connection lifetimes so deploys and routine churn look the same

Drill 5: The Hot Channel#

Prompt: "A live sports channel has 3M concurrent subscribers and 5K comments/s."

Staff Answer

"5K/s × 3M = 15B socket writes/s — impossible and pointless, nobody reads 5K comments/s. So I change the promise: the score and game state are latest-value, pushed at most every 500ms. Comments are sampled — clients get a stream of ~20/s, weighted toward friends and highly-rated comments — batched into one frame per 500ms. Architecturally, the channel is promoted to broadcast class: off the shared channel owner onto a dedicated relay per region, which sends one batched frame per gateway per interval. 3M subscribers × 2 frames/s = 6M writes/s fleet-wide, spread over all hosts — manageable. The full comment firehose is still durably stored for replay."

Why this is L6:

  • Changes the product contract rather than scaling a bad promise
  • Quantifies both the naive and the redesigned fan-out
  • Isolates the hot channel so it can't hurt other channels

What L7 adds:

  • Makes the sampling policy an explicit product decision with an owner
  • Sets org-wide thresholds for automatic promotion so this isn't an incident each time

Drill 6: Reconnect Storm Math#

Prompt: "A network blip disconnects 4M clients at once. What happens?"

Staff Answer

"Clients start full-jitter backoff: first attempt uniform in 0–1s, second 0–2s, and so on, capped at 60s. Even with jitter, ~2M attempts land in the first second. Admission control on each host admits ~2K/s — with 65 hosts that's 130K/s admitted, the rest get a fast 503 with a randomized Retry-After. So the fleet reconnects 4M in ~30–40s without any host falling over. Handshakes use session tickets; resume tokens avoid auth calls; each client makes one batched sync call — 4M sync calls over ~40s is ~100K/s, which the sync API must be sized for, or it rate-limits per client and serves boot payloads from an edge cache."

Why this is L6:

  • Quantifies how jitter and admission interact
  • Identifies the second system in the blast path (sync/boot)

What L7 adds:

  • Runs a quarterly game day that disconnects a cell to verify the math
  • Coordinates with auth and sync owners so every system on the reconnect path has a storm budget

Drill 7: Multi-Tenant Platform#

Prompt: "Five product teams want to push events. How do you let them?"

Staff Answer

"One multiplexed socket per client, one platform. Teams register topics: owner, delivery class (ephemeral, durable, broadcast), max publish rate, payload cap (16 KB — anything bigger is a reference), and schema. The router enforces per-topic quotas and can circuit-break a topic that exceeds it without affecting others. Durable topics must use the platform's seq/sync — teams don't build their own. The platform team owns the fleet and is paged for fleet health; topic owners are paged for their topic's publish errors and quota breaches."

Why this is L6:

  • Avoids N sockets per client and N fleets
  • Puts guardrails (quotas, payload caps) at the platform boundary
  • Splits paging responsibility

What L7 adds:

  • Decides what's out of scope for the platform (e.g., collaborative session state)
  • Sets chargeback by connections-minutes and deliveries so teams see their cost

Drill 8: Build vs Buy#

Prompt: "Why not use a managed real-time service?"

Staff Answer

"Below ~1M concurrent connections, or for a product that isn't real-time-centric, I would — managed services price per connection-minute and message, and they handle storms for you. Three things push toward building: cost at 10M+ concurrent (per-connection pricing scales linearly while self-hosted cost per connection falls with density), control over the reconnect and drain behavior, and integration with our own sequenced log. Middle path: use a managed service for the socket layer but keep the durable log and sync API ours — then the vendor is a replaceable transport, not our system of record."

Why this is L6:

  • Gives a threshold and the cost curve argument
  • Keeps the guarantee in-house regardless of vendor

What L7 adds:

  • Designs the exit: the client SDK abstracts transport so switching vendors is a client flag
  • Uses the Build vs Buy framework with a 3-year TCO

Drill 9: Multi-Region#

Prompt: "Users in Europe complain of 400ms push latency. Go multi-region."

Staff Answer

"Put gateways in each region and let clients connect to the nearest via GeoDNS/anycast. The hard part is routing: a channel with members in three regions. The channel owner (and its seq assignment) lives in one home region; events replicate to regional routers, which fan out locally. Cross-region latency is then paid once per event, not once per subscriber. Sync reads from the regional replica — if its seq is behind the client's cursor, it waits briefly or proxies to home. Region failure: clients reconnect to the next region with jitter; sync from any replica covers gaps."

Why this is L6:

  • Pays cross-region cost once per event, not per subscriber
  • Keeps per-channel ordering with a home region

What L7 adds:

  • Decides home-region placement policy (by creator, by majority membership) and its data residency implications
  • Budgets cross-region replication cost as a line item

Drill 10: Protocol Change Without an Outage#

Prompt: "We want to switch from JSON frames to compressed binary frames."

Staff Answer

"Negotiate at connect: the client advertises supported encodings in the handshake, the gateway picks the best mutual one. Deploy gateway support first (both encodings), then roll the client SDK. Measure bandwidth and CPU per host — compression trades gateway CPU for bandwidth, and for mobile clients bandwidth and battery usually win. Keep JSON supported until old client versions fall below 1% of connections, then deprecate with a server-forced upgrade prompt."

Why this is L6:

  • Uses negotiation to make the change a two-way door
  • Measures the CPU/bandwidth tradeoff before committing

What L7 adds:

  • Defines a client version support policy (e.g., 12 months) across all real-time features
  • Ties the change to a cost outcome (egress $ saved per month)

8. Deep Dive Scenarios#

Deep Dive 1: The Deploy That Took Down Everything#

Context: A config change to the gateway's TLS settings was pushed to all hosts at once. Every host restarted within 90 seconds. 10M clients reconnected; the auth service collapsed; the outage lasted 38 minutes. You're asked to lead the post-incident work.

Questions to Surface First:

  • Why could a config change reach 100% of hosts at once?
  • Did clients back off with jitter, or reconnect in lockstep?
  • Why did reconnection hit the auth service at all?
  • Which systems were on the reconnect path, and what were their storm budgets?

Typical L5 Approach: Adds more auth service capacity and makes clients back off longer. Both help; neither addresses why a single push could restart the whole fleet.

Staff Approach: Fixes each layer on the reconnect path: config changes go through the same cell-by-cell progressive rollout as code; resume tokens remove auth from reconnect; handshake admission control per host; SDK full-jitter backoff verified in tests; boot/sync payload cached at the edge.

Principal Approach: Treats it as an org failure-posture problem: config and code must share one progressive delivery system for every stateful fleet, enforced by tooling. Introduces cells with hard isolation so no single change can touch more than one cell per hour, and schedules a quarterly mass-reconnect game day with auth and sync owners in the room.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Block new connections at LB; admit by region/cell in steps
TriageConfirm auth saturation from reconnect validations; confirm lockstep reconnect timing
Quick fixEnable admission control; raise Retry-After randomization; admit 10% of traffic per 2 minutes
GuardrailsConfig pushes require the progressive rollout pipeline
Post-mortemConfig bypassed the deploy pipeline; auth on the reconnect path; no storm budgets

Metrics to Watch: gateway.new_connections_per_sec, auth.validations_per_sec, gateway.admission_rejected_total, sync.rps, connections_active fleet-wide.

Organizational Follow-up: Every system on the reconnect path (LB, gateway, auth, sync, boot) publishes a storm budget; game day verifies it.

Ownership Question: "Who is allowed to push a config change to all gateways at once?" Staff answer: Nobody, by construction. The pipeline enforces cell-by-cell rollout; the emergency override requires the gateway on-call plus an incident commander and still rolls at most 25% per step.

Key Takeaway: "In a stateful fleet, the blast radius of a change is measured in reconnects. Design the rollout, not just the recovery."

What clears the Staff bar:

  • Maps every system on the reconnect path
  • Removes auth from reconnect rather than scaling it
  • Puts config under the same rollout discipline as code

Deep Dive 2: Messages Silently Missing#

Context: Support tickets say users "sometimes miss messages." Metrics show push success at 99.99%. No errors anywhere.

Questions to Surface First:

  • What does "push success" measure — a write into the kernel send buffer?
  • Do clients report gaps? Do they sync on every reconnect, including resume from background?
  • Is it concentrated by client version, platform or network?

Typical L5 Approach: Adds retries on push and increases heartbeat frequency. Push "success" was never the problem.

Staff Approach: Instruments the contract, not the transport: client.gap_detected_total, client.gap_filled_total, and sync_per_reconnect by client version. Finds an iOS build skipping sync on foreground resume. Fixes in SDK, forces sync server-side for affected versions.

Principal Approach: Establishes an end-to-end delivery SLO measured from the client (e.g., 99.9% of events visible within 2s, 100% after sync) and makes it the platform's headline metric; server-side "push success" is demoted to a debugging metric. Funds a synthetic client fleet that continuously verifies the contract.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateBreak down tickets by client version/platform
TriageCompare sync_per_reconnect across versions
Quick fixServer sends SYNC_REQUIRED after resume for affected versions
GuardrailsSDK contract tests: kill socket mid-stream, assert no gaps
Post-mortemMetrics measured the transport, not the user-visible outcome

Metrics to Watch: client.gap_detected_total − client.gap_filled_total, sync_per_reconnect{client_version}, synthetic event_visible_ms p99.

Organizational Follow-up: Client platform team owns the SDK; app teams can't bypass it.

Ownership Question: "Who owns 'user sees every message'?" Staff answer: The real-time platform team owns the SLO end-to-end, including the client SDK — because the guarantee is implemented half on the client.

Key Takeaway: "A 99.99% push success rate says nothing about whether users saw their messages. Measure at the client."

What clears the Staff bar:

  • Distrusts transport-level success metrics
  • Locates the guarantee in client logic and owns it

Deep Dive 3: Onboarding a Massive Tenant#

Context: A large enterprise customer is moving 400K employees onto the product. They have a single all-company channel with 400K members, and every morning at 9am local time ~300K people log in within 20 minutes.

Questions to Surface First:

  • How many of the 400K are online simultaneously, and how many events/s does the all-company channel get?
  • What happens at 9am — login wave or thundering herd?
  • Are they concentrated in one region/cell?

Typical L5 Approach: Adds gateway capacity for 400K more connections.

Staff Approach: Recognizes two separate problems. The morning wave: 300K connects in 20 min is only ~250/s — fine — but boot payload for a 400K-member workspace is huge; serve it from an edge cache and lazy-load member lists. The all-company channel: pre-promote to broadcast class, restrict posting to admins or rate-limit it, and lazy-load member presence rather than fanning it out.

Principal Approach: Negotiates product limits with the customer (announcement-only channels above 50K members), places the tenant across multiple cells to avoid concentrating risk, and prices the dedicated capacity into the contract.

Staff Approach — Full Reasoning
DimensionStaff Answer
Connections+400K spread across cells; no single cell > 25% of the tenant
Boot payloadEdge-cached, delta-based; member list paginated
Big channelBroadcast class, posting rate-limited, reactions coalesced
PresenceNot fanned out for channels > 10K; fetched on demand

Metrics to Watch: boot_payload_bytes_p99{tenant}, channel_events_per_sec{channel}, cell_connections{tenant}.

Organizational Follow-up: Enterprise onboarding checklist includes a real-time capacity review.

Ownership Question: "Who approves a 400K-member channel with open posting?" Staff answer: Product, with a platform-set hard limit; above it, the channel is announcement-only by default.

Key Takeaway: "Large tenants don't break the fleet with connections. They break it with one giant channel and one giant boot payload."

What clears the Staff bar:

  • Separates connection load from channel fan-out load
  • Uses product limits as an engineering tool

Deep Dive 4: Post-Mortem — Hot Channel Took Down a Shard#

Context: A viral live event channel hashed to the same channel-owner server as 2% of all channels. For 12 minutes those channels had 10–30s delays. Write the post-mortem actions.

Questions to Surface First:

  • Why wasn't the channel promoted automatically?
  • Why did one channel's load affect unrelated channels?
  • How long did detection take?

Typical L5 Approach: Adds more owner servers so each owns fewer channels.

Staff Approach: Automatic promotion at thresholds (> 10K online or > 100 events/s) with migration to a dedicated broadcast pool; per-channel CPU accounting on owners; alert on push_to_visible_ms_p99 for a sample of unrelated channels (collateral damage detector).

Principal Approach: Defines "noisy neighbor isolation" as a platform requirement across all shared tiers (router, gateway, sync), with per-tenant/per-channel budgets; reviews every shared tier for the same failure class.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateManually promote channel; rehash its owner
TriageTop-K channels by owner CPU
Quick fixCoalescing on the hot channel
GuardrailsAutomatic promotion thresholds
Post-mortemNo isolation between channels on an owner

Metrics to Watch: router.owner_cpu{owner}, router.channel_events_per_sec top-K, collateral-latency probe.

Organizational Follow-up: Product signs off on sampling policy for broadcast channels.

Ownership Question: "Who decides that users in huge channels see a sample of messages?" Staff answer: Product, in advance — the platform implements the mechanism and the thresholds.

Key Takeaway: "Consistent hashing distributes channels, not load. The tail needs its own lane."

What clears the Staff bar:

  • Detects collateral damage, not just the hot spot
  • Automates the promotion decision

Deep Dive 5: Multi-Region Expansion#

Context: The company is launching in Asia-Pacific. Leadership wants < 200ms push latency there and regional failover.

Questions to Surface First:

  • What fraction of channels span regions?
  • Where does per-channel ordering (seq assignment) live?
  • Data residency requirements for message content?

Typical L5 Approach: Deploys the full stack in APAC and routes by geography.

Staff Approach: Regional gateways and routers; per-channel home region assigns seq; events replicate to other regions once per event; sync from regional replicas with seq-awareness; regional failover by jittered reconnect to the next region.

Principal Approach: Decides the home-region policy and its residency implications with legal; budgets cross-region egress; sequences the rollout as read-only regional delivery first, then regional writes, with a failover game day before GA.

Staff Approach — Full Reasoning
DimensionStaff Answer
LatencyRegional gateways; cross-region paid once per event
OrderingHome-region sequencer per channel
FailoverClients reconnect to next region with jitter; sync covers gaps
CostEgress ≈ events × avg remote regions × payload size

Metrics to Watch: cross_region_replication_lag_ms, push_to_visible_ms_p99{region}, sync_proxy_to_home_total.

Organizational Follow-up: Regional on-call coverage; residency review.

Ownership Question: "Who decides a channel's home region?" Staff answer: The platform, by an automatic policy (creator's region, rebalanced by majority membership after 30 days), reviewed by legal for residency-bound tenants.

Key Takeaway: "Replicate events between regions, not deliveries."

What clears the Staff bar:

  • Keeps ordering with a clear owner
  • Treats failover as a reconnect storm to design for

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why the socket is a cache of a durable log and design seq + sync with gap/duplicate/race handling
  • Size a gateway fleet by connections, memory and reconnect handshakes, and choose host size for blast radius
  • Design subscription-routed fan-out and promote hot channels to a broadcast class
  • Design against reconnect storms: jitter, drain spreading, admission control, resume tokens, randomized lifetimes
  • Run deploys as drains, cell by cell
  • Define platform vs product ownership with a topic registry and quotas
  • Extend the design to multi-region with a home-region sequencer

The Bar for This Question#

Mid-level (L4): WebSocket server, a pub/sub bus, clients reconnect. Works for a demo.

Senior (L5): Adds a connection registry, horizontal scaling, heartbeats, maybe SSE vs WebSocket analysis. Correct at moderate scale; delivery assumed reliable; reconnect storms and hot channels not considered.

Staff+ (L6): Makes the durable log the source of truth with seq and sync; thin gateways that can be killed; fan-out routed by subscription with a hot-channel lane; reconnect designed as the peak load; deploys as drains; ownership split between platform and topic owners. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "Reliable WebSocket Delivery Is the Wrong Goal"#

EvidenceImplication
Half-open TCP connections accept writes that never arriveServer-side "success" is fiction
Per-connection queues die with the hostReliability requires durable state elsewhere anyway

The Staff position: Make the socket best-effort and the sync path authoritative.

Why this matters in interviews: It collapses a sprawling ack/retry design into one clean invariant.

10.2 "Your Peak Load Is a Deploy, Not a Traffic Spike"#

EvidenceImplication
Idle connections cost memory, not CPUSteady-state CPU is low
Every restart forces reconnects with full handshakes and boot payloadsDeploys and crashes set the CPU headroom

The Staff position: Size for reconnect handshakes/s and design deploys as spread drains.

Why this matters in interviews: It shows you've operated a stateful fleet.

10.3 "Most Apps Should Poll"#

EvidenceImplication
60s polling of a cacheable endpoint is stateless and CDN-friendlyNo fleet, no storms
Many "real-time" features tolerate 30–60s stalenessThe stateful fleet is overkill

The Staff position: Ask what staleness product will accept before building sockets.

Why this matters in interviews: Knowing when not to build is a Staff-level signal.

10.4 "One Million Connections Per Host Is a Vanity Metric"#

EvidenceImplication
Density records are achievable (WhatsApp ~2M in 2012)Memory is not the constraint
Losing a 1M-connection host is a 1M-client stormBlast radius is the constraint

The Staff position: Cap at ~100–200K per host and spend the extra machines on blast radius.

Why this matters in interviews: It reframes capacity as risk management.

10.5 "Presence Is More Expensive Than Messaging"#

EvidenceImplication
Every connect/disconnect/idle change can fan out to every contactPresence events often exceed message events 10×
Users rarely act on precise presenceCoarse, lazy presence loses little value

The Staff position: Presence is lazy, coarse, and scoped to what's on screen.

Why this matters in interviews: It shows you can identify the hidden amplifier.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

At Staff level, the real-time system is a fleet to design. At Principal level, it's shared infrastructure that every product team will depend on — and therefore the org's single largest correlated-failure risk for user-facing freshness. The L7 questions are: what contract does the platform offer, what does it refuse to carry, how is its blast radius partitioned, and how is its cost attributed so that teams don't treat live updates as free.

The Org-Level Fault Line#

One real-time platform vs per-product real-time stacks. One platform means one socket per client (battery, connection cost), one team that has mastered storms, and one contract. It also means one set of cells whose failure is felt by every product at once and a platform backlog that gates every new real-time feature. The L7 position: centralize the socket, the delivery contract, the SDK and the storm engineering; decentralize topics, payloads and product semantics via self-serve registry and quotas; keep collaborative session servers as a separate, affinity-based tier owned closer to the product.

Cost Model#

Assumptions: ~$0.10/host-hour for a 16-vCPU, 32 GB instance class; 200K connections/host; egress at ~$0.02–0.05/GB blended; engineers at $250K fully loaded.

ScaleConcurrentCompute / monthEgress / monthHeadcountOn-call
Small100K~$1–2K (or a managed service at ~$2–5K)~$1K1–2 part-time; buyBusiness hours
Medium10M~$6–10K for ~65 gateways + routers/sync ~$15K~$20–60K5–8 (gateway, router, SDK, sync)24/7; storms dominate pages
Large100M~$150–250K across cells and regions~$300K+15–25 across platform, SDK, regional opsFollow-the-sun, cell-scoped paging

The surprising line item is egress, not compute: frame compression and coalescing are cost projects, not just performance projects.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversibility Cost
Per-channel seq + sync as the delivery contractOne-way (good)Clients in the wild depend on it for years
Client SDK owning reconnect/sync logicOne-wayOld app versions persist 12–24 months
Wire protocol (framing, encoding)Two-way with negotiationCheap if negotiated at connect; expensive if hardcoded
Transport vendor (managed vs self-hosted)Two-way if SDK abstracts itWeeks behind an abstraction; quarters without
Home-region sequencing modelMostly one-wayChanging ordering semantics breaks client assumptions
Host size / cell sizeTwo-wayRebalance via randomized lifetimes

The Standard I'd Write#

RFC: Real-Time Delivery Platform Standard (v1)

Scope: Any feature that pushes server events to clients.

MUST:

  • Use the platform gateway and client SDK; no product-owned sockets.
  • Register every topic with owner, delivery class, max publish rate and payload cap (≤ 16 KB).
  • Durable topics MUST assign a per-channel seq on write and expose sync; the socket MUST NOT be the system of record.
  • Clients MUST implement full-jitter backoff (via SDK) and honor server RECONNECT delays.
  • Fleet changes (code or config) MUST roll out cell by cell.

SHOULD:

  • Broadcast-class topics declare a coalescing interval and sampling policy.
  • Ephemeral events (typing, presence) be droppable under backpressure.

Exceptions: Collaborative session servers (affinity tier) — reviewed by the platform team; time-boxed.

Success metrics: 99.9% of events visible < 2s at the client; zero unfilled gaps; no single change affecting > 1 cell per hour; one socket per client.

What I'd Tell the VP#

Real-time updates are now part of every product we ship, and they all run through connections that have to be re-established whenever anything changes on our side. Today, a single bad deploy can disconnect every user at once, and our biggest outages have been self-inflicted in exactly that way. I'm proposing we consolidate onto one real-time platform, split it into independent cells so any failure touches a few percent of users, and guarantee — measured on users' devices — that nobody misses a message. It's about six engineers for three quarters, and it removes four duplicate systems and our most common cause of company-wide incidents.

Principal Interview Signals#

SignalWhat It Sounds Like
Sees the platform"Five teams will publish through this; I'll give them a registry and a contract, not a socket."
Prices it"At 10M concurrent, egress is the biggest line item — compression is a cost project."
Designs the org's failure posture"Cells of 200K users; no change touches more than one cell per hour."
Knows what not to centralize"Collaborative session state stays in its own affinity tier."
Plans for long-lived clients"The SDK's sync logic is a one-way door — old clients live for two years."

Staff answers that L7 interviewers find insufficient:

  • "Jittered backoff and admission control" — correct, but doesn't address the org-wide correlated risk of one fleet for all products.
  • "Seq + sync" — correct, but doesn't make it a published contract with a client-measured SLO.
  • "Platform owns the gateway" — correct, but no chargeback, quotas or exception process.

🧭 Principal Move: "The technical design is a solved pattern. What I'd spend my time on is the contract: what the platform guarantees, what it refuses to carry, and how its blast radius is partitioned — because every product team's freshness now depends on one fleet."


Appendices

Appendix A: Gateway Mechanics#

A.1 Connection Lifecycle#

Diagram: A.1 Connection Lifecycle

A.2 Gateway Event Loop (pseudocode)#

on_connect(req):
    if not admission.try_take(): return reject(503, retry_after=rand(5,30))
    user = verify_resume_token(req.token) or auth_service.verify(req.jwt)
    conn = Conn(user, device, send_buf_cap=256KB, max_life=rand(2h,4h))
    for ch in membership.channels_for(user):     # cached, batched
        local_index[ch].add(conn)
    router.register_interest_batch(local_index.new_channels())   # flushed every 1s
    conn.send(HELLO{heartbeat_ms=25000})

on_router_event(ch, frame):
    for conn in local_index[ch]:
        if not conn.try_write(frame):             # non-blocking
            if frame.class == EPHEMERAL: drop
            else: conn.close(reason=SLOW)         # client will sync on reconnect

on_tick():
    close conns with 2 missed heartbeats
    for conn past max_life: conn.send(RECONNECT{after_ms=rand(0,30000)})

Appendix B: Data Model and Sequencing#

StoreKeyNotes
Event store(channel_id, seq) clusteredWide-column (e.g., Cassandra); retention 30 days for sync; older via archive
Seq allocatorper channel_idChannel owner in memory with durable high-water mark, or a DB counter; one writer per channel
Membershipuser_id → channels, channel_id → membersSource for subscriptions; cached on gateways
Interest mapchannel_id → gatewaysSoft state, rebuilt by gateway re-registration; TTL 60s
Resume tokenHMAC(user, device, exp)10-minute TTL; verified locally, no auth call

Why per-channel and not global order: global ordering needs a single sequencer — a throughput ceiling and a SPOF. Users only perceive order within a conversation.

Appendix C: Fan-Out Mechanisms — Quick Comparison#

MechanismThroughputHot-Channel BehaviorFailure BehaviorUse For
Redis pub/sub broadcastLimited by fleet × eventsEveryone gets everythingFire-and-forget; loss on disconnect< 20 gateways
Redis pub/sub per channel topicGoodHot topic pins one Redis nodeSubscribe storms on reconnectMedium scale
Kafka + channel-owner serversHigh, replayableOwner hot-spot; needs promotionOwner failover rebuilds from interest mapLarge scale
Relay tier (regional)Very high for broadcastDesigned for itRelay loss → gateways fall back to ownerBroadcast class

Appendix D: Client Contract#

  • Connect with resume token if present; else JWT.
  • Full-jitter backoff: delay = random(0, min(60s, 1s × 2^attempt)); honor Retry-After and RECONNECT{after_ms}.
  • Subscribe → buffer → sync → drain buffer (dedupe by seq).
  • On gap (seq > last + 1): sync that channel; on seq ≤ last: drop.
  • While disconnected > 10s: poll sync every 15s.
  • On app background: close socket; rely on OS push hints; sync on foreground.

Appendix E: Observability#

E.1 Core Metrics#

gateway.connections_active{host,cell}      gateway.new_connections_per_sec
gateway.admission_rejected_total           gateway.heartbeat_timeouts_total{asn}
gateway.conn_send_buffer_bytes (p99)       gateway.event_loop_lag_ms
router.owner_cpu{owner}                    router.channel_events_per_sec (top-K)
push_to_visible_ms (p50/p99, client-measured)
client.gap_detected_total                  client.gap_filled_total
sync.rps / sync.p99                        sync_per_reconnect{client_version}

E.2 Critical Alerts#

AlertThresholdAction
Mass disconnectconnections_active fleet drop > 3% in 1 minPage; halt deploys
Admission saturationrejected > 20% of attempts for 2 minPage; add capacity, check storm source
Unfilled gapsgap_detected − gap_filled > 0.01% of eventsPage client platform
Collateral latencypush_to_visible_ms_p99 > 1s on probe channelsPage; look for hot channel
Sync overloadsync.p99 > 1sPage sync owner; enable edge cache/rate-limit

E.3 Debugging the Silent Failure#

Server metrics can't see missed deliveries. Run a synthetic client fleet (~1K clients across regions and networks) that publishes and subscribes to probe channels, measures publish-to-visible latency, and verifies no gaps — including through forced disconnects.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 100K concurrentOne or two socket servers, Redis pub/sub, or a managed serviceDeploys start disconnecting users noticeably
100K–5MThin gateways, seq + sync, drain-based deploysBroadcast fan-out cost; hot channels
5M–50MChannel owners, hot-channel promotion, cellsMulti-product governance, multi-region
50M+Multi-region cellular platform, topic registry, chargebackOrg coordination; long-lived client versions

What You Don't Build on Day One#

  • Relay tiers (until a channel exceeds ~10K online members)
  • Multi-region sequencing
  • Per-connection durable queues (ever, usually)
  • Custom binary protocol (negotiate later)

Appendix G: Multi-Tenancy and Cost#

  • Chargeback units: connection-minutes (gateway memory) and deliveries (CPU + egress) per topic owner.
  • Quotas: per-topic publish rate; per-tenant maximum channel size for open posting; payload cap 16 KB.
  • Isolation: broadcast-class channels on their own owner pool; large tenants spread across cells.
  • Tradeoff summary: thinner gateways cost a hop (~1–5ms) and buy freedom to kill hosts; smaller hosts cost machines and buy smaller storms; best-effort push costs a sync API and buys stateless gateways. Every one of those trades is worth it above ~1M concurrent.
  1. Loading the index…