Technologies referenced in this case study: Redis · Apache Kafka · Cassandra · API Gateways · ZooKeeper & etcd
Start with the pattern, then come here for the system: Real-time Updates — Cross-Cutting Pattern covers SSE vs WebSocket, the connection registry and presence basics. This case study goes deeper on what the pattern page doesn't: the connection gateway fleet, fan-out at scale, reconnect storms, and delivery guarantees.
Related case studies: Chat Messaging · Notification System · Collaborative Editing · Load Balancer · Message Queues · Degraded Mode
How to Use This Case Study#
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 4, 5 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 → Section 4 (reconnect storm, hot channel) → Deep Dives 1–2 |
| Deep Dive | 3+ hrs | Everything, including the Principal Lens and appendices |
What is a Real-Time Update System? — Why interviewers pick this topic
A real-time update system keeps millions of long-lived client connections open and pushes events — new messages, likes, typing indicators, price ticks, order status — to the right subset of them within a few hundred milliseconds. Slack, Discord, Facebook, Uber's rider app, trading dashboards and every "live" feature run one.
Unlike request/response systems, the expensive resource isn't requests — it's connections. A connection is state that lives on one specific host, for hours, and every one of them must be re-established when that host goes away.
Before vs After — gateway deploy scenario (10M concurrent connections):
Without a Staff-level design:
t=0: Routine deploy restarts 20% of gateway hosts at once.
t=+1s: 2M clients see the socket close. All reconnect immediately.
t=+2s: Load balancer receives 2M TLS handshakes. Remaining hosts' CPU hits 100%.
t=+5s: Handshakes time out. Clients retry with a fixed 1s delay.
t=+10s: Healthy hosts fail health checks under handshake load. More clients disconnect.
t=+30s: 8M clients in a reconnect loop. Auth service sees 400K token validations/s.
t=+4min: Auth service falls over. Full outage — self-inflicted by a deploy.
t=+45min: Recovered by blocking traffic at the LB and admitting it region by region.
With a Staff-level design:
t=0: Deploy drains 2% of hosts per wave, 10 minutes per wave.
t=+0s: Draining host sends 'reconnect' frame with a random 0–120s delay per client.
t=+2min: 200K clients reconnected over 120s = ~1.7K/s. p99 handshake 40ms.
t=+2min: Each client resumes from its last sequence number. Zero events lost.
t=+5h: Deploy complete. Nobody noticed. Nobody was paged.
Why interviewers reach for this question: It tests whether you understand stateful infrastructure — capacity measured in connections not QPS, failure measured in reconnect storms not error rates, and correctness measured in "did the user eventually see everything" not "did this push succeed."
Mechanics Refresher: Transport Options
| Transport | How It Works | Pros | Cons |
|---|---|---|---|
| Short polling | Client asks every N seconds | Trivial, stateless | Latency = N/2; wasteful at scale (10M clients / 5s = 2M req/s of mostly empty responses) |
| Long polling | Server holds request until an event or ~30s timeout | Works everywhere, HTTP semantics | Reconnect per message; ordering and dedupe are your problem |
| Server-Sent Events (SSE) | One-way HTTP stream, text/event-stream, built-in Last-Event-ID resume | Simple, HTTP/2 multiplexed, auto-reconnect in browsers | Server→client only; some proxies buffer |
| WebSocket | Full-duplex framed TCP after HTTP upgrade | Bidirectional, low overhead per message (2–14 byte frames) | Stateful, no built-in resume, proxy/LB idle timeouts |
| Mobile push (APNs/FCM) | OS-level push channel | Works when app is backgrounded | Seconds of latency, no delivery guarantee, payload limits (~4 KB) |
For most production systems: WebSocket (or SSE if the client never sends) for foreground apps, OS push for backgrounded mobile, and a durable sync API underneath all of them. The transport is the easy decision; see Real-time Updates for the comparison. This case study is about what happens behind it.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What This Interview Actually Tests#
Real-time updates is not a WebSocket question. Everyone knows how to open a socket.
It is a stateful-fleet and delivery-contract question that tests:
- Whether you size the system in connections, memory and handshakes — not requests per second
- Whether you know where each event's source of truth lives (hint: not the socket)
- Whether you can design fan-out that survives one channel with 1M subscribers
- Whether you treat reconnection — after deploys, crashes, network blips — as the dominant load event
The key insight: The WebSocket is a cache of a durable log, not the delivery mechanism of record. Push is a latency optimization; the guarantee comes from a per-user or per-channel sequence number and a sync API the client calls on every reconnect. Design it that way and the gateway can be lossy, restartable and boring — which is exactly what you want from a fleet holding 10M connections.
The L5 → L6 → L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "WebSocket servers behind a load balancer, Redis pub/sub between them" | Asks "what's the delivery contract — can the user miss an event, and what's the source of truth?" | Asks "how many product teams will push through this, and are we building a platform or a feature?" |
| Capacity | QPS-based sizing | Sizes in connections (~100–200K/host), memory (~20–50 KB/conn), and reconnect handshakes/s | Prices connections: $/million-concurrent/month, and the cost of blast radius per host size |
| Fan-out | Publish to every server; each filters | Subscription-aware routing: servers subscribe only to channels they have members for; hot channels get tiered fan-out | Sets org limits: max channel size, per-team publish quotas, a topic registry |
| Delivery | "WebSockets are reliable, TCP guarantees delivery" | At-most-once push + sequence numbers + sync-on-reconnect = no gaps | Publishes the delivery contract as an SLO ("99.9% of events visible within 2s, 100% within sync") that product teams build against |
| Failure | "Clients reconnect automatically" | Designs against reconnect storms: jittered backoff, drain with spread, admission control, resumable sessions | Designs the org's failure posture: cells so a bad deploy hits 2% of users, not 100%; game days for mass reconnect |
| Ownership | Each team runs its own socket server | Platform owns the gateway; product teams own topics and payload schemas | Writes the standard; decides what never goes over the socket (large payloads, source of truth) |
Why "delivery" separates levels
L5: "WebSocket runs over TCP, so delivery is reliable." TCP guarantees in-order bytes on one connection while it lives. Messages sent to a socket that's half-dead (mobile client entered a tunnel; NAT dropped the mapping) are acknowledged by the kernel's send buffer and silently lost when the connection eventually closes. There is no end-to-end ack.
L6: "I'll treat the push as best-effort. Every event gets a monotonic sequence number per user (or per channel). The client stores the last sequence it saw. On reconnect, it calls sync(since=seq) against the durable store. Push gives us 200ms latency; sync gives us the guarantee."
L7: "And I'd make that contract the platform's API. Product teams shouldn't each invent their own sequence scheme — the gateway team publishes the guarantee, and a product that needs stronger semantics (payments status, say) is told explicitly that real-time is the wrong channel."
Why "failure" separates levels
L5: Trusts automatic reconnection. Automatic reconnection is the failure: 2M clients reconnecting within the same second is a DDoS against your own load balancers, TLS terminators and auth service.
L6: Designs the reconnect path as the peak load path: exponential backoff with full jitter on clients, server-directed reconnect delays during drains, token-bucket admission on handshakes per host, cheap session resumption (TLS session tickets, a short-lived resume token so auth isn't re-hit), and host sizing chosen partly for blast radius.
L7: Recognizes correlated failure across the org: one gateway fleet serving every product means one bad config push disconnects the entire company's user base. Introduces cells (e.g., 50 cells of 200K users each) and a progressive deploy policy that's enforced by tooling, not by discipline.
Why "capacity" separates levels
L5: "Each server handles 10K requests/s; we need N servers." Idle connections cost almost no CPU, so QPS-based sizing wildly underestimates memory and wildly misses the handshake peak.
L6: Sizes three things independently: steady connections per host (memory-bound), message throughput (CPU-bound, driven by fan-out), and reconnect handshake rate (CPU-bound, driven by failure). The last one sets the headroom.
L7: Chooses host size as a blast-radius decision: 1M connections per host is achievable and cheap, but losing that host means 1M simultaneous reconnects. Prices the extra hosts for 200K/host against the cost of a storm.
The Staff Positions#
| Position | Rationale |
|---|---|
| The socket is not the source of truth | Every event is written durably first; push is a latency optimization over a sync API |
| Sequence numbers + sync-on-reconnect over server-side per-connection queues | Connection-scoped buffers die with the host; a cursor into a durable log doesn't |
| Gateways are dumb and stateless beyond the socket | Subscription and routing state lives in a separate layer so any gateway can be killed |
| Size hosts for blast radius, not max density | 100–200K connections/host, not 1M — a lost host must be a tolerable reconnect wave |
| Reconnect is the peak load, design for it | Full-jitter backoff, server-directed drain delays, handshake admission control |
| Subscription-aware fan-out; special-case hot channels | Broadcasting every event to every gateway is O(events × gateways) and dies at ~100 hosts |
| WebSocket for foreground, OS push for background | Don't hold sockets for backgrounded mobile apps; the OS will kill them anyway |
The Four Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Activity notifications (likes, badges, feed refresh hints) | Huge scale, low value per event | Best-effort push, coalescing, drop under pressure | Missed hint → stale badge until next sync | Eventually visible; loss acceptable |
| Conversational delivery (chat, comments, DMs) | Ordered, no gaps, ~200ms | Sequence per conversation/user, durable log, sync on reconnect | Gaps or duplicates visible to users | Every event visible exactly once in order after sync |
| Collaborative sessions (docs, whiteboards, multiplayer) | Low-latency bidirectional, per-room state | Room affinity: route all members of a room to one session server | Session server loss = room state rebuild | Convergence (OT/CRDT), not just delivery |
| Live broadcast (scores, live-stream comments, auctions) | One publisher, 100K–10M subscribers | Tiered fan-out, sampling, edge relays | Hot channel melts one tier | Latest value wins; sampling acceptable |
🎯 Staff Move: "I'll design a general real-time delivery platform for conversational delivery — chat-like semantics where users must not miss events — because it has the strictest contract. Notifications are the same pipeline with dropping allowed. Collaborative sessions need room affinity, which I'll call out as a different routing mode. Live broadcast I'll handle as a hot-channel special case."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Stateful Gateway vs Stateless Gateway | Keep subscriptions and buffers on the host holding the socket (simple, fast) or externalize them (restartable, more hops)? |
| 2 | Broadcast vs Subscription-Routed Fan-Out | Send every event to every gateway (simple) or route by who's subscribed where (scales, but needs a routing layer)? |
| 3 | Push Guarantees vs Pull Sync | Make the socket reliable (acks, per-connection queues) or keep it best-effort and guarantee via cursor sync? |
| 4 | Reconnect Speed vs Fleet Protection | Let clients reconnect instantly (best UX for one) or throttle and spread (survives mass events)? |
| 5 | Platform Gateway vs Product-Owned Sockets | One shared fleet for every team (efficiency, one blast radius) or per-product sockets (autonomy, duplication)? |
In the Wild: Real Production Systems#
Why this section belongs here: Real-time fleets at scale converge on the same shape. Citing them shows you've seen the operational reality, not just the protocol.
Slack — Gateway Servers, Channel Servers, and an Edge Cache#
Slack has publicly described its real-time architecture: stateful Gateway Servers hold client WebSockets and subscribe on behalf of their clients; Channel Servers own channels via consistent hashing and fan messages out to subscribed gateways; separate services handle presence; and an edge cache ("Flannel") serves the boot/state payload clients need on connect, so reconnects don't hammer the primary databases.
Staff insight: The separation of "who holds the socket" (gateway) from "who owns the channel" (channel server) is the Fault Line 1 and 2 answer in production. And Flannel exists because reconnect is the expensive path — the boot payload, not the socket, is what melts backends.
Discord — Per-Guild Processes on the BEAM#
Discord runs its real-time gateway on Elixir/Erlang, with a process per guild (server) that fans out to member sessions. Their engineering blog has described scaling to millions of concurrent users and the special engineering needed for very large guilds — where a single guild's fan-out becomes the bottleneck — as well as work to compress gateway traffic to cut bandwidth.
Staff insight: Discord's big-guild problem is the hot channel fault line. The general design works until one channel has hundreds of thousands of members; then fan-out for that one channel needs its own architecture (relays, lazy member lists, sampled presence).
Netflix — Zuul Push#
Netflix open-sourced Zuul Push, a push notification service that holds persistent WebSocket/SSE connections for millions of devices. Their public talks describe the operational lessons: connections that live too long pin load to old hosts and make deploys dangerous, so they cap connection lifetime with randomization to continuously rebalance, and they emphasize that the herd of reconnecting clients is the main scaling risk.
Staff insight: Randomized connection lifetimes turn a rare, catastrophic mass reconnect into a constant, gentle trickle. That's a Staff move: convert a tail-risk event into routine background load you're always sized for.
(Also worth knowing: WhatsApp publicly reported handling around 2M concurrent connections on a single Erlang server in 2012. It proves density is achievable; it doesn't mean it's the right blast radius for you.)
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "WebSocket servers + Redis pub/sub" | "One channel has 2M subscribers. What happens?" | Hot-channel fan-out |
| "Clients reconnect automatically" | "You deploy and 2M clients reconnect in the same second." | Reconnect storm awareness |
| "TCP guarantees delivery" | "The phone went into a tunnel for 40s. Which messages did it miss, and how does it know?" | Delivery contract |
| "We store connection → server in Redis" | "That's 10M entries churning on every blip. What's the write rate during a storm?" | Registry cost under failure |
| "We'll scale to 1M connections per host" | "What happens when that host dies?" | Blast-radius sizing |
| "Each team can publish events" | "A team ships a bug that publishes 50K events/s to every user." | Platform guardrails |
System Architecture Overview#
Reading the diagram: Writes go to the durable path first — the Write API assigns a per-channel sequence number and persists before anything is pushed. The log feeds the subscription router, which knows which gateways have members of each channel and sends each event only there. Gateways are deliberately thin: they hold sockets, check resume tokens, and forward frames. On any reconnect, the client calls the Sync API with its last sequence number and fills gaps from the store. Backgrounded mobile clients are served by OS push from the same log.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Delivery | "TCP is reliable" | "Push is best-effort. Per-channel sequence numbers + sync on reconnect make it gap-free." |
| Fan-out | "Redis pub/sub to all servers" | "Subscription-routed: events go only to gateways with subscribers. Hot channels get a relay tier." |
| Capacity | "N servers at 10K QPS" | "~200K connections/host at ~30 KB each; headroom set by reconnect handshakes/s, not steady state." |
| Reconnect | "Auto-reconnect" | "Full-jitter backoff, server-directed drain delays, handshake admission control, resume tokens." |
| Deploys | "Rolling restart" | "Drain 2% per wave with spread reconnects; randomized max connection lifetime keeps load rebalanced." |
| Ownership | "Each service opens sockets" | "One gateway platform; product teams own topics and schemas under quotas." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Memory per idle WebSocket (tuned runtime) | ~10–50 KB | Sets connections per host: 200K × 30 KB ≈ 6 GB |
| Connections per host (practical) | 100K–500K; ~2M demonstrated (WhatsApp 2012) | Density is possible; blast radius is the constraint |
| Heartbeat interval | 20–30s | Must beat LB/NAT idle timeouts (commonly 60s; some mobile NATs shorter) |
| Full TLS handshake CPU (ECDSA) | ~1–2 ms CPU → ~1–3K/s per core | Reconnect storms are CPU-bound on handshakes |
| TLS resumption | ~5–10× cheaper than full handshake | Session tickets are storm insurance |
| Reconnect spread for a drained host | 60–120s | 200K conns / 120s ≈ 1.7K/s — absorbable |
| Target push latency | p99 < 500ms publish → on-screen | Beyond ~1s users perceive "not real-time" |
| WebSocket frame overhead | 2–14 bytes | Why small events are cheap on the wire |
| Mobile push payload limit | ~4 KB | OS push carries hints, not data |
| Registry churn during a 1M-conn storm | ~1M writes + deletes in minutes | The registry must survive the storm it's tracking |
| Fan-out cost of naive broadcast | events/s × gateways | 50K events/s × 200 hosts = 10M deliveries/s of mostly-dropped messages |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on WebSocket vs SSE vs long polling. The interviewer assumes you know. Pick one in 30 seconds and spend the time on the fleet, fan-out, reconnects and the delivery contract.
Phase 1: Requirements & Framing (2–3 minutes)#
Functional scope in one breath:
"Clients keep a live connection and receive events for the channels, conversations and objects they care about within about half a second. Clients can also send small events — typing, acks, presence — upstream."
Then the non-functionals that drive the design:
"The key question is the delivery contract. Can a user miss an event? For chat-like delivery, no — so I'll make the socket a best-effort accelerator over a durable, sequenced log, and guarantee completeness through sync on reconnect. Second, this is a stateful fleet: I'll size it in connections and reconnect handshakes, not QPS."
Commit to numbers:
"Assume 50M DAU, 10M peak concurrent connections, 1M published events/s, average fan-out 10 — so ~10M deliveries/s. Push latency p99 under 500ms. And one hot-channel case: a live event with 2M subscribers."
🎯 Staff Move: "The transport choice is the least interesting part of this problem. I'll use WebSockets for foreground clients and OS push for background, and spend our time on what breaks: fan-out, reconnects and the delivery guarantee."
Phase 2: Core Entities & API (1–2 minutes)#
- Connection (conn_id, user_id, device_id, gateway_host, connected_at, resume_token)
- Subscription (conn_id, channel_id) — derived from membership, not stored per connection durably
- Event (channel_id, seq, type, payload_ref or small payload, created_at)
- Cursor (client-side: last_seq per channel, or a per-user inbox seq)
API surface:
WS /connect?resume_token=… → HELLO { conn_id, heartbeat_ms: 25000 }
WS ← EVENT { channel, seq, type, body } (server → client)
WS → ACK { channel, seq } (optional, for metrics only)
WS ← RECONNECT { after_ms: 0–120000 } (server-directed drain)
GET /sync?cursors=ch1:1042,ch2:77&limit=500 → { events[], more: bool }
POST /channels/{id}/events { … } → 201 { seq } (write path, HTTP not socket)
🎯 Staff Move: "Writes go over HTTP to the write API, not over the socket. The socket is for delivery. That keeps the gateway free of business logic and lets me restart it without losing writes."
Phase 3: High-Level Architecture (≤5 minutes)#
Walk it in 60 seconds:
- A producer writes an event through the write API, which assigns the next per-channel
seqand persists it. - The event is appended to the log; the subscription router consumes it.
- The router looks up which gateways have subscribers for that channel and forwards only to them.
- Each gateway writes the frame to its local sockets for that channel.
- If a client was disconnected, on reconnect it calls
syncwith its cursors and gets anything it missed.
🎯 Staff Move: "That's the whole data path. It's correct on a quiet day. It breaks in three ways at scale: one channel with 2M subscribers, 2M clients reconnecting at once, and gaps nobody notices. Let me take those."
Phase 4: Transition to Depth (1 minute)#
"Three deep dives: fan-out routing including hot channels, the reconnect storm — which is the real peak load of this system — and the delivery contract with sequence numbers. I'd start with reconnects, because that's the failure that turns a routine deploy into a company-wide outage. Preference?"
Phase 5: Deep Dives (25–30 minutes)#
Deep dive 1: Reconnect storms (8–10 min)
"Steady state is easy — idle sockets cost memory, not CPU. The CPU peak is reconnection: a full TLS handshake costs ~1–2ms of CPU, and re-establishing session state costs an auth check plus a boot payload. If 2M clients reconnect in 2 seconds, that's 1M handshakes/s — orders of magnitude over what the fleet is sized for."
Five controls, in order of leverage:
| Control | What It Does | Number |
|---|---|---|
| Client full-jitter backoff | sleep(random(0, min(cap, base × 2^n))) | base 1s, cap 60s |
| Server-directed drain | Draining host sends RECONNECT{after_ms} with a random delay per client | spread over 60–120s |
| Handshake admission control | Per-host token bucket on new connections; excess gets 503 + Retry-After fast | e.g., 2K new conns/s/host |
| Cheap resumption | TLS session tickets + a signed short-lived resume token (no auth service call) | 5–10× cheaper |
| Randomized max connection lifetime | Every connection closes after 2–4h ± jitter | continuous rebalancing |
"The last one is what Netflix described for Zuul Push: if connections recycle constantly, a mass reconnect is just a faster version of normal, and new hosts pick up load without a thundering herd."
Deep dive 2: Fan-out (8–10 min)
"Naive design: every gateway subscribes to one Redis pub/sub bus and filters. That's O(events × gateways): 1M events/s × 50 hosts = 50M messages/s, 95% discarded. Instead, the subscription router tracks channel → set of gateways with at least one subscriber, maintained by gateways as clients join/leave. An event goes to ~k gateways where k is the number of hosts with members — for a 10-person chat, usually 1–10."
"Hot channels flip the math: a live event with 2M subscribers spread across all 50 hosts. Per-event cost is 50 router sends and 2M socket writes — fine for one event, fatal for 1K events/s (2B writes/s). So hot channels get: server-side coalescing (send the latest score every 500ms, not every change), sampling for comments ('show 20/s of 10K/s'), and a relay tier so the router sends once per region."
Deep dive 3: Delivery contract (6–8 min)
"Every event gets a monotonic seq per channel at write time. The client tracks last_seq per channel. Three cases: if a pushed event is last_seq + 1, apply it. If it's higher, there's a gap — call sync(since=last_seq). If it's lower or equal, it's a duplicate — drop it. On every reconnect, sync all channels the client cares about (batched, one request). This gives effectively-once, in-order visibility without any per-connection durable queue on the gateway."
Phase 6: Wrap-Up (2–3 minutes)#
"The design principle is that the socket is a cache of a durable log. That makes gateways stateless enough to kill, which makes deploys and failures routine; the durable path gives the guarantee; and the expensive event — reconnection — is designed for explicitly with jitter, drain spreading, admission control and cheap resumption."
Organizational closer:
"The long-term risk is that every product team pushes through this fleet. I'd run it as a platform with per-topic quotas and a topic registry, and cells so a bad config affects 2% of users. The most important thing I'd publish is the delivery contract — so teams don't put payments status on a best-effort channel and call it real-time."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| Transport debate | 10 min on WebSocket vs SSE vs long poll | Picks in 30s, links the tradeoff, moves on |
| QPS sizing | "10K QPS per server" | Connections/host, memory/conn, handshakes/s |
| Reconnect ignored | "Clients auto-reconnect" | Treats reconnect as the peak and designs for it |
| Delivery hand-waved | "TCP is reliable" | Sequence numbers + sync + gap detection |
| One fan-out design | Pub/sub to all servers | Subscription routing + hot-channel special case |
| No ownership story | Stops at architecture | Platform gateway, topic quotas, cells, delivery contract |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Most system design prompts are stateless request/response problems where scaling means "add more boxes." Real-time delivery is the canonical stateful prompt: each connection is pinned to a host for hours, so every failure, deploy and scaling action has a reconnection cost. Interviewers use it to find candidates who have carried a pager for a stateful fleet — the ones who know that the dangerous moment isn't peak traffic, it's the deploy at 2pm on a Tuesday.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"A user's phone loses signal for 40 seconds in the middle of a busy group chat. Walk me through, message by message, how their client ends up showing exactly the right conversation — no gaps, no duplicates, correct order."
This single question exposes whether the candidate knows that TCP send buffers lie, that the server can't tell a dead connection from a quiet one until heartbeats fail, and that the only robust answer is sequence numbers plus a sync path.
2. Problem Framing & Intent#
2.1 The Four Intents — Explained#
Activity notifications → best-effort, coalesced
- Constraint: very high volume, low value per event (like counts, "someone's typing")
- Strategy: push without durability for ephemeral events; coalesce per client (at most 1 badge update per 2s); drop under backpressure
- Failure mode: stale badge until the next page load
- Who pays for imperfection: nobody meaningfully — product signs off that ephemeral events are lossy
Conversational delivery → gap-free, ordered
- Constraint: user-visible messages must all appear, in order, within ~500ms when online
- Strategy: durable write + per-channel seq + push + sync on reconnect
- Failure mode: gap if the client doesn't detect missing seq; duplicates if it doesn't dedupe
- Who pays for imperfection: users (missed messages), support volume, trust
Collaborative sessions → affinity, convergence
- Constraint: bidirectional, 50–100ms, shared mutable state per room
- Strategy: route all participants of a document/room to the same session server (consistent hash on room_id); the session server is authoritative for ordering; see Collaborative Editing
- Failure mode: session server loss forces room reload from snapshot + op log
- Who pays: users in that room see a 1–3s freeze
Live broadcast → latest-value, sampled
- Constraint: 100K–10M subscribers per channel, one or few publishers
- Strategy: coalescing, sampling, relay tiers, sometimes CDN-delivered SSE or HTTP polling for the long tail
- Failure mode: hot channel saturates router or gateway CPU
- Who pays: every other channel sharing those gateways, unless isolated
🎯 Staff Move: "These four share a gateway fleet but not a delivery contract. I'd tag every topic with its class — ephemeral, durable, session, broadcast — and the platform enforces different rules per class: durable topics must have sequence numbers and a sync endpoint; broadcast topics must declare a coalescing interval."
2.2 When NOT to Use WebSockets#
- Updates every 30s or slower are fine. Poll. 10M clients polling every 60s is ~170K req/s of cacheable GETs — often cheaper and far simpler than a stateful fleet.
- Server→client only, web clients. SSE gives you resume (
Last-Event-ID) and HTTP semantics for free. - The app is backgrounded. Mobile OSes kill background sockets; use APNs/FCM with a hint payload and sync on open.
- Correctness-critical state (payment status, order confirmation). Deliver a hint over the socket if you like, but the client must read the authoritative state from an API. Never make the socket the system of record.
- Very large payloads (files, images). Send a reference; fetch via CDN. See Handling Large Blobs.
2.3 What the Interviewer Leaves Underspecified#
- Delivery semantics. Can users miss events? Must they be ordered? Per channel or globally?
- Multi-device. Does a user have 3 devices each needing every event, with read-state synced?
- Channel size distribution. Median 8 members, but the tail (the company-wide channel, the live event) decides the design.
- Upstream traffic. Typing indicators and presence can exceed message volume 10×.
- Offline behavior. Mobile push? How long do we retain events for sync — 7 days? 30?
- Geography. Single region or global users connecting to the nearest region?
2.4 Precise Terminology#
| Term | What It Means | Why the Precision Matters |
|---|---|---|
| Gateway | Host terminating client sockets | Its only job is holding connections; keep it thin |
| Subscription router | Maps channel → gateways with subscribers | Decides fan-out cost |
| Connection registry | Maps user/conn → gateway | Needed for user-targeted sends; churns during storms |
| Sequence number (seq) | Monotonic per channel (or per user inbox) | Enables gap detection and idempotent apply |
| Sync | Pull API returning events after a cursor | The actual delivery guarantee |
| Drain | Gracefully moving a host's connections elsewhere | Deploys are drains, not restarts |
| Reconnect storm | Many clients reconnecting in a short window | The system's true peak load |
| Hot channel | A channel whose subscriber count or event rate dominates a tier | Needs its own fan-out strategy |
| Cell | An independent slice of the fleet serving a subset of users | Limits blast radius |
3. The Five Fault Lines#
3.1 Fault Line 1: Stateful Gateway vs Stateless Gateway#
The tension: The host holding a socket is the cheapest place to keep per-connection state — subscriptions, unacked buffers, session context. It's also the one place guaranteed to disappear.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Fat gateway (subscriptions, buffers, business logic on the socket host) | Fewest hops, lowest latency, simplest code | Every deploy/crash loses state; gateways can't be killed freely; product logic in the gateway means every team deploys it | On-call (risky deploys), users (lost buffered events) |
| Thin gateway (sockets + auth + frame forwarding; subscriptions derived and re-registered on connect) | Hosts are cattle; deploys are drains; one platform codebase | Extra hop to router; re-registration cost on reconnect | Platform team maintains router; ~1–5ms extra latency |
| Session-affine gateway (for collaborative rooms) | Authoritative per-room ordering in memory | Room reload on host loss | Users in that room (1–3s freeze) |
Staff default: Thin gateways for delivery. The only state on the gateway is the socket and an in-memory index of which local connections subscribe to which channels — rebuildable in seconds from membership when the client reconnects. Collaborative rooms use a separate session tier with explicit affinity, not the delivery gateway.
When to deviate: Small scale (< 100K connections, one team) — a fat gateway with a single deploy unit is fine and faster to build.
🎯 Staff Move: "I want to be able to kill any gateway at any time with no data loss. That's the test for what's allowed to live on it: the socket, yes; anything that isn't reconstructible on reconnect, no."
3.2 Fault Line 2: Broadcast vs Subscription-Routed Fan-Out#
The tension: Broadcasting every event to every gateway is trivially correct and scales as O(events × gateways). Routing only to gateways with subscribers scales with actual interest but needs a routing layer whose state churns with every connect/disconnect.
Back-of-envelope:
Events/s published: 1,000,000
Gateways: 50 (at 200K conns each for 10M)
Broadcast deliveries/s: 50,000,000 gateway messages — ~95% dropped on arrival
Routed (avg 3 gateways/ev): 3,000,000 gateway messages
Router state: channels with ≥1 online member × gateways ≈ 20M entries
Router churn (steady): ~10M conns / 3h avg lifetime ≈ 1K conn events/s × ~20 channels = 20K updates/s
Router churn (storm): 1M reconnects in 120s ≈ 8K conns/s × 20 = 160K updates/s
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Global pub/sub broadcast | Simple; works to ~10–20 gateways | Network and CPU explode linearly with fleet size | Infra bill; eventually latency |
| Pub/sub topic per channel (gateways subscribe to channel topics) | Routing is the pub/sub system's job | Millions of topics; subscribe/unsubscribe storms during reconnects | Pub/sub cluster on-call |
| Channel-owner servers (consistent-hashed owners hold channel → gateways) | Scales horizontally; owner can do per-channel logic (ordering, coalescing) | Owner failover must rebuild membership; hot channel pins one owner | Platform team |
| Hot-channel relay tier | 2M-subscriber channels become 1 send per relay, relays fan to gateways | Extra tier, extra latency (~10–50ms) | Platform team; only for tagged channels |
Staff default: Channel-owner servers (Slack's channel-server shape) for normal channels, sharded by consistent hash on channel_id — see Consistent Hashing. Channels above a threshold (e.g., > 10K online subscribers or > 100 events/s) are promoted to broadcast class: coalesced, sampled, and relayed.
🎯 Staff Move: "Fan-out is designed for the median channel and special-cased for the tail. The median chat has 8 members on 3 gateways. The tail has 2 million on all of them — and for that one I stop promising every event and start promising the latest state."
3.3 Fault Line 3: Push Guarantees vs Pull Sync#
The tension: You can make the socket reliable (server-side per-connection queues, client acks, retransmit) or keep it best-effort and put the guarantee in a pull path.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Fire-and-forget push only | Simplest, fastest | Silent loss on half-open connections, deploys, crashes | Users (missed events), support |
| Per-connection durable queue + acks | Reliable while the host lives | Queue dies with the host; must replicate or persist per connection → heavy; multi-device duplicates | Platform (complexity), storage bill |
| Best-effort push + per-channel seq + sync on reconnect/gap | Guarantee lives in the durable store; gateways stay stateless | Needs a sync API that can handle storm load; clients must implement gap logic | Client teams (logic), store (sync load) |
| Per-user inbox seq (one ordered stream per user) | One cursor per client; simple sync | Write amplification: one event × N recipients inbox writes | Storage; write path cost |
Staff default: Best-effort push + sequence numbers + sync. Use per-channel seq when channels are small-to-medium and the client tracks < ~1K channels; use a per-user inbox seq for notification-style feeds where one cursor per client is valuable and fan-out is bounded.
The subtle part — the race between sync and live push: subscribe first, then sync. If you sync first, events written between the sync read and the subscription land in neither. Subscribe, buffer live events, sync, then apply buffered events deduped by seq.
🎯 Staff Move: "I don't make the socket reliable. I make it unimportant. If every push can be lost and the user still sees everything after sync, then gateways can crash, deploy and drain freely — and that operational freedom is worth more than any ack protocol."
3.4 Fault Line 4: Reconnect Speed vs Fleet Protection#
The tension: The individual user wants to reconnect in 100ms. The fleet wants reconnects spread over minutes. Both can't win during a mass event.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Immediate reconnect | Best single-user UX | Synchronized herds; self-DDoS | Everyone during mass events |
| Fixed-delay retry | Simple | Herds stay synchronized — every client retries at t+1s, t+2s | Everyone |
| Exponential backoff with full jitter | Desynchronizes herds | Unlucky users wait up to the cap (60s) | Some users wait longer |
| Server-directed reconnect delay | Server knows the load; spreads precisely | Only works for graceful drains, not crashes | Users on drained hosts (up to 120s of degraded, push-less state — sync covers it) |
| Admission control on handshakes | Protects healthy hosts from collapse | Rejected clients see errors | Clients beyond the admission rate |
Staff default: All three together — full-jitter backoff (base 1s, cap 60s) in every client SDK, server-directed RECONNECT{after_ms} during drains (spread 0–120s), and per-host admission control (token bucket, e.g., 2K new connections/s/host, fast 503 + Retry-After). Plus cheap resume: TLS session tickets and a signed resume token (HMAC, 10-minute TTL) so reconnection doesn't hit the auth service.
While a client is waiting to reconnect, it isn't blind: it can poll the sync API every 10–30s. The UX degrades from 200ms to ~15s, not to "broken."
🎯 Staff Move: "I'm deliberately making some users wait up to two minutes to reconnect after a deploy, because the alternative is that everyone waits forty-five minutes after the fleet falls over. And because sync covers the gap, 'waiting' means a slightly delayed message, not a missing one."
3.5 Fault Line 5: Platform Gateway vs Product-Owned Sockets#
The tension: Every product wants real-time. If each builds its own socket server, the client holds 4 connections and the company runs 4 fleets with 4 reconnect-storm postures. If one platform owns it, one bad config push disconnects everything.
| Model | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Per-product socket servers | Autonomy, independent deploys | Multiple sockets per client (battery), duplicated storm engineering, inconsistent guarantees | Mobile battery; N teams on-call |
| One platform gateway, shared | One socket per client; one team masters storms | Shared blast radius; platform becomes a bottleneck for new event types | Platform team; everyone during incidents |
| Platform gateway, cellular, with topic registry and quotas | One socket, bounded blast radius, self-serve topics | Needs a registry, schema governance, quota tooling | Platform team (upfront investment) |
Staff default: One multiplexed platform gateway, deployed in cells (e.g., 50 cells × 200K users), with a topic registry: each topic declares owner team, delivery class, max event rate, payload size cap (e.g., 16 KB), and schema. Publish quotas are enforced at the router. Product teams own their topics; the platform owns the fleet and the contract.
🎯 Staff Move: "The platform owns the socket and the guarantee; product teams own the topic and the payload. A team that wants to push 50K events/s to every user files a registry change and gets a quota review — not a surprise at 3am."
4. Failure Modes & Operational Reality#
4.1 Reconnect Storm After a Gateway Crash — Full Timeline#
Scenario: A memory leak in a new gateway build crashes 10 hosts (2M connections) within 30 seconds of each other at peak.
t=0: 10 gateways OOM. 2M sockets close (clients get RST or nothing).
t=+0–5s: Clients that saw the close start jittered backoff (0–1s first attempt).
t=+1s: ~1M reconnect attempts in the first second window (jitter spreads the rest).
t=+1s: LB distributes to 40 healthy hosts → 25K handshakes/s/host vs 2K admission rate.
t=+1s: Admission control rejects excess fast with 503 + Retry-After: 5–30s (randomized).
t=+2s: Healthy hosts' CPU at 70% — handshake work bounded by admission rate.
t=+5s: Clients with dead TCP (no RST) detect it only via missed heartbeat: +25–50s.
t=+30s: Autoscaler adds hosts. Resume tokens mean auth service sees < 5K validations/s.
t=+90s: 1.6M reconnected. Sync API p99 = 300ms at 20K req/s (edge-cached boot payload).
t=+3min: All reconnected. Zero events lost (sync filled gaps). Nobody else disconnected.
Detection: gateway.connections_active (sudden drop by host), gateway.new_connections_per_sec, gateway.admission_rejected_total, sync.requests_per_sec, auth.validations_per_sec.
Mitigation: Admission control and backoff do the work automatically; on-call rolls back the build and ensures the autoscaler isn't adding hosts running the bad build.
Prevention: Canary new gateway builds on one cell for 24h (leaks show up over hours); memory ceiling alerts per host; cap per-host connections so any single crash is ≤ 1% of users.
Owner: Gateway platform on-call. Sync API owner must be looped in — sync is the second system that sees the storm.
4.2 Hot Channel Meltdown#
Scenario: A celebrity joins a public live Q&A channel. 1.5M users subscribe in 10 minutes. The chat rate reaches 20K messages/s.
t=0: Channel has 50K subscribers; owner server handles it fine.
t=+5min: 800K subscribers across all 50 gateways; 8K msgs/s.
Owner server: 8K × 50 gateway sends = 400K sends/s → CPU 100%.
t=+6min: Every other channel hashed to that owner (~2% of all channels) sees 5s delays.
t=+7min: Gateways: 8K msgs/s × 16K local subscribers each = 128M socket writes/s fleet-wide.
t=+8min: Gateway CPU saturates; heartbeats delayed; clients time out and reconnect → storm.
Detection: router.channel_events_per_sec top-K, router.owner_cpu, gateway.socket_writes_per_sec, push_to_visible_ms_p99 for unrelated channels (collateral damage signal).
Mitigation: Promote the channel to broadcast class at runtime: migrate it off the shared owner to a dedicated relay, coalesce to at most 20 displayed messages/s with sampling, and disable per-message push in favor of batched frames every 500ms.
Prevention: Automatic promotion thresholds (> 10K online subscribers or > 100 events/s); isolate broadcast-class channels on their own owner pool; per-channel rate caps in the topic registry.
Owner: Platform (automatic promotion); product (decides sampling policy — "users see a sample of messages in huge channels" is a product decision).
4.3 Silent Gaps — The Client That Never Syncs#
Scenario: A client release has a bug: on reconnect it re-subscribes but skips the sync call when the app resumes from background. Users report "missing messages" days later.
Detection: Server-side: sync.requests_per_reconnect ratio by client version (should be ~1.0; buggy version shows 0.3). Client-side: client.gap_detected_total (events arriving with seq > last_seq + 1) and client.gap_filled_total — the difference is the silent loss.
Mitigation: Force-upgrade or server-side flag that sends a SYNC_REQUIRED frame after every resume.
Prevention: The client SDK owns the seq/sync logic — product client code can't skip it; contract tests in CI that kill the socket and assert no gaps.
Owner: Client platform team (SDK). This is why the delivery contract must be implemented once, in a shared SDK.
4.4 Half-Open Connections and the Heartbeat Trap#
Scenario: A mobile carrier's NAT drops idle mappings after 30s. Your heartbeat is 45s. Connections look alive on the server but packets go nowhere.
Detection: gateway.heartbeat_timeouts_total by ASN/carrier; push_ack_ratio (if clients ack optionally) dropping for specific networks; server-side send-buffer growth per connection.
Mitigation: Lower heartbeat to 20–25s (costs 10M / 25s = 400K tiny frames/s fleet-wide — cheap); detect dead connections within 2 missed heartbeats.
Prevention: Heartbeat interval configurable per client platform; per-connection send buffer cap (e.g., 256 KB) — exceeding it closes the connection rather than letting memory grow.
Owner: Gateway platform team.
4.5 Connection Registry Overload During a Storm#
Scenario: User-targeted sends look up user → gateways in a Redis registry. During a 2M-connection storm, every connect/disconnect is a write.
Detection: registry.write_latency_p99, registry.ops_per_sec, router.lookup_miss_total.
Mitigation: Batch registry updates per gateway (every 1s, not per connection); tolerate stale entries (a send to a gateway without the user is a cheap no-op); rate-limit registry writes during storms.
Prevention: Prefer channel-owner routing (gateway registers interest in channels, batched) over per-user registry for most traffic; shard the registry; TTL entries with heartbeat refresh (every 60s) so crashes self-clean without delete storms.
Owner: Platform team. Key lesson: the system tracking the storm must survive the storm.
4.6 Slow Consumers and Head-of-Line Blocking#
Scenario: A client on a 2G connection can't drain its socket fast enough; the gateway's write buffer for it grows, and if writes are synchronous per channel, other clients wait.
Detection: gateway.conn_send_buffer_bytes p99, gateway.write_queue_depth, per-host event-loop lag.
Mitigation: Non-blocking writes with per-connection bounded buffers; on overflow drop ephemeral events first, then close the connection (client will sync on reconnect).
Prevention: Class-aware buffers: ephemeral events are droppable, durable events are recoverable by sync — so the gateway can always drop.
Owner: Platform team.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Gateway crash | connections_active drop per host | Connections on that host (≤ 1% by design) | Admission control, jittered reconnect, sync | Gateway platform |
| Bad deploy across fleet | Error/crash rate on new build | Up to one cell per wave | Cell-by-cell rollout, auto-rollback | Gateway platform |
| Hot channel | Top-K channel event rate, owner CPU | Channels sharing the owner, gateways with members | Promote to broadcast class, coalesce, sample | Platform + product |
| Silent client gaps | sync_per_reconnect by client version | Users on buggy version | Server-forced sync frame, SDK fix | Client platform |
| Half-open connections | Heartbeat timeouts by ASN | Users on affected networks | Shorter heartbeat, buffer caps | Gateway platform |
| Registry overload | Registry p99, ops/s | User-targeted sends delayed | Batch updates, tolerate stale | Platform |
| Sync API overload after storm | sync.p99, sync.rps | All reconnecting clients | Edge cache boot payload, rate-limit sync per client | Sync API owner |
| Publish flood from one team | router.publish_rate{topic} | Everyone sharing router | Per-topic quotas, circuit-break topic | Topic owner team |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Delivery contract | "TCP/WebSocket is reliable" | Best-effort push + seq + sync; subscribe-then-sync race handled | Publishes the contract as a platform SLO; forbids system-of-record data on the socket |
| Capacity model | QPS per server | Connections/host, memory/conn, handshakes/s as the headroom driver | Host size as a priced blast-radius decision; $/million concurrent |
| Fan-out | Pub/sub broadcast | Subscription-routed with channel owners; hot-channel promotion | Org limits on channel size and publish rates; registry governance |
| Failure | Auto-reconnect | Jitter, drain spreading, admission control, resume tokens, randomized lifetimes | Cells, progressive rollout enforced by tooling, mass-reconnect game days |
| Ownership | Team-owned socket server | Platform gateway; product-owned topics | Decides centralization boundary; funds a client SDK as the contract carrier |
| Evolution | Add servers | Multi-region with regional gateways and sync | 3-year path from single fleet to cellular multi-product platform |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Separates delivery from durability | "The socket is a cache of the log. Push is latency; sync is the guarantee." |
| Sizes for failure, not steady state | "Idle sockets cost memory. The CPU peak is reconnection, so I size for handshakes/s." |
| Treats deploys as the main risk | "A deploy is a drain. I'd spread reconnects over 120 seconds and do 2% per wave." |
| Special-cases the tail | "Median channel has 8 members. The 2M-member channel gets coalescing and a relay tier." |
| Handles the race | "Subscribe first, then sync, then apply buffered events deduped by seq." |
| Names ownership | "Platform owns the socket and the guarantee; product teams own topics under quotas." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| 10+ minutes on WebSocket vs SSE vs polling | Spends time on the least differentiating decision |
| "TCP guarantees delivery" | Misunderstands half-open connections and send buffers |
| Reconnection never discussed | Misses the dominant failure mode of a stateful fleet |
| Every event broadcast to every server | Doesn't scale past ~20 hosts; no awareness of fan-out cost |
| Per-connection queues as the guarantee | State dies with the host; doesn't handle multi-device |
| Sockets used for writes with business logic | Couples gateway deploys to every product change |
5.4 Common False Positives#
- Knowing the WebSocket frame format ≠ designing a fleet. Opcode trivia doesn't predict operational judgment.
- "We'll use a managed service" ≠ design. Fine as a build-vs-buy answer, but the candidate still needs to state the delivery contract and reconnect behavior.
- Redis pub/sub fluency ≠ fan-out design. The question is routing cost at 50+ hosts and hot channels, not the API.
- Mentioning Kafka ≠ durability story. The log is only useful if the client has a cursor into it and a sync path.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing + intent | 0–4 min | Delivery contract, numbers, transport in 30s |
| Entities + API | 4–6 min | Seq per channel, sync API, writes over HTTP |
| High-level design | 6–11 min | Gateway, router, log, store, sync |
| Deep dive: reconnect storms | 11–20 min | Jitter, drain, admission, resume |
| Deep dive: fan-out + hot channels | 20–30 min | Channel owners, promotion, coalescing |
| Deep dive: delivery | 30–40 min | Seq, gaps, dedupe, subscribe-then-sync |
| Wrap-up | 40–45 min | Platform, cells, contract |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Direction |
|---|---|---|
| "Add presence (online/offline)." | Upstream write amplification | Heartbeat-driven TTL, coarse status, fan-out presence only to open conversations; see the pattern page |
| "Add typing indicators." | Ephemeral class | Non-durable, coalesced to 1 per 3s per user, droppable |
| "A user has 5 devices." | Multi-device delivery | Per-device cursors; send to all device connections; read-state as its own synced event |
| "Now go multi-region." | Stateful geo design | Regional gateways; cross-region event replication; sync from nearest replica with seq |
| "What if Kafka is down?" | Durable path dependency | Writes fail (fail-closed for durable topics); ephemeral topics bypass the log |
| "Mobile app in background?" | Transport boundaries | OS push hint + sync on foreground; close the socket |
6.3 What to Deliberately Skip#
- The WebSocket handshake details and frame format. One sentence.
- Message storage schema in depth.
(channel_id, seq)clustered key in a wide-column store is enough; link to Chat Messaging. - Presence internals. Covered in Real-time Updates; mention TTL-based presence and move on.
- Authentication flows. Token at connect, resume token after; that's it.
6.4 Follow-Up Questions to Expect#
- "How do you know a connection is dead?" — (Heartbeats every 25s, dead after 2 misses; send-buffer cap closes stuck connections.)
- "Where does ordering come from?" — (Per-channel seq assigned by the write path — a single sequencer per channel, e.g., the channel owner or a DB counter; no global order.)
- "How big can one host get?" — (Hundreds of thousands to ~1M+ is technically feasible; I'd cap at ~200K for blast radius.)
- "How do you deploy the gateway?" — (Drain cell by cell, 2% of hosts per wave, server-directed reconnect with 0–120s spread.)
- "What does the load balancer do?" — (L4 pass-through or TLS termination with session tickets; least-connections balancing — round robin skews toward new hosts badly after storms.)
- "What if the sync API is down?" — (Clients keep receiving live pushes; gaps are detected and retried with backoff; alert on
sync.error_rate.) - "How do you rebalance after scaling up?" — (Randomized max connection lifetime of 2–4h moves load gradually; never mass-disconnect to rebalance.)
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design a system to push real-time updates to our users."
Staff Answer
"First, the delivery contract: can users miss events? For chat-like updates, no; for likes and typing, yes. I'll design for the strict case and let ephemeral events opt out of durability. The transport is WebSocket for foreground clients and OS push for background — I won't spend time there. The key principle is that the socket is a cache of a durable log: every durable event gets a per-channel sequence number when written, push is best-effort, and clients sync from their last sequence on every reconnect. Numbers: 10M concurrent, 1M events/s, fan-out 10, p99 push latency under 500ms. I'll cover the gateway fleet, fan-out routing including hot channels, and reconnect storms — which are the real peak load."
Why this is L6:
- States the delivery contract before drawing anything
- Dismisses the transport question quickly and correctly
- Names reconnect storms as the peak before being asked
What L7 adds:
- "I'd also ask how many product teams will publish through this — if more than one, it's a platform with a topic registry and quotas."
- Frames host sizing as a blast-radius and cost decision
❌ Common L5 Trap
"We'll use WebSockets because they're bidirectional. Let me compare them with SSE and long polling... Each server subscribes to Redis pub/sub, and when a message comes in we send it to the connected users."
Why this misses: A correct small-scale design; no delivery contract, no reconnect story, fan-out that's O(events × servers). The interviewer now has to pull every Staff-relevant topic out one question at a time.
Drill 2: Capacity Math#
Prompt: "How many gateway hosts do you need for 10M concurrent connections?"
Staff Answer
"Three independent constraints. Memory: ~30 KB per connection with a tuned runtime → 200K connections ≈ 6 GB, fine on a 32 GB host. Steady CPU: 10M deliveries/s across the fleet — at ~200K frame writes/s per host that's ~50 hosts. Reconnect headroom: if the worst tolerated event is losing 5% of the fleet at once (500K clients) and we spread over 60s, that's ~8K handshakes/s, which the rest of the fleet absorbs easily with ECDSA and session tickets. So ~50 hosts at 200K each, plus 20–30% headroom → ~65 hosts. I could run 10 hosts at 1M each, but then one crash is 10% of users reconnecting — I'd rather pay for 55 extra machines."
Why this is L6:
- Separates memory, throughput and reconnect constraints
- Chooses density with blast radius in mind and prices it
What L7 adds:
- Expresses it as cost per million concurrent per month and compares with a managed service
- Aligns host count with cell design (e.g., 13 cells × 5 hosts)
Drill 3: Make the Delivery Guarantee Concrete#
Prompt: "How do you guarantee a user sees every message?"
Staff Answer
"The write path assigns a monotonic seq per channel and persists before pushing. The client SDK keeps last_seq per channel. On push: if seq = last+1 apply; if greater, gap → sync since last; if less or equal, drop as duplicate. On connect: subscribe first, buffer live events, call sync with all cursors in one batched request, apply, then drain the buffer deduped by seq. That closes the race where an event lands between sync and subscribe. Retention for sync is 30 days; beyond that the client does a full reload. Result: effectively-once, in-order per channel, with no per-connection state on the gateway."
Why this is L6:
- Covers gap, duplicate and race cases explicitly
- Puts the guarantee in the durable path, not the socket
What L7 adds:
- Ships the logic in a single client SDK so no product team re-implements it
- Defines a measurable SLO:
gap_detected − gap_filled = 0and 99.9% of events visible < 2s
Drill 4: The Deploy#
Prompt: "You need to deploy a new gateway version to 65 hosts holding 10M connections. Go."
Staff Answer
"Canary one cell (~5 hosts) for 24h — memory leaks in connection servers show up over hours, not minutes. Then waves of ~2% of hosts. For each host: remove from LB for new connections, send RECONNECT{after_ms} with a uniform random 0–120s delay to each client, wait for connections to drain to near zero (hard kill after 5 minutes), restart, re-add. 200K connections over 120s is ~1.7K/s arriving across the remaining fleet — trivial. Clients resume with a resume token so the auth service sees nothing. Total: ~50 waves, a few hours. Auto-halt if connections_active fleet-wide drops more than 3% unexpectedly or push_to_visible_ms_p99 exceeds 1s."
Why this is L6:
- Treats a deploy as a drain with spread, not a restart
- Includes canary duration matched to the failure mode (leaks)
- Has an automatic halt condition
What L7 adds:
- Makes progressive, cell-by-cell deploys enforced by the deploy tooling for every stateful fleet in the org
- Adds randomized connection lifetimes so deploys and routine churn look the same
Drill 5: The Hot Channel#
Prompt: "A live sports channel has 3M concurrent subscribers and 5K comments/s."
Staff Answer
"5K/s × 3M = 15B socket writes/s — impossible and pointless, nobody reads 5K comments/s. So I change the promise: the score and game state are latest-value, pushed at most every 500ms. Comments are sampled — clients get a stream of ~20/s, weighted toward friends and highly-rated comments — batched into one frame per 500ms. Architecturally, the channel is promoted to broadcast class: off the shared channel owner onto a dedicated relay per region, which sends one batched frame per gateway per interval. 3M subscribers × 2 frames/s = 6M writes/s fleet-wide, spread over all hosts — manageable. The full comment firehose is still durably stored for replay."
Why this is L6:
- Changes the product contract rather than scaling a bad promise
- Quantifies both the naive and the redesigned fan-out
- Isolates the hot channel so it can't hurt other channels
What L7 adds:
- Makes the sampling policy an explicit product decision with an owner
- Sets org-wide thresholds for automatic promotion so this isn't an incident each time
Drill 6: Reconnect Storm Math#
Prompt: "A network blip disconnects 4M clients at once. What happens?"
Staff Answer
"Clients start full-jitter backoff: first attempt uniform in 0–1s, second 0–2s, and so on, capped at 60s. Even with jitter, ~2M attempts land in the first second. Admission control on each host admits ~2K/s — with 65 hosts that's 130K/s admitted, the rest get a fast 503 with a randomized Retry-After. So the fleet reconnects 4M in ~30–40s without any host falling over. Handshakes use session tickets; resume tokens avoid auth calls; each client makes one batched sync call — 4M sync calls over ~40s is ~100K/s, which the sync API must be sized for, or it rate-limits per client and serves boot payloads from an edge cache."
Why this is L6:
- Quantifies how jitter and admission interact
- Identifies the second system in the blast path (sync/boot)
What L7 adds:
- Runs a quarterly game day that disconnects a cell to verify the math
- Coordinates with auth and sync owners so every system on the reconnect path has a storm budget
Drill 7: Multi-Tenant Platform#
Prompt: "Five product teams want to push events. How do you let them?"
Staff Answer
"One multiplexed socket per client, one platform. Teams register topics: owner, delivery class (ephemeral, durable, broadcast), max publish rate, payload cap (16 KB — anything bigger is a reference), and schema. The router enforces per-topic quotas and can circuit-break a topic that exceeds it without affecting others. Durable topics must use the platform's seq/sync — teams don't build their own. The platform team owns the fleet and is paged for fleet health; topic owners are paged for their topic's publish errors and quota breaches."
Why this is L6:
- Avoids N sockets per client and N fleets
- Puts guardrails (quotas, payload caps) at the platform boundary
- Splits paging responsibility
What L7 adds:
- Decides what's out of scope for the platform (e.g., collaborative session state)
- Sets chargeback by connections-minutes and deliveries so teams see their cost
Drill 8: Build vs Buy#
Prompt: "Why not use a managed real-time service?"
Staff Answer
"Below ~1M concurrent connections, or for a product that isn't real-time-centric, I would — managed services price per connection-minute and message, and they handle storms for you. Three things push toward building: cost at 10M+ concurrent (per-connection pricing scales linearly while self-hosted cost per connection falls with density), control over the reconnect and drain behavior, and integration with our own sequenced log. Middle path: use a managed service for the socket layer but keep the durable log and sync API ours — then the vendor is a replaceable transport, not our system of record."
Why this is L6:
- Gives a threshold and the cost curve argument
- Keeps the guarantee in-house regardless of vendor
What L7 adds:
- Designs the exit: the client SDK abstracts transport so switching vendors is a client flag
- Uses the Build vs Buy framework with a 3-year TCO
Drill 9: Multi-Region#
Prompt: "Users in Europe complain of 400ms push latency. Go multi-region."
Staff Answer
"Put gateways in each region and let clients connect to the nearest via GeoDNS/anycast. The hard part is routing: a channel with members in three regions. The channel owner (and its seq assignment) lives in one home region; events replicate to regional routers, which fan out locally. Cross-region latency is then paid once per event, not once per subscriber. Sync reads from the regional replica — if its seq is behind the client's cursor, it waits briefly or proxies to home. Region failure: clients reconnect to the next region with jitter; sync from any replica covers gaps."
Why this is L6:
- Pays cross-region cost once per event, not per subscriber
- Keeps per-channel ordering with a home region
What L7 adds:
- Decides home-region placement policy (by creator, by majority membership) and its data residency implications
- Budgets cross-region replication cost as a line item
Drill 10: Protocol Change Without an Outage#
Prompt: "We want to switch from JSON frames to compressed binary frames."
Staff Answer
"Negotiate at connect: the client advertises supported encodings in the handshake, the gateway picks the best mutual one. Deploy gateway support first (both encodings), then roll the client SDK. Measure bandwidth and CPU per host — compression trades gateway CPU for bandwidth, and for mobile clients bandwidth and battery usually win. Keep JSON supported until old client versions fall below 1% of connections, then deprecate with a server-forced upgrade prompt."
Why this is L6:
- Uses negotiation to make the change a two-way door
- Measures the CPU/bandwidth tradeoff before committing
What L7 adds:
- Defines a client version support policy (e.g., 12 months) across all real-time features
- Ties the change to a cost outcome (egress $ saved per month)
8. Deep Dive Scenarios#
Deep Dive 1: The Deploy That Took Down Everything#
Context: A config change to the gateway's TLS settings was pushed to all hosts at once. Every host restarted within 90 seconds. 10M clients reconnected; the auth service collapsed; the outage lasted 38 minutes. You're asked to lead the post-incident work.
Questions to Surface First:
- Why could a config change reach 100% of hosts at once?
- Did clients back off with jitter, or reconnect in lockstep?
- Why did reconnection hit the auth service at all?
- Which systems were on the reconnect path, and what were their storm budgets?
Typical L5 Approach: Adds more auth service capacity and makes clients back off longer. Both help; neither addresses why a single push could restart the whole fleet.
Staff Approach: Fixes each layer on the reconnect path: config changes go through the same cell-by-cell progressive rollout as code; resume tokens remove auth from reconnect; handshake admission control per host; SDK full-jitter backoff verified in tests; boot/sync payload cached at the edge.
Principal Approach: Treats it as an org failure-posture problem: config and code must share one progressive delivery system for every stateful fleet, enforced by tooling. Introduces cells with hard isolation so no single change can touch more than one cell per hour, and schedules a quarterly mass-reconnect game day with auth and sync owners in the room.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Block new connections at LB; admit by region/cell in steps |
| Triage | Confirm auth saturation from reconnect validations; confirm lockstep reconnect timing |
| Quick fix | Enable admission control; raise Retry-After randomization; admit 10% of traffic per 2 minutes |
| Guardrails | Config pushes require the progressive rollout pipeline |
| Post-mortem | Config bypassed the deploy pipeline; auth on the reconnect path; no storm budgets |
Metrics to Watch: gateway.new_connections_per_sec, auth.validations_per_sec, gateway.admission_rejected_total, sync.rps, connections_active fleet-wide.
Organizational Follow-up: Every system on the reconnect path (LB, gateway, auth, sync, boot) publishes a storm budget; game day verifies it.
Ownership Question: "Who is allowed to push a config change to all gateways at once?" Staff answer: Nobody, by construction. The pipeline enforces cell-by-cell rollout; the emergency override requires the gateway on-call plus an incident commander and still rolls at most 25% per step.
Key Takeaway: "In a stateful fleet, the blast radius of a change is measured in reconnects. Design the rollout, not just the recovery."
What clears the Staff bar:
- Maps every system on the reconnect path
- Removes auth from reconnect rather than scaling it
- Puts config under the same rollout discipline as code
Deep Dive 2: Messages Silently Missing#
Context: Support tickets say users "sometimes miss messages." Metrics show push success at 99.99%. No errors anywhere.
Questions to Surface First:
- What does "push success" measure — a write into the kernel send buffer?
- Do clients report gaps? Do they sync on every reconnect, including resume from background?
- Is it concentrated by client version, platform or network?
Typical L5 Approach: Adds retries on push and increases heartbeat frequency. Push "success" was never the problem.
Staff Approach: Instruments the contract, not the transport:
client.gap_detected_total,client.gap_filled_total, andsync_per_reconnectby client version. Finds an iOS build skipping sync on foreground resume. Fixes in SDK, forces sync server-side for affected versions.
Principal Approach: Establishes an end-to-end delivery SLO measured from the client (e.g., 99.9% of events visible within 2s, 100% after sync) and makes it the platform's headline metric; server-side "push success" is demoted to a debugging metric. Funds a synthetic client fleet that continuously verifies the contract.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Break down tickets by client version/platform |
| Triage | Compare sync_per_reconnect across versions |
| Quick fix | Server sends SYNC_REQUIRED after resume for affected versions |
| Guardrails | SDK contract tests: kill socket mid-stream, assert no gaps |
| Post-mortem | Metrics measured the transport, not the user-visible outcome |
Metrics to Watch: client.gap_detected_total − client.gap_filled_total, sync_per_reconnect{client_version}, synthetic event_visible_ms p99.
Organizational Follow-up: Client platform team owns the SDK; app teams can't bypass it.
Ownership Question: "Who owns 'user sees every message'?" Staff answer: The real-time platform team owns the SLO end-to-end, including the client SDK — because the guarantee is implemented half on the client.
Key Takeaway: "A 99.99% push success rate says nothing about whether users saw their messages. Measure at the client."
What clears the Staff bar:
- Distrusts transport-level success metrics
- Locates the guarantee in client logic and owns it
Deep Dive 3: Onboarding a Massive Tenant#
Context: A large enterprise customer is moving 400K employees onto the product. They have a single all-company channel with 400K members, and every morning at 9am local time ~300K people log in within 20 minutes.
Questions to Surface First:
- How many of the 400K are online simultaneously, and how many events/s does the all-company channel get?
- What happens at 9am — login wave or thundering herd?
- Are they concentrated in one region/cell?
Typical L5 Approach: Adds gateway capacity for 400K more connections.
Staff Approach: Recognizes two separate problems. The morning wave: 300K connects in 20 min is only ~250/s — fine — but boot payload for a 400K-member workspace is huge; serve it from an edge cache and lazy-load member lists. The all-company channel: pre-promote to broadcast class, restrict posting to admins or rate-limit it, and lazy-load member presence rather than fanning it out.
Principal Approach: Negotiates product limits with the customer (announcement-only channels above 50K members), places the tenant across multiple cells to avoid concentrating risk, and prices the dedicated capacity into the contract.
Staff Approach — Full Reasoning
| Dimension | Staff Answer |
|---|---|
| Connections | +400K spread across cells; no single cell > 25% of the tenant |
| Boot payload | Edge-cached, delta-based; member list paginated |
| Big channel | Broadcast class, posting rate-limited, reactions coalesced |
| Presence | Not fanned out for channels > 10K; fetched on demand |
Metrics to Watch: boot_payload_bytes_p99{tenant}, channel_events_per_sec{channel}, cell_connections{tenant}.
Organizational Follow-up: Enterprise onboarding checklist includes a real-time capacity review.
Ownership Question: "Who approves a 400K-member channel with open posting?" Staff answer: Product, with a platform-set hard limit; above it, the channel is announcement-only by default.
Key Takeaway: "Large tenants don't break the fleet with connections. They break it with one giant channel and one giant boot payload."
What clears the Staff bar:
- Separates connection load from channel fan-out load
- Uses product limits as an engineering tool
Deep Dive 4: Post-Mortem — Hot Channel Took Down a Shard#
Context: A viral live event channel hashed to the same channel-owner server as 2% of all channels. For 12 minutes those channels had 10–30s delays. Write the post-mortem actions.
Questions to Surface First:
- Why wasn't the channel promoted automatically?
- Why did one channel's load affect unrelated channels?
- How long did detection take?
Typical L5 Approach: Adds more owner servers so each owns fewer channels.
Staff Approach: Automatic promotion at thresholds (> 10K online or > 100 events/s) with migration to a dedicated broadcast pool; per-channel CPU accounting on owners; alert on
push_to_visible_ms_p99for a sample of unrelated channels (collateral damage detector).
Principal Approach: Defines "noisy neighbor isolation" as a platform requirement across all shared tiers (router, gateway, sync), with per-tenant/per-channel budgets; reviews every shared tier for the same failure class.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Manually promote channel; rehash its owner |
| Triage | Top-K channels by owner CPU |
| Quick fix | Coalescing on the hot channel |
| Guardrails | Automatic promotion thresholds |
| Post-mortem | No isolation between channels on an owner |
Metrics to Watch: router.owner_cpu{owner}, router.channel_events_per_sec top-K, collateral-latency probe.
Organizational Follow-up: Product signs off on sampling policy for broadcast channels.
Ownership Question: "Who decides that users in huge channels see a sample of messages?" Staff answer: Product, in advance — the platform implements the mechanism and the thresholds.
Key Takeaway: "Consistent hashing distributes channels, not load. The tail needs its own lane."
What clears the Staff bar:
- Detects collateral damage, not just the hot spot
- Automates the promotion decision
Deep Dive 5: Multi-Region Expansion#
Context: The company is launching in Asia-Pacific. Leadership wants < 200ms push latency there and regional failover.
Questions to Surface First:
- What fraction of channels span regions?
- Where does per-channel ordering (seq assignment) live?
- Data residency requirements for message content?
Typical L5 Approach: Deploys the full stack in APAC and routes by geography.
Staff Approach: Regional gateways and routers; per-channel home region assigns seq; events replicate to other regions once per event; sync from regional replicas with seq-awareness; regional failover by jittered reconnect to the next region.
Principal Approach: Decides the home-region policy and its residency implications with legal; budgets cross-region egress; sequences the rollout as read-only regional delivery first, then regional writes, with a failover game day before GA.
Staff Approach — Full Reasoning
| Dimension | Staff Answer |
|---|---|
| Latency | Regional gateways; cross-region paid once per event |
| Ordering | Home-region sequencer per channel |
| Failover | Clients reconnect to next region with jitter; sync covers gaps |
| Cost | Egress ≈ events × avg remote regions × payload size |
Metrics to Watch: cross_region_replication_lag_ms, push_to_visible_ms_p99{region}, sync_proxy_to_home_total.
Organizational Follow-up: Regional on-call coverage; residency review.
Ownership Question: "Who decides a channel's home region?" Staff answer: The platform, by an automatic policy (creator's region, rebalanced by majority membership after 30 days), reviewed by legal for residency-bound tenants.
Key Takeaway: "Replicate events between regions, not deliveries."
What clears the Staff bar:
- Keeps ordering with a clear owner
- Treats failover as a reconnect storm to design for
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Explain why the socket is a cache of a durable log and design seq + sync with gap/duplicate/race handling
- Size a gateway fleet by connections, memory and reconnect handshakes, and choose host size for blast radius
- Design subscription-routed fan-out and promote hot channels to a broadcast class
- Design against reconnect storms: jitter, drain spreading, admission control, resume tokens, randomized lifetimes
- Run deploys as drains, cell by cell
- Define platform vs product ownership with a topic registry and quotas
- Extend the design to multi-region with a home-region sequencer
The Bar for This Question#
Mid-level (L4): WebSocket server, a pub/sub bus, clients reconnect. Works for a demo.
Senior (L5): Adds a connection registry, horizontal scaling, heartbeats, maybe SSE vs WebSocket analysis. Correct at moderate scale; delivery assumed reliable; reconnect storms and hot channels not considered.
Staff+ (L6): Makes the durable log the source of truth with seq and sync; thin gateways that can be killed; fan-out routed by subscription with a hot-channel lane; reconnect designed as the peak load; deploys as drains; ownership split between platform and topic owners. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "Reliable WebSocket Delivery Is the Wrong Goal"#
| Evidence | Implication |
|---|---|
| Half-open TCP connections accept writes that never arrive | Server-side "success" is fiction |
| Per-connection queues die with the host | Reliability requires durable state elsewhere anyway |
The Staff position: Make the socket best-effort and the sync path authoritative.
Why this matters in interviews: It collapses a sprawling ack/retry design into one clean invariant.
10.2 "Your Peak Load Is a Deploy, Not a Traffic Spike"#
| Evidence | Implication |
|---|---|
| Idle connections cost memory, not CPU | Steady-state CPU is low |
| Every restart forces reconnects with full handshakes and boot payloads | Deploys and crashes set the CPU headroom |
The Staff position: Size for reconnect handshakes/s and design deploys as spread drains.
Why this matters in interviews: It shows you've operated a stateful fleet.
10.3 "Most Apps Should Poll"#
| Evidence | Implication |
|---|---|
| 60s polling of a cacheable endpoint is stateless and CDN-friendly | No fleet, no storms |
| Many "real-time" features tolerate 30–60s staleness | The stateful fleet is overkill |
The Staff position: Ask what staleness product will accept before building sockets.
Why this matters in interviews: Knowing when not to build is a Staff-level signal.
10.4 "One Million Connections Per Host Is a Vanity Metric"#
| Evidence | Implication |
|---|---|
| Density records are achievable (WhatsApp ~2M in 2012) | Memory is not the constraint |
| Losing a 1M-connection host is a 1M-client storm | Blast radius is the constraint |
The Staff position: Cap at ~100–200K per host and spend the extra machines on blast radius.
Why this matters in interviews: It reframes capacity as risk management.
10.5 "Presence Is More Expensive Than Messaging"#
| Evidence | Implication |
|---|---|
| Every connect/disconnect/idle change can fan out to every contact | Presence events often exceed message events 10× |
| Users rarely act on precise presence | Coarse, lazy presence loses little value |
The Staff position: Presence is lazy, coarse, and scoped to what's on screen.
Why this matters in interviews: It shows you can identify the hidden amplifier.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
At Staff level, the real-time system is a fleet to design. At Principal level, it's shared infrastructure that every product team will depend on — and therefore the org's single largest correlated-failure risk for user-facing freshness. The L7 questions are: what contract does the platform offer, what does it refuse to carry, how is its blast radius partitioned, and how is its cost attributed so that teams don't treat live updates as free.
The Org-Level Fault Line#
One real-time platform vs per-product real-time stacks. One platform means one socket per client (battery, connection cost), one team that has mastered storms, and one contract. It also means one set of cells whose failure is felt by every product at once and a platform backlog that gates every new real-time feature. The L7 position: centralize the socket, the delivery contract, the SDK and the storm engineering; decentralize topics, payloads and product semantics via self-serve registry and quotas; keep collaborative session servers as a separate, affinity-based tier owned closer to the product.
Cost Model#
Assumptions: ~$0.10/host-hour for a 16-vCPU, 32 GB instance class; 200K connections/host; egress at ~$0.02–0.05/GB blended; engineers at $250K fully loaded.
| Scale | Concurrent | Compute / month | Egress / month | Headcount | On-call |
|---|---|---|---|---|---|
| Small | 100K | ~$1–2K (or a managed service at ~$2–5K) | ~$1K | 1–2 part-time; buy | Business hours |
| Medium | 10M | ~$6–10K for ~65 gateways + routers/sync ~$15K | ~$20–60K | 5–8 (gateway, router, SDK, sync) | 24/7; storms dominate pages |
| Large | 100M | ~$150–250K across cells and regions | ~$300K+ | 15–25 across platform, SDK, regional ops | Follow-the-sun, cell-scoped paging |
The surprising line item is egress, not compute: frame compression and coalescing are cost projects, not just performance projects.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversibility Cost |
|---|---|---|
| Per-channel seq + sync as the delivery contract | One-way (good) | Clients in the wild depend on it for years |
| Client SDK owning reconnect/sync logic | One-way | Old app versions persist 12–24 months |
| Wire protocol (framing, encoding) | Two-way with negotiation | Cheap if negotiated at connect; expensive if hardcoded |
| Transport vendor (managed vs self-hosted) | Two-way if SDK abstracts it | Weeks behind an abstraction; quarters without |
| Home-region sequencing model | Mostly one-way | Changing ordering semantics breaks client assumptions |
| Host size / cell size | Two-way | Rebalance via randomized lifetimes |
The Standard I'd Write#
RFC: Real-Time Delivery Platform Standard (v1)
Scope: Any feature that pushes server events to clients.
MUST:
- Use the platform gateway and client SDK; no product-owned sockets.
- Register every topic with owner, delivery class, max publish rate and payload cap (≤ 16 KB).
- Durable topics MUST assign a per-channel seq on write and expose sync; the socket MUST NOT be the system of record.
- Clients MUST implement full-jitter backoff (via SDK) and honor server
RECONNECTdelays.- Fleet changes (code or config) MUST roll out cell by cell.
SHOULD:
- Broadcast-class topics declare a coalescing interval and sampling policy.
- Ephemeral events (typing, presence) be droppable under backpressure.
Exceptions: Collaborative session servers (affinity tier) — reviewed by the platform team; time-boxed.
Success metrics: 99.9% of events visible < 2s at the client; zero unfilled gaps; no single change affecting > 1 cell per hour; one socket per client.
What I'd Tell the VP#
Real-time updates are now part of every product we ship, and they all run through connections that have to be re-established whenever anything changes on our side. Today, a single bad deploy can disconnect every user at once, and our biggest outages have been self-inflicted in exactly that way. I'm proposing we consolidate onto one real-time platform, split it into independent cells so any failure touches a few percent of users, and guarantee — measured on users' devices — that nobody misses a message. It's about six engineers for three quarters, and it removes four duplicate systems and our most common cause of company-wide incidents.
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Sees the platform | "Five teams will publish through this; I'll give them a registry and a contract, not a socket." |
| Prices it | "At 10M concurrent, egress is the biggest line item — compression is a cost project." |
| Designs the org's failure posture | "Cells of 200K users; no change touches more than one cell per hour." |
| Knows what not to centralize | "Collaborative session state stays in its own affinity tier." |
| Plans for long-lived clients | "The SDK's sync logic is a one-way door — old clients live for two years." |
Staff answers that L7 interviewers find insufficient:
- "Jittered backoff and admission control" — correct, but doesn't address the org-wide correlated risk of one fleet for all products.
- "Seq + sync" — correct, but doesn't make it a published contract with a client-measured SLO.
- "Platform owns the gateway" — correct, but no chargeback, quotas or exception process.
🧭 Principal Move: "The technical design is a solved pattern. What I'd spend my time on is the contract: what the platform guarantees, what it refuses to carry, and how its blast radius is partitioned — because every product team's freshness now depends on one fleet."
Appendices
Appendix A: Gateway Mechanics#
A.1 Connection Lifecycle#
A.2 Gateway Event Loop (pseudocode)#
on_connect(req):
if not admission.try_take(): return reject(503, retry_after=rand(5,30))
user = verify_resume_token(req.token) or auth_service.verify(req.jwt)
conn = Conn(user, device, send_buf_cap=256KB, max_life=rand(2h,4h))
for ch in membership.channels_for(user): # cached, batched
local_index[ch].add(conn)
router.register_interest_batch(local_index.new_channels()) # flushed every 1s
conn.send(HELLO{heartbeat_ms=25000})
on_router_event(ch, frame):
for conn in local_index[ch]:
if not conn.try_write(frame): # non-blocking
if frame.class == EPHEMERAL: drop
else: conn.close(reason=SLOW) # client will sync on reconnect
on_tick():
close conns with 2 missed heartbeats
for conn past max_life: conn.send(RECONNECT{after_ms=rand(0,30000)})
Appendix B: Data Model and Sequencing#
| Store | Key | Notes |
|---|---|---|
| Event store | (channel_id, seq) clustered | Wide-column (e.g., Cassandra); retention 30 days for sync; older via archive |
| Seq allocator | per channel_id | Channel owner in memory with durable high-water mark, or a DB counter; one writer per channel |
| Membership | user_id → channels, channel_id → members | Source for subscriptions; cached on gateways |
| Interest map | channel_id → gateways | Soft state, rebuilt by gateway re-registration; TTL 60s |
| Resume token | HMAC(user, device, exp) | 10-minute TTL; verified locally, no auth call |
Why per-channel and not global order: global ordering needs a single sequencer — a throughput ceiling and a SPOF. Users only perceive order within a conversation.
Appendix C: Fan-Out Mechanisms — Quick Comparison#
| Mechanism | Throughput | Hot-Channel Behavior | Failure Behavior | Use For |
|---|---|---|---|---|
| Redis pub/sub broadcast | Limited by fleet × events | Everyone gets everything | Fire-and-forget; loss on disconnect | < 20 gateways |
| Redis pub/sub per channel topic | Good | Hot topic pins one Redis node | Subscribe storms on reconnect | Medium scale |
| Kafka + channel-owner servers | High, replayable | Owner hot-spot; needs promotion | Owner failover rebuilds from interest map | Large scale |
| Relay tier (regional) | Very high for broadcast | Designed for it | Relay loss → gateways fall back to owner | Broadcast class |
Appendix D: Client Contract#
- Connect with resume token if present; else JWT.
- Full-jitter backoff:
delay = random(0, min(60s, 1s × 2^attempt)); honorRetry-AfterandRECONNECT{after_ms}. - Subscribe → buffer → sync → drain buffer (dedupe by seq).
- On gap (seq > last + 1): sync that channel; on seq ≤ last: drop.
- While disconnected > 10s: poll sync every 15s.
- On app background: close socket; rely on OS push hints; sync on foreground.
Appendix E: Observability#
E.1 Core Metrics#
gateway.connections_active{host,cell} gateway.new_connections_per_sec
gateway.admission_rejected_total gateway.heartbeat_timeouts_total{asn}
gateway.conn_send_buffer_bytes (p99) gateway.event_loop_lag_ms
router.owner_cpu{owner} router.channel_events_per_sec (top-K)
push_to_visible_ms (p50/p99, client-measured)
client.gap_detected_total client.gap_filled_total
sync.rps / sync.p99 sync_per_reconnect{client_version}
E.2 Critical Alerts#
| Alert | Threshold | Action |
|---|---|---|
| Mass disconnect | connections_active fleet drop > 3% in 1 min | Page; halt deploys |
| Admission saturation | rejected > 20% of attempts for 2 min | Page; add capacity, check storm source |
| Unfilled gaps | gap_detected − gap_filled > 0.01% of events | Page client platform |
| Collateral latency | push_to_visible_ms_p99 > 1s on probe channels | Page; look for hot channel |
| Sync overload | sync.p99 > 1s | Page sync owner; enable edge cache/rate-limit |
E.3 Debugging the Silent Failure#
Server metrics can't see missed deliveries. Run a synthetic client fleet (~1K clients across regions and networks) that publishes and subscribes to probe channels, measures publish-to-visible latency, and verifies no gaps — including through forced disconnects.
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 100K concurrent | One or two socket servers, Redis pub/sub, or a managed service | Deploys start disconnecting users noticeably |
| 100K–5M | Thin gateways, seq + sync, drain-based deploys | Broadcast fan-out cost; hot channels |
| 5M–50M | Channel owners, hot-channel promotion, cells | Multi-product governance, multi-region |
| 50M+ | Multi-region cellular platform, topic registry, chargeback | Org coordination; long-lived client versions |
What You Don't Build on Day One#
- Relay tiers (until a channel exceeds ~10K online members)
- Multi-region sequencing
- Per-connection durable queues (ever, usually)
- Custom binary protocol (negotiate later)
Appendix G: Multi-Tenancy and Cost#
- Chargeback units: connection-minutes (gateway memory) and deliveries (CPU + egress) per topic owner.
- Quotas: per-topic publish rate; per-tenant maximum channel size for open posting; payload cap 16 KB.
- Isolation: broadcast-class channels on their own owner pool; large tenants spread across cells.
- Tradeoff summary: thinner gateways cost a hop (~1–5ms) and buy freedom to kill hosts; smaller hosts cost machines and buy smaller storms; best-effort push costs a sync API and buys stateless gateways. Every one of those trades is worth it above ~1M concurrent.