Technologies that implement this pattern: Redis · Apache Kafka · API Gateways · Cassandra · ZooKeeper & etcd
Why This Matters#
Every product eventually wants to feel alive: the driver's car moving on the map, the "typing…" indicator, the price ticking, the comment appearing without a refresh, the build turning green. HTTP was designed for the client to ask; real-time means the server needs to tell. That inversion is the whole difficulty. The moment the server must push, it has to know which connection belongs to which user, on which machine that connection lives right now, and what the client missed while it was reconnecting through a tunnel.
Most candidates treat real-time as a protocol question: "use WebSockets." Staff engineers treat it as a connection-state ownership and fan-out question. The protocol is the easy 10%. The hard 90% is: a registry of where 5 million connections live that changes thousands of times per second; a routing layer that gets an event from the service that produced it to the one gateway holding the right socket; a fan-out plan for the event that must reach 2 million viewers at once; and a resume protocol so that a dropped connection doesn't silently lose messages.
The second reframe: real-time is a promise about latency and completeness that you have to keep during deploys. A gateway fleet holding 4M long-lived connections cannot be restarted the way a stateless API can. Every deploy, autoscale event and load-balancer timeout is a reconnect storm someone has to design for. The steady state is easy; the transitions are the system.
If you can walk an interviewer from "how fresh does this need to be and does it really need push" to "where connection state lives and how events are routed and fanned out" to "what happens on reconnect, on deploy, and on the event with 2M subscribers", you are answering at Staff level. For the full end-to-end design problem, see the Real-Time Updates case study; this page is the reusable pattern.
The 60-Second Version#
- Ask whether you need push at all. If 5–30 seconds of delay is acceptable and the client is in the foreground, polling with ETags/
If-None-Matchis simpler, cacheable and stateless. 1M clients polling every 10s is 100K requests/sec — mostly cheap 304s. - SSE for server→client streams, WebSockets for bidirectional. SSE is plain HTTP, auto-reconnects with
Last-Event-ID, and passes through proxies; use it for feeds, notifications, dashboards. WebSockets earn their cost when the client sends frequently (chat typing, collaborative cursors, games). - A connection is ~10–50 KB of memory and a heartbeat. A tuned gateway node holds ~100K–500K mostly-idle connections; heartbeat every 25–30s to stay under typical 60s load-balancer and NAT idle timeouts.
- Separate the gateway from the logic. Gateways hold sockets and nothing else; a connection registry (user → gateway) and a pub/sub bus route events. That lets you deploy business logic without dropping connections.
- Fan-out cost is subscribers × events. One chat message to a 10-member room is 10 deliveries; one goal in a match watched by 2M is 2M deliveries in < 1s. Above ~10K subscribers per topic, switch from per-user routing to topic-level broadcast at the gateway tier.
- Design reconnect as a first-class path. Every client reconnects with a cursor (last sequence ID) and the server replays from a short retention buffer (minutes) or falls back to a full fetch. Reconnects use jittered exponential backoff (1s → 30s cap), or a deploy becomes a self-inflicted DDoS.
The Problem#
A delivery app shows 400K couriers moving on maps to 1.2M customers; each courier sends a location every 4 seconds and each customer should see their courier move within ~1 second. A collaboration tool has 3M connected users, each in a handful of channels, some channels with 50K members. A sports app pushes score updates to 2M people watching the same match, and every one of them expects to see the goal before their neighbor's TV does. In all three, a deploy of the gateway fleet disconnects hundreds of thousands of clients who all try to reconnect in the same second, a user on a train drops and resumes four times in ten minutes, and the system must decide what each of them missed. "Open a WebSocket" covers none of that.
Case Studies That Use This Pattern#
- Real-Time Updates — The flagship design problem: gateway fleets, connection registry, pub/sub routing, reconnect storms
- Chat Messaging — Per-conversation ordering, delivery receipts, presence, and offline fallback to push notifications
- Collaborative Editing — Sub-100ms bidirectional updates where the payload is operations that must be ordered and merged
- Ride Hailing & Delivery — High-frequency location streams; latest-value semantics where old updates are worthless
- Notification System — The offline half: when no connection exists, fall back to APNs/FCM/email
- Stock Exchange — Market data fan-out where fairness and ordering of updates matter
- Leaderboard — Live rank changes pushed to many viewers; throttled updates instead of every change
- News Feed — "New posts available" banners: the case for a cheap signal plus a pull, not pushing content
The Four Intents#
"Make it real-time" covers four goals with incompatible designs.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Awareness ("something changed": new comments, build finished, notification badge) | 1–30s acceptable; content fetched separately | Polling, long-poll, or SSE carrying a small signal; client pulls details | Thundering herd of pulls after a broadcast signal | At-least-once signal; content is fetched fresh |
| Conversation (chat, comments, support) | p99 < 500ms; ordered per conversation; no loss | WebSocket/SSE + per-conversation sequence + durable store + resume cursor | Gaps after reconnect; duplicate delivery | Every message delivered once (dedupe by ID), in order |
| Live state (location, presence, prices, cursors) | p99 < 250ms–1s; only latest value matters | Latest-value streams, conflation, drop stale updates under pressure | Slow consumers queue stale data; memory blow-up | Eventually latest; intermediate values may be skipped |
| Broadcast (sports scores, live events, market data) | 100K–10M recipients per event within ~1s | Topic broadcast at gateway tier, edge fan-out, conflation | Fan-out spike saturates egress; uneven delivery latency | Every subscriber sees latest within SLO |
🎯 Staff Move: "I'll split this into two intents: courier location is live state — only the latest point matters, so I'll conflate and drop under pressure — and the order-status changes are conversation-grade: ordered, never lost, replayable on reconnect. Same socket, two very different delivery guarantees."
The Core Tradeoff#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Short polling (every N seconds, ETag/304) | Stateless, cacheable, trivial to scale and debug | Latency = interval/2 on average; wasted requests when nothing changes | Backend (N×clients/interval QPS); battery on mobile |
| Long polling (request held until data or ~30s timeout) | Near-real-time over plain HTTP; works everywhere | Holds a request slot per client; reconnect after every message; ordering across requests is on you | App servers' concurrency; on-call debugging timeouts |
| Server-Sent Events (one-way stream over HTTP) | Simple, proxy-friendly, built-in reconnect with Last-Event-ID | Server→client only; HTTP/1.1 caps ~6 connections per origin in browsers; text-only | Clients that need to send frequently (extra HTTP calls) |
| WebSockets | Full duplex, low overhead per message, binary frames | Stateful connections; LB/proxy config; custom heartbeat, auth refresh and resume protocol | Platform team running a stateful gateway fleet; deploy complexity |
| Mobile push (APNs/FCM) | Reaches backgrounded/offline devices | Best-effort, rate-limited, seconds of delay, no ordering | Product (can't rely on it for completeness) |
| Managed real-time service | Offloads connection fleet and global edge | Per-message/connection pricing at scale; lock-in on client SDK | Finance at scale; platform when migrating off |
Staff Default Position#
Use the least stateful mechanism that meets the freshness budget, and when you do hold connections, keep them in a dumb gateway tier fed by a routing layer — with resume-by-cursor as part of the protocol from day one.
Start with polling for anything that tolerates ≥ 10s. Use SSE for server→client streams (notifications, feeds, dashboards, order status). Use WebSockets when clients send frequently or need sub-100ms round trips. In either push case: gateways hold sockets, authenticate, heartbeat and forward — no business logic; a connection registry maps user/device → gateway with a TTL refreshed by heartbeats; producers publish to a pub/sub bus keyed by user or topic; every message carries a per-stream sequence number, the server keeps a short replay buffer (2–10 minutes), and clients resume from their last sequence or fall back to a full fetch. Mobile push is a parallel path for offline users, never the source of truth.
When to Deviate#
- Very large broadcast audiences — Above ~100K subscribers to one topic, stop routing per user. Subscribe gateways (not users) to the topic, conflate updates (send the latest score every 500ms, not every change), and consider edge/CDN-based streaming (HLS-style or SSE through an edge that supports it).
- Strict bidirectional latency (collaborative editing, multiplayer) — Pin a session to a single server that owns the document/room state (sticky routing by room ID), rather than routing every operation through a shared bus. The bus hop adds 1–5ms and reorders under load.
- Low connection count, internal tools — 500 users on a dashboard? SSE from the app server itself, no separate gateway tier, no registry. Add the architecture when you need it.
- Offline-first mobile — Phones in the background will not hold your socket. The real delivery mechanism is push notification + sync-on-open; the socket is an optimization while the app is foregrounded.
The L5 → L6 → L7 Contrast#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Use WebSockets" | "What's the freshness budget per update type, who sends, and how many recipients per event?" | "How many teams are building socket fleets, and should real-time delivery be one platform with a published contract?" |
| Architecture | App servers hold sockets and push directly | Dumb gateway tier + connection registry + pub/sub routing; logic deploys without dropping sockets | One edge connection per device multiplexing every product's streams; platform owns gateways, products own topics |
| Completeness | "WebSockets are reliable (TCP)" | Per-stream sequence numbers, replay buffer, resume cursor, dedupe by message ID | Delivery semantics (latest-value vs ordered-complete) as a declared per-topic contract with SLOs |
| Failure | "Client reconnects" | Jittered backoff, connection draining on deploy, admission control on reconnect, load-shed by priority | Plans for the regional gateway loss: reconnect capacity reserved, game-days for 30% of connections dropping at once |
| Cost | Not discussed | "500K idle conns per node, 8 nodes per region" | Prices connection-hours and egress across products; build vs managed service decision with a 3-year TCO |
| Ownership | Feature team | Platform owns gateway/registry; feature team owns events and payloads | Defines the platform/product contract, quotas per product, and who can broadcast to 10M devices |
Why "Completeness" separates levels
TCP guarantees order and delivery within one connection. Real users don't have one connection: they have a sequence of connections across Wi-Fi, LTE and tunnels, and messages sent between the old socket dying and the new one opening go nowhere. The L5 answer is not wrong about TCP; it's wrong about the unit. The Staff answer puts sequence numbers and resume in the application protocol. The Principal answer makes the guarantee a declared property of each stream so product teams know whether "missed" is possible.
Why "Failure" separates levels
"Client reconnects" is true and dangerous: 400K clients reconnecting in the same second after a gateway deploy is a self-inflicted DDoS on auth, the registry and the replay store. Staff designs the transition — drain slowly, jitter, admission-control the reconnect path. Principal treats mass reconnect as a capacity requirement that is funded and rehearsed.
Why "Ownership" separates levels
Three product teams building three socket fleets means three heartbeat bugs, three reconnect storms and a phone holding three sockets that drain its battery. Staff makes one product's real-time correct. Principal notices the duplication and makes the device hold one connection that every product multiplexes over — with quotas so one team's broadcast can't starve another's chat.
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Pull vs Push | Stateless simplicity vs freshness and efficiency at low change rates |
| 2 | Signal vs Payload | Push "something changed" and let clients fetch, or push the data itself |
| 3 | Per-User Routing vs Topic Broadcast | Precise delivery vs fan-out cost at large audiences |
| 4 | Complete vs Latest | Every update in order vs only the newest value |
| 5 | Sticky State vs Stateless Gateways | Room state on one server vs any gateway can serve any client |
Fault Line 1: Pull vs Push#
Polling cost is clients / interval, regardless of how often data changes. Push cost is connections held + events × recipients. For a dashboard that changes once a minute viewed by 50K people, polling every 10s is 5K req/s of mostly 304s — trivial behind a CDN or cache. For a chat app where each user gets 2 messages a minute but needs them within 500ms, polling at 0.5s intervals across 3M users would be 6M req/s — absurd; push wins by three orders of magnitude. Who pays: polling — backend QPS, mobile battery, and users waiting on average interval/2; push — the platform team running stateful infrastructure and the on-call during deploys. Staff default: poll when freshness ≥ 10s or change rate is high relative to interval; push when freshness < ~5s and changes are sparse per user. Deviate when: clients are already holding a socket for another reason — then piggyback, the marginal push is nearly free.
Fault Line 2: Signal vs Payload#
Pushing the full payload (the whole new comment, the full order object) saves a round trip but couples the push channel to every schema and every authorization rule. Pushing a signal ("conversation 812 has new messages up to seq 4410") keeps the channel thin; the client fetches through normal, authorized, cacheable APIs. Who pays: payload — the platform (bigger messages, schema coupling, authorization re-checked at push time) and privacy (a payload pushed after access was revoked); signal — latency (+1 round trip, 30–100ms) and a fetch stampede when one signal goes to 100K clients. Staff default: payload for small, high-frequency, conversation-grade updates (chat messages, location); signal for large or permission-sensitive objects, with jittered fetch delay (0–2s) when the audience is large. Deviate when: broadcast scale makes the fetch stampede itself the outage — push the payload (it's identical for everyone) and serve it from the gateway.
Fault Line 3: Per-User Routing vs Topic Broadcast#
Per-user routing looks up each recipient in the registry and sends to their gateway: precise, and the cost scales with recipients. For a 20-member channel that's 20 lookups — fine. For a 2M-viewer match, it's 2M registry lookups and 2M bus messages per goal. Topic broadcast inverts it: each gateway subscribes to the topics its connected clients care about; a goal is one publish to ~200 gateways, which each fan out locally to their ~10K subscribers from memory. Who pays: per-user routing at scale — the registry and bus (millions of operations per event); topic broadcast — gateway memory for subscription tables and the loss of per-user filtering (everyone on the topic gets the same thing). Staff default: per-user routing for small audiences (DMs, small groups); gateway-level topic subscriptions above ~1K–10K subscribers per topic. Deviate when: each recipient needs a personalized payload (e.g., per-user permission redaction) — then precompute variants per audience segment rather than per user.
Fault Line 4: Complete vs Latest#
Chat needs every message, in order: a skipped message is a bug. A courier's location needs only the newest point: delivering 12 queued stale positions to a slow phone is worse than delivering one fresh one. Treating everything as complete-and-ordered makes slow consumers accumulate unbounded buffers on gateways; treating everything as latest-only loses messages. Who pays: complete — gateway memory and the slow client's latency (they get old data first); latest — the user who misses an intermediate state (acceptable for location, not for "order cancelled"). Staff default: declare semantics per stream type. Complete streams get sequence numbers, a bounded per-connection buffer (e.g., 1,000 messages or 1 MB), and on overflow the gateway closes the connection and the client resumes from its cursor. Latest streams get conflation: the gateway keeps one slot per key and overwrites it, sending at most every N ms. Deviate when: a "latest" stream carries a state transition someone acts on — split that transition into a complete stream.
Fault Line 5: Sticky State vs Stateless Gateways#
If any gateway can serve any client, deploys and failures are easy: drain, reconnect anywhere. But a collaborative document or a game room needs its operations ordered by one authority, ideally co-located with the connections. Who pays: stateless gateways + shared bus — latency (an extra hop, 1–5ms) and ordering complexity; sticky rooms — failover (room state must move or be rebuilt, 1–10s) and hot rooms (one server per room caps room size). Staff default: stateless gateways for notifications, chat, feeds and location; a separate stateful "room server" tier, addressed by consistent hashing on room ID, only for workloads that need a single ordering authority. Deviate when: room size is small and latency is critical (multiplayer games) — co-locate connections and room state on the same process.
Common Interview Mistakes#
| What Candidates Say | What Interviewers Hear | What Staff Engineers Say |
|---|---|---|
| "We'll use WebSockets" (first sentence) | "Picked a protocol before the requirements" | "Order status needs ~1s and is server→client only, so SSE. The chat composer needs bidirectional, so that screen uses a WebSocket." |
| "TCP guarantees delivery" | "Hasn't handled mobile reconnects" | "Within one connection. Across reconnects, every message carries a sequence and the client resumes from its cursor." |
| "Each server pushes to its connected users" | "Doesn't know how events find the right server" | "Gateways only hold sockets. A registry maps user to gateway; producers publish through a bus keyed by user or topic." |
| "Push the event to all 2M viewers" | "Per-user fan-out at broadcast scale" | "Gateways subscribe to the match topic; one publish, ~200 gateway deliveries, local fan-out — conflated to every 500ms." |
| "Clients reconnect on disconnect" | "Will build a reconnect storm" | "Jittered exponential backoff from 1s to 30s, drain gateways over 10 minutes on deploy, and admission control on the handshake path." |
| "Keep a message queue per user" | "Unbounded buffers for slow clients" | "Bounded per-connection buffer; overflow closes the socket and the client resumes from the durable store. Location is conflated, not queued." |
Quick Reference#
Staff Sentence Templates#
"Before picking a protocol: [update type] needs [freshness], flows [one-way / both ways], and reaches [N] recipients per event. That makes it [polling / SSE / WebSocket] with [per-user routing / topic broadcast]."
"TCP only covers one connection. Every [message] carries a per-[conversation] sequence; the server keeps [N minutes] of replay; a client that reconnects with a cursor older than that does a full fetch."
"Gateways hold sockets and nothing else, so we can deploy [business logic] without disconnecting anyone. Gateway deploys drain over [10 minutes] with [jittered] reconnects, which keeps handshake rate under [X/s]."
"[Location / prices] are latest-value streams, so the gateway conflates per key and sends at most every [N ms]. [Order status] is a complete stream — never conflated, always replayable."
Implementation Deep Dive#
1. Polling Done Right — ETag + Server-Suggested Interval#
Polling is underrated. Done well, it's stateless, cacheable and debuggable with curl.
GET /v1/orders/8812/status
If-None-Match: "v41"
# Server
function getStatus(order_id, etag):
v = cache.get("order:ver:" + order_id) # tiny version key, ~0.5ms
if etag == quote(v):
return 304, headers={"Retry-After-Ms": interval_for(order_id)}
body = loadStatus(order_id)
return 200, body, headers={"ETag": quote(v), "Retry-After-Ms": interval_for(order_id)}
function interval_for(order_id):
# adaptive: fast while the courier is near, slow otherwise
return state(order_id) == "arriving" ? 2000 : 15000
Numbers: 1M active orders × one poll per 15s = ~67K req/s, >90% of them 304s costing one Redis GET each. The server-suggested interval lets you slow everyone down during an incident with a config change — a load-shedding lever push systems don't get for free.
2. SSE with Resume — Last-Event-ID and a Replay Buffer (Redis Streams)#
GET /v1/stream?topics=orders,notifications
Accept: text/event-stream
Last-Event-ID: 1730301234567-3 # sent automatically by EventSource on reconnect
# Gateway
function stream(user, last_id):
if last_id:
missed = redis.XRANGE("user:" + user + ":events", "(" + last_id, "+", COUNT=500)
if missed is empty and last_id < redis.XINFO_first_id(...):
send("event: resync\ndata: {}\n\n") # cursor too old: client refetches state
for e in missed: send_event(e)
subscribe(user) # registry: user -> this gateway, TTL 90s
every 25s: send(": ping\n\n") # comment line, keeps LB/NAT alive
# Producer side
function publish(user, event):
id = redis.XADD("user:" + user + ":events", MAXLEN="~", 1000, "*", event) # bounded replay
gw = registry.get(user)
if gw: bus.publish("gw:" + gw, {user, id, event})
else: push_notifications.maybe_send(user, event) # offline path
Why it matters: the replay buffer (here the last ~1,000 events per user, or a few minutes) closes the gap between the old socket dying and the new one opening — typically 1–30s on mobile. The resync event handles the rarer case of a long absence by falling back to a full state fetch instead of pretending nothing was missed.
🎯 Staff Insight: Notice the push path writes to the durable buffer before routing to a gateway. Delivery is an optimization on top of the log; the log is the truth. That ordering is what makes reconnects correct rather than lucky.
3. Gateway Tier with Connection Registry — WebSockets + Redis Pub/Sub#
# On connect (after auth via short-lived token)
function onOpen(conn, user, device):
conns[user + ":" + device] = conn
registry.SET("conn:" + user + ":" + device, gateway_id, EX=90) # refreshed on heartbeat
for topic in conn.subscriptions:
if local_subs[topic].size == 0: bus.SUBSCRIBE("topic:" + topic)
local_subs[topic].add(conn)
# Heartbeat: client pings every 25s; gateway closes after 2 missed (60s)
function onPing(conn):
registry.EXPIRE("conn:" + conn.key, 90)
# Per-gateway inbound channel for user-targeted messages
bus.SUBSCRIBE("gw:" + gateway_id)
on bus message (target, payload):
c = conns[target]
if c == null: return # moved; producer will retry via log/resume
if c.buffer_bytes > 1MB: c.close(4008, "slow consumer") # client resumes from cursor
else: c.send(payload)
# Graceful drain on deploy
function drain():
stop_accepting()
for c in conns.values() in batches of 1% every 6s: # ~10 minutes total
c.close(4000, "reconnect", retry_after_ms=random(0, 5000))
Numbers: at ~20 KB per connection (buffers + TLS + bookkeeping), 300K connections ≈ 6 GB of memory per gateway node. Registry writes are dominated by heartbeats: 3M connections / 25s ≈ 120K EXPIRE/s — size the registry for heartbeats, not messages. Draining 1% per 6 seconds keeps a 300K-connection node's reconnect rate around 500/s instead of 300K in one second.
4. Broadcast with Conflation — Gateway-Level Topic Fan-out#
# Upstream: score service publishes every change
bus.PUBLISH("topic:match:991", {seq: 88, score: "2-1", clock: "67:12"})
# Gateway: conflate per topic, flush on a timer
latest = {} # topic -> newest message
on bus message (topic, msg):
if msg.seq > latest[topic].seq: latest[topic] = msg
every 250ms:
for (topic, msg) in latest.dirty():
frame = encode_once(msg) # serialize once, reuse bytes
for c in local_subs[topic]:
c.send_nonblocking(frame) # drop for this client if its buffer is full
latest.mark_clean(topic)
Why it matters: with 200 gateways and 2M subscribers, a goal costs 1 publish + 200 bus deliveries + 2M local socket writes spread across 200 nodes (~10K each, a few ms of CPU). Encoding once per gateway rather than per socket matters: at 2M recipients, per-recipient JSON encoding is the difference between 4ms and 400ms per update. Conflation means a flurry of 10 changes in 250ms costs one frame per client.
Technique Comparison
| Technique | Latency | Server State | Completeness | Operational Burden |
|---|---|---|---|---|
| Short polling (ETag) | interval/2 avg | None | Latest on each poll | Low |
| Long polling | ~RTT after change | Held request per client | Needs cursor per request | Medium |
| SSE + replay buffer | ~RTT | Connection + registry | Complete via Last-Event-ID | Medium |
| WebSocket gateway | ~RTT, bidirectional | Connection + registry + subs | Complete via app-level seq | High |
| Topic broadcast + conflation | ≤ conflation window | Subscription tables | Latest only | Medium–High |
| Mobile push | Seconds, best-effort | Provider-side | None guaranteed | Low (but rate-limited) |
Architecture Diagram#
How to narrate it: the numbered arrows are the contract. A producer appends to the durable per-user log first, then looks up where the user is connected, then publishes. If the user isn't connected, the offline path takes over. Gateways only hold sockets, refresh their registry entries via heartbeats, and replay from the log on resume. That separation is why you can deploy every service on the right without dropping a single connection on the left.
The sequence shows why the log write comes first: the message produced during the reconnect gap was "lost" by routing and recovered by resume.
Failure Scenarios#
1. Gateway Deploy Reconnect Storm — 1.8M Clients in 40 Seconds#
A config change triggers a restart of all 6 gateway nodes in one region within a minute. Clients use a fixed 1-second retry with no jitter.
t=0 Rolling restart begins; orchestrator max-unavailable set to 50 percent.
t=+10s 900K sockets closed. Clients retry at exactly +1s, +2s, +3s.
t=+12s Handshake rate 450K/s. TLS termination CPU 100 percent at the LB.
t=+15s Auth service (token validation on connect) p99 4s; 30 percent of handshakes time out.
t=+20s Timed-out clients retry on 1s cadence; offered handshakes stay above 400K/s.
t=+40s Remaining 900K sockets closed by second half of restart. Storm doubles.
t=+3min Registry Redis at 100 percent CPU from SET + EXPIRE on connect.
t=+11min On-call scales auth 4x and blocks new handshakes at 50K/s via LB rate limit.
t=+24min All clients reconnected. Missed events replayed from log for 92 percent;
8 percent exceeded replay window and did full resync, spiking API reads.
Detection: gateway.handshakes_per_sec > 5× baseline; gateway.connected drop > 20% in 1 min; auth.p99 on connect; registry.cpu; resume.resync_rate.
Blast radius: every real-time user in the region, plus auth and API services that weren't part of the change.
Mitigation: handshake admission control at the LB (token bucket), temporarily extending the server-suggested retry delay.
Prevention: gateway restarts limited to one node at a time with 10-minute drain; client SDK with jittered exponential backoff (1s → 30s, full jitter); auth on connect uses locally verifiable tokens (JWT with short TTL) instead of a remote call; replay window sized to cover a 30-minute drain.
Owner: real-time platform team (gateway, SDK, drain tooling); identity team for connect-path auth capacity.
🎯 Staff Insight: The incident spread to auth and the API tier — services nobody touched. In stateful connection systems, the steady state is never the risk; the transition is. Design and load-test the reconnect path at 100% of connections.
2. Silent Gap — Messages Lost Between Sockets#
A chat product pushes messages directly from the bus to gateways with no per-conversation sequence and no replay. Mobile clients reconnect frequently on cellular handoffs.
t=0 User on a train; LTE handoff every 3-5 minutes. Each drop lasts 2-8s.
t=+2s Message from a teammate routed to old gateway; socket already gone. Dropped.
t=+6s Client reconnects to a new gateway. Shows "connected". No gap detected.
t=+1h User replies to a thread missing two messages. Confusion.
t=+2wk Support volume "missing messages" up 3x after a mobile release that
reduced background keepalive. No metric fired at any point.
Detection: none — the gap. Add delivery.gap_detected (client reports when it receives seq N+2 after N), delivery.undeliverable at gateways for targets not connected, and a periodic client-side reconciliation that compares last-seen seq with the server.
Blast radius: mobile users disproportionately — the most engaged ones.
Mitigation: on app foreground, fetch messages since last seen seq.
Prevention: per-conversation sequence numbers, write-to-log-before-route, resume by cursor on every reconnect, and client-side gap detection that triggers a fetch.
Owner: messaging team owns sequence and resume semantics; platform SDK owns gap detection and metrics.
3. Slow Consumer Memory Blow-up — One Topic Takes Down a Gateway#
A live-trading dashboard subscribes 40K users to a price topic updating 50 times per second. Outbound frames are queued per connection without bound. A subset of clients are on poor networks.
t=0 Market opens. 50 updates/s x 40K subscribers spread over 8 gateways.
t=+2min 3 percent of clients cannot keep up; their outbound queues grow 50 frames/s.
t=+10min Gateway 4 heap from 4 GB to 22 GB; GC pauses 2-6s.
t=+12min GC pauses miss heartbeats; LB marks gateway 4 unhealthy; 300K sockets drop.
t=+13min Those sockets reconnect to gateways 1-3 and 5-8, bringing slow clients along.
t=+20min Two more gateways enter the same spiral.
Detection: gateway.outbound_buffer_bytes p99 per connection; gateway.heap_used; gateway.gc_pause_seconds; conn.closed_slow_consumer count.
Blast radius: every user on affected gateways, including those on unrelated topics.
Mitigation: emergency config: conflate the price topic at 4 updates/s; cap per-connection buffers at 256 KB.
Prevention: declare price streams as latest-value with conflation; bounded per-connection buffers with close-and-resume on overflow; per-topic rate limits enforced at publish.
Owner: platform owns buffer limits and conflation; trading product owns the topic's declared semantics.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Reconnect storm | gateway.handshakes_per_sec > 5× | Region's real-time users + auth | LB handshake limiter, longer retry hint | Real-time platform |
| Messages lost across reconnect | delivery.gap_detected, delivery.undeliverable | Mobile users | Resume by cursor, foreground sync | Messaging team + SDK |
| Slow consumers | gateway.outbound_buffer_bytes p99 | Whole gateway node | Conflate, cap buffers, close-and-resume | Platform + topic owner |
| Stale registry entries | delivery.misrouted_rate | Delayed delivery | Shorter TTL, gateway-side redirect | Real-time platform |
| Hot broadcast topic | bus.topic_publish_rate, egress Gbps | Bus + gateway egress | Conflation, topic rate limit | Topic owner |
| LB idle timeout kills sockets | Spikes of closes at exact intervals (e.g., 60s) | All idle clients | Heartbeat < idle timeout | Platform + network team |
| Push provider throttling | push.rejected by APNs/FCM | Offline users | Collapse keys, priority classes | Notifications team |
The Principal Lens#
Why L7 Sees This Problem Differently#
A Staff engineer builds a correct real-time path for one product. A Principal engineer notices that the mobile app holds three sockets — chat, live location and notifications — built by three teams on three gateway fleets, each with its own heartbeat interval, each draining the battery, each reconnecting independently after a tunnel. The org-level problem is one device, one connection, many products: a shared real-time platform that multiplexes every product's streams over a single connection, with declared delivery semantics per stream, quotas per product, and one team that is excellent at the hardest operational part — moving millions of connections safely during deploys and failures.
The Org-Level Fault Line#
One multiplexed real-time platform vs per-product socket fleets.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Per-product fleets | Autonomy; product-specific protocol tuning | N sockets per device, N reconnect storms, N heartbeat bugs; battery and data cost | Users' batteries; N on-call rotations |
| Managed real-time vendor | No fleet to run; global edge on day one | Per-connection/message pricing grows linearly; SDK lock-in; limited control over resume semantics | Finance at scale; platform during migration |
| Shared platform: one connection per device, topic multiplexing, per-stream contracts | One expert team, consistent resume and backoff, one socket per device | Platform becomes a dependency for every real-time feature; noisy-neighbor risk across products | Platform headcount (4–6); quota governance |
The Principal default is the shared platform once three or more products need push: the platform owns connections, registry, routing, SDK and drain; product teams own topics, payload schemas and their declared semantics (complete vs latest, freshness SLO), within quotas.
Cost Model#
Assumptions: ~20 KB per connection; gateway nodes at ~$600/month holding ~300K connections at < 60% memory; bus/registry clusters ~$2–4K/month per cell; egress ~$0.05/GB; managed-vendor list pricing commonly lands in the range of ~$0.5–2 per thousand peak connections per month plus per-message fees; engineer ~$25K/month.
| Scale | Peak Connections | Self-hosted Infra | Managed Alternative (order of magnitude) | People / On-call | Rough Monthly Total (self-hosted) |
|---|---|---|---|---|---|
| Small | 20K | 2 gateways + small Redis (~$2K) | ~$1–3K | 0.3 FTE; product team on-call | ~$2K + ~$8K people — managed usually wins |
| Growth | 1M | 8 gateways/region × 2 regions, registry + bus cells (~$20K), egress ~$10K | ~$30–60K + messages | 2 FTE; shared rotation | ~$30K + ~$50K people |
| Large | 20M | 150+ gateways across regions, bus cells, replay store (~$180K), egress ~$120K | ~$400K–1M+ | 5–6 engineer platform team; dedicated rotation | ~$300K + ~$150K people |
The Principal observation: below ~100K connections, a managed service is cheaper than the people needed to run gateways well. Above a few million, the vendor bill grows linearly while the self-hosted cost is dominated by a fixed team. The crossover is a decision to make deliberately — and the one-way part is the client SDK, so wrap it from day one.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Heartbeat interval, conflation window, replay length | Two-way | Config |
| SSE vs WebSocket for a stream | Two-way (behind an SDK) | Client release cycle |
| Client wire protocol and resume semantics shipped in mobile apps | One-way-ish | Old app versions live for 1–2 years; server must support both |
| Promising "real-time" completeness in a public/partner API | One-way | Partners build on it; weakening it is a breaking change |
| Vendor SDK called directly throughout app code | One-way-ish | Rip-and-replace across every screen; wrap it on day one |
| Topic naming and authorization model | One-way-ish | Every producer and subscriber depends on it |
The Standard I'd Write#
RFC: Real-Time Delivery Standard (v1)
Scope: Any server-initiated update to clients — web, mobile, partner integrations.
MUST:
- Each stream declares semantics: complete (ordered, sequence-numbered, replayable) or latest (conflated), and a freshness SLO (p99 delivery latency).
- Producers append complete-stream events to the durable per-recipient or per-topic log before routing; delivery is best-effort on top of the log.
- All clients use the platform SDK: jittered exponential backoff (1s → 30s, full jitter), resume by cursor, gap detection.
- Gateways hold no business logic; per-connection outbound buffers are bounded; overflow closes with a resume code.
- Gateway deploys drain at ≤ 1% of connections per 5 seconds per node; reconnect capacity is load-tested at 100% of peak connections twice a year.
SHOULD: Prefer polling for freshness ≥ 10s; use signal-then-fetch for large or permission-sensitive payloads; broadcast topics above 10K subscribers use gateway-level subscription and conflation.
Exceptions: Real-time platform review; stateful room servers (collaboration, games) require a design review for ordering and failover.
Success metrics:
delivery.gap_detected< 0.01% of messages; p99 delivery within declared SLO per stream; one connection per device for platform products; zero Sev-1s from gateway deploys.
What I'd Tell the VP#
Customers increasingly expect our app to update by itself — couriers moving, messages arriving, scores changing. Today three teams each built their own way of doing that, which means three times the infrastructure, three times the battery drain on phones, and three different ways it breaks when we deploy. I'm proposing a single real-time platform that every team uses, with one connection per device and a clear promise for each kind of update about how fast and how complete it is. It's roughly five engineers ongoing, offset by retiring two fleets and avoiding a vendor bill that would grow past $500K a month at our projected scale.
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| One device, one connection | "The phone shouldn't hold a socket per product. I'd multiplex every stream over one connection and give products topics and quotas." |
| Semantics as a contract | "Each stream declares complete or latest and a p99. That's what product teams sign up for, not 'WebSockets'." |
| Transitions funded | "Reconnect capacity for 100% of connections is a requirement we test twice a year, not a hope." |
| Build vs buy crossover | "Managed until about a million connections, behind our own SDK so the switch is a server change, not an app rewrite." |
| Blast radius | "Regional gateway cells, so a bad deploy is one cell's reconnect, not the world's." |
Staff answers that L7 interviewers find insufficient:
- "I'd build a gateway tier with a registry and Redis pub/sub." — Correct for this product; silent on the two other fleets doing the same thing.
- "Clients resume from a cursor." — Right mechanism, but not an org guarantee: what about the teams not using cursors?
- "We'll use a managed service." — No crossover analysis, no SDK wrapper, no exit path.
In the Wild#
Slack: Channel Servers and Gateway Servers#
Slack's engineering blog has described its real-time messaging architecture: stateful channel servers that own channels (mapped by consistent hashing) and hold recent history, gateway servers that hold client WebSocket connections and subscribe to the channels their users need, plus separate admin and presence services. Messages are persisted through the API tier and then fanned out from channel servers to the gateways subscribed to that channel.
Staff insight: This is topic-level routing in production: gateways subscribe to channels on behalf of their users, so a message to a big channel goes to the gateways that care rather than being routed per user. It also separates socket holding from channel state — the gateway/room-server split from Fault Line 5.
Discord: Elixir Gateways and Per-Guild Processes#
Discord has written publicly about building its real-time gateway on Elixir/the Erlang VM, with a process per connected session and a process per guild (server) that fans events out to the sessions in it, and about the work required to scale very large guilds — including relaying fan-out across multiple processes when a single guild's member count made one process the bottleneck.
Staff insight: A per-room ordering authority is elegant until one room is huge. The large-guild work is Fault Line 3 and Fault Line 5 colliding: sticky state gives ordering, and the hottest room needs its fan-out split anyway.
LinkedIn: Hundreds of Thousands of Persistent Connections per Machine#
LinkedIn's engineering blog described its instant messaging and presence infrastructure using Server-Sent Events over long-lived HTTP connections on the Play framework and Akka actors, with each connection represented as a lightweight actor, and reported sustaining hundreds of thousands of concurrent connections per machine after tuning.
Staff insight: A large, mature product chose SSE, not WebSockets, for server-to-client delivery — and the scale achieved per box was a function of connection memory and file-descriptor tuning, not protocol choice. It's a strong reference when an interviewer assumes "real-time means WebSockets".
Practice Drill#
Prompt: "We're building live order tracking for a food delivery app: 1.5M concurrent orders at dinner peak. Customers see the courier move and status changes (accepted, picked up, arriving). Couriers send GPS every 4 seconds. Design the real-time path."
Staff Answer
Two stream types with different semantics. Status is complete: ordered, never lost, replayable — append to a per-order event log (Redis Stream, MAXLEN ~200, TTL 6h), seq per order. Location is latest-value: only the newest point matters — conflate per order, send at most every 1s. Both flow server→client only, so SSE through a gateway tier; the courier app uploads GPS over plain HTTPS batched every 4s (375K req/s at peak, ingestion → Kafka keyed by courier). Routing: a location service consumes courier GPS, maps courier → active order(s), writes the latest point to order:{id}:loc and publishes to gw:{gateway} found via the registry (order:{id} → gateway, TTL 90s, refreshed by 25s heartbeats). Status changes append to the log first, then publish. Sizing: 1.5M customer connections at ~20 KB ≈ 30 GB; ~6–8 gateways per region at 250K each with headroom; location messages ~1.5M/4s ≈ 375K/s across the bus — split bus by order hash into 8 cells. Resume: Last-Event-ID replays status since the cursor; location just sends the current point. Offline: status changes also go to push notifications if no live connection. Deploy safety: 1% per 6s drain, SDK with full-jitter backoff, handshake limiter at the LB. Fallback: if the stream fails, the app polls /orders/{id} every 10s with ETag — a degraded mode product signed off on. Metrics: delivery.p99_ms per stream, gateway.connected, gateway.handshakes_per_sec, resume.resync_rate, location.staleness_seconds.
Why this is L6:
- Splits complete vs latest semantics and designs each separately rather than one queue for everything.
- Writes to the log before routing, so reconnect gaps are recovered, and sizes connections, registry and bus with numbers.
- Designs the transitions — drain, backoff, handshake admission — and a polling degraded mode.
What L7 adds:
- Notices chat-with-courier and promotions also want push and proposes one multiplexed connection with topic quotas instead of a second fleet.
- Prices self-hosted (~$30K/month + 2 FTE) against a managed vendor at 1.5M peak connections and wraps the client SDK either way.
- Sets a quarterly game-day: drop a region's gateways at dinner-peak load in staging and prove reconnect capacity.
Staff Interview Application#
How to Introduce This Pattern#
"Before I pick a protocol, I want to classify the updates: how fresh each needs to be, whether every update matters or only the latest, and how many people receive each one. Then the hard parts are where connections live, how events find them, and what a client missed while it was reconnecting — I'll spend most of my time there."
Lead with freshness and semantics, then the gateway/registry/bus split, then fan-out size, then reconnect and deploy behavior.
When NOT to Use This Pattern#
- Freshness ≥ 10–30s: Poll with ETags. It's cacheable, stateless and debuggable; push adds a stateful fleet for no user-visible gain.
- Background or offline delivery: Sockets don't survive backgrounding on mobile. Use push notifications + sync-on-open.
- Large payloads or downloads: Push a signal, fetch via normal APIs or a CDN; don't stream megabytes through gateways.
- Request/response semantics: If the client asks and waits for one answer, that's an HTTP call, not a stream.
Follow-Up Questions to Anticipate#
| Interviewer Asks | What They Are Testing | How to Respond |
|---|---|---|
| "How does the server know which machine a user is on?" | Routing | "Registry: user/device → gateway with a 90s TTL refreshed by heartbeats. Producers look it up and publish to that gateway's channel." |
| "What if the user disconnects and misses messages?" | Completeness | "Every message is in a log with a sequence. Reconnect sends the last seq; the gateway replays. Too old? Full resync." |
| "How do you deploy gateways?" | Transition design | "Drain 1% per 6s, server-suggested retry with jitter, handshake limiter. Logic lives elsewhere so most deploys don't touch gateways." |
| "A match has 2M viewers?" | Broadcast | "Gateways subscribe to the topic, one publish fans to ~200 gateways, local fan-out, conflated every 250–500ms." |
| "WebSockets or SSE?" | Pragmatism | "SSE for server→client — HTTP-native, auto-resume. WebSocket only where the client sends often." |
| "What about slow clients?" | Backpressure | "Bounded buffers per connection. Latest streams conflate; complete streams close and resume from the log." |
Evaluation Rubric#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | "WebSockets" | Freshness, direction, audience size, complete vs latest per stream | Real-time as a shared platform with per-stream contracts |
| Architecture | Servers push to their own sockets | Gateway + registry + bus; topic broadcast at scale | One connection per device, multiplexed, with quotas and cells |
| Completeness | TCP | Sequence, log-before-route, resume, gap detection | Declared semantics and gap SLO across all products |
| Failure | Client reconnects | Drain, jitter, handshake admission, bounded buffers | Funded reconnect capacity, regional cells, game-days |
| Cost | Not discussed | Connections per node, bus sizing | Build vs managed crossover, SDK as exit path |
Strong Hire Signals
| Signal | What It Sounds Like |
|---|---|
| Questions push itself | "Does this need push, or is a 10s poll fine?" |
| Gap-aware | "TCP covers one connection; users have many." |
| Transition-focused | "The deploy is the incident; I'll design the drain." |
| Fan-out math | "2M recipients is 200 gateway messages, not 2M bus messages." |
Lean No-Hire Signals
| Signal | Why It Misses the Bar |
|---|---|
| Business logic inside the socket servers | Every deploy disconnects everyone |
| No answer for missed messages on reconnect | Silent data loss in the most common mobile case |
| Per-user routing for million-viewer events | Registry and bus collapse at the moment that matters |
Common False Positives: Knowing the WebSocket handshake headers ≠ designing routing. Naming Redis pub/sub ≠ handling slow consumers. "Socket.IO handles it" ≠ a resume protocol.
Capacity Planning Quick Reference#
Sizing the Real-Time Path#
gateway_nodes = ceil(peak_connections / conns_per_node / 0.6) # 60% memory target
conn_memory = peak_connections × 20 KB # 10-50 KB typical
registry_ops = peak_connections / heartbeat_seconds # heartbeats dominate
poll_qps = active_clients / poll_interval_seconds
fanout_deliveries = events_per_sec × avg_recipients # per-user routing
broadcast_bus_msgs = events_per_sec × gateways_subscribed # topic routing
reconnect_rate = connections_per_node × drain_fraction / drain_interval
replay_buffer = recipients × events_per_min × replay_minutes × avg_event_bytes
Key Numbers Worth Memorizing#
| Number | Context |
|---|---|
| 10–50 KB | Memory per idle connection (TLS + buffers + bookkeeping) |
| 100–500K | Mostly idle connections per tuned gateway node |
| 25–30 s | Heartbeat interval; under typical 60s LB/NAT idle timeouts |
| ~6 | Browser HTTP/1.1 connections per origin — the SSE gotcha; HTTP/2 multiplexes |
| 1 s → 30 s | Jittered exponential reconnect backoff range |
| 2–10 min | Replay buffer that covers most mobile reconnect gaps |
| 1–8 s | Typical cellular handoff disconnection |
| 250–1000 ms | Conflation window for broadcast and location streams |
| ~10K | Subscribers per topic where per-user routing should give way to topic broadcast |
| 1% / 6 s | Drain rate that turns a deploy into a trickle, not a storm |
Common Pitfalls Checklist#
- Each stream declares complete vs latest and a freshness SLO
- Polling considered first for freshness ≥ 10s
- Gateways hold no business logic; logic deploys don't drop sockets
- Events written to a log before routing; clients resume by cursor
- Heartbeat interval below every LB/NAT idle timeout on the path
- Client SDK uses jittered exponential backoff
- Per-connection buffers bounded; slow consumers closed or conflated
- Broadcast topics use gateway-level subscription and conflation
- Offline path (push + sync-on-open) defined for backgrounded clients