Hiring BarSupport

Real-time Updates — Cross-Cutting Pattern

Pattern37 min read5 diagrams

Technologies that implement this pattern: Redis · Apache Kafka · API Gateways · Cassandra · ZooKeeper & etcd

Why This Matters#

Every product eventually wants to feel alive: the driver's car moving on the map, the "typing…" indicator, the price ticking, the comment appearing without a refresh, the build turning green. HTTP was designed for the client to ask; real-time means the server needs to tell. That inversion is the whole difficulty. The moment the server must push, it has to know which connection belongs to which user, on which machine that connection lives right now, and what the client missed while it was reconnecting through a tunnel.

Most candidates treat real-time as a protocol question: "use WebSockets." Staff engineers treat it as a connection-state ownership and fan-out question. The protocol is the easy 10%. The hard 90% is: a registry of where 5 million connections live that changes thousands of times per second; a routing layer that gets an event from the service that produced it to the one gateway holding the right socket; a fan-out plan for the event that must reach 2 million viewers at once; and a resume protocol so that a dropped connection doesn't silently lose messages.

The second reframe: real-time is a promise about latency and completeness that you have to keep during deploys. A gateway fleet holding 4M long-lived connections cannot be restarted the way a stateless API can. Every deploy, autoscale event and load-balancer timeout is a reconnect storm someone has to design for. The steady state is easy; the transitions are the system.

If you can walk an interviewer from "how fresh does this need to be and does it really need push" to "where connection state lives and how events are routed and fanned out" to "what happens on reconnect, on deploy, and on the event with 2M subscribers", you are answering at Staff level. For the full end-to-end design problem, see the Real-Time Updates case study; this page is the reusable pattern.

The 60-Second Version#

  • Ask whether you need push at all. If 5–30 seconds of delay is acceptable and the client is in the foreground, polling with ETags/If-None-Match is simpler, cacheable and stateless. 1M clients polling every 10s is 100K requests/sec — mostly cheap 304s.
  • SSE for server→client streams, WebSockets for bidirectional. SSE is plain HTTP, auto-reconnects with Last-Event-ID, and passes through proxies; use it for feeds, notifications, dashboards. WebSockets earn their cost when the client sends frequently (chat typing, collaborative cursors, games).
  • A connection is ~10–50 KB of memory and a heartbeat. A tuned gateway node holds ~100K–500K mostly-idle connections; heartbeat every 25–30s to stay under typical 60s load-balancer and NAT idle timeouts.
  • Separate the gateway from the logic. Gateways hold sockets and nothing else; a connection registry (user → gateway) and a pub/sub bus route events. That lets you deploy business logic without dropping connections.
  • Fan-out cost is subscribers × events. One chat message to a 10-member room is 10 deliveries; one goal in a match watched by 2M is 2M deliveries in < 1s. Above ~10K subscribers per topic, switch from per-user routing to topic-level broadcast at the gateway tier.
  • Design reconnect as a first-class path. Every client reconnects with a cursor (last sequence ID) and the server replays from a short retention buffer (minutes) or falls back to a full fetch. Reconnects use jittered exponential backoff (1s → 30s cap), or a deploy becomes a self-inflicted DDoS.

The Problem#

A delivery app shows 400K couriers moving on maps to 1.2M customers; each courier sends a location every 4 seconds and each customer should see their courier move within ~1 second. A collaboration tool has 3M connected users, each in a handful of channels, some channels with 50K members. A sports app pushes score updates to 2M people watching the same match, and every one of them expects to see the goal before their neighbor's TV does. In all three, a deploy of the gateway fleet disconnects hundreds of thousands of clients who all try to reconnect in the same second, a user on a train drops and resumes four times in ten minutes, and the system must decide what each of them missed. "Open a WebSocket" covers none of that.


Case Studies That Use This Pattern#

  • Real-Time Updates — The flagship design problem: gateway fleets, connection registry, pub/sub routing, reconnect storms
  • Chat Messaging — Per-conversation ordering, delivery receipts, presence, and offline fallback to push notifications
  • Collaborative Editing — Sub-100ms bidirectional updates where the payload is operations that must be ordered and merged
  • Ride Hailing & Delivery — High-frequency location streams; latest-value semantics where old updates are worthless
  • Notification System — The offline half: when no connection exists, fall back to APNs/FCM/email
  • Stock Exchange — Market data fan-out where fairness and ordering of updates matter
  • Leaderboard — Live rank changes pushed to many viewers; throttled updates instead of every change
  • News Feed — "New posts available" banners: the case for a cheap signal plus a pull, not pushing content

The Four Intents#

"Make it real-time" covers four goals with incompatible designs.

IntentConstraintStrategyFailure ModeCorrectness Bar
Awareness ("something changed": new comments, build finished, notification badge)1–30s acceptable; content fetched separatelyPolling, long-poll, or SSE carrying a small signal; client pulls detailsThundering herd of pulls after a broadcast signalAt-least-once signal; content is fetched fresh
Conversation (chat, comments, support)p99 < 500ms; ordered per conversation; no lossWebSocket/SSE + per-conversation sequence + durable store + resume cursorGaps after reconnect; duplicate deliveryEvery message delivered once (dedupe by ID), in order
Live state (location, presence, prices, cursors)p99 < 250ms–1s; only latest value mattersLatest-value streams, conflation, drop stale updates under pressureSlow consumers queue stale data; memory blow-upEventually latest; intermediate values may be skipped
Broadcast (sports scores, live events, market data)100K–10M recipients per event within ~1sTopic broadcast at gateway tier, edge fan-out, conflationFan-out spike saturates egress; uneven delivery latencyEvery subscriber sees latest within SLO

🎯 Staff Move: "I'll split this into two intents: courier location is live state — only the latest point matters, so I'll conflate and drop under pressure — and the order-status changes are conversation-grade: ordered, never lost, replayable on reconnect. Same socket, two very different delivery guarantees."


The Core Tradeoff#

StrategyWhat WorksWhat BreaksWho Pays
Short polling (every N seconds, ETag/304)Stateless, cacheable, trivial to scale and debugLatency = interval/2 on average; wasted requests when nothing changesBackend (N×clients/interval QPS); battery on mobile
Long polling (request held until data or ~30s timeout)Near-real-time over plain HTTP; works everywhereHolds a request slot per client; reconnect after every message; ordering across requests is on youApp servers' concurrency; on-call debugging timeouts
Server-Sent Events (one-way stream over HTTP)Simple, proxy-friendly, built-in reconnect with Last-Event-IDServer→client only; HTTP/1.1 caps ~6 connections per origin in browsers; text-onlyClients that need to send frequently (extra HTTP calls)
WebSocketsFull duplex, low overhead per message, binary framesStateful connections; LB/proxy config; custom heartbeat, auth refresh and resume protocolPlatform team running a stateful gateway fleet; deploy complexity
Mobile push (APNs/FCM)Reaches backgrounded/offline devicesBest-effort, rate-limited, seconds of delay, no orderingProduct (can't rely on it for completeness)
Managed real-time serviceOffloads connection fleet and global edgePer-message/connection pricing at scale; lock-in on client SDKFinance at scale; platform when migrating off

Staff Default Position#

Use the least stateful mechanism that meets the freshness budget, and when you do hold connections, keep them in a dumb gateway tier fed by a routing layer — with resume-by-cursor as part of the protocol from day one.

Start with polling for anything that tolerates ≥ 10s. Use SSE for server→client streams (notifications, feeds, dashboards, order status). Use WebSockets when clients send frequently or need sub-100ms round trips. In either push case: gateways hold sockets, authenticate, heartbeat and forward — no business logic; a connection registry maps user/device → gateway with a TTL refreshed by heartbeats; producers publish to a pub/sub bus keyed by user or topic; every message carries a per-stream sequence number, the server keeps a short replay buffer (2–10 minutes), and clients resume from their last sequence or fall back to a full fetch. Mobile push is a parallel path for offline users, never the source of truth.


When to Deviate#

  • Very large broadcast audiences — Above ~100K subscribers to one topic, stop routing per user. Subscribe gateways (not users) to the topic, conflate updates (send the latest score every 500ms, not every change), and consider edge/CDN-based streaming (HLS-style or SSE through an edge that supports it).
  • Strict bidirectional latency (collaborative editing, multiplayer) — Pin a session to a single server that owns the document/room state (sticky routing by room ID), rather than routing every operation through a shared bus. The bus hop adds 1–5ms and reorders under load.
  • Low connection count, internal tools — 500 users on a dashboard? SSE from the app server itself, no separate gateway tier, no registry. Add the architecture when you need it.
  • Offline-first mobile — Phones in the background will not hold your socket. The real delivery mechanism is push notification + sync-on-open; the socket is an optimization while the app is foregrounded.

The L5 → L6 → L7 Contrast#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Use WebSockets""What's the freshness budget per update type, who sends, and how many recipients per event?""How many teams are building socket fleets, and should real-time delivery be one platform with a published contract?"
ArchitectureApp servers hold sockets and push directlyDumb gateway tier + connection registry + pub/sub routing; logic deploys without dropping socketsOne edge connection per device multiplexing every product's streams; platform owns gateways, products own topics
Completeness"WebSockets are reliable (TCP)"Per-stream sequence numbers, replay buffer, resume cursor, dedupe by message IDDelivery semantics (latest-value vs ordered-complete) as a declared per-topic contract with SLOs
Failure"Client reconnects"Jittered backoff, connection draining on deploy, admission control on reconnect, load-shed by priorityPlans for the regional gateway loss: reconnect capacity reserved, game-days for 30% of connections dropping at once
CostNot discussed"500K idle conns per node, 8 nodes per region"Prices connection-hours and egress across products; build vs managed service decision with a 3-year TCO
OwnershipFeature teamPlatform owns gateway/registry; feature team owns events and payloadsDefines the platform/product contract, quotas per product, and who can broadcast to 10M devices
Why "Completeness" separates levels

TCP guarantees order and delivery within one connection. Real users don't have one connection: they have a sequence of connections across Wi-Fi, LTE and tunnels, and messages sent between the old socket dying and the new one opening go nowhere. The L5 answer is not wrong about TCP; it's wrong about the unit. The Staff answer puts sequence numbers and resume in the application protocol. The Principal answer makes the guarantee a declared property of each stream so product teams know whether "missed" is possible.

Why "Failure" separates levels

"Client reconnects" is true and dangerous: 400K clients reconnecting in the same second after a gateway deploy is a self-inflicted DDoS on auth, the registry and the replay store. Staff designs the transition — drain slowly, jitter, admission-control the reconnect path. Principal treats mass reconnect as a capacity requirement that is funded and rehearsed.

Why "Ownership" separates levels

Three product teams building three socket fleets means three heartbeat bugs, three reconnect storms and a phone holding three sockets that drain its battery. Staff makes one product's real-time correct. Principal notices the duplication and makes the device hold one connection that every product multiplexes over — with quotas so one team's broadcast can't starve another's chat.


The Five Fault Lines#

#Fault LineThe Tension
1Pull vs PushStateless simplicity vs freshness and efficiency at low change rates
2Signal vs PayloadPush "something changed" and let clients fetch, or push the data itself
3Per-User Routing vs Topic BroadcastPrecise delivery vs fan-out cost at large audiences
4Complete vs LatestEvery update in order vs only the newest value
5Sticky State vs Stateless GatewaysRoom state on one server vs any gateway can serve any client

Fault Line 1: Pull vs Push#

Polling cost is clients / interval, regardless of how often data changes. Push cost is connections held + events × recipients. For a dashboard that changes once a minute viewed by 50K people, polling every 10s is 5K req/s of mostly 304s — trivial behind a CDN or cache. For a chat app where each user gets 2 messages a minute but needs them within 500ms, polling at 0.5s intervals across 3M users would be 6M req/s — absurd; push wins by three orders of magnitude. Who pays: polling — backend QPS, mobile battery, and users waiting on average interval/2; push — the platform team running stateful infrastructure and the on-call during deploys. Staff default: poll when freshness ≥ 10s or change rate is high relative to interval; push when freshness < ~5s and changes are sparse per user. Deviate when: clients are already holding a socket for another reason — then piggyback, the marginal push is nearly free.

Fault Line 2: Signal vs Payload#

Pushing the full payload (the whole new comment, the full order object) saves a round trip but couples the push channel to every schema and every authorization rule. Pushing a signal ("conversation 812 has new messages up to seq 4410") keeps the channel thin; the client fetches through normal, authorized, cacheable APIs. Who pays: payload — the platform (bigger messages, schema coupling, authorization re-checked at push time) and privacy (a payload pushed after access was revoked); signal — latency (+1 round trip, 30–100ms) and a fetch stampede when one signal goes to 100K clients. Staff default: payload for small, high-frequency, conversation-grade updates (chat messages, location); signal for large or permission-sensitive objects, with jittered fetch delay (0–2s) when the audience is large. Deviate when: broadcast scale makes the fetch stampede itself the outage — push the payload (it's identical for everyone) and serve it from the gateway.

Fault Line 3: Per-User Routing vs Topic Broadcast#

Per-user routing looks up each recipient in the registry and sends to their gateway: precise, and the cost scales with recipients. For a 20-member channel that's 20 lookups — fine. For a 2M-viewer match, it's 2M registry lookups and 2M bus messages per goal. Topic broadcast inverts it: each gateway subscribes to the topics its connected clients care about; a goal is one publish to ~200 gateways, which each fan out locally to their ~10K subscribers from memory. Who pays: per-user routing at scale — the registry and bus (millions of operations per event); topic broadcast — gateway memory for subscription tables and the loss of per-user filtering (everyone on the topic gets the same thing). Staff default: per-user routing for small audiences (DMs, small groups); gateway-level topic subscriptions above ~1K–10K subscribers per topic. Deviate when: each recipient needs a personalized payload (e.g., per-user permission redaction) — then precompute variants per audience segment rather than per user.

Diagram: Fault Line 3: Per-User Routing vs Topic Broadcast

Fault Line 4: Complete vs Latest#

Chat needs every message, in order: a skipped message is a bug. A courier's location needs only the newest point: delivering 12 queued stale positions to a slow phone is worse than delivering one fresh one. Treating everything as complete-and-ordered makes slow consumers accumulate unbounded buffers on gateways; treating everything as latest-only loses messages. Who pays: complete — gateway memory and the slow client's latency (they get old data first); latest — the user who misses an intermediate state (acceptable for location, not for "order cancelled"). Staff default: declare semantics per stream type. Complete streams get sequence numbers, a bounded per-connection buffer (e.g., 1,000 messages or 1 MB), and on overflow the gateway closes the connection and the client resumes from its cursor. Latest streams get conflation: the gateway keeps one slot per key and overwrites it, sending at most every N ms. Deviate when: a "latest" stream carries a state transition someone acts on — split that transition into a complete stream.

Fault Line 5: Sticky State vs Stateless Gateways#

If any gateway can serve any client, deploys and failures are easy: drain, reconnect anywhere. But a collaborative document or a game room needs its operations ordered by one authority, ideally co-located with the connections. Who pays: stateless gateways + shared bus — latency (an extra hop, 1–5ms) and ordering complexity; sticky rooms — failover (room state must move or be rebuilt, 1–10s) and hot rooms (one server per room caps room size). Staff default: stateless gateways for notifications, chat, feeds and location; a separate stateful "room server" tier, addressed by consistent hashing on room ID, only for workloads that need a single ordering authority. Deviate when: room size is small and latency is critical (multiplayer games) — co-locate connections and room state on the same process.


Common Interview Mistakes#

What Candidates SayWhat Interviewers HearWhat Staff Engineers Say
"We'll use WebSockets" (first sentence)"Picked a protocol before the requirements""Order status needs ~1s and is server→client only, so SSE. The chat composer needs bidirectional, so that screen uses a WebSocket."
"TCP guarantees delivery""Hasn't handled mobile reconnects""Within one connection. Across reconnects, every message carries a sequence and the client resumes from its cursor."
"Each server pushes to its connected users""Doesn't know how events find the right server""Gateways only hold sockets. A registry maps user to gateway; producers publish through a bus keyed by user or topic."
"Push the event to all 2M viewers""Per-user fan-out at broadcast scale""Gateways subscribe to the match topic; one publish, ~200 gateway deliveries, local fan-out — conflated to every 500ms."
"Clients reconnect on disconnect""Will build a reconnect storm""Jittered exponential backoff from 1s to 30s, drain gateways over 10 minutes on deploy, and admission control on the handshake path."
"Keep a message queue per user""Unbounded buffers for slow clients""Bounded per-connection buffer; overflow closes the socket and the client resumes from the durable store. Location is conflated, not queued."

Quick Reference#

Diagram: Quick Reference

Staff Sentence Templates#

"Before picking a protocol: [update type] needs [freshness], flows [one-way / both ways], and reaches [N] recipients per event. That makes it [polling / SSE / WebSocket] with [per-user routing / topic broadcast]."

"TCP only covers one connection. Every [message] carries a per-[conversation] sequence; the server keeps [N minutes] of replay; a client that reconnects with a cursor older than that does a full fetch."

"Gateways hold sockets and nothing else, so we can deploy [business logic] without disconnecting anyone. Gateway deploys drain over [10 minutes] with [jittered] reconnects, which keeps handshake rate under [X/s]."

"[Location / prices] are latest-value streams, so the gateway conflates per key and sends at most every [N ms]. [Order status] is a complete stream — never conflated, always replayable."


Implementation Deep Dive#

1. Polling Done Right — ETag + Server-Suggested Interval#

Polling is underrated. Done well, it's stateless, cacheable and debuggable with curl.

GET /v1/orders/8812/status
If-None-Match: "v41"

# Server
function getStatus(order_id, etag):
    v = cache.get("order:ver:" + order_id)        # tiny version key, ~0.5ms
    if etag == quote(v):
        return 304, headers={"Retry-After-Ms": interval_for(order_id)}
    body = loadStatus(order_id)
    return 200, body, headers={"ETag": quote(v), "Retry-After-Ms": interval_for(order_id)}

function interval_for(order_id):
    # adaptive: fast while the courier is near, slow otherwise
    return state(order_id) == "arriving" ? 2000 : 15000

Numbers: 1M active orders × one poll per 15s = ~67K req/s, >90% of them 304s costing one Redis GET each. The server-suggested interval lets you slow everyone down during an incident with a config change — a load-shedding lever push systems don't get for free.

2. SSE with Resume — Last-Event-ID and a Replay Buffer (Redis Streams)#

GET /v1/stream?topics=orders,notifications
Accept: text/event-stream
Last-Event-ID: 1730301234567-3            # sent automatically by EventSource on reconnect

# Gateway
function stream(user, last_id):
    if last_id:
        missed = redis.XRANGE("user:" + user + ":events", "(" + last_id, "+", COUNT=500)
        if missed is empty and last_id < redis.XINFO_first_id(...):
            send("event: resync\ndata: {}\n\n")   # cursor too old: client refetches state
        for e in missed: send_event(e)
    subscribe(user)                                # registry: user -> this gateway, TTL 90s
    every 25s: send(": ping\n\n")                  # comment line, keeps LB/NAT alive

# Producer side
function publish(user, event):
    id = redis.XADD("user:" + user + ":events", MAXLEN="~", 1000, "*", event)  # bounded replay
    gw = registry.get(user)
    if gw: bus.publish("gw:" + gw, {user, id, event})
    else:  push_notifications.maybe_send(user, event)     # offline path

Why it matters: the replay buffer (here the last ~1,000 events per user, or a few minutes) closes the gap between the old socket dying and the new one opening — typically 1–30s on mobile. The resync event handles the rarer case of a long absence by falling back to a full state fetch instead of pretending nothing was missed.

🎯 Staff Insight: Notice the push path writes to the durable buffer before routing to a gateway. Delivery is an optimization on top of the log; the log is the truth. That ordering is what makes reconnects correct rather than lucky.

3. Gateway Tier with Connection Registry — WebSockets + Redis Pub/Sub#

# On connect (after auth via short-lived token)
function onOpen(conn, user, device):
    conns[user + ":" + device] = conn
    registry.SET("conn:" + user + ":" + device, gateway_id, EX=90)   # refreshed on heartbeat
    for topic in conn.subscriptions:
        if local_subs[topic].size == 0: bus.SUBSCRIBE("topic:" + topic)
        local_subs[topic].add(conn)

# Heartbeat: client pings every 25s; gateway closes after 2 missed (60s)
function onPing(conn):
    registry.EXPIRE("conn:" + conn.key, 90)

# Per-gateway inbound channel for user-targeted messages
bus.SUBSCRIBE("gw:" + gateway_id)
on bus message (target, payload):
    c = conns[target]
    if c == null: return                             # moved; producer will retry via log/resume
    if c.buffer_bytes > 1MB: c.close(4008, "slow consumer")   # client resumes from cursor
    else: c.send(payload)

# Graceful drain on deploy
function drain():
    stop_accepting()
    for c in conns.values() in batches of 1% every 6s:     # ~10 minutes total
        c.close(4000, "reconnect", retry_after_ms=random(0, 5000))

Numbers: at ~20 KB per connection (buffers + TLS + bookkeeping), 300K connections ≈ 6 GB of memory per gateway node. Registry writes are dominated by heartbeats: 3M connections / 25s ≈ 120K EXPIRE/s — size the registry for heartbeats, not messages. Draining 1% per 6 seconds keeps a 300K-connection node's reconnect rate around 500/s instead of 300K in one second.

4. Broadcast with Conflation — Gateway-Level Topic Fan-out#

# Upstream: score service publishes every change
bus.PUBLISH("topic:match:991", {seq: 88, score: "2-1", clock: "67:12"})

# Gateway: conflate per topic, flush on a timer
latest = {}                                        # topic -> newest message
on bus message (topic, msg):
    if msg.seq > latest[topic].seq: latest[topic] = msg

every 250ms:
    for (topic, msg) in latest.dirty():
        frame = encode_once(msg)                   # serialize once, reuse bytes
        for c in local_subs[topic]:
            c.send_nonblocking(frame)              # drop for this client if its buffer is full
        latest.mark_clean(topic)

Why it matters: with 200 gateways and 2M subscribers, a goal costs 1 publish + 200 bus deliveries + 2M local socket writes spread across 200 nodes (~10K each, a few ms of CPU). Encoding once per gateway rather than per socket matters: at 2M recipients, per-recipient JSON encoding is the difference between 4ms and 400ms per update. Conflation means a flurry of 10 changes in 250ms costs one frame per client.

Technique Comparison

TechniqueLatencyServer StateCompletenessOperational Burden
Short polling (ETag)interval/2 avgNoneLatest on each pollLow
Long polling~RTT after changeHeld request per clientNeeds cursor per requestMedium
SSE + replay buffer~RTTConnection + registryComplete via Last-Event-IDMedium
WebSocket gateway~RTT, bidirectionalConnection + registry + subsComplete via app-level seqHigh
Topic broadcast + conflation≤ conflation windowSubscription tablesLatest onlyMedium–High
Mobile pushSeconds, best-effortProvider-sideNone guaranteedLow (but rate-limited)

Architecture Diagram#

Diagram: Architecture Diagram

How to narrate it: the numbered arrows are the contract. A producer appends to the durable per-user log first, then looks up where the user is connected, then publishes. If the user isn't connected, the offline path takes over. Gateways only hold sockets, refresh their registry entries via heartbeats, and replay from the log on resume. That separation is why you can deploy every service on the right without dropping a single connection on the left.

Diagram: Architecture Diagram

The sequence shows why the log write comes first: the message produced during the reconnect gap was "lost" by routing and recovered by resume.


Failure Scenarios#

1. Gateway Deploy Reconnect Storm — 1.8M Clients in 40 Seconds#

A config change triggers a restart of all 6 gateway nodes in one region within a minute. Clients use a fixed 1-second retry with no jitter.

t=0      Rolling restart begins; orchestrator max-unavailable set to 50 percent.
t=+10s   900K sockets closed. Clients retry at exactly +1s, +2s, +3s.
t=+12s   Handshake rate 450K/s. TLS termination CPU 100 percent at the LB.
t=+15s   Auth service (token validation on connect) p99 4s; 30 percent of handshakes time out.
t=+20s   Timed-out clients retry on 1s cadence; offered handshakes stay above 400K/s.
t=+40s   Remaining 900K sockets closed by second half of restart. Storm doubles.
t=+3min  Registry Redis at 100 percent CPU from SET + EXPIRE on connect.
t=+11min On-call scales auth 4x and blocks new handshakes at 50K/s via LB rate limit.
t=+24min All clients reconnected. Missed events replayed from log for 92 percent;
         8 percent exceeded replay window and did full resync, spiking API reads.

Detection: gateway.handshakes_per_sec > 5× baseline; gateway.connected drop > 20% in 1 min; auth.p99 on connect; registry.cpu; resume.resync_rate. Blast radius: every real-time user in the region, plus auth and API services that weren't part of the change. Mitigation: handshake admission control at the LB (token bucket), temporarily extending the server-suggested retry delay. Prevention: gateway restarts limited to one node at a time with 10-minute drain; client SDK with jittered exponential backoff (1s → 30s, full jitter); auth on connect uses locally verifiable tokens (JWT with short TTL) instead of a remote call; replay window sized to cover a 30-minute drain. Owner: real-time platform team (gateway, SDK, drain tooling); identity team for connect-path auth capacity.

🎯 Staff Insight: The incident spread to auth and the API tier — services nobody touched. In stateful connection systems, the steady state is never the risk; the transition is. Design and load-test the reconnect path at 100% of connections.

2. Silent Gap — Messages Lost Between Sockets#

A chat product pushes messages directly from the bus to gateways with no per-conversation sequence and no replay. Mobile clients reconnect frequently on cellular handoffs.

t=0      User on a train; LTE handoff every 3-5 minutes. Each drop lasts 2-8s.
t=+2s    Message from a teammate routed to old gateway; socket already gone. Dropped.
t=+6s    Client reconnects to a new gateway. Shows "connected". No gap detected.
t=+1h    User replies to a thread missing two messages. Confusion.
t=+2wk   Support volume "missing messages" up 3x after a mobile release that
         reduced background keepalive. No metric fired at any point.

Detection: none — the gap. Add delivery.gap_detected (client reports when it receives seq N+2 after N), delivery.undeliverable at gateways for targets not connected, and a periodic client-side reconciliation that compares last-seen seq with the server. Blast radius: mobile users disproportionately — the most engaged ones. Mitigation: on app foreground, fetch messages since last seen seq. Prevention: per-conversation sequence numbers, write-to-log-before-route, resume by cursor on every reconnect, and client-side gap detection that triggers a fetch. Owner: messaging team owns sequence and resume semantics; platform SDK owns gap detection and metrics.

3. Slow Consumer Memory Blow-up — One Topic Takes Down a Gateway#

A live-trading dashboard subscribes 40K users to a price topic updating 50 times per second. Outbound frames are queued per connection without bound. A subset of clients are on poor networks.

t=0      Market opens. 50 updates/s x 40K subscribers spread over 8 gateways.
t=+2min  3 percent of clients cannot keep up; their outbound queues grow 50 frames/s.
t=+10min Gateway 4 heap from 4 GB to 22 GB; GC pauses 2-6s.
t=+12min GC pauses miss heartbeats; LB marks gateway 4 unhealthy; 300K sockets drop.
t=+13min Those sockets reconnect to gateways 1-3 and 5-8, bringing slow clients along.
t=+20min Two more gateways enter the same spiral.

Detection: gateway.outbound_buffer_bytes p99 per connection; gateway.heap_used; gateway.gc_pause_seconds; conn.closed_slow_consumer count. Blast radius: every user on affected gateways, including those on unrelated topics. Mitigation: emergency config: conflate the price topic at 4 updates/s; cap per-connection buffers at 256 KB. Prevention: declare price streams as latest-value with conflation; bounded per-connection buffers with close-and-resume on overflow; per-topic rate limits enforced at publish. Owner: platform owns buffer limits and conflation; trading product owns the topic's declared semantics.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Reconnect stormgateway.handshakes_per_sec > 5×Region's real-time users + authLB handshake limiter, longer retry hintReal-time platform
Messages lost across reconnectdelivery.gap_detected, delivery.undeliverableMobile usersResume by cursor, foreground syncMessaging team + SDK
Slow consumersgateway.outbound_buffer_bytes p99Whole gateway nodeConflate, cap buffers, close-and-resumePlatform + topic owner
Stale registry entriesdelivery.misrouted_rateDelayed deliveryShorter TTL, gateway-side redirectReal-time platform
Hot broadcast topicbus.topic_publish_rate, egress GbpsBus + gateway egressConflation, topic rate limitTopic owner
LB idle timeout kills socketsSpikes of closes at exact intervals (e.g., 60s)All idle clientsHeartbeat < idle timeoutPlatform + network team
Push provider throttlingpush.rejected by APNs/FCMOffline usersCollapse keys, priority classesNotifications team

The Principal Lens#

Why L7 Sees This Problem Differently#

A Staff engineer builds a correct real-time path for one product. A Principal engineer notices that the mobile app holds three sockets — chat, live location and notifications — built by three teams on three gateway fleets, each with its own heartbeat interval, each draining the battery, each reconnecting independently after a tunnel. The org-level problem is one device, one connection, many products: a shared real-time platform that multiplexes every product's streams over a single connection, with declared delivery semantics per stream, quotas per product, and one team that is excellent at the hardest operational part — moving millions of connections safely during deploys and failures.

The Org-Level Fault Line#

One multiplexed real-time platform vs per-product socket fleets.

OptionWhat WorksWhat BreaksWho Pays
Per-product fleetsAutonomy; product-specific protocol tuningN sockets per device, N reconnect storms, N heartbeat bugs; battery and data costUsers' batteries; N on-call rotations
Managed real-time vendorNo fleet to run; global edge on day onePer-connection/message pricing grows linearly; SDK lock-in; limited control over resume semanticsFinance at scale; platform during migration
Shared platform: one connection per device, topic multiplexing, per-stream contractsOne expert team, consistent resume and backoff, one socket per devicePlatform becomes a dependency for every real-time feature; noisy-neighbor risk across productsPlatform headcount (4–6); quota governance

The Principal default is the shared platform once three or more products need push: the platform owns connections, registry, routing, SDK and drain; product teams own topics, payload schemas and their declared semantics (complete vs latest, freshness SLO), within quotas.

Cost Model#

Assumptions: ~20 KB per connection; gateway nodes at ~$600/month holding ~300K connections at < 60% memory; bus/registry clusters ~$2–4K/month per cell; egress ~$0.05/GB; managed-vendor list pricing commonly lands in the range of ~$0.5–2 per thousand peak connections per month plus per-message fees; engineer ~$25K/month.

ScalePeak ConnectionsSelf-hosted InfraManaged Alternative (order of magnitude)People / On-callRough Monthly Total (self-hosted)
Small20K2 gateways + small Redis (~$2K)~$1–3K0.3 FTE; product team on-call~$2K + ~$8K people — managed usually wins
Growth1M8 gateways/region × 2 regions, registry + bus cells (~$20K), egress ~$10K~$30–60K + messages2 FTE; shared rotation~$30K + ~$50K people
Large20M150+ gateways across regions, bus cells, replay store (~$180K), egress ~$120K~$400K–1M+5–6 engineer platform team; dedicated rotation~$300K + ~$150K people

The Principal observation: below ~100K connections, a managed service is cheaper than the people needed to run gateways well. Above a few million, the vendor bill grows linearly while the self-hosted cost is dominated by a fixed team. The crossover is a decision to make deliberately — and the one-way part is the client SDK, so wrap it from day one.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Heartbeat interval, conflation window, replay lengthTwo-wayConfig
SSE vs WebSocket for a streamTwo-way (behind an SDK)Client release cycle
Client wire protocol and resume semantics shipped in mobile appsOne-way-ishOld app versions live for 1–2 years; server must support both
Promising "real-time" completeness in a public/partner APIOne-wayPartners build on it; weakening it is a breaking change
Vendor SDK called directly throughout app codeOne-way-ishRip-and-replace across every screen; wrap it on day one
Topic naming and authorization modelOne-way-ishEvery producer and subscriber depends on it

The Standard I'd Write#

RFC: Real-Time Delivery Standard (v1)

Scope: Any server-initiated update to clients — web, mobile, partner integrations.

MUST:

  1. Each stream declares semantics: complete (ordered, sequence-numbered, replayable) or latest (conflated), and a freshness SLO (p99 delivery latency).
  2. Producers append complete-stream events to the durable per-recipient or per-topic log before routing; delivery is best-effort on top of the log.
  3. All clients use the platform SDK: jittered exponential backoff (1s → 30s, full jitter), resume by cursor, gap detection.
  4. Gateways hold no business logic; per-connection outbound buffers are bounded; overflow closes with a resume code.
  5. Gateway deploys drain at ≤ 1% of connections per 5 seconds per node; reconnect capacity is load-tested at 100% of peak connections twice a year.

SHOULD: Prefer polling for freshness ≥ 10s; use signal-then-fetch for large or permission-sensitive payloads; broadcast topics above 10K subscribers use gateway-level subscription and conflation.

Exceptions: Real-time platform review; stateful room servers (collaboration, games) require a design review for ordering and failover.

Success metrics: delivery.gap_detected < 0.01% of messages; p99 delivery within declared SLO per stream; one connection per device for platform products; zero Sev-1s from gateway deploys.

What I'd Tell the VP#

Customers increasingly expect our app to update by itself — couriers moving, messages arriving, scores changing. Today three teams each built their own way of doing that, which means three times the infrastructure, three times the battery drain on phones, and three different ways it breaks when we deploy. I'm proposing a single real-time platform that every team uses, with one connection per device and a clear promise for each kind of update about how fast and how complete it is. It's roughly five engineers ongoing, offset by retiring two fleets and avoiding a vendor bill that would grow past $500K a month at our projected scale.

Principal Interview Signals#

SignalWhat It Sounds Like
One device, one connection"The phone shouldn't hold a socket per product. I'd multiplex every stream over one connection and give products topics and quotas."
Semantics as a contract"Each stream declares complete or latest and a p99. That's what product teams sign up for, not 'WebSockets'."
Transitions funded"Reconnect capacity for 100% of connections is a requirement we test twice a year, not a hope."
Build vs buy crossover"Managed until about a million connections, behind our own SDK so the switch is a server change, not an app rewrite."
Blast radius"Regional gateway cells, so a bad deploy is one cell's reconnect, not the world's."

Staff answers that L7 interviewers find insufficient:

  • "I'd build a gateway tier with a registry and Redis pub/sub." — Correct for this product; silent on the two other fleets doing the same thing.
  • "Clients resume from a cursor." — Right mechanism, but not an org guarantee: what about the teams not using cursors?
  • "We'll use a managed service." — No crossover analysis, no SDK wrapper, no exit path.

In the Wild#

Slack: Channel Servers and Gateway Servers#

Slack's engineering blog has described its real-time messaging architecture: stateful channel servers that own channels (mapped by consistent hashing) and hold recent history, gateway servers that hold client WebSocket connections and subscribe to the channels their users need, plus separate admin and presence services. Messages are persisted through the API tier and then fanned out from channel servers to the gateways subscribed to that channel.

Staff insight: This is topic-level routing in production: gateways subscribe to channels on behalf of their users, so a message to a big channel goes to the gateways that care rather than being routed per user. It also separates socket holding from channel state — the gateway/room-server split from Fault Line 5.

Discord: Elixir Gateways and Per-Guild Processes#

Discord has written publicly about building its real-time gateway on Elixir/the Erlang VM, with a process per connected session and a process per guild (server) that fans events out to the sessions in it, and about the work required to scale very large guilds — including relaying fan-out across multiple processes when a single guild's member count made one process the bottleneck.

Staff insight: A per-room ordering authority is elegant until one room is huge. The large-guild work is Fault Line 3 and Fault Line 5 colliding: sticky state gives ordering, and the hottest room needs its fan-out split anyway.

LinkedIn: Hundreds of Thousands of Persistent Connections per Machine#

LinkedIn's engineering blog described its instant messaging and presence infrastructure using Server-Sent Events over long-lived HTTP connections on the Play framework and Akka actors, with each connection represented as a lightweight actor, and reported sustaining hundreds of thousands of concurrent connections per machine after tuning.

Staff insight: A large, mature product chose SSE, not WebSockets, for server-to-client delivery — and the scale achieved per box was a function of connection memory and file-descriptor tuning, not protocol choice. It's a strong reference when an interviewer assumes "real-time means WebSockets".


Practice Drill#

Prompt: "We're building live order tracking for a food delivery app: 1.5M concurrent orders at dinner peak. Customers see the courier move and status changes (accepted, picked up, arriving). Couriers send GPS every 4 seconds. Design the real-time path."

Staff Answer

Two stream types with different semantics. Status is complete: ordered, never lost, replayable — append to a per-order event log (Redis Stream, MAXLEN ~200, TTL 6h), seq per order. Location is latest-value: only the newest point matters — conflate per order, send at most every 1s. Both flow server→client only, so SSE through a gateway tier; the courier app uploads GPS over plain HTTPS batched every 4s (375K req/s at peak, ingestion → Kafka keyed by courier). Routing: a location service consumes courier GPS, maps courier → active order(s), writes the latest point to order:{id}:loc and publishes to gw:{gateway} found via the registry (order:{id} → gateway, TTL 90s, refreshed by 25s heartbeats). Status changes append to the log first, then publish. Sizing: 1.5M customer connections at ~20 KB ≈ 30 GB; ~6–8 gateways per region at 250K each with headroom; location messages ~1.5M/4s ≈ 375K/s across the bus — split bus by order hash into 8 cells. Resume: Last-Event-ID replays status since the cursor; location just sends the current point. Offline: status changes also go to push notifications if no live connection. Deploy safety: 1% per 6s drain, SDK with full-jitter backoff, handshake limiter at the LB. Fallback: if the stream fails, the app polls /orders/{id} every 10s with ETag — a degraded mode product signed off on. Metrics: delivery.p99_ms per stream, gateway.connected, gateway.handshakes_per_sec, resume.resync_rate, location.staleness_seconds.

Why this is L6:

  • Splits complete vs latest semantics and designs each separately rather than one queue for everything.
  • Writes to the log before routing, so reconnect gaps are recovered, and sizes connections, registry and bus with numbers.
  • Designs the transitions — drain, backoff, handshake admission — and a polling degraded mode.

What L7 adds:

  • Notices chat-with-courier and promotions also want push and proposes one multiplexed connection with topic quotas instead of a second fleet.
  • Prices self-hosted (~$30K/month + 2 FTE) against a managed vendor at 1.5M peak connections and wraps the client SDK either way.
  • Sets a quarterly game-day: drop a region's gateways at dinner-peak load in staging and prove reconnect capacity.

Staff Interview Application#

How to Introduce This Pattern#

"Before I pick a protocol, I want to classify the updates: how fresh each needs to be, whether every update matters or only the latest, and how many people receive each one. Then the hard parts are where connections live, how events find them, and what a client missed while it was reconnecting — I'll spend most of my time there."

Lead with freshness and semantics, then the gateway/registry/bus split, then fan-out size, then reconnect and deploy behavior.

When NOT to Use This Pattern#

  • Freshness ≥ 10–30s: Poll with ETags. It's cacheable, stateless and debuggable; push adds a stateful fleet for no user-visible gain.
  • Background or offline delivery: Sockets don't survive backgrounding on mobile. Use push notifications + sync-on-open.
  • Large payloads or downloads: Push a signal, fetch via normal APIs or a CDN; don't stream megabytes through gateways.
  • Request/response semantics: If the client asks and waits for one answer, that's an HTTP call, not a stream.

Follow-Up Questions to Anticipate#

Interviewer AsksWhat They Are TestingHow to Respond
"How does the server know which machine a user is on?"Routing"Registry: user/device → gateway with a 90s TTL refreshed by heartbeats. Producers look it up and publish to that gateway's channel."
"What if the user disconnects and misses messages?"Completeness"Every message is in a log with a sequence. Reconnect sends the last seq; the gateway replays. Too old? Full resync."
"How do you deploy gateways?"Transition design"Drain 1% per 6s, server-suggested retry with jitter, handshake limiter. Logic lives elsewhere so most deploys don't touch gateways."
"A match has 2M viewers?"Broadcast"Gateways subscribe to the topic, one publish fans to ~200 gateways, local fan-out, conflated every 250–500ms."
"WebSockets or SSE?"Pragmatism"SSE for server→client — HTTP-native, auto-resume. WebSocket only where the client sends often."
"What about slow clients?"Backpressure"Bounded buffers per connection. Latest streams conflate; complete streams close and resume from the log."

Evaluation Rubric#

DimensionSenior (L5)Staff (L6)Principal (L7)
Framing"WebSockets"Freshness, direction, audience size, complete vs latest per streamReal-time as a shared platform with per-stream contracts
ArchitectureServers push to their own socketsGateway + registry + bus; topic broadcast at scaleOne connection per device, multiplexed, with quotas and cells
CompletenessTCPSequence, log-before-route, resume, gap detectionDeclared semantics and gap SLO across all products
FailureClient reconnectsDrain, jitter, handshake admission, bounded buffersFunded reconnect capacity, regional cells, game-days
CostNot discussedConnections per node, bus sizingBuild vs managed crossover, SDK as exit path

Strong Hire Signals

SignalWhat It Sounds Like
Questions push itself"Does this need push, or is a 10s poll fine?"
Gap-aware"TCP covers one connection; users have many."
Transition-focused"The deploy is the incident; I'll design the drain."
Fan-out math"2M recipients is 200 gateway messages, not 2M bus messages."

Lean No-Hire Signals

SignalWhy It Misses the Bar
Business logic inside the socket serversEvery deploy disconnects everyone
No answer for missed messages on reconnectSilent data loss in the most common mobile case
Per-user routing for million-viewer eventsRegistry and bus collapse at the moment that matters

Common False Positives: Knowing the WebSocket handshake headers ≠ designing routing. Naming Redis pub/sub ≠ handling slow consumers. "Socket.IO handles it" ≠ a resume protocol.


Capacity Planning Quick Reference#

Sizing the Real-Time Path#

gateway_nodes       = ceil(peak_connections / conns_per_node / 0.6)    # 60% memory target
conn_memory         = peak_connections × 20 KB                          # 10-50 KB typical
registry_ops        = peak_connections / heartbeat_seconds              # heartbeats dominate
poll_qps            = active_clients / poll_interval_seconds
fanout_deliveries   = events_per_sec × avg_recipients                   # per-user routing
broadcast_bus_msgs  = events_per_sec × gateways_subscribed              # topic routing
reconnect_rate      = connections_per_node × drain_fraction / drain_interval
replay_buffer       = recipients × events_per_min × replay_minutes × avg_event_bytes

Key Numbers Worth Memorizing#

NumberContext
10–50 KBMemory per idle connection (TLS + buffers + bookkeeping)
100–500KMostly idle connections per tuned gateway node
25–30 sHeartbeat interval; under typical 60s LB/NAT idle timeouts
~6Browser HTTP/1.1 connections per origin — the SSE gotcha; HTTP/2 multiplexes
1 s → 30 sJittered exponential reconnect backoff range
2–10 minReplay buffer that covers most mobile reconnect gaps
1–8 sTypical cellular handoff disconnection
250–1000 msConflation window for broadcast and location streams
~10KSubscribers per topic where per-user routing should give way to topic broadcast
1% / 6 sDrain rate that turns a deploy into a trickle, not a storm

Common Pitfalls Checklist#

  • Each stream declares complete vs latest and a freshness SLO
  • Polling considered first for freshness ≥ 10s
  • Gateways hold no business logic; logic deploys don't drop sockets
  • Events written to a log before routing; clients resume by cursor
  • Heartbeat interval below every LB/NAT idle timeout on the path
  • Client SDK uses jittered exponential backoff
  • Per-connection buffers bounded; slow consumers closed or conflated
  • Broadcast topics use gateway-level subscription and conflation
  • Offline path (push + sync-on-open) defined for backgrounded clients
  1. Loading the index…