Every "make it real-time" requirement ends up here, and most candidates answer "WebSockets" before the question is finished. All three let a server deliver data to a browser or app without the user refreshing. They differ in direction and in how much of the web's infrastructure they keep. Long polling is ordinary HTTP requests that the server holds open until it has something to say. Server-Sent Events (SSE) is one long HTTP response that streams events from server to client, with reconnect and resume built into the browser. WebSockets upgrade an HTTP connection into a full-duplex, message-framed socket that leaves HTTP semantics behind. The question that decides it: does the client need to send frequent messages back on the same channel, or does data flow only from server to client? Most real-time features — feeds, notifications, dashboards, order tracking, live scores — are one-way, and one-way is SSE's job.
The Verdict#
Default to SSE for server-to-client streams, WebSockets when the client sends often (chat typing, cursors, games), long polling only as a fallback or for low-frequency change notifications — and plain polling when 10+ seconds of delay is acceptable.
| Pick SSE when | Pick WebSockets when | Pick long polling when |
|---|---|---|
| Data flows server → client: feeds, notifications, dashboards, progress, LLM token streams | Client sends frequently on the same channel: chat with typing indicators, collaborative cursors, multiplayer games | Updates are infrequent (minutes apart) and you want plain request/response everywhere |
You want auto-reconnect and resume (Last-Event-ID) for free | You need binary frames or very low per-message overhead in both directions | Some clients sit behind proxies that break streaming responses |
| You want standard HTTP: auth headers on the server side, compression, HTTP/2 multiplexing, existing load balancers | Latency both ways matters at < 100 ms and messages are small and frequent | You are building a public API for third parties who can't hold streams (a common pattern for "wait for changes" endpoints) |
| Client → server messages are rare and can be normal HTTP POSTs | You control both ends (your app + your gateway) | You need a fallback when SSE or WebSockets fail |
🎯 Staff Move: "Before I pick a protocol: does the client talk back on this channel? Here it doesn't — order status and courier location only flow down — so I'll use SSE with event IDs for resume, and the rare client actions go over normal HTTPS. I'd switch to WebSockets for the chat screen, where typing indicators go up several times a second."
At a Glance#
| Dimension | WebSockets | SSE | Long polling |
|---|---|---|---|
| Direction | Full duplex | Server → client only | Server → client (client sends via normal requests) |
| Transport | HTTP Upgrade to the WebSocket protocol (RFC 6455); also over HTTP/2 (RFC 8441) and HTTP/3 (RFC 9220) | One long-lived HTTP response, Content-Type: text/event-stream | Repeated HTTP requests, each held until data or timeout |
| Message format | Text or binary frames | UTF-8 text events (id:, event:, data:, retry:) | Whatever the response body is |
| Ordering | In order on one connection | In order on one connection | In order per response; gaps possible between requests unless you use a cursor |
| Reconnect / resume | Build it yourself | Built into browsers: auto-reconnect, sends Last-Event-ID | Natural: every request carries the client's cursor |
| Latency | Lowest both ways (one frame, ~2–14 bytes overhead) | Same as WebSockets for server → client | One round trip after each event; bursts are delivered in batches |
| Per-message overhead | Tiny | Tiny (a few bytes of field names) | Full HTTP request + response headers per batch (~0.5–2 KB) |
| Server state per client | One open connection (~10–50 KB memory) | One open connection (~10–50 KB) | One parked request at a time; cheaper between requests |
| Proxies and load balancers | Need explicit Upgrade support and long idle timeouts | Works through most HTTP infrastructure; buffering proxies must be told not to buffer | Works everywhere HTTP works |
| Browser limits | No per-origin cap like HTTP/1.1's | Counts against ~6 connections per origin on HTTP/1.1; multiplexed on HTTP/2 | Same per-origin cap as any request |
| Scaling model | Stateful gateways, connection registry, pub/sub fan-out | Same as WebSockets | Stateless-ish servers; state lives in the cursor and the backing store |
| Operational burden | Highest: deploy drains, reconnect storms, custom heartbeat and resume | Medium: same gateways, but resume and reconnect are standard | Lowest per server; request rate can be high |
| Managed options | API Gateway WebSocket APIs, Cloudflare, Azure Web PubSub, many hosted real-time services | Any HTTP host or CDN that supports streaming responses | Any HTTP host |
| Cost shape | Connection-minutes + messages | Connection-minutes + bytes | Requests |
How They Actually Differ#
Direction and What the Client Really Sends#
The usual argument for WebSockets is "it's bidirectional". The question is how often the client sends, and whether those sends need to be on the same channel. A news feed client sends nothing between page loads. A notifications client sends "mark as read" a few times a session — a normal POST is fine. A chat client sends typing indicators several times a second and messages that must arrive in order with the server's events. A multiplayer game sends input 20–60 times a second.
| Feature | Client → server rate | Best fit |
|---|---|---|
| News feed "new posts" banner | ~0 | SSE (or polling every 30s) |
| Order tracking | ~0 | SSE |
| Notifications with read receipts | A few per session | SSE + HTTP POST |
| AI assistant streaming tokens | One prompt per response | SSE (one streamed response per request) |
| Chat with typing indicators | Several per second while typing | WebSockets |
| Collaborative cursors / editing | 10–30 per second | WebSockets |
| Multiplayer game input | 20–60 per second | WebSockets (or UDP-based transports in native clients) |
Who pays: choosing WebSockets for one-way data makes the platform team pay for custom reconnect, resume and heartbeat logic that SSE gives for free. Choosing SSE for a high-frequency two-way feature makes the client team pay for a second channel and for ordering between the two.
Reconnect and Resume: Built In vs Build It#
Connections drop constantly: phones change networks, laptops sleep, load balancers time out idle connections, deploys restart gateways. What happens next is the real design.
SSE's browser EventSource reconnects automatically, waits the server-suggested retry: interval, and sends the last event ID it saw in a Last-Event-ID header. If the server keeps a short replay buffer keyed by event ID, resume is a few lines of server code.
WebSockets have no reconnect and no resume. The client library must detect the drop (ping/pong or application heartbeats), back off with jitter, reconnect, re-authenticate, re-subscribe, and send its last sequence number so the server can replay — all custom.
Long polling resumes by construction: every request carries a cursor (?since=evt_9812), and the server answers with everything after it.
What each looks like on the wire:
SSE response (one HTTP response that never ends)
HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache
retry: 5000 ← client waits 5 s before reconnecting
id: 9813
event: status
data: {"order":42,"state":"picked_up"}
: keepalive ← comment line every ~20 s defeats idle timeouts
Long polling (one request per batch)
GET /orders/42/changes?cursor=9812&wait=30
→ held up to 30 s → 200 {"events":[...], "cursor":"9814"} or 200 {"events":[], "cursor":"9812"}
client immediately re-requests with the new cursor (plus jitter after an empty response)
WebSocket client resume (all of it is yours to write)
on open: send {"type":"auth","token":...}; send {"type":"resume","since":last_seq}
on message: if msg.seq != last_seq + 1 → request replay; last_seq = msg.seq
every 25 s: send ping; if no pong in 10 s → close and reconnect
on close: wait random(0, min(30 s, 1 s × 2^attempt)); reconnect
Infrastructure Fit#
SSE is an HTTP response that never ends. It uses the same authentication middleware, the same TLS termination, the same HTTP/2 connection the page already has, and compression where it helps. The traps are buffering proxies (which hold the response until it ends — disable buffering, for example with X-Accel-Buffering: no on NGINX) and idle timeouts (send a comment line every ~15–30 s). One limitation: the browser EventSource API cannot set custom request headers, so auth uses cookies or a short-lived token in the URL — or you use a fetch-based SSE client.
WebSockets need every hop to understand the Upgrade and to tolerate hours-long connections: load balancer idle timeouts (often 60 s by default — heartbeat under that), corporate proxies that strip Upgrade headers, and WAF rules that inspect HTTP but not frames. After the upgrade, HTTP tooling (status codes, caching, per-request logs) no longer applies.
Long polling is ordinary HTTP. Every proxy, CDN and firewall handles it. The cost is request volume: 1M clients with 30-second holds is ~33K requests/s at minimum, more when events are frequent.
Fan-out and Server Cost#
For WebSockets and SSE the server architecture is the same: a fleet of stateful gateways holding connections, a connection registry (user → gateway), and a pub/sub bus that routes events to the right gateway. A tuned gateway holds ~100K–500K mostly idle connections; at ~10–50 KB per connection, 1M connections is ~10–50 GB of memory across the fleet. The protocol choice changes almost nothing about this tier — which is why picking WebSockets "for scale" is a non-argument.
Long polling changes the shape: servers can be closer to stateless (a parked request waits on a subscription; when it returns, the server forgets the client until the next request). It pays in request rate and in latency for bursts: three events 10 ms apart may arrive as one response plus a reconnect.
| 1M connected clients, 1 event per client per minute | WebSockets | SSE | Long polling (30 s hold) |
|---|---|---|---|
| Open connections | 1M | 1M | ~1M parked requests |
| Requests/s | ~0 (after connect) | ~0 (after connect) | ~25K–35K/s (timeouts plus event responses) |
| Bytes per event | ~100 B payload + ~6 B framing | ~100 B + ~20 B field names | ~100 B + ~800 B headers |
| Extra work | Custom resume, heartbeat | Replay buffer for Last-Event-ID | Cursor per request |
Where Each One Breaks#
WebSockets#
| Failure | What happens | Detection | Owner |
|---|---|---|---|
| Deploy reconnect storm | Restarting a gateway drops 250K sockets; all reconnect in the same second; auth and registry overload | Handshakes/s, auth latency | Platform |
| Silent half-open connections | NAT or proxy drops the path without a close; server thinks the client is connected and keeps queuing | Heartbeat failures, messages sent without acks | Platform |
| Gap after reconnect | Messages published during the reconnect window never reach the client | Client-side sequence gap counter | Platform + client team |
| Slow consumer | A client on a bad network can't drain; server-side buffer grows until memory pressure | Per-connection send queue size | Platform |
| Middlebox breaks Upgrade | Some enterprise networks block or strip Upgrade; connection fails or falls back | Connect failure rate by network/ASN | Client team |
SSE#
| Failure | What happens | Detection | Owner |
|---|---|---|---|
| Buffering proxy | Events arrive in bursts minutes late or only when the connection closes | Time from publish to client receipt | Platform |
| HTTP/1.1 connection cap | Several tabs each open a stream; the browser's ~6-per-origin limit blocks other requests | Stalled requests in browser telemetry | Client team (use HTTP/2, or share one stream across tabs) |
| Idle timeout | Quiet streams get cut by the LB; clients reconnect constantly | Reconnect rate, stream duration histogram | Platform (heartbeat comments) |
| Replay buffer too short | Client returns after 20 minutes with a Last-Event-ID older than the buffer | Resume misses → full resync rate | Platform |
| Reconnect storm | Same as WebSockets on deploys — auto-reconnect makes it more synchronised, not less | Handshakes/s | Platform (server retry: with jitter, drain slowly) |
Long Polling#
| Failure | What happens | Detection | Owner |
|---|---|---|---|
| Request flood after a broadcast | One event wakes 1M parked requests; all return and immediately re-request | Request rate spikes in lockstep | Platform (jittered re-poll, server backoff hints) |
| Burst latency | Events closer together than the round trip are batched; per-event latency grows | Publish-to-receipt p99 | Platform |
| Timeouts shorter than hold | Proxy cuts at 30 s but server holds 60 s; clients see errors every cycle | 504 rate | Platform |
| Lost events between polls | Server forgets what a client saw; events between responses are missed | Gaps in client cursors | Platform (always use a cursor) |
A Gateway Deploy, Done Badly and Done Well#
Done badly (restart all gateways at once, clients reconnect immediately)
t=0 8 gateways restart → 2M connections drop
t=+1s 2M reconnect attempts; TLS handshakes + auth checks saturate the edge
t=+5s auth service p99 50ms → 4s; reconnects time out and retry → more load
t=+60s still recovering; events published during the window lost for clients without resume
Done well (drain 1% per 5 s, server-sent retry with jitter, resume from replay buffer)
t=0 gateway 1 stops accepting new connections; closes 1% of its sockets every 5 s
t=+0..8m ~4K reconnects/s fleet-wide, spread by client jitter; auth p99 unchanged
t=+8m all gateways rolled; clients resumed from Last-Event-ID / last seq; zero gaps
Who pays: done badly, users and the auth team pay. Done well, the platform team paid up front for drain tooling and a resume protocol.
The production surprise for all three: the steady state is easy; transitions are the system. Deploys, network changes and broadcasts produce synchronized reconnects that look like a DDoS. Jittered backoff (1 s → 30 s cap), slow gateway drains (1% of connections per few seconds) and a handshake rate limit at the edge matter more than the protocol choice.
Cost and Operations#
| WebSockets | SSE | Long polling | |
|---|---|---|---|
| Who runs it | A real-time platform team owns gateways, registry, client SDK | Same team; smaller client SDK | Any backend team |
| Bill scales with | Concurrent connections (memory) + messages + egress | Concurrent connections + egress | Requests + egress |
| Managed pricing shape | Typically per connection-minute + per message on hosted gateways | Whatever your HTTP hosting charges for long requests | Per request |
| Hidden cost | Custom resume protocol, client SDKs for every platform, deploy tooling | Replay buffer, proxy configuration | Header overhead and request rate at scale |
| People | 1–3 engineers for a large real-time platform | Similar platform team, less client work | Little beyond normal API ownership |
| What on-call watches | WebSockets / SSE | Long polling |
|---|---|---|
| The SLO metric | Publish-to-client latency p99; connected clients | Publish-to-client latency p99; request success rate |
| The "act now" alert | Handshakes/s spike, gateway memory > 80%, resync rate jump | Request rate spike in lockstep, 5xx on the poll endpoint |
| Routine work | Deploy drains, SDK releases, idle-timeout and heartbeat tuning | Hold-time and backoff tuning |
Sizing example. 2M concurrent clients at peak:
Connections: 2M × ~30 KB ≈ 60 GB memory across gateways
Gateways: at ~250K connections each with headroom → 8–10 nodes per region
Heartbeats: every 25 s to stay under 60 s idle timeouts → 80K small frames/s fleet-wide
Deploy: drain 1% per 5 s → a full fleet roll takes ~8 minutes, peak ~4K reconnects/s
Long-poll alt: 2M / 30 s hold ≈ 67K req/s baseline before any events — fine for an API tier,
but every one carries headers and auth checks
🧭 Principal Insight: "The cost of real-time isn't the protocol, it's the gateway tier and the client SDKs. Whatever we choose, I want one connection platform the whole company uses — one registry, one resume protocol, one deploy procedure — rather than each product team running its own WebSocket servers."
Switching Later#
| Move | How | What's hard |
|---|---|---|
| Long polling → SSE | Same cursor becomes the event ID; the endpoint streams instead of returning | Proxy buffering and timeouts on the path |
| SSE → WebSockets | Gateway accepts both; clients upgrade; resume logic ported to the custom protocol | Rebuilding reconnect/resume that EventSource gave for free |
| WebSockets → SSE | Split client sends into normal HTTP; stream down over SSE | Features that relied on ordered client→server messages on one channel |
| Any → a managed real-time service | Gateway replaced by the vendor; your services publish to it | Vendor's message model, limits and pricing; data residency |
One-way doors: the client resume protocol shipped in mobile apps (old app versions live for years); event ID / sequence semantics; the decision to run your own gateway tier. Two-way doors: the transport for web clients (servers can speak several); heartbeat intervals; replay buffer length.
How Real Companies Chose#
Slack: Persistent WebSockets for a Two-Way Product#
Slack's engineering blog describes every client holding a persistent WebSocket connection to receive real-time events, through stateful gateway servers that subscribe to channel servers (mapped by consistent hashing). It reports tens of millions of connected clients and messages delivered across the world in about 500 ms (Slack Engineering).
Staff insight: Chat is the canonical two-way case — presence, typing and message events in both directions — and the architecture's real work is the gateway and channel-server tiers, not the socket.
Shopify: SSE for a One-Way Live Map#
For its Black Friday Cyber Monday live map, Shopify replaced a polling design (with at least a 10-second delay) with a Go SSE server that subscribes to Kafka, scaled horizontally behind NGINX. They chose SSE over WebSockets explicitly because the data only flows server → client, and cited HTTP familiarity and automatic reconnection; the new pipeline put data on screen within about 21 seconds of creation with 100% uptime during the event (Shopify Engineering).
Staff insight: A high-traffic, high-visibility real-time product chose the simpler protocol because the requirement was one-way. That is the default this page argues for.
Dropbox: Long Polling as a Public API Contract#
Dropbox's public API offers a list_folder/longpoll endpoint: the client passes a cursor and the request blocks until changes are available or a timeout (30–480 seconds, default 30) passes, with up to 90 seconds of random jitter added to avoid a thundering herd; the response may include a backoff field telling the client how long to wait before polling again (Dropbox API spec).
Staff insight: For third-party developers on every kind of network, long polling with a cursor is the most compatible push mechanism there is — and the server-controlled jitter and backoff fields are how you keep a million clients from synchronising.
Follow-Ups to Expect#
| After You Say... | They Will Ask... | What They're Testing |
|---|---|---|
| "WebSockets" | "Does the client actually send anything?" | Whether you chose by direction or by habit |
| "SSE" | "How does the client resume after a network drop?" | Last-Event-ID, replay buffer, full-resync fallback |
| "A gateway tier" | "You deploy the gateways. What happens to 2M connections?" | Slow drains, jittered reconnect, handshake limits |
| "Long polling" | "One event goes to 1M users at once." | Lockstep re-poll storm; jitter; server-sent backoff |
| "Heartbeats" | "Why every 25 seconds?" | Load balancer and NAT idle timeouts (~60 s) |
| "SSE through our CDN" | "Events arrive in bursts. Why?" | Proxy buffering; flush and no-buffer headers |
| "Mobile clients too" | "The app is backgrounded." | OS kills sockets; fall back to push notifications |
What to Say in the Interview#
"The deciding question is direction. Order tracking only flows down, so SSE: it's plain HTTP, the browser reconnects on its own, and Last-Event-ID gives us resume with a short replay buffer."
"I'd use WebSockets for the chat screen, where typing indicators and messages go up several times a second, and I'll write the resume protocol explicitly: sequence numbers per conversation and a replay on reconnect."
"Whatever the protocol, the hard part is transitions: gateways drain 1% of connections every few seconds on deploy, clients back off with jitter up to 30 seconds, and the edge rate-limits handshakes."
"If product is happy with 15-second freshness, I'd skip all of this and poll with ETags — it's cacheable, stateless and the cheapest thing to operate."
Related Guides#
- Push vs Poll — the full pattern: registries, fan-out, resume and reconnect storms
- Design Live Updates over WebSockets — the end-to-end design problem
- Design a Chat App — the canonical WebSocket workload
- Latency, Protocols & Tail Amplification — HTTP/1.1, HTTP/2, HTTP/3 and what they change
- Design a Notification Platform — the offline half when no connection exists
- Design a Collaborative Document Editor — high-frequency two-way updates
- Envoy, Kong & NGINX — proxies, idle timeouts and buffering
- Backpressure — slow consumers on the gateway tier