Hiring BarSupport

WebSockets vs SSE vs Long Polling

Comparison16 min read3 diagrams

Every "make it real-time" requirement ends up here, and most candidates answer "WebSockets" before the question is finished. All three let a server deliver data to a browser or app without the user refreshing. They differ in direction and in how much of the web's infrastructure they keep. Long polling is ordinary HTTP requests that the server holds open until it has something to say. Server-Sent Events (SSE) is one long HTTP response that streams events from server to client, with reconnect and resume built into the browser. WebSockets upgrade an HTTP connection into a full-duplex, message-framed socket that leaves HTTP semantics behind. The question that decides it: does the client need to send frequent messages back on the same channel, or does data flow only from server to client? Most real-time features — feeds, notifications, dashboards, order tracking, live scores — are one-way, and one-way is SSE's job.

The Verdict#

Default to SSE for server-to-client streams, WebSockets when the client sends often (chat typing, cursors, games), long polling only as a fallback or for low-frequency change notifications — and plain polling when 10+ seconds of delay is acceptable.

Pick SSE whenPick WebSockets whenPick long polling when
Data flows server → client: feeds, notifications, dashboards, progress, LLM token streamsClient sends frequently on the same channel: chat with typing indicators, collaborative cursors, multiplayer gamesUpdates are infrequent (minutes apart) and you want plain request/response everywhere
You want auto-reconnect and resume (Last-Event-ID) for freeYou need binary frames or very low per-message overhead in both directionsSome clients sit behind proxies that break streaming responses
You want standard HTTP: auth headers on the server side, compression, HTTP/2 multiplexing, existing load balancersLatency both ways matters at < 100 ms and messages are small and frequentYou are building a public API for third parties who can't hold streams (a common pattern for "wait for changes" endpoints)
Client → server messages are rare and can be normal HTTP POSTsYou control both ends (your app + your gateway)You need a fallback when SSE or WebSockets fail

🎯 Staff Move: "Before I pick a protocol: does the client talk back on this channel? Here it doesn't — order status and courier location only flow down — so I'll use SSE with event IDs for resume, and the rare client actions go over normal HTTPS. I'd switch to WebSockets for the chat screen, where typing indicators go up several times a second."

Diagram: The Verdict

At a Glance#

DimensionWebSocketsSSELong polling
DirectionFull duplexServer → client onlyServer → client (client sends via normal requests)
TransportHTTP Upgrade to the WebSocket protocol (RFC 6455); also over HTTP/2 (RFC 8441) and HTTP/3 (RFC 9220)One long-lived HTTP response, Content-Type: text/event-streamRepeated HTTP requests, each held until data or timeout
Message formatText or binary framesUTF-8 text events (id:, event:, data:, retry:)Whatever the response body is
OrderingIn order on one connectionIn order on one connectionIn order per response; gaps possible between requests unless you use a cursor
Reconnect / resumeBuild it yourselfBuilt into browsers: auto-reconnect, sends Last-Event-IDNatural: every request carries the client's cursor
LatencyLowest both ways (one frame, ~2–14 bytes overhead)Same as WebSockets for server → clientOne round trip after each event; bursts are delivered in batches
Per-message overheadTinyTiny (a few bytes of field names)Full HTTP request + response headers per batch (~0.5–2 KB)
Server state per clientOne open connection (~10–50 KB memory)One open connection (~10–50 KB)One parked request at a time; cheaper between requests
Proxies and load balancersNeed explicit Upgrade support and long idle timeoutsWorks through most HTTP infrastructure; buffering proxies must be told not to bufferWorks everywhere HTTP works
Browser limitsNo per-origin cap like HTTP/1.1'sCounts against ~6 connections per origin on HTTP/1.1; multiplexed on HTTP/2Same per-origin cap as any request
Scaling modelStateful gateways, connection registry, pub/sub fan-outSame as WebSocketsStateless-ish servers; state lives in the cursor and the backing store
Operational burdenHighest: deploy drains, reconnect storms, custom heartbeat and resumeMedium: same gateways, but resume and reconnect are standardLowest per server; request rate can be high
Managed optionsAPI Gateway WebSocket APIs, Cloudflare, Azure Web PubSub, many hosted real-time servicesAny HTTP host or CDN that supports streaming responsesAny HTTP host
Cost shapeConnection-minutes + messagesConnection-minutes + bytesRequests

How They Actually Differ#

Direction and What the Client Really Sends#

The usual argument for WebSockets is "it's bidirectional". The question is how often the client sends, and whether those sends need to be on the same channel. A news feed client sends nothing between page loads. A notifications client sends "mark as read" a few times a session — a normal POST is fine. A chat client sends typing indicators several times a second and messages that must arrive in order with the server's events. A multiplayer game sends input 20–60 times a second.

FeatureClient → server rateBest fit
News feed "new posts" banner~0SSE (or polling every 30s)
Order tracking~0SSE
Notifications with read receiptsA few per sessionSSE + HTTP POST
AI assistant streaming tokensOne prompt per responseSSE (one streamed response per request)
Chat with typing indicatorsSeveral per second while typingWebSockets
Collaborative cursors / editing10–30 per secondWebSockets
Multiplayer game input20–60 per secondWebSockets (or UDP-based transports in native clients)

Who pays: choosing WebSockets for one-way data makes the platform team pay for custom reconnect, resume and heartbeat logic that SSE gives for free. Choosing SSE for a high-frequency two-way feature makes the client team pay for a second channel and for ordering between the two.

Reconnect and Resume: Built In vs Build It#

Connections drop constantly: phones change networks, laptops sleep, load balancers time out idle connections, deploys restart gateways. What happens next is the real design.

SSE's browser EventSource reconnects automatically, waits the server-suggested retry: interval, and sends the last event ID it saw in a Last-Event-ID header. If the server keeps a short replay buffer keyed by event ID, resume is a few lines of server code.

WebSockets have no reconnect and no resume. The client library must detect the drop (ping/pong or application heartbeats), back off with jitter, reconnect, re-authenticate, re-subscribe, and send its last sequence number so the server can replay — all custom.

Long polling resumes by construction: every request carries a cursor (?since=evt_9812), and the server answers with everything after it.

Diagram: Reconnect and Resume: Built In vs Build It

What each looks like on the wire:

SSE response (one HTTP response that never ends)
  HTTP/1.1 200 OK
  Content-Type: text/event-stream
  Cache-Control: no-cache

  retry: 5000                       ← client waits 5 s before reconnecting
  id: 9813
  event: status
  data: {"order":42,"state":"picked_up"}

  : keepalive                       ← comment line every ~20 s defeats idle timeouts

Long polling (one request per batch)
  GET /orders/42/changes?cursor=9812&wait=30
  → held up to 30 s → 200 {"events":[...], "cursor":"9814"}   or   200 {"events":[], "cursor":"9812"}
  client immediately re-requests with the new cursor (plus jitter after an empty response)

WebSocket client resume (all of it is yours to write)
  on open:     send {"type":"auth","token":...}; send {"type":"resume","since":last_seq}
  on message:  if msg.seq != last_seq + 1 → request replay; last_seq = msg.seq
  every 25 s:  send ping; if no pong in 10 s → close and reconnect
  on close:    wait random(0, min(30 s, 1 s × 2^attempt)); reconnect

Infrastructure Fit#

SSE is an HTTP response that never ends. It uses the same authentication middleware, the same TLS termination, the same HTTP/2 connection the page already has, and compression where it helps. The traps are buffering proxies (which hold the response until it ends — disable buffering, for example with X-Accel-Buffering: no on NGINX) and idle timeouts (send a comment line every ~15–30 s). One limitation: the browser EventSource API cannot set custom request headers, so auth uses cookies or a short-lived token in the URL — or you use a fetch-based SSE client.

WebSockets need every hop to understand the Upgrade and to tolerate hours-long connections: load balancer idle timeouts (often 60 s by default — heartbeat under that), corporate proxies that strip Upgrade headers, and WAF rules that inspect HTTP but not frames. After the upgrade, HTTP tooling (status codes, caching, per-request logs) no longer applies.

Long polling is ordinary HTTP. Every proxy, CDN and firewall handles it. The cost is request volume: 1M clients with 30-second holds is ~33K requests/s at minimum, more when events are frequent.

Fan-out and Server Cost#

For WebSockets and SSE the server architecture is the same: a fleet of stateful gateways holding connections, a connection registry (user → gateway), and a pub/sub bus that routes events to the right gateway. A tuned gateway holds ~100K–500K mostly idle connections; at ~10–50 KB per connection, 1M connections is ~10–50 GB of memory across the fleet. The protocol choice changes almost nothing about this tier — which is why picking WebSockets "for scale" is a non-argument.

Long polling changes the shape: servers can be closer to stateless (a parked request waits on a subscription; when it returns, the server forgets the client until the next request). It pays in request rate and in latency for bursts: three events 10 ms apart may arrive as one response plus a reconnect.

1M connected clients, 1 event per client per minuteWebSocketsSSELong polling (30 s hold)
Open connections1M1M~1M parked requests
Requests/s~0 (after connect)~0 (after connect)~25K–35K/s (timeouts plus event responses)
Bytes per event~100 B payload + ~6 B framing~100 B + ~20 B field names~100 B + ~800 B headers
Extra workCustom resume, heartbeatReplay buffer for Last-Event-IDCursor per request

Where Each One Breaks#

WebSockets#

FailureWhat happensDetectionOwner
Deploy reconnect stormRestarting a gateway drops 250K sockets; all reconnect in the same second; auth and registry overloadHandshakes/s, auth latencyPlatform
Silent half-open connectionsNAT or proxy drops the path without a close; server thinks the client is connected and keeps queuingHeartbeat failures, messages sent without acksPlatform
Gap after reconnectMessages published during the reconnect window never reach the clientClient-side sequence gap counterPlatform + client team
Slow consumerA client on a bad network can't drain; server-side buffer grows until memory pressurePer-connection send queue sizePlatform
Middlebox breaks UpgradeSome enterprise networks block or strip Upgrade; connection fails or falls backConnect failure rate by network/ASNClient team

SSE#

FailureWhat happensDetectionOwner
Buffering proxyEvents arrive in bursts minutes late or only when the connection closesTime from publish to client receiptPlatform
HTTP/1.1 connection capSeveral tabs each open a stream; the browser's ~6-per-origin limit blocks other requestsStalled requests in browser telemetryClient team (use HTTP/2, or share one stream across tabs)
Idle timeoutQuiet streams get cut by the LB; clients reconnect constantlyReconnect rate, stream duration histogramPlatform (heartbeat comments)
Replay buffer too shortClient returns after 20 minutes with a Last-Event-ID older than the bufferResume misses → full resync ratePlatform
Reconnect stormSame as WebSockets on deploys — auto-reconnect makes it more synchronised, not lessHandshakes/sPlatform (server retry: with jitter, drain slowly)

Long Polling#

FailureWhat happensDetectionOwner
Request flood after a broadcastOne event wakes 1M parked requests; all return and immediately re-requestRequest rate spikes in lockstepPlatform (jittered re-poll, server backoff hints)
Burst latencyEvents closer together than the round trip are batched; per-event latency growsPublish-to-receipt p99Platform
Timeouts shorter than holdProxy cuts at 30 s but server holds 60 s; clients see errors every cycle504 ratePlatform
Lost events between pollsServer forgets what a client saw; events between responses are missedGaps in client cursorsPlatform (always use a cursor)

A Gateway Deploy, Done Badly and Done Well#

Done badly (restart all gateways at once, clients reconnect immediately)
t=0       8 gateways restart → 2M connections drop
t=+1s     2M reconnect attempts; TLS handshakes + auth checks saturate the edge
t=+5s     auth service p99 50ms → 4s; reconnects time out and retry → more load
t=+60s    still recovering; events published during the window lost for clients without resume

Done well (drain 1% per 5 s, server-sent retry with jitter, resume from replay buffer)
t=0       gateway 1 stops accepting new connections; closes 1% of its sockets every 5 s
t=+0..8m  ~4K reconnects/s fleet-wide, spread by client jitter; auth p99 unchanged
t=+8m     all gateways rolled; clients resumed from Last-Event-ID / last seq; zero gaps

Who pays: done badly, users and the auth team pay. Done well, the platform team paid up front for drain tooling and a resume protocol.

The production surprise for all three: the steady state is easy; transitions are the system. Deploys, network changes and broadcasts produce synchronized reconnects that look like a DDoS. Jittered backoff (1 s → 30 s cap), slow gateway drains (1% of connections per few seconds) and a handshake rate limit at the edge matter more than the protocol choice.


Cost and Operations#

WebSocketsSSELong polling
Who runs itA real-time platform team owns gateways, registry, client SDKSame team; smaller client SDKAny backend team
Bill scales withConcurrent connections (memory) + messages + egressConcurrent connections + egressRequests + egress
Managed pricing shapeTypically per connection-minute + per message on hosted gatewaysWhatever your HTTP hosting charges for long requestsPer request
Hidden costCustom resume protocol, client SDKs for every platform, deploy toolingReplay buffer, proxy configurationHeader overhead and request rate at scale
People1–3 engineers for a large real-time platformSimilar platform team, less client workLittle beyond normal API ownership
What on-call watchesWebSockets / SSELong polling
The SLO metricPublish-to-client latency p99; connected clientsPublish-to-client latency p99; request success rate
The "act now" alertHandshakes/s spike, gateway memory > 80%, resync rate jumpRequest rate spike in lockstep, 5xx on the poll endpoint
Routine workDeploy drains, SDK releases, idle-timeout and heartbeat tuningHold-time and backoff tuning

Sizing example. 2M concurrent clients at peak:

Connections:  2M × ~30 KB ≈ 60 GB memory across gateways
Gateways:     at ~250K connections each with headroom → 8–10 nodes per region
Heartbeats:   every 25 s to stay under 60 s idle timeouts → 80K small frames/s fleet-wide
Deploy:       drain 1% per 5 s → a full fleet roll takes ~8 minutes, peak ~4K reconnects/s
Long-poll alt: 2M / 30 s hold ≈ 67K req/s baseline before any events — fine for an API tier,
              but every one carries headers and auth checks

🧭 Principal Insight: "The cost of real-time isn't the protocol, it's the gateway tier and the client SDKs. Whatever we choose, I want one connection platform the whole company uses — one registry, one resume protocol, one deploy procedure — rather than each product team running its own WebSocket servers."


Switching Later#

MoveHowWhat's hard
Long polling → SSESame cursor becomes the event ID; the endpoint streams instead of returningProxy buffering and timeouts on the path
SSE → WebSocketsGateway accepts both; clients upgrade; resume logic ported to the custom protocolRebuilding reconnect/resume that EventSource gave for free
WebSockets → SSESplit client sends into normal HTTP; stream down over SSEFeatures that relied on ordered client→server messages on one channel
Any → a managed real-time serviceGateway replaced by the vendor; your services publish to itVendor's message model, limits and pricing; data residency

One-way doors: the client resume protocol shipped in mobile apps (old app versions live for years); event ID / sequence semantics; the decision to run your own gateway tier. Two-way doors: the transport for web clients (servers can speak several); heartbeat intervals; replay buffer length.

Diagram: Switching Later

How Real Companies Chose#

Slack: Persistent WebSockets for a Two-Way Product#

Slack's engineering blog describes every client holding a persistent WebSocket connection to receive real-time events, through stateful gateway servers that subscribe to channel servers (mapped by consistent hashing). It reports tens of millions of connected clients and messages delivered across the world in about 500 ms (Slack Engineering).

Staff insight: Chat is the canonical two-way case — presence, typing and message events in both directions — and the architecture's real work is the gateway and channel-server tiers, not the socket.

Shopify: SSE for a One-Way Live Map#

For its Black Friday Cyber Monday live map, Shopify replaced a polling design (with at least a 10-second delay) with a Go SSE server that subscribes to Kafka, scaled horizontally behind NGINX. They chose SSE over WebSockets explicitly because the data only flows server → client, and cited HTTP familiarity and automatic reconnection; the new pipeline put data on screen within about 21 seconds of creation with 100% uptime during the event (Shopify Engineering).

Staff insight: A high-traffic, high-visibility real-time product chose the simpler protocol because the requirement was one-way. That is the default this page argues for.

Dropbox: Long Polling as a Public API Contract#

Dropbox's public API offers a list_folder/longpoll endpoint: the client passes a cursor and the request blocks until changes are available or a timeout (30–480 seconds, default 30) passes, with up to 90 seconds of random jitter added to avoid a thundering herd; the response may include a backoff field telling the client how long to wait before polling again (Dropbox API spec).

Staff insight: For third-party developers on every kind of network, long polling with a cursor is the most compatible push mechanism there is — and the server-controlled jitter and backoff fields are how you keep a million clients from synchronising.


Follow-Ups to Expect#

After You Say...They Will Ask...What They're Testing
"WebSockets""Does the client actually send anything?"Whether you chose by direction or by habit
"SSE""How does the client resume after a network drop?"Last-Event-ID, replay buffer, full-resync fallback
"A gateway tier""You deploy the gateways. What happens to 2M connections?"Slow drains, jittered reconnect, handshake limits
"Long polling""One event goes to 1M users at once."Lockstep re-poll storm; jitter; server-sent backoff
"Heartbeats""Why every 25 seconds?"Load balancer and NAT idle timeouts (~60 s)
"SSE through our CDN""Events arrive in bursts. Why?"Proxy buffering; flush and no-buffer headers
"Mobile clients too""The app is backgrounded."OS kills sockets; fall back to push notifications

What to Say in the Interview#

"The deciding question is direction. Order tracking only flows down, so SSE: it's plain HTTP, the browser reconnects on its own, and Last-Event-ID gives us resume with a short replay buffer."

"I'd use WebSockets for the chat screen, where typing indicators and messages go up several times a second, and I'll write the resume protocol explicitly: sequence numbers per conversation and a replay on reconnect."

"Whatever the protocol, the hard part is transitions: gateways drain 1% of connections every few seconds on deploy, clients back off with jitter up to 30 seconds, and the edge rate-limits handshakes."

"If product is happy with 15-second freshness, I'd skip all of this and poll with ETags — it's cacheable, stateless and the cheapest thing to operate."


  1. Loading the index…