Hiring BarSupport

Design WhatsApp (Chat Messaging) — Staff-Level Case Study

Case study64 min read6 diagrams

Technologies referenced in this case study: Cassandra · Redis · Apache Kafka · DynamoDB

Related case studies: Real-Time Updates · Notification System · Message Queues · Blob Storage · File Sync · Real-time Updates pattern

How to Use This Case Study#

Organized for interview use first, reference second. Read once end to end, then return to weak spots.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Active Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failures) → Deep Dives 2 and 5
Deep Dive3+ hrsEverything, including the Principal Lens and appendices
What is Chat Messaging? — Why interviewers pick this topic

Person-to-person and small-group messaging with delivery guarantees: messages sent while the recipient is offline are held and delivered later, in order, exactly once from the user's point of view, with sent/delivered/read receipts, across multiple devices per user — and, in WhatsApp's case, end-to-end encrypted so the server cannot read them.

Before vs After — a phone comes back online after a flight:

Naive design (server-side history table, client polls, server-assigned timestamps):
t=0:       Phone reconnects after 9 hours. 1,400 messages across 60 chats waiting.
t=+1s:     Client asks "give me everything since my last timestamp." Clock skew between
           two chat servers means 30 messages have timestamps earlier than the client's
           last-seen value. They are never delivered.
t=+3s:     Server pushes 1,400 messages; the connection drops at message 900. Client
           reconnects and re-requests from its last timestamp: 900 duplicates render.
t=+5s:     Group messages appear out of order — replies before the questions.
t=+10s:    Sender's phone still shows one grey tick for messages delivered hours ago,
           because receipts were sent to a device that has since been replaced.

Staff design (per-device inbox with sequence numbers, client IDs, acks):
t=0:       Same reconnect.
t=+1s:     Client says "my inbox cursor is 88,214." Server streams 88,215 onward in
           pages of 100. Order within each chat follows the chat's sequence number.
t=+3s:     Connection drops at 900. Client acked through 88,900 in batches; resumes
           from 88,901. Client-generated message IDs dedupe anything re-sent.
t=+4s:     Server deletes acked messages from the inbox (it is a relay, not an archive).
t=+5s:     Delivery receipts flow back to senders' devices as their own inbox entries.

Why interviewers reach for this question: it looks like a WebSocket question. It is actually a delivery-semantics and state-ownership question: where does the durable copy of a message live, who assigns order, how do acknowledgments compose into user-visible ticks, and what happens to all of that when the server is not allowed to read the content (E2E) and a user has five devices.

Mechanics Refresher: Delivery Models
ModelHow It WorksProsCons
PollingClient asks for new messages every N secondsSimple; stateless serversLatency = N/2; battery and request cost scale with users, not messages
Long pollingRequest held open until a message arrives or ~30–60 s timeoutWorks through most proxiesReconnect churn; one message per round trip
Persistent connection (WebSocket, MQTT, custom TCP/TLS)Server pushes over a held socket; heartbeats keep NAT openSub-100 ms push; efficientStateful edge fleet; reconnect storms; connection-to-server routing
OS push (APNs/FCM)Platform wakes the app or shows a notificationWorks when app is backgrounded/killedBest-effort, rate-limited, payload-limited (~4 KB), no ordering
Store-and-forward inboxServer durably queues per recipient until ackedOffline delivery; exactly-once display via dedupeStorage for undelivered backlog; cursor management

For most production systems: persistent connection when foregrounded, OS push to wake when backgrounded, and a durable per-device inbox underneath both. The transport is almost never the interview question — ordering, acks, the inbox, groups and multi-device under E2E are.


Executive Summary

If you only read one section, read this. Everything else in the case study elaborates on the contrast and the positions below.

What This Interview Actually Tests#

Chat is not a WebSocket question. Everyone can hold a socket open.

It is a store-and-forward delivery problem with per-conversation ordering, end-to-end acknowledgments, and — under E2E encryption — a server that must route what it cannot read. It tests:

  • Whether you define delivery semantics precisely: at-least-once transport + client dedupe = exactly-once display
  • Whether you know who assigns order (a per-chat sequencer) and why wall-clock timestamps fail
  • Whether you treat the server as a relay with an inbox, not an archive, and know which intent that implies
  • Whether you understand what E2E encryption moves to the client: group fan-out, multi-device sync, search, backups, abuse detection

The key insight: The server's job in WhatsApp-style chat is to hold an encrypted envelope until every intended device has acknowledged it, then forget it. Every hard design question — ordering, receipts, groups, multi-device — is a question about who holds which cursor and who can read what.

The L5 vs L6 vs L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveWebSocket servers + a messages tableAsks "Is the server the system of record or a relay? Is it E2E?" and commitsAsks what the E2E decision does to every adjacent product — backups, search, moderation, compliance — and who signs that
Delivery"Use a queue so messages aren't lost"At-least-once to a per-device inbox; client-generated message IDs; acks advance a cursor; exactly-once displayDefines delivery SLOs (p99 online delivery, max offline retention) as contracts with product and legal
OrderingSort by timestampPer-chat monotonic sequence from the chat's owner shard; clients order by (chat, seq)Chooses the ordering guarantee per product surface and documents where the company does not promise order
GroupsFan-out on write to all membersFan-out on write for small groups (≤ ~1K), sender-key encryption so the sender encrypts once; caps as design constraintsGroup size caps as product policy with cost and abuse modeling; large broadcast = a different product
Multi-device"Sync messages to all devices"Each device is a first-class endpoint with its own keys and inbox; sender encrypts per deviceDevice-linking security model, history-transfer policy, and key-change UX signed by security and product
StorageKeep all messages foreverDelete after delivery; bounded offline retention (~30 days)Data-minimization as a company position; what that means for law-enforcement requests and for revenue features
Why "first move" separates levels

L5: Designs Slack and calls it WhatsApp: a messages table keyed by conversation, history served from the server, search over server data. That is a valid design for a different intent — server-as-record chat — and it silently assumes the server can read messages.

L6: "There are two very different chat systems. In one, the server is the system of record — history, search, compliance — like Slack or Discord. In the other, the server is a relay — it holds encrypted envelopes until delivered and then deletes them — like WhatsApp or Signal. I'll design the relay with E2E, because that's what 'WhatsApp' implies, and it changes where groups, sync and search live."

L7: "Choosing E2E is a company-level decision: it removes server-side search, server-side spam filtering on content, and easy cross-device history. I'd want Security, Legal, Trust & Safety and Product aligned on that before engineering commits, because it's a one-way door in public perception."

Why "ordering" separates levels

L5: "Each message gets a timestamp; clients sort by timestamp." Server clocks skew by milliseconds to seconds; client clocks by minutes. Two messages sent 5 ms apart in a group can render in either order on different devices, and a skewed server can assign a timestamp earlier than a client's sync cursor so the message is never fetched.

L6: "Order is assigned by the chat's owner — one sequencer per conversation, a monotonic counter stored with the chat. Clients render by (chat_id, seq). Timestamps are display-only. For a 1:1 chat under E2E the server still sequences envelopes, because it can order what it can't read."

L7: Decides where the product explicitly does not promise order — e.g., across different chats, or between a message and a reaction — so teams stop building features that assume it.

Why "multi-device" separates levels

L5: "The server keeps messages and each device syncs." Under E2E, the server cannot re-encrypt for a new device.

L6: "Every device has its own identity key. A sender encrypts a message once per recipient device — and once per their own other devices — so a 1:1 message to someone with 4 devices, sent from a user with 3, is ~6 ciphertexts. Each device has its own inbox and its own cursor. History on a new device is a client-to-client transfer, not a server replay." WhatsApp's 2021 multi-device architecture moved from "phone as the relay for web clients" to exactly this per-device model.

L7: Owns the security model for linking devices (QR-based linking, key-change notifications, max linked devices) as a policy that security and product sign, and prices the extra fan-out it creates.

The Staff Positions#

PositionRationale
Server is a relay, not an archive (for E2E consumer chat)Delete on ack; bounded offline retention; minimizes cost, breach surface and legal exposure
Per-device inbox with a monotonic cursorOffline sync is "give me everything after cursor N"; resume after disconnect is exact
Client-generated message IDsRetries are safe; dedupe at recipient; exactly-once display on at-least-once transport
Per-chat sequencer for orderTimestamps lie; one owner per chat assigns seq
Fan-out on write for groups up to a capRecipients read one inbox; cap (~1K) bounds write amplification; broadcast channels are a different product
Receipts are messagesDelivered/read receipts travel through the same inbox path in reverse; no separate receipt database
Connection gateways are stateless about messagesGateways hold sockets; message state lives in inboxes; a gateway crash loses no data

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Private E2E messaging (WhatsApp, Signal)Server can't read content; offline delivery; small groupsRelay with per-device inboxes; client-side encryption and fan-out; delete on ackLost message on inbox failure; out-of-order delivery; key mismatch after device changeEvery message delivered to every device exactly once (displayed), in per-chat order
Community / workspace chat (Slack, Discord)History, search, huge channels, integrationsServer-side history as system of record; fan-out on read for large channelsHot channels; search index lag; permission leaksHistory durable forever; order per channel
Regulated / enterprise messaging (finance, healthcare)Retention, eDiscovery, legal holdServer-readable storage, immutable archive, compliance exportRetention policy violation; unauthorized accessNothing deleted before retention; auditable

🎯 Staff Move: "I'll design private E2E messaging — WhatsApp's intent. The server is a relay that holds encrypted envelopes per device until they're acked. That choice pushes group fan-out encryption, multi-device sync and search to the client, and I'll call out each place it does. If you want Slack-style history and search, that's a different system and I'd design the storage completely differently."

The Five Fault Lines#

#Fault LineThe Tension
1Relay vs System of RecordDelete after delivery (cheap, private) or keep history server-side (search, sync, compliance)?
2Ordering: Sequencer vs Timestamps vs CausalWho assigns order, and what does it cost per message and on failover?
3Delivery Semantics and ReceiptsAt-least-once + dedupe vs attempts at exactly-once; how acks compose into sent/delivered/read
4Group Fan-Out: Write vs Read, and E2E CostPer-recipient inbox writes vs shared log; pairwise encryption vs sender keys; where the cap sits
5Multi-Device: Primary-Relay vs Per-Device EndpointsPhone forwards to companions, or every device is a peer with its own keys and inbox?

In the Wild: Real Production Systems#

Why this section belongs here: these are well-known, publicly documented systems; citing them shows you understand how production chat actually evolved.

WhatsApp — Erlang Relay, Signal Protocol, Multi-Device#

WhatsApp's backend is famously built on Erlang/BEAM (on FreeBSD in its early years), and its engineers publicly described handling ~2 million concurrent TCP connections on a single server in 2012. WhatsApp completed rollout of end-to-end encryption using the Signal Protocol in 2016, and in 2021 published its multi-device architecture, in which each linked device has its own identity key and senders encrypt per device, so companions no longer depend on the phone being online. WhatsApp states that undelivered messages are deleted from its servers after 30 days. It was also widely reported to have served several hundred million users with a team of ~50 engineers at the time of its 2014 acquisition.

Staff insight: WhatsApp's design is a lesson in doing less on the server: a relay that holds encrypted envelopes, a runtime built for millions of lightweight connection processes, and complexity pushed into clients where E2E requires it. Say "the server is a relay" and you've made the most important decision.

Discord — Server-Side History at Trillion-Message Scale#

Discord has written publicly about storing messages in Cassandra and later migrating to ScyllaDB, with messages partitioned by (channel_id, time bucket) and ordered by Snowflake IDs (time-sortable 64-bit IDs). Their posts describe hot partitions from very large, busy channels and the operational pain of compaction and GC pauses that motivated the move.

Staff insight: This is the system-of-record intent. The partition key embeds a time bucket because unbounded channel partitions become hot and huge. Contrast it with WhatsApp: same product category, opposite storage philosophy, because the intent is different.

Facebook Messenger — Iris, a Per-User Ordered Update Queue#

Facebook Engineering described Iris (2014) for Messenger: a totally ordered queue of messaging updates per user, with each device keeping a pointer (sequence ID) into it. Devices sync by asking for everything after their pointer, which made mobile sync fast and reliable over lossy networks.

Staff insight: "Per-user ordered queue + per-device cursor" is the generalizable pattern. It turns sync into one integer comparison and makes resume-after-disconnect exact. Name it in the interview.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"WebSockets for real-time""The recipient is offline for 3 days. Where is the message?"Store-and-forward thinking
"Messages have timestamps""Two people send at the same millisecond in a group. What order does everyone see?"Sequencer ownership
"We guarantee exactly-once delivery""The ack is lost after the client stored it. What happens?"Precise semantics
"Fan-out to group members""The group has 1,000 members, each with 3 devices, and it's E2E."Write amplification + encryption cost
"Sync to all devices""The server can't decrypt. How does a new laptop get history?"E2E implications
"Store messages in Cassandra""For how long? Why?"Relay vs archive

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: clients own encryption; the core routes opaque envelopes. The Router asks the chat's Sequencer for the next seq, looks up recipient devices (1:1 or group), and appends one entry per device to that device's inbox before acknowledging the sender. If the device is online (per the session registry), the gateway pushes immediately; otherwise the push sender wakes the app. The inbox is the only durable message state, and entries are deleted when the device acks.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Transport"WebSockets""Persistent TLS socket in foreground, OS push to wake, per-device inbox underneath both."
Durability"Store all messages""Durable per-device inbox; the sender's single tick means 'in the inbox, replicated'; delete on device ack."
Semantics"Exactly-once""At-least-once transport, client-generated IDs, dedupe on receipt: exactly-once display."
Ordering"Timestamps""Per-chat sequencer; render by (chat, seq); timestamps for display only."
Receipts"Update a status column""Receipts are small messages sent back through the sender's inbox; ticks are a client-side fold."
Groups"Fan out""Fan-out on write to member devices up to ~1K members; sender keys so the payload is encrypted once."
Multi-device"Sync""Each device is an endpoint with its own keys and inbox; sender encrypts per device; history transfer is client-to-client."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Messages/day (WhatsApp, public order of magnitude, 2020)~100B≈ 1.2M/s average, ~3–5M/s peak (New Year's Eve)
Concurrent connections per server (WhatsApp 2012, public)~2MErlang lightweight processes; modern planning figure ~0.5–1M per host
Heartbeat interval~30–60 s (under typical mobile NAT timeouts)Connection-keepalive cost: 1B devices / 45 s ≈ 22M heartbeats/s
Undelivered message retention30 days (WhatsApp public FAQ)Bounds inbox storage
Group size cap1,024 (WhatsApp, raised from 512 in 2022)Bounds fan-out and E2E cost
Linked companion devicesup to 4 (plus the phone)Per-device fan-out multiplier
Envelope size (text)~200 B–1 KB1.2M/s × 1 KB ≈ 1.2 GB/s inbox write bandwidth before replication
MediaEncrypted blob uploaded once; message carries URL + key (~200 B)Media never flows through the message path
OS push payload limit~4 KBPush carries a wake signal, not the message
Online-to-online deliveryp99 < ~300–500 ms in-regionThe "instant" feel
Inbox entries per group messagemembers × devices (1K × ~2 = ~2K)Write amplification you must budget

Interview Walkthrough

The most common mistake: candidates spend 15 minutes on WebSocket handshakes and load balancer stickiness, then hand-wave "and messages are stored in Cassandra." The transport is two sentences. The phases below get you to the inbox, ordering and E2E by minute 12.


Phase 1: Requirements & Framing (2–3 min)#

Functional core in one breath:

"1:1 and group messages with text and media, delivered in real time when online and stored until delivered when offline, with sent, delivered and read receipts, on multiple devices per user."

Then the framing question that decides the design:

"The big fork: is the server the system of record — history, search, compliance — or a relay that holds messages until delivered? For WhatsApp it's a relay with end-to-end encryption. I'll design that, and I'll call out what E2E pushes to the client."

Then numbers:

"~2B users, ~500M–1B concurrently connected at peak, ~100B messages/day — ~1.2M/s average, ~3–5M/s at peak. Average group message fans out to ~5–10 recipient devices, so inbox writes are ~5–10M/s at peak. Online delivery p99 under ~500 ms in-region. Offline retention 30 days."

🎯 Staff Move: The relay-vs-record question in the first 60 seconds is the level signal. It tells the interviewer you know there are two chat systems and you are choosing one on purpose.


Phase 2: Core Entities & API (1–2 min)#

  • Device: device_id, user_id, identity_public_key, prekeys[], push_token
  • Chat: chat_id (1:1 = hash of sorted user IDs; group = generated), last_seq, owner_shard
  • Envelope: msg_id (client-generated UUIDv7), chat_id, seq, sender_device, recipient_device, ciphertext, type ∈ {message, receipt, key_change, membership}, server_ts
  • InboxEntry: (device_id, inbox_seq) → envelope; deleted on ack
  • GroupMembership: group_id → member user IDs → their device IDs; version

Socket frames (binary, compact):

→ SEND   {msg_id, chat_id, envelopes: [{recipient_device, ciphertext}], membership_version}
← ACK    {msg_id, seq, server_ts}                 # single tick: durably in all inboxes
← DELIVER{inbox_seq, envelope}                    # pushed or fetched
→ RECV_ACK {up_to_inbox_seq}                      # cursor advance; server deletes ≤ cursor
→ SYNC   {after_inbox_seq, limit: 100}            # reconnect / backfill
→ RECEIPT{msg_id, chat_id, kind: delivered|read}  # travels back as an envelope

🎯 Staff Move: "Notice the sender sends one ciphertext per recipient device — the server can't fan out plaintext because it doesn't have any. And the recipient acks a cursor, not individual messages, so one ack covers a batch of 100."


Phase 3: High-Level Architecture (≤ 5 min)#

Diagram: Phase 3: High-Level Architecture (≤ 5 min)

Three sentences:

"Gateways hold sockets and nothing else. The Router gets a sequence number from the chat's sequencer, appends one inbox entry per recipient device, acks the sender only after those appends are durable, then pushes to online devices or wakes offline ones through APNs/FCM. Devices ack a cursor; the server deletes everything up to it."


Phase 4: Transition to Depth#

"Transport is standard — persistent socket plus OS push. The parts I'd want to go deep on are: how the inbox and acks give exactly-once display, how ordering works in groups, and what E2E does to groups and multi-device. I'd start with the inbox because everything else builds on it."


Phase 5: Deep Dives (25–30 min)#

Deep DiveTimeWhat You Must Land
Inbox, acks, exactly-once display6–8 minDurable append before sender ack; client msg IDs; cursor acks; delete on ack; 30-day bound
Ordering4–6 minPer-chat sequencer; seq gaps; failover of the sequencer
Receipts3–4 minReceipts as envelopes; ticks as a fold; read-receipt privacy setting
Groups under E2E6–8 minFan-out on write per device; sender keys; membership versions; cap
Multi-device4–6 minPer-device keys and inboxes; linking; history transfer
Connection edge3–4 minSession registry, heartbeats, reconnect storms

Exactly-once display, as you'd say it:

"The sender's client generates msg_id before sending and retries with the same ID until it gets an ACK. The Router dedupes by (sender_device, msg_id) for a short window so retries don't create duplicate inbox entries. If a duplicate still slips through — say, an ACK was lost after the append — the recipient's local database has a unique constraint on msg_id and drops it. Transport is at-least-once; display is exactly-once."


Phase 6: Wrap-Up (2–3 min)#

"To summarize: a relay with durable per-device inboxes, at-least-once transport with client IDs for exactly-once display, a per-chat sequencer for order, receipts as messages, fan-out on write with sender keys for groups up to ~1K, and per-device keys for multi-device. What I'd monitor first: online delivery p99, inbox backlog age, and reconnect rate per gateway. What I'd build next: client-to-client history transfer for new devices and encrypted backups. What I skipped: channels/broadcast to millions, which is a fan-out-on-read product."


Common Timing Mistakes#

MistakeTime LostFix
WebSocket vs long-poll vs SSE debate5–8 min"Persistent socket + OS push. Moving on."
Designing Slack history and search10+ minCommit to relay vs record in Phase 1
Signal Protocol cryptography details (X3DH, double ratchet math)8 min"Sessions are established from prekeys; each message advances a ratchet. What matters for the system is per-device fan-out."
Presence ("online/last seen") in depth5 min"Presence is lossy, best-effort, subscribed per open chat."
No numbersWhole interview100B/day, 1.2M/s, ×devices in Phase 1

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Chat is everybody's first "real-time" design, which makes it a strong discriminator: nearly every candidate produces something that works in a demo, and the level shows in how precisely they define "delivered" and where they put state.

QuestionL5 InstinctStaff Reality
"Is the message delivered?"Yes, the server has itDelivered = the device acked; the server having it is "sent"
"What order?"Timestamp orderPer-chat sequence; no order across chats
"Where is the message?"In the messages tableIn N device inboxes until each acks; then nowhere on the server
"Can we search?"Elasticsearch over messagesOnly on-device under E2E
"What if a gateway dies?"Messages lost?Nothing lost — gateways hold sockets, not messages

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"When the sender sees two ticks, exactly which device has exactly which bytes?"

Answering this precisely forces every decision:

  • One tick = the envelope is durably appended to every recipient device's inbox (replicated). Not "the server received it."
  • Two ticks = at least one (WhatsApp: all, in the multi-device model it's per-device and folded) recipient device acked receipt.
  • Blue ticks = the recipient opened the chat and the client sent a read receipt — which the user may disable.

🎯 Staff Move: "I'll define the ticks precisely first, because each one is a durability promise. One tick is my promise that the message survives any single server failure. I don't send it until the inbox append is replicated."


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Intent 1 — Private E2E messaging. The server is an untrusted relay. It sees metadata (who, when, size) but not content. Its durable state is per-device inboxes of undelivered envelopes. Clients own history, search, and — for groups — encryption fan-out. Scale is dominated by connection count and inbox write amplification, not by stored history.

Intent 2 — Community/workspace chat. The server is the system of record: history is the product (scrollback, search, pinned messages, integrations, bots). Channels can have 100K+ members, so fan-out on write is impossible for large channels; readers pull from a per-channel log. Storage is partitioned by (channel, time bucket). Content moderation and search are server-side.

Intent 3 — Regulated messaging. Retention and eDiscovery are hard requirements: messages must be retained immutably for years and exportable. The server must be able to read content (or hold escrowed keys). Deletion is governed by policy, not user action.

DimensionPrivate E2ECommunityRegulated
Server reads contentNoYesYes (or escrow)
Durable server stateUndelivered envelopes onlyFull historyFull history + immutable archive
Group fan-outOn write, per device, client-encryptedOn read for large channelsOn read
SearchOn deviceServer indexServer index + legal export
RetentionDelete on ack, ≤ 30 daysForever (product)Policy (e.g., 7 years)
Hardest problemMulti-device + groups under E2EHot channels, history storageRetention correctness, access audit

2.2 When NOT to Use This Design#

SituationWhy the Relay Design Is WrongUse Instead
Users expect full history on any new device instantlyRelay deletes on ack; history lives on clientsSystem-of-record chat (or encrypted server backup as a separate feature)
Channels of 10K–1M membersFan-out on write × devices is millions of writes per messagePer-channel log, fan-out on read (Discord/Slack model)
Compliance retention requiredDeleting on ack violates retentionRegulated archive
One-way broadcast (announcements, notifications)No conversation, no receipts neededNotification System
Live collaborative contentNeeds merge semantics, not message orderCollaborative Editing
In-app support chat at < 10K concurrent usersA database table + polling every 5 s is fineKeep it simple

🎯 Staff Move: "If this were an in-app support chat with a few thousand concurrent users, I'd use a Postgres table and long polling and ship it in a week. The relay-with-inboxes architecture earns its complexity at hundreds of millions of connected devices and when the server must not read content."

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy It MattersWhat to Say
E2E or notEverything about groups, sync, search"WhatsApp implies E2E; I'll design for it."
Group sizeFan-out strategy"Cap ~1K; broadcast channels are a separate product."
History on new devicesRelay vs record"Client-to-client transfer; optional encrypted backup."
Retention of undeliveredInbox storage"30 days, then drop and tell the sender nothing further."
Receipt semanticsPrivacy and fan-out"Delivered always; read receipts user-configurable."
MediaMessage path size"Encrypted blob out of band; message carries pointer + key."
Global vs regionalLatency and residency"Users homed to a region; cross-region routing for cross-region chats."

2.4 Precise Terminology#

TermMeaning HereCommon Confusion
EnvelopeEncrypted payload + routing metadata for one recipient deviceNot "the message" — one message becomes many envelopes
InboxPer-device durable queue of undelivered envelopesNot a mailbox of history
Cursor / inbox_seqMonotonic position in a device's inboxDifferent from the chat seq
Chat seqMonotonic order within one chat, assigned by its sequencerNot a global order
Sent (one tick)Durably appended to all recipient inboxesNot "left the phone"
Delivered (two ticks)Recipient device ackedNot "read"
Read (blue)Recipient opened chat; client sent read receiptOptional, privacy-controlled
Sender keySymmetric key a sender distributes to group members so group messages are encrypted onceNot per-recipient pairwise encryption
PrekeyOne-time public key uploaded by a device so others can start a session while it's offlineConsumed on use; must be replenished
Identity keyLong-term device public keyChange triggers "security code changed" notice
Companion deviceLinked device (web, desktop, tablet) with its own identityNot a mirror of the phone
Session registryMap of device → current gateway, with TTLSoft state; rebuilt on reconnect

3. The Fault Lines#

3.1 Fault Line 1: Relay vs System of Record#

StrategyWhat WorksWhat BreaksWho Pays
Relay: delete on ack, bounded retention (Staff default for E2E)Storage ∝ undelivered backlog, not history; minimal breach and legal surface; cheapNew device has no history; server search impossible; lost phone = lost history unless backed upUsers (history portability); client team (backup, transfer)
System of record: keep all historyAny device sees full history; server search; moderation on contentStorage grows forever (100B msgs/day × 1 KB ≈ 100 TB/day raw); breach surface; incompatible with E2E unless encrypted history + client keysInfra budget; Security; Legal
Relay + optional encrypted backupHistory recoverable; server still can't read itBackup key management is the user's burden; restoring is heavyUsers (remember a key/password); backup platform team

The arithmetic that settles it: at 100B messages/day, a system-of-record design stores ~36 PB/year raw before replication. A relay stores only the backlog — if 95% of messages are delivered within minutes and the tail is bounded to 30 days, the inbox store holds on the order of days of the offline fraction, typically 1–2 orders of magnitude less.

The Staff default: relay with delete-on-ack and a 30-day bound for undelivered envelopes; history on devices; optional encrypted backup as a separate product surface with its own key-management design.

When to deviate: Intent 2 or 3. If the product needs server search, history on any device, or compliance retention, you are building a different system — say so.

🎯 Staff Move: "The cheapest byte to protect is the one we don't keep. With E2E we can't read history anyway, so keeping it server-side buys us cost and risk but no features."

3.2 Fault Line 2: Ordering — Sequencer vs Timestamps vs Causal#

StrategyWhat WorksWhat BreaksWho Pays
Client timestampsNo coordinationClocks skew by minutes; malicious clients reorderUsers (confusing order)
Server timestampsEasySkew across servers (ms–s); ties; sync cursors based on time miss messagesUsers; support
Per-chat sequencer (monotonic seq) (Staff default)Total order per chat; gap detection; exact syncOne owner per chat = hot spot for very busy groups; failover must not reuse seqsMessaging core team owns sequencer and failover
Causal ordering (vector clocks / "reply-to" links)Correct causality without a central sequencerComplex; still needs a tiebreak for displayClient team
Global sequencerTotal order across everythingThroughput ceiling; single point of failureEveryone

How the sequencer works:

# chat owner shard = hash(chat_id) % N  (consistent hashing; see foundations)
on SEND(chat_id, msg_id, envelopes):
    if dedupe.seen(sender_device, msg_id): return previous ACK
    seq = chat.last_seq + 1                  # in-memory, owned by this shard
    persist: chat.last_seq = seq  AND  inbox appends for each envelope   # one durable batch
    ack sender {msg_id, seq}

Failover rule: the new owner reads last_seq from durable storage and fences the old owner (epoch number in every write), so two owners can never issue the same seq. A client seeing a gap (seq 41 then 43) waits briefly (~2 s) for 42 or requests a re-sync; gaps can also be legitimate (a message addressed only to other devices).

When to deviate: very large groups where one sequencer is too hot — split into the community/channel model with a partitioned log per channel. For 1:1 chats, a lighter option is ordering per (sender → recipient) direction, since cross-direction interleaving within milliseconds rarely matters to humans.

🎯 Staff Move: "Order is a promise per chat, never across chats. One sequencer per chat, fenced by epoch on failover. Timestamps are for the UI label, not for sorting."

3.3 Fault Line 3: Delivery Semantics and Receipts#

Diagram: 3.3 Fault Line 3: Delivery Semantics and Receipts
StrategyWhat WorksWhat BreaksWho Pays
At-most-once (fire and forget)SimpleMessages lost on any failureUsers
At-least-once + client IDs + recipient dedupe (Staff default)Robust across retries, reconnects, failovers; exactly-once displayDedupe window and client unique constraint requiredClient team (dedupe), core team (dedupe window)
"Exactly-once" via distributed transactionsSounds goodImpossible end-to-end across a lossy network; costs latencyEveryone, for a promise you can't keep

Receipts are just messages traveling the other direction. That gives them durability, ordering and multi-device delivery for free — and avoids a separate "receipt status" database that would need to be updated ~3× per message (sent, delivered, read) × devices.

Receipt cost control:

  • Batch: one read receipt covers "read up to seq N" in a chat — not one per message.
  • Groups: delivered/read receipts in groups are aggregated client-side into "read by 14 of 20"; for large groups, receipts can be sampled or computed on demand.
  • Privacy: read receipts are user-configurable; turning them off means the client simply doesn't send them.

When to deviate: community chat doesn't need per-recipient delivered receipts at all — "unread count" from a per-user read cursor is sufficient.

🎯 Staff Move: "I don't promise exactly-once delivery — nobody can over a mobile network. I promise at-least-once delivery and exactly-once display, and I can explain every step that makes that true."

3.4 Fault Line 4: Group Fan-Out — Write vs Read, and the E2E Cost#

For a group of M members with an average of d devices each:

StrategyServer Writes per MessageClient Encryption WorkWhat BreaksWho Pays
Fan-out on write, pairwise encryptionM × d inbox entriesSender encrypts M × d timesSender's phone CPU and upload for large groups (1K × 2 = 2K ciphertexts)Sender's battery and bandwidth
Fan-out on write, sender keys (Staff default for groups ≤ ~1K)M × d inbox entries (same ciphertext referenced)Sender encrypts once with its sender key; sender key distributed pairwise once per member device (and rotated on membership change)Membership change requires rekey; removed members must not decrypt future messagesClient team (rekey logic); core (membership versions)
Fan-out on read (shared group log + per-member cursor)1 writeSame sender-key encryptionServer must retain the log until all members read — relay retention becomes history; hot group logsStorage; hot partitions

The server-side optimization with sender keys: store the ciphertext once in a message blob keyed by (chat_id, seq), and write only small pointers (~50 B) into each member device's inbox. That cuts inbox bytes by ~20× for a 1 KB message.

Membership consistency matters for security: a SEND carries the sender's membership_version. If the server's version is newer (someone was added or removed), the Router rejects with stale_membership, and the client fetches the new member list, distributes/rotates sender keys, and resends. That prevents a removed member from receiving new messages and ensures a new member gets the key.

The cap is a design constraint, not a product whim. At 1,024 members × ~2 devices, one message = ~2K inbox pointer writes. A busy group sending 1 msg/s is 2K writes/s — fine. A 100K-member group would be 200K writes per message; that is a broadcast product.

🎯 Staff Move: "The group size cap is where the write amplification and the rekey cost meet. Above ~1K, I'd stop pretending it's a group and build a channel: fan-out on read, admin-only posting, no per-member receipts."

3.5 Fault Line 5: Multi-Device — Primary Relay vs Per-Device Endpoints#

Diagram: 3.5 Fault Line 5: Multi-Device — Primary Relay vs Per-Device Endpoints
StrategyWhat WorksWhat BreaksWho Pays
Primary-device relay (phone decrypts and forwards to web/desktop)Single identity; simple key managementCompanion dead when phone is offline or battery dies; phone bandwidthUsers (reliability), phone battery
Per-device endpoints (Staff default; WhatsApp's 2021 model)Each device works independently; server has one inbox per deviceSender encrypts per device (fan-out × devices); device list must be authentic; history transfer neededSenders' CPU; key-directory team; client team
Server-held history re-encrypted per deviceEasy syncServer has plaintext or keys — breaks E2ESecurity; users' privacy

Mechanics of per-device:

  • Each user has a device list signed by the primary device; senders verify it so the server can't silently add an eavesdropping "device."
  • A 1:1 message is encrypted for each of the recipient's devices and each of the sender's own other devices (so your laptop sees what you sent from your phone).
  • A new companion gets recent history by encrypted transfer from the primary device, not from the server.
  • Linking is a QR scan on the primary — a physical-presence check.

When to deviate: non-E2E products should just use server history — per-device crypto fan-out buys nothing without E2E.

🎯 Staff Move: "With E2E, a 'user' is really a set of devices, and the device list is security-critical. If the server could append a device, it could read messages. So the device list is signed by the user's primary device and verified by senders."


4. Failure Modes & Operational Reality#

4.1 Reconnect Storm After a Gateway Fleet Deploy#

t=0        Rolling deploy restarts 10% of gateways at once (misconfigured batch size).
t=+1s      ~80M sockets drop. Clients reconnect immediately (no jitter in one app version).
t=+5s      TLS handshakes 40× normal; CPU saturated on remaining gateways.
t=+10s     Session registry writes spike 50×; registry p99 2 ms → 900 ms.
t=+30s     Router can't find sessions; treats online devices as offline; OS push spikes
           to 30M/min; APNs/FCM throttle us.
t=+3min    Messages delivered with 1–2 minute delays; "WhatsApp down" trends.
  • Detection: gateway.reconnects_per_s, tls.handshakes_per_s, registry.write_latency_p99, push.throttled_total, delivery.latency_p99.
  • Blast radius: global if the deploy is global; cellular-region-specific for carrier outages.
  • Mitigation: halt the deploy; server-side reconnect hints (retry_after randomized 0–60 s); prioritize handshake capacity; TLS session resumption.
  • Prevention: drain gateways gracefully (send reconnect_in with jitter before closing); deploy batch ≤ 1–2% of connections; client jittered backoff enforced in all versions with a minimum-version policy.
  • Owner: Edge/connection platform.

4.2 Inbox Store Partition Hot Spot#

  • Scenario: a bot account or a very active 1K-member group floods inboxes; or a celebrity with millions of contacts triggers profile/status fan-out.
  • What breaks: inbox partitions keyed by device are fine individually; the sequencer for one very busy group is hot (~thousands of messages/s on one chat owner).
  • Detection: sequencer.ops_per_s_max{shard}, inbox.append_latency_p99.
  • Mitigation: per-chat and per-sender rate limits (see Rate Limiting); spam/abuse throttles on metadata (E2E means content can't be inspected server-side).
  • Owner: Messaging core (sequencer), Integrity team (abuse limits).

4.3 Lost Message After Inbox Failover (The Silent One)#

t=0        Inbox replica set primary fails. Async replication lag was 400 ms.
t=+2s      Replica promoted. Envelopes appended in the last 400 ms — already ACKed to
           senders (one tick) — are gone.
t=+∞       Senders see one tick forever. Recipients never see the messages. No error.
  • Detection: inbox.acked_but_missing — hard to detect directly; proxy: senders' clients report "one tick older than 24 h while recipient was online" via telemetry (client.stuck_single_tick_total).
  • Blast radius: messages in the replication window for affected devices.
  • Mitigation: ack the sender only after quorum (synchronous) replication of the inbox append; accept ~2–5 ms extra latency.
  • Prevention: make "one tick = quorum-durable" a written invariant; chaos-test primary failure under load monthly.
  • Owner: Messaging core / storage team.

4.4 Prekey Exhaustion#

  • Scenario: a device offline for weeks receives many new-session requests (new contacts, group adds). Its uploaded one-time prekeys are consumed.
  • What breaks: senders fall back to the signed prekey (weaker forward-secrecy properties for the first message) or fail session setup.
  • Detection: keys.prekeys_remaining{device} low; keys.fallback_to_signed_prekey_total.
  • Mitigation: clients top up to ~100 prekeys when below a threshold on each connect; server alerts on devices near zero.
  • Owner: Key directory team.

4.5 OS Push Degradation#

  • Scenario: APNs or FCM latency rises to minutes or throttles us.
  • Blast radius: backgrounded/offline devices — messages wait in inboxes (not lost), delivered on next app open.
  • Detection: push.send_latency_p99{provider}, push.throttled_total, delivery.latency_p99{device_state=background}.
  • Mitigation: collapse pushes per device (one wake-up for many messages), priority for 1:1 over group, back off per provider guidance.
  • Owner: Notifications platform (see Notification System).

4.6 New Year's Eve Spike#

  • Scenario: 3–5× message volume within ~5 minutes of midnight in each timezone, dominated by group and media messages.
  • Mitigation: capacity for 5× per-region peak on sequencers, inboxes and media upload; media uploads prioritized below text; follow-the-midnight capacity shifting across regions.
  • Owner: Capacity planning with messaging core.

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Reconnect stormgateway.reconnects_per_sRegion or globalJittered reconnect hints; halt deployEdge platform
Gateway crashgateway.connections drop~1M devices, secondsClients reconnect elsewhere; no data lostEdge platform
Inbox primary failoverclient.stuck_single_tick_totalReplication windowQuorum append before ACKMessaging core
Hot sequencersequencer.ops_per_s_maxOne busy chatRate limits; move chat to less loaded shardMessaging core
Session registry slowregistry.write_latency_p99Online delivery → push fallbackShed registry writes; cacheEdge platform
Prekey exhaustionkeys.prekeys_remainingNew sessions to that deviceClient top-upKey directory
Push provider degradedpush.send_latency_p99Background delivery delayedCollapse, prioritizeNotifications
Stale group membershiprouter.stale_membership_rejectsOne group's sends retryClient refresh + rekeyMessaging core + client
Media store slowmedia.upload_latency_p99Media messagesText unaffected; retry uploadsMedia platform

🎯 Staff Move: "Look at what's absent: no failure in this table loses a message except the inbox failover — and that's exactly why the sender's single tick waits for a quorum write. Everything else delays delivery; nothing else loses it."


5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
ScopingFeatures: 1:1, groups, media, receiptsRelay vs record; commits to E2E relay; names what E2E moves to clientsFrames E2E as a company decision with Security, Legal, T&S and Product sign-off
DeliveryQueue so nothing is lostPer-device inbox, quorum append before one tick, cursor acks, delete on ackDelivery SLOs and retention as external commitments
Semantics"Exactly-once"At-least-once + client IDs = exactly-once displayStandard for idempotent client protocols across products
OrderingTimestampsPer-chat sequencer with epoch fencingDocuments where order is not promised; stops teams from depending on it
GroupsFan-outFan-out on write with sender keys; membership versions; cap rationaleGroup cap and channel product as a portfolio decision with abuse modeling
Multi-deviceSync everywherePer-device keys and inboxes; signed device lists; client history transferDevice-linking security policy; key-change UX; prices fan-out multiplier
OperationsMonitoringReconnect storms, stuck-tick telemetry, push degradationGame days for gateway fleet loss; error budgets per tier; deploy policy

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Relay-vs-record framing"Is the server the record or a relay? For WhatsApp, a relay."
Precise ticks"One tick means quorum-durable in every recipient inbox."
Honest semantics"At-least-once delivery, exactly-once display."
Sequencer ownership"One sequencer per chat, fenced on failover."
E2E consequences"E2E means the sender encrypts per device, search is on-device, and a new laptop gets history from the phone."
Cap as constraint"Above ~1K it's a channel, not a group."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Exactly-once delivery" with no mechanismPromising the impossible
Sorting by timestampOrdering bugs in groups
Server-side search in an E2E designContradiction
Gateways that store messagesGateway crash loses data
No offline pathStore-and-forward is the core of the product
Keep all messages forever with no reasonCost and risk without a feature

5.4 Common False Positives#

  • Cryptography depth ≠ system design. Explaining the double ratchet in detail is impressive and rarely what's being evaluated; the system consequence (per-device fan-out) is.
  • WebSocket scaling tricks ≠ delivery design. Knowing epoll limits doesn't answer "where is the message when the phone is off?"
  • Kafka everywhere ≠ inbox design. A topic per user doesn't scale to billions of devices; a partitioned inbox store does.
  • "Eventually consistent" ≠ acceptable for one tick. The sender's tick is a durability promise.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–4 minRelay vs record, E2E, numbers
Entities and protocol4–7 minEnvelope, inbox, cursor, client msg ID
Architecture7–11 minGateways, router, sequencer, inboxes, push
Inbox and semantics11–18 minQuorum append, dedupe, cursor ack, delete
Ordering and receipts18–24 minSequencer, gaps, receipts as messages
Groups under E2E24–31 minSender keys, membership versions, cap
Multi-device31–37 minPer-device keys, signed device list, history transfer
Edge operations37–42 minReconnect storms, push
Wrap-up42–45 minMetrics, next steps, skipped scope

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response
"Add message search"E2E consistency"On-device index; server search would require server-readable content."
"Add a 100K-member community"Fan-out limits"That's a channel: fan-out on read, admin posting, no per-member receipts."
"Add disappearing messages"Client-enforced policy"A timer attribute in the encrypted payload; clients delete; server deletes on ack anyway."
"Law enforcement asks for message content"Relay implications"We don't have content; we have limited metadata. Legal owns the response process."
"Add voice/video calls"Different transport"Signaling over the message path; media over WebRTC/TURN; separate system."
"Spam is up 5×"Moderation without content"Metadata signals: send rate, new-account fan-out, user reports (which include the reported messages, sent by the reporter)."

6.3 What to Deliberately Skip#

TopicWhy SkipOne-Liner If Asked
Signal Protocol mathNot the systems question"X3DH for session setup from prekeys, double ratchet per message."
Load balancer choiceStandard"L4 LB to gateways; consistent device-to-gateway not required."
Media transcodingSeparate pipeline"Client compresses; encrypted blob upload; see Blob Storage."
Presence detailsLossy, best-effort"Subscribed only for open chats; heartbeat-derived; no durability."
Stickers, reactionsSame envelope path"Just message types."

6.4 Follow-Up Questions to Expect#

  1. "The phone is offline for 40 days. What happens to messages sent on day 1?"
  2. "The sender's app crashes after sending but before the ACK. What does the user see on restart?"
  3. "Two members of a group send at the same moment. What order does everyone see?"
  4. "Someone is removed from a group. How do you guarantee they can't read later messages?"
  5. "A user links a new laptop. What does it show?"
  6. "A gateway with 1M connections dies. What happens?"
  7. "How many writes per second hit the inbox store at peak?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design WhatsApp."

Staff Answer

"The first fork is relay vs system of record. WhatsApp is a relay with end-to-end encryption: the server holds encrypted envelopes per device until acked, then deletes them. I'll design that.

Numbers: ~100B messages/day, ~1.2M/s average, ~3–5M/s peak, and each message fans out to ~5–10 device inboxes on average — so ~10–30M inbox appends/s at peak, but tiny ones if the ciphertext is stored once and inboxes hold pointers. Connections: ~500M–1B concurrent devices; at ~1M per gateway, ~1K gateways.

Core: gateways hold sockets only; a router asks the chat's sequencer for a seq, appends to each recipient device inbox with quorum durability, ACKs the sender — that's one tick — pushes to online devices or wakes offline ones. Devices ack a cursor; the server deletes. E2E pushes encryption fan-out, history and search to the client, and I'll show where."

Why this is L6:

  • Commits to an intent with a reason
  • Sizes connections and write amplification separately
  • Defines one tick precisely

What L7 adds:

  • Names E2E as a cross-functional commitment (Legal, T&S, Product) before engineering
  • Prices the gateway fleet and inbox store and names the dominant cost (connections, not storage)

Drill 2: Core Mechanic — The Lost ACK#

Prompt: "Alice sends a message. The server stores it, but the ACK never reaches Alice. Walk me through it."

Staff Answer

"Alice's client has the message in its local outbox with msg_id = m1 and a clock icon. No ACK within ~10 s or on reconnect, it resends m1. The Router checks its dedupe window keyed by (alice_device, m1) — say a 24-hour TTL in a fast KV — finds it, and returns the original ACK {m1, seq 41} without appending again. One tick appears.

If the dedupe entry was lost (e.g., dedupe store failover), the Router would append a second envelope with a new seq. Bob's client has a unique constraint on msg_id, so the second copy is dropped on arrival and not rendered. The cost is one wasted inbox entry and a seq gap Bob's client can ignore. Display remains exactly-once."

Why this is L6:

  • Two layers of dedupe with clear roles
  • Handles the failure of the dedupe layer itself
  • Explains seq gaps as benign

What L7 adds:

  • Makes the client-ID/idempotency protocol a company-wide standard for all mobile write APIs
  • Measures duplicate-drop rate on clients as a health metric for the server dedupe layer

Drill 3: Make It Concrete — Inbox Sizing#

Prompt: "How big is the inbox store?"

Staff Answer

"Inbox holds only undelivered envelopes. Assume 1.2M messages/s average, ~6 device-deliveries per message → ~7M inbox entries/s. If entries are pointers (~64 B) and the ciphertext blob (~1 KB) is stored once per message, the write bandwidth is ~450 MB/s of pointers + ~1.2 GB/s of blobs.

Residency: ~90% of deliveries complete within a minute (online devices), ~9% within hours (backgrounded), ~1% sit for days. Steady-state entries ≈ 7M/s × (0.9 × 30 s + 0.09 × 3 h + 0.01 × 3 days) ≈ 7M × (27 + 972 + 2,592) s ≈ 25B entries — the long tail dominates. At ~64 B that's ~1.6 TB of pointers plus blobs for undelivered messages (~several TB), ×3 replication. The 30-day cap bounds the worst case. Compare ~100 TB/day raw if we kept history."

Why this is L6:

  • Uses residency time, not throughput, to size storage
  • Notices the long tail dominates
  • Contrasts with system-of-record cost

What L7 adds:

  • Evaluates whether shortening the retention cap (e.g., 30 → 14 days) is worth the product cost, with data on how many messages are delivered after day 14

Drill 4: Dependency Down — Push Provider Outage#

Prompt: "FCM is degraded for an hour. What happens?"

Staff Answer

"Online Android devices are unaffected — they're on our socket. Backgrounded Android devices don't get woken, so their messages wait in inboxes and arrive on next app open. No loss. Mitigations: collapse wake-ups to one per device per window, prioritize 1:1 over group wake-ups when throttled, and show an in-app banner if we detect delivery-latency degradation. Alert on push.send_latency_p99{provider=fcm} and route to the notifications platform. The senders keep seeing one tick until delivery — which is correct."

Why this is L6:

  • Distinguishes online vs background devices
  • No data loss because the inbox is the source of truth
  • Prioritizes within a throttled budget

What L7 adds:

  • Tracks dependency on two platform push providers as a strategic risk; evaluates persistent background connections where OS policy allows

Drill 5: Hot Key — The Viral Group#

Prompt: "A 1,024-member group is sending 50 messages/s during a live sports match."

Staff Answer

"50 msg/s × ~2K device inboxes = ~100K pointer appends/s from one chat, plus 50 seq increments/s on one sequencer — the sequencer is fine (a single in-memory counter handles 100K+/s), the inbox fan-out is spread across devices so it's spread across partitions. The real costs are client-side: every member's phone receives 50 msg/s — battery and data — and delivered/read receipts would be 50 × 1K × 2 per second if sent individually.

So: receipts in large groups are batched as 'read up to seq N' per chat per device, sent at most every few seconds; per-group send rate limits (e.g., 20 msg/s) with slow-mode as an admin option; client coalesces rendering. The inbox pointers are cheap because the ciphertext is stored once."

Why this is L6:

  • Finds the actual hot spots (receipts, client battery), not the imagined one
  • Batching receipts by cursor
  • Product-level rate controls

What L7 adds:

  • Decides whether live-event groups should be a distinct product (channels) with different economics

Drill 6: Multi-Tenant — Business Messaging#

Prompt: "Businesses want to message millions of customers through our platform."

Staff Answer

"Businesses are tenants with very different traffic shapes: bursty campaign sends to millions of 1:1 chats. Isolate them: a separate business ingestion API with per-tenant quotas and templates (pre-approved message types to limit spam), queued sends drained at a rate the inbox tier can absorb without affecting consumer latency, and per-tenant quality scores based on block/report rates that throttle bad senders. Consumer traffic always has priority on shared inbox and push capacity."

Why this is L6:

  • Tenant isolation and priority
  • Abuse controls tied to user signals
  • Protects the consumer SLO

What L7 adds:

  • Frames this as a revenue product with its own P&L, and sets the policy for how business messages interact with E2E (e.g., businesses using hosted endpoints)

Drill 7: Build vs Buy — The Connection Layer#

Prompt: "Should we build our own gateway or use a managed WebSocket service?"

Staff Answer

"At small scale (< ~1M concurrent), a managed WebSocket/pub-sub service is fine and cheaper in people. At hundreds of millions of connections, connection cost dominates: managed services often price per connection-minute and per message, which at 1B devices becomes the largest line item; and we need control over heartbeat intervals, binary framing, and reconnect behavior for mobile battery. WhatsApp's own history — Erlang handling ~2M connections per host — shows what a purpose-built edge can do. I'd buy early, keep the protocol ours, and build when connection cost crosses a threshold we can compute."

Why this is L6:

  • Scale-dependent decision with a trigger
  • Keeps protocol ownership to preserve the option

What L7 adds:

  • Computes the crossover: managed cost/month vs a gateway team (~6–8 engineers) + hosts

Drill 8: Policy Change Without Outage — Rolling Out Multi-Device#

Prompt: "Move from phone-as-relay to per-device endpoints for 2B users."

Staff Answer

"This changes the protocol for every sender, so the rollout is capability-gated. Step 1: clients ship support for per-device encryption and signed device lists but don't use it. Step 2: devices advertise capability; senders encrypt per-device only when all recipient devices support it, else fall back to the old path. Step 3: enable per-device linking for a beta cohort; watch decryption-failure rate, delivery latency and battery metrics. Step 4: expand by percentage; old clients below a minimum version are asked to update. Rollback: disable new linking; existing linked devices keep working through a compatibility path. Key metric: client.decrypt_failures_total must stay at baseline."

Why this is L6:

  • Capability negotiation instead of a flag day
  • Safety metric specific to E2E (decrypt failures)
  • Rollback path defined

What L7 adds:

  • Coordinates security review, public documentation (whitepaper), and support readiness; sets the minimum-client-version policy

Drill 9: Cost#

Prompt: "What dominates the cost of running this, and how would you cut it 30%?"

Staff Answer

"Connections, not storage. 1B devices × heartbeats every ~45 s ≈ 22M tiny frames/s, plus TLS and memory per socket. Levers: longer heartbeat intervals where carrier NATs allow (adaptive per network), more connections per host (runtime efficiency), and fewer reconnects (graceful drains, session resumption). Second: media storage and egress — dedupe forwarded media by content hash of the ciphertext pointer, and expire blobs after all recipients download. Third: push — collapse wake-ups. Storage of the inbox is small by comparison."

Why this is L6:

  • Identifies the non-obvious dominant cost
  • Concrete levers per cost center

What L7 adds:

  • Builds a cost-per-DAU model per component; negotiates CDN/egress contracts for media

Drill 10: Multi-Region#

Prompt: "Users in Brazil message users in India. Design the regions."

Staff Answer

"Each user (and their devices' inboxes) is homed in a region near them; each chat's sequencer is homed where its creator or majority lives. A message from Brazil to India: Brazilian gateway → router → the chat's sequencer (maybe cross-region, ~150–250 ms RTT) → append to the Indian recipient's inbox in India → ACK. To keep the sender's one tick fast, the router can append durably to a local cross-region outbound queue with quorum and ACK, then forward — the tick still means 'durable', just in a different store. Region failure: devices reconnect to another region; inboxes are replicated cross-region asynchronously for disaster recovery, with the understanding that a region loss can delay (not lose, if replicated synchronously in-region) messages."

Why this is L6:

  • Homes users and chats; explains the cross-region path
  • Keeps the one-tick semantics honest
  • Distinguishes in-region sync vs cross-region async replication

What L7 adds:

  • Data residency for metadata by country; decides region footprint by user distribution and regulatory exposure

8. Deep Dive Scenarios#

Deep Dive 1: Peak Traffic — New Year's Eve Delivery Delays#

Context: 00:01 in a large timezone. Online delivery p99 goes from 300 ms to 45 s. Social media says "WhatsApp is down." The on-call escalates to you.

Questions to Surface First:

  • Is the delay at the router/sequencer, inbox append, or push to the recipient gateway?
  • Are messages with media delayed more than text?
  • Are sender ACKs (one tick) delayed, or only delivery (two ticks)?
  • Is the session registry healthy?

Typical L5 Approach: Adds router instances and inbox nodes across the board.

Staff Approach: Splits the latency by stage: send_to_ack_ms (sender side) vs ack_to_deliver_ms. Finds ACKs fine but deliver slow: the session registry is overloaded by reconnects from users toggling airplane mode at midnight, so routers misclassify online devices as offline and fall back to push, which is throttled. Mitigations: registry read cache with 30 s TTL on routers; prioritize text over media on gateway egress; raise push collapse window.

Principal Approach: Treats New Year's as a known annual event with a "follow the midnight" capacity plan: registry and gateway capacity shifted region by region, media upload deprioritization pre-configured, and a load test at 5× last year's peak two weeks ahead. Owns the public status communication plan.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Stage-split latency; confirm no loss (ACK path healthy)
TriageRegistry latency, push fallback rate, gateway egress queues
Quick fixRouter-side registry cache; text priority over media
GuardrailsWatch push.throttled_total, registry p99
Post-mortemRegistry capacity for reconnect bursts; annual event plan

Metrics to Watch: delivery.send_to_ack_ms_p99, delivery.ack_to_deliver_ms_p99, registry.read_latency_p99, router.offline_fallback_ratio, push.throttled_total

Organizational Follow-up: Annual peak events calendar with per-region capacity plans.

Ownership Question: "Who decides to deprioritize media during a peak?" Staff answer: Messaging core on-call, pre-approved in the runbook; product agreed that text-first is the right degradation.

Key Takeaway: "Split latency by stage. A slow two-tick with a fast one-tick is a delivery-path problem, not a durability problem."

What clears the Staff bar:

  • Stage-level latency decomposition
  • Finds the registry → push fallback amplifier
  • Pre-agreed degradation order

Deep Dive 2: The Silent Failure — Messages Stuck on One Tick#

Context: Client telemetry shows client.stuck_single_tick_total rising slowly for 2 weeks in one region: ~0.01% of messages never get a delivered receipt even though recipients are active.

Questions to Surface First:

  • Did the recipient receive the message (check recipient telemetry for msg_id), or did only the receipt get lost?
  • Is it correlated with multi-device users, a client version, or a storage cluster?
  • Any failovers in that region's inbox store?

Typical L5 Approach: Assumes receipts are lossy and suggests resending receipts.

Staff Approach: Correlates by recipient device count: nearly all cases are recipients with a companion device linked in the last 30 days. Root cause: the device list cache in the router had a TTL of 1 h; new companions didn't receive envelopes until cache refresh, and the delivered receipt logic waited for all devices. Fix: device list changes invalidate router caches via an event; senders include the device-list version and the router rejects stale versions (like membership versions).

Principal Approach: Establishes "every cache of security-relevant or routing-relevant lists is versioned and invalidated by event, never TTL-only" as an engineering standard. Adds end-to-end synthetic probes (test accounts with multiple devices in each region sending every minute) so silent delivery failures are detected by the platform, not by telemetry trends.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateConfirm whether messages or receipts are missing
TriageCorrelate by device count, version, cluster
Quick fixInvalidate router device-list caches on change
GuardrailsDevice-list versions in SEND; synthetic probes
Post-mortemTTL-only caches on routing data prohibited

Metrics to Watch: client.stuck_single_tick_total, router.stale_device_list_rejects, probe.e2e_delivery_success{region}

Organizational Follow-up: Synthetic probes owned by messaging core, alerting at 99.99% success.

Ownership Question: "Who owns end-to-end delivery correctness when it spans client, router and key directory?" Staff answer: Messaging core owns the end-to-end SLO and the probes; component teams own their pieces. Someone must own the whole path.

Key Takeaway: "A relay's silent failure is a message that never arrives and never errors. Only end-to-end probes catch it."

What clears the Staff bar:

  • Separates lost message from lost receipt
  • Finds correlation with a routing-cache TTL
  • Adds end-to-end synthetic probes

Deep Dive 3: Large-Customer Onboarding — A Government Uses Groups for Emergency Alerts#

Context: A government agency wants to send emergency updates to 20M citizens via the platform.

Questions to Surface First:

  • Is this conversational (replies expected) or broadcast?
  • Latency requirement — minutes or seconds?
  • Should E2E apply? Who are the "devices" on the sender side?

Typical L5 Approach: Creates many 1,024-member groups and fans out.

Staff Approach: Recognizes broadcast, not chat: use a channel/business-messaging path — the sender posts once to a channel log, subscribers fetch on read or receive batched wake-ups; rate-controlled delivery at a level the inbox and push tiers can absorb (20M × ~1.5 devices = 30M deliveries; at 500K/s, ~1 minute). No per-recipient receipts; aggregate reach metrics instead.

Principal Approach: Weighs the public-interest and political risk of the platform as an emergency channel: SLOs we can publicly commit to, abuse/impersonation of government accounts (verified sender program), and whether to partner or decline. Brings Policy and Legal in before engineering starts.

Staff Approach — Full Reasoning
PhaseWhat to Do
ScopingBroadcast vs conversation
DesignChannel with fan-out on read and batched wake-ups
Capacity30M deliveries within target window; reserve push budget
PilotOne region, test broadcasts
LaunchVerified sender, rate-controlled sends

Metrics to Watch: channel.broadcast_reach_ratio, push.batch_latency_p99, consumer.delivery_latency_p99 (must not regress)

Organizational Follow-up: Verified-sender program owned by Integrity.

Ownership Question: "Who is accountable if an emergency alert is delayed?" Staff answer: We commit only to what we measure in pilots; the agency's contract states our delivery percentile, and Legal signs it.

Key Takeaway: "Don't bend a group feature into a broadcast system. Recognize the intent and use the right fan-out model."

What clears the Staff bar:

  • Identifies intent mismatch
  • Computes delivery time from throughput
  • Protects consumer SLOs

Deep Dive 4: Post-Mortem — Removed Group Member Read Messages#

Context: A security researcher reports that a user removed from a group continued to receive decryptable messages for several minutes.

Questions to Surface First:

  • Did the server keep routing envelopes to the removed member's devices, or did the removed member's client keep the old sender key and receive ciphertexts via another path?
  • Did senders rotate their sender keys on membership change?
  • Were there senders on old client versions?

Typical L5 Approach: Fixes the server to stop routing to removed members immediately.

Staff Approach: Finds two gaps: (1) routers had cached membership for up to 5 minutes; (2) some old clients didn't rotate sender keys on removal. Fix both: membership versions checked on every SEND (stale → reject); sender-key rotation mandatory on removal, enforced by rejecting sends with pre-removal key IDs after a grace period of seconds; old clients below a minimum version blocked from sending in groups.

Principal Approach: Treats it as a security incident with disclosure obligations: coordinates with the researcher (bug bounty), publishes an advisory as appropriate, and makes group-membership semantics part of a formal security model that every client change is reviewed against.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateReproduce; measure the exposure window
TriageServer routing vs client key rotation
Quick fixMembership version checks; forced rotation
GuardrailsMinimum client version for group sends
Post-mortemSecurity model review for membership changes

Metrics to Watch: router.stale_membership_rejects, client.sender_key_rotations_total, group.sends_by_old_clients

Organizational Follow-up: Security review gate for any change touching membership or keys.

Ownership Question: "Who owns the security semantics of group membership?" Staff answer: The messaging security team — with veto on client and server changes that touch keys or membership.

Key Takeaway: "Under E2E, access control is enforced by keys, not by routing. Both must change together."

What clears the Staff bar:

  • Distinguishes routing from cryptographic access
  • Uses versioning to eliminate cache windows
  • Treats as security incident

Deep Dive 5: Multi-Region Expansion — Data Residency for Metadata#

Context: A country requires that messaging metadata for its citizens stays in-country.

Questions to Surface First:

  • What counts as metadata — sender/recipient IDs, timestamps, IPs, group membership?
  • Does content (already E2E) fall under the rule?
  • How are cross-border chats treated?

Typical L5 Approach: Deploys a region in-country and routes those users there.

Staff Approach: Homes those users' devices, inboxes, session registry entries and group memberships in-country. Cross-border chats: the sequencer for a chat is homed based on policy (e.g., the in-country user's region if either participant is in-country); envelopes to foreign recipients leave the country only as delivery requires, with minimal retained metadata. Logs and analytics are aggregated in-country.

Principal Approach: Decides whether to comply, negotiate, or exit the market with Legal and Policy, based on the company's privacy commitments; builds residency as a reusable capability (region-in-a-box with policy-driven data placement) rather than a one-off.

Staff Approach — Full Reasoning
PhaseWhat to Do
ClassifyEnumerate metadata fields and their stores
DesignPolicy-driven homing for users, chats, logs
MigrateMove affected users' inboxes during low traffic with dual-read
VerifyAudit data flows; residency tests in CI
OperateRegion-local on-call; incident process with regulators

Metrics to Watch: residency.cross_border_metadata_bytes, migration.inbox_moves_remaining, delivery.latency_p99{region}

Organizational Follow-up: Data-classification catalog maintained by Privacy engineering.

Ownership Question: "Who decides whether we operate in a market with these rules?" Staff answer: Leadership with Legal and Policy; engineering provides the cost and the technical feasibility, not the decision.

Key Takeaway: "E2E protects content, not metadata. Residency is a metadata-placement problem."

What clears the Staff bar:

  • Separates content from metadata
  • Policy-driven placement
  • Escalates the market decision

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Frame chat as relay vs system of record and commit to one
  • Define one tick, two ticks and blue ticks as precise durability and acknowledgment states
  • Explain at-least-once delivery with exactly-once display using client IDs and dedupe
  • Design per-chat ordering with a fenced sequencer
  • Size inbox storage from residency time and connection cost from heartbeats
  • Explain group fan-out with sender keys, membership versions and the size cap
  • Explain per-device multi-device under E2E with signed device lists
  • At L7: take E2E through a cross-functional decision, price the edge fleet, and write the standard

The Bar for This Question#

Mid-level (L4): WebSocket servers, a messages table, online delivery, some offline storage. Timestamps for ordering.

Senior (L5): Adds a queue per user, Cassandra for messages, presence, push notifications, horizontal gateway scaling. Claims exactly-once, sorts by timestamp, and treats E2E and multi-device as footnotes.

Staff+ (L6): Commits to relay with E2E, builds per-device inboxes with quorum durability and cursor acks, uses a per-chat sequencer, treats receipts as messages, explains sender keys and group caps, and designs per-device multi-device with signed device lists. Every failure has a metric and owner. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "Exactly-Once Delivery Is Marketing"#

EvidenceImplication
Mobile networks drop ACKsRetries are unavoidable
Dedupe at the receiver is cheap and totalExactly-once display is achievable

The Staff position: At-least-once + idempotent display.

Why this matters in interviews: Precision about semantics is the Staff-level signal.

10.2 "The Server Should Know as Little as Possible"#

EvidenceImplication
Undelivered-only storage is orders of magnitude smallerLower cost
Data not held can't be breached or compelledLower risk
WhatsApp's small early team ran a massive relaySimplicity scales

The Staff position: Relay with delete-on-ack for consumer E2E chat.

Why this matters in interviews: It reframes "scale" as "do less."

10.3 "Groups Above ~1K Are a Different Product"#

EvidenceImplication
Fan-out on write × devices grows linearly per messageCost explodes
Sender-key rekey on every membership changeChurn cost
Conversation norms break downProduct mismatch

The Staff position: Cap groups; build channels for broadcast.

Why this matters in interviews: Knowing where a design stops working is Staff judgment.

10.4 "Connections Cost More Than Messages"#

EvidenceImplication
~1B sockets with heartbeats every 30–60 sTens of millions of frames/s with no messages
Reconnect storms are the largest outage class for chat edgesOperational cost concentrates at the edge

The Staff position: Invest in the edge: graceful drains, jittered reconnects, adaptive heartbeats.

Why this matters in interviews: Most candidates optimize message storage, the smaller cost.

10.5 "E2E Is a Product and Policy Decision First"#

EvidenceImplication
E2E removes server search, content moderation and easy historyProduct surface changes
Governments and regulators engage publicly on E2EPolicy exposure

The Staff position: Engineering can implement E2E; the org must choose it.

Why this matters in interviews: Naming the non-engineering owners is a Staff-to-Principal signal.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

A Staff engineer designs a messaging system. A Principal engineer sees that messaging is the substrate for half the company's products — consumer chat, business messaging, notifications, support, payments-in-chat, calling signaling — and that the one decision that shapes all of them is who can read the content. Once consumer chat is end-to-end encrypted, every adjacent product must decide whether it lives inside that boundary (and gives up server-side features) or outside it (and must be clearly distinguished to users). The L7 job is to make that boundary explicit, fund the shared edge and inbox platform, and prevent each product from building its own socket fleet and its own half-compatible delivery semantics.

The Org-Level Fault Line#

One messaging platform with a clear E2E boundary vs per-product messaging stacks.

OptionWhat WorksWhat BreaksWho Pays
Per-product stacks (chat, support, notifications, business each with its own sockets and queues)Local speed3–4 connection fleets on the same phone (battery); inconsistent delivery semantics; duplicated edge on-callUsers' batteries; infra; on-call
One platform, everything E2EStrongest privacy storyBusiness messaging, support and moderation lose server-side capabilities they needRevenue products; T&S
Shared edge + inbox platform; explicit E2E boundary per product (L7 default)One connection per device; one delivery contract; products declare whether they're inside E2ERequires clear UX labeling and governance of the boundaryPlatform team (~15–30 eng); Product (labeling)

🧭 Principal Move: "Every product on the phone shares one connection and one delivery contract. What differs is whether the server can read the payload — and that is declared per product, reviewed by Security and Privacy, and visible to users. Nobody gets to be ambiguous about it."

Cost Model#

Assumptions: cloud list prices order-of-magnitude; ~0.5–1M connections per gateway host; heartbeats every ~45 s; ~6 device deliveries per message; relay storage (undelivered only); media stored encrypted with TTL; loaded engineer ~$250K/yr.

ScaleInfra $/monthDominant CostHeadcount (eng)On-Call Load
10M DAU, ~5M concurrent, ~500M msgs/day~$50–150KGateways, media storage/egress10–201–2 rotations
200M DAU, ~100M concurrent, ~10B msgs/day~$1–3MGateway fleet (~100–200 hosts), media egress, push80–150 across edge, core, media, clients, securityPer-tier rotations
2B users, ~1B concurrent, ~100B msgs/day~$10–30MMedia egress and storage, edge fleet (~1–2K hosts), cross-region trafficHundreds, heavily on clients and securityFollow-the-sun; per-region edge

The line that moves the business case: media. Text inbox storage is a rounding error beside media storage and egress; forwarding dedupe and blob TTLs are worth more than any inbox optimization.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to ReverseWhy
Launching E2E for consumer chatOne-way (trust and public commitment)Reversal is a public-trust crisisUsers and regulators treat it as a promise
Relay (delete on ack) vs recordOne-way-ishCan't retroactively recover deleted history; adding history requires encrypted backup designProduct expectations form around it
Client message ID and cursor protocolOne-way-ishEvery client version in the wild speaks itMobile clients linger for years
Group size capTwo-way upward, one-way downwardLowering breaks existing groupsRaise cautiously
Heartbeat intervalTwo-wayConfigTune per network
Inbox store technologyTwo-way with effortMigration of hot, small dataAbstraction layer
Retention window (30 days)Two-way downward with noticeShortening loses some late deliveriesMeasure first

The Standard I'd Write#

RFC: Device Messaging Delivery & Confidentiality Standard (v1)

Scope: Every product that sends data to user devices over a persistent connection or OS push.

Mandatory requirements:

  • Products MUST use the shared connection gateway; separate persistent connections from the same app are prohibited.
  • Every client-originated write MUST carry a client-generated idempotency ID; recipients MUST dedupe on it.
  • Delivery states exposed to users MUST map to defined durability events (e.g., "sent" = quorum-durable in all recipient inboxes).
  • Each product MUST declare its confidentiality class: E2E (server cannot read) or server-readable; the class MUST be visible in the UI where users could confuse them.
  • Routing-relevant lists (group membership, device lists) MUST be versioned; caches MUST be invalidated by event, not TTL alone.
  • Undelivered payloads MUST NOT be retained beyond 30 days; delivered payloads MUST be deleted on ack for E2E products.
  • Gateway deploys SHOULD drain ≤ 2% of connections per batch with jittered reconnect hints.

Exceptions: Filed with Messaging Platform, Security and Privacy; reviewed quarterly.

Success metrics: One persistent connection per device across products; client.stuck_single_tick_total ≤ 1 in 10⁵ messages; synthetic end-to-end probe success ≥ 99.99% per region; zero unlabeled server-readable surfaces inside E2E apps.

What I'd Tell the VP#

"Our messaging product's biggest promise is simple: when you see one tick, the message won't be lost, and nobody — including us — can read it. Keeping that promise is cheaper than it sounds because we only store messages until they're delivered. The risks are elsewhere: other teams building their own connections that drain batteries, and business features that blur whether we can read a message. I'm proposing one shared messaging platform with a clear, visible line between private and business messaging. It's mostly reorganizing existing teams, and it protects both our battery reputation and our privacy commitment."

Principal Interview Signals#

SignalWhat It Sounds Like
Confidentiality boundary"Each product declares whether it's inside E2E, and users can see it."
Portfolio edge"One connection per device across all products."
Prices it"Media egress is the cost line; inbox storage is a rounding error."
One-way doors"Launching E2E is a public commitment we can't walk back."
Writes the standard"Versioned routing lists, event invalidation, deletion on ack — mandatory."

Staff answers that L7 interviewers find insufficient:

  • "We'll add E2E with the Signal Protocol." — Correct, but ignores the cross-functional decision and the products it constrains.
  • "Per-device inboxes solve multi-device." — Correct, but doesn't own the device-linking security policy or its UX.
  • "We'll cap groups at 1K." — Correct, but doesn't decide what product serves the users who need more, or who owns it.

Appendices

Appendix A: Mechanics in Depth#

A.1 Send Path#

on SEND(dev, msg_id, chat_id, envelopes, membership_version, device_list_versions):
    if dedupe.get(dev, msg_id) as prev: return prev
    chat = sequencer.owner(chat_id)                         # fenced by epoch
    if membership_version < chat.membership_version: return STALE_MEMBERSHIP
    for each recipient user: if device_list_versions[u] < directory.version(u): return STALE_DEVICES
    seq  = chat.next_seq()
    blob = store_once(chat_id, seq, ciphertext_or_per_device)   # group: one blob with sender key
    quorum_write([inbox_append(d, pointer(blob, seq, msg_id)) for d in recipient_devices(chat)])
    dedupe.put(dev, msg_id, {seq, ts}, ttl=24h)
    for d in recipient_devices: if registry.online(d): gateway(d).push(d) else push_sender.wake(d)
    return ACK{msg_id, seq, server_ts}

A.2 Receive Path#

on DELIVER(inbox_seq, envelope):
    if local_db.exists(envelope.msg_id): skip          # exactly-once display
    plaintext = decrypt(envelope)                      # session or sender key
    local_db.insert(chat_id, seq, msg_id, plaintext)
    periodically: send RECV_ACK{up_to_inbox_seq}       # batch, e.g., every 100 msgs or 1 s
    send RECEIPT delivered (batched per chat: up to seq)

A.3 Gap Handling#

A client seeing chat seq 41 then 43 waits ~2 s; if 42 doesn't arrive, it asks the server whether seq 42 was addressed to it. If not (e.g., a message to other devices only, or a sender-side retry gap), it records the gap as benign.

Appendix B: Keys and Data Model#

EntityKeyStoreRetention
Inbox entry(device_id, inbox_seq)Partitioned KV / wide-column (Cassandra-style), quorum writesUntil ack; ≤ 30 days
Message blob(chat_id, seq)Same store or blob storeUntil all recipient devices ack; ≤ 30 days
Chatchat_id → last_seq, epoch, membership_versionStrongly consistent KVWhile chat exists
Dedupe(device_id, msg_id)Fast KV (Redis-style)24 h
Session registrydevice_id → gatewayIn-memory KVTTL ~90 s, refreshed by heartbeat
Key directorydevice_id → identity key, signed prekey, one-time prekeysStrongly consistent KVWhile device linked
Device listuser_id → signed device list + versionStrongly consistent KVWhile account exists
Mediacontent-addressed encrypted blobObject store + CDNTTL after download

Appendix C: Coordination Mechanisms — Quick Comparison#

MechanismPurposeLatencyFailure Behavior
Per-chat sequencer with epoch fencingOrder within chat< 1 ms in-memory + durable batchFailover reads last_seq, bumps epoch
Quorum inbox appendOne-tick durability+2–5 msSurvives single replica loss
Dedupe windowIdempotent sends< 1 msLoss → recipient dedupe catches duplicates
Membership / device-list versionsSecurity-correct routing0 (checked inline)Stale → reject and client refresh
Session registry TTLOnline routing< 1 msStale → push fallback, no loss

Appendix D: Protocol and Client Behavior#

  • Heartbeat every ~30–60 s, adaptive per carrier NAT timeout; server closes idle sockets after ~2 missed heartbeats.
  • Reconnect: jittered exponential backoff (base 1 s, cap 60 s); server may send reconnect_in hints during drains.
  • On reconnect: SYNC{after_inbox_seq} in pages of 100; ack in batches.
  • Outbox: unsent messages persisted locally with msg_id; retried until ACK; UI shows clock icon.
  • Prekeys: top up to ~100 when below ~20.
  • Media: encrypt locally with random key; upload; send message with URL, key and hash.

Appendix E: Observability#

E.1 Core Metrics#

MetricWhy
delivery.send_to_ack_ms_p99, delivery.ack_to_deliver_ms_p99Stage-split latency
client.stuck_single_tick_totalSilent loss proxy
probe.e2e_delivery_success{region}Synthetic end-to-end correctness
gateway.connections, gateway.reconnects_per_sEdge health
inbox.backlog_age_p99Offline tail
router.stale_membership_rejects, router.stale_device_list_rejectsRouting correctness
push.send_latency_p99{provider}Wake-up dependency
client.decrypt_failures_totalE2E health

E.2 Critical Alerts#

AlertThresholdSeverity
E2E probe failure< 99.99% over 10 min in a regionPage messaging core
Online delivery p99> 2 s for 5 minPage messaging core
Reconnect rate> 5× baseline for 2 minPage edge
Decrypt failures> 2× baseline for 15 minPage security + clients
Inbox append errors> 0.01%Page storage

E.3 Control Plane vs Data Plane#

Data plane: sockets, sends, inbox appends, deliveries, acks. Control plane: gateway deploys and drains, heartbeat policy, group caps, rate limits, retention settings, minimum client versions. Control-plane changes are canaried by region and by client cohort.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat You Add
< 100K concurrentOne region; Postgres inbox table; long-polling or WebSocketsClient IDs, per-chat seq
100K–10MGateway fleet; partitioned inbox store; Redis registry; pushQuorum appends, graceful drains
10M–500ME2E, sender keys, per-device multi-device; multi-region homingSynthetic probes, adaptive heartbeats
500M+Shared messaging platform, per-product confidentiality classesChannels product; residency capabilities

F.1 What You Don't Build on Day One#

  • Multi-region active routing
  • Per-device multi-device (start with one device per account, or server-readable companion sync if not E2E yet)
  • Channels/broadcast
  • Custom gateway runtime (use a proven WebSocket stack)

What you do build on day one: client message IDs, per-chat sequence numbers, cursor-based sync and delete-on-ack semantics. Those are the protocol commitments that every client in the wild will carry for years.

Appendix G: Fairness, Abuse and Cost#

  • Abuse without content: send-rate limits per account and per new account, fan-out limits on forwarding (e.g., forward to at most a handful of chats at once), user reports that include reported messages from the reporter's device.
  • Business tenants: per-tenant quotas, templates, quality scores; consumer traffic prioritized on shared tiers.
  • Battery fairness: one connection per device across products; collapse pushes.
  • Cost attribution: per-product share of connections, deliveries, media bytes; monthly showback.
  1. Loading the index…