Technologies referenced in this case study: Cassandra · Redis · Apache Kafka · DynamoDB
Related case studies: Real-Time Updates · Notification System · Message Queues · Blob Storage · File Sync · Real-time Updates pattern
How to Use This Case Study#
Organized for interview use first, reference second. Read once end to end, then return to weak spots.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Active Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failures) → Deep Dives 2 and 5 |
| Deep Dive | 3+ hrs | Everything, including the Principal Lens and appendices |
What is Chat Messaging? — Why interviewers pick this topic
Person-to-person and small-group messaging with delivery guarantees: messages sent while the recipient is offline are held and delivered later, in order, exactly once from the user's point of view, with sent/delivered/read receipts, across multiple devices per user — and, in WhatsApp's case, end-to-end encrypted so the server cannot read them.
Before vs After — a phone comes back online after a flight:
Naive design (server-side history table, client polls, server-assigned timestamps):
t=0: Phone reconnects after 9 hours. 1,400 messages across 60 chats waiting.
t=+1s: Client asks "give me everything since my last timestamp." Clock skew between
two chat servers means 30 messages have timestamps earlier than the client's
last-seen value. They are never delivered.
t=+3s: Server pushes 1,400 messages; the connection drops at message 900. Client
reconnects and re-requests from its last timestamp: 900 duplicates render.
t=+5s: Group messages appear out of order — replies before the questions.
t=+10s: Sender's phone still shows one grey tick for messages delivered hours ago,
because receipts were sent to a device that has since been replaced.
Staff design (per-device inbox with sequence numbers, client IDs, acks):
t=0: Same reconnect.
t=+1s: Client says "my inbox cursor is 88,214." Server streams 88,215 onward in
pages of 100. Order within each chat follows the chat's sequence number.
t=+3s: Connection drops at 900. Client acked through 88,900 in batches; resumes
from 88,901. Client-generated message IDs dedupe anything re-sent.
t=+4s: Server deletes acked messages from the inbox (it is a relay, not an archive).
t=+5s: Delivery receipts flow back to senders' devices as their own inbox entries.
Why interviewers reach for this question: it looks like a WebSocket question. It is actually a delivery-semantics and state-ownership question: where does the durable copy of a message live, who assigns order, how do acknowledgments compose into user-visible ticks, and what happens to all of that when the server is not allowed to read the content (E2E) and a user has five devices.
Mechanics Refresher: Delivery Models
| Model | How It Works | Pros | Cons |
|---|---|---|---|
| Polling | Client asks for new messages every N seconds | Simple; stateless servers | Latency = N/2; battery and request cost scale with users, not messages |
| Long polling | Request held open until a message arrives or ~30–60 s timeout | Works through most proxies | Reconnect churn; one message per round trip |
| Persistent connection (WebSocket, MQTT, custom TCP/TLS) | Server pushes over a held socket; heartbeats keep NAT open | Sub-100 ms push; efficient | Stateful edge fleet; reconnect storms; connection-to-server routing |
| OS push (APNs/FCM) | Platform wakes the app or shows a notification | Works when app is backgrounded/killed | Best-effort, rate-limited, payload-limited (~4 KB), no ordering |
| Store-and-forward inbox | Server durably queues per recipient until acked | Offline delivery; exactly-once display via dedupe | Storage for undelivered backlog; cursor management |
For most production systems: persistent connection when foregrounded, OS push to wake when backgrounded, and a durable per-device inbox underneath both. The transport is almost never the interview question — ordering, acks, the inbox, groups and multi-device under E2E are.
Executive Summary
If you only read one section, read this. Everything else in the case study elaborates on the contrast and the positions below.
What This Interview Actually Tests#
Chat is not a WebSocket question. Everyone can hold a socket open.
It is a store-and-forward delivery problem with per-conversation ordering, end-to-end acknowledgments, and — under E2E encryption — a server that must route what it cannot read. It tests:
- Whether you define delivery semantics precisely: at-least-once transport + client dedupe = exactly-once display
- Whether you know who assigns order (a per-chat sequencer) and why wall-clock timestamps fail
- Whether you treat the server as a relay with an inbox, not an archive, and know which intent that implies
- Whether you understand what E2E encryption moves to the client: group fan-out, multi-device sync, search, backups, abuse detection
The key insight: The server's job in WhatsApp-style chat is to hold an encrypted envelope until every intended device has acknowledged it, then forget it. Every hard design question — ordering, receipts, groups, multi-device — is a question about who holds which cursor and who can read what.
The L5 vs L6 vs L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | WebSocket servers + a messages table | Asks "Is the server the system of record or a relay? Is it E2E?" and commits | Asks what the E2E decision does to every adjacent product — backups, search, moderation, compliance — and who signs that |
| Delivery | "Use a queue so messages aren't lost" | At-least-once to a per-device inbox; client-generated message IDs; acks advance a cursor; exactly-once display | Defines delivery SLOs (p99 online delivery, max offline retention) as contracts with product and legal |
| Ordering | Sort by timestamp | Per-chat monotonic sequence from the chat's owner shard; clients order by (chat, seq) | Chooses the ordering guarantee per product surface and documents where the company does not promise order |
| Groups | Fan-out on write to all members | Fan-out on write for small groups (≤ ~1K), sender-key encryption so the sender encrypts once; caps as design constraints | Group size caps as product policy with cost and abuse modeling; large broadcast = a different product |
| Multi-device | "Sync messages to all devices" | Each device is a first-class endpoint with its own keys and inbox; sender encrypts per device | Device-linking security model, history-transfer policy, and key-change UX signed by security and product |
| Storage | Keep all messages forever | Delete after delivery; bounded offline retention (~30 days) | Data-minimization as a company position; what that means for law-enforcement requests and for revenue features |
Why "first move" separates levels
L5: Designs Slack and calls it WhatsApp: a messages table keyed by conversation, history served from the server, search over server data. That is a valid design for a different intent — server-as-record chat — and it silently assumes the server can read messages.
L6: "There are two very different chat systems. In one, the server is the system of record — history, search, compliance — like Slack or Discord. In the other, the server is a relay — it holds encrypted envelopes until delivered and then deletes them — like WhatsApp or Signal. I'll design the relay with E2E, because that's what 'WhatsApp' implies, and it changes where groups, sync and search live."
L7: "Choosing E2E is a company-level decision: it removes server-side search, server-side spam filtering on content, and easy cross-device history. I'd want Security, Legal, Trust & Safety and Product aligned on that before engineering commits, because it's a one-way door in public perception."
Why "ordering" separates levels
L5: "Each message gets a timestamp; clients sort by timestamp." Server clocks skew by milliseconds to seconds; client clocks by minutes. Two messages sent 5 ms apart in a group can render in either order on different devices, and a skewed server can assign a timestamp earlier than a client's sync cursor so the message is never fetched.
L6: "Order is assigned by the chat's owner — one sequencer per conversation, a monotonic counter stored with the chat. Clients render by (chat_id, seq). Timestamps are display-only. For a 1:1 chat under E2E the server still sequences envelopes, because it can order what it can't read."
L7: Decides where the product explicitly does not promise order — e.g., across different chats, or between a message and a reaction — so teams stop building features that assume it.
Why "multi-device" separates levels
L5: "The server keeps messages and each device syncs." Under E2E, the server cannot re-encrypt for a new device.
L6: "Every device has its own identity key. A sender encrypts a message once per recipient device — and once per their own other devices — so a 1:1 message to someone with 4 devices, sent from a user with 3, is ~6 ciphertexts. Each device has its own inbox and its own cursor. History on a new device is a client-to-client transfer, not a server replay." WhatsApp's 2021 multi-device architecture moved from "phone as the relay for web clients" to exactly this per-device model.
L7: Owns the security model for linking devices (QR-based linking, key-change notifications, max linked devices) as a policy that security and product sign, and prices the extra fan-out it creates.
The Staff Positions#
| Position | Rationale |
|---|---|
| Server is a relay, not an archive (for E2E consumer chat) | Delete on ack; bounded offline retention; minimizes cost, breach surface and legal exposure |
| Per-device inbox with a monotonic cursor | Offline sync is "give me everything after cursor N"; resume after disconnect is exact |
| Client-generated message IDs | Retries are safe; dedupe at recipient; exactly-once display on at-least-once transport |
| Per-chat sequencer for order | Timestamps lie; one owner per chat assigns seq |
| Fan-out on write for groups up to a cap | Recipients read one inbox; cap (~1K) bounds write amplification; broadcast channels are a different product |
| Receipts are messages | Delivered/read receipts travel through the same inbox path in reverse; no separate receipt database |
| Connection gateways are stateless about messages | Gateways hold sockets; message state lives in inboxes; a gateway crash loses no data |
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Private E2E messaging (WhatsApp, Signal) | Server can't read content; offline delivery; small groups | Relay with per-device inboxes; client-side encryption and fan-out; delete on ack | Lost message on inbox failure; out-of-order delivery; key mismatch after device change | Every message delivered to every device exactly once (displayed), in per-chat order |
| Community / workspace chat (Slack, Discord) | History, search, huge channels, integrations | Server-side history as system of record; fan-out on read for large channels | Hot channels; search index lag; permission leaks | History durable forever; order per channel |
| Regulated / enterprise messaging (finance, healthcare) | Retention, eDiscovery, legal hold | Server-readable storage, immutable archive, compliance export | Retention policy violation; unauthorized access | Nothing deleted before retention; auditable |
🎯 Staff Move: "I'll design private E2E messaging — WhatsApp's intent. The server is a relay that holds encrypted envelopes per device until they're acked. That choice pushes group fan-out encryption, multi-device sync and search to the client, and I'll call out each place it does. If you want Slack-style history and search, that's a different system and I'd design the storage completely differently."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Relay vs System of Record | Delete after delivery (cheap, private) or keep history server-side (search, sync, compliance)? |
| 2 | Ordering: Sequencer vs Timestamps vs Causal | Who assigns order, and what does it cost per message and on failover? |
| 3 | Delivery Semantics and Receipts | At-least-once + dedupe vs attempts at exactly-once; how acks compose into sent/delivered/read |
| 4 | Group Fan-Out: Write vs Read, and E2E Cost | Per-recipient inbox writes vs shared log; pairwise encryption vs sender keys; where the cap sits |
| 5 | Multi-Device: Primary-Relay vs Per-Device Endpoints | Phone forwards to companions, or every device is a peer with its own keys and inbox? |
In the Wild: Real Production Systems#
Why this section belongs here: these are well-known, publicly documented systems; citing them shows you understand how production chat actually evolved.
WhatsApp — Erlang Relay, Signal Protocol, Multi-Device#
WhatsApp's backend is famously built on Erlang/BEAM (on FreeBSD in its early years), and its engineers publicly described handling ~2 million concurrent TCP connections on a single server in 2012. WhatsApp completed rollout of end-to-end encryption using the Signal Protocol in 2016, and in 2021 published its multi-device architecture, in which each linked device has its own identity key and senders encrypt per device, so companions no longer depend on the phone being online. WhatsApp states that undelivered messages are deleted from its servers after 30 days. It was also widely reported to have served several hundred million users with a team of ~50 engineers at the time of its 2014 acquisition.
Staff insight: WhatsApp's design is a lesson in doing less on the server: a relay that holds encrypted envelopes, a runtime built for millions of lightweight connection processes, and complexity pushed into clients where E2E requires it. Say "the server is a relay" and you've made the most important decision.
Discord — Server-Side History at Trillion-Message Scale#
Discord has written publicly about storing messages in Cassandra and later migrating to ScyllaDB, with messages partitioned by (channel_id, time bucket) and ordered by Snowflake IDs (time-sortable 64-bit IDs). Their posts describe hot partitions from very large, busy channels and the operational pain of compaction and GC pauses that motivated the move.
Staff insight: This is the system-of-record intent. The partition key embeds a time bucket because unbounded channel partitions become hot and huge. Contrast it with WhatsApp: same product category, opposite storage philosophy, because the intent is different.
Facebook Messenger — Iris, a Per-User Ordered Update Queue#
Facebook Engineering described Iris (2014) for Messenger: a totally ordered queue of messaging updates per user, with each device keeping a pointer (sequence ID) into it. Devices sync by asking for everything after their pointer, which made mobile sync fast and reliable over lossy networks.
Staff insight: "Per-user ordered queue + per-device cursor" is the generalizable pattern. It turns sync into one integer comparison and makes resume-after-disconnect exact. Name it in the interview.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "WebSockets for real-time" | "The recipient is offline for 3 days. Where is the message?" | Store-and-forward thinking |
| "Messages have timestamps" | "Two people send at the same millisecond in a group. What order does everyone see?" | Sequencer ownership |
| "We guarantee exactly-once delivery" | "The ack is lost after the client stored it. What happens?" | Precise semantics |
| "Fan-out to group members" | "The group has 1,000 members, each with 3 devices, and it's E2E." | Write amplification + encryption cost |
| "Sync to all devices" | "The server can't decrypt. How does a new laptop get history?" | E2E implications |
| "Store messages in Cassandra" | "For how long? Why?" | Relay vs archive |
System Architecture Overview#
Reading the diagram: clients own encryption; the core routes opaque envelopes. The Router asks the chat's Sequencer for the next
seq, looks up recipient devices (1:1 or group), and appends one entry per device to that device's inbox before acknowledging the sender. If the device is online (per the session registry), the gateway pushes immediately; otherwise the push sender wakes the app. The inbox is the only durable message state, and entries are deleted when the device acks.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Transport | "WebSockets" | "Persistent TLS socket in foreground, OS push to wake, per-device inbox underneath both." |
| Durability | "Store all messages" | "Durable per-device inbox; the sender's single tick means 'in the inbox, replicated'; delete on device ack." |
| Semantics | "Exactly-once" | "At-least-once transport, client-generated IDs, dedupe on receipt: exactly-once display." |
| Ordering | "Timestamps" | "Per-chat sequencer; render by (chat, seq); timestamps for display only." |
| Receipts | "Update a status column" | "Receipts are small messages sent back through the sender's inbox; ticks are a client-side fold." |
| Groups | "Fan out" | "Fan-out on write to member devices up to ~1K members; sender keys so the payload is encrypted once." |
| Multi-device | "Sync" | "Each device is an endpoint with its own keys and inbox; sender encrypts per device; history transfer is client-to-client." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Messages/day (WhatsApp, public order of magnitude, 2020) | ~100B | ≈ 1.2M/s average, ~3–5M/s peak (New Year's Eve) |
| Concurrent connections per server (WhatsApp 2012, public) | ~2M | Erlang lightweight processes; modern planning figure ~0.5–1M per host |
| Heartbeat interval | ~30–60 s (under typical mobile NAT timeouts) | Connection-keepalive cost: 1B devices / 45 s ≈ 22M heartbeats/s |
| Undelivered message retention | 30 days (WhatsApp public FAQ) | Bounds inbox storage |
| Group size cap | 1,024 (WhatsApp, raised from 512 in 2022) | Bounds fan-out and E2E cost |
| Linked companion devices | up to 4 (plus the phone) | Per-device fan-out multiplier |
| Envelope size (text) | ~200 B–1 KB | 1.2M/s × 1 KB ≈ 1.2 GB/s inbox write bandwidth before replication |
| Media | Encrypted blob uploaded once; message carries URL + key (~200 B) | Media never flows through the message path |
| OS push payload limit | ~4 KB | Push carries a wake signal, not the message |
| Online-to-online delivery | p99 < ~300–500 ms in-region | The "instant" feel |
| Inbox entries per group message | members × devices (1K × ~2 = ~2K) | Write amplification you must budget |
Interview Walkthrough
The most common mistake: candidates spend 15 minutes on WebSocket handshakes and load balancer stickiness, then hand-wave "and messages are stored in Cassandra." The transport is two sentences. The phases below get you to the inbox, ordering and E2E by minute 12.
Phase 1: Requirements & Framing (2–3 min)#
Functional core in one breath:
"1:1 and group messages with text and media, delivered in real time when online and stored until delivered when offline, with sent, delivered and read receipts, on multiple devices per user."
Then the framing question that decides the design:
"The big fork: is the server the system of record — history, search, compliance — or a relay that holds messages until delivered? For WhatsApp it's a relay with end-to-end encryption. I'll design that, and I'll call out what E2E pushes to the client."
Then numbers:
"~2B users, ~500M–1B concurrently connected at peak, ~100B messages/day — ~1.2M/s average, ~3–5M/s at peak. Average group message fans out to ~5–10 recipient devices, so inbox writes are ~5–10M/s at peak. Online delivery p99 under ~500 ms in-region. Offline retention 30 days."
🎯 Staff Move: The relay-vs-record question in the first 60 seconds is the level signal. It tells the interviewer you know there are two chat systems and you are choosing one on purpose.
Phase 2: Core Entities & API (1–2 min)#
- Device:
device_id,user_id,identity_public_key,prekeys[],push_token - Chat:
chat_id(1:1 = hash of sorted user IDs; group = generated),last_seq,owner_shard - Envelope:
msg_id(client-generated UUIDv7),chat_id,seq,sender_device,recipient_device,ciphertext,type ∈ {message, receipt, key_change, membership},server_ts - InboxEntry:
(device_id, inbox_seq)→ envelope; deleted on ack - GroupMembership:
group_id→ member user IDs → their device IDs;version
Socket frames (binary, compact):
→ SEND {msg_id, chat_id, envelopes: [{recipient_device, ciphertext}], membership_version}
← ACK {msg_id, seq, server_ts} # single tick: durably in all inboxes
← DELIVER{inbox_seq, envelope} # pushed or fetched
→ RECV_ACK {up_to_inbox_seq} # cursor advance; server deletes ≤ cursor
→ SYNC {after_inbox_seq, limit: 100} # reconnect / backfill
→ RECEIPT{msg_id, chat_id, kind: delivered|read} # travels back as an envelope
🎯 Staff Move: "Notice the sender sends one ciphertext per recipient device — the server can't fan out plaintext because it doesn't have any. And the recipient acks a cursor, not individual messages, so one ack covers a batch of 100."
Phase 3: High-Level Architecture (≤ 5 min)#
Three sentences:
"Gateways hold sockets and nothing else. The Router gets a sequence number from the chat's sequencer, appends one inbox entry per recipient device, acks the sender only after those appends are durable, then pushes to online devices or wakes offline ones through APNs/FCM. Devices ack a cursor; the server deletes everything up to it."
Phase 4: Transition to Depth#
"Transport is standard — persistent socket plus OS push. The parts I'd want to go deep on are: how the inbox and acks give exactly-once display, how ordering works in groups, and what E2E does to groups and multi-device. I'd start with the inbox because everything else builds on it."
Phase 5: Deep Dives (25–30 min)#
| Deep Dive | Time | What You Must Land |
|---|---|---|
| Inbox, acks, exactly-once display | 6–8 min | Durable append before sender ack; client msg IDs; cursor acks; delete on ack; 30-day bound |
| Ordering | 4–6 min | Per-chat sequencer; seq gaps; failover of the sequencer |
| Receipts | 3–4 min | Receipts as envelopes; ticks as a fold; read-receipt privacy setting |
| Groups under E2E | 6–8 min | Fan-out on write per device; sender keys; membership versions; cap |
| Multi-device | 4–6 min | Per-device keys and inboxes; linking; history transfer |
| Connection edge | 3–4 min | Session registry, heartbeats, reconnect storms |
Exactly-once display, as you'd say it:
"The sender's client generates msg_id before sending and retries with the same ID until it gets an ACK. The Router dedupes by (sender_device, msg_id) for a short window so retries don't create duplicate inbox entries. If a duplicate still slips through — say, an ACK was lost after the append — the recipient's local database has a unique constraint on msg_id and drops it. Transport is at-least-once; display is exactly-once."
Phase 6: Wrap-Up (2–3 min)#
"To summarize: a relay with durable per-device inboxes, at-least-once transport with client IDs for exactly-once display, a per-chat sequencer for order, receipts as messages, fan-out on write with sender keys for groups up to ~1K, and per-device keys for multi-device. What I'd monitor first: online delivery p99, inbox backlog age, and reconnect rate per gateway. What I'd build next: client-to-client history transfer for new devices and encrypted backups. What I skipped: channels/broadcast to millions, which is a fan-out-on-read product."
Common Timing Mistakes#
| Mistake | Time Lost | Fix |
|---|---|---|
| WebSocket vs long-poll vs SSE debate | 5–8 min | "Persistent socket + OS push. Moving on." |
| Designing Slack history and search | 10+ min | Commit to relay vs record in Phase 1 |
| Signal Protocol cryptography details (X3DH, double ratchet math) | 8 min | "Sessions are established from prekeys; each message advances a ratchet. What matters for the system is per-device fan-out." |
| Presence ("online/last seen") in depth | 5 min | "Presence is lossy, best-effort, subscribed per open chat." |
| No numbers | Whole interview | 100B/day, 1.2M/s, ×devices in Phase 1 |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Chat is everybody's first "real-time" design, which makes it a strong discriminator: nearly every candidate produces something that works in a demo, and the level shows in how precisely they define "delivered" and where they put state.
| Question | L5 Instinct | Staff Reality |
|---|---|---|
| "Is the message delivered?" | Yes, the server has it | Delivered = the device acked; the server having it is "sent" |
| "What order?" | Timestamp order | Per-chat sequence; no order across chats |
| "Where is the message?" | In the messages table | In N device inboxes until each acks; then nowhere on the server |
| "Can we search?" | Elasticsearch over messages | Only on-device under E2E |
| "What if a gateway dies?" | Messages lost? | Nothing lost — gateways hold sockets, not messages |
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"When the sender sees two ticks, exactly which device has exactly which bytes?"
Answering this precisely forces every decision:
- One tick = the envelope is durably appended to every recipient device's inbox (replicated). Not "the server received it."
- Two ticks = at least one (WhatsApp: all, in the multi-device model it's per-device and folded) recipient device acked receipt.
- Blue ticks = the recipient opened the chat and the client sent a read receipt — which the user may disable.
🎯 Staff Move: "I'll define the ticks precisely first, because each one is a durability promise. One tick is my promise that the message survives any single server failure. I don't send it until the inbox append is replicated."
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Intent 1 — Private E2E messaging. The server is an untrusted relay. It sees metadata (who, when, size) but not content. Its durable state is per-device inboxes of undelivered envelopes. Clients own history, search, and — for groups — encryption fan-out. Scale is dominated by connection count and inbox write amplification, not by stored history.
Intent 2 — Community/workspace chat. The server is the system of record: history is the product (scrollback, search, pinned messages, integrations, bots). Channels can have 100K+ members, so fan-out on write is impossible for large channels; readers pull from a per-channel log. Storage is partitioned by (channel, time bucket). Content moderation and search are server-side.
Intent 3 — Regulated messaging. Retention and eDiscovery are hard requirements: messages must be retained immutably for years and exportable. The server must be able to read content (or hold escrowed keys). Deletion is governed by policy, not user action.
| Dimension | Private E2E | Community | Regulated |
|---|---|---|---|
| Server reads content | No | Yes | Yes (or escrow) |
| Durable server state | Undelivered envelopes only | Full history | Full history + immutable archive |
| Group fan-out | On write, per device, client-encrypted | On read for large channels | On read |
| Search | On device | Server index | Server index + legal export |
| Retention | Delete on ack, ≤ 30 days | Forever (product) | Policy (e.g., 7 years) |
| Hardest problem | Multi-device + groups under E2E | Hot channels, history storage | Retention correctness, access audit |
2.2 When NOT to Use This Design#
| Situation | Why the Relay Design Is Wrong | Use Instead |
|---|---|---|
| Users expect full history on any new device instantly | Relay deletes on ack; history lives on clients | System-of-record chat (or encrypted server backup as a separate feature) |
| Channels of 10K–1M members | Fan-out on write × devices is millions of writes per message | Per-channel log, fan-out on read (Discord/Slack model) |
| Compliance retention required | Deleting on ack violates retention | Regulated archive |
| One-way broadcast (announcements, notifications) | No conversation, no receipts needed | Notification System |
| Live collaborative content | Needs merge semantics, not message order | Collaborative Editing |
| In-app support chat at < 10K concurrent users | A database table + polling every 5 s is fine | Keep it simple |
🎯 Staff Move: "If this were an in-app support chat with a few thousand concurrent users, I'd use a Postgres table and long polling and ship it in a week. The relay-with-inboxes architecture earns its complexity at hundreds of millions of connected devices and when the server must not read content."
2.3 What the Interviewer Leaves Underspecified#
| Underspecified | Why It Matters | What to Say |
|---|---|---|
| E2E or not | Everything about groups, sync, search | "WhatsApp implies E2E; I'll design for it." |
| Group size | Fan-out strategy | "Cap ~1K; broadcast channels are a separate product." |
| History on new devices | Relay vs record | "Client-to-client transfer; optional encrypted backup." |
| Retention of undelivered | Inbox storage | "30 days, then drop and tell the sender nothing further." |
| Receipt semantics | Privacy and fan-out | "Delivered always; read receipts user-configurable." |
| Media | Message path size | "Encrypted blob out of band; message carries pointer + key." |
| Global vs regional | Latency and residency | "Users homed to a region; cross-region routing for cross-region chats." |
2.4 Precise Terminology#
| Term | Meaning Here | Common Confusion |
|---|---|---|
| Envelope | Encrypted payload + routing metadata for one recipient device | Not "the message" — one message becomes many envelopes |
| Inbox | Per-device durable queue of undelivered envelopes | Not a mailbox of history |
| Cursor / inbox_seq | Monotonic position in a device's inbox | Different from the chat seq |
| Chat seq | Monotonic order within one chat, assigned by its sequencer | Not a global order |
| Sent (one tick) | Durably appended to all recipient inboxes | Not "left the phone" |
| Delivered (two ticks) | Recipient device acked | Not "read" |
| Read (blue) | Recipient opened chat; client sent read receipt | Optional, privacy-controlled |
| Sender key | Symmetric key a sender distributes to group members so group messages are encrypted once | Not per-recipient pairwise encryption |
| Prekey | One-time public key uploaded by a device so others can start a session while it's offline | Consumed on use; must be replenished |
| Identity key | Long-term device public key | Change triggers "security code changed" notice |
| Companion device | Linked device (web, desktop, tablet) with its own identity | Not a mirror of the phone |
| Session registry | Map of device → current gateway, with TTL | Soft state; rebuilt on reconnect |
3. The Fault Lines#
3.1 Fault Line 1: Relay vs System of Record#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Relay: delete on ack, bounded retention (Staff default for E2E) | Storage ∝ undelivered backlog, not history; minimal breach and legal surface; cheap | New device has no history; server search impossible; lost phone = lost history unless backed up | Users (history portability); client team (backup, transfer) |
| System of record: keep all history | Any device sees full history; server search; moderation on content | Storage grows forever (100B msgs/day × 1 KB ≈ 100 TB/day raw); breach surface; incompatible with E2E unless encrypted history + client keys | Infra budget; Security; Legal |
| Relay + optional encrypted backup | History recoverable; server still can't read it | Backup key management is the user's burden; restoring is heavy | Users (remember a key/password); backup platform team |
The arithmetic that settles it: at 100B messages/day, a system-of-record design stores ~36 PB/year raw before replication. A relay stores only the backlog — if 95% of messages are delivered within minutes and the tail is bounded to 30 days, the inbox store holds on the order of days of the offline fraction, typically 1–2 orders of magnitude less.
The Staff default: relay with delete-on-ack and a 30-day bound for undelivered envelopes; history on devices; optional encrypted backup as a separate product surface with its own key-management design.
When to deviate: Intent 2 or 3. If the product needs server search, history on any device, or compliance retention, you are building a different system — say so.
🎯 Staff Move: "The cheapest byte to protect is the one we don't keep. With E2E we can't read history anyway, so keeping it server-side buys us cost and risk but no features."
3.2 Fault Line 2: Ordering — Sequencer vs Timestamps vs Causal#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Client timestamps | No coordination | Clocks skew by minutes; malicious clients reorder | Users (confusing order) |
| Server timestamps | Easy | Skew across servers (ms–s); ties; sync cursors based on time miss messages | Users; support |
| Per-chat sequencer (monotonic seq) (Staff default) | Total order per chat; gap detection; exact sync | One owner per chat = hot spot for very busy groups; failover must not reuse seqs | Messaging core team owns sequencer and failover |
| Causal ordering (vector clocks / "reply-to" links) | Correct causality without a central sequencer | Complex; still needs a tiebreak for display | Client team |
| Global sequencer | Total order across everything | Throughput ceiling; single point of failure | Everyone |
How the sequencer works:
# chat owner shard = hash(chat_id) % N (consistent hashing; see foundations)
on SEND(chat_id, msg_id, envelopes):
if dedupe.seen(sender_device, msg_id): return previous ACK
seq = chat.last_seq + 1 # in-memory, owned by this shard
persist: chat.last_seq = seq AND inbox appends for each envelope # one durable batch
ack sender {msg_id, seq}
Failover rule: the new owner reads last_seq from durable storage and fences the old owner (epoch number in every write), so two owners can never issue the same seq. A client seeing a gap (seq 41 then 43) waits briefly (~2 s) for 42 or requests a re-sync; gaps can also be legitimate (a message addressed only to other devices).
When to deviate: very large groups where one sequencer is too hot — split into the community/channel model with a partitioned log per channel. For 1:1 chats, a lighter option is ordering per (sender → recipient) direction, since cross-direction interleaving within milliseconds rarely matters to humans.
🎯 Staff Move: "Order is a promise per chat, never across chats. One sequencer per chat, fenced by epoch on failover. Timestamps are for the UI label, not for sorting."
3.3 Fault Line 3: Delivery Semantics and Receipts#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| At-most-once (fire and forget) | Simple | Messages lost on any failure | Users |
| At-least-once + client IDs + recipient dedupe (Staff default) | Robust across retries, reconnects, failovers; exactly-once display | Dedupe window and client unique constraint required | Client team (dedupe), core team (dedupe window) |
| "Exactly-once" via distributed transactions | Sounds good | Impossible end-to-end across a lossy network; costs latency | Everyone, for a promise you can't keep |
Receipts are just messages traveling the other direction. That gives them durability, ordering and multi-device delivery for free — and avoids a separate "receipt status" database that would need to be updated ~3× per message (sent, delivered, read) × devices.
Receipt cost control:
- Batch: one read receipt covers "read up to seq N" in a chat — not one per message.
- Groups: delivered/read receipts in groups are aggregated client-side into "read by 14 of 20"; for large groups, receipts can be sampled or computed on demand.
- Privacy: read receipts are user-configurable; turning them off means the client simply doesn't send them.
When to deviate: community chat doesn't need per-recipient delivered receipts at all — "unread count" from a per-user read cursor is sufficient.
🎯 Staff Move: "I don't promise exactly-once delivery — nobody can over a mobile network. I promise at-least-once delivery and exactly-once display, and I can explain every step that makes that true."
3.4 Fault Line 4: Group Fan-Out — Write vs Read, and the E2E Cost#
For a group of M members with an average of d devices each:
| Strategy | Server Writes per Message | Client Encryption Work | What Breaks | Who Pays |
|---|---|---|---|---|
| Fan-out on write, pairwise encryption | M × d inbox entries | Sender encrypts M × d times | Sender's phone CPU and upload for large groups (1K × 2 = 2K ciphertexts) | Sender's battery and bandwidth |
| Fan-out on write, sender keys (Staff default for groups ≤ ~1K) | M × d inbox entries (same ciphertext referenced) | Sender encrypts once with its sender key; sender key distributed pairwise once per member device (and rotated on membership change) | Membership change requires rekey; removed members must not decrypt future messages | Client team (rekey logic); core (membership versions) |
| Fan-out on read (shared group log + per-member cursor) | 1 write | Same sender-key encryption | Server must retain the log until all members read — relay retention becomes history; hot group logs | Storage; hot partitions |
The server-side optimization with sender keys: store the ciphertext once in a message blob keyed by (chat_id, seq), and write only small pointers (~50 B) into each member device's inbox. That cuts inbox bytes by ~20× for a 1 KB message.
Membership consistency matters for security: a SEND carries the sender's membership_version. If the server's version is newer (someone was added or removed), the Router rejects with stale_membership, and the client fetches the new member list, distributes/rotates sender keys, and resends. That prevents a removed member from receiving new messages and ensures a new member gets the key.
The cap is a design constraint, not a product whim. At 1,024 members × ~2 devices, one message = ~2K inbox pointer writes. A busy group sending 1 msg/s is 2K writes/s — fine. A 100K-member group would be 200K writes per message; that is a broadcast product.
🎯 Staff Move: "The group size cap is where the write amplification and the rekey cost meet. Above ~1K, I'd stop pretending it's a group and build a channel: fan-out on read, admin-only posting, no per-member receipts."
3.5 Fault Line 5: Multi-Device — Primary Relay vs Per-Device Endpoints#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Primary-device relay (phone decrypts and forwards to web/desktop) | Single identity; simple key management | Companion dead when phone is offline or battery dies; phone bandwidth | Users (reliability), phone battery |
| Per-device endpoints (Staff default; WhatsApp's 2021 model) | Each device works independently; server has one inbox per device | Sender encrypts per device (fan-out × devices); device list must be authentic; history transfer needed | Senders' CPU; key-directory team; client team |
| Server-held history re-encrypted per device | Easy sync | Server has plaintext or keys — breaks E2E | Security; users' privacy |
Mechanics of per-device:
- Each user has a device list signed by the primary device; senders verify it so the server can't silently add an eavesdropping "device."
- A 1:1 message is encrypted for each of the recipient's devices and each of the sender's own other devices (so your laptop sees what you sent from your phone).
- A new companion gets recent history by encrypted transfer from the primary device, not from the server.
- Linking is a QR scan on the primary — a physical-presence check.
When to deviate: non-E2E products should just use server history — per-device crypto fan-out buys nothing without E2E.
🎯 Staff Move: "With E2E, a 'user' is really a set of devices, and the device list is security-critical. If the server could append a device, it could read messages. So the device list is signed by the user's primary device and verified by senders."
4. Failure Modes & Operational Reality#
4.1 Reconnect Storm After a Gateway Fleet Deploy#
t=0 Rolling deploy restarts 10% of gateways at once (misconfigured batch size).
t=+1s ~80M sockets drop. Clients reconnect immediately (no jitter in one app version).
t=+5s TLS handshakes 40× normal; CPU saturated on remaining gateways.
t=+10s Session registry writes spike 50×; registry p99 2 ms → 900 ms.
t=+30s Router can't find sessions; treats online devices as offline; OS push spikes
to 30M/min; APNs/FCM throttle us.
t=+3min Messages delivered with 1–2 minute delays; "WhatsApp down" trends.
- Detection:
gateway.reconnects_per_s,tls.handshakes_per_s,registry.write_latency_p99,push.throttled_total,delivery.latency_p99. - Blast radius: global if the deploy is global; cellular-region-specific for carrier outages.
- Mitigation: halt the deploy; server-side reconnect hints (
retry_afterrandomized 0–60 s); prioritize handshake capacity; TLS session resumption. - Prevention: drain gateways gracefully (send
reconnect_inwith jitter before closing); deploy batch ≤ 1–2% of connections; client jittered backoff enforced in all versions with a minimum-version policy. - Owner: Edge/connection platform.
4.2 Inbox Store Partition Hot Spot#
- Scenario: a bot account or a very active 1K-member group floods inboxes; or a celebrity with millions of contacts triggers profile/status fan-out.
- What breaks: inbox partitions keyed by device are fine individually; the sequencer for one very busy group is hot (~thousands of messages/s on one chat owner).
- Detection:
sequencer.ops_per_s_max{shard},inbox.append_latency_p99. - Mitigation: per-chat and per-sender rate limits (see Rate Limiting); spam/abuse throttles on metadata (E2E means content can't be inspected server-side).
- Owner: Messaging core (sequencer), Integrity team (abuse limits).
4.3 Lost Message After Inbox Failover (The Silent One)#
t=0 Inbox replica set primary fails. Async replication lag was 400 ms.
t=+2s Replica promoted. Envelopes appended in the last 400 ms — already ACKed to
senders (one tick) — are gone.
t=+∞ Senders see one tick forever. Recipients never see the messages. No error.
- Detection:
inbox.acked_but_missing— hard to detect directly; proxy: senders' clients report "one tick older than 24 h while recipient was online" via telemetry (client.stuck_single_tick_total). - Blast radius: messages in the replication window for affected devices.
- Mitigation: ack the sender only after quorum (synchronous) replication of the inbox append; accept ~2–5 ms extra latency.
- Prevention: make "one tick = quorum-durable" a written invariant; chaos-test primary failure under load monthly.
- Owner: Messaging core / storage team.
4.4 Prekey Exhaustion#
- Scenario: a device offline for weeks receives many new-session requests (new contacts, group adds). Its uploaded one-time prekeys are consumed.
- What breaks: senders fall back to the signed prekey (weaker forward-secrecy properties for the first message) or fail session setup.
- Detection:
keys.prekeys_remaining{device}low;keys.fallback_to_signed_prekey_total. - Mitigation: clients top up to ~100 prekeys when below a threshold on each connect; server alerts on devices near zero.
- Owner: Key directory team.
4.5 OS Push Degradation#
- Scenario: APNs or FCM latency rises to minutes or throttles us.
- Blast radius: backgrounded/offline devices — messages wait in inboxes (not lost), delivered on next app open.
- Detection:
push.send_latency_p99{provider},push.throttled_total,delivery.latency_p99{device_state=background}. - Mitigation: collapse pushes per device (one wake-up for many messages), priority for 1:1 over group, back off per provider guidance.
- Owner: Notifications platform (see Notification System).
4.6 New Year's Eve Spike#
- Scenario: 3–5× message volume within ~5 minutes of midnight in each timezone, dominated by group and media messages.
- Mitigation: capacity for 5× per-region peak on sequencers, inboxes and media upload; media uploads prioritized below text; follow-the-midnight capacity shifting across regions.
- Owner: Capacity planning with messaging core.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Reconnect storm | gateway.reconnects_per_s | Region or global | Jittered reconnect hints; halt deploy | Edge platform |
| Gateway crash | gateway.connections drop | ~1M devices, seconds | Clients reconnect elsewhere; no data lost | Edge platform |
| Inbox primary failover | client.stuck_single_tick_total | Replication window | Quorum append before ACK | Messaging core |
| Hot sequencer | sequencer.ops_per_s_max | One busy chat | Rate limits; move chat to less loaded shard | Messaging core |
| Session registry slow | registry.write_latency_p99 | Online delivery → push fallback | Shed registry writes; cache | Edge platform |
| Prekey exhaustion | keys.prekeys_remaining | New sessions to that device | Client top-up | Key directory |
| Push provider degraded | push.send_latency_p99 | Background delivery delayed | Collapse, prioritize | Notifications |
| Stale group membership | router.stale_membership_rejects | One group's sends retry | Client refresh + rekey | Messaging core + client |
| Media store slow | media.upload_latency_p99 | Media messages | Text unaffected; retry uploads | Media platform |
🎯 Staff Move: "Look at what's absent: no failure in this table loses a message except the inbox failover — and that's exactly why the sender's single tick waits for a quorum write. Everything else delays delivery; nothing else loses it."
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Scoping | Features: 1:1, groups, media, receipts | Relay vs record; commits to E2E relay; names what E2E moves to clients | Frames E2E as a company decision with Security, Legal, T&S and Product sign-off |
| Delivery | Queue so nothing is lost | Per-device inbox, quorum append before one tick, cursor acks, delete on ack | Delivery SLOs and retention as external commitments |
| Semantics | "Exactly-once" | At-least-once + client IDs = exactly-once display | Standard for idempotent client protocols across products |
| Ordering | Timestamps | Per-chat sequencer with epoch fencing | Documents where order is not promised; stops teams from depending on it |
| Groups | Fan-out | Fan-out on write with sender keys; membership versions; cap rationale | Group cap and channel product as a portfolio decision with abuse modeling |
| Multi-device | Sync everywhere | Per-device keys and inboxes; signed device lists; client history transfer | Device-linking security policy; key-change UX; prices fan-out multiplier |
| Operations | Monitoring | Reconnect storms, stuck-tick telemetry, push degradation | Game days for gateway fleet loss; error budgets per tier; deploy policy |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Relay-vs-record framing | "Is the server the record or a relay? For WhatsApp, a relay." |
| Precise ticks | "One tick means quorum-durable in every recipient inbox." |
| Honest semantics | "At-least-once delivery, exactly-once display." |
| Sequencer ownership | "One sequencer per chat, fenced on failover." |
| E2E consequences | "E2E means the sender encrypts per device, search is on-device, and a new laptop gets history from the phone." |
| Cap as constraint | "Above ~1K it's a channel, not a group." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "Exactly-once delivery" with no mechanism | Promising the impossible |
| Sorting by timestamp | Ordering bugs in groups |
| Server-side search in an E2E design | Contradiction |
| Gateways that store messages | Gateway crash loses data |
| No offline path | Store-and-forward is the core of the product |
| Keep all messages forever with no reason | Cost and risk without a feature |
5.4 Common False Positives#
- Cryptography depth ≠ system design. Explaining the double ratchet in detail is impressive and rarely what's being evaluated; the system consequence (per-device fan-out) is.
- WebSocket scaling tricks ≠ delivery design. Knowing epoll limits doesn't answer "where is the message when the phone is off?"
- Kafka everywhere ≠ inbox design. A topic per user doesn't scale to billions of devices; a partitioned inbox store does.
- "Eventually consistent" ≠ acceptable for one tick. The sender's tick is a durability promise.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–4 min | Relay vs record, E2E, numbers |
| Entities and protocol | 4–7 min | Envelope, inbox, cursor, client msg ID |
| Architecture | 7–11 min | Gateways, router, sequencer, inboxes, push |
| Inbox and semantics | 11–18 min | Quorum append, dedupe, cursor ack, delete |
| Ordering and receipts | 18–24 min | Sequencer, gaps, receipts as messages |
| Groups under E2E | 24–31 min | Sender keys, membership versions, cap |
| Multi-device | 31–37 min | Per-device keys, signed device list, history transfer |
| Edge operations | 37–42 min | Reconnect storms, push |
| Wrap-up | 42–45 min | Metrics, next steps, skipped scope |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response |
|---|---|---|
| "Add message search" | E2E consistency | "On-device index; server search would require server-readable content." |
| "Add a 100K-member community" | Fan-out limits | "That's a channel: fan-out on read, admin posting, no per-member receipts." |
| "Add disappearing messages" | Client-enforced policy | "A timer attribute in the encrypted payload; clients delete; server deletes on ack anyway." |
| "Law enforcement asks for message content" | Relay implications | "We don't have content; we have limited metadata. Legal owns the response process." |
| "Add voice/video calls" | Different transport | "Signaling over the message path; media over WebRTC/TURN; separate system." |
| "Spam is up 5×" | Moderation without content | "Metadata signals: send rate, new-account fan-out, user reports (which include the reported messages, sent by the reporter)." |
6.3 What to Deliberately Skip#
| Topic | Why Skip | One-Liner If Asked |
|---|---|---|
| Signal Protocol math | Not the systems question | "X3DH for session setup from prekeys, double ratchet per message." |
| Load balancer choice | Standard | "L4 LB to gateways; consistent device-to-gateway not required." |
| Media transcoding | Separate pipeline | "Client compresses; encrypted blob upload; see Blob Storage." |
| Presence details | Lossy, best-effort | "Subscribed only for open chats; heartbeat-derived; no durability." |
| Stickers, reactions | Same envelope path | "Just message types." |
6.4 Follow-Up Questions to Expect#
- "The phone is offline for 40 days. What happens to messages sent on day 1?"
- "The sender's app crashes after sending but before the ACK. What does the user see on restart?"
- "Two members of a group send at the same moment. What order does everyone see?"
- "Someone is removed from a group. How do you guarantee they can't read later messages?"
- "A user links a new laptop. What does it show?"
- "A gateway with 1M connections dies. What happens?"
- "How many writes per second hit the inbox store at peak?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design WhatsApp."
Staff Answer
"The first fork is relay vs system of record. WhatsApp is a relay with end-to-end encryption: the server holds encrypted envelopes per device until acked, then deletes them. I'll design that.
Numbers: ~100B messages/day, ~1.2M/s average, ~3–5M/s peak, and each message fans out to ~5–10 device inboxes on average — so ~10–30M inbox appends/s at peak, but tiny ones if the ciphertext is stored once and inboxes hold pointers. Connections: ~500M–1B concurrent devices; at ~1M per gateway, ~1K gateways.
Core: gateways hold sockets only; a router asks the chat's sequencer for a seq, appends to each recipient device inbox with quorum durability, ACKs the sender — that's one tick — pushes to online devices or wakes offline ones. Devices ack a cursor; the server deletes. E2E pushes encryption fan-out, history and search to the client, and I'll show where."
Why this is L6:
- Commits to an intent with a reason
- Sizes connections and write amplification separately
- Defines one tick precisely
What L7 adds:
- Names E2E as a cross-functional commitment (Legal, T&S, Product) before engineering
- Prices the gateway fleet and inbox store and names the dominant cost (connections, not storage)
Drill 2: Core Mechanic — The Lost ACK#
Prompt: "Alice sends a message. The server stores it, but the ACK never reaches Alice. Walk me through it."
Staff Answer
"Alice's client has the message in its local outbox with msg_id = m1 and a clock icon. No ACK within ~10 s or on reconnect, it resends m1. The Router checks its dedupe window keyed by (alice_device, m1) — say a 24-hour TTL in a fast KV — finds it, and returns the original ACK {m1, seq 41} without appending again. One tick appears.
If the dedupe entry was lost (e.g., dedupe store failover), the Router would append a second envelope with a new seq. Bob's client has a unique constraint on msg_id, so the second copy is dropped on arrival and not rendered. The cost is one wasted inbox entry and a seq gap Bob's client can ignore. Display remains exactly-once."
Why this is L6:
- Two layers of dedupe with clear roles
- Handles the failure of the dedupe layer itself
- Explains seq gaps as benign
What L7 adds:
- Makes the client-ID/idempotency protocol a company-wide standard for all mobile write APIs
- Measures duplicate-drop rate on clients as a health metric for the server dedupe layer
Drill 3: Make It Concrete — Inbox Sizing#
Prompt: "How big is the inbox store?"
Staff Answer
"Inbox holds only undelivered envelopes. Assume 1.2M messages/s average, ~6 device-deliveries per message → ~7M inbox entries/s. If entries are pointers (~64 B) and the ciphertext blob (~1 KB) is stored once per message, the write bandwidth is ~450 MB/s of pointers + ~1.2 GB/s of blobs.
Residency: ~90% of deliveries complete within a minute (online devices), ~9% within hours (backgrounded), ~1% sit for days. Steady-state entries ≈ 7M/s × (0.9 × 30 s + 0.09 × 3 h + 0.01 × 3 days) ≈ 7M × (27 + 972 + 2,592) s ≈ 25B entries — the long tail dominates. At ~64 B that's ~1.6 TB of pointers plus blobs for undelivered messages (~several TB), ×3 replication. The 30-day cap bounds the worst case. Compare ~100 TB/day raw if we kept history."
Why this is L6:
- Uses residency time, not throughput, to size storage
- Notices the long tail dominates
- Contrasts with system-of-record cost
What L7 adds:
- Evaluates whether shortening the retention cap (e.g., 30 → 14 days) is worth the product cost, with data on how many messages are delivered after day 14
Drill 4: Dependency Down — Push Provider Outage#
Prompt: "FCM is degraded for an hour. What happens?"
Staff Answer
"Online Android devices are unaffected — they're on our socket. Backgrounded Android devices don't get woken, so their messages wait in inboxes and arrive on next app open. No loss. Mitigations: collapse wake-ups to one per device per window, prioritize 1:1 over group wake-ups when throttled, and show an in-app banner if we detect delivery-latency degradation. Alert on push.send_latency_p99{provider=fcm} and route to the notifications platform. The senders keep seeing one tick until delivery — which is correct."
Why this is L6:
- Distinguishes online vs background devices
- No data loss because the inbox is the source of truth
- Prioritizes within a throttled budget
What L7 adds:
- Tracks dependency on two platform push providers as a strategic risk; evaluates persistent background connections where OS policy allows
Drill 5: Hot Key — The Viral Group#
Prompt: "A 1,024-member group is sending 50 messages/s during a live sports match."
Staff Answer
"50 msg/s × ~2K device inboxes = ~100K pointer appends/s from one chat, plus 50 seq increments/s on one sequencer — the sequencer is fine (a single in-memory counter handles 100K+/s), the inbox fan-out is spread across devices so it's spread across partitions. The real costs are client-side: every member's phone receives 50 msg/s — battery and data — and delivered/read receipts would be 50 × 1K × 2 per second if sent individually.
So: receipts in large groups are batched as 'read up to seq N' per chat per device, sent at most every few seconds; per-group send rate limits (e.g., 20 msg/s) with slow-mode as an admin option; client coalesces rendering. The inbox pointers are cheap because the ciphertext is stored once."
Why this is L6:
- Finds the actual hot spots (receipts, client battery), not the imagined one
- Batching receipts by cursor
- Product-level rate controls
What L7 adds:
- Decides whether live-event groups should be a distinct product (channels) with different economics
Drill 6: Multi-Tenant — Business Messaging#
Prompt: "Businesses want to message millions of customers through our platform."
Staff Answer
"Businesses are tenants with very different traffic shapes: bursty campaign sends to millions of 1:1 chats. Isolate them: a separate business ingestion API with per-tenant quotas and templates (pre-approved message types to limit spam), queued sends drained at a rate the inbox tier can absorb without affecting consumer latency, and per-tenant quality scores based on block/report rates that throttle bad senders. Consumer traffic always has priority on shared inbox and push capacity."
Why this is L6:
- Tenant isolation and priority
- Abuse controls tied to user signals
- Protects the consumer SLO
What L7 adds:
- Frames this as a revenue product with its own P&L, and sets the policy for how business messages interact with E2E (e.g., businesses using hosted endpoints)
Drill 7: Build vs Buy — The Connection Layer#
Prompt: "Should we build our own gateway or use a managed WebSocket service?"
Staff Answer
"At small scale (< ~1M concurrent), a managed WebSocket/pub-sub service is fine and cheaper in people. At hundreds of millions of connections, connection cost dominates: managed services often price per connection-minute and per message, which at 1B devices becomes the largest line item; and we need control over heartbeat intervals, binary framing, and reconnect behavior for mobile battery. WhatsApp's own history — Erlang handling ~2M connections per host — shows what a purpose-built edge can do. I'd buy early, keep the protocol ours, and build when connection cost crosses a threshold we can compute."
Why this is L6:
- Scale-dependent decision with a trigger
- Keeps protocol ownership to preserve the option
What L7 adds:
- Computes the crossover: managed cost/month vs a gateway team (~6–8 engineers) + hosts
Drill 8: Policy Change Without Outage — Rolling Out Multi-Device#
Prompt: "Move from phone-as-relay to per-device endpoints for 2B users."
Staff Answer
"This changes the protocol for every sender, so the rollout is capability-gated. Step 1: clients ship support for per-device encryption and signed device lists but don't use it. Step 2: devices advertise capability; senders encrypt per-device only when all recipient devices support it, else fall back to the old path. Step 3: enable per-device linking for a beta cohort; watch decryption-failure rate, delivery latency and battery metrics. Step 4: expand by percentage; old clients below a minimum version are asked to update. Rollback: disable new linking; existing linked devices keep working through a compatibility path. Key metric: client.decrypt_failures_total must stay at baseline."
Why this is L6:
- Capability negotiation instead of a flag day
- Safety metric specific to E2E (decrypt failures)
- Rollback path defined
What L7 adds:
- Coordinates security review, public documentation (whitepaper), and support readiness; sets the minimum-client-version policy
Drill 9: Cost#
Prompt: "What dominates the cost of running this, and how would you cut it 30%?"
Staff Answer
"Connections, not storage. 1B devices × heartbeats every ~45 s ≈ 22M tiny frames/s, plus TLS and memory per socket. Levers: longer heartbeat intervals where carrier NATs allow (adaptive per network), more connections per host (runtime efficiency), and fewer reconnects (graceful drains, session resumption). Second: media storage and egress — dedupe forwarded media by content hash of the ciphertext pointer, and expire blobs after all recipients download. Third: push — collapse wake-ups. Storage of the inbox is small by comparison."
Why this is L6:
- Identifies the non-obvious dominant cost
- Concrete levers per cost center
What L7 adds:
- Builds a cost-per-DAU model per component; negotiates CDN/egress contracts for media
Drill 10: Multi-Region#
Prompt: "Users in Brazil message users in India. Design the regions."
Staff Answer
"Each user (and their devices' inboxes) is homed in a region near them; each chat's sequencer is homed where its creator or majority lives. A message from Brazil to India: Brazilian gateway → router → the chat's sequencer (maybe cross-region, ~150–250 ms RTT) → append to the Indian recipient's inbox in India → ACK. To keep the sender's one tick fast, the router can append durably to a local cross-region outbound queue with quorum and ACK, then forward — the tick still means 'durable', just in a different store. Region failure: devices reconnect to another region; inboxes are replicated cross-region asynchronously for disaster recovery, with the understanding that a region loss can delay (not lose, if replicated synchronously in-region) messages."
Why this is L6:
- Homes users and chats; explains the cross-region path
- Keeps the one-tick semantics honest
- Distinguishes in-region sync vs cross-region async replication
What L7 adds:
- Data residency for metadata by country; decides region footprint by user distribution and regulatory exposure
8. Deep Dive Scenarios#
Deep Dive 1: Peak Traffic — New Year's Eve Delivery Delays#
Context: 00:01 in a large timezone. Online delivery p99 goes from 300 ms to 45 s. Social media says "WhatsApp is down." The on-call escalates to you.
Questions to Surface First:
- Is the delay at the router/sequencer, inbox append, or push to the recipient gateway?
- Are messages with media delayed more than text?
- Are sender ACKs (one tick) delayed, or only delivery (two ticks)?
- Is the session registry healthy?
Typical L5 Approach: Adds router instances and inbox nodes across the board.
Staff Approach: Splits the latency by stage:
send_to_ack_ms(sender side) vsack_to_deliver_ms. Finds ACKs fine but deliver slow: the session registry is overloaded by reconnects from users toggling airplane mode at midnight, so routers misclassify online devices as offline and fall back to push, which is throttled. Mitigations: registry read cache with 30 s TTL on routers; prioritize text over media on gateway egress; raise push collapse window.
Principal Approach: Treats New Year's as a known annual event with a "follow the midnight" capacity plan: registry and gateway capacity shifted region by region, media upload deprioritization pre-configured, and a load test at 5× last year's peak two weeks ahead. Owns the public status communication plan.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Stage-split latency; confirm no loss (ACK path healthy) |
| Triage | Registry latency, push fallback rate, gateway egress queues |
| Quick fix | Router-side registry cache; text priority over media |
| Guardrails | Watch push.throttled_total, registry p99 |
| Post-mortem | Registry capacity for reconnect bursts; annual event plan |
Metrics to Watch: delivery.send_to_ack_ms_p99, delivery.ack_to_deliver_ms_p99, registry.read_latency_p99, router.offline_fallback_ratio, push.throttled_total
Organizational Follow-up: Annual peak events calendar with per-region capacity plans.
Ownership Question: "Who decides to deprioritize media during a peak?" Staff answer: Messaging core on-call, pre-approved in the runbook; product agreed that text-first is the right degradation.
Key Takeaway: "Split latency by stage. A slow two-tick with a fast one-tick is a delivery-path problem, not a durability problem."
What clears the Staff bar:
- Stage-level latency decomposition
- Finds the registry → push fallback amplifier
- Pre-agreed degradation order
Deep Dive 2: The Silent Failure — Messages Stuck on One Tick#
Context: Client telemetry shows client.stuck_single_tick_total rising slowly for 2 weeks in one region: ~0.01% of messages never get a delivered receipt even though recipients are active.
Questions to Surface First:
- Did the recipient receive the message (check recipient telemetry for msg_id), or did only the receipt get lost?
- Is it correlated with multi-device users, a client version, or a storage cluster?
- Any failovers in that region's inbox store?
Typical L5 Approach: Assumes receipts are lossy and suggests resending receipts.
Staff Approach: Correlates by recipient device count: nearly all cases are recipients with a companion device linked in the last 30 days. Root cause: the device list cache in the router had a TTL of 1 h; new companions didn't receive envelopes until cache refresh, and the delivered receipt logic waited for all devices. Fix: device list changes invalidate router caches via an event; senders include the device-list version and the router rejects stale versions (like membership versions).
Principal Approach: Establishes "every cache of security-relevant or routing-relevant lists is versioned and invalidated by event, never TTL-only" as an engineering standard. Adds end-to-end synthetic probes (test accounts with multiple devices in each region sending every minute) so silent delivery failures are detected by the platform, not by telemetry trends.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Confirm whether messages or receipts are missing |
| Triage | Correlate by device count, version, cluster |
| Quick fix | Invalidate router device-list caches on change |
| Guardrails | Device-list versions in SEND; synthetic probes |
| Post-mortem | TTL-only caches on routing data prohibited |
Metrics to Watch: client.stuck_single_tick_total, router.stale_device_list_rejects, probe.e2e_delivery_success{region}
Organizational Follow-up: Synthetic probes owned by messaging core, alerting at 99.99% success.
Ownership Question: "Who owns end-to-end delivery correctness when it spans client, router and key directory?" Staff answer: Messaging core owns the end-to-end SLO and the probes; component teams own their pieces. Someone must own the whole path.
Key Takeaway: "A relay's silent failure is a message that never arrives and never errors. Only end-to-end probes catch it."
What clears the Staff bar:
- Separates lost message from lost receipt
- Finds correlation with a routing-cache TTL
- Adds end-to-end synthetic probes
Deep Dive 3: Large-Customer Onboarding — A Government Uses Groups for Emergency Alerts#
Context: A government agency wants to send emergency updates to 20M citizens via the platform.
Questions to Surface First:
- Is this conversational (replies expected) or broadcast?
- Latency requirement — minutes or seconds?
- Should E2E apply? Who are the "devices" on the sender side?
Typical L5 Approach: Creates many 1,024-member groups and fans out.
Staff Approach: Recognizes broadcast, not chat: use a channel/business-messaging path — the sender posts once to a channel log, subscribers fetch on read or receive batched wake-ups; rate-controlled delivery at a level the inbox and push tiers can absorb (20M × ~1.5 devices = 30M deliveries; at 500K/s, ~1 minute). No per-recipient receipts; aggregate reach metrics instead.
Principal Approach: Weighs the public-interest and political risk of the platform as an emergency channel: SLOs we can publicly commit to, abuse/impersonation of government accounts (verified sender program), and whether to partner or decline. Brings Policy and Legal in before engineering starts.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Scoping | Broadcast vs conversation |
| Design | Channel with fan-out on read and batched wake-ups |
| Capacity | 30M deliveries within target window; reserve push budget |
| Pilot | One region, test broadcasts |
| Launch | Verified sender, rate-controlled sends |
Metrics to Watch: channel.broadcast_reach_ratio, push.batch_latency_p99, consumer.delivery_latency_p99 (must not regress)
Organizational Follow-up: Verified-sender program owned by Integrity.
Ownership Question: "Who is accountable if an emergency alert is delayed?" Staff answer: We commit only to what we measure in pilots; the agency's contract states our delivery percentile, and Legal signs it.
Key Takeaway: "Don't bend a group feature into a broadcast system. Recognize the intent and use the right fan-out model."
What clears the Staff bar:
- Identifies intent mismatch
- Computes delivery time from throughput
- Protects consumer SLOs
Deep Dive 4: Post-Mortem — Removed Group Member Read Messages#
Context: A security researcher reports that a user removed from a group continued to receive decryptable messages for several minutes.
Questions to Surface First:
- Did the server keep routing envelopes to the removed member's devices, or did the removed member's client keep the old sender key and receive ciphertexts via another path?
- Did senders rotate their sender keys on membership change?
- Were there senders on old client versions?
Typical L5 Approach: Fixes the server to stop routing to removed members immediately.
Staff Approach: Finds two gaps: (1) routers had cached membership for up to 5 minutes; (2) some old clients didn't rotate sender keys on removal. Fix both: membership versions checked on every SEND (stale → reject); sender-key rotation mandatory on removal, enforced by rejecting sends with pre-removal key IDs after a grace period of seconds; old clients below a minimum version blocked from sending in groups.
Principal Approach: Treats it as a security incident with disclosure obligations: coordinates with the researcher (bug bounty), publishes an advisory as appropriate, and makes group-membership semantics part of a formal security model that every client change is reviewed against.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Reproduce; measure the exposure window |
| Triage | Server routing vs client key rotation |
| Quick fix | Membership version checks; forced rotation |
| Guardrails | Minimum client version for group sends |
| Post-mortem | Security model review for membership changes |
Metrics to Watch: router.stale_membership_rejects, client.sender_key_rotations_total, group.sends_by_old_clients
Organizational Follow-up: Security review gate for any change touching membership or keys.
Ownership Question: "Who owns the security semantics of group membership?" Staff answer: The messaging security team — with veto on client and server changes that touch keys or membership.
Key Takeaway: "Under E2E, access control is enforced by keys, not by routing. Both must change together."
What clears the Staff bar:
- Distinguishes routing from cryptographic access
- Uses versioning to eliminate cache windows
- Treats as security incident
Deep Dive 5: Multi-Region Expansion — Data Residency for Metadata#
Context: A country requires that messaging metadata for its citizens stays in-country.
Questions to Surface First:
- What counts as metadata — sender/recipient IDs, timestamps, IPs, group membership?
- Does content (already E2E) fall under the rule?
- How are cross-border chats treated?
Typical L5 Approach: Deploys a region in-country and routes those users there.
Staff Approach: Homes those users' devices, inboxes, session registry entries and group memberships in-country. Cross-border chats: the sequencer for a chat is homed based on policy (e.g., the in-country user's region if either participant is in-country); envelopes to foreign recipients leave the country only as delivery requires, with minimal retained metadata. Logs and analytics are aggregated in-country.
Principal Approach: Decides whether to comply, negotiate, or exit the market with Legal and Policy, based on the company's privacy commitments; builds residency as a reusable capability (region-in-a-box with policy-driven data placement) rather than a one-off.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Classify | Enumerate metadata fields and their stores |
| Design | Policy-driven homing for users, chats, logs |
| Migrate | Move affected users' inboxes during low traffic with dual-read |
| Verify | Audit data flows; residency tests in CI |
| Operate | Region-local on-call; incident process with regulators |
Metrics to Watch: residency.cross_border_metadata_bytes, migration.inbox_moves_remaining, delivery.latency_p99{region}
Organizational Follow-up: Data-classification catalog maintained by Privacy engineering.
Ownership Question: "Who decides whether we operate in a market with these rules?" Staff answer: Leadership with Legal and Policy; engineering provides the cost and the technical feasibility, not the decision.
Key Takeaway: "E2E protects content, not metadata. Residency is a metadata-placement problem."
What clears the Staff bar:
- Separates content from metadata
- Policy-driven placement
- Escalates the market decision
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Frame chat as relay vs system of record and commit to one
- Define one tick, two ticks and blue ticks as precise durability and acknowledgment states
- Explain at-least-once delivery with exactly-once display using client IDs and dedupe
- Design per-chat ordering with a fenced sequencer
- Size inbox storage from residency time and connection cost from heartbeats
- Explain group fan-out with sender keys, membership versions and the size cap
- Explain per-device multi-device under E2E with signed device lists
- At L7: take E2E through a cross-functional decision, price the edge fleet, and write the standard
The Bar for This Question#
Mid-level (L4): WebSocket servers, a messages table, online delivery, some offline storage. Timestamps for ordering.
Senior (L5): Adds a queue per user, Cassandra for messages, presence, push notifications, horizontal gateway scaling. Claims exactly-once, sorts by timestamp, and treats E2E and multi-device as footnotes.
Staff+ (L6): Commits to relay with E2E, builds per-device inboxes with quorum durability and cursor acks, uses a per-chat sequencer, treats receipts as messages, explains sender keys and group caps, and designs per-device multi-device with signed device lists. Every failure has a metric and owner. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "Exactly-Once Delivery Is Marketing"#
| Evidence | Implication |
|---|---|
| Mobile networks drop ACKs | Retries are unavoidable |
| Dedupe at the receiver is cheap and total | Exactly-once display is achievable |
The Staff position: At-least-once + idempotent display.
Why this matters in interviews: Precision about semantics is the Staff-level signal.
10.2 "The Server Should Know as Little as Possible"#
| Evidence | Implication |
|---|---|
| Undelivered-only storage is orders of magnitude smaller | Lower cost |
| Data not held can't be breached or compelled | Lower risk |
| WhatsApp's small early team ran a massive relay | Simplicity scales |
The Staff position: Relay with delete-on-ack for consumer E2E chat.
Why this matters in interviews: It reframes "scale" as "do less."
10.3 "Groups Above ~1K Are a Different Product"#
| Evidence | Implication |
|---|---|
| Fan-out on write × devices grows linearly per message | Cost explodes |
| Sender-key rekey on every membership change | Churn cost |
| Conversation norms break down | Product mismatch |
The Staff position: Cap groups; build channels for broadcast.
Why this matters in interviews: Knowing where a design stops working is Staff judgment.
10.4 "Connections Cost More Than Messages"#
| Evidence | Implication |
|---|---|
| ~1B sockets with heartbeats every 30–60 s | Tens of millions of frames/s with no messages |
| Reconnect storms are the largest outage class for chat edges | Operational cost concentrates at the edge |
The Staff position: Invest in the edge: graceful drains, jittered reconnects, adaptive heartbeats.
Why this matters in interviews: Most candidates optimize message storage, the smaller cost.
10.5 "E2E Is a Product and Policy Decision First"#
| Evidence | Implication |
|---|---|
| E2E removes server search, content moderation and easy history | Product surface changes |
| Governments and regulators engage publicly on E2E | Policy exposure |
The Staff position: Engineering can implement E2E; the org must choose it.
Why this matters in interviews: Naming the non-engineering owners is a Staff-to-Principal signal.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
A Staff engineer designs a messaging system. A Principal engineer sees that messaging is the substrate for half the company's products — consumer chat, business messaging, notifications, support, payments-in-chat, calling signaling — and that the one decision that shapes all of them is who can read the content. Once consumer chat is end-to-end encrypted, every adjacent product must decide whether it lives inside that boundary (and gives up server-side features) or outside it (and must be clearly distinguished to users). The L7 job is to make that boundary explicit, fund the shared edge and inbox platform, and prevent each product from building its own socket fleet and its own half-compatible delivery semantics.
The Org-Level Fault Line#
One messaging platform with a clear E2E boundary vs per-product messaging stacks.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Per-product stacks (chat, support, notifications, business each with its own sockets and queues) | Local speed | 3–4 connection fleets on the same phone (battery); inconsistent delivery semantics; duplicated edge on-call | Users' batteries; infra; on-call |
| One platform, everything E2E | Strongest privacy story | Business messaging, support and moderation lose server-side capabilities they need | Revenue products; T&S |
| Shared edge + inbox platform; explicit E2E boundary per product (L7 default) | One connection per device; one delivery contract; products declare whether they're inside E2E | Requires clear UX labeling and governance of the boundary | Platform team (~15–30 eng); Product (labeling) |
🧭 Principal Move: "Every product on the phone shares one connection and one delivery contract. What differs is whether the server can read the payload — and that is declared per product, reviewed by Security and Privacy, and visible to users. Nobody gets to be ambiguous about it."
Cost Model#
Assumptions: cloud list prices order-of-magnitude; ~0.5–1M connections per gateway host; heartbeats every ~45 s; ~6 device deliveries per message; relay storage (undelivered only); media stored encrypted with TTL; loaded engineer ~$250K/yr.
| Scale | Infra $/month | Dominant Cost | Headcount (eng) | On-Call Load |
|---|---|---|---|---|
| 10M DAU, ~5M concurrent, ~500M msgs/day | ~$50–150K | Gateways, media storage/egress | 10–20 | 1–2 rotations |
| 200M DAU, ~100M concurrent, ~10B msgs/day | ~$1–3M | Gateway fleet (~100–200 hosts), media egress, push | 80–150 across edge, core, media, clients, security | Per-tier rotations |
| 2B users, ~1B concurrent, ~100B msgs/day | ~$10–30M | Media egress and storage, edge fleet (~1–2K hosts), cross-region traffic | Hundreds, heavily on clients and security | Follow-the-sun; per-region edge |
The line that moves the business case: media. Text inbox storage is a rounding error beside media storage and egress; forwarding dedupe and blob TTLs are worth more than any inbox optimization.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse | Why |
|---|---|---|---|
| Launching E2E for consumer chat | One-way (trust and public commitment) | Reversal is a public-trust crisis | Users and regulators treat it as a promise |
| Relay (delete on ack) vs record | One-way-ish | Can't retroactively recover deleted history; adding history requires encrypted backup design | Product expectations form around it |
| Client message ID and cursor protocol | One-way-ish | Every client version in the wild speaks it | Mobile clients linger for years |
| Group size cap | Two-way upward, one-way downward | Lowering breaks existing groups | Raise cautiously |
| Heartbeat interval | Two-way | Config | Tune per network |
| Inbox store technology | Two-way with effort | Migration of hot, small data | Abstraction layer |
| Retention window (30 days) | Two-way downward with notice | Shortening loses some late deliveries | Measure first |
The Standard I'd Write#
RFC: Device Messaging Delivery & Confidentiality Standard (v1)
Scope: Every product that sends data to user devices over a persistent connection or OS push.
Mandatory requirements:
- Products MUST use the shared connection gateway; separate persistent connections from the same app are prohibited.
- Every client-originated write MUST carry a client-generated idempotency ID; recipients MUST dedupe on it.
- Delivery states exposed to users MUST map to defined durability events (e.g., "sent" = quorum-durable in all recipient inboxes).
- Each product MUST declare its confidentiality class: E2E (server cannot read) or server-readable; the class MUST be visible in the UI where users could confuse them.
- Routing-relevant lists (group membership, device lists) MUST be versioned; caches MUST be invalidated by event, not TTL alone.
- Undelivered payloads MUST NOT be retained beyond 30 days; delivered payloads MUST be deleted on ack for E2E products.
- Gateway deploys SHOULD drain ≤ 2% of connections per batch with jittered reconnect hints.
Exceptions: Filed with Messaging Platform, Security and Privacy; reviewed quarterly.
Success metrics: One persistent connection per device across products;
client.stuck_single_tick_total≤ 1 in 10⁵ messages; synthetic end-to-end probe success ≥ 99.99% per region; zero unlabeled server-readable surfaces inside E2E apps.
What I'd Tell the VP#
"Our messaging product's biggest promise is simple: when you see one tick, the message won't be lost, and nobody — including us — can read it. Keeping that promise is cheaper than it sounds because we only store messages until they're delivered. The risks are elsewhere: other teams building their own connections that drain batteries, and business features that blur whether we can read a message. I'm proposing one shared messaging platform with a clear, visible line between private and business messaging. It's mostly reorganizing existing teams, and it protects both our battery reputation and our privacy commitment."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Confidentiality boundary | "Each product declares whether it's inside E2E, and users can see it." |
| Portfolio edge | "One connection per device across all products." |
| Prices it | "Media egress is the cost line; inbox storage is a rounding error." |
| One-way doors | "Launching E2E is a public commitment we can't walk back." |
| Writes the standard | "Versioned routing lists, event invalidation, deletion on ack — mandatory." |
Staff answers that L7 interviewers find insufficient:
- "We'll add E2E with the Signal Protocol." — Correct, but ignores the cross-functional decision and the products it constrains.
- "Per-device inboxes solve multi-device." — Correct, but doesn't own the device-linking security policy or its UX.
- "We'll cap groups at 1K." — Correct, but doesn't decide what product serves the users who need more, or who owns it.
Appendices
Appendix A: Mechanics in Depth#
A.1 Send Path#
on SEND(dev, msg_id, chat_id, envelopes, membership_version, device_list_versions):
if dedupe.get(dev, msg_id) as prev: return prev
chat = sequencer.owner(chat_id) # fenced by epoch
if membership_version < chat.membership_version: return STALE_MEMBERSHIP
for each recipient user: if device_list_versions[u] < directory.version(u): return STALE_DEVICES
seq = chat.next_seq()
blob = store_once(chat_id, seq, ciphertext_or_per_device) # group: one blob with sender key
quorum_write([inbox_append(d, pointer(blob, seq, msg_id)) for d in recipient_devices(chat)])
dedupe.put(dev, msg_id, {seq, ts}, ttl=24h)
for d in recipient_devices: if registry.online(d): gateway(d).push(d) else push_sender.wake(d)
return ACK{msg_id, seq, server_ts}
A.2 Receive Path#
on DELIVER(inbox_seq, envelope):
if local_db.exists(envelope.msg_id): skip # exactly-once display
plaintext = decrypt(envelope) # session or sender key
local_db.insert(chat_id, seq, msg_id, plaintext)
periodically: send RECV_ACK{up_to_inbox_seq} # batch, e.g., every 100 msgs or 1 s
send RECEIPT delivered (batched per chat: up to seq)
A.3 Gap Handling#
A client seeing chat seq 41 then 43 waits ~2 s; if 42 doesn't arrive, it asks the server whether seq 42 was addressed to it. If not (e.g., a message to other devices only, or a sender-side retry gap), it records the gap as benign.
Appendix B: Keys and Data Model#
| Entity | Key | Store | Retention |
|---|---|---|---|
| Inbox entry | (device_id, inbox_seq) | Partitioned KV / wide-column (Cassandra-style), quorum writes | Until ack; ≤ 30 days |
| Message blob | (chat_id, seq) | Same store or blob store | Until all recipient devices ack; ≤ 30 days |
| Chat | chat_id → last_seq, epoch, membership_version | Strongly consistent KV | While chat exists |
| Dedupe | (device_id, msg_id) | Fast KV (Redis-style) | 24 h |
| Session registry | device_id → gateway | In-memory KV | TTL ~90 s, refreshed by heartbeat |
| Key directory | device_id → identity key, signed prekey, one-time prekeys | Strongly consistent KV | While device linked |
| Device list | user_id → signed device list + version | Strongly consistent KV | While account exists |
| Media | content-addressed encrypted blob | Object store + CDN | TTL after download |
Appendix C: Coordination Mechanisms — Quick Comparison#
| Mechanism | Purpose | Latency | Failure Behavior |
|---|---|---|---|
| Per-chat sequencer with epoch fencing | Order within chat | < 1 ms in-memory + durable batch | Failover reads last_seq, bumps epoch |
| Quorum inbox append | One-tick durability | +2–5 ms | Survives single replica loss |
| Dedupe window | Idempotent sends | < 1 ms | Loss → recipient dedupe catches duplicates |
| Membership / device-list versions | Security-correct routing | 0 (checked inline) | Stale → reject and client refresh |
| Session registry TTL | Online routing | < 1 ms | Stale → push fallback, no loss |
Appendix D: Protocol and Client Behavior#
- Heartbeat every ~30–60 s, adaptive per carrier NAT timeout; server closes idle sockets after ~2 missed heartbeats.
- Reconnect: jittered exponential backoff (base 1 s, cap 60 s); server may send
reconnect_inhints during drains. - On reconnect:
SYNC{after_inbox_seq}in pages of 100; ack in batches. - Outbox: unsent messages persisted locally with msg_id; retried until ACK; UI shows clock icon.
- Prekeys: top up to ~100 when below ~20.
- Media: encrypt locally with random key; upload; send message with URL, key and hash.
Appendix E: Observability#
E.1 Core Metrics#
| Metric | Why |
|---|---|
delivery.send_to_ack_ms_p99, delivery.ack_to_deliver_ms_p99 | Stage-split latency |
client.stuck_single_tick_total | Silent loss proxy |
probe.e2e_delivery_success{region} | Synthetic end-to-end correctness |
gateway.connections, gateway.reconnects_per_s | Edge health |
inbox.backlog_age_p99 | Offline tail |
router.stale_membership_rejects, router.stale_device_list_rejects | Routing correctness |
push.send_latency_p99{provider} | Wake-up dependency |
client.decrypt_failures_total | E2E health |
E.2 Critical Alerts#
| Alert | Threshold | Severity |
|---|---|---|
| E2E probe failure | < 99.99% over 10 min in a region | Page messaging core |
| Online delivery p99 | > 2 s for 5 min | Page messaging core |
| Reconnect rate | > 5× baseline for 2 min | Page edge |
| Decrypt failures | > 2× baseline for 15 min | Page security + clients |
| Inbox append errors | > 0.01% | Page storage |
E.3 Control Plane vs Data Plane#
Data plane: sockets, sends, inbox appends, deliveries, acks. Control plane: gateway deploys and drains, heartbeat policy, group caps, rate limits, retention settings, minimum client versions. Control-plane changes are canaried by region and by client cohort.
Appendix F: Scale Evolution#
| Scale | What Works | What You Add |
|---|---|---|
| < 100K concurrent | One region; Postgres inbox table; long-polling or WebSockets | Client IDs, per-chat seq |
| 100K–10M | Gateway fleet; partitioned inbox store; Redis registry; push | Quorum appends, graceful drains |
| 10M–500M | E2E, sender keys, per-device multi-device; multi-region homing | Synthetic probes, adaptive heartbeats |
| 500M+ | Shared messaging platform, per-product confidentiality classes | Channels product; residency capabilities |
F.1 What You Don't Build on Day One#
- Multi-region active routing
- Per-device multi-device (start with one device per account, or server-readable companion sync if not E2E yet)
- Channels/broadcast
- Custom gateway runtime (use a proven WebSocket stack)
What you do build on day one: client message IDs, per-chat sequence numbers, cursor-based sync and delete-on-ack semantics. Those are the protocol commitments that every client in the wild will carry for years.
Appendix G: Fairness, Abuse and Cost#
- Abuse without content: send-rate limits per account and per new account, fan-out limits on forwarding (e.g., forward to at most a handful of chats at once), user reports that include reported messages from the reporter's device.
- Business tenants: per-tenant quotas, templates, quality scores; consumer traffic prioritized on shared tiers.
- Battery fairness: one connection per device across products; collapse pushes.
- Cost attribution: per-product share of connections, deliveries, media bytes; monthly showback.