Technologies referenced in this case study: Redis · PostgreSQL · Kafka · ZooKeeper & etcd · DynamoDB
Related case studies: Real-time WebSockets · File Sync · Chat & Messaging · Data Consistency · Blob Storage
Related patterns: Real-time Updates · Contention · Distributed Coordination · Consistency Models
How to Use This Case Study#
Organized for interview use first, reference second. Read front-to-back once, then return to the sections that match your weak spots.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Sections 3–4 → the Deep Dives that scare you |
| Deep Dive | 3+ hrs | Everything, including the Principal Lens (Section 11) and Appendices |
What is Collaborative Editing? — Why interviewers pick this topic
Collaborative editing lets multiple people modify the same document at the same time and see each other's changes within a few hundred milliseconds — Google Docs, Figma, Notion, Microsoft Loop, VS Code Live Share. The hard part is not the text box. It is that two people change the same bytes concurrently, over networks with 50–500ms of latency, sometimes offline for hours, and everybody must converge on the same document without anyone losing work.
Before vs After — the "lost paragraph" incident:
Without a convergence protocol (naive last-write-wins on the whole doc):
t=0: Alice and Bob open the Q3 planning doc (v41)
t=+5s: Alice rewrites the intro paragraph
t=+6s: Bob fixes a typo in section 4
t=+7s: Alice's client PUTs full doc v42
t=+7.2s: Bob's client PUTs full doc v42 (based on v41) — overwrites Alice
t=+3min: Alice notices her intro is gone. No error was ever shown.
t=+1day: Support ticket: "Docs ate my work." Trust is gone.
With an operation-based protocol (OT or CRDT) + single sequencer per doc:
t=0: Both open v41; both subscribe to the doc session
t=+5s: Alice's ops stream: insert(pos 120, "...") at rev 41
t=+6s: Bob's op: delete(pos 2040, 1), insert(pos 2040, "e") at rev 41
t=+6.1s: Server orders Alice's op as rev 42, transforms Bob's op against it → rev 43
t=+6.2s: Both clients converge on rev 43. Both edits survive. Nobody noticed.
Why interviewers reach for this question: It forces you to reason about concurrency semantics (what does "both edits survive" mean?), stateful real-time infrastructure (sticky sessions, per-document ownership), durability (an op acknowledged is an op that must never be lost), and product intent (a legal contract editor and a whiteboard have opposite correctness bars). Every one of those is a Staff-level judgment call, and none of them is "pick OT or CRDT."
Mechanics Refresher: Convergence Strategies
| Strategy | How It Works | Pros | Cons |
|---|---|---|---|
| Lock / check-out | One editor holds a lock per doc or section | Trivially correct | Not collaborative; lock leaks; "who has it?" tickets |
| Last-writer-wins (whole doc) | Latest full save wins | Simple | Silently destroys concurrent work |
| Three-way merge (diff3) | Merge on save against common ancestor | Works for code / async | Conflicts surface to users; no real-time |
| Operational Transformation (OT) | Ops are transformed against concurrent ops before applying; server assigns total order | Compact ops; intent-preserving for text; proven at Google Docs scale | Transform functions are hard to get right; needs a central sequencer in practice |
| CRDT (sequence CRDTs: RGA, YATA/Yjs, Fugue) | Every character has a unique, ordered ID; merges are commutative by construction | Peer-to-peer and offline-friendly; no central transform | Metadata overhead; tombstones; harder to enforce server-side invariants |
| Server-authoritative LWW per property (Figma-style) | Per-object, per-property last-writer-wins, ordered by server | Simple, fast for structured docs | Wrong for rich text runs; needs a separate text strategy |
For most production systems: a central, per-document sequencer with an op log — whether the op format is OT or a CRDT. The sequencer gives you total order, a durable ack point, permission enforcement, and a place to snapshot. The OT-vs-CRDT argument is about the op algebra; the sequencer is about operations, and that is where the interview is won.
Executive Summary
If you only read one section, read this. Everything else in the case study elaborates the contrast below.
What This Interview Actually Tests#
Collaborative editing is not an OT-vs-CRDT question. Candidates who spend 15 minutes on transform functions have told the interviewer they studied a paper, not that they can run a system.
It is a stateful ownership question that tests:
- Whether you pin down the document model and intent before choosing an algorithm
- Whether you can own a stateful, sticky, per-document hot path without making it a single point of failure
- Whether you define exactly when an edit is "safe" (the ack point) and who is paged when that promise breaks
- Whether you treat offline, permissions, and history as first-class — not bolt-ons
The key insight: Every collaborative editor is a replicated log with a UI on top. Once you say "one sequencer per document, ops appended to a durable log, snapshots for fast load, clients as optimistic replicas," the remaining 35 minutes are about what happens when that sequencer moves, lags, or lies.
The L5 vs L6 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "We'll use OT like Google Docs" | Asks what the document is (rich text? structured canvas? code?) and whether offline is a requirement — these pick the algorithm | Asks how many products in the org need real-time collaboration and whether this is a shared sync platform or a one-off |
| Consistency | "Eventually consistent via OT" | Defines the ack point: an op is acknowledged only after it is durably appended to the doc's log; convergence is guaranteed, intent preservation is best-effort | Sets an org-wide durability SLO ("0 acknowledged edits lost per quarter") and funds the audit that proves it |
| Scale | "Shard by doc ID" | Names the hot doc: a 1,000-viewer all-hands doc is a fan-out problem, a 50-editor doc is a sequencing problem — different fixes | Prices the long tail: 99% of docs are cold; storage tiering of op logs saves more money than any hot-path optimization |
| Failure | "Add replicas for the WebSocket servers" | Designs session handoff: doc owner lease in etcd, fencing tokens, clients reconnect with last-acked rev and resend unacked ops | Designs the org's failure posture: cell-based doc placement so one bad deploy touches ≤5% of docs; game days for session-server loss |
| Ownership | "The docs team owns it" | Splits ownership: collaboration infra (sessions, log, snapshots) vs editor product (schema, UX) vs storage platform | Decides whether sync becomes a platform with a contract (op log + presence + permissions) that Docs, Sheets, Slides and Comments all consume |
Why "first move" separates levels
L5: Reaches for OT because Google Docs uses it. This is a correct answer to the wrong question — it skips whether the document is linear text, a tree (rich text with tables and lists), or a graph of objects (a design canvas). Those choices change what a "concurrent edit" even means.
L6: "Before I pick an algorithm, I need to know three things: what's the document model, is offline editing a hard requirement, and do we need server-side validation of every edit — for example, permissions on individual sections or schema invariants? If it's rich text, online-first, with server-enforced permissions, I'll use a central sequencer with OT-style ops. If offline-first is the core product promise, I'd pick a CRDT and accept the metadata overhead."
L7: Recognizes that the company likely has three teams each about to build their own sync engine, and that the one-way door is the op log format and the client SDK contract, not the algorithm.
Why "consistency" separates levels
L5: "OT guarantees convergence." True, and irrelevant to the user whose edit vanished because a session server crashed before persisting it.
L6: Separates three properties and assigns each an owner: convergence (all replicas reach the same state — guaranteed by the algorithm), durability (an acked op survives any single failure — guaranteed by the append-before-ack rule), and intent preservation (the merged result is what both users meant — best-effort, and product signs off on the edge cases like concurrent delete-and-format of the same word).
L7: Turns durability into a measured SLO with an independent auditor: a background job replays op logs against snapshots and pages if any acked revision is missing.
Why "failure" separates levels
L5: Treats the collaboration server as stateless and adds replicas behind a load balancer. But a doc session is stateful: two servers both believing they own doc X will both assign rev 1,042, and clients will diverge permanently.
L6: Makes single ownership explicit: a lease in etcd/ZooKeeper per doc with a fencing token; the op log rejects appends with a stale token. Failover is a lease expiry (5–10s) plus client reconnect with last_acked_rev. "During failover, users keep typing locally; their ops queue client-side and flush on reconnect. The user sees a 'reconnecting' pill, never lost text."
The Staff Positions#
| Position | Rationale |
|---|---|
| One sequencer per document, not per shard of text | Total order per doc is cheap (a doc rarely exceeds ~200 ops/sec) and makes permissions, history, and snapshots trivial |
| Ack only after durable append | The only promise users care about is "my edit is saved." Optimistic local apply gives the speed; durable ack gives the truth |
| OT-style ops with server authority for online-first rich text; CRDT for offline-first or P2P | Algorithm follows intent. Don't pay CRDT metadata tax if you already need a server |
| Presence is ephemeral and lossy — never on the durable path | Cursors at 10–20 updates/sec per user would be 10× the op volume; drop them freely under load |
| Snapshot every ~500 ops or 60s of activity | Bounds cold-load replay to <100ms; op log remains the source of truth for history |
| Permissions checked at session join and on every op, cached with a revocation push | Revoking access must take effect in seconds, not at next page load |
| Offline is a product decision with an explicit horizon | "Offline edits merge cleanly for up to 30 days; beyond that we fork" — say the number |
The Three Intents#
Three intents drive every design decision. Each leads to a different architecture.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Real-time co-authoring (Docs, Sheets) | Online-first, 2–100 concurrent editors, sub-200ms remote visibility | Central per-doc sequencer, OT-style ops, op log + snapshots | Sequencer failover causes brief freeze; divergence if ownership is split | Zero acked-op loss; convergence always; intent best-effort |
| Offline-first / local-first (notes, field apps) | Edits for hours/days disconnected; merge on reconnect | CRDT (Yjs/Automerge-class), server as relay + archival peer | Metadata bloat; surprising merges after long divergence | Convergence always; merges must never lose text, may interleave oddly |
| Structured canvas (Figma-style design, whiteboards) | Objects with properties, not character sequences; 60fps interactions | Server-authoritative per-property LWW, fractional indexing for ordering | Last-writer clobbers a concurrent property change | Per-property atomicity; conflicts rare and acceptable |
🎯 Staff Move: "I'll design for real-time co-authoring of rich text — Google Docs — because that's where concurrency semantics and durability both bite. I'll keep offline as a bounded feature, not the core promise, which lets me keep a central sequencer. If the product were offline-first, I'd switch to a CRDT and I'll call out where the design changes."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | OT vs CRDT | Central transform with compact ops vs decentralized merge with metadata overhead — who pays the complexity? |
| 2 | Stateful Session Ownership vs Stateless Scale-out | One owner per doc gives order; any-server-serves-any-doc gives simple ops — you cannot have both |
| 3 | Latency vs Durability of Acks | Ack on receipt (fast, lossy on crash) vs ack after replicated append (safe, +5–20ms) |
| 4 | Offline Freedom vs Merge Sanity | Longer offline windows mean bigger, weirder merges and permission drift |
| 5 | Storage Cost vs History Fidelity | Keep every keystroke forever (version history, audit) vs compact aggressively (cheap, lossy history) |
In the Wild: Real Production Systems#
Why this section belongs here: Naming real systems — and what they chose — shows you've studied operational reality. Use these as anchors, not trivia.
Google Docs — Server-Centric Operational Transformation#
Google Docs descends from the Jupiter collaboration model (Xerox PARC, 1995) that Google Wave also used: clients send operations tagged with the server revision they were based on, a server transforms them against anything that landed since, assigns the next revision, and broadcasts. Clients only ever transform against one server stream, which collapses the notoriously hard N-way OT problem into a 2-party problem. Google publicly documents that up to 100 people can have a doc open for editing at once, beyond which additional users get view-only access — a product-level concurrency cap.
Staff insight: The 100-editor cap is not a technical embarrassment; it is a Staff decision made visible. Bounding per-doc concurrency bounds sequencer load, transform cost, and presence fan-out. Name the cap in your design.
Figma — Server-Authoritative, CRDT-Inspired Multiplayer#
Figma's engineering blog describes a multiplayer system where each open file is loaded into a dedicated server process that is the authority for that document. Clients apply changes optimistically; the server resolves conflicts with last-writer-wins at the granularity of individual object properties, and uses fractional indexing to order children without reindexing siblings. They explicitly chose not to use a pure CRDT because a central server made the system simpler and let them avoid CRDT overhead.
Staff insight: "Inspired by CRDTs, but with a server" is the most defensible production answer for structured documents. It shows you know the theory and chose operational simplicity on purpose.
Yjs / Automerge — Local-First CRDTs#
Yjs (YATA algorithm) and Automerge are open-source CRDT libraries used by a large number of editors (for example, many ProseMirror/TipTap-based products integrate Yjs). They give convergence without a central transform: any peer can merge any other peer's updates in any order. The price is per-character identity metadata and tombstones, which both projects have invested heavily in compressing.
Staff insight: CRDTs move complexity from the server into the data format. That is the right trade when offline or peer-to-peer is the product; it is the wrong trade when you already need a server to enforce permissions and validate edits.
What Interviewers Probe#
| After You Say... | They Will Ask... | What They're Evaluating |
|---|---|---|
| "We'll use OT" | "Walk me through transforming two concurrent inserts at the same position." | Do you know the tie-break (site ID / server order), or just the acronym? |
| "We'll use a CRDT" | "What happens to document size after a year of edits? How do you enforce permissions?" | Awareness of tombstones, metadata overhead, and the loss of a central enforcement point |
| "Sticky sessions per doc" | "The server holding the doc dies. What does Alice see? Is anything lost?" | Ack point, fencing, client resend, reconnect semantics |
| "We snapshot periodically" | "How long does a 5-year-old doc take to open?" | Snapshot cadence, replay bounds, cold storage tiering |
| "Users can edit offline" | "Alice is offline 3 weeks; her access was revoked yesterday. She reconnects." | Permission re-check at merge, fork semantics, product sign-off |
| "Presence via WebSocket" | "The all-hands doc has 2,000 viewers. What's the fan-out?" | Separating viewers from editors, presence sampling, broadcast tiers |
System Architecture Overview#
Reading the diagram: The client applies its own edits instantly and keeps them in a pending queue. The session router finds the single owner for a document (lease in etcd). The owner transforms each incoming op against ops the client hasn't seen, assigns the next revision, appends it to the replicated op log with its fencing token, and only then acks and broadcasts. Presence travels a separate lossy channel. Snapshots are built off the hot path by a compactor so opening a doc replays at most a few hundred ops.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Algorithm | "OT, like Google Docs" | "Algorithm follows intent. Online rich text with server permissions: central sequencer with OT-style ops. Offline-first: CRDT." |
| Ordering | "Timestamps" | "Server-assigned revision per doc from a single fenced owner. Wall clocks never order edits." |
| Durability | "Save to DB periodically" | "Ack after replicated append to the op log. Snapshots are an optimization, never the source of truth." |
| Scaling | "Shard WebSocket servers" | "Shard docs across session servers by consistent hashing with leases. Scale viewers separately from editors." |
| Failure | "Replicas" | "Lease expiry 10s, fencing token on append, client resends unacked ops idempotently by (client_id, seq)." |
| Offline | "Queue and sync later" | "Bounded offline window, rebase on reconnect, permission re-check at merge, fork if the gap is too large." |
| Presence | "Send cursor on every keystroke" | "Throttle to 5–10 Hz, coalesce, drop under load, never persist." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Google Docs concurrent editor cap | 100 editors (more become viewers) | Public, product-level bound on per-doc concurrency |
| Google Docs max document size | ~1.02M characters | Bounds snapshot size and cold-load cost |
| Active typist op rate | 3–8 ops/sec (keystrokes, batched to ~2–5 msgs/sec) | 50 editors × 5 = 250 ops/sec is the worst-case sequencer load for one doc |
| Remote edit visibility target | <200ms p50, <500ms p99 same-region | Beyond ~500ms, users start typing over each other |
| Durable append latency (replicated, same region) | 3–15ms | The cost of "ack after append" — worth it |
| Presence update rate | 5–10 Hz per user, coalesced | 10× op volume if unthrottled |
| Snapshot cadence | every ~500 ops or 60s idle | Keeps cold-load replay <100ms |
| Ownership lease TTL | 10s, renew every 3s | Failover time floor; shorter causes flapping |
| WebSocket connections per session server | 50K–100K idle, ~10K active docs | Memory ~30–60KB per connection with buffers |
| Naive CRDT metadata overhead | 2–10× raw text size | Compressed encodings (Yjs-style) bring it down substantially; still non-zero |
| Share of docs active in last 30 days | typically <10% | Justifies aggressive cold-tiering of op logs |
Interview Walkthrough
The most common mistake: Candidates spend 20 minutes deriving OT transform functions on the whiteboard and never reach durability, failover, or permissions. The phases below get the skeleton on the board in ~10 minutes and spend the rest on the decisions that set your level.
Phase 1: Requirements & Framing (2–3 minutes)#
State functional requirements in one breath:
"Multiple users open the same document, see each other's edits and cursors in near real time, and nobody's acknowledged edit is ever lost. Plus version history, sharing permissions, and some offline tolerance."
Then spend your time on the three questions that pick the architecture:
"Three things change the design. First, the document model — is it rich text, a spreadsheet grid, or a canvas of objects? Second, is offline a core promise or a degraded mode? Third, do we need the server to validate every edit — permissions, schema, compliance holds? I'll assume rich text like Google Docs, online-first with offline as a bounded feature, and server-enforced permissions."
Commit to non-functional targets out loud:
"Remote edits visible in under 200ms p50 within a region. Local edits apply in under 16ms — one frame — regardless of network. Zero loss of acknowledged edits. Up to ~100 concurrent editors per doc, unbounded viewers via a separate path. Opening a doc under 1 second p95."
🎯 Staff Move: The phrase "zero loss of acknowledged edits" reframes the whole interview. It forces you (and the interviewer) to define the ack point, which is where durability, failover, and client retry design all live.
Phase 2: Core Entities & API (1–2 minutes)#
Name the nouns in 30 seconds:
- Document:
doc_id, owner, ACL, currentrev, latestsnapshot_rev, cell/shard assignment - Operation:
(doc_id, base_rev, client_id, client_seq, ops[])— ops areretain(n) / insert(text, attrs) / delete(n)over a linearized document - Revision: server-assigned monotonically increasing integer per doc;
(doc_id, rev)is the primary key of the op log - Snapshot: full document state at
rev, stored as a blob - Session: the set of clients currently connected to a doc, owned by one session server
Real-time channel (WebSocket, per doc):
→ JOIN { doc_id, last_known_rev, auth }
← SYNC { snapshot_rev, snapshot?, ops[last_known_rev+1 .. head] }
→ SUBMIT { base_rev, client_id, client_seq, ops[] }
← ACK { client_seq, rev }
← REMOTE_OP { rev, author, ops[] }
↔ PRESENCE { cursor, selection, color } (lossy, throttled)
Cold path (REST):
GET /docs/{id}/revisions?from=&to= version history
POST /docs/{id}/restore { rev } restore = new ops, never a rewrite
PUT /docs/{id}/acl { principal, role }
🎯 Staff Move: "
(client_id, client_seq)is the idempotency key. When a client reconnects and resends ops that were actually applied before the crash, the server dedupes them instead of inserting the paragraph twice. Restore is modeled as new forward ops — history is append-only, so audit never breaks."
Phase 3: High-Level Architecture (≤5 minutes)#
Draw ≤8 boxes:
Walk the edit flow in 60 seconds:
- Client applies the edit locally (instant), puts it in a pending queue, sends
SUBMITwithbase_rev. - Gateway routes to the doc's owner — looked up from a lease registry.
- Owner transforms the op against every op in
(base_rev, head], checks permission, assignsrev = head+1. - Owner appends to the op log with its fencing token. Only after the append succeeds, it acks the author and broadcasts to other clients.
- Other clients transform the remote op against their own pending ops and apply.
- A compactor periodically folds ops into a snapshot for fast loads.
🎯 Staff Move: After drawing, say: "This works for a single region with docs up to 100 editors. The interesting parts are what happens when the owner dies mid-append, what happens when a client comes back after a week offline, and how 2,000 viewers on one doc don't melt the owner. Which should I go deep on?" You're at ~9 minutes.
Phase 4: Transition to Depth (1 minute)#
"The skeleton is standard. Three places decide whether this system is trustworthy: session ownership and failover — because split-brain means permanent divergence; the ack point and durability — because users forgive latency but not lost work; and offline plus permissions — because that's where product and security disagree. I'd start with ownership and failover unless you'd prefer otherwise."
If the interviewer asks "OT or CRDT?" first, answer in 90 seconds and bridge back: "Given online-first with server-side permissions, OT-style with a central sequencer. The algorithm matters less than the sequencer's failure semantics — let me show you why."
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → pick a position → quantify → name who pays.
Deep dive A: Session ownership and failover (7–8 min)
"Each doc has exactly one owner at a time. Ownership is a lease in etcd with a 10s TTL, renewed every 3s. Every lease grant carries a monotonically increasing fencing token. The op log accepts an append only if the token is ≥ the highest it has seen for that doc — so a zombie owner that was paused by GC for 15 seconds cannot write rev 1,042 after the new owner already did."
Failure sequence:
- Owner crashes at t=0. Clients' WebSockets drop; they show "Reconnecting…" and keep editing locally.
- Lease expires at t≤10s. Router assigns a new owner, which loads the latest snapshot plus tail ops from the log (~50–200ms).
- Clients reconnect with
last_acked_revand resend every unacked op with its(client_id, client_seq). - New owner dedupes ops already in the log, transforms the rest, appends, acks.
- User-visible impact: a 5–12s "reconnecting" pill; no lost text.
"The victim of a 10s lease is freshness during failover — remote edits freeze for up to 10s. The victim of a 2s lease is stability — GC pauses and network blips cause ownership flapping and needless reloads. I pick 10s and say so."
Deep dive B: The ack point and durability (6–7 min)
"Three places I could ack: on receipt in memory (~1ms, loses edits on crash), after local disk fsync (~2–5ms, loses edits on machine loss), or after a replicated append — quorum of 3 across zones (~5–15ms). I ack after replicated append. 10ms is invisible because the local edit already rendered; the ack only clears the pending queue."
What the op log is: "Could be a partitioned log like Kafka keyed by doc_id, or a table in a strongly consistent KV — DynamoDB or Spanner-class — with (doc_id, rev) as the key and a conditional write attribute_not_exists(rev) AND token >= last_token. I prefer the KV: I need per-doc random reads for history and a conditional write for fencing, which a log doesn't give me cleanly."
Deep dive C: Offline and permissions (6–7 min)
"Offline edits are a queue of ops based on an old revision. On reconnect, the server transforms them against everything since — which for a 3-week-old base could be 50,000 ops. That is expensive and produces weird merges. So: offline edits merge automatically up to a bound — say 10,000 ops of divergence or 30 days. Beyond that, we create a 'your offline copy' fork and show a compare view. Before any merge, we re-check permission at the current ACL, not the ACL at edit time. If access was revoked, the edits are held and the user is told — never silently applied."
🎯 Staff Move: Every deep dive ends with an owner: "Collab infra owns lease tuning and the durability SLO. The editor team owns transform correctness and has a fuzz-test suite that runs 1M random concurrent op pairs per build. Product owns the offline horizon."
Phase 6: Wrap-Up (2–3 minutes)#
"To summarize: a single fenced owner per doc, ops acked after replicated append, snapshots for fast load, presence on a lossy side channel, and bounded offline with permission re-check at merge. What I'd build next: first, a divergence detector — clients periodically send a hash of their doc at a rev, and mismatches page us. Second, cell-based placement so a bad deploy touches a bounded fraction of docs. Third, cold-tiering op logs older than 90 days to object storage — most of our storage bill is history nobody opens."
Common Timing Mistakes#
| Mistake | Time Lost | Fix |
|---|---|---|
| Deriving OT transform matrices on the board | 10–15 min | State the tie-break rule in one sentence; offer the full function if asked |
| Designing the rich-text schema (bold, lists, tables) | 5–8 min | "Linearize to a sequence with attributes; tables are nested sequences" and move on |
| Debating WebSocket vs SSE vs long-poll | 3–5 min | "WebSocket, with long-poll fallback for hostile proxies." Done |
| Ignoring viewers vs editors | discovered at minute 40 | Say it in Phase 3: "viewers get a read-only fan-out tier" |
| Never stating the ack point | fatal | Say "ack after durable append" in Phase 1 or 2 |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Most system design problems let you hide state behind a database and scale stateless servers. Collaborative editing doesn't. The hot path is stateful by necessity — someone must hold the current document and order concurrent edits — and that state lives in memory on one machine at a time. The candidate has to design for a component that cannot simply be load-balanced, and has to reason about what "correct" means when two humans disagree at the same character.
It is also a product-judgment problem in disguise. "Alice deletes a sentence while Bob bolds a word inside it" has no algorithmically correct answer. The Staff candidate says who decides (product), what the default is (delete wins; Bob's formatting is dropped), and how it's tested.
1.2 The L5 vs L6 Contrast — Visual#
The L5 path is technically correct at every step. It fails because it spends its budget on the part of the problem that libraries already solve.
1.3 The Staff Question That Cuts Through Everything#
"When exactly is an edit safe, and who finds out if that promise is broken?"
This single question forces every important decision: the ack point (durable append), the ownership model (fenced single writer), the client retry protocol (idempotent resends), the detection mechanism (divergence hashes, replay audits), and the owner (collab infra on-call). A candidate who asks it in the first five minutes is already interviewing at Staff.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Real-time co-authoring (default). Think Google Docs, Sheets, Slides, Microsoft Word online. Users are almost always online; concurrent editing is the headline feature; the company must enforce sharing permissions, retention policies, and legal holds server-side. The server is already mandatory, so a central sequencer is free — and it buys total order, a single enforcement point, and linear history. OT-style ops fit naturally. Victims of this choice: users on flaky networks (they see "reconnecting" more often), and the collab infra team (they own a stateful fleet).
Offline-first / local-first. Think field-service apps, note-taking apps that must work on a plane, or peer-to-peer editors. Here the network is optional, so correctness cannot depend on a server. CRDTs are the right tool: any replica merges any other in any order. Victims: storage (metadata and tombstones), the security team (enforcing per-edit permissions on data that merged peer-to-peer is hard), and users who get "valid but weird" merges after long divergence.
Structured canvas. Think Figma, Miro, a diagramming tool. The document is a tree of objects with properties; a conflict means two people set fill on the same rectangle. Per-property last-writer-wins, ordered by the server, is simple and good enough — users rarely edit the same property of the same object in the same second, and when they do, "last one wins" matches their mental model. Text boxes inside the canvas still need a sequence strategy. Victims: the occasional user whose property change is overwritten, which product accepts.
2.2 When NOT to Use Real-Time Collaborative Editing#
| Situation | Better Choice | Why |
|---|---|---|
| Documents edited by one person 99% of the time | Autosave + optimistic concurrency (If-Match: etag) + conflict dialog | Paying for sessions, presence, and op logs for a solo editor is waste |
| Source code with review workflows | Git-style branches and 3-way merge | Humans want to review merges; real-time interleaving of code is harmful |
| Regulated records (contracts at signature, medical charts) | Check-out/check-in with explicit locking and audit | Legal wants a single accountable author per version, not a merge |
| Large binary assets (video, CAD meshes) | Locking + blob storage with versioning | No meaningful op algebra for binary diffs |
| Form-like records (CRM fields) | Per-field optimistic concurrency | Row-level conflicts are rare; field-level LWW with audit suffices |
🎯 Staff Move: "If our telemetry says 95% of docs never have two simultaneous editors, the collaborative engine should be lazy: a doc only gets a session server when a second editor joins. Solo edits go through a cheap autosave path." This is how you cut session fleet cost by an order of magnitude.
2.3 What the Interviewer Leaves Underspecified#
| Unstated Assumption | Why It Matters | What to Say |
|---|---|---|
| Document model | Text vs tree vs object graph changes the op algebra | "I'll linearize rich text into a sequence with attributes" |
| Max concurrent editors | Drives sequencer and presence design | "Cap at 100 editors, viewers unbounded via a separate tier" |
| Offline duration | Drives merge cost and fork policy | "Up to 30 days or 10K ops divergence, then fork" |
| Permission granularity | Doc-level vs section-level changes the op validator | "Doc-level roles: owner, editor, commenter, viewer" |
| History retention | Storage cost grows with every keystroke | "Full op history 90 days hot, then snapshots-only at named versions" |
| Multi-region | Where does the sequencer live for a doc with editors on 3 continents? | "Doc homed in one region; move the home if the editor population shifts" |
| Comments/suggestions | Anchoring to text that moves | "Comments anchor to op-transformed positions, same machinery as cursors" |
2.4 Precise Terminology#
| Term | Meaning | Common Confusion |
|---|---|---|
| Convergence | All replicas that have seen the same set of ops reach identical state | Not the same as "correct" — two replicas can converge on garbage |
| Intent preservation | Merged result reflects what each user meant | Best-effort; no algorithm guarantees it for all cases |
| Causality preservation | If op B was generated after seeing op A, every replica applies A before B | Violating it causes "edit appears before its cause" |
| Transform (OT) | T(a, b) → a' such that applying b then a' equals applying a then T(b, a) (TP1) | TP2 (needed for P2P OT) is where most published algorithms were later shown to be wrong |
| Sequence CRDT | Each element has a globally unique, totally ordered ID; inserts reference neighbors by ID | "CRDT" alone includes counters and sets; text needs a sequence CRDT |
| Tombstone | Deleted element kept as a marker so concurrent references still resolve | The reason CRDT docs grow even when text shrinks |
| Base revision | The server revision the client's op was generated against | Transform window = head − base_rev |
| Fencing token | Monotonic number attached to a lease; storage rejects writes with lower tokens | Leases without fencing do not prevent split-brain |
| Snapshot | Materialized doc state at a revision | Cache of the log, not the source of truth |
| Presence | Ephemeral cursor, selection, "who's here" state | Never durable, never ordered with ops |
3. The Five Fault Lines#
Each fault line: the options, who pays, the Staff default, and when to deviate.
3.1 Fault Line 1: OT vs CRDT#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Server-centric OT (Jupiter-style) | Compact ops (~20–50 bytes); linear history; server validates every op; only 2-party transforms | Requires a live sequencer; offline merges are expensive transforms | Collab infra (stateful fleet); offline users (long rebases) |
| Peer-to-peer OT | No server | TP2 correctness is notoriously hard; many published algorithms had counterexamples | Everyone, eventually — via divergence bugs |
| Sequence CRDT (Yjs/Automerge/Fugue class) | Merge in any order; offline and P2P native; server is optional | Per-char IDs + tombstones; server-side validation is awkward; interleaving anomalies on concurrent inserts | Storage and bandwidth; security team (enforcement) |
| Hybrid: CRDT data model + central server ordering | CRDT convergence plus a server for auth, durability, and snapshots | Pays both metadata and server cost | Budget; complexity |
The Staff default: server-centric OT-style ops for online-first rich text. The server is required anyway for permissions, retention, and search indexing, so the key CRDT benefit — no server — is worth nothing here, while its costs (metadata, tombstone GC, enforcement) are real.
When to deviate: choose a CRDT when (a) offline editing for days is the headline promise, (b) you want peer-to-peer or end-to-end encrypted collaboration where the server cannot read ops, or (c) you're a small team that wants to buy a mature library (Yjs) instead of writing transform functions. Point (c) is underrated: "A CRDT library we don't have to maintain beats an OT implementation we'd get subtly wrong."
🎯 Staff Move: "The algorithm choice is a two-way door if we keep the op log format and client protocol abstract. It becomes a one-way door once we have a billion docs stored in that format. So I'll version the op encoding from day one."
3.2 Fault Line 2: Stateful Session Ownership vs Stateless Scale-out#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Single owner per doc (lease + fencing) | Total order in memory; transform against local state; cheap | Owner failure = brief freeze; hot docs pin one machine | Collab infra on-call (failover tuning) |
| Stateless servers + ordered log (e.g., per-doc Kafka partition key) | Any server can accept; log provides order | Transform needs current state → every server reloads; latency of log round-trip on every op | Latency budget (+10–30ms); log ops team |
Database-as-sequencer (conditional write on rev) | No leases; storage enforces order | Contention: concurrent submitters retry CAS; ~5% conflict rate ceiling before retries dominate | Users on busy docs (retry latency) |
The Staff default: single owner per doc with a lease and fencing token, backed by a conditional-write op log. The log's conditional write is the backstop that makes split-brain harmless: even if two servers believe they own the doc, only one can append rev N.
When to deviate: for very low-concurrency products (≤3 editors typical), database-as-sequencer is simpler — no lease service, no router. Contention stays below the ~5% CAS-conflict threshold, beyond which retries dominate. See Contention.
3.3 Fault Line 3: Latency vs Durability of Acks#
| Ack Point | Latency Added | Loss Window | Who Pays |
|---|---|---|---|
| On receipt (memory) | ~1ms | Everything since last flush on owner crash | Users — "Docs ate my work" |
| After local fsync | 2–5ms | Everything on machine/disk loss | Users on rare hardware failures |
| After replicated append (quorum, cross-zone) | 5–15ms | None for single failures | Latency budget — invisible thanks to optimistic local apply |
| After cross-region replication | 60–150ms | None for region loss | Remote-edit latency visibly worse |
The Staff default: ack after quorum append within the doc's home region. Region loss is handled by asynchronous cross-region replication with an RPO of a few seconds — and product signs off that a full region loss may lose the last ~5s of edits.
When to deviate: regulated documents (legal, financial filings) may justify synchronous cross-region acks for a specific doc class. Price it: 100ms+ added to ack latency, only on those docs.
🎯 Staff Move: "The user never waits for the ack to see their own edit — optimistic local apply takes care of perceived latency. The ack only controls when the edit leaves the pending queue. That's why I can afford a replicated append: 10ms here costs nothing the user can see."
3.4 Fault Line 4: Offline Freedom vs Merge Sanity#
| Policy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| No offline editing | Simple; no divergent merges | Flights, trains, flaky networks | Mobile/travel users |
| Unbounded offline with auto-merge | Maximum freedom | 50K-op rebases, interleaved paragraphs, edits by revoked users | Other collaborators; security |
| Bounded offline (time/op-count) + fork beyond | Predictable merges; explicit fallback | Occasional "your offline copy" fork | Users past the bound (they resolve manually) |
The Staff default: bounded offline — 30 days or 10,000 ops of server-side divergence, whichever first — with permission re-check at merge against the current ACL. Beyond the bound, create a sibling doc and show a diff.
When to deviate: local-first products where offline is the promise use CRDTs and unbounded merge, and accept interleaving anomalies. Enterprise docs with DLP (data-loss prevention) might disable offline entirely by admin policy.
3.5 Fault Line 5: Storage Cost vs History Fidelity#
Every keystroke is an op. A heavily edited doc generates 50K–500K ops per year. At ~40 bytes per op that's 2–20MB of history per active doc — tiny per doc, enormous at 1B docs.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Keep every op forever, hot | Perfect history, audit, blame | Storage bill grows linearly forever | Finance |
| Keep ops hot 90 days, then snapshots at named/auto versions | Cheap; history still browsable at coarse grain | Keystroke-level blame lost after 90 days | Users who want fine-grained history of old edits (rare) |
| Compact to snapshots only | Cheapest | No history, no audit | Compliance; users |
The Staff default: ops hot for 90 days, then compacted into periodic snapshots (e.g., one per editing session) moved to cold object storage. Legal-hold docs are exempted from compaction by policy. See Blob Storage for tiering economics.
4. Failure Modes & Operational Reality#
4.1 Split-Brain Ownership — Permanent Divergence#
Scenario: A network partition isolates session server S1 from etcd but not from clients. S1's lease expires; S2 takes ownership. Without fencing, both accept edits.
t=0: Partition: S1 cannot reach etcd; S1 still reaches 12 of 30 clients
t=+10s: Lease expires; router assigns doc to S2; 18 clients reconnect to S2
t=+10s: S1 (no fencing) keeps accepting edits: rev 5001..5040 on S1, 5001..5063 on S2
t=+2min: Partition heals. Two histories claim rev 5001–5040. Snapshot compactor picks S2's.
t=+1hr: 12 users report their last two minutes of work vanished.
Detection: oplog_append_rejected_stale_token_total (should be >0 during partitions — that's fencing working); doc_divergence_detected_total from client hash checks; lease_lost_while_serving_total.
Mitigation: Fencing token on every append — S1's writes are rejected at t=+10s; S1 closes its sockets, clients reconnect to S2 and resend unacked ops. Loss becomes zero.
Prevention: Owner self-fences: if it cannot renew its lease within 70% of TTL, it stops acking before the lease expires. Owner: collab infra.
4.2 The Mega-Doc — Hot Document Fan-out#
Scenario: The CEO shares the all-hands notes doc with 8,000 employees 5 minutes before the meeting. 40 people are editing; 8,000 are watching.
t=0: Link posted in #all-company
t=+30s: 8,000 JOINs on one doc → all routed to one session server
t=+45s: Owner CPU 100%: each op × 8,000 broadcasts + 8,000 presence streams
t=+60s: Ack latency p99 climbs from 40ms to 4s; editors see "Saving…" forever
t=+90s: Owner OOM; failover; 8,000 clients reconnect simultaneously → new owner dies too
Detection: session_subscribers{doc} > 500; op_ack_latency_p99 > 1s; session_failover_total rising for the same doc_id.
Mitigation: Split editors from viewers. The owner serves only editors (capped at 100). Viewers attach to a read-only fan-out tier that subscribes to the owner's op stream once per relay node and rebroadcasts. Presence for viewers is aggregated ("8,012 viewing") rather than per-cursor. Reconnects use jittered backoff (base 1s, max 30s).
Prevention: Admission control at JOIN: beyond 100 editors, new joiners are viewers. Owner: collab infra; product signs off on the cap.
4.3 Transform Bug — Silent Divergence#
Scenario: A new feature adds "suggestion mode" ops. The transform of suggest_delete against a concurrent format has an off-by-one. Clients converge on different documents.
t=0: Deploy editor v412 with suggestion mode to 10% of clients
t=+2hr: ~0.3% of docs with concurrent suggest+format diverge
t=+2hr: Nobody notices: each user sees a coherent doc, just not the same doc
t=+3days: Customer: "My colleague sees a different paragraph than I do"
Detection: Clients send hash(doc_state) with their ack every N revisions; server compares against its own. doc_divergence_detected_total > 0 pages. Fuzz tests in CI: 1M random concurrent op pairs per build, asserting TP1.
Mitigation: Server state is truth. Divergent client discards local state, reloads from snapshot + ops, replays its pending queue. Feature-flag kill switch for the new op type.
Prevention: Every new op type requires transform functions against every existing type (an N² matrix) plus property tests before rollout. Owner: editor team for correctness, collab infra for detection.
4.4 Snapshot Corruption — Poisoned Cold Load#
Scenario: The compactor writes a snapshot at rev 90,000 with a bug that drops table cells. Every subsequent open of the doc loads the bad snapshot.
Detection: Compactor verifies each snapshot by replaying from the previous snapshot and comparing hashes before publishing; snapshot_verify_failed_total. Weekly sampled audit replays 0.1% of docs from rev 0.
Mitigation: Snapshots are cache; delete the bad snapshot and rebuild from the op log. Because the log is the source of truth, this is always recoverable — if you never compacted away the ops the bad snapshot replaced.
Prevention: Never delete ops until the snapshot replacing them has been verified and aged ≥7 days. Owner: collab infra.
4.5 Reconnect Storm After Regional Blip#
Scenario: A 20-second network blip in one zone drops 400K WebSocket connections. All reconnect at once.
t=0: Zone blip; 400K sockets drop
t=+1s: 400K reconnects hit gateways; TLS handshakes saturate CPU
t=+5s: Each JOIN triggers SYNC → snapshot reads spike 50×
t=+30s: Snapshot store throttles; JOINs time out; clients retry → second wave
Detection: ws_connect_rate > 5× baseline; snapshot_read_throttled_total; join_latency_p99.
Mitigation: Jittered exponential backoff in the client (full jitter, base 500ms, cap 30s). Clients that still hold local state send last_known_rev so SYNC returns only tail ops, not a snapshot. Gateway admission control sheds JOINs with 503 + Retry-After.
Prevention: Game day: drop a zone's connections quarterly. Owner: collab infra + edge team. See Real-time WebSockets.
4.6 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Owner crash | session_failover_total, lease expiry | Docs on that server (~10K), 5–12s freeze | Lease expiry, client resend | Collab infra |
| Split-brain | oplog_append_rejected_stale_token_total | One doc per partition | Fencing token | Collab infra |
| Transform bug | doc_divergence_detected_total | Docs using new op types | Kill switch, reload from server | Editor team |
| Mega-doc | session_subscribers > 500 | One doc + co-located docs | Viewer relay tier, editor cap | Collab infra |
| Snapshot corruption | snapshot_verify_failed_total | Docs compacted by bad build | Rebuild from log | Collab infra |
| Op log slow | oplog_append_p99 > 50ms | All docs in a cell | Shed presence, batch ops, fail over log | Storage platform |
| Permission propagation lag | acl_revocation_lag_p99 > 10s | Revoked users keep editing | Push revocation to owner; kick session | Identity + collab |
| Reconnect storm | ws_connect_rate spike | A zone / region | Jitter, tail-only SYNC, admission control | Edge + collab |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Algorithm | Explains OT or CRDT correctly | Chooses based on doc model and offline intent; names the costs of the other | Treats op format as a versioned platform contract across products |
| Ordering | Server assigns versions | Single fenced owner; explains split-brain and fencing | Cell architecture bounding owner failures to ≤5% of docs |
| Durability | "We save to the database" | Ack after replicated append; idempotent resend | Durability SLO with independent replay auditor and error budget |
| Scale | Shards by doc ID | Separates editors from viewers; caps editors; lazy sessions | Prices session fleet vs storage; cold-tiers history; knows where the money goes |
| Offline | Queue and replay | Bounded window, fork, permission re-check | Admin policy per tenant; DLP integration; legal sign-off |
| Operations | Mentions monitoring | Divergence hashes, fuzz tests, named metrics and owners | Game days, org-level incident taxonomy, collaboration SLOs in exec reviews |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Defines the ack point unprompted | "An edit is acknowledged only after quorum append. Local apply gives speed; the ack gives safety." |
| Uses fencing, not just leases | "The lease tells a server it may own the doc. The fencing token makes sure storage agrees." |
| Separates presence from ops | "Cursors are lossy, throttled to 5–10 Hz, and never persisted. I'll drop them first under load." |
| Bounds the ugly cases | "Offline merges are automatic up to 30 days or 10K ops; beyond that we fork and show a diff." |
| Names product sign-off | "Delete-vs-format conflict resolves delete-wins. Product signs off; it's in the test suite." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| 15 minutes of transform math, no failure story | Optimizes the solved part; ignores the part that pages people |
| "Stateless WebSocket servers behind a load balancer" | Doesn't recognize the sequencer is inherently stateful |
| Uses wall-clock timestamps to order edits | Clock skew of 10–100ms across clients reorders keystrokes |
| No distinction between viewers and editors | Mega-doc will melt the design |
| "CRDTs solve everything" with no mention of tombstones or enforcement | Knowledge without cost-awareness |
5.4 Common False Positives#
- Deep OT theory ≠ collaborative editing design. Knowing TP1/TP2 is nice; knowing why the server-centric model avoids TP2 is the signal.
- Name-dropping Yjs ≠ understanding CRDT costs. Ask them what happens to a 1M-character doc after 3 years of edits.
- A beautiful architecture diagram ≠ an owner. If no box has a team name, the design has no pager.
- "We'll use Spanner" ≠ a durability story. Strong storage doesn't help if the owner acks before writing.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Doc model, offline intent, server validation, ack promise |
| Entities & API | 3–5 min | Op shape, rev, idempotency key, WebSocket messages |
| Architecture | 5–10 min | Sequencer, op log, snapshots, presence, leases |
| Deep dive 1 | 10–20 min | Ownership + failover (fencing, resend) |
| Deep dive 2 | 20–30 min | Durability / snapshots / cold load |
| Deep dive 3 | 30–40 min | Offline + permissions or mega-doc fan-out |
| Wrap-up | 40–45 min | Evolution, what you'd build next, owners |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response |
|---|---|---|
| "Now make it work offline for a month." | Do you know OT's weakness and have a policy? | Bounded rebase, fork beyond bound, or switch to CRDT with stated costs |
| "Two regions, editors on both coasts." | Sequencer placement | Home region per doc; migrate home on sustained editor shift; ~70ms cross-country penalty for remote editors |
| "Add comments anchored to text." | Position tracking under concurrent edits | Anchors are transformed like cursors; orphaned anchors resolve to nearest surviving char |
| "Legal needs to see who typed every character." | History fidelity vs cost | Op log with author per op, retained per legal hold policy |
| "Your transform has a bug in prod." | Detection of silent failure | Client state hashes, kill switch, server-is-truth reload |
6.3 What to Deliberately Skip#
- Rich-text schema details (bold/italic run merging) — one sentence.
- The full OT transform table — state insert/insert tie-break and offer more.
- WebSocket vs SSE debate — pick WebSocket.
- Spell-check, autocomplete, export to PDF — out of scope unless asked.
6.4 Follow-Up Questions to Expect#
- "What exactly happens when two users insert at the same position at the same time?"
- "The server holding a doc crashes mid-append. Walk me through the next 15 seconds."
- "How do you open a 5-year-old, 1M-character document in under a second?"
- "A user's access is revoked while they're editing. How fast does it take effect?"
- "How would you detect that two clients are seeing different documents?"
- "How does undo work when other people have edited since?"
- "What does it cost to store every keystroke for 1B documents?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design Google Docs."
Staff Answer
"Before I draw anything: what's the document model, is offline a core promise, and must the server validate each edit? I'll assume rich text, online-first with bounded offline, and server-enforced doc-level permissions — that's Google Docs. Targets: local edits render within one frame (16ms), remote edits visible within 200ms p50 in-region, up to 100 concurrent editors with unbounded viewers on a separate tier, and zero loss of acknowledged edits. That last one is the promise that shapes everything: an edit is acknowledged only after it's durably appended to the doc's op log. I'll use a single fenced owner per doc as the sequencer, OT-style ops, snapshots for fast load, and presence on a lossy side channel."
Why this is L6:
- Asks the three questions that actually select the algorithm, then commits
- States a durability promise with a precise ack point in the first 60 seconds
- Separates editors from viewers before being prompted
What L7 adds:
- "Is this a one-off for Docs, or the sync substrate for Docs, Sheets, Slides and Comments? If the latter, the op log and client SDK are a platform contract and I'd design them for four consumers."
- Names the cost driver up front: history storage, not session compute
❌ Common L5 Trap
"We'll use OT like Google Docs. Clients connect via WebSocket to a server, which transforms ops and saves them to a database."
Why this misses: Correct but unowned. The interviewer asks "what if the server dies after broadcasting but before saving?" and there's no answer, because the ack point was never defined.
Drill 2: Concurrent Inserts at the Same Position#
Prompt: "Alice and Bob both insert at position 10 at the same time, based on rev 50. What happens?"
Staff Answer
Alice's insert(10, "X") arrives first; server assigns rev 51. Bob's insert(10, "Y") arrives with base_rev=50, so the server transforms it against rev 51. Two inserts at the same index need a deterministic tie-break — in the server-centric model the op already sequenced wins the left position, so Bob's op becomes insert(11, "Y") and is assigned rev 52. Result everywhere: ...XY.... On Alice's client, the remote op rev 52 arrives as insert(11,"Y") and applies directly. On Bob's client, rev 51 arrives while his op is pending; the client transforms rev 51 against its pending op — Alice's insert stays at 10, Bob's pending local char shifts to 11. Both converge on XY.
The important part: the tie-break is deterministic and every replica uses the same one. In CRDTs the tie-break is by element ID (e.g., Lamport clock then client ID). The user-visible result may be XY or YX; product doesn't care which, only that it's consistent.
Why this is L6:
- Shows the 2-party transform on both server and client sides
- Explains why server-centric OT avoids the harder N-party problem
- Separates "deterministic" (required) from "which order" (irrelevant)
What L7 adds:
- Insists on a property-based test harness (random op pairs, assert convergence) as a release gate for the whole org's op types
- Notes that the transform matrix grows N² with op types — a governance problem for which teams may add ops
Drill 3: The Owner Dies#
Prompt: "The session server for a doc with 30 editors crashes. Walk me through the next 15 seconds."
Staff Answer
t=0: sockets drop. Clients show "Reconnecting…" and keep applying edits locally into their pending queues — no text is lost locally. t=0–10s: the owner's lease in etcd expires (TTL 10s). Clients retry with jittered backoff (0.5–2s). t≈10s: the router assigns the doc to a new server with fencing token N+1. It loads the latest snapshot (e.g., rev 48,000) plus the tail from the op log (e.g., 312 ops) — ~100ms. t≈10–12s: clients reconnect with last_acked_rev and resend all unacked ops with (client_id, client_seq). The new owner dedupes any that actually landed before the crash (the old owner appended but didn't get to ack), transforms the rest, appends, acks. Remote edits resume. Total: ~12s of frozen remote updates, zero lost acked edits, zero lost local edits.
To shorten the freeze I could have clients detect the crash and ask the router to force-expire — but only if the old owner is provably dead; otherwise fencing protects us but we'd flap. I keep 10s.
Why this is L6:
- Distinguishes local edits (safe in client queue) from acked edits (safe in log)
- Handles the "appended but not acked" gap via idempotent resend
- Justifies the lease TTL as a tradeoff, not a default
What L7 adds:
- Budgets failover into an SLO: "99.9% of doc-minutes have remote-edit latency <1s" — 12s failovers spend the budget, so the owner fleet's crash rate becomes a tracked number
- Cells: a server crash is expected; a cell-wide crash is a SEV1 with a defined blast radius
Drill 4: Make Durability Concrete#
Prompt: "You say 'ack after durable append.' What's the storage, the key, and the write?"
Staff Answer
Table ops in a strongly consistent, horizontally partitioned KV (DynamoDB, Spanner, or a Raft-replicated store). Partition key doc_id, sort key rev. Item: {author, client_id, client_seq, ops_blob, fencing_token, ts}. The write is conditional:
PutItem ops (doc_id, rev=head+1, ...)
condition: attribute_not_exists(rev)
AND doc_meta.fencing_token <= :my_token
In practice I keep doc_meta.{head_rev, fencing_token} in the same partition and use a transaction so rev assignment and token check are atomic. Latency: 5–15ms in-region. Throughput: a doc peaks at ~250 ops/sec (50 heavy typists), so I batch ops arriving within a 10–20ms window into one append — cuts writes 5–10× on busy docs. A second index on (client_id, client_seq) per doc (or a small bounded dedupe window held in the owner) handles resend dedupe.
Why this is L6:
- Concrete keys, concrete condition, concrete latency
- Batching quantified against the per-doc write rate
- Dedupe designed in, not bolted on
What L7 adds:
- Chooses managed storage deliberately: "Owning a Raft log for a billion docs is two engineers forever. DynamoDB costs more per op but removes that headcount."
- Defines the storage contract once so other products reuse it
Drill 5: The Mega-Doc#
Prompt: "A doc has 60 editors and 5,000 viewers. What breaks?"
Staff Answer
Fan-out breaks first, not sequencing. 60 editors × ~3 ops/sec = 180 ops/sec is fine for one owner. But broadcasting each op to 5,060 sockets is ~900K messages/sec, plus presence for 5,060 cursors at 5 Hz squared if naïvely broadcast to everyone — impossible. Fix: editors connect to the owner; viewers connect to a relay tier. The owner publishes the op stream once per relay (say 3 relays), each relay fans out to ~1,700 viewers. Viewers don't get individual cursors — they get the 60 editors' cursors at 2 Hz and an aggregate "5,000 viewing" count. Editor count is capped at 100 via admission control; the 101st becomes a viewer with a "request edit" button.
Why this is L6:
- Identifies fan-out, not sequencing, as the bottleneck with arithmetic
- Separates editor and viewer paths
- Presence downsampled for viewers
What L7 adds:
- The relay tier is the same machinery as live-blog / live-dashboard fan-out elsewhere in the org — build it once
- Product sign-off on the 100-editor cap as a published limit, like Google Docs does
Drill 6: Offline for Three Weeks#
Prompt: "Alice edits offline for 3 weeks. Meanwhile 40,000 ops landed and her edit access was downgraded to comment-only. She reconnects."
Staff Answer
Step 1: re-authorize against the current ACL — she's now a commenter, so her 600 pending edit ops cannot be applied. Step 2: we don't drop them silently. Her client keeps them and the UI offers: "Your offline changes couldn't be applied because your access changed. Save as a copy or send as suggestions." Suggestions mode is the elegant answer — commenters can suggest. Step 3: even if she were still an editor, 40,000 ops exceeds our 10K-op rebase bound, so we'd fork: create "Alice's offline copy" and show a compare view. Transforming 600 ops against 40,000 is ~24M transform steps — seconds of CPU on the owner, blocking everyone else on that doc.
Why this is L6:
- Permission re-check at merge time, not edit time
- Quantified rebase cost justifies the bound
- Never silently drops user work; gives a recoverable path
What L7 adds:
- Makes the offline bound a tenant-admin policy (some enterprises want 0, some want 90 days) with legal/security review
- Rebases run on a separate worker pool so large merges can't starve live sessions — an org-wide principle: batch work never shares a hot path
Drill 7: Build vs Buy the Sync Engine#
Prompt: "Should we build our own OT engine, or use Yjs, or a hosted service?"
Staff Answer
Depends on whether collaboration is our differentiator. If we're a docs company, building gives control over op types, permissions, and storage format — but it's 4–6 engineers for a year to reach production quality plus 2 forever for maintenance, and transform bugs are subtle. If collaboration is a feature of a larger product (a CRM with a notes field), use Yjs with a server-side provider: we get convergence and offline for free and own only persistence, auth, and presence. Hosted services (Liveblocks-style) are fastest to ship; the risk is data residency and lock-in on the doc format. My default for a non-docs company: Yjs + our own persistence, with the stored format being Yjs update blobs plus our own periodic JSON export so we can migrate away.
Why this is L6:
- Frames build vs buy by differentiation and headcount
- Names lock-in (doc format) and a mitigation (neutral export)
What L7 adds:
- Prices it: build ≈ $1.5–2M/year fully loaded vs hosted at usage pricing; crossover point by MAU
- Plans the exit before the entrance — deprecation path written into the vendor decision doc
Drill 8: Changing the Op Format Without an Outage#
Prompt: "We need to add table-merge ops. Old clients are still in the field for 6 months."
Staff Answer
Version the protocol. Clients declare protocol_version on JOIN. The owner knows the minimum version among connected editors. New op types are only emitted when all connected editors support them — otherwise the new feature is disabled in that session ("Update to merge cells"). Rollout: (1) ship server support dark, (2) ship clients that understand but don't emit, (3) enable emission behind a flag at 1% → 10% → 100% of docs with divergence monitoring, (4) after 6 months, enforce a minimum version. Transform functions for the new op against all existing ops are gated by fuzz tests before step 3.
Why this is L6:
- Mixed-version sessions handled explicitly
- Staged rollout with divergence as the canary metric
What L7 adds:
- A written op-type governance process: who may add an op type, required tests, deprecation schedule
- Minimum-client-version policy coordinated with mobile release trains
Drill 9: Cost#
Prompt: "Storage is growing 40% a year. Where's the money going?"
Staff Answer
Almost certainly op history. Back-of-envelope: 1B docs, 5% active monthly, active docs average 20K ops/year at ~40 bytes = 800KB/year each. 50M × 0.8MB = 40TB/year of new ops — plus replicas (×3) and indexes. Snapshots are smaller (median doc ~20KB). Actions: (1) cold-tier ops older than 90 days to object storage at ~1/5 the $/GB, (2) compact old ops into per-session snapshots, keeping author attribution at session granularity, (3) never compact legal-hold docs. Expected: 60–80% storage cost reduction on history with no user-visible change for >99% of opens.
Why this is L6:
- Arithmetic from docs to bytes
- Tiering + compaction with explicit exception for legal hold
What L7 adds:
- Turns retention into a policy decision owned by legal and product, not an infra default
- Tracks $/active-doc as the unit-economics metric reported quarterly
Drill 10: Multi-Region#
Prompt: "We're expanding to EU with data residency. Editors on one doc can be in the US and EU."
Staff Answer
Each doc has a home region where its owner and op log live; residency rules pin EU-tenant docs to EU. A US editor on an EU doc connects to the nearest edge, which proxies the WebSocket to the EU owner — remote edits for them cost ~80–100ms extra, local edits are still instant. For docs without residency constraints, home region follows the majority of editing activity: if >70% of ops over 7 days come from another region, migrate home — drain the session, copy the log tail, flip the lease region, clients reconnect. We don't do multi-master sequencing across regions; cross-region consensus on every keystroke adds 70–150ms and buys nothing users notice.
Why this is L6:
- Home-region model with explicit latency cost for remote editors
- Residency as a hard constraint on placement
- Rejects multi-master with a quantified reason
What L7 adds:
- Region migration as a platform capability used by every stateful product
- Residency compliance audited continuously, with legal as the sign-off owner
8. Deep Dive Scenarios#
Deep Dive 1: Monday-Morning Peak Incident#
Context: At 9:05am Monday, op_ack_latency_p99 jumps from 40ms to 3.5s across one cell. Users see "Saving…" that never clears. On-call escalates to you.
Questions to Surface First:
- Is it every doc in the cell or a subset? (Hot doc vs infrastructure)
- Is the op log slow, or are owners CPU-bound?
- Did anything deploy in the last hour — editor client, session server, storage config?
- Are clients still able to edit locally? (Is this latency or loss?)
Typical L5 Approach: Scale up session servers in the cell and increase op log capacity. Reasonable, but blind — if one mega-doc is the cause, more servers don't help it.
Staff Approach: Check
oplog_append_p99first. If storage is healthy, it's owner-side: look for docs withsession_subscribers> 500 on the slow owners. Found: a company-wide OKR doc with 6,000 joiners pinned to one owner, starving 9,000 co-located docs. Immediate: move viewers of that doc to relay tier; migrate co-located docs off the hot owner. Then fix admission control that should have capped editors at 100.
Principal Approach: Co-location is the systemic bug. One hot doc shouldn't degrade 9,000 innocent docs. Introduce per-doc resource quotas inside owners and bin-packing that isolates docs above a subscriber threshold onto dedicated owners. Add "mega-doc" as a first-class product state with its own UX (viewer mode by default), and make the capacity model assume a weekly peak of N all-hands docs per large tenant.
Staff Approach — Full Reasoning
| Phase | Action |
|---|---|
| Immediate (0–5 min) | Confirm clients edit locally (no loss). Check oplog_append_p99 vs owner_cpu. Identify hot doc_ids by subscriber count. |
| Triage | Hot doc with 6K joins; admission cap config was not applied to that tenant. |
| Quick fix | Force viewer mode beyond 100 editors; migrate 9K co-located docs to other owners (brief reconnect each). |
| Guardrails | Per-owner subscriber ceiling; auto-isolate docs >500 subscribers. |
| Post-mortem | Why was the tenant exempt from the cap? Who owns tenant-level config? |
Metrics to Watch: op_ack_latency_p99 by cell, session_subscribers top-K, owner_cpu, oplog_append_p99, session_migration_total.
Organizational Follow-up: Tenant config changes require review from collab infra; capacity planning includes a "Monday all-hands" peak model.
Ownership Question: Who owns the editor cap — product or infra? Product owns the number; collab infra owns enforcement and alerting when it's bypassed.
Key Takeaway: "Hot docs are a fan-out problem, and co-location turns one hot doc into a cell-wide incident."
What clears the Staff bar:
- Rules out storage before scaling compute
- Protects innocent co-located docs first
- Converts a config gap into an ownership fix
Deep Dive 2: The Silent Divergence#
Context: A customer reports two colleagues see different text in a contract. No alerts fired. It's been happening for 4 days.
Questions to Surface First:
- Do we have client state hashes, and are they being checked or just logged?
- What changed 4–5 days ago in the editor or session server?
- Which op types were involved in the divergent docs?
- Which version is "right" — and does the server's version match either?
Typical L5 Approach: Reproduce the bug, fix the transform, force affected clients to reload. Correct fix for the code bug.
Staff Approach: Treat it as a detection failure first. The transform bug is one bug; the missing alert let it live 4 days. Server state is truth: force reload for divergent sessions, find all affected docs by scanning hash-mismatch logs, kill-switch the new op type. Then wire
doc_divergence_detected_totalto page, and add the offending op pair to the fuzz corpus.
Principal Approach: Divergence is a correctness SLO with zero tolerance, and it had no owner. Establish it as a tier-0 metric with an owning team, require every new op type to pass a convergence fuzz gate owned by a central correctness group, and schedule quarterly "divergence game days" where a deliberately broken transform is injected in staging to prove detection works.
Staff Approach — Full Reasoning
| Phase | Action |
|---|---|
| Immediate (0–5 min) | Kill-switch the new op type. Confirm server state vs both clients. |
| Triage | Scan hash-mismatch logs for 5 days; list affected docs (~0.3% of docs with concurrent suggestions). |
| Quick fix | Force reload for affected sessions; notify affected customers where user-visible text differed. |
| Guardrails | Divergence metric pages at >0 over 5 min. Fuzz tests gate deploys. |
| Post-mortem | Why was the hash computed but not alerted? Who owns correctness metrics? |
Metrics to Watch: doc_divergence_detected_total, client_reload_forced_total, fuzz test pass rate per build.
Organizational Follow-up: Correctness metrics get the same paging tier as availability. New op types need sign-off from the collab infra correctness owner.
Ownership Question: Who gets paged for divergence? Editor team on-call, because they own transform correctness — collab infra owns the detector's uptime.
Key Takeaway: "Silent divergence is worse than an outage — nobody knows to complain. Detection is the feature."
What clears the Staff bar:
- Frames detection gap as the root cause
- Server-is-truth recovery path
- Correctness metric wired to paging
Deep Dive 3: Onboarding a 200K-Seat Enterprise#
Context: A 200,000-employee company is migrating from a competitor next quarter. They require EU data residency, 7-year retention of all edit history, and admin-controlled offline policies.
Questions to Surface First:
- How many docs and how much history are they importing, and in what format?
- Does "7-year edit history" mean keystroke-level or version-level?
- Does residency include presence and relay traffic, or only stored data?
- What's their peak concurrency pattern (all-hands docs, quarter-end planning)?
Typical L5 Approach: Provision more capacity in the EU region and extend retention for their tenant.
Staff Approach: Clarify "edit history" with their compliance team — version-level (one snapshot per session with authors) is ~50× cheaper than keystroke-level. Pin their tenant's docs to EU cells. Import runs as a batch pipeline creating docs with a synthetic single-op history, rate-limited so it can't touch live cells. Offline policy becomes a tenant setting. Capacity: 200K seats, ~10% concurrent at peak, ~2 docs open each → 40K sessions; that's ~4 more owners per EU cell — fine.
Principal Approach: This customer is the forcing function for tenant-level policy as a product: residency, retention, offline, and editor caps become a policy engine consumed by every collaborative product, not per-customer config. Price the retention: 7 years of keystroke-level history for 200K users is a line item sales must see before signing.
Staff Approach — Full Reasoning
| Phase | Action |
|---|---|
| Clarify | Retention granularity, residency scope, import volume |
| Placement | Dedicated EU cells for the tenant; relay nodes in EU |
| Import | Batch pipeline, throttled, isolated worker pool |
| Policy | Tenant-level retention exempt from 90-day compaction; offline policy toggle |
| Verification | Residency audit: no op or snapshot for tenant outside EU |
Metrics to Watch: import_docs_per_sec, cell_owner_cpu{tenant}, residency_violation_total (must be 0), history storage by tenant.
Organizational Follow-up: Sales requires infra sign-off on non-default retention; legal owns residency audit.
Ownership Question: Who owns the tenant's retention cost? The account's P&L — the platform bills retention as a tier, so the decision is priced where it's made.
Key Takeaway: "Enterprise onboarding is a policy problem wearing a capacity costume."
What clears the Staff bar:
- Negotiates the requirement (version vs keystroke) before building
- Isolates import from live traffic
- Turns a one-off into tenant policy
Deep Dive 4: Post-Mortem — Lost Edits During a Deploy#
Context: During a session-server rolling deploy, ~2,000 users lost 5–30 seconds of work. The post-mortem is yours to lead.
Questions to Surface First:
- Were the lost edits acked or unacked?
- Did draining servers ack before appending?
- Did clients resend unacked ops on reconnect — or discard their pending queue?
- Why didn't staging catch it?
Typical L5 Approach: Found: the new server version acked on receipt to "reduce latency," then was SIGTERM'd before flushing. Revert the change and add a flush on shutdown.
Staff Approach: The revert is right, but the root cause is that a single PR could move the ack point without anyone noticing. The ack-after-append rule must be enforced structurally: the ack message is only constructible from an append result. Add a deploy-time invariant test (kill -9 an owner under load in staging; assert zero acked-op loss by replaying logs). Drain protocol: stop accepting new ops, finish in-flight appends, hand off leases, then exit.
Principal Approach: The durability promise needs an independent auditor. Nightly, sample 1% of sessions: compare client-reported acked revs against the op log. Report "acked edits lost" as an org-level SLO with an error budget of zero — breaching it freezes collab deploys until fixed. That turns a norm into a mechanism.
Staff Approach — Full Reasoning
| Phase | Action |
|---|---|
| Timeline | Deploy started 14:02; ack-on-receipt PR merged 3 days earlier; loss during pod termination |
| Root cause | Ack point moved; drain didn't flush; client trusted ack and cleared pending queue |
| Fix | Revert; type-level enforcement of ack-after-append; drain protocol |
| Guardrails | Chaos test in CI; durability auditor |
| Comms | Affected users notified; restore from client crash logs where possible |
Metrics to Watch: acked_ops_missing_from_log_total, drain_duration_seconds, pending_queue_cleared_without_ack_total.
Organizational Follow-up: Changes to ack semantics require design review from collab infra leads.
Ownership Question: Who approves changes to the ack point? Collab infra tech lead — it's a durability contract, not an implementation detail.
Key Takeaway: "If one PR can silently move your ack point, you don't have a durability guarantee — you have a convention."
What clears the Staff bar:
- Enforces invariants structurally, not by review alone
- Adds chaos testing to catch this class
- Makes durability measurable
Deep Dive 5: Multi-Region Expansion#
Context: Leadership wants the product in APAC. Current architecture is single-region US. Latency for APAC users is ~200ms to the US owner.
Questions to Surface First:
- Are APAC users editing APAC-created docs, or collaborating with US teams?
- Any residency requirements (Japan, Australia, India)?
- Is the op log store multi-region capable, or do we deploy new cells?
- What's the budget for region #2?
Typical L5 Approach: Deploy a full stack in APAC and replicate everything bidirectionally.
Staff Approach: Home-region model. New docs are homed where created; each region runs its own cells. Cross-region collaborators connect via edge to the home owner — they pay ~150–200ms on remote edits, local edits still instant. Async replication of op logs to a secondary region gives disaster recovery with ~5s RPO. Home migration when >70% of editing moves. No multi-master sequencing.
Principal Approach: Region expansion is a template, not a project. Codify it: a region "kit" (cells, lease service, op log, relay, snapshots) deployable by one team in weeks; a placement service shared across collaborative products; and a documented DR posture per region with annual failover exercises. Price region #3 and #4 before approving #2.
Staff Approach — Full Reasoning
| Phase | Action |
|---|---|
| Placement | Doc home = creator's region unless residency says otherwise |
| Access | Edge proxy to home owner; relay tier in each region for viewers |
| DR | Async log replication, RPO ~5s, RTO ~5 min with lease re-home |
| Migration | Activity-based home migration, drain + lease flip |
| Validation | Region-kill game day before GA |
Metrics to Watch: remote_edit_latency_p50{editor_region, home_region}, home_migration_total, replication_lag_seconds.
Organizational Follow-up: Placement service owned by collab infra; regional on-call rotation.
Ownership Question: Who decides a doc's home region when residency and latency conflict? Residency always wins; legal owns the policy, collab infra enforces it.
Key Takeaway: "Home the sequencer, proxy the remote editors. Consensus across oceans for every keystroke buys nothing."
What clears the Staff bar:
- Rejects multi-master with a latency argument
- DR with explicit RPO/RTO
- Residency as a hard placement constraint
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Choose OT-style or CRDT based on document model, offline intent, and server validation needs — and name the cost of the one you didn't pick
- Define the ack point precisely and defend "ack after replicated append" with latency numbers
- Design single-owner-per-doc sequencing with leases and fencing tokens, and walk through failover second by second
- Separate editors from viewers and presence from ops, with concrete caps and rates
- Bound offline merges, re-check permissions at merge, and fork beyond the bound
- Build detection for silent divergence and acked-op loss
- Price history storage and propose tiering with a legal-hold exception
- Explain how the design becomes a platform for multiple collaborative products (L7)
The Bar for This Question#
Mid-level (L4): Builds a working real-time editor: WebSockets, a server that applies ops, a database. May use timestamps or whole-doc saves. Understands conflicts exist but resolves them naïvely.
Senior (L5): Correctly explains OT or CRDT mechanics, shards by doc ID, uses sticky sessions, snapshots for load time. The design converges in the happy path. Gaps: ack point undefined, failover hand-waved, viewers and editors treated the same, offline and permissions bolted on.
Staff+ (L6): Starts from intent and commits. Makes durability a stated promise with a precise ack point. Designs fenced single ownership with idempotent client resend. Caps editors and splits viewer fan-out. Bounds offline with permission re-check. Names owners for correctness (editor team), durability and sessions (collab infra), and policy (product/legal). Builds detection for the silent failures. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "OT vs CRDT Is the Least Important Decision in the Design"#
| Evidence | Implication |
|---|---|
| Google Docs (OT) and Figma (CRDT-inspired, server LWW) both ship world-class collaboration | Both algorithm families work in production |
| Real incidents in collaborative products are dominated by session failover, storage, and deploy bugs | The pager rings for operations, not transforms |
| Mature CRDT libraries (Yjs, Automerge) are free to adopt | The algorithm is increasingly a dependency, not a design |
The Staff position: Spend 2 minutes on the algorithm and 30 on ownership, durability, and failure. The choice is reversible if the op format is versioned; the ack semantics are not.
Why this matters in interviews: Interviewers use the OT/CRDT debate as a trap to see whether you'll burn your time budget on it.
10.2 "Most Documents Should Never Get a Session Server"#
| Evidence | Implication |
|---|---|
| Typically <10% of docs are opened in a month; far fewer have 2+ concurrent editors | A stateful session per open doc is mostly wasted |
| Solo editing works fine with autosave + optimistic concurrency | Cheap path covers the majority |
The Staff position: Promote a doc to a live session only when a second participant joins. Demote after 5 minutes solo. Session fleet shrinks by 5–10×.
Why this matters in interviews: Shows you design for the distribution of usage, not the headline feature.
10.3 "Presence Is More Expensive Than Editing"#
| Evidence | Implication |
|---|---|
| Cursor moves happen at mouse/selection rate; edits at keystroke rate | Unthrottled presence is 5–10× op volume |
| Presence fan-out is N² in participants | 100 participants = 10,000 cursor streams |
The Staff position: Presence gets its own lossy channel, throttled to 5–10 Hz, aggregated for viewers, and is the first thing shed under load.
Why this matters in interviews: Candidates who only size ops under-provision by an order of magnitude.
10.4 "Offline Editing Is a Security Feature Request in Disguise"#
| Evidence | Implication |
|---|---|
| Offline copies live on devices outside DLP controls | Security must sign off |
| Permissions change while users are offline | Merge must re-authorize |
The Staff position: Offline is bounded, tenant-configurable, and re-authorized at merge. Never unbounded by default.
Why this matters in interviews: It moves the conversation from algorithms to ownership — which is where Staff lives.
10.5 "Version History Is Your Largest Cost Center"#
The Staff position: Compute is sized by concurrency; storage is sized by every keystroke ever typed. Tier and compact history by policy, or it will dominate the bill within 2–3 years.
Why this matters in interviews: Shows you think about the system at year 3, not day 1.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
At Staff, collaborative editing is a system to design. At Principal, it is a capability the company will need five times: docs, spreadsheets, slides, whiteboards, comments, code notebooks. Each team, left alone, will build its own sequencer, op log, presence channel, and permission check — five fleets, five on-call rotations, five subtly different durability promises. The L7 question is: what is the sync substrate, what's the contract, and which parts must stay product-specific? The op algebra is product-specific. Sessions, logs, presence, permissions, and placement are not.
The Org-Level Fault Line#
One sync platform vs per-product sync engines.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Per-product engines | Each team optimizes for its model; ships fast initially | 5× fleets, inconsistent durability, duplicated presence; every region launch done 5 times | Infra budget; on-call; customers who see inconsistent behavior |
| Central sync platform (log + sessions + presence + ACL), product-owned op types | One durability SLO, one region kit, shared relay tier | Platform becomes a bottleneck; op-type governance needed | Platform team (headcount); product teams (lose some autonomy) |
| Buy (hosted collaboration service) | Fastest; no fleet | Residency, lock-in, per-MAU cost at scale | Finance at scale; legal |
The L7 default: central platform for sessions/log/presence/ACL enforcement, with product teams owning their op types and transforms behind a plugin interface. Don't standardize the document model — that's where products differentiate.
Cost Model#
Assumptions: 40 bytes/op, 3× replication, 5% of docs active monthly, active docs avg 20K ops/year, one session server (~$400/month) holds ~10K active docs, snapshots median 20KB.
| Scale | Docs / Peak Concurrent Sessions | Monthly Infra | Headcount | On-call Load |
|---|---|---|---|---|
| Startup | 1M docs / 5K sessions | ~$3–5K (managed KV + 2–3 session servers + blob) | 2–3 engineers, part-time on-call | ~1–2 pages/month |
| Growth | 100M docs / 500K sessions | ~$80–150K (50 session servers, op log ~200TB with replicas, relay tier) | 8–12 engineers, dedicated rotation | ~5–10 pages/month |
| Hyperscale | 1B+ docs / 5M+ sessions, multi-region | ~$1.5–3M (history storage ~60% of bill before tiering) | 30–50 engineers across platform + per-product | Per-region rotations; SLO-driven |
Where money goes: at growth scale and beyond, history storage outgrows compute. Cold-tiering ops older than 90 days cuts total bill ~30–40%.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Op log storage format (encoding of ops) | One-way (after scale) | Rewrite of billions of stored ops; version it from day one |
| Client protocol (JOIN/SUBMIT/ACK semantics) | One-way-ish | Mobile clients live 6–18 months; needs multi-version server support |
| Ack point semantics | One-way (trust) | Users who lose work don't come back |
| OT vs CRDT algorithm | Two-way if format is versioned | Months of migration, but feasible |
| Lease TTL, snapshot cadence, editor cap | Two-way | Config change |
| Managed KV vs self-run log store | Two-way with effort | Dual-write migration, ~1–2 quarters |
| Retention policy | One-way once data is deleted | Deleted history cannot be restored |
The Standard I'd Write#
RFC: Real-Time Collaboration Platform Standard (v1)
Scope: Any product feature where two or more users edit shared state concurrently with sub-second visibility.
MUST:
- Use the platform session service for ordering; no product-run sequencers.
- Acknowledge edits only after durable replicated append to the platform op log.
- Tag every op with
(client_id, client_seq); servers dedupe on resend.- Enforce ACLs at join and per op; revocations take effect within 10s.
- Emit
doc_divergence_detected_total,op_ack_latency_p99, andacked_ops_missing_from_log_total.- Version op encodings; ship convergence fuzz tests with every new op type.
SHOULD:
- Keep presence on the lossy channel, ≤10 Hz per participant.
- Cap concurrent editors per document (default 100) and route excess to viewer relays.
- Use platform tiering for history older than 90 days unless under legal hold.
Exceptions: Reviewed by the collaboration platform council; granted for ≤2 quarters with a migration plan.
Success metrics: 0 acked edits lost per quarter; remote edit p50 <200ms in-region; ≤1 sync engine per company; new product onboarding ≤6 weeks.
What I'd Tell the VP#
"Real-time collaboration is becoming table stakes across our products, and right now three teams are about to build it separately. I'm proposing one collaboration platform that handles saving, ordering, and sharing permissions, while each product keeps control of its own document features. It costs about 8 engineers for a year, and it replaces roughly 20 engineers' worth of duplicated work across teams over three years. The biggest risk we're buying down is lost user work — we'll commit to zero lost saved edits and measure it. The biggest ongoing cost is edit history storage, and we'll control it with a retention policy that legal signs off on."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Platform framing | "The op algebra is product-specific; sessions, log, presence and ACL are platform." |
| Pricing history | "By year 3 history is 60% of the bill — retention is a policy decision, not an infra default." |
| One-way door awareness | "I'll version the op encoding now because it's the only decision here that's expensive to reverse." |
| Org failure posture | "Cells cap any single deploy or hot doc at 5% of documents; we prove it with quarterly game days." |
| Knowing when not to standardize | "I won't standardize the document model — that's where products compete." |
Staff answers that L7 interviewers find insufficient:
- "Collab infra owns the session servers" — correct, but doesn't address the three other teams building their own.
- "We'll add cold storage for old ops" — a tactic without a retention policy owner or a price.
- "We'll run a failover drill" — one drill for one system, not an org-wide failure posture with error budgets.
🧭 Principal Move: "Before we design Docs' sync engine, let's decide whether it's Docs' or the company's. If Sheets and Whiteboard will need it within 18 months, the right design is a platform with product-owned op types, and the first customer is Docs."
Appendices
Appendix A: Mechanics in Depth#
A.1 Server-Centric OT (Jupiter Model)#
Why it's right for online-first: the client transforms only against the server's stream, and the server transforms only against its own history. That reduces N-way concurrency to repeated 2-party transforms and avoids TP2 entirely.
server.receive(op, base_rev, client_id, client_seq):
if dedupe.contains(client_id, client_seq): return ack(existing_rev)
for r in (base_rev+1 .. head):
op = transform(op, log[r]) # op' against already-sequenced op
authorize(op, current_acl)
rev = head + 1
log.append(doc_id, rev, op, fencing_token) # conditional write
head = rev
ack(client_id, client_seq, rev)
broadcast(rev, op)
client:
pending = [] # sent, not acked
buffer = [] # not yet sent (one in-flight batch at a time)
on remote(rev, op):
for p in pending + buffer: (op, p) = transform_pair(op, p)
apply(op)
Why it's wrong for offline-first: a client offline for weeks must transform against tens of thousands of ops, and the server is a hard dependency for any merge.
A.2 Sequence CRDTs#
Each inserted element gets a unique ID (lamport, replica_id) and references its left neighbor (RGA) or both neighbors (YATA). Deletes mark tombstones. Merges are commutative — apply updates in any order, converge. Costs: per-element metadata (compressed via run-length encoding of consecutive inserts by the same replica), tombstone GC that requires knowing every replica has seen the delete (hard with offline peers), and the interleaving anomaly where two users' concurrent paragraphs get interleaved character by character in naive algorithms (addressed by newer designs like Fugue).
A.3 Per-Property LWW (Canvas Model)#
object[id].props[key] = (value, server_seq)
on op(id, key, value): if accepted by server, assign server_seq, broadcast
ordering children: fractional index strings between neighbors ('a0' < 'a0V' < 'a1')
Right for structured objects; wrong for text runs, where it would drop concurrent characters.
A.4 Undo in a Collaborative Doc#
Undo must be local: undo my last op, not the doc's last op. Implement as generating the inverse of my op, transformed against every op sequenced since. If my inserted text was since deleted by someone else, the inverse becomes a no-op. Never "roll back the log."
Appendix B: Data Model#
| Table | Key | Fields | Notes |
|---|---|---|---|
docs | doc_id | owner, tenant, home_region, cell, head_rev, snapshot_rev, fencing_token, retention_policy | Strongly consistent; small |
ops | (doc_id, rev) | author, client_id, client_seq, op_blob, ts | Append-only; conditional write |
snapshots | (doc_id, rev) | blob_uri, hash, verified_at | Blob in object storage |
acl | (doc_id, principal) | role, granted_by, ts | Revocations pushed to owner |
named_versions | (doc_id, name) | rev | Exempt from compaction |
Appendix C: Coordination Mechanisms — Quick Comparison#
| Mechanism | Ordering | Failover | Split-brain Safety | Best For |
|---|---|---|---|---|
| Lease + fencing + owner | In-memory, fast | 5–12s | Yes (fencing) | Default for busy docs |
| DB conditional write only | Storage CAS | Instant | Yes | Low-concurrency docs |
| Per-doc log partition | Log order | Partition leader election | Yes | Event-sourcing shops already on Kafka |
| CRDT, no sequencer | None needed | N/A | N/A | Offline-first / P2P |
Appendix D: Client Behavior#
- One batch in flight: client sends at most one unacked batch; accumulates new edits in a buffer (reduces transform work and gives natural batching of ~50–100ms).
- Reconnect: full-jitter exponential backoff, base 500ms, cap 30s; send
last_acked_rev; resend pending with originalclient_seq. - Crash safety: pending queue persisted to IndexedDB/local storage every 1s; survives tab close.
- UI states: "Saved" only after ack; "Saving…" while pending; "Offline — changes saved on this device" when disconnected; never "Saved" before ack.
Appendix E: Observability#
E.1 Core Metrics#
op_ack_latency_ms{p50,p99} # submit → ack
remote_op_latency_ms{p50,p99} # author submit → other client apply
oplog_append_ms{p99}
session_failover_total
oplog_append_rejected_stale_token_total
doc_divergence_detected_total # client hash mismatch
acked_ops_missing_from_log_total # auditor
session_subscribers{doc_id} top-K
presence_msgs_dropped_total
acl_revocation_lag_seconds{p99}
E.2 Critical Alerts#
| Alert | Threshold | Action |
|---|---|---|
| Divergence | >0 for 5 min | Page editor team |
| Acked-op loss | >0 | SEV1; freeze collab deploys |
| Ack latency | p99 >1s for 5 min | Page collab infra |
| Failovers | >3× baseline | Page collab infra |
| Revocation lag | p99 >30s | Page identity + collab |
E.3 Control Plane vs Data Plane#
Data plane: session owners, op log, relays — must keep working if the control plane (placement, leases for new docs, admin) is degraded. Existing owners continue serving with their current lease; new doc placements queue.
Appendix F: Scale Evolution#
| Scale | What Works | What Changes |
|---|---|---|
| 1K concurrent sessions | One owner process, Postgres op table | Nothing clever needed |
| 100K sessions | Owner fleet, lease service, managed KV log, snapshots | Router, fencing, compactor |
| 1M+ sessions | Cells, relay tier, lazy sessions, history tiering | Placement service, divergence auditing |
| Multi-region | Home regions, edge proxy, async DR replication | Residency, home migration |
What You Don't Build on Day One#
- Multi-region homing — single region with DR backups until customers demand it
- Custom CRDT or OT library — adopt one or keep op types minimal
- Viewer relay tier — until a doc exceeds ~500 subscribers
- Keystroke-level history for 7 years — start with 90-day hot history
Appendix G: Multi-Tenancy and Cost#
- Noisy tenant: one tenant's all-hands doc should not degrade others — per-tenant cells for the largest 1% of tenants.
- Quotas: per-tenant limits on concurrent sessions and import rate.
- Chargeback: extended retention and dedicated residency billed as plan tiers so cost is priced where it's decided.