Hiring BarSupport

Design Google Docs (Collaborative Editing) — Staff-Level Case Study

Case study70 min read7 diagrams

Technologies referenced in this case study: Redis · PostgreSQL · Kafka · ZooKeeper & etcd · DynamoDB

Related case studies: Real-time WebSockets · File Sync · Chat & Messaging · Data Consistency · Blob Storage

Related patterns: Real-time Updates · Contention · Distributed Coordination · Consistency Models

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once, then return to the sections that match your weak spots.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Sections 3–4 → the Deep Dives that scare you
Deep Dive3+ hrsEverything, including the Principal Lens (Section 11) and Appendices
What is Collaborative Editing? — Why interviewers pick this topic

Collaborative editing lets multiple people modify the same document at the same time and see each other's changes within a few hundred milliseconds — Google Docs, Figma, Notion, Microsoft Loop, VS Code Live Share. The hard part is not the text box. It is that two people change the same bytes concurrently, over networks with 50–500ms of latency, sometimes offline for hours, and everybody must converge on the same document without anyone losing work.

Before vs After — the "lost paragraph" incident:

Without a convergence protocol (naive last-write-wins on the whole doc):
t=0:      Alice and Bob open the Q3 planning doc (v41)
t=+5s:    Alice rewrites the intro paragraph
t=+6s:    Bob fixes a typo in section 4
t=+7s:    Alice's client PUTs full doc v42
t=+7.2s:  Bob's client PUTs full doc v42 (based on v41) — overwrites Alice
t=+3min:  Alice notices her intro is gone. No error was ever shown.
t=+1day:  Support ticket: "Docs ate my work." Trust is gone.

With an operation-based protocol (OT or CRDT) + single sequencer per doc:
t=0:      Both open v41; both subscribe to the doc session
t=+5s:    Alice's ops stream: insert(pos 120, "...") at rev 41
t=+6s:    Bob's op: delete(pos 2040, 1), insert(pos 2040, "e") at rev 41
t=+6.1s:  Server orders Alice's op as rev 42, transforms Bob's op against it → rev 43
t=+6.2s:  Both clients converge on rev 43. Both edits survive. Nobody noticed.

Why interviewers reach for this question: It forces you to reason about concurrency semantics (what does "both edits survive" mean?), stateful real-time infrastructure (sticky sessions, per-document ownership), durability (an op acknowledged is an op that must never be lost), and product intent (a legal contract editor and a whiteboard have opposite correctness bars). Every one of those is a Staff-level judgment call, and none of them is "pick OT or CRDT."

Mechanics Refresher: Convergence Strategies
StrategyHow It WorksProsCons
Lock / check-outOne editor holds a lock per doc or sectionTrivially correctNot collaborative; lock leaks; "who has it?" tickets
Last-writer-wins (whole doc)Latest full save winsSimpleSilently destroys concurrent work
Three-way merge (diff3)Merge on save against common ancestorWorks for code / asyncConflicts surface to users; no real-time
Operational Transformation (OT)Ops are transformed against concurrent ops before applying; server assigns total orderCompact ops; intent-preserving for text; proven at Google Docs scaleTransform functions are hard to get right; needs a central sequencer in practice
CRDT (sequence CRDTs: RGA, YATA/Yjs, Fugue)Every character has a unique, ordered ID; merges are commutative by constructionPeer-to-peer and offline-friendly; no central transformMetadata overhead; tombstones; harder to enforce server-side invariants
Server-authoritative LWW per property (Figma-style)Per-object, per-property last-writer-wins, ordered by serverSimple, fast for structured docsWrong for rich text runs; needs a separate text strategy

For most production systems: a central, per-document sequencer with an op log — whether the op format is OT or a CRDT. The sequencer gives you total order, a durable ack point, permission enforcement, and a place to snapshot. The OT-vs-CRDT argument is about the op algebra; the sequencer is about operations, and that is where the interview is won.


Executive Summary

If you only read one section, read this. Everything else in the case study elaborates the contrast below.

What This Interview Actually Tests#

Collaborative editing is not an OT-vs-CRDT question. Candidates who spend 15 minutes on transform functions have told the interviewer they studied a paper, not that they can run a system.

It is a stateful ownership question that tests:

  • Whether you pin down the document model and intent before choosing an algorithm
  • Whether you can own a stateful, sticky, per-document hot path without making it a single point of failure
  • Whether you define exactly when an edit is "safe" (the ack point) and who is paged when that promise breaks
  • Whether you treat offline, permissions, and history as first-class — not bolt-ons

The key insight: Every collaborative editor is a replicated log with a UI on top. Once you say "one sequencer per document, ops appended to a durable log, snapshots for fast load, clients as optimistic replicas," the remaining 35 minutes are about what happens when that sequencer moves, lags, or lies.

The L5 vs L6 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"We'll use OT like Google Docs"Asks what the document is (rich text? structured canvas? code?) and whether offline is a requirement — these pick the algorithmAsks how many products in the org need real-time collaboration and whether this is a shared sync platform or a one-off
Consistency"Eventually consistent via OT"Defines the ack point: an op is acknowledged only after it is durably appended to the doc's log; convergence is guaranteed, intent preservation is best-effortSets an org-wide durability SLO ("0 acknowledged edits lost per quarter") and funds the audit that proves it
Scale"Shard by doc ID"Names the hot doc: a 1,000-viewer all-hands doc is a fan-out problem, a 50-editor doc is a sequencing problem — different fixesPrices the long tail: 99% of docs are cold; storage tiering of op logs saves more money than any hot-path optimization
Failure"Add replicas for the WebSocket servers"Designs session handoff: doc owner lease in etcd, fencing tokens, clients reconnect with last-acked rev and resend unacked opsDesigns the org's failure posture: cell-based doc placement so one bad deploy touches ≤5% of docs; game days for session-server loss
Ownership"The docs team owns it"Splits ownership: collaboration infra (sessions, log, snapshots) vs editor product (schema, UX) vs storage platformDecides whether sync becomes a platform with a contract (op log + presence + permissions) that Docs, Sheets, Slides and Comments all consume
Why "first move" separates levels

L5: Reaches for OT because Google Docs uses it. This is a correct answer to the wrong question — it skips whether the document is linear text, a tree (rich text with tables and lists), or a graph of objects (a design canvas). Those choices change what a "concurrent edit" even means.

L6: "Before I pick an algorithm, I need to know three things: what's the document model, is offline editing a hard requirement, and do we need server-side validation of every edit — for example, permissions on individual sections or schema invariants? If it's rich text, online-first, with server-enforced permissions, I'll use a central sequencer with OT-style ops. If offline-first is the core product promise, I'd pick a CRDT and accept the metadata overhead."

L7: Recognizes that the company likely has three teams each about to build their own sync engine, and that the one-way door is the op log format and the client SDK contract, not the algorithm.

Why "consistency" separates levels

L5: "OT guarantees convergence." True, and irrelevant to the user whose edit vanished because a session server crashed before persisting it.

L6: Separates three properties and assigns each an owner: convergence (all replicas reach the same state — guaranteed by the algorithm), durability (an acked op survives any single failure — guaranteed by the append-before-ack rule), and intent preservation (the merged result is what both users meant — best-effort, and product signs off on the edge cases like concurrent delete-and-format of the same word).

L7: Turns durability into a measured SLO with an independent auditor: a background job replays op logs against snapshots and pages if any acked revision is missing.

Why "failure" separates levels

L5: Treats the collaboration server as stateless and adds replicas behind a load balancer. But a doc session is stateful: two servers both believing they own doc X will both assign rev 1,042, and clients will diverge permanently.

L6: Makes single ownership explicit: a lease in etcd/ZooKeeper per doc with a fencing token; the op log rejects appends with a stale token. Failover is a lease expiry (5–10s) plus client reconnect with last_acked_rev. "During failover, users keep typing locally; their ops queue client-side and flush on reconnect. The user sees a 'reconnecting' pill, never lost text."

The Staff Positions#

PositionRationale
One sequencer per document, not per shard of textTotal order per doc is cheap (a doc rarely exceeds ~200 ops/sec) and makes permissions, history, and snapshots trivial
Ack only after durable appendThe only promise users care about is "my edit is saved." Optimistic local apply gives the speed; durable ack gives the truth
OT-style ops with server authority for online-first rich text; CRDT for offline-first or P2PAlgorithm follows intent. Don't pay CRDT metadata tax if you already need a server
Presence is ephemeral and lossy — never on the durable pathCursors at 10–20 updates/sec per user would be 10× the op volume; drop them freely under load
Snapshot every ~500 ops or 60s of activityBounds cold-load replay to <100ms; op log remains the source of truth for history
Permissions checked at session join and on every op, cached with a revocation pushRevoking access must take effect in seconds, not at next page load
Offline is a product decision with an explicit horizon"Offline edits merge cleanly for up to 30 days; beyond that we fork" — say the number

The Three Intents#

Three intents drive every design decision. Each leads to a different architecture.

IntentConstraintStrategyFailure ModeCorrectness Bar
Real-time co-authoring (Docs, Sheets)Online-first, 2–100 concurrent editors, sub-200ms remote visibilityCentral per-doc sequencer, OT-style ops, op log + snapshotsSequencer failover causes brief freeze; divergence if ownership is splitZero acked-op loss; convergence always; intent best-effort
Offline-first / local-first (notes, field apps)Edits for hours/days disconnected; merge on reconnectCRDT (Yjs/Automerge-class), server as relay + archival peerMetadata bloat; surprising merges after long divergenceConvergence always; merges must never lose text, may interleave oddly
Structured canvas (Figma-style design, whiteboards)Objects with properties, not character sequences; 60fps interactionsServer-authoritative per-property LWW, fractional indexing for orderingLast-writer clobbers a concurrent property changePer-property atomicity; conflicts rare and acceptable

🎯 Staff Move: "I'll design for real-time co-authoring of rich text — Google Docs — because that's where concurrency semantics and durability both bite. I'll keep offline as a bounded feature, not the core promise, which lets me keep a central sequencer. If the product were offline-first, I'd switch to a CRDT and I'll call out where the design changes."

The Five Fault Lines#

#Fault LineThe Tension
1OT vs CRDTCentral transform with compact ops vs decentralized merge with metadata overhead — who pays the complexity?
2Stateful Session Ownership vs Stateless Scale-outOne owner per doc gives order; any-server-serves-any-doc gives simple ops — you cannot have both
3Latency vs Durability of AcksAck on receipt (fast, lossy on crash) vs ack after replicated append (safe, +5–20ms)
4Offline Freedom vs Merge SanityLonger offline windows mean bigger, weirder merges and permission drift
5Storage Cost vs History FidelityKeep every keystroke forever (version history, audit) vs compact aggressively (cheap, lossy history)

In the Wild: Real Production Systems#

Why this section belongs here: Naming real systems — and what they chose — shows you've studied operational reality. Use these as anchors, not trivia.

Google Docs — Server-Centric Operational Transformation#

Google Docs descends from the Jupiter collaboration model (Xerox PARC, 1995) that Google Wave also used: clients send operations tagged with the server revision they were based on, a server transforms them against anything that landed since, assigns the next revision, and broadcasts. Clients only ever transform against one server stream, which collapses the notoriously hard N-way OT problem into a 2-party problem. Google publicly documents that up to 100 people can have a doc open for editing at once, beyond which additional users get view-only access — a product-level concurrency cap.

Staff insight: The 100-editor cap is not a technical embarrassment; it is a Staff decision made visible. Bounding per-doc concurrency bounds sequencer load, transform cost, and presence fan-out. Name the cap in your design.

Figma — Server-Authoritative, CRDT-Inspired Multiplayer#

Figma's engineering blog describes a multiplayer system where each open file is loaded into a dedicated server process that is the authority for that document. Clients apply changes optimistically; the server resolves conflicts with last-writer-wins at the granularity of individual object properties, and uses fractional indexing to order children without reindexing siblings. They explicitly chose not to use a pure CRDT because a central server made the system simpler and let them avoid CRDT overhead.

Staff insight: "Inspired by CRDTs, but with a server" is the most defensible production answer for structured documents. It shows you know the theory and chose operational simplicity on purpose.

Yjs / Automerge — Local-First CRDTs#

Yjs (YATA algorithm) and Automerge are open-source CRDT libraries used by a large number of editors (for example, many ProseMirror/TipTap-based products integrate Yjs). They give convergence without a central transform: any peer can merge any other peer's updates in any order. The price is per-character identity metadata and tombstones, which both projects have invested heavily in compressing.

Staff insight: CRDTs move complexity from the server into the data format. That is the right trade when offline or peer-to-peer is the product; it is the wrong trade when you already need a server to enforce permissions and validate edits.

What Interviewers Probe#

After You Say...They Will Ask...What They're Evaluating
"We'll use OT""Walk me through transforming two concurrent inserts at the same position."Do you know the tie-break (site ID / server order), or just the acronym?
"We'll use a CRDT""What happens to document size after a year of edits? How do you enforce permissions?"Awareness of tombstones, metadata overhead, and the loss of a central enforcement point
"Sticky sessions per doc""The server holding the doc dies. What does Alice see? Is anything lost?"Ack point, fencing, client resend, reconnect semantics
"We snapshot periodically""How long does a 5-year-old doc take to open?"Snapshot cadence, replay bounds, cold storage tiering
"Users can edit offline""Alice is offline 3 weeks; her access was revoked yesterday. She reconnects."Permission re-check at merge, fork semantics, product sign-off
"Presence via WebSocket""The all-hands doc has 2,000 viewers. What's the fan-out?"Separating viewers from editors, presence sampling, broadcast tiers

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: The client applies its own edits instantly and keeps them in a pending queue. The session router finds the single owner for a document (lease in etcd). The owner transforms each incoming op against ops the client hasn't seen, assigns the next revision, appends it to the replicated op log with its fencing token, and only then acks and broadcasts. Presence travels a separate lossy channel. Snapshots are built off the hot path by a compactor so opening a doc replays at most a few hundred ops.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Algorithm"OT, like Google Docs""Algorithm follows intent. Online rich text with server permissions: central sequencer with OT-style ops. Offline-first: CRDT."
Ordering"Timestamps""Server-assigned revision per doc from a single fenced owner. Wall clocks never order edits."
Durability"Save to DB periodically""Ack after replicated append to the op log. Snapshots are an optimization, never the source of truth."
Scaling"Shard WebSocket servers""Shard docs across session servers by consistent hashing with leases. Scale viewers separately from editors."
Failure"Replicas""Lease expiry 10s, fencing token on append, client resends unacked ops idempotently by (client_id, seq)."
Offline"Queue and sync later""Bounded offline window, rebase on reconnect, permission re-check at merge, fork if the gap is too large."
Presence"Send cursor on every keystroke""Throttle to 5–10 Hz, coalesce, drop under load, never persist."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Google Docs concurrent editor cap100 editors (more become viewers)Public, product-level bound on per-doc concurrency
Google Docs max document size~1.02M charactersBounds snapshot size and cold-load cost
Active typist op rate3–8 ops/sec (keystrokes, batched to ~2–5 msgs/sec)50 editors × 5 = 250 ops/sec is the worst-case sequencer load for one doc
Remote edit visibility target<200ms p50, <500ms p99 same-regionBeyond ~500ms, users start typing over each other
Durable append latency (replicated, same region)3–15msThe cost of "ack after append" — worth it
Presence update rate5–10 Hz per user, coalesced10× op volume if unthrottled
Snapshot cadenceevery ~500 ops or 60s idleKeeps cold-load replay <100ms
Ownership lease TTL10s, renew every 3sFailover time floor; shorter causes flapping
WebSocket connections per session server50K–100K idle, ~10K active docsMemory ~30–60KB per connection with buffers
Naive CRDT metadata overhead2–10× raw text sizeCompressed encodings (Yjs-style) bring it down substantially; still non-zero
Share of docs active in last 30 daystypically <10%Justifies aggressive cold-tiering of op logs

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes deriving OT transform functions on the whiteboard and never reach durability, failover, or permissions. The phases below get the skeleton on the board in ~10 minutes and spend the rest on the decisions that set your level.


Phase 1: Requirements & Framing (2–3 minutes)#

State functional requirements in one breath:

"Multiple users open the same document, see each other's edits and cursors in near real time, and nobody's acknowledged edit is ever lost. Plus version history, sharing permissions, and some offline tolerance."

Then spend your time on the three questions that pick the architecture:

"Three things change the design. First, the document model — is it rich text, a spreadsheet grid, or a canvas of objects? Second, is offline a core promise or a degraded mode? Third, do we need the server to validate every edit — permissions, schema, compliance holds? I'll assume rich text like Google Docs, online-first with offline as a bounded feature, and server-enforced permissions."

Commit to non-functional targets out loud:

"Remote edits visible in under 200ms p50 within a region. Local edits apply in under 16ms — one frame — regardless of network. Zero loss of acknowledged edits. Up to ~100 concurrent editors per doc, unbounded viewers via a separate path. Opening a doc under 1 second p95."

🎯 Staff Move: The phrase "zero loss of acknowledged edits" reframes the whole interview. It forces you (and the interviewer) to define the ack point, which is where durability, failover, and client retry design all live.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Document: doc_id, owner, ACL, current rev, latest snapshot_rev, cell/shard assignment
  • Operation: (doc_id, base_rev, client_id, client_seq, ops[]) — ops are retain(n) / insert(text, attrs) / delete(n) over a linearized document
  • Revision: server-assigned monotonically increasing integer per doc; (doc_id, rev) is the primary key of the op log
  • Snapshot: full document state at rev, stored as a blob
  • Session: the set of clients currently connected to a doc, owned by one session server

Real-time channel (WebSocket, per doc):

→ JOIN      { doc_id, last_known_rev, auth }
← SYNC      { snapshot_rev, snapshot?, ops[last_known_rev+1 .. head] }
→ SUBMIT    { base_rev, client_id, client_seq, ops[] }
← ACK       { client_seq, rev }
← REMOTE_OP { rev, author, ops[] }
↔ PRESENCE  { cursor, selection, color }        (lossy, throttled)

Cold path (REST):

GET  /docs/{id}/revisions?from=&to=        version history
POST /docs/{id}/restore  { rev }           restore = new ops, never a rewrite
PUT  /docs/{id}/acl      { principal, role }

🎯 Staff Move: "(client_id, client_seq) is the idempotency key. When a client reconnects and resends ops that were actually applied before the crash, the server dedupes them instead of inserting the paragraph twice. Restore is modeled as new forward ops — history is append-only, so audit never breaks."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw ≤8 boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk the edit flow in 60 seconds:

  1. Client applies the edit locally (instant), puts it in a pending queue, sends SUBMIT with base_rev.
  2. Gateway routes to the doc's owner — looked up from a lease registry.
  3. Owner transforms the op against every op in (base_rev, head], checks permission, assigns rev = head+1.
  4. Owner appends to the op log with its fencing token. Only after the append succeeds, it acks the author and broadcasts to other clients.
  5. Other clients transform the remote op against their own pending ops and apply.
  6. A compactor periodically folds ops into a snapshot for fast loads.

🎯 Staff Move: After drawing, say: "This works for a single region with docs up to 100 editors. The interesting parts are what happens when the owner dies mid-append, what happens when a client comes back after a week offline, and how 2,000 viewers on one doc don't melt the owner. Which should I go deep on?" You're at ~9 minutes.


Phase 4: Transition to Depth (1 minute)#

"The skeleton is standard. Three places decide whether this system is trustworthy: session ownership and failover — because split-brain means permanent divergence; the ack point and durability — because users forgive latency but not lost work; and offline plus permissions — because that's where product and security disagree. I'd start with ownership and failover unless you'd prefer otherwise."

If the interviewer asks "OT or CRDT?" first, answer in 90 seconds and bridge back: "Given online-first with server-side permissions, OT-style with a central sequencer. The algorithm matters less than the sequencer's failure semantics — let me show you why."


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → pick a position → quantify → name who pays.

Deep dive A: Session ownership and failover (7–8 min)

"Each doc has exactly one owner at a time. Ownership is a lease in etcd with a 10s TTL, renewed every 3s. Every lease grant carries a monotonically increasing fencing token. The op log accepts an append only if the token is ≥ the highest it has seen for that doc — so a zombie owner that was paused by GC for 15 seconds cannot write rev 1,042 after the new owner already did."

Failure sequence:

  1. Owner crashes at t=0. Clients' WebSockets drop; they show "Reconnecting…" and keep editing locally.
  2. Lease expires at t≤10s. Router assigns a new owner, which loads the latest snapshot plus tail ops from the log (~50–200ms).
  3. Clients reconnect with last_acked_rev and resend every unacked op with its (client_id, client_seq).
  4. New owner dedupes ops already in the log, transforms the rest, appends, acks.
  5. User-visible impact: a 5–12s "reconnecting" pill; no lost text.

"The victim of a 10s lease is freshness during failover — remote edits freeze for up to 10s. The victim of a 2s lease is stability — GC pauses and network blips cause ownership flapping and needless reloads. I pick 10s and say so."

Deep dive B: The ack point and durability (6–7 min)

"Three places I could ack: on receipt in memory (~1ms, loses edits on crash), after local disk fsync (~2–5ms, loses edits on machine loss), or after a replicated append — quorum of 3 across zones (~5–15ms). I ack after replicated append. 10ms is invisible because the local edit already rendered; the ack only clears the pending queue."

What the op log is: "Could be a partitioned log like Kafka keyed by doc_id, or a table in a strongly consistent KV — DynamoDB or Spanner-class — with (doc_id, rev) as the key and a conditional write attribute_not_exists(rev) AND token >= last_token. I prefer the KV: I need per-doc random reads for history and a conditional write for fencing, which a log doesn't give me cleanly."

Deep dive C: Offline and permissions (6–7 min)

"Offline edits are a queue of ops based on an old revision. On reconnect, the server transforms them against everything since — which for a 3-week-old base could be 50,000 ops. That is expensive and produces weird merges. So: offline edits merge automatically up to a bound — say 10,000 ops of divergence or 30 days. Beyond that, we create a 'your offline copy' fork and show a compare view. Before any merge, we re-check permission at the current ACL, not the ACL at edit time. If access was revoked, the edits are held and the user is told — never silently applied."

🎯 Staff Move: Every deep dive ends with an owner: "Collab infra owns lease tuning and the durability SLO. The editor team owns transform correctness and has a fuzz-test suite that runs 1M random concurrent op pairs per build. Product owns the offline horizon."


Phase 6: Wrap-Up (2–3 minutes)#

"To summarize: a single fenced owner per doc, ops acked after replicated append, snapshots for fast load, presence on a lossy side channel, and bounded offline with permission re-check at merge. What I'd build next: first, a divergence detector — clients periodically send a hash of their doc at a rev, and mismatches page us. Second, cell-based placement so a bad deploy touches a bounded fraction of docs. Third, cold-tiering op logs older than 90 days to object storage — most of our storage bill is history nobody opens."


Common Timing Mistakes#

MistakeTime LostFix
Deriving OT transform matrices on the board10–15 minState the tie-break rule in one sentence; offer the full function if asked
Designing the rich-text schema (bold, lists, tables)5–8 min"Linearize to a sequence with attributes; tables are nested sequences" and move on
Debating WebSocket vs SSE vs long-poll3–5 min"WebSocket, with long-poll fallback for hostile proxies." Done
Ignoring viewers vs editorsdiscovered at minute 40Say it in Phase 3: "viewers get a read-only fan-out tier"
Never stating the ack pointfatalSay "ack after durable append" in Phase 1 or 2

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Most system design problems let you hide state behind a database and scale stateless servers. Collaborative editing doesn't. The hot path is stateful by necessity — someone must hold the current document and order concurrent edits — and that state lives in memory on one machine at a time. The candidate has to design for a component that cannot simply be load-balanced, and has to reason about what "correct" means when two humans disagree at the same character.

It is also a product-judgment problem in disguise. "Alice deletes a sentence while Bob bolds a word inside it" has no algorithmically correct answer. The Staff candidate says who decides (product), what the default is (delete wins; Bob's formatting is dropped), and how it's tested.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

The L5 path is technically correct at every step. It fails because it spends its budget on the part of the problem that libraries already solve.

1.3 The Staff Question That Cuts Through Everything#

"When exactly is an edit safe, and who finds out if that promise is broken?"

This single question forces every important decision: the ack point (durable append), the ownership model (fenced single writer), the client retry protocol (idempotent resends), the detection mechanism (divergence hashes, replay audits), and the owner (collab infra on-call). A candidate who asks it in the first five minutes is already interviewing at Staff.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Real-time co-authoring (default). Think Google Docs, Sheets, Slides, Microsoft Word online. Users are almost always online; concurrent editing is the headline feature; the company must enforce sharing permissions, retention policies, and legal holds server-side. The server is already mandatory, so a central sequencer is free — and it buys total order, a single enforcement point, and linear history. OT-style ops fit naturally. Victims of this choice: users on flaky networks (they see "reconnecting" more often), and the collab infra team (they own a stateful fleet).

Offline-first / local-first. Think field-service apps, note-taking apps that must work on a plane, or peer-to-peer editors. Here the network is optional, so correctness cannot depend on a server. CRDTs are the right tool: any replica merges any other in any order. Victims: storage (metadata and tombstones), the security team (enforcing per-edit permissions on data that merged peer-to-peer is hard), and users who get "valid but weird" merges after long divergence.

Structured canvas. Think Figma, Miro, a diagramming tool. The document is a tree of objects with properties; a conflict means two people set fill on the same rectangle. Per-property last-writer-wins, ordered by the server, is simple and good enough — users rarely edit the same property of the same object in the same second, and when they do, "last one wins" matches their mental model. Text boxes inside the canvas still need a sequence strategy. Victims: the occasional user whose property change is overwritten, which product accepts.

2.2 When NOT to Use Real-Time Collaborative Editing#

SituationBetter ChoiceWhy
Documents edited by one person 99% of the timeAutosave + optimistic concurrency (If-Match: etag) + conflict dialogPaying for sessions, presence, and op logs for a solo editor is waste
Source code with review workflowsGit-style branches and 3-way mergeHumans want to review merges; real-time interleaving of code is harmful
Regulated records (contracts at signature, medical charts)Check-out/check-in with explicit locking and auditLegal wants a single accountable author per version, not a merge
Large binary assets (video, CAD meshes)Locking + blob storage with versioningNo meaningful op algebra for binary diffs
Form-like records (CRM fields)Per-field optimistic concurrencyRow-level conflicts are rare; field-level LWW with audit suffices

🎯 Staff Move: "If our telemetry says 95% of docs never have two simultaneous editors, the collaborative engine should be lazy: a doc only gets a session server when a second editor joins. Solo edits go through a cheap autosave path." This is how you cut session fleet cost by an order of magnitude.

2.3 What the Interviewer Leaves Underspecified#

Unstated AssumptionWhy It MattersWhat to Say
Document modelText vs tree vs object graph changes the op algebra"I'll linearize rich text into a sequence with attributes"
Max concurrent editorsDrives sequencer and presence design"Cap at 100 editors, viewers unbounded via a separate tier"
Offline durationDrives merge cost and fork policy"Up to 30 days or 10K ops divergence, then fork"
Permission granularityDoc-level vs section-level changes the op validator"Doc-level roles: owner, editor, commenter, viewer"
History retentionStorage cost grows with every keystroke"Full op history 90 days hot, then snapshots-only at named versions"
Multi-regionWhere does the sequencer live for a doc with editors on 3 continents?"Doc homed in one region; move the home if the editor population shifts"
Comments/suggestionsAnchoring to text that moves"Comments anchor to op-transformed positions, same machinery as cursors"

2.4 Precise Terminology#

TermMeaningCommon Confusion
ConvergenceAll replicas that have seen the same set of ops reach identical stateNot the same as "correct" — two replicas can converge on garbage
Intent preservationMerged result reflects what each user meantBest-effort; no algorithm guarantees it for all cases
Causality preservationIf op B was generated after seeing op A, every replica applies A before BViolating it causes "edit appears before its cause"
Transform (OT)T(a, b) → a' such that applying b then a' equals applying a then T(b, a) (TP1)TP2 (needed for P2P OT) is where most published algorithms were later shown to be wrong
Sequence CRDTEach element has a globally unique, totally ordered ID; inserts reference neighbors by ID"CRDT" alone includes counters and sets; text needs a sequence CRDT
TombstoneDeleted element kept as a marker so concurrent references still resolveThe reason CRDT docs grow even when text shrinks
Base revisionThe server revision the client's op was generated againstTransform window = head − base_rev
Fencing tokenMonotonic number attached to a lease; storage rejects writes with lower tokensLeases without fencing do not prevent split-brain
SnapshotMaterialized doc state at a revisionCache of the log, not the source of truth
PresenceEphemeral cursor, selection, "who's here" stateNever durable, never ordered with ops

3. The Five Fault Lines#

Each fault line: the options, who pays, the Staff default, and when to deviate.

3.1 Fault Line 1: OT vs CRDT#

StrategyWhat WorksWhat BreaksWho Pays
Server-centric OT (Jupiter-style)Compact ops (~20–50 bytes); linear history; server validates every op; only 2-party transformsRequires a live sequencer; offline merges are expensive transformsCollab infra (stateful fleet); offline users (long rebases)
Peer-to-peer OTNo serverTP2 correctness is notoriously hard; many published algorithms had counterexamplesEveryone, eventually — via divergence bugs
Sequence CRDT (Yjs/Automerge/Fugue class)Merge in any order; offline and P2P native; server is optionalPer-char IDs + tombstones; server-side validation is awkward; interleaving anomalies on concurrent insertsStorage and bandwidth; security team (enforcement)
Hybrid: CRDT data model + central server orderingCRDT convergence plus a server for auth, durability, and snapshotsPays both metadata and server costBudget; complexity

The Staff default: server-centric OT-style ops for online-first rich text. The server is required anyway for permissions, retention, and search indexing, so the key CRDT benefit — no server — is worth nothing here, while its costs (metadata, tombstone GC, enforcement) are real.

When to deviate: choose a CRDT when (a) offline editing for days is the headline promise, (b) you want peer-to-peer or end-to-end encrypted collaboration where the server cannot read ops, or (c) you're a small team that wants to buy a mature library (Yjs) instead of writing transform functions. Point (c) is underrated: "A CRDT library we don't have to maintain beats an OT implementation we'd get subtly wrong."

🎯 Staff Move: "The algorithm choice is a two-way door if we keep the op log format and client protocol abstract. It becomes a one-way door once we have a billion docs stored in that format. So I'll version the op encoding from day one."

3.2 Fault Line 2: Stateful Session Ownership vs Stateless Scale-out#

StrategyWhat WorksWhat BreaksWho Pays
Single owner per doc (lease + fencing)Total order in memory; transform against local state; cheapOwner failure = brief freeze; hot docs pin one machineCollab infra on-call (failover tuning)
Stateless servers + ordered log (e.g., per-doc Kafka partition key)Any server can accept; log provides orderTransform needs current state → every server reloads; latency of log round-trip on every opLatency budget (+10–30ms); log ops team
Database-as-sequencer (conditional write on rev)No leases; storage enforces orderContention: concurrent submitters retry CAS; ~5% conflict rate ceiling before retries dominateUsers on busy docs (retry latency)

The Staff default: single owner per doc with a lease and fencing token, backed by a conditional-write op log. The log's conditional write is the backstop that makes split-brain harmless: even if two servers believe they own the doc, only one can append rev N.

When to deviate: for very low-concurrency products (≤3 editors typical), database-as-sequencer is simpler — no lease service, no router. Contention stays below the ~5% CAS-conflict threshold, beyond which retries dominate. See Contention.

Diagram: 3.2 Fault Line 2: Stateful Session Ownership vs Stateless Scale-out

3.3 Fault Line 3: Latency vs Durability of Acks#

Ack PointLatency AddedLoss WindowWho Pays
On receipt (memory)~1msEverything since last flush on owner crashUsers — "Docs ate my work"
After local fsync2–5msEverything on machine/disk lossUsers on rare hardware failures
After replicated append (quorum, cross-zone)5–15msNone for single failuresLatency budget — invisible thanks to optimistic local apply
After cross-region replication60–150msNone for region lossRemote-edit latency visibly worse

The Staff default: ack after quorum append within the doc's home region. Region loss is handled by asynchronous cross-region replication with an RPO of a few seconds — and product signs off that a full region loss may lose the last ~5s of edits.

When to deviate: regulated documents (legal, financial filings) may justify synchronous cross-region acks for a specific doc class. Price it: 100ms+ added to ack latency, only on those docs.

🎯 Staff Move: "The user never waits for the ack to see their own edit — optimistic local apply takes care of perceived latency. The ack only controls when the edit leaves the pending queue. That's why I can afford a replicated append: 10ms here costs nothing the user can see."

3.4 Fault Line 4: Offline Freedom vs Merge Sanity#

PolicyWhat WorksWhat BreaksWho Pays
No offline editingSimple; no divergent mergesFlights, trains, flaky networksMobile/travel users
Unbounded offline with auto-mergeMaximum freedom50K-op rebases, interleaved paragraphs, edits by revoked usersOther collaborators; security
Bounded offline (time/op-count) + fork beyondPredictable merges; explicit fallbackOccasional "your offline copy" forkUsers past the bound (they resolve manually)

The Staff default: bounded offline — 30 days or 10,000 ops of server-side divergence, whichever first — with permission re-check at merge against the current ACL. Beyond the bound, create a sibling doc and show a diff.

When to deviate: local-first products where offline is the promise use CRDTs and unbounded merge, and accept interleaving anomalies. Enterprise docs with DLP (data-loss prevention) might disable offline entirely by admin policy.

3.5 Fault Line 5: Storage Cost vs History Fidelity#

Every keystroke is an op. A heavily edited doc generates 50K–500K ops per year. At ~40 bytes per op that's 2–20MB of history per active doc — tiny per doc, enormous at 1B docs.

StrategyWhat WorksWhat BreaksWho Pays
Keep every op forever, hotPerfect history, audit, blameStorage bill grows linearly foreverFinance
Keep ops hot 90 days, then snapshots at named/auto versionsCheap; history still browsable at coarse grainKeystroke-level blame lost after 90 daysUsers who want fine-grained history of old edits (rare)
Compact to snapshots onlyCheapestNo history, no auditCompliance; users

The Staff default: ops hot for 90 days, then compacted into periodic snapshots (e.g., one per editing session) moved to cold object storage. Legal-hold docs are exempted from compaction by policy. See Blob Storage for tiering economics.


4. Failure Modes & Operational Reality#

4.1 Split-Brain Ownership — Permanent Divergence#

Scenario: A network partition isolates session server S1 from etcd but not from clients. S1's lease expires; S2 takes ownership. Without fencing, both accept edits.

t=0:      Partition: S1 cannot reach etcd; S1 still reaches 12 of 30 clients
t=+10s:   Lease expires; router assigns doc to S2; 18 clients reconnect to S2
t=+10s:   S1 (no fencing) keeps accepting edits: rev 5001..5040 on S1, 5001..5063 on S2
t=+2min:  Partition heals. Two histories claim rev 5001–5040. Snapshot compactor picks S2's.
t=+1hr:   12 users report their last two minutes of work vanished.

Detection: oplog_append_rejected_stale_token_total (should be >0 during partitions — that's fencing working); doc_divergence_detected_total from client hash checks; lease_lost_while_serving_total.

Mitigation: Fencing token on every append — S1's writes are rejected at t=+10s; S1 closes its sockets, clients reconnect to S2 and resend unacked ops. Loss becomes zero.

Prevention: Owner self-fences: if it cannot renew its lease within 70% of TTL, it stops acking before the lease expires. Owner: collab infra.

4.2 The Mega-Doc — Hot Document Fan-out#

Scenario: The CEO shares the all-hands notes doc with 8,000 employees 5 minutes before the meeting. 40 people are editing; 8,000 are watching.

t=0:      Link posted in #all-company
t=+30s:   8,000 JOINs on one doc → all routed to one session server
t=+45s:   Owner CPU 100%: each op × 8,000 broadcasts + 8,000 presence streams
t=+60s:   Ack latency p99 climbs from 40ms to 4s; editors see "Saving…" forever
t=+90s:   Owner OOM; failover; 8,000 clients reconnect simultaneously → new owner dies too

Detection: session_subscribers{doc} > 500; op_ack_latency_p99 > 1s; session_failover_total rising for the same doc_id.

Mitigation: Split editors from viewers. The owner serves only editors (capped at 100). Viewers attach to a read-only fan-out tier that subscribes to the owner's op stream once per relay node and rebroadcasts. Presence for viewers is aggregated ("8,012 viewing") rather than per-cursor. Reconnects use jittered backoff (base 1s, max 30s).

Prevention: Admission control at JOIN: beyond 100 editors, new joiners are viewers. Owner: collab infra; product signs off on the cap.

Diagram: 4.2 The Mega-Doc — Hot Document Fan-out

4.3 Transform Bug — Silent Divergence#

Scenario: A new feature adds "suggestion mode" ops. The transform of suggest_delete against a concurrent format has an off-by-one. Clients converge on different documents.

t=0:      Deploy editor v412 with suggestion mode to 10% of clients
t=+2hr:   ~0.3% of docs with concurrent suggest+format diverge
t=+2hr:   Nobody notices: each user sees a coherent doc, just not the same doc
t=+3days: Customer: "My colleague sees a different paragraph than I do"

Detection: Clients send hash(doc_state) with their ack every N revisions; server compares against its own. doc_divergence_detected_total > 0 pages. Fuzz tests in CI: 1M random concurrent op pairs per build, asserting TP1.

Mitigation: Server state is truth. Divergent client discards local state, reloads from snapshot + ops, replays its pending queue. Feature-flag kill switch for the new op type.

Prevention: Every new op type requires transform functions against every existing type (an N² matrix) plus property tests before rollout. Owner: editor team for correctness, collab infra for detection.

4.4 Snapshot Corruption — Poisoned Cold Load#

Scenario: The compactor writes a snapshot at rev 90,000 with a bug that drops table cells. Every subsequent open of the doc loads the bad snapshot.

Detection: Compactor verifies each snapshot by replaying from the previous snapshot and comparing hashes before publishing; snapshot_verify_failed_total. Weekly sampled audit replays 0.1% of docs from rev 0.

Mitigation: Snapshots are cache; delete the bad snapshot and rebuild from the op log. Because the log is the source of truth, this is always recoverable — if you never compacted away the ops the bad snapshot replaced.

Prevention: Never delete ops until the snapshot replacing them has been verified and aged ≥7 days. Owner: collab infra.

4.5 Reconnect Storm After Regional Blip#

Scenario: A 20-second network blip in one zone drops 400K WebSocket connections. All reconnect at once.

t=0:      Zone blip; 400K sockets drop
t=+1s:    400K reconnects hit gateways; TLS handshakes saturate CPU
t=+5s:    Each JOIN triggers SYNC → snapshot reads spike 50×
t=+30s:   Snapshot store throttles; JOINs time out; clients retry → second wave

Detection: ws_connect_rate > 5× baseline; snapshot_read_throttled_total; join_latency_p99.

Mitigation: Jittered exponential backoff in the client (full jitter, base 500ms, cap 30s). Clients that still hold local state send last_known_rev so SYNC returns only tail ops, not a snapshot. Gateway admission control sheds JOINs with 503 + Retry-After.

Prevention: Game day: drop a zone's connections quarterly. Owner: collab infra + edge team. See Real-time WebSockets.

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Owner crashsession_failover_total, lease expiryDocs on that server (~10K), 5–12s freezeLease expiry, client resendCollab infra
Split-brainoplog_append_rejected_stale_token_totalOne doc per partitionFencing tokenCollab infra
Transform bugdoc_divergence_detected_totalDocs using new op typesKill switch, reload from serverEditor team
Mega-docsession_subscribers > 500One doc + co-located docsViewer relay tier, editor capCollab infra
Snapshot corruptionsnapshot_verify_failed_totalDocs compacted by bad buildRebuild from logCollab infra
Op log slowoplog_append_p99 > 50msAll docs in a cellShed presence, batch ops, fail over logStorage platform
Permission propagation lagacl_revocation_lag_p99 > 10sRevoked users keep editingPush revocation to owner; kick sessionIdentity + collab
Reconnect stormws_connect_rate spikeA zone / regionJitter, tail-only SYNC, admission controlEdge + collab

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
AlgorithmExplains OT or CRDT correctlyChooses based on doc model and offline intent; names the costs of the otherTreats op format as a versioned platform contract across products
OrderingServer assigns versionsSingle fenced owner; explains split-brain and fencingCell architecture bounding owner failures to ≤5% of docs
Durability"We save to the database"Ack after replicated append; idempotent resendDurability SLO with independent replay auditor and error budget
ScaleShards by doc IDSeparates editors from viewers; caps editors; lazy sessionsPrices session fleet vs storage; cold-tiers history; knows where the money goes
OfflineQueue and replayBounded window, fork, permission re-checkAdmin policy per tenant; DLP integration; legal sign-off
OperationsMentions monitoringDivergence hashes, fuzz tests, named metrics and ownersGame days, org-level incident taxonomy, collaboration SLOs in exec reviews

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Defines the ack point unprompted"An edit is acknowledged only after quorum append. Local apply gives speed; the ack gives safety."
Uses fencing, not just leases"The lease tells a server it may own the doc. The fencing token makes sure storage agrees."
Separates presence from ops"Cursors are lossy, throttled to 5–10 Hz, and never persisted. I'll drop them first under load."
Bounds the ugly cases"Offline merges are automatic up to 30 days or 10K ops; beyond that we fork and show a diff."
Names product sign-off"Delete-vs-format conflict resolves delete-wins. Product signs off; it's in the test suite."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
15 minutes of transform math, no failure storyOptimizes the solved part; ignores the part that pages people
"Stateless WebSocket servers behind a load balancer"Doesn't recognize the sequencer is inherently stateful
Uses wall-clock timestamps to order editsClock skew of 10–100ms across clients reorders keystrokes
No distinction between viewers and editorsMega-doc will melt the design
"CRDTs solve everything" with no mention of tombstones or enforcementKnowledge without cost-awareness

5.4 Common False Positives#

  • Deep OT theory ≠ collaborative editing design. Knowing TP1/TP2 is nice; knowing why the server-centric model avoids TP2 is the signal.
  • Name-dropping Yjs ≠ understanding CRDT costs. Ask them what happens to a 1M-character doc after 3 years of edits.
  • A beautiful architecture diagram ≠ an owner. If no box has a team name, the design has no pager.
  • "We'll use Spanner" ≠ a durability story. Strong storage doesn't help if the owner acks before writing.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minDoc model, offline intent, server validation, ack promise
Entities & API3–5 minOp shape, rev, idempotency key, WebSocket messages
Architecture5–10 minSequencer, op log, snapshots, presence, leases
Deep dive 110–20 minOwnership + failover (fencing, resend)
Deep dive 220–30 minDurability / snapshots / cold load
Deep dive 330–40 minOffline + permissions or mega-doc fan-out
Wrap-up40–45 minEvolution, what you'd build next, owners

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response
"Now make it work offline for a month."Do you know OT's weakness and have a policy?Bounded rebase, fork beyond bound, or switch to CRDT with stated costs
"Two regions, editors on both coasts."Sequencer placementHome region per doc; migrate home on sustained editor shift; ~70ms cross-country penalty for remote editors
"Add comments anchored to text."Position tracking under concurrent editsAnchors are transformed like cursors; orphaned anchors resolve to nearest surviving char
"Legal needs to see who typed every character."History fidelity vs costOp log with author per op, retained per legal hold policy
"Your transform has a bug in prod."Detection of silent failureClient state hashes, kill switch, server-is-truth reload

6.3 What to Deliberately Skip#

  • Rich-text schema details (bold/italic run merging) — one sentence.
  • The full OT transform table — state insert/insert tie-break and offer more.
  • WebSocket vs SSE debate — pick WebSocket.
  • Spell-check, autocomplete, export to PDF — out of scope unless asked.

6.4 Follow-Up Questions to Expect#

  1. "What exactly happens when two users insert at the same position at the same time?"
  2. "The server holding a doc crashes mid-append. Walk me through the next 15 seconds."
  3. "How do you open a 5-year-old, 1M-character document in under a second?"
  4. "A user's access is revoked while they're editing. How fast does it take effect?"
  5. "How would you detect that two clients are seeing different documents?"
  6. "How does undo work when other people have edited since?"
  7. "What does it cost to store every keystroke for 1B documents?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design Google Docs."

Staff Answer

"Before I draw anything: what's the document model, is offline a core promise, and must the server validate each edit? I'll assume rich text, online-first with bounded offline, and server-enforced doc-level permissions — that's Google Docs. Targets: local edits render within one frame (16ms), remote edits visible within 200ms p50 in-region, up to 100 concurrent editors with unbounded viewers on a separate tier, and zero loss of acknowledged edits. That last one is the promise that shapes everything: an edit is acknowledged only after it's durably appended to the doc's op log. I'll use a single fenced owner per doc as the sequencer, OT-style ops, snapshots for fast load, and presence on a lossy side channel."

Why this is L6:

  • Asks the three questions that actually select the algorithm, then commits
  • States a durability promise with a precise ack point in the first 60 seconds
  • Separates editors from viewers before being prompted

What L7 adds:

  • "Is this a one-off for Docs, or the sync substrate for Docs, Sheets, Slides and Comments? If the latter, the op log and client SDK are a platform contract and I'd design them for four consumers."
  • Names the cost driver up front: history storage, not session compute
❌ Common L5 Trap

"We'll use OT like Google Docs. Clients connect via WebSocket to a server, which transforms ops and saves them to a database."

Why this misses: Correct but unowned. The interviewer asks "what if the server dies after broadcasting but before saving?" and there's no answer, because the ack point was never defined.


Drill 2: Concurrent Inserts at the Same Position#

Prompt: "Alice and Bob both insert at position 10 at the same time, based on rev 50. What happens?"

Staff Answer

Alice's insert(10, "X") arrives first; server assigns rev 51. Bob's insert(10, "Y") arrives with base_rev=50, so the server transforms it against rev 51. Two inserts at the same index need a deterministic tie-break — in the server-centric model the op already sequenced wins the left position, so Bob's op becomes insert(11, "Y") and is assigned rev 52. Result everywhere: ...XY.... On Alice's client, the remote op rev 52 arrives as insert(11,"Y") and applies directly. On Bob's client, rev 51 arrives while his op is pending; the client transforms rev 51 against its pending op — Alice's insert stays at 10, Bob's pending local char shifts to 11. Both converge on XY.

The important part: the tie-break is deterministic and every replica uses the same one. In CRDTs the tie-break is by element ID (e.g., Lamport clock then client ID). The user-visible result may be XY or YX; product doesn't care which, only that it's consistent.

Why this is L6:

  • Shows the 2-party transform on both server and client sides
  • Explains why server-centric OT avoids the harder N-party problem
  • Separates "deterministic" (required) from "which order" (irrelevant)

What L7 adds:

  • Insists on a property-based test harness (random op pairs, assert convergence) as a release gate for the whole org's op types
  • Notes that the transform matrix grows N² with op types — a governance problem for which teams may add ops

Drill 3: The Owner Dies#

Prompt: "The session server for a doc with 30 editors crashes. Walk me through the next 15 seconds."

Staff Answer

t=0: sockets drop. Clients show "Reconnecting…" and keep applying edits locally into their pending queues — no text is lost locally. t=0–10s: the owner's lease in etcd expires (TTL 10s). Clients retry with jittered backoff (0.5–2s). t≈10s: the router assigns the doc to a new server with fencing token N+1. It loads the latest snapshot (e.g., rev 48,000) plus the tail from the op log (e.g., 312 ops) — ~100ms. t≈10–12s: clients reconnect with last_acked_rev and resend all unacked ops with (client_id, client_seq). The new owner dedupes any that actually landed before the crash (the old owner appended but didn't get to ack), transforms the rest, appends, acks. Remote edits resume. Total: ~12s of frozen remote updates, zero lost acked edits, zero lost local edits.

To shorten the freeze I could have clients detect the crash and ask the router to force-expire — but only if the old owner is provably dead; otherwise fencing protects us but we'd flap. I keep 10s.

Why this is L6:

  • Distinguishes local edits (safe in client queue) from acked edits (safe in log)
  • Handles the "appended but not acked" gap via idempotent resend
  • Justifies the lease TTL as a tradeoff, not a default

What L7 adds:

  • Budgets failover into an SLO: "99.9% of doc-minutes have remote-edit latency <1s" — 12s failovers spend the budget, so the owner fleet's crash rate becomes a tracked number
  • Cells: a server crash is expected; a cell-wide crash is a SEV1 with a defined blast radius

Drill 4: Make Durability Concrete#

Prompt: "You say 'ack after durable append.' What's the storage, the key, and the write?"

Staff Answer

Table ops in a strongly consistent, horizontally partitioned KV (DynamoDB, Spanner, or a Raft-replicated store). Partition key doc_id, sort key rev. Item: {author, client_id, client_seq, ops_blob, fencing_token, ts}. The write is conditional:

PutItem ops (doc_id, rev=head+1, ...)
  condition: attribute_not_exists(rev)
  AND doc_meta.fencing_token <= :my_token

In practice I keep doc_meta.{head_rev, fencing_token} in the same partition and use a transaction so rev assignment and token check are atomic. Latency: 5–15ms in-region. Throughput: a doc peaks at ~250 ops/sec (50 heavy typists), so I batch ops arriving within a 10–20ms window into one append — cuts writes 5–10× on busy docs. A second index on (client_id, client_seq) per doc (or a small bounded dedupe window held in the owner) handles resend dedupe.

Why this is L6:

  • Concrete keys, concrete condition, concrete latency
  • Batching quantified against the per-doc write rate
  • Dedupe designed in, not bolted on

What L7 adds:

  • Chooses managed storage deliberately: "Owning a Raft log for a billion docs is two engineers forever. DynamoDB costs more per op but removes that headcount."
  • Defines the storage contract once so other products reuse it

Drill 5: The Mega-Doc#

Prompt: "A doc has 60 editors and 5,000 viewers. What breaks?"

Staff Answer

Fan-out breaks first, not sequencing. 60 editors × ~3 ops/sec = 180 ops/sec is fine for one owner. But broadcasting each op to 5,060 sockets is ~900K messages/sec, plus presence for 5,060 cursors at 5 Hz squared if naïvely broadcast to everyone — impossible. Fix: editors connect to the owner; viewers connect to a relay tier. The owner publishes the op stream once per relay (say 3 relays), each relay fans out to ~1,700 viewers. Viewers don't get individual cursors — they get the 60 editors' cursors at 2 Hz and an aggregate "5,000 viewing" count. Editor count is capped at 100 via admission control; the 101st becomes a viewer with a "request edit" button.

Why this is L6:

  • Identifies fan-out, not sequencing, as the bottleneck with arithmetic
  • Separates editor and viewer paths
  • Presence downsampled for viewers

What L7 adds:

  • The relay tier is the same machinery as live-blog / live-dashboard fan-out elsewhere in the org — build it once
  • Product sign-off on the 100-editor cap as a published limit, like Google Docs does

Drill 6: Offline for Three Weeks#

Prompt: "Alice edits offline for 3 weeks. Meanwhile 40,000 ops landed and her edit access was downgraded to comment-only. She reconnects."

Staff Answer

Step 1: re-authorize against the current ACL — she's now a commenter, so her 600 pending edit ops cannot be applied. Step 2: we don't drop them silently. Her client keeps them and the UI offers: "Your offline changes couldn't be applied because your access changed. Save as a copy or send as suggestions." Suggestions mode is the elegant answer — commenters can suggest. Step 3: even if she were still an editor, 40,000 ops exceeds our 10K-op rebase bound, so we'd fork: create "Alice's offline copy" and show a compare view. Transforming 600 ops against 40,000 is ~24M transform steps — seconds of CPU on the owner, blocking everyone else on that doc.

Why this is L6:

  • Permission re-check at merge time, not edit time
  • Quantified rebase cost justifies the bound
  • Never silently drops user work; gives a recoverable path

What L7 adds:

  • Makes the offline bound a tenant-admin policy (some enterprises want 0, some want 90 days) with legal/security review
  • Rebases run on a separate worker pool so large merges can't starve live sessions — an org-wide principle: batch work never shares a hot path

Drill 7: Build vs Buy the Sync Engine#

Prompt: "Should we build our own OT engine, or use Yjs, or a hosted service?"

Staff Answer

Depends on whether collaboration is our differentiator. If we're a docs company, building gives control over op types, permissions, and storage format — but it's 4–6 engineers for a year to reach production quality plus 2 forever for maintenance, and transform bugs are subtle. If collaboration is a feature of a larger product (a CRM with a notes field), use Yjs with a server-side provider: we get convergence and offline for free and own only persistence, auth, and presence. Hosted services (Liveblocks-style) are fastest to ship; the risk is data residency and lock-in on the doc format. My default for a non-docs company: Yjs + our own persistence, with the stored format being Yjs update blobs plus our own periodic JSON export so we can migrate away.

Why this is L6:

  • Frames build vs buy by differentiation and headcount
  • Names lock-in (doc format) and a mitigation (neutral export)

What L7 adds:

  • Prices it: build ≈ $1.5–2M/year fully loaded vs hosted at usage pricing; crossover point by MAU
  • Plans the exit before the entrance — deprecation path written into the vendor decision doc

Drill 8: Changing the Op Format Without an Outage#

Prompt: "We need to add table-merge ops. Old clients are still in the field for 6 months."

Staff Answer

Version the protocol. Clients declare protocol_version on JOIN. The owner knows the minimum version among connected editors. New op types are only emitted when all connected editors support them — otherwise the new feature is disabled in that session ("Update to merge cells"). Rollout: (1) ship server support dark, (2) ship clients that understand but don't emit, (3) enable emission behind a flag at 1% → 10% → 100% of docs with divergence monitoring, (4) after 6 months, enforce a minimum version. Transform functions for the new op against all existing ops are gated by fuzz tests before step 3.

Why this is L6:

  • Mixed-version sessions handled explicitly
  • Staged rollout with divergence as the canary metric

What L7 adds:

  • A written op-type governance process: who may add an op type, required tests, deprecation schedule
  • Minimum-client-version policy coordinated with mobile release trains

Drill 9: Cost#

Prompt: "Storage is growing 40% a year. Where's the money going?"

Staff Answer

Almost certainly op history. Back-of-envelope: 1B docs, 5% active monthly, active docs average 20K ops/year at ~40 bytes = 800KB/year each. 50M × 0.8MB = 40TB/year of new ops — plus replicas (×3) and indexes. Snapshots are smaller (median doc ~20KB). Actions: (1) cold-tier ops older than 90 days to object storage at ~1/5 the $/GB, (2) compact old ops into per-session snapshots, keeping author attribution at session granularity, (3) never compact legal-hold docs. Expected: 60–80% storage cost reduction on history with no user-visible change for >99% of opens.

Why this is L6:

  • Arithmetic from docs to bytes
  • Tiering + compaction with explicit exception for legal hold

What L7 adds:

  • Turns retention into a policy decision owned by legal and product, not an infra default
  • Tracks $/active-doc as the unit-economics metric reported quarterly

Drill 10: Multi-Region#

Prompt: "We're expanding to EU with data residency. Editors on one doc can be in the US and EU."

Staff Answer

Each doc has a home region where its owner and op log live; residency rules pin EU-tenant docs to EU. A US editor on an EU doc connects to the nearest edge, which proxies the WebSocket to the EU owner — remote edits for them cost ~80–100ms extra, local edits are still instant. For docs without residency constraints, home region follows the majority of editing activity: if >70% of ops over 7 days come from another region, migrate home — drain the session, copy the log tail, flip the lease region, clients reconnect. We don't do multi-master sequencing across regions; cross-region consensus on every keystroke adds 70–150ms and buys nothing users notice.

Why this is L6:

  • Home-region model with explicit latency cost for remote editors
  • Residency as a hard constraint on placement
  • Rejects multi-master with a quantified reason

What L7 adds:

  • Region migration as a platform capability used by every stateful product
  • Residency compliance audited continuously, with legal as the sign-off owner

8. Deep Dive Scenarios#

Deep Dive 1: Monday-Morning Peak Incident#

Context: At 9:05am Monday, op_ack_latency_p99 jumps from 40ms to 3.5s across one cell. Users see "Saving…" that never clears. On-call escalates to you.

Questions to Surface First:

  • Is it every doc in the cell or a subset? (Hot doc vs infrastructure)
  • Is the op log slow, or are owners CPU-bound?
  • Did anything deploy in the last hour — editor client, session server, storage config?
  • Are clients still able to edit locally? (Is this latency or loss?)

Typical L5 Approach: Scale up session servers in the cell and increase op log capacity. Reasonable, but blind — if one mega-doc is the cause, more servers don't help it.

Staff Approach: Check oplog_append_p99 first. If storage is healthy, it's owner-side: look for docs with session_subscribers > 500 on the slow owners. Found: a company-wide OKR doc with 6,000 joiners pinned to one owner, starving 9,000 co-located docs. Immediate: move viewers of that doc to relay tier; migrate co-located docs off the hot owner. Then fix admission control that should have capped editors at 100.

Principal Approach: Co-location is the systemic bug. One hot doc shouldn't degrade 9,000 innocent docs. Introduce per-doc resource quotas inside owners and bin-packing that isolates docs above a subscriber threshold onto dedicated owners. Add "mega-doc" as a first-class product state with its own UX (viewer mode by default), and make the capacity model assume a weekly peak of N all-hands docs per large tenant.

Staff Approach — Full Reasoning
PhaseAction
Immediate (0–5 min)Confirm clients edit locally (no loss). Check oplog_append_p99 vs owner_cpu. Identify hot doc_ids by subscriber count.
TriageHot doc with 6K joins; admission cap config was not applied to that tenant.
Quick fixForce viewer mode beyond 100 editors; migrate 9K co-located docs to other owners (brief reconnect each).
GuardrailsPer-owner subscriber ceiling; auto-isolate docs >500 subscribers.
Post-mortemWhy was the tenant exempt from the cap? Who owns tenant-level config?

Metrics to Watch: op_ack_latency_p99 by cell, session_subscribers top-K, owner_cpu, oplog_append_p99, session_migration_total.

Organizational Follow-up: Tenant config changes require review from collab infra; capacity planning includes a "Monday all-hands" peak model.

Ownership Question: Who owns the editor cap — product or infra? Product owns the number; collab infra owns enforcement and alerting when it's bypassed.

Key Takeaway: "Hot docs are a fan-out problem, and co-location turns one hot doc into a cell-wide incident."

What clears the Staff bar:

  • Rules out storage before scaling compute
  • Protects innocent co-located docs first
  • Converts a config gap into an ownership fix

Deep Dive 2: The Silent Divergence#

Context: A customer reports two colleagues see different text in a contract. No alerts fired. It's been happening for 4 days.

Questions to Surface First:

  • Do we have client state hashes, and are they being checked or just logged?
  • What changed 4–5 days ago in the editor or session server?
  • Which op types were involved in the divergent docs?
  • Which version is "right" — and does the server's version match either?

Typical L5 Approach: Reproduce the bug, fix the transform, force affected clients to reload. Correct fix for the code bug.

Staff Approach: Treat it as a detection failure first. The transform bug is one bug; the missing alert let it live 4 days. Server state is truth: force reload for divergent sessions, find all affected docs by scanning hash-mismatch logs, kill-switch the new op type. Then wire doc_divergence_detected_total to page, and add the offending op pair to the fuzz corpus.

Principal Approach: Divergence is a correctness SLO with zero tolerance, and it had no owner. Establish it as a tier-0 metric with an owning team, require every new op type to pass a convergence fuzz gate owned by a central correctness group, and schedule quarterly "divergence game days" where a deliberately broken transform is injected in staging to prove detection works.

Staff Approach — Full Reasoning
PhaseAction
Immediate (0–5 min)Kill-switch the new op type. Confirm server state vs both clients.
TriageScan hash-mismatch logs for 5 days; list affected docs (~0.3% of docs with concurrent suggestions).
Quick fixForce reload for affected sessions; notify affected customers where user-visible text differed.
GuardrailsDivergence metric pages at >0 over 5 min. Fuzz tests gate deploys.
Post-mortemWhy was the hash computed but not alerted? Who owns correctness metrics?

Metrics to Watch: doc_divergence_detected_total, client_reload_forced_total, fuzz test pass rate per build.

Organizational Follow-up: Correctness metrics get the same paging tier as availability. New op types need sign-off from the collab infra correctness owner.

Ownership Question: Who gets paged for divergence? Editor team on-call, because they own transform correctness — collab infra owns the detector's uptime.

Key Takeaway: "Silent divergence is worse than an outage — nobody knows to complain. Detection is the feature."

What clears the Staff bar:

  • Frames detection gap as the root cause
  • Server-is-truth recovery path
  • Correctness metric wired to paging

Deep Dive 3: Onboarding a 200K-Seat Enterprise#

Context: A 200,000-employee company is migrating from a competitor next quarter. They require EU data residency, 7-year retention of all edit history, and admin-controlled offline policies.

Questions to Surface First:

  • How many docs and how much history are they importing, and in what format?
  • Does "7-year edit history" mean keystroke-level or version-level?
  • Does residency include presence and relay traffic, or only stored data?
  • What's their peak concurrency pattern (all-hands docs, quarter-end planning)?

Typical L5 Approach: Provision more capacity in the EU region and extend retention for their tenant.

Staff Approach: Clarify "edit history" with their compliance team — version-level (one snapshot per session with authors) is ~50× cheaper than keystroke-level. Pin their tenant's docs to EU cells. Import runs as a batch pipeline creating docs with a synthetic single-op history, rate-limited so it can't touch live cells. Offline policy becomes a tenant setting. Capacity: 200K seats, ~10% concurrent at peak, ~2 docs open each → 40K sessions; that's ~4 more owners per EU cell — fine.

Principal Approach: This customer is the forcing function for tenant-level policy as a product: residency, retention, offline, and editor caps become a policy engine consumed by every collaborative product, not per-customer config. Price the retention: 7 years of keystroke-level history for 200K users is a line item sales must see before signing.

Staff Approach — Full Reasoning
PhaseAction
ClarifyRetention granularity, residency scope, import volume
PlacementDedicated EU cells for the tenant; relay nodes in EU
ImportBatch pipeline, throttled, isolated worker pool
PolicyTenant-level retention exempt from 90-day compaction; offline policy toggle
VerificationResidency audit: no op or snapshot for tenant outside EU

Metrics to Watch: import_docs_per_sec, cell_owner_cpu{tenant}, residency_violation_total (must be 0), history storage by tenant.

Organizational Follow-up: Sales requires infra sign-off on non-default retention; legal owns residency audit.

Ownership Question: Who owns the tenant's retention cost? The account's P&L — the platform bills retention as a tier, so the decision is priced where it's made.

Key Takeaway: "Enterprise onboarding is a policy problem wearing a capacity costume."

What clears the Staff bar:

  • Negotiates the requirement (version vs keystroke) before building
  • Isolates import from live traffic
  • Turns a one-off into tenant policy

Deep Dive 4: Post-Mortem — Lost Edits During a Deploy#

Context: During a session-server rolling deploy, ~2,000 users lost 5–30 seconds of work. The post-mortem is yours to lead.

Questions to Surface First:

  • Were the lost edits acked or unacked?
  • Did draining servers ack before appending?
  • Did clients resend unacked ops on reconnect — or discard their pending queue?
  • Why didn't staging catch it?

Typical L5 Approach: Found: the new server version acked on receipt to "reduce latency," then was SIGTERM'd before flushing. Revert the change and add a flush on shutdown.

Staff Approach: The revert is right, but the root cause is that a single PR could move the ack point without anyone noticing. The ack-after-append rule must be enforced structurally: the ack message is only constructible from an append result. Add a deploy-time invariant test (kill -9 an owner under load in staging; assert zero acked-op loss by replaying logs). Drain protocol: stop accepting new ops, finish in-flight appends, hand off leases, then exit.

Principal Approach: The durability promise needs an independent auditor. Nightly, sample 1% of sessions: compare client-reported acked revs against the op log. Report "acked edits lost" as an org-level SLO with an error budget of zero — breaching it freezes collab deploys until fixed. That turns a norm into a mechanism.

Staff Approach — Full Reasoning
PhaseAction
TimelineDeploy started 14:02; ack-on-receipt PR merged 3 days earlier; loss during pod termination
Root causeAck point moved; drain didn't flush; client trusted ack and cleared pending queue
FixRevert; type-level enforcement of ack-after-append; drain protocol
GuardrailsChaos test in CI; durability auditor
CommsAffected users notified; restore from client crash logs where possible

Metrics to Watch: acked_ops_missing_from_log_total, drain_duration_seconds, pending_queue_cleared_without_ack_total.

Organizational Follow-up: Changes to ack semantics require design review from collab infra leads.

Ownership Question: Who approves changes to the ack point? Collab infra tech lead — it's a durability contract, not an implementation detail.

Key Takeaway: "If one PR can silently move your ack point, you don't have a durability guarantee — you have a convention."

What clears the Staff bar:

  • Enforces invariants structurally, not by review alone
  • Adds chaos testing to catch this class
  • Makes durability measurable

Deep Dive 5: Multi-Region Expansion#

Context: Leadership wants the product in APAC. Current architecture is single-region US. Latency for APAC users is ~200ms to the US owner.

Questions to Surface First:

  • Are APAC users editing APAC-created docs, or collaborating with US teams?
  • Any residency requirements (Japan, Australia, India)?
  • Is the op log store multi-region capable, or do we deploy new cells?
  • What's the budget for region #2?

Typical L5 Approach: Deploy a full stack in APAC and replicate everything bidirectionally.

Staff Approach: Home-region model. New docs are homed where created; each region runs its own cells. Cross-region collaborators connect via edge to the home owner — they pay ~150–200ms on remote edits, local edits still instant. Async replication of op logs to a secondary region gives disaster recovery with ~5s RPO. Home migration when >70% of editing moves. No multi-master sequencing.

Principal Approach: Region expansion is a template, not a project. Codify it: a region "kit" (cells, lease service, op log, relay, snapshots) deployable by one team in weeks; a placement service shared across collaborative products; and a documented DR posture per region with annual failover exercises. Price region #3 and #4 before approving #2.

Staff Approach — Full Reasoning
PhaseAction
PlacementDoc home = creator's region unless residency says otherwise
AccessEdge proxy to home owner; relay tier in each region for viewers
DRAsync log replication, RPO ~5s, RTO ~5 min with lease re-home
MigrationActivity-based home migration, drain + lease flip
ValidationRegion-kill game day before GA

Metrics to Watch: remote_edit_latency_p50{editor_region, home_region}, home_migration_total, replication_lag_seconds.

Organizational Follow-up: Placement service owned by collab infra; regional on-call rotation.

Ownership Question: Who decides a doc's home region when residency and latency conflict? Residency always wins; legal owns the policy, collab infra enforces it.

Key Takeaway: "Home the sequencer, proxy the remote editors. Consensus across oceans for every keystroke buys nothing."

What clears the Staff bar:

  • Rejects multi-master with a latency argument
  • DR with explicit RPO/RTO
  • Residency as a hard placement constraint

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Choose OT-style or CRDT based on document model, offline intent, and server validation needs — and name the cost of the one you didn't pick
  • Define the ack point precisely and defend "ack after replicated append" with latency numbers
  • Design single-owner-per-doc sequencing with leases and fencing tokens, and walk through failover second by second
  • Separate editors from viewers and presence from ops, with concrete caps and rates
  • Bound offline merges, re-check permissions at merge, and fork beyond the bound
  • Build detection for silent divergence and acked-op loss
  • Price history storage and propose tiering with a legal-hold exception
  • Explain how the design becomes a platform for multiple collaborative products (L7)

The Bar for This Question#

Mid-level (L4): Builds a working real-time editor: WebSockets, a server that applies ops, a database. May use timestamps or whole-doc saves. Understands conflicts exist but resolves them naïvely.

Senior (L5): Correctly explains OT or CRDT mechanics, shards by doc ID, uses sticky sessions, snapshots for load time. The design converges in the happy path. Gaps: ack point undefined, failover hand-waved, viewers and editors treated the same, offline and permissions bolted on.

Staff+ (L6): Starts from intent and commits. Makes durability a stated promise with a precise ack point. Designs fenced single ownership with idempotent client resend. Caps editors and splits viewer fan-out. Bounds offline with permission re-check. Names owners for correctness (editor team), durability and sessions (collab infra), and policy (product/legal). Builds detection for the silent failures. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "OT vs CRDT Is the Least Important Decision in the Design"#

EvidenceImplication
Google Docs (OT) and Figma (CRDT-inspired, server LWW) both ship world-class collaborationBoth algorithm families work in production
Real incidents in collaborative products are dominated by session failover, storage, and deploy bugsThe pager rings for operations, not transforms
Mature CRDT libraries (Yjs, Automerge) are free to adoptThe algorithm is increasingly a dependency, not a design

The Staff position: Spend 2 minutes on the algorithm and 30 on ownership, durability, and failure. The choice is reversible if the op format is versioned; the ack semantics are not.

Why this matters in interviews: Interviewers use the OT/CRDT debate as a trap to see whether you'll burn your time budget on it.

10.2 "Most Documents Should Never Get a Session Server"#

EvidenceImplication
Typically <10% of docs are opened in a month; far fewer have 2+ concurrent editorsA stateful session per open doc is mostly wasted
Solo editing works fine with autosave + optimistic concurrencyCheap path covers the majority

The Staff position: Promote a doc to a live session only when a second participant joins. Demote after 5 minutes solo. Session fleet shrinks by 5–10×.

Why this matters in interviews: Shows you design for the distribution of usage, not the headline feature.

10.3 "Presence Is More Expensive Than Editing"#

EvidenceImplication
Cursor moves happen at mouse/selection rate; edits at keystroke rateUnthrottled presence is 5–10× op volume
Presence fan-out is N² in participants100 participants = 10,000 cursor streams

The Staff position: Presence gets its own lossy channel, throttled to 5–10 Hz, aggregated for viewers, and is the first thing shed under load.

Why this matters in interviews: Candidates who only size ops under-provision by an order of magnitude.

10.4 "Offline Editing Is a Security Feature Request in Disguise"#

EvidenceImplication
Offline copies live on devices outside DLP controlsSecurity must sign off
Permissions change while users are offlineMerge must re-authorize

The Staff position: Offline is bounded, tenant-configurable, and re-authorized at merge. Never unbounded by default.

Why this matters in interviews: It moves the conversation from algorithms to ownership — which is where Staff lives.

10.5 "Version History Is Your Largest Cost Center"#

The Staff position: Compute is sized by concurrency; storage is sized by every keystroke ever typed. Tier and compact history by policy, or it will dominate the bill within 2–3 years.

Why this matters in interviews: Shows you think about the system at year 3, not day 1.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

At Staff, collaborative editing is a system to design. At Principal, it is a capability the company will need five times: docs, spreadsheets, slides, whiteboards, comments, code notebooks. Each team, left alone, will build its own sequencer, op log, presence channel, and permission check — five fleets, five on-call rotations, five subtly different durability promises. The L7 question is: what is the sync substrate, what's the contract, and which parts must stay product-specific? The op algebra is product-specific. Sessions, logs, presence, permissions, and placement are not.

The Org-Level Fault Line#

One sync platform vs per-product sync engines.

OptionWhat WorksWhat BreaksWho Pays
Per-product enginesEach team optimizes for its model; ships fast initially5× fleets, inconsistent durability, duplicated presence; every region launch done 5 timesInfra budget; on-call; customers who see inconsistent behavior
Central sync platform (log + sessions + presence + ACL), product-owned op typesOne durability SLO, one region kit, shared relay tierPlatform becomes a bottleneck; op-type governance neededPlatform team (headcount); product teams (lose some autonomy)
Buy (hosted collaboration service)Fastest; no fleetResidency, lock-in, per-MAU cost at scaleFinance at scale; legal

The L7 default: central platform for sessions/log/presence/ACL enforcement, with product teams owning their op types and transforms behind a plugin interface. Don't standardize the document model — that's where products differentiate.

Cost Model#

Assumptions: 40 bytes/op, 3× replication, 5% of docs active monthly, active docs avg 20K ops/year, one session server (~$400/month) holds ~10K active docs, snapshots median 20KB.

ScaleDocs / Peak Concurrent SessionsMonthly InfraHeadcountOn-call Load
Startup1M docs / 5K sessions~$3–5K (managed KV + 2–3 session servers + blob)2–3 engineers, part-time on-call~1–2 pages/month
Growth100M docs / 500K sessions~$80–150K (50 session servers, op log ~200TB with replicas, relay tier)8–12 engineers, dedicated rotation~5–10 pages/month
Hyperscale1B+ docs / 5M+ sessions, multi-region~$1.5–3M (history storage ~60% of bill before tiering)30–50 engineers across platform + per-productPer-region rotations; SLO-driven

Where money goes: at growth scale and beyond, history storage outgrows compute. Cold-tiering ops older than 90 days cuts total bill ~30–40%.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Op log storage format (encoding of ops)One-way (after scale)Rewrite of billions of stored ops; version it from day one
Client protocol (JOIN/SUBMIT/ACK semantics)One-way-ishMobile clients live 6–18 months; needs multi-version server support
Ack point semanticsOne-way (trust)Users who lose work don't come back
OT vs CRDT algorithmTwo-way if format is versionedMonths of migration, but feasible
Lease TTL, snapshot cadence, editor capTwo-wayConfig change
Managed KV vs self-run log storeTwo-way with effortDual-write migration, ~1–2 quarters
Retention policyOne-way once data is deletedDeleted history cannot be restored

The Standard I'd Write#

RFC: Real-Time Collaboration Platform Standard (v1)

Scope: Any product feature where two or more users edit shared state concurrently with sub-second visibility.

MUST:

  • Use the platform session service for ordering; no product-run sequencers.
  • Acknowledge edits only after durable replicated append to the platform op log.
  • Tag every op with (client_id, client_seq); servers dedupe on resend.
  • Enforce ACLs at join and per op; revocations take effect within 10s.
  • Emit doc_divergence_detected_total, op_ack_latency_p99, and acked_ops_missing_from_log_total.
  • Version op encodings; ship convergence fuzz tests with every new op type.

SHOULD:

  • Keep presence on the lossy channel, ≤10 Hz per participant.
  • Cap concurrent editors per document (default 100) and route excess to viewer relays.
  • Use platform tiering for history older than 90 days unless under legal hold.

Exceptions: Reviewed by the collaboration platform council; granted for ≤2 quarters with a migration plan.

Success metrics: 0 acked edits lost per quarter; remote edit p50 <200ms in-region; ≤1 sync engine per company; new product onboarding ≤6 weeks.

What I'd Tell the VP#

"Real-time collaboration is becoming table stakes across our products, and right now three teams are about to build it separately. I'm proposing one collaboration platform that handles saving, ordering, and sharing permissions, while each product keeps control of its own document features. It costs about 8 engineers for a year, and it replaces roughly 20 engineers' worth of duplicated work across teams over three years. The biggest risk we're buying down is lost user work — we'll commit to zero lost saved edits and measure it. The biggest ongoing cost is edit history storage, and we'll control it with a retention policy that legal signs off on."

Principal Interview Signals#

SignalWhat It Sounds Like
Platform framing"The op algebra is product-specific; sessions, log, presence and ACL are platform."
Pricing history"By year 3 history is 60% of the bill — retention is a policy decision, not an infra default."
One-way door awareness"I'll version the op encoding now because it's the only decision here that's expensive to reverse."
Org failure posture"Cells cap any single deploy or hot doc at 5% of documents; we prove it with quarterly game days."
Knowing when not to standardize"I won't standardize the document model — that's where products compete."

Staff answers that L7 interviewers find insufficient:

  • "Collab infra owns the session servers" — correct, but doesn't address the three other teams building their own.
  • "We'll add cold storage for old ops" — a tactic without a retention policy owner or a price.
  • "We'll run a failover drill" — one drill for one system, not an org-wide failure posture with error budgets.

🧭 Principal Move: "Before we design Docs' sync engine, let's decide whether it's Docs' or the company's. If Sheets and Whiteboard will need it within 18 months, the right design is a platform with product-owned op types, and the first customer is Docs."


Appendices

Appendix A: Mechanics in Depth#

A.1 Server-Centric OT (Jupiter Model)#

Why it's right for online-first: the client transforms only against the server's stream, and the server transforms only against its own history. That reduces N-way concurrency to repeated 2-party transforms and avoids TP2 entirely.

server.receive(op, base_rev, client_id, client_seq):
  if dedupe.contains(client_id, client_seq): return ack(existing_rev)
  for r in (base_rev+1 .. head):
      op = transform(op, log[r])          # op' against already-sequenced op
  authorize(op, current_acl)
  rev = head + 1
  log.append(doc_id, rev, op, fencing_token)   # conditional write
  head = rev
  ack(client_id, client_seq, rev)
  broadcast(rev, op)

client:
  pending = []        # sent, not acked
  buffer  = []        # not yet sent (one in-flight batch at a time)
  on remote(rev, op):
      for p in pending + buffer: (op, p) = transform_pair(op, p)
      apply(op)

Why it's wrong for offline-first: a client offline for weeks must transform against tens of thousands of ops, and the server is a hard dependency for any merge.

A.2 Sequence CRDTs#

Each inserted element gets a unique ID (lamport, replica_id) and references its left neighbor (RGA) or both neighbors (YATA). Deletes mark tombstones. Merges are commutative — apply updates in any order, converge. Costs: per-element metadata (compressed via run-length encoding of consecutive inserts by the same replica), tombstone GC that requires knowing every replica has seen the delete (hard with offline peers), and the interleaving anomaly where two users' concurrent paragraphs get interleaved character by character in naive algorithms (addressed by newer designs like Fugue).

A.3 Per-Property LWW (Canvas Model)#

object[id].props[key] = (value, server_seq)
on op(id, key, value): if accepted by server, assign server_seq, broadcast
ordering children: fractional index strings between neighbors ('a0' < 'a0V' < 'a1')

Right for structured objects; wrong for text runs, where it would drop concurrent characters.

A.4 Undo in a Collaborative Doc#

Undo must be local: undo my last op, not the doc's last op. Implement as generating the inverse of my op, transformed against every op sequenced since. If my inserted text was since deleted by someone else, the inverse becomes a no-op. Never "roll back the log."

Appendix B: Data Model#

TableKeyFieldsNotes
docsdoc_idowner, tenant, home_region, cell, head_rev, snapshot_rev, fencing_token, retention_policyStrongly consistent; small
ops(doc_id, rev)author, client_id, client_seq, op_blob, tsAppend-only; conditional write
snapshots(doc_id, rev)blob_uri, hash, verified_atBlob in object storage
acl(doc_id, principal)role, granted_by, tsRevocations pushed to owner
named_versions(doc_id, name)revExempt from compaction

Appendix C: Coordination Mechanisms — Quick Comparison#

MechanismOrderingFailoverSplit-brain SafetyBest For
Lease + fencing + ownerIn-memory, fast5–12sYes (fencing)Default for busy docs
DB conditional write onlyStorage CASInstantYesLow-concurrency docs
Per-doc log partitionLog orderPartition leader electionYesEvent-sourcing shops already on Kafka
CRDT, no sequencerNone neededN/AN/AOffline-first / P2P
Diagram: Appendix C: Coordination Mechanisms — Quick Comparison

Appendix D: Client Behavior#

  • One batch in flight: client sends at most one unacked batch; accumulates new edits in a buffer (reduces transform work and gives natural batching of ~50–100ms).
  • Reconnect: full-jitter exponential backoff, base 500ms, cap 30s; send last_acked_rev; resend pending with original client_seq.
  • Crash safety: pending queue persisted to IndexedDB/local storage every 1s; survives tab close.
  • UI states: "Saved" only after ack; "Saving…" while pending; "Offline — changes saved on this device" when disconnected; never "Saved" before ack.

Appendix E: Observability#

E.1 Core Metrics#

op_ack_latency_ms{p50,p99}          # submit → ack
remote_op_latency_ms{p50,p99}       # author submit → other client apply
oplog_append_ms{p99}
session_failover_total
oplog_append_rejected_stale_token_total
doc_divergence_detected_total        # client hash mismatch
acked_ops_missing_from_log_total     # auditor
session_subscribers{doc_id} top-K
presence_msgs_dropped_total
acl_revocation_lag_seconds{p99}

E.2 Critical Alerts#

AlertThresholdAction
Divergence>0 for 5 minPage editor team
Acked-op loss>0SEV1; freeze collab deploys
Ack latencyp99 >1s for 5 minPage collab infra
Failovers>3× baselinePage collab infra
Revocation lagp99 >30sPage identity + collab

E.3 Control Plane vs Data Plane#

Data plane: session owners, op log, relays — must keep working if the control plane (placement, leases for new docs, admin) is degraded. Existing owners continue serving with their current lease; new doc placements queue.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Changes
1K concurrent sessionsOne owner process, Postgres op tableNothing clever needed
100K sessionsOwner fleet, lease service, managed KV log, snapshotsRouter, fencing, compactor
1M+ sessionsCells, relay tier, lazy sessions, history tieringPlacement service, divergence auditing
Multi-regionHome regions, edge proxy, async DR replicationResidency, home migration

What You Don't Build on Day One#

  • Multi-region homing — single region with DR backups until customers demand it
  • Custom CRDT or OT library — adopt one or keep op types minimal
  • Viewer relay tier — until a doc exceeds ~500 subscribers
  • Keystroke-level history for 7 years — start with 90-day hot history

Appendix G: Multi-Tenancy and Cost#

  • Noisy tenant: one tenant's all-hands doc should not degrade others — per-tenant cells for the largest 1% of tenants.
  • Quotas: per-tenant limits on concurrent sessions and import rate.
  • Chargeback: extended retention and dedicated residency billed as plan tiers so cost is priced where it's decided.
  1. Loading the index…