Hiring BarSupport

Design Dropbox (File Sync) — Staff-Level Case Study

Case study78 min read6 diagrams

Technologies referenced in this case study: PostgreSQL · Kafka · Redis · DynamoDB · Cassandra

Related case studies: Blob Storage (S3) · Collaborative Editing · Real-Time Updates · Notification System · Handling Large Blobs

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once. Return to individual sections for targeted review.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → the Deep Dive for your weak spot
Deep Dive3+ hrsEverything, including the Principal Lens and appendices
What is File Sync? — Why interviewers pick this topic

File sync keeps a tree of files identical across N devices and M users, where any device can go offline for weeks, edit anything, and come back. The server holds the canonical history; clients hold a local copy plus a cursor into that history. Dropbox, Google Drive, OneDrive and iCloud Drive all solve the same problem with different answers to the same five questions.

Before vs After — the "laptop comes back from a 3-week trip" scenario:

Naive design (upload whole files, last-writer-wins by mtime):
t=0:       Laptop reconnects after 21 days offline. 4,200 local edits queued.
t=+5s:     Client re-uploads 38 GB of whole files, including a 9 GB video with a 1-byte tag change.
t=+40min:  Upload still running; home uplink saturated at 20 Mbps.
t=+41min:  Laptop's clock is 7 minutes fast. Its stale copy of budget.xlsx "wins" by mtime.
t=+42min:  Colleague's 3 weeks of edits to budget.xlsx silently overwritten on 6 devices.
t=+3 days: Colleague notices. Support ticket. Version history restore. Trust gone.

Staff design (content-addressed blocks, journal cursor, conflicted copies):
t=0:       Laptop reconnects. Client sends its cursor (journal position 88,412) for each namespace.
t=+1s:     Server returns 1,930 remote changes since 88,412. Client reconciles 3-way: base / local / remote.
t=+2s:     Client hashes changed files into 4 MB blocks; asks server which of 9,600 hashes are missing.
t=+3s:     Server: 212 blocks missing (the rest already exist — dedup + unchanged blocks).
t=+90s:    848 MB uploaded, 4,200 commits applied with base-revision checks.
t=+91s:    budget.xlsx commit rejected (base rev 17, server at rev 29) → "budget (Laptop's conflicted copy).xlsx".
t=+92s:    Zero data lost. One visible conflict. Every other device converges within 2s via push.

Why interviewers reach for this question: it is the rare prompt that combines a huge blob-storage problem, a strongly-ordered metadata problem, a push-notification problem, and an offline-first client problem. A candidate who only designs the server fails; a candidate who only talks about chunking fails. The interviewer is watching whether you find the real center of gravity: the metadata journal and the reconciliation protocol.

Mechanics Refresher: Chunking, Dedup, and Delta
TechniqueHow It WorksProsCons
Whole-file uploadRe-upload the full file on any changeTrivial; no block index9 GB video re-uploaded for a 1-byte edit
Fixed-size blocks (e.g., 4 MB)Split at fixed offsets; hash each block (SHA-256)Simple, parallel, cheap block index; dedups identical files and unchanged blocksAn insert at byte 0 shifts every boundary → every block "changes"
Content-defined chunking (CDC)Rolling hash (Rabin/Gear) picks boundaries where hash mod 2^k == 0; avg 1–8 MBInserts only disturb 1–2 chunks; better dedup for shifting contentVariable chunk sizes; ~2–3× more CPU; harder to reason about sizes
rsync-style deltaReceiver sends weak rolling checksum (Adler-32) + strong hash per block; sender finds matches at any offsetTransfers only changed bytes inside a blockNeeds old version on both sides; CPU-heavy on server if done there
Binary diff within a blockUpload a patch against the previous block versionMinimal bytes over the wireServer must materialize patches; patch chains complicate GC

For most production systems: fixed 4 MB content-addressed blocks for storage and dedup, plus an rsync-style delta on the wire for large files that change in place. Content-defined chunking is worth it for backup products, not for consumer sync of mostly-small files — the median synced file is well under one block.


Executive Summary

If you only read one section, read this. Everything else in the case study elaborates on the contrast below.

What This Interview Actually Tests#

File sync is not a storage question. Blob storage is a solved, rentable commodity — S3 will hold your blocks at eleven nines.

This is a distributed state reconciliation question that tests:

  • Whether you separate the block plane (immutable, content-addressed, dumb, huge) from the metadata plane (mutable, ordered, small, the hard part)
  • Whether you can define a commit protocol that is idempotent, resumable, and conflict-detecting across flaky networks
  • Whether you have a conflict policy a product manager has signed off on — and know that "last-writer-wins" silently destroys work
  • Whether you can reason about a fleet of clients you do not control: old versions, wrong clocks, 3-week-offline laptops, 800K-file trees

The key insight: the metadata journal is the product. Blocks are just bytes. Every hard bug in file sync — lost edits, ghost files, infinite re-sync loops, "my folder disappeared" — lives in metadata and in the client's reconciliation of it.

The L5 → L6 → L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws clients → API → S3 + metadata DBAsks "consumer sync, shared team folders, or backup?" — they have incompatible conflict and durability barsAsks what the org already owns: is there a storage platform, an identity/permissions platform, a desktop-client team? Which of those is the real bottleneck?
Chunking"Split into blocks, upload changed ones"Fixed 4 MB content-addressed blocks + commit protocol "which hashes are missing?"; names the shifted-boundary problem and when CDC is worth itPrices the block index: 1 EB / 4 MB = 250B block rows; picks block size as a cost lever for 3 years, knowing it is a one-way door
Consistency"Metadata in a SQL DB, it's consistent"Per-namespace serialized journal with monotonic cursor; strong within a namespace, nothing acrossDesigns namespace as the shard, cell, and blast-radius unit; knows the 1-namespace-with-40M-files customer will arrive
Conflicts"Last-writer-wins by timestamp"Base-revision check on commit; losers become conflicted copies; zero silent lossMakes conflict rate a product metric with an owner; decides which file types get real merge (docs) vs copies (binaries)
Failure"Replicate S3 and the DB"Resumable uploads, idempotent commits, client-side journal replay, kill switch for a bad client releaseOwns the org's client-fleet risk: staged rollout rings, server-side feature gates, a "stop syncing deletes" global brake
Ownership"The storage team owns it"Splits: block storage (infra), metadata/journal (sync core), client (desktop/mobile), notifications (realtime) — names who is paged for "files missing"Redraws boundaries so one team owns the sync protocol end-to-end across client and server, because protocol bugs fall between teams
Why "first move" separates levels

L5: Starts drawing upload flows. The design is fine for consumer sync, but the candidate never learns whether the interviewer meant a 50-seat law firm with shared folders and legal hold, or a personal photo backup that never has two writers. Those two answers differ on conflict handling, retention, permissions, and whether cross-user dedup is even legal.

L6: "Three different products hide under 'Dropbox': personal multi-device sync, shared team folders with many writers, and one-way backup. I'll design multi-device sync with shared folders, because that is where concurrency and conflicts actually happen — backup is a strict subset."

L7: Adds the org lens: "Before I design the block store — does the company already run an object store at exabyte scale? If yes, my design is a metadata system plus a client. If no, the build-vs-buy decision on storage dominates the 3-year cost curve and I'd want to size that first."

Why "conflicts" separates levels

L5: "Last-writer-wins by modified time." This is reasonable in a single-user world and catastrophic in a shared one: client clocks drift by minutes, and the loser's work vanishes silently on every device.

L6: Every commit carries the base revision the client edited from. If the server's current revision ≠ base, the commit is a conflict. The server never overwrites; the losing client materializes a (conflicted copy) file. The product guarantee becomes "we never silently lose an edit" — which is testable and alertable.

L7: Treats conflict rate as a product KPI with a target (e.g., < 0.05% of commits) and an owner, and decides the org-level question: which content deserves real merge (collaborative docs go to an OT/CRDT system — see Collaborative Editing) versus which stays as opaque blobs forever.

Why "ownership" separates levels

L5: Names one owning team. In reality four teams touch every sync: desktop client, sync API/metadata, block storage, notification fan-out. When a user says "my file disappeared," each team's dashboards are green.

L6: Defines the seams and the contract at each seam — the commit API, the journal cursor semantics, the block-existence API — and makes the sync-correctness SLO (e.g., "99.99% of commits visible on all online devices within 10s") owned by one team.

L7: Recognizes that protocol bugs fall between client and server teams, so the org chart should put the sync protocol under one owner with authority over both. Also owns the decision to keep a client team capable of shipping a fix to 100M desktops in 72 hours — the org's most dangerous deploy target.

The Staff Positions#

PositionRationale
Split block plane from metadata planeBlocks are immutable and content-addressed — trivially cacheable, replicable, dedupable. Metadata is mutable and ordered. Mixing them couples the easy problem to the hard one.
Fixed 4 MB blocks, SHA-256 addressedMedian file < 1 block; fixed boundaries keep the index simple. Add wire-level delta for large in-place edits instead of CDC in storage.
Per-namespace serialized journal with a cursorEvery client sync is "give me changes after cursor X." Strong ordering inside a namespace, no cross-namespace transactions.
Conflicted copies, never silent overwriteBase-revision check on every commit. Users hate conflicts less than they hate lost work.
Dedup per-namespace or per-user by default, not globalGlobal dedup leaks "does this file exist anywhere" — a known side channel. Enable cross-user dedup only with server-side proof of possession.
Push a "something changed" hint, pull the actual changesNotifications carry no data; the journal is the source of truth. Lost pushes cost latency, not correctness.
Deletes are soft for 30+ daysA buggy client or ransomware will mass-delete. Recovery must be a server-side rollback, not a support ticket per file.

The Three Intents#

Three intents hide under "Design Dropbox." Each leads to a different system.

IntentConstraintStrategyFailure ModeCorrectness Bar
Personal multi-device syncLatency to other devices; battery; bandwidthThick client, block dedup, delta, push hint + journal pullOffline edits on two devices → conflictNever lose an edit; conflicts rare (< 0.1%) and visible
Shared team foldersMany writers; permissions; audit; admin controlsNamespaces with ACLs, per-namespace journal, mount points, admin recoveryPermission drift; mass delete by one member propagates to 200 peopleNever lose an edit and every change attributable to a user for audit
One-way backup / archiveDurability and cost per GB; restore timeSingle writer, content-defined chunking, aggressive global dedup, cold tiersSilent corruption discovered at restore timeBit-exact restore; periodic scrub; no conflicts by construction

🎯 Staff Move: "I'll design multi-device sync with shared folders. Backup is a strict subset of that design — one writer, no conflicts — so if we get sync right, backup is a configuration. The reverse isn't true: a backup system can't grow into shared folders without rebuilding the metadata layer."

The Five Fault Lines#

#Fault LineThe Tension
1Block Size & Chunking StrategySmall or content-defined chunks → better dedup and delta; large fixed blocks → smaller index, fewer requests, simpler GC
2Dedup ScopeGlobal dedup saves the most storage but leaks existence and complicates deletion/legal holds; per-namespace dedup is safe but saves less
3Metadata Consistency ModelOne serialized journal per namespace (correct, hot-spot prone) vs eventually-consistent per-file metadata (scales, produces ghosts and loops)
4Conflict ResolutionAutomatic resolution (LWW, merge) hides friction but loses data; conflicted copies preserve data but push work to users
5Client Intelligence & Fleet ControlThick clients (hashing, delta, local DB) save bandwidth and server CPU but put the riskiest code on 100M machines you can't roll back

In the Wild: Real Production Systems#

Why this section belongs here: citing well-known production systems shows you've studied operational reality. Use these as one-sentence anchors in the interview.

Dropbox — Content-Addressed Blocks, Then Its Own Storage#

Dropbox's publicly described design splits files into 4 MB blocks addressed by SHA-256 hash, with metadata stored separately from block data. Blocks originally lived in Amazon S3; around 2015–2016 Dropbox migrated the bulk of its data into Magic Pocket, an in-house exabyte-scale block store. Later, Dropbox rewrote its desktop sync engine (publicly called Nucleus) in Rust, explicitly to make the client's sync state model testable and to eliminate classes of reconciliation bugs.

Staff insight: the two biggest Dropbox engineering investments were the storage economics (Magic Pocket) and the client sync engine (Nucleus) — not the upload API. That is exactly the split this case study argues for: the block plane is a cost problem, the sync engine is a correctness problem.

Google Drive and Microsoft OneDrive — Placeholders / Files On-Demand#

Both products (and Dropbox, via its smart-sync feature) ship placeholder files: the directory tree is fully synced, but file contents download only on open. This decouples "the namespace fits on the laptop" from "the bytes fit on the laptop."

Staff insight: placeholders change the client's hardest problem from bandwidth to metadata scale — a 2M-file tree must be listed, watched, and reconciled even when 0 bytes of content are local. Say this when the interviewer asks about enterprise customers with huge shared drives.

rsync — Rolling-Checksum Delta Transfer#

rsync (Tridgell & Mackerras) sends a weak rolling checksum plus a strong hash per block of the receiver's copy, letting the sender find matching blocks at any byte offset and send only literal differences.

Staff insight: rsync proves you can get delta efficiency on the wire without changing your storage format. Keep storage fixed-block and content-addressed; use rsync-style delta only as a transfer optimization for large files edited in place.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Split into 4 MB blocks""I insert one byte at the start of a 2 GB file. What gets uploaded?"Do you know fixed-boundary shift and the CDC/delta remedies?
"Dedup by content hash""Can a user download a file they don't own by claiming its hash?"Security model of dedup; proof of possession
"Metadata in Postgres""One shared folder has 40M files and 3,000 writers. What happens?"Hot namespace; journal sharding; per-namespace limits
"Last-writer-wins""Two offline laptops edit the same file. Which edit survives?"Whether you'll silently destroy work
"Clients poll for changes""100M devices polling every 30s — what's the QPS and cost?"Push-hint + cursor-pull; long-poll economics
"Delete propagates immediately""A buggy client release deletes everyone's Documents folder. Now what?"Soft deletes, rate brakes on deletes, server-side rollback

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: three planes with three different scaling laws. The block plane scales with bytes (exabytes, cheap per GB, immutable, cacheable forever). The metadata plane scales with changes (billions of small ordered writes per day, one writer per namespace). The notification plane scales with connected devices (tens of millions of idle connections). A client never learns content from a notification — it only learns "go ask the journal." That separation is why a dropped notification costs seconds of latency, never correctness.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Chunking"Split files into blocks""Fixed 4 MB SHA-256 blocks; commit is 'here's my blocklist, tell me what's missing.' Wire-level rsync delta for big in-place edits."
Dedup"Dedup globally to save storage""Per-namespace dedup by default. Cross-user dedup only with server-side proof of possession — otherwise hashes become download tokens."
Metadata"Store file rows in SQL""Per-namespace append-only journal, monotonic revision, sharded by namespace_id. Clients sync by cursor."
Conflicts"Last-writer-wins""Base-revision precondition on every commit; losers become conflicted copies. Zero silent loss is the SLO."
Notifications"Clients poll every 30s""Long-poll / WebSocket hint with no payload; client pulls from the journal. Lost hints cost latency, not data."
Deletes"Delete propagates""Soft delete, 30-day retention, server-side delete-rate brake per client, namespace-level rollback."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Block size4 MB (fixed)Balances index size vs dedup granularity; median file fits in one block
Block hashSHA-256, 32 bytesCollision probability negligible (~2⁻¹²⁸ birthday bound); the hash is the address
Block index size at 1 EB1 EB / 4 MB ≈ 2.5 × 10¹¹ entries × ~64 B ≈ 16 TBFits in a sharded KV store; halve block size and you double it
Metadata journal entry~1–2 KB (path, blocklist, rev, author, ts)A 10 GB file = 2,500 block hashes = ~80 KB blocklist — chunk the blocklist too
Commit rate (50M DAU × 20 changes/day)~12K/s average, ~40K/s peakThe metadata plane's real load — not bytes
Idle long-poll connections per host100K–500K (epoll, ~10–20 KB each)30M online devices ≈ 100–300 notification hosts
Polling instead of push (30M devices / 30s)~1M QPS of mostly-empty requestsWhy push hints exist
Convergence targetp99 < 10s for online devicesThe user-visible SLO; measure end-to-end, not per-hop
Conflict rate (healthy)< 0.05–0.1% of commitsA spike means a client bug or clock/cursor bug, not users
Soft-delete retention30 days (consumer) to 180+ days (business)The only defense against a bad client release or ransomware
Single-namespace hot limit~1–5K commits/s per journal shardWhere a monolithic shared folder hits the wall

Interview Walkthrough

The most common mistake: candidates spend 20 minutes on upload APIs, S3 multipart, and a CDN for downloads — the parts that are genuinely easy — and never reach the journal, the reconciliation protocol, or conflicts. Compress the basics to ~10 minutes. Spend the rest where files actually get lost.


Phase 1: Requirements & Framing (2–3 minutes)#

State functional requirements in one breath:

"Users put files in a folder on any device; the same tree appears on every other device and for everyone the folder is shared with. Devices can be offline for a long time. Users can see and restore history."

Then spend the time on the non-functional requirements — that's where the design lives:

"Which product is this: personal multi-device sync, shared team folders, or backup? I'll design sync with shared folders, since backup is a subset. The requirements I'd commit to: one, we never silently lose an edit — conflicts become visible copies. Two, online devices converge within about 10 seconds p99. Three, we upload only bytes the server doesn't already have. Four, a device offline for 30 days reconnects and converges without a full rescan."

Scale it in 20 seconds:

"Say 500M registered, 50M daily active users, ~3 devices each, 20 file changes per active user per day. That's ~1B commits a day — ~12K/s average, ~40K/s peak. Storage in the high hundreds of petabytes to an exabyte. The byte volume is big but boring; the commit rate and the number of connected devices are what shape the design."

🎯 Staff Move: Naming "never silently lose an edit" as requirement #1 reframes the entire interview. It forces a base-revision commit protocol, rules out last-writer-wins, and gives you an SLO you can alert on. The interviewer now knows you think about correctness from the user's point of view.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Namespace — a root of a tree with its own journal: a user's home, or a shared folder. The unit of ordering, sharding, and permissions.
  • Journal entry — (namespace_id, rev, path, file_id, blocklist, size, author, ts, deleted). Append-only; rev is monotonic per namespace.
  • Block — immutable 4 MB byte range, addressed by SHA-256(content).
  • Cursor — a client's position in a namespace journal: (namespace_id, rev).

Block plane (bulk, idempotent):

POST /blocks/missing   { hashes: [h1..hN] }            → { missing: [h3, h7] }
PUT  /blocks/{hash}    <bytes>                          → 201 (server verifies SHA-256)
GET  /blocks/{hash}                                     → bytes (cacheable forever)

Metadata plane (small, ordered):

POST /commit  { ns, path, blocklist, base_rev, client_op_id }
              → 200 { rev } | 409 { conflict, server_rev } | 412 { missing_blocks }
GET  /list_since?ns=..&cursor=88412&limit=2000  → { entries[], new_cursor, has_more }
GET  /notify?cursors=[..]  (long-poll, 60–90s)   → { changed: [ns...] }

🎯 Staff Move: "Two properties matter in this API. base_rev makes every commit a compare-and-set, so the server detects conflicts instead of resolving them by accident. client_op_id makes commits idempotent, so a client that loses the response can retry safely. Without both, flaky networks create either lost edits or duplicate files."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw three planes, then walk the upload and the download.

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk the flow in 90 seconds:

  1. The watcher on Device A sees report.pdf change. The client chunks it into 4 MB blocks and hashes each.
  2. The client asks the block service which hashes are missing; uploads only those. Block PUTs are idempotent — same hash, same bytes.
  3. The client commits (path, blocklist, base_rev=41). The sync service checks permissions, verifies all blocks exist, and does a CAS: if current rev is 41, append rev 42. Otherwise, 409 conflict.
  4. The journal append publishes "namespace 7 changed" to the notification service.
  5. Device B's open long-poll returns with ns 7 changed — no payload.
  6. Device B calls list_since(cursor=41) and gets rev 42's metadata.
  7. Device B fetches only the blocks it doesn't already have locally, assembles the file, and atomically renames it into place.

Hit these explicitly:

  1. Blocks and metadata are separate services with separate scaling laws.
  2. The commit is the linearization point — a file doesn't exist until its journal entry does, and the journal entry can't exist until all its blocks do.
  3. Notifications are hints — correctness comes from the cursor, not from push delivery.
  4. The client holds three trees: last-synced base, local disk, and remote — reconciliation is a 3-way diff.

🎯 Staff Move: After drawing, say: "This is the design everyone draws. It works on the happy path. The interesting part is what the client does when it has been offline for a month and the three trees disagree — and what the server does when a client is wrong."


Phase 4: Transition to Depth (1 minute)#

"The block plane is basically Blob Storage — immutable, content-addressed, I'm happy to hand-wave it. The three places this design actually breaks are: conflict detection and resolution, the metadata journal under a huge shared folder, and the client fleet — especially deletes from a buggy client. Where would you like to go?"

If the interviewer has no preference, lead with conflicts — it's the most product-visible and immediately shows correctness thinking. Then go to the journal.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → pick a position → quantify the cost → name who absorbs it.

Deep dive A: Conflicts and the commit protocol (7–8 min)

"Every commit is a compare-and-set on base revision. Two laptops both edit budget.xlsx from rev 17 while offline. Laptop 1 reconnects first: commit(base=17) succeeds, rev 18. Laptop 2 reconnects: commit(base=17) fails — server is at 18. Laptop 2's client doesn't retry blindly; it downloads rev 18 into place and renames its own version to budget (Laptop 2's conflicted copy 2026-09-29).xlsx, then commits that as a new file. Both edits survive. The user merges by hand."

Name who pays: "The user pays with a bit of friction. The alternative — last-writer-wins — makes the user pay with lost work, and they don't even find out. Product signs off on this; it's a product decision, not an engineering one."

Mention the edge cases:

  • Rename vs edit: Device A renames a.txt → b.txt, Device B edits a.txt. Track file_id separately from path so the edit follows the rename instead of resurrecting a.txt.
  • Delete vs edit: Edit wins. Resurrect the file with the edit rather than deleting unseen work.
  • Directory move vs add: Moving a folder while another device adds files under it — the adds follow the folder by file_id of the parent, not by path string.

Deep dive B: The journal and the huge shared folder (7–8 min)

"The journal is sharded by namespace_id. One namespace = one ordered log on one shard. That gives us strong ordering where it matters — inside a folder tree — without distributed transactions. The failure is the monster namespace: a company-wide shared folder with 40M files and 3,000 active writers. At ~1–5K commits/s per journal shard, one namespace can saturate its shard."

Options and position:

  • Split the namespace into child namespaces at mount points (each team subfolder becomes its own namespace, mounted into the parent). Ordering is only guaranteed within a child — acceptable, because users don't expect cross-folder ordering.
  • Batch commits — a client uploading 10,000 photos sends 100 commits of 100 files, not 10,000.
  • Admission control per namespace — throttle writes on a hot namespace rather than letting it degrade the whole shard's neighbors.

"I'd enforce a soft limit on namespace size — say 5M entries — and have the product guide admins to split. The alternative is a single-namespace design that works for 99.9% of users and fails for our biggest paying customer."

Deep dive C: The client fleet and mass deletes (6–7 min)

"The riskiest code in this system runs on 100M machines we don't control. A client bug that misreads the filesystem — say, an unmounted external drive looks like 'everything deleted' — would faithfully sync a mass delete to every device and every collaborator."

Staff mechanisms:

  • Server-side delete brake: if a single client commits deletes of > 1,000 files or > 20% of a namespace in 10 minutes, pause propagation and ask the user to confirm.
  • Soft deletes with 30-day retention — every delete is a journal entry, and a namespace can be rolled back to a rev.
  • Staged client rollout: internal → 1% → 10% → 100% over ~2 weeks, with server-side feature gates by client version and a server kill switch that forces old behavior.
  • Minimum supported client version with a server-enforced floor — so the protocol can evolve.

Deep dive D (if time): Bandwidth and the 2 GB file (3–4 min)

"Insert one byte at the start of a 2 GB file: with fixed 4 MB blocks, every boundary shifts, all 500 blocks get new hashes, and we'd re-upload 2 GB. Two fixes. Content-defined chunking in storage would re-upload ~1–2 chunks. Or keep fixed blocks in storage and use rsync-style rolling-checksum delta on the wire: the client computes that almost all bytes already exist at shifted offsets and sends a patch; the server materializes new blocks. I'd do the second — storage stays simple, and the case is rare for typical documents, common only for VM images, databases, and PST files."


Phase 6: Wrap-Up (2–3 minutes)#

Synthesize, don't restate:

"The shape is: an immutable, content-addressed block plane that's a cost problem; a per-namespace ordered journal that's a correctness problem; and a notification plane that's a connection-count problem. The product guarantee — never silently lose an edit — comes from base-revision commits and conflicted copies, and the operational guarantee comes from soft deletes and a delete brake, because the client fleet is where the scary bugs live."

Close with evolution:

"What I'd build later: placeholders so the namespace no longer has to fit on disk; LAN sync so two laptops on the same network exchange blocks directly; real merge for document types by handing them to a collaborative editor; and, at exabyte scale, a build-vs-buy review on the block store — at that size the storage bill is the company's biggest line item."

🎯 Staff Move: End on who owns what. "One team should own the sync protocol end-to-end — client and server. Most of the scary incidents I'd expect are protocol bugs that fall between a client team and a server team, each of which is green on its own dashboards."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
10 min on upload mechanicsS3 multipart, presigned URLs, CDN for downloads"Block plane is content-addressed and idempotent — standard blob storage. Moving on."
No conflict storyWaits for "what if two devices edit?"States base-rev CAS + conflicted copies in Phase 2
Journal as a table"files table with path, owner, updated_at"Append-only journal with monotonic rev and cursor-based pulls
Polling"Clients poll every 30s"Computes 1M QPS of empty polls, proposes long-poll hints
Ignores the clientDesigns server onlySpends 5 minutes on local DB, 3-way reconciliation, and fleet rollout
No numbers"Lots of files""12K commits/s average, 16 TB block index at 1 EB, 200K connections per notification host"

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

File sync is a compact test of whether a candidate can identify the center of gravity of a system. Nearly everyone can design blob upload. Few candidates realize that the upload is the least risky part, and that the hard problems are: (1) establishing a total order of changes per tree, (2) detecting — not resolving — concurrent edits, (3) keeping a fleet of offline, clock-skewed, version-skewed clients convergent, and (4) making destructive operations recoverable.

It also tests ownership across a client/server boundary. Most system design interviews stop at the API. This one can't: half the correctness lives in a desktop binary that the company ships and cannot instantly roll back. A Staff candidate owns both halves.

1.2 The L5 vs L6 vs L7 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 vs L7 Contrast — Visual

The L5 path is not wrong — every box in it is a real component. The gap is emphasis: the L5 path spends its effort on bytes; the L6 path spends it on ordering and failure; the L7 path spends it on the org and the money.

1.3 The Staff Question That Cuts Through Everything#

"When two devices disagree about a file, who decides — and does the loser find out?"

Every major design decision follows from the answer:

  • If the server decides by clock, you've chosen LWW and silent loss.
  • If the server detects via base revision and the client materializes a copy, you've chosen visible conflicts, which requires a per-namespace total order (the journal), which requires a sharding unit (the namespace), which requires a limit on namespace size.
  • If a merge engine decides, you've left file sync and entered Collaborative Editing — a different system for a narrower set of file types.

2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Personal multi-device sync. One human, 2–5 devices. Concurrency comes from offline devices, not from simultaneous humans. Conflicts are rare (a laptop and a phone both edited while offline). The hard problems are bandwidth (mobile uplinks, metered data), battery (don't hash a 10 GB folder on a phone), and convergence latency. Correctness bar: never lose an edit.

Shared team folders. Many humans, many devices, one tree. Concurrency is real: two people open the same spreadsheet. Add permissions (view/edit, inherited ACLs), audit ("who deleted the contracts folder?"), admin controls (remote wipe, legal hold, retention policies), and very large namespaces. Correctness bar: never lose an edit, and every change is attributable. This is where most enterprise revenue — and most escalations — come from.

One-way backup. One writer, no concurrent edits by construction. Optimize for cost per GB and restore correctness: content-defined chunking for better dedup across versions, cold storage tiers, periodic scrubbing, and restore drills. Correctness bar: bit-exact restore, proven by testing restores, not by trusting writes.

🎯 Staff Move: "If this were backup, I'd use content-defined chunking and global dedup within a customer — conflicts don't exist, and dedup ratio is the cost lever. For sync with shared folders, I'll keep fixed blocks and spend the complexity budget on the journal and conflicts instead."

2.2 When NOT to Use File Sync#

SituationWhy sync is wrongUse instead
Real-time co-editing of documentsConflicted copies every few seconds; users expect character-level mergeOT/CRDT editor — Collaborative Editing
Application state (settings, game saves)Tiny structured records; file-level conflicts are too coarseA KV store with per-field merge or a sync SDK
Large media pipelines (render farms, ML datasets)Sync of 50 TB trees to every node is wasteful; access is read-mostly and stagedObject storage + manifests, mount on demand
Source codeUsers need branches, diffs, and merges with intentGit — a purpose-built history model
Databases files (SQLite, PST, VM disks) open while syncingSyncing a live DB file mid-write captures a torn stateApplication-level replication, or exclude/lock these types

🎯 Staff Move: "I'd explicitly exclude live database files and VM images from sync, or sync them only when closed. Syncing a SQLite file mid-transaction is a corruption bug we'd own, and no amount of chunking cleverness fixes it."

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy it mattersWhat I'd assume out loud
Consumer vs businessRetention, audit, legal hold, dedup legalityBusiness shared folders — the harder superset
Max file sizeBlocklist size, upload resumability50 GB per file; blocklist itself paginated
Max files per namespaceJournal hot spots, client memorySoft limit 5M entries; guide splits above
Offline durationJournal retention, cursor expiryCursors valid 90 days; after that, a full re-list
VersioningStorage cost, GC complexity30-day version history (consumer), 180 days (business)
MobileBattery, metered dataMobile is on-demand (placeholders), not full sync
E2E encryptionKills server-side dedup and deltaNot in v1; call out as a product tradeoff

2.4 Precise Terminology#

TermMeaningWhy the precision matters
NamespaceA tree root with its own journal (home folder or shared folder)Unit of ordering, sharding, ACLs, and blast radius
Journal / revAppend-only log per namespace; rev is monotonicEnables cursor sync; provides a total order within a tree
CursorClient's last-applied (ns, rev)"Give me everything after X" — makes reconnection O(changes), not O(files)
BlocklistOrdered list of block hashes composing a file versionA file version = metadata + blocklist; the bytes live elsewhere
Base revisionThe rev the client edited fromMakes commits a CAS; enables conflict detection
Conflicted copyLosing concurrent edit saved as a sibling fileConverts silent loss into visible friction
Three-tree modelBase (last synced), Local (disk), Remote (server)Distinguishes "I changed it" from "they changed it" from "both did"
PlaceholderMetadata-only entry; content fetched on openDecouples namespace size from disk size
Proof of possessionServer challenges client to prove it has the bytes, not just the hashPrevents hash-as-download-token attacks under global dedup
Convergence lagTime from commit to all online devices appliedThe user-visible SLO

3. The Five Fault Lines#

Each fault line below gets: the options, who pays, the Staff default, and when to deviate.

3.1 Fault Line 1: Block Size & Chunking Strategy#

StrategyWhat WorksWhat BreaksWho Pays
Whole fileZero index; trivial1-byte edit to 9 GB file = 9 GB uploadUsers on slow uplinks; egress bill
Fixed 4 MB blocksSimple index; parallel upload; dedups unchanged regions and identical filesInserts shift boundaries → full re-upload of the tailUsers editing large files in place
Fixed 64 KB blocksFine-grained dedup and delta64× more index rows; 64× more hash lookups per GBInfra: block index grows from 16 TB to ~1 PB at 1 EB
Content-defined chunking (avg 1–4 MB)Inserts disturb 1–2 chunks; best dedup for backupVariable sizes; 2–3× CPU on client; harder capacity mathClient CPU/battery; the storage team's mental model
Fixed blocks + wire deltaStorage stays simple; bytes on wire ≈ bytes changedServer must reconstruct blocks from patches; CPU on the block serviceBlock-service team (CPU), in exchange for user bandwidth

Staff default: fixed 4 MB blocks in storage, SHA-256 addressed, plus rsync-style delta as a transfer optimization for files > 64 MB that were modified in place.

"Block size is a cost lever that looks like a performance knob. Halving it doubles the block index and the number of has_blocks lookups; it rarely doubles dedup, because most dedup comes from whole identical files, not from shared sub-regions."

When to deviate: backup products (CDC wins: many versions of the same large files, dedup ratio is the business); media workflows with 50 GB files (bigger blocks, e.g., 16 MB, to cut request count).

🧭 Principal Move: "Block size is a one-way door. Once there's an exabyte stored under 4 MB SHA-256 addressing, changing it means re-chunking and re-hashing everything — a multi-quarter migration with double storage cost during the cutover. I'd make this decision with a cost model, not a benchmark."

3.2 Fault Line 2: Dedup Scope#

ScopeStorage SavingsRiskWho Pays
None0%NoneFinance — the storage bill
Per-namespaceModest: repeated files and unchanged blocks across versionsNone meaningfulFinance pays a bit more
Per-user / per-tenantBetter: same file in multiple foldersTenant-level legal hold and deletion must track refcountsStorage team (refcount correctness)
Global (cross-user)Highest for popular contentExistence side channel: "does anyone have this file?" is observable via upload time/bytes. Hash-as-token: knowing a hash may grant the bytes. Deletion/GDPR requires refcounting across tenants.Security, legal, and the users whose files are probed

The side channel is not theoretical: a 2011 open-source tool (publicly known as "Dropship") demonstrated that Dropbox's cross-user dedup let anyone with a file's hashes retrieve the file, and Dropbox changed the protocol in response.

Staff default: dedup within a tenant (user or business account). Cross-tenant dedup only server-side, after the bytes are uploaded — which kills the bandwidth win but keeps the storage win without the side channel. Or require proof of possession: the server asks for hashes of random byte ranges before accepting a "I already have this hash" claim.

When to deviate: backup of public software/OS images where content is public by nature — global dedup of known-public blocks is fine.

🎯 Staff Move: "Dedup scope is a security decision wearing a cost-optimization costume. I'd get security to sign off on anything wider than per-tenant."

3.3 Fault Line 3: Metadata Consistency Model#

ModelWhat WorksWhat BreaksWho Pays
Single global DB, files tableSimple queries; strong consistencyWrite ceiling of one primary; no natural sync cursorEveryone, at scale
Per-namespace serialized journalTotal order per tree; cursor sync; conflict detection is a CASHot namespaces saturate one shard; cross-namespace moves need a 2-step protocolSync-core team (hot-spot handling)
Per-file eventually-consistent rows (Dynamo-style)Scales writes horizontallyNo ordering across files: a folder rename and a file add race; clients see ghosts and re-sync loopsUsers (weird states), support, client team
Event-sourced journal + derived "current state" indexJournal for sync; index for fast listing and searchTwo representations can driftSync-core team (reconciler + drift alarms)

Staff default: per-namespace journal in a sharded relational store (MySQL/PostgreSQL), sharded by namespace_id, with a derived current-state table updated in the same transaction. The commit is:

BEGIN;
SELECT rev FROM ns_head WHERE ns_id = :ns FOR UPDATE;          -- serialize per namespace
-- if file's current rev != :base_rev → ROLLBACK, return 409
INSERT INTO journal (ns_id, rev, file_id, path, blocklist, ...) VALUES (:ns, head+1, ...);
UPSERT INTO current_files (ns_id, file_id, path, rev, ...) ...;
UPDATE ns_head SET rev = head+1 WHERE ns_id = :ns;
COMMIT;

"The row lock on ns_head serializes commits per namespace — that's the whole point. It caps a single namespace at roughly 1–5K commits/s on one shard, which is fine for 99.99% of namespaces and a known limit for the rest."

When to deviate: at extreme scale, the current-state index can move to a separately scaled store (e.g., a KV store per namespace shard), keeping the journal authoritative.

🧭 Principal Move: "The namespace is the unit of everything: ordering, sharding, cells, ACLs, restore, and blast radius. I'd write that down as an invariant — no feature is allowed to require a transaction across namespaces — because the first feature that violates it turns sharding into distributed transactions."

3.4 Fault Line 4: Conflict Resolution#

StrategyWhat WorksWhat BreaksWho Pays
Last-writer-wins by client mtimeNo user frictionClock skew picks the wrong winner; the loser's work is silently goneThe user who lost work, and never learns why
LWW by server arrival orderNo clock problemStill silent loss; the device that reconnected first winsSame
Base-rev CAS + conflicted copyZero silent loss; testable guaranteeUsers see (conflicted copy) files and must merge by handUsers, with visible friction
Automatic content mergeBest UX for mergeable formatsOnly works for known formats; merge bugs corrupt filesThe team that owns every format's merger
Locking (check-out/check-in)Prevents conflicts for binary formats (CAD, InDesign)Stale locks; offline users can't acquire locksUsers waiting on a colleague who left for vacation holding a lock

Staff default: base-rev CAS + conflicted copies for all opaque files. Offer optional advisory locks for business customers with binary formats. Route truly collaborative document types to a real-time editor rather than trying to merge them in the sync engine.

Conflict edge cases the client must handle:

LocalRemoteResolution
EditEditRemote into place; local → conflicted copy
EditDeleteKeep the edit (resurrect); deletion of unseen work is never silent
DeleteEditKeep the remote edit; local delete is dropped
Rename A→BEdit AApply edit to B (track by file_id, not path)
Create xCreate x (different content)Both kept; one becomes conflicted copy
Case-only rename (File → file) on case-insensitive FS—Normalize path keys per platform; a frequent source of loops

🎯 Staff Move: "The rule I'd give the client team: when in doubt, keep both. Duplicate data is a support annoyance. Lost data is a trust incident."

3.5 Fault Line 5: Client Intelligence & Fleet Control#

DesignWhat WorksWhat BreaksWho Pays
Thin client (upload whole files, server does hashing/diff)Logic is on servers you can roll back in minutesBandwidth and server CPU explode; no offline reconciliationUsers (bandwidth), infra (CPU)
Thick client (watcher, local DB, hashing, delta, 3-way reconcile)Minimal bytes; offline-first; fast convergenceThe riskiest code runs on 100M machines; bugs take weeks to drain from the fleetClient team, and every user during a bad release
Thick client + server guardrailsEfficiency of thick, with server-side brakes: version floors, feature gates, delete brakes, quarantine of misbehaving clientsRequires investing in protocol versioning and fleet telemetrySync-core + client teams jointly

Staff default: thick client with server-side guardrails. The server treats every client as potentially buggy:

  • Validates every commit (blocks exist, paths legal, sizes sane).
  • Enforces a minimum client version and can feature-gate by version.
  • Rate-limits destructive operations per client (delete brake).
  • Can put a single device into quarantine — read-only — when its behavior is anomalous (e.g., > 5 conflicts/minute, re-uploading the same file in a loop).

🧭 Principal Move: "Desktop clients are the one deploy target where 'roll back' takes weeks, not minutes. I'd budget for that explicitly: rollout rings, a server-side kill switch for every new client behavior, and a telemetry pipeline that can detect a bad client version in the 1% ring within an hour."


4. Failure Modes & Operational Reality#

4.1 Mass Delete from a Buggy Client — Full Timeline#

The scariest incident in file sync: the system works perfectly and propagates a wrong instruction everywhere.

t=0:        Client v212.4 rolls to 10% of desktops. A regression treats a temporarily
            unmounted network drive as "all files deleted."
t=+6min:    18,000 devices commit mass deletes. Journal faithfully records them.
t=+6min:    Notification fan-out tells collaborators' devices. Their clients delete local copies.
t=+9min:    client.delete_burst_count (deletes > 1,000 per device per 10 min) crosses 50× baseline.
t=+10min:   Page fires to sync-core on-call. Runbook: enable global delete brake.
t=+12min:   Delete brake on: new bulk deletes held pending user confirmation.
t=+15min:   Client v212.4 halted at 10%; server gate disables the new mount-detection path by version.
t=+2h:      Namespace rollback job restores affected namespaces to pre-incident rev, per-file,
            preserving any *legitimate* edits made after the deletes (a 3-way merge server-side).
t=+6h:      All affected users restored. No data lost because deletes are soft for 30 days.

Detection: client.delete_burst_count by client version; journal.delete_ratio (deletes / total commits, baseline ~5–8%); support ticket spike keyed on "files missing."

Blast radius: every namespace touched by an affected device — and every collaborator of those namespaces. A single device in a 3,000-person shared folder can propagate deletes to 3,000 people.

Mitigation: delete brake, version gate, namespace rollback.

Prevention: delete brake enabled by default (threshold: > 1,000 files or > 20% of a namespace in 10 min); staged rollout with a delete-ratio canary check between rings; client-side sanity check "is the sync root mounted?" before any bulk delete.

Owner: sync-core on-call pages first; client team owns the root cause; support owns user comms.

4.2 Re-Sync Storm After a Cursor or Journal Bug#

t=0:        A migration of journal shard 41 renumbers revs (bug: rev not preserved).
t=+1min:    3.2M clients with cursors on shard-41 namespaces call list_since(cursor).
            Server says cursor invalid → clients fall back to FULL re-list.
t=+2min:    Full re-list of 3.2M namespaces × average 30K entries = ~10¹¹ rows read.
t=+3min:    Metadata DB read replicas at 100% CPU. list_since p99 goes from 80ms to 30s.
t=+5min:    Clients time out and retry — without jitter — doubling load.
t=+8min:    Unrelated shards starve because the API tier's connection pools are full.

Detection: sync.cursor_reset_total (baseline near 0), api.list_since.full_relist_ratio, metadata replica CPU.

Mitigation: server returns 503 Retry-After: <jittered 5–30 min> for full re-lists over a budget; admission control of full re-lists per shard (e.g., max 500 concurrent).

Prevention: revs are never renumbered — treat them as part of the public protocol. Migration dry-runs compare cursor validity on a shadow copy. Clients use exponential backoff with full jitter.

Owner: sync-core (journal), with the client team owning retry behavior.

4.3 Hot Namespace#

A company-wide shared folder with 40M entries and 3,000 writers — plus an automated script that writes 200 files/s — saturates its journal shard.

Detection: journal.commit_latency_p99{shard}, journal.lock_wait_ms{ns} — top-N namespaces by commit rate.

Blast radius: every other namespace co-located on that shard (~thousands) sees commit latency climb from 20ms to seconds.

Mitigation: per-namespace commit rate limit (e.g., 500 commits/s with batching encouraged); move the hot namespace to a dedicated shard (namespace migration is a supported operation — copy journal, flip routing, drain).

Prevention: soft namespace size limits with admin warnings; sub-namespace mount points; API batching for automated writers.

Owner: sync-core owns shard placement; the customer's admin owns the folder structure — a success/support engineer carries that conversation.

4.4 Silent Block Corruption#

A disk returns a bit-flipped block; a client downloads it, verifies SHA-256, detects mismatch — good. But an older client version doesn't verify hashes on download, writes corrupted bytes to disk, and then re-uploads a "modified" file.

Detection: block.download_hash_mismatch_total; background scrubber block.scrub_mismatch_total.

Mitigation: the block store fetches from another replica / reconstructs from erasure coding; quarantine the bad copy.

Prevention: every client verifies hashes on download (version floor enforces it); background scrubbing of all blocks every ~2–4 weeks; end-to-end checksums from client to disk.

Owner: storage team (scrub, repair), client team (verification).

4.5 Notification Plane Partial Outage#

One notification cluster (~200 hosts) loses its connection to the change bus. Clients stay connected but receive no hints.

Detection: this is the silent one — no errors anywhere. Use sync.convergence_lag_p99 measured by synthetic canaries: robot accounts on each cluster that commit and time arrival on a second device. Alert if p99 > 60s for 5 min.

Mitigation: clients also run a slow safety-net poll (every 5–10 min) on the cursor; notification hosts health-check their bus subscription, not just their listen socket.

Owner: realtime/notifications team.

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Buggy client mass deleteclient.delete_burst_count by versionAll collaborators of affected namespacesDelete brake, version gate, namespace rollbackSync-core on-call → client team
Cursor invalidation stormsync.cursor_reset_total, full re-list ratioMetadata tier, all shards via shared poolsRe-list admission control, jittered Retry-AfterSync-core
Hot namespacejournal.lock_wait_ms{ns} top-NCo-located namespaces on the shardPer-ns rate limit, move to dedicated shardSync-core + customer success
Block corruptionblock.scrub_mismatch_total, download hash mismatchFiles referencing that block (dedup multiplies this)Repair from replica/EC; quarantineStorage
Notification lagSynthetic sync.convergence_lag_p99All clients on a cluster (latency only)Safety-net poll; bus health checksRealtime team
Conflict spikesync.commit_conflict_rate > 3× baselineUsers of a client version or a platformVersion gate; investigate clock/rename handlingClient team
Block GC deletes live blockblock.get_404_total for referenced hashesEvery file referencing the blockTombstone grace period; restore from GC quarantineStorage
Permission propagation lagacl.propagation_lag_p99Removed users still syncing a shared folderSync API checks ACL on every list/commit, not just on notifyIdentity/permissions

🎯 Staff Move: "Dedup multiplies blast radius. One corrupted or wrongly-GC'd block can be referenced by a million files across a tenant. That's why the block store needs scrubbing, a GC grace period, and a quarantine rather than immediate deletes."


5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingDesigns "upload and download files"Separates sync / shared / backup intents; commits to oneAsks what storage, identity, and client capabilities the org already has; frames the cost curve at 1 EB
Data modelFiles table with path and S3 keyNamespace journal with monotonic rev, blocklists, cursor syncNamespace as the universal unit (shard, cell, ACL, restore); forbids cross-namespace transactions as a standard
CorrectnessMentions consistencyBase-rev CAS, conflicted copies, rename-by-file_id, idempotent commitsDefines data-loss and conflict rate as product KPIs with owners and quarterly review
FailureReplication, retriesSoft deletes, delete brake, cursor storms, hot namespace, notification lagDesigns rollout rings and kill switches for the client fleet as an org capability; runs mass-delete game days
Scale"Shard the DB"Computes commit QPS, block index size, connection counts; hot-namespace planPrices storage build-vs-buy; plans cells so one region's metadata failure hits < 5% of users
OwnershipOne teamFour planes, four owners, one sync-correctness SLOOne team owns the protocol end-to-end; redraws the client/server boundary

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Finds the center of gravity"Blocks are the easy part. The journal and the reconciliation protocol are where files get lost."
Correctness as a product guarantee"Requirement one: we never silently lose an edit. Conflicted copies are the price."
Quantifies the metadata plane"~12K commits/s average, 40K peak. At 1 EB and 4 MB blocks, the block index is ~250B rows."
Treats the client as untrusted"The server validates every commit and can brake deletes per device — clients will be buggy."
Security instinct on dedup"Global dedup turns hashes into download tokens. Per-tenant unless we require proof of possession."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
LWW by timestamp, unpromptedSilent data loss as a design choice; clock skew ignored
Clients poll every few seconds, no math30M devices × 1 poll / 5s = 6M QPS of empty responses
No story for deletesThe most damaging real incident class is absent
Server-only designHalf the correctness lives in the client; ignoring it is ignoring half the system
20 minutes on S3 multipartSpent the time on the commodity layer

5.4 Common False Positives#

  • Deep knowledge of Rabin fingerprinting ≠ file sync design. CDC math is impressive but rarely the deciding factor for sync.
  • Drawing a CDN for downloads ≠ scale thinking. Blocks are cacheable, sure — the hard scale is metadata and connections.
  • "We'll use CRDTs" ≠ solving conflicts. CRDTs don't exist for arbitrary binary files; saying this for a PSD file is a red flag.
  • Mentioning Magic Pocket ≠ build-vs-buy judgment. The question is when building your own storage pays off, not that Dropbox did it.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing & intent0–3 minPick sync + shared folders; "never silently lose an edit"; scale numbers
Entities & API3–5 minNamespace, journal, block, cursor; base_rev + client_op_id
Architecture5–10 minThree planes; upload + download walk-through
Deep dive 110–20 minConflicts & commit protocol, edge cases
Deep dive 220–30 minJournal sharding, hot namespace, cursor storms
Deep dive 330–40 minClient fleet: deletes, rollout, bandwidth/delta
Wrap-up40–45 minOwnership, evolution, what you'd build later

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response
"Make it end-to-end encrypted."Whether you know what E2EE costsServer can't dedup across users or compute deltas; per-file keys, key rotation on share removal; convergent encryption leaks equality
"Now it's 10× the users."Which plane breaks firstNotification connections and commit rate, not bytes; cellular architecture by namespace
"Mobile on a metered connection."Client-side judgmentPlaceholders, on-demand download, upload only on Wi-Fi by default, thumbnails server-side
"An enterprise wants legal hold."Data lifecycle across dedupHold at namespace/rev level; GC must honor holds via refcount; blocks can't be purged while any held rev references them
"Users complain sync is slow."Measurement before optimizingDefine convergence lag end-to-end via synthetic canaries; find the slow hop

6.3 What to Deliberately Skip#

  • Presigned URL mechanics and S3 multipart details — acknowledge, don't design.
  • Thumbnail/preview generation — a downstream consumer of the change bus; one sentence.
  • Search across file contents — a separate indexing pipeline (Search Indexing).
  • Billing and quota — mention the per-user quota check at commit time, move on.

6.4 Follow-Up Questions to Expect#

  1. "How does a client know which blocks it must download?" — diff the new blocklist against its local block cache (keyed by hash).
  2. "What happens if the client crashes mid-upload?" — blocks are idempotent; the commit hasn't happened; retry resumes by asking missing again.
  3. "How do you garbage-collect blocks?" — refcounts from live and retained revisions; mark-and-sweep with a grace period ≥ max in-flight upload time (e.g., 7 days).
  4. "How do you handle moves across namespaces?" — copy blocklist refs into target (no byte copy), commit in target, then delete in source — two commits, idempotent by client_op_id.
  5. "What about very large directories on the client?" — the local DB, not the filesystem, is the source of truth for sync state; watchers overflow, so periodic reconciliation scans are required.
  6. "How do you migrate a namespace between shards?" — copy journal up to rev N, dual-write/lock tail, flip routing, clients keep cursors because revs are preserved.
  7. "How do you test this?" — deterministic simulation of client + server with randomized network, clock, and filesystem faults — the approach Dropbox publicly described for its Nucleus sync engine.

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design Dropbox."

Staff Answer

"Before drawing anything: is this personal multi-device sync, shared team folders, or backup? They differ on conflicts, permissions, and whether cross-user dedup is even allowed. I'll design sync with shared folders — backup is a subset. My top requirement is that we never silently lose an edit; second, online devices converge in ~10s p99; third, we only upload bytes the server doesn't have. Scale: 50M DAU, ~1B commits/day, ~12K/s average. I'll split the system into a block plane, a metadata journal, and a notification plane, and spend most of our time on the journal, conflicts, and the client fleet — that's where files actually get lost."

Why this is L6:

  • Commits to an intent and a correctness guarantee before a single box
  • Names the center of gravity (journal + client) instead of the upload path
  • Scale is expressed in commits and devices, not just bytes

What L7 adds:

  • "Does the org already run an exabyte object store? If not, storage build-vs-buy is the biggest 3-year cost decision, and I'd size it first."
  • Names who owns the sync protocol across client and server teams
❌ Common L5 Trap

"Clients upload files to S3 through presigned URLs, and we store metadata in a files table. Other clients poll for changes..."

Why this misses: It's a competent upload service. It has no concept of ordering, no conflict story, and polling math that doesn't survive 30M devices. The interviewer will spend the rest of the hour extracting what should have been volunteered.


Drill 2: The Commit Protocol#

Prompt: "Walk me through exactly what happens when I save a 20 MB file."

Staff Answer

"The watcher fires; the client debounces ~1–2s for the app to finish writing, then checks the file is stable (size and mtime unchanged across two reads). It chunks into 5 blocks — four 4 MB and one 4 MB remainder — and SHA-256s each. It calls POST /blocks/missing with 5 hashes; say 2 are missing because only the end of the file changed. It uploads those 2 with PUTs; the server re-hashes to verify. Then it commits {ns, path, blocklist:[h1..h5], base_rev:41, client_op_id}. The server, in one transaction per namespace: checks ACL, checks all 5 blocks exist, checks the file's rev is still 41, appends rev 42, updates current state. It returns 200 with rev 42. If the response is lost, the client retries with the same client_op_id and gets the same rev 42 back — no duplicate. If rev was not 41, it gets 409 and runs the conflict path."

Why this is L6:

  • Handles the "file still being written" problem (debounce + stability check) — a real source of torn uploads
  • Blocks before metadata: the commit is the linearization point and can't reference missing bytes
  • Idempotency and conflict detection are explicit, not implied

What L7 adds:

  • The protocol is versioned; the server can gate new commit fields by client version during a rollout
  • Commit latency p99 is a published SLO (e.g., < 200ms) with the sync-core team as owner

Drill 3: Make the Journal Concrete#

Prompt: "You said 'journal.' What's the schema, and how does a device that was offline for 3 weeks catch up?"

Staff Answer

"Journal rows: (ns_id, rev, file_id, path, blocklist_ref, size, author_id, ts, is_deleted), primary key (ns_id, rev), sharded by ns_id. A derived current_files(ns_id, file_id) table holds the latest state. The device sends its cursor per namespace — say rev 88,412. The server pages through journal WHERE ns_id = ? AND rev > 88412 ORDER BY rev LIMIT 2000. Many of those entries supersede each other — a file edited 40 times needs only its latest version — so for big gaps the server compacts on the fly: it returns the current state of files changed since 88,412 rather than every intermediate entry. If the cursor is older than the journal's retention (say 90 days), the server returns 'cursor expired,' and the client does a full re-list of current_files — admission-controlled, because that's expensive."

Why this is L6:

  • Catch-up cost is O(changes since cursor), not O(files in tree)
  • Compaction on catch-up avoids replaying 40 versions of the same file
  • Anticipates cursor expiry and bounds its cost

What L7 adds:

  • Journal retention is a cost vs. client-experience tradeoff priced in storage $ and full-relist load; set by policy (90 days) with a named owner
  • Full re-list capacity is budgeted as part of disaster recovery planning

Drill 4: The Conflict#

Prompt: "Two laptops, both offline, both edit plan.docx. They reconnect 10 minutes apart. What happens?"

Staff Answer

"Both edited from rev 17. Laptop A commits first with base_rev 17 → rev 18. Laptop B commits with base_rev 17 → 409, server at 18. B's client downloads rev 18 into plan.docx and saves its own bytes as plan (Laptop B's conflicted copy 2026-09-29).docx, which it commits as a new file. Every device now sees both. Nothing was lost; the user merges. I would not try to merge .docx in the sync engine. If the product wants real co-editing, those files open in a collaborative editor, where OT/CRDT handles concurrency at the character level."

Why this is L6:

  • Conflict detection by base revision, resolution by preserving both
  • Explicitly rejects timestamps as arbiters
  • Draws the boundary between sync and collaborative editing

What L7 adds:

  • Tracks sync.conflict_copies_created per file type; if Office files drive 60% of conflicts, that's the business case to fund an integration with a co-editing product
  • Product owns the conflict UX; engineering owns the "no silent loss" guarantee — split written down

Drill 5: The Dependency Goes Down#

Prompt: "The metadata database for 1/16 of namespaces is down for 20 minutes. What do users see?"

Staff Answer

"Users on the other 15/16 of shards see nothing. Users on the affected shard can still read and edit files locally — the client is offline-first. Their commits queue in the local DB; the client backs off with jitter. Blocks still upload, because the block plane is independent — so when the shard returns, commits are small metadata writes. Collaborators on the shard see no changes for 20 minutes. The risk is recovery: millions of queued commits arriving at once. I'd have the sync API admission-control per shard, and the client's retry uses full jitter over 1–5 minutes. We'd alert on journal.shard_available == 0 immediately and on sync.convergence_lag_p99 from canaries."

Why this is L6:

  • Blast radius is bounded by the shard and expressed as a fraction
  • Recognizes offline-first means the client is the degraded mode
  • Plans for the recovery thundering herd, not just the outage

What L7 adds:

  • Organizes shards into cells (metadata + notification + API per cell) so an outage is 1/N of users, not 1/16 of namespaces spread across all API hosts
  • Sets a cell size target (e.g., ≤ 5% of users) as an availability standard

Drill 6: The Hot Namespace#

Prompt: "Our largest customer has one shared folder with 40M files and 3,000 editors. It's slow for them and for others on the same shard."

Staff Answer

"The per-namespace lock caps it at a few thousand commits/s, and the co-located namespaces suffer. Short term: move the namespace to a dedicated shard — namespace migration preserves revs so clients keep cursors — and apply a per-namespace commit rate limit that asks bulk writers to batch. Medium term: split it into sub-namespaces at team boundaries mounted into the parent tree; ordering only matters within a team's folder. Also, 40M files won't fit comfortably in a desktop client's local DB — those users should be on placeholders with selective sync."

Why this is L6:

  • Protects the neighbors first, then the customer
  • Uses the data model (sub-namespaces) instead of just bigger hardware
  • Notices the client-side consequence of namespace size

What L7 adds:

  • Creates a large-customer onboarding review: namespaces > 5M entries get a design review with the customer's admin before migration
  • Prices a dedicated shard into the enterprise tier

Drill 7: Dedup and Security#

Prompt: "We could save 30% storage with global dedup. Why not?"

Staff Answer

"Two problems. First, a hash becomes a capability: if the client says 'I have hash X' and the server links it without receiving bytes, anyone who learns X can get the file. Second, an existence oracle: if uploading a file is instant when someone else already has it, an attacker can test whether a specific document exists anywhere. Mitigations: dedup cross-tenant only server-side after bytes are received — we keep the storage savings but not the bandwidth savings — or require proof of possession by challenging for random byte ranges. Also, global dedup couples tenants for deletion and legal hold via refcounts. I'd take the storage savings server-side and get security to sign off."

Why this is L6:

  • Separates the storage win from the bandwidth win and keeps the safe one
  • Names the historical side channel class, not a vague "privacy concerns"
  • Brings in the right sign-off

What L7 adds:

  • Quantifies: 30% of 1 EB at ~$5–10/TB/month on owned hardware = ~$1.5–3M/month — big enough to justify engineering for server-side dedup, not big enough to justify a security risk
  • Makes E2EE tiers explicitly dedup-free and prices that into the product

Drill 8: Build vs Buy the Block Store#

Prompt: "We're on S3. Finance says storage is our largest cost. Should we build our own?"

Staff Answer

"Only above a threshold, and only with the org to run it. The block plane is ideal for in-house storage: immutable, content-addressed, append-heavy, cold after a few weeks. Dropbox's Magic Pocket is the public precedent. But the break-even requires hundreds of petabytes, a hardware supply chain, datacenter capacity, and a storage team of ~20–40 engineers on call for durability. Below ~100 PB, the team cost and durability risk outweigh the savings. I'd start with tiering on S3 — move blocks untouched for 30+ days to infrequent-access tiers — which often gets 30–40% of the savings at near-zero engineering cost."

Why this is L6:

  • Identifies why this workload suits in-house storage (immutable, content-addressed)
  • Gives a threshold and a cheaper first step
  • Names the operational cost, not just the hardware cost

What L7 adds:

  • A 3-year TCO model with migration cost (double-storage during cutover, egress fees) and a reversibility plan
  • Makes it a board-level capex decision, not an engineering one; see Build vs Buy

Drill 9: Policy Change Without Outage — New Client Behavior#

Prompt: "We want clients to start using content-defined chunking for large files. How do you roll it out?"

Staff Answer

"The server must accept both formats indefinitely, because old clients live for years. So: server first — the block service accepts variable-size blocks and the commit API accepts blocklists with sizes. Then clients behind a server-controlled flag: internal dogfood → 1% → 10% → 50% → 100% over 3–4 weeks. Canary metrics between rings: upload bytes per commit (should drop), client CPU, conflict rate, hash mismatch rate. The flag can be turned off server-side instantly; clients fall back to fixed blocks. Existing files are never re-chunked eagerly — only on next modification."

Why this is L6:

  • Server-first, backward-compatible protocol evolution
  • Server-side kill switch for a client behavior
  • Avoids a mass re-chunk migration

What L7 adds:

  • Establishes the rule: every client behavior change ships dark behind a server flag — written into the client team's release standard
  • Tracks the long tail: sets a minimum supported version so formats can eventually be retired

Drill 10: Multi-Region#

Prompt: "EU customers require data residency. How does the design change?"

Staff Answer

"Residency applies to both planes. Each namespace gets a home region, fixed at creation for business accounts per their contract. The journal shard and the blocks for that namespace live in the home region. Dedup scope becomes per-region at most. Cross-region sharing — a US user in an EU folder — reads through the EU region; commits go to the EU journal. The global pieces are just routing (namespace → region) and identity. I would not do active-active multi-region writes for a namespace; per-namespace ordering is exactly what active-active breaks."

Why this is L6:

  • Residency is a namespace attribute, consistent with the namespace-as-unit model
  • Explicitly rejects multi-master for the journal and says why

What L7 adds:

  • Namespace migration between regions becomes a legal/compliance workflow with approvals
  • Region = cell = blast radius, aligned with legal boundaries

8. Deep Dive Scenarios#

Deep Dive 1: Peak Traffic — Monday Morning Reconnect Storm#

Context: Every Monday at 08:00–09:30 local time, 20M laptops wake up. This Monday, after a weekend in which a large customer bulk-imported 2B files, list_since p99 hits 25s and the notification tier drops connections. The on-call escalates to you.

Questions to Surface First:

  • Is the load from catch-up volume (big cursors gaps) or reconnect volume (connection churn)?
  • Are clients retrying with jitter, or synchronized?
  • Which shards hold the bulk-imported namespaces?
  • Is the degradation global or cell-local?

Typical L5 Approach: Scale out the API tier and add read replicas. Maybe increase the notification cluster size. Treats it as a capacity problem.

Staff Approach: Separates the two loads. Admission-controls expensive catch-ups per shard, returns jittered Retry-After to spread reconnection over 10–15 minutes, and serves compacted catch-up (latest state per file) instead of full journal replay for large gaps. Moves the bulk-imported namespaces' shard off shared pools.

Principal Approach: Asks why a single customer's weekend import could affect Monday for everyone — a cell-isolation failure. Funds cell boundaries that include API pools, not just DB shards, and adds a "bulk import" product path that writes through a throttled pipeline with its own capacity, so large customer operations can't consume the interactive tier's budget.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Enable catch-up admission control: max N concurrent large list_since per shard; others get 503 Retry-After: rand(60, 600)s. Clients remain fully usable locally.
TriageBreak down list_since latency by shard and by gap size. Confirm the hot shards hold the bulk-import namespaces.
Quick fixServe compacted catch-up for gaps > 10K revs. Isolate the affected shard's connection pool.
GuardrailsWatch sync.convergence_lag_p99 from canaries in every cell; confirm other cells recovered.
Post-mortemWhy did bulk import share the interactive path? Why do API pools span cells? Why aren't client reconnects jittered by default?

Metrics to Watch: api.list_since.latency_p99{shard}, api.list_since.gap_size_histogram, notify.connections_open, notify.reconnect_rate, sync.convergence_lag_p99

Organizational Follow-up: Bulk-import product path with its own quota; weekly reconnect load test in staging at 1.5× Monday peak; client team adds jittered wake-up.

Ownership Question: "Who approves a customer importing 2B files?" Staff answer: Nobody should have to — the system should make it safe by default. The bulk path is throttled per tenant, and imports above 100M files trigger an automatic capacity reservation reviewed by sync-core.

Key Takeaway: "Catch-up cost is proportional to the gap, and gaps align on Monday morning. Admission-control the expensive path and compact what you replay."

What clears the Staff bar:

  • Distinguishes connection load from catch-up load
  • Uses compaction and admission control, not only more hardware
  • Traces the cross-customer impact to a missing isolation boundary

Deep Dive 2: Silent Failure — Files That Never Arrive#

Context: Support reports a trickle of "my file never showed up on my other computer." All dashboards are green. Error rates are normal. It has been happening for ~2 weeks.

Questions to Surface First:

  • Is the file in the journal? (server-side truth)
  • Is the receiving device's cursor past that rev? (did it think it applied it?)
  • Which client versions and platforms are affected?
  • Is there a common pattern in paths (unicode, case, length)?

Typical L5 Approach: Asks users to reinstall/re-link. Looks at server error logs. Finds nothing, because nothing errored.

Staff Approach: Treats it as a convergence-correctness bug. Builds a consistency checker: for a sample of devices, compare the device's reported tree hash for a namespace (reported via telemetry) against the server's current state at the same rev. Finds a mismatch pattern: files with Unicode names in NFD form on macOS are "applied" but written to a normalized path, so the local DB thinks they're present while the scanner considers them different files and silently ignores them.

Principal Approach: Makes convergence verification a permanent, org-level SLO: every client periodically reports a Merkle-style tree hash per namespace; the server compares. "Divergent namespaces per million" becomes a tracked metric with an owner and a quarterly target. The org stops relying on users to detect correctness bugs.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Pick 20 affected users; verify the missing files exist in the journal and the devices' cursors are past those revs.
TriageSegment by client version, OS, path characteristics. The NFD/NFC pattern emerges.
Quick fixServer-side flag: force affected client versions to re-verify namespaces with non-ASCII paths (targeted re-scan, not full re-list).
GuardrailsShip client fix through rollout rings; monitor divergence rate by version.
Post-mortemWhy did no metric detect this? Add tree-hash convergence checking; add Unicode normalization to the deterministic test simulator.

Metrics to Watch: sync.divergent_namespaces_per_million{client_version,os}, client.local_apply_skipped_total, support tickets tagged "missing file"

Organizational Follow-up: Convergence checking as a standard; property-based test coverage for path normalization on all platforms.

Ownership Question: "Whose bug is this — client or server?" Staff answer: It's a protocol bug: path identity was underspecified. The team that owns the protocol owns it, which is why the protocol should have one owner across client and server.

Key Takeaway: "In sync, the worst bugs don't error. You need a metric that compares what the client has with what the server says it should have."

What clears the Staff bar:

  • Builds a verification signal rather than chasing logs
  • Finds platform-specific path identity issues (Unicode normalization, case sensitivity)
  • Pushes toward a permanent convergence SLO

Deep Dive 3: Large Customer Onboarding#

Context: A 60,000-employee company is migrating from on-prem file servers: 1.2 PB, 900M files, deeply nested permission structures, cutover planned over one weekend.

Questions to Surface First:

  • What's the largest single folder tree, in files and in writers?
  • What permission model do they have (per-folder ACLs, inherited, deny rules)?
  • How many desktops will link on Monday, and with what default sync settings?
  • What are their compliance needs: legal hold, retention, residency?

Typical L5 Approach: Provision more storage and more DB shards. Run the upload over the weekend.

Staff Approach: Re-plans the migration as a data-model exercise. Maps their shares to namespaces with a cap of ~5M entries each; imports via a server-side bulk path (no desktop clients involved); stages the desktop rollout over two weeks by department; defaults every desktop to placeholders so Monday doesn't pull 1.2 PB × 60,000 devices; pre-provisions dedicated shards for the three largest namespaces.

Principal Approach: Turns it into a repeatable enterprise onboarding program: a namespace-shaping tool, a capacity reservation process, a migration runbook owned by professional services, and a commercial guardrail — contracts above 500 TB include a staged migration schedule. The first bespoke migration becomes the template for the next hundred.

Staff Approach — Full Reasoning
PhaseWhat to Do
Plan (weeks before)Namespace mapping; ACL translation dry run; identify namespaces > 5M entries and split.
Bulk importServer-side ingest at a throttled rate (e.g., 20K commits/s across dedicated shards); blocks dedup within the tenant.
VerificationTree-hash comparison against source file servers per namespace.
Client rollout5% of desktops per day with placeholders; monitor notify.connections_open, list_since load.
Post-cutoverKeep source read-only for 30 days as a fallback.

Metrics to Watch: import.commits_per_sec, import.verification_mismatch, journal.lock_wait_ms{tenant}, client.initial_sync_duration_p95

Organizational Follow-up: Onboarding runbook; tenant-level capacity reservations; customer admin training on namespace structure.

Ownership Question: "Who's accountable if Monday goes badly?" Staff answer: One named migration owner with authority across sync-core, storage and customer success — not a shared channel.

Key Takeaway: "A large customer is a data-model problem before it's a capacity problem. Shape the namespaces, then size the shards."

What clears the Staff bar:

  • Uses placeholders to decouple onboarding from bandwidth
  • Imports server-side rather than through a million client uploads
  • Plans verification and rollback, not just ingestion

Deep Dive 4: Post-Mortem — Garbage Collection Deleted Live Blocks#

Context: 0.002% of files return errors on download: their blocks are gone. The block GC job, rewritten last month for performance, is the suspect.

Questions to Surface First:

  • Were the blocks referenced by a live or retained revision at GC time?
  • Was there an in-flight upload that referenced the block before commit?
  • Do we have the blocks in a GC quarantine?

Typical L5 Approach: Roll back the GC job; restore from backups if they exist.

Staff Approach: Identifies the race: a client uploaded a block (refcount 0, not yet committed), the new GC ran with a 1-hour grace window instead of 7 days, and deleted it before a slow client committed 3 hours later. The commit validation checked block existence at commit time — but a different file already referencing the same hash (via dedup) had been committed, which bypassed the check because the index lookup hit a stale cache. Fix: restore from quarantine; GC grace ≥ max upload session lifetime (7 days); GC reads index from primary; deletes go to a 14-day quarantine before physical deletion.

Principal Approach: Classifies GC as a destructive-change process with the same controls as schema migrations: design review, shadow mode (GC computes deletions but logs only, for 2 weeks, compared against an independent mark phase), and staged enablement per cell. Any system that permanently deletes customer data needs two independent implementations to agree.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediatePause GC globally. Stop physical deletion from quarantine.
TriageList affected hashes; check quarantine; check replicas/erasure-coded fragments not yet reclaimed.
Quick fixRestore blocks from quarantine; re-verify all files referencing them.
GuardrailsRe-enable GC with 7-day grace and 14-day quarantine, shadow mode first.
Post-mortemWhy was grace shortened without review? Why did commit validation use a cache?

Metrics to Watch: block.get_404_total{referenced=true}, gc.deleted_blocks_per_hour, gc.quarantine_restores_total

Organizational Follow-up: Destructive-job review standard; independent mark verification.

Ownership Question: "Who signs off on changes to GC?" Staff answer: Storage owns GC, but any change to grace windows or reachability logic needs sync-core sign-off, because sync-core defines what "referenced" means.

Key Takeaway: "Content addressing makes blocks immutable — not immortal. GC is the one place the block plane can lose data, so it gets the most paranoia."

What clears the Staff bar:

  • Finds the upload-before-commit reachability gap
  • Separates logical deletion from physical deletion with quarantine
  • Brings cross-team sign-off to a destructive operation

Deep Dive 5: Multi-Region Expansion#

Context: The product is single-region (US). Leadership wants an EU region within 12 months for residency and latency.

Questions to Surface First:

  • Residency for all users or for business tenants that request it?
  • Can a namespace span regions (US owner, EU collaborators)?
  • What is global today (identity, billing, sharing links) and must stay global?

Typical L5 Approach: Replicate the whole stack to EU; use multi-master DB replication.

Staff Approach: Home-region per namespace. EU namespaces keep journal and blocks in the EU; cross-region collaborators access via their home region's API, which proxies to the namespace's region. Global control plane: identity, namespace → region routing directory (small, strongly consistent, heavily cached). No multi-master journal: it breaks the per-namespace total order. Migration tool moves namespaces between regions with preserved revs.

Principal Approach: Defines region as a cell with a legal boundary and publishes a residency standard: which data classes can leave the region (none of content; limited metadata like billing), how cross-region shares are logged, and how namespace moves are approved. Sequences the build: EU region with new tenants only in months 1–6, migration of existing tenants in 6–12, with a revenue-weighted migration order.

Staff Approach — Full Reasoning
PhaseWhat to Do
DesignNamespace home region; routing directory; cross-region proxy; per-region dedup.
BuildStand up EU cell; new EU tenants land there.
MigratePer-namespace migration with rev preservation; freeze writes for seconds at cutover.
VerifyTree-hash comparison; audit log of every cross-region access.
OperatePer-region on-call rotations; region-level kill switch.

Metrics to Watch: routing.lookup_latency_p99, crossregion.proxy_latency_p99, migration.namespaces_moved, residency.violations_total (must be 0)

Organizational Follow-up: Legal-approved residency spec; per-region capacity planning.

Ownership Question: "Who owns the routing directory?" Staff answer: Sync-core; it's on the critical path of every request, so it gets the highest availability tier and a static fallback cache on every API host.

Key Takeaway: "Multi-region for file sync means home-region namespaces, not multi-master journals. Ordering lives in exactly one place."

What clears the Staff bar:

  • Keeps the per-namespace ordering invariant
  • Minimizes the global control plane
  • Plans migration with preserved cursors

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Split file sync into a block plane, a metadata journal and a notification plane, and say which scaling law governs each
  • Define a commit protocol that is idempotent (client_op_id), conflict-detecting (base_rev) and resumable (blocks before metadata)
  • Explain why last-writer-wins is a data-loss policy, and defend conflicted copies to a product manager
  • Quantify commits/s, block index size, notification connections and polling cost from a DAU assumption
  • Explain the three-tree client model and the rename/delete/case-folding edge cases
  • Name the mass-delete, cursor-storm, hot-namespace, GC and silent-divergence failure modes, each with a metric and an owner
  • Argue dedup scope as a security decision and block size as a one-way door
  • Describe how you'd roll out a client behavior change to 100M devices you can't roll back

The Bar for This Question#

Mid-level (L4): Designs upload/download with blocks and a metadata DB. Mentions dedup by hash. Handles the happy path correctly. Conflicts, deletes and client rollout come up only when prompted.

Senior (L5): Adds delta sync, dedup, a notification mechanism and replication. The design works. Conflicts are handled by timestamp. The client is treated as a black box. Numbers appear for storage but not for commits or connections.

Staff+ (L6): Frames intents, names "never silently lose an edit" as the governing requirement, and builds the per-namespace journal, base-rev CAS and conflicted copies from it. Treats the client fleet as the riskiest deploy target and designs server-side brakes. Quantifies the metadata plane. Assigns owners to every failure mode. The interviewer should learn something from the answer, whether that's the dedup side channel, NFD/NFC path divergence, or why a GC grace window must exceed the upload session lifetime.


10. Staff Insiders: Controversial Opinions#

10.1 "Delta Sync Is Overrated for Most Users"#

EvidenceImplication
The median synced file is far smaller than one 4 MB blockDelta saves nothing: the whole file is one block anyway
Most large files (video, photos, archives) are written once and never edited in placeFixed blocks plus dedup already avoid re-uploads
The files that benefit (VM images, PSTs, databases) are the ones you shouldn't sync liveDelta optimizes a case you'd rather exclude

The Staff position: ship fixed blocks + dedup first. Add wire-level delta only for files > 64 MB with measured in-place edit patterns. The engineering budget belongs in reconciliation correctness.

Why this matters in interviews: candidates who spend 10 minutes on rolling checksums are optimizing the 1% case while the 99% case (conflicts, deletes) goes undesigned.

10.2 "Conflicted Copies Are a Feature, Not a Failure"#

EvidenceImplication
Users forgive visible duplicates; they don't forgive lost workThe UX cost is asymmetric
Silent loss is found days later, on the wrong device, by the wrong personSupport cost and trust damage compound
Conflict rate is measurable and alertable; silent loss isn'tOnly one of these can have an SLO

The Staff position: product should advertise "we never lose your edits." Conflicted copies are the implementation of that promise.

Why this matters in interviews: defending an "ugly" UX on correctness grounds, and naming who signs off, is a clean Staff-level signal.

10.3 "The Client Is the Most Dangerous Component You Own"#

EvidenceImplication
Server rollback takes minutes; draining a bad desktop version takes weeksBlast radius lasts in time as well as spreading in space
The client runs on filesystems, OSes and clocks you don't controlThe test matrix is effectively infinite
Dropbox publicly rewrote its sync engine primarily to make it testableThe industry's most mature player judged the client its main correctness risk

The Staff position: budget for deterministic simulation testing, rollout rings, server-side kill switches, and a delete brake — before adding features.

Why this matters in interviews: most candidates never mention the client. Mentioning its rollout discipline marks you as someone who has operated one.

10.4 "Global Dedup Is a Security Bug With a Cost Justification"#

EvidenceImplication
Hash-based dedup has publicly been exploited as a file-transfer backchannelHashes become capabilities
Upload-timing reveals whether content exists anywhereExistence oracle for leaked documents
Cross-tenant refcounts entangle deletion and legal holdGDPR erasure becomes a distributed GC problem

The Staff position: per-tenant dedup by default. Cross-tenant only server-side, or with proof of possession.

Why this matters in interviews: it shows you reason about a cost optimization's security surface, not only its savings.

10.5 "Building Your Own Storage Is Almost Always Wrong — Until It's Obviously Right"#

EvidenceImplication
Below ~100 PB the storage team outweighs the savingsRent
At exabyte scale, storage is the largest line item and the workload is ideal (immutable, content-addressed)Build, as Dropbox did with Magic Pocket
The migration itself costs double storage for monthsThe decision needs a 3-year model

The Staff position: write the break-even model down and revisit it yearly. Don't argue it from ideology.

Why this matters in interviews: build-vs-buy questions reward thresholds, not opinions.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

A Staff engineer designs a correct sync system. A Principal engineer sees that file sync is really three businesses sharing one protocol: a storage-cost business (the block plane, whose unit economics decide gross margin), a correctness business (the journal and client, whose failures decide churn and trust), and an enterprise business (permissions, residency, legal hold, whose features decide revenue per seat). Each has a different owner, a different time horizon and a different failure posture. The L7 job is to put the invariants that must never break (per-namespace ordering, no silent loss, soft deletes) into standards, and to let everything else evolve independently.

The Org-Level Fault Line#

One sync protocol owned end-to-end vs. separate client and server teams. Conway's law pushes toward a desktop team, a mobile team, a sync-API team and a storage team, each with its own roadmap. Protocol bugs such as path normalization, cursor semantics or rename handling fall into the seams, and every team's dashboard stays green while users lose files. The alternative, one "sync core" team with authority over the protocol spec, the server journal and the client reconciliation engine, concentrates risk and makes that team a bottleneck.

The Principal position: one team owns the protocol spec and the reconciliation engine (shared code across desktop and mobile, as a library), plus the journal. Platform-specific shell teams own UI, OS integration and the file watchers. Storage stays separate behind a narrow, immutable-block API.

🧭 Principal Move: "I'd draw the team boundary where the invariants are. Everything that can lose a file (reconciliation, the journal, GC reachability) sits under one owner. Everything that can only make sync slower or uglier can be distributed."

Cost Model#

Assumptions: 4 MB blocks; 60% of stored bytes cold after 30 days; object storage at ~$21/TB-month hot, ~$10/TB-month infrequent-access list price (owned hardware ~3–5× cheaper at scale); fully loaded engineer ~$300K/year; metadata on sharded SQL.

ScaleUsers / StorageBlock storage $/monthMetadata + notify infra $/monthHeadcountOn-call load
Startup1M users / 1 PB~$15K (tiered S3)~$10K (1 Postgres primary + replicas, 5 notify hosts)6–8 engineers, one team1 rotation, ~2 pages/week
Growth20M users / 50 PB~$600K (tiered S3)~$150K (64 journal shards, 80 notify hosts, Kafka)40–60 across client, sync-core, storage-on-cloud3 rotations
Hyperscale500M users / 1 EB~$3–4M rented, or ~$1–1.5M owned + ~$10M/yr storage team~$1.5M (thousands of shards, cells per region)300+ incl. 30–40 on storage8+ rotations, per-cell

Reading it: at growth scale, headcount costs more than infrastructure. At hyperscale, the rent-vs-own delta (~$2–3M/month) pays for a storage team several times over. That is the trigger for building your own storage.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversibility Cost
Block size and hash functionOne-wayRe-chunk/re-hash every byte; months of double storage
Namespace as the unit of orderingOne-wayEvery client cursor, shard and ACL model depends on it
rev semantics exposed to clientsOne-wayOld clients live for years; revs can never be renumbered
Dedup scope widening (per-tenant → global)One-way in practiceNarrowing later requires re-splitting refcounts and re-uploading
Metadata store engine (Postgres vs MySQL vs KV)Two-way, but expensivePer-shard migration with preserved revs, ~2–4 quarters
Notification transport (long-poll vs WebSocket)Two-wayHints carry no data, so swap behind a version gate
Object store vendorTwo-way at 1 PB, one-way at 1 EBEgress fees and migration time scale with bytes
Conflict UX (copy naming, prompts)Two-wayClient release

The Standard I'd Write#

RFC: Sync Correctness Standard v1

Scope: All clients and services that read or write the namespace journal or block store.

MUST

  1. Every mutating commit carries base_rev and client_op_id. The server rejects commits missing either.
  2. No operation may silently overwrite a revision the client has not observed. Concurrent edits produce conflicted copies.
  3. Deletes are soft for ≥ 30 days (consumer) / ≥ 180 days (business). Physical block deletion requires a ≥ 7-day GC grace period and a ≥ 14-day quarantine.
  4. No feature may require a transaction across namespaces.
  5. Every new client behavior ships behind a server-controlled flag, through ≥ 3 rollout rings, with a delete-ratio and conflict-rate canary gate between rings.
  6. Clients verify SHA-256 on every block download.

SHOULD

  • Clients report per-namespace tree hashes for convergence verification at least daily.
  • Bulk writers (> 100 commits/s) use batched commits.

Exceptions: Filed with sync-core. Approved by the sync-core lead plus one security reviewer. Exceptions expire after 2 quarters.

Success metrics: zero confirmed silent data-loss incidents per quarter; conflict rate < 0.1% of commits; sync.divergent_namespaces_per_million < 10; p99 convergence < 10s.

What I'd Tell the VP#

"Our product's promise is that files are never lost and always show up everywhere. Storage is the biggest cost but the smallest risk. The real risk is the sync software running on our customers' computers, because a bad release there can delete files for millions of people and takes weeks to undo. I'm asking for three things. First, a single team accountable for sync correctness end-to-end. Second, safety brakes that stop mass deletions automatically. Third, a yearly build-vs-rent review of storage, since at our growth rate owning it could save ~$25–35M a year by year 3. In return we commit to a measurable standard: zero silent data-loss incidents, conflicts under 0.1%, and changes visible on every device within 10 seconds."

Principal Interview Signals#

SignalWhat It Sounds Like
Separates the businesses"Blocks are a margin problem, the journal is a trust problem, permissions are a revenue problem. They need different owners."
Prices the decision"At 1 EB, rent vs own is ~$2–3M/month. That funds a 40-person storage team twice over, so we build."
Names one-way doors"Block size and rev semantics are forever. Notification transport isn't. I'll spend review time accordingly."
Designs the org's failure posture"Cells sized to ≤ 5% of users, a global delete brake, and a quarterly mass-delete game day."
Writes the standard"No cross-namespace transactions, ever. I'd put that in the RFC so no feature team negotiates it away."

Staff answers that L7 interviewers find insufficient:

  • "We'll add a delete brake." Correct, but it doesn't say who can override it, how the threshold is governed, or how it's tested org-wide.
  • "We'd consider building our own storage later." That gives no trigger, no cost model and no migration plan.
  • "The client team and server team will coordinate on the protocol." Coordination isn't ownership. It's how protocol bugs fall between teams.

Appendices

Appendix A: Chunking Mechanics in Depth#

A.1 Fixed-Size Blocks#

def chunk_fixed(file, size=4*MB):
    blocks = []
    while (data := file.read(size)):
        h = sha256(data)
        blocks.append(h)
        local_block_cache.put(h, data_location)   # for later dedup/download
    return blocks                                   # the blocklist

Why it's right for sync: O(1) boundary logic, perfectly parallel hashing, and an easy capacity model (bytes / 4 MB = blocks). Why it's wrong for backup: an insert shifts all later boundaries.

A.2 Content-Defined Chunking#

def chunk_cdc(stream, min=1*MB, avg=4*MB, max=8*MB):
    mask = avg - 1                     # boundary when rolling hash & mask == 0
    start, h = 0, 0
    for i, byte in enumerate(stream):
        h = gear_roll(h, byte)
        length = i - start + 1
        if (length >= min and (h & mask) == 0) or length >= max:
            emit(start, i); start = i + 1; h = 0

Boundaries depend on content, so an insert disturbs only the chunk it lands in and possibly its neighbor. It costs ~2–3× the CPU of fixed chunking, and the size distribution needs min/max clamps.

A.3 rsync-Style Wire Delta#

# Receiver (server) has old version's block signatures:
sigs = [(adler32(b), sha256(b)) for b in old_blocks]
# Sender (client) slides a window over the new file:
for each offset: if adler32(window) in weak_set and sha256(window) matches:
        emit COPY(block_index); advance by block
    else: emit LITERAL(byte); advance by 1   # rolling update is O(1)

Use it only as a transfer optimization. The server reconstructs full 4 MB blocks and stores them content-addressed as usual.

Appendix B: Data Model#

B.1 Tables (per journal shard)#

TableKeyColumnsNotes
ns_headns_idrev, region, quota_usedRow-locked per commit, which serializes the namespace
journal(ns_id, rev)file_id, path, blocklist_ref, size, author, ts, deleted, client_op_idAppend-only; retained 90 days for cursors, longer for version history
current_files(ns_id, file_id)path, rev, blocklist_ref, sizeDerived; unique index on (ns_id, lower(path)) for case-insensitive platforms
op_dedup(ns_id, client_op_id)revTTL 7 days; makes commit retries idempotent
blocklistsblocklist_hash[block_hash...]Large files' blocklists stored once, content-addressed

B.2 Block Index (separate KV store)#

KeyValue
block_hash (32 B)locations[], size, refcount_by_tenant, created_at, last_verified_at

At 1 EB that's ~250B entries at ~64 B each, or ~16 TB, sharded by hash prefix across a KV store (see DynamoDB or Cassandra-style stores). Hash prefixes are uniformly distributed, so there are no hot keys except for extremely popular content.

B.3 Path Identity Rules#

  • Paths are keyed by file_id. The path is an attribute, never an identity.
  • Normalize Unicode to NFC on the server. Keep the client's original bytes for display.
  • Track case-insensitive collisions per namespace and surface them as conflicts on case-insensitive platforms.
  • Reject paths that are invalid on any supported OS (e.g., CON, trailing dots on Windows), or map them reversibly.

Appendix C: Client Reconciliation — The Three-Tree Model#

Diagram: Appendix C: Client Reconciliation — The Three-Tree Model

Rules that prevent the classic bugs:

  • Apply remote changes atomically: write to a temp file in the same directory, fsync, then rename, so apps never see torn files.
  • Advance the cursor only after the local DB commits the applied state. A crash between them causes a re-apply (idempotent), never a skip.
  • Don't trust the watcher alone. inotify/FSEvents queues overflow on large trees, so run a periodic full scan (e.g., every 6–24h) that compares disk to the local DB.
  • Ignore echoes. Changes the client itself wrote into place must not be re-detected as local edits (compare against the expected hash).
Diagram: Appendix C: Client Reconciliation — The Three-Tree Model

Appendix D: API Contract & Client Behavior#

SituationServer responseClient behavior
Commit OK200 {rev}Update base tree + cursor
Stale base409 {server_rev}Pull, reconcile, conflicted copy if needed
Missing blocks412 {missing}Upload them, retry the same op_id
Quota exceeded507Stop uploading, surface UI, keep local edits
Overload / re-list budget503 Retry-After: NFull-jitter backoff up to N; keep working locally
Cursor expired410Full re-list (admission-controlled) and rebuild base
Client too old426Prompt upgrade; read-only mode
Device quarantined423Read-only; upload diagnostics

Retries: exponential backoff with full jitter, base 1s, cap 5 min. Long-poll: 60–90s server hold with ±20% jitter so reconnections don't synchronize.

Appendix E: Observability#

E.1 Core Metrics#

# Correctness (the ones that matter)
sync.commit_conflict_rate                 # conflicts / commits, by client_version, os
sync.divergent_namespaces_per_million     # tree-hash mismatch from client reports
sync.convergence_lag_p99                  # synthetic canary: commit on A → visible on B
client.delete_burst_count                 # devices exceeding delete threshold, by version
# Throughput & health
journal.commit_latency_p99{shard}, journal.lock_wait_ms{ns} (top-N)
api.list_since.full_relist_ratio, sync.cursor_reset_total
block.missing_ratio                       # fraction of hashes needing upload (dedup effectiveness)
block.scrub_mismatch_total, block.get_404_total{referenced=true}
notify.connections_open, notify.reconnect_rate

E.2 Critical Alerts#

AlertThresholdSeverity
block.get_404_total{referenced=true}> 0Page (possible data loss)
client.delete_burst_count> 10× 7-day baseline for 10 minPage
sync.convergence_lag_p99> 60s for 5 min in any cellPage
sync.commit_conflict_rate> 3× baseline for a client versionTicket + rollout halt
sync.cursor_reset_total> 1,000/minPage
journal.commit_latency_p99{shard}> 1s for 10 minWarn

E.3 Debugging "My File Is Missing"#

  1. Is it in the journal? If not, the upload/commit never happened, so check client logs and quota.
  2. Is it deleted in the journal? Check author and device. If it came from a burst, look at the delete brake logs.
  3. Is the device's cursor past the rev? If not, it's a notification or catch-up problem.
  4. Cursor past but file absent locally? That's a reconciliation bug (path normalization, ignore rules, watcher overflow).

Appendix F: Scale Evolution#

F.1 What Works at Each Scale#

ScaleMetadataBlocksNotifications
< 1M usersOne Postgres primary, ns_id in every key from day oneS3, one bucket, hash-prefixed keysLong-poll on 5–10 hosts
1–50M usersSharded by ns_id (64–1,024 logical shards)S3 with lifecycle tieringKafka change bus, 50–100 notify hosts
50M+ usersCells per region; namespace migration toolingConsider owned storage past ~300 PBPer-cell notify clusters

F.2 What You Don't Build on Day One#

  • LAN sync, global dedup, CDC chunking, wire delta, owned storage, multi-region.
  • Do build on day one: ns_id in every key, base_rev + client_op_id, soft deletes, the delete brake, rollout rings. These are expensive to retrofit.

Appendix G: Multi-Tenancy, Fairness & Cost#

  • Per-tenant quotas checked at commit time (bytes counted after per-tenant dedup, so users aren't charged for duplicates).
  • Per-namespace commit rate limits protect shard neighbors (see Rate Limiting).
  • Dedicated shards for the largest tenants as a paid enterprise feature.
  • Cost attribution: bytes stored per tenant, commits per tenant and connected devices per tenant together explain ~90% of cost. Report them monthly to finance so pricing tracks cost drivers.
  1. Loading the index…