Technologies referenced in this case study: DynamoDB · Cassandra · PostgreSQL · ZooKeeper & etcd · Kafka
How to Use This Case Study#
Organized for interview use first, reference second. Read front-to-back once, then return to sections for targeted review.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 3, 4 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Appendix A (durability math) |
| Deep Dive | 3+ hrs | Everything, including The Principal Lens and appendices |
What is Blob Storage? — Why interviewers pick this topic
Blob (object) storage stores opaque byte sequences — 1 KB thumbnails to multi-terabyte backups — under a key in a flat namespace, with a simple API: PUT, GET, DELETE, LIST. It's the durable substrate under data lakes, backups, media, ML datasets, logs, and half of every other system design answer ("…and we put the files in S3").
Before vs After — the "three replicas is enough" launch:
With naive 3x replication, random placement, 24h repair:
t=0: 10 PB logical, 30 PB raw on 5,000 disks; storage bill 3x logical
t=+3mo: Finance flags storage as the #2 infra line item
t=+6mo: A rack PDU fails: 40 disks offline at once
t=+6mo: Random placement means ~thousands of objects had 2 replicas in that rack
t=+6mo+2h: Unrelated disk failure elsewhere; 312 objects lose their last copy
t=+6mo+1d: Customer-visible data loss; "11 nines" marketing page becomes a liability
With erasure coding + failure-domain-aware placement:
t=0: Same 10 PB on zone-aware RS(9,6), 5 fragments per AZ: ~16.7 PB raw (1.67x)
t=+6mo: Same PDU failure: at most 1 fragment per stripe in that rack
t=+6mo: Repair prioritizes stripes with the fewest surviving fragments
t=+6mo+2h: Second failure elsewhere: every stripe still has >= 13 of 15 fragments
t=+1yr: Zero loss; storage bill ~53% lower than 3x
Why interviewers reach for this question: It looks like "a key-value store for big values," so it separates candidates who draw an API gateway in front of a database from candidates who understand that blob storage is two systems with opposite properties — a small, strongly consistent, latency-sensitive metadata plane and a huge, throughput-oriented, failure-tolerant data plane — and that durability is a probability you compute, not a feature you claim.
Mechanics Refresher: Redundancy Schemes
| Scheme | How It Works | Storage Overhead | Tolerates | Repair Cost | Read Latency |
|---|---|---|---|---|---|
| 3× replication | Three full copies on distinct failure domains | 3.0× | 2 losses | Copy 1 object | Best — any replica |
| RS(6,3) | 6 data + 3 parity fragments | 1.5× | 3 losses | Read 6 fragments to rebuild 1 | Needs 6 of 9 (or 1 if unsplit) |
| RS(10,4) | 10 data + 4 parity | 1.4× | 4 losses | Read 10 to rebuild 1 | Needs 10 of 14 |
| RS(9,6), 5 per AZ | 9 data + 6 parity over 3 AZs | ~1.67× | A whole AZ + 1 more fragment | Read 9 to rebuild 1 | Needs 9 of 15 |
| LRC (12,2,2) | 12 data, 2 local parities (each over 6), 2 global parities | ~1.33× | Any 3, many 4-loss patterns | Read 6 for a single loss | Similar to RS |
| Geo-EC (cross-region XOR) | XOR of blocks from two regions stored in a third | ~1.5× across 3 regions vs 2–3× geo-replication | Region loss | Cross-region reads | Degraded reads are cross-region |
For most production systems: Replicate small and hot objects (and the first minutes of every write), erasure-code everything else with a zone-aware code — parity ≥ the fragments in any one AZ, e.g. RS(9,6) at ~1.67× — and use RS(10,4) at 1.4× for single-AZ or reduced-redundancy classes. The code choice is not the interview — placement, repair speed, and what the metadata commit guarantees are.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What This Interview Actually Tests#
Blob storage is not a key-value store with big values. Everyone can map a key to a file path.
It is a durability-economics and metadata-scaling question that tests:
- Whether you separate the metadata plane from the data plane and give each the right consistency model
- Whether you can compute durability from failure rates, repair time, and placement — and know what the math leaves out
- Whether you trade storage overhead (money) against repair bandwidth and tail latency (operations) explicitly
- Whether you own the silent failures: bit rot, orphaned bytes, lost metadata, runaway lifecycle bills
The key insight: The data plane is embarrassingly parallel and almost never the interview bottleneck. The metadata plane — ordered keys, strong read-after-write, LIST on hot prefixes, versioning, and the commit protocol that ties bytes to names — is where designs break and where real systems spend their hardest engineering.
The L5 vs L6 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | API gateway → metadata DB → storage nodes | Asks object size distribution, read/write ratio, durability target, and whether it's a general platform or a single-purpose store | Asks what the org already stores where, the $/GB-month target, and which teams will build on this as a one-way dependency |
| Durability | "3 replicas across AZs, 11 nines" | Computes it: AFR, repair window, placement; knows correlated failures and software bugs dominate the math | Designs the org's durability posture: independent failure domains, deletion protection, versioned control-plane changes, and audits |
| Redundancy | Replication | EC for cold/large, replication for small/hot; names repair bandwidth and tail-latency costs | Prices the switch: 3× → 1.67× at 100 PB avoids |
| Consistency | "Eventually consistent is fine for blobs" | Strong read-after-write via metadata commit as the linearization point; data written before metadata | Treats consistency as a customer contract that can't be weakened later without breaking unknown consumers |
| Metadata | "Store it in Postgres/Cassandra" | Range-partitioned, strongly consistent index with auto-splitting for hot prefixes; LIST cost is a first-class concern | Decides the namespace model (flat vs hierarchical) as a multi-year one-way door |
| Ownership | Storage team owns it | Storage owns durability; tenants own lifecycle policy and access; security owns encryption keys | Chargeback and tiering incentives so teams stop storing 40% garbage |
Why "durability" separates levels
L5: Quotes "eleven nines" and three replicas. Not wrong — three replicas with fast repair really can compute to ~11 nines against independent disk failures.
L6: Shows the calculation, then immediately says what it omits: "With a 1.5% annual disk failure rate and 6-hour repair, independent failures give ~10⁻¹¹ annual loss probability per object for 3× replication. But at 10¹² objects that's still ~10 objects a year, and the real risks are correlated — a rack power event, a firmware bug on one drive model, or a bad deploy that deletes metadata. Durability engineering is mostly about those."
Why "consistency" separates levels
L5: "Blobs are immutable, so eventual consistency is fine." Plausible — the bytes are immutable.
L6: The name → bytes mapping is not immutable. Overwrites, deletes, and LIST after PUT are where applications break under eventual consistency: a data pipeline lists a prefix, misses a just-written file, and silently drops a partition. Making the metadata commit the single linearization point gives strong read-after-write for GET, HEAD, and LIST — which is exactly what S3 moved to in December 2020.
Why "metadata" separates levels
L5: Puts object metadata in a sharded database keyed by hash(bucket, key). Point lookups scale perfectly.
L6: Hash partitioning makes LIST prefix=logs/2026/09/ a scatter-gather across every shard. Ordered listing requires range partitioning by (bucket, key) — which reintroduces hot spots when keys are sequential (timestamps). The answer is range partitioning with automatic split-on-load, plus guidance to customers about key design. Public S3 guidance historically quoted per-prefix rates of ~3,500 writes/s and ~5,500 reads/s, scaling by adding prefixes — the visible artifact of this exact design choice.
The Staff Positions#
| Position | Rationale |
|---|---|
| Separate metadata plane from data plane | Different scale (KB vs PB), consistency (linearizable vs append-only), and failure handling |
| Write data first, commit metadata last | The metadata commit is the atomic visibility point; crashes leave only garbage bytes, never dangling names |
| Erasure-code large and cold objects; replicate small and hot | Zone-aware EC is ~1.67× (RS(10,4) 1.4× in one AZ) vs 3×; small objects don't amortize fragment overhead |
| Placement across ≥ 3 AZs with at most one fragment per failure domain per stripe | Correlated failures dominate the loss math |
| Repair speed is a durability lever | Loss probability scales with (repair time)ᵐ; halving repair time buys more than an extra parity fragment |
| Continuous scrubbing with end-to-end checksums | Silent corruption is found by reading, not by waiting for a customer GET |
| Deletes are soft first | Versioning, delayed GC, and object lock protect against the most common loss cause: humans and bugs |
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| General-purpose multi-tenant object store (S3-like) | Durability, strong consistency, unbounded scale, per-tenant isolation | Separate planes; EC across AZs; range-partitioned metadata; lifecycle tiers | Metadata hot spots; noisy tenants; loss from correlated failure or bug | 11 nines designed durability, 99.9–99.99% availability, read-after-write |
| Media/CDN origin (photos, video) | Read throughput and $/GB served; mostly immutable, write-once | Large-object EC, CDN in front, needle-style packing for small files | Origin overload on cache miss storms; small-file metadata overhead | High durability; availability of hot objects |
| Backup / archive | $/GB-month above all; reads rare and slow is OK | Wide EC, dense disks or tape, cold tiers with hours-long retrieval | Surprise retrieval cost/latency during a restore; silent rot never read | Durability over decades; restore time objective |
🎯 Staff Move: "I'll design the general-purpose multi-tenant store, because it contains the other two as tiers. The hard parts are the metadata plane — strong consistency and ordered LIST at trillions of keys — and the durability math, which I want to compute rather than assert. Media origin and archive become storage classes with different placement and codes on the same metadata plane."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Replication vs Erasure Coding | 3× storage and simple repair, or ~1.4–1.67× storage with expensive repair and tail-latency risk? |
| 2 | Strong vs Eventual Metadata Consistency | Linearizable names (coordination cost) or eventual (application bugs you don't see)? |
| 3 | Durability vs Write Latency | Ack after cross-AZ durable writes (slower PUT) or after local write (loss window)? |
| 4 | Flat vs Hierarchical Namespace | Flat keys (scales, cheap) or real directories (atomic rename, expensive to scale)? |
| 5 | Automatic vs Customer-Driven Tiering | Platform moves data to cheaper tiers (surprises on access) or customers opt in (40% of data never tiered)? |
In the Wild: Real Production Systems#
Why this section belongs here: Citing specific, publicly documented object stores shows you've studied real durability engineering, not just the API.
Amazon S3 — Strong Consistency and ShardStore#
S3 is designed for 99.999999999% (11 nines) durability, stores data across a minimum of three AZs for standard classes, and moved from eventual to strong read-after-write consistency for all PUT, GET, LIST, and HEAD operations in December 2020 at no extra cost. AWS has publicly described the approach (a replicated cache-coherence "witness" that tracks metadata freshness) and published a SOSP 2021 paper on ShardStore, the Rust-based storage-node engine validated with lightweight formal methods.
Staff insight: S3 upgraded consistency after a decade of eventual consistency — because customers built workarounds (S3Guard, EMRFS consistent view) that cost more than fixing it at the source. Weak guarantees don't disappear; they move into every consumer.
Facebook — Haystack and f4#
Haystack (OSDI 2010) packed many photos into large append-only volume files so each read needed ~1 disk seek instead of several filesystem metadata seeks — metadata overhead was the bottleneck, not bytes. f4 (OSDI 2014) moved "warm" blobs to Reed-Solomon (10,4) within a region plus XOR across regions, cutting effective replication from 3.6× to ~2.1× for that tier.
Staff insight: Hot/warm separation is a durability-and-cost decision driven by access-age curves. Most bytes stop being read within weeks — the tiering boundary follows the data.
Microsoft Azure Storage — Local Reconstruction Codes#
Azure's LRC (USENIX ATC 2012) adds local parities over subsets of data fragments so the common case — one fragment lost — repairs by reading 6 fragments instead of 12, while keeping ~1.33× overhead and durability comparable to 3× replication.
Staff insight: The code was designed around the repair path, not the storage overhead. Repair bandwidth is the hidden cost of erasure coding, and it's paid continuously.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Three replicas for 11 nines" | "Show me the math. What's the repair time assumption?" | Can you compute durability? |
| "Erasure coding saves storage" | "What does a single-disk repair cost in network? What about tail latency on reads?" | Do you know EC's hidden costs? |
| "Metadata in a sharded DB" | "How does LIST prefix= work? Sharded by what?" | Range vs hash partitioning |
| "Strong consistency" | "A PUT succeeds and the metadata write times out. What does GET return?" | Commit protocol and linearization point |
| "Multipart upload" | "10,000 parts uploaded, client disappears. Who pays for them?" | Orphan lifecycle and ownership |
| "Delete the object" | "A bug deletes 5% of a bucket. Recover it." | Soft delete, versioning, GC delay |
System Architecture Overview#
Reading the diagram: A
PUTwrites data fragments first (step 2) and commits the metadata version last (step 3); the commit is the single point at which the object becomes visible. The metadata plane is small — ~hundreds of bytes per object — but strongly consistent and range-partitioned for orderedLIST. The data plane is enormous, append-only, and tolerant of losing any disk at any time. The background plane — repair, scrub, GC, lifecycle — is where durability is actually earned, day after day.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Architecture | "API servers, metadata DB, storage nodes" | "Two planes: strongly consistent range-partitioned metadata, append-only erasure-coded data. Data first, metadata commit last." |
| Durability | "Three replicas, 11 nines" | "Zone-aware EC like RS(9,6), 5 fragments per AZ, one per rack, ~6h repair — survives an AZ loss plus a disk. Independent failures compute to far beyond 11 nines; correlated failures and bugs are the real budget." |
| Consistency | "Eventual is fine" | "Strong read-after-write for GET/HEAD/LIST; the metadata commit is the linearization point." |
| Large uploads | "Stream it" | "Multipart: parts 5 MiB–5 GiB, up to 10,000, parallel and retryable; complete is an atomic metadata commit; abandoned uploads expire via lifecycle." |
| Listing | "Query the DB" | "Range-partitioned index, auto-split on hot prefixes; LIST is paginated and ordered." |
| Cost | "Disks are cheap" | "Overhead × $/GB × PB. At 100 PB, 3× → 1.67× is ~133 PB of raw capacity avoided." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Designed durability (S3 standard) | 11 nines | Loss of ~1 object per 10¹¹ per year — at 10¹² objects that's ~10/yr |
| Annualized disk failure rate | ~1–2% | Public fleet data (e.g., Backblaze reports) sits in this range |
| 3× replication overhead | 3.0× | Simple, fast repair, fast reads |
| RS(10,4) overhead | 1.4× | Tolerates 4 losses; single repair reads 10 fragments |
| LRC (12,2,2) overhead | ~1.33× | Single-fragment repair reads 6, not 12 |
| Metadata per object | ~0.3–1 KB | 10¹² objects ≈ ~0.3–1 PB of metadata — itself a distributed database |
| Multipart limits (S3) | 5 MiB–5 GiB parts, 10,000 parts | Max object 5 TiB with the classic limits |
| Per-prefix request rate (S3 guidance) | ~3,500 writes/s, ~5,500 reads/s | Hot prefixes split; scale by spreading keys |
| HDD capacity per drive | ~20–30 TB | Rebuilding one drive at 100 MB/s takes ~2–3 days — why repair must be parallel |
| Parallel repair across 1,000 disks | ~minutes–hours per failed drive | Declustered placement turns a days-long rebuild into hours |
| Scrub cycle | ~1–4 weeks per full pass | Bounds how long latent corruption can hide |
| Deep archive retrieval | ~12–48 hours | Restore-time objective lives or dies here |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on REST endpoints and a capacity table, draw "metadata DB + storage nodes," say "3 replicas," and never reach the commit protocol, the durability math, or
LIST. Compress the API to 2 minutes. The level is decided in the planes split and the durability argument.
Phase 1: Requirements & Framing (2–3 minutes)#
Functional scope in one sentence:
"A multi-tenant object store: PUT, GET (with byte ranges), DELETE, and ordered LIST over a flat key namespace inside buckets, with versioning and lifecycle policies."
Then the non-functional requirements that shape the design:
| Question | Why It Matters | Default I'd Assume |
|---|---|---|
| Object size distribution? | Small objects dominate metadata; large dominate bytes | Median ~100 KB, p99 ~1 GB, max 5 TB |
| Total scale? | Metadata and placement design | 100 PB, 10¹¹–10¹² objects, growing ~40%/yr |
| Durability target? | Redundancy scheme and placement | 11 nines designed, against independent failures |
| Availability target? | Replication of metadata, AZ layout | 99.99% for reads, 99.9% for writes |
| Consistency? | Commit protocol | Strong read-after-write, including LIST |
| Request rate? | Front end and metadata partitions | ~10M req/s aggregate, ~70% GET |
"The two numbers I'll design around are 10¹² objects — that makes metadata a distributed database problem — and 11 nines, which I want to derive rather than assume."
Phase 2: Core Entities & API (1–2 minutes)#
Bucket { bucket_id, tenant_id, region, versioning, lifecycle_rules, policy, storage_class_default }
ObjectVersion { bucket_id, key, version_id, size, etag, storage_class, created_at,
is_delete_marker, manifest_ref, checksum, encryption_key_ref }
Manifest { version_id, chunks: [ {chunk_id, offset, length, stripe_id} ] }
Stripe { stripe_id, code: RS(9,6), fragments: [ {node_id, extent_id, offset, crc} x 15 ] }
Upload { upload_id, bucket_id, key, parts: {n -> {etag, size, chunk_refs}}, initiated_at }
Public API (HTTP):
PUT /{bucket}/{key} -> 200, ETag, x-version-id
GET /{bucket}/{key} Range: bytes=a-b -> 206
DELETE /{bucket}/{key} -> 204 (delete marker if versioned)
GET /{bucket}?list-type=2&prefix=&continuation-token=&max-keys=1000
POST /{bucket}/{key}?uploads -> upload_id
PUT /{bucket}/{key}?partNumber=n&uploadId= -> part ETag
POST /{bucket}/{key}?uploadId= -> complete (atomic)
"The key design decision: the ObjectVersion row is the only thing that makes bytes visible. Stripes and chunks can exist without it — that's garbage, collected later. The reverse must never happen."
Phase 3: High-Level Architecture (≤5 minutes)#
Staff candidates spend under 5 minutes here. Draw the two planes, the write order, and move on.
Say it in four sentences:
- "Stateless API servers authenticate, rate-limit per tenant, and stream the body into chunks — say 8 MB each — erasure-coded into stripes."
- "Placement puts at most one fragment of a stripe in any rack and spreads fragments across three AZs. One catch: RS(10,4) over three AZs puts 5 fragments in some AZ, which exceeds 4 parity — so for AZ-loss survival I'd pick a code whose parity covers the largest AZ share, like RS(9,6) at 1.67×, or a zone-aware LRC."
- "After fragments are durable, the API server does a conditional write to the metadata partition that owns (bucket, key). That commit is the linearization point — GET, HEAD, and LIST see it immediately."
- "Repair, scrubbing, GC, and lifecycle run in a background plane with their own budgets."
🎯 Staff Move: Sentence 2 is where candidates get caught. RS(10,4) over 3 AZs puts ≥ 5 fragments in some AZ, so losing that AZ loses the stripe. The rule: parity count ≥ the largest per-AZ fragment count. RS(9,6) at 5-5-5 costs ~1.67× instead of 1.4× — that 0.27× is the price of surviving an AZ. Say that tradeoff out loud; interviewers write it down.
Phase 4: Transition to Depth (1 minute)#
"The API and the byte path are straightforward. Where I'd like to spend time: first, durability — the math, the code choice, and the correlated failures the math leaves out; second, the metadata plane — strong consistency, ordered LIST, and hot prefixes at 10¹² keys; third, the lifecycle of bytes — multipart uploads, deletes, GC, and tiering, which is where money and data actually get lost. Which would you like first?"
Default order if they don't choose: durability → metadata → lifecycle.
Phase 5: Deep Dives (25–30 minutes)#
Deep dive 1 — Durability math (8–10 min).
- Per-disk annual failure rate λ ≈ 1.5%/yr; repair window T. Loss needs m+1 overlapping failures in a stripe of n fragments.
- Approximate annual loss probability per stripe: P ≈ n·λ · C(n−1, m) · (λT)ᵐ.
- 3× replication (n = 3, m = 2), T = 6 h: 3 × 0.015 × 1 × (0.015 × 6.85×10⁻⁴)² ≈ 4.8×10⁻¹² — ~11 nines.
- Same with T = 24 h: ≈ 7.6×10⁻¹¹ — ~10 nines. Repair time enters squared.
- RS(10,4), T = 6 h (for scale; RS(9,6) is lower still): 14 × 0.015 × C(13,4) × (1.03×10⁻⁵)⁴ ≈ 1.7×10⁻¹⁸. The independent-failure math stops being the constraint.
- "So once I'm on a wide code with fast repair, the durability budget is spent on correlated failures — AZ events, firmware bugs, bad deploys, and operator deletes. That's what I design for next."
Deep dive 2 — Metadata plane (8–10 min).
- Range-partitioned by
(bucket_id, key), each partition a Raft/Paxos group of 3–5 replicas across AZs, ~1–10 GB per partition, auto-split at size or load thresholds. PUT= conditional insert of a new version row;GETreads the latest non-delete-marker version from the partition leader (or a follower with a read lease).LIST= ordered range scan with continuation token = last key returned; pages of 1,000.- Hot prefix: sequential keys (
logs/2026-09-30T12:00:00/...) concentrate writes on the last partition; split-on-load moves the boundary, but monotonically increasing keys always hit the tail. Mitigation: split heat detection + guidance to prefix with a hash; the platform can't fully fix a key design.
Deep dive 3 — Lifecycle of bytes (6–8 min).
- Multipart: parts are independent chunks with their own stripes;
completecommits one version row whose manifest references part chunks — no data copy. - Abandoned uploads: bytes exist without a version row. Lifecycle rule "abort incomplete uploads after 7 days" plus a platform-wide backstop.
- Delete: writes a delete marker (versioned) or tombstone; GC reclaims chunks after a delay (e.g., 24–72 h) so a mistaken delete or metadata bug is recoverable.
- Tiering: lifecycle moves versions to cheaper classes by age; the metadata row changes its
storage_classand manifest atomically after bytes are rewritten.
Phase 6: Wrap-Up (2–3 minutes)#
"To summarize: two planes; data written first and metadata committed last so a crash leaves only garbage, never a broken name; erasure coding across failure domains with repair time as the main durability lever; range-partitioned metadata for strong consistency and ordered LIST; soft deletes and delayed GC because humans and bugs are the biggest loss risk. Next I'd build cross-region replication as a per-bucket policy, object lock for compliance, and cost visibility per tenant so lifecycle rules actually get written."
Common Timing Mistakes#
| Mistake | Time Lost | What to Do Instead |
|---|---|---|
| Designing the REST API in detail | 5–8 min | Name 6 operations; move on |
| Capacity math for storage nodes | 5 min | "100 PB at ~1.67× ≈ 167 PB raw ≈ ~7K 24 TB drives" — one line |
| Explaining Reed-Solomon internals | 5 min | "k data, m parity, any k reconstruct" is enough |
| Skipping LIST | level-deciding | LIST drives the metadata partitioning choice |
| Asserting 11 nines without math | level-deciding | Compute it; then name what the math omits |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Almost every system design ends with "and we store the files in S3." This prompt turns that black box inside out. It tests whether a candidate understands the difference between a system that is correct (every PUT readable) and one that is durable over a decade (every PUT readable after 10 years of disk failures, deploys, migrations, and operator mistakes). The second property is not a feature of any component; it emerges from placement, repair, scrubbing, and process.
It's also an economics question. Storage is a line item that grows monotonically. The candidate who says "3×" at 100 PB has just committed ~133 PB of unnecessary disks versus ~1.67× zone-aware EC — tens of millions of dollars a year — and the Staff answer is expected to notice.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"What is the most likely way we lose a customer's object — and what stops it?"
Answering honestly moves the conversation away from disks. At scale, with a wide code and fast repair, the independent-disk path is ~10⁻¹⁸. The likely paths are: a deploy that corrupts metadata; a GC bug that deletes live chunks; a placement bug that puts two fragments in one rack; a firmware issue that fails 3% of one drive model in a week; a customer's own script deleting a prefix. Each needs a named defense — canarying metadata changes, GC with a delay and a reference-count audit, placement verification in the scrubber, drive-model diversity, versioning and object lock.
🎯 Staff Move: "Once the math says the disks aren't the risk, I spend the durability budget on software and humans: delayed GC, versioning, placement audits, and staged deploys of anything that can touch metadata."
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Intent 1 — General-purpose multi-tenant store. Unknown workloads from thousands of tenants: data lakes doing LIST over millions of keys, ML jobs reading at 100 GB/s, backup tools writing TB-sized objects, web apps serving 10 KB avatars. The design must be strongly consistent, isolate noisy tenants, and offer tiers because one price point can't serve all of them. Correctness bar: read-after-write, designed durability, per-tenant fairness.
Intent 2 — Media/CDN origin. Write-once, read-many, heavy-tailed popularity. Most reads are served by a CDN; the origin must survive miss storms (a viral video, a CDN purge) and keep $/GB low for a long tail nobody reads. Small-file packing (Haystack-style) cuts metadata cost. Correctness bar: durability plus availability of popular objects.
Intent 3 — Backup/archive. Writes are large and sequential; reads are rare and often all-at-once during disaster recovery. Wide codes, dense disks or tape, and retrieval measured in hours are fine — until a restore happens under pressure. Correctness bar: decades-long durability and a restore-time objective someone actually tested.
| Dimension | General-Purpose | Media Origin | Backup/Archive |
|---|---|---|---|
| Dominant cost | Metadata ops + storage | Egress + cache misses | $/GB-month |
| Consistency need | Strong incl. LIST | Weak OK (immutable keys) | Weak OK |
| Redundancy | EC across AZs + replicated small objects | EC, CDN absorbs reads | Wide EC / tape |
| Worst silent failure | LIST misses a new file | Origin melts on miss storm | Rot never read until restore |
| Who complains | Data engineers, app teams | End users (slow media) | Execs, during a disaster |
2.2 When NOT to Build Blob Storage#
| Situation | Use Instead | Why |
|---|---|---|
| You're not a cloud provider | S3 / GCS / Azure Blob | 11 nines needs thousands of disks, 3 AZs, and a repair team; you won't beat their $/GB |
| Data < 1 PB on-prem, regulated | MinIO / Ceph on owned hardware | Mature open source; build only the policy layer |
| Small mutable records | A database | Blob stores have no partial update, no transactions across keys |
| POSIX semantics (rename, append, locking) | A distributed filesystem | Emulating directories on flat keys breaks atomicity |
| Sub-millisecond reads | A cache or KV store in front | First-byte latency is ~10–100 ms by design |
🎯 Staff Move: "If the question were 'should we build this' at a normal company, the answer is no — buy it and spend the effort on lifecycle policy and cost governance. I'll design it as if we're the provider."
2.3 What the Interviewer Leaves Underspecified#
| Underspecified | Why It Matters | What to Say |
|---|---|---|
| Object size distribution | Small objects break EC economics | "Replicate or pack objects < ~1 MB; EC above" |
| Durability definition | Per-object vs fleet-wide loss | "11 nines per object per year, against independent failures; correlated risks designed separately" |
| Consistency of LIST | Drives metadata partitioning | "Strong, including LIST" |
| Multi-region | Replication cost and semantics | "Single region, 3 AZs; cross-region is async, per-bucket opt-in" |
| Delete semantics | Recoverability vs compliance | "Versioning on by default for new buckets; object lock available" |
| Encryption | Key management ownership | "Encrypt at rest always; per-tenant keys optional via KMS" |
2.4 Precise Terminology#
| Term | Precise Meaning |
|---|---|
| Durability | Probability an acknowledged object remains readable over a period (usually per year) |
| Availability | Fraction of requests served successfully; a durable object can be unavailable |
| AFR | Annualized failure rate of a component (disk, node) |
| Stripe | A set of k data + m parity fragments encoding one chunk |
| Failure domain | A set of components that can fail together: disk, host, rack, power zone, AZ |
| Repair window (MTTR) | Time from fragment loss detection to redundancy restored |
| Degraded read | Reading a stripe when some fragments are missing, requiring reconstruction |
| Linearization point | The instant an operation takes effect for all observers — here, the metadata commit |
| Delete marker | A version row that hides earlier versions without removing bytes |
| Orphan | Bytes on storage nodes with no metadata reference (failed PUT, abandoned multipart) |
| Scrubbing | Periodically reading and verifying stored data against checksums |
3. The Five Fault Lines#
3.1 Fault Line 1: Replication vs Erasure Coding#
The tension: Replication buys simplicity, fast reads, and cheap repair with storage money. Erasure coding buys storage money with repair bandwidth, CPU, and tail latency.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| 3× replication everywhere | Simplest code path; any replica serves reads; repair copies one object | ~1.8× more raw disk than zone-aware EC | Finance: at 100 PB, ~133 PB of extra disk |
| Zone-aware EC everywhere (e.g., RS(9,6)) | ~1.67×, survives AZ loss | Small objects waste space (min fragment sizes); every repair reads 9 fragments; degraded reads need reconstruction | Tenants with small objects (latency); network team (repair traffic) |
| Hybrid: replicate on write, EC later | Fast PUT ack; EC once object is cold (hours–days) | Two code paths; transcoding job is a new failure mode | Storage team complexity |
| Hybrid by size: replicate < 1 MB, EC ≥ 1 MB | Small objects stay simple; bytes dominated by large objects get EC savings | Small objects are most of the object count at 3× | Small-object tenants pay more per GB (price it that way) |
| LRC (local parities) | Single-fragment repair reads ~half as many fragments | More complex code, slightly higher overhead than RS of same width | Storage team engineering time |
Staff default: Size-based hybrid. Objects < ~1 MB are packed into larger replicated or EC'd containers (Haystack-style) so tiny objects don't each become a stripe; objects ≥ 1 MB are chunked (e.g., 8–64 MB) and erasure-coded with a zone-aware code. Repair traffic is budgeted as a fraction of cluster bandwidth (e.g., ≤ 10–20%) with priority by fragments remaining.
The hidden cost to say out loud: Losing a 24 TB disk under RS(9,6) means reconstructing 24 TB by reading ~9 × 24 = 216 TB across the cluster. Declustered placement spreads that over thousands of disks, but it's continuous background load — at ~1.5% AFR and 10,000 disks, that's ~150 disk failures a year, ~3 a week.
When to deviate: A latency-critical small-object store (avatars behind no CDN) can stay replicated. Archive classes go wider (e.g., RS(17,3) inside an AZ plus cross-region copies) because repair speed matters less than $/GB.
3.2 Fault Line 2: Strong vs Eventual Metadata Consistency#
The tension: Strong consistency needs a single authority per key (consensus, leader leases); eventual consistency lets any replica answer and pushes anomalies onto every consumer.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Eventual everywhere | Highest availability; cheap multi-replica reads | LIST misses new keys; overwrite returns old bytes; delete reappears | Every data engineer writing a workaround (S3Guard-style) |
| Read-after-write for new keys only | Covers the common PUT-then-GET | Overwrites and LIST still anomalous | Pipelines that overwrite or list |
| Strong read-after-write incl. LIST | Applications behave as with a local filesystem (minus rename) | Every read touches the partition leader or a lease-holding follower; caches need coherence | Storage team: consensus, lease, and cache-invalidation engineering |
| Linearizable multi-key transactions | Atomic rename, conditional multi-object ops | Cross-partition coordination (2PC) on the hot path | Latency for everyone; complexity |
Staff default: Strong read-after-write for single-key operations and LIST, via per-partition consensus with the metadata commit as the linearization point. No multi-key transactions; offer conditional writes (If-Match, If-None-Match) for single-key compare-and-swap. Caches in the front end are validated against a per-partition version (a witness-style freshness check) rather than TTLs.
🎯 Staff Move: "Weak consistency doesn't remove the cost — it moves it into every consumer. With thousands of tenants, I'd rather pay it once in the metadata plane."
3.3 Fault Line 3: Durability vs Write Latency#
The tension: When do we say 200 OK? Every additional durable copy before the ack adds latency; every copy after it is a loss window.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Ack after local write | Lowest latency | Node loss before replication loses acked data | Customers — silently |
| Ack after all fragments durable | Maximum durability at ack | One slow node sets p99 for every PUT | All writers' tail latency |
| Ack after write quorum (e.g., 13 of 15), across ≥ 2 AZs | Durable against AZ loss at ack; tail cut by skipping stragglers | Repair must fill missing fragments quickly; quorum logic in placement | Storage team (straggler repair) |
| Ack after fsync on replicated log, EC later | Low latency, durable | Two-phase data lifecycle; log capacity planning | Storage team complexity |
Staff default: Ack after a write quorum that spans at least 2 AZs and still survives an AZ loss plus one disk (e.g., ≥ 13 of 15 fragments in RS(9,6) with ≥ 4 per AZ). Missing fragments are enqueued for immediate repair. Never ack before data is fsynced somewhere durable — "written to page cache" is not durable.
When to deviate: A reduced-redundancy or single-AZ class can ack faster; price and name it so customers choose knowingly.
3.4 Fault Line 4: Flat vs Hierarchical Namespace#
The tension: Flat keys scale linearly and partition cleanly; directories give atomic rename and cheap listing of children, but require a tree with cross-partition operations.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Flat keys, prefix LIST | Range partitioning; unlimited scale; simple | "Rename directory" = copy + delete of N objects; not atomic | Analytics frameworks that commit via rename (job output committers) |
| Hierarchical namespace (real directories) | Atomic rename, fast directory listing | Directory inode hot spots; cross-partition moves need transactions | Storage team; scale limits per directory |
| Flat + optional hierarchical bucket type | Serves both; customers choose | Two metadata engines to run | Storage team headcount |
Staff default: Flat namespace for the general-purpose store. Document that rename is not atomic and push analytics engines toward commit protocols that write a manifest (table formats like Iceberg/Delta commit by writing a metadata file, not renaming). Offer hierarchical as a separate bucket type only if analytics demand justifies a second engine.
Hot prefixes are the practical fault line inside flat namespaces: sequential keys hit one range partition. Split-on-load helps for spread-out heat; monotonically increasing keys always write the tail partition. Publish key-design guidance and expose per-prefix throttling (503 SlowDown) instead of letting one tenant's hot tail slow its neighbors.
3.5 Fault Line 5: Automatic vs Customer-Driven Tiering#
The tension: Cheaper tiers save 50–95% on $/GB-month but add retrieval cost, retrieval latency, and minimum-duration charges. Who decides when data moves?
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Customer lifecycle rules only | Predictable; customer controls tradeoffs | Most teams never write rules; ~30–60% of bytes sit cold in the hot tier | The company's storage bill |
| Platform auto-tiering by access age | Saves money without customer effort | A restore or backfill hits cold data: surprise latency or retrieval fees | Customer during an incident |
| Auto-tiering among instant-access tiers only | Savings without retrieval latency | Smaller savings than archive tiers | Platform (monitoring cost per object) |
| Default rules on new buckets | Good defaults, opt-out possible | Defaults can be wrong for a workload | Teams that don't read defaults |
Staff default: Automatic tiering only between tiers with the same access latency (monitor per-object access, move after 30–90 days idle); archive tiers with hours-long retrieval are strictly opt-in via lifecycle rules, because the victim of a surprise 12-hour restore is a customer mid-incident. Every bucket shows cost by tier and "bytes not read in 90 days" in its dashboard.
4. Failure Modes & Operational Reality#
4.1 Correlated Failure — The Bad Drive Batch#
t=0: Drive model X (18% of fleet) hits a firmware bug at ~30K power-on hours
t=+2d: Disk failure rate for model X: 1.5% AFR -> ~40% annualized
t=+2d: storage.disk_failures_per_hour 3x baseline; repair queue growing
t=+3d: storage.stripes_at_min_redundancy rises from 0 to 1,200
t=+3d: Page: stripes with <= 1 spare fragment > 100
t=+3d+1h: Repair re-prioritized: fewest-surviving-fragments first; bandwidth cap 10% -> 35%
t=+4d: Placement excludes model X for new writes; proactive drain begins
t=+3wk: Model X drained; zero loss
- Detection:
storage.disk_afr{model},storage.stripes_by_surviving_fragments,repair.queue_age_p99. - Blast radius: Any stripe with multiple fragments on model X. Placement diversity by drive model (not just rack/AZ) caps it.
- Mitigation: Priority repair, raise repair bandwidth, stop placing on the bad model, proactive drain.
- Prevention: Treat drive model and firmware version as failure domains in placement; stage firmware updates by cohort.
- Owner: Storage SRE; hardware team for vendor escalation.
4.2 The Metadata Hot Partition#
t=0: Tenant launches IoT ingest writing keys iot/{timestamp}/{device}
t=+10min: All writes land on the last range partition of the bucket
t=+12min: metadata.partition_qps{p=bucket-7:tail} = 18K vs split threshold 5K
t=+13min: Auto-split: new boundary, but monotonic keys keep hitting the new tail
t=+15min: Partition leader CPU 95%; PUT p99 for that tenant 40ms -> 2s
t=+15min: Other tenants unaffected (partitions are per-bucket ranges)
t=+20min: Tenant receives 503 SlowDown; SDK backs off
- Detection:
metadata.partition_qps,metadata.split_rate,api.throttled_total{tenant}. - Blast radius: One tenant's hot prefix — if partitions are per bucket and leaders are spread. If many hot tails share a leader host, collateral damage.
- Mitigation: Throttle with 503 +
Retry-After; tenant adds a hash prefix; temporary pre-split of the key range. - Prevention: Key-design guidance, per-prefix throttling, leader placement spreading.
- Owner: Metadata team for platform; tenant for key design.
4.3 The GC Bug — Deleting Live Data#
The most dangerous failure in any blob store: a garbage collector that decides live chunks are orphans.
t=0: Deploy changes manifest serialization for multipart objects
t=+1h: GC's reference scan can't parse new manifests; treats their chunks as unreferenced
t=+1h: GC marks 2.1M chunks for deletion; grace period 72h starts
t=+6h: Audit job: gc.marked_bytes_ratio 0.03% -> 1.4% of cluster (anomaly)
t=+6h: Page; GC deletion halted by kill switch
t=+8h: Root cause found; marks cleared; zero bytes actually deleted
- Detection:
gc.marked_bytes_ratiovs baseline, independent reference-count audit (a second implementation that samples chunks and verifies references). - Blast radius without delay: Permanent loss of every multipart object written after the deploy.
- Mitigation: GC kill switch; grace period before physical delete; rate limit on deletions per hour.
- Prevention: GC deletes only after two independent scans agree; manifest format changes are backward-compatible and canaried; physical deletion capped at a fraction of cluster per day.
- Owner: Storage team; GC changes require a second reviewer from the durability owner.
🎯 Staff Move: "The garbage collector is the most dangerous component we own, because it's the only one designed to delete data. I want a delay, a rate limit, a kill switch, and an independent audit on it."
4.4 Silent Corruption — Bit Rot Found Late#
- Symptom: A customer
GETfails checksum on an object last read 14 months ago; reconstruction succeeds from parity. - Root cause: Latent sector errors; one fragment corrupted, never read.
- Detection:
scrub.checksum_failures,scrub.cycle_age_max(time since each disk was fully read),get.reconstructed_reads_ratio. - Mitigation: Scrubber rewrites the fragment on detection.
- Prevention: Scrub cycle ≤ 2–4 weeks; end-to-end checksums computed at the client/API and verified at every hop — not just disk-level CRCs.
- Owner: Storage SRE owns the scrub SLO.
4.5 Multipart Orphans — The Invisible Bill#
- Symptom: A tenant's bill includes 800 TB not visible in any
LIST. - Root cause: A backup tool starts multipart uploads, crashes, restarts with a new upload ID; old parts accumulate for 2 years.
- Detection:
multipart.incomplete_bytes{tenant}, age histogram of open uploads. - Mitigation: Abort uploads older than N days; notify the tenant before abort.
- Prevention: Platform default lifecycle rule "abort incomplete uploads after 7 days" on new buckets; incomplete bytes shown in billing.
- Owner: Tenant for the tool; platform for defaults and visibility.
4.6 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Correlated drive failures | disk_afr{model}, stripes at min redundancy | Stripes on the bad cohort | Priority repair, drain, placement exclusion | Storage SRE + hardware |
| AZ loss | AZ health, fragment availability by AZ | All stripes lose ≤ 5 of 15 fragments | Degraded reads, repair after AZ returns or re-protect | Storage SRE |
| Metadata hot partition | partition_qps, split rate | One tenant prefix | Throttle, split, key guidance | Metadata team + tenant |
| GC deletes live data | gc.marked_bytes_ratio, audit mismatch | Potentially all new objects | Grace period, kill switch | Storage team (durability owner) |
| Bit rot | scrub.checksum_failures, scrub age | Single fragments | Rewrite from parity | Storage SRE |
| Metadata replica divergence | Consensus log mismatch, checksum of partition snapshots | One partition | Rebuild replica from log | Metadata team |
| Multipart orphans | multipart.incomplete_bytes | One tenant's bill | Abort policy | Tenant + platform defaults |
| Accidental customer delete | Delete-rate anomaly per bucket | One bucket | Versioning, restore previous version | Tenant (with platform tooling) |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Architecture | API + metadata DB + storage nodes | Two planes with different consistency; data first, metadata commit last | Plans which planes are shared org infrastructure vs per-product |
| Durability | Quotes 11 nines, 3 replicas | Computes it; names repair time as the lever; designs for correlated failures | Sets durability posture: failure-domain taxonomy, delete protection, audits, game days |
| Redundancy | Replication | Size-based hybrid; zone-aware EC; repair bandwidth budget | Prices code choices in $ and headcount at fleet scale |
| Consistency | Eventual for immutable blobs | Strong read-after-write incl. LIST via metadata linearization | Treats the consistency contract as irreversible once published |
| Metadata | Sharded DB | Range partitioning, split-on-load, per-prefix throttling | Namespace model as a multi-year decision; hierarchical as a separate product |
| Lifecycle | Delete = remove bytes | Soft delete, delayed GC, orphan cleanup, opt-in archive | Chargeback and defaults that change org behavior |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Commit protocol | "The metadata commit is the linearization point; a crash before it leaves garbage, never a dangling name." |
| Durability math | "Loss probability scales with repair time to the power of the parity count, so repair speed is my main lever." |
| AZ-aware coding | "RS(10,4) across three AZs doesn't survive an AZ; parity has to cover the largest AZ share." |
| LIST awareness | "Hash partitioning kills ordered LIST; I'll range-partition and split hot ranges." |
| GC paranoia | "The garbage collector is the only component designed to delete data. It gets a delay and a kill switch." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "11 nines because three replicas" without math | Durability as an assertion, not an engineering property |
| Metadata and data in the same store | Couples a KB-scale strongly consistent problem to a PB-scale throughput problem |
| Hash-partitioned metadata with no LIST plan | LIST becomes scatter-gather over every shard |
| Immediate physical delete | Every bug and fat-finger becomes permanent loss |
| No mention of repair or scrubbing | Durability decays silently without them |
5.4 Common False Positives#
- Reed-Solomon math fluency ≠ storage design. Knowing Galois fields doesn't show you know where to put the fragments.
- Big capacity numbers ≠ scale thinking. "140 PB raw" is arithmetic; "a 24 TB disk rebuild reads ~216 TB cluster-wide" is scale thinking.
- Naming Ceph/MinIO ≠ design. Tools implement choices; the interview is the choices.
- "We use checksums" ≠ integrity. Where they're computed (end-to-end vs per-disk) is the point.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Commit to general-purpose store, 100 PB, 10¹² objects, 11 nines, strong consistency |
| Entities + API | 3–5 min | Version row, manifest, stripe; six operations |
| Architecture | 5–10 min | Two planes; PUT sequence; commit last |
| Deep dive: durability | 10–20 min | Math, codes, AZ-aware placement, correlated failures |
| Deep dive: metadata | 20–30 min | Range partitioning, LIST, hot prefixes, consistency |
| Deep dive: lifecycle | 30–40 min | Multipart, delete/GC, tiering |
| Wrap-up | 40–45 min | Ownership, cost, evolution |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Direction |
|---|---|---|
| "An entire AZ goes down" | AZ-aware placement | Parity ≥ per-AZ fragments; degraded reads; don't mass-repair a transient outage |
| "Customers want 10× more small files" | Small-object economics | Packing into containers; metadata becomes the cost |
| "Add cross-region replication" | Async semantics | Per-bucket, async, replication lag metric, conflict rules for versions |
| "How do you know you haven't lost data?" | Observability of durability | Scrub SLO, reference audits, at-risk stripe counts |
| "Cut storage cost 30%" | Economics | Wider codes for cold data, auto-tiering, orphan cleanup, chargeback |
6.3 What to Deliberately Skip#
- Reed-Solomon encoding math beyond "any k of n."
- HTTP details, signing algorithms — "standard signed requests."
- Disk filesystem internals — "append-only extents on raw devices."
- CDN design — a separate problem; see CDN & Edge Caching.
6.4 Follow-Up Questions to Expect#
- "Walk me through a GET when 2 of the needed fragments are unavailable."
- "How does LIST stay consistent while a partition is splitting?"
- "What happens if the client retries a PUT whose first attempt actually succeeded?"
- "How would you implement object lock (WORM) for compliance?"
- "How long does it take to re-protect after an AZ comes back, and do you repair during the outage?"
- "How do you migrate 50 PB to a new storage-node format without downtime?"
- "How would you bill tenants fairly for requests vs bytes?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design S3."
Staff Answer
"A few questions to pin the design: object size distribution, total scale, durability and consistency targets, and whether this is a general platform or a single-purpose store like a media origin or backup target.
I'll assume the general-purpose multi-tenant store: ~100 PB, 10¹¹–10¹² objects, median ~100 KB with a long tail to terabytes, 11 nines designed durability, strong read-after-write including LIST.
That gives two systems. A metadata plane — ~10¹² small rows, strongly consistent, range-partitioned for ordered LIST. And a data plane — petabytes of append-only, erasure-coded fragments. I'll walk the PUT path to show the commit protocol, then spend most of the time on durability math and the metadata plane."
Why this is L6:
- Commits to an intent that contains the others as tiers
- Splits the problem into two planes before drawing boxes
- Names where the depth will go
What L7 adds:
- Asks which internal teams will depend on this and what $/GB-month the business needs to beat the cloud alternative
- Notes that the consistency contract, once published, can never be weakened
❌ Common L5 Trap
"Clients upload through an API gateway to storage servers. We keep a metadata database with the file location and replicate each file three times."
Why this misses: Correct shape, no commitment. The interviewer asks "what if the metadata write fails after the bytes are stored?" and "how does LIST work?" — and the design has no answer yet because it never separated the planes.
Drill 2: Compute the Durability#
Prompt: "Prove 11 nines."
Staff Answer
"Assume disk AFR λ = 1.5%/yr and a repair window T of 6 hours (6.85×10⁻⁴ years). Data is lost only if m+1 fragments of one stripe fail within overlapping repair windows.
For 3× replication: P ≈ 3λ · (λT)² = 0.045 × (1.03×10⁻⁵)² ≈ 4.8×10⁻¹² per object per year — about 11.3 nines. If repair takes 24 hours instead, it's ~7.6×10⁻¹¹ — only about 10 nines. Repair time enters squared.
For RS(10,4): P ≈ 14λ · C(13,4) · (λT)⁴ ≈ 1.7×10⁻¹⁸. Wide codes make independent disk failure a non-issue.
Two caveats. First, 'per object' vs fleet: at 10¹² objects, 10⁻¹¹ still means ~10 lost objects a year. Second, this assumes independent failures. Real loss comes from correlated events — a rack, a drive batch, an AZ — and from software. So the design work is placement across failure domains and protecting against our own bugs."
Why this is L6:
- Derives the number and shows the sensitivity to repair time
- Distinguishes per-object from fleet-level expectations
- Names what the model leaves out and pivots there
What L7 adds:
- Frames durability as an org-level posture with explicit budgets for software-caused loss, verified by independent audits
❌ Common L5 Trap
"Each disk has maybe a 1% chance of failing, so three replicas give 1% × 1% × 1% = one in a million."
Why this misses: It ignores repair — the failures must overlap within the repair window, which is what makes the number 10⁻¹¹ instead of 10⁻⁶. Without repair time in the model, the candidate can't reason about the most important durability lever.
Drill 3: The PUT Commit Protocol#
Prompt: "The bytes are written but the metadata write times out. What does a GET return?"
Staff Answer
"The metadata commit is the linearization point, so the answer depends on whether the commit actually happened — a timeout doesn't tell us.
The API server returns 500 or a timeout to the client. It does not know whether the version row committed. Two cases:
- Commit didn't happen: GET returns the previous version (or 404). The fragments are orphans; GC reclaims them after its grace period.
- Commit happened: GET returns the new version, even though the client saw an error.
Either way, the client retries. For safety, each PUT carries a client-generated request ID or Content-MD5; a retry of an already-committed PUT with the same content produces a new version with identical bytes — harmless. With If-None-Match: * (create-only), the retry gets 412 and the client can HEAD to confirm.
What never happens: a version row pointing at fragments that don't exist, because we only commit after the write quorum of fragments is durable."
Why this is L6:
- Treats timeout as 'unknown', not 'failed'
- Uses write ordering to make the only failure residue garbage
- Gives the client a concrete idempotency path
What L7 adds:
- Makes 'data first, metadata last, GC with delay' an invariant tested by fault-injection in CI for every storage release
❌ Common L5 Trap
"We roll back — delete the bytes that were written."
Why this misses: Rolling back on timeout races the commit: if the metadata write actually succeeded, deleting the bytes creates a visible object with missing data — the one failure mode the protocol exists to prevent.
Drill 4: LIST at Trillions of Keys#
Prompt: "How does
LIST prefix=logs/2026/work with 10¹² objects?"
Staff Answer
"Metadata is range-partitioned by (bucket_id, key), so all keys with that prefix are contiguous across one or a few adjacent partitions. LIST is an ordered range scan starting at the prefix, returning up to 1,000 keys and a continuation token that encodes the last key returned. The next page starts strictly after that key — which also makes pagination stable across partition splits, because the token is a key, not a partition offset.
Delimiter queries (delimiter=/) roll up common prefixes; the scan skips ahead past each rolled-up prefix instead of reading every key under it.
Consistency: each page reads from the partition's leader (or a follower with a valid read lease), so a key committed before the LIST started is visible.
Cost: a LIST page is ~one partition read — far more expensive than a GET's point lookup, so it's priced and rate-limited separately."
Why this is L6:
- Connects LIST to range partitioning
- Continuation token as a key survives splits
- Notes delimiter skip-scan and prices LIST distinctly
What L7 adds:
- Offers an inventory/manifest export for tenants who LIST billions of keys daily — changing an expensive access pattern rather than scaling it
❌ Common L5 Trap
"Query the metadata database with WHERE key LIKE 'logs/2026/%'."
Why this misses: On a hash-sharded store this is a scatter-gather across every shard, sorted in memory, for every page. It works in a demo and melts at 10¹² keys.
Drill 5: The Hot Prefix#
Prompt: "One tenant writes 50K PUT/s to keys starting with the current timestamp. What happens?"
Staff Answer
"Every write lands on the tail partition of that bucket's key range — monotonically increasing keys defeat split-on-load, because after each split the new tail is hot again. A partition leader handles maybe 5–10K writes/s, so the tenant sees throttling.
Platform response: return 503 SlowDown with Retry-After, per prefix, so the throttling is confined to this tenant; ensure partition leaders for hot tails aren't co-located on the same metadata host.
The real fix is key design, and it belongs to the tenant: prefix with a short hash (a7f3/2026-09-30T...) so writes spread over N ranges, or use a device ID first. We publish that guidance and show per-prefix throttling in their dashboard.
I would not make hash-partitioning the default to 'fix' this — it destroys ordered LIST for everyone."
Why this is L6:
- Explains why split-on-load fails for monotonic keys
- Contains blast radius to one tenant
- Assigns the fix to the right owner
What L7 adds:
- Adds key-design review to the onboarding checklist for tenants above a request-rate threshold
❌ Common L5 Trap
"Add more metadata shards."
Why this misses: More shards don't help a single hot key range; the tail is still one partition.
Drill 6: Multipart Upload#
Prompt: "A customer uploads a 2 TB file over a flaky connection."
Staff Answer
"Multipart upload. The client initiates and gets an upload ID, splits the file into parts (say 256 MB each — about 8,000 parts, under the 10,000 limit), and uploads parts in parallel, e.g., 16 at a time. Each part is independently chunked, erasure-coded, and stored; the server returns an ETag per part. A failed part is retried alone.
Complete sends the ordered list of (part number, ETag). The server validates the parts and commits one version row whose manifest references all part chunks — no copying of 2 TB. The object's ETag is a hash of part ETags plus a part count, which is why multipart ETags aren't the MD5 of the file.
Abandoned uploads: parts exist without a version row and cost money. A lifecycle rule aborts incomplete uploads after 7 days; the platform sets it by default on new buckets and shows incomplete bytes in billing."
Why this is L6:
- Right part sizing against the limits
- Completion is a metadata-only atomic commit
- Owns the orphan-cost failure mode
What L7 adds:
- Makes 'incomplete multipart bytes' a line item in every team's chargeback, so the cost has an owner
❌ Common L5 Trap
"Stream it in one request with resumable offsets."
Why this misses: A single stream is serial and ties a server to one connection for hours; resuming requires server-side partial state. Parallel parts are faster and each part's failure is independent.
Drill 7: Small Objects#
Prompt: "80% of objects are under 16 KB. Does erasure coding still work?"
Staff Answer
"Not per object. A 16 KB object in RS(9,6) becomes 15 fragments of ~1.8 KB — each with its own metadata, index entry, and disk IO. Overhead and IOPS dominate.
Pack small objects into large containers — e.g., 1 GB append-only volumes (Haystack-style). The metadata row stores (volume, offset, length). Volumes are sealed when full and then erasure-coded as a unit. Until sealed, the open volume is replicated 3× so writes are fast.
Deletes in packed volumes leave holes; compaction rewrites volumes when garbage exceeds ~30–50%. That's a background cost, budgeted like repair.
The metadata plane is now the bottleneck: at 10¹² small objects it's ~0.3–1 PB of rows. Pricing should reflect it — per-request and per-object fees matter more than per-GB for small-object tenants."
Why this is L6:
- Quantifies why per-object EC fails for small objects
- Packing with replicate-then-seal-then-EC
- Names compaction cost and the metadata bottleneck
What L7 adds:
- Aligns pricing with cost drivers (requests and object count), so small-object workloads fund their own metadata scale
❌ Common L5 Trap
"Use the same erasure coding — it's just smaller fragments."
Why this misses: Ignores per-fragment fixed costs; the IOPS and metadata bill grows 15× for no durability gain.
Drill 8: Build vs Buy#
Prompt: "We run 8 PB on-prem. Build our own object store?"
Staff Answer
"No. At 8 PB, the choices are a public cloud object store or a mature open-source system like Ceph RGW or MinIO on owned hardware. Building a new one means writing an erasure-coding data plane, a consensus-backed metadata plane, repair, scrubbing, and GC — and then carrying a team for it forever.
The decision between cloud and self-hosted: egress patterns (if data is read heavily from on-prem compute, cloud egress fees dominate), regulatory residency, and whether we have ~3–5 engineers for storage operations. At 8 PB, cloud standard storage is roughly $150–200K/month list; self-hosted hardware amortized over 5 years at ~1.5× EC is cheaper per GB, but the people cost closes most of the gap.
What we should build either way: lifecycle defaults, cost attribution per team, and a backup copy in an independent system."
Why this is L6:
- Rejects the build at this scale with reasons
- Compares the realistic options on the right axes
- Identifies the part that's worth building
What L7 adds:
- Plans the exit: data formats and tooling that keep a later cloud-to-on-prem (or reverse) migration feasible
❌ Common L5 Trap
"Build it — we'll have full control and save cloud costs."
Why this misses: No TCO; ignores that durability engineering is a permanent team, not a project.
Drill 9: Changing the Storage Format Without Losing Data#
Prompt: "We need to migrate 50 PB from 3× replication to erasure coding."
Staff Answer
"It's a background transcode with a strict invariant: never remove the old copies until the new stripe is verified and the metadata points at it.
Per object: read a replica, encode, write fragments, verify checksums by reading back a sample, atomically swap the manifest in the metadata row (conditional on the old manifest), then enqueue old replicas for GC after the usual grace period.
Rollout: shadow first — transcode 0.1% and run the scrubber and a read-verify on everything transcoded; then ramp by storage class and age (coldest first, since they're read least and save the most). Rate-limit by cluster bandwidth — maybe 5–10% — so customer latency doesn't move. At 50 PB and ~20 GB/s of spare bandwidth, it takes ~1 month.
Metrics: transcode progress, verify failures (must be 0), manifest-swap conflicts, GET p99 for transcoded vs not."
Why this is L6:
- Explicit safety invariant and atomic manifest swap
- Coldest-first ordering with a bandwidth budget
- Calculates the duration
What L7 adds:
- Makes the savings visible: 50 PB × (3.0 − 1.67) ≈ 66 PB of raw disk freed — a hardware purchase avoided that funds the project
❌ Common L5 Trap
"Write a script that converts each file and deletes the old replicas."
Why this misses: Deleting before verifying and before the metadata points at the new stripe creates a window for permanent loss.
Drill 10: Cross-Region Replication#
Prompt: "A tenant needs a copy in another region for disaster recovery."
Staff Answer
"Per-bucket asynchronous replication. After a local commit, the version is appended to a replication log partitioned by (bucket, key); a replicator copies the bytes and commits the same version ID in the destination bucket. Ordering per key is preserved; versions make it idempotent.
Guarantees to state: RPO equals replication lag — typically seconds, p99 minutes, unbounded during a regional incident. Metric replication.lag_seconds{bucket} and the backlog in bytes; an optional SLA tier (e.g., 99.99% of objects within 15 minutes) is a product decision.
Deletes: replicate delete markers by default, but not permanent version deletes — otherwise a malicious or buggy delete in the source destroys the DR copy. That's the point of DR.
Cost: storage in both regions plus inter-region transfer per GB — make it visible."
Why this is L6:
- Async with explicit RPO and a lag metric
- Idempotency via version IDs
- Delete semantics designed for the DR purpose
What L7 adds:
- Requires DR copies of critical data to be in a separately administered account with different credentials — protection against compromised admins, not just regions
❌ Common L5 Trap
"Write synchronously to both regions."
Why this misses: Every PUT pays a cross-region round trip (~60–150 ms), and a regional outage blocks writes everywhere — availability sacrificed for an RPO most tenants didn't ask for.
8. Deep Dive Scenarios#
Deep Dive 1: Peak Traffic — The Viral Object Miss Storm#
Context: A video embedded in a news story goes viral. The CDN in front of the media bucket is purged by an unrelated config change. 400K req/s hit the origin for one 200 MB object and ~5K related objects. Origin p99 GET climbs from 80 ms to 9 s, affecting every tenant in the cell.
Questions to Surface First:
- Is this one object or broad traffic? Which component saturates — API servers, the metadata partition, or the storage nodes holding those fragments?
- Why did the CDN purge, and can we restore the cache?
- Are other tenants in the same cell affected, and why?
Typical L5 Approach: Scales out API servers. The bottleneck is the 15 storage nodes holding that object's fragments — each serving hundreds of Gbps — so more API servers add more load on the same disks.
Staff Approach: Identifies a hot-object problem, not a capacity problem. Adds a hot-object cache tier in the front end (in-memory, keyed by version ID so it's always consistent), request coalescing so concurrent GETs of one object share one backend read, and per-tenant request isolation so the media tenant's storm doesn't consume shared API capacity. Works with the CDN team to restore caching with a staggered warm-up.
Principal Approach: Asks why a CDN purge could be global and instant. Pushes for purge rate limits and staged purges as a CDN platform standard, and for cell-based isolation in the storage front end so one tenant's storm has a bounded blast radius by construction.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Identify hot keys from front-end sampling; enable request coalescing and hot-object caching for the top 100 version IDs |
| Triage | Confirm saturation on the fragment-holding nodes; check cross-tenant impact |
| Quick fix | Per-tenant concurrency limits; CDN re-enable with origin shield |
| Guardrails | Watch storage.node_egress_gbps and GET p99 for other tenants |
| Post-mortem | Hot-object cache as a permanent tier; purge governance; cell isolation |
Metrics to Watch: api.get_qps_by_version_topk, storage.node_egress_gbps, api.latency_p99{tenant}, cdn.origin_hit_ratio
Organizational Follow-up: CDN team adds staged purges; storage publishes an origin capacity contract per tenant.
Ownership Question: "Who owns origin protection — CDN or storage?" Staff answer: Both, with a written contract: CDN guarantees staged purges and origin shielding; storage guarantees per-tenant isolation so a miss storm degrades only the offending tenant.
Key Takeaway: "Immutable versions make caching trivially consistent — use that. The bottleneck of a hot object is its fragments, not the API tier."
What clears the Staff bar:
- Locates the real bottleneck (fragment holders)
- Uses version IDs for safe caching
- Contains blast radius per tenant
Deep Dive 2: The Silent Failure — Scrubber Stopped Six Weeks Ago#
Context: During a routine review you notice scrub.cycle_age_max is 44 days against a 14-day SLO. Nobody was paged. A config change moved the scrubber to a lower-priority scheduling class and it's been starved by repair and transcode jobs.
Questions to Surface First:
- How many fragments haven't been verified within SLO, and on which disks?
- Were there latent errors in the last completed scrubs, and at what rate?
- Why did a 3× SLO violation not page?
Typical L5 Approach: Restores the scrubber's priority and lets it catch up.
Staff Approach: Restores priority, then computes exposure: with a historical latent-error rate, how many stripes might now have hidden bad fragments plus a missing one? Prioritizes scrubbing stripes that already have a missing fragment (where a hidden error would mean loss) and disks with elevated SMART errors. Adds a page on
scrub.cycle_age_max > 1.5× SLOand a floor on scrubber bandwidth that no scheduler class can take away.
Principal Approach: Recognizes a class of problem: background durability jobs (scrub, repair, GC audits) have no customer-visible symptom when they stop. Creates a "durability SLO" review for all background jobs, with dead-man alerts that fire when a job stops reporting, and reports these SLOs to leadership monthly like availability.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Restore scrubber priority with a guaranteed bandwidth floor |
| Triage | Rank stripes: missing fragments + unscrubbed first |
| Quick fix | Targeted scrub of at-risk stripes within 48 h |
| Guardrails | Dead-man alert on scrub progress; SLO page threshold |
| Post-mortem | Scheduler changes that affect durability jobs require durability-owner review |
Metrics to Watch: scrub.cycle_age_max, scrub.bytes_verified_per_day, scrub.checksum_failures, storage.stripes_by_surviving_fragments
Organizational Follow-up: Durability job inventory with owners and dead-man alerts.
Ownership Question: "Who approves changes to background job scheduling?" Staff answer: The storage durability owner — any job that protects data is part of the durability budget, and its capacity is not up for grabs by other jobs.
Key Takeaway: "Durability jobs fail silently. Alert on their absence, not just their errors."
What clears the Staff bar:
- Computes exposure before declaring victory
- Prioritizes by risk, not by order
- Fixes the alerting class, not just the instance
Deep Dive 3: Large-Customer Onboarding — 5 PB and 20 Billion Objects#
Context: An analytics customer will migrate 5 PB in 20B objects over 60 days, then run jobs that LIST ~200M keys per hour and read at 200 GB/s bursts.
Questions to Surface First:
- What's their key layout? Sequential, hashed, partitioned by date?
- Do they need LIST or could an inventory manifest replace it?
- Peak write rate during migration? Will they use multipart for large files?
Typical L5 Approach: Checks total capacity and adds storage nodes.
Staff Approach: Capacity is the easy part. Reviews key design with them (date-partitioned with a hash component), pre-splits their metadata ranges based on expected key distribution, provides daily inventory exports so their jobs stop LISTing 200M keys an hour, sets per-tenant request limits appropriate to their tier, and places their data across enough storage nodes that 200 GB/s bursts don't concentrate. Migration throttled to protect other tenants.
Principal Approach: Turns this into a repeatable "large tenant onboarding" program with a checklist (key design review, pre-split, inventory, limits, cost forecast) and a capacity reservation contract. Negotiates pricing that reflects metadata and request costs, not just bytes.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Design review | Key layout; LIST patterns; object size distribution |
| Pre-work | Pre-split metadata ranges; reserve capacity; inventory export |
| Migration | Throttled ingest (e.g., 20B objects / 60 days ≈ 3,900 PUT/s average, cap at 10K) |
| Validation | Object count and checksum reconciliation against source |
| Steady state | Per-tenant dashboards: throttles, LIST cost, bytes by tier |
Metrics to Watch: metadata.partition_qps{tenant}, api.throttled_total{tenant}, list.keys_scanned_per_hour{tenant}, storage.cell_utilization
Organizational Follow-up: Account team and storage agree on capacity reservation lead times (e.g., 90 days for > 1 PB).
Ownership Question: "Who decides this customer's request limits?" Staff answer: The storage platform sets limits by tier with a documented exception process; the account team can request, not grant.
Key Takeaway: "Large tenants break metadata, not disks. Onboard their key design, not just their bytes."
What clears the Staff bar:
- Focuses on metadata and access patterns over raw capacity
- Replaces expensive LIST with inventory
- Protects other tenants during migration
Deep Dive 4: Post-Mortem — A Tenant Lost 3% of a Bucket#
Context: A tenant reports 3% of objects in a bucket return 404. Investigation shows their own IAM role, compromised via a leaked CI credential, issued DELETE on 1.2M keys. Versioning was off. GC ran after 24 hours. The data is gone.
Questions to Surface First:
- Was the platform working as designed? (Yes.)
- What defaults allowed an unversioned production bucket?
- Why could one credential delete 1.2M objects without any friction?
Typical L5 Approach: "Customer error; recommend they enable versioning."
Staff Approach: Accepts that "working as designed" isn't a defense when the design makes a common mistake permanent. Changes defaults: versioning on for new buckets; anomaly detection on delete rate per bucket (1.2M deletes in 10 minutes vs a baseline of 100/day) that alerts the tenant; a longer GC grace period for bulk-deleted data (e.g., 7 days) during which support can restore.
Principal Approach: Frames it as a platform-level safety posture: soft-delete by default, object lock for critical data, deletion requiring separate credentials for production buckets, and a recommended cross-account backup. Writes the "data protection defaults" standard and measures adoption across all buckets.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Check whether any fragments remain un-GC'd; halt GC for the bucket |
| Triage | Confirm deletes came from a valid credential; engage security |
| Quick fix | Recover anything still within grace; support the tenant's restore from their own backups |
| Guardrails | Delete-rate anomaly alerts; extended grace for bulk deletes |
| Post-mortem | Default versioning; delete-protection tiers |
t=0: Leaked credential used from unfamiliar IP
t=+3min: DELETE rate on bucket: 2,000/s (baseline ~0.001/s)
t=+10min: 1.2M objects deleted; no alert exists for delete rate
t=+24h: GC reclaims chunks
t=+3d: Tenant notices missing files from user reports
Metrics to Watch: api.delete_rate{bucket} vs baseline, gc.bytes_reclaimed{bucket}, security.credential_anomalies
Organizational Follow-up: Security owns credential anomaly detection; storage owns delete safety defaults.
Ownership Question: "Is this the customer's problem or ours?" Staff answer: The credential leak is theirs; the fact that one credential could permanently delete 1.2M objects in 10 minutes with no alert and a 24-hour grace is ours.
Key Takeaway: "The biggest durability risk is a valid DELETE. Design defaults so the common mistake is recoverable."
What clears the Staff bar:
- Refuses 'working as designed' as a closing statement
- Changes defaults, not just documentation
- Separates security and storage responsibilities cleanly
Deep Dive 5: Multi-Region Expansion — Launching a New Region#
Context: The company is launching the object store in a new region with 3 AZs. Initial demand is 2 PB; tenants want a global namespace for bucket names and cross-region replication from day one.
Questions to Surface First:
- What must be global (bucket-name uniqueness, IAM) vs regional (metadata, data)?
- Is the new region's metadata plane independent of existing regions?
- Can we launch with a smaller cluster without weakening durability?
Typical L5 Approach: Extends the existing metadata cluster to the new region for a single global namespace.
Staff Approach: Keeps every region's metadata and data planes fully independent — a region failure must not affect others. Only the bucket-name registry is global (low write rate, strongly consistent, cached everywhere). Replication is async and per bucket. Small initial footprint uses the same code but checks that enough racks and nodes exist per AZ for "one fragment per failure domain" — if not, launches with replication and transcodes later.
Principal Approach: Codifies regional independence as an architectural rule ("no cross-region calls on the data path") and tests it with region-isolation game days. Plans the capacity ramp and the hardware cost curve, and makes sure the global control plane (bucket names, IAM) has its own availability design, since it's the only shared fate.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Scope | Global: bucket registry, IAM. Regional: everything else |
| Durability at small scale | Verify placement can honor failure domains; otherwise replicate, then EC as capacity grows |
| Replication | Async per bucket; lag SLO; delete-marker semantics |
| Validation | Region-isolation test: sever links, confirm local PUT/GET/LIST unaffected |
| Ramp | Capacity plan for 2 PB → 20 PB over 2 years |
Metrics to Watch: replication.lag_seconds, registry.cache_staleness, region.cross_region_calls_on_data_path (must be 0)
Organizational Follow-up: Each region has its own on-call; the global control plane has a separate team with a higher availability target.
Ownership Question: "Who owns the global bucket registry?" Staff answer: The control-plane team, separate from any region, with a design that tolerates its own outage: existing buckets keep working, only new bucket creation stops.
Key Takeaway: "Regions share only what they must, and what they share must fail gracefully."
What clears the Staff bar:
- Independent regional planes; minimal global state
- Durability honored even at small launch scale
- Tests isolation rather than assuming it
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Split blob storage into a metadata plane and a data plane and justify their different consistency models
- Walk the PUT path and explain why data-first, metadata-last makes crashes leave only garbage
- Compute durability for replication and erasure coding, and explain why repair time is the main lever
- Choose a code whose parity covers the largest per-AZ fragment count
- Explain range partitioning for LIST, why monotonic keys defeat split-on-load, and who fixes it
- Design multipart upload, soft delete, delayed GC, and orphan cleanup
- Name the correlated and software-caused failures that dominate real data loss
- Price redundancy and tiering choices at three scales and name the one-way doors
The Bar for This Question#
Mid-level (L4): Designs an upload service that stores files on servers with a metadata table. Mentions replication. Struggles with large uploads and listing.
Senior (L5): Separate metadata DB and storage nodes, 3× replication across AZs, multipart upload, CDN for reads. A system that works — but durability is asserted, LIST is an afterthought, deletes are immediate, and cost is 1.8× what it needs to be.
Staff+ (L6): Two planes with an explicit commit protocol; durability computed with repair time; zone-aware EC with a size-based hybrid; range-partitioned strongly consistent metadata with hot-prefix handling; soft deletes, delayed GC with a kill switch, scrubbing SLOs; each failure mode with a metric and owner. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "Eleven Nines Is a Marketing Number — the Real Risk Is Software"#
| Evidence | Implication |
|---|---|
| Independent-failure math for wide codes gives ~10⁻¹⁸ | Disks are not the bottleneck on durability |
| GC, metadata, and deploy bugs can touch every object | Software failures are correlated by nature |
| Customer deletes are valid operations | The platform can't distinguish mistake from intent without friction |
The Staff position: Spend durability engineering on delete safety, GC paranoia, staged deploys, and audits — not on adding parity.
Why this matters in interviews: Candidates who say "more replicas" to every durability question signal they've never seen a real data-loss incident.
10.2 "Eventual Consistency Was Never Cheaper — It Was Just Billed to Someone Else"#
| Evidence | Implication |
|---|---|
| Ecosystem tools built consistency layers on top of eventually consistent S3 for years | Every consumer paid the cost separately |
| S3 moved to strong consistency without a price increase | The provider could absorb it once |
| LIST-after-PUT anomalies silently drop data in pipelines | The cost shows up as data quality bugs |
The Staff position: Pay for consistency once in the metadata plane.
Why this matters in interviews: It reframes a "performance vs consistency" debate as "who pays."
10.3 "Repair Speed Matters More Than Parity Count"#
| Evidence | Implication |
|---|---|
| Loss probability ∝ Tᵐ | Halving T with m = 2 gives 4× improvement |
| 24 TB drives take days to rebuild serially | Declustered, parallel repair is mandatory |
| Extra parity costs storage forever | Faster repair costs bandwidth only during failures |
The Staff position: Declustered placement and prioritized repair before wider codes.
Why this matters in interviews: It shows you understand the durability model rather than memorizing codes.
10.4 "Most Blob Storage Cost Is Garbage"#
| Evidence | Implication |
|---|---|
| Incomplete multipart uploads, old versions, and never-read logs accumulate indefinitely | Bytes grow monotonically without lifecycle policy |
| Few teams write lifecycle rules unprompted | Defaults decide the bill |
| Access-age curves show most data goes cold within weeks | Tiering savings are large and mostly untaken |
The Staff position: Defaults (abort multipart after 7 days, noncurrent-version expiration, auto-tier among instant tiers) and chargeback beat any storage-engine optimization.
Why this matters in interviews: Cost questions reward governance answers, not compression answers.
10.5 "Don't Offer Rename"#
| Evidence | Implication |
|---|---|
| Atomic rename across partitions needs distributed transactions | It taxes the metadata plane for one operation |
| Table formats commit via manifests, not renames | The ecosystem already moved on |
| Emulated rename (copy + delete) is non-atomic and misleads users | A fake guarantee is worse than none |
The Staff position: Flat namespace, no rename; offer hierarchical namespaces as a separate product only when demand justifies a second engine.
Why this matters in interviews: It shows you'll decline a feature to protect the core design.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
A Staff engineer designs a durable, consistent object store. A Principal engineer sees that object storage is the organization's system of record by accident: every team dumps data into it, few write lifecycle rules, nobody owns cross-team retention, and the bill grows 40% a year while half the bytes are never read. The durability problem is solved; the governance problem isn't. The L7 question is how to make the org's storage spend and data-protection posture correct by default — and which decisions (namespace, consistency contract, key formats, vendor) will be impossible to reverse once 200 services depend on them.
The Org-Level Fault Line#
One storage platform with enforced defaults vs team-owned buckets with full autonomy.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Full team autonomy | Teams move fast; no platform gate | No lifecycle rules, unversioned prod buckets, inconsistent encryption, unowned buckets after reorgs | Finance (runaway bill), security (exposure), data teams (loss) |
| Central platform with mandatory defaults | Safe and cheap by default | Some workloads fight the defaults; exception process becomes a bottleneck | Platform team (exceptions), outlier teams |
| Paved road: defaults + chargeback + audits | Defaults protect, chargeback motivates, audits catch drift | Requires cost attribution tooling and a policy engine | Platform team headcount (~2–3 engineers) |
The L7 default: The paved road. Every bucket is created through a platform module with versioning, encryption, lifecycle defaults, an owning team tag, and a data classification. Chargeback is per team per month. Quarterly audits flag unowned and non-compliant buckets.
🧭 Principal Move: "I'm not trying to make storage cheaper by engineering the engine — I'm making it cheaper by making every byte have an owner and a lifecycle. That's where the 30–40% is."
Cost Model#
Assumptions: self-hosted at scale; ~$6/TB-month for raw HDD capacity all-in (hardware amortized over 5 years, power, space, network); ~1.67× overhead for zone-aware EC; metadata and front-end fleet ~20–30% of data-plane cost; fully loaded engineer ~$25K/month. Public-cloud standard storage list price for comparison is on the order of ~$21–23/TB-month.
| Scale | Logical Data | Infra $/month | Headcount | On-call Load | Cloud List Equivalent |
|---|---|---|---|---|---|
| Small | 1 PB | ~$40–60K (minimum failure-domain footprint dominates; raw disk alone ~$10K) | 4–6 (~$125K) | 1 rotation, frequent hardware tickets | ~$22K + requests — buy |
| Medium | 100 PB | ~$1.3–1.5M (raw ~$1.0M, metadata/front end ~$0.3–0.5M) | 20–30 (~$600K) | Storage SRE + metadata rotations | ~$2.1M + requests + egress |
| Large | 1 EB | ~$12–15M | 80–150 (~$2–4M) | Multiple rotations, hardware ops team | ~$20M+ — building can pay off |
The number that matters to executives: below roughly tens of petabytes, building your own object store costs more than buying once people are counted. Above hundreds of petabytes, with steady growth and low egress needs, owning can save ~30–50% — but only with a team that treats durability as a permanent product.
The 3-Year Evolution Path#
Triggers: Year 1 when storage becomes a top-3 infra cost; Year 2 when the first tenant needs DR or compliance retention; Year 3 when a second region launches or a drive generation reaches end of life.
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversal Cost |
|---|---|---|
| Consistency contract (strong read-after-write) | One-way | Weakening breaks unknown consumers; can only be strengthened |
| Namespace model (flat vs hierarchical) | One-way | Migrating semantics of billions of keys and every client |
| Public API shape and ETag semantics | One-way | SDKs and customer code depend on it |
| Durability claims published to customers | One-way | Can't lower a promise without trust damage |
| Erasure code parameters | Two-way (expensive) | Background transcode of all data; months of bandwidth |
| Placement policy | Two-way | Rebalance over weeks |
| Metadata storage engine | Two-way (expensive) | Online migration with dual-write and verification |
| Tiering thresholds | Two-way | Config change |
The Standard I'd Write#
RFC: Object Storage Data Protection & Cost Standard v1
Scope: All buckets holding production or customer data.
MUST:
- Create buckets through the platform module; every bucket has an owning team, a data classification, and a cost center tag.
- Enable versioning and encryption at rest on production buckets; noncurrent versions expire after a class-defined period (default 30 days).
- Abort incomplete multipart uploads after 7 days.
- Separate deletion privileges from write privileges for production buckets; bulk deletes above 10K objects/hour trigger an alert to the owner.
- Tier-0 data (customer records, financial ledgers) has a cross-account backup with object lock.
SHOULD:
- Use lifecycle rules to move data not read in 90 days to infrequent-access tiers.
- Use hashed or high-cardinality key prefixes for write rates above 1K/s.
Exceptions: Filed with the storage platform team; approved by platform lead and the data owner's director; reviewed every 6 months.
Success metrics: 100% of buckets owned; < 5% of bytes in incomplete uploads or orphaned versions; storage cost growth ≤ data growth; zero unrecoverable deletions of Tier-0 data.
What I'd Tell the VP#
Our storage is extremely unlikely to lose data because of hardware; the realistic risks are our own software bugs and accidental or malicious deletions, and our defaults today don't protect against those. I'm proposing a standard that turns on versioning and delete protection for production data and gives every bucket an owner. The same change attacks cost: today roughly a third of what we pay for is data nobody reads or has forgotten about, and chargeback plus lifecycle defaults should cut storage spend growth by 25–35% within a year. It needs about three engineers for two quarters. We should keep buying storage from our cloud provider until we're well past a hundred petabytes; building our own engine is not where the savings are.
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Governance over engine | "The bill is driven by unowned bytes, not by the erasure code." |
| Priced tradeoffs | "Zone-aware EC costs 0.27× more than RS(10,4) — at 100 PB that's ~27 PB of disk, the price of surviving an AZ." |
| One-way doors | "The consistency contract and the namespace model are forever; the code parameters aren't." |
| Build vs buy with thresholds | "Below tens of petabytes, buy; the team costs more than the disks." |
| Org failure posture | "Tier-0 data gets a cross-account, object-locked copy — protection against our own admins." |
Staff answers that L7 interviewers find insufficient:
- "We'll add versioning" — right mechanism, no plan to make it the default across 2,000 buckets.
- "Erasure coding saves 44%" — correct math, no decision about whether the org should run the engine at all.
- "Storage team owns durability" — ignores that deletes, lifecycle, and retention are owned by hundreds of tenant teams.
Appendices
Appendix A: Durability Math in Depth#
A.1 The Model#
For a stripe of n fragments tolerating m losses, with per-disk failure rate λ (per year) and repair window T (years):
P_loss_per_year ≈ n·λ · C(n−1, m) · (λT)^m
Read it as: rate at which some fragment fails (n·λ), times the probability that m of the remaining n−1 also fail before repair completes.
| Scheme | n | m | T = 6 h | T = 24 h |
|---|---|---|---|---|
| 3× replication | 3 | 2 | ~4.8×10⁻¹² | ~7.6×10⁻¹¹ |
| RS(6,3) | 9 | 3 | ~8×10⁻¹⁵ | ~5×10⁻¹³ |
| RS(10,4) | 14 | 4 | ~1.7×10⁻¹⁸ | ~4×10⁻¹⁶ |
| RS(9,6) | 15 | 6 | ≪ 10⁻²⁰ | ≪ 10⁻¹⁸ |
(λ = 0.015/yr. Order-of-magnitude only; the model assumes independent failures and exponential lifetimes.)
A.2 What the Model Omits#
| Omitted Factor | Effect | Defense |
|---|---|---|
| Correlated hardware failure (rack, PDU, drive batch) | Multiple fragments fail together | One fragment per failure domain; drive-model diversity |
| Latent sector errors | A "surviving" fragment is actually bad | Scrubbing; end-to-end checksums |
| Repair backlog during mass failure | T grows exactly when failures cluster | Repair priority by surviving fragments; bandwidth headroom |
| Software bugs (GC, metadata, placement) | Can affect all objects at once | Delays, audits, canaries, kill switches |
| Operator and customer error | Valid deletes | Versioning, object lock, grace periods |
A.3 Copysets — Per-Object vs Fleet Loss#
With random placement across thousands of disks, almost any combination of m+1 simultaneous disk failures holds all surviving fragments of some stripe — so fleet-level "we lost something" events become frequent even when per-object probability is tiny. Restricting placement to a limited number of disk groups (copysets, Cidon et al., 2013) makes loss events rarer but larger. Choose deliberately: many small losses (random placement) or rare larger ones (copysets). Most object stores prefer fewer loss events, because each one is an incident with customer notification.
A.4 Repair Bandwidth#
bytes_read_per_disk_loss = disk_capacity × k (RS: read k fragments per rebuilt fragment)
LRC single-loss read = disk_capacity × group_size
24 TB disk, RS(9,6): ~216 TB read cluster-wide
Declustered over 2,000 disks at 50 MB/s each for repair: ~100 GB/s -> ~36 min
Appendix B: Metadata Data Model#
Table object_versions -- range-partitioned by (bucket_id, key)
PK: (bucket_id, key, version_id DESC)
cols: size, etag, checksum_crc64, storage_class, is_delete_marker,
manifest (inline for <= 8 chunks, else pointer), created_at, kms_key_ref
Table uploads -- partitioned by (bucket_id, key, upload_id)
parts: map<part_no, {etag, size, chunk_refs}>, initiated_at
Table chunks -- partitioned by chunk_id (hash)
stripe_id, offset, length, refcount_hint
Table stripes -- partitioned by stripe_id (hash)
code, fragments[ {node_id, extent_id, offset, crc} ], sealed_at
Byte lifecycle — the states that the commit protocol and GC must respect:
Version IDs are time-ordered (e.g., hybrid logical clock + node ID) so "latest" is the first row in the key's range.
Appendix C: Coordination Mechanisms#
| Mechanism | Used For | Notes |
|---|---|---|
| Raft/Paxos per metadata partition | Linearizable version commits | 3–5 replicas across AZs; see Distributed Consensus |
| Leader leases / read leases | Serve strong reads without a consensus round | Bounded clock drift assumption |
Conditional writes (If-Match) | Single-key CAS for clients | No multi-key transactions |
| Partition manager with fencing | Splits, merges, leader moves | Fencing tokens prevent split-brain writes |
| Placement service | Fragment targets | Stateless over a cluster map; membership via ZooKeeper & etcd |
| Witness/version check for caches | Front-end cache coherence | Cache entries validated against partition version |
Quick comparison: A single global metadata database doesn't scale past ~10⁹–10¹⁰ rows comfortably; hash-sharded KV breaks LIST; range-partitioned consensus groups are the standard answer (the same shape as Spanner, CockroachDB, or TiKV).
Appendix D: API Contract and Client Behavior#
| Behavior | Contract |
|---|---|
| Idempotency | PUT of identical content is safe to retry; If-None-Match: * for create-only |
| Integrity | Client sends Content-MD5 or a CRC checksum header; server verifies before commit |
| Throttling | 503 SlowDown with Retry-After; SDK uses exponential backoff with jitter |
| Large objects | Multipart above ~100 MB; parts 5 MiB–5 GiB; ≤ 10,000 parts |
| Range reads | Range: bytes=a-b; parallel ranged GETs for throughput |
| Consistency | Strong read-after-write for PUT/DELETE → GET/HEAD/LIST |
| Pagination | continuation-token is opaque, key-based, stable across splits |
Appendix E: Observability#
E.1 Core Metrics#
Durability: storage.stripes_by_surviving_fragments, repair.queue_age_p99,
scrub.cycle_age_max, scrub.checksum_failures, gc.marked_bytes_ratio
Metadata: metadata.partition_qps, metadata.commit_latency_p99, metadata.split_rate
API: api.latency_p99{op}, api.error_rate{op}, api.throttled_total{tenant}
Cost: storage.bytes{tenant,class}, multipart.incomplete_bytes, versions.noncurrent_bytes
Data path: storage.node_egress_gbps, get.reconstructed_reads_ratio
E.2 Critical Alerts#
| Alert | Threshold | Action |
|---|---|---|
| At-risk stripes | stripes with ≤ 1 spare fragment > 100 | Page |
| Scrub stalled | scrub.cycle_age_max > 1.5× SLO or no progress in 1 h | Page |
| GC anomaly | marked bytes > 3× baseline | Page + auto-pause GC |
| Metadata commit latency | p99 > 100 ms for 10 min | Page |
| Delete anomaly | bucket delete rate > 100× baseline | Notify owner |
E.3 Control Plane vs Data Plane#
Control plane: bucket creation, policy, lifecycle config, placement maps, partition management — may be briefly unavailable. Data plane: GET/PUT/LIST on existing buckets — must keep working with cached control-plane state. A control-plane outage should block new bucket creation, not reads.
E.4 Debugging "Is Any Data at Risk Right Now?"#
storage.stripes_by_surviving_fragmentshistogram — anything at minimum?- Repair queue age and throughput — is repair keeping up?
- Failure-domain view — are failures clustered (rack, model, AZ)?
- GC and scrub health — both running, both within SLO?
Appendix F: Scale Evolution#
F.1 What Works at Each Scale#
| Scale | Architecture |
|---|---|
| < 100 TB | Buy cloud storage. If self-hosted: MinIO/Ceph, replication |
| 100 TB–10 PB | Open-source object store; EC for cold data; lifecycle defaults |
| 10–100 PB | Dedicated storage team; zone-aware EC; small-object packing; chargeback |
| 100 PB+ | Custom metadata plane, declustered repair, hardware co-design, multi-region |
F.2 Multi-Region Path#
Independent regional planes; a global bucket registry and IAM; async per-bucket replication with lag SLOs; no cross-region calls on the data path (Deep Dive 5).
F.3 What You Don't Build on Day One#
- Hierarchical namespace
- Archive tiers with hours-long retrieval
- Cross-region replication
- LRC or custom codes (start with RS)
- Object lock / compliance modes (until a tenant needs it)
Appendix G: Multi-Tenancy, Fairness, and Cost#
| Concern | Mechanism |
|---|---|
| Request fairness | Per-tenant and per-prefix token buckets at the front end; see Rate Limiting |
| Metadata isolation | Partitions are per-bucket ranges; hot tenants split without touching neighbors |
| Data-plane isolation | Cell architecture: tenants assigned to cells; a cell's failure is bounded |
| Bandwidth fairness | Per-tenant egress shaping on storage nodes during contention |
| Billing | Separate dimensions: GB-month by class, requests by type (PUT/LIST priced above GET), egress, retrieval |
| Background budget | Repair, scrub, GC, transcode, compaction each get a guaranteed bandwidth share |
🧭 Principal Insight: "Price what costs us money. If LIST and small objects are what load the metadata plane, bill for them — otherwise we subsidize the workloads that hurt us most."
Related reading: Large Blobs, File Sync, CDN & Edge Caching, Database Sharding, Consistency Models, and Sharding.