Hiring BarSupport

Design Blob Storage (S3) — Staff-Level Case Study

Case study71 min read7 diagrams

Technologies referenced in this case study: DynamoDB · Cassandra · PostgreSQL · ZooKeeper & etcd · Kafka

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once, then return to sections for targeted review.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 3, 4
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Appendix A (durability math)
Deep Dive3+ hrsEverything, including The Principal Lens and appendices
What is Blob Storage? — Why interviewers pick this topic

Blob (object) storage stores opaque byte sequences — 1 KB thumbnails to multi-terabyte backups — under a key in a flat namespace, with a simple API: PUT, GET, DELETE, LIST. It's the durable substrate under data lakes, backups, media, ML datasets, logs, and half of every other system design answer ("…and we put the files in S3").

Before vs After — the "three replicas is enough" launch:

With naive 3x replication, random placement, 24h repair:
t=0:       10 PB logical, 30 PB raw on 5,000 disks; storage bill 3x logical
t=+3mo:    Finance flags storage as the #2 infra line item
t=+6mo:    A rack PDU fails: 40 disks offline at once
t=+6mo:    Random placement means ~thousands of objects had 2 replicas in that rack
t=+6mo+2h: Unrelated disk failure elsewhere; 312 objects lose their last copy
t=+6mo+1d: Customer-visible data loss; "11 nines" marketing page becomes a liability

With erasure coding + failure-domain-aware placement:
t=0:       Same 10 PB on zone-aware RS(9,6), 5 fragments per AZ: ~16.7 PB raw (1.67x)
t=+6mo:    Same PDU failure: at most 1 fragment per stripe in that rack
t=+6mo:    Repair prioritizes stripes with the fewest surviving fragments
t=+6mo+2h: Second failure elsewhere: every stripe still has >= 13 of 15 fragments
t=+1yr:    Zero loss; storage bill ~53% lower than 3x

Why interviewers reach for this question: It looks like "a key-value store for big values," so it separates candidates who draw an API gateway in front of a database from candidates who understand that blob storage is two systems with opposite properties — a small, strongly consistent, latency-sensitive metadata plane and a huge, throughput-oriented, failure-tolerant data plane — and that durability is a probability you compute, not a feature you claim.

Mechanics Refresher: Redundancy Schemes
SchemeHow It WorksStorage OverheadToleratesRepair CostRead Latency
3× replicationThree full copies on distinct failure domains3.0×2 lossesCopy 1 objectBest — any replica
RS(6,3)6 data + 3 parity fragments1.5×3 lossesRead 6 fragments to rebuild 1Needs 6 of 9 (or 1 if unsplit)
RS(10,4)10 data + 4 parity1.4×4 lossesRead 10 to rebuild 1Needs 10 of 14
RS(9,6), 5 per AZ9 data + 6 parity over 3 AZs~1.67×A whole AZ + 1 more fragmentRead 9 to rebuild 1Needs 9 of 15
LRC (12,2,2)12 data, 2 local parities (each over 6), 2 global parities~1.33×Any 3, many 4-loss patternsRead 6 for a single lossSimilar to RS
Geo-EC (cross-region XOR)XOR of blocks from two regions stored in a third~1.5× across 3 regions vs 2–3× geo-replicationRegion lossCross-region readsDegraded reads are cross-region

For most production systems: Replicate small and hot objects (and the first minutes of every write), erasure-code everything else with a zone-aware code — parity ≥ the fragments in any one AZ, e.g. RS(9,6) at ~1.67× — and use RS(10,4) at 1.4× for single-AZ or reduced-redundancy classes. The code choice is not the interview — placement, repair speed, and what the metadata commit guarantees are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What This Interview Actually Tests#

Blob storage is not a key-value store with big values. Everyone can map a key to a file path.

It is a durability-economics and metadata-scaling question that tests:

  • Whether you separate the metadata plane from the data plane and give each the right consistency model
  • Whether you can compute durability from failure rates, repair time, and placement — and know what the math leaves out
  • Whether you trade storage overhead (money) against repair bandwidth and tail latency (operations) explicitly
  • Whether you own the silent failures: bit rot, orphaned bytes, lost metadata, runaway lifecycle bills

The key insight: The data plane is embarrassingly parallel and almost never the interview bottleneck. The metadata plane — ordered keys, strong read-after-write, LIST on hot prefixes, versioning, and the commit protocol that ties bytes to names — is where designs break and where real systems spend their hardest engineering.

The L5 vs L6 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveAPI gateway → metadata DB → storage nodesAsks object size distribution, read/write ratio, durability target, and whether it's a general platform or a single-purpose storeAsks what the org already stores where, the $/GB-month target, and which teams will build on this as a one-way dependency
Durability"3 replicas across AZs, 11 nines"Computes it: AFR, repair window, placement; knows correlated failures and software bugs dominate the mathDesigns the org's durability posture: independent failure domains, deletion protection, versioned control-plane changes, and audits
RedundancyReplicationEC for cold/large, replication for small/hot; names repair bandwidth and tail-latency costsPrices the switch: 3× → 1.67× at 100 PB avoids 133 PB of raw disk ($20M+/yr), paid for with more repair traffic and an EC team
Consistency"Eventually consistent is fine for blobs"Strong read-after-write via metadata commit as the linearization point; data written before metadataTreats consistency as a customer contract that can't be weakened later without breaking unknown consumers
Metadata"Store it in Postgres/Cassandra"Range-partitioned, strongly consistent index with auto-splitting for hot prefixes; LIST cost is a first-class concernDecides the namespace model (flat vs hierarchical) as a multi-year one-way door
OwnershipStorage team owns itStorage owns durability; tenants own lifecycle policy and access; security owns encryption keysChargeback and tiering incentives so teams stop storing 40% garbage
Why "durability" separates levels

L5: Quotes "eleven nines" and three replicas. Not wrong — three replicas with fast repair really can compute to ~11 nines against independent disk failures.

L6: Shows the calculation, then immediately says what it omits: "With a 1.5% annual disk failure rate and 6-hour repair, independent failures give ~10⁻¹¹ annual loss probability per object for 3× replication. But at 10¹² objects that's still ~10 objects a year, and the real risks are correlated — a rack power event, a firmware bug on one drive model, or a bad deploy that deletes metadata. Durability engineering is mostly about those."

Why "consistency" separates levels

L5: "Blobs are immutable, so eventual consistency is fine." Plausible — the bytes are immutable.

L6: The name → bytes mapping is not immutable. Overwrites, deletes, and LIST after PUT are where applications break under eventual consistency: a data pipeline lists a prefix, misses a just-written file, and silently drops a partition. Making the metadata commit the single linearization point gives strong read-after-write for GET, HEAD, and LIST — which is exactly what S3 moved to in December 2020.

Why "metadata" separates levels

L5: Puts object metadata in a sharded database keyed by hash(bucket, key). Point lookups scale perfectly.

L6: Hash partitioning makes LIST prefix=logs/2026/09/ a scatter-gather across every shard. Ordered listing requires range partitioning by (bucket, key) — which reintroduces hot spots when keys are sequential (timestamps). The answer is range partitioning with automatic split-on-load, plus guidance to customers about key design. Public S3 guidance historically quoted per-prefix rates of ~3,500 writes/s and ~5,500 reads/s, scaling by adding prefixes — the visible artifact of this exact design choice.

The Staff Positions#

PositionRationale
Separate metadata plane from data planeDifferent scale (KB vs PB), consistency (linearizable vs append-only), and failure handling
Write data first, commit metadata lastThe metadata commit is the atomic visibility point; crashes leave only garbage bytes, never dangling names
Erasure-code large and cold objects; replicate small and hotZone-aware EC is ~1.67× (RS(10,4) 1.4× in one AZ) vs 3×; small objects don't amortize fragment overhead
Placement across ≥ 3 AZs with at most one fragment per failure domain per stripeCorrelated failures dominate the loss math
Repair speed is a durability leverLoss probability scales with (repair time)ᵐ; halving repair time buys more than an extra parity fragment
Continuous scrubbing with end-to-end checksumsSilent corruption is found by reading, not by waiting for a customer GET
Deletes are soft firstVersioning, delayed GC, and object lock protect against the most common loss cause: humans and bugs

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
General-purpose multi-tenant object store (S3-like)Durability, strong consistency, unbounded scale, per-tenant isolationSeparate planes; EC across AZs; range-partitioned metadata; lifecycle tiersMetadata hot spots; noisy tenants; loss from correlated failure or bug11 nines designed durability, 99.9–99.99% availability, read-after-write
Media/CDN origin (photos, video)Read throughput and $/GB served; mostly immutable, write-onceLarge-object EC, CDN in front, needle-style packing for small filesOrigin overload on cache miss storms; small-file metadata overheadHigh durability; availability of hot objects
Backup / archive$/GB-month above all; reads rare and slow is OKWide EC, dense disks or tape, cold tiers with hours-long retrievalSurprise retrieval cost/latency during a restore; silent rot never readDurability over decades; restore time objective

🎯 Staff Move: "I'll design the general-purpose multi-tenant store, because it contains the other two as tiers. The hard parts are the metadata plane — strong consistency and ordered LIST at trillions of keys — and the durability math, which I want to compute rather than assert. Media origin and archive become storage classes with different placement and codes on the same metadata plane."

The Five Fault Lines#

#Fault LineThe Tension
1Replication vs Erasure Coding3× storage and simple repair, or ~1.4–1.67× storage with expensive repair and tail-latency risk?
2Strong vs Eventual Metadata ConsistencyLinearizable names (coordination cost) or eventual (application bugs you don't see)?
3Durability vs Write LatencyAck after cross-AZ durable writes (slower PUT) or after local write (loss window)?
4Flat vs Hierarchical NamespaceFlat keys (scales, cheap) or real directories (atomic rename, expensive to scale)?
5Automatic vs Customer-Driven TieringPlatform moves data to cheaper tiers (surprises on access) or customers opt in (40% of data never tiered)?

In the Wild: Real Production Systems#

Why this section belongs here: Citing specific, publicly documented object stores shows you've studied real durability engineering, not just the API.

Amazon S3 — Strong Consistency and ShardStore#

S3 is designed for 99.999999999% (11 nines) durability, stores data across a minimum of three AZs for standard classes, and moved from eventual to strong read-after-write consistency for all PUT, GET, LIST, and HEAD operations in December 2020 at no extra cost. AWS has publicly described the approach (a replicated cache-coherence "witness" that tracks metadata freshness) and published a SOSP 2021 paper on ShardStore, the Rust-based storage-node engine validated with lightweight formal methods.

Staff insight: S3 upgraded consistency after a decade of eventual consistency — because customers built workarounds (S3Guard, EMRFS consistent view) that cost more than fixing it at the source. Weak guarantees don't disappear; they move into every consumer.

Facebook — Haystack and f4#

Haystack (OSDI 2010) packed many photos into large append-only volume files so each read needed ~1 disk seek instead of several filesystem metadata seeks — metadata overhead was the bottleneck, not bytes. f4 (OSDI 2014) moved "warm" blobs to Reed-Solomon (10,4) within a region plus XOR across regions, cutting effective replication from 3.6× to ~2.1× for that tier.

Staff insight: Hot/warm separation is a durability-and-cost decision driven by access-age curves. Most bytes stop being read within weeks — the tiering boundary follows the data.

Microsoft Azure Storage — Local Reconstruction Codes#

Azure's LRC (USENIX ATC 2012) adds local parities over subsets of data fragments so the common case — one fragment lost — repairs by reading 6 fragments instead of 12, while keeping ~1.33× overhead and durability comparable to 3× replication.

Staff insight: The code was designed around the repair path, not the storage overhead. Repair bandwidth is the hidden cost of erasure coding, and it's paid continuously.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Three replicas for 11 nines""Show me the math. What's the repair time assumption?"Can you compute durability?
"Erasure coding saves storage""What does a single-disk repair cost in network? What about tail latency on reads?"Do you know EC's hidden costs?
"Metadata in a sharded DB""How does LIST prefix= work? Sharded by what?"Range vs hash partitioning
"Strong consistency""A PUT succeeds and the metadata write times out. What does GET return?"Commit protocol and linearization point
"Multipart upload""10,000 parts uploaded, client disappears. Who pays for them?"Orphan lifecycle and ownership
"Delete the object""A bug deletes 5% of a bucket. Recover it."Soft delete, versioning, GC delay

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: A PUT writes data fragments first (step 2) and commits the metadata version last (step 3); the commit is the single point at which the object becomes visible. The metadata plane is small — ~hundreds of bytes per object — but strongly consistent and range-partitioned for ordered LIST. The data plane is enormous, append-only, and tolerant of losing any disk at any time. The background plane — repair, scrub, GC, lifecycle — is where durability is actually earned, day after day.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Architecture"API servers, metadata DB, storage nodes""Two planes: strongly consistent range-partitioned metadata, append-only erasure-coded data. Data first, metadata commit last."
Durability"Three replicas, 11 nines""Zone-aware EC like RS(9,6), 5 fragments per AZ, one per rack, ~6h repair — survives an AZ loss plus a disk. Independent failures compute to far beyond 11 nines; correlated failures and bugs are the real budget."
Consistency"Eventual is fine""Strong read-after-write for GET/HEAD/LIST; the metadata commit is the linearization point."
Large uploads"Stream it""Multipart: parts 5 MiB–5 GiB, up to 10,000, parallel and retryable; complete is an atomic metadata commit; abandoned uploads expire via lifecycle."
Listing"Query the DB""Range-partitioned index, auto-split on hot prefixes; LIST is paginated and ordered."
Cost"Disks are cheap""Overhead × $/GB × PB. At 100 PB, 3× → 1.67× is ~133 PB of raw capacity avoided."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Designed durability (S3 standard)11 ninesLoss of ~1 object per 10¹¹ per year — at 10¹² objects that's ~10/yr
Annualized disk failure rate~1–2%Public fleet data (e.g., Backblaze reports) sits in this range
3× replication overhead3.0×Simple, fast repair, fast reads
RS(10,4) overhead1.4×Tolerates 4 losses; single repair reads 10 fragments
LRC (12,2,2) overhead~1.33×Single-fragment repair reads 6, not 12
Metadata per object~0.3–1 KB10¹² objects ≈ ~0.3–1 PB of metadata — itself a distributed database
Multipart limits (S3)5 MiB–5 GiB parts, 10,000 partsMax object 5 TiB with the classic limits
Per-prefix request rate (S3 guidance)~3,500 writes/s, ~5,500 reads/sHot prefixes split; scale by spreading keys
HDD capacity per drive~20–30 TBRebuilding one drive at 100 MB/s takes ~2–3 days — why repair must be parallel
Parallel repair across 1,000 disks~minutes–hours per failed driveDeclustered placement turns a days-long rebuild into hours
Scrub cycle~1–4 weeks per full passBounds how long latent corruption can hide
Deep archive retrieval~12–48 hoursRestore-time objective lives or dies here

Interview Walkthrough

The most common mistake: Candidates spend 15 minutes on REST endpoints and a capacity table, draw "metadata DB + storage nodes," say "3 replicas," and never reach the commit protocol, the durability math, or LIST. Compress the API to 2 minutes. The level is decided in the planes split and the durability argument.


Phase 1: Requirements & Framing (2–3 minutes)#

Functional scope in one sentence:

"A multi-tenant object store: PUT, GET (with byte ranges), DELETE, and ordered LIST over a flat key namespace inside buckets, with versioning and lifecycle policies."

Then the non-functional requirements that shape the design:

QuestionWhy It MattersDefault I'd Assume
Object size distribution?Small objects dominate metadata; large dominate bytesMedian ~100 KB, p99 ~1 GB, max 5 TB
Total scale?Metadata and placement design100 PB, 10¹¹–10¹² objects, growing ~40%/yr
Durability target?Redundancy scheme and placement11 nines designed, against independent failures
Availability target?Replication of metadata, AZ layout99.99% for reads, 99.9% for writes
Consistency?Commit protocolStrong read-after-write, including LIST
Request rate?Front end and metadata partitions~10M req/s aggregate, ~70% GET

"The two numbers I'll design around are 10¹² objects — that makes metadata a distributed database problem — and 11 nines, which I want to derive rather than assume."


Phase 2: Core Entities & API (1–2 minutes)#

Bucket        { bucket_id, tenant_id, region, versioning, lifecycle_rules, policy, storage_class_default }
ObjectVersion { bucket_id, key, version_id, size, etag, storage_class, created_at,
                is_delete_marker, manifest_ref, checksum, encryption_key_ref }
Manifest      { version_id, chunks: [ {chunk_id, offset, length, stripe_id} ] }
Stripe        { stripe_id, code: RS(9,6), fragments: [ {node_id, extent_id, offset, crc} x 15 ] }
Upload        { upload_id, bucket_id, key, parts: {n -> {etag, size, chunk_refs}}, initiated_at }

Public API (HTTP):

PUT    /{bucket}/{key}                       -> 200, ETag, x-version-id
GET    /{bucket}/{key}  Range: bytes=a-b      -> 206
DELETE /{bucket}/{key}                       -> 204 (delete marker if versioned)
GET    /{bucket}?list-type=2&prefix=&continuation-token=&max-keys=1000
POST   /{bucket}/{key}?uploads               -> upload_id
PUT    /{bucket}/{key}?partNumber=n&uploadId= -> part ETag
POST   /{bucket}/{key}?uploadId=              -> complete (atomic)

"The key design decision: the ObjectVersion row is the only thing that makes bytes visible. Stripes and chunks can exist without it — that's garbage, collected later. The reverse must never happen."


Phase 3: High-Level Architecture (≤5 minutes)#

Staff candidates spend under 5 minutes here. Draw the two planes, the write order, and move on.

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Say it in four sentences:

  1. "Stateless API servers authenticate, rate-limit per tenant, and stream the body into chunks — say 8 MB each — erasure-coded into stripes."
  2. "Placement puts at most one fragment of a stripe in any rack and spreads fragments across three AZs. One catch: RS(10,4) over three AZs puts 5 fragments in some AZ, which exceeds 4 parity — so for AZ-loss survival I'd pick a code whose parity covers the largest AZ share, like RS(9,6) at 1.67×, or a zone-aware LRC."
  3. "After fragments are durable, the API server does a conditional write to the metadata partition that owns (bucket, key). That commit is the linearization point — GET, HEAD, and LIST see it immediately."
  4. "Repair, scrubbing, GC, and lifecycle run in a background plane with their own budgets."

🎯 Staff Move: Sentence 2 is where candidates get caught. RS(10,4) over 3 AZs puts ≥ 5 fragments in some AZ, so losing that AZ loses the stripe. The rule: parity count ≥ the largest per-AZ fragment count. RS(9,6) at 5-5-5 costs ~1.67× instead of 1.4× — that 0.27× is the price of surviving an AZ. Say that tradeoff out loud; interviewers write it down.


Phase 4: Transition to Depth (1 minute)#

"The API and the byte path are straightforward. Where I'd like to spend time: first, durability — the math, the code choice, and the correlated failures the math leaves out; second, the metadata plane — strong consistency, ordered LIST, and hot prefixes at 10¹² keys; third, the lifecycle of bytes — multipart uploads, deletes, GC, and tiering, which is where money and data actually get lost. Which would you like first?"

Default order if they don't choose: durability → metadata → lifecycle.


Phase 5: Deep Dives (25–30 minutes)#

Deep dive 1 — Durability math (8–10 min).

  • Per-disk annual failure rate λ ≈ 1.5%/yr; repair window T. Loss needs m+1 overlapping failures in a stripe of n fragments.
  • Approximate annual loss probability per stripe: P ≈ n·λ · C(n−1, m) · (λT)ᵐ.
  • 3× replication (n = 3, m = 2), T = 6 h: 3 × 0.015 × 1 × (0.015 × 6.85×10⁻⁴)² ≈ 4.8×10⁻¹² — ~11 nines.
  • Same with T = 24 h: ≈ 7.6×10⁻¹¹ — ~10 nines. Repair time enters squared.
  • RS(10,4), T = 6 h (for scale; RS(9,6) is lower still): 14 × 0.015 × C(13,4) × (1.03×10⁻⁵)⁴ ≈ 1.7×10⁻¹⁸. The independent-failure math stops being the constraint.
  • "So once I'm on a wide code with fast repair, the durability budget is spent on correlated failures — AZ events, firmware bugs, bad deploys, and operator deletes. That's what I design for next."

Deep dive 2 — Metadata plane (8–10 min).

  • Range-partitioned by (bucket_id, key), each partition a Raft/Paxos group of 3–5 replicas across AZs, ~1–10 GB per partition, auto-split at size or load thresholds.
  • PUT = conditional insert of a new version row; GET reads the latest non-delete-marker version from the partition leader (or a follower with a read lease).
  • LIST = ordered range scan with continuation token = last key returned; pages of 1,000.
  • Hot prefix: sequential keys (logs/2026-09-30T12:00:00/...) concentrate writes on the last partition; split-on-load moves the boundary, but monotonically increasing keys always hit the tail. Mitigation: split heat detection + guidance to prefix with a hash; the platform can't fully fix a key design.

Deep dive 3 — Lifecycle of bytes (6–8 min).

  • Multipart: parts are independent chunks with their own stripes; complete commits one version row whose manifest references part chunks — no data copy.
  • Abandoned uploads: bytes exist without a version row. Lifecycle rule "abort incomplete uploads after 7 days" plus a platform-wide backstop.
  • Delete: writes a delete marker (versioned) or tombstone; GC reclaims chunks after a delay (e.g., 24–72 h) so a mistaken delete or metadata bug is recoverable.
  • Tiering: lifecycle moves versions to cheaper classes by age; the metadata row changes its storage_class and manifest atomically after bytes are rewritten.

Phase 6: Wrap-Up (2–3 minutes)#

"To summarize: two planes; data written first and metadata committed last so a crash leaves only garbage, never a broken name; erasure coding across failure domains with repair time as the main durability lever; range-partitioned metadata for strong consistency and ordered LIST; soft deletes and delayed GC because humans and bugs are the biggest loss risk. Next I'd build cross-region replication as a per-bucket policy, object lock for compliance, and cost visibility per tenant so lifecycle rules actually get written."


Common Timing Mistakes#

MistakeTime LostWhat to Do Instead
Designing the REST API in detail5–8 minName 6 operations; move on
Capacity math for storage nodes5 min"100 PB at ~1.67× ≈ 167 PB raw ≈ ~7K 24 TB drives" — one line
Explaining Reed-Solomon internals5 min"k data, m parity, any k reconstruct" is enough
Skipping LISTlevel-decidingLIST drives the metadata partitioning choice
Asserting 11 nines without mathlevel-decidingCompute it; then name what the math omits

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Almost every system design ends with "and we store the files in S3." This prompt turns that black box inside out. It tests whether a candidate understands the difference between a system that is correct (every PUT readable) and one that is durable over a decade (every PUT readable after 10 years of disk failures, deploys, migrations, and operator mistakes). The second property is not a feature of any component; it emerges from placement, repair, scrubbing, and process.

It's also an economics question. Storage is a line item that grows monotonically. The candidate who says "3×" at 100 PB has just committed ~133 PB of unnecessary disks versus ~1.67× zone-aware EC — tens of millions of dollars a year — and the Staff answer is expected to notice.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"What is the most likely way we lose a customer's object — and what stops it?"

Answering honestly moves the conversation away from disks. At scale, with a wide code and fast repair, the independent-disk path is ~10⁻¹⁸. The likely paths are: a deploy that corrupts metadata; a GC bug that deletes live chunks; a placement bug that puts two fragments in one rack; a firmware issue that fails 3% of one drive model in a week; a customer's own script deleting a prefix. Each needs a named defense — canarying metadata changes, GC with a delay and a reference-count audit, placement verification in the scrubber, drive-model diversity, versioning and object lock.

🎯 Staff Move: "Once the math says the disks aren't the risk, I spend the durability budget on software and humans: delayed GC, versioning, placement audits, and staged deploys of anything that can touch metadata."


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Intent 1 — General-purpose multi-tenant store. Unknown workloads from thousands of tenants: data lakes doing LIST over millions of keys, ML jobs reading at 100 GB/s, backup tools writing TB-sized objects, web apps serving 10 KB avatars. The design must be strongly consistent, isolate noisy tenants, and offer tiers because one price point can't serve all of them. Correctness bar: read-after-write, designed durability, per-tenant fairness.

Intent 2 — Media/CDN origin. Write-once, read-many, heavy-tailed popularity. Most reads are served by a CDN; the origin must survive miss storms (a viral video, a CDN purge) and keep $/GB low for a long tail nobody reads. Small-file packing (Haystack-style) cuts metadata cost. Correctness bar: durability plus availability of popular objects.

Intent 3 — Backup/archive. Writes are large and sequential; reads are rare and often all-at-once during disaster recovery. Wide codes, dense disks or tape, and retrieval measured in hours are fine — until a restore happens under pressure. Correctness bar: decades-long durability and a restore-time objective someone actually tested.

DimensionGeneral-PurposeMedia OriginBackup/Archive
Dominant costMetadata ops + storageEgress + cache misses$/GB-month
Consistency needStrong incl. LISTWeak OK (immutable keys)Weak OK
RedundancyEC across AZs + replicated small objectsEC, CDN absorbs readsWide EC / tape
Worst silent failureLIST misses a new fileOrigin melts on miss stormRot never read until restore
Who complainsData engineers, app teamsEnd users (slow media)Execs, during a disaster

2.2 When NOT to Build Blob Storage#

SituationUse InsteadWhy
You're not a cloud providerS3 / GCS / Azure Blob11 nines needs thousands of disks, 3 AZs, and a repair team; you won't beat their $/GB
Data < 1 PB on-prem, regulatedMinIO / Ceph on owned hardwareMature open source; build only the policy layer
Small mutable recordsA databaseBlob stores have no partial update, no transactions across keys
POSIX semantics (rename, append, locking)A distributed filesystemEmulating directories on flat keys breaks atomicity
Sub-millisecond readsA cache or KV store in frontFirst-byte latency is ~10–100 ms by design

🎯 Staff Move: "If the question were 'should we build this' at a normal company, the answer is no — buy it and spend the effort on lifecycle policy and cost governance. I'll design it as if we're the provider."

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy It MattersWhat to Say
Object size distributionSmall objects break EC economics"Replicate or pack objects < ~1 MB; EC above"
Durability definitionPer-object vs fleet-wide loss"11 nines per object per year, against independent failures; correlated risks designed separately"
Consistency of LISTDrives metadata partitioning"Strong, including LIST"
Multi-regionReplication cost and semantics"Single region, 3 AZs; cross-region is async, per-bucket opt-in"
Delete semanticsRecoverability vs compliance"Versioning on by default for new buckets; object lock available"
EncryptionKey management ownership"Encrypt at rest always; per-tenant keys optional via KMS"

2.4 Precise Terminology#

TermPrecise Meaning
DurabilityProbability an acknowledged object remains readable over a period (usually per year)
AvailabilityFraction of requests served successfully; a durable object can be unavailable
AFRAnnualized failure rate of a component (disk, node)
StripeA set of k data + m parity fragments encoding one chunk
Failure domainA set of components that can fail together: disk, host, rack, power zone, AZ
Repair window (MTTR)Time from fragment loss detection to redundancy restored
Degraded readReading a stripe when some fragments are missing, requiring reconstruction
Linearization pointThe instant an operation takes effect for all observers — here, the metadata commit
Delete markerA version row that hides earlier versions without removing bytes
OrphanBytes on storage nodes with no metadata reference (failed PUT, abandoned multipart)
ScrubbingPeriodically reading and verifying stored data against checksums

3. The Five Fault Lines#

3.1 Fault Line 1: Replication vs Erasure Coding#

The tension: Replication buys simplicity, fast reads, and cheap repair with storage money. Erasure coding buys storage money with repair bandwidth, CPU, and tail latency.

StrategyWhat WorksWhat BreaksWho Pays
3× replication everywhereSimplest code path; any replica serves reads; repair copies one object~1.8× more raw disk than zone-aware ECFinance: at 100 PB, ~133 PB of extra disk
Zone-aware EC everywhere (e.g., RS(9,6))~1.67×, survives AZ lossSmall objects waste space (min fragment sizes); every repair reads 9 fragments; degraded reads need reconstructionTenants with small objects (latency); network team (repair traffic)
Hybrid: replicate on write, EC laterFast PUT ack; EC once object is cold (hours–days)Two code paths; transcoding job is a new failure modeStorage team complexity
Hybrid by size: replicate < 1 MB, EC ≥ 1 MBSmall objects stay simple; bytes dominated by large objects get EC savingsSmall objects are most of the object count at 3×Small-object tenants pay more per GB (price it that way)
LRC (local parities)Single-fragment repair reads ~half as many fragmentsMore complex code, slightly higher overhead than RS of same widthStorage team engineering time

Staff default: Size-based hybrid. Objects < ~1 MB are packed into larger replicated or EC'd containers (Haystack-style) so tiny objects don't each become a stripe; objects ≥ 1 MB are chunked (e.g., 8–64 MB) and erasure-coded with a zone-aware code. Repair traffic is budgeted as a fraction of cluster bandwidth (e.g., ≤ 10–20%) with priority by fragments remaining.

The hidden cost to say out loud: Losing a 24 TB disk under RS(9,6) means reconstructing 24 TB by reading ~9 × 24 = 216 TB across the cluster. Declustered placement spreads that over thousands of disks, but it's continuous background load — at ~1.5% AFR and 10,000 disks, that's ~150 disk failures a year, ~3 a week.

When to deviate: A latency-critical small-object store (avatars behind no CDN) can stay replicated. Archive classes go wider (e.g., RS(17,3) inside an AZ plus cross-region copies) because repair speed matters less than $/GB.

3.2 Fault Line 2: Strong vs Eventual Metadata Consistency#

The tension: Strong consistency needs a single authority per key (consensus, leader leases); eventual consistency lets any replica answer and pushes anomalies onto every consumer.

StrategyWhat WorksWhat BreaksWho Pays
Eventual everywhereHighest availability; cheap multi-replica readsLIST misses new keys; overwrite returns old bytes; delete reappearsEvery data engineer writing a workaround (S3Guard-style)
Read-after-write for new keys onlyCovers the common PUT-then-GETOverwrites and LIST still anomalousPipelines that overwrite or list
Strong read-after-write incl. LISTApplications behave as with a local filesystem (minus rename)Every read touches the partition leader or a lease-holding follower; caches need coherenceStorage team: consensus, lease, and cache-invalidation engineering
Linearizable multi-key transactionsAtomic rename, conditional multi-object opsCross-partition coordination (2PC) on the hot pathLatency for everyone; complexity

Staff default: Strong read-after-write for single-key operations and LIST, via per-partition consensus with the metadata commit as the linearization point. No multi-key transactions; offer conditional writes (If-Match, If-None-Match) for single-key compare-and-swap. Caches in the front end are validated against a per-partition version (a witness-style freshness check) rather than TTLs.

🎯 Staff Move: "Weak consistency doesn't remove the cost — it moves it into every consumer. With thousands of tenants, I'd rather pay it once in the metadata plane."

3.3 Fault Line 3: Durability vs Write Latency#

The tension: When do we say 200 OK? Every additional durable copy before the ack adds latency; every copy after it is a loss window.

StrategyWhat WorksWhat BreaksWho Pays
Ack after local writeLowest latencyNode loss before replication loses acked dataCustomers — silently
Ack after all fragments durableMaximum durability at ackOne slow node sets p99 for every PUTAll writers' tail latency
Ack after write quorum (e.g., 13 of 15), across ≥ 2 AZsDurable against AZ loss at ack; tail cut by skipping stragglersRepair must fill missing fragments quickly; quorum logic in placementStorage team (straggler repair)
Ack after fsync on replicated log, EC laterLow latency, durableTwo-phase data lifecycle; log capacity planningStorage team complexity

Staff default: Ack after a write quorum that spans at least 2 AZs and still survives an AZ loss plus one disk (e.g., ≥ 13 of 15 fragments in RS(9,6) with ≥ 4 per AZ). Missing fragments are enqueued for immediate repair. Never ack before data is fsynced somewhere durable — "written to page cache" is not durable.

When to deviate: A reduced-redundancy or single-AZ class can ack faster; price and name it so customers choose knowingly.

3.4 Fault Line 4: Flat vs Hierarchical Namespace#

The tension: Flat keys scale linearly and partition cleanly; directories give atomic rename and cheap listing of children, but require a tree with cross-partition operations.

StrategyWhat WorksWhat BreaksWho Pays
Flat keys, prefix LISTRange partitioning; unlimited scale; simple"Rename directory" = copy + delete of N objects; not atomicAnalytics frameworks that commit via rename (job output committers)
Hierarchical namespace (real directories)Atomic rename, fast directory listingDirectory inode hot spots; cross-partition moves need transactionsStorage team; scale limits per directory
Flat + optional hierarchical bucket typeServes both; customers chooseTwo metadata engines to runStorage team headcount

Staff default: Flat namespace for the general-purpose store. Document that rename is not atomic and push analytics engines toward commit protocols that write a manifest (table formats like Iceberg/Delta commit by writing a metadata file, not renaming). Offer hierarchical as a separate bucket type only if analytics demand justifies a second engine.

Hot prefixes are the practical fault line inside flat namespaces: sequential keys hit one range partition. Split-on-load helps for spread-out heat; monotonically increasing keys always write the tail partition. Publish key-design guidance and expose per-prefix throttling (503 SlowDown) instead of letting one tenant's hot tail slow its neighbors.

3.5 Fault Line 5: Automatic vs Customer-Driven Tiering#

The tension: Cheaper tiers save 50–95% on $/GB-month but add retrieval cost, retrieval latency, and minimum-duration charges. Who decides when data moves?

StrategyWhat WorksWhat BreaksWho Pays
Customer lifecycle rules onlyPredictable; customer controls tradeoffsMost teams never write rules; ~30–60% of bytes sit cold in the hot tierThe company's storage bill
Platform auto-tiering by access ageSaves money without customer effortA restore or backfill hits cold data: surprise latency or retrieval feesCustomer during an incident
Auto-tiering among instant-access tiers onlySavings without retrieval latencySmaller savings than archive tiersPlatform (monitoring cost per object)
Default rules on new bucketsGood defaults, opt-out possibleDefaults can be wrong for a workloadTeams that don't read defaults

Staff default: Automatic tiering only between tiers with the same access latency (monitor per-object access, move after 30–90 days idle); archive tiers with hours-long retrieval are strictly opt-in via lifecycle rules, because the victim of a surprise 12-hour restore is a customer mid-incident. Every bucket shows cost by tier and "bytes not read in 90 days" in its dashboard.

Diagram: 3.5 Fault Line 5: Automatic vs Customer-Driven Tiering

4. Failure Modes & Operational Reality#

4.1 Correlated Failure — The Bad Drive Batch#

t=0:       Drive model X (18% of fleet) hits a firmware bug at ~30K power-on hours
t=+2d:     Disk failure rate for model X: 1.5% AFR -> ~40% annualized
t=+2d:     storage.disk_failures_per_hour 3x baseline; repair queue growing
t=+3d:     storage.stripes_at_min_redundancy rises from 0 to 1,200
t=+3d:     Page: stripes with <= 1 spare fragment > 100
t=+3d+1h:  Repair re-prioritized: fewest-surviving-fragments first; bandwidth cap 10% -> 35%
t=+4d:     Placement excludes model X for new writes; proactive drain begins
t=+3wk:    Model X drained; zero loss
  • Detection: storage.disk_afr{model}, storage.stripes_by_surviving_fragments, repair.queue_age_p99.
  • Blast radius: Any stripe with multiple fragments on model X. Placement diversity by drive model (not just rack/AZ) caps it.
  • Mitigation: Priority repair, raise repair bandwidth, stop placing on the bad model, proactive drain.
  • Prevention: Treat drive model and firmware version as failure domains in placement; stage firmware updates by cohort.
  • Owner: Storage SRE; hardware team for vendor escalation.

4.2 The Metadata Hot Partition#

t=0:       Tenant launches IoT ingest writing keys iot/{timestamp}/{device}
t=+10min:  All writes land on the last range partition of the bucket
t=+12min:  metadata.partition_qps{p=bucket-7:tail} = 18K vs split threshold 5K
t=+13min:  Auto-split: new boundary, but monotonic keys keep hitting the new tail
t=+15min:  Partition leader CPU 95%; PUT p99 for that tenant 40ms -> 2s
t=+15min:  Other tenants unaffected (partitions are per-bucket ranges)
t=+20min:  Tenant receives 503 SlowDown; SDK backs off
  • Detection: metadata.partition_qps, metadata.split_rate, api.throttled_total{tenant}.
  • Blast radius: One tenant's hot prefix — if partitions are per bucket and leaders are spread. If many hot tails share a leader host, collateral damage.
  • Mitigation: Throttle with 503 + Retry-After; tenant adds a hash prefix; temporary pre-split of the key range.
  • Prevention: Key-design guidance, per-prefix throttling, leader placement spreading.
  • Owner: Metadata team for platform; tenant for key design.

4.3 The GC Bug — Deleting Live Data#

The most dangerous failure in any blob store: a garbage collector that decides live chunks are orphans.

t=0:       Deploy changes manifest serialization for multipart objects
t=+1h:     GC's reference scan can't parse new manifests; treats their chunks as unreferenced
t=+1h:     GC marks 2.1M chunks for deletion; grace period 72h starts
t=+6h:     Audit job: gc.marked_bytes_ratio 0.03% -> 1.4% of cluster (anomaly)
t=+6h:     Page; GC deletion halted by kill switch
t=+8h:     Root cause found; marks cleared; zero bytes actually deleted
  • Detection: gc.marked_bytes_ratio vs baseline, independent reference-count audit (a second implementation that samples chunks and verifies references).
  • Blast radius without delay: Permanent loss of every multipart object written after the deploy.
  • Mitigation: GC kill switch; grace period before physical delete; rate limit on deletions per hour.
  • Prevention: GC deletes only after two independent scans agree; manifest format changes are backward-compatible and canaried; physical deletion capped at a fraction of cluster per day.
  • Owner: Storage team; GC changes require a second reviewer from the durability owner.

🎯 Staff Move: "The garbage collector is the most dangerous component we own, because it's the only one designed to delete data. I want a delay, a rate limit, a kill switch, and an independent audit on it."

4.4 Silent Corruption — Bit Rot Found Late#

  • Symptom: A customer GET fails checksum on an object last read 14 months ago; reconstruction succeeds from parity.
  • Root cause: Latent sector errors; one fragment corrupted, never read.
  • Detection: scrub.checksum_failures, scrub.cycle_age_max (time since each disk was fully read), get.reconstructed_reads_ratio.
  • Mitigation: Scrubber rewrites the fragment on detection.
  • Prevention: Scrub cycle ≤ 2–4 weeks; end-to-end checksums computed at the client/API and verified at every hop — not just disk-level CRCs.
  • Owner: Storage SRE owns the scrub SLO.

4.5 Multipart Orphans — The Invisible Bill#

  • Symptom: A tenant's bill includes 800 TB not visible in any LIST.
  • Root cause: A backup tool starts multipart uploads, crashes, restarts with a new upload ID; old parts accumulate for 2 years.
  • Detection: multipart.incomplete_bytes{tenant}, age histogram of open uploads.
  • Mitigation: Abort uploads older than N days; notify the tenant before abort.
  • Prevention: Platform default lifecycle rule "abort incomplete uploads after 7 days" on new buckets; incomplete bytes shown in billing.
  • Owner: Tenant for the tool; platform for defaults and visibility.

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Correlated drive failuresdisk_afr{model}, stripes at min redundancyStripes on the bad cohortPriority repair, drain, placement exclusionStorage SRE + hardware
AZ lossAZ health, fragment availability by AZAll stripes lose ≤ 5 of 15 fragmentsDegraded reads, repair after AZ returns or re-protectStorage SRE
Metadata hot partitionpartition_qps, split rateOne tenant prefixThrottle, split, key guidanceMetadata team + tenant
GC deletes live datagc.marked_bytes_ratio, audit mismatchPotentially all new objectsGrace period, kill switchStorage team (durability owner)
Bit rotscrub.checksum_failures, scrub ageSingle fragmentsRewrite from parityStorage SRE
Metadata replica divergenceConsensus log mismatch, checksum of partition snapshotsOne partitionRebuild replica from logMetadata team
Multipart orphansmultipart.incomplete_bytesOne tenant's billAbort policyTenant + platform defaults
Accidental customer deleteDelete-rate anomaly per bucketOne bucketVersioning, restore previous versionTenant (with platform tooling)

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
ArchitectureAPI + metadata DB + storage nodesTwo planes with different consistency; data first, metadata commit lastPlans which planes are shared org infrastructure vs per-product
DurabilityQuotes 11 nines, 3 replicasComputes it; names repair time as the lever; designs for correlated failuresSets durability posture: failure-domain taxonomy, delete protection, audits, game days
RedundancyReplicationSize-based hybrid; zone-aware EC; repair bandwidth budgetPrices code choices in $ and headcount at fleet scale
ConsistencyEventual for immutable blobsStrong read-after-write incl. LIST via metadata linearizationTreats the consistency contract as irreversible once published
MetadataSharded DBRange partitioning, split-on-load, per-prefix throttlingNamespace model as a multi-year decision; hierarchical as a separate product
LifecycleDelete = remove bytesSoft delete, delayed GC, orphan cleanup, opt-in archiveChargeback and defaults that change org behavior

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Commit protocol"The metadata commit is the linearization point; a crash before it leaves garbage, never a dangling name."
Durability math"Loss probability scales with repair time to the power of the parity count, so repair speed is my main lever."
AZ-aware coding"RS(10,4) across three AZs doesn't survive an AZ; parity has to cover the largest AZ share."
LIST awareness"Hash partitioning kills ordered LIST; I'll range-partition and split hot ranges."
GC paranoia"The garbage collector is the only component designed to delete data. It gets a delay and a kill switch."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"11 nines because three replicas" without mathDurability as an assertion, not an engineering property
Metadata and data in the same storeCouples a KB-scale strongly consistent problem to a PB-scale throughput problem
Hash-partitioned metadata with no LIST planLIST becomes scatter-gather over every shard
Immediate physical deleteEvery bug and fat-finger becomes permanent loss
No mention of repair or scrubbingDurability decays silently without them

5.4 Common False Positives#

  • Reed-Solomon math fluency ≠ storage design. Knowing Galois fields doesn't show you know where to put the fragments.
  • Big capacity numbers ≠ scale thinking. "140 PB raw" is arithmetic; "a 24 TB disk rebuild reads ~216 TB cluster-wide" is scale thinking.
  • Naming Ceph/MinIO ≠ design. Tools implement choices; the interview is the choices.
  • "We use checksums" ≠ integrity. Where they're computed (end-to-end vs per-disk) is the point.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minCommit to general-purpose store, 100 PB, 10¹² objects, 11 nines, strong consistency
Entities + API3–5 minVersion row, manifest, stripe; six operations
Architecture5–10 minTwo planes; PUT sequence; commit last
Deep dive: durability10–20 minMath, codes, AZ-aware placement, correlated failures
Deep dive: metadata20–30 minRange partitioning, LIST, hot prefixes, consistency
Deep dive: lifecycle30–40 minMultipart, delete/GC, tiering
Wrap-up40–45 minOwnership, cost, evolution

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Direction
"An entire AZ goes down"AZ-aware placementParity ≥ per-AZ fragments; degraded reads; don't mass-repair a transient outage
"Customers want 10× more small files"Small-object economicsPacking into containers; metadata becomes the cost
"Add cross-region replication"Async semanticsPer-bucket, async, replication lag metric, conflict rules for versions
"How do you know you haven't lost data?"Observability of durabilityScrub SLO, reference audits, at-risk stripe counts
"Cut storage cost 30%"EconomicsWider codes for cold data, auto-tiering, orphan cleanup, chargeback

6.3 What to Deliberately Skip#

  • Reed-Solomon encoding math beyond "any k of n."
  • HTTP details, signing algorithms — "standard signed requests."
  • Disk filesystem internals — "append-only extents on raw devices."
  • CDN design — a separate problem; see CDN & Edge Caching.

6.4 Follow-Up Questions to Expect#

  1. "Walk me through a GET when 2 of the needed fragments are unavailable."
  2. "How does LIST stay consistent while a partition is splitting?"
  3. "What happens if the client retries a PUT whose first attempt actually succeeded?"
  4. "How would you implement object lock (WORM) for compliance?"
  5. "How long does it take to re-protect after an AZ comes back, and do you repair during the outage?"
  6. "How do you migrate 50 PB to a new storage-node format without downtime?"
  7. "How would you bill tenants fairly for requests vs bytes?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design S3."

Staff Answer

"A few questions to pin the design: object size distribution, total scale, durability and consistency targets, and whether this is a general platform or a single-purpose store like a media origin or backup target.

I'll assume the general-purpose multi-tenant store: ~100 PB, 10¹¹–10¹² objects, median ~100 KB with a long tail to terabytes, 11 nines designed durability, strong read-after-write including LIST.

That gives two systems. A metadata plane — ~10¹² small rows, strongly consistent, range-partitioned for ordered LIST. And a data plane — petabytes of append-only, erasure-coded fragments. I'll walk the PUT path to show the commit protocol, then spend most of the time on durability math and the metadata plane."

Why this is L6:

  • Commits to an intent that contains the others as tiers
  • Splits the problem into two planes before drawing boxes
  • Names where the depth will go

What L7 adds:

  • Asks which internal teams will depend on this and what $/GB-month the business needs to beat the cloud alternative
  • Notes that the consistency contract, once published, can never be weakened
❌ Common L5 Trap

"Clients upload through an API gateway to storage servers. We keep a metadata database with the file location and replicate each file three times."

Why this misses: Correct shape, no commitment. The interviewer asks "what if the metadata write fails after the bytes are stored?" and "how does LIST work?" — and the design has no answer yet because it never separated the planes.


Drill 2: Compute the Durability#

Prompt: "Prove 11 nines."

Staff Answer

"Assume disk AFR λ = 1.5%/yr and a repair window T of 6 hours (6.85×10⁻⁴ years). Data is lost only if m+1 fragments of one stripe fail within overlapping repair windows.

For 3× replication: P ≈ 3λ · (λT)² = 0.045 × (1.03×10⁻⁵)² ≈ 4.8×10⁻¹² per object per year — about 11.3 nines. If repair takes 24 hours instead, it's ~7.6×10⁻¹¹ — only about 10 nines. Repair time enters squared.

For RS(10,4): P ≈ 14λ · C(13,4) · (λT)⁴ ≈ 1.7×10⁻¹⁸. Wide codes make independent disk failure a non-issue.

Two caveats. First, 'per object' vs fleet: at 10¹² objects, 10⁻¹¹ still means ~10 lost objects a year. Second, this assumes independent failures. Real loss comes from correlated events — a rack, a drive batch, an AZ — and from software. So the design work is placement across failure domains and protecting against our own bugs."

Why this is L6:

  • Derives the number and shows the sensitivity to repair time
  • Distinguishes per-object from fleet-level expectations
  • Names what the model leaves out and pivots there

What L7 adds:

  • Frames durability as an org-level posture with explicit budgets for software-caused loss, verified by independent audits
❌ Common L5 Trap

"Each disk has maybe a 1% chance of failing, so three replicas give 1% × 1% × 1% = one in a million."

Why this misses: It ignores repair — the failures must overlap within the repair window, which is what makes the number 10⁻¹¹ instead of 10⁻⁶. Without repair time in the model, the candidate can't reason about the most important durability lever.


Drill 3: The PUT Commit Protocol#

Prompt: "The bytes are written but the metadata write times out. What does a GET return?"

Staff Answer

"The metadata commit is the linearization point, so the answer depends on whether the commit actually happened — a timeout doesn't tell us.

The API server returns 500 or a timeout to the client. It does not know whether the version row committed. Two cases:

  • Commit didn't happen: GET returns the previous version (or 404). The fragments are orphans; GC reclaims them after its grace period.
  • Commit happened: GET returns the new version, even though the client saw an error.

Either way, the client retries. For safety, each PUT carries a client-generated request ID or Content-MD5; a retry of an already-committed PUT with the same content produces a new version with identical bytes — harmless. With If-None-Match: * (create-only), the retry gets 412 and the client can HEAD to confirm.

What never happens: a version row pointing at fragments that don't exist, because we only commit after the write quorum of fragments is durable."

Why this is L6:

  • Treats timeout as 'unknown', not 'failed'
  • Uses write ordering to make the only failure residue garbage
  • Gives the client a concrete idempotency path

What L7 adds:

  • Makes 'data first, metadata last, GC with delay' an invariant tested by fault-injection in CI for every storage release
❌ Common L5 Trap

"We roll back — delete the bytes that were written."

Why this misses: Rolling back on timeout races the commit: if the metadata write actually succeeded, deleting the bytes creates a visible object with missing data — the one failure mode the protocol exists to prevent.


Drill 4: LIST at Trillions of Keys#

Prompt: "How does LIST prefix=logs/2026/ work with 10¹² objects?"

Staff Answer

"Metadata is range-partitioned by (bucket_id, key), so all keys with that prefix are contiguous across one or a few adjacent partitions. LIST is an ordered range scan starting at the prefix, returning up to 1,000 keys and a continuation token that encodes the last key returned. The next page starts strictly after that key — which also makes pagination stable across partition splits, because the token is a key, not a partition offset.

Delimiter queries (delimiter=/) roll up common prefixes; the scan skips ahead past each rolled-up prefix instead of reading every key under it.

Consistency: each page reads from the partition's leader (or a follower with a valid read lease), so a key committed before the LIST started is visible.

Cost: a LIST page is ~one partition read — far more expensive than a GET's point lookup, so it's priced and rate-limited separately."

Why this is L6:

  • Connects LIST to range partitioning
  • Continuation token as a key survives splits
  • Notes delimiter skip-scan and prices LIST distinctly

What L7 adds:

  • Offers an inventory/manifest export for tenants who LIST billions of keys daily — changing an expensive access pattern rather than scaling it
❌ Common L5 Trap

"Query the metadata database with WHERE key LIKE 'logs/2026/%'."

Why this misses: On a hash-sharded store this is a scatter-gather across every shard, sorted in memory, for every page. It works in a demo and melts at 10¹² keys.


Drill 5: The Hot Prefix#

Prompt: "One tenant writes 50K PUT/s to keys starting with the current timestamp. What happens?"

Staff Answer

"Every write lands on the tail partition of that bucket's key range — monotonically increasing keys defeat split-on-load, because after each split the new tail is hot again. A partition leader handles maybe 5–10K writes/s, so the tenant sees throttling.

Platform response: return 503 SlowDown with Retry-After, per prefix, so the throttling is confined to this tenant; ensure partition leaders for hot tails aren't co-located on the same metadata host.

The real fix is key design, and it belongs to the tenant: prefix with a short hash (a7f3/2026-09-30T...) so writes spread over N ranges, or use a device ID first. We publish that guidance and show per-prefix throttling in their dashboard.

I would not make hash-partitioning the default to 'fix' this — it destroys ordered LIST for everyone."

Why this is L6:

  • Explains why split-on-load fails for monotonic keys
  • Contains blast radius to one tenant
  • Assigns the fix to the right owner

What L7 adds:

  • Adds key-design review to the onboarding checklist for tenants above a request-rate threshold
❌ Common L5 Trap

"Add more metadata shards."

Why this misses: More shards don't help a single hot key range; the tail is still one partition.


Drill 6: Multipart Upload#

Prompt: "A customer uploads a 2 TB file over a flaky connection."

Staff Answer

"Multipart upload. The client initiates and gets an upload ID, splits the file into parts (say 256 MB each — about 8,000 parts, under the 10,000 limit), and uploads parts in parallel, e.g., 16 at a time. Each part is independently chunked, erasure-coded, and stored; the server returns an ETag per part. A failed part is retried alone.

Complete sends the ordered list of (part number, ETag). The server validates the parts and commits one version row whose manifest references all part chunks — no copying of 2 TB. The object's ETag is a hash of part ETags plus a part count, which is why multipart ETags aren't the MD5 of the file.

Abandoned uploads: parts exist without a version row and cost money. A lifecycle rule aborts incomplete uploads after 7 days; the platform sets it by default on new buckets and shows incomplete bytes in billing."

Why this is L6:

  • Right part sizing against the limits
  • Completion is a metadata-only atomic commit
  • Owns the orphan-cost failure mode

What L7 adds:

  • Makes 'incomplete multipart bytes' a line item in every team's chargeback, so the cost has an owner
❌ Common L5 Trap

"Stream it in one request with resumable offsets."

Why this misses: A single stream is serial and ties a server to one connection for hours; resuming requires server-side partial state. Parallel parts are faster and each part's failure is independent.


Drill 7: Small Objects#

Prompt: "80% of objects are under 16 KB. Does erasure coding still work?"

Staff Answer

"Not per object. A 16 KB object in RS(9,6) becomes 15 fragments of ~1.8 KB — each with its own metadata, index entry, and disk IO. Overhead and IOPS dominate.

Pack small objects into large containers — e.g., 1 GB append-only volumes (Haystack-style). The metadata row stores (volume, offset, length). Volumes are sealed when full and then erasure-coded as a unit. Until sealed, the open volume is replicated 3× so writes are fast.

Deletes in packed volumes leave holes; compaction rewrites volumes when garbage exceeds ~30–50%. That's a background cost, budgeted like repair.

The metadata plane is now the bottleneck: at 10¹² small objects it's ~0.3–1 PB of rows. Pricing should reflect it — per-request and per-object fees matter more than per-GB for small-object tenants."

Why this is L6:

  • Quantifies why per-object EC fails for small objects
  • Packing with replicate-then-seal-then-EC
  • Names compaction cost and the metadata bottleneck

What L7 adds:

  • Aligns pricing with cost drivers (requests and object count), so small-object workloads fund their own metadata scale
❌ Common L5 Trap

"Use the same erasure coding — it's just smaller fragments."

Why this misses: Ignores per-fragment fixed costs; the IOPS and metadata bill grows 15× for no durability gain.


Drill 8: Build vs Buy#

Prompt: "We run 8 PB on-prem. Build our own object store?"

Staff Answer

"No. At 8 PB, the choices are a public cloud object store or a mature open-source system like Ceph RGW or MinIO on owned hardware. Building a new one means writing an erasure-coding data plane, a consensus-backed metadata plane, repair, scrubbing, and GC — and then carrying a team for it forever.

The decision between cloud and self-hosted: egress patterns (if data is read heavily from on-prem compute, cloud egress fees dominate), regulatory residency, and whether we have ~3–5 engineers for storage operations. At 8 PB, cloud standard storage is roughly $150–200K/month list; self-hosted hardware amortized over 5 years at ~1.5× EC is cheaper per GB, but the people cost closes most of the gap.

What we should build either way: lifecycle defaults, cost attribution per team, and a backup copy in an independent system."

Why this is L6:

  • Rejects the build at this scale with reasons
  • Compares the realistic options on the right axes
  • Identifies the part that's worth building

What L7 adds:

  • Plans the exit: data formats and tooling that keep a later cloud-to-on-prem (or reverse) migration feasible
❌ Common L5 Trap

"Build it — we'll have full control and save cloud costs."

Why this misses: No TCO; ignores that durability engineering is a permanent team, not a project.


Drill 9: Changing the Storage Format Without Losing Data#

Prompt: "We need to migrate 50 PB from 3× replication to erasure coding."

Staff Answer

"It's a background transcode with a strict invariant: never remove the old copies until the new stripe is verified and the metadata points at it.

Per object: read a replica, encode, write fragments, verify checksums by reading back a sample, atomically swap the manifest in the metadata row (conditional on the old manifest), then enqueue old replicas for GC after the usual grace period.

Rollout: shadow first — transcode 0.1% and run the scrubber and a read-verify on everything transcoded; then ramp by storage class and age (coldest first, since they're read least and save the most). Rate-limit by cluster bandwidth — maybe 5–10% — so customer latency doesn't move. At 50 PB and ~20 GB/s of spare bandwidth, it takes ~1 month.

Metrics: transcode progress, verify failures (must be 0), manifest-swap conflicts, GET p99 for transcoded vs not."

Why this is L6:

  • Explicit safety invariant and atomic manifest swap
  • Coldest-first ordering with a bandwidth budget
  • Calculates the duration

What L7 adds:

  • Makes the savings visible: 50 PB × (3.0 − 1.67) ≈ 66 PB of raw disk freed — a hardware purchase avoided that funds the project
❌ Common L5 Trap

"Write a script that converts each file and deletes the old replicas."

Why this misses: Deleting before verifying and before the metadata points at the new stripe creates a window for permanent loss.


Drill 10: Cross-Region Replication#

Prompt: "A tenant needs a copy in another region for disaster recovery."

Staff Answer

"Per-bucket asynchronous replication. After a local commit, the version is appended to a replication log partitioned by (bucket, key); a replicator copies the bytes and commits the same version ID in the destination bucket. Ordering per key is preserved; versions make it idempotent.

Guarantees to state: RPO equals replication lag — typically seconds, p99 minutes, unbounded during a regional incident. Metric replication.lag_seconds{bucket} and the backlog in bytes; an optional SLA tier (e.g., 99.99% of objects within 15 minutes) is a product decision.

Deletes: replicate delete markers by default, but not permanent version deletes — otherwise a malicious or buggy delete in the source destroys the DR copy. That's the point of DR.

Cost: storage in both regions plus inter-region transfer per GB — make it visible."

Why this is L6:

  • Async with explicit RPO and a lag metric
  • Idempotency via version IDs
  • Delete semantics designed for the DR purpose

What L7 adds:

  • Requires DR copies of critical data to be in a separately administered account with different credentials — protection against compromised admins, not just regions
❌ Common L5 Trap

"Write synchronously to both regions."

Why this misses: Every PUT pays a cross-region round trip (~60–150 ms), and a regional outage blocks writes everywhere — availability sacrificed for an RPO most tenants didn't ask for.


8. Deep Dive Scenarios#

Deep Dive 1: Peak Traffic — The Viral Object Miss Storm#

Context: A video embedded in a news story goes viral. The CDN in front of the media bucket is purged by an unrelated config change. 400K req/s hit the origin for one 200 MB object and ~5K related objects. Origin p99 GET climbs from 80 ms to 9 s, affecting every tenant in the cell.

Questions to Surface First:

  • Is this one object or broad traffic? Which component saturates — API servers, the metadata partition, or the storage nodes holding those fragments?
  • Why did the CDN purge, and can we restore the cache?
  • Are other tenants in the same cell affected, and why?

Typical L5 Approach: Scales out API servers. The bottleneck is the 15 storage nodes holding that object's fragments — each serving hundreds of Gbps — so more API servers add more load on the same disks.

Staff Approach: Identifies a hot-object problem, not a capacity problem. Adds a hot-object cache tier in the front end (in-memory, keyed by version ID so it's always consistent), request coalescing so concurrent GETs of one object share one backend read, and per-tenant request isolation so the media tenant's storm doesn't consume shared API capacity. Works with the CDN team to restore caching with a staggered warm-up.

Principal Approach: Asks why a CDN purge could be global and instant. Pushes for purge rate limits and staged purges as a CDN platform standard, and for cell-based isolation in the storage front end so one tenant's storm has a bounded blast radius by construction.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Identify hot keys from front-end sampling; enable request coalescing and hot-object caching for the top 100 version IDs
TriageConfirm saturation on the fragment-holding nodes; check cross-tenant impact
Quick fixPer-tenant concurrency limits; CDN re-enable with origin shield
GuardrailsWatch storage.node_egress_gbps and GET p99 for other tenants
Post-mortemHot-object cache as a permanent tier; purge governance; cell isolation

Metrics to Watch: api.get_qps_by_version_topk, storage.node_egress_gbps, api.latency_p99{tenant}, cdn.origin_hit_ratio

Organizational Follow-up: CDN team adds staged purges; storage publishes an origin capacity contract per tenant.

Ownership Question: "Who owns origin protection — CDN or storage?" Staff answer: Both, with a written contract: CDN guarantees staged purges and origin shielding; storage guarantees per-tenant isolation so a miss storm degrades only the offending tenant.

Key Takeaway: "Immutable versions make caching trivially consistent — use that. The bottleneck of a hot object is its fragments, not the API tier."

What clears the Staff bar:

  • Locates the real bottleneck (fragment holders)
  • Uses version IDs for safe caching
  • Contains blast radius per tenant

Deep Dive 2: The Silent Failure — Scrubber Stopped Six Weeks Ago#

Context: During a routine review you notice scrub.cycle_age_max is 44 days against a 14-day SLO. Nobody was paged. A config change moved the scrubber to a lower-priority scheduling class and it's been starved by repair and transcode jobs.

Questions to Surface First:

  • How many fragments haven't been verified within SLO, and on which disks?
  • Were there latent errors in the last completed scrubs, and at what rate?
  • Why did a 3× SLO violation not page?

Typical L5 Approach: Restores the scrubber's priority and lets it catch up.

Staff Approach: Restores priority, then computes exposure: with a historical latent-error rate, how many stripes might now have hidden bad fragments plus a missing one? Prioritizes scrubbing stripes that already have a missing fragment (where a hidden error would mean loss) and disks with elevated SMART errors. Adds a page on scrub.cycle_age_max > 1.5× SLO and a floor on scrubber bandwidth that no scheduler class can take away.

Principal Approach: Recognizes a class of problem: background durability jobs (scrub, repair, GC audits) have no customer-visible symptom when they stop. Creates a "durability SLO" review for all background jobs, with dead-man alerts that fire when a job stops reporting, and reports these SLOs to leadership monthly like availability.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateRestore scrubber priority with a guaranteed bandwidth floor
TriageRank stripes: missing fragments + unscrubbed first
Quick fixTargeted scrub of at-risk stripes within 48 h
GuardrailsDead-man alert on scrub progress; SLO page threshold
Post-mortemScheduler changes that affect durability jobs require durability-owner review

Metrics to Watch: scrub.cycle_age_max, scrub.bytes_verified_per_day, scrub.checksum_failures, storage.stripes_by_surviving_fragments

Organizational Follow-up: Durability job inventory with owners and dead-man alerts.

Ownership Question: "Who approves changes to background job scheduling?" Staff answer: The storage durability owner — any job that protects data is part of the durability budget, and its capacity is not up for grabs by other jobs.

Key Takeaway: "Durability jobs fail silently. Alert on their absence, not just their errors."

What clears the Staff bar:

  • Computes exposure before declaring victory
  • Prioritizes by risk, not by order
  • Fixes the alerting class, not just the instance

Deep Dive 3: Large-Customer Onboarding — 5 PB and 20 Billion Objects#

Context: An analytics customer will migrate 5 PB in 20B objects over 60 days, then run jobs that LIST ~200M keys per hour and read at 200 GB/s bursts.

Questions to Surface First:

  • What's their key layout? Sequential, hashed, partitioned by date?
  • Do they need LIST or could an inventory manifest replace it?
  • Peak write rate during migration? Will they use multipart for large files?

Typical L5 Approach: Checks total capacity and adds storage nodes.

Staff Approach: Capacity is the easy part. Reviews key design with them (date-partitioned with a hash component), pre-splits their metadata ranges based on expected key distribution, provides daily inventory exports so their jobs stop LISTing 200M keys an hour, sets per-tenant request limits appropriate to their tier, and places their data across enough storage nodes that 200 GB/s bursts don't concentrate. Migration throttled to protect other tenants.

Principal Approach: Turns this into a repeatable "large tenant onboarding" program with a checklist (key design review, pre-split, inventory, limits, cost forecast) and a capacity reservation contract. Negotiates pricing that reflects metadata and request costs, not just bytes.

Staff Approach — Full Reasoning
PhaseWhat to Do
Design reviewKey layout; LIST patterns; object size distribution
Pre-workPre-split metadata ranges; reserve capacity; inventory export
MigrationThrottled ingest (e.g., 20B objects / 60 days ≈ 3,900 PUT/s average, cap at 10K)
ValidationObject count and checksum reconciliation against source
Steady statePer-tenant dashboards: throttles, LIST cost, bytes by tier

Metrics to Watch: metadata.partition_qps{tenant}, api.throttled_total{tenant}, list.keys_scanned_per_hour{tenant}, storage.cell_utilization

Organizational Follow-up: Account team and storage agree on capacity reservation lead times (e.g., 90 days for > 1 PB).

Ownership Question: "Who decides this customer's request limits?" Staff answer: The storage platform sets limits by tier with a documented exception process; the account team can request, not grant.

Key Takeaway: "Large tenants break metadata, not disks. Onboard their key design, not just their bytes."

What clears the Staff bar:

  • Focuses on metadata and access patterns over raw capacity
  • Replaces expensive LIST with inventory
  • Protects other tenants during migration

Deep Dive 4: Post-Mortem — A Tenant Lost 3% of a Bucket#

Context: A tenant reports 3% of objects in a bucket return 404. Investigation shows their own IAM role, compromised via a leaked CI credential, issued DELETE on 1.2M keys. Versioning was off. GC ran after 24 hours. The data is gone.

Questions to Surface First:

  • Was the platform working as designed? (Yes.)
  • What defaults allowed an unversioned production bucket?
  • Why could one credential delete 1.2M objects without any friction?

Typical L5 Approach: "Customer error; recommend they enable versioning."

Staff Approach: Accepts that "working as designed" isn't a defense when the design makes a common mistake permanent. Changes defaults: versioning on for new buckets; anomaly detection on delete rate per bucket (1.2M deletes in 10 minutes vs a baseline of 100/day) that alerts the tenant; a longer GC grace period for bulk-deleted data (e.g., 7 days) during which support can restore.

Principal Approach: Frames it as a platform-level safety posture: soft-delete by default, object lock for critical data, deletion requiring separate credentials for production buckets, and a recommended cross-account backup. Writes the "data protection defaults" standard and measures adoption across all buckets.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateCheck whether any fragments remain un-GC'd; halt GC for the bucket
TriageConfirm deletes came from a valid credential; engage security
Quick fixRecover anything still within grace; support the tenant's restore from their own backups
GuardrailsDelete-rate anomaly alerts; extended grace for bulk deletes
Post-mortemDefault versioning; delete-protection tiers
t=0:       Leaked credential used from unfamiliar IP
t=+3min:   DELETE rate on bucket: 2,000/s (baseline ~0.001/s)
t=+10min:  1.2M objects deleted; no alert exists for delete rate
t=+24h:    GC reclaims chunks
t=+3d:     Tenant notices missing files from user reports

Metrics to Watch: api.delete_rate{bucket} vs baseline, gc.bytes_reclaimed{bucket}, security.credential_anomalies

Organizational Follow-up: Security owns credential anomaly detection; storage owns delete safety defaults.

Ownership Question: "Is this the customer's problem or ours?" Staff answer: The credential leak is theirs; the fact that one credential could permanently delete 1.2M objects in 10 minutes with no alert and a 24-hour grace is ours.

Key Takeaway: "The biggest durability risk is a valid DELETE. Design defaults so the common mistake is recoverable."

What clears the Staff bar:

  • Refuses 'working as designed' as a closing statement
  • Changes defaults, not just documentation
  • Separates security and storage responsibilities cleanly

Deep Dive 5: Multi-Region Expansion — Launching a New Region#

Context: The company is launching the object store in a new region with 3 AZs. Initial demand is 2 PB; tenants want a global namespace for bucket names and cross-region replication from day one.

Questions to Surface First:

  • What must be global (bucket-name uniqueness, IAM) vs regional (metadata, data)?
  • Is the new region's metadata plane independent of existing regions?
  • Can we launch with a smaller cluster without weakening durability?

Typical L5 Approach: Extends the existing metadata cluster to the new region for a single global namespace.

Staff Approach: Keeps every region's metadata and data planes fully independent — a region failure must not affect others. Only the bucket-name registry is global (low write rate, strongly consistent, cached everywhere). Replication is async and per bucket. Small initial footprint uses the same code but checks that enough racks and nodes exist per AZ for "one fragment per failure domain" — if not, launches with replication and transcodes later.

Principal Approach: Codifies regional independence as an architectural rule ("no cross-region calls on the data path") and tests it with region-isolation game days. Plans the capacity ramp and the hardware cost curve, and makes sure the global control plane (bucket names, IAM) has its own availability design, since it's the only shared fate.

Staff Approach — Full Reasoning
PhaseWhat to Do
ScopeGlobal: bucket registry, IAM. Regional: everything else
Durability at small scaleVerify placement can honor failure domains; otherwise replicate, then EC as capacity grows
ReplicationAsync per bucket; lag SLO; delete-marker semantics
ValidationRegion-isolation test: sever links, confirm local PUT/GET/LIST unaffected
RampCapacity plan for 2 PB → 20 PB over 2 years
Diagram: Deep Dive 5: Multi-Region Expansion — Launching a New Region

Metrics to Watch: replication.lag_seconds, registry.cache_staleness, region.cross_region_calls_on_data_path (must be 0)

Organizational Follow-up: Each region has its own on-call; the global control plane has a separate team with a higher availability target.

Ownership Question: "Who owns the global bucket registry?" Staff answer: The control-plane team, separate from any region, with a design that tolerates its own outage: existing buckets keep working, only new bucket creation stops.

Key Takeaway: "Regions share only what they must, and what they share must fail gracefully."

What clears the Staff bar:

  • Independent regional planes; minimal global state
  • Durability honored even at small launch scale
  • Tests isolation rather than assuming it

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Split blob storage into a metadata plane and a data plane and justify their different consistency models
  • Walk the PUT path and explain why data-first, metadata-last makes crashes leave only garbage
  • Compute durability for replication and erasure coding, and explain why repair time is the main lever
  • Choose a code whose parity covers the largest per-AZ fragment count
  • Explain range partitioning for LIST, why monotonic keys defeat split-on-load, and who fixes it
  • Design multipart upload, soft delete, delayed GC, and orphan cleanup
  • Name the correlated and software-caused failures that dominate real data loss
  • Price redundancy and tiering choices at three scales and name the one-way doors

The Bar for This Question#

Mid-level (L4): Designs an upload service that stores files on servers with a metadata table. Mentions replication. Struggles with large uploads and listing.

Senior (L5): Separate metadata DB and storage nodes, 3× replication across AZs, multipart upload, CDN for reads. A system that works — but durability is asserted, LIST is an afterthought, deletes are immediate, and cost is 1.8× what it needs to be.

Staff+ (L6): Two planes with an explicit commit protocol; durability computed with repair time; zone-aware EC with a size-based hybrid; range-partitioned strongly consistent metadata with hot-prefix handling; soft deletes, delayed GC with a kill switch, scrubbing SLOs; each failure mode with a metric and owner. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "Eleven Nines Is a Marketing Number — the Real Risk Is Software"#

EvidenceImplication
Independent-failure math for wide codes gives ~10⁻¹⁸Disks are not the bottleneck on durability
GC, metadata, and deploy bugs can touch every objectSoftware failures are correlated by nature
Customer deletes are valid operationsThe platform can't distinguish mistake from intent without friction

The Staff position: Spend durability engineering on delete safety, GC paranoia, staged deploys, and audits — not on adding parity.

Why this matters in interviews: Candidates who say "more replicas" to every durability question signal they've never seen a real data-loss incident.

10.2 "Eventual Consistency Was Never Cheaper — It Was Just Billed to Someone Else"#

EvidenceImplication
Ecosystem tools built consistency layers on top of eventually consistent S3 for yearsEvery consumer paid the cost separately
S3 moved to strong consistency without a price increaseThe provider could absorb it once
LIST-after-PUT anomalies silently drop data in pipelinesThe cost shows up as data quality bugs

The Staff position: Pay for consistency once in the metadata plane.

Why this matters in interviews: It reframes a "performance vs consistency" debate as "who pays."

10.3 "Repair Speed Matters More Than Parity Count"#

EvidenceImplication
Loss probability ∝ TᵐHalving T with m = 2 gives 4× improvement
24 TB drives take days to rebuild seriallyDeclustered, parallel repair is mandatory
Extra parity costs storage foreverFaster repair costs bandwidth only during failures

The Staff position: Declustered placement and prioritized repair before wider codes.

Why this matters in interviews: It shows you understand the durability model rather than memorizing codes.

10.4 "Most Blob Storage Cost Is Garbage"#

EvidenceImplication
Incomplete multipart uploads, old versions, and never-read logs accumulate indefinitelyBytes grow monotonically without lifecycle policy
Few teams write lifecycle rules unpromptedDefaults decide the bill
Access-age curves show most data goes cold within weeksTiering savings are large and mostly untaken

The Staff position: Defaults (abort multipart after 7 days, noncurrent-version expiration, auto-tier among instant tiers) and chargeback beat any storage-engine optimization.

Why this matters in interviews: Cost questions reward governance answers, not compression answers.

10.5 "Don't Offer Rename"#

EvidenceImplication
Atomic rename across partitions needs distributed transactionsIt taxes the metadata plane for one operation
Table formats commit via manifests, not renamesThe ecosystem already moved on
Emulated rename (copy + delete) is non-atomic and misleads usersA fake guarantee is worse than none

The Staff position: Flat namespace, no rename; offer hierarchical namespaces as a separate product only when demand justifies a second engine.

Why this matters in interviews: It shows you'll decline a feature to protect the core design.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

A Staff engineer designs a durable, consistent object store. A Principal engineer sees that object storage is the organization's system of record by accident: every team dumps data into it, few write lifecycle rules, nobody owns cross-team retention, and the bill grows 40% a year while half the bytes are never read. The durability problem is solved; the governance problem isn't. The L7 question is how to make the org's storage spend and data-protection posture correct by default — and which decisions (namespace, consistency contract, key formats, vendor) will be impossible to reverse once 200 services depend on them.

The Org-Level Fault Line#

One storage platform with enforced defaults vs team-owned buckets with full autonomy.

OptionWhat WorksWhat BreaksWho Pays
Full team autonomyTeams move fast; no platform gateNo lifecycle rules, unversioned prod buckets, inconsistent encryption, unowned buckets after reorgsFinance (runaway bill), security (exposure), data teams (loss)
Central platform with mandatory defaultsSafe and cheap by defaultSome workloads fight the defaults; exception process becomes a bottleneckPlatform team (exceptions), outlier teams
Paved road: defaults + chargeback + auditsDefaults protect, chargeback motivates, audits catch driftRequires cost attribution tooling and a policy enginePlatform team headcount (~2–3 engineers)

The L7 default: The paved road. Every bucket is created through a platform module with versioning, encryption, lifecycle defaults, an owning team tag, and a data classification. Chargeback is per team per month. Quarterly audits flag unowned and non-compliant buckets.

🧭 Principal Move: "I'm not trying to make storage cheaper by engineering the engine — I'm making it cheaper by making every byte have an owner and a lifecycle. That's where the 30–40% is."

Cost Model#

Assumptions: self-hosted at scale; ~$6/TB-month for raw HDD capacity all-in (hardware amortized over 5 years, power, space, network); ~1.67× overhead for zone-aware EC; metadata and front-end fleet ~20–30% of data-plane cost; fully loaded engineer ~$25K/month. Public-cloud standard storage list price for comparison is on the order of ~$21–23/TB-month.

ScaleLogical DataInfra $/monthHeadcountOn-call LoadCloud List Equivalent
Small1 PB~$40–60K (minimum failure-domain footprint dominates; raw disk alone ~$10K)4–6 (~$125K)1 rotation, frequent hardware tickets~$22K + requests — buy
Medium100 PB~$1.3–1.5M (raw ~$1.0M, metadata/front end ~$0.3–0.5M)20–30 (~$600K)Storage SRE + metadata rotations~$2.1M + requests + egress
Large1 EB~$12–15M80–150 (~$2–4M)Multiple rotations, hardware ops team~$20M+ — building can pay off

The number that matters to executives: below roughly tens of petabytes, building your own object store costs more than buying once people are counted. Above hundreds of petabytes, with steady growth and low egress needs, owning can save ~30–50% — but only with a team that treats durability as a permanent product.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

Triggers: Year 1 when storage becomes a top-3 infra cost; Year 2 when the first tenant needs DR or compliance retention; Year 3 when a second region launches or a drive generation reaches end of life.

One-Way Doors vs Two-Way Doors#

DecisionDoorReversal Cost
Consistency contract (strong read-after-write)One-wayWeakening breaks unknown consumers; can only be strengthened
Namespace model (flat vs hierarchical)One-wayMigrating semantics of billions of keys and every client
Public API shape and ETag semanticsOne-waySDKs and customer code depend on it
Durability claims published to customersOne-wayCan't lower a promise without trust damage
Erasure code parametersTwo-way (expensive)Background transcode of all data; months of bandwidth
Placement policyTwo-wayRebalance over weeks
Metadata storage engineTwo-way (expensive)Online migration with dual-write and verification
Tiering thresholdsTwo-wayConfig change

The Standard I'd Write#

RFC: Object Storage Data Protection & Cost Standard v1

Scope: All buckets holding production or customer data.

MUST:

  • Create buckets through the platform module; every bucket has an owning team, a data classification, and a cost center tag.
  • Enable versioning and encryption at rest on production buckets; noncurrent versions expire after a class-defined period (default 30 days).
  • Abort incomplete multipart uploads after 7 days.
  • Separate deletion privileges from write privileges for production buckets; bulk deletes above 10K objects/hour trigger an alert to the owner.
  • Tier-0 data (customer records, financial ledgers) has a cross-account backup with object lock.

SHOULD:

  • Use lifecycle rules to move data not read in 90 days to infrequent-access tiers.
  • Use hashed or high-cardinality key prefixes for write rates above 1K/s.

Exceptions: Filed with the storage platform team; approved by platform lead and the data owner's director; reviewed every 6 months.

Success metrics: 100% of buckets owned; < 5% of bytes in incomplete uploads or orphaned versions; storage cost growth ≤ data growth; zero unrecoverable deletions of Tier-0 data.

What I'd Tell the VP#

Our storage is extremely unlikely to lose data because of hardware; the realistic risks are our own software bugs and accidental or malicious deletions, and our defaults today don't protect against those. I'm proposing a standard that turns on versioning and delete protection for production data and gives every bucket an owner. The same change attacks cost: today roughly a third of what we pay for is data nobody reads or has forgotten about, and chargeback plus lifecycle defaults should cut storage spend growth by 25–35% within a year. It needs about three engineers for two quarters. We should keep buying storage from our cloud provider until we're well past a hundred petabytes; building our own engine is not where the savings are.

Principal Interview Signals#

SignalWhat It Sounds Like
Governance over engine"The bill is driven by unowned bytes, not by the erasure code."
Priced tradeoffs"Zone-aware EC costs 0.27× more than RS(10,4) — at 100 PB that's ~27 PB of disk, the price of surviving an AZ."
One-way doors"The consistency contract and the namespace model are forever; the code parameters aren't."
Build vs buy with thresholds"Below tens of petabytes, buy; the team costs more than the disks."
Org failure posture"Tier-0 data gets a cross-account, object-locked copy — protection against our own admins."

Staff answers that L7 interviewers find insufficient:

  • "We'll add versioning" — right mechanism, no plan to make it the default across 2,000 buckets.
  • "Erasure coding saves 44%" — correct math, no decision about whether the org should run the engine at all.
  • "Storage team owns durability" — ignores that deletes, lifecycle, and retention are owned by hundreds of tenant teams.

Appendices

Appendix A: Durability Math in Depth#

A.1 The Model#

For a stripe of n fragments tolerating m losses, with per-disk failure rate λ (per year) and repair window T (years):

P_loss_per_year ≈ n·λ · C(n−1, m) · (λT)^m

Read it as: rate at which some fragment fails (n·λ), times the probability that m of the remaining n−1 also fail before repair completes.

SchemenmT = 6 hT = 24 h
3× replication32~4.8×10⁻¹²~7.6×10⁻¹¹
RS(6,3)93~8×10⁻¹⁵~5×10⁻¹³
RS(10,4)144~1.7×10⁻¹⁸~4×10⁻¹⁶
RS(9,6)156≪ 10⁻²⁰≪ 10⁻¹⁸

(λ = 0.015/yr. Order-of-magnitude only; the model assumes independent failures and exponential lifetimes.)

A.2 What the Model Omits#

Omitted FactorEffectDefense
Correlated hardware failure (rack, PDU, drive batch)Multiple fragments fail togetherOne fragment per failure domain; drive-model diversity
Latent sector errorsA "surviving" fragment is actually badScrubbing; end-to-end checksums
Repair backlog during mass failureT grows exactly when failures clusterRepair priority by surviving fragments; bandwidth headroom
Software bugs (GC, metadata, placement)Can affect all objects at onceDelays, audits, canaries, kill switches
Operator and customer errorValid deletesVersioning, object lock, grace periods

A.3 Copysets — Per-Object vs Fleet Loss#

With random placement across thousands of disks, almost any combination of m+1 simultaneous disk failures holds all surviving fragments of some stripe — so fleet-level "we lost something" events become frequent even when per-object probability is tiny. Restricting placement to a limited number of disk groups (copysets, Cidon et al., 2013) makes loss events rarer but larger. Choose deliberately: many small losses (random placement) or rare larger ones (copysets). Most object stores prefer fewer loss events, because each one is an incident with customer notification.

A.4 Repair Bandwidth#

bytes_read_per_disk_loss = disk_capacity × k        (RS: read k fragments per rebuilt fragment)
LRC single-loss read      = disk_capacity × group_size

24 TB disk, RS(9,6):  ~216 TB read cluster-wide
Declustered over 2,000 disks at 50 MB/s each for repair: ~100 GB/s -> ~36 min

Appendix B: Metadata Data Model#

Table object_versions  -- range-partitioned by (bucket_id, key)
  PK: (bucket_id, key, version_id DESC)
  cols: size, etag, checksum_crc64, storage_class, is_delete_marker,
        manifest (inline for <= 8 chunks, else pointer), created_at, kms_key_ref

Table uploads          -- partitioned by (bucket_id, key, upload_id)
  parts: map<part_no, {etag, size, chunk_refs}>, initiated_at

Table chunks           -- partitioned by chunk_id (hash)
  stripe_id, offset, length, refcount_hint

Table stripes          -- partitioned by stripe_id (hash)
  code, fragments[ {node_id, extent_id, offset, crc} ], sealed_at

Byte lifecycle — the states that the commit protocol and GC must respect:

Diagram: Appendix B: Metadata Data Model

Version IDs are time-ordered (e.g., hybrid logical clock + node ID) so "latest" is the first row in the key's range.

Appendix C: Coordination Mechanisms#

MechanismUsed ForNotes
Raft/Paxos per metadata partitionLinearizable version commits3–5 replicas across AZs; see Distributed Consensus
Leader leases / read leasesServe strong reads without a consensus roundBounded clock drift assumption
Conditional writes (If-Match)Single-key CAS for clientsNo multi-key transactions
Partition manager with fencingSplits, merges, leader movesFencing tokens prevent split-brain writes
Placement serviceFragment targetsStateless over a cluster map; membership via ZooKeeper & etcd
Witness/version check for cachesFront-end cache coherenceCache entries validated against partition version

Quick comparison: A single global metadata database doesn't scale past ~10⁹–10¹⁰ rows comfortably; hash-sharded KV breaks LIST; range-partitioned consensus groups are the standard answer (the same shape as Spanner, CockroachDB, or TiKV).

Appendix D: API Contract and Client Behavior#

BehaviorContract
IdempotencyPUT of identical content is safe to retry; If-None-Match: * for create-only
IntegrityClient sends Content-MD5 or a CRC checksum header; server verifies before commit
Throttling503 SlowDown with Retry-After; SDK uses exponential backoff with jitter
Large objectsMultipart above ~100 MB; parts 5 MiB–5 GiB; ≤ 10,000 parts
Range readsRange: bytes=a-b; parallel ranged GETs for throughput
ConsistencyStrong read-after-write for PUT/DELETE → GET/HEAD/LIST
Paginationcontinuation-token is opaque, key-based, stable across splits

Appendix E: Observability#

E.1 Core Metrics#

Durability:  storage.stripes_by_surviving_fragments, repair.queue_age_p99,
             scrub.cycle_age_max, scrub.checksum_failures, gc.marked_bytes_ratio
Metadata:    metadata.partition_qps, metadata.commit_latency_p99, metadata.split_rate
API:         api.latency_p99{op}, api.error_rate{op}, api.throttled_total{tenant}
Cost:        storage.bytes{tenant,class}, multipart.incomplete_bytes, versions.noncurrent_bytes
Data path:   storage.node_egress_gbps, get.reconstructed_reads_ratio

E.2 Critical Alerts#

AlertThresholdAction
At-risk stripesstripes with ≤ 1 spare fragment > 100Page
Scrub stalledscrub.cycle_age_max > 1.5× SLO or no progress in 1 hPage
GC anomalymarked bytes > 3× baselinePage + auto-pause GC
Metadata commit latencyp99 > 100 ms for 10 minPage
Delete anomalybucket delete rate > 100× baselineNotify owner

E.3 Control Plane vs Data Plane#

Control plane: bucket creation, policy, lifecycle config, placement maps, partition management — may be briefly unavailable. Data plane: GET/PUT/LIST on existing buckets — must keep working with cached control-plane state. A control-plane outage should block new bucket creation, not reads.

E.4 Debugging "Is Any Data at Risk Right Now?"#

  1. storage.stripes_by_surviving_fragments histogram — anything at minimum?
  2. Repair queue age and throughput — is repair keeping up?
  3. Failure-domain view — are failures clustered (rack, model, AZ)?
  4. GC and scrub health — both running, both within SLO?

Appendix F: Scale Evolution#

F.1 What Works at Each Scale#

ScaleArchitecture
< 100 TBBuy cloud storage. If self-hosted: MinIO/Ceph, replication
100 TB–10 PBOpen-source object store; EC for cold data; lifecycle defaults
10–100 PBDedicated storage team; zone-aware EC; small-object packing; chargeback
100 PB+Custom metadata plane, declustered repair, hardware co-design, multi-region

F.2 Multi-Region Path#

Independent regional planes; a global bucket registry and IAM; async per-bucket replication with lag SLOs; no cross-region calls on the data path (Deep Dive 5).

F.3 What You Don't Build on Day One#

  • Hierarchical namespace
  • Archive tiers with hours-long retrieval
  • Cross-region replication
  • LRC or custom codes (start with RS)
  • Object lock / compliance modes (until a tenant needs it)

Appendix G: Multi-Tenancy, Fairness, and Cost#

ConcernMechanism
Request fairnessPer-tenant and per-prefix token buckets at the front end; see Rate Limiting
Metadata isolationPartitions are per-bucket ranges; hot tenants split without touching neighbors
Data-plane isolationCell architecture: tenants assigned to cells; a cell's failure is bounded
Bandwidth fairnessPer-tenant egress shaping on storage nodes during contention
BillingSeparate dimensions: GB-month by class, requests by type (PUT/LIST priced above GET), egress, retrieval
Background budgetRepair, scrub, GC, transcode, compaction each get a guaranteed bandwidth share

🧭 Principal Insight: "Price what costs us money. If LIST and small objects are what load the metadata plane, bill for them — otherwise we subsidize the workloads that hurt us most."

Related reading: Large Blobs, File Sync, CDN & Edge Caching, Database Sharding, Consistency Models, and Sharding.

  1. Loading the index…