Technologies referenced in this case study: PostgreSQL · DynamoDB · Apache Kafka · Redis · Cassandra · Kubernetes
Related: Object Storage · CDN · Large File Uploads & Delivery · Video Streaming · Cloud File Sync · News Feed · Idempotency & Exactly-Once · Transactional Outbox · Security Fundamentals · Key Management Service
Reading Guide#
Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Design Splits table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal View) and the appendices on the upload state machine, tiering math and deletion |
What is Photo Upload and Storage? — Why interviewers pick this topic
A photo platform takes a 4 MB image from a phone on a flaky cellular link, stores it so it is never lost, turns it into the five or six sizes every screen needs, and serves those sizes to millions of viewers from a CDN — for years. Then, one day, the owner taps delete, and every copy of it, in every cache, replica, derivative and backup, has to actually go away.
The upload is the part everyone draws. The hard parts are around it: an upload that survives a train tunnel, a processing pipeline that never shows a broken image, a metadata record that never points at bytes that don't exist, a storage bill that grows by petabytes per year, and a delete that is honest.
Before vs After — the "half-uploaded album" scenario:
Without resumable uploads and an upload state machine:
t=0: User starts uploading a 40-photo album (160 MB) on a commuter train.
t=+45s: Tunnel. Connection drops at photo 23, 60% through a 4 MB single PUT.
t=+46s: App retries photos 1–40 from scratch. Photos 1–22 now exist twice.
t=+2min: Feed shows the album with 22 duplicates and 3 grey boxes
(metadata committed before thumbnails existed).
t=+1 day: Support tickets: "my album is broken". 18 TB/month of orphaned bytes
from abandoned uploads nobody is cleaning up.
With resumable chunked uploads, client upload IDs and a state machine:
t=0: App creates 40 upload sessions with client-generated upload IDs.
t=+45s: Tunnel. Photo 23 has 2 of 3 chunks acknowledged.
t=+3min: Signal returns. App asks the server which chunks exist, sends chunk 3.
t=+3min: Retried "create" calls return the existing sessions — zero duplicates.
t=+3min: Each photo appears in the album only when its thumbnails are READY;
until then the uploader sees a local preview, followers see nothing.
t=+7 days: Lifecycle rule aborts the 0.4% of sessions that never finished.
Why interviewers reach for this question: It looks like "presigned URL to S3, Lambda makes a thumbnail, CDN in front" — a Senior answer in five minutes. The Staff answer lives in what that picture hides: idempotent, resumable sessions; a state machine that decides when a photo is visible; processing as an untrusted-input pipeline; dedup that doesn't leak privacy; storage tiers priced in dollars per petabyte; and deletion that reaches the CDN, the derivatives and the backups.
Mechanics Refresher: Upload and Storage Primitives
| Primitive | How It Works | Pros | Cons |
|---|---|---|---|
| Proxy upload | Client sends bytes to your API servers, which write to storage | Full control; inspect bytes inline | API fleet carries every byte; slow clients pin connections |
| Presigned direct upload | API issues a short-lived signed URL; client PUTs straight to object storage | API carries no bytes; storage scales the ingest | Policy enforced only through what you sign; harder to inspect inline |
| Multipart / chunked resumable upload | Object split into parts (S3: 5 MiB–5 GiB, up to 10,000 parts); each part retried independently | Survives drops; parallel parts | Session state; orphaned parts cost money until aborted |
| Event-driven processing | Storage emits "object created" → queue → workers generate derivatives | Decoupled, retryable | Eventual: the photo exists before its thumbnails do |
| On-the-fly resizing | CDN miss → resize service renders the requested size from the original, caches it | No wasted derivatives; new sizes free | First-view latency; resize fleet hit by cache misses |
| Content hashing | SHA-256 of bytes identifies identical files | Dedup, integrity checks, immutable URLs | Global dedup leaks "this file exists" |
| Replication vs erasure coding | 3 full copies vs k data + m parity fragments | Replication: fast reads. EC: ~1.2–1.5× overhead | EC: reconstruction cost on failure, slower small reads |
| Signed delivery URLs | CDN validates an expiring signature before serving | Private photos stay private-ish | Leaked URL works until expiry; caching per signature |
For most production systems: Presigned, resumable chunked uploads into a staging area keyed by a client-generated upload ID; an event-driven pipeline that validates, strips location metadata and renders a small fixed set of derivatives; a metadata row that flips to READY only after the thumbnail exists; immutable, content-addressed derivative URLs behind a CDN; age-based tiering of originals from replicated to erasure-coded storage; and a deletion pipeline with a recycle-bin window, CDN purge and per-object or per-user keys so backups can be shredded. The primitives are not the interview — the visibility contract, the cost curve and an honest delete are.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What the Interviewer Is Scoring#
Photo upload is not a file-transfer question. Everyone can draw a presigned URL and a bucket.
It is a lifecycle and cost question that tests:
- Whether an upload is an idempotent, resumable session rather than a single request that either works or doesn't
- Whether you define exactly when a photo becomes visible — and never show a reference to bytes or derivatives that don't exist yet
- Whether you treat every uploaded file as hostile input to a decoder running on your servers
- Whether you price storage over years, not days — because bytes are forever and the bill compounds
- Whether "delete" means the bytes are gone everywhere, including caches, derivatives and backups
The key insight: A photo is written once, read thousands of times in its first week and almost never after a year — and kept for a decade. The design follows that curve: make the write path resumable and the first week fast, then move bytes down tiers as reads fall off, and make deletion a first-class pipeline instead of an afterthought. Senior candidates design the upload; Staff candidates design the photo's whole life.
One Question, Three Levels#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Presigned URL → S3 → Lambda thumbnail → CDN | Asks "Social sharing or personal backup? What's the size distribution, and how long do we keep originals?" | Asks "What's the storage growth curve over five years, and who owns the deletion promise to regulators?" |
| Upload | "Client PUTs to a presigned URL" | "Resumable chunked sessions with a client-generated upload ID; create is idempotent; abandoned sessions expire in 7 days" | Sets one upload SDK and session service for every product that ingests media |
| Visibility | Writes the photo row, then processes async | "The row is PENDING until the thumbnail derivative exists; uploader sees a local preview; followers see it only at READY" | Defines the media visibility contract the feed, search and messaging teams all build against |
| Processing | "Lambda resizes the image" | "Sandboxed decoders with pixel and size limits, EXIF location stripped, fixed derivative set, idempotent outputs keyed by content hash" | Treats media decoding as a security boundary owned jointly with security; budgets decoder CVE response |
| Storage | "S3 Standard" | "Originals replicated while hot, erasure-coded after ~90 days; thumbnails stay hot; lifecycle rules, not cron jobs" | Prices tiers per PB-month over five years; decides build-vs-rent at the exabyte inflection |
| Deletion | "Delete the S3 object" | "Tombstone immediately, CDN purge, 30-day recycle bin, then delete original and derivatives; per-user keys so backups are shredded" | Owns the erasure SLA as a compliance commitment with audit evidence |
Why "visibility" separates levels
L5: "The client uploads to S3, then calls our API to create the photo record, and a worker generates thumbnails." The feed reads the record immediately and renders a grey box for two seconds — or forever, if the worker crashed. A follower's app caches the broken state. Nobody defined what it means for a photo to "exist".
L6: "A photo has a state: PENDING while bytes and derivatives are in flight, READY once the thumbnail and display sizes are written, FAILED or QUARANTINED otherwise. Only READY photos are returned to anyone but the owner. The uploader's app shows a local preview from the bytes it already has, so the owner never waits on the pipeline. The flip to READY is a conditional update — PENDING to READY — so a duplicate processing event can't resurrect a deleted photo."
L7: "Visibility is a contract every downstream team depends on — feed, search, messaging, notifications. I'd publish it: a media ID is referenceable only at READY, and every consumer gets the same event. Otherwise each team invents its own check and half of them get it wrong."
Why "storage" separates levels
L5: "Store everything in S3 Standard with three-way replication; it's durable." At 125 TB of new originals a day, that's ~45 PB a year, and the bill grows linearly forever because nothing ever leaves the hot tier.
L6: "Reads fall off a cliff after the first weeks. Thumbnails and display sizes are small — keep them hot. Originals move to erasure-coded warm storage after ~90 days, cutting effective replication from ~3× to ~1.4×; very old originals in a personal-library product can go colder, with a restore path. The lifecycle policy is configuration, not a cron job, and every tier has a measured read latency so product knows what 'cold' feels like."
L7: "At this growth, storage is the single largest infrastructure line in five years. I'd model cost per PB-month by tier against the access curve, and decide whether renting object storage still makes sense past tens of exabytes — that's a one-way door that takes three years to walk through."
Why "deletion" separates levels
L5: "When the user deletes, we delete the object in S3." The derivatives remain under other keys. The CDN keeps serving the cached thumbnail for its one-year TTL. The backup snapshot keeps the original for 90 days. A deduplicated blob shared with another user is either wrongly deleted or never deleted.
L6: "Delete is a pipeline. Step one, a tombstone: the photo disappears from every API in under a second. Step two, purge the CDN paths. Step three, after a 30-day recycle bin, delete the original and every derivative, decrementing refcounts on deduplicated blobs. For backups, each user's objects are encrypted under a per-user key; erasing the account destroys the key, which shreds every backup copy without rewriting them."
L7: "The erasure promise is a regulatory commitment. I'd want an end-to-end SLA — say, gone from all live systems in 30 days, all backups in 60 — with automated evidence per request, because the auditor will ask us to prove it for a specific user."
Positions to Commit To#
| Position | Rationale |
|---|---|
| Presigned, resumable chunked uploads; never proxy bytes through the API fleet | Storage scales ingest; the API stays a control plane |
| Client-generated upload ID makes create, chunk and complete idempotent | Mobile clients retry everything; retries must not create duplicates |
| A photo is visible to others only at READY, after its thumbnail exists | No grey boxes, no references to missing bytes |
| Decode in a sandbox with pixel, size and time limits; strip location metadata by default | Every upload is untrusted input to a complex parser; GPS coordinates are personal data |
| Small fixed derivative set, rendered eagerly; long tail of sizes on demand | Eager covers 95%+ of views; on-demand avoids storing sizes nobody requests |
| Immutable, content-addressed derivative URLs with year-long CDN TTLs | Cache hit rates of 95%+; invalidation becomes a non-problem except for deletion |
| Tier by age: hot replicated → warm erasure-coded → cold for library originals | Storage cost tracks the read curve instead of growing at hot-tier prices |
| Deletion is a pipeline with an SLA, and per-user keys shred backups | "Delete" must be true, provably, everywhere |
Which Problem Are We Solving?#
Three intents produce three different systems. Name them, then commit.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Social photo sharing (posts, stories, profile photos) | Fast time-to-visible; massive read fan-out in the first days; viral hot objects | Resumable upload, eager small derivatives, CDN-first delivery, age-based tiering | Grey boxes, viral origin overload, slow first view | p95 upload-to-READY < 5s for a 4 MB photo; CDN hit rate > 95% |
| Personal photo library / backup | Full-resolution originals forever; bulk background sync; per-user dedup; rare reads of old photos | Background chunked sync, per-user content-hash dedup, aggressive cold tiering, restore path | Silent data loss; library bloat from re-uploads; "where did my 2014 photos go" | Zero lost originals; 11 nines of durability; dedup of re-uploads per user |
| Marketplace / UGC listing images | Moderation before visibility; uniform aspect ratios; legal takedowns | Upload → moderation queue → publish; strict derivative specs; fast takedown | Unmoderated content visible; takedown leaves cached copies | Nothing visible before moderation; takedown complete in minutes, including CDN |
🎯 Staff Move: "I'll design social photo sharing — upload from phones, visible in feeds within seconds, heavy reads early, kept for years. I'll borrow two things from the backup world: originals stored at full resolution, and a real tiering policy, because the storage curve is what decides cost. If this were a marketplace, moderation would sit between upload and visibility, which changes the state machine but not the storage."
Where the Design Splits#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Upload Path: Proxy vs Presigned, Single vs Resumable | Control and inline inspection vs scalability and survival on bad networks |
| 2 | Derivatives: Eager vs On-Demand | Pay storage and compute up front for every size, or pay first-view latency and a resize fleet on misses |
| 3 | Dedup: Global vs Per-User vs None | Storage savings vs privacy leaks, deletion complexity and refcount bugs |
| 4 | Tiering: How Fast Bytes Move Down | Cost per PB-month vs read latency and restore cost for the long tail |
| 5 | Deletion and Privacy: How "Gone" Is Gone | Fast, cheap soft deletes vs provable erasure across CDN, derivatives, replicas and backups |
How Real Companies Built It#
Why this section belongs here: Two published papers describe how one of the largest photo services stored hot and warm photos differently, and the most widely used object store documents the upload primitives every design leans on. Naming them shows you know where the real constraints come from.
Facebook Haystack — One Disk Operation per Photo Read#
Facebook's OSDI 2010 paper describes Haystack, built because storing each photo as a file on NFS-mounted appliances spent most disk I/O on filesystem metadata: the paper reports more than 10 disk operations to read a single image in directories with thousands of files, and about 3 even after tuning. At the time Facebook stored over 260 billion images (four sizes generated per uploaded photo) — more than 20 PB — with users uploading about one billion new photos (~60 TB) per week. Haystack packs many photos ("needles") into large append-only volume files and keeps the per-photo metadata small enough to hold in memory, aiming for at most one disk operation per read. The paper also explains why a CDN alone was not enough: the long tail of requests for less-popular photos misses the CDN, so the storage tier must serve it efficiently (Haystack paper, OSDI 2010).
Staff insight: The bottleneck at scale was not bandwidth or capacity — it was metadata I/O per object. In an interview, say: "Photos are small, immutable and numerous. I'd store them packed in large objects or rely on an object store that does, and keep the per-photo index in memory or a fast KV store — the long tail misses the CDN, so the origin has to be cheap per read."
Facebook f4 — Warm Storage and Deletion by Key#
Facebook's OSDI 2014 paper describes f4, a warm BLOB store added next to Haystack once it measured that requests and deletes fall sharply with age. Haystack's triple replication with RAID-6 gave an effective replication factor of 3.6; f4 uses Reed-Solomon(10,4) within a datacenter plus XOR coding across geographic regions, bringing the effective replication factor to 2.8 and then 2.1. The paper reports that for most BLOB types about a month was a safe boundary between hot and warm; photos used a three-month threshold and profile photos were never moved. Each BLOB in f4 is encrypted with a per-BLOB key stored in an external database, and deleting the key logically deletes the BLOB — so f4 does not need to reclaim space quickly after deletes (f4 paper, OSDI 2014).
Staff insight: Two interview-grade lessons in one system: tier by measured access curve (three months for photos, not a guess), and make deletion a key operation instead of a byte operation. Say: "Deleting the key is the delete. Compaction can happen whenever it's cheap."
Amazon S3 — Multipart Uploads, Presigned URLs and Abandoned Parts#
Amazon S3's documentation recommends multipart upload once objects reach about 100 MB and specifies parts of 5 MiB to 5 GiB (no minimum for the last part), up to 10,000 parts, and a 48.8 TiB maximum object size (S3 multipart limits). Presigned URLs can be generated with SDKs for up to 7 days, expire early if the signing credentials expire, can be used multiple times until expiry, and an upload to a presigned key replaces any existing object at that key (S3 presigned URLs). S3 recommends a lifecycle rule with AbortIncompleteMultipartUpload so uploads that never complete have their parts deleted after a set number of days (S3 lifecycle for incomplete uploads).
Staff insight: Three operational facts hide in those docs. A presigned URL is a reusable bearer token for its lifetime, so sign minutes, not days, and never sign a key a client chose. A PUT overwrites, so upload to a server-chosen staging key and promote it. And abandoned parts are billed storage until a lifecycle rule aborts them — say so before the interviewer asks.
Follow-Ups to Expect#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "The client uploads with a presigned URL" | "The connection drops at 80% of a 50 MB upload on a phone. Then what?" | Resumable sessions, idempotency |
| "A worker makes thumbnails after upload" | "What does a follower see in the two seconds before the thumbnail exists?" | Visibility contract, state machine |
| "We resize with an image library" | "Someone uploads a 50,000 × 50,000 PNG that's 200 KB on disk." | Decompression bombs, sandboxing |
| "We dedupe by SHA-256" | "Can a user tell whether another user already uploaded a given file?" | Dedup privacy leak |
| "Everything is behind a CDN with long TTLs" | "The user deletes a photo. How long does the CDN keep serving it?" | Purge, signed URLs, cache keys |
| "We store originals in object storage" | "What does year five cost, and what would you change?" | Tiering, cost modeling |
| "We delete the object" | "What about the backups, the derivatives and the blob another user shares?" | Deletion pipeline, crypto-shredding |
System Architecture Overview#
Reading the diagram: The API is a control plane: it creates idempotent upload sessions and signs part URLs, but never carries bytes. Clients upload parts straight to a staging bucket. Completion emits an event; sandboxed processors validate the file, strip location metadata, write the canonical original and a fixed set of derivatives into the hot tier, then flip the metadata row to READY. Viewers fetch immutable derivative URLs through the CDN; rare sizes are rendered on demand. Lifecycle rules move originals to erasure-coded warm storage after ~90 days. The deletion pipeline tombstones, purges and finally destroys. The metric that tells you uploads are healthy is
media.pending_age_p99— how long photos sit between "bytes arrived" and "visible".
One-Minute Recap#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Upload | "Presigned PUT to S3" | "Resumable chunked session with a client upload ID; create, part and complete are all idempotent." |
| Visibility | "Create the row, process async" | "PENDING until the thumbnail exists; READY is a conditional flip; owner sees a local preview." |
| Processing | "Lambda resizes it" | "Sandboxed decode with pixel and time limits, strip GPS, 5 fixed sizes, outputs keyed by content hash." |
| Delivery | "Put a CDN in front" | "Immutable content-addressed URLs, 1-year TTL, signed URLs for private photos, origin shield." |
| Dedup | "SHA-256, store once globally" | "Per-user dedup only; global dedup leaks existence and makes deletion a refcount problem." |
| Storage | "S3 Standard" | "Hot replicated for 90 days, then erasure-coded warm; thumbnails stay hot; library originals can go cold." |
| Deletion | "Delete the object" | "Tombstone in under 1s, CDN purge, 30-day bin, delete everything, per-user key shreds backups." |
Numbers to Bring#
| Metric | Value | Why It Matters |
|---|---|---|
| S3 multipart part size | 5 MiB–5 GiB, up to 10,000 parts; 48.8 TiB max object | Sets chunk size and the largest upload you can accept |
| S3 presigned URL lifetime (SDK) | up to 7 days; reusable until expiry | Sign for minutes; a URL is a bearer token |
| Haystack scale (2010) | 260B images, 20 PB, ~1B uploads/week (~60 TB), 4 sizes per photo | Reference for small-object, read-heavy scale |
| Haystack metadata I/O | >10 disk ops per read on NFS; goal ≤ 1 | Per-object metadata, not bytes, is the bottleneck |
| f4 effective replication | 3.6 → 2.8 → 2.1 | Warm-tier savings from erasure coding |
| f4 photo hot→warm threshold | ~3 months | Tier by measured access curve |
| Typical phone photo | ~2–6 MB (JPEG/HEIC, 12–48 MP) | Most uploads fit in one or a few 8 MB chunks |
| Thumbnail / display sizes | ~10–40 KB thumbnail, ~150–400 KB display | Derivatives add only ~10–20% to original bytes |
| Upload-to-READY target | p95 < 5s for a 4 MB photo | The visibility SLO the pipeline is built around |
| CDN hit rate for photos | typically 90–98%, higher for recent popular media | Sizes origin and resize capacity |
| Mobile chunk size | 4–16 MB typical | Smaller = cheaper retries; larger = fewer requests |
| Erasure coding overhead | ~1.2–1.5× (e.g. 10+4 = 1.4×) vs 3× replication | Halves warm storage cost |
| Abandoned upload sessions | commonly 1–5% of started sessions on mobile | Why abort rules matter |
| Illustrative growth | 50M uploads/day × 2.5 MB ≈ 125 TB/day ≈ 45 PB/year of originals | Storage compounds; it never shrinks on its own |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on the presigned URL flow and the thumbnail Lambda, then run out of time before the interviewer asks the questions that decide the level: "What does a follower see before the thumbnail exists?", "What does storage cost in year five?" and "The user deletes the photo — where does it still exist?" Compress the happy path to ~8 minutes and spend the rest on the session and visibility contract, processing as untrusted input, tiering and deletion.
Phase 1: Requirements & Framing (2–3 minutes)#
State the functional scope in one breath:
"Users upload photos from phones and web, see them in their own grid immediately, and share them to feeds where followers see them within seconds. We store originals at full resolution, serve several display sizes through a CDN, and support delete. Uploads come from mobile networks that drop constantly."
Then the non-functional requirements, which is where the design lives:
"Four constraints drive everything. One: uploads must survive dropped connections without duplicates. Two: nothing references a photo before it's renderable — no grey boxes. Three: every upload is untrusted input to code running on our servers. Four: storage is forever and compounds, so cost per PB-month by tier is a design input. I'll assume 50 million uploads a day, averaging 2.5 MB, peaking around 3× average — about 1,700 uploads a second at peak and roughly 125 TB of new originals a day — with reads around 100× writes, mostly in the first week."
Then name the underspecified parts:
"I'd confirm: do we keep originals at full resolution? Are photos public, followers-only or private? Is there a moderation step before visibility? What's the deletion promise? I'll assume originals kept, mixed visibility with signed URLs for private photos, async moderation after visibility for social posts, and gone from live systems within 30 days of delete."
🎯 Staff Move: Saying "storage is forever and compounds" in the first two minutes reframes the problem from "move bytes" to "own a photo's whole life". It earns you the right to spend time on tiering and deletion later.
Phase 2: Core Entities & API (1–2 minutes)#
Name the nouns in 30 seconds:
- UploadSession:
upload_id(client-generated UUID),owner_id,declared_size,declared_type,part_size,parts_received[],staging_key(server-chosen),state(OPEN / COMPLETE / ABORTED),expires_at - Photo:
photo_id,owner_id,upload_id(unique),state(PENDING / READY / FAILED / QUARANTINED / DELETED),content_sha256,width,height,taken_at,visibility,blob_ref,derivative_version - Blob:
blob_id,sha256,size,tier,locations,owner_scope(per-user dedup),refcount,key_id - Derivative:
photo_id,spec(e.g.thumb_256,display_1080),blob_id,format(WebP/AVIF/JPEG)
Client-facing API:
POST /v1/uploads { upload_id, size, type, sha256? } → session, part_size, signed part URLs (first 4)
GET /v1/uploads/{upload_id} → parts_received[], state (resume)
POST /v1/uploads/{upload_id}/parts { part_numbers[] } → more signed part URLs (lazy, 15-min expiry)
POST /v1/uploads/{upload_id}/complete { parts[{n, etag}], caption?, visibility } → photo_id, state=PENDING
GET /v1/photos/{photo_id} → state, derivative URLs when READY
DELETE /v1/photos/{photo_id} → state=DELETED (tombstone), restorable 30 days
🎯 Staff Move: "The client generates the upload ID before the first request. Create, part uploads and complete are all idempotent on it, so a retry after a timeout returns the same session and the same photo ID — not a second photo. That's the one decision that makes mobile uploads correct."
Phase 3: High-Level Architecture (≤5 minutes)#
Draw at most eight boxes:
Walk one photo in 90 seconds:
- The app generates
upload_id, computes SHA-256 while reading the file, and calls create. The API writes an OPEN session and returns signed URLs for the first parts, each valid 15 minutes, each bound to a server-chosen staging key. - The app PUTs 8 MB parts in parallel (2–4 at a time on cellular), retrying each with backoff. On reconnect it asks which parts landed and sends only the missing ones.
- Complete: the API asks storage to assemble the parts, verifies size and checksum, creates the Photo row as PENDING in the same transaction that closes the session, and emits a media event via an outbox.
- A processor pulls the event, decodes the image inside a sandbox with pixel and time limits, strips GPS and other sensitive EXIF, normalizes orientation, writes the canonical original and five derivatives under content-addressed keys.
- The processor flips the photo PENDING → READY with a conditional update and publishes
photo.ready. The feed fan-out starts from that event, not from the upload. - Viewers request
https://cdn…/d/{sha256-of-derivative}.webp; the CDN caches it for a year because the URL never changes for different bytes.
🎯 Staff Move: Say out loud: "The feed fans out on
photo.ready, not on upload complete. That single choice is why followers never see a broken image." You've now spent ~8 minutes.
Phase 4: Transition to Depth (1 minute)#
"That's the happy path, and it's the Senior-level design. What makes this hard is that mobile uploads fail constantly, every file is a potential exploit against our decoders, storage grows by tens of petabytes a year, and delete has to be true everywhere. I'd like to go deep on resumable sessions and the visibility contract, the processing pipeline, storage tiering and cost, and deletion and privacy. Where would you like to start?"
If no preference: start with sessions and visibility. It's the part where Senior designs quietly break.
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → commit → quantify → name who pays.
Deep dive 1: Resumable sessions and the visibility contract (7–8 min)
"A session is a server-side record keyed by the client's upload ID. Parts are 8 MB — a 4 MB photo is one part, a 60 MB burst-mode file is eight. Part URLs are signed lazily, a few at a time, for 15 minutes, so a leaked URL is useless quickly. Completion is idempotent: if the session is already COMPLETE, return the existing photo ID. The Photo row starts PENDING. Only READY photos are visible to anyone but the owner; the owner's app renders the local file it already has."
Quantify: "At 1,700 uploads a second, sessions are ~150M rows a day with a 7-day TTL — about a billion live rows, small, in a KV store keyed by upload_id. 1–5% never complete; their parts are deleted by a lifecycle abort rule at 7 days. At 2% abandonment and an average 3 MB of parts, that's ~3 TB a day that would otherwise accumulate forever."
Who pays: "The client SDK team carries the complexity of chunking and resume — that's correct, because it's the only place that knows what's on disk."
Deep dive 2: The processing pipeline as an untrusted-input boundary (6–7 min)
"Image decoders are large C libraries parsing attacker-controlled bytes. Processors run in a sandbox — no network except to the blob store, a read-only filesystem, a seccomp profile — with limits: reject anything over 100 megapixels before decoding pixels, 50 MB input, 10 seconds CPU. We sniff the real format from magic bytes, never trust the extension, and re-encode every output so nothing from the original byte stream is served raw. EXIF GPS and device serials are stripped from derivatives by default; the original keeps them encrypted and owner-only."
Quantify: "A 12 MP decode plus five resizes and WebP encodes is ~300–800 ms of CPU. At 1,700 photos a second peak, that's ~1,000 cores busy, so ~1,500 cores with headroom — autoscaled on queue age, not depth."
Deep dive 3: Storage tiering and cost (5–6 min)
"Derivatives are ~15% of bytes and get almost all reads — they stay hot. Originals are read for edits, downloads and re-renders, mostly in the first weeks. Lifecycle moves originals older than 90 days to erasure-coded warm storage at ~1.4× overhead instead of ~3×. Library-style originals older than a year that haven't been opened can go to a cold class with a restore path that takes minutes to hours, and the UI says so."
Quantify: "125 TB a day is ~45 PB of originals a year. At illustrative prices — $20 per TB-month hot, $10 warm, $2 cold — keeping everything hot costs ~$900K a month by the end of year one and ~$4.5M a month at year five; hot-for-90-days-then-warm brings year five to roughly $2.4M a month, and a cold tier for unopened library originals takes it lower." (Run your own assumptions through the cost estimator.)
Deep dive 4: Delivery (3–4 min)
"Derivative URLs are content-addressed — the hash of the derivative bytes is in the path — so they're immutable and cached for a year. Public photos are plain CDN objects behind an origin shield. Private photos use short-lived signed URLs or signed cookies, scoped to a path prefix. A re-render with a new encoder produces new hashes and new URLs; old URLs age out of caches."
Deep dive 5: Deletion and privacy (4–5 min)
"Delete flips the row to DELETED — every API stops returning it within a second — and enqueues a CDN purge for its derivative paths. After a 30-day recycle bin, a deletion job removes the original and derivatives and decrements refcounts on deduplicated blobs. Backups can't be rewritten per photo, so originals are encrypted under per-user keys from the KMS; account erasure destroys the user's key, which makes every backup copy unreadable. Every deletion request has an SLA timer and an audit record."
Phase 6: Wrap-Up (2–3 minutes)#
"The core idea: a photo is written once, read heavily for a week and kept for a decade, so the design follows that curve. Idempotent resumable sessions for the write; a READY state that gates visibility; a sandboxed pipeline because uploads are hostile input; immutable URLs for a CDN-first read path; tiering that tracks the access curve; and a deletion pipeline that reaches the CDN, the derivatives and the backups."
The evolution closer:
"What I'd build later: on-demand rendering for the long tail of sizes and new formats like AVIF; client-side dedup for the backup product; perceptual-hash near-duplicate detection; regional storage for data residency. What I'd not build early: our own blob store. Renting object storage is right until the bill is large enough to fund a storage team — and that's a multi-year decision."
🎯 Staff Move: End on who owns "gone". "The media platform owns the deletion SLA end to end — tombstone, purge, destroy, shred — with per-request evidence. Feed and search consume
photo.deletedand own removing their own copies within the same SLA."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| Presigned URL monologue | 8 min on SigV4 and bucket policies | One sentence: short-lived, server-chosen key, bound size and type |
| Thumbnail obsession | Lists every size and format | "Five fixed sizes eager, the rest on demand" |
| No visibility state | Photo row created at upload, processed later | PENDING → READY gate, fan-out on photo.ready |
| No hostile-input story | Trusts the file extension and the decoder | Sandbox, pixel limits, re-encode, strip GPS |
| Storage as a constant | "S3 is durable and cheap" | Five-year PB curve, tiers, lifecycle rules |
| Delete as one API call | "Delete the object" | Tombstone, purge, bin, destroy, shred — with an SLA |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Photo storage is the interview where time is the hidden axis. The upload lasts seconds; the processing lasts seconds; the photo lasts ten years, through three storage tiers, two encoder generations, a data-residency law and a deletion request. A Senior design is correct at t=0. A Staff design is correct at t=+5 years — when the storage bill is the largest line item, when a decoder CVE forces reprocessing of a billion images, and when a regulator asks to see proof that one specific user's photos are gone.
It also has a deceptive read pattern. Reads look enormous — 100× writes — but the CDN absorbs 95% of them, and the ones that reach origin are disproportionately the long tail: old photos nobody has cached. The expensive system is not the one serving viral photos; it's the one storing and occasionally serving the 99% of photos that are almost never viewed.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"A user uploaded a photo three years ago, shared it publicly, and today deletes their account. Walk me through every place a copy of that photo exists right now, and when each one goes away."
A candidate who lists the metadata row (tombstoned in under a second), the CDN edges and origin shield (purged within minutes; immutable URLs mean nothing else invalidates them), the hot derivatives and warm-tier original (deleted after the recycle window), deduplicated blobs (refcount decremented, bytes freed only at zero, and only within that user's scope), replicas and erasure-coded fragments (deleted with the object), backups (unreadable once the per-user key is destroyed), downstream copies in feed caches and search indexes (removed on photo.deleted within the same SLA), and copies already downloaded by other people (outside our control, and said so honestly) — that candidate has run a media platform. A candidate who says "we delete it from S3" has built an upload form.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Social photo sharing → time-to-visible and read fan-out
- Constraint: seconds from upload to followers seeing it; huge early reads; viral hot objects
- Strategy: resumable upload, eager small derivatives, READY gate, CDN-first delivery, age-based tiering
- Failure mode: grey boxes, duplicates on retry, origin melt on viral photos, slow first view
- Who pays for imperfection: the uploader (embarrassment), followers (broken feed), the CDN bill
Personal library / backup → durability and cost over decades
- Constraint: every original kept at full resolution; background bulk sync of tens of thousands of photos; rare reads of old photos
- Strategy: client-side hashing, per-user dedup before upload, aggressive cold tiering, explicit restore UX
- Failure mode: silent loss, duplicate libraries after reinstall, surprise restore latency on old photos
- Who pays: users (lost memories are not refundable), finance (cold-storage retrieval fees)
Marketplace / UGC listings → moderation and takedown
- Constraint: nothing visible before moderation; uniform derivative specs; legal takedown in minutes
- Strategy: upload → moderation queue → publish state; strict crops; purge-on-takedown
- Failure mode: unmoderated image visible for minutes; takedown leaves cached copies
- Who pays: trust and safety, legal
2.2 When NOT to Build a Photo Platform#
- Images are incidental to your product (avatars, a few attachments). Use managed object storage with presigned uploads and a hosted image-transformation CDN feature. A pipeline team is overkill below a few million images.
- Your files are large and streamed (video). Use the video streaming design: segmenting, adaptive bitrate, packaging. The upload half is similar; the processing and delivery halves are not.
- You need file sync semantics (folders, edits, conflicts across devices). That's cloud file sync: versioning and conflict resolution dominate.
- You'd build your own blob store at < ~10 PB. Rented object storage wins on durability engineering you won't match with a small team; revisit only when the bill funds a storage organization.
🎯 Staff Insight: "I'd only build the pipeline and the lifecycle. The blob store, the CDN and the image codecs are bought or open-source. The parts that are ours are the visibility contract, the cost policy and the deletion promise."
2.3 What the Interviewer Leaves Underspecified#
Interviewers deliberately omit:
- Whether originals are kept — many products only need a 2048-pixel version; keeping originals at full resolution can double storage
- Visibility model — public, followers-only or private changes whether URLs are signed and whether CDN caching is per user
- Moderation timing — before visibility (marketplace) or after (social), which changes the state machine
- Deletion promise — "deleted" in the UI, or provably erased from backups by a deadline
- Size distribution — a mix of 3 MB phone photos and 80 MB RAW files needs a chunked path either way
- Retention of location data — EXIF GPS in a public photo reveals a home address
Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.
2.4 Precise Terminology#
| Term | What It Means | Why It Matters in the Interview |
|---|---|---|
| Upload session | Server-side record of an in-progress upload, keyed by a client ID | Makes retries and resume idempotent |
| Part / chunk | A contiguous slice of the file uploaded independently | Unit of retry |
| Staging key | Server-chosen temporary location for uploaded bytes | Clients never choose final keys |
| Canonical original | The validated, orientation-normalized original kept long-term | Source for every future derivative |
| Derivative | A rendered size or format of the original | Unit of CDN caching |
| READY | State in which a photo is renderable and referenceable by others | The visibility contract |
| Content-addressed key | Object key derived from a hash of the bytes | Immutable URLs, dedup, integrity |
| Erasure coding | k data + m parity fragments; any k reconstruct the object | ~1.4× overhead instead of 3× |
| Tombstone | A deletion marker that hides an object before bytes are removed | Fast user-visible delete |
| Crypto-shredding | Destroying the key that encrypts data so every copy becomes unreadable | Deletion that reaches backups |
🎯 Staff Insight: If the interviewer says "just store the photos", ask: "For how long, at what resolution, and what does delete promise? Those three answers set 80% of the cost and most of the compliance risk."
3. Where the Design Splits#
Every photo decision has a technical side (chunk sizes, codecs, erasure codes) and an organizational side (what users are promised, who owns erasure, who pays for storage growth). Interviewers grade the second side.
3.1 Fault Line 1: Upload Path — Proxy vs Presigned, Single vs Resumable#
The tension: Proxying bytes through your API gives full control and inline inspection, but every slow mobile client pins an API connection for the full upload. Presigned direct upload offloads ingest to storage but enforces policy only through what you sign. Single-request uploads are simple; resumable sessions survive the networks your users actually have.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Proxy through API servers | Inline validation, one auth path | 1,700 uploads/s × ~10s on cellular ≈ 17K pinned connections; API deploys kill uploads | API on-call (saturation), users (failed uploads during deploys) |
| Presigned single PUT | Trivial; storage scales | A drop at 90% restarts from zero; retry may create duplicates | Users on bad networks (data and time) |
| Presigned resumable chunked session | Survives drops; parallel parts; idempotent on upload ID | Session state; orphaned parts; more client code | Client SDK team (complexity) |
| Resumable through an upload edge service | Can inspect while streaming; protocol control | A byte-carrying fleet to run and scale | Platform team (fleet) |
Staff default: "Presigned, resumable chunked sessions. The client generates the upload ID, part size is 8 MB (4 MB on 2G-class networks), part URLs are signed lazily for 15 minutes each and bound to a server-chosen staging key, content type and part length. Completion validates total size and the client-supplied SHA-256. A lifecycle rule aborts sessions after 7 days."
When to deviate:
- Inline scanning is mandatory before bytes land anywhere (some regulated contexts): a dedicated upload edge fleet that streams to storage, not the general API.
- Tiny images only (avatars < 1 MB): a single presigned PUT is fine; resumability buys little.
🧭 Principal Move: "Every product that ingests media — posts, messaging, support attachments, marketplace — should use one upload SDK and one session service. Otherwise five teams build five resumable protocols, and three of them don't abort orphaned parts."
❌ Common L5 Trap: "The client uploads to a presigned URL for
photos/{user_id}/{filename}." The client now chooses part of the final key, a retry with the same name overwrites a different photo, a URL valid for a day is a reusable bearer token for that path, and a drop at 90% restarts the whole file.
3.2 Fault Line 2: Derivatives — Eager vs On-Demand#
The tension: Rendering every size at upload makes first view instant but stores sizes nobody requests and makes adding a format a billion-image backfill. Rendering on demand stores nothing extra but puts a resize fleet behind every CDN miss — and a viral photo's first minute behind a stampede.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Eager: all sizes and formats at upload | Instant first view; origin serves static bytes | Storage for unused sizes; new formats need backfills | Finance (storage), pipeline team (backfills) |
| On-demand: render on CDN miss | No unused derivatives; new sizes are free | First-view latency 200–800ms; resize fleet sized for misses; stampede on viral photos | Viewers (latency), platform (resize fleet) |
| Hybrid: small fixed set eager, long tail on demand | Covers 95%+ of views eagerly; flexible tail | Two code paths | Platform team (both paths) |
Staff default: "Hybrid. Five eager derivatives — a 256-pixel thumbnail, 640, 1080, 1440 and a blurred placeholder of a few hundred bytes — in WebP with a JPEG fallback. They add ~15% to stored bytes and cover almost every view. Anything else renders on demand from the original, with request coalescing at the origin shield and the resizer, so a viral photo triggers one render per size, not ten thousand."
When to deviate:
- Marketplace with strict specs: eager everything — the set is small and fixed, and moderation reviews the derivatives.
- Archive/library product with rare views: on-demand for almost everything beyond the thumbnail.
🎯 Staff Insight: "The eager set is a product decision measured in bytes. I'd check request logs: if 97% of views hit four sizes, those four are eager and the rest aren't worth storing for a billion photos."
❌ Common L5 Trap: "We'll generate twelve sizes in three formats at upload, so every device gets a perfect fit." Thirty-six derivatives per photo multiply object count by 36 and add a backfill of billions of objects every time product wants a new size or format.
3.3 Fault Line 3: Dedup — Global vs Per-User vs None#
The tension: Content-hash dedup saves storage — memes and forwarded images are uploaded millions of times. But global dedup lets one user learn whether another user has a specific file, couples deletions across users, and makes per-user encryption impossible.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Global dedup by SHA-256 | Largest savings for viral/forwarded content | "Already uploaded" timing or response leaks existence; deletion needs global refcounts; per-user keys impossible | Users (privacy), compliance (erasure proofs) |
| Per-user dedup | Catches re-uploads and reinstalls; deletion stays within one user | Smaller savings on social content | Finance (some duplicate bytes) |
| No dedup | Simplest; deletion is trivial | Library re-uploads double storage | Finance |
| Global dedup of derivatives of public content only | Saves on memes without touching private originals | Two storage scopes to reason about | Platform team |
Staff default: "Per-user dedup on originals. The client sends the SHA-256 at session create; if that user already has a READY photo with the same hash, we skip the upload and create a new photo row pointing at the same blob, refcount + 1. No cross-user dedup on private originals: it would tell an attacker whether a target has a specific file, and it would make per-user crypto-shredding impossible."
When to deviate:
- Public, widely forwarded content (stickers, memes in messaging): global dedup scoped to content already public, where existence isn't a secret.
- Backup product at huge scale: per-user dedup is the large win — reinstalls and re-syncs are the common duplicate.
🧭 Principal Move: "Dedup scope and encryption scope must be the same. If I want to shred a user's data by destroying one key, no blob can be shared outside that user. I'd write that down as a rule before someone 'optimizes' storage with global dedup."
❌ Common L5 Trap: "Hash every upload and store each unique file once across the whole platform — huge savings." Then "skip upload, already have it" responses reveal that someone, somewhere, uploaded a given file, and deleting a photo needs a cross-tenant refcount that's one bug away from deleting another user's memories.
3.4 Fault Line 4: Tiering — How Fast Bytes Move Down#
The tension: Hot replicated storage serves anything fast and costs the most. Warm erasure-coded storage halves the cost but small reads and repairs are slower. Cold archive is an order of magnitude cheaper and takes minutes to hours to read, with retrieval fees. The access curve decides where the lines go.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Everything hot forever | Simplest; uniform latency | Cost grows linearly with total history | Finance |
| Age-based: hot 90 days → warm | Predictable; matches social read curves | Old-but-viral photos read from warm | Viewers of old viral photos (slightly slower origin reads) |
| Access-based: move when reads fall below a threshold | Tracks real popularity | Per-object access tracking at billions of objects | Platform (tracking cost) |
| Cold archive for old originals | 5–10× cheaper per TB | Minutes to hours to restore; retrieval fees; surprised users | Users (restore wait), finance (retrieval spikes) |
Staff default: "Age-based, because it's predictable and the read curve for social photos is steep. Derivatives stay hot — they're ~15% of bytes and nearly all reads. Originals move to erasure-coded warm at 90 days. For the library product, originals unopened for a year go cold; the app shows the thumbnail instantly and labels full-resolution as 'restoring'. Tiering is lifecycle configuration in the storage layer, versioned and reviewed, not a cron job."
When to deviate:
- Profile photos and pinned content: never leave hot; they're read constantly regardless of age. (f4 made the same exception for profile photos.)
- Regulated retention: some content can't go to a class or region that violates residency — tier within the allowed region.
🧭 Principal Move: "Tier thresholds are a cost-versus-experience contract. I'd have product sign off on what 'restoring' feels like and finance sign off on retrieval budgets, then revisit the thresholds yearly from measured access curves — the way f4 picked three months for photos from data."
❌ Common L5 Trap: "Move everything older than 30 days to Glacier to save money." A user scrolls back to last summer and every full-resolution tap takes hours, retrieval fees spike when a "memories" feature resurfaces old photos to millions of users at once, and the savings vanish.
3.5 Fault Line 5: Deletion and Privacy — How "Gone" Is Gone#
The tension: Soft deletes are fast and reversible; real erasure has to chase copies across caches, derivatives, replicas, dedup references and backups that were designed to make data hard to lose. Privacy also starts at upload: location metadata and private-photo URLs leak more than users expect.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Delete the original object only | One call | Derivatives, CDN copies, backups survive | Users (privacy), compliance |
| Tombstone + async purge + destroy | Instant user-visible delete; complete over time | Needs a pipeline with retries and an SLA | Platform team |
| Per-user (or per-object) keys + crypto-shredding | Backups unreadable without rewriting them | Key service becomes a dependency of every read; key loss = data loss | Platform + security (KMS dependency) |
| Unguessable public URLs | Cacheable, cheap | URL shared once is public forever until purge | Users (leaked links) |
| Signed URLs / cookies for private photos | Access expires | Lower cache efficiency; signing service on the read path | Platform (cost) |
Staff default: "Delete is a pipeline with an SLA: tombstone within a second, CDN purge within 15 minutes, a 30-day recycle bin, then destruction of original and derivatives; downstream systems consume photo.deleted and must comply within the same SLA. Originals are encrypted under per-user data keys wrapped by the KMS, so account erasure destroys the key and every backup copy becomes unreadable. GPS is stripped from derivatives at processing time; private photos are served with signed cookies scoped to the owner's grant, valid for an hour."
When to deviate:
- Legal hold: a hold flag blocks destruction (not the tombstone) and is visible to the deletion pipeline and the audit trail.
- Abuse evidence: content removed for policy violations may be preserved under a separate, access-controlled retention regime set by legal.
🧭 Principal Move: "I'd make 'every system that copies media must consume
photo.deleted' a platform rule with an audit. The media store can be perfect and we still fail an erasure request because a search index kept a thumbnail."
❌ Common L5 Trap: "Set the CDN TTL short so deletes propagate." Short TTLs crush cache hit rates for the 99.9% of photos that are never deleted. Long TTLs on immutable URLs plus explicit purge on delete gives you both.
4. When It Breaks#
4.1 The Retry Storm of Duplicates — Uploads Without Idempotency#
t=0: A CDN-fronted API region degrades: create-session p99 goes 120ms → 9s.
t=+10s: Mobile client timeout is 8s. Clients retry create with a NEW server-generated ID.
t=+1min: Each slow request eventually succeeds AND its retry succeeds → 2–3 sessions per photo.
t=+5min: Both sessions complete. 340K photos posted twice, some three times.
t=+20min: Users delete duplicates by hand; feed shows double posts; support queue +12K.
t=+1 day: Dedup cleanup job has to guess which duplicate the user meant to keep.
Detection: upload.sessions_per_unique_sha256{owner} > 1.05; photo.duplicate_posts_total; create-session latency vs client timeout.
Mitigation: server-side dedup of PENDING photos with the same owner and SHA-256 within 10 minutes; keep the earliest.
Prevention: client-generated upload_id, unique index on (owner_id, upload_id); create returns the existing session on conflict; client timeout tuned above server p99.9.
Owner: media platform (session service), client SDK team (ID generation and retry policy).
4.2 The Grey-Box Feed — Fan-Out Before READY#
t=0: A deploy changes the processor's WebP encoder; 6% of photos fail encoding.
t=+0: The old design fans out to feeds on upload complete, not on READY.
t=+2min: Followers' feeds render 6% of new photos as grey placeholders.
t=+10min: Clients cache the broken derivative state for the session.
t=+25min: Rollback. Failed photos reprocessed. Clients still show grey until refresh.
Detection: media.pending_age_p99 > 30s; media.failed_total rate; client-reported image.load_error_rate.
Mitigation: roll back the encoder; reprocess FAILED photos from their stored originals; push a photo.ready update so clients refetch.
Prevention: fan-out only on photo.ready; canary processor deploys on 1% of traffic with automatic rollback on failure-rate delta > 0.5%; derivative validation (decode the output) before READY.
Owner: media platform (pipeline); feed team (consumes photo.ready, not upload events).
4.3 The Decompression Bomb#
t=0: Attacker uploads a 180 KB PNG that declares 60,000 × 60,000 pixels.
t=+1s: Processor allocates ~14 GB for the RGBA buffer. OOM-killed.
t=+1s: Event redelivered (at-least-once). Next processor OOMs. Repeat.
t=+3min: Attacker has uploaded 400 such files. 30% of processor pods crash-looping.
t=+5min: media.pending_age_p99: 3s → 6 min for every user.
Detection: processor.oom_kills; queue.redelivery_count{event} > 3; pending age.
Mitigation: read the header and reject declared dimensions over 100 megapixels before allocating pixels; poison-message handling — after 3 failed attempts, mark QUARANTINED and stop redelivering.
Prevention: limits checked before decode (dimensions, frame count for animated formats, input size, CPU time); per-process memory caps; a per-user upload rate limit; redelivery cap with quarantine as the terminal state.
Owner: media platform (pipeline), security (decoder hardening and CVE response).
4.4 The Orphaned Parts Bill#
Month 0: Resumable uploads launch. No AbortIncompleteMultipartUpload rule on staging.
Month 1: 2.3% of sessions abandoned (app killed, user cancels, network gone).
Month 6: Staging bucket holds 610 TB of parts that never became objects.
They don't appear in object listings; cost reports show "storage" growing.
Month 7: Finance asks why storage grew 9% faster than uploads. Nobody owns staging.
Detection: staging.incomplete_upload_bytes; ratio of staging bytes to completed-object bytes; storage-class cost breakdown.
Mitigation: add a lifecycle rule to abort incomplete uploads after 7 days; one-time sweep of existing parts.
Prevention: the abort rule ships with the bucket definition, reviewed in infrastructure code; a dashboard of staging bytes with an owner.
Owner: media platform (staging bucket is ours, not "the storage team's").
4.5 The Delete That Didn't — Cached Derivative Outlives the Photo#
t=0: User deletes a photo after realizing it shows their house number.
t=+1s: API returns 404 for the photo. User relieved.
t=+1h: A friend opens an old share link; the CDN serves the cached 1080 derivative.
t=+3 days: User reports it. Purge job had failed on a CDN API rate limit and never retried.
Detection: delete.purge_pending_age_p99; a canary that deletes a test photo every 5 minutes and probes edge URLs for 404.
Mitigation: retry purges with backoff and a dead-letter queue that pages after 1 hour; purge by surrogate key/tag (all derivatives of a photo) rather than per URL.
Prevention: deletion SLA tracked per request; purge as a durable workflow step, not a fire-and-forget call; signed URLs for non-public photos so a stale cache entry still requires a valid signature.
Owner: media platform (deletion pipeline).
4.6 The Viral Origin Melt#
A celebrity's photo is shared to 80 million followers within minutes. Eager derivatives are cached fine — but a new client version requests a 2160-pixel size that isn't in the eager set. Thousands of edge POPs miss simultaneously; without coalescing, the resizer receives 40,000 identical render requests in a minute and falls over, taking every on-demand size down with it.
Detection: resizer.inflight_per_key max; cdn.origin_requests{photo} top-N; resizer error rate.
Mitigation: origin shield in front of the resizer, request coalescing per (photo, spec), and a degraded fallback that serves the nearest eager size.
Prevention: new client sizes added to the eager set (or pre-warmed) before the client ships; load test a single hot key at 50K requests/minute.
Owner: media platform; client teams own announcing new sizes.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Duplicate uploads on retry | upload.sessions_per_unique_sha256 > 1.05 | Affected users' posts | Client upload IDs; idempotent create | Media platform + client SDK |
| Fan-out before READY | image.load_error_rate, pending age | All viewers of new photos | Fan-out on photo.ready only | Media platform + feed |
| Decompression bomb | processor.oom_kills, redelivery count | Pipeline latency for everyone | Pre-decode limits, quarantine | Media platform + security |
| Orphaned parts | staging.incomplete_upload_bytes | Storage bill | Abort lifecycle rule | Media platform |
| Purge failure | delete.purge_pending_age_p99, delete canary | Deleted photo still visible | Durable purge workflow | Media platform |
| Viral resize stampede | resizer.inflight_per_key | All on-demand sizes | Coalescing, shield, fallback size | Media platform |
| Warm-tier read latency spike | blob.read_p99{tier=warm} | Old-photo views, edits | Promote hot objects, cache at shield | Storage |
| Metadata shard hot spot | meta.shard_qps skew | One shard's owners | Shard by owner hash; cache READY rows | Media platform |
| Processing backlog | queue.oldest_age > 60s | Time-to-visible for everyone | Autoscale on age; shed backfills first | Media platform |
🎯 Staff Insight: The pipeline's health metric is time-to-READY for new uploads, not queue depth. "Backfills and reprocessing can put millions of events on the queue legitimately. I'd run new uploads in their own priority lane and page on
media.pending_age_p99for that lane only."
5. Scorecard#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | "Upload a file, make thumbnails, serve via CDN" | Frames the photo's lifecycle — write once, read early, keep for years, delete honestly; commits to an intent | Frames storage growth and erasure as multi-year business commitments with owners |
| Upload | Presigned single PUT | Resumable chunked sessions; client upload ID; lazy short-lived part URLs; abort rule | One upload SDK and session service for every media-ingesting product |
| Visibility & processing | Row at upload; async resize | PENDING → READY gate; fan-out on ready; sandboxed decode with limits; GPS stripped | Visibility contract published to all consumer teams; decoder security owned jointly with security |
| Storage & cost | One hot tier | Derivatives hot, originals erasure-coded after ~90 days, cold with restore UX; costed per PB-month | Five-year cost model; build-vs-rent inflection; tier thresholds reviewed from access data |
| Deletion & privacy | Delete the object | Tombstone, purge, bin, destroy; per-user keys shred backups; per-user dedup only | Erasure SLA with per-request evidence; platform rule that every copier consumes photo.deleted |
| Operations | Queue depth alerts | Time-to-READY SLO on a priority lane; delete canary; orphan-bytes dashboard | Storage cost per active user as a product KPI; game days for decoder CVEs |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Idempotent sessions | "The client generates the upload ID, so a retry returns the same photo, not a second one." |
| Visibility contract | "Nothing outside the owner sees a photo until it's READY; fan-out happens on photo.ready." |
| Hostile-input awareness | "Check declared dimensions before allocating pixels; three failures and it's quarantined." |
| Cost over time | "45 PB a year of originals; tiering at 90 days roughly halves year-five spend." |
| Honest deletion | "Tombstone, purge, bin, destroy — and per-user keys for the backups." |
| Dedup privacy | "Global dedup tells an attacker whether you have a file. Per-user only." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| Bytes proxied through the API fleet | Ties ingest capacity and deploys to the slowest mobile clients |
| No resume or idempotency | Duplicates and restarts on every flaky connection |
| Photo visible before derivatives exist | Grey boxes and broken caches downstream |
| Trusting file extensions and decoders | Decompression bombs and decoder exploits |
| "Storage is cheap" | Ignores a compounding multi-petabyte bill |
| Delete = one object delete | Derivatives, CDN copies and backups survive |
5.4 Common False Positives#
- Deep S3 API knowledge ≠ upload design. Knowing every SigV4 header doesn't answer what happens on retry or when the photo becomes visible.
- Many derivative formats ≠ good processing. Twelve sizes in three codecs is a storage multiplier, not a feature.
- "We use a CDN" ≠ a delivery design. Without immutable URLs and a purge path, the CDN is either stale or useless.
- Encryption at rest ≠ deletion. Server-side encryption with one bucket key doesn't let you shred one user's backups.
6. The 45 Minutes, Phase by Phase#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Social sharing intent; lifecycle framing; uploads/day, size, read ratio, growth |
| Entities & API | 3–5 min | Session, photo, blob, derivative; client upload ID; resume endpoint |
| Architecture | 5–10 min | ≤ 8 boxes; control plane vs byte path; pipeline; CDN |
| Sessions & visibility | 10–17 min | Resumable chunks, idempotency, PENDING → READY, fan-out on ready |
| Processing pipeline | 17–23 min | Sandbox, limits, EXIF, eager vs on-demand derivatives, priority lanes |
| Storage & cost | 23–30 min | Tiers, erasure coding, lifecycle, five-year cost |
| Deletion & privacy | 30–37 min | Tombstone, purge, bin, destroy, per-user keys, dedup scope |
| Pivot / wrap | 37–45 min | Viral photos, multi-region, build vs rent; close on who owns "gone" |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "Upload drops at 80%" | Resume and idempotency | Session with parts received; client upload ID; lazy re-signing |
| "A follower sees the post before the image" | Visibility contract | READY gate; fan-out on photo.ready; owner-only local preview |
| "Someone uploads a 200 KB file that's 3.6 billion pixels" | Untrusted input | Pre-decode limits, sandbox, quarantine after 3 attempts |
| "What does year five cost?" | Cost modeling | PB/year growth × tier prices; derivatives hot, originals warm |
| "User deletes the account" | Erasure completeness | Every copy enumerated; per-user key shredding for backups |
| "Make it multi-region" | Residency and latency | Upload to nearest region; originals stored in home region; derivatives cached globally |
6.3 What to Deliberately Skip#
- SigV4 internals — "short-lived, bound to key, length and type."
- Codec details — "WebP with JPEG fallback; AVIF later" is enough.
- Image ML (tagging, faces) — consumers of
photo.ready, out of scope. - Feed ranking — the feed consumes events; it's a different interview.
- Building a blob store — say when you would, in one sentence.
6.4 Follow-Up Questions to Expect#
- "The connection drops at 80% of a 50 MB upload. Walk me through the next minute."
- "When exactly can a follower see the photo, and what do they see before that?"
- "How do you stop a malicious image from taking down the processing fleet?"
- "Would you dedupe across users? Why or why not?"
- "What does storage cost in year five, and which three changes cut it most?"
- "The user deletes a photo. List every place a copy exists and when it goes away."
- "A celebrity posts and a new client requests a size you don't pre-render. What happens?"
7. Practice Rounds#
Drill 1: The Opening#
Prompt: "Design photo upload and storage for a social app."
Staff Answer
"Is this social sharing, a personal backup library, or marketplace listings? They differ in what gates visibility and how long we keep originals. I'll assume social sharing with originals kept at full resolution: 50 million uploads a day at ~2.5 MB, peaking near 1,700 a second, reads around 100× writes mostly in the first week — about 45 PB of new originals a year.
Four constraints shape it. Uploads come from phones that drop connections, so they're resumable and idempotent. Nothing is visible before it's renderable. Every file is hostile input to our decoders. And storage compounds, so tiering and deletion are part of the design, not afterthoughts. I'll go: entities and the session API → architecture with the API as a control plane → sessions and visibility → processing → storage tiers and cost → deletion and privacy."
Why this is L6:
- Distinguishes three intents and commits to one with numbers
- Frames the problem as a lifecycle, including storage growth and deletion
- Previews an outline that ends with cost and erasure, not with the CDN
What L7 adds:
- Asks whether other products (messaging, marketplace) ingest media separately — the consolidation question
- Frames storage cost per active user and the erasure SLA as the metrics leadership will hold the platform to
❌ Common L5 Trap
"The client gets a presigned URL, uploads to S3, an S3 event triggers a Lambda that makes thumbnails, and CloudFront serves them."
Why this misses: Each piece is fine; together they ignore retries and duplicates, never define when the photo is visible, trust the decoder, keep everything hot forever and say nothing about deletion. The next three follow-ups have no answer.
Drill 2: The Dropped Upload#
Prompt: "A user uploads a 60 MB burst-mode file on a train. The connection drops at 80%. What happens?"
Staff Answer
"The upload is a session keyed by an upload ID the client generated before its first request, with 8 MB parts — this file is eight parts. Parts 1–6 were acknowledged with ETags; part 7 was mid-flight. When the network returns, the client calls GET /uploads/{id}, learns parts 1–6 exist, asks for fresh signed URLs for parts 7–8 because the originals expired after 15 minutes, uploads them and calls complete with all eight ETags.
If the client's create or complete call timed out and it retries, the server sees the same upload ID and returns the same session or the same photo ID. Nothing is duplicated. If the user never comes back, the session expires and a lifecycle rule aborts the parts after 7 days."
Why this is L6:
- Resume from server-known parts, not from zero
- Idempotency on a client-generated ID covers timeouts on every call
- Names the cleanup path for abandoned sessions
What L7 adds:
- Tracks session completion rate by network type and app version as a product metric
- Ships resume as part of one shared upload SDK so every product gets it
❌ Common L5 Trap
"The client retries the upload with exponential backoff."
Why this misses: Retrying a single 60 MB PUT restarts from zero on every drop, and on a train it may never finish. Retrying the create call without an idempotency key creates a second photo when the first request actually succeeded.
Drill 3: Make It Concrete — Size Ingest, Processing and Storage#
Prompt: "Size the upload path, the processing fleet and year-one storage."
Staff Answer
"Ingest: 50M uploads a day is ~580 a second average, ~1,700 at a 3× peak. At 2.5 MB average, peak ingest is ~4.3 GB/s — about 35 Gbps straight into object storage. The API only handles session calls: ~3–5 requests per upload, so ~8K requests a second at peak, a small fleet.
Processing: decode plus five resizes and encodes is ~0.5 CPU-seconds per photo, so 1,700 a second needs ~850 cores busy; with 50% headroom and a backfill lane, ~1,500 cores.
Storage: 125 TB of originals a day, plus ~15% for derivatives, ≈ 144 TB a day, ~52 PB in year one before replication. Hot replicated at ~3× for 90 days is ~39 PB raw; the rest erasure-coded at 1.4×. Metadata: 18 billion photo rows a year at ~500 bytes ≈ 9 TB — sharded by owner." (The back-of-envelope calculator is a quick way to check the arithmetic.)
Why this is L6:
- Separates the byte path from the control plane in the sizing
- Sizes processing in CPU-seconds and leaves room for backfills
- Includes replication and tiering in storage, plus metadata growth
What L7 adds:
- Turns year-one storage into a five-year cost curve and names the tier thresholds as the biggest lever
- Notes that egress and CDN cost may exceed storage early on, and models both
❌ Common L5 Trap
"50 million uploads a day at 2.5 MB is 125 TB a day, so we need about 125 TB of storage a day."
Why this misses: Ignores derivatives, replication overhead and the fact that storage accumulates — year five is 5× year one — and doesn't size processing or say what the API fleet carries.
Drill 4: When Is the Photo Visible?#
Prompt: "A user posts a photo. When can their followers see it, and what do they see before that?"
Staff Answer
"Followers see nothing until the photo is READY. The Photo row is created PENDING at upload complete. A processor writes the canonical original and the eager derivatives, validates them by decoding each output, then flips PENDING → READY with a conditional update and publishes photo.ready. The feed fans out on that event, so a follower's first fetch always has working derivative URLs.
The owner doesn't wait: their app renders the local file immediately with a 'posting…' badge. Target p95 upload-to-READY is under 5 seconds for a 4 MB photo. If processing fails three times, the photo goes FAILED or QUARANTINED and the owner sees a retry prompt — nobody else ever saw it."
Why this is L6:
- Defines visibility as a state with a precise transition
- Fan-out keyed to READY eliminates grey boxes structurally
- Separates the owner's experience from everyone else's
What L7 adds:
- Publishes the READY contract to every consuming team — search, messaging, notifications
- Tracks time-to-READY as the media platform's headline SLO
❌ Common L5 Trap
"As soon as the upload finishes we create the post; thumbnails are generated in the background within a second or two."
Why this misses: "A second or two" is the median. At p99, or during a processor incident, followers see broken images, and clients cache that state. Without a READY gate, the failure mode is user-visible by design.
Drill 5: The Hostile File#
Prompt: "How do you keep a malicious image from compromising or crashing your processing fleet?"
Staff Answer
"Treat decoders as an attack surface. Before decoding pixels: sniff the real format from magic bytes and allow only a short list (JPEG, PNG, HEIC, WebP, GIF), cap input at 50 MB, read declared dimensions from the header and reject anything over 100 megapixels or 500 animation frames. Decode inside a sandbox — separate process, seccomp, no network except the blob store, read-only filesystem, a memory cap and 10 seconds of CPU.
Re-encode every output, so no original bytes are served. After three failed attempts, the event stops redelivering and the photo is QUARANTINED for review. Decoder libraries are pinned and patched on a CVE SLA, and a reprocessing job can re-render affected derivatives from originals if an encoder bug ever shipped bad output."
Why this is L6:
- Limits checked before allocation, not after a crash
- Sandboxing plus re-encoding contains both exploits and polyglot files
- Poison-message handling keeps one file from stalling the pipeline
What L7 adds:
- Makes decoder CVE response a joint security and media-platform runbook with a patch SLA
- Budgets a yearly game day: "a decoder zero-day lands — how fast do we patch and reprocess?"
❌ Common L5 Trap
"We validate the file extension and MIME type before processing."
Why this misses: Both are attacker-controlled. A 200 KB PNG that declares 3.6 billion pixels passes both checks and OOMs every processor it reaches, and at-least-once redelivery spreads it across the fleet.
Drill 6: Dedup and Privacy#
Prompt: "Should we deduplicate identical photos across all users to save storage?"
Staff Answer
"Not for private originals. Global dedup creates an existence oracle: if uploading a file is instant or skipped because 'we already have it', anyone can test whether some user has a specific image. It also couples deletion across users through a global refcount, and it's incompatible with per-user encryption keys — which I want so that erasing an account shreds its backups.
I'd dedupe per user: the client sends the SHA-256 at create, and a match against that user's own READY photos skips the upload. That catches reinstalls and re-syncs, which are most duplicates in library products. For content that is already public and widely forwarded — stickers, memes in messaging — a separate global store for public derivatives is fine, because existence isn't a secret there."
Why this is L6:
- Names the existence-oracle leak and the deletion coupling
- Ties dedup scope to encryption scope
- Offers the narrow case where global dedup is safe
What L7 adds:
- Quantifies what per-user dedup forgoes — typically single-digit percent of social storage — against the compliance risk
- Writes "dedup scope equals key scope" into the media platform standard
❌ Common L5 Trap
"Yes — hash every file and store each unique one once. Memes are uploaded millions of times; it's a huge saving."
Why this misses: The saving is real for public content and dangerous for private originals: it leaks who has what, and deleting a photo now requires cross-user refcounting that's one bug away from deleting someone else's memory.
Drill 7: Build vs Buy#
Prompt: "Should we build our own blob storage and image pipeline or use cloud services?"
Staff Answer
"Rent the blob store and the CDN; build the pipeline and the lifecycle. Object storage gives eleven nines of designed durability, erasure coding, lifecycle rules and multipart uploads, and the durability engineering behind that takes a large team years. The parts that are specific to us — the session service, the READY contract, the processing sandbox, tiering policy and deletion — are thin layers I'd own.
A hosted image-transformation service is a reasonable start for on-demand sizes, priced per transformation; at billions of renders a month, our own resizer behind the CDN is usually cheaper. Building our own blob store only makes sense when storage spend is large enough to fund a dedicated storage organization — tens of millions a year — and the access pattern is stable enough to design hardware for, the way the paper-published warm-storage systems did."
Why this is L6:
- Draws the line between undifferentiated infrastructure and product-specific logic
- Gives a cost threshold for the expensive decision
- Separates transformation build/buy from storage build/buy
What L7 adds:
- Frames own-storage as a three-year one-way door with a migration plan and a team to fund
- Negotiates committed-use pricing on storage as a lever before building anything
❌ Common L5 Trap
"We should build our own storage on commodity disks — S3 is expensive at our scale."
Why this misses: It compares list price to hardware cost while ignoring durability engineering, repair, on-call, datacenter capacity and the years it takes — and hasn't shown the scale at which the math flips.
Drill 8: Changing the Derivative Set Without an Outage#
Prompt: "Product wants AVIF derivatives for all photos, including the existing 40 billion. How do you ship it?"
Staff Answer
"Don't backfill 40 billion photos. New uploads get AVIF in the eager set behind a flag, rolled out 1% → 10% → 100% while watching encode time, pipeline latency and client decode errors. For existing photos, AVIF renders on demand when a capable client requests it, cached at the shield and edge — old photos are rarely viewed, so most will never need an AVIF version.
For the head of the distribution — the few percent of old photos still getting real traffic — a low-priority backfill lane renders AVIF, driven by CDN logs, and it never shares capacity with new uploads. Derivative URLs are content-addressed, so AVIF variants are new objects and new URLs; nothing existing is invalidated. Clients negotiate via the Accept header, with WebP and JPEG as fallbacks."
Why this is L6:
- Rejects a full backfill in favor of on-demand plus a popularity-driven head backfill
- Isolates backfill capacity from the time-to-READY lane
- Uses immutable URLs so the rollout has no invalidation step
What L7 adds:
- Prices the change: encode CPU and extra bytes against CDN egress saved, per year
- Sets a policy that format additions follow the same playbook, so each one isn't a project
❌ Common L5 Trap
"Run a batch job that re-encodes every photo into AVIF and swap the URLs."
Why this misses: Forty billion encodes is months of compute for photos nobody views, it competes with new uploads for processing capacity, and swapping URLs invalidates caches across the CDN.
Drill 9: The Cost of Storage#
Prompt: "Storage is now the biggest line on the infra bill. Cut it by 40% without losing photos."
Staff Answer
"First, measure where the bytes are by tier, age and type. Typically: originals are ~85% of bytes, derivatives ~15%, and most originals are older than 90 days but still in the hot tier.
Levers, in order of impact: move originals older than 90 days to erasure-coded warm storage — from ~3× to ~1.4× overhead, roughly halving their cost. For the library product, move originals unopened for a year to a cold class with a restore path. Abort orphaned multipart parts and delete derivatives for sizes no client requests anymore. Consider lossless JPEG recompression — savings on the order of 20% are typical for such tools — but only with byte-exact round-trip verification. Per-user dedup for library re-uploads. Together that's well past 40%, and none of it touches what users see except a labeled restore delay on very old full-resolution files."
Why this is L6:
- Starts from a byte breakdown, not a guess
- Orders levers by impact, with overhead numbers
- Keeps user-visible effects explicit
What L7 adds:
- Models the five-year curve so the 40% isn't eaten by growth in 18 months
- Sets storage cost per monthly active user as a tracked KPI with a target
❌ Common L5 Trap
"Compress the photos more aggressively — drop JPEG quality to 70."
Why this misses: Lossy recompression destroys originals users expect to keep, is irreversible, and saves less than moving the same bytes to a cheaper tier.
Drill 10: Multi-Region and Residency#
Prompt: "We're launching in the EU with a data-residency requirement. How do uploads and storage change?"
Staff Answer
"Each user gets a home region at signup; EU users' home is the EU. Uploads go to staging in the home region — the session API returns region-specific signed URLs. Originals, derivatives, metadata and backups for EU users stay in EU storage; processing runs in the EU. The CDN can still cache EU users' public derivatives globally if counsel agrees that edge caching of public content is acceptable; private photos are served with signed URLs and edge caching restricted to EU POPs.
The hard cases are cross-region sharing — a US follower viewing an EU photo reads through the CDN, not by replicating the original — and users who move, which is a migration job with an audit record. Deletion works the same way everywhere, but per-user keys must live in the home region's key service so shredding happens where the data is."
Why this is L6:
- Home-region placement for bytes, metadata, processing and keys
- Distinguishes caching public derivatives from storing originals
- Covers cross-region viewing and user migration
What L7 adds:
- Defines residency as a user attribute every platform honors, with automated proofs
- Prices the extra region's fixed cost against the market it unlocks
❌ Common L5 Trap
"Replicate all photos to an EU bucket so EU users get low latency."
Why this misses: Replication solves latency, not residency — EU data still lives in the US — and it doubles storage for every photo. Residency is about where originals, metadata, processing, backups and keys live.
8. Incident Walkthroughs#
Deep Dive 1: Peak-Traffic Incident — New Year's Eve Upload Surge#
Context: At midnight in each timezone, uploads jump from 600/s to 9,000/s for about 20 minutes. Tonight, media.pending_age_p99 climbs from 3 seconds to 14 minutes in Asia-Pacific, and users complain their posts "aren't showing up". The on-call escalates to you.
Questions to Surface First:
- Is the delay in uploads (session completion), the queue (event lag) or processing (worker throughput)?
- Are new uploads sharing a lane with backfills or reprocessing jobs?
- Is the processing fleet scaling, and on what signal?
- Is one dependency — metadata writes, blob writes, KMS for key wraps — slower than usual?
Typical L5 Approach: Doubles the processor fleet. The new pods start in 4 minutes and pick up backfill events first, because a format backfill had 30M events queued in the same queue. Pending age keeps rising.
Staff Approach: Reads the pipeline stage by stage: sessions completing normally, queue lag 14 minutes, processors saturated — on a re-encode backfill that started at 22:00. Pauses the backfill, confirms new uploads drain within 6 minutes, then splits lanes so new uploads always get priority.
Principal Approach: Treats midnight in every timezone as a known, scheduled peak: a calendar of predictable surges, pre-scaling 30 minutes ahead, a change freeze on backfills during declared peaks, and a capacity review that sizes processing for the worst midnight of the year, not the average day.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Stage metrics: session completion normal, queue oldest age 14 min, processor CPU 98%. Identify backfill events as 70% of throughput. Pause the backfill. |
| Triage | Autoscaler keyed on queue depth, which was already high from the backfill, so it was at max replicas before the surge started. |
| Quick fix | Separate queue lanes: uploads (priority) and backfill (preemptible); processors drain uploads first. |
| Guardrails | Page on pending age for the uploads lane only; autoscale on oldest-message age in the uploads lane. |
| Post-mortem | Backfills run only through the preemptible lane, with a scheduler that pauses them when uploads-lane age exceeds 10s. |
Metrics to Watch: media.pending_age_p99{lane=uploads}, queue.oldest_age{lane}, processor.cpu_util, backfill.events_per_s
Organizational Follow-up: the team running the backfill and the media platform agree that backfills are preemptible by default and registered in a shared calendar.
Ownership Question: "Who owns processing capacity during a backfill?" Staff answer: The media platform owns it and decides priority. Backfill owners get leftover capacity, never shared capacity.
Key Takeaway: "Never let work users are waiting for share a queue with work nobody is waiting for."
What clears the Staff bar:
- Localizes the delay to a pipeline stage before adding capacity
- Finds the competing workload instead of scaling blindly
- Turns lane separation into a permanent rule
Deep Dive 2: Silent Failure — Originals That Were Never Kept#
Context: A user asks to download the full-resolution original of a photo from eight months ago. The download returns the 1440-pixel display version. Investigation finds that for 11 weeks last year, about 3% of uploads stored only derivatives — the canonical original write failed silently.
Questions to Surface First:
- Which code path writes the original, and what happens when it fails?
- Does READY require the original to exist, or only the derivatives?
- How many photos are affected, and do any copies of the originals exist anywhere — staging, clients, backups?
- Why didn't any metric notice?
Typical L5 Approach: Fixes the bug so new uploads store originals. Tells affected users the originals are lost.
Staff Approach: Finds that the original write was best-effort after derivatives, with errors logged but not failing the job. Fixes the READY condition to require the original, its checksum and every eager derivative. Then recovers: staging parts were already aborted, but the backup product still held client-side copies for 40% of affected users; a re-upload prompt reaches the rest whose phones still have the file.
Principal Approach: Establishes that "stored" must be verified, not assumed: a continuous audit samples photos and checks every expected object exists with the right checksum, and the durability SLO is measured by that audit rather than by the storage provider's design number.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Confirm scope: query photos with no original blob reference or a missing object. 1.2B checked over a day; 41M affected. |
| Triage | A storage-client upgrade changed timeout behavior; the original write failed on large files and the error was swallowed. |
| Quick fix | READY requires original + checksum + eager derivatives. Failed original write = job failure = retry. |
| Guardrails | Continuous existence audit: sample 1M photos/day, verify all expected objects and checksums; alert on any miss. |
| Post-mortem | Why did "READY" not mean "fully stored"? Make the state machine's invariants executable tests. |
Metrics to Watch: audit.missing_objects_total (must be 0), processor.original_write_failures, media.ready_without_original (must be 0)
Organizational Follow-up: proactive notice to affected users with a re-upload flow; a durability report to leadership.
Ownership Question: "Who owns durability — the storage provider or us?" Staff answer: The provider owns durability of what we wrote. We own proving we wrote it. Eleven nines of durability don't help an object that was never stored.
Key Takeaway: "Durability starts with the write you didn't verify. Audit existence, not just availability."
What clears the Staff bar:
- Makes READY an invariant over every expected object
- Uses every possible recovery source before declaring loss
- Adds a continuous existence audit
Deep Dive 3: Large-Customer Onboarding — A Photo-Book Partner Uploading 2 Billion Images#
Context: A printing partner will migrate 2 billion customer photos (average 5 MB, 10 PB) into the platform over 60 days through a bulk API, starting next month. Sales has signed. You're asked to make it work without hurting regular users.
Questions to Surface First:
- 10 PB in 60 days is ~2 GB/s sustained — what does that do to processing, storage tiers and cost?
- Do these photos need eager derivatives, or are most never viewed?
- Do we need full sessions, or can the partner push from their own object storage?
- How do deletion and residency obligations transfer with the data?
Typical L5 Approach: Gives the partner an API key and the normal upload endpoint. The partner's 400 parallel workers saturate the processing fleet within an hour and regular users' time-to-READY goes to minutes.
Staff Approach: Treats it as a bulk import, not uploads: the partner stages files in their own bucket; a server-side copy job pulls them in at a scheduled rate into the backfill lane; only thumbnails are rendered eagerly, other sizes on demand; originals go straight to the warm tier because they are historical; per-user keys and residency tags are applied at import.
Principal Approach: Turns bulk ingestion into a product capability with a price — imported bytes, processing and five-year storage costed into the contract — and a standard import path every future partner uses.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (design review) | Cost: 10 PB × warm price × 5 years versus contract value. Capacity: 2B thumbnails ≈ 400 cores for 60 days in the backfill lane. |
| Triage | Regular upload path would put 2 GB/s through session APIs and the priority lane — rejected. |
| Quick fix | Server-side copy importer, rate-limited, preemptible, idempotent per source object key. |
| Guardrails | Import pauses automatically when uploads-lane pending age exceeds 10s; separate cost tag for the partner. |
| Post-mortem (pre-mortem) | Partner deletes their source bucket mid-import: the importer's manifest shows exactly which objects are done. |
Metrics to Watch: import.objects_per_s, import.bytes_per_s, media.pending_age_p99{lane=uploads}, storage.bytes{tag=partner}
Organizational Follow-up: legal confirms deletion and residency obligations; finance books the five-year storage cost.
Ownership Question: "Who owns the partner's import throughput?" Staff answer: We own a committed import rate on preemptible capacity; the partner owns staging their files and a manifest we can reconcile against.
Key Takeaway: "Bulk import is not a lot of uploads. Give it its own path, its own lane and its own price."
What clears the Staff bar:
- Separates bulk import from interactive uploads
- Prices five years of storage, not 60 days of ingest
- Tunes derivatives and tiers to how imported photos are actually used
Deep Dive 4: Post-Mortem — Home Locations Exposed Through Photo Metadata#
Context: A journalist reports that downloading public photos from the platform reveals GPS coordinates in EXIF data, including for accounts belonging to people at risk. You're leading the post-mortem.
Questions to Surface First:
- Which outputs contain EXIF — derivatives, originals, or both? Which endpoints serve originals?
- Since when, and how many public photos?
- Was stripping ever implemented, and where did it stop working?
- Are cached copies on the CDN also affected?
Typical L5 Approach: Adds EXIF stripping to the resize step for new uploads.
Staff Approach: Finds stripping worked for derivatives, but a "download original" feature added last year served the raw original, EXIF included, to any viewer of a public photo. Disables original download for non-owners immediately, purges cached originals from the CDN, re-renders and re-strips derivatives produced by an affected processor version, and makes originals owner-only.
Principal Approach: Treats location metadata as personal data under the privacy program: a data inventory entry, a review gate for any feature that serves originals, and an automated scanner that downloads samples from public endpoints and fails a release if any output contains GPS tags.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Disable original download for non-owners. Purge the /orig/ CDN path prefix. |
| Triage | Feature shipped 14 months ago; served originals with EXIF for public photos; 210M downloads by non-owners. |
| Quick fix | Originals owner-only; derivatives re-validated for metadata; download for others serves a stripped full-size render. |
| Guardrails | Release-blocking scanner on public endpoints; processor test asserting no GPS, serial or owner-name tags in outputs. |
| Post-mortem | Why could a feature serve originals without privacy review? Originals become a classified data type with access rules. |
Metrics to Watch: scanner.gps_tags_found (must be 0), orig.downloads{requester!=owner} (must be 0), cdn.purge_completion
Organizational Follow-up: privacy and legal assess notification obligations; trust and safety contacts high-risk accounts.
Ownership Question: "Who owns what metadata leaves the platform?" Staff answer: The media platform owns the guarantee in every output it serves; privacy owns the classification of which fields are personal data.
Key Takeaway: "Stripping metadata in one code path is not a guarantee. Scan what you actually serve."
What clears the Staff bar:
- Finds the path that bypassed stripping instead of re-fixing the one that worked
- Purges cached copies, not just the source
- Adds output scanning as a release gate
Deep Dive 5: Multi-Region Expansion — Moving 3 PB of EU Users' Photos Home#
Context: The company commits to EU data residency. 40 million EU users' photos — 3 PB of originals plus derivatives, metadata and backups — currently live in a US region. Deadline: six months.
Questions to Surface First:
- What counts as the user's data — originals, derivatives, metadata, backups, logs, keys?
- Can uploads go to the EU before migration finishes, and how do reads find each photo?
- How do we prove nothing remains in the US afterwards?
- What happens to shared content — an EU user's photo on a US user's feed?
Typical L5 Approach: Copies the EU users' buckets to an EU region and switches the read path. Forgets backups, the US-resident key service and analytics exports that include photo metadata.
Staff Approach: Sets a home region per user. New uploads for EU users go to the EU from day one. A migration copies originals and metadata user by user, verifies checksums, flips the user's location pointer, and only then deletes US copies. Per-user keys are re-wrapped under EU-resident KEKs; US backups become unreadable once the US copies of those keys are destroyed.
Principal Approach: Defines residency as a user attribute every platform must honor — feed caches, search, analytics, ML training sets — and stands up an automated residency audit that plants marker photos and proves they exist only in the home region.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (planning) | Inventory: originals, derivatives, metadata, sessions, backups, per-user keys, analytics exports, logs with photo IDs. |
| Triage | Backups can't be rewritten per user → key-based shredding of US copies after migration. |
| Quick fix | Migrate per user: copy, verify checksum, flip location pointer, delete source; derivatives re-rendered in the EU instead of copied. |
| Guardrails | Marker users with known photos; weekly scan of US stores for their hashes and IDs. |
| Post-mortem (pre-launch) | Rate-limit migration to protect the warm tier's read bandwidth; ~3 PB over 120 days ≈ 290 MB/s sustained. |
Metrics to Watch: migration.users_completed, migration.checksum_mismatch (must be 0), residency.marker_found_outside_region (must be 0)
Organizational Follow-up: legal approves the inventory; analytics removes photo metadata from US exports for EU users.
Ownership Question: "Who proves residency for photos?" Staff answer: The media platform proves it for its stores and keys with marker tests; compliance owns the company-wide attestation.
Key Takeaway: "Residency includes backups and keys. Move the key, and the old copies stop mattering."
What clears the Staff bar:
- Enumerates every data class, including backups and keys
- Migrates per user with verification and pointer flips
- Uses key relocation to neutralize unrewritable backups
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Design resumable, chunked upload sessions that are idempotent on a client-generated ID and clean up after themselves
- Define a visibility contract — PENDING to READY — and explain why fan-out keys off READY
- Treat image processing as an untrusted-input boundary with pre-decode limits, sandboxing, re-encoding and quarantine
- Choose eager vs on-demand derivatives from request data, with coalescing for viral misses
- Explain why dedup scope must match encryption scope, and why global dedup leaks existence
- Model storage cost over five years and design age-based tiering with erasure coding
- Enumerate every copy of a photo and design a deletion pipeline with an SLA, CDN purge and crypto-shredding for backups
- Handle residency by home region for bytes, metadata, processing and keys
The Bar for This Question#
Mid-level (L4): Uploads files through the API to object storage and resizes them in a background job. Works in a demo. No resume, no visibility state, no tiering, delete removes one object.
Senior (L5): Presigned uploads, an event-triggered thumbnail function, a CDN, maybe SHA-256 dedup. The gap: retries duplicate photos, followers see broken images during processing, decoders trust input, storage stays hot forever, global dedup leaks privacy, and delete leaves derivatives, cached copies and backups behind.
Staff+ (L6): Frames the problem as a photo's lifecycle within the first five minutes. Designs idempotent resumable sessions, a READY gate with fan-out on ready, a sandboxed pipeline with pre-decode limits, a hybrid derivative strategy, age-based tiering with real cost numbers, per-user dedup, and a deletion pipeline with an SLA and crypto-shredding. Names who pays — client teams for resume logic, users of very old photos for restore latency, the platform for the deletion guarantee. The interviewer should learn something from the answer.
10. Hot Takes#
10.1 "Upload Complete" Is the Wrong Event to Build On#
| Event | What Consumers Get |
|---|---|
| Upload complete | Bytes that may not decode, no derivatives yet |
| Photo READY | Verified original, working derivatives, stripped metadata |
| Photo deleted | The signal every copier must act on |
The Staff position: Downstream systems subscribe to photo.ready and photo.deleted, never to storage events.
Why this matters in interviews: It's the single sentence that removes an entire class of broken-image and privacy bugs.
10.2 Most Derivatives Are Never Viewed#
| Strategy | Stored Objects per Photo | Waste |
|---|---|---|
| Every size × every format | 20–40 | Most never requested |
| Four sizes + placeholder | 5 | Small |
| Thumbnail eager, rest on demand | 1–2 | Near zero, with first-view cost |
The Staff position: The eager set is decided by request logs, not by a design doc's list of devices.
Why this matters in interviews: It shows you think in bytes × objects × years.
10.3 Global Dedup Is a Privacy Bug Wearing a Cost-Saving Costume#
| Benefit | Hidden Cost |
|---|---|
| Fewer bytes for forwarded content | Existence oracle for any file |
| One copy of each meme | Cross-user refcounts on delete |
| Simpler storage accounting | No per-user crypto-shredding |
The Staff position: Per-user dedup for originals; global dedup only for content that is already public.
Why this matters in interviews: Interviewers float global dedup to see whether you notice the leak.
10.4 The Storage Bill Is Decided by Policy, Not Technology#
| Lever | Typical Effect |
|---|---|
| Tier originals at 90 days | ~50% lower cost for aged originals |
| Abort orphaned parts | Removes silent growth of a few percent |
| Retire unused derivative sizes | 5–15% of derivative bytes |
| Keep originals at all? | Up to 2× total storage |
The Staff position: The choices that matter — keep originals, when to tier, what's eager — are product and finance decisions. Make them explicit.
Why this matters in interviews: Candidates who pick storage technologies but not policies miss where the money goes.
10.5 "Delete" Without Crypto-Shredding Is a UI Feature#
The Staff position: If backups can't be rewritten per user, the only honest erasure is destroying the key that encrypts that user's data. Design the key hierarchy for deletion on day one; retrofitting it means re-encrypting petabytes.
Why this matters in interviews: Asking "what about the backups?" is the follow-up that separates a feature from a guarantee.
11. Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
The Staff engineer builds a correct, cost-aware photo pipeline. The Principal engineer notices that posts, messaging, stories, marketplace and customer support each ingest media independently — five upload protocols, three processing pipelines, two of which don't strip location data, and no single place that can answer "delete everything this user ever uploaded". The L7 problem is one media platform for the company: one upload SDK, one visibility contract, one processing sandbox, one tiering policy and one deletion guarantee, with product teams owning only what their media means.
The Org-Level Fault Line#
One media platform vs per-product media handling.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each product handles its own media | Fast to ship; tailored | N upload protocols, N decoders to patch, N deletion paths; erasure can't be proven | Security (CVE surface), privacy (erasure), finance (no tiering) |
| One media platform, products consume events | One sandbox, one deletion pipeline, one cost policy | Platform must support every product's specs; prioritization fights | Platform team (scope), products (less flexibility) |
| Platform plus product-owned derivative specs | Shared plumbing; products declare sizes and moderation gates | Spec governance; capacity planning across products | Platform (tooling), products (spec reviews) |
🧭 Principal Move: "The platform owns upload sessions, processing, storage tiers, delivery and deletion. Products own derivative specs and moderation gates, declared in config. Direct writes to media buckets from product services are blocked after a two-quarter migration, because a bucket nobody tiers and nobody deletes from is a liability."
Cost Model#
Assumptions: rented object storage at illustrative $20/TB-month hot, $10 warm, $2 cold; CDN egress at illustrative blended $0.01–0.02/GB at volume; fully loaded engineer ~$250K/year.
| Scale | Volume | Infra ($/month) | Headcount | On-call Load | Notes |
|---|---|---|---|---|---|
| Startup | 100K uploads/day, ~100 TB stored | ~$3–8K (storage, CDN, a few processors) | 1 eng part-time | Shared rotation | Hosted transformations are fine |
| Growth | 5M uploads/day, ~5 PB stored | ~$100–200K (storage ~60%, CDN ~30%) | 4–6 eng | Dedicated rotation, 1–3 pages/month | Tiering is the biggest lever |
| Giant | 50M uploads/day, 100+ PB stored, multi-region | ~$2–5M | 20–40 eng (pipeline, storage, delivery, privacy tooling) | Per-region rotations | Build-vs-rent for storage becomes a live question |
The pricing insight: cost is dominated by bytes at rest and bytes served, not compute. At growth scale, moving aged originals to warm storage and keeping the CDN hit rate above 95% are worth more than every processing optimization combined. Model the five-year curve: storage cost in year five is roughly five times year one at constant prices, unless tiering bends it.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Keeping originals at full resolution (or not) | One-way | Discarded originals can never be recovered |
| Key hierarchy and encryption scope (per-user keys) | One-way | Retrofitting means re-encrypting every stored byte |
| Dedup scope (global vs per-user) | One-way-ish | Un-sharing blobs requires copying them and rewriting refs |
| Public URL format for derivatives | One-way-ish | Embedded and shared links break |
| Building your own blob store | One-way | Multi-year migration in both directions |
| Tier thresholds | Two-way | Lifecycle config; transition fees |
| Eager derivative set | Two-way | Add on demand; retire sizes slowly |
| Processing implementation | Two-way | Internal |
The Standard I'd Write#
RFC-MEDIA-001: Media Ingestion, Storage and Deletion Standard
Status: Approved Owners: Media Platform + Privacy + Security
Scope
Every system that accepts, stores or serves user-uploaded images.
MUST
1. Ingest through the Media Platform upload SDK and session service;
direct writes to media buckets are blocked.
2. Treat a media ID as referenceable only after photo.ready; consume
photo.deleted and remove derived copies within the deletion SLA.
3. Decode only inside the platform sandbox with published limits.
4. Strip location and device identifiers from every output served to
anyone other than the owner.
5. Encrypt originals under per-user data keys from the company KMS;
dedup scope must not exceed key scope.
SHOULD
1. Declare derivative specs in config; prefer on-demand for rare sizes.
2. Tag all media with a home region at creation.
3. Register backfills in the preemptible lane only.
Exceptions
Filed with Media Platform; privacy review required for any exception
touching originals or metadata; time-boxed to two quarters.
Success metrics
- Time-to-READY p95: < 5s for photos under 10 MB
- Deletion SLA met: > 99.9% of requests (live systems 30 days, backups 60)
- Products ingesting media outside the platform: 0 by end of year 2
- Storage cost per monthly active user: tracked, trending down
What I'd Tell the VP#
"Photos are the fastest-growing cost we have — about 45 petabytes of new originals a year — and five teams handle them five different ways, which means we can't prove a deletion and we patch image-decoder vulnerabilities in five places. I'm proposing one media platform: about six engineers for a year. It roughly halves storage cost growth through tiering, gives us a single, auditable deletion guarantee, and removes most of our exposure to malicious-file attacks. Product teams keep control of how their media looks; the platform owns how it's stored, processed and deleted. The main risk is migration effort, so I'd start with posts, which is 70% of the bytes."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Prices bytes over years | "Year-five storage is five times year one unless tiering bends the curve." |
| Identifies one-way doors | "Keeping originals and per-user keys are forever decisions; tier thresholds aren't." |
| Redraws ownership | "Platform owns lifecycle and deletion; products own specs and moderation gates." |
| Makes erasure provable | "Every deletion request has a timer and evidence; the auditor gets a report, not an explanation." |
| Knows when not to build | "We rent the blob store until the bill funds a storage organization." |
Staff answers that L7 interviewers find insufficient:
- "We'll build a great photo pipeline for posts" — correct, but ignores four other teams ingesting media.
- "Delete runs through a pipeline with a CDN purge" — good, but no SLA or evidence for an auditor.
- "We'll tier originals to save cost" — no five-year model or ownership of the thresholds.
Appendices
Appendix A: Mechanics in Depth#
A.1 The Upload Session#
create(upload_id, owner, size, type, sha256?):
if s = sessions.get(upload_id): # idempotent
assert s.owner == owner
return s
if sha256 and p = photos.find_ready(owner, sha256): # per-user dedup
return new_photo_ref(owner, p.blob_id, upload_id) # refcount + 1, no bytes
s = sessions.put_if_absent(upload_id, {owner, size, type,
part_size = choose_part_size(size, network_hint), # 4–16 MB
staging_key = "stg/" + random_id(), # server-chosen
state = OPEN, expires_at = now + 7d})
return s with sign_parts(s, first=4, ttl=15m)
complete(upload_id, parts[]):
s = sessions.get(upload_id)
if s.state == COMPLETE: return photos.by_upload(upload_id) # idempotent
storage.complete_multipart(s.staging_key, parts)
verify(size == s.size, sha256 matches if supplied)
tx:
sessions.update(upload_id, state=COMPLETE)
photos.insert(photo_id, owner, upload_id UNIQUE, state=PENDING)
outbox.append(media.uploaded{photo_id, staging_key})
A.2 The Processor#
on media.uploaded(photo_id, staging_key):
p = photos.get(photo_id)
if p.state != PENDING: return # duplicate or deleted
hdr = sniff_header(staging_key) # magic bytes, declared dims
reject_if(hdr.format not in ALLOWED or hdr.pixels > 100e6
or hdr.frames > 500 or size > 50MB) → QUARANTINED
in sandbox(mem=2GB, cpu=10s, net=blob_store_only):
img = decode(staging_key); img = apply_orientation(img)
orig_key = put_encrypted(original_bytes, key=user_dek(p.owner))
for spec in EAGER_SPECS: # thumb_256, 640, 1080, 1440, blur
out = encode(resize(img, spec), strip_metadata=True)
verify_decodes(out)
put("d/" + sha256(out) + ext, out) # content-addressed, idempotent
photos.cas(photo_id, PENDING → READY, derivatives=..., orig=orig_key)
emit photo.ready
on failure: retry ≤ 3 with backoff, then FAILED or QUARANTINED
A.3 Deletion Workflow#
delete(photo_id): state → DELETED, deleted_at = now; emit photo.deleted
enqueue purge(surrogate_key = photo_id) # minutes
after 30 days: if not restored and not legal_hold:
delete derivatives; blobs.decref(orig_blob)
if refcount == 0: delete object (all replicas/fragments)
account_erasure(u): run delete for all photos; after bin window:
kms.schedule_destroy(user_kek(u)) # backups unreadable
write audit record; close SLA timer
Appendix B: Data Model#
CREATE TABLE photos (
photo_id BIGINT PRIMARY KEY,
owner_id BIGINT NOT NULL,
upload_id UUID NOT NULL,
state TEXT NOT NULL, -- PENDING | READY | FAILED | QUARANTINED | DELETED
content_sha256 BYTEA NOT NULL,
blob_id BIGINT,
width INT, height INT, taken_at TIMESTAMPTZ,
visibility TEXT NOT NULL, -- public | followers | private
home_region TEXT NOT NULL,
deleted_at TIMESTAMPTZ,
legal_hold BOOLEAN NOT NULL DEFAULT false,
UNIQUE (owner_id, upload_id)
); -- sharded by owner_id
CREATE TABLE blobs (
blob_id BIGINT PRIMARY KEY,
owner_scope BIGINT NOT NULL, -- per-user dedup: blobs never cross owners
sha256 BYTEA NOT NULL,
size_bytes BIGINT NOT NULL,
tier TEXT NOT NULL, -- hot | warm | cold
key_id TEXT NOT NULL, -- wrapped DEK reference
refcount INT NOT NULL,
UNIQUE (owner_scope, sha256)
);
CREATE TABLE derivatives (
photo_id BIGINT NOT NULL,
spec TEXT NOT NULL, -- thumb_256 | w640 | w1080 | w1440 | blur
format TEXT NOT NULL, -- webp | jpeg | avif
object_key TEXT NOT NULL, -- d/{sha256}.{ext}
PRIMARY KEY (photo_id, spec, format)
);
Appendix C: Coordination Mechanisms#
C.1 Visibility and Deletion Across Consumers#
C.2 Quick Comparison#
| Mechanism | Guarantees | Failure Mode | Use For |
|---|---|---|---|
| Client upload ID | No duplicate sessions or photos | Client reuses IDs across files | Every upload |
| Lazy short-lived part URLs | Leaked URL useless in minutes | TTL shorter than slow part uploads | Resumable sessions |
| PENDING → READY CAS | No references to missing derivatives | Processor writes READY before verifying | Visibility |
| Content-addressed derivative keys | Immutable URLs, idempotent writes | Hash collision (negligible with SHA-256) | Delivery |
| Request coalescing | One render per (photo, spec) | Lock leak blocks renders | On-demand sizes |
| Lifecycle abort rule | Orphaned parts deleted | Rule missing on new buckets | Staging |
| Per-user keys | Backup shredding per user | Key loss = data loss | Originals |
| Surrogate-key purge | All derivatives purged together | Purge API rate limits | Deletion |
Appendix D: API Contract & Client Behavior#
- Generate
upload_idonce per file, persist it with the local upload queue, and reuse it on every retry, including after app restarts. - Compute SHA-256 while reading the file; send it at create so per-user dedup can skip the upload.
- Upload 2–4 parts in parallel on cellular, more on Wi-Fi; retry each part with exponential backoff and jitter; on 403 (expired URL), request fresh part URLs.
- After any interruption, call
GET /uploads/{id}and send only missing parts. - Show the local file immediately; poll or subscribe for READY; never construct derivative URLs client-side.
- Request derivatives with
Accept: image/avif, image/webp, image/jpeg; the server picks the best available. - Private photos: use signed cookies issued per session; never embed long-lived signed URLs in shares.
Appendix E: Observability#
Core metrics:
media.pending_age_p99{lane},media.ready_latency_p95,media.failed_total,media.quarantined_totalupload.session_completion_rate{network, app_version},upload.sessions_per_unique_sha256processor.cpu_util,processor.oom_kills,queue.oldest_age{lane}cdn.hit_rate,cdn.origin_requests,resizer.inflight_per_keystaging.incomplete_upload_bytes,storage.bytes{tier},storage.cost_per_maudelete.purge_pending_age_p99,delete.sla_breach_count,audit.missing_objects_total
Critical alerts:
| Alert | Threshold | Severity |
|---|---|---|
| Uploads-lane pending age p99 | > 30s for 5 min | Page |
| Delete canary still served from edge | > 15 min after delete | Page |
| Missing objects in existence audit | > 0 | Page |
| Processor OOM kills | > 10/min | Page |
| Session completion rate drop | > 5 points vs 7-day baseline | Ticket → page at 10 points |
| Staging incomplete bytes growth | > 5% week over week | Ticket |
Debugging the silent failure: overall success rates hide partial writes. Watch the existence audit (every expected object present with the right checksum), the ratio of READY photos to originals stored, and time-to-READY per lane. A photo that's READY without its original looks perfectly healthy until someone asks for the full-resolution file.
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 1M photos | Presigned PUTs, event-triggered resize, one bucket | Duplicates on retry; no visibility state |
| 1M–1B photos | Resumable sessions, READY gate, sandboxed processors, CDN, warm tier | Storage bill; derivative sprawl; deletion completeness |
| 1B–100B photos | Priority lanes, on-demand long tail, per-user keys, deletion SLA, residency | Cross-product sprawl; metadata scale |
| > 100B photos | Company media platform, custom storage economics, multi-region home placement | Org coordination; build-vs-rent decisions |
What you don't build on day one: on-demand rendering, cold tier with restore UX, residency, perceptual near-duplicate detection, a custom blob store. Each has a trigger in Section 11.
Appendix G: Multi-Tenancy, Fairness & Cost#
- Per-user upload rate limits: e.g. 500 photos/hour interactive, higher for verified backup clients; protects the pipeline from abuse and runaway sync bugs. See the rate limiter.
- Priority lanes: interactive uploads > moderation re-renders > backfills and imports; backfills are preemptible.
- Bulk importers run on a separate path with their own cost tag and rate.
- Cost attribution: storage bytes by tier, CDN egress and processing CPU attributed per product; storage cost per monthly active user reported monthly.
- Abuse: per-user quotas on stored bytes for free tiers; quarantine counts per user feed trust and safety signals.