Hiring BarSupport

Handling Large Blobs — Cross-Cutting Pattern

Pattern34 min read6 diagrams

Technologies that implement this pattern: PostgreSQL · DynamoDB · Kafka · Redis · API Gateways

Why This Matters#

Large blobs are where good API designs go to die. The service that handles 20K requests/sec of 2 KB JSON flawlessly falls over at 50 concurrent 2 GB video uploads — connection pools pinned for 15 minutes, memory buffers exhausted, load balancer idle timeouts firing mid-upload, and a database row that says "ready" pointing at a file that never arrived. Blobs break every assumption the request/response path was built on: that payloads are small, requests are short, and a failed request is cheap to retry.

Most candidates treat blob handling as a storage question: "put it in S3." Staff engineers treat it as a data-plane / control-plane separation problem. The bytes should never touch your application servers; your application should only ever handle metadata and permission. The client talks directly to object storage via a short-lived presigned URL, uploads in resumable chunks, and your service learns about the result through an event — and the hard part of the design is keeping the metadata record and the blob consistent when any of those steps fail.

In interviews, this pattern hides inside Dropbox, YouTube, Instagram, WhatsApp media, Slack attachments, and LeetCode submissions. The L5 answer draws "Client → API → S3." The L6 answer says "the API never sees the bytes" and explains the upload state machine, orphan cleanup, and CDN delivery. The L7 answer asks what the egress bill looks like at 10× and which storage tier the 90% of never-read-again objects should be sitting in.

The 60-Second Version#

  • Bytes bypass your servers. Issue a presigned URL (5–15 minute TTL) and let the client PUT directly to object storage. Proxying a 1 GB upload through an app server pins a connection for ~14 minutes on a 10 Mbps mobile uplink.
  • Anything over ~100 MB must be resumable. Use multipart upload: S3 parts are 5 MiB–5 GiB, up to 10,000 parts. 16–64 MiB parts with 4–8 in parallel is a strong default. A dropped connection costs one part, not the file.
  • Metadata and blob are two systems — design the state machine. PENDING → UPLOADED → PROCESSING → READY | FAILED. The DB row is created before upload; the blob's arrival is confirmed by a storage event or a client complete call that the server verifies with HEAD.
  • Orphans are guaranteed; cleanup is mandatory. Abandoned multipart uploads and PENDING rows accumulate forever. Set a lifecycle rule to abort incomplete multipart uploads after 1–7 days and sweep PENDING rows older than 24h.
  • Deliver through a CDN with immutable, content-addressed keys. Cache-Control: max-age=31536000, immutable on hash-named objects gets 95%+ edge hit rates. Origin egress is the bill that grows fastest.
  • Processing is an async pipeline, and it must be idempotent. Storage events are at-least-once and can arrive late or duplicated. Thumbnailing, transcoding, and virus scanning key on (blob_id, version, step) so a replay is a no-op.

The Problem#

A user uploads a 3 GB video from a phone on a train, or a 400 MB design file over hotel Wi-Fi, or 10 photos at once over 4G. The upload takes minutes and the connection will drop at least once. Meanwhile the product needs to show the file in a list immediately, generate a thumbnail, scan it for malware, transcode it into 5 renditions, and serve it to 10 million viewers if it goes viral — without ever showing a broken link, double-charging storage, or letting someone read a file they shouldn't. The naive "POST the file to our API" design fails on memory, timeouts, cost, and consistency simultaneously.


Case Studies That Use This Pattern#

  • Blob Storage — The storage layer itself: chunking, erasure coding, metadata service design
  • File Sync — Block-level chunking, dedup, and resumable sync across devices
  • CDN & Edge Caching — Delivery of large static objects, range requests, origin shielding
  • Chat Messaging — Media attachments with presigned uploads and expiring download links
  • News Feed — Image/video upload, processing into renditions, and CDN delivery at feed scale
  • Code Execution — Test-case bundles and large artifacts moved out of the request path

The Core Tradeoff#

StrategyWhat WorksWhat BreaksWho Pays
Proxy through API serversSimple auth, one code path, easy validationMemory/connection exhaustion, LB timeouts (60s defaults), 2× bandwidth costApp on-call when uploads saturate the fleet; every other API call shares the pain
Presigned single PUTBytes bypass servers, trivially scalableNo resume — a drop at 95% restarts; 5 GiB single-PUT cap on S3Users on flaky networks; mobile retention
Presigned multipart / resumableResume from last part, parallel parts, multi-TB objectsClient complexity, orphaned parts, more API calls to manageClient team owns the SDK; storage bill pays for abandoned parts without lifecycle rules
Chunked + content-addressed (Dropbox-style)Dedup, delta sync, only changed blocks re-uploadMetadata explosion (1 row per block), GC of unreferenced blocksMetadata/storage platform team; GC bugs can delete live data
Direct-from-origin deliverySimple, always freshEgress cost, origin hot spots, latency for distant usersFinance (egress is often the #1 storage-related line item)
CDN with signed URLs/cookies95%+ offload, low latency, access control at edgeInvalidation lag, signed-URL leakage, cache-key mistakesCDN/platform team; security owns leaked-link response

The L5 → L6 → L7 Contrast#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Upload to our API, store in S3, save the URL in the DB""The API issues a presigned URL; bytes never touch our servers""One upload/delivery platform for every product surface, so 20 teams don't each reinvent presigning and scanning"
ReliabilityRetries the whole upload on failureMultipart with per-part retry and resume tokens; server-side verification on completeSets client SDK standards (part size, concurrency, backoff) and measures upload success rate as a product SLO
ConsistencyWrites the DB row after upload succeedsExplicit state machine with PENDING rows, event-driven confirmation, orphan sweepsDefines the org's blob lifecycle contract: retention, legal hold, deletion SLAs, GDPR erasure across replicas and CDN
DeliveryServes files from S3 URLsCDN with immutable content-addressed keys, signed URLs, range requestsOwns the egress cost curve: CDN commits, multi-CDN, and the build-your-own-edge threshold
CostDoesn't mention itTiers cold objects (IA / archive), aborts incomplete uploadsPrices storage × retention × replication across the org; sets default lifecycle policies and chargeback
SecurityMakes the bucket privateShort-TTL scoped presigned URLs, content-type/size constraints, malware scan before READYTreats user uploads as an org-wide attack surface: one scanning pipeline, one abuse-response runbook
Why "First move" separates levels

"Upload through the API" is a competent answer for 100 KB avatars — it is what most frameworks do by default. It fails the moment objects get large or concurrency gets high, because the API tier's scarce resources (connections, memory, request timeouts) are tuned for small payloads. The Staff move is the separation: the API is the control plane (who may upload what, where, how big), object storage is the data plane. The Principal move recognizes that every product team needs exactly this, and that 20 slightly different presigning implementations mean 20 slightly different security bugs.

Why "Consistency" separates levels

"Write the row after the upload" sounds safe, but it means the product can't show the file until the upload finishes, and if the client dies after upload but before the API call, the blob is orphaned forever. Staff engineers accept that metadata and blob live in two systems with no shared transaction, so they design the state machine and the reconciliation job explicitly. Principal engineers extend the lifecycle to the end: deletion. A GDPR erasure request must reach the primary, the replicas, the CDN cache, the thumbnails, the transcodes, and the backups — that's a policy and a platform, not a function call.


Staff Default Position#

The control plane issues permission; the data plane moves bytes; an event reconciles them.

Create the metadata record first in PENDING state, return a short-lived presigned (multipart for anything > ~100 MB) upload URL scoped to one key, content-type, and size range. The client uploads directly to object storage. Completion is confirmed by the storage event (and optionally the client's complete call, which the server verifies with a HEAD). An idempotent async pipeline scans and processes the blob before it transitions to READY. Delivery goes through a CDN using immutable, content-addressed keys. A lifecycle policy aborts incomplete uploads and tiers cold data; a sweeper expires stale PENDING rows.


When to Deviate#

  • Small, high-volume payloads (< ~1 MB) with inline validation needs. Avatars, small JSON documents, or signatures that must be validated synchronously can go through the API — presigning adds a round trip that costs more than the upload itself.
  • Strict content inspection before storage. Some regulated environments require that bytes are scanned before they're persisted anywhere. Route through a dedicated upload-gateway tier (not the general API fleet), sized and isolated for streaming.
  • Heavy dedup / delta-sync workloads. File-sync products gain 30–60% storage savings from block-level content addressing; accept the metadata complexity when the product is sync.
  • Tiny scale or internal tools. At 50 uploads/day, a proxied upload with a 100 MB cap is fine. Don't build a multipart state machine for an admin panel.

Common Interview Mistakes#

What Candidates SayWhat Interviewers HearWhat Staff Engineers Say
"The client uploads the file to our API, which saves it to S3""My app servers will be a bandwidth bottleneck""The API issues a presigned URL; bytes go straight to object storage. Our servers only handle metadata."
"We'll store the file in the database""I haven't done the math on row size or backup time""Blobs go in object storage; the DB holds metadata and a key. 1 TB of blobs in Postgres turns every backup and replica rebuild into a multi-hour event."
"If the upload fails, the user retries""Users on mobile will never finish a 2 GB upload""Multipart with 16–64 MiB parts; a drop costs one part, and the client resumes from ListParts."
"Once uploaded, we save the URL""Orphans and ghost records are someone else's problem""PENDING row first, event-driven confirmation, and a sweeper for PENDING older than 24h plus a lifecycle rule for incomplete multiparts."
"We'll serve files from S3""I haven't looked at the egress bill""CDN in front with immutable content-hashed keys; origin sees < 5% of reads."
"Presigned URLs are secure""I haven't thought about leakage or scope""5–15 minute TTL, scoped to one key with size and content-type conditions; downloads get short-lived signed CDN URLs."

Quick Reference#

Diagram: Quick Reference

Staff Sentence Templates#

"Our API servers should never see the bytes. The API authorizes the upload and returns a presigned URL scoped to [key, size range, content type] with a [N]-minute TTL; the client uploads directly to object storage."

"Metadata and blob live in two systems with no shared transaction, so I'll make the state machine explicit: the row starts PENDING, moves to UPLOADED on the storage event, and only becomes READY after [scan / transcode]. A sweeper handles anything stuck longer than [N hours]."

"For [N GB] files over mobile networks, I'd use multipart with [16–64] MiB parts and [4–8] parallel streams. A dropped connection costs one part, and the client resumes by listing the parts it already sent."

"Reads go through a CDN with content-addressed, immutable keys, so the edge hit rate stays above [95]% and origin egress is a rounding error. Access control is a short-lived signed URL, not a public bucket."


Implementation Deep Dive#

1. Presigned Uploads — S3 SigV4 with Scoped Conditions#

The API authorizes; object storage receives. A presigned URL is a bearer token for exactly one operation on exactly one key, valid for a few minutes.

Diagram: 1. Presigned Uploads — S3 SigV4 with Scoped Conditions
function create_upload(user, req):
    require req.size <= user.plan.max_object_bytes              # e.g. 5 GiB free, 50 GiB paid
    require quota.reserve(user.id, req.size)                     # reserve, release on expiry
    blob_id = uuid7()                                            # time-ordered, index-friendly
    key = "u/" + shard_prefix(blob_id) + "/" + blob_id          # spread across prefixes
    db.insert(blobs, { id: blob_id, owner: user.id, key, state: "PENDING",
                       declared_size: req.size, declared_sha256: req.sha256,
                       expires_at: now() + 24h })
    url = s3.presign_put(key,
            expires_in = 900,                                    # 15 minutes
            conditions = { content_length_range: [req.size, req.size],
                           content_type: req.content_type,
                           checksum_sha256: req.sha256 })         # storage rejects mismatched bytes
    return { blob_id, upload_url: url }

Why each condition matters: the size range stops a client from uploading 50 GB against a 5 MB reservation; the content-type pin stops HTML-as-image stored-XSS tricks; the checksum makes storage reject corrupted or substituted bytes; the short TTL limits a leaked URL to minutes. S3 sustains ~3,500 writes/sec and ~5,500 reads/sec per prefix, so a hashed prefix scheme (shard_prefix) keeps high-volume uploads from concentrating on one partition.

🎯 Staff Insight: The presigned URL is not the security boundary — the policy you signed into it is. A presigned PUT with no size or type condition is an open write endpoint for 15 minutes. Treat the signing function like an authz service: centralized, reviewed, and with every condition mandatory.

2. Multipart and Resumable Uploads — Surviving the Train Tunnel#

For anything above ~100 MB, a single PUT is a bet that the connection survives the whole transfer. On a 10 Mbps uplink, 1 GB takes ~14 minutes; the probability of some interruption in 14 minutes on mobile is high. Multipart turns one fragile transfer into many cheap retries.

# Server: initiate and presign parts lazily (never presign 10,000 URLs up front)
function start_multipart(blob_id):
    upload_id = s3.create_multipart_upload(key_of(blob_id))
    db.update(blobs, blob_id, { upload_id })
    return { upload_id, part_size: choose_part_size(declared_size) }

function choose_part_size(size):
    # S3 limits: 5 MiB min (except last), 5 GiB max per part, 10,000 parts max
    p = max(16 MiB, ceil(size / 9_000))           # leave headroom under 10k parts
    return min(p, 512 MiB)

function presign_parts(blob_id, part_numbers):     # client asks in batches of ~20
    return [ s3.presign_upload_part(key, upload_id, n, expires_in=900) for n in part_numbers ]

# Client: parallel parts, per-part retry, resume from what the server already has
function upload(file, blob_id):
    done = s3.list_parts(upload_id)                  # resume: skip parts already stored
    parallel(max=6) for n in parts(file) if n not in done:
        retry(max=5, backoff=exponential_jitter(base=500ms, cap=30s)):
            etag = http.put(presigned[n], file.slice(n))
            done[n] = etag
    api.complete(blob_id, upload_id, done)           # server calls CompleteMultipartUpload
File SizePart SizePartsCost of a DropNotes
200 MB16 MiB13≤ 16 MiB re-sentSingle PUT also viable on good networks
5 GB16 MiB~300≤ 16 MiBSweet spot; 6 parallel streams saturate most uplinks
100 GB32 MiB~3,000≤ 32 MiBPresign in batches; track progress server-side
1 TB~112 MiB~9,000≤ 112 MiBNear the 10K-part ceiling — size parts from the declared size

Alternatives worth naming: Google Cloud Storage resumable uploads (a session URI; chunks in multiples of 256 KiB; query the offset after a drop) and the open tus protocol (HTTP-based resumable uploads with Upload-Offset headers) when you need a vendor-neutral client.

🎯 Staff Insight: Parallelism is a double-edged knob. Six parallel parts saturate a home connection and cut wall time 3–5×; thirty parallel parts on a phone starve the rest of the app and trigger carrier throttling. Make concurrency adaptive: start at 4, increase while throughput per stream holds, back off on errors.

3. Metadata Consistency — The Upload State Machine#

There is no transaction spanning your database and object storage. Every failure between the two leaves them disagreeing. The fix is to make disagreement a named state with a reconciler, not an accident.

Diagram: 3. Metadata Consistency — The Upload State Machine
# Every transition is a conditional update — duplicates and races become no-ops
function on_object_created(event):
    blob = db.get_by_key(event.key)
    if blob is None:
        metrics.increment("blob.orphan_object")        # object without a row
        return schedule_delete(event.key, after=24h)   # grace period for slow row writes
    head = s3.head(event.key)
    if head.size != blob.declared_size or head.checksum != blob.declared_sha256:
        return db.cas(blob.id, from="PENDING", to="FAILED", reason="integrity")
    if db.cas(blob.id, from="PENDING", to="UPLOADED"):
        pipeline.enqueue(blob.id, version=head.etag)

# Reconciler — hourly
function sweep():
    for b in db.query("state = 'PENDING' AND expires_at < now()"):
        if s3.head(b.key).exists: on_object_created({ key: b.key })   # event was lost
        else: db.cas(b.id, from="PENDING", to="EXPIRED"); quota.release(b)
    for b in db.query("state = 'PROCESSING' AND updated_at < now() - 2h"):
        pipeline.enqueue(b.id)                                       # stuck job replay

Two invariants to state out loud: (1) a READY row always has a verified object — enforced by only transitioning after HEAD and processing; (2) every object eventually has a row or is deleted — enforced by the orphan check and the lifecycle rule. Since late 2020, S3 provides strong read-after-write consistency, so the HEAD immediately after an event is reliable; the event itself remains at-least-once and can be delayed.

🎯 Staff Insight: Never trust the client's complete call alone. Clients lie, crash, and retry. Use the client signal as a hint to speed things up, and the server-side HEAD plus storage event as the truth. That also means a malicious client can't mark a blob READY without the bytes existing.

4. The Processing Pipeline — Idempotent Fan-Out#

After upload, a blob typically needs 3–8 derived steps: malware scan, EXIF stripping, thumbnails, transcodes, content moderation, text extraction for search. These take seconds to hours, fail independently, and must never publish an unscanned file.

STEPS = ["scan", "strip_metadata", "thumbnail", "transcode_720p",
         "transcode_1080p", "moderation"]
GATING = {"scan", "moderation"}                     # must pass before READY

function process(blob_id, step):
    key = (blob_id, blob.version, step)             # idempotency key
    if results.exists(key): return results.get(key) # replay-safe
    with lease(key, ttl=15min, heartbeat=60s):       # long transcodes heartbeat
        out = run_step(step, blob)                  # writes derived/<blob>/<step>/<hash>
        results.put(key, out)
    if all_gating_passed(blob_id) and all_required_done(blob_id):
        db.cas(blob_id, from="PROCESSING", to="READY")

# Transcodes of large videos: split into ~10 s segments, transcode in parallel,
# stitch — a 2-hour video finishes in minutes instead of hours.
StepTypical LatencyFailure Handling
Malware scan (1 GB)10–60 sGating; fail closed → QUARANTINED
Thumbnail (image)100–500 msRetry 3×; fall back to generic icon
Video transcode (1 h, 1080p)5–30 min (segmented, parallel)Heartbeat + resume by segment
Moderation model0.5–5 sGating for public content; async for private
Text extraction / OCR1–60 sNon-gating; search index updates later

🎯 Staff Insight: Split steps into gating and enriching. Gating steps (scan, moderation) block READY and fail closed. Enriching steps (thumbnails, OCR, extra renditions) never block READY — the file is usable, and the UI shows a placeholder until the derivative lands. Candidates who gate on everything make a slow transcode look like a failed upload.

5. CDN Delivery — Immutable Keys and Signed Access#

Delivery is where cost concentrates: a popular object is stored once and read millions of times.

# Content-addressed derived objects: the URL changes when the content changes
url = cdn_host + "/v/" + sha256(content)[0:32] + "/" + rendition + ".mp4"
headers = { "Cache-Control": "public, max-age=31536000, immutable" }

# Private content: short-lived signed URL or signed cookie at the edge
signed = cdn.sign(url, expires = now() + 10min, ip_prefix = optional)

# Range requests: players fetch 2-10 MB byte ranges; the CDN caches ranges,
# so seeking in a 4 GB video never pulls the whole object from origin.
TechniqueEffectWatch Out For
Content-addressed immutable keysNo invalidation ever; 95%+ hit rateMetadata must map logical file → current hash
Origin shield (one mid-tier cache)Collapses N edge misses into 1 origin fetchShield region becomes a hot spot for viral objects
Signed URLs (per object) / signed cookies (per path)Access control at the edge, no origin auth hitLeaked links valid until TTL; keep TTL minutes, not days
Byte-range cachingEfficient seeking, partial downloadsCache key must include range normalization
Tiered storage behind CDNCold objects in IA/archive tiersArchive tiers have retrieval latency (minutes to hours) — never behind a user-facing URL
Diagram: 5. CDN Delivery — Immutable Keys and Signed Access

🎯 Staff Insight: "Invalidate the CDN" is almost always the wrong answer. Invalidations are slow (seconds to minutes to propagate), often rate-limited or billed, and racy. Make objects immutable and change the URL instead — the only thing you ever update is a small metadata pointer.


Architecture Diagram#

Diagram: Architecture Diagram

Reading the diagram: numbered edges are the life of one upload. The control plane never touches bytes; the data plane never makes authorization decisions except by verifying signatures the control plane issued. The sweeper and lifecycle rules are first-class components, not afterthoughts — they are what keep the two planes consistent over months.


Failure Scenarios#

1. Upload Proxy Meltdown — 400 Concurrent Video Uploads Take Down the API#

Context: A marketing launch invites users to upload "your best 60-second clip." The upload endpoint proxies bytes through the general API fleet (40 pods, 8 GB RAM each). Average clip: 350 MB from mobile.

t=0        Campaign email lands; upload rate climbs from 2/s to 40/s
t=+30s     ~400 concurrent uploads; each buffers 16 MB chunks in memory
t=+2min    Pods at 90% memory; connection pools pinned by slow uploads (avg 4 min each)
t=+3min    LB 60s idle timeout kills stalled uploads; clients retry from byte 0
t=+4min    Retries double concurrency; OOM kills 11 pods; login and checkout 5xx at 35%
t=+20min   Upload endpoint disabled by feature flag; API recovers

Detection: http.inflight_requests{route="/upload"}, pod.memory_utilization > 85%, lb.idle_timeout_resets, and the cross-route symptom: http.5xx_rate{route="/login"} rising while login code hasn't changed.

Blast radius: The entire API — login, checkout, and feed — because uploads share the fleet. This is the defining failure of proxying: one heavy route starves every light one.

Mitigation: Kill switch on the upload route; return presigned URLs from a hotfix in the next deploy.

Prevention: Presigned direct uploads; if proxying is unavoidable, a dedicated upload-gateway fleet with its own concurrency limit and bulkhead. Load test uploads at 10× campaign forecast.

Owner: API platform team (fleet isolation); media team (upload path).

🎯 Staff Insight: Byte-heavy and request-heavy traffic must never share a resource pool. The moment a 4-minute upload competes with a 40 ms login for the same connection slots, the upload wins and your SLO loses.

2. The Ghost Files — Lost Events Leave 180K Uploads Stuck in PENDING#

Context: Storage event notifications feed a queue consumed by the metadata service. A misconfigured deployment changed the consumer's IAM role; events were delivered, failed authorization on processing, and went to a DLQ that nobody alerted on.

Day 0 14:00   Deploy changes consumer role; event processing fails silently to DLQ
Day 0 14:05   Users see "Processing..." forever on new uploads
Day 0 16:30   Support tickets spike: "my file disappeared"
Day 0 17:10   On-call finds DLQ depth = 180K
Day 0 17:40   Role fixed; DLQ redriven; conditional updates absorb duplicates
Day 0 19:00   Backlog clear; ~2% of users had re-uploaded (duplicate blobs, double quota)

Detection: blob.pending_age_p99 (normal < 2 min), dlq.depth{queue="blob-events"} > 0 for 5 minutes, and ratio blob.uploaded_rate / blob.created_rate falling below 0.8.

Blast radius: Every upload for 3 hours; user trust ("the app loses files"); quota double-counting from re-uploads.

Mitigation: Redrive DLQ; the hourly sweeper's HEAD fallback would have caught it within an hour had it existed.

Prevention: (1) Alert on DLQ depth > 0 for any tier-1 queue; (2) sweeper that verifies PENDING rows against storage independently of events; (3) dedup re-uploads by sha256 per owner to avoid quota double-count.

Owner: Media platform team owns the event consumer, DLQ, and sweeper.

🎯 Staff Insight: Events are the fast path; reconciliation is the correctness path. Any design that relies only on storage events for state transitions will eventually lose some — to IAM changes, poison messages, or a region incident. The sweeper is not optional.

3. Orphaned Multipart Parts — $38K/Month of Invisible Storage#

Context: A desktop app uploads large design files via multipart. Users often quit mid-upload. There is no lifecycle rule for incomplete multipart uploads. Incomplete parts don't show up in normal object listings.

Month 0     Multipart launched without AbortIncompleteMultipartUpload rule
Month 6     Storage bill growing 12% MoM while object count grows 4% MoM
Month 9     FinOps flags bucket: 1.6 PB billed vs 0.9 PB in object inventory
Month 9     Investigation: ~700 TB in incomplete multipart uploads, oldest 9 months
Month 9+1w  Lifecycle rule (abort after 3 days) applied; bill drops ~$16K/mo immediately

Detection: Divergence between billed storage and inventory-reported object bytes (storage.billed_bytes - storage.inventory_bytes); a count of ListMultipartUploads older than 3 days.

Blast radius: Financial only — but compounding and invisible; ~$150K spent before discovery.

Mitigation/Prevention: Bucket-creation template that mandates the abort-incomplete lifecycle rule (1–7 days) and a monthly storage inventory reconciliation.

Owner: Storage platform team owns bucket templates; FinOps owns the anomaly alert.

🎯 Staff Insight: Cost leaks in blob systems are silent because nothing is broken — uploads work, downloads work, and the only signal is an invoice. Treat storage.billed_bytes vs storage.inventory_bytes divergence as an alert with an owner, exactly like an error rate.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Upload proxy saturationhttp.inflight_requests{route=upload}, pod memoryWhole API fleetPresigned direct upload; isolated gatewayAPI platform
Lost / failed storage eventsblob.pending_age_p99, dlq.depthAll new uploadsDLQ redrive; sweeper with HEAD fallbackMedia platform
Orphaned multipart partsBilled vs inventory bytesStorage billAbort-incomplete lifecycle ruleStorage platform
Processing stuck (transcode)pipeline.step_age_p99{step}, lease expirationsFiles stuck "processing"Heartbeats, lease expiry, segment-level resumeMedia pipeline team
Leaked signed URL / hotlinkingcdn.requests_by_referrer, egress spike per objectEgress bill, content exposureShorten TTL, rotate signing key, per-object blockSecurity + CDN team
Viral object origin overloadcdn.origin_fetch_rate{object}, shield CPUOrigin + shield regionOrigin shield, request collapsing, pre-warmCDN team
Malware served before scanblob.ready_without_scan (should be 0)Users, legalGate READY on scan; quarantineTrust & safety + media platform

The Principal Lens#

Why L7 Sees This Problem Differently#

A Staff engineer designs one excellent upload path. A Principal engineer notices that the company has eleven — chat attachments, profile photos, support-ticket screenshots, data imports, ML dataset uploads, marketing assets — each with its own presigning code, its own (or missing) malware scan, its own lifecycle rules, and its own egress bill. The L7 problem is that user-supplied bytes are simultaneously the company's largest storage cost, one of its largest attack surfaces, and a compliance liability (retention, legal hold, erasure). That makes blob handling a platform with a contract, not a pattern each team reimplements.

The Org-Level Fault Line#

Shared media platform vs per-product blob handling. A central platform gives consistent security (one scanner, one signing policy), consistent lifecycle (erasure that actually reaches every copy), and consolidated CDN/storage commits worth 20–40% in discounts. But product teams have genuinely different needs — video transcoding ladders, sync-style block dedup, ML datasets measured in TB — and a platform that tries to serve all of them becomes a bottleneck. The Principal position: centralize the control plane (signing, scanning, lifecycle, deletion, metering) as a mandatory service; let teams own their processing steps as plug-ins to the pipeline.

Cost Model#

ScaleStoredMonthly EgressStorage $/moDelivery $/moPeopleAssumptions
Small (a startup app)50 TB100 TB~$1.2K~$5–9K (CDN list)0.5 FTEStandard tier, no lifecycle tiering
Mid (consumer app, 10M MAU)3 PB5 PB~$40–70K (tiered: 70% IA/archive)~$100–200K (committed CDN)4–6 FTE media platformLifecycle to IA at 30d; transcodes ~15% of compute budget
Large (video/photo platform)100+ PB100+ PB~$1–1.5M (tiered, negotiated)~$1–3M, or own edge30–60 FTERepatriation or own CDN becomes rational; see build-vs-buy

Assumptions: list-price order of magnitude — standard object storage ~$0.02/GB-month, infrequent access ~$0.01, deep archive ~$0.001; CDN $0.02–0.08/GB depending on commit; $300K fully loaded FTE. Egress, not storage, is usually the dominant and fastest-growing line.

🧭 Principal Move: "Most of what we store is never read again after 30 days. I want default lifecycle policies on every bucket at creation — IA at 30 days, archive at 180 unless a team opts out with a reason — and per-team showback so the teams generating the bytes see the bill."

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to ReverseWhy
Presigned vs proxied uploadTwo-way~1–2 engineer-monthsClient change + API change; no data moves
Part size / concurrency defaultsTwo-wayConfig + SDK releasePurely client behavior
Object key scheme (content-addressed vs path-based)One-way-ishRewriting billions of keys + metadata backfillEvery reference, URL, and cache entry depends on it
Storage provider for the primary blob storeOne-wayEgress to move PBs (~$0.02–0.09/GB) + months of dual-write5 PB at $0.05/GB ≈ $250K egress alone, before engineering
Public vs signed URLsOne-way for leaked contentCan't un-leak; public URLs get embedded everywhereStart signed; making things public later is easy
Retention promises in ToSOne-wayLegal and contractualOnce promised "we keep it forever," deletion becomes a legal question

The Standard I'd Write#

RFC: User-Supplied Blob Handling

Scope: Any service that accepts files from users or partners, or serves stored files to users.

Mandatory requirements:

  • Uploads over 1 MB MUST go directly to object storage via URLs issued by the shared signing service; API fleets MUST NOT proxy bytes.
  • Presigned URLs MUST carry size, content-type, and checksum conditions and a TTL ≤ 15 minutes (upload) or ≤ 60 minutes (download).
  • Objects MUST NOT be served to anyone other than the uploader until the malware scan passes.
  • Every bucket MUST have an abort-incomplete-multipart rule (≤ 7 days) and a documented lifecycle/retention policy.
  • Every blob MUST be registered in the metadata service with owner, retention class, and derived-object links, so erasure reaches all copies within 30 days.
  • Uploads over 100 MB SHOULD use the resumable SDK.

Exceptions: Regulated pre-storage inspection flows use the isolated upload gateway; approved by security and the media platform lead.

Success metrics: Upload success rate ≥ 99% for files ≤ 1 GB; blob.pending_age_p99 < 5 min; zero blobs served unscanned; erasure SLA met for 100% of requests; storage cost per MAU flat or declining year over year.

What I'd Tell the VP#

Every product team currently handles customer files its own way, which means we have several different security gaps and a storage bill growing faster than our user base. I want one shared service that every team uses to accept and serve files. It will scan everything before anyone can see it, automatically move old files to cheaper storage, and make sure that when a customer asks us to delete their data, every copy actually goes away. This should cut storage and delivery cost growth roughly in half over a year and closes a compliance risk we can't currently prove we've addressed.

Principal Interview Signals#

SignalWhat It Sounds Like
Platformization"This is the fourth upload path in the company — I'd make signing, scanning, and lifecycle a shared service and let teams plug in processing."
Egress-first cost thinking"Storage is cheap; egress is where the bill lives. I'd model reads per stored GB before choosing tiers or a CDN contract."
Lifecycle to the end"Deletion is part of the design: primary, replicas, derivatives, CDN caches, backups — with an SLA we can prove to an auditor."
One-way door awareness"The key scheme is the decision I'd spend time on — it's in every URL and cache entry. Part size I'd tune later."
Org-wide attack surface"User uploads are an attack surface for the whole company. One scanner, one signing policy, one abuse runbook."

Staff answers that L7 interviewers find insufficient:

  • "Presigned URLs, multipart, CDN, done" — technically excellent for one product, but silent on the ten other upload paths and their inconsistent security.
  • "We'll add lifecycle rules to move old data to Glacier" — without asking who decides retention per data class, or what the ToS promised.
  • "Delete the object from S3" — ignores derivatives, CDN caches, backups, and the audit trail that proves the deletion happened.

In the Wild#

Dropbox: Block-Level Chunking and Content Addressing#

Dropbox has publicly described splitting files into fixed-size blocks (4 MB) identified by their SHA-256 hash. The client uploads only blocks the server doesn't already have; a file is a list of block hashes in the metadata layer. Editing a large file re-uploads only the changed blocks, and identical blocks across files are stored once.

Staff insight: Dropbox moved the consistency problem from "blob vs row" to "block list vs block store" — a file becomes READY when every referenced block exists, and garbage collection must never delete a block still referenced. In an interview, content addressing is the right answer for sync products and the wrong one for upload-once media, where the metadata cost buys nothing.

Netflix: Processing Pipelines and Owned Delivery#

Netflix has written publicly about per-title encoding — choosing a bitrate ladder per title based on its visual complexity rather than a fixed ladder — and about parallelizing encodes by splitting source video into chunks. Delivery runs on Open Connect, Netflix's own CDN appliances placed inside ISP networks, with popular content pre-positioned during off-peak hours.

Staff insight: At Netflix scale the processing pipeline and delivery network are the product, so both were built. The transferable lessons at any scale: split long transcodes into parallel segments, treat renditions as immutable derived objects, and push popular content to the edge before demand arrives instead of on first miss.

In 2023 Discord announced that attachment links on its CDN would become authenticated and expire after a period, citing abuse — attachments being used to host and distribute malware via permanent, publicly shareable URLs. Links shared inside Discord are refreshed automatically; links pasted elsewhere stop working.

Staff insight: Permanent public URLs for user content are a one-way door: once links are embedded across the internet, the platform effectively becomes a free file host for whoever wants one. Signed, expiring URLs from day one keep that decision reversible — a strong point to make unprompted in any chat or social design.


Practice Drill#

Prompt: "Design video upload for a short-video app: 5M uploads/day, average 80 MB, up to 2 GB, mostly from phones. Videos should appear in the uploader's profile immediately and be publicly viewable within 2 minutes."

Staff Answer

5M × 80 MB ≈ 400 TB/day ingest, ~58 uploads/sec average and ~200/sec at peak — bytes must bypass the API. Client calls POST /videos with size, duration, and SHA-256; the API checks quota, inserts a PENDING row (so the profile can show a local-preview placeholder immediately), and returns a multipart upload with 16 MiB parts, presigned in batches, 15-minute TTLs, size and checksum conditions. The resumable SDK runs 4 parallel parts, retries per part, and resumes via ListParts. The storage event moves the row to UPLOADED after a server-side HEAD; an hourly sweeper covers lost events, and a lifecycle rule aborts incomplete uploads after 3 days. Processing fans out: gating steps (malware scan, moderation model — ~10–30 s) and enriching steps (thumbnail, 3 renditions via segmented parallel transcode). To hit the 2-minute public SLA, publish at the first playable rendition (480p, ~30–60 s for a 60 s clip) once gating passes; higher renditions arrive later. Delivery is CDN with content-addressed keys and immutable caching; public videos use unsigned edge URLs, private ones short-lived signed URLs. Raw originals move to IA after 30 days — they are rarely read after transcoding.

Why this is L6:

  • Does the ingest math first and derives "bytes bypass the API" from it
  • Makes the PENDING → READY state machine explicit, with a reconciler for lost events
  • Separates gating from enriching steps to meet the 2-minute SLA without skipping safety
  • Designs cost in: immutable CDN keys, lifecycle tiering of raw originals

What L7 adds:

  • Prices it: ~12 PB/month ingest means tiering and original-retention policy are million-dollar decisions, so retention per data class needs a product/legal owner
  • Makes upload/scan/lifecycle a shared platform so chat, stories, and ads uploads inherit it
  • Plans multi-CDN and the egress threshold at which owned edge capacity becomes worth evaluating

Prompt: "Users report that files they deleted are still accessible via old links. What happened and how do you fix it?"

Staff Answer

Deletion removed the metadata row and the original object but missed at least one copy: derived renditions under a different prefix, CDN-cached copies with long max-age, or permanent public URLs. Immediate fix: purge CDN paths for the affected objects and delete derived objects found via the metadata links. Durable fix: deletion becomes a workflow (READY → DELETING → deleted) that enumerates every derived object from the metadata record, deletes them, issues CDN purges, and verifies with HEAD before completing; downloads use signed URLs with ≤ 60-minute TTLs so any missed cache expires quickly; backups honor erasure through a tombstone list applied on restore.

Why this is L6:

  • Enumerates every place a copy lives, not just the primary object
  • Turns deletion into a verifiable state machine with an owner

What L7 adds:

  • Sets an org-wide erasure SLA (e.g., 30 days, all copies) with audit evidence for regulators
  • Requires every new derived-data producer to register its outputs with the metadata service — deletion correctness is a platform contract

Staff Interview Application#

How to Introduce This Pattern#

"Before drawing the upload path, I want to separate the control plane from the data plane. Our API decides who may upload what and hands out a short-lived presigned URL; the bytes go directly to object storage. Because metadata and blob then live in two systems with no shared transaction, I'll make the upload lifecycle an explicit state machine with a reconciler, and deliver reads through a CDN with immutable keys."

Say the separation first — it's the single sentence that moves you from L5 to L6 on this topic.

When NOT to Use This Pattern#

  • Payloads under ~1 MB. Avatars and small documents can go through the API with a hard size cap; presigning adds a round trip and a state machine for no benefit.
  • Structured data that needs querying. A 50 MB JSON blob you'll filter and update belongs in a database or columnar format, not an opaque object.
  • Strict pre-storage inspection mandates. Use an isolated streaming gateway instead of direct-to-storage upload.
  • Internal, low-volume tooling. A 100-uploads-a-day admin tool doesn't need multipart, sweepers, and CDNs.

Follow-Up Questions to Anticipate#

Interviewer AsksWhat They Are TestingHow to Respond
"What if the upload succeeds but the client crashes before telling you?"Two-system consistency"The storage event confirms it independently; the sweeper HEADs stale PENDING rows. The client's complete call is only a hint."
"How do you stop someone uploading a 50 GB file?"Presign scoping"The signed policy pins content-length range and type; storage rejects anything else. Quota is reserved at URL issue time."
"How do you resume a 2 GB upload after a drop?"Multipart mechanics"ListParts returns what's stored; the client uploads the missing parts only. At 16 MiB parts, a drop costs at most 16 MiB."
"How do you serve private files through a CDN?"Edge auth"Short-lived signed URLs or signed cookies; the edge verifies the signature, so origin never sees auth traffic."
"A video goes viral — what breaks?"Delivery at scale"Origin fetches for one object. Origin shield plus request collapsing means one fetch per region, and immutable keys keep hit rate above 95%."
"How do you delete a user's data?"Lifecycle completeness"A deletion workflow over every derived object, CDN purge, backup tombstones, and verification before marking complete."

Evaluation Rubric#

DimensionSenior (L5)Staff (L6)Principal (L7)
Upload pathAPI receives and storesPresigned direct, multipart, scoped policyShared signing service and SDK standard for all products
ConsistencyWrite row after uploadPENDING state machine, event + sweeperLifecycle contract through deletion and erasure SLAs
ProcessingSynchronous thumbnailIdempotent async pipeline, gating vs enrichingPluggable pipeline platform; scanning mandated org-wide
Delivery & costServe from bucketCDN, immutable keys, lifecycle tieringEgress cost curve, CDN commits, multi-CDN, own-edge threshold

Strong hire signals: says "bytes never touch our servers" unprompted; names the orphan problem in both directions (rows without objects, objects without rows); separates gating from enriching steps; puts a number on part size and TTL.

Lean no-hire signals: stores blobs in the relational DB; retries whole uploads; marks files ready on the client's word; no mention of malware or access control on reads.

Common false positive: Detailed knowledge of S3 API calls ≠ blob-handling judgment. A candidate can recite CreateMultipartUpload and CompleteMultipartUpload and still miss the lost-event reconciler and the deletion lifecycle — the parts that actually page someone.


Capacity Planning Quick Reference#

NumberValueContext
S3 single PUT max5 GiBUse multipart well before this — anything over ~100 MB
Multipart part size5 MiB min – 5 GiB max, ≤ 10,000 parts16–64 MiB default; size from declared length for huge files
Per-prefix request rate (S3)~3,500 writes/s, ~5,500 reads/sHash-prefix keys for high-throughput ingest
Presigned URL TTL5–15 min upload, ≤ 60 min downloadSigV4 maximum is 7 days — never use it
1 GB over 10 Mbps uplink~14 minWhy mobile needs resumability
1 GB over 100 Mbps~80 sHome broadband; parallel parts help saturate
CDN hit rate target≥ 95% for immutable mediaBelow 90% usually means bad cache keys or short TTLs
Storage price order of magnitude~$0.02 / ~$0.01 / ~$0.001 per GB-monthStandard / infrequent access / deep archive
Egress price order of magnitude~$0.02–0.09 per GBCDN committed vs cloud internet egress list
Abort incomplete multipart1–7 daysMandatory lifecycle rule
Stale PENDING sweep24 hPlus hourly HEAD reconciliation

Pitfalls checklist:

  • Do bytes ever pass through the general API fleet?
  • Does the presigned policy pin size, type, and checksum?
  • What happens to a row whose storage event never arrives?
  • What happens to an object that has no row?
  • Is there an abort-incomplete-multipart lifecycle rule on every bucket?
  • Can anything be served before the malware scan passes?
  • Are delivered objects immutable and content-addressed, or do we rely on CDN invalidation?
  • Does deletion reach derivatives, CDN caches, and backups — and can we prove it?
  1. Loading the index…