Hiring BarSupport

S3 & Object Storage

Technology guide43 min read5 diagrams

Go deeper:

Why This Matters#

Object storage is not a file system and it is not "infinite disk." It is a key-value store for immutable blobs with a billing model attached to every verb. You pay per gigabyte stored, per request made, per gigabyte that leaves the region, per object transitioned between tiers, and per day an object sits in a tier before its minimum duration ends. Almost every production incident with S3 is one of three things: a request pattern the bucket hasn't scaled for yet (503 Slow Down), a correctness assumption the API never made (two writers to one key, an event that arrived twice), or a bill nobody modelled (egress, small objects in archive tiers, millions of abandoned multipart parts).

That is why "and we put the files in S3" is the sentence interviewers lean on. The L5 candidate draws a bucket and moves on. The L6 candidate says "clients upload directly with a presigned multipart URL that expires in 15 minutes; keys are tenant/yyyy/mm/dd/uuid so writes spread over prefixes; the metadata row is written first and flipped to READY by an at-least-once event handler that is idempotent on (key, version); lifecycle moves objects to Standard-IA at 30 days, aborts incomplete multipart uploads after 7, and expires noncurrent versions after 30; delivery goes through a CDN so origin egress stays under 10% of bytes served." The L7 candidate asks which of the company's 4 PB actually gets read, what egress will cost at 10× traffic, and whether the compliance requirement means Object Lock in compliance mode, which nobody, including the root account, can undo.

The L5 → L6 gap is not knowing what a bucket is. It is knowing that S3 gives you strong per-key consistency and nothing else: no cross-key transactions, no locks, no ordering of events, no free requests. Every guarantee above the single key is yours to build.

The L5 → L6 → L7 Contrast#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Store files in S3, URL in the database""What's the object size distribution, read/write ratio and retention? That decides key layout, upload path, storage class and whether a CDN fronts it.""Which data is the system of record, which is derived, and what does each cost per year at 3 scales? Derived data shouldn't be stored for 7 years."
Consistency"S3 is eventually consistent" (outdated) or "S3 is consistent, so we're fine"Strong read-after-write per key, including LIST; concurrent writers are last-writer-wins; conditional writes (If-None-Match, If-Match) for create-once and compare-and-swap; no cross-key atomicitySets the org pattern: metadata in a database as the source of truth, objects immutable and content-addressed so consistency questions mostly disappear
Throughput"S3 scales infinitely"3,500 writes and 5,500 reads per second per partitioned prefix, scaling gradually with 503s on the way; spread keys, back off with jitter, pre-warm for known launchesDecides when a workload has outgrown general-purpose buckets (single-AZ low-latency class, caching tier, or a different store)
Cost"Storage is cheap"Models storage + requests + egress + transitions; lifecycle rules for IA/Glacier with minimum durations; abort incomplete multipart uploadsOwns the storage bill as a portfolio: chargeback by bucket tag, egress architecture (CDN, same-region compute), retention policy signed by legal
Failure"S3 has 11 nines of durability"Durability ≠ protection from your own deletes: versioning, replication, Object Lock, and event-loss reconciliationDesigns the org's data-loss posture: which buckets get WORM, cross-account backups, and a tested restore time
Ownership"The app team owns the bucket"Bucket owner, lifecycle owner and event-consumer owner named; orphan sweeps scheduledStorage platform owns guardrails (Block Public Access, encryption, tagging); product teams own retention and access patterns
Why "Consistency" separates levels

Many candidates still recite "S3 is eventually consistent." That has been false since December 2020: S3 now provides strong read-after-write consistency for PUTs and DELETEs of objects in all regions, and a LIST issued after a successful write includes the new key (S3 consistency model). The Staff answer knows what that guarantee does not cover. Two simultaneous PUTs to the same key resolve as last-writer-wins by timestamp; there is no lock. Updates are per key, so "write the data file, then write the manifest" can be observed half-done. Bucket configuration (versioning, lifecycle, policies) is still eventually consistent; AWS recommends waiting 15 minutes after first enabling versioning before writing. Event notifications are delivered at least once and can take a minute or longer. The Staff move is to use conditional writes for create-once semantics and to keep cross-object invariants in a database, not in S3.

Why "Throughput" separates levels

"S3 scales infinitely" is true in aggregate and false at the moment you need it. AWS documents at least 3,500 PUT/COPY/POST/DELETE or 5,500 GET/HEAD requests per second per partitioned prefix, with no limit on the number of prefixes, and states that scaling to higher rates "happens gradually and is not instantaneous," with 503 Slow Down errors while it does (S3 performance). A launch that goes from 200 to 20,000 writes per second on one new prefix will see 503s for minutes. The Staff answer spreads keys, retries 503s with exponential backoff and jitter, and pre-scales known events with a load ramp. The Principal answer asks whether the request rate itself is the design smell: 20,000 tiny PUTs per second is a batching problem before it is a storage problem.

The 60-Second Pitch#

"User media goes in S3 as immutable objects keyed media/{tenant}/{yyyy}/{mm}/{uuid}, so writes spread over many prefixes and no key is ever overwritten. Clients upload directly with presigned multipart URLs, 16 MiB parts, 15-minute expiry; our API only issues permission and records metadata. A PENDING row is written first; an S3 event into SQS flips it to READY after a HEAD check, and an hourly reconciler catches lost events. Reads go through a CDN with year-long cache headers because keys are content-addressed, which keeps origin egress small. Lifecycle moves objects to Standard-IA at 30 days and Glacier Instant Retrieval at 90, aborts incomplete multipart uploads after 7 days, and versioning plus replication to a second account protects against our own mistakes. S3 gives us 11 nines of designed durability and strong per-key consistency; everything across keys, including 'which objects belong to this album', lives in the database."

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
User content store (photos, video, attachments)Many writers, read-heavy, CDN-fronted, per-user accessPresigned direct upload, immutable keys, metadata DB as truth, CDN deliveryOrphaned objects, lost events leave rows PENDING, egress billNo broken links; no cross-tenant reads
Data lake / analyticsHuge objects, scan-heavy, many parallel readersColumnar files (Parquet) of 128 MiB–1 GiB, table format manifests, partitioned prefixesSmall-file explosion; LIST-heavy planners; 503s on hot partitionsReaders see complete snapshots only
Backup, archive and complianceRarely read, retained for years, must survive deletionVersioning, Object Lock, cross-account replication, Glacier tiersMinimum-duration and per-object fees on tiny files; slow restores; accidental compliance-mode lockRestorable within the stated RTO; immutable for the retention period

🎯 Staff Move: "I'll design for the user content store. The analytics lake and the compliance archive get their own buckets and policies, because their lifecycle rules, key layouts and access patterns conflict. One bucket serving all three is how a lifecycle rule meant for logs ends up archiving profile photos."

The Staff Positions#

PositionRationale
Metadata lives in a database; S3 holds bytesS3 has per-key atomicity only. "Which files are in this folder" and "who can see this" need transactions and indexes.
Objects are immutable; new content gets a new keyNo overwrite races, perfect CDN caching, versioning stays cheap.
Bytes never transit application serversPresigned URLs move upload and download traffic off your fleet; a 2 GB proxy upload pins a connection for minutes.
Every event consumer is idempotent and backed by a reconcilerNotifications are at-least-once and can be delayed; a sweep is the only way to be sure.
Lifecycle rules ship with the bucketAbort incomplete multipart uploads, expire noncurrent versions, tier cold data. A bucket without them grows a bill nobody owns.
Egress is designed, not discoveredCDN in front, compute in the same region, cross-region replication only where the RPO demands it.
Block Public Access on, encryption on, by defaultPublic buckets are a security incident waiting for a headline; exceptions go through review.

Architecture & Internals#

S3's internals are not public in detail, and you don't need them. Five properties change design decisions: the key namespace and partitioning, the consistency model, the multipart protocol, storage classes, and the event and lifecycle machinery. For how such a system is built underneath (erasure coding, placement, repair), see Design an Object Store.

The Namespace: Flat Keys, Partitioned by Prefix#

A bucket is a flat map from key to object. / has no meaning to S3; "folders" are a console convention over shared key prefixes. Internally the index is partitioned by key range, and S3 splits hot ranges as load grows. That is where the documented per-prefix rates come from.

Diagram: The Namespace: Flat Keys, Partitioned by Prefix

Why it matters in design: throughput is bounded per key range, not per bucket. Keys that all start with the same timestamp or the same tenant ID concentrate writes on one range until S3 splits it, and splitting takes time. The modern guidance is not "add random hashes to every key" (that advice predates the 2018 rate increase), but "make sure the leading characters that vary actually vary for the hot workload." A {tenant}/{date}/{uuid} layout spreads load across tenants; {date}/{tenant}/{uuid} sends the whole day's writes to one prefix.

Consistency: Strong Per Key, Nothing Across Keys#

OperationGuarantee todayWhat it doesn't give you
PUT new key, then GETReturns the new object—
PUT overwrite, then GETReturns the new data, never partialProtection from a concurrent writer (last timestamp wins)
DELETE, then GET or LISTObject gone from both—
PUT, then LISTNew key appearsSnapshot isolation across a multi-page LIST of a changing prefix
Two keys updated "together"Each atomic on its ownAny atomicity across them
Bucket config changeEventually consistentImmediate effect (wait ~15 min after first enabling versioning)
If-None-Match: * on PUTCreate only if absent; fails if the key existsLocking for long operations
If-Match: <etag> on PUTCompare-and-swap on the current ETagMulti-key transactions

Conditional writes are supported on PutObject, CompleteMultipartUpload and CopyObject, and conditional deletes on DeleteObject, with no extra charge beyond normal request rates (conditional requests). That turns S3 into a usable primitive for "first writer wins" (idempotent uploads, leader-written manifests) without an external lock.

🎯 Staff Insight: "S3 is strongly consistent per key, so the classic 'read your write' bug is gone. What's left is the multi-key problem: a reader can see the new data file before the new manifest. I keep the manifest pointer in a database or use a conditional write on a single manifest key, so readers flip from one complete snapshot to the next."

Multipart Upload: How Large Objects Actually Move#

Diagram: Multipart Upload: How Large Objects Actually Move
LimitValueDesign consequence
Single PUTUp to 5 GBAbove ~100 MB, use multipart anyway for retries and parallelism
Part size5 MiB to 5 GiB (last part may be smaller)16–64 MiB parts are a good default on mobile and desktop links
Parts per upload10,000Part size × 10,000 caps object size: 5 MiB parts top out near 48.8 GiB
Maximum object size48.8 TiB (~50 TB), raised from 5 TB in December 2025Pick part size from the largest object you accept: ~5 GiB parts for the max
Incomplete uploadsBilled as storage until completed or abortedAbortIncompleteMultipartUpload lifecycle rule, 1–7 days

Sources: multipart limits, 50 TB announcement. The object becomes visible atomically at CompleteMultipartUpload; readers never see a half-assembled object.

Storage Classes and Lifecycle#

ClassDesigned availabilityAZsMinimum durationMinimum billable sizeAccess
Standard99.99%≥ 3NoneNoneMilliseconds
Intelligent-Tiering99.9%≥ 3NoneObjects < 128 KB never tier downMilliseconds (archive tiers opt-in)
Standard-IA99.9%≥ 330 days128 KBMilliseconds, per-GB retrieval fee
One Zone-IA99.5%130 days128 KBMilliseconds; lost with the AZ
Express One Zone99.95%1NoneNoneSingle-digit milliseconds
Glacier Instant Retrieval99.9%≥ 390 days128 KBMilliseconds, higher retrieval fee
Glacier Flexible Retrieval99.99% after restore≥ 390 days40 KB metadata overhead per objectMinutes to hours, restore first
Glacier Deep Archive99.99% after restore≥ 3180 days40 KB metadata overhead per objectHours, restore first

Source: storage class comparison. Every class is designed for 11 nines of durability. What changes is availability, the AZ count, the minimum duration you pay for even if you delete early, the minimum billable size, and the retrieval path.

Diagram: Storage Classes and Lifecycle

The small-object trap: a 10 KB log line object moved to Standard-IA is billed as 128 KB, 12.8× its size, and every lifecycle transition is a per-object request fee. Archiving a billion 10 KB objects can cost more in transition requests and minimum-size padding than it saves. Batch small objects into larger archives (tar or Parquet, 100 MB+) before tiering them.

Events, Versioning and Object Lock#

  • Event notifications go to SNS, SQS, Lambda or EventBridge. AWS states they are "designed to be delivered at least once," typically in seconds but sometimes a minute or longer, and SQS FIFO queues are not a direct destination (event notifications). Treat them as hints that start work, not as a ledger.
  • Versioning keeps every overwritten or deleted version; a simple DELETE adds a delete marker. It is the cheapest protection against application bugs and the most common source of a mysteriously growing bill (noncurrent versions are billed like any object).
  • Object Lock (requires versioning) makes versions write-once-read-many. Governance mode can be bypassed by principals with s3:BypassGovernanceRetention; compliance mode cannot be shortened or removed by anyone, including the root user, until the retention date (Object Lock). Legal holds have no expiry until removed.
  • Presigned URLs are bearer tokens scoped to one operation on one key. With SigV4 and long-lived IAM user credentials they last up to 7 days; signed with role or STS credentials they die when the session does, often within 1–6 hours. Expiry is checked when the request starts, so a long download already in progress completes (presigned URLs).

Data Modeling / Core Usage — "The Entire Game"#

In DynamoDB the game is the partition key. In S3 it is the object contract: the key layout, the immutability rule, and the split between what lives in S3 and what lives in the metadata database. Get these right and the bucket scales, caches and ages out on its own.

Step 1: Classify the Objects#

Class of objectTypical sizeRead patternMutabilityHome
User originals (photos, video, docs)100 KB–5 GBRead-heavy for weeks, then coldImmutableStandard → IA → Glacier IR
Derived renditions (thumbnails, transcodes)5 KB–500 MBVery hot, CDN-cachedRegenerableStandard; expire or regenerate instead of archiving
Analytics files (Parquet)128 MiB–1 GiBScanned in parallelImmutable, compactedStandard or Intelligent-Tiering
Logs and events1 KB–10 MB rawRarely read after 7 daysAppend via new objectsBatch into large objects, then archive
Backups and compliance recordsMBs–TBsAlmost neverWORMSeparate account, Object Lock, Glacier

Step 2: Design the Key#

Good: media/{tenant_id}/{yyyy}/{mm}/{dd}/{uuid}.{ext}
  - leading variation across tenants spreads writes across prefixes
  - date components make lifecycle rules and audits easy
  - uuid (or content hash) makes keys immutable and unguessable

Content-addressed variant for dedupe and caching:
  blobs/{sha256[0:2]}/{sha256}        -> same bytes, same key, written once
  write with If-None-Match: *          -> a duplicate upload is a no-op (412)

Bad:   {yyyy-mm-dd-hh-mm-ss}-{userid}.jpg   -> every write hits the newest range
Bad:   users/{email}/photo.jpg              -> PII in keys, logs and URLs; overwrites
Bad:   one key rewritten every second       -> last-writer-wins races, version bloat

Step 3: Put Cross-Object Truth in the Database#

table objects (
  object_id    uuid primary key,
  owner_id     bigint not null,
  s3_key       text unique not null,
  sha256       bytea,
  size_bytes   bigint,
  status       text check (status in ('PENDING','UPLOADED','READY','FAILED','DELETED')),
  created_at   timestamptz,
  version      int            -- bumped on every state change; events compare it
);

Upload:   insert PENDING -> issue presigned URL -> client uploads
Confirm:  S3 event or client "complete" -> HEAD the key -> conditional update
          UPDATE objects SET status='UPLOADED', version=version+1
          WHERE object_id=$1 AND status='PENDING'          -- duplicates are no-ops
Sweep:    hourly: PENDING older than 24h -> HEAD; present -> UPLOADED, absent -> FAILED
Delete:   mark DELETED in DB first; delete or expire the object asynchronously

The database answers every question that spans objects: listings, permissions, quotas, search. S3 LIST is a sequential, 1,000-keys-per-page scan billed at PUT-tier request rates; it is an inventory tool, not a query engine. For bulk audits use S3 Inventory reports instead of LIST loops. The full upload state machine is in Large File Uploads & Delivery; the metadata store choice is covered in PostgreSQL and DynamoDB.

🎯 Staff Move: "The database row exists before the bytes do, and the object is only visible to users once the row says READY. That way a lost event or a half-finished upload is a stuck row my reconciler can see, not a broken link a user finds."


The Tunable Tradeoff — Cost × Latency × Protection#

Every object storage decision trades three things: what you pay per gigabyte-month, how fast and cheaply you can read it back, and how well it survives mistakes.

SettingCheap endProtected / fast endWho pays at the cheap end
Storage classDeep Archive (~1/20 the price of Standard)Standard or Express One ZoneWhoever needs the data back in under 12 hours
AZ redundancyOne Zone-IA, Express One ZoneMulti-AZ classesThe team whose only copy was in the failed AZ
VersioningOffOn, with noncurrent expirationAnyone whose bug overwrote a million objects
ReplicationNoneCross-region, cross-accountThe business during a regional or account-level incident
Object LockNoneCompliance modeSecurity, after ransomware or a malicious insider
Read pathDirect from originCDN in frontFinance, through egress
Monthly cost model (list prices, US East, illustrative):
  storage   = GB_stored x price_per_GB           (Standard ~$0.023/GB first 50 TB)
  requests  = PUTs/1,000 x $0.005 + GETs/1,000 x $0.0004
  egress    = GB_out_to_internet x ~$0.09/GB     (first 100 GB/month free, tiers drop with volume)
  extras    = transitions + retrievals + minimum-duration and minimum-size padding

Example: photo app, 500 TB stored, 2 B PUTs/month, 40 B GETs/month, 3 PB served
  storage   500,000 GB x ~$0.022    ~= $11K
  PUTs      2e9 / 1,000 x $0.005     = $10K
  GETs      40e9 / 1,000 x $0.0004   = $16K
  egress    3 PB direct at ~$0.05 blended ~= $150K   <- the real bill
  with CDN at 95% hit ratio: origin egress ~150 TB; CDN pricing replaces S3 egress

The lesson candidates miss: at serving scale, egress and requests dwarf storage. Price checks come from the S3 pricing page; plug your own numbers into the Cost Estimator and the Back-of-Envelope Calculator.

🎯 Staff Move: "Before I pick storage classes, I'll split the bill into storage, requests and egress at our projected traffic. For a media product egress is usually 5–10× storage, so the CDN and the key design that makes caching work are the cost decisions; the storage class is a rounding error until the data is cold."

Who Pays for Each Choice#

ChoiceWhat WorksWhat BreaksWho Pays
Everything in Standard foreverSimple, fastBill grows linearly with historyFinance, quietly, every month
Aggressive lifecycle to GlacierStorage cost drops ~80–95%Restores take hours; minimum duration and per-object feesSupport and users when old files are needed
One Zone classes~20% cheaper than multi-AZ IAData lost with an AZWhoever assumed it was regenerable
Versioning without noncurrent expiryUndo for every mistakeStorage silently doubles under churnThe bucket owner at quarter end
Direct-from-S3 servingNo CDN to runEgress and request bills; latency 100–200 ms first byteFinance and users
Compliance-mode Object LockNobody can delete, including attackersNobody can delete, including youLegal and finance if set wrong

Anti-Patterns — What Kills S3 Deployments#

1. Using S3 as a Database#

Listing a prefix to find "all files for user X," storing JSON state that several workers rewrite, or polling LIST for new work. LIST is paginated at 1,000 keys, billed like PUTs, and concurrent writers to one key silently lose updates. Fix: metadata and state in a database; S3 holds immutable bytes.

2. Proxying Bytes Through the Application#

A 2 GB upload through an API server pins a connection and memory for minutes and doubles egress. Fix: presigned direct upload and download; the app handles permission and metadata only.

3. Monotonic Key Prefixes on a Hot Write Path#

2026-10-06T14:00:00Z-... as the key start puts every write on the newest key range. Fix: lead with a high-cardinality component (tenant, shard, hash prefix) for hot write workloads; keep dates later in the key.

4. No Lifecycle Rules#

Incomplete multipart uploads, noncurrent versions and temporary processing outputs accumulate forever. Teams discover terabytes of invisible parts during a cost review. Fix: every bucket ships with abort-incomplete-multipart (1–7 days), noncurrent-version expiry and tiering rules.

5. Tiny Objects in Archive Tiers#

Millions of 2–20 KB objects moved to IA or Glacier pay 128 KB minimum billing or 40 KB metadata overhead each, plus a transition fee per object. Fix: compact into large archive objects first; keep tiny objects in Standard or Intelligent-Tiering, where objects under 128 KB simply stay in the frequent tier.

6. Trusting Events as the Source of Truth#

A Lambda that marks uploads READY on each event, with no reconciler. Events are at-least-once and occasionally delayed past a minute; a misconfigured filter or a disabled consumer drops them entirely. Fix: idempotent handlers keyed on (key, version or etag) plus a periodic sweep comparing the database to the bucket.

7. Long-Lived Presigned URLs as an Access Model#

Seven-day URLs pasted into emails, chat and logs become public links. Fix: short expiry (minutes), issued per request after an authorisation check; signed CDN URLs or cookies for repeated access.

8. Public Buckets and Wide Bucket Policies#

The most common S3 security incident is still a bucket or prefix made world-readable to "fix" an access problem. Fix: Block Public Access on at the account level, CDN with origin access control for public content, and policy-as-code review; see Security Fundamentals.


The Technology Landscape — Head-to-Head Comparison#

DimensionAmazon S3 (general purpose)S3 Express One ZoneGoogle Cloud Storage / Azure BlobSelf-hosted (MinIO, Ceph RGW)Block or file storage (EBS, EFS)
ModelFlat key → immutable object, HTTP APISame API, directory buckets, single AZSame model, different APIs and tiersS3-compatible API on your hardwareDisks and POSIX file systems
ConsistencyStrong per key, LIST includedStrongStrong per objectDepends on deploymentStrong (single attach or NFS semantics)
Latency~100–200 ms first byte for small objectsSingle-digit msSimilar to S3LAN-dependent, often lowerSub-ms to low ms
Throughput modelPer-prefix rates, scales graduallyHundreds of thousands of requests/s per bucketPer-bucket and per-key guidanceBounded by your clusterPer volume IOPS and throughput
Durability11 nines designed, multi-AZ11 nines designed, one AZMulti-zone and multi-region optionsWhat you build and testVolume-level replication, snapshots
Ops burdenNone for the service; yours for lifecycle and costSameSameHigh: disks, repair, upgradesLow–medium
Pick whenDefault for blobs, lakes, backupsHot intermediate data, ML training scratchYou live in that cloudData sovereignty, on-prem, egress economics at PB scaleA process needs a disk, not an API

🎯 Staff Insight: "The S3 API is the industry's de facto blob interface, so portability is about the API surface I depend on, not the vendor. I keep to PUT, GET, ranged GET, multipart, conditional writes and lifecycle, and avoid building core logic on vendor-only event or query features unless the payoff is clear."


Patterns#

Pattern 1: Presigned Direct Upload with a Metadata State Machine#

Client asks the API for permission; the API writes a PENDING row and returns presigned part URLs scoped to one key, a content-length range and a 15-minute expiry; the client uploads directly; an event or a client complete call moves the row forward after a HEAD check; an hourly reconciler fixes what events miss. Full mechanics in Large File Uploads & Delivery.

Pattern 2: Content-Addressed, Immutable Objects#

Key = hash of content. Writes use If-None-Match: *, so duplicate uploads become no-ops and dedupe is free. CDN headers can be max-age=31536000, immutable because a key never changes meaning. This is how Cloud File Sync designs store chunks.

Pattern 3: Event-Driven Processing Pipeline#

Diagram: Pattern 3: Event-Driven Processing Pipeline

Events start the work; the database records the truth; the reconciler closes the gap. Write derived output to a different bucket or prefix than the trigger watches, or a function that writes back into its own trigger prefix loops forever, as the S3 documentation itself warns. This is the standard video pipeline shape.

Pattern 4: Data Lake with Table-Format Manifests#

Write large immutable Parquet files (128 MiB–1 GiB), then commit a new snapshot by atomically swapping a manifest pointer (an open table format, or a conditional write on a single metadata key). Readers resolve the manifest and never LIST the data prefix. Compaction jobs rewrite small files into large ones. See Batch & Stream Pipelines and Real-Time OLAP.

Pattern 5: Backup Vault in a Separate Account#

Production replicates (or a backup job copies) into a bucket owned by a separate account, with versioning, Object Lock and a policy that production principals can write but never delete. Ransomware or a compromised deploy role in production cannot reach the vault.


Scaling#

The Numbers#

ResourceDocumented limit or typical valueDesign note
Request rate≥ 3,500 writes and 5,500 reads per second per partitioned prefix; no prefix limitScales gradually; expect 503 Slow Down while it does
Small-object latency~100–200 ms first byte (documented typical)Cache hot objects; Express One Zone for single-digit ms
Single-instance throughputUp to the instance NIC, e.g. 100 Gb/s, with parallel ranged GETsParallelism, not bigger requests, drives throughput
Single PUT5 GBMultipart above ~100 MB
Multipart10,000 parts, 5 MiB–5 GiB eachPart size × 10,000 must exceed your largest object
Maximum object size48.8 TiB (~50 TB)Raised from 5 TB in December 2025
LIST page1,000 keys1 B keys = 1 M sequential LIST calls; use Inventory
Event deliveryAt least once; seconds, sometimes a minute+Reconcile; don't use for ordering

Request-Rate Math#

Launch: 50 M users upload ~2 photos/day, peak hour is 3x average
  writes/s avg  = 100 M / 86,400       ~= 1,160/s
  writes/s peak = ~3,500/s              -> already at one prefix's write rate
  each photo also writes 3 thumbnails   -> ~14,000 PUTs/s at peak

Key layout media/{tenant_shard 00-ff}/{date}/{uuid}:
  256 leading shards -> ~55 PUTs/s per shard at peak, far under per-prefix rates
  S3 still needs time to split ranges for a brand-new bucket:
  ramp synthetic load over days before launch, or expect 503s in hour one

Run the same numbers through the Back-of-Envelope Calculator and the Sharding Planner when the key layout is the question.

Scaling Moves in Order#

  1. Fix the key layout so the hot workload varies in its leading characters.
  2. Retry 503s with exponential backoff and jitter in every client; SDK defaults help but stacked retries across layers do not.
  3. Parallelise: multipart uploads, ranged GETs (8–16 MiB ranges), many connections. Single-stream throughput is the wrong thing to tune.
  4. Cache in front: CDN for public reads, an in-memory cache for small hot objects, so S3 sees misses only.
  5. Batch small writes into larger objects (logs, events, telemetry). Fewer, bigger objects cut request bills and LIST cost.
  6. Change class or store when latency is the limit: Express One Zone for hot scratch data, or a database or cache for small mutable values that never belonged in object storage.

Failure Modes & Recovery#

1. 503 Slow Down at a Launch or Backfill#

  • Symptom: Error rate on uploads or a batch job jumps to 5–30% with 503 Slow Down; retries make it worse.
  • Root cause: Request rate on a key range climbed faster than S3 repartitions; often a new prefix or a monotonic key.
  • Detection: S3 request metrics 5xxErrors by prefix filter; client-side SlowDown counters; retry rate per caller.
  • Fix: Throttle the producer, add jittered backoff, spread keys; for a backfill, ramp concurrency gradually.
  • Prevention: Key layout review in design; load ramp before known launches; client retry budgets so one job can't amplify.

2. Lost or Delayed Events — Uploads Stuck in PENDING#

  • Symptom: Users see "processing" for hours; support tickets spike; metadata shows thousands of PENDING rows older than an hour.
  • Root cause: Consumer down or erroring into a DLQ nobody watches, a notification filter changed in a deploy, or events delayed.
  • Detection: Age of oldest PENDING row; DLQ depth; ApproximateAgeOfOldestMessage on the queue.
  • Fix: Replay the DLQ; run the reconciler on demand (HEAD each stuck key, advance or fail).
  • Prevention: Reconciler on a schedule from day one; alert on PENDING age, not just consumer errors; notification config in code.

3. The Surprise Bill — Egress, Requests or Invisible Storage#

  • Symptom: S3 spend doubles in a month without a matching traffic change.
  • Root cause: Direct-from-origin serving after a CDN change, a job LISTing or GETting millions of objects in a loop, noncurrent versions piling up, or abandoned multipart parts.
  • Detection: Cost allocation by bucket tag; Storage Lens for noncurrent bytes and incomplete multipart bytes; request metrics by operation.
  • Fix: Restore the CDN path, kill the loop, add noncurrent expiry and abort-incomplete rules.
  • Prevention: Per-bucket budgets and anomaly alerts; lifecycle rules required at bucket creation.

4. Mass Deletion or Overwrite by Your Own Code#

  • Symptom: A deploy or script deletes or overwrites millions of objects; S3's durability is irrelevant because S3 did exactly what it was told.
  • Root cause: Bad prefix in a cleanup job, wrong environment credentials, a lifecycle rule with a too-broad filter.
  • Detection: DeleteObject request spikes; CloudTrail data events; delete-marker counts.
  • Fix: With versioning, remove delete markers or restore prior versions (batch operations at scale); without it, restore from the backup account.
  • Prevention: Versioning on production buckets, deny-delete policies for most roles, a separate-account backup vault, and lifecycle changes reviewed like code.

5. Public Exposure or Leaked Presigned URLs#

  • Symptom: Private files reachable by anyone; found by a researcher, a scanner or the press.
  • Root cause: Block Public Access disabled for a "quick fix," overly broad bucket policy, or long-lived presigned URLs shared outside the product.
  • Detection: Access Analyzer findings; config rules on bucket policies; unusual GET sources in access logs.
  • Fix: Re-enable Block Public Access; rotate the credentials that signed leaked URLs (revoking the signer invalidates its URLs); audit access logs for the exposure window.
  • Prevention: Account-level Block Public Access; short URL expiry; CDN signed URLs for repeat access; policy-as-code checks in CI.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
503 Slow Down5xx by prefix, client SlowDown countOne workload's writes or readsBackoff, throttle producer, spread keysWorkload team
Stuck PENDINGOldest PENDING age, DLQ depthNew uploads for all usersReplay DLQ, run reconcilerUpload service team
Surprise billCost anomaly by bucket tagBudget, not usersLifecycle rules, restore CDN pathBucket owner; platform for guardrails
Self-inflicted deletionDelete spikes, CloudTrailEvery object the job touchedRestore versions or from vaultOwning team; platform owns vault
Public exposureAccess Analyzer, config rulesEvery exposed object, plus trustRe-block, rotate signer, auditSecurity + bucket owner
Region outageElevated 5xx across bucketsAll buckets in the regionFail reads to replica region if designedPlatform + service owners

When to Use vs. Alternatives#

NeedPickWhy
User media, attachments, documentsS3 + CDN + metadata DBCheap, durable, direct upload and cacheable delivery
Analytics lakeS3 with Parquet and a table formatSeparates storage from compute; many engines read it
Backups and compliance archivesS3 with versioning, Object Lock, Glacier tiers, separate accountWORM and cheap long retention
Small mutable records, counters, sessionsA database or cache (DynamoDB, Redis)Per-request pricing and no CAS-heavy workloads in S3
Hot scratch data for training or shuffleExpress One Zone or local NVMeSingle-digit ms and high request rates
A process that needs POSIX semanticsBlock or file storageRenames, appends and locks are file-system features
Multi-PB with massive egress, own data centresSelf-hosted S3-compatible storage, with real headcountUnit economics at very large scale, as Dropbox did below

When NOT to Use S3#

  • As a database or queue. No queries, no transactions, per-request pricing, paginated LIST. Use a database for state and a queue for work.
  • For small mutable values. Rewriting a 2 KB object 100 times a second is a last-writer-wins race and a request bill.
  • For latency-critical reads without a cache. 100–200 ms first byte is fine for media and terrible for a request path with a 50 ms budget; check it against the Latency Budget.
  • As the only copy of irreplaceable data in one account. Durability protects against disks failing, not credentials leaking.
  • For workloads that need appends or renames. S3 has no append and no atomic rename; a "rename" is copy plus delete per object.

Operational Concerns#

What the On-Call Actually Does#

  1. Watches error rates by operation and prefix: 5xx (Slow Down), 4xx spikes (403 from expired URLs or policy changes), and client retry rates.
  2. Watches the async edges: queue depth and age for event consumers, oldest PENDING row, DLQ size.
  3. Reviews the cost dashboard weekly: storage by class, noncurrent bytes, incomplete multipart bytes, requests by operation, egress by bucket.
  4. Runs restores: from versions after bad deploys, from Glacier for support cases (with a stated restore SLA of hours, not minutes).
  5. Reviews bucket and lifecycle changes like code: a lifecycle filter is a scheduled mass deletion.

Key Metrics & Alerts#

MetricHealthyAlert
5xxErrors / AllRequests per bucket or prefix< 0.1%> 1% for 5 min
4xxErrors rateStable baseline3× baseline (policy or URL expiry regression)
FirstByteLatency p99 (request metrics)Stable, ~100–200 ms small objects2× baseline for 15 min
Oldest PENDING upload age< 5 min> 30 min (page)
Event queue ApproximateAgeOfOldestMessage< 60 s> 10 min
Noncurrent version bytes / current bytes< 20%> 50% (ticket)
Incomplete multipart bytes~0Any growth week over week
Daily cost per bucketWithin forecast+30% day over day

Interview Application — Staff-Level Plays#

Which Case Studies Use S3#

Case StudyHow Object Storage Is UsedKey Pattern
Design an Object StoreBuild the thing itselfMetadata/data split, erasure coding, repair
Video StreamingOriginals and renditions behind a CDNEvent-driven transcode pipeline; egress via CDN
Cloud File SyncContent-addressed chunksDedupe with create-once writes; metadata DB as truth
CDNOrigin for static and media contentImmutable keys, origin shielding, origin access control
Web CrawlerRaw page store, WARC-style batchesBatch small pages into large objects
Metrics & MonitoringLong-term block storage for compressed seriesHot data local, cold blocks in object storage
Ad Click AggregationRaw event archive for reprocessingImmutable raw log enables replay and audits

Every System Design Question Has an S3 Moment#

  • Chat app: "Attachments upload directly to S3 with a presigned URL scoped to one key and 10 minutes; the message stores the object ID, and recipients get a short-lived signed URL after an access check. See Chat."
  • URL shortener: "S3 holds nothing on the hot path. It holds nightly exports of the mapping table for analytics and disaster recovery."
  • News feed: "Images are content-addressed in S3 behind the CDN with year-long cache headers. The feed service never touches bytes."
  • Payments or ledger: "Statements and audit exports go to a separate-account bucket with Object Lock in compliance mode for the 7-year retention legal signed off on. See Ledger & Wallet."

What Interviewers Probe#

After You Say...They Will Ask...What They're Evaluating
"Store files in S3""How does the upload work for a 3 GB file on mobile?"Presigned multipart, resumability, bytes off your servers
"S3 is eventually consistent""Are you sure? What about two writers?"Knowing the 2020 change and what remains: LWW, no cross-key atomicity
"Trigger processing on upload events""What if the event never arrives, or arrives twice?"Idempotency and reconciliation
"S3 scales infinitely""What happens when launch traffic hits one prefix?"Per-prefix rates, gradual scaling, 503s, key design
"Lifecycle to Glacier""What does that cost for a billion 10 KB objects?"Minimum size, minimum duration, transition fees
"We have 11 nines""What if someone runs the wrong delete script?"Versioning, Object Lock, separate-account backups

Common Interview Mistakes#

What Candidates SayWhat Interviewers HearWhat Staff Engineers Say
"Upload goes through our API to S3"Will melt the API fleet on large files"Presigned multipart direct to S3; the API issues permission and records metadata."
"List the folder to show the user's files"Using S3 as a database"The metadata DB answers listings; S3 only holds bytes."
"S3 is eventually consistent so add a delay"Five years out of date"Strong per key since 2020; I guard concurrent writers with conditional writes."
"Archive everything to Glacier to save money"Hasn't modelled small-object fees"Compact small objects first; tier only what's large and cold."
"Durability is 11 nines, so backups aren't needed"Confuses hardware failure with human error"Versioning plus a separate-account vault with Object Lock."
"Serve images straight from the bucket"Hasn't priced egress"CDN in front with immutable keys; origin egress drops ~10×."

L5 vs L6 vs L7 Responses#

ScenarioL5 AnswerL6 / Staff AnswerL7 / Principal Answer
"Store user photos"S3 bucket, URL in DBPresigned multipart, tenant-led keys, PENDING→READY with reconciler, CDN, lifecycleRetention policy by data class; chargeback per product; egress architecture across products
"S3 bill tripled"Move to GlacierBreak down storage vs requests vs egress; noncurrent versions, multipart orphans, CDN regressionPer-bucket budgets and anomaly alerts org-wide; lifecycle rules mandatory at creation
"Protect against ransomware"Enable versioningSeparate-account vault, Object Lock, deny-delete roles, tested restoreDecides governance vs compliance mode with legal; restore drills as a quarterly game day
"Leave AWS for cost?"Run MinIOCompare total cost including egress, ops headcount and durability engineeringTreats it as a multi-year one-way door with a staged exit and a break-even point

The Staff S3 Checklist#

  1. Classify objects: "Originals, derived, logs, backups: each gets its own bucket or prefix, class and retention."
  2. Key layout: "Tenant or shard first, date next, UUID or hash last; immutable keys, no overwrites."
  3. Upload path: "Presigned multipart, 16 MiB parts, 15-minute expiry, If-None-Match on complete."
  4. Truth and events: "Metadata DB is the source of truth; events are at-least-once hints; reconciler hourly."
  5. Cost: "Storage, requests and egress modelled separately; CDN in front; lifecycle with abort-incomplete and noncurrent expiry."
  6. Protection: "Versioning, Block Public Access, separate-account vault, Object Lock where legal requires it."

🎯 Staff Insight: Don't use S3 as a database, a queue, a lock service or a low-latency key-value store. The strongest S3 signal is drawing the line between bytes (S3) and truth (the metadata database), and saying what happens when the event that connects them is lost.

Evaluation Rubric#

DimensionSenior (L5)Staff (L6)Principal (L7)
Consistency"Eventually consistent" or no mentionStrong per key; LWW; conditional writes; no cross-key atomicityOrg pattern: immutable content-addressed objects, metadata DB as truth
Throughput"Infinite"Per-prefix rates, gradual scaling, key design, backoffRecognises batching and caching as the real fix
CostStorage price onlyStorage + requests + egress + lifecycle trapsPortfolio view, chargeback, egress architecture
Protection"11 nines"Versioning, replication, Object Lock, restore pathData-loss posture and restore drills across the org
Operations"Monitor the bucket"PENDING age, DLQ, cost anomalies, lifecycle reviewGuardrails at bucket creation enforced by the platform

Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

At Staff level S3 is a well-configured bucket. At Principal level object storage is the company's largest, slowest-growing and least-watched cost line, and the place where retention policy, legal exposure and disaster recovery meet. Data accumulates for years because deleting it needs someone to say it's safe. The L7 question is not "which storage class?" but "what data do we keep, for how long, at what cost, who can delete it, and how fast can we get it back?"

🧭 Principal Move: "Before we tune buckets, I want a data classification with an owner and a retention period for each class. Lifecycle rules are just that policy written in code; without the policy, every bucket keeps everything forever and nobody is allowed to delete anything."

The Org-Level Fault Line#

Central storage platform vs every team owning its buckets.

OptionWhat WorksWhat BreaksWho Pays
Every team creates buckets freelyFast; no gatekeepingInconsistent encryption, public-access mistakes, no lifecycle, unowned spendSecurity and finance
Central team owns all bucketsUniform policyBottleneck; teams wait for buckets; platform can't judge retentionProduct velocity
Platform guardrails, team ownershipTeams self-serve through a template that enforces encryption, Block Public Access, tags and default lifecycleNeeds policy-as-code and exception handlingPlatform headcount (small)
Separate backup and compliance accountsBlast radius containedExtra accounts and replication costPlatform; legal signs retention

The Principal default: guardrails, not gatekeeping. Teams create buckets through a template; account-level Block Public Access and encryption are non-negotiable; tags for owner, data class and cost centre are required; default lifecycle rules ship with the template; backups and compliance data live in separate accounts owned by the platform and security.

🧭 Principal Insight: "The expensive storage isn't what we read. It's what nobody remembers writing. Required owner and data-class tags are how a deletion decision becomes possible three years from now."

Cost Model#

Assumptions: Standard ~$0.021–0.023/GB-month at volume, Standard-IA ~$0.0125, Glacier Instant Retrieval ~$0.004; egress mostly through a CDN; loaded engineer ~$250K/year. Directional only; check current pricing.

ScaleStoredMixStorage/monthRequests + egress/monthHeadcountTotal/month
Startup50 TBAll Standard~$1.1K~$2–5K0.1 FTE~$5–8K
Growth2 PB40% Standard, 40% IA, 20% Glacier IR~$28K~$40–80K0.5 FTE (~$10K)~$80–120K
Enterprise50 PB20% Standard, 30% IA, 50% archive tiers~$400K~$0.5–1.5M3–5 FTE (~$85K)~$1–2M

Two observations. At every scale beyond startup, requests and egress rival or exceed storage, so CDN hit ratio and batching are worth more than storage-class tuning. And at enterprise scale, a retention policy that deletes data nobody needs is the single largest lever; moving the 30% of data with no reader from Standard to deletion beats any tiering scheme.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to Reverse
Object Lock in compliance mode with a long retentionOne-way for the retention periodYou pay for the data until the date; only deleting the account removes it
Key layout baked into clients and URLsOne-way-ishCopy every object to new keys and migrate references
Deleting data without versioning or backupOne-wayGone
Vendor-specific features in core logic (event formats, query features)One-way-ishRewrite when moving providers
Storage class choiceTwo-wayTransition fees and minimum-duration charges
Lifecycle rulesTwo-way going forwardAlready-expired objects stay expired
CDN in front of S3Two-wayConfiguration and cache warm-up

The Standard I'd Write#

RFC-STOR-002: Object Storage Baseline

Scope: Every object storage bucket in every production account.

MUST
  1. Be created from the platform template: account-level Block Public Access,
     default encryption, required tags (owner, data_class, cost_center).
  2. Enable versioning unless data_class is "regenerable".
  3. Include lifecycle rules: abort incomplete multipart uploads after 7 days;
     expire noncurrent versions after 30 days (or the data class retention).
  4. Keep business metadata and access decisions in a database, not in key listings.
  5. Make every event consumer idempotent, with a reconciler that runs at least daily.
  6. Use presigned URLs of at most 1 hour for user access; no long-lived keys on clients.
SHOULD
  7. Front public and semi-public reads with a CDN using origin access control.
  8. Batch objects smaller than 128 KB before transitioning to IA or archive classes.
  9. Replicate regulated or irreplaceable data to the backup account with Object Lock.

Exceptions: platform + security approval, time-boxed to two quarters, listed on a dashboard.
Enforcement: policy checks in CI and account config rules; warn for one quarter, then block.
Success metrics: zero public-bucket findings; incomplete-multipart bytes ~0;
  100% of buckets tagged; restore drill for each regulated dataset twice a year.

What I'd Tell the VP#

"Object storage is cheap per gigabyte, which is exactly why it has become one of our larger and fastest-growing bills: we keep everything forever and serve some of it inefficiently. About half of what we pay is not storage at all but requests and data leaving our cloud. I'm proposing three things: put a content delivery network in front of everything public, give every dataset an owner and a retention period so we can delete what nobody needs, and keep protected backups in a separate account so a mistake or an attacker can't erase our data. That should cut the storage bill by a third within two quarters and close our largest data-loss risk. The cost is roughly half an engineer to build the template and policies."

Principal Interview Signals#

SignalWhat It Sounds Like
Retention as policy"Lifecycle rules encode a retention policy legal signed; no policy, no deletion."
Prices the whole bill"Storage is the small line; requests and egress are where the money goes."
Designs the loss posture"Separate-account vault, Object Lock, restore drills: durability doesn't cover us."
Guardrails over gatekeeping"Teams self-serve buckets from a template that makes the safe thing the default."
Respects one-way doors"Compliance mode is irreversible, so it needs legal sign-off and a test bucket first."

Staff answers that L7 interviewers find insufficient:

  • "Add lifecycle rules" without who decides retention per data class.
  • "Enable versioning" without noncurrent expiry or the cost of doubling storage under churn.
  • "Use Glacier" without restore times, retrieval fees and the small-object trap.

How Real Companies Built It#

Amazon S3 — Strong Consistency Without a Price Increase#

In December 2020 S3 moved from eventual to strong read-after-write consistency for every new and existing object. Werner Vogels described how: S3 added an in-memory "witness" component that tracks only enough metadata to act as a read barrier, so the metadata cache can learn whether its view of an object is stale. The stated goal was strong consistency "with no additional cost" and "no performance or availability tradeoffs," and the team leaned on formal methods and model checking to verify the protocol (All Things Distributed, 2021).

Staff insight: The guarantee changed underneath millions of applications without an API change. In an interview, that is why "S3 is eventually consistent" is a red flag today, and why the remaining gaps (concurrent writers, cross-key updates, bucket configuration) are what you should name.

Amazon S3 — Heat Management Across Millions of Drives#

Andy Warfield's 2023 essay on operating S3 explains that the hard problem is "heat," request load on individual disks. S3 spreads new objects broadly across its disk fleet, deliberately placing different objects on different sets of disks, so that at its scale no single workload can meaningfully move the aggregate peak and a single customer can burst across a very large number of drives. The essay also describes ShardStore, S3's rewritten storage node software, verified with lightweight formal methods in Rust (All Things Distributed, 2023).

Staff insight: Multi-tenancy is the performance feature: aggregation smooths bursty workloads. When you design your own store, as in Design an Object Store, placement and load spreading matter as much as replication.

Dropbox — Magic Pocket, Leaving S3 at 500 PB#

Dropbox built Magic Pocket, its own block storage system, and by October 2015 had moved over 90% of its users' data off S3 onto it, with more than 500 PB of user data under management by early 2016. Dropbox cited end-to-end control of the stack for performance and better unit economics from customising hardware and software for its specific workload, and reported completing the migration without major disruptions or data loss (Dropbox engineering).

Staff insight: Leaving managed object storage is rational only at very large scale with a stable, well-understood workload and a team able to own durability. In an interview, the Staff answer to "should we build our own?" names that threshold and the multi-year engineering cost before saying yes.


Practice Drill#

Prompt: "Your photo-sharing product stores 800 TB in one S3 bucket. In the last quarter the S3 bill grew 2.5× while traffic grew 30%. A new 'burst upload' feature launched last week and uploads fail with 503s for the first hour of every evening peak. Some users also report photos stuck on 'processing' for a day. You own storage. What do you do?"

Staff Answer

Three symptoms, three different mechanisms, so I'd split them. 503s at peak: the burst feature probably writes keys that start with a timestamp or a single new prefix, so all evening writes land on one key range faster than S3 repartitions. I'd confirm with request metrics filtered by prefix, then change new keys to lead with a tenant shard (media/{shard 00-ff}/{date}/{uuid}), make sure the client retries 503s with exponential backoff and jitter under a retry budget, and ramp load on the new layout before the next peak. If each burst is 30 small photos, I'd also batch thumbnail writes. Stuck processing: events are at-least-once and the consumer or its DLQ is the likely gap. I'd alert on oldest PENDING age and DLQ depth, replay the DLQ, and add an hourly reconciler that HEADs stuck keys and advances or fails them, with handlers idempotent on (key, etag). The bill: 2.5× cost on 1.3× traffic means cost per request or per byte rose. I'd break spend into storage by class, noncurrent versions, incomplete multipart bytes, requests by operation and egress. Likely culprits: versioning without noncurrent expiry, abandoned multipart parts from failed burst uploads, and reads bypassing the CDN for the new feature. Fixes: abort-incomplete after 7 days, noncurrent expiry at 30 days, CDN with origin access control on all image reads, then lifecycle to Standard-IA at 30 days for originals over 128 KB. I'd target a 40–50% cost reduction, zero 503s at peak and PENDING age under 5 minutes, and I'd report those three numbers weekly.

Why this is L6:

  • Separates three symptoms into three mechanisms (key range heat, event loss, cost per unit) instead of one blanket fix.
  • Uses the current consistency and event model correctly: at-least-once events need idempotency plus a reconciler.
  • Decomposes cost into storage, requests, egress and invisible bytes before reaching for storage classes.

What L7 adds:

  • Turns the incident into guardrails: a bucket template with mandatory lifecycle rules, prefix-design review for new features and per-bucket cost anomaly alerts.
  • Asks for a retention policy per data class so derived renditions are regenerated or expired instead of stored forever.
  • Prices the CDN and the reconciler against the savings and reports the result to finance as cost per active user.
❌ Common L5 Trap

"Move older photos to Glacier to cut the bill, and increase the client retry count so the 503s go away."

Why this misses: Glacier tiers add retrieval delays and per-object fees without touching the actual growth drivers (noncurrent versions, orphaned parts, egress), and more retries amplify the 503s on an already hot key range. It also ignores the stuck uploads entirely, which need idempotent consumers and a reconciler, not cheaper storage.


Quick Reference Card#

Model:             flat key -> immutable object; "folders" are key prefixes
Consistency:       strong read-after-write per key (PUT, DELETE, LIST) since Dec 2020
                   concurrent writers: last-writer-wins; no cross-key atomicity
                   bucket config: eventually consistent (~15 min after enabling versioning)
Conditional:       If-None-Match: * (create once), If-Match: etag (CAS) on PUT/complete/copy
Request rates:     >= 3,500 writes, 5,500 reads per second per partitioned prefix;
                   scales gradually, 503 Slow Down meanwhile; no prefix limit
Latency:           ~100-200 ms first byte (small objects); Express One Zone single-digit ms
Single PUT:        up to 5 GB; use multipart above ~100 MB
Multipart:         5 MiB-5 GiB parts, 10,000 parts; max object 48.8 TiB (~50 TB, Dec 2025)
Classes:           Standard | Intelligent-Tiering | Standard-IA (30d, 128 KB min)
                   One Zone-IA (1 AZ) | Express One Zone (1 AZ)
                   Glacier IR (90d) | Flexible (90d, restore) | Deep Archive (180d, hours)
Durability:        11 nines designed, every class; not protection from your own deletes
Events:            at least once; seconds, sometimes a minute+; no direct FIFO SQS
Presigned URLs:    up to 7 days (IAM user, SigV4); role/STS creds expire sooner
Object Lock:       needs versioning; governance (bypassable) vs compliance (nobody, not root)
Pricing (approx):  Standard ~$0.023/GB-mo; PUT $0.005/1K; GET $0.0004/1K; egress ~$0.09/GB

RED FLAGS
  - Bytes proxied through application servers
  - LIST used to answer user-facing queries
  - Monotonic timestamp at the start of hot keys
  - No abort-incomplete-multipart or noncurrent-version lifecycle rule
  - Tiny objects tiered to IA or Glacier
  - Event handlers without idempotency or a reconciler
  - Multi-day presigned URLs handed to users
  - Production buckets without versioning or a separate-account backup
  1. Loading the index…