Go deeper:
- To build an object store yourself, see Design an Object Store.
- For the upload and delivery path in front of it, see Large File Uploads & Delivery.
Why This Matters#
Object storage is not a file system and it is not "infinite disk." It is a key-value store for immutable blobs with a billing model attached to every verb. You pay per gigabyte stored, per request made, per gigabyte that leaves the region, per object transitioned between tiers, and per day an object sits in a tier before its minimum duration ends. Almost every production incident with S3 is one of three things: a request pattern the bucket hasn't scaled for yet (503 Slow Down), a correctness assumption the API never made (two writers to one key, an event that arrived twice), or a bill nobody modelled (egress, small objects in archive tiers, millions of abandoned multipart parts).
That is why "and we put the files in S3" is the sentence interviewers lean on. The L5 candidate draws a bucket and moves on. The L6 candidate says "clients upload directly with a presigned multipart URL that expires in 15 minutes; keys are tenant/yyyy/mm/dd/uuid so writes spread over prefixes; the metadata row is written first and flipped to READY by an at-least-once event handler that is idempotent on (key, version); lifecycle moves objects to Standard-IA at 30 days, aborts incomplete multipart uploads after 7, and expires noncurrent versions after 30; delivery goes through a CDN so origin egress stays under 10% of bytes served." The L7 candidate asks which of the company's 4 PB actually gets read, what egress will cost at 10× traffic, and whether the compliance requirement means Object Lock in compliance mode, which nobody, including the root account, can undo.
The L5 → L6 gap is not knowing what a bucket is. It is knowing that S3 gives you strong per-key consistency and nothing else: no cross-key transactions, no locks, no ordering of events, no free requests. Every guarantee above the single key is yours to build.
The L5 → L6 → L7 Contrast#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Store files in S3, URL in the database" | "What's the object size distribution, read/write ratio and retention? That decides key layout, upload path, storage class and whether a CDN fronts it." | "Which data is the system of record, which is derived, and what does each cost per year at 3 scales? Derived data shouldn't be stored for 7 years." |
| Consistency | "S3 is eventually consistent" (outdated) or "S3 is consistent, so we're fine" | Strong read-after-write per key, including LIST; concurrent writers are last-writer-wins; conditional writes (If-None-Match, If-Match) for create-once and compare-and-swap; no cross-key atomicity | Sets the org pattern: metadata in a database as the source of truth, objects immutable and content-addressed so consistency questions mostly disappear |
| Throughput | "S3 scales infinitely" | 3,500 writes and 5,500 reads per second per partitioned prefix, scaling gradually with 503s on the way; spread keys, back off with jitter, pre-warm for known launches | Decides when a workload has outgrown general-purpose buckets (single-AZ low-latency class, caching tier, or a different store) |
| Cost | "Storage is cheap" | Models storage + requests + egress + transitions; lifecycle rules for IA/Glacier with minimum durations; abort incomplete multipart uploads | Owns the storage bill as a portfolio: chargeback by bucket tag, egress architecture (CDN, same-region compute), retention policy signed by legal |
| Failure | "S3 has 11 nines of durability" | Durability ≠ protection from your own deletes: versioning, replication, Object Lock, and event-loss reconciliation | Designs the org's data-loss posture: which buckets get WORM, cross-account backups, and a tested restore time |
| Ownership | "The app team owns the bucket" | Bucket owner, lifecycle owner and event-consumer owner named; orphan sweeps scheduled | Storage platform owns guardrails (Block Public Access, encryption, tagging); product teams own retention and access patterns |
Why "Consistency" separates levels
Many candidates still recite "S3 is eventually consistent." That has been false since December 2020: S3 now provides strong read-after-write consistency for PUTs and DELETEs of objects in all regions, and a LIST issued after a successful write includes the new key (S3 consistency model). The Staff answer knows what that guarantee does not cover. Two simultaneous PUTs to the same key resolve as last-writer-wins by timestamp; there is no lock. Updates are per key, so "write the data file, then write the manifest" can be observed half-done. Bucket configuration (versioning, lifecycle, policies) is still eventually consistent; AWS recommends waiting 15 minutes after first enabling versioning before writing. Event notifications are delivered at least once and can take a minute or longer. The Staff move is to use conditional writes for create-once semantics and to keep cross-object invariants in a database, not in S3.
Why "Throughput" separates levels
"S3 scales infinitely" is true in aggregate and false at the moment you need it. AWS documents at least 3,500 PUT/COPY/POST/DELETE or 5,500 GET/HEAD requests per second per partitioned prefix, with no limit on the number of prefixes, and states that scaling to higher rates "happens gradually and is not instantaneous," with 503 Slow Down errors while it does (S3 performance). A launch that goes from 200 to 20,000 writes per second on one new prefix will see 503s for minutes. The Staff answer spreads keys, retries 503s with exponential backoff and jitter, and pre-scales known events with a load ramp. The Principal answer asks whether the request rate itself is the design smell: 20,000 tiny PUTs per second is a batching problem before it is a storage problem.
The 60-Second Pitch#
"User media goes in S3 as immutable objects keyed media/{tenant}/{yyyy}/{mm}/{uuid}, so writes spread over many prefixes and no key is ever overwritten. Clients upload directly with presigned multipart URLs, 16 MiB parts, 15-minute expiry; our API only issues permission and records metadata. A PENDING row is written first; an S3 event into SQS flips it to READY after a HEAD check, and an hourly reconciler catches lost events. Reads go through a CDN with year-long cache headers because keys are content-addressed, which keeps origin egress small. Lifecycle moves objects to Standard-IA at 30 days and Glacier Instant Retrieval at 90, aborts incomplete multipart uploads after 7 days, and versioning plus replication to a second account protects against our own mistakes. S3 gives us 11 nines of designed durability and strong per-key consistency; everything across keys, including 'which objects belong to this album', lives in the database."
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| User content store (photos, video, attachments) | Many writers, read-heavy, CDN-fronted, per-user access | Presigned direct upload, immutable keys, metadata DB as truth, CDN delivery | Orphaned objects, lost events leave rows PENDING, egress bill | No broken links; no cross-tenant reads |
| Data lake / analytics | Huge objects, scan-heavy, many parallel readers | Columnar files (Parquet) of 128 MiB–1 GiB, table format manifests, partitioned prefixes | Small-file explosion; LIST-heavy planners; 503s on hot partitions | Readers see complete snapshots only |
| Backup, archive and compliance | Rarely read, retained for years, must survive deletion | Versioning, Object Lock, cross-account replication, Glacier tiers | Minimum-duration and per-object fees on tiny files; slow restores; accidental compliance-mode lock | Restorable within the stated RTO; immutable for the retention period |
🎯 Staff Move: "I'll design for the user content store. The analytics lake and the compliance archive get their own buckets and policies, because their lifecycle rules, key layouts and access patterns conflict. One bucket serving all three is how a lifecycle rule meant for logs ends up archiving profile photos."
The Staff Positions#
| Position | Rationale |
|---|---|
| Metadata lives in a database; S3 holds bytes | S3 has per-key atomicity only. "Which files are in this folder" and "who can see this" need transactions and indexes. |
| Objects are immutable; new content gets a new key | No overwrite races, perfect CDN caching, versioning stays cheap. |
| Bytes never transit application servers | Presigned URLs move upload and download traffic off your fleet; a 2 GB proxy upload pins a connection for minutes. |
| Every event consumer is idempotent and backed by a reconciler | Notifications are at-least-once and can be delayed; a sweep is the only way to be sure. |
| Lifecycle rules ship with the bucket | Abort incomplete multipart uploads, expire noncurrent versions, tier cold data. A bucket without them grows a bill nobody owns. |
| Egress is designed, not discovered | CDN in front, compute in the same region, cross-region replication only where the RPO demands it. |
| Block Public Access on, encryption on, by default | Public buckets are a security incident waiting for a headline; exceptions go through review. |
Architecture & Internals#
S3's internals are not public in detail, and you don't need them. Five properties change design decisions: the key namespace and partitioning, the consistency model, the multipart protocol, storage classes, and the event and lifecycle machinery. For how such a system is built underneath (erasure coding, placement, repair), see Design an Object Store.
The Namespace: Flat Keys, Partitioned by Prefix#
A bucket is a flat map from key to object. / has no meaning to S3; "folders" are a console convention over shared key prefixes. Internally the index is partitioned by key range, and S3 splits hot ranges as load grows. That is where the documented per-prefix rates come from.
Why it matters in design: throughput is bounded per key range, not per bucket. Keys that all start with the same timestamp or the same tenant ID concentrate writes on one range until S3 splits it, and splitting takes time. The modern guidance is not "add random hashes to every key" (that advice predates the 2018 rate increase), but "make sure the leading characters that vary actually vary for the hot workload." A {tenant}/{date}/{uuid} layout spreads load across tenants; {date}/{tenant}/{uuid} sends the whole day's writes to one prefix.
Consistency: Strong Per Key, Nothing Across Keys#
| Operation | Guarantee today | What it doesn't give you |
|---|---|---|
| PUT new key, then GET | Returns the new object | — |
| PUT overwrite, then GET | Returns the new data, never partial | Protection from a concurrent writer (last timestamp wins) |
| DELETE, then GET or LIST | Object gone from both | — |
| PUT, then LIST | New key appears | Snapshot isolation across a multi-page LIST of a changing prefix |
| Two keys updated "together" | Each atomic on its own | Any atomicity across them |
| Bucket config change | Eventually consistent | Immediate effect (wait ~15 min after first enabling versioning) |
If-None-Match: * on PUT | Create only if absent; fails if the key exists | Locking for long operations |
If-Match: <etag> on PUT | Compare-and-swap on the current ETag | Multi-key transactions |
Conditional writes are supported on PutObject, CompleteMultipartUpload and CopyObject, and conditional deletes on DeleteObject, with no extra charge beyond normal request rates (conditional requests). That turns S3 into a usable primitive for "first writer wins" (idempotent uploads, leader-written manifests) without an external lock.
🎯 Staff Insight: "S3 is strongly consistent per key, so the classic 'read your write' bug is gone. What's left is the multi-key problem: a reader can see the new data file before the new manifest. I keep the manifest pointer in a database or use a conditional write on a single manifest key, so readers flip from one complete snapshot to the next."
Multipart Upload: How Large Objects Actually Move#
| Limit | Value | Design consequence |
|---|---|---|
| Single PUT | Up to 5 GB | Above ~100 MB, use multipart anyway for retries and parallelism |
| Part size | 5 MiB to 5 GiB (last part may be smaller) | 16–64 MiB parts are a good default on mobile and desktop links |
| Parts per upload | 10,000 | Part size × 10,000 caps object size: 5 MiB parts top out near 48.8 GiB |
| Maximum object size | 48.8 TiB (~50 TB), raised from 5 TB in December 2025 | Pick part size from the largest object you accept: ~5 GiB parts for the max |
| Incomplete uploads | Billed as storage until completed or aborted | AbortIncompleteMultipartUpload lifecycle rule, 1–7 days |
Sources: multipart limits, 50 TB announcement. The object becomes visible atomically at CompleteMultipartUpload; readers never see a half-assembled object.
Storage Classes and Lifecycle#
| Class | Designed availability | AZs | Minimum duration | Minimum billable size | Access |
|---|---|---|---|---|---|
| Standard | 99.99% | ≥ 3 | None | None | Milliseconds |
| Intelligent-Tiering | 99.9% | ≥ 3 | None | Objects < 128 KB never tier down | Milliseconds (archive tiers opt-in) |
| Standard-IA | 99.9% | ≥ 3 | 30 days | 128 KB | Milliseconds, per-GB retrieval fee |
| One Zone-IA | 99.5% | 1 | 30 days | 128 KB | Milliseconds; lost with the AZ |
| Express One Zone | 99.95% | 1 | None | None | Single-digit milliseconds |
| Glacier Instant Retrieval | 99.9% | ≥ 3 | 90 days | 128 KB | Milliseconds, higher retrieval fee |
| Glacier Flexible Retrieval | 99.99% after restore | ≥ 3 | 90 days | 40 KB metadata overhead per object | Minutes to hours, restore first |
| Glacier Deep Archive | 99.99% after restore | ≥ 3 | 180 days | 40 KB metadata overhead per object | Hours, restore first |
Source: storage class comparison. Every class is designed for 11 nines of durability. What changes is availability, the AZ count, the minimum duration you pay for even if you delete early, the minimum billable size, and the retrieval path.
The small-object trap: a 10 KB log line object moved to Standard-IA is billed as 128 KB, 12.8× its size, and every lifecycle transition is a per-object request fee. Archiving a billion 10 KB objects can cost more in transition requests and minimum-size padding than it saves. Batch small objects into larger archives (tar or Parquet, 100 MB+) before tiering them.
Events, Versioning and Object Lock#
- Event notifications go to SNS, SQS, Lambda or EventBridge. AWS states they are "designed to be delivered at least once," typically in seconds but sometimes a minute or longer, and SQS FIFO queues are not a direct destination (event notifications). Treat them as hints that start work, not as a ledger.
- Versioning keeps every overwritten or deleted version; a simple DELETE adds a delete marker. It is the cheapest protection against application bugs and the most common source of a mysteriously growing bill (noncurrent versions are billed like any object).
- Object Lock (requires versioning) makes versions write-once-read-many. Governance mode can be bypassed by principals with
s3:BypassGovernanceRetention; compliance mode cannot be shortened or removed by anyone, including the root user, until the retention date (Object Lock). Legal holds have no expiry until removed. - Presigned URLs are bearer tokens scoped to one operation on one key. With SigV4 and long-lived IAM user credentials they last up to 7 days; signed with role or STS credentials they die when the session does, often within 1–6 hours. Expiry is checked when the request starts, so a long download already in progress completes (presigned URLs).
Data Modeling / Core Usage — "The Entire Game"#
In DynamoDB the game is the partition key. In S3 it is the object contract: the key layout, the immutability rule, and the split between what lives in S3 and what lives in the metadata database. Get these right and the bucket scales, caches and ages out on its own.
Step 1: Classify the Objects#
| Class of object | Typical size | Read pattern | Mutability | Home |
|---|---|---|---|---|
| User originals (photos, video, docs) | 100 KB–5 GB | Read-heavy for weeks, then cold | Immutable | Standard → IA → Glacier IR |
| Derived renditions (thumbnails, transcodes) | 5 KB–500 MB | Very hot, CDN-cached | Regenerable | Standard; expire or regenerate instead of archiving |
| Analytics files (Parquet) | 128 MiB–1 GiB | Scanned in parallel | Immutable, compacted | Standard or Intelligent-Tiering |
| Logs and events | 1 KB–10 MB raw | Rarely read after 7 days | Append via new objects | Batch into large objects, then archive |
| Backups and compliance records | MBs–TBs | Almost never | WORM | Separate account, Object Lock, Glacier |
Step 2: Design the Key#
Good: media/{tenant_id}/{yyyy}/{mm}/{dd}/{uuid}.{ext}
- leading variation across tenants spreads writes across prefixes
- date components make lifecycle rules and audits easy
- uuid (or content hash) makes keys immutable and unguessable
Content-addressed variant for dedupe and caching:
blobs/{sha256[0:2]}/{sha256} -> same bytes, same key, written once
write with If-None-Match: * -> a duplicate upload is a no-op (412)
Bad: {yyyy-mm-dd-hh-mm-ss}-{userid}.jpg -> every write hits the newest range
Bad: users/{email}/photo.jpg -> PII in keys, logs and URLs; overwrites
Bad: one key rewritten every second -> last-writer-wins races, version bloat
Step 3: Put Cross-Object Truth in the Database#
table objects (
object_id uuid primary key,
owner_id bigint not null,
s3_key text unique not null,
sha256 bytea,
size_bytes bigint,
status text check (status in ('PENDING','UPLOADED','READY','FAILED','DELETED')),
created_at timestamptz,
version int -- bumped on every state change; events compare it
);
Upload: insert PENDING -> issue presigned URL -> client uploads
Confirm: S3 event or client "complete" -> HEAD the key -> conditional update
UPDATE objects SET status='UPLOADED', version=version+1
WHERE object_id=$1 AND status='PENDING' -- duplicates are no-ops
Sweep: hourly: PENDING older than 24h -> HEAD; present -> UPLOADED, absent -> FAILED
Delete: mark DELETED in DB first; delete or expire the object asynchronously
The database answers every question that spans objects: listings, permissions, quotas, search. S3 LIST is a sequential, 1,000-keys-per-page scan billed at PUT-tier request rates; it is an inventory tool, not a query engine. For bulk audits use S3 Inventory reports instead of LIST loops. The full upload state machine is in Large File Uploads & Delivery; the metadata store choice is covered in PostgreSQL and DynamoDB.
🎯 Staff Move: "The database row exists before the bytes do, and the object is only visible to users once the row says READY. That way a lost event or a half-finished upload is a stuck row my reconciler can see, not a broken link a user finds."
The Tunable Tradeoff — Cost × Latency × Protection#
Every object storage decision trades three things: what you pay per gigabyte-month, how fast and cheaply you can read it back, and how well it survives mistakes.
| Setting | Cheap end | Protected / fast end | Who pays at the cheap end |
|---|---|---|---|
| Storage class | Deep Archive (~1/20 the price of Standard) | Standard or Express One Zone | Whoever needs the data back in under 12 hours |
| AZ redundancy | One Zone-IA, Express One Zone | Multi-AZ classes | The team whose only copy was in the failed AZ |
| Versioning | Off | On, with noncurrent expiration | Anyone whose bug overwrote a million objects |
| Replication | None | Cross-region, cross-account | The business during a regional or account-level incident |
| Object Lock | None | Compliance mode | Security, after ransomware or a malicious insider |
| Read path | Direct from origin | CDN in front | Finance, through egress |
Monthly cost model (list prices, US East, illustrative):
storage = GB_stored x price_per_GB (Standard ~$0.023/GB first 50 TB)
requests = PUTs/1,000 x $0.005 + GETs/1,000 x $0.0004
egress = GB_out_to_internet x ~$0.09/GB (first 100 GB/month free, tiers drop with volume)
extras = transitions + retrievals + minimum-duration and minimum-size padding
Example: photo app, 500 TB stored, 2 B PUTs/month, 40 B GETs/month, 3 PB served
storage 500,000 GB x ~$0.022 ~= $11K
PUTs 2e9 / 1,000 x $0.005 = $10K
GETs 40e9 / 1,000 x $0.0004 = $16K
egress 3 PB direct at ~$0.05 blended ~= $150K <- the real bill
with CDN at 95% hit ratio: origin egress ~150 TB; CDN pricing replaces S3 egress
The lesson candidates miss: at serving scale, egress and requests dwarf storage. Price checks come from the S3 pricing page; plug your own numbers into the Cost Estimator and the Back-of-Envelope Calculator.
🎯 Staff Move: "Before I pick storage classes, I'll split the bill into storage, requests and egress at our projected traffic. For a media product egress is usually 5–10× storage, so the CDN and the key design that makes caching work are the cost decisions; the storage class is a rounding error until the data is cold."
Who Pays for Each Choice#
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Everything in Standard forever | Simple, fast | Bill grows linearly with history | Finance, quietly, every month |
| Aggressive lifecycle to Glacier | Storage cost drops ~80–95% | Restores take hours; minimum duration and per-object fees | Support and users when old files are needed |
| One Zone classes | ~20% cheaper than multi-AZ IA | Data lost with an AZ | Whoever assumed it was regenerable |
| Versioning without noncurrent expiry | Undo for every mistake | Storage silently doubles under churn | The bucket owner at quarter end |
| Direct-from-S3 serving | No CDN to run | Egress and request bills; latency 100–200 ms first byte | Finance and users |
| Compliance-mode Object Lock | Nobody can delete, including attackers | Nobody can delete, including you | Legal and finance if set wrong |
Anti-Patterns — What Kills S3 Deployments#
1. Using S3 as a Database#
Listing a prefix to find "all files for user X," storing JSON state that several workers rewrite, or polling LIST for new work. LIST is paginated at 1,000 keys, billed like PUTs, and concurrent writers to one key silently lose updates. Fix: metadata and state in a database; S3 holds immutable bytes.
2. Proxying Bytes Through the Application#
A 2 GB upload through an API server pins a connection and memory for minutes and doubles egress. Fix: presigned direct upload and download; the app handles permission and metadata only.
3. Monotonic Key Prefixes on a Hot Write Path#
2026-10-06T14:00:00Z-... as the key start puts every write on the newest key range. Fix: lead with a high-cardinality component (tenant, shard, hash prefix) for hot write workloads; keep dates later in the key.
4. No Lifecycle Rules#
Incomplete multipart uploads, noncurrent versions and temporary processing outputs accumulate forever. Teams discover terabytes of invisible parts during a cost review. Fix: every bucket ships with abort-incomplete-multipart (1–7 days), noncurrent-version expiry and tiering rules.
5. Tiny Objects in Archive Tiers#
Millions of 2–20 KB objects moved to IA or Glacier pay 128 KB minimum billing or 40 KB metadata overhead each, plus a transition fee per object. Fix: compact into large archive objects first; keep tiny objects in Standard or Intelligent-Tiering, where objects under 128 KB simply stay in the frequent tier.
6. Trusting Events as the Source of Truth#
A Lambda that marks uploads READY on each event, with no reconciler. Events are at-least-once and occasionally delayed past a minute; a misconfigured filter or a disabled consumer drops them entirely. Fix: idempotent handlers keyed on (key, version or etag) plus a periodic sweep comparing the database to the bucket.
7. Long-Lived Presigned URLs as an Access Model#
Seven-day URLs pasted into emails, chat and logs become public links. Fix: short expiry (minutes), issued per request after an authorisation check; signed CDN URLs or cookies for repeated access.
8. Public Buckets and Wide Bucket Policies#
The most common S3 security incident is still a bucket or prefix made world-readable to "fix" an access problem. Fix: Block Public Access on at the account level, CDN with origin access control for public content, and policy-as-code review; see Security Fundamentals.
The Technology Landscape — Head-to-Head Comparison#
| Dimension | Amazon S3 (general purpose) | S3 Express One Zone | Google Cloud Storage / Azure Blob | Self-hosted (MinIO, Ceph RGW) | Block or file storage (EBS, EFS) |
|---|---|---|---|---|---|
| Model | Flat key → immutable object, HTTP API | Same API, directory buckets, single AZ | Same model, different APIs and tiers | S3-compatible API on your hardware | Disks and POSIX file systems |
| Consistency | Strong per key, LIST included | Strong | Strong per object | Depends on deployment | Strong (single attach or NFS semantics) |
| Latency | ~100–200 ms first byte for small objects | Single-digit ms | Similar to S3 | LAN-dependent, often lower | Sub-ms to low ms |
| Throughput model | Per-prefix rates, scales gradually | Hundreds of thousands of requests/s per bucket | Per-bucket and per-key guidance | Bounded by your cluster | Per volume IOPS and throughput |
| Durability | 11 nines designed, multi-AZ | 11 nines designed, one AZ | Multi-zone and multi-region options | What you build and test | Volume-level replication, snapshots |
| Ops burden | None for the service; yours for lifecycle and cost | Same | Same | High: disks, repair, upgrades | Low–medium |
| Pick when | Default for blobs, lakes, backups | Hot intermediate data, ML training scratch | You live in that cloud | Data sovereignty, on-prem, egress economics at PB scale | A process needs a disk, not an API |
🎯 Staff Insight: "The S3 API is the industry's de facto blob interface, so portability is about the API surface I depend on, not the vendor. I keep to PUT, GET, ranged GET, multipart, conditional writes and lifecycle, and avoid building core logic on vendor-only event or query features unless the payoff is clear."
Patterns#
Pattern 1: Presigned Direct Upload with a Metadata State Machine#
Client asks the API for permission; the API writes a PENDING row and returns presigned part URLs scoped to one key, a content-length range and a 15-minute expiry; the client uploads directly; an event or a client complete call moves the row forward after a HEAD check; an hourly reconciler fixes what events miss. Full mechanics in Large File Uploads & Delivery.
Pattern 2: Content-Addressed, Immutable Objects#
Key = hash of content. Writes use If-None-Match: *, so duplicate uploads become no-ops and dedupe is free. CDN headers can be max-age=31536000, immutable because a key never changes meaning. This is how Cloud File Sync designs store chunks.
Pattern 3: Event-Driven Processing Pipeline#
Events start the work; the database records the truth; the reconciler closes the gap. Write derived output to a different bucket or prefix than the trigger watches, or a function that writes back into its own trigger prefix loops forever, as the S3 documentation itself warns. This is the standard video pipeline shape.
Pattern 4: Data Lake with Table-Format Manifests#
Write large immutable Parquet files (128 MiB–1 GiB), then commit a new snapshot by atomically swapping a manifest pointer (an open table format, or a conditional write on a single metadata key). Readers resolve the manifest and never LIST the data prefix. Compaction jobs rewrite small files into large ones. See Batch & Stream Pipelines and Real-Time OLAP.
Pattern 5: Backup Vault in a Separate Account#
Production replicates (or a backup job copies) into a bucket owned by a separate account, with versioning, Object Lock and a policy that production principals can write but never delete. Ransomware or a compromised deploy role in production cannot reach the vault.
Scaling#
The Numbers#
| Resource | Documented limit or typical value | Design note |
|---|---|---|
| Request rate | ≥ 3,500 writes and 5,500 reads per second per partitioned prefix; no prefix limit | Scales gradually; expect 503 Slow Down while it does |
| Small-object latency | ~100–200 ms first byte (documented typical) | Cache hot objects; Express One Zone for single-digit ms |
| Single-instance throughput | Up to the instance NIC, e.g. 100 Gb/s, with parallel ranged GETs | Parallelism, not bigger requests, drives throughput |
| Single PUT | 5 GB | Multipart above ~100 MB |
| Multipart | 10,000 parts, 5 MiB–5 GiB each | Part size × 10,000 must exceed your largest object |
| Maximum object size | 48.8 TiB (~50 TB) | Raised from 5 TB in December 2025 |
| LIST page | 1,000 keys | 1 B keys = 1 M sequential LIST calls; use Inventory |
| Event delivery | At least once; seconds, sometimes a minute+ | Reconcile; don't use for ordering |
Request-Rate Math#
Launch: 50 M users upload ~2 photos/day, peak hour is 3x average
writes/s avg = 100 M / 86,400 ~= 1,160/s
writes/s peak = ~3,500/s -> already at one prefix's write rate
each photo also writes 3 thumbnails -> ~14,000 PUTs/s at peak
Key layout media/{tenant_shard 00-ff}/{date}/{uuid}:
256 leading shards -> ~55 PUTs/s per shard at peak, far under per-prefix rates
S3 still needs time to split ranges for a brand-new bucket:
ramp synthetic load over days before launch, or expect 503s in hour one
Run the same numbers through the Back-of-Envelope Calculator and the Sharding Planner when the key layout is the question.
Scaling Moves in Order#
- Fix the key layout so the hot workload varies in its leading characters.
- Retry 503s with exponential backoff and jitter in every client; SDK defaults help but stacked retries across layers do not.
- Parallelise: multipart uploads, ranged GETs (8–16 MiB ranges), many connections. Single-stream throughput is the wrong thing to tune.
- Cache in front: CDN for public reads, an in-memory cache for small hot objects, so S3 sees misses only.
- Batch small writes into larger objects (logs, events, telemetry). Fewer, bigger objects cut request bills and LIST cost.
- Change class or store when latency is the limit: Express One Zone for hot scratch data, or a database or cache for small mutable values that never belonged in object storage.
Failure Modes & Recovery#
1. 503 Slow Down at a Launch or Backfill#
- Symptom: Error rate on uploads or a batch job jumps to 5–30% with
503 Slow Down; retries make it worse. - Root cause: Request rate on a key range climbed faster than S3 repartitions; often a new prefix or a monotonic key.
- Detection: S3 request metrics
5xxErrorsby prefix filter; client-sideSlowDowncounters; retry rate per caller. - Fix: Throttle the producer, add jittered backoff, spread keys; for a backfill, ramp concurrency gradually.
- Prevention: Key layout review in design; load ramp before known launches; client retry budgets so one job can't amplify.
2. Lost or Delayed Events — Uploads Stuck in PENDING#
- Symptom: Users see "processing" for hours; support tickets spike; metadata shows thousands of PENDING rows older than an hour.
- Root cause: Consumer down or erroring into a DLQ nobody watches, a notification filter changed in a deploy, or events delayed.
- Detection: Age of oldest PENDING row; DLQ depth;
ApproximateAgeOfOldestMessageon the queue. - Fix: Replay the DLQ; run the reconciler on demand (HEAD each stuck key, advance or fail).
- Prevention: Reconciler on a schedule from day one; alert on PENDING age, not just consumer errors; notification config in code.
3. The Surprise Bill — Egress, Requests or Invisible Storage#
- Symptom: S3 spend doubles in a month without a matching traffic change.
- Root cause: Direct-from-origin serving after a CDN change, a job LISTing or GETting millions of objects in a loop, noncurrent versions piling up, or abandoned multipart parts.
- Detection: Cost allocation by bucket tag; Storage Lens for noncurrent bytes and incomplete multipart bytes; request metrics by operation.
- Fix: Restore the CDN path, kill the loop, add noncurrent expiry and abort-incomplete rules.
- Prevention: Per-bucket budgets and anomaly alerts; lifecycle rules required at bucket creation.
4. Mass Deletion or Overwrite by Your Own Code#
- Symptom: A deploy or script deletes or overwrites millions of objects; S3's durability is irrelevant because S3 did exactly what it was told.
- Root cause: Bad prefix in a cleanup job, wrong environment credentials, a lifecycle rule with a too-broad filter.
- Detection:
DeleteObjectrequest spikes; CloudTrail data events; delete-marker counts. - Fix: With versioning, remove delete markers or restore prior versions (batch operations at scale); without it, restore from the backup account.
- Prevention: Versioning on production buckets, deny-delete policies for most roles, a separate-account backup vault, and lifecycle changes reviewed like code.
5. Public Exposure or Leaked Presigned URLs#
- Symptom: Private files reachable by anyone; found by a researcher, a scanner or the press.
- Root cause: Block Public Access disabled for a "quick fix," overly broad bucket policy, or long-lived presigned URLs shared outside the product.
- Detection: Access Analyzer findings; config rules on bucket policies; unusual GET sources in access logs.
- Fix: Re-enable Block Public Access; rotate the credentials that signed leaked URLs (revoking the signer invalidates its URLs); audit access logs for the exposure window.
- Prevention: Account-level Block Public Access; short URL expiry; CDN signed URLs for repeat access; policy-as-code checks in CI.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| 503 Slow Down | 5xx by prefix, client SlowDown count | One workload's writes or reads | Backoff, throttle producer, spread keys | Workload team |
| Stuck PENDING | Oldest PENDING age, DLQ depth | New uploads for all users | Replay DLQ, run reconciler | Upload service team |
| Surprise bill | Cost anomaly by bucket tag | Budget, not users | Lifecycle rules, restore CDN path | Bucket owner; platform for guardrails |
| Self-inflicted deletion | Delete spikes, CloudTrail | Every object the job touched | Restore versions or from vault | Owning team; platform owns vault |
| Public exposure | Access Analyzer, config rules | Every exposed object, plus trust | Re-block, rotate signer, audit | Security + bucket owner |
| Region outage | Elevated 5xx across buckets | All buckets in the region | Fail reads to replica region if designed | Platform + service owners |
When to Use vs. Alternatives#
| Need | Pick | Why |
|---|---|---|
| User media, attachments, documents | S3 + CDN + metadata DB | Cheap, durable, direct upload and cacheable delivery |
| Analytics lake | S3 with Parquet and a table format | Separates storage from compute; many engines read it |
| Backups and compliance archives | S3 with versioning, Object Lock, Glacier tiers, separate account | WORM and cheap long retention |
| Small mutable records, counters, sessions | A database or cache (DynamoDB, Redis) | Per-request pricing and no CAS-heavy workloads in S3 |
| Hot scratch data for training or shuffle | Express One Zone or local NVMe | Single-digit ms and high request rates |
| A process that needs POSIX semantics | Block or file storage | Renames, appends and locks are file-system features |
| Multi-PB with massive egress, own data centres | Self-hosted S3-compatible storage, with real headcount | Unit economics at very large scale, as Dropbox did below |
When NOT to Use S3#
- As a database or queue. No queries, no transactions, per-request pricing, paginated LIST. Use a database for state and a queue for work.
- For small mutable values. Rewriting a 2 KB object 100 times a second is a last-writer-wins race and a request bill.
- For latency-critical reads without a cache. 100–200 ms first byte is fine for media and terrible for a request path with a 50 ms budget; check it against the Latency Budget.
- As the only copy of irreplaceable data in one account. Durability protects against disks failing, not credentials leaking.
- For workloads that need appends or renames. S3 has no append and no atomic rename; a "rename" is copy plus delete per object.
Operational Concerns#
What the On-Call Actually Does#
- Watches error rates by operation and prefix: 5xx (Slow Down), 4xx spikes (403 from expired URLs or policy changes), and client retry rates.
- Watches the async edges: queue depth and age for event consumers, oldest PENDING row, DLQ size.
- Reviews the cost dashboard weekly: storage by class, noncurrent bytes, incomplete multipart bytes, requests by operation, egress by bucket.
- Runs restores: from versions after bad deploys, from Glacier for support cases (with a stated restore SLA of hours, not minutes).
- Reviews bucket and lifecycle changes like code: a lifecycle filter is a scheduled mass deletion.
Key Metrics & Alerts#
| Metric | Healthy | Alert |
|---|---|---|
5xxErrors / AllRequests per bucket or prefix | < 0.1% | > 1% for 5 min |
4xxErrors rate | Stable baseline | 3× baseline (policy or URL expiry regression) |
FirstByteLatency p99 (request metrics) | Stable, ~100–200 ms small objects | 2× baseline for 15 min |
| Oldest PENDING upload age | < 5 min | > 30 min (page) |
Event queue ApproximateAgeOfOldestMessage | < 60 s | > 10 min |
| Noncurrent version bytes / current bytes | < 20% | > 50% (ticket) |
| Incomplete multipart bytes | ~0 | Any growth week over week |
| Daily cost per bucket | Within forecast | +30% day over day |
Interview Application — Staff-Level Plays#
Which Case Studies Use S3#
| Case Study | How Object Storage Is Used | Key Pattern |
|---|---|---|
| Design an Object Store | Build the thing itself | Metadata/data split, erasure coding, repair |
| Video Streaming | Originals and renditions behind a CDN | Event-driven transcode pipeline; egress via CDN |
| Cloud File Sync | Content-addressed chunks | Dedupe with create-once writes; metadata DB as truth |
| CDN | Origin for static and media content | Immutable keys, origin shielding, origin access control |
| Web Crawler | Raw page store, WARC-style batches | Batch small pages into large objects |
| Metrics & Monitoring | Long-term block storage for compressed series | Hot data local, cold blocks in object storage |
| Ad Click Aggregation | Raw event archive for reprocessing | Immutable raw log enables replay and audits |
Every System Design Question Has an S3 Moment#
- Chat app: "Attachments upload directly to S3 with a presigned URL scoped to one key and 10 minutes; the message stores the object ID, and recipients get a short-lived signed URL after an access check. See Chat."
- URL shortener: "S3 holds nothing on the hot path. It holds nightly exports of the mapping table for analytics and disaster recovery."
- News feed: "Images are content-addressed in S3 behind the CDN with year-long cache headers. The feed service never touches bytes."
- Payments or ledger: "Statements and audit exports go to a separate-account bucket with Object Lock in compliance mode for the 7-year retention legal signed off on. See Ledger & Wallet."
What Interviewers Probe#
| After You Say... | They Will Ask... | What They're Evaluating |
|---|---|---|
| "Store files in S3" | "How does the upload work for a 3 GB file on mobile?" | Presigned multipart, resumability, bytes off your servers |
| "S3 is eventually consistent" | "Are you sure? What about two writers?" | Knowing the 2020 change and what remains: LWW, no cross-key atomicity |
| "Trigger processing on upload events" | "What if the event never arrives, or arrives twice?" | Idempotency and reconciliation |
| "S3 scales infinitely" | "What happens when launch traffic hits one prefix?" | Per-prefix rates, gradual scaling, 503s, key design |
| "Lifecycle to Glacier" | "What does that cost for a billion 10 KB objects?" | Minimum size, minimum duration, transition fees |
| "We have 11 nines" | "What if someone runs the wrong delete script?" | Versioning, Object Lock, separate-account backups |
Common Interview Mistakes#
| What Candidates Say | What Interviewers Hear | What Staff Engineers Say |
|---|---|---|
| "Upload goes through our API to S3" | Will melt the API fleet on large files | "Presigned multipart direct to S3; the API issues permission and records metadata." |
| "List the folder to show the user's files" | Using S3 as a database | "The metadata DB answers listings; S3 only holds bytes." |
| "S3 is eventually consistent so add a delay" | Five years out of date | "Strong per key since 2020; I guard concurrent writers with conditional writes." |
| "Archive everything to Glacier to save money" | Hasn't modelled small-object fees | "Compact small objects first; tier only what's large and cold." |
| "Durability is 11 nines, so backups aren't needed" | Confuses hardware failure with human error | "Versioning plus a separate-account vault with Object Lock." |
| "Serve images straight from the bucket" | Hasn't priced egress | "CDN in front with immutable keys; origin egress drops ~10×." |
L5 vs L6 vs L7 Responses#
| Scenario | L5 Answer | L6 / Staff Answer | L7 / Principal Answer |
|---|---|---|---|
| "Store user photos" | S3 bucket, URL in DB | Presigned multipart, tenant-led keys, PENDING→READY with reconciler, CDN, lifecycle | Retention policy by data class; chargeback per product; egress architecture across products |
| "S3 bill tripled" | Move to Glacier | Break down storage vs requests vs egress; noncurrent versions, multipart orphans, CDN regression | Per-bucket budgets and anomaly alerts org-wide; lifecycle rules mandatory at creation |
| "Protect against ransomware" | Enable versioning | Separate-account vault, Object Lock, deny-delete roles, tested restore | Decides governance vs compliance mode with legal; restore drills as a quarterly game day |
| "Leave AWS for cost?" | Run MinIO | Compare total cost including egress, ops headcount and durability engineering | Treats it as a multi-year one-way door with a staged exit and a break-even point |
The Staff S3 Checklist#
- Classify objects: "Originals, derived, logs, backups: each gets its own bucket or prefix, class and retention."
- Key layout: "Tenant or shard first, date next, UUID or hash last; immutable keys, no overwrites."
- Upload path: "Presigned multipart, 16 MiB parts, 15-minute expiry,
If-None-Matchon complete." - Truth and events: "Metadata DB is the source of truth; events are at-least-once hints; reconciler hourly."
- Cost: "Storage, requests and egress modelled separately; CDN in front; lifecycle with abort-incomplete and noncurrent expiry."
- Protection: "Versioning, Block Public Access, separate-account vault, Object Lock where legal requires it."
🎯 Staff Insight: Don't use S3 as a database, a queue, a lock service or a low-latency key-value store. The strongest S3 signal is drawing the line between bytes (S3) and truth (the metadata database), and saying what happens when the event that connects them is lost.
Evaluation Rubric#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Consistency | "Eventually consistent" or no mention | Strong per key; LWW; conditional writes; no cross-key atomicity | Org pattern: immutable content-addressed objects, metadata DB as truth |
| Throughput | "Infinite" | Per-prefix rates, gradual scaling, key design, backoff | Recognises batching and caching as the real fix |
| Cost | Storage price only | Storage + requests + egress + lifecycle traps | Portfolio view, chargeback, egress architecture |
| Protection | "11 nines" | Versioning, replication, Object Lock, restore path | Data-loss posture and restore drills across the org |
| Operations | "Monitor the bucket" | PENDING age, DLQ, cost anomalies, lifecycle review | Guardrails at bucket creation enforced by the platform |
Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
At Staff level S3 is a well-configured bucket. At Principal level object storage is the company's largest, slowest-growing and least-watched cost line, and the place where retention policy, legal exposure and disaster recovery meet. Data accumulates for years because deleting it needs someone to say it's safe. The L7 question is not "which storage class?" but "what data do we keep, for how long, at what cost, who can delete it, and how fast can we get it back?"
🧭 Principal Move: "Before we tune buckets, I want a data classification with an owner and a retention period for each class. Lifecycle rules are just that policy written in code; without the policy, every bucket keeps everything forever and nobody is allowed to delete anything."
The Org-Level Fault Line#
Central storage platform vs every team owning its buckets.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Every team creates buckets freely | Fast; no gatekeeping | Inconsistent encryption, public-access mistakes, no lifecycle, unowned spend | Security and finance |
| Central team owns all buckets | Uniform policy | Bottleneck; teams wait for buckets; platform can't judge retention | Product velocity |
| Platform guardrails, team ownership | Teams self-serve through a template that enforces encryption, Block Public Access, tags and default lifecycle | Needs policy-as-code and exception handling | Platform headcount (small) |
| Separate backup and compliance accounts | Blast radius contained | Extra accounts and replication cost | Platform; legal signs retention |
The Principal default: guardrails, not gatekeeping. Teams create buckets through a template; account-level Block Public Access and encryption are non-negotiable; tags for owner, data class and cost centre are required; default lifecycle rules ship with the template; backups and compliance data live in separate accounts owned by the platform and security.
🧭 Principal Insight: "The expensive storage isn't what we read. It's what nobody remembers writing. Required owner and data-class tags are how a deletion decision becomes possible three years from now."
Cost Model#
Assumptions: Standard ~$0.021–0.023/GB-month at volume, Standard-IA ~$0.0125, Glacier Instant Retrieval ~$0.004; egress mostly through a CDN; loaded engineer ~$250K/year. Directional only; check current pricing.
| Scale | Stored | Mix | Storage/month | Requests + egress/month | Headcount | Total/month |
|---|---|---|---|---|---|---|
| Startup | 50 TB | All Standard | ~$1.1K | ~$2–5K | 0.1 FTE | ~$5–8K |
| Growth | 2 PB | 40% Standard, 40% IA, 20% Glacier IR | ~$28K | ~$40–80K | 0.5 FTE (~$10K) | ~$80–120K |
| Enterprise | 50 PB | 20% Standard, 30% IA, 50% archive tiers | ~$400K | ~$0.5–1.5M | 3–5 FTE (~$85K) | ~$1–2M |
Two observations. At every scale beyond startup, requests and egress rival or exceed storage, so CDN hit ratio and batching are worth more than storage-class tuning. And at enterprise scale, a retention policy that deletes data nobody needs is the single largest lever; moving the 30% of data with no reader from Standard to deletion beats any tiering scheme.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse |
|---|---|---|
| Object Lock in compliance mode with a long retention | One-way for the retention period | You pay for the data until the date; only deleting the account removes it |
| Key layout baked into clients and URLs | One-way-ish | Copy every object to new keys and migrate references |
| Deleting data without versioning or backup | One-way | Gone |
| Vendor-specific features in core logic (event formats, query features) | One-way-ish | Rewrite when moving providers |
| Storage class choice | Two-way | Transition fees and minimum-duration charges |
| Lifecycle rules | Two-way going forward | Already-expired objects stay expired |
| CDN in front of S3 | Two-way | Configuration and cache warm-up |
The Standard I'd Write#
RFC-STOR-002: Object Storage Baseline
Scope: Every object storage bucket in every production account.
MUST
1. Be created from the platform template: account-level Block Public Access,
default encryption, required tags (owner, data_class, cost_center).
2. Enable versioning unless data_class is "regenerable".
3. Include lifecycle rules: abort incomplete multipart uploads after 7 days;
expire noncurrent versions after 30 days (or the data class retention).
4. Keep business metadata and access decisions in a database, not in key listings.
5. Make every event consumer idempotent, with a reconciler that runs at least daily.
6. Use presigned URLs of at most 1 hour for user access; no long-lived keys on clients.
SHOULD
7. Front public and semi-public reads with a CDN using origin access control.
8. Batch objects smaller than 128 KB before transitioning to IA or archive classes.
9. Replicate regulated or irreplaceable data to the backup account with Object Lock.
Exceptions: platform + security approval, time-boxed to two quarters, listed on a dashboard.
Enforcement: policy checks in CI and account config rules; warn for one quarter, then block.
Success metrics: zero public-bucket findings; incomplete-multipart bytes ~0;
100% of buckets tagged; restore drill for each regulated dataset twice a year.
What I'd Tell the VP#
"Object storage is cheap per gigabyte, which is exactly why it has become one of our larger and fastest-growing bills: we keep everything forever and serve some of it inefficiently. About half of what we pay is not storage at all but requests and data leaving our cloud. I'm proposing three things: put a content delivery network in front of everything public, give every dataset an owner and a retention period so we can delete what nobody needs, and keep protected backups in a separate account so a mistake or an attacker can't erase our data. That should cut the storage bill by a third within two quarters and close our largest data-loss risk. The cost is roughly half an engineer to build the template and policies."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Retention as policy | "Lifecycle rules encode a retention policy legal signed; no policy, no deletion." |
| Prices the whole bill | "Storage is the small line; requests and egress are where the money goes." |
| Designs the loss posture | "Separate-account vault, Object Lock, restore drills: durability doesn't cover us." |
| Guardrails over gatekeeping | "Teams self-serve buckets from a template that makes the safe thing the default." |
| Respects one-way doors | "Compliance mode is irreversible, so it needs legal sign-off and a test bucket first." |
Staff answers that L7 interviewers find insufficient:
- "Add lifecycle rules" without who decides retention per data class.
- "Enable versioning" without noncurrent expiry or the cost of doubling storage under churn.
- "Use Glacier" without restore times, retrieval fees and the small-object trap.
How Real Companies Built It#
Amazon S3 — Strong Consistency Without a Price Increase#
In December 2020 S3 moved from eventual to strong read-after-write consistency for every new and existing object. Werner Vogels described how: S3 added an in-memory "witness" component that tracks only enough metadata to act as a read barrier, so the metadata cache can learn whether its view of an object is stale. The stated goal was strong consistency "with no additional cost" and "no performance or availability tradeoffs," and the team leaned on formal methods and model checking to verify the protocol (All Things Distributed, 2021).
Staff insight: The guarantee changed underneath millions of applications without an API change. In an interview, that is why "S3 is eventually consistent" is a red flag today, and why the remaining gaps (concurrent writers, cross-key updates, bucket configuration) are what you should name.
Amazon S3 — Heat Management Across Millions of Drives#
Andy Warfield's 2023 essay on operating S3 explains that the hard problem is "heat," request load on individual disks. S3 spreads new objects broadly across its disk fleet, deliberately placing different objects on different sets of disks, so that at its scale no single workload can meaningfully move the aggregate peak and a single customer can burst across a very large number of drives. The essay also describes ShardStore, S3's rewritten storage node software, verified with lightweight formal methods in Rust (All Things Distributed, 2023).
Staff insight: Multi-tenancy is the performance feature: aggregation smooths bursty workloads. When you design your own store, as in Design an Object Store, placement and load spreading matter as much as replication.
Dropbox — Magic Pocket, Leaving S3 at 500 PB#
Dropbox built Magic Pocket, its own block storage system, and by October 2015 had moved over 90% of its users' data off S3 onto it, with more than 500 PB of user data under management by early 2016. Dropbox cited end-to-end control of the stack for performance and better unit economics from customising hardware and software for its specific workload, and reported completing the migration without major disruptions or data loss (Dropbox engineering).
Staff insight: Leaving managed object storage is rational only at very large scale with a stable, well-understood workload and a team able to own durability. In an interview, the Staff answer to "should we build our own?" names that threshold and the multi-year engineering cost before saying yes.
Practice Drill#
Prompt: "Your photo-sharing product stores 800 TB in one S3 bucket. In the last quarter the S3 bill grew 2.5× while traffic grew 30%. A new 'burst upload' feature launched last week and uploads fail with 503s for the first hour of every evening peak. Some users also report photos stuck on 'processing' for a day. You own storage. What do you do?"
Staff Answer
Three symptoms, three different mechanisms, so I'd split them. 503s at peak: the burst feature probably writes keys that start with a timestamp or a single new prefix, so all evening writes land on one key range faster than S3 repartitions. I'd confirm with request metrics filtered by prefix, then change new keys to lead with a tenant shard (media/{shard 00-ff}/{date}/{uuid}), make sure the client retries 503s with exponential backoff and jitter under a retry budget, and ramp load on the new layout before the next peak. If each burst is 30 small photos, I'd also batch thumbnail writes. Stuck processing: events are at-least-once and the consumer or its DLQ is the likely gap. I'd alert on oldest PENDING age and DLQ depth, replay the DLQ, and add an hourly reconciler that HEADs stuck keys and advances or fails them, with handlers idempotent on (key, etag). The bill: 2.5× cost on 1.3× traffic means cost per request or per byte rose. I'd break spend into storage by class, noncurrent versions, incomplete multipart bytes, requests by operation and egress. Likely culprits: versioning without noncurrent expiry, abandoned multipart parts from failed burst uploads, and reads bypassing the CDN for the new feature. Fixes: abort-incomplete after 7 days, noncurrent expiry at 30 days, CDN with origin access control on all image reads, then lifecycle to Standard-IA at 30 days for originals over 128 KB. I'd target a 40–50% cost reduction, zero 503s at peak and PENDING age under 5 minutes, and I'd report those three numbers weekly.
Why this is L6:
- Separates three symptoms into three mechanisms (key range heat, event loss, cost per unit) instead of one blanket fix.
- Uses the current consistency and event model correctly: at-least-once events need idempotency plus a reconciler.
- Decomposes cost into storage, requests, egress and invisible bytes before reaching for storage classes.
What L7 adds:
- Turns the incident into guardrails: a bucket template with mandatory lifecycle rules, prefix-design review for new features and per-bucket cost anomaly alerts.
- Asks for a retention policy per data class so derived renditions are regenerated or expired instead of stored forever.
- Prices the CDN and the reconciler against the savings and reports the result to finance as cost per active user.
❌ Common L5 Trap
"Move older photos to Glacier to cut the bill, and increase the client retry count so the 503s go away."
Why this misses: Glacier tiers add retrieval delays and per-object fees without touching the actual growth drivers (noncurrent versions, orphaned parts, egress), and more retries amplify the 503s on an already hot key range. It also ignores the stuck uploads entirely, which need idempotent consumers and a reconciler, not cheaper storage.
Quick Reference Card#
Model: flat key -> immutable object; "folders" are key prefixes
Consistency: strong read-after-write per key (PUT, DELETE, LIST) since Dec 2020
concurrent writers: last-writer-wins; no cross-key atomicity
bucket config: eventually consistent (~15 min after enabling versioning)
Conditional: If-None-Match: * (create once), If-Match: etag (CAS) on PUT/complete/copy
Request rates: >= 3,500 writes, 5,500 reads per second per partitioned prefix;
scales gradually, 503 Slow Down meanwhile; no prefix limit
Latency: ~100-200 ms first byte (small objects); Express One Zone single-digit ms
Single PUT: up to 5 GB; use multipart above ~100 MB
Multipart: 5 MiB-5 GiB parts, 10,000 parts; max object 48.8 TiB (~50 TB, Dec 2025)
Classes: Standard | Intelligent-Tiering | Standard-IA (30d, 128 KB min)
One Zone-IA (1 AZ) | Express One Zone (1 AZ)
Glacier IR (90d) | Flexible (90d, restore) | Deep Archive (180d, hours)
Durability: 11 nines designed, every class; not protection from your own deletes
Events: at least once; seconds, sometimes a minute+; no direct FIFO SQS
Presigned URLs: up to 7 days (IAM user, SigV4); role/STS creds expire sooner
Object Lock: needs versioning; governance (bypassable) vs compliance (nobody, not root)
Pricing (approx): Standard ~$0.023/GB-mo; PUT $0.005/1K; GET $0.0004/1K; egress ~$0.09/GB
RED FLAGS
- Bytes proxied through application servers
- LIST used to answer user-facing queries
- Monotonic timestamp at the start of hot keys
- No abort-incomplete-multipart or noncurrent-version lifecycle rule
- Tiny objects tiered to IA or Glacier
- Event handlers without idempotency or a reconciler
- Multi-day presigned URLs handed to users
- Production buckets without versioning or a separate-account backup