Hiring BarSupport

Design with DynamoDB — Staff-Level Technology Guide

Technology guide40 min read5 diagrams

Why This Matters#

DynamoDB is not a flexible NoSQL database. It is a contract: tell it every access pattern up front, spread load evenly across partition keys, and in return it delivers single-digit-millisecond latency at any scale with no servers to operate. Every DynamoDB design succeeds or fails on that contract. The interesting question is never "can DynamoDB handle the load?" — it can — but "what are our access patterns, does every one of them map to a key lookup or a single-partition range query, and who pays when a pattern we didn't design for shows up in month nine?"

It shows up in interviews because it forces the query-first discipline interviewers want to see, and because it hides its failure modes behind "fully managed." The managed part is real: no vacuum, no failover runbook, no patching. The unmanaged part is yours: hot partitions capped at ~1,000 writes/sec, GSIs that throttle your base table, Scan bills that arrive at the end of the month, 400 KB item limits, and a data model that can't answer a new question without a backfill.

The L5 answer says "DynamoDB with user_id as the partition key." The L6 answer says "single table; PK = USER#<id>, SK = ORDER#<ts> so 'orders for user newest-first' is one Query; a sparse GSI on status = OPEN for the ops dashboard; conditional writes for idempotency; on-demand until traffic is predictable, then provisioned with auto-scaling at 70% target; streams into Lambda for the search projection." The L7 answer prices the table at three scales, asks what fraction of the org's AWS bill is DynamoDB request units, and decides which workloads are too relational to belong there at all.

The 60-Second Pitch#

"DynamoDB gives us single-digit-millisecond reads and writes at effectively unlimited scale, replicated across three availability zones, with no servers, patches, or failovers to manage. Throughput scales by spreading items across partitions — each handles up to about 3,000 reads and 1,000 writes a second — so the design work is choosing keys that spread load evenly and that answer each access pattern with a single GetItem or Query. We get conditional writes and multi-item transactions for correctness, streams for change data capture, and global tables if we go multi-region. The tradeoff is flexibility: new query shapes need new indexes or backfills, so I'd list every access pattern before writing the schema."

The Staff-level insight: DynamoDB converts operational risk into modeling risk and billing risk. You stop worrying about failover and start worrying about hot keys, index design, and request-unit cost. The teams that thrive treat the access-pattern list as a versioned design artifact with an owner.

IntentConstraintStrategyFailure ModeCorrectness Bar
High-scale key-value / session / cartp99 < 10 ms at 100K+ req/sSimple PK, on-demand or provisioned, DAX if read-hotHot key throttlingPer-item strong consistency where needed
Entity with known access patterns (orders, messages)3–10 query shapes, predictableSingle-table design, composite sort keys, 1–3 GSIsNew query needs a backfillConditional writes for invariants
Event/metadata store with CDC fan-outDownstream consumers need every changeStreams → Lambda/Kinesis → search, analyticsConsumer lag, 24 h stream retentionAt-least-once, idempotent consumers
Global active-activeMulti-region writes, low local latencyGlobal tablesConcurrent writes resolved last-writer-winsDesign writes to be region-homed or commutative

🎯 Staff Move: "Before I write a key, I'll list the access patterns — there are six. Five map to a partition-key lookup or a range query on the sort key. The sixth, 'all open orders across users', gets a sparse GSI. If product adds a seventh we can't model, that's the moment we add a projection to Elasticsearch, not the moment we add a Scan."


Architecture & Internals#

Only the internals that change design decisions.

Partitions: The Unit of Everything#

A table's items are hashed by partition key onto partitions. The published limits per partition drive every design:

LimitValueDesign Consequence
Storage per partition~10 GBLarge item collections split across partitions (by sort-key range)
Read throughput per partition3,000 RCU/s (≈ 3,000 strongly consistent 4 KB reads, 6,000 eventually consistent)A single hot key tops out here
Write throughput per partition1,000 WCU/s (≈ 1,000 writes of ≤ 1 KB)The most common real ceiling
Item size400 KB maxLarge payloads go to S3 with a pointer
Query / Scan page1 MB of data per callPaginate with LastEvaluatedKey

DynamoDB splits partitions automatically as data grows or as a partition runs hot (split for heat), and adaptive capacity shifts unused table throughput to busy partitions. What it cannot do is split a single partition-key value that exceeds per-partition limits when every item shares the same key — "split for heat" can separate items by sort key within a partition key, but a single item taking 2,000 writes/s will throttle no matter what.

Diagram: Partitions: The Unit of Everything

Replication and Consistency#

Each partition is replicated to three storage nodes in three AZs, with a leader elected via Multi-Paxos (per the public 2022 USENIX ATC DynamoDB paper). A write is acknowledged after the leader and at least one other replica durably log it.

Read TypeServed ByGuaranteeCost
Eventually consistent (default)Any replicaMay miss writes from the last ~second0.5 RCU per 4 KB
Strongly consistent (ConsistentRead=true)LeaderReflects all acknowledged writes1 RCU per 4 KB
Transactional (TransactGetItems)Leader, serializable with transactionsIsolated from in-flight transactions2 RCU per 4 KB
GSI queryGSI replicasAlways eventually consistent0.5 RCU per 4 KB

Strong reads double read cost and are not available on GSIs or across regions in classic global tables. Use them where you read-then-write in the same request path; use eventual reads everywhere else.

Capacity Units — The Cost Model Is the Performance Model#

Write cost  = ceil(item_size_KB / 1)  WCU  per write     (×2 for transactional)
Read cost   = ceil(item_size_KB / 4)  RCU  per strong read  (×0.5 eventual, ×2 transactional)
Query cost  = ceil(total_bytes_read_before_filter / 4 KB) RCU   — filters do NOT reduce cost
GSI cost    = every base-table write that changes a projected attribute = another write on the GSI

A 3 KB item costs 3 WCU per write. An update that changes one attribute still rewrites the whole item — 3 WCU. With 2 GSIs projecting ALL, that single write costs up to 9 WCU (plus GSI update-as-delete+insert when the GSI key changes). Item size and index count multiply cost linearly — this is why trimming items and projecting KEYS_ONLY / INCLUDE matters.

Capacity Modes#

ModeBilling (us-east-1, approximate)Scaling BehaviorPick When
On-demand~$0.625 per million write units, ~$0.125 per million read unitsInstantly handles up to 2× previous peak; ramps beyond that over timeNew or spiky workloads; low average utilization
Provisioned + auto-scaling~$0.00065 per WCU-hour, ~$0.00013 per RCU-hourAuto-scaling reacts in minutes; bursts use up to 300 s of unused capacitySteady traffic; typically 3–7× cheaper than on-demand at high, flat utilization
Provisioned + reserved capacityUp to ~50–75% below provisioned for 1–3 year termsSame as provisionedLarge, stable baseline

You can switch modes (with limits on how often). The Staff default: on-demand at launch, move to provisioned once 30 days of traffic show a predictable baseline — and model the crossover: on-demand is cheaper when average utilization of equivalent provisioned capacity would be below roughly 15–30%.

Streams#

DynamoDB Streams emit an ordered, de-duplicated record for every item change (new/old images configurable), retained 24 hours, sharded to match table partitions. Lambda polls with a checkpoint per shard; a failing record blocks its shard (and all later changes to those items) until it succeeds, expires, or is sent to an on-failure destination. Kinesis Data Streams for DynamoDB is the alternative when you need up to 1-year retention and more consumers, at the cost of possible duplicates and ordering you must re-establish with timestamps/versions.


Data Modeling / Core Usage — "The Entire Game"#

DynamoDB modeling is access-pattern-first. You don't design entities and then write queries; you enumerate queries and design keys so each is a GetItem or a single-partition Query. Anything else is a Scan, and a Scan on the hot path is a design defect.

Step 1: Write the Access-Pattern Table#

Example — an e-commerce order service:

#Access PatternFrequencyKey ConditionIndex
1Get customer profile20K/sPK = CUST#<id>, SK = PROFILETable
2List a customer's orders, newest first5K/sPK = CUST#<id>, SK begins_with ORDER#, descendingTable
3Get order with its line items3K/sPK = ORDER#<id> (all SKs)Table
4Look up order by external payment ID500/sGSI1PK = PAY#<payment_id>GSI1
5Ops: all orders in OPEN status older than 1 h1/minGSI2PK = STATUS#OPEN, GSI2SK < <ts>Sparse GSI2 (sharded)
6Enforce unique email per customeron signupPK = EMAIL#<email> conditional putTable

Step 2: Design the Keys (Single-Table)#

EntityPKSKGSI1PK / GSI1SKGSI2PK / GSI2SKAttributes
Customer profileCUST#42PROFILE——name, email, tier
Order (customer view)CUST#42ORDER#2026-09-30T10:04Z#o981——total, status (summary)
Order headerORDER#o981HEADERPAY#pi_77 / ORDER#o981STATUS#OPEN#<shard 0-9> / <created_ts>full order, version
Line itemORDER#o981ITEM#001——sku, qty, price
Email uniquenessEMAIL#a@x.comEMAIL——customer_id
Diagram: Step 2: Design the Keys (Single-Table)

Patterns in this table:

  • Item collections: everything under ORDER#o981 comes back from one Query — the header and all line items in ~5–10 ms, no joins.
  • Composite sort keys: ORDER#<iso_ts>#<id> sorts chronologically, supports begins_with and time-range BETWEEN, and stays unique.
  • Sparse index: only open orders carry GSI2PK; when status changes to SHIPPED, the app removes the attribute and the item disappears from GSI2. The ops query reads only open orders.
  • Write-sharded GSI key: STATUS#OPEN#<0-9> spreads a low-cardinality key across 10 partitions; the reader issues 10 parallel queries and merges.
  • Uniqueness via a sentinel item: a transaction puts CUST#42 and EMAIL#a@x.com with attribute_not_exists(PK) — the database enforces uniqueness.
-- Pattern 2: newest 20 orders for a customer
Query(TableName='app-main',
      KeyConditionExpression='PK = :pk AND begins_with(SK, :pfx)',
      ExpressionAttributeValues={':pk':'CUST#42', ':pfx':'ORDER#'},
      ScanIndexForward=False, Limit=20)

-- Pattern 6: create customer with unique email — atomic, both or neither
TransactWriteItems([
  Put(Item={PK:'CUST#42', SK:'PROFILE', email:'a@x.com', ...}, ConditionExpression='attribute_not_exists(PK)'),
  Put(Item={PK:'EMAIL#a@x.com', SK:'EMAIL', customer_id:'42'},  ConditionExpression='attribute_not_exists(PK)')
])

Single-Table vs Multi-Table#

ApproachWinsLosesStaff Default
Single-tableFetch heterogeneous related items in one Query; fewer tables to provision and monitorHarder to read, harder to evolve, analytics exports need unpacking, all entities share one table's settingsWhen a service has tightly related entities read together on hot paths
Table per entityClear ownership, independent capacity/TTL/streams/backup policies, simpler toolingMultiple round trips for related dataWhen entities have different lifecycles or owners, or the team is new to DynamoDB

Single-table design is an optimization, not a religion. The Staff move is to justify it by a specific hot-path query that reads multiple entity types together.

Conditional Writes: Correctness Without Locks#

-- Optimistic concurrency: update only if nobody else changed it
UpdateItem(Key={PK:'ORDER#o981', SK:'HEADER'},
  UpdateExpression='SET #s = :new, version = version + :one',
  ConditionExpression='version = :expected AND #s = :old',
  ...)
-- ConditionalCheckFailedException → reread and retry (or report conflict)

-- Idempotent create: second attempt is a no-op failure, not a duplicate
PutItem(Item={PK:'IDEM#<key>', SK:'IDEM', response:..., ttl:<now+24h>},
        ConditionExpression='attribute_not_exists(PK)')

-- Inventory decrement that can never go negative
UpdateItem(Key={PK:'SKU#42', SK:'STOCK'},
  UpdateExpression='SET available = available - :q',
  ConditionExpression='available >= :q')

Conditional writes are evaluated atomically on the item's partition leader — they are the DynamoDB equivalent of Postgres's UPDATE ... WHERE. Hot-item contention limits them to the per-item write ceiling (well under 1,000/s for a single item under conflict); a flash-sale SKU needs inventory split into N bucket items. See Dealing with Contention.

Transactions#

TransactWriteItems / TransactGetItems provide ACID across up to 100 items (and 4 MB) in one region, across tables. Costs: 2× capacity units, higher latency (~2–3× a single write), and TransactionCanceledException on conflicts with other transactions or concurrent writes. Use them for invariants spanning items (uniqueness sentinels, transfers, order + inventory), not as a default write path.

Large Items and Hot Collections#

  • Items > ~10–50 KB: split into a header item plus detail items, or store the blob in S3 with a pointer. Every read and write is billed on full item size.
  • Item collections that grow unbounded (a chat channel's messages under one PK) are fine for storage (they split across partitions by sort key) but not for write rate on that PK — time-bucket the key (CHAN#<id>#<yyyymm>) if one channel can exceed ~1,000 writes/s.

The Tunable Tradeoff — Consistency × Cost × Flexibility#

Dial 1: Read Consistency#

ChoiceStalenessCostAvailabilityUse For
Eventually consistent readUsually < 1 s0.5×Highest (any replica)Feeds, catalogs, profile displays
Strongly consistent read0 within region1×Leader onlyRead-modify-write, balances, "did my write land?"
Transactional read0, isolated from txns2×Leader onlyReading an invariant spanning items
GSI readEventually consistent, typically ms–~1 s0.5×Any GSI replicaSecondary access paths — never read-your-write
Global table replica read (classic)Cross-region lag, typically ~1 sLocal priceLocal regionRegion-local reads

Dial 2: Capacity Mode vs Throttle Risk#

on_demand_monthly   ≈ (writes/month × $0.625/M × WCU_per_write) + (reads/month × $0.125/M × RCU_per_read)
provisioned_monthly ≈ (provisioned_WCU × $0.00065 + provisioned_RCU × $0.00013) × 730 h
crossover: provisioned wins when average_utilization of the provisioned capacity > ~15–30%

Example: 2,000 writes/s steady (1 KB) and 10,000 eventually consistent reads/s (4 KB).

  • On-demand: 5.2B writes × $0.625/M ≈ $3,250; 25.9B reads × 0.5 RRU × $0.125/M ≈ $1,620 → ~$4.9K/month.
  • Provisioned at 70% target utilization: ~2,860 WCU × $0.00065 × 730 ≈ $1,360; 7,150 RCU × $0.00013 × 730 ≈ $680 → **$2K/month**, less with reserved capacity.
  • Price of provisioned: throttling if traffic jumps faster than auto-scaling reacts (minutes) and exhausts burst capacity.

Dial 3: Index Flexibility vs Write Amplification#

Every GSI is a second table maintained asynchronously. More GSIs = more query flexibility and proportionally more write cost, storage, and throttling surface.

GSI ProjectionStorageWrite CostQuery Needs Base Fetch?
KEYS_ONLYMinimalLow (small items)Yes — extra GetItem per result
INCLUDE (listed attributes)ModerateModerateOnly for non-projected fields
ALLFull copyUp to 1× base write per GSINo

🎯 Staff Move: "Every GSI is a second write and a second throttling surface. I'll add one only for an access pattern with real volume, project just the fields that query returns, and make it sparse when the pattern only cares about a subset — like open orders."


Anti-Patterns — What Kills DynamoDB Deployments#

1. Modeling Like a Relational Database#

One table per entity, IDs as foreign keys, and "joins" done in the application with N+1 GetItem calls. A page that needs 40 items makes 40 sequential round trips at ~5 ms each = 200 ms. Fix: access-pattern table first; item collections for data read together; BatchGetItem (up to 100 items) when you must fan out.

2. Scan and FilterExpression on the Hot Path#

A Scan reads the whole table at 1 MB per page and bills every byte read — filters apply after capacity is consumed. A 100 GB table scan costs ~12.5M RCU (eventual: ~6.5M) per full pass. Fix: a GSI (often sparse) for the access pattern; exports to S3 + Athena for analytics; Scan only in offline jobs with parallel segments and rate limiting.

3. Low-Cardinality or Monotonic Partition Keys#

PK = status, PK = date, PK = tenant where one tenant is 40% of traffic. Everything lands on a few partitions; per-partition limits throttle while table-level capacity looks idle. Fix: high-cardinality keys; write sharding (DATE#2026-09-30#<0-19>); per-tenant key spreading for large tenants.

4. GSI Back-Pressure#

A GSI's partition key is hot (e.g., GSI1PK = COUNTRY#US) or under-provisioned. When a GSI throttles, base-table writes that update it are throttled too. The symptom looks like a base-table problem. Fix: design GSI keys with the same cardinality discipline as base keys; alarm on GSI throttle metrics separately.

5. Big Items That Change Often#

A 200 KB user document with a last_seen timestamp updated on every request — 200 WCU per heartbeat. Fix: split volatile attributes into their own small item (PK = USER#42, SK = PRESENCE), or move them to Redis.

6. Treating TTL as a Scheduler#

TTL deletes expired items in the background, typically within a few days of expiry, not at the second. Queries still return expired items until deletion. Fix: filter on the TTL attribute in reads (ttl > :now); use EventBridge Scheduler or a queue for anything that must fire on time.

7. Unbounded Retries Without Jitter#

SDK defaults retry throttled requests with exponential backoff, but services wrapping them in their own retries multiply attempts (3 × 3 × 3 = 27 per user request), turning a brief throttle into a sustained one. Fix: retry budget per request, jittered backoff, and one retry layer.

Anti-PatternDetection SignalBlast RadiusWho Pays
Relational modelingp99 grows with page complexity; high request count per API callLatency SLOProduct users
Scan on hot pathConsumedReadCapacityUnits ≫ items returned; Scan in CloudTrail/X-RayCost + throttling for allFinance, neighbors on table
Low-cardinality PKThrottledRequests while consumed ≪ provisioned; Contributor Insights top keysHot partition's itemsCallers of those keys
GSI back-pressureWriteThrottleEvents on the GSIBase-table writesEvery writer to the table
Big volatile itemsWCU per request ≫ 1CostOwning team's budget
TTL as schedulerExpired items in query resultsCorrectness of time-based logicUsers
Retry stormsRequest count ≫ user traffic during throttlingAmplified throttlingEveryone on the table

The Technology Landscape / Head-to-Head Comparison#

DimensionDynamoDBCassandra / ScyllaDBPostgreSQLMongoDB (Atlas)Spanner / CockroachDBRedis
ModelKey-value + sorted item collectionsWide-column, partition + clustering keysRelationalDocumentDistributed relationalIn-memory structures
OpsZero serversHigh (repair, compaction, JVM)MediumLow (managed)Low (managed) / mediumMedium
ScaleEffectively unlimited with good keysLinear, self-managedSingle writer; shard manuallySharded clustersHorizontal writesRAM-bound shards
LatencySingle-digit ms; µs with DAXms1–3 msms5–20 ms writessub-ms
TransactionsUp to 100 items, one regionLWT per partitionFull ACIDMulti-document ACIDFull, distributedSingle-shard atomic
Query flexibilityLow — PK required; GSIs pre-plannedLowHighMedium (secondary indexes, aggregation)HighLow
Multi-regionGlobal tables (multi-active, LWW; newer strong option)Native multi-DCReplicas, one writerGlobal clustersNativeLimited
Lock-inAWS onlyPortablePortablePortable-ishVendorPortable
Pick whenAWS-native, known patterns, spiky or massive scale, tiny ops teamMulti-cloud, huge write throughput, own opsEvolving queries, invariantsDocument shapes, flexible queriesGlobal strong consistencyHot data, sub-ms

DynamoDB vs Cassandra: similar data models (partition key + sorted collection). DynamoDB trades portability and per-request pricing for zero operations; Cassandra trades operations headcount for portability and lower unit cost at very high steady throughput. At ~$100K+/month of steady DynamoDB spend, a self-managed Cassandra/Scylla cluster can be cheaper in infrastructure — but add 2–4 engineers. See Cassandra.

DynamoDB vs Postgres: if access patterns are still changing or you need ad-hoc queries, joins, or aggregates, Postgres wins until a named ceiling. See PostgreSQL and Database Selection.

🎯 Staff Insight: "DynamoDB is the cheapest database to operate and the most expensive database to be wrong about." The ops savings are real; the cost of a missed access pattern is a backfill over billions of items. Say that tradeoff out loud when you pick it.


Patterns#

Pattern 1: Write Sharding for Hot Keys#

-- Writes: spread one logical key across N physical keys
shard = hash(event_id) % 20
PutItem(PK = 'VOTES#contest9#' + shard, SK = event_id, ...)
-- or atomic counter per shard:
UpdateItem(PK = 'COUNT#contest9#' + shard, UpdateExpression='ADD n :one')

-- Reads: query all 20 in parallel and merge / sum

20 shards lift a logical key's ceiling from ~1,000 to ~20,000 writes/s. Cost: reads fan out 20×. Pick N from peak rate ÷ ~500 (50% headroom per partition). See Leaderboard & Counting and Scaling Writes.

Pattern 2: Streams → Projections (CQRS)#

Table stays optimized for OLTP access patterns; a stream consumer maintains projections elsewhere: OpenSearch/Elasticsearch for search, aggregates back into DynamoDB, S3 for analytics. Consumers must be idempotent (use item version in the projection write) because retries replay records.

Diagram: Pattern 2: Streams → Projections (CQRS)

Pattern 3: Idempotency Table#

PutItem with attribute_not_exists(PK) on IDEM#<key>, storing the response and a 24–72 h TTL. On conflict, return the stored response. Cheap (1–2 WCU), durable across 3 AZs, and naturally expires. The DynamoDB-native answer to the idempotency question in Payment Processing.

Pattern 4: Time-Bucketed Keys for Append-Heavy Data#

Chat messages, IoT readings, audit events: PK = DEVICE#<id>#<yyyymmdd>, SK = <ts>#<seq>. Bounds per-key write rate and lets you age data by bucket. Pair with TTL and S3 export for retention. See Chat Messaging.

Pattern 5: DAX for Read-Heavy Hot Items#

DynamoDB Accelerator: an in-VPC write-through cache with microsecond reads for GetItem/Query. Eventually consistent only (strong reads pass through), item cache and query cache invalidate differently (query cache isn't invalidated by item writes — TTL only). Use for read-heavy, hot-item workloads (product pages during a sale); skip for write-heavy or strongly consistent paths.

Pattern 6: Global Tables for Multi-Region#

Multi-active replication, typically sub-second across regions, last-writer-wins on concurrent updates to the same item. Design so each item is homed to one region's writers (user's home region), or so updates are commutative (ADD counters). A newer multi-region strong consistency mode exists for workloads that need it, at higher write latency — evaluate it explicitly rather than assuming classic behavior.

PatternContractPrimary RiskOwner
Write shardingScales hot logical keysRead fan-out cost; wrong NService team
Streams projectionsAt-least-once derived viewsPoison record blocks shard; 24 h retentionService team + projection owner
Idempotency tableExactly-once effectTTL too short vs client retry windowService team
Time bucketsBounded key heatCross-bucket queriesService team
DAXµs readsStale query cacheService team
Global tablesMulti-region availabilityLWW overwrites on concurrent writesService + platform

Scaling#

What Scales Automatically and What Doesn't#

DimensionAutomatic?Your Job
Table storageYes — no practical limitWatch cost ($0.25/GB-month standard; Standard-IA class ~60% cheaper storage, pricier requests)
Table throughputYes, if keys are spreadKeep partition-key cardinality high; watch per-key heat
Single partition-key throughputPartially (split for heat by sort key)Write-shard keys > ~500 writes/s
Single item throughputNoRedesign: bucket items, move counters
On-demand burstsUp to 2× previous peak instantlyPre-warm (switch to provisioned high, or ramp traffic) before a known 10× event
GSIsSame partitioning rulesDesign GSI keys for cardinality too
Account/table quotasNo — soft limits (e.g., default per-table throughput quotas, 20 GSIs per table)Request increases weeks before launches

Pre-Warming for Known Peaks#

On-demand tables handle up to double their previous peak immediately. For a flash sale expected at 20× normal: (1) raise the table's warm throughput (DynamoDB exposes warm throughput settings) or temporarily switch to provisioned at peak levels days ahead so partitions are pre-split, (2) request quota increases, (3) load-test at 1.5× expected peak with realistic key skew. See Flash Sales.

Capacity Sketch#

Chat service: 50M DAU × 40 messages/day = 2B messages/day ≈ 23K writes/s avg, 70K/s peak
Item: 600 bytes → 1 WCU; with 1 GSI (KEYS_ONLY, ~100 bytes) → +1 WCU → 2 WCU per message
Peak write units: 140K WCU/s → spread over ≥ 140 partitions' worth of keys (fine: conversation IDs are high-cardinality)
Hottest conversation: 2K messages/s in a live-event channel → exceeds 1K WCU per key → bucket by (channel, minute) or shard 4×
Monthly on-demand writes: 2B × 30 × 2 WRU × $0.625/M ≈ $75K; provisioned at ~70% utilization ≈ $30–35K; storage 36 TB/yr growth → TTL at 1 year + S3 export

Multi-Region#

OptionWritesConflict HandlingRPOPick When
Single region + PITR / backupsOne regionN/AMinutes (restore from PITR to new table)Default for non-tier-0
Global tables (multi-active, eventual)Any regionLast-writer-wins per item~Seconds of replication lag on region lossRegion-homed users, availability first
Global tables with multi-region strong consistencyAny regionStrongly consistent across regions0When correctness across regions is worth higher write latency
Region-homed tables per cellHome region onlyNone — no shared itemsPer cellData residency, blast-radius cells

🎯 Staff Move: "With global tables I'll home each user to a region and route their writes there. Last-writer-wins is then only a failover-time concern, and I can bound it by what product accepts losing in a regional outage — a few seconds of profile edits, not payments."


Failure Modes & Recovery#

1. Hot Partition Throttling#

Symptom: ProvisionedThroughputExceededException / ThrottlingException for a subset of requests while table-level consumed capacity is far below provisioned (or on-demand isn't near its previous peak). p99 latency climbs as the SDK retries.

Root cause: One partition-key value (a celebrity user, a viral contest, STATUS#OPEN, today's date) exceeds ~1,000 writes/s or ~3,000 reads/s. Adaptive capacity and split-for-heat help when load spreads across sort keys; they can't fix one item or one key taking all writes.

Detection: ThrottledRequests and WriteThrottleEvents / ReadThrottleEvents per table and per GSI; CloudWatch Contributor Insights for DynamoDB (top-N most accessed and most throttled keys); client-side retry counts.

Fix: Immediately: cache reads (DAX or app cache) for the hot key; shed or queue writes to it. Durably: write-shard the key (#<0..N>), bucket by time, move volatile counters to Redis with periodic flush.

Prevention: Load tests with realistic key skew (Zipfian), design review that asks "what's the hottest key's peak rate?" for every PK and GSI PK. Owner: service team; platform provides Contributor Insights by default.

2. GSI Throttling Blocks Base-Table Writes#

Symptom: Writes to the table throttle even though the table has capacity; only writes touching a certain attribute are affected.

Root cause: A GSI is under-provisioned (provisioned mode) or has a hot GSI partition key. DynamoDB applies back-pressure to base writes so the GSI doesn't fall unboundedly behind.

Detection: GSI-level WriteThrottleEvents and OnlineIndexThrottleEvents (during backfill); consumed vs provisioned on the GSI.

Fix: Raise GSI capacity (or move the table to on-demand); change the GSI key to a higher-cardinality one (requires a new GSI + backfill + cutover).

Prevention: GSI capacity in auto-scaling policies alongside the table; alarms per GSI; design GSI keys with the same hot-key review. Owner: service team.

3. Stream Consumer Stuck — Silent Staleness#

Symptom: Search results, aggregates, or downstream caches stop updating for some items; the table itself is healthy.

Root cause: A Lambda consumer throws on a poison record; Lambda retries the batch and blocks that shard until the record expires from the 24 h stream — after which those changes are gone from the stream forever.

Detection: Lambda IteratorAge (alarm at > 60 s for near-real-time projections, page at > 1 h); function error rate; on-failure destination depth.

Fix: Configure BisectBatchOnFunctionError, MaximumRetryAttempts (e.g., 3–5), and an SQS on-failure destination; fix and replay from the DLQ. If records expired, rebuild the projection from a full table export (Export to S3) plus stream catch-up.

Prevention: Idempotent, version-checked consumers; iterator-age alarms mandatory; rebuild path documented. Owner: service team owning the projection.

4. Accidental Data Loss or Bad Deploy Corrupts Items#

Symptom: A bug writes wrong values to millions of items, or a script deletes items it shouldn't.

Root cause: No bounded blast radius on bulk jobs; overly broad IAM permissions; no conditional checks on migrations.

Detection: Business-level anomaly alerts (e.g., orders with total = 0), ConsumedWriteCapacityUnits spike from an unexpected principal, CloudTrail data events.

Fix: Point-in-time recovery restores to any second in the last 35 days — into a new table. Then reconcile: diff the restored table against live, and repair affected items with conditional writes. A full restore of hundreds of GB can take hours; plan the switchover or repair job.

Prevention: PITR enabled on every production table; deletion protection on tables; least-privilege IAM (no dynamodb:* for app roles); bulk jobs rate-limited and dry-run first. Owner: service team; platform enforces PITR and deletion protection via policy.

5. Throttling Storm at Launch or Traffic Spike#

Symptom: A marketing event drives traffic from 5K to 60K writes/s in 2 minutes; an on-demand table throttles above ~2× previous peak; provisioned auto-scaling lags by 5–15 minutes; retries amplify load.

Root cause: Capacity mode assumptions not matched to the traffic shape; no pre-warming; stacked retries.

Detection: ThrottledRequests > 0 sustained, request count ≫ user traffic, SuccessfulRequestLatency p99.

Fix: Temporarily set provisioned capacity (or warm throughput) to the peak; enable request-level load shedding upstream; cap retries.

Prevention: Launch checklist: pre-warm, quota increases, load test at 1.5× peak, single retry layer with jitter. Owner: service team with platform/launch review.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Hot partitionContributor Insights top-throttled keysItems sharing the hot keyCache, write-shard, bucketService team
GSI back-pressureGSI WriteThrottleEventsAll writes touching the GSIRaise GSI capacity, rekeyService team
Stream consumer stuckLambda IteratorAgeDerived views go stale; data loss after 24 hBisect, DLQ, rebuild from exportProjection owner
Bad write / deleteBusiness anomaly alerts, CloudTrailAffected itemsPITR restore to new table + repairService team
Launch spikeThrottles + retry amplificationWhole tablePre-warm, shed load, cap retriesService + launch review
Regional impairmentAWS Health, error rates by regionAll tables in regionGlobal tables failover or restore elsewherePlatform + service
Runaway costDaily spend anomaly, consumed units per requestBudgetKill Scan jobs, fix projections, provisioned modeService team + FinOps
Diagram: Operational Reality Matrix

When to Use vs. Alternatives#

RequirementPickWhyNot DynamoDB Because
AWS-native, known access patterns, zero opsDynamoDBManaged, scales with keys—
Spiky or unpredictable traffic, small teamDynamoDB on-demandPay per request, no capacity planning—
Session, cart, idempotency, metadata at scaleDynamoDBKey lookups, TTL, conditional writes—
Evolving queries, joins, reportingPostgreSQLPlanner, SQLEvery new query needs an index/backfill
Full-text search, facetsElasticsearch via streamsInverted indexNo text search
Sub-ms hot reads with rich structuresRedisRAM, data structuresms latency; DAX is item/query cache only
Multi-cloud or on-premCassandra / ScyllaDBPortableAWS-only
Very high steady throughput where unit cost dominatesCassandra / ScyllaDB (with ops team)Lower $/op at scalePer-request pricing
Analytics over the full datasetExport to S3 → Athena / warehouseColumnar scansScans are costly and slow
Event log with long replayKafkaRetention, offsetsStreams keep 24 h

When NOT to use DynamoDB: when access patterns are unknown or change weekly; when most queries are aggregates or ad-hoc filters; when items are large and frequently updated; when you need portability off AWS; when a single logical key must absorb tens of thousands of writes/s and can't be sharded.


Operational Concerns#

The Settings Every Production Table Should Have#

Point-in-time recovery        ON (35-day window) — restores into a new table
Deletion protection           ON
Encryption                    AWS-owned or customer-managed KMS key (per compliance)
Capacity                      on-demand at launch; provisioned + auto-scaling (target 70%) once steady
Contributor Insights          ON for tables with hot-key risk
TTL attribute                 set for sessions, idempotency keys, ephemeral items
Streams                       NEW_AND_OLD_IMAGES if projections or audit need them
Alarms                        throttles (table + each GSI), SystemErrors, IteratorAge, spend anomaly
IAM                           least privilege per table and action; no Scan/DeleteTable for app roles
Tags                          owner, tier, cost-center — required for chargeback

Schema Evolution#

DynamoDB is schemaless per item, but your access patterns are schema. Evolution patterns:

  1. New attribute: write it going forward; readers default when absent; backfill lazily on read or with a rate-limited job (parallel Scan segments at a fixed WCU budget).
  2. New access pattern: add a GSI (backfills online — watch OnlineIndexPercentageProgress and GSI throttling), or a stream-maintained projection item.
  3. Key change: new table or new key items, dual-write, backfill, switch reads, stop old writes — a multi-week migration for large tables. This is why the access-pattern review happens first.

Version your items (schema_version attribute) so readers can handle multiple shapes during migrations.

Backups, Restores, and Exports#

MechanismWhat It DoesTime / Cost Notes
PITRContinuous backups, restore to any second in 35 daysRestore creates a new table; large tables take hours
On-demand backupsFull snapshots retained until deletedFor compliance retention and pre-migration safety
Export to S3Full or incremental export without consuming table capacityFeeds analytics and projection rebuilds
Import from S3Creates a new table from S3 dataInitial loads and rebuilds without write-capacity cost

The Dashboard the On-Call Actually Uses#

MetricHealthyPage
ThrottledRequests (table + GSIs)0Sustained > 0 for 5 min on tier-0
SuccessfulRequestLatency p99 (GetItem/Query)< 10–20 ms> 50 ms for 10 min
SystemErrors0Any sustained
Lambda IteratorAge on streams< 10 s> 15 min
Consumed / provisioned (provisioned mode)40–70%> 90%
Daily cost vs 7-day baseline± 20%> 50% above
ReplicationLatency (global tables)< 2 s> 60 s

What the On-Call Actually Does#

  1. Throttling? Check table vs GSI vs single key (Contributor Insights) — the fix differs for each.
  2. Check whether request volume matches user traffic — retry storms masquerade as load.
  3. Check streams IteratorAge before blaming the table for "missing" data in downstream systems.
  4. For corrupted data: restore PITR to a new table and repair; never overwrite the live table blindly.
  5. For cost spikes: find the principal and operation (CloudTrail, per-operation metrics) — usually a Scan job or a projection bug looping.

Interview Application — Staff-Level Plays#

Which Case Studies Use DynamoDB#

Case StudyHow DynamoDB Is UsedKey Pattern
URL ShortenerCode → URL mapping at massive read scalePK = CODE#<code>, conditional put for uniqueness, DAX/CDN for hot codes
Chat MessagingMessage storagePK = CONV#<id>#<bucket>, SK = <ts>#<msg_id>, streams for fan-out
Payment ProcessingIdempotency keys and payment stateConditional writes, transactions, TTL on idempotency items
Reservation SystemsHolds and bookingsConditional puts with TTL holds, transactions for multi-seat
Leaderboard & CountingDurable countersSharded atomic counters, stream-driven aggregates
Distributed Job SchedulerJob definitions and run stateTime-bucketed sparse GSI for due jobs
Database SelectionThe managed key-value optionAccess-pattern-first decision

Every System Design Question Has a DynamoDB Moment#

  • Rate limiter: durable per-tenant quotas in DynamoDB, hot counters in Redis — DynamoDB's 1,000 writes/s per key rules it out as the per-request counter.
  • Ticketing: PutItem on SEAT#<event>#<seat> with attribute_not_exists and a TTL hold — double booking is impossible by construction.
  • Notifications: user preferences as items under USER#<id>; the delivery log time-bucketed with a 30-day TTL.
  • File sync: file metadata under PK = USER#<id>, SK = PATH#<path> for directory listing by begins_with; content in S3.

L5 → L6 → L7 Responses#

ScenarioSenior (L5)Staff (L6)Principal (L7)
"Design the schema""Users table and orders table, user_id as PK.""Access-pattern table first: six patterns, five on the base table via CUST# item collections, one sparse GSI for open orders; conditional writes for invariants.""Access-pattern tables are a design-review artifact every team files; the platform checks for Scans and low-cardinality keys before launch."
"Handle a viral item""DynamoDB auto-scales.""Auto-scaling doesn't fix one key above ~1,000 writes/s. Write-shard into 20 keys, cache reads with DAX, sum shards on read.""Hot-key incidents recur across teams; Contributor Insights and a skewed load-test harness become defaults."
"Add a new query""Add a filter expression.""Filters don't reduce cost. New GSI, sparse and INCLUDE-projected, backfilled online; or a streams projection if it's analytic.""If new queries arrive monthly, this data may belong in Postgres or behind a search projection — I'd revisit the store choice, not keep adding GSIs."
"Cost is growing""Switch to provisioned mode.""Provisioned with auto-scaling at 70% cuts ~50–60% here; also shrink items, project fewer attributes, kill the nightly Scan.""DynamoDB is 18% of our AWS bill. Chargeback per table owner, reserved capacity for the baseline, and a quarterly review of the top 10 tables by $/request."
"Go multi-region""Turn on global tables.""Global tables with region-homed writes; LWW only matters during failover; measure ReplicationLatency; idempotency keys stay region-local.""Which tables need multi-region at all? Each costs ~2× writes. I'd tier tables by RTO/RPO and fund multi-region only for tier-0."
Why "Handle a viral item" separates levels

"DynamoDB auto-scales" is what the marketing says and is true at the table level, so it's a reasonable L5 answer. The Staff answer knows the per-partition and per-item ceilings and that auto-scaling operates on the table, not the key — then applies a specific pattern with its read-side cost. The Principal answer notices the same incident will happen to the next team and changes the defaults and test harness.

Why "Add a new query" separates levels

Filter expressions look like SQL WHERE clauses, so reaching for them is natural. The Staff-level signal is knowing they're billed on bytes read before filtering and choosing a GSI or projection instead. The Principal signal is reading the trend: a stream of new query shapes means the data model and the store are mismatched, and the right move may be a different system, not another index.

The Staff DynamoDB Checklist#

  1. List the access patterns with rates: "Six patterns; the hottest is 20K reads/s on customer profile."
  2. Map each to a key: "Profile and orders share CUST#<id>; order details are an item collection under ORDER#<id>; one sparse GSI."
  3. Check the hottest key: "Peak per key is ~50 writes/s — far under 1,000. The live-event channel is the exception; I'll bucket it by minute."
  4. State consistency and invariants: "Eventually consistent reads for display; strong reads before read-modify-write; conditional puts for uniqueness and idempotency."
  5. Price it: "~2K writes/s and 10K reads/s — ~$5K/month on-demand, ~$2K provisioned. Start on-demand, switch after 30 days."
  6. Plan the escape hatches: "Streams to OpenSearch for search, Export to S3 for analytics, PITR on, deletion protection on."

🎯 Staff Insight: What NOT to use DynamoDB for: ad-hoc analytics, relational reporting, full-text search, or any workload whose query shapes you can't list today. "DynamoDB serves the known paths; streams feed everything else" is the sentence that shows you know the contract.


The Principal Lens#

Why L7 Sees This Problem Differently#

At Staff level, DynamoDB is a schema to model and a table to protect. At Principal level, it's a pricing model the whole org is exposed to and an AWS commitment that shapes the next five years. Because there are no servers, there's no natural checkpoint where someone asks "is this still the right store?" — tables accumulate GSIs, Scans, and oversized items, and the bill grows linearly with every modeling shortcut. The L7 job is to make access-pattern design and cost visible, to decide deliberately which workloads belong in DynamoDB versus Postgres or a search projection, and to own the lock-in decision rather than inherit it.

The Org-Level Fault Line#

DynamoDB as the default datastore vs Postgres as the default — and who owns the call.

OptionWhat WorksWhat BreaksWho Pays
DynamoDB default for all servicesZero ops, uniform scaling, tiny platform teamRelational workloads contorted into single tables; analytics via expensive exports; deep AWS lock-inProduct teams on every new query; finance on request-unit bills
Postgres default for all servicesFlexible queries, portableOps burden, sharding cliff for the highest-scale servicesDB platform headcount; teams that hit write ceilings
Decision framework per workload (DynamoDB for known-pattern, high-scale, key-value; Postgres for relational and evolving)Right tool, reviewed onceRequires a review step and two paved roadsArchitecture review time; two platforms to maintain

Cost Model#

Assumptions: us-east-1 approximate pricing — on-demand ~$0.625 per million write units and ~$0.125 per million read units; provisioned ~$0.00065 per WCU-hour and ~$0.00013 per RCU-hour; storage ~$0.25/GB-month; PITR ~$0.20/GB-month; global table replicated writes billed per region. Items ≤ 1 KB, reads eventually consistent. Engineer ~$25K/month loaded.

ScaleTraffic / DataInfra $/monthPeopleOn-call LoadDominant Risk
Startup — one service200 writes/s, 1K reads/s avg; 50 GBOn-demand ~$500–600~0.05 FTE (no DB ops)Near zeroWrong access patterns baked in early
Growth — 15 services20K writes/s, 100K reads/s; 5 TBOn-demand ~$50K; provisioned + reserved ~$15–20K~0.5 FTE platform + review time1–3 pages/month (throttles, stream lag)Cost growth from GSIs, Scans, big items
Enterprise — 200 tables, 3 regions for tier-0300K writes/s, 2M reads/s; 200 TB~$200–400K (reserved baseline + on-demand spikes + global table replication)2–3 platform/FinOps engineers5–10 pages/month fleet-wideCorrelated AWS dependency; lock-in; cost opacity across 200 tables

Enterprise levers: reserved capacity for the steady baseline (often the largest single saving), Standard-IA table class for large, rarely read tables, trimming item size and GSI projections (linear savings), and moving analytics off Scan to S3 exports.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversal CostWhy
Capacity modeTwo-wayMinutes (switch limits apply)Billing setting
Adding a GSITwo-wayDelete GSI; backfill cost onlyOnline build
Partition key designOne-wayNew table + dual-write + backfill: weeks to monthsEvery item and query depends on it
Single-table vs multi-tableOne-way-ishRe-model and migrateClients encode item-type conventions
Global tables with LWW semanticsMostly one-wayRe-architect writes for homingApp correctness assumes the conflict model
Choosing DynamoDB at allOne-way at scaleRe-platform to Cassandra/Postgres: many engineer-quartersProprietary API, data model, operational tooling

The Standard I'd Write#

RFC: DynamoDB Usage Standard (v1)

Scope: All DynamoDB tables in production accounts.

MUST

  • File an access-pattern table (pattern, rate, key condition, index) at design review; update it when patterns change.
  • Enable PITR and deletion protection; tag every table with owner, tier, and cost center.
  • Alarm on throttles for the table and every GSI, on SystemErrors, and on stream IteratorAge for every consumer.
  • Stream consumers use bisect-on-error, bounded retries, and an on-failure destination; projections are rebuildable from Export to S3.
  • No Scan in request-serving code paths; analytics use exports.

SHOULD

  • Launch on on-demand; move to provisioned with auto-scaling when 30 days show a predictable baseline.
  • Keep items under ~10 KB; split volatile attributes into separate items.
  • Use global tables only for tier-0 workloads with region-homed writes.

Exceptions: Architecture review approval; time-boxed to 2 quarters.

Success metrics: 100% of tables tagged with PITR on; hot-key incidents down 50% in a year; DynamoDB $ per 1M requests down 25%; zero projection data loss from expired streams.

What I'd Tell the VP#

DynamoDB lets us run very large systems with almost no database operations staff, which is a real advantage we should keep. The risks are different from a traditional database: a design mistake shows up as either a slowdown for our most popular items or a bill that grows faster than traffic. I'm proposing a lightweight design review for new tables, standard alarms and backups by policy, and per-team cost visibility. We expect to save 30–40% of DynamoDB spend through reserved capacity and design fixes, and to prevent the hot-item outages we've had around launches. I'd also keep relational and reporting workloads on Postgres so we're not paying DynamoDB prices to force-fit them.

Principal Interview Signals#

SignalWhat It Sounds Like
Prices every modeling choice"Projecting ALL on two GSIs triples our write bill — about $20K a month here. INCLUDE with four fields cuts that to ~$8K."
Owns the lock-in decision"DynamoDB is a one-way door onto AWS. That's fine for these three services; for the data platform I want portability."
Workload-level store policy"Known patterns and extreme scale go to DynamoDB; evolving, relational, or reporting-heavy workloads go to Postgres. The review decides, not the team's familiarity."
Correlated-failure awareness"Forty tier-0 tables in one region share a dependency. Multi-region for tier-0 only, tested twice a year."
Knows when not to standardize"I won't mandate single-table design. It's an optimization justified by a specific hot path, not a house style."

Staff answers that L7 interviewers find insufficient:

  • "Write-shard the hot key into 20 partitions." — Correct fix; nothing about why the next team won't hit the same thing.
  • "Switch to provisioned to save money." — A tactical saving; no chargeback, no ownership, no view of the fleet's spend.
  • "Enable global tables for DR." — Doubles write cost fleet-wide without tiering which tables actually need it.

🧭 Principal Move: "The schema decides whether this table is cheap; the review process decides whether the next 200 are. I'd spend the first week on the access-pattern template and cost dashboard, because that's where the org's DynamoDB outcomes are actually set."


In the Wild#

Amazon — Prime Day on DynamoDB#

AWS publishes Prime Day statistics each year; DynamoDB has repeatedly been reported serving peaks above 100 million requests per second across Amazon's retail systems (including Alexa, Amazon.com sites, and fulfillment centers) with single-digit-millisecond latency. The 2022 USENIX ATC paper on DynamoDB describes the architecture behind this: per-partition Multi-Paxos replication, request routers, and admission control that evolved from per-partition provisioning to adaptive and global capacity management.

Staff insight: Even at that scale, the model is unchanged — well-distributed partition keys and known access patterns. The system scales because the data model lets it, not because it's managed.

Zoom — Scaling Through the 2020 Surge#

Zoom publicly described (in AWS case material) relying on DynamoDB among other AWS services as daily meeting participants grew from ~10 million to ~300 million in early 2020. A managed store that scales without capacity migrations let a small team absorb an unprecedented, unplanned surge.

Staff insight: On-demand, key-value workloads are where "no servers to scale" pays off most — the traffic shape was unknowable in advance. That's the argument to make when an interviewer's scenario includes unpredictable growth.

Snap — Moving Core Storage onto DynamoDB#

Snap has spoken publicly (at AWS re:Invent) about migrating significant parts of Snapchat's backend storage, including messaging-related data, onto DynamoDB as part of a move to AWS, citing scale and reduced operational burden.

Staff insight: Migrations like this are one-way doors. The Staff framing is the tradeoff made explicitly: portability and per-request pricing exchanged for operational simplicity at very large scale.


Practice Drill#

Prompt: "Your team's DynamoDB bill tripled in six months while traffic grew 40%. The ops team also reports intermittent throttling during evening peaks, even though table-level consumed capacity is only 35% of provisioned. You own the service. What do you do?"

Staff Answer

These are two problems with probably related causes. Throttling at 35% table utilization means per-key or per-GSI heat, not table capacity. I'd enable Contributor Insights to get the top throttled keys, and check GSI throttle metrics separately — GSI back-pressure throttles base writes. Likely culprits: a low-cardinality GSI key (STATUS#ACTIVE), a date-based key, or a popular entity. Fix: write-shard that key (N ≈ peak rate ÷ 500), or re-key the GSI with a shard suffix and query shards in parallel. Cost tripling vs 40% traffic means cost per request went up about 2×. I'd break spend down by operation and table: common causes are a new GSI projecting ALL, items growing (a list attribute appended forever — each write billed on full item size), a Scan-based job or a FilterExpression-heavy query, or strong reads where eventual reads would do. Fixes: INCLUDE projections, split growing attributes into separate items, move the job to Export to S3 + Athena, eventual reads for display. Then commercial levers: provisioned capacity with auto-scaling at 70% for the steady baseline, reserved capacity once stable. Metrics: WCU per write and RCU per read (target back to ~1–2), throttles by key, $ per 1M requests. I'd expect 50–60% savings and zero evening throttles within two sprints.

Why this is L6:

  • Reads "throttling at 35% utilization" correctly as key or GSI heat and uses the right tool to find it.
  • Frames cost as cost-per-request and decomposes it into the model decisions that drive it.
  • Sequences engineering fixes before billing levers and names the metrics that prove success.

What L7 adds:

  • Treats the regression as a missing guardrail: access-pattern review and per-table cost dashboards so growth in units per request is caught in weeks, not six months.
  • Introduces chargeback so each team sees its own table costs.
  • Uses the moment to reassess whether the new query shapes driving the Scans belong in a search or analytics projection instead.

Quick Reference Card#

Partition limits:  ~10 GB, 3,000 RCU/s, 1,000 WCU/s per partition; item ≤ 400 KB
Units:             1 WCU = 1 KB write; 1 RCU = 4 KB strong read (0.5 eventual, 2 transactional)
Pricing (approx):  on-demand $0.625/M writes, $0.125/M reads; storage $0.25/GB-month
Provisioned:       ~3–7× cheaper at steady high utilization; auto-scaling target ~70%
On-demand burst:   instant up to 2× previous peak; pre-warm for bigger events
Query:             PK equality + SK condition; 1 MB per page; filters don't reduce cost
GSIs:              eventually consistent; own capacity; throttles back-pressure base writes
Transactions:      up to 100 items / 4 MB, 2× cost, single region
Streams:           24 h retention, per-item order; Lambda IteratorAge is the key alarm
TTL:               deletes within days, not seconds — filter expired items on read
PITR:              35 days, restores into a new table
Global tables:     multi-active, typically sub-second replication, LWW — home writes by region
Patterns:          access-pattern table, item collections, sparse GSI, write sharding,
                   conditional puts for uniqueness/idempotency, streams → projections
Red flags:         Scan on hot path, low-cardinality PK, big volatile items, ALL projections
                   everywhere, TTL as scheduler, stacked retries, no PITR
  1. Loading the index…