Technologies referenced in this case study: PostgreSQL · Redis · DynamoDB · Apache Kafka · Flink & Stream Processing · API Gateways
How to Use This Case Study#
Organized for interview use first, reference second. Distributed Coordination introduces idempotency keys and the outbox as tools; Payment Processing applies them to one PSP integration. This case study goes underneath both: the exact semantics of a key, the record's lifecycle, concurrent duplicates, what to replay, how idempotency survives crossing three service boundaries and a message broker, and how to make it the org's default rather than each team's project.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines → Drills 1–4 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Sections 3–4 → Deep Dives 2 and 4 → Appendices A–C |
| Deep Dive | 3+ hrs | Everything, including the Principal Lens and all appendices |
What is Idempotency? — Why interviewers pick this topic
An operation is idempotent if performing it more than once has the same effect as performing it once. In distributed systems that property is not a nicety — it is the only thing that makes retries safe, and retries are the only thing that makes unreliable networks usable. Every timeout is a question with no answer: did it happen? Idempotency lets the client ask again without making it happen twice.
Before vs After — the mobile checkout timeout:
Without idempotency:
t=0: User taps "Place order" on a train; request reaches server
t=+1.2s: Server commits order #A, charges card, enqueues fulfilment
t=+1.3s: Response lost — tunnel. App shows spinner, then "Network error"
t=+4s: App auto-retries (3 attempts policy) → server creates order #B, charges again
t=+9s: User sees "Network error" again, taps "Place order" manually → order #C
t=+2 days: Three parcels arrive. Two chargebacks. Support ticket. App-store review: 1 star.
With an idempotency key per logical operation:
t=0: App generates key k=0190f3…(UUIDv7) when the user taps, persists it locally
t=+1.2s: Server: INSERT key k (in_progress) → commit order #A + charge + outbox → key k completed
t=+1.3s: Response lost
t=+4s: Retry with SAME key k → server finds completed record → replays order #A response
t=+9s: User's manual tap reuses pending key k (button bound to the in-flight operation)
Result: One order, one charge, one parcel. idempotency.replay_total += 2.
Why interviewers reach for this question: "Use an idempotency key" is a one-liner every Senior candidate knows. The level signal lives in the next ten questions: who generates the key, what it is scoped to, how long it lives, what happens when two copies arrive in the same millisecond, what gets replayed when the first attempt failed halfway, how the guarantee survives a downstream call and a Kafka hop, and who owns the platform that makes all of it the default. It is a correctness question with an operational tail and an organizational root.
Mechanics Refresher: The Toolkit
| Mechanism | How It Works | Pros | Cons |
|---|---|---|---|
| Natural idempotency | Operation is a set, not a delta (PUT status=shipped, DELETE) | Free; no state | Most business operations are creates or increments |
| Conditional write | UPDATE … WHERE version = :v / PutItem if attribute_not_exists | No extra table; strong | Caller needs the version; doesn't replay a response |
| Idempotency key + record | Client sends key; server stores key → (fingerprint, status, response) | Generic; replays exact response | Storage, TTL, concurrency and scope must be designed |
| Dedup / inbox table (consumer) | INSERT processed(message_id) in the same transaction as the effect | Exactly-once effect per message | Table growth; window must exceed redelivery horizon |
| Transactional outbox (producer) | Write state + event row in one local transaction; relay publishes | No dual-write gap | At-least-once publish → consumers must dedupe |
| Broker-level dedup | Kafka idempotent producer (PID + sequence); SQS FIFO 5-min dedup window | Removes producer-retry duplicates inside the broker | Says nothing about consumer side effects outside the broker |
| Transactional sink | Stream processor commits output and offsets atomically (Kafka transactions, Flink 2PC sinks) | End-to-end exactly-once within the pipeline | Only for sinks that participate; latency of checkpoint interval |
For most production systems: client-generated idempotency keys on every mutating public API, stored in the same database transaction as the effect; derived keys propagated to every downstream call; transactional outbox for events; consumers that dedupe in the same transaction as their effect or apply versioned writes. Exactly-once delivery does not exist across a network. Exactly-once effect is something you build out of at-least-once delivery and idempotent processing — and you are responsible for every place the chain breaks.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the intents and fault lines below.
What This Interview Actually Tests#
Idempotency is not a header question. Everyone says "Idempotency-Key."
This is an unknown-outcome question that tests:
- Whether you treat a timeout as a third outcome — not success, not failure — and design for it
- Whether you can define a key's scope, lifetime and meaning precisely, not just name it
- Whether you know what happens when two copies of the same request are in flight at once
- Whether the guarantee survives crossing service boundaries and a message broker, or quietly stops at the first hop
- Whether you'd make it a platform default so 40 teams don't each get it 80% right
The key insight: Exactly-once is not a delivery guarantee you buy; it is an effect guarantee you assemble — at-least-once delivery, plus a durable record of what has been done, written atomically with the doing. Every gap between "the record" and "the effect" is where duplicates come from.
The L5 → L6 → L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Add an Idempotency-Key header and dedupe in Redis" | Asks where duplicates come from (client retries, internal retries, redelivery, replays) and which effects are harmful to repeat | Asks how many teams have built their own dedup layer and what the org's retry contract is, because inconsistent semantics across services is the real bug source |
| Storage | Redis SETNX key with 24h TTL | Key record in the same transaction as the effect; Redis only as a cache in front; TTL ≥ max client retry horizon | Picks one platform store and middleware, prices it, and publishes a per-endpoint durability tier (money vs best-effort) |
| Concurrency | Not mentioned | in_progress state with a lease; concurrent duplicate gets 409 + Retry-After; takeover after lease expiry uses recovery points | Standardizes the status codes and SDK behavior org-wide so every client handles "in progress" identically |
| Failure | Retries with backoff | Timeouts are UNKNOWN; same key on every retry; derived keys downstream; resolver for stuck operations; reconciliation for residue | Sets retry budgets and single-layer retry policy across the call graph; owns the "unknown outcome" metric as an org SLO |
| Messaging | "Kafka exactly-once" | At-least-once publish via outbox; consumers dedupe in-transaction or apply versioned writes; dedup window ≥ replay horizon | Makes inbox/outbox a paved road (library + CDC platform) and forbids dual writes in code review lint |
Why "first move" separates levels
L5: Reaches for the header and a Redis SETNX. That stops the easy case — a client retrying 2 seconds later — and fails on every hard case: Redis failover losing keys, a retry 30 hours later, two copies racing, a key reused for a different operation, a downstream call that isn't covered.
L6: "Duplicates come from four places: client retries after a timeout, our own retries between services, broker redelivery, and replays when someone resets a consumer offset. Each needs a different mechanism, and the property I actually need is that the side effect happens once. I'll design the API layer first and then show how the guarantee carries through the downstream call and the event."
L7: "How many idempotency implementations do we have? If the payments team stores keys for 7 days and the orders team for 1 hour, a client SDK that retries for 24 hours is safe against one and dangerous against the other. The design problem is the contract, not the table."
Why "storage" separates levels
L5: Stores keys in Redis because it's fast and has TTLs. But the key and the effect now live in two systems with no transaction between them: crash after the effect and before SET → duplicate on retry; Redis async-replication failover drops the last second of keys → duplicates for every retry in that window.
L6: Puts the key record in the same database transaction as the business write, so "the key says done" and "the effect happened" are the same fact. Uses Redis only as a read-through cache to short-circuit hot replays. When the effect is external (a PSP, an email provider), splits the handler into recorded phases and passes a derived key downstream.
L7: Recognizes that "same transaction as the effect" means the idempotency store must live with each service's database — so the platform ships a library and schema migration, not a central service. Defines durability tiers so that a "like" button doesn't pay the same cost as a wire transfer.
Why "failure" separates levels
L5: "Retry with exponential backoff and jitter." Correct for reads. For writes without a stable key it is the most common source of duplicates.
L6: Splits outcomes into succeeded / failed-definitively / unknown. Unknown keeps the same key forever, gets retried or resolved by query, and is never shown to the user as "failed." Retries happen at one layer, not at every layer.
L7: Sees that retry policy is an emergent property of the whole call graph: four layers each retrying 3× is an 81× amplifier during an incident. Sets a single-layer retry rule, retry budgets (e.g., ≤ 10% of traffic), deadline propagation, and owns the org-wide "unknown outcome" rate as an SLO.
The Staff Positions#
| Position | Rationale |
|---|---|
| Client generates the key per logical operation, before the first attempt | A key minted on retry is useless; a key minted per attempt dedupes nothing |
| Key record lives in the same transaction as the effect | Otherwise there is a crash window where the effect happened and the record didn't |
| Fingerprint the request; reject key reuse with a different payload | Silent replay of a different operation is worse than an error |
| Concurrent duplicates get 409 + Retry-After, never parallel execution | Two in-flight copies must not both run the handler |
| TTL ≥ the longest any client will retry — and reject expired keys explicitly | A retry after the record expires re-executes silently |
| Exactly-once effect = at-least-once delivery + idempotent processing + reconciliation | Exactly-once delivery does not exist across a network |
| Retry at one layer; propagate deadlines and derived keys downstream | Multi-layer retries multiply load and duplicate risk |
The Four Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Client retry safety (public mutating API) | Clients you don't control, unreliable networks, retries up to hours later | Client key + fingerprint + record in same tx + response replay | Key TTL < retry horizon; key reuse across operations | Zero duplicate effects per key within TTL; explicit error after |
| Internal call-chain safety (service-to-service) | Multi-hop RPC with retries at several layers | Derived keys per downstream call; single-layer retry; deadline propagation | Retry amplification; downstream not keyed | Each downstream effect once per parent operation |
| Consumer idempotency (at-least-once messaging) | Redelivery on rebalance, crash, offset reset | Inbox/dedup table in same tx, or versioned upsert | Dedup window shorter than replay horizon | Effect once per message ID forever (or per version) |
| Exactly-once pipelines (stream processing) | High throughput aggregates and sinks | Transactional sink + checkpointed offsets, or idempotent sink keyed on deterministic IDs | Non-transactional sink; non-deterministic IDs on replay | Output equals one processing of each input |
🎯 Staff Move: "I'll focus on the first two — a public mutating API whose handler calls downstream services and emits events — because that's where the unknown-outcome problem bites hardest. Then I'll carry the guarantee through the event to the consumer. Pipeline exactly-once is a real but separate design; I'll touch it at the end."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Same-Transaction vs Separate Store | Key record atomic with the effect (correct, couples to each service's DB) vs a shared fast store (generic, has a crash gap) |
| 2 | Key Scope & Lifetime | Who mints the key, what it is unique within, and how long it is remembered vs storage and the client's retry horizon |
| 3 | Concurrent Duplicates | Reject (409), wait, or let them race into a unique constraint |
| 4 | Replay vs Re-evaluate | Return the stored original response (including errors?) vs return the resource's current state |
| 5 | Generic Middleware vs Domain-Native Idempotency | One layer for every endpoint vs designing each operation to be naturally idempotent with conditional writes |
In the Wild: Real Production Systems#
Why this section belongs here: These are public designs. The tradeoff each one made is the thing to cite.
Stripe — Idempotency Keys as a Public Contract#
Stripe accepts an Idempotency-Key header on POST requests, stores the result of the first request, and returns that same result — status code and body — for subsequent requests with the same key. Its documentation says keys can be pruned after they're at least 24 hours old, that a retry with the same key but different parameters is rejected as an error, and that results are saved only once the endpoint has begun executing — a request that failed validation or collided with a concurrently executing request with the same key does not save a result. Stripe's engineering writing describes the server side as atomic phases with recovery points, so a request interrupted mid-way can be resumed rather than redone.
Staff insight: Stripe's contract has three clauses — same key → same response; different payload → error; concurrent → conflict, nothing saved. Saying all three in an interview is the signal; "we dedupe on the key" says one.
AWS — Client Tokens and "Making Retries Safe"#
Many AWS APIs (EC2 RunInstances, for example) accept a client token. Reusing a token with the same parameters returns the original result; reusing it with different parameters returns an IdempotentParameterMismatch error. The Amazon Builders' Library article on making retries safe with idempotent APIs describes the reasoning: the caller supplies the identifier for the intent, and the service records it atomically with the mutation so retries collapse onto the first execution.
Staff insight: AWS treats the parameter-mismatch error as a first-class part of the API. It is the guard against the most common client bug — reusing a key for a different operation — and it converts a silent wrong answer into a loud one.
Apache Kafka — Idempotent Producer and Transactions#
Kafka's idempotent producer assigns each producer a producer ID and per-partition sequence numbers, so the broker discards duplicates caused by producer retries (enabled by default since Kafka 3.0). Kafka transactions extend this so a consume-transform-produce application can commit its output records and its consumer offsets atomically; downstream consumers in read_committed mode see each result once.
Staff insight: Kafka's exactly-once is real inside Kafka. The moment a consumer writes to Postgres, calls an API or sends an email, the guarantee ends and consumer idempotency begins. Candidates who say "Kafka gives us exactly-once" without that caveat get downleveled.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Idempotency-Key header" | "Who generates it? When? What if the client generates a new one on retry?" | Key semantics, not vocabulary |
| "Store keys in Redis" | "Redis fails over and loses a second of writes. Now what?" | Atomicity of record and effect |
| "TTL of 24 hours" | "A mobile client retries after 30 hours. What happens?" | Lifetime vs retry horizon |
| "We check the key first" | "Two copies arrive 5ms apart. Walk me through it." | Concurrency, in-progress state |
| "We replay the response" | "The first attempt returned 500. Do you replay the 500?" | Replay semantics |
| "The handler calls the inventory service" | "Is that call idempotent? With what key?" | Propagation across boundaries |
| "We publish an event" | "What if the publish succeeds and the commit fails?" | Dual-write, outbox |
| "Kafka exactly-once" | "The consumer sends an email. Exactly once?" | Delivery vs effect |
System Architecture Overview#
Reading the diagram: The idempotency record, the business row and the outbox row live in one database and commit together — that is the whole correctness argument. External calls (inventory, PSP) cannot join that transaction, so each gets a derived key and the handler records a recovery point before and after them. The gateway enforces the presence of keys but does not retry. Events leave via the outbox at-least-once, so every consumer dedupes. The metrics row is how you find out the guarantee is leaking.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This | The L7 Answer — Say This |
|---|---|---|---|
| Key | "Client sends a UUID" | "Per logical operation, minted before the first attempt, persisted by the SDK, scoped to tenant + endpoint, fingerprinted" | "One key contract for every API, enforced by gateway and lint, with SDKs in every language we ship" |
| Storage | "Redis SETNX, 24h" | "Same DB transaction as the effect; Redis as cache only; TTL ≥ retry horizon; expired keys rejected explicitly" | "Durability tiers by endpoint risk; storage cost priced per tier" |
| Concurrency | — | "in_progress with lease; duplicate gets 409 + Retry-After; takeover via recovery points" | "Org-standard status codes so every client behaves the same" |
| Replay | "Return cached response" | "Replay final outcomes, including deterministic errors; never cache 'retry-safe' failures as final" | "Replay semantics published in the API standard" |
| Downstream | "Retry the call" | "Derived key parent:step; single-layer retries; deadlines propagated; timeout = UNKNOWN" | "Retry budgets and single-layer retry rule across the call graph" |
| Events | "Kafka exactly-once" | "Outbox + at-least-once + consumer inbox in same tx; dedup window ≥ replay horizon" | "Outbox/inbox paved road; dual writes blocked in review" |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Stripe key retention | Pruned after ≥ 24h | Industry anchor; your TTL must cover your clients' retry horizon, not Stripe's |
| SQS FIFO dedup interval | 5 minutes | Broker-level dedup windows are short — never your only protection |
| Kafka default retention | 7 days | An offset reset can replay a week — consumer dedup must cover it |
| Retry amplification | 3 attempts × 4 layers = 81× | Why retries belong at one layer |
| Retry budget | ≤ 10% extra traffic | Caps amplification during incidents |
| Key record size | ~200 B without response; 1–4 KB with | Sizing the store |
| Storage at 10M writes/day, 2 KB, 72h | ~60 GB | Trivial in Postgres if partitioned by day and dropped |
| Idempotency check overhead (same DB) | +1 indexed insert, ~0.5–2ms | The tax for correctness on a write path |
| Redis lookup | ~0.3–1ms | Fast — but not atomic with the effect |
| In-progress lease | 2–3× handler p99.9 (e.g., 30s) | Too short → double execution; too long → stuck retries |
| UUIDv4 key collision risk | ~10^-13 at 10^12 keys | Collisions are never the problem; reuse is |
Interview Walkthrough
The most common mistake: Candidates say "idempotency key" in minute three, draw a Redis box, and spend the next twenty minutes on unrelated scaling. The interviewer wanted the twenty minutes on what the key means, what happens to the second copy, and what happens one hop downstream.
Phase 1: Requirements & Framing (2–3 minutes)#
State the functional requirement in one sentence:
"Clients must be able to retry any mutating request after a timeout or error without causing the side effect twice, and the guarantee must hold through our downstream calls and the events we emit."
Then name where duplicates come from and commit:
"Duplicates come from four sources: client retries, our internal retries, broker redelivery, and replays. I'll design for a public POST /orders whose handler reserves inventory, charges a card through a provider, and emits an OrderPlaced event. Clients include mobile apps with offline retry queues — so retries can arrive hours later. Correctness bar: zero duplicate orders, charges or shipments per logical operation."
Commit to the constraint set:
"Keys are generated by the client per logical operation; the key record commits atomically with the order; concurrent duplicates never execute twice; timeouts are unknown outcomes, not failures; and every hop downstream carries a derived key."
🎯 Staff Move: Naming the four duplicate sources up front tells the interviewer you know that the header handles only one of them.
Phase 2: Core Entities & API (1–2 minutes)#
- IdempotencyRecord:
(tenant_id, endpoint, key)PK,fingerprint,status(in_progress/completed),recovery_point,locked_until,response_code,response_body,created_at,expires_at - OutboxEvent:
event_id,aggregate_id,type,payload,created_at,published_at - InboxRecord (consumer side):
(consumer, message_id)PK,processed_at
API contract:
POST /v1/orders
Idempotency-Key: 0190f3a2-7c1e-7b3a-9f21-4d8e6c0b5a11 (required; 16–64 chars)
→ 201 Created first execution, or replay of a completed one (header: Idempotent-Replayed: true)
→ 409 Conflict same key still in progress (Retry-After: 1)
→ 422 Unprocessable same key, different payload (error: idempotency_key_reused)
→ 400 Bad Request missing or malformed key on an endpoint that requires one
→ 410 Gone / 422 key older than retention (error: idempotency_key_expired)
🎯 Staff Move: "The status codes are part of the contract. 409 tells the client 'wait and retry the same key'; 422 tells it 'you have a bug — you reused a key for a different operation.' If those collapse into one generic error, every SDK handles them differently." (The IETF HTTPAPI working group's Idempotency-Key draft uses this same 400/409/422 split.)
Phase 3: High-Level Architecture (≤5 minutes)#
Walk it in 90 seconds:
- Middleware inserts the key record as
in_progress— the unique constraint on(tenant, endpoint, key)is what makes "first one wins" atomic. - The handler runs in phases. Local writes commit with a recovery point. External calls carry a key derived from the parent (
k:charge), so the provider dedupes our retries. - The final phase commits the business state, the outbox event and the completed key record (with the response) in one transaction.
- A retry finds
completedand replays the stored response byte-for-byte. - The outbox relay publishes the event at-least-once; consumers dedupe.
🎯 Staff Move: "That's the happy path and the simple retry. The interesting parts are the second copy arriving while the first is still running, the first attempt dying halfway after charging the card, and what 'the same request' even means. Let me take those in order."
Phase 4: Transition to Depth (1 minute)#
"Three places to go deep: the key's semantics — scope, lifetime, fingerprint, and what happens after it expires; the in-flight and crashed-halfway cases — which is where unknown outcomes live; and carrying the guarantee across service and broker boundaries. Which would you like first?"
If the interviewer has no preference, lead with crashed-halfway: it forces the recovery-point design and the unknown-outcome discussion, which is the deepest Staff-level signal.
Phase 5: Deep Dives (25–30 minutes)#
Deep dive A — Key semantics (6–8 min). Client-generated per logical operation, minted before the first attempt and persisted until a terminal outcome. Scope: (tenant, endpoint, key) — never global, or one tenant's key can collide with or probe another's. Fingerprint: SHA-256 of canonicalized method + path + body (minus volatile fields). TTL ≥ max client retry horizon (here 72h for the offline queue). Expired keys: if the key is a UUIDv7, its embedded timestamp lets the server reject a too-old key explicitly instead of re-executing.
Deep dive B — Concurrency and crash recovery (8–10 min). Second copy arrives while first is in_progress → 409 + Retry-After. Lease (locked_until) on the in-progress record. If the first handler crashed, the next retry after lease expiry takes over and resumes from the last recovery point — re-calling the provider with the same derived key, which returns the original charge. Fencing via a version column so a slow original can't commit after a takeover.
Deep dive C — Replay semantics (4–5 min). Replay completed outcomes, including deterministic 4xx (card declined is an outcome). Do not store transient failures before any side effect (validation error, 503 from our own overload) — release the key so a retry can execute. 5xx after a possible side effect is not final: the record stays in_progress with a recovery point, and resolution happens via takeover or a background resolver.
Deep dive D — Across boundaries (6–8 min). Derived keys for every downstream call. Retries at exactly one layer (the SDK), with deadline propagation so downstream services stop work the caller has abandoned. Outbox for events; consumer inbox in the same transaction as the consumer's effect, or versioned upserts. Dedup window ≥ Kafka retention for consumers that can be replayed.
Deep dive E — Platform (3–5 min). Middleware library + schema migration per service (not a central service, because the record must share the service's transaction); gateway enforces key presence; SDKs in every client language; durability tiers.
Phase 6: Wrap-Up (2–3 minutes)#
"To summarize: client-minted keys per operation, scoped and fingerprinted; the key record commits with the effect; concurrent copies get 409; crashed attempts resume from recovery points with derived downstream keys; timeouts are unknown, never failed; events leave via an outbox and consumers dedupe in-transaction. Day one I'd ship the middleware for the five money-moving endpoints and the SDK change. I'd defer the Redis cache, durability tiers and cross-region key routing until replay volume or a second region justifies them. The metric I'd watch from day one is unknown outcomes older than an hour — that's where real duplicates and real losses hide."
🎯 Staff Move: End on the metric that catches leaks. It shows you expect the design to be imperfect and know where to look.
Common Timing Mistakes#
| Mistake | Time Lost | Fix |
|---|---|---|
| Explaining what idempotency means mathematically | 3–5 min | One sentence; move to unknown outcomes |
| Debating Redis vs DynamoDB for the key store | 5–8 min | "Same DB as the effect"; cache optional |
| Never addressing concurrent duplicates | — (fails probe) | Have the 409 + lease answer ready |
| Treating Kafka EOS as end-to-end | — (fails probe) | "Exactly-once inside Kafka; consumer effects need dedup" |
| Designing payments in depth | 10+ min | Point to Payment Processing; stay generic |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Idempotency is where distributed-systems theory meets an operational reality you can't avoid: networks drop responses, and the client cannot distinguish "never arrived" from "arrived, executed, response lost." Every write API in every company must answer that question, and most answer it inconsistently — one team with a Redis set, another with a DB table, a third with nothing. The problem rewards candidates who think in invariants ("the record and the effect are the same fact"), in failure timelines ("crash after charge, before commit"), and in contracts ("409 means retry the same key"). It also has an obvious organizational dimension: idempotency is only as strong as the weakest hop in the call chain, which makes it a platform problem whether or not anyone admits it.
1.2 The L5 → L6 → L7 Contrast — Visual#
The L5 path handles the retry that arrives two seconds later. The L6 path handles the retry that arrives during execution, after a crash, 30 hours later, or via a broker. The L7 path makes the L6 path the default for 40 teams.
1.3 The Staff Question That Cuts Through Everything#
"If the process crashes right after the side effect and right before recording it, what happens on retry?"
Every idempotency design either has an answer to this or has a duplicate-effect bug. Same-transaction storage answers it for local effects: the record and the effect commit or roll back together. External effects can't join the transaction, so the answer is a derived key passed to the external system (it dedupes our retry) plus a recovery point (we know which phase we reached). If the external system has no idempotency support, the answer is a pre-call record of intent, a lookup-by-reference after the crash, and reconciliation for the residue. Ask this question about every effect in the handler and the design writes itself.
2. Problem Framing & Intent#
2.1 The Four Intents — Explained#
Intent 1: Client retry safety on a public API. Clients you don't control, on networks you don't control, retrying on schedules you don't control. Mobile SDKs with offline queues can retry 1–3 days later. Browser users double-click and open two tabs. Partners run cron jobs that retry failed batches the next morning. The key must be generated by the client (only the client knows that two requests are "the same operation"), and the server's job is to remember long enough and reject misuse loudly.
Intent 2: Internal call-chain safety. Service A calls B calls C. Each hop has a client library with its own retry policy, and a service mesh may add another. Without derived keys, A's single retry becomes two executions in C. Without a single-layer retry policy, a slow C turns one user request into 27–81 requests during an incident. The design here is propagation: one root key from the client, derived deterministically at every hop (root:step), plus deadlines so downstream stops work the caller already gave up on.
Intent 3: Consumer idempotency. Brokers deliver at-least-once. Kafka redelivers on consumer-group rebalance and after a crash before offset commit; an operator resetting offsets replays days. SQS redelivers after visibility timeout. Webhook senders retry for hours. The consumer must record "I processed message M" atomically with its effect — or make the effect naturally idempotent (versioned upsert). The dedup window must outlast the longest possible redelivery, which for a replayable log is the retention period, not "a few minutes."
Intent 4: Exactly-once pipelines. Aggregations and materializations where each input must contribute exactly once to the output. Kafka transactions and Flink's checkpoint-aligned two-phase-commit sinks give this within participating systems. Everywhere else the answer is deterministic output keys (so replays overwrite rather than append) and idempotent sinks. See Stream Processing and Data Pipeline Patterns.
Committing:
🎯 Staff Move: "These four have different owners and different mechanisms. The client-facing API is the hardest because the client is outside our control and retries can arrive days later. I'll design that, carry it through one downstream call and one event, and treat pipeline exactly-once as a separate conversation."
2.2 When NOT to Build Idempotency Keys#
| Situation | Use Instead | Why |
|---|---|---|
Operation is naturally idempotent (PUT full resource, DELETE, set-state) | Nothing extra | Repeating it is already harmless |
Client already holds a version (If-Match: etag) | Conditional write | Optimistic concurrency gives once-only for updates |
Create with a natural unique business key (e.g., external_order_ref unique per tenant) | Unique constraint + return existing | The business key is the idempotency key |
| Duplicates are harmless and cheap (analytics pings, "last seen" updates) | Accept at-least-once | Don't pay 1–2ms and storage for nothing |
| Read-only endpoints | Nothing | Reads are safe to retry by definition |
| Effects that are counted, not applied (metrics) | Approximate aggregation with dedup at query time, or accept error | Exactly-once counting is expensive; ask if ±0.1% is acceptable |
🎯 Staff Move: "If the operation already has a natural key — the partner's own order reference — I'd make that unique per tenant and skip a separate idempotency table. The best idempotency key is one the domain already has."
2.3 What the Interviewer Leaves Underspecified#
| Underspecified | Why It Matters | What to Assume Out Loud |
|---|---|---|
| Client types | Determines retry horizon and who mints keys | "Mobile with offline queue, web, and partners — up to 72h" |
| Side effects in the handler | Local vs external determines atomicity strategy | "One local write, one external charge, one event" |
| Downstream idempotency support | External systems may not dedupe | "Provider supports keys; email vendor doesn't" |
| Duplicate cost | Drives durability tier | "Money and shipments: zero tolerance; notifications: tolerable" |
| Consumer replay | Offset resets replay days | "Kafka retention 7 days; consumers must survive full replay" |
| Multi-region | Same key could hit two regions | "Single-region writes per tenant for now" |
2.4 Precise Terminology#
| Term | Meaning | Common Confusion |
|---|---|---|
| Idempotent operation | f(f(x)) = f(x): repeating has no additional effect | Confused with "safe" (no effect at all) |
| Idempotency key | Client-chosen identifier for one logical operation | Minted per attempt (wrong) or per session (wrong) |
| Scope | The namespace a key is unique within (tenant + endpoint) | Global scope leaks across tenants |
| Fingerprint | Hash of the canonical request, stored with the key | Omitted → key reuse silently replays the wrong response |
| Replay | Returning the stored outcome for a repeat request | Confused with re-executing |
| Unknown outcome | Caller cannot tell whether the effect happened (timeout, crash) | Treated as failure → duplicates on "retry with new key" |
| Exactly-once delivery | Each message delivered once — impossible across an unreliable network | Marketing term for exactly-once processing |
| Exactly-once effect | Each logical operation's effect applied once | The achievable property |
| Outbox | Event row written in the same transaction as state; relayed later | Thought to give exactly-once publish (it's at-least-once) |
| Inbox / dedup table | Consumer's record of processed message IDs, same tx as effect | Kept in a separate store → crash gap |
| Recovery point | Durable marker of the last completed phase of a multi-step handler | Missing → crashed attempts restart from scratch |
3. The Five Fault Lines#
3.1 Fault Line 1: Same-Transaction vs Separate Store#
The tension: The only way to make "the key says done" and "the effect is done" the same fact is to commit them together. That couples the idempotency record to each service's database. A shared fast store (Redis, a central DynamoDB table) is generic and easy to adopt — and has a crash window between effect and record.
| Option | Correctness | Latency | Who Pays |
|---|---|---|---|
| Same DB transaction | Exact for local effects | +0.5–2ms (one indexed insert + update) | Each service team (schema, migration) |
| Separate durable store (DynamoDB, separate Postgres) | Crash gap between effect and record | +2–5ms | Customers (rare duplicates); on-call (reconciliation) |
| Redis | Crash gap + loss on async-replica failover + eviction under memory pressure | +0.3–1ms | Customers (duplicates during every failover); finance |
| Same tx + Redis read-through cache | Exact; cache only short-circuits completed replays | ~0.5ms on replay hits | Platform (cache invalidation is trivial: records are immutable once completed) |
Staff default: Same transaction as the effect. Redis only as a cache of completed records (immutable, so caching is safe). For endpoints whose effects are entirely external, the record lives in the service's own DB and recovery points bracket the external calls.
When to deviate: Low-stakes endpoints (likes, follows, preference toggles) where a rare duplicate is harmless — Redis-only is fine and much cheaper to adopt. Make this an explicit tier, not an accident.
🎯 Staff Move: "I'd rather pay 1ms on the write path than explain to finance why every Redis failover creates a few hundred duplicate charges. Redis can cache completed records — those never change — but it can't be the source of truth."
3.2 Fault Line 2: Key Scope & Lifetime#
The tension: Longer retention protects against later retries and costs storage. Narrower scope prevents cross-tenant collisions and probing but needs the server to know the scope at lookup time.
Who generates:
| Generator | Pros | Cons | Who Pays |
|---|---|---|---|
| Client, per logical op | Only the client knows two requests are the same intent | Clients have bugs (per-attempt, per-session keys) | Client developers (must persist keys) |
| Server-derived from payload hash | No client cooperation | Two intentional identical orders collapse into one | Users (legitimate repeat purchases lost) |
| Domain natural key | Already unique; meaningful | Not every operation has one | Nobody — use when available |
Scope: (tenant_id, endpoint, key). Global scope lets tenant A's key collide with tenant B's (vanishingly rare with UUIDs, common with partner-chosen keys like "order-1"), and lets a malicious tenant probe whether a key exists elsewhere. Endpoint scope prevents a key reused across /orders and /refunds from replaying the wrong resource type.
Lifetime:
| Retry Source | Horizon | Implied TTL |
|---|---|---|
| Interactive web/app retries | seconds–minutes | ≥ 1h |
| SDK automatic retries | minutes | ≥ 1h |
| Mobile offline queues | hours–days | 72h–7d |
| Partner batch re-runs | next day | ≥ 48h |
| Consumer replay (offset reset) | Kafka retention | = retention (for inbox) |
The expiry trap: After the record expires, a late retry executes again. Staff fix: make expiry explicit. If keys are UUIDv7 (see ID Generation), the server can read the key's embedded timestamp and reject any key older than the TTL with idempotency_key_expired instead of executing it. Alternatively the SDK refuses to retry past TTL − margin and surfaces the operation as unknown.
Staff default: Client-minted UUIDv7 per logical operation; scope (tenant, endpoint, key); TTL 72h for public APIs with mobile clients (24h if none); expired keys rejected, never re-executed.
3.3 Fault Line 3: Concurrent Duplicates#
The tension: Two copies of the same request arrive 5ms apart (double-click, SDK retry racing a slow original, load balancer retry). Both pass a naive "check key, not found" read.
| Strategy | Behavior | Who Pays |
|---|---|---|
| Check-then-act (read, then insert) | Both see "not found," both execute | Customers (duplicates) — the L5 bug |
Insert in_progress first; second gets 409 | Exactly one executes; second told to retry | Clients (must honor Retry-After) |
| Insert first; second blocks and polls up to N seconds | Second returns the real result if first finishes quickly | Server (held connections, thread exhaustion under storms) |
| Let both execute into a unique business constraint | DB rejects the second | Developers (only works if the effect has a natural unique key; external effects still duplicate) |
Staff default: Atomic insert of the in_progress record (unique constraint) before any work. A concurrent duplicate gets 409 Conflict with Retry-After: 1. Optionally block for ≤ 1–2s on a cheap notification (e.g., LISTEN/NOTIFY or a short poll) before returning 409, for better UX on double-clicks. Never hold connections longer.
The crash sub-case: The first copy dies while in_progress. Without a lease, every retry gets 409 forever. With locked_until = now + lease (2–3× handler p99.9, e.g., 30s), the next retry after expiry takes over by atomically updating locked_until and incrementing version, then resumes from the recorded recovery_point. The original, if merely slow, will fail its final commit because its version no longer matches — a fencing token for the handler.
3.4 Fault Line 4: Replay vs Re-evaluate#
The tension: On a repeat request, return the original response (byte-identical) or the current state of the resource?
| Option | Pros | Cons | Who Pays |
|---|---|---|---|
| Replay stored response | Client sees exactly what it would have seen; simple contract | Storage for bodies (1–4 KB); response may be stale (order since shipped) | Storage |
| Re-evaluate (return current resource) | Fresh; no body storage | Different responses for the same request confuse clients; errors can't be replayed | Client developers |
Which outcomes to store as final:
| Outcome | Store as final? | Reason |
|---|---|---|
| 2xx success | Yes | The canonical case |
| Deterministic business error after execution (card declined, insufficient stock) | Yes | It is the outcome; re-executing might succeed and surprise the client |
| Validation error before any side effect (400) | No — release key | Nothing happened. A corrected request changes the fingerprint, so the client sends it with a new key; releasing avoids a 422 for an operation that never ran |
| 429 / 503 from our own load shedding before execution | No — release key | Nothing happened; retry should execute |
| 5xx / timeout after a side effect may have happened | No — keep in_progress | Outcome unknown; resolve by takeover or resolver, never by marking final |
Staff default: Replay stored responses for completed outcomes with a header (Idempotent-Replayed: true). Release the key only when you can prove no side effect occurred. Everything else stays in_progress until resolved.
🎯 Staff Move: "The dangerous response is the 500 after we called the provider. If I store that 500 as final, the client is told 'failed' for a charge that succeeded. If I release the key, the retry charges again. Neither — the record stays in progress and the retry resumes from the recovery point."
3.5 Fault Line 5: Generic Middleware vs Domain-Native Idempotency#
The tension: A middleware layer gives every endpoint the same guarantees with zero domain thought. Domain-native idempotency — natural keys, conditional writes, state machines — gives stronger, cheaper guarantees but requires every team to think.
| Approach | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Middleware only | Uniform contract; fast adoption | Can't help with effects inside the handler that the middleware doesn't see (downstream calls, events); teams think they're "covered" | Customers (duplicates in downstream effects) |
| Domain-native only | Precise; no extra storage when natural keys exist | Every team reinvents; inconsistent client contract | Client developers (N contracts) |
| Middleware for the contract + domain-native inside (Staff default) | Uniform external contract; handlers use state machines and derived keys | Needs a library that exposes recovery points and derived keys, not just a filter | Platform (richer library) |
Staff default: The middleware owns the external contract (key validation, scope, fingerprint, 409/422, replay). The handler owns correctness of its effects (state machine transitions with conditional writes, derived keys for downstream, outbox for events). A middleware that wraps a handler doing dual writes gives a false sense of safety.
4. Failure Modes & Operational Reality#
Idempotency failures are almost always silent: a duplicate charge, a second shipment, an event counted twice — or, worse, a legitimate second operation silently swallowed as a "duplicate." Nobody is paged; a customer or a reconciliation job finds it. A Staff design makes each class of leak measurable.
4.1 Key Retention Shorter Than the Retry Horizon#
t=0: Mobile app v5 adds an offline queue: failed POSTs retried on reconnect, up to 7 days
t=0: Order service idempotency TTL: 24h (copied from a blog post)
t=+3 days: User on a long-haul trip; 41 queued "place order" requests replay on landing
t=+3 days: Keys expired → server executes all 41 as new orders
t=+3 days: Fleet-wide: ~0.3% of mobile orders duplicated during holiday travel week
t=+10 days: Support volume up 18%; refunds $410K; app team and order team each think the other owns it
Detection: idempotency.key_age_at_request_seconds histogram (keys older than TTL on arrival should be ~0); idempotency.expired_key_rejections_total; duplicate-order detector (same user, same cart hash, < 10 min apart, different keys).
Blast radius: Every client whose retry horizon exceeds the server TTL — typically one client platform, all its users.
Mitigation: Raise TTL to 7d for the order endpoint (storage: ~10M orders/day × 2 KB × 7 ≈ 140 GB — affordable). Reject keys older than TTL explicitly using the UUIDv7 timestamp, so the client gets idempotency_key_expired and asks the user rather than silently re-ordering.
Prevention: The TTL and the SDK's max retry horizon are one contract, published together; the SDK enforces retry_horizon ≤ TTL − 1h. A contract test in the mobile CI asserts it.
Owner: API platform owns the published TTL; the mobile team owns the SDK's horizon; both sign the contract. Before the incident there was no owner of the pair — that was the bug.
4.2 Idempotency Store Loses Keys on Failover#
t=0: Keys stored only in Redis (primary + async replica), SET NX with 24h TTL
t=+0s: Primary dies; replica promoted; last ~800ms of writes not replicated
t=+0s: ~2,400 keys for in-flight and just-completed requests lost
t=+1s: Clients that timed out during the failover retry with the same keys
t=+1s: Keys not found → handlers re-execute → ~600 duplicate effects
t=+1 week: Reconciliation flags 600 duplicate charges; refunds issued
Detection: redis.replication_offset_lag at failover time; idempotency.store_failover_total; a post-failover reconciliation job that scans effects created in the failover window ± 5 min for duplicates by (user, amount, cart hash).
Blast radius: Every request in flight during the failover — proportional to traffic × replication lag.
Mitigation: Move the source of truth into the service DB transaction; keep Redis as a cache of completed records only. Until then, run a duplicate scan after every failover.
Prevention: A design review rule: an idempotency store that isn't transactional with the effect must declare its duplicate budget per failover and get sign-off from the effect's business owner.
Owner: The service team owns the store choice; the business owner of the effect (finance for charges) signs off on any non-transactional design.
4.3 Stuck In-Progress Records#
t=0: Handler crashes after charging (recovery point: charged) and before final commit
t=0: Record: status=in_progress, no lease (implementation predates locked_until)
t=+1s..+1h: Every client retry gets 409 Conflict, Retry-After: 1
t=+1h: SDK gives up → app shows "Something went wrong"
t=+1h: Customer was charged; no order exists. Support ticket.
Detection: idempotency.stuck_in_progress{age>5m} gauge — should be ~0. Page if > 50 for 10 min.
Blast radius: Each stuck key is one customer with an unknown outcome — money taken, nothing delivered.
Mitigation: Add locked_until + version; a retry after lease expiry takes over and resumes from recovery_point. A background resolver sweeps records stuck > 5 min with no retry, resumes them (re-calls the provider with the derived key → gets the original charge → completes the order), and pages a human queue after N failures.
Prevention: Recovery points are part of the middleware API, not optional; every handler with an external call must declare its phases.
Owner: The service team owns resolution of its stuck records; the platform owns the resolver framework and the gauge.
4.4 False Deduplication — The Key Reused for a Different Operation#
t=0: Partner integration generates Idempotency-Key = customer_id (a "stable" value)
t=0: First order for customer 881 executes; key 881 completed
t=+2h: Second, different order for customer 881 → same key → fingerprint check disabled
("it was causing 422s") → server replays first order's 201
t=+2h: Partner believes second order succeeded. It doesn't exist.
t=+2 weeks: Partner reconciliation: 3,100 "confirmed" orders missing. Escalation to VP.
Detection: idempotency.fingerprint_mismatch_total{client} — with fingerprinting on, this incident is 3,100 loud 422s on day one instead of silent loss. Also idempotency.replay_ratio{client} > 20% is suspicious for a partner that rarely retries.
Blast radius: Silent data loss for every client with the bug. Worse than duplicates: lost operations can't be reconciled from your side because they never reached your database.
Mitigation: Re-enable fingerprinting; contact the partner with the list of 422s. Never disable the guard because it's "noisy" — the noise is the bug report.
Prevention: The fingerprint check cannot be disabled per client without a platform exception with expiry. Partner onboarding docs show correct key generation with a code sample.
Owner: API platform owns the guard; partner engineering owns integration review.
4.5 Retry Amplification Across Layers#
t=0: Database p99 rises from 20ms to 900ms (a slow query plan)
t=+10s: Service B's client times out at 500ms, retries 3× → 3× load on DB
t=+20s: Service A times out on B, retries 3× → 9×; gateway retries → 27×; SDK → 81×
t=+40s: DB CPU 100%; every layer's retries now fail; outage
t=+40s: Requests that DID commit before timeouts get re-executed by upper layers
without derived keys → duplicate side effects in Service B
Detection: rpc.retry_ratio{caller,callee} > 10%; rpc.attempts_per_request p99; request fan-in at the database vs at the edge.
Blast radius: The entire call graph; the retry storm outlives the original cause.
Mitigation: Disable retries at middle layers (feature flag); shed load at the database with a circuit breaker (see Circuit Breakers).
Prevention: Retry at exactly one layer — the outermost client that holds the idempotency key. Inner layers fail fast and propagate deadlines. Retry budgets cap retries at ~10% of requests per client. Every inner call carries a derived key so the retries that do happen are safe.
Owner: The platform team owns client-library retry defaults and the retry-budget rule; service teams own their deadlines.
4.6 Consumer Dedup Window Shorter Than Replay#
t=0: Fulfilment consumer dedupes on event_id using Redis SET with 24h TTL
t=+3 days: Bug in consumer; operator resets consumer-group offset to 3 days ago to reprocess
t=+3 days: Dedup keys for days 1–2 expired → 1.9M events reprocessed as new
t=+3 days: 41K orders get a second "ship" command → warehouse picks duplicates
Detection: consumer.dedup_hit_ratio drops to ~0 during a replay that should be ~100% hits; consumer.lag jump after an offset reset; alert on any manual offset reset in a consumer group tagged "side-effecting."
Blast radius: Every effect produced by the replayed range.
Mitigation: Pause the consumer; reconcile shipments against orders; resume.
Prevention: Dedup window ≥ topic retention for side-effecting consumers (inbox table in the consumer's DB, partitioned by day, dropped at retention + 1 day), or versioned/state-machine writes that are idempotent forever (UPDATE orders SET state='shipped' WHERE id=:id AND state='paid'). Offset resets on side-effecting consumers require a runbook and a second approver.
Owner: Consuming team owns its dedup; the Kafka platform owns the offset-reset guardrail.
4.7 Outbox Relay Stall or Poison Row#
t=0: Outbox relay polls 'SELECT … WHERE published_at IS NULL ORDER BY id LIMIT 500'
t=+0s: One event's payload exceeds broker max.message.bytes → publish fails, relay retries forever
t=+20min: All events behind it are stuck; downstream sees no new orders; search index stale
Detection: outbox.oldest_unpublished_age_seconds > 60 → page; outbox.publish_errors_total.
Mitigation: Move the poison row to an outbox dead-letter table with an alert; relay continues.
Prevention: Payload size validated at write time; per-aggregate ordering only (not global), so one stuck aggregate doesn't block others; CDC-based relay (e.g., Debezium's outbox event router) instead of polling at high volume.
Owner: The CDC/relay platform owns stalls; the producing team owns its poison rows.
4.8 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| TTL < retry horizon | idempotency.key_age_at_request_seconds; duplicate detector | One client platform | Longer TTL; explicit expiry rejection | API platform + client team |
| Store loses keys on failover | Replication lag at failover; post-failover dup scan | In-flight requests | Same-tx store; Redis as cache only | Service team; finance signs off |
| Stuck in-progress | idempotency.stuck_in_progress{age>5m} | One customer per key, money at risk | Lease + takeover + resolver | Service team |
| False dedup (key reuse) | idempotency.fingerprint_mismatch_total | Silent loss of operations | Fingerprint guard non-optional | API platform |
| Retry amplification | rpc.retry_ratio; attempts per request | Whole call graph | Single-layer retries; budgets; deadlines | Platform (client libs) |
| Consumer window too short | consumer.dedup_hit_ratio during replay | Replayed range | Window ≥ retention; versioned writes | Consuming team |
| Outbox stall | outbox.oldest_unpublished_age_seconds | All downstream consumers | Dead-letter poison rows | Relay platform |
| Unknown outcomes accumulate | ops.unknown_outcome_total{age>1h} | Money / inventory discrepancies | Resolver + reconciliation | Service team; ops queue |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | "Add idempotency keys" | Names four duplicate sources; separates delivery from effect | Asks how many implementations and contracts exist; frames it as an org contract |
| Key semantics | Client sends UUID | Per logical op, minted before first attempt, scoped, fingerprinted, TTL ≥ horizon, explicit expiry | Key contract published once, enforced by gateway, SDKs and lint |
| Atomicity | Redis SETNX | Record commits with effect; recovery points around external calls | Durability tiers by business risk, signed off by effect owners |
| Concurrency | Not considered | In-progress + lease + 409; takeover with fencing | Standard status codes and SDK behavior across all APIs |
| Boundaries | Retry with backoff | Derived keys; single-layer retries; deadlines; outbox + inbox | Retry budgets org-wide; dual writes blocked in review |
| Unknown outcomes | Treated as failures | Third state; resolver; reconciliation | Unknown-outcome rate as an SLO with an owning org |
| Operations | — | Metrics for replay, conflict, mismatch, stuck | Cost model, game days for failover, platform roadmap |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Locates the crash window | "What happens if we crash after the charge and before recording it? That's the case the design has to answer." |
| Precise key semantics | "Key per logical operation, minted before the first attempt, scoped to tenant and endpoint, with a request fingerprint." |
| Concurrency | "The second copy can't run the handler. It gets 409 with Retry-After, and a crashed first copy is taken over after its lease." |
| Delivery vs effect | "Kafka's exactly-once ends at Kafka. The consumer that sends the email dedupes in its own transaction." |
| Retry discipline | "Retries happen at one layer, with the key. Everything inside fails fast and propagates the deadline." |
| Unknown outcomes | "A timeout is not a failure. It keeps the same key until we know." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "Check if key exists, then process" | Classic check-then-act race; concurrent duplicates both execute |
| Redis as the only store for money-moving effects, no caveat | Ignores the crash gap and failover loss |
| Server generates the key from a payload hash | Collapses legitimate identical operations; misplaces the responsibility |
| "Kafka exactly-once solves it" | Confuses broker-internal guarantees with end-to-end effects |
| Retries at every layer "for resilience" | Multiplicative amplification; duplicates in un-keyed inner calls |
| No answer for the 500-after-side-effect case | Misses the core unknown-outcome problem |
5.4 Common False Positives#
- Quoting Stripe's header ≠ designing idempotency. The record lifecycle, concurrency and propagation are the design.
- Knowing Kafka transactions ≠ exactly-once effects. Ask what the consumer does next.
- Distributed-transaction vocabulary (2PC, XA) ≠ judgment. External providers won't join your transaction; the Staff answer avoids needing them.
- "We use a dedup table" ≠ consumer idempotency — unless it's in the same transaction as the effect and outlives the replay horizon.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing & intents | 0–4 min | Four duplicate sources; commit to public API + downstream + event |
| Contract & entities | 4–7 min | Key, scope, fingerprint, status codes, record schema |
| Architecture | 7–11 min | Same-tx record, phases, derived keys, outbox |
| Deep dive: concurrency & crashes | 11–21 min | In-progress lease, takeover, recovery points, replay rules |
| Deep dive: boundaries | 21–31 min | Derived keys, single-layer retries, outbox, inbox, replay horizon |
| Deep dive: lifetime & misuse | 31–38 min | TTL vs horizon, expiry rejection, fingerprint, scope |
| Platform / multi-region | 38–43 min | Middleware, SDKs, tiers; key home region |
| Wrap-up | 43–45 min | Deferred items, unknown-outcome metric |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response |
|---|---|---|
| "The downstream service doesn't support idempotency keys" | Handling external effects without help | "Record intent before the call, look up by our reference after a crash, reconcile the residue; if no lookup exists, the effect is at-least-once and the business owner signs off on that." |
| "Make the idempotency store a shared service" | Atomicity vs generality | "Then there's a gap between effect and record. I'd ship a library with a schema instead, so the record shares the service's transaction." |
| "Clients are lazy and won't send keys" | Platform enforcement | "Gateway rejects mutating requests without keys for new endpoints; SDKs generate them automatically; legacy endpoints go through shadow → warn → enforce." |
| "Now it's multi-region active-active" | Key locality | "A key must have one home region: route by tenant home, or use a globally conditional write on the key with its ~100ms tax." |
| "We need exactly-once counting for billing" | Pipeline semantics | "Deterministic event IDs, idempotent aggregation keyed on (window, event_id), and reconciliation against the source ledger." |
| "What if the client retries after 30 days?" | Expiry semantics | "Explicit rejection from the key's embedded timestamp; the SDK never retries past the TTL." |
6.3 What to Deliberately Skip#
- The mathematical definition of idempotence beyond one sentence.
- Two-phase commit internals — say why you don't need it and move on.
- Detailed payment state machines — link Payment Processing.
- Lock-service and leader-election design — link Distributed Coordination.
- Broker internals beyond "at-least-once, redelivery on rebalance, retention-bounded replay."
6.4 Follow-Up Questions to Expect#
- "Who generates the key, and when exactly?"
- "Two identical requests arrive 5ms apart. What happens to each?"
- "The first attempt crashed after charging the card. What does the retry do?"
- "Do you replay a 500? A 402? A 429?"
- "How long do you keep keys, and what happens to a retry after that?"
- "How does the guarantee survive the call to the inventory service and the Kafka event?"
- "How do you get 40 teams to do this consistently?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Make our order API safe to retry."
Staff Answer
"Duplicates come from four places — client retries after timeouts, our own service-to-service retries, broker redelivery, and replays — and the property I need is that the order, the charge and the shipment each happen once per logical operation. Exactly-once delivery isn't available, so I'll build exactly-once effect. Clients mint a key per logical operation before the first attempt; the server stores a record scoped to tenant and endpoint, with a request fingerprint, in the same transaction as the order. Concurrent duplicates get 409; a crashed attempt is resumed from a recovery point; calls to inventory and the payment provider carry derived keys; the event goes out through an outbox and consumers dedupe. I'll start with the record's lifecycle, because that's where the crash window lives."
Why this is L6:
- Separates delivery from effect in the first sentence
- Specifies the key contract (who, when, scope, fingerprint) rather than naming the header
- Previews the crash window, concurrency and propagation before being asked
What L7 adds:
- Asks whether an org-wide idempotency contract already exists, and whether this API should adopt it rather than define its own
- Sets a durability tier for the endpoint with the business owner of the effect (finance for charges)
- Names the unknown-outcome rate as the metric leadership will see
❌ Common L5 Trap
"We'll add an Idempotency-Key header. Before processing, check Redis for the key; if present return the cached response, otherwise process and store the response with a 24-hour TTL."
Why this misses: Check-then-act lets concurrent duplicates both execute; Redis isn't atomic with the order write, so a crash between them duplicates; a 24h TTL is arbitrary relative to the client's retry horizon; and the charge and the event aren't covered at all. Every one of these will be the interviewer's next question.
Drill 2: Key Semantics#
Prompt: "Who generates the key, when, and what is it unique within?"
Staff Answer
"The client, because only the client knows that two requests are the same intent. It's generated when the user commits to the action — the tap on 'Place order' — before the first network attempt, and persisted with the pending operation so a retry after an app restart reuses it. It's unique within (tenant, endpoint): tenant so one tenant's keys can't collide with or probe another's, endpoint so a key accidentally reused across /orders and /refunds can't replay the wrong resource. I'd recommend UUIDv7 — the timestamp lets the server reject keys older than the TTL explicitly. If the request has a natural business key, like the partner's own order reference, I'd make that unique per tenant and use it directly."
Why this is L6:
- Places key generation at the moment of user intent, not at request construction
- Justifies both dimensions of scope with a failure each prevents
- Uses the key's format to close the expiry hole
What L7 adds:
- Ships this as SDK behavior in every client language so app developers can't get it wrong
- Makes the key contract part of the public API standard reviewed by API governance
Drill 3: Concurrent Duplicates#
Prompt: "Two copies of the same request arrive 5ms apart on different instances. Walk me through both."
Staff Answer
"Both try INSERT INTO idempotency_keys (tenant, endpoint, key, fingerprint, status='in_progress', locked_until=now()+30s, version=1). The unique constraint lets exactly one succeed. The winner runs the handler. The loser gets a unique violation, reads the existing row: if the fingerprint differs → 422; if in_progress and the lease is live → wait up to ~1s for completion, then 409 with Retry-After: 1; if completed → replay. If the winner crashes, the row stays in_progress until locked_until; the next retry after that does UPDATE … SET locked_until=now()+30s, version=version+1 WHERE key=… AND locked_until < now() AND version=:v — exactly one taker wins — and resumes from the recorded recovery point. The original, if it was only slow, fails its final commit because its version is stale."
Why this is L6:
- Uses the database's unique constraint as the arbiter rather than a separate lock
- Covers live, crashed and slow-original cases, with fencing
- Keeps waiting bounded so double-click storms can't exhaust threads
What L7 adds:
- Standardizes 409 + Retry-After semantics across every API so SDKs implement it once
- Sizes the lease from fleet-wide handler latency data and reviews it when SLOs change
Drill 4: Make It Concrete — Crash After the Side Effect#
Prompt: "The handler charges the card, then the process dies before committing the order. What happens on retry?"
Staff Answer
"The handler is phased with recovery points. Phase 1, in one transaction: insert the key record and the order in pending_payment, recovery point order_created. Phase 2: call the provider with derived key k:charge. Phase 3, one transaction: mark order paid, write the outbox event, mark the key completed with the response. The crash was between 2 and 3. The record is in_progress at order_created. The retry — after the lease expires, or the background resolver if no retry comes — takes over and resumes at phase 2: calls the provider again with k:charge. The provider's idempotency returns the original successful charge. Phase 3 commits. One charge, one order. If the provider had no idempotency support, phase 2 would first query the provider by our order reference before charging."
Why this is L6:
- Shows the exact phase boundaries and which ones are transactional
- Relies on the downstream's idempotency via a derived key — the only thing that makes the external call safe
- Has a plan for downstreams without idempotency
What L7 adds:
- Makes 'every external call has a derived key and a recovery point' a library-enforced rule, not a code-review hope
- Requires new vendors to support idempotency or lookup-by-reference as a procurement criterion
Drill 5: Dependency Down — The Idempotency Store#
Prompt: "The database holding idempotency records is unavailable. Do you fail open or closed?"
Staff Answer
"If the record lives in the same database as the effect, the question answers itself: the write can't happen either, so we fail with 503 and the client retries with the same key later. That's a feature of co-location — there's no state where the effect can happen without the record. The question only becomes hard with a separate store. Then it's per tier: money-moving and shipment endpoints fail closed — 503, no execution — because a duplicate costs more than a delayed order. Low-stakes endpoints like 'mark notification read' fail open and execute without dedup. The tier is declared per endpoint and signed off by the effect's owner, not decided during the outage."
Why this is L6:
- Shows co-location eliminates the dilemma
- Tiers the decision by business cost of duplicates
- Pre-decides it with a named owner
What L7 adds:
- Publishes durability tiers in the API standard with default fail behavior per tier
- Rehearses the separate-store failure in a game day for any tier that fails open
Drill 6: The Hot Key — A Buggy Client Retry Loop#
Prompt: "A partner's integration retries the same key 5,000 times a second after a 409."
Staff Answer
"Each retry is a primary-key lookup — cheap, but 5K/s on one row is contention on a hot index page and wasted capacity. First, the completed-record cache: once the key completes, replays are served from Redis without touching the DB. For in_progress, the 409 includes Retry-After; the gateway rate-limits per (tenant, key) — say 5 requests per second — returning 429 above that. Then contact the partner: ignoring Retry-After is a contract violation. The SDK we ship honors it, so this is a hand-rolled client — partner engineering adds it to their integration review."
Why this is L6:
- Distinguishes completed (cacheable) from in-progress (must hit the source of truth) hot keys
- Uses the protocol (Retry-After) and the gateway, not heroics in the service
- Treats the client as a party to the contract
What L7 adds:
- Adds per-key rate limiting to the platform's standard gateway policy so no team builds it ad hoc
- Tracks contract-violating clients as a partner-health metric owned by partner engineering
Drill 7: Multi-Tenant#
Prompt: "We're a platform; tenants choose their own idempotency keys. What can go wrong?"
Staff Answer
"Three things. Collisions: tenants will use keys like order-1; with global scope two tenants collide and one gets the other's replayed response — a data leak. So scope is (tenant, endpoint, key) and the tenant comes from authentication, never from the request body. Probing: with global scope a tenant could test whether a key exists; tenant scope removes that. Noisy neighbors: one tenant with a 7-day TTL and 50M writes a day dominates the key table — partition by day and tenant-tier, cap key length at 64 bytes and stored response size at 16 KB, and meter storage per tenant so it shows up in their cost line."
Why this is L6:
- Treats cross-tenant replay as a data-leak, not a curiosity
- Derives tenant from auth, closing a spoofing path
- Puts limits on key and response size, which is where storage abuse hides
What L7 adds:
- Makes retention a priced plan feature for platform customers (24h default, 7d on enterprise) rather than a hidden cost
Drill 8: Build vs Buy#
Prompt: "Should we build idempotency into the gateway, a library, or rely on the database?"
Staff Answer
"Split by responsibility. The gateway enforces presence and format of keys and per-key rate limits — it's stateless and shared, so it's the right place for the contract check. It cannot own the record, because the record must commit with the service's effect, and the gateway isn't in that transaction. So the record lives in a library plus a schema migration in each service's database; the library provides phases, recovery points and derived-key helpers. The database provides the unique constraint and the transaction. A gateway plugin that stores responses itself is fine for the low-stakes tier only. I'd buy nothing extra — every piece rests on the database we already run."
Why this is L6:
- Assigns each responsibility to the layer that can actually provide it
- Explains why a gateway-only solution has a crash window
- Avoids a new stateful dependency
What L7 adds:
- Prices the library across languages: 3 languages × ~0.5 FTE to maintain, versus duplicated incidents across 40 teams
- Plans for version skew: the library's record schema is versioned and migrations are backwards-compatible for two releases
Drill 9: Policy Change Without an Outage#
Prompt: "We want to make Idempotency-Key required on all POST endpoints. 300 integrations don't send it today."
Staff Answer
"Shadow, warn, enforce. Shadow: the gateway logs which clients call mutating endpoints without a key, broken down by client ID and SDK version — that's the migration list. Ship SDK versions that generate keys automatically; most traffic moves with an SDK upgrade. Warn: responses without a key get a Deprecation header and a warning in the developer dashboard; partner engineering contacts the top 50 by volume. Enforce: new API versions require the key from day one; old versions require it on a published date, 6 months out, with a per-client exception list that expires. Throughout, requests without a key still work exactly as before — we don't fake a key for them, because a server-generated key per request dedupes nothing."
Why this is L6:
- Uses data to drive the migration list
- Leverages SDKs to move most traffic for free
- Ties enforcement to API versioning so nothing breaks without notice
What L7 adds:
- Sets the policy in the API standard with a deprecation calendar shared with sales and partner teams
- Tracks coverage (percent of mutating requests carrying keys) as a quarterly platform metric
Drill 10: Multi-Region#
Prompt: "We're going active-active. A retry with the same key lands in a different region."
Staff Answer
"With per-region databases, region B has never seen the key and executes it again — a duplicate. Three options. One: route by tenant home region, so all writes for a tenant — including retries — go to one region; failover moves the home and the key table with it. This is the default; it's how most active-active systems keep writes single-homed per entity. Two: a globally consistent key table, such as Spanner — an asynchronously replicated multi-region table is not enough, because a conditional write only checks the local replica and conflicts are later resolved last-writer-wins. This adds a cross-region round trip, ~70–150ms. Three: accept regional dedup and reconcile duplicates — only for low tiers. I'd pick one: tenant-homed routing, with retries following the home."
Why this is L6:
- Identifies that async multi-region replication silently breaks dedup
- Knows the specific trap with last-writer-wins replicated tables
- Chooses entity-homed writes, aligning idempotency with the multi-region data model
What L7 adds:
- Aligns with the org's Multi-Region cell model so the idempotency store doesn't invent its own topology
- Defines what happens to in-flight keys during a regional failover and puts it in the failover runbook
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Incident — Idempotency Table Becomes the Bottleneck#
Context: During a holiday sale, order-write p99 jumps from 80ms to 2.4s. The order DB's CPU is fine; lock waits on the idempotency_keys table are high. Replay rate has risen from 2% to 35% because clients are timing out and retrying. You're escalated.
Questions to Surface First:
- Are replays hitting the database, or is there a completed-record cache?
- Are retries honoring
Retry-After, or is one client hammering in-progress keys? - Is the table partitioned and its indexes healthy, or did 7-day retention at peak volume bloat it?
Typical L5 Approach: Scales the database vertically and increases client timeouts. Latency recovers partially; the retry-driven load remains and returns at the next peak.
Staff Approach: Recognizes a feedback loop: slow writes → client timeouts → retries → replay lookups and 409s on hot rows → slower writes. Breaks it: enables the completed-record cache in Redis (replays never touch the DB), confirms SDKs honor
Retry-Afterand gateway per-key limits are on, and checks the table is partitioned by day so the hot partition's indexes fit in memory. Raises the client timeout only after the loop is broken.
Principal Approach: Treats it as a retry-policy failure across the org, not a table problem. Mandates retry budgets in all client SDKs (≤ 10% of requests), single-layer retries, and capacity planning that models retry amplification at peak (e.g., plan for 1.3× nominal write load). Adds 'replay ratio at peak' to the launch-readiness review and makes the completed-record cache part of the standard library rather than a per-team optimization.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Enable completed-record cache via flag; turn on gateway per-key rate limit (5 req/s). |
| Triage | Break down replays by client; identify clients ignoring Retry-After. Check partition sizes and index bloat. |
| Quick fix | Disable inner-layer retries; shed low-priority writes. |
| Guardrails | Watch idempotency.replay_ratio, db.lock_wait_ms; keep timeout increases until replay ratio < 5%. |
| Post-mortem | Why was the cache optional? Why did clients retry so aggressively? Model retry load in capacity plans. |
Metrics to Watch: idempotency.replay_ratio, idempotency.conflict_409_total, db.lock_wait_ms{table="idempotency_keys"}, rpc.retry_ratio, orders.write_latency_p99
Organizational Follow-up: Retry budgets in SDKs; capacity planning includes retry amplification.
Ownership Question: "Who decides to turn off retries at inner layers during an incident?" Staff answer: The incident commander, using a pre-approved flag documented in the platform runbook. Inner-layer retries should be off by default anyway; the flag exists for the legacy services that haven't migrated.
Key Takeaway: "Retries are load. An idempotency layer that serves replays from the database turns client panic into a database incident."
What clears the Staff bar:
- Identifies the feedback loop rather than the symptom
- Breaks the loop before tuning timeouts
- Uses contract mechanisms (Retry-After, per-key limits) rather than capacity
Deep Dive 2: Silent Failure — Partners' Second Orders Vanish#
Context: A large partner reports 3,100 orders their system marked "confirmed" over two weeks that don't exist in yours. Your dashboards show no errors. Replay rate for this partner is 22%; for others, 1–3%.
Questions to Surface First:
- How does the partner generate idempotency keys?
- Is fingerprint checking enabled for this partner's traffic?
- Do the "missing" orders share keys with orders that do exist?
Typical L5 Approach: Searches logs for errors, finds none, asks the partner to resend the missing orders. Resent orders succeed. Closes the ticket. The partner's next week loses another 1,500.
Staff Approach: The 22% replay rate is the tell. Pulls replayed requests: the partner uses
customer_idas the key, and fingerprint checking was disabled for them months ago after they complained about 422s. Every second order per customer was replayed as the first. Re-enables fingerprinting immediately (the partner now gets loud 422s instead of silent loss), sends them the list of affected operations with fingerprints, and fixes their key generation with partner engineering.
Principal Approach: Fixes the governance gap that allowed a safety check to be disabled per client: guard rails become non-optional without a time-boxed exception approved by the API platform owner, and every exception appears on a dashboard. Adds
replay_ratioby client as an anomaly alert. Rewrites partner onboarding to include a certification test that sends two distinct operations and verifies two distinct keys.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Re-enable fingerprint check for the partner. |
| Triage | Join replays to fingerprints; enumerate every replay whose fingerprint differed from the original. |
| Quick fix | Send the partner the list; they resubmit with new keys. |
| Guardrails | Alert on per-client replay ratio > 10%; fingerprint check cannot be disabled without expiry. |
| Post-mortem | Why was the guard disabled? Who approved it? Why didn't anything alert? |
Metrics to Watch: idempotency.replay_ratio{client}, idempotency.fingerprint_mismatch_total{client}, idempotency.guard_exceptions_active
Organizational Follow-up: Partner certification tests; exceptions dashboard reviewed monthly.
Ownership Question: "Who should have been able to disable the fingerprint check?" Staff answer: Nobody, without a written exception with an expiry approved by the API platform owner. The engineer who disabled it was solving a support ticket; the system let a support fix silently become a data-loss bug.
Key Takeaway: "False deduplication is worse than duplication — the lost operation never reaches your database, so you can't reconcile it."
What clears the Staff bar:
- Uses replay ratio as a diagnostic signal
- Recognizes the disabled guard as the root cause, not the partner bug alone
- Converts silent loss into loud errors immediately
Deep Dive 3: Large-Customer Onboarding — An Enterprise With a 7-Day Retry Horizon#
Context: An enterprise customer's ERP integration batches 2M mutating API calls nightly and, on any failure, re-runs the entire failed batch for up to 7 days. Your idempotency TTL is 24h. They're about to go live.
Questions to Surface First:
- How does the ERP generate keys — per record, per batch, per attempt?
- Does each record have a natural business key (their document number)?
- What volume of storage does 7-day retention at 2M/night imply?
Typical L5 Approach: Raises the global TTL to 7 days for everyone. Storage grows 7× for all tenants to serve one.
Staff Approach: Uses the natural key: the ERP's document number, unique per tenant, becomes the idempotency key via a unique constraint on
(tenant, external_ref)— dedup is permanent, not TTL-bound, and costs one index. For endpoints without a natural key, offers per-tenant retention (7d for this tenant: 2M × 2 KB × 7 ≈ 28 GB) and a bulk API that accepts a batch ID plus per-item keys, returning per-item outcomes so partial failures re-run only the failed items.
Principal Approach: Turns retention into a product decision: retention tiers (24h standard, 7d enterprise) priced into plans, with the storage cost visible. Makes 'natural external reference as idempotency key' the recommended integration pattern in the enterprise integration guide, because it is stronger and cheaper than any TTL. Adds bulk-API per-item idempotency to the platform roadmap since every ERP customer will need it.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Assess | Map their retry behavior and key generation; size storage. |
| Design | Natural key unique per tenant; per-tenant retention for keyless endpoints. |
| Build | Bulk endpoint with per-item keys and per-item results. |
| Guardrails | Rate-limit nightly batch; per-tenant storage metering. |
| Go-live | Run a replayed batch in staging and verify zero duplicates. |
Metrics to Watch: idempotency.storage_bytes{tenant}, bulk.item_replay_ratio{tenant}, idempotency.key_age_at_request_seconds{tenant} p99
Organizational Follow-up: Retention tiers in pricing; integration guide update.
Ownership Question: "Who pays for 7-day retention?" Staff answer: The customer's plan, through a priced retention tier. If it's free, it becomes everyone's default and the platform pays for it silently.
Key Takeaway: "The strongest idempotency key is the customer's own business identifier — it never expires."
What clears the Staff bar:
- Prefers natural keys over TTL extension
- Sizes storage per tenant
- Designs partial-failure semantics for bulk operations
Deep Dive 4: Post-Mortem — 41,000 Double Shipments After an Offset Reset#
Context: A fulfilment consumer had a bug; an operator reset its Kafka consumer-group offset by 3 days to reprocess. The consumer deduped on event_id in Redis with a 24h TTL. 41,000 orders shipped twice. Cost: ~$1.2M in goods and logistics.
Questions to Surface First:
- Why was the dedup window shorter than the topic's retention?
- Why was a manual offset reset on a side-effecting consumer possible without review?
- Is the "ship" effect itself idempotent at the warehouse (does it accept a shipment ID)?
Typical L5 Approach: Extends the Redis TTL to 7 days. Next replay beyond 7 days, or the next Redis eviction under memory pressure, repeats the incident.
Staff Approach: Makes the effect idempotent forever rather than for a window: the consumer transitions orders with
UPDATE orders SET state='shipping', shipment_id=:sid WHERE id=:id AND state='paid'and calls the warehouse withshipment_idderived fromorder_id, which the warehouse dedupes. The inbox table moves into the consumer's DB, same transaction, retained ≥ topic retention. Offset resets on side-effecting consumers require a runbook and second approver.
Principal Approach: Classifies consumers org-wide as side-effecting or pure, and requires every side-effecting consumer to pass a 'full replay' test in CI: replay the entire retention window and assert no new external effects. Puts the offset-reset guardrail in the Kafka platform, not each team's runbook. Takes the $1.2M to leadership as the price of 'dedup with a TTL' and funds the paved-road inbox library.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Pause the consumer; stop outbound warehouse calls. |
| Triage | Identify all orders with two shipment requests in the replay window; cancel pending pickups. |
| Quick fix | Deploy state-guarded transition; resume. |
| Guardrails | Offset-reset approval; replay test in CI; dedup in-transaction. |
| Post-mortem | Dedup window vs retention; warehouse API idempotency; who approved the reset. |
Metrics to Watch: consumer.dedup_hit_ratio, consumer.offset_reset_events, warehouse.duplicate_shipment_requests, orders.state_transition_rejected_total
Organizational Follow-up: Consumer classification; replay test standard; platform-level reset guardrail.
Ownership Question: "Who owns the guarantee that replays are safe?" Staff answer: The consuming team owns its effects' idempotency. The Kafka platform owns making unsafe replays hard to trigger. The incident needed both to fail; the fix needs both to change.
Key Takeaway: "A dedup window is a bet on how far back anyone will ever replay. State-guarded effects don't need to win that bet."
What clears the Staff bar:
- Replaces windowed dedup with naturally idempotent state transitions
- Pushes idempotency to the external effect (warehouse shipment ID)
- Adds an operational guardrail on the human action that triggered it
Deep Dive 5: Multi-Region Expansion#
Context: The company is moving the order API to active-active in two regions. Today the idempotency table lives in the single-region order database.
Questions to Surface First:
- Is the data model entity-homed (each tenant/order has a home region) or truly multi-writer?
- How does the client route — anycast, geo-DNS, or tenant-aware?
- During a regional failover, what happens to records in
in_progress?
Typical L5 Approach: Replicates the idempotency table with an async multi-region table and assumes conditional writes still work. They do — per region. A retry that lands in the other region within the replication lag executes again; last-writer-wins later overwrites one of the two records.
Staff Approach: Keeps writes single-homed per tenant: the edge routes mutating requests to the tenant's home region; the idempotency record lives with the tenant's data. On failover, the home moves with the data;
in_progressrecords are taken over in the new region via the normal lease-expiry path, and their recovery points tell the resolver which downstream calls to re-issue with derived keys.
Principal Approach: Makes idempotency a first-class requirement of the org's multi-region architecture review: any multi-writer data model must state how idempotency records are homed. Publishes 'tenant-homed writes' as the default cell model, with globally consistent stores reserved for the few entities that need them, priced (~100ms per write) and approved explicitly.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Plan | Confirm tenant-homed routing for mutating requests; idempotency table co-located with tenant data. |
| Failover design | Lease takeover in the new home; resolver re-issues downstream calls with derived keys. |
| Rollout | Migrate tenants region by region; verify replay correctness during cutover. |
| Guardrails | Alert on mutating requests served outside the tenant's home region. |
| Review | Game day: fail a region with in-flight orders; count duplicates (target: zero). |
Metrics to Watch: routing.mutations_outside_home_region, idempotency.takeover_total{region}, ops.unknown_outcome_total{region}, replication lag
Organizational Follow-up: Multi-region review checklist gains an idempotency-homing item.
Ownership Question: "Who owns in-flight operations during a regional failover?"
Staff answer: The order service's resolver in the new home region, following the runbook. The failover isn't complete until in_progress records from the failed region are resolved or escalated.
Key Takeaway: "Asynchronously replicated dedup is regional dedup. Home the key with the data."
What clears the Staff bar:
- Spots the last-writer-wins trap in async multi-region tables
- Aligns idempotency with the data-homing model
- Defines failover behavior for in-flight operations
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Name the four sources of duplicates and the mechanism that handles each
- Explain why exactly-once delivery is impossible and how exactly-once effect is assembled
- Specify an idempotency key contract: who mints it, when, scope, fingerprint, TTL, expiry behavior, status codes
- Design the idempotency record's lifecycle including in-progress leases, takeover, fencing and recovery points
- State which outcomes to replay and which to release, and why the 500-after-side-effect case is neither
- Carry idempotency across a downstream call (derived keys), a broker (outbox, inbox) and a replay (window ≥ retention or state-guarded writes)
- Explain retry amplification and the single-layer retry rule with retry budgets
- Make idempotency a platform default: middleware, SDKs, gateway enforcement, durability tiers
The Bar for This Question#
Mid-level (L4): Knows that POST isn't idempotent and that retries can duplicate. Proposes "check if it already exists." Doesn't reach concurrency or crash windows.
Senior (L5): Proposes Idempotency-Key with a Redis store and TTL, response caching, and exponential backoff. Correct for the simple retry. Under probing, concurrent duplicates, the crash-after-effect window, Redis failover and downstream propagation are improvised.
Staff+ (L6): Frames the problem as unknown outcomes and exactly-once effect. Defines the key contract precisely, co-locates the record with the effect, handles concurrent and crashed attempts with leases, fencing and recovery points, and carries the guarantee through derived keys, outbox and consumer inbox with a window that outlives replay. Has metrics for every leak and an owner for every one. The interviewer should learn something from the answer — often the false-deduplication failure or the expiry-rejection trick with timestamped keys.
10. Staff Insiders: Controversial Opinions#
10.1 "Exactly-Once Delivery Is a Marketing Term"#
| Evidence | Detail |
|---|---|
| Theory | The sender can't distinguish a lost message from a lost acknowledgement (Two Generals); retries are required, so duplicates are possible |
| Kafka EOS | Real within Kafka (idempotent producer, transactions, read_committed); ends at the first non-Kafka side effect |
| Practice | Every "exactly-once" system is at-least-once delivery + deduplicated processing |
The Staff position: Say "exactly-once effect" and name the dedup mechanism at each hop. Never say "exactly-once delivery" without a qualifier.
Why this matters in interviews: It's the fastest way to show you know where guarantees end.
10.2 "Redis Is the Wrong Default Idempotency Store"#
| Evidence | Detail |
|---|---|
| Atomicity | Not in the effect's transaction → crash gap |
| Durability | Async replication loses recent writes on failover |
| Eviction | maxmemory policies can evict keys before TTL under pressure |
| What it's good for | Caching completed records, which are immutable |
The Staff position: Source of truth in the service DB, in the effect's transaction. Redis in front, for completed replays only. Redis-only is acceptable for an explicitly low-stakes tier.
Why this matters in interviews: The Redis answer is the most common L5 answer. Explaining precisely why it leaks — and where Redis still belongs — is a clean level signal.
10.3 "A Fixed TTL Is a Bug Waiting for a Slow Client"#
| Evidence | Detail |
|---|---|
| Client horizons vary | Web: seconds; mobile offline: days; ERP batch re-runs: a week |
| Expiry semantics | After expiry, a retry silently re-executes |
| Fixes | Explicit expiry rejection via timestamped keys; natural business keys with no expiry; per-tenant retention |
The Staff position: TTL is half of a contract whose other half is the client's retry horizon. Publish both, enforce both, and make expiry an error, not a re-execution.
Why this matters in interviews: "24 hours, like Stripe" is cargo cult; the follow-up "and a client that retries at 30 hours?" separates levels.
10.4 "Most Idempotency Bugs Are Retry-Policy Bugs"#
| Evidence | Detail |
|---|---|
| Amplification | Retries at N layers multiply load (3^N) and multiply un-keyed duplicate executions |
| Common cause | Default retries in HTTP clients, service meshes and SDKs stacked without anyone choosing them |
| Fix | Retry at one layer, with the key; propagate deadlines; retry budgets |
The Staff position: Before designing deduplication, find out how many layers retry. Remove all but one.
Why this matters in interviews: Connecting idempotency to retry policy across the call graph shows system-wide thinking.
10.5 "False Deduplication Is Worse Than Duplication"#
| Evidence | Detail |
|---|---|
| Duplicates | Visible in your data; reconcilable; refundable |
| False dedup | The lost operation never reaches your database; only the client knows |
| Cause | Key reuse across different operations with no fingerprint check |
The Staff position: Fingerprinting is not optional. A 422 for key reuse is the system working.
Why this matters in interviews: Most candidates only design against duplicates. Naming the opposite failure is rare and memorable.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
A Staff engineer makes one API safe to retry. A Principal engineer notices that idempotency is a property of the weakest hop: the order API can be perfect, and the customer is still double-charged because the payments client library retries without a derived key, or the fulfilment consumer dedupes with a 24h TTL, or a service mesh adds a retry nobody configured. Across 40 teams there will be 6 storage choices, 5 TTLs, 4 sets of status codes and 3 SDKs that handle 409 differently. At L7 the problem is the org's retry-and-dedup contract: one key standard, one middleware library, one SDK behavior, one retry policy, one outbox/inbox paved road, and a metric — unknown outcomes and duplicate effects — that leadership can see.
🧭 Principal Move: "Before we design this endpoint, how many layers in our stack retry by default, and how many idempotency implementations do we have? I'd expect the answers to be 'four' and 'six', and the highest-leverage work is making it one of each."
The Org-Level Fault Line#
Central idempotency service vs per-team implementations vs platform library + contract.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Per-team implementations | Fits each domain | Inconsistent contracts; Redis-only stores; TTLs mismatched to SDKs; duplicates at seams | Customers; finance (refunds); support |
| Central idempotency service | One implementation, easy audit | Not in any service's transaction → crash gap everywhere; tier-0 dependency on every write | Every team (latency, availability); customers (gap) |
| Platform library + published contract + gateway enforcement (L7 default) | Record co-located with effects; consistent external contract; SDKs uniform | Library maintenance in 3–4 languages; schema migrations per service | Platform: ~2–3 FTE; teams: adoption effort |
The deciding question: can the record commit with the effect? If yes — and for every service with its own database it can — the platform ships code, not a service.
Cost Model#
Assumptions: loaded engineer cost ~$25K/month; Postgres storage ~$0.10–0.25/GB-month with replicas; Redis cache ~$200–$400/node-month; average record 2 KB with response; duplicate-effect cost estimated from refunds and support time.
| Scale | Architecture | Infra $/month | Headcount | On-Call Load |
|---|---|---|---|---|
| ~100K mutating req/day, 1 service | Same-tx table, 72h TTL, daily partition drop | < $50 (~0.6 GB) | ~0.1 FTE (part of service team) | Negligible |
| ~10M/day, 20 services, 2 regions | Library + per-service tables + completed-record cache + outbox/CDC | ~$2K–$5K (60 GB keys; small Redis; CDC infra) | 2 FTE platform (library, SDKs, CDC) | Shared platform rotation; ~1 page/month (stuck records, relay lag) |
| ~500M/day, 100+ services, 5 regions, external partners | Above + tenant-homed routing, retention tiers, resolver framework, replay CI tests | ~$20K–$40K (3 TB keys at 72h; CDC fleet) | 4–5 FTE (library ×4 languages, CDC, resolver, partner tooling) | Own runbook; failover game days; unknown-outcome SLO reviewed monthly |
The line that matters to leadership: The infrastructure is cheap; the absence of it is expensive. A single incident like Deep Dive 4 ($1.2M) or a month of 0.1% duplicate charges on $50M monthly volume ($50K in refunds plus support and chargeback fees) pays for the whole platform team for a year.
The 3-Year Evolution Path#
Each step is triggered by evidence. Building the inbox paved road before any team has a side-effecting consumer is premature; building it after the second replay incident is late.
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversal Cost | Why |
|---|---|---|---|
| Key store implementation (table layout, cache) | Two-way | Weeks | Behind the library interface |
| Fingerprint algorithm | Two-way | Days | Recomputed per request; change with dual-check window |
| Public header name and status-code semantics | One-way | Years | Every external SDK and partner integration encodes them |
| Key scope (global vs tenant + endpoint) | One-way-ish | Quarters | Narrowing scope later changes replay behavior for existing clients |
| Published TTL | One-way downward, two-way upward | Shortening breaks clients that rely on it | Clients build retry horizons on it |
| Replaying errors vs re-executing | One-way | Quarters | Clients write logic around replayed outcomes |
| Same-tx vs separate store | Two-way before scale; costly after | Quarters at 100 services | Migration touches every service's schema |
🧭 Principal Insight: The status codes, header name, scope and TTL are a public API. They get a design review with API governance and partner engineering. The storage engine gets a code review.
The Standard I'd Write#
RFC: Idempotency and Retry Standard (v1)
Scope: Every externally reachable mutating endpoint and every internal endpoint that causes a side effect outside its own database.
Requirements:
- Mutating endpoints MUST accept
Idempotency-Key(16–64 chars), scoped to(tenant, endpoint), and MUST reject reuse with a different request fingerprint (422), concurrent use (409 +Retry-After), and keys older than the published retention (idempotency_key_expired).- Idempotency records MUST commit in the same transaction as the local effect. Separate stores are permitted only for endpoints in the best-effort tier.
- Every call to another service or vendor that causes a side effect MUST carry a derived key and be bracketed by recovery points.
- Retries MUST occur at exactly one layer (the outermost key holder). Inner layers MUST fail fast and propagate deadlines. Client SDKs MUST enforce a retry budget of ≤ 10%.
- Events MUST be published via outbox or CDC; dual writes are blocked by lint.
- Side-effecting consumers MUST dedupe in the same transaction as their effect with a window ≥ topic retention, or use state-guarded writes, and MUST pass a full-replay test.
- Services MUST emit
idempotency.replay_total,idempotency.fingerprint_mismatch_total,idempotency.stuck_in_progress, andops.unknown_outcome_total.Exceptions: Filed with the API platform group; decision within 5 business days; expiry ≤ 2 quarters; listed on a public exceptions dashboard.
Success metrics: 100% of money-moving endpoints compliant in 2 quarters, all mutating endpoints in 4; duplicate-effect incidents → 0 per quarter; unknown outcomes older than 24h < 0.001% of operations;
rpc.retry_ratio< 10% fleet-wide.
What I'd Tell the VP#
"Every time a network hiccups, our systems have to guess whether an action already happened, and right now six teams guess six different ways. That's why we had the double-shipment incident last quarter and why some partners see 'confirmed' orders that don't exist. I'm proposing one standard for how every API handles retries and one shared library that implements it, rolled out to money-moving endpoints first. It costs two to three engineers for a year. It removes the incident class that cost us $1.2M in one night, and it gives us a single number — operations with unknown outcomes — that tells us whether we're getting it right."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Finds the weakest hop | "The endpoint can be perfect and the customer still double-charged if the payments client retries without a derived key." |
| Governs retry policy, not just dedup | "How many layers retry by default? I want that number to be one." |
| Ships code, not a service | "The record has to commit with each service's effect, so the platform ships a library and a schema, not a central store." |
| Prices the absence | "0.1% duplicates on $50M a month is $50K in refunds before support costs — the platform team pays for itself." |
| Treats the contract as public API | "Status codes, scope and TTL are one-way doors; they go through API review with partner engineering." |
Staff answers that L7 interviewers find insufficient:
- "We'll store keys in the same transaction as the effect." — correct, but only for this service; silent on the 39 others and how they adopt it.
- "Consumers dedupe with an inbox table." — correct, but no replay test, no offset-reset guardrail, nothing that stops the next team from using a 24h Redis TTL.
- "Retry with the same key." — correct, but never asks how many layers retry or who sets the retry budget.
Appendices
Appendix A: Mechanics in Depth — The Idempotency Record and Handler
A.1 Schema#
CREATE TABLE idempotency_keys (
tenant_id BIGINT NOT NULL,
endpoint TEXT NOT NULL, -- e.g. 'POST /v1/orders'
idem_key TEXT NOT NULL, -- 16..64 chars
fingerprint BYTEA NOT NULL, -- SHA-256 of canonical request
status TEXT NOT NULL, -- 'in_progress' | 'completed'
recovery_point TEXT NOT NULL DEFAULT 'started',
version INT NOT NULL DEFAULT 1, -- fencing for takeover
locked_until TIMESTAMPTZ NOT NULL,
response_code INT,
response_body BYTEA, -- capped at 16 KB
resource_ref TEXT, -- e.g. 'ord_9' for debugging
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
PRIMARY KEY (tenant_id, endpoint, idem_key)
) PARTITION BY RANGE (created_at); -- daily partitions, drop at TTL + 1 day
A.2 Middleware and Phased Handler#
function handle(req):
key = req.header("Idempotency-Key") or return 400 missing_key
if key_timestamp(key) < now() - TTL: return 422 idempotency_key_expired
fp = sha256(canonical(req.method, req.path, req.body_without_volatile_fields))
scope = (req.auth.tenant_id, req.route, key)
row = INSERT … status='in_progress', locked_until=now()+LEASE
ON CONFLICT DO NOTHING RETURNING *
if row is null: # someone has it
existing = SELECT … WHERE scope
if existing.fingerprint != fp: return 422 idempotency_key_reused
if existing.status == 'completed': return replay(existing) # + Idempotent-Replayed: true
if existing.locked_until > now(): wait ≤ 1s for completion, else return 409 Retry-After: 1
row = UPDATE … SET locked_until=now()+LEASE, version=version+1
WHERE scope AND locked_until < now() AND version = existing.version RETURNING *
if row is null: return 409 Retry-After: 1 # lost takeover race
return run_phases(row, req)
function run_phases(row, req):
try:
if row.recovery_point == 'started':
tx: insert order(pending_payment); set recovery_point='order_created' WHERE version=row.version
if row.recovery_point == 'order_created':
result = psp.charge(key = row.idem_key + ":charge", …) # downstream dedupes
if result.declined:
tx: order→declined; complete(row, 402, body) # deterministic outcome, final
return 402
tx: order→paid; outbox.insert(OrderPlaced); complete(row, 201, body) WHERE version=row.version
return replay(row)
except BeforeAnySideEffect as e:
DELETE row WHERE scope AND version=row.version # release: nothing happened
raise e
except Unknown as e: # timeout, crash, 5xx downstream
leave row in_progress at its recovery point
return 503 Retry-After: 2 # client retries same key
Every tx that touches the record includes WHERE version = row.version; a zero-row update means another attempt took over and this one must stop.
A.3 Canonical Fingerprint#
| Include | Exclude |
|---|---|
| Method, route template, path params | Idempotency-Key itself |
| Body, canonical JSON (sorted keys, normalized numbers) | Trace IDs, request timestamps, client nonces |
Semantically relevant headers (e.g., Stripe-Account-style tenant switches) | User-agent, auth token (tenant is in scope already) |
Change the canonicalization only with a dual-check window (accept either old or new fingerprint) for one TTL.
Appendix B: Keys, Scope and Derived Keys
B.1 Key Generation by Client Type#
| Client | When to Mint | Where to Persist | Retry Horizon |
|---|---|---|---|
| Web app | On user commit (button press), bound to the form instance | Memory + sessionStorage | Minutes |
| Mobile app | On user commit | Local DB with the queued operation | Up to TTL − 1h |
| Partner server | When the business record is created | Their DB next to the record (or use their record ID as key) | Their batch policy, ≤ TTL |
| Internal service | Derived from the parent key | Not needed — deterministic | Parent's horizon |
B.2 Derived Keys#
root = client key k (e.g. 0190f3a2-…)
inventory = k + ":reserve"
payment = k + ":charge"
refund N = k + ":refund:" + N (N from our own state, not attempt count)
email = k + ":email:order_confirmation"
Rules: deterministic from the parent and the step, never from the attempt number; unique per distinct downstream effect; if the downstream limits key length, hash: base64url(sha256(k + ":charge"))[:40].
B.3 Natural Keys#
When the domain has an identifier that already means "this operation" — a partner's order reference, an invoice number, a (user, cart_version) pair — use it with a unique constraint and return the existing resource on conflict. It never expires, costs one index, and needs no separate table.
Appendix C: Outbox, Inbox and Consumer Idempotency
C.1 Producer to Consumer, End to End#
C.2 Consumer Idempotency Options#
| Option | Mechanism | Window | Cost | Use When |
|---|---|---|---|---|
| Inbox table, same tx | INSERT inbox(consumer, msg_id); conflict → skip | Retention you choose (≥ topic retention) | One row per message; partition and drop | Effect is a local DB write |
| State-guarded write | UPDATE … WHERE state = 'expected_prior' | Forever | None | Effect is a state transition |
| Versioned upsert | Apply only if event.version > row.version | Forever | Version column | Materialized views, projections |
| Naturally idempotent effect | Set, not increment | Forever | None | Projections, caches |
| Downstream key | Pass msg_id-derived key to external system | External system's retention | Depends on vendor | Effect is external (email, warehouse, PSP) |
C.3 Decision Tree#
C.4 Outbox Relay Choices#
| Relay | Latency | Ordering | Ops Cost | Use When |
|---|---|---|---|---|
Polling (SELECT … FOR UPDATE SKIP LOCKED) | 100ms–1s | Per aggregate if partitioned by aggregate | Low | < ~5K events/s per DB |
| CDC from WAL/binlog (e.g., Debezium outbox router) | 10–500ms | Commit order | Medium (connector fleet) | High volume; many services |
| Listen/notify + polling fallback | ~10ms | Per aggregate | Low | Low-latency needs on Postgres |
Appendix D: API Contract and Client SDK Behavior
D.1 Response Contract#
| Situation | Status | Headers / Body |
|---|---|---|
| First execution, success | 2xx | — |
| Replay of completed | Original status | Idempotent-Replayed: true |
| Same key in progress | 409 | Retry-After: 1, error: idempotency_key_in_progress |
| Same key, different payload | 422 | error: idempotency_key_reused |
| Missing key on required endpoint | 400 | error: idempotency_key_required |
| Key older than retention | 422 | error: idempotency_key_expired |
| Unknown outcome (server-side timeout after side effect possible) | 503 | Retry-After: 2 — retry with same key |
D.2 SDK Retry Algorithm#
op = Operation(key = uuidv7(), request, created_at = now())
persist(op)
for attempt in 1..MAX_ATTEMPTS (3):
if now() - op.created_at > TTL - 1h: surface UNKNOWN to caller; stop
if !retry_budget.allow(): surface UNKNOWN; stop
resp = send(op.request, Idempotency-Key = op.key, deadline = remaining_budget)
match resp:
2xx, 4xx except 409/429 → terminal; delete(op); return resp
409, 429, 503 → sleep(max(Retry-After, backoff(attempt) with full jitter))
network timeout → sleep(backoff(attempt) with full jitter)
surface UNKNOWN with op.key # caller can poll GET /operations/{key} or retry later
The SDK never mints a new key for an existing operation; a "retry" button in the UI re-sends the same persisted operation.
D.3 Operation Status Endpoint#
GET /v1/idempotency/{key} → { status: in_progress | completed | unknown_key, resource_ref } lets clients resolve UNKNOWN without re-sending a mutation — especially useful near the end of the retention window.
Appendix E: Observability
E.1 Core Metrics#
# Contract health
idempotency.requests_with_key_ratio{endpoint} # coverage
idempotency.replay_total / replay_ratio{client}
idempotency.conflict_409_total{endpoint}
idempotency.fingerprint_mismatch_total{client}
idempotency.expired_key_rejections_total
idempotency.key_age_at_request_seconds (histogram)
# Correctness health
idempotency.stuck_in_progress{age_bucket}
idempotency.takeover_total
ops.unknown_outcome_total{age_bucket}
dup_detector.suspected_duplicates_total{effect} # same user + amount + cart within 10 min, different keys
# Pipeline health
outbox.oldest_unpublished_age_seconds
consumer.dedup_hit_ratio{consumer}
rpc.retry_ratio{caller,callee}
E.2 Critical Alerts#
| Alert | Threshold | Severity |
|---|---|---|
| Stuck in-progress > 5 min | > 50 for 10 min | Page service owner |
| Unknown outcomes > 1h old | > 0.01% of ops | Page |
| Fingerprint mismatches by client | > 1% of client's requests | Ticket to partner engineering |
| Replay ratio by client | > 10% | Warn |
| Outbox oldest unpublished | > 60s | Page relay owner |
| Retry ratio | > 10% for 5 min on any edge | Warn; > 25% page |
| Suspected duplicates | above baseline × 3 | Page effect owner |
E.3 Debugging a Reported Duplicate#
- Do the two effects share an idempotency key? Yes → server-side bug (record not co-located, takeover without fencing, expiry). No → client minted two keys (per-attempt key, manual retry button not bound to the operation) or the dedup was bypassed by an un-keyed inner retry.
- Check derived keys on downstream calls: did the provider see two different keys?
- Check consumer side: was there an offset reset or rebalance near the time; what was
dedup_hit_ratio?
Appendix F: Scale Evolution
F.1 What Works at Each Scale#
| Scale | Approach | Breaks When |
|---|---|---|
| One service, one DB | Same-tx table, daily partitions | Other services call it without keys |
| 10–30 services | Platform library, derived keys, outbox, gateway enforcement | Consumers replay; multi-region |
| 100+ services, partners | Inbox paved road, replay tests, retention tiers, tenant-homed routing | — |
F.2 Multi-Region#
Mutations always land in the tenant's home region, so the idempotency record is authoritative there. Replicas serve reads and disaster recovery; they never accept a mutation for a tenant they don't home.
F.3 What You Don't Build on Day One#
Not on day one: a central idempotency service, retention tiers and per-tenant metering, the operation-status endpoint (until UNKNOWN volume justifies it), a CDC-based relay (polling suffices below ~5K events/s per DB), or cross-region globally consistent key tables.
Appendix G: Cost and Storage Sizing
G.1 Storage Sizing#
records/day × avg_bytes × retention_days
10M × 2 KB × 3 ≈ 60 GB (72h)
10M × 2 KB × 7 ≈ 140 GB (7d)
500M × 2 KB × 3 ≈ 3 TB (72h, org-wide)
Partition by day and drop whole partitions; never DELETE … WHERE created_at < … on a hot table.