Hiring BarSupport

Design a Payment System — Staff-Level Case Study

Case study75 min read9 diagrams

Technologies referenced in this case study: PostgreSQL · Apache Kafka · Redis · DynamoDB · API Gateways

Related: Reservation Systems · Flash Sales & Ticketing · Dealing with Contention · Managing Long-Running Processes · Consistency Models · API Design Patterns

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and deep dives that match your weak spots.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal Lens) and the appendices on ledgers and reconciliation
What is a Payment System? — Why interviewers pick this topic

A payment system moves money on behalf of users: it accepts a charge request, talks to a Payment Service Provider (PSP) such as Stripe, Adyen or Braintree (who talks to card networks and banks), records what happened, and later proves — to finance, auditors and regulators — that every cent is accounted for.

The hard part is not calling POST /charges. The hard part is that the network between you and the PSP can fail after the money moved, and you must never charge twice, never lose a charge, and always be able to explain the balance.

Before vs After — the "timeout during checkout" scenario:

Without idempotency + ledger:
t=0:       User clicks "Pay $120". Checkout service calls PSP.
t=+8s:     PSP authorizes the card. Response is lost (LB reset).
t=+10s:    Client timeout. Service returns 500. User sees "Payment failed".
t=+15s:    User clicks "Pay" again. New PSP call. Second $120 authorization + capture.
t=+2 days: User sees two charges on the statement. Files a dispute.
t=+30 days: Chargeback: $120 refunded + $15 dispute fee. Support ticket. Trust gone.
            Finance finds a $120 hole nobody can explain for 3 weeks.

With idempotency keys + state machine + reconciliation:
t=0:       Checkout creates payment_intent pi_123 (state=CREATED) with key k_abc.
t=+8s:     PSP authorizes. Response lost.
t=+10s:    Timeout → pi_123 moves to UNKNOWN, not FAILED. User sees "Confirming…".
t=+15s:    Retry with the SAME idempotency key k_abc → PSP returns the original result.
t=+15s:    pi_123 → AUTHORIZED. Ledger posts balanced entries. User sees "Paid".
t=+1 day:  Nightly reconciliation matches PSP settlement file to ledger: 0 breaks.

Why interviewers reach for this question: Payments is the purest test of whether a candidate can reason about unknown outcomes. Every other system gets to say "retry". Here, a retry can move money twice, and a missing retry can lose a sale. The candidate who says "exactly-once" without naming the mechanism, the state machine and the reconciliation loop is revealing that they have never owned money movement.

Mechanics Refresher: The Money-Movement Primitives
PrimitiveHow It WorksProsCons
Authorize + Capture (two-step)Auth places a hold on the card (valid ~7 days for most card types); capture moves the money laterCharge only when you ship; void costs nothingHolds expire; re-auth may fail; two calls to reconcile
Sale (auth+capture in one call)Single call, funds captured immediatelySimple; fewer statesRefund instead of void when order cancels — refunds cost fees and take 5–10 days to land
Idempotency keyClient-generated key; server stores key → response; replays return the stored responseMakes retries safeNeeds durable storage with TTL (Stripe retains keys ≥24h); key scope must be designed
Double-entry ledgerEvery movement is ≥2 entries whose debits = credits; balances are derivedInvariant is checkable; audit-proofNeeds account modeling; immutable, so corrections are new entries
Outbox patternWrite business row + event row in one DB transaction; relay publishes eventNo dual-write gapRelay lag (100ms–seconds); consumers must be idempotent
ReconciliationCompare your ledger to PSP reports and bank statements; flag breaksCatches what every other mechanism missedBatch (T+1); needs owner and SLA for breaks
Tokenization / vaultPSP or vault stores the PAN; you store a tokenShrinks PCI DSS scope from SAQ D to SAQ A/A-EPVendor lock-in on stored cards

For most production systems: Two-step auth/capture through one or two PSPs, client-supplied idempotency keys on every mutating call, an internal append-only double-entry ledger, an outbox to publish payment events, and daily three-way reconciliation. The primitives are not the interview — the handling of unknown outcomes is.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What This Interview Actually Tests#

A payment system is not an API integration question. Everyone can call Stripe.

It is a correctness-under-uncertainty question that tests:

  • Whether you treat "timeout" as a third outcome, distinct from success and failure
  • Whether you can make retries safe without distributed transactions
  • Whether you know the difference between your record of money and the PSP's record — and who wins when they disagree
  • Whether you design for the auditor, the finance team and the on-call — not only the checkout button

The key insight: You cannot get exactly-once money movement across a network. You get at-least-once delivery + idempotent effect + reconciliation to catch the residue. Staff candidates name all three; Senior candidates name the first one and call it done.

The L5 vs L6 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws checkout → payment service → Stripe → DBAsks "Are we the merchant, a marketplace, or the processor? Who holds the money?"Asks "Is payments a product feature or a regulated platform? Who is the licensed entity, and which team carries the compliance obligation?"
Retries"Retry with exponential backoff""Retries are only safe with an idempotency key scoped to the payment intent; a timeout moves the payment to UNKNOWN, never FAILED"Sets an org-wide rule: no service may call a PSP without the shared idempotency middleware; enforces it in the PSP client library and in egress policy
Source of truth"The payments table""An append-only double-entry ledger is our truth for balances; the PSP is truth for what the card network did; reconciliation arbitrates"Makes the ledger a company-level system of record with finance as a co-owner; ledger schema changes need controller sign-off
Failure"Add a DB replica and retries""PSP down → queue captures, fail over auth to a second PSP for new payments only, never re-drive an UNKNOWN on a different PSP"Prices multi-PSP: +1 integration team (~3 engineers), +0.1–0.3% fees lost to weaker routing volume discounts, vs expected loss from a 4-hour single-PSP outage
OwnershipPayments team owns everythingPayments owns the state machine; finance owns reconciliation breaks with a 2-business-day SLA; risk owns fraud rulesRedraws boundaries: ledger as a platform, money movement as a platform, product teams consume intents; defines the contract so 12 teams stop integrating PSPs directly
Scale"Shard the payments table""Payments are ~100 TPS for most businesses; the scaling problem is the ledger's hot accounts, not throughput"Plans the 3-year path: single-region Postgres → ledger service with account sharding → multi-entity, multi-currency with regional data residency
Why "retries" separates levels

L5: "We retry failed PSP calls with exponential backoff and jitter." This is correct advice for idempotent reads. For a charge, it is the most common cause of double-charging in production. The L5 answer does not distinguish between failed (PSP said no — safe to show error) and unknown (we don't know — must not create a new charge).

L6: "Every call to the PSP carries an idempotency key derived from our payment intent ID and the attempt's operation, like pi_123:authorize. If we time out, the intent goes to UNKNOWN. A background resolver retries the same call with the same key — the PSP returns the original result — or queries the PSP by our reference. We never show 'failed' to the user for an UNKNOWN, and we never generate a fresh key for the same logical operation."

L7: "The failure mode here is organizational: the fifth team to integrate a PSP will write their own retry loop. I'd make the idempotency middleware the only way to reach a PSP — it lives in the payments client library, and egress to PSP endpoints is blocked for any service that isn't the money-movement service."

Why "source of truth" separates levels

L5: Uses a mutable payments table with a status column and an amount column. Balances are computed by summing rows with status='succeeded'. This works until the first partial refund, currency conversion, or chargeback — and then the balance is "whatever the last UPDATE said".

L6: Separates the state machine (what stage is this payment in?) from the ledger (what money moved, between which accounts?). The ledger is append-only, double-entry, and every entry references the event that caused it. Corrections are reversing entries, never UPDATEs. "If finance asks why the merchant balance is $4,212.50, I can replay the ledger entries and show every cent."

L7: Recognizes that the ledger is a company asset that will outlive this payment system. Designs the chart of accounts with finance, versions it, and makes ledger immutability an audit control (SOX, if public). "The ledger schema is a one-way door. Getting account modeling wrong costs a 9-month migration later."

Why "failure" separates levels

L5: "If Stripe is down, we fail over to Adyen." It sounds resilient. It is dangerous: a payment whose Stripe call is UNKNOWN might have succeeded. Re-driving it on Adyen charges the customer twice across two processors — and neither processor's idempotency can see the other.

L6: "Failover applies only to new payment intents that never touched the primary PSP. Anything UNKNOWN stays pinned to its original PSP until resolved. Captures of already-authorized payments queue until the original PSP recovers — an auth hold lasts days, so a 2-hour queue is fine."

L7: Quantifies whether failover is worth building at all: "At $2M/day GMV, a 4-hour PSP outage costs ~$330K in delayed or lost sales — maybe 30% truly lost, so ~$100K. A second PSP costs ~3 engineers ongoing plus volume-discount erosion. At $2M/day, I'd not build it yet. At $50M/day, it's mandatory."

The Staff Positions#

PositionRationale
Timeout = UNKNOWN, never FAILEDA timeout carries no information about whether money moved; resolving it is a separate, owned process
Idempotency key per logical operation, generated before the first attemptKeys created on retry are useless; the key must exist before the network call
Append-only double-entry ledger, separate from the state machineBalances become derivable and auditable; corrections are new entries, not UPDATEs
Two-step auth/capture for physical goodsVoids are free; refunds cost fees and 5–10 days; capture on fulfillment
Outbox over dual writes"Write DB then publish Kafka" loses events on crash; outbox makes the event part of the transaction
Reconciliation is a product, not a cron jobEvery mechanism above leaks at 0.01–0.1%; reconciliation with an owned break queue is the safety net
Never store PANsTokenize at the PSP or a vault; PCI scope drops from ~300 controls (SAQ D) to ~20 (SAQ A)

The Three Intents#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Merchant checkout (pay-in)Conversion + no double chargeIntegrate 1–2 PSPs, idempotent intents, ledger for revenue/refundsDouble charge, lost sale on timeoutZero duplicate captures; reconciliation breaks < 0.01%
Marketplace (pay-in + payout)Money belongs to third parties; KYC; payout timingLedger with per-party accounts, escrow/holding, scheduled payouts, split paymentsPaying out money you haven't collected; negative balancesLedger must balance to the cent daily; regulatory reporting
Being the processor (PSP)Card-network protocols, 99.99%+ availability, PCI Level 1Direct network connections (ISO 8583), HSMs, own vault, risk engineNetwork-level outages, settlement mismatches at scaleRegulatory audit; four-party settlement

🎯 Staff Move: "I'll assume we're a merchant integrating PSPs for checkout — pay-in only — and I'll design the ledger so payouts can be added later. If we were a marketplace, the ledger becomes the center of gravity; if we were the processor, this is a different interview about ISO 8583 and HSMs. Tell me if you want one of those instead."

The Five Fault Lines#

#Fault LineThe Tension
1Retry vs Reconcile (the Unknown Outcome)Resolve a timeout by retrying fast (risk duplicates if keys are wrong) or by waiting for reconciliation (risk a stuck user)?
2Ledger vs PSP as Source of TruthWhose record wins when your ledger and the PSP disagree — and who fixes it?
3Synchronous Confirmation vs Asynchronous SettlementTell the user "paid" at authorization, or wait for capture/settlement? Conversion vs certainty
4Single PSP vs Multi-PSP RoutingSimplicity and volume pricing vs availability and cost optimization
5Fraud/Risk: Fail-Open vs Fail-ClosedWhen the risk engine is down, lose sales or accept fraud losses? Who signs off?

In the Wild: Real Production Systems#

Why this section belongs here: Naming how real payment companies solved these problems shows you've studied operations, not tutorials.

Stripe — Idempotency Keys as a Public API Contract#

Stripe's API accepts an Idempotency-Key header on every POST. The server stores the key with the first request's result and returns the same response for replays — including errors — and rejects a replay whose parameters differ from the original. Keys are retained for at least 24 hours. Stripe's engineering writing on idempotency describes the server side: record the key and request in a transaction, split the work into atomic phases with recovery points, and have a background process complete or roll back abandoned requests.

Staff insight: Idempotency is not "dedup on the server". It is a contract between client and server: the client promises to reuse the key for the same intent, the server promises the same answer. In an interview, say both halves.

Airbnb — Orpheus, Idempotency as a Library#

Airbnb publicly described avoiding double payments in their SOA migration with an idempotency library (Orpheus). Each request is split into three phases — pre-RPC (record intent in the DB), RPC (call the external processor), post-RPC (record the result) — and the library forbids mixing database writes into the network phase. Failures are classified as retryable or non-retryable, and the classification is persisted so a retry never re-does a non-retryable failure.

Staff insight: The design point is separating local transactions from network calls. You cannot hold a DB transaction open across a 3-second PSP call. Name the three phases and you've shown you've built this.

Uber — Ledger-First Money Movement#

Uber has written publicly about its payments platform and its purpose-built ledger storage (LedgerStore), designed for immutable, verifiable ledger entries at very high volume, including sealing of time ranges so historical entries cannot change. Money movement at Uber is modeled as ledger entries between accounts (riders, drivers, Uber) rather than as mutable balances.

Staff insight: At scale the ledger becomes its own storage system with its own guarantees — immutability, completeness checks, time-range sealing. In an interview, you don't build LedgerStore; you say "the ledger's invariants are append-only and sum-to-zero, and I'd pick storage that makes violating them hard."

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"We'll retry on failure""The PSP call timed out after 10s. Did the customer get charged?"Do you know timeout ≠ failure?
"We use idempotency keys""Who generates the key, when, and what's its scope? What if the user clicks Pay twice in two tabs?"Key design, not just the word
"We store payments in Postgres""Finance says the merchant balance is off by $312. How do you find out why?"Ledger vs mutable status table
"We publish an event to Kafka after saving""The service crashes between the DB commit and the publish. What happens?"Dual-write awareness, outbox
"We fail over to a second PSP""What about the payment that was in-flight on the first PSP?"Pinning UNKNOWNs, cross-PSP duplicates
"Reconciliation runs nightly""Who looks at the breaks? What's the SLA? What happens to unresolved ones?"Reconciliation as an owned process

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: The hot path is short on purpose: gateway → payment service → risk → PSP. Everything that can be async is: ledger posting (via outbox + Kafka), webhook ingestion, resolution of UNKNOWN intents, and reconciliation. The two numbers that matter most on the observability panel are payment.unknown_count (how much money is in limbo right now) and ledger.imbalance_total (must be exactly zero; any non-zero value is a Sev-1).

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Retries"Exponential backoff""Same idempotency key on every retry of the same operation. Timeout → UNKNOWN → resolver. Never a new key."
Exactly-once"Kafka exactly-once semantics""Exactly-once effect: at-least-once delivery + idempotent consumers + reconciliation for the residue."
Data model"payments table with status""Intent state machine + attempts table + append-only double-entry ledger. Balances are derived."
Events"Save to DB, then publish""Transactional outbox. Consumers dedup by event ID."
PSP outage"Fail over to backup PSP""Fail over new intents only. UNKNOWNs stay pinned. Captures queue — holds last ~7 days."
Truth"Our DB is correct""Our ledger is truth for our books; the PSP is truth for the network; daily reconciliation, owned breaks, 2-day SLA."
Compliance"Encrypt card numbers""Never touch PANs. Tokenize client-side via PSP SDK. PCI scope becomes SAQ A."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
PSP authorization latencyp50 300–800ms, p99 2–5sSet PSP client timeouts at 10–30s, not 1s — a short timeout manufactures UNKNOWNs
Card authorization hold validity~7 days typical (varies by network, MCC)Captures can queue for hours during a PSP outage without losing the sale
Stripe idempotency key retention≥ 24 hoursYour own key store needs a TTL at least as long as your longest retry horizon
Typical merchant payment volume10–500 TPS peakThroughput is rarely the problem; correctness and hot ledger accounts are
VisaNet stated capacity~65,000 transaction messages/secEven the largest network is "only" tens of thousands TPS — size claims accordingly
Settlement timingT+1 to T+3 business daysReconciliation runs on settlement files, not real-time responses
Chargeback windowup to ~120 days after transactionLedger and evidence must be retained and queryable at least that long
Dispute-rate monitoring thresholds~0.9–1.5% of transactions (network programs)Crossing them means fines and possible loss of card acceptance
Standard US online card pricing (Stripe list)2.9% + $0.30Why multi-PSP routing and interchange optimization become worth engineers at scale
Unknown-outcome rate~0.01% normally, 1–5% during PSP incidentsSize the resolver and the "Confirming…" UX for the incident, not the steady state
Refund landing time5–10 business daysWhy void (pre-capture) beats refund (post-capture) for cancellations

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on the checkout API and the PSP call, then run out of time before the interviewer asks the only question that matters: "The call timed out. Did we charge them?" Compress the basics to ~10 minutes and spend the rest on unknown outcomes, the ledger and reconciliation.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Users pay for orders with cards and wallets. We authorize at checkout, capture on shipment, support full and partial refunds, and handle disputes. We integrate PSPs — we are not the processor."

Then spend the time on non-functional requirements, which is where the design lives:

"Three constraints drive everything. One: no double charges and no lost charges — correctness beats latency. Two: every cent must be explainable to finance, so I need an auditable record, not a status column. Three: we depend on an external PSP whose p99 is seconds and whose outages we don't control. I'll also assume roughly 200 payments per second at peak, which means throughput is not the hard part."

Then name the underspecified parts:

"A few things I'd confirm: single currency or multi-currency? Are we a marketplace paying out to sellers? Do we store cards for repeat purchase? I'll assume single currency to start, merchant-of-record, stored cards via PSP tokens."

🎯 Staff Move: Saying "throughput is not the hard part" early is a strong signal. It tells the interviewer you won't waste 10 minutes sharding a table that handles 200 TPS, and it redirects the session to correctness.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • PaymentIntent: intent_id, order_id, amount_minor, currency, state, psp, idempotency_key, version
  • PaymentAttempt: one row per PSP call — attempt_id, intent_id, operation (authorize/capture/refund/void), psp_reference, outcome (SUCCEEDED / DECLINED / UNKNOWN), request_hash
  • LedgerEntry: txn_id, account_id, direction (debit/credit), amount_minor, currency, caused_by_event_id, posted_at
  • IdempotencyRecord: key, scope, request_hash, response_blob, status, expires_at

Money is always stored as integer minor units (cents) with an explicit currency. Never floats.

Client-facing API:

POST /v1/payment_intents                Idempotency-Key: <uuid>
     { order_id, amount_minor, currency, payment_method_token }
  → 201 { intent_id, state: "REQUIRES_CONFIRMATION" }

POST /v1/payment_intents/{id}/confirm   Idempotency-Key: <uuid>
  → 200 { state: "AUTHORIZED" | "DECLINED" | "PROCESSING" }

POST /v1/payment_intents/{id}/capture   Idempotency-Key: <uuid>   { amount_minor? }
POST /v1/payment_intents/{id}/refunds   Idempotency-Key: <uuid>   { amount_minor, reason }
GET  /v1/payment_intents/{id}

PROCESSING is the honest client-facing name for UNKNOWN. The client polls GET or receives a push; it never re-submits a new intent.

🎯 Staff Move: "I'm separating intent creation from confirmation. The intent exists before any money moves, so the idempotency key and the order linkage are durable before the first network call to the PSP. That ordering is the whole trick."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk the confirm flow in 90 seconds:

  1. Client confirms intent pi_123 with idempotency key k1. Gateway requires the header.
  2. Payment service looks up k1. Hit with same request hash → return stored response. Hit with different hash → 422. Miss → insert record as IN_PROGRESS (row lock).
  3. Risk scores the payment (p99 < 100ms). Decline → terminal.
  4. Pre-call: write PaymentAttempt(operation=authorize, outcome=PENDING) and commit.
  5. Call: PSP authorize with PSP-level idempotency key pi_123:authorize:1.
  6. Post-call: in one transaction — update attempt outcome, move intent state, write outbox event, store idempotency response.
  7. Outbox relay publishes payment.authorized; ledger consumer posts balanced entries.
  8. On timeout at step 5: intent → UNKNOWN, respond PROCESSING, resolver owns it from here.

🎯 Staff Move: Say out loud: "Notice there's no database transaction open across the PSP call. Pre-call commit, network call, post-call commit. If we crash between them, the PENDING attempt row tells the resolver exactly what to go check." You've now spent ~9 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path, and it's the Senior-level design. What makes this hard is that the PSP call can fail after money moved. I'd like to go deep on three things: how we resolve unknown outcomes without double-charging, how the ledger stays balanced and auditable, and how reconciliation catches what everything else misses. Where would you like to start?"

If no preference: start with the unknown outcome. It's the question that decides the level.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: The unknown outcome (7–8 min)

"A timeout tells us nothing. The PSP may have authorized, declined, or never received the request. Three rules. First, the intent goes to UNKNOWN and the user sees 'Confirming your payment', not 'failed'. Second, the resolver retries the identical call with the identical PSP idempotency key at 1 minute, 5 minutes and 30 minutes — the PSP either returns the original result or processes it for the first time, and both are correct. Third, if the PSP supports lookup by our reference, we query instead of retrying. After 24 hours unresolved, it goes to a human queue owned by payments ops."

Quantify: "Normally ~0.01% of calls end UNKNOWN — at 200 TPS that's ~2 per 1,000 seconds, trivially handled. During a PSP brownout it can hit 5%, which is 10/sec, 36,000/hour. The resolver must be rate-limited so it doesn't DDoS a recovering PSP, and the UX must hold 'Confirming' for minutes."

Who pays: "The customer pays a few minutes of uncertainty. That's the right victim — the alternative victim is the customer charged twice."


Deep dive 2: The ledger (7–8 min)

"State machine answers 'what stage is this payment in'. The ledger answers 'where is the money'. They're different questions. For an authorized-then-captured $100 order with a 2.9% + 30¢ fee:"

txn_1 (capture):   DEBIT  psp_receivable        10000
                   CREDIT customer_payments     10000   -- revenue-pending
txn_2 (fee):       DEBIT  psp_fees                320
                   CREDIT psp_receivable          320
txn_3 (settle):    DEBIT  bank_cash              9680
                   CREDIT psp_receivable         9680
                   -- psp_receivable now nets to 0: settled in full

"Every transaction sums to zero. psp_receivable should net to zero once settlement lands — if it doesn't after T+3, that's a reconciliation break with a name. Corrections are new reversing transactions, never UPDATEs."


Deep dive 3: Reconciliation (5–6 min)

"Three-way: our ledger vs the PSP's settlement report vs the bank statement. Match on PSP reference and amount. Unmatched on either side is a break. Breaks have a type — missing in ledger, missing at PSP, amount mismatch, currency mismatch — and each type has an owner and an auto-resolution rule. Target: < 0.01% of transactions break; breaks older than 2 business days page the finance-ops lead."


Deep dive 4: Events without dual writes (3–4 min)

"The payment service writes the state change and an outbox row in one Postgres transaction. A relay tails the outbox and publishes to Kafka with the outbox row ID as the event ID. Consumers — ledger, notifications, fulfillment — dedup by event ID. Crash anywhere, and the worst case is a duplicate publish, which consumers absorb."


Deep dive 5: Ownership (3–4 min)

"Payments team owns the state machine, PSP integrations and the resolver, and carries the pager for payment.unknown_count. Finance-ops owns reconciliation breaks. Risk owns fraud rules and the fail-open/closed decision for the risk engine — with a CFO-level sign-off because it's a direct loss decision. Product teams never call a PSP directly."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: exactly-once money movement isn't available, so we build exactly-once effect from three parts — idempotency keys that make retries safe, a state machine that treats timeouts as unknown, and reconciliation that catches the residue. The ledger makes the whole thing auditable."

The evolution closer:

"What I'd build later: a second PSP when GMV justifies it, which I'd estimate around $20–50M/day; multi-currency with FX as ledger entries; payouts if we become a marketplace. What I'd never build: our own card-network connectivity, unless payments becomes the product."

🎯 Staff Move: End on who owns the residue. Every Senior candidate ends with "and we'd add monitoring." Staff candidates end with "and finance-ops owns reconciliation breaks with a 2-day SLA, because that's where the money that slips through our mechanisms gets found."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
API over-design10 min on REST resources and field namesNames 4 entities in 30s; spends time on the confirm flow
Card network tourExplains issuer/acquirer/interchange at lengthOne sentence: "PSP abstracts the four-party model; I'll treat it as an unreliable remote"
Sharding theaterShards the payments table for "millions of TPS""~200 TPS. One Postgres primary. The hot spot is ledger accounts, not rows."
No timeout storyWaits for "what if the PSP times out?"Volunteers UNKNOWN state during the architecture phase
No ledgerBalance = SUM(payments.amount)Separates state machine from ledger in the first 15 minutes
No reconciliationAssumes mechanisms are perfect"Reconciliation is the safety net; here's who owns breaks"

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Payments is the interview where "it's eventually consistent" is not an acceptable answer by itself and "use a distributed transaction" is not available. The PSP will not join your two-phase commit. The candidate must invent correctness from local transactions, idempotency and after-the-fact verification — and must name who owns each leak. That is Staff work: reasoning about systems you don't control and outcomes you can't observe directly.

It also has an unusually expensive failure distribution. A lost cache entry costs milliseconds. A duplicate charge costs a chargeback ($15–25 fee plus the amount), a support ticket ($5–15 to handle), and — at the network level — dispute ratios that can trigger monitoring programs. A missing charge is revenue given away. Both are silent: no error fires when a customer is charged twice.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"The PSP call timed out after 10 seconds. The user is looking at a spinner. What do you show them, what state is the payment in, and what happens in the next 24 hours?"

A candidate who answers with a state name (UNKNOWN), a user-facing message ("Confirming"), a resolution mechanism (same-key retry or lookup), a time bound (24h to human queue) and an owner (payments ops) has operated a payment system. A candidate who says "retry with backoff" has integrated one.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Merchant checkout (pay-in) → conversion + no duplicates

  • Constraint: checkout latency adds directly to abandonment; ~1s of added latency is a measurable conversion hit
  • Strategy: PSP-hosted tokenization, two-step auth/capture, idempotent intents, internal ledger for revenue, refunds and fees
  • Failure mode: double charge on retry, lost sale on a timeout shown as failure
  • Who pays for imperfection: customers (duplicates) or the business (lost sales, chargeback fees)

Marketplace (pay-in + payout) → the ledger is the product

  • Constraint: you hold other people's money; payouts, KYC, 1099/tax reporting, possibly money-transmitter licensing depending on structure
  • Strategy: ledger with accounts per buyer, seller and platform; funds flow buyer → holding → seller minus fees; payouts scheduled after risk holds (e.g., 24h after check-in for lodging)
  • Failure mode: paying out before collecting, negative seller balances after chargebacks
  • Who pays: the platform eats every negative balance it can't recover

Being the processor → network protocol and regulatory engineering

  • Constraint: ISO 8583 messages to card networks, HSMs for PIN/key management, PCI DSS Level 1 (>6M transactions/year), four-nines availability
  • Strategy: an entirely different system — acquiring, authorization switching, clearing files, settlement
  • Failure mode: network-wide authorization failures, settlement file mismatches in the millions
  • Who pays: every merchant on the platform

2.2 When NOT to Build a Payment System#

  • You have one product, one currency, one PSP and < $1M/month. Use the PSP's hosted checkout (Stripe Checkout, Adyen Drop-in) plus their dashboard. Your "payment system" is a webhook handler and an orders.paid_at column. Building an internal ledger here is ~2 engineer-quarters for no auditable benefit.
  • You want a ledger for non-money things (loyalty points you can't redeem for cash, internal credits). A double-entry ledger is still a good model, but you don't need PSP reconciliation, PCI scope or UNKNOWN resolution — don't import the machinery.
  • Subscriptions. Billing (plans, proration, dunning, invoices) is a separate system on top of payments. Buy it (Stripe Billing, Chargebee, Zuora) until billing logic is a competitive advantage.
  • You're tempted to become a payment facilitator to save 0.3% in fees. Payfac status brings underwriting, KYC, risk and liability for sub-merchants' chargebacks. That's a company decision, not an architecture decision.

🎯 Staff Insight: "The cheapest payment system is the PSP's hosted page. I'd build our own intent/ledger layer only when we need something the PSP can't give us: multi-PSP routing, marketplace splits, or finance-grade books across processors."

2.3 What the Interviewer Leaves Underspecified#

Interviewers deliberately omit:

  • Who holds the money (merchant vs marketplace vs processor) — changes the ledger and the licensing
  • When the user is told "paid" — at authorization, capture or settlement
  • Currency and geography — multi-currency means FX entries; the EU means Strong Customer Authentication (3-D Secure) and asynchronous confirmation flows
  • Stored credentials — whether you keep cards on file, which changes PCI scope and adds network tokens
  • Refund and dispute policy — partial refunds and chargebacks are where mutable-status designs break
  • Peak shape — a flash sale is 50× baseline for 10 minutes; a subscription renewal batch is 100× at 00:00 UTC

Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.

2.4 Precise Terminology#

TermWhat It MeansWhy It Matters in the Interview
AuthorizationIssuer approves and places a hold; no money moves yetHolds expire (~7 days); void is free
CaptureMerchant claims the authorized fundsCan be partial; can be multiple for split shipments (PSP-dependent)
SettlementFunds actually arrive in the merchant bank account, net of feesT+1 to T+3; what reconciliation matches against
VoidCancel an auth before captureNo fee, no statement line for the customer
RefundReturn captured moneyCosts time (5–10 days) and often the original fee
Chargeback / disputeCardholder disputes via their bankMoney is pulled back, plus a fee; you submit evidence
Idempotency keyClient-generated identifier for a logical operationScope and lifetime are design decisions
Payment intentYour durable record of "we intend to collect X for order Y"Exists before any PSP call; carries the state machine
UNKNOWNA call whose result you didn't observeNot a failure; resolved by retry-with-same-key, lookup, webhook or recon
BreakA reconciliation mismatchNeeds a type, owner and SLA
PANPrimary Account Number (the card number)Touching it puts you in full PCI DSS scope

🎯 Staff Insight: If the interviewer says "when the payment succeeds", ask: "Succeeds at which stage — authorization, capture, or settlement? They happen days apart and fail differently."


3. The Five Fault Lines#

Every payment design decision has a technical side (what state, which store, which retry) and an organizational side (who signs off on losses, who works the break queue, who gets paged). Interviewers grade the second side.

3.1 Fault Line 1: Retry vs Reconcile — The Unknown Outcome#

The tension: Resolving an UNKNOWN quickly (aggressive same-key retries, synchronous lookup) gets the customer an answer in seconds but loads a PSP that may be failing. Waiting for webhooks or reconciliation is safe but can leave a customer staring at "Confirming" for hours — or leave an order unshipped.

ChoiceWhat WorksWhat BreaksWho Pays
Synchronous retry with same key (inline)Most UNKNOWNs resolve in < 30sHolds checkout threads; amplifies load on a degraded PSPInfra/on-call (thread exhaustion)
Async resolver (1m / 5m / 30m backoff)Bounded load; survives PSP brownoutsCustomer waits minutes; needs push or polling UXCustomer (uncertainty)
Webhook-onlyZero extra PSP callsWebhooks are at-least-once and can be delayed hoursSupport (tickets from stuck orders)
Reconciliation-only (T+1)AuthoritativeA day of limbo; fulfillment blockedBusiness (delayed shipping, abandonment)
Diagram: 3.1 Fault Line 1: Retry vs Reconcile — The Unknown Outcome

Staff default: "Layered. One inline retry with the same key after 2 seconds if the checkout budget allows; otherwise hand to the async resolver at 1m, 5m, 30m, 2h; webhooks resolve whichever comes first; reconciliation is the backstop. After 24 hours unresolved, a human queue."

When to deviate:

  • In-person/POS payments: the customer is at the till — resolve synchronously or void-and-retry with explicit reversal messages, because "check back later" isn't a UX.
  • Low-value, high-volume (micro-transactions): skip inline retry entirely; rely on async + batch reconciliation. The cost of a stuck $0.99 purchase is below the cost of the extra PSP load.
  • PSP without idempotency support: lookup-by-reference before any retry. If neither is supported, you cannot retry safely at all — escalate that as a vendor-selection problem.

🧭 Principal Move: "I'd make 'unknown' a first-class concept across every money-moving system in the company — payments, payouts, refunds, gift cards. Same state name, same resolver framework, same dashboard of money-in-limbo by system. The metric money_in_unknown_usd becomes a finance KPI, not a payments-team internal."

❌ Common L5 Trap: "If the call times out, we'll retry with exponential backoff up to 3 times." The interviewer asks: "With a new request each time?" If the key is generated per HTTP request — a common bug in hand-rolled clients — each retry is a new charge. The fix is structural: the key is derived from the persisted attempt row, so a retry cannot produce a new key.


3.2 Fault Line 2: Ledger vs PSP as Source of Truth#

The tension: The PSP knows what the card network did. You know why (which order, which refund, which fee split). Neither alone can answer "what does the business have?" When they disagree, someone has to decide which record is corrected.

ChoiceWhat WorksWhat BreaksWho Pays
PSP is truth (no internal ledger)Zero ledger engineering; PSP dashboard is the booksMulti-PSP impossible to consolidate; no marketplace splits; finance exports CSVsFinance (manual month-end, weeks of close)
Mutable payments tableSimple queriesPartial refunds, disputes and FX corrupt balances; no audit trailAuditors / on-call (forensics by UPDATE log)
Internal double-entry ledger + reconciliationAuditable; multi-PSP; derived balancesLedger modeling effort; recon pipeline to ownPayments + finance-ops (shared ownership)
Diagram: 3.2 Fault Line 2: Ledger vs PSP as Source of Truth

Staff default: "Internal ledger is the truth for our books; the PSP is the truth for what happened on the network. When they disagree, we never edit either — we post an adjusting entry in our ledger with a reference to the break ticket. The break type decides the owner."

When to deviate:

  • Early stage, one PSP, no marketplace: PSP-as-truth with nightly export is fine. Write the intent/state machine now, defer the ledger until month-end close takes more than 3 days or you add a second PSP.
  • Regulated entities (e-money, banking): the ledger must be the legal record with safeguarding reconciliation daily — deviation is in the direction of more rigor, sometimes intraday.

🧭 Principal Move: "Month-end close time is the metric I'd use to justify the ledger to the CFO. If finance spends 8 days closing because payments data doesn't reconcile, a ledger that gets it to 2 days pays for its team."

❌ Common L5 Trap: "We'll compute the merchant balance with SELECT SUM(amount) FROM payments WHERE status='succeeded'." Then: a $100 payment is partially refunded $30, disputed, the dispute is won, and there's a 2.9% fee. What's the balance? The mutable-status model has no answer that survives an audit. The ledger answer is "sum the entries in the account", and every entry has a cause.


3.3 Fault Line 3: Synchronous Confirmation vs Asynchronous Settlement#

The tension: Users and product want "Paid ✓" instantly. Money isn't final until capture (hours–days) or settlement (T+1–T+3), and 3-D Secure, bank transfers and wallets confirm asynchronously anyway.

ChoiceWhat WorksWhat BreaksWho Pays
"Paid" at authorizationFast UX, familiarCapture may fail (expired hold, ~0.1–1%); fulfillment ships before money is certainBusiness (unrecoverable shipped goods)
"Paid" at captureMoney claimedDelays confirmation for auth-then-ship flowsProduct (UX copy, "Order placed" vs "Paid")
"Paid" at settlementCertaintyDays of delay; unacceptable for checkoutCustomer (confusion)
Explicit async states in UXHonest; supports 3DS, ACH, walletsRequires product to design pending statesProduct/design (up-front work)
Diagram: 3.3 Fault Line 3: Synchronous Confirmation vs Asynchronous Settlement

Staff default: "The order is 'placed' at authorization; money is 'collected' at capture; the ledger recognizes revenue per finance's policy. The UI never says 'paid' for an UNKNOWN. Every transition is a conditional update — UPDATE … SET state='CAPTURED', version=version+1 WHERE id=? AND state='AUTHORIZED' AND version=? — so a duplicate webhook can't move a state twice."

When to deviate:

  • Digital goods delivered instantly: use sale (auth+capture) and treat authorization as "paid". Refund risk is lower than the UX cost of a pending state.
  • Long fulfillment (> 7 days, e.g., pre-orders): hold expiry forces either re-authorization (which can fail) or charge-at-order with refund on cancellation. Pick per product with finance.

🎯 Staff Insight: "The state machine is the contract between payments and every other team. Fulfillment listens for CAPTURED, not AUTHORIZED. Support tooling shows UNKNOWN distinctly. If I let each team infer state from different events, I've created five state machines."


3.4 Fault Line 4: Single PSP vs Multi-PSP Routing#

The tension: One PSP means one integration, one reconciliation format, the best volume pricing and one outage away from zero revenue. Multiple PSPs buy availability, authorization-rate optimization and negotiating leverage, at the cost of a routing layer, N reconciliation pipelines and the cross-PSP duplicate risk.

ChoiceWhat WorksWhat BreaksWho Pays
Single PSPSimplest; best tiered pricing; one recon formatPSP outage = checkout outage; no leverageBusiness (revenue during outages)
Active-passive failoverSurvives PSP outage for new intentsStored-card tokens are PSP-specific; secondary gets cold traffic, lower auth ratesPayments team (second integration, drills)
Active-active routing (by cost/auth rate)1–3% auth-rate uplift reported by routing vendors; fee optimizationComplex; needs network tokens or a vault for portability; N recon pipelinesPayments platform (a dedicated team)
Diagram: 3.4 Fault Line 4: Single PSP vs Multi-PSP Routing

Staff default: "Single PSP until the math says otherwise. When we add a second, it's active-passive for new intents only, with stored cards made portable via network tokens or a PCI-compliant vault. UNKNOWN intents are pinned to their original PSP forever."

When to deviate:

  • GMV where a 4-hour outage costs more than a team-year — roughly $20–50M/day for most businesses — build active-passive.
  • Global businesses — local acquiring in the EU, India or Brazil raises authorization rates materially; multi-PSP becomes a revenue lever, not just availability.

🧭 Principal Move: "Portability of stored cards is the one-way door here. If we let the PSP own the vault, switching later means a card-data migration between PCI Level 1 providers — typically months, coordinated with both vendors. I'd negotiate data-portability terms in the contract before signing, not after."


3.5 Fault Line 5: Fraud/Risk — Fail-Open vs Fail-Closed#

The tension: When the risk engine times out or is down, you either approve without a score (conversion preserved, fraud exposure) or decline/hold (fraud prevented, legitimate sales lost). Neither engineering nor product can make this call alone — it's a loss-budget decision.

ChoiceWhat WorksWhat BreaksWho Pays
Fail-open (approve without score)Zero conversion lossFraud rings notice within hours; losses + dispute ratioFinance (fraud losses), risk team (dispute ratio)
Fail-closed (decline)Zero added fraudEvery sale lost during the outageBusiness (revenue)
Degraded rules (fallback)Static rules (velocity, amount caps, BIN country) run locallyWeaker than the model; tuned rarelyRisk team (maintaining fallback rules)
Approve + hold captureCustomer sees success; capture delayed until re-scoreOnly works for auth/capture flowsOps (review queue)

Staff default: "Degraded rules plus hold-capture. Risk scoring has a 150ms timeout; on timeout, a local rule set runs — amount under $200, known device, card previously used successfully — approve; otherwise authorize but don't capture until re-scored. The loss budget for fail-open windows is set by risk and signed by finance, in dollars per hour."

When to deviate:

  • High-risk verticals (gift cards, electronics, crypto on-ramps): fail-closed. The fraud loss rate on unscored gift cards can exceed the margin on the entire category.
  • Subscription renewals of long-tenured customers: fail-open. Fraud on a 3-year-old subscription renewal is near zero.

🧭 Principal Move: "I'd publish a risk posture table per product line — fail-open, degraded or fail-closed — signed by the GM and the head of risk. Engineering implements the table; engineering does not decide it. When an outage hits at 2am, the on-call executes the table instead of guessing."


4. Failure Modes & Operational Reality#

4.1 PSP Brownout — The UNKNOWN Flood#

t=0:       PSP A p99 rises from 3s to 25s. Error rate 0.2% → 4%.
t=+30s:    Our 15s timeout starts firing. payment.unknown_count: 12 → 340.
t=+1min:   Checkout threads saturate waiting on PSP. Gateway p99 → 16s.
t=+2min:   Resolver kicks in at full speed → doubles call volume to the struggling PSP.
t=+3min:   PSP A rate-limits us (429). UNKNOWN count 2,100 and climbing.
t=+5min:   Page: payment.unknown_count > 500 for 3m. Also psp.error_rate > 2%.
t=+8min:   On-call: circuit-open PSP A for NEW intents → route to PSP B (portable tokens only).
t=+10min:  Resolver throttled to 20 req/s against PSP A. UNKNOWNs pinned.
t=+45min:  PSP A recovers. Resolver drains 2,100 UNKNOWNs in ~2 min: 1,870 AUTHORIZED, 230 DECLINED.
t=+1 day:  Reconciliation: 3 breaks (webhook arrived after resolver; dedup handled 2, 1 manual).

Detection: psp.auth_latency_p99, psp.error_rate, payment.unknown_count, payment.unknown_age_p99, checkout.thread_pool_utilization.

Mitigation: bulkhead PSP calls into a dedicated pool (e.g., 200 concurrent per PSP) so a slow PSP can't take checkout threads; adaptive resolver rate (token bucket, 20/s default); circuit breaker on the router for new intents.

Prevention: quarterly failover drills; a resolver that respects Retry-After; contractually agreed PSP status notifications.

Owner: payments on-call (primary), with PSP account manager escalation path in the runbook.

4.2 The Cross-PSP Double Charge#

t=0:       Intent pi_9 sent to PSP A. Timeout. State UNKNOWN.
t=+5s:     Router sees PSP A unhealthy. A "retry on failure" branch re-routes pi_9 to PSP B.
t=+6s:     PSP B authorizes pi_9. State AUTHORIZED (psp=B).
t=+40min:  PSP A recovers; webhook: pi_9 AUTHORIZED at PSP A.
t=+40min:  Webhook handler: intent is already AUTHORIZED → ignores it (dedup by intent).
t=+3 days: PSP A auth hold expires silently. Customer saw a pending hold for 3 days.
           Worse variant: auto-capture was on at PSP A → customer charged twice.

Detection: webhook.conflict_total (webhook for an intent that's terminal on a different PSP) — this must page, not log. Reconciliation break type missing_in_ledger for PSP A.

Mitigation: void/refund the orphan at PSP A immediately; notify the customer proactively.

Prevention: the pin rule is enforced in the router, not by convention — psp on the intent is write-once after the first attempt row exists. Code review checklist item and a property test.

Owner: payments team; incident review with the router's author.

4.3 Ledger Imbalance — The Silent Sev-1#

t=0:       Deploy adds FX fee handling. A code path posts the fee debit but not its credit
           when currency != USD (0.4% of volume).
t=+1h:     ledger.imbalance_total = -$212. No alert — the check runs nightly.
t=+14h:    Nightly invariant job: sum(debits) != sum(credits) by $3,090. Page.
t=+15h:    Revert. 1,140 unbalanced transactions identified by txn_id.
t=+2 days: Correcting entries posted, each referencing the incident. Auditor notified.

Detection: make imbalance impossible at write time: the ledger API accepts a transaction (list of entries) only if debits = credits per currency; reject otherwise. Then ledger.rejected_unbalanced_total pages on the first occurrence, within seconds.

Mitigation: correcting entries, never UPDATE/DELETE.

Prevention: the ledger service is the only writer; posting rules live in versioned config reviewed by finance; property tests on every posting rule.

Owner: ledger team (platform); finance-ops signs off on correcting entries.

4.4 Lost Events — The Dual-Write Gap#

t=0:       Service commits intent CAPTURED, then publishes to Kafka.
t=+1ms:    Pod OOM-killed between commit and publish.
t=+1 day:  Ledger has no capture entry. Fulfillment never shipped. Customer charged.
t=+3 days: Reconciliation flags missing_in_ledger. Support ticket already open.

Detection: outbox.lag_seconds and outbox.unpublished_count if using an outbox; if not, only reconciliation finds it — days late.

Prevention: transactional outbox (Appendix C). This is not optional in payments.

Owner: payments team.

4.5 Webhook Storm and Out-of-Order Delivery#

PSP webhooks are at-least-once and unordered. After a PSP incident, you can receive charge.refunded before charge.captured, or 50,000 retried webhooks in 10 minutes.

Mitigation: acknowledge fast (write raw event to a table keyed by PSP event ID, return 200 in < 100ms), process async; state transitions use conditional updates so an out-of-order event either applies (valid transition) or parks for retry; never trust webhook payloads for amounts without fetching the object from the PSP API when stakes are high.

Detection: webhook.ingest_lag_seconds, webhook.parked_count, webhook.signature_failures_total (spoofing attempts).

Owner: payments team.

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
PSP brownoutpsp.error_rate > 2%, payment.unknown_count > 500All checkouts on that PSPCircuit-break new intents, throttle resolver, bulkheadsPayments on-call
Cross-PSP duplicatewebhook.conflict_total > 0Individual customers, reputationalVoid orphan, pin rulePayments team
Ledger imbalanceledger.rejected_unbalanced_total > 0Books, auditReject at write, correcting entriesLedger platform + finance-ops
Lost eventoutbox.unpublished_count age > 60sFulfillment, ledgerOutbox relay restart; replayPayments team
Webhook stormwebhook.ingest_lag_seconds > 300Delayed state updatesFast-ack + async; scale consumersPayments team
Recon breaks agingrecon.breaks_older_than_2d > 0Month-end close, auditBreak queue triageFinance-ops
Risk engine downrisk.timeout_rate > 5%Fraud exposure or lost salesPosture table: degraded rulesRisk team (decision), payments (execution)
Idempotency store downidem.store_errorsAll mutating callsFail closed — reject with 503; never process without the key checkPayments on-call

🎯 Staff Insight: The idempotency store is the one component that must fail closed. Processing a charge without checking the key trades a few minutes of 503s for an unbounded number of duplicates. "I'd rather take a 5-minute checkout outage than a 5-minute double-charge window."


5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingLists features: charge, refund, webhookNames merchant vs marketplace vs processor; commits; says throughput isn't the hard partAsks whether payments is a platform or a product line, and which legal entity holds funds
Correctness"Use transactions and retries"UNKNOWN state, idempotency key scope, conditional state transitions, outboxMakes idempotency and UNKNOWN handling a company-wide standard with enforcement in the client library and egress
Data modelPayments table with statusState machine separated from double-entry ledger; integer minor unitsChart of accounts co-owned with finance; ledger as a system of record with change control
FailureReplicas and retriesPSP brownout plan: bulkheads, pinning, resolver throttle; idempotency store fails closedPrices multi-PSP vs outage loss; sets the risk posture table with GM sign-off; runs failover game days
Operations"Add monitoring"payment.unknown_count, ledger.rejected_unbalanced_total, recon break SLA with ownermoney_in_unknown_usd and month-end close time as exec-visible KPIs
OrganizationPayments team owns itPayments / finance-ops / risk split with explicit handoffsRedraws boundaries: ledger platform, money-movement platform, product teams as consumers; deprecates direct PSP integrations

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Treats timeout as a third outcome"A timeout moves the intent to UNKNOWN. We show 'Confirming', not 'Failed'."
Designs the key, not just the word"The key is derived from the attempt row: pi_123:capture:2. It exists before the call."
Separates state from money"The state machine says where the payment is. The ledger says where the money is."
Knows the store that must fail closed"If the idempotency store is down, we 503. That's the one place I fail closed."
Owns the residue"Reconciliation breaks go to finance-ops with a 2-day SLA; phantom charges page payments."
Refuses cross-PSP re-drive"An UNKNOWN is pinned to its PSP forever. Failover is for new intents."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Kafka exactly-once solves it"Kafka EOS covers Kafka-to-Kafka; the PSP is outside the transaction
"We'll use 2PC across payment service and PSP"PSPs don't participate in your 2PC; shows no experience with external systems
Floats for moneyRounding errors at scale are real breaks; signals no production money experience
Balance stored as a mutable column, updated in placeNo audit trail; concurrent updates lose money under race
"Retry until it succeeds"Unbounded retries against a charge endpoint are a double-charge generator
No mention of reconciliationAssumes the mechanisms are perfect; they leak at 0.01–0.1%

5.4 Common False Positives#

  • Deep card-network knowledge ≠ payment system design. Explaining interchange, BIN ranges and ISO 8583 fields impresses briefly; if it doesn't lead to the UNKNOWN state, it's trivia.
  • Saga vocabulary ≠ correctness. "We'll use a saga with compensations" is only Staff if the candidate names what the compensation is for a charge (void vs refund), what it costs, and what happens when the compensation itself times out.
  • Microservice count ≠ maturity. Twelve services (wallet, ledger, fraud, router, vault, notifier…) drawn in 5 minutes usually means none of their contracts were thought through.
  • PCI buzzwords ≠ compliance design. "We'll be PCI compliant" is not a design; "the PAN never touches our servers because the PSP SDK tokenizes in the browser" is.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minPick merchant/marketplace/processor; state correctness-first constraints
Entities & API3–5 minIntent, attempt, ledger entry, idempotency record; minor units
Architecture5–10 min≤ 8 boxes; pre-call / call / post-call; outbox
Unknown outcome10–18 minUNKNOWN state, same-key retry, resolver, UX
Ledger18–26 minDouble-entry, posting rules, corrections
Reconciliation + ops26–34 minThree-way match, break types, owners, metrics
Pivot (interviewer's choice)34–42 minMulti-PSP, marketplace payouts, multi-region, fraud posture
Wrap42–45 minExactly-once effect = idempotency + state machine + reconciliation; evolution

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"Now we're a marketplace — pay sellers"Can your ledger model third-party money?Seller accounts, holding account, payout schedule, negative balance policy
"Make it multi-region"Do you know money doesn't want active-active?Home-region per account; single writer per intent; async replication for reads
"Black Friday: 30× traffic"Is throughput the real risk?PSP rate limits and hot ledger accounts, not DB TPS; pre-negotiated PSP limits
"Add subscriptions"Scope disciplineBilling is a separate system that creates intents; buy it first
"A customer says they were charged twice"Operational forensicsQuery attempts by order, PSP references, recon status; one query, not a war room
"How do you test this?"Correctness cultureFault injection at the PSP boundary; PSP sandbox chaos; property tests on ledger posting rules

6.3 What to Deliberately Skip#

  • Card-network internals (issuer/acquirer message flows) — one sentence, then abstract behind the PSP.
  • UI details of checkout — except "Confirming" state for UNKNOWN.
  • Currency conversion math — say "FX is a ledger transaction with a rate reference"; skip rate sourcing.
  • Tax calculation — separate system; mention it exists.
  • Sharding the payments DB — unless the interviewer insists on 100K+ TPS; say why you're skipping it.

6.4 Follow-Up Questions to Expect#

  1. "Who generates the idempotency key, and what happens if two browser tabs submit the same order?"
  2. "The service crashed after the PSP responded but before you wrote the result. Walk me through recovery."
  3. "How do you handle a partial refund followed by a chargeback on the same payment?"
  4. "Finance says the PSP payout was $1,204 short yesterday. How do you find it?"
  5. "How long do you keep idempotency keys, and what happens when a retry arrives after expiry?"
  6. "The risk engine is down. What does checkout do, and who decided?"
  7. "How would you migrate from PSP A to PSP B with 4 million stored cards?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design a payment system for our e-commerce site."

Staff Answer

"Before drawing — are we the merchant collecting for our own goods, a marketplace collecting on behalf of sellers, or are we building the processor itself? Those are three different systems. I'll assume merchant checkout via PSPs, pay-in only, with the ledger designed so payouts can be added.

Constraints I'll commit to: no duplicate charges and no lost charges, so correctness beats latency; every cent explainable to finance; and the PSP is an unreliable dependency with multi-second p99. Volume is ~200 TPS peak, so I won't spend time on sharding. I'll walk through: intent and attempt model → the confirm flow with no transaction across the network call → what a timeout means → the ledger → reconciliation → who owns what."

Why this is L6:

  • Distinguishes three intents with incompatible designs and commits to one
  • Explicitly deprioritizes throughput with a number, freeing time for correctness
  • Previews the outline as a sequence of decisions, ending in ownership

What L7 adds:

  • Asks which legal entity holds funds and whether money-transmission licensing applies
  • Asks whether other business lines already integrate PSPs directly — the consolidation problem
  • Frames the ledger as a finance system of record with shared ownership from day one
❌ Common L5 Trap

"We'll have a payment service that calls Stripe's charge API and stores the result in a payments table, with Kafka for events and retries for reliability."

Why this misses: Every component is reasonable, and the design double-charges on the first lost response. There's no intent that exists before the call, no key, no UNKNOWN, no ledger. The interviewer's next question — "the call timed out" — has no good answer.


Drill 2: Idempotency Key Design#

Prompt: "Walk me through your idempotency keys. Who creates them, what's their scope, how long do they live?"

Staff Answer

"Two layers. Client-to-us: the client generates a UUID when the user lands on the confirm step and reuses it for every retry of that submit; two tabs get two keys, so I also enforce one active intent per order_id with a unique constraint — that's the real guard against double submission. Us-to-PSP: the key is derived from our persisted attempt, {intent_id}:{operation}:{attempt_seq}, created before the call. A retry of the same attempt reuses it; a deliberately new attempt — say, after a decline with a different card — increments the sequence.

Storage: Postgres table keyed by (scope, key) with the request hash and the stored response, inserted in the same transaction as the intent. Retention: 7 days, longer than any resolver retry horizon (24h) plus margin. A replay with a different request hash returns 422, because the client reused a key for a different intent — that's a client bug I want loud."

Why this is L6:

  • Distinguishes the client-facing key from the PSP-facing key and states each one's scope
  • Names the business-level uniqueness constraint (order_id) as the real double-submit guard
  • Sets retention from the retry horizon, not arbitrarily

What L7 adds:

  • Standardizes the key format and middleware across every money-moving service
  • Adds a lint/egress control so no service reaches a PSP without the middleware
  • Tracks idem.hash_mismatch_total per client team as a quality signal fed back to those teams

Drill 3: Crash After PSP Success#

Prompt: "The PSP returned 'authorized'. Your service crashed before writing it. What happens?"

Staff Answer

"Because I committed the attempt row as PENDING before the call, recovery is deterministic. The resolver scans for attempts PENDING older than 2× the PSP timeout — say 30 seconds — and for each one either re-issues the call with the same PSP key, which returns the stored authorization, or queries the PSP by our reference. Either way the post-call transaction runs once: attempt SUCCEEDED, intent AUTHORIZED, outbox row. Meanwhile the PSP's webhook may arrive first; it takes the same conditional transition, so whichever wins, the other is a no-op. The client, if it retries with its key, sees IN_PROGRESS and gets 202 PROCESSING until resolution."

Why this is L6:

  • The pre-call commit is the design element that makes recovery possible — named explicitly
  • Three independent resolution paths (resolver, webhook, client retry) converge via conditional transitions
  • Gives a concrete staleness threshold for "stuck"

What L7 adds:

  • Makes the resolver a shared framework with per-operation plugins (authorize, capture, refund, payout)
  • Defines a company SLO: 99.9% of UNKNOWNs resolved within 1 hour, reported monthly to finance

Drill 4: PSP Is Down#

Prompt: "Your primary PSP is returning errors for 40% of calls. What happens?"

Staff Answer

"Three different populations, three behaviors. New intents: the router's circuit opens at > 2% errors for 60 seconds, and new intents with portable tokens route to the secondary; non-portable cards get 'try another method' or a retry-later message. UNKNOWN intents: pinned to PSP A, resolver throttled to 20 req/s and honoring Retry-After, so we don't worsen the outage. Captures of existing authorizations: queued — holds last about 7 days, so a 3-hour backlog costs nothing but delayed fulfillment. The page is psp.error_rate; the runbook includes the PSP's status page and account-manager escalation."

Why this is L6:

  • Refuses the one-size failover; separates new, unknown and in-flight work
  • Uses the hold duration as a quantified reason to queue rather than panic
  • Protects the recovering PSP from the resolver

What L7 adds:

  • Decides whether the secondary PSP exists at all, from outage-cost math ($/hour of checkout loss vs team cost)
  • Negotiates PSP contracts with incident notification SLAs and credits
  • Runs a twice-yearly failover game day with finance observing reconciliation afterward

Drill 5: Hot Ledger Account#

Prompt: "Every payment credits one platform_revenue account. At 2,000 TPS your ledger writes contend on that row. Fix it."

Staff Answer

"The contention is from storing a running balance on the account row. First fix: don't update a balance row on every entry — entries are inserts, and balances are derived. Inserts to an append-only table don't contend. If we need fast balance reads, keep balance snapshots computed asynchronously every N seconds, and read snapshot + entries since snapshot. If some operations need a synchronous balance check — like 'don't pay out more than the seller has' — those are per-seller accounts, which aren't hot. For the truly hot house account, shard it into 16 sub-accounts, platform_revenue_00..15, chosen by hash of txn_id; the reported balance is the sum. Finance sees one account in reporting; the ledger sees 16."

Why this is L6:

  • Identifies the root cause (mutable balance) rather than scaling the DB
  • Distinguishes accounts that need synchronous balance checks from those that don't
  • Uses sub-account sharding with a reporting roll-up — the standard pattern for hot accounts

What L7 adds:

  • Makes "no synchronous balance on house accounts" a ledger platform rule
  • Aligns the sub-account convention with finance's reporting so month-end isn't affected

Drill 6: Marketplace Payouts#

Prompt: "We're becoming a marketplace. Sellers get paid. What changes?"

Staff Answer

"The ledger becomes the center. Pay-in credits a holding account, not revenue: buyer pays $100 → debit psp_receivable 100, credit seller_123_pending 85, credit platform_fee_revenue 15. After the risk hold — say 24 hours after delivery — a transfer moves seller_123_pending to seller_123_available. Payouts run daily: sum available, create a payout intent with its own idempotency and UNKNOWN handling, debit available, credit payouts_in_transit, and reconcile against the bank. The new failure mode is negative balances: a chargeback after payout pulls $100 back from the platform, and we recover it from the seller's future earnings. Policy — reserve percentage, hold duration — is set by risk; we implement it as ledger rules."

Why this is L6:

  • Models seller money as distinct accounts with states (pending, available, in transit)
  • Reuses the intent/UNKNOWN pattern for payouts
  • Names negative balances as the new risk and assigns the policy owner

What L7 adds:

  • Raises licensing and safeguarding: holding funds for third parties may require a license or a PSP's marketplace product (e.g., connected accounts) to hold them instead
  • Chooses buy (PSP marketplace product) vs build (own payouts) based on licensing cost and geographic expansion

Drill 7: Build vs Buy#

Prompt: "Why not just use Stripe for everything and skip your own ledger?"

Staff Answer

"At our stage, maybe we should. If we have one PSP, one currency, merchant-of-record and finance closes the books in 2–3 days using PSP exports, an internal ledger is ~2 engineers for 2 quarters with no user-visible benefit. I'd build the intent state machine — we need that regardless for order linkage and UNKNOWN handling — and defer the ledger. The triggers to build it: a second PSP, marketplace flows, month-end close > 5 days, or an audit finding. I'd design the event schema now so the ledger can be backfilled from payment.* events when we build it."

Why this is L6:

  • Declines to build by default and names concrete triggers
  • Separates what's always needed (intents) from what's deferrable (ledger)
  • Preserves the option cheaply via event schema

What L7 adds:

  • Evaluates buying a ledger product vs building, with a 3-year TCO
  • Recognizes that the PSP's data model becomes a lock-in vector if it's the only record

Drill 8: Changing Posting Rules Without an Outage#

Prompt: "Finance wants fees recognized differently starting next month. How do you change ledger behavior safely?"

Staff Answer

"Posting rules are versioned config, not code branches. New rule set v14 is authored with finance, reviewed, and deployed in shadow: for every event, the ledger computes v13 (posted) and v14 (shadow, written to a side table). A diff job compares totals daily; finance signs off on the diff. On the cutover date — month boundary, agreed with the controller — v14 becomes active with an effective timestamp. Entries record the rule version that produced them, so historical entries remain explainable. Rollback is flipping the active version; already-posted v14 entries get reversing entries."

Why this is L6:

  • Shadow → review → effective-dated cutover mirrors safe policy rollout
  • Rule version on each entry keeps history auditable
  • Rollback is defined, including already-posted entries

What L7 adds:

  • Establishes a change-control process that satisfies audit (SOX ITGC) without making every change a quarter-long project
  • Separates engineering deploys from accounting policy changes in the release calendar

Drill 9: Multi-Region#

Prompt: "We're expanding to the EU. Make payments multi-region."

Staff Answer

"I won't make a single intent writable in two regions — that's how you get two captures. Each customer (or merchant entity) gets a home region; intents are created and mutated only in the home region. The EU region uses EU PSP acquiring for better authorization rates and data residency. The ledger is per legal entity; cross-entity movements are explicit inter-company transactions. For region failure, I'd accept that EU checkout is down for the duration rather than fail over writes to the US with replication lag — a lagging replica might not know about an UNKNOWN from 2 seconds ago. Read-only order history can serve from the other region."

Why this is L6:

  • Single-writer per intent is the correctness invariant; states it before topology
  • Ties regions to legal entities and local acquiring
  • Chooses availability loss over duplicate risk, explicitly

What L7 adds:

  • Aligns regional architecture with entity structure, tax and data-residency counsel
  • Budgets the cost: second region ~1.6–2× infra cost of payments stack; justified by auth-rate uplift, not availability

8. Deep Dive Scenarios#

Deep Dive 1: Peak-Traffic Incident — Flash Sale Double Charges#

Context: During a flash sale at 40× normal traffic, support receives 300 "charged twice" tickets in an hour. The on-call escalates to you.

Questions to Surface First:

  • Are these two captures, or an auth hold plus a capture (which looks like two charges on a banking app)?
  • Do the duplicates share an order_id? An idempotency key? A PSP?
  • Did anything deploy in the last 24 hours — client, router, PSP SDK?
  • What's payment.unknown_count doing?

Typical L5 Approach: Finds the duplicates, refunds them, adds a check "if order already paid, reject". Fixes the symptom, doesn't find why the key didn't protect.

Staff Approach: First establishes whether they're real duplicates (most "double charge" reports during high load are pending auth holds plus the capture). Then traces a sample: two intents for one order means the order-level unique constraint is missing or bypassed; one intent with two PSP charges means the PSP key changed between retries.

Principal Approach: Treats it as a control failure: the idempotency guarantee should be structurally impossible to bypass. Commissions a cross-team review of every PSP call site, and moves idempotency into the single PSP client with egress restrictions, so the next team can't reintroduce the bug.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Query attempts grouped by order_id having > 1 captured. Classify: real duplicates vs auth-hold confusion. If real and ongoing, disable client auto-retry via flag.
TriageReal duplicates had two intents per order: the mobile client created a new intent on retry after a 15s timeout (new release lowered timeout from 30s).
Quick fixAdd unique index on (order_id) WHERE state NOT IN terminal-failed; revert client timeout; void/refund duplicates proactively with an email.
GuardrailsAlert payment.duplicate_capture_per_order > 0 — paging, not dashboard.
Post-mortemWhy could a client create a second intent for one order? Why did a client timeout change ship without payments review?

Metrics to Watch: payment.intents_per_order_p99, payment.duplicate_capture_per_order, psp.auth_latency_p99, client.checkout_timeout_rate

Organizational Follow-up: payments owns a client SDK; client apps may not set PSP-related timeouts themselves. Proactive refunds within 24 hours cut dispute rate.

Ownership Question: "Who owns the double-charge metric?" Staff answer: Payments owns it and pages on any non-zero value. Client teams consume the SDK and don't own payment correctness.

Key Takeaway: "Idempotency keys protect retries of one intent. Only a business-level uniqueness constraint protects against two intents."

What clears the Staff bar:

  • Distinguishes auth holds from real duplicates before acting
  • Finds the missing business-level constraint, not just the bad key
  • Moves correctness ownership to the payments SDK

Deep Dive 2: Silent Failure — Webhooks Stopped Three Days Ago#

Context: Finance notices captured payments from the last 3 days have no settlement matches for one PSP. No alerts fired. Orders are stuck in AUTHORIZED; fulfillment shipped nothing.

Questions to Surface First:

  • Did captures reach the PSP, or did our capture calls fail silently?
  • Did webhook ingestion stop, or did processing stop?
  • Why didn't payment.unknown_count or state-age metrics fire?

Typical L5 Approach: Finds the webhook endpoint was returning 401 after a secret rotation, fixes the secret, replays webhooks from the PSP dashboard.

Staff Approach: Fixes the secret, then fixes the detection gap: the system relied on webhooks as the only path to advance state. Adds the resolver as a second path for any intent in a non-terminal state older than its expected dwell time, plus an alert on state_age_p99 per state.

Principal Approach: Institutes "dwell-time SLOs" for every state in every money state machine as a platform standard, and moves secret rotation for PSP webhooks to an automated dual-secret process owned by the security platform.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Check webhook.signature_failures_total: spiking since the rotation. Restore old secret alongside new (dual-verify).
Triage41,000 events rejected. PSP retries webhooks for ~3 days, so some are lost for good; resolver re-queries all intents non-terminal > 1h.
Quick fixBackfill via PSP list API by time range; drive state transitions through normal conditional updates.
GuardrailsAlert: any state dwell p99 > 2× baseline; webhook 4xx rate > 1%.
Post-mortemWebhooks are hints, not the source of truth. Every transition needs a pull path.

Metrics to Watch: webhook.signature_failures_total, payment.state_age_p99{state}, recon.unmatched_psp_side

Organizational Follow-up: secret rotation runbook requires overlap windows; recon unmatched counts reviewed daily, not at month-end.

Ownership Question: "Who should have caught this on day 1?" Staff answer: Payments, via a state-dwell alert. Finance caught it on day 3 via reconciliation — which is the backstop working, but three days late.

Key Takeaway: "Push is an optimization. Pull is the guarantee."

What clears the Staff bar:

  • Treats webhooks as at-least-once hints with a pull fallback
  • Alerts on state dwell time, not only on errors
  • Reconciliation framed as the last line, not the first

Deep Dive 3: Large-Customer Onboarding — Enterprise Invoices at $2M Each#

Context: Sales signs an enterprise buyer paying by bank transfer, $2M per invoice, 20 per month. Current system is card-only with 30-minute UNKNOWN resolution.

Questions to Surface First:

  • Push (customer sends a wire) or pull (ACH debit)? Pull debits can be returned for days.
  • How do we match an incoming wire to an invoice when the reference is free text?
  • What's the approval flow for refunding $2M?

Typical L5 Approach: Adds a "bank transfer" payment method type to the existing flow.

Staff Approach: Recognizes a different state machine: AWAITING_FUNDS → FUNDS_RECEIVED → MATCHED → APPLIED, with a virtual account number per customer so matching is deterministic. ACH returns can arrive up to ~2 banking days (longer for some return codes), so "collected" is not final at receipt. Large refunds require dual approval (four-eyes) enforced in the refund API.

Principal Approach: Separates "payments" (card rails) from "treasury operations" (bank rails, cash positioning) as distinct capabilities with distinct owners, and sets thresholds above which money movement requires finance approval regardless of rail.

Staff Approach — Full Reasoning
PhaseWhat to Do
DesignVirtual account numbers per customer from the bank partner; inbound wires auto-match to the customer.
Ledgerbank_cash debit, customer_123_unapplied credit; applying to an invoice moves unapplied → AR settlement.
ControlsRefunds > $50K require a second approver; approvals are ledger-linked events.
ReconDaily bank statement (BAI2/camt.053) import; unmatched deposits > 1 day → treasury queue.
Post-launchTrack unapplied_cash_usd — cash we hold but can't attribute is an audit finding waiting to happen.

Metrics to Watch: bank.unmatched_deposits_count, ledger.unapplied_cash_usd, refund.pending_approval_age

Ownership Question: "Who owns unapplied cash?" Staff answer: Treasury/AR in finance, with payments owning the matching automation and its accuracy metric.

Key Takeaway: "A new rail is a new state machine. Don't bolt it onto card semantics."

What clears the Staff bar:

  • Recognizes different finality semantics per rail
  • Designs deterministic matching instead of fuzzy text matching
  • Adds approval controls proportional to amount

Deep Dive 4: Post-Mortem — $180K Reconciliation Hole at Quarter-End#

Context: Quarter-end close finds $180K of captures in the ledger with no PSP settlement. Auditors are asking. You own the post-mortem.

Questions to Surface First:

  • Were these captures ever sent? Accepted by the PSP? Later reversed?
  • Is the gap concentrated by date, currency, merchant category or code path?
  • How long have the breaks existed, and why did nobody act?

Typical L5 Approach: Traces the specific transactions, finds a bug in partial-capture handling (ledger posted full amount, PSP captured partial), fixes the code, posts corrections.

Staff Approach: Fixes the bug, and then asks why breaks aged 70 days without action: the break queue had no owner after a reorg, and the dashboard showed counts, not dollars. Assigns an owner, adds dollar-weighted aging alerts, and adds a posting-rule property test that ledger amount = PSP-confirmed amount.

Principal Approach: Treats it as a governance gap: reconciliation had no named owner at the org-chart level. Writes reconciliation into the finance operating model — owner, SLA, escalation to controller — and makes "unreconciled $ > 30 days" a quarterly audit committee metric.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateFreeze and snapshot break data; quantify by cause bucket.
Triage94% from partial captures on split shipments since a refactor 11 weeks ago; 6% FX rounding.
Quick fixCorrecting entries per transaction referencing the incident ID; FX rounding rule aligned with PSP.
Guardrailsrecon.break_usd_aged_7d > $10K pages finance-ops lead; > $50K pages the payments manager.
Post-mortemBreak queue ownership lost in reorg; count-based dashboards hid dollar magnitude.

Metrics to Watch: recon.break_usd_by_age_bucket, recon.break_rate_pct, ledger.corrections_posted_total

Ownership Question: "Who owns a break older than 30 days?" Staff answer: The controller's office, by escalation policy. Engineering owns fixing causes; finance owns the books.

Key Takeaway: "Reconciliation without an owner is a report nobody reads."

What clears the Staff bar:

  • Weights breaks by dollars and age, not count
  • Finds the org gap behind the code bug
  • Uses correcting entries, never edits

Deep Dive 5: Multi-Region Expansion — India Launch#

Context: The company launches in India. Local regulation requires payment data to be stored in-country and card-on-file flows use tokenization mandated by the regulator; recurring payments need additional mandate registration.

Questions to Surface First:

  • What data must stay in-country, and can derived/aggregated data leave?
  • Which PSPs have local acquiring and support the mandated token flows?
  • What's our legal entity structure in-country?

Typical L5 Approach: Deploys the existing stack in an India region.

Staff Approach: Deploys a regional payments cell with its own DB and ledger for the Indian entity, a local PSP, and a mandate-aware subscription flow with its own states (mandate pending, mandate active, pre-debit notification sent). Global reporting consumes aggregated, non-restricted ledger summaries.

Principal Approach: Establishes the "payments cell per regulatory jurisdiction" pattern as the paved road — deployable template, entity-scoped ledger, local PSP adapter interface — so the next country is a 2-month launch, not a 9-month project, and puts regulatory review into the launch checklist.

Staff Approach — Full Reasoning
PhaseWhat to Do
ScopingLegal + compliance define restricted data; engineering maps fields to restricted/unrestricted.
ArchitectureRegional cell: intents, attempts, ledger, idempotency in-country.
IntegrationPSP adapter interface; local PSP as plugin; recon adapter for local report formats.
ReportingDaily entity-level ledger summaries exported to the global consolidation ledger.
LaunchShadow traffic with ₹1 test charges; staged rollout 1% → 10% → 100% over 3 weeks.

Metrics to Watch: payment.auth_rate{region}, mandate.registration_failure_rate, recon.break_rate{entity}

Ownership Question: "Who owns the India ledger?" Staff answer: The regional finance controller for books; the payments platform for the system.

Key Takeaway: "Regions follow legal entities, not latency."

What clears the Staff bar:

  • Maps regulation to data classification before architecture
  • Entity-scoped ledgers with consolidated reporting
  • Adapter interfaces so PSPs are replaceable per market

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why exactly-once money movement is unavailable and build exactly-once effect from idempotency, a state machine and reconciliation
  • Design idempotency keys at two layers (client → us, us → PSP) with scope, derivation and retention justified
  • Draw the payment state machine including UNKNOWN, and state which transitions are conditional updates
  • Post a capture, fee, refund and chargeback as balanced double-entry transactions
  • Describe a three-way reconciliation with break types, owners, SLAs and dollar-weighted aging
  • Handle a PSP brownout with separate behavior for new, unknown and in-flight work
  • Name the one store that must fail closed (idempotency) and the decision that engineering must not own alone (fraud fail-open)
  • Price multi-PSP routing against outage cost, and say when you would not build it

The Bar for This Question#

Mid-level (L4): Integrates a PSP correctly on the happy path: tokenization via SDK, a charge call, a payments table, webhook handling. Mentions retries. Doesn't distinguish timeout from failure. Would ship something that works 99.9% of the time and double-charges in the remaining 0.1%.

Senior (L5): Adds idempotency keys, two-step auth/capture, an event bus and reasonable failure handling. Knows reconciliation exists. The gap: treats the idempotency key as a server-side dedup detail rather than a contract; keeps money in a mutable status table; proposes PSP failover without pinning UNKNOWNs. The design is plausible and would pass a code review — and would still produce a quarter-end reconciliation hole.

Staff+ (L6): Frames the problem around unknown outcomes within the first five minutes. Separates state from ledger. Uses pre-call/call/post-call phases with no transaction across the network. Designs reconciliation as an owned product with break types, SLAs and dollar-weighted aging. Names who pays for each tradeoff — customer uncertainty over duplicate charges, finance-ops for breaks, risk and finance for fraud posture. Knows when not to build the ledger. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "Exactly-Once Payments" Is a Marketing Term#

ClaimReality
"Our queue gives exactly-once"Only within the queue's transactional boundary; the PSP is outside it
"The PSP handles idempotency"Only if you send the same key; the PSP can't dedup two keys for one logical charge
"We have no duplicates"You have no detected duplicates unless reconciliation checks for them

The Staff position: Say "exactly-once effect" and list its three components. Anything else is a promise you can't audit.

Why this matters in interviews: Interviewers use "exactly-once" as a trap. Accepting the phrase is a Senior signal; decomposing it is a Staff-level signal.

10.2 Most Companies Should Not Build a Ledger Yet#

StageRight Answer
< $1M/month, one PSPPSP dashboard + exports; intent table for order linkage
$1–50M/month, 1–2 PSPsIntent state machine + events designed for a future ledger
Marketplace, multi-PSP, or > 5-day closeInternal double-entry ledger with finance co-ownership

The Staff position: The ledger is correct and expensive. Build the event schema that makes it cheap to add, and add it when a named trigger fires.

Why this matters in interviews: Candidates who always build the ledger look dogmatic. Candidates who never build it look inexperienced. Name the trigger.

10.3 The Idempotency Store Should Fail Closed — Even Though Everything Else Fails Open#

ComponentFailure PostureWhy
Risk engineDegraded rulesLoss is bounded and priced
Webhook ingestQueue and retryPull path exists
Idempotency storeFail closedDuplicate charges are unbounded and customer-visible

The Staff position: A 5-minute checkout outage costs revenue you can mostly recover (users retry). A 5-minute duplicate window costs chargebacks, refunds and trust you can't.

Why this matters in interviews: Most availability-trained engineers default to fail-open. Knowing where not to is the signal.

10.4 Multi-PSP Is Usually a Cost Project Disguised as a Reliability Project#

DriverTypical Value
Outage protectionMajor PSP outages are rare (a few hours/year)
Authorization-rate uplift via local acquiringOften 1–3 points in international markets
Fee negotiation leverage5–20 bps at scale

The Staff position: Justify the second PSP with auth-rate and fees, not uptime. Availability is the bonus.

Why this matters in interviews: It shows you think about payments as a revenue system, not just an integration.

10.5 Reconciliation Is the Most Important Service Nobody Wants to Own#

The Staff position: Every other mechanism in this case study leaks. Reconciliation is where leaks become visible. If it lacks a named owner with an SLA, the company is flying blind on its own cash.

Why this matters in interviews: Ending with reconciliation ownership is the single most Staff-sounding close available for this question.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

The Staff engineer designs a correct payment system. The Principal engineer notices that the company has four payment systems — checkout, subscriptions, the marketplace team's payouts, and a gift-card service someone wrote in 2021 — each with its own idempotency conventions, its own PSP contract and its own definition of "paid". The L7 problem is not correctness of one flow; it is money-movement governance across the org: one ledger of record, one set of PSP contracts, one resolver framework, and a finance-engineering operating model that survives reorgs.

The Org-Level Fault Line#

Central money-movement platform vs per-product PSP integrations.

OptionWhat WorksWhat BreaksWho Pays
Each product integrates PSPsProduct speed; no platform dependencyN idempotency implementations, N recon pipelines, weak PSP pricing, no consolidated booksFinance (close time), customers (inconsistent bugs), CFO (fees)
Central platform owns all money movementOne ledger, one negotiation, one control surfacePlatform becomes a bottleneck; product roadmap waits on platformProduct teams (velocity)
Platform owns primitives; products own flowsIntents, ledger, PSP adapters, resolver are platform; checkout/subscription logic is productContract design is hard; needs strong API governancePlatform team (API stewardship)

🧭 Principal Move: "Platform owns the primitives that must be correct once — intents, idempotency, PSP adapters, ledger posting, reconciliation. Products own the flows that must be fast to change — checkout UX, subscription logic, promotions. Direct PSP egress from product services gets blocked after a 2-quarter migration window."

Cost Model#

Assumptions: US-centric, card-heavy, fully loaded engineer cost ~$250K/year, PSP fees at list minus negotiated discounts, cloud infra only for the payments stack.

ScaleGMV / yearInfra ($/month)HeadcountOn-call LoadPSP Fees (approx.)
Startup$20M~$2–5K (managed Postgres, small Kafka or none)1–2 eng (part-time), PSP-hosted ledgerShared rotation, < 1 page/month~2.9% + 30¢ → ~$650K/yr
Growth$500M~$20–40K (HA Postgres, Kafka, recon jobs)6–10 eng: payments (4), ledger (2–3), recon tooling (1–2) + 2 finance-opsDedicated payments rotation, 2–4 pages/monthNegotiated ~2.2–2.5% blended → ~$11–12M/yr
Enterprise$10B~$150–300K (multi-region cells, ledger store, data warehouse)40–60 eng across payments, ledger, risk, routing + 15–25 finance-opsFollow-the-sun, per-cell rotationsInterchange-plus, multi-PSP; each 10 bps = $10M/yr

The pricing insight: at enterprise scale, a routing team of 5 engineers (~$1.25M/yr) that improves blended cost by 5 bps saves ~$5M/yr. At growth scale, the same team saves ~$250K — not worth it yet.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Who stores card credentials (PSP vault vs own vault vs network tokens)One-wayCard migration between PCI Level 1 providers: months, vendor cooperation required
Chart of accounts / ledger schemaOne-wayRe-posting history; audit re-certification; 6–12 months
Money representation (integer minor units + currency)One-wayEvery consumer changes; historic data conversion
Legal entity holding funds (merchant vs marketplace vs licensed)One-wayLicensing, contracts, banking relationships
Choice of primary PSP (with portable tokens)Two-way1–2 quarters of integration + recon adapter
Resolver backoff schedule, timeoutsTwo-wayConfig change
Risk fail-open/closed posture per productTwo-wayPolicy change with sign-off; minutes to deploy
Outbox vs CDC for event publishingTwo-wayInternal; consumers see the same events

The Standard I'd Write#

RFC-PAY-001: Money Movement Standard
Status: Approved   Owners: Payments Platform + Office of the Controller

Scope
  Every service that initiates, reverses or records movement of funds
  (card, bank, wallet, payout, stored value).

MUST
  1. Reach PSPs and banks only through the Money Movement API or approved SDK.
  2. Create a durable intent before any external call; carry an idempotency key
     derived from the intent and operation.
  3. Treat timeouts and 5xx as UNKNOWN; never mark FAILED without a definitive response.
  4. Post all financial effects to the company ledger as balanced transactions
     (debits = credits per currency) in integer minor units.
  5. Never store PANs; use approved tokenization.
  6. Define dwell-time SLOs for every non-terminal state.

SHOULD
  1. Use two-step authorize/capture for goods with fulfillment delay.
  2. Emit events via transactional outbox.
  3. Participate in daily reconciliation with named break owners.

Exceptions
  Filed with Payments Platform; reviewed within 10 business days; time-boxed to
  2 quarters; controller sign-off required for any exception to MUST 4.

Success metrics
  - Duplicate captures per million payments: < 1
  - UNKNOWNs resolved within 1h: ≥ 99.9%
  - Reconciliation break rate: < 0.01%; breaks > 30 days: $0
  - Month-end close: ≤ 3 business days
  - Services with direct PSP egress: 0 by end of year 2

What I'd Tell the VP#

"Our payment systems are correct individually but inconsistent collectively — four teams move money four different ways, and finance spends eight days a month reconciling them. I'm proposing a money-movement platform that owns the parts that must be right once: idempotency, the ledger and reconciliation. It costs about six engineers for a year and cuts close time to three days, removes the double-charge class of incidents, and gives us one negotiating position with PSPs worth an estimated 10–20 basis points. Product teams keep ownership of checkout and pricing. The main risk is migration speed; I'd time-box direct PSP integrations to two quarters."

Principal Interview Signals#

SignalWhat It Sounds Like
Prices tradeoffs in dollars"Every 10 bps of blended fee is $10M a year at our volume; routing pays for its team 8×."
Identifies one-way doors"Vault ownership is the decision I'd slow down on. Everything else we can change."
Redraws ownership"Platform owns primitives, products own flows, the controller owns the books."
Sets org-wide failure posture"Risk posture is a signed table per product line, not an on-call judgment."
Knows when not to standardize"Gift cards have different finality; they get the ledger standard but their own state machine."

Staff answers that L7 interviewers find insufficient:

  • "We'll build a great payment service for checkout" — correct, but ignores the three other teams moving money differently.
  • "Reconciliation runs nightly and finance handles breaks" — names the owner but not the operating model, escalation path or audit metric.
  • "We'll add a second PSP for resilience" — no cost model, no mention of token portability as the one-way door.

Appendices

Appendix A: Mechanics in Depth#

A.1 The Three-Phase Request (Pre-call / Call / Post-call)#

def confirm(intent_id, client_key, request):
    # Phase 1: local transaction — record intent to act
    with db.tx() as tx:
        rec = tx.idem_get_for_update(scope="confirm", key=client_key)
        if rec and rec.request_hash != hash(request): raise Http422
        if rec and rec.status == "DONE":         return rec.response
        if rec and rec.status == "IN_PROGRESS":  return Http202("PROCESSING")
        tx.idem_insert(scope="confirm", key=client_key, hash=hash(request), status="IN_PROGRESS")
        intent = tx.intent_get_for_update(intent_id)
        attempt = tx.attempt_insert(intent_id, op="authorize",
                                    psp_key=f"{intent_id}:authorize:{intent.attempt_seq+1}",
                                    outcome="PENDING")
    # Phase 2: network call — no DB transaction open
    try:
        result = psp.authorize(intent, idempotency_key=attempt.psp_key, timeout=15s)
    except (Timeout, Psp5xx):
        result = UNKNOWN
    # Phase 3: local transaction — record outcome
    with db.tx() as tx:
        tx.attempt_update(attempt.id, outcome=result.outcome, psp_ref=result.ref)
        new_state = {"SUCCEEDED": "AUTHORIZED", "DECLINED": "DECLINED", "UNKNOWN": "UNKNOWN"}[result.outcome]
        tx.intent_transition(intent_id, from_state="AUTHORIZING", to_state=new_state)  # conditional
        tx.outbox_insert(event=f"payment.{new_state.lower()}", intent_id=intent_id)
        tx.idem_complete(scope="confirm", key=client_key, response=render(new_state))
    return render(new_state)

A.2 The Resolver#

every 30s:
  for attempt in attempts where outcome in ("PENDING","UNKNOWN")
                           and updated_at < now() - backoff(attempt.tries)
                           limit 500:
      if not psp_budget[attempt.psp].take(): continue      # 20 req/s per PSP
      r = psp.lookup(attempt.psp_ref or attempt.psp_key) or psp.retry(attempt, same key)
      if r.definitive: phase3(attempt, r)
      elif attempt.age > 24h: human_queue.push(attempt)
backoff: 1m, 5m, 30m, 2h, 6h

A.3 Why Not Sagas With Automatic Compensation?#

A saga compensates a failed step by undoing earlier ones. For payments, "undo" is void (cheap, pre-capture) or refund (expensive, customer-visible, 5–10 days). Automatic compensation on UNKNOWN is wrong — you'd refund a charge that may not exist, or void an auth that's about to be captured. Use sagas for the order workflow (reserve inventory → authorize → capture on ship → release inventory on decline) with payment steps that only compensate on definitive outcomes. See Managing Long-Running Processes.

Appendix B: Data Model#

CREATE TABLE payment_intents (
  intent_id      TEXT PRIMARY KEY,
  order_id       TEXT NOT NULL,
  amount_minor   BIGINT NOT NULL CHECK (amount_minor > 0),
  currency       CHAR(3) NOT NULL,
  state          TEXT NOT NULL,
  psp            TEXT,                         -- write-once after first attempt
  attempt_seq    INT NOT NULL DEFAULT 0,
  version        BIGINT NOT NULL DEFAULT 0,
  created_at     TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE UNIQUE INDEX one_live_intent_per_order ON payment_intents(order_id)
  WHERE state NOT IN ('DECLINED','VOIDED','EXPIRED');

CREATE TABLE ledger_entries (
  entry_id       BIGSERIAL PRIMARY KEY,
  txn_id         UUID NOT NULL,
  account_id     TEXT NOT NULL,
  direction      CHAR(1) NOT NULL CHECK (direction IN ('D','C')),
  amount_minor   BIGINT NOT NULL CHECK (amount_minor > 0),
  currency       CHAR(3) NOT NULL,
  caused_by      TEXT NOT NULL,                -- event id; unique per (caused_by, account, direction)
  rule_version   INT NOT NULL,
  posted_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- No UPDATE/DELETE grants on ledger_entries for any application role.

The ledger API accepts a transaction only when Σ debits = Σ credits per currency; the check runs inside the insert transaction. UNIQUE(caused_by, account_id, direction) makes posting idempotent: a replayed event cannot double-post.

Appendix C: Coordination Mechanisms#

C.1 Transactional Outbox#

Diagram: C.1 Transactional Outbox

C.2 Quick Comparison#

MechanismGuaranteesFailure ModeUse For
Idempotency key (client→us)Same response for same requestTwo keys for one intentAPI retries
PSP idempotency keyOne PSP operation per keyKey regenerated on retryExternal calls
Business unique constraintOne live intent per orderConstraint missingDouble-submit
Conditional state transitionEach transition onceBlind UPDATEWebhooks, resolver races
OutboxEvent iff state change committedRelay lagLedger, fulfillment
Ledger caused_by uniquenessEach event posts onceMissing constraintConsumer replays
ReconciliationDetects everything above leakingUnowned breaksBackstop

Appendix D: API Contract & Client Behavior#

  • Idempotency-Key required on every POST; 400 if absent.
  • Replay with same key + same body → original status and body; replay with different body → 422 idempotency_key_reused.
  • 202 PROCESSING for in-progress or UNKNOWN; include Retry-After: 2 and a status URL.
  • Clients poll with exponential backoff capped at 30s, or subscribe to push; clients never create a new intent for the same order after a timeout.
  • Webhooks to merchants/consumers are signed (HMAC with timestamp, 5-minute tolerance), carry a stable event_id, and are retried with backoff for up to 3 days.

Appendix E: Observability#

Core metrics:

  • payment.unknown_count, payment.unknown_age_p99, money_in_unknown_usd
  • payment.state_age_p99{state} — dwell time per state
  • psp.auth_latency_p99{psp}, psp.error_rate{psp}, payment.auth_rate{psp,bin_country}
  • ledger.rejected_unbalanced_total (must stay 0), ledger.post_lag_seconds
  • outbox.unpublished_age_max
  • recon.break_rate_pct, recon.break_usd_by_age_bucket
  • payment.duplicate_capture_per_order (must stay 0)

Critical alerts:

AlertThresholdSeverity
Ledger unbalanced rejection> 0Sev-1 page
Duplicate capture per order> 0Sev-1 page
UNKNOWN count> 500 for 3 minPage
PSP error rate> 2% for 1 minPage + auto circuit
Outbox oldest unpublished> 60sPage
Recon breaks aged > 7d> $10KPage finance-ops lead
Auth rate drop> 3 points vs 7-day baselineWarn → page after 15 min

Debugging the silent failure: a drop in authorization rate with no error-rate change is the most common silent failure (issuer-side declines after a PSP configuration change, a 3DS flow break, a BIN-country rule). Alert on auth rate by BIN country and PSP, not only on errors.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 50 TPSOne Postgres, intents + ledger in same DBNothing — resist complexity
50–1,000 TPSLedger as separate service; hot-account sub-sharding; outbox relayHouse-account contention; recon job runtime
1K–10K TPSLedger store partitioned by account; per-entity cells; streaming reconCross-cell reporting; PSP rate limits
> 10K TPSPurpose-built ledger storage with time-range sealing; multi-PSP routingOrg coordination, not tech

What you don't build on day one: multi-PSP routing, your own vault, multi-region active-active anything, a custom ledger database, real-time reconciliation. Each has a trigger in Section 11.

Appendix G: Multi-Tenancy, Fairness & Cost#

  • PSP rate limits are shared. A batch subscription renewal of 2M charges at 00:00 UTC can consume the PSP's per-account rate limit and starve interactive checkout. Separate PSP accounts or priority lanes: interactive traffic gets ≥ 70% of the PSP budget; batch renewals are spread over 6 hours.
  • Per-merchant isolation (marketplaces/platforms): one merchant's chargeback storm must not raise the platform's overall dispute ratio unnoticed — track dispute ratio per merchant and per MCC, with automated reserves at 0.75%.
  • Cost attribution: tag every ledger fee entry with product line so PSP fees are charged back; product teams that pick expensive payment methods see the cost.
  • Retries cost money: some PSPs and networks charge per authorization attempt, and networks penalize excessive retries on declined cards. Cap retries on hard declines at zero and soft declines at a small, network-compliant number.
  1. Loading the index…