Hiring BarSupport

Design a Ledger and Digital Wallet

Case study86 min read9 diagrams

Technologies referenced in this case study: PostgreSQL · Distributed SQL · Apache Kafka · Redis · OLAP Databases

Related: Payments · Idempotency & Exactly-Once · Dealing with Contention · Hot Keys · Workflows, Sagas & Compensation · Consistency, CAP & PACELC · Schema Design

Go deeper:

  • Treating postings as immutable events with balances as rebuildable projections is covered in depth in CQRS & Event Sourcing.

Reading Guide#

Organized for interview use first, reference second. Read front-to-back once, then return to the fault lines and incidents that match your weak spots.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (When It Breaks) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal View) and the appendices on posting rules, holds and hot accounts
What is a Ledger and Digital Wallet? — Why interviewers pick this topic

A digital wallet holds a balance on behalf of a user — top up from a card or bank, spend at merchants, send money to friends, cash out. Underneath every wallet is a ledger: the system of record that says how much money each account holds and, more importantly, why. A ledger records every movement as balanced entries between accounts; balances are a consequence of entries, not a column someone updates.

The hard part is not storing a number. The hard part is that the number is someone else's money. Two concurrent spends must not both succeed against a $50 balance. A card authorization must reserve funds that are released if the merchant never captures. Every cent the wallet says users hold must match the cash sitting in the company's bank account — to the cent, every day — because regulators call a shortfall by its legal name.

Before vs After — the "two taps, one balance" scenario:

Without a ledger (mutable balance column):
t=0:        Wallet balance = $50.00. User taps "Pay $40" on phone and watch at once.
t=+5ms:     Request A reads balance 50.00, checks 50 >= 40, OK.
t=+6ms:     Request B reads balance 50.00, checks 50 >= 40, OK.
t=+9ms:     A writes balance = 10.00. B writes balance = 10.00.
t=+10ms:    User spent $80. Balance says $10. Lost update; $40 created from nothing.
t=+30 days: Bank reconciliation: customer liabilities exceed cash by $40 × 3,100 incidents.
            Safeguarding shortfall reported to the regulator. Nobody can say which rows are wrong.

With a ledger (entries + conditional balance check + holds):
t=0:        Account wallet:u42 has posted balance $50.00, pending $0.
t=+5ms:     Request A posts txn T1: debit wallet:u42 $40, credit merchant_clearing $40,
            guarded by "available >= 40" in the same transaction. Commits. Available = $10.
t=+6ms:     Request B posts txn T2 with the same guard. Available 10 < 40. Rejected: insufficient funds.
t=+1 day:   Daily reconciliation: sum(customer wallet accounts) = FBO bank balance. 0 breaks.

Why interviewers reach for this question: It is the cleanest test of whether a candidate can hold an invariant under concurrency and failure — money is neither created nor destroyed, and no account goes below its floor — while also designing for an auditor who will ask, two years from now, why one account had $4,212.50 at 14:03 on a Tuesday. Candidates who draw a balances table with an UPDATE have never been on the receiving end of that question.

Mechanics Refresher: Ledger Primitives
PrimitiveHow It WorksProsCons
Double-entry transactionA transaction is ≥2 entries; Σ debits = Σ credits per currencyMoney cannot appear or vanish without an entry; the invariant is checkableRequires a chart of accounts; every flow needs posting rules
Append-only entriesEntries are never updated or deleted; mistakes are fixed by reversing entriesFull audit trail; any past balance is reconstructibleStorage grows forever; corrections are visible (that's the point)
Derived balanceBalance = Σ entries on the accountCan never drift from entriesO(entries) per read unless cached or snapshotted
Materialized running balanceBalance row updated in the same transaction as the entriesO(1) reads; enables overdraft guardHot row under contention; must be provably consistent with entries
Hold (two-phase transfer)Reserve funds as pending; later post (capture) or void; expires on timeoutModels card auths and in-flight withdrawals honestlyAnother balance dimension; expiry needs an owner
Idempotency key per transactionProducer supplies external_id; duplicates return the original resultRetries and event replays can't double-postKey scope and retention must be designed
ReconciliationCompare ledger to external truth (bank, card network, processor)Catches what the ledger got wrong or never sawBatch; needs owned break queue
Period closeLock a period; later corrections post in the current period with an effective dateReports stop changing after sign-offNeeds effective-date vs posted-date modeling

For most production systems: An append-only double-entry ledger in a relational store with serializable-enough guarantees, a materialized per-account balance updated in the same transaction as the entries, holds as a separate pending dimension, an external_id uniqueness constraint for idempotency, and daily reconciliation of customer liabilities against the bank. The primitives are not the interview — holding the invariant while the money is hot, concurrent and audited is.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

A ledger is not a database schema question. Anyone can draw accounts and transactions tables.

It is an invariant-ownership question that tests:

  • Whether balances are a consequence of immutable entries or a mutable number you hope stays right
  • Whether you can prevent overdrafts under concurrent spends without serializing the entire company through one row
  • Whether you model reserved money (holds) as a first-class state rather than a pre-emptive debit
  • Whether every cent in user wallets is matched, daily, by cash in a real bank account — and who gets paged when it isn't

The key insight: A ledger has two invariants, and they fail differently. Conservation (every transaction sums to zero) is enforced at write time and is cheap. Non-negativity (no wallet below its floor) requires reading current state, so it is where concurrency, hot accounts and latency live. Staff candidates separate the two and spend their time on the second; Senior candidates enforce the first and are surprised by the second.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws users, wallets, transactions; balance column on walletAsks "Is this stored value we're liable for, or a view over money held elsewhere? Who is the licensed entity, and where is the cash?"Asks "Is this ledger the company's books of record or one product's wallet? Who in finance co-owns the chart of accounts?"
BalanceUPDATE wallets SET balance = balance - 40"Balance is derived from entries; I materialize it per account in the same transaction, with a guard available >= amount, and prove it equals Σ entries nightly"Makes the derived-equals-materialized check a controlled audit test with evidence retained for the auditor
ConcurrencyRow lock on the wallet"User wallets are low-contention — row lock is fine. The fee and omnibus accounts are hot; those get sub-accounts or deferred aggregation"Sets the platform rule: no account may be on the hot path of more than ~500 posts/s; account design review for any new flow
Holds"Debit now, refund if cancelled""Holds are pending entries: available = posted − pending debits. Capture posts, void releases, expiry is a scheduled job with an owner"Prices hold policy: expiry length vs float exposure vs disputes, signed by risk and finance
CorrectnessUnit tests"Idempotent posting by external_id, reject unbalanced transactions at write, daily three-way reconciliation against the FBO bank account"Defines safeguarding as a regulated control: daily reconciliation evidence, shortfall escalation to compliance in hours, not days
Scale"Shard by user ID""~100–2,000 TPS is one well-tuned primary; I'd shard by account when write volume, not storage, demands it — and cross-shard transfers become two-phase"Plans the 3-year path: single ledger → per-entity ledgers → regional ledgers for data residency, with a consolidated reporting layer
Why "balance" separates levels

L5: Stores balance on the wallet and updates it in place. Adds a transactions table for history. This works on day one and becomes unanswerable on day 400: the history and the balance are two records written by different code paths, and when they disagree nobody knows which is right. A bug that skips the history insert produces a balance with no explanation.

L6: "The entries are the truth. The balance row is a cache I update in the same database transaction as the entries, guarded by a check constraint so it can't go below the account's floor. Every night a job recomputes Σ entries per account and compares it to the balance row; any difference is a Sev-1, because it means some code path wrote one without the other."

L7: "The derived-equals-materialized check isn't a nice-to-have; it's the control an auditor will test. I'd retain its daily output as evidence, and I'd make the ledger API the only writer — no service gets direct table access, so the class of bug can't recur in another team's code."

Why "holds" separates levels

L5: "When a card authorizes, debit the wallet. If the merchant never captures, credit it back." Reasonable — until the user sees a debit that later reverses, two statement lines for one purchase, a refund that looks like income, and an auth that expires silently because nobody wrote the "credit it back" job.

L6: "A card authorization creates a hold — a pending debit. Available balance is posted minus pending. Capture converts the hold into posted entries, possibly for a different amount; void releases it; and a hold with no capture after its expiry (7 days for most card auths) is released by a scheduled sweeper that pages if its backlog ages past an hour."

L7: "Hold expiry is a risk decision. Releasing at 7 days means a late capture can overdraw a wallet; holding for 30 days means users see money locked up and call support. I'd have risk and finance sign a hold-policy table per merchant category, and I'd budget the negative-balance write-off it implies."

Why "concurrency" separates levels

L5: "Lock the wallet row with SELECT … FOR UPDATE." Correct for the user's wallet. Then every transaction also credits fees_revenue or debits omnibus_cash, and those rows are locked by every transaction in the system. Throughput collapses to the rate one row can be updated — roughly 1–3K commits/s on good hardware, far less with network round trips inside the lock.

L6: "I classify accounts by contention. User wallets: thousands of accounts, a few posts per day each — pessimistic lock, no problem. House accounts like fees, settlement and omnibus: every transaction touches them. Those I either split into N sub-accounts chosen by hash, or I don't maintain a running balance for them at all — I append entries and aggregate asynchronously, because nothing ever needs to check that the fee account is non-negative."

L7: Recognizes this as an account-design governance problem: "The hot-account bug gets reintroduced every time a product team designs a new flow with a single clearing account. I'd add a contention class to the account schema and have the ledger reject high-frequency postings to accounts that require a balance guard unless they're sharded."

Positions to Commit To#

PositionRationale
Entries are truth; balances are derived (and cached)A balance without entries cannot be explained; entries without a balance can always be summed
Reject unbalanced transactions at write timeConservation is cheap to enforce once and catastrophic to discover later
Overdraft guard in the same transaction as the postA balance check in one transaction and a debit in another is a race by construction
Holds are a pending dimension, not an early debitAvailable = posted − pending; capture, void and expiry are explicit transitions
Every transaction carries a producer-supplied external_id, uniqueRetries and replayed events must be no-ops; uniqueness is the mechanism
Integer minor units per ISO 4217 currency; never floats; never mix currencies in one balanceJPY has 0 decimals, KWD has 3; FX is a transaction through conversion accounts, not arithmetic
Daily reconciliation of customer liabilities against the safeguarding bank account, with an ownerThe ledger proves internal consistency; only reconciliation proves the money exists

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Stored-value wallet (we hold the money)Real-time spend authorization; no overdraft; regulated safeguardingSynchronous ledger on the spend path; holds; per-account balance guard; daily FBO reconciliationDouble spend, lost update, safeguarding shortfallΣ wallet balances = safeguarded cash, daily, to the cent
Company books of record (ledger behind many products)Ingest events from dozens of systems; finance reporting; auditAsynchronous, event-fed ledger; posting rules per event type; clearing accounts that must net to zeroMissing or late events; unexplained balances at month-endClearing accounts at zero within N days; close in ≤ 3 days
Core banking (deposit accounts)Interest accrual, statements, regulatory reporting, overdraft productsFull core: product engine, accruals, GL integration, regulatory feedsMis-accrued interest, reporting errorsRegulatory exam; deposit insurance reporting

🎯 Staff Move: "I'll assume a stored-value wallet: we hold user funds in a safeguarding account at a partner bank, users top up, spend by card and send money to each other. That puts the ledger on the synchronous spend path and makes no-overdraft a hard invariant. I'll design entries so the same ledger can later feed the company's books — but books-of-record is an asynchronous problem and core banking is a different product. Tell me if you want one of those."

Where the Design Splits#

#Fault LineThe Tension
1Derived vs Materialized BalanceSum entries on read (always right, slow) or maintain a running balance (fast, can drift) — and where does the overdraft check run?
2Hold Model: Pending Dimension vs Early DebitReserve funds as pending (honest, more states) or debit immediately and reverse (simpler, messy statements and expiry bugs)?
3Hot Accounts: Serialize vs Split vs DeferHouse accounts touched by every transaction — lock them, shard them into sub-accounts, or stop maintaining their balance synchronously?
4Synchronous vs Event-Fed PostingLedger on the request path (authoritative, latency-coupled) or downstream of events (decoupled, can't authorize spend)?
5Multi-Currency: Per-Currency Accounts vs Base-Currency LedgerKeep balances per currency with explicit FX transactions, or convert everything to a base currency and lose the trail?

How Real Companies Built It#

Why this section belongs here: Ledgers are one of the few areas where companies have published unusually concrete designs. Citing them shows you've studied the invariants, not just the tables.

Stripe — A Ledger Whose Accounts Must Clear to Zero#

Stripe describes its Ledger as an immutable, double-entry system of record that models money movement as balances flowing between accounts through discrete events, and that underpins its data-quality platform. Its key validation idea is clearing: an intermediate account — for example, one that holds a charge before it reaches a business's balance — should return to zero at steady state, so a missing, late or wrong event shows up as a balance that never clears. Stripe reports processing about five billion events per day, with 99.99% of dollar volume fully ingested and verified within four days (Stripe engineering).

Staff insight: The design point is that non-zero clearing accounts are the alert. In an interview, say: "Every flow passes through a clearing account that should net to zero within a known window; a balance that ages past that window is a named break with an owner." That turns "reconciliation" from a vague batch job into a metric.

Square — Books, an Immutable Double-Entry Database Service#

Square built Books, an immutable double-entry accounting service on Cloud Spanner, with three core tables — books (accounts with balances), journal entries (transactions) and book entries (the debit and credit legs). It is append-only: incorrect entries are fixed by writing corrective entries, never by editing. To scale horizontally it replaced monotonically increasing integer keys with UUIDs, namespaced user-influenced indexes by a tenant "shelf" for multi-tenant isolation, and used interleaving and timestamp-sharding strategies to avoid Spanner hotspots. At publication it managed about 20TB with a team of three (Square engineering).

Staff insight: Even a team running a globally distributed database had to design keys to avoid hotspots — sequential IDs concentrate writes. The ledger's scaling problem is key and account design, not raw capacity. Say: "I'd avoid sequential transaction IDs and single house accounts on the write path for the same reason."

TigerBeetle — Holds as Two-Phase Transfers in the Ledger Itself#

TigerBeetle, a purpose-built financial transactions database, models holds natively: a pending transfer reserves its amount in the accounts' debits_pending and credits_pending fields without changing posted balances; it is later resolved exactly once by posting (fully or partially, with the remainder released), voiding, or timing out, and resolution is itself a new immutable transfer that references the pending one. Balance invariants such as "debits must not exceed credits" are enforced across pending amounts too (TigerBeetle docs).

Staff insight: Holds are not an application-layer afterthought; a ledger built for this treats pending as a balance dimension with exactly-once resolution and built-in expiry. In an interview: "Available equals posted credits minus posted debits minus pending debits, and the overdraft guard runs against available, not posted."

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"We store the balance on the wallet""Two spends arrive at the same millisecond against $50. Walk me through both."Lost update, guard placement
"We use double-entry""Every transaction also credits the fee account. What's your throughput?"Hot-account awareness
"We debit on authorization""The merchant never captures. Who notices, and when?"Holds, expiry ownership
"Posting is idempotent""The producer retries with the same key but a different amount. What happens?"Key + payload hash semantics
"The ledger balances""Your ledger says users hold $41.2M. The bank says $41.1M. Now what?"Reconciliation, safeguarding, ownership
"We support multiple currencies""A user converts €100 to USD. Show me the entries."FX as a transaction, per-currency balancing
"We correct mistakes with new entries""The mistake was in last quarter, which finance already closed."Effective date vs posted date, period locks

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: One service — the ledger — writes entries, and it does so in a single database transaction with the balance update, the overdraft guard, the idempotency record and the outbox row. Everything that doesn't need to block a spend is asynchronous: house-account aggregation, statements, warehouse feeds, hold expiry and reconciliation. The two numbers on the observability panel that must always be zero are ledger.unbalanced_rejects (a producer tried to create or destroy money) and ledger.balance_drift (a balance row disagrees with its entries). The one that wakes compliance is safeguarding.shortfall_usd.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Balance"Balance column on wallet""Entries are truth. Balance row is a cache updated in the same txn, verified against Σ entries nightly."
Overdraft"Check balance, then debit""Conditional update: UPDATE … SET available = available - x WHERE available >= x. Zero rows = insufficient funds."
Holds"Debit and refund later""Pending dimension. Available = posted − pending. Capture, void, expire — each exactly once."
Hot accounts"Lock the row""User wallets lock fine. House accounts get sub-accounts or async aggregation — they have no floor to guard."
Idempotency"Dedupe in the service""external_id unique per producer; replay returns the original txn; same key, different payload = 409."
Multi-currency"Store amount and currency""Balance per (account, currency). FX is a four-leg txn through conversion accounts, rate recorded."
Audit"Keep logs""Append-only entries, posted vs effective date, period close, reconciliation evidence retained 7 years."

Numbers to Bring#

MetricValueWhy It Matters
Entries per business transaction2 minimum; 4–6 typical (principal, fee, FX legs)Size storage and write rate in entries, not transactions
Ledger entry row size~150–300 bytes incl. indexes1B entries ≈ 150–300GB — storage is rarely the constraint
Single hot-row update rate (Postgres, row lock, local txn)~1–3K commits/s best case; ~200–500/s with app round trips inside the lockWhy a single fee or omnibus account caps total throughput
Typical wallet transaction volume10M users × ~20 txns/month ≈ 77 TPS average, ~10× at peakA single primary handles it; correctness and hot accounts are the problem
Card authorization response budget~1–2s end to end, ledger share ≤ 50msThe ledger is on the card network's critical path
Card authorization hold lifetime~7 days typical; longer for some categories (hotels, car rental)Drives hold expiry policy and pending-balance accounting
ISO 4217 minor-unit exponents0 (JPY, KRW), 2 (USD, EUR), 3 (KWD, BHD, OMR)Fixed "cents = amount × 100" is a bug for ~1 in 10 currencies
Stripe Ledger scale (published)~5B events/day; 99.99% of dollar volume verified within 4 daysEven at the top, verification is measured in days, not milliseconds
Financial record retentioncommonly 5–7 years depending on jurisdiction and record typeAppend-only storage must be designed for years of history
Reconciliation cadence for safeguarded fundsdaily (often required), intraday at higher maturityShortfalls must be detected and escalated within a business day
Acceptable balance driftexactly 0 minor unitsAny non-zero drift means a code path writes balances without entries

Interview Walkthrough

The most common mistake: Candidates spend 15 minutes on the wallet API, KYC screens and P2P social features, then run out of time before the interviewer asks the only question that matters: "Two spends hit the same $50 balance at once. And by the way, the fee account is touched by every transaction." Compress the product surface to ~8 minutes and spend the rest on invariants, holds, hot accounts and reconciliation.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Users top up a wallet from a card or bank account, pay merchants with a wallet-linked card, send money to other users, and cash out to a bank. We hold the funds in a safeguarding account at a partner bank. The ledger is the system of record for every balance."

Then the non-functional requirements, which is where the design lives:

"Three constraints drive everything. One: no wallet goes below zero, and no money is created or destroyed — those are invariants, not goals. Two: card authorizations need a ledger answer in tens of milliseconds, because the card network gives the whole chain one to two seconds. Three: every balance must be explainable to an auditor years later, and the sum of all wallets must equal the cash in the bank every day. I'll assume 10 million wallets, roughly 100 transactions per second average and 1,000 at peak — which means throughput is not the hard part, but a few specific accounts will be."

Then name the underspecified parts:

"I'd confirm: single currency or multi-currency? Do we allow overdrafts or negative balances for any account type? Are we the licensed e-money issuer, or does a partner bank hold that obligation? I'll assume multi-currency balances, no consumer overdraft, and that we're responsible for the daily safeguarding reconciliation."

🎯 Staff Move: Saying "a few specific accounts will be" the scaling problem — before drawing anything — tells the interviewer you know where ledgers actually break: house accounts, not user count.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Account: account_id, owner (user, merchant, house), currency, type (asset/liability/revenue/expense), normal_side (debit/credit), floor (minimum available, usually 0 for wallets, none for house accounts), contention_class
  • Transaction: txn_id, external_id (producer idempotency key), producer, posting_rule, rule_version, effective_at, posted_at, request_hash
  • Entry: entry_id, txn_id, account_id, direction (D/C), amount_minor, currency
  • Hold: hold_id, account_id, amount_minor, state (PENDING / POSTED / VOIDED / EXPIRED), expires_at, external_id
  • Balance: account_id, currency, posted, pending_debits, available (= posted − pending_debits), version, last_entry_id

Ledger API (internal — product services never touch tables):

POST /v1/transactions        { external_id, producer, rule, entries[] | params }
  → 201 { txn_id }  |  200 { txn_id } (replay)  |  409 key_reused_with_different_payload
  → 422 unbalanced  |  402 insufficient_available { account_id, available }

POST /v1/holds               { external_id, account_id, amount_minor, currency, expires_at }
POST /v1/holds/{id}/post     { external_id, amount_minor (≤ hold), entries[] }
POST /v1/holds/{id}/void     { external_id }
GET  /v1/accounts/{id}/balance?as_of=2026-03-31T23:59:59Z
GET  /v1/accounts/{id}/entries?cursor=…

Money is integer minor units with an explicit ISO 4217 currency. Never floats. Never one balance across currencies.

🎯 Staff Move: "Producers send business intent — 'P2P transfer of 2,500 cents from u1 to u2' — and the ledger applies a versioned posting rule to produce entries. That way finance reviews posting rules in one place, and no product team hand-writes debits and credits."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk the P2P flow in 90 seconds:

  1. App sends POST /transfers with idempotency key k1: u1 → u2, $25.00.
  2. Wallet service checks KYC tier and daily limits, then calls the ledger with external_id = wallet:k1, rule p2p.v3.
  3. Ledger opens one transaction: insert into transactions (unique on producer, external_id) — a conflict means replay, return the original.
  4. Apply rule → entries: debit wallet:u1:USD 2500, credit wallet:u2:USD 2500. Check Σ D = Σ C.
  5. Conditional update: UPDATE balances SET posted = posted - 2500, available = available - 2500 WHERE account = 'wallet:u1:USD' AND available >= 2500. Zero rows → roll back, 402.
  6. Credit u2's balance, insert entries and an outbox row, commit. One transaction, ~5–15ms.
  7. Outbox relay publishes ledger.txn.posted; notifications, statements and the warehouse consume it.

🎯 Staff Move: Say out loud: "The overdraft guard and the debit are the same statement. There is no read-check-write window. And the idempotency insert is in the same transaction, so a crash anywhere either posts everything or nothing." You've now spent ~8 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path for a user-to-user transfer, and it's the Senior-level design. What makes ledgers hard is three things: card holds, where money is reserved before it moves; house accounts that every transaction touches, which will cap our throughput; and proving every day that wallet balances match real cash. I'd like to go deep on those, plus multi-currency if we have time. Where would you like to start?"

If no preference: start with holds. They break more naive designs than anything else.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: Holds and the card authorization path (7–8 min)

"A card auth arrives: $60 at a gas station. The ledger creates a hold — pending debit of $60 on u1's wallet — guarded by available >= 6000. Posted balance doesn't change; available does. The response goes back in under 50ms. Later, the capture arrives for $43.17 — gas stations authorize high and capture the real amount. Capture is one transaction: post debit $43.17 wallet → merchant clearing, release the remaining $16.83, mark the hold POSTED. If no capture arrives, the sweeper releases the hold after the policy TTL — 7 days for most merchant categories, longer for hotels and car rental."

Quantify: "At 1,000 auths/s peak with 2% never captured, the sweeper releases ~1.7M holds a day. If it stops, user available balances shrink silently — so hold.expired_backlog_age pages at one hour."

Who pays: "If we release too early and the capture arrives late, the wallet can go negative — the business eats it as a write-off. If we hold too long, the user pays with locked funds and support pays with tickets. Risk owns the TTL table."


Deep dive 2: Hot accounts (6–7 min)

"Every card capture credits merchant_clearing; every transaction with a fee credits fee_revenue. If those are single rows with running balances, every transaction serializes on them — a few hundred to a few thousand per second, total, for the company. Two fixes depending on whether the account needs a guard. Accounts with no floor — revenue, clearing — I don't maintain synchronously: entries append, an aggregator rolls them into the balance every second. Accounts with a floor that are still hot — say the omnibus account that funds cash-outs — I split into 16–64 sub-accounts chosen by hash, rebalanced by a background job, and the logical balance is the sum."


Deep dive 3: Idempotency and replay (4–5 min)

"Every producer supplies an external_id. The ledger stores (producer, external_id) → txn_id, request_hash. Same key and same hash returns the original txn with 200. Same key, different hash returns 409 — that's a producer bug and I want it loud. Retention is forever, because the key is on the transaction row itself, not in a TTL cache — ledger replays of a year-old event must still be no-ops."


Deep dive 4: Reconciliation and safeguarding (5–6 min)

"Internally, three checks: every transaction balances (enforced at write), every balance row equals Σ entries (nightly), and clearing accounts return to zero within their window — card clearing within T+2 of capture. Externally: Σ customer wallet balances, per currency, must equal the safeguarding bank balance at end of day, after known in-flight items. A shortfall is a regulatory event: compliance is notified the same business day and the business tops up the account from its own funds while we investigate."


Deep dive 5: Ownership (2–3 min)

"The ledger platform team owns the ledger service, posting-rule engine and drift checks, and carries the pager for ledger.balance_drift and ledger.unbalanced_rejects. Finance owns the chart of accounts and signs off on posting-rule changes. Treasury and compliance own safeguarding reconciliation and the shortfall escalation. Wallet product owns flows and limits but never writes entries."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: conservation is enforced at write time by refusing unbalanced transactions; non-negativity is enforced by a guarded update in the same transaction as the post; holds give reserved money its own dimension; and daily reconciliation proves the ledger matches real cash. Hot accounts are the scaling problem, and they're solved by account design."

The evolution closer:

"What I'd build later: account sharding when write volume — not storage — demands it, at which point cross-shard transfers become two-phase with a clearing account per shard pair; per-legal-entity ledgers when we expand to a second regulated market; and a books-of-record feed for finance. What I'd not build: our own database engine for the ledger, unless we get to tens of thousands of transactions per second sustained."

🎯 Staff Move: End on who answers the regulator. "If the safeguarding reconciliation shows a shortfall, compliance knows within the business day and the company funds it before we finish debugging. That escalation path matters more than any table design."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Product tour10 min on P2P UX, contacts, KYC screensOne sentence of product; jumps to invariants
Schema without invariantsDraws tables, never states what must always be trueStates conservation and non-negativity before any table
Sharding theaterShards by user ID for "billions of users""~1K TPS. One primary. The hot spots are house accounts."
No holdsDebit on auth, refund on cancelPending dimension in the first 15 minutes
No external truthLedger balances, therefore correct"Ledger proves consistency; reconciliation proves existence"
Floating-point moneyDECIMAL(10,2) everywhere, or worse, floatsInteger minor units per ISO 4217 exponent

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

A ledger is the interview where the database is not a detail. The candidate must hold an invariant across concurrent writers, crashes, retries, late events and corrections — and must do it without serializing the company through one lock. There's no "eventually consistent" escape hatch for an overdraft check, and no "we'll fix it in post" for a balance that was wrong for three weeks. That combination — strict invariants, a hot path measured in milliseconds, and an audit trail measured in years — is exactly what separates someone who has used a database from someone who has owned a system of record.

It also has an unforgiving failure distribution. A lost update in a social feed is a missing like. A lost update in a wallet is money created from nothing — and every such cent is a liability the company owes users without the cash to back it. When reconciliation finally catches it, the question is not "how do we fix the bug" but "how much do we owe, to whom, since when" — and only an append-only ledger can answer that.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"Your ledger says customers hold $41,236,118.42 in USD wallets. The bank statement for the safeguarding account says $41,229,904.10. It's 4pm. What do you do, who do you call, and how do you find the $6,214.32?"

A candidate who answers with in-flight items first (top-ups credited in the ledger but not yet settled by the bank, cash-outs sent but not yet debited), then a break taxonomy, then the escalation clock (compliance today, company funds the gap), then the forensic query (entries since last clean reconciliation, grouped by producer and rule version) has operated a ledger. A candidate who says "check the logs" has read about one.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Stored-value wallet → the ledger authorizes spend

  • Constraint: synchronous answers for card auths (≤ 50ms ledger budget), hard no-overdraft, regulated safeguarding of customer funds
  • Strategy: ledger on the request path; guarded balance updates; holds; per-currency accounts; daily reconciliation against the safeguarding account
  • Failure mode: double spend, lost update, safeguarding shortfall, holds that never expire
  • Who pays for imperfection: the company (write-offs, regulatory exposure) and users (locked or missing funds)

Company books of record → the ledger explains money that moved elsewhere

  • Constraint: dozens of upstream producers emitting events at different latencies; finance needs close in days; auditors need lineage
  • Strategy: event-fed ledger; posting rules per event type; clearing accounts that must net to zero; completeness and timeliness metrics per producer
  • Failure mode: missing or late events produce balances that never clear; month-end becomes forensic work
  • Who pays: finance (close time), auditors (sampling more), engineering (backfills)

Core banking → the ledger is the product

  • Constraint: interest accrual, statements, overdraft products, regulatory and deposit-insurance reporting
  • Strategy: a full core-banking platform, usually bought, with a GL integration
  • Failure mode: mis-accrued interest across millions of accounts; reporting errors to regulators
  • Who pays: the bank, in restitution and exam findings

2.2 When NOT to Build a Ledger#

  • You don't hold the money. If a payment processor or partner bank holds funds and gives you balances via API, your "wallet" is a cache of their ledger. Build a read model and reconcile it; don't build a second system of record that will disagree with theirs.
  • You're tracking non-redeemable points. Loyalty points with no cash value still benefit from double-entry, but you don't need safeguarding reconciliation, holds or regulatory retention. A simple append-only points table is fine until points become redeemable for money — then they're stored value and the rules change.
  • You have one revenue stream and a processor dashboard. Revenue reporting from the processor plus a nightly export is enough until month-end close takes more than ~5 days or a second money flow appears.
  • You want "a ledger database" because it sounds rigorous. A purpose-built ledger engine earns its place above ~10K sustained transactions per second or when hot-account contention can't be designed away. Below that, a relational database with the right constraints is the ledger.

🎯 Staff Insight: "Before I build a ledger, I want to know whose balance sheet the money is on. If it's ours, we need a system of record and a safeguarding reconciliation. If it's the partner bank's, we need a well-reconciled read model — and building a second ledger creates a second truth."

2.3 What the Interviewer Leaves Underspecified#

Interviewers deliberately omit:

  • Who holds the funds — us as an e-money issuer, a partner bank, or a processor — which decides whether the ledger is legally authoritative
  • Negative balances — are overdrafts ever allowed (e.g., late card captures, chargebacks)? Every real wallet ends up with some
  • Currencies — multi-currency wallets need per-currency balances and FX transactions; currencies with 0 or 3 decimal places break naive code
  • Card issuing — if the wallet has a card, the ledger is on the authorization path with a hard latency budget
  • Corrections after close — what happens when finance has signed off a month and a bug is found in it
  • Retention and access — who may read balances (support, finance, regulators) and for how many years

Staff engineers surface these and commit. Senior engineers assume them away and get surprised by the follow-up.

2.4 Precise Terminology#

TermWhat It MeansWhy It Matters in the Interview
AccountA named bucket of value in one currency, with a type and normal sideUser wallets are liabilities to us; cash at the bank is an asset
Entry (leg)One debit or credit on one accountImmutable; carries the cause
Transaction (journal entry)A set of entries that sum to zero per currencyThe atomic unit of money movement
Posted balanceΣ of posted entriesWhat has definitively moved
PendingAmount reserved by open holdsReduces available, not posted
Available balancePosted − pending debits (for a liability account)What the guard checks
Hold / authorizationA reservation resolved once by post, void or expiryCard auths, in-flight cash-outs
Clearing / suspense accountIntermediate account that should return to zeroA non-zero aged balance is a break
Omnibus / FBO accountThe real bank account holding all customers' pooled fundsΣ customer wallets must equal it
SafeguardingRegulatory obligation to protect customer funds separately from company fundsShortfalls are reportable
Effective date vs posted dateWhen the economic event happened vs when we recorded itCorrections after period close
Posting ruleVersioned mapping from business event to entriesFinance reviews rules, not code

🎯 Staff Insight: If the interviewer says "the user's balance", ask: "Posted or available? They differ by every open hold, and the card network only cares about available."


3. Where the Design Splits#

Every ledger decision has a technical side (which row, which lock, which constraint) and an organizational side (who signs off on posting rules, who owns the break queue, who tells the regulator). Interviewers grade the second side.

3.1 Fault Line 1: Derived vs Materialized Balance#

The tension: A balance derived by summing entries can never disagree with the entries, but costs O(entries) per read and gives you no row to lock for an overdraft guard. A materialized running balance is O(1) and gives you that row — and becomes a second record that can drift if any code path writes one without the other.

ChoiceWhat WorksWhat BreaksWho Pays
Derived only (Σ entries on read)Zero drift by constructionReads slow as history grows (10K+ entries per active account/year); no row for the guard → needs serializable isolation or advisory locksOn-call (latency), users (slow balance screens)
Derived + periodic snapshotsRead = snapshot + entries since; bounded costStill no guard row; snapshot job is another thing to ownLedger team (snapshot pipeline)
Materialized balance, same transaction as entriesO(1) reads; guard is a conditional UPDATEDrift if any writer bypasses the ledger API; hot rows for busy accountsLedger team (drift check, single-writer discipline)
Materialized balance, updated asynchronouslyNo hot-row contentionCan't authorize spend against it; laggingBusiness (overdrafts) — only acceptable for accounts with no floor
Diagram: 3.1 Fault Line 1: Derived vs Materialized Balance

Staff default: "Materialized balance per (account, currency), updated in the same transaction as the entries, with available >= amount in the WHERE clause for accounts with a floor. The ledger service is the only principal with write grants. A nightly job recomputes Σ entries for every account touched that day and compares; any drift pages the ledger on-call as a Sev-1."

When to deviate:

  • Books-of-record ledger with no spend authorization: derived balances plus daily snapshots are simpler and drift-proof. Nobody needs a guard.
  • Very high-throughput accounts with no floor: asynchronous materialization (see Fault Line 3).
  • Purpose-built ledger engine: engines that enforce balance invariants internally make the drift check redundant — but you still reconcile.

🧭 Principal Move: "The drift check is an audit control, not a test. I'd retain its daily output for the full retention period, because an auditor will ask for evidence that balances matched entries on every day of the year, not just today."

❌ Common L5 Trap: "I'll read the balance, check it in the application, then write the debit — inside a transaction, so it's safe." Under the default READ COMMITTED isolation in most relational databases, two transactions can both read $50 and both write. The fix isn't a stronger isolation level everywhere; it's making the check and the write the same conditional statement, or locking the row first with SELECT … FOR UPDATE.


3.2 Fault Line 2: Hold Model — Pending Dimension vs Early Debit#

The tension: Card authorizations, in-flight cash-outs and pre-authorized rides all reserve money before it moves. You can model that reservation honestly as a pending amount, or pretend it moved and reverse it later.

ChoiceWhat WorksWhat BreaksWho Pays
Early debit + reversalOne balance; simple codeTwo statement lines per purchase; partial captures need a debit, a reversal and a re-debit; expiry bugs leave money "spent" foreverUsers (confusing statements), support (tickets)
Pending dimension (holds)Available reflects reservations; posted stays clean; partial capture is naturalMore states; expiry must be owned; late captures after expiryLedger team (sweeper), risk (TTL policy)
Holds in a separate serviceLedger stays simpleGuard must consider holds the ledger can't see → raceEveryone (double spend)
Diagram: 3.2 Fault Line 2: Hold Model — Pending Dimension vs Early Debit

Staff default: "Holds live in the ledger, in the same transaction scope as balances, because the overdraft guard must see them. A hold reduces available and leaves posted untouched. Capture is one transaction: post entries for the captured amount, release the rest, transition the hold. Every transition is conditional on the current state, so a duplicate capture message can't post twice. Expiry TTL comes from a policy table keyed by merchant category."

When to deviate:

  • Instant P2P and top-ups: no hold — post directly. Holds exist only where there is a gap between authorization and settlement.
  • Late captures: card networks allow captures after the auth has expired in some cases. The ledger must accept a force-post that may take a wallet negative, and route the resulting negative balance to collections. Refusing it doesn't stop the money moving; it only stops you recording it.

🧭 Principal Move: "Hold TTL is a dollar decision. I'd ask risk to model negative-balance write-offs from late captures against support cost from long holds per merchant category, and publish the TTL table with their sign-off. Engineering implements it; engineering doesn't pick 7 days because a blog said so."

❌ Common L5 Trap: "On auth, debit the wallet; if the merchant doesn't capture in a week, a cron job refunds it." Then the cron job fails silently for 4 days during a deploy freeze. Users see money gone that was never spent. With a pending dimension, a stuck sweeper still leaves posted balances correct — and hold.expired_backlog_age makes it loud.


3.3 Fault Line 3: Hot Accounts — Serialize vs Split vs Defer#

The tension: In double-entry, every transaction has a counterparty. User wallets are spread across millions of rows; house accounts — fee revenue, merchant clearing, card settlement, the omnibus funding account — are touched by a large fraction of all transactions. A running balance on a house account is a global lock.

ChoiceWhat WorksWhat BreaksWho Pays
Single row, lock per txnSimple; exactThroughput capped at one row's commit rate (~200–3K/s)Business (throughput ceiling at peak)
Split into N sub-accountsN× throughput; guard still possible per sub-accountLogical balance is a sum; a sub-account can run dry while others have funds → rebalancer neededLedger team (rebalancing, reporting views)
Append entries, aggregate asyncNo contention at allNo synchronous balance → no guardNobody, if the account has no floor
Batch netting (micro-batches)One update per batch of 100–1,000 txnsAdds batch latency (10–100ms); complicates per-txn failureLatency budget
Diagram: 3.3 Fault Line 3: Hot Accounts — Serialize vs Split vs Defer

Staff default: "Classify every account by contention at creation. Revenue, expense and clearing accounts have no floor — their entries append and an aggregator rolls them up every second; their balance is never on the write path. Floored accounts that are hot — the omnibus that funds cash-outs — split into 32 sub-accounts with a rebalancer that moves funds between them when any drops below 2× its median outflow per minute. User wallets stay single rows."

When to deviate:

  • A single extremely hot user account (a large merchant receiving 2,000 payments/s): give the merchant a per-day or per-minute sub-ledger and settle it into the main account in batches. That merchant cares about settled balance, not millisecond accuracy.
  • Low volume (< 100 TPS total): don't split anything. You're adding a rebalancer to solve a problem you don't have.

🧭 Principal Move: "Hot accounts are an account-design problem that recurs every time a product team adds a flow. I'd add a contention_class to the account schema and have the ledger refuse more than a configured post rate on a guarded single-row account — so the design review happens before launch, not during the first sale."

❌ Common L5 Trap: "We'll shard the ledger by user ID to scale." User wallets were never the bottleneck — every shard still credits the same fee_revenue row. Sharding by user without fixing house accounts just turns a local lock into a distributed one.


3.4 Fault Line 4: Synchronous vs Event-Fed Posting#

The tension: If the ledger is on the request path, it can authorize spend — and its latency and availability become the product's. If it consumes events after the fact, it's decoupled and easy to scale — and it can't stop a double spend because it learns about the spend too late.

ChoiceWhat WorksWhat BreaksWho Pays
Synchronous ledger (authoritative)Overdraft prevention; single source of truth for spendLedger outage = wallet outage; must meet card-auth latencyLedger on-call (availability target ≥ 99.99%)
Event-fed ledger (books)Decoupled; producers own their availabilityBalance is lagging; can't guard spend; needs completeness metricsFinance (late or missing events at close)
Hybrid: sync for wallets, events for booksEach path fits its jobTwo posting paths; need a contract on which accounts each may touchLedger platform (contract governance)
Diagram: 3.4 Fault Line 4: Synchronous vs Event-Fed Posting

Staff default: "Hybrid, with account ownership as the contract. Customer wallet accounts are written only by the synchronous path. Event-fed producers post to clearing and house accounts. Where they meet — a card capture arriving by file that must match a hold — the clearing account is the join, and its aged non-zero balance is the reconciliation signal."

When to deviate:

  • No card or real-time spend: pure event-fed. A payout platform that pays sellers weekly doesn't need the ledger in the request path.
  • Ledger availability can't reach the product's target: keep a bounded offline-authorization allowance (e.g., approve card auths under $25 if the ledger times out, and post later) with risk sign-off — the card networks' stand-in processing is the industry precedent for this.

🎯 Staff Insight: "Making the ledger synchronous means its availability target is the wallet's availability target. I'd say that number out loud — 99.99%, about 4 minutes a month — and design the ledger as a small, boring service with one database, not a microservice mesh."

❌ Common L5 Trap: "The wallet service publishes an event and the ledger consumes it." Then the interviewer asks how the wallet service knew the user had funds. It read a balance that was already stale by the time the event landed — the double-spend window is the consumer lag.


3.5 Fault Line 5: Multi-Currency — Per-Currency Accounts vs Base-Currency Ledger#

The tension: Users hold EUR and USD; the business reports in one currency. Converting everything to a base currency at posting time makes reports easy and destroys the ability to say what each user actually holds. Keeping per-currency balances keeps the truth and makes FX an explicit, audited transaction.

ChoiceWhat WorksWhat BreaksWho Pays
Base-currency ledgerSimple reportingUsers' real holdings lost; FX gains/losses invisible; rounding driftFinance (unexplainable FX P&L), users (wrong balances)
Per-currency balances, FX via conversion accountsEach currency balances on its own; FX P&L explicitFour-leg FX transactions; revaluation for reportingLedger team (posting rules), treasury (FX positions)
Per-currency + reporting-currency shadow amountsReports in base currency without re-computationShadow amounts need rate source and revaluation policyFinance (rate policy)

A €100 → USD conversion at 1.0850, with a 0.5% FX fee:

txn_fx_1:
  DEBIT   wallet:u1:EUR             10000 EUR
  CREDIT  fx_position:EUR           10000 EUR     -- EUR leg balances on its own
  DEBIT   fx_position:USD           10850 USD
  CREDIT  wallet:u1:USD             10796 USD     -- 10850 × 0.995, rounded down
  CREDIT  fx_fee_revenue:USD           54 USD     -- USD leg balances on its own
  meta: rate=1.0850 source=treasury_feed quote_id=q_7781 rounding=floor

Staff default: "Balance per (account, currency). A transaction must balance per currency, so FX always goes through position accounts — one leg per currency. The rate, its source and the quote ID are on the transaction. Rounding is a posting-rule decision, written down: floor for the user's credit, remainder to fee revenue, so we never credit a fraction we don't have."

When to deviate:

  • Single-currency product with no plans otherwise: keep the currency column and the per-currency check anyway — they cost nothing and save a migration.
  • High-volume FX: treasury may net FX positions intraday and hedge; the ledger records positions, treasury owns the hedging.

🧭 Principal Move: "Which entity holds which currency is a legal and banking decision before it's a schema decision. EUR balances may need to be safeguarded at an EU bank by an EU entity. I'd design accounts to carry the legal entity from day one, because splitting a shared ledger by entity later is a re-posting project."

❌ Common L5 Trap: "Store amounts as decimals with two places." The first JPY user (0 decimals) or KWD user (3 decimals) produces amounts off by 100× or truncated to the wrong precision. Minor units per the ISO 4217 exponent of the currency, stored as integers, is the only safe default.


4. When It Breaks#

4.1 Payday Rush — The Omnibus Account Becomes a Global Lock#

t=0:       Friday 09:00. Payroll partner pushes 1.8M salary credits over 20 minutes.
           Each credit debits funding_pool (single row, floored) and credits a wallet.
t=+30s:    Lock wait on funding_pool row climbs. ledger.post_p99: 12ms → 900ms.
t=+1min:   Card auths share the DB. Auth ledger budget is 50ms → 31% of auths time out.
t=+2min:   Card processor applies stand-in rules; declines climb to 18%.
t=+4min:   Page: ledger.post_p99 > 200ms for 2m, card.decline_rate > 5%.
t=+9min:   On-call throttles the payroll producer to 200/s. Card auths recover.
t=+3h:     Payroll backlog drains. 40K users complain salaries arrived late.

Detection: ledger.lock_wait_p99{account}, ledger.post_p99, ledger.top_accounts_by_post_rate, card.auth_ledger_timeout_rate.

Mitigation: throttle bulk producers into a separate, lower-priority lane with its own connection pool (bulkhead); split funding_pool into 32 sub-accounts at runtime if the schema supports it.

Prevention: contention classes at account creation; bulk producers get a batch-posting API that nets 1,000 credits against one funding debit per transaction; load test at 3× the largest known batch.

Owner: ledger platform on-call; payroll integration team owns their producer's rate contract.

4.2 The Double Spend — A Check Outside the Transaction#

t=0:       New "split bill" feature ships. It calls GET /balance, checks in app code,
           then POST /transactions with a pre-computed entry list (bypasses guarded rule).
t=+2 days: Users discover that tapping "Settle" on two devices settles twice.
t=+6 days: Reconciliation: 312 wallets negative by a total of $18,440. No alerts — the
           ledger accepted balanced transactions; the floor was never checked.
t=+6 days: Sev-1. Feature flag off. Negative balances routed to collections.

Detection: ledger.accounts_below_floor — a continuous query, not a nightly one. For floored accounts it must always be 0 except for flagged force-posts.

Mitigation: disable the endpoint that accepts raw entry lists for customer accounts; recover funds per policy.

Prevention: the guard lives in the ledger, not in callers. The raw-entries API is restricted to house accounts; customer-account debits go through rules that always apply the floor check. Database CHECK (available >= floor OR allow_negative) as a second line.

Owner: ledger platform (API design), split-bill team (incident follow-ups).

4.3 Stuck Holds — The Sweeper That Stopped#

t=0:       Sweeper deploy introduces a query that times out on a new index.
           Errors are retried silently; no holds expire.
t=+1 day:  Expired-but-pending holds: 1.7M, $41M of available balance locked.
t=+2 days: Support tickets: "money missing but no transaction shown." 6,000 tickets.
t=+2 days: Page finally fires on hold.expired_backlog_age > 24h (threshold was too loose).
t=+2 days: Fix deployed. Sweeper drains backlog in 40 min, rate-limited to protect auth path.

Detection: hold.expired_backlog_count, hold.expired_backlog_age (page at 1h), support.tickets{category=missing_funds}.

Mitigation: drain with rate limiting; proactive in-app message to affected users.

Prevention: sweeper liveness metric (hold.sweeper_last_success_age); available-balance computation optionally ignores holds past expires_at + grace even before the sweeper runs, so a dead sweeper degrades to correct-but-untidy.

Owner: ledger platform.

4.4 Safeguarding Shortfall — The Ledger Balances, the Bank Doesn't#

t=0:       A bank-rail integration change maps "returned" ACH top-ups to status "settled".
t=+1 day:  Ledger credits wallets for 230 top-ups the bank actually returned.
t=+1 day:  Daily recon: customer liabilities exceed FBO balance by $57,900. Break type: unknown.
t=+1 day:  Treasury, compliance notified. Company funds the gap from operating account same day.
t=+2 days: Forensics: breaks join to top-ups with bank return code R01 (insufficient funds).
t=+3 days: Reversing entries posted. Users whose top-ups were spent go negative → collections.

Detection: safeguarding.shortfall_usd{currency} daily (intraday at higher maturity); recon.breaks{type}; topup.return_rate against baseline.

Mitigation: fund the shortfall first, investigate second — the regulatory clock doesn't wait for root cause.

Prevention: top-ups credit available only after the rail's return window for risky rails, or credit with a hold that limits spend until settled; integration changes to rail status mapping require ledger-team review.

Owner: treasury and compliance (escalation), ledger platform (forensics), payments-rails team (fix).

4.5 FX Rounding Drift — A Cent at a Time#

t=0:       FX rule rounds both legs half-up instead of flooring the user credit.
t=+30 days: 9M conversions. fx_position:USD is short $41,600 — each conversion credited
           up to half a cent more than the position received.
t=+31 days: Month-end: treasury's FX P&L doesn't match ledger positions. Escalation.

Detection: per-currency position reconciliation against treasury's executed trades daily; fx.rounding_residual_minor summed per day.

Prevention: rounding is part of the posting rule spec, reviewed by finance; property test: for random amounts and rates, user credit + fee = position debit exactly.

Owner: ledger platform (rule), treasury (position reconciliation).

4.6 The Backdated Correction That Reopened Last Quarter#

A bug in Q1's fee posting is found in Q2, after finance closed Q1 and filed reports. An engineer posts correcting entries with effective_at in Q1. Q1's trial balance changes; the filed report no longer matches the ledger.

Mitigation: period locks — once a period is closed, no entries with posted_at in it; corrections post in the open period with effective_at pointing back, and reports distinguish "as originally reported" from "as restated".

Detection: ledger.posts_into_closed_period must be 0; the API rejects them.

Owner: finance (period close), ledger platform (enforcement).

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Hot account lock contentionledger.lock_wait_p99{account} > 50msAll posts touching that account; card auths if shared DBThrottle bulk lanes; split accountLedger on-call
Floor violation (double spend)ledger.accounts_below_floor > 0Individual wallets; company write-offsDisable bypass path; collectionsLedger platform + feature team
Balance driftledger.balance_drift_accounts > 0Trust in every balanceFreeze writer path; recompute from entriesLedger platform (Sev-1)
Unbalanced transaction attemptledger.unbalanced_rejects > 0One producer's flow failsProducer fixes rule; nothing postedProducing team
Stuck holdshold.expired_backlog_age > 1hUsers' available balancesRestart sweeper; rate-limited drainLedger platform
Safeguarding shortfallsafeguarding.shortfall_usd > 0RegulatoryFund gap same day; forensicsTreasury + compliance
Clearing account not clearingclearing.aged_balance{account, age>T+3}Month-end closeBreak triage by typeFinance ops
Posts into closed periodledger.posts_into_closed_period > 0Filed reportsReject at APIFinance + ledger platform
Ledger DB primary failureledger.availability, failover timeEntire wallet + card authSynchronous replica failover ≤ 30s; card stand-in policyLedger on-call + DBA

🎯 Staff Insight: The ledger is the one component where I'd choose consistency over availability without hesitation for customer accounts. "If the primary is gone, I'd rather decline card auths for 30 seconds during failover than approve them against a replica that might be behind. The card network's stand-in rules exist for exactly this — and they're bounded by limits risk has signed."


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingLists wallet features: top-up, pay, send, cash outAsks who holds the funds; commits to stored value; states conservation and non-negativity as invariantsAsks whether this ledger is one product's or the company's books; names finance as co-owner
CorrectnessTransactions table + balance column; app-level checksAppend-only entries; guarded update in same txn; unique external_id; reject unbalanced writesSingle-writer ledger API enforced org-wide; drift check retained as an audit control
HoldsDebit on auth, refund laterPending dimension; capture/void/expire exactly once; TTL policy; late force-postHold TTL table priced against write-offs and support cost, signed by risk and finance
ScaleShard by user IDIdentifies house accounts as the bottleneck; sub-accounts, async aggregation, batch nettingContention class in the schema; rate guard on single-row floored accounts; per-entity ledger roadmap
Operations"Add monitoring"ledger.balance_drift, accounts_below_floor, hold.expired_backlog_age, safeguarding recon with ownerSafeguarding shortfall escalation to compliance within the business day; reconciliation evidence retention
OrganizationWallet team owns it allLedger platform / finance / treasury / product split with explicit contractsRedraws boundaries: ledger as company system of record; posting rules owned by finance; product teams as producers

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Separates the two invariants"Conservation I check on write. Non-negativity needs current state — that's where concurrency lives."
Puts the guard in the write"WHERE available >= amount is the overdraft check. Zero rows updated means declined."
Models holds honestly"Available is posted minus pending. Capture posts and releases the remainder in one transaction."
Knows where ledgers get hot"User wallets are cold. The fee account is the bottleneck — and it has no floor, so I stop guarding it."
Distinguishes consistency from existence"The ledger proves it's internally consistent. Only the bank reconciliation proves the money exists."
Handles FX as a transaction"Each currency leg balances on its own; the rate and quote ID live on the transaction."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
UPDATE balance = balance - x with no historyNo audit trail; no way to answer "why"
Read balance, check in app, then writeRace by construction under default isolation
Floats or fixed two-decimal storageRounding drift; wrong for 0- and 3-decimal currencies
"Kafka exactly-once" as the idempotency storyProducer retries from outside Kafka still double-post without a unique key
No reconciliation against the bankTreats internal consistency as proof the cash exists
Deletes or updates entries to fix mistakesDestroys the audit trail the ledger exists to provide

5.4 Common False Positives#

  • Accounting vocabulary ≠ ledger design. Fluency in debits, credits and trial balances is valuable; if it doesn't lead to where the overdraft check runs under concurrency, it's bookkeeping, not systems design.
  • "We use serializable isolation" ≠ solved. It's correct and it shifts the problem to retry storms on hot accounts. Staff candidates say what aborts and how often.
  • Event sourcing enthusiasm ≠ a ledger. An event log is not double-entry; "replay the events to get the balance" still needs posting rules, balancing and a guard.
  • A distributed database ≠ scale solved. Globally distributed SQL still serializes writes to one hot row; key and account design matter more than the engine.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minStored value vs books vs core banking; commit; state both invariants
Entities & API3–5 minAccount, transaction, entry, hold, balance; minor units; posting rules
Architecture5–10 min≤ 8 boxes; single-writer ledger; guarded update in one txn
Holds10–17 minPending dimension, capture/void/expire, TTL policy, late capture
Hot accounts17–24 minContention classes; async aggregation; sub-accounts; batch netting
Idempotency + reconciliation24–32 minexternal_id, payload hash; drift check; clearing accounts; safeguarding
Pivot (interviewer's choice)32–42 minMulti-currency, sharding, corrections after close, multi-region
Wrap42–45 minTwo invariants, holds, hot accounts, reconciliation; evolution

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"Now shard it"Cross-shard atomicityShard by account; cross-shard transfers via per-shard clearing accounts, two local txns, idempotent; clearing nets to zero
"Add multi-currency"Per-currency balancingFX as four-leg txn through positions; rate on txn; rounding rule
"A bug corrupted last quarter"Audit modelReversing entries; effective vs posted date; period locks; restatement
"Make it multi-region"Consistency vs availabilityHome region per account; no active-active writes on balances; read replicas for statements
"Users want real-time balance on a smartwatch"Read scalingBalance reads from cache or replica with version; spend guard stays on primary
"How do you test this?"Correctness cultureProperty tests on posting rules (always balanced); concurrency tests on guards; drift check in CI against fixtures

6.3 What to Deliberately Skip#

  • KYC and onboarding flows — one sentence: "KYC tier sets limits; the ledger enforces them via the wallet service."
  • Card network message formats — treat the processor as a source of auth/capture/reversal events.
  • UI and notifications — consumers of ledger events.
  • Interest and accrual — that's core banking; say you're not building it.
  • Database engine internals — unless the interviewer asks why not a purpose-built ledger engine.

6.4 Follow-Up Questions to Expect#

  1. "Two requests spend from the same wallet concurrently. Show me exactly where one of them fails."
  2. "Every transaction credits the fee account. How many transactions per second can you do?"
  3. "A card auth for $60 is captured for $43.17 nine days later, after you released the hold. What happens?"
  4. "The producer retried with the same idempotency key but a different amount. What do you return?"
  5. "Customer liabilities exceed the bank balance by $6,214. Walk me through the next four hours."
  6. "Convert €100 to USD. Show me the entries, the rate and the rounding."
  7. "How do you shard this ledger, and what happens to a transfer between two shards?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design a digital wallet with a ledger."

Staff Answer

"First — do we hold the money? If we're the e-money issuer or the program manager responsible for safeguarding, the ledger is the legal record and sits on the spend path. If a partner bank holds balances and exposes them by API, our wallet is a read model of their ledger and I shouldn't build a second system of record. And is this a wallet ledger or the company's books? I'll assume we hold stored value with a wallet-linked card, multi-currency, no consumer overdraft.

Two invariants I'll protect: every transaction sums to zero per currency, and no customer wallet goes below zero except by explicit force-post. Constraints: card auths need a ledger answer in ~50ms, every balance is explainable years later, and Σ wallets equals the safeguarding bank balance daily. Volume is ~100 TPS average and 1,000 peak, so throughput isn't the hard part — house accounts are. I'll go: entities and posting rules → the guarded post → holds → hot accounts → idempotency → reconciliation → ownership."

Why this is L6:

  • Asks whose balance sheet the money is on before designing anything
  • States both invariants and separates them
  • Deprioritizes throughput with a number and names the real bottleneck

What L7 adds:

  • Asks which legal entities hold which currencies and where they're safeguarded
  • Asks whether other products already run their own ledgers — the consolidation problem
  • Frames finance as co-owner of the chart of accounts from day one
❌ Common L5 Trap

"Users table, wallets table with a balance, transactions table for history. A wallet service updates the balance and inserts a transaction row in a database transaction. Shard by user ID for scale."

Why this misses: Every element is reasonable, and the design has no answer for concurrent spends under default isolation, holds, the fee account every transaction touches, or the auditor's question. Sharding by user ID solves a scale problem that wasn't there.


Drill 2: Where Does the Overdraft Check Run?#

Prompt: "Two $40 spends hit a $50 wallet within a millisecond. Walk me through both."

Staff Answer

"Both requests reach the ledger with different external_ids. Each opens a transaction and runs UPDATE balances SET available = available - 4000, posted = posted - 4000, version = version + 1 WHERE account_id = 'wallet:u1:USD' AND available >= 4000. The database takes a row lock for the first; the second blocks on that row. The first inserts its entries and commits — available is now 1,000. The second's UPDATE re-evaluates the WHERE clause against the committed row, matches zero rows, and the ledger rolls back and returns 402 with the current available balance.

No read-then-write in the application, no reliance on a stronger isolation level. As a second line I keep CHECK (available >= floor OR allow_negative) on the balance row. Lock hold time is the transaction duration, ~5ms, so a single wallet can do ~200 posts per second — far above any human user."

Why this is L6:

  • Makes the check and the write a single conditional statement
  • Explains the row-lock behavior precisely, including the re-evaluation
  • Quantifies per-account throughput and shows why it's fine for wallets

What L7 adds:

  • Makes the guarded rule the only path to debit customer accounts, enforced by API design, not convention
  • Adds ledger.accounts_below_floor as a continuously evaluated control
❌ Common L5 Trap

"Wrap it in a transaction: read the balance, check it's at least $40, then insert the debit."

Why this misses: Under READ COMMITTED both transactions read $50 and both proceed. The candidate is relying on "transaction" to mean "serialized", which it doesn't by default. If pushed, the fix is SELECT … FOR UPDATE or the conditional UPDATE — but the interviewer had to push.


Drill 3: Holds and Partial Capture#

Prompt: "A gas station authorizes $100. It captures $43.17. Show me the ledger."

Staff Answer

"Auth: create hold H1 for 10,000 on wallet:u1:USD, guarded by available >= 10000. Posted unchanged; pending_debits +10,000; available −10,000. No entries yet — nothing has moved.

Capture: one transaction. Conditional transition UPDATE holds SET state='POSTED' WHERE id='H1' AND state='PENDING'. Post entries: debit wallet:u1:USD 4,317, credit card_clearing:USD 4,317. Balance row: posted −4,317, pending_debits −10,000, so available goes up by 5,683 net. The capture message carries its own external_id, so a duplicate capture finds the hold already POSTED and returns the original txn.

Later, the network settlement file debits card_clearing and credits settlement_bank; if card_clearing for this capture hasn't returned to zero within T+2, it's a reconciliation break."

Why this is L6:

  • Shows that holds change available, not posted, and produce no entries
  • Partial capture and release happen atomically, with the transition guarded
  • Connects to clearing and reconciliation

What L7 adds:

  • Notes that fuel and hospitality merchants need longer or different hold policies — a risk-owned table
  • Tracks hold.capture_ratio by merchant category as an input to that policy
❌ Common L5 Trap

"Debit $100 at auth. At capture, credit back $56.83."

Why this misses: The user sees a $100 debit and a $56.83 credit for one purchase; revenue and refund reporting are polluted; and if the capture never arrives, the full $100 stays "spent" unless a separate job remembers to reverse it.


Drill 4: Make It Concrete — The Fee Account#

Prompt: "Every transaction credits fee_revenue. Peak is 3,000 TPS. What breaks, and how do you fix it?"

Staff Answer

"If fee_revenue has a running balance updated in each transaction, every transaction takes its row lock. At ~5ms per transaction, one row supports ~200 per second; even with very short transactions, low thousands. 3,000 TPS doesn't fit.

But fee_revenue has no floor — nothing ever asks whether it can go negative. So I don't update its balance synchronously at all. Entries for it are inserted (inserts don't contend), and an aggregator consumes the outbox and updates its balance every second with one UPDATE per account per second. Its balance lags by ≤ 1–2s, which finance doesn't care about. For a floored hot account — the omnibus that funds cash-outs — I'd split into 32 sub-accounts picked by hash of the transaction ID, each guarded, with a rebalancer moving funds when any sub-account drops below twice its per-minute outflow."

Why this is L6:

  • Computes the ceiling and shows it's below requirement
  • Uses the "does it have a floor?" test to pick between async aggregation and splitting
  • Gives concrete parameters for the split and rebalance

What L7 adds:

  • Adds contention_class to account creation so this is decided in design review, not in an incident
  • Prices it: async aggregation is free; splitting costs a rebalancer and reporting views
❌ Common L5 Trap

"Use a bigger database instance, or move to a distributed SQL database."

Why this misses: A larger machine shortens commits a little; a distributed database still serializes writes to one row, often with higher per-commit latency due to consensus. The bottleneck is the data model, not the hardware.


Drill 5: Idempotency With a Changed Payload#

Prompt: "A producer retries a transfer with the same idempotency key but a different amount. What happens?"

Staff Answer

"The ledger stores (producer, external_id) uniquely with a hash of the canonical request — accounts, amounts, currency, rule. On a duplicate key, it compares hashes. Same hash: return the original txn_id with 200 — a safe replay. Different hash: 409 key_reused_with_different_payload, nothing posted, and ledger.key_conflicts{producer} increments.

It's almost always a producer bug — generating the key from something too coarse, like the user ID and date. I want it loud, not silently 'fixed' by posting the new amount or returning the old one. The key lives on the transaction row forever, so a replay of a year-old event is still a no-op."

Why this is L6:

  • Separates key identity from payload identity and gives both outcomes
  • Refuses to guess which payload is right
  • Retention tied to the transaction row, not a TTL cache

What L7 adds:

  • Publishes key-conflict rates per producing team as a quality signal
  • Standardizes key derivation in the producer SDK so teams can't invent coarse keys
❌ Common L5 Trap

"We'd see the key exists and return success."

Why this misses: If the original was $25 and the retry is $250, returning success tells the producer $250 moved when $25 did. The producer now believes something false about money — the worst kind of bug, because nothing errors.


Drill 6: Safeguarding Shortfall#

Prompt: "At end of day, customer USD wallets total $41,236,118.42. The safeguarding account shows $41,229,904.10. Go."

Staff Answer

"First, adjust for known in-flight items: top-ups credited to wallets but not yet settled into the account, cash-outs debited from wallets but not yet left the account, card settlements pending. Those are expected timing differences, and they sit in clearing accounts I can sum. If the residual is still a shortfall, it's a break and the clock starts: treasury and compliance are notified today, and the company funds the gap from its own money before we finish debugging — that's what safeguarding means.

Forensics: entries posted to customer accounts since the last clean reconciliation, grouped by producer and rule version, joined to bank transactions by reference. A shortfall usually means we credited users for money that never arrived — returned top-ups, reversed card loads — so I'd look at top-up return codes first. Each break gets a type and an owner, and recurring types get a prevention ticket."

Why this is L6:

  • Starts with timing differences rather than assuming a bug
  • Knows the regulatory reflex: fund first, investigate second
  • Forensic query is concrete and scoped from the last clean point

What L7 adds:

  • Moves to intraday reconciliation for rails with high return rates
  • Sets a control: top-ups on reversible rails are spend-limited until their return window passes, signed by risk
❌ Common L5 Trap

"Find the bug, fix it, and post adjusting entries to make them match."

Why this misses: Posting entries "to make them match" hides the break instead of explaining it. And treating it as a bug-fix ticket ignores that a safeguarding shortfall is a regulatory event with a same-day clock.


Drill 7: Build vs Buy#

Prompt: "Should we build our ledger on Postgres, buy a ledger-as-a-service, or use a purpose-built ledger database?"

Staff Answer

"Depends on throughput, team, and how central the ledger is to the product. At ≤ 2,000 TPS with a team that knows Postgres, a well-designed relational ledger — append-only entries, guarded balances, unique external IDs — is boring and correct; I'd build it, roughly two engineers for two quarters plus ongoing ownership. A ledger-as-a-service is attractive if we're early and our flows are standard; the risk is that the chart of accounts and posting rules — the parts that encode our business — end up shaped by the vendor's model, and migrating off a system of record is a multi-quarter project. A purpose-built ledger database earns its place when sustained throughput reaches tens of thousands of TPS or hot-account contention can't be designed away.

My default: Postgres now, with posting rules and the API designed so the storage could change behind them."

Why this is L6:

  • Gives thresholds for each option, not a preference
  • Names the real lock-in: chart of accounts and posting rules, not storage
  • Keeps the door open with an API boundary

What L7 adds:

  • Prices the system-of-record migration explicitly before signing a vendor
  • Requires data export and audit-evidence terms in any vendor contract
❌ Common L5 Trap

"Always build — money is too important to outsource."

Why this misses: It's a values statement, not an analysis. Many correct ledgers run on vendors; the question is cost, control and exit, and the candidate hasn't priced any of them.


Drill 8: Changing a Posting Rule Without an Outage#

Prompt: "Finance wants the FX fee booked to a new revenue account starting next month. How do you ship it?"

Staff Answer

"Posting rules are versioned config, not code. I'd write fx_convert.v5 with the new account, and finance reviews the diff in the rule repository. Before activation, I shadow-run v5 against a day of production events in a sandbox and diff the entries against v4: only the fee leg's account should change, every transaction must still balance per currency. Activation is by effective_at — 00:00 UTC on the first — and every transaction records rule_version, so reports can split before and after. Rollback is activating v4 again; transactions already posted under v5 stay as they are, because entries are immutable, and if needed finance reclassifies with explicit transfer entries."

Why this is L6:

  • Treats posting rules as reviewed, versioned configuration
  • Shadow-diffs before activation and states the invariant the diff must preserve
  • Rollback doesn't pretend to rewrite history

What L7 adds:

  • Makes finance the approver of record for rule changes — a control, not a courtesy
  • Runs the shadow diff automatically on every rule PR
❌ Common L5 Trap

"Change the account in the code and deploy on the first of the month."

Why this misses: No review by the people who own the books, no proof the new rule balances, and a deploy-time cutover that splits transactions mid-flight across versions without recording which one applied.


Drill 9: Sharding the Ledger#

Prompt: "We're at 20,000 TPS. Shard it. What happens to a transfer between two shards?"

Staff Answer

"Shard by account ID, with house accounts split per shard so each shard has its own fee and clearing accounts. A same-shard transfer stays one local transaction. A cross-shard transfer from A (shard 1) to B (shard 2) becomes two local transactions linked by one transfer ID: on shard 1, debit A, credit interShard:1to2 — guarded and committed; on shard 2, debit interShard:2from1, credit B. Each step is idempotent by transfer ID. A transfer coordinator — a durable workflow — drives step 2 and retries until it commits; B's money is 'in flight' in the meantime, visible as a non-zero inter-shard clearing balance.

The invariant: interShard:1to2 on shard 1 and interShard:2from1 on shard 2 net to zero once every transfer completes. Aged non-zero balances are breaks. I don't use 2PC across shards — it holds locks across a network round trip on exactly the accounts that are busiest."

Why this is L6:

  • Splits house accounts per shard so sharding actually removes contention
  • Replaces distributed atomicity with clearing accounts plus idempotent steps
  • Names the invariant that proves cross-shard correctness

What L7 adds:

  • Notes that 20K TPS should trigger the build-vs-buy question for a purpose-built engine before a sharding project
  • Plans shard placement by legal entity and region so sharding and data residency align
❌ Common L5 Trap

"Use a distributed transaction across the two shards."

Why this misses: It's not wrong, but it couples the availability of both shards, holds locks across network round trips, and doesn't scale on hot accounts. The candidate hasn't considered that money can be safely "in flight" in an account built for that purpose.


Drill 10: Multi-Region#

Prompt: "We're launching in the EU. Make the ledger multi-region."

Staff Answer

"First question is legal, not technical: EU customer funds will likely be held by an EU entity and safeguarded at an EU bank, so the natural split is a ledger per legal entity, running in its region — not one global active-active ledger. Each account has a home region; only that region writes its balance. Cross-entity transfers — a US user paying an EU user — are two ledgers, two transactions, linked by an inter-entity clearing account on each side, settled between entities by treasury.

Within a region, synchronous replication across availability zones for the ledger database; failover ≤ 30s. Cross-region replicas exist for disaster recovery and reporting, not for writes. A consolidated reporting layer in the warehouse joins all entity ledgers for group finance. Active-active writes on balances would require conflict resolution on money — I won't do that."

Why this is L6:

  • Leads with the legal-entity split, which determines the architecture
  • Home region per account; no multi-writer balances
  • Cross-entity flows modeled with clearing accounts, settled by treasury

What L7 adds:

  • Builds entity into the account schema from year one so the split isn't a re-posting project
  • Prices a second entity: separate safeguarding, audits, regulatory reporting — often more than the engineering
❌ Common L5 Trap

"Deploy the ledger active-active in two regions with a globally distributed database."

Why this misses: It answers latency and ignores that the regions are probably different legal entities with different safeguarding obligations. And multi-region consensus on every post adds ~70–150ms to the card authorization path.


8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — Card Declines During a Payroll Batch#

Context: It's the first Friday of the month. Card decline rates jump from 2% to 19% at 09:02. The card processor's stand-in is approving small transactions only. The on-call escalates to you.

Questions to Surface First:

  • Are declines "insufficient funds" (ledger said no) or timeouts (ledger didn't answer)?
  • What else is writing to the ledger right now? Any bulk producer?
  • Which accounts have the highest lock wait?
  • Is the card auth path sharing a connection pool or database with bulk work?

Typical L5 Approach: Scales up the ledger service pods and the database instance. Lock waits don't change, because the contention is on one row, not on CPU.

Staff Approach: Reads ledger.top_accounts_by_lock_wait, finds funding_pool at the top, correlates with the payroll producer's ramp, and throttles the producer into a low-priority lane. Card auths recover in minutes. Then fixes the design: batch-netted payroll postings and a split funding pool.

Principal Approach: Treats it as a missing isolation contract between bulk and interactive money flows. Establishes priority lanes in the ledger platform — interactive card auths get reserved capacity — and requires every bulk producer to use the batch API with a negotiated rate.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Classify declines: 85% are ledger timeouts. Lock-wait leaderboard shows funding_pool. Throttle payroll producer to 200/s via its rate-limit config.
TriagePayroll posts one transaction per salary, each debiting the single floored funding_pool row. Card auths share the DB and wait behind them.
Quick fixSeparate connection pools: interactive (card, P2P) and bulk. Bulk pool capped at 20% of connections.
GuardrailsAlert on ledger.lock_wait_p99{account} > 50ms; card auth ledger timeout rate > 1% pages.
Post-mortemWhy could one producer monopolize a floored account? Why did the card path share a pool with batch work? Add batch-posting API: one funding debit nets 1,000 salary credits.

Metrics to Watch: card.auth_ledger_timeout_rate, ledger.lock_wait_p99{account}, ledger.post_rate{producer}, payroll.backlog_count

Organizational Follow-up: payroll integration team signs a producer contract: max rate, batch API, scheduled windows that avoid card peaks.

Ownership Question: "Who decides which producer gets ledger capacity at peak?" Staff answer: The ledger platform owns capacity and priority lanes; producers negotiate rates with it. Card auths are always top priority because the network won't wait.

Key Takeaway: "Ledger capacity is per account, not per cluster. One floored hot account can turn a batch job into a card outage."

What clears the Staff bar:

  • Distinguishes timeouts from real insufficient-funds declines before acting
  • Finds the hot account rather than scaling hardware
  • Separates interactive and bulk paths permanently

Deep Dive 2: Silent Failure — Balances Drifted for Nine Days#

Context: A routine audit sample finds a wallet whose balance row is $12 higher than the sum of its entries. The nightly drift job shows green. You're asked to investigate.

Questions to Surface First:

  • When did the drift job last actually check every account it should?
  • Which code paths can write balance rows? Any outside the ledger service?
  • Is the drift limited to accounts touched by one producer or one rule version?
  • How many accounts are affected, and is any user able to spend money that doesn't exist?

Typical L5 Approach: Fixes the one account by updating its balance to match entries. Assumes it's isolated.

Staff Approach: Discovers the drift job was silently scoped to "accounts with entries today" and a migration script nine days ago updated balances directly without entries. Re-runs the full drift scan across all accounts, finds 4,100 drifted, freezes the migration path, and corrects with explained entries.

Principal Approach: Revokes all direct table write grants outside the ledger service — including for migrations — and makes migrations run through the ledger API with a dedicated producer ID. Drift checks become a controlled audit test with coverage metrics.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Full-scan drift check across all accounts (not just those touched today). Quantify: 4,100 accounts, net +$31K over entries.
TriageAll drifted accounts were touched by a data migration nine days ago that "fixed" balances with a direct UPDATE.
Quick fixRevoke the migration role's write grant. For each account, decide with finance: post entries explaining the migration's intent, or revert the balance to Σ entries.
GuardrailsDrift job reports accounts_checked vs accounts_expected; alert if coverage < 100%. Weekly full scan in addition to nightly incremental.
Post-mortemWhy could anything write balances without entries? Why did the drift job's scope silently exclude untouched accounts?

Metrics to Watch: ledger.balance_drift_accounts, ledger.drift_check_coverage_pct, ledger.direct_writes_total (from DB audit log)

Organizational Follow-up: migrations touching money require a ledger-platform reviewer and must post entries via the API.

Ownership Question: "Who owns the drift check's coverage?" Staff answer: The ledger platform owns the check and its coverage metric; finance audits its evidence monthly.

Key Takeaway: "A check that reports green without reporting what it checked is not a control."

What clears the Staff bar:

  • Questions the monitor before trusting it
  • Finds the bypass path, not just the bad rows
  • Corrects with explained entries, never silent updates

Deep Dive 3: Large-Customer Onboarding — A Marketplace With 40,000 Payouts a Minute#

Context: A large marketplace wants to use the wallet as the payout destination for 2 million sellers, settling every minute. Sales says yes. Engineering is asked whether the ledger can take it.

Questions to Surface First:

  • Which account is debited for each payout? One marketplace funding account?
  • Is the marketplace's funding account floored?
  • Do sellers need per-payout visibility in real time, or per-batch?
  • What's the failure behavior if the marketplace's account runs dry mid-batch?

Typical L5 Approach: "40,000/minute is ~670/s — well within our capacity." True for the cluster; false for the single marketplace funding row that every payout debits.

Staff Approach: Models the batch as one transaction per chunk of 1,000 payouts: one guarded debit on the marketplace funding account for the chunk total, 1,000 credits to seller wallets. 40 transactions a minute on the hot row instead of 40,000. Partial failure is per chunk; chunks are idempotent by batch_id:chunk_no.

Principal Approach: Turns it into a platform capability — a batch-posting API with chunk idempotency and funding pre-checks — priced into the marketplace contract, with a dedicated bulk lane so one customer's payouts can't affect card auths.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (design review)Identify the marketplace funding account as floored and hot (670 debits/s).
TriageSingle-row ceiling ~200–500/s with realistic txn time. Not viable per payout.
Quick fixBatch API: chunk of up to 1,000 payouts → one transaction; guard on chunk total.
GuardrailsPre-flight funding check per batch; reject batch if funding < total; per-customer bulk lane with its own pool.
Post-mortem (pre-mortem)Load test at 3× expected volume with card auth traffic running concurrently.

Metrics to Watch: ledger.batch_chunk_latency_p99, ledger.lock_wait_p99{account=mkt_funding}, payout.batch_partial_failures

Organizational Follow-up: contract specifies batch windows and max rates; marketplace receives chunk-level status via webhooks.

Ownership Question: "Who owns a half-completed batch?" Staff answer: The ledger guarantees each chunk is atomic and idempotent; the payouts service owns driving the batch to completion and reporting chunk status to the customer.

Key Takeaway: "Big customers don't break the cluster; they break one account. Batch the hot side, keep the cold side per-entry."

What clears the Staff bar:

  • Identifies the per-account ceiling, not the cluster average
  • Designs chunked, idempotent batch posting
  • Protects interactive traffic from the new bulk load

Deep Dive 4: Post-Mortem — $57,900 Safeguarding Shortfall#

Context: The daily safeguarding reconciliation shows customer liabilities exceed the bank balance by $57,900. Compliance has been notified. You're leading the post-mortem.

Questions to Surface First:

  • When was the last clean reconciliation?
  • Are the breaks concentrated in one rail, one producer, one rule version?
  • Were affected funds spent by users?
  • Did any integration change ship between the clean day and the break?

Typical L5 Approach: Finds the mapping bug in the ACH integration, fixes it, reverses the bad credits. Closes the ticket.

Staff Approach: Fixes the bug and asks why a reversible rail could credit spendable balance before its return window passed, and why the status mapping could change without ledger review. Introduces a hold on top-ups from reversible rails until settlement plus the return window.

Principal Approach: Treats it as a control design failure in the funds-flow model. Defines rail risk classes (instant final, reversible for N days) and a company-wide rule that spendability depends on rail class, with risk sign-off. Adds intraday safeguarding reconciliation for reversible rails.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Company funds the gap from operating cash. Freeze the ACH status mapping deployment.
Triage230 top-ups with return code R01 mapped to "settled"; users credited; 140 already spent.
Quick fixRevert mapping. Post reversing entries for unspent; spent ones go negative → collections.
GuardrailsAlert on topup.return_rate > 2× baseline; ACH top-ups credited with a hold for the return window.
Post-mortemRail status mappings require ledger-platform review; contract test against the bank's return-code list.

Metrics to Watch: safeguarding.shortfall_usd, topup.return_rate{rail}, wallet.negative_balance_total

Organizational Follow-up: compliance documents the incident and remediation for the regulator; treasury reviews funding buffer.

Ownership Question: "Who decides when top-up money becomes spendable?" Staff answer: Risk, per rail class, signed with finance. The ledger enforces it with holds; product doesn't choose.

Key Takeaway: "Money you credited before it was final is a loan you didn't approve."

What clears the Staff bar:

  • Funds the shortfall before root cause
  • Fixes the policy gap (spendability by rail finality), not just the bug
  • Makes integration mappings a reviewed contract

Deep Dive 5: Multi-Region Expansion — Launching EUR Wallets in the EU#

Context: The company is launching in the EU with an EU-licensed entity. Product wants US users to send money to EU users instantly. The CTO asks for the ledger plan.

Questions to Surface First:

  • Which entity holds EU users' funds, and where are they safeguarded?
  • Must EU customer data and ledger records stay in the EU?
  • Is cross-border P2P a currency conversion, and who bears FX risk?
  • What's the settlement path between the US and EU entities?

Typical L5 Approach: Deploys the existing ledger to an EU region and replicates both ways so either region can post.

Staff Approach: One ledger per legal entity, each in its region, each with its own safeguarding reconciliation. Cross-entity P2P: the US ledger debits the US user and credits interEntity:EU; the EU ledger debits interEntity:US and credits the EU user; treasury settles the inter-entity position daily. FX happens on one side with the rate recorded.

Principal Approach: Establishes the entity-per-ledger model as the company's pattern for every future market, with a group consolidation layer, and prices each new market's fixed costs (safeguarding bank, audits, reporting) before engineering starts.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (planning)Confirm entity and safeguarding structure with legal and treasury.
TriageData residency: EU ledger and entries stay in EU region; US sees only inter-entity positions.
Quick fixInter-entity clearing accounts on both ledgers; transfer workflow idempotent by transfer ID.
GuardrailsinterEntity positions must match across ledgers after settlement; daily check.
Post-mortem (pre-launch)Game day: EU region down — US users see EU-bound transfers pending, not failed.

Metrics to Watch: interentity.position_mismatch, transfer.cross_entity_inflight_age_p99, safeguarding.shortfall_eur

Organizational Follow-up: treasury owns inter-entity settlement; group finance owns consolidation.

Ownership Question: "Who owns a cross-border transfer stuck between ledgers?" Staff answer: The transfer workflow (payments platform) owns driving it to completion; both ledgers show it as in-flight in their inter-entity accounts.

Key Takeaway: "In money systems, regions are usually legal entities. Draw the ledger boundary where the balance sheet boundary is."

What clears the Staff bar:

  • Leads with entity and safeguarding, not replication topology
  • Uses clearing accounts and idempotent steps across ledgers
  • No multi-writer balances

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • State the two ledger invariants — conservation and non-negativity — and explain why they're enforced differently
  • Write a guarded balance update that makes the overdraft check and the debit a single statement
  • Model holds as a pending dimension with exactly-once capture, void and expiry, including late captures
  • Identify hot accounts and choose between async aggregation, sub-accounts and batch netting using the "does it have a floor?" test
  • Design producer idempotency with key and payload-hash semantics and permanent retention
  • Post an FX conversion as per-currency balanced legs with the rate and rounding recorded
  • Run a safeguarding reconciliation with timing differences, break types, owners and a same-day escalation path
  • Shard a ledger with per-shard house accounts and cross-shard clearing instead of distributed transactions

The Bar for This Question#

Mid-level (L4): Builds wallets with a balance column and a transactions table, updated together. Handles the happy path. Doesn't consider concurrent spends, holds or reconciliation. Uses decimals or floats.

Senior (L5): Adds double-entry, row locks, idempotency keys and an event bus. Knows reconciliation exists. The gap: debits on auth instead of holding; no awareness that house accounts serialize everything; treats internal balance as proof the money exists; shards by user ID. The design passes a code review and fails its first payday.

Staff+ (L6): Frames the problem around invariants and who holds the funds. Puts the guard in the write. Holds are first-class. Finds the hot accounts before the interviewer does. Separates "consistent" from "exists" and owns safeguarding reconciliation with a same-day escalation. Names who pays — users for long holds, the business for late-capture write-offs, finance for unclear posting rules. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 "The Balance Column Is a Cache, and You Should Treat It Like One"#

ClaimReality
"The balance is the source of truth"It's a derived value that must be provably equal to Σ entries
"We'll fix drift when we see it"Drift without a daily full-coverage check is invisible until an audit
"Updating the balance is simpler"Until someone asks why it's $4,212.50

The Staff position: Keep the balance row — you need it for the guard — and verify it against entries every day with a coverage metric.

Why this matters in interviews: Calling the balance a cache instantly signals you understand where truth lives.

10.2 Most Hot-Account Problems Are Self-Inflicted#

AccountNeeds a Synchronous Balance?
User walletYes — it has a floor
Fee revenueNo — nothing checks it
Card clearingNo — reconciled, not guarded
Omnibus fundingYes — split it

The Staff position: Before sharding anything, stop maintaining synchronous balances on accounts that have no floor.

Why this matters in interviews: It shows you understand why a lock exists, not just that it hurts.

10.3 Event Sourcing Is Not a Ledger#

Event LogLedger
Records what happenedRecords what happened and where the value went
Replay gives stateReplay gives balanced state, or the rule is wrong
No conservation invariantEvery transaction sums to zero

The Staff position: Feed the ledger from events if you like, but double-entry posting rules are what make it a ledger.

Why this matters in interviews: Candidates who say "event sourcing" and stop have skipped the invariant.

10.4 Holds Are Where Wallets Actually Break#

Bug ClassRoot Cause
Money "missing" without a transactionStuck sweeper, early debit model
Negative balancesLate capture after expiry
Double statement linesDebit-and-reverse instead of pending

The Staff position: A wallet without a pending dimension is a support queue in waiting.

Why this matters in interviews: Interviewers who have run wallets probe holds first.

10.5 Reconciliation Against the Bank Is the Only Proof That Matters#

The Staff position: A perfectly balanced ledger can still owe users money the company doesn't have. Internal consistency is necessary; external reconciliation is sufficient. If the safeguarding reconciliation has no named owner and no same-day escalation path, the ledger is decorative.

Why this matters in interviews: Ending on safeguarding ownership is the most Staff-sounding close available for this question.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer designs a correct wallet ledger. The Principal engineer notices that the company already has four ledgers — the wallet's, the payouts team's, billing's revenue tables, and the finance team's general ledger fed by spreadsheets — each with its own chart of accounts and its own definition of "balance". The L7 problem is one system of record for money across the company: a ledger platform with versioned posting rules owned by finance, a contract for every producer, entity-aware accounts, and controls an auditor can test.

The Org-Level Fault Line#

One company ledger platform vs per-product ledgers.

OptionWhat WorksWhat BreaksWho Pays
Per-product ledgersProduct teams move fast; local modelsN charts of accounts; finance reconciles ledgers against each other; inconsistent controlsFinance (close time), auditors (sampling), engineering (N on-call rotations)
One central ledger owned by a platform teamOne truth; one set of controlsPlatform bottleneck; every new flow waits for posting-rule workProduct teams (velocity)
Platform owns ledger + rule engine; products author rules with finance reviewCentral invariants, distributed authorshipNeeds strong review tooling and rule testingPlatform (tooling), finance (reviews)

🧭 Principal Move: "The platform owns the invariants — balancing, guards, idempotency, immutability, drift checks. Products author posting rules in a shared repository, finance approves them, and automated shadow diffs prove they balance. Nobody writes entries any other way."

Cost Model#

Assumptions: fully loaded engineer ~$250K/year; managed relational database; one region per entity; storage at ~250 bytes per entry including indexes.

ScaleVolumeInfra ($/month)HeadcountOn-call LoadNotes
Startup1M wallets, ~10 TPS~$2–4K (HA Postgres, small Kafka or none)2 eng part-time + 1 finance opsShared rotation, < 1 page/monthSafeguarding recon semi-manual
Growth10M wallets, ~100–1,000 TPS~$15–30K (large HA primary, replicas, warehouse)5–7 eng (ledger 3, recon 2, rails 2) + 3 finance/treasury opsDedicated ledger rotation, 2–4 pages/monthHot-account design matters now
Enterprise100M wallets, 5–20K TPS, 3 entities~$120–250K (sharded or purpose-built engine, per-entity stacks)25–40 eng + 15–20 finance/treasury/compliancePer-entity rotations, follow-the-sunEach new entity adds fixed compliance cost

The pricing insight: at growth scale the ledger's cost is dominated by people — the reconciliation and treasury operations headcount often exceeds engineering. Automating break classification that clears 80% of breaks without a human pays for an engineer within a year.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Chart of accounts and account identity (incl. legal entity)One-wayRe-posting history; audit re-certification; 6–12 months
Money representation (integer minor units + ISO currency)One-wayEvery consumer and every historical row
Entries immutable, corrections by reversalOne-wayAbandoning it destroys audit trail retroactively
Who holds customer funds (us vs partner bank)One-wayLicensing, banking contracts, migration of balances
Storage engine behind the ledger APITwo-way (if the API is clean)1–3 quarters of dual-write and verification
Hold TTL policyTwo-wayPolicy change with risk sign-off
Hot-account strategy per accountTwo-wayMigration of one account's balance
Reconciliation cadence (daily vs intraday)Two-wayOperations and tooling investment

The Standard I'd Write#

RFC-LEDGER-001: Ledger of Record Standard
Status: Approved   Owners: Ledger Platform + Office of the Controller

Scope
  Every system that records, moves or reports balances of money or
  money-equivalent stored value.

MUST
  1. Write financial effects only through the Ledger API; no direct table writes,
     including migrations.
  2. Submit transactions that balance per currency; the API rejects others.
  3. Supply a producer-unique external_id; replays must be no-ops.
  4. Debit floored accounts only through guarded posting rules.
  5. Represent amounts as integers in ISO 4217 minor units with explicit currency.
  6. Correct errors with reversing entries; never update or delete entries.
  7. Post no entries into a closed period; use effective dating for corrections.

SHOULD
  1. Use holds for any reservation that precedes settlement.
  2. Declare a contention class for every new house account.
  3. Route every external flow through a clearing account with a stated clearing window.

Exceptions
  Filed with Ledger Platform; controller sign-off for any exception to MUST 1, 2 or 6;
  time-boxed to one quarter.

Success metrics
  - Balance drift accounts: 0, with 100% daily coverage
  - Safeguarding shortfall detected and funded within 1 business day: 100%
  - Clearing accounts aged beyond window: < 0.01% of volume
  - Month-end close: ≤ 3 business days

What I'd Tell the VP#

"Our wallet ledger is correct, but the company keeps money records in four places, and finance spends a week each month making them agree. I'm proposing we make the wallet ledger the company's system of record for money: products keep building their features, but every financial effect goes through one ledger with rules finance approves. It costs about four engineers for a year, cuts close to three days, and gives auditors one control set to test instead of four. The main risk is migration — I'd move one product per quarter, starting with payouts, and keep the existing wallet path untouched."

Principal Interview Signals#

SignalWhat It Sounds Like
Prices operations, not just infra"Reconciliation headcount outweighs infra 3 to 1 at our scale; automating break classification pays back in a year."
Identifies one-way doors"Account identity with legal entity is the decision I'd slow down on."
Redraws ownership"Platform owns invariants; products author rules; the controller approves them."
Designs controls, not just checks"The drift check is audit evidence — retained, with a coverage metric."
Knows when not to centralize"Loyalty points get the ledger pattern but not the safeguarding machinery until they're redeemable for cash."

Staff answers that L7 interviewers find insufficient:

  • "We'll build a correct wallet ledger" — correct, but ignores the other three ledgers in the company.
  • "Finance handles reconciliation breaks" — names the owner but not the control design or the regulatory escalation.
  • "We'll shard when we need to" — no trigger, no consideration of legal-entity boundaries as the natural shard key.

Appendices

Appendix A: Mechanics in Depth#

A.1 The Guarded Post#

def post(producer, external_id, rule, params):
    entries = rules[rule.name][rule.version].apply(params)      # pure function
    assert_balanced_per_currency(entries)                       # else 422
    req_hash = canonical_hash(rule, entries)
    with db.tx() as tx:
        existing = tx.insert_txn_or_get(producer, external_id, req_hash)
        if existing:
            if existing.req_hash != req_hash: raise Conflict409
            return existing.txn_id                              # replay
        for acct, delta in net_by_account(entries):
            if acct.has_floor and delta < 0:
                rows = tx.exec("""UPDATE balances
                                  SET posted = posted + %s, available = available + %s,
                                      version = version + 1
                                  WHERE account_id = %s AND available + %s >= floor""",
                               delta, delta, acct.id, delta)
                if rows == 0: raise InsufficientAvailable402(acct.id)
            elif acct.contention_class == "async":
                pass                                            # aggregator updates later
            else:
                tx.exec("UPDATE balances SET posted = posted + %s, available = available + %s ...")
        tx.insert_entries(entries)
        tx.insert_outbox("ledger.txn.posted", txn_id)
    return txn_id

Lock ordering: update balance rows in a deterministic order (sorted by account_id) to avoid deadlocks when two transfers touch the same pair of accounts in opposite directions.

A.2 Holds#

create_hold(acct, amt, ttl):   UPDATE balances SET pending_debits += amt, available -= amt
                               WHERE account_id = acct AND available - amt >= floor
                               INSERT holds(state='PENDING', expires_at = now() + ttl)
post_hold(h, amt_captured):    UPDATE holds SET state='POSTED' WHERE id=h AND state='PENDING'
                               post entries for amt_captured
                               UPDATE balances SET pending_debits -= h.amount,
                                      posted -= amt_captured,
                                      available += (h.amount - amt_captured)
void_hold(h):                  UPDATE holds SET state='VOIDED' WHERE id=h AND state='PENDING'
                               UPDATE balances SET pending_debits -= h.amount, available += h.amount
sweeper (every 30s):           holds WHERE state='PENDING' AND expires_at < now() LIMIT 1000
                               → transition to EXPIRED, release, rate-limited

A.3 Why Not Debit-and-Reverse?#

It works mechanically and fails operationally: every authorization that doesn't capture exactly produces at least two posted entries that users and finance must interpret; expiry is a separate job that, when it fails, leaves posted balances wrong rather than merely untidy; and revenue/refund reporting must filter out reversal noise. Holds keep posted balances clean by construction.

Appendix B: Data Model#

CREATE TABLE accounts (
  account_id       TEXT PRIMARY KEY,           -- e.g. wallet:u42:USD, fee_revenue:USD:shard07
  legal_entity     TEXT NOT NULL,
  owner_type       TEXT NOT NULL,              -- user | merchant | house
  currency         CHAR(3) NOT NULL,
  account_type     TEXT NOT NULL,              -- asset | liability | revenue | expense | equity
  floor_minor      BIGINT,                     -- NULL = no floor (not guarded)
  contention_class TEXT NOT NULL DEFAULT 'row' -- row | split | async
);

CREATE TABLE transactions (
  txn_id        UUID PRIMARY KEY,
  producer      TEXT NOT NULL,
  external_id   TEXT NOT NULL,
  request_hash  BYTEA NOT NULL,
  rule_name     TEXT NOT NULL,
  rule_version  INT NOT NULL,
  effective_at  TIMESTAMPTZ NOT NULL,
  posted_at     TIMESTAMPTZ NOT NULL DEFAULT now(),
  UNIQUE (producer, external_id)
);

CREATE TABLE entries (
  entry_id      BIGINT GENERATED ALWAYS AS IDENTITY,
  txn_id        UUID NOT NULL REFERENCES transactions,
  account_id    TEXT NOT NULL REFERENCES accounts,
  direction     CHAR(1) NOT NULL CHECK (direction IN ('D','C')),
  amount_minor  BIGINT NOT NULL CHECK (amount_minor > 0),
  currency      CHAR(3) NOT NULL,
  PRIMARY KEY (account_id, entry_id)           -- clustered by account for statements
);

CREATE TABLE balances (
  account_id     TEXT PRIMARY KEY REFERENCES accounts,
  posted         BIGINT NOT NULL DEFAULT 0,
  pending_debits BIGINT NOT NULL DEFAULT 0,
  available      BIGINT NOT NULL DEFAULT 0,    -- maintained = posted - pending_debits
  version        BIGINT NOT NULL DEFAULT 0
);
-- Application role: INSERT on transactions/entries; UPDATE on balances/holds only via ledger service.
-- No UPDATE or DELETE grant on entries or transactions for any role.

Appendix C: Coordination Mechanisms#

C.1 Cross-Shard Transfer via Clearing Accounts#

Diagram: C.1 Cross-Shard Transfer via Clearing Accounts

C.2 Quick Comparison#

MechanismGuaranteesFailure ModeUse For
Balanced-transaction checkConservationBypass via direct writesEvery post
Guarded conditional UPDATENon-negativityCheck outside the txnFloored accounts
(producer, external_id) uniqueEach business event posts onceCoarse keysRetries, replays
Conditional hold transitionsEach hold resolves onceBlind updatesCapture, void, expiry
Sub-accounts + rebalancerThroughput on floored hot accountsDry shardOmnibus, funding
Async aggregationNo contentionLagging balanceUnfloored house accounts
Clearing accountsDetects missing/late legsUnowned aged balancesExternal flows, cross-shard
Drift checkBalance = Σ entriesPartial coverageDaily control
Safeguarding reconciliationMoney existsUnowned breaksDaily, regulatory

Appendix D: API Contract & Client Behavior#

  • Every mutating call requires external_id (ledger API) or Idempotency-Key (public API); the wallet service derives the ledger key deterministically from the public key.
  • Replay with the same payload returns the original result; different payload returns 409.
  • 402 insufficient_available includes current available so the client can show a precise message.
  • Balance reads return posted, pending, available, as_of_entry_id; statements paginate by (account_id, entry_id) cursor.
  • Historical balance: GET /balance?as_of= computes from the nearest daily snapshot plus entries — never from the live balance row.
  • Ledger events carry txn_id as the event ID; consumers dedupe on it.

Appendix E: Observability#

Core metrics:

  • ledger.post_p99, ledger.post_rate{producer}, ledger.lock_wait_p99{account}
  • ledger.unbalanced_rejects (must stay 0 in steady state), ledger.key_conflicts{producer}
  • ledger.balance_drift_accounts, ledger.drift_check_coverage_pct
  • ledger.accounts_below_floor (excluding flagged force-posts)
  • hold.pending_total_usd, hold.expired_backlog_age, hold.sweeper_last_success_age
  • clearing.aged_balance{account, bucket}, safeguarding.shortfall{currency}

Critical alerts:

AlertThresholdSeverity
Balance drift> 0 accountsSev-1 page
Accounts below floor (unflagged)> 0Sev-1 page
Safeguarding shortfall> 0 after timing adjustmentsPage treasury + compliance
Hold expiry backlog age> 1hPage
Ledger post p99> 100ms for 3 minPage
Lock wait on any account> 50ms p99 for 2 minPage
Drift check coverage< 100%Page

Debugging the silent failure: the most dangerous ledger failures produce no errors — a bypass path that writes balances without entries, a sweeper that stops, a rail mapping that credits returned money. Alert on absence (sweeper last success, drift coverage) and on external disagreement (safeguarding, clearing ages), not only on error rates.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 100 TPSOne Postgres primary; all accounts single-rowNothing — resist complexity
100–2,000 TPSAsync aggregation for unfloored house accounts; bulk lanes; batch APIFloored hot accounts; drift job runtime
2K–20K TPSSplit floored accounts; partition entries by account; per-entity ledgersSingle-primary write ceiling
> 20K TPSShard by account with per-shard house accounts, or a purpose-built ledger engineCross-shard reporting, operations headcount

What you don't build on day one: sharding, a purpose-built ledger engine, multi-region writes, intraday safeguarding reconciliation, a general-purpose posting-rule DSL. Each has a trigger in Section 11.

Appendix G: Multi-Tenancy, Fairness & Cost#

  • Producer lanes: interactive flows (card auths, P2P) get reserved connection capacity; bulk producers (payroll, payouts) use a separate pool capped at ~20% and a batch API.
  • Per-producer rate contracts: each producer has a max post rate; exceeding it returns 429 rather than degrading card auths.
  • Tenant isolation for platform customers: marketplaces using the wallet get their own funding and clearing accounts per tenant so one tenant's shortfall or batch can't consume another's capacity.
  • Cost attribution: ledger storage and post volume tagged by producer; teams that post six entries per transaction where two would do see the cost.
  • Retention cost: at 250 bytes per entry and 1B entries/year, ~250GB/year in the hot store; move closed years to cheaper storage with the same query interface, never delete within the retention period.
  1. Loading the index…