Hiring BarSupport

Design a Stock Exchange — Staff-Level Case Study

Case study66 min read7 diagrams

Technologies referenced in this case study: Kafka · PostgreSQL · ZooKeeper & etcd · Time Series DBs

Related: Distributed Consensus · Payment Processing · Dealing with Contention · Real-time Updates · Network Latency & Protocols · Replicated Data Store

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once, then return to individual sections before a loop.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 2, 4
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failures) → Deep Dives 1 and 4
Deep Dive3+ hrsEverything, including The Principal Lens and appendices (order book structures, sequencer, replay)
What is a Stock Exchange? — Why interviewers pick this topic

An exchange accepts buy and sell orders from members (brokers, market makers), matches them according to published rules — usually price-time priority — and publishes the resulting trades and order book changes to the market. Around that core sit pre-trade risk checks, market data distribution, clearing and settlement hand-off, surveillance, and regulatory reporting.

The core is small. A matching engine for one symbol is a few thousand lines of code. The difficulty is that it must be deterministic, fair, auditable, and fast at the same time, and every one of those properties can be broken by an innocent-looking optimization.

Before vs After — the "ghost fill" incident:

Without a sequencer and deterministic replay:
t=0:        Two matching threads share a symbol's book for throughput
t=+2ms:     Buy 500 @ 100.05 and Sell 500 @ 100.05 arrive 3µs apart on different gateways
t=+2ms:     Thread A matches the buy against a resting sell; thread B matches the same
            resting sell against another order. Lock ordering bug: 500 shares filled twice.
t=+3ms:     Two execution reports go out. Market data prints both trades.
t=+40min:   Primary crashes. Backup rebuilds from logs in a different interleaving;
            its book disagrees with what members were told.
t=+1 day:   Trades busted. Regulator inquiry. Members demand compensation.

With a sequencer + single-threaded deterministic engine:
t=0:        Every inbound message gets a global sequence number and is journaled
t=+2ms:     Engine for the symbol processes them strictly in sequence order, one at a time
t=+2ms:     Exactly one fill. Execution reports and market data carry the sequence number.
t=+40min:   Primary crashes. Hot standby has processed the identical sequence;
            its book hash matches. Failover in < 1s. No bust, no inquiry.

Why interviewers reach for this question: It inverts the usual system-design instincts. Horizontal scaling, eventual consistency, and "just add a queue" are all wrong here. It tests whether a candidate recognizes a problem where a single thread is the correct architecture and can defend that choice under pressure.

Mechanics Refresher: Matching and Order Types
ConceptHow It WorksProsCons
Price-time priority (FIFO)Best price first; at equal price, earliest order firstRewards being first; simple, transparentArms race on latency
Pro-rataAt equal price, fills allocated proportionally to sizeRewards size/liquidity; common in some futures marketsIncentivizes oversized orders
Continuous matchingEach incoming order matches immediately against the bookInstant executionLatency advantage matters most
Call / batch auctionOrders collected, matched at one clearing price at a point in timeNeutralizes microsecond races; used at open/closeExecution delayed until the auction
Limit orderBuy/sell at a price or better; rests if unmatchedPrice controlMay never fill
Market orderExecute immediately at best available pricesCertain executionPrice uncertainty; needs collars
IOC / FOKImmediate-or-cancel / fill-or-killNo resting riskOften zero fill

For most production exchanges: continuous price-time priority for the trading day, with opening and closing auctions, one single-threaded matching engine per symbol partition, and a sequencer in front establishing a total order of inbound events.


Executive Summary

If you only read one section, read this. Every decision in the case study follows from the contrast and the fault lines below.

What This Interview Actually Tests#

A stock exchange is not a throughput question. Matching 1 million orders a second on one core was demonstrated publicly by LMAX more than a decade ago.

It is a determinism and fairness question that tests:

  • Whether you can defend a single-threaded design against the reflex to parallelize
  • Whether you establish one total order of events and make everything downstream a pure function of it
  • Whether recovery is "replay the journal and get the identical state" rather than "restore from a database"
  • Whether you treat fairness — equal access, equal information, provable ordering — as the product
  • Whether you know when the right move is to halt trading

The key insight: An exchange is a replicated state machine whose input log is legally significant. If two replicas given the same log can ever disagree, you don't have an exchange; you have a liability.

The L5 vs L6 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws API → queue → matching service → DBAsks about fairness model, latency target, and regulatory audit; commits to sequencer + deterministic engineAsks which market structure the business is entering (lit, dark, crypto, auction) — the one-way door that dictates everything else
ConcurrencyHorizontal matching workers with locks on the bookOne thread per symbol partition; no locks; parallelism across symbols onlySets the determinism boundary as an org rule: nothing nondeterministic may cross into the engine, ever
DurabilityWrite each trade to a databaseJournal inbound sequenced messages before processing; state = f(journal); snapshots for fast replayTreats the journal as the legal record: retention (7 years), tamper evidence, regulator access
FailoverRestart the service, reload from DBHot standby replaying the same sequence; state-hash comparison; sub-second takeover or controlled haltDesigns the market's failure posture: when to halt, how to reopen via auction, how members are told
RiskValidate orders in the matching servicePre-trade risk in the gateway path with µs budget; kill switches per member; price collarsOwns the market-wide controls: circuit breakers, limit-up/limit-down participation, fat-finger policy as rulebook
Market dataWebSocket push to clientsUDP multicast with sequence numbers and gap recovery; equal dissemination to allBalances fairness vs revenue: co-location, data product tiers, regulatory fair-access obligations
Why "concurrency" separates levels

L5: "Matching is CPU-bound, so we'll run a pool of matching workers and lock the order book per price level." This is the natural distributed-systems reflex — and it destroys determinism. Two arrivals 3µs apart can match in either order depending on scheduler timing; replay produces a different book.

L6: "One symbol's book is processed by exactly one thread, strictly in sequence-number order. A single thread doing in-memory matching handles hundreds of thousands to millions of events per second — far above any single symbol's peak. I scale by partitioning symbols across engines, never by splitting one book. No locks, no nondeterminism, and replay is exact."

L7: "The engine is the one component in the company where 'make it concurrent' is a firing offence. I'd write that down as an engineering principle, with a list of forbidden inputs: wall clock reads, random numbers, hash-map iteration order, anything from the network that isn't sequenced."

Why "failover" separates levels

L5: "Restart the matching service and rebuild the book from the orders table." Minutes of downtime, and the rebuilt book may not match what members saw.

L6: "A hot standby consumes the identical sequenced stream and runs the identical code. Every N messages, primary and standby exchange a hash of book state. On primary failure, the standby is already at the same sequence number; it takes over within hundreds of milliseconds. If hashes ever diverge, we don't fail over — we halt the symbol, because we no longer know which book is true."

L7: "Halting is a legitimate, rehearsed operation, not a failure. The rulebook says what happens: open orders are canceled or preserved, the symbol reopens through an auction, members get a status message. That's a product and regulatory design, not an engineering footnote."

Why "risk" separates levels

L5: Validates order fields and account balance before matching.

L6: "Pre-trade risk runs inline in the gateway: max order size, price collar (e.g., reject limits more than 10% away from last trade), per-member credit limit, message-rate throttle. Budget: under 5µs. Plus a kill switch per member that cancels all their open orders in one command. Knight Capital lost about $440 million in 45 minutes in 2012 from a bad deployment on the participant side — the exchange-side controls are what bound that blast radius."

L7: "Which controls are ours versus the broker's is a regulatory boundary (in the US, the SEC Market Access Rule puts obligations on brokers). I'd design the exchange controls as a defense-in-depth layer and make kill-switch drills a membership requirement."

The Staff Positions#

PositionRationale
Sequencer first, then everything is deterministicA single total order is the foundation of fairness, replay, and replication
Single-threaded matching per symbol; parallelism across symbolsOne core matches far more than any symbol's peak; locks destroy determinism
Journal before process; state is a function of the journalRecovery is replay, audit is replay, testing is replay
Hot standby replays the same stream; compare state hashesSub-second failover with proof the replicas agree
Halt is a valid failure modeA wrong trade is worse than no trade; exchanges halt, then reopen via auction
Pre-trade risk inline, under 5µs, with per-member kill switchesFat-finger and runaway-algo protection must happen before the book changes
Market data via sequenced multicast with gap recoveryEqual, simultaneous dissemination; TCP fan-out is neither fair nor scalable

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Regulated lit exchange (equities, futures)Fairness provable; µs latency; full auditSequencer + single-threaded engines + hot standby + multicast market dataHalt symbol / market; reopen via auctionDeterministic replay; zero busted trades from system error
Crypto / retail exchange24×7 availability, custody of funds, bursty retail loadSame core engine; wallet/custody ledger; API gateways with heavy rate limitingMaintenance windows unacceptable; must degrade (cancel-only mode)Balances never negative; ledger reconciles with engine
Dark pool / periodic auctionMinimize information leakage; neutralize speedBatch matching every N ms at a single price; no public bookSkip auction cycleUniform price per batch; no leakage of resting interest

🎯 Staff Move: "I'll design a regulated lit exchange with continuous price-time matching and opening/closing auctions, because that's where determinism, fairness and latency collide hardest. I'll target under 50µs from order receipt to acknowledgment inside our network, and I'll treat the inbound journal as the legal record."

The Five Fault Lines#

#Fault LineThe Tension
1Determinism vs ParallelismOne thread per book (exact, replayable) or concurrent matching (more throughput per book, nondeterministic)?
2Latency vs FairnessRace to the lowest latency (rewards the fastest firms) or add speed bumps / batch auctions (level the field)?
3Inline Risk vs SpeedEvery check in the hot path costs microseconds; every check removed is a fat-finger waiting to happen
4Availability vs Correctness on FailureFail over and keep trading, or halt and reopen? Who decides, how fast?
5Market Data: Equal Access vs Scale and RevenueMulticast to all at once vs per-client streams; raw feed vs tiered data products

In the Wild: Real Production Systems#

Why this section belongs here: These are public, well-documented designs. Citing them shows you know the single-threaded model is production reality, not a toy.

LMAX Exchange — The Disruptor and a Single Business-Logic Thread#

LMAX publicly described (via Martin Fowler's well-known write-up) an architecture where all business logic for its retail FX exchange ran on a single thread, processing on the order of 6 million orders per second in memory. Input events were journaled and replicated before processing via the Disruptor, a lock-free ring buffer; recovery came from snapshots plus event replay.

Staff insight: The canonical proof that "single-threaded" is a performance strategy, not a limitation. The reason it's fast is the same reason it's correct: no locks, no contention, cache-friendly sequential processing, and every state change is a deterministic function of the input log.

Nasdaq — INET, OUCH and ITCH#

Nasdaq's matching technology (derived from INET) is paired with published protocols: OUCH for order entry and ITCH for market data — a binary, sequence-numbered feed disseminated over UDP multicast (with the MoldUDP64 framing and a retransmission service for gaps). The same message types that describe book changes let any participant rebuild the full order book.

Staff insight: The market data feed is itself a replicated log. Participants are replicas of the book; sequence numbers plus a retransmit service turn unreliable multicast into a reliable, fair, one-to-many log.

IEX — The Speed Bump#

IEX, which became a registered US exchange in 2016, famously introduced a 350-microsecond delay implemented with about 38 miles of coiled fiber on inbound and outbound paths, so that its own price-sensitive order types could react to market changes before fast traders could exploit stale quotes.

Staff insight: Fairness is a design parameter, not an afterthought. IEX deliberately traded latency for a fairness property — and needed regulatory approval to do it. When an interviewer asks "how do you make it faster?", a Staff answer includes "faster for whom, and is that the product?"

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Single-threaded matching""How does that scale? What about a hot symbol?"Whether you know the actual per-core capacity and partition by symbol
"We journal every message""fsync costs 100µs+. How are you at 50µs?"Replicated-memory durability vs disk durability
"Hot standby""How do you know the standby's book is identical?"State hashing, determinism discipline
"Pre-trade risk checks""What do you check, and in how many microseconds?"Concrete controls and budgets
"Publish market data""Two firms subscribe. How do you guarantee neither sees it first?"Multicast, equal cable lengths, fairness
"Store trades in a database""Is the DB the source of truth or the journal?"Event sourcing, replay
"The engine crashed""Do we keep trading or halt?"Correctness over availability, rulebook thinking

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Orders pass session checks and pre-trade risk in the gateway, then get a global sequence number from the sequencer, which replicates the message to at least two journal hosts in memory before releasing it. Matching engines — one thread per symbol partition — consume the sequenced stream; a hot standby consumes the identical stream. Outputs (acks, fills, book updates) are themselves sequenced and published: execution reports on the member's session, market data on multicast with a retransmit service. Market operations can halt or trigger auctions by injecting sequenced control messages, so even halts are replayable.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Concurrency"Worker pool with locks""One thread per symbol partition, strictly in sequence order. Scale across symbols, never within a book."
Ordering"Timestamps""A sequencer assigns a global sequence number. Timestamps are metadata; the sequence number is truth."
Durability"Write trades to Postgres""Replicate the inbound message to 2+ hosts' memory before processing; disk async. State is a function of the journal."
Failover"Restart and reload""Hot standby at the same sequence; state hash comparison; divergence → halt."
Risk"Validate balance""Inline collars, size, credit, rate limits under 5µs; per-member kill switch; market-wide circuit breakers."
Market data"WebSocket to clients""Sequenced UDP multicast, A/B redundant feeds, retransmit service, snapshot for late joiners."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Single-thread in-memory matching~1–6M events/sec per coreWhy one thread per symbol is enough (LMAX publicly cited ~6M orders/sec)
Order-to-ack latency, top venuestens of µs inside the venueThe budget every component must fit
Kernel-bypass NIC receive~1–2µsvs ~10–20µs through the kernel TCP stack
Pre-trade risk budget< 5µsEach check must be O(1) against in-memory state
Sequencer assign + replicate to memory of 2 hosts~5–15µs on a local low-latency networkThe price of durability without disk fsync
NVMe fsync~20–100µs+Why the hot path uses replicated memory, not per-message fsync
Light in fiber~5µs per kmWhy co-location and equal-length cables matter
IEX speed bump350µsA deliberate fairness tax
Order-to-trade ratiooften 20:1 to 100:1+Most traffic is adds and cancels; size capacity for messages, not trades
US equity symbols~8,000–10,000Partitioning unit; a few hundred symbols carry most volume
Knight Capital 2012~$440M loss in ~45 minWhy kill switches and deploy consistency matter
Journal/audit retentioncommonly 5–7 yearsRegulatory record-keeping horizon
Snapshot intervalevery 1–5 min or N million msgsBounds replay time on recovery to seconds

Interview Walkthrough

The most common mistake: Candidates design a web service — REST API, load balancer, stateless workers, a database of orders — and spend 20 minutes on it. Then the interviewer asks "what if two orders arrive at the same time?" and the design has no answer. Establish the sequencer and the single-threaded engine in the first 10 minutes.


Phase 1: Requirements & Framing (2–3 minutes)#

Functional scope in one sentence:

"Members submit, modify and cancel limit and market orders; we match them with price-time priority, send acknowledgments and fills, and publish order-book and trade data to the market."

Then the non-functional contract — this is the Staff investment:

"Four questions: What market structure — continuous lit trading, auctions, or both? What latency class — are members co-located high-frequency firms or retail over the internet? What regulatory regime — do we need to prove order handling to a regulator? And what's the failure policy — is a halt acceptable? I'll assume a regulated lit equities venue, co-located members, a requirement to reproduce any trade from the audit record, and that halting a symbol is acceptable when correctness is in doubt."

Commit to numbers:

"Targets: order-to-ack under 50µs p99 inside our network, 10K symbols, peak 2M inbound messages/sec across the venue, the busiest symbol at ~200K messages/sec, zero system-caused busted trades, failover under 1 second, market data published to all subscribers within the same microsecond window."

🎯 Staff Move: "I'm going to be opinionated about one thing up front: correctness beats availability here. If the engine's state is ever in doubt, we halt the symbol. A wrong trade is a legal event; a halt is an operational event."


Phase 2: Core Entities & API (1–2 minutes)#

  • Order: order_id, member_id, symbol, side, price (integer ticks — never floats), qty, tif (DAY/IOC/FOK), seq
  • Order Book: per symbol, bids and asks, each a sorted map of price level → FIFO queue of orders
  • Execution / Trade: trade_id, buy_order, sell_order, price, qty, seq
  • Sequenced Message: seq, type, payload, gateway_ts — the unit of the journal

Order entry (binary protocol over TCP session, OUCH-like; FIX for slower members):

EnterOrder(client_order_token, symbol, side, price_ticks, qty, tif)  → Accepted(seq, order_id) | Rejected(reason)
CancelOrder(order_id, qty_to_cancel)                                   → Canceled(seq)
ReplaceOrder(order_id, new_price, new_qty)                             → Replaced(seq, new_order_id)   # loses time priority if price changes or qty increases
Outbound: Executed(seq, order_id, qty, price, trade_id)

🎯 Staff Move: "Prices are integers in ticks. A floating-point price in a matching engine is a nondeterminism bug waiting to happen — two machines rounding differently is a divergent book."


Phase 3: High-Level Architecture (≤5 minutes)#

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk the flow in 90 seconds:

  1. Member's order hits a gateway: session validation, message-rate throttle, pre-trade risk.
  2. Gateway forwards to the sequencer, which stamps a monotonically increasing sequence number and replicates the message to at least two hosts' memory.
  3. The matching engine for that symbol group consumes messages in sequence order, one at a time, and updates the in-memory book.
  4. Outputs — ack, fills, book deltas — are emitted in the same order, tagged with the input sequence number.
  5. The hot standby consumes the same stream and produces identical state; outputs are suppressed unless it's promoted.

🎯 Staff Move: "That's the whole core. Everything interesting is about keeping it deterministic, fast and fair. I'd like to go deep on three things: why single-threaded and how we scale it, how recovery and failover work by replay, and how we disseminate market data fairly."


Phase 4: Transition to Depth (1 minute)#

"The architecture is simple on purpose. The fault lines are: determinism — what's allowed to influence the engine; fairness — what 'first' means when two orders are 1µs apart on different gateways; and failure — when we fail over versus halt. Which would you like first?"

If no preference: lead with determinism and replay — it underpins the other two.


Phase 5: Deep Dives (25–30 minutes)#

Deep dive A: Why one thread, and how it scales (6–8 min)

"Matching is a sequence of small, data-dependent operations on a hot, in-memory structure. A single core with the book in L1/L2 cache processes an order in hundreds of nanoseconds to a few microseconds. Adding threads to the same book adds lock acquisition, cache-line bouncing and — worst — nondeterministic interleaving. So: one thread per book. Throughput scales by partitioning symbols across engines: 10K symbols over 16 engine processes on dedicated cores. The busiest symbol at 200K messages/sec uses maybe 20–40% of one core. Partition assignment is by measured volume, not alphabetically, and rebalancing happens overnight, not intraday."

Deep dive B: Sequencer, journal and replay (6–8 min)

"The sequencer is the only place where order is decided. It assigns sequence numbers and doesn't release a message to engines until it's in memory on at least two hosts in different racks — a few microseconds on a low-latency switch. Disk writes are asynchronous and batched. The engine is a pure function: state_n = apply(state_{n-1}, msg_n). It never reads the wall clock — time comes in as sequenced timer messages. Every few minutes we snapshot the book; recovery is 'load snapshot at seq S, replay S+1 through the end'. At 2M msgs/sec and a snapshot every 60s, replay is at most 120M messages — about 20–60 seconds at replay speed, which is why the standby is hot rather than cold."

Deep dive C: Failover versus halt (5–7 min)

"The standby has consumed every sequenced message; it's at the same sequence number as the primary within microseconds. Every 10K messages, both emit a hash of book state onto a side channel. If the primary dies, the standby promotes — within a few hundred ms including detection — and resumes emitting outputs from the next sequence. Members see a brief pause. If the hashes ever mismatch, we halt the affected symbols, because we can't know which book is true. Halt means: stop accepting new orders, keep cancels, publish a halt status on market data, and reopen through an auction once reconciled."

Deep dive D: Fair market data dissemination (4–6 min)

"Book deltas and trades are published as sequenced messages over UDP multicast — one send reaches every subscriber at the switch simultaneously. Two redundant feeds (A and B) over separate networks; subscribers arbitrate by sequence number. Gaps are filled from a retransmit service; late joiners get a snapshot then the incremental feed. In the co-location facility, cross-connects are equal-length so no member is physically closer. Per-client TCP streams would be both unfair — the 1,000th client gets data after the first 999 — and unscalable."


Phase 6: Wrap-Up (2–3 minutes)#

"An exchange is a deterministic state machine behind a sequencer. Fairness comes from a single total order and equal dissemination; resilience comes from replaying that order on a hot standby; correctness is protected by halting rather than guessing. The operational story is market ops owning halts and kill switches, the engine team owning determinism, and a replayable journal that serves recovery, testing and the regulator from the same artifact."

"What I'd build next: closing-auction optimization, a certification environment where members replay their production traffic against new engine builds, and a surveillance pipeline reading the journal."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Web-service framingREST, load balancer, stateless workers, DBSequencer + single-threaded engine in first 10 min
Order book data structures10 min on red-black trees vs skip lists"Sorted price levels, FIFO per level, order-id hash map. Moving on."
Accounts and authDesigns login and user tables"Members are pre-provisioned sessions; not the interesting part."
No failure policy"Restart it"Hot standby, state hash, halt-and-auction
Floating-point pricesNever mentionsInteger ticks, called out explicitly
No numbers"Low latency""< 50µs p99, 5µs risk budget, 200K msgs/sec hot symbol"

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

The exchange question punishes pattern-matching. Every habit from web-scale design — horizontal scaling, eventual consistency, retries with at-least-once semantics, timestamps for ordering — is wrong in the core. Staff candidates recognize that the problem is a replicated state machine with a legally significant log, and design everything around keeping it deterministic. They also recognize the organizational weight: market operations, compliance, and members are all stakeholders in a failure.

What interviewers listen for:

  1. A single total order established before matching
  2. A defended choice of single-threaded matching with a scaling story
  3. Recovery and replication by replay, with state verification
  4. Concrete pre-trade controls and kill switches with a latency budget
  5. "Halt" named as a first-class, designed failure mode

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"Two orders for the same symbol arrive 1 microsecond apart on two different gateways. Which one gets filled, how do you prove it to a regulator a year later, and would your backup have made the same decision?"

A Staff answer: the sequencer decides; the journal records the decision with the sequence number; replaying the journal on any engine build reproduces the same fill; the standby made the same decision because it consumed the same sequence.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Regulated lit exchange → determinism, fairness, auditability

  • Constraint: order handling must be provable and reproducible; tens of µs latency
  • Mechanism: sequencer, single-threaded engines, replicated journal, multicast data
  • Failure mode: halt the symbol; reopen via auction
  • Who pays: members during a halt (can't trade), market ops (running the reopening), the venue's reputation

Crypto / retail exchange → availability, custody, burst handling

  • Constraint: 24×7, no maintenance windows; users hold balances on the platform
  • Mechanism: same core engine plus a custody ledger reconciled against the engine; heavy API rate limiting (see Rate Limiting); cancel-only degraded mode
  • Failure mode: cancel-only mode rather than halt; withdrawals paused on ledger mismatch
  • Who pays: users during cancel-only mode; finance/treasury for reconciliation gaps

Dark pool / periodic auction → information protection, speed neutrality

  • Constraint: don't reveal resting interest; latency advantage shouldn't matter
  • Mechanism: collect orders for N ms, uncross at a single price; no public book
  • Failure mode: skip a batch cycle
  • Who pays: participants waiting for the next batch; the operator if leakage occurs

2.2 When NOT to Build a Matching Engine#

  • Internal crossing for a small broker. If volume is thousands of orders per day, a database with serializable transactions is simpler and auditable enough.
  • Marketplace matching with human timescales (ride-hailing, freelance jobs). These are assignment problems with scoring, not price-time books; see Ride Hailing & Delivery.
  • Ticketing / inventory. "First come first served" on seats is a contention problem (Flash Sales, Contention), not an order book.
  • When a licensed venue or white-label engine exists. Several vendors sell certified matching engines; building one is justified only if matching behavior is the product differentiator.

2.3 What the Interviewer Leaves Underspecified#

Omitted DetailWhy It MattersWhat to Ask
Market structureContinuous vs auction changes the engine"Continuous trading, auctions, or both?"
Latency classCo-located µs vs internet ms changes the entire stack"Who are the members — HFT firms or retail apps?"
Asset class and hours24×7 crypto vs 6.5-hour equities changes maintenance and snapshots"Is there a daily close?"
Regulatory regimeAudit, fair access, circuit breakers"Are we a regulated venue?"
Failure toleranceHalt vs keep trading"Is halting a symbol acceptable?"
SettlementWho holds funds, when trades become final"Do we clear through a CCP or hold balances ourselves?"

Staff engineers surface these. Senior engineers assume a web app.

2.4 Precise Terminology#

TermPrecise MeaningCommon Misuse
SequencerComponent assigning a single total order to all inbound eventsConfused with a message queue
DeterministicSame input sequence → bit-identical state and outputsUsed loosely for "reliable"
Price-time priorityBest price first; FIFO within a price levelAssumed universal (pro-rata exists)
Aggressor / passiveIncoming order that crosses the spread vs resting order it hitsMixed up when describing fees
Tick sizeMinimum price incrementPrices modeled as floats
Crossed / locked marketBid > ask / bid = askAllowed to persist in the book
Kill switchCommand canceling all of a member's orders and blocking new onesAssumed to be a network disconnect
Circuit breaker / LULDMarket-wide or per-symbol trading pause on large movesConfused with the software circuit-breaker pattern
Drop copyReal-time copy of a member's executions sent to their risk/back officeOmitted entirely
Busted tradeTrade canceled after the fact by the venueTreated as a routine correction

🎯 Staff Insight: If the interviewer says "make it distributed," ask: "Distributed across symbols, or distributed within a single book? The first is how we scale; the second breaks price-time priority and determinism."


3. The Fault Lines#

Each fault line has a technical axis and an organizational one. In an exchange, the organizational axis includes a regulator — which makes "who signed off" a literal question.

3.1 Fault Line 1: Determinism vs Parallelism#

The tension: More threads on a book would raise per-symbol throughput. They also make the book's evolution depend on scheduler timing, which breaks replay, replication, and fairness.

ChoiceWhat WorksWhat BreaksWho Pays
Single thread per symbol partitionExact replay; no locks; cache-hot; hot standby by replayPer-symbol ceiling (~1M+ events/sec — rarely binding); one slow message delays the partitionEngine team (must keep every handler O(log n) or better)
Multi-threaded book with locksHigher theoretical per-book throughputNondeterministic interleaving; replay diverges; lock contention at the top of bookMembers (unfair fills), compliance (can't reproduce), on-call (heisenbugs)
Pipeline parallelism around a single-threaded coreParse, risk, and serialization on other cores; core stays single-threadedMore moving parts; needs lock-free handoff (ring buffers)Engine team (complexity), acceptable
Diagram: 3.1 Fault Line 1: Determinism vs Parallelism

L6 answer: "Parallelism lives around the core — decoding, risk, journaling, market-data encoding each get a core connected by lock-free ring buffers, the LMAX Disruptor pattern. The matching logic itself is one thread. We scale out by symbol partition. I'll forbid anything nondeterministic from entering the core: no wall-clock reads, no random numbers, no iteration over unordered hash maps, no floating point."

When to deviate: Never within a single order book for a regulated venue. For cross-symbol strategies (spread instruments, options with legs), co-locate related books in the same partition so implied matching stays single-threaded.

🧭 Principal Move: "I'd enforce the determinism boundary with tooling, not reviews: a CI job replays yesterday's full production journal against every new build and compares output hashes byte-for-byte. A mismatch blocks the release."


3.2 Fault Line 2: Latency vs Fairness#

The tension: Lower latency is what members pay for. But pure speed rewards whoever spends the most on proximity and hardware, and it lets fast participants pick off stale quotes. Every fairness mechanism costs latency or throughput.

MechanismWhat It DoesLatency CostWho Pays
Pure continuous FIFOFirst to the sequencer winsNoneSlower participants (adverse selection)
Equal-length cross-connects in co-locationNo member physically closer than anotherAdds cable to the nearest membersVenue (capex), fastest members lose a few ns
Speed bump (e.g., IEX 350µs)Delay all inbound/outbound symmetrically+350µsFast traders; venue needs regulatory approval
Randomized delayJitter order arrival within a windowRandom 0–N µsEveryone, and determinism must be preserved by sequencing after the delay
Frequent batch auctionsMatch every N ms at a single priceUp to N msParticipants wanting immediacy

L6 answer: "The sequencer defines fairness — first sequenced, first served — so the fairness work is making arrival at the sequencer fair: equal-length cables in co-location, gateways with identical code paths and latency, and monitoring gateway latency skew across the pool. If a gateway is 3µs slower than its peers, members connected to it are disadvantaged; I'd alert when p99 skew across gateways exceeds 1µs."

When to deviate: If the product strategy is to attract long-term investors over HFT flow, a speed bump or batch auction is a deliberate product choice — made by the business and approved by the regulator, not by engineering.

🎯 Staff Insight: "Gateway latency variance is a fairness bug, not just a performance bug. Two members who send at the same instant should reach the sequencer in the same order regardless of which gateway they're assigned to."


3.3 Fault Line 3: Inline Risk vs Speed#

The tension: Each pre-trade check costs microseconds. Each removed check is a runaway algorithm or fat finger waiting to hit the book.

CheckBudgetCatchesWho Pays If Missing
Message rate per session< 0.5µsRunaway loops flooding the venueAll members (venue slowdown)
Max order qty / notional< 0.5µsFat-finger 1,000,000 instead of 1,000Member, and counterparties if busted
Price collar (e.g., ±10% from reference)< 1µsOrders far from marketCounterparties (trades later busted)
Credit / exposure limit per member1–3µsMember exceeding capitalClearing house, venue
Self-trade preventioninside matchingWash tradesCompliance
Kill switch state< 0.2µs (flag lookup)Blocked membersMarket
Diagram: 3.3 Fault Line 3: Inline Risk vs Speed

L6 answer: "All pre-trade state lives in memory in the gateway's risk module, updated from the engine's execution stream so credit usage is current within microseconds. Total budget under 5µs. Plus a per-member kill switch, triggered by the member or market ops, that's itself a sequenced message — so 'all orders canceled at seq N' is replayable."

Who pays: The venue pays latency; the member pays if they disabled their own controls. The Staff move is naming that controls are layered — broker-side (regulatory obligation in many markets), venue-side, and market-wide circuit breakers.


3.4 Fault Line 4: Availability vs Correctness on Failure#

The tension: Trading halts cost members money and cost the venue market share. Trading on a book that may be wrong costs far more — busted trades, lawsuits, regulatory action.

FailureKeep TradingHaltStaff Default
Primary engine crash, standby hash-verifiedPromote standby (< 1s)—Fail over
Standby hash mismatchUnsafeHalt symbols in partitionHalt
Sequencer primary lossPromote sequencer standby with fencingIf journal replication uncertainFail over if last replicated seq agreed; else halt
Market data publisher lagTrading continues but members trade blindHalt if lag > thresholdHalt at ~1s of dissemination lag
Single gateway lossMembers reconnect to another—Keep trading
Diagram: 3.4 Fault Line 4: Availability vs Correctness on Failure

L6 answer: "Fail over when we can prove the standby's book equals the primary's. Halt when we can't. Reopening goes through an auction so the price is re-discovered fairly rather than whoever reconnects first getting the stale quotes."

🧭 Principal Move: "The halt-and-reopen procedure is part of the rulebook filed with the regulator. I'd make sure engineering, market operations and compliance co-own it and rehearse it monthly in the test environment — a halt is where venues lose reputation, and it's almost never practiced."


3.5 Fault Line 5: Market Data — Equal Access vs Scale and Revenue#

The tension: Market data must reach all subscribers simultaneously for fairness. It's also a major revenue line for exchanges, which creates pressure for tiered products (faster, deeper feeds for those who pay more).

ChoiceWhat WorksWhat BreaksWho Pays
UDP multicast, sequenced, A/B feedsOne send reaches all at once; scales to thousandsPacket loss → gaps; needs retransmit and snapshot servicesVenue infra team
Per-client TCP streamsReliable, simpleSerialized sends — the 1,000th client gets data milliseconds late; slow consumers back-pressureLate subscribers (unfairness)
Consolidated / cloud distribution (e.g., WebSocket for retail)Reaches internet usersLatency in ms; must be clearly a derived, slower feedRetail users (acceptably)
Tiered depth products (top-of-book vs full depth)Revenue; smaller feeds for those who don't need depthRegulatory scrutiny on fair accessBusiness manages the policy

L6 answer: "Primary feed is sequenced multicast inside co-location with A/B redundancy, a retransmit service for gaps under N messages, and a snapshot service for late joiners or larger gaps. Retail and cloud consumers get a derived feed via a separate fan-out tier — see Real-Time Updates — explicitly labeled as delayed relative to the direct feed."

🎯 Staff Move: "Market data is a replicated log: every subscriber is a replica of the order book. If they miss a sequence number, their book is wrong until they recover it. That's why the retransmit and snapshot services are part of the core design, not add-ons."


4. Failure Modes & Operational Reality#

4.1 Engine Crash During Peak Trading#

t=0:        09:30:02 — open auction just uncrossed. 1.4M msgs/sec venue-wide.
t=0:        Engine partition 7 (symbols including a top-5 name) segfaults on a new order type.
t=+50µs:    Standby for partition 7 has already processed the same message... and also segfaults.
t=+200ms:   Health monitor: both primary and standby for partition 7 down.
t=+300ms:   Sequencer continues sequencing messages for partition 7; they queue in the journal.
t=+2s:      Market ops halts partition-7 symbols via sequenced HALT message.
t=+4min:    Root cause: new order-type handler dereferences null on an edge case.
t=+25min:   Hotfix disables the order type at the gateway (reject); engines restart,
            replay from snapshot, skip nothing — the offending message is rejected
            deterministically by the new build's validation.
t=+32min:   Reopening auction; trading resumes.

Detection: engine.heartbeat_missing{partition}, engine.last_processed_seq stalled, journal.pending_for_partition growing. Blast radius: All symbols in partition 7 (~600 symbols, ~8% of volume). Key insight: Deterministic replication means a deterministic bug kills the standby too. Hot standby protects against hardware failure, not software bugs. Mitigation: Halt, fix, replay. Gateways can reject the poison message type immediately. Prevention: Journal replay of production traffic in CI; fuzzing new order types; canary new features on low-volume symbols first; feature flags per order type. Owner: Engine team (fix), market operations (halt/reopen), compliance (member notice).

4.2 Sequencer Split Brain#

t=0:       Primary sequencer S1 loses its network link to the coordinator for 800ms.
t=+500ms:  Coordinator lease expires; S2 promoted with epoch 18.
t=+800ms:  S1's link returns. S1 (epoch 17) still has messages in flight.
With fencing:   journal hosts reject epoch 17 writes; S1 steps down. Seqs continuous.
Without fencing: two messages carry sequence 88,120,442 with different content.
                 Engines disagree depending on which they received. Hash mismatch → halt.

Detection: sequencer.epoch_changes, journal.duplicate_seq_rejected, engine state-hash mismatch. Mitigation: Epoch fencing at the journal replicas; lease duration longer than the worst tolerable pause. See Distributed Consensus. Owner: Core infrastructure team.

4.3 Market Data Gap Storm#

t=0:       Microburst: 400K messages in 10ms after a macro news release.
t=+10ms:   Switch buffers overflow on feed A; 2,300 packets dropped.
t=+11ms:   800 subscribers detect gaps; feed B is intact for most — they arbitrate.
t=+12ms:   120 subscribers on a misconfigured host missed both feeds; request retransmit.
t=+15ms:   Retransmit server receives 120 × 2,300 requests; queue grows.
t=+2s:     Subscribers still gapped fall back to snapshot service.

Detection: md.feed_a_gap_rate, md.retransmit_requests_per_sec, md.snapshot_requests. Blast radius: Subscribers trading on stale books — which becomes adverse selection against them. Mitigation: A/B arbitration; retransmit rate limits per subscriber; snapshot fallback. Prevention: Capacity-plan for microbursts (peak per-ms rate, not per-second); conflated feeds for slow consumers. Owner: Market data team.

4.4 Runaway Member Algorithm#

t=0:       Member X deploys a new algo. A flag is misapplied; it sends buy orders in a loop.
t=+1s:     X's session: 45K msgs/sec (normal: 2K). Rate throttle at 10K/sec rejects the excess.
t=+3s:     Orders within the throttle still accumulate: X's net position grows by $8M/sec.
t=+12s:    Credit limit ($100M exposure) reached; gateway rejects all new X orders.
t=+15s:    Market ops sees `member.exposure_utilization{X}` = 100% and calls X.
t=+60s:    X requests kill switch; all X orders canceled at seq N.

Detection: member.msg_rate, member.exposure_utilization, member.reject_rate. Blast radius: Bounded by credit limit; counterparties trade at market prices, not harmed. Mitigation: Throttle, credit limit, kill switch — all pre-agreed with the member. Owner: Member (their algo), market ops (kill switch execution), risk team (limit levels).

4.5 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Engine hardware failureengine.heartbeat_missing1 partition, < 1sPromote hash-verified standbyEngine team
Deterministic engine bugPrimary and standby both stall1 partition, minutesHalt, gateway reject, fix, replayEngine team + market ops
State hash mismatchengine.state_hash_mismatch1 partitionHalt; reconcile from journalEngine team + compliance
Sequencer losssequencer.heartbeat_missingWhole venueFenced standby promotionCore infra
Market data gapsmd.gap_rate, retransmit loadSubscribers on stale booksA/B arbitration, snapshotMarket data team
Gateway latency skewgateway.p99_skew_us > 1Fairness for assigned membersDrain and rebalance gatewayGateway team
Runaway membermember.msg_rate, exposureBounded by credit limitThrottle, kill switchMarket ops + member
Journal disk lagjournal.disk_lag_msgsDurability of last N msgs if both memory copies lostThrottle intake, alertCore infra

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
OrderingTimestamps / queue orderSequencer total order; seq is truthTreats the sequenced journal as legal record with retention and regulator access
ConcurrencyWorker pool + locksSingle thread per book, parallel around the coreCodifies the determinism boundary; CI replay gate for every build
DurabilityTrades in a DBReplicated memory before processing; async disk; snapshotsPrices the durability tiers vs latency; defines disaster-recovery site RPO with regulators
FailureRestartHash-verified standby; halt on doubt; auction reopenDesigns halt/reopen as rulebook, co-owned with compliance and ops, rehearsed
RiskBalance checkInline controls with µs budgets; kill switchLayered market-wide controls; member obligations; incident governance
Market dataPush to clientsSequenced multicast, A/B, retransmit, snapshotFair-access policy vs data revenue; derived feeds strategy
OwnershipEngineering teamEngine / gateway / market-data / market-ops boundariesOrg structure mapping to regulatory accountability

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Defends single-threaded"One core does a million-plus matches a second. Parallelism across symbols, never inside a book."
Sequencer as truth"The sequence number is the ordering; timestamps are annotations."
Replay as the universal tool"Recovery, standby, testing, audit — all replay of the same journal."
Halt is designed"Divergence means halt, then reopen via auction."
Knows the deterministic-bug trap"A hot standby running the same code dies on the same poison message."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
Microservices with a DB per serviceNo total order; no determinism
Floating-point pricesNondeterminism and rounding disputes
"Kafka for everything" without a single partition per bookPartition reordering or consumer-group rebalances break ordering and replay if not designed carefully
Eventual consistency for the bookThe book is the one thing that must be exact
No failure policy beyond restartMinutes of downtime and no proof of state

5.4 Common False Positives#

  • Knows HFT jargon: Talking about FPGAs and kernel bypass ≠ designing a fair, recoverable venue.
  • Order book data-structure depth: A perfect price-level tree with no sequencer is still wrong.
  • Mentions Kafka's ordering guarantee: A single-partition log is a valid sequencer shape, but at 5–10ms p99 it's the wrong latency class for co-located trading — the candidate must say so.
  • "Blockchain for audit": Tamper-evidence is solvable with hash-chained journals; consensus across untrusted parties isn't the requirement.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minMarket structure, latency class, audit, halt acceptability; numbers
Entities + API3–5 minOrder, book, trade, sequenced message; integer ticks
High-level design5–12 minGateway → risk → sequencer → single-threaded engines → outputs
Transition12 minOffer determinism/replay, fairness, failure policy
Deep dives12–40 minDeterminism + scaling → replay + failover → market data fairness → risk
Wrap-up40–45 minOwnership: engine, market ops, compliance; what's next

6.2 How Interviewers Pivot — And What They're Testing#

Interviewer SignalWhat They Care AboutWhere to Go Deep
"How does it scale?"Whether you parallelize the wrong thingSymbol partitioning, per-core capacity, pipeline around core
"What if the engine dies?"Recovery designHot standby, state hash, deterministic-bug trap
"Two orders at the same time"Ordering and fairnessSequencer; gateway latency skew
"Fat finger"Risk controlsCollars, size, credit, kill switch, µs budgets
"How do clients get prices?"Distribution fairnessMulticast, A/B, retransmit, snapshot
"Build it on cloud?"Latency realismCo-location vs cloud; derived feeds in cloud

6.3 What to Deliberately Skip#

TopicWhy L5 Goes HereWhat L6 Says Instead
User accounts and loginFamiliar"Members are provisioned sessions."
Order-book data structure bake-offFeels algorithmic"Array of price levels near the touch, map beyond, FIFO per level, id→order hash. Done."
Settlement internalsSounds complete"Trades go to the clearing house via drop copy; T+1 settlement is theirs."
Charting UIVisible"Derived from market data; not core."

6.4 Follow-Up Questions to Expect#

  1. "Your busiest symbol spikes 10×. What happens?"
  2. "How do you test a new engine build before production?"
  3. "Can the standby ever disagree with the primary? How would you know?"
  4. "What does a cancel-replace do to time priority?"
  5. "How do you handle the market open, when thousands of orders arrive at once?"
  6. "Could we run this in a public cloud?"
  7. "How would a regulator reconstruct a trade from last March?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design a stock exchange."

Staff Answer

"Before boxes: what market structure — continuous trading, auctions, or both? Who are the members — co-located firms or retail apps? Are we regulated, meaning we must reproduce any trade on demand? And is halting acceptable when we're unsure of state? I'll assume a regulated lit equities venue, co-located members, full reproducibility, and halts allowed. That gives me a design built around one idea: a sequencer establishes a single total order of every inbound event, and everything downstream — matching, standby, market data, audit — is a deterministic function of that order. Matching is single-threaded per symbol partition. I'll cover determinism and scaling, then replay-based failover, then fair market data."

Why this is L6:

  • Frames the problem as a deterministic state machine before any component
  • Commits to a failure philosophy (halt over guess) in the opening
  • Previews a deep-dive order that follows from the core idea

What L7 adds:

  • Names market structure as the one-way door the business must choose — continuous vs auction vs dark determines engine, rulebook and regulatory filings
  • Asks for the business model: transaction fees vs data revenue changes the market-data design

Drill 2: Defend Single-Threaded#

Prompt: "A single thread? That can't possibly scale. Why not a thread pool?"

Staff Answer

"Three reasons. Capacity: a matching operation on an in-memory, cache-resident book is hundreds of nanoseconds to a few microseconds, so one core handles on the order of a million events per second — our busiest symbol peaks around 200K. Determinism: with multiple threads, two orders arriving microseconds apart can match in different orders depending on scheduling, which breaks replay, the hot standby, and our ability to prove order handling. Latency: locks and cache-line contention on the top of book — exactly where every order competes — add tail latency. We scale by partitioning symbols across engines, 16 partitions on dedicated cores, and we parallelize around the core: decode, risk, journal and encode run on their own cores connected by lock-free ring buffers."

Why this is L6:

  • Quantifies capacity against actual peak rather than asserting
  • Ties the choice to replay and fairness, not just speed
  • Shows where parallelism does go

What L7 adds:

  • Turns "no concurrency in the core" into an engineering principle with CI enforcement (production-journal replay, byte-identical outputs)
  • Plans for the day a single instrument exceeds one core: splitting it is a market-structure change (e.g., auction), not an engineering one

Drill 3: Make Durability Concrete#

Prompt: "fsync is 50–100µs. Your latency target is 50µs total. How is anything durable?"

Staff Answer

"Durability comes from replication to memory, not from disk on the hot path. The sequencer sends each message to two journal hosts in different racks over a kernel-bypass network; each acknowledges when the message is in memory — a few microseconds round trip. Only then is the message released to engines. Journal hosts write to disk asynchronously in large batches. We lose data only if both memory copies die before their batch hits disk — independent power, racks, and UPS make that a correlated-failure event we accept, and we monitor journal.disk_lag_msgs to keep the window under ~10ms. That's the same trade LMAX described: replicate the input, then process."

Why this is L6:

  • Separates durability mechanism (replicated memory) from persistence (async disk)
  • Names the residual risk and bounds it with a metric
  • Keeps the ordering: durable before processed

What L7 adds:

  • Defines the DR site's RPO with the regulator explicitly (async to a remote site ~ms of loss), and the process to reconcile members' records if invoked
  • Prices a third synchronous replica versus the latency it costs members

Drill 4: The Poison Message#

Prompt: "A new order type crashes the primary engine. What does your hot standby do?"

Staff Answer

"It crashes too — it runs the same code on the same input. Deterministic replication protects against hardware failure, not deterministic bugs. So: the sequencer keeps journaling; market ops halts the affected partition with a sequenced HALT; gateways are switched to reject the new order type (a config flag, seconds); we restart engines from the last snapshot with a build or config that rejects that message deterministically and replay. Reopen with an auction. Prevention: every build replays yesterday's full production journal plus a fuzzed corpus before release; new order types launch disabled, then enabled for a handful of low-volume symbols first."

Why this is L6:

  • Identifies the correlated-failure trap in deterministic replication
  • Has a gateway-level kill for the offending feature
  • Prevention through replay-in-CI and canary symbols

What L7 adds:

  • Considers diversity: some venues run a standby on a different build or keep N-1 build ready — and prices that complexity
  • Makes feature-flagged order types a release policy for all engine changes

Drill 5: Hot Symbol#

Prompt: "A meme stock goes from 20K to 400K messages/sec. What happens?"

Staff Answer

"First, its partition: if 400K/sec pushes that engine core past ~70% busy, the other symbols in the partition suffer queueing delay. Intraday, I don't move symbols between engines — that's a risky state transfer mid-session. Instead I pre-plan: the top ~20 most volatile symbols by recent volume get lightly loaded partitions at start of day. Second, the gateway: per-session throttles keep any one member from dominating. Third, market data: 400K updates/sec on one symbol stresses multicast and subscribers; conflated feeds for slower consumers. Finally, market-wide controls: if the price moves outside limit-up/limit-down bands, the symbol pauses automatically — which is also a capacity relief valve."

Why this is L6:

  • Knows intraday rebalancing is dangerous and plans placement instead
  • Protects neighbors of the hot symbol
  • Connects regulatory volatility controls to capacity

What L7 adds:

  • Defines a daily placement process owned by market ops with data from the previous day's volume and news calendar
  • Capacity SLO stated as "max per-symbol rate at which p99 < 50µs" and published to members

Drill 6: Multi-Tenant Fairness Among Members#

Prompt: "One large market maker sends 60% of all messages. Other members complain about latency."

Staff Answer

"Check whether the latency is in the gateway or the engine. If gateways: the market maker's sessions should be on dedicated gateway instances so their volume doesn't queue other members — gateway assignment is a fairness decision. If the engine: their messages are legitimately sequenced, and the engine processes in order — we can't deprioritize them without violating price-time priority. The lever there is per-session message-rate caps and order-to-trade ratio policies with fees for excessive messaging, which are rulebook changes. Measure gateway.queue_delay_us by member and publish per-member latency percentiles so the complaint becomes data."

Why this is L6:

  • Separates ingress queueing (fixable) from sequenced processing (must stay fair)
  • Uses gateway assignment and rate caps rather than reordering
  • Makes fairness measurable per member

What L7 adds:

  • Frames messaging-fee policy as a business and regulatory decision with engineering data
  • Balances market-maker incentives (liquidity) against fairness for others

Drill 7: Build vs Buy#

Prompt: "Should we build our own matching engine or license one?"

Staff Answer

"License unless matching behavior is our product differentiator. Several vendors provide certified engines with gateways, market data and surveillance; a new venue can be live in 12–18 months. Building a production engine — sequencer, standby, market data, risk, certification tools — is a team of 15–30 specialized engineers for 2–3 years, then permanent ownership of the most scrutinized code in the company. We'd build if we're introducing a novel market structure (new auction type, new asset class) the vendors can't support, or if engine latency is the competitive axis. Either way, own the rulebook, the market-ops tooling and the journal format — those are what the regulator and members depend on."

Why this is L6:

  • Realistic team size and timelines
  • Names the specific conditions under which building wins
  • Keeps ownership of the artifacts that create lock-in

What L7 adds:

  • Evaluates vendor risk: source escrow, regulatory audit access to vendor code, exit path if the vendor is acquired
  • See Build vs Buy — the matching engine is the rare case where "core differentiator" may genuinely be true

Drill 8: Changing Rules Without an Outage#

Prompt: "We're changing tick sizes for 500 symbols next month. How?"

Staff Answer

"Tick size is reference data that the engine reads, so it must enter the engine as a sequenced message — never as a file the engine reloads, or replay would produce different results. Plan: publish the change to members 30+ days ahead (they need to update their systems); load new tick tables into the engine as a sequenced REFDATA_UPDATE effective at the start of a session; resting orders at prices no longer valid are canceled with a specific reason code at the transition. Rehearse by replaying a real day's journal with the new tick table in the test environment and in member certification. Rollback is another sequenced update at the next session."

Why this is L6:

  • Keeps reference data inside the deterministic input stream
  • Handles resting orders explicitly
  • Member communication and certification as part of rollout

What L7 adds:

  • Tick-size policy is regulatory; coordinates with the rule-filing timeline
  • Standardizes all reference-data changes through one sequenced mechanism across products

Drill 9: Cost#

Prompt: "The CFO asks why the exchange infrastructure costs so much."

Staff Answer

"The cost isn't compute — 16 engine cores do the matching. It's the fairness and resilience envelope: co-location facility space and equal-length cabling, redundant low-latency switches, A/B market data networks, a hot standby for every engine and sequencer, a disaster-recovery site, and 7-year journal retention. Rough split: facilities and network ~40%, redundancy ~25%, market data distribution ~20%, storage/compliance ~15%. The lever isn't removing redundancy; it's matching redundancy to risk — e.g., DR site as warm rather than hot for a small venue — and monetizing: co-location and market data are revenue lines that exist because of the same infrastructure."

Why this is L6:

  • Knows where the money actually goes
  • Treats redundancy as a risk decision, not waste
  • Connects cost to revenue streams

What L7 adds:

  • Builds a unit-economics view: cost per million messages vs revenue per million traded shares
  • Frames the DR posture decision for the board with regulatory consequences

Drill 10: Multi-Region#

Prompt: "Should we run the matching engine active-active in two regions?"

Staff Answer

"No. A single book cannot have two sequencers without a cross-region consensus round on every order — 1ms+ even between nearby data centers, 60ms+ cross-country — which destroys both latency and fairness (members near one region win). Exchanges run one primary matching site with a synchronous or near-synchronous standby nearby, and a disaster-recovery site further away replicated asynchronously, activated by a declared, rehearsed procedure. What can be multi-region: market data distribution to remote users, member portals, surveillance and analytics consuming the journal."

Why this is L6:

  • Rejects active-active with the specific latency and fairness reasoning
  • Describes the real primary/DR pattern
  • Separates what can be distributed from what can't

What L7 adds:

  • Defines DR invocation authority (who declares a disaster) and the member-facing procedure
  • Considers regulatory expectations on recovery time for critical market infrastructure

8. Deep Dive Scenarios#

Deep Dive 1: Peak Traffic — Market Open After Major News#

Context: Overnight news moves the whole market. At the 09:30 open, inbound messages spike to 5× normal peak. Order-to-ack p99 goes from 40µs to 2ms for 90 seconds. Members call market ops. The on-call escalates to you.

Questions to Surface First:

  • Where is the queueing — gateways, risk, sequencer, a particular engine partition, or output encoding?
  • Is it uniform across partitions or one partition with hot symbols?
  • Were any messages rejected or throttled? Any member disproportionately affected?
  • Did market data keep pace, or did subscribers trade on stale books?

Typical L5 Approach: Adds more gateway instances and scales up engine hosts for tomorrow. Plausible — but misses that one partition hosting three news-driven symbols saturated while the rest were idle, and adding hosts doesn't change placement.

Staff Approach: Pulls per-stage latency (stage.queue_delay_us for decode, risk, sequencer, engine partition, encode) and finds partition 4's core at 100% for 90s. Rebalances hot symbols to lightly loaded partitions for the next session using overnight volume and news signals, and checks market data lag — found that 3% of subscribers fell back to snapshot, trading blind for ~2s.

Principal Approach: Establishes capacity as a published contract: tested peak per partition and per symbol, pre-open placement driven by a volatility forecast, and quarterly load tests at 10× normal peak replaying real open-auction traffic. Reports p99 at open as a member-facing SLO.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Confirm no correctness impact: state hashes OK, no gaps unrecovered. Latency, not correctness.
TriagePer-stage latency breakdown; per-partition core utilization; top symbols by message rate.
Quick fixNone intraday — moving symbols mid-session is riskier than the latency. Tighten session throttles if one member dominates.
GuardrailsIf market data lag > 1s on any partition, halt that partition per rulebook.
Post-mortemPlacement algorithm; open-auction load testing; subscriber snapshot storms.

Metrics to Watch: stage.queue_delay_us{stage}, engine.core_utilization{partition}, md.publish_lag_us, md.snapshot_requests, order_ack_latency_p99{member}.

Organizational Follow-up: Market ops owns daily placement; engine team owns the tooling; member-facing latency report.

Ownership Question: "Who decides to not rebalance intraday while members are complaining?" Staff answer: Market operations, per a pre-agreed runbook — engineering advises. The rule is written before the incident: no intraday partition moves except under halt.

Key Takeaway: "At the open, capacity is a placement problem, not a hardware problem."

What clears the Staff bar:

  • Per-stage latency decomposition
  • Refuses risky intraday changes
  • Placement as the durable fix

Deep Dive 2: Silent Failure — The Standby Was Wrong for Weeks#

Context: A routine planned failover of partition 9 triggers a state-hash mismatch. Investigation shows the standby's book has diverged on a tiny number of orders for three weeks — hashes were computed but the comparison job had been failing silently since a config change.

Questions to Surface First:

  • Which book is correct? Can we replay the journal on a clean engine to find out?
  • What is the source of nondeterminism?
  • Why did the comparison fail silently? Who owns that job?
  • Have any live decisions depended on the standby (e.g., earlier failovers in the window)?

Typical L5 Approach: Rebuilds the standby from the primary's snapshot, fixes the comparison job, adds an alert on its failures. Fixes the symptom; doesn't find the nondeterminism.

Staff Approach: Replays the journal on a third, clean engine; finds the primary and replay agree, standby differs. Root cause: the standby host has a different CPU microcode/library build that changed iteration order of an unordered map used in an expiry routine. Removes the unordered iteration (determinism bug), makes hash comparison a heartbeat-level check — missing comparisons page, and blocks failover to a standby without a verified hash in the last 60s.

Principal Approach: Treats "verification job silently failing" as an org-wide class: any safety control that can stop running without anyone noticing needs a dead-man's-switch alert. Standardizes environment parity (pinned builds, identical images) across primary and standby as an audited control, and adds determinism linting (banned constructs) to the engine codebase.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateAbort failover; stay on primary. Mark standby unsafe.
TriageClean replay to arbitrate; diff books; bisect by seq to first divergence.
Quick fixRebuild standby from primary snapshot; verify hash.
GuardrailsFailover precondition: verified hash < 60s old. Missing comparison = page.
Post-mortemNondeterministic construct; environment drift; silent control failure.

Metrics to Watch: engine.state_hash_compare_age_s, engine.state_hash_mismatch, build.env_parity_violations.

Organizational Follow-up: Dead-man's switches on all safety jobs; environment parity audit monthly.

Ownership Question: "Who owned the comparison job?" Staff answer: It was owned by nobody after a team reorganization — which is the finding. Safety controls get named owners in the service catalog, reviewed on every reorg.

Key Takeaway: "A standby you haven't verified is a second opinion you can't trust. Verification that can fail silently is not verification."

What clears the Staff bar:

  • Uses a third replay to arbitrate truth
  • Hunts the nondeterministic construct
  • Makes verification freshness a failover precondition

Deep Dive 3: Large Member Onboarding#

Context: A major market maker plans to join, expecting to send 3× the messages of your current largest member and wanting a guaranteed latency SLA. Go-live in 8 weeks.

Questions to Surface First:

  • Message profile: order-to-trade ratio, cancel rates, burst shape at open?
  • Which symbols? Will they concentrate load on a few partitions?
  • What latency are they asking for, and can we offer it without disadvantaging other members?
  • Kill switch and risk-limit arrangements?

Typical L5 Approach: Provisions extra gateway capacity, load tests their expected rate, signs the SLA. Misses that a latency SLA for one member is a fairness question.

Staff Approach: Dedicated gateway instances (so their bursts don't queue others) with the same hardware and code path as all gateways; certification environment testing with their replayed traffic; per-session throttles and credit limits agreed in writing; the SLA offered is the same published venue SLA, not a special one.

Principal Approach: Ensures fair-access obligations: any latency-affecting product (dedicated gateways, co-location tiers) is offered on published, non-discriminatory terms. Works with compliance so the onboarding doesn't create a two-tier market, and updates capacity planning to model the new member's open-auction bursts.

Staff Approach — Full Reasoning
PhaseWhat to Do
Weeks 1–2Traffic profile; partition impact model; risk limits and kill-switch procedures signed.
Weeks 3–5Certification: replay their simulated traffic plus our open-auction replay at 5×.
Weeks 6–7Dedicated gateways (standard product); partition placement adjusted.
Week 8Go-live with conservative throttles; raise after 2 weeks of observed behavior.

Metrics to Watch: member.msg_rate{member}, gateway.queue_delay_us{gateway}, engine.core_utilization{partition}, member.order_to_trade_ratio.

Organizational Follow-up: Member onboarding checklist includes kill-switch drill and certification pass.

Ownership Question: "Who approves a member-specific SLA?" Staff answer: Nobody should — SLAs are venue-wide and published. Any latency-affecting offering goes through the product and compliance review for fair access.

Key Takeaway: "In an exchange, a special deal for one member is a fairness problem before it's a capacity problem."

What clears the Staff bar:

  • Isolates ingress load without changing sequencing fairness
  • Certification by replay
  • Recognizes fair-access constraints on SLAs

Deep Dive 4: Post-Mortem — Runaway Algorithm Moves the Market#

Context: A member's algorithm malfunctioned for 4 minutes, sending aggressive buy orders across 140 symbols. Prices moved 5–15% in several names before limit-up/limit-down pauses kicked in. Counterparties demand trades be busted.

Questions to Surface First:

  • Which controls fired, in what order, and at what seq numbers?
  • Why didn't the member's credit limit stop it sooner?
  • Did our price collars work as configured?
  • What does the rulebook say about clearly erroneous trades?

Typical L5 Approach: Reviews logs, tightens the member's limits, and busts trades that look erroneous. Reasonable, but ad hoc — busting decisions without clear criteria create new disputes.

Staff Approach: Replays the journal to produce the exact control timeline. Finds the credit limit was set on gross notional per symbol, not aggregated across symbols — 140 symbols × per-symbol limit = far too much. Fixes aggregation, adds a cross-symbol velocity check (notional per second per member), and applies the published clearly-erroneous policy mechanically to decide busts.

Principal Approach: Reviews the layered-control model market-wide: which controls are the member's obligation, which the venue's, and which the market's. Proposes default venue-side limits for all members (opt-out requires justification), mandatory kill-switch drills, and a public post-incident summary to maintain market confidence. Knight Capital's 2012 incident is the reference point for why these layers exist.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateKill switch if still active; confirm LULD pauses; stop further harm.
TriageReplay journal; build timeline of each control; identify the gap.
Quick fixAggregate credit limits across symbols; cross-symbol velocity check.
GuardrailsDefault venue limits for all members; alert on member.notional_velocity.
Post-mortemBusts per published rules; member and regulator communications.

Metrics to Watch: member.notional_velocity, member.exposure_utilization, risk.collar_rejects, luld.pauses.

Organizational Follow-up: Annual review of member risk settings; kill-switch drills as a membership condition.

Ownership Question: "Who is responsible for the losses?" Staff answer: Primarily the member whose algorithm malfunctioned — but the venue owns the gap in its own control (per-symbol vs aggregate limits). The post-mortem names both without blame, and busts follow the published rule, not negotiation.

Key Takeaway: "Controls that are per-symbol can be defeated by breadth. Measure exposure the way the damage accumulates."

What clears the Staff bar:

  • Uses journal replay for an exact timeline
  • Finds the aggregation gap
  • Applies published policy rather than case-by-case judgment

Deep Dive 5: Multi-Region Expansion — A Second Market#

Context: The company wants to launch trading in a second country with its own regulator, and asks whether to reuse the existing engine from the current data center.

Questions to Surface First:

  • Does the new regulator require local matching and data residency?
  • Latency: members in the new country are 80ms away from the current site. Is that acceptable?
  • Different trading hours, tick regimes, auction rules?
  • Shared engine codebase vs separate deployments?

Typical L5 Approach: Extends the existing matching site to serve the new market, adding symbols to existing partitions. Simple, but 80ms latency disadvantages local members and likely fails residency rules.

Staff Approach: Separate deployment in the new country — own sequencer, engines, market data, DR site — running the same engine codebase parameterized by market rules as sequenced reference data. Journals stay in-country. Shared tooling and release process.

Principal Approach: Defines the platform/market split: one engine platform team, multiple market operations teams, per-market rulebooks as configuration. Sets the release policy so a change certified for one market can't ship to another without that market's certification. Evaluates whether cross-market products (e.g., cross-listed securities) need a coordination layer — and decides not to build it until a product demands it.

Staff Approach — Full Reasoning
PhaseWhat to Do
Months 0–3Regulatory requirements; site selection; rulebook differences into config schema.
Months 3–9Build site; member certification environment; market-ops team hiring.
Months 9–12Dress rehearsals: full trading days with simulated members; DR invocation drill.
LaunchLimited symbol set; expand after 30 stable days.

Metrics to Watch: Same core metrics per market, plus release.certified_markets per build.

Organizational Follow-up: Platform vs market-ops ownership matrix; per-market change approval.

Ownership Question: "Who owns an engine bug that only affects the new market's rules?" Staff answer: The platform team owns the fix; the market's operations team owns the halt decision and member communications in that market.

Key Takeaway: "Exchanges don't go active-active across regions; they replicate the platform, not the book."

What clears the Staff bar:

  • Rejects cross-region book sharing
  • Parameterizes rules as sequenced data
  • Separates platform and market ownership

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why a sequencer plus single-threaded matching is the correct architecture, with capacity numbers
  • List the constructs forbidden inside a deterministic engine and why
  • Describe durability via replicated memory before processing, and bound the residual risk
  • Walk through failover by replay, state-hash verification, and the deterministic-bug trap
  • Name concrete pre-trade controls with microsecond budgets and a kill switch design
  • Design fair market data dissemination with multicast, A/B feeds, retransmit and snapshots
  • Treat halt-and-reopen as a designed, rehearsed procedure
  • Map engineering ownership to market operations and compliance

The Bar for This Question#

Mid-level (L4): Designs an order service with a database, matches orders with a simple loop, pushes updates via WebSocket. Doesn't address ordering between concurrent orders or recovery.

Senior (L5): Knows price-time priority and a reasonable order book structure, uses a queue to serialize orders, persists trades, adds replicas. The design would work at modest scale. Gaps: parallelizes matching or orders by timestamp, recovers from a database rather than by replay, lacks a failure policy beyond restart, and treats market data as ordinary push notifications.

Staff+ (L6): Builds everything on a sequenced, replicated input log; defends single-threaded matching with numbers; recovers and replicates by deterministic replay with state verification; recognizes that a deterministic bug kills the standby; designs inline risk with µs budgets and kill switches; publishes fair, sequenced market data; and treats halting as a correct outcome. Ownership spans engine, market ops, and compliance. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 The Fastest Exchange Architecture Is the Least Distributed One#

EvidenceDetail
Per-core capacity~1M+ in-memory matches per second on one thread
Coordination costAny cross-node agreement costs µs to ms
ReplayOnly a single ordered stream replays exactly

The Staff position: Distribute around the core (gateways, market data, analytics), never inside it.

Why this matters in interviews: It's the clearest test of whether you apply distributed patterns by reflex or by requirement.

10.2 Halting the Market Is a Feature#

EvidenceDetail
Busted tradesCost trust and money far beyond a short halt
Regulatory designLULD pauses and market-wide circuit breakers exist on purpose
Reopening auctionsRestore fair price discovery after uncertainty

The Staff position: Design halt, cancel-only and auction-reopen states as first-class, rehearsed operations.

Why this matters in interviews: "Keep trading at all costs" is web-scale availability thinking misapplied.

10.3 Hot Standbys Give You Less Protection Than You Think#

EvidenceDetail
Correlated software failureSame code + same input = same crash
Silent divergenceEnvironment drift creates nondeterminism
Verification rotComparison jobs fail quietly

The Staff position: Pair standbys with replay-in-CI, canary symbols, feature flags per order type, and verification freshness as a failover precondition.

Why this matters in interviews: Naming the deterministic-bug trap is a strong senior-to-staff differentiator.

10.4 Fairness Is the Product; Latency Is Just One Way to Sell It#

EvidenceDetail
IEX speed bump350µs deliberately added
Batch auctionsUsed at open/close by major venues
Equal-length cablingVenues spend capex to remove physical advantage

The Staff position: Every latency optimization must be evaluated for who it advantages.

Why this matters in interviews: "Faster for whom?" is a Staff question; "how do we make it faster?" is not.

10.5 Kafka Is Not Your Sequencer (For This Latency Class)#

EvidenceDetail
Kafka produce p99Typically ms-scale with durable acks
Venue budgetTens of µs order-to-ack
Where Kafka fitsDownstream: surveillance, analytics, drop copy fan-out

The Staff position: A single-partition log is the right shape; a purpose-built, kernel-bypass sequencer is the right implementation for co-located trading. For a retail or crypto venue with ms budgets, a log like Kafka can be a legitimate sequencer.

Why this matters in interviews: Shows you calibrate technology to latency class rather than rejecting it wholesale.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

At Staff level, the challenge is building a deterministic, fair, recoverable engine. At Principal level, the exchange is a regulated institution whose software is its rulebook. The engine's behavior is what the venue has promised regulators and members; every change is effectively a rule change. The Principal problem is governance of that promise: who can change matching behavior, how changes are proven identical-or-intended, how failure posture is decided and rehearsed across engineering, operations and compliance, and how the business's revenue incentives (data, co-location) are kept consistent with fair-access obligations.

The Org-Level Fault Line#

Engine velocity vs engine sanctity. Product teams want new order types, new auction mechanics and new asset classes quickly. Every change to the engine is a change to regulated behavior and a risk to determinism.

ChoiceWhat WorksWhat BreaksWho Pays
Engine as a platform with a narrow, versioned rule interfaceFeatures as configuration and plug-in order types with certified semanticsSome features wait for interface extensionsProduct teams (slower novel features)
Every product team modifies the engineFast featuresDeterminism erodes; certification load explodesMembers, regulator relationship, on-call
Fork engines per productIndependenceN engines to certify and staffHeadcount, correlated-knowledge risk

🧭 Principal Move: "The engine team owns a small core and a versioned rules interface. New order types ship as configured behaviors behind flags, certified by full-journal replay and member certification. Nobody else commits to the core."

Cost Model#

Assumptions: fully loaded engineer ~$300K/year; co-location and low-latency network as dominant infrastructure; order-of-magnitude figures.

ScaleProfileInfra $/monthHeadcountOn-call / Ops
Small (crypto or niche venue, ms latency, cloud)100 symbols, 50K msgs/sec$30K–80K8–12 engineersEngineering on-call 24×7; small ops desk
Medium (regional regulated venue, co-located)2K symbols, 500K msgs/sec$300K–800K (co-lo, network, DR site)30–60 engineers + 10–20 market opsMarket-hours ops desk, engineering follow-the-sun
Large (major national venue)10K symbols, millions msgs/sec, multiple markets$2M–6M+150–400 engineers across engine, data, surveillanceDedicated market ops, compliance engineering, DR team

The dominant costs shift from people (small) to the fairness-and-resilience envelope — co-location, redundant networks, DR, retention (large). Compute for matching itself is a rounding error at every scale.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversal Cost
Market structure (continuous vs batch vs dark)One-wayRulebook refiling, member rebuilds, liquidity migration
Order-entry and market-data protocol formatsOne-wayEvery member's software depends on them; multi-year deprecation
Sequence number and journal formatOne-wayAudit history and replay tooling bound to it
Price representation (integer ticks, precision)One-wayTouches every member and every stored record
Engine partition assignmentsTwo-wayOvernight rebalance
Risk limit valuesTwo-wayConfig change with notice
Market data conflation tiersTwo-way-ishMember notice periods

The Standard I'd Write#

RFC: Matching Engine Determinism and Change Standard (v1)

Scope: All code and configuration executed inside the matching core, and all inputs to it.

MUST:

  1. Every input to the engine MUST be a sequenced message, including time, reference data, halts and configuration.
  2. Engine code MUST NOT read wall clocks, generate randomness, use floating-point prices, or iterate unordered collections.
  3. Every release MUST replay the previous 5 trading days' journals and produce byte-identical outputs, except for explicitly declared behavior changes.
  4. New order types and matching behaviors MUST ship disabled and be enabled per symbol via sequenced configuration.
  5. Failover MUST be blocked unless the standby's state hash was verified within the last 60 seconds.
  6. Halt, cancel-only, and auction-reopen procedures MUST be rehearsed at least monthly in the test environment.

SHOULD: keep engine core under a fixed size budget; run member certification for behavior changes; publish latency percentiles per gateway.

Exceptions: Approved jointly by the engine owner and compliance; documented in the rule filing if behavior-affecting.

Success metrics: zero system-caused busted trades; 100% releases passing replay gate; failover drill success ≥ 99%; median time from order-type request to production.

What I'd Tell the VP#

Our exchange's core is deliberately simple: every order is put in a single official sequence and processed one at a time, which is what lets us prove to regulators and members that everyone was treated fairly. The risks that matter aren't raw speed — they're a software change that alters matching behavior, a backup that silently disagrees with the primary, or a member's runaway algorithm. I'm proposing three controls: every release is tested by replaying real trading days, failover is blocked unless the backup is verified, and halts are practiced monthly. The cost is modest — a few engineers and slower delivery of new order types — against the cost of a single day of busted trades, which is measured in reputation and regulatory action.

Principal Interview Signals#

SignalWhat It Sounds Like
Software is the rulebook"An engine change is a rule change; it goes through the same governance."
Correlated risk named"Our standby shares our bugs; replay-in-CI is the real protection."
Priced resilience"The DR site is ~20% of infra spend; warm instead of hot saves half and adds 15 minutes of RTO — that's a board-level trade."
Incentive awareness"Data revenue creates pressure for faster tiers; fair access constrains it."
Knows what not to build"No cross-market coordination layer until a cross-listed product needs one."

Staff answers that L7 interviewers find insufficient:

  • A perfect engine design with no answer for how the fifteenth order type ships safely
  • Failover and halt design without naming who in the organization decides and how it's rehearsed
  • Market data design that ignores the business incentive and regulatory fair-access tension

Appendices

Appendix A: Matching Mechanics in Depth#

A.1 Order Book Structure#

ComponentStructureWhy
Price levels near the touchArray indexed by (price − base) / tickO(1) access; most activity within a few hundred ticks
Distant levelsSorted map (tree)Sparse; rarely touched
Orders within levelIntrusive doubly linked FIFOO(1) append, O(1) cancel with pointer
Order lookuporder_id → node hash map (lookups only, never iterated)O(1) cancel/replace
Best bid / askCached indicesO(1) top-of-book

A.2 Matching Loop#

on_new_limit(order):                      # called in sequence order only
    book = books[order.symbol]
    opp = book.asks if order.side == BUY else book.bids
    while order.remaining > 0 and opp.best_level() crosses order.price:
        level = opp.best_level()
        resting = level.head()
        if resting.member == order.member and stp_enabled:
            apply_self_trade_prevention(order, resting); continue
        qty = min(order.remaining, resting.remaining)
        emit Executed(seq, resting, qty, level.price)     # passive order's price
        emit Executed(seq, order, qty, level.price)
        emit Trade(seq, level.price, qty)
        order.remaining -= qty; resting.remaining -= qty
        if resting.remaining == 0: level.pop_head()
        if level.empty(): opp.remove_level(level)
    if order.remaining > 0 and order.tif == DAY:
        book.side(order.side).level(order.price).append(order)
        emit BookAdd(seq, order)
    elif order.remaining > 0:                              # IOC remainder
        emit Canceled(seq, order, reason=IOC)

A.3 Cancel-Replace and Priority#

ChangeKeeps Time Priority?
Reduce quantityYes
Increase quantityNo — goes to back of queue
Change priceNo — new order at new level

Appendix B: Identity and Data Model#

EntityKeyNotes
Sequenced messageseq (u64, global)Gapless; the ordering truth
Orderorder_id (u64, assigned by engine)Member's client_order_token echoed back
Tradetrade_id (u64)Carries both order_ids and seq
Symbolsymbol_id (u32)Reference data via sequenced update
Priceinteger ticks (i64)Never floating point

Appendix C: Sequencing and Replication Mechanisms — Quick Comparison#

MechanismLatencyDurabilityDeterminismFit
Custom sequencer + replicated memory journal5–15µs2 memory copies + async diskExactCo-located venues
Raft log (e.g., etcd-style)1–5msMajority diskExactRetail/crypto venues
Kafka single partition, acks=all5–20ms p99Replicated diskExact per partitionRetail venues, downstream consumers
Database serializable transactions1–10ms+DiskExact but order is DB-chosenSmall internal crossing

C.1 Recovery by Replay#

recover(partition):
    snap = latest_snapshot(partition)          # book state at seq S, hash H
    engine.load(snap); assert engine.hash() == H
    for msg in journal.read_from(S + 1):
        engine.apply(msg)                      # outputs suppressed (already published)
    compare engine.hash() with primary / checkpoint hashes
    resume live at next seq

Appendix D: Member Protocol and Client Behavior#

ConcernBehavior
Session sequence numbersBoth directions; gap → resend request
Duplicate submissionsclient_order_token unique per session per day; duplicates rejected
ReconnectClient requests outbound replay from last seen seq
Cancel-on-disconnectOptional per session; sequenced cancels on disconnect
Throttle responseExplicit reject with reason; no silent queueing
Market data gapsArbitrate A/B → retransmit request → snapshot + replay

Appendix E: Observability#

E.1 Core Metrics#

order_ack_latency_us{p50,p99,p999,gateway,member}
stage.queue_delay_us{stage}                 # decode, risk, sequencer, engine, encode
engine.core_utilization{partition}
engine.last_processed_seq{partition}
engine.state_hash_mismatch{partition}
engine.state_hash_compare_age_s{partition}
journal.disk_lag_msgs
sequencer.epoch_changes
md.publish_lag_us{partition}
md.gap_rate{feed}
md.retransmit_requests_per_sec
member.msg_rate / member.exposure_utilization / member.notional_velocity
gateway.p99_skew_us

E.2 Critical Alerts#

AlertThresholdAction
State hash mismatchanyPage; block failover; prepare halt
Hash comparison stale> 60sPage; failover blocked
Engine stalledno seq progress 100ms under loadPage; failover or halt
Market data lag> 1sHalt partition per rulebook
Gateway skewp99 skew > 1µsDrain and investigate
Journal disk lag> 10ms of messagesPage infra
Member exposure> 90% of limitNotify market ops

E.3 Control Plane vs Data Plane#

Market-ops actions (halt, kill switch, auction triggers, reference data) are control-plane decisions but data-plane messages: they enter through the sequencer so they're ordered, replayable and auditable alongside orders.

Appendix F: Scale Evolution#

F.1 What Works at Each Scale#

ScaleDesign
< 10K msgs/sec, ms latencySingle process, Raft or Kafka log as sequencer, cloud-hosted
10K–500K msgs/secCustom sequencer, partitioned engines, hot standby, multicast in co-lo
Millions msgs/sec, multiple marketsKernel bypass/FPGA gateways, multi-market platform, DR sites per market

F.2 What You Don't Build on Day One#

  • Active-active matching across sites
  • FPGA-accelerated risk (until kernel-bypass software is measured as the bottleneck)
  • Intraday partition rebalancing
  • Cross-market coordination
  • Custom surveillance ML — start with rule-based alerts on the journal stream

Appendix G: Multi-Member Fairness and Cost#

ConcernMechanism
Ingress fairnessIdentical gateways; skew monitoring; dedicated gateways as a published product
Excessive messagingOrder-to-trade ratio policies, messaging fees
Physical proximityEqual-length cross-connects in co-location
Data accessPublished tiers; direct feed and derived feed clearly labeled
Risk isolationPer-member credit limits, throttles, kill switches
  1. Loading the index…