Hiring BarSupport

Design Ticketmaster (Flash Sales & Virtual Queues) — Staff-Level Case Study

Case study61 min read7 diagrams

Technologies referenced in this case study: Redis · PostgreSQL · DynamoDB · Apache Kafka · API Gateways

Related case studies and patterns: Dealing with Contention · Reservation Systems · Payment Processing · Rate Limiting · CDN & Edge Caching · Degraded Mode · Real-Time Updates

How to Use This Case Study#

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 3, 5
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 → Section 4 → Deep Dives 1–2
Deep Dive3+ hrsEverything, including the Principal Lens and appendices
What is a Flash Sale System? — Why interviewers pick this topic

A flash sale sells a fixed, small inventory to a demand that is 10–1,000× larger, starting at a single announced instant. Concert tickets, sneaker drops, limited game consoles, Singles' Day doorbusters. Ticketing adds a twist: inventory is not fungible — seat 14 in row F is a specific row in a database that exactly one person may buy.

The system's hardest property is that almost everyone who shows up will fail to buy. The design question is how they fail: fast and fairly, or slowly, randomly, and after hours of watching a spinner.

Before vs After — stadium on-sale (60K seats, 2M people):

Without a Staff-level design:
t=-1min:    2M browsers refresh the event page. Origin sees 400K req/s.
t=0:        Sale opens. Every user clicks a seat. 60K seats, 900K hold attempts in 10s.
t=+5s:      Seat rows lock-contend. DB p99 = 12s. Connection pool exhausted.
t=+30s:     Bots with 5,000 sessions each hold 30% of the best seats.
t=+2min:    Checkout errors; users retry, doubling load. Payment gateway rate-limits us.
t=+20min:   Inventory sold out. 1.9M users still in a broken flow, clicking retry.
t=+3h:      Users finally see "sold out". Headlines. Congressional letter.

With a Staff-level design:
t=-2h:      Waiting room opens at the CDN edge. Arrivals before t=0 get randomized positions.
t=0:        Admission controller admits ~2K users/min into the store — matched to checkout capacity.
t=+10s:     Admitted users get server-assigned 'best available' holds: single conditional write each.
t=+8min:    Holds expire for abandoned carts, flow back into inventory within seconds.
t=+25min:   Inventory reaches zero. Queue immediately shows 'sold out' to everyone not yet admitted.
t=+26min:   1.9M users told the truth within one minute of the last seat selling.

Why interviewers reach for this question: It combines extreme contention (many writers, few rows), admission control (protecting a system from its own customers), fairness (a product and PR problem as much as an engineering one), and adversarial users (bots and scalpers with better tooling than fans). There is no configuration in which everyone is happy — only a choice of who is disappointed and how.

Mechanics Refresher: Inventory Contention Strategies
StrategyHow It WorksProsCons
Pessimistic row lock (SELECT … FOR UPDATE)Lock the seat row, check, updateSimple, correctLock queues under contention; p99 explodes at ~hundreds of writers/row
Conditional write / CAS (UPDATE … WHERE status='AVAILABLE')Single atomic statement; loser gets 0 rowsNo lock waits, fast failureHot seats generate many failures users must retry
Atomic counter (Redis DECR, DynamoDB conditional decrement)For fungible stock: decrement if > 0~100K ops/s per keyDoesn't model specific seats; needs reconciliation to durable store
Sharded counters / pre-allocated bucketsSplit stock into N sub-countersScales writes N×Stranded stock in empty-vs-full shards; rebalancing
Server-assigned allocation ("best available")Server picks seats for the user from a poolRemoves the hot-seat collision entirelyUsers lose control of exact seat choice
Queue in front of inventory (single writer per section)Serialize requests per partitionZero contention by constructionThroughput = one writer's speed per partition

For most production ticketing: an admission-controlled waiting room so the inventory service sees a bounded rate, server-assigned holds via conditional writes partitioned by section, and a short hold TTL. The contention strategy matters far less once the admission rate is right. See Dealing with Contention.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What This Interview Actually Tests#

Designing Ticketmaster is not a database locking question. Everyone can prevent a double-booked seat.

It is a demand-shaping and fairness question that tests:

  • Whether you control how many people reach the inventory, rather than making the inventory survive everyone
  • Whether you define fairness explicitly — arrival order, randomness, verified identity — and who signed off on it
  • Whether you treat bots as the primary customer segment by request volume
  • Whether you design the failure experience for the 97% who will not get a ticket

The key insight: The inventory service should never see more demand than there is inventory. The waiting room converts an unbounded, adversarial stampede into a bounded, metered stream sized to checkout capacity — and it's the one component that tells 2M people the truth, fast. Everything behind it becomes an ordinary transactional system.

The L5 → L6 → L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDesigns the seat table and a locking schemeAsks "what's demand vs supply, and what's our fairness policy?"Asks "who owns the fairness policy — us, the artist, the promoter — and what are we willing to defend publicly?"
LoadScales the API and DB horizontallyPuts a waiting room at the CDN edge; admits at checkout capacity; origin sees a bounded rateTreats on-sale capacity as a scheduled product: sells capacity tiers to promoters, runs load tests per mega-event
ContentionRow locks or optimistic concurrency on seatsServer-assigned holds via conditional writes, partitioned by section; hold TTL 5–8 minChooses the inventory model (seat-level vs zone-level) per event type as a product decision
Fairness"First come, first served"Randomizes pre-sale arrivals; FIFO after; per-account limits; verified identity for mega-demandDesigns the allocation mechanism (lottery, registration, dynamic pricing) with business and legal; publishes it
BotsCAPTCHA and rate limitsLayered: edge bot scoring, device/account reputation, verified-fan codes, purchase limits, post-sale cancellationPrices the fraud: scalper markup captured vs fan trust lost; decides whether secondary market is ours
FailureRetries"Sold out" propagates to the queue in < 60s; graceful degradation to queue-only modeDesigns the public communication and refund posture before the sale, not after
Why "first move" separates levels

L5: Starts at the database — seats, holds, orders — and designs correct concurrency control. That's necessary, and it's also the easiest part. At 2M arrivals for 60K seats, the database design is irrelevant if 2M requests reach it.

L6: "Demand is 30× supply. Before I design inventory, I'll design admission: a waiting room at the edge that meters users into the store at the rate checkout can handle. Then the inventory problem is 2K holds per minute, not a million per second."

L7: "And the waiting room's ordering rule is a fairness policy that will be in the news. I'd want the promoter and legal to sign off on whether it's arrival-ordered, randomized, or registration-based before we write code."

Why "fairness" separates levels

L5: "First come, first served" — which sounds fair and rewards whoever has the lowest latency to your servers and the most sessions open: bots.

L6: Defines fairness per phase: everyone who arrives before the sale opens is equal (randomized positions); after the open, FIFO; one queue position per verified account; purchase limit per account and per payment instrument.

L7: Recognizes that when demand is 30×, no queue is fair to the 97% who lose. Proposes changing the mechanism — registration + lottery (Verified Fan-style), or demand-based pricing — and owns the public narrative.

Why "failure" separates levels

L5: Thinks about the system failing — DB down, payment errors. Doesn't think about the product failing: 1.9M users waiting three hours for inventory that was gone in 20 minutes.

L6: Designs the losing experience: queue position and estimated wait, a "remaining inventory is low" signal, and a sold-out broadcast to every queued user within a minute. Also designs hold expiry so abandoned carts return seats quickly.

L7: Owns the post-sale posture: how many tickets were cancelled as bot purchases, how refunds and re-releases work, and what the company says publicly within hours.

The Staff Positions#

PositionRationale
Admission control before inventoryMeter users in at checkout capacity; the inventory tier should never see the stampede
Waiting room lives at the CDN edge2M pollers must not touch origin; signed tokens make it stateless
Randomize pre-open arrivals, FIFO afterRefreshing faster than others before the open shouldn't be a competitive advantage
Server-assigned seats over user-picked seats at peakEliminates the hot-seat collision; users choose section and quantity, not seat
Holds with short TTL (5–8 min), lazily expiredAbandoned carts must return inventory fast; expiry is a read-time predicate, not a cron job
Bots are the majority by volume; design for themLayered bot defense + per-identity limits + post-sale cancellation
Tell losers the truth fastSold-out must reach the queue within 60s; the waiting room is also the notification system

The Four Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Reserved-seat ticketingUnique, non-fungible seats; 5–10 min checkoutWaiting room + section-partitioned holds + TTLDouble-sell (catastrophic) or stranded holdsZero double-sells; holds released < 1 min after expiry
Fungible flash sale (sneakers, consoles, doorbusters)Identical units; very high rateSharded atomic counters + async order creationOversell (refunds) or undersell (stranded shards)Oversell ≤ 0; reconciliation within minutes
Allocation under extreme demand (demand > 20× supply)Fairness perception dominatesRegistration window + lottery for purchase windowsAccusations of riggingAuditable randomness; published rules
Anti-scalpingAdversaries with thousands of accountsIdentity verification, limits, non-transferable ticketsBots win anyway, or fans face heavy frictionMeasured bot share of purchases

🎯 Staff Move: "I'll design reserved-seat ticketing for a mega on-sale — 60K seats, 2M people — because it has the strictest correctness bar and the worst demand ratio. The fungible case is a simplification: replace seat holds with sharded counters. And I'll call out where demand is so extreme that the right answer stops being a queue and becomes a lottery."

The Five Fault Lines#

#Fault LineThe Tension
1Admission Control vs Open DoorMeter users through a waiting room (bounded, fair, slower) or let everyone hit the store (simple, fast, collapses)?
2User-Picked Seats vs Server-Assigned HoldsLet users choose exact seats (UX) or assign best-available (no contention)?
3Hold Duration: Conversion vs Inventory VelocityLong holds help slow payers; short holds recycle abandoned seats
4Arrival-Order vs Randomized vs Registration FairnessWhich definition of fair — and who pays for each?
5Bot Friction vs Fan FrictionEvery anti-bot measure also costs real fans conversion

In the Wild: Real Production Systems#

Why this section belongs here: Flash sales fail in public. The real incidents are the best evidence that the hard part is admission and fairness, not locking.

Ticketmaster — The 2022 Eras Tour Presale#

Ticketmaster publicly described the November 2022 Taylor Swift presale as unprecedented: it reported roughly 3.5 billion total system requests — about four times its previous peak — and said far more users and bots showed up than the ~1.5M fans it had invited through its Verified Fan program. The general public sale was subsequently cancelled due to insufficient remaining inventory, and the episode drew congressional scrutiny.

Staff insight: Verified Fan (registration + invitation codes) is an admission-control mechanism that failed not because the idea was wrong but because non-invited traffic still reached the system at scale. The lesson for interviews: admission control must be enforced at the edge, before any origin work, and demand planning must assume invited ≠ arriving.

Cloudflare Waiting Room and Queue-it — Productized Virtual Queues#

Both Cloudflare (Waiting Room) and Queue-it sell edge-based virtual waiting rooms. They hold users at the CDN edge with a cookie/token, admit them to the origin at a configured rate or concurrency, and support different queueing methods — Cloudflare documents FIFO, random, and other modes; Queue-it popularized randomizing users who arrive before a sale's start time.

Staff insight: The fact that virtual queues are a product category tells you they're a solved, commoditized layer. A Staff answer mentions buy-vs-build explicitly and focuses the custom engineering on what's differentiated: the inventory hold model and the fairness policy.

Alibaba Singles' Day — Fungible Inventory at Extreme Rate#

Alibaba has publicly reported peak order-creation rates of around 583,000 orders per second during the 2020 Singles' Day event. Designs for that scale rely on pre-warming, inventory pre-deduction in fast in-memory tiers with asynchronous persistence, and extensive rehearsal (full-link load testing) ahead of the event.

Staff insight: Fungible inventory scales by splitting counters and making order creation asynchronous; seat inventory can't split that way. Naming the difference shows you understand why ticketing is harder per unit than e-commerce.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Lock the seat row""900K people click in 10 seconds. What's your p99?"Contention awareness
"Add a queue""Where does the queue live, and how do 2M clients know their position without hitting origin?"Edge architecture
"First come, first served""People refreshing since 9:58 vs a bot with 5K sessions — who wins?"Fairness definition
"Hold for 15 minutes""Half of holders abandon. How much inventory is stranded at minute 10?"Hold economics
"CAPTCHA""Bots solve CAPTCHAs for $1 per thousand. What else?"Adversarial depth
"Sold out""How long until the 1.9M queued users find out?"Losing-user experience

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Everything left of the admission controller runs at the CDN edge and scales with the CDN, not with us. The admission controller is the only coupling: it reads remaining inventory and checkout health, and decides how many queued users to admit per minute. Admitted users carry a signed token that the storefront requires. Behind that gate, the system handles thousands of users per minute — a normal transactional workload — with holds as conditional writes on a section-partitioned seat store.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Load"Autoscale the API""Waiting room at the edge; admit at checkout capacity; origin never sees the stampede."
Contention"Row locks""Server-assigned best-available holds via conditional UPDATE … WHERE status='AVAILABLE', partitioned by section."
Holds"Hold 15 minutes""Hold 5–8 minutes, expiry evaluated at read time; one extension allowed during payment."
Fairness"FIFO""Randomize everyone who arrives before open; FIFO after; one position per verified account."
Bots"CAPTCHA""Edge bot scoring, account/device reputation, verified-fan codes, per-account and per-card limits, post-sale cancellation."
Sold out"Show an error at checkout""Admission controller stops at zero and broadcasts sold-out to the queue within 60s."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Stadium capacity50–70K seatsThe supply side of the ratio
Mega on-sale demand1–3M+ concurrent users20–50× supply
Eras Tour presale (public)~3.5B system requests, ~4× prior peakPlan for traffic far above invites
Waiting-room poll interval15–30s2M users / 20s = 100K req/s at the edge — CDN-scale, not origin-scale
Average checkout duration4–8 minSets admission rate: concurrency ÷ duration
Hold TTL5–8 min (+1 extension at payment)Balances slow payers vs stranded inventory
Cart abandonment during on-sales~20–40% of holdsWhy hold expiry speed matters
Conditional write on a single row~1–5msSection-partitioned throughput ~1–2K holds/s per partition
Pessimistic lock collapse~hundreds of waiters per rowWhy "SELECT FOR UPDATE" dies at peak
Redis atomic counter~100K ops/s per keyFungible-inventory ceiling per shard
Payment gateway ceilingoften pre-negotiated, e.g., 500–2K TPSFrequently the real bottleneck; admission must respect it
Sold-out propagation target< 60s to all queued usersThe losing-user SLO
Bot share of on-sale trafficcommonly the majority of requestsDesign for adversaries first

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on the seat schema and locking, and 2 minutes on "we could add a queue." The queue is the design. Flip the ratio.


Phase 1: Requirements & Framing (2–3 minutes)#

Functional scope in one breath:

"Users browse events, pick a section and quantity, get seats held for a few minutes, pay, and receive tickets. Organizers configure events and on-sale times."

Then the non-functionals that matter:

"The defining property is the ratio: a stadium has 60K seats and a mega on-sale draws 2M people at the same instant — 30× demand. Most of them will lose. So the design is about three things: controlling how many people reach inventory, never double-selling a seat, and making the process defensibly fair — including against bots, which will be most of the traffic."

Commit to numbers:

"60K seats, 2M concurrent users arriving over the 30 minutes before open, 4–8 minute checkout, payment gateway limited to ~1K TPS. Zero double-sells. Sold-out visible to everyone within a minute."

🎯 Staff Move: Say the demand ratio out loud. "30× demand means 97% of users will fail. I'm designing how they fail as carefully as how the 3% succeed."


Phase 2: Core Entities & API (1–2 minutes)#

  • Event (event_id, venue_id, onsale_at, per_account_limit, fairness_mode)
  • Seat (event_id, section, row, number, status: AVAILABLE | HELD | SOLD, hold_id, expires_at, version)
  • Hold (hold_id, user_id, seats[], expires_at, extended: bool)
  • QueueTicket (event_id, user_id, position, issued_at, signature) — stateless, signed
  • Order (order_id, hold_id, idempotency_key, payment_status)

API:

GET  /events/{id}                          → cached static page (CDN)
POST /events/{id}/queue                    → { queue_token, position, est_wait }   (edge)
GET  /events/{id}/queue/status             → { position, state: WAITING|ADMITTED|SOLD_OUT }  (edge, poll 20s)
POST /events/{id}/holds  { section, qty }  → { hold_id, seats[], expires_at }   (requires admission token)
POST /holds/{id}/checkout { payment, idempotency_key } → { order_id, status }
DELETE /holds/{id}                         → release

🎯 Staff Move: "The user requests a section and a quantity, not specific seat IDs, at peak. The server picks best-available. That single API decision removes most of the contention problem."


Phase 3: High-Level Architecture (≤5 minutes)#

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk it in 60 seconds:

  1. Users land on a CDN-cached event page; bot scoring runs at the edge.
  2. Before the open, joining the waiting room issues a signed queue token. Everyone who joined before onsale_at gets a random position; after that, FIFO.
  3. The admission controller admits N users per minute — N derived from checkout capacity and remaining inventory — by advancing an "admitted up to position P" watermark.
  4. Admitted users get an admission token (signed, 10-minute TTL) and can create holds: server-assigned best-available seats via conditional writes in their section's partition.
  5. Checkout pays with an idempotency key; on success seats become SOLD; on expiry/failure they return to AVAILABLE.
  6. When remaining inventory hits zero, the controller flips the queue to SOLD_OUT.

🎯 Staff Move: "Everything to the left of the admission controller scales with the CDN. Everything to the right sees a few thousand users a minute. That split is the design."


Phase 4: Transition to Depth (1 minute)#

"Three deep dives: the waiting room and admission math, the hold model under contention, and fairness plus bots. I'd start with admission because if that's wrong, nothing behind it matters. Which is most interesting?"


Phase 5: Deep Dives (25–30 minutes)#

Deep dive 1: Admission math (8–10 min)

"Admission rate is set by the tightest downstream constraint:"

Checkout concurrency budget:   payment gateway 1,000 TPS; ~1 payment attempt per buyer
                               → not the constraint at our rates
Inventory:                     60,000 seats, avg 2.5 tickets/order → ~24,000 orders
Target on-sale duration:       ~20–30 min (don't sell out in 90s — it looks rigged)
Admission rate:                24,000 orders / 25 min ≈ 1,000 buyers/min
                               × 1.3 overbooking for abandonment ≈ 1,300 admits/min
Concurrent shoppers:           1,300/min × ~6 min checkout ≈ 8,000 active in store
Store/inventory load:          8,000 users × a few requests/min → ~1–2K req/s. Boring.

"The controller also watches live signals — checkout error rate, hold success rate, remaining inventory — and slows admission if any degrade. As inventory approaches zero, it stops admitting more users than there are plausible seats for: no point letting 8,000 people into a store with 400 seats left."

Deep dive 2: Holds and contention (8–10 min)

"Seats are partitioned by (event, section) — a stadium has ~100–200 sections, so ~300–600 seats each. A hold request for 'section 112, qty 2' goes to that section's partition and runs one statement:"

UPDATE seats SET status='HELD', hold_id=:h, expires_at=now()+interval '6 min', version=version+1
WHERE event_id=:e AND section=:s
  AND (status='AVAILABLE' OR (status='HELD' AND expires_at < now()))
  AND seat_id IN (:best_available_candidates)
RETURNING seat_id;

"If fewer rows than requested come back, we release them and try the next best block. Expiry is a predicate, not a cron job — an expired hold is available the instant it expires. A sweeper still runs every 30s for cleanliness and metrics. With admission at ~1,300 users/min, even the hottest section sees a handful of concurrent writers."

Deep dive 3: Fairness and bots (6–8 min)

"Fairness policy: arrivals before open are equal — randomized positions — so refreshing at 9:59:59.9 gives no advantage. After open, FIFO. One queue token per verified account; tokens are bound to a device fingerprint and session, so sharing or reselling positions fails. Bots: scored at the edge on IP reputation, device signals and behavior; low-trust sessions get challenges or a lower-priority position. Per-account limit of 4 tickets and per-payment-instrument limit of 4 catch farms at checkout. For mega-demand events, the right answer is registration in advance, verified identity, and randomized invitation codes — admission control days before the on-sale."


Phase 6: Wrap-Up (2–3 minutes)#

"The design principle: the inventory service never sees more demand than inventory. The waiting room at the edge meters users in at checkout capacity with an explicit fairness rule; behind it, holds are conditional writes on section partitions with short, lazily-expired TTLs; and the queue tells losers the truth within a minute of selling out."

Organizational closer:

"The fairness policy isn't mine to set alone. Randomized vs FIFO vs registration is a decision for the promoter and our legal and comms teams, and it should be published before the sale. The engineering job is to make whatever policy they choose enforceable and auditable."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
15 min on lockingCompares pessimistic vs optimistic in depthStates conditional writes in one sentence; moves to admission
Queue as an afterthought"We could add a queue if needed"Leads with the waiting room and its math
Queue at originKafka in front of the APIWaiting room at the CDN edge; origin never sees 2M users
No fairness definition"FIFO"Randomized pre-open, FIFO after, one token per verified account
No losing experienceError at checkoutSold-out broadcast to the queue within 60s
Bots as a footnoteCAPTCHALayered bot defense; limits enforced at checkout

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Ticketmaster looks like a concurrency problem, which is why it's a good level discriminator. L5 candidates solve concurrency correctly and stop. Staff candidates notice that the concurrency problem only exists because the system let 2M people at 60K rows — and that the real design choices are about admission, fairness, adversaries and communication, each of which has a non-engineering owner.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"At 10:00:00 the sale opens. A fan has been on the page since 9:40; a bot operator has 5,000 sessions that joined at 9:59:58. Walk me through, second by second, who gets to the seats first in your design — and defend it to the fan."

This question forces the candidate to define fairness, place the queue, handle identity and bots, and speak to the human cost — all in one answer.


2. Problem Framing & Intent#

2.1 The Four Intents — Explained#

Reserved-seat ticketing → correctness + admission

  • Constraint: each seat sold once; many users want the same few premium sections
  • Strategy: waiting room, section-partitioned conditional holds, short TTL, idempotent checkout
  • Failure mode: double-sell (catastrophic, rare with conditional writes), stranded holds (common, subtle)
  • Who pays for imperfection: fans (lost seats), support (disputes), the promoter (bad press)

Fungible flash sale → throughput + reconciliation

  • Constraint: identical units, very high request rate (e.g., 100K+ attempts/s for 10K units)
  • Strategy: sharded atomic counters in memory, asynchronous order creation, reconciliation to durable store; waiting room still recommended
  • Failure mode: oversell when counters and DB disagree; undersell when stock is stranded in empty shards
  • Who pays: customer support (cancellations), finance (refunds)

Allocation under extreme demand → mechanism design

  • Constraint: demand > 20× supply; any first-come mechanism looks rigged
  • Strategy: registration window (days), identity verification, randomized selection for purchase windows, staggered access by cohort
  • Failure mode: uninvited traffic still overwhelms the system; perception of opacity
  • Who pays: the brand; the artist's relationship with fans

Anti-scalping → adversarial economics

  • Constraint: scalpers profit 2–10× face value; they will spend real money on tooling
  • Strategy: identity-bound tickets, per-identity limits, bot scoring, delayed ticket delivery, controlled resale
  • Failure mode: fans face friction while bots adapt
  • Who pays: fans (friction), trust & safety (arms race), business (resale margin decisions)

🎯 Staff Move: "I'll design for reserved seating at 30× demand, and I'll flag the threshold where we should stop running a queue and start running a lottery. That threshold is a business decision; I'll give them the data."

2.2 When NOT to Build a Waiting Room#

  • Demand ≤ ~2× supply, or spread over days. A normal inventory system with conditional writes and a rate limiter is enough.
  • Fungible inventory with a fast counter and moderate demand (e.g., 10K units, 50K buyers over 10 minutes). Sharded counters handle it; a waiting room adds friction for little benefit.
  • You can buy it. Edge waiting rooms are a commodity product. Build only the admission policy integration and the inventory model.
  • Demand is so extreme that a queue is theater (100× supply). Replace the queue with registration + lottery; the queue then serves only the invited cohort.

2.3 What the Interviewer Leaves Underspecified#

  • Seat selection UX. Pick exact seats vs best-available? Changes contention entirely.
  • Demand ratio and arrival curve. 2× vs 30× vs 100× yields three different designs.
  • Payment constraints. Gateway TPS limits and 3-D Secure flows add minutes to checkout.
  • Fairness definition. Nobody will tell you; you have to propose one and name the owner.
  • Resale. Are tickets transferable? Is there an official secondary market?
  • Presales and tiers. Fan club, credit-card partner and general sale windows — multiple queues on one inventory.

2.4 Precise Terminology#

TermWhat It MeansWhy the Precision Matters
Waiting roomEdge-hosted holding area issuing queue tokensDistinct from a message queue; it's user-facing
AdmissionGranting a user access to the storefrontThe throttle that shapes all downstream load
Admission watermark"All positions ≤ P are admitted"Makes admission O(1) state, not per-user pushes
HoldTemporary exclusive claim on seats with TTLNot a sale; must expire reliably
Best-availableServer-selected seats in a requested sectionRemoves user-driven hot-seat contention
Oversell / undersellSelling more than inventory / stranding inventory unsoldBoth are failures; only one is catastrophic
Verified fan / registrationPre-sale identity check and invitationAdmission control days before the sale
Stranded inventorySeats held by abandoned carts or stuck processesThe common, silent failure

3. The Five Fault Lines#

3.1 Fault Line 1: Admission Control vs Open Door#

The tension: A waiting room adds friction, a new component and a fairness policy to defend. An open door is simpler and faster for users — until demand exceeds what the store can serve, at which point it's slower and less fair for everyone.

StrategyWhat WorksWhat BreaksWho Pays
Open door + autoscalingNo friction; fine at ≤ 2× demandContention collapse at the inventory tier; retries amplify; bots win the raceEveryone at peak; fans most
Queue at origin (Kafka/SQS in front of hold API)Serializes writes2M clients still connect to origin; status polling hits origin; users don't know their positionOrigin fleet; on-call
Edge waiting room with signed tokensOrigin sees only admitted users; positions and status served from CDNNew component; fairness policy must be definedPlatform team (build or buy), fans wait visibly
Registration + randomized invitationsDemand shaped days in advanceUninvited traffic still shows up; complex product flowProduct and marketing; fans who didn't register

Staff default: Edge waiting room for any event whose forecast demand exceeds ~3× checkout capacity in the first 15 minutes; registration + invitation on top for mega events (> 20× supply).

Waiting-room mechanics that keep it stateless:

  • On join, the edge issues token = sign(event_id, user_id, device_hash, position, issued_at). Position comes from a Redis INCR (post-open) or a random draw at open time (pre-open cohort).
  • The controller publishes a single number — admitted_through_position — to the edge every 5–10s (edge KV or a cached JSON with 5s TTL).
  • Status polls are answered at the edge by comparing the token's position to the watermark. 2M users polling every 20s = 100K req/s, all served from cache.
  • When admitted, the user exchanges the queue token for an admission token (signed, 10-minute TTL) that the storefront verifies without a lookup.
Diagram: 3.1 Fault Line 1: Admission Control vs Open Door

🎯 Staff Move: "The waiting room's job is to make 2M users cost the CDN, not us. The only thing that crosses into origin is one number — the admitted-through watermark — and the users it admits."


3.2 Fault Line 2: User-Picked Seats vs Server-Assigned Holds#

The tension: Users want to choose their exact seats. At peak, 5,000 people click the same 200 front-row seats in the same second, and 4,800 of them fail.

StrategyWhat WorksWhat BreaksWho Pays
User-picked seat IDsBest UX, familiarCollisions on hot seats; users retry repeatedly; seat map staleness makes it worseFans (repeated failures), inventory tier (retry load)
Server-assigned best-available in chosen sectionOne attempt usually succeeds; contention spread by algorithmLess control for usersFans who wanted a specific seat
Hybrid: best-available during the first N minutes, pick-your-seat afterwardControls the peak, restores choice when calmTwo code paths; product must communicateProduct team
Zone/GA inventory (count per zone, seats assigned later)Fungible counters; massive throughputRequires assigned seating to be decided later (e.g., at delivery)Fans (don't know seat until later)

Staff default: Hybrid — best-available for the first ~15–30 minutes of a high-demand on-sale, seat map picking afterward. The best-available algorithm randomizes among equally-good blocks so concurrent requests land on different rows.

Seat map freshness: Serve the seat map from a cache (1–2s stale). A stale map causes collisions only for user-picked flows; best-available doesn't depend on it. Staff candidates say this explicitly.

🎯 Staff Move: "The biggest contention fix isn't a lock — it's an API change. If the user asks for 'two in section 112' instead of 'seats F14 and F15', the server can spread concurrent requests across the section and nearly every request succeeds on the first try."


3.3 Fault Line 3: Hold Duration — Conversion vs Inventory Velocity#

The tension: Long holds help slow payers (3-D Secure, entering card details on mobile). Short holds recycle abandoned seats faster.

Back-of-envelope:

Admitted:              1,300 users/min
Abandonment:           30% of holds never check out
Hold TTL 15 min:       at steady state, 1,300 × 0.3 × 15 ≈ 5,850 abandoned holds × 2.5 seats
                       ≈ 14,600 seats (24% of the venue) stranded at any moment
Hold TTL 6 min:        1,300 × 0.3 × 6 × 2.5 ≈ 5,850 seats (~10%) stranded
StrategyWhat WorksWhat BreaksWho Pays
Long TTL (15 min)Slow payers succeed~25% of inventory stranded; "sold out" then "seats available" flappingQueued fans; perception of fairness
Short TTL (5–6 min)Fast inventory recycleUsers with 3DS challenges time out mid-paymentSlow payers; support
Short TTL + one extension at payment startMost users covered; abandoners recycle fastExtension logic; abusable if unboundedMinimal
Adaptive TTL (shorter as inventory depletes)Maximizes velocity near sell-outHarder to explainProduct must sign off

Staff default: 6-minute hold, with a single 3-minute extension granted when payment is initiated (a real payment attempt, not a click). Expiry is evaluated at read time (expires_at < now() is AVAILABLE) so no seat waits for a sweeper.

🎯 Staff Move: "Hold duration is inventory velocity. At a 15-minute hold, a quarter of the stadium is sitting in abandoned carts. I'd rather give a 3-minute extension to people who actually started paying."


3.4 Fault Line 4: Arrival-Order vs Randomized vs Registration Fairness#

The tension: Every fairness definition favors someone.

MechanismWhat WorksWhat BreaksWho Pays
Pure FIFO by arrivalIntuitive; rewards effortRewards low latency and many sessions — bots; incentivizes arriving hours early and refreshingFans without fast connections or automation
Randomize pre-open cohort, FIFO afterRemoves the refresh race; arriving 30 min early = arriving 1 min earlyFans who waited an hour feel cheated if they're placed behind someone who came lateEarly arrivers (perception)
Registration + lotteryControls demand in advance; verifies identity; defensibleLonger process; "I registered and didn't get picked"Unselected fans (but they know early)
Tiered presales (fan club, card partner)Rewards loyalty; monetizes partnershipsPerceived as pay-to-playFans without the tier
Dynamic pricingReduces scalper arbitrage; allocates by willingness to payPublic backlash when prices spikeFans with less money; brand

Staff default: Randomized pre-open cohort + FIFO after, one position per verified account, for most high-demand events. Registration + lottery above ~20× demand. Tiers and pricing are business decisions — the platform supports them as configurable fairness modes per event.

Diagram: 3.4 Fault Line 4: Arrival-Order vs Randomized vs Registration Fairness

🎯 Staff Move: "'First come, first served' sounds fair and rewards whoever has the most bots. I'd make everyone who arrives before the open equal, and I'd want the promoter to sign the fairness mode for each event — because it will be on the news either way."


3.5 Fault Line 5: Bot Friction vs Fan Friction#

The tension: Every anti-bot control costs real fans something — a challenge, a verification step, a lower limit, a delay. Bots adapt; fans churn.

ControlStopsFan CostBot Adaptation
CAPTCHANaive scriptsSeconds; accessibility issuesSolving farms at ~$1–3 per 1K
Edge bot scoring (IP reputation, TLS/device fingerprint, behavior)Datacenter IPs, headless browsersNear zero for most; false positives hurtResidential proxies, real devices
Account age/reputationFreshly created account farmsNew fans deprioritizedAged account markets
Verified identity / phoneMass accountsMinutes of setupSIM farms (expensive)
Per-account + per-card + per-address limitsBulk purchasingGroups must split purchasesMany identities, many cards
Identity-bound, delayed-delivery ticketsResale arbitrageTransfers are restrictedAccount selling
Post-sale cancellation of bot ordersBots that got throughRisk of false cancellationsBetter mimicry

Staff default: Layered, with friction proportional to risk: low-risk sessions see almost nothing; medium-risk get challenges; high-risk get queued behind others or blocked. Hard limits enforced at checkout (where identity and payment instrument are known), not just at the queue.

🎯 Staff Move: "I don't try to stop every bot at the door — that makes fans pay. I push friction to where it's cheap for fans and expensive for bots: limits bound to payment instruments and identity at checkout, and cancellation after the sale when we can see patterns."


4. Failure Modes & Operational Reality#

4.1 The Stampede Past the Waiting Room#

Scenario: The waiting room is configured on /events/123, but the hold API endpoint (/api/holds) wasn't protected. A scalper group discovers it and calls it directly.

t=0:       Sale opens. Waiting room admitting 1,300/min correctly.
t=+10s:    Direct calls to /api/holds: 40K req/s from 20K IPs.
t=+15s:    Section partitions saturate; conditional writes contend; p99 hold latency 8s.
t=+30s:    Admitted fans' hold requests time out. Bots hold 12K seats.
t=+2min:   Alert: holds_created_per_min = 18,000 vs admitted_per_min = 1,300.
t=+4min:   Store API enforces admission token on every endpoint (config push).
t=+30min:  Bot-held seats expire (no checkout) or get cancelled post-sale.

Detection: holds_created_per_min / admitted_per_min ratio (should be ≤ ~1.2), requests_without_admission_token_total, hold latency p99.

Mitigation: Enforce admission token verification on every storefront endpoint; revoke holds created without a valid token.

Prevention: Admission token check lives in the gateway for the whole store route namespace, not per endpoint; pre-sale penetration test of the storefront by the security team.

Owner: Platform (gateway config); security (pre-sale review).


4.2 Stranded Inventory — "Sold Out" Then "Available"#

Scenario: Hold TTL is 15 minutes and expiry is done by a cron job every 5 minutes. At minute 20, inventory shows zero; the queue flips to SOLD_OUT; 1.9M users leave. At minute 30, 9,000 seats return from expired holds.

Detection: inventory_available + inventory_held + inventory_sold = capacity reconciled every minute; held_expired_not_released gauge; soldout_flip_count (SOLD_OUT → AVAILABLE transitions).

Mitigation: Distinguish "no available seats" from "sold out": show "all seats are currently held — X seats may return" and keep the queue alive until held == 0.

Prevention: Read-time expiry predicate; short TTL; controller declares SOLD_OUT only when available == 0 AND held == 0, otherwise "LAST CHANCE" mode that admits at the rate seats are released.

Owner: Inventory service team; product owns the UX copy.


4.3 Payment Gateway Becomes the Bottleneck#

Scenario: Admission rate is tuned for store capacity, but the payment provider's contracted TPS is 500 and 3-D Secure adds 60–90s per payment.

t=+5min:   Checkout attempts 900/s at peak minute. Gateway returns 429 above 500 TPS.
t=+6min:   Users retry payment; holds tick toward expiry.
t=+10min:  Holds expire mid-payment: 'you lost your seats' — for users who were paying.

Detection: payment_attempts_per_sec vs contracted TPS, payment_429_total, hold_expired_during_payment_total.

Mitigation: Admission controller includes payment headroom as an input and slows admission; auto-extend holds when payment is in flight; queue payments with a spinner rather than failing.

Prevention: Pre-negotiate burst TPS with the provider for scheduled on-sales; multi-provider routing; see Payment Processing.

Owner: Payments team (provider contracts); platform (admission inputs).


4.4 Double Sale Through a Retry#

Scenario: A checkout request times out at the client after the charge succeeded; the client retries; a second charge and a second order are created for the same hold.

Detection: orders_per_hold > 1, duplicate_charge_total, reconciliation against the payment provider's settlement report.

Mitigation: Refund duplicates automatically.

Prevention: Idempotency key per checkout attempt (derived from hold_id); order creation is a conditional transition HELD → SOLD on the seats plus a unique constraint on orders(hold_id); the payment call carries the same idempotency key to the provider.

Owner: Checkout team.


4.5 Queue Position Leakage and Resale#

Scenario: Scalpers sell admitted positions: they join with thousands of sessions, then sell the admission URL to fans for $50.

Detection: Admission token used from a device/IP different from the one that joined; admission_token_device_mismatch_total.

Mitigation: Bind queue and admission tokens to device hash and account; reject mismatches.

Prevention: Require login before joining the queue for high-demand events; one token per account.

Owner: Trust & safety with platform.


4.6 The Hot Section#

Scenario: Floor sections (GA pit, front sections) receive 60% of requests but hold 5% of seats.

Detection: Per-partition write rate and conflict rate; hold_retry_count by section.

Mitigation: When a section's conflict rate exceeds ~20%, return "section nearly full — here are alternatives" instead of retrying; reduce requests reaching that partition via UI (grey out sections below a threshold, updated from cache every 2s).

Prevention: Split hot sections into sub-partitions by row range; for GA floor, use a counter (fungible) instead of seat rows.

Owner: Inventory service team.


4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Waiting room bypassholds/admitted ratio > 1.2Whole on-sale fairnessEnforce admission token at gatewayPlatform + security
Stranded holdsreconciliation drift; SOLD_OUT flip-flopPerceived fairness, lost salesRead-time expiry; SOLD_OUT only when held = 0Inventory team
Payment ceilingpayment 429s; holds expiring mid-paymentAll admitted buyersAdmission slowdown; hold auto-extendPayments + platform
Duplicate ordersorders_per_hold > 1Individual buyers; financeIdempotency, unique constraintsCheckout team
Position resaletoken device mismatchFairnessToken binding, login-requiredTrust & safety
Hot sectionpartition conflict rate > 20%Users requesting that sectionAlternatives, sub-partitioningInventory team
Edge/CDN issuewaiting-room error rateEntire on-saleFail-closed to static "queue paused" page; do not open the doorPlatform + CDN vendor
Admission controller downwatermark not advancingSale stalls (safe)Freeze admission, manual watermark advancePlatform on-call

🎯 Staff Insight: The waiting room must fail closed. If the edge or controller breaks, the safe state is "queue paused" — not "let everyone in." This is the opposite of an abuse rate limiter, which fails open. The difference is that here, opening the door creates the very stampede the system exists to prevent.


5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Load shapingAutoscale API and DBEdge waiting room; admission rate from checkout and payment capacityOn-sale capacity as a product with tiers; per-event load tests and go/no-go
ContentionRow locks / optimistic concurrencyServer-assigned holds via conditional writes, section partitions, read-time expiryInventory model (seat vs zone vs lottery) chosen per event type with business
FairnessFIFORandomized pre-open cohort, FIFO after, token binding, per-account limitsDesigns and publishes the allocation mechanism; owns the public narrative
AdversariesCAPTCHALayered bot defense, limits at checkout, post-sale cancellationPrices scalping; decides secondary-market strategy
FailureRetries, replicasFail-closed waiting room; SOLD_OUT only when held = 0; payment-aware admissionPre-sale communication plan; refund and re-release policy
OwnershipEngineering-onlyNames promoter, legal, payments and T&S as decision ownersRedraws who owns on-sale readiness across the company

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Leads with admission"The inventory service should never see more demand than there's inventory for."
Quantifies admission"~1,300 admits a minute: 24K orders over 25 minutes with 30% overbooking."
Removes contention via API design"Users ask for section and quantity; the server picks seats."
Defines fairness and its owner"Randomize everyone who arrives pre-open. The promoter signs the fairness mode."
Designs for losers"Sold-out reaches every queued user within 60 seconds."
Knows fail-closed vs fail-open"The waiting room fails closed — opening the door is the outage."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
All depth spent on locking strategiesSolves the easy part; ignores admission
Queue behind the origin API2M clients still hit origin for status
"First come, first served" with no qualificationUnaware it rewards bots
15-minute holds without mathStrands a quarter of inventory
No idempotency at checkoutDouble charges during retries
No mention of botsIgnores the majority of traffic

5.4 Common False Positives#

  • Distributed lock expertise ≠ ticketing design. Redlock trivia doesn't address admission or fairness.
  • "Use Kafka for the queue" ≠ waiting room. A message queue doesn't give users positions or keep them off origin.
  • Elaborate seat-map rendering ≠ scale. The seat map is a cached read; it's not where on-sales fail.
  • Mentioning dynamic pricing ≠ product judgment. Without naming who decides and the backlash cost, it's a buzzword.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–4 minDemand ratio, correctness bar, fairness as a requirement
Entities + API4–6 minSection + qty holds, signed tokens
High-level design6–11 minEdge waiting room, admission, inventory, checkout
Deep dive: admission11–20 minMath, watermark, fail-closed, sold-out
Deep dive: holds20–30 minConditional writes, partitions, TTL, idempotent checkout
Deep dive: fairness + bots30–40 minRandomized cohort, binding, limits, registration
Wrap-up40–45 minOwners of fairness policy, evolution

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Direction
"Users must pick exact seats."Contention under user choiceSeat-level conditional writes; stale map tolerance; fast-fail with alternatives
"What about GA floor tickets?"Fungible inventoryCounter per zone; sharded if hot; reconcile
"Now 50 concurrent on-sales, one Saturday."Multi-tenant capacityPer-event admission budgets; shared checkout and payment capacity allocation
"The queue has 3M people and 60K seats."Mechanism designRegistration + lottery; tell people early
"The payment provider goes down mid-sale."Degraded modePause admission; extend holds; fail over provider
"Artists want to stop scalpers."Policy + engineeringIdentity-bound tickets, controlled resale, price caps on resale

6.3 What to Deliberately Skip#

  • Event search and discovery. Standard cached reads; see Search Indexing if asked.
  • Ticket delivery (PDF, wallet). Async post-purchase job.
  • Venue map authoring. Admin tooling.
  • Email/SMS confirmations. Mention Notification System and move on.

6.4 Follow-Up Questions to Expect#

  1. "How does a user know their position without hammering you?" — (Signed token + a published watermark served from the edge.)
  2. "What if the Redis sorted set holding positions dies?" — (Positions are in signed tokens; only the watermark is state — reconstructible; pause admission while restoring.)
  3. "Why not just scale the database?" — (Contention is on specific rows; more replicas don't help writes to one seat.)
  4. "How do you prevent one user holding 50 seats?" — (Per-account hold limit enforced at the hold API; per-card limit at checkout.)
  5. "How do you know the sale was fair?" — (Audit log of randomization seed, positions, admissions; publishable summary.)
  6. "What's the SLO?" — (Hold p99 < 500ms for admitted users; sold-out propagation < 60s; zero double-sells.)
  7. "How do you test this?" — (Synthetic on-sale load test with millions of virtual users against the edge before every mega event.)

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design Ticketmaster."

Staff Answer

"The number that drives this design is the demand ratio: a stadium has ~60K seats, and a mega on-sale draws ~2M people at the same instant. 97% will fail. So I'll design three things: admission — an edge waiting room that meters people into the store at the rate checkout can handle; inventory — holds as conditional writes on section partitions with short TTLs, so no seat is ever double-sold; and fairness — an explicit, auditable rule for who gets in first, robust to bots, which will be most of the traffic. I'll also design how the losers find out: sold-out should reach the queue within a minute. Let me sketch the API, then go deep on admission."

Why this is L6:

  • Leads with the ratio and the admission problem
  • Treats fairness and the losing experience as requirements
  • Orders the deep dives by risk

What L7 adds:

  • "The fairness rule is a promoter and legal decision; I'd make it a per-event configuration and publish it."
  • Notes the threshold (~20×) where a lottery beats a queue
❌ Common L5 Trap

"We need a seats table with status. To prevent double booking, we'll use SELECT FOR UPDATE, or optimistic locking with a version column. Then we'll scale with read replicas and a cache for the seat map..."

Why this misses: It's correct for a hotel booking system. At 2M arrivals for 60K seats, the lock strategy is irrelevant — the stampede collapses the API and the database before locking semantics matter.


Drill 2: Admission Math#

Prompt: "How many users do you let into the store per minute?"

Staff Answer

"Start from inventory and checkout: 60K seats at ~2.5 per order is ~24K orders. I want the sale to last ~25 minutes — selling out in 90 seconds looks rigged and overloads everything — so ~1,000 successful buyers/min. With ~30% abandonment, admit ~1,300/min. At ~6 minutes per checkout, that's ~8K concurrent shoppers — a few thousand req/s, trivial. Then cap by the tightest dependency: payment TPS with 3DS, inventory write rate per hot section. The controller adjusts every 30s from hold success rate, checkout errors, and remaining inventory, and throttles down as inventory approaches zero so we don't admit 8K people for 400 seats."

Why this is L6:

  • Derives admission from inventory and business intent (sale duration)
  • Accounts for abandonment
  • Makes the rate adaptive to downstream health

What L7 adds:

  • Turns "sale duration" into a promoter-facing setting with tradeoffs explained
  • Uses historical on-sales to calibrate abandonment by artist/genre

Drill 3: Make the Hold Concrete#

Prompt: "Walk me through holding two seats in section 112."

Staff Answer

"Request: admission token, section 112, qty 2. The gateway verifies the token signature and per-account hold limit. The inventory service routes to the partition for (event, section 112). The best-available algorithm proposes candidate adjacent pairs, randomized among equally-good rows so concurrent requests diverge. One conditional UPDATE sets them to HELD with hold_id and expires_at = now + 6 min only if they're AVAILABLE or their hold has expired. If we get 2 rows, done; if 1, release it and try the next candidate; after 3 attempts, return 'section nearly full' with alternatives. Holds count toward the account's limit. At checkout, a conditional transition HELD(hold_id) → SOLD with an idempotency key; a unique constraint on orders(hold_id) prevents double orders."

Why this is L6:

  • Every write is a single conditional statement — no lock waits
  • Handles partial success and bounded retries
  • Idempotent checkout closes the double-sale path

What L7 adds:

  • Defines inventory correctness invariants (available + held + sold = capacity) as a continuously reconciled SLO
  • Standardizes the hold primitive for other inventory products (merch, parking, add-ons)

Drill 4: Fairness Defense#

Prompt: "A fan who waited 40 minutes got position 800,000. Someone who arrived at 9:59 got position 12. Defend that."

Staff Answer

"Everyone who arrived before the sale opened was treated equally — positions were randomized at 10:00. We do that because arrival-order before the open rewards whoever can refresh fastest or run the most sessions, which is overwhelmingly bots, and it forces fans to sit on the page for hours. We tell people this in advance: 'arriving early doesn't improve your position; arriving before 10:00 does guarantee you're in the randomized group.' After 10:00 it's first-come. It's a policy the promoter approved, and the randomization is logged so we can prove it wasn't rigged."

Why this is L6:

  • States the rule and its rationale in human terms
  • Connects fairness to bot resistance
  • Includes auditability

What L7 adds:

  • Publishes the policy and an audit summary per on-sale
  • Evaluates registration-based alternatives with data on fan satisfaction

Drill 5: Bots#

Prompt: "40% of seats in the last on-sale went to resellers. Fix it."

Staff Answer

"Layers, with friction pushed to where it hurts bots more than fans. At the edge: bot scoring on IP reputation, device/TLS fingerprints and behavior; low-trust sessions challenged or deprioritized. At the queue: login required, one token per account, token bound to device. At checkout — where identity is strongest — per-account, per-card, per-billing-address limits across the event. After the sale: cluster orders by shared payment instruments, addresses and devices; cancel and re-release. And reduce the payoff: identity-bound tickets delivered close to the event, with resale only through an official channel with price caps if the artist wants them. Measure: reseller share of purchases, estimated by listing matches on secondary markets."

Why this is L6:

  • Uses layered controls with measurable outcome
  • Places hard limits at checkout
  • Attacks the economics, not just the traffic

What L7 adds:

  • Frames resale as a business strategy decision (own the secondary market or restrict it)
  • Sets a bot-share target and reports it per on-sale to leadership

Drill 6: Sold Out#

Prompt: "How do the 1.9M people in the queue find out tickets are gone?"

Staff Answer

"The admission controller watches available and held. When available hits zero but held > 0, it switches the queue to 'LAST CHANCE': positions still advance at the rate seats are released, and users see 'all seats are currently in carts — some may return.' When both are zero, it publishes SOLD_OUT to the edge in the same config object as the watermark. Every status poll within 20s returns SOLD_OUT. Propagation < 60s. Late released holds (a failed payment) re-open LAST CHANCE for the head of the queue only, not a global flip-flop."

Why this is L6:

  • Distinguishes available-zero from truly sold out
  • Uses the same edge channel as admission
  • Avoids SOLD_OUT flapping

What L7 adds:

  • Pairs the sold-out state with a waitlist or next-show offer — turns the loss into a product funnel
  • Owns the comms plan for post-sale releases

Drill 7: Fungible Flash Sale#

Prompt: "Same system, but it's 20K identical sneakers and 500K buyers."

Staff Answer

"No seats, so a counter. A single Redis key handles ~100K decrements/s — likely enough behind a waiting room, but I'd shard into 20 sub-counters of 1K each to spread load, route each request to a random shard, and fall through to the next shard on zero. A successful decrement creates a reservation with TTL; order creation is asynchronous via a queue into the durable order DB. Reconciliation every minute: counter state vs reservations vs orders. Stranded units in shards are rebalanced near the end. Oversell is prevented because only a successful decrement creates a reservation. Honestly, with 500K buyers for 20K units, a raffle is the better product."

Why this is L6:

  • Adapts the model to fungible inventory
  • Addresses stranded stock and reconciliation
  • Questions the mechanism itself

What L7 adds:

  • Standardizes "raffle vs queue" as a product configuration across drops
  • Compares costs: waiting room vendor + infrastructure vs raffle operations

Drill 8: Build vs Buy#

Prompt: "Should we build our own waiting room?"

Staff Answer

"Edge waiting rooms are a commodity — Cloudflare and Queue-it sell them. For most companies I'd buy, because the hard parts (serving millions of pollers at the edge, token handling, bot integration) are exactly what CDNs are good at. I'd build the admission policy integration: the controller that computes the rate from our inventory and payment health, and the fairness configuration per event. For a ticketing company where on-sales are the core product, building on top of CDN primitives (edge compute + KV) can be justified to control fairness modes and data — but it's still built on the CDN, not on our origin."

Why this is L6:

  • Separates commodity from differentiating parts
  • Keeps admission decisions in-house

What L7 adds:

  • Evaluates vendor lock-in and multi-CDN failover for mega events
  • Frames it with the Build vs Buy framework and a 3-year cost comparison

Drill 9: Many On-Sales at Once#

Prompt: "Saturday at 10am, 40 events go on sale simultaneously."

Staff Answer

"Each event has its own waiting room and admission controller, but they share checkout and payment capacity. So there's a global budget — e.g., 1,000 payment TPS — allocated across events proportionally to expected demand and seats, with a minimum floor per event. The inventory store is partitioned by event, so events don't contend with each other. The risk is one mega event starving 39 small ones; the allocator prevents that. Scheduling is a product lever too — stagger on-sale times by 15 minutes when we control them."

Why this is L6:

  • Identifies shared downstream capacity as the coupling
  • Allocates with floors to protect small events

What L7 adds:

  • Runs an on-sale calendar with capacity reservations sold to promoters
  • Sets an org policy that no event above a demand threshold shares a slot without review

Drill 10: Degraded Mode#

Prompt: "Your inventory database primary fails 5 minutes into a mega on-sale."

Staff Answer

"Admission pauses immediately — the controller sees hold errors and stops advancing the watermark; the queue shows 'paused, your position is safe.' Existing holds keep their state in the replicated store; we extend all hold TTLs by the outage duration so nobody loses seats they were paying for. Failover to the replica (synchronous replication, so no lost holds) takes ~30–60s. Then admission resumes slowly — 25% of rate, ramping up — because a burst of queued checkouts would hit the new primary cold. Nothing fails open."

Why this is L6:

  • Pauses rather than errors; preserves user state
  • Extends holds for fairness
  • Ramps back carefully

What L7 adds:

  • Requires a failover game day before every mega on-sale
  • Pre-writes the public status message for "sale paused"

8. Deep Dive Scenarios#

Deep Dive 1: The Mega On-Sale That Broke#

Context: A global superstar's tour presale. 1.5M fans were invited by code; 10M+ sessions showed up. The waiting room held, but origin CPU spiked, holds timed out, and the presale was extended twice. The CEO wants a plan before the general sale next week.

Questions to Surface First:

  • What reached origin that shouldn't have — uninvited sessions, status polls, unprotected endpoints?
  • Were invitation codes validated at the edge or at origin?
  • Which dependency hit its ceiling first — inventory, payments, auth?
  • Is the general sale even viable given remaining inventory?

Typical L5 Approach: Scales up origin capacity 5× for the general sale, adds more CAPTCHA.

Staff Approach: Moves invitation-code validation to the edge (signed codes verifiable without origin); requires login before queue join; audits every storefront endpoint for admission-token enforcement; sets admission from the payment ceiling; runs a full synthetic on-sale at 2× the observed traffic before the general sale.

Principal Approach: Questions whether the general sale should happen in this form: with remaining inventory < 5% of demand, a queue will disappoint millions publicly. Proposes converting it to a registration + lottery, or cancelling it with clear comms. Brings promoter, legal and comms into the decision within 24 hours, with engineering providing the capacity and fairness options.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Pause admission; queue shows 'paused, position safe'
TriageBreak down origin traffic by path and token validity
Quick fixEdge validation of codes; enforce tokens on all store routes
GuardrailsAdmission capped by payment TPS; hold auto-extend while paused
Post-mortemInvited ≠ arriving; code checks at origin; no full-scale rehearsal

Metrics to Watch: origin_requests_without_admission_token, edge_code_validation_fail_rate, holds_created_per_min / admitted_per_min, payment_429_total.

Organizational Follow-up: Mega-event readiness review with go/no-go criteria signed by engineering, promoter and comms.

Ownership Question: "Who decides to cancel or change the general sale format?" Staff answer: The business owner with the promoter — engineering provides a written capacity and fairness assessment with a recommendation within 24 hours.

Key Takeaway: "Admission control only works if it's enforced before any origin work — and when demand is 50× supply, the right fix may be the mechanism, not the capacity."

What clears the Staff bar:

  • Finds what reached origin rather than adding capacity
  • Treats rehearsal as mandatory
  • Surfaces the product decision to its owners

Deep Dive 2: Silent Inventory Leak#

Context: After an on-sale reported "sold out," a reconciliation job finds 3,200 seats in HELD state with expired holds and no orders — never released, never sold. Fans are furious that seats appeared on resale sites at 5× while the official site said sold out.

Questions to Surface First:

  • How did seats stay HELD past expiry — was expiry cron-based?
  • Did the SOLD_OUT decision consider held seats?
  • Did any of those seats get sold through a side path?

Typical L5 Approach: Fixes the cron job and re-releases the seats.

Staff Approach: Makes expiry a read-time predicate so a stuck sweeper can't strand seats; changes the SOLD_OUT rule to available == 0 AND held == 0; adds a continuous invariant check (available + held + sold = capacity, with held_expired alerting); investigates whether the stranded seats correlate with resale listings (a possible insider or API abuse).

Principal Approach: Establishes inventory integrity as an auditable control — like financial reconciliation — with a report per on-sale; defines the re-release policy (announced time, queued re-release) so released seats go to fans, not refresh bots.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateFreeze those seats; investigate provenance
TriageWhy sweeper failed; whether any were sold elsewhere
Quick fixRead-time expiry; release via announced, queued re-release
GuardrailsInvariant reconciliation every minute; alert on expired-but-held > 0
Post-mortemSOLD_OUT declared with held seats outstanding

Metrics to Watch: inventory_invariant_drift, held_expired_not_released, soldout_declared_with_held_gt_0.

Organizational Follow-up: Per-on-sale integrity report reviewed by the business.

Ownership Question: "Who owns inventory integrity?" Staff answer: The inventory service team, with an invariant SLO and a reconciliation report — the same way payments owns ledger balance.

Key Takeaway: "Expiry must be a property of the data, not the success of a background job."

What clears the Staff bar:

  • Replaces a job-dependent invariant with a data-level one
  • Connects an engineering bug to a fairness and trust outcome

Deep Dive 3: Onboarding a Stadium Tour Promoter#

Context: A promoter wants to move a 50-date stadium tour (~3M tickets) to your platform, with presales for fan club, a card partner, and a general sale, all on the same inventory.

Questions to Surface First:

  • How is inventory split across presale tiers — fixed allocations or shared pool?
  • What fairness mode does the promoter want per tier?
  • Can our payment and edge capacity handle 50 on-sales in a week?

Typical L5 Approach: Adds a tier column and lets each presale run on the shared pool.

Staff Approach: Models tiers as allocations (e.g., 30% fan club, 20% card partner, remainder general) with unsold allocation rolling forward at defined times; each presale gets its own waiting room and admission; staggered on-sale times across dates; capacity plan for peak days.

Principal Approach: Treats this as a contract: capacity guarantees, fairness modes and bot-share targets written into the agreement; prices premium on-sale capacity; sets up a joint war room for the first three on-sales.

Staff Approach — Full Reasoning
DimensionStaff Answer
InventoryAllocation per tier; roll-forward at scheduled times
AdmissionPer-presale waiting room; shared payment budget allocation
FairnessConfigured per tier; published
CapacityStaggered on-sale times; synthetic load test per peak day

Metrics to Watch: allocation_sold_pct{tier}, rollforward_seats, admission_rate{event}.

Organizational Follow-up: Promoter-facing on-sale configuration tool with guardrails.

Ownership Question: "Who decides the tier allocations?" Staff answer: The promoter, within platform-enforced constraints (e.g., minimum general-sale share if we've committed to one publicly).

Key Takeaway: "Presale tiers are inventory partitions with their own fairness rules — model them explicitly."

What clears the Staff bar:

  • Models tiers as allocations, not labels
  • Plans capacity across the whole tour calendar

Deep Dive 4: Post-Mortem — Bots Won Anyway#

Context: Post-sale analysis shows ~35% of tickets were bought by reseller networks despite CAPTCHA, limits and a waiting room.

Questions to Surface First:

  • Which layer did they defeat — queue join, hold, checkout?
  • What identities did they use — aged accounts, real cards, residential proxies?
  • How much did they profit, and what would change the economics?

Typical L5 Approach: Adds a harder CAPTCHA and lowers purchase limits.

Staff Approach: Analyzes clusters: shared payment instruments, billing addresses, devices, and resale-listing matches. Tightens checkout-level linkage (per-card, per-address, per-device limits across the event), requires verified phone for high-demand events, and cancels clustered orders post-sale with re-release through the queue.

Principal Approach: Addresses the incentive: identity-bound, delayed-delivery tickets and an official resale channel with a price cap remove most of the arbitrage. That's a business model decision — potentially giving up secondary-market revenue — that the Principal frames with numbers for leadership.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateFreeze delivery of suspicious orders
TriageCluster analysis on payment, address, device
Quick fixCancel and re-release clustered orders via queue
GuardrailsCross-dimension limits at checkout; verified phone for top-tier events
Post-mortemControls focused at the door; weakest at checkout

Metrics to Watch: reseller_share_estimate, orders_per_payment_fingerprint, cancelled_bot_orders.

Organizational Follow-up: Quarterly bot-share report; red-team exercise before mega events.

Ownership Question: "Who decides whether to cancel suspected bot orders?" Staff answer: Trust & safety, with a documented evidence threshold and an appeal path — engineering provides the clustering, not the verdict.

Key Takeaway: "You can't out-CAPTCHA a profitable business. Change the economics."

What clears the Staff bar:

  • Moves defenses to where identity is strongest
  • Recognizes the economic root cause

Deep Dive 5: Going Global#

Context: The platform expands to Europe and Asia; a single tour will have on-sales across regions and time zones, and some presales are global.

Questions to Surface First:

  • Is inventory regional (per venue) or global? (Per venue — each show's inventory lives in one region.)
  • Where do global users queue?
  • Data residency and payment methods per region?

Typical L5 Approach: Deploys full stacks per region.

Staff Approach: Inventory for each event is homed in the venue's region (single writer, no cross-region consistency problem). Waiting rooms run at the global CDN edge; admission controllers are co-located with inventory. Payments route to regional providers. Global presales use per-event queues with randomized pre-open cohorts so time zone distance doesn't bias position.

Principal Approach: Defines regional on-sale calendars, local legal constraints (resale price caps vary by country), and cross-region capacity sharing for global mega-events; decides where the fairness policy differs by jurisdiction.

Staff Approach — Full Reasoning
DimensionStaff Answer
InventoryHomed per event in venue region
QueueGlobal CDN edge; watermark from the home region
FairnessRandomized pre-open cohort neutralizes RTT differences
PaymentsRegional providers; local methods

Metrics to Watch: admission_rate{event,region}, edge_status_latency{pop}, payment_success{provider}.

Organizational Follow-up: Regional legal review of resale and fairness rules.

Ownership Question: "Who owns a global presale's fairness policy?" Staff answer: The promoter with our policy team, constrained by the strictest applicable jurisdiction.

Key Takeaway: "Home inventory with the venue; queue at the edge. Global problems stay regional that way."

What clears the Staff bar:

  • Avoids cross-region consistency by homing inventory
  • Uses randomization to neutralize geography

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Lead with the demand ratio and design an edge waiting room with signed tokens and a watermark
  • Derive an admission rate from inventory, abandonment, checkout time and payment ceilings
  • Implement holds as conditional writes on section partitions with read-time expiry and a justified TTL
  • Explain server-assigned best-available as a contention fix
  • Define fairness modes (randomized pre-open, FIFO, registration + lottery) and who owns the choice
  • Layer bot defenses with limits at checkout and post-sale cancellation
  • Design the losing-user experience and SOLD_OUT semantics
  • Explain why the waiting room fails closed

The Bar for This Question#

Mid-level (L4): Seats table, booking API, some locking. Scales with replicas.

Senior (L5): Correct concurrency control (optimistic or pessimistic), holds with TTL, cache for the seat map, rate limiting and CAPTCHA. Correct for moderate demand; collapses under a real on-sale; fairness undefined.

Staff+ (L6): Admission control at the edge sized from downstream capacity; server-assigned conditional holds; short, lazily-expired TTLs; explicit fairness with an owner; layered bot defense; truthful sold-out; fail-closed waiting room. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "The Locking Strategy Doesn't Matter"#

EvidenceImplication
Behind a correct waiting room, inventory sees ~1–2K req/sAny sane concurrency control works
Without one, no concurrency control survives 900K attempts in 10sThe admission rate is the design

The Staff position: Pick conditional writes in one sentence; spend the interview on admission.

Why this matters in interviews: It signals you know where the difficulty actually lives.

10.2 "First Come, First Served Is the Least Fair Option"#

EvidenceImplication
Arrival order rewards latency, automation and session countBots dominate
It incentivizes hours of refreshingFans pay in time

The Staff position: Randomize the pre-open cohort; FIFO only after the open.

Why this matters in interviews: It shows you've thought about fairness as a mechanism, not a slogan.

10.3 "Selling Out in 90 Seconds Is a Failure"#

EvidenceImplication
Instant sell-outs look riggedTrust damage
They concentrate load into the worst possible windowMaximum contention and errors

The Staff position: Pace the sale to ~20–30 minutes via admission rate.

Why this matters in interviews: It reframes throughput as a product parameter.

10.4 "At 20× Demand, Stop Running a Queue"#

EvidenceImplication
A queue makes millions wait hours to learn they lostWorst possible losing experience
Registration + lottery tells losers days aheadBetter outcome for the same inventory

The Staff position: Above ~20× demand, recommend registration + lottery, with the queue serving only selected fans.

Why this matters in interviews: It shows willingness to change the product, not just the system.

10.5 "Bots Are a Business Model Problem"#

EvidenceImplication
Resale markups of several × face value fund sophisticated toolingTechnical defenses are an arms race
Identity-bound tickets and capped resale remove the marginEconomics beats CAPTCHAs

The Staff position: Build the technical layers, but push leadership on the incentive.

Why this matters in interviews: It connects engineering to business levers.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

The Staff engineer builds a system that sells 60K seats fairly and safely. The Principal notices that every on-sale is a public event with the company's reputation at stake, that fairness is a product and legal commitment rather than an algorithm, and that the platform's real product is on-sale readiness — capacity, rehearsal, policy and communication — sold to promoters and artists. The L7 questions: which mechanism do we offer for which demand ratio, who signs it, what do we promise publicly, and how do we make every on-sale boring.

The Org-Level Fault Line#

Engineering-owned fairness vs business-owned fairness. If engineering picks the queue rule, it will be optimized for system stability and defended poorly in public. If business picks it without engineering, it will promise things the system can't enforce. The L7 answer: engineering offers a menu of enforceable, auditable fairness modes with measured tradeoffs; the promoter and business choose per event; legal and comms approve the public description; every on-sale produces an audit summary.

Cost Model#

Assumptions: CDN waiting room vendor pricing or edge compute costs; origin at ~$0.10/host-hour; engineers at $250K fully loaded; load tests with millions of virtual users.

ScaleProfileInfra / monthHeadcountOn-call / operations
SmallRegional venues, < 5K-seat events, rare spikes~$5–15K; buy a waiting room3–5 engineersBusiness-hours; on-sale monitoring by the team
MediumArena tours, weekly high-demand on-sales~$50–150K incl. edge, origin headroom, load testing15–25 (inventory, checkout, edge, T&S)24/7 plus scheduled on-sale war rooms
LargeGlobal stadium tours, mega on-sales~$300K–1M+ incl. reserved capacity, multi-CDN, rehearsals60–120 across platform, payments, T&S, readinessDedicated on-sale operations team; exec-level go/no-go for mega events

The dominant cost is not steady-state infrastructure; it's peak readiness — reserved capacity, rehearsals and people in war rooms. Pricing premium on-sale capacity to promoters makes that visible.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversibility Cost
Publishing a fairness policyOne-way (public)Changing it later looks like moving the goalposts
Identity-bound, non-transferable ticketsOne-way-ishFans and resale partners adapt; reversing reopens arbitrage
Own secondary resale marketOne-way (business model)Contracts, revenue dependence
Waiting room vendorTwo-way with an abstractionWeeks if the admission API is ours
Seat-level vs zone-level inventoryTwo-way per eventConfiguration, if the model supports both
Hold TTLTwo-wayA config change — measure conversion vs velocity

The Standard I'd Write#

RFC: High-Demand On-Sale Standard (v1)

Scope: Any on-sale with forecast demand > 3× supply in the first 15 minutes.

MUST:

  • Use the edge waiting room; storefront routes MUST require a valid admission token.
  • Admission rate MUST be bounded by the lowest of inventory pacing, checkout capacity and contracted payment TPS.
  • The waiting room MUST fail closed.
  • Declare the fairness mode (randomized pre-open, FIFO, registration + lottery) with promoter sign-off and a published description.
  • Enforce per-account and per-payment-instrument limits at checkout.
  • Declare SOLD_OUT only when available = 0 and held = 0.

SHOULD:

  • Run a synthetic on-sale at 2× forecast within 7 days of any event forecast > 1M users.
  • Use registration + lottery above 20× demand.

Exceptions: Approved by the on-sale readiness lead and the business owner; recorded.

Success metrics: Zero double-sells; sold-out propagation < 60s; reseller share < 10%; zero on-sales extended due to platform failure.

What I'd Tell the VP#

On-sale day is when our reputation is decided, and today it's decided by how well our systems survive a crowd that's 30 to 50 times bigger than what we have to sell. I want to change that from a heroic engineering effort into a routine: every big on-sale gets a waiting room that lets people in only as fast as we can serve them, a fairness rule the promoter has approved and we publish, and a rehearsal beforehand. For the biggest events, I recommend we stop running open queues and move to registration and lottery, so fans learn in advance rather than after hours of waiting. The investment is mainly people and reserved capacity around peak days, and we should charge promoters for premium on-sale capacity.

Principal Interview Signals#

SignalWhat It Sounds Like
Fairness as a product decision"Engineering offers enforceable fairness modes; the promoter picks; legal approves the public wording."
Prices readiness"The cost is peak readiness — reserved capacity and rehearsals — so it should be a line item on the promoter contract."
Changes the mechanism"Above 20× demand, the queue is the wrong product."
Owns the public failure posture"The 'sale paused' message is written before the sale."
Attacks incentives"Identity-bound tickets and capped resale beat any CAPTCHA."

Staff answers that L7 interviewers find insufficient:

  • "Randomize pre-open arrivals" — right, but who approved it, and is it published?
  • "Admission from payment TPS" — right, but who negotiates the TPS and who pays for burst capacity?
  • "Cancel bot orders post-sale" — right, but what's the evidence threshold, appeal path and PR plan?

🧭 Principal Move: "The system can be made correct in a quarter. What makes on-sales boring is a readiness process — capacity, rehearsal, fairness sign-off, and a prewritten failure message — owned jointly with the business. That's what I'd build."


Appendices

Appendix A: Inventory Mechanics in Depth#

A.1 Seat State Machine#

Diagram: A.1 Seat State Machine

A.2 Best-Available Selection#

best_available(section, qty):
    blocks = cached_contiguous_blocks(section, qty)          # from seat-map cache, 1-2s stale
    tiers  = group_by_quality(blocks)                         # row distance, centrality
    for tier in tiers:
        for block in shuffle(tier)[:3]:                       # randomize within tier
            got = UPDATE ... WHERE seat_id IN block AND (AVAILABLE OR expired)
            if len(got) == qty: return got
            release(got)
    return SECTION_NEARLY_FULL(alternatives)

A.3 Why Read-Time Expiry#

A seat whose expires_at < now() is treated as AVAILABLE by every query and every conditional write. The sweeper that resets such rows is cosmetic (metrics, cleanliness). If the sweeper stops, correctness doesn't.

Appendix B: Data Model#

TableKeyNotes
seats(event_id, section, seat_id)status, hold_id, expires_at, version; partitioned by (event_id, section)
holdshold_iduser_id, seats[], expires_at, extended
ordersorder_id; unique (hold_id)idempotency_key, payment_status
account_limits(event_id, account_id)tickets held + bought; checked at hold and checkout
queue_audit(event_id, seq)randomization seed, cohort size, watermark history

Sharding: by event_id (each event's inventory in one database shard), sections as partitions or tables within. See Sharding and Data Modeling.

Appendix C: Contention Mechanisms — Quick Comparison#

MechanismThroughput per Hot KeyFailure ModeUse For
SELECT FOR UPDATEHundreds/s before lock queues explodeLatency collapseLow-demand events
Conditional UPDATE … WHERE status~1–2K/s per partitionFast failures to retryDefault for seats
Redis DECR / Lua~100K/s per keyDurability gap; reconcileFungible stock
Sharded countersN × single keyStranded stockVery hot fungible drops
Single-writer queue per sectionOne consumer's speedQueue lagWhen ordering per section matters

Appendix D: Client Contract#

  • Queue status polling every 15–30s with jitter; honor Retry-After.
  • Tokens: queue token (signed, bound to account + device); admission token (signed, 10-minute TTL).
  • Hold response includes expires_at; client shows a countdown; payment start requests the single extension.
  • Checkout carries Idempotency-Key derived from hold_id; client retries are safe.
  • States: WAITING, PAUSED, ADMITTED, LAST_CHANCE, SOLD_OUT — each with prewritten copy.

Appendix E: Observability#

E.1 Core Metrics#

queue_joined_total{event}             admitted_per_min{event}
admission_watermark{event}            holds_created_per_min / admitted_per_min
hold_success_rate{section}            hold_latency_p99
inventory_available / held / sold     inventory_invariant_drift
hold_expired_during_payment_total     payment_attempts_per_sec vs contracted_tps
bot_block_rate / challenge_rate       reseller_share_estimate (post-sale)
soldout_propagation_seconds

E.2 Critical Alerts#

AlertThresholdAction
Bypassholds/admitted > 1.2 for 1 minPage; enforce tokens
Invariant driftavailable + held + sold ≠ capacityPage inventory team
Payment ceilingattempts > 90% of contracted TPSSlow admission
Holds expiring mid-payment> 1% of paymentsAuto-extend; slow admission
Watermark stalledno advance for 2 min while inventory > 0Page platform

E.3 Debugging the Silent Failure#

Stranded inventory and unfair admission don't error. Reconcile the invariant every minute, and audit the queue: sample admitted users' positions vs the watermark history and verify randomization was applied to the pre-open cohort.

Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
Small venuesConditional holds, rate limitingFirst high-demand on-sale
Arena on-salesVendor waiting room + admission tied to inventoryBots, payment ceilings
Stadium toursAdmission from payments, fairness modes, checkout limitsMega-demand beyond queue viability
Global mega-eventsRegistration + lottery, identity-bound tickets, readiness processOrg and business coordination

What You Don't Build on Day One#

  • A custom edge waiting room (buy it)
  • Dynamic pricing
  • An official resale marketplace
  • Multi-region inventory (home each event in one region)

Appendix G: Multi-Tenancy and Cost#

  • Shared resources: payment TPS, edge capacity, checkout fleet — allocated per event with floors.
  • Promoter tiers: standard vs premium on-sale capacity (rehearsal, reserved capacity, war room).
  • Isolation: each event's inventory on its own shard partition; hot events can be moved to dedicated shards days ahead.
  • Tradeoff summary: admission control costs fan wait time and buys system survival and fairness; short holds cost slow payers and buy inventory velocity; checkout-level limits cost groups convenience and buy bot resistance; randomized pre-open cohorts cost early arrivers perceived advantage and buy a defensible, bot-resistant order.
  1. Loading the index…