Technologies referenced in this case study: PostgreSQL · Elasticsearch · Redis · Apache Kafka · DynamoDB
Related: Payment Processing · Flash Sales & Ticketing · Search Indexing · Dealing with Contention · Managing Long-Running Processes · Data Modeling · Sharding & Partitioning
How to Use This Case Study#
Organized for interview use first, reference second. The booking transaction is the part everyone prepares; the split between search and booking consistency is the part that decides the level.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines 1–2 → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Deep Dives 1–3 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal Lens) and Appendix B (inventory data models) |
What is a Reservation System? — Why interviewers pick this topic
A reservation system sells time-bounded, perishable inventory: a room on the night of June 14 can be sold exactly once (or, for a hotel room type, N times), and if it isn't sold by June 14 its value is zero forever. Guests search across millions of listings and date ranges, pick one, and book it; hosts or hotels manage calendars, prices and rules; payments are authorized at booking and captured or paid out later.
The interesting constraint is the asymmetry: search is read-heavy, approximate and must be fast (hundreds of thousands of availability checks per second); booking is write-light, exact and must never double-sell. The same "is it available?" question has two different correctness bars depending on where it's asked.
Before vs After — the "two guests, one cabin" scenario:
Without a correct booking path:
t=0: Guest A and Guest B both see "Lakeside Cabin — available Jun 14–16" in search.
t=+40s: A clicks Book. Service reads calendar: free. Starts payment auth (1.2s).
t=+40.5s: B clicks Book. Service reads calendar: still free. Starts payment auth.
t=+41.2s: A's auth succeeds → INSERT booking. B's auth succeeds → INSERT booking.
t=+2 weeks: Both arrive. One is turned away at 10pm with a toddler.
t=+2 weeks: Host gets penalized, guest gets a refund + rebooking credit, review is 1 star.
With a constraint-enforced booking path:
t=+40s: A: BEGIN; claim nights Jun 14, 15 for listing 88 (unique per night) → OK. Hold 10 min.
t=+40.5s: B: claim same nights → unique violation → "Just booked — here are 6 similar cabins".
t=+41.2s: A's payment auth succeeds → hold converts to CONFIRMED.
t=+41.3s: Availability event → search index updated in ~2s.
Why interviewers reach for this question: It has a clean correctness invariant (no double booking), a massive read path that cannot afford that invariant, a long-running workflow (hold → pay → host accept → stay → payout) and a policy question (overbooking) that engineering alone can't answer. It separates candidates who put a lock on everything from candidates who put the invariant exactly where it belongs.
Mechanics Refresher: Preventing Double Booking
| Mechanism | How It Works | Pros | Cons |
|---|---|---|---|
| Per-night inventory rows + unique constraint | One row per (listing, night); booking inserts/claims each night; DB rejects duplicates | Simple, exact, index-friendly; partial overlaps caught automatically | Rows × nights: 365 rows/listing/year |
| Range exclusion constraint | Postgres EXCLUDE USING gist (listing_id WITH =, stay WITH &&) on a daterange | One row per booking; overlap semantics native | Postgres-specific; GiST index cost; harder to shard across stores |
| Counter per (room type, night) | available = allotment − sold; decrement with a conditional update | Natural for fungible hotel rooms; supports overbooking caps | Hot rows on popular nights; a range needs N conditional updates atomically |
| Pessimistic hold with TTL | Claim inventory at "start checkout", release on expiry | Guest isn't surprised at payment step | Abandoned holds block inventory; bots can squat |
| Optimistic check at commit | Check-and-claim only at the final Book click | No squatting | Guest can lose the room after entering payment details |
| Distributed lock (Redis) around booking | Lock listing, read calendar, write booking, unlock | Familiar | The DB can already do this atomically; the lock adds a failure mode, not safety |
For most production systems: per-night rows (or a range exclusion constraint) with a database-enforced uniqueness guarantee, a short hold (~10 minutes) created when the guest starts checkout, and a search index that is explicitly allowed to be seconds stale. The mechanism is not the interview — where the invariant lives and how stale everything else may be is.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What This Interview Actually Tests#
A reservation system is not a locking question. Everyone knows how to prevent two inserts.
It is a consistency-placement question that tests:
- Whether you enforce the no-double-booking invariant in exactly one place — the inventory store — and let everything else be stale
- Whether you size the read path (search) separately from the write path (booking), with numbers
- Whether you model the booking as a long-running workflow with holds, payment and host acceptance — and know what happens when each step times out
- Whether you recognize that overbooking and hold duration are business policies with owners, not engineering defaults
The key insight: "Available" means two different things. In search, it means probably bookable — seconds stale is fine. At booking, it means claimed, atomically, by exactly one reservation. Staff candidates design two systems with two correctness bars and one reconciling event stream between them.
The L5 vs L6 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Draws search → listing → booking → payment | Asks "Unique listings or pooled room types? Instant book or host approval? Is overbooking allowed?" | Asks "Who owns inventory truth — us, hotels' property systems, or channel managers? That decides whether we can promise anything" |
| Double booking | Distributed lock on the listing during booking | Unique constraint per (listing, night) in the booking DB; no external lock needed | Makes "inventory claims only through the inventory service" a platform invariant across web, mobile, partner APIs and iCal imports |
| Search consistency | Search reads the booking DB | Search index is eventually consistent (~1–5s via CDC); booking re-validates; stale-hit rate is a tracked metric | Sets a product SLO on "booked-but-shown-available" rate and negotiates it with search ranking and supply teams |
| Holds | Lock until payment completes | Hold rows with a 10-minute TTL; sweeper releases; per-user hold caps against squatting | Treats hold duration as a revenue lever with experiment data; owned by product, bounded by engineering |
| Overbooking | "Never overbook" | "Never for unique listings; for pooled room types, a capped overbooking percentage set by revenue management, with a walk policy" | Prices it: expected no-show revenue vs walk cost vs brand damage; delegates the knob to revenue management with guardrails |
| Ownership | Booking team owns it all | Inventory service owns the invariant; search owns freshness; payments owns money; support owns walks and relocations | Draws the supply/demand org boundary: supply (calendars, pricing, external sync) vs demand (search, checkout) with a contract in between |
Why "double booking" separates levels
L5: "Acquire a Redis lock on listing:88, read the calendar, if free write the booking, release the lock." It works in the demo. In production the lock has a TTL; a slow payment call or GC pause outlives it; a second request acquires the lock and both write. Or Redis fails over and the lock vanishes. The lock was never the thing enforcing the invariant — it was a hope.
L6: "The database enforces it. Nights are rows with a unique key on (listing_id, night). A booking claims its nights in one transaction; a conflicting claim fails with a constraint violation, regardless of timing, pauses or lock services. I don't need a distributed lock because the store that holds the truth can do the compare-and-set itself."
L7: "The invariant is only as good as the number of write paths that bypass it. Partner APIs, host-side blocking, iCal imports from other platforms, support-tool overrides — each is a way to put two reservations on one night. I'd make the inventory service the only writer and put the others behind it."
Why "search consistency" separates levels
L5: Either queries the booking database for availability on every search (collapses under a 1000:1 look-to-book ratio) or caches aggressively without saying how stale it is.
L6: "Search availability is a projection. Booking commits emit events via an outbox; a consumer updates the search index within ~1–5 seconds. Search can show a listing that was just booked — the booking path re-validates and offers alternatives. I'd track search.stale_click_rate: the share of listing-page views where the dates shown available are actually taken. Target under 0.5%."
L7: Recognizes that staleness has a revenue cost and a trust cost, and makes it a jointly-owned product metric. "If stale clicks hit 3% during peak, conversion drops and guests stop trusting the calendar. That's a number the search and supply directors share."
Why "overbooking" separates levels
L5: "We never allow overbooking." Correct for Airbnb-style unique listings. Incomplete for hotels, where no-show and late-cancellation rates make a strict zero-overbooking policy leave revenue on the table every night.
L6: Distinguishes: "Unique inventory — zero overbooking, enforced by constraint. Pooled room types — the allotment can exceed physical rooms by a configurable percentage per night, set by revenue management. Engineering enforces the cap exactly; revenue management owns the number; operations owns the walk procedure."
L7: "Overbooking is an options trade. I'd want the expected-value model — no-show rate distribution per property and season, walk cost, lifetime-value impact on a walked guest — reviewed quarterly, and a hard ceiling engineering will enforce regardless of what the model says."
The Staff Positions#
| Position | Rationale |
|---|---|
| The inventory database enforces the invariant; nothing else does | Unique/exclusion constraints are immune to pauses, lock TTLs and failovers |
| Search is a projection, seconds stale by design | Look-to-book is 100:1 to 1000:1+; exact availability on every search is unaffordable |
| Short holds with TTL at checkout start, capped per user | Guests shouldn't lose a room while typing a card; bots shouldn't squat inventory |
| Booking is a saga, not a transaction | Hold → payment auth → (host accept) → confirm spans seconds to 24 hours; each step needs a timeout and a compensation |
| Shard by listing, not by date or guest | All nights of one listing in one shard makes the multi-night claim a single-shard transaction |
| Overbooking is a policy with an owner, not a bug or a feature | Zero for unique listings; capped and owned by revenue management for pooled inventory |
| External calendars are untrusted, lagging inputs | iCal and channel-manager syncs are the #1 source of real double bookings; model them as blocks with provenance |
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Unique-inventory marketplace (Airbnb, vacation rentals) | Each listing is one unit; host controls calendar; instant book or request-to-book | Per-night claims with unique constraint; holds; host-acceptance workflow with 24h expiry | Double booking via external calendars; host cancellations | Zero double bookings from our own paths; external-sync conflicts detected within minutes |
| Pooled / fungible inventory (hotel chains, room types) | N identical rooms per type per night; revenue management; overbooking | Counters per (room type, night) with conditional decrement; allotment including overbooking cap | Hot counters on peak nights; walks when no-shows under-run forecast | Sold ≤ allotment exactly; walks within budget |
| Aggregator / channel distribution (OTAs selling others' inventory) | Truth lives in the supplier's property management system | Cached availability + confirm-with-supplier at booking; allotments pushed by channel managers | Supplier rejects a "confirmed" booking; rate/availability mismatch | Booking confirmed only on supplier ack; mismatch rate tracked per supplier |
🎯 Staff Move: "I'll design the unique-inventory marketplace — Airbnb-shaped — because it has the strictest invariant and the most interesting workflow: holds, instant book vs host approval, external calendar sync. I'll point out where pooled hotel inventory changes the model to counters and allows overbooking, and where an aggregator can't promise anything until the supplier confirms."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Search Freshness vs Booking Correctness | Exact availability in search (expensive, slow) vs stale projection (fast, occasional disappointment)? |
| 2 | Pessimistic Holds vs Optimistic Commit | Reserve at checkout start (guest certainty, inventory squatting) or claim at the final click (no squatting, late disappointment)? |
| 3 | Inventory Model: Nights vs Ranges vs Counters | One row per night, one row per booking with range exclusion, or counters per type — which makes the invariant cheap? |
| 4 | Zero Overbooking vs Managed Overbooking | Leave no-show revenue on the table, or occasionally walk a guest? Who decides, who pays? |
| 5 | Book-Then-Pay vs Pay-Then-Book (Workflow Coupling) | Claim inventory before payment (risk holding unpaid inventory) or authorize payment first (risk charging for a room you can't give)? |
In the Wild: Real Production Systems#
Why this section belongs here: Each of these illustrates a different inventory-truth model. Cite the model, not just the brand.
Airbnb — Instant Book, Request to Book, and Calendar Sync#
Airbnb publicly offers two booking modes: Instant Book, where a guest's reservation is confirmed without host review, and Request to Book, where the host has 24 hours to accept or decline before the request expires. Hosts can also import and export calendars via iCal to synchronize availability with other platforms — a well-known source of double bookings, because imported calendars refresh periodically rather than instantly.
Staff insight: Two booking modes are two state machines sharing one inventory. Request-to-Book must decide whether a pending request holds the nights (blocking other guests for up to 24h) or not (risking the host accepting two requests). And iCal sync proves the invariant is only as strong as the slowest writer that can mark a night unavailable.
Hotel Revenue Management — Overbooking and "Walking" Guests#
Hotels routinely sell more rooms than they physically have for nights with predictable no-shows and late cancellations, a practice that is standard in the lodging and airline industries. When more guests arrive than rooms exist, the hotel "walks" a guest: arranges and pays for a comparable room nearby, transportation, and often a goodwill credit. Airlines do the equivalent with denied boarding, which in the US is regulated with required compensation.
Staff insight: Overbooking is a deliberate, priced risk — not a consistency bug. In an interview, the Staff move is to ask whether the inventory is fungible and whether the business has a walk policy before declaring "never overbook".
Online Travel Agencies and Channel Managers — Truth Lives Elsewhere#
Large OTAs distribute hotel inventory they don't own. A hotel's property management system (PMS) is the source of truth; channel managers push availability and rates to each OTA, often as allotments per room type per night. The OTA's "available" is a cached copy, and a booking is only final when the supplier side accepts it.
Staff insight: When you don't own inventory, you can't enforce the invariant — you can only detect violations and compensate. The design shifts from constraints to reconciliation and supplier-level SLAs.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "We lock the listing during booking" | "The lock expires during a slow payment call. What happens?" | Invariant placement: lock vs constraint |
| "Search shows real-time availability" | "At 50K searches/sec over 7M listings × 30 date combos, how?" | Read/write asymmetry, projections |
| "We hold the room during checkout" | "A bot starts 10,000 checkouts on New Year's Eve. Now what?" | Hold abuse, caps, TTL |
| "Payment and booking in one transaction" | "The payment call takes 3s and times out. Is the room booked?" | Saga, unknown outcomes |
| "We never overbook" | "It's a 400-room hotel with 8% no-shows. Is that the right business call?" | Policy vs engineering; who decides |
| "Hosts update their calendar" | "The host also lists on another site. A booking there takes 3 hours to sync." | External writers, detection, compensation |
System Architecture Overview#
Reading the diagram: Two paths with two correctness bars. The read path serves tens of thousands of searches per second from an index that may be 1–5 seconds stale. The write path runs hundreds of bookings per second through one gate — the inventory service — whose database enforces uniqueness per (listing, night). Every writer of availability (guests, hosts, external calendar sync, hold sweeper) goes through that gate. Changes leave via an outbox to Kafka and update the index.
search.stale_click_ratemeasures the cost of the staleness we chose.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Double booking | "Distributed lock" | "Unique constraint on (listing_id, night). The DB is the lock." |
| Search | "Query availability live" | "Projection via outbox + Kafka, 1–5s stale. Booking re-validates." |
| Holds | "Lock during payment" | "Hold rows, 10-min TTL, sweeper, max 2 active holds per user." |
| Payment | "One transaction with booking" | "Saga: hold → auth → confirm. Auth UNKNOWN keeps the hold alive until resolved." |
| Sharding | "Shard by date" | "Shard by listing_id — a stay is a single-shard transaction." |
| Overbooking | "Never" | "Never for unique listings; capped per night for pooled rooms; revenue management owns the cap." |
| External calendars | "Sync them" | "Untrusted, lagging writers. Import as blocks with provenance; detect conflicts; relocation playbook." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Active listings (Airbnb-scale) | ~7–8 million | 7.5M × 365 nights ≈ 2.7B night-rows/year — fine when sharded by listing |
| Look-to-book ratio (travel) | ~100:1 to 1,000:1+ | Search must be served from a projection, not the booking DB |
| Search QPS (large marketplace, peak) | ~20K–100K | Index sized for this; booking DB never sees it |
| Bookings/sec (large marketplace) | ~50 average, ~500 peak | The write path is small; correctness, not scale, is the challenge |
| Search p99 target | < 400–500ms | Includes geo, filters, availability and pricing |
| Search index staleness target | 1–5s p99 | Stale-click rate stays < 0.5% |
| Hold TTL | ~10 minutes (5–15 typical) | Long enough to enter payment, short enough to limit squatting |
| Request-to-book expiry (Airbnb) | 24 hours | Pending requests either block or don't — a policy decision |
| Payment authorization latency | p50 ~0.5–1s, p99 2–5s | Why holds must outlive the payment call with margin |
| Hotel no-show + late-cancel rates | often ~5–10% | Why pooled inventory overbooks a few percent |
| Walk cost (hotel) | ~1–2 nights' rate + transport + goodwill | The downside side of the overbooking equation |
| External iCal refresh | minutes to hours (platform-dependent) | The window in which a cross-platform double booking can happen |
Interview Walkthrough
The most common mistake: Candidates spend 15 minutes on search ranking and geo-indexing, then draw a booking flow with a distributed lock and run out of time before the interviewer asks "what happens when two people book the last night at the same moment — and what does search show them?" Compress search to a projection in 3 minutes; spend the time on the invariant, holds, the saga and external writers.
Phase 1: Requirements & Framing (2–3 minutes)#
Functional scope in one breath:
"Guests search listings by location, dates and guest count; view a listing's calendar and price; book with payment. Hosts manage calendars, prices and booking rules, and accept or decline requests. We pay hosts after check-in."
Then commit on the questions that change the design:
"Is inventory unique — one listing, one unit — or pooled room types like a hotel? I'll assume unique, Airbnb-style. Both instant book and request-to-book exist. No overbooking for unique listings. Hosts can sync calendars with other platforms."
Then the non-functional split — this is the Staff move:
"Two correctness bars. Booking: zero double bookings from our own write paths, strongly consistent. Search: seconds-stale is fine, as long as booking re-validates. Scale: roughly 7 million listings, tens of thousands of searches per second, a few hundred bookings per second at peak. The read path is 100–1,000× the write path."
🎯 Staff Move: Stating two correctness bars in the first three minutes tells the interviewer you won't try to make search strongly consistent or make booking eventually consistent — the two most common ways this design goes wrong.
Phase 2: Core Entities & API (1–2 minutes)#
- Listing:
listing_id, host, location, capacity, rules (min nights, instant book),timezone - NightInventory:
(listing_id, night)→status(OPEN / HELD / BOOKED / BLOCKED),reservation_id,hold_expires_at,source - Reservation:
reservation_id,listing_id,check_in,check_out(half-open),guest_id,state,price_snapshot,payment_intent_id,version - Hold: part of the reservation in state HELD with
expires_at - CalendarBlock: host or external block with
source(host, ical:vrbo, pms:xyz) andexternal_uid
Nights are local dates in the listing's timezone, and ranges are half-open [check_in, check_out) — a stay Jun 14–16 claims nights 14 and 15, so a checkout on the 16th and a check-in on the 16th don't conflict.
GET /v1/search?bbox=..&check_in=2026-06-14&check_out=2026-06-16&guests=4
→ [{ listing_id, nightly_price, total_estimate, maybe_available: true }]
POST /v1/reservations Idempotency-Key: <uuid>
{ listing_id, check_in, check_out, guests } → 201 { reservation_id, state: HELD, expires_at }
POST /v1/reservations/{id}/confirm Idempotency-Key: <uuid>
{ payment_method_token } → 200 { state: CONFIRMED | PENDING_HOST | PAYMENT_PROCESSING }
POST /v1/reservations/{id}/cancel
PUT /v1/listings/{id}/calendar { blocks: [...], prices: [...] } (host)
🎯 Staff Move: "Search returns
maybe_available, notavailable. That's an honest API — the listing page and the booking call are where availability becomes a promise."
Phase 3: High-Level Architecture (≤5 minutes)#
Walk the booking flow in 90 seconds:
- Guest clicks Reserve. Booking service calls inventory: claim nights 14, 15 for listing 88 as HELD, expires in 10 minutes.
- Inventory runs one transaction on the listing's shard: conditional update of both night rows from OPEN to HELD. Both succeed → hold created. Either fails → rollback, 409 with alternatives.
- Outbox event
inventory.changed→ search index marks nights unavailable in ~1–5s. - Guest enters payment; booking service authorizes (~1s). Success → inventory flips nights HELD → BOOKED for this reservation; reservation CONFIRMED (or PENDING_HOST for request-to-book).
- Hold expiry without confirmation → sweeper flips HELD → OPEN; event back to search.
Key points to hit:
- Invariant in the DB — conditional updates/unique key, not a lock service
- Shard by listing — a stay is a single-shard transaction
- Search is a projection, refreshed by events
- Booking is a saga with timeouts per step
- Every writer goes through inventory — guests, hosts, sync, sweeper
🎯 Staff Move: After drawing: "This handles the core race. The design questions that matter are: how stale search may be and what that costs, how holds behave under abuse and timeouts, and how we deal with writers we don't control — external calendars. Which do you want first?"
Phase 4: Transition to Depth (1 minute)#
"Three areas decide whether this works at Airbnb scale: the search-to-booking consistency gap and how we measure it, the booking saga — holds, payment unknowns, host acceptance — and external calendar sync, which is where real double bookings come from. Where would you like to go?"
Default: the booking saga — it combines the invariant, holds and payments in one story.
Phase 5: Deep Dives (25–30 minutes)#
Deep dive 1: The claim transaction (5–6 min)
"Nights are pre-materialized rows, 365–540 days ahead per listing. The claim:"
BEGIN;
UPDATE night_inventory
SET status = 'HELD', reservation_id = $res, hold_expires_at = now() + interval '10 min'
WHERE listing_id = $l AND night >= $in AND night < $out AND status = 'OPEN';
-- rows_updated must equal (check_out - check_in); otherwise ROLLBACK → 409
INSERT INTO reservations (...) VALUES (..., 'HELD', ...);
INSERT INTO outbox (...) VALUES ('inventory.changed', $l, $in, $out);
COMMIT;
"Two concurrent claims on overlapping nights: the first to update a row takes its row lock; the second blocks, then re-evaluates status = 'OPEN' after the first commits, fails the predicate, updates fewer rows than needed, and rolls back. No distributed lock, no lost update. Contention is per listing — the hottest listing gets maybe a few conflicting attempts per minute on a holiday weekend."
Deep dive 2: Holds and abuse (5–6 min)
"Holds protect the guest who's typing a card number; they also let anyone block inventory for 10 minutes. Controls: max 2 active holds per user and per device fingerprint; holds require an authenticated account; per-listing hold churn alert — a listing held and released 20 times in an hour is being squatted. On peak nights for scarce inventory, shorten TTL to 5 minutes. TTL expiry is enforced two ways: the sweeper releases expired holds every 30 seconds, and the claim predicate treats HELD AND hold_expires_at < now() as OPEN — so a dead sweeper can't strand inventory."
Deep dive 3: Booking saga and payment (6–7 min)
"Steps: hold → payment authorize → confirm (instant) or pending host (request). Compensations: auth declined → release hold; confirm fails after auth → void auth. The dangerous case is auth UNKNOWN — the PSP timed out. I don't release the hold; I extend it until the payment resolves, because releasing and letting someone else book, then discovering the auth succeeded, means charging a guest for nothing. For request-to-book: the host has 24 hours. I hold the nights during that window for instant-bookable conflicts — so the host can't be double-requested into accepting two — and I authorize payment up front so acceptance is instant."
Deep dive 4: Search freshness (4–5 min)
"Index documents carry an availability bitmap per listing — 365–540 bits for the next 12–18 months, a few dozen bytes. Date-range search is a bitwise check. Updates come from the outbox via Kafka; p99 lag target 5s. The residual gap is measured: search.stale_click_rate — listing pages where the searched dates are no longer available. When it's above 0.5%, either lag is up or a hot market is selling out; both are actionable."
Deep dive 5: External calendars and ownership (3–4 min)
"iCal imports are periodic, often hours apart. They're written as BLOCKED nights with source and UID, through the same inventory service. If an import brings a block that overlaps an existing booking, that's a detected double booking — we can't prevent it, because the other platform already sold the night. The response is a playbook: alert host, give them a window to resolve, then relocation support for the guest. Ownership: supply team owns sync; support owns relocations; the host bears the penalty per policy."
Phase 6: Wrap-Up (2–3 minutes)#
"The core idea is putting consistency exactly where it's needed. Booking is strongly consistent because the database enforces one claim per night. Search is deliberately stale and we measure the cost. Booking is a saga with holds that cover payment uncertainty. And the real double bookings come from writers we don't control, so we detect and compensate rather than pretend to prevent."
The evolution closer:
"Later: pooled inventory for hotel partners with counters and capped overbooking, a partner API for channel managers with push-based availability instead of iCal polling, and dynamic hold TTLs by demand. I'd resist building our own global transaction layer — shard by listing keeps every booking single-shard forever."
🎯 Staff Move: End on the external writers and who owns relocations. That's the operational truth of reservation systems, and almost nobody says it.
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| Search ranking rabbit hole | 12 min on ML ranking and geohashes | "Geo + bitmap filter in the index; ranking is a separate team" |
| Lock-centric design | Redis lock + TTL around booking | Unique/conditional update in the inventory DB |
| One consistency level | Everything strong or everything eventual | Two bars, one event stream, measured gap |
| Ignoring payment timing | Books after payment with no hold | Hold → auth → confirm, with UNKNOWN handling |
| Ignoring time zones | Stores UTC timestamps for nights | Local dates per listing, half-open ranges |
| No external writers | Assumes we're the only calendar | iCal/channel sync as untrusted writers with conflict detection |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Reservation systems are a compact test of consistency placement. The invariant is simple and absolute — one guest per night per listing — but the system around it is enormous, read-dominated and full of writers with different trust levels. Senior candidates tend to apply one consistency level everywhere: they either lock too much (and can't serve search) or too little (and double-book). Staff candidates put the invariant in one place, make everything else a projection, and measure the gap.
It also tests workflow thinking. A booking isn't a request; it's a multi-step process that can last from 2 seconds (instant book) to 24 hours (host approval) and crosses payments, notifications and host behavior. Each step can time out. Knowing which compensations are safe — and that "payment unknown" means keep the hold — is Staff-level operational judgment.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"A guest sees the cabin available in search, spends four minutes on the listing page, clicks Reserve, enters a card, and the payment provider times out. Meanwhile another guest clicks Reserve on the same dates. Walk me through what each guest sees, and what state the nights are in at every step."
The answer reveals everything: whether search is understood as stale, whether a hold exists and when it was created, whether the payment timeout is treated as unknown, and whether the second guest is rejected by the database or by luck.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Unique-inventory marketplace → the invariant is per night, per listing
- Constraint: one reservation per night; host rules (min nights, advance notice, instant book)
- Strategy: per-night rows or range exclusion; holds; saga with host acceptance; external calendar sync
- Failure mode: cross-platform double booking; host cancellation
- Who pays for imperfection: the guest (relocation), the host (penalties), the platform (credits, trust)
Pooled / fungible inventory → the invariant is a count
- Constraint:
sold(room_type, night) ≤ allotment(room_type, night)where allotment may include overbooking - Strategy: counter rows per (property, room type, night) with conditional decrement; room assignment deferred to check-in
- Failure mode: hot counters on peak nights; walks when overbooking outruns no-shows
- Who pays: front desk and revenue management (walk costs), the walked guest
Aggregator / channel distribution → the invariant is someone else's
- Constraint: supplier PMS is truth; we hold a cached allotment
- Strategy: sell from cache, confirm with supplier synchronously or within minutes; reconcile rates and availability
- Failure mode: supplier rejects after we've told the guest "confirmed"
- Who pays: the platform's support (rebooking), the supplier (SLA penalties if contracted)
2.2 When NOT to Build a Reservation System#
- A single property or a handful of units. Use a property management system or a booking plugin; your "inventory service" is a calendar in a SaaS tool.
- Truly fungible, non-time-bound inventory (e-commerce stock). That's an inventory counter with backorders, not a reservation system — no nights, no holds of this shape.
- Flash-sale ticketing with seat maps and 100K concurrent buyers for one event. The hard problem is admission control and fairness under a spike on a single hot resource; see Flash Sales & Ticketing. The per-listing sharding that makes Airbnb easy doesn't help when everyone wants the same inventory.
- Resources booked in minutes with complex constraints (clinic appointments with provider/room/equipment combos). The invariant becomes multi-resource; you're closer to a scheduling/constraint-solving system.
- When you don't own the inventory and can't get push updates. Building a strongly consistent booking path on top of a supplier you poll hourly buys nothing; design for confirm-with-supplier and compensation.
🎯 Staff Insight: "The per-listing shard is the gift of this problem — 7 million independent small invariants instead of one big one. If the interviewer turns it into one hot resource with 100,000 buyers, I switch playbooks to a queue-based admission design."
2.3 What the Interviewer Leaves Underspecified#
- Unique vs pooled inventory — changes the data model and the overbooking answer
- Instant book vs host approval — changes the saga and whether pending requests hold nights
- Time zones and night semantics — local dates, half-open ranges, same-day turnover
- External calendars — whether we're the only writer
- Cancellation policy — flexible vs strict changes refund flows and when nights reopen
- Price consistency — whether the price shown in search must be honored at booking
- How far ahead — 12 vs 24 months determines night-row count and index bitmap size
2.4 Precise Terminology#
| Term | What It Means | Why It Matters |
|---|---|---|
| Night | The unit of inventory: a local calendar date for a listing | Store as DATE in listing timezone, not a timestamp |
| Stay range | Half-open [check_in, check_out) | Back-to-back stays don't conflict |
| Hold | Temporary claim with expiry during checkout | Protects the guest; must expire safely |
| Allotment | Units available to sell for a pooled type on a night | May exceed physical units by an overbooking cap |
| Block | Night made unavailable by host or external source | Needs provenance to reconcile |
| Look-to-book | Searches per completed booking | Sizes read vs write paths |
| Stale click | Listing view where searched dates are already gone | The measurable cost of search staleness |
| Walk / relocation | Moving a guest to alternative lodging when inventory fails | The compensation path for any double booking |
| Instant Book / Request to Book | Auto-confirm vs host approval (24h window) | Two state machines on one inventory |
| Rate parity / price snapshot | Price captured at hold time | Guest pays what they saw, even if host changes price |
🎯 Staff Insight: If the interviewer says "check availability", ask "for search ranking or for booking? They have different consistency requirements, and I'd serve them from different stores."
3. The Five Fault Lines#
Each fault line has an engineering choice and a business owner. The Staff answer names both.
3.1 Fault Line 1: Search Freshness vs Booking Correctness#
The tension: Exact availability in search means consulting the source of truth for every candidate listing on every search — at a 1000:1 look-to-book ratio, the inventory DB would serve 1,000× its write load in reads. A projection is cheap and fast but can show booked listings as available.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Search queries inventory DB live | Always correct | 50K QPS × 100s of candidates per search; DB melts; p99 in seconds | Everyone (latency, outages) |
| Projection via CDC/outbox (1–5s) | Fast; scales independently | Stale clicks during sell-outs | Guests (occasional disappointment), search team (lag SLO) |
| Projection + live re-check of top-N results | Top results accurate | Extra reads: ~20 per search to inventory | Inventory team (read replicas) |
| Hourly batch index | Cheapest | Stale clicks 5–20% in hot markets | Conversion (revenue) |
Staff default: "Projection via outbox and Kafka with a 5-second p99 lag SLO. The listing page reads a replica with < 1s lag, so the calendar the guest studies is fresher than search. The booking call is the only place availability is a promise."
When to deviate:
- Sell-out events in a market (festival weekend, eclipse): add a live re-check of the first page of results against a read replica — ~20 point lookups per search, affordable for a few markets.
- Aggregator inventory: you can't get CDC from a supplier; availability in search is a cached allotment with a TTL and a supplier-confirm step.
🧭 Principal Move: "Stale-click rate is a shared KPI between search and supply. When it rises, search wants lower lag and supply wants hosts to keep calendars accurate. I'd put both teams on the same dashboard with a quarterly target — the number, not the architecture, is what aligns them."
❌ Common L5 Trap: "We'll cache availability in Redis with a 60-second TTL." Why 60? What's the stale-click rate at 60s in a market selling out? A TTL without a measured cost is a guess; an event-driven projection with a lag SLO is a design.
3.2 Fault Line 2: Pessimistic Holds vs Optimistic Commit#
The tension: Claiming inventory when the guest starts checkout guarantees they won't lose the room mid-payment, but lets abandoned carts and bots block inventory. Claiming only at the final click prevents squatting but can tell a guest "sorry, just booked" after they've typed their card.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Pessimistic hold at checkout start (10 min) | Guest certainty; clean payment flow | Abandoned holds block ~5–15% of peak inventory-minutes; bot squatting | Hosts (lost sales), other guests |
| Optimistic claim at final click | No squatting | Late rejections after payment entry; conversion hit | Guests (frustration), conversion |
| Hold at final click, before payment auth | Short holds (~seconds–minutes); no squatting | Guest can still lose it between page view and click | Guests (rarely) |
| Adaptive: short holds on scarce inventory, longer elsewhere | Balances both | Complexity; needs demand signal | Product (tuning) |
Staff default: "Hold at the Reserve click, 10-minute TTL, max 2 active holds per user. A payment UNKNOWN extends the hold rather than releasing it. Expiry is enforced by predicate as well as by the sweeper, so the hold can't outlive its TTL even if the sweeper is down."
When to deviate:
- Scarce, high-demand inventory (New Year's Eve, unique properties): 5-minute holds and CAPTCHA/device checks before a hold.
- Request-to-book: the "hold" is the pending request for up to 24h — whether it blocks other guests is a product decision (see 3.5).
🎯 Staff Insight: "The hold TTL is a product knob with an engineering floor. Engineering says 'not shorter than p99 payment latency plus form-fill time'; product decides how much squatting to tolerate above that."
3.3 Fault Line 3: Inventory Model — Nights vs Ranges vs Counters#
The tension: The data model decides how cheap the invariant is. Per-night rows make overlaps trivial but multiply rows; range rows are compact but need overlap constraints; counters suit fungible inventory but not unique listings.
| Model | What Works | What Breaks | Who Pays |
|---|---|---|---|
Per-night rows (listing_id, night) PK | Overlap check = row predicate; works on any SQL or KV store with conditional writes; easy availability bitmaps | ~2.7B rows/year at 7.5M listings; pre-materialization job | Storage (cheap), platform (roll-forward job) |
| Range rows + exclusion constraint | One row per reservation; native overlap semantics in Postgres (daterange, GiST) | Postgres-only; GiST writes slower; harder on KV stores | DB team (engine lock-in) |
| Counters per (type, night) | Natural for pooled rooms; overbooking cap = allotment | Multi-night stay = N conditional decrements in one txn; hot rows on peak nights | Pooled-inventory team |
| Reservation list + application check | Easy to write | Race conditions without serializable isolation or locks | Guests (double bookings) |
Staff default: "Per-night rows. They carry status, price and minimum-stay rules per night, they produce the search bitmap directly, and the claim is a single conditional update on one shard. 2.7 billion small rows sharded by listing is unremarkable."
When to deviate:
- Postgres with modest scale: range + exclusion constraint is elegant and exact; use it if per-night pricing lives elsewhere.
- Hotels: counters per room type and night, with room assignment at check-in.
- DynamoDB: per-night items with
TransactWriteItemsand condition expressions; a transaction covers up to 100 items, which bounds max stay length per transaction — fine for 30-night maximums.
🧭 Principal Move: "The inventory model is the one-way door here. Every downstream — search bitmaps, pricing, analytics, host tools — will bind to it. I'd pick per-night rows because they're store-agnostic, which keeps our database choice a two-way door."
3.4 Fault Line 4: Zero Overbooking vs Managed Overbooking#
The tension: For pooled inventory, no-shows and late cancellations leave rooms empty that could have been sold. Overbooking captures that revenue and occasionally leaves a guest without a room. For unique listings, overbooking is simply double booking.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Zero overbooking (all inventory) | No walks; simple | 5–10% of pooled rooms empty on typical nights | Business (lost revenue) |
| Static cap (e.g., 3% of rooms) | Captures part of no-show revenue; simple to enforce | Walks on nights where no-shows under-run | Walked guests, front desk |
| Forecast-driven cap per night | Maximizes expected revenue | Model errors; complex; needs history | Revenue management (model owner) |
| Overbooking on unique listings | — | Guaranteed relocation for the loser | Never acceptable |
Staff default: "Unique listings: zero, enforced by constraint. Pooled inventory: allotment = physical rooms × (1 + cap), where the cap per night is set by revenue management within a hard ceiling — say 10% — that engineering enforces. Walk policy and budget are owned by operations."
When to deviate:
- Low no-show segments (non-refundable prepaid rates): zero cap for those rate plans; overbooking only against flexible rates.
- Brand-sensitive properties (luxury): very low caps; a walk costs more in reputation than a night's revenue.
🧭 Principal Move: "I'd want the expected-value table on one page: per property, P(no-shows ≥ k), revenue per extra room sold, walk cost, and estimated lifetime-value hit of a walk. Revenue management owns the model; finance signs off the ceiling; engineering makes the ceiling impossible to exceed."
3.5 Fault Line 5: Book-Then-Pay vs Pay-Then-Book (Workflow Coupling)#
The tension: The booking spans inventory, payment and sometimes host approval. Claiming inventory first risks holding unpaid inventory; paying first risks charging a guest for a room they can't have. Request-to-book adds a 24-hour gap.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Claim → authorize → confirm | Never charge without inventory; void on failure is free | Inventory held during payment (seconds) | Hosts (brief holds) |
| Authorize → claim → confirm | No unpaid holds | Claim failure after auth → void; guest sees a pending auth hold on their card | Guests (confusing pending charges) |
| Request-to-book, nights held while pending | Host can't be double-requested | Nights blocked up to 24h for maybe-bookings | Hosts (fewer instant bookings) |
| Request-to-book, nights not held | Other guests can still book | Host accepts a request whose nights were just instant-booked | Guests (declined after waiting) |
Staff default: "Claim → authorize → confirm, orchestrated as a saga with a durable state per reservation. Request-to-book holds the nights for the 24-hour window with the payment already authorized — so acceptance is one click and can't fail on inventory. Capture happens around check-in; payout to the host after check-in per policy."
When to deviate:
- Aggregator inventory: authorize, then confirm with the supplier; supplier rejection voids the auth.
- Long lead times (> 7 days) where auth holds expire: authorize a small verification amount at booking and charge the full amount closer to the stay, per payment policy — see Payment Processing.
🎯 Staff Insight: "The only unsafe compensation here is releasing inventory on a payment UNKNOWN. Everything else — voids, releases on decline, expiry — is safe to automate."
4. Failure Modes & Operational Reality#
4.1 Cross-Platform Double Booking via iCal#
t=0: Host lists cabin on us and on Platform V. Our import of V's iCal runs every 3h.
t=+10min: Guest X books Jun 14–16 on Platform V.
t=+40min: Guest Y instant-books Jun 14–16 on us. Our calendar: open. Claim succeeds.
t=+2h50m: iCal import brings V's block for Jun 14–16. Conflicts with Y's CONFIRMED booking.
t=+2h50m: sync.external_conflicts += 1. Host and guest notified; host has 24h to resolve.
t=+26h: Host cancels on us. Y gets full refund + rebooking credit; host penalized per policy.
Detection: sync.external_conflicts_total, broken down by source platform; sync.import_lag_seconds.
Mitigation: relocation flow for the guest; offer similar listings with credit.
Prevention: push-based APIs with channel managers instead of iCal polling; nudge multi-listed hosts to use a channel manager; for hosts with frequent conflicts, disable instant book.
Owner: supply/calendar-sync team (detection, sync), support (relocation), trust & policy (host penalties).
4.2 Hold Squatting on Peak Nights#
t=0: New Year's Eve inventory in a popular city: 3,000 listings left.
t=+1min: Scripted accounts start holds on 1,800 listings. Each hold 10 min; renewed on expiry.
t=+5min: Search still shows them "maybe available"; guests get 409s. Conversion down 40%.
t=+12min: holds.active_per_account anomaly + listing hold churn alert fires.
t=+15min: Accounts blocked; holds released. Per-account cap 2 and device checks enforced.
Detection: holds.active_count vs baseline, holds.per_account_p99, listing.hold_churn_per_hour, booking.claim_conflict_rate spikes.
Mitigation: release holds from flagged accounts; shorten TTL for scarce markets.
Prevention: holds require verified accounts; per-user and per-device caps; bot detection at the Reserve step.
Owner: booking team (caps), trust & safety (bot detection).
4.3 Hold Sweeper Down — Stranded Inventory#
t=0: Sweeper deploy fails; crashloop. No alert on "sweeper not running".
t=+2h: 60,000 expired holds not released. Search shows these listings unavailable.
t=+3h: Supply team notices hosts complaining about zero bookings.
Detection: holds.expired_unreleased_count (holds past expires_at still HELD), sweeper.last_run_age_seconds.
Mitigation: because the claim predicate treats expired holds as OPEN, bookings still succeed; only the search projection is wrong. Restart sweeper; it emits release events to fix search.
Prevention: expiry enforced in the claim predicate (correctness) plus sweeper (projection hygiene); heartbeat alert on the sweeper.
Owner: booking team.
4.4 Payment Timeout Releases the Hold (the Charged-for-Nothing Bug)#
t=0: Guest confirms. Payment auth times out after 15s.
t=+15s: Bug: booking marks reservation FAILED and releases nights.
t=+30s: Another guest books the same nights.
t=+2min: PSP webhook: first guest's auth succeeded. Reservation is FAILED; money is held.
t=+3 days: Guest sees a $420 pending charge for a trip they don't have. Support ticket.
Detection: booking.auth_succeeded_for_released_reservation_total — must page.
Mitigation: void the auth immediately; apologize proactively.
Prevention: PAYMENT_UNKNOWN state that extends the hold; release only on definitive decline.
Owner: booking team with payments.
4.5 Search Index Lag During a Sell-Out#
t=0: Concert announced; demand in one city spikes 30×.
t=+2min: Kafka consumer for inventory.changed lags: p99 5s → 90s.
t=+5min: search.stale_click_rate in that market: 0.4% → 11%. Conversion falls.
t=+8min: Page: stale_click_rate > 3% in any top-100 market for 5m.
t=+10min: Scale index updaters; enable live re-check of top 20 results for that market.
Detection: index.update_lag_p99, search.stale_click_rate{market}.
Mitigation: scale consumers; coalesce updates per listing; live re-check for hot markets.
Prevention: partition Kafka by listing_id with enough partitions for 10× peak; per-market stale-click alerts.
Owner: search team.
4.6 Time Zone and Date Bugs#
A listing in Honolulu stored with UTC timestamps: a booking for "Jun 14" made from New York at 11pm lands on Jun 15 UTC. Nights shift by one; a back-to-back booking now overlaps or leaves a phantom gap.
Prevention: nights are DATE in the listing's timezone, never timestamps; API accepts dates, not instants; DST changes don't matter because nights are dates.
Detection: booking.checkin_date_mismatch between confirmation email and calendar; host reports.
Owner: inventory service.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| External double booking | sync.external_conflicts_total | Individual stays; trust | Relocation playbook; push integrations | Supply sync + support |
| Hold squatting | holds.per_account_p99, churn | Market-level conversion | Caps, bot checks, shorter TTL | Booking + trust & safety |
| Sweeper down | sweeper.last_run_age_seconds | Search accuracy | Predicate expiry; restart | Booking team |
| Payment UNKNOWN mishandled | auth_succeeded_for_released | Individual guests, disputes | PAYMENT_UNKNOWN extends hold | Booking + payments |
| Index lag | index.update_lag_p99, stale clicks | Market conversion | Scale, coalesce, live re-check | Search team |
| Shard hot spot | Shard CPU, claim latency p99 | Listings on that shard | Rebalance listings; shard split | Inventory platform |
| Inventory DB primary failover | Claim errors, replication lag | Bookings on shard for ~30s | Retry with idempotency; no async-replica promotion without lag check | Inventory platform |
| Host mass-cancellation | cancellations.host_initiated spike | Guests with upcoming stays | Proactive relocation, host policy | Support + trust & policy |
🎯 Staff Insight: "Our own write paths can be made perfect with a constraint. The double bookings that actually reach guests come from writers we don't control. So the Staff investment is in detection speed and the relocation playbook, not in a fancier lock."
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | Lists features: search, book, pay, review | Unique vs pooled vs aggregator; instant vs request; two correctness bars | Asks who owns inventory truth and how many write paths exist across the company |
| Invariant | Lock around booking | Conditional claim / unique key in the inventory DB; shard by listing | One inventory gateway for every channel; bypasses tracked and removed |
| Search | Cache with TTL | Outbox/CDC projection with lag SLO and stale-click metric | Stale-click as a shared cross-team KPI with a quarterly target |
| Workflow | Book then pay, synchronous | Saga with holds, PAYMENT_UNKNOWN, host acceptance, safe compensations | Workflow platform shared by booking, experiences and partner channels |
| Policy | "Never overbook" | Zero for unique, capped for pooled, owned by revenue management | Expected-value model with finance-signed ceiling; walk budget as a line item |
| Operations | "Monitor bookings" | Conflict, hold, sweeper, lag and external-sync metrics with owners | Relocation cost and double-booking rate reported as trust metrics to leadership |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Puts the invariant in the store | "The update's WHERE clause is the lock." |
| Separates two consistency bars | "Search says maybe; booking says yes." |
| Handles payment uncertainty | "A payment timeout extends the hold; it never releases it." |
| Models nights correctly | "Local dates, half-open ranges — checkout and check-in on the same day don't conflict." |
| Treats external writers as untrusted | "iCal is a lagging writer; we detect and relocate." |
| Assigns policy owners | "Revenue management owns the overbooking cap; engineering enforces the ceiling." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| Distributed lock is the only protection | Lock TTLs and failovers break it; the DB can enforce it natively |
| Search reads the booking DB directly | Ignores 100–1000× read amplification |
| Shards by date | Multi-night stays span shards; hot dates create hot shards |
| Releases hold on payment timeout | Charges guests for nothing, or double-books |
| Timestamps for nights | Off-by-one-night bugs across time zones |
| No story for external calendars | Misses the dominant real-world source of double bookings |
5.4 Common False Positives#
- Geo-search sophistication ≠ reservation design. Quadtrees and H3 cells are useful, but they're the search team's problem; spending 15 minutes there skips the invariant.
- "Serializable isolation" ≠ a design. It works, but a candidate who can't say which rows conflict and why a conditional update suffices hasn't reasoned about contention.
- Saga vocabulary ≠ safe compensation. Naming "saga" is cheap; naming PAYMENT_UNKNOWN as the non-compensable state is the signal.
- Big numbers ≠ scale reasoning. "Billions of rows" isn't scary when each shard has millions; the candidate should say why sharding by listing makes it boring.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Unique vs pooled; instant vs request; two correctness bars; numbers |
| Entities & API | 3–5 min | Nights, reservations, holds, blocks; half-open local dates |
| Architecture | 5–10 min | Read path vs write path; inventory gateway; outbox |
| Claim + holds | 10–18 min | Conditional update; TTL; abuse controls |
| Saga + payment | 18–26 min | States, compensations, UNKNOWN, request-to-book |
| Search freshness + external sync | 26–34 min | Projection, lag SLO, stale clicks, iCal conflicts |
| Pivot | 34–42 min | Hotels/overbooking, flash demand, multi-region, partner API |
| Wrap | 42–45 min | Consistency placement; owners; evolution |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "Now it's a hotel with 200 identical rooms" | Model flexibility | Counters per type/night; overbooking cap; assignment at check-in |
| "Taylor Swift concert: 50K people want 3K listings in one city" | Hot-market behavior | Per-listing contention still low; search lag and hold squatting are the risks |
| "Make it multi-region" | Home-region ownership | Listing's home region owns its nights; reads anywhere; writes routed home |
| "Hosts want to change prices constantly" | Price consistency | Price snapshot at hold; search price is an estimate |
| "Add a partner API for channel managers" | Write-path governance | Same inventory gateway; idempotent upserts; conflict reporting |
| "A guest was double-booked. Investigate." | Forensics | Reservation history + block provenance + sync logs in one query |
6.3 What to Deliberately Skip#
- Ranking and personalization — "a separate ML-driven ranking service consumes the candidate set."
- Reviews, messaging, photos — out of scope unless asked.
- Dynamic pricing algorithms — "pricing service produces nightly prices; we snapshot at hold."
- Map tiles and geo indexing internals — one sentence: geo-shape query in the index.
6.4 Follow-Up Questions to Expect#
- "Two guests click Reserve at the same millisecond. Walk me through the database."
- "What does search show for a listing booked two seconds ago?"
- "What stops someone from holding every listing in a city?"
- "The payment provider times out. Is the room booked?"
- "The host is also on another platform. How do you prevent double booking?"
- "How would the model change for a 300-room hotel?"
- "How do you change the hold TTL without breaking in-flight checkouts?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design Airbnb."
Staff Answer
"Airbnb is two systems with different correctness bars: search, which is read-heavy — tens of thousands of QPS — and can be seconds stale, and booking, which is a few hundred per second and must never double-book. I'll assume unique listings, both instant book and request-to-book, and hosts who also sync calendars from other platforms.
I'll focus on booking correctness and the gap between search and booking: nights as rows sharded by listing, a conditional claim that enforces one reservation per night, holds with a TTL, the booking saga with payment and host acceptance, a search projection fed by an outbox, and how we detect and handle double bookings from external calendars. Ranking and reviews I'd treat as separate systems."
Why this is L6:
- Splits into two systems with two correctness bars and numbers
- States inventory assumptions that change the design
- Scopes out ranking explicitly to protect time
What L7 adds:
- Asks how many channels write inventory (web, apps, partner APIs, iCal, support tools)
- Frames supply (calendars, sync) and demand (search, checkout) as separate org domains with a contract
❌ Common L5 Trap
"Let's start with the search — we'll use a geohash-based index, then ranking with ML features like price, reviews and location…"
Why this misses: Fifteen minutes on the part that doesn't need strong consistency, leaving the invariant for the last five minutes, where it gets a Redis lock.
Drill 2: The Race#
Prompt: "Two guests reserve the same nights at the same instant. What happens, exactly?"
Staff Answer
"Both requests reach the inventory service and route to the listing's shard. Each runs UPDATE night_inventory SET status='HELD', reservation_id=? WHERE listing_id=? AND night IN [14,15] AND (status='OPEN' OR (status='HELD' AND hold_expires_at < now())). Postgres takes row locks in the order it touches rows. Request A locks both rows first; B blocks on the first row it needs. A commits. B re-checks the predicate on the now-HELD rows, updates 0 of 2, and — because rows updated ≠ nights requested — rolls back and returns 409 with alternatives. With a partial overlap — A wants 14–15, B wants 15–16 — B updates night 16 but not 15, so it still rolls back. No external lock involved; the guarantee holds through pauses and retries."
Why this is L6:
- Walks through the actual database behavior, including re-evaluation under READ COMMITTED
- Handles partial overlaps explicitly
- Explains why no distributed lock is needed
What L7 adds:
- Ensures every channel uses this same path, and audits for writes that bypass it
- Considers deadlock ordering across multi-night claims (sort nights ascending) as a platform library concern
Drill 3: Search Staleness — Make It Concrete#
Prompt: "How stale can search be, and how would you know if it's too stale?"
Staff Answer
"Target p99 lag of 5 seconds from commit to index. Pipeline: outbox row in the claim transaction → relay → Kafka partitioned by listing_id → updater coalescing per listing over 500ms → index bitmap update. Two metrics: index.update_lag_p99 for the mechanism, and search.stale_click_rate for the outcome — the share of listing views where the searched dates are already taken. Target < 0.5% overall, alert at 3% in any top-100 market. When stale clicks rise but lag doesn't, it's demand, and I'd enable a live re-check of the first page of results against a replica for that market."
Why this is L6:
- Quantifies lag and defines an outcome metric, not just a pipeline metric
- Distinguishes pipeline failure from demand-driven staleness
- Has a targeted mitigation that doesn't make search strongly consistent everywhere
What L7 adds:
- Makes stale-click rate a shared KPI with supply (host calendar accuracy)
- Budgets the replica capacity for live re-checks as a seasonal plan
Drill 4: The Payment Timeout#
Prompt: "The guest confirms, and the payment call times out. What state is everything in?"
Staff Answer
"Reservation moves to PAYMENT_UNKNOWN; nights stay HELD and the hold is extended — say to 30 minutes — because the auth might have succeeded. The guest sees 'Confirming your booking'. A resolver re-queries the payment by intent ID at 1, 5 and 15 minutes, and the PSP webhook may resolve it first. Authorized → nights BOOKED, reservation CONFIRMED. Declined → nights released, guest asked for another card. Still unknown at 30 minutes → extend once more and route to a manual queue; I'd rather block the listing for another half hour than risk charging a guest for a room someone else booked."
Why this is L6:
- Identifies the one non-compensable state and handles it
- Uses idempotent re-query, not a fresh charge
- Makes the tradeoff explicit: host inventory time vs guest charged-for-nothing
What L7 adds:
- Aligns with the payments platform's UNKNOWN standard so booking doesn't invent its own
- Tracks
money_in_unknownfor reservations in finance reporting
Drill 5: Hot Market#
Prompt: "A huge event: 60,000 people searching for 2,500 remaining listings in one city over 10 minutes."
Staff Answer
"Per-listing contention is still small — 60,000 people spread over 2,500 listings is ~24 per listing over 10 minutes; the conditional claim handles that trivially. The real risks are elsewhere: search showing listings already gone, because lag grows and inventory turns over faster than the index; hold squatting; and conversion collapse from repeated 409s. So: live re-check of the top results for that market, 5-minute holds and max 1 active hold per user in that market, 409 responses that immediately suggest the next-best available listings, and a banner that says 'high demand — availability changing quickly' so disappointment is expected."
Why this is L6:
- Does the contention arithmetic instead of assuming it's a hot-key problem
- Finds the actual bottleneck (search accuracy and holds)
- Includes UX as part of the mitigation
What L7 adds:
- Pre-event playbook triggered by demand forecasts, owned by a marketplace ops function
- Coordinates with supply to recruit inventory before the event, since supply is the real constraint
Drill 6: Pooled Hotel Inventory#
Prompt: "We're adding hotels with room types — 150 identical Standard Kings. What changes?"
Staff Answer
"The invariant becomes a count: sold ≤ allotment per (property, room type, night). Model: counter rows with allotment and sold; a 3-night booking does three conditional increments UPDATE … SET sold = sold + 1 WHERE … AND sold < allotment in one transaction, nights sorted to avoid deadlocks. Room numbers are assigned at check-in by the hotel. Overbooking: allotment = 150 × (1 + cap), cap set per night by revenue management within a 10% ceiling. The hot-row risk is real on peak nights for big properties — thousands of bookings on one row — so if contention grows, I'd split the counter into 8 sub-buckets per night and try them in random order, accepting that a sold-out check needs all 8."
Why this is L6:
- Swaps the invariant form cleanly from uniqueness to count
- Places overbooking policy with an owner and a ceiling
- Anticipates hot counters with a concrete mitigation
What L7 adds:
- Recognizes most hotel inventory arrives via channel managers — prioritizes a push partner API over building a native hotel extranet
- Evaluates whether pooled inventory belongs on the same platform or a separate one
Drill 7: Build vs Buy#
Prompt: "Should we build our own calendar-sync integrations with other platforms, or partner with channel managers?"
Staff Answer
"Partner. iCal import is the lowest common denominator with hour-scale lag. Direct platform-to-platform integrations are N² and politically hard. Channel managers already integrate with major platforms and push availability changes in near real time. I'd expose a partner API — idempotent calendar upserts, reservation webhooks, conflict reporting — certify the top channel managers, and keep iCal for small hosts. Success metric: external-conflict rate per 1,000 bookings for channel-managed listings vs iCal listings."
Why this is L6:
- Compares options on lag and integration count
- Defines the API shape we'd need either way
- Measures the outcome by conflict rate
What L7 adds:
- Structures partner certification tiers and SLAs (e.g., push within 60s)
- Considers the competitive dynamics: channel managers also serve competitors
Drill 8: Changing Hold Policy Safely#
Prompt: "Product wants to drop the hold TTL from 10 to 5 minutes. How do you roll it out?"
Staff Answer
"The TTL is read at hold creation and stored on the hold, so in-flight holds keep their 10 minutes — no one loses a room mid-checkout. Roll out as an experiment: 5% of new holds at 5 minutes for a week, measuring hold-expiry-before-confirm rate, conversion, and squatting indicators. Guardrail: 5 minutes must exceed p99 time-from-hold-to-payment-submit plus p99 auth latency; if p99 checkout time is 6 minutes, 5 minutes will cost conversion. Then ramp by market, with a kill switch in config."
Why this is L6:
- Protects in-flight state by snapshotting policy at creation
- Turns a policy change into a measured experiment with an engineering floor
- Defines rollback
What L7 adds:
- Makes hold TTL a per-market, demand-aware policy owned by marketplace product, with engineering floors published
Drill 9: Multi-Region#
Prompt: "Make it multi-region for latency and resilience."
Staff Answer
"Search goes fully multi-region: indexes replicated per region, fed by the global event stream — staleness is already accepted. Booking uses home-region ownership: each listing's nights live in one region, typically near the listing. A guest in Tokyo booking a Paris flat sends the claim to the EU region — 200ms extra RTT once per booking is fine. If the home region fails, bookings for its listings pause; I would not fail over writes to an async replica, because it may be missing the last seconds of claims. For region-level resilience, I'd use synchronous replication within the region across zones and a manual, lag-checked regional failover."
Why this is L6:
- Splits the answer by correctness bar again
- Single-writer ownership per listing avoids cross-region conflicts
- Explicitly chooses unavailability over possible double booking
What L7 adds:
- Aligns home regions with data-residency rules for guest data
- Prices the second-region standby against the revenue at risk per hour of regional outage
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Incident — New Year's Eve Conversion Collapse#
Context: On Dec 28, conversion in the top 20 cities falls 35% over two hours. No errors, latency normal. Escalated to you.
Questions to Surface First:
- Are guests getting 409s at Reserve? At what rate vs normal?
- What's
search.stale_click_rateand index lag in those cities? - Are active holds unusually high? Concentrated in a few accounts?
Typical L5 Approach: Scales search and booking services; checks error rates and latency; finds nothing wrong.
Staff Approach: Looks at outcome metrics, not health metrics: 409 rate at Reserve up from 2% to 22%, stale clicks at 14%, active holds 3× normal. Finds a reseller operation holding listings via 400 accounts. Blocks the accounts, caps holds per device, shortens TTL to 5 minutes in hot markets, enables live re-check.
Principal Approach: Establishes a peak-season readiness process: demand forecasts trigger pre-configured hot-market policies (TTL, caps, re-check) automatically, owned jointly by marketplace ops and trust & safety, reviewed after each peak.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Pull 409 rate, stale-click rate and active holds by market. |
| Triage | 38% of active holds in top cities from 400 accounts created in the last week. |
| Quick fix | Release those holds; block accounts; per-device hold cap; TTL 5 min in top markets. |
| Guardrails | Alerts on hold concentration and per-market 409 rate. |
| Post-mortem | Health dashboards were green while the business was failing — add outcome SLOs. |
Metrics to Watch: booking.claim_conflict_rate{market}, holds.active{market}, holds.per_account_p99, search.stale_click_rate{market}, conversion{market}
Organizational Follow-up: trust & safety adds hold-squatting to its detection models; marketplace ops owns hot-market policy.
Ownership Question: "Who decides to shorten TTLs mid-incident?" Staff answer: The booking on-call, from a pre-approved runbook range (5–10 min). Outside that range, marketplace product signs off.
Key Takeaway: "Green health metrics can hide a failing marketplace. Watch conflict and stale-click rates."
What clears the Staff bar:
- Uses business outcome metrics to find a problem invisible to system health
- Recognizes abuse of a correctness mechanism (holds)
- Pre-approves the mitigation range
Deep Dive 2: Silent Failure — The Import That Stopped#
Context: A host complains of three double bookings in a month. Investigation shows that iCal imports from one partner platform have silently failed for 11 days for 40,000 listings — the partner changed a URL format.
Questions to Surface First:
- How many listings, how many upcoming reservations now conflict with the partner's calendars?
- Why did nothing alert?
- Do we know a feed is stale vs empty?
Typical L5 Approach: Fixes the URL parsing, re-runs imports.
Staff Approach: Fixes parsing, re-runs imports, and computes the conflict set: 312 upcoming stays overlap newly-imported blocks. Kicks off proactive host outreach and guest relocation before arrival dates. Adds
sync.feed_last_success_ageper source and alerts when a source's success rate drops.
Principal Approach: Treats external sync as a tier-1 supply dependency with per-source SLOs and a partner-change communication channel; accelerates migration of high-volume sources to push-based partner APIs.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Fix parser; backfill imports for affected listings. |
| Triage | Compute conflicts; sort by arrival date; stays in next 7 days first. |
| Remediation | Host outreach with 24h resolution window; relocation offers for guests. |
| Guardrails | Per-source success rate and feed age alerts; alert on sudden drop in blocks imported per source. |
| Post-mortem | Import errors were logged, not metered; no per-source view. |
Metrics to Watch: sync.feed_last_success_age{source}, sync.import_error_rate{source}, sync.external_conflicts_total
Ownership Question: "Who owns relocation cost for these 312 stays?" Staff answer: The platform absorbs it — the failure was ours. Support executes; supply sync owns the fix.
Key Takeaway: "An external sync that fails silently is a double-booking generator with a delay."
What clears the Staff bar:
- Computes and prioritizes the blast radius by arrival date
- Acts before guests arrive, not after
- Alerts on data freshness per source, not only on errors
Deep Dive 3: Large-Customer Onboarding — A 12,000-Unit Property Manager#
Context: A professional property manager with 12,000 units wants to onboard via API with real-time calendar updates and bulk pricing changes (all 12,000 units × 365 nights, daily).
Questions to Surface First:
- Write volume: 4.4M night updates/day in bursts?
- Are they the source of truth for availability, or are we?
- How do conflicts get resolved when both sides book the same unit?
Typical L5 Approach: Gives them the host API with higher rate limits.
Staff Approach: Designs a partner API with bulk, idempotent, versioned upserts (
listing, night range, price, status, version), processed per listing in order; pricing writes go to a separate price table so they don't contend with claims; availability blocks from the partner go through the inventory gateway with source provenance; reservation webhooks to the partner with retries. Rate: 4.4M updates/day ≈ 50/s average, spread over shards — fine if batched per listing.
Principal Approach: Uses this onboarding to define the partner platform tier — contract, SLAs, certification tests, conflict responsibility — so the next 50 professional managers onboard through a product, not a project.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Design | Separate price writes from availability writes; batch per listing. |
| Conflict rule | Our CONFIRMED reservations win; partner block on a booked night → conflict event to partner within 60s. |
| Load | Burst test at 10× their average; per-partner write quota. |
| Certification | Partner must handle webhooks idempotently and ack within 5s. |
| Launch | 100 units → 1,000 → 12,000 over 3 weeks, watching conflict rate. |
Metrics to Watch: partner.upsert_lag{partner}, partner.conflicts_per_1k_bookings{partner}, shard write latency
Ownership Question: "Who pays for a double booking caused by the partner's delayed push?" Staff answer: Defined in the partner contract — typically the partner bears it past their SLA; we bear it within ours.
Key Takeaway: "Big partners need a product: bulk APIs, quotas, certification, and a contract for conflicts."
What clears the Staff bar:
- Separates hot claim rows from bulk pricing writes
- Defines conflict precedence explicitly
- Stages the rollout with a conflict metric
Deep Dive 4: Post-Mortem — 180 Double Bookings from a Support Tool#
Context: A support tool that lets agents "force-book" guests into listings during relocations wrote reservations directly to the reservations table, bypassing the inventory service. Over 6 months it produced 180 double bookings.
Questions to Surface First:
- How many other writers bypass the inventory service?
- Why did the database allow it — is the constraint on nights or on reservations?
- Why did 180 incidents over 6 months not trigger a pattern review?
Typical L5 Approach: Changes the support tool to call the inventory service.
Staff Approach: Fixes the tool, then makes bypass impossible: the uniqueness constraint lives on
night_inventory, which only the inventory service's DB role can write; the reservations table gets a foreign-key-like consistency check job; any reservation without claimed nights pages. Audits all writers — finds two more (a migration script, a partner backfill).
Principal Approach: Codifies "single writer for inventory" as an architectural invariant enforced by database permissions and CI checks, and adds double-booking root-cause categorization to the quarterly trust review so patterns are caught at 10, not 180.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Disable force-book; route relocations through the inventory service with a relocation flag. |
| Audit | List every DB role with write access to inventory tables; remove all but the service. |
| Detection | Nightly consistency job: every CONFIRMED reservation has matching BOOKED nights; mismatch pages. |
| Remediation | Contact affected upcoming guests. |
| Post-mortem | Invariant enforced by convention, not by permission. |
Metrics to Watch: inventory.consistency_violations, double_bookings_by_root_cause
Ownership Question: "Who owns the invariant?" Staff answer: The inventory team — and owning it means owning every path that can violate it, including other teams' tools.
Key Takeaway: "An invariant enforced by convention has a bypass. Enforce it by permission."
What clears the Staff bar:
- Moves enforcement from code review to database permissions
- Audits for all bypasses, not just the one found
- Categorizes root causes to catch patterns early
Deep Dive 5: Multi-Region Expansion — Launching in Japan#
Context: The company launches a Japan region for latency and data residency. Some listings will be in Japan; guests worldwide book them; Japanese guests book worldwide.
Questions to Surface First:
- What data must stay in Japan — guest PII, payment data, listing data?
- Where should a Japanese listing's inventory live?
- What happens to Japanese bookings if the Japan region is down?
Typical L5 Approach: Deploys the full stack in Japan with multi-master database replication.
Staff Approach: Home-region ownership: Japanese listings' nights live in the Japan region; claims route there from any guest region. Search indexes are replicated globally. Guest PII for Japanese residents stays in Japan; reservations referencing them store a regional pointer. No multi-master for inventory — single writer per listing.
Principal Approach: Makes "home region per listing, home region per guest" the platform pattern with a routing layer and residency classification, so the next region launch is configuration plus capacity, not architecture.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Classification | Legal classifies fields: guest PII resident; listing and inventory not restricted. |
| Routing | Listing → home region map; booking service routes claims. |
| Migration | Move Japanese listings' shards to Japan region with dual-write verification window. |
| Failure | Japan region down → Japanese listings unbookable; search still shows them with a "temporarily unavailable" flag. |
| Launch | 1% of listings → 100% over a month. |
Metrics to Watch: booking.cross_region_claim_latency_p99, routing.misrouted_claims, search.stale_click_rate{region}
Ownership Question: "Who decides that Japanese listings go unbookable during a regional outage?" Staff answer: Product and trust leadership, informed by the relocation cost of a double booking vs lost bookings per hour.
Key Takeaway: "Replicate reads everywhere; own writes in one place."
What clears the Staff bar:
- Single-writer per listing across regions
- Separates residency of guest data from location of inventory
- Chooses unavailability over conflicting writes, with a named decision-maker
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Split a reservation system into a stale-tolerant read path and a strongly consistent write path, with numbers for each
- Enforce the no-double-booking invariant with a database constraint or conditional update — and explain why no distributed lock is needed
- Model nights as local dates with half-open ranges, and choose per-night rows, range exclusion or counters for the inventory type
- Design holds with TTLs, abuse caps and predicate-based expiry
- Orchestrate the booking saga, including the non-compensable PAYMENT_UNKNOWN state and request-to-book
- Measure search staleness with a stale-click metric and mitigate hot markets
- Treat external calendars and partners as untrusted writers with detection and relocation playbooks
- Place overbooking policy with revenue management, bounded by an engineering ceiling
The Bar for This Question#
Mid-level (L4): Builds search and booking with a single database, checks availability then inserts. Works in low traffic; races under concurrency.
Senior (L5): Adds a distributed lock or serializable transactions, caches availability for search, integrates payments. Reasonable and mostly correct. The gap: one consistency level everywhere, holds without abuse controls, payment timeouts treated as failures, no story for external writers, and "never overbook" as a universal rule.
Staff+ (L6): Places the invariant in the inventory store and makes everything else a measured projection. Designs the booking saga with safe compensations. Shards by listing so every booking is single-shard. Treats external calendar sync as the primary real-world double-booking source and designs detection and relocation. Assigns policy owners for hold TTLs and overbooking. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 You Don't Need a Distributed Lock to Prevent Double Booking#
| Approach | Guarantee |
|---|---|
| Redis lock + read + write | Probabilistic; breaks on TTL expiry, pauses, failover |
| Conditional update on night rows | Exact; the store is both lock and data |
| Exclusion constraint on ranges | Exact; declarative |
The Staff position: When one store holds the truth and supports conditional writes, an external lock is pure added risk.
Why this matters in interviews: It's the fastest way to show you understand where correctness actually comes from.
10.2 Search Should Be Wrong Sometimes — On Purpose#
The Staff position: A search that's never stale costs 100–1,000× the reads of booking. A search that's stale 0.3% of the time, with a graceful "just booked — here are alternatives", costs almost nothing in conversion. Choose the second and measure it.
Why this matters in interviews: Staff candidates choose inconsistency deliberately and put a number on it.
10.3 Most Double Bookings Aren't Race Conditions#
| Source | Share of real incidents (typical pattern) |
|---|---|
| External calendar lag | Dominant |
| Bypassing write paths (tools, scripts, partners) | Significant |
| Host manual error | Significant |
| Concurrent requests on our own path | Near zero once constrained |
The Staff position: After the constraint is in, invest in writer governance and detection, not in fancier concurrency control.
Why this matters in interviews: It reframes the problem from algorithms to operations — the L6 shift.
10.4 Overbooking Is Good Engineering — For the Right Inventory#
The Staff position: For pooled inventory with measurable no-shows, a capped overbooking policy is the revenue-maximizing, customer-respecting choice when paired with a funded walk policy. Refusing it isn't "safe"; it's leaving rooms empty.
Why this matters in interviews: It shows you distinguish business policy from correctness, and that you won't impose an engineering absolute on a business decision.
10.5 The Hold TTL Is a Product Decision Disguised as a Constant#
The Staff position: Engineering sets the floor (checkout p99 + auth p99); product sets the value per market and season. A hard-coded HOLD_TTL = 600 is a policy nobody owns.
Why this matters in interviews: Naming owners for "constants" is a crisp ownership signal.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
The Staff engineer builds a booking path that cannot double-book. The Principal engineer notices that the company now sells stays, experiences, hotel rooms via partners and long-term rentals — four inventory types — through web, apps, partner APIs, iCal and a support tool, and that each product team wants its own inventory service. The L7 problem is inventory as a platform: one gateway for all writes, pluggable inventory models (unique, pooled, time-slot), shared hold and saga machinery, and an org boundary between supply (who owns calendars and partners) and demand (who owns search and checkout) with a contract both can be held to.
The Org-Level Fault Line#
One inventory platform vs per-product inventory services.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Per-product inventory | Each product moves fast | N implementations of holds, sagas, sync; partner APIs differ per product; bypasses multiply | Partners, support, trust (inconsistent behavior) |
| One monolithic inventory service | One invariant owner | Every product waits on one team; model forced to fit all | Product velocity |
| Inventory platform with pluggable models | Shared gateway, holds, saga, events, partner API; product-specific models as plugins | Platform API design is hard; needs strong stewardship | Platform team (4–8 engineers) |
🧭 Principal Move: "The gateway, the hold machinery, the event stream and the partner API are platform. The inventory model — nights, room-type counters, experience time slots — is a plugin each product owns. The invariant is the platform's; the semantics are the product's."
Cost Model#
Assumptions: cloud infrastructure, engineer fully loaded ~$250K/year; relocation cost ~$300 per incident all-in (alternative lodging difference, credits, support time); order-of-magnitude figures.
| Scale | Listings / Bookings per day | Infra ($/month) | Headcount | On-call | Double-booking cost |
|---|---|---|---|---|---|
| Startup | 20K listings / 1K bookings | ~$3–8K (one Postgres, managed search) | 4–6 eng total across search + booking | Shared rotation | ~5/month × $300 ≈ $1.5K |
| Growth | 500K / 50K | ~$40–80K (sharded Postgres, Kafka, search cluster) | 15–25 eng: search (6–8), booking/inventory (5–8), supply sync (3–5) | Per-domain rotations | ~200/month ≈ $60K + trust cost |
| Enterprise | 7M+ / 1M+ | ~$0.5–1.5M (multi-region search, sharded inventory, streaming) | 80–150 eng across search, inventory platform, supply, partners | Follow-the-sun | Rate target < 1 per 10K bookings; each 0.01% ≈ $30K/day |
The pricing insight: at growth scale, a 3-engineer investment in push-based partner integrations that halves external conflicts saves roughly its own cost in relocations alone — before counting trust and host churn.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Inventory unit model (nights as local dates, half-open) | One-way | Every consumer, partner and report binds to it |
| Shard key (listing_id) | One-way-ish | Resharding billions of rows with live bookings |
| Partner API semantics (who wins on conflict) | One-way | Contracts and partner integrations depend on it |
| Single-writer inventory gateway | One-way (in the good direction) | Once enforced, reintroducing bypasses is a choice, not drift |
| Hold TTL, caps | Two-way | Config per market |
| Search index technology | Two-way | Projection can be rebuilt from events |
| Overbooking caps for pooled inventory | Two-way | Policy change with sign-off |
The Standard I'd Write#
RFC-INV-001: Inventory Integrity Standard
Status: Approved Owners: Inventory Platform + Marketplace Trust
Scope
All systems that create, modify or cancel reservations or availability for any
inventory product (stays, hotels, experiences), including tools and partner APIs.
MUST
1. Write availability and reservations only through the Inventory Gateway.
Database write permissions are restricted to the gateway's role.
2. Represent time inventory as local dates (or local slots) with half-open ranges.
3. Enforce the product's invariant (uniqueness or count ≤ allotment) in the store,
not in application locks.
4. Treat payment and supplier timeouts as UNKNOWN; never release claimed inventory
on an UNKNOWN outcome.
5. Record provenance (source, external id) on every block and reservation.
SHOULD
1. Keep search projections within a 5s p99 lag SLO and report stale-click rate.
2. Prefer push-based partner integrations over polling.
3. Snapshot policy (hold TTL, price) at creation time.
Exceptions
Filed with Inventory Platform; overbooking on unique inventory is not an allowed exception.
Overbooking caps for pooled inventory require revenue-management and finance sign-off.
Success metrics
- Double bookings from internal paths: 0
- Total double bookings per 10K bookings: < 1
- Median time to detect external conflict: < 15 minutes
- Stale-click rate: < 0.5%
What I'd Tell the VP#
"Our booking system can't double-book on its own — the database guarantees it. The double bookings guests actually experience come from calendars we sync from other platforms and from internal tools that skip our checks, and each one costs us about $300 plus a guest who may not come back. I'm proposing two things: make one inventory gateway the only way to change availability, and invest in real-time partner integrations to replace hourly calendar polling. Together that's roughly six engineers for two quarters, it should cut double bookings by more than half, and it's the foundation for selling hotel and experience inventory on the same platform next year."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Governs writers, not just writes | "How many systems can mark a night unavailable? Each is a bypass risk." |
| Prices trust failures | "Each double booking is ~$300 plus churn; the partner API pays for itself." |
| Separates platform from product semantics | "Gateway and holds are platform; the unit model is a plugin." |
| Places policy with owners | "Revenue management owns the cap; finance signs the ceiling." |
| Identifies one-way doors | "The inventory unit model and partner conflict semantics are the decisions I'd slow down on." |
Staff answers that L7 interviewers find insufficient:
- "Unique constraint on (listing, night) solves double booking" — solves it on our path; silent on the tools, scripts and partners that bypass it.
- "We'll detect iCal conflicts and relocate guests" — good operations, no strategy for reducing the source (push partners, host nudges).
- "Hotels need counters and overbooking" — correct model, no answer for whether it belongs on the same platform or who owns it.
Appendices
Appendix A: Mechanics in Depth#
A.1 Claim with Per-Night Rows (Postgres)#
-- Claim nights [in, out) for a hold; nights sorted to keep lock order consistent.
WITH target AS (
SELECT night FROM generate_series($in::date, $out::date - 1, '1 day') AS night ORDER BY night
)
UPDATE night_inventory n
SET status = 'HELD', reservation_id = $res, hold_expires_at = now() + $ttl
FROM target t
WHERE n.listing_id = $listing AND n.night = t.night
AND (n.status = 'OPEN' OR (n.status = 'HELD' AND n.hold_expires_at < now()));
-- application: IF row_count <> ($out - $in) THEN ROLLBACK; RETURN 409
A.2 Range Exclusion Alternative#
CREATE EXTENSION IF NOT EXISTS btree_gist;
CREATE TABLE reservations (
reservation_id UUID PRIMARY KEY,
listing_id BIGINT NOT NULL,
stay DATERANGE NOT NULL, -- '[2026-06-14,2026-06-16)'
state TEXT NOT NULL,
EXCLUDE USING gist (listing_id WITH =, stay WITH &&)
WHERE (state IN ('HELD','PAYMENT_PENDING','PAYMENT_UNKNOWN','PENDING_HOST','CONFIRMED'))
);
Expired holds must be moved out of the constrained states (by the sweeper) before the nights are claimable — the predicate trick from A.1 doesn't apply, so sweeper health becomes a correctness dependency. That's one reason to prefer per-night rows.
A.3 Pooled Counters#
UPDATE room_type_inventory
SET sold = sold + 1
WHERE property_id = $p AND room_type = $rt AND night = ANY($nights)
AND sold < allotment; -- allotment includes overbooking cap
-- require row_count = cardinality($nights), else ROLLBACK
A.4 Hold Sweeper#
every 30s:
UPDATE night_inventory SET status='OPEN', reservation_id=NULL
WHERE status='HELD' AND hold_expires_at < now() - interval '5 seconds'
RETURNING listing_id, night -- emit inventory.changed via outbox
batch size 5,000 per shard; alert if last_run_age > 120s
Appendix B: Data Model#
listings(listing_id PK, host_id, timezone, geo, capacity, instant_book, min_nights, home_region)
night_inventory(listing_id, night DATE, status, reservation_id, hold_expires_at,
price_minor, min_stay, source, external_uid, PK(listing_id, night))
reservations(reservation_id PK, listing_id, check_in, check_out, guest_id, state,
price_snapshot JSONB, payment_intent_id, hold_ttl_s, version, created_at)
calendar_blocks(listing_id, night, source, external_uid, imported_at)
outbox(id BIGSERIAL, topic, key, payload, created_at, published_at)
- Shard key:
listing_idfor all tables above → every booking is single-shard. - Pre-materialize nights 540 days ahead; nightly job rolls the window forward.
- Row volume: 7.5M listings × 540 nights ≈ 4B rows; at ~100 bytes each, ~400 GB + indexes — split across 32–64 shards, ~10–20 GB per shard.
Appendix C: Coordination Mechanisms — Quick Comparison#
| Mechanism | Guarantee | Failure Mode | Use For |
|---|---|---|---|
| Conditional update on night rows | Exact per night | Deadlocks without sorted order | Unique listings (default) |
| Exclusion constraint | Exact per range | Sweeper becomes correctness-critical | Postgres, range-centric models |
Counter with sold < allotment | Exact count | Hot rows on peak nights | Pooled rooms |
DynamoDB TransactWriteItems + conditions | Exact, ≤ 100 items | Transaction conflicts under contention | KV-based inventory |
| Redis lock | None under pauses/failover | Double booking | Never as the guarantee |
| Serializable isolation | Exact | Retry storms under contention | Small systems, simple code |
Appendix D: API Contract & Client Behavior#
POST /reservationsrequiresIdempotency-Key; replay returns the same reservation (same hold) — double-clicks don't create two holds.- 409 responses include
alternatives[](similar listings with availability) so the client never dead-ends. PAYMENT_PROCESSINGstate is surfaced as "Confirming your booking"; client polls with backoff (1s, 2s, 5s, cap 15s) or receives push.- Search results carry
availability_as_ofso clients and logs can correlate stale clicks with lag. - Webhooks to partners: signed,
event_idfor dedup, retried with backoff up to 72 hours.
Appendix E: Observability#
Core metrics:
booking.claim_conflict_rate{market}— 409s at Reservesearch.stale_click_rate{market},index.update_lag_p99holds.active{market},holds.per_account_p99,holds.expired_unreleased_count,sweeper.last_run_age_secondsbooking.saga_stuck{state}— reservations in a non-terminal state past its expected dwellbooking.auth_succeeded_for_released_reservation_total(must be 0)sync.external_conflicts_total{source},sync.feed_last_success_age{source}inventory.consistency_violations(reservation without BOOKED nights, or vice versa)
Critical alerts:
| Alert | Threshold | Severity |
|---|---|---|
| Consistency violation | > 0 | Sev-1 page |
| Auth succeeded for released reservation | > 0 | Page |
| Stale-click rate | > 3% in a top-100 market for 5m | Page search on-call |
| Sweeper not running | last run > 120s | Page |
| Feed stale | source success age > 2× interval | Page supply sync |
| Saga stuck | PAYMENT_UNKNOWN > 30 min count > 50 | Page booking |
| Claim conflict rate | > 3× baseline in a market | Warn → investigate squatting |
Debugging the silent failure: a market where bookings drop but errors don't rise usually means holds are squatting inventory or search is showing unavailable listings. Compare holds.active and stale_click_rate for that market against the same weekday last year.
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 50K listings | One Postgres; search via SQL + cache | Search latency with date filters |
| 50K–1M | Search index with bitmaps via outbox; single inventory primary | Write latency, backup size |
| 1M–10M | Inventory sharded by listing (32–64 shards); Kafka; per-market metrics | External sync volume; partner onboarding |
| 10M+ / multi-product | Inventory platform, home regions, pluggable models | Org coordination and partner governance |
What you don't build on day one: sharding, multi-region writes, pooled inventory, partner APIs, dynamic hold TTLs. Build: per-night rows, conditional claims, holds with predicate expiry, the outbox, and stale-click measurement.
Appendix G: Multi-Tenancy, Fairness & Cost#
- Hosts as tenants: large property managers get per-partner write quotas so bulk pricing updates don't delay claims on shared shards; price writes go to a separate table.
- Guest fairness under scarcity: per-user hold caps and verified accounts prevent a few actors from monopolizing peak inventory.
- Market fairness in search capacity: per-market query budgets during events so one city's surge doesn't raise latency globally.
- Cost attribution: relocation costs tagged by root cause (external sync, host cancellation, internal bug) and charged to the owning domain's budget — the fastest way to fund the fixes.