Technologies referenced in this case study: Elasticsearch · Redis · Cassandra · DynamoDB · Apache Kafka
Related case studies: Maps & Geospatial · News Feed · Notification System · Chat Messaging · Leaderboard
How to Use This Case Study#
Organized for interview use first, reference second. Read front-to-back once, then return to the sections you are weakest on.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Active Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failures) → two Deep Dives |
| Deep Dive | 3+ hrs | Everything, including the Principal Lens and appendices |
What is Proximity Matching? — Why interviewers pick this topic
Proximity matching is the class of product where users are shown other users who are physically nearby, express one-sided interest (swipe right / like), and are connected only when interest is mutual. Tinder, Bumble, Hinge, Grindr and Happn are the canonical examples; the same shape appears in "people nearby" features, local marketplaces and gig-worker discovery.
Before vs After — the Saturday-night deck:
Naive design (query-on-open, swipe table keyed by swiper):
t=0: 21:00 local, evening peak. 3× baseline app opens in one metro.
t=+20s: Every app open runs a geo query + 12 filters + ML rank against one geo shard.
t=+60s: The Manhattan shard is at 95% CPU; Montana shard is at 4%.
t=+2min: Deck latency p99 goes 250ms → 4s. Clients retry. Load doubles.
t=+5min: Two users swipe right on each other within 40ms. Both "is there a reverse like?"
reads miss the other's in-flight write. No match is created. Nobody ever knows.
t=+1hr: A blocked user reappears in a deck because the deck cache predates the block.
Trust & Safety ticket. This one is the incident that gets escalated to a VP.
Staff design (precomputed decks, pair-keyed swipes, density-balanced geo shards):
t=0: Same peak.
t=+20s: Decks served from a per-user precomputed list (Redis), ~5ms. Geo query runs
asynchronously when the deck drops below 20 cards.
t=+60s: Geo shards are balanced by *active user count*, not by area. No hot shard.
t=+5min: Simultaneous swipes land on the same pair-keyed row. A conditional write
creates exactly one match. Both users get the "It's a Match" screen.
t=+1hr: Block events invalidate decks synchronously and a final safety filter runs
at serve time. The blocked user never appears. A metric, not a ticket.
Why interviewers reach for this question: it looks like a geo-index question ("use a geohash!") and most candidates spend 20 minutes on quadtrees. The geo index is the easy 10%. The question actually tests whether you can find the three hard problems hiding behind the product: a write path that is 10–50× the size of the match path, a mutual-detection race that silently loses matches, and a two-sided marketplace where ranking decisions shape who gets any attention at all.
Mechanics Refresher: Spatial Indexes
| Index | How It Works | Pros | Cons |
|---|---|---|---|
| Geohash | Interleave lat/lng bits into a base-32 string; prefix = containing cell | Works in any KV store or B-tree; Redis GEOADD/GEOSEARCH uses it | Cells are rectangles that distort with latitude; neighbors can share no prefix across cell edges |
| Quadtree | Recursively split a square into 4 until each leaf holds ≤ K points | Adapts to density (Manhattan splits deep, Montana stays shallow) | In-memory structure; rebalancing under movement is painful |
| Google S2 | Project the sphere onto a cube, walk a Hilbert curve; 64-bit cell IDs at 31 levels | Hilbert locality means nearby cells have nearby IDs → range scans; good shard keys | Library dependency; cells are quadrilaterals, not hexagons |
| Uber H3 | Hierarchical hexagonal grid, 16 resolutions | Every neighbor is equidistant — great for density, surge and heatmaps | Not perfectly hierarchical (children don't exactly tile parents) |
| R-tree / BKD | Bounding-box tree (PostGIS) or block KD-tree (Lucene/Elasticsearch geo_point) | Native in the store you already run; supports compound filters | Tied to that store's scaling model |
For most production systems: Elasticsearch/OpenSearch geo_point (BKD tree) inside a geo-sharded index, with S2 or H3 cells as the shard key. The index structure is almost never the interview question — shard balance, the write path, and mutual-match correctness are.
Executive Summary
If you only read one section, read this. Everything else in the case study elaborates on the contrast and the positions below.
What This Interview Actually Tests#
Tinder is not a geo-query question. Every candidate knows geohashes.
It is a two-sided marketplace with a write-heavy, correctness-sensitive edge (the swipe) and a read-heavy, freshness-tolerant edge (the deck) — and the design is about keeping those two edges from contaminating each other. It tests:
- Whether you notice that swipes outnumber deck fetches ~20:1 and outnumber matches ~50:1, so the write path — not the geo query — sizes the system
- Whether you catch the simultaneous-swipe race that silently loses matches
- Whether you treat safety filters (blocks, reports, bans) as fail-closed while treating everything else as fail-open
- Whether you see that ranking is a policy decision about who gets attention, owned by someone who is not the infrastructure team
The key insight: The deck can be minutes stale and nobody notices. A lost match or a blocked user reappearing is noticed by exactly the person you most need to keep. Staff engineers spend their consistency budget on the swipe → match edge and the safety edge, and spend nothing on the deck.
The L5 vs L6 vs L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Draws a geohash grid and a "nearby users" query | Asks "What's the ratio of swipes to deck fetches to matches?" and sizes the write path first | Asks "Which of these capabilities is shared with the other 4 products in the portfolio — geo, safety, ranking — and which is Tinder-specific?" |
| Deck | Queries the geo index on every app open | Precomputes a deck of ~100–200 candidates per user, refills asynchronously at a low-water mark | Treats candidate generation as a platform shared by dating, friends and events products; prices the recompute fleet |
| Match detection | "When A likes B, check if B liked A" | Names the race; uses a pair-keyed conditional write so exactly one match is created | Makes "exactly one match per pair" an audited invariant with a reconciliation job and an SLO the business signs |
| Safety | "Filter blocked users in the query" | Block/report is fail-closed: synchronous deck invalidation + serve-time filter | Owns a cross-product Trust & Safety contract: every surface that renders a person MUST call the safety filter; enforced in CI and audit |
| Geo sharding | Shards by geohash prefix | Shards by active-user density (S2 cells merged until each shard holds similar load) | Plans the 3-year path: shard map as a versioned, rebalanceable artifact; knows the reshard is the one-way door |
| Ranking | "Sort by distance, then by ML score" | Names the attention-inequality problem and caps impressions per profile | Frames ranking as a governance question — who approves objective changes, how fairness is measured, what regulators may ask |
Why "first move" separates levels
L5: Starts with the visible feature — "show nearby people" — and designs the geo query. It is a competent answer to the wrong bottleneck. At 1.6B swipes/day (the order of magnitude Tinder has publicly cited) the write path is ~18K writes/s average and ~55K/s at evening peak, while deck fetches are an order of magnitude smaller and cacheable.
L6: Sizes the three flows before drawing anything: "Let me get the ratios first. If a session is ~20 swipes on one deck fetch and ~2% of swipes produce a match, then swipe writes dominate, deck reads are cacheable, and matches are rare but must be exact. That tells me where to spend consistency."
L7: Asks which pieces are shared infrastructure. A dating company typically runs several apps; the geo index, the safety graph and the notification pipeline are org assets. The L7 move is deciding what is a platform (safety, geo) versus what is product-owned (ranking objective, deck UX).
Why "match detection" separates levels
L5: "When A swipes right on B, look up whether B already swiped right on A. If yes, create a match." This is correct in a single-threaded world. In a distributed store with two independent writes, both reads can miss.
L6: "Two people swiping on each other within the same ~50ms window is not rare at 55K swipes/s. If each side writes its own like and then reads the other's, both reads can return empty under eventual consistency. I'll put both directions of a pair in the same row — keyed by min(a,b):max(a,b) — and use a conditional write, so the second writer sees the first and exactly one match is created."
L7: Adds the audit loop: a daily reconciliation job that scans pair rows with both directions liked but no match, pages if the count is above zero, and reports the lost-match rate as a business metric. "A lost match is revenue and trust we'll never see a ticket for. I want it counted."
Why "safety" separates levels
L5: Adds a NOT IN blocked_ids clause to the candidate query. Correct at query time, but precomputed decks and caches were generated before the block.
L6: Separates the consistency tiers: decks are eventually consistent, safety is not. On block/report: synchronously purge the blocked ID from both users' decks, write to a safety set, and run a final serve-time filter against that set. "A stale deck costs us a slightly worse card. A stale safety filter costs a harassment report and possibly a news story."
L7: Recognizes that every surface that shows a person — deck, "likes you" grid, chat list, push notifications, shared-contact suggestions — must honor the same safety state, and that this is enforced by contract and audit, not by each team remembering.
The Staff Positions#
| Position | Rationale |
|---|---|
| Precompute decks, don't query on open | Deck reads are ~20× more frequent than geo changes; a 100–200 card deck refilled at a low-water mark turns a 150–300ms query into a ~5ms list pop |
| Pair-keyed swipe storage for match detection | Putting A→B and B→A in the same row turns a distributed race into a single-row conditional write; exactly one match |
| Safety is fail-closed; everything else is fail-open | If the ranking service is down, serve a distance-sorted deck. If the safety filter is down, serve nothing |
| Shard geo by load, not by area | Active users per km² varies by >1,000× between Manhattan and rural areas; equal-area shards guarantee a hot shard |
| Coarsen location before it leaves the server | Precise distances enable trilateration; round displayed distance to ~1 mile and snap stored location to a grid cell |
| Cap impressions per profile | Without a cap, the top ~10% of profiles absorb most of the likes and most of the deck slots; the marketplace starves |
| Swipes go through a log, not a synchronous fan-out | Swipe → Kafka → consumers (match, ranking features, analytics); only the match check is on the synchronous path |
The Three Intents#
Three intents hide inside "design Tinder." Each leads to a different system.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Discovery deck (browse ranked nearby profiles) | Latency + relevance; minutes of staleness fine | Precomputed candidate list per user; geo filter → exclusion → rank | Stale/irrelevant cards; hot geo shard | Eventual; ~5–15 min staleness acceptable |
| Mutual matching (swipe → match → chat) | Exactly one match per mutual pair; low latency on the "It's a Match" moment | Pair-keyed conditional write; event log for downstream | Lost match (silent) or duplicate match (visible) | Exactly-once per pair; audited |
| Live proximity ("who is near me right now", crossed paths) | Location freshness in seconds; privacy | Streaming location updates into cell-keyed in-memory index; TTL on presence | Location leakage, battery drain, stalking risk | Freshness ≤ 30–60s; privacy fail-closed |
🎯 Staff Move: "I'll design the discovery deck plus mutual matching, because that is Tinder's core loop and it forces the interesting split: a freshness-tolerant read path and an exactness-required write path. Live proximity is a different product with a different privacy posture — I'd scope it out explicitly rather than let it leak location precision into the deck design."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Precompute vs Query-on-Open | Precomputed decks are fast and cheap per read but stale and costly to maintain for inactive users; live queries are fresh but put the geo index on the hot path |
| 2 | Swipe Keying: Swiper vs Pair | Keying by swiper makes "my history" cheap; keying by pair makes match detection exact. You usually need both, and one is the source of truth |
| 3 | Location Precision vs Privacy | Precise location improves relevance and "distance away"; it also enables trilateration and stalking |
| 4 | Engagement Ranking vs Marketplace Health | Ranking by predicted swipe-right maximizes short-term engagement and concentrates attention on few profiles |
| 5 | Area-Based vs Load-Based Geo Sharding | Area shards are simple and stable; density varies >1,000× so they produce hot shards. Load-balanced shards need a shard map and resharding |
In the Wild: Real Production Systems#
Why this section belongs here: citing publicly documented systems shows you've studied operational reality, not a whiteboard idealization.
Tinder — Geosharded Recommendations on Elasticsearch#
Tinder's engineering blog has described moving its recommendation candidate search to a geosharded Elasticsearch layout: the world is partitioned using Google's S2 cells, and cells are grouped into shards so that each shard carries a comparable load rather than a comparable area. A user's query hits only the shard(s) covering their search radius instead of scattering to every shard. Tinder has also publicly said that its early "Elo score" desirability ranking is no longer how it ranks.
Staff insight: The lesson isn't "use Elasticsearch." It's that the shard key is geographic but the shard boundary is chosen by load. Say that sentence in the interview and you've skipped ten minutes of geohash discussion.
Hinge — Stable-Matching-Inspired Recommendations#
Hinge has publicly said its "Most Compatible" feature draws on the Gale–Shapley stable matching algorithm: instead of showing everyone the most-liked profiles, it considers both sides' likely preferences to suggest pairs likely to be mutual. This is an explicit two-sided optimization rather than one-sided engagement ranking.
Staff insight: Dating is a two-sided market. Ranking only by "will A swipe right on B" ignores "will B swipe right on A," which is where the mutual-match rate — the actual product outcome — comes from. Naming this is an L6 ranking signal.
Uber H3 and Redis GEO — The Index Is a Commodity#
Uber open-sourced H3, a hexagonal hierarchical grid, for density and pricing; Redis ships GEOADD/GEOSEARCH, which store points as 52-bit geohash scores in a sorted set. Both are off-the-shelf.
Staff insight: When the spatial index is a library call, designing one from scratch in an interview signals misallocated time. Say "S2 or H3 for the cell ID, a sorted set or BKD tree for the lookup — now let's talk about the swipe path."
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Geohash the users and query neighbors" | "Manhattan has 1,000× the density of Wyoming. What does your shard look like at 9pm Saturday?" | Do you see load skew, not just correctness? |
| "When A likes B, check if B liked A" | "What if they swipe at the same moment?" | Do you know the race, and do you fix it structurally? |
| "Precompute decks" | "User moves from SF to NYC on a flight. What do they see when they land?" | Invalidation triggers, not just caching |
| "Filter blocked users" | "The deck was built an hour ago. The block happened a minute ago." | Consistency tiers — is safety fail-closed? |
| "Rank by ML score" | "The top 5% of profiles get 60% of the likes. Is that your problem?" | Marketplace thinking and ownership |
| "Store swipes in Cassandra" | "How big is that table in a year? What do you ever read from it?" | Retention, compaction, read patterns |
System Architecture Overview#
Reading the diagram: Two synchronous paths and one asynchronous brain. The deck path pops cards from a precomputed Redis list and passes every card through the fail-closed safety filter. The swipe path does exactly one correctness-critical thing — the pair-keyed conditional write — and hands everything else to Kafka. The candidate generator is the only component that touches the geo index, and it runs off the critical path when a user's deck drops below its low-water mark.
match.lost_pairs_totalis the metric that catches the silent failure.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Geo index | "Geohash + neighbor cells" | "S2/H3 cell as the shard key, BKD or sorted set for lookup. Shards are balanced by active users, not area." |
| Deck | "Query nearby users on open" | "Precompute ~150 candidates; refill at 20 remaining. The geo index is never on the tap-to-card path." |
| Match | "Check if they liked me back" | "Pair-keyed row, conditional write. Exactly one match. Reconciliation job counts lost pairs." |
| Seen filter | "Query the swipe table" | "Per-user Bloom filter of swiped IDs — ~12 KB for 10K swipes at 1% FPR — plus the pair store as authority." |
| Safety | "Exclude blocked users in the query" | "Fail-closed serve-time filter plus synchronous deck purge. Different consistency tier from the deck." |
| Privacy | "Show distance" | "Snap to a grid, round displayed distance to 1 mile, never return coordinates. Trilateration is a known attack." |
| Ranking | "ML score" | "Two-sided: P(A likes B) × P(B likes A), with impression caps. Ranking objective changes need a named owner." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Swipes per day (Tinder-scale, publicly cited order) | ~1.6B | ≈ 18.5K/s average, ~55K/s evening peak (3×) — the dominant write load |
| Swipes per deck fetch | ~20–50 | Why deck reads are cacheable and swipes are the sizing input |
| Swipe-right → match rate | ~1–3% of right swipes | Matches are rare, so match creation can afford a strongly consistent write |
| Swipe record size | ~40–60 bytes | 1.6B × 50 B ≈ 80 GB/day raw, ~240 GB/day at RF 3; ~90 TB/yr before TTL |
| Deck size precomputed | 100–200 candidates | ~1 KB of IDs per user; 20M DAU × 1.5 KB ≈ 30 GB of Redis |
| Deck low-water mark | ~20 cards | Refill before the user notices a spinner |
| Bloom filter for 10K seen IDs @ 1% FPR | ~12 KB (9.6 bits/element) | "Don't show me people I've already swiped on" without a DB read |
| Geohash precision 5 / 6 | ~4.9 × 4.9 km / ~1.2 × 0.6 km | Precision 5 is a reasonable "search cell" for a 10–25 km radius |
| S2 level 12 / 13 cell | ~3.3 km² / ~0.8 km² (varies ~2× by location) | Typical grouping level for geo shards |
| Density skew metro vs rural | >1,000× active users per km² | Why area-based sharding creates hot shards |
| Displayed distance granularity | 1 mile / 1 km | Anti-trilateration floor |
| Deck p99 target | < 50 ms (precomputed) vs 150–300 ms (live query) | The precompute payoff |
| "It's a Match" notification | p99 < 2 s to both parties | The emotional peak of the product |
Interview Walkthrough
The most common mistake: candidates spend 20 minutes on geohash math and quadtree splitting, then run out of time before the swipe race, the safety tier, or the marketplace question. The phases below compress the geo index to about two minutes so the remaining 30+ go to the parts that decide your level.
Phase 1: Requirements & Framing (2–3 min)#
State the functional core in one breath:
"Users set a location, age range, gender preference and max distance. They see a deck of nearby profiles, swipe left or right, and when two people both swipe right, they match and can chat."
Then spend the remaining time on the non-functionals — that is where the design lives:
"Before I draw anything, let me get the ratios. I'll assume ~20M daily actives, ~80 swipes per active per day — that's ~1.6B swipes/day, roughly 18K/s average and ~55K/s at the evening peak. If a deck fetch returns ~50 cards, deck fetches are ~20× fewer than swipes. And if ~2% of right swipes are mutual, matches are ~1–2K/s at peak. So: swipes size the write path, decks are cacheable, and matches are rare but must be exact."
Commit to an intent and a consistency split:
"I'll design the discovery deck plus mutual matching. The deck can be minutes stale. The match must be exactly-once per pair. Safety — blocks and reports — is fail-closed and strongly consistent. Those three consistency tiers drive everything."
🎯 Staff Move: Say the ratios out loud before drawing. The interviewer learns in 30 seconds that you size from the workload, not from the feature list — and you have just justified why you won't spend time on the geo index.
Phase 2: Core Entities & API (1–2 min)#
Name the nouns and move on:
- Profile:
user_id, prefs (age range, genders, max_distance_km),geo_cell(S2 level 13),last_active_at - Deck: ordered list of
candidate_idper user,generated_at,generated_cell - PairState: key
min(a,b):max(a,b),a_to_b ∈ {none, like, pass},b_to_a,matched_at,version - Match:
match_id = hash(pair_key), both user IDs,created_at,chat_id - SafetyEdge:
(blocker, blocked),reason,created_at— never expires by default
GET /v1/deck?limit=20 → [{candidate_id, display_distance_mi, photos[], ...}]
POST /v1/swipes {target_id, direction, client_swipe_id}
→ {matched: bool, match_id?}
POST /v1/location {lat, lng, accuracy_m} → 204 (server snaps to cell; never echoes coordinates)
POST /v1/blocks {target_id, reason} → 204 (synchronous, fail-closed)
GET /v1/matches?cursor=... → [{match_id, user, chat_id, created_at}]
🎯 Staff Move: "
client_swipe_idmakes the swipe idempotent — mobile networks retry, and I don't want a retry to double-count a like for ranking features or send a second match notification. And the location endpoint never returns coordinates — only a rounded display distance ever leaves the server."
Phase 3: High-Level Architecture (≤ 5 min)#
Staff candidates spend under five minutes here. Draw two paths and a brain:
Walk it in three sentences:
"The deck path pops from a precomputed list and runs a fail-closed safety check per card. The swipe path does one conditional write on a pair-keyed row, returns matched: true if this write completed the pair, and emits an event. Everything expensive — geo query, ranking, notifications — is asynchronous and driven by the event log."
Phase 4: Transition to Depth#
The sentence that steers the interviewer to where you are strongest:
"The geo index is a solved problem — S2 cells, sharded by load. I think the three interesting parts are the simultaneous-swipe race, keeping safety fail-closed while decks are precomputed, and hot geo shards at peak. Which would you like first? If you don't mind, I'll start with the swipe race because it's the one that silently loses data."
🎯 Staff Move: Offering a menu with a recommendation is the Staff pattern. It shows you know where the bodies are buried and still lets the interviewer steer.
Phase 5: Deep Dives (25–30 min)#
Budget your deep-dive time across these, in the order the interviewer allows:
| Deep Dive | Time | What You Must Land |
|---|---|---|
| Match race | 6–8 min | Pair-keyed row, conditional write (LWT / DynamoDB condition / single-partition transaction), idempotent match creation, reconciliation job |
| Deck freshness & safety | 5–7 min | Low-water refill, invalidation triggers (move > 25 km, pref change, block), serve-time safety filter, fail-closed |
| Geo sharding at peak | 5–7 min | Load-balanced S2 shard map, radius → shard fan-out ≤ 3, boundary queries, rebalancing |
| Ranking & marketplace | 4–6 min | Two-sided scoring, impression caps, who owns the objective |
| Privacy | 2–3 min | Grid snapping, rounded distance, trilateration, no coordinates in API |
Match race, as you'd say it:
"If A→B and B→A are stored in separate partitions keyed by swiper, each swipe does write-own then read-other, and under eventual consistency both reads can miss. Instead, both directions live in one row keyed min:max. A's swipe is UPDATE pair SET a_to_b='like' IF version=v. If the row now shows both likes and matched_at is null, the same conditional write sets matched_at. Exactly one writer wins that condition. The loser retries, reads the row, sees the match already exists, and returns matched: true too — both users get the screen, one match row exists."
Phase 6: Wrap-Up (2–3 min)#
Close with constraints, evolution, and what you deliberately did not build:
"To summarize: swipes size the system and go through a pair-keyed conditional write; decks are precomputed and eventually consistent; safety is a separate fail-closed tier. What I'd monitor first is match.lost_pairs_total from the reconciliation job and geo shard load skew. What I'd build next is two-sided ranking with impression caps, because that's the lever on mutual-match rate. What I skipped: live proximity, which has a different privacy posture and I'd want a separate review for it."
Common Timing Mistakes#
| Mistake | Time Lost | Fix |
|---|---|---|
| Deriving geohash bit interleaving on the whiteboard | 8–12 min | "S2 cell IDs; nearby cells have nearby IDs. Moving on." |
| Designing the chat system in depth | 10+ min | "Match creates a chat thread; chat is its own design — happy to go there if you want." |
| Listing 15 functional requirements (super likes, boosts, passport) | 5 min | Name the core loop; say monetization features reuse it |
| Drawing a microservice per noun | 5 min | Two paths and a brain. Seven boxes max in Phase 3 |
| Never saying a number | Whole interview | Ratios in Phase 1; latency targets on every box |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Proximity matching is a disguise test. The prompt names a geographic feature, and the geographic part is the easiest thing in the system. Interviewers pick it because it lets them watch whether you follow the prompt's framing or the workload's shape.
The workload has three edges with incompatible requirements:
| Edge | Volume (Tinder-scale) | Consistency Need | Cost of Getting It Wrong | Who Notices |
|---|---|---|---|---|
| Deck reads | ~1–3K/s avg, cacheable | Eventual (minutes) | A slightly worse card | Nobody, individually |
| Swipe writes | ~18K/s avg, ~55K/s peak | Per-pair linearizable | Lost or duplicate match | The two users — silently or loudly |
| Safety writes | ~10s/s | Strong, fail-closed | Harassment, press, regulator | Trust & Safety, Legal, the VP |
A Senior engineer who designs all three at the same consistency level either over-pays on decks (strongly consistent geo queries on every open) or under-pays on safety (eventually consistent block lists). Staff engineers make the split explicit and name who pays for each.
1.2 The L5 vs L6 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"Which mistake would a user notice, and which would they never know about?"
Every decision in this system falls out of that question:
- A card that is 10 minutes stale — never noticed. Cache it.
- A profile shown twice — noticed, mildly annoying. Bloom filter; accept ~1% false positives (which hide a card, never re-show one).
- A lost match — never noticed by the users, but it is the product's core outcome silently failing. Make it structurally impossible and measure it anyway.
- A blocked user reappearing — noticed immediately, by the person most harmed. Fail-closed, synchronous, audited.
🎯 Staff Move: "The failure I worry about most is the one no user will ever report: two people who liked each other and never matched. I'm going to design that out structurally and then count it with a reconciliation job, because it will never show up in support tickets."
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Intent 1 — Discovery deck. The user wants a stream of plausible, nearby, not-yet-seen people. The system's job is candidate generation (geo + preferences + exclusions) and ranking. Latency matters at the moment of the swipe animation (the next card must already be on the device); freshness matters on the order of minutes. The deck is a recommendation feed with a geographic filter — structurally closer to News Feed than to Maps & Geospatial.
Intent 2 — Mutual matching. The user expresses interest; the system must detect mutual interest exactly once and trigger the match moment (notification, chat creation). This is a contention problem on a two-party key — the same family as seat reservation in Flash Sales & Ticketing, where two writers race for one logical resource. See Dealing with Contention.
Intent 3 — Live proximity. "Who is within 500 m right now" or "whom did I cross paths with today" (Happn's model). Requires continuous location updates, second-level freshness, and a much stricter privacy posture because it reveals current location. This is closer to the driver-location problem in Ride Hailing & Delivery.
| Dimension | Discovery Deck | Mutual Matching | Live Proximity |
|---|---|---|---|
| Location update frequency | On open, or on > 1 km move | N/A | Every 30–60 s while active |
| Location precision stored | Cell (~1 km) | N/A | ~100 m, short TTL |
| Primary store | Redis lists + geo index | Pair-keyed KV with conditional writes | In-memory cell index with TTL |
| Hardest problem | Hot shards, freshness vs cost | Simultaneous-swipe race | Privacy and battery |
| Who owns the risk | Recommendations team | Core platform / matching team | Privacy + Trust & Safety |
2.2 When NOT to Use Proximity Matching#
| Situation | Why Proximity Matching Is Wrong | Use Instead |
|---|---|---|
| Matching is one-sided (marketplace search: "restaurants near me") | No mutual consent step; no pair race | Geo search + ranking (Search Indexing) |
| Supply is assigned, not chosen (drivers to riders) | The platform decides the pair; nobody swipes | Dispatch/assignment (Ride Hailing & Delivery) |
| Distance doesn't matter (professional networking, remote jobs) | Geo filter adds cost and shard complexity for no relevance gain | Graph- or embedding-based recommendation |
| Population is small (< ~100K users in a region) | A single Postgres + PostGIS instance answers every query in < 20 ms | One database, no geo-sharding |
| Mutual consent is legally required to be auditable (e.g., regulated introductions) | Swipe-based implicit consent may be insufficient | Explicit request/accept workflow with an audit log |
🎯 Staff Move: "At under a few million users per metro I would not geo-shard at all — PostGIS with a GiST index on one primary handles this. Geo-sharding is a scaling tool, and it brings a shard map that someone has to own. I'd adopt it when a single region's index stops fitting in memory or its write rate exceeds one primary."
2.3 What the Interviewer Leaves Underspecified#
| Underspecified | Why It Matters | What to Say |
|---|---|---|
| Scale (users, swipes/day) | Determines whether you shard at all | "I'll assume ~20M DAU and ~1.6B swipes/day; I'll call out what changes at 10× smaller." |
| Deck freshness tolerance | Determines precompute vs live | "I'll assume minutes of staleness is fine for ranking, but not for safety." |
| What "nearby" means | Radius drives shard fan-out | "User-set max distance 2–160 km; default ~50 km." |
| Location update policy | Battery, privacy, write load | "On app open and on significant movement — not continuous." |
| Global vs regional | Data residency, cross-border matching | "Users match within a region; travel mode ('Passport') is a remote-location query, not a cross-region write." |
| What a "match" triggers | Chat, notification, analytics | "Match creates a chat thread and notifies both parties; I'll treat chat as its own system." |
| Safety requirements | Consistency tier | "Blocks must take effect immediately on every surface. I'll assume that's a hard requirement." |
2.4 Precise Terminology#
| Term | Meaning Here | Common Confusion |
|---|---|---|
| Deck | Ordered, precomputed list of candidate IDs for one user | Not a query result; it's a materialized feed |
| Candidate generation | Geo + preference + exclusion filtering that produces the pool | Distinct from ranking, which orders the pool |
| Pair key | min(user_a, user_b) + ":" + max(user_a, user_b) | Not the swiper's key; symmetric by construction |
| Mutual | Both directions of a pair are like | "Match" is the record created when mutual is detected |
| Seen set | IDs the user has already swiped on | Must never be re-shown; Bloom filter is an optimization, pair store is authority |
| Safety edge | A block, report or ban that removes a person from another's surfaces | Not a preference filter — fail-closed, never cached without invalidation |
| Geo cell | S2/H3 cell ID at a fixed level representing a user's snapped location | Not the user's coordinates; coordinates should not be stored long-term |
| Shard map | Versioned mapping from cell ranges → index shards | An artifact with an owner, not a hash function |
| Low-water mark | Deck length at which refill is triggered (~20) | Too low → spinner; too high → wasted recompute |
| Impression cap | Maximum times a profile is placed in decks per window | Marketplace-health control, not a rate limit on the user |
3. The Fault Lines#
Five structural tensions. For each: the options, who pays, the Staff default, and when to deviate.
3.1 Fault Line 1: Precompute vs Query-on-Open#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Query-on-open (geo + filters + rank on every deck request) | Always fresh; no deck storage; simple invalidation | Geo index on the hot path; 150–300 ms p99; peak-hour hot shards; ML ranking cost per open | Users (latency at peak), infra (index sized for peak opens) |
| Precompute for all users nightly | Cheapest serve path; predictable batch cost | Wasted work for the ~60–70% of MAU who don't open today; stale for travelers | Finance (batch compute for inactive users), travelers (wrong-city decks) |
| Precompute on demand at a low-water mark (Staff default) | Only active users pay; refill is async; serve path is a list pop | Invalidation triggers must be explicit (move, pref change, block); first open after long absence pays a cold start | Recs team owns the trigger list; cold-start users see ~300 ms once |
| Hybrid: precomputed pool + live re-rank at serve | Fresh ordering with cached candidates | Serve path depends on the ranking service | Ranking team gets paged for deck latency |
The Staff default: on-demand precompute. The deck holds ~150 IDs; when it drops below ~20, the Deck Service emits a deck.refill event; the Candidate Generator runs geo → preference → exclusion → rank and writes a new list. Invalidation triggers, owned and documented:
| Trigger | Action | Latency Budget |
|---|---|---|
| Location moves > 25 km (or to a new region) | Discard deck, regenerate | Next open |
| Preference change (age, distance, gender) | Discard deck, regenerate | Next open |
| Block / report / ban involving the user | Purge the specific ID from both decks synchronously | Before the POST /blocks returns |
| Candidate deletes account / pauses | Lazy: dropped at serve time by safety + existence check | Serve time |
| Deck age > 6 h | Regenerate on next open | Next open |
When to deviate: at < 1M DAU, query-on-open against a single PostGIS or Elasticsearch cluster is simpler and cheap enough. Precompute buys latency and peak-load isolation; it costs a trigger list someone must own.
🎯 Staff Move: "Precompute is only safe if the invalidation triggers are a documented contract. The dangerous one is block — that can't wait for a refill. So block does a synchronous purge plus a serve-time filter; everything else rides the low-water mark."
3.2 Fault Line 2: Swipe Keying — Swiper vs Pair#
This is the fault line that separates levels most sharply.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Keyed by swiper, write-then-read-reverse | "My swipes" is one partition; cheap history | Race loses matches under eventual consistency; with quorum reads it creates duplicate matches instead | The two users (silent loss) or client team (dedupe duplicate matches) |
| Keyed by swiper, strongly consistent reads (QUORUM/QUORUM) | At least one side sees the other → no loss | Both may see each other → double match creation; 2× latency on every swipe | Latency budget; match service must be idempotent anyway |
Keyed by pair min:max, conditional write (Staff default) | One row, one condition; exactly one writer completes the pair | "All of Alice's swipes" requires a secondary index or a second table | Storage (dual-write to a swiper-keyed history via the log) |
| Single-writer per pair via routing (hash pair key → owner node) | No store-level conditional needed | Requires sticky routing and failover handling | Platform team owns the router |
The Staff default: pair-keyed conditional write as the source of truth for matching; a swiper-keyed history table populated asynchronously from Kafka for "my swipes," undo, and analytics.
# Swipe handler (pair store supports compare-and-set on a version)
pair = key(min(a,b), max(a,b))
loop up to 3 times:
row = read(pair) # consistent read on the owning partition
if row.dir(swiper) == direction: return idempotent_ok(row) # retry of same swipe
new = row.with(dir(swiper) = direction)
if new.a_to_b == LIKE and new.b_to_a == LIKE and row.matched_at is null:
new.matched_at = now(); new.match_id = hash(pair)
if cas(pair, expected=row.version, new):
emit(swipe_event); if new.matched_at and not row.matched_at: emit(match_event)
return {matched: new.matched_at != null, match_id: new.match_id}
# fall through → 503, client retries with same client_swipe_id
The contention on a single pair row is effectively two writers, so CAS conflict rates are tiny — this is the regime where optimistic concurrency shines (well under ~5% conflict). Cassandra LWT costs ~4 round trips (Paxos) — acceptable only because it is per-swipe on a small row; DynamoDB conditional writes cost one round trip and are the cheaper fit.
When to deviate: if your store has no conditional writes and you can't route by pair, use quorum reads and make match creation idempotent on match_id = hash(pair). You trade silent loss for harmless duplicates that collapse on the idempotency key.
🎯 Staff Move: "Keying by swiper feels natural because that's how the UI thinks. Keying by pair is how the invariant thinks. I key the source of truth by the invariant and derive the UI view from the log."
3.3 Fault Line 3: Location Precision vs Privacy#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Store and return precise coordinates | Maximum relevance; exact distance | Trilateration: three spoofed vantage points reveal a user's location to ~tens of meters | Users (stalking risk), Legal, the brand |
| Store precise, display rounded distance | Good relevance; UI looks safe | Rounding alone is attackable if the server computes on precise data and the attacker can move freely | Same, with a longer attack |
| Snap to a grid cell server-side, compute distance from cell centers, round display to 1 mile (Staff default) | Trilateration yields only the cell; relevance loss negligible at 2–160 km radii | Users near cell edges get slightly odd distances | Product accepts ~0.5–1 km relevance fuzz |
| Coarse location only (city-level) | Strongest privacy | "Nearby" becomes meaningless in dense metros | Product relevance |
Precise-distance leakage in location-based dating apps is a well-publicized class of vulnerability — security researchers have repeatedly demonstrated trilateration against apps that returned precise or finely rounded distances. The fix is structural: the server should not know more precision than it needs, and the API should never compute on more precision than it reveals.
The Staff default: client sends lat/lng; the Location Service snaps to an S2 level-13 cell (~0.8 km²) and stores only the cell; displayed distances are computed between cell centers and rounded to whole miles/km with a floor of "< 1 mile"; raw coordinates are discarded within seconds and never logged.
When to deviate: live-proximity features need ~100 m. Scope them as a separate, opt-in product with its own privacy review, short retention (minutes), and no distance display.
🎯 Staff Move: "Privacy here isn't a policy page, it's a data-minimization decision in the Location Service. If we never store more than a 1 km cell, no bug, insider, or subpoena can reveal more than that."
3.4 Fault Line 4: Engagement Ranking vs Marketplace Health#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Rank by distance only | Transparent, cheap, fair-ish | Low relevance; low mutual rate | Users (swipe fatigue), revenue |
| Rank by P(viewer likes candidate) | Maximizes right-swipes | Attention concentrates on a small set of profiles; most users get few likes and churn | The majority of users (invisible), long-term retention |
| Desirability score (Elo-like) | Simple, stable | Rich-get-richer feedback loop; Tinder has publicly said it moved away from Elo | Same as above |
| Two-sided: P(A likes B) × P(B likes A), with impression caps (Staff default) | Optimizes for mutual likes — the product outcome; distributes attention | More model complexity; harder to explain; needs a fairness metric | Ranking team (complexity), product (short-term right-swipe rate may drop) |
The Staff default: score = P(viewer→cand) × P(cand→viewer) with a per-profile impression cap per 24 h (e.g., a profile appears in at most ~N decks where N scales with the local active population), and an exploration budget (~5–10% of deck slots for new or low-exposure profiles).
The owner question is the Staff-level signal: the ranking objective is a product decision with marketplace consequences. Infrastructure implements it; a named product owner signs off on objective changes, and every change ships behind an A/B test with mutual-match rate and like-distribution Gini as guardrail metrics.
When to deviate: early-stage apps with small populations should rank by distance and recency — there isn't enough data to train two-sided models, and exploration is the whole deck.
🎯 Staff Move: "Optimizing for right-swipes and optimizing for matches are different objectives. I'd rank for mutual probability and put a guardrail metric on how concentrated likes become. Whoever owns the ranking objective owns that guardrail."
3.5 Fault Line 5: Area-Based vs Load-Based Geo Sharding#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Hash by user ID | Perfect balance | Every geo query scatters to every shard; fan-out × N | Infra (N× query cost), users (tail latency = slowest shard) |
| Fixed geohash prefix per shard | Simple, no map | 1,000×+ density skew → hot shards; Saturday 9pm in big metros | On-call for the hot shard |
| S2 cell ranges merged by load (Staff default) | Balanced; Hilbert locality keeps queries to 1–3 shards | Shard map must be versioned and rebalanced; boundary queries hit 2+ shards | Platform team owns the shard map and rebalance tooling |
| Per-metro clusters | Operationally obvious; aligns with data residency | Metros grow unevenly; boundary cities | Same, coarser |
The Staff default: compute the S2 level-13 cell for each user; walk cells in Hilbert order and cut shard boundaries every ~K active users (weighted by query load, not just population). Store the map as a versioned config. A query computes the cells covering its radius (an S2 RegionCoverer call), maps them to shards — typically 1–3 — and queries those in parallel.
Rebalancing: when a shard's p99 or doc count exceeds ~1.5× the median for 7 days, split it; reindex the moving range into the new shard, dual-read during migration, flip the map version. The map version must be in every query log so you can debug "why did this user see nobody?"
When to deviate: under a few million users, one cluster with a geo_point field and no geo-sharding. Geo-sharding is justified when fan-out cost or shard size forces it.
🎯 Staff Move: "The shard key is geography; the shard boundary is load. And the shard map is a production artifact — versioned, owned, and in every query log — not a constant in the code."
4. Failure Modes & Operational Reality#
4.1 The Silent Lost Match#
Scenario: the pair store is eventually consistent and match detection does write-then-read-reverse keyed by swiper.
t=0 Evening peak, 55K swipes/s. Store replica lag p99 = 80 ms (normally 5 ms)
because a compaction storm is running on 2 of 12 nodes.
t=+1min Pairs that swipe on each other within the lag window: both reverse reads miss.
t=+10min ~0.4% of mutual pairs lost. No errors. No alerts. Match rate dashboard
dips 0.4% — inside normal daily variance.
t=+3 days A data scientist notices mutual-like-without-match rows in the warehouse.
t=+3 days Nobody can notify the lost pairs retroactively without it being creepy.
- Detection:
match.lost_pairs_totalfrom a reconciliation job that scans pair rows (or joins swipe history) for mutual-like withoutmatched_at. Alert: > 0 for 15 min → page. - Blast radius: a fraction of new matches during the lag window; invisible to users.
- Mitigation: switch to pair-keyed conditional writes; reconciliation job creates missed matches within minutes (acceptable UX: "You have a new match").
- Prevention: make the invariant structural (Section 3.2), then keep the reconciliation job as a canary.
- Owner: Matching platform team.
4.2 The Hot Geo Shard#
t=0 New Year's Eve, 23:30 local, one metro. App opens 6× baseline.
t=+30s Deck refills cluster: every user who swiped through their deck triggers a
candidate generation against the same 2 shards.
t=+2min Shard CPU 98%; candidate generation p99 300 ms → 6 s. Refill queue grows.
t=+4min Decks run dry. Clients show "There's no one new around you."
t=+6min Users widen distance to 160 km → queries now cover 3–4 shards each. Worse.
- Detection:
geo.shard_cpu_skew(max/median > 3),recs.refill_queue_depth,deck.empty_responses_total. - Blast radius: one metro, all users in it — and the metros sharing its shards.
- Mitigation: serve stale decks (skip
generated_atexpiry during overload); cap refill concurrency per shard; extend deck size to 300 on refill to cut refill frequency; drop to distance-only ranking to shed ML cost. - Prevention: load-balanced shard map; scheduled pre-scaling for known events (NYE, festivals); replicas of hot shards for read fan-out.
- Owner: Recommendations (refill policy) + Search platform (shard map, replicas).
4.3 Blocked User Reappears#
t=0 Alice blocks Bob at 20:00.
t=+0s Block written to safety store. Bob removed from Alice's *current* deck.
t=+10min Alice's deck is regenerated by a candidate generator replica that read a
cached copy of the safety set loaded at 19:50.
t=+12min Bob appears in Alice's deck. Alice reports the app for harassment.
- Detection:
safety.serve_filter_hits_total(serve-time filter catching an ID that the generator should have excluded — should be ~0; non-zero means an upstream cache is stale), plus user reports tagged "blocked user reappeared." - Blast radius: one user — but the most harmful one-user incident the product has.
- Mitigation: serve-time filter against the authoritative safety store on every card (batch lookup of ~20 IDs, ~2 ms); fail-closed: if the safety store is unreachable, return an empty deck with a retry hint.
- Prevention: safety set is never cached by candidate generators beyond a version check; every surface that renders a person calls the same filter (deck, likes-you, matches, notifications).
- Owner: Trust & Safety platform.
4.4 The Celebrity Swipee (Hot Key)#
A creator with a large following joins; they receive ~500K right swipes in an hour.
- What breaks: anything keyed by swipee — a "likes you" counter, a reverse index for "who liked me" — becomes a hot partition (~140 writes/s on one key sustained, bursts of thousands). Pair-keyed storage is unaffected because each pair is a separate row.
- Detection:
store.partition_write_rate_max,likes_you.counter_contention. - Mitigation: sharded counters (
likes_you:{user}:{0..15}), summed on read; cap the displayed count ("99+"); the "likes you" grid reads from a sampled list, not the full set. - Prevention: impression caps mean a celebrity profile is shown to a bounded number of people per window — fixing the marketplace problem also fixes the hot key.
- Owner: Matching platform (storage), Recommendations (caps).
4.5 The Deck Refill Thundering Herd#
After a Redis failover, all deck lists are gone. Every active user's next request triggers a refill.
t=0 Redis primary for deck lists fails; replica promoted but with a
replication gap of 90 s (async replication).
t=+5s ~15% of decks empty. 300K refills enqueued in 30 s.
t=+1min Candidate generation at 10× normal; geo index saturates.
- Detection:
deck.cache_miss_rate> 10%,recs.refill_ratespike. - Mitigation: singleflight per user; refill admission control (token bucket per shard); serve a cheap fallback deck (distance-sorted, no ML, capped at 20 cards) while the full refill queues.
- Prevention: the fallback deck path is tested monthly in a game day; deck store is treated as a cache — losing it degrades quality, never availability.
- Owner: Recommendations.
4.6 Location Spoofing and Bot Swarms#
Bots spoof GPS into dense metros and swipe right on everyone to harvest matches for scams.
- Detection:
swipe.right_ratioper account > 95% over > 200 swipes; impossible travel (> 900 km/h between location updates); device attestation failures. - Mitigation: rate-limit right swipes per account (a product cap also serves as abuse protection — see Rate Limiting); shadow-ban: accept swipes but exclude from others' decks.
- Owner: Trust & Safety, with the rate-limit policy owned by the API platform.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Lost match (race) | match.lost_pairs_total > 0 | New mutual pairs during lag window | Pair-keyed CAS; reconciliation creates late match | Matching platform |
| Duplicate match | match.duplicates_total (same pair key, 2 rows) | Two chat threads, confused users | Idempotent match_id = hash(pair) | Matching platform |
| Hot geo shard | geo.shard_cpu_skew > 3 | One metro | Stale decks, refill caps, distance-only rank | Search platform |
| Blocked user reappears | safety.serve_filter_hits_total > 0 | One user, severe | Serve-time filter, fail-closed | Trust & Safety |
| Safety store down | safety.filter_errors | All deck serving | Serve empty deck (fail-closed) | Trust & Safety |
| Deck store loss | deck.cache_miss_rate > 10% | All active users, quality only | Fallback deck, singleflight | Recommendations |
| Celebrity hot key | store.partition_write_rate_max | "Likes you" for one profile | Sharded counters, sampling | Matching platform |
| Ranking service down | ranking.errors | Deck quality | Distance + recency ordering (fail-open) | Ranking team |
| Location spoofing / bots | swipe.right_ratio, impossible travel | Scam exposure | Shadow-ban, swipe caps | Trust & Safety |
| Notification backlog | notify.queue_lag_s > 30 | Delayed "It's a Match" | Priority lane for match pushes | Notifications platform |
🎯 Staff Move: "Look at the mitigation column: ranking fails open, the deck store fails to a cheap fallback, but safety fails closed. The system has three different degraded modes on purpose — and each has a named owner who agreed to it."
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Scoping | Lists features: swipe, match, chat, super like | Names three intents, commits to deck + match, scopes out live proximity with a reason | Identifies which capabilities are org platforms (safety, geo, notifications) vs product-specific |
| Sizing | Estimates users and storage | Derives swipe:deck:match ratios and sizes the write path from them | Converts sizing into $/month and team load; knows which component dominates cost |
| Correctness | Detects mutual likes with a read | Names the race; pair-keyed CAS; idempotent match ID | Defines exactly-once-per-pair as an audited invariant with an SLO and reconciliation |
| Consistency tiers | One consistency level for everything | Deck eventual, pair linearizable, safety fail-closed — stated and justified | Codifies the safety tier as a cross-product contract enforced by tooling |
| Geo | Geohash + neighbors | Load-balanced S2 shard map, 1–3 shard fan-out, rebalance procedure | Shard map as versioned infrastructure; reshard as a planned one-way door with a migration budget |
| Marketplace | Ranks by ML score | Two-sided scoring, impression caps, guardrail metrics | Ranking governance: objective owner, fairness reporting, regulator-readiness |
| Operations | "Add monitoring" | Named metrics and alerts per failure; degraded mode per component | Game days for deck-store loss and safety-store outage; error budget per tier |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Sizes from the workload | "Swipes are ~20× deck fetches, so the write path sizes the system." |
| Finds the silent failure | "Two people swiping on each other at once can both miss. I'll make that structurally impossible and count it anyway." |
| Separates consistency tiers | "The deck can be ten minutes stale. The block list can't be one second stale." |
| Owns the marketplace | "Ranking for right-swipes concentrates attention. I'd rank for mutual probability with an impression cap — and product owns that objective." |
| Treats privacy as data minimization | "If we only store a 1 km cell, no bug can leak more than 1 km." |
| Names the shard map as an artifact | "The shard map is versioned and in every query log." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| 15+ minutes on geohash/quadtree internals | Misallocated time on a commodity component |
| "Check if the other person liked me" with no race discussion | Misses the core correctness problem |
| Block list filtered only at candidate generation | Precomputed decks make this stale; safety failure |
| Returns coordinates or precise distance in the API | Privacy failure that has real-world precedent |
| No numbers for swipe volume or storage growth | Cannot justify any sharding or caching decision |
| "We'll use a graph database for matches" | Pattern-matching; matches are a KV invariant, not a traversal |
5.4 Common False Positives#
- Deep S2/H3 knowledge ≠ proximity-matching design. Knowing Hilbert curves is nice; knowing where to put the shard boundary is the signal.
- ML ranking vocabulary ≠ marketplace thinking. "Two-tower model with embeddings" without "who gets zero likes" misses the Staff point.
- "Eventually consistent" everywhere ≠ understanding consistency. The signal is choosing different tiers per edge.
- Drawing Kafka ≠ async design. The question is what stays synchronous (the CAS) and why.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing and ratios | 0–4 min | Intents, swipe:deck:match ratios, three consistency tiers |
| Entities and API | 4–6 min | Pair key, client_swipe_id, no coordinates out |
| High-level architecture | 6–11 min | Deck path, swipe path, async brain |
| Deep dive 1: match race | 11–19 min | Pair-keyed CAS, reconciliation |
| Deep dive 2: deck freshness & safety | 19–26 min | Triggers, fail-closed filter |
| Deep dive 3: geo sharding at peak | 26–33 min | Load-balanced shard map, hot shard mitigation |
| Deep dive 4: ranking / privacy | 33–40 min | Two-sided score, caps, grid snapping |
| Wrap-up | 40–45 min | Metrics, evolution, what was skipped |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response |
|---|---|---|
| "Now add Passport — swipe in another city" | Whether your deck is keyed by user location or by search location | "The deck key includes the search cell. Passport is just a different search cell; swipes and matches are unchanged." |
| "Add 'likes you' for paid users" | Reverse index, hot keys | "Swiper-keyed history from the log, plus a swipee-keyed sampled list with sharded counters." |
| "Add video profiles" | Blob path isolation | "Media goes through object storage + CDN; the deck carries URLs. Nothing in the match path changes." See Handling Large Blobs. |
| "Launch in the EU" | Data residency, right to erasure | "Regional clusters; erasure deletes pair rows and tombstones the user in every deck on next serve." |
| "Your match rate dropped 10%" | Debugging a silent metric | "Check lost pairs first (correctness), then ranking changes (policy), then supply (fewer active users nearby)." |
| "Make it real-time — show who's near now" | Scope discipline, privacy | "That's a separate intent. It needs a separate privacy review and 100 m precision with minute-level TTL." |
6.3 What to Deliberately Skip#
| Topic | Why Skip | One-Liner If Asked |
|---|---|---|
| Geohash bit interleaving | Commodity | "Library call; S2 cell IDs." |
| Chat internals | Separate system | "Match creates a chat; see Chat Messaging." |
| Photo storage | Standard blob path | "Object store + CDN, signed URLs, moderation pipeline async." |
| Payment for boosts | Standard billing | "Boost is a ranking multiplier with an expiry; billing is a separate system." |
| Model architecture | Not the systems question | "Two-sided score; I'll leave the model to the ranking team." |
6.4 Follow-Up Questions to Expect#
- "What happens if both users swipe at exactly the same millisecond?"
- "A user swipes through 200 cards in 3 minutes. What's the load on your system?"
- "How do you avoid showing someone a person they already swiped left on six months ago?"
- "How much storage do swipes take after a year, and what do you delete?"
- "The ranking service is down. What does the user see?"
- "A user blocks someone. Walk me through every surface that has to change."
- "How do you rebalance a geo shard without downtime?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design Tinder. You have 45 minutes."
Staff Answer
"Let me frame this before drawing. There are three intents hiding in 'Tinder': a discovery deck, mutual matching, and live proximity. I'll design the first two — they're the core loop — and scope out live proximity because it needs a different privacy posture.
Ratios first. ~20M DAU, ~80 swipes each → ~1.6B swipes/day, ~18K/s average, ~55K/s at the evening peak. One deck fetch serves ~20–50 swipes, so deck reads are an order of magnitude smaller and cacheable. ~2% of right swipes are mutual, so matches are ~1–2K/s at peak.
That gives me three consistency tiers: deck eventual (minutes), pair exact (per-pair linearizable), safety fail-closed. The architecture follows: precomputed decks in Redis, a pair-keyed store with conditional writes for swipes, Kafka for everything downstream, and a fail-closed safety filter at serve time. The geo index is only touched by the async candidate generator."
Why this is L6:
- Sizes from ratios before drawing, which justifies every later decision
- Commits to intents and explicitly scopes one out with a reason
- Names consistency tiers per edge rather than one global choice
What L7 adds:
- Notes which components are shared org platforms (safety, geo, notifications) and which are Tinder-specific
- Prices the tiers: the pair store's conditional writes and the recompute fleet dominate cost
- Flags live proximity as a legal/privacy review gate, not just an engineering scope cut
Drill 2: The Core Mechanic — Simultaneous Swipes#
Prompt: "Alice and Bob swipe right on each other within 10 ms. Walk me through it."
Staff Answer
"Both requests land on different Swipe Service nodes. Both compute the same pair key, min(alice,bob):max(alice,bob), so both target the same row on the same partition. Each does a consistent read, gets version 7 with both directions empty, and attempts CAS to version 8. One wins — say Alice's: row now has alice→bob = like. Bob's CAS fails on the version check. Bob's handler re-reads: version 8, Alice's like present. Bob's write sets bob→alice = like, sees both likes and matched_at = null, sets matched_at and match_id = hash(pair), CAS to version 9 succeeds. Bob's response says matched: true; a match_created event goes to Kafka; Alice gets the match via push within ~1–2 s.
Two properties matter: exactly one handler transitioned matched_at from null, so exactly one match event is emitted; and match_id is deterministic, so even a bug that emits twice collapses downstream. The reconciliation job is the belt to the suspenders."
Why this is L6:
- Walks the race at the row level with versions, not hand-waving "use a transaction"
- Shows the retry path and why the loser still returns the right answer
- Adds deterministic IDs and reconciliation as defense in depth
What L7 adds:
- Makes "lost match rate" a reported business metric with an owner
- Decides the pair-store technology across products (DynamoDB conditional writes vs Cassandra LWT) based on cost per million writes and ops load
- Writes the invariant into the platform contract so other teams' features (e.g., "mutual friend" suggestions) reuse it
Drill 3: Make It Concrete — Storage Growth#
Prompt: "How big is the swipe data after a year, and what do you keep?"
Staff Answer
"1.6B swipes/day × ~50 bytes = ~80 GB/day raw, ~240 GB/day at RF 3, ~29 TB/year raw or ~88 TB replicated — before indexes and the swiper-keyed history copy, so call it ~150–200 TB/year.
What do we actually read? (1) The pair row for match detection — only matters until a match or ~6–12 months of inactivity. (2) The seen-set to avoid re-showing — only IDs, and a Bloom filter covers it. (3) Analytics — belongs in the warehouse, not the online store.
So: online pair store with a TTL of ~180 days on pass-only rows (left swipes are ~70–80% of volume and never produce a match); right-swipe rows keep until matched or 12 months; matched rows move to the match table. Full history goes to cold object storage via the Kafka log for analytics. That cuts the online footprint by ~60–70%. The re-show risk after TTL is a product decision — 'you might see someone again after 6 months' — and I'd get product to sign off."
Why this is L6:
- Numbers with replication factor, not just raw bytes
- Asks what is read before deciding what to keep
- Turns retention into an explicit product sign-off
What L7 adds:
- Prices the storage: ~$X/TB-month online vs ~$0.02/GB-month cold, and the delta in dollars
- Aligns retention with privacy law (erasure requests must reach the cold archive too)
Drill 4: Dependency Down — Ranking Service Outage#
Prompt: "The ranking model service is returning 500s. What happens?"
Staff Answer
"Existing decks are unaffected — they're precomputed. Only refills are hit. Refills fail open: the candidate generator falls back to distance + recency + a random exploration shuffle, and tags the deck ranked=false so it's regenerated when ranking recovers. Users see slightly worse cards; nobody sees an empty deck. Alert on ranking.fallback_decks_ratio > 20% for 10 minutes — that's a page for the ranking team, not for matching.
What must not fail open is the safety filter. If that store is down, we serve nothing and show a retry — better an empty deck than a blocked person."
Why this is L6:
- Distinguishes existing decks from refills — blast radius reasoning
- Fail-open for ranking, fail-closed for safety, stated side by side
- Routes the page to the right owner
What L7 adds:
- Defines an error budget per tier so ranking outages don't consume the platform's budget
- Runs a quarterly game day on "ranking down" to verify the fallback path actually works
Drill 5: Hot Key — The Celebrity#
Prompt: "A celebrity joins and gets 2 million right swipes in a day."
Staff Answer
"Pair-keyed storage doesn't care — 2M separate pair rows. What breaks is anything keyed by the celebrity: a 'likes you' counter, a reverse list for the paid 'see who liked you' grid, and their own match notifications if many are mutual.
Counter: shard to 16 sub-keys and sum on read; display '99+' above a threshold so exact counts aren't needed. Reverse list: bounded sample of the most recent ~1K likers, not the full set. And upstream: impression caps mean the celebrity appears in a bounded number of decks per hour. That's the real fix — the hot key is a symptom of unbounded exposure, which is also a marketplace problem."
Why this is L6:
- Identifies which keys are hot and which aren't
- Uses product degradation ("99+") to avoid an engineering problem
- Connects the hot key to the ranking policy
What L7 adds:
- Establishes a "verified public figure" policy with Trust & Safety — celebrities attract impersonation and harassment, a governance issue as much as a storage one
Drill 6: Multi-Tenant — Multiple Apps on One Platform#
Prompt: "The company runs three dating apps. Should they share infrastructure?"
Staff Answer
"Share the things where divergence is dangerous or duplicative; separate the things that are the product.
Share: the safety platform (a person banned for fraud on one app is a risk on all three — subject to privacy law and consent), the geo index platform, notifications, and the pair-store pattern. Separate: ranking objectives, deck UX, and user pools (unless product wants cross-app matching, which is a legal/consent conversation).
Tenancy model: each app is a tenant with its own index namespace and its own impression caps; shared shard map; per-tenant quotas so one app's launch can't starve the others."
Why this is L6:
- Draws the platform/product line by risk and differentiation
- Addresses noisy-neighbor isolation
- Raises consent for cross-app safety sharing
What L7 adds:
- Makes the org decision: a Trust & Safety platform team with a contract all apps must call, funded centrally
- Plans the migration: which app moves first and what the rollback is
Drill 7: Build vs Buy — The Geo Index#
Prompt: "Would you build a custom geo index?"
Staff Answer
"No. Elasticsearch/OpenSearch geo_point handles radius + compound filters (age, gender, last active) in one query, which is exactly the candidate query. Redis GEO handles pure radius. PostGIS handles it at smaller scale. What I build is the shard map and the refill policy — that's where our workload differs from anyone else's.
The build-custom threshold: if candidate generation becomes > ~40% of infra cost and our query shape is fixed, an in-memory cell index (S2 cell → sorted user list) can be 5–10× cheaper. That's a year-2 optimization with a measured business case, not a day-one choice."
Why this is L6:
- Buys the commodity, builds the differentiator
- States a quantitative trigger for revisiting
What L7 adds:
- Evaluates vendor/license risk (Elasticsearch vs OpenSearch licensing history) as a lock-in consideration
- Plans who staffs a custom index for 3 years before approving it
Drill 8: Policy Change Without an Outage — New Ranking Objective#
Prompt: "Product wants to switch ranking from right-swipe probability to mutual probability. Ship it."
Staff Answer
"This changes who gets attention, so it's a marketplace change, not a model swap. Shadow first: compute both scores for 7 days, compare deck overlap and projected like distribution. Then A/B at 1% → 5% → 25% by metro, with guardrails: mutual-match rate (primary), right-swipe rate (may drop — expected), like-distribution Gini (must not rise), and 7-day retention for the bottom-50% of profiles by exposure. Rollback is a config flag that regenerates decks on next refill — no deploy. Product owns the go/no-go; I own the guardrail dashboards."
Why this is L6:
- Shadow → canary → enforce with explicit guardrails
- Rollback without deploy
- Splits ownership between product decision and engineering guardrails
What L7 adds:
- Establishes a ranking-change review board with fairness reporting for all objective changes
- Anticipates external scrutiny (regulators increasingly ask about recommender systems) and keeps the audit trail
Drill 9: Cost#
Prompt: "Finance says infra cost per DAU doubled this year. Where do you look?"
Staff Answer
"Three suspects, in order. (1) Refill rate: if deck size shrank or the low-water mark rose, candidate generation runs more often — check refills per active user per day; target ~2–4. (2) Swipe storage growth without TTL on pass rows — the online store grows linearly forever. (3) Geo index replicas added during peaks and never removed.
Quick wins: TTL pass-only rows at 180 days (~60% storage cut), increase deck size from 100 to 200 (halves refills), and autoscale index replicas on a schedule. I'd report cost per DAU by component monthly so this is a trend, not a surprise."
Why this is L6:
- Knows which knobs drive cost and their direction
- Concrete targets (refills per user per day)
What L7 adds:
- Builds a unit-economics model (cost per match) the business uses for pricing decisions
- Assigns cost ownership to teams via showback
Drill 10: Multi-Region#
Prompt: "Expand from the US to the EU and India."
Staff Answer
"Users match with people near them, so geography gives us natural partitioning: each region owns its users, its geo shards and its pair store. Cross-region pairs only happen with travel mode — route those swipes to the swipee's home region so the pair row still has one owner. EU data stays in the EU for residency; erasure requests fan out to the swiper-keyed history and cold archive.
Global pieces: the safety graph (bans must follow a person across regions — replicated with strong consistency on writes, which are rare), identity, and payments. Region failover: deck serving fails over to a stale read-only copy; swipes queue on the client for up to ~5 minutes rather than being written to a non-owning region."
Why this is L6:
- Uses geography as a natural partition and names the one cross-region case
- Keeps single ownership of pair rows under travel mode
- Distinguishes regional vs global state
What L7 adds:
- Makes the residency and erasure design a signed contract with Legal
- Decides the region-launch sequence based on revenue and regulatory cost, not just engineering readiness
8. Deep Dive Scenarios#
Deep Dive 1: Saturday Night Peak — Empty Decks in One Metro#
Context: Saturday, 22:00 local. Users in one large metro report "There's no one new around you." Global dashboards look green. The on-call escalates to you.
Questions to Surface First:
- Is this one metro or several? Do the affected metros share geo shards?
- Are decks empty because candidate generation is slow, or because it returns zero results?
- Did anything change today — shard map version, ranking model, a filter?
- Is the refill queue growing, and is it per-shard or global?
Typical L5 Approach: Checks the Deck Service error rate (zero — it is returning empty lists successfully), scales the Deck Service, then scales the Elasticsearch cluster uniformly. Adds capacity to shards that are idle while the hot one stays hot.
Staff Approach: Goes straight to per-shard skew:
geo.shard_cpu_skewandrecs.refill_queue_depth{shard}. Confirms the affected metro maps to 1–2 shards at 95%+ CPU. Enables the pre-built overload mode: serve stale decks past their 6 h expiry, cap refill concurrency on the hot shards, switch refills there to distance-only ranking, and double refill size to 300. Then adds read replicas to the hot shards only.
Principal Approach: Asks why a predictable weekly peak can page anyone. Makes shard capacity a function of the weekly load curve per metro, with scheduled replica scaling owned by the Search platform, and adds "metro peak headroom ≥ 40%" to the quarterly capacity review. Pushes Recommendations and Search to agree on a single overload contract (what degrades first) so the next incident is a runbook step, not a debate.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Confirm per-shard skew; enable stale-deck serving (no expiry) for the affected metro |
| Triage | Is the hot shard hot from refills or from live queries? Check refills_per_active_user — if > 10/h, clients are burning decks faster than normal (UI change? bot swarm?) |
| Quick fix | Refill concurrency cap per shard; distance-only ranking; deck size 300; add 2 read replicas to the hot shard |
| Guardrails | Watch deck.empty_responses_total fall below 1%; don't let replica addition trigger a rebalance storm |
| Post-mortem | Split the hot shard (it has been > 1.5× median for weeks?); add scheduled replica scaling for weekend peaks |
Metrics to Watch: geo.shard_cpu_skew, recs.refill_queue_depth{shard}, deck.empty_responses_total, refills_per_active_user, deck.served_stale_ratio
Organizational Follow-up: Shard split criteria (1.5× median for 7 days) become an automated ticket, not a judgment call. Weekend peak scaling becomes a scheduled job owned by Search platform.
Ownership Question: "Who decides to degrade ranking quality during an overload?" Staff answer: Recommendations owns the degrade ladder and pre-approved it; the on-call executes step 1–3 without asking. Step 4 (disable refills entirely) requires the Recommendations lead.
Key Takeaway: "An empty deck is an availability incident even when every service returns 200. Measure empty responses, not errors."
What clears the Staff bar:
- Looks at per-shard skew before global capacity
- Has a pre-agreed degrade ladder
- Converts a recurring peak into scheduled capacity
Deep Dive 2: The Silent Failure — Match Rate Down 3%#
Context: The weekly business review shows mutual-match rate down 3% week over week. No incidents were declared. The VP of Product asks engineering whether something is broken.
Questions to Surface First:
- Did right-swipe volume change, or did mutual detection change?
- Is
match.lost_pairs_totalnon-zero? When did the reconciliation job last run successfully? - Any ranking, shard map, or store changes in the window?
- Is the drop uniform across metros and platforms (iOS vs Android)?
Typical L5 Approach: Checks error rates on the Swipe and Match services — all normal. Suggests it's seasonal or a ranking effect and hands it to data science.
Staff Approach: Separates correctness from policy. First, correctness: queries pair rows with both likes and no
matched_at— finds 1.1% of mutual pairs unmatched since a store upgrade 9 days ago that changed the default read consistency on one client library fromLOCAL_QUORUMtoONE, and the reconciliation job had been silently failing on a permissions change. Then policy: the remaining ~2% is a ranking experiment that shifted exposure. Two causes, two owners.
Principal Approach: Treats the reconciliation job failing silently as the real incident. Requires every correctness canary to have a heartbeat alert ("job has not reported in 2 h") and makes consistency level an enforced, reviewed setting in the shared client library — not a per-call default. Adds "lost match rate" to the weekly business review so it is a number executives see, not one engineers discover.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Check the reconciliation job's last successful run; run it manually on the last 24 h |
| Triage | Split the drop by cause: lost pairs vs exposure change vs supply change (active users per metro) |
| Quick fix | Restore consistency level; backfill missed matches from the last 9 days — product decides whether to surface them ("You have new matches") |
| Guardrails | Heartbeat alert on the reconciliation job; match.lost_pairs_total > 0 for 15 min pages |
| Post-mortem | Library defaults for consistency are pinned and tested; experiment guardrails include mutual-match rate |
Metrics to Watch: match.lost_pairs_total, reconcile.last_success_age_s, match.rate_by_experiment_arm, swipe.right_rate
Organizational Follow-up: Client-library changes that touch consistency require Matching platform review.
Ownership Question: "Who decides whether to notify the 9 days of missed matches?" Staff answer: Product, with Trust & Safety consulted — a late match between people who've since blocked or deleted must be filtered first. Engineering provides the list and the safety filter, not the decision.
Key Takeaway: "The silent failure in this system is a missing row. Only a job that looks for missing rows will ever find it — and that job needs its own alert."
What clears the Staff bar:
- Separates correctness regressions from policy/ranking effects
- Suspects the canary before trusting it
- Routes the customer-facing decision to product
Deep Dive 3: Large-Customer Onboarding — A University Campus Launch#
Context: Marketing is launching a campus program: 40 universities, ~800K new signups expected in 72 hours, heavily concentrated in small towns that currently map to low-traffic shards.
Questions to Surface First:
- Which shards cover these towns, and what's their current headroom?
- Are new users' decks cold-started (no ranking features) — what does candidate generation cost for them?
- Does the campus program have its own ranking rules (e.g., students-only filter)?
- What is the expected swipe rate per new user in the first session? (New users swipe 2–3× more.)
Typical L5 Approach: Scales the whole Elasticsearch cluster by 30% ahead of launch.
Staff Approach: Projects load per shard: small-town shards will go from ~10K to ~60K actives. Pre-splits those shards before launch using the shard map tooling, pre-warms the pair store partitions, and sets a cold-start deck path (distance + recency + exploration) that doesn't need ML features. Adds a campus filter as a tag on documents, not a separate index. Watches new-user swipe velocity against right-swipe caps to avoid bots riding the launch.
Principal Approach: Establishes a "launch readiness" process between Marketing and Platform: any campaign projecting > 5% regional growth files a capacity request 2 weeks ahead, with per-shard projections generated by a tool. Makes shard pre-splitting self-service for Recommendations so Platform isn't a bottleneck.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| T-14 days | Per-shard projection; identify 6 shards exceeding 1.5× median post-launch |
| T-7 days | Split shards, dual-read verify, flip shard map version |
| T-1 day | Scale pair store; enable cold-start deck path; lower refill low-water mark to 30 for new users |
| Launch | Watch per-shard CPU, deck.empty_responses_total, swipe.right_ratio for new accounts |
| T+7 days | Remove temporary replicas; review shard map |
Metrics to Watch: geo.shard_active_users{shard}, newuser.first_session_swipes, deck.cold_start_latency_p99, abuse.new_account_right_ratio
Organizational Follow-up: Launch-readiness checklist owned by Platform PM; Marketing's campaign calendar feeds capacity planning.
Ownership Question: "Who pays for the pre-provisioned capacity if the launch underperforms?" Staff answer: It's a marketing launch cost, charged back — which is exactly why Marketing should see the projection and sign it.
Key Takeaway: "Growth lands on shards, not clusters. Project per shard."
What clears the Staff bar:
- Per-shard projection instead of uniform scaling
- Cold-start path designed for new users
- Bot risk tied to a launch
Deep Dive 4: Post-Mortem — Blocked User Reappeared, Press Inquiry#
Context: A user posted on social media that a person she blocked appeared in her deck two days later. A journalist has asked for comment. You're leading the post-mortem.
Questions to Surface First:
- Which surface showed the blocked person — deck, likes-you grid, or a push notification?
- Was the block recorded in the safety store? When?
- Did the blocked person create a new account (ban evasion) or is it the same account?
- Is the serve-time safety filter on that surface?
Typical L5 Approach: Finds the bug — the "likes you" grid was built by a different team and read from a cache that didn't check the block list — and fixes that cache.
Staff Approach: Fixes the bug, then asks why a surface could render a person without calling the safety filter at all. Audits every surface that renders a person (7 found, 2 without the filter). Moves the filter into the shared "render a profile card" API so it cannot be bypassed. Adds
safety.serve_filter_hits_totalper surface. Separately investigates ban evasion: same device fingerprint, new account.
Principal Approach: Writes the Trust & Safety contract: every surface that shows a person MUST call the safety API; CI rejects services that fetch profile cards without it; quarterly audit. Funds a device-level identity signal for ban evasion. Coordinates the external response with Comms and Legal with an accurate technical timeline.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Confirm the account/surface; hard-remove blocked ID from all caches for this user; verify |
| Triage | Audit all surfaces; measure how many blocked pairs were exposed in the last 30 days |
| Quick fix | Add the filter to the two missing surfaces |
| Guardrails | Structural: profile-card API enforces safety; per-surface metrics |
| Post-mortem | Blameless; the failure is a missing contract, not a missing line of code |
Metrics to Watch: safety.serve_filter_hits_total{surface}, safety.blocked_exposure_total, ban_evasion.suspected_accounts
Organizational Follow-up: T&S platform owns the contract; each product surface owner signs off on compliance.
Ownership Question: "Who is accountable when a new surface ships without the filter?" Staff answer: The shipping team — but the platform must make it impossible to do by default. Accountability without guardrails is just blame.
Key Takeaway: "A safety filter that each team must remember to call is a safety filter that will be skipped."
What clears the Staff bar:
- Generalizes from one bug to a class of surfaces
- Makes the correct path the only path
- Distinguishes bug from ban evasion
Deep Dive 5: Multi-Region Expansion — India Launch#
Context: The company launches in India. Data residency requirements and a very different density profile (extremely dense metros, mobile networks with high packet loss).
Questions to Surface First:
- What data must stay in-country? Profiles, swipes, messages, location?
- What are p99 RTTs from major Indian cities to the nearest current region? (~150–250 ms to Singapore is common.)
- Is cross-border matching (travel mode) allowed?
- What's the expected density — how many actives per S2 level-13 cell in Mumbai vs current top metros?
Typical L5 Approach: Deploys a copy of the stack in an Indian cloud region.
Staff Approach: Deploys a regional cell with its own geo shards, pair store and deck store; safety graph replicated globally (bans follow a person). Recalibrates the shard map for density (Mumbai cells may need level 14–15 granularity). Client-side: larger deck prefetch (40 cards) and swipe batching (send up to 10 swipes per request) for lossy networks, with
client_swipe_ididempotency making retries safe. Travel-mode swipes route to the swipee's home region.
Principal Approach: Treats the region as a template: "region-in-a-box" with a residency checklist, a shard map generator, and a launch runbook — so the next 5 regions cost weeks, not quarters. Makes a portfolio decision about which global services (identity, safety, payments) must become multi-region before more launches.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Design | Regional cell; global safety and identity; residency map signed by Legal |
| Pre-launch | Density-calibrated shard map from pre-registration data; load test at 3× projection |
| Launch | Swipe batching and deck prefetch tuned for 3–5% packet loss |
| Steady state | Per-region SLOs; region-local on-call |
| Evolution | Region template for next launches |
Metrics to Watch: swipe.batch_retry_rate, deck.prefetch_hit_ratio, region.cross_region_swipe_ratio, geo.shard_cpu_skew{region}
Organizational Follow-up: Region-local on-call or follow-the-sun; residency compliance audit.
Ownership Question: "Who owns the global safety graph's availability SLO?" Staff answer: T&S platform — and because it's fail-closed, its availability SLO is effectively the deck's SLO in every region. It needs to be the most reliable service we run.
Key Takeaway: "A fail-closed global dependency sets the availability ceiling for every region. Budget for it."
What clears the Staff bar:
- Regional vs global state separated explicitly
- Network-conditions-aware client design
- Notices fail-closed dependencies cap availability
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Size the system from swipe:deck:match ratios and justify why the write path dominates
- Explain the simultaneous-swipe race and fix it with a pair-keyed conditional write
- Assign consistency tiers: deck eventual, pair exact, safety fail-closed
- Design load-balanced geo sharding with S2 cells and a versioned shard map
- Name the marketplace problem in ranking and propose impression caps and two-sided scoring
- Coarsen location to defeat trilateration
- Name the owner and metric for every failure mode
- At L7: price the system, define the safety contract across products, and plan the shard-map evolution
The Bar for This Question#
Mid-level (L4): Designs a geo query, a swipe table and a match check. Works in the demo. Doesn't estimate load, doesn't see the race, doesn't mention safety or privacy.
Senior (L5): Adds caching, sharding by geohash, Kafka for notifications, and reasonable storage estimates. Knows eventual consistency exists. Misses the pair-keyed invariant, treats all data at one consistency level, and ranks by ML score without thinking about who gets attention.
Staff+ (L6): Frames the three intents, sizes from ratios, builds exactly-once matching structurally, separates safety as fail-closed, shards geo by load, and names owners for ranking, safety and the shard map. Every failure has a metric. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "Tinder Is Not a Geo Problem"#
| Evidence | Implication |
|---|---|
Geo indexing is a library call (S2, H3, Redis GEO, ES geo_point) | Designing it is misallocated time |
| Swipes are ~20× deck fetches | Write path sizes the system |
| The core correctness bug is a two-row race | The hard problem is contention, not geometry |
The Staff position: Spend 2 minutes on geo, 25 on swipes, safety and marketplace.
Why this matters in interviews: Interviewers who ask this prompt have usually seen 50 geohash walkthroughs. The candidate who reframes early stands out.
10.2 "Precise Location Is a Liability, Not a Feature"#
| Evidence | Implication |
|---|---|
| Trilateration attacks on dating apps are publicly documented | Precision leaks |
| Relevance at 2–160 km radii is insensitive to sub-km precision | Precision buys nothing |
| Stored precise location is discoverable in breaches and legal process | Liability grows with retention |
The Staff position: Store cells, not coordinates. Round everything that leaves the server.
Why this matters in interviews: Privacy by data minimization is a Staff-level signal; "we'll encrypt it" is not.
10.3 "Ranking for Engagement Kills Dating Apps"#
| Evidence | Implication |
|---|---|
| Likes concentrate heavily on a minority of profiles in one-sided ranking | Most users get little attention and churn |
| The product outcome is a mutual match, not a right swipe | One-sided objectives optimize the wrong thing |
| Hinge publicly frames its compatibility feature around stable matching | Two-sided thinking is a known industry direction |
The Staff position: Rank for mutual probability with exposure caps; measure concentration.
Why this matters in interviews: It shows you think about the marketplace, not just the model.
10.4 "The Reconciliation Job Is More Important Than the Algorithm"#
| Evidence | Implication |
|---|---|
| Lost matches produce no errors and no tickets | Only a scan finds them |
| Config drift (consistency levels, library defaults) reintroduces races | Structural fixes regress |
| Canaries fail silently too | The canary needs a heartbeat |
The Staff position: Structural correctness plus an audited, alerting reconciliation job.
Why this matters in interviews: Staff engineers design for the day their design is subtly broken.
10.5 "Don't Geo-Shard Until It Hurts"#
| Evidence | Implication |
|---|---|
| A single PostGIS or ES cluster handles millions of users in a region | Sharding is premature for most apps |
| A shard map is a new production artifact with rebalancing tooling | Ongoing ops cost |
| Resharding later is a one-way-ish migration | Do it once, when forced |
The Staff position: One cluster until fan-out or size forces geo-sharding; then load-balanced S2 ranges.
Why this matters in interviews: Knowing when not to build is as strong a signal as knowing how.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
A Staff engineer designs Tinder. A Principal engineer notices that the company doesn't run one proximity product — it runs a portfolio (dating apps, friend-finding, events, local marketplaces), and every one of them rebuilds the same four capabilities: a geo candidate index, a mutual-consent invariant, a safety graph and a notification fan-out. The L7 question is which of those are platforms with contracts and which are product differentiation, and what it costs — in dollars, headcount and trust incidents — to leave that line undrawn. The safety graph in particular is not an engineering nicety: it is the one capability where inconsistency between products becomes a public incident.
The Org-Level Fault Line#
One shared "people-nearby" platform vs per-product stacks.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Per-product stacks | Product speed; each app tunes its own ranking and deck | Safety state diverges (a ban on one app isn't seen on another); 3–5× duplicated geo infra; inconsistent privacy posture | Trust & Safety, Legal, Finance |
| Full shared platform (geo + decks + matching + safety) | One safety posture, one cost base | Platform becomes a bottleneck for product ranking experiments; lowest-common-denominator UX | Product velocity |
| Shared primitives, product-owned policy (L7 default) | Safety graph, geo index, pair-invariant store and notifications are platforms; ranking objective and deck UX are product-owned | Needs clear contracts and a funded platform team | Platform headcount (~6–10 engineers) |
🧭 Principal Move: "Centralize the things where divergence creates harm — safety and privacy — and the things that are pure commodity — geo and notifications. Leave ranking to products. The platform's contract is the safety API and the pair invariant; everything above that line is product."
Cost Model#
Assumptions: cloud list prices order-of-magnitude, ~80 swipes/DAU/day, RF 3, 180-day TTL on pass rows, precomputed decks, loaded engineer cost ~$250K/yr.
| Scale | Infra $/month | Dominant Cost | Headcount (eng) | On-Call Load |
|---|---|---|---|---|
| 1M DAU (~80M swipes/day, ~3K/s peak) | ~$25–50K | Single ES cluster + managed KV | 4–6 across one team | 1 rotation, ~1 page/week |
| 10M DAU (~800M swipes/day, ~28K/s peak) | ~$200–400K | Pair store writes, candidate generation | 15–25 across 3 teams (matching, recs, T&S) | 3 rotations, ~3–5 pages/week total |
| 50M DAU (~4B swipes/day, ~140K/s peak) | ~$1–2M | Pair store (~40%), recs compute (~30%), geo index (~15%) | 50–80 across 6+ teams, incl. platform | Per-team rotations + region on-call |
The line that changes the business case: at 10M+ DAU, pair-store writes and candidate recompute dominate. Every 2× increase in deck size roughly halves recompute; every TTL cut on pass rows reduces storage linearly. These two knobs are worth more than any index optimization.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse | Why |
|---|---|---|---|
| Pair-keyed vs swiper-keyed source of truth | One-way-ish | Full data migration of 100+ TB, dual-write period | Invariant lives in the key; pick right on day one |
| Storing precise coordinates | One-way | Can't un-leak historical data; breach liability persists | Data minimization only works if done before collection |
| Shard map granularity (S2 level) | Two-way with effort | Reindex per shard, weeks | Versioned map makes it tractable |
| Geo index vendor (ES vs OpenSearch vs custom) | Two-way with effort | Months; query translation | Keep candidate generation behind an interface |
| Ranking objective | Two-way | Config flag + deck regeneration | Treat as experiment, never as schema |
| Deck size / low-water mark | Two-way | Minutes | Tune freely |
| Region residency boundaries | One-way | Legal commitments to regulators | Legal sign-off required |
The Standard I'd Write#
RFC: People-Surface Safety & Proximity Data Standard (v1)
Scope: Every service that renders, recommends, notifies about, or stores the location of a person, across all apps.
Mandatory requirements:
- Services MUST fetch profile cards through the Profile Card API, which applies the safety filter. Direct profile-table reads for rendering are rejected in CI.
- The safety filter MUST fail closed. A surface that cannot reach the safety service MUST render nothing.
- Services MUST NOT store raw coordinates longer than 60 seconds. Persisted location MUST be an S2 cell at level ≤ 13 (or H3 resolution ≤ 8).
- APIs MUST NOT return coordinates of another user. Displayed distances MUST be rounded to ≥ 1 mile/km.
- Mutual-consent features MUST use the pair-invariant store and MUST have a reconciliation job with a heartbeat alert.
- Ranking objective changes SHOULD ship with exposure-concentration guardrail metrics.
Exceptions: Filed with the T&S platform and Privacy; time-boxed to 90 days; reviewed monthly.
Success metrics: 0 blocked-user exposures per quarter;
match.lost_pairs_total= 0 for 30 consecutive days; 100% of people-rendering surfaces on the Profile Card API within 2 quarters.
What I'd Tell the VP#
"The part of this product that can hurt us most isn't the matching algorithm — it's whether a person someone blocked can ever reappear, and whether we hold location data we don't need. I want one safety and location standard across all our apps, owned by one platform team of about eight engineers. It costs roughly $2M a year in people and saves us duplicate infrastructure in each app. It also turns our worst-case headline into something we can prove doesn't happen. Ranking and product experience stay with each app team so they keep their speed."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Portfolio view | "Which of these four capabilities do our other apps rebuild?" |
| Prices the tradeoff | "Deck size is the biggest cost knob: 2× deck size halves recompute, about $X/month at 10M DAU." |
| One-way door awareness | "Storing precise coordinates is the decision we can't take back — I'd kill it before launch." |
| Writes the standard | "Profile cards go through one API that fails closed, and CI enforces it." |
| Governance of ranking | "Ranking objective changes are marketplace policy. They get a review and a fairness metric." |
Staff answers that L7 interviewers find insufficient:
- "We'll add the safety filter to every surface." — Correct, but relies on every team remembering; L7 makes bypass impossible.
- "We'll shard by load with S2." — Correct, but doesn't address who owns the shard map across products or when reshards happen.
- "Ranking should be two-sided." — Correct, but doesn't say who approves objective changes or how the org measures marketplace health over years.
Appendices
Appendix A: Mechanics in Depth#
A.1 Candidate Generation Pipeline#
Preference filtering must be bidirectional: A's preferences must admit B, and B's preferences must admit A. A one-directional filter wastes deck slots on people who will never see A.
candidates = geo_index.search(
cells = s2_cover(user.cell, user.max_distance_km),
filter = age in user.age_range AND gender in user.genders
AND user.age in cand.age_range AND user.gender in cand.genders
AND last_active > now - 30d,
size = 1500)
candidates = [c for c in candidates if not seen_bloom.might_contain(c.id)]
candidates = safety.filter(user.id, candidates) # authoritative, fail-closed
scored = rank(user, candidates) # P(u→c) × P(c→u)
deck = apply_caps_and_exploration(scored, n=150)
redis.del(deck_key); redis.rpush(deck_key, *deck); redis.expire(deck_key, 6h)
A.2 Why Bloom Filters Fit the Seen Set#
A false positive hides a candidate the user hasn't seen (harmless at ~1%). A false negative is impossible, so a seen candidate is never re-shown by the filter. At 10K swipes and 1% FPR: m ≈ 9.6 bits × 10K ≈ 12 KB, k = 7 hashes. Heavy swipers (100K lifetime) roll to a new filter per 6-month epoch, matching the pass-row TTL.
A.3 Distance Rounding#
display_km = max(1, round(haversine(cell_center(u), cell_center(c))))
Computed from cell centers, not raw coordinates, so no amount of repeated queries converges below the cell size.
Appendix B: Keys and Data Model#
| Entity | Key | Store | TTL |
|---|---|---|---|
| PairState | pair:{min}:{max} | DynamoDB / Cassandra (conditional writes) | 180 d pass-only; 12 mo like; permanent if matched |
| SwipeHistory | hist:{swiper} clustering by ts | Cassandra (from Kafka) | 12 mo |
| Deck | deck:{user}:{search_cell} | Redis list | 6 h |
| SeenBloom | seen:{user}:{epoch} | Redis string | 6 mo |
| Match | match:{hash(pair)} | Relational or KV | Permanent until unmatch/delete |
| SafetyEdge | block:{blocker}:{blocked} + reverse | Strongly consistent KV | Permanent |
| GeoDoc | user_id in index shard by S2 range | Elasticsearch | Updated on location change |
Appendix C: Coordination Mechanisms — Quick Comparison#
| Mechanism | Latency | Exactly-One Match | Ops Cost | Fit |
|---|---|---|---|---|
| DynamoDB conditional write on pair item | 1 RTT, ~5–10 ms | Yes | Low (managed) | Best default |
| Cassandra LWT on pair row | ~4 RTT Paxos, ~15–30 ms | Yes | Medium | OK at swipe volume per row |
| Redis Lua script on pair key | < 1 ms | Yes while Redis is up | Durability risk on failover | Only with durable log replay |
| Single-writer routing by pair hash | < 1 ms in-memory | Yes | High (router, failover) | Very high scale |
| Quorum reads + idempotent match ID | 2 RTT | Duplicates collapse | Low | Fallback when no CAS |
Appendix D: API Contract and Client Behavior#
POST /swipescarriesclient_swipe_id; retries within 24 h return the original result.- Clients batch up to 10 swipes per request on poor networks; server processes in order per pair.
- Deck prefetch: request more when ≤ 5 cards remain locally; server refill is triggered at ≤ 20 server-side.
- On 503 from
/swipes, clients retry with jittered exponential backoff (base 200 ms, cap 5 s) and keep swipes in a local queue for up to 5 minutes. /locationaccepted at most once per 5 minutes unless movement > 1 km; the server returns 204 and never echoes location.
Appendix E: Observability#
E.1 Core Metrics#
| Metric | Why |
|---|---|
deck.latency_p99, deck.empty_responses_total | Availability of the core UX |
match.lost_pairs_total, reconcile.last_success_age_s | The silent failure |
safety.serve_filter_hits_total{surface}, safety.filter_errors | Safety tier health |
geo.shard_cpu_skew, recs.refill_queue_depth{shard} | Hot shards |
swipe.cas_conflict_rate | Should be < 0.1%; spikes mean retry storms or bugs |
ranking.fallback_decks_ratio | Degraded ranking |
E.2 Critical Alerts#
| Alert | Threshold | Severity |
|---|---|---|
| Lost pairs | > 0 for 15 min | Page matching |
| Reconciliation heartbeat | no success in 2 h | Page matching |
| Safety filter errors | > 0.1% for 2 min | Page T&S |
| Empty decks | > 2% for 5 min in any metro | Page recs |
| Shard skew | max/median > 3 for 10 min | Page search platform |
E.3 Control Plane vs Data Plane#
Data plane: deck serve, swipe CAS, safety filter. Control plane: shard map, ranking config, caps, triggers. Control-plane changes are versioned, canaried by metro, and logged with the version in every data-plane log line.
Appendix F: Scale Evolution#
| Scale | What Works | What You Add |
|---|---|---|
| < 1M users | Postgres + PostGIS, query-on-open, pair-keyed table with row locks | Nothing else |
| 1–10M | ES geo_point, Redis decks, DynamoDB/Cassandra pair store, Kafka | Precompute, safety filter at serve |
| 10–50M | Load-balanced S2 shard map, two-sided ranking, reconciliation SLO | Regional cells |
| 50M+ | Shared people-nearby platform, region-in-a-box | Custom in-memory cell index if cost justifies |
F.1 What You Don't Build on Day One#
- Custom spatial index
- Geo-sharding
- ML ranking (distance + recency + exploration is enough early)
- Multi-region active-active pair store
- Live proximity
What you do build on day one: the pair-keyed invariant, cell-only location storage, and the fail-closed safety filter. Those are the one-way doors.
Appendix G: Fairness, Multi-Tenancy and Cost#
- Exposure fairness: track the Gini coefficient of impressions and likes per metro weekly; guardrail on ranking changes.
- Exploration budget: 5–10% of deck slots for new or under-exposed profiles; new users get a time-boxed boost of ~48 h.
- Tenancy across apps: per-app index namespace, per-app caps, shared shard map, per-app quota on candidate generation to prevent one launch from starving others.
- Showback: cost per DAU and cost per match by app and component, reported monthly.