Technologies referenced in this case study: DynamoDB · Cassandra · PostgreSQL · ZooKeeper & etcd
Related: Replicated Data Store covers replication and quorum internals — this case study links to it rather than repeating it · Consistency Models · Load Balancer · Degraded Mode · Distributed Lock Service
How to Use This Case Study#
Organized for interview use first, reference second. Read front-to-back once, then return to individual sections for targeted review.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines → Active Drills 1–4 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Fault Lines → Failure Modes → weak-spot Deep Dives |
| Deep Dive | 3+ hrs | Everything, including the Principal Lens and appendices |
What is Multi-Region Active-Active? — Why interviewers pick this topic
Multi-region active-active means two or more geographically separate regions serve production traffic at the same time, each able to take over the others' load. It differs from active-passive, where a standby region sits idle (or near idle) until a failover. The hard part is never the load balancer. It's the data: a write accepted in Virginia and a conflicting write accepted in Frankfurt 80 ms later — which one wins, who decided, and did a customer lose money?
Before vs After — primary region loss at 14:05 on a weekday:
Single region, backups in another region:
t=0: us-east loses a core network layer; 100% of requests fail
t=+5min: Incident declared. Decision: wait for the provider or restore elsewhere?
t=+45min: Team starts restoring last night's snapshot in us-west
t=+3h: Restore done, DNS cut over. 9 hours of writes missing (RPO = snapshot age)
t=+2 days: Manual reconciliation of orders. Finance and support absorb the cost.
Active-active, home-region ownership, rehearsed evacuation:
t=0: us-east error rate spikes to 60%
t=+45s: Health-based routing marks us-east unhealthy; global LB shifts new traffic
t=+2min: Evacuation runbook: us-west promotes followers for us-east-homed users
t=+4min: us-west at 190% of normal load — pre-provisioned headroom absorbs it
t=+6min: Error rate < 0.5%. Writes for us-east-homed users resume in us-west
with RPO ≈ replication lag at failure (~800 ms of writes at risk)
t=+3 days: us-east healthy. Failback is a planned, throttled re-homing — not a reflex.
Why interviewers reach for this question: Multi-region exposes whether a candidate can reason about the speed of light, partial failure, conflicting writes and money at the same time. It's also the question where "more" is easy to say and hard to justify: two regions roughly double infrastructure, add a replication system nobody has operated, and create a new class of bugs (split brain, lost writes on failback). Interviewers want to see you justify why active-active, which data goes active-active, and who pays.
Mechanics Refresher: Topologies and Replication Modes
| Option | How It Works | Pros | Cons |
|---|---|---|---|
| Backup & restore | Snapshots shipped to another region | Cheapest | RTO hours, RPO = snapshot age |
| Pilot light / warm standby | Data replicated async; minimal compute in standby, scaled on failover | 1.1–1.3× cost | RTO 15–60 min; failover path rarely exercised, so it rots |
| Active-passive (hot standby) | Full stack in standby, async replication, traffic only on failover | RTO minutes; simple data model (one writer) | ~2× cost for idle capacity; failover still a rare, scary event |
| Active-active, read-local / write-home | Every region serves reads; each entity has a home region that takes its writes | Low read latency everywhere; no write conflicts | Writes from far users pay cross-region RTT; requires an ownership map |
| Active-active, write-anywhere (multi-master) | Any region accepts writes for any entity; async replication + conflict resolution | Lowest write latency; no single owner | Conflicts: LWW loses data; CRDTs constrain the data model |
| Active-active, synchronous (global consensus) | Writes commit on a cross-region quorum (Spanner, CockroachDB multi-region) | Strong consistency, RPO = 0 | Every write pays ≥ 1 cross-region RTT (70–200 ms) |
For most production systems: active-active for the stateless tier and reads everywhere; home-region ownership for writes, with async replication to followers in other regions; synchronous global consensus only for the small set of data where RPO must be zero (ledgers, uniqueness). Write-anywhere only for data that is commutative by nature (counters, sets, presence).
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What This Interview Actually Tests#
Multi-region is not a load-balancing question. Everybody can draw two regions behind GeoDNS.
It is a data-ownership and blast-radius question: when two regions can both accept writes, who owns each piece of data, what happens when they disagree, and who decides to move traffic when one region is sick. It tests:
- Whether you ask why multi-region (latency, availability, residency) before choosing a topology
- Whether you classify data by its consistency need instead of picking one replication mode for everything
- Whether you know the speed of light sets a floor (NY–London ≈ 56 ms RTT theoretical, ~70–80 ms real)
- Whether you treat failover as a capacity, routing and human-decision problem — not just a database flag
- Whether you price it: active-active is rarely cheaper than 1.5× single-region cost
The key insight: Active-active is not "every region does everything". It's "every region can serve, and every piece of data has exactly one owner at a time". Conflict resolution is the fallback for the small set of data that cannot have an owner — not the default.
The L5 → L6 → L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Draws two regions, GeoDNS, and "multi-master replication" | Asks which intent — latency, availability or residency — and what RTO/RPO the business signed | Asks whether the business case survives the bill: active-active costs 1.5–2× infra plus a platform team, and many outages it prevents are not regional |
| Data | One replication mode for the whole database | Classifies data: strongly consistent (ledger) vs home-region owned (profiles, orders) vs commutative (counters, presence) | Makes data classification an org-level schema attribute; every new table declares its multi-region class at creation |
| Conflicts | "Last-writer-wins" | Avoids conflicts with home-region ownership; uses LWW only where loss is acceptable, CRDTs where data is commutative | Writes the rule: no LWW on data with financial or legal meaning; exceptions need sign-off from the data owner |
| Failover | "DNS fails over automatically" | Failover = routing + data promotion + capacity headroom; pre-provisioned N+1; evacuation drill quarterly | Designs org failure posture: cells smaller than regions, static stability (no control-plane calls during failover), evacuation as a routine monthly operation |
| Ownership | Platform/infra owns "the multi-region thing" | Product teams own data classification; platform owns routing and replication; incident commander owns the evacuation decision | Redraws the boundary: a regional-independence standard that every tier-1 service must meet, audited by game days, with a funded exceptions list |
| Cost | Not raised | Mentions 2× infra | Prices N/(N−1) capacity, cross-region egress, replication tooling and on-call; compares against the expected cost of regional outages |
Why "first move" separates levels
L5: Starts with topology. Two regions, GeoDNS, a multi-master database. Every component is real, but the design hasn't been told what it's for — so it can't decide whether 80 ms of cross-region write latency is acceptable or whether losing 1 second of writes on failover is.
L6: Starts with intent and numbers. "Are we doing this for latency to far-away users, to survive a region outage, or because EU data has to stay in the EU? They pull in different directions. For availability I need an RTO and RPO — say 5 minutes and under 1 second. That number decides between async replication with home regions and synchronous global consensus."
L7: Starts with whether to do it at all. "Most of our last ten incidents were bad deploys and dependency failures, not region loss. Cells inside one region would have prevented eight of them at a third of the cost. If the driver is a contractual 99.99% or a regulator, active-active is justified — and then I want the budget and the team before the architecture."
Why "data" separates levels
L5: Picks one database mode for everything — usually multi-master with LWW, because it sounds most available. That silently converts every concurrent update into a possible lost write.
L6: Splits by data class. Ledger and uniqueness constraints (usernames, inventory) need a single writer or synchronous consensus. User-owned data (profile, cart, orders) gets a home region. Counters, likes and presence are commutative — CRDTs or per-region shards summed on read. Most of the database ends up home-region owned; conflict resolution covers a few percent.
L7: Makes the classification durable. A table's multi-region class is declared in its schema metadata, enforced by the data platform (a "global-sync" table can't be created without approval), and reviewed when data is repurposed — because a "commutative" table that gains a balance column is how LWW eats money.
Why "failover" separates levels
L5: Trusts automation: health checks fail, DNS flips, the replica is promoted. Doesn't ask whether the surviving region has capacity, whether clients honor DNS TTLs, or whether the failover tooling itself depends on the failed region.
L6: Treats failover as three problems — routing, data promotion, capacity — each with a number: routing shift in < 2 min, promotion with RPO = replication lag, survivor provisioned to absorb 100% of the failed region's load. And it's practiced: a quarterly evacuation drill with real traffic.
L7: Designs for static stability: during failover, nothing calls a control plane — no autoscaling, no new instances, no config pushes from the failed region. Capacity is already there; routing changes use data-plane mechanisms. Evacuation becomes boring because it happens monthly.
The Staff Positions#
| Position | Rationale |
|---|---|
| Home-region ownership over write-anywhere | One writer per entity removes conflicts for 90%+ of data; conflict resolution is reserved for commutative data |
| Async replication by default, sync for a named few tables | Sync costs ≥ 1 cross-region RTT (70–200 ms) per write; reserve it for ledgers and uniqueness |
| No last-writer-wins on money or legal data | LWW is silent data loss under concurrent writes; clock skew makes "last" arbitrary |
| Pre-provisioned N+1 capacity | Survivors must absorb the failed region without calling autoscaling — which is the first thing to break |
| Static stability during failover | Failover must use only data-plane actions in surviving regions; no dependency on the failed region's control plane |
| Cells within regions | Most outages are deploys and bad config, not region loss; cells cap blast radius at a fraction of a region |
| Evacuate regularly, fail back deliberately | A failover path not exercised quarterly will not work; failback is a planned re-homing, not an automatic flip |
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Latency (serve far users fast) | p99 within budget for users 5,000+ km away | Read replicas + edge caching everywhere; writes to home region near the user | Stale reads from lagging replicas; far-from-home writes pay RTT | Read-your-writes for the writing user; bounded staleness for others |
| Availability (survive region loss) | RTO minutes, RPO ≈ seconds; 99.99%+ target | Active-active compute, home-region data with async followers, N+1 capacity, rehearsed evacuation | Split brain during partition; lost writes within the lag window; survivor overload | Explicit RPO per data class; no double-ownership |
| Data residency (legal placement) | Certain data must not be stored or processed outside a jurisdiction | Region-pinned home for regulated users; no replication of regulated fields outside the boundary | Failover blocked by law; a "global" service quietly copies regulated data | Zero residency violations; audit trail per field |
🎯 Staff Move: "I'll design for availability first — survive the loss of a full region with an RTO under 5 minutes and RPO under a second for most data — because that's what forces the hard decisions on ownership and conflicts. Latency improves as a side effect of serving reads locally. Residency I'll treat as a placement constraint on the home-region map, and I'll call out where it conflicts with failover."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Active-Passive vs Active-Active | Simple single-writer with an idle standby that rots, or always-on regions with ownership complexity? |
| 2 | Synchronous vs Asynchronous Replication | RPO = 0 at 70–200 ms per write, or millisecond writes with a lag-sized loss window? |
| 3 | Write-Anywhere vs Home-Region Ownership | Lowest write latency with conflicts, or one writer per entity with cross-region writes for travelers? |
| 4 | DNS vs Anycast / Global Load Balancer | Cheap and universal but slow to converge, or fast and precise but vendor-coupled? |
| 5 | Global Shared Services vs Regional Independence | One global auth/config/control plane (simple, single blast radius) or everything regional (independent, harder to keep consistent)? |
In the Wild: Real Production Systems#
Why this section belongs here: These are public, battle-tested designs. Citing them shows you know what active-active costs in practice.
Netflix — Active-Active Across AWS Regions with Routine Evacuation#
Netflix runs its streaming control plane active-active across multiple AWS regions, with Cassandra replicating data across regions and the Zuul edge tier able to steer traffic between regions. It practices region evacuation through "Chaos Kong" exercises, and its public "Project Nimble" write-up describes reducing evacuation time from close to an hour to under ten minutes by keeping capacity pre-scaled in the receiving regions rather than scaling up during the failover.
Staff insight: The breakthrough wasn't a database feature; it was capacity. Evacuation was slow because survivors had to scale up mid-incident. Pre-provisioned headroom turned failover from a scramble into a routing change — the static-stability principle in practice.
Facebook TAO — Leader Region per Shard, Followers Everywhere#
Facebook's TAO (USENIX ATC 2013) serves the social graph from caches in many regions. Each shard has a leader region holding the primary database; follower regions serve reads from their own caches and replicas, and forward writes to the leader. Writes are async-replicated to followers, giving read-after-write for the writer via the leader path and eventual consistency for everyone else.
Staff insight: This is home-region ownership at planet scale. Reads are local everywhere; writes have exactly one owner per shard; conflicts are avoided rather than resolved. When an interviewer says "active-active", this is the shape most real systems actually have.
AWS — Cell-Based Architecture and Shuffle Sharding#
AWS has publicly described (Builders' Library, re:Invent talks, Well-Architected guidance) building services as many independent cells, each a complete stack serving a subset of customers, with a thin routing layer mapping customers to cells. Combined with shuffle sharding — assigning each customer a random subset of resources — a bad deploy or poison request affects one cell's customers, not the region.
Staff insight: Cells move the blast-radius unit below the region. Most outages are self-inflicted; a region-level active-active design does nothing for a bad config pushed to every region at once. Cells plus staggered deploys do.
DynamoDB Global Tables / CockroachDB / Spanner — Three Database Answers#
DynamoDB Global Tables replicate items across regions with multi-active writes and last-writer-wins reconciliation (typically sub-second lag). CockroachDB offers REGIONAL BY ROW tables where each row has a home region and is written there with low latency. Google Spanner commits writes on a cross-region Paxos quorum with TrueTime, giving external consistency at the cost of cross-region commit latency.
Staff insight: The three map to the three strategies — write-anywhere with LWW, home-region ownership, and synchronous global consensus. Naming which one fits each data class is the Staff answer; picking one for the whole company is not.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Active-active with multi-master" | "Two regions update the same account balance concurrently. What's the result?" | Conflict awareness; LWW = loss |
| "Async replication" | "us-east dies with 2 seconds of lag. What happened to those writes?" | RPO honesty; reconciliation plan |
| "DNS failover" | "TTL is 60 s. How long until 99% of clients move?" | Real-world DNS behavior |
| "Home region per user" | "User flies Tokyo → London. What's their write latency?" | Ownership tradeoffs; re-homing |
| "Each region at 50% capacity" | "One region dies at peak. Does the survivor hold?" | N+1 math |
| "Automated failover" | "Network blip for 40 s. Did you just fail over? Who decided?" | Flapping, split brain, human-in-loop |
| "EU data stays in EU" | "EU region is down. Do EU users get service from the US?" | Residency vs availability conflict |
System Architecture Overview#
Reading the diagram: Traffic lands in the nearest healthy region. The home-region router looks up the user's (or tenant's) home region from a locally replicated ownership map and either serves the write locally or forwards it to the home region. Reads are served locally from leader or follower data. A small set of tables — ledger, uniqueness — use synchronous cross-region consensus and pay ~80 ms per write. Cells inside each region cap the blast radius of deploys. The ownership map is read locally everywhere, so routing still works when any one region — including the one that "owns" the map's writes — is down.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This | The L7 Answer — Say This |
|---|---|---|---|
| Why | "For high availability" | "Which intent — latency, availability, residency? I'll design for availability with RTO 5 min, RPO < 1 s." | "Is region loss our top outage cause? If not, cells in one region buy more availability per dollar." |
| Data | "Multi-master database" | "Home-region ownership for most data, sync consensus for the ledger, CRDTs for counters." | "Multi-region class is a required schema attribute; LWW on money is banned by standard." |
| Conflicts | "Last-writer-wins" | "Avoid them with ownership; LWW only where losing a concurrent write is acceptable." | "Every LWW table has a named owner who signed for silent loss." |
| Failover | "DNS fails over" | "Routing + promotion + capacity. Survivors pre-provisioned to N+1. Drill quarterly." | "Static stability: zero control-plane calls during evacuation. Evacuate monthly." |
| Failback | "Flip DNS back" | "Reconcile the lag window first, then re-home gradually with throttling." | "Failback has its own runbook and owner; it causes as many incidents as failover." |
| Cost | — | "N/(N−1) capacity: 2× for two regions, 1.5× for three." | "Priced against expected regional-outage cost; three regions is the efficiency sweet spot." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Speed of light in fiber | ~200 km/ms (≈ 5 µs/km) | The floor no architecture beats |
| US East ↔ US West RTT | ~60–75 ms | Sync replication across the US costs this per write |
| US East ↔ EU West RTT | ~70–90 ms | NY–London theoretical minimum ~56 ms |
| US ↔ Southeast Asia RTT | ~180–230 ms | Why APAC users need a local home region |
| Async replication lag (healthy) | ~100 ms – 1 s | Your RPO in the normal case |
| Async replication lag (under stress) | seconds to minutes | Your RPO when it actually matters |
| Capacity multiplier N/(N−1) | 2 regions = 2×, 3 = 1.5×, 4 = 1.33× | Why three regions is the common sweet spot |
| DNS failover long tail | TTL + 5–15 min for stragglers | Resolvers and clients cache beyond TTL |
| Anycast / global LB shift | ~30 s – 2 min | Health checks + BGP / LB convergence |
| Inter-region data transfer | ~$0.01–0.02 per GB (major clouds, intra-continent) | 50 TB/day replication ≈ $15–30K/month |
| Evacuation target | < 10 min (Netflix public figure) | A rehearsed, pre-scaled evacuation benchmark |
| Typical tier-1 RTO / RPO | 5–15 min / < 1–5 s | What most businesses actually sign |
Interview Walkthrough
The most common mistake: Candidates spend 20 minutes drawing two identical regions and a replication arrow, then get asked "two users update the same row in two regions — what happens?" with 10 minutes left. Compress topology to under 10 minutes. Spend the rest on data classification, conflicts, failover mechanics, capacity and who decides to evacuate.
Phase 1: Requirements & Framing (2–3 minutes)#
Functional scope in 30 seconds:
"We have a user-facing product — say an e-commerce platform with accounts, carts, orders and payments — running in one region today. We want it to run in at least two regions concurrently."
Then intent and numbers:
"Why are we going multi-region? If it's latency for users in Europe and Asia, reads near the user plus a CDN gets us most of it. If it's availability, I need the RTO and RPO the business signed. If it's residency, some data can't leave a jurisdiction, which constrains failover. I'll assume availability is the driver: survive the loss of a full region with RTO under 5 minutes and RPO under 1 second for most data, zero for payments."
Then constraints that shape everything:
"Physics: US-East to EU-West is ~80 ms round trip. Any write that must be acknowledged by both regions pays that. So I'll keep synchronous cross-region writes to the few tables that need RPO zero, and give everything else an owner region."
🎯 Staff Move: Pinning RTO/RPO in the first two minutes turns "multi-region" from an aspiration into an engineering target. Every later decision — sync vs async, capacity headroom, routing mechanism — is then justified against a number.
Phase 2: Core Entities & API (1–2 minutes)#
| Entity | Fields | Note |
|---|---|---|
| Region | id, status (ACTIVE / DRAINING / EVACUATED), capacity | Status is data-plane state, readable without a control plane |
| Cell | id, region, assigned partitions | Unit of deploy and blast radius |
| Ownership record | partition_id → home_region, epoch | Epoch increments on every re-home; fences stale owners |
| Data class | table → GLOBAL_SYNC / HOME_REGION / COMMUTATIVE / REGION_PINNED | Declared per table |
Internal routing contract:
Route(request):
partition = hash(user_id or tenant_id) mod P # P = 4096
home, epoch = ownership_map.get(partition) # local replica, cached
if request.is_write and home != local_region:
forward to home region (+1 cross-region RTT)
else:
serve locally (reads may be follower reads; add read-your-writes token)
"Two contract details matter. The ownership map is replicated to every region and read locally, so routing never depends on a remote region. And writes carry the epoch; the home region's data store rejects writes with a stale epoch, so after a re-home an old region can't keep writing."
Phase 3: High-Level Architecture (≤5 minutes)#
"Anycast global load balancer sends users to the nearest healthy region. Each region has a router and a handful of cells — full stacks serving a slice of partitions. Every partition of user data has a home region; the home holds the leader, other regions hold async followers. Writes go to the home, reads are local. Payments and uniqueness live in a synchronously replicated store across three regions — a quorum tolerates one region loss with RPO zero, at ~80 ms per write. That's the topology. Details on replication internals are in the data-store design; I want to spend the time on ownership, conflicts and failover."
Phase 4: Transition to Depth (1 minute)#
"The boxes are the easy part. Four places where this design succeeds or fails: how we classify data so most of it never conflicts; what happens to the async lag window when a region dies; how we move traffic and ownership without split brain; and whether the surviving regions can carry the load. I'd start with data classification because it determines how hard the others are."
🎯 Staff Move: Announcing that "most data never conflicts" is a strong opener — it signals you're going to avoid the hard problem before you solve it.
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → pick a position → quantify the cost → name who absorbs it.
Deep Dive 1: Data classification and conflicts (7–8 min)
| Data Class | Examples | Replication | Conflict Strategy | Write Latency |
|---|---|---|---|---|
| GLOBAL_SYNC | Ledger entries, username uniqueness, inventory for scarce items | Synchronous quorum across 3 regions | None — consensus prevents conflicts | +70–150 ms |
| HOME_REGION | Profile, cart, orders, settings | Async to followers | None — single writer per partition | Local for home users; +1 RTT for travelers |
| COMMUTATIVE | Like counts, view counters, presence, tags | Async multi-writer | CRDT merge (G-Counter, OR-Set) | Local everywhere |
| REGION_PINNED | EU health records, India payment data | No replication out of jurisdiction | None — single region | Local in jurisdiction |
| CACHE / DERIVED | Recommendations, search index | Rebuilt per region | Regenerate | Local |
"In a typical product, 80–90% of tables are HOME_REGION. Conflict resolution only exists for the COMMUTATIVE class, where merges are mathematically safe. I explicitly don't use last-writer-wins on anything with money or legal meaning — two concurrent updates to a balance under LWW means one update silently disappears, and with clock skew 'last' isn't even well defined."
🎯 Staff Move: "Conflict resolution is a sign the ownership model failed. I want it on two or three tables, not two hundred."
Deep Dive 2: The replication lag window and RPO (5 min)
"With async replication, every region failure loses the writes that hadn't replicated. If lag p99 is 800 ms and we're doing 20K writes/s in that region, that's ~16K writes at risk. Three things make that acceptable:"
- "We measure it:
replication_lag_secondsper partition, alert at 5 s, page at 30 s — because lag under stress is what defines our real RPO." - "We don't throw them away: when the failed region comes back, its unreplicated writes are extracted from its log and reconciled — replayed if they don't conflict with post-failover writes, queued for human review if they do."
- "Payments aren't in that window at all — they're GLOBAL_SYNC."
"The honest sentence for the design doc: 'For HOME_REGION data, RPO equals replication lag at the moment of failure — typically under 1 s, worst case observed 40 s during a replication backlog. Lost writes are recovered on failback where possible.' Product and legal sign that sentence."
Deep Dive 3: Failover — routing, promotion, fencing (6–7 min)
"Failover is three separate actions with three owners:"
| Step | Mechanism | Time | Owner |
|---|---|---|---|
| 1. Stop sending traffic | Global LB marks region unhealthy (or operator drains) | 30 s – 2 min | Traffic platform |
| 2. Move ownership | Ownership map: partitions of failed region → survivors, epoch + 1 | < 1 min | Data platform |
| 3. Promote followers | Followers for those partitions become leaders; reject epoch−1 writes | < 1 min | Data platform |
| 4. Absorb load | Pre-provisioned capacity; no scale-up calls | Immediate | Capacity owner |
"Step 2 is where split brain lives. If the 'failed' region is actually alive but partitioned, its local writers still think they own those partitions. The epoch is the fence: any write from the old home carries epoch N, the new leaders only accept N+1. Combined with the old region's routers noticing they've been drained, that bounds the damage to in-flight requests."
"Who decides? Automatic routing shift on hard health signals — 5xx > 25% for 60 s across multiple cells. Ownership moves are human-approved by the incident commander, because a false positive here is expensive: moving 2,000 partitions and then moving them back costs two lag windows."
Deep Dive 4: Capacity (4–5 min)
"If each of two regions runs at 60% utilization at peak, losing one puts 120% on the survivor — the failover causes the second outage. The rule is N/(N−1): with two regions each must handle 100% of global peak, so we pay 2×. With three regions each handles 50%, total 1.5×. That's why I'd rather run three regions than two."
"And capacity must be there — not 'autoscaling will add it'. During a regional event, every customer of the cloud provider is scaling into the same surviving regions, and the control plane you'd call may be degraded. Netflix's public evacuation work made exactly this point: pre-scaling the receivers is what got evacuation under ten minutes."
Deep Dive 5: Global dependencies (3–4 min)
"The most common reason multi-region failover fails is a dependency nobody listed: auth tokens validated against a single-region service, a feature-flag service in us-east, a secrets manager, a CI/CD system you need to push the 'failover' config, the cloud provider's own control-plane APIs. My rule: list every runtime dependency, mark each as regional or global, and for every global one either make it regional or make it statically stable — readable from a local replica with a long TTL, so its outage doesn't block requests."
Phase 6: Wrap-Up (2–3 minutes)#
"Summary: active-active stateless tier, home-region ownership for most data with async followers, synchronous consensus only for the ledger and uniqueness, CRDTs for commutative counters. Failover is routing, ownership move with epoch fencing, and pre-provisioned capacity — rehearsed quarterly. What I'd build next: cells inside each region, because they protect against the far more common failure, a bad deploy."
The organizational closer:
"The hardest part long-term isn't replication — it's keeping every team regionally independent. One team adding a call to a single-region service breaks evacuation for everyone. That needs a standard, a dependency audit in design review, and a drill that would catch it."
🎯 Staff Move: Close on the drill and the dependency audit. Senior candidates close on "and we'd add a third region".
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| Topology marathon | 15 min on VPC peering, transit gateways, GeoDNS records | One diagram in 5 min, then data classes |
| One replication mode | "Multi-master for everything" | Five data classes with different strategies |
| Conflicts on demand | Waits for "what about conflicts?" | Volunteers ownership as conflict avoidance |
| Capacity ignored | Two regions at 70% each | N/(N−1) math in one sentence |
| No failback | Ends at "traffic fails over" | Names the lag-window reconciliation and throttled re-homing |
| No numbers | "Low latency, high availability" | "80 ms RTT, RPO < 1 s, RTO 5 min, 1.5× cost for 3 regions" |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Multi-region is where architecture meets physics, money and law at once. It can't be answered by naming components, because every component works in the happy path. The test is whether a candidate sees that the real design artifacts are a data classification, an ownership map, a capacity plan and an evacuation runbook — and that each has a different owner.
1.2 The L5 → L6 → L7 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"It's 2 PM on your busiest day. One region's error rate is 30% and climbing, but it isn't down. Walk me through the next ten minutes: who decides to evacuate, what moves first, what data is at risk, and how do you know the other region can take it?"
This single question tests routing, ownership, RPO, capacity and decision rights. A candidate who answers only "DNS fails over" has not operated a multi-region system.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Latency → read local, write home, cache everything
- Constraint: p99 page load for far users within budget (e.g., < 300 ms from Sydney)
- Strategy: CDN for static and cacheable API responses; read replicas per region; user's home region chosen near them at signup
- Failure mode: stale reads, cross-region writes for travelers
- Who pays: users who travel or whose home is far (one RTT per write); product (explaining "my change didn't show up")
- Often solvable without active-active writes — see CDN & Edge Caching
Availability → survive a region, rehearse it
- Constraint: RTO 5–15 min, RPO < 1–5 s, 99.99% annual (≈ 52 min/year downtime budget)
- Strategy: active-active compute, home-region data with async followers, sync for zero-RPO tables, N+1 capacity, evacuation drills
- Failure mode: split brain, lag-window loss, survivor overload, hidden global dependencies
- Who pays: finance (1.5–2× infra), platform teams (replication and routing ownership), product teams (classification work)
Data residency → placement is law
- Constraint: specified data (GDPR personal data under certain contracts, financial data in some jurisdictions, health records) must be stored or processed only in-jurisdiction
- Strategy: region-pinned home for regulated users; field-level classification; replicate only non-regulated or pseudonymized data out
- Failure mode: no failover target inside the jurisdiction; global services (logging, analytics, support tools) leak regulated data
- Who pays: legal and compliance (sign-off), the business (a second in-jurisdiction region, or accepting lower availability for pinned users)
2.2 When NOT to Go Active-Active#
| Situation | Better Option | Why |
|---|---|---|
| Outages are mostly bad deploys and config | Cells + staggered deploys in one region | Region-level redundancy doesn't help when the same bad config ships everywhere |
| Users are in one geography | Multi-AZ in one region + warm standby elsewhere | Multi-AZ already survives datacenter loss; regions add ~2× cost for a rarer event |
| Only latency matters | CDN + read replicas, single write region | Reads dominate; writes can tolerate one RTT |
| RTO of hours is contractually fine | Pilot light / backup-restore | 1.1–1.3× cost instead of 1.5–2× |
| Team has never operated cross-region replication | Active-passive with monthly failover drills first | Active-active multiplies failure modes; build muscle on the simpler topology |
| Strongly coupled transactional workload across all data | Single-region primary with sync standby | Partitioning ownership is impossible without a data-model rewrite |
🎯 Staff Insight: "Multi-AZ survives the failures that happen monthly. Multi-region survives the failure that happens every few years. I'd want to see that region loss is actually in our top risks — or that a contract or regulator demands it — before paying 1.5–2× for it."
2.3 What the Interviewer Leaves Underspecified#
- The intent — latency, availability or residency
- RTO and RPO — and whether they differ by data class
- Write ratio and geography — 95% reads near users, or global collaborative writes?
- Consistency expectations — read-your-writes? Cross-user ordering?
- Regulatory scope — which data, which jurisdictions
- Team maturity — has anyone run cross-region replication in anger?
- Budget — is 2× infra acceptable?
Staff engineers surface these. Senior engineers assume them away. The two that change the design most: RPO per data class and whether residency forbids failover.
2.4 Precise Terminology#
| Term | What It Means | Common Confusion |
|---|---|---|
| Active-active | Multiple regions serve production traffic simultaneously | Doesn't imply every region writes every entity |
| Active-passive | One region serves; another stands by | "Hot standby" still has a failover event |
| Home region | The region that owns writes for a partition/user/tenant | Not necessarily where the user is right now |
| RTO | Time until service is restored | Measured from impact start, not from the decision |
| RPO | Amount of data (time) that may be lost | With async replication, RPO = lag at failure |
| Evacuation | Deliberately moving all traffic and ownership out of a region | Planned, rehearsed; differs from emergency failover only in urgency |
| Failback | Returning traffic/ownership to a recovered region | Requires reconciling the lag window first |
| Cell | Independent full stack serving a subset of users | Smaller than a region; the unit of deploy blast radius |
| Static stability | System keeps working in failure without making control-plane changes | Capacity must already exist |
| Split brain | Two regions both act as owner for the same data | Prevented by epochs/fencing, not by health checks |
🎯 Staff Insight: If the interviewer says "active-active", ask: "Active-active for serving, or for writes to the same data from multiple regions? The first is common and safe. The second is rare and needs a conflict strategy per data class."
3. The Five Fault Lines#
Each fault line is a technical choice with an ownership consequence. In Staff interviews, naming who absorbs lost writes, idle capacity and failover risk is what gets scored.
3.1 Fault Line 1: Active-Passive vs Active-Active#
The tension: Active-passive keeps a single writer and a simple data model, but the passive region is a path that's only exercised during disasters — and paths that aren't exercised rot. Active-active exercises every region every second, at the cost of ownership complexity.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Active-passive (hot standby) | One writer; no conflicts; simple mental model | Standby drifts (config, capacity, dependencies); failover is a rare, high-stakes event; ~2× cost for idle capacity | On-call during the one failover a year that doesn't work |
| Active-active, home-region writes (Staff default) | Every region proven by live traffic; local reads; no conflicts for owned data | Ownership map, forwarding, re-homing; travelers pay RTT on writes | Platform (routing + ownership); far users (write latency) |
| Active-active, write-anywhere | Lowest write latency everywhere | Conflicts on every shared entity; LWW data loss; hard-to-debug anomalies | Customers (lost updates), support (explaining them) |
L6 answer: "If the data partitions cleanly by user or tenant, active-active with home regions beats a hot standby: same cost, but both regions are proven every day. If it doesn't partition — a single global inventory, say — I'd keep a single writer and make the standby hot, with failover drilled monthly so it can't rot."
L7 answer: Asks what fraction of outages in the last two years were regional. If it's under 10%, invests first in cells and deploy safety, and treats multi-region as phase two — a decision the VP should make knowing the 1.5–2× bill.
🧭 Principal Insight: The passive region in active-passive is the least-tested production system you own. If you must run it, make it serve something real — read traffic, batch jobs, a canary slice — so it's never truly idle.
3.2 Fault Line 2: Synchronous vs Asynchronous Replication#
The tension: Synchronous replication across regions gives RPO = 0 and no conflicts, at the cost of a cross-region round trip on every write and reduced availability during partitions. Async gives local-speed writes with a loss window equal to replication lag.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Synchronous (quorum across ≥ 3 regions) | RPO = 0; linearizable; no conflicts | +70–200 ms per write; minority-side region can't write during partition | Users (write latency), finance (3-region footprint) |
| Asynchronous | Local write latency; region keeps writing during partition | RPO = lag at failure; lag spikes under load; reconciliation on failback | Data owners (lost or reconciled writes) |
| Per-class: sync for a few tables, async for the rest (Staff default) | RPO 0 where it matters, fast elsewhere | Two replication systems; cross-class transactions need care | Platform (two systems), product (classification) |
The latency math for sync writes:
Write commit latency (sync, majority of 3 regions, leader in us-east):
us-east → us-west ~65 ms RTT
us-east → eu-west ~80 ms RTT
quorum = leader + fastest follower ≈ 65 ms + local fsync ~2 ms ≈ 67 ms
User in Frankfurt writing: Frankfurt→us-east 80 ms + 67 ms ≈ 150 ms
Checkout flow with 4 sequential writes ≈ 600 ms of pure physics
"That's why I don't put the cart on synchronous replication. I put the payment ledger entry — one write per checkout — on it."
L6 answer: "Sync for the ledger, uniqueness and anything else where a lost write is a legal or financial event. Async for everything else, with lag measured per partition and the RPO written into the design doc in seconds, not adjectives."
L7 answer: Treats the sync tier as a scarce, governed resource. It has a small table allowlist, a latency SLO, and a cost per write that's charged back, so teams don't put their whole schema on it "to be safe".
🎯 Staff Move: "Sync replication is a tax on every write. I want the list of tables paying it to fit on one slide."
See Replicated Data Store for quorum mechanics and Consistency Models for the guarantees each mode gives.
3.3 Fault Line 3: Write-Anywhere vs Home-Region Ownership#
The tension: Write-anywhere gives every user local writes but turns every concurrent update into a conflict. Home-region ownership eliminates conflicts but makes writes from outside the home region pay a round trip, and requires a protocol to move ownership.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Write-anywhere + LWW | Simple; local writes | Concurrent updates silently lost; clock skew picks the "winner" | Customers (lost data), support |
| Write-anywhere + CRDTs | Mathematically convergent; local writes | Only for commutative types (counters, sets, registers with merge rules); can't express "balance ≥ 0" | Engineering (data model constraints) |
| Write-anywhere + app-level merge | Rich semantics (e.g., merge two carts) | Per-entity merge code; hard to test | Product teams (custom logic) |
| Home-region ownership (Staff default) | No conflicts; single writer; simple invariants | Travelers pay RTT on writes; re-homing protocol; ownership map is critical | Far-from-home users (latency), platform (map + re-homing) |
Choosing the home:
| Partition Key | Home Selection | Re-homing Trigger |
|---|---|---|
| User | Region nearest signup location, or residency-required | User's traffic is > 80% from another region for 30 days |
| Tenant (B2B) | Contracted region | Contract change / tenant request |
| Document / room | Region of creator, or of most active editor | Sustained majority of edits from elsewhere |
| Inventory item | Region of the warehouse | Rarely — warehouse moves |
L6 answer: "Home-region ownership for user and tenant data. Travelers pay one round trip per write, which for most products is a handful of writes per session. CRDTs for counters and presence. LWW only on fields where I can say out loud 'losing a concurrent write here is fine' — like 'last viewed item'."
L7 answer: Notices the ownership map is now the most critical piece of global state in the company. It gets its own SLO, is replicated with synchronous consensus (writes are rare), read locally everywhere, and has a change process — a mass re-homing is a production change, not a script.
❌ Common L5 Trap: "DynamoDB Global Tables handles conflicts for us." It does — with last-writer-wins. That's a resolution policy, not correctness. The interviewer will ask what happens to two concurrent increments to a balance.
3.4 Fault Line 4: DNS vs Anycast / Global Load Balancer#
The tension: DNS-based routing is universal and cheap, but failover speed is controlled by caches you don't own. Anycast and global load balancers shift traffic in seconds with fine-grained weights, at the cost of vendor coupling and a new global component.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| GeoDNS / latency-based DNS | Works with any client; cheap; no extra hop | TTL honored loosely — 5–15 min stragglers; resolver location ≠ user location | Users on stale resolvers during failover |
| Anycast (single IP, BGP) | One IP worldwide; network routes to nearest PoP; shift in seconds by withdrawing routes | BGP convergence can flap; coarse control without an L7 layer behind it | Network team (BGP operations) |
| Global L7 load balancer (anycast front + health-weighted backends) (Staff default) | Seconds-level shifts; per-region weights; draining; health-based | Vendor coupling; the GLB is itself global — its config plane is a dependency | Traffic platform; vendor lock-in |
| Client-side routing (SDK picks region) | Fast failover for apps you ship; no DNS dependency | Only for first-party clients; old app versions linger for years | Mobile/client teams (release cadence) |
DNS failover in practice (TTL = 60 s):
t=0 Region A marked unhealthy, DNS answers switch to B
t=+60s Well-behaved resolvers refresh: ~70–85% of traffic on B
t=+5min ~95%: some ISP resolvers enforce minimum TTLs
t=+15min ~99%: long-lived connections, clients that cache DNS per process
t=+hours Embedded devices / old SDKs still hitting A's IPs
L6 answer: "Anycast front door with health-weighted regional backends for the main product — failover in under 2 minutes, and I can drain a region gradually by weight. DNS-based routing as a fallback for partners who need stable hostnames. And the mobile app has a region list baked in so it can retry another region on connection failure."
L7 answer: Treats the global routing layer as the one component that is necessarily global, and engineers its own failure posture: config changes staged region by region, a "last known good" config that keeps serving if the config plane is down, and a manual DNS escape hatch rehearsed annually.
🧭 Principal Insight: You can't regionalize the front door. So make it the most boring, most statically stable component you own — and know exactly what happens when its control plane is unreachable.
3.5 Fault Line 5: Global Shared Services vs Regional Independence#
The tension: Shared global services (auth, config, feature flags, secrets, billing) are simpler to build and keep consistent. Every one of them is a way for a single region's failure — or a single bad push — to take down every region.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Global services, single home region | One source of truth; easy consistency | Home region outage = global outage; failover blocked by its own dependencies | Everyone — correlated failure |
| Fully regional services | Independent failure domains | Config/flag drift between regions; duplicated operations | Platform teams (N copies to operate) |
| Global write, regional read, statically stable (Staff default) | Rare writes to a global store; each region reads a local replica and keeps last-known-good indefinitely | Stale config during outages; needs explicit "frozen" mode | Platform (replication + freeze semantics) |
The dependency audit that every Staff candidate should describe:
| Dependency | Typical Hidden Coupling | Regionalize How |
|---|---|---|
| Auth / token validation | Central token service in one region | Signed tokens (JWT) verified locally with replicated keys |
| Feature flags / config | SaaS or service in one region | Local cache with last-known-good; no fail-closed on fetch error |
| Secrets | Single-region secrets store | Replicated secrets per region |
| Service discovery | Global registry | Per-region registry — see Service Discovery |
| CI/CD and deploy tooling | Build system in one region | Pre-built artifacts replicated; deploy agents per region |
| Cloud control plane (DNS changes, IAM, scaling APIs) | Provider's global endpoints | Pre-provisioned capacity; failover via data-plane actions only |
| Observability | Metrics/logs shipped to one region | Regional collectors; on-call can see the surviving region without the failed one |
L6 answer: "For every runtime dependency I ask: if region A disappears, does region B still serve? Global writes are fine; global reads on the request path are not. Each region keeps a local, last-known-good replica of config, flags and keys, and keeps serving if the global source is unreachable."
L7 answer: Turns this into a standard with automated enforcement: a dependency graph from tracing data flags any cross-region synchronous call on a tier-1 path, and the quarterly evacuation drill is the audit that proves it.
4. Failure Modes & Operational Reality#
4.1 Full Region Loss at Peak — Full Timeline#
t=0: eu-west loses its regional load balancer tier; EU error rate 70%
t=+20s: Global LB health checks fail for eu-west across 3 of 3 cells
t=+45s: Auto-shift: eu-west weight → 0. EU users land in us-east (+80 ms)
t=+60s: us-east at 175% of its normal load. Pre-provisioned to 200% → holding
t=+2min: Incident commander approves ownership move for EU-homed partitions
(residency-pinned partitions excluded — see 4.6)
t=+3min: Ownership map: 1,800 partitions eu-west → us-east, epoch +1
t=+4min: us-east followers promoted. replication_lag at failure: p99 0.9s
~11K EU writes unreplicated — held in eu-west's log
t=+5min: EU users can write again. Error rate 0.4%. RTO met.
t=+2h: eu-west recovered. NOT failed back yet.
t=+1 day: Lag-window writes extracted from eu-west log: 10.6K replayed,
400 conflicted with post-failover writes → review queue
t=+2 days: Throttled re-homing back to eu-west, 100 partitions/min
Detection: region_error_rate{region}, glb_backend_health{region}, replication_lag_seconds{partition} at failure time, evacuation_headroom_pct in survivors.
Blast radius: In-flight requests (seconds); lag-window writes (~11K, recoverable); EU users' latency +80 ms until failback; residency-pinned users unavailable (by policy).
Mitigation: Pre-provisioned headroom; epoch-fenced promotion; lag-window reconciliation.
Prevention: Quarterly evacuation drill at peak-like load; headroom SLO.
Owner: Traffic platform (routing), data platform (ownership + promotion), incident commander (decision), data owners (conflict review queue).
4.2 Split Brain During a Partition#
t=0: Network partition between us-east and eu-west; both regions healthy
t=+30s: Automation in us-east sees eu-west as dead; moves EU partitions to us-east
t=+30s: eu-west, unaware, keeps serving EU users locally — it still believes it owns them
t=+5min: Both regions accept writes for the same 1,800 partitions
t=+20min: Partition heals. Replication delivers conflicting histories.
Detection: ownership_epoch_conflicts_total, writes rejected for stale epoch, both regions reporting leadership for the same partition (partition_leader_count > 1).
Mitigation: Ownership moves require (a) a majority view — 3 regions or an external witness — not a single region's opinion, and (b) epoch fencing at the store: once the map says epoch 13, eu-west's leaders at epoch 12 refuse writes as soon as they see the new map; if they can't see it, they must stop writing after their ownership lease expires.
Prevention: Ownership is a lease — regions hold ownership of their partitions with a lease renewed against the global map. Losing contact with the map's quorum means stopping writes after lease expiry. This is a locking problem; see Distributed Lock Service.
Owner: Data platform.
4.3 Silent Replication Lag Growth#
Day 0: Lag p99 300 ms
Day 9: New batch job doubles write volume in us-east at 02:00 daily
Day 9-30: Lag p99 at 02:00–04:00 reaches 45 s; daytime normal
Day 31: Region failure at 03:10. RPO = 45 s of writes, not the 1 s in the design doc
Detection: replication_lag_seconds p99 by hour, not daily average; alert at 5 s for 10 min; rpo_budget_burn (fraction of time lag exceeds the signed RPO).
Mitigation: Throttle bulk writers when lag exceeds threshold; separate replication channels for bulk vs interactive traffic.
Owner: Data platform owns lag; the batch job's owner owns its write rate.
4.4 Failback Data Loss#
The most underestimated failure. After failover, the old region comes back holding writes the new owner never saw. Flipping traffic back without reconciliation either drops those writes or overwrites post-failover writes with older ones.
Rule: Failback is a separate procedure: (1) old region rejoins as follower for everything; (2) its unreplicated log segment is extracted and reconciled; (3) partitions are re-homed back gradually with epoch bumps; (4) at any point it can stop.
Detection: failback_reconciliation_pending, rehoming_rate, error rates per re-homed partition.
Owner: Data platform, with data owners signing off on conflict resolution.
4.5 Survivor Overload — The Second Outage#
If survivors aren't pre-provisioned, failover moves the outage. Autoscaling fails in exactly this scenario: the cloud provider's control plane is under load from every customer scaling at once, and capacity in the surviving region may be constrained.
Detection: evacuation_headroom_pct = (provisioned − peak current) / failed region's peak; alert if < 100% in any 2-region topology.
Mitigation: Load shedding with priority — checkout before recommendations; degraded mode for non-critical features — see Degraded Mode.
Owner: Capacity planning owns headroom; product owns shed priorities. See Capacity Planning.
4.6 Residency Blocks Failover#
EU-pinned data can't be served from us-east. If the EU has one region, EU-pinned users have no failover. Options: a second in-jurisdiction region (cost), accepting reduced availability for pinned users (signed by legal and product), or serving read-only from an in-jurisdiction backup.
Owner: Legal and product sign the posture; platform implements it.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Region loss | region_error_rate, glb_backend_health | One region's users for RTO; lag-window writes | Auto routing shift; approved ownership move; headroom | Traffic + data platform; IC |
| Split brain | partition_leader_count > 1, ownership_epoch_conflicts_total | Partitions moved during partition | Majority-based moves; epoch fencing; ownership leases | Data platform |
| Lag growth | replication_lag_seconds p99 by hour, rpo_budget_burn | Real RPO ≫ signed RPO | Throttle bulk writers; separate channels | Data platform + job owners |
| Failback loss | failback_reconciliation_pending | Lag-window writes | Follower rejoin; reconcile; gradual re-home | Data platform + data owners |
| Survivor overload | evacuation_headroom_pct, survivor saturation | All users after failover | N+1 headroom; priority shedding | Capacity planning + product |
| Hidden global dependency | Cross-region sync calls in traces; drill failures | All regions | Regionalize or make statically stable | Owning team; enforced by platform |
| DNS stragglers | Traffic still arriving at drained region | Users on stale resolvers | Anycast front; client-side region list | Traffic platform |
| Residency vs failover | Pinned users' error rate during regional event | Pinned users | In-jurisdiction second region or signed reduced SLO | Legal + product |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | Builds two regions | Names intent, RTO/RPO, and constraints from physics | Asks whether region loss is the top risk; compares against cells; prices the program |
| Data | One replication mode | Data classes with different strategies; ownership to avoid conflicts | Data class as schema metadata, enforced by the data platform |
| Conflicts | LWW | Ownership first, CRDTs for commutative, LWW only with explicit acceptance | Bans LWW on financial/legal data via standard |
| Failover | DNS flip | Routing + epoch-fenced ownership + N+1 capacity; human decision for ownership moves | Static stability; evacuation as routine; cells below regions |
| Failback | Not raised | Reconciliation and throttled re-homing | Failback owned and drilled like failover |
| Dependencies | Not raised | Audits global dependencies on the request path | Automated enforcement from tracing; independence standard |
| Cost | Not raised | N/(N−1), egress | Full program cost vs expected outage cost; 3-region sweet spot |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Pins RTO/RPO early | "RTO 5 minutes, RPO under a second for most data, zero for payments." |
| Classifies data | "80–90% of tables are home-region; sync is a one-slide allowlist." |
| Fences ownership | "Every ownership move bumps the epoch; the store rejects stale epochs." |
| Does capacity math | "Two regions means each must carry 100% of peak — 2×. Three is 1.5×." |
| Names hidden dependencies | "Auth, flags and deploy tooling must work from the surviving region." |
| Treats failback as dangerous | "Failback reconciles the lag window before re-homing anything." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "Multi-master with LWW" for everything | Silent data loss presented as availability |
| "DNS handles failover" | Ignores TTL reality, capacity, and data promotion |
| Sync replication everywhere "for consistency" | 150 ms+ per write; doesn't know the physics |
| No RPO number | Can't evaluate any replication choice |
| Regions at 70% utilization each | Failover causes the second outage |
| Automated ownership moves on single-region health checks | Builds split brain into the design |
5.4 Common False Positives#
- Deep CRDT theory ≠ multi-region design. A whiteboard of OR-Set semantics doesn't answer which tables need them — usually two.
- Cloud networking detail ≠ Staff. Transit gateways, VPC peering and BGP communities are table stakes, not judgment.
- Naming Spanner ≠ solving it. "Use Spanner" without the write-latency cost and the table allowlist is a buzzword.
- Chaos-engineering vocabulary ≠ drills. Saying "Chaos Kong" isn't the same as describing who decides to evacuate and what the receiving region's headroom is.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Intent and numbers | 0–4 min | Intent, RTO/RPO, physics |
| Entities and routing | 4–7 min | Home region, ownership map, epoch |
| Architecture | 7–12 min | Anycast front, regions, cells, data tiers |
| Data classes and conflicts | 12–20 min | Five classes; LWW only with consent |
| Failover and split brain | 20–30 min | Routing, ownership move, fencing, decision rights |
| Capacity, dependencies, failback | 30–38 min | N/(N−1); audit; reconciliation |
| Org, cost, evolution | 38–45 min | Ownership boundaries, drills, phased rollout |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Direction |
|---|---|---|
| "Two regions update the same balance" | Conflict reasoning | Ownership or sync; never LWW for money |
| "Region is at 30% errors, not down" | Decision rights under ambiguity | Drain by weight; IC approves ownership move |
| "User moves to another continent" | Re-homing | Threshold-based, epoch-fenced, async migration |
| "EU region is down" | Residency vs availability | Pre-signed posture per data class |
| "Make it cheaper" | Cost levers | 3 regions; tiering; non-critical services single-region |
| "How do you test this?" | Operational maturity | Evacuation drills with real traffic, quarterly then monthly |
6.3 What to Deliberately Skip#
Replication protocol internals (link Replicated Data Store), VPC/transit networking, cloud-vendor product comparisons beyond one sentence, CRDT proofs, and per-service deployment pipelines.
6.4 Follow-Up Questions to Expect#
- "What's your RPO, and how do you know it's true at 3 AM?"
- "How does a write from a traveling user work, and what does it cost?"
- "What stops two regions from both owning a partition?"
- "Walk me through failback."
- "Which services are still global, and what happens when they're down?"
- "How much does this cost compared to single-region?"
- "How often do you evacuate a region, and who decides?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Make our application multi-region active-active."
Staff Answer
"First, why? Latency, surviving a region outage, or data residency — they lead to different designs. I'll assume availability: RTO under 5 minutes, RPO under 1 second for most data and zero for payments. Physics sets the frame: ~80 ms between US-East and EU-West, so anything synchronous across regions pays that per write. My plan: active-active stateless tier behind an anycast load balancer; every user or tenant has a home region that owns their writes, with async followers elsewhere; a small synchronously replicated tier for the ledger and uniqueness; CRDTs for counters. Then failover — routing, epoch-fenced ownership moves, pre-provisioned capacity — and failback. I'll finish with cost and who owns what."
Why this is L6:
- Asks for intent and pins RTO/RPO before topology
- Uses physics to justify the data tiers
- Lays out a sequence of decisions, not a list of boxes
What L7 adds:
- Asks whether region loss is actually the top outage cause; proposes cells first if not
- Frames the budget: 1.5–2× infra plus a platform team, against the expected cost of regional outages
- Names the decision owner for evacuations before designing the automation
❌ Common L5 Trap
"We'll deploy the same stack in two regions, use GeoDNS to route users to the nearest one, and use a multi-master database so both regions can write."
Why this misses: Every component is plausible, but the interviewer's next question — "two users update the same record in both regions" — has no answer except LWW, and the design has no RPO, no capacity plan and no failover decision rights.
Drill 2: Concurrent Writes to the Same Entity#
Prompt: "A shared team account has a credit balance. Two admins, one in New York and one in London, spend credits at the same moment in different regions. What happens?"
Staff Answer
"With write-anywhere and LWW, one debit silently disappears — the account ends up with the balance from whichever write had the later timestamp, and clock skew between regions decides which. That's not acceptable for money. Two correct options. Home-region ownership: the account's home is, say, us-east; the London write is forwarded (+80 ms) and both debits serialize on one leader with a conditional check balance >= amount. Or, if credits are the core of the business, put balances in the synchronous tier — every debit pays a cross-region quorum, but RPO is zero and the account survives region loss without reconciliation. I'd start with home-region ownership, because the London admin pays 80 ms once per purchase, which nobody notices."
Why this is L6:
- Shows exactly how LWW loses data and why clocks make it worse
- Offers two correct strategies with costs
- Picks one based on how often the cost is paid
What L7 adds:
- Makes "no LWW on balances" an org rule, enforced by the data platform's table classification
- Notes that B2B accounts are a natural ownership key and aligns home region with the contract's region
Drill 3: Home-Region Routing — Make It Concrete#
Prompt: "You said users have a home region. A user homed in EU spends three months in Singapore. Walk me through their reads and writes."
Staff Answer
"Day one in Singapore: the anycast front lands them in the APAC region. Reads are served from APAC's follower replica — for their own data we need read-your-writes, so each write returns a replication position token the client sends back; if the APAC follower is behind that position, we wait up to 50 ms or forward the read to the EU. Writes are forwarded to EU: Singapore↔Frankfurt is ~160–200 ms RTT, so each write takes ~200 ms. For a typical session with a handful of writes, acceptable. After 30 days with > 80% of traffic from APAC, the re-homing job migrates them: copy state to APAC, briefly pause writes for that partition (< 1 s), bump the epoch, switch the ownership map, resume. If the user's data is EU-residency-pinned, we never re-home; they keep paying the RTT."
Why this is L6:
- Handles read-your-writes explicitly with a position token
- Quantifies traveler cost and sets a re-homing threshold
- Respects residency as a constraint on re-homing
What L7 adds:
- Partitions at a granularity (user-group or tenant) where re-homing is cheap; per-user re-homing at 100M users is a constant background migration with its own SLO
- Tracks
cross_region_write_ratioas a product-quality metric — rising values mean the home-assignment policy is stale
Drill 4: The Region Is Sick, Not Dead#
Prompt: "us-east error rate is 30% and climbing. It isn't down. What do you do in the next 10 minutes?"
Staff Answer
"Gray failure is the common case. Minute 0–2: confirm it's regional — are errors spread across all us-east cells or one? If one cell, drain that cell, not the region. If regional, shift traffic first: drop us-east's weight in the global LB to 50%, then 0%, watching the survivors' saturation. Traffic shift is cheap and reversible. Minute 2–5: decide on ownership. If us-east's data tier is healthy and the problem is the serving tier, I don't move ownership — writes for us-east-homed users forward from the survivors to us-east's healthy store. If the data tier is sick, the incident commander approves an ownership move with epoch bump, accepting the lag-window RPO. Minute 5–10: confirm survivors are within headroom, start shedding non-critical features if not, and post status. The decision rule is in the runbook: traffic shifts are automatic or on-call; ownership moves need the IC."
Why this is L6:
- Separates traffic shift (cheap, reversible) from ownership move (expensive, RPO-bearing)
- Checks cell vs region scope before evacuating
- Names decision rights
What L7 adds:
- Gray-failure detection uses multiple vantage points (synthetic probes from other regions), not just self-reported health
- Evacuation is routine enough — monthly — that executing it at 30% errors is low-risk and fast, so the bias is to evacuate early
Drill 5: The Global Hot Entity#
Prompt: "A live event page gets comments and reactions from users in every region — 200K writes/s at peak, all to one entity. Home-region ownership puts all those writes on one region."
Staff Answer
"This entity doesn't fit home-region ownership, and it doesn't need to. Reactions are counters — commutative — so each region keeps its own counter shard and the global count is the sum, replicated async; a G-Counter CRDT per region converges without conflicts. Comments are an append-only set: each region appends locally with a globally unique ID and a timestamp; replication merges by union; display order is by timestamp with ties broken by region ID. Readers in each region see their own region's comments immediately and others within ~1 s. Nothing here needs single ownership because no invariant spans concurrent writes. Moderation deletes are tombstones — an OR-Set — so a delete in one region wins over a concurrent replication of the same comment."
Why this is L6:
- Recognizes the data class (commutative) and applies the right tool
- Avoids forcing all writes through one region
- Handles the hard case (deletes) explicitly
What L7 adds:
- Offers CRDT-backed counters and append-only feeds as platform primitives so teams don't hand-roll merge logic
- Keeps the commutative class narrow and documented, so it doesn't spread to data with invariants
Drill 6: Multi-Tenant Cells#
Prompt: "We're B2B with 20,000 tenants. The largest 10 tenants generate 40% of traffic. How do you lay out regions and cells?"
Staff Answer
"Tenant is the ownership key. Each region runs several cells — say 8 — each a full stack with its own data partitions. A tenant is assigned a home region (contract or proximity) and a home cell. Small tenants are packed into shared cells. The 10 large tenants each get a dedicated cell or are spread so no cell holds more than ~15% of regional traffic. Each cell has a designated failover cell in another region, sized to absorb it — so evacuation is cell-by-cell, not region-wide. A bad deploy goes to one cell first and bakes for an hour; it can hurt at most ~1/8 of a region's tenants. Cell assignment lives in the ownership map, and moving a tenant between cells is the same epoch-fenced re-homing as moving between regions."
Why this is L6:
- Uses tenant as the natural ownership key
- Isolates large tenants to cap blast radius
- Unifies cell and region moves into one mechanism
What L7 adds:
- Prices dedicated cells into enterprise contracts — isolation is a product tier, not a favor
- Sets the cell size policy (max tenants, max % of traffic) as a platform standard tied to the blast-radius budget
Drill 7: Build vs Buy#
Prompt: "Should we use Spanner, CockroachDB, DynamoDB Global Tables, or build replication on our PostgreSQL?"
Staff Answer
"It depends on the data class, and different classes can use different answers. For the synchronous tier — ledger, uniqueness — a managed globally consistent database (Spanner, or CockroachDB with a multi-region config) is worth paying for; building cross-region consensus on PostgreSQL is a multi-year project. For home-region data, CockroachDB's regional-by-row tables, or PostgreSQL per region with logical replication to followers and our own ownership map, both work; the latter keeps our existing PostgreSQL expertise but means we own failover and reconciliation tooling — 2–4 engineers ongoing. DynamoDB Global Tables is excellent for data where last-writer-wins is acceptable, like session state or preferences, and wrong for balances. I'd buy the sync tier, and choose for the home-region tier based on how much PostgreSQL we already run."
Why this is L6:
- Maps products to data classes rather than picking one
- Prices the build option in headcount
- Names where each product's conflict policy is wrong
What L7 adds:
- Weighs vendor lock-in: a globally consistent managed DB is the hardest dependency to leave — a one-way door that needs an exit analysis
- Limits the number of distinct replication technologies the org operates to two
See Database Selection and Build vs Buy.
Drill 8: Migrating from Active-Passive Without an Outage#
Prompt: "We're active-passive today with us-east primary and us-west standby. Move to active-active."
Staff Answer
"Phases, each reversible. (1) Serve reads from us-west — the standby becomes useful and gets real traffic, which finds the drift we didn't know about. (2) Build the ownership map with every partition homed in us-east; route all writes through it. Nothing changes behaviorally, but every write now goes through the routing layer. (3) Re-home 1% of partitions — internal employee accounts first — to us-west, with epoch fencing. Watch cross_region_write_ratio, conflict and epoch-rejection metrics for a week. (4) Ramp: 10%, 25%, 50%, by tenant cohort. (5) First full evacuation drill of us-east at low traffic. The data classification work happens in parallel: every table must be classified before its partitions can be re-homed, so unclassified tables keep us-east as home until their owner signs."
Why this is L6:
- Each step is observable and reversible
- Starts with reads to flush out drift
- Gates re-homing on data classification by owners
What L7 adds:
- Publishes a migration SLO and a date after which new services must be region-independent by default
- Funds the classification effort as part of the program, not as unplanned work for 30 product teams
Drill 9: Cost#
Prompt: "The CFO asks why infrastructure costs went up 70% after going multi-region."
Staff Answer
"Three components. Capacity: with three regions, each must absorb half the global peak if one fails — 1.5× the single-region footprint; that's most of the increase. Replication: cross-region transfer at ~$0.02/GB — at 30 TB/day that's ~$18K/month — plus the synchronous tier's 3-region footprint. Duplication of fixed costs: per-region observability, NAT, load balancers, ~5–10%. What it buys: our contract's 99.99% SLA, which we couldn't meet with a single region, and protection against a regional event that last time cost us ~$2M in a 4-hour outage. Levers if we need to cut: keep non-critical services (analytics, internal tools) single-region with warm standby, and move from two regions to three to drop the headroom multiplier from 2× to 1.5×."
Why this is L6:
- Breaks cost into capacity, transfer and fixed duplication with numbers
- Ties cost to a contractual requirement and an incident cost
- Offers concrete levers
What L7 adds:
- Tiers services by availability requirement so only tier-1 pays the multi-region premium
- Reports cost per region per tier quarterly, so the premium stays a deliberate choice
Drill 10: The Evacuation Drill#
Prompt: "How would you run a region evacuation drill in production?"
Staff Answer
"Announce it; run it during business hours, first at low traffic, later at peak-like load. Pre-checks: survivors' headroom > 110% of the evacuated region's current load; replication lag < 1 s; no in-progress deploys. Step 1: shift traffic by weight, 25% steps every 5 minutes, watching survivor latency and error budgets — abort criteria written down. Step 2: move ownership for the region's partitions with epoch bumps, measuring time-to-writable per partition. Step 3: hold for 1 hour. Step 4: failback via the real failback procedure — follower rejoin, reconciliation (should be empty), gradual re-home. Outputs: evacuation time, max survivor saturation, every dependency that broke, and a list of owners with action items. Cadence: quarterly to start, monthly once it's boring."
Why this is L6:
- Pre-checks and abort criteria defined up front
- Exercises failback as well as failover
- Produces measurable outputs and owners
What L7 adds:
- Makes evacuation readiness a reported SLO (time-to-evacuate, last successful drill date) for every tier-1 service
- Treats a failed drill as a success of the program and protects the teams who surfaced the failure
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Incident — Region Degradation on the Biggest Sales Day#
Context: Black Friday, 11:00 ET. us-east's managed cache tier is degraded; p99 latency in us-east is 4 s and error rate 18%. The design is two regions, active-active, each running at 65% utilization at peak. On-call asks whether to evacuate us-east.
Questions to Surface First:
- Can us-west absorb us-east's load? (65% + 65% = 130% — no.)
- Is the degradation in one cell or region-wide?
- What can we shed — recommendations, personalization, reviews — to cut load per request?
- What's the per-minute revenue impact of 18% errors vs a full us-west overload?
Typical L5 Approach: Evacuates us-east because it's unhealthy. us-west hits 130% of its capacity, latency collapses there too, and the incident becomes global. Autoscaling requests in us-west are slow because every retailer is scaling at the same moment.
Staff Approach: Does the capacity math first. Shifts only what us-west can carry (~35% of us-east's traffic), turns on degraded mode in both regions — disabling recommendations and personalization cuts per-request cost ~40% — then shifts more as headroom appears. Works the cache problem in us-east in parallel. Accepts 5–8% errors in us-east for 20 minutes rather than 100% degradation everywhere.
Principal Approach: Treats "two regions at 65%" as the root cause: the architecture promised failover it couldn't deliver at peak. Moves to three regions (1.5× instead of 2× for real N+1), makes
evacuation_headroom_pct ≥ 100%a pre-peak release gate, and creates a peak-readiness review that includes an evacuation drill at the projected peak load two weeks before the event.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Check evacuation_headroom_pct for us-west; determine cell vs region scope |
| Triage | Is it one cell? Drain it. Region-wide? Partial shift only |
| Quick fix | Degraded mode both regions; shift us-east weight by 15% steps while us-west stays < 85% CPU |
| Guardrails | Priority shedding: checkout and payment paths protected; browse degraded |
| Post-mortem | Why was a 2-region design at 65% utilization at peak? Who signed off on headroom? |
Metrics to Watch: evacuation_headroom_pct{region}, region_error_rate, checkout_success_rate, cpu_utilization{region}, shed_requests_total{feature}.
Organizational Follow-up: Headroom SLO owned by capacity planning; peak-event readiness includes a load-level evacuation drill; product agrees in advance on shed priorities.
Ownership Question: "Who decided the regions could run at 65%?" Staff answer: Usually nobody explicitly — it happened through cost optimization. Capacity planning should own the headroom number, finance should see its cost, and the VP who owns the SLA signs the tradeoff. If nobody signed it, the post-mortem's first action is to make someone sign it.
Key Takeaway: "Failover to a region without headroom isn't failover — it's spreading the outage."
What clears the Staff bar:
- Capacity math before evacuation
- Partial shifts plus degraded mode
- Names headroom as an owned, signed number
Deep Dive 2: Silent Failure — Last-Writer-Wins Ate Six Months of Updates#
Context: Support notices customers saying "my address change didn't stick". Investigation shows the customer-profile table is on a multi-master store with last-writer-wins. A background job in us-west re-syncs profiles from the CRM nightly, writing every profile with a fresh timestamp. Any customer edit in us-east within the replication window before the job's write is overwritten. It's been happening for six months.
Questions to Surface First:
- How many writes were overwritten? Can we reconstruct them from logs or CDC streams?
- Which other tables use LWW, and which have background writers?
- Do any affected fields have legal or financial meaning (billing address → tax)?
- Why didn't any metric show it?
Typical L5 Approach: Changes the batch job to only write changed fields and adds a check to skip recently updated profiles. Fixes this instance.
Staff Approach: Reclassifies the profile table as HOME_REGION: one writer per user, and the batch job writes through the router to each user's home region, where the conditional write
WHERE updated_at < :crm_snapshot_timeprevents stale overwrites. Recovers overwritten values from the CDC stream and notifies affected users whose billing addresses changed. Addslww_overwrites_total— counting replicated writes that replaced a newer-in-wall-clock local write within 5 s — on every remaining LWW table.
Principal Approach: Treats LWW as a governance failure: the table was put on multi-master without anyone signing that concurrent writes could be lost. Introduces the data-class attribute on every table, bans LWW on fields with financial or legal meaning, and requires a named owner who accepts silent loss for any LWW table. Audits all LWW tables within a quarter.
Staff Approach — Full Reasoning
| Dimension | Staff Answer |
|---|---|
| Root cause | LWW on a table with two writers (users and a batch job) in different regions |
| Immediate action | Pause the batch job; reconstruct overwritten values from CDC |
| System fix | HOME_REGION ownership; batch writes routed through home; conditional writes |
| Detection fix | lww_overwrites_total, per-table writer inventory |
| Broader question | Every LWW table and every background writer |
Metrics to Watch: lww_overwrites_total{table}, profile_update_reverted_total (user edits followed by a different value within 24 h), support tickets tagged "change didn't save".
Organizational Follow-up: Data-class attribute required; LWW sign-off; background-writer registry.
Ownership Question: "The batch job team didn't know the table was multi-master. Whose incident is it?" Staff answer: The platform made a lossy mode available without requiring the data owner to accept its semantics. The batch team owns the job; the data platform owns making LWW an explicit, signed choice. Both have action items; neither gets blamed.
Key Takeaway: "Last-writer-wins doesn't fail — it succeeds at losing data. Without a metric for overwrites, you'll find out from customers."
What clears the Staff bar:
- Moves the table to single ownership instead of patching the job
- Builds detection for a failure that's silent by construction
- Recovers data and notifies affected customers
Deep Dive 3: Large Customer Onboarding — EU Enterprise Wants Residency and 99.99%#
Context: Sales is closing a large EU bank. Contract terms: all customer data stored and processed in the EU, 99.99% availability, RPO 0 for transactions. You have one EU region.
Questions to Surface First:
- Does "processed in the EU" include logs, support tooling, analytics and backups?
- With one EU region, what happens to them during an EU regional outage?
- Is RPO 0 for all data or for transactions only?
- What's the deal worth versus the cost of a second EU region?
Typical L5 Approach: Pins the tenant to eu-west, enables replication to us-east for availability. That violates residency the moment replication starts.
Staff Approach: States the conflict: one EU region cannot deliver both residency and 99.99% with region-loss tolerance. Options: (a) add a second EU region (e.g., Frankfurt + Ireland) — tenant's home in one, sync tier for transactions across the two plus a third in-EU location for quorum, async followers for the rest; (b) keep one region and negotiate the SLA to 99.95% with a regional-loss exclusion. Audits every pipeline that touches tenant data — logs, metrics labels, support tools, backups — for out-of-EU flows.
Principal Approach: Turns the deal into a product tier: "EU sovereign" with a defined architecture (two EU regions, EU-only operations tooling, EU-resident on-call access controls), priced to cover the second region. Works with legal and sales so future residency commitments come from a standard menu the platform can deliver. Decides whether the second EU region becomes the general EU footprint, amortizing it across all EU customers.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Scope | Define residency per data type with legal: primary data, backups, logs, telemetry, support access |
| Architecture | Second EU region; sync tier with in-EU quorum; tenant cell dedicated |
| Leak audit | Trace every flow of tenant data; remove PII from global logs and metrics labels |
| SLA | 99.99% only once the second EU region passes an evacuation drill |
| Timeline | Contract signature does not equal go-live; residency audit is a gate |
Metrics to Watch: residency_violations_total{tenant} (data-flow scanners), cross_jurisdiction_egress_bytes, tenant availability per month.
Organizational Follow-up: Residency commitments require platform sign-off; standard residency tiers; data-flow scanning in CI.
Ownership Question: "Sales already promised 99.99% and residency. Who owns the gap?" Staff answer: The commitment requires an architecture that doesn't exist yet. Sales owns re-setting expectations, the platform owns the cost and timeline of the second region, and the executive sponsoring the deal decides whether to fund it. Engineering doesn't absorb an unfunded contractual promise silently.
Key Takeaway: "Residency and region-loss tolerance conflict inside a single-region jurisdiction. Resolve it with a second in-jurisdiction region or a signed SLA exception — never with quiet replication."
What clears the Staff bar:
- Spots the residency vs availability conflict immediately
- Audits secondary data flows (logs, support, backups)
- Refuses to absorb an unfunded commitment silently
Deep Dive 4: Post-Mortem — Failover Took Three Hours#
Context: us-east suffered a 4-hour regional event. The multi-region design promised a 5-minute RTO. Actual recovery took 3 hours. You present the post-mortem.
Questions to Surface First:
- Which step took the time — routing, ownership, capacity, or something else?
- Which dependencies did the failover procedure itself need?
- When was the last successful evacuation drill?
- Who had authority to start the failover, and when did they decide?
Typical L5 Approach: Fixes each blocker: replicate the auth service, move the deploy tool, pre-create break-glass credentials. All necessary; none prevents the next unlisted dependency.
Staff Approach: Names the root causes as systemic: (1) failover depended on control-plane services in the failed region — violating static stability; (2) decision rights were unclear, costing 35 minutes; (3) the last drill was 11 months ago. Fixes: failover executes only data-plane actions (weights in the global LB, ownership map in the replicated store); a written decision rule — the on-call IC evacuates on defined signals without executive approval; capacity pre-provisioned; drills quarterly.
Principal Approach: Presents the incident to leadership as an independence debt, not a set of bugs. Establishes a regional-independence standard for tier-1 services, a dependency inventory derived from tracing, and evacuation readiness as a reported SLO. Funds a small team to own evacuation tooling and the drill program. Sets the expectation that drills will fail at first, and that failed drills are the program working.
Staff Approach — Full Reasoning
| Section | Content |
|---|---|
| What happened | 35 min to decide, 2h25m blocked on dependencies in the failed region, capacity shortfall in the survivor |
| Why | Failover path untested for 11 months; depended on control planes; no pre-provisioned capacity |
| Immediate actions | Replicate auth signing keys; regional deploy agents; break-glass credentials per region |
| Systemic fix | Static stability for failover; written decision rule; headroom SLO; quarterly drills |
| Ownership boundary | Evacuation tooling owned by a named team; every tier-1 service owns its independence |
Metrics to Watch: time_to_evacuate_seconds (from drills), days_since_last_successful_drill, cross_region_sync_dependencies{service}, evacuation_headroom_pct.
Organizational Follow-up: Independence standard; drill calendar; decision-rights document signed by the CTO.
Ownership Question: "Who should have known the failover depended on us-east?" Staff answer: No individual could have — the dependency graph changed weekly. That's the point: only an automated dependency inventory and a regular drill can know. The accountability gap is that nobody owned the drill program.
Key Takeaway: "A failover path that hasn't run recently has dependencies you don't know about. Drill it, and make it use only data-plane actions."
What clears the Staff bar:
- Identifies static-stability violations as the root cause class
- Fixes decision rights, not just tooling
- Turns drills into the audit mechanism
Deep Dive 5: Multi-Region Expansion — From Two Regions to Three Plus APAC#
Context: You run us-east and eu-west active-active with home-region ownership. Leadership wants APAC presence (Singapore) for latency and asks whether to add a third region in the US too.
Questions to Surface First:
- What fraction of traffic is APAC, and what's the current p99 for APAC users?
- Is the sync tier's quorum currently across 2 or 3 locations? (2 can't tolerate one loss.)
- Can existing regions absorb a region loss today — what's headroom?
- Do any APAC jurisdictions require local residency?
Typical L5 Approach: Clones the stack into Singapore and extends multi-master replication to three regions. APAC users get local reads and writes; replication fan-out and conflict volume grow.
Staff Approach: Sequences it. First add a third region near the bulk of traffic (us-west) — this makes N+1 cheaper (each region from 100% to ~50% spare) and gives the sync tier a 3-location quorum that tolerates one region loss. Then add APAC as a home region for APAC users, re-homing them gradually. The sync tier stays on the three Western regions — APAC writes to the ledger pay ~180 ms, which is acceptable for one write per checkout. Lag and conflict metrics are watched per new region.
Principal Approach: Builds a repeatable region-launch process: a region is a product with a readiness checklist (capacity, dependency audit, evacuation drill passed, residency review) and a cost model per added region. Decides that region #4+ should be launched by the platform team in weeks from templates, not by a project team in quarters. Re-evaluates whether APAC needs a full region or whether edge caching plus a read-only follower meets the latency goal at a third of the cost.
Staff Approach — Full Reasoning
| Option | Tradeoffs |
|---|---|
| Add APAC only (3 regions: US, EU, APAC) | APAC latency fixed; capacity N+1 needs each region at ~50% spare; sync quorum spans continents (higher latency for all) |
| Add us-west first, then APAC | Cheaper N+1 sooner; sync quorum within ~80 ms; APAC added as home region later |
| APAC as read-only follower + CDN | Lower cost; APAC writes pay RTT; no APAC failover target |
Metrics to Watch: p99_latency{user_geo}, cross_region_write_ratio{region}, sync_tier_commit_latency, evacuation_headroom_pct per region.
Organizational Follow-up: Region-launch checklist; per-region cost model reviewed by finance; APAC residency review with legal.
Ownership Question: "Who decides where the next region goes?" Staff answer: A joint decision — product brings latency and market data, finance brings cost, legal brings residency, and the platform brings the readiness plan. The platform owns the recommendation; the business owner of the availability SLA signs.
Key Takeaway: "Adding regions is a capacity and quorum decision before it's a latency one. The third region pays for itself in headroom."
What clears the Staff bar:
- Sequences expansion by capacity and quorum benefit
- Keeps the sync tier compact
- Considers a cheaper APAC option before a full region
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Distinguish latency, availability and residency intents and explain how each shapes the design
- Pin RTO and RPO and derive replication choices from them
- Classify data into sync, home-region, commutative, region-pinned and derived classes
- Explain why last-writer-wins loses data and when it's acceptable
- Walk through failover as routing, epoch-fenced ownership move and capacity — and failback as reconciliation plus throttled re-homing
- Compute N/(N−1) capacity and price cross-region replication
- Audit global dependencies and explain static stability
- Run an evacuation drill and name who decides to evacuate
The Bar for This Question#
Mid-level (L4/E4): Knows regions, availability zones and DNS-based routing. Proposes deploying the stack twice. Doesn't reason about data conflicts or capacity.
Senior (L5/E5): Designs a credible two-region topology with replication and health-based failover. Knows sync vs async tradeoffs and names LWW as a conflict policy. Gets to split brain and capacity when prompted. Treats the problem as infrastructure.
Staff+ (L6/E6+): Starts from intent and RTO/RPO, uses physics to bound the design, and avoids conflicts through home-region ownership with epoch fencing. Classifies data, does capacity math, separates traffic shifts from ownership moves, treats failback as its own procedure, and audits global dependencies. Names decision rights and drills. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "Most Companies Don't Need Active-Active"#
| Outage Cause (typical mix) | Does Multi-Region Help? | What Helps More |
|---|---|---|
| Bad deploy / config push | Rarely — it ships to every region | Cells, staggered rollout, fast rollback |
| Dependency failure (DB, third party) | Sometimes | Circuit breakers, degraded mode |
| Capacity / traffic spike | No | Load shedding, autoscaling, headroom |
| Full region loss | Yes | Multi-region |
The Staff position: Region loss is real but rare. If you haven't built cells, staggered deploys and degraded modes, multi-region is buying insurance against the less likely fire. Do it when a contract, regulator or measured risk demands it.
Why this matters in interviews: Questioning the premise — with numbers — is the clearest seniority signal available in this question.
10.2 "Active-Active Is a Capacity Problem Before It's a Data Problem"#
Most failed failovers in public post-mortems weren't caused by replication bugs. They were caused by survivors that couldn't carry the load, or by autoscaling that couldn't add capacity in time.
The Staff position: If evacuation_headroom_pct isn't a tracked, owned number, you don't have active-active — you have two single-region deployments that hope.
10.3 "Last-Writer-Wins Is Data Loss with Good PR"#
| Property | LWW Promise | Reality |
|---|---|---|
| Convergence | All replicas agree | Yes — on one of the writes |
| "Last" | Most recent write wins | The write with the larger timestamp, subject to clock skew |
| Data loss | "Conflicts resolved" | The other write is gone, silently |
The Staff position: LWW is fine for "last viewed", "theme preference", session state. It's wrong for anything someone would be upset to lose. Every LWW table needs an owner who has said that sentence out loud.
10.4 "If You Haven't Evacuated a Region This Quarter, You Don't Have Multi-Region"#
Dependencies change weekly. Capacity drifts. Runbooks rot. The only proof the failover path works is running it.
The Staff position: Quarterly evacuation with production traffic is the minimum; monthly is the goal. Netflix publicly treats evacuation as a routine operation — that's what turns a 3-hour failover into a 10-minute one.
10.5 "Cells Beat Regions as the Unit of Blast Radius"#
A region is a very large, very expensive failure domain. A cell — a full stack serving a few percent of users — limits the far more common self-inflicted failures, and makes region evacuation a series of small cell moves.
The Staff position: Build cells inside one region first. Multi-region becomes "cells in more places" rather than a new architecture.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
A Staff engineer designs a correct active-active system. A Principal engineer sees that multi-region is not a system — it's a property every tier-1 service must keep, and it decays every time any team adds a synchronous call to a single-region dependency. The L7 questions are: Is region loss our biggest risk, or are we buying the wrong insurance? Which services must be region-independent, and who audits that? What does each additional region cost per year, and what does an hour of regional outage cost? Multi-region at org scale is a governance program — data classification, independence standards, a drill calendar, decision rights — with a platform underneath.
🧭 Principal Move: "Before we commit to active-active, I want the last two years of incidents classified by cause. If region loss is a small slice, I'd fund cells and deploy safety first and do multi-region in phase two — for tier-1 services only, with a written independence standard and a drill program. The budget conversation is 1.5× infra plus a small platform team, against the cost of our worst plausible regional outage."
The Org-Level Fault Line#
Multi-region as a platform property vs a per-team project.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each team makes its own service multi-region | Teams choose fitting designs | 30 different replication and failover approaches; evacuation requires 30 coordinated runbooks | On-call during evacuation; customers when one team's path fails |
| Central platform makes everything multi-region transparently | Consistency | Platform can't know each table's data class; hides conflicts; becomes a bottleneck | Platform team; product teams surprised by semantics |
| Platform provides primitives + standard; teams classify and own independence (L7 default) | One routing layer, one ownership map, one drill; teams own their data classes | Requires a standard, audits, and funded classification work | Platform budget; ~1–2 weeks per product team for classification |
The deciding question: can the org evacuate a region with one command? If evacuation requires coordinating many teams, the multi-region property isn't real.
Cost Model#
Assumptions: cloud on-demand pricing; single-region baseline infra cost B; inter-region transfer ~$0.02/GB; loaded engineer ~$25K/month.
| Scale | Architecture | Infra $/month | Headcount | On-Call Load |
|---|---|---|---|---|
| ~$50K/month baseline, one geography | Multi-AZ + cells + warm standby in second region | ~1.2–1.3× B (+$10–15K) | ~0.5 FTE | Monthly failover drill; shared rotation |
| ~$500K/month baseline, 2 continents | 3 regions active-active, home-region ownership, sync tier for ledger | ~1.6–1.8× B (+$300–400K incl. ~$20–40K transfer) | 4–6 FTE (routing, ownership/replication, evacuation tooling) | Dedicated multi-region rotation; quarterly evacuation |
| ~$5M/month baseline, global | 5+ regions, cells everywhere, sovereign tiers, monthly evacuation | ~1.4–1.6× B (more regions → lower N/(N−1)) | 15–25 FTE across traffic, data and resilience platforms | Evacuation is routine; game days monthly |
The line that matters to leadership: the premium shrinks as you add regions (2× for two, 1.5× for three, 1.25× for five) but the fixed platform cost doesn't. Below a few hundred thousand dollars a month of infra, the platform team often costs more than the extra capacity — another reason small companies should start with warm standby.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversal Cost | Why |
|---|---|---|---|
| Routing vendor / global LB | Two-way-ish | 1–2 quarters | Abstracted behind hostnames |
| Partition key for ownership (user vs tenant vs entity) | One-way | Years | Every table, router and re-homing tool depends on it |
| Sync-tier database product | One-way | Multi-year migration | Globally consistent stores have unique semantics and pricing |
| Conflict semantics exposed to users (e.g., LWW on shared docs) | One-way | Product redesign | Users build workflows around observed behavior |
| Residency commitments in contracts | One-way | Contract renegotiation | Legal obligations |
| Number of regions | Two-way (expensive) | Quarters | Adding is easier than removing; data must be re-homed |
| Headroom percentage | Two-way | Weeks | Budget decision |
🧭 Principal Insight: The partition key for ownership is the one-way door most teams choose by accident. Choose it in a design review with product and data owners, because it determines what can be re-homed, what can be residency-pinned and what has to be global.
The Standard I'd Write#
RFC: Regional Independence Standard (v1)
Scope: All tier-1 services (customer-facing, revenue-bearing, or in the checkout/auth/payment path).
Requirements:
- Services MUST serve requests in any active region with no synchronous dependency on another region.
- Every table MUST declare a data class (GLOBAL_SYNC / HOME_REGION / COMMUTATIVE / REGION_PINNED / DERIVED) and an owner.
- Last-writer-wins MUST NOT be used on fields with financial, legal or security meaning.
- Ownership moves MUST be epoch-fenced; stores MUST reject writes with stale epochs.
- Failover and evacuation MUST use only data-plane actions; no autoscaling, deploys or control-plane API calls on the critical path.
- Each region MUST maintain headroom to absorb its designated failover share at peak (
evacuation_headroom_pct ≥ 100%).- Services MUST participate in quarterly evacuation drills; failures produce owned action items within 2 weeks.
- Global configuration SHOULD be readable from a local replica with last-known-good semantics.
Exceptions: Reviewed by the resilience architecture group within 5 business days; time-boxed to 2 quarters; listed on the evacuation runbook as known gaps.
Success metrics: Time-to-evacuate < 10 minutes; 100% of tier-1 services pass drills; zero cross-region synchronous dependencies on tier-1 paths; measured RPO within signed RPO 99.9% of the time.
What I'd Tell the VP#
"Today, if our main region has a serious outage, we're down for hours and could lose recent orders. I'm proposing we run in three regions so any one can fail and customers barely notice. It raises infrastructure costs by roughly 60% and needs a team of about five to run it, and every product team will spend a week or two classifying their data. The biggest risk isn't the technology — it's that teams add hidden dependencies over time, so we'll practice evacuating a region every quarter and treat failures in practice as progress. Before we commit, I'd like to confirm that region outages are a meaningful share of our downtime; if they're not, a cheaper step comes first."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Questions the premise with data | "What share of our incidents were regional? If it's under 10%, cells first." |
| Prices the program | "Three regions is ~1.5× infra plus five engineers; one regional outage costs us ~$2M." |
| Treats independence as a decaying property | "Every new dependency can break evacuation. Drills and trace-based audits are the control." |
| Picks the one-way doors | "The ownership partition key and the sync-tier vendor get a design review; the GLB vendor doesn't." |
| Designs decision rights | "The IC evacuates on defined signals without waiting for executives." |
Staff answers that L7 interviewers find insufficient:
- "Home-region ownership with epoch fencing." — correct for one service; silent on how 40 services adopt it and who audits them.
- "We'll run quarterly evacuation drills." — no owner for the program, no consequence when a team fails a drill, no SLO.
- "Three regions gives 1.5× capacity." — right math, but no comparison to the cost of outages or to cheaper alternatives like cells.
Appendices
Appendix A: Replication and Conflict Mechanics#
A.1 Home-Region Write Path#
handle_write(req):
p = partition(req.owner_key) # hash mod 4096
home, epoch = ownership_map.local_get(p)
if home != LOCAL_REGION:
return forward(home, req, epoch) # +1 RTT
txn = store.begin(partition=p)
txn.require_epoch(p, epoch) # store rejects if its epoch > epoch
txn.apply(req)
pos = txn.commit() # async replicate to followers
return ok(read_token=pos) # read-your-writes token
A.2 Read-Your-Writes on Followers#
handle_read(req):
if req.read_token and follower.position(p) < req.read_token:
wait_until(follower.position(p) >= req.read_token, max=50ms) \
or forward_to_home(req)
return follower.read(req)
A.3 CRDTs for the Commutative Class#
| CRDT | Use | Merge |
|---|---|---|
| G-Counter | Views, likes (increment-only) | Per-region counters; value = sum; merge = max per region |
| PN-Counter | Up/down votes | Two G-Counters |
| OR-Set | Tags, membership, comments with deletes | Adds tagged with unique IDs; removes tombstone observed IDs |
| LWW-Register | Single preference fields | Highest (timestamp, region_id) wins — explicit, accepted loss |
A.4 Why LWW Loses Data#
us-east t=10:00:00.120 (clock +40ms skew) balance = 100 - 30 = 70
eu-west t=10:00:00.100 balance = 100 - 50 = 50
Replication: us-east timestamp larger → final balance = 70
Correct: 20. One debit of 50 vanished. No error anywhere.
Appendix B: Ownership Map and Partitioning#
B.1 Partition Design#
- Fixed partition count (e.g., 4,096) hashed from owner key; partitions are the unit of homing and re-homing
- Ownership record:
partition → {home_region, home_cell, epoch, state: ACTIVE | MOVING | FROZEN} - Map stored in a small synchronously replicated store (writes rare); every region reads a local replica
B.2 Re-homing Protocol#
1. state = MOVING; destination starts as follower, catches up (lag < 100ms)
2. source stops accepting writes for p (FROZEN), drains in-flight (< 1s)
3. destination confirms position == source final position
4. map: home = destination, epoch += 1, state = ACTIVE
5. source becomes follower; any write with old epoch is rejected
Write pause per partition: ~0.2–1 s. Rate-limit re-homing (e.g., 100 partitions/min) to cap replication load.
Appendix C: Topology and Database Options — Quick Comparison#
| Option | Write Latency | RPO | Conflicts | Ops Burden | Best For |
|---|---|---|---|---|---|
| Active-passive, async | Local | Lag | None | Medium | Single-writer workloads |
| Home-region, async followers | Local for home users | Lag | None | Medium–High (map, re-homing) | Most user/tenant data |
| Write-anywhere + LWW (e.g., DynamoDB Global Tables) | Local | Lag | Silent loss | Low (managed) | Preferences, sessions |
| Write-anywhere + CRDT | Local | Lag (convergent) | Merged | Medium | Counters, sets, presence |
| Sync consensus (Spanner, CockroachDB multi-region) | +1 cross-region RTT | 0 | None | Low–Medium (managed) | Ledgers, uniqueness |
Replication internals — quorums, leader election, anti-entropy — are covered in Replicated Data Store.
Appendix D: Routing and Client Behavior#
| Mechanism | Failover Speed | Granularity | Notes |
|---|---|---|---|
| GeoDNS / latency DNS | TTL + 5–15 min tail | Per resolver | Universal; slow tail |
| Anycast + global L7 LB | ~30 s – 2 min | Per request, weighted | Best default for web/API |
| Client region list (mobile/SDK) | Seconds | Per client | Retry another region on connect failure |
Client rules: retries across regions use jittered backoff; idempotency keys on writes so a retried write in another region doesn't double-apply (see Idempotency); sticky region per session to avoid read-your-writes violations from region hopping.
Appendix E: Observability#
E.1 Core Metrics#
# Routing
region_request_share{region}, glb_backend_health{region}, region_error_rate{region}
# Data
replication_lag_seconds{partition,src,dst} p50/p99 by hour
rpo_budget_burn{class} time fraction lag > signed RPO
ownership_epoch_conflicts_total, partition_leader_count{partition}
lww_overwrites_total{table}
cross_region_write_ratio{region}
# Readiness
evacuation_headroom_pct{region}
days_since_last_successful_drill
cross_region_sync_dependencies{service} from tracing
E.2 Critical Alerts#
| Alert | Threshold | Severity |
|---|---|---|
| Replication lag | p99 > 5 s for 10 min | Page data platform |
| Split brain | partition_leader_count > 1 | Page immediately |
| Headroom | evacuation_headroom_pct < 100% | Ticket capacity owner; page before peak events |
| Region errors | > 5% for 5 min | Page traffic on-call; runbook: drain by weight |
| Drill staleness | > 100 days | Ticket to resilience program owner |
| New cross-region sync dependency | Any on tier-1 path | Ticket owning team |
E.3 Control Plane vs Data Plane#
During failover, only data-plane actions run: global LB weights (pre-authorized), ownership map writes (replicated store), and feature degradation flags read from local caches. Deploys, autoscaling, IAM changes and DNS record edits are explicitly not on the failover path.
Appendix F: Scale Evolution#
| Stage | Topology | What You Don't Build Yet |
|---|---|---|
| Startup | One region, multi-AZ, backups elsewhere | Any cross-region writes |
| Growth | Warm standby → hot; reads in region 2; cells | Write-anywhere, CRDT platform |
| Scale | 3 regions, home-region ownership, sync tier | Per-user re-homing automation (start with tenant/cohort) |
| Global | 5+ regions, sovereign tiers, region templates | Custom consensus database |
Appendix G: Multi-Tenancy, Fairness and Cost#
- Cell assignment: large tenants capped at ~15% of a cell; enterprise isolation priced as a tier
- Failover fairness: during evacuation, shed by tier — free tier degrades before enterprise — agreed with product in advance
- Residency tiers: standard (global), regional (home pinned, async followers in-continent), sovereign (in-jurisdiction only, two in-jurisdiction regions)
- Chargeback: charge sync-tier writes and cross-region egress to owning teams so the premium stays visible