Technologies referenced in this case study: API Gateways & Service Mesh · Redis · Kafka · ZooKeeper & etcd
Related case studies: Rate Limiting · Load Balancer · Service Discovery · API Gateway · Degraded Mode
How to Use This Case Study#
Organized for interview use first, reference second. Read front-to-back once, then come back to the sections where you are weak.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 3, 4 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → §3 Fault Lines → §4 Failure Modes → two Deep Dives |
| Deep Dive | 3+ hrs | Everything, including §11 Principal Lens and the appendices |
What is Fault Tolerance with Circuit Breakers? — Why interviewers pick this topic
A circuit breaker is a small state machine that sits between a caller and a dependency. When the dependency starts failing or slowing down, the breaker opens and the caller stops sending it traffic for a while — failing fast or serving a fallback instead of waiting. After a cool-down it lets a few probe requests through (half-open) and closes again if they succeed.
The breaker is the most famous piece of a larger toolkit: timeouts, retries with budgets, bulkheads, load shedding, and backpressure. Interviewers use "circuit breaker" as the entry point to that whole toolkit, because the real question is: how does one slow dependency avoid taking down the whole company?
Before vs After — the slow recommendations service:
Without fault tolerance:
t=0: Recommendations service p99 goes from 40ms to 8s (GC death spiral)
t=+5s: Product-page service has 200 worker threads; all 200 now blocked on recs
t=+10s: Product-page health checks time out; load balancer marks instances unhealthy
t=+15s: Remaining instances absorb the traffic, block on recs, go unhealthy too
t=+30s: Clients retry 3x; product-page ingress QPS is now 4x normal
t=+2min: Checkout (which calls product-page for prices) starts failing
t=+40min: Full site outage caused by a non-critical widget
With fault tolerance:
t=0: Same recs slowdown
t=+0.3s: Product-page's 300ms timeout on recs fires; bulkhead caps recs to 20 threads
t=+5s: Breaker sees 60% slow calls over 100 requests, opens
t=+5s: Product page renders with a static 'Popular items' fallback
t=+35s: Half-open probes: 10 requests; recs still slow; breaker stays open
t=+6min: Recs team rolls back; probes succeed; breaker closes
Result: Checkout never noticed. A dashboard blip and a page to the recs team only.
Why interviewers reach for this question: It is the purest test of whether you think in systems of systems. Every candidate can draw the closed/open/half-open diagram. Very few can explain why their retries made the outage worse, what the fallback actually returns, who owns the threshold, or why a breaker that opens on every instance at once can be as dangerous as no breaker at all.
Mechanics Refresher: The Fault-Tolerance Toolkit
| Mechanism | How It Works | Pros | Cons |
|---|---|---|---|
| Timeout | Abandon a call after N ms | Bounds the worst case; the single most important control | Too long = thread exhaustion; too short = false failures and retry storms |
| Retry (with backoff + jitter) | Re-issue a failed idempotent call after a randomized delay | Masks transient faults (a packet drop, one bad host) | Multiplies load exactly when the dependency is weakest |
| Retry budget | Cap retries to a % of successful traffic (e.g. 10–20%) | Retries help when failures are rare, disappear when they are widespread | Needs per-client accounting; often forgotten in libraries |
| Circuit breaker | Track failure/slow-call rate; stop calling when it crosses a threshold | Fails fast; frees caller resources; gives dependency room to recover | Per-instance view is noisy; thresholds are hard to tune; fallback code is rarely tested |
| Bulkhead | Separate thread pools / connection pools / semaphores per dependency | One slow dependency cannot consume all caller capacity | Capacity fragmentation; pool sizing is a guess until measured |
| Load shedding | Callee rejects work it cannot finish in time (by priority/criticality) | Protects the server from the outside; works even when clients misbehave | Needs a cheap rejection path and a notion of request priority |
| Adaptive concurrency limit | Server/client learns max in-flight from latency (AIMD, gradient) | Tracks real capacity without hand-tuned thresholds | Oscillation; harder to reason about during incidents |
| Backpressure | Propagate "slow down" upstream (queue depth, credits, 429/503 + Retry-After) | Fixes the cause, not the symptom | Only works if every hop honors it |
For most production systems: Timeouts and deadline propagation on every call, a retry budget instead of a retry count, bulkheads for every non-critical dependency, server-side load shedding by criticality — and a circuit breaker last, as the component that ties these together. The breaker is the least important of the six. That sentence alone will surprise most interviewers.
Executive Summary
If you only read one section, read this. Every later section expands one row of what follows.
What This Interview Actually Tests#
Circuit breakers are not a state-machine question. Everyone can draw closed → open → half-open.
This is a cascading-failure containment question that tests:
- Whether you know that the slow dependency, not the dead one, is what kills systems
- Whether you can reason about load amplification — retries, fan-out, and fallbacks that multiply traffic at the worst moment
- Whether you decide, per dependency, what "degraded" means to a user — and who signed off on it
- Whether you put protection in the right place: caller, callee, mesh, or all three
The key insight: A circuit breaker is a policy about which failures you are willing to show users in exchange for keeping everything else alive. The mechanism is 50 lines of code. The policy is a negotiation between product, the calling team, and the dependency team.
The L5 vs L6 vs L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Wraps every client call in a breaker with default thresholds | Classifies dependencies as critical / degradable / optional and designs a failure behavior for each | Asks which dependencies should not exist on the request path at all, and sets an org-wide criticality taxonomy |
| Timeouts | "We'll set a 5s timeout" | Derives timeouts from the caller's deadline and the callee's p99.9; propagates the remaining deadline | Makes deadline propagation a platform default (RPC framework), so no team can forget it |
| Retries | "Retry 3 times with exponential backoff" | Retries at one layer only, with a 10% budget and jitter; knows 3 retries × 4 layers = 256× amplification | Treats retry policy as a shared resource with a fleet-wide budget and a post-incident review item |
| Failure | Breaker opens → return 503 | Breaker opens → per-dependency fallback (cached, static, omitted section) that product has approved | Designs the org's failure posture: cells, criticality tiers, and regular game days that prove fallbacks run |
| Ownership | Each team configures its own library | Mesh/sidecar owns mechanics, calling team owns fallback semantics, callee owns load shedding | Decides library vs mesh for 300 services, funds the platform team, retires the three competing libraries |
| Scale | Tunes per-instance thresholds | Recognizes per-instance breakers across 500 callers act like a synchronized herd; adds jitter and adaptive limits | Prices a correlated-failure event in $/minute and uses that to justify cell architecture spend |
Why "first move" separates levels
L5: Reaches for the mechanism. "Put Resilience4j around the payment client, the recommendations client and the inventory client, 50% failure threshold, 30s open." Reasonable — and it treats every dependency the same. A payment breaker that opens and returns "payment failed" and a recommendations breaker that opens and hides a carousel are not the same decision.
L6: Starts with the dependency graph and classifies. "Before I tune anything: which of these calls can the page live without? Recommendations are optional — we omit the section. Pricing is degradable — we serve the last cached price with a 5-minute staleness cap. Payments are critical — there is no fallback; we fail the request clearly and never retry non-idempotently." The mechanism is identical; the policy is completely different per row.
L7: Asks why a critical path has twelve synchronous dependencies to begin with, and whether some should be made asynchronous or precomputed. Establishes criticality tiers org-wide so a new service declares "I am tier-2 degradable" at creation time, rather than every team rediscovering this in an incident.
Why "retries" separates levels
L5: "Retry 3 times with exponential backoff" is correct for a single client talking to a single server. It is dangerous in a 5-deep call graph where every layer does the same thing: a failure at the bottom produces up to 4⁵ = 1,024 attempts per user request.
L6: "Retries happen at exactly one layer — the one closest to the failure that knows the call is idempotent. Everywhere else, a failure propagates. And retries are budgeted: at most 10% extra load over successful traffic, so when the dependency is healthy retries mask blips, and when it is sick they vanish automatically."
L7: Recognizes this cannot be enforced by code review across 300 services. Puts retry policy into the mesh or RPC framework, removes per-call retry knobs from application code, and adds "retry amplification factor" to the standard incident review template.
Why "ownership" separates levels
L5: Each team adds the library, picks thresholds, and moves on. Six months later there are Hystrix, Resilience4j, Polly, and a hand-rolled Go breaker in production, with four different metrics formats and no fleet view.
L6: Splits the concern: "Mechanics — timeouts, retry budgets, outlier ejection, connection limits — live in the sidecar, owned by the platform team. Semantics — what the fallback returns — live in the calling service, owned by the product team, because only they know whether a stale price is acceptable. Protection of the callee — load shedding — lives in the callee, owned by that team, because they know their capacity."
L7: Writes that split down as a standard and funds it: platform headcount, a migration off legacy libraries, and a criticality registry that the mesh reads.
The Staff Positions#
| Position | Rationale |
|---|---|
| Timeouts before breakers | A breaker without a timeout never sees the slow call finish — it cannot trip on what it cannot measure |
| Slow is worse than down | Dead dependencies fail fast (connection refused in <1ms). Slow ones hold threads for seconds. Design for slowness first |
| Retry budgets, not retry counts | A count multiplies load under failure; a budget (≤10–20% of successes) caps it |
| Retry at one layer only | Nested retries multiply geometrically; pick the layer that knows idempotency |
| Server-side load shedding is mandatory | Client-side breakers are a courtesy; you cannot trust 400 clients to behave. The callee must protect itself |
| Every fallback is a product decision | "Return cached data" is a correctness choice; product signs off on staleness, not engineering |
| Mechanics in the mesh, semantics in the service | Mechanism consistency across the fleet; fallback meaning where the domain knowledge lives |
The Three Intents#
"Add circuit breakers" hides three different goals. They lead to different placements, different thresholds and different failure semantics.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Contain the cascade (protect the caller) | Caller threads, connections and latency budget | Timeouts, bulkheads, breaker at the caller, fail fast | Users see a degraded feature, not a dead page | Caller p99 stays within SLO while dependency is down |
| Let the dependency recover (protect the callee) | Callee capacity; recovery needs load to drop | Retry budgets, jittered backoff, server-side load shedding, adaptive concurrency | Some requests rejected early with 503/429 | Callee goodput stays ≥ 90% of capacity under 3× overload |
| Preserve the user experience (degrade gracefully) | Product semantics; what users can tolerate | Per-dependency fallbacks: cache, default, omit, queue for later | Users see stale or partial data | Staleness/partiality bounds signed off by product |
🎯 Staff Move: "I'll design for containing the cascade first, because that's what turns a single-team incident into a company outage. But I'll say up front that the breaker alone doesn't let the dependency recover — that needs server-side shedding and retry budgets — and it doesn't decide what users see — that's a per-dependency fallback product has to approve. Three intents, three owners."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Fail Fast vs Degrade | Return an error immediately (honest, simple) or serve a fallback (better UX, but a code path that rarely runs and may be wrong)? |
| 2 | Caller Protection vs Callee Protection | Put the intelligence in 400 clients (breakers) or in 1 server (load shedding)? Who do you trust? |
| 3 | Retry for Success vs Retry Amplification | Retries fix transient faults and multiply sustained ones. Where, how many, and under what budget? |
| 4 | Static Thresholds vs Adaptive Limits | Hand-tuned "50% errors over 10s" is legible but wrong as load changes; adaptive limits track capacity but are harder to reason about |
| 5 | Library vs Mesh (Ownership) | In-process library (rich semantics, polyglot drift) or sidecar/mesh (uniform mechanics, no domain knowledge)? |
In the Wild: Real Production Systems#
Why this section belongs here: Citing specific systems shows you've studied operational reality. Use one of these in the first ten minutes.
Netflix — Hystrix and the Bulkhead Model#
Netflix built and open-sourced Hystrix (2012) after learning that a single slow dependency among hundreds could saturate every request thread in the API tier. Hystrix wrapped each dependency in its own thread pool (a bulkhead — default 10 threads), enforced a timeout, tracked failures in a rolling 10-second window, and opened the circuit when at least 20 requests had been seen and ≥50% failed, probing again after 5 seconds. Every command had an explicit fallback. Hystrix went into maintenance mode in 2018; Netflix moved toward adaptive concurrency limits (their open-source concurrency-limits library) that infer capacity from latency rather than using static thresholds.
Staff insight: The lasting contribution was not the state machine — it was the bulkhead per dependency and the fallback per call. And the move away from static thresholds is the interview-worthy lesson: fixed thresholds are wrong the moment load patterns change.
Google — Client-Side Adaptive Throttling and Criticality#
The Google SRE book describes client-side adaptive throttling: each client tracks requests and accepts over the last two minutes and locally rejects new requests with probability max(0, (requests − K × accepts) / (requests + 1)), with K typically 2. Clients stop sending work the backend is going to reject anyway — a breaker with a continuous, not binary, response. Requests also carry a criticality (e.g. CRITICAL_PLUS, CRITICAL, SHEDDABLE_PLUS, SHEDDABLE) so an overloaded backend sheds the least important traffic first, plus per-request retry caps (3 attempts) and a per-client retry budget (~10%).
Staff insight: "Circuit breaker" is a special case of adaptive throttling with K → ∞ and a binary output. Citing the formula signals you know the continuous version exists and why it avoids the thundering-herd close.
Amazon — Timeouts, Jittered Retries, and Cells#
The Amazon Builders' Library publishes the company's defaults: timeouts derived from downstream latency percentiles, exponential backoff with jitter, retry token buckets in the AWS SDKs so retries stop when failures are widespread, and load shedding that rejects early rather than letting queued work time out. Amazon's architecture guidance also pushes cell-based architecture and shuffle sharding so a poison request or a bad customer affects a bounded fraction of capacity. The public post-mortem of the September 2015 DynamoDB disruption described exactly the failure this case study is about: storage servers' metadata requests timed out, retried, and kept the metadata service overloaded until load was manually reduced.
Staff insight: Amazon treats retries as a system-level risk, not a client convenience. The DynamoDB event is the canonical example of a metastable failure: the trigger was brief, the retry-driven overload was self-sustaining.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Add a circuit breaker" | "What does the user see when it's open?" | Whether fallback semantics are designed or hand-waved |
| "Retry 3 times with backoff" | "The call graph is 4 deep. What's the worst-case amplification?" | Whether you model load, not just one call |
| "50% error threshold" | "Errors or latency? What about a dependency that returns 200 in 9 seconds?" | Whether you know slow-call detection matters more than error rate |
| "The breaker protects the service" | "Which service? The caller or the callee?" | Caller vs callee protection clarity |
| "We'll use a service mesh" | "The mesh can't know your fallback. Who writes it?" | Ownership split between platform and product |
| "Each instance has its own breaker" | "500 instances all open at t=5s and close at t=35s. Then what?" | Correlated behavior and thundering herd on recovery |
System Architecture Overview#
Reading the diagram: Protection is layered, and each layer has a different owner. The gateway sets the overall deadline and criticality. The calling service owns bulkheads, breakers and — critically — the fallback semantics, because only it knows what a stale price means. The sidecar (platform team) owns uniform mechanics: connect timeouts, retry budgets, outlier ejection. The callee owns load shedding, because it is the only component that knows its own capacity and cannot trust 400 callers to behave. Observability emits breaker state and fallback counts, because a breaker that is silently open for three days looks exactly like a healthy system on a latency graph.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Timeouts | "5-second timeout on every call" | "Deadline propagation from the edge. Each hop's timeout = min(remaining deadline, ~p99.9 × 1.5). No hop waits longer than its caller will." |
| Retries | "3 retries with exponential backoff" | "Retries at one layer, idempotent calls only, full jitter, and a 10% retry budget so they disappear under widespread failure." |
| Breaker trigger | "Open at 50% errors" | "Open on error rate or slow-call rate over a minimum volume of ~20–100 calls. Slow calls are the real killer." |
| When open | "Return 503" | "Per-dependency fallback: omit, default, cached-with-staleness-cap, or fail — each approved by product." |
| Callee protection | "Clients have breakers" | "Callee sheds load by criticality and queue-wait time. Never trust clients to protect you." |
| Ownership | "Each team adds the library" | "Mechanics in the mesh (platform), fallback semantics in the caller (product), shedding in the callee (dependency owner)." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Hystrix classic defaults | 20 req min volume, 50% errors, 10s window, 5s sleep | The reference point every interviewer knows; say why you'd change them |
| Resilience4j defaults | 50% failure rate, 100-call window, 60s open, 10 half-open calls | Shows the industry shifted toward larger windows and longer open times |
| Retry amplification, 3 retries × N layers | 4ᴺ: 4 layers → 256×, 5 layers → 1,024× | The single most important number in this case study |
| Retry budget (Google SRE, Envoy) | ~10% per client (Google); 20% default budget in Envoy | Caps retry load at 1.1–1.2× instead of 4× |
| Per-request attempt cap (Google SRE) | 3 attempts | Bounded even for a single request |
| Adaptive throttle multiplier K | 2 (accept ≤ 2× what backend accepts) | Continuous alternative to binary breaker |
| Little's Law | in-flight = RPS × latency: 1,000 RPS × 50ms = 50; × 5s = 5,000 | Why slowness exhausts threads 100× faster than errors |
| Connection refused vs timeout | <1ms vs full timeout (often 1–30s) | A dead dependency is cheap; a slow one is expensive |
| Envoy circuit-breaker defaults | 1,024 max connections / pending / requests; 3 max concurrent retries | Mesh breakers are concurrency caps, not error-rate state machines |
| Envoy outlier detection defaults | 5 consecutive 5xx → eject 30s; max 10% of hosts ejected | Host-level breaker; the 10% cap prevents ejecting the whole cluster |
| Timeout rule of thumb | p99.9 of downstream × 1.2–1.5, bounded by remaining deadline | Timeouts should be derived, not guessed |
| Load-shed cost vs serve cost | Rejecting should cost ≤ 1–5% of serving | If rejection is expensive, shedding cannot save you |
Interview Walkthrough
The prompt usually arrives as one of: "Design fault tolerance for a microservice platform", "Our checkout went down because recommendations were slow — design it so that can't happen", or "Design a circuit breaker library." The third is a trap: if you spend 30 minutes on a state machine you will be leveled Senior. Spend 5 on the mechanism and 30 on placement, amplification, fallbacks and ownership.
Phase 1: Requirements & Framing (2–3 min)#
Say this, nearly verbatim:
"Before I design anything, I want to know the shape of the dependency graph and what 'failure' means to the user. Three questions. First: which dependencies are on the critical path — is there any call the page cannot render without? Second: how deep is the call graph — are we 2 hops or 6? That determines how dangerous retries are. Third: what does the business want when a dependency is sick — an error, stale data, or a partial page? I'll assume a product page with ~8 synchronous dependencies, a 4-deep call graph, 20K RPS at peak, and a 1.5s end-to-end deadline. I'll optimize for containing cascades first."
What you have established in 30 seconds:
- Criticality is per dependency, not global
- Depth determines amplification risk
- Degraded semantics are a product decision
- A deadline exists, and you will propagate it
Phase 2: Core Entities & API (1–2 min)#
Name the entities the design manipulates. Keep it short.
| Entity | Fields That Matter | Owner |
|---|---|---|
| Dependency policy | name, criticality (critical / degradable / optional), timeout, retry policy, bulkhead size, breaker thresholds, fallback id | Calling team (semantics) + platform (defaults) |
| Breaker state | CLOSED / OPEN / HALF_OPEN, window counters (calls, failures, slow calls), opened_at | In-process or sidecar, per instance |
| Request context | deadline (absolute), criticality, attempt number, retry-allowed flag | Propagated on every hop |
| Fallback | kind (omit / static / cached / queue), staleness cap, product sign-off | Product + calling team |
The request context is the API. Everything else is local:
Headers on every internal RPC:
x-request-deadline: 2026-09-29T10:15:03.412Z # absolute, not relative
x-criticality: CRITICAL | DEGRADABLE | SHEDDABLE
x-attempt: 1 # retries set 2, 3...
x-retry-allowed: false # set by the retrying layer so lower layers don't retry
🎯 Staff Move: "The most important API here is the one nobody draws: the request context. Deadline, criticality and attempt count on every hop are what let a service four layers down make the right decision without knowing the call graph."
Phase 3: High-Level Architecture (≤5 min)#
Staff candidates spend under 5 minutes here. Draw the layered picture from the Executive Summary and name the owner of each layer:
- Edge/gateway — sets deadline, tags criticality, is the only place user-facing retries happen for non-idempotent flows (usually: none).
- Calling service — bulkhead per dependency, breaker per dependency, fallback per dependency.
- Sidecar/mesh — connect timeout, retry budget, outlier ejection, max concurrent requests.
- Callee — load shedding by queue-wait and criticality; adaptive concurrency limit.
- Observability — breaker state, fallback served, retry ratio, shed count, deadline-exceeded count.
"That's the whole architecture. The interesting part is not the boxes — it's how the policies interact, especially retries and correlated breaker behavior. Let me go there."
Phase 4: Transition to Depth#
The sentence that steers the interviewer toward your strengths:
"There are three places this design usually fails in production, and I'd like to go through them in order of how often they cause real outages: retry amplification, slow-call detection, and correlated recovery — every breaker in the fleet closing at the same moment. Then I'll cover fallbacks and ownership. Does that order work for you?"
If the interviewer wants the state machine, give it in 90 seconds and pivot back.
Phase 5: Deep Dives (25–30 min)#
Deep dive 1 — Timeouts and deadlines (5 min). Derive, don't guess:
End-to-end deadline: 1,500 ms (set at gateway)
Gateway → product-page: remaining ≈ 1,480 ms
product-page own work: ~100 ms
product-page → pricing: timeout = min(remaining − reserve, p99.9_pricing × 1.5)
= min(1,380 − 100, 120 × 1.5) = 180 ms
product-page → recs: timeout = min(…, p99.9_recs × 1.5) = 300 ms, OPTIONAL
If remaining deadline < 50 ms: don't call at all — fail/fallback immediately
"If a callee receives a request whose deadline has already passed, it should drop it without doing work. That single check prevents the 'zombie work' that keeps overloaded services overloaded."
Deep dive 2 — Retries and amplification (7 min). State the math, then the policy:
Naive: each of 4 layers retries 3x → up to 4^4 = 256 attempts at the bottom per user request
Staff: retry at 1 layer (the one that knows idempotency), 2 attempts max,
budget: retries ≤ 10% of successful calls over a 10s window,
backoff: full jitter, base 25ms, cap 250ms,
never retry on: deadline exceeded, 429/503 without Retry-After, non-idempotent POST
Result: worst-case load at the bottom under total failure ≈ 1.1× normal
Deep dive 3 — Breaker trigger and slow calls (5 min). Error-rate breakers miss the dangerous case: a dependency returning 200 OK in 8 seconds. Trip on slow_call_rate too, with a minimum volume so a service getting 3 RPS doesn't flap on one failure.
Deep dive 4 — Correlated recovery (5 min). 500 caller instances each run a local breaker. All see the same failure, all open within ~1s, all go half-open 30s later, all send 10 probes = 5,000 simultaneous probes into a service that just recovered. Fixes: jitter the open duration (30s ± 20%), ramp traffic after close (10% → 25% → 50% → 100% over 60s), and prefer continuous adaptive throttling over binary open/close.
Deep dive 5 — Fallbacks and ownership (5 min). Walk the criticality table; state that product signed off on each fallback; state that fallbacks are exercised weekly in a game day or they are assumed broken.
Phase 6: Wrap-Up (2–3 min)#
"To summarize: timeouts from propagated deadlines, retries at one layer under a 10% budget, bulkheads and breakers per dependency with slow-call detection, fallbacks product has approved, and load shedding in every callee. Mechanics live in the mesh; semantics live in the service. What I'd build next: adaptive concurrency limits to replace hand-tuned thresholds, a criticality registry the mesh reads, and a monthly game day that forces every optional dependency's breaker open in production for 10 minutes. The biggest remaining risk is fallbacks that have never run — I'd measure that as 'days since fallback last served traffic' per dependency."
Common Timing Mistakes#
| Mistake | Time Lost | What to Do Instead |
|---|---|---|
| Drawing the closed/open/half-open diagram in detail | 8–10 min | 90 seconds; it is table stakes |
| Debating Hystrix vs Resilience4j vs Polly | 5 min | One sentence: "any mature library or the mesh; the policy matters more" |
| Designing a distributed shared breaker state store | 10 min | Say local breakers + server-side shedding; shared state adds a dependency to your fault-tolerance layer |
| Never mentioning retries | — | Retries are where the outages come from; lead with amplification |
| Skipping what the user sees | — | Every breaker needs a fallback row in the criticality table |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Senior engineers own services. Staff engineers own the interactions between services. Fault tolerance is the canonical interaction problem: every individual component can be well-built, and the system can still collapse because of how they combine — retries that multiply, timeouts that nest the wrong way, breakers that synchronize, fallbacks that shift load onto the database that was already struggling. No single team's code review catches any of it.
That is why interviewers use it. It separates candidates who think "my service is robust" from candidates who think "my service's behavior under failure is part of someone else's failure."
1.2 The L5 vs L6 vs L7 Contrast — Visual#
The L5 path is not wrong — a breaker on recs would have helped. It fixes this incident. The L6 path asks why the architecture allowed it, and fixes the class. The L7 path makes the fix the default for the next 300 services.
1.3 The Staff Question That Cuts Through Everything#
"When this dependency is slow — not down, slow — what happens to the caller's threads, and what does the user see?"
Ask this for every dependency on the whiteboard. It forces:
- A timeout (otherwise the answer is "threads block forever")
- A bulkhead (otherwise the answer is "all threads block")
- A fallback (otherwise the answer is "an error page")
- A named owner for the fallback semantics
If you only remember one sentence from this case study, remember that one.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Intent 1: Contain the cascade (protect the caller). The failure you are preventing is resource exhaustion in the caller: threads, connections, memory for queued requests, and latency budget. By Little's Law, a caller at 1,000 RPS whose dependency goes from 50ms to 5s needs 5,000 in-flight slots instead of 50. No thread pool survives that. The tools are timeouts, bulkheads and caller-side breakers. Success metric: the caller's own p99 and availability stay inside SLO while the dependency is at 0%.
Intent 2: Let the dependency recover (protect the callee). The failure you are preventing is a metastable state: the dependency could recover if load dropped, but retries and queued work keep it saturated. Caller-side breakers help if every caller has one — but you don't control every caller. The tools are server-side load shedding (reject early and cheaply), retry budgets, and adaptive concurrency. Success metric: under 3× offered load, the callee's goodput stays near capacity instead of collapsing to zero.
Intent 3: Preserve the user experience (degrade gracefully). The failure you are preventing is a binary outcome — full page or error page — when a partial page would do. The tools are fallbacks: omit the section, show a static default, serve cached data with a staleness cap, or accept the write and process later. Success metric: conversion or task completion during a dependency outage, measured against baseline.
🎯 Staff Move: "These three have different owners. The caller team owns containment, the dependency team owns their own recovery, and product owns what the degraded experience is. If I design only the breaker, I've only designed the first one."
2.2 When NOT to Use Circuit Breakers#
Breakers are not free. They add a state machine, thresholds that need tuning, and a code path (the open state) that almost never runs. Don't add one when:
| Situation | Why a Breaker Is Wrong | What to Use Instead |
|---|---|---|
| Single-instance dependency with no alternative (e.g., the primary DB for a write) | Opening the breaker = failing every write; the timeout already fails fast | Timeout + connection pool limit + clear error |
| Low-volume calls (< 1 RPS per caller instance) | Not enough samples; one failure flips a 50% threshold | Timeouts + retry with budget; aggregate at the mesh |
| Asynchronous consumers (queue workers) | The queue is the buffer; a breaker just stops consuming | Consumer backoff, pause partition, DLQ (Message Queue) |
| Per-host failures behind a load balancer | Service-level breaker opens for a problem with 1 of 50 hosts | Outlier detection / host ejection at the LB or sidecar |
| Non-idempotent critical writes (payment capture) | Fallback is meaningless; you must not pretend success | Idempotency keys, fail clearly, reconcile async (Payment Processing) |
| Dependencies you could remove from the request path | A breaker manages a coupling you should eliminate | Precompute, cache asynchronously, or move to an event |
"The best circuit breaker is a dependency you took off the synchronous path."
2.3 What the Interviewer Leaves Underspecified#
| Unstated Assumption | Why It Matters | What to Say |
|---|---|---|
| Call graph depth | Amplification is exponential in depth | "I'll assume 4 hops; that's why I'll restrict retries to one layer" |
| Idempotency of each call | Determines where retries are legal | "Reads are idempotent; writes carry idempotency keys or aren't retried" |
| Criticality of each dependency | Determines fallback vs fail | "I'll classify each as critical / degradable / optional" |
| Who owns the callers | If you don't control callers, caller-side protection is optional for them | "External and legacy callers won't have breakers, so callees must shed" |
| Latency distribution of dependencies | Timeouts derive from p99.9, not averages | "I'll set timeouts from measured p99.9 × 1.5, capped by remaining deadline" |
| Synchronous vs async boundary | Async edges need different tools | "Anything that can be eventual, I'll move to a queue before adding a breaker" |
2.4 Precise Terminology#
| Term | Precise Meaning | Common Confusion |
|---|---|---|
| Timeout | Max time a caller waits for one attempt | Confused with deadline; a timeout is per-attempt |
| Deadline | Absolute time by which the whole request must complete, propagated across hops | Using relative timeouts per hop, which nest incorrectly |
| Circuit breaker | Caller-side state machine that stops calls after a failure/slow threshold | Used loosely for any protection, including LB host ejection |
| Outlier detection | Ejecting individual hosts from a pool based on their failures | Not a breaker for the service; operates one level down |
| Bulkhead | Isolated capacity (threads, connections, semaphore permits) per dependency | Confused with rate limiting; bulkheads cap concurrency, not rate |
| Load shedding | Callee rejects work it cannot complete in time, before doing it | Confused with rate limiting, which enforces per-client policy regardless of server health |
| Backpressure | Signal propagated upstream to slow producers | Often claimed but not implemented — a 503 nobody honors is not backpressure |
| Retry budget | Cap on retries as a fraction of successful requests | Confused with max attempts per request; you need both |
| Goodput | Requests completed successfully within deadline per second | Throughput counts work that timed out and was wasted |
| Metastable failure | Overload that persists after the trigger is removed, sustained by the system's own reaction (retries, cache misses) | Treated as "the dependency was down for 40 minutes" when it was down for 30 seconds |
3. The Five Fault Lines#
These are the tensions where good engineers disagree. For each: options, who pays, the Staff default, and when to deviate.
3.1 Fault Line 1: Fail Fast vs Degrade#
When the breaker is open, you either return an error or serve something else.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Fail fast (error) | Honest; no hidden correctness risk; simple to test | User sees an error for a non-essential feature | The user; product's conversion metric |
| Omit the feature | Page renders; the missing section is rarely noticed | Requires the UI to handle absence | Front-end team (layout must tolerate gaps) |
| Static default | Always available; zero dependency | Generic, possibly irrelevant content | Product (lower engagement), but bounded |
| Cached / stale data | Close to real UX | Stale data can be wrong (prices, inventory, permissions) | The business, if a stale price is honored; security if stale auth is used |
| Queue for later | Accepts writes; completes eventually | User believes it's done; failure surfaces hours later | Support team handling "my order vanished" |
The Staff default: Classify first. Optional → omit. Degradable → cached with an explicit staleness cap and a visible indicator if the staleness matters. Critical → fail fast with a clear message; never fake success.
The trap: Fallbacks that call another dependency. "If the cache is down, fall back to the database" converts a cache outage into a database outage — the database was sized for a 5% miss rate and now takes 100%. A fallback must be cheaper and more available than the primary, or it is not a fallback.
🎯 Staff Move: "Every fallback I propose has three properties: it's cheaper than the primary, it doesn't depend on anything the primary depends on, and product has signed off on what it shows. A fallback that fails any of those is a second outage waiting to happen."
When to deviate: For read-heavy catalogs where staleness is harmless (product descriptions, images), serve stale aggressively — even hours old. For anything involving money, entitlements, or authorization, fail fast; a stale "allow" is a security incident.
3.2 Fault Line 2: Caller Protection vs Callee Protection#
Breakers live in the caller. Load shedding lives in the callee. Which do you rely on?
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Caller-only (breakers) | Stops traffic before it leaves; saves network and caller resources | Requires every caller to behave; one legacy client with aggressive retries still kills you | Callee team, paged for an overload they cannot control |
| Callee-only (shedding) | Protects the callee regardless of who calls; one place to tune | Callers still waste resources sending doomed requests; rejection still costs something | Callers (wasted threads/latency) |
| Both | Defense in depth; callers save resources, callee guarantees survival | Two sets of thresholds that can interact (callers' breakers open on callee's 503s — which is correct) | Platform team maintains two mechanisms |
The Staff default: Both, with the callee's shedding as the guarantee and the caller's breaker as the optimization. A service that depends on its callers' good behavior for survival does not have fault tolerance.
What callee-side shedding looks like:
on request arrival:
if now() > request.deadline: drop, no response work # zombie
if queue_wait_p50 > 50ms and criticality == SHEDDABLE: reject 503 + Retry-After
if inflight > adaptive_limit: reject by lowest criticality first
else: admit
cost of rejection must be ≤ ~1–5% of cost of serving
When to deviate: In a small system with 2–3 callers you control, caller-side breakers alone are acceptable for a first version — say so explicitly and name the trigger (first external or legacy caller) for adding shedding.
3.3 Fault Line 3: Retry for Success vs Retry Amplification#
Retries are the most dangerous fault-tolerance feature because they work perfectly in testing.
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| No retries | Zero amplification | Transient faults (one bad host, a dropped packet) become user errors — often 0.1–1% of requests | Users, for blips that were trivially recoverable |
| Fixed count at every layer | Masks transient faults locally | Geometric amplification: 3 retries × 4 layers = 256× | The callee, then everyone |
| Fixed count, one layer | Bounded: max 3–4× load under total failure | Still 3–4× load exactly when the dependency is weakest | The callee during an incident |
| Budgeted retries (≤10% of successes) | Near-zero amplification under widespread failure; still masks blips | Needs per-client accounting; budget needs a floor for low-traffic clients | Platform team (implementation) |
| Hedged requests (send a 2nd copy after p95) | Cuts tail latency for reads | Adds ~5% load constantly; dangerous without a budget | Callee capacity |
The Staff default: Retries only at one layer (usually the one closest to the failing dependency, or the sidecar), only for idempotent operations, full jitter, and a budget. Signal "don't retry" downstream with a header.
retry_allowed(req, resp):
if not req.idempotent: return false
if resp.status in {400..499} except 429: return false
if deadline_remaining(req) < min_useful: return false
if retries_10s / successes_10s > 0.10: return false # budget exhausted
return true
backoff = random(0, min(cap=250ms, base=25ms * 2^attempt)) # full jitter
🎯 Staff Move: "Three retries is a good policy for one client and a terrible policy for a fleet. I'd replace 'max attempts' with a budget: retries can add at most 10% to the load. When failures are rare, that's plenty; when they're widespread, retries disappear automatically, which is exactly when we need them to."
When to deviate: At the very edge (mobile client over a flaky network), one retry on connection failure is almost always right — the failure is likely the network, not the server. Hedging is worth it for read paths with strict tail latency SLOs (search, ads), with its own budget.
3.4 Fault Line 4: Static Thresholds vs Adaptive Limits#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Static breaker thresholds (50% over 100 calls, open 30s) | Legible; easy to reason about in an incident | Wrong as traffic shifts; 3 AM traffic trips on noise, peak traffic trips late | On-call, tuning thresholds after each incident |
| Static concurrency caps (bulkhead = 40) | Hard ceiling; simple | Caps are guesses; too low throttles healthy traffic, too high doesn't protect | Calling team, via mysterious rejections |
| Adaptive concurrency (AIMD / gradient on latency) | Tracks real capacity continuously; no hand tuning | Oscillation; less intuitive; needs a stable latency signal | Platform team (owns the algorithm) |
| Client-side adaptive throttling (reject with p = f(requests, accepts)) | Continuous response; no binary cliff; no synchronized reopen | Requires callee to reject cleanly so "accepts" is meaningful | Platform |
The Staff default: Static breakers with a minimum call volume and slow-call detection as a starting point, bulkhead sizes derived from Little's Law, and a roadmap item to move to adaptive concurrency on the highest-traffic edges. Say the math:
bulkhead size ≈ peak RPS to dependency × p99 latency × headroom
= 400 RPS × 0.080 s × 1.5 ≈ 48 → round to 50
When the dependency slows to 1s: 400 × 1.0 = 400 needed > 50 → 350 RPS rejected fast
That rejection is the feature: it caps the damage at the bulkhead.
When to deviate: Go adaptive first when traffic is highly variable (10× diurnal swing) or the dependency's capacity changes with autoscaling — static thresholds will be wrong half the day.
3.5 Fault Line 5: Library vs Mesh (Ownership)#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| In-process library (Resilience4j, Polly, custom) | Rich semantics: typed fallbacks, per-method policies, access to domain context | Polyglot drift: 4 languages, 4 libraries, 4 metric formats; upgrades take quarters | Every product team maintains it; platform can't see the fleet |
| Sidecar / service mesh (Envoy, Istio, Linkerd) | Uniform timeouts, retry budgets, outlier ejection, concurrency caps; language-agnostic; central config | No domain knowledge: cannot serve a cached price; adds ~0.5–2ms per hop and a proxy to operate | Platform team (operates mesh); latency budget |
| Both, split by concern | Mechanics uniform, semantics local | Two places to look during an incident; must avoid double retries | Platform + product, with a written contract |
The Staff default: Split by concern. Mechanics — connect timeouts, retry budgets, outlier detection, max concurrency — in the mesh, configured by platform with per-service overrides reviewed in code. Semantics — per-call timeouts tied to the deadline, fallbacks, criticality — in the service, via a thin, standard library that emits standard metrics.
The critical rule: Retries must be configured in exactly one of the two. The most common real-world amplification bug is a library that retries 3× sitting on top of a sidecar that also retries 3× — 16 attempts per call, and neither team knows.
🎯 Staff Move: "The mesh owns how we fail. The service owns what the user sees when we do. And retries live in the mesh only — I'd have the service library refuse to retry if it detects a sidecar."
When to deviate: No mesh and fewer than ~20 services in 1–2 languages: a standard library is fine. Say what would trigger the move (a third language, or the first incident caused by library version drift).
4. Failure Modes & Operational Reality#
The breaker state machine itself, for reference — then the ways the whole system actually fails.
4.1 Retry Storm → Metastable Failure — Full Timeline#
The most common way fault-tolerance features cause an outage.
Setup: 4-layer call graph (gateway → BFF → orders → inventory), each layer retries 3x
on timeout. inventory capacity: 12,000 RPS. Normal load: 8,000 RPS.
t=0: inventory DB failover; inventory p99 goes 30ms → 2s for 20 seconds
t=+1s: orders' 500ms timeouts fire; orders retries 3x → inventory offered 32,000 RPS
t=+2s: BFF's calls to orders time out; BFF retries 3x → orders offered 4x load,
each of which fans out 4x to inventory → inventory offered up to 128,000 RPS
t=+20s: DB failover completes. inventory is healthy — but offered 10x its capacity
t=+25s: inventory queues grow; every request waits > timeout; all work is wasted
goodput → ~0 even though every inventory host is "up"
t=+5min: on-call scales inventory 2x; new hosts are immediately saturated
t=+18min: someone disables retries in BFF via config push; load drops to 2x
t=+22min: inventory drains queues and recovers
Trigger lasted 20 seconds. Outage lasted 22 minutes.
Detection: rpc.retry_ratio (retries ÷ first attempts) jumping from ~0.01 to > 1.0; inventory.goodput diverging from inventory.throughput; deadline_exceeded_total rising at the callee while CPU is pegged.
Mitigation (now): Kill-switch retries fleet-wide via mesh config (should take < 60s). Enable aggressive load shedding at inventory (drop requests whose deadline has passed — this alone often breaks the loop).
Prevention: One retry layer; retry budgets; callee drops expired-deadline work; a documented, tested fleet-wide "retries off" switch.
Owner: Platform team owns the retry policy and the kill switch. Inventory team owns shedding. The incident commander owns the decision to flip the switch — pre-approved in the runbook.
4.2 The Slow Dependency That Never Trips the Breaker#
t=0: recs-svc starts returning 200 OK in 6–9 seconds (a lock contention bug)
t=+10s: product-page's breaker is error-rate based: 0% errors → stays CLOSED
t=+15s: product-page has a 10s timeout (the library default nobody changed)
t=+20s: all 200 request threads blocked on recs; product-page latency 9s
t=+30s: LB health checks to product-page time out → instances marked unhealthy
t=+45s: product-page fully down. recs is "healthy" by every dashboard.
Detection: breaker.slow_call_rate, threadpool.active / threadpool.max per bulkhead, dependency.latency_p99 vs its SLO.
Mitigation: Emergency config: drop recs timeout to 300ms, force recs breaker open.
Prevention: Slow-call threshold as a first-class trip condition (Resilience4j supports slowCallDurationThreshold / slowCallRateThreshold); timeouts derived from p99.9; bulkhead for every non-critical dependency so exhaustion is local; a lint rule that fails the build on library-default timeouts.
Owner: product-page team (their timeout, their bulkhead). Platform owns the lint rule.
🎯 Staff Move: "Error-rate breakers protect you from dead dependencies, which were already cheap. Slow-call breakers protect you from sick ones, which are the ones that kill you."
4.3 Correlated Breakers — The Synchronized Herd#
Why it happens: Identical thresholds, identical open durations, identical failure signal. Independent breakers become a synchronized oscillator.
Detection: breaker.state_transitions_total spiking at a fixed period; callee load graph showing a sawtooth with period = open duration.
Prevention: Jitter the open duration (±20%); limit half-open probes per instance (1–5) and ramp traffic after close (10% → 25% → 50% → 100% over 30–60s); prefer continuous client-side adaptive throttling so there is no cliff; callee warms caches before advertising healthy.
Owner: Platform (breaker defaults). Callee team (warm-up before readiness).
4.4 The Fallback That Became the Outage#
t=0: Redis cache cluster for user profiles loses a shard
t=+1s: profile-svc breaker on Redis opens; fallback = "read from Postgres"
t=+2s: Postgres, sized for a 3% cache-miss rate (~600 QPS), receives 20,000 QPS
t=+10s: Postgres connection pool exhausted; p99 30s; replicas fall behind
t=+30s: Every service that reads Postgres directly (billing, auth) starts failing
Cache incident became a database incident became a company incident.
Detection: fallback.served_total{dependency} by fallback target; db.connections_in_use / max.
Prevention: Fallback targets must be sized for fallback load or rate-limited (e.g., fallback to DB capped at 2× normal miss traffic, the rest served a static default); request coalescing on miss; see Distributed Caching for thundering-herd controls.
Owner: profile-svc team — they chose the fallback. The DB team should have been consulted on the fallback's load; that consultation is the organizational fix.
4.5 The Breaker That Hid a Bug#
A deploy introduces a bug that makes 15% of calls to the tax service fail with 400 (bad request). The breaker counts 4xx as failures, opens repeatedly, and the fallback (a flat "estimated tax") serves quietly for two weeks. Finance discovers under-collected tax at month-end.
Lessons: Don't count caller-caused 4xx as dependency failures. Alert on fallback.served_ratio sustained above baseline for > 15 minutes. A fallback serving for days is an incident, not a success.
Owner: Calling team for classification; product/finance for the decision that estimated tax was ever an acceptable fallback.
4.6 Silent Open — Nobody Knows the Breaker Is Open#
The breaker's whole purpose is to hide failures from users. Which means it also hides them from dashboards that measure user-facing error rate. Mandatory metrics:
breaker_state{caller, dependency} gauge: 0 closed, 1 half-open, 2 open
breaker_open_seconds_total{caller, dependency} counter
fallback_served_total{caller, dependency, kind} counter
fallback_last_served_timestamp{dependency} gauge — "days since fallback ran"
Alert: any CRITICAL-tier dependency breaker open > 60s → page caller owner and dependency owner. Any DEGRADABLE breaker open > 15 min → ticket + notify product.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Retry storm / metastable overload | rpc.retry_ratio > 0.5; goodput ≪ throughput | Whole call graph below the trigger | Fleet retry kill switch; shed expired-deadline work | Platform (switch), callee (shedding) |
| Slow dependency, breaker closed | threadpool.active/max → 1.0; slow_call_rate | Caller and all its callers | Force-open breaker; cut timeout | Calling team |
| Synchronized breaker oscillation | Sawtooth callee load at period = open duration | Callee + all callers | Jitter, probe limits, ramp-up | Platform |
| Fallback overloads its target | fallback.served_total ↑ + target saturation | Fallback target and its other clients | Cap fallback rate; static default | Calling team + target owner |
| Breaker hides a correctness bug | fallback.served_ratio above baseline for days | Business correctness (revenue, compliance) | Alert on sustained fallback; exclude 4xx | Calling team + product |
| Breaker flapping on low volume | state_transitions_total high, low RPS | One caller-dependency pair | Raise min volume; aggregate at mesh | Calling team |
| Mesh config push breaks timeouts fleet-wide | deadline_exceeded_total ↑ across many services at once | Entire fleet | Roll back config; staged config rollout | Platform |
| Load shedder rejects critical traffic | shed_total{criticality=CRITICAL} > 0 | Revenue paths | Fix criticality tagging; reserve capacity for CRITICAL | Callee team |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | Treats it as "add breakers" | Classifies dependencies by criticality; names three intents | Questions whether dependencies belong on the synchronous path; sets org taxonomy |
| Timeouts | Picks round numbers | Derives from p99.9 and propagated deadlines | Makes deadline propagation a framework default; audits for library defaults |
| Retries | Count + backoff | One layer, idempotent only, jitter, budget; computes amplification | Fleet-wide retry governance, kill switch, amplification in incident reviews |
| Failure analysis | Dependency down | Dependency slow; metastable states; correlated breakers | Correlated failure across teams; cell architecture to bound blast radius |
| Fallbacks | "Return a default" | Per-dependency, product-approved, sized, and tested | Game-day program that proves fallbacks run; "days since last fallback" SLO |
| Ownership | Each team configures its library | Mechanics/semantics/shedding split with named owners | Funds platform, retires competing libraries, writes the standard |
| Cost awareness | Not discussed | Mentions mesh latency and resource overhead | Prices outage minutes vs resilience investment; decides where not to invest |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Leads with slowness, not failure | "The dependency that returns 200 in 8 seconds is more dangerous than the one that's down." |
| Quantifies amplification | "Three retries at four layers is 256 attempts at the bottom. I'll retry at one layer under a 10% budget." |
| Protects the callee independently | "I can't trust every caller to have a breaker, so the callee sheds by criticality." |
| Makes fallbacks a product decision | "Stale price for up to 5 minutes — that needs product and finance sign-off, not mine." |
| Anticipates correlated behavior | "500 instances with identical breakers will reopen together; I'll jitter and ramp." |
| Splits ownership cleanly | "Mesh owns mechanics, service owns semantics, callee owns shedding." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| 15 minutes on the state machine | Table stakes presented as depth |
| "Retry 3 times" at every layer, unexamined | The most common cause of real cascading outages |
| No timeouts, or a single global timeout | Breakers can't trip on calls that never finish |
| Fallback = "call the database instead" | Moves the outage to a component sized for 3% of the load |
| Breaker as the only protection | Ignores callers you don't control |
| No metrics for breaker state | A silently open breaker is an invisible outage |
5.4 Common False Positives#
- Knowing Hystrix configuration keys ≠ understanding fault tolerance. Configuration trivia signals you've used a library, not that you've debugged a retry storm.
- Drawing a service mesh ≠ solving ownership. The mesh can't write a fallback.
- "We'll use chaos engineering" ≠ resilience. Chaos without hypotheses, blast-radius limits and owners is just random outages.
- A distributed, consistent breaker state store ≠ sophistication. It adds a dependency to your fault-tolerance layer — the one layer that must not have dependencies.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Criticality, depth, degraded semantics, deadline |
| Entities & request context | 3–5 min | Deadline/criticality/attempt headers |
| Layered architecture | 5–10 min | Gateway, caller, mesh, callee, observability — with owners |
| Retries & amplification | 10–17 min | The math, the budget, one layer |
| Timeouts & slow calls | 17–22 min | Derivation, slow-call detection, bulkhead sizing |
| Correlated recovery & shedding | 22–30 min | Jitter, ramp, callee protection |
| Fallbacks & ownership | 30–38 min | Criticality table; product sign-off; mesh vs library |
| Failure scenario / pivot | 38–43 min | Whatever the interviewer throws |
| Wrap-up | 43–45 min | Next steps; biggest remaining risk |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Direction |
|---|---|---|
| "Design the breaker library itself" | Can you do mechanics without losing the plot? | State machine in 90s, then sliding-window counting, then say why thresholds and fallbacks are the hard part |
| "The dependency is our own database" | When NOT to use a breaker | Timeout + pool limit; breaker on the only DB is a self-inflicted outage |
| "What if callers are external partners?" | Callee protection | Shed by client and criticality; rate limit (Rate Limiting); Retry-After contract |
| "Make it multi-region" | Blast radius | Regional isolation; don't fail over on a single breaker signal; failover capacity math |
| "How do you test this?" | Operational maturity | Fault injection in staging, then production game days with abort criteria |
| "What if the breaker state should be shared?" | Resisting unnecessary coupling | Local state + server-side signals; shared state is a new SPOF |
6.3 What to Deliberately Skip#
- Library selection debates. One sentence.
- Distributed breaker state. Mention and reject.
- Exact sliding-window data structure. Ring buffer of buckets; move on unless asked.
- Every possible fallback type. Show the criticality table with three rows; that's enough.
6.4 Follow-Up Questions to Expect#
- "What's the difference between a circuit breaker and a rate limiter?" — Breaker reacts to the dependency's health; limiter enforces a client's policy regardless of health.
- "How do you pick the breaker thresholds?" — Minimum volume from traffic (≥ 20–100 calls per window), slow-call duration from p99.9, error rate starting at 50%, then validate with fault injection.
- "Should 4xx errors trip the breaker?" — No, except 429. They're caller errors, not dependency health.
- "How do you know a fallback works?" — It serves production traffic regularly: game days or a small permanent percentage.
- "Where do retries live if we have a mesh and a library?" — One place — the mesh — with the library's retries disabled.
- "How does the callee tell callers to back off?" — 503/429 with Retry-After, and gRPC status + pushback metadata; callers' budgets honor it.
- "What happens when half the fleet has a new breaker config and half doesn't?" — Staged config rollout with canary and automatic rollback on deadline-exceeded or fallback rate.
7. Active Drills#
Drill 1: The Opening#
Prompt: "Our site went down last week because the recommendations service got slow. Design fault tolerance so it doesn't happen again."
Staff Answer
"First, the fact that a slow optional dependency took down the site tells me three things were missing: a timeout short enough to matter, a bulkhead isolating recs from the rest of the request threads, and a fallback for when recs is unavailable. Before designing, I want to classify every dependency on the page: critical, degradable, or optional. Recs is optional — the page renders without it.
For recs specifically: timeout derived from its p99.9 — say 300ms, bounded by the remaining request deadline — a bulkhead of ~20 concurrent calls, a breaker that trips on slow calls as well as errors, and a fallback that omits the section or shows a static popular list. Product signs off on that.
But I'd fix the class, not the instance: deadline propagation on every hop, retries at one layer under a budget, and every callee shedding load by criticality. Then a game day that forces the recs breaker open in production for 10 minutes to prove the page survives."
Why this is L6:
- Diagnoses why the architecture allowed a slow optional call to become fatal
- Leads with slowness and bulkheads, not the state machine
- Generalizes the fix to the class of dependencies and proves it with a game day
What L7 adds:
- Asks whether recs belongs on the synchronous path at all — precomputing recommendations per user into a cache removes the dependency entirely
- Proposes an org-wide criticality registry so every new dependency declares its tier and gets default protection from the mesh
- Adds "which optional dependency could take down a critical page?" to architecture review
❌ Common L5 Trap
"I'll add a circuit breaker around the recs client with a 50% error threshold and a 30-second open duration."
Why this misses: Recs was slow, not erroring. An error-rate breaker with a default 10s timeout would not have tripped. And "breaker opens" doesn't say what the page shows.
Drill 2: Retry Math#
Prompt: "Every service in our call graph retries 3 times with exponential backoff. Is that a problem?"
Staff Answer
"Yes, if the graph is deeper than one hop. With 3 retries — 4 attempts — at each of N layers, the bottom service can see up to 4ᴺ attempts per user request: 16 at two layers, 256 at four. Worse, the amplification activates precisely when the bottom service is struggling, turning a 20-second brownout into a self-sustaining overload.
The fix has three parts. One: retry at a single layer — the one nearest the failing dependency that knows the call is idempotent — and propagate a 'don't retry' flag so upper layers fail through. Two: replace counts with a budget: retries may add at most 10% over successful calls in a rolling 10-second window. When 1% of calls fail, every failure gets retried; when 60% fail, almost none do. Three: full jitter on backoff, and no retries once the remaining deadline can't fit another attempt.
Net effect: worst-case load at the bottom goes from 256× to about 1.1×."
Why this is L6:
- Computes the amplification rather than asserting "it's bad"
- Distinguishes retries that help (rare failures) from retries that hurt (widespread failures) and designs a mechanism that does both automatically
- Connects retries to deadlines
What L7 adds:
- Moves retry policy into the mesh so no service can opt into nested retries
- Adds
retry_ratioto fleet dashboards and "retry amplification" to the incident template - Builds and tests a fleet-wide retry kill switch with a < 60s propagation SLO
❌ Common L5 Trap
"Exponential backoff solves it — retries get spread out over time."
Why this misses: Backoff spreads retries in time; it does not reduce their number. Total work is still multiplied, and with synchronized clients and no jitter, backoff can make retries arrive in waves.
Drill 3: Make It Concrete — Timeouts#
Prompt: "What timeout would you put on the call from product-page to pricing?"
Staff Answer
"It's derived, not picked. Two inputs: pricing's latency distribution and the remaining deadline. If pricing's p99.9 is 120ms, I'd set ~180ms (1.5×) so we cut the pathological tail without failing healthy calls — at most ~0.1% false timeouts. Then cap it by the remaining deadline: if the gateway gave us 1,500ms and 1,200ms remain after our own work, 180ms fits; if only 100ms remain, we use 100ms, or skip the call and use the fallback if 100ms isn't enough to be useful.
Separately, a connect timeout of ~100–250ms, because connection establishment failing is a different failure than a slow response. And I'd revisit the timeout when pricing's p99.9 changes — an alert when timeout rate exceeds 0.5% tells us it's stale."
Why this is L6:
- Derives the number from data and the deadline, with an explicit false-timeout rate
- Separates connect and request timeouts
- Treats the timeout as a maintained value with an alert
What L7 adds:
- Makes the RPC framework compute
min(configured, remaining_deadline)automatically, so the deadline is never violated regardless of team config - Publishes a per-service latency SLO registry that timeouts are generated from
Drill 4: The Dependency Is Down — What Does the User See?#
Prompt: "Pricing is down. The breaker is open. What does the product page show?"
Staff Answer
"Pricing is degradable, not optional — we can't show a product with no price, and we can't let someone check out at a wrong one. So: show the last cached price if it's under 5 minutes old, which covers the vast majority of brownouts; beyond that, show the product with 'price unavailable' and disable Add to Cart. Checkout independently re-prices from the authoritative service and fails clearly if it can't — the cached price is display-only and never honored.
That 5-minute number isn't mine to choose. Product and finance own it, because a stale price during a sale could be off by 50%. I'd bring them a proposal with the tradeoff: every minute of staleness allowed covers more of our historical pricing brownouts, and increases the risk of a display mismatch."
Why this is L6:
- Fallback semantics differ for display vs transaction
- Explicit staleness cap and explicit owner
- Clear failure (disable purchase) instead of fake success
What L7 adds:
- Makes "display-only stale, authoritative at transaction" a documented pattern for every money-adjacent fallback
- Quantifies the tradeoff in revenue: minutes of pricing brownout per quarter × conversion loss vs mismatch risk
Drill 5: Callee Can't Trust Callers#
Prompt: "Our inventory service has 40 callers. Some are legacy services with aggressive retries and no breakers. How do you protect inventory?"
Staff Answer
"Inventory has to protect itself; client behavior is a hope, not a guarantee. Four layers:
- Drop expired work: every request carries a deadline; if it's passed on arrival or dequeue, drop it with no work. Under overload, often 30–50% of queued work is already dead.
- Shed by criticality: requests tagged SHEDDABLE (batch, prefetch) are rejected first when queue wait exceeds ~50ms; CRITICAL (checkout) last. Reserve ~20% of concurrency for CRITICAL.
- Adaptive concurrency limit: cap in-flight requests at a limit that tracks latency (gradient/AIMD); excess is rejected with 503 + Retry-After in microseconds.
- Per-caller fairness: a per-caller concurrency quota so one legacy caller's retry storm consumes only its own share.
Rejection must be cheap — under 1–5% of the cost of serving — or shedding can't save you. Then I'd put the legacy callers behind the mesh so their retries get budgeted without code changes."
Why this is L6:
- Designs callee protection that works regardless of caller behavior
- Uses deadline, criticality and per-caller fairness together
- Names the cost-of-rejection requirement
What L7 adds:
- Makes criticality tagging mandatory at the gateway so every request in the company has one
- Uses mesh onboarding as the lever to fix 40 callers without 40 code changes
Drill 6: The Synchronized Herd#
Prompt: "We have 600 instances of the checkout service, each with a breaker on the fraud service. After a fraud outage, fraud recovered but immediately fell over again, three times. Why?"
Staff Answer
"The breakers synchronized. All 600 saw the same failure, opened within a second of each other, and — with identical 30-second open durations — went half-open together. Each sent its probes plus pent-up retries, fraud went from ~0 to full load plus a burst in one second with cold caches, failed, and every breaker re-opened. The period of the sawtooth equals the open duration; that's the fingerprint.
Fixes: jitter the open duration ±20–30% to spread half-open over ~15 seconds; cap probes per instance at 1–3; after closing, ramp from 10% to 100% of traffic over 60 seconds; and have fraud warm caches and only report ready when warm. Longer term, move from binary breakers to client-side adaptive throttling where rejection probability tracks the accept rate — there's no cliff to synchronize on."
Why this is L6:
- Recognizes emergent behavior from independent components
- Identifies the diagnostic fingerprint (sawtooth period = open duration)
- Fixes both sides: caller ramp and callee warm-up
What L7 adds:
- Adds "recovery load" to capacity planning: every critical service must survive 1.5–2× normal load for 60 seconds after an outage
- Makes jitter and ramp-up non-optional platform defaults
Drill 7: Library or Mesh?#
Prompt: "We have 250 services in Java, Go, Node and Python. Some use Resilience4j, some Polly-style ports, some nothing. What should we standardize on?"
Staff Answer
"Split by concern. Mechanics that don't need domain knowledge — connect timeouts, retry budgets, outlier ejection, max concurrency, circuit breaking as concurrency caps — move into the mesh, where one team configures them uniformly across all four languages. Semantics that need domain knowledge — per-call deadlines, fallbacks, criticality — stay in-process behind a thin standard library per language that only does timeouts, fallbacks, and standard metrics. It explicitly does not retry.
Migration: shadow first — mesh collects metrics without enforcing. Then enable mesh timeouts and budgets for tier-3 services, then tier-2, then critical. Disable library retries in the same change as mesh retries go live, per service, so there's never a window of double retries. Success metric: fleet retry_ratio p99 under 0.1 and zero services with nested retry policies."
Why this is L6:
- Splits by what needs domain knowledge, not by preference
- Plans the migration to avoid double retries
- Defines a measurable success criterion
What L7 adds:
- Prices the options: a mesh adds ~1 vCPU and 50–100MB per sidecar-heavy pod fleet-wide versus N engineer-quarters per language per year maintaining libraries
- Sets a deprecation date for legacy libraries and owns the exception process
Drill 8: Change a Threshold Without an Outage#
Prompt: "You want to lower the default request timeout in the mesh from 15s to 2s for all services. How do you roll it out?"
Staff Answer
"This is one of the most dangerous changes you can make: any service with a legitimately slow endpoint breaks instantly. So: first, measure — pull per-route p99.9 for every service from the mesh; any route above 1.5s is an exception candidate. Second, shadow: emit would_have_timed_out_total for 1 week at the new value. Third, notify owners of routes with nonzero shadow timeouts and give them an override path with a sunset date. Fourth, canary: enforce for 5% of instances of tier-3 services, watch deadline_exceeded_total and error rate, then widen by tier over 2–3 weeks. Automated rollback if any service's error rate rises > 0.5 percentage points.
Owner: platform proposes and operates; each service owner signs off on their overrides; the change is announced with a date, not a surprise."
Why this is L6:
- Treats a config value as a production change with shadow → canary → enforce
- Uses data to find exceptions before they break
- Names owners and the override process
What L7 adds:
- Makes shadow-mode evaluation a built-in capability of the config system, so every future policy change gets it for free
- Tracks exception count as a health metric for the standard itself — if 30% of services need overrides, the default is wrong
Drill 9: Cost#
Prompt: "Leadership asks: we spend a lot on resilience features. Is it worth it?"
Staff Answer
"Frame it as outage minutes avoided per dollar. Last year's incidents tell us the baseline: suppose 6 cascading incidents averaging 35 minutes, with revenue at risk of ~$40K/minute at peak — ~$8M exposure. The resilience work — mesh, shedding, game days — costs roughly 3–4 platform engineers plus ~5% more compute for sidecars and headroom, maybe $1.5–2M/year. If it converts 35-minute cascades into 5-minute single-feature degradations, the return is several-fold.
But not every dependency deserves the same investment. Tier-1 paths get the full treatment and quarterly game days. Tier-3 internal tools get timeouts and mesh defaults only."
Why this is L6:
- Converts resilience into outage minutes and dollars
- Tiers the investment rather than applying it uniformly
What L7 adds:
- Tracks the metric over time (cascading-incident count and minutes per quarter) as the platform's success measure
- Identifies opportunity cost: what the platform engineers aren't building
Drill 10: Multi-Region#
Prompt: "We're active-active in two regions. If a dependency in us-east is failing, should callers fail over to us-west's copy?"
Staff Answer
"Rarely per-dependency, and never automatically on a single breaker signal. Cross-region calls add 60–80ms RTT per hop, and if every us-east caller fails over to us-west's pricing, us-west pricing sees 2× load and may fall over too — converting a regional incident into a global one. Also, dependencies often share failure causes (a bad deploy rolled to both regions), so failover moves traffic from one sick copy to another.
Default: contain failures within the region with local fallbacks. Fail over whole user traffic at the edge (DNS/anycast/GSLB) when a region is broadly unhealthy, with a human or a well-tested automated decision, and only if the target region has headroom — which means each region runs at ≤ 50% of combined capacity or you pre-scale. Staggered regional deploys so both copies aren't broken at once."
Why this is L6:
- Recognizes failover as load transfer that can cascade
- Distinguishes dependency-level from region-level failover
- Connects failover to capacity headroom
What L7 adds:
- Designs for cells within regions so most incidents never need regional failover
- Prices the headroom: N+1 regions at 50–67% utilization is a real budget line that leadership must choose
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Cascade#
Context: It's the first hour of a major sale. Traffic is 6× normal. The search service's p99 went from 80ms to 1.2s, and within four minutes the home page, category pages and cart are all erroring at 30–40%. Checkout is at 12% errors. The incident commander pulls you in.
Questions to Surface First:
- Is search slow or failing? Is its CPU saturated, or is it waiting on its own dependency?
- Which callers of search have bulkheads and breakers, and are any breakers open?
- What is the fleet-wide
retry_ratioright now versus baseline? - Why is cart failing — does cart call search, or is cart sharing a thread pool or database with something that does?
Typical L5 Approach: Scales search horizontally, bumps its thread pool, and restarts unhealthy instances. Reasonable — and it treats search as the problem, when the outage is really everyone else failing because of search. New search instances saturate instantly because offered load includes retries.
Staff Approach: Contains first, fixes second. Force-opens the search breakers in non-critical callers (home page carousels, category "related searches") so those pages render with fallbacks, freeing threads immediately. Cuts retries to search to zero via mesh config. Asks search to enable shedding of SHEDDABLE traffic (autocomplete, prefetch) to reserve capacity for the actual search results page. Then scales search, now that scaling can actually help.
Principal Approach: Asks why a sale — a scheduled, predictable event — found search at 6× load without a pre-scaled, load-tested posture. Institutes an event-readiness review for tier-1 paths (load test at 1.5× forecast, verified fallbacks, pre-approved force-open list). Looks at the dependency graph that let cart fail because of search, and makes "no optional dependency shares a resource pool with a critical path" an architecture-review rule.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Mesh: retries to search → 0. Force-open search breakers in all OPTIONAL callers. Verify fallbacks are serving (fallback_served_total ↑). |
| Triage | Why is cart failing? Likely a shared thread pool or shared DB connection pool with a search-calling path. Check bulkhead saturation per dependency. |
| Quick fix | Search sheds SHEDDABLE (autocomplete, prefetch) → frees ~40% of its capacity. Scale search 2×. Close breakers gradually, OPTIONAL callers last. |
| Guardrails | Keep retries off until search.p99 < 200ms for 10 min. Ramp breaker close 10% → 50% → 100% over 5 min. |
| Post-mortem | Missing bulkhead in cart; no event load test; autocomplete tagged CRITICAL by default; retries at 3 layers. |
Metrics to Watch: search.latency_p99, rpc.retry_ratio{dest=search}, bulkhead.active/max{dependency=search}, breaker_state{dependency=search}, fallback_served_total, checkout.error_rate.
Organizational Follow-up: Event-readiness checklist owned by the SRE lead; criticality tagging audit for search's callers; bulkhead requirement for every OPTIONAL dependency added to the service template.
Ownership Question: "Who decides to force-open breakers for other teams' services during an incident?" Staff answer: The incident commander, from a pre-approved list of OPTIONAL dependencies whose fallbacks product has already signed off on. Forcing open a DEGRADABLE or CRITICAL breaker needs the owning team's on-call. The list lives in the runbook, not in someone's head.
Key Takeaway: "In a cascade, contain first — free the callers — then fix the dependency. Scaling a service whose offered load includes a retry storm is pouring water into a bucket with no bottom."
What clears the Staff bar:
- Treats callers' resource exhaustion as the outage, not the slow dependency
- Turns off retries before scaling
- Has a pre-approved force-open list and uses it
Deep Dive 2: The Silent Fallback#
Context: Finance reports that shipping-cost revenue is 9% below forecast for the month. Investigation shows the checkout service's breaker on the shipping-rates service has been opening intermittently for 23 days, and the fallback — a flat $5.99 rate — has served ~18% of orders. No alert fired. User-facing error rate was 0%.
Questions to Surface First:
- Why did the breaker open — is shipping-rates actually unhealthy, or is the breaker counting something it shouldn't (4xx, a new slow endpoint)?
- Who approved "flat $5.99" as a fallback, and was it approved as a brief-outage measure or indefinitely?
- Why is there no alert on fallback rate?
- Are there other breakers in the same state right now?
Typical L5 Approach: Fixes the root cause in shipping-rates (it turned out a new carrier integration made 20% of calls take 2.5s, over the 2s slow-call threshold), closes the ticket. Correct, but leaves the systemic hole: any breaker can serve a fallback indefinitely without anyone knowing.
Staff Approach: Fixes the root cause, then fixes observability and policy. Adds
fallback_served_ratioalerting per dependency with a duration threshold. Adds a "max fallback duration" to each fallback's definition: flat-rate shipping is acceptable for 30 minutes, after which it pages. Audits all breakers fleet-wide for sustained open or fallback states — finds two more.
Principal Approach: Reframes fallbacks as business policy with an expiry. Every fallback in a revenue- or compliance-adjacent path gets a registered owner in product/finance, a maximum duration, and an estimated $/hour cost. The fallback registry becomes an input to monthly business reviews. The org learns that "0% errors" is not the same as "working."
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Confirm current fallback rate. If shipping-rates is healthy enough, raise slow-call threshold for the new carrier path to stop spurious opens. |
| Triage | Separate spurious opens (threshold wrong) from real ones. Quantify revenue impact by day. |
| Quick fix | Per-route slow-call thresholds; new carrier path gets its own timeout and bulkhead. |
| Guardrails | Alert: fallback_served_ratio > 2× baseline for 15 min → ticket; > 5% for 30 min on revenue paths → page. |
| Post-mortem | No max duration on fallbacks; no alert on fallback rate; breaker treated as success because errors were 0. |
Metrics to Watch: fallback_served_total{dependency=shipping_rates}, breaker_open_seconds_total, fallback_last_served_timestamp, orders.shipping_revenue_per_order.
Organizational Follow-up: Fallback registry: dependency, fallback kind, owner, max duration, estimated cost/hour. Quarterly review with finance for revenue-adjacent entries.
Ownership Question: "Who owns the $5.99 fallback?" Staff answer: Engineering owns the mechanism; the shipping product manager owns the decision to show $5.99 and for how long. Nobody could have answered this question before the incident — that's the actual root cause.
Key Takeaway: "A breaker's job is to hide failures from users. Your job is to make sure it doesn't hide them from you."
What clears the Staff bar:
- Treats a sustained fallback as an incident, not a success
- Adds duration limits and business owners to fallbacks
- Audits the fleet for the same class of problem
Deep Dive 3: Onboarding a Large Internal Caller#
Context: The data-platform team wants to call the user-profile service from a new batch enrichment job: 50M lookups nightly, ideally in 2 hours (~7,000 RPS). Profile serves 15,000 RPS of interactive traffic at peak with ~35% headroom at night. The profile team asks you to review.
Questions to Surface First:
- Can the batch job read from a replica, a snapshot, or a CDC-fed copy instead of the online service?
- If it must call online: what criticality will its requests carry, and what does it do when shed?
- What's profile's actual nighttime headroom in RPS, and what happens if an interactive spike coincides with the batch?
Typical L5 Approach: Adds a rate limit for the batch caller at 7,000 RPS and a breaker in the batch job. Works on a normal night — and on a bad night the batch still consumes 7,000 RPS of capacity while interactive users suffer.
Staff Approach: Prefers moving the batch off the online path (a nightly snapshot export or a CDC-fed read store). If it must be online: tag all batch requests SHEDDABLE, give the batch a per-caller concurrency quota, and use adaptive client-side throttling so it slows automatically when profile rejects. The batch deadline (2 hours) is soft — it's fine if it takes 3 hours on a busy night. Interactive traffic always wins.
Principal Approach: Establishes a rule for the company: online services are not bulk data sources. Batch consumers get data through the platform's snapshot/CDC path, which the data platform funds. Prevents the next ten teams from making the same request, and removes a whole class of capacity conflicts.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (review) | Ask for an offline path first. Estimate profile's nighttime headroom: 15K peak × 0.35 ≈ 5K RPS — less than the 7K asked. |
| Triage | If online: batch gets max 3,000 concurrent RPS quota, SHEDDABLE tag, adaptive throttle. Expected runtime ~4.5h. |
| Quick fix | Batch job honors 503 + Retry-After; no retries beyond a 5% budget. |
| Guardrails | Profile sheds SHEDDABLE first when queue wait > 30ms. Alert profile on-call only if CRITICAL traffic is shed. |
| Post-mortem-proofing | Load test with batch + synthetic interactive spike before go-live. |
Metrics to Watch: profile.shed_total{criticality}, profile.latency_p99{caller}, batch.throughput, batch.throttled_ratio.
Organizational Follow-up: A documented "bulk access" pattern; online services can reject bulk callers who don't use it.
Ownership Question: "If the batch doesn't finish by morning, whose problem is it?" Staff answer: The data-platform team's. The profile team's commitment is interactive SLO; the batch's completion time is the batch owner's risk. Writing that down before go-live is the review's main output.
Key Takeaway: "Criticality tags turn 'who wins under contention' from an incident-time argument into a design-time decision."
What clears the Staff bar:
- Prefers removing the coupling over protecting it
- Makes the batch shed first by construction
- States who bears the risk of the batch running late
Deep Dive 4: Post-Mortem — The Mesh Config That Took Down Everything#
Context: A platform engineer pushed a mesh config change intended to lower the default retry count. A typo set per_try_timeout to 10ms fleet-wide. Within 90 seconds, ~70% of inter-service calls were timing out. Rollback took 14 minutes because the config pipeline itself depended on a service that was now failing. You're running the post-mortem.
Questions to Surface First:
- Why did a fleet-wide change go to 100% at once?
- Why did no validation catch a 10ms timeout?
- Why did the rollback path depend on the data plane it controls?
- What was the blast radius by tier, and did any CRITICAL path have a safeguard?
Typical L5 Approach: Adds input validation for timeout values (min 50ms). Necessary, insufficient — the next bad config will be a different field.
Staff Approach: Treats config as code with a deployment pipeline: schema validation plus semantic checks (no timeout below observed p50 for that route), shadow evaluation, staged rollout by tier and region with automatic rollback on
deadline_exceeded_total, and a rollback path that works when the mesh is broken (control plane talks to proxies over a path that doesn't traverse the proxies' own routing, and the last-known-good config is retained locally).
Principal Approach: Recognizes the resilience platform is itself the largest correlated-failure risk in the company — one config reaches every service. Requires cell-based rollout for all platform config (cell 1 → wait → cell 2…), tracks "fleet-wide changes per quarter" as a risk metric, and gives the platform team an error budget of its own. The fault-tolerance layer must be more conservative than anything it protects.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | (During incident) Revert via out-of-band path; if unavailable, restart proxies with baked-in last-known-good. |
| Triage | Timeline of config propagation; identify the dependency loop in the rollback path. |
| Quick fix | Semantic validation; rollback path independent of mesh routing. |
| Guardrails | Staged rollout: 1% of one cell → 1 cell → region → fleet, each gated on automated health checks, ≥10 min dwell. |
| Post-mortem | Themes: blast radius of platform config, circular dependency in recovery tooling, lack of shadow mode. |
Metrics to Watch: mesh.config_version{instance} distribution, deadline_exceeded_total by service, control_plane.push_errors.
Organizational Follow-up: Platform config changes follow the same change-management tier as a database schema migration. Quarterly "break the control plane" game day.
Ownership Question: "Who approves a fleet-wide mesh config change?" Staff answer: Platform owns it, but the pipeline approves it — automated staged rollout with health gates. A human approving a diff is not a safeguard against a typo that looks plausible.
Key Takeaway: "The component that protects every service is the component that can break every service. Roll it out like it."
What clears the Staff bar:
- Identifies the circular dependency in recovery tooling
- Proposes a pipeline, not a validation rule
- Applies blast-radius thinking to the resilience platform itself
Deep Dive 5: Multi-Region Expansion#
Context: The company is going from one region to two, active-active. The resilience design so far is regional. Leadership asks: "If a service fails in one region, can we just send its traffic to the other?"
Questions to Surface First:
- What's each region's utilization at peak? Can either absorb the other's load?
- Are deploys staggered across regions, or can a bad release break both?
- Which data is regional and which is global — does cross-region failover of a service also mean cross-region data access?
- What is the cross-region RTT, and how many hops would a cross-region failover add to a request?
Typical L5 Approach: Configures the mesh to fail over to the other region's instances when local instances are unhealthy. Sounds robust; creates a path by which one region's failure doubles load on the other region's copy — correlated global failure.
Staff Approach: Contain within region; fail over at the edge for whole-user traffic, not per-dependency. Per-dependency cross-region failover only for a small allowlist of stateless, idempotent read services with verified headroom, and capped (e.g., ≤ 20% of the remote capacity). Staggered deploys (region A, bake 1 hour, region B) so both copies are rarely broken at once. Capacity: each region provisioned to take ~100% of global peak for failover, or explicitly accept brownout with shedding.
Principal Approach: Decides the company's failure-domain strategy: regions as blast-radius boundaries with cells inside them. Prices the headroom — two regions at 50% utilization cost ~2× the compute of one at 100% — and puts that choice to leadership as an availability-vs-cost decision with numbers. Standardizes "failover is an edge decision" so 200 service teams don't each invent cross-region fallbacks.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Design | Regional isolation by default; edge (GSLB/anycast) moves users between regions. |
| Capacity | Each region at ≤ 50% of combined peak, or explicit shedding plan for the failover case. |
| Deploy safety | Region-staggered deploys with 1-hour bake; mesh config staged per region. |
| Allowlist | Cross-region dependency failover only for stateless reads with remote headroom, capped at 20% of remote capacity. |
| Testing | Quarterly regional evacuation drill at off-peak, then at peak with abort criteria. |
Metrics to Watch: region.utilization_pct, cross_region.request_ratio, gslb.traffic_split, deadline_exceeded_total{region}.
Organizational Follow-up: Region evacuation runbook owned by SRE; product sign-off on what degrades during evacuation.
Ownership Question: "Who decides to evacuate a region?" Staff answer: The incident commander, using criteria written in advance (e.g., > 25% error rate on tier-1 for > 5 min and no mitigation in progress). Automation can recommend; a human confirms unless the tested automation has earned trust through drills.
Key Takeaway: "Failover is load transfer. If the destination can't take the load, you've built a mechanism for global outages."
What clears the Staff bar:
- Treats cross-region failover as a capacity decision
- Moves failover to the edge, whole-user granularity
- Staggers deploys to break correlation
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Explain why a slow dependency is more dangerous than a dead one, using Little's Law with numbers
- Compute worst-case retry amplification for an N-layer graph and design a retry budget that bounds it to ~1.1×
- Derive a timeout from p99.9 latency and a propagated deadline
- Size a bulkhead from RPS × latency × headroom
- Classify dependencies as critical / degradable / optional and specify a product-approved fallback for each
- Explain why callees must shed load independently of caller breakers, and how criticality ordering works
- Diagnose synchronized breaker oscillation and fix it with jitter, probe limits and ramp-up
- Split fault-tolerance ownership among platform (mechanics), calling team (semantics) and callee team (shedding)
- Explain why the resilience platform's own config is the biggest correlated-failure risk, and how to roll it out safely
The Bar for This Question#
Mid-level (L4): Knows the breaker states and can implement one. Mentions timeouts and exponential backoff. Designs for the dependency being down.
Senior (L5): Configures breakers, timeouts and retries sensibly for a single service. Mentions fallbacks and bulkheads. Misses amplification across layers, slow-call detection, correlated behavior, and who owns fallback semantics.
Staff+ (L6): Starts from dependency criticality, designs for slowness, quantifies retry amplification and bounds it with budgets, puts protection on both sides of every call, makes fallbacks a product decision with sign-off and duration limits, anticipates correlated recovery, and splits ownership between platform and product. Closes with how the system is tested (game days) and what can still go wrong. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "The Circuit Breaker Is the Least Important Part of Fault Tolerance"#
| Mechanism | Prevents | Without It |
|---|---|---|
| Timeouts / deadlines | Threads blocking indefinitely | Breakers never see failures complete |
| Bulkheads | One dependency consuming all capacity | A breaker trips after exhaustion has begun |
| Retry budgets | Load amplification | Breakers open, but upstream retries keep hammering |
| Load shedding | Callee collapse from any caller | You rely on every caller being well-behaved |
| Circuit breaker | Wasted calls to a known-sick dependency | Marginal cost: some wasted calls that timeouts already bound |
The Staff position: Build timeouts, bulkheads, budgets and shedding first. The breaker is an optimization on top that saves wasted work and gives a clean point to attach fallbacks.
Why this matters in interviews: Leading with the breaker state machine signals you've memorized the famous pattern. Leading with timeouts and amplification signals you've been paged.
10.2 "Most Fallbacks Are Untested Code That Will Fail When You Need It"#
| Evidence | Implication |
|---|---|
| Fallback paths run only during incidents — minutes per year | Bugs survive for months |
| Fallbacks often depend on something (cache, DB, config) the primary also depends on | Correlated failure |
| Fallback load is rarely capacity-tested | Fallback overloads its target |
The Staff position: A fallback that hasn't served production traffic in the last 30 days is assumed broken. Exercise them in game days or route a small permanent percentage through them.
Why this matters in interviews: Saying "and we'd verify the fallback runs monthly in production" is a single sentence that separates people who design from people who operate.
10.3 "Retries Cause More Outages Than They Prevent — At the Fleet Level"#
| Retry benefit | Retry cost |
|---|---|
| Masks ~0.1–1% transient failures per call | Amplifies sustained failures 4ᴺ× in deep graphs |
| Invisible when things work | Converts 20-second triggers into 20-minute metastable outages |
The Staff position: Retries are fine at one layer, with a budget. Unbudgeted retries at every layer are the single most common amplifier in large-scale outage post-mortems — the September 2015 DynamoDB event and the Metastable Failures research (Bronson et al., HotOS 2021) both describe retry-sustained overload.
Why this matters in interviews: It's a counterintuitive claim you can defend with arithmetic in 20 seconds.
10.4 "Don't Share Breaker State Across Instances"#
| Shared state | Local state |
|---|---|
| Faster, more accurate trip decision | Each instance learns independently in ~1–10s |
| New dependency (Redis, etcd) inside the fault-tolerance layer | No new dependency |
| Global flip = global synchronized reopen | Natural spread in open times (plus jitter) |
The Staff position: Keep breakers local. If you want a fleet-level signal, get it from the callee (503 + Retry-After, load reports) or from the mesh control plane's outlier data — not from a shared store in the hot path.
Why this matters in interviews: Interviewers often suggest shared state to see whether you'll add a dependency to your resilience layer. Declining with a reason is the signal.
10.5 "Service Mesh Breakers Aren't Circuit Breakers"#
Mesh "circuit breaking" (Envoy's max_connections, max_pending_requests, max_requests) is a set of concurrency caps, closer to a bulkhead. Outlier detection ejects hosts, not services. Neither serves a fallback. Teams that "turned on circuit breaking in Istio" often believe they have Hystrix-style protection and fallbacks when they have neither.
The Staff position: Use the mesh for caps and host ejection; you still need service-level semantics in code.
Why this matters in interviews: Precision about what the mesh does and doesn't do is a fast credibility signal.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
A Staff engineer makes one call graph survive one slow dependency. A Principal engineer notices that the company has 400 services, 3,000 dependency edges, four resilience libraries, one mesh, and no shared definition of what "critical" means — and that the next large outage will come not from a missing breaker but from a correlated event: a mesh config push, a shared database, a common library bug, a region-wide deploy. At L7, fault tolerance is a portfolio of blast-radius bets, priced in outage minutes and engineer-years, and the most leveraged decisions are the ones that make good behavior the default for teams who will never read this case study.
The Org-Level Fault Line#
Centralized resilience platform vs per-team autonomy.
| Option | What It Buys | What It Costs | Who Pays |
|---|---|---|---|
| Central platform (mesh + standard library + criticality registry) | Uniform mechanics, fleet visibility, one kill switch, enforced defaults | A single team whose config can break everything; slower for teams with unusual needs | Platform team carries correlated-failure risk; product teams lose some control |
| Per-team choice | Teams tune for their domain; no central bottleneck | Four libraries, nested retries, no fleet view, every incident rediscovers the same lessons | Every on-call rotation, in the next cascade |
| Paved road + exceptions (the L7 default) | Defaults that are right for ~85% of services; audited exceptions for the rest | Needs an exception process and someone to say no | Platform (governance), exception owners (their own risk) |
🧭 Principal Move: "I'd centralize mechanics and the criticality taxonomy, and deliberately not centralize fallbacks. The day the platform team starts writing product fallbacks is the day it becomes the bottleneck for every launch."
Cost Model#
Assumptions: mesh sidecar ~0.1–0.25 vCPU and ~50–100MB per pod at moderate load; blended compute ~$30/vCPU-month; fully loaded engineer ~$300K/year; revenue at risk scales with company size. Rough, order-of-magnitude.
| Scale | Setup | Monthly Infra | Headcount | On-call Load | Outage Exposure Avoided |
|---|---|---|---|---|---|
| Startup (20 services, 1 region) | Standard library, timeouts, one retry layer, no mesh | ~$0–500 | ~0.25 FTE (part of infra) | Shared rotation; ~1 cascade/quarter | Tens of $K per incident |
| Growth (150 services, 2 regions) | Mesh, criticality tags, callee shedding on tier-1, quarterly game days | ~$15–40K (sidecars ~1,500 pods) | 2–4 FTE platform | Platform rotation; cascades → ~1/half | $0.5–2M/yr |
| Large (1,000+ services, 3+ regions, cells) | Mesh + adaptive concurrency + cells + fallback registry + monthly game days | ~$150–400K | 8–15 FTE across platform + SRE | Dedicated rotations; most incidents single-cell | $10M+/yr |
The biggest line item at scale isn't compute — it's headroom. Running every tier-1 service at ≤ 60% utilization so it survives recovery surges and regional failover can cost 30–70% more than running at 85%. That is the decision leadership actually has to make.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse | Why |
|---|---|---|---|
| Breaker thresholds, timeouts, retry budgets | Two-way | Minutes (config) | Tune freely — with staged rollout |
| Criticality taxonomy (tiers and names) | Mostly one-way | Quarters: every service, dashboard and runbook references it | Get the tiers right early; keep them to 3–4 |
| Request context headers (deadline, criticality) | One-way | Every client and service in the fleet | Wire format outlives everyone who designed it |
| Library vs mesh as the enforcement point | Expensive two-way | 2–4 engineer-quarters per migration | Decide once you have ~3 languages or ~50 services |
| Cell architecture | One-way | Re-partitioning data and routing | Decide before a correlated outage makes it urgent |
| Specific mesh vendor | Expensive two-way | Proxy config dialects, CRDs, tooling | Keep policy in your own abstraction if you can |
The Standard I'd Write#
RFC: Inter-Service Fault Tolerance Standard v1
Scope: All synchronous RPC between production services. Asynchronous messaging is covered by the Messaging Standard.
Requirements:
- Every request MUST carry an absolute deadline and a criticality (CRITICAL / DEGRADABLE / SHEDDABLE). The gateway sets defaults.
- Every outbound call MUST have a timeout ≤ remaining deadline. Library default timeouts are prohibited (lint-enforced).
- Retries MUST occur in the mesh only, for idempotent methods, with a budget ≤ 10% of successful requests per destination. Application-level retries require an exception.
- Every service MUST drop requests whose deadline has passed and SHOULD shed SHEDDABLE traffic when queue wait exceeds its SLO.
- Every DEGRADABLE or OPTIONAL dependency MUST have a registered fallback with an owner, maximum duration, and product sign-off.
- Every tier-1 fallback MUST serve production traffic at least once per 30 days (game day or permanent sampling).
- Services MUST emit
breaker_state,fallback_served_total,retry_ratio,shed_total{criticality}.Exceptions: Filed with the platform team, reviewed by the architecture group, expire after 6 months unless renewed with data.
Success metrics: Cascading incidents (≥ 2 services affected by one trigger) per quarter; fleet
retry_ratiop99 < 0.1; % of tier-1 fallbacks exercised in 30 days = 100%; exception count trending down.
What I'd Tell the VP#
"Most of our big outages last year started as a small problem in one service and spread because of how our services retry and wait on each other. We're going to make the platform handle that by default: every request carries a deadline, retries are capped centrally, and every service can say 'no' when it's overloaded. Product teams will decide what customers see when a feature is degraded, and we'll test those decisions monthly instead of discovering them during incidents. It costs about four platform engineers and a few percent more compute. The goal is that next year's incidents stay the size of one feature, not the size of the company."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Prices resilience | "This is about $2M a year against roughly $8M of cascade exposure; the headroom is the real cost." |
| Sees correlated risk in the platform itself | "The mesh is now our largest shared failure domain; its config rolls out by cell." |
| Knows what not to centralize | "Platform owns mechanics; it will never own fallbacks." |
| Identifies one-way doors | "The request-context headers are forever — let's get them right and keep criticality to three tiers." |
| Measures the standard | "If 30% of services need exceptions, the default is wrong, not the teams." |
Staff answers that L7 interviewers find insufficient:
- A perfect per-service design with no plan for the other 399 services.
- "We'll add chaos engineering" without abort criteria, owners, or a metric that shows it's working.
- Ownership boundaries named for one system, but no view on how the platform team is funded or what it refuses to do.
Appendices
Appendix A: Mechanics in Depth#
A.1 Sliding-Window Breaker (Count-Based)#
class Breaker:
state = CLOSED
window = RingBuffer(size=100) # last 100 outcomes: SUCCESS / FAILURE / SLOW
min_calls = 20
failure_threshold = 0.50
slow_threshold = 0.50, slow_call_ms = 200
open_ms = 30_000 * uniform(0.8, 1.2) # jitter
half_open_permits = 5
def allow():
if state == OPEN and now() - opened_at >= open_ms: state = HALF_OPEN; permits = half_open_permits
if state == OPEN: return False
if state == HALF_OPEN: return permits.try_acquire()
return True
def record(outcome, latency_ms):
if outcome == CALLER_ERROR_4XX: return # not dependency health (except 429)
o = SLOW if latency_ms > slow_call_ms and outcome == SUCCESS else outcome
window.add(o)
if state == HALF_OPEN: evaluate_probe(o); return
if window.count() >= min_calls:
if window.rate(FAILURE) >= failure_threshold or window.rate(SLOW) >= slow_threshold:
state = OPEN; opened_at = now()
Why count-based vs time-based: count-based windows adapt to traffic (100 calls is 1s at peak, 1 minute at night). Time-based (e.g., 10 × 1s buckets) gives consistent reaction time but needs a minimum volume guard. Use time-based for high-traffic edges, count-based with min volume otherwise.
A.2 Client-Side Adaptive Throttling#
# per client, per backend, over a trailing 2-minute window
p_reject = max(0, (requests - K * accepts) / (requests + 1)) # K = 2
if random() < p_reject: fail locally (counts as a request, not an accept)
No states, no cliff. At K=2 the client sends at most ~2× what the backend is accepting, so the backend still sees enough traffic to signal recovery. Lower K is more aggressive; higher K wastes more backend rejection work.
A.3 Adaptive Concurrency (AIMD sketch)#
limit = 20
on response(rtt):
if rtt <= rtt_min * 1.5 and inflight >= limit * 0.9: limit += 1 # additive increase
if rtt > rtt_min * 2.0 or dropped: limit = max(5, limit * 0.9) # multiplicative decrease
admit if inflight < limit else reject 503
Gradient-based variants (as in Netflix's concurrency-limits) compare long-term vs short-term RTT and converge faster; all share the idea that latency growth means queueing, and queueing means you are past capacity.
A.4 Bulkhead Variants#
| Variant | Mechanism | Overhead | Use When |
|---|---|---|---|
| Thread-pool bulkhead | Separate executor per dependency | Context switch + thread memory (~1MB stack each) | Blocking I/O clients (classic Java) |
| Semaphore bulkhead | Permit count per dependency, caller thread | Near zero | Async/non-blocking clients; the modern default |
| Connection-pool bulkhead | Separate pool per dependency | Connections | Databases, HTTP/1.1 clients |
| Mesh concurrency cap | max_requests per upstream cluster | Proxy-level | Uniform ceiling across languages |
Appendix B: Criticality and Dependency Classification#
| Tier | Timeout Posture | Retries | Breaker | Fallback | Shed Order |
|---|---|---|---|---|---|
| CRITICAL | p99.9 × 1.5, capped by deadline | Mesh budget, idempotent only | Yes, alert on open > 60s | None — fail clearly | Last |
| DEGRADABLE | p99.9 × 1.2 | Mesh budget | Yes | Cached with cap | Middle |
| OPTIONAL | p99 × 1.2, often ≤ 300ms | None | Yes, force-openable in incidents | Omit / static | First |
Appendix C: Coordination Mechanisms — Where Protection Lives#
| Mechanism | Location | Signal | Reaction Time | Owner |
|---|---|---|---|---|
| Timeout / deadline | Caller + every hop | Elapsed time | Per call | Framework / calling team |
| Bulkhead | Caller | In-flight count | Instant | Calling team |
| Circuit breaker | Caller | Error & slow rate | 1–10s | Calling team (semantics), platform (defaults) |
| Retry budget | Sidecar | Retry/success ratio | 10s window | Platform |
| Outlier ejection | Sidecar / LB | Per-host consecutive errors | Seconds | Platform |
| Load shedding | Callee | Queue wait, in-flight, CPU | Instant | Callee team |
| Adaptive throttling | Caller | Accept ratio | ~Seconds to minutes | Platform |
C.1 Quick Comparison#
| Question | Breaker | Load Shedding | Rate Limiting |
|---|---|---|---|
| Reacts to | Dependency health | Own capacity | Client policy |
| Lives in | Caller | Callee | Gateway / callee |
| Fails | Fast, with fallback | Fast, with 503 | 429 + Retry-After |
| Protects | Caller resources | Callee survival | Fairness, abuse, quota |
See Rate Limiting for the policy side.
Appendix D: API Contract and Client Behavior#
D.1 Deadline Propagation#
Use absolute deadlines on the wire (gRPC converts to a relative timeout per hop and back); relative "timeout" headers accumulate skew and ignore queueing time.
D.2 Rejection Responses#
| Situation | Status | Headers | Client Should |
|---|---|---|---|
| Callee overloaded (shed) | 503 | Retry-After: 1–5 | Retry once after delay if budget allows |
| Client over quota | 429 | Retry-After | Back off per header; don't retry immediately |
| Deadline already expired | 504 / DEADLINE_EXCEEDED | — | Never retry |
| Breaker open locally | (no call made) | — | Serve fallback |
D.3 Retry Thundering Herd#
Exponential backoff without jitter synchronizes clients: all retry at 100ms, 200ms, 400ms. Full jitter (sleep = random(0, min(cap, base × 2^n))) spreads them uniformly; the AWS Builders' Library analysis shows full jitter completes work with markedly fewer total calls than un-jittered backoff under contention.
Appendix E: Observability#
E.1 Core Metrics — Non-Negotiable#
# Caller side
breaker_state{caller, dependency} gauge
breaker_transitions_total{caller, dependency, to} counter
bulkhead_active{caller, dependency} / bulkhead_max gauge
fallback_served_total{caller, dependency, kind} counter
fallback_last_served_timestamp{dependency} gauge
deadline_exceeded_total{caller, dependency} counter
# Mesh
rpc_retry_ratio{src, dst} gauge (retries / first attempts)
outlier_ejections_active{cluster} gauge
# Callee side
shed_total{service, criticality, reason} counter
expired_on_arrival_total{service} counter
goodput_rps{service} vs throughput_rps{service} gauge
queue_wait_ms_p99{service} histogram
E.2 Critical Alerts#
| Alert | Condition | Action |
|---|---|---|
| Retry storm | rpc_retry_ratio{dst} > 0.5 for 2 min | Page platform; consider retry kill switch |
| Critical breaker open | CRITICAL dependency breaker_state == 2 for > 60s | Page caller + dependency owners |
| Sustained fallback | fallback_served_ratio > 2× baseline for 15 min | Ticket; page if revenue path > 30 min |
| Goodput collapse | goodput / throughput < 0.7 for 3 min | Page callee; enable shedding |
| Shedding critical traffic | shed_total{criticality=CRITICAL} > 0 | Page callee |
| Stale fallback | now − fallback_last_served > 30d for tier-1 | Ticket for game day |
E.3 Control Plane vs Data Plane#
The data plane (proxies, in-process breakers) must keep working with the last-known-good config when the control plane is down. A mesh that fails closed when its control plane is unreachable turns a platform outage into a fleet outage. Test it: kill the control plane in staging monthly.
E.4 Debugging "Everything Is Up but Nothing Works"#
- Compare goodput to throughput at each tier — find where work is being done and discarded.
- Check
rpc_retry_ratioalong the call graph — find the amplifier. - Check
expired_on_arrival_total— a high value means queues are full of dead work. - Check breaker transitions for periodicity — a sawtooth means synchronized breakers.
Appendix F: Scale Evolution#
F.1 What Works at Each Scale#
| Scale | Enough | Add When |
|---|---|---|
| < 20 services, 1 language | Library timeouts, retries at one layer, a few breakers | First cascade or first external caller |
| 20–150 services | Standard library + deadlines + callee shedding for tier-1 | Third language, or nested-retry incident |
| 150–1,000 services | Mesh mechanics, criticality tags, fallback registry, game days | Threshold toil, silent fallback incidents |
| 1,000+ services, multi-region | Adaptive concurrency, cells, staged platform config, resilience SLOs | Correlated failures from shared infra |
F.2 Multi-Region Path#
Regional isolation → edge-level whole-user failover → cells within regions → shuffle sharding for multi-tenant services. Per-dependency cross-region failover stays a narrow, capped exception.
F.3 What You Don't Build on Day One#
- A shared, distributed breaker state store
- Custom adaptive-concurrency algorithms (use a library)
- Per-method breaker tuning for every endpoint (start per dependency)
- Automated regional evacuation (start with a tested manual runbook)
- Hedged requests (add only for read paths with strict tail SLOs, with a budget)
Appendix G: Multi-Tenancy, Fairness and Cost#
G.1 Noisy Neighbor via Retries#
One tenant's malformed requests fail, their SDK retries aggressively, and the shared service sheds — hitting every tenant. Fixes: per-tenant concurrency quotas at the callee; shed by tenant and criticality; per-tenant retry budgets at the gateway.
G.2 Shuffle Sharding#
Assign each tenant to a random subset of k hosts out of n (e.g., 2 of 16). A poison tenant takes down only its 2 hosts; the probability another tenant shares both is 1/C(16,2) = 1/120. Bounded blast radius without a cell per tenant. Amazon has written publicly about using this for Route 53.
G.3 Cost of Resilience Features#
| Feature | Cost | Worth It When |
|---|---|---|
| Sidecar mesh | ~0.1–0.25 vCPU + 50–100MB per pod; +0.5–2ms per hop | ≥ 3 languages or ≥ 50 services |
| Headroom for recovery surges | 20–40% extra capacity on tier-1 | Always for tier-1 |
| Regional failover capacity | Up to 2× compute | Revenue per minute justifies it |
| Hedged requests | ~2–5% extra load | Tail-latency-sensitive reads |
| Game days | ~1–2 engineer-days per service per quarter | Tier-1 and tier-2 |
G.4 Tradeoff Summary#
| Choice | Default | Who Pays If Wrong |
|---|---|---|
| Retry placement | Mesh only, 10% budget | Callee, then everyone |
| Breaker state | Local, jittered | Callee on synchronized reopen |
| Fallback semantics | Per dependency, product-approved, time-limited | Business (silent revenue loss) |
| Callee protection | Deadline drop + criticality shed + adaptive limit | Callee on-call |
| Platform config | Staged by cell with health gates | Entire fleet |