Technologies that implement this pattern: API Gateways · Redis · Apache Kafka · ZooKeeper & etcd · DynamoDB
Why This Matters#
Every system degrades. The only choice you get is whether it degrades in the order you chose in advance or the order the failure chooses for you at 3am. Uncontrolled degradation is what an outage is: the recommendation service slows down, holds threads, exhausts the pool that checkout also uses, and a non-essential widget takes the revenue path with it.
Most candidates treat degraded mode as a resilience-technique question — circuit breakers, retries, timeouts. Staff engineers treat it as a product-priority question decided before the incident. Which features are critical, which can get worse, which can vanish? What does the user see in each tier? Who is allowed to push the system down a tier, and does that decision need a human? Circuit breakers are the mechanism; the degradation ladder — signed off by product, wired to automated triggers, and rehearsed — is the design.
The second reframe: a fallback you haven't exercised is a second outage waiting to happen. Fallback code runs rarely, so it rots: it calls the same failed dependency, it was never load-tested, the static snapshot is six months old. Degraded modes are only real if they run in production on purpose — brownouts, game days, a small percentage of traffic permanently on the fallback path.
If you can present a tiered ladder with entry/exit criteria, map each feature to a criticality class, say who decides each transition, and explain how you know the fallbacks work, you are answering at Staff level.
The 60-Second Version#
- Classify every feature into three buckets: Critical, Degradable, Sheddable. In a typical product page, ~10–20% of calls are Critical (auth, price, add-to-cart), ~30% Degradable (search, recommendations with a generic fallback), ~50% Sheddable (counts, badges, "people also viewed"). Shedding that 50% frees roughly half the fan-out capacity.
- Define 4–5 tiers with numeric entry and exit criteria. E.g., enter Tier 1 when p99 > 2× SLO for 2 minutes or CPU > 80%; exit only after 10 minutes below threshold. Exit hysteresis of 3–5× the entry window prevents flapping.
- Automate the first two tiers; require a human for the last two. Shedding sheddable features should happen in <30 seconds without a page. Going read-only or to a static site is a business decision — the incident commander owns it under a pre-signed policy.
- Static fallbacks must be precomputed and dependency-free. A "last known good" snapshot on object storage/CDN, refreshed every 5–15 minutes, served with zero calls to the failing tier. If the fallback path needs the database, it isn't a fallback.
- Brownouts test and relieve at the same time. Disable optional components for a dynamic fraction of requests (0–100%) under load, or deliberately for 1–5% of traffic in steady state, so the fallback path is always warm and always proven.
- Kill switches must not depend on the thing that's failing. The flag/config system needs its own availability ≥ the services it controls, a local cached value on every host, and a break-glass path that works when the control plane is down.
The Problem#
Demand exceeds capacity, or a dependency fails, and the system has to keep doing something useful. Without a plan, degradation is decided by thread pools and timeouts: whichever dependency is slowest consumes the shared resources and every feature fails together. A ride-hailing app that can't compute surge pricing should still dispatch at the last known multiplier. A streaming service whose personalization is down should still play video with a generic row of popular titles. A checkout whose fraud model is timing out needs someone to have decided — in daylight, with finance in the room — whether to approve small orders or block them all. The engineering problem is building those modes; the Staff problem is making sure the decisions are made before anyone needs them.
Case Studies That Use This Pattern#
- Circuit Breakers — The mechanism that trips a dependency into its fallback; degraded mode is what the fallback does
- Rate Limiting — Fail-open vs fail-closed when the limiter's own state store is down
- Payment Processing — Fraud/risk dependency down: approve-below-threshold vs block decisions owned by finance
- News Feed — Ranking down → chronological feed; the canonical degradable feature
- Flash Sales & Ticketing — Waiting rooms as a planned Tier 4 under demand spikes
- API Gateway — Where criticality headers are enforced and load is shed first
- Notification System — Deferring non-urgent notifications while keeping security alerts flowing
The Four Intents#
"Add graceful degradation" can mean four different problems. They need different mechanisms and different owners.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Survive overload (demand > capacity) | Goodput must stay near capacity, not collapse | Priority-based load shedding, brownouts, waiting rooms | Shedding the wrong traffic; retry storms amplify load 3–10× | Critical requests keep SLO |
| Survive dependency failure | A downstream is slow or dead | Timeouts + circuit breakers + per-feature fallbacks | Fallback calls the same dead dependency; shared pools exhausted | Page renders with degraded components |
| Survive infrastructure loss (AZ/region/control plane) | Can't provision or reconfigure during the event | Static stability: pre-provisioned capacity, cached config, no control-plane calls on the data path | Recovery needs the control plane that's down | Existing capacity keeps serving |
| Protect correctness under failure | Degrading must not create money/security incidents | Fail-closed for auth/payments; bounded fail-open with explicit limits | Fail-open leaks access or approves fraud | Zero unauthorized actions; bounded financial exposure |
🎯 Staff Move: "I'll design for overload and dependency failure first, since that's where 80% of incidents come from. But I want to flag up front that auth and payments fail closed — or fail open only under a limit finance has signed — and that's not an engineering call I'll make alone."
The Degradation Ladder#
The artifact at the center of this pattern: a small number of tiers, each with a user-visible description, entry and exit criteria, and a decision owner.
| Tier | Name | What the User Sees | Entry Trigger (example) | Exit | Who Decides |
|---|---|---|---|---|---|
| 0 | Full service | Everything | — | — | — |
| 1 | Trim | No recommendations, counts, badges; lower image quality | p99 > 2× SLO for 2 min, or CPU > 80% fleet-wide | 10 min healthy | Automatic |
| 2 | Cached / static fallback | Stale-but-complete pages (≤15 min old), generic content instead of personalized | Critical-path error rate > 1% for 1 min, or dependency breaker open | 15 min healthy + dependency green | Automatic, IC notified |
| 3 | Core only / read-only | Browse and checkout work; edits, uploads, reviews queued or disabled | Primary datastore degraded, or Tier 2 insufficient | IC declares | Incident commander under pre-signed policy |
| 4 | Waiting room / maintenance | Queue page with position, or static status page | Capacity < 30% of demand, or data-integrity risk | IC + business owner | IC + VP/business owner |
🎯 Staff Insight: Recovery is the most dangerous part of the ladder. Re-enabling all features at once is a self-inflicted thundering herd — cold caches, reconnecting clients, and queued writes all hit together. Step up one tier at a time, with a dwell time, and ramp traffic within a tier (10% → 50% → 100%).
The Core Tradeoff#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Feature shedding (flags per feature) | Frees capacity fast; users keep core flows | Revenue from shed features (recs can be 10–30% of engagement on some surfaces) | Product/growth metrics during the incident |
| Static / cached fallback | Zero dependency on failing tier; survives total backend loss | Staleness; can't personalize; snapshot pipeline to maintain | Users see stale data; a team owns the snapshot job forever |
| Priority load shedding | Keeps goodput at capacity; protects critical paths | Requires criticality tagging across every service | Sheddable callers get 503s; every team must tag requests |
| Brownout (dimmer) | Smooth, proportional degradation; continuously tests fallbacks | Harder to reason about; user experience varies per request | Product accepts inconsistent UX; platform owns the controller |
| Read-only mode | Protects data integrity when writes are unsafe | Users can't act; queued writes need replay and conflict handling | Business (lost transactions); support (confused users) |
| Fail-open | Availability of the protected flow | Security/financial exposure while open | Security, finance, or trust & safety |
| Fail-closed | No unauthorized action | Availability loss for legitimate users | Users and revenue |
Staff Default Position#
Decide the degradation order in daylight, automate the reversible tiers, and prove every fallback in production.
Classify features into Critical / Degradable / Sheddable with product sign-off. Build a 4–5 tier ladder with numeric entry and exit criteria. Automate Tier 1–2 transitions (reversible, low business risk) with hysteresis; require the incident commander for Tier 3+ under a pre-approved policy so nobody is waiting on a VP at 3am. Make fallbacks dependency-free (precomputed snapshots, cached config, local defaults). Run the fallback path continuously for a small slice of traffic or on a regular brownout schedule so it cannot silently rot.
When to Deviate#
- Correctness-critical flows — Auth, payments, permissions, medical or safety systems. Default to fail-closed; any fail-open needs explicit limits (e.g., approve orders < $50 for ≤15 minutes) signed by the owning business function.
- Single-purpose systems — A system with one feature (a payment ledger, a DNS resolver) has nothing to shed. Invest in redundancy and static stability instead of a ladder.
- Early-stage products — Under ~10 services and one team, a full ladder is overhead. Start with timeouts, one kill switch per non-critical dependency, and a static maintenance page.
- Regulated freshness — Where showing stale data is itself a violation (trading prices, drug dosage), a stale fallback is worse than an error. Degrade to "unavailable" with a clear message.
The L5 → L6 → L7 Contrast#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Wrap calls in circuit breakers with fallbacks" | "Which features are critical, degradable, sheddable — and has product signed off on the order?" | "Does the org have one degradation vocabulary and control plane, or does every team invent its own ladder?" |
| Fallbacks | Return empty list / default value | Precomputed, dependency-free fallbacks with a staleness bound; exercised continuously | Fallback coverage is a tracked org metric; services without tested fallbacks can't be tagged Critical-path dependencies |
| Triggers | Manual flag flips during the incident | Automated Tier 1–2 with hysteresis; IC-owned Tier 3+ under a pre-signed policy | Decision rights codified org-wide: what on-call may do without asking, what needs the business owner |
| Failure | "The breaker opens and we return a default" | "The fallback must not share a thread pool, a dependency, or a config service with the primary path" | Designs against correlated degradation: the flag system, the CDN and the auth service as shared fate across 50 teams |
| Testing | Unit tests for fallback code | Brownouts and game days; 1–5% of traffic on fallback path | Quarterly org-wide game days; error budgets spent deliberately on degradation drills |
| Ownership | The service team | Product owns the ladder order; SRE owns triggers; service teams own fallbacks | Reliability org owns the framework; product leadership owns the criticality map; exceptions reviewed quarterly |
Why "First move" separates levels
A circuit breaker answers "how do I stop calling a broken dependency?" It doesn't answer "what does the page look like then, and is that acceptable?" The Staff candidate recognizes the second question is a product decision and gets it made before the incident. The Principal candidate recognizes that 40 teams answering it independently produces 40 incompatible ladders — and an incident commander who can't reason about any of them.
Why "Failure" separates levels
The most common degraded-mode failure isn't a missing fallback; it's a fallback that shares fate with the primary path — the same connection pool, the same config service, the same region. Staff checks shared fate within one service. Principal checks it across the org: if every team's kill switch lives in one flag service, that flag service is the most critical system in the company and must be engineered like it.
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Fail-Open vs Fail-Closed | Availability of the flow vs exposure while unprotected |
| 2 | Automated vs Human Triggers | Speed (seconds) vs judgment (minutes) |
| 3 | Precomputed vs Live Degraded Fallbacks | Independence from failure vs quality of the degraded experience |
| 4 | Server-Driven vs Client-Driven Degradation | Central control vs resilience when the server can't be reached |
| 5 | Practice in Production vs Hope | Deliberate user impact now vs untested fallbacks during the real event |
Fault Line 1: Fail-Open vs Fail-Closed#
When a protective dependency (rate limiter, fraud model, authz service) fails, you either let traffic through or block it. Who pays: fail-open pays in exposure (fraud, abuse, data access); fail-closed pays in availability for legitimate users. Staff default: fail-open for abuse protection and rate limiting (the harm of blocking everyone exceeds the harm of 15 minutes without limits), fail-closed for authorization and money movement, with bounded fail-open (amount caps, time caps, known-good users only) where the business signs it. Deviate when: the protective system is itself the product (a WAF vendor, a compliance gate).
Fault Line 2: Automated vs Human Triggers#
An automated trigger reacts in 10–30 seconds; a paged human takes 5–15 minutes to acknowledge, orient, and act. Who pays: automation pays in false positives (shedding recommendations because of a noisy metric); humans pay in minutes of full outage. Staff default: automate reversible, low-business-impact tiers; human-in-the-loop for tiers that stop revenue or risk data. Every automated trigger has hysteresis and a rate limit on transitions (e.g., max 1 tier change per 5 minutes). Deviate when: the event is predictable (a ticket drop) — pre-enter Tier 1 on a schedule.
Fault Line 3: Precomputed vs Live Degraded Fallbacks#
A live fallback computes a simpler answer (popular items from a lighter query); a precomputed fallback serves a snapshot from object storage. Who pays: live fallbacks still consume capacity on the degraded tier — sometimes the same capacity that's exhausted; precomputed fallbacks cost a snapshot pipeline and staleness. Staff default: precomputed for Tier 2 (must survive total backend loss), live for Tier 1 (still have most capacity). Deviate when: the content is per-user and can't be meaningfully precomputed — degrade to a generic version instead.
Fault Line 4: Server-Driven vs Client-Driven Degradation#
Servers can flag degradation in responses; clients can also degrade themselves (cached last response, hide widgets on timeout). Who pays: server-driven depends on reaching the server; client-driven means old app versions behave differently and mobile teams own part of the reliability story. Staff default: both — the server sends a degradation hint header, and clients have a local timeout budget per component and render cached content when the call fails. Deviate when: clients are third parties you can't update (public APIs) — degrade only server-side and document it in the API contract.
Fault Line 5: Practice in Production vs Hope#
Brownouts and game days impose small, deliberate user impact. Who pays: product accepts a measurable dip (typically <0.1% of sessions affected) in exchange for knowing the fallback works. Staff default: continuous low-percentage fallback traffic for Tier 1–2; quarterly game days for Tier 3–4. Deviate when: a regulated flow can't be degraded deliberately — test in a production-shaped staging environment with production traffic replay.
Common Interview Mistakes#
| What Candidates Say | What Interviewers Hear | What Staff Engineers Say |
|---|---|---|
| "We'll add circuit breakers with fallbacks" | "Mechanism, no policy" | "The breaker is the switch. The ladder — what each tier looks like and who approves it — is the design." |
| "If recommendations fail we return an empty list" | "Didn't ask product what the page should show" | "Recs degrade to a precomputed popular-items list refreshed every 10 minutes; product approved that as the Tier 2 experience." |
| "On-call can flip the kill switch" | "Kill switch depends on what's broken?" | "Flags are cached on every host with a last-known-good value; there's a break-glass path if the flag service is down." |
| "We'll fail open so users aren't affected" | "Hasn't thought about exposure" | "Fail-open for the rate limiter, fail-closed for authz. The fraud fallback approves under $50 for 15 minutes — finance signed that." |
| "We'll bring everything back when it recovers" | "Recovery storm not considered" | "Recovery steps up one tier at a time with a 10-minute dwell, ramping 10% → 50% → 100% inside each tier." |
| "We tested the fallback in staging" | "Fallback will rot in production" | "1% of traffic runs the fallback permanently, so it's always warm, always load-tested, and alerts if it breaks." |
Quick Reference#
Staff Sentence Templates#
"Before I add breakers, I want the degradation order signed off: [features A, B] are sheddable, [C] degrades to [a precomputed fallback refreshed every N minutes], and [D, E] are critical. Product owns that list; I own making each tier actually work."
"Tiers 1 and 2 are automatic because they're reversible and cost only [engagement metric]. Tier 3 stops [revenue flow], so the incident commander decides under a policy [business owner] pre-approved — nobody waits for a VP at 3am."
"The fallback for [feature] must not share [thread pool / database / config service / region] with the primary path, otherwise it fails in exactly the situation it exists for."
"For [protective control], we fail [open/closed]. If open, exposure is capped at [amount/time] and [team] is paged the moment we enter that mode."
Implementation Deep Dive#
1. Criticality Propagation + Priority Load Shedding#
Google's SRE book describes tagging every request with a criticality (e.g., CRITICAL_PLUS, CRITICAL, SHEDDABLE_PLUS, SHEDDABLE) that propagates through the call graph, so an overloaded backend rejects sheddable work first. The same idea works with a single header.
# At the edge (API gateway): classify once, propagate everywhere
CRITICALITY = {
"POST /checkout": 3, # critical
"GET /product/:id": 2, # degradable (core content)
"GET /recommendations": 1, # sheddable_plus
"POST /analytics/beacon":0, # sheddable
}
function edge(req):
req.headers["x-criticality"] = CRITICALITY.get(route(req), 1)
req.headers["x-deadline-ms"] = now_ms() + 800 # end-to-end budget
# In every service: admission control before doing work
function admit(req):
c = int(req.headers["x-criticality"])
util = concurrency_in_use / concurrency_limit # adaptive limit
threshold = { 0: 0.60, 1: 0.75, 2: 0.90, 3: 1.00 }[c] # shed lowest first
if util > threshold:
metrics.incr("shed.rejected", tags=["crit:" + c])
return reject(503, retry_after = 2 + rand(0, 3)) # jittered, bounded
if now_ms() > int(req.headers["x-deadline-ms"]):
metrics.incr("shed.expired")
return reject(504) # don't work on dead requests
return process(req)
Why it matters: without tags, an overloaded service rejects uniformly — 30% of checkouts fail alongside 30% of analytics beacons. With tags, analytics goes to 100% rejection before a single checkout fails. Downstream calls inherit the caller's criticality; a batch job can't escalate itself to critical.
🎯 Staff Insight: The hard part isn't the admission code — it's getting 40 services to tag honestly. If every team marks its traffic critical, the scheme collapses. Criticality is assigned at the edge by route, reviewed by the reliability team, and can only be lowered downstream.
2. The Ladder Controller — Automated Tiers with Hysteresis#
TIERS = [
{ tier: 1, enter: "p99_ms > 2*SLO for 120s OR cpu > 0.80 for 120s", exit_dwell: 600 },
{ tier: 2, enter: "crit_error_rate > 0.01 for 60s OR breaker_open(core_dep)", exit_dwell: 900 },
]
MAX_TRANSITIONS_PER_5MIN = 1
every 10s:
signals = read_metrics()
target = highest tier whose enter condition holds # auto tiers only (1-2)
if target > current and transitions_recent() < MAX_TRANSITIONS_PER_5MIN:
set_tier(target) # step down immediately
page(IC, "entered tier " + target, signals)
elif target < current and healthy_for(current.exit_dwell):
set_tier(current - 1) # step up ONE tier at a time
ramp_features(current, steps=[0.1, 0.5, 1.0], interval=120)
function set_tier(t):
flags.publish("degrade.tier", t) # pushed to all hosts
audit_log.write(actor="ladder-controller", tier=t, signals=signals)
Hosts read the tier locally: every host caches degrade.tier in memory with a last-known-good value persisted to local disk. If the flag service is unreachable, hosts keep the last value — they do not reset to Tier 0.
3. Dependency-Free Static Fallback#
# Snapshot job: every 10 minutes, render fallback content from healthy state
every 10m:
for surface in ["home", "category/*", "product/top-50K"]:
html_or_json = render(surface, personalization=false)
s3.put("fallback/" + surface, html_or_json, cache_control="max-age=900")
s3.put("fallback/_manifest", { generated_at: now(), count: n })
metrics.gauge("fallback.snapshot_age_s", 0)
# Serving path: consult tier BEFORE calling any backend
function serve(req):
if local_tier() >= 2 or breaker("catalog").open:
body = local_disk_cache.get(req.surface) or cdn_fetch("fallback/" + req.surface)
return respond(body, headers={"x-degraded": "tier2", "Cache-Control": "max-age=60"})
return live_render(req)
Rules: the snapshot is stored where the failing tier isn't (object storage + CDN), refreshed often enough that staleness is product-approved (here 10 min, max 25 min with a missed run), and alerts when fallback.snapshot_age_s > 1,800. The serving path checks the tier first, so in Tier 2 not even a timed-out call is made.
4. Brownout Dimmer#
The brownout idea (Klein et al., ICSE 2014) is a "dimmer" in [0, 1] that sets the probability of executing optional content per request, adjusted continuously from a latency target.
dimmer = 1.0 # 1.0 = all optional features on
TARGET_P95_MS = 300
every 5s:
err = (TARGET_P95_MS - p95_ms()) / TARGET_P95_MS
dimmer = clamp(dimmer + 0.3 * err, 0.02, 1.0) # floor 2%: fallback always exercised
metrics.gauge("brownout.dimmer", dimmer)
function render(req):
page = core(req)
if rand() < dimmer: page.recs = recommendations(req) else: page.recs = popular_fallback()
if rand() < dimmer: page.counts = counts(req)
return page
Why the 2% floor: the fallback path runs for at least 2% of traffic all the time, so it is load-tested and monitored continuously — a broken fallback surfaces as an alert on a Tuesday afternoon, not during a Sev-1.
5. Read-Only Mode with Deferred Writes#
function handleWrite(req):
if local_tier() >= 3:
if req.kind in DEFERRABLE: # reviews, profile edits, uploads metadata
queue.put(req, idempotency_key=req.key) # durable queue outside failing DB
return 202, {"status": "queued", "message": "Saved; will appear shortly"}
return 503, {"message": "Temporarily read-only", "retry_after": 300}
return db.write(req)
Replay on recovery goes through the same idempotency keys and rate-limits itself to ~20% of primary write capacity so recovery doesn't become the next overload.
Mechanism Comparison
| Mechanism | Reaction Time | Granularity | Needs Product Sign-off | Tests Itself |
|---|---|---|---|---|
| Priority load shedding | ms (per request) | Request class | Criticality map only | Yes, under any overload |
| Ladder controller (auto tiers) | 10–120s | Feature set | Yes, tier contents | Only if tiers are drilled |
| Static fallback | Instant once tier set | Surface | Yes, staleness bound | Only with floor traffic |
| Brownout dimmer | 5–30s | Per-feature probability | Yes, UX variance | Yes, with a floor |
| Read-only mode | Minutes (human) | Whole write path | Yes, business owner | Game days only |
Architecture Diagram#
How to narrate it: the data plane never asks the control plane for permission at request time. Tiers and dimmers are pushed and cached locally; fallback assets live outside the failing tier; checkout has its own pool so a slow recommendations call can't consume its threads.
Failure Scenarios#
1. The Fallback That Called the Outage — 100% Error Rate From a 5% Feature#
Personalization's fallback was "fetch the popular list from the catalog DB." When the catalog replica fleet slowed down, both the primary path and the fallback hit the same replicas.
t=0 Catalog replicas p99 40ms -> 900ms (bad query plan after stats refresh).
t=+20s Personalization times out; breaker opens; fallback fires.
t=+25s Fallback queries the same catalog replicas: load +60%.
t=+60s Page service thread pool (shared) saturated waiting on fallback.
t=+90s Home page error rate 100%, including checkout links rendered there.
t=+9min On-call disables personalization AND fallback via flag; page renders.
t=+40min Query plan fixed; features restored one at a time.
Detection: fallback.invocations spike correlated with dep.latency_p99{dep=catalog}; threadpool.saturation{pool=page} > 90%.
Blast radius: entire home page — a sheddable feature took down a critical surface.
Mitigation: fallback moved to precomputed snapshot in object storage; separate bulkhead pool per dependency.
Prevention: design review checklist: "What does the fallback depend on? Is any of it shared with the primary path?"; 2% brownout floor so the fallback runs under real load daily.
Owner: personalization team owns the fallback; reliability team owns the checklist.
🎯 Staff Insight: The rule is shared fate audit: a fallback must differ from the primary path in every resource that can fail — dependency, pool, region, config. If you can't name the difference, you don't have a fallback.
2. Ladder Flapping — Tier Oscillation Every 90 Seconds#
An automated ladder entered Tier 1 at CPU > 80% and exited at CPU < 80%, with no dwell. Shedding dropped CPU to 65%, which exited Tier 1, which restored load, which re-entered Tier 1.
t=0 CPU 82% -> Tier 1 (recs, counts shed). CPU drops to 65%.
t=+30s Exit condition met -> Tier 0. Cold recs caches refill, CPU 95%.
t=+60s Tier 1 again. Users see widgets appear and vanish.
t=+15min 20 transitions. Recs cache never warms; CDN hit ratio falls on page variants.
t=+17min On-call pins Tier 1 manually. Stable.
Detection: degrade.tier_transitions > 3 per 10 min.
Blast radius: all users on affected surfaces; cache churn raised backend load 25% above the no-ladder baseline.
Mitigation: manual pin; exit dwell of 10 min; max 1 transition per 5 min; exit threshold 20 points below entry (80% in, 60% out).
Owner: reliability team owns the controller and its parameters.
3. The Kill Switch That Needed the Outage to Work#
A regional networking event made the flag service unreachable from half the fleet. On-call flipped the kill switch for a failing payment-risk dependency; only hosts that could reach the flag service received it. Worse, hosts configured to "fetch flags on startup" were being restarted by auto-healing and came up with defaults — Tier 0, risk checks enabled, calling the dead dependency.
t=0 Risk service unavailable in region A. Checkout p99 3s, errors 35%.
t=+4min IC approves pre-signed policy: bypass risk for orders < $50 for 15 min.
t=+5min Flag flipped. 48% of hosts receive it. Error rate falls to 18%.
t=+8min Auto-healing restarts unhealthy hosts; they boot with defaults (risk on).
t=+20min Break-glass: env override pushed via deployment pipeline. Errors 0.5%.
t=+35min Risk recovers; override removed; exposure review: $41K in un-scored orders.
Detection: flags.fetch_failures per host; flags.version_skew (hosts on different flag versions) > 5%.
Blast radius: half of checkout for 20 minutes.
Mitigation: break-glass override via deploy pipeline; bypass bounded by amount and time.
Prevention: flags persisted to local disk and loaded before startup defaults; flag service deployed per region with static stability; break-glass path documented and drilled.
Owner: platform team owns the flag system's availability SLO (set higher than any service it controls); payments owns the bounded-bypass policy with finance.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Fallback shares fate with primary | fallback.error_rate during incident; shared-pool saturation | Surface-wide | Bulkheads, precomputed fallback | Service team |
| Ladder flapping | degrade.tier_transitions > 3/10min | All users on surface | Dwell time, asymmetric thresholds | Reliability team |
| Stale fallback snapshot | fallback.snapshot_age_s > 1,800 | Users in Tier 2 see very old data | Alert, re-run job, block Tier 2 if > 2h | Snapshot job owner |
| Kill switch unreachable | flags.version_skew > 5% | Portion of fleet un-degraded | Local persistence, break-glass | Platform team |
| Recovery thundering herd | Backend QPS step on tier exit; cache miss spike | Can re-trigger the incident | One tier at a time, 10/50/100% ramp | Reliability team |
| Fail-open exposure overrun | bypass.approved_amount vs cap | Financial / security loss | Auto-expire bypass after cap | Business owner + service |
| Criticality inflation | Share of traffic tagged critical > 30% | Shedding ineffective | Edge-assigned tags, quarterly review | Reliability team |
Who Decides — The Decision Rights Matrix#
Degraded mode fails organizationally before it fails technically: on-call knows how to flip the switch but doesn't know whether they're allowed to turn off checkout for 5% of users. Decide this in advance and write it down.
| Action | Who May Do It | Needs Approval From | Max Duration Before Re-approval | Pre-signed By |
|---|---|---|---|---|
| Tier 1 (shed sheddable) | Automation, any on-call | Nobody | Unlimited | Product lead for each surface |
| Tier 2 (static fallback) | Automation, any on-call | Nobody; IC notified | 4 hours | Product lead |
| Fail-open a rate limiter | On-call | Nobody; security notified | 1 hour | Security lead |
| Bounded fraud bypass (< $X) | Incident commander | Nobody at incident time | 15–30 min, auto-expire | Finance / risk owner |
| Tier 3 (read-only) | Incident commander | Nobody at incident time | 1 hour | VP Engineering + product |
| Tier 4 (waiting room / maintenance) | Incident commander | Business owner on call | Until resolved | VP + business owner |
| Fail-open authorization | Nobody | — | — | Not permitted |
🎯 Staff Move: "The key design output isn't the flag system — it's this table, signed before launch. At 3am the incident commander should be executing a decision, not making one."
The Principal Lens#
Why L7 Sees This Problem Differently#
A Staff engineer designs a ladder for their system. A Principal engineer sees that the company's real failure posture is the composition of fifty ladders — and that the composition has properties no single team designed. Tier 2 on the product page may be fine; Tier 2 on product page + search + cart at the same time may leave users unable to buy anything. Every team's kill switch lives in the same flag service; every fallback snapshot lives in the same bucket in the same region. At org scale, degraded mode becomes a question of correlated failure and shared decision rights: one vocabulary of tiers, one control plane engineered to outlive everything it controls, and a practice budget (game days, brownouts) that the org actually spends.
The Org-Level Fault Line#
A central degradation control plane with org-wide tiers vs per-team fallbacks.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Per-team fallbacks and flags | Local knowledge, fast iteration | No global "go to Tier 2" button; incident commanders must know 50 systems; inconsistent UX | IC during every large incident; users seeing half-degraded pages |
| One central "big red button" | One action degrades everything | Blunt; a single mistake degrades the whole company; control plane is a massive SPOF | Everyone, when it misfires |
| Shared vocabulary + federated ladders + one resilient control plane | Org-wide tier semantics; teams own what "Tier 2" means for them; IC can move a surface or the whole site | Requires a standard, a platform team, and quarterly drills | Reliability platform (3–6 engineers) + ~1 day/quarter per team |
Meta's published Defcon system (OSDI 2023) is a public example of the third shape: feature "knobs" registered by product teams, grouped by business criticality into levels, and actuated centrally during capacity crises.
Cost Model#
Assumptions: fully loaded engineer ~$25K/month; snapshot storage and CDN at object-storage pricing; game-day cost measured in engineer-hours; revenue at risk estimated as hourly revenue × outage minutes avoided.
| Scale | Services | What You Build | Infra $/month | People | On-call / Practice Load |
|---|---|---|---|---|---|
| Startup (~$10M ARR) | 5–15 | Timeouts, a kill switch per non-critical dependency, static maintenance page | ~$200 | 0.1 FTE | 1 tabletop drill per half |
| Growth (~$200M ARR) | 50–150 | Criticality tags at gateway, auto Tier 1–2 ladder, snapshot fallbacks for top surfaces, flag service with local persistence | ~$5–10K (snapshots, flag infra, spare pools) | 1–2 FTE reliability | Quarterly game day (~40 eng-hours) |
| Large (~$5B ARR) | 1,000+ | Federated ladder standard, central control plane per region, continuous brownout floor, reserved headroom for recovery | ~$100–300K (per-region control plane, bulkhead headroom, snapshot fleets) | 5–8 FTE platform | Monthly drills per org; annual full-region failover (~500 eng-hours) |
The Principal framing: at $5B ARR, revenue is ~$570K/hour. If the degradation program turns two 60-minute full outages per year into 60-minute Tier 2 events that retain ~80% of revenue, it preserves roughly $900K/year in direct revenue — before counting trust and SLA credits. That's the number that funds the platform team.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Adding a kill switch or fallback | Two-way | Hours |
| Tier thresholds and dwell times | Two-way | Config change |
| Fail-open policy for payments or auth | One-way (per incident) | Money or data exposed can't be recovered |
| Public API promise of "degraded responses" (e.g., partial results flag) | One-way | Clients build on it; removing it breaks them |
| Criticality tagging protocol across services | One-way-ish | Every service and client library changes; 2–3 quarters to redo |
| Putting the flag system on the same infra as what it controls | Two-way to fix, one-way in the incident | Discovered only when it's too late |
The Standard I'd Write#
RFC: Degradation Tiers and Fallback Standard (v1)
Scope: All user-facing services and any service with a Critical-tagged caller.
MUST:
- Each service declares its features/endpoints as Critical, Degradable, or Sheddable, approved by the owning product lead and reviewed annually.
- Each service defines its behavior at org tiers T1–T4; "no change" is allowed but must be explicit.
- Fallbacks for Degradable features MUST NOT share a dependency, thread pool, or region with the primary path; the design doc names the difference.
- Degradation state MUST be read from local cache with last-known-good persistence; no request-path calls to the control plane.
- Requests MUST carry an edge-assigned criticality; services may lower but never raise it.
- Fail-open of authorization is prohibited. Fail-open of financial controls requires a signed cap (amount, duration) and auto-expiry.
SHOULD: Run fallback paths for ≥1% of production traffic continuously; participate in a quarterly game day.
Exceptions: Reliability review, time-boxed to 2 quarters, owner named.
Success metrics: % of Critical-path dependencies with tested fallbacks (target > 90%); minutes at Tier 2+ vs minutes of full outage per quarter; zero incidents where a fallback shared fate with its primary; median time-to-degrade < 60s.
What I'd Tell the VP#
When something breaks today, the whole site tends to break with it, because nothing decides what's allowed to fail first. I want us to agree — product and engineering together, ahead of time — on which features are essential, which can get simpler, and which can switch off, and then automate the first steps so we degrade in seconds instead of having an outage for minutes. The cost is a small platform team and about a day a quarter from each product team to practice. The payoff is that our worst incidents become "the site was plainer for an hour" instead of "the site was down for an hour." I need product leaders to sign off on the essentials list and finance to set the limits for the payment exceptions.
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Correlated-failure awareness | "Every team's kill switch lives in one flag service. That service needs a higher availability target than anything it controls — and it needs to work when its region doesn't." |
| Decision rights as design output | "The artifact I'd ship first is the decision-rights table, signed by product, finance and security." |
| Composed degradation | "Tier 2 on search plus Tier 2 on cart might mean nobody can buy. I'd test tiers in combination, not per service." |
| Pricing resilience | "Converting a one-hour outage into a one-hour Tier 2 retains ~80% of revenue; at our size that funds the team twice over." |
| Practice as a budget | "We spend a slice of error budget on brownouts on purpose. An unexercised fallback isn't an asset — it's a liability we haven't discovered yet." |
Staff answers that L7 interviewers find insufficient:
- "I'd add a fallback and a kill switch for each dependency." — Right per service, silent on shared fate across services and on who's allowed to use the switch.
- "We'll automate the ladder so humans aren't needed." — Ignores that the expensive tiers are business decisions that need pre-delegated authority, not automation.
- "We run chaos tests in staging." — Doesn't commit the org to practicing in production where fallbacks actually rot.
In the Wild#
Google: Request Criticality and Load Shedding#
The Google SRE book's chapters on handling overload describe tagging requests with a criticality level (from CRITICAL_PLUS down to SHEDDABLE) that propagates across RPCs, so that overloaded backends reject the least important work first. It also describes client-side adaptive throttling, where clients locally reject a fraction of requests when their backend has been rejecting them, so retries don't amplify overload.
Staff insight: The mechanism that keeps critical traffic alive is not a clever algorithm; it's an org-wide convention (criticality propagated in every call) that every service honors. That's the argument for making degradation a shared vocabulary rather than a per-team feature.
AWS: Static Stability#
The Amazon Builders' Library describes "static stability": systems designed so that the data plane keeps working without the control plane — for example, pre-provisioning enough capacity across Availability Zones that losing one AZ doesn't require launching new instances during the event. The data plane keeps serving with what it already has.
Staff insight: This is the design principle behind "the kill switch must not depend on what's failing." Degraded modes that require provisioning, deploying, or reaching a control plane during the incident are fallbacks in name only.
Meta: Defcon — Graceful Feature Degradation#
Meta's OSDI 2023 paper on Defcon describes a system where product teams register "knobs" that turn off or reduce features, grouped into levels by business criticality, which can be actuated during capacity crises (such as losing a datacenter's worth of capacity) to cut load while keeping core functionality. The paper reports that turning off lower-priority features freed substantial capacity during real events, and that knobs are tested regularly.
Staff insight: Defcon is the "federated ladder with central actuation" shape: teams own what their knobs do, the org owns the levels and the button. It's the strongest public citation for the Principal answer.
Practice Drill#
Prompt: "Our e-commerce home page calls 14 services. Last month, the reviews service had a memory leak and the whole home page went down for 35 minutes, including links to checkout. Design degraded mode so this can't happen again."
Staff Answer
A reviews leak taking down the home page is a missing degradation order, not a missing breaker. First, classify the 14 calls with product: Critical (session, cart count, top navigation, prices — ~4 calls), Degradable (hero banners, category tiles, search suggestions — fall back to a 10-minute snapshot), Sheddable (reviews summary, recently viewed, recommendations, badges — ~7 calls). Second, isolate: each dependency gets its own bulkhead (bounded concurrency, e.g., 50 in-flight) and a per-call timeout within an 800ms page deadline — reviews gets 120ms. A slow reviews service can then consume at most its own 50 slots. Third, fallbacks: sheddable components render empty with layout preserved; degradable components render from a snapshot in object storage, never from a service on the same path. Fourth, the ladder: Tier 1 (drop sheddable) auto-enters when page p99 > 1.6s for 2 min, exits after 10 min healthy; Tier 2 (full static home page from CDN) auto-enters when critical error rate > 1% for 1 min. Fifth, practice: 2% of home page traffic always runs with sheddable components off, so the degraded layout is known-good; quarterly game day kills reviews at peak. Metrics: bulkhead.rejected{dep}, page.p99_ms, degrade.tier, fallback.snapshot_age_s, page.critical_error_rate. Ownership: product owns the classification; each component team owns its fallback; reliability owns the ladder controller.
Why this is L6:
- Treats the incident as a missing priority decision and gets product to own the classification.
- Bounds blast radius with bulkheads and per-call timeouts inside a page-level deadline, with concrete numbers.
- Makes fallbacks dependency-free and continuously exercised, with entry/exit criteria and hysteresis.
What L7 adds:
- Asks whether the other top surfaces (search, product page, cart) have the same failure shape and proposes an org-wide tier vocabulary so an IC can degrade the site coherently.
- Checks correlated fate: are all 14 fallback snapshots and all kill switches in one bucket and one flag service in one region?
- Prices it: at the company's revenue per hour, 35 minutes of home page outage vs 35 minutes of Tier 1 — and uses that to fund the snapshot pipeline and game days.
Staff Interview Application#
How to Introduce This Pattern#
"Before I add circuit breakers, I want to decide the order in which this system is allowed to get worse. I'll split features into critical, degradable, and sheddable with product, define three or four tiers with numeric entry and exit criteria, automate the reversible ones, and name who's allowed to trigger the rest. Then I'll make sure every fallback runs in production a little bit every day, so we know it works."
When NOT to Use This Pattern#
- Nothing to shed: Single-purpose systems (a ledger, a lock service) should invest in redundancy and static stability, not ladders.
- Stale is worse than nothing: Regulated or safety data where a stale fallback misleads. Degrade to an explicit "unavailable."
- Tiny scale: Under ~10 services, timeouts plus a handful of kill switches and a maintenance page deliver most of the value; a ladder controller is overhead.
- Batch/offline systems: Pipelines degrade by delaying, not by shedding features — use backpressure and SLAs on lateness instead (Data Pipeline Patterns).
Follow-Up Questions to Anticipate#
| Interviewer Asks | What They Are Testing | How to Respond |
|---|---|---|
| "What happens when the fallback fails too?" | Shared-fate reasoning | "It shouldn't share any resource with the primary path — precomputed in object storage, separate pool. If it does fail, the component renders empty; the page still ships." |
| "Who decides to go read-only?" | Decision rights | "The IC, under a policy the VP and product signed before launch. Tiers 1–2 are automatic; 3–4 are human but pre-authorized." |
| "Fail open or fail closed?" | Risk ownership | "Open for rate limiting and abuse controls with a time cap; closed for authorization; bounded open for fraud under a finance-signed limit." |
| "How do you avoid flapping?" | Control-loop maturity | "Asymmetric thresholds, 10-minute exit dwell, max one transition per 5 minutes, recover one tier at a time." |
| "How do you know the fallbacks work?" | Operational honesty | "They run for 1–2% of traffic permanently, and we game-day the big tiers quarterly." |
| "What if the flag service is down?" | Control-plane dependency | "Flags are cached in memory and on disk per host; no request-path calls. Break-glass override goes through the deploy pipeline." |
Evaluation Rubric#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | Breakers and retries | Criticality classes + ladder, product-signed | Org-wide tier vocabulary, composed degradation |
| Fallbacks | Defaults / empty responses | Dependency-free, precomputed, continuously exercised | Fallback coverage as an org KPI |
| Triggers | Manual flips | Auto Tiers 1–2 with hysteresis, IC for 3–4 | Decision rights codified across org |
| Failure | Single dependency | Shared fate within the service | Correlated fate across teams (flag service, snapshots, regions) |
| Cost | Not discussed | Features' revenue impact per tier | Resilience ROI in $/year funds the program |
Strong Hire Signals
| Signal | What It Sounds Like |
|---|---|
| Priority before mechanism | "What must keep working? Let's get product to rank it." |
| Numeric tiers | "Enter at 2× SLO for 2 minutes; exit after 10 minutes healthy." |
| Shared-fate audit | "The fallback can't touch the same replicas." |
| Pre-delegated authority | "The IC executes a pre-signed decision." |
Lean No-Hire Signals
| Signal | Why It Misses the Bar |
|---|---|
| "Return a default value" everywhere | No UX or product decision; defaults can be wrong (price = 0) |
| No recovery plan | Re-enabling everything at once re-triggers the incident |
| Fail-open for auth "for availability" | Creates a security incident to avoid an availability one |
Common False Positives: Knowing Hystrix/resilience4j configuration ≠ designing degraded modes. A chaos-engineering vocabulary ≠ having exercised fallbacks. "We have feature flags" ≠ having a degradation ladder.
Capacity Planning Quick Reference#
Sizing the Degraded Paths#
freed_capacity(tier) = Σ fanout_cost(features shed at tier) / total_fanout_cost
required_tier(load) = lowest tier where capacity × (1 + freed_capacity) ≥ load
snapshot_staleness_max = refresh_interval + job_runtime + one missed run
bulkhead_slots(dep) = dep_qps × dep_timeout_s × 1.5
recovery_ramp = 10% → 50% → 100% over ≥ 6 min per tier
Key Numbers Worth Memorizing#
| Number | Context |
|---|---|
| ~50% | Share of page fan-out typically sheddable (counts, badges, recs) |
| 10–30 s | Automated tier reaction time |
| 5–15 min | Human page → acknowledge → act |
| 3–5× | Exit dwell relative to entry window, to prevent flapping |
| 1–2% | Brownout floor: traffic always on the fallback path |
| 5–15 min | Static snapshot refresh interval |
| 3–10× | Load amplification from naive retries during overload |
| < 30% | Capacity-to-demand ratio that justifies a waiting room |
| 15–30 min | Typical auto-expiry for a bounded fail-open |
Common Pitfalls Checklist#
- Every feature has a criticality class signed by product
- Every tier has numeric entry/exit criteria and hysteresis
- Fallbacks share no dependency, pool, or region with the primary path
- Degradation state is cached locally with last-known-good persistence
- Recovery steps up one tier at a time with a traffic ramp
- Fail-open decisions have signed caps and auto-expiry; authz never fails open
- Fallback paths run in production continuously, and game days happen on a calendar
- The decision-rights table exists and on-call has read it