Hiring BarSupport

Design for Auto-Scaling & Capacity Planning — Staff-Level Case Study

Case study71 min read8 diagrams

Technologies referenced in this case study: Kafka · Redis · PostgreSQL · DynamoDB · Time Series DBs · API Gateways

Related case studies: Load Balancer · Circuit Breakers · Rate Limiting · Flash Sales · Metrics & Monitoring · Degraded Mode · Back-of-Envelope Estimation

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once. Return to individual sections for targeted review.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → one Deep Dive
Deep Dive3+ hrsEverything, including the Principal Lens and appendices
What is Auto-Scaling & Capacity Planning? — Why interviewers pick this topic

Capacity planning decides how much infrastructure you own and when. Auto-scaling decides how fast that amount changes in response to demand. Together they answer one question: when demand arrives, will the capacity already be there, and who paid for it while it was idle?

The prompt arrives in several forms: "Design an auto-scaler," "How would you prepare for Black Friday?", "Your service falls over every Monday at 9am," or as a follow-up inside any other design ("what happens at 10× traffic?").

Before vs After — the product-launch scenario:

Reactive-only autoscaling, CPU target 70%, no pre-warming:
t=0:       Launch email to 40M users. RPS climbs from 20K to 140K in 90 seconds.
t=+30s:    CPU crosses 70%. Autoscaler (60s evaluation window) hasn't fired yet.
t=+75s:    Scale-out requested: +80 instances.
t=+90s:    p99 latency 4s. Clients time out and retry, so effective RPS is 210K.
t=+3min:   New instances boot. JVM cold, caches cold. They take full traffic and fail health checks.
t=+4min:   LB ejects "unhealthy" new instances; autoscaler replaces them. Death spiral.
t=+6min:   New app instances open 6,400 DB connections. Postgres max_connections 5,000 → DB refuses.
t=+25min:  Manual intervention: shed load, cap scale-out, raise pool limits. Launch is a headline.

Staff design (scheduled pre-scale + reactive on concurrency + protected dependencies):
t=-60min:  Scheduled action raises min capacity to forecast peak × 1.3 (load-tested last week).
t=-30min:  Warm pools filled; caches primed by synthetic traffic.
t=0:       Same 140K RPS. Utilization peaks at 58%. p99 stays at 180ms.
t=+10min:  Actual traffic exceeds forecast by 20%. Reactive policy on in-flight requests adds 15%.
t=+12min:  Admission control at the gateway sheds 2% of low-priority traffic during the step.
t=+3h:     Scale-in, gradually, over 45 minutes. Cost of pre-scale: ~$4K. Cost of the alternative: the launch.

Why interviewers reach for this question: it separates engineers who think about capacity as a control loop with delays from those who think of it as "set a CPU threshold." It also exposes whether you see the dependencies that don't scale — databases, third-party APIs, license limits — and whether you can talk about money.

Mechanics Refresher: Scaling Mechanisms
MechanismHow It WorksProsCons
Threshold / step scalingIf metric > X for N minutes, add K instancesSimple, predictableOscillates; step sizes are guesses
Target trackingController keeps metric near target (e.g., CPU 60%), computes desired = current × metric / targetSelf-sizing, few knobsOnly as good as the metric; lags on spikes
Scheduled scalingMin/max capacity changed on a calendarZero lag for known eventsUseless for surprises; stale schedules waste money
Predictive scalingForecast (seasonality model) sets capacity ahead of demandBeats boot time on diurnal curvesWrong on anomalies; needs weeks of history
Vertical scaling (VPA)Change CPU/memory per instanceHelps stateful or single-threaded workOften requires restart; hard ceiling
Queue-driven (KEDA-style)Workers scale on backlog / (throughput per worker × target drain time)Right signal for async workNeeds accurate per-worker throughput
Scale-to-zeroNo instances when idle; start on first requestLowest costCold start on the first request, 100ms–10s+

For most production systems: target tracking on a demand signal (in-flight requests or RPS per instance), with a scheduled or predictive floor for known peaks. CPU is the default everyone picks and the wrong signal for I/O-bound services.


Executive Summary

If you only read one section, read this. Everything else in the case study elaborates on the contrast below.

What This Interview Actually Tests#

Auto-scaling is not a controller question. The formula desired = current × (metric / target) fits on one line.

This is a risk pre-payment question that tests:

  • Whether you understand that a control loop with a 3–5 minute actuation delay cannot react to a 90-second spike, and that someone must pay for capacity before demand arrives
  • Whether you pick a scaling signal that measures demand, not a side effect of demand
  • Whether you find the components that can't scale (databases, connection limits, quotas, third parties) and protect them from the parts that can
  • Whether you can price headroom in dollars and name who pays for idle capacity versus who pays for the outage

The key insight: every capacity decision trades idle dollars for incident risk. Autoscaling narrows the gap between them without closing it. Staff engineers state the headroom they are buying, the scenario it covers, and who signed off on the scenarios it doesn't.

The L5 → L6 → L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Set up an ASG / HPA on CPU at 70%"Asks: cost efficiency, known peaks, or survival of unknown spikes and failures? What's the spike shape vs. our boot time?Asks what the capacity budget is, who owns it, and how capacity is planned across 200 services that share databases and quotas
SignalCPU utilizationDemand signal: in-flight requests / RPS per instance, queue backlog ÷ drain rate; CPU only for CPU-bound workStandardizes signals per workload class across the org; makes the "capacity unit" (RPS per core at SLO) a published number per service
TimingReactive scalingReactive + scheduled/predictive floor; computes spike rise time vs. provisioning time; warm poolsPlans quarterly and annual capacity with finance: reservations, commitments, lead times for hardware and quota increases
Dependencies"The DB will scale too"Identifies non-elastic dependencies; caps scale-out by the DB's connection budget; adds pooling, admission control, load sheddingOwns the shared-dependency capacity model: which teams draw on which databases, and who arbitrates contention
Headroom"Add some buffer"Headroom as a number: N+1 zone, region-failover factor, burst absorption; load-tested to 1.5–2× forecastPrices headroom per tier in $/month vs. expected outage cost; sets org-wide utilization targets by criticality tier
OwnershipPlatform team runs the autoscalerService owners own their scaling policy and capacity model; platform owns the mechanism; a named approver for event pre-scalesDesigns the operating model: capacity reviews, chargeback/showback, error budgets tied to headroom decisions
Why "timing" separates levels

L5: Trusts the autoscaler to react. That works for diurnal curves that rise over hours. It fails for anything that rises faster than the provisioning delay: metric collection (15–60s) + evaluation window (60–180s) + instance boot (60–300s) + warm-up (30–180s). That's 3–10 minutes from spike to useful capacity.

L6: Compares the two time constants out loud. "If the spike rises in 90 seconds and our capacity arrives in 5 minutes, reactive scaling covers nothing. The first 5 minutes must be absorbed by headroom, load shedding, or capacity scheduled in advance. For known events I schedule. For unknown ones I keep headroom and shed."

L7: Extends the same reasoning to longer time constants: quota increases (days), reserved-capacity purchases (a 1–3 year commitment), and datacenter/hardware lead times (months). Capacity planning is the same control-loop problem at a slower frequency.

Why "dependencies" separates levels

L5: Scales the stateless tier and assumes everything behind it keeps up.

L6: Knows that scaling the stateless tier is often what causes the outage. 200 app instances × 30 pool connections = 6,000 DB connections against a 5,000 limit. Caps max instances by the downstream budget, puts a connection pooler in front of the DB, and prefers shedding load at the edge over stampeding the DB.

L7: Recognizes that across an org, 40 services autoscaling independently against one shared database is a tragedy of the commons. Introduces per-consumer quotas on shared dependencies and a capacity review before any service raises its max.

Why "ownership" separates levels

L5: "The platform team manages autoscaling."

L6: Platform owns the mechanism (controllers, node provisioning, warm pools). Service owners own their capacity model: RPS per instance at SLO, max safe instances, dependencies' limits, and the event calendar. Someone specific approves pre-scaling for launches.

L7: Adds money and governance. Showback per team, utilization targets per tier, and a quarterly capacity review where teams defend their headroom. Idle capacity is visible on a dashboard with a team name on it.

The Staff Positions#

PositionRationale
Scale on demand, not on symptomsIn-flight requests or RPS per instance lead; CPU lags and misleads for I/O-bound services
Headroom covers the actuation gapWhatever arrives faster than capacity can be provisioned must be absorbed by pre-paid headroom or shed
Schedule the knowns, react to the unknownsLaunches, sales and Monday mornings are on a calendar. Pre-scale them. Reserve reactive scaling for surprises
Scale out fast, scale in slowlyAdding capacity late causes an outage. Removing it late costs a few dollars. Asymmetric cooldowns (e.g., 1 min out, 10–15 min in)
Cap scale-out by the weakest dependencyMax instances = min(budget, DB connections / pool size, downstream quota / per-instance rate)
Load shedding is part of capacityWhen demand exceeds capacity, the choice is which requests fail, not whether some fail
Load test to 1.5–2× forecast peakForecasts are wrong by 20–50% on launches. Find the knee before customers do

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Cost efficiencyMinimize idle spend on variable loadAggressive target tracking (65–75% util), scale-in, spot/preemptible, scale-to-zero for devUnder-provisioned at spikes; latency SLO breaches during rampSLO met ≥ 99% of minutes; idle ≤ 25%
Reliability through known peaksLaunches, sales, sports events, Monday 9amForecast + scheduled pre-scale + load test + warm pools + admission controlForecast wrong; dependency ceiling hitZero SLO breach during planned events
Survival of unknown spikes and failuresViral traffic, retry storms, zone/region lossStanding headroom (N+1 zone, region-failover factor), load shedding, priority tiersPaying for idle; still overwhelmed by > headroom spikesGraceful degradation: priority traffic served, rest shed

🎯 Staff Move: "I'll design for reliability through known peaks with a survival floor. That means a user-facing stateless tier with a strong diurnal curve plus planned launches, and enough standing headroom to lose one availability zone without paging anyone. Pure cost efficiency is a tuning exercise on top of that design, and the reverse doesn't hold."

The Five Fault Lines#

#Fault LineThe Tension
1Scaling SignalCPU (universal, lagging, misleading for I/O) vs. demand signals (RPS, concurrency, backlog: accurate, service-specific) vs. latency (what users feel, but noisy and ambiguous)
2Reactive vs PredictiveReactive is always right about the present and always late. Predictive is early and sometimes wrong
3Cold Start vs Warm CapacityScale-from-cold is cheap but slow (minutes). Warm pools and provisioned concurrency are fast but cost money while idle
4Headroom vs CostEvery 10% of headroom is 10% of the compute bill. Every 10% less is a bigger spike you can't absorb
5Central Capacity vs Service OwnershipA capacity team with global view and budget authority vs. service teams who know their workloads and move fast

In the Wild: Real Production Systems#

Why this section belongs here: citing well-known production practice shows you've studied operational reality. Use these as one-sentence anchors in the interview.

Netflix — Predictive Scaling and Regional Evacuation#

Netflix publicly described Scryer, a predictive autoscaling engine that forecasts demand from historical patterns and provisions capacity ahead of the diurnal ramp, because reactive scaling lagged their steep evening rise. Separately, Netflix's regional failover exercises (Chaos Kong) evacuate a whole AWS region and shift its traffic to the others, which requires each remaining region to absorb a large traffic increase within minutes.

Staff insight: two ideas worth citing. Predictive scaling exists because boot time exceeds ramp time. Region evacuation means your headroom target is set by the failover scenario, not by normal peak. Say both.

Facebook — Kraken: Load Testing With Live Traffic#

Facebook's Kraken (published at OSDI 2016) shifts live production traffic toward a subset of servers, clusters or regions to measure their true capacity limits, guided by health metrics that stop the test before users are hurt.

Staff insight: synthetic load tests miss real request mixes, cache hit rates and dependency behavior. Kraken's lesson is that the most accurate capacity number comes from production, under controlled conditions. Mention it when asked "how do you know your capacity?"

Kubernetes & AWS — Controllers With Built-In Asymmetry#

The Kubernetes Horizontal Pod Autoscaler evaluates metrics every 15 seconds by default and applies a default 5-minute scale-down stabilization window, so it scales out quickly and in cautiously. AWS Auto Scaling offers target tracking, scheduled actions, warm pools and a predictive scaling mode driven by historical load. Karpenter and the Cluster Autoscaler add nodes when pods can't be scheduled.

Staff insight: the defaults encode the Staff position. Scale out fast because under-provisioning is expensive. Scale in slowly because flapping is expensive and idle is cheap. Being able to quote these defaults shows you've operated these systems.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Autoscale on CPU at 70%""The service is I/O-bound and CPU sits at 25% while latency is 3s. What now?"Whether you pick signals that measure demand
"The autoscaler handles spikes""Traffic triples in 60 seconds. How long until new capacity serves traffic?"Whether you know the actuation delay
"We'll scale the app tier""What happens to the database when you go from 50 to 300 instances?"Non-elastic dependency awareness
"Keep 30% headroom""Why 30? What does it cost and what does it cover?"Can you price and justify headroom?
"Predictive scaling""The forecast was wrong by 2×. What happens?"Fallbacks and guardrails on predictions
"Load test before launch""Your load test passed, but production fell over. Why?"Test realism: data, cache state, dependencies, request mix

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: the controller has three inputs. It measures demand (reactive), takes a floor from forecasts and the event calendar (predictive), and takes a ceiling from the capacity model's dependency budgets. The gateway's admission control is the safety valve for the gap between demand and capacity. The warm pool shortens the actuation delay from minutes to ~30 seconds. The database sits outside the elastic tier on purpose: it scales in hours, and the design protects it rather than pretending it's elastic.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Signal"CPU at 70%""In-flight requests per instance for request/response work, backlog ÷ drain rate for queues. CPU only when the service is CPU-bound."
Timing"Autoscaler reacts""Capacity arrives 3–5 minutes after the spike. Headroom and shedding cover that gap. Known events are pre-scaled."
Cold start"Instances boot quickly""Boot plus warm-up is 2–5 min on VMs and 10–60s on warm nodes. Warm pools, slow-start at the LB, pre-warmed caches."
Headroom"Add a buffer""Target 60% utilization. That covers losing 1 of 3 zones (×1.5) with a small burst margin. It costs ~40% idle, priced at $X/month."
Dependencies"DB scales too""Max instances is capped by DB connections ÷ pool size. A pooler, load shedding and per-consumer quotas protect it."
Validation"Load test once""Load test to 1.5–2× forecast with production-like data. Find the knee. Rerun after every major change."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
HPA metric sync period15s (default)Reactive loop's fastest reaction
HPA scale-down stabilization300s (default)Built-in asymmetry: in slowly
VM boot to in-service60–300sWhy reactive scaling can't catch fast spikes
Container start on an existing node2–30s (+ app warm-up)Why node-level headroom matters
New node provisioning (cluster autoscaler)60–180sPods pending while nodes boot
JVM/JIT warm-up to steady-state latency30–180sNew instances serve slowly at first
Serverless cold start~100ms–1s typical; several seconds for heavy runtimesScale-to-zero's hidden latency tax
Target utilization (user-facing)50–65%Leaves room for zone loss and bursts
Target utilization (batch / async)80–90%Queues absorb bursts; latency isn't the SLO
Zone-loss headroom (3 AZs)Each AZ must carry 1.5× normal → run at ≤ 66%N+1 at zone granularity
Region-failover headroom (3 regions)Each region takes +50% → run at ≤ 60–66%Set by the evacuation scenario
Load-test multiple1.5–2× forecast peakForecast error on launches is 20–50%
Latency kneeUsually 70–85% utilizationQueueing delay grows as 1/(1−ρ): 80% → 5×, 90% → 10× service time
Retry amplification2–3× offered load during brownoutsWhy capacity math must include retries

Interview Walkthrough

The most common mistake: candidates draw an autoscaling group, write "CPU > 70% → add instance," and stop there. Everything that decides whether the launch survives comes after that sentence: the actuation delay, the dependency that doesn't scale, the load test that lied, and who paid for the headroom.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional goal in one sentence:

"We want the service to meet its latency SLO as demand changes, without paying for peak capacity 24/7."

Then frame the real requirements:

"Three intents pull in different directions: minimizing cost on a variable curve, surviving known peaks like launches, and surviving unknown spikes and failures like a zone outage. I'll design for known peaks with a survival floor. My constraints are p99 under 300ms, at most 1% of requests shed at the worst minute, the ability to lose one availability zone without paging, and a cost target of no more than ~40% idle averaged over a day."

Establish the shape of demand, because it determines everything:

"What does the curve look like? I'll assume a diurnal pattern from 20K RPS at night to 100K at peak, rising over about 2 hours, plus planned launches that step from 100K to 250K in under 2 minutes. The diurnal curve is slow enough for reactive scaling. The launch step isn't."

🎯 Staff Move: Asking for the rise time of demand, not just the peak, is the question that separates levels. Compare it to your provisioning time and you know which tool you need before drawing anything.


Phase 2: Core Entities & API (1–2 minutes)#

The nouns:

  • Capacity unit — the measured throughput of one instance at SLO (e.g., 400 RPS per 4-vCPU instance at p99 < 300ms). The single most important number in the design.
  • Scaling policy — signal, target, min, max, cooldowns, step limits.
  • Capacity floor — scheduled or predicted minimum for a time window.
  • Dependency budget — max connections/RPS each downstream grants this service.
  • Event — a calendar entry with expected multiplier, owner, and pre-scale window.

The control API (platform-facing, not end-user):

PUT  /services/{svc}/scaling-policy
     { signal: "inflight_per_instance", target: 60, min: 40, max: 300,
       scale_out: { max_step_pct: 50, cooldown_s: 60 },
       scale_in:  { max_step_pct: 10, stabilization_s: 600 } }

POST /services/{svc}/capacity-floors
     { start: "2026-11-27T13:00Z", end: "...", min_instances: 220, reason: "BF launch", approver: "..." }

GET  /services/{svc}/capacity-model
     → { rps_per_instance_at_slo: 400, max_safe_instances: 280,
         limiting_dependency: "orders-db connections", headroom_pct: 38 }

🎯 Staff Move: "The capacity model endpoint matters more than the policy. If a team can't tell me their RPS per instance at SLO and their limiting dependency, their autoscaling policy is a guess. Everything else I design assumes that number exists and is re-measured regularly."


Phase 3: High-Level Architecture (≤5 minutes)#

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk it in 90 seconds:

  1. The service emits in-flight requests per instance every 15 seconds.
  2. The reactive controller computes desired = ceil(current × observed / target) and clamps it between the floor and the ceiling.
  3. The floor comes from the forecast (trailing-weeks seasonality) and the event calendar (pre-scales for launches).
  4. The ceiling comes from the capacity model: the smallest of budget, DB connection budget ÷ per-instance pool size, and third-party quota ÷ per-instance call rate.
  5. New capacity comes from the warm pool (~30s) before cold boot (~3 min). The LB slow-starts new targets over 60 seconds.
  6. If demand exceeds what the fleet can serve, the gateway sheds lowest-priority traffic first. It returns 503 with Retry-After rather than letting latency collapse for everyone.

🎯 Staff Move: "That's the design most people draw, minus the floor and ceiling. The floor handles what we know is coming. The ceiling stops us from DDoSing our own database. Admission control handles what neither covers. Now let me show where it breaks."


Phase 4: Transition to Depth (1 minute)#

"Four places this gets interesting: choosing the signal, the gap between spike rise time and provisioning time, the dependencies that don't scale, and how much headroom to buy and who pays for it. Which do you want first?"

If the interviewer has no preference, lead with the actuation gap. It's the most counter-intuitive point, and the other topics follow from it.


Phase 5: Deep Dives (25–30 minutes)#

Deep dive A: The actuation gap (6–8 min)

"Add up the delays. Metric scrape 15s, evaluation window 60s, API call and scheduling 10s, VM boot 90s, app start 30s, JIT and cache warm-up 60s, LB health checks and slow-start 30–60s. That's about 5 minutes from spike to useful capacity. A launch that goes from 100K to 250K RPS in 90 seconds is over before the autoscaler contributes anything."

Three tools for the gap, in order of preference:

  1. Know it's coming and pre-scale. The event calendar and forecast floor cover launches and diurnal ramps.
  2. Shorten the delay. Warm pools (pre-booted, stopped instances attach in ~30s), pre-pulled images, node headroom so pods don't wait on new nodes, and pre-warmed caches.
  3. Absorb the remainder with standing headroom plus admission control that sheds the lowest-priority traffic.

"For a truly surprise spike, say a viral post, the formula is: headroom must cover (spike rate × actuation delay). If we can tolerate a 2× surprise within 5 minutes, we run at ~50%. That's expensive, so I'd do it only for tier-0 services and shed for the rest."

Deep dive B: Choosing the signal (5–6 min)

SignalUse whenFailure
CPUCPU-bound compute (encoding, crypto, rendering)I/O-bound services sit at 20% CPU while latency explodes
In-flight requests / concurrencyRequest/response servicesNeeds a measured "good concurrency per instance"
RPS per instanceHomogeneous request costMix shifts (e.g., more expensive search queries) mislead it
Queue backlog ÷ drain rateAsync workersNeeds per-worker throughput; poison messages look like load
LatencyLast-resort guardrailRises for reasons scaling can't fix: a slow DB, GC. Scaling out makes a slow DB worse

"Latency is the worst primary signal, because the most common cause of high latency is a slow dependency. Scaling out on latency then adds connections to the thing that's already struggling. I use concurrency as the primary signal and treat latency as an alert, not a scaling input."

Deep dive C: The dependencies that don't scale (6–7 min)

"Our app tier goes from 40 to 300 instances. Each has a 20-connection DB pool, so that's 6,000 connections against a Postgres primary configured for 5,000, whose sweet spot is much lower because each Postgres connection is a process. Scaling out converts an app-tier problem into a database outage."

Mitigations:

  • Connection pooler (PgBouncer in transaction mode) multiplexes thousands of client connections onto ~200–500 server connections.
  • Max instances derived from the dependency budget, recomputed when pool sizes change.
  • Per-consumer quotas on shared dependencies, so one service's scale-out can't starve another.
  • Cache and queue in front of writes to decouple app-tier elasticity from DB write capacity (see Scaling Writes).
  • Third-party quotas as hard ceilings, with degraded modes when they're hit (see Degraded Mode).

Deep dive D: Headroom and who pays (5–6 min)

"Three AZs, each carrying a third. Lose one and the other two must carry 1.5× their load, so steady-state utilization can't exceed ~66% of the knee. The knee for this service is about 80% CPU-equivalent, so the target is ~0.66 × 80% ≈ 53%, call it 55%. On a $400K/month compute bill that's roughly $180K/month of 'idle.' It isn't idle. It's the price of zone failure being a non-event. Finance should see that line item with that label."


Phase 6: Wrap-Up (2–3 minutes)#

"The design has three layers: a forecast and calendar floor for what we know, reactive target tracking on concurrency for what we don't, and admission control for the gap. It's capped by a ceiling that protects dependencies that can't scale. Headroom is set by the zone-failure scenario and priced explicitly. The capacity model, meaning RPS per instance at SLO and the limiting dependency, is re-measured every quarter and after every major release."

The org close:

"Mechanically, the platform team owns the controllers. Each service team owns its capacity model and its event calendar entries. A capacity review each quarter makes idle spend and headroom visible per team. Most capacity incidents I'd expect aren't controller bugs. They're a team that didn't tell anyone about a launch, or a shared database nobody owned the budget for."

🎯 Staff Move: End with where the next outage comes from. For capacity, it's almost always a process gap (an unannounced event, a stale capacity model) or a shared dependency, rarely the autoscaler itself.


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
10 min on the scaling formulaDerives target-tracking math, cooldown tuningStates the formula in one line, moves to delays and dependencies
No demand shape"Traffic spikes""Rise time 90s vs. provisioning 5 min, so the gap is covered by X"
Only the elastic tierDesigns ASG/HPA and stopsSpends 5+ min on the DB, pools and third-party quotas
Headroom as vibes"Some buffer""55% target, covers AZ loss, costs $180K/month"
Load test as a checkbox"We'll load test"Describes realistic data, cache state, request mix, and finding the knee
No org storyEnds on autoscaler configEnds on capacity reviews, event calendar, dependency budgets

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Capacity is where engineering meets finance, and where a design meets the rest of the company. The autoscaler itself is a commodity: every cloud and Kubernetes distribution ships one. What isn't a commodity is the judgment around it. Which signal actually measures demand for this workload? Which time constants make reactive scaling useless? Which components silently won't scale? How much should the company pay to make a zone failure boring?

It's also the purest test of the Staff thesis. The L5 answer ("autoscale on CPU") is correct most of the time. The L6 answer is opinionated about the minute it isn't: the launch spike, the zone loss, the DB connection cliff. And it names who gets paged and who paid for the headroom.

1.2 The L5 vs L6 vs L7 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 vs L7 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"How fast does demand rise compared to how fast capacity arrives, and who pays for the difference?"

  • If demand rises slower than provisioning (diurnal curves over hours), reactive scaling works and the question is tuning.
  • If demand rises faster and it's known (launches, sales), you pre-scale and the question is process: who files the event and who approves it.
  • If demand rises faster and it's unknown, you choose between paying for standing headroom (finance pays) and shedding load (some users pay). You say which, for which traffic tier, and who signed off.

2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Cost efficiency. Common for internal services, batch processing, dev/test and variable workloads without strict latency SLOs. Run hot (70–90% utilization for async), scale in aggressively, use spot/preemptible instances with interruption handling, and scale to zero outside business hours. The risk is latency SLO breaches during ramps, which the team explicitly accepts.

Reliability through known peaks. Consumer products with launches, flash sales, sports events, tax deadlines and Monday mornings. The work is forecasting, calendaring, load testing and pre-scaling, and making sure every dependency on the path is sized for the event, not only the stateless tier. Failures here are usually process failures: marketing sent the email early, the forecast used last year's numbers, or the DB wasn't included in the load test. See Flash Sales.

Survival of unknown spikes and failures. Tier-0 services such as login, checkout, and core APIs other services depend on. They need standing headroom for zone and region loss, load shedding with priority tiers, and protection from retry storms. Here the design question is how the service degrades, not how it scales.

🎯 Staff Move: "These aren't mutually exclusive. They're tiers. I'd classify each service: tier-0 gets survival headroom, consumer-facing gets known-peak planning, and internal/batch gets cost efficiency. One autoscaling policy for all 200 services is how you get either an outage or a bill."

2.2 When NOT to Auto-Scale#

SituationWhy autoscaling is wrongDo instead
Stateful primaries (DB primaries, Kafka brokers, ZooKeeper)Adding a node means rebalancing data, which takes hours and adds load mid-incidentPlan capacity quarterly; scale vertically ahead of need; shard deliberately
Workloads with long warm-up (large caches, ML models loading 20 GB)New instances aren't useful for minutes, and cold caches hammer the backendFixed capacity + pre-warming; scale on schedule
Spiky load that's shorter than provisioning timeCapacity arrives after the spike ends, so you pay for it and it doesn't helpHeadroom + shedding + queueing
License-bound or quota-bound dependenciesMore instances just hit the external limit fasterCap at the quota; degrade gracefully
Very small fleets (≤ 3 instances)Scale-in to 1 removes redundancy; each step is a 33% changeFixed N+1 with a scheduled floor
Latency-critical services with tight tails where warm-up hurts p99New cold instances inflate tail latencyOver-provision; scale on schedule only

🎯 Staff Move: "The database is where I stop autoscaling and start capacity planning. I'll scale the stateless tier elastically and plan the stateful tier a quarter ahead, with growth tracking and a 6-month runway alert."

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy it mattersWhat I'd assume out loud
Demand shape and rise timeDecides reactive vs. scheduled vs. headroomDiurnal 5× over 2h; launches 2.5× in 90s
Latency SLOSets the utilization target (knee)p99 < 300ms
Workload typeCPU-bound vs. I/O-bound → signal choiceI/O-bound API calling a DB and one third party
Failure scenarios to surviveSets headroomLose 1 of 3 AZs without paging
Cost constraintsBounds headroom≤ 40% idle daily average; tier-0 exempt
Dependency limitsSets the ceilingPostgres 5,000 connections; third party 2,000 RPS
Cloud vs. own hardwareActuation time: minutes vs. monthsPublic cloud, with reserved commitments for the baseline

2.4 Precise Terminology#

TermMeaningWhy the precision matters
Capacity unitThroughput of one instance at SLOEverything else is derived from it; must be measured, not assumed
KneeUtilization where latency starts rising non-linearlyTarget utilization sits below it with margin
Headroom(Capacity − demand) ÷ capacity at a point in timeMust be tied to a scenario (AZ loss, 2× surprise)
Actuation delayTime from demand change to useful new capacityCompared against rise time
Floor / ceilingMin capacity from forecast/schedule; max from budget/dependenciesClamps the reactive controller
Stabilization windowPeriod the controller considers before scaling inPrevents flapping
Warm poolPre-initialized but idle/stopped capacityCuts actuation delay at a small cost
Load sheddingRejecting some requests on purpose to protect the restThe honest answer when demand > capacity
Retry amplificationOffered load multiplied by client retries during slownessCapacity math must include it
Reserved / committed capacityCapacity bought for 1–3 years at a discountThe baseline; autoscaling handles the variable part above it

3. The Five Fault Lines#

3.1 Fault Line 1: Scaling Signal#

StrategyWhat WorksWhat BreaksWho Pays
CPU utilizationUniversal, no instrumentation, correct for CPU-bound workI/O-bound services look idle while saturated; GC and noisy neighbors distort itUsers during brownouts the scaler never saw
Concurrency (in-flight per instance)Directly measures work in progress; Little's Law ties it to RPS × latencyRises when a dependency slows, so it can scale out into a slow DBThe DB, unless the ceiling protects it
RPS per instanceLeading indicator, easy to reason aboutWrong when request mix shifts (expensive queries)Users when mix shifts
Queue backlog ÷ (drain rate per worker × target drain time)The right signal for async work; directly expresses the SLO ("drain within 5 min")Poison messages and retry loops inflate backlogThe downstream that the workers hammer
LatencyWhat users feelAmbiguous cause; scaling out on dependency-induced latency amplifies the problemThe dependency and every other consumer of it
Custom business signal (e.g., active sessions)Leading, domain-accurateNeeds a model mapping signal → capacityThe team maintaining the model

Staff default: concurrency for synchronous services, backlog-over-drain-rate for async, CPU only for CPU-bound workers. Latency is an alert, never a primary scaling input. Always pair the reactive signal with a ceiling, so that "concurrency rising because the DB slowed down" can't turn into 300 instances hammering the DB.

"Little's Law gives me the target. If one instance meets SLO at 400 RPS with 150ms average latency, it holds 400 × 0.15 = 60 requests in flight. I target 60 and scale on it. If latency doubles because the DB slows, in-flight doubles too, and the ceiling is what stops me from making it worse."

🧭 Principal Move: "I'd publish a signal standard per workload class (request/response, stream consumer, batch, GPU inference) so 200 teams don't rediscover that CPU is wrong for I/O-bound services one incident at a time."

3.2 Fault Line 2: Reactive vs Predictive#

StrategyWhat WorksWhat BreaksWho Pays
Reactive onlyAlways tracks reality; no model to maintainLate by the actuation delay; fails on fast rampsUsers in the first 3–5 minutes of every ramp
ScheduledZero lag for calendared eventsStale schedules waste money; unannounced events uncoveredFinance (stale floors), users (missed events)
Predictive (seasonality forecast)Pre-provisions diurnal/weekly rampsWrong on holidays, launches, anomalies; needs weeks of historyUsers when the forecast is low; finance when it's high
Predictive floor + reactive aboveForecast sets the minimum; reactive handles the surprise upsideTwo systems to reason about; the floor must never cap the reactive sidePlatform team's complexity budget

Staff default: forecast and schedule set the floor. Reactive scaling tracks the actual signal above it. The controller takes max(floor, reactive_desired), clamped by the ceiling. A wrong forecast therefore costs money (too high) or falls back to reactive behavior (too low). It never removes the reactive path.

When to deviate: services with flat load can skip prediction. Services with extreme predictable spikes (a daily 9:00:00 batch trigger, a sports kickoff) should use scheduled capacity, since forecasting adds nothing there.

🎯 Staff Move: "Predictive scaling should only ever raise the floor. The day it's allowed to lower capacity below what reactive wants is the day a bad forecast becomes an outage."

3.3 Fault Line 3: Cold Start vs Warm Capacity#

Diagram: 3.3 Fault Line 3: Cold Start vs Warm Capacity
StrategyWhat WorksWhat BreaksWho Pays
Cold scale-outNo idle cost3–5 min actuation; cold instances serve slowly and can fail health checksUsers during ramps
Warm pools (pre-booted, stopped)Attach in ~20–40s; storage cost only while stoppedPool can be exhausted; images drift if not refreshedSmall standing cost
Node overprovisioning (placeholder pods with low priority)Pods schedule instantly on spare nodesPaying for idle nodesFinance (~5–15% of cluster)
Provisioned concurrency (serverless)Eliminates cold starts for N concurrentPay per provisioned unit whether used or notFinance
Pre-warming (synthetic traffic, cache priming, JIT warm-up)New instances hit steady-state latency before real trafficNeeds representative warm-up trafficThe service team's effort
LB slow-startNew targets ramp traffic over 30–120sSlows how fast new capacity absorbs loadSlight delay in relief

Staff default: warm pool for VM fleets, low-priority placeholder pods for Kubernetes (~10% node headroom), pre-warming on startup (hit the top 1,000 cache keys, run a synthetic request loop until p99 stabilizes) before registering with the LB, and LB slow-start of 60s.

🎯 Staff Move: "A new instance that isn't warm is worse than no instance. It takes a full share of traffic, serves it slowly, fails health checks, and gets replaced by another cold instance. Warm-up has to finish before registration."

3.4 Fault Line 4: Headroom vs Cost#

Target utilizationCoversMonthly cost of headroom on a $500K fleetWho Pays if Wrong
85%Almost nothing; latency near the knee~$75KUsers: latency spikes, zone loss is an outage
70%Moderate bursts; not zone loss~$150KUsers during AZ failure
60%Loss of 1 of 3 AZs (×1.5) with a thin margin~$200KBalanced
50%AZ loss + ~1.3× surprise, or region failover in a 3-region setup~$250KFinance
35%Region loss in a 2-region active-active setup~$325KFinance, heavily

Headroom cost ≈ fleet cost × (1 − utilization), assuming the fleet is sized to peak. Autoscaling reduces it off-peak but not at peak.

Staff default: set the target by the failure scenario you've committed to survive, then check it against the latency knee. For a 3-AZ user-facing service, ~55–60%. For async workers, 80–90% (the queue is the headroom). For tier-0 services with region-failover requirements, whatever the failover math says, even if it's 45%.

"I'd rather say '60%, because losing one of three zones multiplies load by 1.5 and our knee is at ~85%' than '60% is industry standard.' The first is a decision. The second is a habit."

🧭 Principal Move: "Headroom should be a line item with a scenario attached: '$200K/month buys AZ-loss tolerance for checkout.' Then leadership can make the trade explicitly. Most orgs pay for headroom they can't explain and skip headroom they need."

3.5 Fault Line 5: Central Capacity vs Service Ownership#

ModelWhat WorksWhat BreaksWho Pays
Central capacity team plans everythingGlobal view; purchasing leverage; shared-dependency arbitrationBottleneck; can't know each workload; teams hide growthService teams waiting on tickets
Each team fully owns its capacityFast; knowledgeableNo one owns shared DBs and quotas; tragedy of the commons; duplicated toolingWhoever shares a dependency with the loudest team
Platform mechanism + service-owned models + central reviewPlatform owns controllers and tooling; teams own their capacity models and events; a central group owns shared budgets and the quarterly reviewNeeds discipline and a review cadenceEveryone, a little

Staff default: the hybrid. Specifically:

  • Platform team: autoscaling controllers, warm pools, node provisioning, load-test tooling, capacity dashboards.
  • Service team: capacity model (RPS per instance at SLO, limiting dependency), scaling policy, event calendar entries, load tests before launches.
  • Shared-dependency owners (DB platform, API platform): publish per-consumer budgets and enforce them.
  • Capacity review (quarterly): growth vs. runway, headroom vs. cost, top risks.

🎯 Staff Move: "The failure I'd design against is the unannounced launch. The fix is process, not technology: an event calendar that marketing and product write to, and a pre-scale approval with a named owner."


4. Failure Modes & Operational Reality#

4.1 The Scale-Out Death Spiral#

Diagram: 4.1 The Scale-Out Death Spiral
t=0:       Spike: 2× traffic in 60s.
t=+45s:    p99 from 200ms to 2s. Health checks (1s timeout) begin failing on busy instances.
t=+60s:    LB ejects 30% of instances. The rest go to 140% of capacity.
t=+90s:    Autoscaler adds 60 instances. Clients retrying, so offered load is 3×.
t=+4min:   New instances register cold, get a full share, fail health checks, get replaced.
t=+5min:   New app instances add 1,200 DB connections. DB CPU 100%. All instances slow.
t=+8min:   Every instance is "unhealthy" at some point. Effective capacity near zero.

Detection: lb.healthy_host_ratio falling while autoscaler.desired_capacity rises; db.active_connections near max; client.retry_ratio > 1.5.

Mitigation: freeze scale-in and instance replacement (stop the churn); enable gateway admission control to cut offered load to known capacity; separate liveness from readiness checks (a slow instance is busy, not dead); cap scale-out at the DB budget.

Prevention: health checks that test "the process is alive," not "the process is fast"; LB slow-start; warm-up before registration; retry budgets in clients (≤ 10% extra load); circuit breakers (Circuit Breakers); ceiling from dependency budgets.

Owner: service on-call leads; platform on-call freezes the controllers; DB on-call guards the database.

4.2 Scaling on a Symptom: Latency-Driven Scale-Out into a Slow Database#

A slow query plan makes DB latency jump from 5ms to 80ms. App concurrency rises 16×, and the reactive controller (on concurrency) scales from 60 to 300 instances. The DB now has 5× the connections and is slower still.

Detection: autoscaler.scale_out_events correlated with db.query_latency_p99 rising before request rate did. The signature is capacity rising while RPS stays flat.

Mitigation: freeze scaling; fix or roll back the plan; shed.

Prevention: a flat-RPS guard: don't scale out if RPS hasn't risen more than 20% in the window, since rising concurrency with flat RPS means a dependency slowed. Also the ceiling from the dependency budget.

Owner: service team (policy), DB team (query plan).

4.3 Scale-In Flapping and Connection Drops#

A too-short stabilization window makes the fleet oscillate between 80 and 120 instances every 4 minutes. Each scale-in terminates instances holding long-lived connections (WebSockets, gRPC streams), which reconnect and spike load on the rest, which triggers scale-out.

Detection: autoscaler.direction_changes_per_hour > 4; connections.reset_total spikes aligned with scale-in.

Mitigation: lengthen the scale-in stabilization window to 10–15 min; cap the scale-in step at 10% per window; use connection draining (up to 5 min) before termination.

Owner: service team.

4.4 Capacity Ceiling Nobody Knew About#

Launch day: the fleet scales perfectly to 280 instances, and the payment provider's API rate limit of 2,000 RPS fails 30% of checkouts. Or the cloud account's vCPU quota in the region stops scale-out at 180 instances. Or the NAT gateway's port allocation is exhausted.

Detection: autoscaler.scale_out_failures{reason=quota}; third-party 429 rates; nat.port_allocation_errors.

Mitigation: emergency quota increase (hours to days, which is too slow for today); shed low-priority traffic; queue payments for async capture.

Prevention: a limits inventory per service: cloud quotas, third-party limits, NAT/IP, license counts. Checked in the pre-launch review, with alerts at 70% of every limit.

Owner: service team for the inventory; platform for cloud quotas.

4.5 Forecast Wrong by 2×#

Predictive floor built from the last 4 weeks; a holiday week cuts the actual load in half. Paying 2× for a week is a cost incident, not an outage. The opposite case, a viral event at 2× the forecast, falls back to reactive plus headroom plus shedding.

Detection: forecast.error_pct (actual/forecast − 1) tracked daily; alert when |error| > 30% for 2 hours.

Owner: platform (forecasting), service team (event calendar hygiene).

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Death spirallb.healthy_host_ratio ↓ while desired_capacity ↑Whole service + shared DBFreeze replacement, shed, cap at DB budgetService on-call + platform
Scale into slow dependencyScale-out with flat RPSDependency + all its consumersFlat-RPS guard, freeze, fix queryService + DB team
Flappingdirection_changes_per_hour > 4Connection-heavy clientsLonger stabilization, smaller steps, drainingService team
Hidden ceiling (quota, third party)scale_out_failures{reason}, third-party 429sFeature or whole serviceShed, queue, emergency quotaService team + platform
Forecast errorforecast.error_pctCost (high) or latency (low)Reactive covers low; review floorsPlatform
Warm pool exhaustedwarm_pool.available == 0Ramp latencyCold fallback; resize poolPlatform
Unannounced eventSudden step in RPS with no calendar entryService, sometimes company-wideShed; emergency pre-scaleProduct/marketing + service team
Retry stormclient.retry_ratio > 1.5All services sharing the pathRetry budgets, jittered backoff, shedClient owners + gateway team

🎯 Staff Move: "When I read this matrix, the autoscaler is at fault in maybe one row. The rest are process, dependencies and client behavior. That's why I'd spend my design time on those."


5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Framing"Autoscale on CPU"Demand shape, rise time vs. actuation delay; intent by tierCapacity as a budget and a portfolio: baseline commitments + elastic + headroom by tier
SignalCPU/memoryConcurrency/backlog with Little's Law; latency as alert onlyOrg-wide signal standards per workload class
TimingReactive + cooldownsFloor (forecast/schedule) + reactive + warm pools + shedding for the gapPlans quarterly and annual capacity with finance; hardware and quota lead times
DependenciesAdds replicasCeiling from DB/quotas; pooling; flat-RPS guardShared-dependency budgets and arbitration across teams
Headroom"Some buffer"Target tied to a failure scenario; pricedUtilization targets by tier as policy; headroom as a line item with a scenario
Validation"Load test"Production-like load tests to 1.5–2×; find the kneeContinuous capacity measurement in production (Kraken-style); game days
OwnershipPlatform teamPlatform mechanism, service-owned models, event approverOperating model: capacity reviews, showback, error-budget links

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Compares time constants"Spike rises in 90s, capacity arrives in 5 min. Reactive covers nothing here, so I pre-scale."
Right signal"It's I/O-bound. I scale on in-flight requests: 400 RPS × 150ms = 60 per instance."
Protects the non-elastic"Max instances is 5,000 DB connections ÷ 20 per pool, minus a margin: 220."
Prices headroom"60% target covers AZ loss and costs ~$200K/month. That's the price of AZ failure being a non-event."
Knows the process failure"The outage I'd fear is the launch nobody told us about. I'd build an event calendar with an approver."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
CPU for everythingMisreads I/O-bound saturation, which is the most common case
"The autoscaler will handle spikes"Ignores the actuation delay
Scales the DB like the app tierDoesn't know stateful systems rebalance slowly
No numbers on headroom or costCan't reason about the tradeoff the question is about
No load sheddingAssumes capacity always catches up with demand

5.4 Common False Positives#

  • Knowing HPA YAML fields ≠ capacity judgment. Configuration fluency doesn't mean the candidate knows which signal to pick.
  • Queueing theory derivations ≠ design. M/M/1 math is useful for one sentence ("latency ∝ 1/(1−ρ)"), not for ten minutes.
  • "We'll use serverless" ≠ solving scaling. It moves the problem to cold starts, concurrency limits and downstream protection.
  • Kubernetes cluster autoscaler internals ≠ capacity planning. Node provisioning is one delay in the chain.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minIntents by tier; demand shape; SLO; failure scenario to survive
Entities3–5 minCapacity unit, policy, floor, ceiling, event
Architecture5–10 minController with floor/ceiling; warm pool; admission control
Deep dive 110–18 minActuation gap: delays, pre-scale, warm-up
Deep dive 218–26 minSignal choice; scaling into a slow dependency
Deep dive 326–36 minDependencies & ceiling; headroom & cost
Deep dive 436–41 minLoad testing and validating the capacity model
Wrap-up41–45 minOwnership, process, capacity review

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response
"Traffic triples in 30 seconds, no warning."Unknown-spike postureHeadroom + admission control by priority; reactive recovers minutes later; state who is shed
"Cut compute cost by 30%."Cost levers without breaking reliabilityTier services; raise async targets to 85%; scale-to-zero non-prod; commit the baseline with reservations; spot for stateless batch
"It's a Kafka consumer, not an API."Signal for asyncLag ÷ (per-consumer throughput × target drain time); ceiling = partition count
"Now it's GPU inference."Scarce, slow-to-provision capacityMinutes-to-hours provisioning; queue + batch; capacity reservations; admission by priority
"How do you know 400 RPS per instance is right?"Measurement disciplineLoad test to the knee in a prod-like env; verify in prod with traffic shifting; re-measure after releases

6.3 What to Deliberately Skip#

  • The exact target-tracking algorithm and PID tuning. One sentence is enough.
  • Instance type selection details. Mention only if CPU:memory ratio matters.
  • Cloud-provider-specific policy syntax.
  • Detailed queueing theory beyond the knee.

6.4 Follow-Up Questions to Expect#

  1. "What's your scale-in policy?" Slow (10–15 min stabilization), small steps (≤ 10%), connection draining, never below the floor.
  2. "How do you avoid scaling on noise?" Evaluate over 1–3 min windows, require N consecutive breaches for scale-in, but allow a single strong breach to trigger scale-out.
  3. "How do you handle multi-tenant load on shared services?" Per-tenant quotas at the gateway (Rate Limiting) so one tenant's spike doesn't drive everyone's capacity.
  4. "What about stateful services?" Plan quarterly; vertical first; add read replicas ahead of need; shard deliberately (Database Sharding).
  5. "How do you load test safely?" A prod-like environment with production-sized data, or production traffic shifting with automated abort thresholds.
  6. "How do you forecast?" Weekly seasonality × growth trend × event multipliers; compare forecast error daily.
  7. "Who decides to pre-scale for Black Friday?" The service owner files it; the capacity reviewer approves; finance sees the cost; the platform team executes.

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design auto-scaling for our API."

Staff Answer

"First, what does the demand curve look like? How fast does it rise, and are the big peaks known in advance? I'll assume a diurnal 5× swing over ~2 hours plus planned launches that step 2.5× in 90 seconds. Second, what are we optimizing: cost, surviving known peaks, or surviving failures? I'll design for known peaks with a survival floor of losing one AZ without paging. Third, what's the workload? If it's I/O-bound, CPU is the wrong signal. The design is a controller with three inputs: a floor from forecast and event calendar, a reactive target on in-flight requests per instance, and a ceiling from our dependencies' budgets. Admission control at the gateway covers the gap. I'll spend most of our time on the actuation gap, the dependencies that don't scale, and what the headroom costs."

Why this is L6:

  • Asks for rise time, not just peak
  • Chooses the signal based on workload type
  • Introduces floor/ceiling and shedding before being asked

What L7 adds:

  • "Which tier is this service? Tier-0 gets region-failover headroom. Internal services run hot."
  • Frames baseline capacity as a reserved-commitment purchase with finance
❌ Common L5 Trap

"I'll set up an auto-scaling group with a target of 70% CPU, min 10, max 100, and a 5-minute cooldown."

Why this misses: It's a valid configuration. But it doesn't know whether CPU measures this workload's demand, whether 5 minutes of lag is survivable, or whether 100 instances will overrun the database. The interviewer will extract all three with follow-ups.


Drill 2: Compute the Actuation Gap#

Prompt: "Traffic doubles in 60 seconds. Walk me through the timeline of your autoscaler."

Staff Answer

"t=0 spike. Metrics scraped at 15s granularity; the controller needs the target breached over a 60s window, so the decision lands at ~t+60–75s. The cloud API and scheduling take ~10s. VM boot 60–120s. App start plus image pull 20–40s. Warm-up (JIT, connection pools, cache) 60s. LB slow-start ramps over 60s. Useful capacity arrives around t+4–5 minutes. For those 4–5 minutes, the existing fleet absorbs 2× load. If we're at 55% utilization, 2× puts us at 110%, past the knee. So either we run at ≤ 40% (expensive), or the gateway sheds ~15–20% of low-priority traffic for 4 minutes, or we cut the delay: a warm pool brings it to ~90s, and node headroom for containers brings it to ~45s. I'd use a warm pool plus shedding, and accept a few minutes of degraded low-priority traffic for surprise doublings."

Why this is L6:

  • Adds up each delay explicitly
  • Converts the gap into a utilization requirement
  • Chooses a combination and names who pays (low-priority traffic)

What L7 adds:

  • Turns "surprise 2× in 60s" into a written survivability requirement per tier, with product sign-off on what gets shed
  • Tracks actuation delay as a platform SLO (e.g., p95 < 90s for warm-pool attach)

Drill 3: Pick the Signal#

Prompt: "Our service averages 25% CPU but p99 latency hits 2s at peak. The autoscaler never fires. What's wrong?"

Staff Answer

"The service is I/O-bound. It spends its time waiting on a DB or downstream, so threads or event loops saturate long before CPU does. CPU is measuring the wrong thing. I'd switch the signal to in-flight requests per instance. From a load test, find the concurrency where p99 hits the SLO. Say an instance meets SLO at 400 RPS × 150ms average = 60 in flight, so the target is 60. Before switching, I'd confirm the latency isn't caused by a slow dependency. If DB latency rose first, scaling out adds connections and makes it worse. So I'd add a flat-RPS guard and a ceiling from the DB's connection budget."

Why this is L6:

  • Diagnoses the I/O-bound mismatch
  • Uses Little's Law to derive a target instead of guessing
  • Anticipates the danger of the new signal

What L7 adds:

  • Publishes a per-workload-class signal standard so this isn't rediscovered team by team
  • Adds "capacity unit measured in the last 90 days" as a launch-readiness check

Drill 4: The Dependency Ceiling#

Prompt: "The app tier scaled from 50 to 250 instances during a sale and the database fell over. Design so this can't happen."

Staff Answer

"250 instances × 25 pool connections = 6,250 connections against a 5,000 limit, and Postgres degrades long before its limit because each connection is a process. Four changes. One, a connection pooler (PgBouncer, transaction mode) so app connections multiplex onto ~300 server connections. Two, max instances derived from the dependency budget: min(budget, pooler_client_limit ÷ pool_size, db_write_capacity ÷ writes_per_instance_at_target). Three, load shedding at the gateway once the fleet hits its ceiling, so excess demand is rejected at the edge instead of queuing in the DB. Four, move non-critical writes behind a queue so the DB sees smoothed write load. The rule I'd state: the elastic tier's maximum is a function of the least elastic thing it calls."

Why this is L6:

  • Quantifies the connection math
  • Derives the ceiling instead of picking a number
  • Moves the failure to the edge, where it's cheap

What L7 adds:

  • Per-consumer budgets on the shared DB, published by the DB platform team, enforced at the pooler
  • Adds the DB to every launch's load-test scope as a policy

Drill 5: Headroom Math#

Prompt: "Why do you run at 55% and not 80%? Finance wants 80."

Staff Answer

"80% gives us nothing to lose. Three AZs, and if one fails the other two must absorb 1.5× their load: 80% × 1.5 = 120%, an outage. And latency rises non-linearly: queueing delay goes as 1/(1−ρ), so at 80% it's ~5× the service time and at 90% ~10×. Our measured knee is ~80%. To survive AZ loss below the knee we need 80% ÷ 1.5 ≈ 53%, so 55%. On a $400K/month fleet, 55% vs 80% costs ~$125K/month. What finance is buying is 'an AZ failure is a non-event instead of a 1–2 hour customer-facing outage.' If an outage of that length costs more than ~$125K times its probability over a month, and for checkout it does, we pay. For internal batch services I'd happily run at 85%, since the queue absorbs everything."

Why this is L6:

  • Ties the number to a failure scenario and the latency knee
  • Prices it and frames the trade for finance
  • Differentiates by tier

What L7 adds:

  • Establishes utilization targets as an org policy by tier, reviewed annually with finance
  • Compares with reserved-instance/commitment discounts that lower the cost of headroom by 30–50%

Drill 6: Queue Workers#

Prompt: "Scale a fleet of workers consuming from a Kafka topic with 64 partitions."

Staff Answer

"Signal: consumer lag in time, not messages. Desired workers = ceil(lag_msgs ÷ (throughput_per_worker × target_drain_seconds)), plus workers for the steady-state arrival rate. The hard ceiling is 64 consumers, since Kafka assigns at most one consumer per partition in a group. Beyond that, extra workers sit idle. If we need more, we raise partition count (painful, changes key distribution) or make each consumer process concurrently within a partition while preserving per-key order. Scale-in carefully: every consumer group membership change triggers a rebalance that pauses consumption for seconds, so the stabilization window should be 10+ minutes. Poison messages inflate lag without real load, so I'd route to a DLQ after N retries and exclude DLQ'd work from the signal."

Why this is L6:

  • Uses the right signal and formula for async work
  • Knows the partition-count ceiling and rebalance cost
  • Guards against lag caused by poison messages

What L7 adds:

  • Partition count is a capacity decision made at topic creation. Set a standard (e.g., sized for 3-year peak ÷ per-consumer throughput)
  • Downstream protection: the workers' target system needs a budget too

Drill 7: Load Testing That Tells the Truth#

Prompt: "Your load test showed we could handle 3× peak. On launch day we fell over at 1.4×. Why?"

Staff Answer

"The usual suspects. One, data: the test ran against a small or warm dataset, so indexes fit in memory. Production doesn't. Two, cache hit rate: synthetic traffic with 1,000 repeated user IDs gives a 99% hit rate, while real traffic across 10M users gives 85%, which is 15× the DB reads. Three, request mix: the test hit the cheap endpoints, but launch traffic is 40% search. Four, dependencies stubbed: the payment provider was mocked. Five, the test client was the bottleneck, or ran from one AZ. Six, no retries: real clients retried under slowness and tripled load. Fixes: replay a sampled production traffic log, use production-scale data (anonymized), include real dependencies or their measured limits, verify the load generator isn't saturated, and model retries. Better still, measure capacity in production by shifting live traffic to a subset of hosts with automatic abort thresholds, the way Facebook's Kraken does."

Why this is L6:

  • Lists concrete, common causes, not "the test was wrong"
  • Knows cache hit rate is the silent multiplier
  • Proposes production-based measurement

What L7 adds:

  • Makes production capacity measurement a continuous platform service, not a pre-launch event
  • Load-test results feed a published capacity model per service that the capacity review reads

Drill 8: Build vs Buy#

Prompt: "Should we build our own autoscaler?"

Staff Answer

"No. The controllers in Kubernetes and the clouds are good, and the controller isn't where the risk is. What we should build are the pieces around it that are specific to us: the event calendar and pre-scale workflow, the capacity model per service (measured RPS per instance at SLO, limiting dependency), dependency budget enforcement, and the custom metrics adapter that feeds concurrency and lag into the standard controller. Buy the loop, build the inputs. The exception is at very large scale with owned hardware, where you do build capacity planning systems, because provisioning takes months and there's no cloud API to call."

Why this is L6:

  • Places the build effort where the differentiation is
  • Names concrete components

What L7 adds:

  • Evaluates Karpenter-style node provisioning vs. managed options by the cost of the platform team to run them
  • Sunsets bespoke per-team scaling scripts as part of standardization

Drill 9: Changing the Policy Without an Outage#

Prompt: "You want to move 150 services from CPU-based to concurrency-based scaling. How?"

Staff Answer

"Shadow first. Run the new signal in recommendation mode, computing desired capacity without acting, for 2 weeks per service. Compare against the actual. Where recommendations diverge by > 30%, investigate before switching. Then roll out by tier: internal services first, then consumer-facing, tier-0 last. During the switch, run both policies with max(cpu_desired, concurrency_desired) so the new signal can only add capacity, never remove it. After 2 clean weeks, drop CPU. Each service needs a measured concurrency target from a load test, and the platform provides a default load-test job to make that cheap."

Why this is L6:

  • Shadow → max-of-both → cut over, a safe migration pattern
  • Ordered by risk tier
  • Makes it cheap for teams to comply

What L7 adds:

  • Tracks migration as an org program with a completion metric and an exception process
  • Defines the success metric: fewer latency SLO breaches during ramps, and cost flat or down

Drill 10: Multi-Region#

Prompt: "We're going active-active in 3 regions. How does capacity change?"

Staff Answer

"The headroom target is now set by region evacuation, not AZ loss. With 3 regions each serving a third, losing one means the others each take +50%. Without failover headroom, each region must run at ≤ ~66% of its knee at peak. And the failover happens in minutes, faster than autoscaling, so the headroom must be standing. Alternatively, accept a degraded failover: shift traffic, shed non-critical features, and let autoscaling catch up over 10 minutes. That lets regions run at ~75% and costs a brief brownout on failover. Also: follow-the-sun peaks mean regional peaks don't coincide, so global capacity can be lower than 3× regional peak. Data tiers need the failover math too: a region's DB replica becomes primary and takes write load it's never seen."

Why this is L6:

  • Recomputes headroom from the new failure scenario
  • Offers a cheaper degraded alternative and names the tradeoff
  • Remembers the data tier

What L7 adds:

  • Prices full vs. degraded failover posture and takes it to leadership as a business choice
  • Mandates regular evacuation game days to verify the headroom actually exists

8. Deep Dive Scenarios#

Deep Dive 1: Peak-Traffic Incident — Launch Day Brownout#

Context: A highly promoted feature launches at 10:00. At 10:02, p99 latency is 6s and 12% of requests fail. The autoscaler shows desired capacity climbing from 80 to 240. On-call escalates to you.

Questions to Surface First:

  • Is the new capacity serving, or stuck in boot/warm-up/health-check failure?
  • Which dependency is at its limit: DB connections, third party, cache?
  • Is offered load inflated by retries?
  • Was this launch on the event calendar, and was it pre-scaled?

Typical L5 Approach: Raise max instances; increase instance size; wait for the autoscaler to catch up.

Staff Approach: Stop the churn: freeze instance replacement, since health-check failures are killing slow-but-alive instances. Enable gateway admission control to cap offered load at current real capacity, shedding low-priority traffic first. Check the DB connection count and cap max instances below the dependency ceiling. Once stable, let warm capacity join with slow-start. Then ask why the launch wasn't pre-scaled.

Principal Approach: Treats it as a planning-process failure. Institutes a launch-readiness gate: any launch with expected traffic > 1.5× baseline requires an event calendar entry, a load test at 2× forecast including dependencies, a pre-scale plan and a named approver. Adds launch readiness to the product review checklist, so marketing can't send the email without an infra sign-off.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Freeze replacement/scale-in. Enable shedding of low-priority routes. Confirm DB and third-party headroom.
TriageAre new instances healthy? Is readiness failing because of cold caches? What's the retry ratio?
Quick fixRaise the floor to the forecast × 1.3. Loosen readiness timeouts. Prime caches.
GuardrailsCap max instances at the DB budget. Watch db.active_connections and p99.
Post-mortemWhy no pre-scale? Why did health checks kill busy instances? Why were retries unbounded?

Metrics to Watch: lb.healthy_host_ratio, autoscaler.desired_vs_inservice, gateway.shed_rate{priority}, db.active_connections, client.retry_ratio, app.p99_latency

Organizational Follow-up: Launch-readiness gate; event calendar integrated with the marketing tool; liveness/readiness separation standard.

Ownership Question: "Who decides which traffic gets shed?" Staff answer: product decides priority tiers in advance (e.g., checkout > browse > recommendations). Engineering executes during the incident. Deciding priorities at 10:02 in an incident channel is too late.

Key Takeaway: "Known events are a process problem. If a launch needs the autoscaler to save it, the planning already failed."

What clears the Staff bar:

  • Stops the death spiral before adding capacity
  • Uses shedding with pre-agreed priorities
  • Traces root cause to the missing pre-scale process

Deep Dive 2: Silent Failure — Headroom That Wasn't There#

Context: A zone outage lasts 40 minutes. The service was "designed for AZ loss" at a 60% target, yet it browned out for the whole outage. Dashboards showed 60% average utilization before the event.

Questions to Surface First:

  • 60% of what? Averaged across the fleet, or at the hottest instances?
  • Did the remaining zones have instance quota and warm capacity?
  • Did traffic distribute evenly, or did one zone take the surge?
  • Was the capacity unit (RPS per instance at SLO) still accurate after recent releases?

Typical L5 Approach: Raise the target headroom to 50% and move on.

Staff Approach: Finds the failures in the assumption chain. The 60% was an average, but the hottest AZ ran at 72% because of uneven LB weights. A release 6 weeks earlier made requests 25% more expensive, so the real capacity unit had dropped from 400 to 300 RPS per instance and the actual utilization vs. SLO was ~80%. And the regional vCPU quota prevented scale-out to compensate. Fixes: measure utilization against the current capacity unit, re-measured after each major release; monitor per-AZ peak, not the average; keep quota headroom ≥ 1.5× peak; run an AZ-evacuation game day quarterly.

Principal Approach: Institutionalizes headroom verification: every tier-0/tier-1 service runs a quarterly AZ-drain exercise, removing a zone's capacity in production for 30 minutes. A service that fails loses its tier-0 designation until fixed. Headroom is no longer a number in a config. It's a tested property.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateShed low-priority traffic; request emergency quota; raise floor in surviving zones.
TriagePer-AZ utilization; capacity unit from recent load tests vs. production; quota status.
Quick fixCorrect LB weights; raise target instance counts to match the real capacity unit.
GuardrailsAlert on per-AZ utilization vs. current capacity unit; alert when quota headroom < 1.5× peak.
Post-mortemWhy did the capacity unit drift unnoticed? Why wasn't AZ loss tested?

Metrics to Watch: capacity.utilization_vs_unit{az}, capacity.unit_rps_measured (trend), cloud.quota_headroom_ratio, lb.weight_skew

Organizational Follow-up: Capacity-unit regression check in the release pipeline (a perf test comparing RPS per core with the previous release); quarterly AZ-drain game days.

Ownership Question: "Who owns the capacity unit number?" Staff answer: the service team. It's their code that changes it. The platform provides the test harness and flags regressions over 10%.

Key Takeaway: "Headroom you haven't tested is a hypothesis. Efficiency regressions quietly erase it between incidents."

What clears the Staff bar:

  • Questions the averaging and the capacity unit instead of just raising headroom
  • Finds the quota ceiling
  • Turns headroom into a tested property

Deep Dive 3: Large Customer Onboarding#

Context: A B2B platform signs a customer that will send 40% of current total traffic, ramping over 3 weeks starting next month. Their traffic is bursty: nightly batch syncs of 5M records in 20 minutes.

Questions to Surface First:

  • What's the peak shape: sustained, or 20-minute bursts at a fixed time?
  • Which endpoints, and what's their cost relative to the average request?
  • Which shared dependencies will they hit hardest?
  • Can we shape their traffic contractually (rate limits, batch windows, a bulk API)?

Typical L5 Approach: Increase max instances by 40%; scale the database vertically.

Staff Approach: Models the customer's load explicitly: 5M records ÷ 20 min ≈ 4,200 writes/s at 02:00 UTC, which is 3× the DB's current peak write rate. Rather than sizing everything for that burst, offers a bulk import API that goes through a queue drained at a controlled rate (e.g., 1,500/s over ~55 min), with per-tenant rate limits at the gateway. Pre-scales workers on a schedule for the batch window. Load-tests the bulk path at 2× their stated volume. Ramps their traffic weekly with checkpoints.

Principal Approach: Adds a commercial-capacity interface. Deals above a threshold (e.g., > 10% of traffic) trigger a capacity review before signing, and contracts include rate limits and batch windows. Capacity becomes part of the sales process, and the cost of serving a customer shows up in deal margin.

Staff Approach — Full Reasoning
PhaseWhat to Do
ModelRequests by endpoint × cost; burst shape; dependency impact (DB writes/s, cache footprint).
ShapeBulk API via queue; per-tenant limits; agreed batch window.
ProvisionScheduled worker floor for the window; DB capacity for the drained rate, not the raw burst.
ValidateLoad test at 2×; shadow their traffic in week 1 at 10%.
Ramp10% → 30% → 60% → 100% weekly with go/no-go checks.

Metrics to Watch: tenant.rps{tenant}, bulk_queue.lag_seconds, db.write_iops, gateway.tenant_throttled_total

Organizational Follow-up: Capacity review in the deal desk; tenant-level cost attribution.

Ownership Question: "Who says no if the customer wants unthrottled real-time sync?" Staff answer: the product owner of the API, with the capacity model as evidence. The alternative is pricing that covers dedicated capacity.

Key Takeaway: "Don't size for a customer's burst. Shape the burst, then size for the shaped load."

What clears the Staff bar:

  • Quantifies the burst against current capacity
  • Changes the demand curve rather than only the supply
  • Stages the ramp with checkpoints

Deep Dive 4: Post-Mortem — The Retry Storm That Scaled to the Ceiling#

Context: A 90-second DB failover caused errors. After the DB recovered, the service stayed down for 25 more minutes. The autoscaler had scaled to its max of 400 instances, and offered load was 4× normal.

Questions to Surface First:

  • What was the retry policy in clients and in intermediate services?
  • Did retries stack across layers (client × gateway × service)?
  • Why didn't load fall once the DB recovered?

Typical L5 Approach: Raise the max instances; make the DB failover faster.

Staff Approach: Identifies multiplicative retries: mobile clients retried 3×, the gateway retried 2×, and the service retried its DB calls 3×, so up to 18× amplification on failing paths. After recovery, the backlog of retries plus normal load kept the DB saturated (a metastable failure). Fixes: retry only at one layer, with retry budgets (≤ 10% extra), exponential backoff with full jitter, and circuit breakers that stop retries during failure. Add admission control keyed on DB health so the fleet sheds load quickly and lets the DB recover.

Principal Approach: Publishes an org-wide retry standard: retries at the outermost layer only, budgets enforced by the service mesh, and a required "metastability test" for tier-0 services (inject a 2-minute dependency outage and verify recovery within 2 minutes after it ends).

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateShed 50% at the gateway to break the storm; the DB recovers; ramp admission back up.
TriageRetry counts per layer; amplification factor.
Quick fixDisable gateway retries; cut service retries to 1 with jitter.
GuardrailsRetry budget metrics; circuit breakers on DB calls.
Post-mortemWhy did autoscaling scale on retry-inflated load? Add a flat-organic-RPS guard (organic = first attempts only).

Metrics to Watch: requests.attempt_number histogram, retry.budget_exhausted_total, db.cpu, gateway.shed_rate

Organizational Follow-up: Retry standard in the service mesh; metastability test in game days. See Circuit Breakers.

Ownership Question: "Who owns retry policy?" Staff answer: the platform, via the mesh, as a default. Teams can lower it, not raise it, without review.

Key Takeaway: "Autoscaling can't outrun a retry storm. It scales the attack. Retry budgets and shedding end it."

What clears the Staff bar:

  • Computes multiplicative retry amplification
  • Names metastable failure
  • Moves retry policy into a platform standard

Deep Dive 5: Multi-Region Expansion#

Context: The company is adding two regions to its single US-East deployment for latency and resilience. Leadership asks what it will cost and how capacity should be planned.

Questions to Surface First:

  • Is the goal latency (serve users nearby) or resilience (survive region loss), or both?
  • Must a region failure be invisible (full headroom), or is a degraded failover acceptable?
  • Do regional peaks coincide or follow the sun?
  • Where does the data live, and can each region's data tier take failover writes?

Typical L5 Approach: Replicate the current fleet 3×. Cost triples.

Staff Approach: Models regional demand curves. With follow-the-sun peaks, the global peak is ~1.4× the largest region's peak rather than 3×. For full region-loss tolerance, each region needs capacity for its own peak plus its share of the failed region's load at that hour. Computes the two postures: (a) invisible failover, where each region runs at ~55% at local peak (~1.8× current total spend), and (b) degraded failover with shedding and a 10-minute autoscale catch-up, where regions run at ~70% (~1.4× spend). Recommends (b) for most services and (a) for tier-0.

Principal Approach: Takes the decision to leadership as a business choice with prices: "Invisible regional failover costs $X/year more than degraded failover. Degraded failover means ~10 minutes of reduced functionality once or twice a year." Pairs it with a reservation strategy: commit ~60–70% of baseline for 1–3 years at a 30–50% discount, and keep the rest on-demand for elasticity.

Staff Approach — Full Reasoning
PhaseWhat to Do
ModelHourly demand per region; overlap; failover redistribution matrix.
DecideFailover posture per service tier.
ProvisionPer-region floors; quota ≥ 1.5× regional peak; data tier sized for failover writes.
ValidateRegion evacuation game day before GA, then quarterly.
OperatePer-region capacity dashboards; forecast by region.

Metrics to Watch: capacity.headroom_for_evacuation{region}, region.peak_overlap, db.replica_write_capacity_ratio

Organizational Follow-up: Tier-based failover policy; finance-approved reservation plan.

Ownership Question: "Who decides invisible vs. degraded failover?" Staff answer: the VP-level owner of the product, informed by an engineering cost model. It's a business risk decision.

Key Takeaway: "Multi-region capacity is set by the evacuation scenario, and follow-the-sun traffic makes it much cheaper than N× the current fleet."

What clears the Staff bar:

  • Uses non-coincident peaks to cut cost
  • Offers two priced postures
  • Includes the data tier and quotas

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Compare demand rise time with the actuation delay, and pick schedule, predict, react, headroom or shed accordingly
  • Choose a scaling signal per workload class and derive its target with Little's Law
  • Compute a ceiling from dependency budgets (DB connections, third-party quotas, partitions)
  • Set a utilization target from a failure scenario and the latency knee, and price it
  • Describe the death spiral, scaling into a slow dependency, flapping, hidden ceilings and retry storms, each with a metric and an owner
  • Design a load test that tells the truth, and explain why production traffic shifting is better
  • Describe the org model: platform-owned mechanism, service-owned capacity models, event calendar, quarterly review

The Bar for This Question#

Mid-level (L4): Configures an autoscaling group with CPU targets, min/max and cooldowns. Knows horizontal vs. vertical. Treats the autoscaler as the solution.

Senior (L5): Adds scheduled scaling for known events, read replicas for the DB and a load test. The design is reasonable and handles the diurnal curve. It doesn't quantify the actuation gap, uses CPU without questioning it, and doesn't derive a ceiling.

Staff+ (L6): Frames capacity as pre-paid risk. Compares time constants, picks a demand signal, clamps the controller between a forecast floor and a dependency ceiling, covers the gap with warm pools and shedding, ties headroom to a failure scenario with a price, and puts the owner of every piece in the answer. The interviewer should learn something from the answer, whether that's the flat-RPS guard, the capacity-unit drift after a release, or the retry amplification math.


10. Staff Insiders: Controversial Opinions#

10.1 "CPU-Based Autoscaling Is Wrong for Most Services"#

EvidenceImplication
Most web services are I/O-bound: they wait on databases, caches and other servicesCPU stays low while threads saturate
CPU rises after latency doesLate signal on a system that's already late
GC, noisy neighbors and CPU throttling distort itNoise that triggers flapping

The Staff position: concurrency or backlog by default. CPU for proven CPU-bound work.

Why this matters in interviews: saying "CPU at 70%" without justification is the most common L5 tell in this topic.

10.2 "Autoscaling Doesn't Handle Spikes — Headroom Does"#

EvidenceImplication
Actuation delay is 3–5 minutes on VMsSpikes shorter than that are absorbed by what's already running
Launch traffic rises in seconds to minutesReactive scaling helps only after the damage
Autoscaling's real benefit is scaling in at nightIt's a cost tool more than a reliability tool

The Staff position: autoscaling saves money on slow curves. Headroom, pre-scaling and shedding handle spikes.

Why this matters in interviews: reframing autoscaling as a cost tool shows you understand the time constants.

10.3 "Load Tests Lie; Production Tells the Truth"#

EvidenceImplication
Synthetic tests miss cache hit rates, data size, request mix and retriesCapacity is overstated by 1.5–3×
Facebook's Kraken measures capacity with live trafficIndustry precedent for production measurement
Capacity drifts with every releaseA one-time test goes stale in weeks

The Staff position: continuous production capacity measurement with automatic abort, plus pre-launch synthetic tests at 2× forecast.

Why this matters in interviews: it shows you've been burned by a passing load test.

10.4 "Scale-to-Zero Is a Dev Environment Feature"#

EvidenceImplication
Cold starts add 100ms–several seconds to the first requestsUser-facing tail latency suffers
Provisioned concurrency to hide cold starts reintroduces idle costThe savings shrink
The savings matter most where traffic is low, which is where cost matters leastOptimizes the wrong thing

The Staff position: scale to zero for dev/test, internal tools and truly sporadic jobs. Keep a warm minimum for user-facing paths.

Why this matters in interviews: it shows you weigh the latency cost of a cost optimization.

10.5 "Most Capacity Outages Are Calendar Failures"#

EvidenceImplication
Launches, sales and campaigns are scheduled weeks aheadThey're predictable in principle
The outage happens when infra wasn't toldThe failure is communication
Fixing the autoscaler doesn't fix the calendarThe investment belongs in process

The Staff position: an event calendar with an owner and a pre-scale approval step is worth more than any controller tuning.

Why this matters in interviews: ending on process instead of config is a clear Staff-level signal.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

A Staff engineer makes one service scale correctly. A Principal engineer sees capacity as the company's largest controllable operating expense and its most correlated reliability risk at once. Every service's headroom, commitment level and scaling policy adds up to a portfolio: a baseline bought with 1–3 year commitments, an elastic layer bought on demand, and a headroom layer that is insurance. The L7 job is to set the rules for that portfolio (utilization targets by tier, commitment coverage, shared-dependency budgets, failure scenarios the company pays to survive) and to make those costs visible so leadership trades them deliberately.

The Org-Level Fault Line#

Central capacity authority vs. team autonomy over spend. Teams that own their budgets move fast but over-provision (fear of pages) or under-provision (pressure on cost). They also ignore shared dependencies. A central capacity group sees the whole picture but becomes a ticket queue and loses workload knowledge.

The Principal position: federated with guardrails. Teams own their capacity models and scaling policies within org-wide targets by tier. A small central capacity team (3–6 people at a 1,000-engineer company) owns forecasting, commitments, shared-dependency budgets, the quarterly review and showback. Exceptions to targets are allowed, but they're visible and priced.

🧭 Principal Move: "I don't need every team to be efficient. I need every team's inefficiency to be visible, intentional and attached to a reason. Showback does most of the work; the review does the rest."

Cost Model#

Assumptions: cloud compute ~$0.04–0.05 per vCPU-hour on-demand; commitments save ~30–50%; engineer fully loaded ~$300K/year; utilization figures are time-weighted averages.

ScaleCompute spendTypical waste without disciplineAchievable with Staff designHeadcount for capacity/platformOn-call load
Startup (20 services)~$50K/month50–60% idle (fixed peak sizing)~35% idle; ~$15K/month saved0.5 FTE inside the platform teamCapacity pages ~2/month
Growth (200 services)~$2M/month45–55% idle; 20% on-demand premium35–40% idle, 60–70% commitment coverage: ~$500–700K/month saved3–5 FTE capacity + autoscaling platform1 platform rotation
Hyperscale (2,000 services)~$40M/monthHeadroom duplicated across tiers and regionsTiered targets + follow-the-sun + commitments: 20–30% saved (~$8–12M/month)15–30 FTE: forecasting, capacity eng, efficiencyDedicated capacity on-call during events

Reading it: past growth stage, the capacity team pays for itself 10–50× over. The biggest single lever is usually commitment coverage on the baseline, not autoscaling.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversibility Cost
3-year compute commitmentsOne-wayPaid whether used or not. Over-commit by 20% and that's pure waste for 3 years
Kafka partition counts / DB shard countsOne-way-ishRe-partitioning changes key ordering and needs migration
Region count and failover postureOne-way for ~1–2 yearsData residency, replication topology and contracts
Scaling signal per serviceTwo-wayShadow-mode switch in weeks
Utilization targets by tierTwo-wayPolicy change, then gradual rebalancing
Autoscaler vendor/controllerTwo-wayReplace the loop, keep the inputs
Instance families / hardware generationTwo-way on cloud, one-way on owned hardwareOwned hardware depreciates over 4–5 years

The Standard I'd Write#

RFC: Capacity & Scaling Standard v1

Scope: All production services on shared compute platforms.

MUST

  1. Every service publishes a capacity model: RPS (or throughput) per instance at SLO, limiting dependency, max safe instances. Re-measured every 90 days and after any release flagged by the perf regression check.
  2. Scaling policies use a demand signal appropriate to the workload class (request/response: concurrency or RPS; async: backlog ÷ drain rate). CPU only with a documented justification.
  3. Max instances ≤ the ceiling derived from dependency budgets.
  4. Scale-in stabilization ≥ 5 minutes; scale-in step ≤ 20% per window.
  5. Any planned event > 1.5× baseline has an event calendar entry ≥ 5 business days ahead, a load test at 2× forecast including dependencies, and a named approver.
  6. Tier-0/1 services pass an AZ-drain exercise each quarter.

SHOULD

  • Warm pools or node headroom for services with actuation delay > 2 minutes.
  • Retries only at the outermost layer, with budgets ≤ 10%.

Utilization targets (peak, vs. measured knee): Tier-0: ≤ 50%. Tier-1: ≤ 60%. Tier-2: ≤ 70%. Async/batch: ≤ 85%.

Exceptions: Requested through the capacity review; approved by the capacity lead and the service's director; expire after 2 quarters.

Success metrics: zero capacity-caused SEV-1s at planned events; org-wide peak utilization within ±10% of tier targets; commitment coverage 60–75% of baseline; forecast error < 15% at the weekly level.

What I'd Tell the VP#

"We spend about $2M a month on compute, and roughly half of it sits idle at any moment. Some of that is insurance we need, so a data center failure doesn't take us down. A lot of it is just unmanaged. I'm proposing three things: each team sees its own idle cost monthly, we set utilization targets by how critical a service is, and we commit to about two-thirds of our baseline usage for discounts. That should save $500–700K a month within two quarters. Separately, our last two launch-day outages happened because infrastructure wasn't told about the launch. A simple event calendar with a sign-off step fixes that for almost no cost."

Principal Interview Signals#

SignalWhat It Sounds Like
Portfolio thinking"Baseline on commitments, the variable layer on-demand, headroom as priced insurance by tier."
Prices the decision"Invisible region failover costs ~$X/year more than degraded failover. Leadership picks."
Makes waste visible"Showback per team, with idle cost next to the scenario it buys."
Governs shared dependencies"Per-consumer budgets on the shared DB, enforced at the pooler, reviewed quarterly."
Tests the posture"Headroom you haven't drained a zone to prove is a hypothesis."

Staff answers that L7 interviewers find insufficient:

  • "We'll set utilization at 60%." Correct for one service, but no tiering, no pricing and no governance across 200 services.
  • "We'll load test before big launches." This relies on heroics per event instead of continuous measurement and a readiness gate.
  • "Autoscaling will reduce cost." This ignores that commitments on the baseline usually save more than autoscaling does.

Appendices

Appendix A: Scaling Mechanics in Depth#

A.1 Target Tracking#

every 15s:
    observed = avg(metric over last 60s)           # e.g., in-flight per instance
    desired  = ceil(current_ready × observed / target)
    desired  = max(desired, floor(now))            # forecast / schedule
    desired  = min(desired, ceiling)               # dependency budget
    if desired > current: scale_out(min(desired, current × 1.5))   # step cap
    if desired < current and stable_for(600s):    scale_in(max(desired, current × 0.9))

The asymmetric step caps and stabilization window are the whole safety story. Kubernetes HPA's defaults (15s loop, 300s scale-down window) encode the same idea.

A.2 Little's Law Target#

in_flight_target = rps_per_instance_at_slo × avg_latency_seconds
e.g. 400 RPS × 0.150s = 60 concurrent requests per instance

When latency rises for external reasons, in-flight rises without real demand growth. Hence the flat-RPS guard:

if concurrency_breach and organic_rps_growth_5m < 20%:
    suppress scale_out; alert "dependency slowdown suspected"

A.3 Queue Worker Sizing#

arrival_workers = arrival_rate / throughput_per_worker
drain_workers   = backlog / (throughput_per_worker × target_drain_s)
desired         = ceil(arrival_workers + drain_workers)
desired         = min(desired, partition_count, downstream_budget / per_worker_rate)

A.4 Predictive Floor#

floor(t) = seasonal_baseline(t, weeks=4)          # same hour-of-week, median
         × growth_trend                            # e.g., 1.02 week-over-week
         × event_multiplier(t)                     # from calendar
         × safety_factor                           # 1.1–1.3
         ÷ rps_per_instance_at_slo × (1 / target_util)

The floor may only raise capacity. Reactive scaling always operates above it.

Appendix B: Capacity Model Construction#

B.1 Finding the Knee#

Diagram: B.1 Finding the Knee

B.2 What Goes in the Model#

FieldExampleSource
rps_per_instance_at_slo400Load test at the knee × 0.9
instance_shape4 vCPU / 8 GBDeployment config
limiting_dependencyorders-db connectionsLoad test / architecture
ceiling_instances220Dependency budget ÷ per-instance use
actuation_delay_p9595s (warm pool)Platform telemetry
tier1Service catalog
last_measured2026-08-14Must be < 90 days

B.3 Dependency Ceiling Worksheet#

DependencyBudgetPer-instance useCeiling
Postgres via PgBouncer4,000 client conns15266
Payments API2,000 RPS8 RPS250
Redis cluster400K ops/s1,800 ops/s222
Regional vCPU quota2,000 vCPU4500
Effective ceiling222 (Redis)

Appendix C: Load Shedding & Admission Control#

MechanismWhereBehavior
Concurrency limit per instanceServiceReject (503) when in-flight > 1.2× target
Priority tiersGatewayShed tier-3 at 90% fleet concurrency, tier-2 at 95%, never tier-0
Adaptive (latency-gradient) limitsService/meshLower concurrency limit when latency rises (AIMD-style)
Per-tenant quotasGatewayIsolate noisy tenants (Rate Limiting)
Queue with deadlineAsync pathDrop work older than its usefulness
Diagram: Appendix C: Load Shedding & Admission Control

Appendix D: Client Behavior#

  • Retries: at the outermost layer only. Exponential backoff (base 100ms, cap 20s) with full jitter. Retry budget ≤ 10% of requests.
  • Honor Retry-After on 429/503.
  • Deadlines propagate. Don't do work whose caller has already given up.
  • Idempotency keys on retried writes.

Appendix E: Observability#

E.1 Core Metrics#

capacity.utilization_vs_unit{service,az}     # demand ÷ (instances × capacity unit)
capacity.headroom_pct{service}               # 1 − utilization at peak
autoscaler.desired / .inservice / .pending   # gap = actuation in progress
autoscaler.scale_out_failures{reason}        # quota, capacity, image pull
autoscaler.direction_changes_per_hour        # flapping
warm_pool.available
gateway.shed_rate{priority}
forecast.error_pct{service}
client.retry_ratio                            # attempts ÷ organic requests
cost.idle_dollars{team}                       # showback

E.2 Critical Alerts#

AlertThresholdSeverity
Headroom at peak < scenario requirement< 33% for a 3-AZ tier-1 serviceTicket (weekly)
autoscaler.scale_out_failures> 0 for 5 minPage
Desired − in-service gap> 20% for 10 minPage
gateway.shed_rate{priority<=1}> 0.1%Page
client.retry_ratio> 1.3 for 5 minPage
Quota headroom< 1.5× peakTicket
Capacity unit stalelast_measured > 90 daysTicket

Appendix F: Scale Evolution#

StageWhat worksWhat you don't build yet
< 20 servicesHPA/ASG with correct signals; manual pre-scale; one load-test scriptPredictive floors, showback, capacity team
20–200 servicesEvent calendar; dependency ceilings; warm pools; showback; commitmentsContinuous prod capacity measurement
200+ servicesTiered targets; quarterly review; predictive floors; AZ-drain game days—
Multi-region / hyperscaleEvacuation posture per tier; follow-the-sun planning; Kraken-style measurement—

What you don't build on day one: a custom autoscaler, a forecasting ML model, a central capacity team. Do build on day one: the right signal, a ceiling, and a pre-scale habit for launches.

Appendix G: Multi-Tenancy, Fairness & Cost#

  • Tenant-driven demand: per-tenant quotas keep one tenant from driving everyone's scale-out. The capacity bill follows the tenant through cost attribution.
  • Shared-dependency budgets: each consuming service gets a connection/RPS budget from the dependency owner, enforced at the pooler or mesh.
  • Showback: monthly per-team report of spend, peak utilization, idle dollars, and the scenario headroom buys.
  • Spot/preemptible: fine for stateless batch and fault-tolerant async workers (typically 60–90% cheaper). Never for tier-0 synchronous paths without on-demand fallback.
  1. Loading the index…