Technologies referenced in this case study: Kafka · Redis · PostgreSQL · DynamoDB · Time Series DBs · API Gateways
Related case studies: Load Balancer · Circuit Breakers · Rate Limiting · Flash Sales · Metrics & Monitoring · Degraded Mode · Back-of-Envelope Estimation
How to Use This Case Study#
Organized for interview use first, reference second. Read front-to-back once. Return to individual sections for targeted review.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → one Deep Dive |
| Deep Dive | 3+ hrs | Everything, including the Principal Lens and appendices |
What is Auto-Scaling & Capacity Planning? — Why interviewers pick this topic
Capacity planning decides how much infrastructure you own and when. Auto-scaling decides how fast that amount changes in response to demand. Together they answer one question: when demand arrives, will the capacity already be there, and who paid for it while it was idle?
The prompt arrives in several forms: "Design an auto-scaler," "How would you prepare for Black Friday?", "Your service falls over every Monday at 9am," or as a follow-up inside any other design ("what happens at 10× traffic?").
Before vs After — the product-launch scenario:
Reactive-only autoscaling, CPU target 70%, no pre-warming:
t=0: Launch email to 40M users. RPS climbs from 20K to 140K in 90 seconds.
t=+30s: CPU crosses 70%. Autoscaler (60s evaluation window) hasn't fired yet.
t=+75s: Scale-out requested: +80 instances.
t=+90s: p99 latency 4s. Clients time out and retry, so effective RPS is 210K.
t=+3min: New instances boot. JVM cold, caches cold. They take full traffic and fail health checks.
t=+4min: LB ejects "unhealthy" new instances; autoscaler replaces them. Death spiral.
t=+6min: New app instances open 6,400 DB connections. Postgres max_connections 5,000 → DB refuses.
t=+25min: Manual intervention: shed load, cap scale-out, raise pool limits. Launch is a headline.
Staff design (scheduled pre-scale + reactive on concurrency + protected dependencies):
t=-60min: Scheduled action raises min capacity to forecast peak × 1.3 (load-tested last week).
t=-30min: Warm pools filled; caches primed by synthetic traffic.
t=0: Same 140K RPS. Utilization peaks at 58%. p99 stays at 180ms.
t=+10min: Actual traffic exceeds forecast by 20%. Reactive policy on in-flight requests adds 15%.
t=+12min: Admission control at the gateway sheds 2% of low-priority traffic during the step.
t=+3h: Scale-in, gradually, over 45 minutes. Cost of pre-scale: ~$4K. Cost of the alternative: the launch.
Why interviewers reach for this question: it separates engineers who think about capacity as a control loop with delays from those who think of it as "set a CPU threshold." It also exposes whether you see the dependencies that don't scale — databases, third-party APIs, license limits — and whether you can talk about money.
Mechanics Refresher: Scaling Mechanisms
| Mechanism | How It Works | Pros | Cons |
|---|---|---|---|
| Threshold / step scaling | If metric > X for N minutes, add K instances | Simple, predictable | Oscillates; step sizes are guesses |
| Target tracking | Controller keeps metric near target (e.g., CPU 60%), computes desired = current × metric / target | Self-sizing, few knobs | Only as good as the metric; lags on spikes |
| Scheduled scaling | Min/max capacity changed on a calendar | Zero lag for known events | Useless for surprises; stale schedules waste money |
| Predictive scaling | Forecast (seasonality model) sets capacity ahead of demand | Beats boot time on diurnal curves | Wrong on anomalies; needs weeks of history |
| Vertical scaling (VPA) | Change CPU/memory per instance | Helps stateful or single-threaded work | Often requires restart; hard ceiling |
| Queue-driven (KEDA-style) | Workers scale on backlog / (throughput per worker × target drain time) | Right signal for async work | Needs accurate per-worker throughput |
| Scale-to-zero | No instances when idle; start on first request | Lowest cost | Cold start on the first request, 100ms–10s+ |
For most production systems: target tracking on a demand signal (in-flight requests or RPS per instance), with a scheduled or predictive floor for known peaks. CPU is the default everyone picks and the wrong signal for I/O-bound services.
Executive Summary
If you only read one section, read this. Everything else in the case study elaborates on the contrast below.
What This Interview Actually Tests#
Auto-scaling is not a controller question. The formula desired = current × (metric / target) fits on one line.
This is a risk pre-payment question that tests:
- Whether you understand that a control loop with a 3–5 minute actuation delay cannot react to a 90-second spike, and that someone must pay for capacity before demand arrives
- Whether you pick a scaling signal that measures demand, not a side effect of demand
- Whether you find the components that can't scale (databases, connection limits, quotas, third parties) and protect them from the parts that can
- Whether you can price headroom in dollars and name who pays for idle capacity versus who pays for the outage
The key insight: every capacity decision trades idle dollars for incident risk. Autoscaling narrows the gap between them without closing it. Staff engineers state the headroom they are buying, the scenario it covers, and who signed off on the scenarios it doesn't.
The L5 → L6 → L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Set up an ASG / HPA on CPU at 70%" | Asks: cost efficiency, known peaks, or survival of unknown spikes and failures? What's the spike shape vs. our boot time? | Asks what the capacity budget is, who owns it, and how capacity is planned across 200 services that share databases and quotas |
| Signal | CPU utilization | Demand signal: in-flight requests / RPS per instance, queue backlog ÷ drain rate; CPU only for CPU-bound work | Standardizes signals per workload class across the org; makes the "capacity unit" (RPS per core at SLO) a published number per service |
| Timing | Reactive scaling | Reactive + scheduled/predictive floor; computes spike rise time vs. provisioning time; warm pools | Plans quarterly and annual capacity with finance: reservations, commitments, lead times for hardware and quota increases |
| Dependencies | "The DB will scale too" | Identifies non-elastic dependencies; caps scale-out by the DB's connection budget; adds pooling, admission control, load shedding | Owns the shared-dependency capacity model: which teams draw on which databases, and who arbitrates contention |
| Headroom | "Add some buffer" | Headroom as a number: N+1 zone, region-failover factor, burst absorption; load-tested to 1.5–2× forecast | Prices headroom per tier in $/month vs. expected outage cost; sets org-wide utilization targets by criticality tier |
| Ownership | Platform team runs the autoscaler | Service owners own their scaling policy and capacity model; platform owns the mechanism; a named approver for event pre-scales | Designs the operating model: capacity reviews, chargeback/showback, error budgets tied to headroom decisions |
Why "timing" separates levels
L5: Trusts the autoscaler to react. That works for diurnal curves that rise over hours. It fails for anything that rises faster than the provisioning delay: metric collection (15–60s) + evaluation window (60–180s) + instance boot (60–300s) + warm-up (30–180s). That's 3–10 minutes from spike to useful capacity.
L6: Compares the two time constants out loud. "If the spike rises in 90 seconds and our capacity arrives in 5 minutes, reactive scaling covers nothing. The first 5 minutes must be absorbed by headroom, load shedding, or capacity scheduled in advance. For known events I schedule. For unknown ones I keep headroom and shed."
L7: Extends the same reasoning to longer time constants: quota increases (days), reserved-capacity purchases (a 1–3 year commitment), and datacenter/hardware lead times (months). Capacity planning is the same control-loop problem at a slower frequency.
Why "dependencies" separates levels
L5: Scales the stateless tier and assumes everything behind it keeps up.
L6: Knows that scaling the stateless tier is often what causes the outage. 200 app instances × 30 pool connections = 6,000 DB connections against a 5,000 limit. Caps max instances by the downstream budget, puts a connection pooler in front of the DB, and prefers shedding load at the edge over stampeding the DB.
L7: Recognizes that across an org, 40 services autoscaling independently against one shared database is a tragedy of the commons. Introduces per-consumer quotas on shared dependencies and a capacity review before any service raises its max.
Why "ownership" separates levels
L5: "The platform team manages autoscaling."
L6: Platform owns the mechanism (controllers, node provisioning, warm pools). Service owners own their capacity model: RPS per instance at SLO, max safe instances, dependencies' limits, and the event calendar. Someone specific approves pre-scaling for launches.
L7: Adds money and governance. Showback per team, utilization targets per tier, and a quarterly capacity review where teams defend their headroom. Idle capacity is visible on a dashboard with a team name on it.
The Staff Positions#
| Position | Rationale |
|---|---|
| Scale on demand, not on symptoms | In-flight requests or RPS per instance lead; CPU lags and misleads for I/O-bound services |
| Headroom covers the actuation gap | Whatever arrives faster than capacity can be provisioned must be absorbed by pre-paid headroom or shed |
| Schedule the knowns, react to the unknowns | Launches, sales and Monday mornings are on a calendar. Pre-scale them. Reserve reactive scaling for surprises |
| Scale out fast, scale in slowly | Adding capacity late causes an outage. Removing it late costs a few dollars. Asymmetric cooldowns (e.g., 1 min out, 10–15 min in) |
| Cap scale-out by the weakest dependency | Max instances = min(budget, DB connections / pool size, downstream quota / per-instance rate) |
| Load shedding is part of capacity | When demand exceeds capacity, the choice is which requests fail, not whether some fail |
| Load test to 1.5–2× forecast peak | Forecasts are wrong by 20–50% on launches. Find the knee before customers do |
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Cost efficiency | Minimize idle spend on variable load | Aggressive target tracking (65–75% util), scale-in, spot/preemptible, scale-to-zero for dev | Under-provisioned at spikes; latency SLO breaches during ramp | SLO met ≥ 99% of minutes; idle ≤ 25% |
| Reliability through known peaks | Launches, sales, sports events, Monday 9am | Forecast + scheduled pre-scale + load test + warm pools + admission control | Forecast wrong; dependency ceiling hit | Zero SLO breach during planned events |
| Survival of unknown spikes and failures | Viral traffic, retry storms, zone/region loss | Standing headroom (N+1 zone, region-failover factor), load shedding, priority tiers | Paying for idle; still overwhelmed by > headroom spikes | Graceful degradation: priority traffic served, rest shed |
🎯 Staff Move: "I'll design for reliability through known peaks with a survival floor. That means a user-facing stateless tier with a strong diurnal curve plus planned launches, and enough standing headroom to lose one availability zone without paging anyone. Pure cost efficiency is a tuning exercise on top of that design, and the reverse doesn't hold."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Scaling Signal | CPU (universal, lagging, misleading for I/O) vs. demand signals (RPS, concurrency, backlog: accurate, service-specific) vs. latency (what users feel, but noisy and ambiguous) |
| 2 | Reactive vs Predictive | Reactive is always right about the present and always late. Predictive is early and sometimes wrong |
| 3 | Cold Start vs Warm Capacity | Scale-from-cold is cheap but slow (minutes). Warm pools and provisioned concurrency are fast but cost money while idle |
| 4 | Headroom vs Cost | Every 10% of headroom is 10% of the compute bill. Every 10% less is a bigger spike you can't absorb |
| 5 | Central Capacity vs Service Ownership | A capacity team with global view and budget authority vs. service teams who know their workloads and move fast |
In the Wild: Real Production Systems#
Why this section belongs here: citing well-known production practice shows you've studied operational reality. Use these as one-sentence anchors in the interview.
Netflix — Predictive Scaling and Regional Evacuation#
Netflix publicly described Scryer, a predictive autoscaling engine that forecasts demand from historical patterns and provisions capacity ahead of the diurnal ramp, because reactive scaling lagged their steep evening rise. Separately, Netflix's regional failover exercises (Chaos Kong) evacuate a whole AWS region and shift its traffic to the others, which requires each remaining region to absorb a large traffic increase within minutes.
Staff insight: two ideas worth citing. Predictive scaling exists because boot time exceeds ramp time. Region evacuation means your headroom target is set by the failover scenario, not by normal peak. Say both.
Facebook — Kraken: Load Testing With Live Traffic#
Facebook's Kraken (published at OSDI 2016) shifts live production traffic toward a subset of servers, clusters or regions to measure their true capacity limits, guided by health metrics that stop the test before users are hurt.
Staff insight: synthetic load tests miss real request mixes, cache hit rates and dependency behavior. Kraken's lesson is that the most accurate capacity number comes from production, under controlled conditions. Mention it when asked "how do you know your capacity?"
Kubernetes & AWS — Controllers With Built-In Asymmetry#
The Kubernetes Horizontal Pod Autoscaler evaluates metrics every 15 seconds by default and applies a default 5-minute scale-down stabilization window, so it scales out quickly and in cautiously. AWS Auto Scaling offers target tracking, scheduled actions, warm pools and a predictive scaling mode driven by historical load. Karpenter and the Cluster Autoscaler add nodes when pods can't be scheduled.
Staff insight: the defaults encode the Staff position. Scale out fast because under-provisioning is expensive. Scale in slowly because flapping is expensive and idle is cheap. Being able to quote these defaults shows you've operated these systems.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Autoscale on CPU at 70%" | "The service is I/O-bound and CPU sits at 25% while latency is 3s. What now?" | Whether you pick signals that measure demand |
| "The autoscaler handles spikes" | "Traffic triples in 60 seconds. How long until new capacity serves traffic?" | Whether you know the actuation delay |
| "We'll scale the app tier" | "What happens to the database when you go from 50 to 300 instances?" | Non-elastic dependency awareness |
| "Keep 30% headroom" | "Why 30? What does it cost and what does it cover?" | Can you price and justify headroom? |
| "Predictive scaling" | "The forecast was wrong by 2×. What happens?" | Fallbacks and guardrails on predictions |
| "Load test before launch" | "Your load test passed, but production fell over. Why?" | Test realism: data, cache state, dependencies, request mix |
System Architecture Overview#
Reading the diagram: the controller has three inputs. It measures demand (reactive), takes a floor from forecasts and the event calendar (predictive), and takes a ceiling from the capacity model's dependency budgets. The gateway's admission control is the safety valve for the gap between demand and capacity. The warm pool shortens the actuation delay from minutes to ~30 seconds. The database sits outside the elastic tier on purpose: it scales in hours, and the design protects it rather than pretending it's elastic.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Signal | "CPU at 70%" | "In-flight requests per instance for request/response work, backlog ÷ drain rate for queues. CPU only when the service is CPU-bound." |
| Timing | "Autoscaler reacts" | "Capacity arrives 3–5 minutes after the spike. Headroom and shedding cover that gap. Known events are pre-scaled." |
| Cold start | "Instances boot quickly" | "Boot plus warm-up is 2–5 min on VMs and 10–60s on warm nodes. Warm pools, slow-start at the LB, pre-warmed caches." |
| Headroom | "Add a buffer" | "Target 60% utilization. That covers losing 1 of 3 zones (×1.5) with a small burst margin. It costs ~40% idle, priced at $X/month." |
| Dependencies | "DB scales too" | "Max instances is capped by DB connections ÷ pool size. A pooler, load shedding and per-consumer quotas protect it." |
| Validation | "Load test once" | "Load test to 1.5–2× forecast with production-like data. Find the knee. Rerun after every major change." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| HPA metric sync period | 15s (default) | Reactive loop's fastest reaction |
| HPA scale-down stabilization | 300s (default) | Built-in asymmetry: in slowly |
| VM boot to in-service | 60–300s | Why reactive scaling can't catch fast spikes |
| Container start on an existing node | 2–30s (+ app warm-up) | Why node-level headroom matters |
| New node provisioning (cluster autoscaler) | 60–180s | Pods pending while nodes boot |
| JVM/JIT warm-up to steady-state latency | 30–180s | New instances serve slowly at first |
| Serverless cold start | ~100ms–1s typical; several seconds for heavy runtimes | Scale-to-zero's hidden latency tax |
| Target utilization (user-facing) | 50–65% | Leaves room for zone loss and bursts |
| Target utilization (batch / async) | 80–90% | Queues absorb bursts; latency isn't the SLO |
| Zone-loss headroom (3 AZs) | Each AZ must carry 1.5× normal → run at ≤ 66% | N+1 at zone granularity |
| Region-failover headroom (3 regions) | Each region takes +50% → run at ≤ 60–66% | Set by the evacuation scenario |
| Load-test multiple | 1.5–2× forecast peak | Forecast error on launches is 20–50% |
| Latency knee | Usually 70–85% utilization | Queueing delay grows as 1/(1−ρ): 80% → 5×, 90% → 10× service time |
| Retry amplification | 2–3× offered load during brownouts | Why capacity math must include retries |
Interview Walkthrough
The most common mistake: candidates draw an autoscaling group, write "CPU > 70% → add instance," and stop there. Everything that decides whether the launch survives comes after that sentence: the actuation delay, the dependency that doesn't scale, the load test that lied, and who paid for the headroom.
Phase 1: Requirements & Framing (2–3 minutes)#
State the functional goal in one sentence:
"We want the service to meet its latency SLO as demand changes, without paying for peak capacity 24/7."
Then frame the real requirements:
"Three intents pull in different directions: minimizing cost on a variable curve, surviving known peaks like launches, and surviving unknown spikes and failures like a zone outage. I'll design for known peaks with a survival floor. My constraints are p99 under 300ms, at most 1% of requests shed at the worst minute, the ability to lose one availability zone without paging, and a cost target of no more than ~40% idle averaged over a day."
Establish the shape of demand, because it determines everything:
"What does the curve look like? I'll assume a diurnal pattern from 20K RPS at night to 100K at peak, rising over about 2 hours, plus planned launches that step from 100K to 250K in under 2 minutes. The diurnal curve is slow enough for reactive scaling. The launch step isn't."
🎯 Staff Move: Asking for the rise time of demand, not just the peak, is the question that separates levels. Compare it to your provisioning time and you know which tool you need before drawing anything.
Phase 2: Core Entities & API (1–2 minutes)#
The nouns:
- Capacity unit — the measured throughput of one instance at SLO (e.g., 400 RPS per 4-vCPU instance at p99 < 300ms). The single most important number in the design.
- Scaling policy — signal, target, min, max, cooldowns, step limits.
- Capacity floor — scheduled or predicted minimum for a time window.
- Dependency budget — max connections/RPS each downstream grants this service.
- Event — a calendar entry with expected multiplier, owner, and pre-scale window.
The control API (platform-facing, not end-user):
PUT /services/{svc}/scaling-policy
{ signal: "inflight_per_instance", target: 60, min: 40, max: 300,
scale_out: { max_step_pct: 50, cooldown_s: 60 },
scale_in: { max_step_pct: 10, stabilization_s: 600 } }
POST /services/{svc}/capacity-floors
{ start: "2026-11-27T13:00Z", end: "...", min_instances: 220, reason: "BF launch", approver: "..." }
GET /services/{svc}/capacity-model
→ { rps_per_instance_at_slo: 400, max_safe_instances: 280,
limiting_dependency: "orders-db connections", headroom_pct: 38 }
🎯 Staff Move: "The capacity model endpoint matters more than the policy. If a team can't tell me their RPS per instance at SLO and their limiting dependency, their autoscaling policy is a guess. Everything else I design assumes that number exists and is re-measured regularly."
Phase 3: High-Level Architecture (≤5 minutes)#
Walk it in 90 seconds:
- The service emits in-flight requests per instance every 15 seconds.
- The reactive controller computes
desired = ceil(current × observed / target)and clamps it between the floor and the ceiling. - The floor comes from the forecast (trailing-weeks seasonality) and the event calendar (pre-scales for launches).
- The ceiling comes from the capacity model: the smallest of budget, DB connection budget ÷ per-instance pool size, and third-party quota ÷ per-instance call rate.
- New capacity comes from the warm pool (~30s) before cold boot (~3 min). The LB slow-starts new targets over 60 seconds.
- If demand exceeds what the fleet can serve, the gateway sheds lowest-priority traffic first. It returns 503 with Retry-After rather than letting latency collapse for everyone.
🎯 Staff Move: "That's the design most people draw, minus the floor and ceiling. The floor handles what we know is coming. The ceiling stops us from DDoSing our own database. Admission control handles what neither covers. Now let me show where it breaks."
Phase 4: Transition to Depth (1 minute)#
"Four places this gets interesting: choosing the signal, the gap between spike rise time and provisioning time, the dependencies that don't scale, and how much headroom to buy and who pays for it. Which do you want first?"
If the interviewer has no preference, lead with the actuation gap. It's the most counter-intuitive point, and the other topics follow from it.
Phase 5: Deep Dives (25–30 minutes)#
Deep dive A: The actuation gap (6–8 min)
"Add up the delays. Metric scrape 15s, evaluation window 60s, API call and scheduling 10s, VM boot 90s, app start 30s, JIT and cache warm-up 60s, LB health checks and slow-start 30–60s. That's about 5 minutes from spike to useful capacity. A launch that goes from 100K to 250K RPS in 90 seconds is over before the autoscaler contributes anything."
Three tools for the gap, in order of preference:
- Know it's coming and pre-scale. The event calendar and forecast floor cover launches and diurnal ramps.
- Shorten the delay. Warm pools (pre-booted, stopped instances attach in ~30s), pre-pulled images, node headroom so pods don't wait on new nodes, and pre-warmed caches.
- Absorb the remainder with standing headroom plus admission control that sheds the lowest-priority traffic.
"For a truly surprise spike, say a viral post, the formula is: headroom must cover (spike rate × actuation delay). If we can tolerate a 2× surprise within 5 minutes, we run at ~50%. That's expensive, so I'd do it only for tier-0 services and shed for the rest."
Deep dive B: Choosing the signal (5–6 min)
| Signal | Use when | Failure |
|---|---|---|
| CPU | CPU-bound compute (encoding, crypto, rendering) | I/O-bound services sit at 20% CPU while latency explodes |
| In-flight requests / concurrency | Request/response services | Needs a measured "good concurrency per instance" |
| RPS per instance | Homogeneous request cost | Mix shifts (e.g., more expensive search queries) mislead it |
| Queue backlog ÷ drain rate | Async workers | Needs per-worker throughput; poison messages look like load |
| Latency | Last-resort guardrail | Rises for reasons scaling can't fix: a slow DB, GC. Scaling out makes a slow DB worse |
"Latency is the worst primary signal, because the most common cause of high latency is a slow dependency. Scaling out on latency then adds connections to the thing that's already struggling. I use concurrency as the primary signal and treat latency as an alert, not a scaling input."
Deep dive C: The dependencies that don't scale (6–7 min)
"Our app tier goes from 40 to 300 instances. Each has a 20-connection DB pool, so that's 6,000 connections against a Postgres primary configured for 5,000, whose sweet spot is much lower because each Postgres connection is a process. Scaling out converts an app-tier problem into a database outage."
Mitigations:
- Connection pooler (PgBouncer in transaction mode) multiplexes thousands of client connections onto ~200–500 server connections.
- Max instances derived from the dependency budget, recomputed when pool sizes change.
- Per-consumer quotas on shared dependencies, so one service's scale-out can't starve another.
- Cache and queue in front of writes to decouple app-tier elasticity from DB write capacity (see Scaling Writes).
- Third-party quotas as hard ceilings, with degraded modes when they're hit (see Degraded Mode).
Deep dive D: Headroom and who pays (5–6 min)
"Three AZs, each carrying a third. Lose one and the other two must carry 1.5× their load, so steady-state utilization can't exceed ~66% of the knee. The knee for this service is about 80% CPU-equivalent, so the target is ~0.66 × 80% ≈ 53%, call it 55%. On a $400K/month compute bill that's roughly $180K/month of 'idle.' It isn't idle. It's the price of zone failure being a non-event. Finance should see that line item with that label."
Phase 6: Wrap-Up (2–3 minutes)#
"The design has three layers: a forecast and calendar floor for what we know, reactive target tracking on concurrency for what we don't, and admission control for the gap. It's capped by a ceiling that protects dependencies that can't scale. Headroom is set by the zone-failure scenario and priced explicitly. The capacity model, meaning RPS per instance at SLO and the limiting dependency, is re-measured every quarter and after every major release."
The org close:
"Mechanically, the platform team owns the controllers. Each service team owns its capacity model and its event calendar entries. A capacity review each quarter makes idle spend and headroom visible per team. Most capacity incidents I'd expect aren't controller bugs. They're a team that didn't tell anyone about a launch, or a shared database nobody owned the budget for."
🎯 Staff Move: End with where the next outage comes from. For capacity, it's almost always a process gap (an unannounced event, a stale capacity model) or a shared dependency, rarely the autoscaler itself.
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| 10 min on the scaling formula | Derives target-tracking math, cooldown tuning | States the formula in one line, moves to delays and dependencies |
| No demand shape | "Traffic spikes" | "Rise time 90s vs. provisioning 5 min, so the gap is covered by X" |
| Only the elastic tier | Designs ASG/HPA and stops | Spends 5+ min on the DB, pools and third-party quotas |
| Headroom as vibes | "Some buffer" | "55% target, covers AZ loss, costs $180K/month" |
| Load test as a checkbox | "We'll load test" | Describes realistic data, cache state, request mix, and finding the knee |
| No org story | Ends on autoscaler config | Ends on capacity reviews, event calendar, dependency budgets |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Capacity is where engineering meets finance, and where a design meets the rest of the company. The autoscaler itself is a commodity: every cloud and Kubernetes distribution ships one. What isn't a commodity is the judgment around it. Which signal actually measures demand for this workload? Which time constants make reactive scaling useless? Which components silently won't scale? How much should the company pay to make a zone failure boring?
It's also the purest test of the Staff thesis. The L5 answer ("autoscale on CPU") is correct most of the time. The L6 answer is opinionated about the minute it isn't: the launch spike, the zone loss, the DB connection cliff. And it names who gets paged and who paid for the headroom.
1.2 The L5 vs L6 vs L7 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"How fast does demand rise compared to how fast capacity arrives, and who pays for the difference?"
- If demand rises slower than provisioning (diurnal curves over hours), reactive scaling works and the question is tuning.
- If demand rises faster and it's known (launches, sales), you pre-scale and the question is process: who files the event and who approves it.
- If demand rises faster and it's unknown, you choose between paying for standing headroom (finance pays) and shedding load (some users pay). You say which, for which traffic tier, and who signed off.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Cost efficiency. Common for internal services, batch processing, dev/test and variable workloads without strict latency SLOs. Run hot (70–90% utilization for async), scale in aggressively, use spot/preemptible instances with interruption handling, and scale to zero outside business hours. The risk is latency SLO breaches during ramps, which the team explicitly accepts.
Reliability through known peaks. Consumer products with launches, flash sales, sports events, tax deadlines and Monday mornings. The work is forecasting, calendaring, load testing and pre-scaling, and making sure every dependency on the path is sized for the event, not only the stateless tier. Failures here are usually process failures: marketing sent the email early, the forecast used last year's numbers, or the DB wasn't included in the load test. See Flash Sales.
Survival of unknown spikes and failures. Tier-0 services such as login, checkout, and core APIs other services depend on. They need standing headroom for zone and region loss, load shedding with priority tiers, and protection from retry storms. Here the design question is how the service degrades, not how it scales.
🎯 Staff Move: "These aren't mutually exclusive. They're tiers. I'd classify each service: tier-0 gets survival headroom, consumer-facing gets known-peak planning, and internal/batch gets cost efficiency. One autoscaling policy for all 200 services is how you get either an outage or a bill."
2.2 When NOT to Auto-Scale#
| Situation | Why autoscaling is wrong | Do instead |
|---|---|---|
| Stateful primaries (DB primaries, Kafka brokers, ZooKeeper) | Adding a node means rebalancing data, which takes hours and adds load mid-incident | Plan capacity quarterly; scale vertically ahead of need; shard deliberately |
| Workloads with long warm-up (large caches, ML models loading 20 GB) | New instances aren't useful for minutes, and cold caches hammer the backend | Fixed capacity + pre-warming; scale on schedule |
| Spiky load that's shorter than provisioning time | Capacity arrives after the spike ends, so you pay for it and it doesn't help | Headroom + shedding + queueing |
| License-bound or quota-bound dependencies | More instances just hit the external limit faster | Cap at the quota; degrade gracefully |
| Very small fleets (≤ 3 instances) | Scale-in to 1 removes redundancy; each step is a 33% change | Fixed N+1 with a scheduled floor |
| Latency-critical services with tight tails where warm-up hurts p99 | New cold instances inflate tail latency | Over-provision; scale on schedule only |
🎯 Staff Move: "The database is where I stop autoscaling and start capacity planning. I'll scale the stateless tier elastically and plan the stateful tier a quarter ahead, with growth tracking and a 6-month runway alert."
2.3 What the Interviewer Leaves Underspecified#
| Underspecified | Why it matters | What I'd assume out loud |
|---|---|---|
| Demand shape and rise time | Decides reactive vs. scheduled vs. headroom | Diurnal 5× over 2h; launches 2.5× in 90s |
| Latency SLO | Sets the utilization target (knee) | p99 < 300ms |
| Workload type | CPU-bound vs. I/O-bound → signal choice | I/O-bound API calling a DB and one third party |
| Failure scenarios to survive | Sets headroom | Lose 1 of 3 AZs without paging |
| Cost constraints | Bounds headroom | ≤ 40% idle daily average; tier-0 exempt |
| Dependency limits | Sets the ceiling | Postgres 5,000 connections; third party 2,000 RPS |
| Cloud vs. own hardware | Actuation time: minutes vs. months | Public cloud, with reserved commitments for the baseline |
2.4 Precise Terminology#
| Term | Meaning | Why the precision matters |
|---|---|---|
| Capacity unit | Throughput of one instance at SLO | Everything else is derived from it; must be measured, not assumed |
| Knee | Utilization where latency starts rising non-linearly | Target utilization sits below it with margin |
| Headroom | (Capacity − demand) ÷ capacity at a point in time | Must be tied to a scenario (AZ loss, 2× surprise) |
| Actuation delay | Time from demand change to useful new capacity | Compared against rise time |
| Floor / ceiling | Min capacity from forecast/schedule; max from budget/dependencies | Clamps the reactive controller |
| Stabilization window | Period the controller considers before scaling in | Prevents flapping |
| Warm pool | Pre-initialized but idle/stopped capacity | Cuts actuation delay at a small cost |
| Load shedding | Rejecting some requests on purpose to protect the rest | The honest answer when demand > capacity |
| Retry amplification | Offered load multiplied by client retries during slowness | Capacity math must include it |
| Reserved / committed capacity | Capacity bought for 1–3 years at a discount | The baseline; autoscaling handles the variable part above it |
3. The Five Fault Lines#
3.1 Fault Line 1: Scaling Signal#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| CPU utilization | Universal, no instrumentation, correct for CPU-bound work | I/O-bound services look idle while saturated; GC and noisy neighbors distort it | Users during brownouts the scaler never saw |
| Concurrency (in-flight per instance) | Directly measures work in progress; Little's Law ties it to RPS × latency | Rises when a dependency slows, so it can scale out into a slow DB | The DB, unless the ceiling protects it |
| RPS per instance | Leading indicator, easy to reason about | Wrong when request mix shifts (expensive queries) | Users when mix shifts |
| Queue backlog ÷ (drain rate per worker × target drain time) | The right signal for async work; directly expresses the SLO ("drain within 5 min") | Poison messages and retry loops inflate backlog | The downstream that the workers hammer |
| Latency | What users feel | Ambiguous cause; scaling out on dependency-induced latency amplifies the problem | The dependency and every other consumer of it |
| Custom business signal (e.g., active sessions) | Leading, domain-accurate | Needs a model mapping signal → capacity | The team maintaining the model |
Staff default: concurrency for synchronous services, backlog-over-drain-rate for async, CPU only for CPU-bound workers. Latency is an alert, never a primary scaling input. Always pair the reactive signal with a ceiling, so that "concurrency rising because the DB slowed down" can't turn into 300 instances hammering the DB.
"Little's Law gives me the target. If one instance meets SLO at 400 RPS with 150ms average latency, it holds 400 × 0.15 = 60 requests in flight. I target 60 and scale on it. If latency doubles because the DB slows, in-flight doubles too, and the ceiling is what stops me from making it worse."
🧭 Principal Move: "I'd publish a signal standard per workload class (request/response, stream consumer, batch, GPU inference) so 200 teams don't rediscover that CPU is wrong for I/O-bound services one incident at a time."
3.2 Fault Line 2: Reactive vs Predictive#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Reactive only | Always tracks reality; no model to maintain | Late by the actuation delay; fails on fast ramps | Users in the first 3–5 minutes of every ramp |
| Scheduled | Zero lag for calendared events | Stale schedules waste money; unannounced events uncovered | Finance (stale floors), users (missed events) |
| Predictive (seasonality forecast) | Pre-provisions diurnal/weekly ramps | Wrong on holidays, launches, anomalies; needs weeks of history | Users when the forecast is low; finance when it's high |
| Predictive floor + reactive above | Forecast sets the minimum; reactive handles the surprise upside | Two systems to reason about; the floor must never cap the reactive side | Platform team's complexity budget |
Staff default: forecast and schedule set the floor. Reactive scaling tracks the actual signal above it. The controller takes max(floor, reactive_desired), clamped by the ceiling. A wrong forecast therefore costs money (too high) or falls back to reactive behavior (too low). It never removes the reactive path.
When to deviate: services with flat load can skip prediction. Services with extreme predictable spikes (a daily 9:00:00 batch trigger, a sports kickoff) should use scheduled capacity, since forecasting adds nothing there.
🎯 Staff Move: "Predictive scaling should only ever raise the floor. The day it's allowed to lower capacity below what reactive wants is the day a bad forecast becomes an outage."
3.3 Fault Line 3: Cold Start vs Warm Capacity#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Cold scale-out | No idle cost | 3–5 min actuation; cold instances serve slowly and can fail health checks | Users during ramps |
| Warm pools (pre-booted, stopped) | Attach in ~20–40s; storage cost only while stopped | Pool can be exhausted; images drift if not refreshed | Small standing cost |
| Node overprovisioning (placeholder pods with low priority) | Pods schedule instantly on spare nodes | Paying for idle nodes | Finance (~5–15% of cluster) |
| Provisioned concurrency (serverless) | Eliminates cold starts for N concurrent | Pay per provisioned unit whether used or not | Finance |
| Pre-warming (synthetic traffic, cache priming, JIT warm-up) | New instances hit steady-state latency before real traffic | Needs representative warm-up traffic | The service team's effort |
| LB slow-start | New targets ramp traffic over 30–120s | Slows how fast new capacity absorbs load | Slight delay in relief |
Staff default: warm pool for VM fleets, low-priority placeholder pods for Kubernetes (~10% node headroom), pre-warming on startup (hit the top 1,000 cache keys, run a synthetic request loop until p99 stabilizes) before registering with the LB, and LB slow-start of 60s.
🎯 Staff Move: "A new instance that isn't warm is worse than no instance. It takes a full share of traffic, serves it slowly, fails health checks, and gets replaced by another cold instance. Warm-up has to finish before registration."
3.4 Fault Line 4: Headroom vs Cost#
| Target utilization | Covers | Monthly cost of headroom on a $500K fleet | Who Pays if Wrong |
|---|---|---|---|
| 85% | Almost nothing; latency near the knee | ~$75K | Users: latency spikes, zone loss is an outage |
| 70% | Moderate bursts; not zone loss | ~$150K | Users during AZ failure |
| 60% | Loss of 1 of 3 AZs (×1.5) with a thin margin | ~$200K | Balanced |
| 50% | AZ loss + ~1.3× surprise, or region failover in a 3-region setup | ~$250K | Finance |
| 35% | Region loss in a 2-region active-active setup | ~$325K | Finance, heavily |
Headroom cost ≈ fleet cost × (1 − utilization), assuming the fleet is sized to peak. Autoscaling reduces it off-peak but not at peak.
Staff default: set the target by the failure scenario you've committed to survive, then check it against the latency knee. For a 3-AZ user-facing service, ~55–60%. For async workers, 80–90% (the queue is the headroom). For tier-0 services with region-failover requirements, whatever the failover math says, even if it's 45%.
"I'd rather say '60%, because losing one of three zones multiplies load by 1.5 and our knee is at ~85%' than '60% is industry standard.' The first is a decision. The second is a habit."
🧭 Principal Move: "Headroom should be a line item with a scenario attached: '$200K/month buys AZ-loss tolerance for checkout.' Then leadership can make the trade explicitly. Most orgs pay for headroom they can't explain and skip headroom they need."
3.5 Fault Line 5: Central Capacity vs Service Ownership#
| Model | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Central capacity team plans everything | Global view; purchasing leverage; shared-dependency arbitration | Bottleneck; can't know each workload; teams hide growth | Service teams waiting on tickets |
| Each team fully owns its capacity | Fast; knowledgeable | No one owns shared DBs and quotas; tragedy of the commons; duplicated tooling | Whoever shares a dependency with the loudest team |
| Platform mechanism + service-owned models + central review | Platform owns controllers and tooling; teams own their capacity models and events; a central group owns shared budgets and the quarterly review | Needs discipline and a review cadence | Everyone, a little |
Staff default: the hybrid. Specifically:
- Platform team: autoscaling controllers, warm pools, node provisioning, load-test tooling, capacity dashboards.
- Service team: capacity model (RPS per instance at SLO, limiting dependency), scaling policy, event calendar entries, load tests before launches.
- Shared-dependency owners (DB platform, API platform): publish per-consumer budgets and enforce them.
- Capacity review (quarterly): growth vs. runway, headroom vs. cost, top risks.
🎯 Staff Move: "The failure I'd design against is the unannounced launch. The fix is process, not technology: an event calendar that marketing and product write to, and a pre-scale approval with a named owner."
4. Failure Modes & Operational Reality#
4.1 The Scale-Out Death Spiral#
t=0: Spike: 2× traffic in 60s.
t=+45s: p99 from 200ms to 2s. Health checks (1s timeout) begin failing on busy instances.
t=+60s: LB ejects 30% of instances. The rest go to 140% of capacity.
t=+90s: Autoscaler adds 60 instances. Clients retrying, so offered load is 3×.
t=+4min: New instances register cold, get a full share, fail health checks, get replaced.
t=+5min: New app instances add 1,200 DB connections. DB CPU 100%. All instances slow.
t=+8min: Every instance is "unhealthy" at some point. Effective capacity near zero.
Detection: lb.healthy_host_ratio falling while autoscaler.desired_capacity rises; db.active_connections near max; client.retry_ratio > 1.5.
Mitigation: freeze scale-in and instance replacement (stop the churn); enable gateway admission control to cut offered load to known capacity; separate liveness from readiness checks (a slow instance is busy, not dead); cap scale-out at the DB budget.
Prevention: health checks that test "the process is alive," not "the process is fast"; LB slow-start; warm-up before registration; retry budgets in clients (≤ 10% extra load); circuit breakers (Circuit Breakers); ceiling from dependency budgets.
Owner: service on-call leads; platform on-call freezes the controllers; DB on-call guards the database.
4.2 Scaling on a Symptom: Latency-Driven Scale-Out into a Slow Database#
A slow query plan makes DB latency jump from 5ms to 80ms. App concurrency rises 16×, and the reactive controller (on concurrency) scales from 60 to 300 instances. The DB now has 5× the connections and is slower still.
Detection: autoscaler.scale_out_events correlated with db.query_latency_p99 rising before request rate did. The signature is capacity rising while RPS stays flat.
Mitigation: freeze scaling; fix or roll back the plan; shed.
Prevention: a flat-RPS guard: don't scale out if RPS hasn't risen more than 20% in the window, since rising concurrency with flat RPS means a dependency slowed. Also the ceiling from the dependency budget.
Owner: service team (policy), DB team (query plan).
4.3 Scale-In Flapping and Connection Drops#
A too-short stabilization window makes the fleet oscillate between 80 and 120 instances every 4 minutes. Each scale-in terminates instances holding long-lived connections (WebSockets, gRPC streams), which reconnect and spike load on the rest, which triggers scale-out.
Detection: autoscaler.direction_changes_per_hour > 4; connections.reset_total spikes aligned with scale-in.
Mitigation: lengthen the scale-in stabilization window to 10–15 min; cap the scale-in step at 10% per window; use connection draining (up to 5 min) before termination.
Owner: service team.
4.4 Capacity Ceiling Nobody Knew About#
Launch day: the fleet scales perfectly to 280 instances, and the payment provider's API rate limit of 2,000 RPS fails 30% of checkouts. Or the cloud account's vCPU quota in the region stops scale-out at 180 instances. Or the NAT gateway's port allocation is exhausted.
Detection: autoscaler.scale_out_failures{reason=quota}; third-party 429 rates; nat.port_allocation_errors.
Mitigation: emergency quota increase (hours to days, which is too slow for today); shed low-priority traffic; queue payments for async capture.
Prevention: a limits inventory per service: cloud quotas, third-party limits, NAT/IP, license counts. Checked in the pre-launch review, with alerts at 70% of every limit.
Owner: service team for the inventory; platform for cloud quotas.
4.5 Forecast Wrong by 2×#
Predictive floor built from the last 4 weeks; a holiday week cuts the actual load in half. Paying 2× for a week is a cost incident, not an outage. The opposite case, a viral event at 2× the forecast, falls back to reactive plus headroom plus shedding.
Detection: forecast.error_pct (actual/forecast − 1) tracked daily; alert when |error| > 30% for 2 hours.
Owner: platform (forecasting), service team (event calendar hygiene).
4.6 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Death spiral | lb.healthy_host_ratio ↓ while desired_capacity ↑ | Whole service + shared DB | Freeze replacement, shed, cap at DB budget | Service on-call + platform |
| Scale into slow dependency | Scale-out with flat RPS | Dependency + all its consumers | Flat-RPS guard, freeze, fix query | Service + DB team |
| Flapping | direction_changes_per_hour > 4 | Connection-heavy clients | Longer stabilization, smaller steps, draining | Service team |
| Hidden ceiling (quota, third party) | scale_out_failures{reason}, third-party 429s | Feature or whole service | Shed, queue, emergency quota | Service team + platform |
| Forecast error | forecast.error_pct | Cost (high) or latency (low) | Reactive covers low; review floors | Platform |
| Warm pool exhausted | warm_pool.available == 0 | Ramp latency | Cold fallback; resize pool | Platform |
| Unannounced event | Sudden step in RPS with no calendar entry | Service, sometimes company-wide | Shed; emergency pre-scale | Product/marketing + service team |
| Retry storm | client.retry_ratio > 1.5 | All services sharing the path | Retry budgets, jittered backoff, shed | Client owners + gateway team |
🎯 Staff Move: "When I read this matrix, the autoscaler is at fault in maybe one row. The rest are process, dependencies and client behavior. That's why I'd spend my design time on those."
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | "Autoscale on CPU" | Demand shape, rise time vs. actuation delay; intent by tier | Capacity as a budget and a portfolio: baseline commitments + elastic + headroom by tier |
| Signal | CPU/memory | Concurrency/backlog with Little's Law; latency as alert only | Org-wide signal standards per workload class |
| Timing | Reactive + cooldowns | Floor (forecast/schedule) + reactive + warm pools + shedding for the gap | Plans quarterly and annual capacity with finance; hardware and quota lead times |
| Dependencies | Adds replicas | Ceiling from DB/quotas; pooling; flat-RPS guard | Shared-dependency budgets and arbitration across teams |
| Headroom | "Some buffer" | Target tied to a failure scenario; priced | Utilization targets by tier as policy; headroom as a line item with a scenario |
| Validation | "Load test" | Production-like load tests to 1.5–2×; find the knee | Continuous capacity measurement in production (Kraken-style); game days |
| Ownership | Platform team | Platform mechanism, service-owned models, event approver | Operating model: capacity reviews, showback, error-budget links |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Compares time constants | "Spike rises in 90s, capacity arrives in 5 min. Reactive covers nothing here, so I pre-scale." |
| Right signal | "It's I/O-bound. I scale on in-flight requests: 400 RPS × 150ms = 60 per instance." |
| Protects the non-elastic | "Max instances is 5,000 DB connections ÷ 20 per pool, minus a margin: 220." |
| Prices headroom | "60% target covers AZ loss and costs ~$200K/month. That's the price of AZ failure being a non-event." |
| Knows the process failure | "The outage I'd fear is the launch nobody told us about. I'd build an event calendar with an approver." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| CPU for everything | Misreads I/O-bound saturation, which is the most common case |
| "The autoscaler will handle spikes" | Ignores the actuation delay |
| Scales the DB like the app tier | Doesn't know stateful systems rebalance slowly |
| No numbers on headroom or cost | Can't reason about the tradeoff the question is about |
| No load shedding | Assumes capacity always catches up with demand |
5.4 Common False Positives#
- Knowing HPA YAML fields ≠ capacity judgment. Configuration fluency doesn't mean the candidate knows which signal to pick.
- Queueing theory derivations ≠ design. M/M/1 math is useful for one sentence ("latency ∝ 1/(1−ρ)"), not for ten minutes.
- "We'll use serverless" ≠ solving scaling. It moves the problem to cold starts, concurrency limits and downstream protection.
- Kubernetes cluster autoscaler internals ≠ capacity planning. Node provisioning is one delay in the chain.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Intents by tier; demand shape; SLO; failure scenario to survive |
| Entities | 3–5 min | Capacity unit, policy, floor, ceiling, event |
| Architecture | 5–10 min | Controller with floor/ceiling; warm pool; admission control |
| Deep dive 1 | 10–18 min | Actuation gap: delays, pre-scale, warm-up |
| Deep dive 2 | 18–26 min | Signal choice; scaling into a slow dependency |
| Deep dive 3 | 26–36 min | Dependencies & ceiling; headroom & cost |
| Deep dive 4 | 36–41 min | Load testing and validating the capacity model |
| Wrap-up | 41–45 min | Ownership, process, capacity review |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response |
|---|---|---|
| "Traffic triples in 30 seconds, no warning." | Unknown-spike posture | Headroom + admission control by priority; reactive recovers minutes later; state who is shed |
| "Cut compute cost by 30%." | Cost levers without breaking reliability | Tier services; raise async targets to 85%; scale-to-zero non-prod; commit the baseline with reservations; spot for stateless batch |
| "It's a Kafka consumer, not an API." | Signal for async | Lag ÷ (per-consumer throughput × target drain time); ceiling = partition count |
| "Now it's GPU inference." | Scarce, slow-to-provision capacity | Minutes-to-hours provisioning; queue + batch; capacity reservations; admission by priority |
| "How do you know 400 RPS per instance is right?" | Measurement discipline | Load test to the knee in a prod-like env; verify in prod with traffic shifting; re-measure after releases |
6.3 What to Deliberately Skip#
- The exact target-tracking algorithm and PID tuning. One sentence is enough.
- Instance type selection details. Mention only if CPU:memory ratio matters.
- Cloud-provider-specific policy syntax.
- Detailed queueing theory beyond the knee.
6.4 Follow-Up Questions to Expect#
- "What's your scale-in policy?" Slow (10–15 min stabilization), small steps (≤ 10%), connection draining, never below the floor.
- "How do you avoid scaling on noise?" Evaluate over 1–3 min windows, require N consecutive breaches for scale-in, but allow a single strong breach to trigger scale-out.
- "How do you handle multi-tenant load on shared services?" Per-tenant quotas at the gateway (Rate Limiting) so one tenant's spike doesn't drive everyone's capacity.
- "What about stateful services?" Plan quarterly; vertical first; add read replicas ahead of need; shard deliberately (Database Sharding).
- "How do you load test safely?" A prod-like environment with production-sized data, or production traffic shifting with automated abort thresholds.
- "How do you forecast?" Weekly seasonality × growth trend × event multipliers; compare forecast error daily.
- "Who decides to pre-scale for Black Friday?" The service owner files it; the capacity reviewer approves; finance sees the cost; the platform team executes.
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design auto-scaling for our API."
Staff Answer
"First, what does the demand curve look like? How fast does it rise, and are the big peaks known in advance? I'll assume a diurnal 5× swing over ~2 hours plus planned launches that step 2.5× in 90 seconds. Second, what are we optimizing: cost, surviving known peaks, or surviving failures? I'll design for known peaks with a survival floor of losing one AZ without paging. Third, what's the workload? If it's I/O-bound, CPU is the wrong signal. The design is a controller with three inputs: a floor from forecast and event calendar, a reactive target on in-flight requests per instance, and a ceiling from our dependencies' budgets. Admission control at the gateway covers the gap. I'll spend most of our time on the actuation gap, the dependencies that don't scale, and what the headroom costs."
Why this is L6:
- Asks for rise time, not just peak
- Chooses the signal based on workload type
- Introduces floor/ceiling and shedding before being asked
What L7 adds:
- "Which tier is this service? Tier-0 gets region-failover headroom. Internal services run hot."
- Frames baseline capacity as a reserved-commitment purchase with finance
❌ Common L5 Trap
"I'll set up an auto-scaling group with a target of 70% CPU, min 10, max 100, and a 5-minute cooldown."
Why this misses: It's a valid configuration. But it doesn't know whether CPU measures this workload's demand, whether 5 minutes of lag is survivable, or whether 100 instances will overrun the database. The interviewer will extract all three with follow-ups.
Drill 2: Compute the Actuation Gap#
Prompt: "Traffic doubles in 60 seconds. Walk me through the timeline of your autoscaler."
Staff Answer
"t=0 spike. Metrics scraped at 15s granularity; the controller needs the target breached over a 60s window, so the decision lands at ~t+60–75s. The cloud API and scheduling take ~10s. VM boot 60–120s. App start plus image pull 20–40s. Warm-up (JIT, connection pools, cache) 60s. LB slow-start ramps over 60s. Useful capacity arrives around t+4–5 minutes. For those 4–5 minutes, the existing fleet absorbs 2× load. If we're at 55% utilization, 2× puts us at 110%, past the knee. So either we run at ≤ 40% (expensive), or the gateway sheds ~15–20% of low-priority traffic for 4 minutes, or we cut the delay: a warm pool brings it to ~90s, and node headroom for containers brings it to ~45s. I'd use a warm pool plus shedding, and accept a few minutes of degraded low-priority traffic for surprise doublings."
Why this is L6:
- Adds up each delay explicitly
- Converts the gap into a utilization requirement
- Chooses a combination and names who pays (low-priority traffic)
What L7 adds:
- Turns "surprise 2× in 60s" into a written survivability requirement per tier, with product sign-off on what gets shed
- Tracks actuation delay as a platform SLO (e.g., p95 < 90s for warm-pool attach)
Drill 3: Pick the Signal#
Prompt: "Our service averages 25% CPU but p99 latency hits 2s at peak. The autoscaler never fires. What's wrong?"
Staff Answer
"The service is I/O-bound. It spends its time waiting on a DB or downstream, so threads or event loops saturate long before CPU does. CPU is measuring the wrong thing. I'd switch the signal to in-flight requests per instance. From a load test, find the concurrency where p99 hits the SLO. Say an instance meets SLO at 400 RPS × 150ms average = 60 in flight, so the target is 60. Before switching, I'd confirm the latency isn't caused by a slow dependency. If DB latency rose first, scaling out adds connections and makes it worse. So I'd add a flat-RPS guard and a ceiling from the DB's connection budget."
Why this is L6:
- Diagnoses the I/O-bound mismatch
- Uses Little's Law to derive a target instead of guessing
- Anticipates the danger of the new signal
What L7 adds:
- Publishes a per-workload-class signal standard so this isn't rediscovered team by team
- Adds "capacity unit measured in the last 90 days" as a launch-readiness check
Drill 4: The Dependency Ceiling#
Prompt: "The app tier scaled from 50 to 250 instances during a sale and the database fell over. Design so this can't happen."
Staff Answer
"250 instances × 25 pool connections = 6,250 connections against a 5,000 limit, and Postgres degrades long before its limit because each connection is a process. Four changes. One, a connection pooler (PgBouncer, transaction mode) so app connections multiplex onto ~300 server connections. Two, max instances derived from the dependency budget: min(budget, pooler_client_limit ÷ pool_size, db_write_capacity ÷ writes_per_instance_at_target). Three, load shedding at the gateway once the fleet hits its ceiling, so excess demand is rejected at the edge instead of queuing in the DB. Four, move non-critical writes behind a queue so the DB sees smoothed write load. The rule I'd state: the elastic tier's maximum is a function of the least elastic thing it calls."
Why this is L6:
- Quantifies the connection math
- Derives the ceiling instead of picking a number
- Moves the failure to the edge, where it's cheap
What L7 adds:
- Per-consumer budgets on the shared DB, published by the DB platform team, enforced at the pooler
- Adds the DB to every launch's load-test scope as a policy
Drill 5: Headroom Math#
Prompt: "Why do you run at 55% and not 80%? Finance wants 80."
Staff Answer
"80% gives us nothing to lose. Three AZs, and if one fails the other two must absorb 1.5× their load: 80% × 1.5 = 120%, an outage. And latency rises non-linearly: queueing delay goes as 1/(1−ρ), so at 80% it's ~5× the service time and at 90% ~10×. Our measured knee is ~80%. To survive AZ loss below the knee we need 80% ÷ 1.5 ≈ 53%, so 55%. On a $400K/month fleet, 55% vs 80% costs ~$125K/month. What finance is buying is 'an AZ failure is a non-event instead of a 1–2 hour customer-facing outage.' If an outage of that length costs more than ~$125K times its probability over a month, and for checkout it does, we pay. For internal batch services I'd happily run at 85%, since the queue absorbs everything."
Why this is L6:
- Ties the number to a failure scenario and the latency knee
- Prices it and frames the trade for finance
- Differentiates by tier
What L7 adds:
- Establishes utilization targets as an org policy by tier, reviewed annually with finance
- Compares with reserved-instance/commitment discounts that lower the cost of headroom by 30–50%
Drill 6: Queue Workers#
Prompt: "Scale a fleet of workers consuming from a Kafka topic with 64 partitions."
Staff Answer
"Signal: consumer lag in time, not messages. Desired workers = ceil(lag_msgs ÷ (throughput_per_worker × target_drain_seconds)), plus workers for the steady-state arrival rate. The hard ceiling is 64 consumers, since Kafka assigns at most one consumer per partition in a group. Beyond that, extra workers sit idle. If we need more, we raise partition count (painful, changes key distribution) or make each consumer process concurrently within a partition while preserving per-key order. Scale-in carefully: every consumer group membership change triggers a rebalance that pauses consumption for seconds, so the stabilization window should be 10+ minutes. Poison messages inflate lag without real load, so I'd route to a DLQ after N retries and exclude DLQ'd work from the signal."
Why this is L6:
- Uses the right signal and formula for async work
- Knows the partition-count ceiling and rebalance cost
- Guards against lag caused by poison messages
What L7 adds:
- Partition count is a capacity decision made at topic creation. Set a standard (e.g., sized for 3-year peak ÷ per-consumer throughput)
- Downstream protection: the workers' target system needs a budget too
Drill 7: Load Testing That Tells the Truth#
Prompt: "Your load test showed we could handle 3× peak. On launch day we fell over at 1.4×. Why?"
Staff Answer
"The usual suspects. One, data: the test ran against a small or warm dataset, so indexes fit in memory. Production doesn't. Two, cache hit rate: synthetic traffic with 1,000 repeated user IDs gives a 99% hit rate, while real traffic across 10M users gives 85%, which is 15× the DB reads. Three, request mix: the test hit the cheap endpoints, but launch traffic is 40% search. Four, dependencies stubbed: the payment provider was mocked. Five, the test client was the bottleneck, or ran from one AZ. Six, no retries: real clients retried under slowness and tripled load. Fixes: replay a sampled production traffic log, use production-scale data (anonymized), include real dependencies or their measured limits, verify the load generator isn't saturated, and model retries. Better still, measure capacity in production by shifting live traffic to a subset of hosts with automatic abort thresholds, the way Facebook's Kraken does."
Why this is L6:
- Lists concrete, common causes, not "the test was wrong"
- Knows cache hit rate is the silent multiplier
- Proposes production-based measurement
What L7 adds:
- Makes production capacity measurement a continuous platform service, not a pre-launch event
- Load-test results feed a published capacity model per service that the capacity review reads
Drill 8: Build vs Buy#
Prompt: "Should we build our own autoscaler?"
Staff Answer
"No. The controllers in Kubernetes and the clouds are good, and the controller isn't where the risk is. What we should build are the pieces around it that are specific to us: the event calendar and pre-scale workflow, the capacity model per service (measured RPS per instance at SLO, limiting dependency), dependency budget enforcement, and the custom metrics adapter that feeds concurrency and lag into the standard controller. Buy the loop, build the inputs. The exception is at very large scale with owned hardware, where you do build capacity planning systems, because provisioning takes months and there's no cloud API to call."
Why this is L6:
- Places the build effort where the differentiation is
- Names concrete components
What L7 adds:
- Evaluates Karpenter-style node provisioning vs. managed options by the cost of the platform team to run them
- Sunsets bespoke per-team scaling scripts as part of standardization
Drill 9: Changing the Policy Without an Outage#
Prompt: "You want to move 150 services from CPU-based to concurrency-based scaling. How?"
Staff Answer
"Shadow first. Run the new signal in recommendation mode, computing desired capacity without acting, for 2 weeks per service. Compare against the actual. Where recommendations diverge by > 30%, investigate before switching. Then roll out by tier: internal services first, then consumer-facing, tier-0 last. During the switch, run both policies with max(cpu_desired, concurrency_desired) so the new signal can only add capacity, never remove it. After 2 clean weeks, drop CPU. Each service needs a measured concurrency target from a load test, and the platform provides a default load-test job to make that cheap."
Why this is L6:
- Shadow → max-of-both → cut over, a safe migration pattern
- Ordered by risk tier
- Makes it cheap for teams to comply
What L7 adds:
- Tracks migration as an org program with a completion metric and an exception process
- Defines the success metric: fewer latency SLO breaches during ramps, and cost flat or down
Drill 10: Multi-Region#
Prompt: "We're going active-active in 3 regions. How does capacity change?"
Staff Answer
"The headroom target is now set by region evacuation, not AZ loss. With 3 regions each serving a third, losing one means the others each take +50%. Without failover headroom, each region must run at ≤ ~66% of its knee at peak. And the failover happens in minutes, faster than autoscaling, so the headroom must be standing. Alternatively, accept a degraded failover: shift traffic, shed non-critical features, and let autoscaling catch up over 10 minutes. That lets regions run at ~75% and costs a brief brownout on failover. Also: follow-the-sun peaks mean regional peaks don't coincide, so global capacity can be lower than 3× regional peak. Data tiers need the failover math too: a region's DB replica becomes primary and takes write load it's never seen."
Why this is L6:
- Recomputes headroom from the new failure scenario
- Offers a cheaper degraded alternative and names the tradeoff
- Remembers the data tier
What L7 adds:
- Prices full vs. degraded failover posture and takes it to leadership as a business choice
- Mandates regular evacuation game days to verify the headroom actually exists
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Incident — Launch Day Brownout#
Context: A highly promoted feature launches at 10:00. At 10:02, p99 latency is 6s and 12% of requests fail. The autoscaler shows desired capacity climbing from 80 to 240. On-call escalates to you.
Questions to Surface First:
- Is the new capacity serving, or stuck in boot/warm-up/health-check failure?
- Which dependency is at its limit: DB connections, third party, cache?
- Is offered load inflated by retries?
- Was this launch on the event calendar, and was it pre-scaled?
Typical L5 Approach: Raise max instances; increase instance size; wait for the autoscaler to catch up.
Staff Approach: Stop the churn: freeze instance replacement, since health-check failures are killing slow-but-alive instances. Enable gateway admission control to cap offered load at current real capacity, shedding low-priority traffic first. Check the DB connection count and cap max instances below the dependency ceiling. Once stable, let warm capacity join with slow-start. Then ask why the launch wasn't pre-scaled.
Principal Approach: Treats it as a planning-process failure. Institutes a launch-readiness gate: any launch with expected traffic > 1.5× baseline requires an event calendar entry, a load test at 2× forecast including dependencies, a pre-scale plan and a named approver. Adds launch readiness to the product review checklist, so marketing can't send the email without an infra sign-off.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Freeze replacement/scale-in. Enable shedding of low-priority routes. Confirm DB and third-party headroom. |
| Triage | Are new instances healthy? Is readiness failing because of cold caches? What's the retry ratio? |
| Quick fix | Raise the floor to the forecast × 1.3. Loosen readiness timeouts. Prime caches. |
| Guardrails | Cap max instances at the DB budget. Watch db.active_connections and p99. |
| Post-mortem | Why no pre-scale? Why did health checks kill busy instances? Why were retries unbounded? |
Metrics to Watch: lb.healthy_host_ratio, autoscaler.desired_vs_inservice, gateway.shed_rate{priority}, db.active_connections, client.retry_ratio, app.p99_latency
Organizational Follow-up: Launch-readiness gate; event calendar integrated with the marketing tool; liveness/readiness separation standard.
Ownership Question: "Who decides which traffic gets shed?" Staff answer: product decides priority tiers in advance (e.g., checkout > browse > recommendations). Engineering executes during the incident. Deciding priorities at 10:02 in an incident channel is too late.
Key Takeaway: "Known events are a process problem. If a launch needs the autoscaler to save it, the planning already failed."
What clears the Staff bar:
- Stops the death spiral before adding capacity
- Uses shedding with pre-agreed priorities
- Traces root cause to the missing pre-scale process
Deep Dive 2: Silent Failure — Headroom That Wasn't There#
Context: A zone outage lasts 40 minutes. The service was "designed for AZ loss" at a 60% target, yet it browned out for the whole outage. Dashboards showed 60% average utilization before the event.
Questions to Surface First:
- 60% of what? Averaged across the fleet, or at the hottest instances?
- Did the remaining zones have instance quota and warm capacity?
- Did traffic distribute evenly, or did one zone take the surge?
- Was the capacity unit (RPS per instance at SLO) still accurate after recent releases?
Typical L5 Approach: Raise the target headroom to 50% and move on.
Staff Approach: Finds the failures in the assumption chain. The 60% was an average, but the hottest AZ ran at 72% because of uneven LB weights. A release 6 weeks earlier made requests 25% more expensive, so the real capacity unit had dropped from 400 to 300 RPS per instance and the actual utilization vs. SLO was ~80%. And the regional vCPU quota prevented scale-out to compensate. Fixes: measure utilization against the current capacity unit, re-measured after each major release; monitor per-AZ peak, not the average; keep quota headroom ≥ 1.5× peak; run an AZ-evacuation game day quarterly.
Principal Approach: Institutionalizes headroom verification: every tier-0/tier-1 service runs a quarterly AZ-drain exercise, removing a zone's capacity in production for 30 minutes. A service that fails loses its tier-0 designation until fixed. Headroom is no longer a number in a config. It's a tested property.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Shed low-priority traffic; request emergency quota; raise floor in surviving zones. |
| Triage | Per-AZ utilization; capacity unit from recent load tests vs. production; quota status. |
| Quick fix | Correct LB weights; raise target instance counts to match the real capacity unit. |
| Guardrails | Alert on per-AZ utilization vs. current capacity unit; alert when quota headroom < 1.5× peak. |
| Post-mortem | Why did the capacity unit drift unnoticed? Why wasn't AZ loss tested? |
Metrics to Watch: capacity.utilization_vs_unit{az}, capacity.unit_rps_measured (trend), cloud.quota_headroom_ratio, lb.weight_skew
Organizational Follow-up: Capacity-unit regression check in the release pipeline (a perf test comparing RPS per core with the previous release); quarterly AZ-drain game days.
Ownership Question: "Who owns the capacity unit number?" Staff answer: the service team. It's their code that changes it. The platform provides the test harness and flags regressions over 10%.
Key Takeaway: "Headroom you haven't tested is a hypothesis. Efficiency regressions quietly erase it between incidents."
What clears the Staff bar:
- Questions the averaging and the capacity unit instead of just raising headroom
- Finds the quota ceiling
- Turns headroom into a tested property
Deep Dive 3: Large Customer Onboarding#
Context: A B2B platform signs a customer that will send 40% of current total traffic, ramping over 3 weeks starting next month. Their traffic is bursty: nightly batch syncs of 5M records in 20 minutes.
Questions to Surface First:
- What's the peak shape: sustained, or 20-minute bursts at a fixed time?
- Which endpoints, and what's their cost relative to the average request?
- Which shared dependencies will they hit hardest?
- Can we shape their traffic contractually (rate limits, batch windows, a bulk API)?
Typical L5 Approach: Increase max instances by 40%; scale the database vertically.
Staff Approach: Models the customer's load explicitly: 5M records ÷ 20 min ≈ 4,200 writes/s at 02:00 UTC, which is 3× the DB's current peak write rate. Rather than sizing everything for that burst, offers a bulk import API that goes through a queue drained at a controlled rate (e.g., 1,500/s over ~55 min), with per-tenant rate limits at the gateway. Pre-scales workers on a schedule for the batch window. Load-tests the bulk path at 2× their stated volume. Ramps their traffic weekly with checkpoints.
Principal Approach: Adds a commercial-capacity interface. Deals above a threshold (e.g., > 10% of traffic) trigger a capacity review before signing, and contracts include rate limits and batch windows. Capacity becomes part of the sales process, and the cost of serving a customer shows up in deal margin.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Model | Requests by endpoint × cost; burst shape; dependency impact (DB writes/s, cache footprint). |
| Shape | Bulk API via queue; per-tenant limits; agreed batch window. |
| Provision | Scheduled worker floor for the window; DB capacity for the drained rate, not the raw burst. |
| Validate | Load test at 2×; shadow their traffic in week 1 at 10%. |
| Ramp | 10% → 30% → 60% → 100% weekly with go/no-go checks. |
Metrics to Watch: tenant.rps{tenant}, bulk_queue.lag_seconds, db.write_iops, gateway.tenant_throttled_total
Organizational Follow-up: Capacity review in the deal desk; tenant-level cost attribution.
Ownership Question: "Who says no if the customer wants unthrottled real-time sync?" Staff answer: the product owner of the API, with the capacity model as evidence. The alternative is pricing that covers dedicated capacity.
Key Takeaway: "Don't size for a customer's burst. Shape the burst, then size for the shaped load."
What clears the Staff bar:
- Quantifies the burst against current capacity
- Changes the demand curve rather than only the supply
- Stages the ramp with checkpoints
Deep Dive 4: Post-Mortem — The Retry Storm That Scaled to the Ceiling#
Context: A 90-second DB failover caused errors. After the DB recovered, the service stayed down for 25 more minutes. The autoscaler had scaled to its max of 400 instances, and offered load was 4× normal.
Questions to Surface First:
- What was the retry policy in clients and in intermediate services?
- Did retries stack across layers (client × gateway × service)?
- Why didn't load fall once the DB recovered?
Typical L5 Approach: Raise the max instances; make the DB failover faster.
Staff Approach: Identifies multiplicative retries: mobile clients retried 3×, the gateway retried 2×, and the service retried its DB calls 3×, so up to 18× amplification on failing paths. After recovery, the backlog of retries plus normal load kept the DB saturated (a metastable failure). Fixes: retry only at one layer, with retry budgets (≤ 10% extra), exponential backoff with full jitter, and circuit breakers that stop retries during failure. Add admission control keyed on DB health so the fleet sheds load quickly and lets the DB recover.
Principal Approach: Publishes an org-wide retry standard: retries at the outermost layer only, budgets enforced by the service mesh, and a required "metastability test" for tier-0 services (inject a 2-minute dependency outage and verify recovery within 2 minutes after it ends).
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Shed 50% at the gateway to break the storm; the DB recovers; ramp admission back up. |
| Triage | Retry counts per layer; amplification factor. |
| Quick fix | Disable gateway retries; cut service retries to 1 with jitter. |
| Guardrails | Retry budget metrics; circuit breakers on DB calls. |
| Post-mortem | Why did autoscaling scale on retry-inflated load? Add a flat-organic-RPS guard (organic = first attempts only). |
Metrics to Watch: requests.attempt_number histogram, retry.budget_exhausted_total, db.cpu, gateway.shed_rate
Organizational Follow-up: Retry standard in the service mesh; metastability test in game days. See Circuit Breakers.
Ownership Question: "Who owns retry policy?" Staff answer: the platform, via the mesh, as a default. Teams can lower it, not raise it, without review.
Key Takeaway: "Autoscaling can't outrun a retry storm. It scales the attack. Retry budgets and shedding end it."
What clears the Staff bar:
- Computes multiplicative retry amplification
- Names metastable failure
- Moves retry policy into a platform standard
Deep Dive 5: Multi-Region Expansion#
Context: The company is adding two regions to its single US-East deployment for latency and resilience. Leadership asks what it will cost and how capacity should be planned.
Questions to Surface First:
- Is the goal latency (serve users nearby) or resilience (survive region loss), or both?
- Must a region failure be invisible (full headroom), or is a degraded failover acceptable?
- Do regional peaks coincide or follow the sun?
- Where does the data live, and can each region's data tier take failover writes?
Typical L5 Approach: Replicate the current fleet 3×. Cost triples.
Staff Approach: Models regional demand curves. With follow-the-sun peaks, the global peak is ~1.4× the largest region's peak rather than 3×. For full region-loss tolerance, each region needs capacity for its own peak plus its share of the failed region's load at that hour. Computes the two postures: (a) invisible failover, where each region runs at ~55% at local peak (~1.8× current total spend), and (b) degraded failover with shedding and a 10-minute autoscale catch-up, where regions run at ~70% (~1.4× spend). Recommends (b) for most services and (a) for tier-0.
Principal Approach: Takes the decision to leadership as a business choice with prices: "Invisible regional failover costs $X/year more than degraded failover. Degraded failover means ~10 minutes of reduced functionality once or twice a year." Pairs it with a reservation strategy: commit ~60–70% of baseline for 1–3 years at a 30–50% discount, and keep the rest on-demand for elasticity.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Model | Hourly demand per region; overlap; failover redistribution matrix. |
| Decide | Failover posture per service tier. |
| Provision | Per-region floors; quota ≥ 1.5× regional peak; data tier sized for failover writes. |
| Validate | Region evacuation game day before GA, then quarterly. |
| Operate | Per-region capacity dashboards; forecast by region. |
Metrics to Watch: capacity.headroom_for_evacuation{region}, region.peak_overlap, db.replica_write_capacity_ratio
Organizational Follow-up: Tier-based failover policy; finance-approved reservation plan.
Ownership Question: "Who decides invisible vs. degraded failover?" Staff answer: the VP-level owner of the product, informed by an engineering cost model. It's a business risk decision.
Key Takeaway: "Multi-region capacity is set by the evacuation scenario, and follow-the-sun traffic makes it much cheaper than N× the current fleet."
What clears the Staff bar:
- Uses non-coincident peaks to cut cost
- Offers two priced postures
- Includes the data tier and quotas
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Compare demand rise time with the actuation delay, and pick schedule, predict, react, headroom or shed accordingly
- Choose a scaling signal per workload class and derive its target with Little's Law
- Compute a ceiling from dependency budgets (DB connections, third-party quotas, partitions)
- Set a utilization target from a failure scenario and the latency knee, and price it
- Describe the death spiral, scaling into a slow dependency, flapping, hidden ceilings and retry storms, each with a metric and an owner
- Design a load test that tells the truth, and explain why production traffic shifting is better
- Describe the org model: platform-owned mechanism, service-owned capacity models, event calendar, quarterly review
The Bar for This Question#
Mid-level (L4): Configures an autoscaling group with CPU targets, min/max and cooldowns. Knows horizontal vs. vertical. Treats the autoscaler as the solution.
Senior (L5): Adds scheduled scaling for known events, read replicas for the DB and a load test. The design is reasonable and handles the diurnal curve. It doesn't quantify the actuation gap, uses CPU without questioning it, and doesn't derive a ceiling.
Staff+ (L6): Frames capacity as pre-paid risk. Compares time constants, picks a demand signal, clamps the controller between a forecast floor and a dependency ceiling, covers the gap with warm pools and shedding, ties headroom to a failure scenario with a price, and puts the owner of every piece in the answer. The interviewer should learn something from the answer, whether that's the flat-RPS guard, the capacity-unit drift after a release, or the retry amplification math.
10. Staff Insiders: Controversial Opinions#
10.1 "CPU-Based Autoscaling Is Wrong for Most Services"#
| Evidence | Implication |
|---|---|
| Most web services are I/O-bound: they wait on databases, caches and other services | CPU stays low while threads saturate |
| CPU rises after latency does | Late signal on a system that's already late |
| GC, noisy neighbors and CPU throttling distort it | Noise that triggers flapping |
The Staff position: concurrency or backlog by default. CPU for proven CPU-bound work.
Why this matters in interviews: saying "CPU at 70%" without justification is the most common L5 tell in this topic.
10.2 "Autoscaling Doesn't Handle Spikes — Headroom Does"#
| Evidence | Implication |
|---|---|
| Actuation delay is 3–5 minutes on VMs | Spikes shorter than that are absorbed by what's already running |
| Launch traffic rises in seconds to minutes | Reactive scaling helps only after the damage |
| Autoscaling's real benefit is scaling in at night | It's a cost tool more than a reliability tool |
The Staff position: autoscaling saves money on slow curves. Headroom, pre-scaling and shedding handle spikes.
Why this matters in interviews: reframing autoscaling as a cost tool shows you understand the time constants.
10.3 "Load Tests Lie; Production Tells the Truth"#
| Evidence | Implication |
|---|---|
| Synthetic tests miss cache hit rates, data size, request mix and retries | Capacity is overstated by 1.5–3× |
| Facebook's Kraken measures capacity with live traffic | Industry precedent for production measurement |
| Capacity drifts with every release | A one-time test goes stale in weeks |
The Staff position: continuous production capacity measurement with automatic abort, plus pre-launch synthetic tests at 2× forecast.
Why this matters in interviews: it shows you've been burned by a passing load test.
10.4 "Scale-to-Zero Is a Dev Environment Feature"#
| Evidence | Implication |
|---|---|
| Cold starts add 100ms–several seconds to the first requests | User-facing tail latency suffers |
| Provisioned concurrency to hide cold starts reintroduces idle cost | The savings shrink |
| The savings matter most where traffic is low, which is where cost matters least | Optimizes the wrong thing |
The Staff position: scale to zero for dev/test, internal tools and truly sporadic jobs. Keep a warm minimum for user-facing paths.
Why this matters in interviews: it shows you weigh the latency cost of a cost optimization.
10.5 "Most Capacity Outages Are Calendar Failures"#
| Evidence | Implication |
|---|---|
| Launches, sales and campaigns are scheduled weeks ahead | They're predictable in principle |
| The outage happens when infra wasn't told | The failure is communication |
| Fixing the autoscaler doesn't fix the calendar | The investment belongs in process |
The Staff position: an event calendar with an owner and a pre-scale approval step is worth more than any controller tuning.
Why this matters in interviews: ending on process instead of config is a clear Staff-level signal.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
A Staff engineer makes one service scale correctly. A Principal engineer sees capacity as the company's largest controllable operating expense and its most correlated reliability risk at once. Every service's headroom, commitment level and scaling policy adds up to a portfolio: a baseline bought with 1–3 year commitments, an elastic layer bought on demand, and a headroom layer that is insurance. The L7 job is to set the rules for that portfolio (utilization targets by tier, commitment coverage, shared-dependency budgets, failure scenarios the company pays to survive) and to make those costs visible so leadership trades them deliberately.
The Org-Level Fault Line#
Central capacity authority vs. team autonomy over spend. Teams that own their budgets move fast but over-provision (fear of pages) or under-provision (pressure on cost). They also ignore shared dependencies. A central capacity group sees the whole picture but becomes a ticket queue and loses workload knowledge.
The Principal position: federated with guardrails. Teams own their capacity models and scaling policies within org-wide targets by tier. A small central capacity team (3–6 people at a 1,000-engineer company) owns forecasting, commitments, shared-dependency budgets, the quarterly review and showback. Exceptions to targets are allowed, but they're visible and priced.
🧭 Principal Move: "I don't need every team to be efficient. I need every team's inefficiency to be visible, intentional and attached to a reason. Showback does most of the work; the review does the rest."
Cost Model#
Assumptions: cloud compute ~$0.04–0.05 per vCPU-hour on-demand; commitments save ~30–50%; engineer fully loaded ~$300K/year; utilization figures are time-weighted averages.
| Scale | Compute spend | Typical waste without discipline | Achievable with Staff design | Headcount for capacity/platform | On-call load |
|---|---|---|---|---|---|
| Startup (20 services) | ~$50K/month | 50–60% idle (fixed peak sizing) | ~35% idle; ~$15K/month saved | 0.5 FTE inside the platform team | Capacity pages ~2/month |
| Growth (200 services) | ~$2M/month | 45–55% idle; 20% on-demand premium | 35–40% idle, 60–70% commitment coverage: ~$500–700K/month saved | 3–5 FTE capacity + autoscaling platform | 1 platform rotation |
| Hyperscale (2,000 services) | ~$40M/month | Headroom duplicated across tiers and regions | Tiered targets + follow-the-sun + commitments: 20–30% saved (~$8–12M/month) | 15–30 FTE: forecasting, capacity eng, efficiency | Dedicated capacity on-call during events |
Reading it: past growth stage, the capacity team pays for itself 10–50× over. The biggest single lever is usually commitment coverage on the baseline, not autoscaling.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversibility Cost |
|---|---|---|
| 3-year compute commitments | One-way | Paid whether used or not. Over-commit by 20% and that's pure waste for 3 years |
| Kafka partition counts / DB shard counts | One-way-ish | Re-partitioning changes key ordering and needs migration |
| Region count and failover posture | One-way for ~1–2 years | Data residency, replication topology and contracts |
| Scaling signal per service | Two-way | Shadow-mode switch in weeks |
| Utilization targets by tier | Two-way | Policy change, then gradual rebalancing |
| Autoscaler vendor/controller | Two-way | Replace the loop, keep the inputs |
| Instance families / hardware generation | Two-way on cloud, one-way on owned hardware | Owned hardware depreciates over 4–5 years |
The Standard I'd Write#
RFC: Capacity & Scaling Standard v1
Scope: All production services on shared compute platforms.
MUST
- Every service publishes a capacity model: RPS (or throughput) per instance at SLO, limiting dependency, max safe instances. Re-measured every 90 days and after any release flagged by the perf regression check.
- Scaling policies use a demand signal appropriate to the workload class (request/response: concurrency or RPS; async: backlog ÷ drain rate). CPU only with a documented justification.
- Max instances ≤ the ceiling derived from dependency budgets.
- Scale-in stabilization ≥ 5 minutes; scale-in step ≤ 20% per window.
- Any planned event > 1.5× baseline has an event calendar entry ≥ 5 business days ahead, a load test at 2× forecast including dependencies, and a named approver.
- Tier-0/1 services pass an AZ-drain exercise each quarter.
SHOULD
- Warm pools or node headroom for services with actuation delay > 2 minutes.
- Retries only at the outermost layer, with budgets ≤ 10%.
Utilization targets (peak, vs. measured knee): Tier-0: ≤ 50%. Tier-1: ≤ 60%. Tier-2: ≤ 70%. Async/batch: ≤ 85%.
Exceptions: Requested through the capacity review; approved by the capacity lead and the service's director; expire after 2 quarters.
Success metrics: zero capacity-caused SEV-1s at planned events; org-wide peak utilization within ±10% of tier targets; commitment coverage 60–75% of baseline; forecast error < 15% at the weekly level.
What I'd Tell the VP#
"We spend about $2M a month on compute, and roughly half of it sits idle at any moment. Some of that is insurance we need, so a data center failure doesn't take us down. A lot of it is just unmanaged. I'm proposing three things: each team sees its own idle cost monthly, we set utilization targets by how critical a service is, and we commit to about two-thirds of our baseline usage for discounts. That should save $500–700K a month within two quarters. Separately, our last two launch-day outages happened because infrastructure wasn't told about the launch. A simple event calendar with a sign-off step fixes that for almost no cost."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Portfolio thinking | "Baseline on commitments, the variable layer on-demand, headroom as priced insurance by tier." |
| Prices the decision | "Invisible region failover costs ~$X/year more than degraded failover. Leadership picks." |
| Makes waste visible | "Showback per team, with idle cost next to the scenario it buys." |
| Governs shared dependencies | "Per-consumer budgets on the shared DB, enforced at the pooler, reviewed quarterly." |
| Tests the posture | "Headroom you haven't drained a zone to prove is a hypothesis." |
Staff answers that L7 interviewers find insufficient:
- "We'll set utilization at 60%." Correct for one service, but no tiering, no pricing and no governance across 200 services.
- "We'll load test before big launches." This relies on heroics per event instead of continuous measurement and a readiness gate.
- "Autoscaling will reduce cost." This ignores that commitments on the baseline usually save more than autoscaling does.
Appendices
Appendix A: Scaling Mechanics in Depth#
A.1 Target Tracking#
every 15s:
observed = avg(metric over last 60s) # e.g., in-flight per instance
desired = ceil(current_ready × observed / target)
desired = max(desired, floor(now)) # forecast / schedule
desired = min(desired, ceiling) # dependency budget
if desired > current: scale_out(min(desired, current × 1.5)) # step cap
if desired < current and stable_for(600s): scale_in(max(desired, current × 0.9))
The asymmetric step caps and stabilization window are the whole safety story. Kubernetes HPA's defaults (15s loop, 300s scale-down window) encode the same idea.
A.2 Little's Law Target#
in_flight_target = rps_per_instance_at_slo × avg_latency_seconds
e.g. 400 RPS × 0.150s = 60 concurrent requests per instance
When latency rises for external reasons, in-flight rises without real demand growth. Hence the flat-RPS guard:
if concurrency_breach and organic_rps_growth_5m < 20%:
suppress scale_out; alert "dependency slowdown suspected"
A.3 Queue Worker Sizing#
arrival_workers = arrival_rate / throughput_per_worker
drain_workers = backlog / (throughput_per_worker × target_drain_s)
desired = ceil(arrival_workers + drain_workers)
desired = min(desired, partition_count, downstream_budget / per_worker_rate)
A.4 Predictive Floor#
floor(t) = seasonal_baseline(t, weeks=4) # same hour-of-week, median
× growth_trend # e.g., 1.02 week-over-week
× event_multiplier(t) # from calendar
× safety_factor # 1.1–1.3
÷ rps_per_instance_at_slo × (1 / target_util)
The floor may only raise capacity. Reactive scaling always operates above it.
Appendix B: Capacity Model Construction#
B.1 Finding the Knee#
B.2 What Goes in the Model#
| Field | Example | Source |
|---|---|---|
rps_per_instance_at_slo | 400 | Load test at the knee × 0.9 |
instance_shape | 4 vCPU / 8 GB | Deployment config |
limiting_dependency | orders-db connections | Load test / architecture |
ceiling_instances | 220 | Dependency budget ÷ per-instance use |
actuation_delay_p95 | 95s (warm pool) | Platform telemetry |
tier | 1 | Service catalog |
last_measured | 2026-08-14 | Must be < 90 days |
B.3 Dependency Ceiling Worksheet#
| Dependency | Budget | Per-instance use | Ceiling |
|---|---|---|---|
| Postgres via PgBouncer | 4,000 client conns | 15 | 266 |
| Payments API | 2,000 RPS | 8 RPS | 250 |
| Redis cluster | 400K ops/s | 1,800 ops/s | 222 |
| Regional vCPU quota | 2,000 vCPU | 4 | 500 |
| Effective ceiling | 222 (Redis) |
Appendix C: Load Shedding & Admission Control#
| Mechanism | Where | Behavior |
|---|---|---|
| Concurrency limit per instance | Service | Reject (503) when in-flight > 1.2× target |
| Priority tiers | Gateway | Shed tier-3 at 90% fleet concurrency, tier-2 at 95%, never tier-0 |
| Adaptive (latency-gradient) limits | Service/mesh | Lower concurrency limit when latency rises (AIMD-style) |
| Per-tenant quotas | Gateway | Isolate noisy tenants (Rate Limiting) |
| Queue with deadline | Async path | Drop work older than its usefulness |
Appendix D: Client Behavior#
- Retries: at the outermost layer only. Exponential backoff (base 100ms, cap 20s) with full jitter. Retry budget ≤ 10% of requests.
- Honor
Retry-Afteron 429/503. - Deadlines propagate. Don't do work whose caller has already given up.
- Idempotency keys on retried writes.
Appendix E: Observability#
E.1 Core Metrics#
capacity.utilization_vs_unit{service,az} # demand ÷ (instances × capacity unit)
capacity.headroom_pct{service} # 1 − utilization at peak
autoscaler.desired / .inservice / .pending # gap = actuation in progress
autoscaler.scale_out_failures{reason} # quota, capacity, image pull
autoscaler.direction_changes_per_hour # flapping
warm_pool.available
gateway.shed_rate{priority}
forecast.error_pct{service}
client.retry_ratio # attempts ÷ organic requests
cost.idle_dollars{team} # showback
E.2 Critical Alerts#
| Alert | Threshold | Severity |
|---|---|---|
| Headroom at peak < scenario requirement | < 33% for a 3-AZ tier-1 service | Ticket (weekly) |
autoscaler.scale_out_failures | > 0 for 5 min | Page |
| Desired − in-service gap | > 20% for 10 min | Page |
gateway.shed_rate{priority<=1} | > 0.1% | Page |
client.retry_ratio | > 1.3 for 5 min | Page |
| Quota headroom | < 1.5× peak | Ticket |
| Capacity unit stale | last_measured > 90 days | Ticket |
Appendix F: Scale Evolution#
| Stage | What works | What you don't build yet |
|---|---|---|
| < 20 services | HPA/ASG with correct signals; manual pre-scale; one load-test script | Predictive floors, showback, capacity team |
| 20–200 services | Event calendar; dependency ceilings; warm pools; showback; commitments | Continuous prod capacity measurement |
| 200+ services | Tiered targets; quarterly review; predictive floors; AZ-drain game days | — |
| Multi-region / hyperscale | Evacuation posture per tier; follow-the-sun planning; Kraken-style measurement | — |
What you don't build on day one: a custom autoscaler, a forecasting ML model, a central capacity team. Do build on day one: the right signal, a ceiling, and a pre-scale habit for launches.
Appendix G: Multi-Tenancy, Fairness & Cost#
- Tenant-driven demand: per-tenant quotas keep one tenant from driving everyone's scale-out. The capacity bill follows the tenant through cost attribution.
- Shared-dependency budgets: each consuming service gets a connection/RPS budget from the dependency owner, enforced at the pooler or mesh.
- Showback: monthly per-team report of
spend,peak utilization,idle dollars, and thescenarioheadroom buys. - Spot/preemptible: fine for stateless batch and fault-tolerant async workers (typically 60–90% cheaper). Never for tier-0 synchronous paths without on-demand fallback.