Hiring BarSupport

Design a CI/CD Deployment System

Case study97 min read10 diagrams

Technologies referenced in this case study: Kubernetes · PostgreSQL · Apache Kafka · ZooKeeper & etcd · Envoy, Kong & NGINX

Related: Feature Flags · Multi-Region Active-Active · Autoscaling & Capacity · Metrics & Alerting Platform · Degraded Mode Framework · Schema Design · Build vs Buy Framework · Content Delivery Network

Reading Guide#

Organized for interview use first, reference second. This page covers getting a change from a merged commit into every production region safely, and getting it back out fast. Pod-level rollout mechanics — readiness probes, maxUnavailable, disruption budgets — live in Kubernetes. The step-by-step schema migration pattern lives in Schema Design. Separating deploy from release with runtime switches lives in Feature Flags. This page links to them rather than repeating them.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (Failure Modes) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal Lens) and the appendices on gate evaluation, rollback safety and the pipeline's own failure posture
What is a Deployment System? — Why interviewers pick this topic

A deployment system takes a change someone has merged — code, configuration, a schema migration, a rules file, a model — builds it into an immutable artifact, proves it as cheaply as possible before production, and then exposes it to production traffic in steps, watching at each step whether it made things worse. When it did, the system undoes the change, ideally before a human has finished reading the page.

The hard part is not running a build or calling kubectl apply. The hard part is that most outages are caused by changes, and the deployment system is the one place where the company decides how much of production a bad change can touch before anyone notices. Google's SRE book puts the share of outages caused by changes to a live system at roughly 70% (Google SRE book). The pipeline is the outage-prevention system that most companies don't call by that name.

Before vs After — the "bad change at 4 p.m." scenario:

Without a staged pipeline:
t=0:       Engineer merges a change to the pricing service. CI green.
t=+8min:   Deploy script rolls all 12 regions in parallel, 25% of pods at a time.
t=+14min:  100% of pods on the new version. A rarely used currency path now throws.
t=+40min:  Error rate for EUR checkouts at 31%. Global average error rate 0.6% → 0.9%:
           under the 1% alert threshold. Nothing pages.
t=+3h:     Finance notices EUR revenue down. Incident declared.
t=+3h20m:  "What changed?" — 41 deploys in the last 4 hours across the company.
t=+3h50m:  Culprit found. Rollback runs. Old version can't read rows the new version
           wrote with a new enum value. Second outage.

With waves, evidence-based gates and pre-verified rollback:
t=0:       Change merged. Pipeline runs an N-1 compatibility test: old code reads
           new-format rows. Pass.
t=+25min:  One-box in the first wave's region (≤ 10% of that cell's traffic).
t=+40min:  Gate holds: 'POST /checkout currency=EUR' has seen 38 requests, needs 500.
t=+2h10m:  500 EUR requests seen. Canary error rate on that endpoint 30% vs 0.2% baseline.
t=+2h11m:  Automatic rollback of the one box. Change blocked from promotion.
           Owner notified with the failing endpoint and diff. Blast radius: one box,
           one cell, ~150 failed checkouts.

Why interviewers reach for this question: Deployment is the purest test of blast-radius and time-to-undo thinking. Every company has a pipeline; few have one where the answer to "what's the worst a bad change can do before we notice?" is a number. The candidate who answers "blue-green with automated tests" has described a mechanism. The candidate who answers "a bad change reaches at most one cell in one low-traffic region, is detected by a gate that waits for evidence rather than minutes, and is rolled back automatically in under 10 minutes — and here's how I guarantee the rollback itself is safe" has run one.

Mechanics Refresher: Deployment Strategies
StrategyHow It WorksProsCons
Big bang / recreateStop old version everywhere, start newSimple; no mixed versionsDowntime; 100% blast radius instantly
Rolling updateReplace instances in batches (e.g. 25–33% at a time)No extra capacity; standard in orchestratorsWhole fleet converts in minutes; the batch is not a canary unless you pause and judge
Blue-greenStand up a full new fleet, switch the load balancer, keep old fleet warmInstant cutover and instant rollback2× capacity during deploy; the switch moves 100% at once; shared databases aren't switched
Canary (one-box)Send a small slice of traffic to the new version, compare against a baselineReal traffic, small blast radiusNeeds enough traffic on the changed path to judge; needs per-version metrics
Waves / staged regional rolloutPromote cell → AZ → region → groups of regions, with bake time betweenBounds blast radius at every step; catches slow-burn faultsA change takes hours to days to reach everywhere
Dark launch / shadow trafficMirror requests to the new version, discard its responsesZero user impact; catches crashes and latencySide-effecting calls must be stubbed; doubles load for mirrored paths
Feature flags (deploy ≠ release)Ship code dark, turn behavior on by percentage or cohortRelease is a runtime switch, reversible in secondsFlag debt; flag config is itself a production change — see Feature Flags

For most production systems: an immutable artifact promoted through pre-production, then one-box canaries and waves of increasing size across cells and regions, with gates that require evidence (traffic counts on the changed paths, canary-vs-baseline comparison) and automatic rollback on alarm. Behavior changes ride behind flags. The strategy names are not the interview — blast radius per step, detection time, and whether rollback is safe are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

A deployment system is not a speed question. Nobody is impressed that your pipeline deploys in eight minutes.

It is a blast-radius and time-to-undo question that tests:

  • Whether you can state, as a number, how much of production a bad change can reach before it is detected
  • Whether your gates wait for evidence — requests on the changed code path — or just for a clock
  • Whether you know that a rollback is itself a deploy, and can fail, usually because of data written by the new version
  • Whether you treat configuration, rules and content pushes as deploys, because the worst public outages were not code
  • Whether the pipeline has a failure posture of its own — what happens when the thing that ships fixes is the thing that's broken

The key insight: The damage a bad change does is roughly exposure fraction × (time to detect + time to undo). Speed of deployment appears nowhere in that product. Staff candidates design every stage to shrink one of the three terms — smaller first steps, gates that detect sooner on the right signals, rollbacks that are pre-verified and fast — and treat deploy speed as the budget they spend to buy those, not the goal.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws repo → CI → artifact → deploy to Kubernetes, adds "blue-green"Asks "What kinds of change ship through this — code, config, schema, content? How many regions and cells? What's the worst acceptable blast radius for a bad change, and who signs that?"Asks "What fraction of our incidents start with a change, which change types bypass the pipeline today, and who owns the pipeline as a product with an SLO?"
Rollout"Rolling update, 25% at a time, health checks""One-box in one cell of a low-traffic region, then a high-traffic region one AZ at a time, then parallel waves. Bake gates require ≥ N requests on the changed endpoints, not just 60 minutes"Sets the wave plan, bake minimums and deployment windows as org defaults that pipelines inherit; teams can be stricter, never looser, without an exception
Detection"If health checks fail, stop the rollout""Canary vs baseline per endpoint, per version label; global averages hide a 30% failure on one path. Synthetic canaries cover low-traffic paths"Makes change failure rate and failed-deploy recovery time org-level metrics reviewed monthly, broken down by team and by change type
Rollback"Redeploy the previous version""Rollback is automatic on alarm and pre-verified: an N-1 compatibility stage proves old code can read data the new code writes. Schema changes ship expand → migrate → contract"Declares 'every change must be reversible or explicitly marked one-way with a second reviewer' as a standard; audits one-way changes quarterly
Config & contentOut of scope — "config is in a separate system""Config, rules and content go through the same wave machinery, faster bake, never global-at-once except a signed emergency path"Inventories every channel that can change production behavior — flags, rules, models, DNS, IAM — and puts each on staged rails or on a list with a named owner
Pipeline failureNot considered"The pipeline is tier 0 during incidents. A break-glass path that doesn't depend on CI, the artifact cache or the orchestrator's primary region exists and is drilled"Prices pipeline availability as part of every service's recovery time; staffs the platform team with an on-call and an error budget
Why "rollout" separates levels

L5: "We'll do a rolling deployment, 25% of pods at a time, with readiness probes, and if the probes fail the rollout stops." Correct and incomplete. Readiness probes catch "the process doesn't start". They don't catch "the process starts and returns wrong prices for EUR". A rolling update with no pause converts the whole fleet in minutes, so the 25% batch is a pacing mechanism, not a canary.

L6: "I'll separate pacing from judgment. Pacing is the rolling batch inside a cell. Judgment happens at stage boundaries: after the one-box, after the first region, after each wave. The gate at each boundary asks two questions — did any alarm fire, and have we seen enough traffic on the code paths this change touched to have an opinion? The second is the one people forget. A one-box at 3 a.m. in a small region can see zero requests to the endpoint that's broken and still pass a 60-minute bake."

L7: "The wave plan is an org-wide decision about how much correlated risk we accept, so I'd publish it as the default every pipeline inherits: which regions go first, minimum bake, minimum evidence, deployment windows. Teams can tighten it. Loosening it needs a named reviewer. That's how one team's aggressive pipeline stops being everyone's outage."

Why "rollback" separates levels

L5: "If something goes wrong, we redeploy the previous image. It's immutable, so rollback is safe." The artifact is immutable; the data isn't. The new version may have written rows, messages or cache entries in a format the old version can't read, and the rollback becomes the second outage.

L6: "Rollback safety is a property I test before production, not one I hope for. The pipeline runs an N-1 stage: old binary, new data. Any change to a persisted format ships in two deploys — first teach every reader the new format, then start writing it. Schema changes go expand → migrate → contract, and only the contract step is one-way. A change that can't pass N-1 is flagged one-way and needs a second reviewer and a roll-forward plan."

L7: "I'd make 'is this reversible?' a field in every change record, with the pipeline enforcing it. Then I can report how many one-way changes we shipped last quarter, which teams, and how many of those caused incidents — and fund tooling where it clusters."

Why "config & content" separates levels

L5: "Config is in a key-value store; changes take effect immediately." That sentence describes the fastest, least-guarded path to every server in the company.

L6: "A config change, a WAF rule, a detection-content file and a code change are all production changes. They get different bake times — config might bake for minutes, code for hours — but the same shape: validate, stage to a small slice, judge, widen. The emergency global push exists, but it's a separate, audited path."

L7: "I'd inventory every channel that can change production behavior without a code deploy and ask who owns its blast radius. In most companies the answer for at least one channel is 'nobody', and that's the next big outage."

Positions to Commit To#

PositionRationale
Every production change goes through staged rails — code, config, schema, rules, contentChanges cause most outages; a channel that bypasses staging is where the next global outage comes from
Waves sized by blast radius: one-box → one cell → one region → groups of regionsBounds the worst case at each step; the first production exposure is ≤ 10% of one cell in a low-traffic region
Gates require evidence, not just timeA 60-minute bake with zero requests on the changed path proves nothing; gates count requests per touched endpoint
Automatic rollback by default; humans approve going forward, not going backRolling back a healthy change costs minutes; waiting for a human to approve rollback of a bad one costs the incident
Rollback is pre-verified, or the change is marked one-wayImmutable artifacts don't make rollbacks safe; data written by the new version does or doesn't
The pipeline is a tier-0 system with a break-glass path that doesn't depend on itThe worst time for the pipeline to be down is the hour you need to ship the fix
Freeze is a risk budget, not a calendarBlanket freezes batch up risk for the thaw; restrict by risk class, keep fixes and rollbacks flowing

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Service fleet deploys100s–1,000s of services, 10–30 regions, many cells; you control the serversImmutable artifacts, one-box canaries, regional waves with evidence-based bake, automatic rollbackSlow-burn faults that pass bake; rollback blocked by data format; correlated multi-region pushChange failure rate, failed-deploy recovery time ≤ 30 min, zero multi-region incidents from a single change
Config, rules and content pushesChanges must reach the fleet in minutes (security rules, kill switches, pricing tables)Schema-validated payloads, staged by cell with short bake, client-side guards (bounds checks, last-known-good fallback), signed emergency global pathA malformed payload crashes every consumer at once because it went everywhere in secondsNo config change reaches > 1 cell before a health signal; every consumer survives a malformed payload
Client and edge softwareCode runs on devices you don't control — mobile apps, agents, firmware; rollback is slow or impossibleStaged percentage rollouts, internal rings first, crash-rate gates, remote kill switches, compatibility windows for old clientsA crash-on-start bug you can't remotely roll back; months of old versions in the fieldCrash-free sessions by version; time to halt a rollout ≤ 15 min; every client can be told to stop using a feature

🎯 Staff Move: "I'll design the service deployment pipeline — that's where the scale and the multi-region risk are. But I'll put config and content pushes on the same wave machinery with shorter bake times, because the biggest deployment outages in public post-mortems were not code deploys. Client software I'll treat as a separate problem with a shared gate engine. Tell me if you'd rather go deep on one of the others."

Where the Design Splits#

#Fault LineThe Tension
1Deploy Speed vs Bake TimeEvery hour of bake catches slow-burn faults and delays every fix, feature and security patch by an hour; batch size grows with pipeline duration
2Region-by-Region Waves vs Global PushWaves bound blast radius but leave regions on different versions for days; a global push is consistent and fast and makes every bad change a global outage
3Automatic vs Human-Approved RollbackAutomation reacts in minutes and sometimes rolls back healthy changes; humans add judgment and 15–40 minutes of page-to-decision time
4One Central Pipeline vs Per-Team PipelinesOne paved road gives consistent safety and becomes a bottleneck and a single point of failure; per-team pipelines are fast and inconsistent
5Freeze Policy: What Ships During an Incident or a PeakFreezing everything reduces change risk and blocks the fix; allowing everything adds risk to an already burning system

How Real Companies Built It#

Why this section belongs here: Deployment safety is unusually well documented, because the companies that got it wrong published post-mortems and the companies that got it right published their pipelines. Each entry shows a fault line from this page in production.

Amazon — Waves, One-Box Stages and Evidence-Based Bake#

Amazon's Builders' Library describes production pipelines split into waves: the first wave deploys to a single low-traffic Region one Availability Zone or cell at a time, the second to a high-traffic Region, and later waves to growing groups of Regions in parallel. Each wave starts with a one-box stage that typically serves at most 10% of the Region or AZ's requests; a typical pipeline then waits at least 1 hour after each one-box stage, at least 12 hours after the first regional wave and 2–4 hours after later waves, and the bake includes waiting for a number of data points (the article's example is at least 100 requests to a Create API). Deployments roll back automatically on alarm, avoid nights, weekends and holidays, and by default reach all Regions in about four or five business days (Amazon Builders' Library). A companion article describes the two-phase "prepare, then activate" technique for changing data formats so a rollback stays safe (Amazon Builders' Library).

Staff insight: Two numbers in that article do most of the work in an interview: "at most 10% of one AZ" as the first production exposure, and "wait for at least N requests" as the bake condition. Say both. The four-to-five-day global rollout is the price of fault line 1, paid deliberately — and the article also describes an expedited path for urgent fixes that still goes through every stage, with more review rather than fewer steps.

CrowdStrike — Channel File 291 and Unstaged Content (2024)#

CrowdStrike's root-cause analysis says that a sensor Template Type introduced in February 2024 defined 21 input fields while the integration code supplied only 20, and that the mismatch passed testing because early content used wildcard matching for the 21st field. On 19 July 2024 two new Template Instances were deployed as Rapid Response Content; one used a non-wildcard criterion for the 21st input, the Content Validator — which had a logic error — passed it, and sensors that received the new Channel File 291 performed an out-of-bounds memory read and crashed. Among the findings: each Template Instance should be deployed in a staged rollout through canary and successively wider rings, and customers should get control over when Rapid Response Content is deployed (CrowdStrike root cause analysis).

Staff insight: This is intent 2 — content, not code — on the client-software side, where rollback means touching each device. The lessons map directly onto this design: content is a deploy and needs rings; the validator must test against what the consumer actually does, not the schema's declaration; and the consumer must survive a malformed payload (bounds checks, last-known-good). In an interview, say "content that runs on a million machines gets a canary ring, no matter how urgent it is."

Cloudflare — A WAF Rule Deployed Globally (2019)#

Cloudflare's post-mortem of 2 July 2019 describes a new WAF managed rule containing a regular expression that backtracked enormously, pushing CPUs serving HTTP traffic to nearly 100% across its network. WAF rules were deployed through its distributed key-value store and reached the whole fleet within seconds, unlike its staged software releases; the rule went out at 13:42 UTC and the WAF was globally disabled at 14:07 UTC. Afterwards Cloudflare changed its procedure to stage rule rollouts the way it stages other software while keeping an emergency global path for active attacks (Cloudflare post-mortem).

Staff insight: Fault line 2 in one incident. The software pipeline was staged; the rules channel wasn't, because rules "had to be fast". The fix wasn't to make rules slow — it was to make them staged by default with an emergency lane that is explicit, audited and rare. See Content Delivery Network for the edge side.

Meta — Quasi-Continuous Release With Employee and Percentage Tiers#

Meta's engineering blog describes moving the Facebook web front end from three pushes a day — plus 500 to 700 cherry-picks a day — to a quasi-continuous system that pushes tens to hundreds of diffs every few hours, first to employees, then to 2% of production, then to 100%. By then the master branch was taking more than 1,000 diffs a day; Gatekeeper lets code releases and feature launches roll out independently, and an emergency stop button can halt a release from going further (Meta engineering).

Staff insight: At very high change volume, the unit of deployment becomes a batch of diffs, and the employee tier is a canary with unusually good error reporting. The decoupling of deploy from release (Gatekeeper) is what makes frequent deploys survivable: most behavior changes are flag flips, not binary changes. Cite it when an interviewer asks how deploy frequency and safety can both go up.

Knight Capital — One Server Out of Eight (2012)#

The SEC's order describes Knight Capital deploying new order-routing code to its SMARS servers, during which a technician did not copy the new code to one of eight servers, and no second technician reviewed the deployment; Knight had no written procedures requiring such a review. The new code repurposed a flag that, on the eighth server, triggered old "Power Peg" code still present there. Over about 45 minutes the system sent millions of orders, obtaining over 4 million executions, and Knight lost over $460 million (SEC order).

Staff insight: Two deployment lessons. First, the system must verify the fleet converged to the intended version — "deploy succeeded" must mean every target reports the expected artifact digest. Second, reusing a flag or field that old code interprets differently is a rollback-safety and mixed-fleet hazard; mixed versions exist during every rollout, so the change has to be safe for them.

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"Canary for an hour, then roll out""The canary ran 2–3 a.m. in your smallest region. The broken endpoint got 4 requests. Did it pass?"Evidence-based gates vs clock-based gates
"We roll back automatically""The new version added a column value the old version can't parse. Rollback starts. What happens?"Rollback safety, N-1 testing, expand/contract
"Waves, region by region""A security fix needs to be everywhere in an hour. Your pipeline takes four days."Expedited path design without bypassing safety
"Config is stored in etcd and pushed live""Someone pushes a malformed config. How many servers have it after 10 seconds?"Config as a deploy; staged config rollout
"The platform team owns CI/CD""CI is down during a SEV-1. The fix is merged. How does it get to production?"Pipeline as tier 0, break-glass path
"We freeze deploys during the holidays""A memory leak needs a fix on December 23. And what lands on January 3?"Freeze as risk budget, thaw risk

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Three loops at three speeds. The build loop (merge → artifact, ~12 minutes) runs thousands of times a day and is a throughput problem. The promotion loop (artifact → every cell, 1–5 days depending on risk class) is a safety problem: the orchestrator moves a release one stage at a time, and the gate evaluator decides each step from metrics labeled by version. The undo loop (alarm → previous version, target ≤ 10 minutes per cell) is the one the whole design is judged by. The change log turns "what changed?" during an incident from a 30-minute search into a query. The break-glass path deliberately bypasses CI and the orchestrator's primary region.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Rollout shape"Rolling update with health checks""One-box ≤ 10% of one cell, low-traffic region first, then high-traffic region, then parallel waves; judge at every stage boundary."
Bake"Wait an hour""Wait for evidence: ≥ N requests on every endpoint the diff touched, plus a minimum wall-clock floor for slow burns."
Detection"Alert on error rate""Canary vs baseline per endpoint and version label; synthetic probes for low-traffic paths; high-severity alarm auto-rolls back."
Rollback"Redeploy the previous image""Automatic, ≤ 10 min per cell, previous version pre-pulled. Safety pre-verified by an N-1 stage; persisted-format changes ship in two deploys."
Config"Pushed live from a KV store""Same wave machinery, shorter bake, consumers keep last-known-good; global push only on a signed emergency path."
Urgent fix"Skip the pipeline""Expedited risk class: same stages, shorter bake, second reviewer; break-glass only if the pipeline itself is down."
Freeze"No deploys in December""Restrict by risk class; fixes and rollbacks always flow; cap the thaw batch."

Numbers to Bring#

MetricValueWhy It Matters
Share of outages from changes to a live systemRoughly 70% (Google SRE book)Why the pipeline is an outage-prevention system
First production exposureOne box, ≤ 10% of one AZ's or cell's traffic (Amazon Builders' Library)Bounds the worst case of the first step
Typical Amazon bake≥ 1 h after one-box, ≥ 12 h after first regional wave, 2–4 h after later waves (same source)Early waves bake longest; later waves move faster
Amazon default time to all RegionsAbout 4–5 business days (same source)The price of fault line 1, paid on purpose
Rolling batch inside a cell≤ 33% of boxes per batch, ≥ 66% capacity kept (same source)Pacing, not judgment
Requests needed to judge a 2× error-rate regression at 0.1% baseline~20K per arm (~20 vs ~40 errors) — rough Poisson estimateA 1% canary on a 200 rps service needs ~3 hours, not 1
Rollback target per cell≤ 10 min, alarm to previous version servingUndo time is half the damage formula
Page-to-human-decision time15–40 min at nightWhy rollback should not wait for a human
CI build p50 / p9512 min / 25 min (design target)Over ~30 min, engineers batch changes and bypass
Change-log lookup during an incident< 1 min for "what changed in this region in the last 6 h"Replaces a 20–40 minute hunt across tools
Break-glass artifact cacheLast 30 artifacts per tier-0 serviceThe fix or the rollback target must be deployable with CI down
Kayenta judge confidenceMann-Whitney U test, 98% confidence before flagging a metric (Spinnaker docs)Statistical canary judgment is a solved, open-source problem

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on the CI half — build caching, test parallelism, Docker layers, the Git branching model — and never reach the question the interviewer cares about: "A bad change is merged at 4 p.m. How much of production does it touch, how do you find out, and how fast is it gone?" Compress CI to ~5 minutes. Spend the rest on promotion, gates, rollback safety and the pipeline's own failure posture.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Engineers merge changes; the system builds immutable artifacts, tests them, promotes them through pre-production and then through production in stages, decides at each stage whether to continue, and rolls back automatically when a change makes things worse. It also records every production change so incident responders can answer 'what changed?' in seconds."

Then the non-functional requirements, which is where the design lives:

"Four constraints drive everything. One: blast radius — a bad change must reach no more than one cell in one region before a gate has an opinion about it. Two: time to undo — automatic rollback within 10 minutes of an alarm, and the rollback itself must be safe, which means data written by the new version can't break the old one. Three: every channel that changes production behavior — code, config, schema, rules — goes through staged rails; nothing is global-at-once by default. Four: the pipeline is tier 0 during incidents, so there's a path to ship a fix when CI is down. Scale: I'll assume 1,500 engineers, ~800 services, 12 regions with 3 cells each, ~2,500 merges a day and ~600 releases a day reaching production."

Then name the underspecified parts:

"A few things I'd confirm: monorepo or many repos? Are there client apps or agents on customer devices, or only servers? Is there a regulated environment that needs change approval records? What's the current change failure rate — is this a greenfield pipeline or a fix for a company that keeps breaking itself? I'll assume servers only, polyrepo with a shared pipeline definition, and some services that handle payments."

🎯 Staff Move: Putting a number on blast radius in the first three minutes — "one cell, one region, ≤ 10% of that cell's traffic for the first step" — tells the interviewer that you see the pipeline as a safety system. Everything you draw next can be judged against that number.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Change: change_id, repo, commit_sha, author, reviewers, type (code, config, schema, rules, flag), reversible (true, one_way), risk_class (standard, low, expedited, emergency)
  • Artifact: artifact_digest (content hash), build_id, provenance (source sha, builder identity, signature), created_at — never mutated, never re-tagged
  • Release: release_id, service, artifact_digest, config_version, schema_version, changes[], pipeline_id
  • Stage execution: release_id, stage (preprod, n_minus_1, one_box, wave_1 … wave_5), target (region, cell), status (pending, deploying, baking, passed, held, rolled_back), started_at, gate_evidence
  • Gate evaluation: stage_execution_id, alarms_checked[], endpoints_required[], requests_seen{}, canary_score, decision, decided_by (system or person)
  • Change event: the immutable log entry for every production change, from any source

API (what teams and the orchestrator call):

POST /v1/releases                      { service, artifact_digest, config_version, changes[] }
  → 201 { release_id, pipeline_plan: [stages...], risk_class }

GET  /v1/releases/{id}                 → stage-by-stage status, gate evidence, current holds
POST /v1/releases/{id}/halt            { reason }          (anyone on the owning team)
POST /v1/releases/{id}/rollback        { scope: cell | region | all }
POST /v1/releases/{id}/expedite        { justification, second_reviewer }
POST /v1/freeze                        { scope, risk_classes_blocked, until, owner }
GET  /v1/changes?region=eu-west-1&since=6h   (incident responders: what changed here?)

Release creation returns the stage plan the pipeline will follow; teams see their pipeline as data, not as a script.

🎯 Staff Move: "I'm separating the change, the artifact and the release. The artifact is what we built, immutable by digest. The release is the artifact plus the config and schema versions it expects. The change carries the reversibility flag. When something goes wrong, the question 'can we roll this back?' is answered by data on the release, not by someone reading a diff at 3 a.m."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk the flow in 90 seconds:

  1. A merge triggers a hermetic build; unit and contract tests run; the output is an artifact identified by digest and signed with provenance.
  2. The orchestrator creates a release and its stage plan from the service's pipeline definition, which inherits the org's default wave plan.
  3. Pre-production: deploy to an integration environment; run integration tests against real dependencies. Then the N-1 stage: run the old binary against data written by the new one.
  4. Production wave 1: one-box in one cell of a low-traffic region. The gate waits for a minimum wall-clock bake and minimum request counts on the endpoints the diff touched, comparing the one-box with baseline boxes.
  5. On pass, the cell rolls 33% at a time, then the next cell, then wave 2 in a high-traffic region, then waves 3–5 widening across regions.
  6. At any stage, a high-severity alarm triggers automatic rollback of that stage's scope and blocks promotion. The owner is notified with evidence.
  7. Every stage transition writes a change event to the log.

🎯 Staff Move: Say out loud: "Notice the three loops: build at thousands a day is throughput; promotion over hours to days is safety; undo in under 10 minutes is the one I'll be judged on. I'm going to spend the time on the second and third." You've now spent ~9 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path, and it's the Senior-level design. What makes deployment hard is the bad change that looks fine. I'd like to go deep on four things: how a gate decides it has seen enough evidence, why rollback can fail and how to make it safe, how config and content pushes ride the same rails, and what happens when the pipeline itself is down during an incident. Where would you like to start?"

If no preference: start with gates. It's where the bad change that "passed canary" lives.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: Gates that wait for evidence (7–8 min)

"A bake timer answers 'has enough time passed?'. The question that matters is 'has the new code done enough work for us to judge it?'. So each gate has three conditions: no high-severity alarm in the stage's scope; a minimum wall-clock floor — 60 minutes for the one-box, 12 hours for the first region — to catch slow burns like leaks; and minimum evidence — for each endpoint and message type the diff touched, at least N requests served by the new version. Then canary-versus-baseline comparison on those endpoints, with a statistical test rather than a fixed threshold."

Quantify: "Say an endpoint's baseline error rate is 0.1% and I want to catch a doubling. That's roughly 20 errors versus 40 — about 20,000 requests per side. A one-box at 10% of a cell serving 200 requests a second on that endpoint gets 20 a second: about 17 minutes. An endpoint at 2 requests a second needs 3 hours. That's why bake time is a per-endpoint calculation, and why low-traffic endpoints need synthetic traffic or a longer gate."

Who pays: "Engineers pay in latency — a change to a rare endpoint takes longer to promote. The platform team pays for the per-endpoint, per-version metrics, which roughly doubles metric cardinality for services in mid-deploy. I'd rather pay that than ship a bad change that passed on zero traffic."


Deep dive 2: Rollback safety (7–8 min)

"The artifact is immutable; the world it wrote into isn't. Rollback breaks in three ways: the new version wrote data the old one can't read — new enum values, new message fields, a new cache format; the new version ran a destructive migration; or the new version changed an external contract that clients already adopted. So: every persisted-format change ships in two deploys — first teach every reader the new format, then start writing it — and the pipeline's N-1 stage runs the previous binary against data written by the new one. Schema changes go expand → migrate → contract, as in Schema Design, and the contract step is its own release, flagged one-way."

Quantify: "Rollback time target is 10 minutes per cell: previous image pre-pulled on every node, rollout at 33% batches, readiness gates on. For a 300-pod cell that's three batches at ~2 minutes each plus detection. If rollback isn't safe, the alternative is roll-forward, which is a full pipeline run — even expedited, that's hours, not minutes."


Deep dive 3: Config and content on the same rails (5–6 min)

"Config changes are the most common cause of fast global outages because they're built to propagate in seconds. I'd keep the propagation mechanism and change the default: a config change is a release with its own stage plan — one cell, then one region, then widening — with bake measured in minutes, not hours, because the config is small and its effects show fast. Every consumer validates the payload against a schema and keeps last-known-good, so a malformed payload is rejected locally rather than crashing the process. The global push still exists for emergencies: signed, two-person, logged, and reviewed afterwards."

Who pays: "The security team pays a few minutes of propagation on a new rule. I'd agree with them on a number — say 15 minutes to 100% for standard rules, 2 minutes on the emergency lane — and that's a signed policy."


Deep dive 4: The pipeline is down during an incident (5 min)

"The worst time for the pipeline to fail is the hour you need it. Its dependencies — source control, CI runners, the artifact registry, the orchestrator — are probably in one region, and maybe affected by the same incident. So the pipeline gets a degraded mode: rollbacks don't need CI, because the previous artifact is already cached on nodes and in a regional registry mirror; a break-glass path can deploy one of the last 30 artifacts for a tier-0 service with two-person approval from an out-of-band runner; and the orchestrator's state is replicated so a region loss doesn't lose in-flight releases. We drill it every quarter by turning CI off."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: the damage of a bad change is exposure times detection time plus undo time, so I shrank each term — the first exposure is one box in one cell, gates wait for evidence on the changed paths, and rollback is automatic and pre-verified. Config and content ride the same rails with shorter bake. The pipeline has its own degraded mode because it's what ships the fix."

The evolution closer:

"What I'd build later: automatic per-endpoint evidence requirements derived from the diff and the call graph; progressive delivery tied to Feature Flags so most behavior changes are flag ramps, not deploys; cell-based architecture so a 'cell' really is a blast-radius boundary; and change-failure-rate dashboards by team. What I'd not build: a custom CI system — I'd buy or adopt one and own the promotion layer, which is where the safety lives."

🎯 Staff Move: End on the damage formula and who owns each term. Senior candidates end with "and we'd add more tests." Staff candidates end with "the platform owns the wave plan and the gate engine, each service owns its alarms and evidence requirements, and change failure rate is a number we review monthly — because the next outage is most likely a change."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
CI deep diveSpends 10 minutes on build caching and test shardingOne sentence: "Hermetic builds, remote cache, p50 12 min — I'll come back if you want"
Strategy catalogueLists blue-green, canary, rolling, A/B as a menuPicks waves + one-box + evidence gates and says why
Branching model debateTrunk-based vs GitFlow for 5 minutes"Trunk-based, short-lived branches, flags for unfinished work"
No blast-radius numberWaits for "what if the canary misses it?"States the first-step exposure in Phase 1
Rollback as an afterthought"And we can always roll back" at minute 40Brings rollback safety into the deep dives with N-1 testing
Ignores configTreats config as someone else's systemPuts config and content on the same rails, unprompted

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Deployment is the dependency every change has. A database design affects one system; the deployment system decides how every change to every system meets production. When it works, nobody notices — which is why its quality is invisible to Senior engineers and obvious to anyone who has run incident review for a year and watched "a deploy" appear in the root cause of most of them.

It also has a deceptive happy path. A Senior engineer can build a pipeline that ships good changes quickly on day one. The design is judged entirely by the bad changes: the one that only breaks a rare code path, the one that leaks memory for 18 hours before it falls over, the one that writes data the old version can't read, the config push that reaches every server in four seconds. None show up in a demo. Staff candidates design for them before drawing the first box.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"A change that breaks one code path used by 0.5% of requests is merged at 4 p.m. on a Thursday. Walk me through every stage it reaches, what each gate sees, when it's detected, who is paged, and what happens to the data it wrote when you roll it back."

A candidate who answers with the first exposure (one box, one cell), the evidence the gate requires on that path (and how long a low-traffic path takes to accumulate it), the comparison method (canary vs baseline, per endpoint), the rollback mechanism (automatic, pre-pulled, ≤ 10 minutes), the data question (N-1 stage, two-phase format changes) and the owner (service on-call for the alarm, platform for the gate engine) has operated a deployment system. A candidate who says "the canary would catch it" has built a pipeline that ships good changes.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Service fleet deploys. You control the machines. Changes are binaries or containers deployed to cells across regions. The design centers on staged exposure, version-labeled metrics, automatic rollback and the data-compatibility discipline that makes rollback safe. Lead time from merge to all regions is a business tradeoff, typically one to five days depending on risk class. Correctness bar: change failure rate, failed-deploy recovery time, and the absence of multi-region incidents caused by a single change. This is where most interviews go, and the rollout mechanics at the pod level are in Kubernetes.

Config, rules and content pushes. Feature configuration, rate-limit tables, WAF rules, fraud models, detection content, routing weights. These changes are designed to be fast — that's their reason for existing — and usually skip the code pipeline. They are also data that every server parses, which means a malformed payload is a fleet-wide crash with no build step in between. The design centers on schema validation against the consumer's real behavior, cell-staged propagation with short bake, consumer-side last-known-good, and an emergency lane that is explicit. Correctness bar: no payload reaches more than one cell before a health signal; every consumer survives a malformed payload.

Client and edge software. Mobile apps, desktop agents, kernel modules, firmware, browser extensions. You don't control the machines, the user decides when to update, and rollback is either slow (ship a new version and wait for adoption) or impossible (the device no longer boots). The design centers on rings (employees, beta, 1%, 10%, 100%), crash-rate gates by version, remote kill switches for every feature, and a long compatibility window because old versions live for months. Correctness bar: crash-free rate by version; time from "stop" to rollout halted; percentage of features that can be disabled remotely.

🎯 Staff Move: "These share a gate engine and a change log, and almost nothing else. Fleet deploys optimize undo time. Config pushes optimize propagation safety. Client software optimizes for the fact that undo may not exist. If you ask me to build one pipeline for all three, I'll share the evaluation and audit layers and keep the delivery mechanics separate."

2.2 When NOT to Build a Staged Deployment Platform#

SituationWhat to Do InsteadWhy
Fewer than ~30 engineers, one region, a handful of servicesHosted CI + the orchestrator's built-in rolling update + one canary step + a manual promote buttonWaves across regions you don't have are ceremony; a staged platform is 2–4 engineers you can't spare
Stateless internal tools with no external usersPlain rolling deploy with readiness probesBlast radius is small and the users are your colleagues; rollback cost is low
A batch or data pipeline jobVersion the job, run the new version on a sample or shadow output, compare, then switchStaged traffic doesn't apply; the "canary" is a diff of outputs
You are on a managed platform with progressive delivery built inUse it; own the gate definitions and alarms, not the machineryThe value is in what you measure, not in the rollout controller
A single-tenant appliance shipped to customers quarterlyRelease trains with long soak in staging and customer-controlled upgradesNo production traffic to canary against; the customer owns the window

And within the design, things you should not build even at scale:

  • Don't build your own CI execution engine. Builds are a solved commodity; hermeticity and caching are configuration problems. Own the promotion layer.
  • Don't build per-team deployment scripts. Every bespoke script is a pipeline without gates, and it will be the one used during an incident.
  • Don't make the canary analysis a custom ML model on day one. A statistical comparison of canary vs baseline per metric — what Kayenta-style judges do — is enough for years.
  • Don't add a human approval step to every production deploy. Approvals that happen 600 times a day become a rubber stamp; put humans on exceptions, expedites and one-way changes.

The Staff signal is knowing that the promotion and gate layer is the part worth owning, and everything under it — build execution, container runtime, rollout primitives — is worth buying or adopting. See Buy or Build: The Total-Cost Test and fault line 4.

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy It MattersWhat to Say
Which change typesCode-only pipelines leave config and content as the unguarded path"I'll treat code, config, schema and rules as releases on the same rails."
Number of regions and cellsDetermines the wave plan and the minimum blast radius"12 regions, 3 cells each — the cell is my smallest blast-radius unit."
Acceptable lead timeBake time trades directly against it"Standard changes reach every region in 1–2 days; tier-0 services in 3–5."
Who can approve an expediteWithout an owner, every change becomes urgent"Expedite needs a second reviewer from a small, named group."
Regulated workloadsMay require change records and segregation of duties"The change log and approvals are the audit record; no separate ticket system."
Mono vs polyrepoChanges cross services in one commit, or across repos in order"I'll assume polyrepo; cross-service changes must be backward-compatible in either order."
Existing change failure rateTells you whether this is greenfield or a safety retrofit"If one in five deploys causes an incident today, gates come before speed."

2.4 Precise Terminology#

TermMeaningCommon Confusion
DeployPutting a new artifact on machinesTreated as the same as release; with flags, new code can be deployed dark
ReleaseExposing new behavior to usersConflated with deploy, so every behavior change becomes a binary rollout
ArtifactImmutable build output identified by content digestMutable tags like latest treated as artifacts
Canary / one-boxA small slice of production running the new version, compared against a baselineTreated as a time delay rather than an experiment with a comparison
BaselineInstances running the current version, started at the same time and size as the canaryCompared against the whole fleet, whose long-running instances differ in warm caches and uptime
Bake timeTime a stage is observed before promotionAssumed to imply evidence; zero traffic for an hour is still zero evidence
WaveA group of targets promoted together after the previous group's gateConfused with a rolling batch inside one target
CellAn independent slice of a region's capacity with its own dependenciesAssumed to exist when the "cells" share a database
RollbackReturning a target to the previous artifactAssumed always safe because artifacts are immutable
Roll-forwardFixing by shipping a new changeAssumed fast; it's a full pipeline run unless expedited
N-1 compatibilityThe previous version works correctly alongside, and on data written by, the new versionTested only in one direction (new reads old)
Change failure rateShare of deployments that need immediate interventionCounted only when someone files an incident
FreezeA period when some or all changes are blockedImplemented as "no deploys", which also blocks fixes

3. Where the Design Splits#

Each fault line below follows the same shape: the options, who pays for each, the Staff default, and when to deviate.

3.1 Fault Line 1: Deploy Speed vs Bake Time#

The tension: Every hour of bake catches faults that only appear under load, after cache warm-up or after a leak accumulates. Every hour of bake also delays every fix, every feature and every security patch — and, less obviously, makes each release bigger.

The hidden cost is batch size. By Little's law, the changes in flight in a pipeline equal the arrival rate times the time a change spends in it. A service merging 6 changes a day through a 2-day pipeline has ~12 changes in flight; if releases are cut every 4 hours, each one bundles about one change. Stretch the pipeline to 5 days and cut releases daily, and each release bundles 6 changes — so when one fails, the gate tells you a batch is bad and someone has to bisect it.

StrategyWhat WorksWhat BreaksWho Pays
Minutes: deploy everywhere within an hourFast fixes; tiny batches; engineers see results the same daySlow-burn faults (leaks, daily batch jobs, cache expiry at 24h) reach every region before they showEvery customer in every region when a slow burn lands; on-call
Fixed long bake: 24h per stageCatches daily-cycle faults5+ days to global; fixes wait as long as features; releases bundle many changesEngineers (latency); customers waiting for fixes; whoever bisects a 15-change batch
Graduated bake: long early, short lateMost risk is found in the first two waves; later waves move fastNeeds a real wave plan and per-wave alarmsPlatform team (wave plan ownership)
Evidence-based gates + wall-clock floor (Staff default)Bake ends when the changed paths have been exercised enough, not after an arbitrary interval; floor catches slow burnsLow-traffic paths take longer; needs per-endpoint, per-version metricsEngineers changing rare paths; platform pays metric cardinality

The Staff default: graduated, evidence-based gates. One-box: floor 60 minutes, plus ≥ 20,000 requests across touched endpoints or synthetic coverage of each one. First region: floor 12 hours, so the change sees a full daily peak in at least one region. Later waves: floor 2–4 hours. This matches the public Amazon pattern quoted earlier. Standard changes reach all regions in about 1–2 days for most services, 3–5 days for tier 0.

"I want the first region to see a full daily peak before anyone else gets the change — that's the 12-hour floor. After that, I'm mostly buying protection against correlated failure, so later waves can move in hours. And the gate counts requests on the paths the diff touched, so a change to a rarely used endpoint doesn't pass on silence."

When to deviate:

  • Security fixes and incident mitigations: expedited risk class — same stages, floors cut to ~15 minutes per wave, second reviewer required. Still not global-at-once.
  • Internal tools and low-blast-radius services: shorter floors; the cost of a bad deploy is small.
  • Changes behind a flag that is off: the deploy itself changes no behavior, so a shorter bake is fine; the flag ramp gets the careful staging. See Feature Flags.

3.2 Fault Line 2: Region-by-Region Waves vs Global Push#

The tension: Waves bound blast radius but leave regions running different versions for days, which complicates debugging and any cross-region protocol. A global push is consistent and fast, and makes every bad change a global outage.

Diagram: 3.2 Fault Line 2: Region-by-Region Waves vs Global Push
StrategyWhat WorksWhat BreaksWho Pays
Global pushOne version everywhere; simple reasoning; fastestAny bad change is a global incident; no comparison populationEvery customer, every time it goes wrong
Region by region, seriallySmallest blast radius at every step12 regions × 12h ≈ 6 days; long-lived version skewEngineers (latency), anyone debugging cross-region behavior
Waves of increasing size (Staff default)First exposure is small; later waves parallel; at most one cell per region in flight at onceVersion skew for 1–2 days; needs mixed-version compatibilityEngineers (mixed-version discipline); platform (wave plan)
Cell-first within regionRegion never fully affected by one stepNeeds real cell isolation, or "cells" share the failureInfra teams building cell boundaries

The Staff default: waves, and inside each region, one cell at a time. The rule that matters: no single step may touch more than one cell in any region, so a bad change can't take a whole region down, even in the late parallel waves. The first wave goes to a low-traffic region, the second to a high-traffic one where all code paths get exercised.

Version skew is a design constraint, not a bug. During a 2-day rollout, every interaction between services, regions and replicas may involve two versions. That's why the mixed-fleet compatibility rule — new code must work alongside old code, in both directions — is part of this fault line. See Multi-Region Active-Active for cross-region replication with mixed versions.

When to deviate: a global push is justified only for changes whose absence is the outage — revoking a leaked credential, blocking an active attack — and then through the emergency path: validated, signed, two-person approved, with consumers that keep last-known-good. Cloudflare's post-2019 procedure — staged by default, global for active attacks — is exactly this split.

🎯 Staff Move: "Fast propagation is a feature for the emergency and a hazard for everything else. I'll keep the global lane, but it's a different door, with a different key, and every use is reviewed."

3.3 Fault Line 3: Automatic vs Human-Approved Rollback#

The tension: Automatic rollback reacts in minutes and occasionally rolls back a healthy change because of a noisy alarm or an unrelated dependency blip. Human-approved rollback adds judgment and 15–40 minutes of page-to-decision time at night, during which the bad change keeps serving.

StrategyWhat WorksWhat BreaksWho Pays
Human decides every rollbackNo false rollbacks; context-awareDetection-to-undo measured in tens of minutes; fatigue at nightCustomers during the decision time; on-call
Automatic on any alarm in scopeFast undo, often before the page is readFalse rollbacks when a dependency blips; flapping alarms block all promotionEngineers whose healthy changes are rolled back; platform (alarm hygiene)
Automatic on high-severity aggregate alarm + canary verdict; humans approve resuming (Staff default)Fast undo on real regressions; humans decide whether to retry, not whether to stopNeeds well-scoped alarms; rollback must be safe (fault line 3 depends on section 4.2)Service owners (alarm quality)

The Staff default: roll back automatically, resume manually. The asymmetry is the argument: rolling back a healthy change costs minutes of delay; leaving a bad one costs the incident. The rollback decision uses two signals — the service's high-severity aggregate alarm scoped to the stage's targets, and the canary-versus-baseline verdict. A rolled-back release is held; the owner reviews the evidence and either fixes it or resumes with a justification.

Diagram: 3.3 Fault Line 3: Automatic vs Human-Approved Rollback

Guarding against false rollbacks: compare canary against a baseline started at the same time, not against the long-running fleet, so dependency blips affect both arms equally. Require the regression to be in the canary's metrics relative to the baseline, not in the global metric. Track deploy.auto_rollback_false_positive_rate (rollbacks later resumed unchanged) and keep it under ~10%; above that, engineers start asking for overrides.

When to deviate: for changes flagged one-way (a contract migration already executed), automatic rollback is disabled for that release and the plan is roll-forward; the pipeline pages a human instead. For client software, "rollback" is a halt — stop the rollout and flip kill switches — and that also happens automatically on crash-rate gates.

3.4 Fault Line 4: One Central Pipeline vs Per-Team Pipelines#

The tension: A central pipeline platform gives every service the same safety — waves, gates, rollback, change log — and becomes a bottleneck for teams with unusual needs and a single point of failure for every change. Per-team pipelines are fast and tailored, and each one is a separate opinion about how much risk is acceptable.

StrategyWhat WorksWhat BreaksWho Pays
Each team scripts its ownSpeed; fits each serviceN levels of safety; the weakest pipeline sets the company's outage rate; no single change logEvery customer of the weakest team; incident responders hunting changes
One central pipeline, platform-definedConsistent safety; one change log; one place to fixEvery unusual case is a ticket; one platform bug blocks every deployProduct teams (velocity); platform team (queue of requests)
Paved road: platform-owned engine + team-owned pipeline definitions that inherit defaults (Staff default)Safety floor is uniform; teams tune above it in code they ownInheritance rules and exceptions need governancePlatform (engine, defaults, exception process); teams (their alarms and evidence)
Monorepo with one global pipelineAtomic cross-service changes; one build graphRelease unit is ambiguous; one broken test blocks everyoneBuild/infra team; everyone when trunk is red

The Staff default: the paved road. The platform owns the orchestrator, gate engine, change log, wave plan defaults and the break-glass path. Each team owns a pipeline definition in its own repo that inherits org defaults and adds service-specific alarms, evidence requirements and extra stages. Teams can make their pipeline stricter freely; making it looser than the org floor needs an exception with an owner and an expiry.

"Safety is a floor the platform owns. Speed above the floor is the team's choice. A team that wants to skip the first region files an exception that someone has to sign and that expires."

Monorepo note: in a monorepo, build and test are central, but deployment should still be per service. The commit is atomic; the rollout isn't. Cross-service changes still need mixed-version compatibility, because services don't reach a region at the same moment.

3.5 Fault Line 5: Freeze Policy — What Ships During an Incident or a Peak#

The tension: Change causes most outages, so freezing changes during peaks and incidents reduces risk. Freezing also blocks the fix, accumulates changes for a risky thaw, and teaches engineers to bypass the pipeline.

Diagram: 3.5 Fault Line 5: Freeze Policy — What Ships During an Incident or a Peak
PolicyWhat WorksWhat BreaksWho Pays
No freezeNo thaw risk; fixes flowRisky features land during peak trafficCustomers during peaks
Blanket calendar freeze (e.g. 3 weeks in December)Fewer changes at peakFixes blocked; bypasses grow; thaw week ships 3 weeks of changes at onceEngineers; customers in the thaw week
Risk-class freeze (Staff default)Features wait; fixes, rollbacks and security flow through the expedited laneNeeds risk classes people classify honestly; needs an approverIncident commander and release managers (approvals)
Incident-scoped freezeDuring an active incident, block changes to the affected regions or services onlyRequires knowing the blast radius of the incidentIncident commander

The Staff default: freezes restrict by risk class and scope, never by "all deploys". Rollbacks always flow. During an active incident, standard changes to the affected regions and their dependencies are held automatically, so responders aren't debugging someone else's deploy. After a freeze, the thaw is metered: the pipeline releases queued changes one batch per service per wave, not all at once.

🎯 Staff Move: "A freeze moves risk in time — it doesn't remove it. I'd rather restrict risky change classes during the peak and meter the thaw than freeze everything and ship three weeks of changes on January 3rd."


4. When It Breaks#

4.1 The Canary That Saw No Relevant Traffic#

t=0 (Thu 01:10 UTC): Release R88 of the orders service enters wave 1: one-box in
           one cell of ap-southeast-2 (lowest-traffic region, 02:00-ish local lull).
           Diff changes the bulk-cancel endpoint (0.3% of traffic globally).
t=+60min:  Bake floor reached. Global error rate on the one-box: 0.08% vs baseline
           0.09%. Bulk-cancel requests served by the one-box: 3. Gate: PASS
           (clock-based gate, no evidence requirement).
t=+13h:    Wave 1 floor reached; ap-southeast-2 fully on R88. Bulk-cancel still rare.
t=+17h:    Wave 2 reaches us-east-1 (high traffic) cell 1. Merchants run end-of-day
           bulk cancels at 17:00 local. Bulk-cancel error rate 0% → 64% in that cell.
t=+17h05m: Aggregate alarm on the service: 0.4% → 0.6%. Below rollback threshold 1%.
t=+17h40m: Merchant support tickets. Manual investigation finds R88.
t=+17h55m: Manual rollback of us-east-1 cell 1 and all of ap-southeast-2.
           ~41,000 failed bulk cancels; 1,200 merchants affected.

Why it happened: the gate measured time, not evidence. The one-box saw 3 requests on the changed endpoint, the aggregate alarm averaged a 64% failure on one endpoint into a 0.2-point blip, and nobody had told the pipeline which endpoint the change touched.

Detection: gate.evidence_requests{endpoint} versus requirement; per-endpoint canary-vs-baseline error rate; deploy.passed_with_low_evidence_total as a platform metric.

Mitigation: roll back; add the endpoint to the release's evidence requirement; replay synthetic bulk-cancel traffic against the one-box.

Prevention: evidence requirements derived from the diff — endpoints and handlers whose code changed, mapped through the service's route table — with a default floor of N requests per touched endpoint; endpoints that can't reach N in the floor time get synthetic traffic or an explicit owner waiver. Per-endpoint alarms for the top revenue paths. Schedule wave 1 so its floor overlaps the region's daily peak; the Google SRE workbook makes the same point that canaries must see enough traffic to be representative and that performance defects tend to show under heavy load (Google SRE workbook).

Owner: service team (alarms and route coverage); platform team (evidence engine and the "passed with low evidence" report).

4.2 Rollback Fails Because of a Data Change#

t=0:       Release R120 of the billing service adds payment_state = 'PARTIALLY_REFUNDED'.
           R120 writes it; R119 deserializes payment_state into a closed enum.
t=+3h:     R120 in wave 2. 18,000 rows now carry PARTIALLY_REFUNDED.
t=+3h10m:  Unrelated latency regression in R120 trips the high-severity alarm.
t=+3h11m:  Automatic rollback to R119 in us-east-1 cell 2.
t=+3h12m:  R119 pods throw on read: unknown enum value. Invoice listing 500s for
           every customer with a partial refund. Error rate 2% → 11% in the cell.
t=+3h14m:  Rollback alarm on the rollback. Orchestrator halts: rollback is now
           making things worse. Pages billing on-call and platform on-call.
t=+3h40m:  Decision: roll forward to R120 in that cell, accept the latency issue.
t=+6h:     Hotfix R121 (latency fix) through expedited lane.

Why it happened: the change was deployable forward and not backward. The artifact was immutable; the data was not. This is the classic rollback trap: the new version wrote a value the old version can't read.

Detection: the N-1 stage would have caught it — run R119 against a database snapshot after R120's integration tests wrote to it. In production: deploy.rollback_error_rate_delta (error rate after rollback vs before).

Mitigation: halt automatic rollback when the rollback's own error rate rises; roll forward with an expedited fix.

Prevention: the two-phase pattern. Deploy 1 ("prepare"): every reader accepts the new value — R119.5 maps unknown enum values to a safe default and logs them. Deploy 2 ("activate"): start writing the new value, only after deploy 1 is in every region. Amazon's rollback-safety article describes exactly this prepare-then-activate sequence (Amazon Builders' Library). For schema, expand → migrate → contract with the contract step as a separate, one-way release (Schema Design). The pipeline refuses to start an "activate" release while any target still runs a version without "prepare".

Owner: service team (format discipline); platform team (N-1 stage and the "rollback made it worse" halt).

Diagram: 4.2 Rollback Fails Because of a Data Change

Each box is its own release through every wave. Releases A–D can each roll back safely; E cannot, and the pipeline knows it.

4.3 The Pipeline Is Down When the Fix Is Ready#

t=0:       SEV-1: a memory leak in the session service (from yesterday's release)
           is OOM-killing pods in 4 regions every ~40 minutes.
t=+10min:  Decision: roll back to the previous release. Automatic rollback didn't
           fire: the leak passed the 12h floor in wave 1 (leak takes ~18h to OOM).
t=+12min:  Rollback request fails: the orchestrator's primary region is one of the
           4 affected, its database pods are on nodes under memory pressure.
t=+15min:  Fallback: build the old version again? CI runners are in the same
           region and their artifact cache is cold. Estimated 50 min.
t=+20min:  Someone proposes kubectl from a laptop. Incident commander says no:
           no audit, no wave, and nobody knows the exact old image digest.
t=+35min:  Break-glass runner (out-of-band, different region) deploys the cached
           previous digest to the 4 regions, one cell at a time, two-person approved.
t=+55min:  OOM kills stop. Orchestrator recovers at t=+1h20m; break-glass actions
           are imported into the change log.

Why it happened: the pipeline shared a failure domain with the services it deploys. The worst-case scenario for a deployment system — being needed most when it is weakest — is predictable and usually unplanned.

Detection: pipeline.orchestrator_availability, pipeline.rollback_request_latency_p99, synthetic "deploy a no-op release" probe every 15 minutes per region.

Mitigation: break-glass path; pre-cached artifacts on nodes and in regional registry mirrors.

Prevention: the orchestrator runs active-standby across two regions with replicated state; rollback needs no build, because previous digests are pinned in every regional mirror; the break-glass runner lives outside the main platform with its own credentials, caches the last 30 artifacts per tier-0 service, deploys only digests already in the release store, and still goes cell by cell. Quarterly game day: turn off CI and the primary orchestrator, ship a rollback and a fix. Treat a leak like this one as an argument for per-version memory-growth alerts during bake, not only error alarms.

Owner: platform team (pipeline availability, break-glass); incident commander authorizes break-glass use.

4.4 A Config Push Takes Down Every Region in Seconds#

t=0:       An engineer updates the rate-limit config: adds a rule with a regex
           intended to match bot user agents. Config service validates JSON schema.
t=+4s:     Config replicated to all 12 regions via watch streams; every API gateway
           node compiles the rule.
t=+6s:     The regex backtracks catastrophically on long user agents. Gateway CPU
           at 100% fleet-wide. Global error rate 0.1% → 78%.
t=+3min:   Pages from every region. "What changed?" — config change log shows the push.
t=+9min:   Revert pushed. Gateways recover as they pick up the revert and as
           pods restart.

Why it happened: the config channel was built for speed and treated as data, not as a deploy. Schema validation checked shape, not behavior. This is the Cloudflare 2019 shape of failure on your own infrastructure.

Detection: gateway CPU by config version; config.apply_errors_total; per-cell health after each config version.

Mitigation: revert; consumers fall back to last-known-good when a new config version fails a local health check (CPU, error rate) within 60 seconds of applying it.

Prevention: config goes through the same wave engine with short floors — one cell (5 minutes), one region (10 minutes), then widening; total ~30–45 minutes to everywhere for standard rules. Behavioral validation: compile and run the rule against a corpus of recorded requests with a CPU budget before publish; regex engines with linear-time guarantees for user-supplied patterns. Emergency global lane exists, signed and two-person.

Owner: config platform (staging, last-known-good); rule-owning team (behavioral tests).

4.5 Slow Burn Passes Every Gate#

t=0:       Release R7 of the search service introduces a cache keyed by query
           that never evicts.
t=+12h:    Wave 1 passes: memory 41% → 52%. No alarm (threshold 85%).
t=+28h:    Waves 2–3 pass. Wave 1 region at 79%.
t=+31h:    Wave 1 region pods start OOM-restarting in sequence; p99 latency 3×.
t=+33h:    Waves 1–3 regions all affected; rollback across 7 regions.

Why it happened: gates measured instantaneous health, not trends. A leak is a slope; the first region's floor was shorter than the time to failure.

Detection: per-version resource slope alarms during bake — memory growth per hour normalized by traffic, compared canary vs baseline; file descriptors, threads, connection counts.

Mitigation: roll back the oldest regions first; restart pods to buy time.

Prevention: include slope metrics in the canary comparison; keep the first region's floor ≥ one daily cycle; when a release is in late waves, the pipeline keeps watching early regions and can recall the release from every wave.

Owner: service team (resource alarms); platform (slope metrics in the default canary config).

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Canary passes on no relevant trafficdeploy.passed_with_low_evidence_total, per-endpoint error by versionEvery region the release reaches before the path gets trafficEvidence gates from the diff; synthetic trafficService team + platform
Rollback breaks on new datadeploy.rollback_error_rate_delta > 0Cells being rolled backHalt rollback, roll forward via expediteService team
Pipeline down during incidentNo-op deploy probe fails; rollback_request_latency_p99Every service's recovery timeBreak-glass runner, cached digestsPlatform team
Global config pushCPU / errors by config version, all regions at onceGlobalLast-known-good fallback, revertConfig platform + rule owner
Slow-burn leakMemory slope canary vs baselineEvery region past the floorRecall release from all wavesService team
Fleet drift (target not on intended version)deploy.version_mismatch_targets > 0 after completionMixed behavior, Knight-styleRe-converge; block "complete" until digests matchPlatform team
Flapping alarm blocks all promotiongate.hold_reason{alarm} repeated across releasesOne service's delivery stallsFix alarm; time-boxed owner waiverService team
Approval bottleneckrelease.wait_for_approval_minutes p90Delivery latency org-wideMove approvals to exceptions onlyPlatform + engineering leadership
Thaw stampede after freezeReleases started per hour > 3× normalMany services changing at onceMetered thaw, batch capsRelease management

🎯 Staff Insight: The most dangerous deployment failure is the one that passed. Every other row pages someone. A release that passed its gates with no evidence on the changed path pages nobody until customers do. That's why deploy.passed_with_low_evidence_total is a platform metric with a weekly review, not a dashboard nobody opens.


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingLists stages: build, test, deployNames change types and intents; states first-step blast radius and undo time as requirementsAsks what fraction of incidents start with a change and which channels bypass the pipeline
RolloutRolling update, health checks, maybe blue-greenWaves by cell and region; evidence-based gates with wall-clock floors; canary vs baseline per endpointWave plan, floors and windows as org defaults pipelines inherit; exceptions with expiry
Rollback"Redeploy the previous image"Automatic on alarm, pre-verified with N-1; two-phase format changes; one-way changes flaggedReversibility as a field on every change; quarterly report of one-way changes and their incident rate
Config & contentOut of scopeSame rails, short bake, last-known-good, signed emergency laneInventory of every behavior-changing channel with an owner and a staging policy
Pipeline failureNot consideredBreak-glass path, cached artifacts, orchestrator outside the blast radius, drilledPipeline availability priced into every service's recovery objective; platform has an error budget
Operations"Add alerts"deploy.passed_with_low_evidence_total, rollback error delta, version-mismatch check, change logChange failure rate and recovery time as org metrics by team and change type

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
States blast radius as a number"The first production step is one box, at most 10% of one cell, in a low-traffic region."
Gates on evidence, not time"The gate waits for 20,000 requests on the endpoints the diff touched, with a 60-minute floor."
Knows rollback can fail"Rollback is a deploy too. I prove it's safe with an N-1 stage before production."
Treats config as a deploy"A config change gets the same waves with a 5-minute floor, and consumers keep last-known-good."
Designs the pipeline's own failure"When CI is down during a SEV-1, rollback still works from cached digests, and break-glass is drilled."
Uses the asymmetry"Roll back automatically, resume manually. A false rollback costs minutes; a slow one costs the incident."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Blue-green, so rollback is instant"Ignores data written by the new version and the 100% switch
"Canary for an hour, then everywhere"Clock-based gate; no regional staging; one bad hour = global outage
"We have good tests, so production staging is overkill"Tests catch what someone predicted; production catches the rest
A human approval on every deployApprovals at 600 a day become rubber stamps and add hours of latency
"Config changes are instant, that's the point"Builds the fastest path to a global outage
No mention of mixed versionsEvery rollout has two versions running; incompatibility is a rollout-time bug

5.4 Common False Positives#

  • CI tooling fluency ≠ deployment design. Knowing every pipeline YAML keyword is table stakes; the Staff question is what happens between the first box and the last region.
  • "We use canaries" ≠ canary analysis. A canary without a baseline, a statistical comparison and an evidence requirement is a delay.
  • High deploy frequency ≠ safety. Frequency is good only with small batches, staged exposure and fast undo; frequency alone multiplies incidents.
  • GitOps vocabulary ≠ a rollout design. A Git repo as the source of desired state is a good control-plane choice; it says nothing about waves or gates.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minChange types, intents; state blast-radius and undo-time requirements
Entities & API3–5 minChange, artifact, release, stage execution, gate evaluation, change event
Architecture5–10 min≤ 8 boxes; build, promotion and undo loops
Gates10–18 minEvidence requirements, canary vs baseline, floors, slow burns
Rollback safety18–25 minN-1 stage, two-phase format changes, expand/contract, one-way flag
Config + pipeline failure25–34 minConfig on rails, last-known-good, break-glass path
Pivot (interviewer's choice)34–42 minFreeze policy, monorepo, client software, urgent fixes, org metrics
Wrap42–45 minDamage formula; owner of each term; evolution

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"A zero-day needs patching everywhere in an hour"Expedited path without bypassing safetyExpedited class: same stages, 10–15 min floors, second reviewer, still cell by cell
"We deploy mobile apps too"Client software where rollback doesn't existRings, crash-rate gates, kill switches, compatibility windows
"Make it a monorepo"Release unit vs commit unitAtomic commit, per-service rollout, mixed-version compatibility still required
"The canary passed and the release broke production anyway"Evidence and representativenessPer-endpoint evidence, peak-overlap scheduling, slope metrics
"How do you know your pipeline is any good?"Org-level measurementChange failure rate, recovery time, lead time, frequency by team and change type
"Black Friday is in two weeks"Freeze as risk budgetRisk-class freeze, fixes flow, metered thaw

6.3 What to Deliberately Skip#

  • Build caching internals. One sentence: "hermetic builds with a remote cache; p50 12 minutes."
  • Test pyramid debates. "Unit, contract and integration tests before production; production gates for the rest."
  • Branching models. "Trunk-based with flags."
  • Container image layering. Not the interview.
  • Choosing a CI vendor. "Buy or adopt; we own the promotion layer."

6.4 Follow-Up Questions to Expect#

  1. "Your canary passed. The release broke production in the second region. What did the gate miss?"
  2. "The new version wrote data the old one can't read. Rollback has started. What happens next?"
  3. "A config change took down every region in 10 seconds. How do you make sure it can't happen again without making config slow?"
  4. "CI is down during a SEV-1. How does the fix reach production?"
  5. "How long does a standard change take to reach every region, and what would you trade to make it faster?"
  6. "Who decides whether a change can skip the first region?"
  7. "How would you measure whether the deployment system is improving the company's reliability?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design a CI/CD system for our company."

Staff Answer

"Before drawing — what kinds of change does this need to ship: service code only, or config, schema, rules and client apps too? And how many regions and cells do we run? I'll assume ~800 services, 12 regions with 3 cells each, and that config and schema changes are in scope; client apps I'll treat as a separate pipeline sharing the gate engine.

The constraints I'll commit to: the first production exposure of any change is one box in one cell of a low-traffic region; gates promote on evidence — traffic on the code paths the change touched — not only on time; rollback is automatic, under 10 minutes per cell, and pre-verified to be safe; config rides the same rails; and the pipeline itself has a degraded mode for incidents. I'll walk through: entities → the build, promotion and undo loops → gates → rollback safety → config → the pipeline's own failure → ownership."

Why this is L6:

  • Defines scope by change type and commits
  • States blast radius and undo time as requirements before drawing
  • Includes the pipeline's own failure in the outline

What L7 adds:

  • Asks what fraction of incidents today start with a change and which channels bypass the pipeline
  • Frames the pipeline as a product with an SLO and an on-call
  • Asks how change failure rate is measured and who reviews it
❌ Common L5 Trap

"Developers push to Git, CI builds a Docker image and runs tests, then we deploy to Kubernetes with a rolling update. We'll use blue-green for zero downtime and roll back if health checks fail."

Why this misses: Every piece works and the design has no staging across regions, no judgment between steps, no answer for data written by the new version, and no config story. The next question — "a change breaks one endpoint" — has no good answer.


Drill 2: Core Mechanic — The Gate#

Prompt: "Exactly how does your pipeline decide a one-box is healthy enough to continue?"

Staff Answer

"Three conditions, all required. First, no high-severity alarm scoped to the one-box's cell has fired, and the one-box's own error rate and latency aren't worse than its baseline — a set of boxes running the current version, started at the same time, receiving the same traffic share. Second, a wall-clock floor of 60 minutes to catch warm-up and short slow burns. Third, evidence: for each endpoint or message handler the diff touched, at least N requests served by the new version — N set so a doubling of the baseline error rate is detectable, ~20,000 at a 0.1% baseline. If an endpoint can't reach N in the floor, the pipeline drives synthetic traffic to it or the owner signs a waiver.

The comparison is statistical per metric — a Mann-Whitney-style test with a tolerance band — and a critical metric failing makes the verdict fail regardless of the aggregate score, which is how Spinnaker's open-source judge works (Spinnaker docs)."

Why this is L6:

  • Distinguishes the floor (time) from evidence (requests on changed paths)
  • Compares against a contemporaneous baseline, not the whole fleet
  • Gives a defensible way to choose N

What L7 adds:

  • Makes the evidence requirement derived automatically from the diff and route table as a platform feature
  • Reports 'passed with low evidence' by team weekly
  • Sets the canary configuration defaults org-wide so teams start from a good one

Drill 3: Make It Concrete — Capacity#

Prompt: "Size the system: builds, deploy operations and the data the gates need."

Staff Answer

"Builds: 2,500 merges a day plus pre-merge runs — say 9,000 CI runs a day, peaking at ~1,500 an hour. At 12 minutes and 8 vCPUs each, peak concurrency is 300 builds, ~2,400 vCPUs; with the remote cache hit rate at 80%, average cost drops a lot, so I'd plan ~3,000 vCPUs of autoscaled runners.

Promotion: ~600 releases a day, each touching 36 cells — ~21,600 cell deployments a day, ~15 a minute average. That's trivial for the orchestrator; the hard part is concurrency limits: no more than one cell per region per release in flight.

Gates: per-version, per-endpoint metrics. 800 services × ~40 endpoints × 4 metrics × 2 versions during a deploy × 36 cells is ~9M series at worst, but only services mid-deploy carry two versions — ~15% at any time — so the extra is ~1M series. That's a real cost for the metrics platform, worth negotiating with Metrics & Alerting Platform.

Change log: every deploy, config change, flag flip and migration — ~50,000 events a day, under 100 MB a day. Trivial to store, priceless during an incident."

Why this is L6:

  • Separates throughput (builds) from safety (promotion concurrency)
  • Finds the non-obvious cost: metric cardinality for version labels
  • Sizes the change log as cheap, which makes it an easy yes

What L7 adds:

  • Prices CI compute against engineer waiting time — a 10-minute faster build at 1,500 engineers is worth more than the runners
  • Sets build-time SLOs because slow builds cause batching and bypasses
  • Plans capacity with Autoscaling & Capacity for the metrics platform

Drill 4: Dependency Down — The Orchestrator Is Unreachable#

Prompt: "Your pipeline orchestrator is down for two hours. What still works?"

Staff Answer

"Running services: everything — the orchestrator isn't in any request path. In-flight releases freeze where they are; deploy agents finish their current cell batch and stop, and automatic rollback still works because each cell's agent holds the rollback rule locally: if the high-severity alarm fires during a deploy it started, it reverts to the previous digest without asking the orchestrator. New releases queue. Rollbacks of completed releases go through the break-glass runner, which deploys digests already in the regional mirrors with two-person approval. When the orchestrator returns, it reconciles: reads each cell's actual digest and imports break-glass actions into the change log."

Why this is L6:

  • Separates the data plane (services) from the control plane (pipeline)
  • Pushes the rollback decision down to the agent so it survives control-plane loss
  • Reconciles actual vs intended state afterwards

What L7 adds:

  • Sets pipeline availability as an SLO with an error budget the platform team owns
  • Requires the break-glass path to be drilled quarterly and measured
  • Refuses to let any team keep a private deploy script as their 'backup'

Drill 5: Hot Spot — One Service Deploys 60 Times a Day#

Prompt: "The web front end merges 60 changes a day and the pipeline takes a day and a half. What happens?"

Staff Answer

"Little's law: 60 changes a day × 1.5 days is ~90 changes in flight. If each change were its own release, 90 releases would be stacked in waves, and a failure in wave 2 would block 80 others. So for high-volume services, the unit is a batch: the pipeline cuts a release every 2 hours from whatever's merged — ~5 changes — and a release can only enter wave 1 after the previous one has left it, as in Amazon's description of promotion. When a batch fails, the pipeline bisects automatically: it re-runs the one-box with halves of the batch while other changes wait. Most behavior changes ship dark behind flags, so the binary deploy is low-risk and the flag ramps carry the real exposure — the pattern Meta describes with Gatekeeper."

Why this is L6:

  • Quantifies in-flight work and batch size
  • Chooses batching with automatic bisection
  • Moves behavior risk to flags so binary deploys stay boring

What L7 adds:

  • Recognizes a service with 60 merges a day as an org design signal — maybe it should split
  • Funds flag tooling and flag-debt cleanup as part of the deploy platform
  • Measures batch failure rate as a separate metric

Drill 6: Multi-Tenant — A Big Customer's Cell#

Prompt: "Your largest customer, 18% of revenue, lives in its own cell. Should their cell get changes first, last or in the middle?"

Staff Answer

"Last among their region's cells, and never in the first two waves. The first waves exist to find problems with small exposure; putting 18% of revenue in them defeats the purpose. Their cell also gets a longer floor and a contract-level change window if the agreement says so. The risk is that their usage pattern differs — a feature only they use. For that, I'd require evidence on their top endpoints in a synthetic or shadow environment that replays their traffic shape before their cell gets the release."

Why this is L6:

  • Places blast radius by business impact, not only traffic
  • Notices the representativeness gap and closes it with replay
  • Keeps the contract window explicit

What L7 adds:

  • Makes cell placement in waves a policy tied to customer tier, reviewed with sales and support
  • Offers premium customers a ring they control — the same ask CrowdStrike committed to for content
  • Prices dedicated cells in the contract

Drill 7: Build vs Buy#

Prompt: "Should we build our own deployment platform or use an off-the-shelf CD product?"

Staff Answer

"Split it. Build execution: buy or adopt — a hosted CI or an open-source runner fleet. Rollout primitives: use the orchestrator's — Kubernetes deployments, or a progressive-delivery controller. The promotion layer — wave plan, gate definitions, evidence requirements, change log, freeze policy — is where safety lives and where our topology, cells and alarms are specific to us. I'd adopt an open-source CD system for the orchestration if it can express waves and external gates, and own the gate configuration and the change log. Building the whole thing from scratch is 6–10 engineers for a year before it's as good as what we could adopt."

Why this is L6:

  • Separates commodity from differentiating parts
  • Keeps the safety logic under the company's control
  • Gives a cost for the build option

What L7 adds:

  • Decides the exit plan: pipeline definitions in a neutral format so a vendor change is a migration
  • Treats CI vendor outage as a correlated-failure risk across every team
  • Retires bespoke scripts on a deadline

Drill 8: Policy Change Without an Outage#

Prompt: "You want to require a 12-hour bake in the first region for every service. Today most have 30 minutes. How do you roll this out?"

Staff Answer

"The policy is a change too, and it changes delivery latency for 800 teams. Shadow first: compute for one month what the new rule would have done — how many releases would have been held, how many incidents it would have caught. Then warn: pipelines show 'this would have been held' for two weeks. Then enforce for tier-0 services, with an exception process, then everyone. In parallel, raise build speed and switch high-volume services to batching so the extra bake doesn't stack changes. And I'd publish the data: 'releases that would have been caught' is the argument."

Why this is L6:

  • Applies shadow → warn → enforce to policy, not only to code
  • Anticipates the latency side effect and mitigates it
  • Uses data to win the argument

What L7 adds:

  • Ties the policy to the change failure rate target the org committed to
  • Has engineering leadership sign the exception process
  • Reviews after a quarter whether the floor is set correctly

Drill 9: Cost#

Prompt: "Finance says CI costs $180K a month. Cut it."

Staff Answer

"First, where it goes: usually 60–70% is test execution, and a large share of that is tests unrelated to the change. Three levers: test selection from the build graph — run only tests whose dependencies changed, typically 40–60% fewer test-minutes; remote cache hit rate — every 10 points of hit rate is a direct cut; and spot or preemptible runners for retryable jobs. I would not cut the production gates or the N-1 stage — they cost almost nothing next to an outage. And I'd measure engineer wait time alongside dollars: 1,500 engineers waiting 5 extra minutes per build costs more than the runners."

Why this is L6:

  • Finds the cost drivers before cutting
  • Picks levers that keep safety
  • Counts engineer time

What L7 adds:

  • Presents CI cost per engineer per month as a standard metric
  • Chooses the build-cache investment as a platform bet with a payback date
  • Says no to cuts that move cost into incidents

Drill 10: Multi-Region#

Prompt: "We're going from 3 regions to 15. What changes in the deployment system?"

Staff Answer

"The wave plan. With 3 regions, serial region by region is fine — 3 × 12 hours. With 15, serial is a week, so I group them: wave 1 a low-traffic region, wave 2 a high-traffic one, then 3, 5 and 6 regions in parallel, still one cell per region at a time. Second, the orchestrator and artifact mirrors become multi-region so no region's deploy depends on another being healthy. Third, version skew lasts longer, so cross-region protocols need to tolerate two versions for up to 3–4 days. Fourth, regional regulations may impose change windows; the policy service handles per-region windows."

Why this is L6:

  • Reworks the wave plan with numbers
  • Removes cross-region dependencies from the pipeline
  • Calls out longer version skew

What L7 adds:

  • Uses region count growth to invest in cells, which make waves meaningful
  • Sets one org-wide wave map so all services agree which regions go first
  • Prices the slower global lead time and gets product to accept it

8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — A Deploy During the Biggest Sale of the Year#

Context: It's the evening peak of a seasonal sale. Checkout latency p99 jumps from 400 ms to 3 seconds in two regions. Twelve releases are mid-pipeline, three of them in those regions. The incident commander asks you: "Is it a deploy?"

Questions to Surface First:

  • What does the change log say changed in those two regions in the last 6 hours — code, config, flags, migrations?
  • Do the latency increases line up with a specific version label?
  • Is anything shared between the two regions that isn't shared with healthy ones?
  • Is the freeze policy in effect, and did anything bypass it?

Typical L5 Approach: Rolls back all three releases in those regions at once. Latency doesn't recover — the cause was a feature flag ramp to 50% that isn't in the deploy pipeline, and the triple rollback added churn during peak.

Staff Approach: Queries the change log first: three deploys, one config change, one flag ramp to 50% at 18:02. Latency by version label shows old and new versions equally slow — not a deploy. The flag ramp matches the onset. Flips the flag to 0%; latency recovers in 2 minutes. Freezes standard releases in those regions for the rest of the peak.

Principal Approach: Asks why flag ramps weren't subject to the peak freeze policy and why the change log was the only place they appeared together. Makes flags, config and deploys one change stream with one freeze policy, and adds 'ramp during peak' to the risk classes.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Query change log for the two regions; compare latency by version label; check flag ramps
TriageVersion labels show no difference → not a deploy; flag ramp correlates with onset
Quick fixFlag to 0%; hold pipeline releases in affected regions
GuardrailsFlag ramps above 10% blocked during peak freeze; ramps gated on the same latency alarms
Post-mortemWhy were flags outside the freeze policy? Why was "is it a deploy?" answerable only by one person?

Metrics to Watch: checkout.latency_p99{version}, flag.ramp_events, change_log.events{region}, deploy.in_flight{region}

Organizational Follow-up: flags join the change stream and the freeze policy; incident runbook starts with the change-log query.

Ownership Question: "Who decides to freeze flag ramps during peak?" Staff answer: The release policy owner — the platform team — sets the default; product can request exceptions that the incident commander for the peak approves.

Key Takeaway: "'Is it a deploy?' is answered by version labels and one change log, not by rolling back everything that moved."

What clears the Staff bar:

  • Checks the change log before acting
  • Uses version-labeled metrics to rule deploys in or out
  • Extends freeze to flags and config

Deep Dive 2: Silent Failure — Gates Have Been Passing on No Evidence for Months#

Context: An audit after an incident finds that 23% of releases in the last quarter passed the one-box gate with fewer than 100 requests on any changed endpoint. Three incidents trace to such releases. Nothing ever paged about it.

Questions to Surface First:

  • Do gates have evidence requirements at all, or only time floors?
  • Which services and which hours are most affected?
  • Is the one-box stage scheduled in a region's off-peak?
  • Do services' route tables let us map diffs to endpoints?

Typical L5 Approach: Raises every bake floor from 60 minutes to 4 hours. Lead time grows for everyone; low-traffic endpoints still see almost nothing in 4 hours.

Staff Approach: Adds evidence requirements derived from the diff, schedules wave 1 floors to overlap the region's peak, adds synthetic traffic generation for endpoints below N, and publishes deploy.passed_with_low_evidence_total by team.

Principal Approach: Treats "the gate is green but proves nothing" as a measurement-integrity problem across the org: every gate must report its evidence, gates without evidence are labeled as such in dashboards, and the platform's quarterly review tracks the share of releases promoted on real evidence.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Not an outage; open a reliability review; identify the three incidents' releases
Triage80% of low-evidence passes are wave-1 one-boxes between 00:00 and 06:00 local
Quick fixWave-1 starts only during a regional traffic window; evidence report on every release
GuardrailsDiff-derived evidence requirements; synthetic traffic for rare endpoints; waivers expire
Post-mortemWhy did a gate that measures nothing look identical to one that measures everything?

Metrics to Watch: deploy.passed_with_low_evidence_total{team}, gate.evidence_requests{endpoint}, gate.waivers_active

Organizational Follow-up: gate evidence becomes part of the release record and visible in incident reviews.

Ownership Question: "Who owns the evidence requirement for a given endpoint?" Staff answer: The service team owns its endpoints' requirements; the platform owns the defaults and the report that shows who's below them.

Key Takeaway: "A green gate that measured nothing is worse than no gate, because it looks like proof."

What clears the Staff bar:

  • Recognizes clock-based gates as the root cause
  • Fixes with evidence, not longer clocks
  • Makes the silent failure visible as a metric

Deep Dive 3: Large-Customer Onboarding — A Regulated Bank Requires Change Control#

Context: A bank signs a contract requiring that changes to its dedicated cell be approved, logged with a reviewer, and deployed only in a weekly window. Sales signed it. Engineering finds out at kickoff.

Questions to Surface First:

  • What exactly does the contract require — approval per change or per release? Evidence of testing?
  • Can a change record in our system satisfy their auditors?
  • What happens to security fixes outside the window?
  • How long can their cell lag the rest of the fleet?

Typical L5 Approach: Creates a separate manual process: a ticket per change, a human runs the deploy in the window. Their cell drifts weeks behind and becomes the snowflake every incident involves.

Staff Approach: Models the bank's cell as a wave with its own policy: last in its region, weekly window, approval recorded on the release by a named reviewer. The change log already has the audit trail. Security fixes use an expedited class pre-agreed in the contract. Mixed-version compatibility window extends to 2 weeks for that cell.

Principal Approach: Creates a standard "controlled cell" product tier with pricing, change window, approval semantics and security-fix terms — so sales sells a defined offering instead of bespoke promises.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Read the contract clauses; list which the existing release record already satisfies
TriageGaps: weekly window, named approver, security-fix exception terms
Quick fixPolicy service gets a per-cell window; approvals become release-record fields
GuardrailsMax lag of the controlled cell: 2 releases or 14 days; compatibility tests cover that span
Post-mortemHow did a deploy-policy commitment get signed without engineering review?

Metrics to Watch: cell.version_lag_days{cell}, release.approval_wait_hours{cell}, expedite.count{cell}

Organizational Follow-up: deployment terms become a standard contract schedule reviewed by the platform team.

Ownership Question: "Who approves changes for the bank's cell?" Staff answer: A named reviewer rotation in the owning service team, recorded on the release; the bank's auditors read our change log, not a separate ticket system.

Key Takeaway: "Change control is a policy on the same pipeline, not a second, manual pipeline."

What clears the Staff bar:

  • Expresses the requirement as wave and policy configuration
  • Bounds the cell's version lag
  • Pre-negotiates security-fix terms

Deep Dive 4: Post-Mortem — Rollback Made It Worse#

Context: Last night a release was automatically rolled back after a latency alarm. The rollback raised errors from 1% to 14% in two cells for 40 minutes, because the new version had written records in a new format. You are leading the post-mortem.

Questions to Surface First:

  • Did the pipeline's N-1 stage exist for this service? Did it cover this data store?
  • Was the change flagged as reversible? By whom?
  • Did the orchestrator detect that the rollback was worsening things?
  • Why did roll-forward take 40 minutes?

Typical L5 Approach: Action item: "be careful with data format changes." Adds a checklist item to the PR template.

Staff Approach: Action items with owners: N-1 stage made mandatory for services with persisted state, covering every store the service writes; format changes must be two-phase and the pipeline blocks an "activate" release until "prepare" is in every cell; the orchestrator halts a rollback when the rollback's error rate rises; expedited roll-forward rehearsed.

Principal Approach: Adds reversibility to the change record and reports one-way changes per quarter; makes "rollback made it worse" a named incident category tracked across the org; funds a shared serialization library that tolerates unknown values by default.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)(During the incident) halt rollback; roll forward in affected cells
TriageNew enum written by v2; v1 deserializer rejects unknown values
Quick fixv1.1 tolerant reader shipped to all cells via expedite
GuardrailsN-1 stage mandatory; rollback error-delta halt; two-phase enforcement
Post-mortemWhy was a one-way change shipped as reversible, and why did nothing check?

Metrics to Watch: deploy.rollback_error_rate_delta, release.one_way_total{team}, n_minus_1.failures_total

Organizational Follow-up: platform owns the N-1 stage; service teams own their data-format compatibility tests.

Ownership Question: "Who decides a change is reversible?" Staff answer: The author declares it; the N-1 stage verifies it; a one-way declaration needs a second reviewer from the service's owning team.

Key Takeaway: "Rollback safety is tested before production, or it's a hope."

What clears the Staff bar:

  • Names the data, not the binary, as the cause
  • Adds a mechanical check rather than a checklist
  • Teaches the pipeline to notice that rollback is making things worse

Deep Dive 5: Multi-Region Expansion — From 3 Regions to 12 With Cells#

Context: The company is expanding from 3 regions to 12 and introducing cells. Today, every deploy goes to all 3 regions in sequence with a 30-minute bake each. Leadership wants "the same speed" in 12 regions.

Questions to Surface First:

  • Are cells truly isolated, or do they share databases or caches?
  • How many services can tolerate a multi-day version skew?
  • Do all services agree on which regions go first?
  • What's the current change failure rate and lead time?

Typical L5 Approach: Keeps the serial plan with shorter bake — 12 regions × 15 minutes. Blast radius per step grows relative to bake, and slow burns now reach all 12 regions in 3 hours.

Staff Approach: Introduces an org-wide wave map: wave 1 (low-traffic region, cell by cell), wave 2 (high-traffic region), waves 3–5 parallel groups of 3, 4 and 3 regions, one cell per region at a time. Lead time ~1.5 days for standard, with an expedited class at ~3 hours. Verifies cell isolation before treating cells as blast-radius units.

Principal Approach: Tells leadership "same speed" is the wrong goal and offers the trade in numbers: lead time 1.5 days vs today's 1.5 hours, in exchange for bounding any bad change to one cell per region. Makes the wave map a company standard and aligns region launch order with it.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Draft wave map; inventory services with cross-region protocols
TriageFind "cells" that share a database — they're one blast-radius unit
Quick fixWave map in the default pipeline template; expedite class defined
GuardrailsOne cell per region per step; version-skew tests for 4-day spans
Post-mortem(Planning review) Are lead-time and change failure rate both tracked before and after?

Metrics to Watch: release.lead_time_hours{risk_class}, deploy.cells_in_flight{region}, incident.multi_region_from_single_change_total

Organizational Follow-up: region launch checklist includes "position in the wave map" and "cell isolation verified."

Ownership Question: "Who decides which region goes first?" Staff answer: The platform team owns the wave map, reviewed with SRE and product; individual services can't reorder it, only add stages.

Key Takeaway: "More regions means a deliberate wave plan, not the old plan run faster."

What clears the Staff bar:

  • Uses parallel waves with per-region cell limits
  • Verifies that cells are real before relying on them
  • Gives leadership a lead-time number and what it buys

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • State the damage of a bad change as exposure × (detection time + undo time) and show how each stage shrinks a term
  • Design a wave plan by cell and region with a first exposure of one box in one cell
  • Define gates with wall-clock floors and evidence requirements on the code paths a change touched
  • Explain why rollback fails, and make it safe with an N-1 stage, two-phase format changes and expand → migrate → contract
  • Put config, rules and content on the same rails with short bake and last-known-good consumers
  • Design the pipeline's degraded mode: agent-local rollback, cached digests, a drilled break-glass path
  • Set freeze policy by risk class and meter the thaw
  • Decide what to buy (CI, rollout primitives) and what to own (promotion, gates, change log)

The Bar for This Question#

Mid-level (L4): Builds a working CI pipeline: build, unit tests, push an image, rolling deploy, maybe a staging environment. Doesn't consider staged exposure, rollback safety or config. Would ship good changes quickly and bad changes everywhere.

Senior (L5): Adds canaries, blue-green, health checks and automated rollback. Knows deploys cause incidents. The gap: canaries are clock-based and compared against nothing, rollback is assumed safe because artifacts are immutable, config is somebody else's problem, and the pipeline has no failure posture. The design is plausible and would pass review — and would still let a change to a rare endpoint pass canary on three requests.

Staff+ (L6): Frames deployment as blast radius and time to undo in the first five minutes. Builds waves by cell and region, gates on evidence, compares canary against baseline per endpoint, and proves rollback safety before production. Puts config on the same rails, designs the break-glass path, and sets freeze by risk class. Names who pays: engineers carry lead time, the metrics platform carries version-label cardinality, service teams own alarms and evidence. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 Blue-Green Is a Rollback Story That Ignores the Database#

ClaimReality
"Instant rollback: flip the load balancer back"Instant for the binary; the shared database already holds whatever green wrote
"Zero-risk cutover"The switch moves 100% of traffic at once — it's a global push with a fast undo
"We test green before switching"Tests without production traffic; the first real request hits 100%

The Staff position: Blue-green is a fine mechanism inside a cell. It is not a rollout strategy. Pair it with staged traffic shifting and the same data-compatibility rules as any other deploy.

Why this matters in interviews: "Blue-green" is the answer that most often ends the candidate's thinking about rollback. Say what it doesn't cover.

10.2 Most Canaries Are Delays, Not Experiments#

Canary TypeWhat It Proves
One box for 30 minutes, no baseline, global alarms onlyThat the process starts and doesn't crash immediately
One box vs contemporaneous baseline, per-metric comparisonThat common paths aren't measurably worse
Plus evidence requirements on changed pathsThat the changed code was exercised and isn't worse

The Staff position: A canary without a baseline and without evidence on the changed path is a timer. Name the comparison and the evidence or don't call it a canary.

Why this matters in interviews: "The canary would catch it" is the sentence interviewers attack first.

10.3 Human Approval on Every Deploy Makes Deploys Less Safe#

EffectMechanism
Rubber-stampingHundreds of approvals a day get approved without reading
Bigger batchesApproval latency encourages bundling more changes per release
BypassesEngineers route urgent changes around the slow path

The Staff position: Automate the common case with gates. Spend human attention on the rare cases — expedites, one-way changes, exceptions to the floor — where it changes outcomes.

Why this matters in interviews: Proposing "a manager approves each production deploy" as a safety measure signals process thinking without systems thinking.

10.4 Config Is the Most Dangerous Code You Ship#

The Staff position: Config is code without a compiler, a test suite or a staged rollout — and it's designed to reach every server in seconds. The Cloudflare 2019 rule and CrowdStrike's Channel File 291 were both data, not binaries. Put config on staged rails with short bake, validate it against what the consumer actually does, and make every consumer survive a malformed payload.

Why this matters in interviews: Bringing up config unprompted is one of the clearest Staff signals on this question.

10.5 A Holiday Freeze Moves the Outage to January#

The Staff position: Blanket freezes stop fixes, encourage bypasses and turn the first week back into the riskiest week of the year. Freeze risky change classes, keep rollbacks and fixes flowing, and meter the thaw.

Why this matters in interviews: Interviewers ask about freezes to see whether you think about risk as a quantity that moves in time.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer designs a safe pipeline. The Principal engineer notices that the company has six ways to change production — the main pipeline, a legacy deploy script two teams still use, the config service, the feature-flag console, the database migration tool run by hand, and the security team's rule push — and that only the first has waves, gates and a change log. The L7 problem is not one good pipeline; it is change itself as a company-wide risk: one inventory of every channel that alters production behavior, one change stream that incident responders can query, one set of org metrics for how often change hurts us and how fast we recover, and a platform team that owns the paved road as a product with an SLO.

🧭 Principal Move: "I'm not going to ask how fast our pipeline is. I'm going to ask how many ways there are to change production without going through it, and what fraction of last year's incidents came through those. That tells me where the next year goes."

The Org-Level Fault Line#

One deployment platform for every change type vs a pipeline per change type.

OptionWhat WorksWhat BreaksWho Pays
Each change type has its own system (code CD, config service, flag console, migration runner)Each optimized for its speedN safety models; no single "what changed?"; the fastest channel has the least safetyIncident responders; customers when the unguarded channel fails
One platform for everythingOne change log, one policy, one set of gatesBecomes a bottleneck; config and flag owners resist slower propagationPlatform team (scope); config owners (latency)
Shared spine, separate deliveryOne change log, one policy service, one gate engine; each channel keeps its own delivery mechanism and speedIntegration work for every channel; contracts between platform and channel ownersPlatform team (spine); channel owners (integration)

🧭 Principal Move: "Every channel that changes production behavior must write to the change log, must respect freeze policy, and must stage through at least one cell with a health check before going wider. How it delivers — containers, watch streams, flag SDKs — is its owner's choice. New channels don't launch without that contract."

The org metrics. DORA's software delivery metrics — change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate — are the right vocabulary for the VP conversation (DORA). The Principal adds two breakdowns the averages hide: by team, and by change type (code, config, flag, migration). A company whose code change failure rate is low and whose config change failure rate is unknown has a measurement problem, not a safety record.

Cost Model#

Assumptions: fully loaded engineer ~$250K/year; cloud list prices; CI cost dominated by test execution; metrics cost includes per-version labels during deploys; estimates, not quotes.

ScaleEngineers / Services / RegionsCI Compute ($/month)Gate & Metrics Overhead ($/month)Platform HeadcountOn-call LoadNotes
Startup50 / 30 / 1~$3–8K hosted CI~$1K0.5–1 eng part-timeBusiness hours; pipeline issues waitRolling deploy + one canary step + manual promote
Growth500 / 300 / 4~$40–80K~$10–20K5–8 eng: orchestration (2), gates and canary analysis (2), CI and build cache (2), developer experience (1–2)Dedicated rotation; 2–5 pages/month, mostly CI capacity and flaky gatesWaves, evidence gates, change log; config joins the rails
Enterprise5,000 / 3,000 / 15+~$250–500K~$60–150K30–50 eng across CI, release orchestration, config platform, flags, migrations toolingFollow-the-sun; pipeline has its own SLO and error budgetCells, org-wide wave map, break-glass drills, change metrics by team

The pricing insight: at every scale, the pipeline's dollar cost is small next to what it controls. At growth scale, one avoided multi-region outage a year typically pays for the whole platform team. The cost that leadership underestimates is engineer waiting time: 500 engineers each waiting an extra 10 minutes a day for builds is ~80 engineer-hours a day — roughly 10 full-time engineers — so build speed is a headcount line, not an infrastructure line.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Artifact identity by content digest vs mutable tagsOne-way-ishEvery deploy record, cache and audit trail keys on it; switching later means rewriting history
Cell boundaries (what shares a database)One-wayRe-cutting cells is an infrastructure migration across every service
Wave map order that customers and contracts depend onOne-way-ishContract terms reference it; changing it needs customer communication
Release record schema (reversible flag, approvals, evidence)One-way-ishAudit history and dashboards depend on it
CI vendorTwo-wayPainful but bounded if pipeline definitions are in a neutral format
Bake floors and evidence thresholdsTwo-wayConfig, with shadow → warn → enforce
Freeze policyTwo-wayPolicy change with leadership sign-off
Canary judge implementationTwo-waySwap behind the gate interface

🧭 Principal Insight: The decisions I'd slow down on are the cell boundaries and the release record. Everything about gates and bake can change in a quarter; cells and the audit trail take years.

The Standard I'd Write#

RFC-REL-001: Production Change Standard
Status: Approved   Owners: Release Platform + SRE

Scope
  Every mechanism that changes production behavior: service deploys, config,
  feature flags, rules and content, schema migrations, infrastructure changes.

MUST
  1. Every production change is recorded in the change log with author,
     reviewer, artifact or payload digest, targets and timestamp.
  2. No change reaches more than one cell in any region before a health gate
     has evaluated it. Emergency global changes use the signed emergency path.
  3. Service deploys inherit the org wave map and bake floors; teams may add
     stages or lengthen floors, not remove or shorten them without an exception.
  4. Gates evaluate evidence on changed paths, not only elapsed time.
  5. Changes to persisted formats are two-phase (prepare, then activate).
     Services with persisted state pass the N-1 stage.
  6. Rollback is automatic on high-severity alarm; resuming a rolled-back
     release requires the owning team's approval.
  7. Every consumer of dynamic config keeps last-known-good and rejects
     payloads that fail validation.

SHOULD
  1. Ship behavior changes behind flags; ramp flags through the same cells.
  2. Run synthetic traffic for endpoints below evidence thresholds.
  3. Keep the last 30 artifacts of tier-0 services in break-glass caches.

Exceptions
  Filed with the Release Platform; reviewed within 5 business days;
  time-boxed to 1 quarter; VP Engineering sign-off for exceptions to MUST 2 or 5.

Success metrics
  - Change fail rate and failed-deployment recovery time, by team and change type
  - Multi-region incidents caused by a single change: 0
  - Releases promoted with low evidence: < 5%
  - Break-glass drill: quarterly, rollback shipped with CI off in ≤ 30 min

What I'd Tell the VP#

"Most of our outages start with a change, and we have six ways to change production, only one of which is staged and logged. I'm proposing one standard that every change channel must meet — staged through a single cell first, recorded in one change log, automatically rolled back on alarm — while each team keeps its delivery tools. It's about eight engineers for a year on the platform team, it adds roughly a day to how long a standard change takes to reach every region, and in exchange no single bad change can take down more than one slice of one region before we catch it. We'll report change failure rate and recovery time by team every month so you can see whether it's working. The main risk is friction from teams who own the fast channels; I'd bring them in by giving them an emergency lane rather than taking speed away."

Principal Interview Signals#

SignalWhat It Sounds Like
Counts the channels, not the pipeline"How many ways can production change without going through the pipeline?"
Prices lead time against blast radius"A day of lead time buys us a one-cell blast radius. Here's what an unstaged change cost us last year."
Uses org metrics with breakdowns"Change fail rate by team and by change type — config is where our number is unknown."
Sets the floor, lets teams go above it"The platform owns the floor; teams own everything stricter."
Knows when not to standardize delivery"Config and flags keep their delivery mechanisms; they share the log, the policy and the gate."

Staff answers that L7 interviewers find insufficient:

  • "We'll build a safe pipeline with waves and canaries" — correct for one channel; ignores config, flags and migrations.
  • "The platform team owns deployments" — names an owner but not the contract between the platform and the teams, nor the exception process.
  • "We'll track deploy frequency" — frequency without change failure rate and recovery time rewards the wrong thing.

Appendices

Appendix A: Mechanics in Depth#

A.1 Release Lifecycle#

Diagram: A.1 Release Lifecycle

Complete requires every target to report the expected digest — the Knight Capital lesson. A release that "finished" with one target on the old version is not complete.

A.2 Gate Evaluation#

function evaluate_gate(stage, release):
    if high_severity_alarm(stage.scope, since=stage.started_at):
        return ROLL_BACK("alarm", evidence)

    if now() - stage.deployed_at < stage.floor:
        return WAIT("floor")

    for endpoint in release.touched_endpoints:          # derived from diff + route table
        n = requests_served(endpoint, version=release.digest, scope=stage.scope)
        if n < required(endpoint):                      # default ~20K at 0.1% baseline
            if synthetic_available(endpoint): drive_synthetic(endpoint)
            elif waiver_active(endpoint, release): continue
            else: return WAIT("evidence", endpoint, n)

    verdict = canary_judge(canary=release.digest, baseline=stage.baseline_digest,
                           metrics=stage.canary_config)  # per-metric statistical test
    if verdict.critical_failed or verdict.score < pass_threshold:
        return ROLL_BACK("canary", verdict)
    if verdict.score < marginal_threshold:
        return HOLD_FOR_HUMAN(verdict)
    return PROMOTE

A.3 Choosing the Evidence Threshold#

For an error rate p and a regression you want to catch of k×p, a rough rule: you want ~20 baseline-expected errors in each arm, so N ≈ 20 / p. At p = 0.1%, N ≈ 20,000; at p = 1%, N ≈ 2,000. For latency regressions, a few thousand samples per arm usually suffice for p99 shifts of 20%+. These are planning numbers; the judge's statistical test makes the final call.

Appendix B: Data Model#

change            (change_id PK, type, repo, commit_sha, author, reviewers[],
                   reversible BOOL, risk_class, created_at)
artifact          (digest PK, build_id, source_sha, builder_id, signature, created_at)
release           (release_id PK, service, artifact_digest FK, config_version,
                   schema_version, change_ids[], pipeline_id, risk_class, created_at)
stage_execution   (id PK, release_id FK, stage, region, cell, status,
                   started_at, deployed_at, finished_at, baseline_digest)
gate_evaluation   (id PK, stage_execution_id FK, decision, reason,
                   evidence JSONB, canary_score, decided_by, at)
change_event      (event_id PK, source {deploy|config|flag|migration|infra|breakglass},
                   service, region, cell, ref, actor, at)   -- append-only
freeze            (freeze_id PK, scope, blocked_risk_classes[], start, end, owner)

The release store is small — ~600 releases a day × ~40 stage executions — and fits comfortably in PostgreSQL with a replica in a second region. The change-event stream goes through Kafka to the incident tooling and an indexed store queried by region, service and time.

Appendix C: Coordination Mechanisms#

C.1 Concurrency Rules#

  • Per region: at most one cell in a non-steady state per release; at most K releases (default 3) concurrently mid-deploy in one cell, so a regression can be attributed.
  • Per release: a release enters a wave only after the previous release of the same service has left it.
  • Leases: each deploy agent holds a lease per cell (in etcd & ZooKeeper or the orchestrator's database), so two orchestrator instances can't deploy to the same cell.
  • Agent-local rollback: the agent that started a deploy can revert it without the orchestrator if the alarm fires mid-deploy.

C.2 Quick Comparison#

MechanismGuaranteesCostUse For
Orchestrator-driven promotionGlobal ordering, policy enforcementControl-plane dependencyAll standard promotion
Agent-local rollbackUndo survives control-plane lossAgent must hold rule and previous digestRollback during deploy
Break-glass runnerDeploy with CI and orchestrator downSeparate credentials and audit importIncidents only
GitOps reconciliationDesired state in Git, drift correctedReconcile lag; Git becomes a dependencySteady-state convergence

Appendix D: API Contract & Client Behavior#

  • Idempotent release creation: POST /v1/releases with an Idempotency-Key = artifact digest + config version; retries return the same release.
  • Halt is cheap and open: any member of the owning team or the incident commander can halt; halting never needs approval.
  • Expedite requires a reason and a second reviewer, recorded on the release.
  • Webhooks to owners on hold, rollback and completion, with the gate evidence attached.
  • Status API returns the plan, current stage, and the reason for any wait — "waiting for 14,200 more requests on POST /orders/bulk-cancel" is a message an engineer can act on.

Appendix E: Observability#

MetricMeaningAlert
deploy.passed_with_low_evidence_totalStages promoted below evidence requirement (waived)Weekly review by team
deploy.rollback_error_rate_deltaError change after a rollback> 0 → halt rollback, page
deploy.version_mismatch_targetsTargets not on intended digest after "complete"> 0 → page platform
deploy.auto_rollback_false_positive_rateRollbacks resumed unchanged> 10% → alarm hygiene review
pipeline.noop_deploy_probe_successSynthetic no-op release per regionFailure 2× → page platform
release.lead_time_hours{risk_class}Merge to all regionsTrend; SLO per class
ci.build_duration_p95Build time> 25 min → capacity
change_log.ingest_lag_secondsFreshness of "what changed?"> 60 s → page

Control plane vs data plane: the pipeline is never in a request path, so its outages don't take services down — but they extend every service's recovery time. Alert on that: the no-op probe measures the pipeline's ability to deploy, which is what incident responders need.

Debugging the silent failure: every gate decision stores its evidence. When a bad release passes, the first question is "what did the gate actually see?" — and the answer is in gate_evaluation.evidence, not in someone's memory.

Appendix F: Scale Evolution#

StageWhat WorksWhat Breaks Next
1 region, 20 servicesRolling deploy + one canary + manual promoteSecond region doubles the blast radius of each push
3–4 regions, 200 servicesSerial region rollout, time-based bake, auto rollbackCanaries pass on no traffic; config incidents
10+ regions, 1,000 servicesWave map, evidence gates, N-1 stage, change logLead time; teams bypass; cells not truly isolated
15+ regions, cells, 3,000+ servicesOrg-wide wave map, diff-derived evidence, every channel on the spineGovernance of exceptions; flag debt
Diagram: Appendix F: Scale Evolution

What you don't build on day one: automatic bisection of batches, diff-derived evidence, a statistical canary judge of your own (adopt one), per-customer rings, or a cross-channel change spine. Start with waves, time floors plus a manual evidence check, automatic rollback and a change log; add the rest when the triggers in the evolution path fire.

Appendix G: Multi-Tenancy, Fairness and Cost#

  • Fair CI scheduling: per-team concurrency quotas on runners with borrowing when idle, so one team's 2,000-test suite doesn't starve everyone at 10 a.m.
  • Fair promotion: no team's releases block another's; per-service queues, with global limits only on cells in flight per region.
  • Tenant-aware waves: large or regulated tenants in dedicated cells sit late in their region's order, with contractually defined windows (Deep Dive 3).
  • Cost attribution: CI minutes and metrics overhead charged back per team monthly; the expensive teams are usually the ones with slow, unselective test suites — a conversation the platform team should start with data.
  • Client-software rings: employees → beta → 1% → 10% → 50% → 100%, with crash-rate gates at each step and a remote kill switch; for agents running privileged code on customer machines, customer-controlled rings — the control CrowdStrike committed to for Rapid Response Content — are part of the product.
  1. Loading the index…