Hiring BarSupport

Design a Feature Flag Service

Case study97 min read10 diagrams

Technologies referenced in this case study: etcd & ZooKeeper · PostgreSQL · Redis · Kubernetes · Envoy, Kong & NGINX

Related: Degraded Mode Framework · Service Registry · Distributed Consensus · Content Delivery Network · Build vs Buy Framework · Deployment System · Hot Keys · Backpressure

Reading Guide#

Organized for interview use first, reference second. This page designs the service that stores, distributes and evaluates feature flags and dynamic configuration. How individual services use a kill switch inside a degradation ladder lives in Degraded Mode Framework; how code itself is shipped lives in Deployment System. This page links to both rather than repeating them.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Design Splits table → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Design Splits) → Section 4 (Failure Modes) → Deep Dives 1–2
Deep Dive3+ hrsEverything, including Section 11 (Principal Lens) and the appendices on evaluation, bucketing and distribution
What is a Feature Flag & Config Service? — Why interviewers pick this topic

A feature flag service lets engineers change what running software does without deploying it: turn a feature on for 1% of users, run an A/B test, switch traffic to a new database, raise a timeout, or kill a misbehaving dependency in seconds. Dynamic configuration is the same machinery with richer values — numbers, lists, JSON blobs — instead of booleans. Together they are the company's runtime control plane.

The hard part is not storing {"new_checkout": true}. The hard part is that every service in the company reads it, on every request, and almost nobody load-tests or failure-tests that dependency. A flag service that is down, slow, or serving a bad value is not one outage; it is every outage at once. And a flag flip is a production change that skips the deploy pipeline, the code review and the canary — unless the flag service itself provides them.

Before vs After — the "bad value pushed everywhere" scenario:

Without a designed flag service:
t=0:        Engineer raises search.max_results from 50 to 5000 in the flag UI. No review.
t=+3s:      Value streams to all 4,000 search pods in every region at once.
t=+40s:     Search p99 goes from 120ms to 9s. Heap pressure on every pod. GC storms.
t=+2min:    Pods OOM and restart. On restart, the SDK blocks waiting for the flag
            service, which is now overloaded by 4,000 reconnects. Pods fail readiness.
t=+6min:    Search is down globally. On-call can't find what changed: no audit diff.
t=+25min:   Someone remembers the flag. Reverts it. Pods still crash-loop on startup.
t=+48min:   Flag service connection storm subsides. Search recovers.

With a designed flag service:
t=0:        Change request: search.max_results 50 → 5000. Schema says max 500. Rejected.
            Engineer files exception; value 400 approved by search owner.
t=+1min:    Staged rollout starts: canary cell (1% of pods) for 10 minutes.
t=+4min:    Canary cell p99 up 6×, heap +70%. Health predicate fails. Auto-revert.
t=+4min:    Blast radius: 1% of pods for 3 minutes. Audit log names the change.
            Pods that restart load the last-known-good ruleset from disk in 20ms.

Why interviewers reach for this question: It is the cleanest test of shared-fate thinking. Every candidate can design a key-value store with a UI. Few notice that they have just built the one dependency every other system shares, that a configuration change is a deploy in disguise, and that the most important behavior of the system is what its clients do when it is unreachable.

Mechanics Refresher: The Flag Primitives
PrimitiveHow It WorksProsCons
Static config file in the deployValues baked into the artifact; change = redeployReviewed, versioned, canaried for freeMinutes to hours to change; useless as a kill switch
Remote evaluation APIService calls GET /evaluate?flag=x&user=y per checkRules stay server-side; always freshNetwork call per flag per request; flag service in every request's critical path
Local evaluation SDK (server)SDK downloads the ruleset, evaluates in-process~0.2–1 µs per check; works when the service is downRuleset must be distributed to every process; staleness window
Client-side evaluated valuesServer evaluates all flags for one user, ships results to the appRules and other users' targeting never leave the serverValues are a snapshot; refresh cadence decides staleness
PollingSDK asks "anything newer than v1234?" every N secondsSimple, stateless server, self-healingFreshness = poll interval; N × fleet = constant load
Streaming (SSE / gRPC stream)Server pushes changes over a long-lived connectionSeconds-level propagation; no idle loadConnection state at scale; reconnect storms
Percentage rollout (bucketing)hash(flag_salt + user_key) mod 100000 → variationSticky, deterministic, no stored assignmentChanging the salt or the key reshuffles every user
Targeting rulesOrdered rules over context attributes and segmentsExpress "employees, then 5% of EU iOS users"Rich rules cost CPU and memory on every evaluation
Last-known-good (LKG) cacheSDK persists the latest valid ruleset to diskRestarts and outages serve real values, not defaultsCan be stale; needs a version and an age metric

For most production systems: a central control plane with change review and staged rollout, a versioned ruleset distributed by push with a polling backstop through regional relays, local in-process evaluation on servers with an on-disk last-known-good copy, and server-side evaluation for untrusted clients (browsers, mobile apps). The primitives are not the interview — the blast radius of a bad change and the behavior when the service is gone are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What the Interviewer Is Scoring#

A feature flag service is not a key-value store question. Anyone can store a boolean.

It is a shared-fate and change-safety question that tests:

  • Whether you notice that every service reads this system and nobody load-tests it, so its failure mode is the company's failure mode
  • Whether you design what the SDK does when the flag service is unreachable — before you design the flag service
  • Whether you treat a flag flip as a production change that needs the same blast-radius controls as a deploy
  • Whether you see the flag inventory as debt that compounds: 30,000 flags, a third of them dead, each a branch in production code

The key insight: A flag service has two products with opposite requirements. The data plane — evaluation inside every process — must be boringly available and keep working when everything behind it is gone. The control plane — changing a flag — must be deliberately slow and gated for changes that add risk, and instant for changes that remove it. Staff candidates separate the two in the first five minutes and design the data plane to survive the control plane's death.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws flag UI → flag API → database → services call the APIAsks "Server-side, client-side or both? How fast must a kill switch take effect, and what must still work when this service is down?"Asks "How many config systems does the company already run, how many outages last year started with a config change, and who is allowed to change production without a deploy?"
Evaluation"Services call /evaluate and we cache in Redis""Servers evaluate locally from a pushed ruleset — about 1 µs a check, no network. Browsers and apps get server-evaluated values so targeting rules never leave our infrastructure."Publishes one SDK contract (evaluation semantics, defaults, telemetry) every language must pass; retires the three homegrown config clients
Failure"Run the flag service with replicas in three zones""When the service is unreachable the SDK serves last-known-good from disk, then code defaults. It never blocks startup longer than 2 seconds. The data plane survives a total control-plane outage."Runs the flag control plane as a separate failure domain from everything it controls, with a written fail-static contract and a quarterly 'flag service off' game day
Change safety"Changes are audited""A flag change is a deploy: schema validation, owner approval for tier-0 flags, staged rollout by cell with automatic revert. Kill switches that only reduce functionality skip the stages."Makes staged config rollout a company-wide rule for every config system, not only flags, and tracks config-caused incidents as an org metric
Lifecycle"Engineers clean up old flags""Every release flag has an owner and an expiry; at 100% for 30 days it raises a removal ticket; stale count is on the team's dashboard"Treats flag debt as a budget: each team has a flag quota; quarterly debt reviews; removal is part of the definition of done
Scale"Shard the flags table""30K flags, 80K SDK instances, 80M evaluations/s — all local. The hard number is 80K streaming connections and the reconnect storm after a control-plane deploy, so I put relays in front."Prices it: build at ~6 engineers plus on-call vs a vendor at per-seat or per-connection pricing; decides which parts must stay in-house
Why "evaluation" separates levels

L5: "Each service calls the flag service's evaluate endpoint. We cache results in Redis for 30 seconds so we don't hammer it." Reasonable, and it puts the flag service and Redis on every request path in the company. A checkout request that checks 40 flags is now 40 cache lookups or one batched call — 1–3 ms added — and when Redis is unhealthy, every product is unhealthy.

L6: "Server SDKs download the ruleset for their service — a few hundred KB — and evaluate in-process. A check is a hash and a rule walk, about a microsecond. At 2 million requests a second and 40 checks each, that's 80 million evaluations a second and zero network calls. The flag service is never on the request path; it is on the change path. For browsers and mobile apps I do the opposite: the client sends its context, the server evaluates every flag for that user and returns a 3 KB map. Shipping our targeting rules to a phone would leak unreleased feature names and every customer segment."

L7: "We have a flag vendor for product teams, an in-house config system for infra, and a YAML-in-a-bucket thing the data team built. Three evaluation semantics, three ideas of 'default'. I'd standardize the SDK contract — evaluation order, defaults, LKG, telemetry — and let the backends converge over two years. The contract is the one-way door; the backend isn't."

Why "failure" separates levels

L5: "The flag service is highly available — three replicas, a replicated database." True, and it ignores that the same bad value is replicated perfectly to all three replicas, that a bad control-plane deploy hits all zones, and that the candidate hasn't said what 80,000 SDK instances do when they can't connect.

L6: "I design from the client inward. The SDK starts by loading the last-known-good ruleset from local disk, serves it immediately, then connects. If it can't connect it keeps serving LKG and reports its age. If there's no LKG — a brand-new pod in a brand-new cluster — it waits at most 2 seconds and then serves code defaults and emits flags.sdk.serving_defaults. It never blocks the process from starting. That makes a flag-service outage a 'flags are frozen' event, not an outage."

L7: Recognizes the flag service as the company's largest control dependency: "Every team's kill switch lives here, which means the day we most need it is the day something else is broken. I'd run it in its own failure domain — separate cluster, separate database, separate deploy pipeline — and require a break-glass path that works with the control plane dead: a signed override file pushed through the host agent."

Why "change safety" separates levels

L5: "Every change is logged with who made it, and we can roll back." Logging is necessary and changes nothing about blast radius. A bad value still reaches every pod in every region in three seconds.

L6: "Propagation speed is a feature for kill switches and a liability for everything else. So I make change speed asymmetric. Turning a feature off, or a kill switch on — changes that remove risk — propagate globally in under 10 seconds. Everything else goes through validation against a schema, approval by the flag's owner for tier-0 flags, and a staged rollout: one canary cell for 10 minutes, then 10%, 50%, 100%, with health predicates that auto-revert."

L7: "A flag flip, a WAF rule, a routing table and a threat-signature file are the same thing: a production change that bypasses the deploy pipeline. Public post-mortems keep pointing at these paths. I'd write one standard — every system that can change production behavior at runtime must stage it — and audit every config system against it, not only the flag service."

Positions to Commit To#

PositionRationale
Evaluate locally on servers; never put the flag service on the request path80M evaluations/s at ~1 µs in-process vs a network call each; the service is on the change path, not the request path
Last-known-good beats code defaults; code defaults beat blockingA frozen-but-real config is safer than defaults written two years ago, and both are safer than pods that can't start
Push for speed, poll as a backstopStreaming gives ~5 s propagation; a 60 s poll catches every missed event and dropped connection
Asymmetric change speedChanges that reduce risk (kill, disable) go global in seconds; changes that add risk go through staged rollout
A flag change is a deploySchema validation, ownership, approval for tier-0 flags, canary by cell, auto-revert, audit
Untrusted clients get values, not rulesRules reveal unreleased features and customer segments; the server evaluates and returns a per-user map
Every flag has an owner, a type and an expiryRelease flags are temporary by design; without an expiry they become permanent untested branches

Which Problem Are We Solving?#

Three intents produce three different systems. Name them, then commit.

IntentConstraintStrategyFailure ModeCorrectness Bar
Release & kill switches (server-side)1–3K services, ~80K processes, kill must work during incidentsLocal evaluation SDK, pushed ruleset, LKG on disk, asymmetric change speedBad value pushed globally; SDK blocks startup; kill switch can't reach the fleetKill propagation p99 ≤ 10 s; zero request-path dependency; config-caused incidents trend down
Experimentation & targeting (client + server)10–500M users, sticky assignment, exposure logging for analysisDeterministic bucketing, server-evaluated values for clients, exposure events to the analytics pipelineBucketing reshuffle invalidates experiments; exposure logs missing or double-countedAssignment stable across sessions and devices; sample-ratio mismatch < 0.1%
Dynamic operational configTimeouts, pool sizes, rate limits, routing weights; typed values, not booleansTyped schemas, validators, staged rollout, config-as-code with reviewValid-looking value that is wrong at scale (timeout 30 s, pool size 10,000)Every value schema-checked; every change canaried; revert in ≤ 1 min

🎯 Staff Move: "I'll design the server-side release and kill-switch path first — that's where the shared-fate risk lives, because every service depends on it. Dynamic config rides the same distribution with stricter validation. Client-side experimentation reuses the rules engine but evaluates on our servers, and I'll treat exposure logging as an analytics integration rather than part of the core. Tell me if you'd rather go deep on experimentation."

Where the Design Splits#

#Fault LineThe Tension
1Push vs PollSeconds-level propagation and connection state vs simple, self-healing, constant load
2Where Evaluation Happens: Server SDK vs Central Service vs EdgeLatency and independence vs rule secrecy, freshness and SDK sprawl
3Store Unreachable: Last-Known-Good vs Code Default vs BlockStale-but-real vs fresh-but-ancient vs correct-but-down; who signs the posture?
4Per-Request Evaluation Cost: Rich Targeting vs Cheap ChecksExpressive rules and huge segments vs microseconds, memory and GC on every request
5Who Can Flip a Flag in Production — and How FastSpeed for incident response vs review, approval and staged rollout for everything else

How Real Companies Built It#

Why this section belongs here: Configuration systems are where public post-mortems are most consistent. Each of these shows a fault line from this page playing out in production.

Meta — Configerator and Gatekeeper (SOSP 2015)#

Facebook's SOSP 2015 paper describes a configuration stack that, at publication, managed hundreds of thousands of configs distributed to hundreds of thousands of servers and more than a billion mobile devices, with thousands of config changes a day. Gatekeeper, built on it, gates feature rollouts with tens of thousands of "projects" composed from hundreds of predefined restraints; config changes reach production after automated canaries (for example 20 servers, then a full cluster of thousands); commit-to-server latency was about 14.5 seconds at baseline; each server's proxy keeps an on-disk cache so applications can still read configs if every Configerator component fails; and a manual review of three months of high-impact incidents found 16% related to configuration management (Meta / SOSP 2015 paper).

Two details from the paper matter most here. Of the config-related incidents, 22% were valid config changes that exposed latent code bugs — a correct value that took the code down an untested path. And 35% of configs had not been updated once in the previous 300 days, which is flag debt measured at scale.

Staff insight: This is the reference architecture in one paper: config-as-code with review, validators and automated canaries; push distribution through a tree; local on-disk caching so the data plane outlives the control plane; mobile clients that poll (the paper cites once an hour as an example) with push reserved for emergencies. In an interview, cite the 22% — it is why "the value passed validation" is not a safety argument, and why canaries are needed even for flags.

Knight Capital — A Repurposed Flag (2012)#

The SEC's order describes how on August 1, 2012, Knight deployed new code that repurposed a flag formerly used to activate "Power Peg", functionality it had stopped using in 2003 but never removed. A technician did not copy the new code to one of eight servers; orders carrying the repurposed flag reached that server and activated the old code, which sent millions of child orders in about 45 minutes — over 4 million executions in 154 stocks — and Knight lost more than $460 million. The SEC fined Knight $12 million (SEC order).

Staff insight: This is flag debt with a price tag. Dead code behind a live flag is not dead; reusing a flag name ties new meaning to old behavior on any host that didn't get the new code. The rules that follow are cheap: flag keys are never reused, removed flags are tombstoned forever, and a flag's removal ships with the removal of the code it guarded.

Google Cloud — A Feature Without a Flag (2025)#

Google's incident report for June 12, 2025 says a new quota-policy feature added to Service Control on May 29 contained a code path that wasn't exercised during rollout and was not feature flag protected. A policy change with unintended blank fields was replicated globally within seconds, hit the null-pointer path, and crashed Service Control in every region; a "red-button" to disable the serving path was fully rolled out about 40 minutes in, and the incident lasted about three hours. Google committed to require that all changes to critical binaries be feature flag protected and disabled by default (Google Cloud incident report).

Staff insight: Two lessons, one for each plane. The code lesson: new behavior in a critical binary ships dark behind a flag so it can be enabled region by region. The data lesson: the trigger was policy data that replicated globally in seconds — the same speed this page wants for kill switches. Fast global propagation is safe only for changes that reduce risk.

Cloudflare — Two Global Configuration Pushes (2019, 2025)#

On July 2, 2019, a WAF rule with a catastrophic-backtracking regex was distributed globally at once through Quicksilver — which Cloudflare said propagates a change to every machine worldwide at a p99 of 2.29 seconds — and pinned CPU across the network; Cloudflare moved WAF rules to staged rollouts while keeping emergency global deployment for active attacks (Cloudflare July 2019 post-mortem). On November 18, 2025, a database permissions change made the query generating the Bot Management "feature file" — refreshed every few minutes and published to the entire network — return duplicate rows; the file exceeded the proxy's hard limit of 200 features (normally about 60) and the proxies failed. Remediations included hardening ingestion of Cloudflare-generated configuration files the same way as user input and enabling more global kill switches for features (Cloudflare November 2025 post-mortem).

Staff insight: The 2019 incident is covered from the CDN side in Content Delivery Network; read the pair here as config-system lessons. Machine-generated config is still config: validate its size and shape at the consumer, and keep the last good version when the new one fails. And note the remediation that keeps an emergency fast path — staged by default, global when it reduces risk.

CrowdStrike — Content Configuration as Code (2024)#

CrowdStrike's preliminary review says that on July 19, 2024 at 04:09 UTC it released a content configuration update for Windows sensors that passed its Content Validator because of a validator bug, and reverted it at 05:27 UTC; it committed to staggered deployment of this content, starting with a canary and moving through wider rings (CrowdStrike PIR). Microsoft estimated 8.5 million Windows devices were affected (Microsoft).

Staff insight: "It's only content, not code" is how a config path ends up with less protection than a deploy. The fix CrowdStrike committed to — canary, then rings — is exactly the staged config rollout this page makes mandatory. Validators are one layer; they have bugs too.

Follow-Ups to Expect#

After You Say...They Will Ask...(What They're Evaluating)
"Services call the flag API""The flag service is down. What does checkout do?"Request-path dependency; fail-static posture
"The SDK caches flags locally""A pod restarts while the flag service is down. What values does it serve?"LKG on disk vs defaults vs blocking startup
"Changes stream to every SDK in seconds""Someone pushes a bad value. How many pods get it before anyone notices?"Change speed as a liability; staged rollout
"We use percentage rollouts""We change the rollout from 10% to 20%. Do the original 10% stay in?"Deterministic, monotonic bucketing
"The mobile app gets the flags""What can an attacker learn by reading the app's flag payload?"Rules vs values; leaking unreleased features
"Engineers clean up old flags""We have 30,000 flags. How many are dead, and how do you know?"Flag debt as a measured, owned inventory
"Anyone on the team can flip it""It's 3 a.m. and the owner is asleep. Who can kill the feature?"Approval vs incident speed; break-glass

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Two planes with opposite jobs. The control plane (top) is where humans and automation change flags; it is allowed to be slow, gated and occasionally down. The data plane (bottom) is evaluation inside every process; it must work when everything above it is gone, which is why each SDK keeps a last-known-good ruleset on local disk. The distribution layer in between turns one change into 80,000 updates — through relays, so the origin sees ~450 connections instead of 80,000 — and is the layer that decides how fast a bad value can spread.

One-Minute Recap#

TopicThe L5 AnswerThe L6 Answer — Say This
Evaluation"Call the flag API, cache in Redis""Local in-process evaluation on servers, ~1 µs. Server-evaluated values for browsers and apps."
Distribution"Poll every 30 seconds""Push via regional relays for ~5 s propagation, 60 s poll as backstop, versioned snapshots for cold start."
Service down"Replicas""SDK serves LKG from disk, then code defaults; never blocks startup more than 2 s. Flags freeze, nothing breaks."
Bad change"Roll back""Schema validation, owner approval for tier-0, staged by cell with auto-revert. Kills skip the stages."
Rollouts"Random 10%""hash(salt + user_key) mod 100000, monotonic: 10% → 20% keeps the first 10%."
Debt"Clean up later""Owner, type and expiry on every flag; removal ticket at 30 days at 100%; stale count per team."
Who flips"Anyone with access""Owners for normal changes, two-person rule for tier-0, on-call can kill anything, all audited."

Numbers to Bring#

MetricValueWhy It Matters
Local flag evaluation~0.2–1 µs per check (hash + rule walk)40 checks per request cost ~40 µs; a remote call per check would cost milliseconds
Remote evaluation call~0.5–3 ms in-region, plus a new dependencyWhy the flag service must not be on the request path
Flag checks per request10–100 in a mature serviceMultiplies any per-check cost and any per-check failure
Server ruleset per service~100 KB–1 MB scoped; full environment ~10–50 MBScoping by service keeps memory and cold-start download small
Streaming propagation targetp99 ≤ 10 s for kill switchesCloudflare reported a p99 of 2.29 s for its global KV in 2019; seconds are achievable
Poll backstop interval30–60 s servers; 15–60 min mobileUpper bound on staleness when the stream silently breaks
SDK init timeout≤ 2 s, then serve LKG or defaultsAnything longer turns a flag outage into a startup outage
Facebook config propagation (2014)~14.5 s commit to production servers at baseline (SOSP 2015)Even "fast" at Meta was double-digit seconds, and that was fine because canaries took ~10 min
Config share of high-impact incidents16% over three months at Facebook (SOSP 2015)The number to quote when arguing config changes deserve deploy-grade safety
Bucketing resolution100,000 buckets → 0.001% stepsLets you start a rollout at 0.01% of 200M users = 20K users
Staged config rollout1 canary cell 10 min → 10% → 50% → 100%, ~45 minBounds a bad value to ~1% of the fleet for the time it takes to detect
Stale-flag thresholdRelease flag at 100% (or 0%) for 30 daysThe point at which the flag is debt, not a release tool

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on the flag admin UI, the database schema and the targeting-rule language, then run out of time before the only question that matters: "The flag service is down and someone just pushed a bad value. What does the rest of the company experience?" Compress the CRUD to ~8 minutes and spend the rest on evaluation location, distribution, the unreachable-store posture and change safety.


Phase 1: Requirements & Framing (2–3 minutes)#

State the functional scope in one breath:

"Engineers create flags and typed config values, target them by user, tenant, region, app version or percentage, change them at runtime without a deploy, and kill features instantly during incidents. Services evaluate flags on the hot path; web and mobile apps receive values for their user. Every change is audited and reversible."

Then the non-functional requirements, which is where the design lives:

"Four constraints drive everything. One: the flag service is a dependency of every service, so evaluation must never need the flag service to be up — it's on the change path, not the request path. Two: a kill switch must reach the fleet fast; I'll target 10 seconds p99. Three: every other change must be slow on purpose — validated, approved where it matters, and rolled out in stages, because a flag change is a production change. Four: flags are temporary by default and the system has to fight their accumulation. Scale: I'll assume 2,500 engineers, 1,800 services, ~80,000 server processes, ~30,000 flags, 2 million requests a second across the fleet, and 50 million daily active app users."

Then name the underspecified parts:

"A few things I'd confirm: is this replacing existing config systems or greenfield? Are experiments in scope — do we need exposure logging for analysis? Are there regulated features where a flag flip needs approval on record? I'll assume server-side release and kill switches first, typed operational config second, client-side experimentation third."

🎯 Staff Move: Saying "the flag service is on the change path, not the request path" in the first three minutes tells the interviewer you've already decided the most important architectural property. Everything you draw next can be judged against it.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • Flag: key (never reused), kind (release, experiment, ops/kill, permission, config), value_type (bool, string, int, json + schema), owner_team, tier (0–3), created_at, expires_at, variations[], tags
  • Environment config: flag_key, env, on, targets (explicit keys), rules[] (clauses → variation or rollout), fallthrough, off_variation, salt, version
  • Segment: key, rules[], included_keys (capped), version — reusable audiences like "employees" or "EU enterprise tenants"
  • ChangeRequest: id, flag_key, env, diff, author, approvers[], risk_class (reduce / add), rollout_plan, state
  • Ruleset snapshot: env, scope (service), version (monotonic), checksum, flags{}, segments{}
  • AuditEvent: event_id, flag_key, actor, change_request_id, before, after, at

Control-plane API (humans and automation):

POST  /v1/flags                                   { key, kind, value_type, owner_team, tier, expires_at }
POST  /v1/change-requests                         { flag_key, env, diff, rollout_plan }
  → 201 { id, risk_class: "add" | "reduce", required_approvals: 0 | 1 | 2 }
POST  /v1/change-requests/{id}/approve
POST  /v1/kill/{flag_key}                         (emergency: may only set off_variation; on-call role)
GET   /v1/flags/{key}/history                     (audit trail with diffs)

Data-plane API (SDKs and clients):

GET   /v1/rulesets/{env}?scope=checkout&since=v81234     → delta or full snapshot + version
GET   /v1/stream/{env}?scope=checkout                    (SSE: ruleset.updated {version})
POST  /v1/client/evaluate      { context: { key, kind, attrs } }
  → 200 { values: { flag_key: value, ... }, version, ttl: 900 }

The risk_class on a change request is the most important field on this page: it decides whether the change goes out in seconds or in 45 minutes.

🎯 Staff Move: "I'm classifying every change as risk-reducing or risk-adding at submit time. Turning a feature off, or raising a kill switch, reduces risk and gets the fast path. Turning something on, widening a rollout or changing a number adds risk and gets the staged path. That one field is how the same system can be both a kill switch and a safe release tool."


Phase 3: High-Level Architecture (≤5 minutes)#

Draw at most eight boxes:

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk a flag change in 90 seconds:

  1. Engineer submits a change request: checkout.new_tax_engine from 5% to 25% in prod. The admin API validates it against the flag's schema and policy, classifies it as risk-adding, and requires the owner's approval because the flag is tier 0.
  2. On approval the rollout orchestrator stages it: canary cell (~1% of checkout pods) for 10 minutes, watching checkout.error_rate and checkout.p99 against the control cells.
  3. For each stage the publisher writes a new ruleset version for that scope and cell, stores the snapshot, and emits ruleset.updated.
  4. Regional relays receive the event and push it to the SDKs subscribed to that scope. SDKs fetch the delta, verify the checksum, swap the ruleset atomically, and write it to local disk as last-known-good.
  5. On the request path, checkout calls flags.bool("checkout.new_tax_engine", ctx, false) — a hash and a rule walk in-process, about a microsecond.
  6. Mobile apps call the client evaluation service at launch and every 15 minutes, receive a 3 KB map of values for their user, and cache it on the device.
  7. If health predicates fail at any stage, the orchestrator reverts to the previous version — itself a risk-reducing change on the fast path.

🎯 Staff Move: Say out loud: "Notice the flag store never appears on a request path. Evaluation is local. If the entire control plane disappears, every service keeps running on the last ruleset it saw. The only thing we lose is the ability to change flags — which is exactly why the kill path needs its own break-glass." You've now spent ~9 minutes.


Phase 4: Transition to Depth (1 minute)#

"That's the happy path, and it's the Senior-level design. What makes a flag service hard is that it's shared by everything and changes production without a deploy. I'd like to go deep on four things: what every SDK does when the flag service is unreachable, how a change propagates and how we stop a bad one from going global, the cost of evaluation on the hot path when targeting gets rich, and how we keep 30,000 flags from becoming 30,000 untested branches. Where would you like to start?"

If no preference: start with the unreachable-store posture. It's the question that decides the level.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → commit → quantify → name who pays.

Deep dive 1: The flag service is unreachable (7–8 min)

"Three cases, and the SDK handles each without asking anyone. A running process that loses its connection keeps serving its in-memory ruleset and reports flags.sdk.ruleset_age_seconds; nothing changes for the request path. A process that restarts loads the last-known-good ruleset from local disk — written atomically on every successful update — and serves it within ~20 ms of startup, then tries to connect in the background. A process with no LKG — a new node, a fresh container without a persistent volume — waits up to 2 seconds for the relay, then serves code defaults and emits flags.sdk.serving_defaults. It never blocks readiness beyond that."

Quantify: "A 30-minute control-plane outage during a normal day sees maybe 4,000 pod restarts from deploys and autoscaling. With LKG they all come up with real values. Without LKG and with blocking init, every one of those pods fails readiness — autoscaling stops working, and a deploy during the outage takes the service down. With defaults but no LKG, 4,000 pods come up running a configuration nobody has seen in production for months."

Who pays: "With LKG, users pay staleness — a flag changed during the outage doesn't take effect until we're back. Product signs off on that. The platform pays for keeping LKG on a volume that survives container restarts — in Kubernetes, a node-local cache populated by a per-node agent rather than per-pod storage."


Deep dive 2: Propagation and change safety (7–8 min)

"Propagation speed is a tool with two edges. I want kills to reach every process in under 10 seconds, so distribution is push: the publisher emits a versioned event, ~450 regional relays hold the 80,000 SDK streams, and every SDK also polls every 60 seconds with its current version in case the stream silently died. Then I deliberately slow down everything that adds risk: schema validation at submit, owner approval for tier-0 flags, and a staged rollout by cell. The SDK knows its cell, and a stage is just a rule — 'cell in {canary}' — so staging costs nothing extra in distribution."

Quantify: "A risk-adding change takes ~45 minutes to reach 100%: 10 minutes in a canary cell at ~1%, 10 at 10%, 15 at 50%, then everywhere. If the change is bad, ~1% of capacity sees it for the 3–5 minutes it takes predicates to fire. A risk-reducing change takes under 10 seconds globally."

Who pays: "Engineers pay 45 minutes of patience on every risky change. That's the deal: product teams lose minutes, the company stops losing hours. The incident commander keeps the global fast path for anything that turns things off."


Deep dive 3: Evaluation cost on the hot path (5–6 min)

"A flag check is a hash and an ordered rule walk; with simple clauses it's under a microsecond. Two things blow that up: segments with huge inline key lists, and rules that need data the SDK doesn't have. I cap inline segment lists at 10,000 keys — bigger audiences become a hashed set compiled into a Bloom filter or a server-side attribute — and I forbid network calls during evaluation: anything that needs a lookup becomes an attribute the caller puts in the context, or a precomputed membership joined into the context upstream. Meta's Gatekeeper did allow one lookup restraint against a flash- or memory-backed key-value store; that's a deliberate, budgeted exception, not a default."

Quantify: "80 million evaluations a second at 1 µs is ~80 cores fleet-wide — noise. At 50 µs because someone put a 2-million-key segment in a regex clause, it's 4,000 cores and a GC problem. The SDK reports flags.eval.duration_us per flag so we can find that one flag."


Deep dive 4: Flag debt (5–6 min)

"Every release flag is a fork in production code with two paths, only one of which is tested daily. I make expiry mandatory at creation: release flags 90 days, experiments 60, ops and permission flags have no expiry but a yearly review. A flag at 100% or 0% for 30 days opens a removal ticket for the owning team with a code reference from the SDK's usage telemetry. Keys are never reused; removed flags are tombstoned so an old binary asking for them gets its code default and an alert, not someone else's new meaning."


Phase 6: Wrap-Up (2–3 minutes)#

"The core idea: the flag service is on the change path, not the request path. Servers evaluate locally and survive the control plane's death on last-known-good. Change speed is asymmetric — kills in seconds, everything else staged by cell with automatic revert. Untrusted clients get values, never rules. And flags have owners and expiry dates, because the inventory is the long-term risk."

The evolution closer:

"What I'd build later: config-as-code for operational config with the same staged pipeline, an org-wide rule that every runtime-config system — WAF rules, routing tables, ML model pushes — goes through staged rollout, and the OpenFeature SDK interface so we can swap backends without touching call sites. What I'd not build: a custom experimentation statistics engine — we'd export exposures to the analytics stack."

🎯 Staff Move: End on the outage you designed out and who owns it. Senior candidates end with "and there's an audit log." Staff candidates end with "and flags.sdk.serving_defaults and config-caused incidents per quarter are numbers the platform team reviews monthly, because those tell us whether the flag service is making outages smaller or causing them."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
UI tourDesigns the flag dashboard and its filters for 8 minutesOne sentence: "UI and config-as-code both create change requests"
Rule-language designSpecifies every operator in the targeting DSL"Ordered rules, clauses over context attributes, segments, percentage rollout"
Database shardingShards a 30,000-row table"30K flags fit in one PostgreSQL primary; the scaling problem is distribution"
No outage postureWaits for "what if it's down?"States "evaluation is local; LKG on disk" in Phase 1
Change speed only as a virtue"Changes propagate in 2 seconds!""Kills in seconds; everything else staged"
Debt as afterthought"And we'd clean up old flags" at minute 44Makes owner and expiry required fields in Phase 2

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

A flag service is the purest example of a shared-fate dependency that hides. Nobody draws it on their architecture diagram. It doesn't serve user traffic. It has no dashboard on the NOC screen. Yet every service in the company imports its SDK, checks it dozens of times per request, and relies on it as the emergency brake. That combination — invisible during design reviews, universal at runtime, and most needed during somebody else's incident — produces exactly the kind of correlated failure Staff engineers are hired to see before it happens.

It also has a deceptive happy path. A Senior engineer can build a flag service that works perfectly on day one: a table, an API, a cache, a UI. The design is judged entirely by three days nobody demos: the day the flag service is down while pods are restarting, the day a valid-looking value is pushed to every region in three seconds, and the day someone reuses a flag key that an old binary still understands. Staff candidates design for those days before drawing the first box.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"The flag service's database has just been corrupted and the control plane is down. In the same minute, a checkout pod crash-loops and a new region is being brought up. What value does checkout.new_tax_engine have in each of those processes — and who decided that?"

A candidate who answers with the running pods (in-memory, last version, unchanged), the restarted pod (last-known-good from local disk, version N, age reported), the new region (no LKG: 2-second wait, then the code default — and the code default had better be the behavior that's live in production, not the behavior from two years ago), and the owner (the flag's owning team chose the default; platform enforces that defaults are reviewed when a flag reaches 100%) has operated a flag system. A candidate who says "the flag service has replicas" has built a CRUD app.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Release and kill switches (server-side). The population is the company's own services: 1,800 services, 80,000 processes, every language the company uses. Flags here are mostly booleans that hide new code paths until they're ready, ramp them by percentage, and turn them off when something goes wrong. The threat is self-inflicted: a bad flip, a missing kill, an SDK that misbehaves when the service is away. The design centers on local evaluation, last-known-good, fast propagation for kills and staged propagation for everything else. Correctness bar: kill propagation p99, zero request-path dependency, and the number of incidents caused by config changes.

Experimentation and targeting (client and server). The population is users: 50 million daily actives on web and mobile, plus server-side experiments. The design centers on deterministic, sticky bucketing — the same user sees the same variant across sessions, devices and services — and on exposure logging, the event that says "user U saw variant B of experiment X at time T", which the analysis pipeline joins with outcomes. Rules must not ship to devices. The failure modes are statistical rather than operational: a salt change reshuffles every user mid-experiment; exposure events are logged on evaluation instead of on actual exposure and dilute the effect; a sample-ratio mismatch silently invalidates a month of results. Correctness bar: assignment stability and sample-ratio mismatch under 0.1%. The analytics side pairs with Ad Click Aggregator-style event pipelines and with Real-Time OLAP for results.

Dynamic operational config. Timeouts, retry budgets, pool sizes, rate-limit thresholds, routing weights, model versions. Values are typed and often structured; mistakes are rarely booleans flipped the wrong way and usually plausible numbers that are wrong at scale: a timeout raised from 2 s to 30 s that turns a slow dependency into thread-pool exhaustion, a retry count of 10 that turns a blip into a storm (see Backpressure). The design centers on schemas with bounds, validators written by the owning team, config-as-code review, and the same staged rollout. Correctness bar: every value schema-checked, every change canaried, revert in under a minute.

🎯 Staff Move: "These share the distribution and evaluation layers and almost nothing else. Release flags need speed and a kill path; experiments need statistical hygiene and exposure logs; operational config needs schemas and bounds. If I'm asked to build one system for all three, I'll share the pipe and the SDK, and give each kind its own validation and lifecycle rules."

2.2 When NOT to Build a Feature Flag Service#

SituationWhat to Do InsteadWhy
Under ~50 engineers, a handful of servicesBuy a hosted flag product, or use environment variables plus a deployYou'd spend 2–4 engineers rebuilding SDKs in five languages that a vendor already maintains
Values that change once a quarterStatic config in the repo, shipped with the deployIt gets code review, canary and rollback for free; runtime mutability only adds risk
Secrets (API keys, DB passwords)A secrets managerFlag payloads are cached on disk, logged in audit diffs and shipped widely; secrets must not be
Per-user entitlements and billing plansThe entitlement or billing system of record"Pro plan has feature X" is a product contract, not a rollout; it needs durability and joins, not flags
Authorization ("can user U edit doc D")An authorization service or policy libraryFlags are coarse and cached; permissions need correctness per request (see Auth & Identity)
Large data (ML models, blocklists of millions of entries)Artifact distribution with a flag holding only the version pointerFacebook's paper separates large-config bulk delivery from the metadata tree for the same reason
Service endpoints and membershipService RegistryMembership changes every second and needs health semantics; flags are human-paced

And within the design, some things you should not build even when you own the service:

  • Don't let flags call out during evaluation. No database or HTTP lookup inside a rule. If a rule needs data, it's an attribute the caller supplies.
  • Don't ship targeting rules to browsers or apps. They expose unreleased features, partner names and customer segments to anyone who opens dev tools.
  • Don't build a general-purpose scripting language for rules. Meta's paper explains choosing structured restraints over arbitrary code for safety and usability; expressive rules are where per-request cost and outages hide.
  • Don't make the flag service your coordination primitive. Leader election and locks belong in ZooKeeper & etcd or a Distributed Lock Service; flags are eventually consistent across the fleet by design.

The Staff signal is knowing that flags are the right tool for temporary, human-paced, runtime decisions and the wrong tool for durable product state, secrets and coordination — and that most flag-service pain comes from using it for the second list. See Build vs Buy Framework and Drill 7.

2.3 What the Interviewer Leaves Underspecified#

UnderspecifiedWhy It MattersWhat to Say
Server, client, or bothDecides local evaluation vs server-evaluated values"Both: local evaluation on servers, values-only for clients."
Kill-switch latencyDecides push vs poll"I'll target 10 s p99 for kills; 60 s worst case via poll."
Experiments in scopeAdds bucketing guarantees and exposure logging"I'll design bucketing for experiments and export exposures; the stats engine is out of scope."
Existing config systemsMigration is often the real project"Is there an existing config store I'm replacing, and how many SDKs call it?"
Regulated featuresSome flips need approval on record (payments, health, trading)"Tier-0 flags need a second approver; the audit log is retained for 7 years."
Multi-regionControl plane location; whether a region can change flags in isolation"One control plane, regional relays, and each region survives isolation on LKG."
Who may flipDecides RBAC and break-glass"Owners change; on-call can kill; nobody changes tier 0 alone."

2.4 Precise Terminology#

TermMeaningCommon Confusion
Feature flag / toggleA runtime switch that selects a code pathUsed as a synonym for config, permissions and experiments, which have different lifecycles
Release flagTemporary flag hiding unfinished or ramping codeTreated as permanent; becomes debt
Kill switch / ops flagLong-lived flag that disables a feature or dependency in an incidentWired the wrong way round, so "off" enables the dangerous path
Experiment flagFlag whose variants are assigned for statistical comparisonChanged mid-experiment, invalidating results
Dynamic configTyped runtime value (number, string, JSON)Assumed safe because it's "just a number"
ContextThe attributes evaluation runs against (user key, tenant, region, app version)Missing attributes silently match no rule
RulesetThe compiled flags and segments an SDK evaluates againstConfused with the flag store; the ruleset is a versioned snapshot
VariationOne of the possible values a flag can serveBooleans assumed; multivariate flags exist
FallthroughThe variation or rollout when no rule matchesForgotten; everyone not targeted gets an unintended value
Off variationValue served when the flag is turned off in an environmentAssumed equal to the code default; often isn't
Code defaultThe value passed at the call site, used when no ruleset is availableWritten once, never updated as production moves on
Last-known-good (LKG)The most recent valid ruleset persisted locallyAbsent in containers with ephemeral storage
BucketingDeterministic mapping of a context key to a rollout bucketImplemented with random(), so users flip variants per request
ExposureRecord that a user actually saw a variantLogged on every evaluation, including ones that didn't affect the UI
Flag debtFlags (and code branches) that no longer serve a purposeMeasured only when someone trips over one

3. Where the Design Splits#

Each fault line below follows the same shape: the options, who pays for each, the Staff default, and when to deviate.

3.1 Fault Line 1: Push vs Poll#

The tension: Push (streaming) delivers a change in seconds and costs nothing when nothing changes, but it holds tens of thousands of long-lived connections and can fail silently. Poll is stateless and self-healing, but freshness equals the interval and the load is constant whether or not anything changed.

StrategyWhat WorksWhat BreaksWho Pays
Poll only (30 s)Trivial server; every client self-heals each interval; CDN-cacheable snapshotsKill switch takes up to 30 s + snapshot build; 80K clients × 2/min = ~2,700 req/s foreverOn-call waits during incidents; infra pays idle load
Poll only (5 s)Faster kills~16K req/s of mostly "304 Not Modified"; still not instantFlag platform capacity
Push only (stream direct from origin)~1–3 s propagation; no idle load80K connections on the origin; a silently dead stream means a pod never hears the killPlatform (connection state); whoever relied on the kill
Push via relays + poll backstop (Staff default)~5 s p99; origin sees ~450 connections; 60 s poll catches dead streamsTwo code paths in the SDK; relays are a new tierPlatform team owns relays and SDK complexity
Pull from a consensus store's watch API (etcd, ZooKeeper)Ordered, versioned updates; existing infraWatch fan-out to 80K clients exceeds what these stores are designed forThe infra team running the store — see etcd & ZooKeeper
Diagram: 3.1 Fault Line 1: Push vs Poll

The arithmetic: 80,000 SDK instances streaming directly from the origin is 80,000 connections and, after any origin deploy, 80,000 simultaneous reconnects each asking for a full snapshot of a few hundred KB — up to ~20 GB of egress in seconds. With ~450 relays (three per cluster across ~150 clusters), the origin holds 450 connections and a reconnect wave is 450 snapshot fetches. Relays keep the current snapshot in memory, so SDK reconnects are served locally and never reach the origin. This is the same fan-out shape as Live Updates, with a far smaller payload.

Why the poll backstop is not optional: streams fail in ways that don't close the socket — a proxy that buffers, a relay that stopped receiving from upstream but still holds client connections. A 60-second poll with the client's current version means the worst-case staleness for any process is bounded even when every push mechanism is lying. The SDK reports which path delivered each version; flags.sdk.updates_via_poll_ratio above 1% means the stream is broken somewhere.

The Staff default: push through regional relays with a 60-second poll backstop for servers; deltas keyed by monotonic version; full snapshots in object storage for cold start. Meta's paper made the same push choice — a leader → observer → proxy tree — and argued that pull is wasteful when servers need tens of thousands of configs.

When to deviate:

  • Mobile and browsers: poll on launch, on foreground and every 15–60 minutes; reserve push (silent push notification) for emergency kills. Meta's MobileConfig did exactly this because push notification is unreliable.
  • Small fleets (< 2,000 processes): poll every 15–30 s from a CDN-cached snapshot and skip relays entirely. Fewer moving parts beats 10-second kills you will never need.
  • Edge / CDN config: push, but with staged rollout by PoP group — see Content Delivery Network.

3.2 Fault Line 2: Where Evaluation Happens — Server SDK vs Central Service vs Edge#

The tension: Evaluating locally makes every check a microsecond and removes the flag service from the request path, but it means distributing rules to every process, in every language, and trusting each SDK to implement identical semantics. Evaluating centrally keeps rules secret and semantics single-sourced, but puts a network hop and a shared dependency into every request.

StrategyWhat WorksWhat BreaksWho Pays
Remote evaluation per checkOne implementation; rules never leave the service; always fresh1–3 ms × N checks per request; flag outage = company outageEvery product's latency and availability
Remote evaluation, batched per request + short cacheOne call per request instead of NStill on the request path; cache TTL is the new staleness; cache stampede on expiryProduct latency; platform (cache tier)
Local evaluation in server SDK (Staff default for servers)~1 µs; no request-path dependency; survives control-plane lossRules shipped to every process; N language SDKs must agree byte-for-byte on bucketingPlatform (SDK maintenance in 5–8 languages)
Local evaluation via sidecar / node agentOne implementation per node; language-agnostic~50–200 µs IPC per check; another process to keep alivePlatform; latency-sensitive services pay IPC
Server-evaluated values for clients (Staff default for clients)Rules stay private; client payload small; one evaluation per sessionValues are a snapshot; context changes need re-fetchClient evaluation service capacity (~30K req/s peak)
Edge evaluation (CDN worker evaluates for the request)Values decided before origin; personalized cached pagesRules at every PoP; another evaluation implementationEdge platform team
Diagram: 3.2 Fault Line 2: Where Evaluation Happens — Server SDK vs Central Service vs Edge

The cross-SDK consistency problem. The moment there are SDKs in Go, Java, Python, Node, Swift and Kotlin, "10% of users" must mean the same 10% in all of them — otherwise a user is in the new checkout on web and the old one in the payment service, and the experiment is measuring a bug. The Staff answer is a conformance suite: a shared file of ~5,000 (ruleset, context, expected value) cases every SDK must pass in CI, including Unicode keys, missing attributes, and boundary buckets. The hashing algorithm, salt format and bucket count are written into the SDK contract and never change for an existing flag.

The Staff default: local evaluation for every server-side process, a node agent only for languages the platform won't support natively, server-evaluated value maps for untrusted clients. "The flag service is on the change path, not the request path" is the sentence.

When to deviate: serverless functions with 50 ms lifetimes can't hold a stream or load a ruleset per invocation — use a regional evaluation endpoint with a per-container cache, and accept the dependency explicitly.

3.3 Fault Line 3: Store Unreachable — Last-Known-Good vs Code Default vs Block#

The tension: When an SDK can't reach the flag service, it has three choices: serve the last ruleset it saw (possibly stale), serve the default written at the call site (possibly ancient), or wait (correct but down). Each choice moves the risk to a different victim. See the general framework in Graceful Degradation: Fail Open or Closed.

Diagram: 3.3 Fault Line 3: Store Unreachable — Last-Known-Good vs Code Default vs Block
PostureWhat WorksWhat BreaksWho Pays
Block until the service answersNever serves an unexpected valueFlag outage = no pod can start; autoscaling and deploys stopEvery service's availability
Code defaults immediatelyAlways starts; simpleDefaults drift from production reality; a fleet restart silently reverts months of rolloutsUsers (old behavior or worse); the team that wrote the default years ago
LKG from disk, then defaults after a 2 s timeout (Staff default)Restarts come up with real values; outages freeze flags rather than revert themStale during outages; needs storage that survives restarts; new nodes still hit defaultsProduct (staleness, signed off); platform (node-local cache)
LKG with maximum age (e.g., refuse LKG older than 7 days)Prevents a long-dead pod from resurrecting ancient valuesWhen the age limit trips during an outage, you fall to defaults anywaySame as above, with a sharper edge

Code defaults are a time machine. A flag created in 2023 with default false ramps to 100% in 2023 and stays there; the old path is deleted in a refactor in 2024 or, worse, isn't. In 2026 a new region comes up during a flag outage and every process there serves false — code paths that haven't run in production for two years. The Staff mitigations: LKG first; the SDK reports flags.default_divergence (flags whose production value differs from the code default observed in telemetry); a flag at 100% for 30 days must either be removed or have its code default flipped in the same change.

Kill-switch polarity matters. Write kill switches so that the code default is the safe state for that switch. A flag named recommendations.enabled with default true keeps recommendations on when flags are unreachable — fine if recommendations are healthy, wrong if you were in the middle of killing them. A kill switch is only as good as its behavior when the flag service is gone, which is exactly when an incident is most likely. For tier-0 kills, provide a break-glass override: a signed local file the node agent can write from a deploy-independent channel, which the SDK checks before the ruleset.

Who signs off: the owning team picks each flag's default and declares its polarity at creation; the platform enforces that tier-0 flags declare a safe state; incident command can use break-glass. That's a policy table, not a 3 a.m. judgment call.

🎯 Staff Move: "A flag service outage should be a 'flags are frozen' event. I'd rather serve a ruleset that's 20 minutes old than defaults that are two years old, and I'd rather serve defaults than have pods that can't start. That ordering — LKG, then defaults, never block — is the most important line in the SDK."

3.4 Fault Line 4: Per-Request Evaluation Cost — Rich Targeting vs Cheap Checks#

The tension: Product and growth teams want expressive targeting: "users in segment X who signed up after Y on app version ≥ Z, except tenants on the enterprise contract list". Every clause costs CPU on every evaluation, every segment costs memory in every process, and every attribute the rule needs must be in the context at the call site.

Targeting FeatureCost per EvaluationMemory per ProcessRisk
Boolean on/off~50 nsBytesNone
Percentage rollout~200 ns (one hash)BytesSalt or key changes reshuffle
Attribute clauses (equals, in, semver)~0.2–1 µsKBMissing attributes silently fall through
Inline segment, ≤ 10K keys (hash set)~100 ns lookup~1 MB per 10K keysLists grow; copied into every scope
Inline segment, 2M keys~100 ns lookup, but GC and load time~150–250 MB per process × 80K processesMemory blowup, slow cold start, huge deltas
Regex clauses1 µs to unbounded (backtracking)KBA catastrophic pattern pins CPU — the Cloudflare 2019 failure shape
Remote lookup in a rule0.5–3 ms + a dependency—Puts a database on every request path

The Staff default: rules are a small, structured language — ordered rules, clauses over context attributes, segments, percentage rollouts — with no network calls and only a linear-time regex engine (or none). Segments above 10,000 keys are not inline lists: they become an attribute the caller supplies (computed by the owning service), or a precomputed Bloom filter with a published false-positive rate for non-critical targeting. The admin API rejects rules whose compiled size exceeds a budget (e.g., 64 KB per flag) and SDKs report flags.eval.duration_us per flag key, so a slow flag is a named flag, not a mystery.

Who pays: product teams pay for rich targeting by computing attributes themselves and passing them in the context — which also makes the cost visible in their own service. The platform pays for the budgets and the per-flag telemetry. A flag that hits a hot key-style concentration — one flag evaluated in every request of every service — is fine locally; it's a disaster if anyone ever moves evaluation back to a remote call.

When to deviate: experimentation platforms that genuinely need million-user audiences (e.g., a holdout of 2% of users for a year) should express them as a percentage of a hash, not a list — the hash costs nothing and needs no distribution.

3.5 Fault Line 5: Who Can Flip a Flag in Production — and How Fast#

The tension: Fast, unreviewed flips are what make flags valuable during incidents and launches. Fast, unreviewed flips are also how a single engineer pushes a production change to every region without review, canary or a second pair of eyes. The deploy pipeline won't catch it; it never sees it.

PolicyWhat WorksWhat BreaksWho Pays
Anyone with access, instant, globalMaximum speedConfig becomes the largest unreviewed change path in the companyEvery user during the next bad flip
Every change reviewed and stagedSafestA kill switch that takes 45 minutes is not a kill switchIncident duration
Asymmetric by risk class (Staff default)Risk-reducing changes in seconds; risk-adding staged and approved by tierClassification must be right; some changes are bothEngineers wait ~45 min for risky changes; platform maintains classification
Config-as-code only (PR + CI + pipeline)Full review and historyToo slow for kills; UI users bypass it with shadow toolsProduct teams (friction)
Diagram: 3.5 Fault Line 5: Who Can Flip a Flag in Production — and How Fast

Classification rules, precisely: a change is risk-reducing only if it moves a flag toward its declared safe state — turning a kill switch on, turning a release flag off, lowering a rollout percentage, or reverting to the immediately previous version. Everything else is risk-adding, including "lowering" a timeout (which can cause failures) and any change to a JSON config. When in doubt the system classifies as risk-adding; the fast path is opt-in by declaration, not by guess.

Health predicates: the stage compares the treated cell against control cells on the owning service's golden signals — error rate, p99 latency, saturation — plus any metric the flag owner names. A stage fails if the treated cell is worse by more than a configured margin (e.g., error rate +0.5 percentage points or p99 +20%) for 3 consecutive minutes. Meta's paper describes the same shape: canary specs with phases, target servers, health metrics and pass predicates.

Who signs off: the platform owns the classification and the pipeline; each flag's owner owns its tier, safe state and predicates; security and compliance own which flags are tier 0 (payments, auth, data deletion). The incident commander can force any change through the fast path with a reason, and that action pages the flag's owner.

🎯 Staff Move: "The speed of a flag change should depend on which direction it moves risk. Off is fast. On is staged. That's how I keep the kill switch a kill switch without turning the flag service into the biggest unreviewed deploy path in the company."


4. When It Breaks#

4.1 A Bad Value Pushed Globally in Seconds#

t=0:       An automation job regenerates the fraud-rules config (a JSON flag) from a
           warehouse query. A schema change upstream makes the query return duplicates:
           the payload grows from 180 KB to 2.1 MB and the rule count from 900 to 11,000.
t=+4s:     Without staging: the publisher accepts it (valid JSON), relays push it,
           all 6,000 payment pods swap it in.
t=+20s:    Payment pods' rule evaluation goes from 40 µs to 3 ms per transaction.
           CPU 95%. Payment authorization p99 from 300 ms to 6 s. Timeouts at the gateway.
t=+3min:   Pods restart under memory pressure — and load the bad ruleset from LKG,
           because it was written to disk as "last known good" after it parsed.
t=+9min:   On-call identifies the config. Reverts. Pods recover as the revert arrives.

With the Staff design:
t=0:       Same regenerated payload submitted as a change request by automation.
t=+1s:     Admin API: size 2.1 MB exceeds the flag's 512 KB budget; rule count delta
           +1,122% exceeds the 50% change guard. Rejected. Automation pages its owner.
           (Had it passed: canary cell only, CPU predicate fails in 3 min, auto-revert.)
           The SDK also enforces the size limit at load and keeps the previous ruleset.

Why it was bad: the config was machine-generated, so nobody thought of it as a human change needing review; it was valid, so parsing didn't stop it; and LKG was written on parse rather than on health. The Cloudflare November 2025 post-mortem describes the same shape — a generated file that doubled in size and exceeded a consumer limit — and its first remediation was to treat internally generated config like user input.

Detection: flags.change.size_delta_ratio, flags.sdk.ruleset_bytes by scope, flags.eval.duration_us p99 by flag, and a deploy-marker style annotation on every service dashboard for every flag version change in its scope.

Prevention: per-flag size budgets and change-delta guards at the admin API and in the SDK; generated config goes through the same staged rollout as human config; LKG is promoted only after the ruleset has been live for N minutes without the process's own health degrading — a two-slot LKG (current, previous_good) instead of one.

Owner: the flag's owning team (the automation job); the flag platform owns the guards.

4.2 The SDK That Fails Closed#

t=0:       Flag control plane deploy has a bad migration. Relays lose upstream. Origin
           returns 503 for snapshots.
t=+1min:   Running pods unaffected: in-memory rulesets, ruleset_age climbing.
t=+4min:   Routine: cluster autoscaler adds 300 nodes for the evening peak; an unrelated
           deploy of the search service begins rolling 1,200 pods.
t=+5min:   SDK init in the Java SDK v3 waits for the first ruleset with no timeout
           (the "wait_for_init" option was copied from a quick-start guide).
           New search pods never become ready. Old ones are terminated by the rollout.
t=+12min:  Search capacity at 40%. p99 30 s. Autoscaler adds more pods; none become ready.
t=+25min:  Control plane fixed. 2,000 SDKs request full snapshots at once from the origin
           (relays were cold). Origin overloaded for another 8 minutes.
t=+35min:  Search recovers.

The Staff design: the SDK's init contract is serve within 2 seconds, always: LKG from a node-local cache populated by a per-node agent (so new pods on existing nodes get it), otherwise defaults after 2 s. wait_for_init without a timeout does not exist in the API. Rollouts of services are paused automatically when flags.sdk.serving_defaults rises in their scope. Relays keep the last snapshot in memory and on disk, so a relay restart doesn't need the origin.

Detection: flags.sdk.init_duration_ms p99, flags.sdk.serving_defaults count, flags.relay.upstream_connected, readiness failures correlated with SDK version.

Owner: flag platform (SDK contract and relays); service owners for SDK version currency — the platform enforces a minimum SDK version in CI.

4.3 Stale Flags and the Reused Key#

t=-3 years: payments.route_v2 created, ramped to 100%. Old route code left in place.
t=-1 year:  Flag "removed" from the UI. Code still reads it with default false.
t=0:        New team creates payments.route_v2 for an unrelated feature (same key —
            the UI allows it because the old one was deleted). Sets it to false
            for most tenants while they build.
t=+1min:    Services still on the old binary read payments.route_v2=false and take the
            three-year-old route. It calls a processor endpoint decommissioned last month.
t=+6min:    Payment failures 4% for tenants on those services. The new team sees nothing:
            "our flag isn't even used yet".

This is the Knight Capital shape at small scale: a flag whose meaning changed while some binaries still carried the old meaning.

Prevention: flag keys are never reused — deletion leaves a tombstone forever; deleting a flag requires the SDK usage telemetry to show zero evaluations for 14 days; flags past expiry open removal tickets; a flag at 100% for 30 days opens a ticket to delete the losing code path; and code-reference scanning in CI fails a build that references a tombstoned key.

Detection: flags.evaluations_of_unknown_key (SDKs report keys they're asked for that aren't in the ruleset), flags.stale_count per team, flags.default_divergence.

Owner: owning team for each flag; platform for the tombstones, scanning and the debt dashboard.

4.4 Reconnect Storm After a Control-Plane Deploy#

t=0:       Relay fleet deploy restarts relays in one region without connection draining.
t=+2s:     22,000 SDKs reconnect within the same second, each requesting a full snapshot
           because their delta cursor isn't accepted by a fresh relay. Relays fetch from
           the origin: 150 relays × 30 MB environment snapshot = 4.5 GB in seconds.
t=+10s:    Origin egress saturated; snapshot requests time out; SDKs retry with no jitter.
t=+1min:   Retry rate 3× the initial wave. Other regions' relays can't reach the origin.

Staff mitigations: relays persist their snapshot to local disk and serve SDK deltas from it after restart; deltas are accepted for any version within the last 24 hours; SDK reconnects use exponential backoff with full jitter (0–30 s) since running SDKs lose nothing by waiting; relay deploys drain connections at a bounded rate (e.g., 500/s per relay); the origin sheds snapshot requests with Retry-After instead of queueing. These are standard Backpressure tools applied to the distribution tier.

Detection: flags.relay.connections rate of change, flags.origin.snapshot_requests_per_s, flags.sdk.reconnects by region.

Owner: flag platform.

4.5 The Kill Switch That Couldn't Reach the Fleet#

A regional network event isolates the flag control plane from half the fleet at the same moment a dependency starts failing. On-call flips the kill switch; only hosts that can reach a relay hear it. This is described from the consumer side in Degraded Mode Framework; the flag-service-side design answers are: relays in every region that keep serving the last snapshot and accept break-glass overrides injected locally; a signed override file distributed through the host agent's own channel (independent of the flag service) for tier-0 kills; and a measured flags.kill.propagation_p99 from a synthetic kill flipped every 10 minutes in each region against a canary service.

4.6 Targeting Segment Blowup#

A growth team pastes 1.8 million user IDs into a segment for a promotion. The segment is referenced by a flag in the web-frontend scope. Each frontend process now loads ~200 MB more, cold start goes from 4 s to 40 s, and GC pauses add 300 ms to p99. The admin API should have refused it: inline segments are capped at 10,000 keys and per-scope ruleset size at a budget (e.g., 5 MB); large audiences become attributes computed by the owning service or percentage holdouts.

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Bad value pushedflags.change.size_delta_ratio; owning service golden signals in canary cellCanary cell (~1%) if staged; whole fleet if notAuto-revert via fast path; size and delta guards; two-slot LKGFlag owner + platform
SDK blocks startupflags.sdk.init_duration_ms p99 > 2 s; readiness failuresEvery service that restarts during the outage2 s init contract; node-local LKG; pause rolloutsPlatform (SDK)
Serving ancient defaultsflags.sdk.serving_defaults > 0; flags.default_divergenceNew nodes and regions during outagesLKG first; flip defaults when flags reach 100%Flag owner
Reused / stale flagflags.evaluations_of_unknown_key; flags.stale_countBinaries holding the old meaningTombstones; usage-gated deletion; CI scanningFlag owner + platform
Reconnect stormflags.origin.snapshot_requests_per_s; relay connection churnDistribution freezes fleet-wideRelay disk snapshots; jittered backoff; drained deploysPlatform
Kill doesn't propagateflags.kill.propagation_p99 synthetic > 10 sIncident lasts longerRegional relays; break-glass override filePlatform + incident command
Segment / rule blowupflags.sdk.ruleset_bytes; flags.eval.duration_us by flagMemory and latency for every process in the scopeSize budgets; segment caps; linear-time regexFlag owner + platform
Experiment reshuffleSample-ratio mismatch alert; assignment-change rateWeeks of experiment resultsImmutable salt per flag; bucketing conformance suiteExperimentation team
Control-plane outageAdmin API availability; flags.sdk.ruleset_age_secondsNo changes possible; evaluation unaffectedSeparate failure domain; break-glassPlatform

🎯 Staff Insight: The dangerous flag failures don't page. A pod serving two-year-old defaults returns 200s. A flag nobody removed works fine until somebody reuses its key. flags.sdk.serving_defaults, flags.default_divergence and flags.evaluations_of_unknown_key are the three metrics that turn silent flag failures into tickets — make them first-class from day one.


5. Scorecard#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingLists features: UI, targeting, rollouts, auditNames release / experiment / ops-config intents; commits; states "change path, not request path"Asks how many runtime-config systems exist and what share of incidents started with a config change
EvaluationRemote API with a cacheLocal SDK evaluation for servers, values-only for clients; conformance suite across languagesOwns the SDK contract as a company standard; decides which languages get native SDKs vs a node agent
FailureReplicas and failoverLKG → defaults → never block; 2 s init contract; break-glass for tier-0 killsFlag control plane as its own failure domain; quarterly "flag service off" game day; fail-static contract published
Change safetyAudit log and rollbackAsymmetric change speed; tier-based approval; staged by cell with health predicates and auto-revertStaged rollout mandatory for every runtime-config system; config-caused incidents as an org metric
Lifecycle"Clean up old flags"Owner, kind, expiry required; tombstoned keys; usage-gated deletion; stale count per teamFlag-debt budgets per team; removal in the definition of done; debt reported alongside reliability
Operations"Add monitoring"flags.kill.propagation_p99, flags.sdk.serving_defaults, flags.default_divergence, per-flag eval costError-budget accounting that attributes incidents to config changes across all config systems

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Takes the flag service off the request path"Evaluation is local; the flag service is on the change path, not the request path."
Designs the client before the server"First question: what does the SDK do when it can't reach us? LKG, then defaults, never block."
Treats propagation speed as two-edged"Kills reach the fleet in 10 seconds. Everything else is staged, because a bad value would too."
Knows defaults rot"A code default written two years ago is the most dangerous value in the system."
Separates rules from values for clients"Phones get values, not rules — rules leak our roadmap."
Measures debt"35% of Meta's configs hadn't changed in 300 days. I'd measure ours and give each team a budget."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Each request calls the flag service; Redis makes it fast"Puts a shared dependency on every request path in the company
"Changes go out instantly everywhere" as a selling pointNo notion that propagation speed is also blast-radius speed
No answer for "the flag service is down and pods are restarting"Hasn't thought about the only failure that matters
Random-number percentage rolloutsUsers flip variants per request; experiments are invalid
Shipping the full ruleset to mobile appsLeaks unreleased features and customer segments
No owner or expiry on flagsBuilds the debt that Knight Capital paid for

5.4 Common False Positives#

  • Targeting-DSL fluency ≠ flag-service design. A candidate who designs a beautiful rule language with 30 operators has usually designed the per-request cost problem, not solved it.
  • "We use LaunchDarkly / Unleash / an in-house tool" ≠ understanding it. The Staff question is what that SDK does with no connection on a cold start.
  • Consensus-store knowledge ≠ distribution design. Knowing etcd watches is good; proposing 80,000 watchers on a 3-node etcd cluster is not.
  • "Everything is audited" ≠ change safety. An audit log tells you who broke production; staging decides how much of it broke.

6. The 45 Minutes, Phase by Phase#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minPick intents; "change path, not request path"; kill-switch latency; scale
Entities & API3–5 minFlag, environment config, segment, change request with risk_class, ruleset version
Architecture5–10 min≤ 8 boxes; control plane, distribution, data plane
Unreachable store10–18 minLKG, defaults, init timeout, polarity, break-glass
Propagation & change safety18–26 minPush + poll via relays; asymmetric speed; staged rollout; predicates
Eval cost + debt26–34 minRule budgets, segment caps; owner, expiry, tombstones
Pivot (interviewer's choice)34–42 minExperiments, multi-region, build vs buy, migration, mobile
Wrap42–45 minTwo planes; asymmetric speed; metrics and owners

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"Make it work for A/B tests"Statistical hygieneImmutable salt, sticky bucketing, exposure on actual exposure, SRM alert
"A bad flag took down checkout"Change safetyStaged rollout, predicates, auto-revert, size and delta guards
"Our mobile app is 40% of revenue"Client distributionValues not rules; launch + foreground + 60 min poll; silent push for kills; cached on device
"Go multi-region"Control plane placementOne writer, regional relays, each region survives isolation on LKG; staging by region
"Should we just buy this?"Build vs buyBuy below ~200 engineers; own the SDK contract via OpenFeature; keep break-glass in-house
"We have three config systems"Migration and standardsOne SDK contract first; backends converge; dual-read with diff telemetry

6.3 What to Deliberately Skip#

  • The admin UI. "UI and config-as-code both produce change requests."
  • Every rule operator. "Equals, in, semver, numeric compare, segment, percentage."
  • Database choice. "30,000 flags with history fit in PostgreSQL; the store isn't the hard part."
  • Experiment statistics. "We export exposures; the analytics team owns the stats engine."
  • Authentication of the admin API. "SSO plus RBAC by team and tier."

6.4 Follow-Up Questions to Expect#

  1. "The flag service is down for 30 minutes. What happens to running pods, restarting pods and a new region?"
  2. "How do you guarantee 10% → 20% keeps the original 10%?"
  3. "Someone pushes timeout_ms = 30000 to every service. What stops it?"
  4. "How do you kill a feature in under 10 seconds when the flag service is degraded?"
  5. "How do you know how many of your 30,000 flags are dead?"
  6. "What's in the payload the mobile app receives, and what's not?"
  7. "Who can flip a tier-0 flag at 3 a.m., and how do you audit it?"

7. Practice Rounds#

Drill 1: The Opening#

Prompt: "Design a feature flag system for our company."

Staff Answer

"Before drawing — is this for server-side releases and kill switches, client-side experiments, operational config, or all three? They share distribution and evaluation but need different safety rules. I'll start with server-side release and kill switches, add typed operational config on the same pipe, and treat client experiments as server-evaluated values.

The constraints I'll commit to: evaluation never needs the flag service to be up — it's on the change path, not the request path; kill switches reach the fleet in under 10 seconds p99; every other change is validated, approved by tier and staged by cell with automatic revert; and every flag has an owner and an expiry. Scale: ~1,800 services, ~80,000 processes, ~30,000 flags, 2 million requests a second. I'll go: entities → control plane, distribution, data plane → what the SDK does when we're down → how changes propagate and how bad ones are stopped → evaluation cost → flag debt and ownership."

Why this is L6:

  • Distinguishes intents and commits
  • States the availability property (change path, not request path) before any boxes
  • Treats change speed as asymmetric from the start

What L7 adds:

  • Asks how many config systems exist and how many incidents started with a config change last year
  • Frames the flag service as the company's largest control dependency with its own failure domain
  • Asks who in the org can make production changes without a deploy today
❌ Common L5 Trap

"A flags table in Postgres, an API to read and write flags, a Redis cache in front, a UI for product managers, and services call the API to check flags. We'll add percentage rollouts and audit logging."

Why this misses: Every part works and the design puts Postgres-behind-Redis on every request path in the company, pushes every change globally at once, and has no answer for a pod that restarts while the API is down.


Drill 2: Core Mechanic — Sticky Percentage Rollouts#

Prompt: "Explain exactly how a user ends up in the 10% that gets the new feature, and what happens when we go to 20%."

Staff Answer

"Bucket = sha1(flag.salt + '.' + context.key), take the first 32 bits mod 100,000. The variation weights are cumulative ranges over those buckets: treatment is [0, 10,000) at 10%. Going to 20% extends the range to [0, 20,000), so everyone in the first 10% stays in — rollouts are monotonic. The salt is generated once per flag and never changes; changing it reshuffles everyone. The context key is chosen per flag — user ID for user-facing features, tenant ID when the whole tenant must see the same behavior, device ID for logged-out traffic — and it can't change mid-rollout either.

Every SDK must produce identical buckets, so the algorithm, byte encoding and bucket count are in the SDK contract and a conformance suite of thousands of cases runs in each SDK's CI."

Why this is L6:

  • Deterministic and monotonic, with the reason
  • Names the salt and context-key choice as one-way decisions per flag
  • Cross-language consistency as a tested contract

What L7 adds:

  • Uses a separate salt namespace for experiments so a user's assignment in one experiment is uncorrelated with another
  • Defines a company-wide holdout (e.g., 1% of users excluded from all launches for a year) by hash range, not by list
  • Owns the bucketing spec as a one-way door and versions it explicitly

Drill 3: Make It Concrete — Capacity#

Prompt: "Size the system. Connections, bandwidth, CPU."

Staff Answer

"Evaluation: 2M requests/s × ~40 checks = 80M evaluations/s, ~1 µs each, ~80 cores fleet-wide spread across every service — noise, but only because it's local.

Distribution: 80,000 SDK instances. Through ~450 relays (3 per cluster × 150 clusters), the origin holds ~450 streams. ~4,000 flag changes/day average to ~0.05/s, but bursts of a few per second during business hours. A delta is ~1–5 KB, so steady-state bandwidth is trivial. The expensive event is cold start: a full environment snapshot is ~30 MB; scoped per service it's ~300 KB median. Fleet-wide restart = 80,000 × 300 KB = 24 GB, served by relays from memory, not the origin. Poll backstop: 80,000 / 60 s = ~1,300 req/s of mostly 304s at the relays.

Client evaluation: 50M DAU × ~3 fetches/day = 150M/day, ~1,700/s average, ~30K/s at peak with app-launch spikes; each evaluates ~300 client-visible flags in ~300 µs → ~10 cores at peak plus headroom. Payload ~3 KB.

Store: 30,000 flags × ~2 KB + history ~4,000 changes/day × 2 KB = ~3 GB/year. One PostgreSQL primary with a replica."

Why this is L6:

  • Sizes evaluation, distribution and client paths separately
  • Spots that cold start, not steady state, sizes distribution
  • Puts the storage tier in proportion: not the problem

What L7 adds:

  • Prices a vendor alternative per seat / per connection / per MAU against ~6 engineers in-house
  • Notices that SDK count grows with the fleet, not with flags, so cost scales with infra, not product
  • Plans relay capacity per cell so a region can be isolated without losing distribution

Drill 4: The Dependency Goes Down#

Prompt: "The flag service's database is corrupted and the control plane is down for two hours. Walk me through the company's experience."

Staff Answer

"Running processes: unaffected; they hold their ruleset in memory. flags.sdk.ruleset_age_seconds climbs and pages the flag platform, not every service. Restarting processes: load LKG from the node-local cache in ~20 ms; deploys and autoscaling work. New nodes: the node agent seeds LKG from the regional relay, which still has the last snapshot in memory and on disk. Only a brand-new cluster with no relay snapshot falls back to code defaults after 2 s, and we'd pause bringing up new clusters during the outage.

What we lose: the ability to change flags. Kill switches for tier-0 features still work through break-glass — a signed override injected through the node agent's channel, which doesn't touch the control plane. Meanwhile we restore the database from the point-in-time backup; the rulesets in relays are the source of truth for 'what's live' and we diff the restored store against them before re-enabling publishing, so we don't publish an older version over a newer one."

Why this is L6:

  • Walks each class of process separately
  • Keeps kill capability through a path independent of the control plane
  • Thinks about restore safety — monotonic versions, diff before publish

What L7 adds:

  • Makes "flag service off for two hours" a scheduled game day so this answer is tested, not hoped
  • Puts the control plane in its own failure domain: separate database, cluster and pipeline from the services it controls
  • Publishes the fail-static contract so every product team knows what "flags frozen" means for them

Drill 5: Hot Key — The Flag Every Request Checks#

Prompt: "One flag — global.maintenance_mode — is checked by every request in every service. Is that a hot key problem?"

Staff Answer

"Locally, no: it's a boolean in memory, 50 ns, evaluated 2 million times a second across 80,000 processes with no shared state. It would be the worst hot key in the company the day anyone moved evaluation to a remote call or a shared cache — 2M req/s on one key. So the design rule is that evaluation is always local, and this flag's real risk is different: changing it affects everything at once. It's tier 0, its safe state is false, turning it on is risk-adding and needs two approvers unless incident command invokes it, and its change is staged by region even in an incident, because maintenance mode on in every region simultaneously is itself an outage."

Why this is L6:

  • Distinguishes read hotness (irrelevant locally) from change blast radius (the real risk)
  • Applies tiering and staging to a global flag
  • Names the scenario where it becomes a hot key

What L7 adds:

  • Questions whether a global maintenance flag should exist at all versus per-cell switches
  • Requires that any flag referenced by > 50% of services is reviewed as infrastructure, not product config

Drill 6: Multi-Tenant B2B — Per-Customer Flags#

Prompt: "We're B2B with 20,000 tenants. Enterprise customers want features enabled for them only, sometimes for years."

Staff Answer

"First, separate rollouts from entitlements. 'Tenant X is on the Enterprise plan and has SSO' is a product contract; it lives in the entitlement system with the billing record, and services check entitlements. Flags are for temporary rollout: 'new reporting engine for 50 design-partner tenants'. If a flag targets a tenant list for more than its expiry, that's a signal it's become an entitlement and should migrate.

Mechanically, context.kind = tenant and the bucketing key is the tenant ID, so all users in a tenant see the same behavior. Tenant lists above 10,000 entries aren't inline; they're an attribute the request carries from the tenant record. And per-tenant flag changes are still staged: a change for one large tenant is canaried on that tenant's traffic in one cell first."

Why this is L6:

  • Draws the entitlements vs flags boundary
  • Picks the right bucketing key for B2B
  • Keeps segment size bounded

What L7 adds:

  • Gives sales and support a governed path for per-customer exceptions with expiry, instead of engineers adding tenant IDs by hand
  • Tracks "long-lived tenant-targeted flags" as a product-debt metric that feeds the entitlement roadmap

Drill 7: Build vs Buy#

Prompt: "Should we build this or buy a vendor product?"

Staff Answer

"Under ~200 engineers, buy. A vendor gives SDKs in every language, relays, an audit trail and an experimentation UI; building that is 4–6 engineers for a year plus an on-call. Between ~200 and ~2,000 engineers, still usually buy — but verify the SDK's behavior with no connection and no cache, require a relay you can run inside your network so evaluation never depends on the vendor's uptime, and check that the vendor supports staged changes and approvals. Above ~2,000 engineers, or when flags are deeply tied to an existing in-house config system as at Meta, building becomes defensible.

Either way, I'd put the OpenFeature API — a CNCF project that standardizes the flag-evaluation interface and lets backends plug in as providers — between our code and the vendor (OpenFeature specification), and keep break-glass kill switches in-house. Then a vendor change is a provider swap, not 40,000 call-site edits."

Why this is L6:

  • Gives thresholds, not "it depends"
  • Names the non-negotiable vendor requirements (offline behavior, in-network relay)
  • Protects reversibility with a standard interface

What L7 adds:

  • Prices it: vendor fees vs headcount vs the cost of a vendor outage freezing changes during your incident
  • Decides which capability must stay in-house (break-glass, ops config for tier-0 infra) regardless
  • See Buy or Build: The Total-Cost Test

Drill 8: Changing Policy Without an Outage — Introducing Staged Rollout#

Prompt: "Today any engineer can flip any flag globally. You want staged rollouts and approvals. How do you introduce that without a revolt or an outage?"

Staff Answer

"Shadow → warn → enforce, by tier. Month 1: classify every change as risk-adding or risk-reducing and log what would have required approval or staging — no behavior change. Publish the data: how many changes per week, how many incidents were preceded by a flag change. Month 2: tier-0 flags (payments, auth, data deletion — about 300 flags) get staged rollout enforced, with the fast path untouched for kills. Warn on everything else. Month 3: tier-1 enforced. Tier 3 stays self-serve forever with staging but no approval.

The key concession: the fast path for risk-reducing changes is never slowed. If engineers believe the new system slows down incident response, they'll build shadow flags in environment variables."

Why this is L6:

  • Uses shadow mode to gather evidence before enforcing
  • Starts with the highest-risk tier
  • Protects the kill path explicitly

What L7 adds:

  • Extends the same policy to every runtime-config path (WAF rules, routing, ML model pushes), not only flags
  • Reports config-caused incidents before and after as the success metric to leadership

Drill 9: Multi-Region#

Prompt: "We're in 6 regions. Design the flag service for multi-region."

Staff Answer

"One logical control plane with a single writer — flag changes are human-paced, ~0.05/s, so there's no reason for multi-writer conflict resolution. Its store is replicated to a standby region with failover measured in minutes, which is acceptable because evaluation doesn't need it (see Multi-Region Active-Active for why I wouldn't pay for active-active here). Each region has its own relays and snapshot cache, so a region isolated from the control plane keeps serving LKG. Staged rollouts include region as a stage dimension: canary cell in one region, then the rest of that region, then other regions one at a time. A flag can be scoped to a region for regional kills. Break-glass works per region without the control plane."

Why this is L6:

  • Rejects multi-writer complexity with a reason (write rate)
  • Isolated regions keep working
  • Regions are a staging dimension

What L7 adds:

  • Sets the order of regional waves company-wide (smallest-revenue region first) in coordination with the Deployment System
  • Considers data-residency: context attributes sent to the client evaluation service may be personal data and must stay in-region

Drill 10: Cost#

Prompt: "Finance asks what this costs and whether it's worth it."

Staff Answer

"Infrastructure is small: a PostgreSQL pair, ~450 relays at ~0.5 vCPU each, a client evaluation service at ~30K req/s peak, snapshot storage — roughly $15–30K/month at our scale. The real cost is people: ~5–6 engineers for the platform and SDKs, plus an on-call rotation. The value side: if config-related changes cause a meaningful share of our high-impact incidents — Meta measured 16% over one quarter — then staged rollout and fast kills that shorten even a few incidents a quarter pay for the team. I'd measure config-caused incident minutes before and after, and report that number."

Why this is L6:

  • Separates infra cost from headcount
  • Ties value to a measurable reliability outcome
  • Uses a public number as a prior, then commits to measuring our own

What L7 adds:

  • Compares against vendor pricing and the cost of flag debt (engineer hours lost to dead branches)
  • Frames the platform as reducing every team's on-call load, and asks for that to be counted

8. Incident Walkthroughs#

Deep Dive 1: Peak-Traffic Incident — Black Friday Kill Switch Takes 11 Minutes#

Context: During the evening peak, the recommendations service starts timing out and dragging checkout p99 to 4 s. On-call flips recs.enabled off. Half the checkout fleet stops calling recommendations in 5 seconds; the other half keeps calling for 11 minutes. The incident commander escalates to you.

Questions to Surface First:

  • Which hosts got the change late — by region, cluster, SDK version, or relay?
  • Did the stream fail silently, and did the poll backstop work?
  • Was the flag service itself under load from the same peak?

Typical L5 Approach: Restarts the slow checkout pods so they pick up the new value. The restart wave hits the control plane for snapshots, slowing the propagation further.

Staff Approach: Checks flags.sdk.version by host: the slow half is behind two relays that lost upstream during a relay autoscale but kept client streams open. The SDKs' poll backstop was configured at 10 minutes in an old SDK version. Fixes by restarting the two relays (SDKs reconnect to healthy relays with jitter), and pushes an SDK config to cut poll to 60 s.

Principal Approach: Treats kill latency as an SLO with a synthetic probe per region, makes minimum SDK version an enforced platform requirement, and adds "kill a non-critical feature at peak" to the pre-peak game day for every tier-0 service.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Identify lagging hosts via version telemetry; restart the two stale relays; confirm the kill arrives
TriageRelays held client streams without upstream for 11 min; old SDK poll interval 10 min
Quick fixRelay health check includes upstream freshness; relays close client streams when upstream is stale > 30 s
GuardrailsSynthetic kill every 10 min per region; page if flags.kill.propagation_p99 > 10 s
Post-mortemWhy could a relay be "healthy" while serving stale data? Why were SDKs older than 12 months allowed?

Metrics to Watch: flags.kill.propagation_p99, flags.relay.upstream_lag_seconds, flags.sdk.updates_via_poll_ratio, SDK version distribution

Organizational Follow-up: minimum SDK version enforced in CI; relays' readiness tied to upstream freshness.

Ownership Question: "Who owns kill latency — the flag platform or the service that needed the kill?" Staff answer: The flag platform owns propagation and the synthetic probe; service owners own having a kill switch with a safe polarity. The probe measures the end-to-end result, and the platform carries the page.

Key Takeaway: "A kill switch has a latency, and an unmeasured latency is a guess."

What clears the Staff bar:

  • Uses version telemetry to localize the lag instead of restarting blindly
  • Finds the silent-stream failure and the weak backstop
  • Turns kill latency into a measured SLO

Deep Dive 2: Silent Failure — A Region Has Been Serving Defaults for Six Days#

Context: A product manager notices that an experiment's results in one region look nothing like the others. Investigation shows the region's new Kubernetes cluster, brought up six days ago, has every pod serving code defaults: its relay was never configured, and the SDK's init timeout fell back to defaults silently.

Questions to Surface First:

  • Which flags have production values different from their code defaults? What did users in that region experience?
  • Why didn't flags.sdk.serving_defaults page?
  • Were any kill switches active elsewhere that this region didn't honor?

Typical L5 Approach: Configures the relay, confirms values arrive, closes the ticket.

Staff Approach: Fixes the relay, then quantifies impact: 214 flags had divergent defaults, including two kill switches active elsewhere and a release flag whose default pointed at a deprecated pricing path. Makes serving_defaults > 0 for 5 minutes a paging alert and adds "relay healthy and ruleset age < 2 min" to the cluster bring-up checklist, enforced by automation.

Principal Approach: Recognizes defaults drift as systemic: requires the default to be updated (or the flag removed) when a flag reaches 100%, tracks flags.default_divergence per team, and makes "new region/cluster serves real flags" a gate in the platform's cluster provisioning, not a checklist line.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Configure relay; verify ruleset age drops; confirm kill switches now honored
Triage214 divergent flags; 2 active kills not honored; 1 pricing path deprecated
Quick fixReview pricing-path transactions with finance; invalidate the region's experiment data for 6 days
GuardrailsPage on serving defaults; provisioning gate; divergence report per team weekly
Post-mortemWhy was "defaults" a silent, normal-looking state? Why did 214 flags have stale defaults?

Metrics to Watch: flags.sdk.serving_defaults, flags.default_divergence, flags.sdk.ruleset_age_seconds by cluster

Ownership Question: "Who fixes the 214 stale defaults?" Staff answer: Each flag's owning team, with tickets generated by the platform from divergence telemetry and a 30-day SLA; the platform publishes the count per team.

Key Takeaway: "Serving defaults is a failure state, not a fallback state. Alert on it."

What clears the Staff bar:

  • Quantifies divergence instead of assuming defaults are fine
  • Turns a silent state into a paging signal
  • Fixes provisioning so the class of failure can't recur

Deep Dive 3: Large-Customer Onboarding — A Bank Requires Change Approval on Record#

Context: A bank is signing a contract that requires every production change affecting its tenant to have documented approval, a 7-year audit trail and the ability to freeze changes during its quarter-end. Sales asks if the flag service can comply.

Questions to Surface First:

  • Which flags affect this tenant — tenant-targeted only, or every global flag too?
  • What counts as a "change": a percentage widening that newly includes them?
  • Does a freeze block kill switches?

Typical L5 Approach: Adds an approval checkbox to the UI for flags that target the bank.

Staff Approach: Computes affected flags per tenant (any change whose evaluated value for any of the tenant's contexts would differ), routes those changes through a tenant-aware approval step, retains the audit log 7 years in write-once storage, and implements change freezes as a policy in the admin API — freezes block risk-adding changes only; kills are always allowed and audited.

Principal Approach: Generalizes into a compliance tier offered to every regulated customer, priced into the enterprise plan, and negotiates contract language that explicitly exempts risk-reducing changes — because a contract that blocks your kill switch is a reliability risk you've signed.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateInventory which flags evaluate differently for the bank's contexts (dry-run evaluation over sample contexts)
DesignTenant-aware approval policy; freeze windows; risk-reducing exemption
ImplementationWORM audit retention; per-tenant change report exportable monthly
GuardrailsFreeze-violation attempts logged and reported; kill usage reviewed post-hoc
ReviewLegal reviews the exemption clause before signature

Ownership Question: "Who approves a flag change that affects the bank?" Staff answer: The flag owner plus the bank's designated account engineer for tier-0 flags; the platform enforces it, and the audit export goes to the bank monthly.

Key Takeaway: "Compliance requirements belong in the change policy engine, not in a UI checkbox — and never block the kill path."

What clears the Staff bar:

  • Defines "affects the tenant" by evaluation, not by targeting syntax
  • Keeps kills outside every freeze
  • Builds it once as policy for all regulated tenants

Deep Dive 4: Post-Mortem — A Valid Config Change Took Down Payments#

Context: An engineer changed payments.retry.max_attempts from 3 to 8 to reduce card-decline false negatives. Validation passed (allowed range 1–10). It went out globally via an old "urgent" path that skipped staging. During a processor slowdown an hour later, retries amplified traffic 8×, the processor rate-limited the company, and payments failed for 24 minutes.

Questions to Surface First:

  • Why did an "urgent" path exist for a risk-adding change?
  • Why did validation allow a value that was only safe under normal conditions?
  • What would staging have shown — would a canary cell have caught a problem that only appears under dependency stress?

Typical L5 Approach: Lowers the allowed range to 1–5 and reverts the value.

Staff Approach: Removes the urgent bypass for risk-adding changes; reclassifies retry, timeout and concurrency configs as tier 0 requiring owner approval; adds a predicate on downstream call volume to the canary stage; and pairs retry settings with a retry budget (e.g., retries ≤ 10% of requests) so no single number can amplify unbounded — see Backpressure.

Principal Approach: Notes this is Meta's "valid config exposing a latent problem" category (22% in their study) and that validators can't catch it; invests in canaries that include synthetic dependency stress, and makes "amplification-capable config" (retries, fan-out, batch sizes) a named class with stricter review across every config system.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Revert via fast path; confirm processor rate limit clears
TriageChange was valid and went global via a bypass; retries amplified 8× under stress
Quick fixRemove bypass; retry budget in the payments client
GuardrailsTier-0 class for amplification configs; downstream-volume predicate in canary
Post-mortemWhy did the bypass exist? Who approved its creation? What else uses it?

Ownership Question: "Who owns the retry count — payments or the platform?" Staff answer: Payments owns the value and its predicates; the platform owns the rule that it can't skip staging and the classification of amplification configs.

Key Takeaway: "Validators catch invalid values. Staging catches valid values that are wrong in production. You need both."

What clears the Staff bar:

  • Recognizes a valid value can still be an outage
  • Removes the bypass, not just the value
  • Adds structural limits (retry budgets) instead of tighter ranges alone

Deep Dive 5: Multi-Region Expansion — Flags in a Sovereign Region#

Context: The company is launching in a sovereign-cloud region with a network that may be isolated from the main control plane for days, and a rule that no user data leaves the region.

Questions to Surface First:

  • Can flag changes flow in from the main control plane? Must some be made locally?
  • What context attributes does the client evaluation service receive, and are they personal data?
  • How are kills handled during isolation?

Typical L5 Approach: Deploys a second, independent flag service in the region.

Staff Approach: Runs a regional relay tier and client evaluation service inside the region; the main control plane pushes rulesets in (rules contain no user data, segments with personal IDs are excluded or resolved in-region); exposure events stay in-region; during isolation the region runs on LKG with a local break-glass for kills operated by in-region on-call.

Principal Approach: Defines a "region-sovereign flag" class whose authoritative control plane lives in the region, decides which flags must be globally consistent vs regionally controlled, and puts that boundary in the standard so every new sovereign region is a configuration, not a project.

Staff Approach — Full Reasoning
PhaseWhat to Do
DesignIn-region relays, client eval, exposure pipeline; rulesets flow in, data never flows out
IsolationLKG for days; ruleset age alert tuned to expected isolation; local break-glass
SegmentsPersonal-ID segments resolved in-region or replaced with attributes
StagingRegion is its own wave, always last for risk-adding global changes
AuditIn-region audit log of break-glass actions, reconciled when connectivity returns

Ownership Question: "Who can change a flag in the sovereign region during isolation?" Staff answer: In-region on-call, kills only, via break-glass; everything else waits for reconnection. That's written into the region's runbook and signed by the region's compliance owner.

Key Takeaway: "Rules can travel; user data can't. Design the flag payload so it never needs to."

What clears the Staff bar:

  • Keeps personal data out of rules and in-region
  • Plans for long isolation on LKG
  • Limits local authority to kills

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Explain why the flag service belongs on the change path and never the request path, with the evaluation arithmetic
  • Design SDK behavior when the service is unreachable: LKG, then defaults, never block, with a 2-second init contract
  • Choose push, poll or both for servers and for clients, and size the relay tier
  • Implement deterministic, monotonic percentage rollouts consistent across languages
  • Make change speed asymmetric: kills in seconds, risk-adding changes staged with predicates and auto-revert
  • Bound per-request evaluation cost with rule budgets, segment caps and no network calls
  • Run flag lifecycle: owner, kind, expiry, tombstones, usage-gated deletion, debt metrics
  • Decide build vs buy and keep it reversible through a standard SDK interface

The Bar for This Question#

Mid-level (L4): Builds a flags table, an API, a cache and a UI with on/off and a percentage. Services call the API. Works in the demo; has no answer for outages, bad values or flags that never go away.

Senior (L5): Adds local caching in the SDK, audit logs, targeting rules, and replicas. Knows remote calls are slow. The gap: a cache that expires is still a request-path dependency, changes still go global at once, defaults are assumed safe, and flag cleanup is "process". The design would pass review and would still turn a flag-service deploy into a company-wide startup failure.

Staff+ (L6): Separates control plane from data plane in the first five minutes. Evaluates locally, keeps last-known-good on disk, never blocks startup. Makes change speed asymmetric with staged rollout and auto-revert, bounds evaluation cost, and treats flags as owned, expiring inventory. Names who pays: product accepts staleness during outages and minutes of staging, the platform owns propagation and SDK contracts, owners own defaults and polarity. The interviewer should learn something from the answer.


10. Hot Takes#

10.1 Fast Config Propagation Is a Liability You Opt Into#

ClaimReality
"Changes reach every server in 2 seconds"So does a bad change — Cloudflare 2019, Google Cloud 2025
"Flags are safer than deploys"Only if they get a deploy's safety: staging, canary, revert
"Validation catches bad values"22% of Meta's config incidents were valid changes exposing code bugs

The Staff position: Global-in-seconds is a privilege reserved for changes that reduce risk. Everything else earns its way out in stages.

Why this matters in interviews: Candidates brag about propagation speed. Interviewers hear blast-radius speed.

10.2 The Code Default Is the Most Dangerous Value in Your System#

The Staff position: It's the value nobody has seen in production for longest, and it's served exactly when things are already going wrong — outages, new regions, cold starts. Prefer last-known-good, alert on serving defaults, and update the default when a flag reaches 100%.

Why this matters in interviews: "We fall back to the default" sounds safe. Asking "how old is the default?" is a Staff question.

10.3 Most Flags Should Die Within 90 Days#

Flag KindExpected LifetimeDefault Expiry
ReleaseWeeks90 days
ExperimentWeeks to a quarter60 days after the decision
Ops / kill switchYearsAnnual review
Permission / entitlementShouldn't be a flagMigrate out

The Staff position: A release flag older than a quarter isn't a release tool; it's a permanent untested branch. Measure the inventory, give teams budgets, and make removal part of done.

Why this matters in interviews: Bringing up flag debt unprompted signals you've lived with a flag system for years, not weeks.

10.4 Don't Build Experimentation Into the Flag Service#

The Staff position: Share the bucketing and the SDK; export exposures. Statistics, metric definitions and result pages belong to an analytics team with its own warehouse and review process. Flag services that grow a stats engine end up with two teams' on-call and neither team's expertise.

Why this matters in interviews: Scoping experimentation down — clearly — is a stronger signal than designing a t-test.

10.5 "It's Only Config" Is How Outages Skip the Pipeline#

The Staff position: Any artifact that changes production behavior at runtime — flags, WAF rules, routing tables, threat-signature files, ML models — is a deploy. CrowdStrike's 2024 content update and Cloudflare's generated feature file in 2025 were both "only config". The company needs one rule for all of them, not one for code and none for the rest.

Why this matters in interviews: Generalizing from flags to every runtime-config path is the bridge from Staff to Principal.


11. Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

The Staff engineer designs a safe flag service. The Principal engineer notices that the company has five runtime-change paths — the flag vendor product teams use, an in-house config system for infrastructure, YAML in a bucket for the data platform, edge rules managed in the CDN console, and a model-push pipeline for ML — and that only one of them has staged rollout. Public post-mortems keep showing the same root cause from different paths: a runtime change that reached everything at once. The L7 problem is not one flag service; it is runtime change as a company-wide risk class: one safety standard for every path that can change production without a deploy, one SDK contract for evaluation and defaults, and a debt budget that keeps the inventory from outgrowing the code.

🧭 Principal Move: "I'm not going to ask whether our flag service is reliable. I'm going to ask how many ways an engineer can change production behavior without a deploy, which of those are staged, and what share of last year's incidents started on the unstaged ones. That number tells me where the next big outage comes from."

The Org-Level Fault Line#

One runtime-config platform vs per-team config systems.

OptionWhat WorksWhat BreaksWho Pays
Each team owns its config pathFits each use case; no central bottleneckN safety models; most paths unstaged; no single audit of "what changed before the incident"Incident responders; the next post-mortem
One platform owns every config pathOne pipeline, one audit, one standardPlatform becomes a bottleneck for edge, ML and infra teams with special needsSpecialized teams (velocity)
Platform owns the standard and the pipeline primitives; teams own their payloadsStaging, approval, audit, LKG and kill are shared services; payload formats and validators stay with ownersContract design and adoption effort; needs a credible platform teamPlatform (stewardship); teams adopt the pipeline API

🧭 Principal Move: "The platform owns everything that must be right once — change classification, staging, approvals, audit, last-known-good and break-glass. Teams own what must be specific — payload schemas, validators, health predicates. Any new runtime-config path must use the pipeline; the four existing ones migrate within six quarters, staged-rollout first, because that's where the incidents are."

Cost Model#

Assumptions: fully loaded engineer ~$250K/year, cloud list prices, vendor pricing as rough ranges that vary by contract (estimates, not quotes).

ScaleEngineers / ProcessesInfra ($/month)HeadcountOn-call LoadBuy Alternative (approx.)
Startup50 / ~500~$0.5–2K self-hosted open source0.25–0.5 engVendor or none; < 1 page/quarterVendor ~$1–5K/month — buy
Growth500 / ~15K~$5–15K (relays, client eval, store)3–4 eng: SDKs, pipeline, relaysShared platform rotation; 1–3 pages/monthVendor ~$20–80K/month — buy, with in-network relay
Enterprise2,500 / ~80K~$15–40K5–8 eng plus experimentation integrationDedicated rotation; kill-latency SLOVendor fees comparable to the team; build or hybrid

The pricing insight: infrastructure is never the cost driver; headcount and incidents are. At enterprise scale, the platform's budget is justified by config-caused incident minutes avoided and by engineer hours not lost to flag debt — so those are the two numbers to measure from the first quarter.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Bucketing algorithm, salt format, bucket countOne-wayChanging it reshuffles every user in every running rollout and experiment
Flag key namespace and the no-reuse ruleOne-wayOld binaries keep old meanings forever; reuse is unsafe once allowed
SDK evaluation semantics (rule order, fallthrough, defaults)One-way-ishEvery call site in every language depends on them; versioning takes quarters
Context schema (key kinds, attribute names)One-way-ishRules across thousands of flags reference attribute names
Rules shipped to clients vs values onlyOne-way-ishOnce rules are on devices, old app versions keep them for years
Push vs poll mechanismTwo-waySDK and relay change; behind the same API
Flag store technologyTwo-way30K rows; dual-write and cut over in a week
Vendor vs in-house backendTwo-way if behind a standard SDK interface; one-way otherwise40,000 call sites on a vendor SDK is a multi-quarter migration

🧭 Principal Insight: The bucketing spec and the SDK contract are the decisions I'd slow down on. Backends, relays and stores can change in a quarter; the meaning of "10% of users" and "the default" can't change at all once 1,800 services depend on them.

The Standard I'd Write#

RFC-CFG-001: Runtime Change Safety Standard
Status: Approved   Owners: Runtime Config Platform + SRE

Scope
  Every system that can change production behavior without a code deploy:
  feature flags, dynamic config, edge/WAF rules, routing tables, ML model and
  rule-file pushes.

MUST
  1. Classify each change as risk-reducing or risk-adding; only changes toward a
     declared safe state may use the fast path.
  2. Stage risk-adding changes: canary cell, then wider waves, with health
     predicates and automatic revert.
  3. Validate payloads at submit AND at the consumer (schema, size, delta guards);
     consumers keep the previous good version when a new one fails.
  4. Evaluate on the request path without network calls; serve last-known-good,
     then code defaults; never block process startup longer than 2 seconds.
  5. Record every change with actor, approver, diff and change request ID.
  6. Tier-0 flags: owner plus second approver for risk-adding changes; declared
     safe state; break-glass path independent of the control plane.
  7. Release and experiment flags carry an owner and an expiry; keys are never reused.

SHOULD
  1. Use the OpenFeature-compatible SDK interface.
  2. Update the code default when a flag reaches 100% for 30 days, or remove it.
  3. Run a synthetic kill per region every 10 minutes.

Exceptions
  Filed with the platform; reviewed within 5 business days; time-boxed to one
  quarter; SRE director sign-off for any exception to MUST 1, 2 or 4.

Success metrics
  - Config-caused high-impact incidents: tracked quarterly, target -50% in year 1
  - flags.kill.propagation_p99 ≤ 10 s in every region
  - flags.sdk.serving_defaults: 0 outside declared outages
  - Release flags past expiry: ≤ 5% of inventory
  - Unstaged runtime-change paths: 0 by end of year 2

What I'd Tell the VP#

"We have five ways to change production without a deploy, and only one of them rolls changes out gradually. That's where several of our worst outages started, and it's the same pattern behind well-known public outages at large cloud providers. I'm proposing one change pipeline that every runtime-config system uses: changes that turn things off stay instant, changes that turn things on roll out in stages and undo themselves if metrics drop. It's about six engineers for a year, mostly reusing what the flag platform already has. The measure of success is fewer config-caused incidents and shorter ones; I'll report that number every quarter. The main cost is that risky changes take about 45 minutes instead of 3 seconds, and I'll make sure emergency kills are never slowed."

Principal Interview Signals#

SignalWhat It Sounds Like
Counts change paths, not flags"How many ways can someone change production without a deploy, and which are staged?"
Prices the platform by incidents avoided"Infra is $30K a month; the case is config-caused incident minutes, which I'll measure."
Sets the org's failure posture"Every team's kill switch lives here, so it runs in its own failure domain and we game-day it quarterly."
Identifies the real one-way doors"The bucketing spec and SDK semantics are forever; the backend isn't."
Knows when not to standardize"Experimentation shares our bucketing and SDK, not our stats; I won't centralize that."

Staff answers that L7 interviewers find insufficient:

  • "We'll build a reliable flag service with staged rollouts" — correct for flags; silent on the four other runtime-change paths.
  • "Teams will clean up their flags" — names a behavior, not a budget, metric or owner.
  • "We'll buy a vendor" — no SDK contract, no in-network relay, no break-glass, no exit plan.

Appendices

Appendix A: Mechanics in Depth#

A.1 Evaluation Order#

Diagram: A.1 Evaluation Order

A.2 Bucketing#

def bucket(flag, context):
    key = context.key_for(flag.bucket_by)            # user, tenant or device key
    if key is None: return None                      # rule falls through; counted
    h = sha1(f"{flag.salt}.{key}".encode("utf-8")).digest()
    return int.from_bytes(h[:4], "big") % 100_000    # 0.001% resolution

def rollout(flag, weights, context):                 # weights in buckets, sum 100_000
    b = bucket(flag, context)
    if b is None: return flag.fallthrough_default
    acc = 0
    for variation, w in weights:                     # cumulative ranges: monotonic
        acc += w
        if b < acc: return variation
    return weights[-1][0]

Monotonicity comes from cumulative ranges in a fixed variation order: widening treatment from 10,000 to 20,000 buckets only adds users. Re-ordering variations or changing the salt breaks stickiness, so the admin API forbids both on a live flag.

A.3 SDK Update and LKG Promotion#

on ruleset_event(version):
    if version <= current.version: return            # monotonic, ignore old/dup
    delta = relay.fetch_delta(since=current.version) # falls back to full snapshot
    candidate = apply(current, delta)
    if not verify_checksum(candidate) or candidate.bytes > SCOPE_BUDGET:
        metric("flags.sdk.rejected_ruleset"); return  # keep current
    swap_atomic(current, candidate)
    write_file(lkg_dir / "pending", candidate)

every 60s:
    if pending.age > 5 min and process_health_ok_since(pending.applied_at):
        rename(lkg_dir / "pending", lkg_dir / "good")  # two-slot LKG
    poll_backstop()

Appendix B: Data Model and Flag Lifecycle#

flags(key PK, kind, value_type, schema_json, owner_team, tier, safe_state,
      bucket_by, salt, created_at, expires_at, status)        -- status: active | archived | tombstoned
env_configs(flag_key, env, on, targets_json, rules_json, fallthrough_json,
            off_variation, version, updated_at, PK(flag_key, env))
segments(key PK, env, rules_json, included_keys_count, version)  -- included_keys ≤ 10,000
change_requests(id PK, flag_key, env, diff_json, risk_class, author, approvers,
                rollout_plan_json, state, created_at)
audit_events(id PK, flag_key, env, actor, change_request_id, before_json,
             after_json, at)                                    -- append-only, WORM for tier 0
ruleset_versions(env, scope, version, checksum, snapshot_uri, created_at)
Diagram: Appendix B: Data Model and Flag Lifecycle

Appendix C: Coordination Mechanisms#

MechanismUsed ForWhy Not Something Else
Single-writer PostgreSQL with optimistic version checkFlag edits and change requestsHuman-paced writes (~0.05/s); conflicts are rare and must be shown to humans, not merged
Monotonic ruleset version per (env, scope)Ordering deltas; ignoring stale eventsLets SDKs and relays be idempotent; no consensus needed downstream
Object storage snapshotsCold start, relay restoreCheap, CDN-cacheable, survives control-plane loss
SSE streams via relaysChange notificationFan-out to 80K processes; consensus-store watches aren't built for that fan-out
Poll backstopSilent stream failureBounds staleness independent of push health
Signed break-glass file via host agentTier-0 kills when the control plane is downIndependent channel; verified by signature, not by reachability

Consensus stores like etcd are a good fit for small, strongly ordered infrastructure config with few readers — see Consensus Service and ZooKeeper & etcd. Meta's Configerator used ZooKeeper-derived Zeus at the root of its tree and fanned out through observers and per-host proxies for the same reason: consensus for ordering, a tree for fan-out.

Appendix D: API Contract & Client Behavior#

  • Server SDK: init(timeout=2s) returns immediately-usable client; bool/string/number/json(key, ctx, default); never throws on evaluation; reports evaluation counts per key every 60 s (powers usage-gated deletion).
  • Reconnect: exponential backoff with full jitter, base 1 s, cap 30 s; delta cursor accepted for 24 h.
  • Client evaluation: POST /v1/client/evaluate with context; response ttl (default 900 s) and version; clients refresh on launch, foreground and TTL; Retry-After honored; on failure the device keeps its cached map.
  • Payload rules: only flags marked client_visible; values only; no rules, segment names or other users' targeting.
  • Exposure: client calls track_exposure(key) when the variant is actually rendered, not on evaluation.

Appendix E: Observability#

MetricAlertWhy
flags.kill.propagation_p99 (synthetic)> 10 s for 3 runsThe kill switch's real latency
flags.sdk.ruleset_age_seconds p99 by cluster> 5 minDistribution broken somewhere
flags.sdk.serving_defaults> 0 for 5 minSilent failure state
flags.sdk.init_duration_ms p99> 2,000Startup contract violated
flags.default_divergence by teamWeekly reportDefaults rotting
flags.evaluations_of_unknown_key> 0 for a tombstoned keyReuse or stale binaries
flags.eval.duration_us p99 by flag> 20 µsExpensive rule or segment
flags.stale_count by teamOver budgetFlag debt
flags.change.auto_revertsAny on tier 0A bad change was caught; review it

Control plane vs data plane: page the platform on data-plane signals (propagation, age, defaults) at all hours; control-plane availability pages during business hours unless a kill is pending.

Appendix F: Scale Evolution#

StageWhat WorksWhat You Add
< 50 engineersVendor or env vars + deployOwner and expiry from day one
50–500Vendor with local eval SDKs; pollIn-network relay; LKG; staged rollouts for risky changes
500–2,500Relays + push; asymmetric change speed; debt dashboardConformance suite; tier-0 approvals; synthetic kill
2,500+Runtime-change platform shared by all config pathsOrg standard; config incidents as KPI; per-team debt budgets

What you don't build on day one: a rules DSL beyond equality, membership, semver and percentage; a stats engine; multi-writer control planes; edge evaluation; sidecars for languages you don't use.

  1. Loading the index…