Technologies referenced in this case study: ZooKeeper & etcd · API Gateways & Service Mesh · Redis
Related case studies: Load Balancer · Circuit Breakers · API Gateway · Distributed Consensus · Distributed Coordination · Consistent Hashing
How to Use This Case Study#
Organized for interview use first, reference second.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 4, 5 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → §3 Fault Lines → §4 Failure Modes → Deep Dives 1 and 4 |
| Deep Dive | 3+ hrs | Everything, including §11 Principal Lens and appendices |
What is Service Discovery? — Why interviewers pick this topic
Service discovery answers one question millions of times per second: "Which network addresses can serve requests for service X right now?" In a world of autoscaling, containers and rolling deploys, instance addresses change constantly — a 2,000-pod service might see hundreds of endpoint changes during a single deploy. Discovery is the system that tracks those changes (the registry), decides which instances are healthy, and distributes that view to every caller (the data plane: DNS resolvers, client libraries, sidecars, load balancers).
Before vs After — the deploy that routed to ghosts:
Without a well-designed discovery system:
t=0: Rolling deploy of payments-svc begins; old pods terminate as new ones start
t=+5s: Old pods are killed immediately; registry still lists them (TTL 90s)
t=+10s: Callers cached the endpoint list 30s ago; ~40% of their endpoints are dead
t=+15s: Connection-refused and timeouts; callers retry into other dead endpoints
t=+90s: Registry finally expires old entries; callers refresh over the next 30s
t=+2min: Deploy "succeeded" — and payments had a 2-minute 35% error rate
With discovery designed for churn:
t=0: Same deploy; pod receives SIGTERM → fails readiness → deregistered in <2s
t=+2s: Control plane pushes updated endpoints to all proxies (p99 < 3s)
t=+5s: Pod drains in-flight requests for 20s, then exits
t=+5s: Callers' proxies already stopped sending new requests to it
Result: Zero-error deploy; the same churn, handled as a routine event
Why interviewers reach for this question: It looks like a lookup table and is actually a distributed-systems problem with a nasty twist — the system that tells everyone where everything is must itself be more available than everything it describes. It tests CAP reasoning under realistic conditions, control-plane vs data-plane separation, failure detection, and — at Staff level — who owns the platform every service depends on.
Mechanics Refresher: Discovery Patterns
| Pattern | How It Works | Pros | Cons |
|---|---|---|---|
| DNS-based (A/SRV records, low TTL) | Callers resolve a name; registry updates DNS | Universal; every language supports it | TTL caching (often ignored); no health detail; no per-request LB; slow convergence (30s–minutes) |
| Client-side registry (Eureka, Consul API, ZooKeeper) | Client library fetches/watches the instance list and load-balances itself | Fast, rich metadata; no proxy hop | Library per language; version drift; every client talks to the registry |
| Server-side / proxy (L4/L7 LB, Kubernetes Service VIP) | Callers hit a stable address; a proxy knows the backends | Clients stay dumb; language-agnostic | Extra hop; the LB is a scaling and failure point |
| Sidecar / service mesh (Envoy + xDS) | Local proxy receives pushed endpoints from a control plane | Language-agnostic, fast push, rich LB and health | Operational complexity; control-plane fan-out at scale |
| Proxyless xDS (gRPC with xDS) | Client library speaks the mesh control-plane protocol directly | No sidecar hop; mesh-grade config | Limited to supported languages/features |
For most production systems: A platform-registered registry (e.g., Kubernetes endpoints fed by readiness), a push-based control plane feeding sidecars or proxyless clients, DNS as the universal fallback for things that can't speak xDS — and a data plane that fails static: if the control plane dies, keep routing to the last-known-good endpoints.
Executive Summary
If you only read one section, read this.
What This Interview Actually Tests#
Service discovery is not a key-value lookup question. Everyone can draw "service registers in ZooKeeper, client reads ZooKeeper."
This is a control-plane availability and freshness question that tests:
- Whether you know the registry must be more available than every service it describes — and what that implies for CP vs AP
- Whether you separate the control plane (who knows the truth) from the data plane (who routes requests) so one can fail without the other
- Whether you reason about who decides an instance is healthy — and what happens when that judgment is wrong for the whole fleet at once
- Whether you design for churn — deploys, autoscaling, preemption — as the normal case, not the exception
The key insight: Discovery data is eventually consistent no matter what you pick, because the world changes faster than any registry can observe it. The design question is how stale a caller's view may get, what it does when the view is wrong, and whether a registry outage can ever take down traffic that was working fine.
The L5 vs L6 vs L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | Picks a registry (Consul/ZooKeeper/etcd) and draws register → lookup | Asks about churn rate, fleet size, languages, and what happens when the registry is down | Asks how many discovery systems already exist in the org and which one becomes the paved road |
| Consistency | "Use ZooKeeper for strong consistency" | Registry can be CP for writes, but the data plane must be AP: cache and fail static | Separates what must be consistent (identity, config) from what must be available (endpoints) across the platform |
| Health | Heartbeats with a TTL | Combines platform readiness (orchestrator) with data-plane outlier detection; guards against mass deregistration | Owns the fleet-wide policy: panic thresholds, max-eject %, and who can change health semantics |
| Propagation | Clients poll every 30s | Push via watch/xDS with incremental updates; bounds convergence (p99 < 5s) and fan-out | Plans control-plane scaling for 10× services and cross-cluster federation; sets convergence SLOs |
| Failure | "Run 3 registry nodes" | Registry outage must not stop existing traffic; static fallback; staggered reconnect | Treats discovery as the org's largest shared failure domain; cells, blast-radius limits, independent fallback path |
| Ownership | Each service registers itself | Platform registers instances from orchestrator truth; teams own readiness semantics | Decides one platform vs many, funds it, deprecates legacy registries |
Why "consistency" separates levels
L5: "Discovery needs to be accurate, so I'll use a strongly consistent store like ZooKeeper or etcd." Sensible — and incomplete. A CP registry refuses writes (and sometimes reads) during a partition or quorum loss. If callers must read the registry per request, a quorum loss becomes a fleet outage.
L6: "The registry's write path can be CP — I want one authoritative record of which instances exist. But callers never depend on the registry being reachable to route. Every data-plane component holds a local copy and keeps using it if the control plane disappears. Staleness is bounded by how fast instances die, and I cover that with client-side health checks and outlier ejection." This is the insight behind Netflix choosing an AP design (Eureka) and behind Envoy's "last-known-good" behavior.
L7: Recognizes this as a platform principle, not a design choice: "Control planes may fail; data planes must not fail with them." Writes it into the standard for every control plane — discovery, config, certificates, feature flags.
Why "health" separates levels
L5: Instances heartbeat; missing 3 heartbeats means deregistered. Works for crashes — and the heartbeat measures "the process can reach the registry," not "the instance can serve traffic."
L6: Uses multiple signals with clear owners: the orchestrator's readiness (the service team defines what ready means), the data plane's observed errors (outlier detection per caller), and a registry-side sanity guard. Crucially, adds a panic mode: if more than ~50% of a service's endpoints are marked unhealthy at once, assume the health system is wrong and route to all of them.
L7: Recognizes that a health-check change deployed fleet-wide can deregister every instance of every service in minutes — and puts health-semantics changes through the same staged rollout as code.
Why "ownership" separates levels
L5: Each service embeds a registration library and registers itself on startup.
L6: "Self-registration means every service in every language implements registration, heartbeating and graceful deregistration correctly — they won't. I'd have the platform register instances from the orchestrator, which already knows what's running, and let service teams own only their readiness endpoint."
L7: Drives the organization to a single discovery platform, with a migration plan for the three legacy registries and a clear contract: platform owns availability and convergence SLOs, service teams own readiness semantics and graceful shutdown.
The Staff Positions#
| Position | Rationale |
|---|---|
| Data plane fails static | A registry outage must never stop traffic that was flowing; keep last-known-good endpoints |
| Registry can be CP; reads must be cached | One authoritative write path, but no request depends on registry reachability |
| Platform registers, services don't | The orchestrator knows what's running; self-registration drifts across languages |
| Push over poll for internal RPC | Poll-based convergence is 30–90s; deploys and autoscaling need < 5s |
| Health is layered | Orchestrator readiness + caller-side outlier ejection + panic threshold |
| Deregister before shutdown | Fail readiness, wait for propagation, drain, then exit — zero-error deploys |
| DNS is the fallback, not the primary | Universal, but TTL caching makes it slow and blind to health |
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Intra-cluster RPC routing | High churn (hundreds of changes/min), low latency, thousands of callers | Platform registration + push (xDS/watch) + sidecar or proxyless LB, fail static | Brief routing to dead endpoints, covered by retries and outlier ejection | Convergence p99 < 5s; no request depends on registry reachability |
| Cross-cluster / multi-region discovery | WAN latency, partial connectivity, locality preference | Federated registries, locality-aware routing, per-cluster authority | Failover floods a remote cluster; split views | Local first; remote only with capacity; bounded failover % |
| Coordination discovery (leaders, config, singletons) | Correctness over availability; low volume | Consensus store (etcd/ZooKeeper), leases, fencing tokens | Two leaders if fencing is missing | Linearizable; exactly one owner per role |
🎯 Staff Move: "I'll design for intra-cluster RPC routing — that's the high-churn, high-fan-out case where the hardest tradeoffs live. I'll call out that leader and config discovery are a different problem that needs a consensus store with fencing, and that I would not put both on the same system with the same availability posture."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | CP vs AP Registry | Refuse to serve during partitions (accurate) or serve possibly-stale data (available)? |
| 2 | Client-Side vs Server-Side Discovery | Smart clients (fast, rich, polyglot drift) or smart proxies (uniform, extra hop)? |
| 3 | Push vs Pull Propagation | Watches/xDS (fast, fan-out load) or DNS/polling (simple, slow)? |
| 4 | Who Decides Health | Self-heartbeat, orchestrator readiness, or caller-observed errors? What stops a mass deregistration? |
| 5 | Fail Static vs Fail Accurate | When the control plane is down or suspect, route on stale data or stop routing? |
In the Wild: Real Production Systems#
Netflix — Eureka and the Case for AP#
Netflix built Eureka for AWS, where instances come and go and network partitions between availability zones are real. Instances renew a lease every 30 seconds; entries expire after 90 seconds without renewal; clients cache the full registry and refresh every 30 seconds. Eureka servers replicate peer-to-peer without consensus. Its signature feature is self-preservation: if renewals across the fleet drop below ~85% of expected, the server assumes it is partitioned, not that the fleet died, and stops expiring instances.
Staff insight: Eureka is the explicit argument that discovery should prefer availability. Self-preservation is a panic threshold at the registry level — "if too much looks dead at once, the observer is probably wrong."
Airbnb — SmartStack (Nerve + Synapse)#
Airbnb's SmartStack (open-sourced ~2013) put a small health-checking agent (Nerve) next to each service instance, registering it in ZooKeeper only while healthy, and a local agent (Synapse) on each caller host that watched ZooKeeper and rewrote a local HAProxy config. Applications just called localhost:port. If ZooKeeper became unavailable, HAProxy kept routing with its current config.
Staff insight: SmartStack is a sidecar service mesh before the term existed: language-agnostic, local proxy, fail-static by construction. The pattern — control plane watches a registry, data plane is a local proxy — is what Envoy and xDS later standardized.
Kubernetes + Envoy/xDS — Platform-Registered, Push-Based#
In Kubernetes, services don't register themselves: the EndpointSlice controller derives endpoints from pods whose readiness probes pass (EndpointSlices hold up to 100 endpoints each by default, to keep update fan-out small). kube-proxy or a mesh control plane (e.g., Istio's istiod) watches those objects and pushes endpoints to proxies over xDS (EDS for endpoints). Envoy, created at Lyft, keeps serving with its last-received config if the control plane disappears, and applies a panic threshold (default 50%): if fewer than half of hosts are healthy, it load-balances across all of them.
Staff insight: This is the modern reference design: orchestrator is the source of truth, readiness is the service team's contract, push-based distribution with incremental updates, and a data plane that fails static.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Use ZooKeeper / etcd" | "Quorum is lost. Can services still call each other?" | Control plane vs data plane separation |
| "Services register themselves" | "The process is up but deadlocked. Who removes it?" | Health semantics and ownership |
| "Clients poll every 30s" | "A deploy replaces 500 pods in 2 minutes. What error rate do callers see?" | Convergence under churn |
| "Use DNS" | "Java caches DNS; clients ignore TTL. Now what?" | Real-world DNS behavior |
| "Push via watches" | "10,000 proxies, each watching 2,000 services. What happens on a deploy?" | Control-plane fan-out |
| "Deregister unhealthy instances" | "A bad health-check config marks 100% unhealthy. What happens?" | Panic thresholds, mass-deregistration guards |
System Architecture Overview#
Reading the diagram: Truth flows one way — orchestrator → registry → controller → discovery servers → proxies — and requests never flow through the control plane. Each sidecar holds a last-known-good endpoint list, so if etcd loses quorum or every discovery server crashes, traffic continues on the current view; only changes stop propagating. Pod 2 is draining: it failed readiness, was removed from endpoints, and is finishing in-flight requests before exiting. DNS serves clients that can't speak xDS, at worse freshness. The key metrics are convergence time and the age of each proxy's config.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Registry | "ZooKeeper for consistency" | "CP store for writes is fine — as long as no request depends on reaching it. Data plane caches and fails static." |
| Registration | "Services register on startup" | "Platform registers from orchestrator state; services own a readiness endpoint and graceful shutdown." |
| Propagation | "Poll every 30 seconds" | "Push incremental updates via watch/xDS; convergence p99 < 5s; debounce and batch to control fan-out." |
| Health | "Heartbeat with TTL" | "Readiness from the orchestrator, outlier ejection in callers, panic threshold at 50% so a bad health check can't empty a service." |
| Registry down | "Run it with 3 replicas" | "Replicas help, but the real answer is that existing traffic continues on cached endpoints; only changes pause." |
| Deploys | "Rolling deploy" | "Fail readiness → wait ≥ propagation p99 → drain → exit. preStop hook covers the gap." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Eureka lease renewal / expiry / client refresh | 30s / 90s / 30s | Worst-case ~2–3 min to stop routing to a dead instance via registry alone |
| Eureka self-preservation threshold | ~85% of expected renewals | Registry-level panic mode |
| Envoy panic threshold (default) | 50% healthy | Below it, route to all hosts rather than overload the few "healthy" |
| Envoy outlier ejection max (default) | 10% of hosts | Caps how much the data plane can remove on its own judgment |
| EndpointSlice default max | 100 endpoints per slice | Bounds update size; a 5,000-pod service is ~50 slices, not one 5,000-entry object |
| Kubernetes readiness probe defaults | period 10s, failureThreshold 3 | ~30s to mark unready unless tuned; tune to 2–5s for fast removal |
| Kubernetes termination grace (default) | 30s | Drain budget after deregistration |
| Typical internal DNS TTL | 5–60s (and many clients ignore it) | DNS convergence is minutes in practice |
| JVM DNS cache (no security manager) | 30s default; "forever" if a security manager is installed | Classic source of stale endpoints |
| etcd storage quota | 2GB default, ~8GB suggested max | Registry size limits — don't store blobs |
| Raft cluster size | 3 (tolerates 1 failure) or 5 (tolerates 2) | More nodes = slower writes, not more capacity |
| Push convergence target | p99 < 5s, ideally < 2s | Sets the drain wait during deploys |
| ZooKeeper session timeout | 2× to 20× tickTime (commonly 2–20s) | How long an ephemeral node lingers after a crash |
Interview Walkthrough
The prompt usually arrives as "Design service discovery for a microservices platform", "How do services find each other?", or "Design something like Consul." The trap is to spend the interview on the registry's data model. The registry is the easy part. The hard parts are propagation under churn, health semantics, and what happens when the control plane fails.
Phase 1: Requirements & Framing (2–3 min)#
"A few questions before I draw anything. How big is the fleet — hundreds of services or thousands, and how many instances? How much churn — are we autoscaling and deploying dozens of times a day? How many languages do callers use? And is this only for RPC routing, or also for things like leader election and config? I'll assume 1,500 services, 60,000 instances, 3 clusters across 2 regions, heavy autoscaling so ~500 endpoint changes per minute at peak, 4 languages, and RPC routing as the primary use. Leader election I'll treat as a separate problem."
What you've established:
- Scale and churn set the propagation design
- Languages decide client-side vs sidecar
- Scope separates endpoint discovery (AP-friendly) from coordination (CP-required)
Phase 2: Core Entities & API (1–2 min)#
| Entity | Fields That Matter | Owner |
|---|---|---|
| Service | name, namespace, ports, protocol, owner team, criticality | Service team (declared), platform (validated) |
| Endpoint | ip:port, service, zone, region, cluster, ready, weight, version, labels | Platform (derived from orchestrator) |
| Health state | ready (orchestrator), ejected (per caller, local), draining | Orchestrator + data plane |
| Subscription | caller identity, services of interest, last version acked | Discovery server |
API — the control plane's contract, not a per-request lookup:
Registration (platform-internal; not called by services):
UpsertEndpoint(service, endpoint, ready, zone, metadata)
RemoveEndpoint(service, endpoint)
Subscription (proxies / proxyless clients):
Watch(services[], known_versions{}) → stream of Delta{service, added[], removed[], changed[], version}
Ack(version) / Nack(version, error) # client rejected a bad config
Fallback:
DNS: payments.svc.cluster.local → A/AAAA or SRV records, TTL 30s
🎯 Staff Move: "Notice there's no 'Lookup(service)' on the request path. Callers subscribe once and receive deltas. Every request is routed from local memory. That's the property that makes a registry outage survivable."
Phase 3: High-Level Architecture (≤5 min)#
Draw the overview diagram and name five components with owners:
- Source of truth — the orchestrator (scheduler + node agent readiness). Owner: platform/compute.
- Registry store — consensus-backed (etcd/Raft), 3–5 nodes per cluster. Owner: platform.
- Discovery servers — stateless, horizontally scaled, watch the store and push deltas to subscribers. Owner: platform/networking.
- Data plane — sidecars or proxyless clients: local endpoint cache, load balancing, outlier ejection, fail static. Owner: platform (mechanics), service teams (per-service overrides).
- DNS fallback — for legacy and third-party clients. Owner: platform/networking.
"The registry is the least interesting box. The interesting questions are how fast changes reach callers, what happens when health signals are wrong, and what the data plane does when the control plane is gone."
Phase 4: Transition to Depth#
"I'd like to go deep on three things, in order of how often they cause real incidents: deploy-time churn and convergence, health signals and mass deregistration, and control-plane failure — including the fan-out problem when thousands of proxies reconnect at once. Then ownership and multi-cluster."
Phase 5: Deep Dives (25–30 min)#
Deep dive 1 — Convergence under churn (7 min).
Target: an endpoint removed from service must stop receiving new requests within 5s (p99)
Pipeline budget:
readiness failure detected ≤ 2s (probe period 1s, failureThreshold 2 during shutdown via preStop)
controller updates slice ≤ 0.5s
discovery server debounce 100ms
push to 10,000 subscribers ≤ 2s (incremental delta, only subscribers of this service)
total ≈ 4.6s → drain wait = 5s + in-flight max (e.g. 10s) → grace period 20s
Graceful shutdown: SIGTERM → preStop: fail readiness, sleep 5s → stop accepting → drain 10s → exit
Deep dive 2 — Health and mass deregistration (7 min). Three layers: orchestrator readiness (service-defined), caller-side outlier ejection (e.g., 5 consecutive 5xx → eject 30s, max 10%), and panic threshold (< 50% healthy → route to all). Add a control-plane guard: if a single update would remove > 30% of a service's endpoints, delay and alert rather than apply instantly.
Deep dive 3 — Control-plane failure and fan-out (7 min). If etcd loses quorum, no updates propagate; data plane keeps last-known-good. Endpoints that die during the outage are handled by outlier ejection and retries. When the control plane recovers, 10,000 proxies reconnect — add jittered reconnect backoff, resume from known versions (send deltas, not full state), and rate-limit full pushes.
Deep dive 4 — Scoping fan-out (4 min). A proxy that subscribes to every service receives every change: 1,500 services × churn → most updates are irrelevant. Scope subscriptions to declared dependencies (Istio's Sidecar resource is a real example). Push cost drops from O(proxies × changes) to O(interested proxies × changes).
Deep dive 5 — Multi-cluster (3 min). Each cluster is authoritative for its own endpoints. Cross-cluster discovery is federated read-only; routing prefers local zone → local region → remote with capacity caps.
Phase 6: Wrap-Up (2–3 min)#
"To summarize: the orchestrator is the source of truth, a consensus store holds endpoints per cluster, discovery servers push scoped deltas to sidecars and proxyless clients, and the data plane fails static with outlier ejection and a panic threshold. Deploys drain after deregistration using a wait derived from convergence p99. What I'd build next: convergence SLOs per cluster, a mass-deregistration guard, and federation with locality-aware failover. The biggest remaining risk is the control plane as a shared failure domain — I'd cell it per cluster and test 'control plane dead for 1 hour' quarterly."
Common Timing Mistakes#
| Mistake | Time Lost | What to Do Instead |
|---|---|---|
| Designing the registry schema in detail | 8 min | A table of entities; move on |
| Explaining Raft or ZAB internals | 10 min | "Consensus store, 3–5 nodes"; link to Distributed Consensus if asked |
| Comparing Consul vs etcd vs ZooKeeper vs Eureka feature by feature | 5 min | One sentence each on CP vs AP posture |
| Ignoring deploys | — | Deploys are the #1 source of discovery-related errors |
| Never saying "fail static" | — | It is the single most important property |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Discovery is the ultimate shared dependency. Every service depends on it; no service team owns it; and its failure modes are correlated — when it breaks, it breaks for everyone at once. That makes it a perfect Staff question: the technical design is modest, but the consequences of each decision are fleet-wide. A Senior engineer designs a registry. A Staff engineer designs a registry whose worst day doesn't take down the company.
It also forces precision about CAP. Candidates who say "we need strong consistency" are usually right about the registry and wrong about the system: the data plane must be available even when the registry isn't.
1.2 The L5 vs L6 vs L7 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"If the registry disappeared for an hour right now, what would stop working?"
The correct answer is: "Nothing that's currently working. New deploys and scale-ups wouldn't become routable, and instances that die would be caught by outlier ejection instead of the registry — but no existing request path depends on the registry being reachable." If your design can't say that, redesign it until it can.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Intent 1: Intra-cluster RPC routing. High churn, high fan-out, latency-sensitive. Staleness of a few seconds is fine because retries and outlier ejection cover dead endpoints. What's not fine: a registry outage that blocks routing, or a health bug that empties a service. Design for availability of the data plane and bounded convergence.
Intent 2: Cross-cluster and multi-region discovery. The new problem is that clusters fail independently and WAN links partition. Each cluster should be authoritative for its own endpoints, with other clusters' endpoints federated read-only. The danger is failover: if cluster A's payments dies and every caller in A starts calling cluster B's payments, B receives 2× load. Locality-aware routing with capacity caps.
Intent 3: Coordination discovery. "Who is the leader of shard 7?" "Which instance owns the cron scheduler?" These need linearizable answers and fencing, not eventual consistency — two leaders means corrupted data. Use a consensus store directly (leases in etcd, ephemeral sequential nodes in ZooKeeper) and fencing tokens; see Distributed Coordination.
🎯 Staff Move: "Endpoint discovery tolerates staleness and must be available; leader discovery tolerates unavailability and must be correct. Putting both behind one API with one failure posture is how you get either a fleet outage or a split brain."
2.2 When NOT to Build (or Use) a Discovery System#
| Situation | Why It's Wrong | What to Use Instead |
|---|---|---|
| < 10 services, stable hosts | Static config or a load balancer VIP is simpler and more reliable | LB with DNS name; config file |
| Already on Kubernetes | Kubernetes Services + EndpointSlices are discovery | Use the platform's; add a mesh only if you need L7 features |
| Managed platform services (cloud DB, queue) | Provider gives you a stable endpoint | Provider DNS name |
| Batch jobs reading data | They need data locations, not RPC endpoints | Metadata catalog / object store paths |
| Leader election | Endpoint discovery's AP posture causes split brain | Consensus store + leases + fencing tokens |
| Public clients | You don't want the internet seeing your topology | API gateway + DNS/anycast (API Gateway) |
2.3 What the Interviewer Leaves Underspecified#
| Unstated Assumption | Why It Matters | What to Say |
|---|---|---|
| Churn rate | Determines push vs poll | "I'll assume hundreds of changes per minute — push is required" |
| Language diversity | Client libraries vs sidecar | "Four languages → sidecar or proxyless xDS, not per-language libraries" |
| Orchestrator presence | Who's the source of truth | "Kubernetes-like orchestrator; I'll register from its state" |
| Number of clusters / regions | Federation and locality | "Three clusters; each authoritative for its own endpoints" |
| Health semantics | Liveness vs readiness | "Readiness is the service's contract; liveness only restarts" |
| Legacy clients | DNS fallback needed | "Legacy and third-party clients use DNS at worse freshness" |
2.4 Precise Terminology#
| Term | Precise Meaning | Common Confusion |
|---|---|---|
| Registry | Authoritative store of service → endpoints | Confused with the data plane that uses it |
| Control plane | Components that compute and distribute routing state | Treated as in the request path |
| Data plane | Components that route requests (proxies, clients) | Assumed to need the control plane per request |
| Self-registration | Instance registers itself and heartbeats | Assumed to be the only pattern |
| Third-party registration | A platform component registers instances it observes | Seen as "extra infrastructure" rather than simplification |
| Readiness | "Send me traffic" — service-defined | Confused with liveness ("restart me if false") |
| Outlier detection | Caller-local ejection of misbehaving hosts | Confused with registry health |
| Panic threshold | Below X% healthy, ignore health and use all hosts | Unknown to most candidates |
| Fail static | Keep using last-known-good state when the source is unavailable | Confused with "stale cache bug" |
| Convergence time | Time from a change until all subscribers apply it | Only measured on the happy path |
| xDS | Envoy's family of discovery APIs (LDS, RDS, CDS, EDS, SDS) | Thought to be Istio-specific |
3. The Five Fault Lines#
3.1 Fault Line 1: CP vs AP Registry#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| CP registry, clients query per request | Always-accurate answers | Quorum loss = no routing = fleet outage | Every service, simultaneously |
| CP registry, clients cache + fail static | Authoritative writes; routing survives registry outage | Stale during outage; changes pause | Services deploying/scaling during the outage |
| AP registry (Eureka-style, peer replication) | Registry itself stays writable during partitions | Divergent views; ghosts linger (90s+); no linearizable ops | Callers hitting stale endpoints; covered by retries |
| Orchestrator as registry (Kubernetes API/etcd) | No second source of truth; readiness built in | Kubernetes API server load at scale; per-cluster scope | Platform team scaling the API server |
The Staff default: CP store for the write path (etcd behind the orchestrator), strictly cached data plane that fails static. You get a single authoritative history and an available routing layer. The CAP choice is made separately for the two planes.
🎯 Staff Move: "I'm happy with a CP registry, as long as nothing in the request path reads it. The CAP decision I care about is the data plane's, and there I choose availability: route on the last good view."
When to deviate: Environments without an orchestrator and with frequent partitions between zones (classic early-cloud VMs) — an AP registry like Eureka with self-preservation is reasonable. Say so, and note the cost: longer ghost-endpoint windows.
3.2 Fault Line 2: Client-Side vs Server-Side Discovery#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Client library (Eureka client, Finagle, Consul API) | No extra hop; rich LB (least-request, zone-aware) | One library per language; version drift; upgrades take quarters | Every service team; platform can't enforce behavior |
| Central LB / VIP (L4/L7 load balancer, kube-proxy VIP) | Dumb clients; one place to configure | Extra hop (~0.2–1ms); LB capacity and failure domain; L4 can't balance gRPC streams per request | Platform (LB scaling); callers (latency, imbalance) |
| Sidecar proxy (Envoy) | Language-agnostic; fast push; full L7 features | Per-pod CPU/memory (~0.1–0.25 vCPU, 50–100MB); +0.5–2ms per hop; proxy operations | Platform (mesh ops); everyone (resource tax) |
| Proxyless xDS (gRPC xDS) | Mesh-grade config without sidecar overhead | Only supported languages; feature gaps | Service teams in unsupported languages |
The Staff default: Sidecar for the general fleet (polyglot), proxyless xDS for the highest-QPS services where the sidecar tax is measurable, and DNS/VIP for legacy clients. One control plane feeds all three.
When to deviate: A mostly single-language shop (e.g., all JVM) with < 100 services can run a good client library and skip the mesh. Name the trigger for revisiting: the second language, or the first incident caused by library version skew.
3.3 Fault Line 3: Push vs Pull Propagation#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| DNS with TTL | Universal; zero client code | Clients ignore TTL; 30s–minutes convergence; no health or weights; resolver caching layers | Callers during deploys (errors) |
| Polling the registry (every 30s) | Simple; load is predictable | Convergence = poll interval + registry lag; N clients × 1/30s steady load | Callers (stale), registry (steady load) |
| Watch / push, full state | Fast | Every change sends full lists to every subscriber: O(subscribers × endpoints) per change | Control plane and network during deploys |
| Watch / push, incremental and scoped | Fast and efficient — only deltas to interested subscribers | Complex versioning; resync logic on reconnect | Platform (complexity) |
Fan-out math, full-state push:
10,000 proxies subscribed to all services
payments has 2,000 endpoints (~100KB of endpoint data)
deploy changes payments ~40 times per minute (batched)
full push: 10,000 × 100KB × 40/min ≈ 40 GB/min from the control plane
Scoped + incremental:
800 proxies actually call payments; delta ~2KB per batch
800 × 2KB × 40/min ≈ 64 MB/min — ~600× less
The Staff default: Push, incremental, scoped to declared dependencies, with 100ms–1s debounce to batch changes during deploys. DNS stays as the fallback for clients that can't subscribe.
🎯 Staff Move: "Push is only better than poll if it's scoped and incremental. Full-state push to every proxy is how control planes melt during the biggest deploy of the day."
3.4 Fault Line 4: Who Decides Health#
| Signal | Measures | Blind Spot | Owner |
|---|---|---|---|
| Self-heartbeat | Process can reach the registry | Deadlocked request threads still heartbeat; partition from registry ≠ dead | Service (library) |
| Orchestrator readiness probe | Service's own definition of "ready" | Probe may check the wrong thing (e.g., dependency DB) → correlated failure | Service team defines; platform runs |
| Registry active checks | Registry can reach the instance | Registry's network view ≠ callers' view | Platform |
| Caller-side outlier detection | Actual request outcomes from this caller | Local only; can't see instances it doesn't call | Data plane (platform config) |
The critical failure: readiness probes that check dependencies. If every payments pod's readiness checks the payments DB and the DB blips, every pod goes unready simultaneously, the service has zero endpoints, and a 5-second DB blip becomes a full outage that lasts until probes pass again (plus propagation). Readiness should check this instance's ability to serve, not shared dependencies.
The Staff default: Orchestrator readiness as the registry's health truth (service-owned, checks only local state), caller-side outlier detection for fast per-caller reaction (capped at 10% ejection), panic threshold at 50%, and a control-plane guard against removing > 30% of a service's endpoints in one update.
3.5 Fault Line 5: Fail Static vs Fail Accurate#
When the data plane can't reach the control plane, or the control plane sends something suspicious, what happens?
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Fail closed (no routing without fresh data) | Never routes to a stale endpoint | Control-plane outage = fleet outage | Everyone |
| Fail static (last-known-good, indefinitely) | Traffic continues through any control-plane outage | Endpoints that died during the outage linger; new ones never appear | Services scaling or deploying during the outage |
| Fail static with expiry (e.g., drop endpoints after 24h without update) | Bounds how stale it gets | If outage > expiry, you fail closed after all | Everyone, on long outages |
| Reject bad config (NACK + keep previous) | A bad push can't empty routing tables | Needs validation rules that detect "bad" | Platform (validation) |
The Staff default: Fail static without expiry for endpoints (combined with outlier ejection to cover dead hosts), validate every push (NACK suspicious updates such as an empty endpoint set for a service that had 200), and alert on stale_config_age_seconds > 60. Persist last-known-good to local disk so proxy restarts during a control-plane outage don't come up empty.
🎯 Staff Move: "Stale data with outlier ejection routes most requests correctly. No data routes none. I'll take stale."
When to deviate: Security-sensitive routing (e.g., revoked tenants, decommissioned hosts whose IPs get reused) may need shorter bounds — handle that with mTLS identity checks at the destination rather than by failing closed.
4. Failure Modes & Operational Reality#
4.1 Deploy-Time Errors — Routing to Terminated Instances#
The most frequent discovery failure is not an outage. It's a 0.5–5% error spike on every deploy, which teams learn to ignore.
Setup: payments, 400 pods, rolling deploy 25% at a time, callers poll registry every 30s.
Pod killed on SIGTERM immediately; deregistration via heartbeat expiry (30s).
t=0: 100 pods receive SIGTERM and exit within 1s
t=+1s: Registry still lists them; callers' caches list them
t=+1–30s: ~25% of new connections to payments are refused; in-flight requests are cut
t=+30s: Heartbeats expire; registry removes 100 pods
t=+30–60s: Callers poll and converge
t=+60s: Next batch of 100 pods... repeat 4×
Result: ~4 minutes of 5–25% errors on callers without retries; retry load on the rest
Detection: upstream_cx_connect_fail, upstream_rq_reset by destination correlated with deploy events; deploy.error_rate_delta.
Mitigation: Pause deploy; enable caller retries on connect-failure (safe — the request never reached the server).
Prevention: Shutdown sequence:
Owner: Platform owns the preStop defaults and convergence SLO; service teams own graceful drain in their server code.
4.2 Mass Deregistration — The Health Check That Emptied a Service#
t=0: Team adds a readiness check that verifies connectivity to the shared config service
t=+2min: Deploy reaches 100%; config service has a 20-second GC pause
t=+2min: All 600 pods fail readiness within 10s; endpoints → 0
t=+2min: Control plane dutifully pushes 'payments has no endpoints' to 8,000 proxies
t=+2m10s: Every caller gets 'no healthy upstream'; checkout at 100% errors
t=+2m20s: Config service recovers; readiness passes after failureThreshold/successThreshold windows
t=+3min: Endpoints repopulate; convergence completes
A 20-second blip in a non-critical dependency caused a full payments outage.
Detection: endpoints_ready{service} dropping > 30% in < 60s; no_healthy_upstream_total.
Mitigation: Panic mode in proxies (route to all hosts when < 50% healthy) would have kept serving through the blip. Control-plane guard would have held the update.
Prevention: Readiness checks only local state; dependency health belongs in circuit breakers (Circuit Breakers), not readiness. Health-check changes roll out with canary like code.
Owner: Service team (readiness semantics); platform (panic threshold, mass-removal guard).
4.3 Registry Quorum Loss#
t=0: 2 of 3 etcd members lose disks in a maintenance event; quorum lost
t=+0s: Writes fail; the orchestrator can't update endpoints; watches stall
t=+0s: Data plane: every proxy continues with last-known-good — traffic unaffected
t=+5min: Autoscaler adds 200 pods for rising traffic — they're not routable
t=+10min: 30 pods die from a node failure — outlier ejection removes them per caller
t=+45min: etcd restored from snapshot; controllers reconcile; watches resume
t=+46min: 8,000 proxies reconnect at once → discovery servers CPU-bound
Detection: etcd_server_has_leader == 0, xds_push_age_seconds rising on every proxy, registry_write_errors_total.
Mitigation: Freeze deploys during the outage; scale by adding capacity to existing pods if possible; after recovery, stagger reconnects.
Prevention: 5-member etcd for large clusters, spread across failure domains; tested restore runbook (< 30 min); jittered proxy reconnect; discovery servers that serve from their own cache while the store is down.
Owner: Platform on-call; incident commander freezes deploys.
4.4 Reconnect Storm After Control-Plane Recovery#
Detection: xds_connected_clients sawtooth; discovery_server_cpu pegged; xds_push_errors.
Prevention: Reconnect with exponential backoff and full jitter; resume by version so most proxies need only a delta; discovery servers build one snapshot and serve it to many clients rather than querying the store per client; admission control on full-state requests.
Owner: Platform.
4.5 Stale DNS — The Endpoint That Wouldn't Go Away#
A legacy JVM service with a security manager caches DNS forever. A database proxy's IP changes during a migration; the service keeps connecting to the old IP for 11 days until its next restart — the old host was reused by another service, which now receives (and rejects) the traffic.
Detection: connections_to_decommissioned_ips (from flow logs); destination-side mTLS identity mismatches.
Prevention: Explicit networkaddress.cache.ttl in base images; mTLS so a reused IP can't impersonate a service; decommissioned IPs quarantined for 24h before reuse.
Owner: Platform (base images, IP reuse policy); service team (legacy upgrade).
4.6 Zombie Instance / IP Reuse#
A node partition leaves a pod running but disconnected from the orchestrator. The orchestrator reschedules the pod elsewhere and the old IP is reassigned to a new pod of a different service. Callers with stale views send payments traffic to an inventory pod.
Prevention: mTLS with workload identity (SPIFFE-style) so the caller verifies it reached payments, not whoever holds the IP; endpoint metadata includes a pod UID; servers reject requests addressed to a different service name.
Owner: Platform (identity), security.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Deploy routes to dead pods | upstream_cx_connect_fail during deploys | One service's callers, per deploy | preStop drain, push convergence, retry on connect fail | Platform + service team |
| Mass deregistration | endpoints_ready drop > 30% in 60s | One service, all callers | Panic threshold, mass-removal guard | Service (readiness), platform (guards) |
| Registry quorum loss | etcd_server_has_leader == 0 | No changes fleet-wide; traffic continues | Fail static; freeze deploys; restore | Platform |
| Reconnect storm | xds_connected_clients sawtooth | Control plane; delays recovery | Jittered reconnect, snapshot serving | Platform |
| Stale DNS | Flows to decommissioned IPs | Individual legacy clients | TTL config, IP quarantine | Platform + service owners |
| Zombie / IP reuse | mTLS identity mismatches | Misrouted requests | Workload identity, UID checks | Platform + security |
| Full-state push fan-out | Control-plane egress GB/min spikes on deploys | Control plane, then convergence | Scoped, incremental push; debounce | Platform |
| Bad config push | Proxies NACK or no_healthy_upstream spike | Potentially fleet | Validation, staged rollout, NACK-and-keep | Platform |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | Picks a registry | Clarifies churn, scale, languages; separates endpoint vs coordination discovery | Asks what exists today and what becomes the paved road |
| CAP reasoning | "Strong consistency" | CP writes, AP data plane; fail static | Codifies "control planes may fail, data planes must not" for all control planes |
| Propagation | Poll | Scoped incremental push; convergence budget | Convergence SLOs per cluster, capacity model for control-plane fan-out |
| Health | Heartbeats | Layered health, panic threshold, mass-removal guard | Health-semantics changes governed like code; fleet-wide policy |
| Deploys | Rolling deploy | Deregister → wait → drain → exit, with numbers | Deploy error budget as a platform SLO |
| Ownership | Services self-register | Platform registers; service owns readiness and drain | Single platform; legacy registry deprecation; funding |
| Security | Not mentioned | mTLS identity to defeat IP reuse | Identity-based routing and authorization as the long-term contract |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Separates planes | "Nothing in the request path reads the registry." |
| Designs for churn | "Deploys are the common case; I'll budget convergence at 5 seconds and drain after it." |
| Guards against health bugs | "Readiness checks only local state, and a panic threshold keeps a bad check from emptying the service." |
| Quantifies fan-out | "Full-state push is 40GB a minute during a big deploy; scoped deltas are ~64MB." |
| Owns the ownership split | "Platform registers and distributes; service teams own readiness and graceful shutdown." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| Registry lookup on every request | Makes the registry a per-request SPOF |
| "ZooKeeper is strongly consistent, so we're safe" | Confuses registry consistency with system correctness |
| Readiness probes that check the database | Converts dependency blips into full outages |
| No shutdown/drain story | Every deploy produces errors |
| DNS with a 5s TTL as the whole answer | Ignores client caching behavior and lack of health data |
5.4 Common False Positives#
- Knowing Raft in detail ≠ designing discovery. Consensus is one box; propagation and health are the problem.
- "We'll use Istio" ≠ a design. The mesh doesn't decide readiness semantics or panic thresholds for you, and it has its own fan-out limits.
- Comparing Consul/etcd/ZooKeeper feature matrices ≠ judgment. State the CP/AP posture and move on.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Scale, churn, languages, scope |
| Entities & API | 3–6 min | Subscribe-and-push contract; no per-request lookup |
| Architecture | 6–11 min | Orchestrator → store → discovery servers → data plane |
| Churn & convergence | 11–18 min | Budget, drain sequence |
| Health & mass deregistration | 18–25 min | Layers, panic, guard |
| Control-plane failure & fan-out | 25–33 min | Fail static, reconnect storm, scoping |
| Multi-cluster / ownership | 33–41 min | Federation, locality, owners |
| Wrap-up | 41–45 min | Next steps, biggest risk |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Direction |
|---|---|---|
| "Make it work without Kubernetes" | Registration without an orchestrator | Agent-based registration (SmartStack/Consul agent) with local health checks |
| "Now add leader election" | CP vs AP separation | Separate consensus-backed leases + fencing tokens |
| "Callers are in 4 languages" | Client vs sidecar | Sidecar or proxyless xDS; one control plane |
| "What about 50,000 pods per service?" | Fan-out and object sizes | Endpoint slicing, subsetting (each caller sees a subset), delta push |
| "Multi-region?" | Locality and failover | Per-cluster authority, federation, capacity-capped failover |
| "How do you secure it?" | IP reuse, identity | mTLS workload identity; registry write ACLs |
6.3 What to Deliberately Skip#
- Consensus internals (Raft log replication) unless asked
- Registry storage schema beyond an entity table
- Detailed DNS protocol (SRV record fields)
- Vendor comparisons — one line each
6.4 Follow-Up Questions to Expect#
- "How do you choose the drain wait?" — Convergence p99 plus margin; measured, not guessed.
- "What happens if the discovery server pushes an empty endpoint list?" — Validation NACKs it; guard holds large removals; panic threshold routes to all.
- "How does the caller pick among 2,000 endpoints?" — Subsetting (each caller gets ~50–100 deterministic endpoints) plus least-request/P2C; see Load Balancer.
- "How do you handle a service that's in both old and new registries during migration?" — Dual-register via the platform, callers read the new system first with fallback to the old, migrate by caller cohort.
- "Why not just use a load balancer?" — At small scale, do. At thousands of services, per-service LBs are costly and add a hop; client-side LB with pushed endpoints scales better.
- "What's the difference between liveness and readiness?" — Liveness restarts; readiness routes. Never make liveness depend on dependencies — that's a restart storm.
- "How do you know convergence is healthy?" — Synthetic changes (canary endpoint flipped every minute) and measure time to observe in sampled proxies.
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design service discovery for our microservices platform."
Staff Answer
"Three clarifying questions: scale and churn, language mix, and whether this includes coordination like leader election. I'll assume ~1,500 services, 60K instances, heavy autoscaling, four languages, and RPC routing as the scope — leader election is a separate, consensus-backed problem.
The core design: the orchestrator is the source of truth; a platform controller derives endpoints from ready instances into a consensus-backed store; stateless discovery servers push scoped, incremental updates to sidecars and proxyless clients; the data plane routes from local memory and fails static. No request ever reads the registry.
Then I'd go deep on three things: convergence during deploys, health signals and how we prevent a bad health check from emptying a service, and what happens when the control plane fails or recovers."
Why this is L6:
- Separates endpoint discovery from coordination up front
- States the plane separation and fail-static property as the core design
- Outlines depth in order of real-world incident frequency
What L7 adds:
- Asks what discovery systems exist today and frames the design as the paved road with a migration plan
- Names convergence and control-plane availability SLOs the platform will be held to
❌ Common L5 Trap
"Services register with Consul on startup and send heartbeats every 10 seconds. Clients query Consul to get the list of healthy instances and pick one at random."
Why this misses: Correct building blocks, but it's unclear whether clients query per request (making Consul a SPOF), there's no churn story, and health is only heartbeats.
Drill 2: Registry Down#
Prompt: "Your registry cluster has lost quorum. What happens to traffic?"
Staff Answer
"Existing traffic is unaffected by design: every proxy and client holds last-known-good endpoints in memory and on local disk, and routes from that. What stops is change: new pods from deploys or autoscaling aren't routable, and pods that die aren't removed from the registry.
The second part is covered by the data plane: outlier detection ejects endpoints that fail (e.g., 5 consecutive connect failures → eject 30s), and retries on connect-failure are safe. The first part is operational: freeze deploys, and if we need capacity, scale vertically or rely on existing headroom.
The recovery is the risky moment — thousands of proxies reconnecting. Jittered reconnect over ~30 seconds, resume-by-version so most get a small delta, and discovery servers that serve a shared snapshot rather than hitting the store per client."
Why this is L6:
- Distinguishes loss of state changes from loss of routing
- Covers the gap with data-plane mechanisms
- Anticipates the recovery storm
What L7 adds:
- Makes 'control plane dead for 1 hour' a quarterly game day for every control plane
- Tracks "time since last successful restore" for the registry as a platform health metric
Drill 3: Make It Concrete — Zero-Error Deploys#
Prompt: "Every deploy of payments causes a 2% error spike for 3 minutes. Fix it."
Staff Answer
"The errors are almost certainly callers sending to pods that already stopped. Fix the ordering: on SIGTERM, the pod fails readiness but keeps serving; the control plane removes it and pushes to callers; the pod waits long enough for that push to land everywhere — our convergence p99, say 4 seconds, plus margin, so a preStop sleep of ~8s; then it stops accepting new connections, drains in-flight for up to 15s, and exits well within the 30s grace period.
Then measure: I'd add upstream_cx_connect_fail and upstream_rq_reset by destination, annotated with deploy events, and set a deploy error budget — e.g., < 0.01% errors attributable to deploys. Callers get a retry on connect-failure, which is always safe because the request never reached the server."
Why this is L6:
- Diagnoses the ordering problem precisely
- Derives the wait from convergence p99
- Adds a measurable deploy error budget
What L7 adds:
- Bakes the preStop pattern into the platform's service template so no team configures it by hand
- Publishes convergence p99 as an SLO so the preStop wait can be derived automatically
Drill 4: The Health Check That Emptied a Service#
Prompt: "A readiness check change caused all instances of a service to go unhealthy at once. How do you prevent this class of problem?"
Staff Answer
"Three layers. First, semantics: readiness must check only this instance's ability to serve — warmed caches, loaded config, thread pool not saturated — never shared dependencies. Dependency health belongs in circuit breakers. That's a standard with a lint rule on probe definitions.
Second, the control plane: if an update would remove more than ~30% of a service's endpoints within a minute, hold it and page platform on-call. Legitimate large removals (scale-down) are rate-limited anyway.
Third, the data plane: panic threshold. If fewer than 50% of a service's endpoints are healthy, route to all of them. If the health signal is wrong, we keep serving; if it's right, we're no worse off than routing to the few survivors, who'd be overloaded anyway.
And the organizational piece: readiness changes ship via canary like code, because they are code."
Why this is L6:
- Fixes semantics, control plane and data plane independently
- Explains why panic mode is safe in both cases
- Treats health config as code
What L7 adds:
- Establishes the principle fleet-wide: any automated system that can remove capacity has a cap on how much it can remove at once
- Applies it beyond discovery — autoscalers, outlier detection, node drainers
Drill 5: Control-Plane Fan-Out#
Prompt: "Our mesh control plane falls over during large deploys. We have 12,000 sidecars and 2,000 services."
Staff Answer
"Likely cause: every sidecar subscribes to every service and gets full-state pushes. A deploy of a 1,000-pod service produces hundreds of changes; each one is pushed to 12,000 sidecars.
Four fixes in order of impact. Scope subscriptions: each sidecar only receives services its workload actually calls — typically 10–50, not 2,000; this alone cuts push volume by ~40–200×. Incremental pushes: send endpoint deltas, not full lists. Debounce: batch changes over 100ms–1s windows during deploys. Horizontal scale: discovery servers are stateless — add replicas and shard subscribers across them, with a shared snapshot so the store is read once per change, not once per subscriber.
I'd measure push latency p99 and control-plane egress per change, and set an SLO: convergence p99 < 5s during the largest deploy of the week."
Why this is L6:
- Identifies the O(subscribers × changes) problem
- Prioritizes fixes by impact
- Sets an SLO under peak churn, not the average
What L7 adds:
- Makes declared dependencies (the scope list) part of the service manifest, which also feeds security policy and the dependency graph for architecture review
- Plans for 3× growth: a capacity model for the control plane as a function of services × churn × subscribers
Drill 6: Hot Service With 20,000 Instances#
Prompt: "One service has 20,000 instances and 3,000 callers. Every caller connecting to every instance means 60M connections. What do you do?"
Staff Answer
"Subsetting. Each caller gets a deterministic subset of ~50–100 instances, chosen so that load is balanced across instances — e.g., shuffle instances with a seed, divide into subsets, and assign callers round-robin across subsets. Connections drop to 3,000 × 100 = 300K. Discovery pushes only the caller's subset, which also shrinks update fan-out.
Tradeoffs: subset changes cause connection churn, so the algorithm must be stable when instances are added — only a small fraction of assignments move. And with small subsets, a few slow instances affect a caller more; outlier ejection and least-request LB within the subset handle it. Google's SRE book describes deterministic subsetting for exactly this reason."
Why this is L6:
- Quantifies the connection problem
- Knows subsetting and its stability requirement
- Connects subset size to failure sensitivity
What L7 adds:
- Makes subsetting a platform capability with a default size, rather than something the hot service team builds
- Considers whether the hot service should be split, not just subsetted
Drill 7: Build vs Buy#
Prompt: "Should we build our own discovery system?"
Staff Answer
"Almost certainly not. If we're on Kubernetes, EndpointSlices plus a mesh control plane (Istio, Linkerd, or a managed mesh) cover intra-cluster discovery. If we're on VMs, Consul is a mature choice. What we'd build is the thin layer that's specific to us: the service manifest format, criticality and dependency declarations, multi-cluster federation policy, and dashboards and SLOs.
Building a registry and xDS server from scratch is 2–4 engineers for a year and a permanent on-call; a managed or open-source option costs integration time and some flexibility. The case for building is only when scale exceeds what the open-source control planes handle — tens of thousands of services — and even then companies usually extend rather than replace. For vendor-level comparison, see the API gateway & mesh technology guide."
Why this is L6:
- Identifies what's differentiating vs commodity
- Prices the build option in headcount and on-call
- Names the condition that changes the answer
What L7 adds:
- Evaluates lock-in: keep the service manifest and policy in our own format so the underlying mesh can be swapped
- Plans the exit path for whatever is chosen
Drill 8: Changing Discovery Without an Outage#
Prompt: "We need to migrate 800 services from Eureka to Kubernetes-native discovery with a mesh. How?"
Staff Answer
"Dual-run, migrate callers before callees, one cohort at a time. Step one: a bridge that syncs Eureka registrations into the new registry and vice versa, so both systems see all endpoints. Step two: migrate callers to the mesh in cohorts — lowest criticality first — so they discover both old and new instances through the bridge. Step three: migrate callees to platform registration once all their callers are on the mesh. Step four: remove the bridge per service when no Eureka clients remain for it.
Each cohort has a rollback: flip back to the Eureka client. Success metrics per cohort: error rate and p99 unchanged within noise for 1 week. The bridge is the riskiest piece — it must handle Eureka's 90s expiry vs the new system's second-level updates without creating flapping."
Why this is L6:
- Orders the migration by dependency direction
- Uses a bridge for coexistence and a per-cohort rollback
- Identifies the semantic mismatch in the bridge as the key risk
What L7 adds:
- Sets a hard deprecation date and tracks the long tail explicitly — the last 10% of services take as long as the first 90%
- Budgets the migration (engineer-quarters) against the ongoing cost of running two systems
Drill 9: Cost#
Prompt: "The mesh sidecars cost us a lot of CPU. Is service discovery worth it?"
Staff Answer
"Discovery and the mesh are separable costs. Discovery itself — a registry and control plane — is cheap: a few dozen cores for a large fleet. The sidecar tax is the bigger number: at ~0.1–0.25 vCPU and ~50–100MB per pod, 30,000 pods is 3,000–7,500 vCPUs — perhaps $90–225K/month at ~$30/vCPU-month.
Options: proxyless xDS for the highest-traffic services removes the sidecar where it costs the most, per-node proxies instead of per-pod sidecars (the ambient-mesh direction), or tune sidecar resources — most are overprovisioned. Compare against what the mesh replaces: per-language discovery libraries, per-service load balancers, and mTLS implementation in every service."
Why this is L6:
- Separates the cheap part (discovery) from the expensive part (data-plane proxies)
- Quantifies with stated assumptions
- Offers architectural options, not just 'tune it'
What L7 adds:
- Frames the decision as a 2–3 year data-plane strategy and names the one-way doors in each option
Drill 10: Multi-Region#
Prompt: "We're adding a second region. How does discovery work across regions?"
Staff Answer
"Each cluster stays authoritative for its own endpoints; no cross-region consensus in the write path, because a WAN partition would then stall both regions. Cross-region visibility is federated and read-only: each region's discovery servers import remote endpoints with locality labels.
Routing: prefer same zone, then same region, then remote — and only for services explicitly allowed to fail over cross-region, with a cap (e.g., at most 20% of a remote service's capacity), because unrestricted failover doubles load on the remote copy. Whole-region failover is an edge decision (DNS/anycast), not something each service improvises. And the remote view must fail static too: if the link drops, the local region keeps working with local endpoints only."
Why this is L6:
- Keeps per-cluster authority, avoiding cross-region consensus
- Caps cross-region failover by capacity
- Distinguishes service-level and region-level failover
What L7 adds:
- Defines cells as the unit of failure and discovery scope
- Prices N+1 regional headroom for leadership
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Autoscaling Doesn't Help#
Context: During a product launch, traffic to the search service triples. The autoscaler adds 300 pods within 4 minutes, but search latency keeps climbing and callers see timeouts for 12 minutes. The new pods show near-zero traffic. The on-call escalates to you.
Questions to Surface First:
- Are the new pods passing readiness? How long do they take to become ready?
- What's the control plane's push latency right now — is it backlogged by the scale-up itself?
- Are callers subset-based? Do new pods appear in any caller's subset?
- Are callers holding long-lived connections (HTTP/2, gRPC) that never rebalance to new pods?
Typical L5 Approach: Checks that pods are running and ready, restarts a few callers to see if traffic shifts, adds more pods. Restarts help anecdotally, which hides the real cause.
Staff Approach: Measures each pipeline stage: pod ready time (warm-up took 90s due to a large index load), control-plane push latency (backlogged to 45s by 300 simultaneous changes pushed as full-state updates), and caller connection behavior (gRPC channels with long-lived HTTP/2 connections never opened connections to new endpoints because the client's LB policy was pick-first). Fixes the LB policy to round-robin/least-request over all endpoints, enables incremental pushes with debounce, and moves index loading before readiness.
Principal Approach: Recognizes that "autoscaling works" was never tested end-to-end: time from scale decision to useful traffic on new capacity. Defines a platform SLO — capacity-to-traffic time p99 < 60s — and a quarterly test that scales a tier-1 service 2× under synthetic load and measures it. Makes the gRPC LB policy a platform default rather than a per-service choice.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Check endpoints_ready{service=search} vs pods running; check xds_push_latency_p99; sample a caller's endpoint table. |
| Triage | Three suspects: slow readiness, slow propagation, sticky connections. Measure each. |
| Quick fix | Rolling restart of the top 10 callers by QPS forces new connections (buys time). |
| Guardrails | Switch gRPC LB policy; set max connection age (e.g., 5 min) on servers so clients rebalance. |
| Post-mortem | Warm-up before readiness; incremental push; capacity-to-traffic SLO. |
Metrics to Watch: pod_ready_seconds, xds_push_latency_p99, upstream_rq_total{endpoint} distribution (new vs old pods), search.latency_p99.
Organizational Follow-up: Launch-readiness checklist includes a scale-up test; platform owns the default client LB policy.
Ownership Question: "Whose bug was the pick-first LB policy?" Staff answer: The platform's. A default that silently defeats autoscaling is a platform defect, even if each team technically chose it. The fix is a platform default and a lint rule, not 200 tickets to service teams.
Key Takeaway: "Discovery isn't done when the registry knows about new instances. It's done when traffic reaches them."
What clears the Staff bar:
- Measures each stage of scale-to-traffic separately
- Finds the connection-level cause (sticky HTTP/2 connections)
- Converts the fix into a platform default
Deep Dive 2: Silent Staleness#
Context: A security review discovers that for 6 weeks, ~3% of sidecars in one cluster have been running with endpoint data 2–9 days old. They were routing mostly correctly because the underlying services were stable. No alert fired. How did this happen and what do you change?
Questions to Surface First:
- Why did these sidecars stop receiving updates — a crashed watch, a NACKed config, a bug in resume logic?
- What does the data plane do when it NACKs a push — keep the previous config forever?
- Is there any metric for config age per proxy? Is anyone looking at it?
- Were any requests misrouted — e.g., to reused IPs?
Typical L5 Approach: Finds the bug (a specific config in one namespace was rejected by an older sidecar version, which NACKed every subsequent push containing it), upgrades the sidecars, closes the ticket.
Staff Approach: Fixes the bug, then the detection gap: every proxy reports
config_versionandlast_update_age_seconds; alert when > 1% of proxies are more than 5 minutes behind the latest version, or any single proxy is > 1 hour behind. Makes NACK a first-class signal with a per-resource error, so one bad resource doesn't block every update. Adds a synthetic "canary endpoint" flipped every minute to measure end-to-end convergence continuously.
Principal Approach: Treats fail-static's hidden cost explicitly: the property that saves you during outages also hides failures during normal operation. Every fail-static system in the org (discovery, config, certificates, feature flags) must report staleness, and the platform's quarterly review includes a staleness distribution. Sets a policy for version skew between control plane and data plane (e.g., N−2 max).
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Identify all stale proxies; force resync or restart in batches. |
| Triage | Audit flow logs for requests to IPs no longer in the registry (misroutes, especially to reused IPs). |
| Quick fix | Upgrade old sidecars; make NACK per-resource. |
| Guardrails | Staleness alert; canary endpoint convergence probe; version skew policy. |
| Post-mortem | Fail static without visibility is silent failure. |
Metrics to Watch: proxy_config_age_seconds distribution, xds_nack_total{resource}, canary_endpoint_convergence_seconds, sidecar_version distribution.
Organizational Follow-up: Version skew policy; staleness dashboards for all control planes.
Ownership Question: "Who is accountable for a proxy that's been stale for a week?" Staff answer: Platform. The service team can't see it and shouldn't need to. A platform that fails static owes the org a staleness SLO.
Key Takeaway: "Fail static is a feature during outages and a bug when nobody measures staleness."
What clears the Staff bar:
- Connects the resilience property to its failure mode
- Adds per-proxy staleness telemetry and a continuous convergence probe
- Handles NACK granularity
Deep Dive 3: Onboarding a Large Service Into the Mesh#
Context: The ads-serving team (8,000 pods, 400K RPS, p99 budget of 15ms) wants to join the mesh for mTLS and uniform discovery. They're worried about the sidecar's latency and the control-plane load their churn will add — they autoscale by ±30% daily.
Questions to Surface First:
- What's their current discovery mechanism and latency overhead?
- What's the sidecar's measured p99 overhead at their request size and QPS?
- How many callers do they have, and how many services do they call?
- Can they use proxyless xDS instead of a sidecar?
Typical L5 Approach: Deploys sidecars with default resources, measures, finds +1.5ms p99, and the ads team rejects the mesh.
Staff Approach: Proposes proxyless gRPC xDS for ads (same control plane, no extra hop), with mTLS via the xDS-provided certificates. Handles their churn with subsetting (each caller sees ~100 of 8,000 pods) and debounced incremental pushes. Validates in shadow: control-plane load test simulating their daily ±30% swing, before onboarding.
Principal Approach: Recognizes this as a platform product question: the mesh must offer tiers (sidecar, proxyless, per-node proxy) with clear latency and feature tradeoffs, or the highest-value services will opt out and the platform fragments. Owns the published tier matrix and the latency SLO for each.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Assessment | Benchmark sidecar vs proxyless at 400K RPS; measure control-plane load from ±30% churn (2,400 pod changes/day, mostly in 2 ramps). |
| Design | Proxyless xDS, subsetting, debounced deltas. |
| Shadow | Control plane computes ads endpoints and pushes to a shadow subscriber; compare to current system. |
| Rollout | 1% of callers → 10% → 50% → 100% over 3 weeks, rollback per cohort. |
| Success | p99 overhead < 0.3ms; convergence p99 < 5s during daily ramps. |
Metrics to Watch: ads.latency_p99, xds_push_latency_p99, control_plane_cpu during ramps.
Organizational Follow-up: Tier matrix published by platform; ads becomes the reference customer for proxyless.
Ownership Question: "If the mesh adds latency that costs ads revenue, who decides whether to proceed?" Staff answer: The ads team owns their latency budget and can veto; the platform owns offering an option that fits. Security requirements (mTLS) are the lever that makes 'opt out forever' not an option — so the platform must deliver a low-overhead path.
Key Takeaway: "A platform that only has one shape loses its most important customers."
What clears the Staff bar:
- Offers proxyless as an alternative rather than forcing the default
- Quantifies churn impact on the control plane before onboarding
- Uses shadow mode and cohort rollout
Deep Dive 4: Post-Mortem — Multi-Day Outage in the Discovery Layer#
Context: Your company's discovery and configuration platform — a Consul-like system that almost every service depends on for both endpoints and key/value config — degraded after enabling a new streaming feature under record load. The store's performance collapsed, and because services required it at startup and for config reads, recovery took days. (The public Roblox post-mortem from October 2021 describes a 73-hour outage with a similar shape: a Consul streaming change plus a storage-layer pathology under high load.) You're leading the post-mortem.
Questions to Surface First:
- Which services required the discovery system at startup versus only for updates?
- Was the same cluster used for discovery, config, and leader election — one failure domain for three functions?
- Why couldn't the team roll back quickly — was the feature change entangled with data format changes?
- Did monitoring depend on the system that failed?
Typical L5 Approach: Fixes the specific bug (disable the new feature, tune the storage layer), adds more capacity to the cluster, adds alerts on its latency.
Staff Approach: Separates functions and removes startup dependencies. Endpoint discovery, dynamic config, and coordination move to separate clusters with separate failure postures. Services boot with a cached/baked config and last-known endpoints, so a discovery outage doesn't prevent restarts. Monitoring and incident tooling get an independent path. Feature changes to the platform roll out per-cluster with dwell time and a tested rollback.
Principal Approach: Frames the discovery/config platform as the company's single largest correlated-failure risk. Requires cell architecture — multiple independent clusters, each serving a subset of services — so no single platform failure takes out more than a fraction of the business. Sets a "cold start without control plane" requirement for tier-1 services and tests it quarterly. Budgets the headcount for it explicitly, because this is an insurance purchase that leadership must approve.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (during) | Stop the bleeding: disable the new feature, shed load from the store (reduce watchers, pause non-critical clients). |
| Triage | Map which services were blocked on startup vs degraded; that list is the post-mortem's core artifact. |
| Quick fix | Cached config at boot; last-known endpoints persisted to disk. |
| Guardrails | Separate clusters for discovery / config / coordination; independent monitoring path. |
| Post-mortem | Themes: shared failure domain, startup dependency, rollback path, load testing at record scale. |
Metrics to Watch: store_commit_latency_p99, raft_leader_changes, watchers_total, services_blocked_on_startup.
Organizational Follow-up: Cell architecture proposal; quarterly "cold start without control plane" game day; change management for platform features.
Ownership Question: "Who decided that every service should read config from the discovery cluster at startup?" Staff answer: Nobody decided — it accreted, one convenient integration at a time. That's the finding: platform dependencies need an explicit registry and review, because convenience creates coupling nobody signed off on.
Key Takeaway: "The worst discovery outages aren't about routing. They're about everything else that quietly started depending on the discovery cluster."
What clears the Staff bar:
- Separates functions by failure posture
- Removes startup dependencies
- Makes monitoring independent of the failed system
Deep Dive 5: Multi-Region Expansion#
Context: The company is adding two regions to its one. Leadership wants "any service can call any service in any region." The platform team proposes one global discovery cluster stretched across all three regions.
Questions to Surface First:
- What's inter-region RTT (typically 30–150ms)? A stretched consensus cluster pays that on every write.
- What happens to each region when the WAN partitions?
- Do services actually need cross-region calls, or is it "just in case"?
- Which services hold regional data, and would a cross-region call even return the right answer?
Typical L5 Approach: Accepts a stretched cluster with 5 nodes across 3 regions for global consistency.
Staff Approach: Rejects the stretched cluster: writes pay WAN RTT on every endpoint change, and a partition can leave a region without quorum — freezing its discovery for reasons unrelated to its own health. Proposes one registry per cluster (authoritative locally), federated read-only views of other regions, locality-first routing, and cross-region calls only for an allowlist with capacity caps.
Principal Approach: Pushes back on the requirement: "any service can call any region" becomes an implicit global dependency graph that makes every region's availability depend on every other's. Proposes regional independence as the default contract, with explicit, reviewed cross-region dependencies. Prices the design options — per-region stacks cost roughly N× platform infrastructure, but a stretched design risks correlated global failure.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Design | Per-cluster registries; federation for remote read-only views with locality labels. |
| Routing | Zone → region → remote allowlist with 20% capacity caps. |
| Partition behavior | Each region continues with local endpoints; remote endpoints fail static then age out after hours. |
| Rollout | Federation enabled for a few services first; measure cross-region call rate. |
| Testing | WAN partition game day: sever inter-region links for 30 min at low traffic. |
Metrics to Watch: cross_region_rq_ratio, federation_sync_lag_seconds, region_partition_detected.
Organizational Follow-up: Cross-region dependency registry with architecture review.
Ownership Question: "Who approves a new cross-region dependency?" Staff answer: The architecture review board, with the calling team demonstrating the capacity cap and failure behavior. Cross-region calls are a global coupling; they get the same scrutiny as a new database.
Key Takeaway: "Consensus across regions makes every region's discovery depend on the WAN. Keep authority local; share views, not quorums."
What clears the Staff bar:
- Rejects stretched consensus with a concrete reason (WAN RTT, partition quorum)
- Keeps authority local and federation read-only
- Caps cross-region failover by capacity
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Explain why the data plane must fail static and never read the registry per request
- Distinguish endpoint discovery (AP-friendly) from coordination discovery (CP-required) and design each
- Derive a deploy drain wait from convergence p99 and describe the full shutdown sequence
- Explain layered health (readiness, outlier ejection, panic threshold) and why readiness must not check shared dependencies
- Compute control-plane fan-out and fix it with scoped, incremental, debounced pushes
- Explain subsetting for very large services
- Design multi-cluster discovery with local authority, federation and capacity-capped failover
- Split ownership: platform registers and distributes; service teams own readiness and graceful shutdown
The Bar for This Question#
Mid-level (L4): Knows registries exist; describes register-heartbeat-lookup. Might suggest DNS.
Senior (L5): Picks a reasonable registry, explains heartbeats and client caching, mentions health checks and load balancing. Misses fail-static as a principle, deploy-time errors, mass deregistration, control-plane fan-out, and ownership.
Staff+ (L6): Separates control and data planes, designs for churn with a convergence budget, layers health with a panic threshold, handles control-plane failure and recovery storms, scopes fan-out, keeps authority per cluster, and splits ownership between platform and service teams. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "Service Discovery Should Be AP — Even If the Registry Is CP"#
| Evidence | Implication |
|---|---|
| Endpoints change faster than any registry observes (crashes take seconds to detect) | Every view is already stale; "strong consistency" is an illusion at the data plane |
| Quorum loss in a CP registry blocks writes | Routing that depends on writes stops |
| Retries + outlier ejection handle stale endpoints cheaply | Staleness costs a few retries; unavailability costs everything |
The Staff position: Let the registry be CP if it's convenient; make the system AP by caching everything in the data plane.
Why this matters in interviews: It resolves the classic "Eureka vs ZooKeeper" debate by separating the planes — and shows you know the debate exists.
10.2 "Self-Registration Is an Anti-Pattern on Modern Platforms"#
Every language needs a correct implementation of registration, heartbeating and graceful deregistration. The orchestrator already knows what's running and whether it passed readiness. Third-party registration from the orchestrator has one implementation, one owner, and one set of bugs.
The Staff position: Platform registers; services expose readiness. Self-registration only where no orchestrator exists.
Why this matters in interviews: It's a clean ownership argument that interviewers rarely hear.
10.3 "DNS Is a Terrible Service Discovery System That You Must Support Forever"#
| DNS Weakness | Consequence |
|---|---|
| Clients and resolvers cache beyond TTL | Minutes of staleness |
| No health or weight semantics beyond basic records | Can't do outlier ejection or zone-aware LB |
| Response size limits | Large services truncated or require TCP fallback |
The Staff position: DNS is the compatibility layer for everything that can't speak your control-plane protocol — databases' clients, third-party tools, legacy code. Support it; don't design around it.
Why this matters in interviews: Saying "DNS with a short TTL" as the whole answer is a common L5 signal; saying why it's only a fallback is L6.
10.4 "Health Checks Cause More Outages Than They Prevent"#
Readiness checks that probe shared dependencies, liveness checks that restart healthy-but-busy processes, and fleet-wide health config changes have caused more full-service outages than the dead instances health checks remove — which outlier detection would have caught anyway.
The Staff position: Health checks should be local, conservative, and capped (panic thresholds, removal guards). Caller-observed failure is the more reliable signal.
Why this matters in interviews: It shows you've debugged a mass deregistration — or read enough post-mortems to know the pattern.
10.5 "Your Discovery Cluster Is Secretly Your Config System, Lock Service, and Monitoring Dependency"#
Convenience accretes: "it's already there, let's store feature flags in it", "let's use it for leader election", "let's have the metrics agent discover targets from it". Each is reasonable; together they make one cluster a single point of failure for functions with different availability needs.
The Staff position: One cluster per failure posture. Audit what depends on the discovery cluster every quarter.
Why this matters in interviews: It's the Principal-flavored insight you can deliver at Staff level.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
At L6, discovery is a well-designed system with a fail-static data plane. At L7, discovery is the nervous system of the company's infrastructure — the one control plane every service touches, the place where security identity, traffic policy, config and dependency graphs converge, and therefore the org's largest shared failure domain. The Principal questions are about scope and time: how many discovery systems should exist (usually one, sometimes one per failure posture), what it should refuse to become (a config store, a lock service), which of its contracts are forever (service naming, identity format, the manifest), and how to make it survivable when — not if — it fails.
The Org-Level Fault Line#
One discovery platform vs many.
| Option | What It Buys | What It Costs | Who Pays |
|---|---|---|---|
| One global platform for everything (discovery + config + locks) | One integration, one team, one dashboard | A single correlated failure domain; startup dependencies everywhere | Every service, in the multi-day outage |
| Per-team or per-stack systems (Eureka for JVM, Consul for VMs, K8s for new services) | Local fit and speed | No global dependency graph, no uniform identity, three on-calls, bridges | Platform (bridges), security (inconsistent mTLS), on-call |
| One platform, cellular, split by posture (the L7 default) | Uniform contract and identity; failures bounded to a cell; coordination separated from endpoints | Higher infra cost; more clusters to operate; cell assignment decisions | Platform headcount; finance approves the insurance premium |
🧭 Principal Move: "I want exactly one way to name, identify and discover a service — and more than one cluster implementing it. Standardize the contract; distribute the failure domain."
Cost Model#
Assumptions: control-plane nodes ~8 vCPU each; ~$30/vCPU-month; sidecar ~0.15 vCPU + 75MB per pod where used; fully loaded engineer ~$300K/year. Order of magnitude only.
| Scale | Setup | Monthly Infra | Headcount | On-call Load |
|---|---|---|---|---|
| Startup (30 services, 500 pods, 1 cluster) | Kubernetes Services + DNS; no mesh | ~$0 incremental (bundled in the orchestrator) | ~0.1 FTE | Shared infra rotation; rare pages |
| Growth (400 services, 15K pods, 4 clusters) | Mesh control plane per cluster, sidecars for most, DNS fallback | Control plane ~$3–6K; sidecars ~$65K | 3–5 FTE networking/platform | Dedicated platform rotation; ~2–4 pages/month |
| Large (3,000 services, 150K pods, 20+ clusters, cells) | Cellular control planes, proxyless for hot paths, federation, staleness SLOs | Control plane ~$30–60K; data plane $300–600K (proxyless cuts this) | 12–20 FTE | 24/7 follow-the-sun; control-plane changes staged by cell |
The line item that surprises leadership is the data plane, not the registry. The strategic decision is sidecar vs proxyless vs per-node proxies, which can swing data-plane cost by 3–5×.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse | Why |
|---|---|---|---|
Service naming scheme (svc.namespace.cluster) | One-way | Every config, dashboard, policy and client | Names outlive services and teams |
| Workload identity format (e.g., SPIFFE IDs) | One-way | Every authz policy and certificate | Security policy is written against it |
| Registry implementation behind the platform API | Two-way (if abstracted) | Months | Keep an abstraction so it can change |
| Sidecar vs proxyless default | Expensive two-way | Fleet-wide rollouts | Revisit every ~2 years as tech evolves |
| Putting config/locks in the discovery cluster | Easy to do, hard to undo | Quarters of untangling | Say no early |
| Stretched cross-region consensus | Hard to undo | Re-architecture | Avoid; federate instead |
The Standard I'd Write#
RFC: Service Discovery & Routing Standard v1
Scope: All production services and clients in all clusters.
Requirements:
- Services MUST be registered by the platform from orchestrator state; self-registration is prohibited except on approved non-orchestrated hosts.
- Services MUST expose a readiness endpoint that checks only local state. Readiness MUST NOT depend on shared downstream services.
- Services MUST implement graceful shutdown: fail readiness on SIGTERM, continue serving for ≥ the published convergence p99 + 2s, then drain.
- Data-plane components MUST route from local state and MUST continue with last-known-good state if the control plane is unreachable. Last-known-good MUST survive process restart.
- The control plane MUST NOT remove more than 30% of a service's endpoints within 60s without human confirmation; proxies MUST use a panic threshold of 50%.
- Services SHOULD declare their dependencies in the service manifest; proxies subscribe only to declared dependencies.
- Discovery clusters MUST NOT be used for application config, locks or leader election.
Exceptions: Filed with the platform team; reviewed monthly; 6-month expiry.
Success metrics: Convergence p99 < 5s during peak deploys; deploy-attributable error rate < 0.01%; % of proxies with config age > 5 min < 0.1%; zero incidents where control-plane failure stopped existing traffic.
What I'd Tell the VP#
"Every service we run finds every other service through one system, which makes it both essential and our biggest single risk. We've designed it so that if it fails, the traffic that's already flowing keeps flowing — we lose the ability to deploy for a while, not the business. Over the next year we'll retire our three older discovery systems into one, and split it into independent cells so a failure affects a fraction of services, not all of them. The main cost is the proxy fleet, and we have a plan to cut that by moving our heaviest services to a lighter-weight option. The outcome you'll see is deploys that don't cause customer errors and no company-wide outages from this layer."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Standardizes contracts, distributes failure | "One naming and identity scheme; many cells implementing it." |
| Refuses scope creep | "The discovery cluster will not store config or locks — different failure posture." |
| Knows the forever decisions | "Service names and identity formats are one-way doors; the registry implementation isn't." |
| Prices the data plane | "The registry is cheap; sidecars are the six-figure monthly line, and proxyless changes it." |
| Makes staleness a platform SLO | "Fail static is only safe if we publish config-age distributions." |
Staff answers that L7 interviewers find insufficient:
- An excellent single-cluster design with no answer for how 20 clusters and 3 regions share naming and identity.
- "We'll migrate everything to the mesh" without a plan for the long tail or the cost of running two systems meanwhile.
- Treating the discovery cluster's other uses (config, locks) as someone else's problem.
Appendices
Appendix A: Mechanics in Depth#
A.1 Registration Patterns#
| Pattern | Flow | When Right | When Wrong |
|---|---|---|---|
| Self-registration + heartbeat | Instance → registry: register, renew every N s | No orchestrator; small fleets | Polyglot; deadlocked instances keep heartbeating |
| Agent registration (sidecar/agent on host) | Local agent health-checks instance, registers it | VMs; SmartStack/Consul-agent style | Agent failure = instance invisible |
| Orchestrator registration | Controller derives endpoints from ready pods | Kubernetes-like platforms | Orchestrator API scale limits at huge fleets |
A.2 Failure Detection Math#
Heartbeat-based removal time ≈ heartbeat_interval × missed_threshold + propagation
Eureka-style: 30s × 3 + client refresh 30s ≈ up to 120s
Readiness-based removal (tuned): probe_period 2s × failureThreshold 2 + push 2s ≈ 6s
Outlier ejection (caller-local): 5 consecutive 5xx at 500 RPS ≈ 10ms; at 1 RPS ≈ 5s
Graceful shutdown: 0s detection (the pod fails its own readiness) + push p99
Caller-side outlier detection is the fastest signal for high-traffic callers and the slowest for low-traffic ones; the registry is the opposite. Use both.
A.3 Incremental Push With Versions#
Subscriber state: {service → version}
On connect: Watch(services, known_versions)
Server: for each service: if known_version == current: nothing
elif delta_log has known_version: send delta
else: send full state
On change: debounce 100ms–1s per service, then fan out delta to subscribers of that service
Client: apply; ACK(version) or NACK(version, error) and keep previous
A.4 Deterministic Subsetting#
def subset(backends, client_id, subset_size):
subset_count = ceil(len(backends) / subset_size)
round = client_id // subset_count
shuffled = shuffle(sorted(backends), seed=round) # same shuffle for a round of clients
subset_id = client_id % subset_count
return shuffled[subset_id * subset_size : (subset_id + 1) * subset_size]
Each "round" of clients covers every backend exactly once, so load is balanced across backends while each client holds only subset_size connections.
Appendix B: Naming, Identity and Data Model#
| Field | Example | Notes |
|---|---|---|
| Service name | payments.checkout.prod | name.namespace.env; immutable once published |
| Endpoint | 10.4.17.22:8443 | Never the identity — IPs are reused |
| Workload identity | spiffe://corp.example/ns/checkout/sa/payments | What mTLS verifies |
| Locality | region=us-east-1, zone=us-east-1b, cluster=use1-a | Drives routing preference |
| Metadata | version=2026.09.3, canary=false, weight=100 | For traffic splitting |
| Readiness | ready / not_ready / terminating | Terminating ≠ gone: still draining |
Rule: Route by name, verify by identity, never trust an IP.
Appendix C: Coordination Mechanisms — Registry Options#
| System | Consistency | Health Model | Propagation | Best For |
|---|---|---|---|---|
| etcd (via Kubernetes) | CP (Raft) | Orchestrator readiness | Watch | Kubernetes-native platforms |
| ZooKeeper | CP (ZAB) | Ephemeral nodes tied to sessions | Watch (one-shot, re-arm) | Coordination; legacy discovery |
| Consul | CP servers (Raft) + gossip membership | Agent checks + gossip failure detection | Blocking queries / streaming | VM fleets; multi-platform |
| Eureka | AP (peer replication) | Client heartbeats, self-preservation | Client poll (30s) | Legacy JVM on cloud VMs |
| DNS | Eventually consistent | Minimal | Resolver TTL | Universal fallback |
C.1 Quick Comparison#
| Question | CP registry | AP registry |
|---|---|---|
| During partition | Minority side can't write; may not read | Both sides accept writes; views diverge |
| Ghost endpoints | Short (sessions expire) | Longer (lease expiry, self-preservation) |
| Coordination primitives | Yes (leases, locks, sequences) | No |
| Data plane still needs caching? | Yes | Yes |
See ZooKeeper & etcd and Distributed Consensus.
Appendix D: Client Behavior Contract#
D.1 Client/Proxy Requirements#
| Behavior | Requirement |
|---|---|
| Routing source | Local endpoint table only |
| Control plane unreachable | Keep last-known-good; persist to disk; reconnect with jittered backoff (1s → 60s) |
| Bad update | Validate; NACK; keep previous |
| Connect failure | Retry on another endpoint (safe: request not sent) |
| Long-lived connections | Honor server max-connection-age (e.g., 5 min) to rebalance to new endpoints |
| LB policy | Least-request / power-of-two-choices over the subset; zone-aware |
D.2 Reconnect Storm Prevention#
backoff = min(60s, 1s × 2^attempt) × random(0.5, 1.5)
on reconnect: send known_versions → receive deltas, not full state
server: admission control — at most N concurrent full-state syncs per replica
D.3 Graceful Shutdown Contract#
SIGTERM →
set readiness = false
sleep(convergence_p99 + 2s) # e.g., 7s
stop accepting new connections; send GOAWAY / Connection: close
wait for in-flight (max 15s)
exit
terminationGracePeriodSeconds ≥ sleep + drain + margin (e.g., 30s)
Appendix E: Observability#
E.1 Core Metrics — Non-Negotiable#
# Control plane
registry_write_latency_p99, registry_has_leader
xds_push_latency_seconds{p50,p99}
xds_connected_clients, xds_nack_total{resource}
endpoint_changes_per_minute{service}
mass_removal_guard_triggered_total{service}
# Data plane
proxy_config_age_seconds (distribution across proxies)
endpoints_healthy_ratio{service} (per caller)
panic_mode_active{service}
outlier_ejections_active{service}
upstream_cx_connect_fail{service}, no_healthy_upstream_total{service}
# End-to-end
canary_endpoint_convergence_seconds # synthetic change, observed in sampled proxies
deploy_attributable_error_rate{service}
E.2 Critical Alerts#
| Alert | Condition | Action |
|---|---|---|
| Registry leaderless | registry_has_leader == 0 for 30s | Page platform; freeze deploys |
| Convergence degraded | canary_endpoint_convergence_seconds p99 > 10s for 10 min | Page platform |
| Stale proxies | > 1% of proxies with config age > 5 min | Ticket; page if > 5% |
| Panic mode | panic_mode_active for a tier-1 service | Page service owner + platform |
| Mass removal held | Guard triggered | Page platform to confirm or reject |
| No healthy upstream | no_healthy_upstream_total rising | Page service owner |
E.3 Control Plane vs Data Plane#
Control-plane dashboards tell you whether changes are flowing; data-plane dashboards tell you whether traffic is flowing. During a control-plane outage, the first set is red and the second should be green. If both go red together, the fail-static property is broken — that is a Sev-1 design bug, not just an outage.
E.4 Debugging "Traffic Isn't Reaching New Instances"#
- Is the pod ready? (
pod_ready_seconds) - Is it in the registry? (endpoint slice contents)
- Did the push reach callers? (
proxy_config_age, sampled endpoint tables) - Are callers opening connections to it? (long-lived HTTP/2 connections, pick-first LB, subsetting)
- Is it in any caller's subset?
Appendix F: Scale Evolution#
F.1 What Works at Each Scale#
| Scale | Enough | Add When |
|---|---|---|
| < 50 services | Orchestrator Services + DNS | Deploy errors, second language |
| 50–500 services | Mesh control plane per cluster, sidecars, drain standard | Control-plane load during deploys |
| 500–3,000 services | Scoped incremental pushes, subsetting, staleness SLOs | Multi-region, correlated platform risk |
| 3,000+ / multi-region | Cellular control planes, federation, proxyless tiers | — |
F.2 Multi-Region Path#
Per-cluster authority → read-only federation → locality-aware routing with capped remote failover → edge-level regional evacuation. Never stretched consensus.
F.3 What You Don't Build on Day One#
- Your own registry or xDS server
- Cross-region federation before a second region exists
- Subsetting before any service exceeds ~1,000 instances
- Proxyless xDS before the sidecar tax is measured
- A custom service manifest format before you have 50+ services
Appendix G: Multi-Tenancy, Fairness and Cost#
G.1 Control-Plane Fairness#
One namespace's churn (a team autoscaling every 30 seconds) can dominate control-plane capacity. Rate-limit endpoint changes per namespace (e.g., ≤ 200/min before debounce extends), and prioritize tier-1 services' pushes.
G.2 Registry Write Quotas#
Per-namespace quotas on registered objects and write rate protect the consensus store (etcd performs best well under its 2–8GB size range).
G.3 Cost Levers#
| Lever | Savings | Tradeoff |
|---|---|---|
| Proxyless for top 10% QPS services | 30–60% of data-plane CPU | Language/feature limits |
| Scoped subscriptions | 10–100× control-plane egress | Requires declared dependencies |
| Sidecar right-sizing | 20–40% of sidecar requests | Needs per-service measurement |
| Per-node proxy model | Fewer proxy instances | Weaker per-workload isolation |
G.4 Tradeoff Summary#
| Choice | Default | Who Pays If Wrong |
|---|---|---|
| Registry consistency | CP writes, cached AP data plane | Everyone, if the data plane reads the registry per request |
| Registration | Platform from orchestrator | Every team, if self-registration drifts |
| Propagation | Scoped incremental push | Platform, during large deploys |
| Health | Local readiness + outlier ejection + panic | Service owners, in mass deregistration |
| Control-plane outage | Fail static, visible staleness | Security/correctness if staleness is silent |