Technologies referenced in this case study: Redis · API Gateways · ZooKeeper & etcd · Apache Kafka
Related: API Gateway · Circuit Breakers · Idempotency & Exactly-Once · Multi-Region Active-Active · Degraded Mode Framework · Dealing with Contention · Consistent Hashing · API Design Patterns
How to Use This Case Study#
Built for the interview first and the pager second. Read it once end to end, then come back to the fault line or deep dive that matches the question you dread.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → 30-Second Cheat Sheet → Interview Walkthrough → Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Deep Dives 1 and 4 |
| Deep Dive | 3+ hrs | Everything, including Section 11 (Principal Lens) and Appendices A (algorithms), C (coordination) and G (multi-tenancy and cost) |
What is a Distributed Rate Limiter? — Why interviewers pick this topic
A rate limiter decides, for each incoming unit of work, whether to admit it now, delay it, or refuse it — based on how much work the same identity (an API key, a tenant, a user, an IP, a route) has already consumed in some window. "Distributed" means the decision is made on dozens or hundreds of nodes that each see only a slice of the traffic, while the limit is defined over the whole fleet.
The hard part is not the counting. The hard part is that the limiter sits on the hot path of every request, so it must be cheaper than the work it protects, it must keep working when its own state store is sick, and every "no" it says lands on a real customer who may have paid for that capacity.
Before vs After — the "partner integration goes into a retry loop" scenario:
Without a deliberate limiter (only a per-node in-memory counter, 100 req/s per node):
t=0: A partner's sync job ships a bug: it retries every 5xx immediately, no backoff.
t=+20s: Partner traffic climbs from 300 req/s to 9,000 req/s. Spread over 40 gateway
nodes, that is 225 req/s per node — each node sees "only" 2.25× its local cap,
drops half, and the other 4,500 req/s reach the orders service.
t=+1min: Orders DB connection pool (400 conns) saturates. p99 for ALL tenants: 80ms → 6s.
t=+2min: Timeouts cause more partner retries. 14,000 req/s. Every tenant sees 5xx.
t=+25min: On-call finds the partner by grepping access logs, blocks the key by hand.
Status page: "degraded API performance" for 31 minutes, 1,900 tenants affected.
With a tenant-scoped fleet limit + concurrency guard + explicit 429 contract:
t=0: Same bug ships.
t=+2s: Partner exceeds its 500 req/s contract limit. Gateway returns 429 with
Retry-After: 2. Partner's tenant concurrency cap (64 in-flight) also holds.
t=+5s: limiter.rejected{tenant=partner_77} spikes. Alert to the partner-success channel,
not the pager — the system is doing its job.
t=+10min: Partner support emails the partner with the exact rejected volume and the fix.
Other 1,899 tenants: p99 unchanged at 80ms. No status page.
Why interviewers reach for this question: It looks like an algorithms question (token bucket vs sliding window) and almost every candidate answers it as one. It is actually a question about shared state on the hot path, failure posture, and who is allowed to say "no" to a customer. It separates people who have read about rate limiting from people who have been paged by a rate limiter.
Mechanics Refresher: The Five Counting Schemes
| Scheme | How It Works | Pros | Cons |
|---|---|---|---|
| Fixed window counter | INCR key:{minute}; reject when count > limit; key expires after the window | 1 op per request; trivial to shard; cheap memory (1 integer per key per window) | Up to 2× the limit at a window boundary (limit at 0:59, limit again at 1:00); synchronized client resets at :00 |
| Sliding window log | Store every request timestamp in a sorted set; count entries newer than now − window | Exact | Memory O(limit) per key — 10K/min limit = 10K entries per key; ZREMRANGEBYSCORE cost grows with rate |
| Sliding window counter (weighted) | Keep current + previous fixed-window counts; estimate = prev × (overlap fraction) + current | 2 integers per key; boundary burst error typically within a few percent for smooth traffic | An approximation — assumes the previous window's traffic was uniform |
| Token bucket | Bucket of capacity B refills at rate R; each request takes 1 (or cost) tokens; state = (tokens, last_refill_ts) | Explicit burst (B) and sustained rate (R) as two independent, explainable knobs; supports weighted cost | Two fields to update atomically; needs a server-side script or CAS in a shared store |
| GCRA (generic cell rate algorithm) | Store one timestamp, the "theoretical arrival time" (TAT); admit if now ≥ TAT − burst_tolerance, then TAT = max(TAT, now) + emission_interval | Token-bucket semantics with one field; constant memory; easy to compute Retry-After exactly | Less intuitive to explain; still needs an atomic read-modify-write |
A sixth mechanism is not a counting scheme at all, and it matters more than the five above for protecting backends: the concurrency limiter — cap the number of in-flight requests per identity (a semaphore with a lease), released on completion or timeout. Rate limits cap arrivals; concurrency limits cap occupancy, and occupancy is what actually exhausts thread pools and DB connections.
For most production systems: token bucket (or GCRA, its one-field equivalent) for per-tenant contract limits, a concurrency cap per tenant for expensive routes, and a separate adaptive load shedder that ignores identity and protects the service itself. The algorithm choice takes 60 seconds of the interview. Spend the remaining 40 minutes on where the state lives, what happens when it is unavailable, and who owns the numbers.
Executive Summary
If you only read one section, read this. The rest of the case study expands these pages.
What This Interview Actually Tests#
A distributed rate limiter is not a counting question. Counting is solved.
It is a capacity-allocation question: when demand exceeds what you are willing to serve, who gets served, who gets told no, how they are told, and what happens when the thing doing the telling is itself broken. It tests:
- Whether you can tell apart three jobs that all get called "rate limiting" — enforcing a commercial contract, repelling an adversary, and keeping a service alive — and refuse to build one mechanism for all three
- Whether you put shared mutable state on the hot path of every request knowingly, with a latency budget and a fallback, or by default
- Whether you choose the failure posture (admit or reject when the limiter is blind) per intent, and name who signed off
- Whether you treat the
429response as an API contract that clients will build against, not as an error
The key insight: A limiter's precision is worth less than its availability. A limiter that is 3% imprecise and never adds more than 1ms is a feature; a limiter that is exact and occasionally adds 200ms is an outage generator. Staff candidates trade precision for locality on purpose, put a number on the overshoot, and get the product owner to accept it.
The L5 vs L6 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "I'll use a token bucket in Redis keyed by user ID" | Asks "Is this protecting revenue contracts, repelling abuse, or keeping a backend alive? Those want different keys, different failure modes and different owners." | Asks "How many teams already run their own limiter, and which of them would page someone else if they got it wrong?" — treats limiting as a fleet-wide policy surface |
| State placement | Central Redis call on every request | Local token buckets on each node, synchronized with a shared store every 100–500ms; overshoot bounded and stated as a number | Sets a platform rule: no synchronous remote call on the request path for any limit with a tolerance > 5%; exact limits require a written justification |
| Failure | "Redis is replicated, so it's fine" | "If the store is slow beyond 2ms, contract limits fail open to the last-known local budget; abuse limits on login fail closed; load shedding never depends on the store at all" | Owns the org's limiter failure posture table, signed by product and security leadership, and game-days it twice a year |
| Client contract | Return 429 | 429 with Retry-After, remaining-quota headers, a stable error body, and documented backoff — plus SDKs that honor it | Treats quota as a product: plans, upgrade paths, quota dashboards for customers; 429 rate per tenant is a revenue signal reviewed with sales |
| Policy change | Edit the config and deploy | Shadow → log-only → 1% canary → enforce, with per-tenant override and a kill switch that needs no deploy | Separates who proposes limits (product), who approves them (capacity/SRE), and who executes them (platform); limit changes on top-50 customers need account-team sign-off |
| Scale | "Shard Redis" | "At 2M req/s the store sees ~20K sync ops/s because nodes batch; the real scaling problem is hot tenants, not total traffic" | Prices the limiter: $/month of store + gateway CPU per million requests, and the cost of a false 429 on an enterprise account |
Why "state placement" separates levels
L5: "Every gateway node calls Redis with a Lua script that does the token-bucket math atomically." This is correct, exact, and it puts a network hop — 0.3ms median, 2–5ms p99 inside an AZ, much worse during a failover — in front of every request in the company. At 500K req/s it also means 500K script executions per second against a store whose single-shard ceiling is in the low hundreds of thousands of simple ops.
L6: "Each node keeps a local bucket per active key and periodically reconciles with the shared store: it reports what it consumed and receives its share of the remaining budget. The request path never blocks on the network. The cost is overshoot — with 40 nodes and a 250ms sync interval, a tenant can exceed its limit by at most what the nodes admit between syncs, which I'll cap by giving each node a lease of a slice of the budget rather than the full bucket."
L7: "I'd make that the paved road. The platform ships the local-plus-sync library; teams that want an exact, synchronous limit — say for a one-time-password endpoint — get it, but they file the latency budget and the failure posture with the platform team, because their choice affects everyone sharing the store."
Why "failure" separates levels
L5: Treats the store as infrastructure that should not fail; adds replicas. When asked "Redis is down, now what?", improvises: "We'd allow the requests."
L6: Has the answer before the question, and it differs by intent. Contract limits: fail open to a local, conservative per-node budget (limit ÷ node count × 1.5), because a few minutes of over-admission costs some backend headroom, while failing closed turns a cache incident into a total API outage. Login and signup abuse limits: fail closed to a strict local cap, because the attacker is waiting for exactly that window. Backend-protection shedding: never depends on the shared store in the first place.
L7: Recognizes that "fail open" is a business decision, not an engineering one. "Security owns the login posture; the API product owner owns the contract posture; the platform owns making both postures executable without a deploy. I'd want those three names in a table before launch."
Why "policy change" separates levels
L5: Limits live in a YAML file in the gateway repo. Changing one is a code review and a deploy.
L6: Limits are data, not code — stored in a policy service, versioned, pushed to nodes within ~30 seconds, and every change rolls through shadow mode: the new policy computes its verdict alongside the old one and logs would_reject without enforcing. "I want to know which 14 customers a new limit would have rejected yesterday before it rejects anyone today."
L7: Recognizes that the dangerous changes are not code bugs but unit and scope mistakes — per-second typed as per-minute, per-tenant applied as global. Requires every policy to declare its expected rejection rate and blocks any change whose shadow rejection rate exceeds 10× the declared value.
The Staff Positions#
| Position | Rationale |
|---|---|
| One mechanism per intent — contract quotas, abuse defense and load shedding are separate layers | They differ in key, failure posture, precision need and owner; merging them makes every change a three-team negotiation |
| The request path never waits on a remote store for a tolerant limit | Local buckets + periodic sync keep added latency under 50µs; the store becomes a coordinator, not a dependency |
| State the overshoot as a number and get it accepted | "Up to 5% over contract for up to 2s" is a product decision someone signs; "approximately enforced" is not |
| Failure posture is chosen per intent and written down | Contract: fail open to a local budget. Credential endpoints: fail closed. Shedding: no store dependency at all |
| Concurrency caps protect backends; rate caps protect contracts | A tenant at 50 req/s of 30-second report exports holds 1,500 workers; a rate limit never sees it |
| The 429 is an API contract | Retry-After, remaining/reset headers and a stable error code — clients will build retry logic on exactly what you return |
| Every policy change ships through shadow mode | The cheapest outage to prevent is the one you cause to your own biggest customer with a typo |
The Three Intents#
Three systems hide under the same name. Name them, pick one to lead with, and keep the others as separate layers.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Contract enforcement (paid API tiers, partner quotas) | Customer paid for N req/s or M calls/month; violations either way cost money or trust | Per-tenant token bucket with burst; local counting + periodic sync; usage metered separately for billing | False 429 on a paying customer (support escalation, churn); silent over-admission (lost upsell, backend cost) | Within ±5% of contract over any 10s window; billing uses the metered log, never the limiter's counters |
| Abuse defense (login, signup, OTP, scraping, card testing) | Adversary rotates identities and probes for the failure window | Multi-key limits (account, IP /24, device fingerprint, ASN), exact counting on small hot sets, reputation-based throttling | Attacker gets free attempts during a limiter outage; legitimate users behind shared NAT get blocked | Zero unthrottled attempts on credential endpoints, even during store failure |
| Capacity protection (keeping a service or dependency alive) | The backend has a hard ceiling (DB connections, GPU slots, a vendor's quota) independent of who is calling | Adaptive concurrency limits and priority-based load shedding at each service; identity used only for fairness ordering | Shedding the wrong traffic (checkout instead of analytics); limiter oscillation | Service stays within its saturation point; p99 latency of admitted work stays inside SLO |
🎯 Staff Move: "I'll lead with contract enforcement for a multi-tenant public API, because that's where the distributed-state tradeoff is sharpest — the limit is fleet-wide but the traffic is spread across 40 nodes. I'll keep abuse defense and capacity protection as separate layers with different keys and different failure postures, and I'll come back to them. If you'd rather I start with login abuse or with load shedding, the design changes, so tell me now."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Where the Count Lives | Exactness of one shared counter vs latency and availability of local counting — and the leasing designs in between |
| 2 | Blind-Limiter Posture | When the limiter cannot see its state, admit (risk overload or abuse) or reject (turn a limiter incident into an outage)? |
| 3 | What Unit You Limit | Requests per second vs weighted cost units vs in-flight concurrency — the unit decides which incidents you can prevent |
| 4 | Static Contract vs Adaptive Capacity | Limits as fixed promises customers can plan around vs limits that move with backend health — predictability vs survival |
| 5 | Who Owns the Number | Platform, product, SRE or the customer's account team — and how a limit change ships without becoming an incident |
In the Wild: Real Production Systems#
Why this section belongs here: Naming mechanisms that real companies have described publicly shows you have studied operated systems, not blog diagrams.
Stripe — Four Limiters, Not One#
Stripe's engineering blog post on scaling its API with rate limiters describes four distinct mechanisms running together: a per-user request rate limiter (token bucket, Redis-backed), a per-user concurrent requests limiter for expensive endpoints, a fleet-wide load shedder that reserves capacity for critical request types, and a worker-utilization shedder that sheds lower-priority traffic first when workers saturate. The post also stresses rolling limiters out in dark mode first, and returning clear 429s so well-behaved clients can back off.
Staff insight: This is the clearest public evidence that "rate limiting" is several layers with different keys and purposes. Saying "I'd separate contract limits from load shedding, the way Stripe describes" lands better than any algorithm comparison.
Envoy and Lyft's Global Rate Limit Service — Local First, Global Second#
Envoy proxy offers two modes: a local rate limit filter (a token bucket inside each proxy, no network call) and a global mode that calls an external gRPC rate limit service — Lyft open-sourced its implementation, which counts in Redis using descriptor-based keys (e.g., domain + remote_address + path). Envoy's documentation recommends combining them: the local limiter absorbs bursts cheaply and the global service enforces fleet-wide quotas, with a configurable failure mode if the service is unreachable.
Staff insight: The two-tier shape — cheap local check that sheds most excess, expensive global check for what remains — is the production pattern. It also shows that the failure mode of the global call is a configuration choice, which is exactly the posture conversation interviewers want.
Netflix — Concurrency Limits Instead of Rate Limits#
Netflix open-sourced a concurrency-limits library that adapts a service's in-flight limit automatically based on observed latency, borrowing ideas from TCP congestion control (Vegas- and gradient-style algorithms). Instead of a human picking "500 req/s", the service measures how latency grows as concurrency rises and shrinks the limit when queuing appears, shedding excess immediately rather than letting it queue.
Staff insight: For the capacity-protection intent, a static req/s number is guaranteed to be wrong after the next deploy changes the cost per request. An adaptive concurrency limit tracks the backend's actual capacity. Naming this is how you show you know rate limiting and load shedding are different tools.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
| "Token bucket in Redis" | "That's a Redis round trip on every request. What's your latency budget, and what happens at p99.9?" | Do you know the cost of state on the hot path? |
| "We'll keep counters locally" | "40 nodes, limit 1,000 req/s. How far over can a tenant get, and for how long?" | Can you bound the overshoot with a formula? |
| "Fail open if Redis is down" | "Including on the login endpoint? Who agreed to that?" | Per-intent posture and sign-off |
| "Key by user ID" | "What about unauthenticated traffic? What about 30,000 users behind one carrier NAT?" | Key design under real network topology |
| "Return 429" | "The client retries immediately. Now what? What does your SDK do?" | Client contract and retry amplification |
| "We'll set the limit at 1,000" | "Where did 1,000 come from? Who changes it, and how do you know a change won't hurt your largest customer?" | Policy ownership and safe rollout |
| "Shard by tenant" | "One tenant is 40% of all traffic. Which shard melts?" | Hot-key awareness |
System Architecture Overview#
Reading the diagram: Three layers, three owners. The edge layer is about adversaries and holds no shared counters. The gateway layer enforces contracts using local buckets that a background loop keeps roughly in sync — the store is never on the request path. The service layer protects itself with an adaptive concurrency limit that doesn't care who the caller is except for priority ordering. Billing reads the metering stream, never the limiter. The single most important metric is
limiter.sync_staleness_ms— if it climbs, enforcement silently degrades to per-node limits.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This |
|---|---|---|
| Algorithm | "Sliding window log for accuracy" | "Token bucket or GCRA — two knobs, burst and rate, O(1) state. The algorithm is a 60-second decision." |
| State | "Redis call per request" | "Local buckets, 250ms sync, per-node leases. Overshoot ≤ 5% for ≤ 1s, stated up front." |
| Store down | "Allow everything" | "Contract limits fall back to limit ÷ N × 1.5 per node. Login fails closed. Shedding never needed the store." |
| Keys | "User ID or IP" | "Tenant × route class for contracts; account + IP /24 + device for abuse; no identity for capacity." |
| Expensive endpoints | "Lower their rate limit" | "Concurrency cap per tenant, or charge cost units per request — rate can't see a 30-second request." |
| Client | "Return 429" | "429 + Retry-After + remaining/reset headers + SDK with jittered backoff. The 429 is a contract." |
| Changes | "Update config" | "Shadow → log-only → canary → enforce; per-tenant overrides; kill switch without deploy." |
| Billing | "Use the limiter's counters" | "Never. The limiter is approximate by design. Billing reads an at-least-once usage stream with dedup." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
| Redis round trip, same AZ | ~0.2–0.5ms p50, 1–5ms p99 | The floor cost of a synchronous limiter; p99 is what customers feel |
| Redis cross-AZ round trip | ~0.5–2ms p50 | Gateways in 3 AZs talking to one primary pay this on 2/3 of calls |
| Single Redis shard throughput | ~100K–200K simple ops/s; ~50–100K/s for small Lua scripts | Why one shard cannot be the per-request limiter for a 1M req/s API |
| Local token-bucket check | ~50–200ns in-process | Five orders of magnitude cheaper than a network call |
| State per token-bucket key | ~2 fields; ~100–200 bytes with Redis overhead | 10M active keys ≈ 1–2 GB — memory is rarely the constraint |
| Fixed-window boundary burst | Up to 2× the limit across the boundary | Why fixed windows are wrong for anything with a hard ceiling |
| Overshoot with N nodes, sync interval S | ≤ N × per-node lease admitted per S | 40 nodes × 25 tokens = 1,000 extra per 250ms worst case — so cap leases |
| Policy propagation target | ≤ 30s fleet-wide | Long enough to batch, short enough that a bad limit is reversible in a minute |
| Healthy 429 rate | < 0.1% of requests overall; per-tenant varies | A fleet-wide 429 rate of 2% is either an attack or a broken policy |
| Retry amplification without backoff | 3 retries = up to 4× load during an overload | Why the 429 contract and SDK behavior matter as much as the limiter |
| Cross-region RTT | ~60–150ms | Why global quotas across regions must be allocated, never checked synchronously |
| GitHub REST API default (authenticated) | 5,000 requests/hour per user | A well-known public example of a contract limit with reset headers |
Interview Walkthrough
A 45-minute interview has room for one strong opinion per phase. The phases below are timed so the depth lands on state placement and failure posture, where the signal is.
Phase 1: Requirements & Framing (2–3 minutes)#
Say this, close to verbatim:
"Before I draw anything — rate limiting means three different things in practice. One: enforcing what a customer's plan allows, which is a contract. Two: stopping attackers on endpoints like login, which is adversarial. Three: keeping a service alive when demand exceeds capacity, which is about the backend, not the caller. They want different keys, different precision and different behavior when the limiter itself fails. I'll design the contract layer for a multi-tenant public API as the core, and show where the other two plug in."
Then pin down the numbers:
| Question | Assumption I'll State | Why It Matters |
|---|---|---|
| Traffic | 500K req/s average, 2M req/s peak, across 40–200 gateway nodes in 3 AZs, one region to start | Decides whether a per-request remote call is even possible |
| Tenants | ~50K tenants; top 20 account for 45% of traffic | Hot tenants, not total traffic, break the store |
| Limit shapes | Per tenant per route class: sustained req/s + burst; some routes have concurrency caps; monthly quotas for billing tiers | Two time scales: seconds (protection) and months (commerce) |
| Precision | ±5% over any 10s window is acceptable; billing is exact but separate | Unlocks local counting |
| Latency budget | Limiter adds ≤ 1ms p99 | Rules out synchronous remote calls at p99 |
| Failure | Limiter failure must not take down the API | Forces a fallback design |
🎯 Staff Move: "The number that drives everything is the precision tolerance. If product says exact, I'll put a store on the hot path and we'll talk about what that costs. If product accepts ±5% for a few seconds, I can make the limiter essentially free and immune to store outages. I'll assume ±5% and say who has to agree."
Phase 2: Core Entities & API (1–2 minutes)#
Policy { policy_id, version, scope: tenant|plan|global, match: {route_class, method},
rate_per_s, burst, concurrency_cap?, cost_fn?, mode: shadow|enforce,
fail_posture: open_local|closed_local, owner, expires_at? }
Override { tenant_id, policy_id, rate_per_s, burst, reason, approved_by, expires_at }
BucketState { key = "{tenant}:{route_class}", tokens, last_refill_ts } // or GCRA: { tat }
Lease { key, node_id, granted_tokens, valid_until }
Decision { allow|reject|shadow_reject, retry_after_ms, remaining, reset_s, reason }
UsageEvent { event_id, tenant_id, route_class, cost_units, ts } // for billing only
The limiter's external surface is the HTTP contract, not a "check" endpoint:
HTTP/1.1 429 Too Many Requests
Retry-After: 2
RateLimit-Policy: "tenant-standard";q=500;w=1
RateLimit: "tenant-standard";r=0;t=2
Content-Type: application/json
{ "error": { "type": "rate_limited", "code": "tenant_rate_exceeded",
"limit_scope": "tenant", "route_class": "write",
"retry_after_ms": 1840, "docs": "/docs/rate-limits" } }
"The response headers follow the IETF HTTP API working group's RateLimit header draft; many APIs use the older X-RateLimit-* family. What matters is that the shape is stable and documented, because SDKs will parse it."
Phase 3: High-Level Architecture (≤5 minutes)#
Staff candidates spend less than five minutes here. Draw three layers and one background loop.
"The request path is edge → gateway → service, and none of those hops waits on the quota store. The store is a coordinator the gateways talk to in the background. Billing never reads the limiter."
Phase 4: Transition to Depth (1 minute)#
The sentence that steers the interviewer toward your strongest ground:
"The box diagram is the easy part. The three decisions that make or break this design are: where the count lives and how much overshoot we accept, what the limiter does when it can't see its state, and how we change a limit without hurting our biggest customer. I'd like to go deep on the first two, then cover policy rollout. Does that match what you want to explore?"
Phase 5: Deep Dives (25–30 minutes)#
Deep dive A — where the count lives (10 min). Walk the spectrum: per-request remote check → local bucket with periodic sync → leased budgets. Give the overshoot formula. Then name the hot-tenant problem: the top tenant's key is on one shard, so use leases that let nodes count locally for seconds at a time, cutting that shard's load by 100–1,000×.
"Each node asks for a lease: 'give me 40 tokens for tenant 77 for the next 250ms.' The store decrements the shared bucket by 40 and hands them out. The node spends them locally. Unused tokens expire with the lease. Overshoot is bounded by the leases outstanding, which I control."
Deep dive B — blind-limiter posture (8 min). Present the posture table per intent, the fallback budget formula, and the circuit around the store: if sync calls exceed 5ms p99 or 1% errors for 10s, stop calling, enforce local fallback budgets, alert. "The limiter must degrade into a slightly less accurate limiter — never into an outage and never into no limiter."
Deep dive C — unit and fairness (5 min). Introduce cost units and concurrency caps. "A search request and a bulk export both count as one request under a req/s limit. One takes 8ms, the other 30 seconds. I'll charge cost units by route class and cap in-flight exports per tenant at 8."
Deep dive D — policy rollout (5 min). Shadow mode, per-tenant overrides with expiry, kill switch, and the "declared rejection rate" guard.
Phase 6: Wrap-Up (2–3 minutes)#
"To recap: three layers, three owners — abuse at the edge, contracts at the gateway, capacity at the service. The contract layer counts locally and syncs every 250ms, accepting up to 5% overshoot for a second in exchange for zero store latency on the request path. When the store is blind, contract limits fall back to per-node budgets, login fails closed, and shedding never needed the store. Changes ship through shadow mode. What I'd build next: global quotas across regions using allocated budgets, customer-visible quota dashboards, and adaptive concurrency limits replacing hand-tuned numbers on the service side."
Common Timing Mistakes#
| Mistake | Cost | Fix |
|---|---|---|
| 10 minutes comparing five algorithms | No time for state placement or failure — the actual signal | One table, one pick, move on in 60 seconds |
| Writing the Lua script line by line | Shows coding, hides judgment | Say "atomic read-refill-decrement in a server-side script" and move to what happens when the script is slow |
| Never stating the overshoot | Interviewer assumes you think local counting is exact | Say the formula before they ask |
| Treating login abuse and API quotas as one limiter | Inevitable contradiction when asked about fail-open | Separate the layers in the first 3 minutes |
| Leaving the client contract to the last minute | Misses retry amplification, a top failure cause | Mention Retry-After and SDK behavior when you draw the gateway |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
The rate limiter is the smallest system that contains every Staff-level concern at once. It is shared mutable state on the hottest path in the company, so it forces a latency-vs-correctness decision with a number attached. It fails in a way that cascades: a limiter that rejects too much is an outage; one that admits too much lets the next outage through. It says no to paying customers, so engineering cannot own its policy alone. And it is usually a platform, so the design has an organizational shape: who sets limits, who approves them, who is paged when they misfire.
A Senior candidate can build a correct limiter. The interview is checking whether you know that correctness is the least interesting property it has.
1.2 The L5 vs L6 Contrast — Visual#
The Senior path is sequential and reactive — each answer is prompted by an interviewer question. The Staff path is front-loaded: the tolerance and the intent decide the architecture, so the failure question has an answer before it is asked.
1.3 The Staff Question That Cuts Through Everything#
"When this limiter says no, who is on the other end — and when it can't decide, who gets the benefit of the doubt?"
Every major decision falls out of that sentence. If the other end is a paying customer, false rejections are expensive and you lean toward admitting. If it is an attacker, a false admission is the expensive mistake and you lean toward rejecting. If it is "whoever happens to arrive while the database is drowning", identity barely matters and you shed by priority. The algorithm never changes the answer; the intent always does.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Intent 1: Contract enforcement. The limit is a promise in two directions: the customer may send up to N, and the platform will serve N. Breaking it either way has a cost — false 429s become support tickets with the customer's account manager attached, and silent over-admission means the free tier is consuming paid-tier capacity. The keys are business identities (tenant, API key, plan) that the gateway resolves after authentication. Precision needs are moderate: customers care about "I paid for 500/s and I'm getting ~500/s", not about the 501st request in a particular second. Monthly quotas ("1M calls/month on the Starter plan") are the same intent at a different time scale and belong to metering, not to the real-time limiter.
Intent 2: Abuse defense. The caller is adversarial and adaptive. Credential stuffing at 10 attempts per account per hour, spread across 200K residential proxy IPs, never trips a per-IP limit of 100/min. The right keys are combinations: per target account (attempts against one victim), per IP prefix, per device fingerprint, per ASN, with reputation scores feeding thresholds. Counts are small and need to be close to exact — the 6th failed password attempt matters. And the posture flips: when the limiter is blind, the default is to restrict, because attackers notice.
Intent 3: Capacity protection. Nobody did anything wrong; demand just exceeds supply. The backend's limit is physical — 400 DB connections, 64 GPU slots, a vendor that allows 1,000 calls/s. Identity matters only for ordering: which work to drop first. The best tools here are concurrency limits and priority-based shedding, ideally adaptive, running inside each service with no shared state at all.
| Dimension | Contract | Abuse | Capacity |
|---|---|---|---|
| Key | Tenant × route class | Account, IP /24, device, ASN — combined | None (priority class only) |
| Precision | ±5% over 10s | Near-exact on small counts | Irrelevant — track saturation |
| Blind posture | Admit to local fallback budget | Restrict to strict local cap | No shared state to lose |
| Owner of the number | Product + platform | Security / trust & safety | Service owner + SRE |
| Feedback to caller | 429 + headers + docs | Often silent: CAPTCHA, delay, generic error | 503 + Retry-After, or priority degrade |
2.2 When NOT to Build a Distributed Rate Limiter#
- When your gateway or CDN already does it. Managed API gateways and edge providers ship per-key and per-IP limiting. If your limits are coarse (per API key, per minute) and your traffic is under ~50K req/s, configuration beats construction. See API Gateways.
- When the real problem is capacity. If the incident you are trying to prevent is "the database fell over", per-tenant limits are the wrong first tool. An adaptive concurrency limit inside the service and a circuit breaker around its dependencies will prevent more outages per engineer-week.
- When a queue is the right answer. If work can be deferred — webhooks, emails, exports — accept it into a queue and drain at the backend's rate. Rejecting deferrable work moves the retry problem onto your customers.
- When a per-node limit is good enough. If you have 8 nodes behind a round-robin load balancer, limit ÷ 8 per node is within ~10–20% of correct for steady traffic. Build the distributed part when that error costs something someone can name.
2.3 What the Interviewer Leaves Underspecified#
| Gap | Why It's Left Open | What to Say |
|---|---|---|
| Which intent | Tests whether you notice there are three | Name all three, commit to one, layer the rest |
| Precision | Tests whether you trade it deliberately | "±5% over 10s, and billing is separate — tell me if product needs exact" |
| Authenticated or not | Changes keys completely | "Contract limits post-auth; pre-auth traffic goes through edge abuse limits keyed on IP prefix and fingerprint" |
| Number of regions | Global limits are a different problem | "One region now; global quota later via allocation, not synchronous checks" |
| Client population | Retry behavior is half the system | "We publish SDKs; third-party clients may not honor Retry-After, so the limiter must survive retry storms" |
| Who sets limits | Org question disguised as config | "Product proposes, capacity approves, platform executes" |
2.4 Precise Terminology#
| Term | Meaning in this case study |
|---|---|
| Limit | The policy: rate R, burst B, scope, unit, posture |
| Quota | A limit over a long period (day/month), usually commercial and metered, not enforced in real time to the request |
| Overshoot | Requests admitted above the policy because of distributed counting; expressed as % over limit for up to T seconds |
| Lease | A slice of a shared budget granted to one node for a bounded time, spendable locally |
| Sync staleness | Time since a node last successfully reconciled with the store; the effective precision of the limiter |
| Blind mode | The limiter cannot reach shared state and is enforcing its fallback posture |
| Shadow mode | A policy computes decisions and logs would_reject without enforcing |
| Load shedding | Rejecting work to protect a service's capacity regardless of caller quota |
| Cost unit | A weight per request reflecting its resource cost; limits are expressed in units, not requests |
| Retry amplification | Additional load generated by clients retrying rejected or failed requests |
3. The Five Fault Lines#
3.1 Fault Line 1: Where the Count Lives#
Every distributed limiter answers one question: the limit is defined over the fleet, but each node sees a fraction of the traffic — where is the authoritative count?
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Synchronous shared counter (Redis script per request) | Exact; simple mental model; one place to look | +0.3–5ms per request; store is a hard dependency; hot tenant → hot shard; at 2M req/s needs ~20+ shards just for the limiter | Every customer pays latency; on-call pays for every store blip |
| Pure local (limit ÷ N per node) | Zero latency; no dependency; trivially available | Wrong when traffic is skewed across nodes (sticky connections, few long-lived HTTP/2 clients); wrong whenever N changes during autoscaling | Small tenants with 1–2 connections get 1/N of their limit — false 429s |
| Local + periodic sync (report consumption, receive global view) | Store off the request path; error bounded by sync interval | Overshoot during the sync interval; more code; needs staleness monitoring | Platform team owns complexity; backends absorb bounded overshoot |
| Leased budgets (store hands each node a slice of tokens) | Overshoot bounded by outstanding leases; hot keys hit the store once per lease, not per request | Leases stranded on idle nodes cause under-admission; needs lease sizing logic | Low-traffic tenants on many nodes may see slight under-admission |
The overshoot math. With N nodes, a sync interval S and a per-node lease L, a tenant can be admitted at most global_remaining + N × L in an interval where every node holds a fresh lease. With L sized from each node's own recent demand (e.g., ceil(1.2 × consumed_last_interval)), outstanding leases track the tenant's real traffic distribution, and the worst-case overshoot is roughly 20% of one interval's traffic — under 5% over a 1-second window for a tenant at steady state. Unleased nodes that see a first request use a tiny default (e.g., max(1, R × S ÷ N)) and ask for more at the next sync.
Staff default: leased budgets with a 250ms sync for contract limits; synchronous exact checks only for small, high-stakes counters (login attempts, OTP sends) where traffic per key is tiny and an extra millisecond is irrelevant.
When to deviate: a limit that protects a third-party vendor quota you are billed overage for (say $0.01 per call over 1,000/s) justifies a synchronous check — the overshoot has a price tag. A fleet with fewer than 10 nodes and smooth traffic can live with pure local limits.
🎯 Staff Move: "I'm deliberately trading exactness for locality. With 250ms sync and demand-sized leases, a tenant can run up to ~5% over contract for about a second. In exchange, the limiter adds no network latency and keeps working when the store doesn't. I'd want the API product owner to sign that tolerance."
3.2 Fault Line 2: Blind-Limiter Posture#
The limiter will lose its state — store failover, network partition, a bad deploy of the sync loop. What it does while blind is the most consequential decision in the design, and the most frequently improvised one.
| Posture | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Fail open (admit all) | API stays up; no false 429s | No protection exactly when something else may already be wrong; abuse windows | Backends (overload), security (attack window) |
| Fail closed (reject all) | No over-admission | Limiter incident becomes a total API outage | Every customer; the on-call; the status page |
| Fail to local budget | Bounded, imperfect enforcement continues | Under-admits skewed tenants; over-admits if N is miscounted | Some tenants see extra 429s; platform owns fallback math |
| Fail to last-known lease, then local | Smoothest transition; enforcement nearly unchanged for short blips | Most complex; lease expiry must be handled | Platform team |
Why ×1.5 in the fallback? Traffic is never perfectly balanced. A tenant with 4 long-lived connections might land on 4 of 40 nodes; a strict limit ÷ 40 would cut it to 10% of its contract. The 1.5× multiplier accepts that a perfectly spread tenant could get 150% of contract during blind mode, in exchange for not crushing a concentrated one. Blind mode should last minutes, so the platform accepts the overshoot. If it lasts longer, the posture table escalates.
Staff default: posture is a field on every policy (fail_posture: open_local | closed_local), not a global flag. Contract limits: open to local fallback. Credential and OTP endpoints: closed to a strict local cap. Capacity shedding: designed with no shared state.
When to deviate: if the protected backend is the thing that is failing — say the store outage is part of a broader incident and the database is already at 95% — the service-level shedder, not the contract limiter, is what saves you. That's the argument for keeping them separate.
🎯 Staff Move: "I want the blind-mode posture written next to each policy, with the name of the person who chose it. Security picks for login. The API product owner picks for contract limits. If nobody signs, the default is fail-to-local-budget, never fail-open-unbounded."
3.3 Fault Line 3: What Unit You Limit#
A limit is only as good as its unit. "Requests per second" quietly assumes all requests cost the same.
| Unit | Prevents | Misses | Who Pays |
|---|---|---|---|
| Requests/second | Floods of cheap requests; scripted loops | Expensive requests at low rate (exports, search with wildcards, LLM calls with 100K tokens) | Backend owners — their capacity is consumed by "low-rate" tenants |
| Cost units/second (weighted by route, payload, or measured cost) | Cost-heavy usage; aligns limits with pricing | Needs a cost model per route; measured cost is only known after the request | Platform maintains the cost table; product must explain it to customers |
| Concurrency (in-flight) | Slow requests holding threads, connections, GPU slots | Fast floods (a 2ms request at 10K req/s holds only 20 in flight) | Clients with legitimately slow work may be throttled sooner than expected |
| Bytes/second | Bulk upload/download abuse | Small expensive requests | Storage/egress owners |
The two-phase cost pattern. When cost is only known after execution (rows scanned, tokens generated), charge an estimate up front from the route class, then reconcile: after the response, debit or credit the bucket by actual − estimate. The bucket can go negative, which delays the tenant's next admission instead of interrupting work in progress. Public LLM APIs that limit on both requests per minute and tokens per minute use this shape.
Staff default: contract limits in requests/second with a per-route-class multiplier (reads 1, writes 2, search 5, export 50); concurrency caps per tenant on any route whose p99 exceeds ~1s; post-hoc cost reconciliation only where cost variance per request exceeds ~10×.
When to deviate: if you sell capacity in a non-request unit (tokens, compute-seconds, rows), the limit unit must match what you bill. Otherwise customers will discover the arbitrage between them before you do.
🎯 Staff Move: "A rate limit can't see a 30-second request. For the export route I'd add a tenant concurrency cap of 8 — that bounds the worker pool any single tenant can hold to 8 out of 200, regardless of rate."
3.4 Fault Line 4: Static Contract vs Adaptive Capacity#
Should the limit move when the backend is struggling?
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Static limits | Predictable; customers can plan; contract is honest | When backend capacity drops (failover to half the replicas), static limits admit more than the backend can take | Backend and every tenant during degraded periods |
| Adaptive limits (scale all tenants' limits by backend health) | Survives capacity loss | Customers see limits change without notice; hard to debug "why did I get 429 at 40% of my plan?" | Customers' trust; support |
| Static contracts + separate adaptive shedder | Contract stays honest when healthy; shedder protects in degradation, by priority | Two systems to reason about; shed requests look like 429s/503s to customers | Platform owns both; status page must explain shedding |
Staff default: keep contract limits static and honest; put adaptivity in the service-level shedder, which returns 503 with Retry-After (not 429) so customers and dashboards can distinguish "you exceeded your plan" from "we are degraded". Priority classes decide what is shed first: critical (checkout, auth) last, batch (exports, analytics) first.
When to deviate: for internal service-to-service limits where callers are your own teams, adaptive limits are fine — there is no contract to honor, only a shared resource to protect.
🎯 Staff Move: "429 means you exceeded what you're entitled to. 503 means we can't serve what you're entitled to right now. Mixing them hides our outages inside customers' quota graphs, and it makes our 429 rate useless as a signal."
3.5 Fault Line 5: Who Owns the Number#
Every limit is a number someone picked. The fault line is between speed of change and safety of change.
| Model | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Platform team owns all limits | Consistent; one review path | Bottleneck; platform lacks context on each product's capacity | Product teams wait days for changes; platform on-call fields every customer escalation |
| Each service team owns its limits | Fast; local context | Inconsistent semantics; no one sees global effect; unit mistakes | Customers integrating multiple APIs; the next incident's on-call |
| Product proposes, capacity approves, platform executes | Separation of concerns; safe rollout built in | Process overhead if not automated | Small, predictable cost in review latency |
Staff default: limits are data in a policy service, each with an owner, a declared expected rejection rate, an expiry for overrides, and a mandatory shadow period (≥ 24h for changes affecting the top 100 tenants). The platform's tooling makes the safe path the fast path: a limit change is a PR to a policy repo, the bot attaches yesterday's shadow results — "this change would have rejected 0.3% of tenant 77's traffic at peak" — and approvers act on data.
When to deviate: emergency blocks during an attack skip shadow — but they go through a separate, audited break-glass path with automatic expiry (e.g., 4 hours) so a 3 a.m. block doesn't become a permanent policy nobody remembers.
🎯 Staff Move: "The outage I worry about most isn't Redis failing — it's someone typing 50 instead of 5,000 into a policy for our largest customer. So every change goes through shadow with a declared expected rejection rate, and the tooling blocks enforcement if the shadow rate is 10× higher than declared."
4. Failure Modes & Operational Reality#
4.1 The Limiter Becomes the Outage — Store Latency on the Hot Path#
A synchronous limiter turns every blip in its store into an API-wide latency event.
t=0: Quota store primary for shard 7 starts a background save; fork + copy-on-write
on a 24 GB instance. Command latency p99: 0.8ms → 140ms.
t=+5s: Gateways call the store synchronously with a 200ms timeout. Request threads
block. Gateway p99: 45ms → 230ms for every route, every tenant (1/16 of keys
on shard 7, but threads are shared).
t=+20s: Gateway worker pools at 100%. Health checks start timing out.
t=+40s: Load balancer ejects 9 of 40 gateway nodes as unhealthy. Remaining 31 nodes
get 29% more traffic each. More blocking.
t=+90s: Page: gateway.5xx_rate > 2%. On-call sees "Redis latency" — 6 minutes to
connect the gateway outage to a limiter shard.
t=+8min: Save completes; latency recovers. Ejected nodes take 3 min to rejoin.
Total: 11 minutes of elevated errors across the whole API because of one shard's fork.
Detection: limiter.decision_latency_p99 (should be < 1ms), limiter.store_call_latency_p99 by shard, gateway.worker_pool_utilization, limiter.blind_mode gauge.
Mitigation: move the store off the request path (leased budgets); if a synchronous check must exist, give it a tight timeout (≤ 5ms) plus a circuit breaker that trips to local fallback after 1% timeouts over 10s, and run it on a bounded, separate executor so it can't consume request threads.
Prevention: disable persistence on limiter-only stores (limiter state is regenerable in seconds), or run persistence on replicas only; load-test the gateway with injected store latency (100ms, 1s, black hole) before launch and after every limiter release.
Owner: API platform team (limiter library and gateway); store operations team for the shard.
4.2 The NAT Cliff — Legitimate Users Behind One Address#
t=0: Security adds a pre-auth limit: 300 req/min per source IP on /api/session.
Shadow mode skipped — "it's just a login limit".
t=+2h: University campus exam day; 6,000 students reach the app through 4 egress IPs.
Every IP exceeds 300/min within 10 seconds of the hour.
t=+2h05: 9 of 10 students get 429 on login. The app shows "Something went wrong".
t=+2h40: Social media posts; the support queue grows 40×.
t=+3h10: On-call rolls back. Mobile carrier CGNAT users in two countries were also
affected — 0.8% of all logins failed for 70 minutes.
Detection: limiter.rejected by key cardinality — a single key rejecting thousands of distinct downstream identities (user-agents, device IDs, cookies) is a shared-address signature; auth.login_success_rate by ASN.
Mitigation: move to composite keys — IP alone only for very high thresholds; per-account attempt limits; per-device-fingerprint limits; challenge (CAPTCHA, proof-of-work) instead of hard reject for ambiguous traffic.
Prevention: shadow mode is mandatory for abuse rules too; maintain an allowlist/higher-tier list of known shared egress ranges (campuses, enterprise proxies, carrier CGNAT ranges) with a reviewed process.
Owner: trust & safety owns the rule; API platform owns the shadow-mode tooling that would have caught it.
4.3 The Retry Storm at the Window Boundary#
t=0: A popular open-source client library implements "on 429, sleep until
X-RateLimit-Reset, then retry". The API uses fixed 60s windows aligned to :00.
t=+0:59 Thousands of throttled clients compute the same reset time: hh:mm:00.
t=+1:00 Gateway traffic jumps from 40K to 310K req/s in under 200ms.
Every client's first request after reset is admitted — window counters are empty.
t=+1:00.4 Backend p99 → 4s. Some admitted requests time out, clients retry on 5xx too.
t=+1:02 Load shedder sheds 35% for 15s. Recovers. Repeats every minute for 2 hours.
Detection: gateway.requests_per_second at 100ms resolution shows a sawtooth with period = window length; limiter.admitted_in_first_second_of_window / average.
Mitigation: replace fixed windows with token buckets or GCRA (no shared reset instant); return Retry-After with a jittered value per client (e.g., base + random 0–20%); per-tenant window offsets (hash tenant ID into the window phase) if fixed windows must stay.
Prevention: own the SDKs and the docs: publish the backoff algorithm (exponential with full jitter, capped at 3 retries for writes), and test SDKs against a simulated 429 storm.
Owner: API platform (algorithm and headers); developer experience team (SDKs and docs).
4.4 Silent Over-Admission — Nine Days of No Limits#
Day 0: Refactor renames the route class "search" to "search_v2" in the gateway.
The policy service still holds limits for "search". No policy matches →
default "no limit" path. No errors, no alerts — admitted requests look normal.
Day 3: Search cluster CPU 55% → 70%. Attributed to "growth".
Day 6: One scraper tenant at 8× its contract. Search p99 creeps 120ms → 260ms.
Day 9: Capacity review notices one tenant is 31% of search load. Investigation finds
the unmatched route class. Search cluster was scaled up 40% in the meantime:
~$38K of unnecessary spend over the month.
Detection: limiter.unmatched_requests (requests that matched no policy) — should be ~0 for authenticated traffic and must alert; limiter.policy_match_rate per route class; a daily diff of route classes seen by the gateway vs route classes with policies.
Mitigation: default policy for unmatched authenticated traffic is a conservative plan-level limit, never "unlimited".
Prevention: route classes are a shared enum validated at deploy time against the policy service; CI fails if a route class has no policy.
Owner: API platform. This is the failure the platform's own dashboards exist to catch — an absence of rejections is not evidence of health.
4.5 The Unit Mistake — Throttling the Biggest Customer#
t=0: Sales negotiates a burst increase for tenant 12 (largest customer, 9% of
revenue): "burst 5,000 per minute". Engineer enters burst=5000, rate=5000 in a
policy whose unit is per-second... then "corrects" rate to 83/s (5000 / 60).
Their existing sustained rate was 1,200/s.
t=+30s: Policy propagates. Tenant 12's admitted rate drops from 1,150/s to 83/s.
t=+2min: Tenant 12's order sync falls behind; their on-call pages ours via account team.
t=+14min: Override applied; root cause understood at t=+40min.
Detection: limiter.rejected by tenant with anomaly detection against the tenant's 7-day baseline; any top-50 tenant crossing 1% rejected pages the platform on-call.
Prevention: shadow is mandatory for top-tier tenants; policies carry an explicit unit field and the UI shows the result in three units (per second / minute / hour) before submit; the tooling refuses a change that lowers a tenant's effective limit by more than 50% without a second approver.
Owner: API platform for tooling; account management owns the customer conversation.
4.6 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Store latency on hot path | limiter.decision_latency_p99 > 1ms | Entire API if synchronous | Leased budgets; 5ms timeout + breaker to local fallback | API platform |
| Store unreachable | limiter.blind_mode = 1, limiter.sync_staleness_ms > 2000 | Precision degrades fleet-wide | Per-policy blind posture | API platform; security for credential endpoints |
| Shared-address false positives | Rejected key fan-out to many device IDs | Campus, enterprise, carrier users | Composite keys; challenge instead of reject | Trust & safety |
| Retry storm | 100ms-resolution sawtooth; client.retry_ratio > 0.3 | All tenants during peaks | GCRA/token bucket; jittered Retry-After; SDK backoff | API platform + DX |
| Unmatched policy | limiter.unmatched_requests > 0 | Backend cost and capacity | Default conservative policy; CI check | API platform |
| Unit/typo in policy | Tenant rejection anomaly vs baseline | One (often large) customer | Shadow; second approver; 3-unit preview | API platform + account team |
| Hot tenant overloads store shard | store.shard_ops skew > 5× median | Every tenant on that shard | Leases (per-node aggregation); split hot key into sub-buckets | API platform |
| Clock skew between nodes | node.clock_offset_ms > 50 | Refill errors on local buckets | Use monotonic clocks locally; store-side time for shared state | Infrastructure |
| Policy push stalls | policy.version_lag by node | Mixed enforcement across fleet | Nodes keep last-good policy; alert at 2 min lag | API platform |
🎯 Staff Insight: The limiter's most dangerous failures are quiet. A store outage pages someone; a policy that matches nothing, or a typo that throttles one customer, does not — until revenue or an account manager notices. "I'd alert on the absence of expected rejections as hard as on the presence of unexpected ones."
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | "Limit each user to N req/s" | Separates contract, abuse and capacity intents; commits; states precision tolerance | Asks how many limiters already exist across the org and whether limits are part of the product's pricing |
| State & latency | Synchronous Redis script, sharded | Leased local budgets with bounded overshoot formula; store off hot path | Sets the org rule for when synchronous limiting is allowed and who pays for the store |
| Failure | "Fail open" when asked | Per-intent blind posture table; local fallback math; breaker on store calls | Signed posture table with product/security owners; game days; posture reviewed after every incident |
| Client contract | Returns 429 | 429 vs 503 distinction; Retry-After; headers; SDK backoff with jitter | Quota as a product: dashboards, upgrade flows, 429 rate reviewed as a revenue/experience metric |
| Policy lifecycle | Edit config, deploy | Policy service; shadow → canary → enforce; overrides with expiry; kill switch | Governance: proposer/approver/executor split; declared rejection rates; audit trail |
| Operations | "Monitor 429s" | limiter.unmatched_requests, staleness, per-tenant anomaly alerts, owners per layer | Org-level error budget for false rejections; quarterly review of top-50 tenants' limits with account teams |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| Separates intents up front | "Login abuse and API quotas want opposite failure postures, so they're separate layers." |
| Quantifies overshoot | "40 nodes, demand-sized leases, 250ms sync — under 5% over contract for about a second." |
| Keeps the store off the hot path | "The store is a coordinator the nodes talk to in the background, not a dependency of the request." |
| Distinguishes rate from occupancy | "A rate limit can't see a 30-second export. I'd cap concurrency per tenant on that route." |
| Treats 429 as a contract | "Retry-After, remaining quota, a stable error code — and our SDK honors them with jitter." |
| Owns policy safety | "Every limit change runs in shadow first, and the tool shows which customers it would have hit." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| Spends 15 minutes on algorithm comparison | The algorithm is the least differentiated decision in the design |
| "Redis is highly available, so the limiter can't fail" | Replication doesn't help latency spikes, failovers or network partitions |
| Same failure posture for every limit | Shows no model of who is on the other end of a rejection |
| Uses the limiter's counters for billing | Conflates an approximate control mechanism with a financial record |
| Keys everything by IP | Breaks behind NAT and does nothing against rotating proxies |
| No answer for "who changes the limit" | Misses the most common real-world limiter incident: a bad policy |
5.4 Common False Positives#
- Writing a correct Lua token-bucket script ≠ designing a limiter. It shows Redis fluency. The question is whether that script should be on the request path at all.
- Naming five algorithms ≠ judgment. Listing fixed window, sliding log, sliding counter, token bucket and leaky bucket in a table is useful for 60 seconds, then becomes a signal of avoiding the hard parts.
- "Consistent hashing to shard counters" ≠ hot-key handling. Consistent hashing spreads keys, not load; one tenant at 40% of traffic is still one key on one shard.
- Mentioning "adaptive rate limiting" ≠ understanding it. Unless the candidate separates adaptive capacity shedding from static contract limits, it usually means "limits change when things are bad", which breaks the customer contract.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Framing | 0–3 min | Three intents; commit to contract limits; precision tolerance; latency budget |
| Entities & API | 3–6 min | Policy, bucket state, lease, decision; 429 response contract |
| Architecture | 6–10 min | Edge / gateway / service layers; sync loop; policy push; metering |
| State placement | 10–20 min | Sync vs local vs leases; overshoot formula; hot tenants |
| Blind posture | 20–27 min | Per-intent table; fallback math; breaker on store |
| Units & fairness | 27–32 min | Cost units; concurrency caps; 429 vs 503 |
| Pivot (interviewer's choice) | 32–42 min | Multi-region, login abuse, policy rollout, billing quotas |
| Wrap | 42–45 min | Layers and owners; what you'd build next |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Shape |
|---|---|---|
| "Make the limit global across 3 regions" | Do you avoid synchronous cross-region checks? | Allocate per-region budgets from demand; rebalance every few seconds; accept regional overshoot |
| "Now protect the login endpoint" | Posture flip; composite keys | Per-account + per-IP-prefix + device; fail closed; challenge instead of reject |
| "A customer complains about false 429s" | Operational forensics | Decision logs with reason codes, per-node view, staleness at decision time |
| "We need exact limits for billing" | Separating control from accounting | Metering stream with dedup for billing; limiter stays approximate |
| "Traffic is 10× tomorrow" | Where does the design break first? | Hot-tenant shards and sync fan-in, not total req/s; leases scale with node count |
| "How do you test this?" | Correctness and failure culture | Simulated fleets with skewed traffic; injected store latency; shadow comparisons |
6.3 What to Deliberately Skip#
- Leaky bucket as a separate algorithm — say it's a token bucket viewed as a queue, and move on.
- Lua script internals — one sentence: "atomic refill-and-decrement on the server".
- DDoS at the network layer — volumetric attacks belong to the edge provider; say so and draw the edge box.
- Exact sliding-window logs — only for tiny-count abuse limits; not worth the memory for contracts.
- Building your own consensus — the policy service can sit on etcd or a database; it doesn't need novel coordination.
6.4 Follow-Up Questions to Expect#
- "How far over its limit can one tenant get with your design, and for how long?"
- "The quota store is down for 10 minutes. Walk me through what each kind of limit does."
- "One tenant is 40% of your traffic. What happens to the store shard holding its key?"
- "How does a client know when to retry, and what stops all clients retrying at once?"
- "Product wants to lower the free tier from 100/s to 20/s. How do you ship that?"
- "How do you limit expensive requests differently from cheap ones?"
- "How would you enforce one quota across three regions?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design a rate limiter for our public API."
Staff Answer
"Three questions first, because they lead to different designs. Is this about enforcing customer plans, stopping abuse, or protecting our backends from overload? I'll assume plan enforcement for a multi-tenant API — around 50K tenants, 500K req/s average, 2M peak across ~40–200 gateway nodes — with abuse defense at the edge and load shedding in services as separate layers I'll come back to.
The constraint that shapes everything: how exact must it be? I'll propose ±5% over any 10-second window, added latency under 1ms p99, and the API must stay up if the limiter's store fails. That lets me count locally on each node, sync leased budgets with a shared store every 250ms, and keep the store off the request path. I'll cover: the policy model and 429 contract, state placement and overshoot, blind-mode posture, cost units for expensive routes, and how limits change safely."
Why this is L6:
- Names three intents with different designs and commits to one with numbers
- Converts "rate limiter" into a precision tolerance and a latency budget before any box
- The outline ends with policy change — the operational reality, not the algorithm
What L7 adds:
- Asks whether other teams already run limiters and whether this should become the shared platform
- Asks whether limits are part of pricing — if so, product and finance co-own the numbers
- Frames the false-429 rate on enterprise tenants as a revenue-protection metric
❌ Common L5 Trap
"I'll compare token bucket, leaky bucket, fixed window and sliding window, pick sliding window for accuracy, and store counters in Redis with a Lua script, sharded by user ID."
Why this misses: It's correct and it answers the wrong question. Nothing says what happens at 2M req/s, when Redis is slow, or when the user is a company with 4,000 employees. The next 30 minutes become the interviewer dragging those out one at a time.
Drill 2: Bound the Overshoot#
Prompt: "You count locally and sync. Tenant limit is 1,000 req/s, burst 1,000. 50 nodes, sync every 500ms. How far over can the tenant get?"
Staff Answer
"Depends on what each node may spend between syncs. If every node holds the full remaining budget — the naive version — the worst case is 50 × 1,000 = 50,000 requests in one interval: 50× over. That's why I don't do that. With leases, each node may spend only what the store granted it. If I size leases from the node's own recent demand — say 1.2 × what it used last interval — then outstanding leases sum to about 1.2 × the tenant's per-interval traffic, and worst-case overshoot is ~20% of one interval's worth: at 1,000/s and 500ms, about 100 extra requests per interval, so roughly 10% over for that second. Shorten the sync to 250ms and it's ~5%. A node that sees a tenant for the first time gets a floor lease, max(1, R × S ÷ N) = 10 tokens, so cold nodes can't stampede.
The tradeoff runs the other way too: leases stranded on nodes the tenant stopped hitting cause under-admission until they expire — at most one interval. I'd monitor limiter.lease_utilization; below 50% means leases are oversized."
Why this is L6:
- Shows the naive local design's failure with a number (50×), then fixes it
- Gives the formula and computes it for two sync intervals
- Names the opposite error (stranded leases) and how it's observed
What L7 adds:
- Turns the tolerance into a product SLA ("within 10% over any 10-second window") published in the API docs
- Makes lease sizing a platform-owned library default so teams don't each invent one
Drill 3: Make It Concrete — Size the Store#
Prompt: "Size the quota store for 2M req/s peak."
Staff Answer
"With per-request checks: 2M script executions per second. At ~50–80K small-script ops/s per shard, that's 25–40 primary shards plus replicas, and every request pays a round trip. With leases at 250ms: each node syncs once per interval per active key it holds. 100 nodes × 4 syncs/s × active keys per node. If a node sees ~500 active tenants per second, that's 200K key-syncs per second — but I batch: one pipelined call per node per interval carrying all its keys, so 400 round trips per second fleet-wide, carrying ~200K key operations. Keys themselves: 50K tenants × ~5 route classes × ~150 bytes ≈ 40 MB. So: a 3-shard cluster with replicas, mostly idle. The sizing driver isn't volume, it's the hot tenant: its key gets one update per node per interval — 400 updates/s — trivial with leases, catastrophic without them."
Why this is L6:
- Compares sync-per-request vs lease designs with concrete shard counts
- Notices batching changes round trips by 500×
- Identifies the hot key as the real sizing driver
What L7 adds:
- Notes that a mostly idle 3-shard cluster can be shared with other platform state — or that it shouldn't be, because sharing couples blast radius
- Prices it: ~$1–2K/month for the store vs ~$25–40K/month for the per-request version, before counting gateway CPU spent waiting
Drill 4: The Store Is Down#
Prompt: "Your quota store just became unreachable from all gateways. What happens in the next 10 minutes?"
Staff Answer
"Nothing on the request path waits, because the store wasn't on it. Within one sync interval, nodes notice failed syncs; after 3 consecutive failures — under a second — they enter blind mode. Each node spends its remaining lease, then enforces its fallback per policy. Contract policies: limit ÷ N × 1.5 per node, where N comes from the membership list, not a hard-coded constant. Credential-endpoint policies: strict local caps — 3 failed attempts per account per node per 10 minutes — which, across 40 nodes, is worse precision but still bounded. Service-level shedders: unaffected.
limiter.blind_mode goes to 1 on every node: that's a ticket immediately and a page if it lasts 15 minutes, because by then over-admission across 1.5× fallbacks starts to matter for backend capacity. When the store returns, nodes resume syncing; buckets start fresh — I don't try to replay consumption from blind mode, because a few minutes of approximate counting doesn't justify the complexity."
Why this is L6:
- Answers per intent, with the specific fallback numbers
- Detection is fast and automatic; escalation is time-based
- Deliberately skips reconciliation of blind-mode usage and says why
What L7 adds:
- Game-days blind mode quarterly with the store black-holed in one AZ, and tracks how long the fleet actually stays in it
- Reviews whether the 1.5× multiplier still matches traffic skew data each quarter
Drill 5: The Hot Tenant#
Prompt: "One tenant generates 40% of all API traffic — 800K req/s at peak. What breaks?"
Staff Answer
"With per-request checks, their key lives on one shard and that shard takes 800K ops/s — it falls over and takes every other tenant on that shard with it. With leases, the store sees ~400 lease updates per second for that key, which is fine. So the store isn't the problem; three other things are. First, contract math: at 800K req/s the tenant almost certainly has a custom contract, and I'd give them their own policy with their own burst. Second, their traffic dominates node-level resources — their connections may concentrate on a subset of nodes, so I watch per-node skew. Third, the backend: 40% of load from one tenant means their bad deploy is our incident. I'd put them on a dedicated capacity pool or cell, with their own concurrency caps, so their retry storm can't consume the shared pool."
Why this is L6:
- Shows how the lease design dissolves the store hot-key problem
- Moves the conversation to where the risk actually sits — shared backend capacity
- Proposes isolation (cells/pools), not just higher limits
What L7 adds:
- Asks whether the commercial contract prices that isolation — dedicated capacity is a product, not a favor
- Plans for the tenant's growth: 40% today may be 60% in 18 months; cell architecture is the one-way door to decide now
Drill 6: Noisy Neighbors on Expensive Routes#
Prompt: "All tenants are within their req/s limits, but the reporting service is saturated and everyone's reports are slow."
Staff Answer
"Request rate is the wrong unit for that route. A report request at 2/s that takes 45 seconds holds 90 workers; three tenants doing that hold the whole 256-worker pool. I'd add two things. A per-tenant concurrency cap on the report route — 8 in flight on the standard plan, 32 on enterprise — implemented as a semaphore with a lease timeout of 2× the route's p99 so crashed requests release their slot. And fair queuing in front of the worker pool: a short per-tenant queue with round-robin dequeue, so one tenant's backlog doesn't delay everyone else's single request. Requests that wait more than 5 seconds in the queue get 503 with Retry-After. Longer term, reports become async jobs with a status URL — the right shape for 45-second work."
Why this is L6:
- Diagnoses the unit mismatch — rate vs occupancy — with arithmetic
- Concurrency cap with lease timeouts handles crashed holders
- Moves long work to an async contract, not just a bigger pool
What L7 adds:
- Aligns plan tiers with concurrency entitlements so the limit is something sales can sell
- Sets a platform rule: any synchronous route with p99 > 5s must have a concurrency cap and an async alternative
Drill 7: Build vs Buy#
Prompt: "Why not just use our cloud API gateway's built-in rate limiting?"
Staff Answer
"Often we should. Managed gateways handle per-key throttling with burst and rate, usage plans and quotas, at no engineering cost. If our needs are per-API-key limits, a handful of plans and under a few hundred thousand req/s, buy it. The triggers to build are specific: limits that depend on our own data the gateway can't see — tenant hierarchies, per-route-class cost units, concurrency caps; a blind-mode posture we need to choose per policy; shadow-mode rollout with per-tenant impact reports; or cost — managed per-request pricing at billions of requests a month can exceed a small platform team. Even then, I'd build on Envoy's local and global rate-limit hooks rather than writing a proxy."
Why this is L6:
- Defaults to buying and names concrete triggers
- Identifies capabilities, not vague "flexibility", as the reason to build
- Suggests a middle path (extend an open-source proxy)
What L7 adds:
- Computes break-even: managed per-request fees vs a 2–3 engineer team (~$600K/year fully loaded)
- Weighs lock-in: the policy model is the asset; keep it portable across gateways
Drill 8: Changing Limits Without an Outage#
Prompt: "Product wants to cut the free tier from 100 req/s to 20 req/s next week."
Staff Answer
"First, how many free-tier tenants exceed 20/s today, and who are they? The policy goes into shadow immediately: we compute both verdicts and log would_reject per tenant. After a few days of shadow data I can tell product: '1,840 free tenants exceed 20/s at peak; 37 of them are in an active sales conversation; 6 are integration partners.' Product decides on exceptions. Then: email affected tenants with their actual peak usage, two weeks' notice; headers advertise the new limit during the notice period. Enforcement rolls out on 1% of nodes, then 10%, then 100% over a day, with a kill switch that reverts in 30 seconds. Partners get time-boxed overrides that expire."
Why this is L6:
- Uses shadow data to turn a policy change into a list of named affected customers
- Communication is part of the rollout, with headers carrying the advance notice
- Canary and kill switch make enforcement reversible
What L7 adds:
- Recognizes this is a pricing change with revenue and churn impact — product owns the decision; finance should see the conversion forecast
- Establishes that tier changes follow a standard notice period, written into the API terms
Drill 9: The Cost of the Limiter#
Prompt: "Finance says the gateway fleet is too expensive. How much of that is the limiter?"
Staff Answer
"Measure, don't guess: profile gateway CPU per request with and without the limiter in a canary. With local buckets, the limiter is typically a small single-digit percentage of gateway CPU — hashing a key, a map lookup and some arithmetic. If we were doing synchronous store calls, the cost is mostly waiting: threads parked on the network, which shows up as more nodes needed for the same throughput — easily 10–20% of fleet size at high request rates. The store itself is small. The more interesting cost is what the limiter saves: rejected requests that never touched a backend. If abusive and over-plan traffic is 3% of requests, rejecting at the gateway in microseconds avoids 3% of backend cost, which usually dwarfs the limiter's own cost."
Why this is L6:
- Proposes measurement with a canary
- Separates CPU cost from waiting cost of synchronous designs
- Frames the limiter's net cost including avoided backend work
What L7 adds:
- Builds a per-tenant cost-to-serve model that the limiter feeds — rejected volume, admitted cost units — for pricing decisions
- Connects limiter data to capacity planning so fleet sizing uses admitted, not offered, load
Drill 10: Multi-Region Global Quota#
Prompt: "A tenant's limit of 10,000 req/s must hold globally across three regions."
Staff Answer
"I won't check a global counter synchronously — that's 60–150ms per request. Each region gets an allocation of the global limit, sized from its recent demand share, and a regional coordinator rebalances every 2–5 seconds through a lightweight global aggregator. If the tenant is 60/30/10 across US/EU/APAC, they get roughly 6,000/3,000/1,000 plus a 10% shared reserve any region can claim on the next rebalance. Inside a region, the same leased-budget design runs. Overshoot is bounded by the reserve plus one rebalance period of demand shift — I'd tell the customer 'within 10% globally, averaged over 10 seconds'. If the aggregator or a cross-region link fails, each region keeps its last allocation — the sum still equals the global limit, so a partition can't create over-admission beyond the reserve."
Why this is L6:
- Refuses synchronous global state with the latency number
- Allocation-from-demand plus a reserve handles shifting traffic
- Partition behavior is designed: last allocations sum to the limit
What L7 adds:
- Questions whether global limits are needed at all — most contracts can be per-region, which removes the whole coordination layer
- Treats data-residency rules as a constraint on what the aggregator may see (counts are fine; request contents are not)
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Incident — The Launch-Day Lockout#
Context: A consumer product launch drives 12× normal traffic. Within 20 minutes, 18% of authenticated API calls receive 429 — including from enterprise tenants who are nowhere near their limits. The on-call has already doubled the gateway fleet, which made it worse. You are pulled in.
Questions to Surface First:
- Are the 429s coming from the contract layer, the edge, or the service shedder? What reason codes do the decision logs show?
- Did the fleet scaling change N, and does the fallback or lease math depend on N?
- What is
limiter.sync_staleness_msdoing — are nodes in blind mode? - Are rejected tenants actually over their limits according to the store's view?
Typical L5 Approach: Raises all limits by 3× through a config push to stop the bleeding, then investigates. The 429s drop; the backend, now unprotected during a 12× peak, starts timing out ten minutes later.
Staff Approach: Reads reason codes before touching limits. Finds that new nodes start with no leases and enter blind mode immediately because the store's connection limit (10K clients) was exhausted by the doubled fleet's sync connections; in blind mode their fallback divides by N, which doubled — so every tenant's per-node allowance halved exactly when traffic peaked. Fix: raise the store connection limit, cap sync connections per node to 1 pooled connection, temporarily pin N in fallback math to the pre-scale count.
Principal Approach: Treats "scaling the gateway made rate limiting stricter" as a design flaw in the platform contract: limiter correctness must not depend on fleet size. Commissions a change so fallback budgets derive from each node's observed share of traffic rather than 1/N, and adds limiter behavior under 2× scale-out to the launch-readiness checklist every product launch must pass.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Group rejections by reason code: 81% blind_fallback_exceeded, 12% contract_exceeded, 7% shed. Confirms limiter self-inflicted, not real over-use. |
| Triage | limiter.blind_mode = 1 on all 80 new nodes; store shows rejected_connections. New nodes never got leases. |
| Quick fix | Raise store maxclients; sync uses one pipelined connection per node; set fallback_n_override=40 via policy flag (no deploy). 429 rate drops to 0.4% in 3 minutes. |
| Guardrails | Alert on limiter.blind_mode fraction > 5% of nodes; gateway autoscaling checks store connection headroom before adding nodes. |
| Post-mortem | Why did blind fallback scale with N? Why did adding capacity reduce admitted throughput? Why didn't load tests include scale-out? |
Metrics to Watch: limiter.rejected by reason, limiter.blind_mode fraction of fleet, store.connected_clients vs maxclients, limiter.lease_grant_latency_p99.
Organizational Follow-up: launch-readiness review includes the limiter; product launches notify the API platform 2 weeks ahead with expected traffic multipliers so leases and store capacity are pre-sized.
Ownership Question: "Who decides to raise limits during an incident?" Staff answer: The platform on-call may apply temporary, auto-expiring overrides within a pre-agreed envelope (up to 2× for 4 hours). Anything beyond that needs the API product owner, because it trades backend safety for customer experience.
Key Takeaway: "Read reason codes before changing numbers. Most launch-day 429 storms are the limiter's own failure mode, not customers exceeding plans."
What clears the Staff bar:
- Diagnoses from decision reason codes rather than raising limits blindly
- Finds the N-dependence in fallback math and the store connection ceiling
- Uses a policy flag, not a deploy, to fix it in minutes
Deep Dive 2: Silent Failure — The Limiter That Stopped Limiting#
Context: A cost review shows the search cluster grew from 60 to 96 nodes over six weeks. The search team blames organic growth. The data team notices that three tenants' search traffic grew 9× in the same period, but none of them shows a single 429 since a gateway release six weeks ago.
Questions to Surface First:
- What does
limiter.unmatched_requestsshow since that release? - Did the release change route classification, tenant resolution, or policy matching?
- Is the policy service reachable, and what version do nodes report?
- What was the rejected volume for search before the release?
Typical L5 Approach: Adds explicit limits for the three tenants and rolls them out. The underlying gap remains — any future rename or new route will silently bypass policy again.
Staff Approach: Finds that the release changed tenant resolution for API keys created after a migration: those keys resolve to
tenant_id = null, and null tenants fall through to "no limit". 14% of authenticated traffic is unmatched. Fixes resolution, adds a conservative default policy for unmatched authenticated traffic, and makesunmatched_requests > 0.1%a paging alert.
Principal Approach: Recognizes a class of failure — fail-open by omission — that no runtime alert fully prevents. Requires a contract test in CI: a representative sample of production keys and routes replayed against the policy engine on every gateway build, failing the build if match rates drop. Assigns limiter effectiveness a weekly review alongside cost, owned by the platform, with the search team as a consumer.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Query decision logs: rejection rate by route since release vs before. Search: 0.9% → 0.0%. |
| Triage | unmatched_requests = 14% of authenticated traffic since release; all unmatched have tenant_id = null. |
| Quick fix | Hotfix tenant resolution; deploy default policy unmatched_authenticated at the Starter plan limit. |
| Guardrails | Page on unmatched_requests > 0.1% for 10 minutes; daily report: route classes and tenants with zero policy matches. |
| Post-mortem | Why was "no match" equivalent to "no limit"? Why was zero rejections not anomalous? Who owns limiter effectiveness? |
Metrics to Watch: limiter.unmatched_requests, limiter.rejected per route class vs 28-day baseline, per-tenant admitted cost units, search cluster cost per tenant.
Organizational Follow-up: the search team gets a per-tenant usage dashboard; the platform adds "expected rejection rate" to each policy so absence of rejections becomes detectable.
Ownership Question: "Who should have noticed — search or the API platform?" Staff answer: The API platform. Search consumes protection as a service; the platform owns proving it's working. Search owns noticing unexplained growth, which they did, slowly.
Key Takeaway: "A limiter that rejects nothing looks exactly like a healthy one. Monitor the match rate, not just the reject rate."
What clears the Staff bar:
- Checks policy match coverage, not just per-tenant limits
- Makes "no match" resolve to a conservative default, never unlimited
- Puts effectiveness ownership on the platform with a concrete alert
Deep Dive 3: Large-Customer Onboarding — "We Need 50,000 Requests per Second, Guaranteed"#
Context: Sales is closing a contract with a logistics company that will send 50K req/s sustained, 120K req/s bursts during end-of-day batch windows, and wants a contractual guarantee of no throttling within those limits. Your current largest tenant is 8K req/s. The deal is 6% of next year's revenue.
Questions to Surface First:
- What does "guaranteed" mean in the contract — an SLA with credits, or just a limit?
- Which routes, and what's the cost profile? Are the bursts reads or writes?
- Can the 120K bursts be scheduled or smoothed — e.g., a bulk endpoint?
- What backend capacity exists today, and what does 120K req/s of their mix consume?
Typical L5 Approach: Creates a tenant policy at 50K/s with 120K burst. The limiter now admits the traffic; the backend, sized for ~40K req/s total, can't serve it. The limit is honored and the service falls over.
Staff Approach: Treats it as a capacity question first. A limit is only a guarantee if capacity backs it. Proposes: a dedicated cell (gateway pool + service replicas + DB read replicas) sized for 120K req/s of their mix; a bulk endpoint that accepts 1,000 items per request, cutting end-of-day bursts to ~120 req/s; a contract limit of 50K req/s with 120K burst only on the bulk endpoint window; shadow-mode run during their pilot with weekly reports.
Principal Approach: Turns the deal into a product decision: "dedicated capacity" becomes a priced tier with a standard architecture (cells), not a one-off. Brings finance in on the cost: the cell costs ~$90K/month; the contract must price it. Writes the playbook so the next large customer takes 3 weeks, not 3 months.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (week 1) | Get the request mix from their pilot; compute backend cost per request class; identify which bursts are batch-able. |
| Triage | 80% of burst volume is per-item status updates → bulk endpoint candidate. Sustained 50K is mostly reads → cacheable. |
| Quick fix | Pilot at 5K req/s with the tenant's policy in shadow above that; collect real skew across nodes and regions. |
| Guardrails | Tenant concurrency cap sized to the cell's worker pool; cell-level shedder with the tenant's own priority classes so their batch can't starve their interactive traffic. |
| Post-mortem / review | 30-day review: actual vs projected; contract SLA measured from gateway-admitted latency, not client-observed. |
Metrics to Watch: tenant's admitted req/s and cost units, cell saturation, limiter.rejected{tenant} (must be ~0 within contract), bulk endpoint adoption rate.
Organizational Follow-up: account team, platform, and service owners agree on a joint escalation path; the customer gets a named technical contact and a shared dashboard.
Ownership Question: "Who signs the guarantee?" Staff answer: The API product owner signs the commercial term; the platform signs that the limiter will admit it; the service owners sign that capacity backs it. If any of the three can't sign, the contract number changes.
Key Takeaway: "A rate limit is a promise only if capacity backs it. For large customers, sell isolation, not a bigger number."
What clears the Staff bar:
- Reframes the limit request as a capacity and isolation question
- Uses API shape (bulk endpoint) to change the traffic, not just the limit
- Assigns three distinct sign-offs for the guarantee
Deep Dive 4: Post-Mortem — The 429s That Caused an Outage#
Context: Last Tuesday, a 4-minute partial outage of the orders service turned into a 47-minute full API brownout. The orders service started shedding (503) correctly. Mobile clients, receiving 503, retried immediately up to 5 times. The gateway rate limiter then started returning 429 to those same clients; the mobile app treated 429 as a generic failure and also retried. You are running the post-mortem.
Questions to Surface First:
- What was offered load vs admitted load during each phase?
- What retry policy does each client (iOS, Android, web, SDKs) implement for 429 and 503?
- Did
Retry-Afterexist on both responses? Did clients read it? - Why did the 4-minute orders issue spread to other routes?
Typical L5 Approach: Recommends exponential backoff in the mobile app and moves on. Correct, and it takes 6 weeks for 80% of users to update the app.
Staff Approach: Builds the load timeline: offered load reached 6.2× normal at minute 9, of which 84% was retries. Retries hit all routes because the app's network layer retried the whole screen's request set. Three fixes on different timescales: server-side within a day — a retry budget at the gateway (per client install, at most 10% of requests may be retries; excess retries get 429 with long
Retry-After) andRetry-Afteron all 503s; client within one release — honorRetry-After, full-jitter backoff, retry only idempotent requests, max 2 retries; platform within a quarter — a shared HTTP client library all apps must use.
Principal Approach: Identifies that no one owned the client side of the reliability contract. Establishes a cross-platform client reliability standard with the mobile, web and SDK teams, owned jointly by API platform and mobile platform. Adds a server-side kill switch that can tell old app versions to back off for N minutes via a config endpoint, because the long tail of app versions lives for a year.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (in the incident) | Should have been: enable gateway-level retry budget; return long Retry-After to clients exceeding it. It didn't exist. |
| Triage (post-mortem data) | X-Retry-Attempt header (added by the app) shows 84% of traffic at minute 9 was attempts 2–5. |
| Quick fix | Retry-After on all 503/429; gateway retry budget per install ID. |
| Guardrails | Alert on gateway.retry_ratio > 0.2; load tests include client retry behavior, not just raw load. |
| Post-mortem | Why did 429 and 503 trigger identical client behavior? Why did one service's shedding cascade to all routes? Who owns client retry policy? |
Metrics to Watch: gateway.retry_ratio, offered vs admitted load, client.retry_attempt distribution by app version, Retry-After compliance per client type.
Organizational Follow-up: mobile and API platform co-own a client networking library; app releases require passing a "429/503 storm" test.
Ownership Question: "Who owns retry behavior?" Staff answer: The server owns telling clients what to do; the client platform owns doing it. Until both exist, the server must defend itself with retry budgets.
Key Takeaway: "Every 429 is an instruction. If clients don't follow it, the limiter amplifies the outage it was built to prevent."
What clears the Staff bar:
- Quantifies retries as a share of offered load
- Fixes on three timescales — server today, client next release, platform this quarter
- Assigns ownership of the client half of the contract
Deep Dive 5: Multi-Region Expansion — The Quota That Doubled#
Context: The API launches in a second region (EU) with active-active traffic. Two weeks later, finance notices several tenants' metered usage is ~1.8× their plan limits. Each region enforces its own limits correctly. Nobody designed global limits.
Questions to Surface First:
- Are contracts written as global or per-region limits?
- How do tenants route — geo-DNS, or can one client hit both regions?
- Is metering global (billing correct) while enforcement is regional?
- Which tenants are affected, and is it deliberate (clients splitting traffic) or incidental?
Typical L5 Approach: Proposes a global Redis with cross-region replication for counters. Every request in the EU now pays 80ms to check a US primary, or uses an eventually-consistent replica that doesn't solve the problem.
Staff Approach: Confirms contracts are global. Implements demand-based allocation: each region's coordinator reports per-tenant consumption to a small global aggregator every 2 seconds; the aggregator allocates the tenant's global limit across regions in proportion to recent demand, holding back a 10% reserve. Regions enforce their allocation with the existing leased design. On partition, each region keeps its last allocation. Overshoot: bounded by reserve plus one rebalance period.
Principal Approach: Questions the contract shape: per-region limits are simpler to enforce, match data-residency boundaries, and are easier for customers to reason about. Proposes moving new contracts to per-region limits with a global monthly quota enforced by metering, and only keeping real-time global enforcement for the ~20 tenants whose contracts require it.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate | Quantify: 212 tenants above 1.2× plan; most are clients with endpoints configured for both regions. |
| Triage | No over-capacity risk yet (EU is 15% of traffic); revenue risk is that heavy users are on cheap plans. |
| Quick fix | Temporarily set each region's limit to plan × regional share from the last 7 days — crude static allocation. |
| Guardrails | Metering compares global usage to plan daily; alert per tenant above 1.1×. |
| Long-term | Global aggregator with demand-based allocation; contracts revised to specify scope. |
Metrics to Watch: quota.global_utilization per tenant, allocation.rebalance_lag_s, allocation.reserve_claimed, regional limiter.rejected deltas after rebalances.
Organizational Follow-up: legal and product update API terms to define limit scope; the multi-region readiness checklist gains a "limits and quotas" line.
Ownership Question: "Who decides whether limits are global or regional?" Staff answer: Product, because it's a contract term. Engineering presents the cost of each: global real-time enforcement is a new coordination system; regional limits are free.
Key Takeaway: "Global limits are an allocation problem, not a counting problem. Never put a cross-region round trip on the request path."
What clears the Staff bar:
- Rejects synchronous global counting with the latency number
- Designs allocation with a reserve and partition behavior
- Asks whether the contract should be global at all
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Separate contract enforcement, abuse defense and capacity protection, and explain why each wants different keys, precision and failure posture
- Choose token bucket or GCRA in under a minute and move on
- Design leased local budgets with a stated overshoot bound, and size the store for 2M req/s
- Give a per-intent blind-mode posture with fallback math and named owners
- Explain when rate is the wrong unit and apply cost units or concurrency caps
- Define the 429 contract — headers, codes,
Retry-After, 429 vs 503 — and the client behavior it requires - Ship a limit change through shadow, canary and enforcement with a kill switch
- Enforce a global quota across regions via allocation rather than synchronous checks
- Price the limiter and argue when to buy instead of build
The Bar for This Question#
Mid-level (L4): Implements a correct token bucket or sliding window in Redis with an atomic script, keyed by user or IP, returning 429. Handles the happy path. Hasn't thought about what happens when Redis is slow, when the key is shared by thousands of people, or when the client retries.
Senior (L5): Adds sharding, replication, a fail-open answer when asked, and probably a local cache of counters. Knows the fixed-window boundary problem. The gap: treats all limiting as one mechanism; can't bound the overshoot of local counting; doesn't separate 429 from 503; edits limits through deploys. The design works in steady state and causes the next incident during a launch or a policy change.
Staff+ (L6): Frames the problem by intent within the first three minutes and commits. States the precision tolerance and latency budget, then designs to them — leases, sync, bounded overshoot. Has the blind-mode posture table ready, with owners. Uses concurrency caps for slow routes and distinguishes contract limits from load shedding. Treats the 429 as an API contract and policy changes as risky deploys with shadow mode. Names who pays for every choice: customers for false 429s, backends for overshoot, security for fail-open windows. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "Your Limiter Should Be Wrong on Purpose"#
| Common Belief | Reality |
|---|---|
| "A limiter must be accurate" | Customers can't tell 500/s from 520/s; they can tell 1ms from 50ms |
| "Approximate means unsafe" | Approximate with a stated bound is a spec; exact with a remote dependency is a risk |
| "We'll add local caching later" | Retrofitting local counting changes failure behavior; customers will notice |
The Staff position: Pick an error bound — say 5% for a second — and design to it. Precision beyond what any customer or backend can perceive is paid for with latency and availability that everyone perceives.
Why this matters in interviews: Saying "I'm choosing to be ~5% wrong, here's the bound, here's who agreed" is the clearest single signal that you design to requirements rather than to an ideal.
10.2 "The 429 Is a Product Surface, Not an Error Code"#
| What Teams Do | What Happens |
|---|---|
| Generic error body on 429 | Clients can't tell plan limits from abuse blocks; support tickets ask "why" |
No Retry-After | Clients guess; guesses synchronize; retry storms |
| 429 for both quota and overload | Customers' dashboards blame themselves for your outages, or vice versa |
The Staff position: Design the rejection like an API endpoint: documented codes, headers that tell the client exactly when to return, a link to the plan's limits, and SDKs that behave. Product owns the wording; platform owns the semantics.
Why this matters in interviews: Most candidates stop at "return 429". Spending one minute on the response contract shows you've watched real clients react to a limiter.
10.3 "Per-IP Limits Are Mostly Security Theater for Authenticated APIs"#
| Attack / Situation | Per-IP Limit Effect |
|---|---|
| Credential stuffing via residential proxies | Each IP sends 2–5 attempts; never trips |
| Corporate or campus NAT | Thousands of legitimate users share one IP; false positives |
| IPv6 | One client can rotate across a /64 trivially; per-address keys are meaningless without prefix aggregation |
| Volumetric L3/L4 floods | Handled by the edge provider long before your limiter |
The Staff position: For authenticated traffic, key on the identity you authenticated. For pre-auth endpoints, combine target-account limits, IP prefix aggregation, device signals and challenges. Per-IP limits remain useful as a high-threshold backstop, not as the primary control.
Why this matters in interviews: "Rate limit by IP" is the default answer for abuse. Explaining why it fails — and what replaces it — moves the conversation to adversarial thinking.
10.4 "A Limit Without Shadow Mode Is an Untested Deploy to Production"#
| Change | Typical Outcome Without Shadow |
|---|---|
| New abuse rule | NAT cliff for campuses and carriers |
| Tier reduction | Partner integrations broken without notice |
| Unit change | Largest customer throttled to 1/60th |
The Staff position: Every limiter must support computing a decision without enforcing it, and every policy change uses it. Shadow mode costs one extra evaluation per request — nanoseconds locally — and converts "what will this do?" from opinion into a list of affected tenants.
Why this matters in interviews: Naming shadow mode with declared expected rejection rates shows operational maturity that algorithm knowledge never will.
10.5 "For Protecting Backends, Concurrency Beats Rate Every Time"#
| Situation | Rate Limit Sees | Concurrency Limit Sees |
|---|---|---|
| Backend slows from 20ms to 2s | Same request rate; admits as before → queue explodes | In-flight count rises 100×; limit trips immediately |
| One tenant's requests get 50× more expensive | Nothing | Their in-flight share grows; capped |
| Deploy changes per-request cost | Static number now wrong | Adaptive limit re-learns within seconds |
The Staff position: Rate limits are for contracts. For keeping a service alive, use concurrency — ideally adaptive, ideally per service, with no shared state. Little's law makes the case: in-flight = rate × latency, and latency is the variable that changes during an incident.
Why this matters in interviews: Distinguishing the tools — and knowing that Little's law is the reason — is a crisp, defensible Staff opinion that most candidates don't have.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
The Staff engineer designs one excellent limiter. The Principal engineer discovers there are already nine: the public gateway's, two in-house service meshes with different semantics, a login limiter security wrote in 2022, a homegrown one in the payments service, the CDN's, a per-vendor limiter around the SMS provider, and two queue consumers that sleep in a loop. Each was correct when written. Together they produce incidents nobody can explain — a request rejected by three layers, each with a different error code, none of them visible to the others. The L7 problem is admission control as an organizational capability: one policy model, one vocabulary of rejection, one place to see why any request was refused, and a clear line between what the platform standardizes and what teams keep.
The Org-Level Fault Line#
One admission-control platform vs limiters embedded in each team's service.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Every team builds its own | Fast; tailored to each service | Inconsistent 429 semantics; no cross-layer visibility; each team re-learns blind-mode lessons through incidents | Customers (confusing errors), on-call (multi-layer debugging), the company (N× the same mistakes) |
| Central platform enforces everything at the gateway | One policy model; one decision log | Gateway can't see service-internal costs (rows scanned, GPU seconds); platform becomes a bottleneck for every limit change | Product teams (velocity); platform team (owns everyone's capacity knowledge) |
| Platform owns primitives and contract limits; services own capacity protection with platform libraries | Contracts consistent at the edge; capacity protection where the cost is known; shared libraries and telemetry | Requires a well-designed library and a shared decision-log schema | Platform (library stewardship); service teams (adopt and tune) |
🧭 Principal Move: "The platform owns three things: the policy model, the contract limiter at the gateway, and the libraries for concurrency limiting and shedding. Services own their capacity numbers and run the library in-process. Every rejection anywhere emits the same decision record with a reason code, so 'why was this request refused?' has one query, not nine."
Cost Model#
Assumptions: ~$250K fully loaded per engineer-year; cloud list prices; a 30% on-call burden counted as headcount; limiter infra counted separately from the gateway fleet it runs inside.
| Scale | Traffic | Infra ($/month) | Headcount | On-call Load | Cost of a False 429 |
|---|---|---|---|---|---|
| Startup | ~5K req/s | ~$0–500 (managed gateway feature, or a small Redis) | 0.25 eng (part of the API team) | Shared; limiter pages < 1/quarter | Mostly invisible; a frustrated developer |
| Growth | ~200K req/s, 20K tenants | ~$3–8K (3-shard store, policy service, decision logs at 1% sampling) | 2–3 eng platform + 0.5 security | Platform rotation; 1–3 limiter-related pages/month | A support ticket per incident; occasional enterprise escalation |
| Enterprise | ~2M+ req/s, multi-region, 200+ services | ~$30–80K (regional stores, global aggregator, full decision logging for rejected requests, shadow evaluation) | 6–10 eng (gateway limiter, libraries, policy tooling) + trust & safety partnership | Dedicated platform rotation; game days quarterly | A wrongly throttled top-10 customer can cost a renewal worth millions — weigh tooling against that |
The pricing insight: the limiter's own infrastructure is never the big number. The big numbers are (a) backend spend it avoids by rejecting over-plan and abusive traffic early — often 2–5% of total request volume — and (b) the revenue risk of false rejections on large accounts. A 3-engineer investment in shadow-mode tooling and per-tenant anomaly alerts is justified by preventing one enterprise escalation per year.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
| Limit semantics published to customers (per-second vs per-minute, global vs regional, what counts as a request) | One-way | Changing them is a contract change; customers build integrations around the published numbers |
| 429 response format and header names | One-way (once SDKs ship) | Old SDK versions parse it for years; changes need dual-emission for 12+ months |
| Pricing tiers expressed in limits | One-way | Downgrading entitlements needs notice periods and causes churn |
| Cell / dedicated-capacity architecture for large tenants | One-way-ish | Moving tenants between cells is a migration project |
| Algorithm (token bucket vs GCRA vs sliding counter) | Two-way | Internal; swap behind the same policy model |
| Store technology for leases | Two-way | State is regenerable in seconds |
| Sync interval, lease sizing, fallback multiplier | Two-way | Config |
| Blind-mode posture per policy | Two-way | Policy change with sign-off; minutes |
The Standard I'd Write#
RFC-ADM-001: Admission Control Standard
Status: Approved Owners: API Platform + Security + SRE
Scope
Every component that rejects, delays or sheds requests or messages on behalf
of a customer, tenant or internal caller.
MUST
1. Classify each limit as CONTRACT, ABUSE or CAPACITY, and declare its key,
unit, expected rejection rate and blind-mode posture.
2. Not place a synchronous remote call on the request path for limits with
tolerance above 5%. Exceptions require platform review.
3. Return 429 only for CONTRACT/ABUSE rejections and 503 for CAPACITY
rejections, both with Retry-After.
4. Emit a decision record (policy_id, version, key hash, reason, node,
sync_staleness_ms) for every rejection.
5. Roll out every new or tightened policy in shadow mode for at least 24h,
or 7 days if it affects any top-100 tenant.
6. Expire every override automatically (maximum 90 days).
SHOULD
1. Use the platform concurrency-limit library for routes with p99 above 1s.
2. Use the platform SDK retry policy: full-jitter exponential backoff,
honor Retry-After, max 2 retries for non-idempotent requests.
3. Prefer cost units over request counts where per-request cost varies 10x+.
Exceptions
Filed with API Platform; reviewed within 5 business days; time-boxed to
2 quarters. Security may bypass shadow for active-attack blocks via the
break-glass path, with automatic 4-hour expiry.
Success metrics
- False-rejection escalations from top-100 tenants: 0 per quarter
- Limiter-attributed latency: under 1ms p99 on the gateway
- Unmatched authenticated requests: under 0.1%
- Teams running non-standard limiters: 0 by end of year 2
- Incidents where a rejection's cause took more than 15 min to find: 0
What I'd Tell the VP#
"Today nine different systems can reject a customer's request, each with its own rules and error messages, and when something goes wrong it takes us hours to work out which one did it. Twice this year, a rate-limit mistake throttled a top-20 customer. I'm proposing one admission-control platform: a shared policy model, a standard way to roll out limit changes safely, and one place to see why any request was refused. It's about four engineers for three quarters. It removes a recurring class of customer-facing incidents, cuts backend spend by rejecting abusive traffic earlier, and lets sales offer dedicated-capacity tiers we can actually guarantee."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Sees the class of systems | "We have nine limiters. The problem isn't any one of them — it's that a rejection can come from any of them with no shared record." |
| Prices the tradeoff | "The store costs $5K a month. The tooling that stops us throttling a $3M account is the real investment." |
| Identifies one-way doors | "The limit semantics we publish are the one-way door. The algorithm behind them isn't." |
| Redraws ownership | "Platform owns the contract layer and the libraries; services own their capacity numbers; security owns abuse posture." |
| Knows when not to standardize | "The SMS vendor limiter stays bespoke — it mirrors a vendor contract, not our policy model." |
Staff answers that L7 interviewers find insufficient:
- "We'll build a great limiter at the gateway" — correct, and it leaves eight other limiters with different semantics in place.
- "Failure posture is configurable per policy" — right mechanism; missing who signs each posture and how it's audited.
- "We'll make limits global across regions" — no question of whether the contract should be global, and no cost of the coordination layer.
Appendices
Appendix A: Mechanics in Depth#
A.1 GCRA — Token Bucket With One Field#
GCRA stores a single timestamp per key: the theoretical arrival time (TAT) of the next request if the client sent at exactly the sustained rate.
emission_interval T = 1 / rate # e.g. 500/s → 2ms
burst_tolerance τ = T × (burst - 1) # how far ahead of schedule a client may run
def admit(key, now, cost=1):
tat = load(key) or now
tat = max(tat, now)
new_tat = tat + T × cost
allow_at = new_tat - τ - T # earliest time this request may be admitted
if now < allow_at:
return REJECT, retry_after = allow_at - now
store(key, new_tat, ttl = new_tat - now + τ)
return ALLOW, remaining = floor((now + τ + T - new_tat) / T)
Why it's right for contracts: one field, exact Retry-After (not a guess), natural support for weighted cost, no window boundaries. Why it's wrong for cost known only afterward: it debits up front; use the two-phase reconcile pattern (Fault Line 3) with a bucket that may go negative.
A.2 Token Bucket — The Version Everyone Can Explain#
def admit(bucket, now, cost=1):
elapsed = now - bucket.last_refill
bucket.tokens = min(bucket.capacity, bucket.tokens + elapsed × bucket.rate)
bucket.last_refill = now
if bucket.tokens >= cost:
bucket.tokens -= cost
return ALLOW
return REJECT, retry_after = (cost - bucket.tokens) / bucket.rate
Locally, now must come from a monotonic clock; wall-clock jumps (NTP step corrections) otherwise mint or destroy tokens. In a shared store, use the store's clock inside the atomic script so nodes with skewed clocks can't disagree about refill.
A.3 Sliding Window Counter — Where It Still Wins#
Two fixed-window counters and a weighted estimate: estimate = prev_count × (1 − elapsed_fraction) + curr_count. Two INCRs per request and natural expiry make it the cheapest shared-store design for coarse, long windows (per-hour, per-day caps) where token-bucket state per key would be overkill and the boundary error is irrelevant.
A.4 Concurrency Limiter With Leases#
def acquire(tenant, route, now):
slots = inflight[tenant, route] # map: request_id → expires_at
purge(slots, now) # drop holders past lease (crashed requests)
if len(slots) >= cap(tenant, route):
return REJECT_503_OR_QUEUE
rid = new_request_id()
slots[rid] = now + 2 × p99(route)
return rid
def release(tenant, route, rid):
inflight[tenant, route].pop(rid, None)
Per-node concurrency caps are usually sufficient — concurrency is a local resource (threads, connections). A fleet-wide concurrency cap needs leases in a shared store and is rarely worth it.
A.5 Adaptive Concurrency (Capacity Layer)#
The service measures its no-load latency (minimum RTT over a window) and current latency, and adjusts its limit: grow additively when latency stays near the minimum, shrink multiplicatively when latency rises, e.g. new_limit = limit × (min_rtt / current_rtt) + headroom. Requests above the limit are rejected immediately with 503 — queuing them would only add latency. Priority classes are admitted in order; batch traffic is rejected first.
Appendix B: Keys and Identity#
| Layer | Key Construction | Notes |
|---|---|---|
| Contract | {tenant_id}:{route_class} | Route classes are a small enum (read, write, search, export, admin) — never raw paths, which explode cardinality |
| Contract hierarchy | {org_id} and {org_id}:{api_key_id} both checked | An org's limit caps all its keys; a key limit protects the org from one runaway integration |
| Abuse — login | acct:{username_hash}, ip4:{/24}, ip6:{/56 or /64}, dev:{fingerprint}, asn:{asn} | Evaluate all; reject or challenge if any exceeds; thresholds differ per key type |
| Abuse — signup / OTP | phone:{e164_hash}, email_domain:{domain}, ip:{prefix} | OTP limits are cost limits too — each SMS costs money |
| Capacity | priority_class only | Identity used for fairness ordering, not admission |
Hashing user-supplied values (usernames, phone numbers) before using them as keys keeps PII out of the store and the decision logs. Cardinality bounds: for unauthenticated keys, cap tracked keys per node (e.g., 1M with LRU eviction) — an attacker spraying random keys must not exhaust memory; evicted keys reset to "new", which is acceptable for coarse limits.
Appendix C: Coordination Mechanisms#
C.1 Lease Sync Protocol#
The store-side operation per key is one atomic script: refill, credit unspent tokens from the expired lease, grant up to min(requested, available), decrement. Batching all keys from a node into one round trip is what keeps store load proportional to node count, not request rate.
C.2 Multi-Region Allocation#
On partition, a region keeps its last allocation; allocations always sum to at most the global limit, so a partition can only cause under-use, plus whatever reserve a region claimed.
C.3 Quick Comparison#
| Mechanism | Request-Path Latency | Overshoot | Store Load | Survives Store Outage | Best For |
|---|---|---|---|---|---|
| Synchronous script per request | +0.3–5ms | ~0 | 1 op/request | No (needs fallback) | Small, exact, high-stakes counters |
| Pure local (limit ÷ N) | ~0 | High under skew; wrong when N changes | None | Yes | Small fleets, smooth traffic |
| Local + periodic sync (full view) | ~0 | Up to N × remaining per interval | 1 batch/node/interval | Degrades | Rarely right — use leases |
| Leased budgets | ~0 | ~20% of one interval's traffic | 1 batch/node/interval | Degrades gracefully | Contract limits at scale |
| Regional allocation + leases | ~0 | Reserve + one rebalance period | Tiny, cross-region | Yes (keeps last allocation) | Global quotas |
Appendix D: API Contract & Client Behavior#
| Element | Contract |
|---|---|
429 | You exceeded your entitlement (contract or abuse). Body code says which policy. |
503 + Retry-After | We are shedding load. Not your fault; not counted against your plan. |
Retry-After | Seconds until a retry is likely to succeed; servers add jitter so clients don't synchronize |
| Remaining / reset headers | Let well-behaved clients self-throttle before hitting 429 |
| Retries of writes | Only with an idempotency key; otherwise a retry after a lost response can duplicate the write |
| SDK default | Honor Retry-After; otherwise full-jitter exponential backoff from 200ms, cap 30s, max 2–3 retries |
| Server-side retry budget | If a client's retry ratio exceeds ~10%, further retries get long Retry-After — defends against clients that ignore the contract |
Appendix E: Observability#
E.1 Core Metrics#
| Metric | Meaning | Alert |
|---|---|---|
limiter.decision_latency_p99 | Time spent deciding, per request | > 1ms for 5 min: page |
limiter.rejected{reason, policy, tenant_tier} | Rejections by cause | Top-50 tenant > 1% rejected vs baseline: page |
limiter.unmatched_requests | Authenticated requests matching no policy | > 0.1% for 10 min: page |
limiter.sync_staleness_ms | Age of last successful sync per node | p99 > 2s: ticket; > 60s fleet-wide: page |
limiter.blind_mode | Nodes in fallback | > 5% of nodes: ticket; > 15 min: page |
limiter.shadow_would_reject{policy} | Shadow policy impact | Feeds rollout tooling; no alert |
limiter.lease_utilization | Granted tokens actually spent | < 50%: leases oversized (tuning ticket) |
gateway.retry_ratio | Share of requests that are retries | > 0.2: page |
shed.rejected{priority} | Capacity-layer rejections | Any critical shed: page |
E.2 Control Plane vs Data Plane#
The data plane (local buckets, concurrency caps, shedders) must keep working with the control plane completely down: nodes cache the last-good policy set on local disk and start with it after a restart. The control plane (policy service, lease store, aggregator) may be unavailable for minutes; its failure degrades precision, never availability. If a node restarts while the policy service is down and has no cached policy, it starts with a built-in conservative default — the posture is "limited", never "unlimited".
E.3 Debugging the Silent Failure#
A customer says "I'm getting 429s but I'm under my limit." The decision record answers it in one query: which policy and version, which node, the node's local view of the bucket, its sync staleness at decision time, and whether it was in blind mode. Without that record the investigation takes a day. Store rejected-request decision records at 100%; sample admitted ones at ~0.1%.
Appendix F: Scale Evolution#
| Scale | What Works | What Breaks Next |
|---|---|---|
| < 10K req/s, < 10 nodes | Managed gateway limits or per-node limits; synchronous Redis is fine | Per-node limits under skew; per-minute granularity |
| 10K–200K req/s | Synchronous store checks with tight timeouts + local fallback, or local + lease sync | Hot tenants on one shard; latency tail |
| 200K–2M req/s | Leased budgets, batched sync, policy service with shadow mode | Multiple teams' limiters; cross-layer debugging |
| 2M+ req/s, multi-region | Regional allocation, cells for large tenants, admission-control platform | Organizational consistency, contract semantics |
What You Don't Build on Day One:
- Global cross-region quotas — start with per-region limits and global monthly metering
- Adaptive contract limits — keep contracts static; put adaptivity in the shedder
- Exact sliding-window logs — only for small abuse counters
- A custom proxy — extend the gateway or Envoy
- Per-endpoint limits for 400 routes — start with 5 route classes
Appendix G: Multi-Tenancy, Fairness & Cost#
Hierarchical limits. Org → API key → route class. A request must pass every level. This protects the org from one runaway key and protects the platform from an org that creates 500 keys to multiply its limit.
Fair sharing under scarcity. When the capacity layer is shedding, tenants within the same priority class should be shed proportionally to their share, not first-come. Weighted fair queuing with weights from plan tier is the standard approach; simpler variants drop a tenant's requests with probability proportional to how far that tenant exceeds its fair share.
Noisy-neighbor isolation. Rate limits cap arrival; they don't isolate. For tenants above ~5% of total traffic, give them a dedicated cell or capacity pool. The limiter's job then becomes enforcing each cell's contract, and a tenant's incident stays inside its cell.
Cost attribution. Admitted cost units per tenant, from the limiter, are the cheapest available cost-to-serve signal. Feed them to capacity planning and pricing — but never to invoicing, which reads the deduplicated metering stream.
Pricing alignment. If plans are sold in requests, limit in requests. If you charge by compute or tokens, limit in the same unit. A mismatch between the limit unit and the billing unit is an arbitrage customers will find — and a support conversation nobody enjoys.