Why This Matters#
An API gateway is not a product you pick. It is a decision to put one piece of shared code in front of every request your company serves — and then a decision about which policies live in that code and who is allowed to change them. Envoy vs Kong vs AWS API Gateway is the easy part of the conversation. The hard part is that the gateway becomes the single largest shared blast radius in your architecture, and its configuration becomes a file that 40 teams want to edit on the same afternoon.
Gateways show up in nearly every Staff loop because every design has a front door. "Put an API gateway in front" is the most common box candidates draw, and the least examined. The L5 candidate draws the box and lists features: auth, rate limiting, routing, logging. The L6 candidate says "Envoy-based edge tier, stateless, terminating TLS and validating JWTs locally against a cached JWKS; coarse rate limits locally per pod with a global limit for paid tiers; no business logic and no response aggregation — that's a BFF owned by the product team; routes are per-team HTTPRoute objects in their own repos, and a bad route can only break that team's paths." The L7 candidate asks what the org's policy is for what may be added to the gateway, because a gateway with no such policy turns into an enterprise service bus within two years.
The L5 → L6 gap is not knowing what xDS stands for. It is knowing that every feature you put in the gateway is a feature every request pays for, and every config change is a change to every team's availability. The design problem — building a gateway for a specific system — is covered in the API Gateway case study. This page is about the technology choices and how to run them.
The L5 → L6 → L7 Contrast#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Add an API gateway for auth and routing" | "Which concerns are cross-cutting enough to centralize, and which belong in services? I'll keep the gateway thin." | "What's our org rule for what goes into the gateway, and who approves exceptions? Without one it becomes an ESB." |
| Product choice | "Kong, because it has plugins" | "Envoy data plane for performance and xDS; a managed gateway only for the public partner API where we need keys, quotas and a portal" | Standardizes on one data plane across edge, ingress and mesh to share skills and tooling; keeps the vendor layer replaceable |
| Auth | "Gateway checks the token" | "Local JWT validation at the gateway (no network call), fine-grained authz in services; fail-closed on authn, documented behavior if JWKS refresh fails" | Defines the identity contract: what claims the gateway guarantees to services, and which services may trust them |
| Failure | "Run multiple instances" | "Config is the outage vector: validate, canary to 1% of the fleet, auto-rollback on 5xx" | Cell-based gateway fleets so one bad config or noisy tenant can't take down everything |
| Ownership | "Platform team owns the gateway" | "Platform owns the data plane and global policies; each team owns its routes, with blast radius scoped to its paths" | Redraws boundaries: platform-owned guardrails, team-owned routes, a paved road with an exception process |
| Cost | "It's open source" | "~2–4 cores per 10K rps with TLS and auth; managed API Gateway is ~$1–3.50 per million requests" | Prices managed vs self-run at 3 scales, including the on-call rotation |
Why "First move" separates levels
"Auth, rate limiting, routing and logging at the gateway" is a reasonable list. What it misses is that each item has a cost and an owner. JWT validation is cheap and stateless — centralize it. Fine-grained authorization ("can user X edit document Y?") needs domain data — centralizing it couples the gateway to every service's data model. Response aggregation looks convenient and makes the gateway team the bottleneck for every product launch. The Staff candidate sorts concerns by how much domain knowledge they need and centralizes only the ones that need none.
Why "Failure" separates levels
Gateways are stateless and horizontally scaled, so "run more instances" handles hardware failure. But gateway outages are rarely hardware. They are configuration pushes — a malformed route, an over-broad regex, a plugin bug, an expired certificate — and a config push reaches every instance in seconds. The Staff answer treats config like code deployment (validation, canary, rollback). The Principal answer shards the gateway itself into cells so that even a successful bad push has a bounded blast radius.
The 60-Second Pitch#
"I'd run an Envoy-based gateway tier as the single public entry point. It terminates TLS, validates JWTs locally against cached signing keys, applies per-client rate limits, and routes by host and path to service clusters — adding about 1ms at p50. It's stateless, so it scales horizontally behind an L4 load balancer and an anycast or DNS front. It does not contain business logic or aggregate responses; product teams own BFFs for that. Each team owns its routes as declarative config in its own repo, validated in CI and rolled out to a canary slice of the fleet first, so a bad route breaks one team's paths for a minute rather than the whole site. For the external partner API, where we need API keys, usage plans and a developer portal, I'd put a managed API-management product in front of the same backends rather than build those features."
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Edge / ingress traffic management (TLS, routing, canaries, protection for your own apps) | Low added latency, high throughput, fast safe config changes | Envoy / NGINX / HAProxy-class proxy, declarative routes, K8s Gateway API | Bad config push takes down every route | Every request routed correctly; no added outage risk beyond the fleet |
| External API product management (partners, public developers) | API keys, quotas, monetization, versioning, developer portal, analytics | Kong / Apigee / Azure APIM / AWS API Gateway in front of internal ingress | Quota or key store outage locks out paying customers | Metering accurate enough to bill; contract stability for years |
| Aggregation / BFF (one client call fanning out to many services) | Client-specific shapes, fewer round trips on mobile | Service owned by the product team (GraphQL or REST BFF) behind the gateway | Gateway becomes the bottleneck for every product change | Correct partial-failure semantics per client |
🎯 Staff Move: "I'm designing for the ingress intent: a thin, fast, safe front door. API-product features like keys and billing I'll buy for the partner API only, and aggregation lives in a BFF the mobile team owns — I don't want the gateway team on the critical path of every feature launch."
The Staff Positions#
| Position | Rationale |
|---|---|
| Thin gateway, thick services | Centralize only what needs no domain knowledge: TLS, authn, coarse limits, routing, telemetry. |
| Config is the outage vector | Treat route and policy changes like deploys: validate, canary, auto-rollback. |
| Authenticate at the edge, authorize in the service | Token validation is generic; "can X do Y to Z" needs the service's data. |
| No synchronous dependency without a timeout and a failure mode | Every external call from the gateway (authz, rate limit, key lookup) has a deadline and a declared fail-open/closed. |
| Retry at one layer | The gateway retries only idempotent requests, with a budget; services below it don't also retry the same failure. |
| Teams own their routes; platform owns the guardrails | Scoped config ownership keeps one team's mistake from becoming everyone's outage. |
| Buy API-management features, run the data plane | Portals, key lifecycle and metering are commodity; the hot path is where you need control. |
Architecture & Internals#
Only the internals that change design decisions.
Data Plane vs Control Plane#
Every modern gateway splits into a data plane (the proxies that handle requests) and a control plane (the system that computes and distributes configuration). Most gateway outages live on the boundary between them.
| Design Question | Why It Matters |
|---|---|
| What happens to the data plane if the control plane dies? | Good gateways keep serving the last-known-good config. Verify this — it's the difference between "can't change routes" and "site down." |
| How fast does config propagate? | Envoy via xDS: seconds. NGINX reload: seconds, but spawns new workers and can double memory briefly. Managed gateways: often tens of seconds to minutes per deploy. |
| Is config pushed as a whole or as deltas? | Full snapshots of 10K routes to 500 proxies is a lot of CPU and bandwidth; incremental xDS (delta) scales better. |
| Can one tenant's config break another's? | A single global config file means a syntax error anywhere rejects the whole push. |
The Request Pipeline (Filter Chain)#
A gateway is a sequence of filters. Each one adds latency, may call out to another system, and may fail. Order matters.
Put cheap rejections first. Reject unauthenticated or over-limit requests before doing anything expensive. A request that fails JWT validation should never consume a global rate-limit call or an upstream connection.
Configuration Models by Product#
| Product | Config Model | Dynamic Changes | What It Implies |
|---|---|---|---|
| Envoy | Protobuf/YAML; dynamic via xDS APIs (listeners, routes, clusters, endpoints, secrets) | Yes, no restart | You need a control plane (Istio, Envoy Gateway, Gloo, or your own) |
| NGINX (OSS) | Static config files | Reload (graceful; new workers) | Fine for tens of routes; painful for thousands that change hourly |
| NGINX Plus / OpenResty | Files + API/Lua for dynamic upstreams | Partial | Commercial features or Lua code for dynamism |
| Kong | Admin API backed by Postgres, or DB-less declarative YAML | Yes | DB-less + GitOps (decK) is the safer mode for production |
| HAProxy | Config file + Runtime API | Partial (servers, maps, weights) | Extremely efficient; less built-in API-management |
| Traefik | Auto-discovery from Kubernetes/Docker labels | Yes | Convenient; discovery-from-labels means app teams change edge config implicitly |
| AWS API Gateway | Console/API/IaC; deploy to stages | Per-deployment | No servers; limits and pricing are the design constraints |
| K8s Gateway API | CRDs: GatewayClass, Gateway, HTTPRoute, ... | Yes (via implementation) | Vendor-neutral spec with role separation built in |
Concurrency and Connection Handling#
Envoy, NGINX and HAProxy all use event loops, roughly one worker per core, each handling thousands of concurrent connections. Consequences for design:
- A blocking call in a filter blocks every request on that worker. Plugins (Lua, Wasm, JavaScript) must be non-blocking; a synchronous HTTP call in a Lua plugin is a classic self-inflicted outage.
- Upstream connection pools are per worker. 500 proxies × 16 workers × 10 connections = 80,000 connections into a service. Use HTTP/2 to upstreams or cap pool sizes, or the gateway becomes the service's largest connection storm on restart.
- Long-lived client connections (gRPC, WebSockets) pin to one proxy. Draining for deploys takes as long as the longest connection allows; set a max connection age (e.g., 10–30 min) to keep fleets rebalanced.
- Request buffering costs memory. Buffering large bodies (uploads) at the gateway turns a 2GB upload into 2GB of proxy memory. Stream bodies or send uploads directly to object storage with pre-signed URLs.
Where State Lives#
Gateways should be stateless, but several features need state:
| Feature | State | Typical Home | Failure Behavior to Decide |
|---|---|---|---|
| JWT validation | Signing keys (JWKS) | In-memory, refreshed every ~5–15 min | If refresh fails, keep last-known keys (fail-static) |
| Global rate limits | Counters | Redis or a rate-limit service | Fail-open for abuse protection; fail-closed for paid quota enforcement only with care |
| API keys / usage plans | Key → plan mapping | DB with local cache | Cache with a TTL so a DB blip doesn't lock out customers |
| Session affinity | Hash ring / cookie | In-proxy | Rebalance on scale events |
| Response cache | Cached bodies | In-proxy or CDN | Usually better at the CDN — see Caching Fundamentals |
Core Usage — "The Entire Game": Deciding What Goes in the Gateway#
The quality of a gateway design is determined almost entirely by what you refuse to put in it. Every concern has a natural home; the gateway is the right home only for concerns that are uniform across services and need no domain data.
The Placement Test#
| Concern | Gateway | Service / BFF | Mesh (east-west) | Reasoning |
|---|---|---|---|---|
| TLS termination (public) | Yes | — | mTLS internally | Centralized certs, modern ciphers, one rotation process |
| Authentication (token validity, signature, expiry) | Yes | Re-verify if zero-trust | Workload identity | Generic, stateless, cheap |
Coarse authorization (scope orders:read present?) | Yes | — | — | Needs only token claims |
| Fine-grained authorization (owns order 123?) | No | Yes | — | Needs domain data |
| Per-client rate limiting / quotas | Yes | Per-resource limits | — | Client identity is known at the edge |
| Request schema validation | Partial (size, content-type) | Yes (semantics) | — | Full schema validation at the edge couples deploys |
| Routing, canaries, traffic splitting | Yes | — | Yes internally | Pure traffic concern |
| Retries, timeouts | Edge deadline, idempotent retries | Own downstream calls | Yes internally | One retry layer per hop |
| Response aggregation / composition | No | BFF | — | Product logic; changes weekly |
| Protocol translation (REST → gRPC) | OK if mechanical | — | — | gRPC-JSON transcoding is fine; hand-written mapping is not |
| Caching | Rarely | Domain caches | — | CDN for public, service caches for domain data |
| Business rules, pricing, feature flags | Never | Yes | — | The ESB trap |
🎯 Staff Move: "The gateway gets anything I could configure identically for every service — TLS, token validation, client rate limits, routing, telemetry. If a policy needs to know what an order or a document is, it lives in the service."
Route Ownership as Config#
The Kubernetes Gateway API makes ownership explicit: infrastructure owns the GatewayClass, the platform team owns the Gateway (listeners, certificates, which namespaces may attach), and each application team owns HTTPRoute objects in its own namespace.
# Owned by the platform team
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: public-edge
namespace: edge-system
spec:
gatewayClassName: envoy-edge
listeners:
- name: https
protocol: HTTPS
port: 443
hostname: "*.example.com"
tls:
certificateRefs: [{ name: wildcard-example-com }]
allowedRoutes:
namespaces:
from: Selector
selector: { matchLabels: { edge-access: "public" } }
---
# Owned by the orders team, in its own namespace and repo
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: orders-api
namespace: orders
spec:
parentRefs: [{ name: public-edge, namespace: edge-system }]
hostnames: ["api.example.com"]
rules:
- matches: [{ path: { type: PathPrefix, value: /v1/orders } }]
backendRefs:
- { name: orders-v41, port: 8080, weight: 95 }
- { name: orders-v42, port: 8080, weight: 5 } # canary
timeouts: { request: 3s }
The design properties this buys: the orders team can canary its own release without a ticket; it cannot claim /v1/payments (CI rejects overlapping prefixes across namespaces); and a malformed HTTPRoute is rejected for that route only rather than blocking the whole gateway's config.
Authentication at the Edge#
Validate tokens locally; never make a network call per request for authentication.
# Pseudocode for the edge authn filter
def authenticate(req):
tok = parse_bearer(req.headers.get("authorization"))
if tok is None:
return reject(401)
key = jwks_cache.get(tok.header.kid) # refreshed every 10 min in background
if key is None:
jwks_cache.refresh_async() # rotated key? refresh once, rate-limited
return reject(401)
if not verify_sig(tok, key) or tok.exp < now() or tok.aud != THIS_API:
return reject(401)
req.headers["x-auth-subject"] = tok.sub # strip any client-supplied copies first
req.headers["x-auth-scopes"] = ",".join(tok.scopes)
return next_filter(req)
Two details that interviewers probe:
- Strip inbound identity headers. If services trust
x-auth-subject, the gateway must delete any client-supplied version before setting it — or any client can impersonate any user. - Revocation. Locally validated JWTs can't be revoked before expiry. Keep access-token lifetimes short (5–15 min) and, if instant revocation matters, check a small denylist (bloom filter or in-memory set pushed from the control plane) rather than calling the identity service per request.
When you need external decisions (Envoy ext_authz, Kong plugins, AWS Lambda authorizers), cache the decision by (token, route) for 30–300s and set a deadline of ~20–50ms with an explicit failure mode.
Extensibility — Code in Every Request's Path#
| Product | Extension Mechanism | Risk |
|---|---|---|
| Envoy | C++ filters (compiled in), Lua, Wasm, ext_proc / ext_authz (external gRPC) | Wasm/Lua bugs crash or slow every worker; ext calls add a network hop |
| Kong | Lua plugins (OpenResty); Go/Python/JS plugin servers | Plugin ordering and blocking I/O |
| NGINX | C modules, njs (JavaScript), Lua via OpenResty | Custom modules tie you to build pipelines |
| AWS API Gateway | Lambda authorizers, mapping templates (VTL) | Authorizer cold starts and per-invocation cost |
| Apigee / Azure APIM | Policy XML/JS executed per request | Policies grow into business logic |
Rule of thumb: any extension that runs on every request needs the same review, testing, canarying, and on-call ownership as the gateway itself. A plugin written by a product team for its own convenience is a liability on everyone's hot path.
The Tunable Tradeoff — Centralization vs. Blast Radius#
Every policy you move into the gateway buys consistency and costs coupling and shared risk. Two quantities make this concrete.
Added latency is additive across filters:
gateway_overhead ≈ TLS (amortized) + route match + Σ local filters + Σ remote calls
typical thin Envoy edge: 0.5–1.5ms p50, 2–5ms p99
with global rate-limit RPC (+1–2ms) and ext_authz (+2–10ms): 4–15ms p99
Availability is multiplicative across synchronous dependencies:
A(path) = A(gateway) × A(authz) × A(rate_limit_store) × A(service)
0.9999 × 0.9995 × 0.999 × 0.9995 ≈ 0.9979
-> ~18 hours/year of failed requests caused by the front door's dependencies
fail-open on rate limit + cached authz: ≈ 0.9989-0.9994
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Thin gateway (TLS, authn, coarse limits, routing) | Fast, rarely changes, small blast radius | Services re-implement some policy (fine-grained authz, validation) | Service teams — via a shared library |
| Thick gateway (aggregation, transforms, business rules) | One place to change behavior; clients simple | Gateway team becomes a bottleneck; changes risk every route | Every product team waiting on gateway changes; the on-call |
| Fail-closed on dependencies (authz, rate limits) | No unauthorized or unmetered traffic | Dependency outage = full outage | Every user, during the dependency's bad hour |
| Fail-open with bounds | Front door stays up | Brief window of unenforced limits | Security/finance accept a bounded risk, in writing |
| Fail-static (keep last-known keys/decisions) | Best of both for authn | Stale revocations for the cache window | Security accepts "revocation within N minutes" |
🎯 Staff Move: "Authentication is fail-static — we keep the last-known signing keys if the identity provider is down, so revocation can lag by up to 15 minutes, which security has agreed to. Abuse rate limits fail-open with a local per-pod limit as a backstop. Only the partner billing quota fails closed, because over-serving it is a contractual problem."
Anti-Patterns — What Kills Gateway Deployments#
1. The Gateway as an Enterprise Service Bus#
It starts with one header transform, then a response field rename for an old client, then aggregation for the home screen, then a pricing rule "just for this campaign." Two years later the gateway has thousands of lines of per-route logic, one team understands it, and every product launch has a gateway ticket on its critical path. Prevention: a written placement policy and a review for any per-route code.
2. One Global Config, Pushed All at Once#
A single YAML file for 2,000 routes, edited by 40 teams, pushed to every proxy simultaneously. One typo — or one regex that matches everything — and the whole edge is down in seconds. Prevention: per-team route objects, CI conflict detection, staged rollout to a canary slice, automated rollback on 5xx rate.
3. Synchronous Calls Without Deadlines#
An authz or API-key lookup per request with the HTTP client's default 30-second timeout. When that dependency slows, every gateway worker fills with waiting requests and the gateway stops accepting new connections. Prevention: 20–50ms deadlines, cached decisions, declared failure mode per dependency.
4. Retry Amplification#
Client SDK retries 3×, gateway retries 3×, service retries its database 3× — up to 27× load during the exact incident where the backend is struggling. Prevention: gateway retries only idempotent methods, once, within a retry budget (~10% of requests); see Circuit Breakers.
5. Regex-Heavy Routing at Scale#
Hundreds of ordered regex routes evaluated linearly per request: CPU grows with route count, and one catastrophic-backtracking pattern can pin a core. Prevention: prefix and exact matches by default, host-based partitioning, regex only with an engine that guarantees linear time (Envoy uses RE2).
6. A Gateway Per Team#
Every team runs its own Kong or NGINX "for autonomy." Now there are 25 front doors with 25 TLS configurations, 25 ways of validating tokens, and 25 places a CVE patch must land. Prevention: one data plane, shared control plane, per-team route ownership — autonomy in config, not in infrastructure.
7. Buffering Large Bodies#
Uploads and large downloads proxied and buffered through the gateway. Memory spikes, slow clients hold workers, and the gateway's capacity plan is now driven by file sizes. Prevention: stream, set body-size limits (e.g., 10MB default), and route large uploads directly to object storage via pre-signed URLs — see Handling Large Blobs.
8. Treating the Managed Gateway's Limits as Someone Else's Problem#
Building on AWS API Gateway and discovering in production the regional account throttle (10,000 rps steady with 5,000 burst by default), the 29-second default integration timeout, or the 10MB payload cap. Prevention: list the provider's quotas in the design doc and request increases before launch.
The Technology Landscape — Head-to-Head Comparison#
The market sorts into three families. Mixing them up is the most common product-selection mistake.
| Family | Products | Core Strength | Typical Buyer |
|---|---|---|---|
| Programmable proxies (data planes) | Envoy, NGINX / OpenResty, HAProxy, Traefik | Raw performance, fine control, open source | Platform/infra teams running their own edge and ingress |
| API-management platforms | Kong, Apigee, Azure API Management, Tyk, Gravitee | Keys, quotas, portals, analytics, monetization, policy UI | Teams exposing APIs as a product to partners or developers |
| Cloud-managed gateways | AWS API Gateway, Google API Gateway / Cloud Endpoints, Azure APIM consumption tier, Cloudflare API Shield | Zero servers, pay per request, cloud-native integrations | Serverless backends, small teams, single-cloud shops |
Product Comparison#
| Product | Data Plane | Config / Control | Extensibility | Strengths | Weaknesses | Cost Shape |
|---|---|---|---|---|---|---|
| Envoy (+ Envoy Gateway, Istio ingress, Gloo, Emissary) | C++, event loop per core | xDS APIs — needs a control plane | C++, Lua, Wasm, ext_authz/ext_proc | Best-in-class observability, HTTP/2 + gRPC, dynamic config, retries/outlier detection | Complex config; you must pick or build a control plane | Compute + people |
| NGINX / NGINX Plus | C, event loop per worker | Files + reload; Plus adds API | C modules, njs | Ubiquitous, efficient, well understood | Dynamic config is a commercial or Lua story; reloads at high churn | Compute (+ Plus license) |
| Kong Gateway | NGINX/OpenResty + Lua | Admin API (Postgres) or DB-less declarative | Lua, Go/Python/JS plugin servers | Large plugin ecosystem; API-management features; hybrid CP/DP mode | Plugin performance varies; enterprise features licensed | Compute + license at scale |
| HAProxy | C, multi-threaded | Config + Runtime / Data Plane API | Lua, SPOE agents | Extremely high performance at L4/L7; mature | Fewer API-management features | Compute (+ enterprise license) |
| Traefik | Go | Auto-discovery from K8s/Docker labels | Middlewares, plugins | Very easy for small/medium K8s setups | Less control at large scale; label-driven config spreads ownership | Compute |
| AWS API Gateway | Managed | Console/IaC, stages | Lambda authorizers, VTL mapping | No servers; usage plans and keys built in; IAM/Cognito integration | Per-request price at scale; quotas; added latency (~10–30ms typical); AWS-only | REST ~$3.50/M, HTTP ~$1.00/M requests (first tier) |
| Apigee | Managed (Google) | Proxies + policies | JS/Java policies | Full API-product lifecycle, analytics, monetization | Expensive; heavy for internal traffic | Subscription / per-call |
| Azure API Management | Managed (+ self-hosted gateway) | Policies (XML) | Policy expressions | Strong for Azure shops; portal | Tier limits; policy sprawl | Tier-based |
| Spring Cloud Gateway / Zuul | JVM (Netty) | Code + config | Java filters | Team already in JVM; app-level routing | Gateway logic drifts into app code | Compute |
How to choose in one breath: your own edge and Kubernetes ingress → an Envoy-based implementation of the Gateway API (or NGINX/HAProxy if your team already runs them well). Public or partner API that needs keys, quotas and a portal → an API-management product in front of that ingress. Small serverless backend on one cloud → the cloud's managed gateway, until request volume makes per-request pricing hurt.
🎯 Staff Insight: The Kubernetes
IngressAPI is frozen; the Gateway API (GA at v1.0 in October 2023) is its successor and is implemented by Envoy Gateway, Istio, Kong, NGINX, HAProxy, Traefik and the major clouds. The community ingress-nginx controller announced its retirement in late 2025. Choosing a Gateway API–conformant implementation makes the vendor a two-way door.
Patterns#
Pattern 1: Two-Tier Front Door (Edge + Internal Ingress)#
The edge tier is global, rarely changes, and owned by the platform team. Cluster ingresses are per region or cell and carry team-owned routes. A bad route push affects one cell's ingress, not the global edge. API management sits in front of the edge only for partner traffic.
Pattern 2: Federated Gateways With a Shared Control Plane#
Large orgs run one gateway fleet per domain (payments, commerce, media) using the same data plane and control-plane tooling. Each domain's fleet has its own capacity, change calendar and on-call; the platform team owns the build, the base policies, and upgrades. This is the answer to "one gateway is a bottleneck" that doesn't become "25 different gateways."
Pattern 3: Backend-for-Frontend Behind the Gateway#
Aggregation, client-specific shapes and partial-failure rules live in a BFF that the client team deploys on its own schedule. The gateway stays generic. GraphQL federation is a variant of this pattern — the router is a BFF, not a gateway concern.
Pattern 4: Edge Gateway + Service Mesh#
The gateway handles north-south traffic (client → platform); a mesh (Istio, Linkerd, or an Envoy-based sidecarless mode) handles east-west (service → service) with mTLS, retries and telemetry. Using the same data plane (Envoy) for both means one set of skills, metrics and failure modes. Don't route internal service-to-service traffic back out through the edge gateway — it adds a hop, a shared bottleneck and confusing blast radius.
Pattern 5: Managed Gateway for Serverless#
AWS API Gateway (HTTP API flavor where features allow) → Lambda. No servers, scale-to-zero, built-in authorizers. Watch: per-request cost above ~1–2B requests/month, cold starts in authorizers, and quota limits. The migration path out is a self-run Envoy/ALB layer pointing at the same functions or containers.
Pattern 6: Progressive Delivery at the Gateway#
Weighted routing (1% → 5% → 25% → 100%), header-based routing for internal testers, and shadow traffic (mirror a copy of requests to a new version, discard responses). Pair weights with automated analysis of error rate and latency per backend version so rollback is automatic. Shadowing writes is dangerous — mirror only idempotent reads, or point mirrors at a sandbox.
Pattern 7: Strangler Migration#
Put the gateway in front of a monolith, then move routes one prefix at a time to new services. The gateway is the switch: /v1/orders/* → new service, everything else → monolith. Rollback is a route change. This is one of the strongest justifications for introducing a gateway in an existing system.
Scaling#
The Numbers#
| Metric | Planning Value | Notes |
|---|---|---|
| Envoy / NGINX proxy throughput | ~5–20K rps per core for small HTTP/1.1 requests with keep-alive | Less with heavy filters, Lua/Wasm, large bodies |
| Added latency, thin self-run gateway | ~0.5–1.5ms p50, 2–5ms p99 | Mostly TLS amortization and filter work |
| Added latency, managed cloud gateway | ~10–30ms typical | Varies by provider, region, and authorizer |
| Full TLS handshakes | ~1–2K/s per core (RSA-2048); several × more with ECDSA P-256 | Session resumption and keep-alive are the main levers |
| JWT verify (RS256) | ~50–200µs CPU | Cheap; cache parsed keys |
| Global rate-limit call | +0.5–2ms in-region | Local token buckets avoid it for most traffic |
| Config propagation (xDS) | ~1–10s fleet-wide | Delta xDS for 10K+ routes |
| Connection draining | Request timeout + longest stream | Max connection age 10–30 min for gRPC/WebSocket |
| AWS API Gateway defaults | 10,000 rps steady / 5,000 burst per account per region; 29s integration timeout; 10MB payload | Request increases before launch |
| Managed price (AWS, first tier) | REST ~$3.50 per million; HTTP ~$1.00 per million | Plus data transfer |
Sizing Example#
Peak: 120K rps, 70% on keep-alive connections, TLS 1.3, avg 4KB response
Per core: ~6K rps with authn + local rate limit + access logging
Cores: 120K ÷ 6K = 20 cores busy -> target 50% utilization = 40 cores
Instances: 10 × 4 vCPU per region, + N+2 for AZ loss and deploys = 12
Handshakes: 30% new connections ≈ 36K handshakes/s spike on reconnect storms
-> ECDSA certs + resumption, or 36K ÷ ~5K/core = 7+ cores just for TLS
Managed alt: avg ~50K rps × 2.6M s/month ≈ 130B requests/month
× ~$0.90–1.00 per million ≈ $120–130K/month (HTTP API pricing)
vs self-run ≈ $3–6K/month compute per region + people
At high volume, per-request pricing is the decisive factor; at low volume, people are. The crossover for many teams lands somewhere around a few hundred million to a couple of billion requests per month.
Scaling Moves in Order#
- Remove work from the hot path — cache authz decisions, local rate limiting, kill unneeded filters, drop full request-body logging.
- Keep connections warm — client keep-alive, TLS resumption, HTTP/2 to upstreams.
- Scale horizontally — stateless proxies behind L4; autoscale on CPU and active connections, not just rps.
- Partition the fleet — per-region and per-cell gateway fleets so one noisy tenant or bad config has a bounded blast radius.
- Push work to the edge — WAF, bot filtering, and cacheable responses at the CDN before they reach your gateway.
Failure Modes & Recovery#
1. The Bad Config Push#
- Symptom: 5xx or 404 rate jumps across many routes within seconds of a deploy; sometimes every route.
- Root cause: malformed or over-broad route (
PathPrefix: /attached to the wrong backend), conflicting routes, a plugin config error, or a control-plane bug sending a partial snapshot. - Detection:
gateway.responses{code=5xx|404}rate by route, correlated withconfig.versionchange; synthetic probes on top-20 routes every 10s. - Fix: automated rollback to last-known-good config version; data plane keeps serving old config if the new one is rejected.
- Prevention: CI validation (schema, overlap detection, "does this route match the health probes of other teams?"), staged rollout to 1% → 10% → 100% of the fleet with automated analysis, per-team route isolation.
t=0 Team merges route change; control plane pushes to all 400 proxies.
t=+4s Prefix "/" now matches before "/v1/payments"; payments traffic hits the wrong backend.
t=+20s 404 rate on /v1/payments 0% -> 100%. Checkout fails globally.
t=+90s Synthetic probe alert fires; on-call correlates with config version 8812.
t=+6min Manual rollback to 8811. Six minutes of failed checkouts.
Fix: CI rejects routes that shadow another namespace's prefix; staged rollout
would have contained it to 1% of traffic for ~30s before auto-rollback.
Owner: Platform team owns the pipeline; the route's team owns the change.
2. Synchronous Dependency Brownout (Authz or Rate-Limit Store)#
- Symptom: gateway p99 climbs from 5ms to seconds; active connections and pending requests per worker climb; eventually connection refusals.
- Root cause: the external authz service or rate-limit Redis slows down; gateway workers wait on it.
- Detection:
ext_authz.latency_p99,ratelimit.call_errors,gateway.downstream_rq_activeper worker. - Fix: enforce 20–50ms deadlines; switch to declared failure mode (fail-open for abuse limits with local backstop, fail-static for authz decisions in cache).
- Prevention: decision caching; circuit breaker on the dependency; game-day the dependency's outage.
3. Retry Storm Through the Gateway#
- Symptom: a backend slowing down sees its request rate rise 3–10× instead of fall; recovery stalls.
- Root cause: retries at client, gateway, and service layers multiply; non-idempotent requests retried.
- Detection:
upstream_rq_retryrate vsupstream_rq_total; retry ratio > 10%. - Fix: cut gateway retries to zero for the affected cluster; enable retry budgets.
- Prevention: retry only idempotent methods, once, with a budget (e.g., ≤ 10% of active requests); outlier detection to eject bad hosts rather than retrying into them.
4. Certificate Expiry#
- Symptom: TLS handshake failures for all clients of a hostname at a precise moment; mobile apps with pinned certs fail hardest.
- Root cause: manual renewal missed, or automated renewal failed silently (DNS challenge broken, secret not reloaded).
- Detection:
tls.cert_expiry_daysper hostname, alert at 21 and 7 days; synthetic TLS checks from outside. - Fix: emergency renewal and hot reload (Envoy SDS reloads without restart).
- Prevention: automated issuance (ACME or a cloud certificate manager), expiry alerts owned by the platform team, no manual certs in the edge.
5. Connection Exhaustion and Slow Upstreams#
- Symptom: one slow upstream causes errors for unrelated routes on the same gateway.
- Root cause: shared worker or connection-pool resources; slow responses hold connections; no per-cluster concurrency limits.
- Detection:
upstream_cx_activeandupstream_rq_pending_overflowper cluster; worker event-loop lag. - Fix: per-cluster circuit breakers (max connections, max pending, max requests), shorter timeouts on the slow cluster.
- Prevention: per-upstream limits as defaults in the route template; bulkheading by fleet partitioning for critical paths.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Bad config push | 5xx/404 by route vs config.version | Up to every route | Auto-rollback; staged rollout | Platform (pipeline) + route-owning team |
| Authz / rate-limit dependency brownout | Dependency latency, active rq per worker | All authenticated traffic | Deadlines, declared fail mode, caching | Platform + identity team |
| Retry storm | Retry ratio > 10% | Struggling backend → its callers | Retry budgets, outlier ejection | Platform sets defaults; service teams own retries below |
| Certificate expiry | Days-to-expiry per hostname | Every client of the hostname | Automated renewal, SDS hot reload | Platform team |
| Slow upstream contaminating others | Pending overflow per cluster | Routes sharing workers/fleet | Per-cluster circuit breakers, bulkheads | Platform defaults; service team fixes upstream |
| DDoS / abusive client | rps per client key/IP, WAF signals | Edge capacity | CDN/WAF filtering, per-client limits | Security + platform |
When to Use vs. Alternatives#
| Situation | Pick | Why |
|---|---|---|
| Kubernetes platform, many teams, own the edge | Envoy-based Gateway API implementation | Dynamic config, role-separated ownership, strong telemetry |
| Existing NGINX/HAProxy expertise, modest route churn | NGINX / HAProxy | Proven, efficient; don't migrate for fashion |
| Public/partner API product with keys, quotas, portal, analytics | Kong / Apigee / Azure APIM | Buy the API-management features |
| Serverless on one cloud, < ~1B requests/month | Cloud-managed gateway | No servers; cost acceptable at this volume |
| Simple app, one service, one team | Cloud L7 load balancer (ALB, GCLB) | A gateway adds a tier you don't need |
| Service-to-service policy (mTLS, retries) | Service mesh | Not a gateway's job |
| Aggregating many calls for one client | BFF / GraphQL router | Product logic, owned by product teams |
When NOT to Add an API Gateway#
- One backend, one client. A load balancer with TLS and a WAF is enough; a gateway is an extra hop, an extra on-call, and an extra config pipeline.
- To fix internal service-to-service problems. Hairpinning east-west traffic through the edge gateway creates a bottleneck and blurs blast radius. Use a mesh or client libraries.
- As the place to "quickly" add business logic. If the reason is "the service team is too slow," the gateway will become the slow team.
- When the real need is a CDN. Caching, DDoS absorption, and static content belong at the edge network, ahead of your gateway.
Operational Concerns#
What the On-Call Actually Does#
- Correlates error spikes with config versions and rolls back — this is the most common gateway page.
- Distinguishes "gateway is broken" from "a backend is broken and the gateway is the messenger" using per-cluster upstream metrics. Most 5xx at the gateway are upstream 5xx.
- Watches certificate expiry dashboards and renewal job health.
- Handles abusive clients: identifies by key/IP/ASN, applies targeted limits or blocks at the CDN/WAF.
- Runs data-plane upgrades: canary instances first, then by AZ, with connection draining.
Key Metrics & Alerts#
| Metric | Alert Threshold | Why |
|---|---|---|
gateway.5xx_rate split by gateway-generated vs upstream | Gateway-generated > 0.1% for 5 min | Separates "us" from "them" |
| Added latency (gateway time minus upstream time) p99 | > 10ms for 10 min | Filter or dependency regression |
config.push.rejected / version skew across fleet | Any rejection; skew > 5 min | Control-plane health |
| Retry ratio per upstream cluster | > 10% | Storm in progress |
| Pending overflow / circuit-breaker opens | > 0 sustained | Upstream saturation |
| Cert days-to-expiry | < 21 days warn, < 7 days page | Avoidable outage |
| CPU per instance, active connections | > 60% / approaching limits | Capacity |
| Rate-limit decisions by client | Sudden top-client shift | Abuse or broken client release |
Config Deployment Pipeline#
Upgrades#
Data-plane upgrades are frequent (security fixes in proxies are common) and should be boring: immutable images, canary instances, AZ-by-AZ rollout, connection draining with a bounded max connection age. Keep the control plane and data plane within the version skew the vendor supports, and never upgrade both in the same change.
Interview Application — Staff-Level Plays#
Which Case Studies Use API Gateways#
| Case Study | Where the Gateway Fits | The Staff Detail |
|---|---|---|
| API Gateway | The design problem itself | Placement policy, config safety, cells, multi-tenancy |
| Rate Limiting | Enforcement point for client limits | Local buckets + global limits; fail-open for abuse protection |
| Load Balancer | L4 in front, L7 gateway behind | Who balances long-lived gRPC connections |
| Circuit Breakers | Retry budgets, outlier ejection at the edge | One retry layer; per-cluster limits |
| Service Discovery | Gateway as an xDS/endpoint consumer | Stale endpoint behavior when discovery is down |
| CDN & Edge Caching | Layer in front of the gateway | What to push to the CDN before it reaches you |
| Auto-Scaling & Capacity | Gateway fleet sizing | Scale on CPU and connections, not just rps |
| Multi-Region Active-Active | Regional gateway fleets, global steering | Failover is a routing decision at DNS/anycast, not in the gateway |
Every System Design Question Has a Gateway Moment#
- Any public API: "The gateway terminates TLS, validates the JWT locally and applies per-client limits. Authorization on the actual resource happens in the service."
- Mobile app with many screens: "Aggregation goes in a mobile BFF behind the gateway — the mobile team owns it and ships on their own cadence."
- Migration off a monolith: "The gateway is my strangler switch — I move one path prefix at a time, and rollback is a route change."
- Partner/public developer API: "I'll buy API management for keys, usage plans and the portal, and put it in front of the same ingress our own apps use."
What Interviewers Probe#
| After You Say... | They Will Ask... | What They're Evaluating |
|---|---|---|
| "The gateway handles auth" | "Authentication or authorization? What happens when the identity provider is down?" | Placement and failure-mode judgment |
| "We'll use Kong / Envoy" | "Why that one? What does it cost to switch?" | Product reasoning vs brand preference |
| "Rate limiting at the gateway" | "Per pod or global? What if Redis is slow?" | State, latency, fail-open/closed |
| "The gateway aggregates the responses" | "Who changes it when the home screen changes?" | ESB trap awareness |
| "We run 20 instances for HA" | "What's the most likely way the gateway takes down the site?" | Config as the outage vector |
| "Retries at the gateway" | "And in the client? And in the service?" | Retry amplification math |
| "AWS API Gateway" | "What are its limits, and what does it cost at 50K rps?" | Managed-service due diligence |
Common Interview Mistakes#
| What Candidates Say | What Interviewers Hear | What Staff Engineers Say |
|---|---|---|
| "The gateway handles security" | Undefined boundary | "The gateway authenticates and checks scopes; services authorize against their own data" |
| "Put the aggregation in the gateway to save round trips" | Will create a bottleneck team | "A BFF behind the gateway, owned by the client team" |
| "The gateway is stateless, so it's highly available" | Ignores the config and dependency risk | "Hardware isn't the risk; config pushes and synchronous dependencies are — here's how I contain each" |
| "Kong because it has the most plugins" | Feature-list selection | "Envoy data plane for our ingress, an API-management product only for the partner API" |
| "We'll check the token with the auth service on every request" | Adds a dependency and 5–20ms to everything | "Local JWT validation with cached keys; short-lived tokens; a pushed denylist for urgent revocation" |
| "One gateway for everything, internal and external" | Hairpinning, mixed blast radius | "North-south at the gateway, east-west in the mesh" |
L5 vs L6 vs L7 Responses#
| Scenario | L5 Answer | L6 / Staff Answer | L7 / Principal Answer |
|---|---|---|---|
| "Pick a gateway" | "Kong — it's popular and has plugins" | "Envoy-based Gateway API implementation for ingress; managed API management for partners; criteria: config model, extensibility, ops cost" | "Standardize on one data plane for edge, ingress and mesh; Gateway API as the portability contract so the vendor is a two-way door" |
| "A route change took the site down" | "Add code review for config" | "Per-team routes, CI overlap checks, staged rollout with auto-rollback" | "Split the fleet into cells, and make 'gateway config' a change class with the same SLO accounting as code deploys" |
| "Teams want custom logic at the gateway" | "Write a plugin for them" | "Does it need domain data? Then it's a service or BFF. Generic? Then a reviewed, platform-owned filter" | "Publish the placement policy and an exception process; track plugin count as a health metric" |
| "Partner API launch" | "Expose the same gateway with API keys" | "API-management layer for keys, quotas, metering, portal; versioning policy" | "Treat the partner API as a product with a multi-year deprecation policy and a revenue owner" |
| "Gateway cost is rising" | "Use bigger instances" | "Profile filters, cache authz, local limits, keep-alive/TLS resumption" | "Price managed vs self-run at our volume and headcount; decide per traffic class" |
The Staff API Gateway Checklist#
- "The gateway is thin: TLS, authentication, scope checks, client rate limits, routing, telemetry."
- "Authorization on resources lives in the services; aggregation lives in BFFs owned by client teams."
- "Every synchronous dependency has a 20–50ms deadline and a declared fail-open, fail-closed or fail-static mode."
- "Retries happen once, at the gateway, for idempotent requests only, inside a retry budget."
- "Routes are owned per team, validated in CI, and rolled out to a canary slice with automatic rollback."
- "The platform team owns the data plane, base policies, certificates and upgrades; the page goes to them for gateway-generated errors and to the service team for upstream errors."
🎯 Staff Insight: Say what the gateway won't do. "No business logic, no response composition, no fine-grained authorization, no internal east-west traffic." The refusals are what convince an interviewer you've run one.
Evaluation Rubric#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Placement | Lists gateway features | Sorts concerns by domain knowledge; thin gateway, BFFs, mesh | Writes the placement policy and exception process |
| Product choice | Picks a familiar product | Matches product family to intent; knows managed-service limits and prices | Standardizes data plane org-wide; keeps vendor reversible |
| Failure | Redundant instances | Config safety pipeline, dependency deadlines, retry budgets | Cell-based fleets; config changes under SLO accounting |
| Ownership | "Platform owns it" | Platform owns guardrails; teams own routes; paging split by error source | Redraws team contracts across edge, ingress and mesh |
| Cost | "Open source is free" | Sizes cores and instances; compares per-request pricing | Prices build vs buy vs managed with headcount at 3 scales |
Strong Hire Signals
| Signal | What It Sounds Like |
|---|---|
| Thin-gateway discipline | "If it needs to know what an order is, it's not a gateway concern." |
| Config as the main risk | "Our gateway outages will be pushes, not crashes." |
| Declared failure modes | "Authn fail-static, abuse limits fail-open, partner quota fail-closed." |
| Ownership split | "Gateway-generated 5xx page platform; upstream 5xx page the service." |
| Managed-service realism | "At 50K rps average that's ~$120K a month in request fees." |
Lean No-Hire Signals
| Signal | Why It Misses the Bar |
|---|---|
| Gateway does aggregation and business rules | ESB in the making |
| Auth service called per request with no cache or deadline | Gateway availability capped by a dependency |
| Retries everywhere | Retry storms |
| No config rollout story | The most likely outage is unaddressed |
| Chooses by plugin count or popularity | No criteria |
Common False Positives
- Knowing every Envoy filter name ≠ knowing what belongs in the gateway.
- A long feature list for the gateway box ≠ a design — it's usually a sign of the ESB trap.
- "We use a service mesh" ≠ having a north-south strategy.
The Principal Lens#
Why L7 Sees This Problem Differently#
At Staff level, the gateway is a component to configure safely. At Principal level it is the place where the org's policies become enforceable — and therefore the place every team will try to put its problems. Security wants WAF rules, finance wants metering, product wants experiments, legal wants regional routing, and each request is reasonable. The L7 job is not choosing Envoy over Kong; it is deciding what the org centralizes at the front door, what it federates, and how a policy gets in or out of the hot path — and pricing the shared risk each addition creates.
The Org-Level Fault Line#
One central gateway platform vs. domain-owned gateways vs. no shared gateway (every team behind its own load balancer).
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| One central gateway, one team | Consistent security and telemetry; one patch path | Gateway team becomes a bottleneck; shared blast radius | Product teams waiting; everyone during the bad push |
| Federated: shared data plane + control plane, domain-owned fleets and routes | Isolation by domain; teams move independently; consistent base policy | Platform must support multi-fleet tooling; drift between domains | Platform headcount (~1 FTE per ~10 fleets with automation) |
| Each team its own gateway product | Maximum autonomy | 20 security postures, 20 CVE patch paths, no common telemetry | Security, incident command |
| No gateway; cloud LBs + libraries | Minimal infra | Policy re-implemented in every service and language | Every service team, forever |
The Principal default: federated — one data plane (Envoy-class), one control plane and pipeline, Gateway API as the contract, domain-owned fleets and routes, and a short list of org-mandated filters (authn, rate-limit hooks, telemetry, WAF integration). API management is bought, scoped to external API products only.
Cost Model#
Assumptions: self-run on cloud VMs at ~$0.04/vCPU-hr; ~6K rps per core with authn and local limits; 2 regions with N+2 headroom; managed pricing at ~$1.00 per million (HTTP-API-class) or ~$3.50 (REST-API-class); API-management licenses vary widely and are shown as ranges; loaded engineer ~$250K/yr. Directional only.
| Scale | Traffic | Self-Run Compute/month | Self-Run People | Managed Gateway/month | API-Management License (partner APIs) |
|---|---|---|---|---|---|
| Startup | 500 rps avg (~1.3B req/month) | ~$500 | 0.25 FTE (~$5K) | ~$1.3–4.5K | Usually unnecessary |
| Growth | 20K rps avg (~52B req/month) | ~$3–8K (incl. L4 LBs) | 1–2 FTE (~$20–40K) | ~$45–150K | ~$5–30K |
| Enterprise | 500K rps avg (~1.3T req/month), 4 regions | ~$100–200K | 6–10 FTE platform | Not viable at list price | ~$50–300K+ |
At startup scale, the managed gateway is cheapest once people are counted. At growth scale, request fees overtake a small platform team. At enterprise scale, the dominant costs are the platform team, cross-AZ traffic between gateway and services, and the cost of outages — which is why cells pay for themselves.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse |
|---|---|---|
| Public API URL structure, auth scheme, and versioning promise | One-way | External clients you don't control; multi-year deprecation |
| Putting business logic into gateway plugins | One-way in practice | Untangling logic from hundreds of routes; owners long gone |
| Identity headers services trust from the gateway | One-way-ish | Every service's trust model changes |
| Proprietary policy language (vendor XML/VTL) for many routes | One-way-ish | Rewrite every policy to migrate |
| Data-plane product behind Gateway API | Two-way | Conformant implementations are swappable route-by-route |
| Managed vs self-run for the internal ingress | Two-way | Route migration behind DNS weights |
| Rate-limit values, timeouts, retry policies | Two-way | Config |
🧭 Principal Move: "I'll let domains choose their fleet sizes, their canary cadence, and even their gateway implementation as long as it speaks Gateway API. I won't let anyone put business logic in a plugin or change what identity headers mean — those are the doors we can't walk back through."
The Standard I'd Write#
RFC-EDGE-003: What Belongs at the API Gateway
Scope: All north-south HTTP/gRPC traffic entering production.
MUST
1. Enter through a platform-operated gateway fleet; no service exposes a public
endpoint directly.
2. Authenticate at the gateway (local token validation); services MUST NOT trust
identity headers from any other source; the gateway strips inbound copies.
3. Declare routes as Gateway API resources in the owning team's repo; CI rejects
overlapping prefixes across teams.
4. Roll out config changes to <= 5% of the fleet first with automated rollback.
5. Every synchronous gateway dependency declares a deadline (<= 50ms) and a
failure mode (fail-open / fail-closed / fail-static), approved by security.
6. No business logic, response composition, or resource-level authorization in
gateway filters or plugins.
SHOULD
7. Retry only idempotent methods, once, within a 10% retry budget.
8. Use BFFs owned by client teams for aggregation.
9. Keep partner/public API products behind an API-management layer with a
published deprecation policy (>= 12 months).
Exceptions: Requested via the edge platform team; approved by the platform lead
and the requesting director; reviewed every two quarters; custom filters require
an owning team on the gateway on-call rotation.
Success metrics: zero Sev1s caused by config pushes; gateway-added p99 < 10ms;
custom filter count flat or decreasing; route change lead time < 1 hour.
What I'd Tell the VP#
"Every customer request passes through our API gateway, so it is both our best place to enforce security and our biggest single point of failure. Two of our last three major outages were gateway configuration changes, not broken servers. I'm proposing three things: each team manages its own routes with automatic safety checks and gradual rollout, the gateway is split into independent sections so one mistake can't take everything down, and we publish a clear rule that product logic stays out of the gateway. That's about two engineers for two quarters. It reduces our most common cause of site-wide outages and stops the gateway team from becoming a bottleneck for every launch."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Governs placement, not products | "The policy on what enters the gateway matters more than which gateway." |
| Bounds shared blast radius | "Cells and staged config — a bad push should cost one cell for 30 seconds." |
| Keeps vendors reversible | "Gateway API is our contract; implementations are replaceable." |
| Prices the hot path | "Every filter is paid by every request — 50K rps × 2ms is 100 cores of waiting." |
| Knows when not to centralize | "Partner metering is centralized; internal service-to-service policy goes to the mesh, not the edge." |
Staff answers that L7 interviewers find insufficient:
- "Use canary rollouts for config" — without cells, so a successful bad push still reaches everyone eventually.
- "Keep the gateway thin" — as a preference, not as a published policy with an exception process.
- "Managed is simpler" — without pricing per-request fees at the org's actual volume.
In the Wild#
Lyft — Envoy#
Lyft built Envoy to give a polyglot microservice fleet consistent networking — retries, timeouts, circuit breaking, outlier detection, and observability — without re-implementing them in every language, and open-sourced it in 2016. It became a CNCF graduated project and the data plane under many gateways and meshes, including Istio, Envoy Gateway, and several commercial products. Its xDS APIs became a de facto standard for dynamic proxy configuration.
Staff insight: Envoy's justification was organizational — inconsistency across teams — not raw speed. Make the same argument for a gateway: it pays for itself by making policy uniform, so keep in it only what should be uniform.
Netflix — Zuul#
Netflix built Zuul as the front door for its streaming API traffic, using filters for routing, authentication, and resiliency, and later rewrote it as Zuul 2 on an asynchronous, non-blocking Netty core to handle very large numbers of persistent connections more efficiently. Netflix has written about using the gateway for regional traffic steering and for testing via traffic shaping and fault injection.
Staff insight: The move from blocking to non-blocking was driven by connection counts and resilience, not latency. Gateways fail by exhausting concurrency when something behind them slows — design for that first.
The Kubernetes Gateway API — Ownership as a Spec#
The Kubernetes SIG-Network community designed the Gateway API (v1.0 GA in October 2023) to replace the limited Ingress resource. Its central idea is role separation: infrastructure providers own GatewayClass, cluster operators own Gateway (listeners, certificates, which namespaces may attach routes), and application developers own HTTPRoute and similar resources. Many vendors implement it, which makes the data-plane choice largely portable.
Staff insight: The spec encodes the ownership model Staff candidates should describe anyway — platform owns the door, teams own their routes. Citing it shows you think about who can change what, not just what the box does.
Practice Drill#
Prompt: "A 300-engineer company has 60 services behind a single NGINX config file maintained by one infra team. Changes take three days, two outages last quarter were config mistakes, and the business is about to launch a public partner API with paid tiers. What do you do over the next two quarters?"
Staff Answer
Diagnose first: the problems are ownership and change safety, not NGINX's performance — so I won't sell a migration as a speed fix. Quarter 1: (1) adopt the Kubernetes Gateway API with an Envoy-based implementation for new ingress, run it alongside NGINX behind the same L4 load balancer, and migrate routes prefix by prefix (strangler pattern), starting with low-risk internal tools; (2) each team owns HTTPRoute resources in its own repo, CI enforces schema, rejects cross-team prefix overlaps, and runs synthetic probes; (3) config rolls to a 5% canary slice of proxies with automatic rollback if 5xx or 404 rates move for 5 minutes. Target: route change lead time from 3 days to under an hour, and config-caused Sev1s to zero. Gateway scope stays thin — TLS, local JWT validation with cached JWKS, scope checks, per-client local rate limits, routing, telemetry; fine-grained authorization and aggregation stay in services and BFFs. Quarter 2: the partner API goes behind a bought API-management layer (keys, usage plans, metering, developer portal) that forwards to the same Envoy ingress, so partner traffic gets identical security and telemetry. Paid-tier quota enforcement fails closed only for the metered dimension, with a 50ms deadline and a cached plan lookup; abuse limits fail open with a local backstop. Publish a 12-month deprecation policy before the first partner signs. Ownership: platform team owns data plane, pipeline, certificates and base policies and is paged for gateway-generated errors; service teams own routes and are paged for upstream errors; a partner-API product owner owns versioning and tiers. Metrics: gateway-added p99 < 10ms, config rollback count, route lead time, partner quota accuracy against billing.
Why this is L6:
- Names the real problem (ownership and change safety) and doesn't sell a tool migration as the fix.
- Migrates incrementally behind the same front door with rollback at every step.
- Keeps the gateway thin and puts the partner API's product features in a bought layer.
- Declares failure modes per dependency and splits paging by error source.
What L7 adds:
- Publishes the placement policy and exception process so the new gateway doesn't become the next 3-day bottleneck in a different form.
- Plans gateway cells by domain for year two, bounding the blast radius even for config pushes that pass canary.
- Prices it: request-fee and license costs for the partner layer against expected partner revenue, and platform headcount against the outage minutes it removes — presented to leadership as one investment case.
Quick Reference Card#
Job: thin front door — TLS, authn, scope check, client limits, routing, telemetry
Not its job: business logic, aggregation (BFF), resource authz, east-west traffic (mesh)
Families: proxies (Envoy, NGINX, HAProxy, Traefik) | API mgmt (Kong, Apigee, Azure APIM)
| cloud-managed (AWS API Gateway, Google API Gateway)
Standard: K8s Gateway API v1 (GA Oct 2023): GatewayClass / Gateway / HTTPRoute
Added latency: self-run thin ~0.5–1.5ms p50, 2–5ms p99; managed ~10–30ms
Throughput: ~5–20K rps per core (lighter filters = higher)
TLS: ~1–2K full RSA handshakes/s/core; prefer ECDSA + resumption + keep-alive
AuthN: local JWT verify (~50–200µs), JWKS refresh ~10 min, fail-static
Dependencies: deadline 20–50ms; declare fail-open / fail-closed / fail-static
Retries: idempotent only, once, budget <= 10%
Config: per-team routes, CI overlap checks, 5% canary, auto-rollback
AWS API GW: 10K rps / 5K burst per account-region default; 29s integration timeout;
10MB payload; REST ~$3.50/M, HTTP ~$1.00/M
Long connections: max connection age 10–30 min for gRPC/WebSockets
RED FLAGS
- Aggregation or business rules in the gateway
- Per-request call to the auth service with no cache or deadline
- One global config file pushed to the whole fleet at once
- Retries at client + gateway + service
- Large uploads buffered through the gateway
- A different gateway product per team