Hiring BarSupport

Design an API Gateway — Staff-Level Case Study

Case study67 min read6 diagrams

Technologies referenced in this case study: API Gateways — Technology Guide (Kong, Envoy, NGINX, AWS API Gateway) · Redis · ZooKeeper & etcd

Related case studies: Rate Limiting · Circuit Breakers · Service Discovery · Load Balancer · CDN & Edge Caching · API Design · Build vs Buy

Scope note: This case study is the design problem — what the gateway should own, how it fails, and who runs it. For product-by-product comparisons (Kong vs Envoy vs NGINX vs AWS API Gateway), see the technology guide; this page deliberately doesn't repeat them.

How to Use This Case Study#

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Drills 1, 3, 7
Targeted Study1–2 hrsExecutive Summary → Walkthrough → §3 Fault Lines → §4 Failure Modes → Deep Dives 1 and 4
Deep Dive3+ hrsEverything, including §11 Principal Lens and appendices
What is an API Gateway? — Why interviewers pick this topic

An API gateway is the single entry point between clients (browsers, mobile apps, partners) and a fleet of backend services. It terminates client connections, authenticates requests, applies rate limits and quotas, routes each request to the right service (sometimes the right version or region), and emits the logs and metrics that describe what customers actually experienced. Some gateways also transform protocols (REST ↔ gRPC), aggregate multiple backend calls into one response, or serve cached results.

Before vs After — the auth bug in 40 services:

Without a gateway:
t=0:      A JWT library vulnerability is disclosed (algorithm confusion)
t=+1d:    Security finds 40 services validating tokens with 6 different libraries
t=+2w:    31 services patched; 9 owned by teams busy with a launch
t=+5w:    Last service patched. Five weeks of exposure; a public API endpoint exploited in week 3

With a gateway:
t=0:      Same disclosure
t=+2h:    Gateway team patches the one validation path; staged rollout by region
t=+6h:    Fully deployed. Services receive a verified identity header over mTLS and never parse raw tokens
Result:   One change, one owner, six hours.

Why interviewers reach for this question: The gateway is where cross-cutting concerns concentrate — and where organizations make their most consequential "centralize or not" decisions. It tests whether you know what belongs at the edge (authentication, coarse rate limiting, routing) and what doesn't (business logic), whether you can keep a component that touches 100% of traffic from becoming a 100% outage, and whether you can design a platform that 50 teams change every day without breaking each other.

Mechanics Refresher: What a Gateway Does Per Request
StageWhat HappensTypical CostFailure If Wrong
Connection & TLSTerminate TLS/HTTP2/HTTP3, keep-aliveTLS handshake 1 RTT (TLS 1.3) for new connectionsCert expiry = total outage
Request normalizationSize limits, header sanitation, request IDµsSmuggling, oversized payloads reach services
AuthenticationValidate JWT locally (JWKS cache) or introspect opaque tokenLocal: ~0.05–0.5ms; introspection: 2–20ms network callBypass or total lockout
Authorization (coarse)Scopes, tenant, route-level policyµs–1msOver- or under-permissive
Rate limit / quotaPer key/tenant/route; local-first~0 local; 1–5ms if centralized per requestAbuse gets through or legit traffic blocked
RoutingMatch host/path/method/header → upstream cluster; canary splitsµs (trie/radix match)Traffic to wrong service/version
Upstream callDiscovery, LB, timeout, retry budgetNetwork + backendCascades (Circuit Breakers)
Response processingHeader filtering, compression, (optional) transformµs–msData leaks via internal headers
TelemetryAccess log, metrics, trace spanµs (async)Blind to customer experience

For most production systems: A thin, horizontally scaled, stateless proxy (usually Envoy- or NGINX-based, or a managed service) that owns TLS, authentication, coarse authorization, rate limiting, routing and telemetry — with configuration managed as code through a self-service pipeline, and no business logic.


Executive Summary

If you only read one section, read this.

What This Interview Actually Tests#

An API gateway is not a "reverse proxy with plugins" question. Everyone can draw a box between the client and the services.

This is a centralization and blast-radius question that tests:

  • Whether you can decide what to centralize (security, identity, traffic policy) and what to keep out (business logic, aggregation for one team)
  • Whether you design a component in 100% of request paths so that it is more reliable than anything behind it
  • Whether you can let 50 teams change routing daily without any of them being able to take down the other 49
  • Whether you understand the gateway as an organizational boundary: the contract between a platform team and every product team

The key insight: The gateway's value comes from being shared, and its risk comes from being shared. Every feature you add to it gets consistency across all services — and puts every service behind that feature's bugs. Staff engineers keep the gateway thin, keep its config changes staged, and keep asking "does this need to be in the one component everyone depends on?"

The L5 vs L6 vs L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws gateway → services; lists features (auth, rate limit, routing, caching, transforms)Asks who the clients are (first-party apps, partners, public developers) and what must be centralized vs owned by teamsAsks how many gateways the org already has and what the platform contract should be across all of them
ScopeAdds features as needed — aggregation, transforms, validationKeeps the gateway thin: identity, coarse policy, routing, telemetry; business logic and aggregation go to services/BFFsWrites the scope charter; says no to features that would make the gateway the new monolith
Reliability"Run multiple gateway instances behind a load balancer"Stateless data plane, fail-static config, dependency-free hot path (local JWT validation, local rate limits), cell/regional isolationTreats the gateway as the org's largest correlated failure domain; cells, staged config, error budget for the platform itself
AuthValidates tokens at the gatewayGateway authenticates and issues an internal identity; services still authorize (zero trust, mTLS)Owns the identity model across edge and internal (token exchange, workload identity), with security sign-off
Config changesGateway team edits routesSelf-service route manifests per team, validated and staged, with blast radius limited to the team's routesGovernance model: ownership of routes, review requirements, exception process, audit
Ownership"The platform team owns the gateway"Platform owns runtime and pipeline; product teams own their routes, quotas and API contractsDecides central vs federated gateways, funds the platform, sets deprecation paths for legacy gateways
Why "scope" separates levels

L5: Treats the gateway as a place to put useful things: request validation against schemas, response transformation, aggregation of three backend calls for the mobile home screen, a bit of feature-flag logic. Each addition is reasonable. Together they turn the gateway into a monolith that every team must change — and deploy — to ship features.

L6: "The gateway owns concerns that are the same for every service: TLS, authentication, coarse authorization, rate limiting, routing, and telemetry. Anything that's specific to one product — aggregation, response shaping, business validation — belongs in a service or a backend-for-frontend that the product team owns and deploys independently."

L7: Writes that boundary down as a charter and enforces it through the plugin/extension review process. Knows the failure pattern — a gateway team with a 6-week backlog of feature requests from product teams is the symptom that the boundary was lost.

Why "reliability" separates levels

L5: "Run N instances behind a load balancer, autoscale on CPU." That handles instance failure. It does nothing about the failure modes that actually take gateways down: a bad config push to every instance at once, a hot-path dependency (auth service, rate-limit Redis) that becomes slow, a certificate that expires everywhere simultaneously.

L6: "The gateway's hot path has no synchronous dependencies it can't live without: JWTs are validated locally against cached signing keys, rate limits are local-first, routing config is in memory and fails static. Config rolls out by region and cell with automatic rollback. The gateway fails in pieces, never all at once."

L7: Prices the gateway's own availability target — it must be higher than any service behind it — and budgets for cells, multi-region, and a platform team whose on-call is the most important in the company.

Why "config changes" separates levels

L5: The gateway team maintains the route table. Product teams file tickets. It works for 10 services and becomes the bottleneck at 100.

L6: Each team owns a route manifest for its hostnames/paths in its own repo. A pipeline validates it (schema, conflicts, ownership of path prefixes, policy requirements like "all routes require auth unless exempted"), then rolls it out gradually. A mistake in team A's manifest can only affect team A's routes.

L7: Defines who can own which prefixes, how conflicts are resolved, and how exceptions (unauthenticated routes, raised quotas) are approved and audited.

The Staff Positions#

PositionRationale
Thin gateway, no business logicBusiness logic in the gateway couples every team's release to the platform's
Authenticate at the edge, authorize in services tooEdge auth gives consistency; service auth gives defense in depth (zero trust)
No synchronous dependency on the hot pathLocal JWT validation, local-first rate limits, in-memory routes; a slow auth service must not slow every request
Config is code, staged by cell/regionThe most common gateway outage is a config push, not a crash
Self-service routes with scoped blast radiusTeams move fast; one team's mistake can't break another's routes
Aggregation belongs in BFFs, owned by product teamsClient-specific composition changes weekly; the gateway shouldn't
Buy the runtime, build the platformProxies are commodity; the self-service pipeline, policy and governance are yours

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
First-party client ingress (web + mobile → microservices)Latency, reliability, many teams shippingThin edge gateway: TLS, auth, rate limit, routing, telemetry; BFFs behind itGateway outage = total outage; config push breaks routesGateway overhead p99 ≤ 5ms; availability ≥ 99.99%; zero cross-team config blast radius
Public API product (partners, third-party developers)Contract stability, quotas, keys, billing, versioningAPI management: keys, plans, quotas, developer portal, versioned routes, usage meteringBreaking change ships to partners; quota miscounts bill wronglyVersioned contracts with deprecation windows; quota accuracy for billing
Aggregation / protocol translation (BFF, GraphQL, REST↔gRPC)Client-specific shapes, fewer round tripsSeparate BFF or GraphQL layer owned by client teams, behind the gatewayAggregation layer becomes a distributed monolith; fan-out latencyPer-client SLOs; partial responses when backends fail

🎯 Staff Move: "I'll design for first-party client ingress — web and mobile into a microservice fleet — because that's where the centralization tradeoffs are sharpest. If there's a public API product, I'd add API-management features on a separate hostname and plane, and I'd put aggregation in BFFs behind the gateway rather than in it."

The Five Fault Lines#

#Fault LineThe Tension
1Thin vs Smart GatewayCentralize more logic (consistency, reuse) or keep it minimal (independence, smaller blast radius)?
2One Central Gateway vs Federated GatewaysOne platform for all traffic, or per-domain gateways owned closer to teams?
3Auth at the Edge vs Auth in Every ServiceTrust the gateway's verdict, or re-verify everywhere?
4Gateway Aggregation vs BFF vs Client CompositionWhere do multi-service responses get assembled — and who owns them?
5Self-Service Velocity vs Global Blast RadiusLet teams change routing freely, or gate every change through a central team?

In the Wild: Real Production Systems#

Netflix — Zuul, Then Federated GraphQL#

Netflix's edge gateway Zuul routed all streaming API traffic into its microservices, with filters for authentication, routing, insights and resiliency testing (e.g., routing a slice of traffic to test failure behavior). Zuul 2 moved to an asynchronous, non-blocking architecture on Netty to handle large numbers of persistent connections efficiently. Separately, Netflix's engineering blog has described adopting federated GraphQL for parts of its product (notably Studio applications), where domain teams own their subgraphs and a gateway composes them — moving aggregation ownership to domain teams rather than a central API team.

Staff insight: The trajectory is instructive: a central edge for cross-cutting concerns stayed; client-specific aggregation moved to a model where domain teams own their slice. That's the thin-gateway-plus-owned-composition split, arrived at through scale.

Uber — Edge Gateway as an API Lifecycle Platform#

Uber's engineering blog describes its Edge Gateway as more than a proxy: a platform for the full API lifecycle — configuration-driven endpoint definitions, schema and protocol handling, authentication, rate limiting, and a self-service workflow so product teams can create and modify endpoints without the gateway team in the loop, with reviews and validation built into the pipeline.

Staff insight: At scale, the gateway's hard problem is not the proxy — it's the workflow by which thousands of engineers safely change it. Self-service with validation is the product.

AWS API Gateway — The Managed Contract#

Amazon API Gateway is a managed gateway whose published defaults teach good instincts: an account-level throttle of 10,000 requests per second with a 5,000-request burst per region by default (adjustable), per-stage and per-method throttles, usage plans with API keys and quotas for public APIs, and a 29-second default integration timeout. Stages and canary deployments are built into the model.

Staff insight: The managed model encodes what a gateway should own — throttling, keys and plans, staged deployment, timeouts — and makes the cost visible per request. When you design your own, those defaults are a sanity check: if your design has no throttle or no maximum timeout, you've left something out.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"The gateway handles auth""Auth service is slow. What happens to every request?"Hot-path dependencies
"Aggregate calls for mobile at the gateway""Five teams need different aggregations, weekly. Who changes the gateway?"Scope creep and ownership
"Services trust the gateway's identity header""What stops an internal caller from forging that header?"Zero trust, mTLS
"Run multiple instances""A bad route config reaches every instance. Now what?"Config blast radius
"Rate limit at the gateway""Per instance or global? What does a partner with a 1,000 RPS quota actually get?"Distributed quota accuracy
"One gateway for everything""Partners, mobile and internal tools share it. One of them spikes."Isolation between client classes

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: The gateway data plane is stateless and has no synchronous dependency on the request path: tokens are validated locally against cached signing keys, rate limits are enforced locally with async sync to a quota store, and routes live in memory from the last good config snapshot. Each team owns a route manifest in its own repo; the pipeline validates prefix ownership and policy before the config server rolls it out cell by cell. Aggregation lives in BFFs owned by the web and mobile teams, not in the gateway. Domain services still verify the caller's mTLS identity and make their own authorization decisions — the gateway is the first check, not the only one.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Scope"Auth, routing, rate limiting, caching, transforms, aggregation""Identity, coarse policy, rate limits, routing, telemetry. Aggregation and business logic go to team-owned BFFs and services."
Auth"Gateway validates the token""Gateway validates locally via cached JWKS and forwards a verified identity; services re-verify via mTLS and authorize."
Reliability"Multiple instances, autoscaling""No hot-path dependencies, fail-static config, staged rollout by cell, regional isolation."
Config"Gateway team manages routes""Self-service manifests, prefix ownership, validation, canary, auto-rollback; blast radius limited to the owning team."
Rate limiting"Redis counter per request""Local-first with async sync; per-key quotas for partners are approximate within a stated bound." See Rate Limiting.
Public API"Same gateway""Separate hostname and plane with keys, plans, versioning and deprecation windows."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Gateway latency overhead targetp50 < 1ms, p99 ≤ 2–5msAnything more is a tax on every request in the company
Local JWT verification~0.05–0.5ms (signature check with cached key)Why local validation beats per-request introspection
Token introspection call~2–20ms network + auth service availabilityA hot-path dependency on every request
JWKS cache TTL5–15 min, with refresh-ahead and key rotation overlapKey rotation must overlap by ≥ cache TTL
Access-token lifetime5–15 min typicalBounds revocation lag for locally validated JWTs
AWS API Gateway default throttle10,000 RPS steady / 5,000 burst per account per region (adjustable)Published reference for "a gateway needs a ceiling"
AWS API Gateway integration timeout29s defaultGateways need a max timeout well below client timeouts
Envoy-class proxy throughputTens of thousands of RPS per core for simple routingGateways are rarely CPU-bound unless you add heavy plugins
Connections per gateway node10K–100K+ concurrent (HTTP/2, WebSockets)Memory, file descriptors, and drain time on deploy
Route table sizeThousands of routes; trie/radix match in µsRegex-heavy routing is where CPU goes
Config propagation targetp99 < 30–60s per stage, staged over 30–60 min fleet-wideSpeed vs safety
Gateway availability target≥ 99.99% (≤ ~4.3 min/month)Must exceed every service behind it

Interview Walkthrough

Typical prompts: "Design an API gateway", "Design the entry point for our microservices", "We have 200 services and every team implements auth differently — fix it." The trap is to enumerate gateway features for 20 minutes. Spend 5 minutes on the request pipeline, then go to scope, hot-path dependencies, config safety and ownership.

Phase 1: Requirements & Framing (2–3 min)#

"First — who are the clients? First-party web and mobile apps, external partners, public developers, internal tools? They have different contracts. Second, how many services and teams sit behind it, and how often do routes change? Third, what's the auth model — who issues tokens? I'll assume web and mobile apps plus a partner API, 300 services owned by 60 teams, route changes several times a day, 150K RPS peak, and an existing OAuth identity provider issuing JWTs. Latency budget for the gateway itself: under 5ms p99."

Established:

  • Client classes (first-party vs partner) → separate planes/policies
  • Team count and change rate → self-service config is a requirement, not a nice-to-have
  • Identity provider exists → the gateway validates, doesn't issue
  • Overhead budget → forces a dependency-free hot path

Phase 2: Core Entities & API (1–2 min)#

EntityFields That MatterOwner
Routehost, path prefix, methods, upstream cluster, timeout, retry policy, auth requirement, rate-limit policy, owner teamProduct team (manifest)
Upstream clusterservice name (via discovery), LB policy, health, circuit limitsPlatform defaults, team overrides
Consumer (partner/app)API key / client ID, plan, quotas, allowed scopesAPI product team
Policyauthn mode, required scopes, rate limit, CORS, size limitsPlatform (defaults) + team (per route)
Config snapshotversion, routes, clusters, policies, signatureControl plane

A route manifest, owned by a product team:

# repo: orders-service/gateway/routes.yaml   (owner: team-orders)
host: api.example.com
routes:
  - prefix: /v1/orders            # prefix ownership registered to team-orders
    methods: [GET, POST]
    upstream: orders-svc
    timeout_ms: 800
    auth: { required: true, scopes: [orders:read, orders:write] }
    rate_limit: { per_user: 20/s, burst: 40 }
    retries: { idempotent_only: true }          # budget enforced by platform

🎯 Staff Move: "The most important API here isn't what clients call — it's the route manifest. That's the contract between the platform team and 60 product teams, and the pipeline that validates it is what keeps one team's typo from becoming everyone's outage."

Phase 3: High-Level Architecture (≤5 min)#

Draw the overview and name owners:

  1. CDN/WAF in front — volumetric protection, coarse bot filtering (CDN & Edge Caching).
  2. Gateway data plane — stateless proxies per region, in cells; TLS, auth, rate limit, routing, telemetry. Owner: platform.
  3. Gateway control plane — manifests → validation pipeline → versioned snapshots → staged push. Owner: platform.
  4. BFFs — per-client aggregation, owned by client teams.
  5. Services — verify mTLS identity, authorize, own their business logic.

Phase 4: Transition to Depth#

"The request pipeline is standard; I'd rather spend time where gateways actually fail. Three places: hot-path dependencies — auth and rate limiting — that make the gateway only as available as its slowest dependency; configuration changes, which cause most gateway outages; and scope creep, which turns the gateway into a monolith. Then multi-region and the public API plane."

Phase 5: Deep Dives (25–30 min)#

Deep dive 1 — Dependency-free hot path (7 min).

Per request, no network calls except to the upstream:
  TLS:            session resumption for returning clients
  Authn:          JWT signature check with cached JWKS (refresh every 10 min, refresh-ahead)
                  opaque tokens (partners): cache introspection result for min(token TTL, 60s)
  Rate limit:     local token bucket per key; async sync to quota store every 500ms
  Routing:        in-memory radix tree from last good snapshot
  Identity out:   gateway strips client-supplied identity headers, adds signed x-identity,
                  upstream call over mTLS
Failure posture:
  IdP down        → keep validating with cached keys (tokens valid until expiry)
  Quota store down → local limits only (fail-open for abuse protection, conservative caps)
  Control plane down → serve last good config indefinitely

Deep dive 2 — Config safety (7 min). Validation (schema, prefix ownership, conflicts, policy: "no unauthenticated routes without an exception"), then staged rollout: canary cell (1–5% traffic) → 1 region → all regions, with automatic rollback on 5xx_by_route, 404_rate, or latency regression. A team's manifest can only affect its own prefixes.

Deep dive 3 — Scope (5 min). Criteria for whether a feature belongs in the gateway: same for every service? needed before routing? stateless? If not all three, it goes to a service or BFF.

Deep dive 4 — Isolation (5 min). Separate listeners/clusters (or separate gateways) for first-party and partner traffic so a partner spike can't exhaust capacity for the mobile app; per-client-class concurrency limits; cells so a poison request pattern takes out one cell.

Deep dive 5 — Multi-region (3 min). Gateways per region, anycast/GeoDNS to the nearest healthy region, routes prefer in-region upstreams; config staged by region.

Phase 6: Wrap-Up (2–3 min)#

"To summarize: a thin, stateless gateway per region and cell that owns TLS, identity, rate limits, routing and telemetry, with no synchronous dependencies on the hot path. Teams own their routes via manifests; a pipeline validates ownership and policy and rolls changes out by cell with automatic rollback. Aggregation lives in team-owned BFFs; services re-verify identity. Next I'd add per-route SLO dashboards for teams, a policy engine for exceptions, and the partner API plane with keys, plans and versioning. Biggest risk: a gateway-wide change — TLS cert, auth library, global plugin — that bypasses the staged pipeline. I'd make the pipeline the only path, including for the platform team."

Common Timing Mistakes#

MistakeTime LostWhat to Do Instead
Listing every gateway feature10 minPipeline table in 2 minutes
Comparing Kong vs Envoy vs AWS5–8 min"Any mature proxy; see the technology guide" — then move on
Designing the rate-limiting algorithm10 minReference local-first + async sync; link to Rate Limiting
Ignoring config changes—The #1 cause of gateway outages
Putting aggregation in the gateway without discussing ownership—Raise the BFF question explicitly

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

The gateway is the most literal example of a platform decision: one team builds something that every other team depends on and changes. Senior engineers can build a proxy. Staff engineers decide what the proxy is for — and, just as importantly, what it will refuse to do — then design the workflow that lets dozens of teams use it without coordination meetings, and the rollout discipline that keeps the one component on every request path from being the thing that takes everything down.

1.2 The L5 vs L6 vs L7 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 vs L7 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"If we add this to the gateway, what happens to every other team the day it has a bug?"

Apply it to every proposed gateway feature. JWT validation: a bug affects everyone — but so does 40 different implementations; centralize, with extreme care. Aggregating the mobile home screen: a bug affects everyone to benefit one team; put it in the mobile BFF. Request body validation against 300 schemas: a bug or slow schema affects everyone; leave it in services.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Intent 1: First-party client ingress. Web and mobile apps call your backend. You control both sides, can ship client updates (slowly, for mobile), and care most about latency, reliability and team velocity. The gateway owns TLS, authentication, coarse authorization, rate limiting against abuse, routing (including canaries), and telemetry. Versioning is flexible because you own clients — but old mobile versions live for years, so routes stay backward compatible.

Intent 2: Public API product. Partners and developers integrate against a contract you can't change unilaterally. The gateway needs API keys and OAuth client credentials, plans and quotas (often tied to billing), usage metering, versioned routes (/v1, /v2), deprecation headers and windows (often 6–12+ months), and a developer portal. Quotas used for billing need accuracy that abuse-protection limits don't. Treat it as a product with a product manager, on its own hostname and capacity pool.

Intent 3: Aggregation and protocol translation. Clients want one call instead of seven; internal services speak gRPC; the web wants GraphQL. These are real needs, but they're client-specific and change at client-team velocity. Put them in a BFF per client (or a GraphQL layer with federated ownership) behind the gateway — not in the gateway.

🎯 Staff Move: "The gateway serves every team; a BFF serves one client. If a feature's requirements come from one client team, it belongs in something that team can deploy without asking the platform team."

2.2 When NOT to Use an API Gateway (or a New One)#

SituationWhy It's WrongWhat to Use Instead
Service-to-service (east-west) trafficHairpinning internal calls through a central gateway adds a hop and a shared bottleneckService mesh / client-side LB with mTLS (Service Discovery)
A monolith with one backendA reverse proxy or LB with TLS is enoughLoad balancer (Load Balancer)
Streaming bulk data / large uploadsGateway buffers and timeouts break large transfersPre-signed URLs direct to object storage (Blob Storage)
Ultra-low-latency paths (trading, real-time bidding)Even 1ms overhead mattersDirect connections with auth in the service or a sidecar
A new team-specific gatewayFragments policy and identityRoutes on the shared gateway, or a BFF behind it
Business workflowsOrchestration in the gateway is a monolith in disguiseServices, workflow engines

2.3 What the Interviewer Leaves Underspecified#

Unstated AssumptionWhy It MattersWhat to Say
Client classesDifferent contracts and isolation needs"First-party and partner, isolated"
Number of teams and route churnSelf-service config vs central team"60 teams, daily changes → self-service"
Identity providerValidate vs issue"Existing IdP issues JWTs; gateway validates"
ProtocolsHTTP/1.1, HTTP/2, gRPC, WebSockets"HTTP/2 and gRPC; WebSockets on a separate listener with long-lived connection handling"
Latency budgetHot-path design"≤ 5ms p99 overhead"
RegionsIsolation and failover"Two regions; gateways per region"

2.4 Precise Terminology#

TermPrecise MeaningCommon Confusion
North-south trafficClient → backend, crossing the edgeConfused with east-west (service ↔ service)
Edge gatewayPublic-facing gatewayAssumed to be the same as a mesh ingress
BFF (Backend for Frontend)Per-client aggregation service owned by that client's teamAssumed to be a gateway feature
AuthenticationWho is calling (verify token)Confused with authorization
AuthorizationMay this caller do this action on this resourceAssumed to be fully doable at the gateway (it isn't — resource-level checks need domain data)
Token introspectionAsking the IdP whether an opaque token is validAdds a per-request network dependency
JWKSPublished set of public keys for verifying JWT signaturesCache misunderstanding → rotation outages
Route manifestDeclarative per-team route configConfused with the gateway's global config
Data plane / control planeProxies serving traffic / system that computes and distributes configControl plane put in the request path
CellAn isolated slice of gateway capacity serving a subset of trafficConfused with region

3. The Five Fault Lines#

3.1 Fault Line 1: Thin vs Smart Gateway#

StrategyWhat WorksWhat BreaksWho Pays
Thin (TLS, authn, coarse authz, rate limit, routing, telemetry)Small blast radius; easy to reason about; platform team isn't a feature bottleneckSome duplication in services/BFFsProduct teams (own their logic)
Smart (+ transforms, validation, aggregation, caching, business rules)One place for many concerns; less code in servicesEvery team's features ride the platform's release train; bugs affect everyone; plugin latency accumulatesEveryone, in the outage; platform team, in the backlog
Pluggable with governance (extensions allowed, reviewed, sandboxed, per-route)Flexibility for genuinely cross-cutting needsGovernance overhead; plugins still run in the shared processPlatform (review), owners of plugins

The Staff default: Thin, with a short, written list of what the gateway owns, and an extension mechanism for truly cross-cutting needs that goes through review, runs with a CPU/latency budget, and can be disabled per route.

The scope test: A feature belongs in the gateway only if it is (1) identical for all or most services, (2) needed before routing, and (3) stateless or cacheable locally. Authentication passes. Aggregation fails (1). Business validation fails (1) and (3).

🎯 Staff Move: "The gateway becoming the new monolith is the default outcome, not a risk. It happens one reasonable feature request at a time. I'd write the scope charter on day one and make every addition justify itself against it."

3.2 Fault Line 2: One Central Gateway vs Federated Gateways#

StrategyWhat WorksWhat BreaksWho Pays
One central gateway (one platform, one config, all traffic)Uniform policy, one identity model, one telemetry formatLargest blast radius; all client classes share capacity; central team bottleneckEveryone (correlated outage)
Per-domain gateways (each business unit runs its own)Autonomy; independent release; smaller blast radiusDivergent auth, rate limiting and logging; duplicated effort; security gapsSecurity, platform consistency; users see inconsistent APIs
One platform, many deployments (shared runtime and pipeline; separate gateway instances per client class/domain/cell)Uniform contract with isolated failure domainsMore deployments to operate; routing to the right instancePlatform headcount

The Staff default: One gateway platform — same runtime, same pipeline, same policy engine — deployed as multiple isolated instances: at least first-party vs partner, and cells within each. Autonomy comes from self-service manifests, not from separate gateways.

When to deviate: Acquired companies or regulated business units may need separate gateways temporarily; plan convergence rather than accepting permanent divergence.

3.3 Fault Line 3: Auth at the Edge vs Auth in Every Service#

StrategyWhat WorksWhat BreaksWho Pays
Edge only (gateway validates; services trust a header)Simple; one implementationAny internal caller can forge the identity header; lateral movement after one compromiseSecurity; every service
Services only (each validates tokens)Defense in depth40 implementations; library drift; vulnerabilities patched over weeksEvery team; security
Edge authenticates, services authorize with verified identity (Staff default)One token-validation path; services still enforce access; zero-trust transportNeeds mTLS/workload identity and a signed internal identity formatPlatform (identity plumbing)
Token exchange (gateway exchanges external token for a short-lived internal token)Internal tokens scoped per audience; external tokens never enter the fleetMore moving parts; token service on a near-pathPlatform + security

The Staff default: The gateway validates the external token locally (cached JWKS), strips any client-supplied identity headers, and forwards a signed internal identity (or an exchanged internal token) over mTLS. Services verify the mTLS peer is the gateway (or an authorized caller) and perform resource-level authorization — "can user 42 read order 9001?" — because only they have the data to answer it.

Client → Gateway:   Authorization: Bearer <external JWT, 15 min>
Gateway:            verify signature (cached JWKS), exp, aud, iss; check coarse scopes
                    strip: x-user-id, x-identity, x-tenant (anything client-supplied)
Gateway → Service:  mTLS (gateway workload identity)
                    x-identity: <signed internal assertion: sub, tenant, scopes, exp 60s>
Service:            verify mTLS peer ∈ allowed callers; verify assertion signature;
                    authorize resource access with domain data

Revocation tradeoff: Locally validated JWTs can't be revoked instantly. Keep lifetimes short (5–15 min), and for high-risk events (account compromise), push a small deny-list of token IDs or subjects to the gateway via the config channel.

3.4 Fault Line 4: Gateway Aggregation vs BFF vs Client Composition#

StrategyWhat WorksWhat BreaksWho Pays
Aggregation in the gatewayFewer client round tripsClient-specific logic in the shared component; fan-out failures handled centrally for everyonePlatform team (backlog), everyone (blast radius)
BFF per client (web BFF, iOS/Android BFF)Client team owns shape and release; tailored payloadsSome duplication across BFFs; more services to runClient teams
GraphQL with federated subgraphsClients request exactly what they need; domain teams own subgraphsQuery cost control, N+1 fan-out, schema governancePlatform (router), domain teams (subgraphs)
Client composition (client calls services directly via gateway)No middle layerMany round trips on mobile networks (each 50–300ms on cellular)Users on slow networks

The Staff default: BFFs owned by client teams for first-party apps; federated GraphQL when many clients need flexible composition and the org can invest in schema governance. The gateway routes to them; it doesn't compose.

Diagram: 3.4 Fault Line 4: Gateway Aggregation vs BFF vs Client Composition

3.5 Fault Line 5: Self-Service Velocity vs Global Blast Radius#

StrategyWhat WorksWhat BreaksWho Pays
Central team edits all configConsistency; expert reviewBottleneck: days to add a route; platform team becomes a ticket queueProduct teams (velocity)
Teams edit a shared config freelyFastOne team's bad regex or wildcard route breaks others; global pushEveryone
Self-service manifests + validation + staged rollout (Staff default)Fast and safe; blast radius scoped to the owner's prefixesPipeline to build and maintain; ownership registryPlatform

Validation that matters:

  • Prefix ownership: a manifest may only define routes under prefixes registered to that team; conflicts are rejected.
  • Policy: routes require auth unless an approved exception exists; timeouts ≤ platform max (e.g., 30s); retries only for idempotent methods.
  • Safety: regex routes are linted for catastrophic backtracking; wildcard routes need review.
  • Rollout: canary cell → region → global, each with health gates and automatic rollback, ~30–60 min end to end; emergency path exists but still staged, just faster.

🎯 Staff Move: "Self-service doesn't mean unreviewed. It means the pipeline does the review — ownership, policy, safety — so humans only review exceptions. And the platform team uses the same pipeline; there is no side door."


4. Failure Modes & Operational Reality#

4.1 The Config Push That Took Down Every Route#

t=0:      Team-search merges a manifest with route prefix '/' (meant '/search/')
t=+1min:  Pipeline v1 has no prefix-ownership check; config server pushes globally
t=+1m30s: All gateway nodes load the new snapshot; '/' matches before longer prefixes
          due to an ordering bug → 70% of requests routed to search-svc
t=+2min:  404s and 400s everywhere; checkout, login, feed all failing
t=+6min:  On-call identifies config version change; rollback initiated
t=+9min:  Rollback propagates; errors subside

Detection: gw_5xx_by_route and gw_404_rate spike correlated with config_version change; per-route traffic share shifts abruptly.

Mitigation: One-click rollback to previous snapshot; config server keeps the last N signed snapshots.

Prevention: Prefix ownership registry; longest-prefix-match semantics tested in CI; staged rollout (canary cell 5% for 10 min with automatic rollback on route-level error or traffic-share anomalies); diff reports showing "routes whose upstream changed" for review.

Owner: Platform owns the pipeline and rollback; team-search owns the manifest. The post-mortem finding is the platform's: the pipeline allowed a change outside the team's ownership.

Diagram: 4.1 The Config Push That Took Down Every Route

4.2 The Hot-Path Dependency — Auth Service Slowdown#

t=0:      Gateway validates partner opaque tokens via introspection on every request
t=+0s:    Identity service DB has lock contention; introspection p99 2ms → 1.5s
t=+10s:   Gateway worker connections pile up waiting on introspection;
          Little's law: 20K RPS × 1.5s = 30K in-flight vs ~8K capacity
t=+30s:   Gateway latency for ALL routes rises (shared event loop/connection pools)
t=+60s:   First-party mobile traffic (JWT, no introspection) also timing out
          because it shares gateway capacity with partner traffic
t=+8min:  Identity team fixes DB; recovery

Detection: gw_auth_latency_p99, gw_upstream_pending{cluster=idp}, overall gw_request_latency_overhead_ms.

Mitigation: Cache introspection results (min(token TTL, 60s)); bulkhead introspection calls (dedicated connection pool with a concurrency cap and a 50ms timeout); on timeout for a previously-seen valid token, allow within cache grace period.

Prevention: Prefer JWT for first-party (local validation); separate partner and first-party gateway deployments; never let a per-request dependency share capacity with traffic that doesn't need it.

Owner: Platform (gateway design); identity team (IdP performance).

4.3 Certificate Expiry#

t=0:      Wildcard cert for *.api.example.com expires at 00:00 UTC Saturday
t=+0s:    Every client TLS handshake fails; mobile apps show 'cannot connect'
t=+20min: Someone notices social media reports; on-call paged by synthetic monitors
t=+50min: New cert issued, deployed to all gateway nodes

Detection: cert_expiry_seconds per hostname per node; external synthetic TLS checks.

Prevention: Automated issuance and renewal (ACME or internal CA) at ≤ 2/3 of lifetime; alerts at 30, 14, 7 days; certs deployed through the staged config pipeline like everything else; inventory of every hostname's certificate.

Owner: Platform. This is a 100%-blast-radius failure with a 100%-preventable cause — which is why it keeps happening across the industry.

4.4 Key Rotation Lockout#

The IdP rotates its JWT signing key and immediately stops publishing the old one. Gateways that cached JWKS 8 minutes ago don't have the new key and reject every newly issued token; gateways that refresh see the new key but tokens signed with the old key (still valid for 15 minutes) now fail.

Prevention: Rotation protocol: publish new key → wait ≥ JWKS cache TTL → start signing with new key → keep old key published until the longest token lifetime passes → remove. Gateways refresh JWKS on unknown kid (rate-limited) rather than failing immediately.

Owner: Identity team (rotation protocol); platform (JWKS refresh behavior).

4.5 Noisy Client Class#

A partner's batch job sends 30K RPS of valid, within-quota requests to a partner endpoint that's expensive upstream. The shared gateway deployment's connection pools and CPU are consumed; first-party app latency rises 4×.

Prevention: Separate deployments (or at least separate listeners, worker pools and upstream connection pools) for first-party vs partner; per-client-class concurrency caps; quotas that account for cost, not just request count.

Owner: Platform (isolation); API product team (partner quotas).

4.6 Retry Amplification Through the Gateway#

The gateway retries 5xx on all routes twice; the BFF retries twice; services retry. A downstream brownout becomes a 27× load multiplier. See Circuit Breakers for the full analysis. Gateway rule: retries only on connection failure and idempotent methods, under a budget (≤ 10% of requests), and never when the upstream returned Retry-After or the remaining deadline is too short.

Owner: Platform.

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Bad route configgw_5xx_by_route + config_version changePotentially all routesRollback; staged rollout; ownership validationPlatform (pipeline), team (manifest)
Hot-path auth dependency slowgw_auth_latency_p99, pending requestsAll traffic sharing the gatewayCache, bulkhead, local JWT, separate planesPlatform + identity
Certificate expirycert_expiry_seconds, TLS syntheticsHostname — often everythingAutomated renewal, alertsPlatform
Key rotation lockoutauth_failures{reason=unknown_kid} spikeAll authenticated trafficRotation protocol; refresh on unknown kidIdentity + platform
Noisy client classPer-class latency divergenceOther client classesSeparate deployments and poolsPlatform
Retry amplificationgw_retry_ratioDownstream then allRetry budget, idempotent onlyPlatform
Plugin CPU regressiongw_cpu, gw_overhead_ms after plugin changeAll routes using the pluginPer-plugin latency budget; disable per routePlugin owner + platform
Gateway region outageRegional syntheticsUsers routed to that regionDNS/anycast failover with capacityPlatform/SRE

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
FramingFeature listClient classes, teams, change rate, overhead budgetExisting gateways, platform contract, scope charter
ScopeGrows to include aggregation and transformsThin; scope test; BFFs owned by client teamsCharter enforced via governance; says no
ReliabilityMore instancesDependency-free hot path; fail-static config; cells; staged rolloutGateway error budget, cells, correlated-failure analysis
SecurityValidate tokens at gatewayEdge authn + service authz, mTLS, header stripping, rotation protocolOrg identity model: token exchange, workload identity, audit
ConfigCentral teamSelf-service manifests, ownership, validation, canaryGovernance, exceptions, audit, deprecation of side doors
IsolationOne deploymentSeparate planes for client classes; per-class poolsCell architecture tied to business units/tenants
CostNot discussedPer-request overhead, plugin budgetsBuild vs buy vs managed at scale; cost allocation per team

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Draws a scope boundary"Identity, policy, routing, telemetry. Aggregation goes to BFFs."
Eliminates hot-path dependencies"JWTs validated locally; if the IdP is down, cached keys keep working."
Treats config as the main risk"Most gateway outages are config pushes; every change is staged by cell with automatic rollback."
Zero trust"Gateway strips client identity headers; services verify mTLS and authorize."
Isolates client classes"A partner spike can't consume mobile's capacity."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
Gateway does aggregation, validation and business rulesBuilds the next monolith
Per-request introspection for all traffic, unexaminedMakes the IdP a dependency of every request
"Services trust the X-User-Id header"Header forgery; no defense in depth
No plan for config changesIgnores the most common failure
Gateway handles east-west traffic tooCentral bottleneck for internal calls

5.4 Common False Positives#

  • Detailed plugin knowledge for one vendor ≠ gateway design. The questions are scope, dependencies and ownership.
  • "Add GraphQL" ≠ solving aggregation. It moves the problem to schema governance and query cost control.
  • Drawing a mesh and a gateway ≠ understanding their split. North-south vs east-west, and who owns each.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minClient classes, teams, change rate, auth model
Entities & manifest3–6 minRoute manifest as the contract
Architecture6–11 minData plane, control plane, BFFs, services
Hot path & auth11–19 minLocal validation, identity propagation, failure posture
Config safety19–26 minValidation, ownership, staged rollout
Scope & aggregation26–32 minScope test, BFF vs GraphQL
Isolation & multi-region32–40 minClient classes, cells, regions
Wrap-up40–45 minNext steps, biggest risk

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Direction
"Add response caching at the gateway"Scope and correctnessOnly for public, non-personalized GETs with explicit headers; prefer CDN (CDN & Edge Caching)
"Support WebSockets"Long-lived connectionsSeparate listener/deployment; connection draining on deploy; per-node connection caps
"Partners need usage-based billing"Quota accuracySeparate metering pipeline from abuse limits; exact counts from logs, not gateway counters
"Version the API"Contract managementPath or header versioning, deprecation headers, sunset dates, usage tracking by version
"We want to build our own gateway"Build vs buyBuy/adopt the proxy, build the pipeline and policy
"Gateway latency is 20ms"Performance debuggingPer-filter timing; plugins, regex routes, sync dependencies, TLS without resumption

6.3 What to Deliberately Skip#

  • Vendor feature comparisons (use the technology guide)
  • Rate-limit algorithm details (link to Rate Limiting)
  • OAuth flow diagrams beyond "IdP issues tokens; gateway validates"
  • Load balancer algorithms for upstreams

6.4 Follow-Up Questions to Expect#

  1. "How do you revoke a token immediately?" — Short lifetimes plus a pushed deny-list of token IDs/subjects for emergencies.
  2. "How do services know the request came through the gateway?" — mTLS peer identity plus a signed internal assertion with a short expiry.
  3. "How do you do canary releases of a service through the gateway?" — Weighted routes in the manifest, header-based overrides for testers, automatic promotion on SLOs.
  4. "Where do CORS and request size limits live?" — Gateway — they're uniform and needed before routing.
  5. "What if the gateway itself needs to call a service to decide routing?" — Avoid; if unavoidable, cache aggressively and fail to a safe default route.
  6. "How do you debug a slow request?" — Request ID at the gateway, trace span with per-filter timings, upstream timing header.
  7. "How do you deprecate a route?" — Deprecation and Sunset headers, usage tracking per consumer, outreach, then 410 after the window.

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design an API gateway for our company."

Staff Answer

"Three questions first: who are the clients — our own apps, partners, public developers? How many teams are behind it and how often do routes change? And who issues identity today? I'll assume web and mobile apps plus a partner API, ~60 teams, daily changes, and an existing OAuth IdP.

My design principle is a thin gateway: TLS, authentication, coarse authorization, rate limiting, routing and telemetry — things that are identical for every service and needed before routing. Aggregation goes to BFFs owned by client teams; business logic stays in services.

Then the three things I'd go deep on: keeping the hot path free of synchronous dependencies so the gateway is more available than anything behind it; making config changes self-service but safe, since config pushes cause most gateway outages; and isolating first-party from partner traffic so one can't starve the other."

Why this is L6:

  • Frames by clients, teams and change rate before features
  • States a scope principle with a reason
  • Leads with the real failure sources

What L7 adds:

  • Asks what gateways exist today and positions this as the platform contract with a retirement plan for the rest
  • Names the gateway's own availability target and how it's funded
❌ Common L5 Trap

"The gateway will handle authentication, rate limiting, request validation, response transformation, caching, aggregation and logging. I'll use Kong with plugins for each."

Why this misses: A feature list with no scope boundary, no ownership model, and no discussion of how the gateway fails. It's a description of the next monolith.


Drill 2: Auth Without a Hot-Path Dependency#

Prompt: "How does the gateway authenticate requests without depending on the identity service for every call?"

Staff Answer

"For first-party clients, JWT access tokens with 10–15 minute lifetimes. The gateway validates signatures locally with the IdP's public keys, cached from the JWKS endpoint for ~10 minutes with refresh-ahead and an on-demand refresh when it sees an unknown key ID, rate-limited. Validation is ~0.1ms and needs no network call. If the IdP is down, validation continues until tokens expire; new logins fail, existing sessions keep working.

For partners with opaque tokens, cache introspection results for min(token TTL, 60s), bulkhead the introspection calls with a 50ms timeout and a concurrency cap, and run partner traffic on a separate gateway deployment so an IdP slowdown can't affect first-party traffic.

Revocation: short token lifetimes plus a deny-list of subjects or token IDs pushed through the config channel for emergencies — seconds, not minutes."

Why this is L6:

  • Designs for the IdP being slow or down
  • Quantifies validation cost and cache behaviour
  • Handles revocation, the weakness of local validation

What L7 adds:

  • Owns the key-rotation protocol across IdP and gateway teams as a documented standard
  • Considers token exchange so external tokens never enter the internal fleet

Drill 3: Make It Concrete — Self-Service Routes#

Prompt: "Sixty teams need to add and change routes daily. How do they do it without breaking each other?"

Staff Answer

"Each team owns a route manifest in its own repo. A prefix registry maps path prefixes and hostnames to owning teams; the pipeline rejects any route outside the team's prefixes. Validation checks schema, policy — auth required unless an approved exception, timeouts ≤ 30s, retries only for idempotent methods — and safety, like linting regexes for catastrophic backtracking.

Approved manifests compile into a versioned, signed snapshot. Rollout: canary cell at ~5% of traffic for 10 minutes, then one region, then all, with automatic rollback if that team's routes show 5xx, 404 or latency regressions, or if any other team's traffic share shifts — which catches overlapping routes. End to end ~30–60 minutes; an expedited path exists for incidents but is still staged.

Blast radius of a team's mistake: its own routes, in the canary cell, for up to 10 minutes."

Why this is L6:

  • Ownership enforced by the pipeline, not by review
  • Staged rollout with route-level health gates
  • States the blast radius explicitly

What L7 adds:

  • Defines governance for prefix disputes, exceptions and audit
  • Makes the platform team use the same pipeline — no side doors

Drill 4: The Identity Service Is Down#

Prompt: "The identity provider has been down for 30 minutes. What works?"

Staff Answer

"Requests with valid, unexpired JWTs keep working — the gateway validates locally with cached keys. New logins and token refreshes fail, so users whose tokens expire during the outage get logged out, gradually: with 15-minute tokens, by the 30-minute mark most active sessions need a refresh.

Mitigations, pre-decided: the IdP's refresh endpoint is its most critical path and should be the last thing to fail; clients refresh tokens early (at ~75% of lifetime) so a short outage is invisible; and for the partner plane, cached introspection results cover up to 60s. What I would not do is extend token validity at the gateway on the fly — that's a security decision with a named owner, pre-approved, if at all."

Why this is L6:

  • Walks through the timeline of degradation
  • Identifies refresh behaviour as the lever
  • Refuses to improvise security relaxation

What L7 adds:

  • Sets the IdP's availability target above the gateway's, since everything depends on it, and funds it accordingly

Drill 5: Aggregation Request#

Prompt: "The mobile team wants the gateway to combine five service calls into one home-screen response."

Staff Answer

"The need is real — five round trips on a cellular connection could cost a second — but the gateway is the wrong place. The home screen's shape is mobile-specific and changes with mobile releases; putting it in the gateway means the mobile team waits on the platform team, and a bug in their aggregation affects every route.

So: a mobile BFF owned by the mobile team, behind the gateway. It calls the five services in parallel with per-call timeouts inside the request deadline, returns partial results when optional sections fail, and ships on the mobile team's schedule. The gateway routes /mobile/* to it and handles auth and rate limiting as usual. If several clients later need flexible composition, that's the point to evaluate federated GraphQL."

Why this is L6:

  • Validates the need, rejects the location, with reasons
  • Designs the BFF's failure behaviour
  • Names the trigger for GraphQL

What L7 adds:

  • Adds BFFs to the platform's paved road — a template with deadline propagation and partial-response helpers — so each client team doesn't build one from scratch

Drill 6: One Partner Dominates#

Prompt: "One partner is 60% of partner-API traffic and its batch jobs spike to 10× at midnight."

Staff Answer

"Three layers. Isolation: partner traffic already runs on a separate gateway deployment from first-party, so mobile is safe. Within the partner plane, per-partner concurrency limits and quotas so this partner's spike consumes its own allocation; for their size, possibly a dedicated cell — shuffle-sharded if we have many large partners. Contract: their midnight spike is predictable, so negotiate a batch endpoint or a scheduled window with a higher quota, rather than enforcing the interactive quota against a batch workload.

Quotas for billing need accuracy; I'd meter usage from access logs in a separate pipeline, not from the gateway's approximate local counters."

Why this is L6:

  • Isolation at deployment, partner and cell levels
  • Treats the problem as partly a contract question
  • Separates billing accuracy from enforcement

What L7 adds:

  • Prices dedicated capacity into the partner's plan — large customers pay for isolation

Drill 7: Build vs Buy#

Prompt: "Should we build our own gateway?"

Staff Answer

"Build the platform, not the proxy. The data plane — TLS, HTTP/2, connection management, routing, filters — is mature in open-source and managed options; writing one means years of catching up on protocol edge cases and security fixes. The comparison between those options lives in the technology guide.

What's specific to us, and what we'd build: the route manifest format, the prefix-ownership registry, the validation and staged rollout pipeline, the policy defaults, identity propagation, and dashboards per team. That's where the value — and the safety — is.

Managed vs self-hosted depends on scale and control: managed gateways charge per request, which is great at low volume and expensive at billions of requests per month; self-hosted costs a platform team. I'd estimate the crossover with our volume before choosing."

Why this is L6:

  • Separates commodity (proxy) from differentiating (pipeline, policy)
  • Links out rather than re-litigating vendors
  • Frames managed vs self-hosted as a volume question

What L7 adds:

  • Keeps manifests vendor-neutral so the runtime can be swapped — the exit path is designed in
  • Builds the 3-year cost curve for managed vs self-hosted

Drill 8: Policy Change Without an Outage#

Prompt: "Security wants every route to require authentication by default. Today 15% of routes are unauthenticated. Roll it out."

Staff Answer

"First inventory: which of the 15% are intentionally public (health checks, login, public catalog) and which are accidents? Pull traffic data per route. Then shadow mode: the gateway evaluates 'would this request fail auth?' and logs it for 2 weeks. Owners of routes with real unauthenticated traffic either add auth or file an exception — with security approval and an expiry.

Then enforce in stages — by team, lowest-traffic first, canary cell first — with automatic rollback if 401s exceed the shadow prediction. Finally, flip the default in the manifest schema so new routes are authenticated unless they declare auth: public with an exception ID. Success metric: unauthenticated routes without exceptions = 0."

Why this is L6:

  • Inventory, shadow, exceptions, staged enforcement, then change the default
  • Uses data to predict impact

What L7 adds:

  • Makes shadow mode a built-in capability of the policy engine for every future policy
  • Reports exception count to security leadership as a standing metric

Drill 9: Cost#

Prompt: "Our managed gateway bill is $180K/month. Is that reasonable?"

Staff Answer

"Work out the unit cost. At, say, ~$1–3.50 per million requests on a managed gateway, $180K/month implies roughly 50–180 billion requests a month — 20–70K RPS average. Self-hosting an Envoy-class gateway at that volume needs perhaps a few dozen instances across regions — low tens of thousands of dollars in compute — plus a platform team of 3–5 engineers ($1–1.5M/year) who we may already have.

So at this volume self-hosting is probably cheaper in direct cost, but the real question is whether we have the team and whether managed features (usage plans, developer portal) are load-bearing. Before migrating, cheaper wins: move internal east-west calls that hairpin through the gateway onto the mesh — often a big share of requests — and move large downloads to pre-signed URLs."

Why this is L6:

  • Derives the unit economics
  • Compares with headcount, not just compute
  • Finds waste before migrating

What L7 adds:

  • Allocates gateway cost per team by route so owners see their usage
  • Plans the migration as a two-way door with a runtime-agnostic manifest

Drill 10: Multi-Region#

Prompt: "We're going active-active in two regions. How does the gateway work?"

Staff Answer

"A gateway deployment in each region, with users routed to the nearest healthy region via anycast or latency-based DNS. Each regional gateway routes to in-region upstreams by default; cross-region routing only for services that exist in one region, with explicit latency budgets.

Config rolls out region by region, never simultaneously, so a bad push takes out at most one region, which the edge can route around. Failover capacity: each region must absorb the other's traffic, so either run each at ≤ 50% peak utilization or have shedding rules that protect critical routes during failover. Sessions must not be pinned to a region — stateless JWTs make that easy — and write paths depend on the data layer's multi-region model, which the gateway must not paper over."

Why this is L6:

  • Regional isolation with staged config
  • Failover capacity math
  • Knows the gateway can't solve data-layer consistency

What L7 adds:

  • Cells within each region so most incidents never need regional failover
  • Prices the headroom for leadership

8. Deep Dive Scenarios#

Deep Dive 1: Peak-Traffic Latency Collapse#

Context: During the holiday peak (4× normal traffic), gateway p99 overhead rises from 3ms to 400ms and error rates climb to 8% across all routes. Backend services report normal latency. The gateway fleet autoscaled from 40 to 120 nodes, but it didn't help. You're pulled in.

Questions to Surface First:

  • Is the time spent inside the gateway (filters) or waiting on something (upstream connections, a dependency)?
  • Did anything change recently — a plugin, a route with a regex, a logging config?
  • Are upstream connection pools exhausted, and did autoscaling multiply connection counts to backends?
  • Is a shared dependency (rate-limit store, auth) in the hot path?

Typical L5 Approach: Scales further, raises instance sizes. The added nodes each open new connection pools to every upstream, making backend connection counts explode, and the true bottleneck — a synchronous call to a central rate-limit Redis per request, added two weeks earlier by a well-meaning plugin update — gets worse with more nodes.

Staff Approach: Uses per-filter timing to find that the rate-limit filter accounts for ~390ms: its Redis cluster is saturated at 4× load. Switches the rate-limit filter to local-first mode via config (staged, canary cell first), cutting overhead back to ~3ms within 15 minutes. Caps upstream connection pools per node to prevent the autoscaling connection explosion.

Principal Approach: Asks how a synchronous per-request dependency entered the gateway hot path without review. Establishes a rule that any filter making network calls requires platform architecture review, a latency budget, and a failure mode (fail-open with local fallback) — and adds peak-traffic load testing of the gateway with all production filters to the holiday readiness checklist.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Per-filter latency breakdown; identify the rate-limit filter.
TriageRedis CPU 100%; per-request calls = gateway RPS; scaling gateways increased Redis load.
Quick fixLocal-first limiter mode; staged rollout.
GuardrailsUpstream connection caps per node; stop autoscaling beyond need.
Post-mortemPlugin update introduced a hot-path dependency without review or load test.

Metrics to Watch: gw_filter_duration_ms{filter}, gw_overhead_ms_p99, ratelimit_store_latency_p99, upstream_cx_active{cluster}.

Organizational Follow-up: Filter review policy; per-filter latency budgets; peak load test.

Ownership Question: "Who approved the plugin update?" Staff answer: It went through normal code review, which checks correctness, not latency under 4× load. The gap is that gateway filters need a different review — performance and failure mode — owned by the platform's architecture reviewers.

Key Takeaway: "When the gateway is slow and backends are fine, look for a filter that makes a network call."

What clears the Staff bar:

  • Per-filter timing before scaling
  • Recognizes scaling can worsen a shared-dependency bottleneck
  • Turns the fix into a review policy

Deep Dive 2: The Silent Auth Bypass#

Context: A penetration test finds that GET /internal/admin/users is reachable from the internet without authentication. Investigation shows a team added the route with auth: none eight months ago "temporarily for testing," and the downstream service trusted the gateway to have authenticated the call.

Questions to Surface First:

  • How many other routes have auth: none, and which were approved?
  • Did the service perform any authorization itself?
  • Logs: was the route accessed by anyone other than the testers?
  • Why does the pipeline allow auth: none without an exception?

Typical L5 Approach: Removes the route, adds auth, closes the finding.

Staff Approach: Fixes the class. Makes unauthenticated routes require an exception ID with security approval and an expiry; adds a CI test enumerating all public routes against the exception list; requires services to authorize using the verified identity (a request without a valid internal identity assertion is rejected by the service, regardless of gateway config); separates internal/admin routes onto a hostname not exposed to the internet at all.

Principal Approach: Treats "services trusted the gateway" as an organizational design flaw. Writes the zero-trust principle into the gateway standard: the gateway is a first check, never the only one. Makes the public-route inventory a quarterly security review artifact.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateBlock the route at the gateway; check logs for exposure.
TriageInventory all auth: none routes; classify intended vs accidental.
Quick fixException-required policy in the pipeline.
GuardrailsServices reject missing/invalid identity assertions; admin routes on internal-only hostname.
Post-mortemSingle layer of defense; "temporary" config with no expiry.

Metrics to Watch: routes_unauthenticated_total, exceptions_expiring_30d, service_rejected_missing_identity_total.

Organizational Follow-up: Quarterly public-route review with security; expiry on all exceptions.

Ownership Question: "Whose vulnerability was this?" Staff answer: The team's config, the platform's missing guardrail, and the service's missing authorization. Defense in depth means each should have caught it — the platform's fix is the one that makes the whole class impossible.

Key Takeaway: "The gateway is the first lock, not the only one."

What clears the Staff bar:

  • Fixes the class with a pipeline policy
  • Requires service-side authorization
  • Adds expiry to exceptions

Deep Dive 3: Onboarding the Partner API Product#

Context: The business is launching a public partner API with three plans (Free, Pro, Enterprise), usage-based billing, and a promise of 12 months' notice before breaking changes. Today all traffic goes through one first-party gateway.

Questions to Surface First:

  • Must partner traffic be isolated from first-party traffic? (Yes — different risk profile.)
  • How accurate must usage counts be for billing?
  • How will versions be expressed and tracked per consumer?
  • Who is the product owner for the API contract?

Typical L5 Approach: Adds API-key auth and per-key rate limits to the existing gateway; uses gateway counters for billing.

Staff Approach: Separate hostname (api.partners.example.com) and separate gateway deployment on the same platform. API keys/OAuth client credentials mapped to plans; rate limits for protection enforced locally; billing from an exact metering pipeline built on access logs (idempotent aggregation per request ID) — gateway counters are approximate by design. Path versioning (/v1, /v2) with Deprecation/Sunset headers, per-consumer version usage dashboards, and a developer portal. Contract tests in CI block breaking changes to published versions.

Principal Approach: Establishes that the partner API is a product with an owner, SLAs, and an on-call — not a set of routes. Decides whether to use a managed API-management layer for keys, plans and portal (commodity) while keeping the gateway runtime shared, and sets the deprecation policy as a company commitment that legal and sales understand.

Staff Approach — Full Reasoning
PhaseWhat to Do
IsolationSeparate deployment and capacity pool; per-plan concurrency caps.
IdentityKeys for Free, OAuth client credentials for Pro/Enterprise.
MeteringAccess logs → stream → per-consumer counts, reconciled daily.
VersioningContract tests; deprecation headers; usage per version per consumer.
LaunchPrivate beta with 5 partners; shadow metering compared to billing.

Metrics to Watch: partner_requests_by_plan, metering_vs_gateway_count_delta, requests_by_version{consumer}, partner_5xx_rate.

Organizational Follow-up: API product owner; deprecation policy approved by legal; support rotation.

Ownership Question: "Who decides when /v1 is turned off?" Staff answer: The API product owner, following the published policy — 12 months after the deprecation announcement, and only after usage by paying customers is below an agreed threshold or those customers have been contacted. Engineering provides the data; the product owner makes the call.

Key Takeaway: "A public API is a product contract; the gateway enforces it, but someone has to own it."

What clears the Staff bar:

  • Isolates partner traffic on a shared platform
  • Separates exact metering from approximate enforcement
  • Designs versioning and deprecation as first-class

Deep Dive 4: Post-Mortem — The Global Outage From a Gateway Upgrade#

Context: The platform team upgraded the gateway runtime to a new minor version. It was tested in staging and rolled to all regions over 20 minutes. A change in default header-size limits caused ~12% of requests (those with large cookies from a marketing tool) to be rejected with 431 errors. It took 2 hours to identify because error dashboards were aggregated across all routes and the error rate looked like "elevated but not critical."

Questions to Surface First:

  • Why did a 20-minute global rollout pass without detecting a 12% error rate?
  • Why didn't staging catch it — what traffic does staging see?
  • Were the health gates measuring the right things (status codes including 4xx)?
  • Why did detection take 2 hours?

Typical L5 Approach: Raises the header limit, rolls forward, adds a staging test for large headers.

Staff Approach: Fixes the rollout. Runtime upgrades go through the same staged pipeline as config: canary cell with 1–5% of production traffic for at least an hour, health gates that include 4xx classes the gateway itself generates (431, 413, 400 from parsing), compared against the control cells. Adds shadow traffic replay (mirrored production requests to the new version, responses compared) before canary. Dashboards break out gateway-generated errors from upstream-generated errors.

Principal Approach: Treats runtime upgrades as the highest-risk change class for the platform, with their own error budget and change calendar (not during peak seasons). Asks how many teams would have noticed in minutes if they had per-route SLO alerts — and makes per-team route SLO dashboards part of the platform offering.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (during)Roll back runtime version; confirm 431s disappear.
TriageDiff default configs between versions; identify header limits.
Quick fixPin explicit limits in config rather than relying on defaults.
GuardrailsTraffic replay; longer canary; gateway-generated 4xx in health gates.
Post-mortemStaging lacks production traffic shape; aggregate dashboards hid the signal.

Metrics to Watch: gw_local_reply_total{code} (errors generated by the gateway itself), gw_version distribution, per-route error rate vs control cell.

Organizational Follow-up: Change calendar; explicit config for all security-relevant defaults; per-team route SLOs.

Ownership Question: "Who owns detecting a 12% error rate on some routes?" Staff answer: The platform's rollout gates should have caught it in the canary; route owners' SLO alerts should have caught it within minutes. Both were missing — the platform owns providing both.

Key Takeaway: "Upgrading the gateway is a deploy to every service at once. Canary it with real traffic and measure the errors the gateway itself generates."

What clears the Staff bar:

  • Distinguishes gateway-generated from upstream errors
  • Pins defaults explicitly
  • Extends the staged pipeline to runtime upgrades

Deep Dive 5: Multi-Region Expansion#

Context: The company is expanding to the EU with a new region. Leadership wants a single global API hostname, EU data residency for EU users, and no increase in latency for US users.

Questions to Surface First:

  • How is a user's home region determined — by account, by location, by token claim?
  • Which services are region-local (with residency requirements) and which are global?
  • What happens when an EU user travels to the US?
  • How does config roll out across three regions without correlated failure?

Typical L5 Approach: Deploys a gateway in the EU, uses GeoDNS to route EU IPs to it.

Staff Approach: Routes by account home region, not client IP: the token carries a home_region claim; any gateway can receive the request (nearest by anycast/latency DNS), and if the home region differs, forwards it to the home region's gateway over the backbone — so a traveling EU user's data stays in the EU. Region-local services are only routed within their region; global services (catalog) are served locally everywhere. Config staged per region; each region can absorb failover only for non-resident traffic.

Principal Approach: Aligns gateway routing with the company's data-residency architecture as a whole — identity, logging, analytics — and gets legal sign-off on the routing rules. Recognizes that "single global hostname" plus residency constraints means the gateway now enforces compliance, which raises its change-management bar.

Staff Approach — Full Reasoning
PhaseWhat to Do
RoutingNearest gateway terminates; forwards by home_region claim for resident services.
Service classesResident (region-pinned) vs global (served anywhere).
FailoverResident traffic fails within region; global traffic can fail over.
ConfigStaged per region, 1 region at a time with bake time.
ValidationSynthetic EU and US users from both geographies; assert data-plane region.

Metrics to Watch: gw_forwarded_cross_region_total, resident_request_served_outside_home_total (must be 0), per-region latency p99.

Organizational Follow-up: Residency routing rules reviewed by legal; region added to the change calendar.

Ownership Question: "If an EU request is served from a US service, who is accountable?" Staff answer: The gateway platform, because routing rules enforce residency — which is why residency violations are a metric with an alert, not an audit finding a year later.

Key Takeaway: "Route by where the user's data lives, not where the user is standing."

What clears the Staff bar:

  • Uses token claims for home-region routing
  • Separates resident and global services
  • Makes residency violations observable

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • State and defend a thin-gateway scope with a concrete scope test
  • Design a hot path with no synchronous dependencies: local JWT validation, local-first rate limits, fail-static config
  • Propagate identity safely: strip client headers, signed internal identity, mTLS, service-side authorization
  • Design self-service route manifests with prefix ownership, policy validation and staged rollout
  • Separate first-party and partner traffic, and meter partner usage accurately
  • Explain when aggregation belongs in a BFF, a GraphQL layer, a service, or the client
  • Plan runtime upgrades, certificate renewal and key rotation as the high-risk changes they are
  • Reason about multi-region routing, including residency

The Bar for This Question#

Mid-level (L4): Describes a reverse proxy that authenticates, rate limits and routes.

Senior (L5): Designs a functional gateway with auth, rate limiting, routing, logging and horizontal scaling. Tends to add aggregation and transformation. Misses hot-path dependencies, config blast radius, zero trust and ownership.

Staff+ (L6): Draws a scope boundary, removes hot-path dependencies, makes config self-service and staged, isolates client classes, propagates identity with defense in depth, puts aggregation in team-owned BFFs, and names owners for routes, policies and the platform. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "Your Gateway Is Becoming a Monolith, and You Did It One Reasonable Request at a Time"#

RequestWhy It Seemed ReasonableWhat It Cost
"Validate request bodies at the edge"Reject bad input early300 schemas in the gateway; schema changes need platform deploys
"Aggregate the home screen"Fewer mobile round tripsMobile features wait on the platform team
"Add feature flags to routing"Faster rolloutsBusiness logic in the shared path

The Staff position: Write the scope charter first; every addition must pass the scope test. Say no to good ideas that belong elsewhere.

Why this matters in interviews: Proposing to keep things out of the gateway is a stronger signal than proposing features to add.

10.2 "Most Gateway Outages Are Config Changes, Not Traffic"#

Gateways are stateless and scale horizontally; traffic spikes rarely kill them. Config pushes, certificate expiry, runtime upgrades and key rotation — changes applied to every instance — are the dominant causes of gateway-wide outages.

The Staff position: Invest in the change pipeline before the scaling story.

Why this matters in interviews: Candidates spend their time on autoscaling; interviewers are waiting to hear about config.

10.3 "The Gateway Should Not Authorize"#

Coarse checks (scopes, tenant match) at the gateway are fine. Resource-level authorization — "can this user edit this document?" — needs domain data the gateway doesn't have and shouldn't fetch. Gateways that try (calling a policy service per request with resource lookups) add a hot-path dependency and still get it wrong.

The Staff position: Authenticate and coarse-filter at the gateway; authorize in services, possibly using a shared policy library or engine they call.

Why this matters in interviews: It shows precision about authentication vs authorization and about where data lives.

10.4 "You Probably Don't Need GraphQL at the Gateway — You Need Owned BFFs"#

GraphQL solves flexible composition for many clients, at the cost of schema governance, query-cost control and new failure modes (expensive queries, N+1 fan-out). Many organizations with 2–3 first-party clients get the benefit more cheaply with one BFF per client.

The Staff position: BFFs first; federated GraphQL when the number of clients and the variety of their needs justify the governance.

Why this matters in interviews: It shows you evaluate a popular technology on cost and ownership, not fashion.

10.5 "Per-Request Pricing Makes Managed Gateways a Trap at Scale — and a Gift Below It"#

At low volume, a managed gateway eliminates a platform team. At tens of billions of requests per month, per-request pricing can exceed the cost of the team and the compute to self-host. The crossover moves with your volume, not with vendor features.

The Staff position: Choose managed early, keep manifests portable, and recompute the crossover yearly.

Why this matters in interviews: Framing build vs buy as a curve rather than a verdict is an L6/L7 signal.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

At L6, the gateway is a thin, reliable proxy with a safe change pipeline. At L7, it is the organizational boundary between the platform and every product team — the place where security policy, identity, API contracts, traffic governance and cost converge. The Principal decisions are about how many gateways the company should have, what the platform will and won't own, how 60 teams change it without coordination, which contracts (URL structure, identity format, versioning policy) are forever, and whether the company should rent or run the runtime. Get the boundary right and the gateway accelerates every team; get it wrong and it becomes the company's slowest-moving monolith and its largest correlated failure domain.

The Org-Level Fault Line#

Central API platform team vs domain-owned gateways.

OptionWhat It BuysWhat It CostsWho Pays
Central team owns the gateway and all routesConsistency, security review on everythingTicket queue; product velocity capped by one teamProduct teams
Each domain runs its own gatewayAutonomyDivergent auth, logging, rate limiting; security gaps; duplicated on-callSecurity, SRE, customers (inconsistent APIs)
Central platform, domain-owned routes and BFFs (the L7 default)Uniform runtime and policy; autonomy via self-service; blast radius by cell and teamPlatform must build self-service, governance and per-team observabilityPlatform headcount

🧭 Principal Move: "Centralize the mechanism and the policy; decentralize the routes and the composition. The platform team's success metric is how rarely product teams need to talk to it."

Cost Model#

Assumptions: self-hosted proxy instances ~8 vCPU at ~$30/vCPU-month; managed per-request pricing on the order of ~$1–3.50 per million requests; fully loaded engineer ~$300K/year. Order of magnitude only.

ScaleTrafficManaged OptionSelf-Hosted ComputeHeadcountOn-call
Startup~1B req/month (~400 RPS)~$1–4K/month~$1K + time~0.2 FTEShared
Growth~50B req/month (~20K RPS)~$50–175K/month~$10–25K/month3–5 FTE platformPlatform rotation
Large~1T req/month (~400K RPS)~$1–3.5M/month (list)~$100–250K/month10–20 FTE (runtime, pipeline, API product, security)24/7, cell-aware

The pattern: managed is cheapest at startup scale, self-hosted almost always wins at large scale, and the growth tier is where the decision is genuinely close — dominated by whether you already have a platform team and whether managed API-management features are load-bearing.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to ReverseWhy
Public URL structure and hostnamesOne-wayEvery client, partner, SDK and docURLs are forever
Public API versioning scheme and deprecation policyOne-wayPartner trust and contractsA promise to customers
Internal identity propagation formatOne-way-ishEvery service's auth codeStandardize early
Gateway runtime (vendor/proxy)Two-way if manifests are portableMonthsKeep the manifest your own format
Managed vs self-hostedTwo-wayMigration quarterRecompute yearly
Putting aggregation or business logic in the gatewayEasy in, very hard outQuarters of extractionSay no early

The Standard I'd Write#

RFC: API Gateway Platform Standard v1

Scope: All north-south HTTP/gRPC traffic entering production. East-west traffic is out of scope (see the Service Mesh standard).

Requirements:

  • All public hostnames MUST be served by the gateway platform. New standalone gateways require architecture review.
  • Routes MUST be declared in team-owned manifests under registered prefixes; changes MUST go through the validation and staged rollout pipeline. There is no manual path, including for the platform team.
  • Routes MUST require authentication unless an approved exception with an expiry exists.
  • The gateway MUST strip client-supplied identity headers and forward a signed internal identity over mTLS. Services MUST verify it and MUST perform resource-level authorization.
  • Gateway filters MUST NOT make synchronous network calls on the request path without platform architecture review, a latency budget and a defined failure mode.
  • The gateway MUST NOT contain client-specific aggregation or business logic; these belong in team-owned BFFs or services.
  • Partner and first-party traffic MUST run on isolated deployments.
  • Runtime upgrades, certificates and signing-key rotations MUST follow the staged change process with production canaries.

Exceptions: Filed with the platform team; security approval for auth exceptions; 6-month expiry.

Success metrics: Gateway availability ≥ 99.99%; p99 overhead ≤ 5ms; median route change lead time < 1 hour; zero cross-team config incidents; unauthenticated routes without exceptions = 0; number of non-platform gateways trending to zero.

What I'd Tell the VP#

"Every customer request goes through our API gateway, so it's both our strongest security control and a single point of failure. We're keeping it deliberately simple — it checks who you are, enforces limits and sends you to the right service — and moving everything else back to the teams that own it, so they can ship without waiting on us. Teams will change their own routes through an automated pipeline that tests every change on a small slice of traffic first; that's where most gateway outages come from today. We'll separate partner traffic from our apps so a partner's spike can't slow our customers down. At our current volume, running it ourselves is cheaper than the managed service by roughly a factor of five, and I'd like to revisit that decision in a year with real numbers."

Principal Interview Signals#

SignalWhat It Sounds Like
Defines the platform boundary"Centralize mechanism and policy; teams own routes and composition."
Treats change as the main risk"No side doors — the platform team ships through the same staged pipeline."
Prices build vs buy as a curve"Managed is right until ~20K RPS; beyond that, the team pays for itself."
Knows the forever decisions"URLs, versioning policy and identity format are one-way doors; the proxy isn't."
Measures the platform by independence"Success is product teams rarely needing to talk to us."

Staff answers that L7 interviewers find insufficient:

  • A well-designed gateway with no plan for the four other gateways the company already runs.
  • "We'll use managed" or "we'll self-host" without a volume-based cost comparison and a portability plan.
  • Self-service routes without a governance model for prefix ownership, exceptions and audit.

Appendices

Appendix A: Mechanics in Depth#

A.1 Request Pipeline Pseudocode#

handle(req):
  enforce_limits(req.headers_size ≤ 16KB, req.body_size ≤ route.max_body)
  req.id = req.header('x-request-id') or uuid()
  route = trie.match(req.host, req.path, req.method)         # in-memory snapshot
  if not route: return 404
  strip(req, ['x-identity', 'x-user-id', 'x-tenant', ...])   # never trust client identity
  if route.auth.required:
      claims = verify_jwt(req.bearer, jwks_cache)             # local; ~0.1ms
      if not claims or not scopes_ok(claims, route.auth.scopes): return 401/403
  if not local_bucket(route.limit_key(claims, req)).take(): return 429 + Retry-After
  deadline = min(route.timeout, platform_max)
  resp = upstream(route.cluster, req + signed_identity(claims), deadline,
                  retry = idempotent(req) and budget_ok(route.cluster))
  scrub(resp, internal_headers)
  emit(access_log, metrics, span)                             # async
  return resp

A.2 Route Matching#

StrategyCostRisk
Exact / prefix via radix treeO(path length)None
Regex routesPer-regex evaluation; backtracking riskCPU spikes, ReDoS
Header/weight-based (canary)SmallMisconfigured weights

Rule: prefix routes by default; regex requires review and a linter check for catastrophic backtracking.

A.3 Connection Management#

  • Client side: HTTP/2 multiplexing, TLS session resumption, idle timeouts (e.g., 60s), max connection age to rebalance.
  • Upstream side: pooled keep-alive connections per cluster, capped per node (autoscaling multiplies pools — cap total upstream connections).
  • Deploys: drain long-lived connections (HTTP/2 GOAWAY, WebSocket close with retry hints) over a window, not all at once.

Appendix B: Identity and Data Model#

B.1 Identity Propagation#

LayerCredentialVerified By
Client → gatewayExternal JWT (15 min) or API key / client credentialsGateway
Gateway → servicemTLS workload identity + signed internal assertion (60s)Service
Service → servicemTLS workload identity (+ propagated end-user assertion if needed)Callee

B.2 Rate-Limit and Quota Keys#

TrafficKeyEnforcement
First-party authenticateduser_id + route classLocal-first, approximate
First-party anonymousIP / device fingerprint + route classLocal-first; CDN/WAF for volumetric
Partnerclient_id + plan + routeLocal-first for protection; exact metering for billing

See Rate Limiting for algorithms and drift bounds.

B.3 Route Manifest Fields#

FieldRequiredValidated Against
host, prefixYesPrefix ownership registry
upstreamYesService registry (Service Discovery)
authYes (default required)Exception list
timeout_msYes≤ platform max
retriesNoIdempotent methods only
rate_limitNo (defaults apply)Plan and platform ceilings
ownerYesTeam registry and on-call

Appendix C: Where Each Concern Lives#

Diagram: Appendix C: Where Each Concern Lives

C.1 Quick Comparison#

ConcernCDN/WAFGatewayBFFServiceMesh
Volumetric DDoS✓
Authentication✓verify identity
Resource authorization✓
Abuse rate limitingcoarse✓
Routing / canary✓internal
Aggregation✓
East-west retries/timeouts✓

Appendix D: API Contract and Client Behavior#

D.1 Standard Error Responses#

StatusWhenHeaders
400Malformed requestx-request-id
401Missing/invalid tokenWWW-Authenticate
403Valid token, insufficient scope—
404No route—
413 / 431Body / headers too large—
429Rate limit or quotaRetry-After, RateLimit-*
502 / 503 / 504Upstream failure / overload / timeoutRetry-After on 503

Distinguish gateway-generated errors from upstream errors in both metrics and a response header, so clients and dashboards can tell them apart.

D.2 Versioning and Deprecation#

GET /v1/orders/123
Deprecation: true
Sunset: Wed, 30 Sep 2027 00:00:00 GMT
Link: <https://developer.example.com/migrate-v2>; rel="deprecation"

Track requests per version per consumer; contact consumers directly before sunset.

D.3 Client Retry Guidance#

Clients retry only idempotent requests, on 503 with Retry-After or connection failure, with jittered backoff and a cap of 1–2 retries. Mobile SDKs should include an idempotency key on POSTs so the gateway/service can deduplicate (API Design).


Appendix E: Observability#

E.1 Core Metrics — Non-Negotiable#

gw_requests_total{route, code, origin=gateway|upstream}
gw_overhead_ms{quantile}                 # time spent in gateway, excluding upstream
gw_filter_duration_ms{filter}
gw_upstream_latency_ms{cluster}
gw_auth_failures_total{reason}           # expired, bad_sig, unknown_kid, scope
gw_ratelimit_rejections_total{route, key_class}
gw_config_version{node}                  # rollout visibility
gw_cert_expiry_seconds{hostname}
gw_upstream_cx_active{cluster}

E.2 Critical Alerts#

AlertConditionAction
Gateway-generated errorsgw_requests_total{origin=gateway, code=~"4xx or 5xx"} > 2× baseline for 5 minPage platform
Overhead regressiongw_overhead_ms p99 > 10ms for 10 minPage platform
Auth failure spikeunknown_kid or bad_sig > 1%Page platform + identity
Cert expiry< 14 daysTicket; page at < 7 days
Config skew> 1 config version live for > 1 hour outside a rolloutPage platform
Route SLO breachPer-route error or latency SLOPage route owner

E.3 Control Plane vs Data Plane#

The config server and pipeline can be down without affecting traffic — nodes serve the last signed snapshot. Alert on the control plane separately, and never let a node start without a valid snapshot (bake a last-known-good into the image or local disk).

E.4 Debugging "The Gateway Is Slow"#

  1. Overhead vs upstream: is the time inside the gateway?
  2. Per-filter timings: which filter?
  3. Does that filter make network calls?
  4. Regex routes or large route tables recently added?
  5. TLS: handshake rate spike (resumption broken, client reconnect storm)?

Appendix F: Scale Evolution#

F.1 What Works at Each Scale#

ScaleEnoughAdd When
< 20 servicesManaged gateway or single proxy config; JWT at edgeRoute ticket backlog
20–200 servicesSelf-service manifests, staged rollout, mTLS identityPartner API, client-class interference
200–1,000 servicesSeparate planes, BFF paved road, policy engineCorrelated outages, multi-region
1,000+ / multi-regionCells, residency routing, per-team SLO dashboards—

F.2 Multi-Region Path#

Single region → gateway per region with staged config → home-region routing via token claims → cells within regions.

F.3 What You Don't Build on Day One#

  • Your own proxy
  • GraphQL federation
  • A policy engine (start with manifest validation rules)
  • Per-tenant cells
  • Response caching at the gateway (use the CDN)

Appendix G: Multi-Tenancy, Fairness and Cost#

G.1 Isolation Levels#

LevelMechanismUse For
Per-key limitsLocal token bucketsEveryone
Per-class deploymentSeparate gateway fleetsFirst-party vs partner
Per-large-tenant cellDedicated capacity, shuffle shardingEnterprise partners

G.2 Cost Allocation#

Attribute gateway cost per route owner by request count and bytes; publish monthly. Hairpinned internal traffic and large file transfers through the gateway are usually the first things teams remove once they see the cost.

G.3 Tradeoff Summary#

ChoiceDefaultWho Pays If Wrong
ScopeThin, with a charterEveryone (monolith, blast radius)
Hot pathNo synchronous dependenciesAll traffic, when a dependency slows
AuthEdge authn + service authz + mTLSSecurity (lateral movement)
ConfigSelf-service, validated, stagedOther teams (cross-team outages)
AggregationBFFs owned by client teamsPlatform backlog; client velocity
RuntimeManaged early, self-host at scale, portable manifestsFinance (per-request pricing) or platform (ops)
  1. Loading the index…