Why This Matters#
API design is not a question about REST vs gRPC. It is a question about contracts you cannot take back. The moment an endpoint ships to an external client, a mobile app pinned to an old version, or a sibling team's batch job, its shape becomes load-bearing for people you will never meet. Every field name, every status code, every pagination cursor is a promise with an indefinite expiry date.
Most candidates treat the API step of an interview as a formality: sketch POST /orders, GET /orders/{id}, move on. Staff candidates use those 90 seconds to make three decisions the rest of the design depends on — what is idempotent, what is paginated, and what is asynchronous — because those three choices determine the retry behavior of every client, the load shape on every database, and the failure semantics the on-call will debug at 3am.
The Principal version goes one step further: an API is an organizational boundary. The API between the payments team and the checkout team is the org chart rendered in protobuf. Whoever owns the contract owns the deprecation schedule, the error taxonomy, and the pager when a client misuses it. Get the API wrong and you have not built a bad endpoint — you have built a permanent coordination tax between teams.
The 60-Second Version#
- An API is a contract, not an implementation. Internal refactors are two-way doors; public field removals are one-way doors with a 12–24 month deprecation tail.
- REST for resources crossing org or company boundaries, gRPC for service-to-service hot paths, GraphQL for client-driven aggregation over many backends. Pick by who the consumer is, not by what's fashionable.
- Every mutating endpoint that a client might retry needs an idempotency key. Networks drop ~0.1–1% of responses at the edge; without idempotency, those become duplicate charges, duplicate orders, duplicate emails.
- Cursor pagination by default. Offset pagination is O(offset) on the database and skips/duplicates rows under concurrent writes.
OFFSET 1000000on Postgres reads and discards a million rows. - Anything that can take > ~2s is a long-running operation. Return
202 Acceptedwith an operation resource; never hold an HTTP connection open for 45 seconds behind a load balancer with a 60s idle timeout. - Errors are part of the contract. A stable, machine-readable error code (
INSUFFICIENT_FUNDS, not "Error 500: something went wrong") tells the client whether to retry. Retryable vs non-retryable is the most important bit in the response. - Additive change only. Add optional fields freely; never rename, never repurpose, never tighten validation on an existing field. Versioning is the escape hatch, not the plan.
How API Design Works#
The Basic Idea#
An API defines four things: the nouns (resources or messages), the verbs (operations), the guarantees (idempotency, ordering, consistency, latency), and the failure vocabulary (errors, retry semantics, rate limits). Candidates spend 95% of their time on the first two. Production incidents live almost entirely in the last two.
| Layer | L5 Focus | What Actually Breaks in Production |
|---|---|---|
| Nouns | Resource naming, URL shape | Leaking internal IDs (auto-increment) that enable enumeration and couple clients to one database |
| Verbs | GET/POST/PUT/DELETE mapping | Non-idempotent POSTs retried by clients, proxies, and SDKs |
| Guarantees | Rarely discussed | Clients assuming read-after-write on an eventually consistent read path |
| Failure vocabulary | "Return 500 on error" | Clients retrying non-retryable errors, amplifying an outage 3–10× |
Key Terms#
| Term | Meaning | Why It Matters |
|---|---|---|
| Idempotency | Same request applied N times has the same effect as once | Makes retries safe; the foundation of at-least-once delivery at the edge |
| Idempotency key | Client-generated unique ID attached to a mutating request | Server dedupes on it for a retention window (Stripe: 24 hours) |
| Cursor | Opaque token encoding a position in a stable sort order | O(log n) seek instead of O(offset) scan; stable under inserts |
| Long-running operation (LRO) | An async job represented as a pollable resource | Decouples request lifetime from work lifetime |
| Field mask | Client-specified subset of fields to return or update | Prevents over-fetching and accidental overwrites on partial updates |
| ETag / If-Match | Version token for optimistic concurrency | Prevents lost updates on read-modify-write over HTTP |
| Breaking change | Any change that can cause a correct existing client to fail | Includes removing fields, adding required fields, changing enum semantics, tightening validation |
Where API Contracts Live#
| Boundary | Typical Protocol | Consumer Count | Change Cost | Who Owns the Contract |
|---|---|---|---|---|
| Public / partner API | REST + JSON, OpenAPI spec | 100s–100,000s of external devs | Months to years (you cannot force upgrades) | Product API team + developer relations |
| Mobile / web client API | REST or GraphQL | Every installed app version (often 6–18 months of versions in the wild) | Weeks to months (app store adoption curves) | Client platform / BFF team |
| Internal service-to-service | gRPC + protobuf | 5–50 internal teams | Days to weeks (can coordinate, but rarely do) | Owning service team |
| Event / async contracts | Avro / protobuf on Kafka | Every downstream consumer, including ones you do not know exist | Very high — events are replayed for years | Producing team + schema registry owners |
🎯 Staff Insight: "The protocol choice matters less than the consumer count. A gRPC API with 40 internal consumers is harder to change than a REST API with 3. I size change cost by who I'd have to coordinate with, not by what the wire format is."
Core Strategies#
Strategy 1: REST (Resource-Oriented HTTP)#
Model the domain as resources with stable URLs; use HTTP verbs for standard operations and custom methods (POST /orders/{id}:cancel) for everything else.
GET /v1/accounts/{account_id}/orders?page_size=50&page_token=abc # list
GET /v1/orders/{order_id} # read
POST /v1/orders Idempotency-Key: 7f3c... # create
PATCH /v1/orders/{order_id} If-Match: "v17" (field mask in body) # partial update
POST /v1/orders/{order_id}:cancel Idempotency-Key: 91ab... # custom verb
DELETE /v1/orders/{order_id} # idempotent by definition
When to use: anything crossing a company boundary or consumed by browsers. Universal tooling, cacheable GETs (CDNs honor Cache-Control), debuggable with curl, self-describing enough for a partner to integrate in an afternoon.
Failure mode: chatty clients. A mobile screen that needs 7 resources makes 7 round trips; at 150ms RTT on cellular that is ~1 second of waterfall before first render. The fix is a composite endpoint or a BFF, not GraphQL by reflex.
Strategy 2: gRPC (RPC over HTTP/2 + Protobuf)#
Define services and messages in .proto; generate typed clients in every language. Binary encoding, multiplexed streams, deadlines propagated across hops.
service OrderService {
rpc CreateOrder(CreateOrderRequest) returns (Order); // request_id field = idempotency key
rpc ListOrders(ListOrdersRequest) returns (ListOrdersResponse); // page_token / next_page_token
rpc WatchOrder(WatchOrderRequest) returns (stream OrderEvent); // server streaming
}
message CreateOrderRequest {
string request_id = 1; // client-generated UUID, required for retries
string account_id = 2;
repeated LineItem items = 3;
// field 4 reserved: was 'coupon_code', removed 2025-03. NEVER reuse the number.
reserved 4;
}
When to use: internal service-to-service calls on hot paths. Protobuf payloads are typically 3–10× smaller than equivalent JSON and 5–20× faster to serialize; deadline propagation (grpc-timeout) prevents a 30s downstream wait from outliving a 2s upstream budget.
Failure mode: field-number reuse. Reusing a deleted field number means an old client's bytes get decoded as the new field's type — silent data corruption, no error. Mark removed fields reserved. Second failure mode: HTTP/2 long-lived connections defeat L4 load balancers — one backend gets all the streams. You need L7 (client-side or proxy) load balancing.
Strategy 3: GraphQL (Client-Specified Query Over a Typed Graph)#
The client sends a query describing exactly the shape it wants; a gateway resolves each field against backend services.
query OrderScreen($id: ID!) {
order(id: $id) {
id status total { amount currency }
items(first: 20) { edges { node { sku name thumbnail(size: SMALL) } } pageInfo { endCursor hasNextPage } }
shipment { carrier eta }
}
}
When to use: many client surfaces (iOS, Android, web, TV) with divergent data needs over many backend services, where the aggregation layer would otherwise become a dozen hand-built BFF endpoints. Facebook built it for exactly this: mobile News Feed over dozens of data sources.
Failure mode: unbounded query cost. A client can ask for friends { friends { friends { posts } } } and fan out to millions of resolver calls. You need query cost analysis (depth limits ~10, complexity budgets per request), persisted queries in production, and the N+1 resolver problem solved with batching (DataLoader-style). Second failure mode: HTTP caching is nearly useless — everything is POST /graphql.
The Protocol Tradeoff#
| Protocol | What Works | What Breaks | Who Pays |
|---|---|---|---|
| REST/JSON | Universal tooling, CDN-cacheable GETs, easy partner onboarding | Over/under-fetching, chatty mobile clients, loose typing drift | Mobile users pay in latency; client teams pay in glue code |
| gRPC | Typed contracts, 3–10× smaller payloads, deadlines, streaming | Browser support needs a proxy (gRPC-Web), L4 load balancing, opaque on the wire | Platform team pays for L7 LB and service mesh; debugging on-call pays in tooling |
| GraphQL | One round trip per screen, client autonomy, schema as a catalog | Query-cost DoS, N+1 resolvers, no HTTP caching, authorization per field | Gateway team owns every slow query any client writes; backend teams pay in unpredictable load |
🎯 Staff Move: "Public API is REST with OpenAPI, because partners integrate with curl and we need CDN-cacheable reads. Internally it's gRPC with deadlines. I'd only introduce GraphQL if we have 3+ client surfaces and a team willing to own query cost limits — otherwise it's a gateway nobody owns."
Idempotency, Pagination, and Long-Running Operations#
These three mechanics are the hard sub-problem. They are where "the API looks fine" becomes "we double-charged 4,000 customers."
Idempotency Keys#
The client generates a unique key per logical operation and sends it on every retry. The server records the key with the result; a repeat returns the stored result instead of re-executing.
handle_create(request):
key = request.header("Idempotency-Key")
if key is null: return 400 "IDEMPOTENCY_KEY_REQUIRED" # for money-moving endpoints
record = idem_store.get(scope=request.account_id, key=key)
if record:
if record.request_hash != hash(request.body):
return 422 "IDEMPOTENCY_KEY_REUSED" # same key, different payload = client bug
if record.state == IN_PROGRESS:
return 409 "REQUEST_IN_FLIGHT" # concurrent retry; client backs off
return record.stored_response # replay, do not re-execute
# atomic claim: INSERT ... ON CONFLICT DO NOTHING
if not idem_store.claim(account_id, key, hash(request.body), ttl=24h):
return 409 "REQUEST_IN_FLIGHT"
result = execute_business_logic(request) # same DB txn as the claim, ideally
idem_store.complete(account_id, key, result)
return result
The three rules that make it correct:
- Scope the key to the caller (account/tenant). A global namespace lets one tenant's collision affect another.
- Hash the payload. Same key + different body is a client bug; reject it rather than silently returning the first result.
- Write the idempotency record in the same transaction as the side effect when possible. If the charge commits and the record write fails, the retry charges again. When the side effect is external (card network, email provider), propagate the key downstream so they dedupe.
Retention window: 24 hours is the common production default (Stripe documents 24h). It must exceed the client's maximum retry horizon — if a mobile client retries queued requests after 3 days offline, 24h is too short.
Pagination#
| Approach | Query Cost | Stable Under Writes? | Random Access? | Use When |
|---|---|---|---|---|
| Offset/limit | O(offset + limit) — OFFSET 100000 scans 100K rows | No — inserts shift pages, causing skips and duplicates | Yes ("jump to page 40") | Admin UIs over < 10K rows |
| Keyset (seek) | O(log n + limit) via index | Yes | No | Default for large, append-heavy collections |
| Opaque cursor | Same as keyset; cursor encodes (sort_key, tiebreak_id) + filter hash | Yes | No | Public APIs — lets you change the implementation without breaking clients |
| Snapshot / consistent read | Keyset + read timestamp | Fully consistent point-in-time view | No | Exports, reconciliation, anything that must not miss a row |
# Keyset pagination — sort by (created_at DESC, id DESC) with a composite index
SELECT id, created_at, total FROM orders
WHERE account_id = $1
AND (created_at, id) < ($cursor_created_at, $cursor_id)
ORDER BY created_at DESC, id DESC
LIMIT 51; -- fetch page_size + 1 to compute has_next without COUNT(*)
next_page_token = base64(encrypt({created_at, id, filter_hash, expires_at}))
Rules: make cursors opaque (encrypted or signed) so clients cannot construct them and you can change the encoding; always include a unique tiebreaker (the ID) in the sort; cap page_size server-side (e.g., max 100–1,000); never return total counts on large collections by default — COUNT(*) on a 50M-row filtered set can take seconds.
Long-Running Operations#
If work can exceed ~2 seconds p99 — report generation, video transcoding, bulk imports, provisioning — do not hold the connection. Return an operation handle.
POST /v1/exports Idempotency-Key: e1d2...
→ 202 Accepted
Location: /v1/operations/op_8f2a
{ "name": "operations/op_8f2a", "done": false, "metadata": {"progress_pct": 0} }
GET /v1/operations/op_8f2a
→ 200 { "done": true, "response": { "download_url": "...", "expires_at": "..." } }
or
→ 200 { "done": true, "error": { "code": "QUOTA_EXCEEDED", "retryable": false } }
Why it matters: load balancers and proxies commonly idle-timeout at 60s (AWS ALB default); mobile networks drop connections on handoff; a client that times out and retries a synchronous 45s operation starts it again. The operation resource makes the work addressable — the client can poll, cancel, or resume after a crash. Offer webhooks or server streaming for completion, but keep polling as the fallback because webhooks fail silently when the receiver is down.
Visual Guide#
Choosing the Protocol#
Idempotent Retry Across a Dropped Response#
Long-Running Operation Lifecycle#
Versioning and Errors#
Versioning: Plan to Not Need It#
| Change | Breaking? | Why |
|---|---|---|
| Add optional response field | No (if clients ignore unknown fields) | Requires the "tolerant reader" rule — document it on day 1 |
| Add optional request field with safe default | No | Old clients get the old behavior |
| Add a new enum value | Often yes | Clients with exhaustive switch statements crash or mis-route on unknown values |
| Rename a field | Yes | Old clients stop receiving it |
| Make an optional field required | Yes | Old clients fail validation |
| Tighten validation (max length 255 → 100) | Yes | Previously valid requests now 400 |
| Change default sort order or pagination size | Yes (behavioral) | Clients silently get different data — the worst kind of break, because nothing errors |
| Change error code for an existing condition | Yes | Client retry logic keys off codes |
Versioning mechanisms:
| Mechanism | Example | Strength | Weakness |
|---|---|---|---|
| URL major version | /v1/, /v2/ | Obvious, routable at the gateway, cacheable | Coarse; v2 migrations take years |
| Header / date-pinned version | API-Version: 2025-06-01 | Fine-grained; each account pinned to its signup date (Stripe's public model) | Server must maintain version-transform layers for every past version |
| Protobuf package version | orders.v1, orders.v2 | Explicit in generated code | Two services during migration |
| No versioning, additive only | Internal gRPC | Zero overhead | Requires discipline and lint tooling (buf breaking) in CI |
🎯 Staff Move: "URL major version for the public API, but I'd treat a v2 as a failure of evolution discipline — it costs a year of running both. Day to day we make only additive changes, enforced by a breaking-change linter in CI, not by code review."
Errors: The Retry Contract#
The most important thing an error tells a client is whether to retry, and when.
| HTTP | gRPC | Meaning | Client Should |
|---|---|---|---|
| 400 | INVALID_ARGUMENT | Request is malformed | Never retry — fix the request |
| 401 / 403 | UNAUTHENTICATED / PERMISSION_DENIED | Auth failure | Refresh credentials once, then stop |
| 404 | NOT_FOUND | Resource absent | Do not retry (unless eventual consistency after create — document it) |
| 409 | ABORTED / ALREADY_EXISTS | Conflict, concurrent modification | Re-read and retry the read-modify-write |
| 412 | FAILED_PRECONDITION | ETag mismatch | Re-read, re-apply, retry |
| 422 | FAILED_PRECONDITION | Business rule violated (INSUFFICIENT_FUNDS) | Do not retry; surface to user |
| 429 | RESOURCE_EXHAUSTED | Rate limited | Retry after Retry-After seconds |
| 500 | INTERNAL | Server bug | Retry with backoff only if the operation is idempotent |
| 503 | UNAVAILABLE | Overloaded / transient | Retry with exponential backoff + jitter |
| 504 | DEADLINE_EXCEEDED | Timed out — outcome unknown | Retry only with the same idempotency key |
{
"error": {
"code": "INSUFFICIENT_FUNDS", // stable, machine-readable, documented, never changes
"message": "Account balance is 12.00 USD; charge requires 40.00 USD.", // human, may change
"retryable": false,
"request_id": "req_9f2c...", // for support tickets and log correlation
"details": [{"type": "balance", "available": "12.00", "currency": "USD"}]
}
}
Rules: the code string is part of the contract and never changes; the message is for humans and can; every response carries a request_id; never leak stack traces or internal hostnames; 504 must be documented as "outcome unknown", because the write may have committed.
Implementation Patterns#
Optimistic Concurrency with ETags#
Prevent lost updates when two clients read-modify-write the same resource.
GET /v1/documents/d1 → 200, ETag: "v42"
PATCH /v1/documents/d1 If-Match: "v42" { "title": "New" }
→ server: UPDATE documents SET title=$1, version=43 WHERE id='d1' AND version=42
→ rows_affected == 0 ? 412 Precondition Failed : 200, ETag: "v43"
Use for any resource edited by more than one actor (config, documents, settings). Conflict rates below ~5% make this far cheaper than locks.
Field Masks for Partial Updates#
PUT replaces the whole resource — a client built against v1 that doesn't know about a v2 field will null it out on every save. PATCH with an explicit field mask (update_mask: "title,description") updates only named fields and is forward-compatible.
Rate Limit Headers as Contract#
Return RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset (IETF draft) or X-RateLimit-*, and Retry-After on every 429. Well-behaved clients self-throttle; you shed 30–50% of retry-storm load just by telling them when to come back. See Rate Limiting.
Deadline Propagation#
Every inbound request carries a deadline; every outbound call gets min(remaining_budget − safety_margin, per-call timeout). Without it, a 2s user request can trigger a 30s downstream query that keeps running (and holding a connection) long after the user gave up. gRPC propagates grpc-timeout automatically; for HTTP, pass an explicit header.
Resource Naming and Opaque IDs#
Use prefixed, opaque, non-sequential IDs (ord_2Nf8x..., cus_...). Prefixes make logs and support tickets self-describing; randomness prevents enumeration attacks and decouples clients from your database's auto-increment (which you will outgrow when you shard — see Sharding & Partitioning).
Webhooks as the Outbound Contract#
Outbound events need the same discipline: sign payloads (HMAC with a rotating secret), include an event ID for receiver-side dedupe, deliver at-least-once with exponential backoff for ~3 days, and expose a "list events" API so receivers can reconcile what they missed.
The Numbers in Context#
| Number | Value | What It Means for Your Design |
|---|---|---|
| Response loss at the edge | ~0.1–1% of requests on mobile | At 10K writes/s that is 10–100 ambiguous outcomes per second — idempotency is mandatory, not optional |
| Idempotency key retention | 24 hours (common default) | Must exceed the client's maximum retry horizon; 24h × 10K writes/s ≈ 864M records — size the store |
| LB idle timeout | 60s (AWS ALB default) | Any operation that might exceed ~30s must be async |
| LRO threshold | p99 > ~2s | Above this, users and clients time out and retry, duplicating work |
| Offset pagination cost | Linear in offset | OFFSET 1,000,000 reads 1M rows; keyset stays ~1ms at any depth |
| Page size cap | 100–1,000 items | Prevents a single request from holding a DB connection for seconds |
| Protobuf vs JSON | 3–10× smaller, 5–20× faster encode | Matters at > 10K RPS internal fan-out; irrelevant for a 50 RPS admin API |
| Mobile app version tail | 6–18 months of versions in the wild | Any client-facing breaking change needs a deprecation window of at least that long |
| Public API deprecation | 12–24 months notice | Budget two years of running both versions when planning a v2 |
| GraphQL depth limit | ~7–10 levels | Beyond this, fan-out is almost always accidental or malicious |
| Retry amplification | 3 layers × 3 retries = 27× | Retries must happen at one layer only, usually the outermost client, with budgets |
How This Shows Up in Interviews#
Scenario 1: "Define the API for this system"#
This is the 2-minute Phase 2 of every case study. The L5 answer lists CRUD endpoints. The Staff answer lists them and annotates three things: which mutations carry idempotency keys, which lists are cursor-paginated, and which operations are async. "POST /rides takes an idempotency key because riders double-tap. GET /rides is cursor-paginated by (created_at, id). Driver matching is not in the request path — the create returns REQUESTED and the client subscribes for the match."
Scenario 2: "A client retried and the customer was charged twice" (Full Walkthrough)#
Step 1 — Locate the ambiguity. "A double charge means the client didn't know whether the first request succeeded. That happens on a timeout, a connection reset after commit, or a 5xx from a proxy. The server executed twice because nothing told it the second request was the same logical operation."
Step 2 — Fix the contract, not the client. "I'd require an Idempotency-Key on every money-moving endpoint — returning 400 without one — scoped per merchant account, retained 24 hours, with the payload hashed so a reused key with a different amount is rejected with 422."
Step 3 — Make the dedupe atomic with the side effect. "The claim is an INSERT ... ON CONFLICT DO NOTHING into an idempotency table inside the same transaction that writes the ledger entry. If the transaction rolls back, the key is free to retry. For the external card network call, I pass our key downstream as their idempotency key, so a retry after a crash between 'called PSP' and 'committed locally' is deduped at the PSP."
Step 4 — Handle the concurrent retry. "If the retry arrives while the first is still in flight, it sees IN_PROGRESS and gets a 409 with Retry-After: 1. It must not execute in parallel."
Step 5 — Detect, reconcile, and own it. "I'd emit idempotency.replay_count and idempotency.key_mismatch_count per client SDK version. A spike in mismatches means an SDK is generating keys per attempt instead of per operation — that's a client bug the SDK team owns. Nightly reconciliation against the PSP settlement file catches anything the key didn't. Payments platform owns the idempotency store and gets paged on its availability, because if it's down, we fail closed on charges."
Why this is a Staff answer: it treats the double charge as a contract gap, puts dedupe in the same transaction as the effect, extends it across the external boundary, and names the metric and owner for the residual risk.
Scenario 3: "The list endpoint is timing out for our biggest customer"#
Probe: offset pagination on a 40M-row tenant, plus a COUNT(*) for the total. Staff answer: switch to keyset with an opaque cursor over a (tenant_id, created_at, id) index; drop the exact total (return has_more or an estimated count from table stats); cap page size at 200. Name the migration: accept both page and page_token for one release, log callers still using page, notify them, then return 400 for deep offsets beyond 10K.
Scenario 4: "We need to rename a field in a public API"#
You don't rename. You add the new field, populate both, mark the old one deprecated in the spec and in a Deprecation / Sunset response header, measure usage of the old field per API key, reach out to the top consumers directly, and remove it only when usage is below an agreed threshold (e.g., < 0.1% of calls) — or never, if a top-10 customer still depends on it. The cost of carrying a dead field is ~zero; the cost of breaking a large customer is a churn conversation.
Advanced Patterns#
| Pattern | How It Works | When to Use |
|---|---|---|
| Backend-for-Frontend (BFF) | One aggregation service per client surface | 2–3 client types with different needs; cheaper than GraphQL governance |
| Persisted queries | Clients send a query hash; server only executes pre-registered queries | GraphQL in production — kills query-cost DoS and enables caching |
| Date-based version pinning | Each account pinned to its first-call date; transforms upgrade/downgrade payloads | Public APIs with many long-lived integrations |
| Request hedging | Send a duplicate read after p95 latency; take the first response | Idempotent reads with long tails; cap at ~5% extra load |
| Batch endpoints | POST /items:batchGet with up to N IDs | Replacing N+1 client loops; cap N at ~100–500 |
| Async request-reply | 202 + operation resource + webhook/stream | Anything with p99 > 2s |
| Consumer-driven contract tests | Consumers publish expectations; provider CI verifies them | Internal APIs with 5+ consumers |
| API linting in CI | Spectral / buf breaking rejects breaking diffs | Every org with more than a handful of API-owning teams |
The Principal Lens#
Why L7 Sees This Problem Differently#
At Staff level, an API is a well-designed interface for one system. At Principal level, the set of APIs is the organization's architecture — the org's ability to ship depends on whether 200 service contracts look alike, fail alike, and evolve alike. A company where every team invents its own pagination, error envelope, and auth header pays an invisible tax: every integration is bespoke, every SDK is hand-written, every incident starts with "what does this error mean?" The Principal question is not "is this API good?" but "what is the paved road that makes the next 500 APIs good by default, and what am I willing to not standardize?"
The Org-Level Fault Line#
Central API governance vs team autonomy. A central API review board produces consistency but becomes a 2–3 week bottleneck and gets routed around. Pure autonomy produces 40 dialects of pagination. The Principal resolution is automate the standard, review only the exceptions: a style guide encoded as linters (OpenAPI/Spectral, buf lint, buf breaking) in CI, generated SDKs, and human review reserved for public surface area and one-way-door decisions (new top-level resources, new auth models).
| Option | Consistency | Team Velocity | Who Pays |
|---|---|---|---|
| Central review board for every API | High | Low — 1–3 week queue | Product teams pay in cycle time; they route around it |
| Style guide as a doc | Low (unenforced) | High | Consumers pay in integration cost; SDK team hand-writes adapters |
| Linters + generated SDKs + review for public/one-way changes | High on mechanics, flexible on domain | High | Platform team pays ~2–4 engineers to own tooling |
Cost Model#
Assumptions: fully loaded engineer ≈ $300K/yr (~$25K/month); cloud prices rough list; "integration" = one team consuming one other team's API.
| Scale | API Surface | Governance Cost | Cost of No Standard | On-Call Impact |
|---|---|---|---|---|
| Startup (20 eng, 10 services) | ~10 internal APIs, 1 public | ~0.25 FTE writing a style guide ≈ $6K/mo | Low — everyone knows everyone | Shared on-call; errors debugged by the author |
| Growth (200 eng, 80 services) | ~80 internal, public API with 1K partners | 2 FTE platform (linters, SDK gen, gateway) ≈ $50K/mo + gateway ~$5–15K/mo | ~1–2 eng-weeks per new integration × ~100/yr ≈ $500K–1M/yr | Idempotency and error-code bugs cause ~1–2 Sev2s/quarter |
| Large (2,000 eng, 800 services) | ~800 internal, public API with 50K+ developers | 6–10 FTE API platform + DevRel ≈ $150–250K/mo | Integration tax scales with N² team pairs; a single breaking change can hit thousands of partners | Retry storms from inconsistent error semantics are a top-5 cause of cascading incidents |
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse |
|---|---|---|
| Public resource names and ID format | One-way | Every partner integration; effectively permanent |
Error code strings | One-way | Client retry logic breaks silently |
| Pagination token semantics (opaque vs structured) | One-way if structured | Clients parse what you exposed; opaque from day 1 is free |
| Idempotency key retention window (shortening) | One-way | Retries beyond the new window double-execute |
| Internal gRPC field additions | Two-way | Free |
| REST vs gRPC for a new internal service | Two-way (mostly) | A few weeks to add a second transport |
| Adopting GraphQL for client aggregation | Two-way early, one-way after ~6 months | Every client screen rewritten against it |
| Auth model for public API (API keys vs OAuth) | One-way | Every partner re-onboards |
The Standard I'd Write#
RFC-API-001: API Contract Standard
Scope: All synchronous APIs exposed across a team boundary (internal gRPC and external REST). Event contracts are covered by RFC-EVT-001.
Mandatory (MUST):
- Every non-GET mutating operation MUST accept an idempotency key (
Idempotency-Keyheader orrequest_idfield), scoped per caller, retained ≥ 24h.- Every list operation MUST use opaque cursor pagination with a server-enforced max page size.
- Every error MUST use the shared error model: stable
code,retryableboolean,request_id.- Operations with p99 > 2s MUST be exposed as long-running operations.
- Breaking changes MUST fail CI (
buf breaking/ OpenAPI diff); overrides require an approved exception.Recommended (SHOULD): propagate deadlines; support field masks on update; emit
Deprecation/Sunsetheaders ≥ 12 months before removal of public fields.Exceptions: filed as a 1-page doc to the API platform channel; decided within 5 business days; time-boxed to 2 quarters.
Success metrics: % of APIs passing lint (target 95% in 2 quarters); median new-integration time (target < 3 days); Sev2+ incidents attributed to retry/duplicate semantics (target −50% YoY).
What I'd Tell the VP#
"Every team currently invents its own API conventions, so each new integration between teams takes one to two engineer-weeks of translation, and we've had duplicate-charge incidents because retries aren't safe by default. I'm proposing a small API platform team — about three engineers — to encode our standards into automated checks and generated client libraries. That makes the right thing the default without adding a review committee. We expect integration time to drop from weeks to days and to eliminate an entire class of payment incidents. The main risk is teams seeing it as bureaucracy, so it's enforced by tooling, with a five-day exception process."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Treats APIs as an org-wide system | "The problem isn't this endpoint — it's that we have 40 pagination dialects. I'd fix that with a linter, not a meeting." |
| Prices the tradeoff | "A v2 costs us about two engineers for 18 months of dual-running. Additive evolution costs nothing. That's why v2 is the last resort." |
| Names the one-way doors | "ID format and error codes are forever. Transport choice internally is not — I'd spend my review time on the former." |
| Knows when not to standardize | "I would not mandate GraphQL. I'd mandate the error model and idempotency, and let teams pick transport within the paved road." |
| Designs the deprecation machine | "We need per-client usage telemetry on every field before we can deprecate anything. That's the prerequisite, not the afterthought." |
Staff answers that L7 interviewers find insufficient:
- "We'll version it as v2" — without pricing the dual-running cost or the partner migration campaign.
- "Each team should follow REST best practices" — with no enforcement mechanism, which in practice means no standard.
- "Add idempotency keys to this endpoint" — correct locally, but misses that the fix belongs in the gateway or framework so the next 200 endpoints get it for free.
🧭 Principal Move: "I'd put idempotency, pagination, and the error model into the service framework and the gateway, so a team that writes a new endpoint gets them without thinking. Standards that depend on every engineer remembering them are not standards."
Failure Modes & Operational Reality#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Duplicate side effects from retries | idempotency.replay_count flat while payments.duplicate_detected rises; customer complaints | Every retried mutation; money, emails, orders | Mandatory idempotency keys; same-txn dedupe; downstream key propagation | Owning service team; payments platform for money paths |
| Retry storm amplifying an outage | http.requests_per_sec rising while success_rate falls; retries/original ratio > 1 | Entire dependency chain; 3 layers × 3 retries = 27× | Retry budgets (≤ 10% extra), retry only at the edge, Retry-After on 429/503 | Platform (client libraries) |
| Deep offset pagination saturating DB | db.rows_examined / rows_returned > 100; slow query log with high OFFSET | One tenant degrades a shared cluster | Keyset cursors; cap offsets; per-tenant query budget | Owning service team |
| Silent breaking change | Client error rate by sdk_version jumps after deploy; no server errors | All clients of the changed field | Breaking-change linter; contract tests; staged rollout by client | API platform (tooling) + service owner |
| GraphQL query-cost explosion | graphql.resolver_calls_per_request p99 > 10K; gateway CPU | Gateway and every backend it fans out to | Persisted queries, complexity limits, per-client cost budgets | Gateway team |
| LRO operations orphaned | operations.pending_age_seconds p99 climbing; workers idle | Users waiting forever on "processing" | Lease-based workers, heartbeat timeouts, operation GC | Owning service team |
Failure Scenario: The Retry Storm That Wasn't a Traffic Spike#
t=0 Inventory service p99 rises from 40ms to 900ms (a slow query after an index was dropped).
t=+15s Checkout's HTTP client times out at 500ms, retries 3×. Gateway also retries 5xx 2×.
t=+30s Inventory sees 6× normal RPS. Connection pool exhausted. p99 → 5s.
t=+1min Mobile apps retry on their own after 10s. Effective amplification ≈ 18×.
t=+2min Inventory is 100% saturated; checkout success rate 12%. Dashboards show "traffic spike."
t=+8min On-call disables gateway retries via flag; restores index. Recovery at t=+14min.
Detection: retry_ratio = retried_requests / original_requests per edge — should be < 0.1; alert at > 0.5. Prevention: retry budgets in the shared client library (retry only if the retry rate over the last 10s is < 10% of requests), retries at exactly one layer, 503 + Retry-After from the overloaded service. Owner: the platform team that owns the client library — this is not fixable service-by-service.
In the Wild#
Stripe: Idempotency Keys and Date-Pinned Versions#
Stripe's public API accepts an Idempotency-Key header on POST requests, stores the result for 24 hours, and returns the saved response on replay — including errors — and rejects reuse of a key with different parameters. Separately, Stripe pins each account to the API version current at its first request and maintains a chain of version transformations so old integrations keep working while the API evolves.
Staff insight: Stripe made retry safety and backward compatibility properties of the platform, not of individual endpoints. In an interview, cite this when you argue the idempotency layer belongs in shared middleware.
Google: API Improvement Proposals (AIPs)#
Google publishes its API design guidance as AIPs — resource-oriented design, standard methods, page_token/next_page_token pagination, long-running operations as a standard Operation resource, and field masks for partial updates. These conventions are enforced across a very large number of Google Cloud APIs with lint tooling.
Staff insight: This is the Principal pattern in public: one written standard, machine-enforced, adopted across thousands of APIs. Borrowing AIP conventions in an interview ("LRO as an operation resource, AIP-151 style") signals you know the solved problems.
Facebook/Meta: GraphQL for Mobile#
Facebook created GraphQL (open-sourced 2015) because its mobile News Feed needed data from many backends and REST endpoints were either over-fetching over cellular networks or multiplying into bespoke endpoints per screen. Persisted queries — shipping query IDs rather than query text — became a standard production practice.
Staff insight: GraphQL solved a specific problem: many client surfaces, many backends, bandwidth-constrained networks, and a team large enough to own the gateway. Without those conditions, a BFF is cheaper.
Staff Calibration#
What Staff Engineers Say (That Seniors Don't)#
| Concept | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Protocol choice | "gRPC is faster, let's use it" | "REST for partners, gRPC internally with deadlines; the consumer decides the protocol" | "The paved road is gRPC internally; I'd spend governance effort on error models and idempotency, not transport debates" |
| Retries | "Client retries on failure" | "Retries only with idempotency keys, only at the edge, with a 10% retry budget" | "Retry policy lives in the shared client library so no team can build a 27× amplifier by accident" |
| Pagination | "?page=3&limit=50" | "Opaque keyset cursors over (created_at, id), max page 200, no totals" | "Pagination is in the API standard and linted; opaque cursors keep our storage migration a two-way door" |
| Versioning | "We'll ship v2" | "Additive-only with a breaking-change linter; v2 only for a model change" | "A v2 is ~2 FTE × 18 months of dual-running plus a partner campaign; I'd need a business case" |
| Errors | "Return 500 with a message" | "Stable codes with a retryable flag; 504 documented as outcome unknown" | "One error model org-wide so incident response doesn't start with 'what does this code mean'" |
| Ownership | "My team owns the endpoint" | "We own the contract, its deprecation timeline, and per-client usage telemetry" | "API ownership is registered in a catalog with SLOs; unowned APIs get a deprecation owner assigned" |
Why "Retries" separates levels
The L5 answer is not wrong — clients should retry transient failures. The gap is the system view: L6 sees that retries without idempotency cause duplicates and retries at every layer cause amplification. L7 sees that no amount of per-team discipline fixes this across 300 services; only a shared library with retry budgets does, and that library needs an owner.
Why "Versioning" separates levels
L5 treats versioning as the change mechanism. L6 treats it as the escape hatch and invests in additive evolution. L7 prices it: dual-running infrastructure, documentation, SDKs, support load, and the partner migration campaign — and asks whether the business outcome justifies that spend.
Common Interview Traps#
- Listing endpoints without guarantees.
POST /ordersis incomplete until you say what happens when it's retried. - Offset pagination on an unbounded collection. It works in the demo and fails for your largest customer first.
- Synchronous endpoints for slow work. "The endpoint generates the report and returns it" — for a 90-second report behind a 60s LB timeout.
- Leaking auto-increment IDs. Enumerable, reveals business volume, and couples clients to a single-database ID generator you'll replace when you shard.
- GraphQL as a default. Choosing it without mentioning query cost limits, N+1 batching, or who owns the gateway.
- Treating 504 as failure. A timed-out write may have committed. Clients must retry with the same idempotency key, not a new one.
- "We'll just version it." Without naming the dual-running cost and the deprecation telemetry you'd need.
Practice Drill#
Prompt: "Design the API for a bulk import feature: enterprise customers upload CSVs of up to 5 million contacts. The current implementation is a synchronous
POST /contacts/importthat times out for anything over ~50K rows."
Staff Answer
I'd split upload from processing and make processing a long-running operation. Step 1: POST /v1/imports with an Idempotency-Key returns a pre-signed upload URL for object storage — the file never passes through our API servers, so a 2 GB CSV doesn't pin a web worker. Step 2: POST /v1/imports/{id}:start returns 202 with operations/{op_id}; workers process in chunks of ~10K rows, each chunk upserted by a natural key (email within the tenant) so a crashed worker can resume without duplicates. GET /v1/operations/{op_id} returns progress_pct, rows_processed, rows_failed, and a link to a paginated error report (GET /v1/imports/{id}/errors?page_token=) rather than one giant error payload. Completion fires a signed webhook, with polling as the fallback. Per-tenant concurrency is capped at 2 active imports and a row-rate budget (~5K rows/s) so one customer's 5M-row import doesn't starve the shared database. Metrics: import.rows_per_sec, import.operation_age_p99, import.chunk_retry_count. The contacts team owns the operation lifecycle; the platform owns the LRO framework and the upload service.
Why this is L6:
- Separates transport (upload) from work (processing) and uses the LRO pattern instead of stretching timeouts.
- Makes chunk processing idempotent via natural-key upserts, so retries and crashes are safe.
- Names per-tenant fairness limits, metrics, and ownership.
What L7 adds:
- Notes that "bulk import" is the fourth team building upload + LRO + error reports, and proposes a shared bulk-operations framework with a standard operation resource.
- Prices it: shared framework ≈ 2 engineers × 1 quarter, versus each team spending ~1 quarter building its own.
- Makes the error-report format and operation resource part of the org's API standard so customers learn one pattern across every bulk feature.
Where This Appears#
- API Gateway — where idempotency, auth, rate-limit headers, and version routing are enforced centrally
- Payment Processing — idempotency keys end to end, including downstream to card networks
- Rate Limiting —
429+Retry-Afteras a contract, and retry storms - URL Shortener — opaque ID design and read-path caching of GETs
- Distributed Job Scheduler — long-running operations and job status resources
- Real-Time Updates — when polling an operation should become a push stream
- Managing Long-Running Processes — the async half of the LRO pattern
- Dealing with Contention — ETags and optimistic concurrency behind
If-Match
Related Technologies: API Gateways · Apache Kafka · PostgreSQL