Hiring BarSupport

API Design Patterns

Foundation31 min read4 diagrams

Why This Matters#

API design is not a question about REST vs gRPC. It is a question about contracts you cannot take back. The moment an endpoint ships to an external client, a mobile app pinned to an old version, or a sibling team's batch job, its shape becomes load-bearing for people you will never meet. Every field name, every status code, every pagination cursor is a promise with an indefinite expiry date.

Most candidates treat the API step of an interview as a formality: sketch POST /orders, GET /orders/{id}, move on. Staff candidates use those 90 seconds to make three decisions the rest of the design depends on — what is idempotent, what is paginated, and what is asynchronous — because those three choices determine the retry behavior of every client, the load shape on every database, and the failure semantics the on-call will debug at 3am.

The Principal version goes one step further: an API is an organizational boundary. The API between the payments team and the checkout team is the org chart rendered in protobuf. Whoever owns the contract owns the deprecation schedule, the error taxonomy, and the pager when a client misuses it. Get the API wrong and you have not built a bad endpoint — you have built a permanent coordination tax between teams.

The 60-Second Version#

  • An API is a contract, not an implementation. Internal refactors are two-way doors; public field removals are one-way doors with a 12–24 month deprecation tail.
  • REST for resources crossing org or company boundaries, gRPC for service-to-service hot paths, GraphQL for client-driven aggregation over many backends. Pick by who the consumer is, not by what's fashionable.
  • Every mutating endpoint that a client might retry needs an idempotency key. Networks drop ~0.1–1% of responses at the edge; without idempotency, those become duplicate charges, duplicate orders, duplicate emails.
  • Cursor pagination by default. Offset pagination is O(offset) on the database and skips/duplicates rows under concurrent writes. OFFSET 1000000 on Postgres reads and discards a million rows.
  • Anything that can take > ~2s is a long-running operation. Return 202 Accepted with an operation resource; never hold an HTTP connection open for 45 seconds behind a load balancer with a 60s idle timeout.
  • Errors are part of the contract. A stable, machine-readable error code (INSUFFICIENT_FUNDS, not "Error 500: something went wrong") tells the client whether to retry. Retryable vs non-retryable is the most important bit in the response.
  • Additive change only. Add optional fields freely; never rename, never repurpose, never tighten validation on an existing field. Versioning is the escape hatch, not the plan.

How API Design Works#

The Basic Idea#

An API defines four things: the nouns (resources or messages), the verbs (operations), the guarantees (idempotency, ordering, consistency, latency), and the failure vocabulary (errors, retry semantics, rate limits). Candidates spend 95% of their time on the first two. Production incidents live almost entirely in the last two.

LayerL5 FocusWhat Actually Breaks in Production
NounsResource naming, URL shapeLeaking internal IDs (auto-increment) that enable enumeration and couple clients to one database
VerbsGET/POST/PUT/DELETE mappingNon-idempotent POSTs retried by clients, proxies, and SDKs
GuaranteesRarely discussedClients assuming read-after-write on an eventually consistent read path
Failure vocabulary"Return 500 on error"Clients retrying non-retryable errors, amplifying an outage 3–10×

Key Terms#

TermMeaningWhy It Matters
IdempotencySame request applied N times has the same effect as onceMakes retries safe; the foundation of at-least-once delivery at the edge
Idempotency keyClient-generated unique ID attached to a mutating requestServer dedupes on it for a retention window (Stripe: 24 hours)
CursorOpaque token encoding a position in a stable sort orderO(log n) seek instead of O(offset) scan; stable under inserts
Long-running operation (LRO)An async job represented as a pollable resourceDecouples request lifetime from work lifetime
Field maskClient-specified subset of fields to return or updatePrevents over-fetching and accidental overwrites on partial updates
ETag / If-MatchVersion token for optimistic concurrencyPrevents lost updates on read-modify-write over HTTP
Breaking changeAny change that can cause a correct existing client to failIncludes removing fields, adding required fields, changing enum semantics, tightening validation

Where API Contracts Live#

BoundaryTypical ProtocolConsumer CountChange CostWho Owns the Contract
Public / partner APIREST + JSON, OpenAPI spec100s–100,000s of external devsMonths to years (you cannot force upgrades)Product API team + developer relations
Mobile / web client APIREST or GraphQLEvery installed app version (often 6–18 months of versions in the wild)Weeks to months (app store adoption curves)Client platform / BFF team
Internal service-to-servicegRPC + protobuf5–50 internal teamsDays to weeks (can coordinate, but rarely do)Owning service team
Event / async contractsAvro / protobuf on KafkaEvery downstream consumer, including ones you do not know existVery high — events are replayed for yearsProducing team + schema registry owners

🎯 Staff Insight: "The protocol choice matters less than the consumer count. A gRPC API with 40 internal consumers is harder to change than a REST API with 3. I size change cost by who I'd have to coordinate with, not by what the wire format is."

Core Strategies#

Strategy 1: REST (Resource-Oriented HTTP)#

Model the domain as resources with stable URLs; use HTTP verbs for standard operations and custom methods (POST /orders/{id}:cancel) for everything else.

GET    /v1/accounts/{account_id}/orders?page_size=50&page_token=abc   # list
GET    /v1/orders/{order_id}                                           # read
POST   /v1/orders            Idempotency-Key: 7f3c...                  # create
PATCH  /v1/orders/{order_id} If-Match: "v17"   (field mask in body)    # partial update
POST   /v1/orders/{order_id}:cancel  Idempotency-Key: 91ab...          # custom verb
DELETE /v1/orders/{order_id}                                           # idempotent by definition

When to use: anything crossing a company boundary or consumed by browsers. Universal tooling, cacheable GETs (CDNs honor Cache-Control), debuggable with curl, self-describing enough for a partner to integrate in an afternoon.

Failure mode: chatty clients. A mobile screen that needs 7 resources makes 7 round trips; at 150ms RTT on cellular that is ~1 second of waterfall before first render. The fix is a composite endpoint or a BFF, not GraphQL by reflex.

Strategy 2: gRPC (RPC over HTTP/2 + Protobuf)#

Define services and messages in .proto; generate typed clients in every language. Binary encoding, multiplexed streams, deadlines propagated across hops.

service OrderService {
  rpc CreateOrder(CreateOrderRequest) returns (Order);          // request_id field = idempotency key
  rpc ListOrders(ListOrdersRequest) returns (ListOrdersResponse); // page_token / next_page_token
  rpc WatchOrder(WatchOrderRequest) returns (stream OrderEvent); // server streaming
}

message CreateOrderRequest {
  string request_id = 1;      // client-generated UUID, required for retries
  string account_id = 2;
  repeated LineItem items = 3;
  // field 4 reserved: was 'coupon_code', removed 2025-03. NEVER reuse the number.
  reserved 4;
}

When to use: internal service-to-service calls on hot paths. Protobuf payloads are typically 3–10× smaller than equivalent JSON and 5–20× faster to serialize; deadline propagation (grpc-timeout) prevents a 30s downstream wait from outliving a 2s upstream budget.

Failure mode: field-number reuse. Reusing a deleted field number means an old client's bytes get decoded as the new field's type — silent data corruption, no error. Mark removed fields reserved. Second failure mode: HTTP/2 long-lived connections defeat L4 load balancers — one backend gets all the streams. You need L7 (client-side or proxy) load balancing.

Strategy 3: GraphQL (Client-Specified Query Over a Typed Graph)#

The client sends a query describing exactly the shape it wants; a gateway resolves each field against backend services.

query OrderScreen($id: ID!) {
  order(id: $id) {
    id status total { amount currency }
    items(first: 20) { edges { node { sku name thumbnail(size: SMALL) } } pageInfo { endCursor hasNextPage } }
    shipment { carrier eta }
  }
}

When to use: many client surfaces (iOS, Android, web, TV) with divergent data needs over many backend services, where the aggregation layer would otherwise become a dozen hand-built BFF endpoints. Facebook built it for exactly this: mobile News Feed over dozens of data sources.

Failure mode: unbounded query cost. A client can ask for friends { friends { friends { posts } } } and fan out to millions of resolver calls. You need query cost analysis (depth limits ~10, complexity budgets per request), persisted queries in production, and the N+1 resolver problem solved with batching (DataLoader-style). Second failure mode: HTTP caching is nearly useless — everything is POST /graphql.

The Protocol Tradeoff#

ProtocolWhat WorksWhat BreaksWho Pays
REST/JSONUniversal tooling, CDN-cacheable GETs, easy partner onboardingOver/under-fetching, chatty mobile clients, loose typing driftMobile users pay in latency; client teams pay in glue code
gRPCTyped contracts, 3–10× smaller payloads, deadlines, streamingBrowser support needs a proxy (gRPC-Web), L4 load balancing, opaque on the wirePlatform team pays for L7 LB and service mesh; debugging on-call pays in tooling
GraphQLOne round trip per screen, client autonomy, schema as a catalogQuery-cost DoS, N+1 resolvers, no HTTP caching, authorization per fieldGateway team owns every slow query any client writes; backend teams pay in unpredictable load

🎯 Staff Move: "Public API is REST with OpenAPI, because partners integrate with curl and we need CDN-cacheable reads. Internally it's gRPC with deadlines. I'd only introduce GraphQL if we have 3+ client surfaces and a team willing to own query cost limits — otherwise it's a gateway nobody owns."

Idempotency, Pagination, and Long-Running Operations#

These three mechanics are the hard sub-problem. They are where "the API looks fine" becomes "we double-charged 4,000 customers."

Idempotency Keys#

The client generates a unique key per logical operation and sends it on every retry. The server records the key with the result; a repeat returns the stored result instead of re-executing.

handle_create(request):
    key = request.header("Idempotency-Key")
    if key is null: return 400 "IDEMPOTENCY_KEY_REQUIRED"      # for money-moving endpoints

    record = idem_store.get(scope=request.account_id, key=key)
    if record:
        if record.request_hash != hash(request.body):
            return 422 "IDEMPOTENCY_KEY_REUSED"                # same key, different payload = client bug
        if record.state == IN_PROGRESS:
            return 409 "REQUEST_IN_FLIGHT"                      # concurrent retry; client backs off
        return record.stored_response                          # replay, do not re-execute

    # atomic claim: INSERT ... ON CONFLICT DO NOTHING
    if not idem_store.claim(account_id, key, hash(request.body), ttl=24h):
        return 409 "REQUEST_IN_FLIGHT"

    result = execute_business_logic(request)                    # same DB txn as the claim, ideally
    idem_store.complete(account_id, key, result)
    return result

The three rules that make it correct:

  1. Scope the key to the caller (account/tenant). A global namespace lets one tenant's collision affect another.
  2. Hash the payload. Same key + different body is a client bug; reject it rather than silently returning the first result.
  3. Write the idempotency record in the same transaction as the side effect when possible. If the charge commits and the record write fails, the retry charges again. When the side effect is external (card network, email provider), propagate the key downstream so they dedupe.

Retention window: 24 hours is the common production default (Stripe documents 24h). It must exceed the client's maximum retry horizon — if a mobile client retries queued requests after 3 days offline, 24h is too short.

Pagination#

ApproachQuery CostStable Under Writes?Random Access?Use When
Offset/limitO(offset + limit) — OFFSET 100000 scans 100K rowsNo — inserts shift pages, causing skips and duplicatesYes ("jump to page 40")Admin UIs over < 10K rows
Keyset (seek)O(log n + limit) via indexYesNoDefault for large, append-heavy collections
Opaque cursorSame as keyset; cursor encodes (sort_key, tiebreak_id) + filter hashYesNoPublic APIs — lets you change the implementation without breaking clients
Snapshot / consistent readKeyset + read timestampFully consistent point-in-time viewNoExports, reconciliation, anything that must not miss a row
# Keyset pagination — sort by (created_at DESC, id DESC) with a composite index
SELECT id, created_at, total FROM orders
WHERE account_id = $1
  AND (created_at, id) < ($cursor_created_at, $cursor_id)
ORDER BY created_at DESC, id DESC
LIMIT 51;                         -- fetch page_size + 1 to compute has_next without COUNT(*)

next_page_token = base64(encrypt({created_at, id, filter_hash, expires_at}))

Rules: make cursors opaque (encrypted or signed) so clients cannot construct them and you can change the encoding; always include a unique tiebreaker (the ID) in the sort; cap page_size server-side (e.g., max 100–1,000); never return total counts on large collections by default — COUNT(*) on a 50M-row filtered set can take seconds.

Long-Running Operations#

If work can exceed ~2 seconds p99 — report generation, video transcoding, bulk imports, provisioning — do not hold the connection. Return an operation handle.

POST /v1/exports         Idempotency-Key: e1d2...
→ 202 Accepted
  Location: /v1/operations/op_8f2a
  { "name": "operations/op_8f2a", "done": false, "metadata": {"progress_pct": 0} }

GET /v1/operations/op_8f2a
→ 200 { "done": true, "response": { "download_url": "...", "expires_at": "..." } }
   or
→ 200 { "done": true, "error": { "code": "QUOTA_EXCEEDED", "retryable": false } }

Why it matters: load balancers and proxies commonly idle-timeout at 60s (AWS ALB default); mobile networks drop connections on handoff; a client that times out and retries a synchronous 45s operation starts it again. The operation resource makes the work addressable — the client can poll, cancel, or resume after a crash. Offer webhooks or server streaming for completion, but keep polling as the fallback because webhooks fail silently when the receiver is down.

Visual Guide#

Choosing the Protocol#

Diagram: Choosing the Protocol

Idempotent Retry Across a Dropped Response#

Diagram: Idempotent Retry Across a Dropped Response

Long-Running Operation Lifecycle#

Diagram: Long-Running Operation Lifecycle

Versioning and Errors#

Versioning: Plan to Not Need It#

ChangeBreaking?Why
Add optional response fieldNo (if clients ignore unknown fields)Requires the "tolerant reader" rule — document it on day 1
Add optional request field with safe defaultNoOld clients get the old behavior
Add a new enum valueOften yesClients with exhaustive switch statements crash or mis-route on unknown values
Rename a fieldYesOld clients stop receiving it
Make an optional field requiredYesOld clients fail validation
Tighten validation (max length 255 → 100)YesPreviously valid requests now 400
Change default sort order or pagination sizeYes (behavioral)Clients silently get different data — the worst kind of break, because nothing errors
Change error code for an existing conditionYesClient retry logic keys off codes

Versioning mechanisms:

MechanismExampleStrengthWeakness
URL major version/v1/, /v2/Obvious, routable at the gateway, cacheableCoarse; v2 migrations take years
Header / date-pinned versionAPI-Version: 2025-06-01Fine-grained; each account pinned to its signup date (Stripe's public model)Server must maintain version-transform layers for every past version
Protobuf package versionorders.v1, orders.v2Explicit in generated codeTwo services during migration
No versioning, additive onlyInternal gRPCZero overheadRequires discipline and lint tooling (buf breaking) in CI

🎯 Staff Move: "URL major version for the public API, but I'd treat a v2 as a failure of evolution discipline — it costs a year of running both. Day to day we make only additive changes, enforced by a breaking-change linter in CI, not by code review."

Errors: The Retry Contract#

The most important thing an error tells a client is whether to retry, and when.

HTTPgRPCMeaningClient Should
400INVALID_ARGUMENTRequest is malformedNever retry — fix the request
401 / 403UNAUTHENTICATED / PERMISSION_DENIEDAuth failureRefresh credentials once, then stop
404NOT_FOUNDResource absentDo not retry (unless eventual consistency after create — document it)
409ABORTED / ALREADY_EXISTSConflict, concurrent modificationRe-read and retry the read-modify-write
412FAILED_PRECONDITIONETag mismatchRe-read, re-apply, retry
422FAILED_PRECONDITIONBusiness rule violated (INSUFFICIENT_FUNDS)Do not retry; surface to user
429RESOURCE_EXHAUSTEDRate limitedRetry after Retry-After seconds
500INTERNALServer bugRetry with backoff only if the operation is idempotent
503UNAVAILABLEOverloaded / transientRetry with exponential backoff + jitter
504DEADLINE_EXCEEDEDTimed out — outcome unknownRetry only with the same idempotency key
{
  "error": {
    "code": "INSUFFICIENT_FUNDS",          // stable, machine-readable, documented, never changes
    "message": "Account balance is 12.00 USD; charge requires 40.00 USD.",   // human, may change
    "retryable": false,
    "request_id": "req_9f2c...",            // for support tickets and log correlation
    "details": [{"type": "balance", "available": "12.00", "currency": "USD"}]
  }
}

Rules: the code string is part of the contract and never changes; the message is for humans and can; every response carries a request_id; never leak stack traces or internal hostnames; 504 must be documented as "outcome unknown", because the write may have committed.

Implementation Patterns#

Optimistic Concurrency with ETags#

Prevent lost updates when two clients read-modify-write the same resource.

GET /v1/documents/d1        → 200, ETag: "v42"
PATCH /v1/documents/d1      If-Match: "v42"   { "title": "New" }
   → server: UPDATE documents SET title=$1, version=43 WHERE id='d1' AND version=42
   → rows_affected == 0 ?  412 Precondition Failed  :  200, ETag: "v43"

Use for any resource edited by more than one actor (config, documents, settings). Conflict rates below ~5% make this far cheaper than locks.

Field Masks for Partial Updates#

PUT replaces the whole resource — a client built against v1 that doesn't know about a v2 field will null it out on every save. PATCH with an explicit field mask (update_mask: "title,description") updates only named fields and is forward-compatible.

Rate Limit Headers as Contract#

Return RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset (IETF draft) or X-RateLimit-*, and Retry-After on every 429. Well-behaved clients self-throttle; you shed 30–50% of retry-storm load just by telling them when to come back. See Rate Limiting.

Deadline Propagation#

Every inbound request carries a deadline; every outbound call gets min(remaining_budget − safety_margin, per-call timeout). Without it, a 2s user request can trigger a 30s downstream query that keeps running (and holding a connection) long after the user gave up. gRPC propagates grpc-timeout automatically; for HTTP, pass an explicit header.

Resource Naming and Opaque IDs#

Use prefixed, opaque, non-sequential IDs (ord_2Nf8x..., cus_...). Prefixes make logs and support tickets self-describing; randomness prevents enumeration attacks and decouples clients from your database's auto-increment (which you will outgrow when you shard — see Sharding & Partitioning).

Webhooks as the Outbound Contract#

Outbound events need the same discipline: sign payloads (HMAC with a rotating secret), include an event ID for receiver-side dedupe, deliver at-least-once with exponential backoff for ~3 days, and expose a "list events" API so receivers can reconcile what they missed.

The Numbers in Context#

NumberValueWhat It Means for Your Design
Response loss at the edge~0.1–1% of requests on mobileAt 10K writes/s that is 10–100 ambiguous outcomes per second — idempotency is mandatory, not optional
Idempotency key retention24 hours (common default)Must exceed the client's maximum retry horizon; 24h × 10K writes/s ≈ 864M records — size the store
LB idle timeout60s (AWS ALB default)Any operation that might exceed ~30s must be async
LRO thresholdp99 > ~2sAbove this, users and clients time out and retry, duplicating work
Offset pagination costLinear in offsetOFFSET 1,000,000 reads 1M rows; keyset stays ~1ms at any depth
Page size cap100–1,000 itemsPrevents a single request from holding a DB connection for seconds
Protobuf vs JSON3–10× smaller, 5–20× faster encodeMatters at > 10K RPS internal fan-out; irrelevant for a 50 RPS admin API
Mobile app version tail6–18 months of versions in the wildAny client-facing breaking change needs a deprecation window of at least that long
Public API deprecation12–24 months noticeBudget two years of running both versions when planning a v2
GraphQL depth limit~7–10 levelsBeyond this, fan-out is almost always accidental or malicious
Retry amplification3 layers × 3 retries = 27×Retries must happen at one layer only, usually the outermost client, with budgets

How This Shows Up in Interviews#

Scenario 1: "Define the API for this system"#

This is the 2-minute Phase 2 of every case study. The L5 answer lists CRUD endpoints. The Staff answer lists them and annotates three things: which mutations carry idempotency keys, which lists are cursor-paginated, and which operations are async. "POST /rides takes an idempotency key because riders double-tap. GET /rides is cursor-paginated by (created_at, id). Driver matching is not in the request path — the create returns REQUESTED and the client subscribes for the match."

Scenario 2: "A client retried and the customer was charged twice" (Full Walkthrough)#

Step 1 — Locate the ambiguity. "A double charge means the client didn't know whether the first request succeeded. That happens on a timeout, a connection reset after commit, or a 5xx from a proxy. The server executed twice because nothing told it the second request was the same logical operation."

Step 2 — Fix the contract, not the client. "I'd require an Idempotency-Key on every money-moving endpoint — returning 400 without one — scoped per merchant account, retained 24 hours, with the payload hashed so a reused key with a different amount is rejected with 422."

Step 3 — Make the dedupe atomic with the side effect. "The claim is an INSERT ... ON CONFLICT DO NOTHING into an idempotency table inside the same transaction that writes the ledger entry. If the transaction rolls back, the key is free to retry. For the external card network call, I pass our key downstream as their idempotency key, so a retry after a crash between 'called PSP' and 'committed locally' is deduped at the PSP."

Step 4 — Handle the concurrent retry. "If the retry arrives while the first is still in flight, it sees IN_PROGRESS and gets a 409 with Retry-After: 1. It must not execute in parallel."

Step 5 — Detect, reconcile, and own it. "I'd emit idempotency.replay_count and idempotency.key_mismatch_count per client SDK version. A spike in mismatches means an SDK is generating keys per attempt instead of per operation — that's a client bug the SDK team owns. Nightly reconciliation against the PSP settlement file catches anything the key didn't. Payments platform owns the idempotency store and gets paged on its availability, because if it's down, we fail closed on charges."

Why this is a Staff answer: it treats the double charge as a contract gap, puts dedupe in the same transaction as the effect, extends it across the external boundary, and names the metric and owner for the residual risk.

Scenario 3: "The list endpoint is timing out for our biggest customer"#

Probe: offset pagination on a 40M-row tenant, plus a COUNT(*) for the total. Staff answer: switch to keyset with an opaque cursor over a (tenant_id, created_at, id) index; drop the exact total (return has_more or an estimated count from table stats); cap page size at 200. Name the migration: accept both page and page_token for one release, log callers still using page, notify them, then return 400 for deep offsets beyond 10K.

Scenario 4: "We need to rename a field in a public API"#

You don't rename. You add the new field, populate both, mark the old one deprecated in the spec and in a Deprecation / Sunset response header, measure usage of the old field per API key, reach out to the top consumers directly, and remove it only when usage is below an agreed threshold (e.g., < 0.1% of calls) — or never, if a top-10 customer still depends on it. The cost of carrying a dead field is ~zero; the cost of breaking a large customer is a churn conversation.

Advanced Patterns#

PatternHow It WorksWhen to Use
Backend-for-Frontend (BFF)One aggregation service per client surface2–3 client types with different needs; cheaper than GraphQL governance
Persisted queriesClients send a query hash; server only executes pre-registered queriesGraphQL in production — kills query-cost DoS and enables caching
Date-based version pinningEach account pinned to its first-call date; transforms upgrade/downgrade payloadsPublic APIs with many long-lived integrations
Request hedgingSend a duplicate read after p95 latency; take the first responseIdempotent reads with long tails; cap at ~5% extra load
Batch endpointsPOST /items:batchGet with up to N IDsReplacing N+1 client loops; cap N at ~100–500
Async request-reply202 + operation resource + webhook/streamAnything with p99 > 2s
Consumer-driven contract testsConsumers publish expectations; provider CI verifies themInternal APIs with 5+ consumers
API linting in CISpectral / buf breaking rejects breaking diffsEvery org with more than a handful of API-owning teams

The Principal Lens#

Why L7 Sees This Problem Differently#

At Staff level, an API is a well-designed interface for one system. At Principal level, the set of APIs is the organization's architecture — the org's ability to ship depends on whether 200 service contracts look alike, fail alike, and evolve alike. A company where every team invents its own pagination, error envelope, and auth header pays an invisible tax: every integration is bespoke, every SDK is hand-written, every incident starts with "what does this error mean?" The Principal question is not "is this API good?" but "what is the paved road that makes the next 500 APIs good by default, and what am I willing to not standardize?"

The Org-Level Fault Line#

Central API governance vs team autonomy. A central API review board produces consistency but becomes a 2–3 week bottleneck and gets routed around. Pure autonomy produces 40 dialects of pagination. The Principal resolution is automate the standard, review only the exceptions: a style guide encoded as linters (OpenAPI/Spectral, buf lint, buf breaking) in CI, generated SDKs, and human review reserved for public surface area and one-way-door decisions (new top-level resources, new auth models).

OptionConsistencyTeam VelocityWho Pays
Central review board for every APIHighLow — 1–3 week queueProduct teams pay in cycle time; they route around it
Style guide as a docLow (unenforced)HighConsumers pay in integration cost; SDK team hand-writes adapters
Linters + generated SDKs + review for public/one-way changesHigh on mechanics, flexible on domainHighPlatform team pays ~2–4 engineers to own tooling

Cost Model#

Assumptions: fully loaded engineer ≈ $300K/yr (~$25K/month); cloud prices rough list; "integration" = one team consuming one other team's API.

ScaleAPI SurfaceGovernance CostCost of No StandardOn-Call Impact
Startup (20 eng, 10 services)~10 internal APIs, 1 public~0.25 FTE writing a style guide ≈ $6K/moLow — everyone knows everyoneShared on-call; errors debugged by the author
Growth (200 eng, 80 services)~80 internal, public API with 1K partners2 FTE platform (linters, SDK gen, gateway) ≈ $50K/mo + gateway ~$5–15K/mo~1–2 eng-weeks per new integration × ~100/yr ≈ $500K–1M/yrIdempotency and error-code bugs cause ~1–2 Sev2s/quarter
Large (2,000 eng, 800 services)~800 internal, public API with 50K+ developers6–10 FTE API platform + DevRel ≈ $150–250K/moIntegration tax scales with N² team pairs; a single breaking change can hit thousands of partnersRetry storms from inconsistent error semantics are a top-5 cause of cascading incidents

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to Reverse
Public resource names and ID formatOne-wayEvery partner integration; effectively permanent
Error code stringsOne-wayClient retry logic breaks silently
Pagination token semantics (opaque vs structured)One-way if structuredClients parse what you exposed; opaque from day 1 is free
Idempotency key retention window (shortening)One-wayRetries beyond the new window double-execute
Internal gRPC field additionsTwo-wayFree
REST vs gRPC for a new internal serviceTwo-way (mostly)A few weeks to add a second transport
Adopting GraphQL for client aggregationTwo-way early, one-way after ~6 monthsEvery client screen rewritten against it
Auth model for public API (API keys vs OAuth)One-wayEvery partner re-onboards

The Standard I'd Write#

RFC-API-001: API Contract Standard

Scope: All synchronous APIs exposed across a team boundary (internal gRPC and external REST). Event contracts are covered by RFC-EVT-001.

Mandatory (MUST):

  1. Every non-GET mutating operation MUST accept an idempotency key (Idempotency-Key header or request_id field), scoped per caller, retained ≥ 24h.
  2. Every list operation MUST use opaque cursor pagination with a server-enforced max page size.
  3. Every error MUST use the shared error model: stable code, retryable boolean, request_id.
  4. Operations with p99 > 2s MUST be exposed as long-running operations.
  5. Breaking changes MUST fail CI (buf breaking / OpenAPI diff); overrides require an approved exception.

Recommended (SHOULD): propagate deadlines; support field masks on update; emit Deprecation/Sunset headers ≥ 12 months before removal of public fields.

Exceptions: filed as a 1-page doc to the API platform channel; decided within 5 business days; time-boxed to 2 quarters.

Success metrics: % of APIs passing lint (target 95% in 2 quarters); median new-integration time (target < 3 days); Sev2+ incidents attributed to retry/duplicate semantics (target −50% YoY).

What I'd Tell the VP#

"Every team currently invents its own API conventions, so each new integration between teams takes one to two engineer-weeks of translation, and we've had duplicate-charge incidents because retries aren't safe by default. I'm proposing a small API platform team — about three engineers — to encode our standards into automated checks and generated client libraries. That makes the right thing the default without adding a review committee. We expect integration time to drop from weeks to days and to eliminate an entire class of payment incidents. The main risk is teams seeing it as bureaucracy, so it's enforced by tooling, with a five-day exception process."

Principal Interview Signals#

SignalWhat It Sounds Like
Treats APIs as an org-wide system"The problem isn't this endpoint — it's that we have 40 pagination dialects. I'd fix that with a linter, not a meeting."
Prices the tradeoff"A v2 costs us about two engineers for 18 months of dual-running. Additive evolution costs nothing. That's why v2 is the last resort."
Names the one-way doors"ID format and error codes are forever. Transport choice internally is not — I'd spend my review time on the former."
Knows when not to standardize"I would not mandate GraphQL. I'd mandate the error model and idempotency, and let teams pick transport within the paved road."
Designs the deprecation machine"We need per-client usage telemetry on every field before we can deprecate anything. That's the prerequisite, not the afterthought."

Staff answers that L7 interviewers find insufficient:

  • "We'll version it as v2" — without pricing the dual-running cost or the partner migration campaign.
  • "Each team should follow REST best practices" — with no enforcement mechanism, which in practice means no standard.
  • "Add idempotency keys to this endpoint" — correct locally, but misses that the fix belongs in the gateway or framework so the next 200 endpoints get it for free.

🧭 Principal Move: "I'd put idempotency, pagination, and the error model into the service framework and the gateway, so a team that writes a new endpoint gets them without thinking. Standards that depend on every engineer remembering them are not standards."

Failure Modes & Operational Reality#

FailureDetection SignalBlast RadiusMitigationOwner
Duplicate side effects from retriesidempotency.replay_count flat while payments.duplicate_detected rises; customer complaintsEvery retried mutation; money, emails, ordersMandatory idempotency keys; same-txn dedupe; downstream key propagationOwning service team; payments platform for money paths
Retry storm amplifying an outagehttp.requests_per_sec rising while success_rate falls; retries/original ratio > 1Entire dependency chain; 3 layers × 3 retries = 27×Retry budgets (≤ 10% extra), retry only at the edge, Retry-After on 429/503Platform (client libraries)
Deep offset pagination saturating DBdb.rows_examined / rows_returned > 100; slow query log with high OFFSETOne tenant degrades a shared clusterKeyset cursors; cap offsets; per-tenant query budgetOwning service team
Silent breaking changeClient error rate by sdk_version jumps after deploy; no server errorsAll clients of the changed fieldBreaking-change linter; contract tests; staged rollout by clientAPI platform (tooling) + service owner
GraphQL query-cost explosiongraphql.resolver_calls_per_request p99 > 10K; gateway CPUGateway and every backend it fans out toPersisted queries, complexity limits, per-client cost budgetsGateway team
LRO operations orphanedoperations.pending_age_seconds p99 climbing; workers idleUsers waiting forever on "processing"Lease-based workers, heartbeat timeouts, operation GCOwning service team

Failure Scenario: The Retry Storm That Wasn't a Traffic Spike#

t=0       Inventory service p99 rises from 40ms to 900ms (a slow query after an index was dropped).
t=+15s    Checkout's HTTP client times out at 500ms, retries 3×. Gateway also retries 5xx 2×.
t=+30s    Inventory sees 6× normal RPS. Connection pool exhausted. p99 → 5s.
t=+1min   Mobile apps retry on their own after 10s. Effective amplification ≈ 18×.
t=+2min   Inventory is 100% saturated; checkout success rate 12%. Dashboards show "traffic spike."
t=+8min   On-call disables gateway retries via flag; restores index. Recovery at t=+14min.

Detection: retry_ratio = retried_requests / original_requests per edge — should be < 0.1; alert at > 0.5. Prevention: retry budgets in the shared client library (retry only if the retry rate over the last 10s is < 10% of requests), retries at exactly one layer, 503 + Retry-After from the overloaded service. Owner: the platform team that owns the client library — this is not fixable service-by-service.

In the Wild#

Stripe: Idempotency Keys and Date-Pinned Versions#

Stripe's public API accepts an Idempotency-Key header on POST requests, stores the result for 24 hours, and returns the saved response on replay — including errors — and rejects reuse of a key with different parameters. Separately, Stripe pins each account to the API version current at its first request and maintains a chain of version transformations so old integrations keep working while the API evolves.

Staff insight: Stripe made retry safety and backward compatibility properties of the platform, not of individual endpoints. In an interview, cite this when you argue the idempotency layer belongs in shared middleware.

Google: API Improvement Proposals (AIPs)#

Google publishes its API design guidance as AIPs — resource-oriented design, standard methods, page_token/next_page_token pagination, long-running operations as a standard Operation resource, and field masks for partial updates. These conventions are enforced across a very large number of Google Cloud APIs with lint tooling.

Staff insight: This is the Principal pattern in public: one written standard, machine-enforced, adopted across thousands of APIs. Borrowing AIP conventions in an interview ("LRO as an operation resource, AIP-151 style") signals you know the solved problems.

Facebook/Meta: GraphQL for Mobile#

Facebook created GraphQL (open-sourced 2015) because its mobile News Feed needed data from many backends and REST endpoints were either over-fetching over cellular networks or multiplying into bespoke endpoints per screen. Persisted queries — shipping query IDs rather than query text — became a standard production practice.

Staff insight: GraphQL solved a specific problem: many client surfaces, many backends, bandwidth-constrained networks, and a team large enough to own the gateway. Without those conditions, a BFF is cheaper.


Staff Calibration#

What Staff Engineers Say (That Seniors Don't)#

ConceptSenior (L5)Staff (L6)Principal (L7)
Protocol choice"gRPC is faster, let's use it""REST for partners, gRPC internally with deadlines; the consumer decides the protocol""The paved road is gRPC internally; I'd spend governance effort on error models and idempotency, not transport debates"
Retries"Client retries on failure""Retries only with idempotency keys, only at the edge, with a 10% retry budget""Retry policy lives in the shared client library so no team can build a 27× amplifier by accident"
Pagination"?page=3&limit=50""Opaque keyset cursors over (created_at, id), max page 200, no totals""Pagination is in the API standard and linted; opaque cursors keep our storage migration a two-way door"
Versioning"We'll ship v2""Additive-only with a breaking-change linter; v2 only for a model change""A v2 is ~2 FTE × 18 months of dual-running plus a partner campaign; I'd need a business case"
Errors"Return 500 with a message""Stable codes with a retryable flag; 504 documented as outcome unknown""One error model org-wide so incident response doesn't start with 'what does this code mean'"
Ownership"My team owns the endpoint""We own the contract, its deprecation timeline, and per-client usage telemetry""API ownership is registered in a catalog with SLOs; unowned APIs get a deprecation owner assigned"
Why "Retries" separates levels

The L5 answer is not wrong — clients should retry transient failures. The gap is the system view: L6 sees that retries without idempotency cause duplicates and retries at every layer cause amplification. L7 sees that no amount of per-team discipline fixes this across 300 services; only a shared library with retry budgets does, and that library needs an owner.

Why "Versioning" separates levels

L5 treats versioning as the change mechanism. L6 treats it as the escape hatch and invests in additive evolution. L7 prices it: dual-running infrastructure, documentation, SDKs, support load, and the partner migration campaign — and asks whether the business outcome justifies that spend.

Common Interview Traps#

  • Listing endpoints without guarantees. POST /orders is incomplete until you say what happens when it's retried.
  • Offset pagination on an unbounded collection. It works in the demo and fails for your largest customer first.
  • Synchronous endpoints for slow work. "The endpoint generates the report and returns it" — for a 90-second report behind a 60s LB timeout.
  • Leaking auto-increment IDs. Enumerable, reveals business volume, and couples clients to a single-database ID generator you'll replace when you shard.
  • GraphQL as a default. Choosing it without mentioning query cost limits, N+1 batching, or who owns the gateway.
  • Treating 504 as failure. A timed-out write may have committed. Clients must retry with the same idempotency key, not a new one.
  • "We'll just version it." Without naming the dual-running cost and the deprecation telemetry you'd need.

Practice Drill#

Prompt: "Design the API for a bulk import feature: enterprise customers upload CSVs of up to 5 million contacts. The current implementation is a synchronous POST /contacts/import that times out for anything over ~50K rows."

Staff Answer

I'd split upload from processing and make processing a long-running operation. Step 1: POST /v1/imports with an Idempotency-Key returns a pre-signed upload URL for object storage — the file never passes through our API servers, so a 2 GB CSV doesn't pin a web worker. Step 2: POST /v1/imports/{id}:start returns 202 with operations/{op_id}; workers process in chunks of ~10K rows, each chunk upserted by a natural key (email within the tenant) so a crashed worker can resume without duplicates. GET /v1/operations/{op_id} returns progress_pct, rows_processed, rows_failed, and a link to a paginated error report (GET /v1/imports/{id}/errors?page_token=) rather than one giant error payload. Completion fires a signed webhook, with polling as the fallback. Per-tenant concurrency is capped at 2 active imports and a row-rate budget (~5K rows/s) so one customer's 5M-row import doesn't starve the shared database. Metrics: import.rows_per_sec, import.operation_age_p99, import.chunk_retry_count. The contacts team owns the operation lifecycle; the platform owns the LRO framework and the upload service.

Why this is L6:

  • Separates transport (upload) from work (processing) and uses the LRO pattern instead of stretching timeouts.
  • Makes chunk processing idempotent via natural-key upserts, so retries and crashes are safe.
  • Names per-tenant fairness limits, metrics, and ownership.

What L7 adds:

  • Notes that "bulk import" is the fourth team building upload + LRO + error reports, and proposes a shared bulk-operations framework with a standard operation resource.
  • Prices it: shared framework ≈ 2 engineers × 1 quarter, versus each team spending ~1 quarter building its own.
  • Makes the error-report format and operation resource part of the org's API standard so customers learn one pattern across every bulk feature.

Where This Appears#

Related Technologies: API Gateways · Apache Kafka · PostgreSQL

  1. Loading the index…