Hiring BarSupport

gRPC

Technology guide39 min read5 diagrams

Why This Matters#

gRPC is not a faster REST. It is a contract-first RPC system whose real features are a typed schema that evolves safely, deadlines that propagate across hops, and cancellation that stops wasted work. Binary encoding and HTTP/2 make it efficient, but efficiency is rarely why it wins or loses an interview. What interviewers probe is the sharp edge that comes with HTTP/2: a gRPC client holds a few long-lived connections and multiplexes every request over them, so connection-level (L4) load balancing quietly sends all of a client's traffic to one backend, and a scale-out adds pods that receive nothing.

That is why "services talk over gRPC" is a sentence interviewers push on. The L5 candidate says "gRPC is faster than JSON." The L6 candidate says "protobuf contracts in a shared repo with breaking-change checks in CI; every call carries a deadline derived from the caller's remaining budget; retries only on UNAVAILABLE, at one layer, capped by a retry-throttling token bucket; per-request load balancing through a mesh or client-side round robin over a headless service, because L4 balancing pins HTTP/2 connections; and servers set a max connection age so clients rebalance after a deploy." The L7 candidate asks who owns the schema repository for 400 services, how breaking changes are governed, and whether the company standardises on proxyless gRPC or a sidecar mesh, because that choice sets the platform team's size for years.

The L5 → L6 gap is not knowing that protobuf is binary. It is knowing that gRPC moves reliability policy (deadlines, retries, balancing) into the client, so a missing deadline or a retry at every layer is a fleet-wide behaviour, not a local bug.

The L5 → L6 → L7 Contrast#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Use gRPC for internal services, it's faster""Who are the callers? Internal services get gRPC with protobuf contracts; browsers and partners get REST or gRPC-Web at the edge.""One RPC standard for the company, one schema registry, one governance process for breaking changes, and an edge story for external callers."
Deadlines"Set a timeout on the client"Every call has a deadline; propagated with elapsed time deducted; servers check cancellation and stop workMandates deadlines in the framework (calls without one fail lint or default to a platform value), as Dropbox did with Courier
Load balancing"Put it behind the load balancer"L4 pins HTTP/2 connections; per-request balancing via L7 proxy/mesh or client-side policy; max connection age to rebalanceChooses proxyless vs sidecar mesh fleet-wide, priced in CPU, latency and platform headcount
Retries"Retry failed calls 3 times"Retry only idempotent methods on retryable codes, at one layer, with throttling; hedging only for idempotent, latency-critical readsOwns the retry budget as a fleet policy: retries per layer, throttling defaults, and who may enable hedging
Schema"Define a proto file"Field numbers are forever; add fields, never change types or reuse numbers; reserved for deletions; breaking-change detection in CIRuns the schema registry and deprecation process; decides who may approve a breaking change
Ownership"Each team writes its own client"Generated clients; interceptors for auth, tracing, metrics, deadlines owned by platformDefines the RPC platform contract: what every service gets for free and what it must declare
Why "Load balancing" separates levels

"Put it behind the load balancer" works for HTTP/1.1, where each request tends to take a connection from a pool and connections churn. A gRPC client opens one HTTP/2 connection per backend address it knows and multiplexes thousands of concurrent RPCs over it. Behind an L4 balancer (a Kubernetes ClusterIP, a TCP load balancer), the balancer sees one long-lived connection and pins it to one pod. Scale from 10 to 20 pods during a spike and the new 10 receive almost nothing, because existing clients never reconnect. The gRPC project's own load-balancing guidance says exactly this: L3/L4 balancers treat each connection as a unit and can't distribute individual requests (gRPC blog). The Staff answer picks per-request balancing (an L7 proxy or mesh, or client-side round robin over a headless service) and sets a server-side max connection age so even pinned clients rebalance. The Principal answer decides that fleet-wide, because it determines whether every pod carries a sidecar.

The 60-Second Pitch#

"Internal service-to-service calls use gRPC with protobuf contracts in a shared schema repo, generated clients, and a CI check that rejects breaking changes. Every call carries a deadline: the edge sets the user-facing budget, say 800ms, and each hop propagates what's left. Retries happen in one place, the gRPC client's service config, only for idempotent methods on UNAVAILABLE, at most 3 attempts with exponential backoff and jitter, capped by retry throttling so a struggling backend sees fewer retries, not more. Balancing is per request: a sidecar mesh or client-side round robin over a headless service, never plain L4, and servers set a max connection age of a few minutes so clients rebalance after scale-outs. For streaming, I'd use server streaming for push with a resume token so reconnects don't lose events. Browsers reach us through REST or gRPC-Web at the edge, because browsers can't speak native gRPC."

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Internal service-to-service RPCMany services, many languages, low latency, frequent deploysUnary gRPC, protobuf contracts, deadlines, per-request LB, interceptors for auth/tracingL4 pinning → hot pods; missing deadlines → cascading failure; retry amplificationp99 within budget; no cascade from one slow dependency
Streaming and pushServer-to-client events, long-lived connections, mobile networksServer or bidi streaming, keepalives, resume tokens, max connection age, graceful GOAWAY on deployStreams pinned to old pods; lost events on reconnect; keepalive abuse → too_many_pingsEvery event delivered or detectably missed; reconnect within seconds
External or browser-facing APIBrowsers, partners, caching, human-readable debuggingREST/JSON or gRPC-Web via a proxy; JSON transcoding at the edge; gRPC behindBrowser can't do client or bidi streaming; partners can't read protobuf; HTTP caches don't applyContract stable for external clients for years

🎯 Staff Move: "I'll use gRPC for the internal mesh, the first intent, and keep the public API as REST at the gateway. External clients need caching, curl-ability and a contract they can read; internal services need deadlines, typed contracts and efficient multiplexing. The gateway translates, so neither side compromises."

The Staff Positions#

PositionRationale
Every RPC has a deadline, propagated across hopsgRPC's default is no deadline; one slow dependency otherwise exhausts every caller upstream.
Per-request load balancing, never plain L4HTTP/2 connections are long-lived and multiplexed; L4 pins all of a client's traffic to one backend.
Retries at exactly one layer, with throttling3 layers × 3 attempts is 27 calls per user request at the bottom during an outage.
Field numbers are foreverProtobuf identifies fields by number on the wire; reuse or type changes silently corrupt data across versions.
Breaking-change detection in CIHumans miss renumbered fields; a schema linter doesn't.
Servers cap connection ageForces clients to reconnect and rebalance after scale-outs and deploys.
Browsers and partners get REST or gRPC-Web at the edgeBrowsers can't speak native gRPC; partners need a stable, readable contract.

Architecture & Internals#

Only five internals change design decisions: HTTP/2 multiplexing, protobuf's wire format, the four call types, deadlines and status codes, and channels and load-balancing policies. The comparison with REST and GraphQL is in REST vs gRPC vs GraphQL.

HTTP/2: One Connection, Many Streams#

Each RPC is an HTTP/2 stream: a HEADERS frame (method path like /orders.v1.OrderService/GetOrder, deadline as grpc-timeout, metadata), length-prefixed protobuf messages in DATA frames, and trailing headers carrying grpc-status. Many streams share one TCP connection; each has flow-control windows so a slow reader on one stream doesn't stall the others at the HTTP/2 layer.

Diagram: HTTP/2: One Connection, Many Streams

Why it matters in design: connection setup (TCP + TLS) happens once, then every RPC is a few frames, which is where gRPC's latency and CPU advantage over JSON-over-HTTP/1.1 comes from. The price is that the connection is the unit an L4 balancer sees, and it lives for hours. HTTP/2 still runs on TCP, so a lost packet stalls every stream on that connection (TCP head-of-line blocking); on lossy mobile networks that is one reason teams look at HTTP/3, as Uber did for its push platform (below).

Protobuf: The Wire Format Is Field Numbers#

A protobuf message on the wire is a sequence of (field number, wire type, value) entries. Names never appear on the wire. That single fact sets every schema-evolution rule:

syntax = "proto3";
package orders.v1;

message Order {
  string id = 1;
  int64  amount_minor = 2;           // money as integer minor units
  string currency = 3;
  OrderStatus status = 4;
  google.protobuf.Timestamp created_at = 5;
  reserved 6;                        // was 'coupon_code', removed in 2025
  reserved "coupon_code";
  optional string note = 7;          // proto3 optional: presence is tracked
}

enum OrderStatus {
  ORDER_STATUS_UNSPECIFIED = 0;      // zero value must mean 'unknown'
  ORDER_STATUS_PENDING = 1;
  ORDER_STATUS_PAID = 2;
}
ChangeSafe?Why
Add a new field with a new numberYesOld readers skip unknown fields; new readers see the default when absent
Rename a field (same number, same type)Wire-safe; breaks JSON and generated codeJSON mapping uses names; callers' code stops compiling
Delete a fieldOnly with reserved number and nameOtherwise someone reuses the number later
Reuse a field numberNoOld data decodes into the new field with the wrong meaning
Change a field's type (e.g. int32 → string)NoWire types differ; decoding fails or silently misreads
Add an enum valueYes, if readers handle unknown valuesOld readers see an unrecognised number
Change a method's request or response typeNoIt's a different contract; add a new method

🎯 Staff Insight: "Protobuf compatibility is a promise about field numbers, not names. So deletions get reserved, type changes become new fields, and CI runs a breaking-change check against the last released schema, because the failure mode is silent data corruption between two services deployed a day apart."

The Four Call Types#

TypeShapeUse forWatch out for
Unary1 request → 1 response90%+ of service callsNothing special: deadlines, retries, LB all behave as expected
Server streaming1 request → N responsesPush feeds, large result sets, watch APIsStream pinned to one backend for its lifetime; resume tokens for reconnect
Client streamingN requests → 1 responseUploads, batched telemetryRetries are hard once data has been sent; partial-progress semantics needed
Bidirectional streamingN ↔ M interleavedChat, real-time sync, push with acksNot supported by gRPC-Web in browsers; ordering per direction only; long-lived load

Streams are not a message queue. A stream lives on one connection to one backend; if that backend restarts, the stream ends and the client must reconnect and resume. Anything that must survive that needs an offset or resume token, and usually a durable log behind the server. The push design in Live Updates and the transport comparison in WebSockets vs SSE vs Long Polling cover the alternatives.

Deadlines, Cancellation and Status Codes#

Diagram: Deadlines, Cancellation and Status Codes

When a server acts as a client, gRPC can propagate the incoming deadline, converting it to a timeout with elapsed time already deducted so clock skew between hosts doesn't matter; some languages do this by default, others need it enabled (deadlines guide). When the deadline passes, the client fails with DEADLINE_EXCEEDED and the server's call is cancelled; the application still has to check for cancellation and stop work.

Status codeMeaningRetry?
OKSuccess—
UNAVAILABLETransient: connection failure, server shutting down, overloadedYes, with backoff (the canonical retryable code)
DEADLINE_EXCEEDEDRan out of timeRarely: the budget is spent; retrying usually just adds load
RESOURCE_EXHAUSTEDQuota or rate limit hitOnly after the server's pushback delay
INVALID_ARGUMENT, NOT_FOUND, PERMISSION_DENIEDCaller's problemNever
INTERNAL, UNKNOWNServer bug or unexpected failureNo by default; a bug retried is a bug amplified

Channels, Resolvers and Load-Balancing Policies#

A channel is a client's handle to a logical service. It uses a resolver (DNS, a registry, or xDS from a control plane) to get addresses, opens a subchannel per address, and applies an LB policy per RPC. With pick_first (the default in most implementations) the client sends everything to one address; with round_robin it spreads RPCs across all ready subchannels. Proxyless setups get the address list and policy (weighted round robin, outlier ejection) from an xDS control plane, the same API Envoy uses; see Envoy, Kong & NGINX.

Keepalive defaults matter for long-lived connections: client keepalive pings are disabled by default, the server's keepalive time defaults to 2 hours, the ping timeout is 20 seconds, and servers by default reject pings more frequent than every 5 minutes, answering abusive clients with GOAWAY (keepalive guide). Client authors must coordinate keepalive settings with service owners.


Core Usage — "The Entire Game": The Service Contract#

In Kafka the game is the partition key. In gRPC it is the service contract: five things every service declares, most of which live in the client. Get them right and the RPC layer degrades gracefully; get them wrong and every caller inherits the mistake.

The service contract (per method):
  1. schema     -> request/response messages, field numbers, versioned package
  2. deadline   -> the budget this method gets, and what callers propagate
  3. retry      -> is it idempotent? which codes? how many attempts? hedging?
  4. balancing  -> how each RPC picks a backend, and how connections rebalance
  5. signals    -> status codes it returns, metrics and traces it emits

Step 1: Design the Schema for Evolution#

service OrderService {
  rpc GetOrder(GetOrderRequest) returns (Order);
  rpc ListOrders(ListOrdersRequest) returns (ListOrdersResponse);
  rpc PlaceOrder(PlaceOrderRequest) returns (PlaceOrderResponse);
}

message PlaceOrderRequest {
  string request_id = 1;        // idempotency key: makes PlaceOrder safe to retry
  string customer_id = 2;
  repeated LineItem items = 3;
}

message ListOrdersRequest {
  string customer_id = 1;
  int32  page_size = 2;         // server caps at e.g. 100
  string page_token = 3;        // opaque cursor, never an offset
  google.protobuf.FieldMask read_mask = 4;  // caller asks only for fields it needs
}

Rules that hold up over years: a request and response message per method (never a bare primitive, so fields can be added); a versioned package (orders.v1), with v2 as a new service rather than in-place breakage; idempotency keys on mutating methods so they can be retried; cursor pagination with server-capped page sizes; and field masks for wide resources. The contract discipline is the same as in API Contracts; the schema-change process is in Schema Design.

Step 2: Budget the Deadline#

User-facing budget: 800 ms at the edge (p99 target 600 ms)

  edge -> checkout                  deadline 800 ms
  checkout local work               ~40 ms
  checkout -> pricing   (parallel)  remaining ~760 ms, pricing p99 120 ms
  checkout -> inventory (parallel)  remaining ~760 ms, inventory p99 150 ms
  checkout -> payments  (serial)    remaining ~600 ms, payments p99 300 ms

Rules:
  - each hop propagates (remaining - local reserve), never a fresh fixed timeout
  - a callee whose p99 exceeds its share is a design problem, not a timeout problem
  - servers check ctx/cancellation before expensive steps and stop when cancelled

The Latency Budget calculator does this split; the Latency Numbers table grounds each hop's share.

Step 3: Configure Retries and Hedging in One Place#

gRPC clients read a service config (JSON, from DNS, xDS or local configuration) that defines per-method retry or hedging policies. Neither is enabled unless configured, and maxAttempts values above 5 are treated as 5 (retry design A6).

{
  "methodConfig": [
    {
      "name": [{ "service": "orders.v1.OrderService", "method": "GetOrder" }],
      "timeout": "0.3s",
      "retryPolicy": {
        "maxAttempts": 3,
        "initialBackoff": "0.05s",
        "maxBackoff": "0.5s",
        "backoffMultiplier": 2,
        "retryableStatusCodes": ["UNAVAILABLE"]
      }
    },
    {
      "name": [{ "service": "catalog.v1.CatalogService", "method": "GetProduct" }],
      "hedgingPolicy": {
        "maxAttempts": 2,
        "hedgingDelay": "0.02s",
        "nonFatalStatusCodes": ["UNAVAILABLE"]
      }
    }
  ],
  "retryThrottling": { "maxTokens": 10, "tokenRatio": 0.1 }
}

Three behaviours to know (retry guide, hedging guide):

  1. Commit point. Once response headers arrive, the RPC is committed and won't be retried. Retries cover failures before the server started answering.
  2. Throttling. Each failure costs a token and each success earns back tokenRatio; when the count falls below half of maxTokens, retries and hedges stop. During an outage the client automatically backs off to roughly one attempt per request.
  3. Hedging sends a second copy after hedgingDelay without waiting for a failure and takes the first success, cancelling the rest. It cuts tail latency and multiplies load; use only on idempotent reads, with a delay near the method's p95. Servers can push back with grpc-retry-pushback-ms.
Retries at several layers multiply: three layers of three attempts can turn one user request into 27 calls on the failing service, which is why retries belong at one layer with a budget.

Step 4: Pick Per-Request Balancing#

OptionHowPer-request?CostPick when
L4 balancer / ClusterIPKernel or TCP LB picks a backend per connectionNoLowestNever for gRPC alone; OK only with max connection age and many clients
L7 proxy (gateway, Envoy)Proxy terminates HTTP/2 and balances each RPCYesExtra hop (~sub-ms to ~1 ms), proxy CPUEdge, trust boundaries, mixed clients
Sidecar meshA proxy beside every podYesSidecar CPU/memory per pod, extra hop each sidePolyglot fleets wanting uniform mTLS, retries, telemetry
Client-side (round_robin over headless service)Client resolves all pod IPs, balances itselfYesNo hop; client complexity per languageHigh-volume internal traffic, few languages
Proxyless xDSClient gets endpoints and policy from a control planeYesControl plane to run; library support per languageLarge fleets wanting mesh features without sidecars

Whatever the choice, set max connection age on servers (minutes, with a grace period) so long-lived connections get recycled and rebalanced, and drain with GOAWAY on shutdown so in-flight RPCs finish. The design space is in Load Balancing and Service Registry.

Step 5: Signals — Status Codes, Metrics, Traces#

Return precise status codes (INVALID_ARGUMENT vs UNAVAILABLE decides whether callers retry). Emit per-method RED metrics (rate, errors by code, duration histogram) from server and client interceptors, and propagate trace context in metadata. Without client-side metrics you cannot see L4 pinning, because each server looks healthy; the skew only shows when you compare per-backend request rates. The trace plumbing is in Distributed Tracing.

🎯 Staff Move: "Every method gets a deadline share from the edge budget, an idempotency decision that determines whether retries are allowed, and a retry policy in service config with throttling. Retries live only in the gRPC client, not also in the mesh and the application, because three layers of three attempts is 27× load during exactly the outage we're trying to survive."


The Tunable Tradeoff — Where Reliability Logic Lives#

Every gRPC platform decision moves along one axis: how much reliability policy lives in the client library vs in proxies?

SettingClient-heavy endProxy-heavy endWho pays at the client-heavy end
Load balancingClient-side / proxyless xDSSidecar or central L7 proxyLibrary maintainers in every language
Retries and hedgingService config in the clientProxy retry policyService teams if configs drift
mTLS and identityLibrary-managed certificatesSidecar-terminated TLSPlatform, for rotation in N runtimes
TelemetryInterceptorsProxy access logs and metricsTeams on languages without good interceptors
Latency and CPUNo extra hop+1–2 hops per call, sidecar CPU per podAt the proxy end: every request, and the compute bill

🎯 Staff Move: "For a polyglot fleet of 300 services I'd take the sidecar mesh: uniform mTLS, retries and telemetry without four client libraries to keep in sync. For one or two very high-QPS paths in a single language, I'd go proxyless with xDS to remove the hops. Same control plane, two data planes."


Anti-Patterns — What Kills gRPC Deployments#

1. No Deadlines#

The default is to wait forever. One slow downstream holds every caller's threads, and the outage climbs the call graph. Fix: framework-enforced deadlines; propagate remaining budget; servers honour cancellation.

2. L4 Load Balancing for Long-Lived Connections#

ClusterIP or a TCP balancer in front of gRPC pins each client connection to one pod; after a scale-out the new pods idle while old ones melt. Fix: per-request balancing (L7 proxy, mesh or client-side) plus server max connection age.

3. Retries at Every Layer#

Application retries 3×, the generated client retries 3×, the mesh retries 3×: 27 attempts at the bottom per user request during a brownout. Fix: one layer retries, with throttling; others pass errors through.

4. Retrying Non-Idempotent Methods#

PlaceOrder retried on DEADLINE_EXCEEDED creates two orders if the first actually succeeded. Fix: retries only on methods with an idempotency key; see Idempotency.

5. Reusing Field Numbers or Changing Types#

A deleted field's number is reused for a new field; services deployed a day apart now read each other's data with the wrong meaning, with no error. Fix: reserved on every deletion; breaking-change checks in CI against the last release.

6. Large Messages#

A 50 MB response hits the common 4 MB default receive limit, or worse, gets raised past it and spikes server memory while blocking other streams on the connection. Fix: paginate, stream in chunks, or return a pointer to object storage.

7. Streams Treated as Durable#

A bidirectional stream carries events with no resume token; every deploy, keepalive timeout or mobile network change silently drops whatever was in flight. Fix: offsets or resume tokens, a durable source behind the stream, client reconnect with backoff and jitter.


The Technology Landscape — Head-to-Head Comparison#

DimensiongRPCREST + JSONGraphQLApache ThriftAsync messaging (queues/logs)
ContractProtobuf IDL, generated codeOpenAPI (optional), hand-written or generatedSchema with typed queriesThrift IDL, generated codeMessage schemas (Avro, protobuf)
TransportHTTP/2HTTP/1.1 or 2HTTP (usually POST)Pluggable (TCP, HTTP)Broker
StreamingUnary, server, client, bidiSSE or WebSocket on the sideSubscriptions (extra infra)LimitedNative, async
Browser supportVia gRPC-Web proxy (unary and server streaming)NativeNativePoorNot directly
CachingNo HTTP cachingHTTP caches and CDNsHard (POST, per-query)NoN/A
Deadlines/cancellationBuilt in, propagatedPer-client timeoutsPer-client timeoutsPer-clientMessage TTLs
Pick whenInternal service mesh, polyglot, low latency, streamingPublic APIs, partners, cacheable readsClient-driven aggregation for many UIsExisting Thrift estatesWork that can wait; decoupled producers

🎯 Staff Insight: "gRPC versus REST is a question about who the caller is. Internal callers I control get gRPC for typed contracts, deadlines and efficiency. Callers I don't control get REST, because they need caching, curl and a contract they can read without a code generator."


Patterns#

Pattern 1: gRPC Inside, REST at the Edge#

Diagram: Pattern 1: gRPC Inside, REST at the Edge

The gateway owns external contracts, caching and auth; internal services speak gRPC with propagated deadlines. JSON transcoding (mapping HTTP routes to RPCs via annotations) lets one proto serve both. The edge design is in Edge Gateway.

Pattern 2: Sidecar Mesh, or Proxyless xDS#

Envoy-style sidecars terminate mTLS, balance each RPC, apply retry and outlier-ejection policy, and emit uniform telemetry; the price is a proxy per pod (illustratively 0.1–0.25 vCPU and 50–150 MB each, so hundreds of vCPU across 2,000 pods) and extra hops. Proxyless gRPC consumes the same xDS config inside the library: no sidecar, no hop, but only for languages with mature support.

Pattern 3: Server-Streaming Push With Resume#

The client opens Subscribe(SubscribeRequest{resume_token}); the server streams events from a durable log, each carrying a token. On reconnect (deploy, network change, max connection age) the client resumes from its last token. Keepalives detect dead connections; the server drains with GOAWAY and clients reconnect with jitter. See Push vs Poll.


Scaling#

The Numbers#

ResourceDefault or documented figureDesign note
DeadlineNone by defaultAlways set one; propagate remaining budget
Retry / hedging maxAttemptsValues above 5 treated as 52–3 is typical for retries, 2 for hedging
Retry throttlingOff unless configured; retries pause below half of maxTokensExample config: maxTokens 10, tokenRatio 0.1
Transparent retriesDon't count toward maxAttemptsCover RPCs that never reached server application logic
Max receive message size4 MB default in the main implementationsRaise deliberately or stream/paginate
Client keepaliveDisabled by default; timeout 20sDon't set below ~1 minute without the service owner
Server keepalive / permit time2 hours / 5 minutesFaster client pings get GOAWAY too_many_pings
Max connection ageUnlimited by defaultSet to minutes for rebalancing, with a grace period
gRPC-WebUnary + server streaming (text mode); via a proxyNo client or bidi streaming from browsers

Scaling Moves in Order#

  1. Deadlines and per-request balancing everywhere; max connection age on servers.
  2. One retry layer with throttling, configured centrally.
  3. Schema registry and breaking-change gates once more than a few teams share protos.
  4. Connection subsetting when clients × backends makes full meshes expensive.
  5. Mesh or proxyless xDS for uniform policy; proxyless for the hottest paths.
  6. HTTP/3 for lossy client networks (mobile push) where TCP head-of-line blocking hurts.

🎯 Staff Move: "Past a few hundred pods per service I'd subset: each client connects to, say, 20 backends chosen deterministically, rather than all 300. Full-mesh connections cost memory and health-check traffic and buy no balance a subset doesn't."


Failure Modes & Recovery#

1. Hot Pods After Scale-Out (Connection Pinning)#

  • Symptom: Autoscaler doubles a service from 10 to 20 pods; p99 doesn't improve; 10 pods at 90% CPU, 10 nearly idle.
  • Root cause: Clients hold long-lived HTTP/2 connections balanced at L4; new pods get no existing connections.
  • Detection: Per-pod request rate skew (max/median > 2); per-pod CPU spread; client-side per-backend metrics.
  • Fix: Roll clients to force reconnects; enable round robin or route through the mesh.
  • Prevention: Per-request balancing; server max connection age of a few minutes; alert on per-pod RPS skew.

2. Retry Storm#

  • Symptom: A dependency brownout turns into a hard outage; its request rate triples while its success rate falls.
  • Root cause: Retries at several layers, no throttling, retrying DEADLINE_EXCEEDED.
  • Detection: Ratio of attempts to logical calls (client metrics); request rate on the dependency vs upstream traffic.
  • Fix: Disable retries at all but one layer; enable throttling; shed load at the dependency.
  • Prevention: Central retry policy with throttling; lint for retry config in more than one layer.

3. Cascading Exhaustion From Missing Deadlines#

  • Symptom: One downstream slows from 50ms to 5s; within minutes every upstream service's thread pool or goroutine count is maxed.
  • Root cause: Calls without deadlines, or fixed long timeouts not derived from the caller's budget.
  • Detection: In-flight RPC count per service; latency of calls to the slow dependency; caller saturation.
  • Fix: Short emergency deadlines via config; circuit-break the dependency; shed.
  • Prevention: Framework-enforced deadlines; propagation; cancellation checks in handlers. See Cascading Failures.

4. Silent Data Corruption From a Breaking Schema Change#

  • Symptom: After a deploy, some orders show wrong amounts or missing fields; no errors logged.
  • Root cause: A field number reused or a type changed; old and new binaries decode each other's messages differently.
  • Detection: Data-quality checks; mismatch between producer and consumer schema versions; contract tests.
  • Fix: Roll back the producer; repair affected records; reintroduce the change as a new field.
  • Prevention: Breaking-change detection in CI; reserved for deletions; schema owner approval.

5. Streams Stuck on Old Pods During a Deploy#

  • Symptom: After a rollout, half the clients' push streams still point at pods being terminated; events drop when those pods exit.
  • Root cause: Long-lived streams with no graceful drain or resume; termination grace shorter than stream life.
  • Detection: Active streams per pod over time; reconnect rate spikes; missed-event reports.
  • Fix: Send GOAWAY, let clients reconnect with jitter and resume from tokens.
  • Prevention: Max connection age; resume tokens; termination grace that covers drain; reconnect jitter so 100K clients don't return in the same second.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Connection pinningPer-pod RPS skewOne service's p99Per-request LB, max connection agePlatform (LB); service team (config)
Retry stormAttempts per logical callThe dependency, then its callersOne retry layer, throttlingPlatform policy; calling teams
Missing deadlinesIn-flight RPCs climbingEntire call graph upstreamEmergency deadlines, circuit breakingCalling service; platform for enforcement
Breaking schema changeData-quality checksEvery consumer of the messageRoll back, repair dataSchema owner; producing team
Stuck streams on deployStreams per terminating podPush clients of that serviceGOAWAY drain, resumeStreaming service team
Keepalive abusetoo_many_pings GOAWAYsClients of that serverAlign keepalive settingsClient team with service owner

When to Use vs. Alternatives#

NeedPickWhy
Internal service-to-service calls, many languagesgRPCTyped contracts, deadlines, efficient multiplexing
Public API for partners and developersREST + JSON (OpenAPI)Readable, cacheable, universal tooling
Many UIs needing different shapes of the same dataGraphQL at the edgeClient-driven queries; see REST vs gRPC vs GraphQL
Browser calling services directlyREST, or gRPC-Web via a proxyBrowsers can't speak native gRPC
Server push to clientsgRPC server streaming, SSE or WebSocketsDepends on client platform and proxy support
Work that can wait, decoupled producersA queue or logSee RabbitMQ & SQS and Kafka

When NOT to Use gRPC#

  • Public, cacheable APIs. HTTP caching and CDNs don't apply; partners want JSON they can read.
  • Browser-first products without a proxy layer. gRPC-Web supports unary and server streaming only, and always needs a proxy.
  • Async work. A call that can wait belongs on a queue; RPC couples the caller's availability to the callee's.
  • Tiny systems with one language and two services. The schema pipeline and LB setup cost more than they save.
  • Environments where L7 balancing isn't possible and clients can't do it either. You'll fight pinning forever.
Diagram: When NOT to Use gRPC

Operational Concerns#

What the On-Call Actually Does#

  1. Reads errors by status code, not just error rate: a jump in UNAVAILABLE is a capacity or connectivity issue; INVALID_ARGUMENT after a deploy is a contract issue; DEADLINE_EXCEEDED points at a slow dependency.
  2. Checks per-pod request skew after every scale-out and deploy; rebalances by recycling connections.
  3. Watches attempts per logical call to catch retry storms early, and flips retry policy centrally if needed.
  4. Traces slow requests across hops to find which callee ate the budget.
  5. Manages streaming drains during deploys: GOAWAY, grace periods, reconnect jitter.
  6. Reviews schema changes flagged by the breaking-change gate and approves or rejects exceptions.

Key Metrics & Alerts#

MetricHealthyAlert
grpc_server_handled_total by grpc_codeErrors < SLOUNAVAILABLE or INTERNAL rate above SLO for 5 min
Server latency histogram p99 per methodWithin the method's budget share> budget share for 10 min
Per-pod request rate max / median< 1.5> 2 (pinning)
Client attempts per logical RPC~1.0> 1.2 sustained (retry pressure)
In-flight RPCs per serverSteadyClimbing with flat throughput (stuck callee)
Active streams per podEven across podsConcentrated on few pods
DEADLINE_EXCEEDED rate per caller → calleeLowSpike (callee slow or budget wrong)

Interview Application — Staff-Level Plays#

Which Case Studies Use gRPC#

Case StudyHow gRPC Is UsedKey Pattern
Load BalancingPer-request vs per-connection balancingL4 pinning; L7 or client-side balancing
Service RegistryResolvers and xDS feed channel address listsClients fail static when the registry is wrong
Edge GatewayREST at the edge, gRPC behind; transcodingThe gateway owns external contracts
Cascading FailuresDeadlines, retries and throttling as cascade controlsOne retry layer; propagated budgets
Distributed TracingTrace context in gRPC metadata via interceptorsUniform instrumentation from the framework
Live UpdatesServer streaming for pushResume tokens; drain on deploy

Every System Design Question Has a gRPC Moment#

  • Ride hailing: "Driver location updates come in over a client stream from the gateway to the location service; offers go out over a bidirectional stream with an ack per message, so we know within a second whether a driver saw the offer."
  • Checkout: "The edge sets an 800ms deadline; checkout fans out to pricing and inventory in parallel with the remaining budget and calls payments serially; only reads retry, and PlaceOrder carries a request ID so a retry can't double-charge."
  • Autoscaling: "New pods only help gRPC traffic if clients rebalance, so servers cap connection age at 5 minutes and the mesh balances per request."

What Interviewers Probe#

After You Say...They Will Ask...What They're Evaluating
"Services use gRPC""You scale from 10 to 20 pods. Does load spread?"L4 pinning; per-request LB; max connection age
"We set timeouts""Service C is slow. What happens to A?"Deadline propagation and cancellation
"Retry on failure""Which codes? Which methods? At which layer?"Idempotency, retryable codes, amplification, throttling
"Protobuf handles versioning""You delete a field. What could go wrong next year?"Field-number reuse; reserved; CI checks
"Stream events to clients""What happens during a deploy?"GOAWAY drains, resume tokens, reconnect jitter
"gRPC for the public API""How do browsers and partners call it?"gRPC-Web limits; REST or transcoding at the edge

L5 vs L6 vs L7 Responses#

ScenarioL5 AnswerL6 / Staff AnswerL7 / Principal Answer
"Design internal communication for 50 services"gRPC with protobufShared proto repo with CI gates, deadlines from edge budget, one retry layer with throttling, per-request LB, interceptors for auth and tracingRPC platform contract: framework defaults, schema governance, mesh vs proxyless decision, and the edge strategy for external callers
"Scale-out didn't help latency"Add more podsDiagnoses connection pinning from per-pod skew; enables round robin or mesh; max connection ageMakes per-pod skew a platform SLO and bans L4-only gRPC in the paved road
"A slow service took down checkout"Increase timeoutsPropagated deadlines, cancellation checks, circuit breaker, retry throttlingFramework refuses calls without deadlines; game days that slow a dependency by 5s
"Should the public API be gRPC?"Yes, for performanceREST for partners, gRPC internally, transcoding at the gatewayOwns the external contract lifecycle: versioning policy, deprecation windows, SDK generation

The Staff gRPC Checklist#

  1. Caller first: "Internal callers get gRPC; external and browser callers get REST or gRPC-Web at the edge."
  2. Deadlines: "Edge budget of 800ms, propagated minus elapsed at every hop; handlers stop on cancellation."
  3. Retries: "One layer, UNAVAILABLE only, max 3 attempts with backoff and jitter, throttled; idempotent methods only."
  4. Balancing: "Per-request via mesh or client round robin; max connection age of a few minutes."
  5. Schema: "Field numbers forever, reserved on deletes, breaking-change gate in CI, new package for v2."
  6. Signals: "Errors by status code, per-pod request skew, attempts per call, traces through metadata."

🎯 Staff Insight: Don't use gRPC for public cacheable APIs, for browsers without a proxy, as a substitute for a queue, or behind plain L4 balancing. The strongest gRPC signal is explaining why doubling the pods didn't help, and what you'd change so it does.

Evaluation Rubric#

DimensionSenior (L5)Staff (L6)Principal (L7)
ContractsProto filesEvolution rules, CI gates, versioned packagesSchema governance and registry for the org
ReliabilityTimeouts and retriesPropagated deadlines, single-layer throttled retries, hedging used sparinglyFramework-enforced defaults; fleet retry budget
Load balancing"Behind the LB"Pinning understood; per-request LB; connection ageMesh vs proxyless decision, priced
Streaming"Use bidi streaming"Resume tokens, drains, keepalive coordinationPush platform standard across products
ChoicegRPC for everythinggRPC inside, REST/GraphQL at the edgeExternal API lifecycle and SDK strategy

Strong hire signals

SignalWhat It Sounds Like
Pinning awareness"HTTP/2 connections are long-lived, so L4 balancing pins them; new pods need clients to reconnect."
Budget thinking"The edge has 800ms; every hop gets what's left, and handlers stop when it's gone."
Retry discipline"One layer retries, with a token bucket, and only methods with an idempotency key."
Schema discipline"Field numbers are forever; deletions are reserved."
Caller-based choice"gRPC for callers I control, REST for callers I don't."

Lean no-hire signals

SignalWhy It Misses the Bar
"gRPC because it's faster" with nothing elseBenchmark choice, no contract or reliability reasoning
No deadlinesWill ship cascading failures
Retries everywhereWill amplify outages
gRPC exposed straight to browsersDoesn't know the platform limits

Common false positives

  • Protobuf encoding trivia ≠ contract governance. Ask how they'd stop a field-number reuse.
  • Mesh enthusiasm ≠ judgment. Ask what the sidecars cost across 2,000 pods and who upgrades them.

Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

At Staff level gRPC is a well-configured client and server. At Principal level it is the company's RPC platform: the framework every service is generated from, the place where deadlines, retries, identity and telemetry become defaults, and the schema estate that 400 services depend on. The L7 question is "What does every service get for free from the RPC framework, who governs the schemas, and do we put reliability logic in libraries or in a mesh?"

🧭 Principal Move: "I want the RPC framework to make the right thing the default: no call without a deadline, retries only where the method is marked idempotent, mTLS identity on every connection, metrics by caller and method out of the box. Then reliability stops depending on every team remembering four rules."

The Org-Level Fault Line#

Smart clients (libraries, proxyless) vs smart network (sidecar mesh).

OptionWhat WorksWhat BreaksWho Pays
Libraries per languageNo extra hops, lowest latency and CPUFeatures drift across languages; every fix needs N releases and every team to upgradePlatform, maintaining N libraries; teams on lagging languages
Sidecar meshUniform policy in any language; upgrades without app redeploysProxy per pod (CPU, memory), extra hops, mesh upgrades become platform-criticalCompute budget; platform team for the mesh
Proxyless xDS for hot paths + mesh elsewhereBest latency where it matters, uniformity elsewhereTwo data planes to reason aboutPlatform, in complexity
No standardEach team choosesInconsistent deadlines, retries and telemetry; incidents debugged from scratchIncident responders

The Principal default: a mesh for the polyglot majority with one control plane, proxyless gRPC on the few highest-volume paths, one schema registry with breaking-change gates, and framework defaults (deadline required, retry throttling on, interceptors for auth/metrics/tracing) shipped in a thin library per supported language.

🧭 Principal Insight: "The question isn't library versus mesh. It's how many languages we commit to supporting well. Every supported language is a library to maintain; a mesh is how you buy your way out of that, priced in sidecar CPU."

Cost Model#

Assumptions: sidecar ~0.15 vCPU and ~100 MB per pod; vCPU ~$25/month; loaded engineer ~$21K/month. Directional only.

ScaleFleetMesh compute/monthPlatform peopleNotes
Startup15 services, 100 pods~$4000.25 FTE (~$5K)Libraries and a gateway suffice; mesh is optional
Growth150 services, 2,000 pods~$7.5K2 FTE (~$42K) for RPC framework, schema registry, meshSchema governance starts paying for itself
Enterprise1,000 services, 20,000 pods~$75K6–10 FTE (~$170K)Proxyless for hot paths saves a share of sidecar cost; subsetting needed

At every scale the people cost dominates the proxy cost, and the hidden cost is incidents caused by inconsistent defaults: one missing deadline in a hot path can cost more in an afternoon than the platform team costs in a quarter.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to Reverse
Field numbers and wire types in shipped protosOne-wayData written with them is decoded that way forever
Public API protocol (REST vs gRPC) for external clientsOne-way-ishYears of client support and SDKs
Mesh vs proxylessTwo-way-ishFleet migration of data planes; policy reimplementation
Package versioning schemeOne-way-ishEvery generated client
Retry and deadline defaultsTwo-wayConfig change, but behaviour changes fleet-wide

The Standard I'd Write#

RFC-RPC-001: Internal RPC Contract

Scope: All synchronous service-to-service calls.

MUST
  1. Use gRPC with protobuf schemas in the central registry; packages versioned (x.v1).
  2. Pass breaking-change checks against the last release; deletions use `reserved`.
  3. Set a deadline on every call, derived from the caller's remaining budget;
     handlers check cancellation before expensive work.
  4. Retry only methods marked idempotent (or carrying an idempotency key), only on
     UNAVAILABLE, in the client service config, with retry throttling enabled.
     No retries in application code or mesh for the same call.
  5. Use per-request load balancing (mesh or client-side); servers set max connection age.
  6. Emit per-method metrics by status code and propagate trace context.
SHOULD
  7. Use hedging only for idempotent reads, with delay near the method's p95.
  8. Use subsetting when a service exceeds ~200 backends.

Exceptions: platform review; time-boxed.
Success metrics: 100% of calls with deadlines; attempts per logical call < 1.1 at p99;
  zero breaking schema changes reaching production; per-pod RPS skew < 1.5.

What I'd Tell the VP#

"Our services talk to each other millions of times a second, and today each team decides how long to wait, how often to retry and how to describe its data. When one service slows down, those choices decide whether we have a small slowdown or a full outage. I'm proposing one standard framework that makes the safe choices automatic, plus a check that stops a team from changing a data format in a way that silently breaks another team. It costs a small platform team, about two engineers, and it removes the most common way our outages spread."

Principal Interview Signals#

SignalWhat It Sounds Like
Defaults over advice"The framework fails a call without a deadline; we don't rely on code review."
Prices the mesh"Sidecars across 20,000 pods are a real line item; proxyless on the hottest paths pays for itself."
Schema governance"One registry, one breaking-change gate, one approver per package."
Edge strategy"External callers get REST with a deprecation policy; internal gets gRPC; the gateway translates."
Language commitment"Every supported language is a library to maintain; I'd support three well."

Staff answers that L7 interviewers find insufficient:

  • "Use a service mesh" without the sidecar cost, upgrade ownership or which paths go proxyless.
  • "Add deadlines everywhere" as guidance rather than a framework default with enforcement.
  • "Use buf/protolint in CI" without who approves breaking changes and how versions are retired.

How Real Companies Built It#

Google — From Stubby to gRPC#

The gRPC project's design-principles post explains that Google ran an internal RPC system called Stubby for over a decade, handling tens of billions of requests per second across its data centers, and built gRPC as a standards-based redesign on HTTP/2 because Stubby was too coupled to internal infrastructure to release. The principles it lists include deadlines and cancellation, flow control, streaming, metadata exchange for cross-cutting concerns such as auth and tracing, and pluggable load balancing and monitoring (gRPC blog).

Staff insight: Deadlines and cancellation were design principles, not add-ons. In an interview, leading with propagated deadlines is using gRPC the way it was designed to be used.

Dropbox — Courier#

Dropbox described Courier, its gRPC-based RPC framework, built while migrating hundreds of services that exchange millions of requests per second from a legacy HTTP/1.1 RPC protocol with protobuf payloads. Courier integrates mTLS service identity, per-method ACLs and rate limits, mandatory deadlines for every method, a time-bounded LIFO queue to shed overload, per-client and per-method stats and tracing. Among its lessons: observability is a feature, and migration takes far longer than the development (Dropbox tech blog).

Staff insight: The value wasn't gRPC itself but the framework around it that made deadlines, identity and metrics mandatory. That's the Principal move: defaults, not guidelines.

Uber — Push Platform on gRPC Bidirectional Streaming#

Uber moved its real-time push platform (RAMEN) from Server-Sent Events over HTTP/1.1 to gRPC bidirectional streaming over QUIC/HTTP/3. Under SSE, acknowledgements came back via separate RPCs every 30 seconds, so a message's delivery state could be unknown for up to 30 seconds, a problem for driver offers valid for about that long. With acks on the same stream, Uber reported p95 connection latency improving by at least 45% and push success rates rising by at least 1–2%, and described handling head-of-line blocking from large messages and keeping an SSE fallback (Uber engineering).

Staff insight: Bidi streaming earned its complexity because acks were the requirement. Pick streaming for a concrete protocol need (acks, flow control), not because it sounds real-time.


Practice Drill#

Drill 1: The Scale-Out That Didn't Help#

Prompt: "Your checkout service calls pricing over gRPC through a Kubernetes ClusterIP Service. During a sale, pricing's p99 went from 40ms to 2s. The autoscaler doubled pricing from 12 to 24 pods, nothing improved, and checkout then ran out of worker threads and started failing. Retries are configured in the checkout code and in the mesh. What happened, and what do you change?"

Staff Answer

Three failures stacked. First, connection pinning: checkout pods hold long-lived HTTP/2 connections to pricing through ClusterIP, which balances per connection, so the 12 new pricing pods received almost no traffic while the original 12 stayed saturated. Per-pod request-rate skew would have shown it immediately. Second, no propagated deadlines: checkout called pricing with a fixed long timeout or none, so when pricing slowed, checkout's handlers piled up waiting until its worker pool was exhausted, turning a pricing slowdown into a checkout outage. Third, retry amplification: retries in the application and the mesh multiply attempts during exactly the brownout, adding load to pricing. Changes: per-request balancing, either by routing through the mesh's L7 balancing or by a headless Service with client-side round_robin, plus a server-side max connection age of ~5 minutes with a grace period so clients rebalance after any scale-out. Deadlines from the edge budget: checkout has 800ms, gives pricing ~250ms (well above its normal 40ms p99), propagates remaining time, and pricing handlers check cancellation so abandoned work stops. Retries in one place only: the gRPC client service config, UNAVAILABLE only, max 3 attempts, 50ms initial backoff, retry throttling on (maxTokens 10, tokenRatio 0.1); remove application and mesh retries. Degradation: if pricing times out, checkout uses the cached quote for items whose price hasn't changed in the last minute, or fails fast with a clear error, rather than holding threads. Signals: per-pod RPS skew alert at max/median > 2, attempts per logical call, in-flight RPCs in checkout, DEADLINE_EXCEEDED by callee. Then a game day: add 2s of latency to pricing in staging and verify checkout sheds cleanly and new pricing pods take traffic within minutes.

Why this is L6:

  • Identifies L4 pinning as the reason the scale-out did nothing, and fixes both balancing and rebalancing.
  • Ties checkout's thread exhaustion to missing deadline propagation and cancellation.
  • Collapses retries to one throttled layer and removes amplification.
  • Adds a degradation path and a test that proves the fix.

What L7 adds:

  • Makes per-request balancing, deadlines and single-layer throttled retries framework defaults, so the next service doesn't rediscover this.
  • Bans L4-only gRPC in the paved road and adds per-pod skew to the platform's SLOs.
  • Prices the mesh vs proxyless choice for the hot checkout path.
❌ Common L5 Trap

"Pricing needs more capacity, so raise the autoscaler's maximum to 50 pods, increase checkout's timeout to 10 seconds so requests don't fail, and add one more retry so transient errors recover."

Why this misses: More pods don't receive traffic while connections are pinned at L4, so capacity isn't the bottleneck. A 10-second timeout makes checkout hold threads longer and exhausts its pool faster. Another retry adds load to the struggling service at the worst moment. The fix is per-request balancing with connection recycling, propagated deadlines with cancellation, and retries in exactly one throttled layer.


Quick Reference Card#

Model:           contract-first RPC over HTTP/2; protobuf on the wire; reliability in the client
Call types:      unary, server streaming, client streaming, bidirectional
Wire format:     fields identified by NUMBER; add fields freely; never reuse numbers or
                 change types; delete with `reserved`; enum zero = UNSPECIFIED
Deadlines:       none by default; propagate remaining budget (converted to timeout);
                 DEADLINE_EXCEEDED on client, CANCELLED on server -> stop work
Retryable:       UNAVAILABLE (canonical); never INVALID_ARGUMENT / NOT_FOUND / PERMISSION_DENIED
Retry config:    service config retryPolicy; maxAttempts capped at 5; committed once
                 response headers arrive; retryThrottling pauses below half of maxTokens
Hedging:         hedgingPolicy (maxAttempts <= 5, hedgingDelay); idempotent reads only
Load balancing:  L4 pins HTTP/2 connections -> per-request LB (L7 proxy, mesh,
                 client round_robin, proxyless xDS) + server max connection age
Keepalive:       client off by default; timeout 20s; server 2h; permit 5 min
                 (faster pings -> GOAWAY too_many_pings)
Messages:        4 MB default receive limit in main implementations; paginate or stream
gRPC-Web:        needs a proxy (Envoy default); unary + server streaming only
Edge:            REST/JSON or GraphQL for external callers; gRPC inside

RED FLAGS
  - Calls without deadlines
  - gRPC behind ClusterIP / L4 only, no connection age limit
  - Retries in app + client + mesh
  - Retrying non-idempotent methods
  - Reused field numbers or changed field types
  - Streams with no resume token
  - Native gRPC promised to browsers
  1. Loading the index…