Why This Matters#
gRPC is not a faster REST. It is a contract-first RPC system whose real features are a typed schema that evolves safely, deadlines that propagate across hops, and cancellation that stops wasted work. Binary encoding and HTTP/2 make it efficient, but efficiency is rarely why it wins or loses an interview. What interviewers probe is the sharp edge that comes with HTTP/2: a gRPC client holds a few long-lived connections and multiplexes every request over them, so connection-level (L4) load balancing quietly sends all of a client's traffic to one backend, and a scale-out adds pods that receive nothing.
That is why "services talk over gRPC" is a sentence interviewers push on. The L5 candidate says "gRPC is faster than JSON." The L6 candidate says "protobuf contracts in a shared repo with breaking-change checks in CI; every call carries a deadline derived from the caller's remaining budget; retries only on UNAVAILABLE, at one layer, capped by a retry-throttling token bucket; per-request load balancing through a mesh or client-side round robin over a headless service, because L4 balancing pins HTTP/2 connections; and servers set a max connection age so clients rebalance after a deploy." The L7 candidate asks who owns the schema repository for 400 services, how breaking changes are governed, and whether the company standardises on proxyless gRPC or a sidecar mesh, because that choice sets the platform team's size for years.
The L5 → L6 gap is not knowing that protobuf is binary. It is knowing that gRPC moves reliability policy (deadlines, retries, balancing) into the client, so a missing deadline or a retry at every layer is a fleet-wide behaviour, not a local bug.
The L5 → L6 → L7 Contrast#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Use gRPC for internal services, it's faster" | "Who are the callers? Internal services get gRPC with protobuf contracts; browsers and partners get REST or gRPC-Web at the edge." | "One RPC standard for the company, one schema registry, one governance process for breaking changes, and an edge story for external callers." |
| Deadlines | "Set a timeout on the client" | Every call has a deadline; propagated with elapsed time deducted; servers check cancellation and stop work | Mandates deadlines in the framework (calls without one fail lint or default to a platform value), as Dropbox did with Courier |
| Load balancing | "Put it behind the load balancer" | L4 pins HTTP/2 connections; per-request balancing via L7 proxy/mesh or client-side policy; max connection age to rebalance | Chooses proxyless vs sidecar mesh fleet-wide, priced in CPU, latency and platform headcount |
| Retries | "Retry failed calls 3 times" | Retry only idempotent methods on retryable codes, at one layer, with throttling; hedging only for idempotent, latency-critical reads | Owns the retry budget as a fleet policy: retries per layer, throttling defaults, and who may enable hedging |
| Schema | "Define a proto file" | Field numbers are forever; add fields, never change types or reuse numbers; reserved for deletions; breaking-change detection in CI | Runs the schema registry and deprecation process; decides who may approve a breaking change |
| Ownership | "Each team writes its own client" | Generated clients; interceptors for auth, tracing, metrics, deadlines owned by platform | Defines the RPC platform contract: what every service gets for free and what it must declare |
Why "Load balancing" separates levels
"Put it behind the load balancer" works for HTTP/1.1, where each request tends to take a connection from a pool and connections churn. A gRPC client opens one HTTP/2 connection per backend address it knows and multiplexes thousands of concurrent RPCs over it. Behind an L4 balancer (a Kubernetes ClusterIP, a TCP load balancer), the balancer sees one long-lived connection and pins it to one pod. Scale from 10 to 20 pods during a spike and the new 10 receive almost nothing, because existing clients never reconnect. The gRPC project's own load-balancing guidance says exactly this: L3/L4 balancers treat each connection as a unit and can't distribute individual requests (gRPC blog). The Staff answer picks per-request balancing (an L7 proxy or mesh, or client-side round robin over a headless service) and sets a server-side max connection age so even pinned clients rebalance. The Principal answer decides that fleet-wide, because it determines whether every pod carries a sidecar.
The 60-Second Pitch#
"Internal service-to-service calls use gRPC with protobuf contracts in a shared schema repo, generated clients, and a CI check that rejects breaking changes. Every call carries a deadline: the edge sets the user-facing budget, say 800ms, and each hop propagates what's left. Retries happen in one place, the gRPC client's service config, only for idempotent methods on UNAVAILABLE, at most 3 attempts with exponential backoff and jitter, capped by retry throttling so a struggling backend sees fewer retries, not more. Balancing is per request: a sidecar mesh or client-side round robin over a headless service, never plain L4, and servers set a max connection age of a few minutes so clients rebalance after scale-outs. For streaming, I'd use server streaming for push with a resume token so reconnects don't lose events. Browsers reach us through REST or gRPC-Web at the edge, because browsers can't speak native gRPC."
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Internal service-to-service RPC | Many services, many languages, low latency, frequent deploys | Unary gRPC, protobuf contracts, deadlines, per-request LB, interceptors for auth/tracing | L4 pinning → hot pods; missing deadlines → cascading failure; retry amplification | p99 within budget; no cascade from one slow dependency |
| Streaming and push | Server-to-client events, long-lived connections, mobile networks | Server or bidi streaming, keepalives, resume tokens, max connection age, graceful GOAWAY on deploy | Streams pinned to old pods; lost events on reconnect; keepalive abuse → too_many_pings | Every event delivered or detectably missed; reconnect within seconds |
| External or browser-facing API | Browsers, partners, caching, human-readable debugging | REST/JSON or gRPC-Web via a proxy; JSON transcoding at the edge; gRPC behind | Browser can't do client or bidi streaming; partners can't read protobuf; HTTP caches don't apply | Contract stable for external clients for years |
🎯 Staff Move: "I'll use gRPC for the internal mesh, the first intent, and keep the public API as REST at the gateway. External clients need caching, curl-ability and a contract they can read; internal services need deadlines, typed contracts and efficient multiplexing. The gateway translates, so neither side compromises."
The Staff Positions#
| Position | Rationale |
|---|---|
| Every RPC has a deadline, propagated across hops | gRPC's default is no deadline; one slow dependency otherwise exhausts every caller upstream. |
| Per-request load balancing, never plain L4 | HTTP/2 connections are long-lived and multiplexed; L4 pins all of a client's traffic to one backend. |
| Retries at exactly one layer, with throttling | 3 layers × 3 attempts is 27 calls per user request at the bottom during an outage. |
| Field numbers are forever | Protobuf identifies fields by number on the wire; reuse or type changes silently corrupt data across versions. |
| Breaking-change detection in CI | Humans miss renumbered fields; a schema linter doesn't. |
| Servers cap connection age | Forces clients to reconnect and rebalance after scale-outs and deploys. |
| Browsers and partners get REST or gRPC-Web at the edge | Browsers can't speak native gRPC; partners need a stable, readable contract. |
Architecture & Internals#
Only five internals change design decisions: HTTP/2 multiplexing, protobuf's wire format, the four call types, deadlines and status codes, and channels and load-balancing policies. The comparison with REST and GraphQL is in REST vs gRPC vs GraphQL.
HTTP/2: One Connection, Many Streams#
Each RPC is an HTTP/2 stream: a HEADERS frame (method path like /orders.v1.OrderService/GetOrder, deadline as grpc-timeout, metadata), length-prefixed protobuf messages in DATA frames, and trailing headers carrying grpc-status. Many streams share one TCP connection; each has flow-control windows so a slow reader on one stream doesn't stall the others at the HTTP/2 layer.
Why it matters in design: connection setup (TCP + TLS) happens once, then every RPC is a few frames, which is where gRPC's latency and CPU advantage over JSON-over-HTTP/1.1 comes from. The price is that the connection is the unit an L4 balancer sees, and it lives for hours. HTTP/2 still runs on TCP, so a lost packet stalls every stream on that connection (TCP head-of-line blocking); on lossy mobile networks that is one reason teams look at HTTP/3, as Uber did for its push platform (below).
Protobuf: The Wire Format Is Field Numbers#
A protobuf message on the wire is a sequence of (field number, wire type, value) entries. Names never appear on the wire. That single fact sets every schema-evolution rule:
syntax = "proto3";
package orders.v1;
message Order {
string id = 1;
int64 amount_minor = 2; // money as integer minor units
string currency = 3;
OrderStatus status = 4;
google.protobuf.Timestamp created_at = 5;
reserved 6; // was 'coupon_code', removed in 2025
reserved "coupon_code";
optional string note = 7; // proto3 optional: presence is tracked
}
enum OrderStatus {
ORDER_STATUS_UNSPECIFIED = 0; // zero value must mean 'unknown'
ORDER_STATUS_PENDING = 1;
ORDER_STATUS_PAID = 2;
}
| Change | Safe? | Why |
|---|---|---|
| Add a new field with a new number | Yes | Old readers skip unknown fields; new readers see the default when absent |
| Rename a field (same number, same type) | Wire-safe; breaks JSON and generated code | JSON mapping uses names; callers' code stops compiling |
| Delete a field | Only with reserved number and name | Otherwise someone reuses the number later |
| Reuse a field number | No | Old data decodes into the new field with the wrong meaning |
Change a field's type (e.g. int32 → string) | No | Wire types differ; decoding fails or silently misreads |
| Add an enum value | Yes, if readers handle unknown values | Old readers see an unrecognised number |
| Change a method's request or response type | No | It's a different contract; add a new method |
🎯 Staff Insight: "Protobuf compatibility is a promise about field numbers, not names. So deletions get
reserved, type changes become new fields, and CI runs a breaking-change check against the last released schema, because the failure mode is silent data corruption between two services deployed a day apart."
The Four Call Types#
| Type | Shape | Use for | Watch out for |
|---|---|---|---|
| Unary | 1 request → 1 response | 90%+ of service calls | Nothing special: deadlines, retries, LB all behave as expected |
| Server streaming | 1 request → N responses | Push feeds, large result sets, watch APIs | Stream pinned to one backend for its lifetime; resume tokens for reconnect |
| Client streaming | N requests → 1 response | Uploads, batched telemetry | Retries are hard once data has been sent; partial-progress semantics needed |
| Bidirectional streaming | N ↔ M interleaved | Chat, real-time sync, push with acks | Not supported by gRPC-Web in browsers; ordering per direction only; long-lived load |
Streams are not a message queue. A stream lives on one connection to one backend; if that backend restarts, the stream ends and the client must reconnect and resume. Anything that must survive that needs an offset or resume token, and usually a durable log behind the server. The push design in Live Updates and the transport comparison in WebSockets vs SSE vs Long Polling cover the alternatives.
Deadlines, Cancellation and Status Codes#
When a server acts as a client, gRPC can propagate the incoming deadline, converting it to a timeout with elapsed time already deducted so clock skew between hosts doesn't matter; some languages do this by default, others need it enabled (deadlines guide). When the deadline passes, the client fails with DEADLINE_EXCEEDED and the server's call is cancelled; the application still has to check for cancellation and stop work.
| Status code | Meaning | Retry? |
|---|---|---|
OK | Success | — |
UNAVAILABLE | Transient: connection failure, server shutting down, overloaded | Yes, with backoff (the canonical retryable code) |
DEADLINE_EXCEEDED | Ran out of time | Rarely: the budget is spent; retrying usually just adds load |
RESOURCE_EXHAUSTED | Quota or rate limit hit | Only after the server's pushback delay |
INVALID_ARGUMENT, NOT_FOUND, PERMISSION_DENIED | Caller's problem | Never |
INTERNAL, UNKNOWN | Server bug or unexpected failure | No by default; a bug retried is a bug amplified |
Channels, Resolvers and Load-Balancing Policies#
A channel is a client's handle to a logical service. It uses a resolver (DNS, a registry, or xDS from a control plane) to get addresses, opens a subchannel per address, and applies an LB policy per RPC. With pick_first (the default in most implementations) the client sends everything to one address; with round_robin it spreads RPCs across all ready subchannels. Proxyless setups get the address list and policy (weighted round robin, outlier ejection) from an xDS control plane, the same API Envoy uses; see Envoy, Kong & NGINX.
Keepalive defaults matter for long-lived connections: client keepalive pings are disabled by default, the server's keepalive time defaults to 2 hours, the ping timeout is 20 seconds, and servers by default reject pings more frequent than every 5 minutes, answering abusive clients with GOAWAY (keepalive guide). Client authors must coordinate keepalive settings with service owners.
Core Usage — "The Entire Game": The Service Contract#
In Kafka the game is the partition key. In gRPC it is the service contract: five things every service declares, most of which live in the client. Get them right and the RPC layer degrades gracefully; get them wrong and every caller inherits the mistake.
The service contract (per method):
1. schema -> request/response messages, field numbers, versioned package
2. deadline -> the budget this method gets, and what callers propagate
3. retry -> is it idempotent? which codes? how many attempts? hedging?
4. balancing -> how each RPC picks a backend, and how connections rebalance
5. signals -> status codes it returns, metrics and traces it emits
Step 1: Design the Schema for Evolution#
service OrderService {
rpc GetOrder(GetOrderRequest) returns (Order);
rpc ListOrders(ListOrdersRequest) returns (ListOrdersResponse);
rpc PlaceOrder(PlaceOrderRequest) returns (PlaceOrderResponse);
}
message PlaceOrderRequest {
string request_id = 1; // idempotency key: makes PlaceOrder safe to retry
string customer_id = 2;
repeated LineItem items = 3;
}
message ListOrdersRequest {
string customer_id = 1;
int32 page_size = 2; // server caps at e.g. 100
string page_token = 3; // opaque cursor, never an offset
google.protobuf.FieldMask read_mask = 4; // caller asks only for fields it needs
}
Rules that hold up over years: a request and response message per method (never a bare primitive, so fields can be added); a versioned package (orders.v1), with v2 as a new service rather than in-place breakage; idempotency keys on mutating methods so they can be retried; cursor pagination with server-capped page sizes; and field masks for wide resources. The contract discipline is the same as in API Contracts; the schema-change process is in Schema Design.
Step 2: Budget the Deadline#
User-facing budget: 800 ms at the edge (p99 target 600 ms)
edge -> checkout deadline 800 ms
checkout local work ~40 ms
checkout -> pricing (parallel) remaining ~760 ms, pricing p99 120 ms
checkout -> inventory (parallel) remaining ~760 ms, inventory p99 150 ms
checkout -> payments (serial) remaining ~600 ms, payments p99 300 ms
Rules:
- each hop propagates (remaining - local reserve), never a fresh fixed timeout
- a callee whose p99 exceeds its share is a design problem, not a timeout problem
- servers check ctx/cancellation before expensive steps and stop when cancelled
The Latency Budget calculator does this split; the Latency Numbers table grounds each hop's share.
Step 3: Configure Retries and Hedging in One Place#
gRPC clients read a service config (JSON, from DNS, xDS or local configuration) that defines per-method retry or hedging policies. Neither is enabled unless configured, and maxAttempts values above 5 are treated as 5 (retry design A6).
{
"methodConfig": [
{
"name": [{ "service": "orders.v1.OrderService", "method": "GetOrder" }],
"timeout": "0.3s",
"retryPolicy": {
"maxAttempts": 3,
"initialBackoff": "0.05s",
"maxBackoff": "0.5s",
"backoffMultiplier": 2,
"retryableStatusCodes": ["UNAVAILABLE"]
}
},
{
"name": [{ "service": "catalog.v1.CatalogService", "method": "GetProduct" }],
"hedgingPolicy": {
"maxAttempts": 2,
"hedgingDelay": "0.02s",
"nonFatalStatusCodes": ["UNAVAILABLE"]
}
}
],
"retryThrottling": { "maxTokens": 10, "tokenRatio": 0.1 }
}
Three behaviours to know (retry guide, hedging guide):
- Commit point. Once response headers arrive, the RPC is committed and won't be retried. Retries cover failures before the server started answering.
- Throttling. Each failure costs a token and each success earns back
tokenRatio; when the count falls below half ofmaxTokens, retries and hedges stop. During an outage the client automatically backs off to roughly one attempt per request. - Hedging sends a second copy after
hedgingDelaywithout waiting for a failure and takes the first success, cancelling the rest. It cuts tail latency and multiplies load; use only on idempotent reads, with a delay near the method's p95. Servers can push back withgrpc-retry-pushback-ms.
Step 4: Pick Per-Request Balancing#
| Option | How | Per-request? | Cost | Pick when |
|---|---|---|---|---|
L4 balancer / ClusterIP | Kernel or TCP LB picks a backend per connection | No | Lowest | Never for gRPC alone; OK only with max connection age and many clients |
| L7 proxy (gateway, Envoy) | Proxy terminates HTTP/2 and balances each RPC | Yes | Extra hop (~sub-ms to ~1 ms), proxy CPU | Edge, trust boundaries, mixed clients |
| Sidecar mesh | A proxy beside every pod | Yes | Sidecar CPU/memory per pod, extra hop each side | Polyglot fleets wanting uniform mTLS, retries, telemetry |
Client-side (round_robin over headless service) | Client resolves all pod IPs, balances itself | Yes | No hop; client complexity per language | High-volume internal traffic, few languages |
| Proxyless xDS | Client gets endpoints and policy from a control plane | Yes | Control plane to run; library support per language | Large fleets wanting mesh features without sidecars |
Whatever the choice, set max connection age on servers (minutes, with a grace period) so long-lived connections get recycled and rebalanced, and drain with GOAWAY on shutdown so in-flight RPCs finish. The design space is in Load Balancing and Service Registry.
Step 5: Signals — Status Codes, Metrics, Traces#
Return precise status codes (INVALID_ARGUMENT vs UNAVAILABLE decides whether callers retry). Emit per-method RED metrics (rate, errors by code, duration histogram) from server and client interceptors, and propagate trace context in metadata. Without client-side metrics you cannot see L4 pinning, because each server looks healthy; the skew only shows when you compare per-backend request rates. The trace plumbing is in Distributed Tracing.
🎯 Staff Move: "Every method gets a deadline share from the edge budget, an idempotency decision that determines whether retries are allowed, and a retry policy in service config with throttling. Retries live only in the gRPC client, not also in the mesh and the application, because three layers of three attempts is 27× load during exactly the outage we're trying to survive."
The Tunable Tradeoff — Where Reliability Logic Lives#
Every gRPC platform decision moves along one axis: how much reliability policy lives in the client library vs in proxies?
| Setting | Client-heavy end | Proxy-heavy end | Who pays at the client-heavy end |
|---|---|---|---|
| Load balancing | Client-side / proxyless xDS | Sidecar or central L7 proxy | Library maintainers in every language |
| Retries and hedging | Service config in the client | Proxy retry policy | Service teams if configs drift |
| mTLS and identity | Library-managed certificates | Sidecar-terminated TLS | Platform, for rotation in N runtimes |
| Telemetry | Interceptors | Proxy access logs and metrics | Teams on languages without good interceptors |
| Latency and CPU | No extra hop | +1–2 hops per call, sidecar CPU per pod | At the proxy end: every request, and the compute bill |
🎯 Staff Move: "For a polyglot fleet of 300 services I'd take the sidecar mesh: uniform mTLS, retries and telemetry without four client libraries to keep in sync. For one or two very high-QPS paths in a single language, I'd go proxyless with xDS to remove the hops. Same control plane, two data planes."
Anti-Patterns — What Kills gRPC Deployments#
1. No Deadlines#
The default is to wait forever. One slow downstream holds every caller's threads, and the outage climbs the call graph. Fix: framework-enforced deadlines; propagate remaining budget; servers honour cancellation.
2. L4 Load Balancing for Long-Lived Connections#
ClusterIP or a TCP balancer in front of gRPC pins each client connection to one pod; after a scale-out the new pods idle while old ones melt. Fix: per-request balancing (L7 proxy, mesh or client-side) plus server max connection age.
3. Retries at Every Layer#
Application retries 3×, the generated client retries 3×, the mesh retries 3×: 27 attempts at the bottom per user request during a brownout. Fix: one layer retries, with throttling; others pass errors through.
4. Retrying Non-Idempotent Methods#
PlaceOrder retried on DEADLINE_EXCEEDED creates two orders if the first actually succeeded. Fix: retries only on methods with an idempotency key; see Idempotency.
5. Reusing Field Numbers or Changing Types#
A deleted field's number is reused for a new field; services deployed a day apart now read each other's data with the wrong meaning, with no error. Fix: reserved on every deletion; breaking-change checks in CI against the last release.
6. Large Messages#
A 50 MB response hits the common 4 MB default receive limit, or worse, gets raised past it and spikes server memory while blocking other streams on the connection. Fix: paginate, stream in chunks, or return a pointer to object storage.
7. Streams Treated as Durable#
A bidirectional stream carries events with no resume token; every deploy, keepalive timeout or mobile network change silently drops whatever was in flight. Fix: offsets or resume tokens, a durable source behind the stream, client reconnect with backoff and jitter.
The Technology Landscape — Head-to-Head Comparison#
| Dimension | gRPC | REST + JSON | GraphQL | Apache Thrift | Async messaging (queues/logs) |
|---|---|---|---|---|---|
| Contract | Protobuf IDL, generated code | OpenAPI (optional), hand-written or generated | Schema with typed queries | Thrift IDL, generated code | Message schemas (Avro, protobuf) |
| Transport | HTTP/2 | HTTP/1.1 or 2 | HTTP (usually POST) | Pluggable (TCP, HTTP) | Broker |
| Streaming | Unary, server, client, bidi | SSE or WebSocket on the side | Subscriptions (extra infra) | Limited | Native, async |
| Browser support | Via gRPC-Web proxy (unary and server streaming) | Native | Native | Poor | Not directly |
| Caching | No HTTP caching | HTTP caches and CDNs | Hard (POST, per-query) | No | N/A |
| Deadlines/cancellation | Built in, propagated | Per-client timeouts | Per-client timeouts | Per-client | Message TTLs |
| Pick when | Internal service mesh, polyglot, low latency, streaming | Public APIs, partners, cacheable reads | Client-driven aggregation for many UIs | Existing Thrift estates | Work that can wait; decoupled producers |
🎯 Staff Insight: "gRPC versus REST is a question about who the caller is. Internal callers I control get gRPC for typed contracts, deadlines and efficiency. Callers I don't control get REST, because they need caching, curl and a contract they can read without a code generator."
Patterns#
Pattern 1: gRPC Inside, REST at the Edge#
The gateway owns external contracts, caching and auth; internal services speak gRPC with propagated deadlines. JSON transcoding (mapping HTTP routes to RPCs via annotations) lets one proto serve both. The edge design is in Edge Gateway.
Pattern 2: Sidecar Mesh, or Proxyless xDS#
Envoy-style sidecars terminate mTLS, balance each RPC, apply retry and outlier-ejection policy, and emit uniform telemetry; the price is a proxy per pod (illustratively 0.1–0.25 vCPU and 50–150 MB each, so hundreds of vCPU across 2,000 pods) and extra hops. Proxyless gRPC consumes the same xDS config inside the library: no sidecar, no hop, but only for languages with mature support.
Pattern 3: Server-Streaming Push With Resume#
The client opens Subscribe(SubscribeRequest{resume_token}); the server streams events from a durable log, each carrying a token. On reconnect (deploy, network change, max connection age) the client resumes from its last token. Keepalives detect dead connections; the server drains with GOAWAY and clients reconnect with jitter. See Push vs Poll.
Scaling#
The Numbers#
| Resource | Default or documented figure | Design note |
|---|---|---|
| Deadline | None by default | Always set one; propagate remaining budget |
Retry / hedging maxAttempts | Values above 5 treated as 5 | 2–3 is typical for retries, 2 for hedging |
| Retry throttling | Off unless configured; retries pause below half of maxTokens | Example config: maxTokens 10, tokenRatio 0.1 |
| Transparent retries | Don't count toward maxAttempts | Cover RPCs that never reached server application logic |
| Max receive message size | 4 MB default in the main implementations | Raise deliberately or stream/paginate |
| Client keepalive | Disabled by default; timeout 20s | Don't set below ~1 minute without the service owner |
| Server keepalive / permit time | 2 hours / 5 minutes | Faster client pings get GOAWAY too_many_pings |
| Max connection age | Unlimited by default | Set to minutes for rebalancing, with a grace period |
| gRPC-Web | Unary + server streaming (text mode); via a proxy | No client or bidi streaming from browsers |
Scaling Moves in Order#
- Deadlines and per-request balancing everywhere; max connection age on servers.
- One retry layer with throttling, configured centrally.
- Schema registry and breaking-change gates once more than a few teams share protos.
- Connection subsetting when clients × backends makes full meshes expensive.
- Mesh or proxyless xDS for uniform policy; proxyless for the hottest paths.
- HTTP/3 for lossy client networks (mobile push) where TCP head-of-line blocking hurts.
🎯 Staff Move: "Past a few hundred pods per service I'd subset: each client connects to, say, 20 backends chosen deterministically, rather than all 300. Full-mesh connections cost memory and health-check traffic and buy no balance a subset doesn't."
Failure Modes & Recovery#
1. Hot Pods After Scale-Out (Connection Pinning)#
- Symptom: Autoscaler doubles a service from 10 to 20 pods; p99 doesn't improve; 10 pods at 90% CPU, 10 nearly idle.
- Root cause: Clients hold long-lived HTTP/2 connections balanced at L4; new pods get no existing connections.
- Detection: Per-pod request rate skew (max/median > 2); per-pod CPU spread; client-side per-backend metrics.
- Fix: Roll clients to force reconnects; enable round robin or route through the mesh.
- Prevention: Per-request balancing; server max connection age of a few minutes; alert on per-pod RPS skew.
2. Retry Storm#
- Symptom: A dependency brownout turns into a hard outage; its request rate triples while its success rate falls.
- Root cause: Retries at several layers, no throttling, retrying
DEADLINE_EXCEEDED. - Detection: Ratio of attempts to logical calls (client metrics); request rate on the dependency vs upstream traffic.
- Fix: Disable retries at all but one layer; enable throttling; shed load at the dependency.
- Prevention: Central retry policy with throttling; lint for retry config in more than one layer.
3. Cascading Exhaustion From Missing Deadlines#
- Symptom: One downstream slows from 50ms to 5s; within minutes every upstream service's thread pool or goroutine count is maxed.
- Root cause: Calls without deadlines, or fixed long timeouts not derived from the caller's budget.
- Detection: In-flight RPC count per service; latency of calls to the slow dependency; caller saturation.
- Fix: Short emergency deadlines via config; circuit-break the dependency; shed.
- Prevention: Framework-enforced deadlines; propagation; cancellation checks in handlers. See Cascading Failures.
4. Silent Data Corruption From a Breaking Schema Change#
- Symptom: After a deploy, some orders show wrong amounts or missing fields; no errors logged.
- Root cause: A field number reused or a type changed; old and new binaries decode each other's messages differently.
- Detection: Data-quality checks; mismatch between producer and consumer schema versions; contract tests.
- Fix: Roll back the producer; repair affected records; reintroduce the change as a new field.
- Prevention: Breaking-change detection in CI;
reservedfor deletions; schema owner approval.
5. Streams Stuck on Old Pods During a Deploy#
- Symptom: After a rollout, half the clients' push streams still point at pods being terminated; events drop when those pods exit.
- Root cause: Long-lived streams with no graceful drain or resume; termination grace shorter than stream life.
- Detection: Active streams per pod over time; reconnect rate spikes; missed-event reports.
- Fix: Send
GOAWAY, let clients reconnect with jitter and resume from tokens. - Prevention: Max connection age; resume tokens; termination grace that covers drain; reconnect jitter so 100K clients don't return in the same second.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Connection pinning | Per-pod RPS skew | One service's p99 | Per-request LB, max connection age | Platform (LB); service team (config) |
| Retry storm | Attempts per logical call | The dependency, then its callers | One retry layer, throttling | Platform policy; calling teams |
| Missing deadlines | In-flight RPCs climbing | Entire call graph upstream | Emergency deadlines, circuit breaking | Calling service; platform for enforcement |
| Breaking schema change | Data-quality checks | Every consumer of the message | Roll back, repair data | Schema owner; producing team |
| Stuck streams on deploy | Streams per terminating pod | Push clients of that service | GOAWAY drain, resume | Streaming service team |
| Keepalive abuse | too_many_pings GOAWAYs | Clients of that server | Align keepalive settings | Client team with service owner |
When to Use vs. Alternatives#
| Need | Pick | Why |
|---|---|---|
| Internal service-to-service calls, many languages | gRPC | Typed contracts, deadlines, efficient multiplexing |
| Public API for partners and developers | REST + JSON (OpenAPI) | Readable, cacheable, universal tooling |
| Many UIs needing different shapes of the same data | GraphQL at the edge | Client-driven queries; see REST vs gRPC vs GraphQL |
| Browser calling services directly | REST, or gRPC-Web via a proxy | Browsers can't speak native gRPC |
| Server push to clients | gRPC server streaming, SSE or WebSockets | Depends on client platform and proxy support |
| Work that can wait, decoupled producers | A queue or log | See RabbitMQ & SQS and Kafka |
When NOT to Use gRPC#
- Public, cacheable APIs. HTTP caching and CDNs don't apply; partners want JSON they can read.
- Browser-first products without a proxy layer. gRPC-Web supports unary and server streaming only, and always needs a proxy.
- Async work. A call that can wait belongs on a queue; RPC couples the caller's availability to the callee's.
- Tiny systems with one language and two services. The schema pipeline and LB setup cost more than they save.
- Environments where L7 balancing isn't possible and clients can't do it either. You'll fight pinning forever.
Operational Concerns#
What the On-Call Actually Does#
- Reads errors by status code, not just error rate: a jump in
UNAVAILABLEis a capacity or connectivity issue;INVALID_ARGUMENTafter a deploy is a contract issue;DEADLINE_EXCEEDEDpoints at a slow dependency. - Checks per-pod request skew after every scale-out and deploy; rebalances by recycling connections.
- Watches attempts per logical call to catch retry storms early, and flips retry policy centrally if needed.
- Traces slow requests across hops to find which callee ate the budget.
- Manages streaming drains during deploys: GOAWAY, grace periods, reconnect jitter.
- Reviews schema changes flagged by the breaking-change gate and approves or rejects exceptions.
Key Metrics & Alerts#
| Metric | Healthy | Alert |
|---|---|---|
grpc_server_handled_total by grpc_code | Errors < SLO | UNAVAILABLE or INTERNAL rate above SLO for 5 min |
| Server latency histogram p99 per method | Within the method's budget share | > budget share for 10 min |
| Per-pod request rate max / median | < 1.5 | > 2 (pinning) |
| Client attempts per logical RPC | ~1.0 | > 1.2 sustained (retry pressure) |
| In-flight RPCs per server | Steady | Climbing with flat throughput (stuck callee) |
| Active streams per pod | Even across pods | Concentrated on few pods |
DEADLINE_EXCEEDED rate per caller → callee | Low | Spike (callee slow or budget wrong) |
Interview Application — Staff-Level Plays#
Which Case Studies Use gRPC#
| Case Study | How gRPC Is Used | Key Pattern |
|---|---|---|
| Load Balancing | Per-request vs per-connection balancing | L4 pinning; L7 or client-side balancing |
| Service Registry | Resolvers and xDS feed channel address lists | Clients fail static when the registry is wrong |
| Edge Gateway | REST at the edge, gRPC behind; transcoding | The gateway owns external contracts |
| Cascading Failures | Deadlines, retries and throttling as cascade controls | One retry layer; propagated budgets |
| Distributed Tracing | Trace context in gRPC metadata via interceptors | Uniform instrumentation from the framework |
| Live Updates | Server streaming for push | Resume tokens; drain on deploy |
Every System Design Question Has a gRPC Moment#
- Ride hailing: "Driver location updates come in over a client stream from the gateway to the location service; offers go out over a bidirectional stream with an ack per message, so we know within a second whether a driver saw the offer."
- Checkout: "The edge sets an 800ms deadline; checkout fans out to pricing and inventory in parallel with the remaining budget and calls payments serially; only reads retry, and
PlaceOrdercarries a request ID so a retry can't double-charge." - Autoscaling: "New pods only help gRPC traffic if clients rebalance, so servers cap connection age at 5 minutes and the mesh balances per request."
What Interviewers Probe#
| After You Say... | They Will Ask... | What They're Evaluating |
|---|---|---|
| "Services use gRPC" | "You scale from 10 to 20 pods. Does load spread?" | L4 pinning; per-request LB; max connection age |
| "We set timeouts" | "Service C is slow. What happens to A?" | Deadline propagation and cancellation |
| "Retry on failure" | "Which codes? Which methods? At which layer?" | Idempotency, retryable codes, amplification, throttling |
| "Protobuf handles versioning" | "You delete a field. What could go wrong next year?" | Field-number reuse; reserved; CI checks |
| "Stream events to clients" | "What happens during a deploy?" | GOAWAY drains, resume tokens, reconnect jitter |
| "gRPC for the public API" | "How do browsers and partners call it?" | gRPC-Web limits; REST or transcoding at the edge |
L5 vs L6 vs L7 Responses#
| Scenario | L5 Answer | L6 / Staff Answer | L7 / Principal Answer |
|---|---|---|---|
| "Design internal communication for 50 services" | gRPC with protobuf | Shared proto repo with CI gates, deadlines from edge budget, one retry layer with throttling, per-request LB, interceptors for auth and tracing | RPC platform contract: framework defaults, schema governance, mesh vs proxyless decision, and the edge strategy for external callers |
| "Scale-out didn't help latency" | Add more pods | Diagnoses connection pinning from per-pod skew; enables round robin or mesh; max connection age | Makes per-pod skew a platform SLO and bans L4-only gRPC in the paved road |
| "A slow service took down checkout" | Increase timeouts | Propagated deadlines, cancellation checks, circuit breaker, retry throttling | Framework refuses calls without deadlines; game days that slow a dependency by 5s |
| "Should the public API be gRPC?" | Yes, for performance | REST for partners, gRPC internally, transcoding at the gateway | Owns the external contract lifecycle: versioning policy, deprecation windows, SDK generation |
The Staff gRPC Checklist#
- Caller first: "Internal callers get gRPC; external and browser callers get REST or gRPC-Web at the edge."
- Deadlines: "Edge budget of 800ms, propagated minus elapsed at every hop; handlers stop on cancellation."
- Retries: "One layer,
UNAVAILABLEonly, max 3 attempts with backoff and jitter, throttled; idempotent methods only." - Balancing: "Per-request via mesh or client round robin; max connection age of a few minutes."
- Schema: "Field numbers forever,
reservedon deletes, breaking-change gate in CI, new package for v2." - Signals: "Errors by status code, per-pod request skew, attempts per call, traces through metadata."
🎯 Staff Insight: Don't use gRPC for public cacheable APIs, for browsers without a proxy, as a substitute for a queue, or behind plain L4 balancing. The strongest gRPC signal is explaining why doubling the pods didn't help, and what you'd change so it does.
Evaluation Rubric#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Contracts | Proto files | Evolution rules, CI gates, versioned packages | Schema governance and registry for the org |
| Reliability | Timeouts and retries | Propagated deadlines, single-layer throttled retries, hedging used sparingly | Framework-enforced defaults; fleet retry budget |
| Load balancing | "Behind the LB" | Pinning understood; per-request LB; connection age | Mesh vs proxyless decision, priced |
| Streaming | "Use bidi streaming" | Resume tokens, drains, keepalive coordination | Push platform standard across products |
| Choice | gRPC for everything | gRPC inside, REST/GraphQL at the edge | External API lifecycle and SDK strategy |
Strong hire signals
| Signal | What It Sounds Like |
|---|---|
| Pinning awareness | "HTTP/2 connections are long-lived, so L4 balancing pins them; new pods need clients to reconnect." |
| Budget thinking | "The edge has 800ms; every hop gets what's left, and handlers stop when it's gone." |
| Retry discipline | "One layer retries, with a token bucket, and only methods with an idempotency key." |
| Schema discipline | "Field numbers are forever; deletions are reserved." |
| Caller-based choice | "gRPC for callers I control, REST for callers I don't." |
Lean no-hire signals
| Signal | Why It Misses the Bar |
|---|---|
| "gRPC because it's faster" with nothing else | Benchmark choice, no contract or reliability reasoning |
| No deadlines | Will ship cascading failures |
| Retries everywhere | Will amplify outages |
| gRPC exposed straight to browsers | Doesn't know the platform limits |
Common false positives
- Protobuf encoding trivia ≠ contract governance. Ask how they'd stop a field-number reuse.
- Mesh enthusiasm ≠ judgment. Ask what the sidecars cost across 2,000 pods and who upgrades them.
Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
At Staff level gRPC is a well-configured client and server. At Principal level it is the company's RPC platform: the framework every service is generated from, the place where deadlines, retries, identity and telemetry become defaults, and the schema estate that 400 services depend on. The L7 question is "What does every service get for free from the RPC framework, who governs the schemas, and do we put reliability logic in libraries or in a mesh?"
🧭 Principal Move: "I want the RPC framework to make the right thing the default: no call without a deadline, retries only where the method is marked idempotent, mTLS identity on every connection, metrics by caller and method out of the box. Then reliability stops depending on every team remembering four rules."
The Org-Level Fault Line#
Smart clients (libraries, proxyless) vs smart network (sidecar mesh).
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Libraries per language | No extra hops, lowest latency and CPU | Features drift across languages; every fix needs N releases and every team to upgrade | Platform, maintaining N libraries; teams on lagging languages |
| Sidecar mesh | Uniform policy in any language; upgrades without app redeploys | Proxy per pod (CPU, memory), extra hops, mesh upgrades become platform-critical | Compute budget; platform team for the mesh |
| Proxyless xDS for hot paths + mesh elsewhere | Best latency where it matters, uniformity elsewhere | Two data planes to reason about | Platform, in complexity |
| No standard | Each team chooses | Inconsistent deadlines, retries and telemetry; incidents debugged from scratch | Incident responders |
The Principal default: a mesh for the polyglot majority with one control plane, proxyless gRPC on the few highest-volume paths, one schema registry with breaking-change gates, and framework defaults (deadline required, retry throttling on, interceptors for auth/metrics/tracing) shipped in a thin library per supported language.
🧭 Principal Insight: "The question isn't library versus mesh. It's how many languages we commit to supporting well. Every supported language is a library to maintain; a mesh is how you buy your way out of that, priced in sidecar CPU."
Cost Model#
Assumptions: sidecar ~0.15 vCPU and ~100 MB per pod; vCPU ~$25/month; loaded engineer ~$21K/month. Directional only.
| Scale | Fleet | Mesh compute/month | Platform people | Notes |
|---|---|---|---|---|
| Startup | 15 services, 100 pods | ~$400 | 0.25 FTE (~$5K) | Libraries and a gateway suffice; mesh is optional |
| Growth | 150 services, 2,000 pods | ~$7.5K | 2 FTE (~$42K) for RPC framework, schema registry, mesh | Schema governance starts paying for itself |
| Enterprise | 1,000 services, 20,000 pods | ~$75K | 6–10 FTE (~$170K) | Proxyless for hot paths saves a share of sidecar cost; subsetting needed |
At every scale the people cost dominates the proxy cost, and the hidden cost is incidents caused by inconsistent defaults: one missing deadline in a hot path can cost more in an afternoon than the platform team costs in a quarter.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse |
|---|---|---|
| Field numbers and wire types in shipped protos | One-way | Data written with them is decoded that way forever |
| Public API protocol (REST vs gRPC) for external clients | One-way-ish | Years of client support and SDKs |
| Mesh vs proxyless | Two-way-ish | Fleet migration of data planes; policy reimplementation |
| Package versioning scheme | One-way-ish | Every generated client |
| Retry and deadline defaults | Two-way | Config change, but behaviour changes fleet-wide |
The Standard I'd Write#
RFC-RPC-001: Internal RPC Contract
Scope: All synchronous service-to-service calls.
MUST
1. Use gRPC with protobuf schemas in the central registry; packages versioned (x.v1).
2. Pass breaking-change checks against the last release; deletions use `reserved`.
3. Set a deadline on every call, derived from the caller's remaining budget;
handlers check cancellation before expensive work.
4. Retry only methods marked idempotent (or carrying an idempotency key), only on
UNAVAILABLE, in the client service config, with retry throttling enabled.
No retries in application code or mesh for the same call.
5. Use per-request load balancing (mesh or client-side); servers set max connection age.
6. Emit per-method metrics by status code and propagate trace context.
SHOULD
7. Use hedging only for idempotent reads, with delay near the method's p95.
8. Use subsetting when a service exceeds ~200 backends.
Exceptions: platform review; time-boxed.
Success metrics: 100% of calls with deadlines; attempts per logical call < 1.1 at p99;
zero breaking schema changes reaching production; per-pod RPS skew < 1.5.
What I'd Tell the VP#
"Our services talk to each other millions of times a second, and today each team decides how long to wait, how often to retry and how to describe its data. When one service slows down, those choices decide whether we have a small slowdown or a full outage. I'm proposing one standard framework that makes the safe choices automatic, plus a check that stops a team from changing a data format in a way that silently breaks another team. It costs a small platform team, about two engineers, and it removes the most common way our outages spread."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Defaults over advice | "The framework fails a call without a deadline; we don't rely on code review." |
| Prices the mesh | "Sidecars across 20,000 pods are a real line item; proxyless on the hottest paths pays for itself." |
| Schema governance | "One registry, one breaking-change gate, one approver per package." |
| Edge strategy | "External callers get REST with a deprecation policy; internal gets gRPC; the gateway translates." |
| Language commitment | "Every supported language is a library to maintain; I'd support three well." |
Staff answers that L7 interviewers find insufficient:
- "Use a service mesh" without the sidecar cost, upgrade ownership or which paths go proxyless.
- "Add deadlines everywhere" as guidance rather than a framework default with enforcement.
- "Use buf/protolint in CI" without who approves breaking changes and how versions are retired.
How Real Companies Built It#
Google — From Stubby to gRPC#
The gRPC project's design-principles post explains that Google ran an internal RPC system called Stubby for over a decade, handling tens of billions of requests per second across its data centers, and built gRPC as a standards-based redesign on HTTP/2 because Stubby was too coupled to internal infrastructure to release. The principles it lists include deadlines and cancellation, flow control, streaming, metadata exchange for cross-cutting concerns such as auth and tracing, and pluggable load balancing and monitoring (gRPC blog).
Staff insight: Deadlines and cancellation were design principles, not add-ons. In an interview, leading with propagated deadlines is using gRPC the way it was designed to be used.
Dropbox — Courier#
Dropbox described Courier, its gRPC-based RPC framework, built while migrating hundreds of services that exchange millions of requests per second from a legacy HTTP/1.1 RPC protocol with protobuf payloads. Courier integrates mTLS service identity, per-method ACLs and rate limits, mandatory deadlines for every method, a time-bounded LIFO queue to shed overload, per-client and per-method stats and tracing. Among its lessons: observability is a feature, and migration takes far longer than the development (Dropbox tech blog).
Staff insight: The value wasn't gRPC itself but the framework around it that made deadlines, identity and metrics mandatory. That's the Principal move: defaults, not guidelines.
Uber — Push Platform on gRPC Bidirectional Streaming#
Uber moved its real-time push platform (RAMEN) from Server-Sent Events over HTTP/1.1 to gRPC bidirectional streaming over QUIC/HTTP/3. Under SSE, acknowledgements came back via separate RPCs every 30 seconds, so a message's delivery state could be unknown for up to 30 seconds, a problem for driver offers valid for about that long. With acks on the same stream, Uber reported p95 connection latency improving by at least 45% and push success rates rising by at least 1–2%, and described handling head-of-line blocking from large messages and keeping an SSE fallback (Uber engineering).
Staff insight: Bidi streaming earned its complexity because acks were the requirement. Pick streaming for a concrete protocol need (acks, flow control), not because it sounds real-time.
Practice Drill#
Drill 1: The Scale-Out That Didn't Help#
Prompt: "Your checkout service calls pricing over gRPC through a Kubernetes ClusterIP Service. During a sale, pricing's p99 went from 40ms to 2s. The autoscaler doubled pricing from 12 to 24 pods, nothing improved, and checkout then ran out of worker threads and started failing. Retries are configured in the checkout code and in the mesh. What happened, and what do you change?"
Staff Answer
Three failures stacked. First, connection pinning: checkout pods hold long-lived HTTP/2 connections to pricing through ClusterIP, which balances per connection, so the 12 new pricing pods received almost no traffic while the original 12 stayed saturated. Per-pod request-rate skew would have shown it immediately. Second, no propagated deadlines: checkout called pricing with a fixed long timeout or none, so when pricing slowed, checkout's handlers piled up waiting until its worker pool was exhausted, turning a pricing slowdown into a checkout outage. Third, retry amplification: retries in the application and the mesh multiply attempts during exactly the brownout, adding load to pricing. Changes: per-request balancing, either by routing through the mesh's L7 balancing or by a headless Service with client-side round_robin, plus a server-side max connection age of ~5 minutes with a grace period so clients rebalance after any scale-out. Deadlines from the edge budget: checkout has 800ms, gives pricing ~250ms (well above its normal 40ms p99), propagates remaining time, and pricing handlers check cancellation so abandoned work stops. Retries in one place only: the gRPC client service config, UNAVAILABLE only, max 3 attempts, 50ms initial backoff, retry throttling on (maxTokens 10, tokenRatio 0.1); remove application and mesh retries. Degradation: if pricing times out, checkout uses the cached quote for items whose price hasn't changed in the last minute, or fails fast with a clear error, rather than holding threads. Signals: per-pod RPS skew alert at max/median > 2, attempts per logical call, in-flight RPCs in checkout, DEADLINE_EXCEEDED by callee. Then a game day: add 2s of latency to pricing in staging and verify checkout sheds cleanly and new pricing pods take traffic within minutes.
Why this is L6:
- Identifies L4 pinning as the reason the scale-out did nothing, and fixes both balancing and rebalancing.
- Ties checkout's thread exhaustion to missing deadline propagation and cancellation.
- Collapses retries to one throttled layer and removes amplification.
- Adds a degradation path and a test that proves the fix.
What L7 adds:
- Makes per-request balancing, deadlines and single-layer throttled retries framework defaults, so the next service doesn't rediscover this.
- Bans L4-only gRPC in the paved road and adds per-pod skew to the platform's SLOs.
- Prices the mesh vs proxyless choice for the hot checkout path.
❌ Common L5 Trap
"Pricing needs more capacity, so raise the autoscaler's maximum to 50 pods, increase checkout's timeout to 10 seconds so requests don't fail, and add one more retry so transient errors recover."
Why this misses: More pods don't receive traffic while connections are pinned at L4, so capacity isn't the bottleneck. A 10-second timeout makes checkout hold threads longer and exhausts its pool faster. Another retry adds load to the struggling service at the worst moment. The fix is per-request balancing with connection recycling, propagated deadlines with cancellation, and retries in exactly one throttled layer.
Quick Reference Card#
Model: contract-first RPC over HTTP/2; protobuf on the wire; reliability in the client
Call types: unary, server streaming, client streaming, bidirectional
Wire format: fields identified by NUMBER; add fields freely; never reuse numbers or
change types; delete with `reserved`; enum zero = UNSPECIFIED
Deadlines: none by default; propagate remaining budget (converted to timeout);
DEADLINE_EXCEEDED on client, CANCELLED on server -> stop work
Retryable: UNAVAILABLE (canonical); never INVALID_ARGUMENT / NOT_FOUND / PERMISSION_DENIED
Retry config: service config retryPolicy; maxAttempts capped at 5; committed once
response headers arrive; retryThrottling pauses below half of maxTokens
Hedging: hedgingPolicy (maxAttempts <= 5, hedgingDelay); idempotent reads only
Load balancing: L4 pins HTTP/2 connections -> per-request LB (L7 proxy, mesh,
client round_robin, proxyless xDS) + server max connection age
Keepalive: client off by default; timeout 20s; server 2h; permit 5 min
(faster pings -> GOAWAY too_many_pings)
Messages: 4 MB default receive limit in main implementations; paginate or stream
gRPC-Web: needs a proxy (Envoy default); unary + server streaming only
Edge: REST/JSON or GraphQL for external callers; gRPC inside
RED FLAGS
- Calls without deadlines
- gRPC behind ClusterIP / L4 only, no connection age limit
- Retries in app + client + mesh
- Retrying non-idempotent methods
- Reused field numbers or changed field types
- Streams with no resume token
- Native gRPC promised to browsers