People compare these three as if they were rival wire formats. They are not. REST is a set of conventions that lets the web's infrastructure — caches, CDNs, proxies, browsers — understand your API. gRPC is a contract-first RPC system that makes service-to-service calls typed, compact and deadline-aware. GraphQL is a query layer that moves the shape of the response to the client, so many screens can share one backend without a new endpoint each. Most mature systems run two of them at once. The question that decides it is: who is on the other end of the wire, and who controls their release cycle? If it is a stranger's code or a browser, you want HTTP semantics. If it is your own service deployed by your own CI, you want a compiled contract. If it is a fleet of your own clients whose screens change weekly, you want the client to choose its fields.
Go deeper:
- Deadlines, retry and hedging config, and why L4 balancing pins gRPC connections are covered in gRPC.
The Verdict#
Default: REST (JSON over HTTP) at the public edge, gRPC between internal services, and GraphQL only when many first-party client screens aggregate data from many backends.
| Pick REST when | Pick gRPC when | Pick GraphQL when |
|---|---|---|
The API is public or partner-facing and must be callable with curl | Both ends are services you own and deploy | You own several client apps (iOS, Android, web) with fast-changing screens |
| Responses are cacheable and a CDN or browser cache can absorb 80–99% of reads | Call volume is high (10K+ RPS per service) and payload/CPU cost matters | A single screen needs data from 5+ backends and round trips dominate latency |
| You need broad tooling: API gateways, WAFs, logs, HTTP debuggers | You need streaming (server, client or bidirectional) or strict deadlines | Over-fetching is measurable: clients discard most of each response |
| Resource semantics (create, read, update, delete) map cleanly | Polyglot services need generated, type-safe clients | A platform team can own the schema, resolvers and query cost controls |
🎯 Staff Move: "I'll keep the public API as versioned REST because partners integrate with curl and our CDN caches the catalog reads. Inside, services talk gRPC with deadlines. I'd only add GraphQL if we end up with three client teams each building their own aggregation endpoints."
When the non-default wins:
- gRPC at the public edge wins when your "public" clients are a small set of sophisticated integrators (trading firms, IoT fleets you ship firmware to) who value streaming and compact payloads over curl-ability.
- REST internally wins when the org has fewer than ~10 services, low RPS, and no platform team to own a mesh. JSON over HTTP with a shared client library is cheaper to run than a gRPC stack nobody owns.
- GraphQL for a single web app rarely wins. One client and one backend team gain little that a well-designed REST endpoint does not already give them.
At a Glance#
| Dimension | REST (JSON/HTTP) | gRPC (Protobuf/HTTP/2) | GraphQL (JSON over HTTP) |
|---|---|---|---|
| Data model | Resources addressed by URL; verbs from HTTP | Services and methods defined in .proto; typed messages | One typed graph schema; client selects fields per query |
| Contract | Optional (OpenAPI); often drifts from reality | Mandatory .proto; code generated for client and server | Mandatory schema; introspection for tooling |
| Wire format | Text JSON, typically 1.5–5× larger than protobuf for the same data | Binary protobuf; field numbers instead of names | JSON; response shaped exactly like the query |
| Transport | HTTP/1.1, HTTP/2 or HTTP/3 | HTTP/2 required; long-lived multiplexed connections | Usually HTTP POST to one endpoint; subscriptions over WebSocket or SSE |
| Caching | Native: Cache-Control, ETag, CDN keys on URL | None built in; cache inside the service | Hard at the HTTP layer (POST, one URL); needs persisted queries or client-side normalized caches |
| Streaming | Not native; use SSE, WebSocket or long polling | Native server, client and bidirectional streams | Subscriptions (separate transport); incremental delivery is emerging |
| Browser support | Universal | Not directly; needs gRPC-Web or a compatible protocol plus a proxy | Universal |
| Latency overhead | Serialization ~10–100µs for small payloads; one round trip per resource | Serialization ~1–20µs; connection reuse avoids handshakes | Resolver fan-out adds one backend hop per nested level unless batched |
| Throughput (per core, small messages) | ~5–20K RPS typical framework | ~20–50K RPS typical framework | ~1–5K queries/s; cost depends on query shape, not request count |
| Errors | HTTP status codes + body | Status codes (DEADLINE_EXCEEDED, UNAVAILABLE...) + details | HTTP 200 with an errors array; partial success is normal |
| Versioning | URL or header versions; additive changes | Field numbers make additive changes safe; never reuse a number | Deprecate fields; schema evolves without versions |
| Scaling model | Stateless servers behind any L4/L7 load balancer | Needs L7 or client-side balancing; L4 pins a connection to one backend | Stateless gateway; scaling is in the resolvers and backends |
| Operational burden | Lowest; every tool understands it | Medium: proxies, LB, debugging tooling, schema registry | Highest: query cost limits, N+1 control, schema governance, per-field observability |
| Managed options | Every API gateway | Most service meshes and cloud load balancers | Managed GraphQL gateways from cloud vendors; or self-host |
| Cost shape | Scales with requests; CDN offload makes reads near-free | Scales with RPS; ~2–5× less CPU and bandwidth than JSON at high volume | Scales with resolver calls per query; one query can cost 1 or 10,000 backend calls |
The throughput and serialization figures are orders of magnitude for small (under ~1KB) messages on a modern core; they move a lot with payload size and framework. The ratios hold better than the absolutes.
How They Actually Differ#
1. Who Owns the Shape of the Response#
This is the deepest difference, and it is organizational, not technical.
| Protocol | Who decides the response shape | What changes when a screen changes |
|---|---|---|
| REST | The server team, per endpoint | Server adds fields or a new endpoint; client waits for a deploy |
| gRPC | Whoever owns the .proto, usually the server team | Same as REST, but a compile error tells everyone |
| GraphQL | The client, per query, within the schema | Client edits its query; no server deploy if the fields exist |
REST and gRPC put the server team on the critical path of every UI change. That is fine when the clients are few and stable, or when the consumer is another service. It becomes a bottleneck when 3 client teams ship weekly and each needs a slightly different view of the same entities. The usual REST workaround is a backend-for-frontend (BFF) per client — which works, and is often the right answer for 2 clients. GraphQL is what you reach for when the BFF count heads toward 5+ and they are all reimplementing the same joins.
🎯 Staff Insight: GraphQL does not remove the aggregation work. It moves it into resolvers owned by a platform team. Adopt it when you are ready to staff that team.
2. Caching: The Web's Free Infrastructure#
REST inherits 30 years of HTTP caching. A GET /products/123 with Cache-Control: max-age=60 and an ETag is cacheable by the browser, the CDN and any reverse proxy, with conditional revalidation (304 Not Modified) costing a few hundred bytes. For a read-heavy public API, a CDN hit ratio of 90%+ means the origin sees one request in ten.
gRPC has no HTTP-level caching story: every call is a POST-like stream over HTTP/2. Caching lives in the service (a local LRU or Redis), which is fine because internal callers are not geographically spread.
GraphQL is the hard case. Queries are usually POSTed to /graphql, so CDNs see one URL with varying bodies. The fixes are persisted queries (the client sends a hash, the server maps it to a known query, and the request can become a cacheable GET) and normalized client caches keyed by type and ID. Both work; both are extra machinery someone must own.
The diagram is the shape most large systems converge on: HTTP semantics where the internet is involved, compiled contracts where only your services are.
3. The Contract and How It Evolves#
| Change | REST/JSON | gRPC/Protobuf | GraphQL |
|---|---|---|---|
| Add a field | Safe if clients ignore unknown fields | Safe; old clients skip unknown field numbers | Safe; clients only see fields they ask for |
| Rename a field | Breaking | Safe on the wire (numbers, not names), breaking in code | Add new, deprecate old, track usage, remove |
| Remove a field | Breaking for anyone reading it | Mark reserved; never reuse the number | Remove after per-field usage drops to zero |
| Know who uses a field | Log analysis, guesswork | Hard; binary payloads | Easy; every query names its fields |
GraphQL's underrated advantage is field-level usage telemetry: you can see that user.legacyAvatarUrl was requested 4 times last week and by which client version. Protobuf's advantage is that the compiler catches drift before deploy. REST's JSON has neither by default; teams bolt on OpenAPI and contract tests.
4. Connections, Load Balancing and Streaming#
gRPC runs over HTTP/2, multiplexing thousands of concurrent calls on one TCP connection. That saves handshakes (a TLS handshake costs 1–2 round trips; on a 50ms cross-region link that is 50–100ms per new connection) but breaks L4 load balancing: an L4 balancer spreads connections, and if a client holds one connection for hours, every request goes to one backend. You need per-request (L7) balancing in a proxy or sidecar, client-side balancing with a resolver, or a maximum connection age (commonly 5–30 minutes) that forces reconnection.
| gRPC balancing option | How it works | Cost | Pick it when |
|---|---|---|---|
| L7 proxy / sidecar | Proxy terminates HTTP/2 and balances each call | ~0.2–1ms and some CPU per hop | You already run a mesh |
| Client-side balancing | Client resolves all endpoints and round-robins calls | Smart clients in every language | Few languages, latency-critical paths |
| Lookaside balancer | Client asks a balancer service which backend to use | Another service to run | Very large fleets with load-aware routing |
| Max connection age | Server closes connections after N minutes | Brief reconnect churn | Always, as a backstop to the above |
Streaming is native in gRPC: a server stream is the natural shape for "send me price updates" between services. REST needs SSE or WebSocket for that. GraphQL subscriptions usually ride WebSocket and bring their own operational concerns (sticky connections, fan-out, reconnect storms) — see Live Updates.
5. Where the Cost Hides#
REST cost is visible: one request, one handler, one bill. gRPC cost is visible too. GraphQL cost is a function of the query, and a client can write a query that fans out to 10,000 backend calls. Without per-request batching (a loader that collapses getItems(id) calls into batchGetItems(ids) within one tick) every list field becomes an N+1. Without query cost analysis — depth limits (commonly 7–10), a complexity budget per query, and timeouts — a single malicious or careless query can take down shared backends. Public GraphQL APIs typically rate-limit by computed query cost, not by request count.
Where Each One Breaks#
| Protocol | Failure mode | What it looks like in production | Mitigation | Owner |
|---|---|---|---|---|
| REST | Chatty clients | A mobile screen makes 8 sequential calls; on a 300ms cellular RTT that is 2.4s before render | Composite endpoints, a BFF, or embed/expand parameters | Client + API team |
| REST | Contract drift | OpenAPI says price is a number; one handler returns a string; a partner's parser breaks | Generate the spec from code or code from the spec; contract tests in CI | API team |
| REST | Version sprawl | /v1, /v2, /v3 all live; each fix ships three times | Additive changes by default; date-based versions with a sunset policy | API platform |
| gRPC | Load pinned to one backend | One pod at 95% CPU, its siblings at 10%; p99 doubles after a deploy | L7/client-side balancing, max connection age | Platform / mesh |
| gRPC | Missing deadlines | A slow dependency holds threads for 30s; callers queue; cascade | Deadline on every call, propagated down the call graph | Every service team |
| gRPC | Opaque debugging | Binary payloads; on-call cannot read traffic with standard tools | Reflection in non-prod, structured logging, protocol-aware proxies | Platform |
| gRPC | Browser blocked | Frontend cannot call services directly | gRPC-Web or a compatible protocol through a proxy, or a REST/GraphQL edge | Edge team |
| GraphQL | N+1 resolver storms | One query issues 2,000 backend calls; the order service's p99 triples | Per-request batching loaders; resolver-level tracing | GraphQL platform |
| GraphQL | Expensive queries | Deeply nested query pins gateway CPU; shared backends saturate | Depth and cost limits, persisted-query allowlist for first-party clients | GraphQL platform |
| GraphQL | Partial failure confusion | HTTP 200 with half the fields null; dashboards show 0% error rate | Alert on the errors array per field; nullable design reviewed | Platform + clients |
| GraphQL | Schema as a monolith | 40 teams edit one schema; reviews block shipping | Federation with clear type ownership; schema linting and checks in CI | Principal-level decision |
The production surprise for each:
- REST: the bottleneck is usually not the API, it is the round trips. A 6-call screen on mobile is a latency bug no server optimization fixes.
- gRPC: teams adopt it for speed and get burned by load balancing. The first incident is almost always "one pod is hot".
- GraphQL: dashboards lie. Request count is flat while backend load triples because a client shipped a heavier query.
Incident Sketch: The Hot Pod After a Deploy#
t=0 Deploy rolls the pricing service from 12 to 12 new pods, one at a time
t=+4min Callers' HTTP/2 connections reconnected to the first 3 pods that came up
t=+5min Pod 1: 92% CPU, p99 340ms. Pods 4–12: 8% CPU
t=+6min Upstream deadlines (100ms) start firing; checkout error rate 3%
t=+9min On-call scales to 20 pods. Nothing improves: new pods get no connections
t=+14min Rolling restart of callers spreads connections; error rate back to 0.1%
Detection: per-pod CPU skew (max/median > 3), grpc_server_handled_total by pod, client-side DEADLINE_EXCEEDED rate. Root cause: L4 balancing of long-lived connections. Prevention: client-side round-robin over resolved endpoints or an L7 sidecar, plus a max connection age of ~10 minutes so connections rebalance after every deploy. Owner: the platform team that owns the mesh defaults — not each service team.
Incident Sketch: The Query That Shipped in a Mobile Release#
t=0 App release adds 'recent orders with items and reviews' to the home screen
t=+2h Adoption reaches 20% of DAU. GraphQL request rate: flat
t=+3h Review service p99 from 40ms to 900ms; it now gets 25 calls per home-screen load
t=+3h10m Review service sheds load; home screen renders with null review fields, HTTP 200
t=+4h A product manager notices empty review stars. No alert fired
Detection that would have caught it: resolver-level call counts per operation name, backend RPS split by calling GraphQL operation, and an alert on the per-field errors rate rather than HTTP 5xx. Prevention: a loader batching per request, a query cost budget enforced in CI against persisted queries, and a load test for every new persisted query above a cost threshold.
Cost and Operations#
| REST | gRPC | GraphQL | |
|---|---|---|---|
| Who runs it | Each service team; an API gateway team for the public edge | Service teams plus a platform team for mesh, LB, schema registry | A dedicated GraphQL platform team (commonly 3–8 engineers once 10+ teams contribute) |
| What the bill scales with | Origin requests after CDN offload; egress bytes | RPS × CPU per call; ~2–5× cheaper than JSON per call at volume | Resolver calls per query; gateway CPU for parsing, validation and planning |
| Hidden cost | Duplicate BFFs; version maintenance | Proxy/sidecar CPU and memory; tooling for humans | Query cost governance, schema review, client cache complexity |
| When the cheap option wins | Read-heavy public data behind a CDN | High-volume internal traffic (100K+ RPS fleet-wide) | Many client teams; the alternative is N hand-written BFFs |
Rough numbers for a fleet doing 200K internal RPS: moving from JSON/HTTP to protobuf/gRPC commonly saves 30–60% of serialization CPU and 50–70% of bytes on the wire for structured payloads. That is real money at 200K RPS and irrelevant at 200 RPS. At low volume, pick for developer experience, not efficiency.
| Scale | Sensible stack | Rough people cost | What dominates the bill |
|---|---|---|---|
| Startup (5 services, 1 web app, ~500 RPS) | REST everywhere, one shared client library | 0 dedicated engineers | Engineer time, not compute |
| Growth (50 services, 3 client apps, ~20K RPS) | REST edge, gRPC internally, BFFs per client | 1–2 platform engineers for RPC tooling | Proxy/sidecar overhead, duplicated BFF logic |
| Large (500+ services, 5+ client apps, 500K+ RPS) | REST public, gRPC internal, federated GraphQL for first-party clients | 4–10 engineers across RPC platform and GraphQL platform | Serialization CPU, resolver fan-out, schema governance |
Assumptions: cloud compute list prices, small structured payloads, people cost dominating until roughly 50K RPS.
🧭 Principal Insight: The cost of GraphQL is not the gateway's CPU. It is the permanent team that owns the schema and the governance process 30 product teams must follow. Price that headcount before you adopt it.
Switching Later#
| Migration | Difficulty | What's hard to undo |
|---|---|---|
| REST → gRPC internally | Moderate; service by service | Running both stacks during the migration (commonly 6–24 months for hundreds of services); client libraries in every language |
| gRPC → REST edge for external clients | Easy; transcoding proxies map HTTP/JSON to gRPC methods | Little; the .proto stays the source of truth |
| REST → GraphQL for first-party clients | Moderate; put GraphQL in front of existing REST/gRPC services | Client caches and query patterns spread across many app versions you cannot force-upgrade |
| GraphQL → REST | Hard | Every shipped mobile app version is a client of the schema; old queries must keep working for years |
| Public REST v1 → anything | Very hard | Partners integrated once and will not return; breaking changes cost relationships |
The one-way doors:
- A public API contract. Whatever partners integrate against, you support for years. Start with REST and additive-only changes.
- GraphQL in shipped mobile apps. Old app versions keep sending old queries. Persisted queries help: you know exactly which queries exist and can keep them working.
- Protobuf field numbers. Reusing a number corrupts data for any old binary still running. Mark removed fields
reservedforever.
The safe migration pattern for internal gRPC adoption is the one large companies report: freeze new features on the legacy RPC stack, define the shared interface layer first (auth, discovery, metrics, deadlines), then migrate clients incrementally. The migration takes longer than building the new framework.
Adopting GraphQL without a big bang:
- Stand up the gateway in front of existing REST and gRPC services; resolvers call them, nothing below changes.
- Move one high-traffic screen; measure round trips and payload bytes before and after.
- Make persisted queries mandatory for first-party apps from day one — retrofitting them later means supporting arbitrary queries from every old app version.
- Assign type ownership before the second team contributes.
How Real Companies Chose#
Facebook — GraphQL for Mobile News Feed#
Facebook started GraphQL in 2012 while rebuilding its native iOS and Android apps. News Feed had only been delivered as HTML, and the REST and FQL approaches made developers write a lot of code on the server to prepare data and on the client to parse it, with a mismatch between what the app needed and what the queries returned. By 2015 GraphQL served millions of requests per second from nearly 1,000 shipped app versions (Facebook Engineering).
Staff insight: GraphQL was born for many app versions with fast-changing screens. "Nearly 1,000 shipped versions" is the condition under which it wins — not a single web app talking to a single backend.
GitHub — Adding a GraphQL API Next to REST#
GitHub launched its GraphQL API because its REST API generated over 60% of the requests to its database tier, responses carried too much data and still missed what consumers needed, and integrators often needed two or three calls to assemble one view. GraphQL also gave them a typed schema for documentation, scopes and client generation. The REST API stayed (GitHub Blog).
Staff insight: Even the company that went all-in on GraphQL kept REST for its public integrators. You add GraphQL for the clients it helps; you do not rip out the API your partners depend on.
Dropbox — Migrating Internal RPC to gRPC#
Dropbox replaced a legacy HTTP/1.1 RPC framework with protobuf encoding by Courier, a gRPC-based layer adding Dropbox's auth, service discovery, stats, logging and tracing, across hundreds of services and millions of requests per second. They chose gRPC because it kept their existing protobuf investment and added HTTP/2 multiplexing and bidirectional streaming. The biggest reliability win: requiring deadlines in service definitions removed whole classes of problems. They also note the migration took far longer than development (Dropbox Tech).
Staff insight: The value was not the binary format. It was a single framework where deadlines, observability and auth are mandatory. Say that, and you sound like someone who has run the migration.
Follow-Ups to Expect#
| After You Say... | They Will Ask... | What They're Testing |
|---|---|---|
| "gRPC internally" | "How do you load-balance long-lived HTTP/2 connections?" | Whether you know L4 pins connections and how to fix it |
| "gRPC internally" | "How does the web frontend call these services?" | Browser limits; edge translation layer |
| "REST for the public API" | "How do you evolve it without breaking partners?" | Additive changes, versioning policy, deprecation and sunset |
| "GraphQL for the mobile apps" | "What stops one query from taking down the backends?" | Depth/cost limits, persisted queries, timeouts |
| "GraphQL for the mobile apps" | "How do you cache it?" | Persisted queries as GETs, normalized client caches, per-field cache hints |
| "GraphQL" | "Who owns the schema when 30 teams contribute?" | Federation, type ownership, schema review process |
| "Protobuf is faster" | "How much faster, and does it matter here?" | Quantifying: 2–5× CPU and bytes, relevant only at high RPS |
| "We'll use streaming" | "What happens when a stream breaks mid-way?" | Resume tokens, idempotent replay, backpressure |
| "Deadlines on every call" | "How do deadlines propagate through 4 hops?" | Remaining-budget propagation, not a fixed timeout per hop |
| "Persisted queries" | "How do you ship a new query to an app already in the store?" | Build-time query registration, allowlists, old versions staying valid |
| "BFF per client" | "When does the BFF approach stop scaling?" | Recognizing duplicated aggregation logic as the GraphQL trigger |
| "HTTP 200 with errors" | "How does your alerting see GraphQL failures?" | Per-field error metrics instead of status-code dashboards |
What to Say in the Interview#
"The deciding question is who sits on the other end of the wire. Partners and browsers get REST, because they need HTTP caching and curl-level tooling. Our own services get gRPC with a deadline on every call."
"I'd only introduce GraphQL if we have several client teams building their own aggregation endpoints — and then I'd staff a platform team to own the schema, batching and query cost limits."
"With gRPC I'd plan load balancing up front: client-side or L7 balancing plus a maximum connection age, because an L4 balancer will pin each long-lived connection to one pod."
"Protobuf saves roughly 2 to 5 times the CPU and bytes of JSON, which matters at a few hundred thousand requests per second and not at all at a few hundred."
Related Guides#
- API Contracts — idempotency, pagination, versioning and error contracts for all three styles
- Latency and Protocols — HTTP/2, HTTP/3 and the gRPC load-balancing problem in depth
- Envoy, Kong and NGINX — the proxies that do L7 balancing and protocol translation
- Design an Edge Gateway — where the REST/GraphQL edge lives
- Design Live Updates — streaming and subscriptions at scale
- Push vs Poll — choosing between streams, SSE and polling
- Design a News Feed — the classic aggregation-heavy client screen