These three get compared because each one is sold as "the relational database that scales", and each one means something different by it. Aurora scales storage and reads behind a single writer: it is still one primary, just one with a distributed, self-healing disk. Spanner and CockroachDB scale writes by splitting the keyspace into ranges, each replicated by its own consensus group, and running distributed transactions across them. The question that decides it is: will one writer instance carry your write load and your availability target for the next three years? If yes, Aurora is cheaper, faster per query and far more compatible. If no — because writes outgrow the biggest instance, or because the business needs writes to survive a region loss with zero data loss — you are choosing between Spanner and CockroachDB, and that choice is mostly about cloud and operating model.
The Verdict#
Default: Aurora (or plain managed Postgres) until a single writer is a measured problem. Then Spanner if you are on Google Cloud and want it fully managed, CockroachDB if you need to run on several clouds or on your own hardware.
| Pick Aurora when | Pick Spanner when | Pick CockroachDB when |
|---|---|---|
| Peak writes fit one large instance (roughly tens of thousands of simple writes/s) | Writes outgrow one machine and you are on Google Cloud | Writes outgrow one machine and you must stay cloud-portable or self-host |
| You need full PostgreSQL or MySQL compatibility: extensions, tooling, ORMs | You want 99.999% multi-region availability as a managed product | You need row-level data domiciling (EU rows stay in EU) inside one logical database |
| Single-region writes with cross-region read replicas and DR are enough | You need external consistency (strict serializability) across regions | You want PostgreSQL wire compatibility with horizontal writes |
| Latency per query matters: 1–5ms commits, no consensus round per write | Your team will design keys around interleaving and avoid hot ranges | You have a platform team, or budget for the vendor's managed service |
🎯 Staff Move: "I'll start on Aurora Postgres: 8K writes per second at peak fits one writer with headroom, and a single primary gives 2ms commits. I'll keep keys and transactions shard-friendly — UUID keys, no cross-tenant transactions — so moving to distributed SQL later is a data migration, not a rewrite."
When the non-default wins:
- Spanner or CockroachDB on day one wins when the requirement is regulatory or contractual from the start — RPO 0 across regions for a ledger, or rows legally pinned to a jurisdiction — because retrofitting those onto a single-writer design is a rewrite.
- Aurora past "one writer" can still win if the write-heavy part is a small, separable slice (events, audit logs, counters) that can move to a log or a key-value store, leaving the relational core on one primary.
- CockroachDB on AWS beats Aurora when the team needs horizontal writes but not Google Cloud, and accepts a different engine with retry semantics.
At a Glance#
| Dimension | Aurora (PostgreSQL / MySQL) | Spanner | CockroachDB |
|---|---|---|---|
| Architecture | One writer + up to 15 readers sharing a distributed storage volume | Range-sharded (splits), Paxos per split, TrueTime | Range-sharded (ranges), Raft per range, hybrid logical clocks |
| Data model / SQL | Full PostgreSQL or MySQL engine | GoogleSQL or PostgreSQL dialect; interleaved tables | PostgreSQL wire protocol and most of its SQL; not the Postgres engine |
| Consistency | Writer is strongly consistent; readers lag, usually well under 100ms | External consistency (strict serializability) | Serializable by default; strong reads from leaseholders |
| Transactions | Local ACID on one writer; no distributed commit | Distributed 2PC across splits, automatic | Distributed transactions with parallel commits, automatic |
| Ordering / clocks | Single writer orders everything | TrueTime with ~1–7ms uncertainty; commit wait | HLC with configured max offset (500ms default); read restarts |
| Write throughput | One instance: ~10K–100K simple writes/s depending on size and row shape | Scales roughly linearly with nodes; each node commonly thousands of writes/s | Scales with nodes; each vCPU commonly hundreds to low thousands of writes/s |
| Commit latency (in-region) | ~1–5ms | ~5–15ms (consensus round + commit wait overlap) | ~5–15ms (consensus round) |
| Multi-region writes | One primary region; up to 10 read-only secondary regions, replication typically under 1s | Multi-region configs with quorum across regions; 99.999% SLA | Survival goals (zone or region) and per-row or per-table locality |
| Scaling model | Scale up the writer; add readers; storage grows automatically to 128 or 256 TiB depending on version (as of 2026) | Add nodes or processing units; 10 TiB storage per node (as of 2026) | Add nodes; ranges split and rebalance automatically |
| Failover | Promote a reader: typically tens of seconds | Automatic per split; no instance-level failover event | Automatic per range; lease moves in seconds |
| Operational burden | Low: managed by the cloud provider | Lowest: fully managed, no versions to pick | Medium on the managed service; high self-hosted |
| Managed options | AWS only | Google Cloud only | Vendor-managed cloud service on AWS, GCP, Azure; or self-host anywhere |
| Cost shape | Instance-hours + storage + I/O (or I/O-optimized pricing) | Node or processing-unit hours (edition and topology dependent) + storage | Vendor service by usage or vCPU; self-hosted pays nodes + license above the free threshold |
Throughput figures are orders of magnitude for small rows and simple transactions; contention, secondary indexes and row size move them by 3–10×.
Numbers to bring:
| Figure | Value | Condition |
|---|---|---|
| Aurora replica lag | Usually well under 100ms | Same region; rises with write rate |
| Aurora Global Database secondaries | Up to 10 read-only regions, replication typically under 1s | As of 2026 |
| Aurora cluster volume | 128 TiB, or 256 TiB on recent engine versions | As of 2026 |
| Spanner commit limits | 80,000 mutations, 100 MiB per commit | As of 2026 |
| Spanner storage per node | 10 TiB | As of 2026 |
| CockroachDB default max clock offset | 500ms | Lower to 250ms for well-synced multi-region |
| In-region consensus commit | ~5–10ms p50 | 3 replicas across zones |
| Cross-region quorum commit | +20–70ms | US-scale distances |
How They Actually Differ#
1. Shared Storage vs Sharded Consensus#
Aurora's insight is that the bottleneck in a cloud database is the network, so the engine ships only redo log records to a storage fleet that keeps 6 copies across 3 availability zones, with writes acknowledged by 4 of 6 and reads needing 3 of 6. Readers attach to the same volume, which is why replica lag is milliseconds rather than seconds. But there is still exactly one writer. Aurora makes the single-primary model durable, fast to recover and cheap to read-scale. It does not make writes horizontal.
Spanner and CockroachDB cut the keyspace into ranges (Spanner: splits) of a few hundred MiB, give each range its own consensus group, and move ranges between nodes as load shifts. Writes scale with nodes because different ranges have different leaders. The price: every write pays a consensus round, and any transaction that touches two ranges pays a distributed commit.
🎯 Staff Insight: "Aurora scales" and "Spanner scales" describe different axes. Aurora scales reads and storage. Spanner and CockroachDB scale writes. Say which axis your workload needs before naming a product.
2. Who Pays for Clock Uncertainty#
| Spanner | CockroachDB | Aurora | |
|---|---|---|---|
| Clock source | TrueTime: GPS and atomic-clock-backed interval | NTP-class clocks + hybrid logical clock | Irrelevant: one writer orders all commits |
| Uncertainty bound | ~1–7ms (reported in the 2012 paper) | Configured max offset, 500ms default (250ms suggested for well-synced multi-region) | N/A |
| Who pays | Every read-write transaction waits out the uncertainty before acknowledging (overlaps with replication) | Reads that land on recently written keys may restart | Nobody |
| Guarantee | External consistency | Serializable; strict serializability for single-key operations, not across arbitrary keys | Serializable available on one node |
This is the cleanest Staff-level distinction between the two distributed systems. Spanner buys a global commit order with special hardware and a few milliseconds of wait. CockroachDB runs on any hardware and pushes the cost to occasional read restarts. In practice both surface as retryable transaction errors your client must handle; Cockroach's 40001 errors appear more often under contention and with long-running transactions.
3. Multi-Region Write Semantics#
- Aurora Global Database: one primary region, up to 10 read-only secondary regions as of 2026. Replication is asynchronous at the storage layer, typically under a second. A planned switchover loses nothing; an unplanned regional failover can lose the last seconds of writes. That is the honest RPO.
- Spanner multi-region: writes go to a leader region and need a quorum across regions. RPO is zero for region loss; every write pays a cross-region round trip.
- CockroachDB: you choose per database whether to survive a zone (fast writes, region outage makes that region's rows unavailable) or a region (cross-region quorum on every write), and per table or row where data lives (
REGIONAL BY ROW,GLOBAL). It is the most flexible locality model of the three, and the easiest to misconfigure.
4. Compatibility and Lock-In#
Aurora runs the real PostgreSQL and MySQL engines on new storage. Extensions, the query planner, EXPLAIN output and ORMs behave as on community Postgres or MySQL, with documented exceptions. Leaving Aurora for RDS or self-hosted Postgres is a dump-and-restore or logical replication project, not a rewrite.
CockroachDB speaks the PostgreSQL wire protocol and much of its SQL, but it is a different engine: different planner, different locking, no Postgres extensions, and behaviors like serializable-by-default and transaction retries that applications must handle. Spanner offers a PostgreSQL dialect, but schema design (interleaving, key choice, avoiding monotonically increasing keys) is Spanner-specific, and it exists only on Google Cloud.
Licensing matters for CockroachDB: since version 24.3 (November 2024) there is a single Enterprise edition, free for individuals and businesses under $10M annual revenue, paid above it. Self-hosting is no longer "free open source" for most companies. Check current terms before you count it as zero license cost.
5. Key Design Is Physical Design#
On Aurora a bad primary key is a slow index. On Spanner and CockroachDB it is a hot range: a monotonically increasing key (timestamp, sequence) sends every insert to the last range, so one node takes 100% of writes while the cluster of 30 idles. Both systems document the fix — UUIDs, hash-sharded indexes (CockroachDB), bit-reversed sequences or UUIDs (Spanner) — and both reward colocating parent and child rows so a business transaction touches one range. Plan the keys first; see Distributed SQL for the full key-design method.
Where Each One Breaks#
| System | Failure mode | Symptom | Detection | Mitigation | Owner |
|---|---|---|---|---|---|
| Aurora | Writer saturation | CPU pinned on the writer, commit p99 from 3ms to 200ms, readers idle | Writer CPU, commit latency, lock waits | Scale up (one-time 2×), move reads off, then shard or migrate | Platform + product |
| Aurora | Failover gap | 20–60s of write errors when the writer fails; connection storms on reconnect | Failover events; client error rate | Connection proxy, retry with jitter, DNS TTL awareness | Platform |
| Aurora | Unplanned regional failover loses tail writes | Last ~1s of acknowledged writes missing in the new primary | Global replication lag metric | Reconcile from an outbox or event log; accept and document RPO | Business sign-off |
| Aurora | Long transactions block cleanup | Reader queries holding old snapshots; storage and vacuum lag grow | Oldest transaction age, MaximumUsedTransactionIDs style metrics | Statement timeouts, separate analytics replica or warehouse | Product team |
| Spanner | Hot split from sequential keys | One split leader at 100%, write latency climbs, total throughput flat as you add nodes | Key Visualizer hot bands, per-split CPU | Re-key with UUIDs or bit-reversed sequence; this is a migration | Product team |
| Spanner | Commit size limits | Bulk jobs fail: 80,000 mutations or 100 MiB per commit (as of 2026) | Job error rate | Batch writes, partitioned DML for large updates | Product team |
| Spanner | Cross-region latency surprise | Multi-region config adds 20–70ms to every write; checkout p99 regresses | Commit latency by region | Regional config plus async DR if RPO 0 is not truly required | Business sign-off |
| CockroachDB | Retry storms under contention | 40001 errors spike; ORMs surface them as 500s | Restart and retry metrics | Bounded client retries with backoff; shorten transactions; avoid hot rows | Product team |
| CockroachDB | Clock skew | Node exceeding max offset shuts itself down; with bad NTP, several do | Clock offset metric per node | Reliable time sync (cloud time service); alert well under the max | Platform |
| CockroachDB | Misconfigured locality | "EU" rows actually have voters in the US; writes cross the Atlantic | Replica placement reports, write latency by region | Locality reviews in schema CI; survival goals chosen deliberately | Platform |
| CockroachDB | Self-hosted upgrade and rebalancing load | Range rebalancing after a node loss saturates disks; p99 doubles for hours | Rebalance rate, disk throughput | Rate-limit rebalancing; keep 30–40% headroom | Platform |
The production surprise for each:
- Aurora: the writer ceiling arrives as a cliff, not a slope. At 70% CPU it is fine; at 90% lock waits and checkpoint stalls compound and p99 multiplies. Plan the exit at 50%.
- Spanner: adding nodes does nothing for a hot key. Teams discover their key design only after the cluster is big.
- CockroachDB: the database is correct, and the application is not. Retry handling is the most common production defect in migrations from Postgres.
Incident Sketch: The Aurora Writer Cliff#
t=0 Marketing campaign starts; write rate climbs from 6K to 11K/s
t=+8min Writer CPU 88%; commit p99 from 4ms to 60ms; lock waits rising
t=+12min Connection pool exhausted on app side; request queueing; checkout p99 4s
t=+15min On-call scales the writer up one size: 6-10 minute instance swap via failover
t=+16min Failover: ~30s of write errors; reconnect storm doubles the pain
t=+26min New writer at 50% CPU; queue drains
Detection that should have fired earlier: writer CPU over 50% at normal peak for a week, commit p99 trend, and headroom projections. Lesson: scaling a single writer up is a one-time, disruptive move. Owner: platform for capacity, product for the plan to shed or split writes.
Incident Sketch: The Retry Storm After a Migration to CockroachDB#
t=0 Service migrated from Postgres; ORM treats SQLSTATE 40001 as fatal
t=+3d Flash sale: 2,000 transactions/s update the same inventory rows
t=+3d+1m Serialization conflicts spike; 12% of checkouts return HTTP 500
t=+3d+4m Clients retry blindly with no backoff; contention worsens
t=+3d+20m Hotfix: bounded retry with jittered backoff, inventory split into 16 sub-counters
Root cause: single-node habits (no retry loop, a hot row) on a serializable distributed database. Prevention: retry handling as a migration gate, contention tests in staging, and hot-row patterns reviewed before cutover.
Cost and Operations#
| Aurora | Spanner | CockroachDB | |
|---|---|---|---|
| Who runs it | Cloud provider runs storage and patching; your platform team picks instance sizes, parameters, upgrades | Cloud provider runs everything; your team owns schema, keys and capacity in nodes | Vendor's managed service, or your platform team (commonly 2–4 engineers for a serious self-hosted fleet) |
| What the bill scales with | Writer and reader instance size and count, storage GB, I/O requests (or flat I/O-optimized pricing) | Node or processing-unit hours (regional cheapest, multi-region several times higher), storage GB | vCPUs or request units on the managed service; nodes plus license when self-hosted |
| Minimum sensible footprint | 1 writer + 1 reader across 2 AZs | Can start under one node (processing units); production commonly 3+ nodes | 3 nodes across 3 zones; 5+ for region survival |
| Cost cliff | Moving from one writer to sharding or another database | Multi-region configurations; over-provisioned nodes for peak | Region survival (5 replicas instead of 3, ~1.7× replicas) |
Rough shape, as of 2026 list prices and order of magnitude only: a production Aurora cluster with a large writer and two readers commonly costs low thousands of dollars a month; a regional Spanner deployment of 3 nodes is in the same low-thousands range before storage, and multi-region configurations multiply that; a self-hosted CockroachDB cluster of 9 mid-size nodes costs similar infrastructure plus license and the engineers to run it. The real difference shows at scale: distributed SQL costs more per transaction than a single primary, and you pay it to remove the ceiling.
| Scale | Sensible choice | Rough monthly infrastructure | People |
|---|---|---|---|
| Startup (1K writes/s, one region) | Aurora or managed Postgres | Hundreds to low thousands of dollars | Part of one engineer's time |
| Growth (10K writes/s, DR region) | Aurora with Global Database, shard-friendly schema | Low to mid thousands | 1–2 platform engineers |
| Large (100K+ writes/s, RPO 0 or residency) | Spanner or CockroachDB, multi-region for the core only | Tens of thousands and up | Platform team plus migration effort measured in engineer-quarters |
Assumptions: list prices as of 2026, order of magnitude only; storage and I/O excluded; the people column dominates below the large tier.
🧭 Principal Insight: The expensive part of distributed SQL is not the nodes. It is the application changes: retries, key redesign, transaction scoping, and teams relearning a planner. Budget engineer-quarters, not just dollars per node.
Switching Later#
| Migration | Difficulty | What's hard to undo |
|---|---|---|
| Postgres/MySQL → Aurora | Easy: logical replication or the provider's migration path; same engine | Little; you can move back to RDS or self-hosted the same way |
| Aurora → CockroachDB | Moderate to hard: wire-compatible, but schema, keys, extensions and retry logic change | Application assumptions about single-node behavior; sequences and hot keys |
| Aurora → Spanner | Hard: new dialect choices, interleaving, key redesign, cloud move | Cloud commitment; Spanner-specific schema |
| Spanner → anything | Hard: no other place runs Spanner | The cloud and the schema |
| CockroachDB → Postgres | Hard if you rely on multi-region locality; moderate if single region | Locality-based designs have no single-node equivalent |
| Single region → multi-region | Easy to configure, hard to live with | Per-write latency becomes a product property; rolling back means moving data |
The one-way doors:
- Spanner is a cloud commitment. Choose it when you have already chosen Google Cloud, not as a way to choose it.
- Multi-region write configurations become product behavior: once checkout has absorbed 40ms of cross-region commit, nobody remembers how fast it used to be, but every feature depends on that durability.
- Key design on any range-sharded system. Re-keying a 5 TB table is a dual-write migration measured in quarters.
A migration off a single writer, in order:
- Add CDC from Aurora into a log (Kafka) so the new database can be filled and kept current without touching the application.
- Backfill, then run shadow reads: every read goes to both, results compared, differences logged and counted.
- Move writes table group by table group, starting with the one causing the writer pressure, behind a feature flag with a rollback path.
- Keep the old writer read-only for a soak period of weeks before decommissioning it.
The cheap two-way door is the one most teams skip: design for sharding while still on Aurora. Random keys, no cross-tenant transactions, idempotent writes, an outbox for side effects. Those choices cost almost nothing on one writer and turn a later migration from a rewrite into a data move.
How Real Companies Chose#
Uber — Spanner for the Fulfillment Platform#
Uber rebuilt its fulfillment platform, which had used Cassandra and Redis with best-effort consistency and a peer-to-peer coordination layer that had physical scaling limits. It evaluated evolving the NoSQL stack, an in-house MySQL-based store, and NewSQL options (CockroachDB, FoundationDB and Spanner), and chose Google Cloud Spanner in a North America multi-region configuration for external consistency, multi-row and multi-table transactions, and horizontal scaling by key range, after benchmarking on availability SLA, operational overhead, transactions, schema and shard management (Uber Engineering).
Staff insight: The trigger was multi-entity transactions that the application had been faking with sagas and best-effort logic. Distributed SQL earns its cost when it deletes application-level consistency code.
Figma — Staying on RDS Postgres and Sharding It#
Figma evaluated CockroachDB, TiDB, Spanner and Vitess and chose to horizontally shard its existing RDS Postgres instead. Its reasons: any alternative required a complex migration to ensure consistency, the team would have had to rebuild years of operational expertise, it had only months of runway at its growth rate, and it preferred known low-risk solutions it controlled. A custom solution could also support a much smaller feature set than a general database (Figma Blog).
Staff insight: Distributed SQL is not automatically the answer to "Postgres is full". Migration risk, team expertise and runway are first-class inputs, and saying so is a Staff signal.
Amazon — Why Aurora Separated Compute from Storage#
The Aurora paper argues that in the cloud the network, not compute or storage, is the binding constraint, and so the engine pushes only redo log records to a multi-tenant storage service. Data is kept as 6 copies across 3 availability zones with quorum writes and reads, giving fast crash recovery, replica failover without data loss, and self-healing storage (SIGMOD 2017 paper).
Staff insight: Aurora was designed to make the single-writer model durable and fast to recover, not to remove it. Knowing what a system was built for tells you exactly where it stops.
Follow-Ups to Expect#
| After You Say... | They Will Ask... | What They're Testing |
|---|---|---|
| "Aurora to start" | "What's your trigger to move off a single writer?" | A measured threshold (writer CPU, commit p99) and a plan before the cliff |
| "Aurora Global Database for DR" | "What's your RPO on an unplanned regional failover?" | Knowing replication is async: about a second of writes at risk |
| "Spanner" | "What does a write cost in a multi-region config, and who signed off?" | Cross-region quorum latency and the business decision behind it |
| "Spanner" | "Your insert rate doubled but throughput didn't. Why?" | Hot splits from sequential keys |
| "CockroachDB" | "How does your application handle transaction retries?" | 40001 handling, idempotent transaction bodies, no side effects inside |
| "CockroachDB" | "How do you keep EU data in the EU?" | REGIONAL BY ROW, survival goals, verifying replica placement |
| "Distributed SQL scales writes" | "What happens to a transaction touching 5 ranges in 3 regions?" | Distributed commit cost; colocating related rows |
| "Strong consistency everywhere" | "Which reads can be stale?" | Follower or stale reads for dashboards and reporting |
| "Region survival" | "Which tables actually need it?" | Scoping expensive guarantees to the data that needs them |
| "Spanner is fully managed" | "What do you still own?" | Schema, keys, capacity, and the cloud commitment |
| "Aurora failover is fast" | "What does the app see during those 30 seconds?" | Connection handling, retries, proxies |
What to Say in the Interview#
"Aurora scales reads and storage behind one writer; Spanner and CockroachDB scale writes by sharding the keyspace into consensus groups. The first question is whether one writer carries us for three years."
"I'd start on Aurora with UUID keys and short, single-tenant transactions, and set a move trigger at 50% writer CPU at peak, because the single-writer ceiling arrives as a cliff."
"If we need writes to survive a region loss with zero data loss, every write pays a cross-region round trip — 20 to 70 milliseconds — and I want the business to sign off on that, not discover it."
"Between Spanner and CockroachDB, the deciding factor is usually the cloud: Spanner if we're on Google Cloud and want it fully managed, CockroachDB if we need portability or row-level data residency."
Related Guides#
- Distributed SQL — internals of Spanner and CockroachDB: ranges, TrueTime, HLC, locality
- PostgreSQL — the single-primary baseline and when it runs out
- Choosing a Database — the broader decision across data models
- Design a Sharded Database — what you build if you shard Postgres yourself
- Multi-Region Architecture — survival goals, RPO and active-active tradeoffs
- Replication — synchronous vs asynchronous replication and quorum math
- Consistency, CAP and PACELC — the vocabulary behind strict serializability
- Design Payments — a workload where RPO 0 is a real requirement