Hiring BarSupport

Spanner vs CockroachDB vs Aurora

Comparison19 min read4 diagrams

These three get compared because each one is sold as "the relational database that scales", and each one means something different by it. Aurora scales storage and reads behind a single writer: it is still one primary, just one with a distributed, self-healing disk. Spanner and CockroachDB scale writes by splitting the keyspace into ranges, each replicated by its own consensus group, and running distributed transactions across them. The question that decides it is: will one writer instance carry your write load and your availability target for the next three years? If yes, Aurora is cheaper, faster per query and far more compatible. If no — because writes outgrow the biggest instance, or because the business needs writes to survive a region loss with zero data loss — you are choosing between Spanner and CockroachDB, and that choice is mostly about cloud and operating model.

The Verdict#

Default: Aurora (or plain managed Postgres) until a single writer is a measured problem. Then Spanner if you are on Google Cloud and want it fully managed, CockroachDB if you need to run on several clouds or on your own hardware.

Pick Aurora whenPick Spanner whenPick CockroachDB when
Peak writes fit one large instance (roughly tens of thousands of simple writes/s)Writes outgrow one machine and you are on Google CloudWrites outgrow one machine and you must stay cloud-portable or self-host
You need full PostgreSQL or MySQL compatibility: extensions, tooling, ORMsYou want 99.999% multi-region availability as a managed productYou need row-level data domiciling (EU rows stay in EU) inside one logical database
Single-region writes with cross-region read replicas and DR are enoughYou need external consistency (strict serializability) across regionsYou want PostgreSQL wire compatibility with horizontal writes
Latency per query matters: 1–5ms commits, no consensus round per writeYour team will design keys around interleaving and avoid hot rangesYou have a platform team, or budget for the vendor's managed service

🎯 Staff Move: "I'll start on Aurora Postgres: 8K writes per second at peak fits one writer with headroom, and a single primary gives 2ms commits. I'll keep keys and transactions shard-friendly — UUID keys, no cross-tenant transactions — so moving to distributed SQL later is a data migration, not a rewrite."

Diagram: The Verdict

When the non-default wins:

  • Spanner or CockroachDB on day one wins when the requirement is regulatory or contractual from the start — RPO 0 across regions for a ledger, or rows legally pinned to a jurisdiction — because retrofitting those onto a single-writer design is a rewrite.
  • Aurora past "one writer" can still win if the write-heavy part is a small, separable slice (events, audit logs, counters) that can move to a log or a key-value store, leaving the relational core on one primary.
  • CockroachDB on AWS beats Aurora when the team needs horizontal writes but not Google Cloud, and accepts a different engine with retry semantics.

At a Glance#

DimensionAurora (PostgreSQL / MySQL)SpannerCockroachDB
ArchitectureOne writer + up to 15 readers sharing a distributed storage volumeRange-sharded (splits), Paxos per split, TrueTimeRange-sharded (ranges), Raft per range, hybrid logical clocks
Data model / SQLFull PostgreSQL or MySQL engineGoogleSQL or PostgreSQL dialect; interleaved tablesPostgreSQL wire protocol and most of its SQL; not the Postgres engine
ConsistencyWriter is strongly consistent; readers lag, usually well under 100msExternal consistency (strict serializability)Serializable by default; strong reads from leaseholders
TransactionsLocal ACID on one writer; no distributed commitDistributed 2PC across splits, automaticDistributed transactions with parallel commits, automatic
Ordering / clocksSingle writer orders everythingTrueTime with ~1–7ms uncertainty; commit waitHLC with configured max offset (500ms default); read restarts
Write throughputOne instance: ~10K–100K simple writes/s depending on size and row shapeScales roughly linearly with nodes; each node commonly thousands of writes/sScales with nodes; each vCPU commonly hundreds to low thousands of writes/s
Commit latency (in-region)~1–5ms~5–15ms (consensus round + commit wait overlap)~5–15ms (consensus round)
Multi-region writesOne primary region; up to 10 read-only secondary regions, replication typically under 1sMulti-region configs with quorum across regions; 99.999% SLASurvival goals (zone or region) and per-row or per-table locality
Scaling modelScale up the writer; add readers; storage grows automatically to 128 or 256 TiB depending on version (as of 2026)Add nodes or processing units; 10 TiB storage per node (as of 2026)Add nodes; ranges split and rebalance automatically
FailoverPromote a reader: typically tens of secondsAutomatic per split; no instance-level failover eventAutomatic per range; lease moves in seconds
Operational burdenLow: managed by the cloud providerLowest: fully managed, no versions to pickMedium on the managed service; high self-hosted
Managed optionsAWS onlyGoogle Cloud onlyVendor-managed cloud service on AWS, GCP, Azure; or self-host anywhere
Cost shapeInstance-hours + storage + I/O (or I/O-optimized pricing)Node or processing-unit hours (edition and topology dependent) + storageVendor service by usage or vCPU; self-hosted pays nodes + license above the free threshold

Throughput figures are orders of magnitude for small rows and simple transactions; contention, secondary indexes and row size move them by 3–10×.

Numbers to bring:

FigureValueCondition
Aurora replica lagUsually well under 100msSame region; rises with write rate
Aurora Global Database secondariesUp to 10 read-only regions, replication typically under 1sAs of 2026
Aurora cluster volume128 TiB, or 256 TiB on recent engine versionsAs of 2026
Spanner commit limits80,000 mutations, 100 MiB per commitAs of 2026
Spanner storage per node10 TiBAs of 2026
CockroachDB default max clock offset500msLower to 250ms for well-synced multi-region
In-region consensus commit~5–10ms p503 replicas across zones
Cross-region quorum commit+20–70msUS-scale distances

How They Actually Differ#

1. Shared Storage vs Sharded Consensus#

Aurora's insight is that the bottleneck in a cloud database is the network, so the engine ships only redo log records to a storage fleet that keeps 6 copies across 3 availability zones, with writes acknowledged by 4 of 6 and reads needing 3 of 6. Readers attach to the same volume, which is why replica lag is milliseconds rather than seconds. But there is still exactly one writer. Aurora makes the single-primary model durable, fast to recover and cheap to read-scale. It does not make writes horizontal.

Spanner and CockroachDB cut the keyspace into ranges (Spanner: splits) of a few hundred MiB, give each range its own consensus group, and move ranges between nodes as load shifts. Writes scale with nodes because different ranges have different leaders. The price: every write pays a consensus round, and any transaction that touches two ranges pays a distributed commit.

Diagram: 1. Shared Storage vs Sharded Consensus

🎯 Staff Insight: "Aurora scales" and "Spanner scales" describe different axes. Aurora scales reads and storage. Spanner and CockroachDB scale writes. Say which axis your workload needs before naming a product.

2. Who Pays for Clock Uncertainty#

SpannerCockroachDBAurora
Clock sourceTrueTime: GPS and atomic-clock-backed intervalNTP-class clocks + hybrid logical clockIrrelevant: one writer orders all commits
Uncertainty bound~1–7ms (reported in the 2012 paper)Configured max offset, 500ms default (250ms suggested for well-synced multi-region)N/A
Who paysEvery read-write transaction waits out the uncertainty before acknowledging (overlaps with replication)Reads that land on recently written keys may restartNobody
GuaranteeExternal consistencySerializable; strict serializability for single-key operations, not across arbitrary keysSerializable available on one node

This is the cleanest Staff-level distinction between the two distributed systems. Spanner buys a global commit order with special hardware and a few milliseconds of wait. CockroachDB runs on any hardware and pushes the cost to occasional read restarts. In practice both surface as retryable transaction errors your client must handle; Cockroach's 40001 errors appear more often under contention and with long-running transactions.

3. Multi-Region Write Semantics#

Diagram: 3. Multi-Region Write Semantics
  • Aurora Global Database: one primary region, up to 10 read-only secondary regions as of 2026. Replication is asynchronous at the storage layer, typically under a second. A planned switchover loses nothing; an unplanned regional failover can lose the last seconds of writes. That is the honest RPO.
  • Spanner multi-region: writes go to a leader region and need a quorum across regions. RPO is zero for region loss; every write pays a cross-region round trip.
  • CockroachDB: you choose per database whether to survive a zone (fast writes, region outage makes that region's rows unavailable) or a region (cross-region quorum on every write), and per table or row where data lives (REGIONAL BY ROW, GLOBAL). It is the most flexible locality model of the three, and the easiest to misconfigure.

4. Compatibility and Lock-In#

Aurora runs the real PostgreSQL and MySQL engines on new storage. Extensions, the query planner, EXPLAIN output and ORMs behave as on community Postgres or MySQL, with documented exceptions. Leaving Aurora for RDS or self-hosted Postgres is a dump-and-restore or logical replication project, not a rewrite.

CockroachDB speaks the PostgreSQL wire protocol and much of its SQL, but it is a different engine: different planner, different locking, no Postgres extensions, and behaviors like serializable-by-default and transaction retries that applications must handle. Spanner offers a PostgreSQL dialect, but schema design (interleaving, key choice, avoiding monotonically increasing keys) is Spanner-specific, and it exists only on Google Cloud.

Licensing matters for CockroachDB: since version 24.3 (November 2024) there is a single Enterprise edition, free for individuals and businesses under $10M annual revenue, paid above it. Self-hosting is no longer "free open source" for most companies. Check current terms before you count it as zero license cost.

5. Key Design Is Physical Design#

On Aurora a bad primary key is a slow index. On Spanner and CockroachDB it is a hot range: a monotonically increasing key (timestamp, sequence) sends every insert to the last range, so one node takes 100% of writes while the cluster of 30 idles. Both systems document the fix — UUIDs, hash-sharded indexes (CockroachDB), bit-reversed sequences or UUIDs (Spanner) — and both reward colocating parent and child rows so a business transaction touches one range. Plan the keys first; see Distributed SQL for the full key-design method.

Where Each One Breaks#

SystemFailure modeSymptomDetectionMitigationOwner
AuroraWriter saturationCPU pinned on the writer, commit p99 from 3ms to 200ms, readers idleWriter CPU, commit latency, lock waitsScale up (one-time 2×), move reads off, then shard or migratePlatform + product
AuroraFailover gap20–60s of write errors when the writer fails; connection storms on reconnectFailover events; client error rateConnection proxy, retry with jitter, DNS TTL awarenessPlatform
AuroraUnplanned regional failover loses tail writesLast ~1s of acknowledged writes missing in the new primaryGlobal replication lag metricReconcile from an outbox or event log; accept and document RPOBusiness sign-off
AuroraLong transactions block cleanupReader queries holding old snapshots; storage and vacuum lag growOldest transaction age, MaximumUsedTransactionIDs style metricsStatement timeouts, separate analytics replica or warehouseProduct team
SpannerHot split from sequential keysOne split leader at 100%, write latency climbs, total throughput flat as you add nodesKey Visualizer hot bands, per-split CPURe-key with UUIDs or bit-reversed sequence; this is a migrationProduct team
SpannerCommit size limitsBulk jobs fail: 80,000 mutations or 100 MiB per commit (as of 2026)Job error rateBatch writes, partitioned DML for large updatesProduct team
SpannerCross-region latency surpriseMulti-region config adds 20–70ms to every write; checkout p99 regressesCommit latency by regionRegional config plus async DR if RPO 0 is not truly requiredBusiness sign-off
CockroachDBRetry storms under contention40001 errors spike; ORMs surface them as 500sRestart and retry metricsBounded client retries with backoff; shorten transactions; avoid hot rowsProduct team
CockroachDBClock skewNode exceeding max offset shuts itself down; with bad NTP, several doClock offset metric per nodeReliable time sync (cloud time service); alert well under the maxPlatform
CockroachDBMisconfigured locality"EU" rows actually have voters in the US; writes cross the AtlanticReplica placement reports, write latency by regionLocality reviews in schema CI; survival goals chosen deliberatelyPlatform
CockroachDBSelf-hosted upgrade and rebalancing loadRange rebalancing after a node loss saturates disks; p99 doubles for hoursRebalance rate, disk throughputRate-limit rebalancing; keep 30–40% headroomPlatform

The production surprise for each:

  • Aurora: the writer ceiling arrives as a cliff, not a slope. At 70% CPU it is fine; at 90% lock waits and checkpoint stalls compound and p99 multiplies. Plan the exit at 50%.
  • Spanner: adding nodes does nothing for a hot key. Teams discover their key design only after the cluster is big.
  • CockroachDB: the database is correct, and the application is not. Retry handling is the most common production defect in migrations from Postgres.

Incident Sketch: The Aurora Writer Cliff#

t=0        Marketing campaign starts; write rate climbs from 6K to 11K/s
t=+8min    Writer CPU 88%; commit p99 from 4ms to 60ms; lock waits rising
t=+12min   Connection pool exhausted on app side; request queueing; checkout p99 4s
t=+15min   On-call scales the writer up one size: 6-10 minute instance swap via failover
t=+16min   Failover: ~30s of write errors; reconnect storm doubles the pain
t=+26min   New writer at 50% CPU; queue drains

Detection that should have fired earlier: writer CPU over 50% at normal peak for a week, commit p99 trend, and headroom projections. Lesson: scaling a single writer up is a one-time, disruptive move. Owner: platform for capacity, product for the plan to shed or split writes.

Incident Sketch: The Retry Storm After a Migration to CockroachDB#

t=0        Service migrated from Postgres; ORM treats SQLSTATE 40001 as fatal
t=+3d      Flash sale: 2,000 transactions/s update the same inventory rows
t=+3d+1m   Serialization conflicts spike; 12% of checkouts return HTTP 500
t=+3d+4m   Clients retry blindly with no backoff; contention worsens
t=+3d+20m  Hotfix: bounded retry with jittered backoff, inventory split into 16 sub-counters

Root cause: single-node habits (no retry loop, a hot row) on a serializable distributed database. Prevention: retry handling as a migration gate, contention tests in staging, and hot-row patterns reviewed before cutover.

Cost and Operations#

AuroraSpannerCockroachDB
Who runs itCloud provider runs storage and patching; your platform team picks instance sizes, parameters, upgradesCloud provider runs everything; your team owns schema, keys and capacity in nodesVendor's managed service, or your platform team (commonly 2–4 engineers for a serious self-hosted fleet)
What the bill scales withWriter and reader instance size and count, storage GB, I/O requests (or flat I/O-optimized pricing)Node or processing-unit hours (regional cheapest, multi-region several times higher), storage GBvCPUs or request units on the managed service; nodes plus license when self-hosted
Minimum sensible footprint1 writer + 1 reader across 2 AZsCan start under one node (processing units); production commonly 3+ nodes3 nodes across 3 zones; 5+ for region survival
Cost cliffMoving from one writer to sharding or another databaseMulti-region configurations; over-provisioned nodes for peakRegion survival (5 replicas instead of 3, ~1.7× replicas)

Rough shape, as of 2026 list prices and order of magnitude only: a production Aurora cluster with a large writer and two readers commonly costs low thousands of dollars a month; a regional Spanner deployment of 3 nodes is in the same low-thousands range before storage, and multi-region configurations multiply that; a self-hosted CockroachDB cluster of 9 mid-size nodes costs similar infrastructure plus license and the engineers to run it. The real difference shows at scale: distributed SQL costs more per transaction than a single primary, and you pay it to remove the ceiling.

ScaleSensible choiceRough monthly infrastructurePeople
Startup (1K writes/s, one region)Aurora or managed PostgresHundreds to low thousands of dollarsPart of one engineer's time
Growth (10K writes/s, DR region)Aurora with Global Database, shard-friendly schemaLow to mid thousands1–2 platform engineers
Large (100K+ writes/s, RPO 0 or residency)Spanner or CockroachDB, multi-region for the core onlyTens of thousands and upPlatform team plus migration effort measured in engineer-quarters

Assumptions: list prices as of 2026, order of magnitude only; storage and I/O excluded; the people column dominates below the large tier.

🧭 Principal Insight: The expensive part of distributed SQL is not the nodes. It is the application changes: retries, key redesign, transaction scoping, and teams relearning a planner. Budget engineer-quarters, not just dollars per node.

Switching Later#

MigrationDifficultyWhat's hard to undo
Postgres/MySQL → AuroraEasy: logical replication or the provider's migration path; same engineLittle; you can move back to RDS or self-hosted the same way
Aurora → CockroachDBModerate to hard: wire-compatible, but schema, keys, extensions and retry logic changeApplication assumptions about single-node behavior; sequences and hot keys
Aurora → SpannerHard: new dialect choices, interleaving, key redesign, cloud moveCloud commitment; Spanner-specific schema
Spanner → anythingHard: no other place runs SpannerThe cloud and the schema
CockroachDB → PostgresHard if you rely on multi-region locality; moderate if single regionLocality-based designs have no single-node equivalent
Single region → multi-regionEasy to configure, hard to live withPer-write latency becomes a product property; rolling back means moving data
Diagram: Switching Later

The one-way doors:

  1. Spanner is a cloud commitment. Choose it when you have already chosen Google Cloud, not as a way to choose it.
  2. Multi-region write configurations become product behavior: once checkout has absorbed 40ms of cross-region commit, nobody remembers how fast it used to be, but every feature depends on that durability.
  3. Key design on any range-sharded system. Re-keying a 5 TB table is a dual-write migration measured in quarters.

A migration off a single writer, in order:

  1. Add CDC from Aurora into a log (Kafka) so the new database can be filled and kept current without touching the application.
  2. Backfill, then run shadow reads: every read goes to both, results compared, differences logged and counted.
  3. Move writes table group by table group, starting with the one causing the writer pressure, behind a feature flag with a rollback path.
  4. Keep the old writer read-only for a soak period of weeks before decommissioning it.

The cheap two-way door is the one most teams skip: design for sharding while still on Aurora. Random keys, no cross-tenant transactions, idempotent writes, an outbox for side effects. Those choices cost almost nothing on one writer and turn a later migration from a rewrite into a data move.

How Real Companies Chose#

Uber — Spanner for the Fulfillment Platform#

Uber rebuilt its fulfillment platform, which had used Cassandra and Redis with best-effort consistency and a peer-to-peer coordination layer that had physical scaling limits. It evaluated evolving the NoSQL stack, an in-house MySQL-based store, and NewSQL options (CockroachDB, FoundationDB and Spanner), and chose Google Cloud Spanner in a North America multi-region configuration for external consistency, multi-row and multi-table transactions, and horizontal scaling by key range, after benchmarking on availability SLA, operational overhead, transactions, schema and shard management (Uber Engineering).

Staff insight: The trigger was multi-entity transactions that the application had been faking with sagas and best-effort logic. Distributed SQL earns its cost when it deletes application-level consistency code.

Figma — Staying on RDS Postgres and Sharding It#

Figma evaluated CockroachDB, TiDB, Spanner and Vitess and chose to horizontally shard its existing RDS Postgres instead. Its reasons: any alternative required a complex migration to ensure consistency, the team would have had to rebuild years of operational expertise, it had only months of runway at its growth rate, and it preferred known low-risk solutions it controlled. A custom solution could also support a much smaller feature set than a general database (Figma Blog).

Staff insight: Distributed SQL is not automatically the answer to "Postgres is full". Migration risk, team expertise and runway are first-class inputs, and saying so is a Staff signal.

Amazon — Why Aurora Separated Compute from Storage#

The Aurora paper argues that in the cloud the network, not compute or storage, is the binding constraint, and so the engine pushes only redo log records to a multi-tenant storage service. Data is kept as 6 copies across 3 availability zones with quorum writes and reads, giving fast crash recovery, replica failover without data loss, and self-healing storage (SIGMOD 2017 paper).

Staff insight: Aurora was designed to make the single-writer model durable and fast to recover, not to remove it. Knowing what a system was built for tells you exactly where it stops.

Follow-Ups to Expect#

After You Say...They Will Ask...What They're Testing
"Aurora to start""What's your trigger to move off a single writer?"A measured threshold (writer CPU, commit p99) and a plan before the cliff
"Aurora Global Database for DR""What's your RPO on an unplanned regional failover?"Knowing replication is async: about a second of writes at risk
"Spanner""What does a write cost in a multi-region config, and who signed off?"Cross-region quorum latency and the business decision behind it
"Spanner""Your insert rate doubled but throughput didn't. Why?"Hot splits from sequential keys
"CockroachDB""How does your application handle transaction retries?"40001 handling, idempotent transaction bodies, no side effects inside
"CockroachDB""How do you keep EU data in the EU?"REGIONAL BY ROW, survival goals, verifying replica placement
"Distributed SQL scales writes""What happens to a transaction touching 5 ranges in 3 regions?"Distributed commit cost; colocating related rows
"Strong consistency everywhere""Which reads can be stale?"Follower or stale reads for dashboards and reporting
"Region survival""Which tables actually need it?"Scoping expensive guarantees to the data that needs them
"Spanner is fully managed""What do you still own?"Schema, keys, capacity, and the cloud commitment
"Aurora failover is fast""What does the app see during those 30 seconds?"Connection handling, retries, proxies

What to Say in the Interview#

"Aurora scales reads and storage behind one writer; Spanner and CockroachDB scale writes by sharding the keyspace into consensus groups. The first question is whether one writer carries us for three years."

"I'd start on Aurora with UUID keys and short, single-tenant transactions, and set a move trigger at 50% writer CPU at peak, because the single-writer ceiling arrives as a cliff."

"If we need writes to survive a region loss with zero data loss, every write pays a cross-region round trip — 20 to 70 milliseconds — and I want the business to sign off on that, not discover it."

"Between Spanner and CockroachDB, the deciding factor is usually the cloud: Spanner if we're on Google Cloud and want it fully managed, CockroachDB if we need portability or row-level data residency."

  1. Loading the index…