Hiring BarSupport
← All calculators

Sharding planner

From data size, growth, read and write rates and per-node limits, work out shards and replicas now and at your horizon, and when you will need to reshard.

Inputs

Workload

One copy, before replication
Peak, not average
Peak, after caching

Per node

Plan

Headroom for spikes and node loss

Plan

Shards now
721 nodes · write throughput-bound
Shards at 24 months
1545 nodes · storage-bound
Replicas per shard
32 followers each
Reshard with today’s count
month 3inside your horizon
Suggested logical shards
64≈4× the horizon count, power of two
015nowmonth 24today: 7

Data at horizon: 17 TB per copy · 32.2K writes/s · 161K reads/s

How this is calculated

Usable capacity per shard
storage = node TB × utilisation, writes = node writes/s × utilisation (every replica applies every write), reads = node reads/s × utilisation × readable copies.
Shards
shards = ceil(max(data ÷ storage cap, writes ÷ write cap, reads ÷ read cap)). The largest term is what the design is bound by.
Growth
data(m) = data + growth × m and traffic(m) = traffic × (1 + g)ᵐ. Resharding month is the first month the shard count you need exceeds today’s.
Logical shards
Map keys to many more logical shards than machines (here ≈4× the horizon count, rounded to a power of two), so growing means moving whole shards rather than re-splitting keys.

What to say in the interview

  • “At 60% target utilisation, one shard holds about 1.2 TB and takes 3K writes a second, so today we need 7 shards, 21 nodes with 3× replication, and we’re write throughput-bound.”
  • “Over 24 months that grows to 15 shards. If we start at 7, we’d have to reshard around month 3, so I’d rather provision for the horizon or plan the split now.”
  • “To keep resharding cheap I’d hash keys into 64 logical shards and map those to machines, so adding nodes moves whole shards instead of rehashing every key.”
  1. Loading the index…