Hiring BarSupport

Cell-Based Architecture

Pattern47 min read5 diagrams

Technologies that implement this pattern: Kubernetes · API Gateway · DynamoDB · etcd & ZooKeeper · PostgreSQL · Apache Kafka

Why This Matters#

Every large outage post-mortem has the same sentence in it: "a change intended for one component affected all of them." A bad config push, a poison request that crashes every node that touches it, a deploy with a memory leak, a database migration that locks a hot table, one customer's traffic spike that saturates a shared queue. Redundancy does not help — five replicas of the same binary with the same config and the same poison request fail together. Availability zones do not help when the failure is in your software rather than the data center. What helps is having many independent copies of the whole system, each serving a fixed slice of customers, so that whatever goes wrong goes wrong in one slice.

Most candidates treat this as a scaling question: "we'll shard the service." Staff engineers treat it as a blast-radius question. The design question is not how do I handle more load but when — not if — something goes wrong, what fraction of customers is affected, and how do I make that fraction small, known in advance and enforceable? A 99.95% system with one cell has a worst day that hits 100% of customers. The same system as 20 cells has a worst day that hits 5%, and most changes reach that 5% first, while the other 95% are still on the old version.

The second reframe: a cell architecture is only as isolated as the things the cells share. The router in front of every cell, the global identity service every cell calls, the control plane that pushes config to every cell, the one Kafka cluster they all publish to — each of those is a path by which one failure reaches every cell at once. Drawing the cells is easy. The Staff work is the inventory of what is still shared, the decision about which of those to cell-ify and which to harden, and the discipline to keep the router so simple it almost cannot fail.

If you can walk an interviewer from "what's the tolerable blast radius" to "what is the partition key and how big is a cell" to "how does a request find its cell when the control plane is down" to "how do we move a tenant and deploy a change one cell at a time," you are answering at Staff level.

The 60-Second Version#

  • Blast radius is 1/N, so pick N from the business. If losing 5% of customers for an hour is survivable and losing 25% is not, you need at least 20 cells. Work back from that, not from hardware.
  • Cap cell size and test at the cap. A cell has a published maximum — e.g. 500 tenants, 20K requests/s, 2 TB — and a load test that proves it. Past the cap you add a cell; you never grow one into an untested regime.
  • The router must be the thinnest layer in the system. Partition key in, cell out, from a locally cached map. No business logic, no synchronous dependency on the control plane; if the directory is down, route on the last-known map for hours.
  • Deploy cell by cell, with bake time. A typical wave plan: 1 canary cell → 5% → 25% → 50% → 100%, with 30–60 minutes of bake per wave and automatic rollback on cell-level SLO breach. A bad change caught in wave one touches ~1–5% of customers.
  • Tenant migration is a routine operation, not a project. Copy, catch up via change stream, pause writes for ~10–60 seconds, verify, flip the directory. Rehearse it weekly on synthetic tenants; you'll need it for whales, hot spots and evacuating a sick cell.
  • Shared dependencies are the real blast radius. List every component all cells call. Each one either gets cell-ified, gets a static fallback the cells can run on for hours, or is accepted as a global risk with a name, an owner and an error budget.

The Problem#

A B2B communications platform serves 9,000 customer organizations from one regional stack: a fleet of API servers behind a load balancer, a sharded PostgreSQL cluster keyed by organization ID, a Redis cluster, a Kafka cluster and a config service. In one quarter it has three major incidents. A feature-flag change with a malformed rule is evaluated by every API server within 40 seconds and crashes them all: a 35-minute full outage. A large customer's integration starts sending messages with a 2 MB attachment field that triggers a pathological parse; every server that receives one stalls, and the load balancer spreads them evenly, so all servers stall. And a schema migration on the shared database locks a table that every request touches for 11 minutes. In each case the post-mortem says the same thing: the system was redundant but not isolated. Every server ran the same code and config and saw a uniform mix of every customer's traffic, so every failure was total. The job is to restructure the platform so that a bad change, a poison request or a misbehaving customer can only affect a bounded, known slice of customers; so that changes reach a small slice first; so that customers can be moved between slices; and so that the few components that remain shared are hardened enough not to undo all of that.

Diagram: The Problem

Case Studies That Use This Pattern#

  • Multi-Region — Regions are the coarsest cells; cells inside a region bound the failures that regions can't
  • Deployment System — Waves of cells, bake times and automatic rollback are what make cells pay off for change safety
  • Cascading Failures — Cells are the structural answer to a failure that spreads through shared pools and retries
  • Sharded Database — A shard is a data partition; a cell is the whole stack for a partition, and the tenant directory is shared between the two ideas
  • Edge Gateway — The cell router usually lives here, and must be the simplest component in the request path
  • Feature Flags — Config pushed globally is the classic cell-escaping failure; per-cell rollout of flags is part of the design
  • Service Registry — Discovery scoped per cell, so a registry failure or bad entry stays inside one cell
  • Load Balancing — Balancing within a cell and routing between cells are different problems with different keys

Which Problem Are We Solving?#

"Make it cellular" hides four different goals. They lead to different cell boundaries and different shared components.

IntentConstraintStrategyFailure ModeCorrectness Bar
Blast-radius containment for a multi-tenant serviceThousands of tenants; deploys daily; poison requests and noisy tenants happenFull-stack cells keyed by tenant ID; router plus directory; cell-by-cell deploysShared dependency or router takes all cells down; whale outgrows a cellAny single fault affects ≤ 1/N of tenants; changes reach the canary cell first
Availability-zone fault isolationGray failures in one zone degrade cross-zone trafficZonal cells: each zone serves its own traffic end to end; drain a zone by reweightingServices that can't be zone-siloed; capacity in remaining zonesA sick zone can be drained in minutes with no customer-visible errors
Bounded scale unitGrowth is unbounded; scaling cliffs appear at untested sizesCells with a capped, load-tested size; add cells to growCells drift above the tested cap; cell count explodes operational loadNo cell exceeds the size it was tested at
Residency and dedicated isolationSome tenants must stay in a region or on dedicated capacityCells pinned to regions; dedicated cells for whales and regulated tenantsResidency violated by a shared global componentTenant data and processing provably stay in the assigned cell

🎯 Staff Move: "I'll treat this as intent one: a multi-tenant service where the incidents were a global config push, a poison request and a shared migration. So I want full-stack cells keyed by org ID, about 20 of them so the worst case is 5% of orgs, a router that runs on a cached map, and deploys that go to an internal canary cell first. Then I'll go through every shared component and decide what happens to it."


The Core Tradeoff#

StrategyWhat WorksWhat BreaksWho Pays
One big stack, horizontally scaledSimplest; best utilization; one thing to operateEvery bad change, poison request and noisy tenant affects everyoneEvery customer on the worst day
Sharded data onlyData scales; some data-layer isolationStateless tier, cache, queue and config still shared; poison requests hit all serversCustomers during app-tier incidents
Shuffle sharding within one stackBounds overlap between tenants on shared nodes; cheapGlobal config and code still shared; helps noisy neighbours more than bad deploysOps complexity; not a deploy-safety story
Zonal cellsAligns with infrastructure failure domains; fast drainsBlast radius is a third of traffic; software faults still cross zones if code is globalCapacity headroom (each zone sized to absorb others)
Full-stack cells by tenantBounded blast radius for software, data, tenants and deploysRouter and shared services become critical; cross-cell features get hard; more fixed cost per cellPlatform team (router, directory, migration); finance (per-cell overhead)
Dedicated cell per tenantMaximum isolation; simple story for enterprise buyersCost explodes; thousands of stacks to deploy and patchThe customer (priced in) or the margin

Staff Default Position#

Partition the whole stack into capped, identical cells keyed by the tenant, route with the thinnest possible layer from a cached map, deploy one cell at a time behind a canary cell, make tenant migration a rehearsed routine, and keep an explicit, shrinking inventory of every component still shared across cells.

The default shape: 10–50 cells per region, each a complete copy of the serving stack — API, workers, database, cache, queue — sized with a hard cap (e.g. ≤ 25% of the cell's tested capacity for any single tenant, and a maximum tenant count) and load-tested at that cap. A tenant directory (tenant → cell) owned by the control plane, replicated to the router as a versioned, cached map; the router does a lookup and forwards, nothing else, and keeps working on its last-known map if the control plane is down. New tenants are placed by the control plane on the least-loaded cell under its cap. Deploys and config changes go to one internal canary cell, then waves, with bake time and automatic rollback on cell-level SLOs. Whales get dedicated cells on the same template. Cross-cell needs — global search, admin views, billing aggregation — run asynchronously off per-cell event streams, not as synchronous calls between cells.


When to Deviate#

  • Small scale or single-team service. Below a few hundred tenants or a handful of servers, cells multiply fixed cost and operational surface for little gain. Use shuffle sharding and careful staged deploys first; adopt a cell router when the directory and migration tooling would pay for themselves.
  • No natural partition key. A social graph where every request touches many users, or a global marketplace where any buyer meets any seller, has no grain that keeps requests inside one cell. Cells may still work for a subset (per-region order processing), but forcing them onto the graph creates cross-cell chatter that is worse than the shared risk.
  • Zone failures are the dominant risk. If incidents are mostly infrastructure gray failures rather than bad deploys or tenants, zonal cells with fast drains buy most of the value with far less migration machinery.
  • Strong global invariants. A global uniqueness constraint (usernames) or a global counter cannot live in a cell. Keep a small, hardened global service for it, with a cached or degraded mode, and say it is the cell architecture's known shared risk.

One Question, Three Levels#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Shard the database by customer and add more replicas""What fraction of customers can we lose at once? That sets the cell count. Then partition key, router and the list of shared dependencies.""Is cell-based the org's default failure posture? Which services must be cellular, who runs the router and directory as a platform, and what does it cost per cell?"
Cell size"Make cells big so they're efficient"Capped and load-tested at the cap; largest tenant ≤ ~25% of a cell; size chosen from blast radius and test costStandard cell sizes across services so cells line up; capacity planning in cells, not servers
RoutingLoad balancer with hash on customer IDThin router with a cached directory map, last-known-good on control-plane failure, no business logicRouter and directory as a tier-0 platform with its own deploy discipline and game days
Shared dependenciesNot inventoriedInventory every shared component; cell-ify, give a static fallback, or accept with an ownerError budgets for shared services set tighter than any cell; global services minimized by policy
DeploysRolling deploy across the fleetCanary cell, then waves with bake time, automatic rollback on cell SLOs; config and flags follow the same wavesOrg-wide wave calendar; no global push path exists, including for emergencies, without approval
Migration"We'd rebalance if needed"Copy, catch up, short write pause, verify, flip — rehearsed weekly; used for whales, hot spots and evacuationMigration throughput as a capacity metric; cell retirement and rebuild as routine operations
Why "First move" separates levels

Sharding the database is a scale answer, and it leaves the failures in the problem untouched: a global config push, a poison request and a shared migration would all still be total. The Staff candidate starts from the tolerable blast radius — a business number — and derives the architecture from it. The Principal candidate asks whether every critical service will become cellular, because a cellular API in front of a non-cellular dependency is still one failure away from a global outage.

Why "Routing" separates levels

The router is the one component every request crosses, so it is the place where a cell architecture is most likely to fail globally. The L5 instinct is to make it smart — hash, plus overrides, plus feature logic, plus a synchronous call to the directory. The Staff answer makes it boring: a lookup in a local copy of a small map, versioned and pushed asynchronously, with a last-known-good fallback that can run for hours. If the directory service is down, new tenants can't be placed and migrations pause, but every existing tenant keeps being routed.

Why "Shared dependencies" separates levels

Cells are only independent if what they depend on is independent. Seniors draw cells and stop. Staff engineers list identity, configuration, secrets, the container registry, DNS, the observability pipeline, the global billing service, and decide for each one: make it per-cell, give cells a cached or static mode so they can run for hours without it, or accept it as a named global risk with a tighter error budget than any cell. That inventory is the actual blast-radius analysis; the cell diagram is just the picture of it.


Where the Design Splits#

#Fault LineThe Tension
1Few Large Cells vs Many Small CellsUtilization and fewer things to operate vs smaller blast radius and easier full-scale testing
2Hash Routing vs Directory RoutingA router with no state to lose vs placement you can change tenant by tenant
3Per-Cell vs Shared DependenciesIsolation for every component vs the cost and impossibility of cell-ifying global concerns
4Static Placement vs Active RebalancingSimple, predictable cells vs moving tenants to fix hot spots and whales
5Zonal Cells vs Tenant CellsMatching infrastructure failure domains vs matching software and tenant failure domains

Fault Line 1: Few Large Cells vs Many Small Cells#

Blast radius is roughly 1/N. Five cells mean a cell failure affects 20% of customers; 50 cells mean 2%. But every cell has fixed cost — its own database primary and replicas, cache, queue, minimum API capacity, monitoring — and fixed operational weight: 50 databases to patch, 50 deploy targets, 50 dashboards. Who pays: large cells make customers pay a larger share of every outage, and make the platform pay when a cell drifts into a size nobody has tested; small cells make finance pay fixed overhead (often 10–30% more infrastructure at small cell sizes) and make the platform pay in automation it can no longer skip. Staff default: choose N from the tolerable blast radius (often 5% → ~20 cells per region), then pick the cell cap so the largest normal tenant is ≤ ~25% of a cell and the cell can be load-tested at full size in staging for a reasonable cost. Deviate when: fixed cost per cell is very low (stateless services on shared compute) — then go smaller; or when the largest tenants are huge — give them dedicated cells instead of making every cell large.

Fault Line 2: Hash Routing vs Directory Routing#

Hash routing computes cell = hash(tenant_id) mod N or a consistent-hash ring; the router needs no state. Directory routing looks up tenant → cell in a table the control plane owns. Who pays: hash routing makes operations pay when anything needs to move — adding a cell re-homes a fraction of all tenants, and a single hot tenant cannot be moved alone; directory routing makes the platform pay for a directory service, map distribution, caching and consistency during migrations, and adds a lookup the router must survive without. Staff default: directory routing, with the map small enough to replicate to every router instance (10K tenants × ~50 bytes ≈ 500 KB; even 10M tenants ≈ 500 MB fits in memory or a local store), pushed as versioned snapshots plus deltas, and a last-known-good fallback. Use hash routing as the default placement for new tenants and the directory for exceptions — this is the best of both. Deviate when: tenants are tiny, numerous and uniform (per-device telemetry) — then hash routing with a hash ring and bulk movement on cell add is simpler.

Fault Line 3: Per-Cell vs Shared Dependencies#

Some dependencies cell-ify naturally: databases, caches, queues, workers. Others resist: identity (a user may belong to several tenants), billing aggregation, global search, the container registry, DNS, the deploy system itself. Who pays: cell-ifying everything makes the platform pay in duplicated services and makes cross-cell features slow or eventually consistent; sharing makes every customer pay when the shared service fails, which undoes the cell boundary. Staff default: for each shared dependency, pick one of three: (1) cell-ify it; (2) make cells able to run without it for hours — cache tokens and validate locally with rotating public keys, cache config with last-known-good, pre-pull images; or (3) accept it as a named global risk with an error budget tighter than any single cell's and its own staged-change discipline. The request path should have zero synchronous cross-cell calls and as few synchronous global calls as possible. Deviate when: the global service is read-mostly and tiny (a feature catalogue) — then replicate it into every cell rather than calling it.

Fault Line 4: Static Placement vs Active Rebalancing#

Static placement assigns a tenant to a cell at signup and leaves it. Active rebalancing moves tenants to even out load, isolate whales and evacuate sick cells. Who pays: static placement makes customers on the hottest cell pay as tenants grow unevenly — tenant growth follows a power law, so within a year some cells run at 80% while others idle at 30%; active rebalancing makes the platform pay for a migration system that copies data, catches up changes, pauses writes and flips routing safely, and makes migrated tenants pay a short write pause and a risk of something going wrong mid-move. Staff default: place statically, but build migration on day one and use it deliberately — for tenants crossing ~25% of a cell, for cells crossing ~70% of their cap, and for evacuation — with a human approving large moves. Rehearse weekly with synthetic tenants so the tool works on the day you need it. Deviate when: the data is huge per tenant and moves take days; then plan headroom per cell for tenant growth and prefer splitting new work into new cells.

Fault Line 5: Zonal Cells vs Tenant Cells#

Zonal cells align with availability zones: each zone runs a full copy of the stack serving the traffic that lands in it, and draining a zone shifts traffic to the others. Tenant cells partition by customer, regardless of zone; each tenant cell typically spans zones for its own redundancy. Who pays: zonal cells make capacity pay — each zone must absorb its share of a drained zone, so you run ~50% headroom with three zones — and leave software faults global, since every zone runs the same code at once; tenant cells make the platform pay for routing, placement and migration machinery, and leave zone failures to each cell's internal redundancy. Staff default: tenant cells for blast-radius containment of software, data and tenant faults; zone awareness inside each cell for infrastructure faults. Deviate when: the dominant incidents are zone gray failures and the services are hard to partition by tenant — zonal cells with fast drains are a large improvement at lower cost, and can be the first step toward tenant cells.


Common Interview Mistakes#

What Candidates SayWhat Interviewers HearWhat Staff Engineers Say
"We'll shard by customer for isolation""Data partitioning, not failure isolation""Sharding the data still leaves code, config and the app tier shared. A cell is the whole stack for a slice of customers."
"The router hashes the customer ID and calls the directory to check overrides""The router now has a synchronous dependency""The router reads a local, versioned copy of the map. If the directory is down, it routes on last-known-good for hours."
"Each cell is independent""Hasn't inventoried shared dependencies""Each cell is independent except identity, config and the deploy system. Here's what each cell does when those are down."
"Make cells as large as possible for efficiency""Will grow cells into untested territory""Cells have a hard cap that we load-test. Past the cap we add a cell; we don't grow one."
"We'll rebalance customers if a cell gets hot""Has never moved a tenant with live writes""Migration is copy, catch up, a 10–60 second write pause, verify, flip — and we run it weekly on synthetic tenants."
"Deploy to all cells with a rolling update""Throws away the main benefit""Canary cell first, then 5%, 25%, 50%, 100% with bake time, and config and flags follow the same waves."

Quick Reference#

Diagram: Quick Reference

Staff Sentence Templates#

"The business can tolerate losing about [X]% of customers for [duration], so I want at least [N] cells. Each cell is capped at [tenants / requests per second / storage], load-tested at that cap, and no tenant exceeds [25]% of one."

"The router does one thing: [partition key] to cell, from a local copy of the directory map at version [v]. If the control plane is down, it routes on last-known-good; new signups and migrations pause, existing tenants are unaffected."

"These are the components every cell still shares: [list]. [A] is cell-ified, [B] has a cached mode that lasts [hours], and [C] is an accepted global risk owned by [team] with an error budget of [X]."

"Changes reach the canary cell first, then [waves] with [minutes] of bake each. Automatic rollback triggers if any cell's [SLO metric] degrades by [threshold] relative to the others."


Cell Size: The Arithmetic#

Cell size is a negotiation between three numbers: the blast radius the business accepts, the largest tenant you must fit, and the fixed cost of a cell. Run them explicitly.

inputs (illustrative B2B SaaS, one region):
  tenants                      = 9,000
  peak load                    = 120K requests/s
  largest tenant               = 3.5K requests/s (~3% of total)
  tolerable blast radius       = 5% of tenants  ->  N >= 20 cells
  fixed cost per cell          = DB primary + 2 replicas, cache, queue, min API = ~$9K/month

per-cell at N = 20:
  tenants per cell             = 450
  average load per cell        = 6K requests/s; plan cap at 2× average = 12K requests/s
  largest tenant share of cap  = 3.5K / 12K ≈ 29%   -> above 25%: give it a dedicated cell
  fixed overhead               = 20 × $9K = $180K/month vs ~$110K for one big stack of equal capacity

per-cell at N = 50:
  blast radius                 = 2%; fixed overhead 50 × $9K = $450K/month
  largest tenant share of cap  = 3.5K / 4.8K ≈ 73%  -> whales must be dedicated

Use the availability calculator to turn cell count into "customer-minutes lost per incident," and the cost estimator to price fixed overhead per cell. The shape of the result is almost always the same: pooled cells sized for the long tail, dedicated cells for the few tenants that would dominate a pooled one, and a cell count driven by blast radius rather than efficiency.

Cell Count (per region)Blast RadiusFixed OverheadTest Cost at Full SizeTypical Fit
3 (zonal)~33%LowHigh (each cell is huge)Zone fault isolation
5–1010–20%Low–mediumMediumEarly adoption, few whales
20–502–5%Medium–highLowMature multi-tenant SaaS
100+< 1%High; needs full automationVery lowHyperscale services with many small tenants

🎯 Staff Insight: "Same 30 hosts, now in cells" is a valid starting point. Cells don't have to mean more hardware on day one — they mean the existing hardware is grouped, routed and deployed in independent slices. The overhead grows with per-cell stateful components, so that is where to look when finance asks.


Implementation Deep Dive#

1. The Cell Router — Thinnest Possible Layer, Cached Map, Last-Known-Good#

# Router (Envoy filter, edge service, or gateway plugin) — runs in every edge instance
state:
  map_version      = 18342
  tenant_to_cell   = {...}            # replicated snapshot + deltas, ~50 bytes per tenant
  default_placement = consistent_hash(tenant_id, cells_active)  # for tenants not in the map yet
  cell_endpoints   = {cell_07: "cell-07.internal", ...}

route(request):
  tenant = extract_partition_key(request)       # from host, path or signed token claim — never the body
  cell   = tenant_to_cell.get(tenant) or reject_unknown(tenant)
  if cell.state == "migrating_write_pause" and request.is_write:
      return 503 + Retry-After: 5                 # bounded pause during a move
  forward(request, cell_endpoints[cell])

map updates:
  control plane publishes versioned snapshots to object storage + deltas on a stream
  router applies deltas in version order; on gap, reloads snapshot
  if control plane unreachable: keep serving map_version; alert after 5 min; no expiry for hours
  router changes deploy cell-by-cell too: router fleet split into groups, canary group first

What stays out of the router: authentication beyond extracting a signed tenant claim, rate limiting by plan, A/B logic, request transformation, any synchronous call. Each feature added to the router is code that runs in front of every cell at once — the exact thing cells exist to avoid. AWS's guidance on cell-based architecture calls the router the thinnest possible layer, responsible for routing requests to the right cell "and only that" (What is a cell-based architecture?).

🎯 Staff Move: "If someone wants to add rate limiting to the router, I'd push it into the cell's own gateway. The router's only job is tenant to cell; everything else runs inside a blast radius."

2. Tenant Placement and Migration — Copy, Catch Up, Pause, Verify, Flip#

place(new_tenant):
  candidates = cells where state == active and tenants < cap and load < 70% of cap
                     and region == tenant.residency
  cell = argmin(candidates, projected_load)        # or consistent-hash default, then check caps
  directory.put(tenant, cell, version++)

migrate(tenant, src, dst):                          # used for whales, hot cells, evacuation
  1. preflight: dst has headroom >= tenant peak × 1.5; same schema version; same code version
  2. bulk copy tenant rows/objects src -> dst       (online, throttled to protect src)
  3. catch up via change stream (CDC) filtered by tenant_id until lag < 1s
  4. directory: tenant state = write_pause (router returns 503 + Retry-After on writes)
  5. drain in-flight writes on src; apply final changes; verify row counts + checksums
  6. directory: tenant -> dst, state = active, version++   (routers pick it up in ~1–5s)
  7. src keeps tenant data read-only for 7 days, then deletes
  typical write pause: 10–60s; abort and roll back to src at any step before 6
  rehearse: weekly on synthetic tenants in every cell; track p99 pause and failure rate

The parts that bite: background jobs and queues holding work for the tenant in the source cell (drain or re-enqueue them in the destination), caches in the destination that start cold (pre-warm for large tenants), external webhooks or partners that cached the source cell's endpoint (route everything through the router, never expose cell addresses), and IDs that encode the cell (never put the cell in a tenant-visible identifier). Large tenants make all of these worse, which is why moving a whale is a scheduled, human-approved operation.

Diagram: 2. Tenant Placement and Migration — Copy, Catch Up, Pause, Verify, Flip

3. Cell-by-Cell Deployment — Waves, Bake Time, Relative Health#

wave plan (code, config, feature flags and schema changes all follow it):
  wave 0: canary cell (internal orgs + synthetic tenants)   bake 60 min
  wave 1: 1 customer cell (~5%)                              bake 60 min
  wave 2: 4 cells (~25% cumulative)                          bake 30 min
  wave 3: 5 cells (~50%)                                     bake 30 min
  wave 4: remaining cells                                    bake 30 min, then done
  total: ~4 hours for a routine change; emergency fixes may compress bake, never skip waves

automatic rollback per cell if, vs the cells not yet deployed:
  error_rate     > 2× baseline cells for 5 min
  p99_latency    > 1.5× baseline cells for 5 min
  crash_loops    > 0 on more than 10% of hosts
  -> halt pipeline globally, roll back the affected cell(s), page the change owner

Why compare against undeployed cells: comparing a cell to its own history catches regressions only when they are large; comparing a deployed cell to cells still on the old version during the same minutes cancels out time-of-day traffic and shared-dependency noise. This is the payoff of having many identical cells — every wave is a controlled experiment.

Config is code: the incident that started this page was a feature-flag change. Flags, dynamic config, secrets rotation and schema migrations must flow through the same waves. A config system with a "push to all" button is a global blast-radius path; if it must exist for emergencies, it needs two-person approval and an audit trail. See Feature Flags and Deployment System.

4. Evacuating a Sick Cell — Drain, Fail, Rebuild#

cell health states: active -> degraded -> draining -> evacuated -> rebuilding -> active

trigger evacuation when:
  cell SLO burn rate > 10× for 15 min AND not explained by a single tenant
  or infrastructure fault that cannot be repaired in < 1 hour

evacuate(cell):
  1. mark cell 'draining' in directory: no new tenants placed
  2. migrate tenants out, largest first, to cells with headroom (parallelism limited by
     destination headroom; with 20 cells at <= 70% load, ~1 cell's worth fits elsewhere)
  3. once empty: rebuild from the template (infrastructure as code), rejoin as active

The headroom rule: if every cell runs at 90% of its cap, no cell can be evacuated without overloading the rest. Planning each cell to run at ≤ 70% of its tested cap leaves room to absorb one cell's tenants across 19 others with margin. For zonal designs, the same logic means each zone must be able to absorb its share of a drained zone.

Technique Comparison

TechniqueBlast Radius ControlChange SafetyTenant MobilityFixed CostOps BurdenBest For
One stack, rolling deploysNoneLowN/ALowestLowSmall services
Sharded data, shared app tierData onlyLowShard movesLowMediumData scale without isolation needs
Shuffle shardingBounded tenant overlapLowReassign shard setsLowMediumNoisy neighbours on shared fleets
Zonal cells~1/3 per zoneMedium (zone-wave deploys)N/AMedium (headroom)MediumInfrastructure gray failures
Tenant cells, hash routing1/NHighPoor (bulk moves only)Medium–highMediumMany small uniform tenants
Tenant cells, directory routing1/NHighPer-tenant migrationMedium–highHighB2B SaaS with whales
Dedicated cellsPer tenantHighTenant is the cellHighestHighEnterprise and regulated tenants

Architecture Diagram#

Diagram: Architecture Diagram

How to narrate it: solid arrows are the request path — DNS, router, one cell, and nothing else. Every control-plane arrow is dashed: the directory, provisioning, migration and deploys can all be down without dropping a request. The shared services are drawn separately on purpose: identity is called only to refresh signing keys, which cells cache, and billing consumes cell event streams asynchronously. The platform team owns the router, directory and migration service as tier-0; each cell is operated identically by automation; product teams own their service's behaviour inside every cell.


Failure Scenarios#

1. The Router Became the Monolith — One Bad Rule Takes Down Every Cell#

The router accumulated features over two years: plan-based rate limits, a header rewrite for a partner, and a regex-based override list.

t=0       Override list update adds a pattern with catastrophic backtracking.
t=0       Router fleet updated globally (router deploys never went through waves).
t=+20s    Router CPU at 100% on every instance; p99 routing latency 8s.
t=+1min   All 20 cells healthy and idle; 100% of customers see timeouts.
t=+18min  Override list reverted; routers recover.

Detection: router.cpu and router.route_latency_p99 across all instances; healthy cells with near-zero traffic is the tell. Blast radius: every customer — the cell architecture bypassed by its own front door. Mitigation: revert the override list; restart routers. Prevention: strip the router back to lookup-and-forward; move rate limits and rewrites into each cell's gateway; deploy router code and data in groups with a canary group first; validate map and rule updates before publishing. Owner: platform team owns the router and its change process.

🎯 Staff Insight: Every line of logic in the router is code in front of every cell at once. The router's feature list is the most important design document in a cell architecture, and it should be nearly empty.

2. The Shared Identity Service — All Cells Up, Nobody Can Log In#

t=0       Identity service database fails over; token issuance and introspection time out.
t=0       Every cell validates tokens by calling identity's introspection endpoint per request.
t=+30s    All 20 cells return 401 or time out on authenticated requests.
t=+2min   Cell retries triple load on identity, delaying its recovery.
t=+26min  Identity recovers. 26 minutes of global outage with every cell healthy.

Detection: auth.introspection_latency and error rate across cells; correlated 401 spikes in all cells at once. Blast radius: 100% of authenticated traffic. Mitigation: none quick without a local validation path. Prevention: cells validate signed tokens locally with cached public keys (rotated with overlap so a key fetch failure doesn't matter for hours); identity issuance has an error budget tighter than any cell and a cell-aware rollout of its own; client retry budgets. Owner: identity team owns issuance and key rotation; platform owns the shared-dependency inventory that should have flagged this.

3. The Migrated Poison — Automatic Rebalancing Spreads a Bad Tenant#

An automatic rebalancer moves tenants off cells whose CPU exceeds 80%.

t=0       Tenant 9120's new integration sends requests that trigger a pathological parse.
t=+5min   Cell 04 CPU at 95%; rebalancer picks the hottest tenant, 9120, moves it to cell 11.
t=+12min  Cell 11 CPU at 95%; rebalancer moves 9120 to cell 15.
t=+20min  Cells 04, 11 and 15 have each been degraded; 15% of customers affected
          instead of 5%. Rebalancer paused manually.

Detection: migration.count_per_tenant > 1 per hour; correlated degradation following one tenant across cells. Blast radius: grew with each move — the opposite of what cells are for. Mitigation: pause rebalancer; quarantine tenant 9120 to an isolated cell; rate-limit its integration. Prevention: never auto-migrate a tenant whose load is anomalous relative to its own history — that's a poison or abuse signal, not growth; cap moves per tenant per day; human approval for moving the tenant that is causing the load; a quarantine cell for suspected poison tenants. Owner: platform owns rebalancer policy; the tenant's account team owns the customer conversation.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Bad deploy in a cellCell error rate vs undeployed cellsOne wave's cellsAuto-rollback; halt pipelineChange owner + platform
Router faultRouter latency; all cells idleAll customersRevert; group-wise router deploysPlatform
Shared dependency downCorrelated errors in every cellAll cells using itLocal validation, cached config, static fallbacksDependency owner
Control plane downMap version stale; placements failingNew signups and migrations onlyLast-known-good map; alert after 5 minPlatform
Cell over capcell.load_pct_of_cap > 70%Cell's tenants at riskMigrate largest tenants; open new cellPlatform capacity
Whale outgrows cellTenant share of cell > 25%Neighbours in that cellDedicated cell, scheduled migrationPlatform + account team
Poison tenantOne tenant dominating CPU or errorsOne cell, more if auto-movedQuarantine cell; no automatic movesPlatform + tenant owner
Cell infrastructure faultCell SLO burn rate > 10×One cellEvacuate and rebuild from templatePlatform

Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

A Staff engineer turns one service into cells and makes the router thin. A Principal engineer notices that cells only bound blast radius if the whole request path is cellular: a cellular API that calls a non-cellular messaging service, a global feature-flag system and a shared identity introspection endpoint is still one bad change away from a global outage. Cells are therefore an org-level decision about failure posture — which services must be cellular, what the standard cell is, who runs the router and directory, how cells line up across services so that cell 07 of the API talks to cell 07 of the messaging service — and about money, because cells raise fixed cost by a meaningful percentage at every scale. At L7 the work is making cells the paved road for tier-0 and tier-1 services, aligning cell boundaries across teams, keeping a public inventory of the global services that remain, and pricing the overhead against the customer-minutes of outage it removes.

🧭 Principal Move: "I'd make the cell the unit of capacity, deployment and failure for every tier-1 service, aligned on the same tenant-to-cell map, with one platform team owning the router, directory and migration tooling. And I'd publish the list of global services with tighter error budgets — because those are now where our worst days come from."

The Org-Level Fault Line#

One cell platform with aligned cells across services vs each team cell-ifying on its own vs staying with zonal isolation only.

OptionWhat WorksWhat BreaksWho Pays
Each team builds its own cellsFits each service; no central dependencyDifferent partition keys and cell maps; cross-service calls fan out across cells; N routers and N migration toolsCustomers when misaligned cells turn one failure into many partial ones; every team's on-call
Zonal isolation onlyCheap; matches infrastructure; fast drainsSoftware and tenant faults stay global; blast radius ~33% at bestCustomers on bad-deploy days
Platform cells, aligned mapOne router, directory and migration tool; services share the cell assignment so a tenant lives in the same cell everywherePlatform must support diverse stacks; alignment constrains teams' sharding choicesPlatform headcount (4–8 engineers); teams' migration effort

The Principal default is the third row: one tenant-to-cell map that every tier-1 service consumes, so a cell is a complete vertical slice of the product, not a slice of one service.

Cost Model#

Assumptions: fully loaded engineer ~$25K/month; fixed per-cell stateful overhead (database primary plus replicas, cache, queue, minimum compute, monitoring) ~$6–12K/month; variable cost proportional to load is unchanged by cells.

ScaleCellsMachineryExtra Infra vs One StackPeople / On-callRough Monthly Extra
Early4 cells, 1 regionRouter on the gateway, directory table, manual migration script, wave deploys~4 × $6K, partly offset by existing shards ≈ +$15K1 FTE platform~$15K + ~$25K people
Growth20 cells, 2 regionsRouter fleet, directory service, automated migration, per-cell dashboards, canary cell~40 × $9K minus shared capacity ≈ +$150K4 FTE platform; cell-aware on-call~$150K + ~$100K people
Large150 cells, 6 regions, 10 services alignedPlatform cells, cell template as code, evacuation and rebuild automation, global-service error budgets≈ +$900K (overhead shrinks as a share of a much larger bill)8 FTE platform~$900K + ~$200K people

The Principal observation: at growth stage the cell overhead might be 20–30% of the serving bill, which is a real number to defend. The defence is customer-minutes: one global 35-minute outage across 9,000 organizations is ~315,000 org-minutes; the same incident in a 5% cell is ~16,000. If a global outage costs SLA credits, churn risk and a week of executive attention, the overhead pays for itself in one or two avoided incidents a year.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

The Year 3 trigger is the sign the architecture is working: once cells contain deploys and tenants, the remaining global outages all come from the shared services, and the program shifts to shrinking that list.

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Partition key (tenant, user, region)One-wayChanging it re-homes every record in every cell
Directory routing vs pure hashOne-way-ishMoving from hash to directory is a backfill; the reverse loses per-tenant placement
Cell size capTwo-wayNew cells at a new size; old cells migrate gradually
Exposing cell identity to customers or in IDsOne-wayCustomers and partners hard-code it; migrations break them
Aligning cells across servicesOne-way-ishRe-aligning later needs coordinated migrations across teams
Wave plan and bake timesTwo-wayPipeline config
Which services remain globalTwo-way-ishCell-ifying a global service later is a project, not a config change

The Standard I'd Write#

RFC: Cell-Based Isolation for Tier-1 Services (v1)

Scope: Every tier-0 and tier-1 service in the customer request path.

MUST:

  1. Services MUST partition serving infrastructure into cells using the platform tenant-to-cell map, with a published per-cell cap load-tested at least quarterly.
  2. The request path MUST contain no synchronous cross-cell calls. Global dependencies MUST be listed in the shared-service inventory with an owner and a fallback the cell can run on for at least 4 hours.
  3. Code, configuration, feature flags and schema changes MUST deploy through cell waves starting with the canary cell, with automatic rollback on cell-relative SLO degradation.
  4. No customer-visible identifier, endpoint or document MAY reveal a tenant's cell.
  5. Each cell MUST run at no more than 70% of its tested cap, and tenant migration MUST be rehearsed weekly.

SHOULD: Keep any single pooled tenant under 25% of a cell; rebuild each cell from its template at least yearly; keep the router free of business logic.

Exceptions: Filed with the platform team, time-boxed, with a named global-risk owner.

Success metrics: share of incidents affecting more than one cell; customer-minutes lost per quarter; p99 migration write pause; time to evacuate a cell.

What I'd Tell the VP#

Today, when something goes wrong in our product — a bad update, one customer's unusual traffic — it tends to go wrong for every customer at once, which is why our last three incidents were total outages. I'm proposing we split the product into about twenty independent copies, each serving a fixed group of customers, and roll out every change to one copy first. The worst realistic incident then affects about one customer in twenty, and most bad changes are caught before they reach anyone outside our own test accounts. It adds roughly a fifth to our infrastructure bill at our current size and needs a small platform team. One avoided global outage a year covers the cost in SLA credits and customer trust alone.

Principal Interview Signals#

SignalWhat It Sounds Like
Blast radius as a business number"Five percent of customers for an hour is survivable; that's twenty cells."
Whole-path thinking"A cellular API on a global flag service is still a global outage."
Prices the overhead"Cells add about 25% to serving cost; one avoided global incident a year pays it."
Aligns across teams"One tenant-to-cell map, so a cell is a vertical slice of the product."
Shrinks the shared list"Our remaining global outages come from five shared services. That list is the roadmap."

Staff answers that L7 interviewers find insufficient:

  • "We'll cell-ify the API." — Silent on the dependencies the API calls, which decide the real blast radius.
  • "Cells are independent." — No inventory, owners or error budgets for what they still share.
  • "We'll add cells as we grow." — No cost model, no cap discipline, no alignment across services.

How Real Companies Built It#

AWS: Cell-Based Architecture as Well-Architected Guidance#

AWS publishes guidance on reducing the scope of impact with cell-based architecture, extending the bulkhead idea from Availability Zones and Regions into the workload itself. A cell is a complete, independent instance of the workload that does not share state with other cells and handles a subset of requests, selected by a partition key aligned with the service's natural grain — customer ID or resource ID. A cell router, described as the thinnest possible layer, routes requests to the right cell and only that; a control plane provisions and de-provisions cells and migrates customers. The guidance frames the payoff simply — with 10 cells serving 100 requests, a failure in one cell leaves 90% of requests unaffected — and names bad deployments and poison-pill requests as failures cells contain (What is a cell-based architecture?). On sizing, it recommends capping the maximum cell size, keeping sizes consistent across installations, and weighing three forces: big enough for the largest workloads, small enough to test at full scale, and big enough for economies of scale (Cell sizing).

Staff insight: "Small enough to test at full scale" is the line to quote. A cell's real value as a scale unit is that its maximum is known and proven; a system that grows by adding identical, tested cells never meets an untested scaling cliff.

Slack: Zonal Cells and a Five-Minute Drain#

After a June 2021 network disruption in one US-East-1 availability zone caused slowness and degraded connections between Slack's servers — a gray failure, not a clean outage — Slack spent about 1.5 years moving its most critical user-facing services to a cellular architecture in which each availability zone is a cell. Services are siloed so they receive traffic only from within their zone and send traffic upstream only to servers in the same zone. The goal is to remove as much traffic as possible from a zone within 5 minutes: a signal goes through Slack's control plane to the edge Envoy load balancers, which reweight per-zone target clusters, allowing gradual drains with 1% granularity and propagating in seconds (Slack's Migration to a Cellular Architecture).

Staff insight: This is the zonal-cell end of the spectrum, chosen for the failure Slack actually had: infrastructure gray failure. Note the mechanism — drain by reweighting at the edge, in fine increments — which is the same "thin router, decisions elsewhere" discipline applied to zones.

Roblox: Cells of ~1,400 Machines After a 73-Hour Outage#

Roblox's October 2021 outage lasted 73 hours; it began with an issue in one component in one data center and spread. In response, Roblox built cells — sets of about 1,400 machines — as blast walls inside each data center, replicating services within and across cells. A failing cell can be deactivated while replication across cells keeps the service running, and cells can be fully reprovisioned; the definition of everything in a cell lives in source control so it can be rebuilt from scratch with automated tools, and the deployment tool automatically stripes a service across cells so owners don't manage replication themselves. Roblox reported that at peak more than 70% of back-end service traffic was served from cells, with close to 30,000 machines managed by cells (Making Roblox's infrastructure more efficient and resilient).

Staff insight: "Rebuild the cell from its definition" is the operational payoff people forget. Once a cell is code, evacuating and rebuilding it is routine maintenance rather than surgery — and that is what makes a sick cell a minor incident instead of a long one.


Practice Drill#

Prompt: "We run a project-management SaaS: 12,000 customer workspaces, one region, a Kubernetes fleet of stateless API pods, a sharded PostgreSQL cluster keyed by workspace ID, Redis, Kafka and a global config service. In six months we've had four full outages: a config push, a bad deploy with a memory leak, a customer whose API usage triggered a pathological query on every pod, and a Redis failover. Our biggest customer is 8% of traffic. Leadership asks for 'cells.' Design the move."

Staff Answer

All four outages were total because every pod ran the same code and config and saw every workspace's traffic. The fix is to bound what any one change, request or component failure can reach. First the business number: product and support agree that losing 5% of workspaces for an hour is survivable, so I want about 20 cells. Partition key is workspace ID — it's already the shard key and almost every request carries it. Each cell is a full stack: its own API deployment, worker pool, PostgreSQL primary and replicas, Redis and Kafka topics — about 600 workspaces per cell, capped at 2× average load and load-tested at that cap. The biggest customer at 8% of traffic would be about 80% of a pooled cell's cap (each cell averages 5% of traffic, capped at 10%), so it gets a dedicated cell on the same template, as do the next few largest; nobody pooled exceeds 25% of a cell. Routing: a tenant directory in the control plane, replicated as a versioned map to a thin router in our gateway — workspace ID from the signed token, lookup, forward. On control-plane failure it routes on last-known-good; only new signups and migrations pause. Migration: copy the workspace's rows, catch up via CDC filtered by workspace, pause writes for 10–60 seconds behind a 503 with Retry-After, checksum, flip the directory, keep the source read-only for 7 days. We rehearse it weekly on synthetic workspaces. Deploys: an internal canary cell with our own workspaces, then 1, 4, 5 and the remaining cells with 30–60 minutes of bake, rolling back automatically if a deployed cell's error rate exceeds 2× the undeployed cells. The config service moves into the same waves — no global push. Shared-dependency inventory: identity becomes local token validation with cached keys; config is per-cell with last-known-good; billing and search consume cell event streams asynchronously; the router and directory are the accepted global risk, owned by platform with a tighter error budget and group-wise deploys. Rollout: introduce the directory first with every workspace mapped to "cell 0," carve out the canary cell, then migrate ~600 workspaces at a time into new cells over two quarters. Each cell runs at ≤ 70% of cap so one cell can be evacuated into the others. Owners: platform owns router, directory, migration and wave pipeline; service teams own behaviour inside cells; account teams own dedicated-cell conversations.

Why this is L6:

  • Derives the cell count from a tolerable blast radius and handles the whale with a dedicated cell rather than oversizing every cell.
  • Keeps the router to lookup-and-forward with a last-known-good map, and puts config and flags through the same cell waves that caught none of the four incidents before.
  • Inventories shared dependencies and assigns each a fate, and makes migration a rehearsed routine with a bounded write pause and headroom for evacuation.

What L7 adds:

  • Prices the change — roughly 20 cells of fixed stateful overhead, perhaps 20–30% more serving cost — against four global outages' customer-minutes, SLA credits and churn risk.
  • Insists the tenant-to-cell map is shared by every tier-1 service so cells are vertical slices of the product, and sets an org rule against global push paths.
  • Publishes the remaining global services with error budgets as the next reliability roadmap.
❌ Common L5 Trap

"Split the fleet into cells by hashing workspace ID across 4 large Kubernetes clusters, each with its own database shards. Put a smart gateway in front that hashes the ID, applies per-plan rate limits and checks the config service for overrides. Do rolling deploys to each cluster in turn."

Why this misses: Four large cells still mean every incident hits 25% of customers, and the 8% customer dominates whichever cell it hashes to. The gateway is now the monolith — rate limits, overrides and a synchronous call to the config service sit in front of every cell, so a config outage or bad rule is still global. Hash routing can't move a single hot workspace, there's no migration path, and nothing says what happens to the shared Redis, config and identity paths that caused the incidents.


Staff Interview Application#

How to Introduce This Pattern#

"I want to start from how much of the customer base we can afford to lose at once — that sets the number of cells. Each cell is a full copy of the stack for a fixed set of tenants, capped and load-tested. A very thin router maps tenant to cell from a cached directory. Every change goes to a canary cell first, then waves. Then I'll go through everything the cells still share, because that's where the real blast radius is."

Lead with the tolerable blast radius, then partition key and cell size, then the router, then deploy waves and migration, then the shared-dependency inventory and owners.

When NOT to Use This Pattern#

  • Small systems: a few hundred tenants on a handful of servers get more from staged deploys and graceful degradation than from a router, directory and migration tool.
  • No partition key: if most requests span many users or tenants (social graphs, global matching), cells force cross-cell calls that are worse than the shared risk.
  • The problem is noisy neighbours, not total outages: per-tenant quotas and shuffle sharding in multi-tenancy are cheaper and target the actual failure.
  • One hot entity: a single viral key is a hot keys problem; moving it between cells just moves the fire.
  • Global invariants on the hot path: uniqueness across all tenants or a global counter can't live in a cell; keep a small hardened global service and name it as the shared risk.

Follow-Up Questions to Anticipate#

Interviewer AsksWhat They Are TestingHow to Respond
"How many cells, and how big?"Blast-radius reasoning"From the tolerable blast radius — 5% means about 20 — then a cap the largest pooled tenant fits under at 25%, load-tested at full size."
"What if the router fails?"Router discipline"It's lookup-and-forward on a cached map, deployed in groups. If the directory is down it routes on last-known-good for hours."
"How do you move a customer between cells?"Migration mechanics"Copy, catch up with CDC, 10–60 seconds of write pause, verify, flip the directory. Source stays read-only for a week; rehearsed weekly."
"What still takes down all cells?"Shared-dependency inventory"The router, identity issuance and the deploy system. Each has a fallback or a tighter error budget, and that list is our roadmap."
"How does a deploy work?"Change safety"Canary cell, then waves with bake, comparing deployed cells to undeployed ones and rolling back automatically. Config follows the same waves."
"What about the biggest customer?"Whale handling"Over 25% of a cell gets a dedicated cell on the same template, priced into their contract."

Scorecard#

DimensionSenior (L5)Staff (L6)Principal (L7)
FramingShards data for scaleBlast radius as 1/N from a business tolerance; full-stack cellsCells as the org's failure posture across services
RoutingSmart gateway with hashing and logicThinnest router, cached directory, last-known-goodRouter and directory as tier-0 platform with group-wise deploys
SizingLarge cells for efficiencyCapped, load-tested cells; whales dedicated; ≤ 70% utilization for evacuationStandard cell sizes and aligned maps across services; overhead priced
Change safetyRolling deploysCanary cell and waves; config and flags follow; relative-health rollbackNo global push paths by policy
DependenciesNot consideredInventory with a fate per shared serviceError budgets for global services; shrinking list as roadmap

Strong Hire Signals

SignalWhat It Sounds Like
Starts from tolerance"Five percent for an hour is survivable. Twenty cells."
Keeps the router boring"Tenant to cell from a local map, nothing else."
Knows migration is routine"Copy, catch up, a 30-second write pause, flip. We rehearse it weekly."
Hunts shared paths"Identity introspection would make every cell fail together, so tokens are validated locally."

Lean No-Hire Signals

SignalWhy It Misses the Bar
Cells with a global config pushThe most common global outage path is still open
No migration storyWhales and hot cells have no remedy
Router with business logic and synchronous callsRebuilds the monolith in front of the cells

Common False Positives: Drawing many boxes labelled "cell" ≠ isolation, if they share config and identity. Knowing Kubernetes multi-cluster tooling ≠ having a routing and migration design. "We use multiple availability zones" ≠ cells — zones don't contain software faults.


Capacity Planning Quick Reference#

Sizing Cells#

min_cells               = 1 / tolerable_blast_radius                 # 5% => 20
cell_cap_load           = (total_peak / cells) × 2                   # headroom for growth and skew
max_pooled_tenant       = 0.25 × cell_cap_load                       # above => dedicated cell
target_utilization      ≤ 0.70 × cell_cap_load                       # leaves room to evacuate one cell
evacuation_fit          = (cells - 1) × (cap - current) >= one cell's load
directory_size          = tenants × ~50 bytes                        # 12K => 600 KB; 10M => 500 MB
fixed_overhead          = cells × per_cell_stateful_cost              # 20 × $9K = $180K/month
customer_minutes_lost   = incident_minutes × customers / cells        # 35 × 9,000 / 20 ≈ 16K
migration_pause         = final_catch_up + drain + verify            # 10–60 s typical
wave_duration           = sum(bake per wave)                         # ~4 h for 5 waves

Key Numbers Worth Memorizing#

NumberContext
1/NBlast radius of a single-cell failure with N equal cells
20 cellsRoughly 5% blast radius per cell
≤ 25%Largest pooled tenant as a share of one cell's cap
≤ 70%Target cell utilization, so one cell can be evacuated into the rest
10–60 sTypical write pause during a tenant migration
30–60 minBake time per deploy wave
~5 minSlack's target to drain traffic from an availability zone
~1,400 machinesRoblox's reported cell size
10 cells, 90% unaffectedAWS's illustration of cell blast radius
20–30%Rough extra serving cost of cells at growth stage (illustrative)

Common Pitfalls Checklist#

  • Cell count derived from a tolerable blast radius agreed with the business
  • Every cell has a published cap and a load test at that cap
  • Router does lookup-and-forward only, from a cached, versioned map, with last-known-good
  • Router code and map changes deploy in groups, canary first
  • Code, config, flags and schema changes all flow through cell waves
  • Automatic rollback compares deployed cells against undeployed cells
  • Shared-dependency inventory with a fate and owner for each entry
  • Tenant migration rehearsed weekly; no automatic moves of anomalous tenants
  • Cells run at ≤ 70% of cap; whales on dedicated cells
  • No cell identity in customer-visible IDs or endpoints
  1. Loading the index…