Technologies that implement this pattern: PostgreSQL · Redis · Kubernetes · API Gateway · Apache Kafka · DynamoDB
Go deeper:
- A shared log index is a sharp multi-tenancy example: one team's dynamic fields can reject everyone's writes, as Log Aggregation explains.
- Shared IP reputation in Email Delivery Service is a noisy-neighbor resource that must be isolated by each tenant's track record, not just by throughput.
- Sizing cells, keeping the cell router thin, migrating tenants and deploying cell by cell are covered in depth in Cell-Based Architecture.
Why This Matters#
Every shared system is one customer away from an outage it didn't cause. The bulk export that pins the shared database at 100% CPU, the integration that retries in a tight loop, the tenant whose working set evicts everyone else's cache, the customer who uploads a 40 GB file into a queue sized for 4 MB messages — none of them is a bug in your code, and all of them become incidents for the other 4,999 tenants on the same hardware. Tenant traffic follows a power law: in most B2B products the top 1% of tenants generate 30–50% of the load, and the biggest one is 10–100× the median.
Most candidates treat multi-tenancy as a schema question: "shared table with a tenant_id column, or a database per tenant?" Staff engineers treat it as a blast-radius and fairness question. The design question is not where does the tenant's data live but when one tenant misbehaves — by accident, by success or by malice — how many other tenants notice, which resource do they notice it on, and who decided that was acceptable? The storage layout is one answer among several; quotas, scheduling, cells and placement are the others, and you need all of them because tenants contend on CPU, connections, cache, queues and locks, not just rows.
The second reframe: isolation is a product tier, not an architecture. Pooled, shared-everything infrastructure is the cheapest per tenant and the noisiest. Dedicated infrastructure is quiet and costs 5–20× more per tenant. Real systems run both at once and move tenants between them as they grow. The skill is deciding who gets which level, what triggers a move, and who pays for it — usually the tenant, through their plan.
If you can walk an interviewer from "what does a tenant share" to "how is each shared resource divided" to "how big is the blast radius when division fails" to "how does a tenant graduate to its own capacity," you are answering at Staff level.
The 60-Second Version#
- Assume a power law, and size for the biggest tenant, not the average. Top 1% of tenants ≈ 30–50% of load is common. If the largest tenant can exceed one shard's ceiling (e.g. ~10K writes/s on one PostgreSQL primary), the tenant needs to be split or given its own capacity before day one of its contract.
- Every shared resource needs a per-tenant limit — not just the API. Rate limits at the gateway don't stop one tenant's query from holding 40 of 100 database connections. Budget per tenant at each layer: requests/s, concurrent requests, connections, queue share, cache share, storage.
tenant_idgoes in every key, every row, every log line and every metric. Without per-tenant attribution you cannot find the noisy neighbour in under 5 minutes, bill for usage, or move a tenant. Cardinality matters: emit per-tenant metrics for the top ~100 tenants and roll the rest into "other."- Cells cap blast radius at a known fraction. 20 cells of ~250 tenants each means a cell-wide failure touches 5% of tenants. That fraction, not the uptime of any component, is the number to bring.
- Shuffle sharding beats plain sharding for noisy neighbours. With 100 nodes and 5 per tenant, there are ~75M possible node combinations; two tenants rarely share more than 1–2 nodes, so one tenant's poison request degrades a small slice of everyone else's capacity instead of a whole shard.
- Isolation tiers cost real money. Pooled tenant: cents to a few dollars a month of infrastructure. Dedicated database: ~$500–2,000/month minimum. A dedicated cell: $10K+/month. Price the tier to the customer, and make the move between tiers a routine, tested operation.
The Problem#
A B2B analytics product serves 5,000 companies from one PostgreSQL cluster, one Redis cache, one Kafka topic for ingestion and one pool of query workers. Tenant 4417 — a new enterprise customer — runs an unfiltered 18-month report at 9:00 on Monday. Their query plan scans 400M rows, holds 30 of the pool's 100 connections, and pushes every other tenant's dashboard p99 from 300ms to 6 seconds. The same morning a small tenant's broken integration retries a failing ingest call 2,000 times a second, filling the shared ingestion partition so every tenant's data arrives 20 minutes late. Nobody can tell which tenant is responsible for 15 minutes, because the metrics aggregate across tenants. And when the enterprise customer's security team asks how their data is separated from everyone else's, the honest answer is "a WHERE clause in application code." The job is to divide every shared resource so that one tenant's worst day is bounded, to know within minutes which tenant is the cause, to make cross-tenant data access structurally impossible rather than conventionally avoided, and to give large tenants a path to more isolation that doesn't require a rewrite.
Case Studies That Use This Pattern#
- Rate Limiter — Per-tenant quotas at the edge: the first, coarsest layer of noisy-neighbour protection
- Code Runner — Running untrusted tenant code on shared machines; isolation is a security boundary, not just a fairness one
- Metrics Platform — One tenant's cardinality explosion can take down ingestion and queries for all
- Search Engine — Index-per-tenant vs shared indices with tenant routing; one tenant's heavy query starving a shard
- Message Broker — Shared partitions and per-client quotas; head-of-line blocking across tenants
- Notifications — One tenant's campaign burst delaying everyone's password resets
- Sharded Database — Tenant ID as the shard key, and what happens when one tenant outgrows a shard
- Multi-Region — Data residency puts specific tenants in specific regions; tenant placement becomes a compliance decision
- Distributed Cache — Shared caches where one tenant's working set evicts everyone else's hot keys
Which Problem Are We Solving?#
"Make it multi-tenant" hides four goals that lead to different designs. Name them and commit before drawing a box.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| B2B SaaS with contracts (CRM, analytics, collaboration) | Hundreds to tens of thousands of tenants; a few whales; enterprise buyers ask about isolation and residency | Pooled by default, tenant-keyed sharding, cells; dedicated tier for whales and regulated tenants | Whale saturates a shared shard; cross-tenant leak via missing filter | Per-tenant SLO; zero cross-tenant reads; documented isolation per tier |
| B2C or long-tail platform (millions of small accounts, free tier) | Per-tenant cost must be cents; abuse is common | Shared everything, strict per-account quotas, aggressive abuse detection | Free-tier abuse consumes capacity paid tenants need | Paid tenants protected from free-tier load; cost per tenant bounded |
| Internal platform (teams as tenants of a shared Kafka, database or compute platform) | Tenants are other teams; chargeback and fairness matter more than security | Quotas per team, priority classes, usage attribution for chargeback | One team's batch job starves another's serving path | Every team gets its reserved share; usage attributable to a cost centre |
| Untrusted code or data (CI, serverless, notebooks) | Tenants actively hostile; escape is a security incident | Hard isolation: microVMs or sandboxes, per-tenant network policy, no shared kernel for high-risk work | Sandbox escape; side channels; resource exhaustion on the host | No tenant can read or affect another's execution |
🎯 Staff Move: "I'll treat this as intent one: a B2B product with a long tail and a handful of enterprise whales. So the default is pooled infrastructure with tenant-keyed sharding and per-tenant limits at every shared layer, organized into cells so a bad day touches 5% of tenants. The top tenants and anyone with a residency clause get a dedicated tier, priced into their contract."
The Core Tradeoff#
| Strategy | What Works | What Breaks | Who Pays |
|---|---|---|---|
Pool: shared tables with tenant_id | Cheapest per tenant; one schema to migrate; easy onboarding (a row insert) | Noisy neighbours on every shared resource; isolation depends on every query being correct; one big tenant can dominate a shard | Small tenants during a whale's peak; security team on the day a filter is missed |
| Pool + row-level security | Database enforces the tenant filter even when application code forgets | Policy overhead (~1–5% on simple queries, more on complex plans); superuser and owner roles bypass unless forced; still shares capacity | DBAs debugging plans; platform team owning connection-level tenant context |
| Schema per tenant | Logical separation; per-tenant backup and restore; easier to explain to auditors | Catalog bloat beyond ~5–10K schemas; migrations run N times; connection pooling gets awkward | The team running migrations across thousands of schemas |
| Database per tenant (silo) | Strong isolation; per-tenant tuning, maintenance windows and residency | 5–20× cost per tenant; fleet management for thousands of databases; cross-tenant analytics is an ETL problem | The business (infrastructure cost) and the platform team (fleet operations) |
| Cells (N independent stacks, tenants assigned to one) | Blast radius = 1/N of tenants; deploys and incidents contained | Router becomes critical; cross-cell features (global search, shared users) are hard; uneven cell fill | Platform team owning routing and tenant moves |
| Shuffle sharding | One tenant's poison traffic hits only its own random subset of nodes; overlap with any other tenant is small | Needs stateless or replicated backends; harder to reason about capacity per node | Capacity planners; debugging is per-combination |
| Dedicated capacity for whales | Whale can't hurt the pool, pool can't hurt whale; price matches cost | Two operational paths; moving a tenant between tiers is a migration | The whale (via contract) and the team that owns tenant moves |
Staff Default Position#
Pool by default, put tenant_id in every key and every metric, enforce a per-tenant budget at every shared resource, cap blast radius with cells, and give tenants a tested path to dedicated capacity — isolation is a tier you sell, not a rewrite you do under pressure.
The default stack: tenant identity resolved once at the edge and propagated through every hop as request context; data keyed by tenant_id (shard key and leading index column) with database-enforced row-level security as a second line behind application filters; per-tenant token buckets at the gateway, per-tenant concurrency caps in services, weighted fair queuing for background work, and per-tenant connection and statement-timeout limits at the database; per-tenant metrics for the top ~100 tenants by usage; tenants grouped into cells of a size that makes one cell's failure a tolerable fraction (5–10%) of the customer base; and a tenant-move tool that can relocate a tenant between cells or into a dedicated cell with minutes of write pause, rehearsed monthly.
When to Deviate#
- Regulated or contractual isolation — healthcare, government, or a customer contract that names dedicated infrastructure or a region. Silo them from day one; retrofitting a tenant out of a pool under audit pressure is the most expensive version of this work.
- Few, large tenants — 20 enterprise customers, each 5% of revenue. Database or cell per tenant is affordable and makes every conversation with their security teams shorter. Pooling saves little when there is no long tail.
- Untrusted execution — tenant-supplied code, plugins or queries that can run arbitrary computation. Fairness controls aren't enough; you need a hard sandbox per tenant (microVM, gVisor-style kernel isolation) and per-tenant network policy.
- Very early product, single shared cluster — below ~100 tenants and ~500 req/s, a gateway quota and
tenant_iddiscipline are enough. Build cells when the largest outage you can tolerate is smaller than "everyone," or when the biggest tenant approaches a shard's ceiling.
One Question, Three Levels#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Shared tables with a tenant_id column, or a database per tenant?" | "What does a tenant share, how is each shared resource divided, and how many tenants does one bad tenant hurt?" | "What isolation tiers do we sell, what does each cost us, and who owns moving tenants between them?" |
| Noisy neighbour | Rate limit per tenant at the API gateway | Per-tenant budgets at every layer: requests, concurrency, connections, queue share, cache; plus cells | Makes per-tenant budgets a platform default so 40 services don't each forget the database layer |
| Data isolation | WHERE tenant_id = ? in the ORM base query | App filter plus database-enforced row-level security; tenant context set per connection; tested with cross-tenant probes | Sets isolation guarantees per tier in writing, signed off by security and sales; audits them |
| Large tenants | "They'll get a bigger instance" | Detects tenants approaching a shard's ceiling; moves them to dedicated capacity with a tested tool | Prices tiers so whale isolation pays for itself; contracts specify what dedicated means |
| Observability | Global p99 and error rate | Per-tenant latency and usage for top tenants; find the noisy neighbour in < 5 minutes | Per-tenant cost attribution feeding pricing and capacity planning |
| Ownership | Each service team handles its own tenants | Platform owns tenant context, quotas and placement; services own per-tenant limits on their resources | Defines the platform/product contract and the tenant lifecycle (onboard, grow, move, offboard, delete) |
Why "First move" separates levels
The L5 question — shared tables or separate databases — is a real decision, and answering it well still gets downleveled because it addresses only one shared resource. The noisy neighbour usually shows up on something else: connection pools, a shared queue partition, cache memory, a worker pool. The Staff candidate inventories what tenants share and assigns each a division mechanism. The Principal candidate notices that isolation is something the company sells, and that the architecture has to support moving a tenant between tiers as a routine operation.
Why "Noisy neighbour" separates levels
A gateway rate limit counts requests. It cannot see that one tenant's 50 requests/s are each a 30-second report holding a database connection, while another tenant's 2,000 requests/s are 5ms point reads. Staff answers budget the expensive dimension at the layer where it is spent: concurrency at the service, connections and statement timeouts at the database, share-of-workers in the job queue. Fairness by request count is the L5 answer; fairness by resource consumed is the Staff answer.
Why "Data isolation" separates levels
"Every query has a tenant filter" is a convention, and conventions fail at the 300th query written by the 40th engineer. A cross-tenant data leak is a security incident with notification obligations, not a bug. Staff engineers make the database enforce the filter as a second line — row-level security keyed on a per-connection tenant setting — and add an automated probe that tries to read tenant A's data as tenant B on every deploy.
Where the Design Splits#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Pool vs Silo | Cheap, shared and noisy vs isolated, expensive and operationally heavy |
| 2 | Isolation by Quota vs Isolation by Placement | Software limits on shared capacity vs separate capacity that can't be contended |
| 3 | Static Placement vs Live Rebalancing | Tenants pinned to a shard or cell forever vs the ability to move them as they grow |
| 4 | Application-Enforced vs Database-Enforced Tenancy | Filters in code (flexible, forgettable) vs policies in the store (structural, with overhead and bypass rules) |
| 5 | Equal Share vs Paid Priority | Every tenant gets the same slice vs higher tiers get more of the shared resource under contention |
Fault Line 1: Pool vs Silo#
Pooled tenants share tables, processes and machines; siloed tenants get their own database, cluster or cell. Most real systems are a bridge: pooled compute with siloed data, or pooled everything with silos for the top tier. Who pays: pooling makes small tenants pay with noisy-neighbour exposure and makes the security team pay with a weaker isolation story; silos make the business pay 5–20× per tenant in infrastructure and the platform team pay in fleet management — 3,000 databases means 3,000 upgrades, backups and connection pools. Staff default: pool the long tail, silo the top tier and regulated tenants, and use the same code and schema for both so a tenant can move between them. Deviate when: there are fewer than ~50 tenants and each is large — then silo everything; the fleet is small enough to manage and every customer's security review gets easier.
Fault Line 2: Isolation by Quota vs Isolation by Placement#
Quotas divide shared capacity in software: token buckets, concurrency caps, weighted fair queues, cgroups. Placement divides capacity physically: this tenant's requests only run on these nodes. Who pays: quotas pay in imperfection — they bound the rate of consumption, but a tenant within quota can still trigger a pathological query, a lock convoy or a cache-eviction storm; placement pays in utilization, because reserved capacity sits idle when its tenant is quiet (dedicated capacity commonly runs at 15–30% utilization versus 50–70% for pooled). Staff default: quotas everywhere as the first line, placement (shuffle shards or cells) as the second, so that the failures quotas can't stop — poison requests, crash-inducing inputs — are contained to a subset. Deviate when: the failure you fear is a security boundary, not a fairness one — then placement (separate VMs, separate keys) is mandatory and quotas are a nicety.
Fault Line 3: Static Placement vs Live Rebalancing#
Hashing tenant_id to a shard is simple until one tenant grows to 30% of a shard, or a cell fills unevenly. A directory (tenant → cell/shard lookup) costs one cached lookup per request and buys the ability to move a tenant. Who pays: static hashing pays at the worst moment — the day a whale outgrows its shard you have no lever; a directory pays in a critical-path dependency (cache it with a 30–60s TTL and serve stale on failure) and in building a tenant-move tool, which is a migration with a short write pause. Staff default: directory-based placement from the start, even if every tenant initially maps to cell 1, so the move tool exists before you need it. See Online Schema & Data Migrations for the copy–catch-up–cutover mechanics. Deviate when: tenants are uniformly small and numerous (B2C accounts) — hash and never move individuals; rebalance by splitting shards instead.
Fault Line 4: Application-Enforced vs Database-Enforced Tenancy#
Application enforcement means every query includes tenant_id, typically via an ORM default scope. Database enforcement means a row-level security policy compares each row's tenant_id to a session variable set when the connection is checked out. Who pays: application-only pays when one raw query, report or admin tool forgets the filter — a breach rather than a bug; database enforcement pays in per-query policy evaluation, in subtle plan changes, in the need to set and reset tenant context on pooled connections, and in remembering that table owners and BYPASSRLS roles skip policies unless forced. Staff default: both — the application filter for index-friendly plans, RLS as the backstop, and a CI probe that queries as tenant B for tenant A's known rows and must get zero. Deviate when: tenants are siloed in separate databases with separate credentials; then the connection string is the boundary.
Fault Line 5: Equal Share vs Paid Priority#
Under contention, does a free-tier tenant get the same slice of workers as an enterprise tenant? Who pays: equal share makes paying customers subsidize free-tier spikes; strict priority can starve the free tier entirely during a sustained enterprise burst, which becomes a growth and PR problem. Staff default: weighted fair share — enterprise weight 10, pro 3, free 1 — with a guaranteed minimum for every tier so nobody starves, and hard quotas on top so weights only matter under contention. Deviate when: a tier is explicitly best-effort and product has signed off that it can be paused during incidents — then it's a shedding class, see Graceful Degradation: Fail Open or Closed.
Common Interview Mistakes#
| What Candidates Say | What Interviewers Hear | What Staff Engineers Say |
|---|---|---|
| "We'll add a tenant_id column to every table" | "Data layout is the whole answer" | "tenant_id everywhere, yes — and then a budget per tenant on connections, workers, queue share and cache, because that's where neighbours actually collide." |
| "Rate limit each tenant at the gateway" | "Counts requests, not cost" | "Gateway limits stop floods. A 30-second report at 5 requests/s still holds a third of the pool, so the service caps concurrency per tenant and the database caps statement time." |
| "Give every tenant its own database for isolation" | "Hasn't priced 5,000 databases" | "Silo the top 1% and regulated tenants; pool the rest on the same schema so anyone can move up a tier when they pay for it." |
| "The ORM always adds the tenant filter" | "Isolation by convention" | "The ORM filter is for plans. Row-level security is the backstop, and a deploy-time probe proves tenant B can't read tenant A." |
| "We'll shard by tenant_id hash" | "No plan for the whale" | "Directory-based placement, so when one tenant hits 30% of a shard we move it without rehashing everyone." |
| "We'll monitor overall latency" | "Can't find the noisy neighbour" | "Per-tenant p99 and usage for the top 100 tenants, so on-call names the tenant in five minutes, not fifty." |
Quick Reference#
Staff Sentence Templates#
"Tenants share [resources]. Each one gets a per-tenant budget: [requests/s at the gateway], [concurrency in the service], [connections and statement timeout in the database], [share of workers in the queue]. If a tenant exceeds it, [they get 429 / they queue behind themselves], not behind everyone."
"Cells hold about [N] tenants each, so a cell-wide failure touches [1/cells] of customers. The top [K] tenants and anyone with [a residency clause] get a dedicated cell, and that's priced into the [enterprise] tier."
"Isolation of data is enforced twice: [the application filter] for plans, and [row-level security] as the backstop. A probe on every deploy proves tenant B gets zero rows for tenant A."
"When a tenant reaches [30%] of a shard's capacity, the move tool relocates it to [a dedicated cell] with a write pause under [60 seconds]. We rehearse that monthly so it's routine when we need it."
Implementation Deep Dive#
1. Tenant Context and Row-Level Security — PostgreSQL#
Resolve the tenant once at the edge (from the auth token, never from a request parameter), carry it in request context, and set it on the database connection at checkout. The database then refuses to return another tenant's rows even if a query forgets its filter.
-- every tenant-scoped table: tenant_id leads the primary key and every index
CREATE TABLE reports (
tenant_id uuid NOT NULL,
report_id bigint NOT NULL,
...
PRIMARY KEY (tenant_id, report_id)
);
ALTER TABLE reports ENABLE ROW LEVEL SECURITY;
ALTER TABLE reports FORCE ROW LEVEL SECURITY; -- owners don't silently bypass
CREATE POLICY tenant_isolation ON reports
USING (tenant_id = current_setting('app.tenant_id')::uuid)
WITH CHECK (tenant_id = current_setting('app.tenant_id')::uuid);
-- the application role must not have BYPASSRLS; migrations run as a separate role
on connection checkout(ctx):
conn = pool.acquire()
conn.execute("SET app.tenant_id = $1", ctx.tenant_id) # session-scoped
conn.execute("SET statement_timeout = $1", budget_for(ctx.tenant_tier)) # 5s pooled, 60s reports
return conn
on connection release(conn):
conn.execute("RESET app.tenant_id") # never leak context to the next request
pool.release(conn)
# with transaction-mode poolers, use SET LOCAL inside each transaction instead,
# because a session setting can outlive the transaction on a shared server connection
The probe that makes it real: a CI and post-deploy job creates rows for synthetic tenant A, connects as tenant B, and runs every read endpoint and a list of raw queries. Any non-zero result fails the deploy. RLS without a probe is an assumption; with one it is a tested guarantee.
🎯 Staff Insight: Keep the application-level
WHERE tenant_id = ?even with RLS. The policy is a filter the planner adds, and explicit predicates on the leading index column keep plans predictable. RLS is the seatbelt, not the steering.
2. Per-Tenant Budgets at Every Layer — Gateway, Service, Queue#
One limit at the edge is not isolation. Each shared resource needs its own per-tenant budget, in the unit that resource is actually spent in.
# Gateway: request rate, per tenant and tier (token bucket in Redis, local cache for hot tenants)
limits = {free: 20 rps burst 40, pro: 200 rps burst 400, enterprise: contract}
if not bucket(tenant_id).take(1): return 429, Retry-After
# Service: concurrency, per tenant — protects threads and downstream connections
MAX_INFLIGHT_PER_TENANT = max(4, 0.2 × service_concurrency_limit) # no tenant > 20% of the box
if inflight[tenant_id] >= MAX_INFLIGHT_PER_TENANT:
metrics.incr("tenant.concurrency_rejected", tags=[tenant_bucket(tenant_id)])
return 429
inflight[tenant_id] += 1 ... finally inflight[tenant_id] -= 1
# Background jobs: weighted fair queuing instead of one FIFO
queues = per-tenant FIFO queues
weights = {enterprise: 10, pro: 3, free: 1}
worker loop:
t = pick tenant with smallest (virtual_time[t]) # deficit round robin variant
job = queues[t].pop()
run(job)
virtual_time[t] += job.cost / weights[tier(t)] # cost = seconds of worker time
# Database: connection share and statement timeouts per tier (see section 1)
Why it matters: with one FIFO, tenant A's 10,000 jobs put tenant C's 3 jobs behind ~50 minutes of work. With weighted fair queuing, C's jobs start within a few scheduling rounds — seconds — while A still gets most of the pool when nobody else is waiting. Fair queuing only costs anything when there is contention; otherwise it is a FIFO with extra bookkeeping.
3. Shuffle Sharding — Assigning Each Tenant a Random Subset of Nodes#
Plain sharding puts tenants into fixed groups; a tenant whose requests crash or saturate nodes takes down its whole group. Shuffle sharding gives each tenant its own pseudo-random subset of nodes, so the overlap between any two tenants is small.
NODES = 100 # stateless workers behind a tenant-aware router
SHARD_SIZE = 5 # nodes per tenant
function nodes_for(tenant_id):
rng = seeded_prng(hash(tenant_id)) # deterministic: every router agrees
return rng.sample(range(NODES), SHARD_SIZE)
route(request):
candidates = nodes_for(request.tenant_id)
healthy = [n for n in candidates if health[n] is OK]
return least_loaded(healthy) if healthy else reject(503)
# combinations: C(100, 5) = 75,287,520
# a poison tenant that kills all 5 of its nodes leaves every other tenant
# with at least 0-2 of their own 5 nodes affected in almost all cases
| Layout | Tenants Affected When One Tenant Kills Its Nodes | Capacity Lost by Others |
|---|---|---|
| No sharding (all tenants on all 100 nodes) | All | 100% if the poison input crashes every node |
| 20 fixed shards of 5 nodes | Every tenant in that shard (~5%) | 100% for them |
| Shuffle shards of 5 from 100 | Many tenants share 1 node, few share 2, almost none share all 5 | Typically 20–40% of their capacity, and retries land on healthy nodes |
The precondition people skip: shuffle sharding works when backends are stateless or the data is replicated to all candidate nodes. For a stateful sharded database it does not apply directly — use cells there.
4. Cells and the Tenant Directory — Placement You Can Change#
A cell is a complete, independent copy of the stack (services, database shards, caches, queues) serving a subset of tenants. A thin router looks up each tenant's cell in a directory.
directory (strongly consistent store, replicated; cached at routers, TTL 30s):
tenant_id -> { cell: "cell-07", tier: "pooled", state: "ACTIVE" | "MOVING" | "PAUSED" }
router(request):
entry = cache.get(request.tenant_id) or directory.get(request.tenant_id)
if entry.state == "PAUSED": return 503, Retry-After: 5 # brief write pause during a move
forward(request, cell=entry.cell)
move_tenant(tenant_id, from_cell, to_cell):
1. copy: snapshot tenant rows from from_cell -> to_cell (chunked, throttled)
2. catch up: stream tenant's changes via CDC filtered on tenant_id until lag < 2s
3. pause: directory.state = PAUSED; wait for router cache TTL to expire everywhere
4. drain: wait until CDC lag = 0 for this tenant
5. verify: row counts and checksums per table match
6. flip: directory = {cell: to_cell, state: ACTIVE}
7. cleanup: after 7 days, delete tenant rows from from_cell
# typical write pause: 10-60s; rehearse monthly on synthetic tenants
Sizing cells: choose cell count from the blast radius you can tolerate, then check that the largest tenant fits comfortably in one cell (under ~30% of its capacity). 20 cells → 5% blast radius. If the largest tenant would be 60% of a cell, it gets a dedicated cell.
Technique Comparison
| Technique | Stops | Doesn't Stop | Cost | Best For |
|---|---|---|---|---|
| Gateway rate limit per tenant | Request floods, retry loops | Expensive requests within the rate | Low | Every system, first line |
| Per-tenant concurrency cap | One tenant holding threads and connections | Pathological single requests | Low | Request/response services |
| Statement timeout per tier | Runaway queries | Many medium queries | Low | Shared databases |
| Weighted fair queuing | Head-of-line blocking in job queues | Jobs that crash workers | Medium | Background and batch work |
| Row-level security | Missing-filter data leaks | Capacity contention | Low–medium | Pooled data |
| Shuffle sharding | Poison requests, per-tenant overload on stateless tiers | Stateful shard saturation | Medium | Stateless fleets, DNS, gateways |
| Cells | Blast radius of deploys, bad config, stateful failures | Cross-cell features | High | Mature B2B platforms |
| Dedicated capacity | Everything between tenants | Nothing (it costs money) | Highest | Whales, regulated tenants |
Architecture Diagram#
How to narrate it: identity and quotas are enforced before any cell sees the request. The router is deliberately thin — one cached lookup — so it is cheap to make highly available. Each pooled cell is a complete stack with its own per-tenant limits; a bad deploy or a runaway tenant is contained to its cell's ~250 tenants. The dedicated cell runs the same code, which is what makes the move tool possible: moving a tenant is copying rows and flipping a directory entry, not porting an application.
Failure Scenarios#
1. The Monday Report — One Enterprise Tenant Pins the Shared Database#
A pooled analytics cell serves 300 tenants from one PostgreSQL primary with a 200-connection pool. There are gateway rate limits but no per-tenant concurrency caps and a global 5-minute statement timeout.
09:00 Tenant 4417 schedules 60 dashboards to refresh at once; each runs an 18-month scan.
09:01 60 queries x 90s each. They hold 60 connections; CPU 95%; shared buffers churn.
09:02 Other tenants' dashboard p99: 300ms -> 6s. Pool waits climb. Gateway limits not hit:
4417 is sending 1 request/s.
09:05 API timeouts at 10s; clients retry; pool exhausted. Error rate 30% for all 300 tenants.
09:17 On-call finds the cause by sorting pg_stat_activity by duration. No per-tenant metrics.
09:19 Kills 4417's queries. Recovery in 2 minutes. 19 minutes of cell-wide degradation.
Detection: per-tenant db.active_connections share > 20%; db.query_seconds by tenant; per-tenant p99 divergence from cell median; pool wait time.
Blast radius: every tenant in the cell — 300 customers — for 19 minutes.
Mitigation: kill the queries; temporarily cap tenant 4417's concurrency to 4.
Prevention: per-tenant concurrency cap (no tenant > 20% of connections), statement timeout by tier (30s interactive, long reports via an async job queue with weighted fair scheduling), and a read replica or OLAP store for heavy reports (OLAP Databases).
Owner: analytics platform owns per-tenant limits; the account team owns telling 4417 that scheduled reports moved to the async path.
🎯 Staff Insight: The gateway limit never fired because the expensive dimension wasn't requests. Ask, for each shared resource, "what is the unit a tenant spends here?" — for a database it's connection-seconds, not calls.
2. The Missing Filter — Cross-Tenant Data Exposure#
A new "export to CSV" endpoint is written with a raw SQL query for performance, bypassing the ORM's default tenant scope. There is no row-level security.
Day 0 Endpoint ships. Query filters by report_id only; report IDs are sequential integers.
Day 3 A customer's admin edits the report_id in the URL and downloads another company's report.
Day 3 Customer reports it to support. Security incident declared.
Day 4 Logs show 37 cross-tenant exports across 11 tenants over 3 days.
Day 5-30 Breach notification to affected tenants; two enterprise renewals put on hold.
Detection: cross-tenant probe in CI (would have caught it before shipping); audit log alert where resource.tenant_id != request.tenant_id; sequential-ID enumeration patterns in access logs.
Blast radius: 11 tenants' data exposed; contractual and legal exposure across the whole customer base.
Mitigation: disable the endpoint; rotate any exposed credentials in the reports; notify affected tenants.
Prevention: row-level security with FORCE, tenant-scoped composite keys (tenant_id, report_id) so lookups by bare ID are impossible, non-enumerable IDs, and the cross-tenant probe as a deploy gate.
Owner: the platform team owns RLS and the probe; the product team owns the endpoint; security owns the incident and notification.
3. Cache Eviction Storm — One Tenant's Working Set Evicts Everyone#
A shared 64 GB Redis cache with LRU eviction serves all tenants in a cell. Hit rate is normally 96%.
t=0 Tenant 9120 starts a migration that reads 900K distinct records once each.
t=+2min 9120 inserts ~40 GB of one-time cache entries; LRU evicts other tenants' hot keys.
t=+4min Cell hit rate 96% -> 61%. Database read load 10x. p99 for all tenants 80ms -> 1.4s.
t=+15min Migration ends, but hit rate recovers only as keys are refilled: 25 more minutes.
Detection: cache.hit_rate drop > 10 points in 5 min; per-tenant share of cache inserts > 30%; cache.evictions spike.
Blast radius: every tenant in the cell, for ~40 minutes; database nearly saturated.
Mitigation: throttle 9120's batch path; bypass cache for bulk reads.
Prevention: per-tenant cache quotas (key prefix with a byte budget, or separate cache pools per tier); bulk and migration traffic flagged to skip the cache; admission policy (only cache keys read twice).
Owner: cache platform owns quotas and admission; tenant 9120's integration team owns the bulk-read flag.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Expensive-query neighbour | Tenant share of db.active_connections > 20% | Whole cell | Per-tenant concurrency cap, tiered statement timeouts | Service team + DB platform |
| Request flood or retry loop | Tenant gateway.rejected_429 rate; request share | Gateway and downstream | Token bucket per tenant; client retry budgets | Edge platform |
| Cross-tenant data leak | Cross-tenant probe failure; audit tenant_mismatch | Security incident, all tenants' trust | RLS with FORCE, composite keys, deploy gate | Platform + security |
| Cache eviction by one tenant | Tenant share of inserts > 30%; hit-rate drop | Whole cell; DB load | Per-tenant cache budgets; bulk bypass | Cache platform |
| Job queue head-of-line blocking | Per-tenant queue.wait_ms divergence | All tenants on the queue | Weighted fair queuing; per-tenant queue caps | Jobs platform |
| Whale outgrows its shard | Tenant > 30% of shard CPU or storage | That shard's tenants | Move to dedicated cell via directory | Platform team |
| Bad deploy or config | Cell error rate vs other cells | One cell (5%) if cells deploy in waves | Halt wave, roll back cell | Release owner |
| Metric cardinality blow-up from per-tenant labels | Metrics series count growth | Observability stack | Top-N tenants labelled, rest as "other" | Observability team |
Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
A Staff engineer makes one system fair and contained. A Principal engineer sees that multi-tenancy is where engineering, pricing, sales and security meet, and that most of the org's pain comes from those seams: sales promising "dedicated infrastructure" that doesn't exist, a free tier that consumes 40% of capacity and 0% of revenue, enterprise security questionnaires answered differently by three teams, 40 services that each implemented per-tenant limits on the API but none on their database. The L7 work is turning isolation into a small number of named tiers with written guarantees, a cost per tenant per tier that finance can see, a platform that makes the default tier safe without every team re-solving it, and a tenant lifecycle — onboard, grow, move, offboard, delete — that is an operation, not a project.
🧭 Principal Move: "I want three isolation tiers with written guarantees — pooled, pooled-in-a-dedicated-cell, and fully dedicated — each with a cost per tenant that finance signs off on and a move between them that we rehearse monthly. Sales sells from that menu; engineering builds only that menu."
The Org-Level Fault Line#
One tenancy platform that every service inherits vs per-service tenancy.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Every service handles tenants its own way | Fast for the first few services | Inconsistent limits; one service with no database caps sinks the cell; isolation story differs per service | On-call for whichever service is weakest; sales in every security review |
| Separate stack per enterprise customer, built by hand | Each big customer gets what they asked for | Snowflake deployments; upgrades lag months; every customer is a special case | Platform team forever; customers on stale versions |
| Tenancy platform: context, quotas, placement, metering as shared libraries and control plane | Every service gets tenant context, per-tenant limits and cell routing by default; tiers are configuration | Platform team must support every runtime; a platform bug is cross-cutting | Platform headcount (4–6 engineers) |
The Principal default is the third row: mechanism (context propagation, quota enforcement, the directory, the move tool, metering) is centralized; policy (tier weights, per-tier limits, which customers get which tier) is configuration owned by product and sales operations.
Cost Model#
Assumptions: fully loaded engineer ~$25K/month; pooled cost per tenant from shared cluster cost divided by tenant count; a dedicated cell's minimum footprint is ~$8–15K/month (two app nodes per service, a primary plus replica, cache, queue, observability).
| Scale | Tenants | Isolation Machinery | Infra Cost | People / On-call | Rough Monthly Total |
|---|---|---|---|---|---|
| Startup | 300 tenants, 1 cluster | tenant_id everywhere, gateway quotas, statement timeouts | 0.3 FTE; product on-call | ~$8K + ~$8K people | |
| Growth | 5,000 tenants, 8 cells, 6 dedicated | Per-tenant concurrency caps, RLS, weighted fair queues, directory and move tool | Pooled | 2.5 FTE platform; shared rotation | ~$160K + ~$60K people |
| Large | 60,000 tenants, 40 cells, 80 dedicated | Full tenancy platform, shuffle-sharded edge, per-tenant metering, residency regions | Pooled | 6-person platform team; dedicated rotation | ~$1.7M + ~$150K people |
The Principal observation: at scale, the 80 dedicated cells cost more than the 60,000 pooled tenants combined. That is fine only if those customers pay for it. The number to put in front of finance is cost per tenant per tier next to revenue per tenant per tier; if dedicated customers are on the same price as pooled ones, the architecture is subsidizing sales discounts. Equally, a free tier consuming 30–40% of pooled capacity is a pricing decision disguised as an infrastructure bill.
The 3-Year Evolution Path#
The Year 2 trigger is the one teams underestimate: quotas protect tenants from each other, but nothing in Year 1 protects tenants from you. A bad config push or schema change still lands on everyone at once until cells exist and deploys roll through them in waves.
One-Way Doors vs Two-Way Doors#
| Decision | Door Type | Reversibility Cost |
|---|---|---|
tenant_id in every primary key and shard key | One-way | Adding it later means rewriting every table and every query; do it on day one even with one tenant |
| Hash-based vs directory-based tenant placement | One-way-ish | Moving from hash to directory later means building the move tool under pressure with live whales |
| Per-tenant limits and weights | Two-way | Configuration |
| Pool vs silo for a given tenant | Two-way (if same code and schema) | A tenant move with a short write pause |
| Promising "dedicated infrastructure" in a contract | One-way | Contractual; you must support that tier for the contract's life |
| Data residency per region | One-way | Once a tenant's data is legally pinned, moving it needs their consent and a migration |
| Cell boundaries crossing user identity (users in many tenants) | One-way-ish | A global identity service becomes a cross-cell dependency you'll carry forever |
The Standard I'd Write#
RFC: Tenant Isolation Requirements (v1)
Scope: Every service that stores or processes customer data or shares capacity across customers.
MUST:
- Tenant identity MUST be derived from authenticated credentials at the edge and propagated as request context; services MUST NOT accept tenant ID from request parameters.
- Every tenant-scoped table MUST lead its primary key with
tenant_idand MUST have database-enforced row-level isolation (RLS with FORCE, or a per-tenant database).- Every shared resource a service owns (threads, connections, queues, caches) MUST have a per-tenant budget, and no single tenant MAY exceed 20% of a pooled resource without an explicit exception.
- Services MUST emit per-tenant latency, error and usage metrics for the top 100 tenants by usage.
- Every tier-1 service MUST pass a cross-tenant read probe in CI and after each deploy.
SHOULD: Deploy through cells in waves; route through the tenant directory rather than hashing; support the platform tenant-move tool.
Exceptions: Filed with the tenancy platform team, time-boxed to 2 quarters, with security sign-off for any data-isolation exception.
Success metrics: zero cross-tenant data exposures; no incident where one tenant degrades more than one cell; noisy-neighbour root cause identified in < 10 minutes median; cost per tenant per tier published monthly.
What I'd Tell the VP#
Our customers share infrastructure, which keeps costs low, but today one large customer having a busy morning can slow down everyone else, and our answer to "how is my data kept separate?" depends on who you ask. I'm proposing three clearly defined service levels — shared, shared with a protected section, and fully dedicated — each with a guaranteed level of protection and a known cost per customer, so sales can sell them and finance can price them. It needs a small platform team of about five engineers over the next year. In return, one customer's spike stops being everyone's outage, our largest customers get a dedicated option we can actually deliver, and security reviews get a single, true answer.
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Isolation as product | "Isolation is a tier on the price list. Engineering builds three, not one per customer." |
| Cost per tenant | "Dedicated cells cost $12K a month each. If the contract doesn't cover that, the architecture is subsidizing a discount." |
| Blast radius from you, not just tenants | "Quotas protect tenants from each other. Cells and wave deploys protect them from us." |
| Lifecycle thinking | "Onboarding, moving, offboarding and deleting a tenant are operations with runbooks, not projects." |
| One-way door awareness | "tenant_id leads every key on day one. Residency promises are forever." |
Staff answers that L7 interviewers find insufficient:
- "We'll add per-tenant limits to this service." — Correct locally; silent on the other 39 services and the platform that would make it default.
- "Big customers get their own database." — Doesn't price it, doesn't say whether the contract pays for it, doesn't say how they get there.
- "We'll use cells." — Doesn't size them by tolerable blast radius or say how a tenant moves between them.
How Real Companies Built It#
Amazon Route 53: Shuffle Sharding to Contain One Customer's Attack#
The Amazon Builders' Library describes shuffle sharding with a small example — eight workers, each customer assigned two — which yields 28 unique combinations and roughly a 7× reduction in impact compared with plain sharding. Route 53 applies the idea at scale: 2,048 virtual name servers, with every customer domain assigned a shuffle shard of four, for about 730 billion possible combinations. No customer domain shares more than two virtual name servers with any other, so when one domain is targeted by a DDoS, its four name servers spike while other customers' domains don't notice, and the targeted customer can be moved to dedicated attack capacity (Workload isolation using shuffle-sharding).
Staff insight: This is the clearest public example of "a noisy neighbour should degrade a slice, not a shard." The combinatorics are the interview line: overlap between any two tenants is bounded by design, not hoped for.
AWS: Quotas, Token Buckets and Admission Control for Fairness#
The Builders' Library article on fairness in multi-tenant systems describes rate-based quotas per customer, enforced with token-bucket admission control: an admitted request takes a token, an empty bucket means rejection, and bucket capacity allows bursting when the system has spare capacity — with the warning that too large a burst defeats the protection. When one customer spikes, the unplanned portion of that workload is rejected rather than degrading everyone. It covers three ways to enforce quotas across a fleet (dividing locally by server count, routing to dedicated rate-tracking servers by consistent hashing, and asynchronously sharing observed rates), names DynamoDB provisioned throughput, Lambda per-function concurrency and API Gateway usage plans as examples, and returns HTTP 429 for quota rejections while exposing quota usage to customers through metrics (Fairness in multi-tenant systems).
Staff insight: Note that quotas are framed as rejecting the unplanned portion of a workload. That is the Staff framing of a noisy neighbour: the tenant isn't bad, its burst is unplanned, and the system's job is to make that burst the tenant's problem — visibly, with a 429 and a metric — rather than everyone's.
Salesforce: Governor Limits Per Transaction and Per Org#
Salesforce describes its platform as a multitenant, metadata-driven architecture in which each customer's org is a tenant instance on shared resources. To keep tenants well behaved, it imposes governor and execution limits calculated per transaction and per org, so that no single execution of customer code or database transaction can monopolize shared compute — and the limits apply to installed packages and platform code running in the same context, not just the customer's own code (Salesforce architecture basics).
Staff insight: Governor limits are per-tenant budgets on the expensive dimension — queries, DML, CPU time per transaction — enforced inside the runtime, not at the edge. That's the answer to "the gateway didn't fire because the tenant sent one request that did a million things."
Notion: Workspace ID as the Partition Key#
When Notion sharded its PostgreSQL monolith, it chose the workspace ID as the partition key because every block belongs to exactly one workspace, which keeps most queries on a single shard and avoids cross-shard joins. The result was 480 logical shards spread evenly across 32 physical databases — 15 logical shards, each a Postgres schema, per physical database (Sharding Postgres at Notion).
Staff insight: Tenant as shard key is the default for B2B data, and a logical-shard layer between tenants and hosts is what keeps placement changeable: in that layout, rebalancing moves whole logical shards rather than re-keying tenants. The question to raise next is what happens when one workspace outgrows its logical shard — the whale problem this page is about.
Practice Drill#
Prompt: "We run a B2B workflow-automation product: 8,000 customer companies on one Kubernetes cluster, one PostgreSQL primary with two replicas, and one job queue. Our three largest customers are 35% of all job volume. Last month one of them imported 2M records and every customer's automations ran 40 minutes late; separately, a security review flagged that tenant separation is 'enforced only in application code.' Redesign for isolation."
Staff Answer
There are three different problems here and they need three different mechanisms. First, fairness on the job queue: one FIFO means a 2M-record import sits in front of everyone. I'd move to per-tenant queues with weighted fair scheduling across a shared worker pool — weights by tier (enterprise 10, pro 3, free 1), cost measured in worker-seconds — plus a per-tenant cap of 20% of workers so even an enterprise import can't take the whole pool. With 400 workers, the importing tenant gets up to 80 workers, finishes 2M records in a few hours instead of the pool finishing it in one, and a small tenant's job starts within seconds. Imports and other bulk work go to a separate bulk queue with its own worker pool, so interactive automations never share a scheduler with bulk loads. Second, database contention: per-tenant concurrency caps in the API (no tenant above 20% of connections), statement timeouts by path (5s interactive, 120s on the bulk path), and bulk reads served from a replica. Third, data isolation: tenant_id already leads our keys, so I'd add row-level security with FORCE, set the tenant per transaction from the authenticated token, and add a cross-tenant probe to CI and post-deploy that must return zero rows. For blast radius, I'd introduce a tenant directory now — every tenant maps to cell 1 — then split into 8 cells of ~1,000 tenants (12.5% blast radius) over two quarters, deploying in waves. The three largest customers each need about 12% of total capacity; that's too large to share a pooled cell comfortably, so they become dedicated cells on the same code, priced into renewals. Metrics: per-tenant queue.wait_ms p99, tenant share of workers and connections, cross-tenant probe status, cell error rate vs fleet. Owners: platform owns scheduler, directory, RLS and probe; product owns tier weights; account teams own dedicated-tier pricing conversations.
Why this is L6:
- Separates fairness (queue scheduling), contention (connections, timeouts) and data isolation (RLS plus probe) instead of answering all three with "rate limit."
- Budgets each shared resource in the unit it's spent in — worker-seconds, connections — and caps any tenant's share.
- Introduces the directory before cells so placement becomes changeable, and sizes whales out of the pool with numbers.
What L7 adds:
- Turns the outcome into named isolation tiers with written guarantees that sales can sell and security can cite.
- Prices dedicated cells (~$12K/month each) against the three customers' contract value, and flags if renewals don't cover it.
- Proposes the tenancy mechanisms as a platform so the next 20 services inherit per-tenant limits and RLS by default.
Staff Interview Application#
How to Introduce This Pattern#
"Before I pick a data layout, I want to list what tenants share — threads, connections, queues, cache, deploys — and give each one a per-tenant budget in the unit it's spent in. Then I'll cap blast radius with cells, enforce data isolation in the database, not just the ORM, and give the biggest tenants a path to dedicated capacity they pay for."
Lead with what's shared, then how each shared thing is divided, then blast radius, then the tier ladder and who owns moves.
When NOT to Use This Pattern#
- Single-customer or internal tools: an internal admin app for one company has one tenant.
tenant_ideverywhere and RLS are cost without benefit. - A handful of large customers: with 10–20 tenants, silo each one. Pooling machinery costs more than the hardware it saves.
- Untrusted code: quotas and fair queuing don't stop a sandbox escape. Start from hard isolation and see Code Runner.
- Abuse rather than tenancy: if the noisy party is an attacker, not a customer, per-identity Rate Limiter and abuse detection come first.
- When the real problem is one hot entity, not one tenant: a single viral record inside a tenant is a Hot Keys problem; tenant isolation won't split it.
Follow-Up Questions to Anticipate#
| Interviewer Asks | What They Are Testing | How to Respond |
|---|---|---|
| "Shared database or database per tenant?" | Tiering judgment | "Pooled for the long tail, dedicated for the top tier and regulated tenants, same schema so tenants can move." |
| "One tenant is slowing everyone down. What do you do now?" | Operational depth | "Per-tenant metrics name them in minutes; cap their concurrency, kill runaway queries, then fix the missing budget at that layer." |
| "How do you guarantee tenants can't see each other's data?" | Defence in depth | "Tenant from the token, composite keys, RLS with FORCE, and a cross-tenant probe on every deploy." |
| "How big should a cell be?" | Blast-radius reasoning | "Start from the fraction of customers I can afford to lose at once — 5% means 20 cells — then check the largest tenant fits under ~30% of one." |
| "What about tenants that grow huge?" | Placement design | "Directory-based placement and a move tool: copy, catch up via CDC, pause writes ~30 seconds, verify, flip." |
| "How do you charge for this?" | Cost attribution | "Per-tenant usage metering on the expensive dimensions; dedicated tiers priced at least at their infrastructure cost." |
Scorecard#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Framing | Chooses shared tables vs separate databases | Inventories shared resources and divides each | Defines isolation tiers as products with written guarantees |
| Noisy neighbour | Gateway rate limit | Per-tenant budgets at every layer in the right unit; fair queuing; cells | Platform defaults so every service inherits limits |
| Data isolation | ORM tenant scope | RLS backstop, composite keys, cross-tenant probe | Security sign-off per tier; audited guarantees |
| Growth | Bigger instance for big tenants | Directory placement and a rehearsed move tool | Tier pricing matched to cost; lifecycle runbooks |
| Observability | Global metrics | Per-tenant metrics for top tenants; < 5 min to name the neighbour | Per-tenant cost attribution feeding pricing |
Strong Hire Signals
| Signal | What It Sounds Like |
|---|---|
| Right unit of fairness | "The database spends connection-seconds, so that's what I budget per tenant." |
| Blast radius as a number | "20 cells, so a bad deploy touches 5% of customers." |
| Isolation enforced twice | "ORM filter for plans, RLS for safety, and a probe that proves it." |
| Plans for the whale | "The top tenant is 12% of capacity. They get a dedicated cell, priced in." |
Lean No-Hire Signals
| Signal | Why It Misses the Bar |
|---|---|
| "Rate limiting solves noisy neighbours" | Ignores expensive requests within the limit and every non-request resource |
| Isolation by application filter only | One forgotten query becomes a breach |
| No plan for a tenant that outgrows a shard | Hash placement with no move tool leaves no lever on the worst day |
Common False Positives: Knowing PostgreSQL RLS syntax ≠ having a tenancy model. Drawing cells ≠ sizing them by blast radius. "Database per tenant" ≠ isolation, if they all share one connection pool and one deploy.
Capacity Planning Quick Reference#
Sizing Isolation#
tenant_share(resource) = tenant_usage / pooled_capacity # cap at 0.20 by default
cell_count ≥ 1 / tolerable_blast_fraction # 5% → 20 cells
tenants_per_cell = total_tenants / cell_count
whale_test = largest_tenant_load / cell_capacity # > 0.3 → dedicated cell
shuffle_combinations = C(nodes, shard_size) # C(100,5) ≈ 75.3M
fair_queue_wait(small) ≈ scheduling_rounds × avg_job_cost / workers # seconds, not backlog
per_tenant_concurrency_cap = max(min_floor, 0.2 × service_limit)
pooled_cost_per_tenant = pooled_infra_cost / pooled_tenants
dedicated_floor ≈ $8–15K/month per cell # price the tier above this
Key Numbers Worth Memorizing#
| Number | Context |
|---|---|
| Top 1% ≈ 30–50% of load | Typical B2B tenant skew; size for the whale |
| 20% | Default cap on any one tenant's share of a pooled resource |
| 5–10% | Tolerable blast radius per cell for most B2B products |
| ~30% of a cell | Above this, a tenant should get dedicated capacity |
| 5–20× | Cost per tenant of silo vs pool |
| C(100,5) ≈ 75M | Shuffle-shard combinations for 5 of 100 nodes |
| 2,048 / 4 / ~730B | Route 53 virtual name servers / per domain / combinations |
| 30–60 s | Directory cache TTL; also roughly the write pause in a tenant move |
| Top ~100 tenants | Per-tenant metric labels; roll the rest into "other" |
| < 5 min | Target time for on-call to name the noisy tenant |
Common Pitfalls Checklist#
- Tenant identity comes from credentials, never from request parameters
-
tenant_idleads every primary key, shard key and cache key - Row-level security with FORCE (or per-tenant databases) backs up application filters
- A cross-tenant read probe runs in CI and after every deploy
- Every shared resource has a per-tenant budget in the unit it's spent in
- Background work uses weighted fair queuing, not one FIFO
- Per-tenant metrics exist for the top tenants; on-call can name the neighbour in minutes
- Placement goes through a directory; the tenant-move tool is rehearsed
- Cells are sized by tolerable blast radius and deployed in waves
- Dedicated tiers are priced at or above their infrastructure cost