Why This Matters#
Back-of-envelope estimation is not an arithmetic test. Interviewers don't care whether you compute 2B ÷ 86,400 to three significant figures. Estimation is a decision tool: its only job is to tell you which parts of the design are hard. If your estimate doesn't change a design decision — shard or don't, cache or don't, one region or three, build or buy — it was a ritual, not an estimate.
Most candidates perform estimation as a toll booth at the start of the interview: QPS, storage, bandwidth, then never mention the numbers again. The design that follows would have been identical with numbers 100× larger or smaller. Interviewers notice. The Staff version is shorter and sharper: "40 writes a second is nothing — one Postgres primary handles that for a decade. The interesting number is 12K reads/second at peak with a hot tail, so the design effort goes into the read path." The estimate chose where to spend the next 30 minutes.
At Principal level, the estimate has one more step: turn it into dollars and people. 8PB of photos is not a storage fact; it's ~$190K/month on standard object storage, ~$25–30K/month if 90% of it moves to a deep-archive tier, and a business decision about which of those the product can live with. The engineer who can say that sentence gets pulled into planning meetings; the one who can't gets handed the plan.
The 60-Second Version#
- Memorize a small set of anchors, not a big table. ~1 day ≈ 10⁵ seconds (86,400). 1M requests/day ≈ 12/s. 1B/day ≈ 12K/s. Peak ≈ 2–3× average (10× for launches and events).
- Latency anchors: memory ~100ns, NVMe SSD random read ~20–100µs, same-AZ round trip ~0.1–0.5ms, cross-region ~60–150ms. Every tier is roughly 100–1,000× the one before.
- Throughput anchors (per node, orders of magnitude): Redis ~100K+ ops/s; Postgres ~10K+ simple writes/s and ~50K+ indexed reads/s; Kafka broker ~100s of MB/s; a stateless app instance ~1–10K RPS; a tuned connection server ~100K–1M sockets.
- Size anchors: a row ~100B–1KB, a post ~1KB with metadata, a phone photo ~2–5MB, 1 minute of 1080p video ~40–60MB.
- Round aggressively, one significant figure. 86,400 → 10⁵. 2.6M seconds/month → 2.5 × 10⁶. Being off by 20% never changes a design; being off by 100× always does.
- Every estimate ends with "so…". "…so one primary is enough." "…so storage, not compute, is the cost driver." "…so we need sharding by year 2, not day 1."
- L7 converts to money: ~$30–35 per vCPU-month on demand, ~$0.023/GB-month object storage, ~$0.05–0.09/GB internet egress, ~$0.01/GB each way cross-AZ, ~$25–35K/month per fully loaded engineer. People are usually the biggest line item until you pass ~$1M/month in infrastructure.
How Estimation Works#
The Method (4 Steps, Under 5 Minutes)#
- State the drivers. Users (DAU), actions per user per day, object sizes, read:write ratio, retention. Say them out loud as assumptions; invite correction.
- Compute the rates. Average QPS = daily actions ÷ 10⁵. Peak = 2–3× average (say which multiplier and why).
- Compute the volumes. Storage/day = writes/day × size × replication factor. Multiply by retention. Bandwidth = QPS × response size.
- Draw the conclusion. Compare each number against the per-node anchors. Anything within 10× of a single node's limit is the design's hard part; anything 1,000× below is not worth discussing.
Time Constants#
| Quantity | Exact | Use |
|---|---|---|
| Seconds per day | 86,400 | ≈ 10⁵ |
| Seconds per month | ~2.6M | ≈ 2.5 × 10⁶ |
| Seconds per year | ~31.5M | ≈ 3 × 10⁷ (π × 10⁷ is a handy mnemonic) |
| 1M per day | ~11.6/s | ≈ 12/s |
| 100M per day | ~1,160/s | ≈ 1.2K/s |
| 1B per day | ~11,600/s | ≈ 12K/s |
| 1M per month | ~0.4/s | Barely a trickle |
Powers of Two and Ten#
| Power | Exact | Approx | Name |
|---|---|---|---|
| 2¹⁰ | 1,024 | 10³ | KB (kilo) |
| 2²⁰ | 1,048,576 | 10⁶ | MB (mega) |
| 2³⁰ | ~1.07 × 10⁹ | 10⁹ | GB (giga) |
| 2⁴⁰ | ~1.1 × 10¹² | 10¹² | TB (tera) |
| 2⁵⁰ | ~1.13 × 10¹⁵ | 10¹⁵ | PB (peta) |
| 2³² | ~4.3 × 10⁹ | 4B | Max unsigned 32-bit int — the reason IDs overflow |
| 62⁷ | ~3.5 × 10¹² | 3.5T | Base62 IDs of length 7 |
Latency Numbers#
Approximate, modern server hardware. Precision is not the point — the ratios are.
| Operation | Latency | Relative to Memory Access |
|---|---|---|
| L1 cache reference | ~1ns | 0.01× |
| Branch mispredict | ~3ns | 0.03× |
| L2 cache reference | ~4ns | 0.04× |
| Uncontended mutex lock/unlock | ~20ns | 0.2× |
| Main memory reference | ~100ns | 1× |
| Compress 1KB (LZ4/Snappy) | ~1–3µs | 10–30× |
| Context switch / syscall | ~0.1–5µs | 1–50× |
| Read 1MB sequentially from memory | ~10–50µs | 100–500× |
| NVMe SSD random 4KB read | ~20–100µs | 200–1,000× |
| Round trip within an AZ | ~100–500µs | 1,000–5,000× |
| Redis GET over the network | ~0.2–1ms | 2,000–10,000× |
| Read 1MB sequentially from NVMe | ~0.3–1ms | ~5,000× |
| Indexed DB point query (hot pages) | ~1–5ms | 10,000–50,000× |
| HDD seek | ~5–10ms | 50,000–100,000× |
| Read 1MB sequentially from HDD | ~5–10ms | ~75,000× |
| Cross-region round trip | ~60–150ms | ~1,000,000× |
| Cold HTTPS request across an ocean | ~250–600ms | Several million× |
The three facts that drive most designs: memory is ~1,000× faster than a network hop; a network hop within a region is ~100× faster than a cross-region hop; and a disk seek on spinning media costs as much as ~20 in-region round trips.
Throughput Numbers (Per Node, Order of Magnitude)#
| Component | Rough Ceiling | Caveat |
|---|---|---|
| Stateless app instance (4–8 vCPU, JSON API) | ~1K–10K RPS | Dominated by what each request does |
| NGINX / Envoy proxy | ~10K–50K+ RPS | TLS termination and logging cost the most |
| Redis (single instance) | ~100K–200K ops/s; 1M+ with pipelining | Single-threaded command execution |
| Memcached | ~several hundred K ops/s | Multithreaded |
| PostgreSQL / MySQL primary | ~10K–20K simple write TPS; ~50K–100K+ indexed reads/s | fsync, index count, and row size dominate |
| Cassandra node | ~10K–50K writes/s | LSM writes are cheap; reads and compaction are not |
| Kafka broker | ~100s of MB/s; a partition ~10+ MB/s | Disk and network bound, not CPU |
| Elasticsearch node | ~thousands–tens of thousands of docs/s indexed | Mapping and refresh interval dominate |
| WebSocket / connection server | ~100K–1M+ concurrent connections | Memory per connection and ephemeral ports |
| 10 Gbps NIC | ~1.25 GB/s | Often the real ceiling for caches and proxies |
Size Numbers#
| Object | Size | Note |
|---|---|---|
| Int / long / timestamp | 4 / 8 / 8 bytes | |
| UUID | 16 bytes binary, 36 as string | Strings in indexes cost 2×+ |
| Typical relational row | ~100B–1KB | Plus ~20–30B per-row overhead and indexes |
| Short text post with metadata | ~1KB | Text is small; metadata and indexes dominate |
| JSON vs binary encoding | JSON ~2–5× larger | Matters at >10K RPS or TB-scale storage |
| Thumbnail image | ~20–50KB | |
| Phone photo (original) | ~2–5MB | |
| 1 minute of 1080p video | ~40–60MB (at ~5–8 Mbps) | 4K is ~3–4× that |
| 1 hour of audio (128 kbps) | ~60MB | |
| Metric sample (compressed) | ~1–2 bytes | Raw is ~16 bytes (timestamp + value) |
Core Strategies#
Strategy 1: Rate Estimation (QPS)#
avg_qps = DAU × actions_per_user_per_day ÷ 86,400
peak_qps = avg_qps × peak_factor # 2–3× diurnal; 5–10× launches/events
read_qps = peak_qps × read_share
write_qps = peak_qps × write_share
Example: 20M DAU × 50 reads/day = 1B reads/day → ~12K/s avg → ~30K/s peak (2.5×)
When to use: Every interview. It decides whether the hot path is a single node, a cache, or a sharded tier.
Failure mode: Using the average. Systems are sized for peak, and SLOs are broken at peak. A service sized for 12K/s average falls over at 7pm every day.
Strategy 2: Storage Estimation#
storage_per_day = writes_per_day × object_size
storage_total = storage_per_day × retention_days × replication_factor × (1 + index_overhead)
Example: 100M posts/day × 1KB = 100GB/day
× 365 days × 3 replicas × 1.3 index overhead ≈ 140TB/year
When to use: Whenever data is retained. It decides sharding, tiering, and whether storage or compute is the cost driver.
Failure mode: Forgetting replication (×3), indexes (+20–50%), and secondary copies (search index, analytics lake, backups — often another ×2–3). "100GB/day" becomes ~1TB/day of provisioned storage across the company.
Strategy 3: Bandwidth Estimation#
egress_bytes_per_sec = read_qps × avg_response_size
Example: 30K/s × 200KB image = 6GB/s ≈ 48 Gbps → CDN, not origin servers
When to use: Media, file sync, video, large API payloads.
Failure mode: Ignoring bandwidth because QPS looked small. 3K RPS of 2MB responses is 48 Gbps — a network problem that no amount of CPU will solve.
Strategy 4: Fleet Sizing#
instances = peak_qps ÷ (per_instance_qps × target_utilization)
Example: 30K RPS ÷ (2K RPS × 0.6) = 25 instances, + N+2 for failures/deploys → ~27–30
When to use: When the interviewer asks "how many servers?" — and to sanity-check cost.
Failure mode: Planning at 100% utilization. Latency climbs sharply above ~70% utilization (queueing), and you need headroom to lose an AZ: with 3 AZs, each must absorb 50% more load when one fails, so steady-state utilization should sit near ~60%.
Strategy 5: Memory / Cache Sizing#
cache_bytes = hot_objects × (object_size + per_key_overhead)
Example: top 20% of 500M URLs = 100M keys × (500B + ~50B overhead) ≈ 55GB
→ fits in a 3-node cache cluster with room for replicas
When to use: Deciding whether the working set fits in memory — the single largest latency lever in most read-heavy designs.
Failure mode: Caching the whole dataset. Access is usually Zipfian: the top 1–20% of keys serve 80–99% of reads. Size for the working set, not the table.
From Estimate to Decision: The Hard Sub-Problem#
The arithmetic is easy. Three things make estimates wrong in ways that matter:
Peak Is Not a Constant#
| Pattern | Peak ÷ Average | Example |
|---|---|---|
| Diurnal consumer traffic | ~2–3× | Social feed, messaging |
| Business-hours B2B | ~3–5× (near zero at night) | Collaboration tools, CRMs |
| Scheduled spikes | ~5–10× for minutes | Top-of-hour cron jobs, push notification sends |
| Launches and events | ~10–100× | Ticket drops, flash sales, live sports |
| Retry storms | Unbounded | Every client retrying after an outage |
Skew Beats Averages#
Average load per shard means little when one key is hot. If 1% of users generate 30% of traffic (common for celebrities, large tenants, popular products), the busiest shard sees several times the average. Estimate the p99 shard, not the mean shard.
10M users, 100 shards → 100K users/shard on average
Top tenant: 2M users on one shard → 20× the average shard
Conclusion: shard by something finer than tenant, or isolate the whale.
Headroom Has Three Components#
provisioned = peak × (1 / target_utilization) × failure_domain_factor × growth_buffer
= peak × 1/0.7 × 1.5 (lose 1 of 3 AZs) × 1.3 (6 months growth)
≈ 2.8 × peak
Candidates who size to peak are ~3× under-provisioned for a real production system. Candidates who size to 10× peak "to be safe" have tripled the bill for no reason. Say the multiplier and justify each factor.
🎯 Staff Move: "Peak is about 30K reads a second. With 70% target utilization, surviving an AZ loss, and six months of growth, I'd provision for ~85K. At ~2K per instance that's ~40–45 four-vCPU instances — about $5–6K a month of compute. Compute is not the problem here. Storage at 140TB a year is, so I'll spend the deep dive on tiering."
Visual Guide#
The Estimation Loop#
Where Each Scale Threshold Forces a Decision#
These thresholds are deliberately conservative "stop and think" markers, not hard limits — a tuned Postgres primary can exceed 5K writes/s. Their job is to tell you when the default architecture stops being obviously enough.
Implementation Patterns: Worked Estimations#
Worked Example 1: URL Shortener#
Assumptions: 100M new URLs/month, 100:1 read:write, 500B per record, 5-year retention.
Writes: 100M / 2.5M s ≈ 40/s avg, ~100/s peak
Reads: 40 × 100 ≈ 4K/s avg, ~10–12K/s peak
Storage: 100M × 12 × 5 = 6B records × 500B = 3TB (×3 replicas = 9TB)
Key space: base62, 7 chars = 3.5T IDs → 6B uses ~0.2%
Cache: top 20% of daily-accessed URLs; ~10M hot keys × ~600B ≈ 6GB
So: writes are trivial (one primary for years); 3TB fits on one node's disk but 9TB replicated argues for a partitioned KV store by year 2–3; reads are the design problem, and 6GB of cache absorbs most of them. Don't shard on day 1. Do put a cache in front on day 1.
Worked Example 2: Chat Messaging#
Assumptions: 50M DAU, 40 messages sent/user/day, 300B/message incl. metadata,
20% of DAU connected at peak, retention forever.
Messages: 50M × 40 = 2B/day → ~23K/s avg → ~60–70K/s peak
Deliveries (avg ~2 recipients incl. multi-device): ~50K/s avg, ~140K/s peak
Storage: 2B × 300B = 600GB/day → ×3 replicas = 1.8TB/day → ~650TB/year
Connections: 10M concurrent → at ~200K per gateway host ≈ 50 hosts (+ headroom ≈ 75)
So: write throughput (~70K/s) and storage growth (~650TB/year) rule out a single relational primary — this is a partitioned, write-optimized store (wide-column/LSM) keyed by conversation. The connection tier is a separate stateful fleet. The design depth goes into fan-out and the connection tier, not the API.
Worked Example 3: Photo Sharing#
Assumptions: 10M uploads/day, 2MB original + ~300KB of resized variants,
500M photo views/day at ~200KB average served size.
Upload storage: 10M × 2.3MB = 23TB/day → ~8.4PB/year (object storage handles durability)
Upload bandwidth: 23TB / 10^5 s ≈ 230MB/s avg ingress
View egress: 500M × 200KB = 100TB/day ≈ 1.2GB/s avg ≈ 10 Gbps avg, ~25 Gbps peak
So: this is a storage-and-bandwidth problem, not a compute problem. Egress (~3PB/month) goes through a CDN; the storage design is about tiering the ~95% of photos that are rarely viewed after 30 days. See the Principal Lens for what that costs.
Worked Example 4: Metrics Ingestion#
Assumptions: 100K hosts × 500 time series × 1 sample every 10s.
Samples: 100K × 500 / 10 = 5M samples/s
Raw size: 16B/sample (8B timestamp + 8B value) → 80MB/s → ~6.9TB/day
Compressed (delta-of-delta timestamps + XOR'd floats): ~1.4B/sample → ~0.6TB/day
Retention: 15 days raw at full resolution ≈ 9TB; 1 year downsampled to 5-min ≈ ~7TB
So: compression is the design. A 12× compression ratio is the difference between a small cluster and a large one. Cardinality (number of series), not sample rate, is the thing that will page someone — one team adding a user_id label multiplies series by millions.
Where the Money Goes: The Video Example#
Using the video walkthrough below (1M uploads/day, 100M views/day), the year-1 monthly bill splits roughly like this — which is why the design effort belongs in egress and storage, not in the upload API:
Failure Scenario: The Estimate That Missed a Multiplier#
Design review: "10K events/s, 200B each → 170GB/day. One Kafka cluster, 7-day retention."
Launch day: Mobile SDK batches events and retries on flaky networks → 2.5× duplicates.
Week 2: A product team adds a 'scroll_depth' event firing every 250ms → 8× volume.
Week 3: Brokers at 85% disk; retention cut to 2 days; downstream jobs miss replays.
Week 4: Emergency cluster expansion (+$40K/month) and a sampling hotfix.
Detection: kafka.disk.used_pct > 75%; ingest.events_per_sec vs forecast ratio > 2×; per-event-type volume dashboard.
Blast radius: every consumer that relied on 7-day replay; every team sharing the cluster.
Mitigation: sampling at the SDK for high-frequency types; per-producer quotas.
Prevention: estimates include duplication/retry factors and a per-event-type budget; new high-frequency events require a capacity review; the 90-day actuals check.
Owner: data platform owns the cluster and quotas; each producing team owns its event budget.
| Estimation Miss | Typical Multiplier | Who Catches It Too Late |
|---|---|---|
| Replication and secondary copies ignored | 3–10× storage | Storage on-call when disks fill |
| Average used instead of peak | 2–10× rate | Service on-call at the daily peak |
| Client retries / duplicates ignored | 1.5–3× rate | Ingest pipeline on-call |
| Hot keys / whales ignored | 10–100× on one shard | Whoever owns the hottest shard |
| Growth not compounded | 2–5× by year 2 | Finance, at budget time |
The Numbers in Context#
Unit Costs (Approximate Public Cloud List Prices)#
Approximate AWS us-east-1 on-demand list prices; they change and commitments discount them. Use them for orders of magnitude, never for a budget.
| Resource | Unit Cost | What It Means for Your Design |
|---|---|---|
| General-purpose vCPU | ~$30–35/vCPU-month on demand; ~30–60% less with 1–3 year commitments | 100 instances × 8 vCPU ≈ $25–30K/month. Compute is rarely the top line item until large scale. |
| Memory (memory-optimized instances) | ~$3–5/GB-month | A 500GB cache tier ≈ $2K/month. Caching the working set is almost always cheaper than scaling the DB. |
| Block storage (SSD, gp3-class) | ~$0.08/GB-month | 10TB ×3 replicas on EBS ≈ $2.4K/month. |
| Object storage (standard) | ~$0.023/GB-month ($23/TB) | 1PB ≈ $23K/month. Durable and cheap per GB; request charges matter for small objects. |
| Object storage (infrequent / instant-retrieval archive) | ~$0.0125 / ~$0.004 per GB-month | Tiering cold data cuts storage cost 2–6×. |
| Object storage (deep archive) | ~$0.001/GB-month | ~23× cheaper than standard; retrieval takes hours. For compliance, not for serving. |
| Object storage requests | PUT ~$0.005/1K; GET ~$0.0004/1K | 1B small PUTs/month ≈ $5K — more than storing them. Batch small objects. |
| Internet egress | ~$0.05–0.09/GB (volume tiers) | 1PB/month egress ≈ $50–90K. Egress is often the largest bill for media. |
| CDN delivery | ~$0.01–0.08/GB depending on volume and commit | At PB scale, negotiated CDN rates beat origin egress by 2–5×. |
| Cross-AZ transfer | ~$0.01/GB each direction | Chatty east-west traffic becomes a real line item at PB/month. |
| Serverless functions | ~$0.20 per 1M requests + duration | Great below ~10–50M requests/month; a steady 5K RPS workload is usually cheaper on instances. |
| Fully loaded engineer | ~$25–35K/month (US big-tech-ish) | Three engineers ≈ $1M/year. Below ~$1M/month of infra, people are usually the biggest cost. |
Unit Economics: The Number That Survives the Interview#
Raw totals don't compare across designs; cost per unit of value does:
| Metric | How to Compute | Healthy Ballpark (consumer web, varies widely) |
|---|---|---|
| $ per 1M requests (compute) | monthly compute $ ÷ (monthly requests ÷ 1M) | ~$0.05–1 for simple APIs |
| $ per active user per month | total infra $ ÷ MAU | ~$0.01–0.50 for social/content; more for video |
| $ per GB stored per month (blended) | storage $ ÷ total GB (all tiers, all replicas) | ~$0.005–0.03 |
| Infra $ ÷ revenue | total infra ÷ revenue | Commonly ~5–15% for SaaS; lower is a product of scale and discipline |
🧭 Principal Insight: "I don't ask whether the design is expensive. I ask whether cost per active user goes up or down as we grow. A design whose unit cost rises with scale is a design that will be rewritten under duress."
How This Shows Up in Interviews#
Scenario 1: "Estimate the storage for a Twitter-like service."#
The interviewer wants the drivers, the math, and the conclusion — in about two minutes. "Say 200M DAU, 20% post once a day: 40M posts/day × ~1KB with metadata ≈ 40GB/day, ×3 replicas ≈ 120GB/day, ~45TB/year. Media dominates: if 10% of posts carry a 500KB image, that's 2TB/day, ~730TB/year in object storage. So the text store is a partitioned DB of modest size; the media is an object-storage and CDN problem. I'll design those separately."
Scenario 2: "We want to build a video-sharing feature. Is that feasible?" (Full Walkthrough)#
Step 1 — Drivers, stated as assumptions. "Let's say 1M uploads per day, average 3 minutes at 1080p, and 100M views per day, average 4 minutes watched. I'll invite correction on those."
Step 2 — Storage. "3 minutes × ~50MB/min = 150MB per original. We transcode into ~4 renditions that together are roughly the original's size again, so ~300MB per upload. 1M uploads/day × 300MB = 300TB/day — ~110PB/year. That's the headline: this is a storage company."
Step 3 — Bandwidth. "100M views × 4 minutes × ~40MB/min at a blended rendition ≈ 16PB/day of egress, ~190GB/s average, ~1.5 Tbps. Only a CDN — with an edge footprint or a large commit — can carry that."
Step 4 — Compute. "Transcoding 3 minutes of 1080p into 4 renditions costs on the order of several CPU-minutes per upload. 1M uploads/day × ~10 CPU-minutes ≈ 7K CPU-cores busy continuously — real, but small next to storage and egress."
Step 5 — Convert to money. "Year-1 storage at standard object pricing averages ~55PB × $23/TB ≈ $1.3M/month and grows; egress at a negotiated ~$0.005–0.01/GB on 480PB/month is ~$2.5–5M/month. Compute is ~$150–250K/month. So egress and storage are ~95% of the bill."
Step 6 — Decide. "That changes the design: aggressive tiering (most videos get almost no views after 30 days), delete renditions that aren't watched and re-transcode on demand, a per-title bitrate ladder, and a CDN strategy — possibly our own edge caches inside ISPs at this scale. Before building any of it, product and finance need to see the per-view cost, because it sets the ad or subscription economics."
Why this is a Staff answer: The estimate identified storage and egress as the dominant problems, the design effort went there, and the output was a decision (tiering, on-demand transcoding, CDN strategy) plus a question for the business. Nobody spent 10 minutes on the upload API.
Scenario 3: "Your estimate is off — we have 10× more users than that."#
This tests whether your design is sensitive to the number. "Good — let me see what changes. 10× writes puts us at ~400/s, still one primary. 10× reads is 100K/s peak, so the cache tier grows from 3 nodes to ~15, which is fine. 10× storage is 30TB replicated to 90TB — that crosses my single-node threshold, so partitioning moves from year 3 to day 1. So one decision flips; the rest holds." Showing which decisions are sensitive to which inputs is more valuable than the original estimate.
Scenario 4: "How many servers do we need?"#
"Peak is 30K RPS, and a load test shows ~2K RPS per 4-vCPU instance at our p99 target. Target 60–70% utilization, survive losing one of three AZs, six months of growth: ~2.8× peak ≈ 85K RPS of capacity → ~43 instances. I'd round to 45 across three AZs, autoscale between 30 and 60, and re-derive the per-instance number after every major release because it drifts."
Advanced Patterns#
| Pattern | How It Works | When to Use |
|---|---|---|
| Sensitivity analysis | Vary each driver 10× and see which decisions flip | When inputs are uncertain — which is always |
| Fermi decomposition | Break an unknown into knowable factors (users × actions × size) | Any estimate with no direct data |
| Load-test-derived per-node numbers | Replace anchors with measured RPS/instance at the p99 target | Anything past the whiteboard — anchors are for interviews |
| Little's Law | in-flight = arrival rate × latency (L = λW) | Sizing connection pools, thread pools, queues: 5K RPS × 200ms = 1K in flight |
| Working-set analysis | Measure the fraction of keys serving 80/95/99% of reads | Cache sizing; tiering boundaries |
| Unit economics | Divide cost by users, requests, or GB | Comparing designs and justifying investments |
| Growth curves | Model storage as cumulative (it never shrinks) and traffic as rate | Multi-year capacity planning; storage costs compound |
The Principal Lens#
Why L7 Sees This Problem Differently#
A Staff engineer uses estimates to find the hard part of a design. A Principal engineer uses them to decide what the company should build, fund, and stop doing. At L7 the estimate's output is not "we need 45 servers"; it's "this feature costs $0.04 per active user per month, rising to $0.07 by year 3 because storage compounds, and here's the tiering investment that bends it back to $0.03." The unit of analysis moves from the system to the portfolio: which systems' unit costs grow faster than their value, and which teams need a capacity model before they need another engineer.
The Org-Level Fault Line#
Central capacity/FinOps function vs team-owned budgets with unit economics.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Central capacity team owns forecasts and budgets | Consistent models; negotiated commitments; one throat to choke | Teams have no incentive to be efficient; forecasts lag product reality | Central team becomes a bottleneck; product teams overspend invisibly |
| Each team owns its bill (showback/chargeback) | Incentives aligned; teams optimize what they see | Inconsistent estimation quality; local optimizations that raise global cost (e.g., everyone buys their own cache) | Teams without cost expertise; duplicated platforms |
| Central models + team-owned unit-cost targets | Shared tooling, per-team accountability on $/unit | Requires good attribution (tagging) and a review cadence | A small FinOps/platform group plus every team's tech lead |
The Principal position: centralize the model (unit prices, forecasting tooling, commitment purchasing) and decentralize the accountability (each product's $/unit target, reviewed quarterly). Mandate that every design doc include an estimate that ends in a unit cost.
Cost Model#
Using the photo-sharing workload from Worked Example 3, scaled. Assumptions: standard object storage $23/TB-month, instant-retrieval archive $4/TB-month, CDN $0.02–0.05/GB, compute at list, engineers at ~$30K/month fully loaded. Storage shown at the end of year 1 (it compounds after that).
| Scale | Uploads/Day | Storage (end of yr 1) | Storage $/month | Egress $/month | Compute $/month | Team | People $/month |
|---|---|---|---|---|---|---|---|
| Small | 100K | ~84TB | ~$2K | ~$1.5K (30TB) | ~$2K | 3 engineers | ~$90K |
| Medium | 1M | ~840TB | ~$19K; ~$7K tiered | ~$9K (300TB) | ~$15K | 15 engineers | ~$450K |
| Large | 10M | ~8.4PB | ~$190K; ~$65K tiered (20% hot) | ~$60K (3PB at negotiated rates) | ~$100K | 50–60 engineers | ~$1.5–1.8M |
Two Principal conclusions: (1) at every scale here, people cost more than infrastructure — so the highest-leverage decision is often not building something (buy the CDN, buy the transcoder); (2) storage is the only line that compounds — by year 3 the Large row holds ~25PB, and the tiering decision is worth ~$4M/year.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversal Cost |
|---|---|---|
| Instance counts and autoscaling bounds | Two-way | Minutes |
| On-demand vs 1-year commitment | Two-way after a year | Up to a year of committed spend |
| 3-year reserved commitment | One-way for 3 years | Paying for unused capacity if the architecture changes |
| Retaining all user data forever | One-way in practice | Deleting later requires product, legal, and user-trust sign-off |
| Storage format and object layout (e.g., millions of tiny objects) | Mostly one-way | Re-writing PBs to change layout; request costs to migrate |
| Building your own CDN / edge | One-way | Hardware, peering contracts, and a team for years |
The Standard I'd Write#
RFC: Capacity & Cost Estimation in Design Reviews (v1)
Scope: All design docs for new systems or changes expected to add >$10K/month or >10% to an existing system's cost.
MUST:
- Include a capacity section: drivers (stated as assumptions), peak QPS (with peak factor), storage at 1 and 3 years (with replication and secondary copies), egress, and headroom multiplier.
- End with a unit cost ($ per 1M requests, per active user, or per GB) at launch and at year 3.
- Name the top three cost drivers and the one-way-door commitments (retention, commitments, formats).
- Include a data lifecycle policy for any retained data (tiering and deletion dates).
- Re-check actuals against the estimate 90 days after launch; a >2× miss triggers a brief review.
SHOULD: use the org's published unit-price sheet and load-test-derived per-node numbers rather than personal anchors.
Exceptions: Prototypes under $2K/month; approved by the area's tech lead.
Success metrics: ≥ 90% of in-scope designs include a unit cost; median estimate within 2× of actuals at 90 days; org-wide cost per active user flat or declining year over year.
What I'd Tell the VP#
"Most of our systems are sized by gut feel and we only find out what they cost when the bill arrives. We want every new design to show, on one page, what it will cost per user now and in three years, and to check that number 90 days after launch. This doesn't need a new team — it needs a standard price sheet and a line in the design review template. Where we've done this informally, it has found savings of 30–60% on storage alone, mostly by moving old data to cheaper tiers. More importantly, it tells us early which products get more expensive per user as they grow, so we fix them before they're big."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Ends estimates in dollars | "8PB is ~$190K a month on standard storage. Tiering 80% of it gets us under $70K." |
| Tracks unit cost over time | "Cost per user rises from 4 to 7 cents by year 3 because storage compounds. That's the curve I want to bend." |
| Weighs people against infrastructure | "The whole infra bill is less than the three engineers it would take to build our own transcoder. Buy it." |
| Names commitments as one-way doors | "I'd take a 1-year commitment on the baseline, not 3-year — the architecture will change before then." |
| Makes estimation a standard | "Every design doc should end in a unit cost, checked 90 days after launch." |
Staff answers that L7 interviewers find insufficient:
- "We need about 45 servers." — Correct, and stops before the question the business actually has: what does it cost per user, and how does that change?
- "Storage is cheap, so we'll keep everything." — True per GB, false cumulatively; ignores compounding and a one-way retention door.
- "We'll optimize cost later." — Some cost decisions (formats, retention, commitments) are one-way doors made on day one.
In the Wild#
These are public, documented examples.
Jeff Dean's "Numbers Everyone Should Know"#
Jeff Dean popularized a short table of latency figures — L1 cache reference, main memory reference, disk seek, round trip within a data center, a packet from California to the Netherlands and back — in talks around 2009 (including at LADIS and Stanford). The exact values have shifted with SSDs and faster networks, and several people have published updated versions, but the ratios remain the reason the table endures: memory versus network versus disk versus cross-continent differ by orders of magnitude.
Staff insight: Interviewers don't check your nanoseconds. They check whether you know that a cross-region call costs about a million memory references, and whether your design reflects that.
Stack Overflow: Scale Up Before You Scale Out#
Stack Overflow has publicly described serving one of the web's most-visited sites from a small number of on-premises servers — on the order of ten web servers and a couple of large SQL Server machines — with heavy use of caching. Their engineers have written about choosing vertical scale and careful performance work over a large distributed architecture.
Staff insight: Honest estimation often concludes "one big database is enough." Saying that confidently, with numbers, is a stronger signal than reflexively sharding.
WhatsApp: Two Million Connections per Server#
WhatsApp's engineering team published in 2012 that they had tuned FreeBSD and Erlang to hold over 2 million concurrent TCP connections on a single server. The company was widely reported to be serving hundreds of millions of users with an engineering team of only a few dozen when Facebook acquired it in 2014.
Staff insight: Per-node anchors decide fleet size and team size. At 2M connections per host, 100M concurrent users need ~50 hosts, not 5,000 — and a small team can run 50 hosts. Estimating the per-node limit correctly changes the organization, not just the diagram.
Staff Calibration#
What Staff Engineers Say (That Seniors Don't)#
| Concept | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| QPS | "1B requests/day is about 11,574 QPS" | "~12K/s average, ~30K/s peak — within a cache cluster's range, so the read path is fine; writes at 40/s don't need discussion" | "At 30K/s the compute is ~$6K/month; the interesting cost is egress, so that's what we optimize" |
| Storage | "We'll need 100TB" | "100TB replicated ×3 plus search and backups ≈ 500TB provisioned; that forces partitioning and tiering" | "Storage compounds; by year 3 it's the top line item. Lifecycle policy is a launch requirement, not a follow-up" |
| Servers | "We need 15 servers" | "15 at peak; ×2.8 for utilization, AZ loss, and growth ≈ 42, autoscaled 30–60" | "Baseline on a 1-year commitment, burst on demand; re-derive per-node numbers each release" |
| Precision | Computes to 4 significant figures | Rounds to 1 significant figure and moves on | Says which inputs the decision is sensitive to and asks product to firm those up |
| Conclusion | Numbers presented, design proceeds unchanged | "So…" — each number chooses a design direction | "So…" — each number chooses a design direction and a unit cost the business can plan around |
Why "Conclusion" separates levels
The failure mode interviewers see most often is an accurate estimate that changes nothing. Seniors compute; Staff engineers conclude — "so writes are trivial, reads are the problem." Principals add the business translation: what it costs per user, how that changes with growth, and which of the resulting decisions are hard to reverse. The arithmetic is identical at all three levels; what differs is what the number is used for.
Why "Storage" separates levels
Storage is the one resource that only grows. A Senior estimate is a snapshot; a Staff estimate includes replication and secondary copies; a Principal estimate is a curve with a lifecycle policy attached, because the cost of not deciding retention compounds every month.
Staff Sentence Templates#
"That's roughly [N] per second at peak, which is [well under / close to] what one [component] handles, so [decision]."
"The number that matters here is [storage / egress / connections], not QPS — so I'll spend the deep dive on [area]."
"If I'm off by 10× on [input], [decision] flips from [A] to [B]; everything else holds."
"At list prices that's about $[X] a month, or [Y] cents per active user — and it grows [faster / slower] than usage because [reason]."
Common Interview Traps#
- Estimating without concluding. Every number needs a "so…".
- Using averages for sizing. Size for peak; name the peak factor.
- Forgetting replication and secondary copies. ×3 replicas, plus search, analytics, and backups.
- False precision. 11,574 QPS signals arithmetic, not judgment. Say ~12K.
- Ignoring bandwidth. Small QPS × large responses = a network problem.
- Sizing to 100% utilization. Queueing wrecks p99 above ~70%, and you must survive an AZ loss.
- Treating storage as a snapshot. It compounds; state year 1 and year 3.
- Spending 10 minutes on estimation. Two to five minutes, then design.
Practice Drill#
Prompt: "We're launching a feature that stores every user's full activity history — every click — for personalization. We have 30M DAU. Estimate it and tell me if we should build it."
Staff Answer
Assume 30M DAU × ~500 events per active day (clicks, views, scrolls instrumented at a moderate granularity) ≈ 15B events/day → ~175K events/s average, ~450K/s peak at a 2.5× diurnal factor. At ~200 bytes per event (IDs, timestamp, type, a few attributes) in a compressed columnar format ~50 bytes, that's ~3TB/day raw or ~0.75TB/day compressed; ×3 replicas for the online store, ~2.2TB/day provisioned, ~800TB/year. Ingest at 450K/s peak is a partitioned log (Kafka-class, ~90MB/s at 200B/event — a handful of brokers) into a streaming job and then two sinks: an online feature store that keeps only aggregates the model actually reads (last 100 events, per-category counts over 7/30 days — ~KBs per user, ~100GB total, fits in a small KV cluster) and an offline lake in object storage for training (~270TB/year compressed, ~$6K/month in year 1 at standard pricing, less with tiering). So the answer to "should we store every click online forever" is no: the online store needs aggregates, not raw history, and raw history belongs in cheap object storage with a retention limit (e.g., 13 months) that privacy/legal must sign off on. Owners: the data platform team owns ingest and the lake; the personalization team owns which features are materialized online; privacy owns retention.
Why this is L6:
- Converts DAU into rate, volume, and replicated storage with stated assumptions.
- Uses the numbers to split the design (online aggregates vs offline raw), rejecting the literal requirement.
- Names retention as a decision with a sign-off owner, not a default.
What L7 adds:
- States the unit cost (roughly $15–30K/month all-in for ingest, streaming, and storage — about $0.0005–0.001 per DAU per month) and asks product for the expected lift in engagement to compare against it.
- Flags retention of raw behavioral data as a one-way door with regulatory exposure, and proposes an org-wide event retention standard rather than a per-feature decision.
- Checks whether an existing company event pipeline can absorb this instead of building a second one — the cheapest system is the one you don't build.
Where This Appears#
- URL Shortener — Key-space math, read:write ratios, and why writes don't need sharding on day 1
- Chat Messaging — Message rates, connection counts per host, and storage growth
- Blob Storage — Storage tiering, request costs, and egress economics
- Metrics & Monitoring — Sample rates, cardinality, and compression as the design
- Auto-Scaling & Capacity — Headroom multipliers, forecasting, and commitment decisions
- News Feed — Fan-out math and working-set sizing
- Web Crawler — Pages/day, bandwidth, and politeness budgets
Related Foundations & Patterns: Network Latency & Protocols · Caching Fundamentals · Build vs Buy Framework · Handling Large Blobs
Related Technologies: Redis · PostgreSQL · Apache Kafka