Hiring BarSupport

Design a Distributed Cache — Staff-Level Case Study

Case study88 min read11 diagrams

Technologies referenced in this case study: Redis · PostgreSQL · Apache Kafka · DynamoDB

Related: Caching Fundamentals · Consistent Hashing · Scaling Reads · CDN & Edge Caching · Replicated Data Store · Multi-Region Active-Active · Circuit Breakers · Degraded Mode Framework · Consistency Models

How to Use This Case Study#

This case study assumes you already know what cache-aside and LRU are — Caching Fundamentals covers those. Here the subject is the cache as a distributed system you operate: its correctness contract, its failure modes, and the database it is quietly holding up.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → 30-Second Cheat Sheet → Interview Walkthrough → Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → Deep Dives 2 and 4
Deep Dive3+ hrsEverything, including Section 11 (Principal Lens) and Appendices A (invalidation mechanics), C (coordination) and G (multi-tenancy and cost)
What is a Distributed Cache? — Why interviewers pick this topic

A distributed cache is a fleet of memory-resident key-value servers — Memcached, Redis, or something built on them — that stores copies of data whose source of truth lives elsewhere: a database, a service, an expensive computation. Keys are spread across nodes (usually by consistent hashing in the client or a proxy), reads try the cache first, and some mechanism keeps the copies from drifting too far from the truth.

The hard part is not storing bytes in RAM. The hard part is that the cache is a second copy of your data with weaker guarantees, and over time the rest of the system starts depending on it in ways nobody wrote down: the database is sized for the miss rate, product features assume reads are 1ms, and a stale copy can outlive the fix that should have replaced it.

Before vs After — the "cache cluster rolling upgrade" scenario:

Without load-bearing awareness (cache treated as optional):
t=0:       Ops starts a rolling upgrade of the 12-node cache cluster, 1 node per 2 minutes.
           Each restarted node comes back empty. Hit rate 97% → 89% after 3 nodes.
t=+6min:   DB read QPS: 9K → 36K. DB was provisioned for 15K. p99: 4ms → 900ms.
t=+8min:   App threads block on DB. API p99 → 6s. Health checks fail; app pods recycle,
           losing their in-process caches too. Hit rate 81%.
t=+10min:  DB connection pool exhausted. Full outage of product pages.
t=+45min:  Upgrade paused, traffic shed at the edge, DB recovers, cache refills over 30 min.

With the cache treated as a capacity tier:
t=0:       Upgrade tooling checks DB headroom: at 97% hit rate the DB can absorb loss of
           at most 1 node (≈ +8% of reads) safely. Upgrade proceeds 1 node at a time.
t=+0:      Before restart, node N's hot keys are pre-copied to its replacement (warm handoff);
           reads for N's keyspace go to the replica pool during the swap.
t=+3min:   Hit rate dips 97% → 95.8%, recovers within 90s. DB peak 13K QPS, inside budget.
t=+40min:  12 nodes upgraded. No customer-visible impact. Tooling records the minimum hit
           rate and peak DB QPS for the change log.

Why interviewers reach for this question: Every candidate can put Redis in front of a database. Very few can say what the cache promises — how stale, for how long, with what bound — or what happens to the database the day the cache forgets everything. The question separates candidates who have used a cache from those who have been on call for one.

Mechanics Refresher: Fill, Write and Eviction Strategies
StrategyHow It WorksProsCons
Cache-aside (lazy loading)App reads cache; on miss, reads DB and populates cache; on write, updates DB then deletes the cache keyOnly caches what is read; cache failure degrades to DB reads; simplestMiss storms on cold start or hot-key expiry; a delete/fill race can pin stale data
Read-throughCache layer itself loads from the source on missFill logic in one place; easier to coalesce concurrent missesCache layer needs to know data sources; harder to evolve per service
Write-throughWrites go to cache and DB synchronouslyCache always holds latest written value for written keysWrite latency includes both; caches data that may never be read; partial failures leave them inconsistent
Write-behind (write-back)Writes go to cache; flushed to DB asynchronouslyVery fast writes; absorbs write burstsData loss on cache node failure; the cache becomes the source of truth for a while
Refresh-aheadEntries near expiry are refreshed in the background before they expireHot keys never missRefreshes data nobody reads anymore; needs access tracking
CDC-driven invalidationA pipeline tails the DB's change log and deletes or updates affected keysCatches every write path, including ones that bypass the appAdds a pipeline with its own lag and failure modes
EvictionBehaviorWhen It Fits
LRUEvict least recently usedGeneral purpose; vulnerable to scans that flush the hot set
LFU / TinyLFU-style admissionTrack frequency; admit new items only if likelier to be reused than the victimSkewed workloads; resists one-time scans
TTL-onlyItems expire by ageWhen freshness, not memory, is the binding constraint
Slab / size-class aware (Memcached)Memory divided into size classesPredictable allocation; needs rebalancing when value sizes shift

For most production systems: cache-aside with a TTL as a safety net and explicit invalidation on writes, with concurrent-miss coalescing in the client. Add CDC-driven invalidation when more than one service writes to the same data. Avoid write-behind unless you are prepared to treat the cache as a database. The strategies are a 2-minute topic; the interview is about the staleness contract and the database behind it.


Executive Summary

If you only read one section, read this. The rest of the case study expands these pages.

What This Interview Actually Tests#

A distributed cache is not a performance question. Everyone knows RAM is faster than disk.

It is a dependency-and-correctness question: you are creating a second copy of data with weaker guarantees, and the system will start depending on that copy in ways that are invisible until it disappears. It tests:

  • Whether you know which job the cache is doing — trimming latency, carrying load the database cannot, or avoiding expensive recomputation — because each one fails differently
  • Whether you can state a staleness contract (how wrong, for how long, who agreed) instead of "eventually consistent"
  • Whether you know the database is now sized for the miss rate, and design for the day the hit rate drops
  • Whether you can keep the cache correct when there are many writers, many readers and an unreliable network in between

The key insight: The moment a cache absorbs enough reads that the database could not serve them alone, it stops being an optimization and becomes a capacity tier with no durability. From then on, cold starts, node loss and hot keys are availability events for the database. Staff candidates notice the crossover point and design warmth, headroom and miss-path protection as first-class features.

The L5 vs L6 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Put Redis in front of Postgres with cache-aside and a TTL"Asks "Is this cache trimming latency, or is it carrying load the DB can't? What staleness can product accept, and on which fields?"Asks "How many teams run their own cache clusters, and which databases would fall over if those clusters were flushed tonight?"
Consistency"TTL of 5 minutes, and we delete on write"States a staleness budget per data class, names the delete/fill race, and uses versioned sets or lease tokens so a stale fill can't land after an invalidationSets an org rule: any cache of user-visible mutable data declares its staleness bound and invalidation source in a registry; audits it
Failure"Redis has replicas""Cache loss is a DB capacity event. At 95% hit rate, losing 1 of 10 nodes adds +10% of reads to the DB — fine; losing all of them is 20× — so we shed at the edge and warm before we admit"Prices headroom: DB capacity to survive full cache loss vs cost of cache redundancy, and decides per tier with the DB owners
Hot keys"Shard by key with consistent hashing""Consistent hashing spreads keys, not load. Hot keys get an in-process near cache with a 1–5s TTL, or replicated copies across shards"Makes hot-key detection and near-caching a feature of the shared client library, so every team gets it without designing it
OwnershipThe service team runs its RedisCache cluster owned by the service; invalidation pipeline owned by the data's writer; DB owners informed of the miss budgetRedraws ownership: caching platform owns clusters, clients and tooling; data owners own invalidation contracts; DB owners own headroom policy
Scale"Add more nodes""Scale by working-set size and per-node network — a 25 Gbps NIC caps a node near ~1.5M small GETs/s; hot keys hit that ceiling first"Plans the 3-year path: per-service clusters → shared platform with tenancy → regional pools with cross-region invalidation
Why "consistency" separates levels

L5: "On write we update the database and delete the cache key. Next read repopulates it. TTL is a safety net." This is the standard, correct starting point. It has a known race: reader A misses, reads the old value from the DB (or a lagging replica), stalls; writer B commits the new value and deletes the key; reader A then sets the old value. The stale entry now lives for the full TTL — 24 hours if someone chose a long one.

L6: Names the race and closes it. Options: a lease or version token issued on miss, which the delete invalidates so the late set is rejected; a versioned set (SET key value IF version > current); or a short delete-delay (delete again 1–2 seconds after the write, longer than replica lag). Then sets a staleness budget per data class: prices and inventory 5 seconds, profile names 5 minutes, recommendations 1 hour. "Each budget is a product decision. I'd write them down next to the cache key schema."

L7: Recognizes that the race exists in every team's cache, and that most teams will not close it themselves. "I'd put the lease-token protocol in the shared cache client so the safe pattern is the default. Then the question for each team is only 'what's your staleness budget', not 'did you implement leases correctly'."

Why "failure" separates levels

L5: Treats a cache outage as a performance degradation: "If Redis is down, we read from the DB; it's slower but correct."

L6: Does the arithmetic first. At a 95% hit rate, the DB serves 5% of reads; with no cache it serves 100% — 20× its normal read load. If the DB is provisioned at 2× normal, a full cache loss is a DB outage. So the cache-down path needs a plan: circuit-break the cache client fast (≤ 50ms timeouts), cap concurrent DB fills per key (coalescing), shed or serve degraded responses for non-critical reads, and warm the cache before restoring full traffic.

L7: Turns it into a posture decision with a price. "We can make the DB survive total cache loss — roughly 10× more DB capacity, call it +$80K/month — or make total cache loss very unlikely with replicas across 3 AZs, warm standbys and slow, guarded rollouts, for about +$15K/month. I'd pick the second and pay for a game day each quarter that proves the shed path works."

Why "hot keys" separates levels

L5: Consistent hashing with virtual nodes spreads keys evenly. Correct — and irrelevant when one key receives 30% of all reads. A key requested 600K times per second with a 2 KB value needs ~1.2 GB/s — about 10 Gbps — from a single node.

L6: Detects hot keys (client-side sampling, or the server's hot-key reporting), then serves them from an in-process near cache with a 1–5 second TTL, which cuts remote traffic for that key by the number of app instances times the requests per instance per TTL. For keys that must be fresher, replicates the key under N suffixes (key#0..key#7) across shards and reads a random one.

L7: Recognizes the pattern repeats every launch, every sports final, every viral post. Builds it into the platform client: automatic hot-key detection, automatic near-caching above a threshold, with per-namespace opt-outs for data that cannot tolerate even 1 second of staleness.

The Staff Positions#

PositionRationale
Name the cache's job before choosing its designLatency shield, load shield and compute shield have different failure postures and different owners
Every cached data class has a written staleness budget"Eventually consistent" is not a requirement; "≤ 5s for price, ≤ 5 min for display name" is
Close the delete/fill race with leases or versionsTTL alone means a stale entry can live for the full TTL after the write that should have replaced it
Hit rate is a capacity SLO once the DB depends on itA drop from 95% to 90% doubles DB read load; alert on it like a capacity metric
Coalesce concurrent misses per keyOne fill per key per node turns a 10,000-request miss storm into a handful of DB reads
Near-cache the hot keysIn-process caches with second-scale TTLs solve hot keys more cheaply than any resharding
Invalidation follows the writer, not the readerThe team that writes the data owns the invalidation path — ideally CDC, so no write path can skip it

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Latency shield (DB could serve all reads, slower)p99 read latency; modest staleness toleratedCache-aside, moderate TTLs, small cluster; cache loss = slower, not downSilent staleness; latency regressions when hit rate driftsStaleness within budget; p99 SLO met at the expected hit rate
Load shield (DB is sized for the miss rate)DB capacity; hit rate must stay above a floorReplicated cluster across AZs, miss coalescing, warm handoffs, guarded rollouts, shed path when cache is lostCold start / node loss → DB overload → outage; hot-key node saturationHit rate ≥ floor (e.g., 94%) as an SLO; DB survives loss of any single node
Compute shield (results of expensive work: rendered pages, ML features, aggregations, model outputs)Recompute cost ($ and seconds); entries large; staleness often generousLong TTLs, refresh-ahead, size-aware eviction, sometimes persistence (SSD-backed cache)Recompute storm after eviction or deploy; cost spikes; serving results computed by an old model/versionCost per hit; bounded recompute rate; version-tagged entries

🎯 Staff Move: "I'll assume this is a load shield: the read path is around 500K reads per second and the database is provisioned for the misses, not the total. That's the version where the cache's failure modes become the database's outages, so it's where the interesting tradeoffs are. If it's only trimming latency, the design gets simpler and I'll say which parts I'd drop."

The Five Fault Lines#

#Fault LineThe Tension
1Who Keeps the Cache HonestTTL-only vs application-driven deletes vs CDC-driven invalidation — simplicity vs coverage of every write path
2Staleness Budget vs Hit RateLonger TTLs raise hit rate and protect the DB; shorter TTLs reduce wrongness — who decides per data class?
3Optional vs Load-BearingIs the cache an accelerator the system can lose, or a tier the DB can't live without — and is that posture deliberate?
4Where the Bytes LiveIn-process near cache vs remote cluster vs both — hot-key relief and latency vs coherence and memory duplication
5Shared Cluster vs DedicatedCost efficiency and one platform vs blast radius and eviction interference between tenants

In the Wild: Real Production Systems#

Why this section belongs here: These designs are publicly documented, and each one solves a problem interviewers will steer you toward.

Facebook — Memcache at Global Scale#

Facebook's NSDI 2013 paper, Scaling Memcache at Facebook, describes a look-aside cache serving billions of requests per second. Several mechanisms from it are now standard vocabulary: leases — a token handed out on a miss that both throttles concurrent fills (thundering herd) and rejects a set that arrives after the key was invalidated (stale sets); mcrouter, a proxy that routes, batches and replicates requests; a gutter pool of spare servers that absorbs traffic for failed nodes so their load doesn't hit the database; and invalidations driven by tailing the database's commit log and broadcasting deletes, rather than trusting every application write path to delete correctly.

Staff insight: Leases solve two of the hardest interview follow-ups — miss storms and the delete/fill race — with one mechanism. Naming them and explaining both effects is a strong signal.

Netflix — EVCache, Replicated Across Zones#

Netflix built EVCache on Memcached, with a client that writes each item to a copy of the cache in every availability zone and reads from the local zone, falling back to another zone on a miss or failure. Netflix has described it as handling very large request volumes across regions, with cross-region replication for some use cases and warming tools to fill new clusters from existing ones before they take traffic.

Staff insight: Zone-replicated writes trade memory (one copy per zone) for the ability to lose a whole zone's cache without a miss storm against the database. That's the load-shield posture made explicit — pay for redundancy so the database never sees a cold cache.

Uber — CacheFront, Integrated Cache With CDC Invalidation#

Uber has written about CacheFront, a caching layer integrated into its Docstore database's query engine, serving tens of millions of reads per second from Redis. Rather than relying on each service to invalidate correctly, it uses Docstore's change-data-capture stream to invalidate or refresh cached rows after writes commit, and it measures cache staleness directly by comparing samples of cached values against the database.

Staff insight: Two ideas worth stealing: put invalidation where every write passes (the change log), and measure staleness instead of assuming it. "I'd sample 0.1% of cache hits, compare against the DB asynchronously, and alert on the mismatch rate."

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Cache-aside with a TTL""A user updates their name and still sees the old one 10 minutes later. Why?"Do you know the delete/fill race and replica lag?
"We delete on write""Three services write to this table. Do all of them delete? What about the nightly backfill script?"Invalidation coverage; CDC
"We get a 95% hit rate""The cache cluster restarts. What does the database see?"Load-bearing awareness
"Consistent hashing across 20 nodes""One key gets 40% of reads during a live event. Which node dies?"Hot keys vs key distribution
"Redis has replicas""The primary fails over. What happens to writes in flight and to the cache's contents?"Replication semantics, async loss
"TTL of 1 hour""Who chose an hour? What's wrong if the data is 59 minutes stale?"Staleness budget as a product decision
"We share the platform Redis""Another team's batch job fills it with 200 GB of new keys. What happens to your hit rate?"Multi-tenant eviction interference

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: Reads go near cache → remote cache → read replicas, with at most one in-flight fill per key per pod. Writes never touch the cache directly; the commit log drives invalidation, so every write path — services, migrations, admin scripts — is covered. The spare pool keeps a failed shard's traffic off the database. The two numbers on the dashboard that matter most: db.read_qps against its budget (the cache's real SLO) and cache.stale_sample_rate (whether the staleness contract holds).

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Pattern"Cache-aside""Cache-aside with leases for fills, TTL as a backstop, CDC for invalidation once there's more than one writer."
Consistency"Eventually consistent, TTL 5 min""Staleness budget per data class, signed by product. Delete/fill race closed with lease tokens."
Cache down"Fall back to the DB""At 95% hit rate that's 20× DB load. Breaker in 50ms, coalesce fills, shed non-critical reads, warm before reopening."
Hot keys"Consistent hashing""Hashing spreads keys, not load. Near-cache hot keys for 1–5s, or replicate under N suffixes."
Sizing"Enough RAM for the data""Size for the working set at the target hit rate, plus per-node network: one hot key can saturate a NIC."
Multi-tenant"Shared Redis""Shared clusters need per-namespace memory quotas or separate pools — otherwise one batch job evicts everyone's hot set."
Multi-region"Replicate the cache""Each region caches locally; invalidations follow the DB's replication, applied after the regional replica has the write."
Measure"Hit rate""Hit rate per namespace, DB QPS vs budget, invalidation lag, and sampled staleness."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
In-process cache read~50–200ns1,000× cheaper than any network hop — the hot-key weapon
Remote cache GET, same AZ~0.2–0.5ms p50, 1–2ms p99The baseline the DB is being compared against
DB indexed point read~1–5ms; complex queries 10–100ms+The miss penalty — and why miss cost matters as much as miss rate
Single Redis shard~100K–200K ops/s (more with pipelining)Single-threaded command execution caps a shard
Memcached on a large multi-core node~1M+ simple ops/sMultithreaded; the NIC often saturates first
25 Gbps NIC with 2 KB values~1.5M GETs/s ceilingOne hot key can own a node's entire network
Hit rate 95% → 90%DB read load doublesMiss rate, not hit rate, is the number the DB feels
Hit rate 99% → 98%DB read load doubles againEvery "nine" of hit rate halves DB load
Redis per-key overhead~50–100 bytes beyond the value1B small keys = ~100 GB of overhead alone
Managed in-memory cache costroughly $5–15 per GB-month1 TB of RAM ≈ $5–15K/month — memory is the budget line
Invalidation lag target< 1s p99 for user-visible mutable dataBeyond a few seconds users notice their own edits missing
Near-cache TTL for hot keys1–5sBounds staleness while cutting remote load by orders of magnitude

Interview Walkthrough

Forty-five minutes, one cache, and a database you must not hurt. The timing below puts most of the depth on staleness and on the miss path, where interviewers look for judgment.

Phase 1: Requirements & Framing (2–3 minutes)#

Say this, close to verbatim:

"A cache can be doing one of three jobs: trimming latency on reads the database could serve anyway, carrying read load the database can't, or saving expensive recomputation. They fail differently — in the first, losing the cache makes us slower; in the second, it takes the database down. So: how much read traffic, how much of it can the database serve on its own, and how stale can each kind of data be?"

Then commit to numbers:

QuestionAssumption I'll StateWhy It Matters
Read traffic500K reads/s peak, 200K average; 50:1 read/writeSets cache throughput and the miss budget
DB capacityRead replicas provisioned for ~60K reads/sAt 500K reads/s, the cache must hit ≥ 88%; I'll target 94% to leave 2× headroom
Data~2B entities; working set (read in last hour) ~300M items, avg 1.5 KB ≈ 450 GB + overheadCluster sizing from working set, not total data
StalenessPrice/inventory ≤ 5s; profile/display ≤ 60s; recommendations ≤ 1hDifferent invalidation strategies per class
Read-your-writesUsers must see their own edits immediatelyNeeds a session-level bypass or write-through for the author
RegionsOne region, three AZs; second region next yearInvalidation design must extend across regions

🎯 Staff Move: "Here's the number that matters: the DB can serve about 60K reads a second, and we have 500K. So the cache is load-bearing — below an 88% hit rate we're down. I'll treat hit rate as a capacity SLO with a floor of 94%, and design cold start, node loss and hot keys around protecting that floor."

Phase 2: Core Entities & API (1–2 minutes)#

Namespace      { name: "product", owner_team, staleness_budget_s, ttl_s, max_memory_gb,
                 invalidation: cdc|app|ttl_only, near_cache: auto|off, value_schema_version }
CacheKey       = "{namespace}:v{schema_version}:{entity_id}[:{variant}]"
CacheEntry     { value, version (source row version or commit LSN), written_at, ttl }
FillLease      { key, token, issued_at }      // issued on miss; invalidated by delete
Invalidation   { namespace, entity_id, source_lsn, emitted_at }

Client API the services see — note what's missing: there is no set exposed to application code for load-shield namespaces.

get(ns, id, loader, opts) -> value
    // near cache → remote → (lease) → loader() → versioned set
    // opts: max_staleness_s, bypass_cache (read-your-writes), timeout_ms
get_many(ns, ids, loader_batch) -> map
invalidate(ns, id)          // app-driven; CDC is preferred

"Applications pass a loader instead of calling set themselves. That's how the client can coalesce misses, enforce leases and attach versions — the safe protocol is the only protocol."

Phase 3: High-Level Architecture (≤5 minutes)#

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

"Reads go near cache, remote cache, replicas. Writes go only to the database, and the commit log drives invalidation. Every arrow on the miss path has a concurrency limit, because that's the arrow that kills the database."

Phase 4: Transition to Depth (1 minute)#

"The boxes are standard. Three things decide whether this works in production: how we keep entries from staying wrong — the invalidation path and the delete/fill race; what happens to the database when the cache loses data — cold starts, node loss, hot keys; and how many teams share this cluster. I'd like to start with correctness, then the miss path. Sound right?"

Phase 5: Deep Dives (25–30 minutes)#

Deep dive A — correctness (10 min). Staleness budgets per class; the delete/fill race on a sequence diagram; leases or versioned sets; why CDC beats app deletes when there are many writers; read-your-writes for the author via a short-lived per-user bypass marker.

"The race is: reader misses and reads v1, writer commits v2 and deletes, reader sets v1. With leases, the delete invalidates the reader's token, so its set is rejected. Without them, v1 lives for the full TTL."

Deep dive B — the miss path (10 min). Arithmetic of hit rate vs DB load; miss coalescing per key; cache-down posture (50ms timeouts, breaker, shed non-critical reads, serve stale where allowed); warm handoff for planned node replacement; spare pool for unplanned loss; hot-key detection and near-caching.

Deep dive C — tenancy and cost (5 min). Per-namespace memory quotas or pool separation; eviction interference; cost per GB vs DB cost saved.

Phase 6: Wrap-Up (2–3 minutes)#

"To recap: the cache is load-bearing, so hit rate is a capacity SLO with a floor of 94%, and the miss path is protected by coalescing, breakers and a spare pool. Correctness comes from CDC-driven invalidation, lease tokens against the delete/fill race, and per-class staleness budgets product signed. Hot keys are absorbed by a near cache. Next I'd build sampled staleness measurement, per-namespace quotas for the shared cluster, and the cross-region invalidation path for the second region."

Common Timing Mistakes#

MistakeCostFix
Explaining LRU vs LFU for 5 minutesSpends the depth budget on the least differentiating topicOne sentence: "LRU with scan resistance; eviction isn't the hard part"
Never computing DB load at the miss rateCan't answer "cache restarts" except with "it's slower"Do the arithmetic in Phase 1
Saying "eventually consistent"Interviewer hears "I don't know how stale"Give per-class budgets in seconds
Designing the cluster topology in detail (proxy vs client hashing) before correctnessBox-drawing instead of judgmentOne sentence on topology; depth on invalidation and the miss path
Forgetting writers outside the serviceThe nightly backfill bypasses app deletesMention CDC as soon as you mention invalidation

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Caching is where small, reasonable local decisions accumulate into an unowned global risk. One team adds a cache to fix a slow page; a year later the database has been right-sized to the miss rate; two years later nobody can restart the cluster during business hours. The question tests whether you see the cache as a dependency that changes the system around it: it changes how big the database is, what consistency users see, what an upgrade costs, and who must be in the room when it fails.

It also tests correctness under concurrency with nothing exotic — no consensus, no transactions — just two copies of data and an unreliable ordering between writes, deletes and fills. Candidates who can reason precisely about that race usually reason well about every other distributed race.

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

The Senior path is correct and reactive. The Staff path computes the dependency first — the hit-rate floor — and every later decision protects it.

1.3 The Staff Question That Cuts Through Everything#

"If this cache were empty right now, what would happen — and how wrong is it allowed to be when it's full?"

The first half finds the capacity dependency: if the answer is "the database falls over", every operational procedure — upgrades, failovers, deploys, flushes — must be designed around warmth. The second half forces the staleness contract: if nobody can say "how wrong", nobody can say whether the invalidation design is good enough, and the first stale-price incident will be argued rather than diagnosed.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Intent 1: Latency shield. The database can serve every read, at a cost in latency: say 8ms p50 and 60ms p99 against a 1ms cache hit. The cache improves user experience but nothing breaks without it. These caches are usually small, have moderate TTLs, and can be flushed whenever convenient. The main risk is correctness drift — stale data nobody notices because the cache is "just an optimization". The design can stay simple: cache-aside, TTLs, app-driven deletes.

Intent 2: Load shield. The database was sized — or has been allowed to shrink — to serve the misses. At 500K reads/s and a 60K reads/s database budget, the cache is required infrastructure. Its risks are all availability risks: cold start after a deploy or flush, a shard failure, a hot key exceeding a node's network, a mass expiry of keys that were all written at the same time. The design must include replicas across AZs, miss coalescing, spare capacity for failed nodes, warm handoffs, and a shed path for when all else fails.

Intent 3: Compute shield. The cached thing is expensive to produce: a rendered page (50ms of CPU), a personalized feed (300ms fan-in across 40 services), an ML feature vector, a model response that costs real money per call. Entries are larger (10 KB–1 MB), staleness tolerances are generous (minutes to hours), and the risk is cost — an eviction storm or a version change that forces recomputing millions of entries. Design leans on long TTLs, refresh-ahead for popular entries, size-aware eviction, version-tagged entries so a model change can roll out gradually, and sometimes SSD-backed storage because RAM is too expensive for the working set.

DimensionLatency ShieldLoad ShieldCompute Shield
Loss of cacheSlowerOutageCost spike + slower
Binding constraintStalenessHit-rate floorRecompute cost
Typical TTLMinutesMinutes + explicit invalidationHours, refresh-ahead
ReplicationOptionalRequired (cross-AZ)Often optional; persistence sometimes
Owner of the miss budgetService teamDB owners + service teamService team + finance (cost)
Key metricStale sample ratedb.read_qps vs budgetCost per miss × miss rate

2.2 When NOT to Add a Distributed Cache#

  • When the database can serve it with an index or a replica. A missing index or an N+1 query is a cheaper fix than a cache that must be kept correct forever. Profile first.
  • When reads are not repeated. Long-tail access (each key read once a day) yields hit rates under 30%; you pay for RAM and invalidation and get little. Check reuse distance before building.
  • When correctness can't tolerate any staleness and invalidation can't be made synchronous. Account balances, inventory at checkout, permission checks after revocation — read from the source, or cache with version checks that fall through on doubt.
  • When an in-process cache is enough. If the working set fits in ~100 MB and staleness of a few seconds is fine, a per-pod cache has no network hop and no cluster to operate.
  • When a CDN can do it. Public, cacheable-by-URL content belongs at the edge — see CDN & Edge Caching.

2.3 What the Interviewer Leaves Underspecified#

GapWhy It's Left OpenWhat to Say
Whether the DB survives without the cacheTests load-bearing awarenessCompute DB budget vs read load; state the hit-rate floor
Staleness toleranceTests whether you trade it explicitlyPer-class budgets with a named approver
Who writes the dataInvalidation coverage"If more than one writer — or any batch job — I'll invalidate from the change log"
Read-your-writesUsers notice their own edits first"Author bypasses the cache for that entity for ~10s after a write"
Value sizes and access skewHot keys and memory sizingAsk for p50/p99 value size and top-key share; assume Zipfian
TenancyShared cluster interference"Dedicated pool for load-bearing namespaces; shared for the rest, with quotas"

2.4 Precise Terminology#

TermMeaning in this case study
Working setDistinct keys read within a window (e.g., an hour) — what must fit for a target hit rate
Hit-rate floorThe minimum hit rate at which the DB stays within its budget; a capacity SLO
Staleness budgetMaximum time a cached value may differ from the source of truth, per data class
Delete/fill raceA fill computed from old data lands after the invalidation, pinning stale data until TTL
Lease (fill token)Token issued on miss; required to set; invalidated by a delete; also rate-limits concurrent fills
Miss coalescingOnly one request per key per process (or cluster) fetches from the source; others wait for it
Near cacheIn-process cache in each app instance, usually for hot keys with second-scale TTLs
Warm handoffPre-copying a node's hot keys to its replacement before it takes traffic
Spare poolIdle capacity that absorbs a failed node's keyspace temporarily, so misses don't hit the DB
Invalidation lagTime from DB commit to the cache delete being applied

3. The Five Fault Lines#

3.1 Fault Line 1: Who Keeps the Cache Honest#

Every cached entry is a claim that "this is still true". Something must retract the claim when the source changes.

OptionWhat WorksWhat BreaksWho Pays
TTL onlyZero coupling; trivialEvery write is invisible for up to TTL; long TTLs to protect the DB mean long stalenessUsers (stale reads); product (bug reports)
App-driven delete after writeFast invalidation; simple to startEvery write path must remember — batch jobs, admin tools, migrations, another team's service often don't; the delete/fill raceThe team whose data is stale, discovering another team's write path
Write-through / update-in-placeCache holds the new value immediatelyConcurrent writers can apply updates to the cache in a different order than to the DB; caches unread dataOn-call debugging ordering bugs
CDC-driven invalidation (tail the commit log)Covers every write path; ordered per row; decoupled from app codeA pipeline with its own lag, failures and owner; needs a mapping from row changes to cache keysData-platform or caching team owns the pipeline
Diagram: 3.1 Fault Line 1: Who Keeps the Cache Honest

Staff default: cache-aside with lease tokens on fill, TTL as a backstop (sized to the staleness budget × ~10, not as the primary mechanism), and CDC-driven deletes once there is more than one writer or any out-of-band write path. App-driven deletes remain as a fast path for the author's own writes.

When to deviate: for data written by exactly one service with no batch paths, app-driven deletes plus leases are enough — CDC is a pipeline you'll have to run. For compute-shield caches, version-tag entries (model_v17) instead of invalidating: a new version is a new key, and old entries age out.

🎯 Staff Move: "Invalidation should follow the writer, not the reader. If I only delete from the service's write path, the first batch job that updates the table directly will create stale entries nobody can explain. So I'd invalidate from the commit log, and the team that owns the table owns that pipeline's SLO."

3.2 Fault Line 2: Staleness Budget vs Hit Rate#

TTL is the knob that trades wrongness for database protection. Turning it is a product decision disguised as a config value.

OptionWhat WorksWhat BreaksWho Pays
One global TTLSimpleToo long for prices, too short for recommendationsEither users (stale) or DB (misses) — usually both, on different data
Per-class TTL from staleness budgetsEach class gets the freshness it needsNeeds product decisions and a registryProduct signs budgets; platform enforces
Short TTL + refresh-ahead for hot keysFresh and high hit rate for popular keysRefreshes cost DB reads even when nothing changedDB budget
Long TTL + reliable invalidationHighest hit rate; freshness from invalidationInvalidation bugs become long-lived stale dataWhoever owns the invalidation pipeline
Diagram: 3.2 Fault Line 2: Staleness Budget vs Hit Rate

The arithmetic product needs to see. If a key is written once a day and read 10,000 times a day, a 60-second TTL produces ~1,440 misses/day for that key; a 1-hour TTL produces 24. With reliable invalidation, the TTL can be long and the real staleness is the invalidation lag (< 1s). Without it, staleness is the TTL. So the investment in invalidation buys hit rate — that's the argument that funds the CDC pipeline.

Staff default: per-namespace staleness budgets in a registry; TTL derived from the budget when invalidation is TTL-only, or set long (hours) with CDC invalidation and a sampled staleness monitor proving the budget holds.

When to deviate: data that must never be stale in a specific flow — the price at checkout, a permission after revocation — bypasses the cache in that flow, or reads the cached value with its version and confirms the version against the source. Cache it for browsing; verify it at the moment of commitment.

🎯 Staff Move: "I'd ask product to sign three numbers: price can be 5 seconds stale on browse pages and 0 at checkout; display names 60 seconds; recommendations an hour. Those numbers decide the TTLs and where we bypass. Without them we're guessing, and we'll argue about every stale-data bug."

3.3 Fault Line 3: Optional vs Load-Bearing#

The most important property of a cache is whether the system survives its absence. Most teams find out during an incident.

PostureWhat WorksWhat BreaksWho Pays
Optional (DB sized for full load)Cache loss = slower; simple opsExpensive DB capacity sitting idle; drift — someone later downsizes the DB "since the cache handles it"Finance (idle DB capacity)
Load-bearing, undeclaredCheap — until it isn'tCold starts and node loss become DB outages; upgrades are frighteningEvery user during the incident; on-call
Load-bearing, declared and engineeredDB right-sized; cache has replicas, spare pool, warm handoffs, shed pathMore cache cost and engineering; hit rate becomes an SLO with alertsCache platform team (ops burden); +~30–50% cache cost for redundancy
Diagram: 3.3 Fault Line 3: Optional vs Load-Bearing

The math to say out loud. DB load = read QPS × miss rate. At 500K reads/s: 94% hit → 30K/s to the DB; 88% → 60K/s, the limit; 80% → 100K/s, overload. Losing one of 24 shards with no replica and no spare pool drops hit rate by ~4 points → +20K/s. Losing an AZ's third of the cluster with no cross-AZ replicas drops it ~31 points → outage.

Staff default: declare the posture. For load shields: replicas across AZs, a spare pool sized for one failed shard, fill budgets per pod (a token bucket on DB fills, e.g., 200/s per pod × 300 pods = 60K/s ceiling — the DB budget), stale-on-error for classes that allow it, and a shed path for the rest.

When to deviate: if the DB is cheap to over-provision (a managed key-value store such as DynamoDB with on-demand capacity), sizing the DB for full load and keeping the cache optional can be cheaper than engineering a load-bearing cache. Do the price comparison.

🎯 Staff Move: "The fill path gets its own rate limit, sized to the DB's budget. If the cache fails, the worst the application can do to the database is the budget — beyond that, we serve stale or degraded responses. The cache can fail; the database must not fail because of it."

3.4 Fault Line 4: Where the Bytes Live#

OptionWhat WorksWhat BreaksWho Pays
Remote cluster onlyOne copy; invalidation is one delete; memory efficientNetwork hop on every read; hot keys saturate one nodeHot-key nodes; latency-sensitive callers
Near cache only (in-process)~100ns reads; no clusterMemory duplicated per pod (300 pods × N); invalidation must reach every pod; cold after every deployPod memory; correctness if invalidation is TTL-only
Two tiers: near for hot keys, remote for the restHot keys absorbed locally; remote handles the long tailTwo staleness layers; near-cache TTL adds to total stalenessStaleness budget must cover both tiers
Replicated hot keys (key#0..#7 across shards)Fresher than near cache; spreads loadWrites and deletes multiply by N; must know which keys are hotCache client complexity
Diagram: 3.4 Fault Line 4: Where the Bytes Live

Near-cache arithmetic. A key read 600K/s across 300 pods is 2,000 reads/s per pod. With a 2-second near-cache TTL, each pod fetches it once per 2s: 300 pods × 0.5/s = 150 remote reads/s instead of 600K. The cost: up to 2 seconds of extra staleness for that key, and a few KB of memory per pod.

Staff default: remote cluster for everything; near cache only for keys the client detects as hot (above ~1% of a pod's reads or a fixed rate threshold), with a 1–5s TTL that counts against the namespace's staleness budget. Namespaces with zero staleness tolerance opt out.

When to deviate: small, read-mostly reference data (feature flags, currency tables, country lists, under ~100 MB) belongs fully in-process with push-based refresh — no remote tier at all.

🎯 Staff Move: "Consistent hashing gives every key a home; it doesn't stop one key from being popular. For hot keys I'd rather spend 2 seconds of staleness in a near cache than rebuild the cluster topology."

3.5 Fault Line 5: Shared Cluster vs Dedicated#

OptionWhat WorksWhat BreaksWho Pays
One shared cluster for all teamsHigh utilization; one team operates it; cheap per GBEviction interference (one tenant's churn evicts another's hot set); one tenant's hot key or big values hurt everyone; upgrades need every team's sign-offLoad-bearing tenants (their DB absorbs another tenant's evictions)
Dedicated cluster per serviceIsolation; independent upgrades40 clusters at 30% utilization; 40 teams learning cache ops; inconsistent clientsFinance (idle RAM); each team's on-call
Platform-operated pools with tenancy (dedicated pools for load-bearing namespaces, shared pool with per-namespace quotas for the rest)Isolation where it matters, efficiency elsewhere; one operator, one clientNeeds quota enforcement (separate instances, or memory accounting per namespace) and a placement policyPlatform team (operator); tenants accept placement rules

Why quotas are hard in a shared cluster. Redis and Memcached evict across the whole instance; there is no per-prefix memory limit. Real isolation means separate instances (or separate Redis databases on separate processes), not key prefixes. A "quota" enforced by convention is a request, not a limit.

Staff default: load-bearing namespaces get their own pool; latency-shield namespaces share a pool with per-tenant memory budgets enforced by placing tenants on separate instances within it; the platform owns placement, the client, upgrades and dashboards.

When to deviate: at small scale (< 100 GB total, a handful of services), one shared cluster with good dashboards is fine — the interference risk is lower than the cost of operating many clusters.

🎯 Staff Move: "Prefixes aren't isolation. If another team's batch job can evict our hot set, our database's capacity depends on their release schedule. Load-bearing namespaces get their own pool."


4. Failure Modes & Operational Reality#

4.1 The Cold Restart — A Deploy That Emptied the Cache#

t=0:       A cache client library upgrade changes the key serializer (hash of a struct
           now includes a new field). Every key the new pods compute is different.
           Effective hit rate for upgraded pods: 0%.
t=+2min:   Canary at 5% of pods. DB read QPS 30K → 38K. Within budget; canary "healthy".
t=+10min:  Rollout to 50%. DB QPS 30K → 165K. Replica CPU 100%, p99 3ms → 2.4s.
t=+11min:  App pods time out on fills and retry — fills double. Old pods' hits keep half
           the traffic alive. Error rate 22%.
t=+14min:  Page: db.read_qps > budget, api.error_rate > 5%. On-call rolls back.
t=+16min:  Rolled-back pods hit the old keys — still warm. Recovery in ~2 minutes.
Total:     6 minutes of severe degradation. Root cause: a key change is a cache flush.

Detection: cache.hit_rate per deploy version (not just fleet-wide) — the canary's hit rate dropped to 0% at t=+2min, which a per-version panel would have shown; db.read_qps vs budget.

Mitigation: fill-rate budget per pod caps DB load at the budget no matter how cold the cache is; stale-on-error and shed paths take the excess.

Prevention: key schema version is explicit (ns:v3:id) and changing it requires a migration plan: dual-read (new key, fall back to old key, write new), warm the new keyspace from the old, then cut over. Canary analysis includes per-version hit rate with a hard gate (e.g., canary hit rate within 2 points of baseline).

Owner: caching platform (client library and canary gates); service team for its key schema.

4.2 The Celebrity Key — One Key, One NIC#

t=0:       A live event starts. Key event:9921:summary (3 KB) goes from 2K to 450K reads/s.
           It lives on shard 14.
t=+20s:    Shard 14 network egress: 450K × 3 KB ≈ 1.35 GB/s ≈ 10.8 Gbps on a 10 Gbps NIC.
           Packet drops; every key on shard 14 (~4% of keyspace) times out.
t=+40s:    Clients treat timeouts as misses → fill from DB. Coalescing is per pod, so 300
           pods each fill event:9921 and every other shard-14 key. DB QPS +35K.
t=+60s:    Breaker opens for shard 14 in clients. Shard 14's keyspace now all misses.
t=+3min:   On-call identifies the key from the server's hot-key sampling and hand-copies
           it to a near-cache allowlist via config push. Load drops in 30 seconds.

Detection: client-side hot-key sampler (count top keys per pod over 10s windows, report keys above 0.5% of reads); cache.node_network_bytes vs NIC capacity; cache.shard_timeout_rate by shard.

Mitigation: automatic near-caching for keys above a threshold (no human in the loop); replicate the hottest keys under N suffixes across shards; a shard that times out routes to the spare pool rather than straight to the DB.

Prevention: for scheduled events, pre-mark event keys as hot (the product knows when the final starts); size NICs for the hottest-key scenario, not the average; keep large values out of hot keys (split summary from detail).

Owner: caching platform (client auto-near-cache, hot-key detection); event product team for pre-marking.

4.3 The Stale Fill — A Price That Wouldn't Update#

t=0:       Merchant updates price 49.99 → 39.99. Writer commits to primary, deletes key.
t=-40ms:   (Concurrently) A reader missed, read from a replica lagging 300ms → 49.99.
t=+260ms:  Reader sets price=49.99 with TTL 24h, after the delete.
t=+1h:     Merchant reports the sale isn't showing. Support says "cache, try again later".
t=+6h:     340 orders at 49.99 shown, charged at 39.99 (checkout reads the source). Shoppers
           who saw 49.99 didn't buy; the merchant's promotion underperformed.
t=+24h:    TTL expires; price correct. Ticket closed "resolved itself".

Detection: sampled staleness — for 0.1% of cache hits, asynchronously read the source row and compare versions; cache.stale_sample_rate per namespace with alert above the budget (e.g., > 0.01% of samples older than 5s for prices).

Mitigation: delete the key (or the namespace for that merchant) via an admin tool; never "wait for TTL" on a reported stale value.

Prevention: lease tokens so a fill that started before the delete can't land after it; versioned sets (store the source row version; reject sets with older versions); CDC invalidation applied after replica lag — or a second, delayed delete ~1–2s after the first, longer than p99 replica lag. TTL sized to the staleness budget × ~10 as a backstop, not 24h.

Owner: the service team owning price data; caching platform for the lease protocol in the client.

4.4 The Neighbor's Batch Job — Eviction Interference#

Mon 02:00: Analytics team's new backfill writes 280 GB of per-user aggregates into the
           shared cluster (TTL 7 days, 40 KB values). Cluster memory: 85% → 100%.
02:10:     LRU evicts older keys — including the session and catalog namespaces' hot set,
           which were "older" only because they're long-lived.
02:15:     Catalog hit rate 97% → 79%. Catalog DB QPS 6× normal. It's 2 a.m.; DB copes.
08:30:     Morning traffic ramps. Catalog DB saturates. p99 product page 4s.
09:05:     Page. On-call sees hit rate drop; doesn't see why — no per-namespace memory view.
10:20:     Root cause found by scanning key prefixes. Backfill keys deleted. Recovery 11:00.

Detection: per-namespace memory and eviction counts (cache.evictions{namespace}); cache.hit_rate{namespace} with alerts per load-bearing namespace; cache.bytes_written{namespace} anomaly detection.

Mitigation: delete or TTL-shorten the offending namespace; temporarily raise fill budgets for affected services if the DB has headroom.

Prevention: load-bearing namespaces on dedicated pools; shared pools enforce per-tenant memory by instance placement; new tenants above 10 GB require a capacity review; large-value bulk loads go to a compute-shield pool (or an SSD-backed store), not the hot-path cluster.

Owner: caching platform (placement and quotas); analytics team for their backfill; the incident review includes both.

4.5 Invalidation Lag — The Pipeline Nobody Watched#

t=0:       A schema change adds a column to the orders table. The CDC consumer's
           deserializer fails on the new column type for 0.3% of rows and retries forever
           on the first bad record — partition 7 stalls.
t=+5min:   Invalidations for 1/32 of entities stop. No errors user-side.
t=+2h:     Support tickets: "I changed my address and it shows the old one" — read-your-
           writes bypass covers the author for 10s, then the stale entry returns.
t=+3h:     Engineer notices consumer lag on partition 7 = 3h. Skips the record, lag drains.
           Entries written in the gap remain stale until TTL (6h).

Detection: inval.lag_ms per partition with alert at p99 > 5s; inval.consumer_errors; sampled staleness (4.3) catches it independently of the pipeline's own metrics.

Mitigation: dead-letter poison records after N retries instead of blocking the partition; on lag recovery, issue a namespace-wide (or time-window) bulk invalidation for entities written during the gap.

Prevention: schema changes run through the CDC consumer's contract tests; the pipeline has an owner and an SLO (p99 < 1s); DLQ has an owner who reviews within a business day.

Owner: the data's writer team owns the invalidation pipeline's correctness; the platform owns the CDC infrastructure.

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Cold start (deploy, flush, key change)cache.hit_rate by version; db.read_qps > budgetDB and every readerFill budgets; warm before cut-over; rollbackCaching platform + service
Hot keyHot-key sampler; node NIC > 70%All keys on that nodeAuto near-cache; key replication; spare poolCaching platform
Stale fill racecache.stale_sample_rateIndividual entities, often high-value (prices)Leases / versioned sets; delayed second deleteData-owning service
Eviction interferencecache.evictions{namespace}, per-namespace hit rateLoad-bearing tenants of a shared clusterDedicated pools; placement quotasCaching platform
Invalidation laginval.lag_ms p99 > 5sEvery entity in stalled partitionsDLQ poison records; bulk invalidate after gapWriter team + data platform
Shard failurecache.shard_timeout_rate1/N of keyspaceReplica promotion; spare poolCaching platform
AZ lossAZ-level error rate~1/3 of cache if not cross-AZ replicatedCross-AZ replicas; fill budgetsCaching platform + SRE
Mass expiry (synchronized TTLs)Periodic miss spikes aligned to load timeDB spikes every TTL periodTTL jitter ±10–20%Service team
Memory fragmentation / big valuesmem_fragmentation_ratio > 1.5, value size p99Node OOM, evictionsValue size limits (e.g., 512 KB); defragCaching platform

🎯 Staff Insight: The cache's loud failures — a node dies, the hit rate drops — are the easy ones. The expensive ones are quiet: a stale price that "resolves itself" after 24 hours, an invalidation partition stuck for three hours. "I don't trust a cache whose staleness I'm not measuring. Sampled comparison against the source is the only metric that tells me the contract holds."


5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framing"Add Redis with cache-aside"Names latency/load/compute shields; computes hit-rate floor from DB budget; commitsAsks which databases across the org are already sized for cache hit rates, and who knows
CorrectnessTTL + delete on writePer-class staleness budgets; delete/fill race closed with leases or versions; CDC for multiple writers; read-your-writes for authorsStaleness contracts in a registry; lease protocol in the shared client; sampled staleness as an org-wide metric
Failure"Fall back to DB"Hit-rate arithmetic; fill budgets; stale-on-error; spare pool; warm handoffsPrices DB headroom vs cache redundancy per tier; runs cold-cache game days
Hot keysConsistent hashingDetection + near cache + key replication; NIC mathAutomatic in the platform client; event pre-marking process with product
Operations"Monitor hit rate"Per-namespace hit rate, db.read_qps vs budget, invalidation lag, stale sample rate, ownersCache SLOs tied to DB capacity plans; upgrade tooling with hit-rate gates
OrganizationTeam runs its own RedisWriter owns invalidation; platform owns client; DB owners know the miss budgetCaching platform with pools and placement; namespace registry; retires per-team clusters

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Computes the dependency"The DB can do 60K reads; we have 500K. Below 88% hit rate we're down. I'll target 94%."
States staleness in seconds"Price 5 seconds on browse, zero at checkout; display names 60 seconds."
Names the race precisely"Reader reads v1 from a lagging replica, writer deletes, reader sets v1. Leases reject that set."
Protects the miss path"Fills have their own budget sized to the DB. A cold cache can't push the DB past it."
Separates key spread from load spread"Hashing spreads keys, not popularity. Hot keys go to a near cache."
Measures correctness"I'd sample 0.1% of hits against the source and alert on staleness beyond budget."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Cache everything with a 24-hour TTL"No staleness model; stale data lives a day after any race
"Write-through keeps it consistent"Ignores concurrent writers applying out of order and partial failures
"If the cache dies we read from the DB" with no arithmeticDoesn't see that the DB may be sized for the miss rate
"Use Redis transactions to keep cache and DB consistent"Redis can't participate in a DB transaction; shows a mental model gap
No mention of writers outside the serviceMisses the most common real-world source of stale data
Spends the interview on eviction policiesDepth on the least differentiated topic

5.4 Common False Positives#

  • Redis internals fluency ≠ cache design. Knowing about RDB vs AOF, skiplists and cluster slots is useful; if it doesn't lead to a staleness contract and a miss-path plan, it's trivia.
  • "Cache invalidation is one of the two hard problems" ≠ insight. Quoting the joke is not the same as naming the race and closing it.
  • High hit-rate targets ≠ good design. "We'll get 99.9%" without asking about reuse distance and working-set size is optimism, not engineering.
  • Many cache tiers ≠ sophistication. Browser + CDN + near + remote + DB buffer pool, each with its own TTL, often means total staleness nobody can compute.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minThree intents; DB budget vs read load; hit-rate floor; staleness classes
Entities & API3–6 minNamespace registry; key schema with version; loader-based get; leases
Architecture6–10 minNear → remote → replicas; CDC invalidation; spare pool
Correctness10–20 minDelete/fill race; leases/versions; CDC; read-your-writes
Miss path20–30 minCoalescing; fill budgets; cache-down posture; hot keys
Tenancy and cost30–34 minPools vs shared; sizing; $ per GB vs DB saved
Pivot34–42 minMulti-region, hot event, key schema migration, compute cache
Wrap42–45 minFloor, contract, owners, next steps

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Shape
"Make it multi-region"Cross-region invalidation and replica lagRegional caches; invalidate after regional replica applies the write; never cross-region reads on the hot path
"A celebrity posts — one key gets 500K reads/s"Hot keys vs hashingNIC math; near cache; key replication; pre-marking
"Users complain they don't see their own edits"Read-your-writesAuthor bypass marker; or write-through for the author's session
"We're changing the value format"Key migration without a cold startVersioned keys; dual-read; pre-warm; gated cutover
"Cut cache cost by 40%"EconomicsWorking-set analysis; compress values; evict low-value namespaces; tier to SSD-backed cache
"How do you know it's correct?"Measurement cultureSampled staleness against the source; per-namespace budgets

6.3 What to Deliberately Skip#

  • Eviction algorithm details — "LRU with scan resistance" and move on.
  • Redis vs Memcached feature comparison — one sentence: data structures and replication vs multithreaded simplicity.
  • Persistence (RDB/AOF) tuning — for a cache of a DB, persistence is usually off or replica-only; say why.
  • Building your own cache server — use Redis or Memcached; the design is in the client and the operating model.
  • Distributed transactions between cache and DB — they don't exist in practice; say so and move to leases and CDC.

6.4 Follow-Up Questions to Expect#

  1. "A user updates their profile and sees the old one. Walk me through every way that can happen."
  2. "The cache cluster restarts. What does the database see in the first 60 seconds?"
  3. "How do you size the cluster? What determines the number of nodes?"
  4. "One key gets 30% of all reads. What happens and what do you do?"
  5. "Three services and a nightly batch job write to the same table. How is the cache invalidated?"
  6. "How do you change the cached value's format without an outage?"
  7. "How do you know your cache is serving correct data right now?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Our product catalog service is slow and the database is struggling. Design a distributed cache for it."

Staff Answer

"First, which job is the cache doing? If the database can serve all reads and we want lower latency, that's one design. If the database can't serve the load, the cache becomes required infrastructure and its failures become the database's outages. I'll assume the second: 500K catalog reads/s at peak, the replicas can do about 60K, so we need at least 88% hits; I'll target 94% for headroom.

Before topology I want staleness budgets from product: price and stock on browse pages maybe 5 seconds — and checkout reads the source; descriptions and images, minutes. Then: cache-aside through a shared client with leases, invalidation from the database's change log because catalog data has several writers, a near cache for hot products, fill budgets so a cold cache can't overload replicas, and cross-AZ replicas. I'll go deep on correctness and on the miss path."

Why this is L6:

  • Classifies the cache's job and derives the hit-rate floor from numbers
  • Gets staleness budgets per field before drawing anything
  • Plans depth on the two places production caches fail

What L7 adds:

  • Asks why the database is struggling — a missing index or a query pattern may be the cheaper fix
  • Asks whether a caching platform exists so this team doesn't operate its own cluster
❌ Common L5 Trap

"I'll put a Redis cluster in front of the database using cache-aside with a 10-minute TTL, consistent hashing across nodes, and replicas for availability. On writes we delete the key."

Why this misses: Everything is reasonable and nothing is quantified. The interviewer's next three questions — "what if Redis restarts?", "why do users see old prices?", "who else writes this table?" — each expose a gap the design didn't consider.


Drill 2: The Hit-Rate Arithmetic#

Prompt: "Your hit rate dropped from 96% to 92%. Is that a big deal?"

Staff Answer

"It's a doubling of database read load — the miss rate went from 4% to 8%. At 500K reads/s that's 20K → 40K reads/s on replicas budgeted at 60K. So we went from 3× headroom to 1.5×. Not an outage today, but one more drop — a node failure, a traffic peak — and we're at the limit. I'd look at which namespace dropped, because a fleet-wide average hides it: a 30-point drop in one namespace looks like 4 points overall. Typical causes: a deploy changed keys, a new tenant is evicting our hot set, TTLs were shortened, or traffic shifted toward the long tail. The alert should be on db.read_qps vs budget and on per-namespace hit rate, not on the fleet average."

Why this is L6:

  • Converts hit rate into miss rate and DB load immediately
  • Expresses risk as remaining headroom
  • Warns against fleet-wide averages and lists likely causes

What L7 adds:

  • Ties the hit-rate floor into the DB capacity plan, so DB downsizing requires checking it
  • Makes "hit rate per namespace" part of every service's SLO dashboard by default

Drill 3: Make It Concrete — Size the Cluster#

Prompt: "How many cache nodes do you need?"

Staff Answer

"Three constraints; the largest wins. Memory: working set for a 94% hit rate — from access logs, the top 300M of 2B items cover ~94% of reads; at 1.5 KB average plus ~100 bytes overhead, that's ~480 GB; with 30% headroom for fragmentation and growth, ~620 GB. On 64 GB nodes with ~50 GB usable, that's 13 primaries. Throughput: 500K reads/s plus fills and deletes, ~550K ops/s; at a conservative 80K ops/s per Redis shard, 7 shards. Network: 550K × 1.5 KB ≈ 825 MB/s ≈ 6.6 Gbps cluster-wide — fine on average, but a hot key can put several Gbps on one node, which the near cache handles. So memory dominates: 13 primaries; I'd round to 16 for even hashing and headroom, each with a replica in another AZ — 32 nodes. Roughly $10–15K/month on a managed service at typical per-GB pricing."

Why this is L6:

  • Sizes from working set and target hit rate, not total data
  • Checks all three constraints and names the binding one
  • Includes replicas, headroom, and an order-of-magnitude cost

What L7 adds:

  • Compares against the DB capacity the cache replaces — 500K reads/s on replicas would need ~8–10× the replica fleet
  • Asks whether values can be compressed or trimmed (often 2–4× on JSON) — the cheapest capacity available

Drill 4: The Cache Is Down#

Prompt: "The entire cache cluster is unreachable. Walk me through the next five minutes."

Staff Answer

"In the first second, client timeouts at 50ms start firing; within ~5 seconds, breakers open per pod, so we stop paying the timeout on every request. Now three behaviors by data class. Classes that allow stale-on-error: serve from the near cache, even past its normal TTL, up to a hard maximum — say 5 minutes for product descriptions — and mark responses degraded. Classes that need fresh data: fill from replicas, but each pod has a fill budget — 200/s × 300 pods = 60K/s, the replicas' budget — so the DB can't be pushed past capacity. Requests beyond the budget get degraded responses: a product page without the recommendations panel, or a 503 with Retry-After for non-critical APIs.

When the cluster returns, I don't snap the breakers closed: half-open with 5% of traffic, warm the hottest keys first from the access-log top-N, then ramp. Without that, the moment it returns empty, every pod fills at full budget for minutes."

Why this is L6:

  • Bounded detection with specific timeouts and breaker behavior
  • Per-class behavior: stale-on-error vs budgeted fills vs degradation
  • Recovery is designed, not assumed — warming before ramping

What L7 adds:

  • Decides the org posture: either the DB tier survives full cache loss, or the cache tier is engineered to make it extremely rare — priced both ways
  • Runs a quarterly game day that empties a cache pool in production during low traffic

Drill 5: The Hot Key#

Prompt: "During a product launch, one item's key receives 40% of all reads."

Staff Answer

"40% of 500K is 200K reads/s on one node; at a 4 KB value that's ~800 MB/s, ~6.4 Gbps — most of a node's NIC, and every other key on that node suffers. Near cache: with 300 pods and a 2-second TTL, remote reads for that key drop to ~150/s. That costs up to 2 extra seconds of staleness; for a product page during a launch that's acceptable — and price at checkout reads the source anyway. If the key's data changes often during the event, replicate it under 8 suffixes across shards; reads pick a random suffix, deletes go to all 8. For a planned launch, product tells us the item ID ahead of time and we pre-mark it hot. The spare pool covers us if the node degrades before detection kicks in."

Why this is L6:

  • Computes per-node bandwidth and the collateral damage to co-located keys
  • Near cache with arithmetic and an explicit staleness cost
  • Pre-marking for planned events — operational foresight

What L7 adds:

  • Builds automatic hot-key detection into the shared client for all teams
  • Creates a launch-readiness process where product events notify the caching platform

Drill 6: Many Writers#

Prompt: "Three services and a nightly batch job update the products table. How is the cache invalidated?"

Staff Answer

"Not from the writers. Four write paths means four places to forget a delete, and the batch job will forget first. I'd tail the database's change log — logical replication or a CDC connector into Kafka, keyed by product ID so changes for one product stay ordered — and an invalidation worker maps each row change to cache keys: product:v3:{id}, plus derived keys like category:{cat}:top if the row affects them. The mapping lives with the products team, because they know which keys a row change affects. The pipeline has an SLO — p99 lag under 1s — and poison records go to a DLQ instead of blocking a partition. If replicas lag, the worker waits until the replica's applied position passes the change's position before deleting, or deletes twice. The writer services can still issue a fast delete for their own writes; CDC is the guarantee."

Why this is L6:

  • Moves invalidation to where all writes pass
  • Handles derived keys, ordering per entity, and replica lag
  • Gives the pipeline an owner, an SLO and a DLQ

What L7 adds:

  • Makes CDC-to-invalidation a platform capability with a per-table mapping config, not per-team code
  • Standardizes "every load-bearing cache of a table invalidates from that table's change stream"

Drill 7: The Shared Cluster#

Prompt: "Twelve teams share one large cache cluster. Two teams' hit rates dropped overnight. What's going on and what do you change?"

Staff Answer

"Most likely eviction interference: someone wrote a lot of new data and LRU evicted the other teams' long-lived hot keys. I'd check per-prefix memory and eviction counts — if we don't have them, that's the first gap. Short-term: identify the namespace, shorten its TTL or delete it, and alert per-namespace on hit rate. Structurally: prefixes aren't quotas — Redis evicts across the whole instance. So I'd split the cluster into pools: the two or three load-bearing namespaces each get dedicated instances; the rest share a pool where each tenant is placed on its own instances with a memory budget. New tenants above ~10 GB go through a capacity review. The platform owns placement; tenants own their namespace's staleness and TTLs."

Why this is L6:

  • Diagnoses eviction interference and the missing visibility
  • Knows prefix "quotas" don't exist in the engine
  • Isolation by placement; a review gate for large tenants

What L7 adds:

  • Prices dedicated pools vs incidents: one evening's interference incident vs the cost of 3 extra pools
  • Introduces chargeback per GB to make tenants think about what they cache

Drill 8: Build vs Buy#

Prompt: "Should we run Redis ourselves on VMs or use a managed cache?"

Staff Answer

"Managed by default. The managed services handle failover, patching, backups and scaling operations that otherwise cost us a part-time on-call rotation. Running it ourselves makes sense when: the cluster is large enough that the managed premium exceeds ~1–2 engineers' cost — a 30–40% premium on $100K/month of cache is real money; we need versions or modules the managed offering doesn't support; or we need control over failover behavior the managed service doesn't expose. Regardless, the parts that matter for correctness — the client library with leases, coalescing and near cache, the CDC invalidation pipeline, the namespace registry — are ours either way. The server is the least differentiated part."

Why this is L6:

  • Defaults to managed and names cost and capability triggers
  • Separates the commodity (server) from the differentiated parts (client, invalidation, registry)

What L7 adds:

  • Builds a 3-year TCO with growth projections, including engineer time and incident cost
  • Keeps the client protocol portable so switching providers is a two-way door

Drill 9: Changing the Value Format#

Prompt: "We need to change the cached product object's serialization format. How do you roll it out?"

Staff Answer

"A format change is a key change, and a key change is a cache flush if done naively. Keys carry a schema version: product:v3:{id} → product:v4:{id}. Phase 1: new code reads v4, falls back to v3 on miss — deserializing v3 and converting — and writes v4. Old code still reads and writes v3. Invalidation deletes both versions during the transition. Phase 2: pre-warm v4 for the top keys from the access log, so the hit rate on v4 starts near the v3 rate. Phase 3: canary with a gate — v4 hit rate within 2 points of v3 and db.read_qps within budget — then roll out. Phase 4: after a full TTL period, stop reading v3 and let it expire. Memory: dual keys temporarily need up to 2× for hot items; I'd confirm headroom first."

Why this is L6:

  • Recognizes format change = cold start without a plan
  • Dual-read fallback, pre-warming and gated canary
  • Accounts for invalidation of both versions and memory headroom

What L7 adds:

  • Bakes versioned keys and dual-read into the client so every format change follows the same path
  • Makes per-version hit rate a mandatory canary gate across all services

Drill 10: Multi-Region#

Prompt: "We're adding a second region with its own replicas. How does caching work?"

Staff Answer

"Each region gets its own cache cluster; no cross-region reads on the hot path — 60–150ms would defeat the point. Writes go to the primary region's database and replicate to the secondary with some lag — say 200ms–2s. The trap: if we invalidate the secondary region's cache as soon as the primary commits, a reader in the secondary region can miss, read its local replica that hasn't applied the write yet, and re-fill the old value. So invalidation for each region must happen after that region's replica has applied the change: drive the secondary's invalidations from the secondary replica's own change stream, or attach the commit position to the invalidation and have the worker wait until the local replica passes it. For read-your-writes after a user writes in the secondary region, route that user's reads for the entity to the primary briefly, or bypass the cache for ~2× replication lag."

Why this is L6:

  • Refuses cross-region cache reads with a number
  • Identifies the cross-region version of the stale-fill race and fixes it with ordering
  • Handles read-your-writes explicitly

What L7 adds:

  • Decides which namespaces even need the second region warm — compute caches may run cold there initially
  • Plans the active-active future: where writes become regional, per-entity home regions keep invalidation tractable

8. Deep Dive Scenarios#

Deep Dive 1: Peak-Traffic Incident — The Midnight Drop#

Context: A limited sneaker release goes live at midnight. Traffic to the product page jumps from 20K to 900K reads/s in 90 seconds. Cache hit rate stays at 97%, yet product pages time out for 40% of users and the catalog database is at 100% CPU. The on-call has added cache nodes, which didn't help. You're pulled in.

Questions to Surface First:

  • If the hit rate is 97%, where is the DB load coming from — which queries, from which callers?
  • Is the release product's key on one node? What is that node's network utilization?
  • Are timeouts being treated as misses — so cache slowness converts into DB load?
  • What changed in the release flow — any keys with short TTLs, like stock counts?

Typical L5 Approach: Adds more cache nodes and more DB replicas. New cache nodes reshuffle part of the keyspace (consistent hashing moves ~1/N of keys), causing a fresh wave of misses; replicas take 20 minutes to provision. The drop is over before they help.

Staff Approach: Finds two interacting problems. The product key (8 KB with all size variants) lives on one node pushing ~9 Gbps; that node's timeouts are counted as misses by clients, so 300 pods fill the same key from the DB repeatedly — per-pod coalescing doesn't help across 300 pods. Separately, the stock count has a 1-second TTL and 900K reads/s, so it misses ~300 times a second per pod. Fix: push the product key to the near-cache allowlist (config, no deploy); serve stock from a dedicated counter key with a near-cache TTL of 1s and an explicit "stock may lag 1s" label on browse pages; checkout continues to read the source.

Principal Approach: Treats it as a missing process: launches with known hot items must be pre-registered so the platform pre-marks keys hot, pre-sizes, and runs a load test with the actual value sizes. Also changes the client default: a timeout from a single overloaded node routes to the spare pool or serves stale, never directly to the DB.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Top DB queries by caller: 85% are SELECT product WHERE id = release_item and stock lookups. Hot-key sampler: one key at 380K reads/s. Node NIC at 92%.
TriageTimeouts on the hot node are treated as misses → fills from 300 pods. Stock TTL 1s → constant misses.
Quick fixNear-cache allowlist for the product key (2s TTL); stock served from near cache with 1s TTL; fill budget lowered for the hot namespace. DB CPU 100% → 35% in 2 minutes.
GuardrailsAutomatic near-caching above 0.5% of pod reads; timeouts from an overloaded node never count as misses for DB fills.
Post-mortemWhy did "hit rate 97%" hide the problem? Why did the client convert cache timeouts into DB load? Why was the launch not pre-registered?

Metrics to Watch: cache.hot_key_reads top-10, cache.node_network_bytes, cache.timeout_as_miss_total, db.queries_by_caller, db.read_qps vs budget.

Organizational Follow-up: launch calendar shared with the caching platform; release items pre-marked hot 24 hours ahead; a launch load test with real value sizes is part of readiness.

Ownership Question: "Who decides that stock may be 1 second stale on browse pages?" Staff answer: The commerce product owner, in writing, with the explicit carve-out that checkout always reads the source. Engineering proposes; product signs.

Key Takeaway: "A high hit rate can coexist with a database meltdown when one key's misses — or timeouts disguised as misses — are concentrated. Look at where DB load comes from, not just at the hit rate."

What clears the Staff bar:

  • Looks past the average hit rate to per-key load and node network
  • Spots timeouts-as-misses as the amplifier
  • Fixes with config (near-cache allowlist), not more nodes

Deep Dive 2: Silent Failure — Three Days of Wrong Permissions#

Context: A customer's security team reports that an employee who was removed from a workspace three days ago could still read its documents for about 70 minutes after removal — the audit log shows reads. Nobody noticed internally. Permission checks are cached with a 1-hour TTL and app-driven invalidation.

Questions to Surface First:

  • What wrote the membership change — the main app, an admin API, an SCIM sync job?
  • Does that write path delete the permission cache key? Which key — user-scoped, workspace-scoped, or both?
  • Is there a delete/fill race window with replica lag?
  • How many other removals took the same path, and over what period?

Typical L5 Approach: Shortens the permission TTL to 5 minutes and adds a delete call to the SCIM job. Fixes this path; the next new write path repeats the bug, and DB load for permission checks rises 12×.

Staff Approach: Finds the SCIM sync job writes membership directly via a bulk SQL update, bypassing the service that invalidates. Every SCIM-driven removal for 8 months stayed authorized for up to the TTL plus one race window. Fix: CDC-driven invalidation from the membership table, covering all write paths; permission entries carry the membership row version; revocation-sensitive checks (document reads after a membership change) verify the version with a cheap lookup. Notify affected customers per the security incident process.

Principal Approach: Reclassifies permission data: authorization decisions are a correctness-critical cache with a revocation SLO (e.g., ≤ 5 seconds), not a performance cache. Establishes a rule that any cache of authorization state must invalidate from the source's change stream and report revocation latency as a security metric reviewed by the security team.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Flush permission namespace for affected workspaces; confirm the user is now denied. Open a security incident.
TriageAudit: 1,140 SCIM removals in 8 months; reads after removal occurred for 63 of them, max 71 minutes.
Quick fixSCIM job calls the invalidation API after each batch; TTL temporarily 5 min while CDC is built.
GuardrailsCDC invalidation for membership and role tables; revocation latency probe: synthetic user removed every 5 min, measure time to deny; alert above 10s.
Post-mortemWhy was authorization cached with performance-cache rules? Why did a new write path not trigger an invalidation review?

Metrics to Watch: authz.revocation_latency_p99 (synthetic), inval.lag_ms{table=membership}, cache.stale_sample_rate{ns=authz}, authz.cache_hit_rate.

Organizational Follow-up: security team co-owns the authz cache's staleness budget; new write paths to authz tables require a review; customer notification through the security incident process.

Ownership Question: "Who owns revocation latency?" Staff answer: The identity team owns it as an SLO, with security as the approver of the budget. The caching platform provides the CDC invalidation mechanism.

Key Takeaway: "Caches of authorization state are security controls. Their staleness budget is a revocation SLO, and their invalidation must cover every write path."

What clears the Staff bar:

  • Finds the bypassing write path rather than just shortening TTL
  • Moves invalidation to the change stream and adds version checks for sensitive reads
  • Creates a synthetic revocation probe — measured correctness, not assumed

Deep Dive 3: Large-Customer Onboarding — The ML Team Wants 4 TB#

Context: The ML platform team wants to serve real-time features (user embeddings and aggregates, ~2 KB each, 2B users) from the shared cache cluster that currently holds 600 GB for twelve services. They need p99 under 5ms at 300K reads/s and plan to refresh all entries every 6 hours via a batch job.

Questions to Surface First:

  • What's the access distribution — do all 2B users' features get read, or a recent-active 200M?
  • What happens on a miss — is there a source to fill from, or is the batch job the only writer?
  • What's the cost of serving a stale or default feature vector — degraded ranking, or broken product?
  • How does the 6-hourly refresh write — 4 TB in a burst?

Typical L5 Approach: Adds 70 nodes to the shared cluster to fit 4 TB. The 6-hourly bulk write churns memory and evicts other tenants; the cluster's upgrade and failover procedures now involve a tenant whose data can't be refilled on miss.

Staff Approach: Recognizes this isn't a cache — there's no source to fill from on a miss; it's a feature store serving tier. Different intent, different system: a dedicated pool (or an SSD-backed key-value store, since 4 TB of RAM is ~$20–60K/month while NVMe-backed storage meets 5ms p99 at a fraction of the cost), loaded by bulk swap — write a new version of the dataset, flip a pointer — not by overwriting live keys. Recently active users (~200M, ~400 GB) can sit in RAM; the long tail on SSD.

Principal Approach: Uses the request to define the platform's tenancy rules: the shared cache pool is for caches of a source of truth; datasets without a fill path go to a serving store with bulk-load semantics. Puts a price on each tier per GB so teams can make the tradeoff themselves.

Staff Approach — Full Reasoning
PhaseWhat to Do
Week 1Access analysis: 91% of reads hit 180M users active in the last 7 days. Miss behavior: default vector → ~2% ranking quality loss, acceptable for inactive users.
DesignHot tier: dedicated in-memory pool for active users (~400 GB). Cold tier: SSD-backed KV for the rest. Dataset versions loaded side-by-side; a pointer flips after validation.
GuardrailsLoad jobs throttled to a write budget; dataset version validation (row counts, null rates) before flip; rollback = flip back.
RolloutShadow reads for a week comparing features from the new tier vs current source; then cut over by traffic percentage.
Review30 days: p99 latency, cost per million reads, ranking metrics.

Metrics to Watch: features.read_latency_p99, features.default_vector_rate, features.dataset_version_age, pool memory, cost per million reads.

Organizational Follow-up: ML platform owns the feature serving tier's SLO; caching platform provides the in-memory pool as a service; finance gets a per-GB cost model.

Ownership Question: "Who's paged when features are stale?" Staff answer: The ML platform on-call. Staleness here comes from the batch pipeline, not the cache, and only they can fix it.

Key Takeaway: "If there's no source to fill from on a miss, it isn't a cache — it's a database with a cache's durability. Put it in a store designed for bulk-loaded serving."

What clears the Staff bar:

  • Identifies that the request is a serving store, not a cache, from the missing fill path
  • Uses access distribution to split RAM and SSD tiers with cost
  • Protects existing tenants from bulk-load churn

Deep Dive 4: Post-Mortem — The Rolling Upgrade That Took Down Checkout#

Context: A routine minor-version upgrade of the session and cart cache pool (18 nodes, primary/replica pairs) caused a 25-minute checkout outage. The runbook said: upgrade replicas, fail over, upgrade old primaries. You are running the post-mortem.

Questions to Surface First:

  • Were replicas fully synchronized before failover? How was that checked?
  • How did clients discover the new primaries — and how long did they keep writing to the old ones?
  • What happened to writes acknowledged by old primaries but not replicated?
  • Why did a cart cache problem become a checkout outage — is the cart stored only in the cache?

Typical L5 Approach: Adds a check for replication offset before failover and slows the upgrade. Necessary; it misses that carts lived only in the cache.

Staff Approach: Timeline shows three compounding issues. Failovers happened while replication lag was ~4 seconds under peak load, so acknowledged cart writes were lost (asynchronous replication). Client topology refresh took up to 30 seconds, so some pods wrote to demoted nodes. And the cart service used write-behind: carts were written to cache and flushed to the DB every 60 seconds — so lost cache writes were lost carts, and checkout failed validation against stale carts. Fixes: carts written to the database first (cache-aside), the cache is a cache again; upgrade tooling waits for replica offset parity and runs off-peak; clients refresh topology on MOVED/connection errors immediately.

Principal Approach: Names the class: a cache silently promoted to a source of truth. Introduces a namespace registry field — source_of_truth: external | cache — where cache requires a durability review and is disallowed for money-adjacent data. Upgrade tooling refuses to proceed on pools whose namespaces declare themselves as source of truth until a review signs off.

Staff Approach — Full Reasoning
PhaseWhat to Do
Timeline reconstructionFailover at 14:02 with replica offset lag of ~1.1 MB (~4s). Client errors spike 14:02–14:03. Cart validation failures 14:03–14:27.
Root causesWrite-behind made cache the source of truth for carts; async replication loses acknowledged writes on failover; slow client topology refresh.
Quick fixesPause upgrades; cart writes switched to DB-first behind a flag (latency +6ms p99).
GuardrailsUpgrade tooling: wait for offset parity, max one failover per 5 min, abort if hit rate drops > 2 points; off-peak window.
PreventionNamespace registry declares source of truth; durability review for any cache-as-truth namespace.

Metrics to Watch: cache.replication_offset_lag_bytes, cache.failover_events, client.topology_refresh_latency, cart.validation_failures, cache.writes_lost_estimate.

Organizational Follow-up: cart team owns the DB-first migration; caching platform owns upgrade tooling; change management requires a hit-rate and durability gate for cache maintenance.

Ownership Question: "Who approved putting carts only in the cache?" Staff answer: Nobody explicitly — that's the finding. Write-behind was a performance change reviewed as a performance change. Durability decisions need a durability review, owned by the data's team with platform consultation.

Key Takeaway: "Write-behind turns a cache into a database with asynchronous replication and no backups. If you choose that, choose it in writing."

What clears the Staff bar:

  • Finds the write-behind design as the root cause beneath the upgrade trigger
  • Explains asynchronous replication loss on failover precisely
  • Adds a durability declaration so the class of failure is visible before the next upgrade

Deep Dive 5: Multi-Region Expansion — Stale Fills From the Lagging Replica#

Context: After launching an EU region (reads local, writes forwarded to the US primary, DB replication ~300ms–3s), EU users report that edits "revert" — they save a change, see it, then see the old value for minutes. US users don't see it. Invalidations are broadcast to both regions' caches as soon as the US primary commits.

Questions to Surface First:

  • When the EU cache entry is deleted, what does the next EU reader read — the EU replica? Has it applied the write yet?
  • How does the author see their change initially — a read-your-writes bypass, or the response of the write itself?
  • What is EU replica lag at p99 during peak?
  • Are invalidations and replication on independent paths with independent delays?

Typical L5 Approach: Adds a short TTL in the EU region. Staleness drops to the TTL, but EU DB load triples and the bug persists within the TTL window.

Staff Approach: Diagnoses the cross-region stale fill: the invalidation (fast, ~100ms) beats replication (~1–3s); the next EU reader misses, reads the EU replica that doesn't yet have the write, and fills the old value, which then lives for the TTL. Fix: drive EU invalidations from the EU replica's own change stream, so the delete happens only after the local replica has the write. Interim: a second delete in the EU 5 seconds after the first (above p99 replica lag). For authors, a per-user, per-entity bypass marker for 10 seconds after their write.

Principal Approach: Turns it into a standard for multi-region caching: invalidation is always derived from the local replica's change stream, never from a remote commit. Includes it in the multi-region readiness checklist, and plans for per-entity home regions as writes become regional.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediateConfirm with tracing: EU fill timestamps fall between invalidation receipt and replica apply time.
Triage0.7% of EU writes followed by a stale fill; median stale duration = remaining TTL (~20 min).
Quick fixDelayed second delete in EU (5s); author bypass marker (10s). Stale fills drop 95%.
Long-termEU invalidation worker consumes from EU replica CDC; delayed delete removed.
GuardrailsSampled staleness per region; alert on EU-specific staleness > budget.

Metrics to Watch: cache.stale_sample_rate{region}, db.replica_lag_ms{region} p99, inval.lag_ms{region}, cache.fill_before_replica_apply_total.

Organizational Follow-up: the multi-region program's checklist gains a caching item; the caching platform owns per-region invalidation workers.

Ownership Question: "Who owns cross-region staleness?" Staff answer: The caching platform owns the mechanism; each data owner owns their namespace's regional staleness budget.

Key Takeaway: "In multi-region, an invalidation that outruns replication is worse than no invalidation — it invites a stale fill. Invalidate from the local replica's change stream."

What clears the Staff bar:

  • Explains the race between invalidation and replication across regions
  • Uses local CDC for ordering instead of shorter TTLs
  • Handles read-your-writes for the author explicitly

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Identify whether a cache is a latency, load or compute shield, and derive the hit-rate floor from DB capacity
  • Convert hit-rate changes into DB load and headroom in your head
  • Write per-class staleness budgets and pick an invalidation strategy for each
  • Draw the delete/fill race and close it with leases, versioned sets or ordered invalidation
  • Design invalidation from the change log for multi-writer data, with lag SLOs and DLQs
  • Protect the miss path: coalescing, fill budgets, stale-on-error, spare capacity, warm handoffs
  • Handle hot keys with near caches and replication, with the bandwidth arithmetic
  • Size a cluster from working set, throughput and network, and estimate its cost
  • Operate a shared cache platform with isolation by placement, and explain why prefixes are not quotas
  • Extend caching to a second region without cross-region stale fills

The Bar for This Question#

Mid-level (L4): Puts Redis in front of the database with cache-aside and a TTL; deletes on write. Knows LRU. Works in steady state; hasn't considered races, cold starts or hot keys.

Senior (L5): Adds consistent hashing, replicas, TTL jitter, maybe request coalescing; knows the thundering-herd problem. The gap: doesn't compute the DB's dependency on hit rate; staleness is "eventually consistent"; invalidation relies on every writer remembering; treats a cache outage as "slower". The design is sound until the first upgrade, launch or batch job.

Staff+ (L6): Starts from the cache's job and the hit-rate floor. States staleness in seconds per class with a product owner. Closes the delete/fill race with leases or versions and invalidates from the change log. Protects the database with fill budgets and a degraded path. Handles hot keys with near caches and bandwidth math. Separates load-bearing tenants. Measures staleness instead of assuming it. Names who pays: users for staleness, the DB team for misses, other tenants for shared-cluster churn. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "A Cache You Can't Restart Is a Database You Forgot to Back Up"#

SymptomWhat It Means
Upgrades only at 3 a.m. with a war roomThe DB can't survive a cold cache
"Don't flush that cluster" in the runbookThe cache is load-bearing and undeclared
Write-behind for performanceThe cache holds data nowhere else

The Staff position: If restarting the cache is dangerous, either engineer it as a capacity tier — replicas, warm handoffs, fill budgets, game days — or make the DB survive without it. Leaving it in between is the most common caching risk in production.

Why this matters in interviews: Asking "can we restart this cache at noon?" is a fast way to show you think about the cache as part of the system's capacity, not a bolt-on.

10.2 "Most Invalidation Bugs Are Ordering Bugs, Not Logic Bugs"#

BugReal Cause
Stale value after updateFill from old read landed after delete
Edits "revert" in a regionInvalidation arrived before replication
Write-through shows older valueTwo writers updated cache in a different order than the DB

The Staff position: The question isn't "did we delete?" but "did the delete happen after every read that could re-populate the old value?" Leases, versions and change-log ordering answer that; more deletes and shorter TTLs don't.

Why this matters in interviews: Candidates who frame invalidation as an ordering problem can reason through any variant the interviewer invents.

10.3 "Near Caches Fix More Hot-Key Incidents Than Resharding Ever Will"#

ApproachEffect on a 400K reads/s key
Add cache nodesNone — the key still has one home
Better hash functionNone
Key replication ×8~50K reads/s per replica — helps, costs 8× deletes
Near cache, 2s TTL, 300 pods~150 remote reads/s total

The Staff position: Popularity is concentrated by nature. Moving the hottest bytes into the process that needs them — for a bounded few seconds — beats any change to the cluster.

Why this matters in interviews: "Consistent hashing" is the reflexive answer to hot keys; correcting it with arithmetic is memorable.

10.4 "Shared Cache Clusters Are a Tragedy of the Commons"#

Team IncentiveShared Outcome
"RAM is free here — cache more, longer"Memory pressure, evictions for everyone
"Our batch job is just a one-time load"Hot sets of other teams evicted
"We need that upgrade postponed"Cluster stuck on old versions

The Staff position: Share operations, not memory. A platform team should run the pools and the client; load-bearing tenants get isolated instances; everyone gets per-namespace dashboards and a cost per GB.

Why this matters in interviews: Showing you've thought about the cache as a multi-tenant platform is a strong L6-to-L7 bridge.

10.5 "Hit Rate Is a Vanity Metric Until You Weight It by Miss Cost"#

NamespaceHit RateMiss CostDB Time per Second
User profile99%2ms point readLow
Search facets80%120ms aggregate queryVery high
Fleet average96%—Hides the problem

The Staff position: Report hit rate per namespace and weight misses by their cost: misses/s × cost per miss = DB seconds per second. That's the number the database feels, and it points to which namespace to optimize.

Why this matters in interviews: Moving from "our hit rate is 96%" to "our misses cost the DB 1.8 CPU-seconds per second, 70% from facets" shows you measure what matters.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

The Staff engineer builds a correct, resilient cache for one service. The Principal engineer counts the organization's caches and finds 37 Redis and Memcached clusters run by 22 teams, with four client libraries, three different ideas of what a TTL protects, and no inventory of which databases have been right-sized to which hit rates. Each cluster is reasonable; together they are a hidden capacity dependency of unknown size. The L7 problem is turning caching from a per-team performance trick into a governed capacity tier: one client that makes the safe protocol the default, declared staleness contracts, declared load-bearing status, and a platform that can upgrade, resize and fail over clusters without a war room.

The Org-Level Fault Line#

A caching platform vs caches owned by each service team.

OptionWhat WorksWhat BreaksWho Pays
Every team runs its ownAutonomy; tuned to each workload; no platform dependency22 teams relearning failover, hot keys and the fill race; idle RAM at 30% utilization; nobody knows which DBs depend on which cachesFinance (idle capacity), on-call (repeat incidents), DB teams (hidden dependencies)
One central shared clusterCheapest per GB; one operatorNoisy neighbors; upgrades need everyone's consent; one bad tenant affects allLoad-bearing tenants; platform on-call
Platform operates pools; teams own namespacesShared client, tooling and on-call; isolation where it matters; teams own staleness and keysRequires a namespace registry, placement policy and chargebackPlatform team (5–8 engineers at scale); tenants accept standards

🧭 Principal Move: "The platform owns the servers, the client library and the upgrade tooling. Teams own their namespaces — keys, staleness budgets, invalidation mappings — declared in a registry. Load-bearing namespaces get isolated pools and a hit-rate SLO the DB team signs. Prefix-sharing in one big cluster is what we're migrating away from."

Cost Model#

Assumptions: managed in-memory cache at ~$8–12 per GB-month including replicas' share; ~$250K fully loaded per engineer-year; DB read capacity priced as replica instances that would be needed to serve the traffic without the cache.

ScaleRead TrafficCache Infra ($/month)HeadcountOn-call LoadDB Cost Avoided
Startup~10K reads/s~$300–1K (one small managed instance + replica)~0.1 eng (part of the service team)Rare; mostly upgradesSmall; often the cache is a latency shield only
Growth~300K reads/s, 10–20 namespaces~$10–25K (1–2 TB across pools, cross-AZ replicas)2–3 eng caching platform (client, tooling) + service teams own namespaces1–3 cache pages/month; upgrades monthly~$50–150K/month of replicas that would otherwise serve the reads
Enterprise~5M+ reads/s, 100+ namespaces, multi-region~$150–400K (10–30 TB, multiple regions, spare pools)6–10 eng platform (client, CDC invalidation, pools, registry)Dedicated rotation; game days quarterlyMillions per month; the cache is a primary capacity tier

The pricing insight: the cache usually pays for itself many times over in DB capacity — which is exactly why its failure is expensive. The platform's budget should be justified against incident cost and DB headroom, not against the RAM bill. A 3-engineer investment in warm handoffs, fill budgets and per-namespace isolation is cheaper than one hour-long outage of a revenue path.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoor TypeReversibility Cost
Letting the DB be sized for the miss rateOne-way-ishUndoing it means buying back 5–20× DB read capacity; in practice you're committed to a load-bearing cache
Write-behind / cache as source of truth for a namespaceOne-wayData that lived only in the cache may be lost; migration to DB-first needs a durability project
Staleness contracts published to product teams and customersOne-way-ishTightening later costs DB capacity and invalidation investment
Key schema without a version componentOne-way (cheaply avoided)Every future format change is a cold start
Cache engine (Redis vs Memcached vs managed variant)Two-wayClient abstraction makes it a migration of servers, not services
TTLs, near-cache thresholds, fill budgetsTwo-wayConfig
Proxy vs client-side routingTwo-wayClient library swap

The Standard I'd Write#

RFC-CACHE-001: Caching Standard
Status: Approved   Owners: Caching Platform + Database Reliability + Security

Scope
  Every remote cache (and every near cache holding user-visible mutable data)
  used by production services.

MUST
  1. Register each namespace with: owner, data source, staleness budget,
     invalidation mechanism, load-bearing (yes/no), source_of_truth
     (external | cache).
  2. Use the platform cache client (leases on fill, miss coalescing, timeouts,
     breaker, per-pod fill budgets).
  3. Include a schema version in every key.
  4. Load-bearing namespaces: run on isolated pools with cross-AZ replicas, a
     declared hit-rate floor, and a DB fill budget agreed with the DB owner.
  5. Data with more than one writer, or any out-of-band write path, invalidates
     from the source's change stream.
  6. Caches of authorization state declare a revocation SLO approved by Security.
  7. source_of_truth = cache requires a durability review and is prohibited for
     payments, carts and authorization.

SHOULD
  1. Apply TTL jitter of ±10–20%.
  2. Enable sampled staleness checks (0.1% of hits) for namespaces with budgets
     under 60 seconds.
  3. Keep values under 100 KB; never exceed 512 KB.

Exceptions
  Filed with Caching Platform; reviewed within 5 business days; time-boxed to
  2 quarters. Security must approve exceptions to MUST 6.

Success metrics
  - Outages attributed to cold caches or cache maintenance: 0 per half-year
  - Namespaces with measured staleness within budget: ≥ 99%
  - Unregistered caches in production: 0 by end of year 2
  - Cache maintenance during business hours without incident: routine
  - Fleet RAM utilization: ≥ 60% (from ~30%)

What I'd Tell the VP#

"Our caches carry most of our read traffic — by my estimate, several of our databases would fall over within minutes if certain caches emptied, and nobody has a list of which ones. We've had two outages this year from routine cache maintenance and one security finding from a stale permission cache. I'm proposing a caching platform: one client library that makes the safe patterns automatic, a registry of what each cache holds and how stale it may be, and isolated pools for the caches our databases depend on. It's about five engineers for a year. It removes a recurring outage class, lets us do maintenance in daylight, and should cut cache spend by consolidating 37 clusters running at 30% utilization."

Principal Interview Signals#

SignalWhat It Sounds Like
Sees the hidden dependency"Which databases are sized for a hit rate, and who knows the number? That list doesn't exist yet."
Prices the posture"Making the DB survive full cache loss is ~$80K a month; making cache loss rare is ~$15K plus game days."
Identifies one-way doors"Write-behind and right-sizing the DB to the miss rate are the doors that don't swing back."
Redraws ownership"Platform owns servers and the client; data owners own invalidation; DB owners sign the fill budget."
Knows when not to standardize"The ML feature store isn't a cache — it gets a serving store, not our pool."

Staff answers that L7 interviewers find insufficient:

  • "We'll engineer this cache with replicas, leases and near caching" — right for one service; ignores the 36 other clusters with none of it.
  • "Hit rate is our SLO" — correct; missing that the DB owner should sign the floor and the fill budget.
  • "We'll use CDC for invalidation" — right mechanism; missing who owns the pipeline's lag SLO and the per-table key mapping.

Appendices

Appendix A: Invalidation Mechanics in Depth#

A.1 Leases (Fill Tokens)#

def get(ns, key, loader):
    v, token = cache.get_or_lease(key)          # server returns value, or a lease token
    if v is not None:
        return v
    if token is None:                            # someone else holds the lease
        sleep(10ms); return get_retry(ns, key, loader, max_wait=200ms)  # or serve stale
    value, version = loader(key)
    cache.set_with_lease(key, value, version, token, ttl=jitter(ns.ttl))
    # rejected if a delete happened after the lease was issued
    return value

def invalidate(key):
    cache.delete(key)                            # also invalidates any outstanding lease

Two effects with one mechanism: only one caller per key fills at a time (the rest wait briefly or get a stale copy), and a fill based on data read before the delete can't land after it. Redis doesn't ship leases natively; implement with a small script: on miss, SET lease:{key} token NX PX 2000; on set, compare-and-set only if lease:{key} still equals the token; on delete, delete both.

A.2 Versioned Sets#

Store the source row's version (or commit LSN) with the value. A set succeeds only if its version is ≥ the version already cached — and deletes leave a tombstone with the deleting version for a few seconds so an older fill can't recreate the key.

SET_IF_NEWER(key, value, version):
    current = HGET key version  (or tombstone version)
    if current is None or version > current:
        HSET key value version; EXPIRE key ttl

A.3 Delayed Double Delete#

Delete on write, then delete again after a delay longer than p99 replica lag plus fill time (e.g., 1–2s in one region, 5s cross-region). Cheap, needs no server-side support, and narrows — but doesn't close — the race window. A reasonable interim fix; not the end state.

A.4 Read-Your-Writes for the Author#

After a user writes entity E, set a short-lived marker ryw:{user}:{E} (TTL ~2× replica lag, e.g., 10s). Reads by that user for E bypass the cache and read the primary. Everyone else sees the cache, within the staleness budget.

A.5 Miss Coalescing and Fill Budgets#

Diagram: A.5 Miss Coalescing and Fill Budgets

Appendix B: Keys and Data Model#

ElementConventionWhy
Namespaceproduct, profile, authzMaps to owner, budget and pool in the registry
Schema versionv3Format changes become new keys with dual-read migration
Entity IDSource primary keyLets CDC map row changes to keys deterministically
Variant:locale=fr, :size=thumbExplicit fan-out; invalidation must delete every variant (keep a variant index or bounded set)
Derived/aggregate keyscategory:{id}:top20Invalidation mapping must list them, or they rely on TTL — declare which

Values. Store the source version with the value. Compress values over ~1 KB (2–4× is typical for JSON). Avoid giant composite values for hot keys — split the frequently-read summary from the rarely-read detail.

Negative caching. Cache "not found" with a short TTL (e.g., 30–60s) to stop repeated lookups for missing IDs — including enumeration attacks — from reaching the DB. Invalidate negatives on create.


Appendix C: Coordination Mechanisms#

C.1 Warm Handoff for Planned Node Replacement#

Diagram: C.1 Warm Handoff for Planned Node Replacement

C.2 Cross-Region Invalidation Ordering#

Diagram: C.2 Cross-Region Invalidation Ordering

Driving each region's invalidations from its own replica guarantees the delete happens after the local data is new, so a refill can't pick up the old row.

C.3 Quick Comparison#

MechanismCloses Fill Race?Covers All Writers?Added LatencyOperational Cost
TTL onlyNoYes (eventually)NoneNone
App delete on writeNoNo~0.5ms on writesLow
Delayed double deleteNarrows itNoNone (async)Low
LeasesYesNo (needs a delete source)Retry wait on contentionClient support
Versioned sets + tombstonesYesNoSmallClient + script
CDC invalidationWith local-replica orderingYesPipeline lag (~100ms–1s)A pipeline with an owner
CDC + leases/versionsYesYesPipeline lagHighest — the load-shield default

Appendix D: API Contract & Client Behavior#

Client BehaviorDefaultWhy
Remote GET timeout20–50msAbove p99.9 of healthy reads; well below user-visible latency budgets
BreakerOpen at > 1% timeouts over 5s per shard; half-open probes every 2sStop paying timeouts; avoid hammering a recovering node
Timeout handlingNever treated as a plain miss when the shard is overloaded; route to spare pool or stalePrevents converting cache distress into DB load
Fill budgetPer-pod token bucket sized so pods × rate ≤ DB budgetHard ceiling on DB damage from a cold cache
Stale-on-errorPer namespace, max staleness (e.g., 5 min)Keeps pages up when the remote tier is gone
Near cacheAuto for keys > 0.5% of pod reads; TTL 1–5s; opt-out per namespaceHot-key protection without a human
Batch readsget_many with pipelining; fan-out capped per requestAvoid one request touching all shards
TTL jitter±10–20%Prevents synchronized expiry after bulk loads

Appendix E: Observability#

E.1 Core Metrics#

MetricMeaningAlert
cache.hit_rate{namespace}Per-namespace hit rateLoad-bearing namespace below floor for 5 min: page
db.read_qps vs db.read_budgetWhat the DB actually feels> 80% of budget: page
cache.miss_cost_seconds{namespace}Misses × average fill latencyTrend; points to which namespace to fix
cache.stale_sample_rate{namespace}Sampled hits older than budget> 0.1% for budgets ≤ 60s: page owner
inval.lag_ms{table, partition}Commit → delete appliedp99 > 5s: page writer team
cache.evictions{namespace}Memory pressure by tenantSpike on load-bearing namespace: ticket
cache.hot_key_readsTop keys by read rateAuto near-cache; alert if NIC > 70%
cache.node_network_bytesPer-node bandwidth> 70% NIC: page
cache.timeout_rate{shard}Shard health from the client's view> 1%: breaker + ticket
cache.replication_offset_lagReplica freshnessBlock failovers above threshold

E.2 Control Plane vs Data Plane#

The data plane — client library, cache servers — must serve reads with the control plane down: clients cache topology and namespace config locally and keep working with the last-good version. The control plane — registry, upgrade tooling, CDC workers, warm-handoff orchestration — can be unavailable for minutes. Its failure stops changes and delays invalidations (raising staleness), but must never stop reads. Invalidation lag during a control-plane outage is the measurable cost; sampled staleness shows it.

E.3 Debugging the Silent Failure#

"Users see stale data sometimes" is the hardest cache ticket. The checklist: (1) sampled staleness for that namespace — is it over budget, and since when? (2) invalidation lag per partition — a stalled partition explains 1/N of entities; (3) write paths — any new writer bypassing the change stream? (4) regional skew — does it happen only in a replica region? (5) TTLs — did someone raise a TTL relying on invalidation that doesn't exist? Each step is a dashboard query when the metrics above exist, and a day of guesswork when they don't.


Appendix F: Scale Evolution#

ScaleWhat WorksWhat Breaks Next
< 20K reads/sOne managed instance + replica; cache-aside, TTL, app deletesStale data from out-of-band writers
20K–300K reads/sClustered cache, shared client with coalescing and leases, per-namespace metricsCold starts become DB outages; hot keys
300K–3M reads/sIsolated pools, CDC invalidation, near caches, fill budgets, warm handoffsMulti-tenant governance, cost, regional consistency
3M+ reads/s, multi-regionRegional pools, local-replica invalidation, RAM+SSD tiers, registry and chargebackOrganizational consistency; one-way doors

What You Don't Build on Day One:

  • CDC invalidation — until there's a second writer or an out-of-band write path
  • A proxy tier — client-side routing is enough until connection counts or multi-language clients force it
  • Cross-region cache replication — regional caches with local invalidation are simpler
  • SSD-backed tiers — until the working set's RAM cost passes ~$20K/month
  • A custom cache server — the leverage is in the client and the operating model

Appendix G: Multi-Tenancy, Fairness & Cost#

Isolation by placement. Memcached and Redis evict across the whole process; namespaces sharing an instance share fate. Real isolation means separate instances (or separate processes) per tenant group. A pool can contain many small instances and pack tenants onto them by size and criticality.

Chargeback. Charge namespaces per GB-month of reserved memory, and show them their hit rate and DB cost avoided next to it. Teams caching low-reuse data see a high cost with little avoided DB load and remove it themselves — the most effective cost optimization is usually deletion.

Value economics. Compression cuts memory 2–4× on text formats for ~10–50µs of CPU per value; for large values it's almost always worth it. Trimming fields nobody reads from cached objects is the next cheapest win.

Tiering. For compute-shield and long-tail data, SSD-backed caches (flash-backed Memcached variants, or key-value stores on NVMe) serve sub-millisecond to low-millisecond reads at a fraction of RAM cost. Keep the hot set in RAM and the long tail on flash.

Fairness under scarcity. When a pool is short of memory, evict from namespaces that exceed their reservation first. When the DB is short of capacity, fill budgets should be per namespace, with load-bearing, user-facing namespaces getting priority over background ones.

  1. Loading the index…