Why This Matters#
Redis is not a cache. It is a single-threaded, in-memory data-structure server that happens to be very good at caching — and most Redis incidents come from teams that forgot the first half of that sentence. Every design that touches Redis inherits three facts: one command runs at a time per shard, the working set must fit in RAM, and replication is asynchronous. Those three facts decide whether Redis is your fastest component or your largest blast radius.
It keeps showing up in interviews because it is the default answer to half the hard questions: rate limiting, leaderboards, sessions, distributed locks, presence, deduplication, fan-out timelines, queues. Interviewers know the default, so they do not score you for saying "Redis." They score what comes next: which data structure, what the key looks like, what happens to the 0.5–1 second of acknowledged writes lost on failover, what the hot key does to one shard, and who gets paged when used_memory crosses maxmemory at 2am.
The L5 answer says "we'll put it in Redis." The L6 answer says "sorted set per region keyed lb:{region}:{day}, ZADD on score change, ZREVRANGE 0 99 for top-100, 16 shards, AOF everysec because we can rebuild from Kafka, and we accept ~1s of loss on failover because product signed off that a leaderboard glitch is not a refund." The L7 answer asks whether the org should run 40 unmanaged Redis clusters at all — and what the license change of 2024 means for the paved road.
The 60-Second Pitch#
"Redis gives us sub-millisecond reads and writes — roughly 0.1–0.3 ms server time, 0.5–1 ms round trip in the same AZ — on rich data structures, not just blobs. That matters because the operations we need — increment a counter, add to a sorted set, pop from a list, append to a stream — execute atomically on the server, so we get correctness without distributed locks. One shard does 100K–200K simple ops/sec, over 1M with pipelining. I'd treat it as derived state: the source of truth lives in Postgres or Kafka, Redis can be rebuilt, and we design the degraded mode for when it's gone. If the data can't be rebuilt, I either pay for a durable variant like MemoryDB or I don't put it in Redis."
The Staff-level insight: the decision is never "Redis or not" — it is which Redis contract. Cache (evictable, rebuildable), ephemeral state (sessions, rate-limit counters — loss is tolerable), coordination primitive (locks, leases — correctness depends on timing assumptions), or system of record (durability required). Each contract has a different persistence setting, eviction policy, failover posture, and owner. Mixing two contracts in one cluster is the most common root cause in Redis post-mortems.
| Contract | Example | Eviction Policy | Persistence | Loss Tolerance | Who Pays on Loss |
|---|---|---|---|---|---|
| Cache | Product page fragments | allkeys-lru / allkeys-lfu | None or RDB | Total — rebuild from source | Origin DB absorbs a miss storm |
| Ephemeral state | Sessions, rate-limit buckets | volatile-ttl | RDB or AOF everysec | Minutes of state | Users re-login; limits reset briefly |
| Coordination | Locks, leases, idempotency keys | noeviction | AOF everysec | Must fail safe | Double-execution risk owned by caller |
| System of record | Leaderboard with no upstream log | noeviction | AOF always or MemoryDB | Zero | Customer-visible data loss |
🎯 Staff Move: "Before I pick a data structure, I want to name the contract. If this Redis is a cache, eviction is a feature. If it's holding idempotency keys, eviction is a correctness bug — so it gets its own cluster with
noevictionand a different alert."
Architecture & Internals#
Only the internals that change design decisions.
The Single-Threaded Event Loop#
Redis executes commands on one main thread per process. Since 6.0, network reads/writes can be offloaded to I/O threads (io-threads 4), but command execution is still serial. This is the source of both its superpower and its worst failure mode:
- Superpower: every command is atomic.
INCR,ZADD,LPUSH,SET NX PXneed no locks. A Lua script orMULTI/EXECblock runs without interleaving. - Failure mode: one slow command stalls every client on that shard.
KEYS *on 50M keys,SMEMBERSon a 2M-member set,DELon a 1 GB hash, or a Lua script looping for 800 ms — each is a shard-wide outage for its duration.
The mental model: Redis throughput is ~100K–200K ops/sec per core, and latency is the sum of everything queued ahead of you. An O(N) command on a large N is not slow for one client; it is slow for all of them.
Encodings: Why Small Collections Are Cheap#
Redis stores small collections in compact contiguous encodings and converts them to pointer-heavy structures past a threshold. The thresholds are tunable and they change memory by 5–10×:
| Type | Compact Encoding | Converts When (defaults, 7.x) | Large Encoding | Memory Impact |
|---|---|---|---|---|
| Hash | listpack | > 128 fields or any value > 64 bytes | hashtable | ~5–10× more per field |
| Sorted set | listpack | > 128 members or member > 64 bytes | skiplist + hashtable | ~3–5× more; O(log N) ops |
| Set | intset (integers) / listpack | > 512 integers / > 128 members | hashtable | ~4–8× more |
| List | quicklist of listpacks | — (always quicklist) | quicklist | Nodes of ~8 KB |
| String | int / embstr (≤ 44 bytes) / raw | length | raw SDS | ~50–70 bytes overhead per key |
Why it matters: Instagram publicly described storing ~300M photo→user mappings by bucketing them into hashes of ~1,000 fields each (HSET media:{id/1000} {id} {user}), cutting memory from ~21 GB to ~5 GB. The trick works because each top-level key carries ~50–90 bytes of overhead (dict entry, robj, expire entry), while a field inside a listpack hash costs a few bytes plus the payload.
Persistence: RDB, AOF, and the fork()#
| Mode | Mechanism | Loss Window on Crash | Cost | When |
|---|---|---|---|---|
| None | Pure memory | Everything | Zero | Pure cache, rebuilt from origin |
| RDB | fork() + child writes a point-in-time snapshot | Since last snapshot (often 5–15 min) | Fork pause + copy-on-write memory | Cache with warm restart; backups |
AOF everysec | Append every write; fsync once per second | ~1 s (up to 2 s under disk stall) | ~5–15% throughput; AOF rewrite forks | The production default for state |
AOF always | fsync per write batch | ~0 on the node | 2–10× lower write throughput | Rarely worth it — replication loss dominates |
| RDB + AOF (preamble) | AOF rewrite starts with RDB snapshot | ~1 s | Faster restarts than pure AOF | Default for state since 4.0 |
The fork() is the hidden cost. The kernel must copy page tables before the child runs — roughly 10–20 ms per GB of resident memory on typical cloud VMs, worse with some hypervisors and without huge-page tuning. A 50 GB instance can pause for ~0.5–1 s every snapshot. Copy-on-write then duplicates every page the parent modifies while the child writes: a write-heavy 50 GB instance can need 20–40 GB of extra RAM during BGSAVE. Rule: size instances so maxmemory is ≤ 50–60% of host RAM when persistence or replication is on, and prefer many 10–25 GB shards over few 100 GB ones.
🎯 Staff Insight: "Persistence" on a replicated Redis mostly protects against simultaneous loss of primary and replicas. A single-node crash is handled by failover — and failover loses whatever the replica hadn't received, regardless of
appendfsync alwayson the primary. Local fsync does not fix asynchronous replication.
Replication and Failover#
Replication is asynchronous. The primary acknowledges the client, then streams the command to replicas. Replicas track an offset; a short disconnect resumes via PSYNC from the replication backlog (repl-backlog-size, default 1 MB — far too small; set 256 MB–1 GB for busy primaries). If the replica falls behind the backlog, it needs a full resync: the primary forks, writes an RDB, ships it, then streams buffered writes. Full resync of a 25 GB dataset over 10 Gbps takes minutes and costs the primary a fork plus a replica output buffer that can reach GBs.
Two knobs bound loss:
WAIT numreplicas timeout— block until N replicas acknowledge the offset. Reduces loss, does not make Redis linearizable (a failover can still pick a replica that did not ack).min-replicas-to-write 1+min-replicas-max-lag 10— the primary refuses writes if no replica is within 10 s. Bounds split-brain loss to ~10 s instead of unbounded.
Sentinel (for non-cluster deployments) runs 3+ processes that vote a primary down after down-after-milliseconds (commonly 5,000–30,000) and promote a replica. Redis Cluster does the same internally via its gossip bus, with cluster-node-timeout (default 15,000 ms). Either way, expect 5–30 s of write unavailability per failover plus the replication-lag loss window.
Redis Cluster: 16,384 Slots#
Cluster mode partitions keys by CRC16(key) mod 16384. Each primary owns a range of slots; clients cache the slot map and follow MOVED (permanent) and ASK (mid-migration) redirects.
| Fact | Consequence for Design |
|---|---|
Multi-key commands (MGET, SUNIONSTORE, Lua, MULTI) require all keys in one slot | Use hash tags: cart:{user42}:items and cart:{user42}:meta hash only user42 |
| Hash tags co-locate by design | A hash tag with low cardinality ({global}) recreates a single-node bottleneck |
Resharding moves slots key-by-key with MIGRATE | Big keys (> 100 MB) block migration and the source shard |
| Practical ceiling ~ 500–1,000 nodes | Gossip overhead grows with N²; beyond that, run multiple clusters |
| Cross-slot atomicity does not exist | Design operations to be single-key or single-tag |
Lua Scripts and Functions: Atomicity Has a Budget#
EVAL/EVALSHA (and Redis 7 FUNCTIONs) run a script atomically on the shard — nothing interleaves. That makes Lua the right tool for read-modify-write sequences (check-then-decrement inventory, token-bucket refill, compare-and-delete lock release). It is also a stop-the-world event for that shard.
| Rule | Why | Number |
|---|---|---|
| Keep scripts O(1) or O(small N) | The shard serves nothing else while the script runs | Target < 1 ms; alert > 5 ms |
Declare every key in KEYS[] | Cluster routing and slot checks rely on it | All keys must share one slot |
| No unbounded loops over collections | A loop over 1M members blocks for ~100s of ms | Cap iterations or use SCAN from the client |
| Script timeout is not a kill switch | After busy-reply-threshold (default 5 s) clients get BUSY; a script that already wrote can only be stopped with SHUTDOWN NOSAVE | Treat 5 s as an outage, not a limit |
| Load once, call by SHA / function name | Avoids resending script bodies | Saves bandwidth on 100K+ calls/s |
MULTI/EXEC gives atomic batching without logic; WATCH adds optimistic concurrency (the EXEC aborts if a watched key changed). Under contention above ~5–10% conflict rate, WATCH retry loops waste more than a Lua script costs. See Dealing with Contention.
Pub/Sub vs Streams#
Both move messages; their contracts are opposite.
| Property | Pub/Sub (PUBLISH / SPUBLISH) | Streams (XADD / XREADGROUP) |
|---|---|---|
| Storage | None — delivered to connected subscribers, then gone | Persisted in the keyspace (RAM), replicated, AOF-able |
| Offline subscriber | Misses everything | Resumes from last-delivered ID |
| Acknowledgment | None | Per-consumer Pending Entries List + XACK |
| Slow consumer | Disconnected at output-buffer limit (default hard 32 MB) | Falls behind; memory grows until MAXLEN trims |
| Fan-out cost in cluster | Classic Pub/Sub broadcasts across the whole cluster bus; sharded Pub/Sub (7.0+) stays on one shard | Stream lives on one shard (one key) |
| Throughput | 100K+ msgs/s per shard, fan-out multiplies cost | ~50K–150K XADD/s per shard |
| Right for | Cache-invalidation hints, presence pings, "something changed" nudges | Job queues, activity feeds, small event pipelines |
Memory: Where the Bytes Actually Go#
used_memory is not your payload. For 100M small string keys (20-byte key, 20-byte value), expect ~8–10 GB, not 4 GB: dict entries, robj headers, SDS headers, expire-dict entries for TTL keys (~30–40 more bytes each), and allocator rounding. Add fragmentation (mem_fragmentation_ratio = RSS / used; 1.0–1.3 healthy, > 1.5 means 33%+ waste — activedefrag reclaims it slowly), client buffers (output buffers for slow readers and replicas can reach GBs), and the replication backlog. Budget maxmemory for data and headroom separately, and measure with MEMORY USAGE on samples before launch rather than estimating from payload size.
Data Modeling / Core Usage — "The Entire Game"#
In Redis, the data structure is the algorithm. Picking the right type turns a distributed-systems problem into one atomic command. Picking the wrong one turns it into an O(N) scan on the hot path.
Pick the Structure from the Operation#
| You Need To... | Structure | Core Commands | Complexity | Watch Out For |
|---|---|---|---|---|
| Cache a blob / counter | String | SET k v EX 300, INCR | O(1) | Values > 100 KB hurt network and latency |
| Store an object with partial updates | Hash | HSET, HINCRBY, HGETALL | O(1) per field | HGETALL on 10K-field hashes |
| Rank / top-K / time-window | Sorted set | ZADD, ZREVRANGE, ZRANGEBYSCORE, ZREMRANGEBYSCORE | O(log N + M) | Unbounded growth — trim on write |
| Simple FIFO work queue | List | LPUSH, BRPOP, LMOVE | O(1) | No ack — crash between pop and process loses work |
| Durable log with consumer groups | Stream | XADD, XREADGROUP, XACK, XAUTOCLAIM | O(1) append | Pending-entries list grows if consumers don't ack |
| Membership / dedup | Set | SADD, SISMEMBER | O(1) | SMEMBERS on big sets blocks |
| Cardinality estimate | HyperLogLog | PFADD, PFCOUNT | O(1), 12 KB max | 0.81% standard error — not for billing |
| Per-user boolean flags | Bitmap | SETBIT, BITCOUNT | O(1) / O(N) | Sparse IDs waste memory — offset = ID |
| Nearby search | Geo (sorted set) | GEOADD, GEOSEARCH | O(N + log M) | Radius search cost grows with density |
Key Design Rules#
- Namespace with colons, put the tenant or entity first:
rl:{tenant}:{route}:{window}. Keys are grep-able in--bigkeysoutput and debuggable at 3am. - Hash-tag only what must be atomic together.
{user42}for a user's cart and its metadata; never{hot}for global counters. - Every key gets a TTL unless it is explicitly a system-of-record key. Untagged keys are how a cache cluster quietly becomes a database nobody owns.
- Bound every collection on write.
ZADDthenZREMRANGEBYRANK k 0 -1001keeps a top-1000;XADD ... MAXLEN ~ 100000caps a stream. - Keep values under ~10 KB; hard ceiling ~1 MB. A 5 MB value takes ~4 ms of network time on 10 Gbps and blocks the event loop while serializing.
Worked Pattern: Sliding-Window Rate Limiter#
-- KEYS[1] = rl:{tenant42}:POST:/orders ARGV = now_ms, window_ms, limit, request_id
-- Runs atomically: no other command interleaves on this shard.
local key, now, window, limit = KEYS[1], tonumber(ARGV[1]), tonumber(ARGV[2]), tonumber(ARGV[3])
redis.call('ZREMRANGEBYSCORE', key, 0, now - window) -- drop expired entries
local count = redis.call('ZCARD', key)
if count >= limit then
return {0, count} -- reject
end
redis.call('ZADD', key, now, ARGV[4]) -- record this request
redis.call('PEXPIRE', key, window) -- bound memory
return {1, count + 1} -- allow
Cost: 4 commands, one round trip, ~50–100 µs server time. Memory: one sorted-set member (~40–60 bytes) per request in window — at 1,000 req/min per tenant × 10K tenants that's ~500 MB, which is why most production limiters use a fixed-window counter (INCR + EXPIRE, 1 key per tenant per window) or a token bucket in a hash (2 fields). See Rate Limiting.
Worked Pattern: Leaderboard#
ZADD lb:{global}:2026-09-29 4210 user:77 -- O(log N)
ZINCRBY lb:{global}:2026-09-29 15 user:77 -- atomic score change
ZREVRANGE lb:{global}:2026-09-29 0 99 WITHSCORES -- top 100, O(log N + 100)
ZREVRANK lb:{global}:2026-09-29 user:77 -- my rank, O(log N)
EXPIRE lb:{global}:2026-09-29 691200 -- keep 8 days
A 10M-member sorted set is ~1 GB and ZREVRANK stays ~10–20 µs. The limit is not the structure — it is that one key lives on one shard. At 200K score updates/sec you shard the leaderboard (lb:{global}:{shard_0..15}) and merge top-K at read time: each shard returns its top 100, the app merges 1,600 entries. See Leaderboard & Counting.
🎯 Staff Move: "The sorted set solves ranking in one command. My real design question is the single-shard ceiling: one key tops out around 100K writes a second. Past that, I split the key into 16 sub-leaderboards and merge the top-K on read — exact top-100, approximate global rank."
Worked Pattern: Distributed Lock (and Its Limits)#
SET lock:invoice:991 <uuid> NX PX 30000 -- acquire with 30s lease
-- do work, must finish well under 30s, check elapsed time
-- release only if we still own it (Lua, atomic compare-and-delete):
if redis.call('GET', KEYS[1]) == ARGV[1] then return redis.call('DEL', KEYS[1]) else return 0 end
This lock is efficiency-grade, not correctness-grade. A GC pause longer than the lease, or a failover that loses the SET, lets two holders proceed. Redlock (acquire on a majority of 5 independent primaries) narrows the window but still relies on bounded clock drift and pause times — the public Kleppmann/antirez debate is the canonical reference. If double execution corrupts money or inventory, you need a fencing token checked by the resource (a monotonically increasing version in Postgres: UPDATE ... WHERE version < $token) or a consensus system like etcd. See Distributed Coordination and ZooKeeper & etcd.
The Tunable Tradeoff — Durability × Memory × Latency#
Redis exposes three dials, and every production cluster is a point in this space. The Staff move is to set them per contract, not per company.
Dial 1: Durability (how much acknowledged data can vanish)#
worst_case_loss ≈ max(replication_lag, fsync_interval) + failover_detection_window × write_rate (if split-brain)
| Setting | Typical Loss on Failover | Throughput Cost | Pick When |
|---|---|---|---|
| No persistence, 1 replica | Replication lag (~1–100 ms normally; seconds under load) | None | Cache |
| AOF everysec + replica | Same as above — replication dominates | 5–15% | Ephemeral state |
+ min-replicas-to-write 1, max-lag 10 | Bounded ≤ ~10 s in split-brain | Writes fail when replicas lag | Coordination |
+ WAIT 1 50 per critical write | Near-zero unless failover picks the non-acking replica | +0.5–2 ms per write | Idempotency keys, payment dedup |
| MemoryDB (multi-AZ transaction log) | Zero acknowledged-write loss by design | Writes ~ single-digit ms | System of record |
Dial 2: Memory policy (what happens at maxmemory)#
| Policy | Behavior | Right For | Wrong For |
|---|---|---|---|
noeviction (default) | Writes return OOM errors | Locks, queues, source-of-truth data | A cache — writes start failing at 100% |
allkeys-lru | Evict approx. least-recently-used (samples 5 keys) | General cache | Mixed clusters — evicts your locks |
allkeys-lfu | Evict least-frequently-used, with decay | Skewed caches with a stable hot set | Bursty new content |
volatile-lru / volatile-ttl | Evict only keys with TTL | Mixed cache + state (if you must) | Clusters where nobody sets TTLs — degrades to noeviction |
Dial 3: Read consistency (where reads go)#
Reading from replicas doubles or triples read capacity but reads can be milliseconds to seconds stale and do not see your own writes. Staff default: primary reads for anything that was just written by the same user (read-your-writes), replica reads for fan-out data where staleness under 1 s is invisible.
🎯 Staff Move: "I'll set durability per cluster, not per company. The cache cluster runs no AOF and
allkeys-lfu. The idempotency cluster runs AOF everysec,noeviction,min-replicas-to-write 1, and critical writes useWAIT. Same product, two contracts, two alert policies, two owners."
Anti-Patterns — What Kills Redis Deployments#
1. The Shared "Company Redis"#
One cluster serving sessions, feature flags, rate limits, job queues, and a cache for 12 services. Every team's worst command becomes everyone's latency. A cache fill evicts the job queue under allkeys-lru. Nobody can change maxmemory-policy without a 12-team meeting. Fix: one cluster per contract per owning team; a platform-provided provisioning path makes this cheap.
2. O(N) Commands on the Hot Path#
KEYS, SMEMBERS, HGETALL, LRANGE 0 -1, ZRANGE 0 -1 on large collections; DEL on big keys (freeing 1M elements takes ~100s of ms). Fix: SCAN/HSCAN/SSCAN with COUNT 100–1000; UNLINK (lazy free in a background thread) instead of DEL; lazyfree-lazy-eviction yes; rename or ACL-deny KEYS, FLUSHALL, DEBUG in production.
3. Big Keys#
A single key over ~10 MB, or a collection over ~100K elements. It pins one shard's memory, stalls replication and slot migration, and each read ships megabytes. Detection: redis-cli --bigkeys / --memkeys in off-peak, MEMORY USAGE key. Fix: split by sub-key (feed:{user}:{page}), cap on write.
4. No TTL, No Owner#
Keys written without expiry by a service that was decommissioned 18 months ago. Common finding in audits: 30–60% of memory in a long-lived shared cluster belongs to keys nobody reads. Fix: a lint in the client library that rejects SET without EX/PX for cache-contract clusters; periodic OBJECT IDLETIME sampling.
5. Redis as the Only Copy of Money-Adjacent State#
Balance counters, inventory decrements, or order state in Redis with no upstream log. One failover drops the last ~500 ms of acknowledged writes and there is nothing to reconcile against. Fix: Redis holds the fast path; Postgres or Kafka holds the ledger; a reconciler compares them. See Flash Sales and Payment Processing.
6. Pub/Sub as a Message Queue#
Pub/Sub is fire-and-forget: a subscriber that disconnects for 200 ms misses every message published in that window, and a slow subscriber is disconnected once its output buffer exceeds client-output-buffer-limit pubsub 32mb 8mb 60. Fix: Streams for anything that must be delivered; Kafka for anything that must be replayed across days.
7. Unbounded Connection Counts#
2,000 pods × 50-connection pools = 100K connections to a shard with maxclients 10000. Each connection also costs ~20–40 KB of buffers. Fix: pools of 5–20 per pod, a proxy tier (Envoy, twemproxy-style, or managed proxy) in front of clusters with > 5K clients.
| Anti-Pattern | Detection Signal | Blast Radius | Who Pays |
|---|---|---|---|
| Shared company Redis | latency_percentiles_usec_* p99 jumps correlated with one client | Every tenant service | All on-calls, none owns root cause |
| O(N) on hot path | SLOWLOG GET entries > 10 ms | Entire shard | Latency SLO of every caller |
| Big keys | --bigkeys, replication lag spikes | One shard + its replicas | Owner of co-located keys |
| No TTL | expired_keys flat while used_memory grows | Cluster OOM / eviction | Whoever runs out of memory first |
| Redis as sole ledger | Reconciliation gaps after failover | Customer data | Finance, support, trust |
| Pub/Sub as queue | client_output_buffer disconnects | Silent message loss | Downstream consumers |
The Technology Landscape / Head-to-Head Comparison#
| Dimension | Redis OSS / Redis 8 | Valkey | Memcached | AWS MemoryDB | DragonflyDB | KeyDB |
|---|---|---|---|---|---|---|
| Model | Data structures, single-threaded execution | Redis 7.2 fork, same model, multithreaded I/O improvements | Strings only, multithreaded | Redis/Valkey API + durable multi-AZ log | Redis API, shared-nothing multithreaded | Redis fork, multithreaded |
| Throughput / node | 100K–200K ops/s/core; ~1M pipelined | Similar; higher with I/O threads | 1M+ ops/s on many cores | Similar reads; writes ~ms | Millions/s on many cores (vendor claims) | Higher than Redis on many cores |
| Durability | Async replication, AOF | Same | None | Durable acknowledged writes | Snapshots | Async |
| Clustering | Redis Cluster, 16,384 slots | Same | Client-side hashing | Managed cluster | Single-node vertical focus | Active-replica option |
| License (2026) | RSALv2/SSPLv1, AGPLv3 option from Redis 8 | BSD-3, Linux Foundation | BSD | Managed service | BSL | BSD |
| Pick when | Richest feature set, Redis Stack modules | Default OSS choice after the 2024 relicense; ElastiCache/Memorystore support it | Pure cache, huge multi-core boxes | Redis semantics as system of record | Very large single-node working sets | Legacy |
Redis vs Kafka: Redis Streams handle 10K–100K msgs/sec with retention bounded by RAM (hours to days). Kafka handles millions/sec with retention bounded by disk (days to forever) and replay by offset. If the words "replay", "audit", or "reprocess last week" appear, it's Kafka.
Redis vs DynamoDB: both answer key-value lookups. Redis: 0.5 ms, RAM-priced ($5–10 per GB-month managed), async durability. DynamoDB: 5–10 ms, disk-priced ($0.25 per GB-month plus request units), durable across 3 AZs. The crossover is working-set size and read rate: hot, small, high-QPS → Redis; large, cold, durable → DynamoDB (optionally with DAX).
🎯 Staff Insight: Since the March 2024 license change, "Redis" in an interview means the protocol and data model, not the vendor. Saying "Redis-compatible — Valkey on ElastiCache, or MemoryDB where we need durability" signals you track the ecosystem and understand that the license is now an architectural input for anyone offering Redis as a service.
Patterns#
Pattern 1: Cache-Aside with Stampede Protection#
The default. App reads Redis, on miss reads the DB and populates with TTL. Add three guards: TTL jitter (±10–20%) so keys don't expire together, single-flight (SET lock:k NX PX 2000; losers wait 50 ms and re-read), and stale-while-revalidate (store soft_expiry inside the value; serve stale while one caller refreshes). Full treatment in Caching Fundamentals and Distributed Caching.
Pattern 2: Write-Behind Counters#
High-frequency increments (INCRBY views:{post}) in Redis; a flusher drains to the DB every 5–30 s via GETDEL or a snapshot. Cuts DB writes by 100–1,000×. Loss contract: up to one flush interval plus replication lag. Fine for view counts, wrong for inventory.
Pattern 3: Reliable Queue with Streams#
XADD jobs:{q1} MAXLEN ~ 1000000 * type resize img 123
XGROUP CREATE jobs:{q1} workers $ MKSTREAM
XREADGROUP GROUP workers w-7 COUNT 10 BLOCK 5000 STREAMS jobs:{q1} >
XACK jobs:{q1} workers 1727600000000-0
XAUTOCLAIM jobs:{q1} workers w-9 60000 0-0 COUNT 100 -- steal jobs idle > 60s
At-least-once delivery, per-consumer pending lists, redelivery after crash. Consumers must be idempotent. Poison messages: check XPENDING delivery count; after 5 deliveries move to jobs:{q1}:dlq. See Message Queues.
Pattern 4: Fan-Out Timeline Cache#
Precomputed home timelines as capped lists or sorted sets per user (LPUSH tl:{user} post_id + LTRIM tl:{user} 0 799). Twitter publicly described serving home timelines from Redis this way, with fan-out-on-write for most users and fan-out-on-read for high-follower accounts. Memory math: 300M active users × 800 IDs × ~10 bytes ≈ 2.4 TB before overhead — which is why only active users get materialized timelines. See News Feed.
Pattern 5: Presence and Ephemeral Sessions#
SET presence:{user} {node_id} EX 60 refreshed by heartbeat every 20–30 s; expiry is the "offline" signal. Keyspace notifications are tempting for "user went offline" events but are fire-and-forget Pub/Sub — use them as hints, not truth. See Real-Time Updates.
Pattern 6: Near-Cache in Front of Redis#
For keys read > 50K times/sec (config, feature flags, a celebrity profile), put a 1–5 s in-process cache in front of Redis. Redis 6+ client-side caching (CLIENT TRACKING) pushes invalidations to clients. This is the standard hot-key fix — it moves load from one shard to N app pods.
| Pattern | Contract | Loss Tolerance | Primary Risk | Owner |
|---|---|---|---|---|
| Cache-aside | Cache | Total | Stampede on origin | Service team |
| Write-behind counters | Ephemeral state | One flush interval | Double-count on flush retry | Service team |
| Streams queue | Coordination | Zero after ack | PEL growth, poison messages | Service team + platform alerting |
| Fan-out timelines | Cache (rebuildable) | Total, rebuild cost high | Memory growth with users | Feed team |
| Presence | Ephemeral | Seconds | Flapping on network blips | Real-time team |
| Near-cache | Cache | Total | Staleness up to TTL | Service team |
Scaling#
Vertical First, Then Shard#
A single Redis primary on a modern core serves ~100K–200K ops/sec with sub-ms p99, and 25 GB of data comfortably. Most services never need more. When you do:
| Pressure | First Move | Second Move | Third Move |
|---|---|---|---|
| Read QPS | Pipelining / MGET batching (5–10× fewer round trips) | Read replicas (2–5) | Near-cache for hot keys |
| Write QPS | Pipelining | Redis Cluster, more primaries | Split hot keys (key:{0..15}) |
| Memory | Encoding tuning (listpack thresholds), shorter TTLs | Cluster with more shards | Tier cold data to DynamoDB/SSD-backed store |
| Connections | Smaller pools | Proxy tier | Fewer, larger app processes |
Cluster Sizing Heuristics#
- Shard size: 10–25 GB per primary. Smaller shards fork faster, resync faster, and migrate faster.
- Shard count: start at 3–6 primaries; plan ~50% headroom on memory and CPU.
- Replicas: 1 per primary for ephemeral, 2 for coordination (survive one replica loss during maintenance).
- Network: a shard doing 150K ops/sec × 1 KB values moves ~1.2 Gbps — value size, not op count, often saturates the NIC first.
Hot Keys#
A hot key lives on one shard and scales with nothing. Detection: redis-cli --hotkeys (requires LFU policy), per-shard CPU skew > 2×, client-side key sampling. Mitigations, in order: near-cache (1–5 s), replica reads for that key, key splitting (counter:{k}:{0..N} with read-time sum), and request coalescing at the app tier.
Multi-Region#
Open-source Redis has no multi-primary replication. Options:
| Option | Mechanism | Consistency | Pick When |
|---|---|---|---|
| Regional caches, independent | Each region caches from its own DB replica | Per-region; invalidation via CDC/Kafka | Default for caches |
| Primary in one region, replicas elsewhere | Cross-region async replication | Remote reads stale by 50–200 ms | Read-mostly global data |
| ElastiCache Global Datastore | Managed cross-region replication, one writer region | Async; promote on regional failure | Managed DR |
| Redis Enterprise Active-Active | CRDT-based multi-primary | Eventually consistent, conflict-free types | Truly global writes, vendor-accepted |
🎯 Staff Move: "I don't replicate caches across regions. Each region's cache fills from that region's database and gets invalidated by the same change stream. Replicating a cache means paying twice to replicate derived data and debugging two invalidation paths."
Failure Modes & Recovery#
1. Fork Stall and Copy-on-Write OOM#
Symptom: p99 latency spikes to 500 ms–2 s every 15 minutes; occasionally the kernel OOM-killer terminates redis-server during BGSAVE.
Root cause: RDB snapshot or AOF rewrite calls fork() on a 60 GB process. Page-table copy stalls the main thread (latest_fork_usec ≈ 900,000). Heavy writes during the snapshot duplicate pages via copy-on-write, pushing RSS toward 2× and past host memory.
Detection: latest_fork_usec > 100,000; rdb_last_cow_size / aof_last_cow_size growing; mem_fragmentation_ratio and host MemAvailable < 20%.
Fix: Disable RDB on the primary and snapshot from a replica; cap shards at ~25 GB; disable transparent huge pages (THP makes CoW copy 2 MB pages instead of 4 KB); set vm.overcommit_memory=1.
Prevention: Platform standard: maxmemory ≤ 60% of host RAM when persistence is enabled; alert on latest_fork_usec > 250 ms. Owner: platform team for instance shape; service team for data growth.
2. Full-Resync Loop (Replication Storm)#
Symptom: Replica never becomes healthy; primary CPU and network saturated; sync_full counter increments every few minutes.
Root cause: Replica disconnects (network blip, replica restart). Backlog of 1 MB covers ~10 ms of writes at 100 MB/s, so partial resync fails. Full resync starts: fork, RDB transfer takes 4 minutes, meanwhile the replica output buffer exceeds client-output-buffer-limit replica 256mb 64mb 60, the primary drops the replica, and it starts over.
Detection: sync_full increasing; master_link_status:down on replica; connected_slaves flapping; replica master_last_io_seconds_ago rising.
Fix: Raise repl-backlog-size to cover 60+ s of write volume (e.g., 1 GB); raise replica output buffer hard limit to 1–2 GB; use diskless replication (repl-diskless-sync yes) on fast networks.
Prevention: Size backlog = peak write bytes/sec × 60–120 s as a platform default. Owner: platform.
3. Hot Key Melts One Shard#
Symptom: One primary at 100% CPU, the other 11 at 15%; timeouts for every key on that shard, not just the hot one.
Root cause: A viral post, a global feature-flag key, or a shared rate-limit key with a low-cardinality hash tag ({global}) taking 300K reads/sec — above a single core's ceiling.
Detection: Per-shard instantaneous_ops_per_sec skew > 3×; redis-cli --hotkeys; client-side top-K key sampling. The cluster-wide average looks fine — always alert on the max shard, not the mean.
Fix: Immediately: enable in-process near-cache (1–2 s TTL) for that key; route its reads to replicas. Durably: split the key into N sub-keys, add client-side caching (CLIENT TRACKING).
Prevention: Load-test with Zipfian key distribution (s ≈ 1.0), not uniform. Owner: service team that owns the key; platform provides hot-key telemetry.
4. Failover Loses Acknowledged Writes / Split-Brain#
Symptom: After a failover, idempotency checks miss, some orders process twice, lock holders overlap.
Root cause: Async replication. The old primary acknowledged writes the replica never received; or a network partition isolated the primary with some clients, which kept writing to it for cluster-node-timeout (15 s) before it demoted itself — those writes vanish when it rejoins as a replica.
Detection: master_repl_offset gap between old primary and promoted replica (log it at failover); business-level reconciliation diffs; duplicate-execution counters downstream.
Fix: min-replicas-to-write 1 + min-replicas-max-lag 10; WAIT on critical writes; fencing tokens for locks; the durable copy of truth lives outside Redis.
Prevention: Game-day: kill the primary under load and measure lost writes. If the number isn't acceptable to the business owner, the data doesn't belong in async Redis. Owner: service team (correctness), with product sign-off on the loss budget.
5. Memory Exhaustion → Eviction Storm or OOM Errors#
Symptom (allkeys-lru): Hit rate drops from 97% to 60% in minutes; origin DB CPU climbs to 100% — a cache miss storm. Symptom (noeviction): writes fail with OOM command not allowed; locks can't be acquired; queues stop accepting jobs.
Root cause: A new feature writes 10× larger values, a TTL-less key family grows unbounded, or fragmentation (mem_fragmentation_ratio > 1.5) wastes 30%+ of RAM.
Detection: used_memory / maxmemory > 85%; evicted_keys rate > 0 on a coordination cluster (should always be 0); keyspace_hits / (hits + misses) drop > 5 points in 10 min.
Fix: Add shards (online reshard takes minutes to hours); enable activedefrag yes; shorten TTLs for the offending key family; shed load at the origin with degraded mode.
Prevention: Per-key-family memory budgets reviewed at launch; alert at 75%. Owner: service team for key growth; platform for capacity headroom.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Fork stall | latest_fork_usec > 250 ms | One shard, periodic | Snapshot from replica, smaller shards | Platform |
| Resync loop | sync_full rising | Shard loses redundancy | Bigger backlog + buffers | Platform |
| Hot key | Max-shard ops/sec skew > 3× | All keys on that shard | Near-cache, key split | Service team |
| Failover write loss | Offset gap at promotion | Correctness of recent writes | WAIT, min-replicas, fencing | Service + product |
| Memory exhaustion | used_memory > 85% of max | Cluster or origin DB | Reshard, TTL fix, shed load | Service + platform |
| Slow command | SLOWLOG > 10 ms | Entire shard | Kill client, ACL-deny command | Service team |
| Connection storm | connected_clients > 80% of maxclients | Shard refuses clients | Proxy tier, smaller pools | Service team |
When to Use vs. Alternatives#
| Requirement | Pick | Why | Not Redis Because |
|---|---|---|---|
| Sub-ms reads of hot, rebuildable data | Redis | RAM + O(1) structures | — |
| Atomic counters, rate limits, top-K | Redis | Server-side atomic ops, no locks | — |
| Short-lived job queue (< 1M in flight) | Redis Streams | Consumer groups, ack, simple ops | — |
| Event log with replay, weeks of retention | Kafka | Disk-backed, offset replay | RAM-bounded retention |
| Leader election, config with correctness | etcd / ZooKeeper | Linearizable, consensus | Async replication, lease timing |
| Durable key-value > 500 GB | DynamoDB / Cassandra | Disk-priced, durable | $5–10/GB-month RAM |
| Transactions across entities | PostgreSQL | ACID | No cross-slot atomicity |
| Full-text search | Elasticsearch | Inverted index, relevance | RediSearch exists but ops/licensing differ |
| Pure large-object cache on 32-core boxes | Memcached | Multithreaded, simpler | Single-threaded per shard |
When NOT to use Redis: when the data cannot be rebuilt and you won't pay for a durable variant; when the working set is > 1 TB and mostly cold; when you need cross-key transactions; when correctness depends on a lock never being held twice.
Operational Concerns#
The Config Every Production Cluster Should Have#
maxmemory <= 60% of host RAM when persisting, 80% when not
maxmemory-policy set explicitly per contract — never inherit the default silently
repl-backlog-size 60–120 s of peak write bytes (256 MB – 1 GB)
client-output-buffer-limit replica 1gb 256mb 120
timeout 300 # drop idle clients
tcp-keepalive 60
lazyfree-lazy-eviction yes
lazyfree-lazy-expire yes
activedefrag yes (if fragmentation ratio > 1.3 sustained)
rename-command / ACLs deny KEYS, FLUSHALL, FLUSHDB, DEBUG, CONFIG for app users
latency-monitor-threshold 10
slowlog-log-slower-than 10000 # 10 ms
The Dashboard the On-Call Actually Uses#
| Metric | Healthy | Page |
|---|---|---|
| Max-shard p99 command latency | < 1 ms | > 5 ms for 5 min |
used_memory / maxmemory (max shard) | < 75% | > 90% |
evicted_keys rate on coordination clusters | 0 | Any |
| Cache hit ratio | > 90% (workload-specific) | Drop > 10 points in 10 min |
master_link_status | up | down > 60 s |
latest_fork_usec | < 100 ms | > 500 ms |
connected_clients / maxclients | < 50% | > 80% |
rejected_connections | 0 | Any sustained |
Upgrades and Maintenance#
Rolling upgrade in cluster mode: upgrade replicas first, CLUSTER FAILOVER (manual, coordinated — zero loss because the replica catches up before promotion), upgrade the old primary, repeat per shard. Expect 1–3 s of client-visible redirects per shard. Managed services (ElastiCache, Memorystore) do the same but pick the maintenance window — pin it.
What the On-Call Actually Does#
- Check which shard — never debug the cluster average.
SLOWLOG GET 20,CLIENT LISTsorted bycmdandomem— find the offending client.INFO memory,INFO persistence,INFO replication— rule out fork, fragmentation, resync.- Kill the client (
CLIENT KILL ID), not the server. Restarting a primary triggers failover and a cold replica. - If the origin DB is melting from a miss storm, turn on request coalescing or serve stale — see Degraded Mode.
Interview Application — Staff-Level Plays#
Which Case Studies Use Redis#
| Case Study | How Redis Is Used | Key Pattern |
|---|---|---|
| Rate Limiting | Shared counters for distributed limits | Lua token bucket / fixed window, local-first with Redis sync |
| Distributed Caching | The cache tier itself | Cache-aside, stampede control, hot-key near-cache |
| Leaderboard & Counting | Rankings and counters | Sorted sets, sharded top-K merge |
| Flash Sales | Inventory gate and virtual queue | DECR with floor check in Lua, ledger in Postgres |
| Real-Time Updates | Presence, connection routing | TTL keys, sharded Pub/Sub for fan-out hints |
| Chat Messaging | Online status, recent-message cache | TTL presence, capped lists |
| News Feed | Materialized timelines | Capped lists per user, hybrid fan-out |
| Proximity Matching | Driver locations | GEOADD / GEOSEARCH with TTL cleanup |
Every System Design Question Has a Redis Moment#
- URL shortener: cache the top 1% of short codes — they take ~90% of redirects;
allkeys-lfuwith a 24 h TTL. - Ticketing: hold seats with
SET hold:{event}:{seat} {user} NX EX 600— expiry releases abandoned holds; the purchase commits in Postgres. - Job scheduler: sorted set by due time (
ZADD due now+delay job), workersZRANGEBYSCORE ... LIMIT 0 100+ atomic claim in Lua. - Notification dedup:
SET dedup:{event_id} 1 NX EX 86400— duplicates within 24 h are dropped, and the dedup cluster runsnoeviction.
L5 → L6 → L7 Responses#
| Scenario | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| "Add a cache" | "Cache-aside in Redis with a TTL." | "Cache-aside, TTL 5 min ±15% jitter, single-flight on miss, allkeys-lfu, alert on hit-rate drop. The DB must survive a cold cache at 30% capacity or we need warming." | "What does the cache buy us in $ and SLO? At 40 services each running their own cluster, we pay ~$60K/month and 3 incidents a quarter. I'd offer a paved cache SKU with these defaults and make the cold-cache test a launch gate." |
| "Distributed lock" | "Redlock across 5 nodes." | "SET NX PX with a UUID and Lua release — efficiency only. For correctness, fencing tokens checked by the DB, or etcd." | "Locks are a symptom. I'd push teams to idempotent writes with conditional updates, and make the lock library refuse to exist without a fencing-token parameter." |
| "Redis goes down" | "Replicas fail over." | "Failover is 5–30 s of write loss and ~1 s of acked-data loss. Cache path degrades to DB with a load-shed cap; rate limiter fails open to local limits." | "Which tier-0 services have Redis as a hard dependency? I'd publish a dependency map and require a tested degraded mode for each before the next peak season." |
| "Scale to 10× traffic" | "Add shards." | "Shard, but first check hot keys — max-shard CPU, not average. Split hot keys and add near-cache." | "Is RAM the right tier at 10×? Price the cold 80% on DynamoDB vs RAM: if it saves $40K/month, we tier." |
| "Which Redis?" | "Redis." | "Valkey on ElastiCache for cache; MemoryDB for the idempotency store that needs durability." | "The 2024 relicense is a vendor-risk event. Standardize on the protocol, keep modules out of the paved road, and keep a documented exit to Valkey." |
Why "Redis goes down" separates levels
The L5 answer describes the mechanism (replicas, failover) and is correct. The L6 answer quantifies the mechanism (5–30 s unavailable, ~1 s loss) and designs each caller's behavior during that window — which is where outages actually happen. The L7 answer notices that the question isn't about one Redis: it's about how many critical paths share an undeclared hard dependency on a best-effort component, and it turns that into an org-wide requirement with a deadline.
Why "Distributed lock" separates levels
Redlock is a real, documented algorithm, so the L5 answer isn't wrong — it just doesn't ask what happens when the lock is held twice. The L6 answer separates efficiency locks (duplicate work is tolerable) from correctness locks (duplicate work corrupts state) and moves correctness to the resource via fencing. The L7 answer changes the default so teams can't accidentally build the unsafe version.
The Staff Redis Checklist#
- Name the contract: "This cluster is a cache — evictable, rebuildable, no persistence."
- Pick the structure from the operation: "Top-K by score is a sorted set, one
ZADDper update,ZREVRANGEfor reads." - Design the key and bound it: "
lb:{region}:{day}, trimmed to 10K members on write, 8-day TTL." - Check the single-shard ceiling: "Peak is 40K writes/sec on the hottest key — under the ~100K ceiling. If it grows 3×, I split into 16 sub-keys."
- State the loss budget: "Failover loses under ~1 s of writes. We rebuild from Kafka, so that's acceptable."
- Design the degraded mode: "If Redis is unavailable, the leaderboard serves the last snapshot from S3 and writes buffer in Kafka."
🎯 Staff Insight: What NOT to use Redis for: the only copy of money, a lock whose double-holding corrupts data, or an event log anyone will want to replay next month. Saying this unprompted is the strongest Redis signal in an interview — it shows you know where the tool's contract ends.
The Principal Lens#
Why L7 Sees This Problem Differently#
A Staff engineer designs a Redis cluster. A Principal engineer notices the company has 60 of them, provisioned by 25 teams with 25 different maxmemory-policy settings, and that three tier-0 checkout paths take a hard dependency on a component everyone describes as "just a cache." At org scale Redis is not a technology choice; it is a fleet with an implicit contract problem. The work is making the contract explicit — cache, state, coordination, or record — so that durability, eviction, alerting, and ownership follow from a label instead of from whoever created the cluster.
The Org-Level Fault Line#
One managed Redis platform vs. per-team clusters.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Per-team clusters, no standard | Autonomy, no platform bottleneck | 25 config dialects; fork/OOM incidents repeat; no fleet-wide upgrade path | Every on-call rotation, repeatedly |
| Central shared "company Redis" | One team, one config | Noisy neighbors; nobody can change policy; one outage takes out 12 services | All consumers + the platform team |
| Paved road: platform-provisioned, per-team clusters by contract SKU | Isolation + consistent defaults + fleet upgrades | Platform team becomes a dependency; needs 2–3 engineers | Platform headcount, repaid by fewer incidents |
The Principal default is the third row: isolation per team, standardization per contract. The platform owns shapes, configs, telemetry, and upgrades; service teams own keys, TTLs, memory budgets, and degraded modes.
Cost Model#
Assumptions: managed Redis-compatible service, on-demand pricing in a major US region, memory-optimized Graviton-class nodes (~13 GB ≈ $160/month, ~26 GB ≈ $320/month, ~52 GB ≈ $640/month per node, approximate), 1 replica per primary, ~20% headroom. Engineer loaded cost ~$25K/month.
| Scale | Footprint | Infra $/month | People | On-call Load | Dominant Risk |
|---|---|---|---|---|---|
| Startup — 1 service, 10 GB, 50K ops/s | 1 primary + 1 replica (13 GB nodes) | ~$330 | ~0.1 FTE | < 1 page/quarter | Treating it as a database |
| Growth — 8 services, 300 GB, 1M ops/s | ~15 shards × 2 (26 GB nodes) | ~$10K + ~$1–2K cross-AZ transfer | ~0.5–1 FTE | 1–3 pages/month | Hot keys, fork stalls, shared clusters |
| Enterprise — 40 clusters, 5 TB, 10M ops/s | ~125 shards × 2 (52 GB nodes) | ~$160K + ~$10–20K transfer | 2–3 FTE platform + 0.2 FTE per owning team | 5–10 pages/month fleet-wide | Correlated failure: shared config bug, license/vendor event, AZ loss |
Levers at enterprise scale: moving the cold 60–80% of keys to a disk-backed store saves $50–100K/month; reserved nodes save ~30–50%; self-managing on EC2 saves another ~20–30% of infra but costs ~2 FTE — rarely worth it below $300K/month.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversal Cost | Why |
|---|---|---|---|
| Cache TTLs, eviction policy | Two-way | Config push, minutes | Tune freely with metrics |
| Cluster mode vs standalone | Mostly one-way | Client changes, hash-tag audit, multi-key rewrite: weeks | Cross-slot constraints leak into app code |
| Key naming / hash-tag scheme | One-way at scale | Dual-write migration of every key family: months | Every client encodes it |
| Redis as system of record (no upstream log) | One-way | Building the missing ledger retroactively; data already lost | You can't reconcile what you never logged |
| Adopting vendor modules (search, JSON, time-series) | One-way-ish | Rewrite features on another store | Modules tie you to a license and vendor |
| Valkey vs Redis at the protocol level | Two-way (today) | Low while you avoid divergent features | Keep it that way by standard |
The Standard I'd Write#
RFC: In-Memory Data Store Standard (v1)
Scope: All Redis-protocol data stores in production, managed or self-hosted.
MUST
- Declare a contract label (
cache,state,coordination,durable) at provisioning; the label setsmaxmemory-policy, persistence, and alert policy.- Set a TTL on every key in
cacheandstateclusters; client library rejects writes without one.- Keep shard memory ≤ 25 GB and
maxmemory≤ 60% of host RAM when persistence is on.- Document and game-day a degraded mode for every tier-0 caller, re-tested annually.
- Store no money, inventory, or entitlement state in a
cache/statecluster without an authoritative upstream record.SHOULD
- Use protocol-level features only; modules require an architecture review.
- Isolate clusters per owning team; shared clusters require a named owner and quotas.
Exceptions: Filed with the platform team, time-boxed to 2 quarters, with a named VP-level approver for tier-0 services.
Success metrics: Redis-attributed Sev-1/Sev-2 incidents down 50% in 4 quarters; 100% of clusters labeled; zero untagged tier-0 hard dependencies; RAM $/request down 20%.
What I'd Tell the VP#
We spend about $170K a month on in-memory data stores across 40 clusters, and they caused 9 incidents last year — mostly the same three failure patterns repeating in different teams. The fix is not more hardware; it's a standard: every cluster declares whether it's a disposable cache or holds data we can't lose, and the platform sets the safe defaults automatically. That needs two platform engineers for two quarters. In return I expect incidents to halve and about $50K a month in savings from moving rarely used data to cheaper storage. It also removes our exposure to the 2024 Redis license change by keeping us on the open protocol.
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Contracts over technology | "The question isn't Redis or not — it's which of four contracts this data has, and each gets its own cluster and defaults." |
| Prices the tradeoff | "RAM is ~$5–10 per GB-month managed versus ~$0.25 on DynamoDB. The cold 70% of this data is costing us $40K a month to be fast for nobody." |
| Correlated-failure thinking | "Three tier-0 paths depend on one cluster labeled 'cache'. That's a single point of failure disguised as an optimization." |
| Vendor and license risk | "After the relicense I'd standardize on the protocol and keep modules off the paved road so Valkey stays a real exit." |
| Knows when not to standardize | "I wouldn't mandate a shared cluster. Standardize configs and telemetry; keep blast radius per team." |
Staff answers that L7 interviewers find insufficient:
- "I'd set up Sentinel with 3 nodes and AOF everysec." — Correct for one cluster; says nothing about the 40 others or why they diverged.
- "We'll add a near-cache for the hot key." — Fixes today's incident; doesn't create the hot-key telemetry and load-test standard that prevents the next one.
- "Redis is cheap, it's just a cache." — Ignores that fleet RAM spend and incident cost are both line items someone owns.
🧭 Principal Move: "Before we design this cluster, I'd like to know how many clusters like it we already run and why each one exists. If the answer is 'nobody knows,' the highest-leverage design work is the provisioning standard, not this key schema."
In the Wild#
Twitter — Timelines in Redis#
Twitter has publicly described serving home timelines from large Redis clusters: each active user's timeline is a capped list of tweet IDs, populated by fan-out-on-write, with high-follower accounts merged at read time instead. They also modified Redis internally for memory efficiency on these lists. The design trades very large RAM spend for predictable millisecond reads on the hottest page in the product.
Staff insight: The Redis structure is the easy part; the hybrid fan-out rule (who gets pushed, who gets pulled) is the real design decision, and it's driven by the single-key and RAM ceilings of the cache tier.
Instagram — Hash Bucketing for Memory#
Instagram engineering published how they stored ~300M media-ID→user-ID mappings. Naive string keys needed ~21 GB; bucketing IDs into hashes of ~1,000 fields each let Redis use its compact encoding and cut memory to ~5 GB — roughly 4× savings for a change in key design, not infrastructure.
Staff insight: Encoding thresholds are a data-modeling lever. In an interview, "I'd bucket into hashes to stay under the listpack threshold" shows you understand Redis's memory model, not just its API.
Stack Overflow — Redis as a Shared L2 Cache#
Stack Overflow has publicly documented a small, heavily optimized infrastructure in which Redis serves as a shared L2 cache behind per-server in-memory L1 caches, with Pub/Sub used to broadcast cache invalidations. A small number of Redis servers handle the site's caching at low CPU.
Staff insight: The two-tier L1/L2 pattern is the hot-key answer: the in-process cache absorbs the heaviest reads, Redis absorbs the rest, and invalidation is a broadcast hint — acceptable because the data is rebuildable.
Practice Drill#
Prompt: "Your checkout service uses Redis for idempotency keys (
SET idem:{key} {response} NX EX 86400). During a regional network event, Redis failed over and finance found 140 duplicate charges. The team proposes switching toappendfsync always. What do you do?"
Staff Answer
appendfsync always doesn't address the cause. The duplicates came from asynchronous replication: the old primary acknowledged the SET NX, the promoted replica never received it, and retried requests found no key and charged again. Local fsync protects against a single-node crash, not replication loss. Immediate: refund and reconcile the 140 charges against the payment processor's idempotency (Stripe-style providers accept an idempotency key — pass ours through so the provider dedups even if we don't). Short term: move idempotency to a durable store — a Postgres table with a unique constraint on the key (~2–5 ms, acceptable on a checkout path of ~300 ms), or MemoryDB if we need Redis latency with durable acknowledged writes. If Redis must remain as a fast pre-check, keep it but make the database constraint the authority. Guardrails: min-replicas-to-write 1 on any remaining coordination clusters, a game day that kills the primary under load and counts duplicates, and a reconciliation job that alerts within 15 minutes on duplicate charge IDs. Owner: the payments team owns correctness; the platform team owns the contract label that should have flagged this cluster as coordination, not cache.
Why this is L6:
- Correctly identifies async replication — not fsync — as the loss mechanism.
- Moves the correctness guarantee to a component whose contract supports it, with a latency number.
- Adds defense in depth at the provider boundary and a detection loop with a time bound.
What L7 adds:
- Audits the fleet for every other correctness-critical use of async Redis — this is a class of bug, not an incident.
- Makes the contract label mandatory at provisioning so
coordinationclusters get durable defaults automatically. - Prices it for the business: 140 duplicates × average order value + support cost vs ~$2K/month for a durable store.
Quick Reference Card#
Execution: single-threaded commands per shard; 100K–200K ops/s/core; ~1M pipelined
Latency: ~0.1–0.3 ms server, 0.5–1 ms same-AZ round trip
Key overhead: ~50–90 bytes per top-level key; bucket small values into hashes
Encodings: listpack under 128 entries / 64-byte values (hash, zset) — 5–10× smaller
Persistence: AOF everysec ≈ 1 s local loss; failover loss = replication lag
Fork cost: ~10–20 ms per GB RSS; CoW can need up to 2× RAM — maxmemory ≤ 60% host
Cluster: 16,384 slots, CRC16; hash tags {x} for multi-key; ~25 GB per shard
Failover: 5–30 s write unavailability (Sentinel / cluster-node-timeout 15 s)
Backlog: repl-backlog-size = 60–120 s of peak write bytes (not 1 MB)
HyperLogLog: 12 KB, 0.81% error
Pub/Sub: fire-and-forget; Streams for at-least-once; Kafka for replay
Red flags: KEYS *, DEL on big keys, no TTL, shared company cluster,
Redis as sole ledger, Redlock for correctness, alerts on mean shard
Defaults: label the contract → pick policy; noeviction for coordination;
allkeys-lfu for cache; near-cache for keys > 50K reads/s