Hiring BarSupport

Design with Redis — Staff-Level Technology Guide

Technology guide42 min read6 diagrams

Why This Matters#

Redis is not a cache. It is a single-threaded, in-memory data-structure server that happens to be very good at caching — and most Redis incidents come from teams that forgot the first half of that sentence. Every design that touches Redis inherits three facts: one command runs at a time per shard, the working set must fit in RAM, and replication is asynchronous. Those three facts decide whether Redis is your fastest component or your largest blast radius.

It keeps showing up in interviews because it is the default answer to half the hard questions: rate limiting, leaderboards, sessions, distributed locks, presence, deduplication, fan-out timelines, queues. Interviewers know the default, so they do not score you for saying "Redis." They score what comes next: which data structure, what the key looks like, what happens to the 0.5–1 second of acknowledged writes lost on failover, what the hot key does to one shard, and who gets paged when used_memory crosses maxmemory at 2am.

The L5 answer says "we'll put it in Redis." The L6 answer says "sorted set per region keyed lb:{region}:{day}, ZADD on score change, ZREVRANGE 0 99 for top-100, 16 shards, AOF everysec because we can rebuild from Kafka, and we accept ~1s of loss on failover because product signed off that a leaderboard glitch is not a refund." The L7 answer asks whether the org should run 40 unmanaged Redis clusters at all — and what the license change of 2024 means for the paved road.

The 60-Second Pitch#

"Redis gives us sub-millisecond reads and writes — roughly 0.1–0.3 ms server time, 0.5–1 ms round trip in the same AZ — on rich data structures, not just blobs. That matters because the operations we need — increment a counter, add to a sorted set, pop from a list, append to a stream — execute atomically on the server, so we get correctness without distributed locks. One shard does 100K–200K simple ops/sec, over 1M with pipelining. I'd treat it as derived state: the source of truth lives in Postgres or Kafka, Redis can be rebuilt, and we design the degraded mode for when it's gone. If the data can't be rebuilt, I either pay for a durable variant like MemoryDB or I don't put it in Redis."

The Staff-level insight: the decision is never "Redis or not" — it is which Redis contract. Cache (evictable, rebuildable), ephemeral state (sessions, rate-limit counters — loss is tolerable), coordination primitive (locks, leases — correctness depends on timing assumptions), or system of record (durability required). Each contract has a different persistence setting, eviction policy, failover posture, and owner. Mixing two contracts in one cluster is the most common root cause in Redis post-mortems.

ContractExampleEviction PolicyPersistenceLoss ToleranceWho Pays on Loss
CacheProduct page fragmentsallkeys-lru / allkeys-lfuNone or RDBTotal — rebuild from sourceOrigin DB absorbs a miss storm
Ephemeral stateSessions, rate-limit bucketsvolatile-ttlRDB or AOF everysecMinutes of stateUsers re-login; limits reset briefly
CoordinationLocks, leases, idempotency keysnoevictionAOF everysecMust fail safeDouble-execution risk owned by caller
System of recordLeaderboard with no upstream lognoevictionAOF always or MemoryDBZeroCustomer-visible data loss

🎯 Staff Move: "Before I pick a data structure, I want to name the contract. If this Redis is a cache, eviction is a feature. If it's holding idempotency keys, eviction is a correctness bug — so it gets its own cluster with noeviction and a different alert."


Architecture & Internals#

Only the internals that change design decisions.

The Single-Threaded Event Loop#

Redis executes commands on one main thread per process. Since 6.0, network reads/writes can be offloaded to I/O threads (io-threads 4), but command execution is still serial. This is the source of both its superpower and its worst failure mode:

  • Superpower: every command is atomic. INCR, ZADD, LPUSH, SET NX PX need no locks. A Lua script or MULTI/EXEC block runs without interleaving.
  • Failure mode: one slow command stalls every client on that shard. KEYS * on 50M keys, SMEMBERS on a 2M-member set, DEL on a 1 GB hash, or a Lua script looping for 800 ms — each is a shard-wide outage for its duration.

The mental model: Redis throughput is ~100K–200K ops/sec per core, and latency is the sum of everything queued ahead of you. An O(N) command on a large N is not slow for one client; it is slow for all of them.

Diagram: The Single-Threaded Event Loop

Encodings: Why Small Collections Are Cheap#

Redis stores small collections in compact contiguous encodings and converts them to pointer-heavy structures past a threshold. The thresholds are tunable and they change memory by 5–10×:

TypeCompact EncodingConverts When (defaults, 7.x)Large EncodingMemory Impact
Hashlistpack> 128 fields or any value > 64 byteshashtable~5–10× more per field
Sorted setlistpack> 128 members or member > 64 bytesskiplist + hashtable~3–5× more; O(log N) ops
Setintset (integers) / listpack> 512 integers / > 128 membershashtable~4–8× more
Listquicklist of listpacks— (always quicklist)quicklistNodes of ~8 KB
Stringint / embstr (≤ 44 bytes) / rawlengthraw SDS~50–70 bytes overhead per key

Why it matters: Instagram publicly described storing ~300M photo→user mappings by bucketing them into hashes of ~1,000 fields each (HSET media:{id/1000} {id} {user}), cutting memory from ~21 GB to ~5 GB. The trick works because each top-level key carries ~50–90 bytes of overhead (dict entry, robj, expire entry), while a field inside a listpack hash costs a few bytes plus the payload.

Persistence: RDB, AOF, and the fork()#

ModeMechanismLoss Window on CrashCostWhen
NonePure memoryEverythingZeroPure cache, rebuilt from origin
RDBfork() + child writes a point-in-time snapshotSince last snapshot (often 5–15 min)Fork pause + copy-on-write memoryCache with warm restart; backups
AOF everysecAppend every write; fsync once per second~1 s (up to 2 s under disk stall)~5–15% throughput; AOF rewrite forksThe production default for state
AOF alwaysfsync per write batch~0 on the node2–10× lower write throughputRarely worth it — replication loss dominates
RDB + AOF (preamble)AOF rewrite starts with RDB snapshot~1 sFaster restarts than pure AOFDefault for state since 4.0

The fork() is the hidden cost. The kernel must copy page tables before the child runs — roughly 10–20 ms per GB of resident memory on typical cloud VMs, worse with some hypervisors and without huge-page tuning. A 50 GB instance can pause for ~0.5–1 s every snapshot. Copy-on-write then duplicates every page the parent modifies while the child writes: a write-heavy 50 GB instance can need 20–40 GB of extra RAM during BGSAVE. Rule: size instances so maxmemory is ≤ 50–60% of host RAM when persistence or replication is on, and prefer many 10–25 GB shards over few 100 GB ones.

🎯 Staff Insight: "Persistence" on a replicated Redis mostly protects against simultaneous loss of primary and replicas. A single-node crash is handled by failover — and failover loses whatever the replica hadn't received, regardless of appendfsync always on the primary. Local fsync does not fix asynchronous replication.

Replication and Failover#

Replication is asynchronous. The primary acknowledges the client, then streams the command to replicas. Replicas track an offset; a short disconnect resumes via PSYNC from the replication backlog (repl-backlog-size, default 1 MB — far too small; set 256 MB–1 GB for busy primaries). If the replica falls behind the backlog, it needs a full resync: the primary forks, writes an RDB, ships it, then streams buffered writes. Full resync of a 25 GB dataset over 10 Gbps takes minutes and costs the primary a fork plus a replica output buffer that can reach GBs.

Two knobs bound loss:

  • WAIT numreplicas timeout — block until N replicas acknowledge the offset. Reduces loss, does not make Redis linearizable (a failover can still pick a replica that did not ack).
  • min-replicas-to-write 1 + min-replicas-max-lag 10 — the primary refuses writes if no replica is within 10 s. Bounds split-brain loss to ~10 s instead of unbounded.
Diagram: Replication and Failover

Sentinel (for non-cluster deployments) runs 3+ processes that vote a primary down after down-after-milliseconds (commonly 5,000–30,000) and promote a replica. Redis Cluster does the same internally via its gossip bus, with cluster-node-timeout (default 15,000 ms). Either way, expect 5–30 s of write unavailability per failover plus the replication-lag loss window.

Redis Cluster: 16,384 Slots#

Cluster mode partitions keys by CRC16(key) mod 16384. Each primary owns a range of slots; clients cache the slot map and follow MOVED (permanent) and ASK (mid-migration) redirects.

FactConsequence for Design
Multi-key commands (MGET, SUNIONSTORE, Lua, MULTI) require all keys in one slotUse hash tags: cart:{user42}:items and cart:{user42}:meta hash only user42
Hash tags co-locate by designA hash tag with low cardinality ({global}) recreates a single-node bottleneck
Resharding moves slots key-by-key with MIGRATEBig keys (> 100 MB) block migration and the source shard
Practical ceiling ~ 500–1,000 nodesGossip overhead grows with N²; beyond that, run multiple clusters
Cross-slot atomicity does not existDesign operations to be single-key or single-tag
Diagram: Redis Cluster: 16,384 Slots

Lua Scripts and Functions: Atomicity Has a Budget#

EVAL/EVALSHA (and Redis 7 FUNCTIONs) run a script atomically on the shard — nothing interleaves. That makes Lua the right tool for read-modify-write sequences (check-then-decrement inventory, token-bucket refill, compare-and-delete lock release). It is also a stop-the-world event for that shard.

RuleWhyNumber
Keep scripts O(1) or O(small N)The shard serves nothing else while the script runsTarget < 1 ms; alert > 5 ms
Declare every key in KEYS[]Cluster routing and slot checks rely on itAll keys must share one slot
No unbounded loops over collectionsA loop over 1M members blocks for ~100s of msCap iterations or use SCAN from the client
Script timeout is not a kill switchAfter busy-reply-threshold (default 5 s) clients get BUSY; a script that already wrote can only be stopped with SHUTDOWN NOSAVETreat 5 s as an outage, not a limit
Load once, call by SHA / function nameAvoids resending script bodiesSaves bandwidth on 100K+ calls/s

MULTI/EXEC gives atomic batching without logic; WATCH adds optimistic concurrency (the EXEC aborts if a watched key changed). Under contention above ~5–10% conflict rate, WATCH retry loops waste more than a Lua script costs. See Dealing with Contention.

Pub/Sub vs Streams#

Both move messages; their contracts are opposite.

PropertyPub/Sub (PUBLISH / SPUBLISH)Streams (XADD / XREADGROUP)
StorageNone — delivered to connected subscribers, then gonePersisted in the keyspace (RAM), replicated, AOF-able
Offline subscriberMisses everythingResumes from last-delivered ID
AcknowledgmentNonePer-consumer Pending Entries List + XACK
Slow consumerDisconnected at output-buffer limit (default hard 32 MB)Falls behind; memory grows until MAXLEN trims
Fan-out cost in clusterClassic Pub/Sub broadcasts across the whole cluster bus; sharded Pub/Sub (7.0+) stays on one shardStream lives on one shard (one key)
Throughput100K+ msgs/s per shard, fan-out multiplies cost~50K–150K XADD/s per shard
Right forCache-invalidation hints, presence pings, "something changed" nudgesJob queues, activity feeds, small event pipelines
Diagram: Pub/Sub vs Streams

Memory: Where the Bytes Actually Go#

used_memory is not your payload. For 100M small string keys (20-byte key, 20-byte value), expect ~8–10 GB, not 4 GB: dict entries, robj headers, SDS headers, expire-dict entries for TTL keys (~30–40 more bytes each), and allocator rounding. Add fragmentation (mem_fragmentation_ratio = RSS / used; 1.0–1.3 healthy, > 1.5 means 33%+ waste — activedefrag reclaims it slowly), client buffers (output buffers for slow readers and replicas can reach GBs), and the replication backlog. Budget maxmemory for data and headroom separately, and measure with MEMORY USAGE on samples before launch rather than estimating from payload size.


Data Modeling / Core Usage — "The Entire Game"#

In Redis, the data structure is the algorithm. Picking the right type turns a distributed-systems problem into one atomic command. Picking the wrong one turns it into an O(N) scan on the hot path.

Pick the Structure from the Operation#

You Need To...StructureCore CommandsComplexityWatch Out For
Cache a blob / counterStringSET k v EX 300, INCRO(1)Values > 100 KB hurt network and latency
Store an object with partial updatesHashHSET, HINCRBY, HGETALLO(1) per fieldHGETALL on 10K-field hashes
Rank / top-K / time-windowSorted setZADD, ZREVRANGE, ZRANGEBYSCORE, ZREMRANGEBYSCOREO(log N + M)Unbounded growth — trim on write
Simple FIFO work queueListLPUSH, BRPOP, LMOVEO(1)No ack — crash between pop and process loses work
Durable log with consumer groupsStreamXADD, XREADGROUP, XACK, XAUTOCLAIMO(1) appendPending-entries list grows if consumers don't ack
Membership / dedupSetSADD, SISMEMBERO(1)SMEMBERS on big sets blocks
Cardinality estimateHyperLogLogPFADD, PFCOUNTO(1), 12 KB max0.81% standard error — not for billing
Per-user boolean flagsBitmapSETBIT, BITCOUNTO(1) / O(N)Sparse IDs waste memory — offset = ID
Nearby searchGeo (sorted set)GEOADD, GEOSEARCHO(N + log M)Radius search cost grows with density

Key Design Rules#

  1. Namespace with colons, put the tenant or entity first: rl:{tenant}:{route}:{window}. Keys are grep-able in --bigkeys output and debuggable at 3am.
  2. Hash-tag only what must be atomic together. {user42} for a user's cart and its metadata; never {hot} for global counters.
  3. Every key gets a TTL unless it is explicitly a system-of-record key. Untagged keys are how a cache cluster quietly becomes a database nobody owns.
  4. Bound every collection on write. ZADD then ZREMRANGEBYRANK k 0 -1001 keeps a top-1000; XADD ... MAXLEN ~ 100000 caps a stream.
  5. Keep values under ~10 KB; hard ceiling ~1 MB. A 5 MB value takes ~4 ms of network time on 10 Gbps and blocks the event loop while serializing.

Worked Pattern: Sliding-Window Rate Limiter#

-- KEYS[1] = rl:{tenant42}:POST:/orders    ARGV = now_ms, window_ms, limit, request_id
-- Runs atomically: no other command interleaves on this shard.
local key, now, window, limit = KEYS[1], tonumber(ARGV[1]), tonumber(ARGV[2]), tonumber(ARGV[3])
redis.call('ZREMRANGEBYSCORE', key, 0, now - window)      -- drop expired entries
local count = redis.call('ZCARD', key)
if count >= limit then
  return {0, count}                                        -- reject
end
redis.call('ZADD', key, now, ARGV[4])                      -- record this request
redis.call('PEXPIRE', key, window)                         -- bound memory
return {1, count + 1}                                      -- allow

Cost: 4 commands, one round trip, ~50–100 µs server time. Memory: one sorted-set member (~40–60 bytes) per request in window — at 1,000 req/min per tenant × 10K tenants that's ~500 MB, which is why most production limiters use a fixed-window counter (INCR + EXPIRE, 1 key per tenant per window) or a token bucket in a hash (2 fields). See Rate Limiting.

Worked Pattern: Leaderboard#

ZADD   lb:{global}:2026-09-29  4210  user:77        -- O(log N)
ZINCRBY lb:{global}:2026-09-29  15   user:77        -- atomic score change
ZREVRANGE lb:{global}:2026-09-29 0 99 WITHSCORES     -- top 100, O(log N + 100)
ZREVRANK lb:{global}:2026-09-29 user:77              -- my rank, O(log N)
EXPIRE lb:{global}:2026-09-29 691200                 -- keep 8 days

A 10M-member sorted set is ~1 GB and ZREVRANK stays ~10–20 µs. The limit is not the structure — it is that one key lives on one shard. At 200K score updates/sec you shard the leaderboard (lb:{global}:{shard_0..15}) and merge top-K at read time: each shard returns its top 100, the app merges 1,600 entries. See Leaderboard & Counting.

🎯 Staff Move: "The sorted set solves ranking in one command. My real design question is the single-shard ceiling: one key tops out around 100K writes a second. Past that, I split the key into 16 sub-leaderboards and merge the top-K on read — exact top-100, approximate global rank."

Worked Pattern: Distributed Lock (and Its Limits)#

SET lock:invoice:991 <uuid> NX PX 30000     -- acquire with 30s lease
-- do work, must finish well under 30s, check elapsed time
-- release only if we still own it (Lua, atomic compare-and-delete):
if redis.call('GET', KEYS[1]) == ARGV[1] then return redis.call('DEL', KEYS[1]) else return 0 end

This lock is efficiency-grade, not correctness-grade. A GC pause longer than the lease, or a failover that loses the SET, lets two holders proceed. Redlock (acquire on a majority of 5 independent primaries) narrows the window but still relies on bounded clock drift and pause times — the public Kleppmann/antirez debate is the canonical reference. If double execution corrupts money or inventory, you need a fencing token checked by the resource (a monotonically increasing version in Postgres: UPDATE ... WHERE version < $token) or a consensus system like etcd. See Distributed Coordination and ZooKeeper & etcd.


The Tunable Tradeoff — Durability × Memory × Latency#

Redis exposes three dials, and every production cluster is a point in this space. The Staff move is to set them per contract, not per company.

Dial 1: Durability (how much acknowledged data can vanish)#

worst_case_loss ≈ max(replication_lag, fsync_interval) + failover_detection_window × write_rate (if split-brain)
SettingTypical Loss on FailoverThroughput CostPick When
No persistence, 1 replicaReplication lag (~1–100 ms normally; seconds under load)NoneCache
AOF everysec + replicaSame as above — replication dominates5–15%Ephemeral state
+ min-replicas-to-write 1, max-lag 10Bounded ≤ ~10 s in split-brainWrites fail when replicas lagCoordination
+ WAIT 1 50 per critical writeNear-zero unless failover picks the non-acking replica+0.5–2 ms per writeIdempotency keys, payment dedup
MemoryDB (multi-AZ transaction log)Zero acknowledged-write loss by designWrites ~ single-digit msSystem of record

Dial 2: Memory policy (what happens at maxmemory)#

PolicyBehaviorRight ForWrong For
noeviction (default)Writes return OOM errorsLocks, queues, source-of-truth dataA cache — writes start failing at 100%
allkeys-lruEvict approx. least-recently-used (samples 5 keys)General cacheMixed clusters — evicts your locks
allkeys-lfuEvict least-frequently-used, with decaySkewed caches with a stable hot setBursty new content
volatile-lru / volatile-ttlEvict only keys with TTLMixed cache + state (if you must)Clusters where nobody sets TTLs — degrades to noeviction

Dial 3: Read consistency (where reads go)#

Reading from replicas doubles or triples read capacity but reads can be milliseconds to seconds stale and do not see your own writes. Staff default: primary reads for anything that was just written by the same user (read-your-writes), replica reads for fan-out data where staleness under 1 s is invisible.

🎯 Staff Move: "I'll set durability per cluster, not per company. The cache cluster runs no AOF and allkeys-lfu. The idempotency cluster runs AOF everysec, noeviction, min-replicas-to-write 1, and critical writes use WAIT. Same product, two contracts, two alert policies, two owners."


Anti-Patterns — What Kills Redis Deployments#

1. The Shared "Company Redis"#

One cluster serving sessions, feature flags, rate limits, job queues, and a cache for 12 services. Every team's worst command becomes everyone's latency. A cache fill evicts the job queue under allkeys-lru. Nobody can change maxmemory-policy without a 12-team meeting. Fix: one cluster per contract per owning team; a platform-provided provisioning path makes this cheap.

2. O(N) Commands on the Hot Path#

KEYS, SMEMBERS, HGETALL, LRANGE 0 -1, ZRANGE 0 -1 on large collections; DEL on big keys (freeing 1M elements takes ~100s of ms). Fix: SCAN/HSCAN/SSCAN with COUNT 100–1000; UNLINK (lazy free in a background thread) instead of DEL; lazyfree-lazy-eviction yes; rename or ACL-deny KEYS, FLUSHALL, DEBUG in production.

3. Big Keys#

A single key over ~10 MB, or a collection over ~100K elements. It pins one shard's memory, stalls replication and slot migration, and each read ships megabytes. Detection: redis-cli --bigkeys / --memkeys in off-peak, MEMORY USAGE key. Fix: split by sub-key (feed:{user}:{page}), cap on write.

4. No TTL, No Owner#

Keys written without expiry by a service that was decommissioned 18 months ago. Common finding in audits: 30–60% of memory in a long-lived shared cluster belongs to keys nobody reads. Fix: a lint in the client library that rejects SET without EX/PX for cache-contract clusters; periodic OBJECT IDLETIME sampling.

5. Redis as the Only Copy of Money-Adjacent State#

Balance counters, inventory decrements, or order state in Redis with no upstream log. One failover drops the last ~500 ms of acknowledged writes and there is nothing to reconcile against. Fix: Redis holds the fast path; Postgres or Kafka holds the ledger; a reconciler compares them. See Flash Sales and Payment Processing.

6. Pub/Sub as a Message Queue#

Pub/Sub is fire-and-forget: a subscriber that disconnects for 200 ms misses every message published in that window, and a slow subscriber is disconnected once its output buffer exceeds client-output-buffer-limit pubsub 32mb 8mb 60. Fix: Streams for anything that must be delivered; Kafka for anything that must be replayed across days.

7. Unbounded Connection Counts#

2,000 pods × 50-connection pools = 100K connections to a shard with maxclients 10000. Each connection also costs ~20–40 KB of buffers. Fix: pools of 5–20 per pod, a proxy tier (Envoy, twemproxy-style, or managed proxy) in front of clusters with > 5K clients.

Anti-PatternDetection SignalBlast RadiusWho Pays
Shared company Redislatency_percentiles_usec_* p99 jumps correlated with one clientEvery tenant serviceAll on-calls, none owns root cause
O(N) on hot pathSLOWLOG GET entries > 10 msEntire shardLatency SLO of every caller
Big keys--bigkeys, replication lag spikesOne shard + its replicasOwner of co-located keys
No TTLexpired_keys flat while used_memory growsCluster OOM / evictionWhoever runs out of memory first
Redis as sole ledgerReconciliation gaps after failoverCustomer dataFinance, support, trust
Pub/Sub as queueclient_output_buffer disconnectsSilent message lossDownstream consumers

The Technology Landscape / Head-to-Head Comparison#

DimensionRedis OSS / Redis 8ValkeyMemcachedAWS MemoryDBDragonflyDBKeyDB
ModelData structures, single-threaded executionRedis 7.2 fork, same model, multithreaded I/O improvementsStrings only, multithreadedRedis/Valkey API + durable multi-AZ logRedis API, shared-nothing multithreadedRedis fork, multithreaded
Throughput / node100K–200K ops/s/core; ~1M pipelinedSimilar; higher with I/O threads1M+ ops/s on many coresSimilar reads; writes ~msMillions/s on many cores (vendor claims)Higher than Redis on many cores
DurabilityAsync replication, AOFSameNoneDurable acknowledged writesSnapshotsAsync
ClusteringRedis Cluster, 16,384 slotsSameClient-side hashingManaged clusterSingle-node vertical focusActive-replica option
License (2026)RSALv2/SSPLv1, AGPLv3 option from Redis 8BSD-3, Linux FoundationBSDManaged serviceBSLBSD
Pick whenRichest feature set, Redis Stack modulesDefault OSS choice after the 2024 relicense; ElastiCache/Memorystore support itPure cache, huge multi-core boxesRedis semantics as system of recordVery large single-node working setsLegacy

Redis vs Kafka: Redis Streams handle 10K–100K msgs/sec with retention bounded by RAM (hours to days). Kafka handles millions/sec with retention bounded by disk (days to forever) and replay by offset. If the words "replay", "audit", or "reprocess last week" appear, it's Kafka.

Redis vs DynamoDB: both answer key-value lookups. Redis: 0.5 ms, RAM-priced ($5–10 per GB-month managed), async durability. DynamoDB: 5–10 ms, disk-priced ($0.25 per GB-month plus request units), durable across 3 AZs. The crossover is working-set size and read rate: hot, small, high-QPS → Redis; large, cold, durable → DynamoDB (optionally with DAX).

🎯 Staff Insight: Since the March 2024 license change, "Redis" in an interview means the protocol and data model, not the vendor. Saying "Redis-compatible — Valkey on ElastiCache, or MemoryDB where we need durability" signals you track the ecosystem and understand that the license is now an architectural input for anyone offering Redis as a service.


Patterns#

Pattern 1: Cache-Aside with Stampede Protection#

The default. App reads Redis, on miss reads the DB and populates with TTL. Add three guards: TTL jitter (±10–20%) so keys don't expire together, single-flight (SET lock:k NX PX 2000; losers wait 50 ms and re-read), and stale-while-revalidate (store soft_expiry inside the value; serve stale while one caller refreshes). Full treatment in Caching Fundamentals and Distributed Caching.

Pattern 2: Write-Behind Counters#

High-frequency increments (INCRBY views:{post}) in Redis; a flusher drains to the DB every 5–30 s via GETDEL or a snapshot. Cuts DB writes by 100–1,000×. Loss contract: up to one flush interval plus replication lag. Fine for view counts, wrong for inventory.

Pattern 3: Reliable Queue with Streams#

XADD jobs:{q1} MAXLEN ~ 1000000 * type resize img 123
XGROUP CREATE jobs:{q1} workers $ MKSTREAM
XREADGROUP GROUP workers w-7 COUNT 10 BLOCK 5000 STREAMS jobs:{q1} >
XACK jobs:{q1} workers 1727600000000-0
XAUTOCLAIM jobs:{q1} workers w-9 60000 0-0 COUNT 100   -- steal jobs idle > 60s

At-least-once delivery, per-consumer pending lists, redelivery after crash. Consumers must be idempotent. Poison messages: check XPENDING delivery count; after 5 deliveries move to jobs:{q1}:dlq. See Message Queues.

Pattern 4: Fan-Out Timeline Cache#

Precomputed home timelines as capped lists or sorted sets per user (LPUSH tl:{user} post_id + LTRIM tl:{user} 0 799). Twitter publicly described serving home timelines from Redis this way, with fan-out-on-write for most users and fan-out-on-read for high-follower accounts. Memory math: 300M active users × 800 IDs × ~10 bytes ≈ 2.4 TB before overhead — which is why only active users get materialized timelines. See News Feed.

Pattern 5: Presence and Ephemeral Sessions#

SET presence:{user} {node_id} EX 60 refreshed by heartbeat every 20–30 s; expiry is the "offline" signal. Keyspace notifications are tempting for "user went offline" events but are fire-and-forget Pub/Sub — use them as hints, not truth. See Real-Time Updates.

Pattern 6: Near-Cache in Front of Redis#

For keys read > 50K times/sec (config, feature flags, a celebrity profile), put a 1–5 s in-process cache in front of Redis. Redis 6+ client-side caching (CLIENT TRACKING) pushes invalidations to clients. This is the standard hot-key fix — it moves load from one shard to N app pods.

PatternContractLoss TolerancePrimary RiskOwner
Cache-asideCacheTotalStampede on originService team
Write-behind countersEphemeral stateOne flush intervalDouble-count on flush retryService team
Streams queueCoordinationZero after ackPEL growth, poison messagesService team + platform alerting
Fan-out timelinesCache (rebuildable)Total, rebuild cost highMemory growth with usersFeed team
PresenceEphemeralSecondsFlapping on network blipsReal-time team
Near-cacheCacheTotalStaleness up to TTLService team

Scaling#

Vertical First, Then Shard#

A single Redis primary on a modern core serves ~100K–200K ops/sec with sub-ms p99, and 25 GB of data comfortably. Most services never need more. When you do:

PressureFirst MoveSecond MoveThird Move
Read QPSPipelining / MGET batching (5–10× fewer round trips)Read replicas (2–5)Near-cache for hot keys
Write QPSPipeliningRedis Cluster, more primariesSplit hot keys (key:{0..15})
MemoryEncoding tuning (listpack thresholds), shorter TTLsCluster with more shardsTier cold data to DynamoDB/SSD-backed store
ConnectionsSmaller poolsProxy tierFewer, larger app processes

Cluster Sizing Heuristics#

  • Shard size: 10–25 GB per primary. Smaller shards fork faster, resync faster, and migrate faster.
  • Shard count: start at 3–6 primaries; plan ~50% headroom on memory and CPU.
  • Replicas: 1 per primary for ephemeral, 2 for coordination (survive one replica loss during maintenance).
  • Network: a shard doing 150K ops/sec × 1 KB values moves ~1.2 Gbps — value size, not op count, often saturates the NIC first.

Hot Keys#

A hot key lives on one shard and scales with nothing. Detection: redis-cli --hotkeys (requires LFU policy), per-shard CPU skew > 2×, client-side key sampling. Mitigations, in order: near-cache (1–5 s), replica reads for that key, key splitting (counter:{k}:{0..N} with read-time sum), and request coalescing at the app tier.

Multi-Region#

Open-source Redis has no multi-primary replication. Options:

OptionMechanismConsistencyPick When
Regional caches, independentEach region caches from its own DB replicaPer-region; invalidation via CDC/KafkaDefault for caches
Primary in one region, replicas elsewhereCross-region async replicationRemote reads stale by 50–200 msRead-mostly global data
ElastiCache Global DatastoreManaged cross-region replication, one writer regionAsync; promote on regional failureManaged DR
Redis Enterprise Active-ActiveCRDT-based multi-primaryEventually consistent, conflict-free typesTruly global writes, vendor-accepted

🎯 Staff Move: "I don't replicate caches across regions. Each region's cache fills from that region's database and gets invalidated by the same change stream. Replicating a cache means paying twice to replicate derived data and debugging two invalidation paths."


Failure Modes & Recovery#

1. Fork Stall and Copy-on-Write OOM#

Symptom: p99 latency spikes to 500 ms–2 s every 15 minutes; occasionally the kernel OOM-killer terminates redis-server during BGSAVE.

Root cause: RDB snapshot or AOF rewrite calls fork() on a 60 GB process. Page-table copy stalls the main thread (latest_fork_usec ≈ 900,000). Heavy writes during the snapshot duplicate pages via copy-on-write, pushing RSS toward 2× and past host memory.

Detection: latest_fork_usec > 100,000; rdb_last_cow_size / aof_last_cow_size growing; mem_fragmentation_ratio and host MemAvailable < 20%.

Fix: Disable RDB on the primary and snapshot from a replica; cap shards at ~25 GB; disable transparent huge pages (THP makes CoW copy 2 MB pages instead of 4 KB); set vm.overcommit_memory=1.

Prevention: Platform standard: maxmemory ≤ 60% of host RAM when persistence is enabled; alert on latest_fork_usec > 250 ms. Owner: platform team for instance shape; service team for data growth.

2. Full-Resync Loop (Replication Storm)#

Symptom: Replica never becomes healthy; primary CPU and network saturated; sync_full counter increments every few minutes.

Root cause: Replica disconnects (network blip, replica restart). Backlog of 1 MB covers ~10 ms of writes at 100 MB/s, so partial resync fails. Full resync starts: fork, RDB transfer takes 4 minutes, meanwhile the replica output buffer exceeds client-output-buffer-limit replica 256mb 64mb 60, the primary drops the replica, and it starts over.

Detection: sync_full increasing; master_link_status:down on replica; connected_slaves flapping; replica master_last_io_seconds_ago rising.

Fix: Raise repl-backlog-size to cover 60+ s of write volume (e.g., 1 GB); raise replica output buffer hard limit to 1–2 GB; use diskless replication (repl-diskless-sync yes) on fast networks.

Prevention: Size backlog = peak write bytes/sec × 60–120 s as a platform default. Owner: platform.

3. Hot Key Melts One Shard#

Symptom: One primary at 100% CPU, the other 11 at 15%; timeouts for every key on that shard, not just the hot one.

Root cause: A viral post, a global feature-flag key, or a shared rate-limit key with a low-cardinality hash tag ({global}) taking 300K reads/sec — above a single core's ceiling.

Detection: Per-shard instantaneous_ops_per_sec skew > 3×; redis-cli --hotkeys; client-side top-K key sampling. The cluster-wide average looks fine — always alert on the max shard, not the mean.

Fix: Immediately: enable in-process near-cache (1–2 s TTL) for that key; route its reads to replicas. Durably: split the key into N sub-keys, add client-side caching (CLIENT TRACKING).

Prevention: Load-test with Zipfian key distribution (s ≈ 1.0), not uniform. Owner: service team that owns the key; platform provides hot-key telemetry.

4. Failover Loses Acknowledged Writes / Split-Brain#

Symptom: After a failover, idempotency checks miss, some orders process twice, lock holders overlap.

Root cause: Async replication. The old primary acknowledged writes the replica never received; or a network partition isolated the primary with some clients, which kept writing to it for cluster-node-timeout (15 s) before it demoted itself — those writes vanish when it rejoins as a replica.

Detection: master_repl_offset gap between old primary and promoted replica (log it at failover); business-level reconciliation diffs; duplicate-execution counters downstream.

Fix: min-replicas-to-write 1 + min-replicas-max-lag 10; WAIT on critical writes; fencing tokens for locks; the durable copy of truth lives outside Redis.

Prevention: Game-day: kill the primary under load and measure lost writes. If the number isn't acceptable to the business owner, the data doesn't belong in async Redis. Owner: service team (correctness), with product sign-off on the loss budget.

5. Memory Exhaustion → Eviction Storm or OOM Errors#

Symptom (allkeys-lru): Hit rate drops from 97% to 60% in minutes; origin DB CPU climbs to 100% — a cache miss storm. Symptom (noeviction): writes fail with OOM command not allowed; locks can't be acquired; queues stop accepting jobs.

Root cause: A new feature writes 10× larger values, a TTL-less key family grows unbounded, or fragmentation (mem_fragmentation_ratio > 1.5) wastes 30%+ of RAM.

Detection: used_memory / maxmemory > 85%; evicted_keys rate > 0 on a coordination cluster (should always be 0); keyspace_hits / (hits + misses) drop > 5 points in 10 min.

Fix: Add shards (online reshard takes minutes to hours); enable activedefrag yes; shorten TTLs for the offending key family; shed load at the origin with degraded mode.

Prevention: Per-key-family memory budgets reviewed at launch; alert at 75%. Owner: service team for key growth; platform for capacity headroom.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Fork stalllatest_fork_usec > 250 msOne shard, periodicSnapshot from replica, smaller shardsPlatform
Resync loopsync_full risingShard loses redundancyBigger backlog + buffersPlatform
Hot keyMax-shard ops/sec skew > 3×All keys on that shardNear-cache, key splitService team
Failover write lossOffset gap at promotionCorrectness of recent writesWAIT, min-replicas, fencingService + product
Memory exhaustionused_memory > 85% of maxCluster or origin DBReshard, TTL fix, shed loadService + platform
Slow commandSLOWLOG > 10 msEntire shardKill client, ACL-deny commandService team
Connection stormconnected_clients > 80% of maxclientsShard refuses clientsProxy tier, smaller poolsService team
Diagram: Operational Reality Matrix

When to Use vs. Alternatives#

RequirementPickWhyNot Redis Because
Sub-ms reads of hot, rebuildable dataRedisRAM + O(1) structures—
Atomic counters, rate limits, top-KRedisServer-side atomic ops, no locks—
Short-lived job queue (< 1M in flight)Redis StreamsConsumer groups, ack, simple ops—
Event log with replay, weeks of retentionKafkaDisk-backed, offset replayRAM-bounded retention
Leader election, config with correctnessetcd / ZooKeeperLinearizable, consensusAsync replication, lease timing
Durable key-value > 500 GBDynamoDB / CassandraDisk-priced, durable$5–10/GB-month RAM
Transactions across entitiesPostgreSQLACIDNo cross-slot atomicity
Full-text searchElasticsearchInverted index, relevanceRediSearch exists but ops/licensing differ
Pure large-object cache on 32-core boxesMemcachedMultithreaded, simplerSingle-threaded per shard

When NOT to use Redis: when the data cannot be rebuilt and you won't pay for a durable variant; when the working set is > 1 TB and mostly cold; when you need cross-key transactions; when correctness depends on a lock never being held twice.


Operational Concerns#

The Config Every Production Cluster Should Have#

maxmemory                <= 60% of host RAM when persisting, 80% when not
maxmemory-policy         set explicitly per contract — never inherit the default silently
repl-backlog-size        60–120 s of peak write bytes (256 MB – 1 GB)
client-output-buffer-limit replica 1gb 256mb 120
timeout                  300          # drop idle clients
tcp-keepalive            60
lazyfree-lazy-eviction   yes
lazyfree-lazy-expire     yes
activedefrag             yes (if fragmentation ratio > 1.3 sustained)
rename-command / ACLs    deny KEYS, FLUSHALL, FLUSHDB, DEBUG, CONFIG for app users
latency-monitor-threshold 10
slowlog-log-slower-than  10000        # 10 ms

The Dashboard the On-Call Actually Uses#

MetricHealthyPage
Max-shard p99 command latency< 1 ms> 5 ms for 5 min
used_memory / maxmemory (max shard)< 75%> 90%
evicted_keys rate on coordination clusters0Any
Cache hit ratio> 90% (workload-specific)Drop > 10 points in 10 min
master_link_statusupdown > 60 s
latest_fork_usec< 100 ms> 500 ms
connected_clients / maxclients< 50%> 80%
rejected_connections0Any sustained

Upgrades and Maintenance#

Rolling upgrade in cluster mode: upgrade replicas first, CLUSTER FAILOVER (manual, coordinated — zero loss because the replica catches up before promotion), upgrade the old primary, repeat per shard. Expect 1–3 s of client-visible redirects per shard. Managed services (ElastiCache, Memorystore) do the same but pick the maintenance window — pin it.

What the On-Call Actually Does#

  1. Check which shard — never debug the cluster average.
  2. SLOWLOG GET 20, CLIENT LIST sorted by cmd and omem — find the offending client.
  3. INFO memory, INFO persistence, INFO replication — rule out fork, fragmentation, resync.
  4. Kill the client (CLIENT KILL ID), not the server. Restarting a primary triggers failover and a cold replica.
  5. If the origin DB is melting from a miss storm, turn on request coalescing or serve stale — see Degraded Mode.

Interview Application — Staff-Level Plays#

Which Case Studies Use Redis#

Case StudyHow Redis Is UsedKey Pattern
Rate LimitingShared counters for distributed limitsLua token bucket / fixed window, local-first with Redis sync
Distributed CachingThe cache tier itselfCache-aside, stampede control, hot-key near-cache
Leaderboard & CountingRankings and countersSorted sets, sharded top-K merge
Flash SalesInventory gate and virtual queueDECR with floor check in Lua, ledger in Postgres
Real-Time UpdatesPresence, connection routingTTL keys, sharded Pub/Sub for fan-out hints
Chat MessagingOnline status, recent-message cacheTTL presence, capped lists
News FeedMaterialized timelinesCapped lists per user, hybrid fan-out
Proximity MatchingDriver locationsGEOADD / GEOSEARCH with TTL cleanup

Every System Design Question Has a Redis Moment#

  • URL shortener: cache the top 1% of short codes — they take ~90% of redirects; allkeys-lfu with a 24 h TTL.
  • Ticketing: hold seats with SET hold:{event}:{seat} {user} NX EX 600 — expiry releases abandoned holds; the purchase commits in Postgres.
  • Job scheduler: sorted set by due time (ZADD due now+delay job), workers ZRANGEBYSCORE ... LIMIT 0 100 + atomic claim in Lua.
  • Notification dedup: SET dedup:{event_id} 1 NX EX 86400 — duplicates within 24 h are dropped, and the dedup cluster runs noeviction.

L5 → L6 → L7 Responses#

ScenarioSenior (L5)Staff (L6)Principal (L7)
"Add a cache""Cache-aside in Redis with a TTL.""Cache-aside, TTL 5 min ±15% jitter, single-flight on miss, allkeys-lfu, alert on hit-rate drop. The DB must survive a cold cache at 30% capacity or we need warming.""What does the cache buy us in $ and SLO? At 40 services each running their own cluster, we pay ~$60K/month and 3 incidents a quarter. I'd offer a paved cache SKU with these defaults and make the cold-cache test a launch gate."
"Distributed lock""Redlock across 5 nodes.""SET NX PX with a UUID and Lua release — efficiency only. For correctness, fencing tokens checked by the DB, or etcd.""Locks are a symptom. I'd push teams to idempotent writes with conditional updates, and make the lock library refuse to exist without a fencing-token parameter."
"Redis goes down""Replicas fail over.""Failover is 5–30 s of write loss and ~1 s of acked-data loss. Cache path degrades to DB with a load-shed cap; rate limiter fails open to local limits.""Which tier-0 services have Redis as a hard dependency? I'd publish a dependency map and require a tested degraded mode for each before the next peak season."
"Scale to 10× traffic""Add shards.""Shard, but first check hot keys — max-shard CPU, not average. Split hot keys and add near-cache.""Is RAM the right tier at 10×? Price the cold 80% on DynamoDB vs RAM: if it saves $40K/month, we tier."
"Which Redis?""Redis.""Valkey on ElastiCache for cache; MemoryDB for the idempotency store that needs durability.""The 2024 relicense is a vendor-risk event. Standardize on the protocol, keep modules out of the paved road, and keep a documented exit to Valkey."
Why "Redis goes down" separates levels

The L5 answer describes the mechanism (replicas, failover) and is correct. The L6 answer quantifies the mechanism (5–30 s unavailable, ~1 s loss) and designs each caller's behavior during that window — which is where outages actually happen. The L7 answer notices that the question isn't about one Redis: it's about how many critical paths share an undeclared hard dependency on a best-effort component, and it turns that into an org-wide requirement with a deadline.

Why "Distributed lock" separates levels

Redlock is a real, documented algorithm, so the L5 answer isn't wrong — it just doesn't ask what happens when the lock is held twice. The L6 answer separates efficiency locks (duplicate work is tolerable) from correctness locks (duplicate work corrupts state) and moves correctness to the resource via fencing. The L7 answer changes the default so teams can't accidentally build the unsafe version.

The Staff Redis Checklist#

  1. Name the contract: "This cluster is a cache — evictable, rebuildable, no persistence."
  2. Pick the structure from the operation: "Top-K by score is a sorted set, one ZADD per update, ZREVRANGE for reads."
  3. Design the key and bound it: "lb:{region}:{day}, trimmed to 10K members on write, 8-day TTL."
  4. Check the single-shard ceiling: "Peak is 40K writes/sec on the hottest key — under the ~100K ceiling. If it grows 3×, I split into 16 sub-keys."
  5. State the loss budget: "Failover loses under ~1 s of writes. We rebuild from Kafka, so that's acceptable."
  6. Design the degraded mode: "If Redis is unavailable, the leaderboard serves the last snapshot from S3 and writes buffer in Kafka."

🎯 Staff Insight: What NOT to use Redis for: the only copy of money, a lock whose double-holding corrupts data, or an event log anyone will want to replay next month. Saying this unprompted is the strongest Redis signal in an interview — it shows you know where the tool's contract ends.


The Principal Lens#

Why L7 Sees This Problem Differently#

A Staff engineer designs a Redis cluster. A Principal engineer notices the company has 60 of them, provisioned by 25 teams with 25 different maxmemory-policy settings, and that three tier-0 checkout paths take a hard dependency on a component everyone describes as "just a cache." At org scale Redis is not a technology choice; it is a fleet with an implicit contract problem. The work is making the contract explicit — cache, state, coordination, or record — so that durability, eviction, alerting, and ownership follow from a label instead of from whoever created the cluster.

The Org-Level Fault Line#

One managed Redis platform vs. per-team clusters.

OptionWhat WorksWhat BreaksWho Pays
Per-team clusters, no standardAutonomy, no platform bottleneck25 config dialects; fork/OOM incidents repeat; no fleet-wide upgrade pathEvery on-call rotation, repeatedly
Central shared "company Redis"One team, one configNoisy neighbors; nobody can change policy; one outage takes out 12 servicesAll consumers + the platform team
Paved road: platform-provisioned, per-team clusters by contract SKUIsolation + consistent defaults + fleet upgradesPlatform team becomes a dependency; needs 2–3 engineersPlatform headcount, repaid by fewer incidents

The Principal default is the third row: isolation per team, standardization per contract. The platform owns shapes, configs, telemetry, and upgrades; service teams own keys, TTLs, memory budgets, and degraded modes.

Cost Model#

Assumptions: managed Redis-compatible service, on-demand pricing in a major US region, memory-optimized Graviton-class nodes (~13 GB ≈ $160/month, ~26 GB ≈ $320/month, ~52 GB ≈ $640/month per node, approximate), 1 replica per primary, ~20% headroom. Engineer loaded cost ~$25K/month.

ScaleFootprintInfra $/monthPeopleOn-call LoadDominant Risk
Startup — 1 service, 10 GB, 50K ops/s1 primary + 1 replica (13 GB nodes)~$330~0.1 FTE< 1 page/quarterTreating it as a database
Growth — 8 services, 300 GB, 1M ops/s~15 shards × 2 (26 GB nodes)~$10K + ~$1–2K cross-AZ transfer~0.5–1 FTE1–3 pages/monthHot keys, fork stalls, shared clusters
Enterprise — 40 clusters, 5 TB, 10M ops/s~125 shards × 2 (52 GB nodes)~$160K + ~$10–20K transfer2–3 FTE platform + 0.2 FTE per owning team5–10 pages/month fleet-wideCorrelated failure: shared config bug, license/vendor event, AZ loss

Levers at enterprise scale: moving the cold 60–80% of keys to a disk-backed store saves $50–100K/month; reserved nodes save ~30–50%; self-managing on EC2 saves another ~20–30% of infra but costs ~2 FTE — rarely worth it below $300K/month.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversal CostWhy
Cache TTLs, eviction policyTwo-wayConfig push, minutesTune freely with metrics
Cluster mode vs standaloneMostly one-wayClient changes, hash-tag audit, multi-key rewrite: weeksCross-slot constraints leak into app code
Key naming / hash-tag schemeOne-way at scaleDual-write migration of every key family: monthsEvery client encodes it
Redis as system of record (no upstream log)One-wayBuilding the missing ledger retroactively; data already lostYou can't reconcile what you never logged
Adopting vendor modules (search, JSON, time-series)One-way-ishRewrite features on another storeModules tie you to a license and vendor
Valkey vs Redis at the protocol levelTwo-way (today)Low while you avoid divergent featuresKeep it that way by standard

The Standard I'd Write#

RFC: In-Memory Data Store Standard (v1)

Scope: All Redis-protocol data stores in production, managed or self-hosted.

MUST

  • Declare a contract label (cache, state, coordination, durable) at provisioning; the label sets maxmemory-policy, persistence, and alert policy.
  • Set a TTL on every key in cache and state clusters; client library rejects writes without one.
  • Keep shard memory ≤ 25 GB and maxmemory ≤ 60% of host RAM when persistence is on.
  • Document and game-day a degraded mode for every tier-0 caller, re-tested annually.
  • Store no money, inventory, or entitlement state in a cache/state cluster without an authoritative upstream record.

SHOULD

  • Use protocol-level features only; modules require an architecture review.
  • Isolate clusters per owning team; shared clusters require a named owner and quotas.

Exceptions: Filed with the platform team, time-boxed to 2 quarters, with a named VP-level approver for tier-0 services.

Success metrics: Redis-attributed Sev-1/Sev-2 incidents down 50% in 4 quarters; 100% of clusters labeled; zero untagged tier-0 hard dependencies; RAM $/request down 20%.

What I'd Tell the VP#

We spend about $170K a month on in-memory data stores across 40 clusters, and they caused 9 incidents last year — mostly the same three failure patterns repeating in different teams. The fix is not more hardware; it's a standard: every cluster declares whether it's a disposable cache or holds data we can't lose, and the platform sets the safe defaults automatically. That needs two platform engineers for two quarters. In return I expect incidents to halve and about $50K a month in savings from moving rarely used data to cheaper storage. It also removes our exposure to the 2024 Redis license change by keeping us on the open protocol.

Principal Interview Signals#

SignalWhat It Sounds Like
Contracts over technology"The question isn't Redis or not — it's which of four contracts this data has, and each gets its own cluster and defaults."
Prices the tradeoff"RAM is ~$5–10 per GB-month managed versus ~$0.25 on DynamoDB. The cold 70% of this data is costing us $40K a month to be fast for nobody."
Correlated-failure thinking"Three tier-0 paths depend on one cluster labeled 'cache'. That's a single point of failure disguised as an optimization."
Vendor and license risk"After the relicense I'd standardize on the protocol and keep modules off the paved road so Valkey stays a real exit."
Knows when not to standardize"I wouldn't mandate a shared cluster. Standardize configs and telemetry; keep blast radius per team."

Staff answers that L7 interviewers find insufficient:

  • "I'd set up Sentinel with 3 nodes and AOF everysec." — Correct for one cluster; says nothing about the 40 others or why they diverged.
  • "We'll add a near-cache for the hot key." — Fixes today's incident; doesn't create the hot-key telemetry and load-test standard that prevents the next one.
  • "Redis is cheap, it's just a cache." — Ignores that fleet RAM spend and incident cost are both line items someone owns.

🧭 Principal Move: "Before we design this cluster, I'd like to know how many clusters like it we already run and why each one exists. If the answer is 'nobody knows,' the highest-leverage design work is the provisioning standard, not this key schema."


In the Wild#

Twitter — Timelines in Redis#

Twitter has publicly described serving home timelines from large Redis clusters: each active user's timeline is a capped list of tweet IDs, populated by fan-out-on-write, with high-follower accounts merged at read time instead. They also modified Redis internally for memory efficiency on these lists. The design trades very large RAM spend for predictable millisecond reads on the hottest page in the product.

Staff insight: The Redis structure is the easy part; the hybrid fan-out rule (who gets pushed, who gets pulled) is the real design decision, and it's driven by the single-key and RAM ceilings of the cache tier.

Instagram — Hash Bucketing for Memory#

Instagram engineering published how they stored ~300M media-ID→user-ID mappings. Naive string keys needed ~21 GB; bucketing IDs into hashes of ~1,000 fields each let Redis use its compact encoding and cut memory to ~5 GB — roughly 4× savings for a change in key design, not infrastructure.

Staff insight: Encoding thresholds are a data-modeling lever. In an interview, "I'd bucket into hashes to stay under the listpack threshold" shows you understand Redis's memory model, not just its API.

Stack Overflow — Redis as a Shared L2 Cache#

Stack Overflow has publicly documented a small, heavily optimized infrastructure in which Redis serves as a shared L2 cache behind per-server in-memory L1 caches, with Pub/Sub used to broadcast cache invalidations. A small number of Redis servers handle the site's caching at low CPU.

Staff insight: The two-tier L1/L2 pattern is the hot-key answer: the in-process cache absorbs the heaviest reads, Redis absorbs the rest, and invalidation is a broadcast hint — acceptable because the data is rebuildable.


Practice Drill#

Prompt: "Your checkout service uses Redis for idempotency keys (SET idem:{key} {response} NX EX 86400). During a regional network event, Redis failed over and finance found 140 duplicate charges. The team proposes switching to appendfsync always. What do you do?"

Staff Answer

appendfsync always doesn't address the cause. The duplicates came from asynchronous replication: the old primary acknowledged the SET NX, the promoted replica never received it, and retried requests found no key and charged again. Local fsync protects against a single-node crash, not replication loss. Immediate: refund and reconcile the 140 charges against the payment processor's idempotency (Stripe-style providers accept an idempotency key — pass ours through so the provider dedups even if we don't). Short term: move idempotency to a durable store — a Postgres table with a unique constraint on the key (~2–5 ms, acceptable on a checkout path of ~300 ms), or MemoryDB if we need Redis latency with durable acknowledged writes. If Redis must remain as a fast pre-check, keep it but make the database constraint the authority. Guardrails: min-replicas-to-write 1 on any remaining coordination clusters, a game day that kills the primary under load and counts duplicates, and a reconciliation job that alerts within 15 minutes on duplicate charge IDs. Owner: the payments team owns correctness; the platform team owns the contract label that should have flagged this cluster as coordination, not cache.

Why this is L6:

  • Correctly identifies async replication — not fsync — as the loss mechanism.
  • Moves the correctness guarantee to a component whose contract supports it, with a latency number.
  • Adds defense in depth at the provider boundary and a detection loop with a time bound.

What L7 adds:

  • Audits the fleet for every other correctness-critical use of async Redis — this is a class of bug, not an incident.
  • Makes the contract label mandatory at provisioning so coordination clusters get durable defaults automatically.
  • Prices it for the business: 140 duplicates × average order value + support cost vs ~$2K/month for a durable store.

Quick Reference Card#

Execution:     single-threaded commands per shard; 100K–200K ops/s/core; ~1M pipelined
Latency:       ~0.1–0.3 ms server, 0.5–1 ms same-AZ round trip
Key overhead:  ~50–90 bytes per top-level key; bucket small values into hashes
Encodings:     listpack under 128 entries / 64-byte values (hash, zset) — 5–10× smaller
Persistence:   AOF everysec ≈ 1 s local loss; failover loss = replication lag
Fork cost:     ~10–20 ms per GB RSS; CoW can need up to 2× RAM — maxmemory ≤ 60% host
Cluster:       16,384 slots, CRC16; hash tags {x} for multi-key; ~25 GB per shard
Failover:      5–30 s write unavailability (Sentinel / cluster-node-timeout 15 s)
Backlog:       repl-backlog-size = 60–120 s of peak write bytes (not 1 MB)
HyperLogLog:   12 KB, 0.81% error
Pub/Sub:       fire-and-forget; Streams for at-least-once; Kafka for replay
Red flags:     KEYS *, DEL on big keys, no TTL, shared company cluster,
               Redis as sole ledger, Redlock for correctness, alerts on mean shard
Defaults:      label the contract → pick policy; noeviction for coordination;
               allkeys-lfu for cache; near-cache for keys > 50K reads/s
  1. Loading the index…