Hiring BarSupport

Redis vs Memcached

Comparison16 min read3 diagrams

Both are in-memory key-value stores with sub-millisecond reads, both sit in front of a database, and both appear in every "add a cache" answer. They are built on opposite philosophies. Memcached is a deliberately dumb, multi-threaded, volatile cache: strings in, strings out, evict whatever is least recently used, no persistence, no replication, no opinions. Redis (and its open-source fork Valkey) is an in-memory data structure server: sorted sets, hashes, streams, atomic scripts, replication, optional persistence — a cache that can also be a leaderboard, a rate limiter, a lock service and a queue. The question that decides it: is this purely a cache of values you can recompute from somewhere else, or does the data structure or the data itself matter? If it is purely a cache and the fleet is large, Memcached's simplicity is a feature. If you need anything beyond get/set — or you will within a year — Redis.

Go deeper:

  • The miss path in depth (leases, gutter pools, routing tiers and the arithmetic of a miss storm when a node dies) is in Memcached.

The Verdict#

Default to Redis (or Valkey) — it does everything Memcached does well enough and many things Memcached cannot. Pick Memcached for very large, pure look-aside caches where multi-threaded throughput per node and operational simplicity matter more than features.

Pick Redis / Valkey whenPick Memcached when
You need data structures: sorted sets (leaderboards), counters, hashes, sets, streams, geoEvery operation is get/set/delete of opaque blobs (rendered fragments, serialized objects)
You need atomic multi-step operations (Lua/Functions, MULTI) for rate limits, locks, inventoryYou want each node to use all its cores for one cache with no sharding inside the node
Losing the cache on restart would hurt (warm-up takes too long, or it holds session data)Losing a node is fine: the database can absorb the misses, or the fleet is big enough that one node is a few percent
You need replication and automatic failover from the managed serviceYou already shard on the client (consistent hashing) and want the server to stay simple
You need pub/sub, keyspace notifications or TTL-based expiry with precise semanticsYou are running a cache fleet of hundreds of nodes where per-node efficiency is money

🎯 Staff Move: "For a pure look-aside cache either works, so I'll pick Redis — or Valkey on the managed service — because the next three asks from product will be a rate limiter, a leaderboard and a session store, and I'd rather run one caching technology than two. If we were at hundreds of cache nodes doing nothing but get and set, I'd revisit Memcached for cost per GB and cores per node."

Diagram: The Verdict

At a Glance#

DimensionRedis / ValkeyMemcached
Data modelStrings, hashes, lists, sets, sorted sets, streams, bitmaps, HyperLogLog, geo; modules (JSON, search) in Redis 8Opaque byte strings with flags and TTL
Atomic operationsEvery command is atomic; INCR, ZADD, Lua scripts / Functions, MULTI/EXECincr/decr, cas (check-and-set with a version token)
ConsistencyPrimary–replica, asynchronous replication; replicas can serve stale reads; WAIT for acknowledged replicasNone across nodes; each key lives on exactly one node chosen by the client
PersistenceOptional: RDB snapshots, AOF log (fsync every second is the usual setting)None (data lost on restart); extstore can spill values to flash, still a cache
ThreadingCommand execution on one main thread per shard; I/O threads for network (Redis 6+, much improved in Valkey 8)Fully multi-threaded; scales with cores on one node
Throughput (order of magnitude)~100K–300K simple ops/s per shard (one core); ~1M+ with pipelining; scale out with Cluster~1M+ simple ops/s per large multi-core node
LatencySub-ms p50; p99 sensitive to slow commands blocking the main threadSub-ms p50 and p99 for get/set
Scaling modelRedis Cluster: 16,384 hash slots across primaries, client-aware routing, online reshardingClient-side consistent hashing across independent nodes
EvictionConfigurable maxmemory-policy (allkeys-lru, allkeys-lfu, volatile-ttl, noeviction, …)LRU per slab class (segmented LRU in modern versions)
Max value size512 MB per string (keep values ≤ ~100 KB in practice)1 MB default item size (configurable)
Operational burdenMedium: persistence, replication, failover, memory fragmentation, big keysLow: stateless processes; the hard part is the client-side routing and warm-up
Managed optionsElastiCache and MemoryDB (Valkey, Redis OSS), Redis Cloud, Google Memorystore, Azure, othersElastiCache for Memcached, Google Memorystore for Memcached
Cost shapeMemory-hours; replicas double it; Valkey on ElastiCache priced below Redis OSS (as of 2026)Memory-hours; no replicas to pay for
LicenseRedis 8: AGPLv3 / RSALv2 / SSPLv1 options; Valkey: BSDBSD

Redis or Valkey? For design purposes they are the same engine family; the choice is licensing and vendor, not architecture.

RedisValkey
LicenceAGPLv3 option added in Redis 8 (May 2025), alongside RSALv2 and SSPLv1BSD, Linux Foundation project forked from Redis 7.2.4
ExtrasRedis 8 folds JSON, time series, probabilistic types, the query engine and vector sets into coreFocus on core performance, I/O threading, memory efficiency
Where you'll meet itRedis Cloud, self-hostedElastiCache and MemoryDB, where it is the lowest-priced engine (as of 2026)

How They Actually Differ#

Data Structures: The Cache That Became a Database#

Memcached stores bytes. To add an item to a cached list you read the whole list, deserialize, append, serialize and write it back — with cas to avoid losing a concurrent update. Redis does it server-side in one atomic command.

Use caseRedisMemcached
Leaderboard top 100ZADD scores 4210 user:42, ZREVRANGE scores 0 99 WITHSCORES — O(log N)Recompute in the app or the database; cache the rendered result
Rate limiterINCR + EXPIRE or a Lua token bucket — atomic per requestincr with a TTL'd key works for fixed windows; anything smarter needs a round trip per step
Session with 10 fieldsHSET session:abc field value — update one fieldSerialize the whole session on every change
Recent activity feedLPUSH + LTRIM to keep the last 500Read-modify-write the whole blob with cas
Distributed lockSET key token NX PX 30000 + compare-and-delete scriptadd (only if absent) with TTL — workable, but no fencing story

The same "append to a user's recent-items list, keep 50" in each:

Memcached (client does the work, retries on conflict):
  loop up to 5 times:
      value, cas_token = gets("recent:42")
      items = deserialize(value) or []
      items = [new_item] + items[:49]
      if cas("recent:42", serialize(items), cas_token, ttl=3600) == STORED: break
      # NOT_STORED/EXISTS → someone else wrote first; retry
  # under contention on a hot key, retries pile up and some updates are dropped

Redis (server does the work, atomically, one round trip):
  MULTI
    LPUSH recent:42 new_item
    LTRIM recent:42 0 49
    EXPIRE recent:42 3600
  EXEC

Who pays: with Memcached the application team pays for every data-structure feature in serialization code and CAS retry loops. With Redis the on-call pays when someone runs ZRANGE over 10 million members or KEYS * on the main thread and blocks every other client.

Threading: One Fast Core vs All the Cores#

Redis executes commands on a single main thread per process. That is why every command is atomic without locks, and why one slow command (a large DEL, an unbounded LRANGE, a heavy Lua script) stalls every client on that shard. Network I/O can be spread over threads, and Valkey 8 made that substantially more effective, but data access stays single-threaded. To use 16 cores you run many shards.

Memcached is multi-threaded across a shared hash table with fine-grained locking. One process on a 32-core box uses all 32 cores for one cache. Fewer, bigger nodes; fewer things to route to.

Diagram: Threading: One Fast Core vs All the Cores

Who pays: a Redis fleet has more processes to monitor and fail over; a Memcached fleet has fewer, larger blast radii per node.

Durability and Replication#

Memcached has none, on purpose. A node restart is a cold cache for its slice of keys; a node loss sends that slice of traffic to the database. Large operators handle this with redundant pools or regional replicas built in the client or a proxy layer.

Redis offers asynchronous replication with automatic failover (Sentinel, Cluster or the managed service), plus optional persistence. Replication is async, so a failover can lose the last few hundred milliseconds of writes — fine for a cache, worth stating for a session store or a rate limiter.

🎯 Staff Insight: "Redis has persistence" does not make it a system of record. AOF with fsync every second can lose a second of writes, async replication can lose more on failover, and recovery time grows with dataset size. If the data must not be lost, it belongs in a database with Redis in front.

Memory Efficiency and Eviction#

Both pay per-key overhead (tens of bytes) on top of the value. Memcached's slab allocator groups items by size class; if the item size distribution shifts, memory gets stuck in the wrong slab class until it is rebalanced (modern versions do this automatically). Redis allocates per object with jemalloc; fragmentation after heavy churn can leave resident memory well above the dataset size (mem_fragmentation_ratio > 1.5 is a warning sign).

Redis's compact encodings for small hashes, sets and sorted sets can store many small values far more densely than one key per value — a well-known trick for large maps of small IDs. Memcached has no equivalent; one key, one item.

Eviction differs in control: Redis lets you choose the policy — and noeviction turns it into an error-returning store when full, which is the right setting for data you must not silently drop and the wrong one for a cache. Memcached always evicts by LRU within the slab class.

Scaling Out: Who Knows Where a Key Lives#

Memcached servers do not know about each other. The client hashes the key onto a ring of servers (consistent hashing, so adding a node remaps ~1/N of keys). Proxies (for example mcrouter at Facebook scale) add pools, replication and failover outside the servers.

Redis Cluster puts the routing in the protocol: 16,384 hash slots assigned to primaries; clients cache the slot map and follow MOVED redirects; resharding moves slots online. Multi-key commands work only when all keys share a slot ({user:42}:cart, {user:42}:profile via hash tags).


Where Each One Breaks#

Redis / Valkey#

FailureWhat happensDetectionOwner
Slow command blocks the shardKEYS *, a 5M-member SMEMBERS, or a heavy script stalls every client for 100s of msSLOWLOG, latency spikes on one shardApp team
Big or hot keyOne key gets 50% of traffic or holds 500 MB; one shard saturates while others idlePer-shard CPU skew, --hotkeys / --bigkeys scansApp team
Fork for persistenceRDB/AOF rewrite forks the process; copy-on-write doubles memory under heavy writes; latency spikes or OOMlatest_fork_usec, memory near limitPlatform
Failover loses writesAsync replication; acknowledged writes on the old primary vanishReplication offset gap at failoverPlatform + app (if using it as a store)
noeviction at maxmemoryWrites return errors; app treats cache errors as fatalused_memory vs maxmemory, OOM errorsApp team
Licence / fork confusionTeams pin different engines and versions; managed service moves defaultsEngine inventoryPlatform

Memcached#

FailureWhat happensDetectionOwner
Node loss → database stampede1 of 10 nodes dies; 10% of keys miss at once; database load jumpsMiss rate per node, DB QPSPlatform + app
Cold restart after deployRolling restart empties each node; hit rate drops for minutes to hoursHit ratio by node over timePlatform
Slab calcificationValue size distribution shifts; memory stuck in old slab classes; evictions while memory looks freeEvictions per slab classPlatform
Client hashing mismatchTwo services use different hashing or server lists; same key on two nodes; stale readsInconsistent reads across servicesApp teams
Value too largeItems above the item size limit silently fail to store; every read missesSet failures, miss rate on large objectsApp team

One Cache Node Dies at Peak#

Ten nodes, 500K reads/s, 95% hit rate, database sized for 40K reads/s.

                 Memcached (10 nodes, no replicas)              Redis (10 primaries + replicas)
t=0              node 7 dies                                    primary 7 dies
t=+1s            10% of keys now miss → +50K DB reads/s         clients get errors / timeouts for slot range
                 DB at 75K reads/s (sized for 40K)              of shard 7 (10% of keys)
t=+5s            DB p99 10ms → 400ms; app threads pile up       replica promoted (managed failover:
                 on slow queries                                ~seconds to tens of seconds)
t=+30s           client ring removes node 7 → its keys          shard 7 serving again from replica's data,
                 rehash to 9 nodes, all cold → still missing    hit rate back to ~95%
t=+10min         hit rate recovers as keys refill               replacement replica syncing in background
Mitigation       request coalescing on miss, stale-while-       client retry with short timeout,
                 revalidate, DB headroom, a second pool         replicas in another AZ

Who pays: with Memcached the database team absorbs the incident; with Redis the cache budget paid for it in advance with replicas. Neither is wrong — but the choice must be explicit, and the database owner must have signed off on the Memcached version.

The production surprise for both: the cache becomes load-bearing. Once the database is sized for a 95% hit rate, a cold cache is an outage, not a slowdown. Both need a warm-up plan and request coalescing on miss.


Cost and Operations#

Redis / ValkeyMemcached
Who runs itPlatform team or managed service; app teams own key design and command hygienePlatform team owns the fleet and the client library / proxy
Bill scales withGB of memory × hours, × (1 + replicas)GB of memory × hours
Hidden costsReplicas (2× memory for HA), headroom for fork copy-on-write (often 25–50%), cross-AZ replication trafficDatabase capacity to absorb a node's worth of misses
Managed pricing noteOn ElastiCache, Valkey is priced 20% below other node-based engines and 33% below Redis OSS for serverless (as of 2026)Same node prices as other non-Valkey engines
People0.25–1 engineer for a multi-team fleet; more if used as a primary store0.25–0.5 engineer; more at hundreds of nodes
What on-call watchesRedis / ValkeyMemcached
The SLO metricHit ratio, p99 latency per shardHit ratio, p99 latency per node
The "act now" alertMemory > 90% of maxmemory, replication broken, slowlog spikesNode down, eviction rate jump, DB QPS spike from misses
Routine workEngine upgrades, resharding, big-key cleanup, persistence tuningRolling restarts with warm-up, slab rebalancing, client config rollouts

Sizing example. 200 GB of cached objects, 500K reads/s at peak:

Memcached:  4 × 64 GB nodes (with headroom), no replicas      ≈ 256 GB of memory paid for
            each node ~125K ops/s — well within one multi-core node
            lose a node → 25% of keys miss → DB must handle ~125K extra reads/s briefly

Redis:      8 shards × 32 GB + 1 replica each                 ≈ 512 GB of memory paid for
            ~62K ops/s per shard — comfortable on one core
            lose a primary → replica promoted in seconds, no miss storm

The Memcached fleet is roughly half the memory bill. The Redis fleet buys failover without a miss storm. Who pays is the choice: the cache budget, or the database during an incident.

🧭 Principal Insight: "Memcached is cheaper per cached gigabyte; Redis is cheaper per feature. At most companies the feature count wins because every team that needs a counter or a lock would otherwise build one. At a handful of very large companies the gigabyte wins, and they run both."


Switching Later#

MoveHowWhat's hard
Memcached → RedisPoint the client at Redis for get/set; warm it by double-writing for one TTL period; cut reads overEasy for pure caching. Client-side hashing logic is replaced by Cluster or a proxy
Redis → MemcachedOnly for the pure-cache keys; move data-structure usage elsewhere firstAnything using sorted sets, scripts, pub/sub or persistence has no equivalent
Redis → ValkeyWire-compatible for the shared command set; managed services offer in-place engine upgradesFeatures added only in Redis after the fork (some Redis 8 modules) won't exist; test client libraries
Redis as a store → a databaseDual-write, backfill, move reads, keep Redis as a cacheCode that relied on atomic data-structure ops needs a new concurrency story

One-way doors: using Redis as the system of record (every feature built on it assumes its semantics); hash-tag schemes that pin related keys to one slot. Two-way doors: engine choice for a pure cache, eviction policy, node size, number of shards.

Diagram: Switching Later

How Real Companies Chose#

Facebook: Memcached as a Planet-Scale Cache#

Facebook's NSDI 2013 paper describes turning memcached into a distributed cache tier that handles billions of requests per second and holds trillions of items for over a billion users, accepting imprecise semantics — eviction, data loss, no transactions — in exchange for performance and simplicity, and putting the intelligence (consistency with the database, failover, regional replication) around the servers rather than in them (USENIX NSDI '13).

Staff insight: At that scale the server's simplicity is the point. Every feature in the cache process is multiplied by thousands of machines; Facebook moved the hard parts into clients and proxies they controlled.

Netflix: EVCache, Built on Memcached#

Netflix's EVCache is an open-source caching solution built on memcached and the spymemcached client, designed for AWS; the name stands for Ephemeral, Volatile Cache — data persists only for its TTL and can be evicted at any time (Netflix EVCache on GitHub).

Staff insight: Netflix chose a volatile engine and designed the system around that volatility — the name is a contract with every client team: don't store anything here you can't lose.

GitHub: Moving Persistent Data Out of Redis#

GitHub described disabling Redis persistence and moving data that had been stored durably in Redis — configuration, counters, rate-limit data, spam signals and activity feeds — into a MySQL-backed key-value store, citing operational cost and the team's deeper MySQL expertise. Activity feeds alone accounted for about 67% of over 350 million daily Redis writes before the migration cut write volume (GitHub Blog).

Staff insight: Redis's features invite teams to use it as a database. The bill arrives later as persistence operations nobody planned for. Decide early which Redis clusters are caches and which are stores, and staff them differently.


Follow-Ups to Expect#

After You Say...They Will Ask...What They're Testing
"Redis as the cache""A cache node dies. What does the database see?"Replicas vs miss storm; request coalescing; warm-up
"Memcached for simplicity""Now product wants a real-time leaderboard."Whether you planned for data-structure needs or a second system
"Redis Cluster""How do you do a multi-key operation across users?"Hash slots, hash tags, cross-slot limits
"Redis for sessions""What if a failover loses the last second of writes?"Async replication; whether sessions can be re-created
"Persistence on""What happens to latency during a snapshot?"Fork, copy-on-write, memory headroom
"LRU eviction""One key gets 40% of reads."Hot keys: local in-process cache, key replication, read replicas
"Valkey instead of Redis""Why does that matter?"Licensing, managed pricing, compatibility — and knowing it's not a design decision for most systems

What to Say in the Interview#

"For a pure look-aside cache both work. I'll use Redis — or Valkey on the managed service — because the same cluster type will serve the rate limiter and the leaderboard, and running one technology is cheaper than running two."

"This cluster is a cache, not a store: allkeys-LRU, no persistence, and the database is sized to survive losing one shard's worth of hits for a few minutes."

"The rate-limit counters are different — losing them on failover means a burst gets through, which product has signed off on — so they get replicas but still no persistence."

"If we were running hundreds of nodes of nothing but get and set, I'd look at Memcached seriously: multi-threaded nodes and no replicas roughly halve the memory bill."


  1. Loading the index…