Hiring BarSupport

Design a Distributed Lock Service — Staff-Level Case Study

Case study82 min read8 diagrams

Technologies referenced in this case study: ZooKeeper & etcd · Redis · PostgreSQL · DynamoDB

Related: Consensus / Coordination Service covers Raft internals · Distributed Coordination covers the broad pattern · Contention covers lock-free alternatives. This case study stays on the lock service as a product: its API, its guarantees, and who gets paged when it lies.

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once, then return to individual sections for targeted review.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines → Active Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Fault Lines → Failure Modes → weak-spot Deep Dives
Deep Dive3+ hrsEverything, including the Principal Lens and appendices
What is a Distributed Lock Service? — Why interviewers pick this topic

A distributed lock lets one process among many claim exclusive (or shared) access to a named resource — a file, a shard, a job, an account — across machines that share no memory. Because machines crash, pause and lose network, every practical distributed lock is a lease: a lock that expires unless renewed. That single fact — the lock can be taken away from a holder who does not know it yet — is the entire interview.

Before vs After — nightly settlement job scenario:

Without a correct lock (Redis SETNX, 30s TTL, no fencing):
t=0:       Worker A acquires "settle:2026-10-01", starts writing payouts
t=+12s:    Worker A enters a 41-second stop-the-world GC pause
t=+30s:    Lease expires. Worker B acquires the same lock legitimately
t=+31s:    Worker B starts writing payouts for the same batch
t=+53s:    Worker A wakes up, still believes it holds the lock, keeps writing
t=+2min:   4,112 merchants paid twice. $1.8M in duplicate transfers.
t=+3 days: Finance reconciliation finds it. Nobody was paged.

With a lease + fencing token checked by the ledger:
t=0:       Worker A acquires lock, receives fencing token 7741
t=+12s:    Same 41-second GC pause
t=+30s:    Lease expires. Worker B acquires lock, receives token 7742
t=+31s:    Worker B writes with token 7742 — ledger records max_token = 7742
t=+53s:    Worker A wakes, writes with token 7741 — ledger REJECTS (7741 < 7742)
t=+53s:    Worker A logs lock_fencing_rejections_total, aborts. Zero duplicates.

Why interviewers reach for this question: Locking looks like a solved problem — "use Redis SETNX" or "use ZooKeeper". It is actually the cleanest probe for whether a candidate understands that distributed systems cannot guarantee mutual exclusion by themselves. The lock service can only promise "at most one holder according to me". Whether that becomes "at most one writer to the resource" depends on something outside the lock service — the fencing check — and on someone owning it. Interviewers want to see you find that gap without being led to it.

Mechanics Refresher: Lock Implementations
ImplementationHow It WorksProsCons
Redis SET key val NX PX ttlSingle atomic set-if-absent with expiry; release via Lua compare-and-delete~0.2–1 ms acquire; trivial to runAsync replication: failover can lose the lock; no native fencing token
RedlockAcquire on majority (3 of 5) independent Redis masters within a validity windowNo single Redis SPOFSafety depends on bounded clock drift and bounded pauses; still no fencing token
ZooKeeper ephemeral sequential znodesCreate /locks/x/lock-000042; lowest sequence holds; others watch predecessorLinearizable; session-bound; zxid/sequence usable as token2–10 ms per acquire (quorum write + fsync); ops burden of a ZK ensemble
etcd lease + transactionTxn(If create_revision==0) Then Put(key, lease); key dies with leaseLinearizable; revision is a free monotonic fencing token; Kubernetes-grade tooling~10–30K writes/s ceiling per cluster; 8 GB recommended data limit
DynamoDB conditional writePutItem with attribute_not_exists(pk) OR expires_at < :now plus a record version numberServerless, multi-AZ, pay per requestExpiry check uses client clock; ~5–10 ms; you build heartbeats yourself
PostgreSQL advisory lockpg_try_advisory_lock(key) held by session or transactionZero new infrastructure; dies with the connectionCoupled to one DB primary; lost on failover; connection-pool footguns
Chubby-style lock serviceCoarse-grained locks on a Paxos cell; sessions + KeepAlives; sequencers as tokensDesigned for exactly this; lock-delay protects un-fenced resourcesBuilt, not bought — Google-scale investment

For most production systems: a lease on a consensus store (etcd or ZooKeeper — or a managed equivalent) that returns a monotonic fencing token, with the token checked by the protected resource. Redis SET NX is fine — and the right answer — when the lock is only an efficiency optimization.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What This Interview Actually Tests#

A distributed lock is not a data-structure question. Everybody knows SETNX.

It is a safety-ownership question: when the lock lies — and it will, on the first long GC pause — who stops the second writer? It tests:

  • Whether you ask what a double-holder costs before choosing an implementation
  • Whether you know a lease can expire under a holder that is still running
  • Whether you push the safety check to the resource (fencing) and name who owns that check
  • Whether you can argue against the lock — conditional writes, single-writer partitioning, queues
  • Whether you treat the lock service as a shared product with an API contract, quotas and an on-call

The key insight: A lock service can only guarantee "at most one holder according to the lock service." Mutual exclusion at the resource needs a fencing token that the resource checks. If the resource cannot check, you do not have a correctness lock — you have a strong hint, and your design must be safe when the hint is wrong.

The L5 → L6 → L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Redis SET NX PX with a TTL, release with a Lua script"Asks "What breaks if two holders exist for 30 seconds?" and splits efficiency locks from correctness locksAsks how many lock implementations the org already runs, which ones guard money or data, and whether any of those locks can be deleted entirely
SafetyPicks a TTL "longer than the job" and adds renewalSays leases will expire under live holders (GC, VM pause, partition) and requires fencing tokens checked by the resourceMakes fencing a platform contract: storage and ledger teams expose token-checked writes as a paved-road API so product teams cannot forget it
StoreRedlock for HA "so Redis isn't a SPOF"Redis for efficiency, etcd/ZooKeeper for correctness; explains why Redlock helps neitherPicks one consensus-backed lock substrate for the org, prices it, and sets a deprecation date for the other three
Failure"Add replicas; renew the lease in a background thread"Defines lock-service-down behavior per intent: efficiency locks fail-open, correctness locks fail-closed, with lock_acquire_errors_total pagingDesigns correlated-failure posture: the lock cluster is a tier-0 dependency, so it is cell-local, never cross-region on the hot path, and its outage is a rehearsed game-day scenario
OwnershipEach service wires its own lock clientLock service owned by platform; resource owners own fencing checks; one runbookRedraws the boundary: platform owns the substrate and SDK, storage teams own fencing enforcement, product teams must justify every new correctness lock in design review
AlternativesNot raisedProposes conditional writes / single-writer partitions before reaching for a lockWrites the standard that makes "lock-free first" the default and treats every new lock as technical debt with an owner
Why "first move" separates levels

L5: Reaches for an implementation. SET lock:job NX PX 30000 is correct Redis, and the Lua compare-and-delete release shows real knowledge. But it commits to a mechanism before knowing whether a double-holder costs nothing (a duplicate cache rebuild) or $1.8M (a duplicate payout run).

L6: Separates intents out loud. "Before I pick a store — if two workers hold this lock at once, does something get computed twice, or does something get corrupted? Those need different systems. Efficiency locks can live on a single Redis. Correctness locks need a consensus store and a fencing token the resource checks."

L7: Asks whether the lock should exist. "Most locks I've audited guard something a conditional write or a single-writer partition would make safe by construction. Before we design a lock service, which of our locks guard invariants, and which are just deduplication we could do with idempotency keys?"

Why "safety" separates levels

L5: Believes a long enough TTL plus background renewal makes the lock safe. It reduces the probability of overlap; it does not bound it. A 41-second GC pause, a VM live-migration stall, or a network partition that blocks renewals will expire a lease under a holder that is still about to write.

L6: States the impossibility: "No lease protocol can stop a paused process from waking up and acting on stale belief. The only defense is at the resource: every write carries the fencing token, and the resource rejects tokens lower than the highest it has seen." Then names who owns that check — the storage team, not the lock team.

L7: Turns the check into infrastructure. If 40 teams each implement if token < max_token: reject, 10 will get it wrong. The ledger, the blob store and the job-state table expose fenced writes as a first-class API, and the design review checklist asks "which resource checks your token?" for every correctness lock.

Why "alternatives" separates levels

L5: Treats "we need mutual exclusion" as the requirement. Designs the lock.

L6: Treats mutual exclusion as one implementation of the real requirement — usually "this invariant must hold" or "this work must happen once". Offers cheaper constructions: UPDATE ... WHERE version = :v, partitioning so each key has exactly one writer, idempotency keys, or a queue with a single consumer per partition.

L7: Tracks locks as an org-level smell. Every correctness lock is a cross-team runtime dependency on a tier-0 service; the standard requires a written justification for why a lock-free construction does not work.

The Staff Positions#

PositionRationale
Leases, never indefinite locksA crashed holder must not block the world forever; every lock has a TTL (typ. 10–30 s) and a renewal protocol
Fencing token on every correctness lockLeases expire under live holders; only the resource can reject the stale writer
Consensus store for correctness, Redis for efficiencyetcd/ZooKeeper give linearizable acquire and a monotonic revision; Redis async replication can lose a lock on failover
Redlock is not the answer to either intentOverkill for efficiency, unsafe for correctness without fencing — and with fencing you do not need it
Coarse-grained locks only on the lock serviceLock services are built for ~100s of acquires/s per hot key and minutes-to-hours holds, not per-row locking at 50K/s
Fail-closed for correctness, fail-open for efficiencyLosing the lock service must stop money-moving work but must not stop the cache warmer
Lock-free firstConditional writes and single-writer partitions are safe by construction and remove a tier-0 dependency

The Three Intents#

Three intents drive every design decision. Each leads to a different system.

IntentConstraintStrategyFailure ModeCorrectness Bar
Efficiency (dedupe work)Cheap and fast; occasional double execution is fineSingle Redis SET NX PX, short TTL, no fencingTwo workers do the same work; wasted CPUOverlap acceptable; cost of a double run < cost of a consensus store
Correctness (protect an invariant)Two writers corrupt data or move money twiceConsensus-backed lease + monotonic fencing token checked by the resourceStale holder writes after lease expiryZero accepted stale writes; overlap in belief allowed, overlap in effect forbidden
Ownership / leader election (coarse, long-lived)One scheduler, one shard owner, one primary for minutes to daysSession-based lease on etcd/ZK, epoch number as fencing token, watch for handoffSplit brain during handoff; failover takes TTL + election secondsOne effective leader per epoch; downstream rejects old epochs

🎯 Staff Move: "I'll assume a correctness lock — two holders would corrupt data — because that's where fencing, lease expiry and the Redlock debate actually matter. If it turns out we only need efficiency, I'll downgrade to a single Redis and save us a consensus cluster. And before either, I want to check whether a conditional write removes the lock entirely."

The Five Fault Lines#

#Fault LineThe Tension
1Efficiency vs CorrectnessIs a double-holder a wasted CPU-minute or a corrupted ledger? The answer picks the store, the failure posture and the owner.
2Short vs Long Leases (Liveness vs Safety)Short TTL = fast takeover after a crash but more false expiries under pauses; long TTL = fewer false expiries but minutes of stall when a holder dies.
3Consensus Store vs Fast Storeetcd/ZooKeeper (linearizable, 2–10 ms, ops-heavy) vs Redis (sub-ms, cheap, loses locks on failover).
4Lock vs Lock-Free ConstructionA lock serializes and adds a dependency; conditional writes / single-writer partitions are safe by construction but reshape the data model.
5Shared Lock Platform vs Embedded LocksOne governed service with fencing as a contract vs every team wiring its own Redis — velocity now vs correlated failures later.

In the Wild: Real Production Systems#

Why this section belongs here: Citing real systems shows you know these designs were paid for with incidents.

Google Chubby — Coarse-Grained Locks with Sequencers#

Google's Chubby (Burrows, OSDI 2006) is the archetype: a Paxos-replicated cell of five replicas serving coarse-grained locks held for hours or days — leader election for GFS and Bigtable, not per-request locking. Clients hold sessions renewed via KeepAlives (default lease on the order of 12 seconds). Chubby exposes sequencers — opaque byte strings carrying the lock name, mode and generation number — that servers can verify before acting. For resources that cannot check sequencers, Chubby offers lock-delay: after a holder fails, the lock cannot be re-granted for up to a minute, giving stale holders time to die.

Staff insight: The paper explicitly designs for "the lock can be lost while the holder still thinks it has it" and ships two answers: fencing (sequencers) when the resource can check, and a time-based buffer (lock-delay) when it cannot. That is exactly the answer interviewers want, from the people who built it.

Kubernetes — Leader Election on Lease Objects#

Every Kubernetes controller manager and scheduler runs leader election over a Lease object in the coordination.k8s.io API, backed by etcd. The defaults — leaseDuration 15 s, renewDeadline 10 s, retryPeriod 2 s — encode a tradeoff: a dead leader is replaced within ~15–17 s, and a leader that cannot renew within 10 s steps down before its lease expires. The resourceVersion on the Lease acts as optimistic concurrency for takeover.

Staff insight: The leader voluntarily abdicating at renewDeadline < leaseDuration is the client-side half of safety: stop acting before others may legitimately start. It narrows the overlap window but does not close it — which is why controllers are written to be idempotent against the API server's own optimistic concurrency.

Redlock and the Kleppmann–antirez Debate#

In 2016 Martin Kleppmann published "How to do distributed locking", arguing that Redlock (Redis's multi-master lock algorithm) is unsafe for correctness because it depends on bounded network delay, bounded process pauses and bounded clock drift — and offers no fencing token. Salvatore Sanfilippo (antirez) replied in "Is Redlock safe?", arguing the timing assumptions are reasonable in practice and that checks after acquisition mitigate pauses. Kleppmann's conclusion has become the industry heuristic: for efficiency, one Redis is enough; for correctness, use a consensus system and fencing tokens.

Staff insight: You do not need to win the debate in an interview. You need to show it collapses once you separate intents: Redlock is more machinery than efficiency needs and less safety than correctness needs.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Redis SET NX PX""The holder pauses 40 seconds. What happens?"Do you know leases expire under live holders?
"We'll renew the lease in a background thread""Renewal thread is fine, main thread is stuck on I/O. Now what?"Liveness of the renewer ≠ liveness of the work
"Use Redlock for HA""Walk me through a clock jump on one of the five nodes."Timing assumptions; fencing
"Fencing token""Who checks it? What if the resource is an S3 bucket or a third-party API?"Ownership of the check; fallback when unfenceable
"etcd for correctness""etcd loses quorum. What do the 400 jobs that need locks do?"Failure posture per intent; blast radius
"We'll lock per order ID""Flash sale: one SKU, 20K req/s. What's your lock throughput?"Contention math: throughput ≤ 1 / hold time
"Lock service for everyone""Who owns it? How do you stop one team creating 10M locks?"Platform thinking, quotas, governance

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: The lock service is a thin product layer over a consensus store. Its job is linearizable acquire, lease management and a monotonic token. The safety guarantee is completed outside the lock service: the resource tier rejects stale tokens. lock_fencing_rejections_total is emitted by the resource, not the lock service — which is why the resource team must own it. The control plane registers every namespace with an owner and an intent, so an efficiency lock cannot silently become load-bearing.

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say ThisThe L7 Answer — Say This
Store"Redis, or Redlock for HA""Redis for efficiency locks; etcd/ZooKeeper for correctness — linearizable acquire and the revision is a free fencing token.""One consensus-backed substrate for the org with a paved-road SDK; the three ad-hoc Redis lock libraries get a deprecation date."
Safety"TTL longer than the job + renewal""Leases expire under live holders. The resource rejects writes with tokens lower than the max it has seen.""Fenced writes are an API the storage platform provides. Design review rejects correctness locks without a named token checker."
Lease length"30 seconds""TTL = max tolerable takeover delay; renew at TTL/3; holder self-aborts at 2/3 TTL without a successful renewal.""TTL bounds are policy per namespace, set by the business owner of the stall cost, not by the developer."
Failure"Replicas""Lock store down → correctness work stops (fail-closed), efficiency work proceeds (fail-open). Both paged on lock_acquire_errors_total.""Lock store is tier-0 and cell-local. Its outage is a quarterly game-day scenario with a pre-approved list of what stops."
Contention"Lock per row""Lock throughput ≤ 1 / hold time. A 20 ms hold caps one key at 50 acquires/s. Hot keys need partitioning or no lock.""If a lock is hot, the data model is wrong. Hot-lock reports go to the owning team's quarterly review as debt."
Avoidance—"Can a WHERE version = :v conditional write or a single-writer partition replace this lock?""Lock-free first is the written standard; every new correctness lock needs a justification and an owner."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Redis SET NX PX acquire, same AZ~0.2–1 msWhy Redis is tempting; fine for efficiency
etcd / ZooKeeper acquire, same region~2–10 msOne quorum write + fsync; dominated by disk latency
Cross-region consensus acquire~70–150 msWhy global locks on a hot path are a non-starter
etcd practical write ceiling~10–30K writes/s per clusterEvery acquire, renew and release is a write
etcd recommended data size≤ 8 GBMillions of fine-grained locks will not fit gracefully
Typical lease TTL10–30 sTakeover delay after a crash ≈ TTL
Renewal cadenceTTL / 3Survive two lost renewals before expiry
JVM stop-the-world pauses on large heapsseconds; tens of seconds pathologicallyLonger than many TTLs — the classic zombie holder
ZooKeeper session timeout bounds2× to 20× tickTime (4–40 s at 2 s tick)Session expiry = ephemeral lock deletion
Kubernetes leader election defaults15 s lease / 10 s renew deadline / 2 s retryA production-tested ratio to quote
Max lock throughput on one key1 / hold time (10 ms hold → ~100/s)The contention ceiling no hardware fixes
Chubby lock-delayup to 1 minTime buffer for resources that cannot check tokens

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on SETNX syntax, Lua release scripts and Redlock quorum math, then run out of time before the only question that matters: what happens when the lock is wrong? Compress the mechanism to 10 minutes. Spend the rest on lease expiry, fencing, failure posture and whether the lock should exist.


Phase 1: Requirements & Framing (2–3 minutes)#

State functional requirements in 30 seconds:

"Clients acquire a named lock with a lease, renew it, release it, and optionally wait for it. Exclusive mode first; shared/read mode if needed."

Then spend the time on intent and the cost of failure:

"The first question is what a double-holder costs. If two workers rebuild the same cache, we waste a minute of CPU — that's an efficiency lock and a single Redis is fine. If two workers both run the payout batch, we move money twice — that's a correctness lock and it needs a different system. I'll assume correctness, because that's where the hard problems are."

Commit to non-functional constraints:

"For correctness: zero accepted stale writes at the resource, acquire latency under 10 ms in-region, takeover after a crashed holder within 30 s, and the service keeps working through a single node or AZ loss. I'm explicitly not promising that two clients never believe they hold the lock — no lease system can promise that. I'm promising that the resource never accepts work from the stale one."

🎯 Staff Move: The sentence "I'm not promising two clients never believe they hold the lock" is the one interviewers remember. It shows you know the impossibility result before they spring it on you, and it moves the guarantee to where it can actually be enforced.


Phase 2: Core Entities & API (1–2 minutes)#

Entities in 30 seconds:

EntityFieldsNote
Namespacename, owner_team, intent, ttl_min/max, quotaRegistered in the control plane; no anonymous locks
Sessionsession_id, client_identity, lease_id, ttlOne lease per client process; covers all its locks
Lockns/name, holder_session, mode, token, acquired_attoken = store revision at acquire, strictly increasing
Waiterns/name, session_id, seqFIFO queue, each waiter watches only its predecessor

API:

Acquire(ns, name, mode=EXCLUSIVE, wait_timeout=0s, request_id)
    -> { token: uint64, session_id, lease_expires_at }   | LOCK_HELD | TIMEOUT
Renew(session_id)                      -> { lease_expires_at }  | SESSION_EXPIRED
Release(ns, name, token)               -> OK | NOT_HOLDER        (idempotent)
Describe(ns, name)                     -> { holder, token, held_for, waiters }
Watch(ns, name)                        -> stream of { ACQUIRED | RELEASED | EXPIRED, token }

"Three contract details matter. Acquire is idempotent on request_id, so a retried acquire after a timeout doesn't double-queue. Release requires the token, so a stale holder can't release the new holder's lock. And there is no ForceUnlock in the data-plane API — breaking a lock is an audited admin operation, because it's exactly the action that creates two holders."


Phase 3: High-Level Architecture (≤5 minutes)#

Staff candidates spend under 5 minutes here. The boxes are not the interview.

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

"Stateless frontends in each AZ, a 5-node etcd cluster spread over 3 AZs — it tolerates 2 node failures or one AZ. Acquire is a single etcd transaction: if the key's create_revision is 0, put it with the session's lease attached. The revision of that write is the fencing token, monotonic for free. Lease expiry deletes the key; waiters watch their predecessor, so one release wakes one waiter, not a thousand. The protected resource checks the token. That's the whole hot path."

"Why frontends instead of clients talking to etcd directly? Three reasons: quotas per namespace, so one team can't create 10M keys and push etcd past its 8 GB comfort zone; session multiplexing, so 5,000 workers share fewer leases; and an API we can keep stable while we change the store underneath."


Phase 4: Transition to Depth (1 minute)#

"The happy path is simple. What makes this hard is that the lock can be revoked from a holder that doesn't know it. Three places worth going deep: lease expiry and fencing — how we stay safe when the holder pauses; failure posture — what happens when etcd loses quorum; and contention — whether some of these locks should exist at all. I'd start with fencing, since it determines whether the rest is even correct."

🎯 Staff Move: Name the impossibility first, then offer the menu. You are steering toward the topic where you are strongest and that the interviewer is guaranteed to care about.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → pick a position → quantify the cost → name who absorbs it.

Deep Dive 1: Lease expiry and fencing (7–8 min)

Walk the failure explicitly:

  1. Client A acquires, token 33. Starts a write that will take 200 ms.
  2. A's JVM enters a 25 s stop-the-world pause. Renewals stop.
  3. At TTL (15 s), etcd expires the lease and deletes the key.
  4. Client B acquires, token 34, writes to the resource. Resource records max_token = 34.
  5. A resumes. From A's point of view no time passed. A sends its write with token 33.
  6. Resource rejects: 33 < 34. A receives STALE_TOKEN, logs, aborts.

"Step 5 is unfixable from the client side. A checked 'do I still hold the lock' right before step 2 and the answer was yes. Any check-then-act on the client has a window. The resource is the only party that sees both writes in order, so it is the only party that can enforce the order."

What the resource must do:

-- Fenced write on a relational resource
UPDATE settlement_batches
   SET status = 'PAID', paid_by = :worker, fence = :token
 WHERE batch_id = :id
   AND fence <= :token;          -- 0 rows updated => stale holder, abort

"For an object store: a conditional put keyed on a metadata version, or write to a path that includes the token and promote atomically. For a third-party payment API that can't check our token, we fall back to an idempotency key derived from the batch ID — the provider deduplicates, so a stale holder's call is a no-op. If neither exists, I'd say plainly that the lock is a hint and we need reconciliation downstream."

🎯 Staff Move: Name who owns the check. "The ledger team owns fence <= :token. If they don't sign up for that, this isn't a correctness lock and I'll write that in the design doc."


Deep Dive 2: Lease length and renewal (5 min)

"TTL is the time we're willing to stall if a holder dies. Renewal is TTL/3, so we survive two missed renewals. And the client stops doing work when it hasn't renewed within 2/3 of TTL — Kubernetes uses 15 s lease, 10 s renew deadline for exactly this reason."

TTLTakeover after crashFalse expiries under 10 s pausesUse
5 s~5–7 sFrequent on JVM servicesOnly with fencing and fast work units
15 s~15–17 sRareDefault for job-level locks
60 s~60 sVery rareCoarse leader election where a 1-minute stall is acceptable

"With fencing, false expiries are a liveness cost — some work is wasted and retried. Without fencing, they're a safety violation. That's why I'd never pick a TTL to achieve safety; I pick it to bound stall time and use fencing for safety."

The subtle bug to volunteer: "A background renewal thread proves the process is alive, not that the work is progressing. If the worker thread is stuck on a 10-minute I/O, the renewer happily keeps the lock forever. So the SDK renews only if the worker has heartbeated progress within the last TTL — and every lock has a maximum hold time, say 15 minutes, after which the service refuses to renew."


Deep Dive 3: Store choice — and Redlock (5 min)

"For correctness I need linearizable acquire and a monotonic token. etcd gives both: a transaction on a Raft log and a cluster-wide revision. ZooKeeper gives both via ephemeral sequential nodes and zxid. Single Redis doesn't: replication is async, so if the primary acks my SET and dies before replicating, the replica is promoted without my lock and someone else acquires it. That's a correctness bug with no client-visible signal."

"Redlock fixes the single-node failover problem by acquiring on 3 of 5 independent masters, but its safety still assumes bounded clock drift and bounded pauses, and it has no token. If I add a fencing token to fix the pause problem, I need a monotonic counter — which is a consensus problem — and then I don't need Redlock. So Redlock is more machinery than efficiency needs and less safety than correctness needs."

"For efficiency locks, single Redis with SET NX PX and a compare-and-delete release is the right answer. Cheap, sub-millisecond, and losing a lock on failover just means some duplicate work."


Deep Dive 4: Failure posture when the lock service is down (4–5 min)

"etcd loses quorum — say two nodes in one AZ plus a bad deploy. Acquires and renewals fail. Existing leases can't be renewed, so after TTL every holder must assume it lost the lock. Per intent:"

  • Correctness locks: fail-closed. Payout batch stops. That's a page, but a stopped batch is recoverable; a double payout is not.
  • Efficiency locks: fail-open. The cache warmer runs without a lock and some work is duplicated. No page, just a counter.
  • Leader election: current leader steps down at renewDeadline; nobody leads until quorum returns. Every controller that depends on a leader stops reconciling — so the lock cluster is tier-0 and its SLO is tighter than any consumer's.

"The SDK makes the posture explicit: Acquire(..., on_unavailable=FAIL_CLOSED) is required for correctness namespaces and validated against the namespace registry, so nobody accidentally fails a payout lock open."


Deep Dive 5: Contention and lock-free alternatives (4–5 min)

"If a lock is held for 20 ms, one key supports at most 50 acquires per second — and queueing theory says latency explodes well before that, around 70–80% utilization. A flash sale with 20K req/s on one SKU through a per-SKU lock is mathematically impossible. The fix is not a faster lock; it's removing it: an atomic decrement with a condition stock > 0, or partitioning inventory into 32 sub-buckets with a single writer each."

"My general test: if the critical section is a read-modify-write on one record, a conditional write replaces the lock. If it's work spanning many records over seconds, a lock — or better, a single-writer partition owned via leader election — is justified."


Phase 6: Wrap-Up (2–3 minutes)#

"To summarize: the lock service guarantees at most one holder according to itself, using leases on etcd with the revision as a fencing token. Safety at the resource comes from the fencing check, owned by the resource team. Correctness locks fail closed, efficiency locks fail open, and that posture is declared per namespace. What I'd build next: a namespace registry with quotas, a max-hold-time policy, and a quarterly audit of which locks could become conditional writes."

The organizational closer:

"The hard part long-term isn't the lock service — it's that every team will want to use it for things it wasn't built for. Per-row locking at 50K/s, cross-region locks, locks with no fencing guarding money. The platform team needs a design-review gate for correctness locks and the authority to say no."

🎯 Staff Move: End on the governance gap. Senior candidates end with "and we could use Redlock for more availability."


Common Timing Mistakes#

MistakeL5 Does ThisL6 Does This Instead
Mechanism marathon10 min on Redlock quorum math and clock-drift formula"Redlock is the wrong answer for both intents — here's why in 60 seconds"
No intentDesigns one lock for cache warming and payoutsSplits efficiency vs correctness in the first 2 minutes
Fencing only when promptedWaits for "what about GC pauses?"Volunteers the paused-holder timeline in Phase 4
No resource owner"The service checks the token""The ledger team owns fence <= :token; here's the SQL"
No failure posture"etcd is highly available""Quorum loss stops payouts by design; cache warmer fails open"
No numbers"Locks are fast""2–10 ms acquire, 1/hold-time throughput ceiling, TTL 15 s, renew at 5 s"

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Distributed locking is the smallest problem that contains the whole distributed-systems curriculum: partial failure, unbounded pauses, unreliable clocks, consensus, and the gap between what a component guarantees and what the business needs. It also contains a clean ownership trap: the lock team cannot deliver the guarantee alone. Candidates who see that the guarantee is completed by another team's code are thinking at Staff level.

1.2 The L5 → L6 → L7 Contrast — Visual#

Diagram: 1.2 The L5 → L6 → L7 Contrast — Visual

1.3 The Staff Question That Cuts Through Everything#

"Your lock holder pauses for longer than the lease. Another client acquires. The first one wakes up and writes. Which component rejects that write — and which team owns that component?"

If the answer is "the lock service", the candidate has not understood leases. If the answer names the resource and its owner, the rest of the interview is about tradeoffs, not correctness.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Efficiency → cheap, fail-open, no fencing

  • Constraint: avoid redundant work (cache rebuild, report generation, crawl of the same URL); a double run wastes resources but corrupts nothing
  • Store: single Redis SET NX PX, TTL ≈ expected work time × 2
  • Failure posture: fail-open — if Redis is down, run without the lock
  • Who pays for imperfection: the infra budget (duplicate CPU), nobody's data

Correctness → consensus store, fencing, fail-closed

  • Constraint: two concurrent effects violate an invariant — double payout, two writers on a file, two primaries for a shard
  • Store: etcd / ZooKeeper / a Chubby-like service; token = revision / zxid
  • Failure posture: fail-closed — no lock, no work; page on acquire errors
  • Who pays for imperfection: finance, customers, the data team cleaning up — so the token-checking resource team must sign on

Ownership / leader election → coarse, long-lived, epoch-fenced

  • Constraint: exactly one active scheduler / shard owner / primary for long periods; handoff must be clean
  • Store: session-based lease on etcd/ZK; epoch = token, carried on every downstream call
  • Failure posture: leader abdicates on renew failure; no leader during store outage
  • Who pays for imperfection: every consumer of the leader's output during a split-brain window — see Distributed Consensus for the election internals

2.2 When NOT to Use a Distributed Lock#

SituationUse InsteadWhy
Read-modify-write of one row/documentConditional write (WHERE version = :v, DynamoDB ConditionExpression)Atomic in the store; no extra dependency; no lease to expire
Counter / inventory decrementAtomic UPDATE ... SET n = n - 1 WHERE n > 0 or Redis DECR with checkLock throughput is capped at 1/hold-time; atomics run at store speed
"Run this job exactly once"Idempotency key + unique constraint on the resultDuplicate execution becomes harmless; see Idempotency
Per-entity serialization at high ratePartition by key, one consumer per partition (Kafka partition, actor)Single writer by construction; no acquire per message
Cross-service workflowSaga / workflow engine with state machineHolding a lock across network calls for seconds is a contention and failure magnet
Global uniqueness (usernames, IDs)Unique index or ID generation serviceThe database already provides linearizable uniqueness
Cross-region mutual exclusion on a hot pathHome-region ownership — see Multi-Region70–150 ms per acquire; partition = no progress anywhere

🎯 Staff Insight: "Every lock I don't build is a tier-0 dependency I don't have to keep alive at 3 AM. I reach for a lock when the critical section spans multiple resources or long-running work — not when one conditional write would do." See Contention for the full lock-free toolbox.

2.3 What the Interviewer Leaves Underspecified#

Interviewers deliberately omit:

  • What the lock protects — and whether that resource can check a token
  • Hold time — 5 ms critical section and 2-hour batch job are different systems
  • Acquire rate and key cardinality — 10 locks at 1/s or 10M locks at 50K/s
  • Waiting semantics — try-lock, block with timeout, fair FIFO queue?
  • Shared vs exclusive — readers/writer locks multiply complexity
  • Region scope — region-local or global?
  • Who runs it — platform service or library in each team

Staff engineers surface these. Senior engineers assume them away. The two that change the design most: can the resource check a token and what is the hold time.

2.4 Precise Terminology#

TermWhat It MeansCommon Confusion
LockExclusive right to act on a named resourceUsed loosely to mean "lease" — in distributed systems, there is no other kind
LeaseA lock with an expiry, renewed by the holderExpiry is decided by the lock server's clock, not the holder's
SessionA client's liveness relationship with the lock service; locks die with itOne session can carry many locks; losing it drops all of them
Fencing tokenMonotonically increasing number issued at each acquireNot the lease ID, not a UUID — it must be ordered
SequencerChubby's name for a token carrying lock name, mode and generationSame idea, richer payload
Epoch / termFencing token for leadershipWhat downstream systems check to reject a deposed leader
Lock-delayRefusing to re-grant a lock for a period after holder failureA time-based fallback when resources cannot check tokens
Advisory lockLock that is only honored by cooperating clientsAll distributed locks are advisory to the resource unless it fences
Herd effectAll waiters wake on release and stampede the storeFixed by watching only the predecessor

🎯 Staff Insight: If the interviewer says "lock", ask: "Do you mean exclusive access while I do work, or exactly-once execution of the work?" The second one usually does not need a lock.


3. The Five Fault Lines#

Each fault line has a technical side and an ownership side. In a Staff interview the ownership side — who absorbs a double-holder, who owns the token check, who gets paged when the lock store is down — is what's being scored.

3.1 Fault Line 1: Efficiency vs Correctness#

The tension: The same word "lock" covers "please don't duplicate this work" and "never let two writers touch this ledger". The first wants cheap and available. The second wants linearizable and fail-closed. One implementation cannot be both.

ChoiceWhat WorksWhat BreaksWho Pays
Treat everything as efficiency (Redis, no fencing)Cheap, sub-ms, trivially availablePayout/primary/file locks silently allow overlap on failover or pauseFinance / customers — duplicate effects discovered days later
Treat everything as correctness (etcd + fencing everywhere)SafeCache warmers and crawlers hammer a consensus cluster; etcd at 30K writes/s ceilingPlatform on-call — tier-0 cluster overloaded by non-critical traffic
Split by intent, declared per namespace (Staff default)Each lock gets the guarantee it needsRequires a registry and discipline; misclassification riskPlatform (registry), namespace owner (declares intent and signs)
Diagram: 3.1 Fault Line 1: Efficiency vs Correctness

L6 answer: "I'll make the intent a required field when a namespace is created. Efficiency namespaces route to Redis; correctness namespaces route to etcd and the SDK refuses to hand out a lock without returning the token, so callers can't ignore it."

L7 answer: Treats misclassification as the main risk. An efficiency lock created for a cron job three years ago becomes load-bearing when someone adds a ledger write inside it. The registry requires re-attestation of intent annually and flags namespaces whose holders write to fenced resources without passing a token.

🧭 Principal Insight: The expensive failure is not choosing the wrong store on day one. It's the drift from efficiency to correctness without anyone re-reviewing. Make intent a governed attribute with an owner and an expiry, the same way you treat data classification.

❌ Common L5 Trap: "We'll use Redlock for everything so it's both fast and safe." It is neither the fastest option for efficiency nor safe for correctness, and it makes every lock depend on five Redis masters.


3.2 Fault Line 2: Short vs Long Leases (Liveness vs Safety)#

The tension: A lease must be short enough that a crashed holder doesn't stall the system, and long enough that a merely slow holder doesn't lose it. GC pauses, VM live migration, page-cache stalls and network partitions put these two requirements in direct conflict.

ChoiceWhat WorksWhat BreaksWho Pays
Short TTL (≤5 s)Takeover within secondsFrequent false expiries; wasted work; renew traffic 3× higherPlatform (renew QPS), job owners (retries)
Long TTL (≥60 s)Pauses rarely cause expiryCrash = 60 s+ stall; pipelines miss SLAsDownstream consumers waiting on the stalled holder
Medium TTL + fencing + self-abort (Staff default)15 s takeover; overlap harmless because fencedRequires resource cooperationResource team (implements fencing)

The renewal math that makes medium TTLs work:

TTL            = 15s          # max stall the business tolerates after a crash
renew_interval = TTL / 3 = 5s # survive 2 lost renewals
self_abort_at  = 2/3 TTL = 10s since last successful renew
max_hold       = 15 min       # service refuses renewal beyond this
overlap window (unfenced) ≤ pause_duration − (TTL − time_since_last_renew)

"Without fencing, the overlap window is bounded only by the longest pause you'll ever see — which is unbounded. With fencing, overlap in belief is harmless, so TTL becomes purely a stall-time knob."

Diagram: 3.2 Fault Line 2: Short vs Long Leases (Liveness vs Safety)

L6 answer: "TTL is a stall budget, not a safety mechanism. I'd set 15 s, renew every 5 s, self-abort at 10 s without a successful renewal, and cap hold time at 15 minutes so a stuck worker can't hold a lock forever through a healthy renewer thread."

L7 answer: The stall budget is a business number: a 15 s stall on a payout batch is fine; on a trading-session leader it's not. TTL bounds per namespace are set by the namespace's business owner and reviewed with the SLO, not picked by whoever wrote the client.

🧭 Principal Insight: Org-wide JVM and runtime settings (heap sizes, GC algorithms, container CPU throttling) silently change the effective pause distribution. A platform-wide move to larger heaps can turn a safe 10 s TTL into a weekly overlap. Track lock_lease_expired_while_held_total as a fleet-wide SLI and review it when runtime defaults change.


3.3 Fault Line 3: Consensus Store vs Fast Store#

The tension: Consensus stores give linearizability and a free monotonic token at the cost of quorum writes, fsync latency and a cluster that needs specialist operation. Fast stores are cheap and familiar but replicate asynchronously.

ChoiceWhat WorksWhat BreaksWho Pays
Single Redis~0.5 ms, 100K+ ops/s, everyone knows itFailover loses acknowledged locks (async replication); no tokenCorrectness consumers if misused
Redlock (5 masters)Survives single-master lossTiming assumptions (drift, pauses); 5× infra; no tokenEveryone — false sense of safety
etcd / ZooKeeperLinearizable; revision/zxid token; watches; sessions2–10 ms; ~10–30K writes/s; 8 GB data; quorum ops expertisePlatform (operations), callers (latency)
Database row / advisory lockNo new infra; fencing via same transactionLock lost on DB failover; connection-bound; adds load to the primaryDB owners (connection and primary load)
Managed (DynamoDB conditional writes, cloud lock APIs)No ops; multi-AZClient-clock expiry checks; vendor-specific semanticsVendor dependency; app team builds heartbeats

Why single-Redis failover breaks correctness:

t=0     Client A: SET lock:x A NX PX 15000 → OK (acked by primary)
t=+1ms  Primary crashes before replicating to the replica
t=+3s   Sentinel promotes the replica — lock:x does not exist there
t=+3.1s Client B: SET lock:x B NX PX 15000 → OK
        Two holders. Neither saw an error. No metric fired.

L6 answer: "Redis for efficiency namespaces, etcd for correctness. If the company already runs a well-operated ZooKeeper for Kafka or HBase, I'd use that instead of standing up etcd — the deciding factor is who already has the on-call expertise. A second consensus system nobody knows how to restore is worse than either."

L7 answer: Picks one consensus substrate for the org based on operational maturity, not feature lists, and provides it as a managed internal service. Considers the database the team already runs: if the protected resource is a PostgreSQL table, a row with a version column inside the same transaction is both the lock and the fence — no separate service at all.

🧭 Principal Insight: The cheapest correct lock is often the resource's own transaction. A lock service is justified when the critical section spans resources that don't share a transaction boundary.


3.4 Fault Line 4: Lock vs Lock-Free Construction#

The tension: Locks are easy to bolt on and easy to reason about locally. They also serialize throughput, add a tier-0 dependency and introduce lease hazards. Lock-free constructions are safe by construction but require changing the data model or the processing topology.

ChoiceWhat WorksWhat BreaksWho Pays
Distributed lock around the critical sectionMinimal code change; works across heterogeneous resourcesThroughput ≤ 1/hold time; lease hazards; dependency on lock serviceOn-call (two systems to debug), users (contention latency)
Optimistic concurrency (CAS / version check)No dependency; store-speed throughput below ~5% conflict rateRetry storms above ~20–30% conflict rateApp team (retry logic)
Single-writer partitioningNo per-op coordination; scales with partitionsRebalancing; requires ownership (which itself needs leader election)Platform (partition assignment)
Idempotent effects + dedupeDuplicate execution harmlessNeeds a dedupe store and stable keysApp team (key design)

The contention ceiling:

Hold TimeMax Acquires/s per KeyAt 70% utilizationExample
1 ms1,000~700In-memory critical section — use an atomic instead
10 ms100~70One DB round-trip under lock
100 ms10~7Cross-service call under lock — red flag
10 s0.1~0.07Batch job — fine, that's what locks are for

L6 answer: "Locks are for long, multi-resource critical sections at low rates. For per-entity updates at high rate I'd use a conditional write or route each key to a single owner. The lock service should see hundreds of acquires per second fleet-wide, not hundreds of thousands."

L7 answer: Notices that single-writer partitioning moves the lock rather than removing it: partition ownership is itself a lease, but one acquired per partition per minutes instead of per operation. That's a 10,000× reduction in lock traffic and the pattern the org should standardize.

🎯 Staff Move: "Single-writer partitions still need a lease for ownership — but it's one acquire per partition per rebalance, not one per request. I've turned 20K lock ops/s into about 2."


3.5 Fault Line 5: Shared Lock Platform vs Embedded Locks#

The tension: Every team can add a Redis lock in an afternoon. A shared platform takes quarters and becomes a tier-0 dependency for many teams at once. But N home-grown lock libraries means N subtly different semantics, none fenced, and no one who can answer "what holds this lock right now?" during an incident.

ChoiceWhat WorksWhat BreaksWho Pays
Each team embeds its ownFast to start; no cross-team dependency4–6 incompatible libraries; no fencing; no visibility; same bug found 4 timesCustomers (correctness incidents), on-call (no shared runbook)
Central lock service, platform-ownedOne semantics, one runbook, quotas, auditCorrelated failure: one outage stops many teams; platform bottleneckPlatform (tier-0 on-call), consumers (shared fate)
Platform substrate + SDK, cell-local deployment (Staff/L7 default)Consistent semantics; blast radius limited to a cell/regionNeeds per-cell clusters and automationPlatform budget up front

L6 answer: "Platform owns the lock service and SDK; resource teams own fencing; namespace owners own TTLs and intent. One runbook. And I'd deploy one cluster per region — never a global lock cluster on the hot path."

L7 answer: Draws the boundary as a contract: the platform guarantees linearizable acquire, monotonic tokens, a 99.99% availability SLO per region and P99 acquire < 10 ms; it does not guarantee mutual exclusion at any resource. That sentence goes in the service's README, because it's the one consumers will misremember.

🧭 Principal Insight: Shared fate is the price of consistency. Limit it with cells: one lock cluster per region (or per cell), with namespaces pinned to the cell of the resources they protect, so an etcd incident in one cell stops one cell's work.


4. Failure Modes & Operational Reality#

4.1 The Zombie Holder — Full Timeline#

The canonical failure: a holder pauses past its lease and resumes acting.

t=0:       Settlement worker A acquires "settle:batch-9921", token 50311, TTL 15s
t=+4s:     A renews successfully (next renew due t=+9s)
t=+6s:     A's 28 GB heap triggers a full GC. Stop-the-world begins.
t=+21s:    Lease expires on the lock server. Key deleted. Waiter B notified.
t=+21.01s: B acquires, token 50312, starts the batch from its checkpoint
t=+24s:    B writes 1,200 ledger rows with fence 50312
t=+38s:    A's GC ends after 32s. A's clock says ~32s passed, but A never
           checks the clock before its next write — it's mid-loop.
t=+38s:    A writes the next ledger row with fence 50311
           → fenced ledger: REJECTED (50311 < 50312). A aborts.
           → unfenced ledger: ACCEPTED. Duplicate payouts begin.
t=+40s:    A's renew fails with SESSION_EXPIRED — too late to matter.

Detection: lock_lease_expired_while_held_total (SDK reports renew failure after a pause), lock_fencing_rejections_total (resource), jvm_gc_pause_seconds_max > 0.5 × TTL, lock_holder_overlap_seconds (derived from token timeline).

Blast radius: every effect A performed between t=+21s and t=+40s; unbounded without fencing.

Mitigation: fencing at every resource the holder writes; SDK checks time_since_last_renew < self_abort_at before each effectful step (narrows but does not close the window); idempotent writes keyed on batch item.

Prevention: GC tuning or smaller heaps for lock-holding services; alert when gc_pause_max exceeds 30% of TTL; architecture review rejects correctness locks on unfenced resources.

Owner: the resource team owns the rejection; the job owner owns the abort path; the platform owns the expired-while-held metric.

4.2 Renewer Alive, Work Dead#

t=0:       Worker acquires "reindex:shard-17", TTL 15s, background renewer thread
t=+30s:    Main thread blocks on a socket read to a dead downstream (no timeout)
t=+30s →   Renewer keeps renewing every 5s. Lock held. No progress.
t=+6h:     Reindex SLA missed. Nobody else can take the shard. No alert fired,
           because from the lock service's perspective everything is healthy.

Detection: lock_hold_time_seconds P99 vs namespace's expected hold time; lock_held_without_progress_seconds (SDK tracks the worker's progress heartbeat).

Mitigation: renewals gated on a progress heartbeat from the worker; a hard max_hold after which the server refuses renewal; admin "break lock" increments the token so the zombie is fenced if it ever resumes.

Owner: platform (max-hold policy), job owner (progress heartbeat + timeouts on every I/O).

4.3 Clock Jumps and Redlock#

Redlock computes validity as TTL − elapsed − drift, assuming clocks on the five masters advance at roughly the same rate. A step change breaks it:

t=0:     Client A acquires on masters 1, 2, 3 (majority of 5). TTL 10s.
t=+2s:   NTP on master 3 steps its clock forward 9s (or an operator fixes it manually)
t=+2s:   Master 3 expires A's key — from its view 11s elapsed
t=+3s:   Client B acquires on masters 3, 4, 5 — also a majority
         Two holders. Each believes it holds a majority. No component errored.

Detection: node_timex_offset_seconds jumps, ntp_step_events_total; nothing in the lock path itself.

Mitigation: for efficiency, accept it. For correctness, don't use Redlock — use a consensus store whose leases are measured on the leader with monotonic clocks, plus fencing.

Owner: platform; the fix is a design decision, not an operational one.

4.4 Lock Store Loses Quorum#

t=0:      Routine etcd upgrade rolls node 4 in AZ-b
t=+40s:   AZ-b network event isolates nodes 3 and 4 (node 4 still restarting)
t=+40s:   Cluster has 3 of 5 reachable... until node 5 OOMs under the
          compaction backlog. 2 of 5 → no quorum.
t=+41s:   All Acquire/Renew calls fail. lock_acquire_errors_total spikes.
t=+56s:   TTL passes. Every correctness holder self-aborts. 380 jobs stop.
t=+56s:   Efficiency namespaces fail-open; cache warmers continue unlocked.
t=+12min: Quorum restored. Waiters re-acquire in FIFO order. Jobs resume
          from checkpoints. Payout batch finishes 14 minutes late. Zero dupes.

Detection: etcd_server_has_leader == 0, lock_acquire_errors_total rate, lock_renew_failures_total.

Blast radius: every correctness namespace in the cell. This is why the cluster is cell-local and why non-critical traffic is kept off it.

Mitigation: fail-closed by design; jobs checkpoint so resumption is cheap; staggered re-acquire with jitter to avoid a thundering herd when quorum returns.

Owner: platform on-call (restore quorum); namespace owners (their runbook says "wait, don't break locks").

4.5 Herd Effect and Hot Locks#

Two different problems that look alike in dashboards:

ProblemCauseSymptomFix
Herd effect2,000 waiters watch the same key; release wakes all; 2,000 acquire attempts hit etcdWrite spike on each release; P99 acquire jumps to 500 ms+Sequential waiter nodes, each watching only its predecessor (O(1) wakeups)
Hot lock / convoyArrival rate approaches 1/hold-timeQueue grows without bound; every waiter times out; renewals compete with acquiresShrink hold time, partition the resource, or replace with conditional writes

Detection: lock_waiters{ns,name} > 100, lock_acquire_wait_seconds P99, etcd_mvcc_put_total rate around releases.

Owner: namespace owner — a hot lock is a data-model problem the platform cannot fix.

4.6 Leaked Locks and Force-Unlock Accidents#

Locks without TTLs (database rows written as "locked=true", Redis keys set without PX) outlive their holders and stall work until a human intervenes. The human fix — deleting the key — is the most dangerous operation in the system, because if the holder is alive, there are now two.

Rule: "break lock" is an admin API that (a) records who and why, (b) bumps the token, so any surviving holder is fenced, and (c) requires the namespace's on-call to acknowledge.

Detection: lock_age_seconds max per namespace vs max_hold; lock_admin_break_total.

Owner: namespace owner approves; platform logs.

4.7 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Zombie holder (pause > TTL)lock_lease_expired_while_held_total, lock_fencing_rejections_totalEffects between expiry and abortFencing at resource; self-abortResource team + job owner
Renewer alive, work stucklock_hold_time_seconds P99 vs expectedOne lock, stalled indefinitelyProgress-gated renewal; max holdJob owner + platform
Redis failover drops lockNone in lock path (silent)Overlap until next acquire cycleDon't use Redis for correctnessPlatform (policy)
Clock step on Redlock nodentp_step_events_totalOverlap for lock lifetimeConsensus store + fencingPlatform
Lock store quorum lossetcd_server_has_leader, lock_acquire_errors_totalAll correctness work in the cellFail-closed, checkpoints, jittered resumePlatform on-call
Herd on releaseetcd_mvcc_put_total spikes, lock_waitersAll namespaces sharing the clusterPredecessor-watch queuePlatform (SDK)
Hot lock convoylock_acquire_wait_seconds P99, timeoutsOne namespace; can starve clusterPartition / lock-free redesignNamespace owner
Leaked locklock_age_seconds > max_holdStalled work on one resourceTTL mandatory; audited break that bumps tokenNamespace owner
Quota exhaustionlock_namespace_keys near quotaOne tenant; cluster if unenforcedPer-namespace quotasPlatform

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
Problem framingDesigns a lockSplits efficiency vs correctness vs leadership; asks what overlap costsAsks whether the lock should exist and how many lock systems the org already runs
Safety reasoningTTL + renewalLeases expire under live holders; fencing token checked by resourceFenced writes as a platform capability; design review requires a named token checker
Store choiceRedis or RedlockRedis for efficiency, consensus for correctness, explains Redlock gapOne substrate org-wide chosen on operational maturity; deprecation plan for the rest
Failure posture"HA cluster"Fail-closed vs fail-open per intent; quorum loss runbookCell-local lock clusters; quorum loss as a game-day scenario; tier-0 SLO
ContentionNot raisedThroughput ≤ 1/hold-time; proposes conditional writes / partitioningHot locks reported as data-model debt to owning teams
OwnershipImplicitPlatform / resource team / namespace owner splitContract language: what the platform guarantees and explicitly does not
CostNot raisedMentions infra footprintPrices clusters per cell, on-call, and the incident cost of unfenced locks

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
States the impossibility unprompted"No lease protocol can stop a paused holder from waking up and acting — the resource has to reject it."
Splits intents"Cache warming gets Redis and fails open; payouts get etcd and fail closed."
Names the token checker"The ledger team owns the fence <= :token predicate. Without it this is a hint."
Quantifies contention"A 20 ms hold caps us at 50 acquires/s per key — the flash sale needs a different design."
Offers lock-free alternatives"A conditional update on the version column replaces this lock entirely."
Defines the service contract"We guarantee linearizable acquire and monotonic tokens, not mutual exclusion at your database."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Redlock makes it safe"Doesn't understand timing assumptions or the absence of fencing
"Set TTL longer than the job"Treats probability reduction as a guarantee
"Background thread renews, so the lock never expires"Ignores pauses that stop the renewer and stuck workers that don't
Global lock for a multi-region hot pathIgnores 70–150 ms cross-region quorum latency and partition behavior
No failure posture for lock-store outageLeaves the most common real incident undesigned
Locks per row at high QPSDoesn't know the 1/hold-time ceiling

5.4 Common False Positives#

  • Deep Raft/Paxos internals ≠ lock service design. Explaining log replication for 10 minutes while never mentioning fencing is a miss. Point to Distributed Consensus and move on.
  • Redlock clock-drift math ≠ safety. Precisely computing TTL − elapsed − drift shows study, not judgment.
  • Quoting the Kleppmann post ≠ understanding it. The test is applying "efficiency vs correctness" to the interviewer's scenario.
  • Elaborate reader/writer lock design ≠ Staff. Shared locks are rarely the interview; lease safety is.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Intent & framing0–4 minEfficiency vs correctness; commit; state the impossibility
API & entities4–7 minAcquire/Renew/Release with tokens; no data-plane force-unlock
Architecture7–12 minFrontends + etcd + fenced resource; one diagram
Fencing & leases12–22 minPaused-holder timeline; who checks; TTL math
Failure posture22–30 minQuorum loss; Redis failover; per-intent behavior
Contention / alternatives30–38 min1/hold-time; conditional writes; partitions
Ownership & evolution38–45 minPlatform contract, quotas, multi-region, what I'd build next

6.2 How Interviewers Pivot — And What They're Testing#

PivotWhat They're TestingStrong Response Direction
"What if the holder pauses for a minute?"Lease safetyFencing timeline; resource owner
"Why not Redlock?"Depth on timing assumptionsIntents split; token makes Redlock unnecessary
"The resource is S3 / a partner API"Unfenceable resourcesIdempotency keys, conditional puts, lock-delay, reconciliation
"Make it global"Multi-region judgmentHome-region ownership instead of global locks
"10M locks, 50K acquires/s"Scale vs fitThat's not a lock-service workload; partition or CAS
"etcd is down"Failure postureFail-closed for correctness, fail-open for efficiency, jittered resume

6.3 What to Deliberately Skip#

Raft log replication internals (link Distributed Consensus and move on), Redlock's drift formula beyond one sentence, reader/writer fairness algorithms, byte-level etcd/ZK API differences, and any hint of implementing your own consensus.

6.4 Follow-Up Questions to Expect#

  1. "How is the fencing token generated, and why must it be monotonic rather than unique?"
  2. "What does the client do between losing its lease and finding out?"
  3. "How do waiters get notified without a herd?"
  4. "How would you implement this on DynamoDB with no consensus service?"
  5. "What happens to held locks during a rolling upgrade of the lock cluster?"
  6. "A team wants per-user locks for 50M users. What do you tell them?"
  7. "How do you lock across two regions?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design a distributed lock service."

Staff Answer

"Before the store — what does it cost if two holders overlap for 30 seconds? If it's duplicate work, that's an efficiency lock: single Redis, fail-open, done in an afternoon. If it's corrupted data or double-moved money, that's a correctness lock, and no lease system can guarantee exclusivity on its own — the resource has to check a fencing token. I'll design for correctness: etcd-backed leases, the revision as a monotonic token, 15 s TTL, fail-closed when the store is unavailable. I'll go: API → architecture → paused-holder timeline and fencing → failure posture → contention and when not to lock → ownership."

Why this is L6:

  • Asks for the cost of overlap before choosing a mechanism
  • States the impossibility and where the guarantee must live
  • Commits to a posture (fail-closed) and a store with reasons

What L7 adds:

  • Asks how many lock implementations already exist and which guard money
  • Frames the service contract — what the platform guarantees and what it explicitly does not
  • Proposes that the first deliverable is an audit of locks that can be deleted

Drill 2: Fencing Tokens — Mechanics#

Prompt: "Explain fencing tokens. Why isn't a UUID lock ID enough?"

Staff Answer

"The resource needs to answer one question: is this write from a holder older than one I've already accepted? A UUID tells you who, not when. A fencing token is strictly increasing across acquisitions, so the resource stores max_token and rejects anything lower. In etcd, the revision of the acquire transaction is cluster-wide monotonic — free. In ZooKeeper, the sequence number or zxid. On a single SQL database, a bigserial or the row's version column. The check has to be atomic with the write — UPDATE ... WHERE fence <= :token — otherwise you've reintroduced check-then-act. And tokens must survive store failover monotonically, which is why Redis INCR isn't enough: an async-replicated counter can go backwards after failover."

Why this is L6:

  • Distinguishes identity from ordering
  • Names token sources per store and the atomic-check requirement
  • Spots that failover can make a naive counter non-monotonic

What L7 adds:

  • Makes "fenced write" a storage-platform primitive with a standard column name and SDK helper
  • Treats token propagation across services (headers, job metadata) as a cross-team contract, not per-team code

Drill 3: "Use Redis" — Make It Concrete#

Prompt: "Our team already has Redis. Can we just use it for the payout-batch lock?"

Staff Answer

"For the cache warmer, yes. For payouts, not by itself. Redis replication is asynchronous — if the primary acks our SET NX and dies before replicating, the replica gets promoted without the lock and a second worker acquires. Redlock on five masters closes that hole but depends on bounded pauses and clock drift and still has no token. The cheapest correct option is probably already in our stack: the payout ledger is PostgreSQL, so the batch row itself is the lock — UPDATE batches SET owner = :w, fence = fence + 1 WHERE id = :b AND (owner IS NULL OR lease_until < now()) returning the new fence, and every ledger write checks fence. Same database, same transaction boundary, no new infrastructure. If the critical section spans the ledger and an external payout API, I'd add an idempotency key on the API call derived from batch ID and item."

Why this is L6:

  • Says yes to the right use, no to the wrong one, with the failure mechanism
  • Finds that the resource's own database is the cheapest correct lock
  • Handles the unfenceable external API with idempotency

What L7 adds:

  • Notes lease_until < now() uses the DB server clock — one clock, not N — and makes this "lease row" pattern a documented paved road
  • Counts how many teams have the same "Redis for payouts" pattern and schedules a sweep

Drill 4: The Lock Service Is Down#

Prompt: "etcd loses quorum for 12 minutes during peak. Walk me through what happens."

Staff Answer

"Acquires and renewals fail immediately; lock_acquire_errors_total pages platform on-call within a minute. Existing holders keep working until they can't renew — at 10 s since last renewal they self-abort, before the 15 s TTL. Correctness namespaces stop: payouts, shard reassignments, and leader-elected controllers all pause. Efficiency namespaces fail-open and keep running with duplicated work. When quorum returns, leases from before the outage are gone; waiters re-acquire with jittered backoff (0–5 s) so we don't send 20K acquires in one second. Jobs resume from checkpoints, so the cost is 12–15 minutes of delay, not data. The runbook explicitly says: do not break locks or bypass the lock service to 'unblock' payouts — that's how an outage becomes a double-payout."

Why this is L6:

  • Per-intent posture with numbers
  • Recovery herd anticipated
  • Runbook names the dangerous human action

What L7 adds:

  • Cell-local clusters so the blast radius is one region/cell, not the company
  • Quorum loss in the game-day calendar with consumer teams participating
  • Tier-0 SLO (99.99%) and change freeze rules for the lock cluster during peak events

Drill 5: Hot Lock#

Prompt: "During a flash sale, the per-SKU lock for one product has 8,000 waiters and acquire P99 is 30 seconds."

Staff Answer

"The lock is the bottleneck by construction. Hold time is ~15 ms (read stock, write stock, write order), so the ceiling is ~65 acquires/s, and we're getting thousands. Immediate: shed — return 'sold out / try again' above a waiter threshold rather than queueing for 30 s. Then remove the lock: UPDATE inventory SET stock = stock - 1 WHERE sku = :s AND stock > 0 is atomic at the database and runs at thousands/s on one row; or split the SKU's stock across 32 buckets with a single writer each. Orders are written asynchronously after the decrement succeeds. The lock was never needed for a single-row invariant."

Why this is L6:

  • Does contention math before proposing fixes
  • Sheds load immediately, redesigns afterwards
  • Replaces the lock rather than tuning it — see Flash Sales

What L7 adds:

  • Adds "lock on a hot path" to the launch-readiness checklist for peak events
  • Tracks hot-lock incidents as data-model debt owned by the product team, not the lock platform

Drill 6: Multi-Tenant Abuse of the Lock Service#

Prompt: "One team created 6 million lock keys and etcd's DB size alarm fired. Everyone's locks are failing."

Staff Answer

"etcd raised its space quota alarm and went read-only for writes — every acquire across every team fails. Immediate: identify the namespace by key count, revoke its sessions so leases expire and keys are deleted, compact and defragment, disarm the alarm. Then prevent it: per-namespace quotas on keys (say 10K default) and acquires/s (say 200/s), enforced at the frontend before etcd sees the write. That team's use case — per-user locks for 6M users — isn't a lock-service workload; per-user serialization should be a conditional write or a partitioned consumer. I'd sit with them on the redesign, not just block them."

Why this is L6:

  • Knows the actual failure (space quota → cluster-wide write refusal)
  • Quotas at the frontend, not trust
  • Redirects the workload instead of just rejecting it

What L7 adds:

  • Namespace onboarding requires declared cardinality and rate; anything above thresholds gets an architecture review
  • Considers dedicated cells for high-volume tenants rather than one shared cluster

Drill 7: Build vs Buy#

Prompt: "Should we build a lock service, or use etcd, ZooKeeper, DynamoDB or a cloud service directly?"

Staff Answer

"Never build the consensus part. The choice is which substrate and how thin our wrapper is. If we already operate ZooKeeper for Kafka or etcd for Kubernetes and have people who can restore it from a snapshot at 3 AM, use that. If we're all-in on AWS and don't run either, DynamoDB conditional writes with a lease-row pattern give us multi-AZ durability and a monotonic version for fencing with zero ops — at ~5–10 ms per operation, which is fine for coarse locks. What we do build is the thin layer: namespace registry, quotas, SDK with renew/self-abort/progress-gating, metrics, and the fenced-write helpers. That's 1–2 engineers for two quarters, not a consensus project."

Why this is L6:

  • Draws the build line above consensus, below policy
  • Picks based on existing operational expertise
  • Sizes the build

What L7 adds:

  • Counts lock-in: DynamoDB-based locks tie job orchestration to one cloud; acceptable if the data is already there
  • Writes the decision as an ADR with a revisit trigger (e.g., "if we exceed 5K acquires/s or go multi-cloud")
  • See Build vs Buy

Drill 8: Changing Lock Semantics Without an Outage#

Prompt: "We're migrating 40 namespaces from a Redis lock library to the etcd-backed service. How do you do it safely?"

Staff Answer

"The danger is a window where old clients lock in Redis and new clients lock in etcd — two lock systems means no mutual exclusion. Per namespace: (1) deploy the SDK in dual-acquire mode — acquire etcd first, then Redis, hold both, release both; old and new clients still exclude each other via Redis. (2) Once 100% of the namespace's clients are on the new SDK — verified by lock_client_version metrics — (3) start passing the etcd token to the resource and turn on fencing in shadow mode, logging would-be rejections. (4) Enforce fencing. (5) Drop the Redis acquire. Each step is a flag, reversible in minutes. Namespaces migrate in order of risk: efficiency first, payouts last."

Why this is L6:

  • Identifies the split-lock hazard during migration
  • Dual-acquire gives a safe overlap; shadow fencing de-risks enforcement
  • Ordered, flag-driven, reversible

What L7 adds:

  • Migration SLO and a published date the Redis lock library stops receiving fixes
  • Tracks the long tail (the last 5 namespaces always take as long as the first 35) with named owners

Drill 9: Cost#

Prompt: "Finance asks why the lock platform needs 15 dedicated nodes."

Staff Answer

"Three regions × one 5-node etcd cluster, each node a small compute instance with fast local SSD — roughly $250–500/month per node, so ~$4–7K/month in infra. The real cost is the 1.5 engineers who operate it. What it buys: the payout batch, shard ownership for 3 storage systems, and 60 controllers run with correct mutual exclusion. One duplicate-payout incident — like the $1.8M one in our timeline — costs more than five years of the platform. The cheaper alternative isn't Redis; it's deleting locks we don't need, which is why half of my roadmap is a lock-free migration."

Why this is L6:

  • Gives a number with assumptions
  • Compares against incident cost, not against zero

What L7 adds:

  • Shows unit economics: cost per namespace and per million acquires, and who should be charged back
  • Positions lock deletion as the main cost lever

Drill 10: Global Locks#

Prompt: "We're going multi-region. The payout lock needs to be global."

Staff Answer

"A global lock is a consensus cluster spread across regions: every acquire and renew pays a cross-region quorum round trip, 70–150 ms, and a partition isolating the minority regions stops their locked work entirely. Instead, I'd make the payout batch owned by one region — the home region of the ledger shard it writes to. The lock stays region-local and fast; cross-region coordination happens only when ownership moves, which is a rare, planned operation with a token bump. If the ledger is globally replicated with a single leader per shard, the lock lives next to that leader. Global locks exist — Spanner-backed systems do it — but I'd only accept that latency for coarse, rare operations like 'who runs the monthly close'."

Why this is L6:

  • Quantifies the cost of a global lock
  • Replaces it with ownership — see Multi-Region
  • Names the narrow case where global locking is acceptable

What L7 adds:

  • Aligns lock placement with data placement as an org rule: "locks live with the leader of the data they protect"
  • Plans failover: ownership transfer is part of the region-evacuation runbook and rehearsed

8. Deep Dive Scenarios#

Deep Dive 1: Peak-Traffic Incident — Lock Convoy During a Product Launch#

Context: At a ticketing launch, seat-hold requests time out. The seat-hold service takes a per-section lock (held ~40 ms) through the shared lock service. lock_acquire_wait_seconds P99 is 22 s, etcd write latency has climbed from 4 ms to 180 ms, and unrelated namespaces (payouts, controllers) are now timing out too. On-call escalates to you.

Questions to Surface First:

  • Is etcd slow because of this namespace's volume, or is something else wrong (disk, compaction)?
  • What's the acquire rate on the hot sections versus the 1/hold-time ceiling (~25/s)?
  • Are other namespaces' renewals failing — i.e., are we about to drop correctness locks across the cell?
  • Can seat holds be served without the lock for the next hour?

Typical L5 Approach: Scales up etcd nodes, raises the frontend timeout, and increases the lock TTL so holders don't lose locks while etcd is slow. Each change is defensible; together they make the queue longer and keep the shared cluster overloaded.

Staff Approach: Protects the cell first: rate-limits the seat-hold namespace at the frontend to its quota so payout and controller renewals recover. Then removes the lock from the hot path: seat holds become a conditional write on the seat row (WHERE status = 'AVAILABLE'), which the database handles at thousands/s. Post-incident, the namespace's quota is set from contention math, not trust.

Principal Approach: Treats the incident as evidence that a tier-0 dependency had no admission control and that a product team could put a high-rate workload on it without review. Introduces namespace classes (coarse / medium / never-on-hot-path) with enforced rate ceilings, adds lock usage to launch-readiness review, and splits the shared cluster so control-plane locks (controllers, shard ownership) never share fate with product locks.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Throttle the seat-hold namespace at the frontend to 100 acquires/s; watch lock_renew_failures_total for other namespaces drop
TriageConfirm etcd latency is load-driven (etcd_disk_wal_fsync_duration_seconds normal, put rate 10× baseline)
Quick fixFeature-flag seat holds to conditional-write path; fail fast with "section busy, try another" above 200 waiters
GuardrailsPer-namespace acquire quotas enforced; a dedicated cluster for platform-critical namespaces
Post-mortemWhy could a 40 ms-hold lock be put on a 5K req/s path without review?

Metrics to Watch: lock_acquire_wait_seconds{ns}, lock_waiters{ns}, etcd_request_duration_seconds P99, lock_renew_failures_total by namespace, seat_hold_success_rate.

Organizational Follow-up: Lock usage review in launch checklists; namespace classes with rate ceilings; the seat-hold team owns its redesign.

Ownership Question: "Who decides to throttle a product team's namespace during their launch?" Staff answer: Platform on-call, pre-authorized by the runbook — protecting tier-0 locks for payouts and controllers outranks one product's throughput. The product team is paged in parallel and owns the fallback UX.

Key Takeaway: "A shared lock service without admission control turns one team's contention into everyone's outage."

What clears the Staff bar:

  • Protects shared correctness locks before helping the hot namespace
  • Recognizes contention math makes the lock unfixable by scaling
  • Replaces the lock with a conditional write rather than tuning TTLs

Deep Dive 2: Silent Failure — Two Holders for Three Weeks#

Context: A data team finds that 0.3% of nightly export files in object storage are corrupted — interleaved content from two writers. The export job uses a Redis lock with a 60 s TTL. No alerts fired in three weeks.

Questions to Surface First:

  • Did Redis fail over during those nights? (Async replication loses locks.)
  • Do the export workers have long pauses — GC, container CPU throttling?
  • Does the object store write check anything, or is it last-writer-wins on overlapping multipart uploads?
  • Which downstream consumers read the corrupted files?

Typical L5 Approach: Finds two Redis failovers and several 70 s GC pauses, increases TTL to 300 s, adds a check that the lock is still held before each part upload. Reduces frequency; the check-then-act window remains.

Staff Approach: Classifies this as a correctness lock that was implemented as an efficiency lock. Moves it to the etcd-backed service; writes each export to a token-suffixed path (export/2026-10-01/v50312/) and promotes it with a conditional pointer update that only succeeds if the token is the highest seen. Stale writers produce orphan objects that lifecycle rules delete — they can no longer corrupt the published file.

Principal Approach: Asks how many other locks were classified efficiency-but-actually-correctness. Runs a sweep: every namespace whose holders write to durable stores gets reviewed. Makes "intent" a required, owned attribute and adds lock_holder_overlap_seconds (computed from token timelines) as a fleet SLI so overlap stops being silent anywhere.

Staff Approach — Full Reasoning
DimensionStaff Answer
Root causeUnfenced lock on a correctness path; Redis failover and GC pauses both create overlap
Immediate actionIdentify and regenerate the corrupted files; notify consumers
System fixConsensus-backed lock; token-suffixed write path; conditional promotion
Detection fixlock_lease_expired_while_held_total, export_promotion_conflicts_total, content checksums validated by readers
Broader questionWhich other "efficiency" locks guard durable writes?

Metrics to Watch: lock_holder_overlap_seconds, export_promotion_conflicts_total, jvm_gc_pause_seconds_max, redis_failover_total.

Organizational Follow-up: Lock-intent audit across namespaces; readers validate checksums; storage team publishes a "fenced publish" helper.

Ownership Question: "The export team chose Redis. Is this their incident?" Staff answer: Their implementation, but the platform made the unsafe option the easiest one. Platform owns making the correct path the default — the SDK should have required an intent declaration and refused unfenced correctness locks.

Key Takeaway: "Unfenced locks don't fail loudly — they fail as data corruption weeks later. Overlap must be measured, not assumed away."

What clears the Staff bar:

  • Reclassifies the lock rather than tuning TTL
  • Uses write-then-promote with tokens for an unfenceable store
  • Turns overlap into a measured SLI

Deep Dive 3: Large Customer Onboarding — A Team Wants 50K Locks per Second#

Context: The orders team plans to adopt the lock service for per-order locks: 50K acquires/s at peak, 20M distinct keys/day, hold time ~30 ms. They want to go live in a month.

Questions to Surface First:

  • What invariant does the per-order lock protect? Single-row or multi-resource?
  • Can concurrent updates to one order actually happen, and how often?
  • What is the current lock cluster's headroom? (~10–30K writes/s total, shared.)
  • What's the failure posture they need?

Typical L5 Approach: Sizes a bigger etcd cluster or a dedicated one, benchmarks 50K acquires/s, and proposes sharding etcd by key hash across several clusters.

Staff Approach: Explains this is 2–5× the write ceiling of a well-tuned etcd cluster and that every acquire, renew and release is a write — 150K writes/s, effectively. Finds that 99.9% of orders are updated by one service at a time; the real risk is rare concurrent updates from refunds and fulfillment. Proposes optimistic concurrency on the order row (version column) with retry, and Kafka partitioning by order ID for the event pipeline — zero lock-service traffic.

Principal Approach: Uses the request to publish the lock service's workload envelope: coarse locks, ≤ 200 acquires/s per namespace, ≤ 10K keys, holds ≥ 100 ms. Anything outside that gets an architecture consult, not a quota bump. Adds per-entity concurrency patterns (CAS, single-writer partitions) to the internal design handbook so teams find them before they find the lock service.

Staff Approach — Full Reasoning
PhaseWhat to Do
Capacity math50K acquires + 50K releases + renew traffic ≈ 100–150K writes/s vs ~20K/s ceiling
Invariant analysisSingle-row invariant → CAS; event ordering → partition by order ID
Conflict rateMeasure: if < 1% conflicts, optimistic concurrency is near free
Alternative designVersion column + bounded retries (3, jittered); Kafka key = order_id
CommitmentLock service declines the workload; platform team pairs on the redesign

Metrics to Watch: order_update_conflicts_total, order_update_retries_total, kafka_consumer_lag{topic=orders}.

Organizational Follow-up: Publish the workload envelope; onboarding form requires rate, cardinality, hold time and intent.

Ownership Question: "The orders VP says the lock service is 'blocking' their launch. Who decides?" Staff answer: The platform owns the envelope of what the lock service can safely carry — saying yes would endanger every tier-0 consumer. The orders team owns their data model. The escalation goes to the architecture review with the capacity math attached, and the platform offers engineers to help with the CAS redesign so "no" comes with a path.

Key Takeaway: "A lock service is for coarse coordination. Per-entity concurrency at high rate belongs in the data model."

What clears the Staff bar:

  • Converts the request into write math against a known ceiling
  • Finds the real invariant and the right lock-free construction
  • Says no with a path

Deep Dive 4: Post-Mortem — Redlock Allowed a Double Primary#

Context: A sharded storage system uses Redlock (5 Redis masters) to elect the primary for each shard. During a maintenance window, an operator corrected clock skew on two Redis hosts. Within a minute, shard 412 had two primaries accepting writes for 47 seconds; 3,900 writes diverged. You present the post-mortem.

Questions to Surface First:

  • Did both primaries hold a "majority"? Which masters' keys expired early?
  • Do replicas or clients reject writes from a stale primary? (Is there an epoch?)
  • How were the divergent writes detected, and are they reconcilable?
  • How many other systems use Redlock for leadership?
Diagram: Deep Dive 4: Post-Mortem — Redlock Allowed a Double Primary

Typical L5 Approach: Recommends disabling manual clock changes, using slew mode for NTP, and alerting on clock offset. All good hygiene; the design still depends on clocks for safety.

Staff Approach: States the root cause as design: leadership was decided by an algorithm whose safety depends on timing, and nothing downstream checked an epoch. Moves shard leadership to etcd leases with a monotonically increasing epoch; replicas and the client routing layer reject writes from lower epochs. Reconciles the 3,900 writes with a merge job and customer-facing comms where needed.

Principal Approach: Presents to leadership as a class of risk: "systems where correctness depends on clocks". Inventories every Redlock and time-based lock in the org, schedules their migration with owners and dates, and sets a standard: leadership must be epoch-fenced at the data path. Adds clock-step injection to the chaos program so the next timing-dependent design is found in a game day, not production.

Staff Approach — Full Reasoning
SectionContent
What happenedClock steps on 2 of 5 masters expired a majority-held lock early; a second node acquired a different majority
Why it wasn't caughtNo epoch at the data path; both primaries' writes were valid to replicas and clients
Immediate actionsFreeze writes to shard 412; reconcile 3,900 writes; identify affected customers
Systemic fixetcd-based leadership with epoch; storage replicas reject lower epochs; routing layer caches epoch
Ownership boundaryStorage team owns epoch checks; platform owns the lock substrate; SRE owns clock change policy

Metrics to Watch: shard_primary_count{shard} (must be ≤ 1 per epoch), storage_stale_epoch_rejections_total, ntp_step_events_total.

Organizational Follow-up: Redlock deprecation plan; leadership-fencing standard; chaos experiments for clock steps and process pauses.

Ownership Question: "The operator followed the runbook to fix clock skew. Are they responsible?" Staff answer: No. The runbook was correct for a system that shouldn't depend on clocks for safety. The accountable decision is the original design choice and the review that approved it — fix the design and the review checklist, not the operator.

Key Takeaway: "If correctness depends on clocks, an ordinary clock fix becomes a data-loss event. Fence leadership with epochs."

What clears the Staff bar:

  • Root cause is design, not operations
  • Epoch fencing at the data path, not just a better lock
  • Blameless framing with a systemic sweep

Deep Dive 5: Multi-Region Expansion — "Make the Lock Service Global"#

Context: The company is adding EU and APAC regions. Several teams ask for a single global lock namespace so a job runs "exactly once worldwide". The lock service today is one etcd cluster per region.

Questions to Surface First:

  • Which jobs truly need global exclusivity vs per-region exclusivity?
  • Where does the data each job writes live? Is it globally replicated with a single leader, or per region?
  • What happens to those jobs if a region is partitioned for 30 minutes?
  • What latency can the critical sections tolerate?
Diagram: Deep Dive 5: Multi-Region Expansion — "Make the Lock Service Global"

Typical L5 Approach: Proposes a 5-node etcd cluster spread across 3 regions (2-2-1) so locks are global, and measures ~120 ms acquires. Accepts the latency as the price of correctness.

Staff Approach: Rejects global locks on the data path. Assigns each job and each data shard a home region; locks stay regional and fast. A small global ownership registry (cross-region consensus, rare writes) records which region owns which shard, and moving ownership bumps the fencing epoch. Region evacuation transfers ownership as a scripted step.

Principal Approach: Makes "locks live with the leader of the data they protect" an org standard, aligned with the data-residency and home-region model. Funds the global ownership registry as part of the multi-region platform, not the lock team, and adds ownership transfer to the quarterly region-evacuation drill.

Staff Approach — Full Reasoning
OptionTradeoffs
Global etcd across regionsCorrect; 70–150 ms per acquire and renew; minority-region partition stops all locked work there
Per-region clusters, home-region ownershipFast and partition-tolerant; requires ownership model and transfer protocol
Per-region with "exactly once worldwide" via idempotencyJobs may run in two regions, effects deduplicated globally; needs global idempotency store

Metrics to Watch: ownership_transfer_duration_seconds, lock_acquire_latency_seconds{region}, cross_region_registry_write_latency.

Organizational Follow-up: Each team declares per job: regional or global; global requires review. Evacuation runbook includes ownership transfer.

Ownership Question: "Who owns the global ownership registry?" Staff answer: The multi-region platform team, because its correctness is a region-failover concern. The lock team consumes it.

Key Takeaway: "Don't make locks global. Make ownership regional and the transfer of ownership rare, fenced and rehearsed."

What clears the Staff bar:

  • Quantifies global lock cost and partition behavior
  • Substitutes ownership for coordination
  • Ties lock placement to data placement

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Split lock requests into efficiency, correctness and leadership intents and pick a store for each
  • Walk the paused-holder timeline and explain why no lease protocol can prevent it
  • Design fencing at the resource for SQL, object storage and unfenceable external APIs
  • Explain the Kleppmann–antirez debate in two sentences and apply it to a scenario
  • Set TTL, renewal cadence, self-abort deadline and max hold time with reasons
  • Compute the 1/hold-time contention ceiling and propose lock-free alternatives
  • Define failure posture per intent when the lock store loses quorum
  • Draw the ownership line between the lock platform, resource teams and namespace owners

The Bar for This Question#

Mid-level (L4/E4): Knows SETNX and that locks need a TTL. Can write the Lua compare-and-delete release. Doesn't consider pauses, failover or what the lock protects.

Senior (L5/E5): Builds a clean lease-based lock on Redis or ZooKeeper, handles renewal and release correctly, knows Redlock exists and roughly why. Reaches fencing when prompted with "what if the holder pauses?" Treats the lock service as the owner of safety.

Staff+ (L6/E6+): Asks what overlap costs before choosing a store. States that leases cannot guarantee exclusivity and moves the guarantee to the resource with fencing, naming its owner. Separates fail-open and fail-closed by intent. Computes contention ceilings and proposes deleting locks where conditional writes suffice. Treats the lock service as a tier-0 shared product with quotas, an envelope and a contract. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "Most Distributed Locks Should Be Deleted"#

Audit any mid-sized codebase and the majority of distributed locks protect a single-record read-modify-write, a "run once" job, or a counter.

Lock UsageLock-Free ReplacementDependency Removed
Update one row safelyVersion column / conditional writeLock service on the hot path
Run a job onceIdempotency key + unique constraintLease hazards
Decrement inventoryAtomic conditional decrementContention ceiling
Serialize per-user eventsPartition by user IDPer-event acquires

The Staff position: A lock is the right tool for long, multi-resource critical sections at low rates. Everywhere else it's a tier-0 dependency added to avoid a data-model change.

Why this matters in interviews: Proposing to remove the lock — and showing the conditional write — is the single strongest signal you can give.

10.2 "Redlock Is the Wrong Answer to Both Questions"#

IntentWhat It NeedsWhat Redlock Gives
EfficiencyCheap, fast, usually right5 masters, majority round trips — more than needed
CorrectnessLinearizable acquire + monotonic tokenTiming-dependent safety, no token

The Staff position: Use one Redis for efficiency. Use a consensus store plus fencing for correctness. Once you add fencing, the extra Redis masters buy nothing. You don't have to say antirez is wrong about typical clocks; you have to say your design shouldn't depend on typical.

Why this matters in interviews: It shows you can resolve a famous debate by reframing intent instead of picking a side.

10.3 "A Lock Without a Fencing Token Is a Hint"#

Every distributed lock is advisory to the resource. If the resource doesn't check, the lock guarantees what the lock service's clock believed — nothing about the resource.

ResourceCan Fence?Fallback
SQL tableYes — predicate on fence column—
Object storePartially — conditional puts, token-suffixed pathsWrite-then-promote
KafkaProducer epochs (transactional producers fence zombies)Idempotent consumers
External payment APINoIdempotency keys
Filesystem via NFSRarelyLock-delay + reconciliation

The Staff position: Call an unfenced lock what it is in the design doc. Hints are fine for efficiency; for correctness, a hint must be paired with idempotency or reconciliation.

10.4 "Lease TTL Is a Business Decision"#

Engineers pick TTLs by gut feel — 30 s because it's round. But TTL is the guaranteed stall after a crash, and with fencing it's only that.

The Staff position: The namespace owner states the maximum tolerable stall; TTL follows. A payout batch can stall 60 s; a trading-session leader can't stall 5 s and should probably not be using a general-purpose lock service at all.

10.5 "The Lock Service Should Be Boring and Slow"#

Chubby's authors designed for coarse-grained locks held for hours and explicitly discouraged fine-grained use. Teams that optimize a lock service for 100K acquires/s invite exactly the workloads that should have been redesigned.

The Staff position: Publish a workload envelope (≤ 200 acquires/s per namespace, holds ≥ 100 ms, ≤ 10K keys). A lock service that's too fast to notice becomes a dependency of everything.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

A Staff engineer designs a correct lock service. A Principal engineer notices the org already has five: a Redis lock library in the Python monorepo, Redlock in the storage team, PostgreSQL advisory locks in billing, Kubernetes leases for controllers, and a ZooKeeper recipe nobody remembers writing. Two guard money, none are fenced, and no one can list what holds what during an incident. The L7 question is "which of these locks protect invariants, which can be deleted, and what single contract do the survivors run on?" Locking at org scale is a correctness governance problem: classifying intents, making fenced writes the path of least resistance, and limiting the shared fate of a tier-0 dependency.

🧭 Principal Move: "Before we build anything, I'd inventory every lock that guards durable writes or money. I expect most can become conditional writes. The survivors move onto one consensus-backed substrate with fencing enforced by the storage platform, and the lock libraries get a deprecation date."

The Org-Level Fault Line#

Central lock platform vs locks-in-the-data-layer.

OptionWhat WorksWhat BreaksWho Pays
Every team embeds locksSpeed; no shared fateUnfenced, inconsistent semantics; same incident repeats per teamCustomers, finance, scattered on-call
Central lock service for everythingOne contract, one runbookTier-0 shared fate; becomes a dumping ground for per-entity lockingPlatform on-call; every consumer during outages
Data layer owns concurrency; lock service only for coarse cross-resource coordination (L7 default)Most concurrency handled by CAS / partitions where data lives; lock service small and stableRequires storage platform investment in fenced writes and CAS helpersStorage platform budget; pays back in fewer correctness incidents

The deciding question: how many correctness incidents in the last 2 years trace to locking? If any involved money, the investment is justified by one avoided incident.

Cost Model#

Assumptions: cloud pricing; etcd nodes on small compute with local NVMe at ~$250–500/month; loaded engineer ~$25K/month; costs are for the lock platform only.

ScaleArchitectureInfra $/monthHeadcountOn-Call Load
Startup — ~20 namespaces, 1 regionUse existing PostgreSQL lease rows or managed DynamoDB locks; Redis for efficiency~$0–$500 (incremental)~0.1 FTEShared with DB on-call; < 1 page/quarter
Growth — ~150 namespaces, 2 regions, ~2K acquires/sPer-region 5-node etcd behind a thin lock API; SDK with fencing helpers~$3K–$6K1–1.5 FTETier-0 rotation shared with coordination services; ~1 page/month
Enterprise — ~800 namespaces, 4+ regions, cellsCell-local clusters (2 per region: platform-critical vs product), namespace registry, quotas, global ownership registry~$15K–$35K3–4 FTE (substrate, SDK, fenced-write helpers)Own rotation; quarterly quorum-loss game day

The line that matters to leadership: at every scale, the platform costs less than one duplicate-payout or split-brain incident. The bigger lever is deletion: each lock replaced by a conditional write removes a tier-0 dependency from a product's critical path for free.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversal CostWhy
Consensus substrate (etcd vs ZK vs managed)Two-way if wrapped1–2 quartersHidden behind the lock API
Fencing token semantics (64-bit monotonic, per-namespace or global)One-wayQuartersStored in every fenced resource's columns and checked by their code
Lock API contract (no data-plane force-unlock, token required on release)One-wayYearsEvery client SDK and runbook depends on it
Default TTL and renew ratiosTwo-wayHoursPolicy per namespace
Global lock namespacesOne-way-ishPainfulTeams build "exactly once worldwide" assumptions on them
Allowing per-entity high-rate useOne-wayQuarters to unwindWorkloads become dependent; removing them is a migration

🧭 Principal Insight: The irreversible decisions are contracts, not technology. Spend review time on token semantics and the API, not on etcd vs ZooKeeper.

The Standard I'd Write#

RFC: Distributed Locking and Mutual Exclusion Standard (v1)

Scope: Any mechanism that grants one process exclusive rights over a shared resource across machines.

Requirements:

  • New designs MUST document why a conditional write, single-writer partition or idempotency key is insufficient before adopting a lock.
  • Every lock namespace MUST declare intent (efficiency / correctness / leadership), owner, maximum stall time and expected rate in the registry.
  • Correctness and leadership locks MUST use the platform lock service and MUST pass the fencing token to every durable write; the resource MUST reject lower tokens atomically with the write.
  • Correctness locks MUST fail closed on lock-service unavailability; efficiency locks SHOULD fail open.
  • Locks MUST NOT span regions on a request path; cross-region exclusivity uses the ownership registry.
  • Redlock and Redis-based locks MUST NOT be used for correctness.
  • Breaking a lock MUST use the audited admin API, which bumps the token.

Exceptions: Architecture review within 5 business days; expiry ≤ 2 quarters.

Success metrics: Zero correctness incidents attributed to locking; 100% of correctness namespaces fenced within 3 quarters; lock-service P99 acquire < 10 ms; namespaces within workload envelope ≥ 95%.

What I'd Tell the VP#

"Several of our teams built their own locking, and at least two of those protect payments without the safeguard that stops a stalled server from writing twice. I'm proposing we remove the locks we don't need — most can be replaced with database features we already pay for — and move the rest onto one shared service with that safeguard built in. It costs about 1.5 engineers for three quarters and a few thousand dollars a month in infrastructure. The risk we're retiring is a duplicate-payment or data-corruption incident that we wouldn't detect for days. The main tradeoff is that the shared service becomes critical infrastructure, so we'll run one per region and rehearse its failure."

Principal Interview Signals#

SignalWhat It Sounds Like
Inventories before building"How many lock systems do we run, and which guard money?"
Moves concurrency into the data layer"Most of these become version checks; the lock service shrinks to coarse coordination."
Names the contract"We promise linearizable acquire and monotonic tokens — not exclusivity at your database."
Prices shared fate"One cluster per cell: an etcd outage costs one cell 15 minutes, not the company."
Governs drift"Intent is an owned attribute, re-attested yearly, because efficiency locks quietly become correctness locks."

Staff answers that L7 interviewers find insufficient:

  • "We'll use etcd with fencing tokens." — correct for one system; silent on the four other lock libraries and who migrates them.
  • "Correctness locks fail closed." — right, but no tier-0 SLO, no cell boundary, no game day.
  • "The resource checks the token." — which resources, owned by whom, and what's the paved-road helper so 40 teams don't each get it wrong?

Appendices

Appendix A: Mechanics in Depth#

A.1 Redis Single-Node Lock (Efficiency Only)#

# Acquire
SET lock:{ns}:{name} {owner_uuid} NX PX 15000      -> OK | nil

# Release — compare-and-delete, atomic via Lua
if redis.call("GET", KEYS[1]) == ARGV[1] then
    return redis.call("DEL", KEYS[1])
else
    return 0
end

# Renew — compare-and-extend
if redis.call("GET", KEYS[1]) == ARGV[1] then
    return redis.call("PEXPIRE", KEYS[1], ARGV[2])
else
    return 0
end

Right when: overlap is harmless. Wrong when: anything durable depends on exclusivity — async replication can lose the key on failover.

A.2 Redlock (For Recognition, Not Recommendation)#

N = 5 independent masters, quorum = 3
start = now()
for m in masters: SET key val NX PX ttl  (short per-node timeout, ~5-50ms)
elapsed  = now() - start
validity = ttl - elapsed - drift_allowance        # drift ~1% of ttl + 2ms
acquired = successes >= 3 and validity > 0
if not acquired: release on all masters

Why it fails for correctness: validity is computed with clocks that can step; a pause after acquired and before the write is not covered; no token for the resource to check.

A.3 ZooKeeper Lock Recipe#

create /locks/{name}/lock-  EPHEMERAL_SEQUENTIAL   -> lock-0000000042
children = getChildren(/locks/{name})               # sorted
if mine is lowest: acquired; token = my sequence (or czxid)
else: watch the node immediately before mine; wait for its deletion; re-check

Session expiry deletes the ephemeral node. Watching only the predecessor gives O(1) wakeups per release instead of a herd.

A.4 etcd Lease + Transaction#

lease = LeaseGrant(ttl=15)
KeepAlive(lease) every 5s
txn = Txn()
  .If(CreateRevision("/locks/ns/name") == 0)
  .Then(Put("/locks/ns/name", owner, lease=lease))
  .Else(Get("/locks/ns/name"))
resp = txn.Commit()
token = resp.Header.Revision        # monotonic across the cluster
# waiters: put a key under /locks/ns/name/waiters/ with the lease,
# watch the key with the next-lower create_revision

The etcd concurrency package implements this pattern; its Mutex exposes the key's revision for fencing.

A.5 Lease Row in the Resource's Own Database#

-- Acquire / take over
UPDATE job_leases
   SET owner = :me, fence = fence + 1, lease_until = now() + interval '15 seconds'
 WHERE job = :job AND (owner IS NULL OR lease_until < now() OR owner = :me)
RETURNING fence;

-- Fenced work
UPDATE payouts SET status = 'SENT'
 WHERE batch = :job AND item = :i
   AND (SELECT fence FROM job_leases WHERE job = :job) = :my_fence;

One clock (the DB server), one transaction boundary, fencing for free. The best lock when the protected data lives in the same database. Note: lost on DB failover only if replication is async and the lease update was lost — the fence still prevents regression if the fence column replicates with the data.

A.6 DynamoDB Conditional-Write Lock#

PutItem(lock_table, {pk: name, owner: me, rvn: uuid(), version: v+1, lease_ms: 15000},
        ConditionExpression = "attribute_not_exists(pk) OR rvn = :observed_stale_rvn")

Contenders observe the RVN for one full lease period before taking over, avoiding cross-machine clock comparison. Use version (incrementing) as the fencing token for downstream resources.

Appendix B: Lock Naming, Granularity and Data Model#

B.1 Key Construction#

/locks/{namespace}/{resource_type}/{resource_id}
  e.g. /locks/payments/settlement-batch/2026-10-01-eu
  • Namespace first: quotas, ACLs and metrics aggregate by prefix
  • Never embed user-controlled strings without normalization — key cardinality explosions start here

B.2 Granularity#

GranularityContentionLock CountDeadlock Risk
Coarse (per job / shard)Higher per key, low overall rateSmallLow
Fine (per entity)Low per keyHuge — exceeds lock-service envelopeHigher when holding several
Hierarchical (intent locks)BalancedMediumRequires strict ordering

Multiple locks: acquire in a global canonical order (sorted keys) to avoid deadlock; bound total hold with a single session; prefer one coarser lock over three fine ones held together.

Deadlock in a leased world: leases convert deadlock into stall-then-timeout: two holders waiting on each other release after TTL. That's liveness by expiry, not by design — it shows up as lock_acquire_wait_seconds spikes and periodic expiries. Canonical ordering removes it.

Appendix C: Coordination Mechanisms — Quick Comparison#

MechanismLinearizableFencing TokenAcquire LatencyOps BurdenBest For
Redis single nodeNo (failover)No~0.5 msLowEfficiency
RedlockTiming-dependentNo~2–5 msMedium (5 masters)Nothing you can't do better
ZooKeeperYeszxid / sequence~2–10 msHighExisting ZK shops, leadership
etcdYesRevision~2–10 msMedium–HighCorrectness, K8s-adjacent orgs
DynamoDB conditionalYes (per item)Version attribute~5–10 msNoneAWS-native, coarse locks
DB lease rowYes (per DB)Fence column~1–5 msNone extraData in the same DB
Chubby-style serviceYesSequencer~msVery high (build)Hyperscalers

See ZooKeeper & etcd for operational detail and Distributed Coordination for leader election and membership patterns.

Appendix D: API Contract and Client Behavior#

D.1 Error Semantics#

ResponseMeaningClient Action
LOCK_HELDAnother holderBack off or wait via watch
TIMEOUTWaited wait_timeoutReturn to caller; do not retry tightly
SESSION_EXPIREDLease goneStop all work under this session immediately
UNAVAILABLEStore unreachableApply namespace posture: fail-closed or fail-open
QUOTA_EXCEEDEDNamespace over limitPage namespace owner; do not retry
STALE_TOKEN (from resource)Fenced outAbort, log, do not retry with same token

D.2 SDK Responsibilities#

acquire(ns, name):
    resp = api.Acquire(ns, name, request_id=uuid())
    start renew loop at ttl/3, gated on worker.progress_heartbeat < ttl
    on renew failure for > 2/3 ttl: cancel worker context, emit lease_lost
    return Lock(token=resp.token, ctx=cancellable_context)

Retries and herds:

  • Acquire retries: exponential backoff 50 ms → 2 s with full jitter
  • Waiting: use server-side wait queues (predecessor watch), not client polling
  • Post-outage resume: random 0–5 s delay before first acquire

Appendix E: Observability#

E.1 Core Metrics#

# Service health
lock_acquire_latency_seconds{ns}        P50/P99
lock_acquire_errors_total{ns,code}
lock_renew_failures_total{ns}
etcd_server_has_leader, etcd_disk_wal_fsync_duration_seconds

# Safety
lock_lease_expired_while_held_total{ns}   # SDK: lost lease while working
lock_fencing_rejections_total{resource}   # emitted by resource owners
lock_holder_overlap_seconds{ns}           # derived from token timeline

# Contention and hygiene
lock_waiters{ns,name}, lock_acquire_wait_seconds{ns}
lock_hold_time_seconds{ns}, lock_age_seconds_max{ns}
lock_namespace_keys{ns}, lock_admin_break_total{ns}

E.2 Critical Alerts#

AlertThresholdSeverity
No leaderetcd_server_has_leader == 0 for 30 sPage platform
Acquire errors> 1% for 2 min on correctness namespacesPage platform
Lease lost while held> 0 on correctness namespaceTicket owner; page if repeating
Fencing rejections> 0Page resource owner — a zombie just tried to write
Hold time> max_holdPage namespace owner
Namespace keys> 80% of quotaTicket namespace owner

E.3 Debugging "Two Holders" Reports#

  1. Pull the token timeline for the lock: acquire, renew, expiry events with server timestamps
  2. Correlate with GC/pause metrics and node events for the old holder
  3. Check the resource's fencing rejections — if zero and both wrote, the resource isn't fencing
  4. If tokens never overlapped, the report is a client bug (not using the lock or the token)

Appendix F: Scale Evolution#

ScaleApproachDon't Build Yet
< 10 locks, one DBLease rows in the databaseAny lock service
Tens of namespacesManaged or existing etcd/ZK + thin SDKReader/writer locks, fair queuing
Hundreds of namespacesLock API, registry, quotas, per-region clustersGlobal namespaces, a lock-breaking UI
Thousands, multi-regionCell-local clusters, ownership registry (see Multi-Region), workload envelopeA custom consensus implementation — ever

Appendix G: Multi-Tenancy, Fairness and Cost#

One namespace can exhaust etcd's write capacity or storage quota for all. Enforce at the frontend: keys per namespace, acquires/s per namespace, max waiters per lock.

TierWhoIsolation
Platform-criticalControllers, shard ownership, payout leadershipDedicated cluster per cell
Product correctnessBatch jobs, exportsShared cluster, strict quotas
EfficiencyCache warmers, crawlersRedis, separate entirely

Chargeback: charge by namespace on acquires + key-hours. The goal isn't revenue; it's making a 6M-key namespace show up on someone's budget before it shows up in an incident.

  1. Loading the index…