Technologies referenced in this case study: ZooKeeper & etcd · Redis · PostgreSQL · DynamoDB
Related: Consensus / Coordination Service covers Raft internals · Distributed Coordination covers the broad pattern · Contention covers lock-free alternatives. This case study stays on the lock service as a product: its API, its guarantees, and who gets paged when it lies.
How to Use This Case Study#
Organized for interview use first, reference second. Read front-to-back once, then return to individual sections for targeted review.
| Mode | Time | What to Read |
|---|---|---|
| Quick Review | 15 min | Executive Summary → Interview Walkthrough → Fault Lines → Active Drills 1–3 |
| Targeted Study | 1–2 hrs | Executive Summary → Walkthrough → Fault Lines → Failure Modes → weak-spot Deep Dives |
| Deep Dive | 3+ hrs | Everything, including the Principal Lens and appendices |
What is a Distributed Lock Service? — Why interviewers pick this topic
A distributed lock lets one process among many claim exclusive (or shared) access to a named resource — a file, a shard, a job, an account — across machines that share no memory. Because machines crash, pause and lose network, every practical distributed lock is a lease: a lock that expires unless renewed. That single fact — the lock can be taken away from a holder who does not know it yet — is the entire interview.
Before vs After — nightly settlement job scenario:
Without a correct lock (Redis SETNX, 30s TTL, no fencing):
t=0: Worker A acquires "settle:2026-10-01", starts writing payouts
t=+12s: Worker A enters a 41-second stop-the-world GC pause
t=+30s: Lease expires. Worker B acquires the same lock legitimately
t=+31s: Worker B starts writing payouts for the same batch
t=+53s: Worker A wakes up, still believes it holds the lock, keeps writing
t=+2min: 4,112 merchants paid twice. $1.8M in duplicate transfers.
t=+3 days: Finance reconciliation finds it. Nobody was paged.
With a lease + fencing token checked by the ledger:
t=0: Worker A acquires lock, receives fencing token 7741
t=+12s: Same 41-second GC pause
t=+30s: Lease expires. Worker B acquires lock, receives token 7742
t=+31s: Worker B writes with token 7742 — ledger records max_token = 7742
t=+53s: Worker A wakes, writes with token 7741 — ledger REJECTS (7741 < 7742)
t=+53s: Worker A logs lock_fencing_rejections_total, aborts. Zero duplicates.
Why interviewers reach for this question: Locking looks like a solved problem — "use Redis SETNX" or "use ZooKeeper". It is actually the cleanest probe for whether a candidate understands that distributed systems cannot guarantee mutual exclusion by themselves. The lock service can only promise "at most one holder according to me". Whether that becomes "at most one writer to the resource" depends on something outside the lock service — the fencing check — and on someone owning it. Interviewers want to see you find that gap without being led to it.
Mechanics Refresher: Lock Implementations
| Implementation | How It Works | Pros | Cons |
|---|---|---|---|
Redis SET key val NX PX ttl | Single atomic set-if-absent with expiry; release via Lua compare-and-delete | ~0.2–1 ms acquire; trivial to run | Async replication: failover can lose the lock; no native fencing token |
| Redlock | Acquire on majority (3 of 5) independent Redis masters within a validity window | No single Redis SPOF | Safety depends on bounded clock drift and bounded pauses; still no fencing token |
| ZooKeeper ephemeral sequential znodes | Create /locks/x/lock-000042; lowest sequence holds; others watch predecessor | Linearizable; session-bound; zxid/sequence usable as token | 2–10 ms per acquire (quorum write + fsync); ops burden of a ZK ensemble |
| etcd lease + transaction | Txn(If create_revision==0) Then Put(key, lease); key dies with lease | Linearizable; revision is a free monotonic fencing token; Kubernetes-grade tooling | ~10–30K writes/s ceiling per cluster; 8 GB recommended data limit |
| DynamoDB conditional write | PutItem with attribute_not_exists(pk) OR expires_at < :now plus a record version number | Serverless, multi-AZ, pay per request | Expiry check uses client clock; ~5–10 ms; you build heartbeats yourself |
| PostgreSQL advisory lock | pg_try_advisory_lock(key) held by session or transaction | Zero new infrastructure; dies with the connection | Coupled to one DB primary; lost on failover; connection-pool footguns |
| Chubby-style lock service | Coarse-grained locks on a Paxos cell; sessions + KeepAlives; sequencers as tokens | Designed for exactly this; lock-delay protects un-fenced resources | Built, not bought — Google-scale investment |
For most production systems: a lease on a consensus store (etcd or ZooKeeper — or a managed equivalent) that returns a monotonic fencing token, with the token checked by the protected resource. Redis SET NX is fine — and the right answer — when the lock is only an efficiency optimization.
Executive Summary
If you only read one section, read this. Everything in the case study flows from the contrast below.
What This Interview Actually Tests#
A distributed lock is not a data-structure question. Everybody knows SETNX.
It is a safety-ownership question: when the lock lies — and it will, on the first long GC pause — who stops the second writer? It tests:
- Whether you ask what a double-holder costs before choosing an implementation
- Whether you know a lease can expire under a holder that is still running
- Whether you push the safety check to the resource (fencing) and name who owns that check
- Whether you can argue against the lock — conditional writes, single-writer partitioning, queues
- Whether you treat the lock service as a shared product with an API contract, quotas and an on-call
The key insight: A lock service can only guarantee "at most one holder according to the lock service." Mutual exclusion at the resource needs a fencing token that the resource checks. If the resource cannot check, you do not have a correctness lock — you have a strong hint, and your design must be safe when the hint is wrong.
The L5 → L6 → L7 Contrast — Start Here#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Redis SET NX PX with a TTL, release with a Lua script" | Asks "What breaks if two holders exist for 30 seconds?" and splits efficiency locks from correctness locks | Asks how many lock implementations the org already runs, which ones guard money or data, and whether any of those locks can be deleted entirely |
| Safety | Picks a TTL "longer than the job" and adds renewal | Says leases will expire under live holders (GC, VM pause, partition) and requires fencing tokens checked by the resource | Makes fencing a platform contract: storage and ledger teams expose token-checked writes as a paved-road API so product teams cannot forget it |
| Store | Redlock for HA "so Redis isn't a SPOF" | Redis for efficiency, etcd/ZooKeeper for correctness; explains why Redlock helps neither | Picks one consensus-backed lock substrate for the org, prices it, and sets a deprecation date for the other three |
| Failure | "Add replicas; renew the lease in a background thread" | Defines lock-service-down behavior per intent: efficiency locks fail-open, correctness locks fail-closed, with lock_acquire_errors_total paging | Designs correlated-failure posture: the lock cluster is a tier-0 dependency, so it is cell-local, never cross-region on the hot path, and its outage is a rehearsed game-day scenario |
| Ownership | Each service wires its own lock client | Lock service owned by platform; resource owners own fencing checks; one runbook | Redraws the boundary: platform owns the substrate and SDK, storage teams own fencing enforcement, product teams must justify every new correctness lock in design review |
| Alternatives | Not raised | Proposes conditional writes / single-writer partitions before reaching for a lock | Writes the standard that makes "lock-free first" the default and treats every new lock as technical debt with an owner |
Why "first move" separates levels
L5: Reaches for an implementation. SET lock:job NX PX 30000 is correct Redis, and the Lua compare-and-delete release shows real knowledge. But it commits to a mechanism before knowing whether a double-holder costs nothing (a duplicate cache rebuild) or $1.8M (a duplicate payout run).
L6: Separates intents out loud. "Before I pick a store — if two workers hold this lock at once, does something get computed twice, or does something get corrupted? Those need different systems. Efficiency locks can live on a single Redis. Correctness locks need a consensus store and a fencing token the resource checks."
L7: Asks whether the lock should exist. "Most locks I've audited guard something a conditional write or a single-writer partition would make safe by construction. Before we design a lock service, which of our locks guard invariants, and which are just deduplication we could do with idempotency keys?"
Why "safety" separates levels
L5: Believes a long enough TTL plus background renewal makes the lock safe. It reduces the probability of overlap; it does not bound it. A 41-second GC pause, a VM live-migration stall, or a network partition that blocks renewals will expire a lease under a holder that is still about to write.
L6: States the impossibility: "No lease protocol can stop a paused process from waking up and acting on stale belief. The only defense is at the resource: every write carries the fencing token, and the resource rejects tokens lower than the highest it has seen." Then names who owns that check — the storage team, not the lock team.
L7: Turns the check into infrastructure. If 40 teams each implement if token < max_token: reject, 10 will get it wrong. The ledger, the blob store and the job-state table expose fenced writes as a first-class API, and the design review checklist asks "which resource checks your token?" for every correctness lock.
Why "alternatives" separates levels
L5: Treats "we need mutual exclusion" as the requirement. Designs the lock.
L6: Treats mutual exclusion as one implementation of the real requirement — usually "this invariant must hold" or "this work must happen once". Offers cheaper constructions: UPDATE ... WHERE version = :v, partitioning so each key has exactly one writer, idempotency keys, or a queue with a single consumer per partition.
L7: Tracks locks as an org-level smell. Every correctness lock is a cross-team runtime dependency on a tier-0 service; the standard requires a written justification for why a lock-free construction does not work.
The Staff Positions#
| Position | Rationale |
|---|---|
| Leases, never indefinite locks | A crashed holder must not block the world forever; every lock has a TTL (typ. 10–30 s) and a renewal protocol |
| Fencing token on every correctness lock | Leases expire under live holders; only the resource can reject the stale writer |
| Consensus store for correctness, Redis for efficiency | etcd/ZooKeeper give linearizable acquire and a monotonic revision; Redis async replication can lose a lock on failover |
| Redlock is not the answer to either intent | Overkill for efficiency, unsafe for correctness without fencing — and with fencing you do not need it |
| Coarse-grained locks only on the lock service | Lock services are built for ~100s of acquires/s per hot key and minutes-to-hours holds, not per-row locking at 50K/s |
| Fail-closed for correctness, fail-open for efficiency | Losing the lock service must stop money-moving work but must not stop the cache warmer |
| Lock-free first | Conditional writes and single-writer partitions are safe by construction and remove a tier-0 dependency |
The Three Intents#
Three intents drive every design decision. Each leads to a different system.
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Efficiency (dedupe work) | Cheap and fast; occasional double execution is fine | Single Redis SET NX PX, short TTL, no fencing | Two workers do the same work; wasted CPU | Overlap acceptable; cost of a double run < cost of a consensus store |
| Correctness (protect an invariant) | Two writers corrupt data or move money twice | Consensus-backed lease + monotonic fencing token checked by the resource | Stale holder writes after lease expiry | Zero accepted stale writes; overlap in belief allowed, overlap in effect forbidden |
| Ownership / leader election (coarse, long-lived) | One scheduler, one shard owner, one primary for minutes to days | Session-based lease on etcd/ZK, epoch number as fencing token, watch for handoff | Split brain during handoff; failover takes TTL + election seconds | One effective leader per epoch; downstream rejects old epochs |
🎯 Staff Move: "I'll assume a correctness lock — two holders would corrupt data — because that's where fencing, lease expiry and the Redlock debate actually matter. If it turns out we only need efficiency, I'll downgrade to a single Redis and save us a consensus cluster. And before either, I want to check whether a conditional write removes the lock entirely."
The Five Fault Lines#
| # | Fault Line | The Tension |
|---|---|---|
| 1 | Efficiency vs Correctness | Is a double-holder a wasted CPU-minute or a corrupted ledger? The answer picks the store, the failure posture and the owner. |
| 2 | Short vs Long Leases (Liveness vs Safety) | Short TTL = fast takeover after a crash but more false expiries under pauses; long TTL = fewer false expiries but minutes of stall when a holder dies. |
| 3 | Consensus Store vs Fast Store | etcd/ZooKeeper (linearizable, 2–10 ms, ops-heavy) vs Redis (sub-ms, cheap, loses locks on failover). |
| 4 | Lock vs Lock-Free Construction | A lock serializes and adds a dependency; conditional writes / single-writer partitions are safe by construction but reshape the data model. |
| 5 | Shared Lock Platform vs Embedded Locks | One governed service with fencing as a contract vs every team wiring its own Redis — velocity now vs correlated failures later. |
In the Wild: Real Production Systems#
Why this section belongs here: Citing real systems shows you know these designs were paid for with incidents.
Google Chubby — Coarse-Grained Locks with Sequencers#
Google's Chubby (Burrows, OSDI 2006) is the archetype: a Paxos-replicated cell of five replicas serving coarse-grained locks held for hours or days — leader election for GFS and Bigtable, not per-request locking. Clients hold sessions renewed via KeepAlives (default lease on the order of 12 seconds). Chubby exposes sequencers — opaque byte strings carrying the lock name, mode and generation number — that servers can verify before acting. For resources that cannot check sequencers, Chubby offers lock-delay: after a holder fails, the lock cannot be re-granted for up to a minute, giving stale holders time to die.
Staff insight: The paper explicitly designs for "the lock can be lost while the holder still thinks it has it" and ships two answers: fencing (sequencers) when the resource can check, and a time-based buffer (lock-delay) when it cannot. That is exactly the answer interviewers want, from the people who built it.
Kubernetes — Leader Election on Lease Objects#
Every Kubernetes controller manager and scheduler runs leader election over a Lease object in the coordination.k8s.io API, backed by etcd. The defaults — leaseDuration 15 s, renewDeadline 10 s, retryPeriod 2 s — encode a tradeoff: a dead leader is replaced within ~15–17 s, and a leader that cannot renew within 10 s steps down before its lease expires. The resourceVersion on the Lease acts as optimistic concurrency for takeover.
Staff insight: The leader voluntarily abdicating at renewDeadline < leaseDuration is the client-side half of safety: stop acting before others may legitimately start. It narrows the overlap window but does not close it — which is why controllers are written to be idempotent against the API server's own optimistic concurrency.
Redlock and the Kleppmann–antirez Debate#
In 2016 Martin Kleppmann published "How to do distributed locking", arguing that Redlock (Redis's multi-master lock algorithm) is unsafe for correctness because it depends on bounded network delay, bounded process pauses and bounded clock drift — and offers no fencing token. Salvatore Sanfilippo (antirez) replied in "Is Redlock safe?", arguing the timing assumptions are reasonable in practice and that checks after acquisition mitigate pauses. Kleppmann's conclusion has become the industry heuristic: for efficiency, one Redis is enough; for correctness, use a consensus system and fencing tokens.
Staff insight: You do not need to win the debate in an interview. You need to show it collapses once you separate intents: Redlock is more machinery than efficiency needs and less safety than correctness needs.
What Interviewers Probe#
| After You Say... | They Will Ask... | (What They're Evaluating) |
|---|---|---|
"Redis SET NX PX" | "The holder pauses 40 seconds. What happens?" | Do you know leases expire under live holders? |
| "We'll renew the lease in a background thread" | "Renewal thread is fine, main thread is stuck on I/O. Now what?" | Liveness of the renewer ≠ liveness of the work |
| "Use Redlock for HA" | "Walk me through a clock jump on one of the five nodes." | Timing assumptions; fencing |
| "Fencing token" | "Who checks it? What if the resource is an S3 bucket or a third-party API?" | Ownership of the check; fallback when unfenceable |
| "etcd for correctness" | "etcd loses quorum. What do the 400 jobs that need locks do?" | Failure posture per intent; blast radius |
| "We'll lock per order ID" | "Flash sale: one SKU, 20K req/s. What's your lock throughput?" | Contention math: throughput ≤ 1 / hold time |
| "Lock service for everyone" | "Who owns it? How do you stop one team creating 10M locks?" | Platform thinking, quotas, governance |
System Architecture Overview#
Reading the diagram: The lock service is a thin product layer over a consensus store. Its job is linearizable acquire, lease management and a monotonic token. The safety guarantee is completed outside the lock service: the resource tier rejects stale tokens.
lock_fencing_rejections_totalis emitted by the resource, not the lock service — which is why the resource team must own it. The control plane registers every namespace with an owner and an intent, so an efficiency lock cannot silently become load-bearing.
Quick-Reference: The 30-Second Cheat Sheet#
| Topic | The L5 Answer | The L6 Answer — Say This | The L7 Answer — Say This |
|---|---|---|---|
| Store | "Redis, or Redlock for HA" | "Redis for efficiency locks; etcd/ZooKeeper for correctness — linearizable acquire and the revision is a free fencing token." | "One consensus-backed substrate for the org with a paved-road SDK; the three ad-hoc Redis lock libraries get a deprecation date." |
| Safety | "TTL longer than the job + renewal" | "Leases expire under live holders. The resource rejects writes with tokens lower than the max it has seen." | "Fenced writes are an API the storage platform provides. Design review rejects correctness locks without a named token checker." |
| Lease length | "30 seconds" | "TTL = max tolerable takeover delay; renew at TTL/3; holder self-aborts at 2/3 TTL without a successful renewal." | "TTL bounds are policy per namespace, set by the business owner of the stall cost, not by the developer." |
| Failure | "Replicas" | "Lock store down → correctness work stops (fail-closed), efficiency work proceeds (fail-open). Both paged on lock_acquire_errors_total." | "Lock store is tier-0 and cell-local. Its outage is a quarterly game-day scenario with a pre-approved list of what stops." |
| Contention | "Lock per row" | "Lock throughput ≤ 1 / hold time. A 20 ms hold caps one key at 50 acquires/s. Hot keys need partitioning or no lock." | "If a lock is hot, the data model is wrong. Hot-lock reports go to the owning team's quarterly review as debt." |
| Avoidance | — | "Can a WHERE version = :v conditional write or a single-writer partition replace this lock?" | "Lock-free first is the written standard; every new correctness lock needs a justification and an owner." |
Key Numbers Worth Memorizing#
| Metric | Value | Why It Matters |
|---|---|---|
Redis SET NX PX acquire, same AZ | ~0.2–1 ms | Why Redis is tempting; fine for efficiency |
| etcd / ZooKeeper acquire, same region | ~2–10 ms | One quorum write + fsync; dominated by disk latency |
| Cross-region consensus acquire | ~70–150 ms | Why global locks on a hot path are a non-starter |
| etcd practical write ceiling | ~10–30K writes/s per cluster | Every acquire, renew and release is a write |
| etcd recommended data size | ≤ 8 GB | Millions of fine-grained locks will not fit gracefully |
| Typical lease TTL | 10–30 s | Takeover delay after a crash ≈ TTL |
| Renewal cadence | TTL / 3 | Survive two lost renewals before expiry |
| JVM stop-the-world pauses on large heaps | seconds; tens of seconds pathologically | Longer than many TTLs — the classic zombie holder |
| ZooKeeper session timeout bounds | 2× to 20× tickTime (4–40 s at 2 s tick) | Session expiry = ephemeral lock deletion |
| Kubernetes leader election defaults | 15 s lease / 10 s renew deadline / 2 s retry | A production-tested ratio to quote |
| Max lock throughput on one key | 1 / hold time (10 ms hold → ~100/s) | The contention ceiling no hardware fixes |
| Chubby lock-delay | up to 1 min | Time buffer for resources that cannot check tokens |
Interview Walkthrough
The most common mistake: Candidates spend 20 minutes on
SETNXsyntax, Lua release scripts and Redlock quorum math, then run out of time before the only question that matters: what happens when the lock is wrong? Compress the mechanism to 10 minutes. Spend the rest on lease expiry, fencing, failure posture and whether the lock should exist.
Phase 1: Requirements & Framing (2–3 minutes)#
State functional requirements in 30 seconds:
"Clients acquire a named lock with a lease, renew it, release it, and optionally wait for it. Exclusive mode first; shared/read mode if needed."
Then spend the time on intent and the cost of failure:
"The first question is what a double-holder costs. If two workers rebuild the same cache, we waste a minute of CPU — that's an efficiency lock and a single Redis is fine. If two workers both run the payout batch, we move money twice — that's a correctness lock and it needs a different system. I'll assume correctness, because that's where the hard problems are."
Commit to non-functional constraints:
"For correctness: zero accepted stale writes at the resource, acquire latency under 10 ms in-region, takeover after a crashed holder within 30 s, and the service keeps working through a single node or AZ loss. I'm explicitly not promising that two clients never believe they hold the lock — no lease system can promise that. I'm promising that the resource never accepts work from the stale one."
🎯 Staff Move: The sentence "I'm not promising two clients never believe they hold the lock" is the one interviewers remember. It shows you know the impossibility result before they spring it on you, and it moves the guarantee to where it can actually be enforced.
Phase 2: Core Entities & API (1–2 minutes)#
Entities in 30 seconds:
| Entity | Fields | Note |
|---|---|---|
| Namespace | name, owner_team, intent, ttl_min/max, quota | Registered in the control plane; no anonymous locks |
| Session | session_id, client_identity, lease_id, ttl | One lease per client process; covers all its locks |
| Lock | ns/name, holder_session, mode, token, acquired_at | token = store revision at acquire, strictly increasing |
| Waiter | ns/name, session_id, seq | FIFO queue, each waiter watches only its predecessor |
API:
Acquire(ns, name, mode=EXCLUSIVE, wait_timeout=0s, request_id)
-> { token: uint64, session_id, lease_expires_at } | LOCK_HELD | TIMEOUT
Renew(session_id) -> { lease_expires_at } | SESSION_EXPIRED
Release(ns, name, token) -> OK | NOT_HOLDER (idempotent)
Describe(ns, name) -> { holder, token, held_for, waiters }
Watch(ns, name) -> stream of { ACQUIRED | RELEASED | EXPIRED, token }
"Three contract details matter. Acquire is idempotent on request_id, so a retried acquire after a timeout doesn't double-queue. Release requires the token, so a stale holder can't release the new holder's lock. And there is no ForceUnlock in the data-plane API — breaking a lock is an audited admin operation, because it's exactly the action that creates two holders."
Phase 3: High-Level Architecture (≤5 minutes)#
Staff candidates spend under 5 minutes here. The boxes are not the interview.
"Stateless frontends in each AZ, a 5-node etcd cluster spread over 3 AZs — it tolerates 2 node failures or one AZ. Acquire is a single etcd transaction: if the key's create_revision is 0, put it with the session's lease attached. The revision of that write is the fencing token, monotonic for free. Lease expiry deletes the key; waiters watch their predecessor, so one release wakes one waiter, not a thousand. The protected resource checks the token. That's the whole hot path."
"Why frontends instead of clients talking to etcd directly? Three reasons: quotas per namespace, so one team can't create 10M keys and push etcd past its 8 GB comfort zone; session multiplexing, so 5,000 workers share fewer leases; and an API we can keep stable while we change the store underneath."
Phase 4: Transition to Depth (1 minute)#
"The happy path is simple. What makes this hard is that the lock can be revoked from a holder that doesn't know it. Three places worth going deep: lease expiry and fencing — how we stay safe when the holder pauses; failure posture — what happens when etcd loses quorum; and contention — whether some of these locks should exist at all. I'd start with fencing, since it determines whether the rest is even correct."
🎯 Staff Move: Name the impossibility first, then offer the menu. You are steering toward the topic where you are strongest and that the interviewer is guaranteed to care about.
Phase 5: Deep Dives (25–30 minutes)#
For each: state the tradeoff → pick a position → quantify the cost → name who absorbs it.
Deep Dive 1: Lease expiry and fencing (7–8 min)
Walk the failure explicitly:
- Client A acquires, token 33. Starts a write that will take 200 ms.
- A's JVM enters a 25 s stop-the-world pause. Renewals stop.
- At TTL (15 s), etcd expires the lease and deletes the key.
- Client B acquires, token 34, writes to the resource. Resource records
max_token = 34. - A resumes. From A's point of view no time passed. A sends its write with token 33.
- Resource rejects:
33 < 34. A receivesSTALE_TOKEN, logs, aborts.
"Step 5 is unfixable from the client side. A checked 'do I still hold the lock' right before step 2 and the answer was yes. Any check-then-act on the client has a window. The resource is the only party that sees both writes in order, so it is the only party that can enforce the order."
What the resource must do:
-- Fenced write on a relational resource
UPDATE settlement_batches
SET status = 'PAID', paid_by = :worker, fence = :token
WHERE batch_id = :id
AND fence <= :token; -- 0 rows updated => stale holder, abort
"For an object store: a conditional put keyed on a metadata version, or write to a path that includes the token and promote atomically. For a third-party payment API that can't check our token, we fall back to an idempotency key derived from the batch ID — the provider deduplicates, so a stale holder's call is a no-op. If neither exists, I'd say plainly that the lock is a hint and we need reconciliation downstream."
🎯 Staff Move: Name who owns the check. "The ledger team owns
fence <= :token. If they don't sign up for that, this isn't a correctness lock and I'll write that in the design doc."
Deep Dive 2: Lease length and renewal (5 min)
"TTL is the time we're willing to stall if a holder dies. Renewal is TTL/3, so we survive two missed renewals. And the client stops doing work when it hasn't renewed within 2/3 of TTL — Kubernetes uses 15 s lease, 10 s renew deadline for exactly this reason."
| TTL | Takeover after crash | False expiries under 10 s pauses | Use |
|---|---|---|---|
| 5 s | ~5–7 s | Frequent on JVM services | Only with fencing and fast work units |
| 15 s | ~15–17 s | Rare | Default for job-level locks |
| 60 s | ~60 s | Very rare | Coarse leader election where a 1-minute stall is acceptable |
"With fencing, false expiries are a liveness cost — some work is wasted and retried. Without fencing, they're a safety violation. That's why I'd never pick a TTL to achieve safety; I pick it to bound stall time and use fencing for safety."
The subtle bug to volunteer: "A background renewal thread proves the process is alive, not that the work is progressing. If the worker thread is stuck on a 10-minute I/O, the renewer happily keeps the lock forever. So the SDK renews only if the worker has heartbeated progress within the last TTL — and every lock has a maximum hold time, say 15 minutes, after which the service refuses to renew."
Deep Dive 3: Store choice — and Redlock (5 min)
"For correctness I need linearizable acquire and a monotonic token. etcd gives both: a transaction on a Raft log and a cluster-wide revision. ZooKeeper gives both via ephemeral sequential nodes and zxid. Single Redis doesn't: replication is async, so if the primary acks my SET and dies before replicating, the replica is promoted without my lock and someone else acquires it. That's a correctness bug with no client-visible signal."
"Redlock fixes the single-node failover problem by acquiring on 3 of 5 independent masters, but its safety still assumes bounded clock drift and bounded pauses, and it has no token. If I add a fencing token to fix the pause problem, I need a monotonic counter — which is a consensus problem — and then I don't need Redlock. So Redlock is more machinery than efficiency needs and less safety than correctness needs."
"For efficiency locks, single Redis with SET NX PX and a compare-and-delete release is the right answer. Cheap, sub-millisecond, and losing a lock on failover just means some duplicate work."
Deep Dive 4: Failure posture when the lock service is down (4–5 min)
"etcd loses quorum — say two nodes in one AZ plus a bad deploy. Acquires and renewals fail. Existing leases can't be renewed, so after TTL every holder must assume it lost the lock. Per intent:"
- Correctness locks: fail-closed. Payout batch stops. That's a page, but a stopped batch is recoverable; a double payout is not.
- Efficiency locks: fail-open. The cache warmer runs without a lock and some work is duplicated. No page, just a counter.
- Leader election: current leader steps down at renewDeadline; nobody leads until quorum returns. Every controller that depends on a leader stops reconciling — so the lock cluster is tier-0 and its SLO is tighter than any consumer's.
"The SDK makes the posture explicit: Acquire(..., on_unavailable=FAIL_CLOSED) is required for correctness namespaces and validated against the namespace registry, so nobody accidentally fails a payout lock open."
Deep Dive 5: Contention and lock-free alternatives (4–5 min)
"If a lock is held for 20 ms, one key supports at most 50 acquires per second — and queueing theory says latency explodes well before that, around 70–80% utilization. A flash sale with 20K req/s on one SKU through a per-SKU lock is mathematically impossible. The fix is not a faster lock; it's removing it: an atomic decrement with a condition stock > 0, or partitioning inventory into 32 sub-buckets with a single writer each."
"My general test: if the critical section is a read-modify-write on one record, a conditional write replaces the lock. If it's work spanning many records over seconds, a lock — or better, a single-writer partition owned via leader election — is justified."
Phase 6: Wrap-Up (2–3 minutes)#
"To summarize: the lock service guarantees at most one holder according to itself, using leases on etcd with the revision as a fencing token. Safety at the resource comes from the fencing check, owned by the resource team. Correctness locks fail closed, efficiency locks fail open, and that posture is declared per namespace. What I'd build next: a namespace registry with quotas, a max-hold-time policy, and a quarterly audit of which locks could become conditional writes."
The organizational closer:
"The hard part long-term isn't the lock service — it's that every team will want to use it for things it wasn't built for. Per-row locking at 50K/s, cross-region locks, locks with no fencing guarding money. The platform team needs a design-review gate for correctness locks and the authority to say no."
🎯 Staff Move: End on the governance gap. Senior candidates end with "and we could use Redlock for more availability."
Common Timing Mistakes#
| Mistake | L5 Does This | L6 Does This Instead |
|---|---|---|
| Mechanism marathon | 10 min on Redlock quorum math and clock-drift formula | "Redlock is the wrong answer for both intents — here's why in 60 seconds" |
| No intent | Designs one lock for cache warming and payouts | Splits efficiency vs correctness in the first 2 minutes |
| Fencing only when prompted | Waits for "what about GC pauses?" | Volunteers the paused-holder timeline in Phase 4 |
| No resource owner | "The service checks the token" | "The ledger team owns fence <= :token; here's the SQL" |
| No failure posture | "etcd is highly available" | "Quorum loss stops payouts by design; cache warmer fails open" |
| No numbers | "Locks are fast" | "2–10 ms acquire, 1/hold-time throughput ceiling, TTL 15 s, renew at 5 s" |
1. The Staff Lens#
1.1 Why This Problem Exists in Staff Interviews#
Distributed locking is the smallest problem that contains the whole distributed-systems curriculum: partial failure, unbounded pauses, unreliable clocks, consensus, and the gap between what a component guarantees and what the business needs. It also contains a clean ownership trap: the lock team cannot deliver the guarantee alone. Candidates who see that the guarantee is completed by another team's code are thinking at Staff level.
1.2 The L5 → L6 → L7 Contrast — Visual#
1.3 The Staff Question That Cuts Through Everything#
"Your lock holder pauses for longer than the lease. Another client acquires. The first one wakes up and writes. Which component rejects that write — and which team owns that component?"
If the answer is "the lock service", the candidate has not understood leases. If the answer names the resource and its owner, the rest of the interview is about tradeoffs, not correctness.
2. Problem Framing & Intent#
2.1 The Three Intents — Explained#
Efficiency → cheap, fail-open, no fencing
- Constraint: avoid redundant work (cache rebuild, report generation, crawl of the same URL); a double run wastes resources but corrupts nothing
- Store: single Redis
SET NX PX, TTL ≈ expected work time × 2 - Failure posture: fail-open — if Redis is down, run without the lock
- Who pays for imperfection: the infra budget (duplicate CPU), nobody's data
Correctness → consensus store, fencing, fail-closed
- Constraint: two concurrent effects violate an invariant — double payout, two writers on a file, two primaries for a shard
- Store: etcd / ZooKeeper / a Chubby-like service; token = revision / zxid
- Failure posture: fail-closed — no lock, no work; page on acquire errors
- Who pays for imperfection: finance, customers, the data team cleaning up — so the token-checking resource team must sign on
Ownership / leader election → coarse, long-lived, epoch-fenced
- Constraint: exactly one active scheduler / shard owner / primary for long periods; handoff must be clean
- Store: session-based lease on etcd/ZK; epoch = token, carried on every downstream call
- Failure posture: leader abdicates on renew failure; no leader during store outage
- Who pays for imperfection: every consumer of the leader's output during a split-brain window — see Distributed Consensus for the election internals
2.2 When NOT to Use a Distributed Lock#
| Situation | Use Instead | Why |
|---|---|---|
| Read-modify-write of one row/document | Conditional write (WHERE version = :v, DynamoDB ConditionExpression) | Atomic in the store; no extra dependency; no lease to expire |
| Counter / inventory decrement | Atomic UPDATE ... SET n = n - 1 WHERE n > 0 or Redis DECR with check | Lock throughput is capped at 1/hold-time; atomics run at store speed |
| "Run this job exactly once" | Idempotency key + unique constraint on the result | Duplicate execution becomes harmless; see Idempotency |
| Per-entity serialization at high rate | Partition by key, one consumer per partition (Kafka partition, actor) | Single writer by construction; no acquire per message |
| Cross-service workflow | Saga / workflow engine with state machine | Holding a lock across network calls for seconds is a contention and failure magnet |
| Global uniqueness (usernames, IDs) | Unique index or ID generation service | The database already provides linearizable uniqueness |
| Cross-region mutual exclusion on a hot path | Home-region ownership — see Multi-Region | 70–150 ms per acquire; partition = no progress anywhere |
🎯 Staff Insight: "Every lock I don't build is a tier-0 dependency I don't have to keep alive at 3 AM. I reach for a lock when the critical section spans multiple resources or long-running work — not when one conditional write would do." See Contention for the full lock-free toolbox.
2.3 What the Interviewer Leaves Underspecified#
Interviewers deliberately omit:
- What the lock protects — and whether that resource can check a token
- Hold time — 5 ms critical section and 2-hour batch job are different systems
- Acquire rate and key cardinality — 10 locks at 1/s or 10M locks at 50K/s
- Waiting semantics — try-lock, block with timeout, fair FIFO queue?
- Shared vs exclusive — readers/writer locks multiply complexity
- Region scope — region-local or global?
- Who runs it — platform service or library in each team
Staff engineers surface these. Senior engineers assume them away. The two that change the design most: can the resource check a token and what is the hold time.
2.4 Precise Terminology#
| Term | What It Means | Common Confusion |
|---|---|---|
| Lock | Exclusive right to act on a named resource | Used loosely to mean "lease" — in distributed systems, there is no other kind |
| Lease | A lock with an expiry, renewed by the holder | Expiry is decided by the lock server's clock, not the holder's |
| Session | A client's liveness relationship with the lock service; locks die with it | One session can carry many locks; losing it drops all of them |
| Fencing token | Monotonically increasing number issued at each acquire | Not the lease ID, not a UUID — it must be ordered |
| Sequencer | Chubby's name for a token carrying lock name, mode and generation | Same idea, richer payload |
| Epoch / term | Fencing token for leadership | What downstream systems check to reject a deposed leader |
| Lock-delay | Refusing to re-grant a lock for a period after holder failure | A time-based fallback when resources cannot check tokens |
| Advisory lock | Lock that is only honored by cooperating clients | All distributed locks are advisory to the resource unless it fences |
| Herd effect | All waiters wake on release and stampede the store | Fixed by watching only the predecessor |
🎯 Staff Insight: If the interviewer says "lock", ask: "Do you mean exclusive access while I do work, or exactly-once execution of the work?" The second one usually does not need a lock.
3. The Five Fault Lines#
Each fault line has a technical side and an ownership side. In a Staff interview the ownership side — who absorbs a double-holder, who owns the token check, who gets paged when the lock store is down — is what's being scored.
3.1 Fault Line 1: Efficiency vs Correctness#
The tension: The same word "lock" covers "please don't duplicate this work" and "never let two writers touch this ledger". The first wants cheap and available. The second wants linearizable and fail-closed. One implementation cannot be both.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Treat everything as efficiency (Redis, no fencing) | Cheap, sub-ms, trivially available | Payout/primary/file locks silently allow overlap on failover or pause | Finance / customers — duplicate effects discovered days later |
| Treat everything as correctness (etcd + fencing everywhere) | Safe | Cache warmers and crawlers hammer a consensus cluster; etcd at 30K writes/s ceiling | Platform on-call — tier-0 cluster overloaded by non-critical traffic |
| Split by intent, declared per namespace (Staff default) | Each lock gets the guarantee it needs | Requires a registry and discipline; misclassification risk | Platform (registry), namespace owner (declares intent and signs) |
L6 answer: "I'll make the intent a required field when a namespace is created. Efficiency namespaces route to Redis; correctness namespaces route to etcd and the SDK refuses to hand out a lock without returning the token, so callers can't ignore it."
L7 answer: Treats misclassification as the main risk. An efficiency lock created for a cron job three years ago becomes load-bearing when someone adds a ledger write inside it. The registry requires re-attestation of intent annually and flags namespaces whose holders write to fenced resources without passing a token.
🧭 Principal Insight: The expensive failure is not choosing the wrong store on day one. It's the drift from efficiency to correctness without anyone re-reviewing. Make intent a governed attribute with an owner and an expiry, the same way you treat data classification.
❌ Common L5 Trap: "We'll use Redlock for everything so it's both fast and safe." It is neither the fastest option for efficiency nor safe for correctness, and it makes every lock depend on five Redis masters.
3.2 Fault Line 2: Short vs Long Leases (Liveness vs Safety)#
The tension: A lease must be short enough that a crashed holder doesn't stall the system, and long enough that a merely slow holder doesn't lose it. GC pauses, VM live migration, page-cache stalls and network partitions put these two requirements in direct conflict.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Short TTL (≤5 s) | Takeover within seconds | Frequent false expiries; wasted work; renew traffic 3× higher | Platform (renew QPS), job owners (retries) |
| Long TTL (≥60 s) | Pauses rarely cause expiry | Crash = 60 s+ stall; pipelines miss SLAs | Downstream consumers waiting on the stalled holder |
| Medium TTL + fencing + self-abort (Staff default) | 15 s takeover; overlap harmless because fenced | Requires resource cooperation | Resource team (implements fencing) |
The renewal math that makes medium TTLs work:
TTL = 15s # max stall the business tolerates after a crash
renew_interval = TTL / 3 = 5s # survive 2 lost renewals
self_abort_at = 2/3 TTL = 10s since last successful renew
max_hold = 15 min # service refuses renewal beyond this
overlap window (unfenced) ≤ pause_duration − (TTL − time_since_last_renew)
"Without fencing, the overlap window is bounded only by the longest pause you'll ever see — which is unbounded. With fencing, overlap in belief is harmless, so TTL becomes purely a stall-time knob."
L6 answer: "TTL is a stall budget, not a safety mechanism. I'd set 15 s, renew every 5 s, self-abort at 10 s without a successful renewal, and cap hold time at 15 minutes so a stuck worker can't hold a lock forever through a healthy renewer thread."
L7 answer: The stall budget is a business number: a 15 s stall on a payout batch is fine; on a trading-session leader it's not. TTL bounds per namespace are set by the namespace's business owner and reviewed with the SLO, not picked by whoever wrote the client.
🧭 Principal Insight: Org-wide JVM and runtime settings (heap sizes, GC algorithms, container CPU throttling) silently change the effective pause distribution. A platform-wide move to larger heaps can turn a safe 10 s TTL into a weekly overlap. Track
lock_lease_expired_while_held_totalas a fleet-wide SLI and review it when runtime defaults change.
3.3 Fault Line 3: Consensus Store vs Fast Store#
The tension: Consensus stores give linearizability and a free monotonic token at the cost of quorum writes, fsync latency and a cluster that needs specialist operation. Fast stores are cheap and familiar but replicate asynchronously.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Single Redis | ~0.5 ms, 100K+ ops/s, everyone knows it | Failover loses acknowledged locks (async replication); no token | Correctness consumers if misused |
| Redlock (5 masters) | Survives single-master loss | Timing assumptions (drift, pauses); 5× infra; no token | Everyone — false sense of safety |
| etcd / ZooKeeper | Linearizable; revision/zxid token; watches; sessions | 2–10 ms; ~10–30K writes/s; 8 GB data; quorum ops expertise | Platform (operations), callers (latency) |
| Database row / advisory lock | No new infra; fencing via same transaction | Lock lost on DB failover; connection-bound; adds load to the primary | DB owners (connection and primary load) |
| Managed (DynamoDB conditional writes, cloud lock APIs) | No ops; multi-AZ | Client-clock expiry checks; vendor-specific semantics | Vendor dependency; app team builds heartbeats |
Why single-Redis failover breaks correctness:
t=0 Client A: SET lock:x A NX PX 15000 → OK (acked by primary)
t=+1ms Primary crashes before replicating to the replica
t=+3s Sentinel promotes the replica — lock:x does not exist there
t=+3.1s Client B: SET lock:x B NX PX 15000 → OK
Two holders. Neither saw an error. No metric fired.
L6 answer: "Redis for efficiency namespaces, etcd for correctness. If the company already runs a well-operated ZooKeeper for Kafka or HBase, I'd use that instead of standing up etcd — the deciding factor is who already has the on-call expertise. A second consensus system nobody knows how to restore is worse than either."
L7 answer: Picks one consensus substrate for the org based on operational maturity, not feature lists, and provides it as a managed internal service. Considers the database the team already runs: if the protected resource is a PostgreSQL table, a row with a version column inside the same transaction is both the lock and the fence — no separate service at all.
🧭 Principal Insight: The cheapest correct lock is often the resource's own transaction. A lock service is justified when the critical section spans resources that don't share a transaction boundary.
3.4 Fault Line 4: Lock vs Lock-Free Construction#
The tension: Locks are easy to bolt on and easy to reason about locally. They also serialize throughput, add a tier-0 dependency and introduce lease hazards. Lock-free constructions are safe by construction but require changing the data model or the processing topology.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Distributed lock around the critical section | Minimal code change; works across heterogeneous resources | Throughput ≤ 1/hold time; lease hazards; dependency on lock service | On-call (two systems to debug), users (contention latency) |
| Optimistic concurrency (CAS / version check) | No dependency; store-speed throughput below ~5% conflict rate | Retry storms above ~20–30% conflict rate | App team (retry logic) |
| Single-writer partitioning | No per-op coordination; scales with partitions | Rebalancing; requires ownership (which itself needs leader election) | Platform (partition assignment) |
| Idempotent effects + dedupe | Duplicate execution harmless | Needs a dedupe store and stable keys | App team (key design) |
The contention ceiling:
| Hold Time | Max Acquires/s per Key | At 70% utilization | Example |
|---|---|---|---|
| 1 ms | 1,000 | ~700 | In-memory critical section — use an atomic instead |
| 10 ms | 100 | ~70 | One DB round-trip under lock |
| 100 ms | 10 | ~7 | Cross-service call under lock — red flag |
| 10 s | 0.1 | ~0.07 | Batch job — fine, that's what locks are for |
L6 answer: "Locks are for long, multi-resource critical sections at low rates. For per-entity updates at high rate I'd use a conditional write or route each key to a single owner. The lock service should see hundreds of acquires per second fleet-wide, not hundreds of thousands."
L7 answer: Notices that single-writer partitioning moves the lock rather than removing it: partition ownership is itself a lease, but one acquired per partition per minutes instead of per operation. That's a 10,000× reduction in lock traffic and the pattern the org should standardize.
🎯 Staff Move: "Single-writer partitions still need a lease for ownership — but it's one acquire per partition per rebalance, not one per request. I've turned 20K lock ops/s into about 2."
3.5 Fault Line 5: Shared Lock Platform vs Embedded Locks#
The tension: Every team can add a Redis lock in an afternoon. A shared platform takes quarters and becomes a tier-0 dependency for many teams at once. But N home-grown lock libraries means N subtly different semantics, none fenced, and no one who can answer "what holds this lock right now?" during an incident.
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Each team embeds its own | Fast to start; no cross-team dependency | 4–6 incompatible libraries; no fencing; no visibility; same bug found 4 times | Customers (correctness incidents), on-call (no shared runbook) |
| Central lock service, platform-owned | One semantics, one runbook, quotas, audit | Correlated failure: one outage stops many teams; platform bottleneck | Platform (tier-0 on-call), consumers (shared fate) |
| Platform substrate + SDK, cell-local deployment (Staff/L7 default) | Consistent semantics; blast radius limited to a cell/region | Needs per-cell clusters and automation | Platform budget up front |
L6 answer: "Platform owns the lock service and SDK; resource teams own fencing; namespace owners own TTLs and intent. One runbook. And I'd deploy one cluster per region — never a global lock cluster on the hot path."
L7 answer: Draws the boundary as a contract: the platform guarantees linearizable acquire, monotonic tokens, a 99.99% availability SLO per region and P99 acquire < 10 ms; it does not guarantee mutual exclusion at any resource. That sentence goes in the service's README, because it's the one consumers will misremember.
🧭 Principal Insight: Shared fate is the price of consistency. Limit it with cells: one lock cluster per region (or per cell), with namespaces pinned to the cell of the resources they protect, so an etcd incident in one cell stops one cell's work.
4. Failure Modes & Operational Reality#
4.1 The Zombie Holder — Full Timeline#
The canonical failure: a holder pauses past its lease and resumes acting.
t=0: Settlement worker A acquires "settle:batch-9921", token 50311, TTL 15s
t=+4s: A renews successfully (next renew due t=+9s)
t=+6s: A's 28 GB heap triggers a full GC. Stop-the-world begins.
t=+21s: Lease expires on the lock server. Key deleted. Waiter B notified.
t=+21.01s: B acquires, token 50312, starts the batch from its checkpoint
t=+24s: B writes 1,200 ledger rows with fence 50312
t=+38s: A's GC ends after 32s. A's clock says ~32s passed, but A never
checks the clock before its next write — it's mid-loop.
t=+38s: A writes the next ledger row with fence 50311
→ fenced ledger: REJECTED (50311 < 50312). A aborts.
→ unfenced ledger: ACCEPTED. Duplicate payouts begin.
t=+40s: A's renew fails with SESSION_EXPIRED — too late to matter.
Detection: lock_lease_expired_while_held_total (SDK reports renew failure after a pause), lock_fencing_rejections_total (resource), jvm_gc_pause_seconds_max > 0.5 × TTL, lock_holder_overlap_seconds (derived from token timeline).
Blast radius: every effect A performed between t=+21s and t=+40s; unbounded without fencing.
Mitigation: fencing at every resource the holder writes; SDK checks time_since_last_renew < self_abort_at before each effectful step (narrows but does not close the window); idempotent writes keyed on batch item.
Prevention: GC tuning or smaller heaps for lock-holding services; alert when gc_pause_max exceeds 30% of TTL; architecture review rejects correctness locks on unfenced resources.
Owner: the resource team owns the rejection; the job owner owns the abort path; the platform owns the expired-while-held metric.
4.2 Renewer Alive, Work Dead#
t=0: Worker acquires "reindex:shard-17", TTL 15s, background renewer thread
t=+30s: Main thread blocks on a socket read to a dead downstream (no timeout)
t=+30s → Renewer keeps renewing every 5s. Lock held. No progress.
t=+6h: Reindex SLA missed. Nobody else can take the shard. No alert fired,
because from the lock service's perspective everything is healthy.
Detection: lock_hold_time_seconds P99 vs namespace's expected hold time; lock_held_without_progress_seconds (SDK tracks the worker's progress heartbeat).
Mitigation: renewals gated on a progress heartbeat from the worker; a hard max_hold after which the server refuses renewal; admin "break lock" increments the token so the zombie is fenced if it ever resumes.
Owner: platform (max-hold policy), job owner (progress heartbeat + timeouts on every I/O).
4.3 Clock Jumps and Redlock#
Redlock computes validity as TTL − elapsed − drift, assuming clocks on the five masters advance at roughly the same rate. A step change breaks it:
t=0: Client A acquires on masters 1, 2, 3 (majority of 5). TTL 10s.
t=+2s: NTP on master 3 steps its clock forward 9s (or an operator fixes it manually)
t=+2s: Master 3 expires A's key — from its view 11s elapsed
t=+3s: Client B acquires on masters 3, 4, 5 — also a majority
Two holders. Each believes it holds a majority. No component errored.
Detection: node_timex_offset_seconds jumps, ntp_step_events_total; nothing in the lock path itself.
Mitigation: for efficiency, accept it. For correctness, don't use Redlock — use a consensus store whose leases are measured on the leader with monotonic clocks, plus fencing.
Owner: platform; the fix is a design decision, not an operational one.
4.4 Lock Store Loses Quorum#
t=0: Routine etcd upgrade rolls node 4 in AZ-b
t=+40s: AZ-b network event isolates nodes 3 and 4 (node 4 still restarting)
t=+40s: Cluster has 3 of 5 reachable... until node 5 OOMs under the
compaction backlog. 2 of 5 → no quorum.
t=+41s: All Acquire/Renew calls fail. lock_acquire_errors_total spikes.
t=+56s: TTL passes. Every correctness holder self-aborts. 380 jobs stop.
t=+56s: Efficiency namespaces fail-open; cache warmers continue unlocked.
t=+12min: Quorum restored. Waiters re-acquire in FIFO order. Jobs resume
from checkpoints. Payout batch finishes 14 minutes late. Zero dupes.
Detection: etcd_server_has_leader == 0, lock_acquire_errors_total rate, lock_renew_failures_total.
Blast radius: every correctness namespace in the cell. This is why the cluster is cell-local and why non-critical traffic is kept off it.
Mitigation: fail-closed by design; jobs checkpoint so resumption is cheap; staggered re-acquire with jitter to avoid a thundering herd when quorum returns.
Owner: platform on-call (restore quorum); namespace owners (their runbook says "wait, don't break locks").
4.5 Herd Effect and Hot Locks#
Two different problems that look alike in dashboards:
| Problem | Cause | Symptom | Fix |
|---|---|---|---|
| Herd effect | 2,000 waiters watch the same key; release wakes all; 2,000 acquire attempts hit etcd | Write spike on each release; P99 acquire jumps to 500 ms+ | Sequential waiter nodes, each watching only its predecessor (O(1) wakeups) |
| Hot lock / convoy | Arrival rate approaches 1/hold-time | Queue grows without bound; every waiter times out; renewals compete with acquires | Shrink hold time, partition the resource, or replace with conditional writes |
Detection: lock_waiters{ns,name} > 100, lock_acquire_wait_seconds P99, etcd_mvcc_put_total rate around releases.
Owner: namespace owner — a hot lock is a data-model problem the platform cannot fix.
4.6 Leaked Locks and Force-Unlock Accidents#
Locks without TTLs (database rows written as "locked=true", Redis keys set without PX) outlive their holders and stall work until a human intervenes. The human fix — deleting the key — is the most dangerous operation in the system, because if the holder is alive, there are now two.
Rule: "break lock" is an admin API that (a) records who and why, (b) bumps the token, so any surviving holder is fenced, and (c) requires the namespace's on-call to acknowledge.
Detection: lock_age_seconds max per namespace vs max_hold; lock_admin_break_total.
Owner: namespace owner approves; platform logs.
4.7 Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Zombie holder (pause > TTL) | lock_lease_expired_while_held_total, lock_fencing_rejections_total | Effects between expiry and abort | Fencing at resource; self-abort | Resource team + job owner |
| Renewer alive, work stuck | lock_hold_time_seconds P99 vs expected | One lock, stalled indefinitely | Progress-gated renewal; max hold | Job owner + platform |
| Redis failover drops lock | None in lock path (silent) | Overlap until next acquire cycle | Don't use Redis for correctness | Platform (policy) |
| Clock step on Redlock node | ntp_step_events_total | Overlap for lock lifetime | Consensus store + fencing | Platform |
| Lock store quorum loss | etcd_server_has_leader, lock_acquire_errors_total | All correctness work in the cell | Fail-closed, checkpoints, jittered resume | Platform on-call |
| Herd on release | etcd_mvcc_put_total spikes, lock_waiters | All namespaces sharing the cluster | Predecessor-watch queue | Platform (SDK) |
| Hot lock convoy | lock_acquire_wait_seconds P99, timeouts | One namespace; can starve cluster | Partition / lock-free redesign | Namespace owner |
| Leaked lock | lock_age_seconds > max_hold | Stalled work on one resource | TTL mandatory; audited break that bumps token | Namespace owner |
| Quota exhaustion | lock_namespace_keys near quota | One tenant; cluster if unenforced | Per-namespace quotas | Platform |
5. Evaluation Rubric#
5.1 Level-Based Signals#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Problem framing | Designs a lock | Splits efficiency vs correctness vs leadership; asks what overlap costs | Asks whether the lock should exist and how many lock systems the org already runs |
| Safety reasoning | TTL + renewal | Leases expire under live holders; fencing token checked by resource | Fenced writes as a platform capability; design review requires a named token checker |
| Store choice | Redis or Redlock | Redis for efficiency, consensus for correctness, explains Redlock gap | One substrate org-wide chosen on operational maturity; deprecation plan for the rest |
| Failure posture | "HA cluster" | Fail-closed vs fail-open per intent; quorum loss runbook | Cell-local lock clusters; quorum loss as a game-day scenario; tier-0 SLO |
| Contention | Not raised | Throughput ≤ 1/hold-time; proposes conditional writes / partitioning | Hot locks reported as data-model debt to owning teams |
| Ownership | Implicit | Platform / resource team / namespace owner split | Contract language: what the platform guarantees and explicitly does not |
| Cost | Not raised | Mentions infra footprint | Prices clusters per cell, on-call, and the incident cost of unfenced locks |
5.2 Strong Hire Signals#
| Signal | What It Sounds Like |
|---|---|
| States the impossibility unprompted | "No lease protocol can stop a paused holder from waking up and acting — the resource has to reject it." |
| Splits intents | "Cache warming gets Redis and fails open; payouts get etcd and fail closed." |
| Names the token checker | "The ledger team owns the fence <= :token predicate. Without it this is a hint." |
| Quantifies contention | "A 20 ms hold caps us at 50 acquires/s per key — the flash sale needs a different design." |
| Offers lock-free alternatives | "A conditional update on the version column replaces this lock entirely." |
| Defines the service contract | "We guarantee linearizable acquire and monotonic tokens, not mutual exclusion at your database." |
5.3 Lean No-Hire Signals#
| Signal | Why It Misses the Bar |
|---|---|
| "Redlock makes it safe" | Doesn't understand timing assumptions or the absence of fencing |
| "Set TTL longer than the job" | Treats probability reduction as a guarantee |
| "Background thread renews, so the lock never expires" | Ignores pauses that stop the renewer and stuck workers that don't |
| Global lock for a multi-region hot path | Ignores 70–150 ms cross-region quorum latency and partition behavior |
| No failure posture for lock-store outage | Leaves the most common real incident undesigned |
| Locks per row at high QPS | Doesn't know the 1/hold-time ceiling |
5.4 Common False Positives#
- Deep Raft/Paxos internals ≠ lock service design. Explaining log replication for 10 minutes while never mentioning fencing is a miss. Point to Distributed Consensus and move on.
- Redlock clock-drift math ≠ safety. Precisely computing
TTL − elapsed − driftshows study, not judgment. - Quoting the Kleppmann post ≠ understanding it. The test is applying "efficiency vs correctness" to the interviewer's scenario.
- Elaborate reader/writer lock design ≠ Staff. Shared locks are rarely the interview; lease safety is.
6. Interview Flow & Pivots#
6.1 Typical 45-Minute Shape#
| Phase | Time | Goal |
|---|---|---|
| Intent & framing | 0–4 min | Efficiency vs correctness; commit; state the impossibility |
| API & entities | 4–7 min | Acquire/Renew/Release with tokens; no data-plane force-unlock |
| Architecture | 7–12 min | Frontends + etcd + fenced resource; one diagram |
| Fencing & leases | 12–22 min | Paused-holder timeline; who checks; TTL math |
| Failure posture | 22–30 min | Quorum loss; Redis failover; per-intent behavior |
| Contention / alternatives | 30–38 min | 1/hold-time; conditional writes; partitions |
| Ownership & evolution | 38–45 min | Platform contract, quotas, multi-region, what I'd build next |
6.2 How Interviewers Pivot — And What They're Testing#
| Pivot | What They're Testing | Strong Response Direction |
|---|---|---|
| "What if the holder pauses for a minute?" | Lease safety | Fencing timeline; resource owner |
| "Why not Redlock?" | Depth on timing assumptions | Intents split; token makes Redlock unnecessary |
| "The resource is S3 / a partner API" | Unfenceable resources | Idempotency keys, conditional puts, lock-delay, reconciliation |
| "Make it global" | Multi-region judgment | Home-region ownership instead of global locks |
| "10M locks, 50K acquires/s" | Scale vs fit | That's not a lock-service workload; partition or CAS |
| "etcd is down" | Failure posture | Fail-closed for correctness, fail-open for efficiency, jittered resume |
6.3 What to Deliberately Skip#
Raft log replication internals (link Distributed Consensus and move on), Redlock's drift formula beyond one sentence, reader/writer fairness algorithms, byte-level etcd/ZK API differences, and any hint of implementing your own consensus.
6.4 Follow-Up Questions to Expect#
- "How is the fencing token generated, and why must it be monotonic rather than unique?"
- "What does the client do between losing its lease and finding out?"
- "How do waiters get notified without a herd?"
- "How would you implement this on DynamoDB with no consensus service?"
- "What happens to held locks during a rolling upgrade of the lock cluster?"
- "A team wants per-user locks for 50M users. What do you tell them?"
- "How do you lock across two regions?"
7. Active Drills#
Drill 1: The Opening#
Prompt: "Design a distributed lock service."
Staff Answer
"Before the store — what does it cost if two holders overlap for 30 seconds? If it's duplicate work, that's an efficiency lock: single Redis, fail-open, done in an afternoon. If it's corrupted data or double-moved money, that's a correctness lock, and no lease system can guarantee exclusivity on its own — the resource has to check a fencing token. I'll design for correctness: etcd-backed leases, the revision as a monotonic token, 15 s TTL, fail-closed when the store is unavailable. I'll go: API → architecture → paused-holder timeline and fencing → failure posture → contention and when not to lock → ownership."
Why this is L6:
- Asks for the cost of overlap before choosing a mechanism
- States the impossibility and where the guarantee must live
- Commits to a posture (fail-closed) and a store with reasons
What L7 adds:
- Asks how many lock implementations already exist and which guard money
- Frames the service contract — what the platform guarantees and what it explicitly does not
- Proposes that the first deliverable is an audit of locks that can be deleted
Drill 2: Fencing Tokens — Mechanics#
Prompt: "Explain fencing tokens. Why isn't a UUID lock ID enough?"
Staff Answer
"The resource needs to answer one question: is this write from a holder older than one I've already accepted? A UUID tells you who, not when. A fencing token is strictly increasing across acquisitions, so the resource stores max_token and rejects anything lower. In etcd, the revision of the acquire transaction is cluster-wide monotonic — free. In ZooKeeper, the sequence number or zxid. On a single SQL database, a bigserial or the row's version column. The check has to be atomic with the write — UPDATE ... WHERE fence <= :token — otherwise you've reintroduced check-then-act. And tokens must survive store failover monotonically, which is why Redis INCR isn't enough: an async-replicated counter can go backwards after failover."
Why this is L6:
- Distinguishes identity from ordering
- Names token sources per store and the atomic-check requirement
- Spots that failover can make a naive counter non-monotonic
What L7 adds:
- Makes "fenced write" a storage-platform primitive with a standard column name and SDK helper
- Treats token propagation across services (headers, job metadata) as a cross-team contract, not per-team code
Drill 3: "Use Redis" — Make It Concrete#
Prompt: "Our team already has Redis. Can we just use it for the payout-batch lock?"
Staff Answer
"For the cache warmer, yes. For payouts, not by itself. Redis replication is asynchronous — if the primary acks our SET NX and dies before replicating, the replica gets promoted without the lock and a second worker acquires. Redlock on five masters closes that hole but depends on bounded pauses and clock drift and still has no token. The cheapest correct option is probably already in our stack: the payout ledger is PostgreSQL, so the batch row itself is the lock — UPDATE batches SET owner = :w, fence = fence + 1 WHERE id = :b AND (owner IS NULL OR lease_until < now()) returning the new fence, and every ledger write checks fence. Same database, same transaction boundary, no new infrastructure. If the critical section spans the ledger and an external payout API, I'd add an idempotency key on the API call derived from batch ID and item."
Why this is L6:
- Says yes to the right use, no to the wrong one, with the failure mechanism
- Finds that the resource's own database is the cheapest correct lock
- Handles the unfenceable external API with idempotency
What L7 adds:
- Notes
lease_until < now()uses the DB server clock — one clock, not N — and makes this "lease row" pattern a documented paved road - Counts how many teams have the same "Redis for payouts" pattern and schedules a sweep
Drill 4: The Lock Service Is Down#
Prompt: "etcd loses quorum for 12 minutes during peak. Walk me through what happens."
Staff Answer
"Acquires and renewals fail immediately; lock_acquire_errors_total pages platform on-call within a minute. Existing holders keep working until they can't renew — at 10 s since last renewal they self-abort, before the 15 s TTL. Correctness namespaces stop: payouts, shard reassignments, and leader-elected controllers all pause. Efficiency namespaces fail-open and keep running with duplicated work. When quorum returns, leases from before the outage are gone; waiters re-acquire with jittered backoff (0–5 s) so we don't send 20K acquires in one second. Jobs resume from checkpoints, so the cost is 12–15 minutes of delay, not data. The runbook explicitly says: do not break locks or bypass the lock service to 'unblock' payouts — that's how an outage becomes a double-payout."
Why this is L6:
- Per-intent posture with numbers
- Recovery herd anticipated
- Runbook names the dangerous human action
What L7 adds:
- Cell-local clusters so the blast radius is one region/cell, not the company
- Quorum loss in the game-day calendar with consumer teams participating
- Tier-0 SLO (99.99%) and change freeze rules for the lock cluster during peak events
Drill 5: Hot Lock#
Prompt: "During a flash sale, the per-SKU lock for one product has 8,000 waiters and acquire P99 is 30 seconds."
Staff Answer
"The lock is the bottleneck by construction. Hold time is ~15 ms (read stock, write stock, write order), so the ceiling is ~65 acquires/s, and we're getting thousands. Immediate: shed — return 'sold out / try again' above a waiter threshold rather than queueing for 30 s. Then remove the lock: UPDATE inventory SET stock = stock - 1 WHERE sku = :s AND stock > 0 is atomic at the database and runs at thousands/s on one row; or split the SKU's stock across 32 buckets with a single writer each. Orders are written asynchronously after the decrement succeeds. The lock was never needed for a single-row invariant."
Why this is L6:
- Does contention math before proposing fixes
- Sheds load immediately, redesigns afterwards
- Replaces the lock rather than tuning it — see Flash Sales
What L7 adds:
- Adds "lock on a hot path" to the launch-readiness checklist for peak events
- Tracks hot-lock incidents as data-model debt owned by the product team, not the lock platform
Drill 6: Multi-Tenant Abuse of the Lock Service#
Prompt: "One team created 6 million lock keys and etcd's DB size alarm fired. Everyone's locks are failing."
Staff Answer
"etcd raised its space quota alarm and went read-only for writes — every acquire across every team fails. Immediate: identify the namespace by key count, revoke its sessions so leases expire and keys are deleted, compact and defragment, disarm the alarm. Then prevent it: per-namespace quotas on keys (say 10K default) and acquires/s (say 200/s), enforced at the frontend before etcd sees the write. That team's use case — per-user locks for 6M users — isn't a lock-service workload; per-user serialization should be a conditional write or a partitioned consumer. I'd sit with them on the redesign, not just block them."
Why this is L6:
- Knows the actual failure (space quota → cluster-wide write refusal)
- Quotas at the frontend, not trust
- Redirects the workload instead of just rejecting it
What L7 adds:
- Namespace onboarding requires declared cardinality and rate; anything above thresholds gets an architecture review
- Considers dedicated cells for high-volume tenants rather than one shared cluster
Drill 7: Build vs Buy#
Prompt: "Should we build a lock service, or use etcd, ZooKeeper, DynamoDB or a cloud service directly?"
Staff Answer
"Never build the consensus part. The choice is which substrate and how thin our wrapper is. If we already operate ZooKeeper for Kafka or etcd for Kubernetes and have people who can restore it from a snapshot at 3 AM, use that. If we're all-in on AWS and don't run either, DynamoDB conditional writes with a lease-row pattern give us multi-AZ durability and a monotonic version for fencing with zero ops — at ~5–10 ms per operation, which is fine for coarse locks. What we do build is the thin layer: namespace registry, quotas, SDK with renew/self-abort/progress-gating, metrics, and the fenced-write helpers. That's 1–2 engineers for two quarters, not a consensus project."
Why this is L6:
- Draws the build line above consensus, below policy
- Picks based on existing operational expertise
- Sizes the build
What L7 adds:
- Counts lock-in: DynamoDB-based locks tie job orchestration to one cloud; acceptable if the data is already there
- Writes the decision as an ADR with a revisit trigger (e.g., "if we exceed 5K acquires/s or go multi-cloud")
- See Build vs Buy
Drill 8: Changing Lock Semantics Without an Outage#
Prompt: "We're migrating 40 namespaces from a Redis lock library to the etcd-backed service. How do you do it safely?"
Staff Answer
"The danger is a window where old clients lock in Redis and new clients lock in etcd — two lock systems means no mutual exclusion. Per namespace: (1) deploy the SDK in dual-acquire mode — acquire etcd first, then Redis, hold both, release both; old and new clients still exclude each other via Redis. (2) Once 100% of the namespace's clients are on the new SDK — verified by lock_client_version metrics — (3) start passing the etcd token to the resource and turn on fencing in shadow mode, logging would-be rejections. (4) Enforce fencing. (5) Drop the Redis acquire. Each step is a flag, reversible in minutes. Namespaces migrate in order of risk: efficiency first, payouts last."
Why this is L6:
- Identifies the split-lock hazard during migration
- Dual-acquire gives a safe overlap; shadow fencing de-risks enforcement
- Ordered, flag-driven, reversible
What L7 adds:
- Migration SLO and a published date the Redis lock library stops receiving fixes
- Tracks the long tail (the last 5 namespaces always take as long as the first 35) with named owners
Drill 9: Cost#
Prompt: "Finance asks why the lock platform needs 15 dedicated nodes."
Staff Answer
"Three regions × one 5-node etcd cluster, each node a small compute instance with fast local SSD — roughly $250–500/month per node, so ~$4–7K/month in infra. The real cost is the 1.5 engineers who operate it. What it buys: the payout batch, shard ownership for 3 storage systems, and 60 controllers run with correct mutual exclusion. One duplicate-payout incident — like the $1.8M one in our timeline — costs more than five years of the platform. The cheaper alternative isn't Redis; it's deleting locks we don't need, which is why half of my roadmap is a lock-free migration."
Why this is L6:
- Gives a number with assumptions
- Compares against incident cost, not against zero
What L7 adds:
- Shows unit economics: cost per namespace and per million acquires, and who should be charged back
- Positions lock deletion as the main cost lever
Drill 10: Global Locks#
Prompt: "We're going multi-region. The payout lock needs to be global."
Staff Answer
"A global lock is a consensus cluster spread across regions: every acquire and renew pays a cross-region quorum round trip, 70–150 ms, and a partition isolating the minority regions stops their locked work entirely. Instead, I'd make the payout batch owned by one region — the home region of the ledger shard it writes to. The lock stays region-local and fast; cross-region coordination happens only when ownership moves, which is a rare, planned operation with a token bump. If the ledger is globally replicated with a single leader per shard, the lock lives next to that leader. Global locks exist — Spanner-backed systems do it — but I'd only accept that latency for coarse, rare operations like 'who runs the monthly close'."
Why this is L6:
- Quantifies the cost of a global lock
- Replaces it with ownership — see Multi-Region
- Names the narrow case where global locking is acceptable
What L7 adds:
- Aligns lock placement with data placement as an org rule: "locks live with the leader of the data they protect"
- Plans failover: ownership transfer is part of the region-evacuation runbook and rehearsed
8. Deep Dive Scenarios#
Deep Dive 1: Peak-Traffic Incident — Lock Convoy During a Product Launch#
Context: At a ticketing launch, seat-hold requests time out. The seat-hold service takes a per-section lock (held ~40 ms) through the shared lock service. lock_acquire_wait_seconds P99 is 22 s, etcd write latency has climbed from 4 ms to 180 ms, and unrelated namespaces (payouts, controllers) are now timing out too. On-call escalates to you.
Questions to Surface First:
- Is etcd slow because of this namespace's volume, or is something else wrong (disk, compaction)?
- What's the acquire rate on the hot sections versus the 1/hold-time ceiling (~25/s)?
- Are other namespaces' renewals failing — i.e., are we about to drop correctness locks across the cell?
- Can seat holds be served without the lock for the next hour?
Typical L5 Approach: Scales up etcd nodes, raises the frontend timeout, and increases the lock TTL so holders don't lose locks while etcd is slow. Each change is defensible; together they make the queue longer and keep the shared cluster overloaded.
Staff Approach: Protects the cell first: rate-limits the seat-hold namespace at the frontend to its quota so payout and controller renewals recover. Then removes the lock from the hot path: seat holds become a conditional write on the seat row (
WHERE status = 'AVAILABLE'), which the database handles at thousands/s. Post-incident, the namespace's quota is set from contention math, not trust.
Principal Approach: Treats the incident as evidence that a tier-0 dependency had no admission control and that a product team could put a high-rate workload on it without review. Introduces namespace classes (coarse / medium / never-on-hot-path) with enforced rate ceilings, adds lock usage to launch-readiness review, and splits the shared cluster so control-plane locks (controllers, shard ownership) never share fate with product locks.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Immediate (0–5 min) | Throttle the seat-hold namespace at the frontend to 100 acquires/s; watch lock_renew_failures_total for other namespaces drop |
| Triage | Confirm etcd latency is load-driven (etcd_disk_wal_fsync_duration_seconds normal, put rate 10× baseline) |
| Quick fix | Feature-flag seat holds to conditional-write path; fail fast with "section busy, try another" above 200 waiters |
| Guardrails | Per-namespace acquire quotas enforced; a dedicated cluster for platform-critical namespaces |
| Post-mortem | Why could a 40 ms-hold lock be put on a 5K req/s path without review? |
Metrics to Watch: lock_acquire_wait_seconds{ns}, lock_waiters{ns}, etcd_request_duration_seconds P99, lock_renew_failures_total by namespace, seat_hold_success_rate.
Organizational Follow-up: Lock usage review in launch checklists; namespace classes with rate ceilings; the seat-hold team owns its redesign.
Ownership Question: "Who decides to throttle a product team's namespace during their launch?" Staff answer: Platform on-call, pre-authorized by the runbook — protecting tier-0 locks for payouts and controllers outranks one product's throughput. The product team is paged in parallel and owns the fallback UX.
Key Takeaway: "A shared lock service without admission control turns one team's contention into everyone's outage."
What clears the Staff bar:
- Protects shared correctness locks before helping the hot namespace
- Recognizes contention math makes the lock unfixable by scaling
- Replaces the lock with a conditional write rather than tuning TTLs
Deep Dive 2: Silent Failure — Two Holders for Three Weeks#
Context: A data team finds that 0.3% of nightly export files in object storage are corrupted — interleaved content from two writers. The export job uses a Redis lock with a 60 s TTL. No alerts fired in three weeks.
Questions to Surface First:
- Did Redis fail over during those nights? (Async replication loses locks.)
- Do the export workers have long pauses — GC, container CPU throttling?
- Does the object store write check anything, or is it last-writer-wins on overlapping multipart uploads?
- Which downstream consumers read the corrupted files?
Typical L5 Approach: Finds two Redis failovers and several 70 s GC pauses, increases TTL to 300 s, adds a check that the lock is still held before each part upload. Reduces frequency; the check-then-act window remains.
Staff Approach: Classifies this as a correctness lock that was implemented as an efficiency lock. Moves it to the etcd-backed service; writes each export to a token-suffixed path (
export/2026-10-01/v50312/) and promotes it with a conditional pointer update that only succeeds if the token is the highest seen. Stale writers produce orphan objects that lifecycle rules delete — they can no longer corrupt the published file.
Principal Approach: Asks how many other locks were classified efficiency-but-actually-correctness. Runs a sweep: every namespace whose holders write to durable stores gets reviewed. Makes "intent" a required, owned attribute and adds
lock_holder_overlap_seconds(computed from token timelines) as a fleet SLI so overlap stops being silent anywhere.
Staff Approach — Full Reasoning
| Dimension | Staff Answer |
|---|---|
| Root cause | Unfenced lock on a correctness path; Redis failover and GC pauses both create overlap |
| Immediate action | Identify and regenerate the corrupted files; notify consumers |
| System fix | Consensus-backed lock; token-suffixed write path; conditional promotion |
| Detection fix | lock_lease_expired_while_held_total, export_promotion_conflicts_total, content checksums validated by readers |
| Broader question | Which other "efficiency" locks guard durable writes? |
Metrics to Watch: lock_holder_overlap_seconds, export_promotion_conflicts_total, jvm_gc_pause_seconds_max, redis_failover_total.
Organizational Follow-up: Lock-intent audit across namespaces; readers validate checksums; storage team publishes a "fenced publish" helper.
Ownership Question: "The export team chose Redis. Is this their incident?" Staff answer: Their implementation, but the platform made the unsafe option the easiest one. Platform owns making the correct path the default — the SDK should have required an intent declaration and refused unfenced correctness locks.
Key Takeaway: "Unfenced locks don't fail loudly — they fail as data corruption weeks later. Overlap must be measured, not assumed away."
What clears the Staff bar:
- Reclassifies the lock rather than tuning TTL
- Uses write-then-promote with tokens for an unfenceable store
- Turns overlap into a measured SLI
Deep Dive 3: Large Customer Onboarding — A Team Wants 50K Locks per Second#
Context: The orders team plans to adopt the lock service for per-order locks: 50K acquires/s at peak, 20M distinct keys/day, hold time ~30 ms. They want to go live in a month.
Questions to Surface First:
- What invariant does the per-order lock protect? Single-row or multi-resource?
- Can concurrent updates to one order actually happen, and how often?
- What is the current lock cluster's headroom? (~10–30K writes/s total, shared.)
- What's the failure posture they need?
Typical L5 Approach: Sizes a bigger etcd cluster or a dedicated one, benchmarks 50K acquires/s, and proposes sharding etcd by key hash across several clusters.
Staff Approach: Explains this is 2–5× the write ceiling of a well-tuned etcd cluster and that every acquire, renew and release is a write — 150K writes/s, effectively. Finds that 99.9% of orders are updated by one service at a time; the real risk is rare concurrent updates from refunds and fulfillment. Proposes optimistic concurrency on the order row (version column) with retry, and Kafka partitioning by order ID for the event pipeline — zero lock-service traffic.
Principal Approach: Uses the request to publish the lock service's workload envelope: coarse locks, ≤ 200 acquires/s per namespace, ≤ 10K keys, holds ≥ 100 ms. Anything outside that gets an architecture consult, not a quota bump. Adds per-entity concurrency patterns (CAS, single-writer partitions) to the internal design handbook so teams find them before they find the lock service.
Staff Approach — Full Reasoning
| Phase | What to Do |
|---|---|
| Capacity math | 50K acquires + 50K releases + renew traffic ≈ 100–150K writes/s vs ~20K/s ceiling |
| Invariant analysis | Single-row invariant → CAS; event ordering → partition by order ID |
| Conflict rate | Measure: if < 1% conflicts, optimistic concurrency is near free |
| Alternative design | Version column + bounded retries (3, jittered); Kafka key = order_id |
| Commitment | Lock service declines the workload; platform team pairs on the redesign |
Metrics to Watch: order_update_conflicts_total, order_update_retries_total, kafka_consumer_lag{topic=orders}.
Organizational Follow-up: Publish the workload envelope; onboarding form requires rate, cardinality, hold time and intent.
Ownership Question: "The orders VP says the lock service is 'blocking' their launch. Who decides?" Staff answer: The platform owns the envelope of what the lock service can safely carry — saying yes would endanger every tier-0 consumer. The orders team owns their data model. The escalation goes to the architecture review with the capacity math attached, and the platform offers engineers to help with the CAS redesign so "no" comes with a path.
Key Takeaway: "A lock service is for coarse coordination. Per-entity concurrency at high rate belongs in the data model."
What clears the Staff bar:
- Converts the request into write math against a known ceiling
- Finds the real invariant and the right lock-free construction
- Says no with a path
Deep Dive 4: Post-Mortem — Redlock Allowed a Double Primary#
Context: A sharded storage system uses Redlock (5 Redis masters) to elect the primary for each shard. During a maintenance window, an operator corrected clock skew on two Redis hosts. Within a minute, shard 412 had two primaries accepting writes for 47 seconds; 3,900 writes diverged. You present the post-mortem.
Questions to Surface First:
- Did both primaries hold a "majority"? Which masters' keys expired early?
- Do replicas or clients reject writes from a stale primary? (Is there an epoch?)
- How were the divergent writes detected, and are they reconcilable?
- How many other systems use Redlock for leadership?
Typical L5 Approach: Recommends disabling manual clock changes, using
slewmode for NTP, and alerting on clock offset. All good hygiene; the design still depends on clocks for safety.
Staff Approach: States the root cause as design: leadership was decided by an algorithm whose safety depends on timing, and nothing downstream checked an epoch. Moves shard leadership to etcd leases with a monotonically increasing epoch; replicas and the client routing layer reject writes from lower epochs. Reconciles the 3,900 writes with a merge job and customer-facing comms where needed.
Principal Approach: Presents to leadership as a class of risk: "systems where correctness depends on clocks". Inventories every Redlock and time-based lock in the org, schedules their migration with owners and dates, and sets a standard: leadership must be epoch-fenced at the data path. Adds clock-step injection to the chaos program so the next timing-dependent design is found in a game day, not production.
Staff Approach — Full Reasoning
| Section | Content |
|---|---|
| What happened | Clock steps on 2 of 5 masters expired a majority-held lock early; a second node acquired a different majority |
| Why it wasn't caught | No epoch at the data path; both primaries' writes were valid to replicas and clients |
| Immediate actions | Freeze writes to shard 412; reconcile 3,900 writes; identify affected customers |
| Systemic fix | etcd-based leadership with epoch; storage replicas reject lower epochs; routing layer caches epoch |
| Ownership boundary | Storage team owns epoch checks; platform owns the lock substrate; SRE owns clock change policy |
Metrics to Watch: shard_primary_count{shard} (must be ≤ 1 per epoch), storage_stale_epoch_rejections_total, ntp_step_events_total.
Organizational Follow-up: Redlock deprecation plan; leadership-fencing standard; chaos experiments for clock steps and process pauses.
Ownership Question: "The operator followed the runbook to fix clock skew. Are they responsible?" Staff answer: No. The runbook was correct for a system that shouldn't depend on clocks for safety. The accountable decision is the original design choice and the review that approved it — fix the design and the review checklist, not the operator.
Key Takeaway: "If correctness depends on clocks, an ordinary clock fix becomes a data-loss event. Fence leadership with epochs."
What clears the Staff bar:
- Root cause is design, not operations
- Epoch fencing at the data path, not just a better lock
- Blameless framing with a systemic sweep
Deep Dive 5: Multi-Region Expansion — "Make the Lock Service Global"#
Context: The company is adding EU and APAC regions. Several teams ask for a single global lock namespace so a job runs "exactly once worldwide". The lock service today is one etcd cluster per region.
Questions to Surface First:
- Which jobs truly need global exclusivity vs per-region exclusivity?
- Where does the data each job writes live? Is it globally replicated with a single leader, or per region?
- What happens to those jobs if a region is partitioned for 30 minutes?
- What latency can the critical sections tolerate?
Typical L5 Approach: Proposes a 5-node etcd cluster spread across 3 regions (2-2-1) so locks are global, and measures ~120 ms acquires. Accepts the latency as the price of correctness.
Staff Approach: Rejects global locks on the data path. Assigns each job and each data shard a home region; locks stay regional and fast. A small global ownership registry (cross-region consensus, rare writes) records which region owns which shard, and moving ownership bumps the fencing epoch. Region evacuation transfers ownership as a scripted step.
Principal Approach: Makes "locks live with the leader of the data they protect" an org standard, aligned with the data-residency and home-region model. Funds the global ownership registry as part of the multi-region platform, not the lock team, and adds ownership transfer to the quarterly region-evacuation drill.
Staff Approach — Full Reasoning
| Option | Tradeoffs |
|---|---|
| Global etcd across regions | Correct; 70–150 ms per acquire and renew; minority-region partition stops all locked work there |
| Per-region clusters, home-region ownership | Fast and partition-tolerant; requires ownership model and transfer protocol |
| Per-region with "exactly once worldwide" via idempotency | Jobs may run in two regions, effects deduplicated globally; needs global idempotency store |
Metrics to Watch: ownership_transfer_duration_seconds, lock_acquire_latency_seconds{region}, cross_region_registry_write_latency.
Organizational Follow-up: Each team declares per job: regional or global; global requires review. Evacuation runbook includes ownership transfer.
Ownership Question: "Who owns the global ownership registry?" Staff answer: The multi-region platform team, because its correctness is a region-failover concern. The lock team consumes it.
Key Takeaway: "Don't make locks global. Make ownership regional and the transfer of ownership rare, fenced and rehearsed."
What clears the Staff bar:
- Quantifies global lock cost and partition behavior
- Substitutes ownership for coordination
- Ties lock placement to data placement
9. Level Expectations Summary#
After studying this case study, you should be able to:
- Split lock requests into efficiency, correctness and leadership intents and pick a store for each
- Walk the paused-holder timeline and explain why no lease protocol can prevent it
- Design fencing at the resource for SQL, object storage and unfenceable external APIs
- Explain the Kleppmann–antirez debate in two sentences and apply it to a scenario
- Set TTL, renewal cadence, self-abort deadline and max hold time with reasons
- Compute the 1/hold-time contention ceiling and propose lock-free alternatives
- Define failure posture per intent when the lock store loses quorum
- Draw the ownership line between the lock platform, resource teams and namespace owners
The Bar for This Question#
Mid-level (L4/E4): Knows SETNX and that locks need a TTL. Can write the Lua compare-and-delete release. Doesn't consider pauses, failover or what the lock protects.
Senior (L5/E5): Builds a clean lease-based lock on Redis or ZooKeeper, handles renewal and release correctly, knows Redlock exists and roughly why. Reaches fencing when prompted with "what if the holder pauses?" Treats the lock service as the owner of safety.
Staff+ (L6/E6+): Asks what overlap costs before choosing a store. States that leases cannot guarantee exclusivity and moves the guarantee to the resource with fencing, naming its owner. Separates fail-open and fail-closed by intent. Computes contention ceilings and proposes deleting locks where conditional writes suffice. Treats the lock service as a tier-0 shared product with quotas, an envelope and a contract. The interviewer should learn something from the answer.
10. Staff Insiders: Controversial Opinions#
10.1 "Most Distributed Locks Should Be Deleted"#
Audit any mid-sized codebase and the majority of distributed locks protect a single-record read-modify-write, a "run once" job, or a counter.
| Lock Usage | Lock-Free Replacement | Dependency Removed |
|---|---|---|
| Update one row safely | Version column / conditional write | Lock service on the hot path |
| Run a job once | Idempotency key + unique constraint | Lease hazards |
| Decrement inventory | Atomic conditional decrement | Contention ceiling |
| Serialize per-user events | Partition by user ID | Per-event acquires |
The Staff position: A lock is the right tool for long, multi-resource critical sections at low rates. Everywhere else it's a tier-0 dependency added to avoid a data-model change.
Why this matters in interviews: Proposing to remove the lock — and showing the conditional write — is the single strongest signal you can give.
10.2 "Redlock Is the Wrong Answer to Both Questions"#
| Intent | What It Needs | What Redlock Gives |
|---|---|---|
| Efficiency | Cheap, fast, usually right | 5 masters, majority round trips — more than needed |
| Correctness | Linearizable acquire + monotonic token | Timing-dependent safety, no token |
The Staff position: Use one Redis for efficiency. Use a consensus store plus fencing for correctness. Once you add fencing, the extra Redis masters buy nothing. You don't have to say antirez is wrong about typical clocks; you have to say your design shouldn't depend on typical.
Why this matters in interviews: It shows you can resolve a famous debate by reframing intent instead of picking a side.
10.3 "A Lock Without a Fencing Token Is a Hint"#
Every distributed lock is advisory to the resource. If the resource doesn't check, the lock guarantees what the lock service's clock believed — nothing about the resource.
| Resource | Can Fence? | Fallback |
|---|---|---|
| SQL table | Yes — predicate on fence column | — |
| Object store | Partially — conditional puts, token-suffixed paths | Write-then-promote |
| Kafka | Producer epochs (transactional producers fence zombies) | Idempotent consumers |
| External payment API | No | Idempotency keys |
| Filesystem via NFS | Rarely | Lock-delay + reconciliation |
The Staff position: Call an unfenced lock what it is in the design doc. Hints are fine for efficiency; for correctness, a hint must be paired with idempotency or reconciliation.
10.4 "Lease TTL Is a Business Decision"#
Engineers pick TTLs by gut feel — 30 s because it's round. But TTL is the guaranteed stall after a crash, and with fencing it's only that.
The Staff position: The namespace owner states the maximum tolerable stall; TTL follows. A payout batch can stall 60 s; a trading-session leader can't stall 5 s and should probably not be using a general-purpose lock service at all.
10.5 "The Lock Service Should Be Boring and Slow"#
Chubby's authors designed for coarse-grained locks held for hours and explicitly discouraged fine-grained use. Teams that optimize a lock service for 100K acquires/s invite exactly the workloads that should have been redesigned.
The Staff position: Publish a workload envelope (≤ 200 acquires/s per namespace, holds ≥ 100 ms, ≤ 10K keys). A lock service that's too fast to notice becomes a dependency of everything.
11. The Principal Lens (L7)#
Why L7 Sees This Problem Differently#
A Staff engineer designs a correct lock service. A Principal engineer notices the org already has five: a Redis lock library in the Python monorepo, Redlock in the storage team, PostgreSQL advisory locks in billing, Kubernetes leases for controllers, and a ZooKeeper recipe nobody remembers writing. Two guard money, none are fenced, and no one can list what holds what during an incident. The L7 question is "which of these locks protect invariants, which can be deleted, and what single contract do the survivors run on?" Locking at org scale is a correctness governance problem: classifying intents, making fenced writes the path of least resistance, and limiting the shared fate of a tier-0 dependency.
🧭 Principal Move: "Before we build anything, I'd inventory every lock that guards durable writes or money. I expect most can become conditional writes. The survivors move onto one consensus-backed substrate with fencing enforced by the storage platform, and the lock libraries get a deprecation date."
The Org-Level Fault Line#
Central lock platform vs locks-in-the-data-layer.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Every team embeds locks | Speed; no shared fate | Unfenced, inconsistent semantics; same incident repeats per team | Customers, finance, scattered on-call |
| Central lock service for everything | One contract, one runbook | Tier-0 shared fate; becomes a dumping ground for per-entity locking | Platform on-call; every consumer during outages |
| Data layer owns concurrency; lock service only for coarse cross-resource coordination (L7 default) | Most concurrency handled by CAS / partitions where data lives; lock service small and stable | Requires storage platform investment in fenced writes and CAS helpers | Storage platform budget; pays back in fewer correctness incidents |
The deciding question: how many correctness incidents in the last 2 years trace to locking? If any involved money, the investment is justified by one avoided incident.
Cost Model#
Assumptions: cloud pricing; etcd nodes on small compute with local NVMe at ~$250–500/month; loaded engineer ~$25K/month; costs are for the lock platform only.
| Scale | Architecture | Infra $/month | Headcount | On-Call Load |
|---|---|---|---|---|
| Startup — ~20 namespaces, 1 region | Use existing PostgreSQL lease rows or managed DynamoDB locks; Redis for efficiency | ~$0–$500 (incremental) | ~0.1 FTE | Shared with DB on-call; < 1 page/quarter |
| Growth — ~150 namespaces, 2 regions, ~2K acquires/s | Per-region 5-node etcd behind a thin lock API; SDK with fencing helpers | ~$3K–$6K | 1–1.5 FTE | Tier-0 rotation shared with coordination services; ~1 page/month |
| Enterprise — ~800 namespaces, 4+ regions, cells | Cell-local clusters (2 per region: platform-critical vs product), namespace registry, quotas, global ownership registry | ~$15K–$35K | 3–4 FTE (substrate, SDK, fenced-write helpers) | Own rotation; quarterly quorum-loss game day |
The line that matters to leadership: at every scale, the platform costs less than one duplicate-payout or split-brain incident. The bigger lever is deletion: each lock replaced by a conditional write removes a tier-0 dependency from a product's critical path for free.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Door | Reversal Cost | Why |
|---|---|---|---|
| Consensus substrate (etcd vs ZK vs managed) | Two-way if wrapped | 1–2 quarters | Hidden behind the lock API |
| Fencing token semantics (64-bit monotonic, per-namespace or global) | One-way | Quarters | Stored in every fenced resource's columns and checked by their code |
| Lock API contract (no data-plane force-unlock, token required on release) | One-way | Years | Every client SDK and runbook depends on it |
| Default TTL and renew ratios | Two-way | Hours | Policy per namespace |
| Global lock namespaces | One-way-ish | Painful | Teams build "exactly once worldwide" assumptions on them |
| Allowing per-entity high-rate use | One-way | Quarters to unwind | Workloads become dependent; removing them is a migration |
🧭 Principal Insight: The irreversible decisions are contracts, not technology. Spend review time on token semantics and the API, not on etcd vs ZooKeeper.
The Standard I'd Write#
RFC: Distributed Locking and Mutual Exclusion Standard (v1)
Scope: Any mechanism that grants one process exclusive rights over a shared resource across machines.
Requirements:
- New designs MUST document why a conditional write, single-writer partition or idempotency key is insufficient before adopting a lock.
- Every lock namespace MUST declare intent (efficiency / correctness / leadership), owner, maximum stall time and expected rate in the registry.
- Correctness and leadership locks MUST use the platform lock service and MUST pass the fencing token to every durable write; the resource MUST reject lower tokens atomically with the write.
- Correctness locks MUST fail closed on lock-service unavailability; efficiency locks SHOULD fail open.
- Locks MUST NOT span regions on a request path; cross-region exclusivity uses the ownership registry.
- Redlock and Redis-based locks MUST NOT be used for correctness.
- Breaking a lock MUST use the audited admin API, which bumps the token.
Exceptions: Architecture review within 5 business days; expiry ≤ 2 quarters.
Success metrics: Zero correctness incidents attributed to locking; 100% of correctness namespaces fenced within 3 quarters; lock-service P99 acquire < 10 ms; namespaces within workload envelope ≥ 95%.
What I'd Tell the VP#
"Several of our teams built their own locking, and at least two of those protect payments without the safeguard that stops a stalled server from writing twice. I'm proposing we remove the locks we don't need — most can be replaced with database features we already pay for — and move the rest onto one shared service with that safeguard built in. It costs about 1.5 engineers for three quarters and a few thousand dollars a month in infrastructure. The risk we're retiring is a duplicate-payment or data-corruption incident that we wouldn't detect for days. The main tradeoff is that the shared service becomes critical infrastructure, so we'll run one per region and rehearse its failure."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Inventories before building | "How many lock systems do we run, and which guard money?" |
| Moves concurrency into the data layer | "Most of these become version checks; the lock service shrinks to coarse coordination." |
| Names the contract | "We promise linearizable acquire and monotonic tokens — not exclusivity at your database." |
| Prices shared fate | "One cluster per cell: an etcd outage costs one cell 15 minutes, not the company." |
| Governs drift | "Intent is an owned attribute, re-attested yearly, because efficiency locks quietly become correctness locks." |
Staff answers that L7 interviewers find insufficient:
- "We'll use etcd with fencing tokens." — correct for one system; silent on the four other lock libraries and who migrates them.
- "Correctness locks fail closed." — right, but no tier-0 SLO, no cell boundary, no game day.
- "The resource checks the token." — which resources, owned by whom, and what's the paved-road helper so 40 teams don't each get it wrong?
Appendices
Appendix A: Mechanics in Depth#
A.1 Redis Single-Node Lock (Efficiency Only)#
# Acquire
SET lock:{ns}:{name} {owner_uuid} NX PX 15000 -> OK | nil
# Release — compare-and-delete, atomic via Lua
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("DEL", KEYS[1])
else
return 0
end
# Renew — compare-and-extend
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("PEXPIRE", KEYS[1], ARGV[2])
else
return 0
end
Right when: overlap is harmless. Wrong when: anything durable depends on exclusivity — async replication can lose the key on failover.
A.2 Redlock (For Recognition, Not Recommendation)#
N = 5 independent masters, quorum = 3
start = now()
for m in masters: SET key val NX PX ttl (short per-node timeout, ~5-50ms)
elapsed = now() - start
validity = ttl - elapsed - drift_allowance # drift ~1% of ttl + 2ms
acquired = successes >= 3 and validity > 0
if not acquired: release on all masters
Why it fails for correctness: validity is computed with clocks that can step; a pause after acquired and before the write is not covered; no token for the resource to check.
A.3 ZooKeeper Lock Recipe#
create /locks/{name}/lock- EPHEMERAL_SEQUENTIAL -> lock-0000000042
children = getChildren(/locks/{name}) # sorted
if mine is lowest: acquired; token = my sequence (or czxid)
else: watch the node immediately before mine; wait for its deletion; re-check
Session expiry deletes the ephemeral node. Watching only the predecessor gives O(1) wakeups per release instead of a herd.
A.4 etcd Lease + Transaction#
lease = LeaseGrant(ttl=15)
KeepAlive(lease) every 5s
txn = Txn()
.If(CreateRevision("/locks/ns/name") == 0)
.Then(Put("/locks/ns/name", owner, lease=lease))
.Else(Get("/locks/ns/name"))
resp = txn.Commit()
token = resp.Header.Revision # monotonic across the cluster
# waiters: put a key under /locks/ns/name/waiters/ with the lease,
# watch the key with the next-lower create_revision
The etcd concurrency package implements this pattern; its Mutex exposes the key's revision for fencing.
A.5 Lease Row in the Resource's Own Database#
-- Acquire / take over
UPDATE job_leases
SET owner = :me, fence = fence + 1, lease_until = now() + interval '15 seconds'
WHERE job = :job AND (owner IS NULL OR lease_until < now() OR owner = :me)
RETURNING fence;
-- Fenced work
UPDATE payouts SET status = 'SENT'
WHERE batch = :job AND item = :i
AND (SELECT fence FROM job_leases WHERE job = :job) = :my_fence;
One clock (the DB server), one transaction boundary, fencing for free. The best lock when the protected data lives in the same database. Note: lost on DB failover only if replication is async and the lease update was lost — the fence still prevents regression if the fence column replicates with the data.
A.6 DynamoDB Conditional-Write Lock#
PutItem(lock_table, {pk: name, owner: me, rvn: uuid(), version: v+1, lease_ms: 15000},
ConditionExpression = "attribute_not_exists(pk) OR rvn = :observed_stale_rvn")
Contenders observe the RVN for one full lease period before taking over, avoiding cross-machine clock comparison. Use version (incrementing) as the fencing token for downstream resources.
Appendix B: Lock Naming, Granularity and Data Model#
B.1 Key Construction#
/locks/{namespace}/{resource_type}/{resource_id}
e.g. /locks/payments/settlement-batch/2026-10-01-eu
- Namespace first: quotas, ACLs and metrics aggregate by prefix
- Never embed user-controlled strings without normalization — key cardinality explosions start here
B.2 Granularity#
| Granularity | Contention | Lock Count | Deadlock Risk |
|---|---|---|---|
| Coarse (per job / shard) | Higher per key, low overall rate | Small | Low |
| Fine (per entity) | Low per key | Huge — exceeds lock-service envelope | Higher when holding several |
| Hierarchical (intent locks) | Balanced | Medium | Requires strict ordering |
Multiple locks: acquire in a global canonical order (sorted keys) to avoid deadlock; bound total hold with a single session; prefer one coarser lock over three fine ones held together.
Deadlock in a leased world: leases convert deadlock into stall-then-timeout: two holders waiting on each other release after TTL. That's liveness by expiry, not by design — it shows up as lock_acquire_wait_seconds spikes and periodic expiries. Canonical ordering removes it.
Appendix C: Coordination Mechanisms — Quick Comparison#
| Mechanism | Linearizable | Fencing Token | Acquire Latency | Ops Burden | Best For |
|---|---|---|---|---|---|
| Redis single node | No (failover) | No | ~0.5 ms | Low | Efficiency |
| Redlock | Timing-dependent | No | ~2–5 ms | Medium (5 masters) | Nothing you can't do better |
| ZooKeeper | Yes | zxid / sequence | ~2–10 ms | High | Existing ZK shops, leadership |
| etcd | Yes | Revision | ~2–10 ms | Medium–High | Correctness, K8s-adjacent orgs |
| DynamoDB conditional | Yes (per item) | Version attribute | ~5–10 ms | None | AWS-native, coarse locks |
| DB lease row | Yes (per DB) | Fence column | ~1–5 ms | None extra | Data in the same DB |
| Chubby-style service | Yes | Sequencer | ~ms | Very high (build) | Hyperscalers |
See ZooKeeper & etcd for operational detail and Distributed Coordination for leader election and membership patterns.
Appendix D: API Contract and Client Behavior#
D.1 Error Semantics#
| Response | Meaning | Client Action |
|---|---|---|
LOCK_HELD | Another holder | Back off or wait via watch |
TIMEOUT | Waited wait_timeout | Return to caller; do not retry tightly |
SESSION_EXPIRED | Lease gone | Stop all work under this session immediately |
UNAVAILABLE | Store unreachable | Apply namespace posture: fail-closed or fail-open |
QUOTA_EXCEEDED | Namespace over limit | Page namespace owner; do not retry |
STALE_TOKEN (from resource) | Fenced out | Abort, log, do not retry with same token |
D.2 SDK Responsibilities#
acquire(ns, name):
resp = api.Acquire(ns, name, request_id=uuid())
start renew loop at ttl/3, gated on worker.progress_heartbeat < ttl
on renew failure for > 2/3 ttl: cancel worker context, emit lease_lost
return Lock(token=resp.token, ctx=cancellable_context)
Retries and herds:
- Acquire retries: exponential backoff 50 ms → 2 s with full jitter
- Waiting: use server-side wait queues (predecessor watch), not client polling
- Post-outage resume: random 0–5 s delay before first acquire
Appendix E: Observability#
E.1 Core Metrics#
# Service health
lock_acquire_latency_seconds{ns} P50/P99
lock_acquire_errors_total{ns,code}
lock_renew_failures_total{ns}
etcd_server_has_leader, etcd_disk_wal_fsync_duration_seconds
# Safety
lock_lease_expired_while_held_total{ns} # SDK: lost lease while working
lock_fencing_rejections_total{resource} # emitted by resource owners
lock_holder_overlap_seconds{ns} # derived from token timeline
# Contention and hygiene
lock_waiters{ns,name}, lock_acquire_wait_seconds{ns}
lock_hold_time_seconds{ns}, lock_age_seconds_max{ns}
lock_namespace_keys{ns}, lock_admin_break_total{ns}
E.2 Critical Alerts#
| Alert | Threshold | Severity |
|---|---|---|
| No leader | etcd_server_has_leader == 0 for 30 s | Page platform |
| Acquire errors | > 1% for 2 min on correctness namespaces | Page platform |
| Lease lost while held | > 0 on correctness namespace | Ticket owner; page if repeating |
| Fencing rejections | > 0 | Page resource owner — a zombie just tried to write |
| Hold time | > max_hold | Page namespace owner |
| Namespace keys | > 80% of quota | Ticket namespace owner |
E.3 Debugging "Two Holders" Reports#
- Pull the token timeline for the lock: acquire, renew, expiry events with server timestamps
- Correlate with GC/pause metrics and node events for the old holder
- Check the resource's fencing rejections — if zero and both wrote, the resource isn't fencing
- If tokens never overlapped, the report is a client bug (not using the lock or the token)
Appendix F: Scale Evolution#
| Scale | Approach | Don't Build Yet |
|---|---|---|
| < 10 locks, one DB | Lease rows in the database | Any lock service |
| Tens of namespaces | Managed or existing etcd/ZK + thin SDK | Reader/writer locks, fair queuing |
| Hundreds of namespaces | Lock API, registry, quotas, per-region clusters | Global namespaces, a lock-breaking UI |
| Thousands, multi-region | Cell-local clusters, ownership registry (see Multi-Region), workload envelope | A custom consensus implementation — ever |
Appendix G: Multi-Tenancy, Fairness and Cost#
One namespace can exhaust etcd's write capacity or storage quota for all. Enforce at the frontend: keys per namespace, acquires/s per namespace, max waiters per lock.
| Tier | Who | Isolation |
|---|---|---|
| Platform-critical | Controllers, shard ownership, payout leadership | Dedicated cluster per cell |
| Product correctness | Batch jobs, exports | Shared cluster, strict quotas |
| Efficiency | Cache warmers, crawlers | Redis, separate entirely |
Chargeback: charge by namespace on acquires + key-hours. The goal isn't revenue; it's making a 6M-key namespace show up on someone's budget before it shows up in an incident.