Hiring BarSupport

Clocks, Ordering & Time

Foundation36 min read6 diagrams

Why This Matters#

Time is not a timestamp question. Every candidate can call now(). The question hiding underneath is: when two machines disagree about what happened first, whose answer wins, and what breaks when it's the wrong one? A clock that is 80ms fast is harmless on a log line and catastrophic on a last-write-wins register, a lease, or a "this coupon expires at midnight" check that runs on 400 hosts.

The textbook answer — "use NTP, it keeps clocks in sync" — is correct and incomplete. NTP over the public internet keeps hosts within tens of milliseconds most of the time. It says nothing about the host whose daemon silently stopped syncing three days ago and is now 1.3 seconds behind, the leap second that made a wall clock step backwards, the VM that was paused for 4 seconds during live migration, or the garbage-collection pause that let a lock holder keep writing 20 seconds after its lease expired. Clocks fail quietly, and the failures look like data corruption, not like clock errors.

Staff engineers treat time as three different tools with three different contracts: a monotonic clock for measuring durations on one machine, a logical clock for ordering events across machines, and a wall clock with a stated error bound for anything a human or a contract will read. They never use one where another belongs. They also know the single most important rule in this space: a lease or lock based on time is a performance optimization, never a correctness guarantee — correctness comes from a fencing token that the storage layer checks.

This page covers the clock types, the ordering mechanisms (Lamport, vector, hybrid logical clocks, bounded-uncertainty time), and the lease-and-fencing problem that ties them together. Consistency guarantees built on top of these live in Consistency, CAP & PACELC; a full lock-service design lives in Distributed Lock Service.

The 60-Second Version#

  • Never measure a duration with the wall clock. Wall clocks step forward and backward (NTP corrections, leap seconds, manual changes). Use the monotonic clock (CLOCK_MONOTONIC, System.nanoTime, Go's monotonic reading) for timeouts, latency and rate windows. A negative duration is a bug waiting for a leap second.
  • Clock skew is real and unbounded unless you bound it. Public-internet NTP: ~1–50ms typical, with outliers in the hundreds. A well-run datacenter with local stratum-1 sources: ~0.1–1ms. A host whose sync daemon died: drifts ~10–200ppm, i.e. 1–17 seconds per day.
  • Last-write-wins with wall-clock timestamps silently loses data. If two writes to the same key land within the skew window (say 30ms), the "later" write is wrong roughly half the time. Fine for a profile photo, not for a balance.
  • Lamport clocks give a total order consistent with causality in 8 bytes per message — but they can't tell you whether two events were concurrent.
  • Vector clocks detect concurrency (so you can keep both siblings and merge), at O(N) size per value where N = number of writers. Past ~10–50 writers they need pruning or a different design.
  • Hybrid logical clocks (HLC) pack wall time and a logical counter into one 64-bit timestamp: close to real time, monotonic, causality-preserving, and tolerant of bounded skew. The default for modern distributed SQL.
  • Bounded-uncertainty time (TrueTime-style) returns an interval [earliest, latest]. Wait out the interval width (~1–7ms in Spanner's published numbers) and you get external consistency without a central sequencer.
  • Leases need fencing tokens. A process paused for 30 seconds still believes it holds the lock. Only a monotonically increasing token, checked by the storage layer on every write, makes the stale holder harmless.

How Clocks Work (for System Designers)#

The Three Clocks on Every Machine#

ClockWhat It AnswersCan Go Backwards?Comparable Across Machines?Use It For
Wall clock (CLOCK_REALTIME, System.currentTimeMillis)"What time is it in UTC?"Yes — NTP steps, leap seconds, operator changesRoughly, within skew (ms to seconds)Timestamps humans read, expiry dates, audit logs, TTLs measured in minutes or more
Monotonic clock (CLOCK_MONOTONIC, System.nanoTime)"How much time has elapsed since some arbitrary point?"No (may be slewed, never stepped back)No — the origin is per bootTimeouts, latency measurement, rate-limit windows, retry backoff, lease countdowns on the holder
Logical clock (Lamport, vector, HLC)"Did A happen before B?"No, by constructionYes — that's the pointOrdering writes, conflict detection, snapshot reads, replication

🎯 Staff Move: "I'll keep three clocks separate. Timeouts and latency use the monotonic clock. Ordering of writes uses a hybrid logical clock stamped by the database, not the app server. Wall time only appears where a human reads it or a contract defines it — and there I'll state the error bound."

How Wall Clocks Drift and Get Corrected#

A server's clock is a quartz oscillator counted by the kernel. Quartz drifts with temperature and age — typically 10–100 parts per million, with a conservative worst case near 200ppm. At 50ppm a free-running clock gains or loses ~4.3 seconds a day.

Synchronization daemons (ntpd, chrony, or a cloud provider's time agent) poll reference servers and correct the clock in one of two ways:

CorrectionHowSide Effect
SlewSpeed the clock up or slow it down slightly (up to ~500ppm) until it convergesDurations measured by the wall clock are slightly wrong for minutes; the clock never jumps
StepJump the clock to the right valueWall time can move backwards. Anything computing end - start with wall time sees a negative or huge number

Leap seconds add a third hazard: a 61-second minute that some kernels implement as a repeated second (time goes backwards by 1s) and others smear over hours. The rule that survives all three: wall time is a label, not a ruler.

Where Skew Comes From#

SourceTypical ContributionNotes
Network asymmetry to the time server0.1–10msNTP assumes symmetric paths; it can't detect asymmetry
Time-server qualityµs (GPS/atomic, local) to tens of ms (public pool)Stratum matters less than path length
Poll interval × drift rate64s × 50ppm ≈ 3ms; 1,024s × 50ppm ≈ 51msLonger poll intervals mean bigger sawtooth error
VM pause / live migration0.1–5s on the guestThe guest wakes up behind until it resyncs
Daemon stopped or misconfiguredUnbounded: seconds per day, foreverThe most common real-world cause of large skew

Key Terms#

TermMeaning
Skew (offset)The difference between two clocks at the same instant
DriftThe rate at which a clock gains or loses time, in ppm
Happens-before (→)A → B if A and B are on the same process and A came first, or A is a send and B the matching receive, or transitively through either
Concurrent (∥)Neither A → B nor B → A. Concurrent events have no "correct" order
Lamport timestampA counter that guarantees A → B implies L(A) < L(B) — but not the converse
Vector clockOne counter per node; compares element-wise to decide before / after / concurrent
HLCHybrid logical clock: (physical_ms, logical) pair that tracks the max wall time seen plus a counter
Uncertainty interval[now − ε, now + ε]: the window within which true time is known to lie
External consistencyIf T1 commits before T2 starts in real time, T1's timestamp is lower than T2's
LeaseTime-bounded ownership: "you own X until t"
Fencing tokenA monotonically increasing number issued with each lease; the resource rejects requests carrying an older token

Where Time Lives in a Design#

  • Write ordering — last-write-wins registers (Cassandra-style cell timestamps), MVCC snapshot timestamps, replication logs.
  • Leases and leader election — lock services, partition leaders, primary-replica failover, scheduler ownership.
  • Expiry — cache TTLs, session and token lifetimes, rate-limit windows, idempotency-key retention.
  • IDs — time-prefixed IDs (Snowflake-style, ULID, UUIDv7) that sort roughly by creation time. See ID Generation.
  • Business deadlines — "sale ends at 12:00:00," auction close, option expiry, SLA timers.
  • Observability — trace spans and log lines across hosts; skew makes child spans appear to start before their parents.

Core Strategies#

Strategy 1: Wall-Clock Timestamps with Last-Write-Wins#

write(key, value):
    ts = wall_clock_micros()               # stamped by the client or coordinator
    for replica in replicas(key):
        replica.put(key, value, ts)

replica.put(key, value, ts):
    if ts > stored[key].ts:                # higher timestamp wins
        stored[key] = (value, ts)
    # else: silently discard — the "older" write loses

When to use: Data where losing one of two near-simultaneous writes is acceptable: user preferences, presence status, "last seen," cache entries, sensor readings where the newest sample is all that matters.

Failure mode: If node A's clock is 200ms fast, every write it stamps beats every write from correctly synced nodes for the next 200ms — even writes that happened after it in real time. A delete stamped by a slow clock can be ignored forever because the row it targets carries a higher timestamp. No error is raised; the data is just wrong. Monitor clock.offset_ms per host and reject writes from any host whose offset exceeds a bound (e.g., 50ms).

Strategy 2: Lamport Clocks#

on local event:        L = L + 1
on send(msg):          L = L + 1;  msg.ts = L
on receive(msg):       L = max(L, msg.ts) + 1

total_order(a, b):     compare (a.ts, a.node_id) < (b.ts, b.node_id)   # ties broken by node id

When to use: You need a single total order that respects causality and you don't care which of two concurrent events goes first — ordering commands in a replicated log, deterministic tie-breaking, distributed mutual exclusion algorithms.

Failure mode: L(A) < L(B) does not mean A happened before B. Two independent writes get an arbitrary order, so Lamport clocks can't detect a conflict — they quietly pick a winner, which is last-write-wins with better causality hygiene. They also carry no relation to wall time, so "show me the state as of 14:02" is impossible.

Strategy 3: Vector Clocks (and Version Vectors)#

# One counter per writer. VC is a map node_id -> counter.
on write at node i:    VC[i] = VC[i] + 1
merge(a, b):           for each n: out[n] = max(a[n], b[n])

compare(a, b):
    if all a[n] <= b[n] and a != b:   return BEFORE
    if all a[n] >= b[n] and a != b:   return AFTER
    if a == b:                        return EQUAL
    return CONCURRENT                 # keep both as siblings; app or CRDT merges

When to use: Multi-writer data where silently dropping a concurrent write is unacceptable and the application can merge — shopping carts, collaborative state, offline-first sync. Version vectors (one entry per replica, not per client) keep the size bounded by replica count.

Failure mode: Size grows with the number of distinct writers. With per-client entries, a popular key touched by thousands of clients carries thousands of counters, so systems truncate the oldest entries — and truncation can turn "before" into "concurrent," creating false siblings. The deeper cost is organizational: every reader must handle siblings. If the product team won't write merge logic, vector clocks just move the conflict to a place nobody owns.

Strategy 4: Hybrid Logical Clocks (HLC)#

# State: (pt, c) — pt = highest physical time seen (ms), c = logical counter
now():
    wall = physical_clock_ms()
    if wall > pt:  pt, c = wall, 0
    else:          c = c + 1                # wall clock behind or equal: bump counter
    return (pt, c)

on receive(msg_pt, msg_c):
    wall = physical_clock_ms()
    new_pt = max(pt, msg_pt, wall)
    if new_pt == pt == msg_pt: c = max(c, msg_c) + 1
    elif new_pt == pt:         c = c + 1
    elif new_pt == msg_pt:     c = msg_c + 1
    else:                      c = 0
    pt = new_pt
    if pt - wall > MAX_OFFSET_MS: crash_or_alert()     # our view drifted too far from local wall time
    return (pt, c)

# Encoding: 48 bits of milliseconds + 16 bits of counter fit one uint64.

When to use: The default for multi-node databases and replication logs that need causality, monotonicity, and timestamps that mean something to humans ("read as of 14:02:31"). Distributed SQL engines use HLCs to stamp MVCC versions without a central timestamp oracle.

Failure mode: HLC preserves causality but does not by itself give real-time ordering between unrelated transactions. If node A's clock is 300ms ahead, a transaction on A can get a timestamp higher than a later, causally unrelated transaction on B. Systems handle this with a configured maximum clock offset: reads treat any value stamped within (read_ts, read_ts + max_offset] as uncertain and restart above it, and nodes that drift past the bound remove themselves. Set the bound too low and healthy nodes crash on an NTP hiccup; too high and uncertainty restarts spike on hot keys.

Strategy 5: Bounded-Uncertainty Time and Commit Wait#

TT.now()  -> [earliest, latest]          # true time is guaranteed inside the interval
epsilon   = (latest - earliest) / 2       # ~1–7ms with GPS + atomic references per DC

commit(txn):
    s = TT.now().latest                   # pick a timestamp no earlier than true time
    replicate_via_paxos(txn, s)
    while TT.now().earliest <= s:          # commit wait: ~2 × epsilon
        sleep()
    release_locks_and_ack(txn)            # any later txn anywhere gets a higher timestamp

When to use: Globally distributed transactional stores that must be externally consistent (strict serializable) across regions without routing every commit through one sequencer. It also makes lock-free snapshot reads at a timestamp safe on any replica.

Failure mode: Correctness now rests on the interval being honest. Commit latency rises with ε, so the time infrastructure (GPS receivers, atomic clocks, a fleet of time masters, liar detection) becomes a tier-0 dependency. If ε widens to 100ms during a time-master outage, every commit waits ~200ms. If ε lies — the interval is narrower than the true error — external consistency silently breaks.

🎯 Staff Move: "I don't need global real-time ordering for this feed, so I won't pay for it. Per-key ordering comes from the partition leader's log sequence number, causality across services comes from an HLC in the request context, and wall time is display-only. If this were a ledger spanning regions, I'd reach for a store with bounded-uncertainty timestamps and accept the commit-wait cost."

Leases, Locks and Fencing: The Hard Sub-Problem#

Ordering events is the easy half. The hard half is time-bounded ownership: "node A is the leader until 12:00:10," "worker 7 owns job 991 for 30 seconds," "this client holds the lock." Every one of these quietly assumes that the holder's sense of elapsed time matches the grantor's. That assumption fails in four routine ways:

What HappensHow LongEffect on a Lease Holder
Stop-the-world GC pause50ms–30s+ on large heapsHolder resumes and writes as if nothing happened
VM paused / live-migrated0.1–5sSame, plus the wall clock is now behind
Process swapped out / CPU throttled (container limits)Seconds under memory or CPU pressureHeartbeats stop; lease expires; holder doesn't know
Network partition between holder and lock serviceUntil the partition healsHolder can still reach the database; lock service already granted the lease elsewhere

The Unfenced Lease#

worker_A: acquire lease (30s)            -> ok, at grantor time t=0
worker_A: begins writing batch...
worker_A: [GC pause 35s]
grantor:  lease expires at t=30, grants to worker_B
worker_B: writes batch v2 to storage at t=31
worker_A: wakes at t=35, still "holds" the lease in its own mind
worker_A: writes stale batch v1 over v2  -> corruption, no error anywhere

Checking "is my lease still valid?" before each write doesn't help: the pause can happen between the check and the write. Any scheme where the holder decides it is still the owner is unsafe under pauses.

The Fix: Fencing Tokens Checked by the Resource#

lock_service.acquire(resource) -> (lease, token)    # token strictly increases per grant: 33, 34, 35...

storage.write(resource, data, token):
    if token < resource.highest_token_seen:
        reject("stale token")                         # A's token 33 loses to B's 34
    resource.highest_token_seen = token
    apply(data)

The token can be a ZooKeeper zxid or znode version, an etcd revision, a database sequence, or a Raft term. What matters is that the resource being protected enforces it — a lock service alone can never stop a paused holder.

ApproachSafe Under Pauses?RequirementTypical Use
Lease, no fencingNoNothingEfficiency only: avoid duplicate work where duplicates are harmless
Lease + holder self-check before writeNoMonotonic clock on holderReduces, doesn't eliminate, the window
Lease + fencing token checked by storageYesStorage must support conditional writes on a tokenLeader election, job ownership, single-writer partitions
Consensus-backed writes (every write goes through the log)YesAll writes go through Raft/PaxosMetadata stores, configuration

Lease Timing Rules That Survive Drift#

  1. The holder counts down with its monotonic clock, starting from before it sent the request. If acquisition took 200ms, the holder believes it has lease − 200ms − drift margin, never the full lease.
  2. The grantor waits out the full lease plus a skew margin before re-granting. For a 10s lease and 200ppm worst-case drift, the margin is ~2ms; add the max expected pause you're willing to tolerate unfenced, or just fence.
  3. Renew at one-third of the lease. A 10s lease renewed every ~3s survives two missed renewals.
  4. Shorter is not safer, only faster. A 2s lease fails over faster but flaps more under GC pauses; 10–30s is a common range for leader leases.

🎯 Staff Move: "The lease decides who should be working. The fencing token decides whose writes count. I'll pass the token through to the storage layer and make every write conditional on it, so a paused ex-leader gets a rejection instead of corrupting data."

When NOT to Trust a Timestamp#

  • Ordering financial events across hosts. Use a sequence from a single writer or a consensus log, not app-server wall time.
  • Expiry checks with sub-second precision across hosts. "Coupon valid until 12:00:00.000" evaluated on 400 hosts with 50ms skew means some users get it at 12:00:00.040 and some lose it at 11:59:59.960. Decide in one place, or publish the tolerance.
  • Comparing timestamps from client devices. Phone clocks are off by minutes or set deliberately. Record server receipt time alongside client time and trust the server's for ordering.
  • Deriving causality from trace timestamps. Use the span parent-child relationship; skew routinely makes children appear to start before their parents.

Visual Guide#

Happens-Before, Concurrency and Lamport Timestamps#

Diagram: Happens-Before, Concurrency and Lamport Timestamps

Choosing an Ordering Mechanism#

Diagram: Choosing an Ordering Mechanism

Commit Wait with Bounded Uncertainty#

Diagram: Commit Wait with Bounded Uncertainty

The Fenced Lease#

Diagram: The Fenced Lease

Clock State on a Host#

Diagram: Clock State on a Host

Implementation Patterns#

Stamp at the Store, Not the Client#

Let the component that serializes writes assign the order: the partition leader's log sequence number, the database's HLC, the broker's offset. App servers number in the hundreds and their clocks are managed by many teams; database nodes number in the tens and are managed by one. Fewer clocks, one owner, smaller skew.

Carry Causality in the Request Context#

When service B must observe the effect of a write from service A, pass A's commit timestamp (or log position) in the request — a "causal token." B reads with min_timestamp = token, waiting or redirecting to a replica that has caught up. This is how read-your-writes survives a hop through a queue or a second service, without global ordering.

# Write path returns its position
resp = orders.create(...)            # resp.commit_ts = (1733412000123, 4)

# Downstream read honors it
inventory.get(sku, min_ts = resp.commit_ts)   # replica waits up to 50ms or forwards to leader

Time-Ordered IDs Without Trusting Time for Correctness#

Snowflake-style IDs (41 bits of milliseconds, 10 bits of worker, 12 bits of sequence ≈ 4,096 IDs per ms per worker) and UUIDv7 sort approximately by creation time — useful for index locality and pagination. Two rules: refuse to issue IDs if the clock goes backwards (wait it out, or keep issuing from the last-seen millisecond), and never use ID order as proof of causal order across workers. Skew between workers means ID order is only "roughly" time order.

Expiry with Tolerances#

For TTLs and token lifetimes, accept a skew allowance explicitly: a JWT verifier with leeway = 30–60s for exp and nbf avoids rejecting fresh tokens from an issuer whose clock is slightly ahead. Write the allowance down; an undocumented 5-minute leeway is a security finding.

Monitor Clocks Like You Monitor Disks#

MetricAlert Threshold (example)Why
clock.offset_ms (from chrony/ntpd)> 10ms for 5 min (DB nodes); > 100ms (app hosts)Catches drift before it reaches the database's max-offset bound
clock.sync_stateNot synced for > 10 minA stopped daemon is the most common cause of large skew
clock.max_error_bound_ms> 50% of configured max offsetEarly warning before nodes self-fence
db.uncertainty_restarts_per_sec> 2× baselineSkew or hot keys are inflating HLC uncertainty
lease.rejected_stale_tokenAny sustained rateA paused or partitioned holder exists — working as designed, but investigate why

Failure Scenario: The Fast Clock That Ate the Deletes#

t=0       A config-management change disables the time daemon on 1 of 40 app hosts (host-17).
t=+3d     host-17 drifts to +1.8s ahead. Nothing alerts: no one monitors app-host offset.
t=+3d     host-17 stamps writes to a last-write-wins store (client-supplied timestamps).
t=+3d     Users "delete" saved items via other hosts. Deletes carry timestamps 1.8s lower
          than the rows host-17 wrote seconds earlier. Tombstones lose. Items reappear.
t=+4d     Support tickets: "deleted items keep coming back." 0.3% of active users affected.
t=+4d+2h  Engineer finds rows whose write timestamp is in the future relative to ingest time.
t=+4d+3h  host-17 drained; daemon re-enabled. Rows already written with future timestamps
          keep winning until a cleanup job rewrites them.

Detection: clock.offset_ms per host (absent here); write.ts_minus_ingest_ts_ms histogram on the store — future-dated writes are a direct signal; spike in "deleted item reappeared" tickets. Blast radius: every key written by host-17 for a day — invisible until a later write disagreed. Mitigation: drain the host; reject writes whose timestamp exceeds the coordinator's clock by more than a bound (e.g., 500ms); one-off job to re-stamp future-dated cells. Prevention: server-side timestamps assigned by the store; offset alerting on every host; config-management test that the time daemon is running. Owner: infrastructure/compute owns host time sync and its alerting; the storage team owns the "reject future timestamps" guard; the product team owns the user-facing cleanup.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Host clock drifts (daemon dead)clock.offset_ms, clock.sync_stateEvery LWW write from that hostDrain host; reject future-dated writesCompute / infra
Clock steps backwards (leap second, step correction)Negative durations; timer.negative_elapsed counterTimeouts, rate limiters, latency metrics, ID generatorsMonotonic clocks for durations; leap smearingEvery service team; platform library owners
Paused lease holder writes after expirylease.rejected_stale_token; duplicate job runsData guarded by that leaseFencing tokens enforced at storageOwning service + storage team
Skew exceeds DB max offsetNode self-terminates; db.clock_offset_violationCapacity loss; possible quorum loss if many nodesFix time source; never "fix" by raising max offset in an incidentDatabase platform
Time-master / reference outageclock.max_error_bound_ms rising fleet-wideCommit latency (commit wait) or uncertainty restarts everywhereRedundant references across DCs; alert on bound growthTime infrastructure owner
Trace spans out of orderChildren start before parents in UIDebugging confusion onlyOrder by span relationships, not timestampsObservability team

The Numbers in Context#

NumberValueWhat It Means for Your Design
Quartz drift, typical10–100ppm (~1–9s/day)A host that stops syncing is seconds off within a day. Alert on sync state, not just offset.
Conservative worst-case drift~200ppm (0.2ms per second)Use this to size lease margins: a 10s lease needs ~2ms of drift allowance.
Public-internet NTP offset~1–50ms typical, 100ms+ outliersNever order cross-host writes by app-server wall time.
Datacenter time with local references~0.1–1msGood enough for HLC max-offset bounds in the low milliseconds to hundreds of ms.
TrueTime ε (published)~1–7ms sawtooth, ~4ms typicalCommit wait costs single-digit milliseconds when the time infrastructure is healthy.
Leap second1s step (or smeared over hours)Wall-clock durations can go negative exactly once in a while — that's when they break.
GC pause, large JVM heap50ms typical, 10–30s+ worst caseLonger than many leases. Fence, don't trust.
VM live-migration pause~0.1–5sGuest clock lands behind; leases held by the guest are suspect.
Lamport timestamp size8 bytesCheap enough for every message.
Vector clock size8–16 bytes × writers1,000 writers per key ≈ 8–16KB per value — prune or redesign.
HLC encoding48-bit ms + 16-bit counter in 64 bits65,536 events per ms per node before the counter overflows.
JWT clock leeway30–60s commonDocument it; it extends every token's effective lifetime.
Leader lease10–30s common, renew at ~1/3Shorter fails over faster and flaps more under pauses.
Snowflake-style ID41-bit ms, 10-bit worker, 12-bit seq~4,096 IDs per ms per worker; ~69 years of milliseconds.

How This Shows Up in Interviews#

Scenario 1: "Two users edit the same document field at the same time. Who wins?"#

Don't say "the later timestamp." Say: "Within a single region, the field's owning partition serializes writes, so the leader's log order decides — no clocks involved. If edits can come from two regions or offline clients, wall-clock last-write-wins will drop one edit whenever they land within the skew window. For a title field I'd accept that and surface 'edited by X' history. For the document body, I'd keep concurrent versions with version vectors or use a CRDT so both edits survive." Then link it to Collaborative Documents.

Scenario 2: "Design leader election for a job scheduler so a job never runs twice." (Full Walkthrough)#

Step 1 — Reframe the guarantee. "'Never runs twice' is impossible with timeouts alone — a paused leader can't tell it was deposed. What I can guarantee is that a stale leader's side effects are rejected. So the design is a lease for liveness plus fencing for safety."

Step 2 — Pick the lease source. "I'll use etcd: the scheduler takes a lease with a 15s TTL and renews every 5s. The etcd revision at acquisition is the fencing token — it strictly increases across leaders." Link etcd & ZooKeeper.

Step 3 — Count time correctly on the leader. "The leader measures its lease with the monotonic clock, starting from when it sent the request. It stops dispatching at 12s, leaving 3s of margin for drift and for in-flight work to drain."

Step 4 — Fence the side effects. "Every job-state write is conditional: UPDATE jobs SET state='running', leader_token=$t WHERE id=$j AND leader_token < $t. Downstream effects carry an idempotency key of (job_id, run_id), so a duplicate dispatch to a worker is deduplicated even if fencing is bypassed."

Step 5 — Size failover. "Failover time is lease TTL plus election: ~15–20s worst case. If product needs faster, I'd drop the TTL to 5s — and I'd show them the flap rate under GC pauses first."

Step 6 — Name owners and alerts. "Scheduler team owns lease.rejected_stale_token and scheduler.leader_changes_per_hour (alert above 6). Infra owns host clock offset. Any job type that can't be made idempotent needs sign-off from its owning team that a rare duplicate is acceptable."

Why this is a Staff answer: It refuses the impossible guarantee, replaces it with one that can be enforced, puts enforcement in the storage layer, quantifies failover, and names who accepts the residual risk. See Job Scheduler.

Scenario 3: "Our flash sale must end at exactly 12:00:00 for everyone."#

This tests whether you'll decide time in one place. "Exactly is only meaningful at one decision point. I'd make the inventory service the authority: the sale window is evaluated against its clock, on the write path, when the order reserves stock. Clients show a countdown synced to a server timestamp, but the client clock never decides eligibility. Across 400 app hosts with ±50ms skew, evaluating at the edge would give a 100ms window where different users get different answers." See Ticket Drops & Flash Sales.

Scenario 4: "Why can't we just use timestamps to get a consistent snapshot across shards?"#

"You can, if you bound the skew and pay for it. With HLC timestamps and a max-offset of 250ms, a snapshot read at time t must treat any value stamped in (t, t+250ms] as uncertain and either wait or restart. With bounded-uncertainty clocks at ~4ms, you wait ~4ms instead. Without either, a 'snapshot' at t can include a write from a fast-clocked shard that really happened after a write it excludes. The cost of the timestamp approach is the size of your clock error."

Advanced Patterns#

PatternHow It WorksWhen to Use
Leap smearingSpread the leap second over ~24h by slewing all clocks identicallyFleets that can't tolerate a 1s step; must be uniform across every host you compare
Dotted version vectorsVersion vector plus a single "dot" for the latest event to avoid sibling explosionMulti-writer KV stores with many clients per key
Causal tokens / session timestampsReturn commit position to client; subsequent reads demand at least that positionRead-your-writes across replicas and service hops
Timestamp oracleOne service hands out strictly increasing timestamps (batched, ~1M/s)Single-region transactional stores; simple but a single point of latency
Read leases / follower reads at a past timestampServe reads at now − max_offset from any replica without coordinationLow-latency stale-but-consistent reads in multi-region setups
Clock error bound APIsTime daemon exposes [earliest, latest] to applications via shared memoryApps that want TrueTime-style reasoning on commodity cloud hosts
Self-fencing nodesNode stops serving when its clock offset exceeds the boundAny store whose correctness assumes a max offset

Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

A Staff engineer picks the right ordering mechanism for one system and fences its leases. A Principal engineer notices that the company's correctness quietly depends on clocks that nobody owns. The database team assumes offsets under 250ms. The auth team assumes 60s of JWT leeway is enough. The billing pipeline orders events by app-server wall time. Five lock implementations exist, two of them unfenced. Time sync is "configured by the base image" — which means the team that owns the base image owns the correctness of every last-write-wins store in the company without knowing it. At L7, time stops being a library call and becomes infrastructure with an SLO: a published error bound, monitoring on every host, and a short list of approved ordering and locking primitives.

🧭 Principal Move: "Every team is making an assumption about clock skew, and none of those assumptions are written down. I'd publish a time SLO — say, 99.99% of host-minutes within 10ms on database tiers and 100ms elsewhere — alert on it, and require any design that depends on wall-clock ordering to cite it."

The Org-Level Fault Line#

Shared time and coordination platform vs every team rolling its own clocks and locks.

OptionWhat WorksWhat BreaksWho Pays
Each team chooses (Redis locks, DB advisory locks, wall-clock LWW)Fast, local decisionsUnfenced locks; LWW on app-server time; skew assumptions nobody verifiesProduct on-call during subtle data-loss incidents; customers whose data silently reverts
Shared coordination library (fenced leases over etcd/ZooKeeper, HLC in request context)One correct implementation; consistent tokensLibrary adoption is slow; storage layers still must enforce tokensPlatform team maintaining language ports
Time as a platform with an SLO (redundant references, offset monitoring, error-bound API)Skew becomes a measured, alerted quantityNew tier-0 dependency; needs ownership and budgetInfra team; every database team depends on it

The Principal position: fund time infrastructure and offset monitoring centrally (it's cheap), ship one fenced-lease primitive, and ban client-supplied timestamps for ordering in shared stores. Leave CRDT and merge semantics to product teams — that's domain logic, not platform.

Cost Model#

Assumptions: engineer ~$25K/month fully loaded; commodity GPS/PTP time appliance ~$5–15K each; managed cloud time services are typically free; one data-loss incident costs ~2–4 engineer-weeks plus customer impact.

ScaleSetupMonthly CostHeadcountMain Risk Avoided
Small (1 region, ~100 hosts)Cloud provider time service, chrony everywhere, offset alert~$0 infra + a few hours0 dedicatedSilent LWW corruption from one drifting host
Medium (3 regions, ~5K hosts, 10 stateful systems)Above + fleet-wide offset dashboards, shared fenced-lease library~$5K/month in tooling + 0.5 engineer ($12K)0.5–1 on infra platformUnfenced leader elections; duplicate job runs; billing order bugs
Large (10+ regions, ~100K hosts, globally consistent store)Redundant GPS/atomic references per region, error-bound API, time SLO with paging$20–40K/month hardware amortized + 2–3 engineers ($60K)2–3 dedicatedCommit-latency spikes and correctness violations from ε excursions

At small scale the right spend is nearly zero — just monitoring. The money only appears when a store's correctness depends on bounded uncertainty.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversal Cost
Time daemon and reference sourcesTwo-wayConfig change; verify offsets
Lease TTL and renewal intervalTwo-wayConfig change; watch flap rate
HLC max-offset settingMostly two-wayMust be uniform cluster-wide; lowering it can crash healthy nodes
Last-write-wins vs siblings/CRDT in the data modelOne-wayEvery reader changes; historical data has already lost writes
Timestamp format in IDs and keys (ms vs µs, epoch)One-waySort order and index layout of every existing row
Client-supplied vs server-assigned write timestampsMostly one-wayClients in the wild keep sending timestamps; long deprecation

The Standard I'd Write#

RFC: Time, Ordering and Leases (v1)

Scope: Any service that measures durations, orders writes across hosts, or grants time-bounded ownership.

MUST:

  1. Measure durations, timeouts and rate windows with a monotonic clock. Wall-clock subtraction is a lint error.
  2. Shared stores MUST assign write-ordering timestamps server-side (log position or HLC). Client- or app-server-supplied timestamps MUST NOT decide conflicts.
  3. Any lease protecting a write MUST issue a fencing token, and the protected resource MUST reject stale tokens.
  4. Every host MUST export clock.offset_ms and clock.sync_state; alert at 10ms (stateful tiers) / 100ms (stateless) for 5 minutes.
  5. Designs that assume a skew bound MUST cite the published time SLO.

SHOULD: use the platform coordination library for leader election; carry causal tokens across service hops where read-your-writes matters.

Exceptions: Display-only timestamps and logs. Approved by the platform architecture group for anything else.

Success metrics: zero incidents attributed to unfenced leases or wall-clock ordering per year; 99.99% of host-minutes within the SLO bound.

What I'd Tell the VP#

"Our systems use clocks to decide which change came last and who is allowed to do a job. Right now, if one server's clock drifts, customer data can silently revert, and we found out about the last one from support tickets. The fix is mostly cheap: watch every server's clock, have the database decide ordering instead of hundreds of app servers, and use one shared, safe locking tool. That's about half an engineer for a year. The expensive version, precise time hardware in every region, is only needed if we build a globally consistent ledger, and we'd decide that with that project."

Principal Interview Signals#

SignalWhat It Sounds Like
Treats skew as a measured SLO"Which number are we assuming for skew, who publishes it, and who gets paged when it's violated?"
Separates liveness from safety"Leases make failover fast. Fencing makes it correct. I'm only standardizing the second."
Prices bounded uncertainty"Commit wait costs ~2ε per transaction. Our hardware decides ε, so time infrastructure is now part of our latency budget."
Finds hidden owners"The base-image team owns our data correctness without knowing it. That has to change."
Knows what not to centralize"Merge semantics for carts and docs stay with product teams. The platform owns clocks and tokens, not business conflict rules."

Staff answers that L7 interviewers find insufficient:

  • "We'll use a fenced lease in this service." — Correct locally; silent on the four other lock implementations in the company.
  • "NTP keeps clocks within a few ms." — An assumption, not a measured guarantee with an owner.
  • "We'll use HLCs." — Doesn't say what max offset is assumed or what happens to the fleet when it's exceeded.

How Real Companies Built It#

These are public, documented examples.

Google Spanner: TrueTime and Commit Wait#

Spanner's OSDI 2012 paper describes TrueTime, an API that returns an interval guaranteed to contain true time, backed by GPS receivers and atomic clocks in each datacenter, with daemons that detect and reject "liar" time masters. The paper reports that ε is typically a sawtooth from about 1 to 7ms over each 30-second poll interval — about 4ms most of the time — and that a commit leader waits until the chosen timestamp is guaranteed to be in the past before making the commit visible, which gives externally consistent transactions across regions (OSDI '12 paper).

Staff insight: TrueTime doesn't eliminate clock uncertainty — it measures it and makes commits wait it out. In an interview, the speakable version is "correctness costs ~2ε per commit, so the time hardware is part of the latency budget."

CockroachDB: Hybrid Logical Clocks and a Maximum Offset#

CockroachDB documents that it implements hybrid logical clocks with a physical component close to local wall time and a logical component, and that a node which detects its clock is out of sync with at least half the cluster by 80% of the configured maximum offset crashes immediately to protect consistency (architecture docs). Its engineering blog explains how, without atomic clocks, a read that encounters a value inside its uncertainty window performs an "uncertainty restart" at a higher timestamp instead of waiting (Cockroach Labs blog).

Staff insight: Same problem as Spanner, different bill: instead of commit wait on every write, you pay occasional read restarts on contended keys and accept that bad clocks take nodes offline. Naming that tradeoff is the signal.

Cloudflare: The 2017 Leap Second#

Cloudflare's post-mortem describes how the leap second at midnight UTC on January 1, 2017 made time appear to go backwards on its DNS servers. Code that measured upstream response times by subtracting timestamps produced negative values, which flowed into a weighted random selection that panics on negative input. At peak about 0.2% of DNS queries were affected; the fix was a check that discards negative time differences (Cloudflare blog).

Staff insight: The bug was not exotic distributed-systems theory — it was a duration computed from a non-monotonic clock. That's the most common clock bug in production, and the cheapest one to prevent with a lint rule.

AWS ClockBound: Error Bounds for Applications#

AWS publishes ClockBound, an open-source daemon and client library that exposes a clock error bound to applications: instead of a single timestamp it returns an (earliest, latest) window within which true time lies, plus the synchronization status, read through shared memory without system calls (GitHub).

Staff insight: Bounded-uncertainty reasoning is no longer exclusive to one company's custom hardware. If a design needs "this happened definitely after that," ask for the error bound, not just the timestamp.


Staff Calibration#

What Staff Engineers Say (That Seniors Don't)#

ConceptSenior (L5)Staff (L6)Principal (L7)
Ordering writes"Use the timestamp; latest wins""The partition leader's log order decides; LWW only where losing a concurrent write is acceptable, and the store stamps it, not the client""Client-supplied timestamps are banned for conflict resolution org-wide; the standard says who may assign order"
Clock sync"NTP keeps clocks in sync""NTP gives ~1–50ms; I'll alert at 10ms on DB nodes and the store self-fences beyond its max offset""Skew is a published SLO with an owner and a pager; designs cite it"
Locks"Take a Redis lock with a TTL""The TTL gives liveness; the fencing token checked by storage gives safety""One fenced-lease primitive for the company; unfenced locks fail design review"
Durations"end - start""Monotonic clock for durations; wall time steps backwards on leap seconds and corrections""Lint rule plus a library; this class of bug should be impossible, not reviewed"
Global consistency"Use a strongly consistent DB""HLC with bounded offset gives causality; external consistency needs bounded-uncertainty time and commit wait""We buy precise time only for the systems whose revenue depends on strict serializability; everything else runs on causality"
Why "Locks" separates levels

The Senior answer works almost all the time, which is exactly why it's dangerous: the failure needs a pause longer than the TTL, which happens on a large heap or a busy VM perhaps once a month. Staff engineers recognize that no holder-side check can close the gap and move enforcement to the resource. Principal engineers notice the pattern repeating across teams and remove the unsafe option from the menu.

Why "Ordering writes" separates levels

"Latest timestamp wins" sounds like a definition of correctness. Staff candidates ask whose clock and how wrong it can be, and usually discover that a single writer per key makes clocks irrelevant. Principal candidates see that the riskiest timestamps are the ones from hundreds of app servers managed by many teams, and standardize where ordering is decided.

Common Interview Traps#

  • Using wall-clock subtraction for timeouts or latency. It goes negative on a step correction or leap second.
  • Assuming NTP means "in sync." Say a number, and say what happens when a host exceeds it.
  • Taking a lock with a TTL and calling it mutual exclusion. A GC pause longer than the TTL breaks it. Name the fencing token and who checks it.
  • Last-write-wins on data that can't lose writes. Balances, inventory counts, and document bodies need a single writer, CAS, or merge semantics.
  • Trusting client device clocks. Phones are minutes off. Record server receipt time.
  • Thinking Lamport timestamps detect conflicts. They order; they don't detect concurrency. That's what vector clocks are for.
  • Raising the database's max clock offset during an incident. It hides the broken clock and weakens the guarantee for every transaction.
  • Ordering events by trace or log timestamps across hosts. Use causal links (span parents, log positions).

Practice Drill#

Prompt: "Our nightly billing job runs on one of three workers, elected via a Redis key with a 60-second TTL. Last month a customer was billed twice. The team wants to shorten the TTL to 10 seconds. What do you do?"

Staff Answer

Shortening the TTL makes double-billing more likely, not less: a 10-second TTL expires during an ordinary multi-second GC pause or VM migration, at which point a second worker acquires the key while the first still believes it is the leader. The root cause is that the lock grants liveness but nothing enforces safety — the paused worker's writes still land. Fix it in layers. (1) Fence: replace the Redis key with a lease from etcd (or a database row) that yields a strictly increasing token — the etcd revision works — and make every billing-run state transition conditional: UPDATE billing_runs SET state='charging', token=$t WHERE run_date=$d AND token < $t. A stale worker's update affects zero rows and it aborts. (2) Idempotency at the payment edge: each charge carries an idempotency key of (customer_id, billing_period), so even a duplicate that slips past fencing is deduplicated by the payment provider. This is the layer that actually protects the customer. (3) Count time correctly: the leader measures its lease on the monotonic clock from before the acquire request and stops starting new charges with 20% of the lease remaining. (4) Keep the TTL at 30–60s — the job runs nightly, so failover speed is worth little and stability is worth a lot. (5) Detect: alert on billing.duplicate_charge_attempts (rejected by idempotency) and lease.rejected_stale_token; both should be zero on a normal night, and a nonzero value is the early warning. Owner: the billing team owns the fencing and idempotency keys; the platform team owns the lease service; finance signs off on the refund for the affected customer and on whether to audit the past 90 days for other duplicates.

Why this is L6:

  • Rejects the proposed fix with a mechanism (pauses longer than the TTL), not an opinion.
  • Separates liveness (lease) from safety (fencing token checked by storage) and adds idempotency as the customer-facing guarantee.
  • Defines detection signals that turn a silent failure into an alert, and names who owns each layer.

What L7 adds:

  • Asks how many other jobs use the same Redis-TTL pattern and makes the fenced-lease library the only approved election mechanism.
  • Requires idempotency keys on every money-moving call as an org standard, with the payments platform rejecting calls that lack one.
  • Prices the audit: a 90-day duplicate scan costs a few engineer-days, far less than one regulator or chargeback escalation.

Where This Appears#

Related Foundations & Patterns: Consistency, CAP & PACELC · Replication · Concurrency Control · Coordination Strategies

Related Technologies: etcd & ZooKeeper · Distributed SQL · Cassandra · Redis

  1. Loading the index…