Hiring BarSupport

Design a Distributed Job Scheduler — Staff-Level Case Study

Case study73 min read6 diagrams

Technologies referenced in this case study: PostgreSQL · Kafka · Redis · Cassandra · DynamoDB · ZooKeeper & etcd

Related: Message Queues · Long-Running Processes · Distributed Coordination · Contention · Degraded Mode · Distributed Consensus

How to Use This Case Study#

Organized for interview use first, reference second. Read front-to-back once. Return to individual sections for targeted review.

ModeTimeWhat to Read
Quick Review15 minExecutive Summary → Interview Walkthrough → Fault Lines table → Active Drills 1–3
Targeted Study1–2 hrsExecutive Summary → Walkthrough → Section 3 (Fault Lines) → Section 4 (Failure Modes) → two Deep Dives
Deep Dive3+ hrsEverything, including the Principal Lens (Section 11) and Appendices
What is a Distributed Job Scheduler? — Why interviewers pick this topic

A distributed job scheduler accepts work that must run later — at a specific time ("send the renewal reminder at 09:00 local on March 3"), after a delay ("expire this cart in 30 minutes"), or on a recurrence ("rebuild the search index every hour") — and guarantees that work is handed to an executor near that time, even if the machines that accepted it have since died.

It is the thing that replaces crontab on one box once one box is no longer acceptable.

Before vs After — the "cron box" incident:

Without a real scheduler (cron on one VM):
t=0:        Host cron-prod-01 runs 340 crontab entries for 22 teams
t=+0s:      Hypervisor maintenance reboots the VM at 01:58
t=+2min:    02:00 jobs (billing export, cert renewal, GDPR purge) never fire
t=+6h:      Finance notices missing export; nobody knows which other jobs missed
t=+9h:      Someone re-runs billing export by hand — it runs twice, double-charges 4,100 invoices
t=+3 days:  Cert renewal job never ran; TLS cert expires; partner API integration down

With a distributed scheduler:
t=0:        Scheduler shard 7 owner dies at 01:58
t=+10s:     Lease on shard 7 expires; shard reassigned to node 3
t=+12s:     Node 3 scans shard 7's due window, finds 02:00 jobs pending
t=+2:00:04: Jobs fire 4s late; fire_lag_p99 alert stays green (threshold 30s)
t=+2:00:05: Billing export runs with run_id idempotency key — exactly one export written

Why interviewers reach for this question: It looks like a CRUD-plus-timer problem, so it filters candidates who stop at "store jobs in a table and poll it." The real content — lease-based ownership, the duplicate-vs-missed tradeoff, the top-of-the-hour thundering herd, and the fact that "exactly-once" is a property of the job not the scheduler — is where Senior and Staff answers diverge sharply. It also has a strong organizational dimension: a scheduler is a platform that 30 teams depend on and nobody wants to own.

Mechanics Refresher: How Timers Are Stored and Fired
MechanismHow It WorksProsCons
DB pollingSELECT … WHERE run_at <= now() … FOR UPDATE SKIP LOCKED LIMIT 100 every 1sSimple, transactional, one storeIndex hot spot on run_at; ~2–10K dequeues/s per Postgres primary
Time-bucketed partitionsTimers written to bucket = floor(run_at / 60s) × shard; scheduler reads whole bucket when dueHorizontal scale; sequential reads; no global indexBucket granularity bounds precision; late writes into past buckets need handling
In-memory timing wheelHierarchical ring of slots (e.g., 1s × 60, 1m × 60, 1h × 24); O(1) insert/fireMicrosecond insert, millions of timers per nodeVolatile — must be rebuilt from durable store on failover
Redis sorted setZADD timers <run_at> <job_id>; ZRANGEBYSCORE 0 now to find dueFast, simple, O(log N)Memory-bound; durability depends on AOF config; one hot key per set
Broker delay featureSQS DelaySeconds, RabbitMQ delayed exchange, Kafka tiered delay topicsNo scheduler to buildSQS caps delay at 15 minutes; tiered topics give coarse precision
Durable execution engineTemporal / Cadence persist workflow state; timers are first-classRetries, timeouts, history built inHeavy platform; learning curve; you now operate a database cluster

For most production systems: durable timers in time-bucketed partitions (or Postgres with SKIP LOCKED below ~5K jobs/s), an in-memory timing wheel per shard owner for the next few minutes, and pull-based workers with leases feeding off a ready queue. The storage mechanism is almost never the interview question — ownership, duplicates, and the herd are.


Executive Summary

If you only read one section, read this. Everything in the case study flows from the contrast below.

What This Interview Actually Tests#

A job scheduler is not a cron question. Everyone can store a run_at column.

It is a "what happens to work when the thing holding it dies" question that tests:

  • Whether you pick between duplicate execution and missed execution out loud — because you cannot eliminate both
  • Whether you know that exactly-once is a contract the job signs (idempotency), not a guarantee the scheduler gives
  • Whether you see the top-of-the-hour herd coming before the interviewer points at it
  • Whether you define who owns a failed job — the platform team or the team that wrote it

The key insight: The scheduler's only real promise is "every due job is dispatched at least once, within a bounded lag of its scheduled time, and every dispatch is observable." Everything else — retries, idempotency, timeouts, dead-lettering — is a contract between the platform and job owners, and Staff candidates design that contract explicitly.

The L5 vs L6 vs L7 Contrast — Start Here#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First moveDraws a jobs table + a poller + workersAsks "cron replacement, delayed one-shots, or workflow DAGs?" and commits to oneAsks "how many schedulers does this org already run, and which ones do we retire?"
Guarantee"Exactly-once execution""At-least-once dispatch + idempotent jobs. Exactly-once is the job's contract, enforced by a run_id key."Makes idempotency a platform onboarding gate: no idempotency key, no scheduler access
Failure"Run a standby scheduler"Lease-based shard ownership with fencing tokens; quantifies failover gap (lease TTL ≈ 10s)Designs cell-based schedulers so one bad shard map can't stall every team's jobs; runs quarterly failover game days
Load shape"Autoscale the workers"Names the :00 herd — 40–60% of cron jobs fire on the minute — and adds jitter windows with owner opt-inSets org policy: "precision is a paid tier"; default jobs get a ±5 min window; exact-time requires justification
Ownership"The scheduler retries failures"Separates dispatch (platform) from execution (job team); DLQ per owning team; pages the job owner, not the platformRedraws the boundary: platform SLO on fire-lag only; job success rate is the product team's SLO, reported on a shared scorecard
Scale"Shard the jobs table"Time-bucket × hash-shard partitioning; one owner per shard; rebalance by lease handoffKnows when to buy (Temporal Cloud, EventBridge Scheduler) vs build; prices the migration of 2,000 legacy crontabs
Why "guarantee" separates levels

L5: Says "exactly-once" because it sounds rigorous. Under probing ("the worker finished the job and crashed before acking — what now?") the design either re-runs the job (at-least-once, contradicting the claim) or never re-runs it (at-most-once, silently losing work on every worker crash).

L6: States the impossibility up front and designs around it. "A worker can die between doing the side effect and recording that it did it. No scheduler can close that window. So I'll guarantee at-least-once dispatch, give every run a stable run_id = hash(job_id, scheduled_time), and require jobs to use it as an idempotency key against their own side effects. The scheduler's dedup reduces duplicates to rare; the job's idempotency makes them harmless."

L7: Turns the contract into policy. "Idempotency is not optional documentation. Job registration requires declaring idempotency: natural | keyed | unsafe. unsafe jobs get at-most-once semantics — no automatic retries — and the owning director signs off that a missed run is acceptable."

Why "load shape" separates levels

L5: Plans capacity for the average: 50M jobs/day ≈ 580 jobs/s. Provisions for 2× that.

L6: Knows humans write 0 * * * * and 0 0 * * *. In real fleets, a large fraction of recurring jobs are scheduled exactly on the minute, hour, or midnight UTC. "Average is 580/s, but at 00:00:00 UTC we'll see ~400K jobs due in the same second. Peak-to-average is 100×+ for one second. I'll smear non-precise jobs across a declared tolerance window — deterministically, by hash(job_id) — so the peak becomes ~7K/s spread over the first minute."

L7: Recognizes the herd hits downstream systems too — 300 teams' midnight jobs all hit the same shared Postgres and the same payments API. The fix is an org default (tolerance windows on by default, precision opt-in) and a calendar of "known herd moments" reviewed with capacity planning.

Why "ownership" separates levels

L5: The scheduler retries failed jobs 3 times and logs errors. Who reads the logs is unspecified — in practice, the scheduler on-call gets paged for every team's broken job.

L6: Draws the line: "The platform owns 'the job was dispatched on time to a healthy worker pool.' The job team owns 'the job succeeded.' Failed runs go to a per-team DLQ with the team's on-call rotation attached at registration time. The platform pager should never ring because someone's SQL query has a typo."

L7: Makes that boundary measurable: two SLOs, two scorecards, and a quarterly review where teams with >5% job failure rate lose the right to page the platform.

The Staff Positions#

PositionRationale
At-least-once dispatch + idempotent jobsMissed jobs are silent; duplicates are loud and preventable with a run_id key. Choose the failure you can detect.
Pull with leases over pushWorkers that pull only take what they can run; leases with heartbeats give automatic recovery when a worker dies
Shard ownership via leases + fencing tokensOne owner per time-shard avoids double-firing; fencing stops a zombie owner from dispatching after losing the lease
Time-bucketed storage over one global run_at indexGlobal index is a write hot spot at the "now" edge; buckets turn firing into sequential reads
Jitter by default, precision by exceptionThe :00 herd is the #1 scaling failure; most jobs don't care about ±60s
Platform owns dispatch; teams own executionThe platform pager must not ring for job logic bugs; DLQ + alerting route to the job owner
Separate the timer service from the execution runtimeFiring a timer and running a 6-hour batch have different scaling, isolation, and failure properties

The Three Intents#

Three intents drive every design decision. Each leads to a fundamentally different architecture.

IntentConstraintStrategyFailure ModeCorrectness Bar
Recurring cron (infra & batch)Low count (10K–1M definitions), strong expectations on "did it run"Durable definitions, next-fire computation, run history, overlap policyMissed run goes unnoticed for daysEvery scheduled occurrence accounted for: ran, skipped by policy, or failed-and-alerted
Delayed one-shot timers at scaleHigh volume (10K–1M timers/s created), mostly cancelled or fired onceTime-bucketed durable timers, sharded ownership, ready queue, workersBacklog grows at the herd; fire-lag balloonsFire-lag p99 < 1–5s; no timer lost; duplicates rare and harmless
Workflow / DAG orchestrationMulti-step, long-running (minutes → days), dependencies, human waitsDurable execution (Temporal/Cadence), Airflow-style DAG schedulerStuck workflow, poisoned history, non-deterministic replayEach step's state persisted; resumes after any crash

🎯 Staff Move: "These three share a word — 'schedule' — and nothing else. Cron needs auditability, timers need throughput, workflows need durable state. I'll design the delayed-timer and recurring-job core — a time-triggered dispatch service — because that's where the distributed ownership problem lives, and I'll treat workflow orchestration as a consumer that sits on top of it rather than something the timer service tries to be."

The Five Fault Lines#

#Fault LineThe Tension
1Duplicate vs Missed ExecutionAt-least-once (risk running twice) or at-most-once (risk never running)? You must pick per job class.
2Polling a Database vs Time-Partitioned TimersOne transactional table (simple, ~5K/s ceiling) vs bucketed partitions + in-memory wheel (scales, more moving parts)?
3Push vs Pull DispatchScheduler assigns work to workers (fast, needs worker health tracking) vs workers lease from a queue (self-balancing, lease tuning)?
4Precision vs Herd SmoothingFire exactly at :00 (what users typed) or smear across a window (what infrastructure can survive)?
5Central Platform vs Team-Owned SchedulersOne scheduler with governance, or every team runs its own cron/Airflow/Quartz?

In the Wild: Real Production Systems#

Why this section belongs here: Citing specific production systems demonstrates you've studied operational reality, not textbook designs.

Uber Cadence → Temporal — Timers as Durable Workflow State#

Uber built Cadence to orchestrate long-running business processes; its authors later founded Temporal around the same model. Workflow code calls sleep(30 days) and the engine persists a durable timer; if every worker dies, the workflow resumes from its event history when the timer fires. Timers, retries, and activity timeouts are platform primitives rather than application code.

Staff insight: Temporal draws the exact line this case study argues for: the platform guarantees the timer fires and the step is retried; activities must be idempotent because the platform delivers them at-least-once. Cite it when the interviewer asks "can you make it exactly-once?" — the most successful durable-execution engine explicitly says no.

Apache Airflow (originated at Airbnb) — The Scheduler as a Bottleneck#

Airflow's scheduler parses DAG files, computes which task instances are due, and writes them to a metadata database; executors pull from there. For years a single scheduler process was the ceiling; Airflow 2.0 added support for running multiple schedulers concurrently, coordinating through row-level locks (SELECT … FOR UPDATE SKIP LOCKED) in the metadata database.

Staff insight: This is the "DB polling" design at real scale, and its evolution shows the ceiling: the metadata database becomes the coordination point. When you propose SKIP LOCKED, also say where it stops scaling — and that Airflow is a DAG scheduler, which is a different intent from a high-volume timer service.

Kubernetes CronJob — Overlap and Missed-Run Policy as First-Class Config#

Kubernetes CronJob exposes concurrencyPolicy (Allow, Forbid, Replace) and startingDeadlineSeconds. If the controller misses too many schedules (more than 100 missed start times, per the documentation), it stops trying to catch up and logs an error.

Staff insight: Kubernetes forced every user to answer two questions most candidates never ask: "If the previous run is still going, do we skip, run in parallel, or kill it?" and "If we missed the window, do we catch up or skip?" Naming these policies — and who chooses them — is a strong Staff-level signal.

What Interviewers Probe#

After You Say...They Will Ask...(What They're Evaluating)
"Store jobs in a table, poll every second""What's the ceiling? What happens at midnight UTC?"Whether you know the herd and the index hot spot
"Exactly-once execution""Worker crashes after charging the card but before acking. Now what?"Whether you understand the side-effect/ack window
"Leader election for the scheduler""The old leader is paused by GC for 20s and wakes up still thinking it's leader."Fencing tokens, not just election
"Workers retry on failure""Who gets paged when a job fails 3 times? The platform team?"Ownership boundary between platform and job owner
"We'll shard by job_id""How does a shard owner find what's due now without scanning everything?"Time-dimension partitioning
"Autoscale workers""Autoscaling takes 3 minutes. The herd lasts 10 seconds."Whether you smooth demand vs chase it

System Architecture Overview#

Diagram: System Architecture Overview

Reading the diagram: The API writes two things: a definition (small, relational, owned by a team) and a timer (huge, append-mostly, partitioned by time first). Scheduler nodes never scan a global index — each owns a slice of shards through an etcd lease, pre-loads the next five minutes of timers into an in-memory timing wheel, and dispatches to a ready queue stamped with a deterministic run_id. Workers are owned by the job teams, pull at their own pace, and report outcomes. The platform alerts on fire-lag (its promise); teams alert on failure rate (theirs).

Quick-Reference: The 30-Second Cheat Sheet#

TopicThe L5 AnswerThe L6 Answer — Say This
Guarantee"Exactly-once""At-least-once dispatch, idempotent jobs keyed on run_id = hash(job_id, scheduled_at). Exactly-once is the job's contract."
Storage"Jobs table with an index on run_at""Postgres + SKIP LOCKED up to ~5K/s. Beyond that, time-bucketed partitions with shard owners."
Coordination"Leader election""1,024 shards, each owned via a 10s etcd lease. Every dispatch carries the lease's fencing token; stale tokens are rejected."
Dispatch"Scheduler sends job to a worker""Scheduler enqueues; workers pull and hold a visibility lease with heartbeats. Dead worker → lease expires → re-dispatch."
Herd"Autoscale""40–60% of cron fires on :00. Deterministic jitter within a declared tolerance; precision is opt-in."
Failure ownership"Scheduler retries and logs""Platform owns fire-lag SLO. Job owners own success rate, their DLQ, and their pager."

Key Numbers Worth Memorizing#

MetricValueWhy It Matters
Postgres SKIP LOCKED dequeue ceiling~2–10K jobs/s per primaryPast this, the run_at index and vacuum become the bottleneck
Shard lease TTL10s (renew every ~3s)Sets the worst-case failover gap for firing
Fire-lag SLO (delayed timers)p99 < 5sThe platform's actual promise
Fire-lag SLO (cron, default tolerance)p99 < 60sBuys the jitter window that kills the herd
Share of cron jobs on :00commonly 40–60%Why average-based capacity planning fails
Worker visibility lease1–5 min, heartbeat at lease/3Too short → duplicate runs of slow jobs; too long → slow recovery
Timer record size~200–500 bytes1B pending timers ≈ 200–500 GB before replication
Retry scheduleexp backoff 30s × 2ⁿ, cap 1h, 5 attempts, ±20% jitterRetries must not re-create the herd
SQS max delay15 minutesWhy "just use SQS delay" fails for long timers
NTP clock skew (healthy fleet)~1–10 ms; alarm at > 100 msWhy fencing tokens, not timestamps, decide ownership
Timing wheel insert/fireO(1), millions of timers per node in RAMWhy the in-memory wheel covers only the next N minutes

Interview Walkthrough

The most common mistake: Candidates spend 20 minutes on the jobs table schema, cron-expression parsing, and a REST API, then run out of time before the interviewer asks the only question that matters: "A worker died mid-job. Walk me through exactly what happens." The phases below compress the basics to ~10 minutes.


Phase 1: Requirements & Framing (2–3 minutes)#

State functional requirements in 30 seconds:

"Clients schedule work to run at a time, after a delay, or on a recurrence. They can cancel or update it. The system fires the work to an executor near the scheduled time and records what happened."

Then spend the remaining time on intent and non-functionals — this is the Staff move:

"'Scheduler' covers three different systems: a cron replacement for ~100K recurring jobs, a high-volume delayed-timer service — think 'expire this reservation in 15 minutes' at tens of thousands per second — and a workflow engine for multi-step processes. I'll design the timer-and-cron core, because the distributed ownership problem lives there. Workflows can sit on top of it."

Commit to a constraint set with numbers:

"Assume 50M timer creations a day, ~580/s average, 50K/s burst; 1B pending timers at any time because many are days out; 200K recurring definitions. Fire-lag p99 under 5 seconds for timers, under 60 seconds for default cron. At-least-once dispatch; jobs must be idempotent. No timer may be silently lost — a missed fire is worse than a duplicate."

🎯 Staff Move: The sentence "a missed fire is worse than a duplicate" is the single most load-bearing statement in this interview. It chooses a side of Fault Line 1 in the first three minutes, and every later decision — leases, retries, idempotency keys — follows from it.


Phase 2: Core Entities & API (1–2 minutes)#

Name the nouns in 30 seconds:

  • JobDefinition: job_id, owner_team, schedule (cron expr / one-shot run_at), tolerance_s, payload_ref, retry_policy, overlap_policy, idempotency_class
  • Timer: (bucket, shard, run_at, job_id) — the durable "fire this at T" record, one per upcoming occurrence
  • Run: run_id = hash(job_id, scheduled_at), attempt, status, lease_owner, started_at, finished_at
POST   /jobs                 { schedule, target, payload, tolerance_s, retry_policy, client_key } → { job_id }
DELETE /jobs/{job_id}        → cancels future occurrences (idempotent)
GET    /jobs/{job_id}/runs   → run history, attempts, last error

Worker protocol (pull, not REST callbacks):

Lease(pool, max=10, lease_s=300) → [{ run_id, job_id, attempt, payload, fencing_token }]
Heartbeat(run_id, fencing_token) → extends lease
Complete(run_id, fencing_token, status, result_ref)

🎯 Staff Move: "run_id is deterministic — hash of job and scheduled time — not a random UUID. That's what makes every retry and every duplicate dispatch of the same occurrence collapse onto one idempotency key downstream. A random run ID would make duplicates indistinguishable from distinct runs."


Phase 3: High-Level Architecture (≤5 minutes)#

Diagram: Phase 3: High-Level Architecture (≤5 minutes)

Walk the flow in 90 seconds:

  1. API validates, writes the definition, computes the next occurrence, writes a timer into bucket = floor(run_at/60s), shard = hash(job_id) % 1024
  2. Each scheduler node holds leases on ~64 shards; every 30s it loads timers for the next 5 minutes of its shards into an in-memory timing wheel
  3. When a slot fires, the node writes a dispatch record to the ready queue with run_id and its fencing token
  4. Workers pull, take a visibility lease, heartbeat, and complete; outcome goes to run history
  5. For recurring jobs, completion of dispatch (not execution) triggers computing and writing the next timer
  6. Lease loss → another node picks up the shard, re-reads the durable bucket, and skips anything already in run history as dispatched

Key points to state explicitly:

  1. Time-first partitioning — firing is a sequential read of "what's due in this minute for my shards", never a global ORDER BY run_at
  2. Ownership via leases with fencing — exactly one node fires each shard; a stale node's dispatches are rejected
  3. At-least-once with deterministic run IDs — duplicates are possible and harmless
  4. Workers pull — the scheduler never needs to know which workers are alive
  5. Two SLOs — platform: fire-lag; teams: success rate

🎯 Staff Move: Draw ≤8 boxes and then say: "This is the design that works on a good day. Let me show you the three days it doesn't: a scheduler node pauses, a worker dies mid-job, and midnight UTC."


Phase 4: Transition to Depth (1 minute)#

"The architecture is straightforward. What makes this hard is failure. Three places I'd go deep: the duplicate-vs-missed guarantee and how idempotency closes it; shard ownership when a node is paused rather than dead; and the top-of-the-hour herd. Which do you want first?"

If there's no preference, lead with the guarantee — it anchors everything else.


Phase 5: Deep Dives (25–30 minutes)#

For each: state the tradeoff → pick a position → quantify → name who pays.

Deep dive A: Duplicate vs missed (6–8 min)

"There's a window between a worker performing the side effect and the scheduler learning it did. If the worker dies in that window, the scheduler has two options: re-dispatch (possible duplicate) or not (possible miss). I re-dispatch. Duplicates are made harmless by run_id: the billing job writes INSERT … ON CONFLICT (run_id) DO NOTHING before charging. For the rare 'unsafe' job that genuinely can't be idempotent — say, a legacy integration that emails a PDF — the owner opts into at-most-once and signs off that a crash means a missed run."

Quantify the window: "With a 5-minute visibility lease and 60s heartbeats, a dead worker's run is re-dispatched after at most 5 minutes. At a 0.1% worker crash rate per run and 50M runs/day, that's ~50K re-dispatches per day — which is why idempotency cannot be optional."

Deep dive B: Shard ownership and zombies (6–8 min)

"Leader election isn't enough. A node can be paused — a 20-second GC pause, a VM live-migration — lose its lease, and wake up still believing it owns shard 42. Meanwhile node 3 took over and fired those timers. If node 7 now fires them too, we double-dispatch a whole minute of a shard. Fix: etcd lease revision is a monotonically increasing fencing token. Every dispatch carries it; the ready-queue writer (or run-history insert) rejects any token lower than the highest it has seen for that shard. The zombie's writes bounce."

Quantify failover: "Lease TTL 10s, renewal every 3s. Worst case: ~10s where a shard has no owner, plus ~2s to load its window. Fire-lag during failover is ~12s — inside a 30s alert threshold, outside a 5s SLO for those shards for one event. I'll accept that and burn error budget rather than shrink the TTL to 3s and get false failovers on every network blip."

Deep dive C: The herd (6–8 min)

"Of 200K recurring jobs, expect ~100K to be on 0 * * * * or 0 0 * * *. At 00:00:00 that's 100K dispatches in one second plus all the hourly ones, against a 580/s average. Autoscaling doesn't help — it reacts in minutes. Instead every job declares tolerance_s; default 300. The scheduler places each occurrence at scheduled_at + hash(job_id) % tolerance_s. Deterministic, so the same job fires at the same offset every day — predictable for owners, spread for us. 100K jobs across 300s is ~330/s. Jobs that truly need :00 — market open, a contractual export — set tolerance_s: 0 and go through review."

🎯 Staff Move: Every deep dive ends with an owner. "Tolerance defaults are set by the platform team; opting out is approved by the platform on-call lead; the downstream DB team is told the herd calendar."


Phase 6: Wrap-Up (2–3 minutes)#

"To summarize: time-bucketed durable timers, lease-owned shards with fencing tokens, at-least-once dispatch with deterministic run IDs, pull-based workers with visibility leases, and jitter-by-default to kill the herd. The platform's promise is fire-lag; job success is the owner's. What I'd build next: per-tenant quotas on the ready queue so one team's 10M-job backfill can't starve everyone, and a migration path to pull the remaining crontabs off VMs. What I'd deliberately not build: a workflow engine inside the timer service — if we need one, we adopt Temporal on top."

Common Timing Mistakes#

MistakeTime LostFix
Designing cron-expression parsing5–8 min"Standard cron library; next-fire computed at write time."
Full DB schema with indexes5 minName three entities, one sentence on partition keys
Debating Kafka vs RabbitMQ for the ready queue5 min"Any durable queue with per-partition ordering. Kafka, because we already run it."
Worker autoscaling details4 min"Workers scale on queue lag; the herd is solved upstream, not by autoscaling."
Never reaching "worker dies mid-job"Interview-endingGet to the guarantee by minute 12

1. The Staff Lens#

1.1 Why This Problem Exists in Staff Interviews#

Every company above ~50 engineers has a scheduler problem, and almost none of them decided to have one. It starts as a crontab on a bastion host. Then it's three crontabs, a Jenkins job, a Kubernetes CronJob per team, an Airflow cluster the data team runs, and a Quartz instance embedded in the monolith. Each of these fails differently, alerts differently, and has a different answer to "did last night's run happen?"

The interview compresses that history into 45 minutes. The interviewer isn't checking whether you can build a timer — they're checking whether you see that the scheduler is the place where the company's "did it run?" accountability lives, and whether you design that accountability instead of letting it emerge.

The distributed-systems content is real and dense: leases, fencing, clock skew, idempotency, partitioned time-series storage, backpressure. But every one of those mechanisms exists to answer an ownership question: who is responsible for this occurrence right now, and how does everyone else know?

1.2 The L5 vs L6 Contrast — Visual#

Diagram: 1.2 The L5 vs L6 Contrast — Visual

The L5 path is correct until step 5. It stalls because it built mechanism before choosing a guarantee, so the first hard failure question has no pre-committed answer. The L6 path chooses the guarantee at step 2; every later mechanism is justified by it.

1.3 The Staff Question That Cuts Through Everything#

"When a job runs twice or doesn't run at all, who finds out first — and is it a customer?"

Ask this out loud. It forces three decisions at once:

  • Guarantee: Which failure (duplicate or miss) do we prefer, per job class?
  • Detection: What metric shows a miss? (runs.missed_total requires the scheduler to know what should have run — which means occurrences are materialized, not computed lazily.)
  • Ownership: Whose pager rings — the platform's or the job owner's?

If the answer is "a customer, via a support ticket, three days later," the design isn't done.


2. Problem Framing & Intent#

2.1 The Three Intents — Explained#

Intent 1: Recurring cron (infra & batch). Low volume, high stakes per occurrence. A nightly ledger reconciliation runs once; if it doesn't, finance closes the month on wrong numbers. The correctness bar is accountability: every occurrence must end in a recorded state — succeeded, failed-and-alerted, or skipped-by-policy (e.g., overlap forbidden). The hard problems are overlap policy (previous run still going), catch-up policy (scheduler was down for 3 hours — run 3 missed hourly occurrences, 1, or 0?), time zones and DST (does a 02:30 local job run twice, once, or never on DST days?), and making "did it run?" answerable in one query.

Intent 2: Delayed one-shot timers at scale. High volume, low stakes per timer, most cancelled. "Release this seat hold in 10 minutes." "Send the abandoned-cart email in 24 hours unless they check out." "Retry this webhook in 30 seconds." Tens of thousands created per second, a billion pending, and 60–90% cancelled before firing in cancel-heavy use cases like holds and reminders. The correctness bar is no lost timers + bounded fire-lag. The hard problems are write throughput, cheap cancellation, and firing without a global index.

Intent 3: Workflow / DAG orchestration. A multi-step process where step 3 waits on step 2 and step 5 waits for a human approval that may take a week. The correctness bar is resumability: after any crash, the workflow continues from its last durable step. Here the scheduler is a subcomponent — timers inside workflows — and the hard problems are state persistence, deterministic replay, and versioning running workflows when code changes. This is Long-Running Processes territory.

🎯 Staff Move: "I'll build intents 1 and 2 as one timer service — a cron occurrence is just a timer that schedules its successor — and explicitly not build intent 3. If we need durable workflows, we run Temporal and let it use its own timers. Trying to make the timer service also a workflow engine is how you get a 4-year platform rewrite."

2.2 When NOT to Build a Distributed Scheduler#

SituationUse InsteadWhy
< 1K jobs, one team, tolerance for a missed runKubernetes CronJob or managed cron (EventBridge Scheduler, Cloud Scheduler)Zero platform to operate; built-in history
Delays under 15 minutes, AWS-nativeSQS DelaySeconds / message timersQueue already provides durable delay
Delays of seconds for retriesRetry with backoff in the client / queue redeliveryA scheduler round-trip for a 2s retry is overhead
Multi-step business processesTemporal / Cadence / Step FunctionsDurable state and replay are the hard part, not timing
Data pipeline dependenciesAirflow / DagsterDAG semantics, backfills, and lineage matter more than precision
"Every 100ms" polling loopsAn in-process loop or a stream processorSchedulers are for seconds-to-days, not sub-second cadence

"The best scheduler is often the one you don't build. I'd need to see >50K timers/s or a real multi-tenant governance need before I'd staff a team for a custom one."

2.3 What the Interviewer Leaves Underspecified#

Unstated AssumptionWhy It MattersWhat to Say
Guarantee (exactly / at-least / at-most once)Determines retries, leases, and whether idempotency is mandatory"At-least-once dispatch, idempotent jobs."
Precision±1s vs ±5 min changes storage and the herd strategy"p99 < 5s for timers; cron gets a tolerance window by default."
Cancellation rate80% cancelled → cheap cancel matters more than cheap fire"I'll assume most delayed timers are cancelled — tombstones, not deletes."
Job duration1s callbacks vs 6-hour batches need different executors"The scheduler dispatches; execution runtime is separate and team-owned."
Missed-run policy after outageCatch-up can create its own herd"Default: run once for the latest missed occurrence, record the rest as skipped."
Time zones / DSTLocal-time cron is a correctness bug factory"Store in UTC plus IANA zone; compute next-fire with the zone library; DST ambiguity policy explicit."
Multi-tenancyOne team's backfill starving others"Per-team quotas on the ready queue."

2.4 Precise Terminology#

TermPrecise Meaning
OccurrenceOne scheduled instance of a job at a specific scheduled_at
DispatchThe scheduler handing an occurrence to the ready queue — the platform's promise ends here
Run / attemptAn execution of an occurrence by a worker; one occurrence may have several attempts
Fire-lagdispatched_at − effective_scheduled_at (after jitter). The platform SLO.
Visibility leaseTime a worker holds a run exclusively; expires → the run becomes available again
Fencing tokenMonotonic number attached to ownership; downstream rejects writes carrying a lower token
Tolerance windowOwner-declared range within which firing is acceptable; the scheduler uses it to smooth load
Overlap policyWhat happens if occurrence N+1 is due while N is still running: allow, skip, replace
Catch-up policyWhat happens to occurrences missed during downtime: all, latest, none
Idempotency classnatural (safe to repeat), keyed (safe with run_id), unsafe (at-most-once only)

3. The Fault Lines#

3.1 Fault Line 1: Duplicate vs Missed Execution#

The tension: A worker performs a side effect and then reports completion. Between those two steps, anything can die. The system must decide, without knowing, whether to try again.

Diagram: 3.1 Fault Line 1: Duplicate vs Missed Execution
StrategyWhat WorksWhat BreaksWho Pays
At-most-once (mark dispatched before running, never retry)No duplicate side effects everEvery worker crash silently loses a runThe customer whose reminder never sent; finance with a missing export
At-least-once, no idempotencyNothing is lostDuplicate charges, double emails, double exportsCustomers (double charges), support, the job team
At-least-once + run_id idempotencyNothing lost; duplicates collapseRequires every side effect to accept a key; legacy targets can'tJob teams — they write the dedup; platform provides the key
"Exactly-once" via 2PC with side effectTheoretically correctMost side effects (email, third-party APIs) don't do 2PCEveryone — it doesn't actually exist end to end

Staff default: At-least-once dispatch + deterministic run_id + an idempotency_class declared per job. unsafe jobs get at-most-once with an explicit sign-off.

When to deviate: For jobs whose duplicates are catastrophic and whose target cannot dedupe (a legacy wire-transfer file drop), run at-most-once and add a reconciliation job that detects misses the next morning. You trade a real-time miss for a next-day correction — and the finance owner signs that trade.

"Duplicates are loud — someone notices a double email. Misses are silent — nobody notices a reminder that didn't send. I'd rather design for the loud failure and then make it harmless."

3.2 Fault Line 2: Polling a Database vs Time-Partitioned Timers#

The tension: A single transactional table is the simplest possible correct design. It also has a hard ceiling, and the ceiling is hit exactly when the business succeeds.

StrategyWhat WorksWhat BreaksWho Pays
Postgres + FOR UPDATE SKIP LOCKEDTransactional create/cancel/fire; one store; easy "did it run" queries~2–10K dequeues/s; run_at index is a hot right edge; dead tuples from updates → vacuum pressureDB team at scale; everyone during vacuum stalls
Redis sorted set per shard100K+ ops/s per shard; ZRANGEBYSCORE is cheapRAM-bound (1B timers ≈ 100+ GB); durability depends on AOF fsync settingPlatform on-call after a failover loses the last second of timers
Time-bucketed wide rows (Cassandra/DynamoDB)Partition = (minute, shard); fire = read one partition; linear write scalePrecision bounded by bucket; cancel is a tombstone; late writes into a "past" bucket must be caughtPlatform team maintaining bucket-edge logic
Buckets + in-memory timing wheelDurable buckets for correctness, wheel for sub-second firingWheel is volatile — rebuilt from buckets on failover (~1–3s)Platform team — two representations to keep consistent

Staff default: Start on Postgres with SKIP LOCKED if the target is <5K jobs/s — it's a legitimate answer, and saying so is a Staff-level signal. Move to time-bucketed partitions + per-shard timing wheels when either sustained dispatch passes ~5K/s or pending timers pass ~100M rows.

The subtle bug at bucket edges: a timer created at 12:04:59.900 for run_at = 12:05:00.100 lands in bucket 12:05, which the shard owner may have already loaded. Fix: the API writes the timer and notifies the current shard owner if run_at < now + load_horizon, so the owner inserts it into its wheel directly. If the notify fails, the owner's next 30s refresh rereads the bucket. Worst-case extra lag: 30s — so short-delay timers need the direct path.

3.3 Fault Line 3: Push vs Pull Dispatch#

The tension: Who decides which worker runs what?

StrategyWhat WorksWhat BreaksWho Pays
Push (scheduler → worker RPC)Lowest latency; scheduler knows exactly where each run isScheduler must track worker health and capacity; slow workers back-pressure the schedulerPlatform — scheduler becomes a worker-fleet manager
Push via HTTP callback to job ownerLanguage-agnostic; teams just expose an endpointRetries against a down endpoint pile up; callback timeouts vs long jobsJob teams with flaky endpoints; platform with a retry storm
Pull from ready queue with visibility leaseWorkers self-balance; dead workers recover by lease expiry; scheduler knows nothing about workersLease tuning: too short → duplicates of slow jobs; too long → slow recoveryJob teams tune their lease; platform sets bounds

Staff default: Pull with visibility leases for anything longer than ~1s. HTTP callback push is acceptable for short webhooks with a strict 10s timeout, circuit breaker per target, and per-target concurrency caps.

Lease sizing rule: lease = p99_job_duration × 1.5, heartbeat every lease / 3, and a heartbeat that fails twice means the worker stops working (it can no longer prove ownership). A 6-hour batch uses a 5-minute lease with heartbeats — never a 6-hour lease.

3.4 Fault Line 4: Precision vs Herd Smoothing#

The tension: Owners type 0 0 * * * out of habit, not need. Infrastructure pays for that habit at midnight.

Diagram: 3.4 Fault Line 4: Precision vs Herd Smoothing
StrategyWhat WorksWhat BreaksWho Pays
Exact time for everyoneMatches what owners typed100×+ peak/average; downstream DBs melt; retries re-herdPlatform + every shared downstream at midnight
Random jitter per occurrenceSmooths loadJob fires at a different time every day; owners can't reason about itJob owners debugging "why did it run at 00:03 today and 00:01 yesterday?"
Deterministic jitter: hash(job_id) % toleranceSmooth and stable per jobOwners must declare tolerance; some jobs genuinely need :00Platform maintains the review process for tolerance: 0
Admission cap per secondHard ceiling protects the queueExcess slides later; long tails for low-priority jobsLow-priority job owners

Staff default: Deterministic jitter with tolerance_s defaulting to 300 for cron and 0 for delayed timers (they're already naturally spread), plus a per-second admission cap as a backstop. tolerance: 0 requires a reason field and is reviewed.

"I'd rather fire every job at a slightly surprising time consistently than fire all of them on time into a dead database."

3.5 Fault Line 5: Central Platform vs Team-Owned Schedulers#

The tension: A central scheduler gives governance, visibility, and one "did it run?" answer. It also becomes a shared dependency and a queue of feature requests.

StrategyWhat WorksWhat BreaksWho Pays
Every team runs its own (CronJobs, Quartz, Airflow)Autonomy; no platform bottleneck6 schedulers, 6 failure modes, no org-wide view; nobody knows about missed runsIncident responders who can't answer "what didn't run last night?"
One central scheduler, platform-ownedUniform guarantees, audit, quotas, one dashboardPlatform becomes a bottleneck; blast radius is every teamPlatform team on-call; all teams during its outages
Central timer service + team-owned executionPlatform owns firing; teams own workers, code, DLQsNeeds a clear contract and per-team quotasBalanced — the contract is the work

Staff default: Central timer/dispatch service, team-owned execution. The platform publishes one SLO (fire-lag p99), per-team quotas, and a registration contract (owner, pager, idempotency class, tolerance, retry policy). Teams own worker pools, DLQs, and success-rate alerting.

"The platform pager rings when timers fire late. The team pager rings when their jobs fail. If those are the same pager, the platform team will burn out in a quarter and the scheduler will get rewritten by someone angrier."


4. Failure Modes & Operational Reality#

4.1 The Zombie Shard Owner — Double Dispatch of a Whole Minute#

t=0:       Node 7 owns shards 384–447 (lease revision 90211)
t=+1s:     Node 7 enters a 22s stop-the-world GC pause
t=+10s:    Lease TTL expires; etcd revokes node 7's lease
t=+11s:    Node 3 acquires shards 384–447 (lease revision 90344), loads next 5 min
t=+12s:    Node 3 fires 12:00:00–12:00:12 occurrences (8,400 runs) → ready queue with token 90344
t=+23s:    Node 7 resumes. Its in-memory wheel still has the same slots. It fires 8,400 runs with token 90211
t=+23s:    WITHOUT fencing: 8,400 duplicate dispatches → 8,400 duplicate reminder emails
t=+23s:    WITH fencing: dispatch writer rejects token 90211 < 90344 for those shards → 0 duplicates
t=+24s:    Node 7 attempts renewal, finds lease gone, drops its wheel, rejoins as idle

Detection: scheduler.dispatch_rejected_stale_token_total (should be rare; a spike means a zombie or clock/lease bug), scheduler.shard_owner_changes_total, runs.duplicate_total measured at run-history insert conflicts.

Blast radius: Up to shards_owned × dispatch_rate × pause_duration — one node's whole slice for the length of the pause.

Mitigation: Fencing tokens checked at the dispatch sink; node self-checks lease validity against a local monotonic clock before each batch (lease_expiry_local = acquire_time + TTL − safety_margin(2s)) and stops firing if past it.

Prevention: GC tuning (sub-100ms pauses), alarm on pauses > 1s, and never rely on wall-clock comparison for ownership.

Owner: Platform scheduler team.

4.2 The Midnight Herd — Plus the Retry Echo#

t=00:00:00  142,000 occurrences due (hourly + daily + all "0 0 * * *")
t=+1s       Ready queue receives 142K messages; team worker pools pull at their own rates
t=+4s       38 teams' jobs open DB connections to shared-reporting Postgres (max_connections 500)
t=+6s       Connection pool exhausted; 60% of jobs fail with "too many connections"
t=+36s      Retry policy (30s fixed, no jitter) re-dispatches 85K failures at the same instant
t=+40s      Second wave fails identically; third wave queued at t+66s
t=+3min     Reporting DB CPU pinned; unrelated dashboards timing out for all users

Detection: scheduler.due_per_second histogram (look at max, not mean), runs.failure_rate by downstream dependency, scheduler.retry_dispatch_per_second correlated with first-wave failures.

Blast radius: Every job touching the shared downstream, plus the downstream's other users.

Mitigation (in-incident): Enable admission cap (e.g., 2K dispatch/s), pause retries for the affected job group, raise retry backoff.

Prevention: Tolerance windows by default; exponential backoff with ±20% jitter; per-downstream concurrency tags (resource: reporting-db, max_concurrency: 50) enforced by the scheduler.

Owner: Platform for jitter/admission defaults; the reporting-DB owner for declaring its concurrency budget.

4.3 Silent Miss — The Recurring Job That Stopped Recurring#

The most dangerous scheduler failure produces no errors. A recurring job computes its next occurrence after dispatching the current one. If that "write next timer" step fails — a transient timer-store error, a deploy that killed the process between dispatch and write — the chain breaks. Nothing fails. Nothing alerts. The job simply never runs again.

Day 0 02:00  nightly-gdpr-purge dispatched; process restarts before writing Day 1 timer
Day 1 02:00  nothing due; nothing fires; no error anywhere
Day 30       Privacy audit asks for purge logs. Last run: Day 0.

Detection: A reconciler that, every 5 minutes, checks every active recurring definition has a pending timer with run_at ≤ next_expected + tolerance. Metric: scheduler.definitions_without_next_timer — alert if > 0 for 10 minutes. And a per-job "dead man's switch": job.last_success_age_seconds > 2 × period pages the owner.

Mitigation: Write the next timer before or atomically with dispatching the current one (transactional outbox or a single-partition batch write).

Owner: Platform owns the reconciler; job owners own the dead man's switch threshold.

4.4 Worker Lease Too Short — The Job That Runs Four Times#

A report job's p99 grows from 3 to 7 minutes after its dataset doubles. Its visibility lease is 5 minutes with no heartbeat. At minute 5 the lease expires and a second worker starts the same run. At minute 10, a third. Four copies run concurrently, each hammering the warehouse, each slower because of the others.

Detection: worker.lease_expired_while_running_total, runs.concurrent_attempts_max > 1 for the same run_id.

Mitigation: Mandatory heartbeats; worker stops work if heartbeat fails (it has lost ownership); Complete with a stale fencing token is rejected.

Owner: Job team (lease config), platform (heartbeat protocol and rejection of stale completions).

4.5 Catch-Up Storm After Scheduler Outage#

The scheduler is down 3 hours. On recovery, catch-up policy all fires 3 hourly occurrences for 40K jobs at once — 120K dispatches, most useless (three copies of "refresh the cache").

Prevention: Default catch-up policy latest (run the most recent missed occurrence once, record the others as skipped_outage), and catch-up dispatch rate-limited to 20% of normal capacity. Jobs that need every occurrence (hourly billing aggregation) declare catch_up: all and are replayed in order, throttled.

Owner: Platform (defaults), job owners (declaring all).

4.6 Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Zombie shard ownerdispatch_rejected_stale_token_total spikeOne node's shards × pause lengthFencing tokens; local lease-expiry self-checkPlatform
Shard owner crashshard_unowned_seconds > 15sFire-lag on ~6% of shards (64/1024)Lease expiry → reassignment in ~10–12sPlatform
Midnight herddue_per_second max ≫ capacity; downstream errorsShared downstreamsTolerance windows; admission cap; jittered retriesPlatform + downstream owner
Broken recurrence chaindefinitions_without_next_timer > 0One job, silently, foreverReconciler; atomic next-timer writePlatform
Lease too shortlease_expired_while_running_totalOne job × N concurrent copiesHeartbeats; stale-token rejectionJob team
Timer store partition unavailabletimer_store_read_errors for bucketTimers in affected partitions lateReplica reads (RF=3, LOCAL_QUORUM); lag alertPlatform + storage team
Ready queue backlogready_queue_lag_seconds > 60All teams sharing the lanePriority lanes; per-team quotasPlatform
Team job failure spikeruns.failure_rate{team} > 5%One teamTeam DLQ + team pagerJob team
Clock skew on scheduler nodesntp_offset_ms > 100Early/late fires on one nodeNTP alarms; fire based on shard-owner clock onlyInfra

5. Evaluation Rubric#

5.1 Level-Based Signals#

DimensionSenior (L5)Staff (L6)Principal (L7)
GuaranteeClaims exactly-onceAt-least-once + deterministic run_id; explains the side-effect/ack windowIdempotency class as a registration gate; unsafe jobs need director sign-off
StorageJobs table + indexKnows the SKIP LOCKED ceiling; time-bucketed partitions past ~5K/sPicks build vs buy (Temporal Cloud, EventBridge Scheduler) with a TCO and migration cost
CoordinationLeader electionSharded leases + fencing tokens; quantifies failover gapCell-based scheduler fleets; limits blast radius of a bad shard map to one cell
LoadAutoscale workersNames the :00 herd; deterministic jitter; admission capOrg policy: precision is opt-in; "herd calendar" reviewed with capacity planning
Failure ownershipScheduler retriesPlatform fire-lag SLO vs team success SLO; per-team DLQShared scorecard; teams with chronic failure rates lose platform-paging rights
Silent failureNot mentionedReconciler for broken recurrence chains; dead man's switchMakes "every scheduled occurrence is accounted for" an auditable compliance control

5.2 Strong Hire Signals#

SignalWhat It Sounds Like
Chooses the failure"A missed fire is silent; a duplicate is loud. I'll take duplicates and make them harmless."
Fencing, not just election"The lease revision is the fencing token. A paused node's dispatches bounce at the sink."
Sees the herd"Average is 580/s; midnight is 140K in a second. Autoscaling reacts in minutes; I smooth in advance."
Silent-miss detection"The reconciler checks every recurring definition has a next timer. That's the only way to catch a chain that just stopped."
Ownership contract"Registration requires owner, pager, idempotency class, and tolerance. The platform pager never rings for job logic."

5.3 Lean No-Hire Signals#

SignalWhy It Misses the Bar
"Exactly-once" without qualificationDoesn't understand the side-effect/ack window — the core of the problem
Single leader polling one table at any scaleNo ceiling awareness; no failover gap analysis
Random UUID per runMakes duplicates indistinguishable from distinct runs; idempotency impossible
Retries with fixed delay and no jitterRe-creates the herd 30 seconds later
Platform owns job failuresOwnership model that burns out the platform team in a quarter

5.4 Common False Positives#

  • Deep cron-expression knowledge ≠ scheduler design. Knowing L, W, and # modifiers says nothing about ownership under failure.
  • Naming Temporal ≠ understanding durable execution. Can they explain why activities must be idempotent in Temporal?
  • Raft/Paxos detail ≠ correct coordination. A candidate can explain Raft perfectly and still forget that a paused leader keeps writing without fencing.
  • Elaborate priority queues ≠ fairness. Five priority levels without per-tenant quotas still let one team starve the rest within a level.

6. Interview Flow & Pivots#

6.1 Typical 45-Minute Shape#

PhaseTimeGoal
Framing0–3 minName 3 intents; commit to timers + cron; state "missed is worse than duplicate"
Entities + API3–5 minDefinition, Timer, Run; deterministic run_id; pull protocol
High-level design5–12 minBucketed timers → leased shards → ready queue → pull workers
Transition12 minOffer: guarantee, zombie owners, herd
Deep dives12–38 minGuarantee → fencing → herd → silent miss
Org & evolution38–43 minSLO split, registration contract, migration from crontabs
Wrap-up43–45 minSummary; what you'd build next; what you won't build

6.2 How Interviewers Pivot — And What They're Testing#

Interviewer PivotWhat They're TestingWhere to Go
"The worker crashed after charging the card."Side-effect semanticsrun_id idempotency; at-least-once; unsafe class
"The leader is paused for 20 seconds."Distributed coordination depthFencing tokens; local lease-expiry self-check
"Everyone schedules at midnight."Load shaping vs reactive scalingTolerance windows; deterministic jitter; admission cap
"One team submits 10M jobs for a backfill."Multi-tenancyPer-team quotas; separate backfill lane; weighted fair dequeue
"The scheduler was down for 3 hours."Recovery semanticsCatch-up policy latest vs all; throttled replay
"How would you know if a job silently stopped?"Observability maturityReconciler; dead man's switch; definitions_without_next_timer
"Could we just use Airflow?"Build vs buy judgmentDifferent intent (DAGs); where Airflow's metadata DB tops out

6.3 What to Deliberately Skip#

TopicWhy L5 Goes HereWhat L6 Says Instead
Cron syntax parsingConcrete and safe"Standard library. Next-fire computed at write time."
Worker container orchestrationFeels complete"Workers are team-owned; they run on our normal compute platform."
Admin UIEasy to describe"CRUD over definitions and run history. Not interesting here."
Exact queue productBikeshed"Any durable partitioned queue. Kafka because we run it."
Payload storageDetails"Payloads > 64KB go to blob storage; the timer holds a reference."

6.4 Follow-Up Questions to Expect#

  1. "How do you cancel a timer that's already loaded into a node's in-memory wheel?"
  2. "A job is scheduled at 02:30 local time. What happens on the DST spring-forward day?"
  3. "How do you rebalance shards when you add scheduler nodes without double-firing?"
  4. "How do you support a job that must never run concurrently with itself, across regions?"
  5. "What's your SLO, and what burns the error budget?"
  6. "How do you migrate 2,000 existing crontab entries without a missed run?"
  7. "How would this work multi-region — does a timer created in us-east fire if us-east is down?"

7. Active Drills#

Drill 1: The Opening#

Prompt: "Design a distributed job scheduler."

Staff Answer

"Before I draw anything: are we replacing cron for recurring infra jobs, running a high-volume delayed-timer service, or orchestrating multi-step workflows? The first needs auditability, the second throughput, the third durable state — they're incompatible optimizations. I'll build the timer-and-cron core: a cron occurrence is just a timer that schedules its successor. Workflows sit on top.

Constraints: 50M timers/day, 1B pending, fire-lag p99 < 5s for timers and < 60s for default cron. At-least-once dispatch with idempotent jobs — because a missed run is silent and a duplicate is loud. I'll walk through: storage partitioned by time → shard ownership with leases → dispatch and worker leases → the midnight herd → who owns failures."

Why this is L6:

  • Separates intents that share a name but not a design
  • Commits to a guarantee and justifies it with detectability
  • The outline is a sequence of decisions, not a component list

What L7 adds:

  • Asks how many schedulers the org already runs and makes consolidation part of the scope
  • Frames the SLO split (fire-lag vs job success) as an org contract in the first five minutes
❌ Common L5 Trap

"I'll have a jobs table with run_at, a poller that selects due jobs every second, and a pool of workers. For HA, a standby poller with leader election."

Why this misses: Correct for 1K jobs/s on a good day. It hasn't chosen a guarantee, so "worker dies after side effect" has no answer, and leader election without fencing double-fires during pauses.


Drill 2: Core Mechanic — The Worker Dies Mid-Job#

Prompt: "A worker picks up a job that charges a customer, charges them, then crashes before reporting success. What happens?"

Staff Answer

"The worker held a 5-minute visibility lease with heartbeats. Heartbeats stop, the lease expires, and the run becomes available again as attempt 2 with the same run_id — which is hash(job_id, scheduled_at), so every attempt of this occurrence shares it. The charge call passes run_id as the payment provider's idempotency key; attempt 2 gets 'already processed' and completes. If the target doesn't support idempotency keys, the job writes INSERT INTO charges_done(run_id) ON CONFLICT DO NOTHING in the same transaction as recording the charge intent. The scheduler can't close this window — only the job can — which is why idempotency class is declared at registration."

Why this is L6:

  • Names the window and admits the scheduler can't close it
  • Deterministic run ID is the mechanism that makes idempotency possible
  • Pushes responsibility to the right owner with a contract, not a hope

What L7 adds:

  • Platform ships an SDK helper (with run_guard(run_id):) so 300 teams don't each reinvent the dedup table
  • Tracks runs.attempts > 1 per team as a quality metric on the platform scorecard

Drill 3: "Leader Election" — Make It Concrete#

Prompt: "You said each shard has one owner. The owner is paused by GC for 20 seconds. Walk me through it."

Staff Answer

"Lease TTL is 10s; at t+10s etcd expires node 7's lease; node 3 acquires shard 42 with a higher lease revision — say 90344 vs 90211 — loads the next 5 minutes, and fires. At t+20s node 7 wakes with a full in-memory wheel. Two defenses. First, node 7 checks its own lease validity against a local monotonic clock before each dispatch batch — acquired_at + TTL − 2s — sees it's past, and drops the wheel. Second, in case that check races, every dispatch carries the lease revision as a fencing token; the dispatch sink keeps max_token[shard] and rejects anything lower. Node 7's writes bounce. Worst case: zero duplicates, ~12s fire-lag for shard 42's timers during the handover."

Why this is L6:

  • Goes past election to fencing — the part that actually prevents double-firing
  • Uses monotonic time, not wall clock, for self-checks
  • Quantifies the handover lag instead of claiming zero

What L7 adds:

  • Runs quarterly game days that inject 30s pauses into a scheduler node in production
  • Tracks dispatch_rejected_stale_token_total as a fleet-wide leading indicator of GC and VM-migration problems

Drill 4: Dependency Down — Timer Store Unavailable#

Prompt: "The Cassandra cluster holding timers loses a rack. What happens to firing and to new timer creation?"

Staff Answer

"RF=3 across racks with LOCAL_QUORUM means losing one rack keeps both reads and writes available. If we lose quorum for some partitions: new timer writes for those partitions fail — the API returns 503 with Retry-After, and callers retry; that's fail-closed on creation because accepting a timer we can't persist is the one unforgivable sin. Firing continues from the in-memory wheels for the next 5 minutes of already-loaded timers, so the fire path degrades gracefully. Beyond the horizon, affected shards show fire-lag growth; scheduler.bucket_read_errors pages platform, and we alert when the wheel's loaded horizon drops below 60s."

Why this is L6:

  • Distinguishes the create path (fail-closed) from the fire path (degrade on in-memory horizon)
  • Explains why the 5-minute wheel is also a resilience buffer, not just a latency optimization

What L7 adds:

  • Sets the wheel horizon based on the storage team's recovery SLO (if storage restores in 15 min p99, horizon should be ≥ 15 min)
  • Negotiates a joint SLO with the storage platform team rather than assuming availability

Drill 5: Hot Key — One Job, a Million Occurrences#

Prompt: "A marketing team schedules one 'campaign send' job that fans out into 20M per-user timers all at 09:00."

Staff Answer

"Two problems: storage skew and dispatch herd. Storage: shard is hash(timer_id), not hash(job_id), so 20M timers spread across 1,024 shards — ~20K each. Dispatch: 20M in one second is 400× our normal peak. The campaign is a bulk intent; it shouldn't be 20M timers. It should be one timer that fires a fan-out job, which streams the audience through a rate-limited producer at, say, 50K/s — 20M in ~7 minutes. That's also what the downstream notification provider can absorb. If they genuinely need per-user local-time delivery (09:00 in each user's zone), that spreads naturally across 24 hours."

Why this is L6:

  • Refuses to model a bulk operation as 20M independent timers
  • Ties dispatch rate to downstream capacity, not scheduler capacity

What L7 adds:

  • Creates a separate "bulk" lane and quota class with its own cost center, so campaigns can't erode the timer SLO for transactional users
  • Coordinates with the Notification Systems platform so campaign rate limits live in one place

Drill 6: Multi-Tenant — The 10M-Job Backfill#

Prompt: "Team A enqueues 10M backfill jobs at 10:00. Team B's hourly jobs are now 40 minutes late."

Staff Answer

"Our ready queue lane is shared and FIFO, so A's 10M is ahead of B. Fix: per-team quotas at dispatch — each team has a guaranteed share (say, dispatch weight proportional to its reserved capacity) with borrowing when idle. The scheduler dispatches with deficit round robin across teams rather than global FIFO. And backfills go to a separate bulk lane with lower priority, admitted at most 20% of dispatch capacity when the primary lanes are busy. Team A's backfill takes 3 hours instead of 40 minutes; Team B's hourly jobs stay on time. A signed up for that at registration by choosing the bulk lane."

Why this is L6:

  • Identifies FIFO as the bug; replaces with weighted fair dequeue
  • Makes the backfill owner pay in latency for their volume

What L7 adds:

  • Quotas map to budget: teams buy reserved dispatch capacity; overage is billed back
  • Publishes a lane taxonomy (P0 transactional / P1 standard / P2 bulk) org-wide and forbids per-team custom lanes

Drill 7: Build vs Buy#

Prompt: "Why not use Temporal / EventBridge Scheduler / Airflow instead of building?"

Staff Answer

"Default is buy. EventBridge Scheduler covers one-shot and recurring schedules with managed durability; Temporal covers durable timers inside workflows; Airflow covers DAGs. I'd build only if: (1) volume or cost breaks managed pricing — at 50M timers/day, per-invocation pricing may run to thousands of dollars a month, which is fine, while at 5B/day it becomes a six-figure line item that justifies a team; (2) we need org-wide governance features — per-team quotas, tolerance-window policy, fencing into our own queue — that the managed service doesn't expose; or (3) latency/precision below what managed services guarantee. Otherwise building a scheduler is a 3–5 engineer team forever, and the first year is spent rediscovering the failure modes in Section 4."

Why this is L6:

  • States a concrete threshold for building rather than a preference
  • Counts the ongoing team, not just the build

What L7 adds:

  • Evaluates lock-in: the job contract (run_id, idempotency class, tolerance) is ours; the engine behind it is swappable
  • Writes the exit plan from the managed service before adopting it

Drill 8: Policy Change Without Outage — Changing Retry Defaults#

Prompt: "You want to change the default retry policy from 3 fixed retries at 30s to 5 exponential retries with jitter. 4,000 jobs use the default."

Staff Answer

"Changing a default silently changes 4,000 jobs' behavior — some of which may not be idempotent enough for 5 attempts. Rollout: (1) shadow — compute what the new policy would do and log would_retry_at; (2) publish the list of affected jobs to owners with a 2-week opt-out; (3) enable for idempotency_class = natural | keyed jobs first, 5% → 25% → 100% over a week, watching runs.duplicate_side_effect reports and downstream error rates; (4) unsafe jobs keep at-most-once regardless. Policy is versioned; each run records the policy version it used so post-incident review knows which rules applied. Rollback is a flag."

Why this is L6:

  • Treats a default change as a fleet-wide behavior change needing staged rollout
  • Uses the idempotency class to scope blast radius

What L7 adds:

  • Sets the org rule that platform defaults change at most quarterly, with a changelog and owner notification — "defaults are an API"

Drill 9: Cost#

Prompt: "Finance says the scheduler costs $40K/month. Where does the money go and what would you cut?"

Staff Answer

"At 1B pending timers and 50M fires/day: timer store (RF=3, ~1.5 TB with replication and tombstones) is likely the largest line; then the ready queue (Kafka, 7-day retention is overkill — dispatch messages need hours); then scheduler nodes (16 × mid-size). Cuts: (1) drop ready-queue retention from 7 days to 24h; (2) TTL timers at run_at + 3d instead of 7; (3) compact cancelled timers faster — if 70% are cancelled, tombstones dominate; (4) move run history older than 7 days to object storage. Realistically 30–40% saving. What I would not cut: RF=3 on the timer store, because a lost timer is the unforgivable failure."

Why this is L6:

  • Knows which component dominates cost and why (tombstones from cancellations)
  • Protects the correctness-critical spend explicitly

What L7 adds:

  • Shows cost per 1M timers as a unit metric and charges it back to teams, which changes behavior (teams stop creating-then-cancelling timers every page view)

Drill 10: Multi-Region#

Prompt: "We're going active-active in two regions. Timer created in us-east — does it fire if us-east is down?"

Staff Answer

"Two options. (A) Timers are region-homed: created and fired in their home region; if the region is down, they're late until it recovers — simple, no cross-region coordination, fire-lag SLO excludes regional outages. (B) Timers replicate to the peer region asynchronously; on declared regional failover, the peer takes ownership of the failed region's shards. Risk: replication lag means the last ~1–5s of timers may be missing, and anything already dispatched in the dead region may dispatch again — so idempotency carries us. I'd default to A for most jobs and offer B as a tier for jobs whose owners can justify it, like payment retries. Failover is a declared operator action, not automatic, to avoid split-brain double-firing on a network partition."

Why this is L6:

  • Keeps failover explicit, avoiding automated split-brain
  • Uses the at-least-once contract to make cross-region takeover safe

What L7 adds:

  • Prices tier B (2× storage + cross-region transfer) and makes owners opt in with budget
  • Aligns regional failover runbooks with the org's broader DR program so the scheduler isn't the only system that fails over

8. Deep Dive Scenarios#

Deep Dive 1: Peak-Traffic Incident — New Year's Midnight#

Context: It's 00:04 UTC on January 1. Fire-lag p99 is 11 minutes. Ready-queue lag is climbing. Three product teams have escalated: renewal emails, subscription billing retries, and a partner data export are all late. You're pulled in.

Questions to Surface First:

  • Is the bottleneck firing (scheduler nodes, timer reads) or consumption (worker pools, downstreams)?
  • Which lanes are late — P0 transactional or bulk?
  • Are failures from the first wave being retried into a second wave?
  • Which downstream is saturated, and whose jobs are hitting it?

Typical L5 Approach: Scales worker pools 3× and scheduler nodes 2×. Worker autoscaling takes 4 minutes; when the new workers arrive they hit the already-saturated billing DB and increase failures.

Staff Approach: Separates fire-lag from consumption lag within 2 minutes. Finds scheduler fire-lag is fine (2s) but ready-queue lag is 11 min because billing-retry jobs are failing against a saturated DB and being retried at fixed 30s intervals. Pauses the retry lane for billing, applies admission cap to bulk lane, and lets P0 drain.

Principal Approach: Treats this as a predictable calendar event that the org failed to plan for. Establishes a "herd calendar" (midnight UTC, month-end, New Year, Black Friday) reviewed by capacity planning 2 weeks prior, and makes tolerance windows mandatory for new jobs so next year's midnight is structurally smaller.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Split scheduler.fire_lag vs ready_queue.lag by lane. Identify the failing downstream via runs.failure_rate{dependency}.
TriageBilling retries failing at 70% against billing DB (connection saturation). Retry policy is fixed 30s — every failure returns as a new wave.
Quick fixPause billing-retry lane dispatch for 5 min; set bulk lane admission to 10%; let billing DB recover; resume with exponential backoff + jitter.
GuardrailsPer-dependency concurrency cap: billing-db max_concurrency 200. Scheduler withholds dispatch above cap instead of letting jobs fail.
Post-mortemWhy were 180K jobs on 00:00:00? Why did billing retries have no jitter? Why wasn't New Year on a capacity calendar?

Metrics to Watch: scheduler.fire_lag_seconds{lane}, ready_queue.lag_seconds{lane}, runs.failure_rate{dependency}, scheduler.retry_dispatch_per_second, billing_db.active_connections

Organizational Follow-up: Tolerance windows default-on for all new recurring jobs; a one-quarter migration campaign to add tolerance to the top 500 midnight jobs; dependency concurrency tags required for any shared database.

Ownership Question: "Who decides to pause the billing retry lane at 00:05 on New Year's?" Staff answer: The platform on-call, under a pre-approved runbook that allows pausing any non-P0 lane for up to 15 minutes without escalation. Billing's on-call is notified automatically. Longer pauses require billing's sign-off.

Key Takeaway: "Scaling consumers into a saturated downstream makes the herd worse. Shape the demand — don't chase it."

What clears the Staff bar:

  • Distinguishes firing lag from consumption lag immediately
  • Identifies retry echo as the amplifier
  • Converts a one-night fix into a default policy

Deep Dive 2: Silent Failure — The Purge That Stopped#

Context: A privacy audit finds the nightly GDPR deletion job last ran 34 days ago. No alerts ever fired. Legal is involved. The job owner says "the scheduler should have run it."

Questions to Surface First:

  • Does the definition still exist and is it enabled?
  • Is there a pending timer for its next occurrence?
  • What happened at the last successful run — any deploy, restart, or store error?
  • Who owns "did it run" — platform or the privacy team?

Typical L5 Approach: Re-runs the job, finds a bug in the recurring-job code path where the next timer write can be skipped on process restart, fixes it, adds a unit test.

Staff Approach: Fixes the bug and the class of bug. Makes next-timer creation atomic with dispatch (single batch write), adds the reconciler (definitions_without_next_timer), and adds a per-job dead man's switch: last_success_age > 2 × period pages the job owner. Scans all 200K definitions for other broken chains — finds 312.

Principal Approach: Reclassifies scheduling of compliance jobs as a controlled process. Compliance-tagged jobs require a dead man's switch, an owner of record, and a monthly attestation report generated from run history. The scheduler becomes evidence for auditors, not just infrastructure.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (0–5 min)Manually trigger the purge with a fresh run_id; confirm it completes. Notify legal of the gap window.
TriageQuery run history: last dispatch 34 days ago; no timer exists for any later occurrence. Deploy logs show a restart 3s after that dispatch.
Quick fixBackfill next timers for every active definition missing one (312 found).
GuardrailsReconciler every 5 min; alert definitions_without_next_timer > 0 for 10m. Dead man's switch per job with owner-set thresholds.
Post-mortemWhy was "write next" a separate step after dispatch? Why did no one-layer-up check exist? Why did the job owner assume the platform owned "did it run"?

Metrics to Watch: scheduler.definitions_without_next_timer, job.last_success_age_seconds{job}, reconciler.repairs_total

Organizational Follow-up: Registration requires a max_silence field for any job tagged compliance or financial. Platform publishes a weekly "jobs that didn't run" report to each team.

Ownership Question: "Who is accountable that the purge runs?" Staff answer: The privacy team owns that the purge succeeds and owns the dead man's switch. The platform owns that every enabled definition has a next timer — and the reconciler proves it. Both failed here; both get action items.

Key Takeaway: "Schedulers fail silently by not doing things. The only defense is something that knows what should have happened and checks."

What clears the Staff bar:

  • Fixes the class of bug, not the instance, and scans for siblings
  • Designs detection for absence of events
  • Splits accountability explicitly

Deep Dive 3: Large-Customer Onboarding — The Tenant With 300M Timers#

Context: A new enterprise tenant (a logistics company) will create a timer per package-scan SLA — 300M pending timers, 15K creates/s, 80% cancelled when the scan arrives on time. That's 3× your current pending volume. Sales signed the contract; go-live is in 6 weeks.

Questions to Surface First:

  • What precision do they need — ±1s or ±1 min?
  • What's their cancel pattern — cancel by ID, or by prefix (all timers for package X)?
  • Do their fires hit our infrastructure or their webhooks?
  • What SLO did sales promise?

Typical L5 Approach: Adds timer store nodes to triple capacity and more scheduler nodes.

Staff Approach: Recognizes cancellation as the dominant cost: 80% of 15K/s = 12K cancels/s, each a tombstone. Proposes storing their timers with run_at-bucketed partitions and a short TTL so tombstones age out with the bucket, or a cancel-by-marker (check cancelled_set at fire time) which avoids tombstones entirely. Puts the tenant in its own shard range (a cell) so their load can't affect existing tenants' fire-lag. Webhook dispatch gets a per-tenant concurrency cap and circuit breaker.

Principal Approach: Uses this deal to establish tenant cells as the scaling unit and a pricing model where large tenants pay for dedicated capacity. Pushes back on sales that SLOs must be set with platform before contracts are signed.

Staff Approach — Full Reasoning
PhaseWhat to Do
Immediate (week 1)Load-model: 300M × 300B = 90 GB raw, ~270 GB replicated; tombstones: 12K/s × 86,400 ≈ 1B/day.
TriageTombstone-heavy partitions degrade reads; fire-time reads for their buckets would scan 4× more tombstones than live rows.
Quick fixCancel-by-marker: cancellations write to a compact cancelled(run_id) set with TTL; at fire time, check the set before dispatch. No tombstones in the timer partitions.
GuardrailsTenant cell: dedicated shard range + scheduler nodes; per-tenant quota on creates (20K/s) and dispatch (10K/s).
Post-mortem-style reviewLoad test at 2× in staging for 72h; game day killing a scheduler node in the tenant's cell.

Metrics to Watch: timer_store.tombstones_per_read, scheduler.fire_lag_seconds{tenant}, api.create_throttled_total{tenant}, webhook.circuit_open{tenant}

Organizational Follow-up: Sales-engineering intake form for any tenant > 10% of current volume; platform sign-off on SLOs before contract.

Ownership Question: "Who pays for the dedicated cell?" Staff answer: The tenant, via contract pricing. If sales discounted it away, the business unit that owns the account absorbs the cost — not the platform budget.

Key Takeaway: "For timer services, cancellations — not fires — are often the dominant load. Model the cancel path first."

What clears the Staff bar:

  • Finds the non-obvious dominant cost (tombstones)
  • Isolates the big tenant structurally
  • Moves SLO-setting upstream of the sales process

Deep Dive 4: Post-Mortem — 11,000 Duplicate Invoice Emails#

Context: Customers received duplicate invoice emails — 11,000 of them — during a scheduler rebalance when two nodes were added. The billing team is furious. You're writing the post-mortem.

Questions to Surface First:

  • Did two nodes own the same shard at once, or did a shard's timers dispatch twice from one owner?
  • Were fencing tokens checked at the sink for that path?
  • Why didn't the email job's idempotency catch it?

Typical L5 Approach: Finds a race in the rebalance code where the old owner released a shard after the new owner acquired it; adds a lock.

Staff Approach: Finds two failures, not one. (1) The rebalance released-then-acquired correctly, but the new owner re-loaded the current minute and re-fired timers the old owner had already dispatched, because "already dispatched" was checked in a cache, not run history. (2) The email job used a random message ID, not run_id, as its dedup key — so even with duplicates, it should have been safe, and wasn't. Fixes both: new owner checks run history (INSERT … ON CONFLICT (run_id)) before dispatching; the email job adopts run_id.

Principal Approach: Adds an org-level control: jobs with external side effects (email, SMS, payments) cannot register without idempotency_class: keyed, and the platform verifies it in a pre-production harness that dispatches every job twice and diffs side effects.

Staff Approach — Full Reasoning
PhaseWhat to Do
ImmediatePause rebalancing; freeze shard map.
TriageCorrelate duplicates by run_id: every duplicate pair has the same run_id, different dispatch nodes, 1–4s apart, all in shards moved during rebalance.
Quick fixDispatch sink enforces unique (run_id, attempt=1); the second first-attempt dispatch is dropped.
GuardrailsRebalance protocol: old owner stops firing, flushes, records high-water mark (last_dispatched_slot); new owner starts at high-water + 1.
Post-mortemTwo layers of defense both failed. Blameless; the fix is structural — dedup at the sink plus mandatory keyed idempotency for side-effecting jobs.

Metrics to Watch: runs.duplicate_first_attempt_total, scheduler.rebalance_duration_seconds, dispatch_rejected_stale_token_total

Organizational Follow-up: Billing gets a customer apology template; platform adds the "dispatch twice" test harness to CI for side-effecting jobs.

Ownership Question: "Whose bug was it?" Staff answer: Both. The platform violated its own 'duplicates are rare' promise during rebalance; billing violated the 'duplicates are harmless' contract. The post-mortem assigns an action item to each and doesn't let either hide behind the other.

Key Takeaway: "At-least-once works only if both halves hold: the platform keeps duplicates rare, the job keeps them harmless."

What clears the Staff bar:

  • Looks for the second failure behind the first
  • Uses run_id correlation as the forensic tool
  • Refuses single-team blame for a two-contract failure

Deep Dive 5: Multi-Region Expansion#

Context: The company is launching an EU region for data residency. EU customers' jobs must run in the EU, and timers must never be stored outside it. Leadership also wants "the scheduler to survive a regional outage."

Questions to Surface First:

  • Is residency about the timer payload, the job execution, or both?
  • Does "survive a regional outage" apply to EU jobs, which can't fail over to the US?
  • What's the fire-lag SLO during a regional outage?

Typical L5 Approach: Deploys a second scheduler cluster in the EU and replicates timers bidirectionally for failover.

Staff Approach: Notes that bidirectional replication violates residency. Deploys independent regional schedulers; each timer is homed by its tenant's region. EU survives EU-zonal failures (3 AZs) but not a full EU-region outage — and says so explicitly. For non-residency tenants, offers an async-replicated tier with declared failover. Makes "home region" part of the job definition.

Principal Approach: Takes the residency conflict to leadership as a decision, not an engineering problem: "EU jobs can be resident or regionally-failover-capable, not both, unless we add a second EU region at ~$X/month." Makes the region topology a product-level choice documented in the data-residency standard.

Staff Approach — Full Reasoning
PhaseWhat to Do
DesignRegion is a property of the tenant; API routes timer creation to the home region; cross-region creates are proxied, never stored locally.
Failure modelAZ failure: survive (RF=3 across AZs, shard leases re-home in ~10s). Region failure: EU timers late until recovery; US tier-B timers fail over by declared action.
GuardrailsResidency check at API: refuse to persist timers whose tenant home ≠ local region. Audit metric residency_violation_attempts_total.
CostSecond region ≈ +100% scheduler infrastructure for the EU slice; tier-B replication ≈ +60–80% storage + transfer.
RolloutMigrate EU tenants' pending timers via dual-write for 7 days, cut over reads, then delete US copies.

Metrics to Watch: scheduler.fire_lag_seconds{region}, replication.lag_seconds, residency_violation_attempts_total

Organizational Follow-up: Legal/privacy sign-off on the residency design; DR runbook updated with "EU is not failover-capable" in bold.

Ownership Question: "Who decides that EU jobs don't survive an EU-region outage?" Staff answer: Not the platform team. It's a product and legal decision, documented with the cost of the alternative. The platform's job is to make the tradeoff visible and priced.

Key Takeaway: "Residency and cross-region failover are in tension. Name the conflict and send it to the people who own the risk."

What clears the Staff bar:

  • Spots the residency vs failover conflict immediately
  • Keeps failover declared, not automatic
  • Prices the alternative instead of promising both

9. Level Expectations Summary#

After studying this case study, you should be able to:

  • Separate cron, delayed timers, and workflow orchestration in the first 2 minutes and commit to one
  • State "at-least-once dispatch, idempotent jobs" and explain the side-effect/ack window that makes exactly-once impossible at the scheduler
  • Explain why run_id must be deterministic (hash(job_id, scheduled_at)), not random
  • Design shard ownership with leases and fencing tokens, and quantify the failover gap (~10–12s)
  • Predict the midnight herd and smooth it with deterministic jitter inside declared tolerance windows
  • Detect the silent miss (broken recurrence chain) with a reconciler and a dead man's switch
  • Draw the ownership line: platform owns fire-lag, teams own job success, DLQs, and their pager
  • Argue build vs buy with a volume threshold and a team-cost estimate

The Bar for This Question#

Mid-level (L4): Designs a jobs table, a poller, and workers. Handles retries. Scales by adding pollers with a lock. Correct on a good day; no answer when the worker dies after the side effect.

Senior (L5): Adds leader election, partitions jobs, uses a queue between scheduler and workers, and mentions idempotency when prompted. Plans capacity for average load. Owns the mechanism well but leaves the guarantee, the herd, and the ownership boundary implicit — so each probe from the interviewer forces a redesign.

Staff+ (L6): Chooses the guarantee first and derives everything from it. Uses fencing tokens, deterministic run IDs, and pull-based leases as a coherent set. Sees the herd before being asked, designs detection for silent misses, and separates platform and job-owner accountability with two SLOs. The interviewer should learn something from the answer.


10. Staff Insiders: Controversial Opinions#

10.1 "Exactly-Once Scheduling" Is a Marketing Term#

ClaimReality
"Our scheduler guarantees exactly-once"It guarantees exactly-once dispatch records at best. Side effects outside its transaction can still repeat.
"We use Kafka transactions, so it's exactly-once"Kafka EOS covers read-process-write within Kafka. An email sent from inside the loop is outside it.
"Temporal is exactly-once"Temporal's workflow state transitions are; activities are at-least-once and must be idempotent — by design.

The Staff position: Say "effectively-once = at-least-once delivery + idempotent effect." Anyone promising more hasn't defined the side-effect boundary.

Why this matters in interviews: Saying "exactly-once" is the single fastest way to get a follow-up you can't answer. Saying "effectively-once, and here's the idempotency key" ends the line of questioning in your favor.

10.2 Most Jobs Don't Need to Run On Time#

JobTyped ScheduleActual Requirement
Nightly report0 0 * * *"Before 07:00 when people read it"
Cache warm0 * * * *"Roughly hourly"
Cleanup of expired sessions*/5 * * * *"Eventually, within a day"
Market-open order release30 13 * * 1-5Truly exact — seconds matter

The Staff position: Precision is a product requirement that owners almost never state. Make them state it; default to tolerance. The rare exact job is easier to protect when it's not competing with 100K habitual midnight jobs.

Why this matters in interviews: It shows you design the demand, not just the capacity.

10.3 The Scheduler Team Should Not Be Paged for Job Failures#

The Staff position: If a platform team's pager carries every product team's job failures, the signal-to-noise collapses within weeks and real platform incidents (fire-lag) get missed in the noise. Route failures to owners at registration. A scheduler team with a clean pager is a scheduler team that notices when the scheduler itself breaks.

Why this matters in interviews: Ownership boundaries are the clearest Staff-level signal in this question — and the one most candidates never mention.

10.4 Postgres Is a Fine Scheduler Until It Isn't — And That Line Is Far Away#

ScalePostgres + SKIP LOCKED
< 1K jobs/s, < 10M pendingExcellent. One store, transactional cancel, SQL for "did it run".
1–5K jobs/s, < 100M pendingWorks with partitioning by run_at day and aggressive autovacuum.
> 5–10K jobs/s or > 100M pendingIndex hot spot and vacuum churn dominate; move to bucketed timers.

The Staff position: Proposing Postgres for a 500 jobs/s scheduler is not a Senior answer — it's a mature one, as long as you name the ceiling and the migration trigger.

Why this matters in interviews: Over-engineering for 50× scale on day one is a lean-no-hire signal. Knowing where the simple design breaks is a hire signal.

10.5 Cron Is Where Compliance Goes to Die#

The Staff position: Retention purges, key rotations, certificate renewals, and regulatory exports are almost always cron jobs, and almost never have a "did it run" alarm. The scheduler is a compliance control whether anyone admits it or not. Treat compliance-tagged jobs as audited controls with dead man's switches and attestations.

Why this matters in interviews: Connecting a scheduler to audit risk shows you think about what the system is for in the business, not just what it does.


11. The Principal Lens (L7)#

Why L7 Sees This Problem Differently#

A Staff engineer designs a scheduler. A Principal engineer notices the organization already has seven — Kubernetes CronJobs in 40 namespaces, two Airflow clusters, a Quartz instance inside the monolith, a Jenkins box, EventBridge rules in three AWS accounts, and a Redis sorted set someone built for cart expiry — and that the real failure mode is not "the scheduler goes down" but "nobody can answer what didn't run last night across the company." The L7 problem is consolidation, governance, and the contract: one way to register time-triggered work, one place to see its history, one set of guarantees, and a deliberate decision about which of the seven to retire and which to leave alone.

The Org-Level Fault Line#

One scheduling platform vs a federation of schedulers behind one contract.

OptionWhat It BuysWhat It Costs
Mandate one platformUniform guarantees; one dashboard; one on-call12–18 month migration; teams with Airflow DAGs lose features; platform becomes a bottleneck
Federation + common contractKeep Airflow for DAGs, Temporal for workflows, the timer service for timers; all emit the same run-history events and honor run_id/idempotency classContract enforcement across engines; a shared run-history pipeline
Laissez-faireZero coordination costCompliance gaps; incidents where nobody can enumerate what missed

The Principal position: Federation behind one contract. Standardize the interface and the evidence (registration fields, run-history schema, SLO definitions), not the engine. Retire only the engines with no owner (the Jenkins box, the bastion crontab).

Cost Model#

Assumptions: AWS-like on-demand pricing, RF=3 timer store, Kafka ready queue shared with other workloads (allocated share), fully loaded engineer cost ~$250K/year.

ScaleVolumeInfra $/monthHeadcountOn-call Load
Small1M jobs/day, 10M pending, Postgres + SKIP LOCKED~$1.5–3K0.5 FTE (part of a platform team)~1 page/month
Medium50M jobs/day, 1B pending, bucketed timers + 16 scheduler nodes~$25–45K3–4 engineers2–4 pages/month; shared rotation
Large5B jobs/day, 50B pending, multi-cell, multi-region~$250–400K8–12 engineersDedicated 24×7 rotation

At Medium scale, the people cost (~$75–85K/month) exceeds infrastructure. That's the number that decides build vs buy: if a managed service costs less than ~$100K/month total at your volume and meets governance needs, buying wins.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionDoorReversibility Cost
run_id format (deterministic hash)One-wayEvery job team's dedup keys depend on it; changing it creates a wave of duplicates
Guarantee (at-least-once)One-wayTeams build idempotency around it; switching to at-most-once silently loses work
Registration contract fieldsOne-way-ishAdding fields is cheap; removing or re-meaning them breaks 300 teams
Timer store engineTwo-wayDual-write + backfill migration; 1–2 quarters
Shard count (1,024)Two-way if chosen as a power of two and larger than node countSplits are mechanical; too few shards is painful
Ready queue productTwo-wayDispatch sink abstraction hides it
Default tolerance windowTwo-wayFlag + rollout; communicate as an API change

The Standard I'd Write#

RFC: Time-Triggered Work Standard (v1)

Scope: Any code that runs on a schedule, after a delay, or on a recurrence in production, regardless of engine.

MUST:

  1. Register with owner team, pager rotation, idempotency_class, tolerance_s, retry policy, and max_silence.
  2. Use run_id = hash(job_id, scheduled_at) as the idempotency key for all external side effects.
  3. Emit run-history events (dispatched, started, succeeded, failed, skipped) to the shared run-history topic.
  4. Jobs tagged compliance or financial MUST have a dead man's switch alert routed to the owner.

SHOULD:

  1. Use tolerance_s ≥ 60 unless a documented business reason requires exact timing.
  2. Use exponential backoff with ≥ 20% jitter for retries.
  3. Declare a per-dependency concurrency tag for any shared datastore.

Exceptions: Filed with the scheduling platform team; approved by its tech lead; expire after 12 months.

Success metrics: 100% of production schedules registered within 4 quarters; definitions_without_next_timer = 0; zero audit findings for missing runs; platform pages for job-logic failures < 1/month.

What I'd Tell the VP#

We currently run seven different schedulers, and none of them can tell us what failed to run last night. That's a compliance and reliability risk: last quarter a data-deletion job silently stopped for 34 days. I'm proposing one contract that all schedulers must meet — registration, idempotency, run history, and ownership — rather than forcing every team onto one engine. It costs roughly three engineers for a year plus ~$30K/month in infrastructure, and it retires two unowned systems. The outcome is a single answer to "did it run?" and a platform on-call that isn't paged for other teams' bugs.

Principal Interview Signals#

SignalWhat It Sounds Like
Sees the fleet, not the system"Before designing a scheduler, I'd inventory the ones we already have."
Standardizes the contract, not the engine"Airflow and Temporal can stay. What they share is run_id, run history, and ownership fields."
Prices the tradeoff"At 50M jobs/day, headcount dominates infra. That's the build-vs-buy number."
Identifies one-way doors"The run_id format is the decision I'd spend a week on; the storage engine I'd spend a day on."
Connects to business risk"Cron is where compliance controls live. The scheduler is audit evidence."

Staff answers that L7 interviewers find insufficient:

  • "We'll build a great central scheduler and migrate everyone" — no migration cost, no plan for DAG/workflow users who'd lose features.
  • "Teams own their job failures" — correct, but with no mechanism (scorecard, quotas, exceptions process) to make it stick across 300 teams.
  • "Multi-region active-active for all timers" — no price, no residency analysis, no tiering.

Appendices

Appendix A: Mechanics in Depth#

A.1 Postgres SKIP LOCKED Dequeue#

-- Claim up to 100 due jobs; concurrent pollers skip rows others hold
WITH due AS (
  SELECT job_id, scheduled_at
  FROM timers
  WHERE run_at <= now() AND state = 'pending'
  ORDER BY run_at
  LIMIT 100
  FOR UPDATE SKIP LOCKED
)
UPDATE timers t SET state = 'dispatched', dispatched_at = now()
FROM due WHERE t.job_id = due.job_id AND t.scheduled_at = due.scheduled_at
RETURNING t.job_id, t.scheduled_at;

Why it's right below ~5K/s: transactional, no separate coordination service. Why it breaks: every update creates a dead tuple; the run_at index's right edge is contended by all pollers; autovacuum falls behind at the herd.

A.2 Hierarchical Timing Wheel#

wheel_seconds[60], wheel_minutes[60]           # covers next 60 minutes in RAM
insert(timer):
  delta = timer.run_at - now
  if delta < 60s:  wheel_seconds[(now_s + delta) % 60].append(timer)
  else:            wheel_minutes[(now_m + delta/60) % 60].append(timer)
tick every 1s:
  fire all in wheel_seconds[now_s % 60]
  on minute boundary: cascade wheel_minutes[now_m % 60] into wheel_seconds

O(1) insert and fire. Volatile: on ownership change, rebuild from the durable buckets for the load horizon (~1–3s for 64 shards × 5 min).

A.3 Recurring Job: Atomic Next-Timer#

on_fire(occurrence):
  next = cron.next_after(occurrence.scheduled_at, tz=job.tz)
  batch_write([
    insert run_history(run_id, status='dispatched') if not exists,
    insert timer(bucket(next), shard(job_id), next, job_id)
  ])                                   # single partition-batch or outbox txn
  enqueue(ready_queue, run_id, token)  # idempotent at sink on run_id

Writing the next timer before enqueuing closes the broken-chain bug in Section 4.3.

Appendix B: Keys and Data Model#

RecordPartition KeyClusteringNotes
Timer(bucket_minute, shard)run_at, timer_idTTL run_at + 3d; one partition read per shard per minute
Cancel markertimer_id—TTL = timer TTL; checked at fire time to avoid tombstones
Definitionjob_id—Postgres; owner, policy, tz, tolerance
Run historyjob_idscheduled_at DESC, attempt30d hot, then object storage
Shard lease/sched/shards/{n} in etcd—Lease TTL 10s; revision = fencing token

run_id = sha256(job_id || scheduled_at_utc_iso)[:16] — scheduled time is the pre-jitter time so jitter changes don't change identity.

Appendix C: Coordination Mechanisms — Quick Comparison#

MechanismFailover GapDuplicate RiskComplexityUse When
Single leader + standbyLease TTL (~10s)High without fencingLow< 1K jobs/s
DB row locks (SKIP LOCKED)Immediate (row-level)Low (transactional)Low< 5K jobs/s
Sharded leases + fencingLease TTL per shardLowMedium5K–500K jobs/s
Consistent hashing of shards to nodes (gossip)Membership convergence (~seconds)Medium during churnMedium–HighVery large fleets; tolerant jobs
Durable execution engineEngine-definedActivities at-least-onceHigh (operate the engine)Workflows

Appendix D: Worker Contract and Retries#

  • Lease: lease_s = 1.5 × p99_duration, min 30s, max 15 min; heartbeat every lease_s / 3.
  • Lost heartbeat twice: worker must stop and not call Complete.
  • Retry schedule: delay = min(30s × 2^attempt, 1h) × uniform(0.8, 1.2); max 5 attempts, then team DLQ.
  • Non-retryable errors: worker returns FAILED_PERMANENT (e.g., validation) → straight to DLQ, no retries.
  • DLQ replay: owner-initiated, rate-limited, with the original run_id preserved.

Appendix E: Observability#

Core metrics:

scheduler.fire_lag_seconds{lane,shard}        # platform SLO
scheduler.due_per_second                      # herd visibility, watch max
scheduler.definitions_without_next_timer      # silent miss
scheduler.dispatch_rejected_stale_token_total # zombie owners
ready_queue.lag_seconds{lane,team}
runs.failure_rate{team,job}                   # team SLO
runs.attempts_gt_1_total{team}                # duplicate pressure
job.last_success_age_seconds{job}             # dead man's switch

Critical alerts:

AlertThresholdRoutes To
Fire-lag p99> 30s for 5 minPlatform
Unowned shardsany shard unowned > 30sPlatform
Broken chainsdefinitions_without_next_timer > 0 for 10 minPlatform
Team failure rate> 5% over 30 minOwning team
Dead man's switchlast_success_age > max_silenceOwning team

Control plane vs data plane: registration, policy, and quotas are control plane (can be down for minutes without missing fires); timer store reads, shard leases, and dispatch are data plane (must stay up).

Appendix F: Scale Evolution#

ScaleDesign
< 1K jobs/sPostgres SKIP LOCKED, one poller pool, K8s workers
1–5K jobs/sPartitioned Postgres tables by day; multiple pollers; tolerance windows
5–500K jobs/sBucketed timers (Cassandra/DynamoDB), 1,024 leased shards, timing wheels, Kafka ready queue
> 500K jobs/s or multi-regionCells by tenant tier; region-homed timers; tier-B replication

What you don't build on day one: multi-region failover, tenant cells, chargeback, custom workflow semantics, a UI beyond run history.

Appendix G: Multi-Tenancy and Fairness#

  • Lanes: P0 transactional, P1 standard, P2 bulk/backfill. P2 capped at 20% of dispatch when P0/P1 have backlog.
  • Within a lane: deficit round robin across teams, weights from reserved capacity.
  • Quotas: create rate, pending count, and dispatch rate per team; exceeding create quota returns 429 with Retry-After.
  • Chargeback: cost per 1M timers created + per 1M dispatched; cancellations counted, because tombstones cost money.
  • Noisy neighbor detection: scheduler.dispatch_share{team} > 50% of a lane for 10 min → notify owner and platform.
  1. Loading the index…