Go deeper:
- For the saga and compensation patterns these engines run, see Workflows, Sagas & Compensation.
- For time-triggered work at scale, see Distributed Job Scheduler.
Why This Matters#
A workflow engine is not a job queue with extra steps. It is a database for the program counter: it durably records every step a piece of business logic has taken, so the logic can crash, be redeployed or sleep for 30 days and resume exactly where it left off. Temporal calls this durable execution. The interesting failures are not "the worker died" (the engine handles that); they are the ones the engine enforces: workflow code that isn't deterministic and breaks on replay, a code change deployed under 50,000 running workflows, an activity retried forever because nobody set a timeout, an event history that hits 51,200 events and terminates the workflow.
That is why "we'll orchestrate it with Temporal" is a sentence interviewers push on. The L5 candidate draws a box labelled "workflow engine" between services. The L6 candidate says "the order workflow is deterministic code; every side effect is an activity with a Start-To-Close timeout of 30 seconds, a retry policy capped at 10 attempts with exponential backoff, and an idempotency key derived from the workflow ID; payment capture and inventory reservation have compensations; cancellation arrives as a signal; the 14-day return window is a durable timer; I version changes with patching and drain old workers before removing the old branch." The L7 candidate asks whether the company should run one engine or three, who owns the cluster, and what it costs when every team starts encoding business processes in a system only a few people know how to operate.
The L5 → L6 gap is not knowing what an activity is. It is knowing that the engine replays your workflow code from history, so determinism and versioning are not style rules; they are the correctness contract.
The L5 → L6 → L7 Contrast#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Use Temporal to orchestrate the services" | "Is this a long-running, multi-step process with waits, retries and compensations? If it's one step plus a retry, a queue is enough." | "Which processes deserve durable execution, which engine is the company standard, and who runs it?" |
| Code | Writes the workflow like normal code | Workflow code deterministic; all I/O, time and randomness through activities or SDK APIs; replay tests in CI | Org lint rules and replay-test gates; workflow code reviewed as a distinct discipline |
| Failure | "Temporal retries automatically" | Every activity has Start-To-Close, a bounded retry policy, non-retryable error types and idempotency at the callee | Retry budgets across the fleet so a downstream outage doesn't become a retry storm from 100K workflows |
| Change | Deploys new code | Patching or worker versioning for running workflows; keeps old workers until old executions drain | Version-lifecycle policy: max workflow age, Continue-As-New cadence, deprecation windows |
| Scale | "It scales" | History under ~10K events via Continue-As-New; payloads under 2 MB with references to blobs; task-queue backlog and schedule-to-start latency as signals | Capacity and cost of the engine as a platform: namespaces, actions per second, persistence store sizing |
| Ownership | "Platform team runs Temporal" | Platform owns the cluster; service teams own workflows, activities, timeouts and versioning | Defines the paved road and the exceptions: when Step Functions, a queue or a state table is the right answer |
Why "Code" separates levels
A Temporal worker doesn't keep your workflow's stack in memory forever. When a worker restarts, or a workflow's state is evicted from cache, the SDK re-executes the workflow function from the beginning and feeds it the recorded results from the event history. The docs are explicit: emitted commands are compared with the existing history, and if a command doesn't match, the execution fails with a non-determinism error (workflow definition). So if time.Now().Hour() < 12, rand.Int(), iterating a hash map in random order, or calling an HTTP API directly inside workflow code all produce different commands on replay. The Senior answer writes natural code and is surprised in production; the Staff answer routes all non-determinism through activities or SDK-provided time and random APIs and runs replay tests against real histories in CI; the Principal answer makes those tests a merge gate for every team.
Why "Failure" separates levels
Temporal's default activity retry policy is an initial interval of 1 second, backoff coefficient 2.0, maximum interval 100 seconds and unlimited attempts (retry policies). That is a sensible default for transient faults and a dangerous one for a permanent failure: a validation error retried forever, or 100,000 workflows each retrying a down payment provider every 100 seconds. The Staff answer sets a Start-To-Close timeout on every activity (the docs strongly recommend it, because it's how the server detects a crashed worker), caps attempts or Schedule-To-Close, marks business errors non-retryable, and makes every activity idempotent at the callee because "retried" means "possibly executed twice." The Principal answer adds fleet-level protection: rate limits per task queue and circuit breakers so the engine's reliability doesn't become the downstream's outage.
The 60-Second Pitch#
"Order fulfilment runs as a Temporal workflow per order, with the order ID as the workflow ID so duplicate starts are rejected. The workflow is deterministic code: reserve inventory, authorise payment, create shipment, then wait on a durable timer for the 14-day return window. Every call to another service is an activity with a 30-second Start-To-Close timeout, exponential backoff capped at 10 attempts, and an idempotency key of workflow ID plus step name, so a retried capture never charges twice. Failures after payment trigger compensations in reverse order. Customer cancellation arrives as a signal; the support UI reads status through a query. Large payloads go to object storage and only references go into history. Code changes ship with patching, and we keep old worker builds running until workflows on the old path complete. We run the engine as a platform service or use the managed cloud offering; service teams own their workflows. For a single fire-and-forget step, I'd use a queue instead."
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Business process orchestration (orders, payments, onboarding, KYC) | Multi-step, minutes to months, compensations, human waits | Workflow per entity, activities per side effect, signals for external events, durable timers | Non-idempotent activities double-charge; un-versioned changes break running workflows | Every process reaches a terminal state; no duplicate side effects |
| Infrastructure automation (deploys, provisioning, database operations) | Long steps, external systems, operator approvals | Workflows wrapping cloud APIs with heartbeating activities; signals for approvals | Activity hangs without heartbeat; retries hammer a cloud API's rate limit | Safe to resume after any crash; operations auditable |
| High-volume pipelines and fan-out (per-message processing, batch jobs) | Millions of short executions, throughput over duration | Often a queue or stream instead; if Temporal, batch work per workflow and use child workflows | Actions per second and history writes become the bottleneck and the bill | Throughput per dollar; at-least-once with idempotent consumers |
🎯 Staff Move: "I'll use the engine for the order lifecycle, because it spans services, days and compensations. The per-event analytics stream stays on Kafka: millions of independent one-step messages don't need a durable program counter, and paying per workflow action for them would be the wrong trade."
The Staff Positions#
| Position | Rationale |
|---|---|
| Workflow code is deterministic; all I/O lives in activities | Replay re-executes workflow code; any divergence fails the execution. |
| Every activity is idempotent and has a Start-To-Close timeout | Retries mean at-least-once; the timeout is how crashed workers are detected. |
| Retries are bounded and business errors are non-retryable | Default retries are unlimited; a permanent failure should fail fast and compensate. |
| The workflow ID is the business key | Duplicate starts for the same order are rejected; the ID is the natural idempotency anchor. |
| History stays small | Continue-As-New well before the 51,200-event limit; references, not payloads, in history. |
| Every change to running workflow code is versioned | Patching or worker versioning; replay tests on production histories gate the deploy. |
| Use the engine for long, multi-step processes, not everything | A queue is cheaper and simpler for one-step, high-volume work. |
Architecture & Internals#
Five internals change design decisions: workflows vs activities, event history and replay, task queues and workers, timers, signals, queries and updates, and the server and its persistence.
Workflows vs Activities#
| Workflow | Activity | |
|---|---|---|
| What it is | Orchestration logic: the sequence, branching, waits, compensations | A single unit of work with side effects: call an API, write a DB, send an email |
| Determinism | Must be deterministic | May do anything |
| Durability | State reconstructed from event history | Result recorded in history once it completes |
| Duration | Seconds to years (no timeout by default) | Bounded by Start-To-Close; long ones heartbeat |
| Failure handling | Doesn't retry by default; handles activity failures in code | Retried by policy (default: unlimited attempts) |
| Where it runs | Your worker process, replayed as needed | Your worker process, possibly on a different host each attempt |
Event History and Replay#
Every state change is an event appended to the workflow's history: started, activity scheduled, activity completed, timer fired, signal received. On replay the SDK re-runs workflow code and, instead of executing activities again, returns their recorded results. That is how a workflow survives a crash in the middle of a 30-day wait without anyone writing checkpoint code.
| Limit (Temporal Cloud and server defaults) | Value | Design consequence |
|---|---|---|
| Event history warning | 10,240 events | Plan Continue-As-New well before this |
| Event history hard limit | 51,200 events or 50 MB | Execution is terminated beyond it |
| Single payload / request | 2 MB | Store blobs in object storage, pass references |
| Incomplete activities, child workflows, signals per execution | 2,000 | Fan out with child workflows in batches |
| Signals per execution | 10,000 | High-frequency signals need batching or a new run |
| Updates in history | 2,000 total, 10 concurrent | Use update IDs and validators |
Sources: event history limits, Temporal Cloud limits.
Continue-As-New closes the current run and starts a fresh one with the same workflow ID, carrying state forward as input. A subscription workflow that bills monthly forever, or an entity workflow processing a signal stream, continues-as-new every N iterations to keep history small and replay fast.
Task Queues and Workers#
Workers are your processes. They long-poll task queues on the server for workflow tasks and activity tasks, run your code, and report results. The server never calls your code directly, so workers can sit in private networks and scale independently.
Why it matters in design: a backlog on a task queue means workers can't keep up, which you see as rising schedule-to-start latency before you see failures. Separate task queues per activity type let you scale and rate-limit a slow dependency (a payment API capped at 200 requests per second) without starving everything else.
Timeouts#
| Timeout | Measures | Default | Staff guidance |
|---|---|---|---|
| Activity Start-To-Close | One attempt's duration | Same as Schedule-To-Close | Always set; p99 of the call × 2–3 |
| Activity Schedule-To-Close | All attempts, including retries | ∞ | Set to the business deadline for the step |
| Activity Schedule-To-Start | Time waiting in the queue | ∞ | Rarely set; alert on the latency metric instead |
| Activity Heartbeat | Max gap between heartbeats | Disabled | Set for anything over ~1 minute; enables fast crash detection and progress checkpoints |
| Workflow Execution / Run | Whole workflow / one run | ∞ | Usually leave unset; model deadlines as timers in code |
| Workflow Task | Worker processing one workflow task | 10 s | Keep workflow code fast; no blocking I/O |
Sources: activity timeouts, workflow timeouts.
Timers, Signals, Queries and Updates#
| Primitive | Direction | Sync? | Mutates state? | Recorded in history? | Use for |
|---|---|---|---|---|---|
Durable timer (sleep, await with timeout) | Internal | — | Yes | Yes | Return windows, retries with long waits, reminders |
| Signal | Into workflow | Async, fire-and-forget | Yes | Yes | Cancellation, approvals, external events |
| Query | Read from workflow | Sync | No | No | Status pages, debugging; works on completed workflows |
| Update | Into workflow, with a result | Sync | Yes | Yes | "Add item to cart and return the new total"; validators reject bad requests before they're recorded |
Source: message passing.
🎯 Staff Insight: "A durable timer is one of the most underrated primitives. 'Wait 14 days, then close the return window unless a return signal arrived' is five lines of workflow code instead of a cron job, a state table, a sweeper and the bugs between them."
Data Modeling / Core Usage — "The Entire Game"#
In DynamoDB the game is the access-pattern table. In a workflow engine it is the workflow contract: what one workflow execution represents, where the side-effect boundaries are, and how the code will change while executions are still running.
Step 1: One Workflow per Business Entity Lifecycle#
Good workflow IDs (the business key, so duplicate starts are rejected):
order-{order_id} order lifecycle: reserve -> pay -> ship -> return window
subscription-{account_id} long-lived entity; Continue-As-New every billing cycle
deploy-{service}-{build} infrastructure automation with approvals
Bad:
random UUID per start -> a retried "start" creates a second order workflow
one workflow for all orders -> unbounded history, a single hot execution
one workflow per HTTP call -> durable execution for work that never waits
Step 2: Draw the Activity Boundaries#
Every side effect is an activity, and every activity is idempotent at the callee using a key the workflow can reproduce on replay.
// Workflow: deterministic orchestration only
func OrderWorkflow(ctx workflow.Context, order Order) error {
ao := workflow.ActivityOptions{
StartToCloseTimeout: 30 * time.Second,
RetryPolicy: &temporal.RetryPolicy{
InitialInterval: time.Second,
BackoffCoefficient: 2.0,
MaximumInterval: time.Minute,
MaximumAttempts: 10,
NonRetryableErrorTypes: []string{"CardDeclined", "InvalidAddress"},
},
}
ctx = workflow.WithActivityOptions(ctx, ao)
key := workflow.GetInfo(ctx).WorkflowExecution.ID // stable across replays
var resv Reservation
if err := workflow.ExecuteActivity(ctx, ReserveInventory, key+"-reserve", order).Get(ctx, &resv); err != nil {
return err // nothing to compensate yet
}
var auth Auth
if err := workflow.ExecuteActivity(ctx, AuthorizePayment, key+"-auth", order).Get(ctx, &auth); err != nil {
_ = workflow.ExecuteActivity(ctx, ReleaseInventory, key+"-release", resv).Get(ctx, nil)
return err // compensate in reverse order
}
// ... capture, ship; then a durable 14-day timer races a "return" signal
returnCh := workflow.GetSignalChannel(ctx, "return-requested")
sel := workflow.NewSelector(ctx)
sel.AddReceive(returnCh, func(c workflow.ReceiveChannel, more bool) { /* start refund */ })
sel.AddFuture(workflow.NewTimer(ctx, 14*24*time.Hour), func(f workflow.Future) { /* close window */ })
sel.Select(ctx)
return nil
}
| Inside workflow code | Inside an activity |
|---|---|
| Branching on activity results and signals | HTTP calls, database writes, queue publishes |
workflow.Now(), workflow.Sleep(), SDK timers | time.Now(), real sleeps, polling loops |
| SDK side-effect APIs for one-off random values | Random IDs, UUID generation for external systems |
| Deterministic iteration over sorted collections | Reading files, environment variables, config services |
| Small state (IDs, statuses) | Large payloads: write to object storage, pass the key |
Step 3: Plan for Change Before the First Deploy#
Running workflows replay against whatever code the worker has. Changing the sequence of commands for executions already in flight causes non-determinism errors. Two approaches:
| Approach | How it works | Use when | Cost |
|---|---|---|---|
Patching (GetVersion in Go and Java, patched() in TypeScript and Python) | Code branches on a marker recorded in history: old executions take the old path, new ones the new path | Small changes to long-running workflows | Branches accumulate; remove old ones only after old executions finish |
| Worker versioning | Workers are tagged with a build; executions stay pinned to (or are deliberately moved between) builds | Frequent deploys, many short-to-medium workflows | Old worker builds must keep running until their executions drain |
| New workflow type | OrderWorkflowV2; new starts use it, old ones finish on V1 | Large rewrites | Two codepaths live until V1 drains |
The Temporal docs recommend worker versioning and describe patching as the alternative (workflow definition). Either way the discipline is the same: replay tests that run new code against histories exported from production, as a CI gate.
🎯 Staff Move: "Activity code changes freely, because activities aren't replayed. Workflow code changes are versioned, replay-tested against a sample of production histories in CI, and the old branch is deleted only after a visibility query shows no open executions on it."
Step 4: Keep History Small#
History budget: stay under ~10K events per run (warning at 10,240; hard stop at 51,200)
Each activity costs ~3 events (scheduled, started, completed); each timer ~2; each signal 1.
Order workflow: ~12 activities + 2 timers + few signals -> ~45 events: fine
Entity workflow receiving 50 signals/day -> ~200 days to warning
-> Continue-As-New every 1,000 signals, carrying current state as input
Batch workflow over 100K items as activities -> 300K events: impossible
-> child workflows of 1,000 items each, or activities that process pages
The Tunable Tradeoff — Durability × Latency × Cost#
Every step you make durable is a write to the engine's persistence store. Durability is the product; it is also the bill and the latency.
| Setting | Light end | Heavy end | Who pays at the light end |
|---|---|---|---|
| Granularity | Few large activities | One activity per tiny step | On-call, when a large activity fails halfway and isn't idempotent |
| Engine vs code | Queue + state table, hand-rolled | Full workflow engine | Engineers maintaining sweepers and state transitions |
| History size | Continue-As-New often | One run for the whole lifecycle | Replay latency and history limits |
| Activity timeouts | Long, generous | Tight, with heartbeats | Detection time when a worker dies mid-activity |
| Retries | Unlimited (default) | Bounded with non-retryable errors | Downstream services during an outage |
| Hosting | Self-hosted cluster | Managed cloud service | Platform team's pager; or finance, per action |
Latency floor (illustrative, typical self-hosted or managed deployments):
each activity round trip = schedule + dispatch + execute + record
engine overhead per step ~ 10-50 ms on top of the work itself
10-step workflow in the request path -> 100-500 ms of orchestration overhead
-> keep synchronous user-facing paths short; start the workflow, return,
and let the UI poll a query or subscribe to updates
Cost shape (managed services bill per action or state transition):
Step Functions Standard: $0.025 per 1,000 state transitions (US East list price)
1 M orders/day x 15 transitions = 15 M/day -> ~$375/day -> ~$11K/month
same volume as per-message processing at 50 M messages/day x 3 steps -> ~$112K/month
-> durable execution is cheap per business process, expensive per event
Step Functions pricing is from the AWS pricing page; Temporal Cloud bills per action with a different rate card, so model your own volumes in the Cost Estimator.
🎯 Staff Move: "I put one activity around each external side effect, not around each line of code. That's where idempotency keys and compensations naturally live, and it keeps both history and the per-action bill proportional to business steps, not to code size."
Who Pays for Each Choice#
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Unlimited default retries | Transient faults heal themselves | Permanent errors retried forever; retry storms on outages | Downstream teams and their on-call |
| Non-idempotent activities | Simpler activity code | Double charges, duplicate emails on retry | Customers and finance |
| No versioning discipline | Faster first deploys | Non-determinism errors stall in-flight workflows | Service on-call during deploys |
| Huge payloads in history | Convenient | 2 MB limit, slow replay, persistence growth | Platform team |
| Synchronous workflows in the request path | Simple API | Engine latency added to user requests | Users |
Anti-Patterns — What Kills Workflow Engine Deployments#
1. Non-Deterministic Workflow Code#
time.Now(), random numbers, map iteration order, direct HTTP calls or reading config inside workflow code. Works in testing; fails on the first replay after a worker restart. Fix: SDK APIs for time and randomness, all I/O in activities, replay tests in CI.
2. Unversioned Changes Under Running Workflows#
Reordering two activities and deploying while 40,000 orders are mid-flight. Every one of them hits a non-determinism error on its next workflow task. Fix: patching or worker versioning; replay tests on production histories; keep old builds until executions drain.
3. Activities Without Timeouts or Idempotency#
No Start-To-Close means a crashed worker is detected late or never; no idempotency key means a retry after a timeout charges the card twice. Fix: Start-To-Close on every activity, idempotency keys from workflow ID plus step, heartbeats for long work. See Idempotency.
4. Unbounded Retries Against a Down Dependency#
The default policy retries forever with a 100-second ceiling; 200,000 workflows retrying a payment provider during its outage is a self-inflicted DDoS when it comes back. Fix: bounded attempts or Schedule-To-Close, rate-limited task queues per dependency, circuit breaking in the activity.
5. The Workflow as a Database#
Storing large documents, growing lists or every message of a chat in workflow state. History balloons toward 50 MB and replay slows to seconds. Fix: state belongs in a database; the workflow holds IDs and statuses; Continue-As-New for long-lived entities.
6. Durable Execution for Everything#
Wrapping every API call or every Kafka message in a workflow. Throughput is capped by actions per second, and the bill scales with event volume. Fix: queues and streams for one-step high-volume work; workflows for multi-step processes. See Kafka vs SQS vs RabbitMQ.
7. Blocking Work in Workflow Code#
CPU-heavy computation or blocking calls in the workflow function exceed the 10-second workflow task timeout and cause repeated task failures. Fix: move heavy work into activities; keep workflow code a fast decision function.
8. No Owner for Stuck and Failed Workflows#
Executions that fail or sit waiting on a signal that will never come accumulate silently. Fix: visibility queries for failed and long-open executions per workflow type, a named owner, and runbooks for terminate, reset or signal.
The Technology Landscape — Head-to-Head Comparison#
| Dimension | Temporal (self-hosted or cloud) | AWS Step Functions | Hand-rolled state machine (DB + queue + sweeper) | Netflix Conductor-style JSON orchestrators | DAG schedulers (Airflow and similar) |
|---|---|---|---|---|---|
| Model | Workflows as code, replayed from history | JSON/ASL state machine, managed | Status column, transitions in code, queue for work | Workflow definitions in JSON, workers poll tasks | Batch DAGs on a schedule |
| Long waits | Durable timers, signals; no execution limit by default | Standard: up to 1 year; Express: 5 minutes | Rows plus a sweeper | Supported | Not the model |
| History limit | 51,200 events / 50 MB per run | Standard: 25,000 events | Whatever you build | Engine-specific | N/A |
| Payload limit | 2 MB per payload | 256 KiB per state input/output | Your DB | Engine-specific | XCom-style small values |
| Versioning running work | Patching, worker versioning | Versions and aliases; executions keep the version they started on | Your migration code | Definition versions | DAG versions per run |
| Ops burden | High self-hosted (persistence store, upgrades); low on cloud | None | Medium, spread across every team | Medium–high | Medium |
| Pick when | Complex, long-running logic best expressed in code; multi-cloud | AWS-native, mostly service integrations, modest logic | One or two simple flows, strong DB skills | Existing investment, JSON-defined flows | Scheduled data pipelines |
Step Functions limits are from the service quotas page.
🎯 Staff Insight: "Every team that hand-rolls state machines ends up building a worse workflow engine: a status column, a retry counter, a sweeper cron and a dead-letter table. I'll hand-roll for one simple flow; by the third, I want a real engine and one team that's good at running it."
Patterns#
Pattern 1: Saga with Compensations#
Each forward step registers its compensation; on failure the workflow runs compensations in reverse. The engine guarantees the compensations themselves are retried until they succeed or escalate. Full treatment in Workflows, Sagas & Compensation; payment specifics in Payments.
Pattern 2: Entity Workflow#
One long-lived workflow per entity (an account, a cart, a device) receives signals and updates, holds a small state, and continues-as-new every N events. It serialises operations per entity without locks, the same property a per-key partition gives in Kafka.
Pattern 3: Human-in-the-Loop Approval#
The workflow sends a notification activity, then waits on an approval signal raced against a timer (escalate after 48 hours, auto-reject after 7 days). This replaces an approvals table, a reminder cron and an expiry sweeper.
Pattern 4: Fan-Out with Child Workflows#
Each level keeps history bounded and pending children under the 2,000 limit; activities process pages, not single items.
Pattern 5: Outbox to Workflow Start#
The service writes its business row and an outbox row in one transaction; a relay starts the workflow with the business key as workflow ID, so relay retries are deduplicated by the engine. See Transactional Outbox.
Scaling#
The Numbers#
| Resource | Documented limit or default | Design note |
|---|---|---|
| Event history per run | Warning 10,240 events; terminated at 51,200 events or 50 MB | Continue-As-New early |
| Payload per request | 2 MB | References to blobs, not blobs |
| Pending activities / children / signals per execution | 2,000 | Batch fan-out through child workflows |
| Workflow task timeout | 10 s default | Workflow code must be fast and non-blocking |
| Activity retry default | 1 s initial, ×2, max 100 s, unlimited attempts | Always set your own bounds |
| Temporal Cloud namespace rate | 500 actions per second default, scaling with capacity mode | Size namespaces per domain; request increases ahead of launches |
| Step Functions Standard | 1-year max duration, 25,000 history events, 256 KiB payloads | Express for < 5 min, high-volume flows |
| Step Functions Standard transitions | 5,000/s bucket in the largest regions (soft quota) | Throttling shows as ExecutionThrottled |
Scaling Moves in Order#
- Shrink per-workflow cost: fewer, coarser activities around real side effects; references instead of payloads; Continue-As-New for long-lived entities.
- Scale workers horizontally per task queue; tune poller counts and concurrent-execution slots; watch schedule-to-start latency.
- Split task queues by dependency so a slow or rate-limited downstream gets its own workers and limits.
- Split namespaces by domain (payments, fulfilment, infra) for rate limits, retention and blast radius.
- Scale the engine: more history shards and persistence capacity when self-hosting, or a higher capacity tier on a managed service. History shard count is set at cluster creation, so pick it for year-three volume.
- Move high-volume one-step work off the engine onto queues or streams.
Failure Modes & Recovery#
1. Non-Determinism Errors After a Deploy#
- Symptom: Minutes after a deploy, workflow tasks fail repeatedly for a workflow type; executions stop progressing.
- Root cause: Workflow code changed the command sequence without a version marker, or introduced non-deterministic calls.
- Detection: Workflow task failure rate by type and build; non-determinism error counts in worker logs; executions stuck with pending workflow tasks.
- Fix: Roll back the worker build (executions resume on the old code); re-ship with patching.
- Prevention: Replay tests on production histories in CI; worker versioning; review checklist for workflow code.
2. Retry Storm on a Recovering Dependency#
- Symptom: A payment provider recovers from an outage and immediately falls over again; activity failure rates stay at 100%.
- Root cause: Tens of thousands of executions retrying in sync under an unbounded policy.
- Detection: Activity attempt count distribution; downstream request rate vs normal; activity failure rate by type.
- Fix: Throttle the task queue's activity rate; pause or scale down activity workers for that dependency; let backlog drain at a controlled rate.
- Prevention: Bounded retries, task-queue rate limits per dependency, circuit breakers in activities.
3. History Limit Termination#
- Symptom: Long-lived workflows terminate unexpectedly; the warning at 10,240 events was in logs nobody read.
- Root cause: Entity workflow accumulating signals, or a loop of activities with no Continue-As-New.
- Detection: History length percentiles by workflow type; server warnings for large histories.
- Fix: Reset or restart affected entities from their last known state; add Continue-As-New.
- Prevention: History-length alert at 5,000 events; design reviews that budget events per lifecycle.
4. Worker Backlog — Schedule-to-Start Latency Climbs#
- Symptom: Orders take minutes instead of seconds to progress; no errors.
- Root cause: Too few workers or slots, a deploy that reduced pollers, or one slow activity type hogging shared workers.
- Detection:
schedule_to_startlatency p95 per task queue; task-queue backlog; worker slot utilisation. - Fix: Scale workers; move the slow activity to its own task queue.
- Prevention: Autoscale workers on backlog and schedule-to-start latency; separate queues per dependency.
5. Engine or Persistence Outage#
- Symptom: Starts, signals and queries fail; all workflows pause.
- Root cause: Persistence store (Cassandra, MySQL, Postgres) overloaded or down, or the managed service region impaired.
- Detection: Frontend error rates and latency; persistence latency; synthetic start-and-complete canary workflow every minute.
- Fix: Restore the store; workflows resume from history once the engine is back, because state is durable and timers fire late rather than not at all.
- Prevention: Persistence sized with headroom; managed service for teams without database operators; callers degrade gracefully (accept the order, start the workflow later from an outbox).
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Non-determinism after deploy | Workflow task failures by build | All running executions of the type | Roll back worker build, re-ship patched | Service team |
| Retry storm | Attempt counts, downstream RPS | Downstream service and every workflow using it | Rate-limit task queue, pause workers | Service team + downstream owner |
| History limit | History length percentiles | Long-lived executions of the type | Continue-As-New, reset | Service team |
| Worker backlog | Schedule-to-start latency | Latency for that task queue | Scale workers, split queues | Service team; platform for autoscaling |
| Engine outage | Canary workflow, frontend errors | Every workflow in the cluster or namespace | Restore persistence; outbox buffering at callers | Platform team |
When to Use vs. Alternatives#
| Need | Pick | Why |
|---|---|---|
| Multi-step business process spanning services, days and compensations | Temporal (or a similar code-first engine) | Durable state, timers, retries and versioning built in |
| AWS-native glue between managed services with simple logic | Step Functions | No infrastructure; direct service integrations |
| One background task with retry | A queue with a DLQ | Cheaper and simpler; no orchestration needed |
| High-volume event processing | Kafka plus stream processing | Throughput per dollar; ordering per key; see Kafka |
| Scheduled data pipelines | A DAG scheduler | Built for batch dependencies and backfills |
| One simple flow, strong DB team, no engine yet | State table + outbox + sweeper | Fine for one or two flows; plan to replace by the third |
When NOT to Use a Workflow Engine#
- Single-step jobs. "Send this email with retries" is a queue with a DLQ, not a durable program.
- Sub-50 ms request paths. Engine round trips add tens of milliseconds per step; check against the Latency Budget.
- Per-event processing at millions per second. Per-action pricing and history writes make streams the right tool.
- Teams that won't learn the determinism rules. The engine enforces them in production whether the team learned them or not.
- As the system of record. Workflow state is retained for a configured period (Temporal Cloud: 1–90 days after close); durable business data belongs in a database.
Operational Concerns#
What the On-Call Actually Does#
- Watches failed and stuck executions per workflow type through visibility queries; decides reset, signal, terminate or fix-forward.
- Watches schedule-to-start latency and backlog per task queue; scales workers or splits queues.
- Gates deploys on replay tests and rolls back worker builds on non-determinism errors.
- Rate-limits activities against struggling dependencies instead of letting retries pile up.
- Runs the canary workflow (start, one activity, timer, complete every minute) as the engine's end-to-end health check.
Key Metrics & Alerts#
| Metric | Healthy | Alert |
|---|---|---|
| Canary workflow end-to-end latency | < 2 s | Failure or > 10 s for 3 runs |
| Schedule-to-start latency p95 per task queue | < 1 s | > 30 s for 5 min |
| Workflow task failures (non-determinism) | 0 | Any after a deploy (page) |
| Activity failure rate by type | Baseline | 3× baseline for 10 min |
| Executions open longer than expected per type | Near 0 | Growth day over day |
| History length p99 per type | < 5,000 events | > 10,000 |
Interview Application — Staff-Level Plays#
Which Case Studies Use Workflow Engines#
| Case Study | How the Engine Is Used | Key Pattern |
|---|---|---|
| Payments | Authorise, capture, refund as a saga | Idempotency keys per step; compensations |
| Hotel Booking | Hold → pay → confirm with expiring holds | Durable timer releases the hold |
| Ride-Hailing | Trip lifecycle from request to payout | Entity workflow per trip; signals for state changes |
| Deployment System | Waves, bake times, approvals, rollback | Human-in-the-loop signals; long activities with heartbeats |
| Distributed Job Scheduler | Execution of scheduled multi-step jobs | Scheduler of record plus workflow execution |
| Webhook Delivery | Retry schedules over hours or days | Bounded retries, per-endpoint rate limits |
| Ledger & Wallet | Transfers across internal and external rails | Saga with reconciliation, ledger as source of truth |
Every System Design Question Has a Workflow Moment#
- E-commerce checkout: "The order is a workflow keyed by order ID: reserve, authorise, capture, ship, then a 14-day return timer. Each step is an idempotent activity with a bounded retry policy."
- User onboarding: "KYC is a workflow that waits up to 7 days for document upload signals, with reminder activities at 24 and 72 hours."
- Data deletion (GDPR): "A deletion workflow fans out to 15 services as activities, retries each with a deadline, and records completion per service for the audit."
What Interviewers Probe#
| After You Say... | They Will Ask... | What They're Evaluating |
|---|---|---|
| "Temporal handles retries" | "What if the activity charged the card and then timed out?" | Idempotency at the callee |
| "Write the workflow in normal code" | "What happens when the worker restarts mid-workflow?" | Replay and determinism |
| "We'll deploy a fix" | "What about the 30,000 workflows already running?" | Versioning: patching, worker builds |
| "The workflow waits for the user" | "What if they never respond?" | Timers raced against signals; terminal states |
| "Use a workflow engine for every message" | "What does that cost at 50 M messages a day?" | Choosing queues vs workflows |
Common Interview Mistakes#
| What Candidates Say | What Interviewers Hear | What Staff Engineers Say |
|---|---|---|
| "Temporal gives us exactly-once" | Doesn't know activities are at-least-once | "Workflow logic runs effectively once; activities are at-least-once, so the callee dedupes on a key." |
| "Retries are automatic" | Unbounded retries against a down dependency | "Bounded attempts, non-retryable business errors, rate-limited task queues." |
| "We'll just redeploy the workflow" | Will break running executions | "Patched or worker-versioned, replay-tested against production histories." |
| "Keep the cart in the workflow state" | Workflow as a database | "The workflow holds IDs; the cart lives in the database." |
L5 vs L6 vs L7 Responses#
| Scenario | L5 Answer | L6 / Staff Answer | L7 / Principal Answer |
|---|---|---|---|
| "Orchestrate a 6-service order flow" | Temporal workflow calling each service | Workflow ID = order ID; idempotent activities with timeouts and bounded retries; compensations; signals and timers; versioning plan | Decides engine standard, hosting model and the platform contract for every team |
| "Deploy broke running workflows" | Roll back | Roll back the worker build, re-ship with a patch, add replay tests on production histories | Makes replay tests a merge gate org-wide and tracks non-determinism incidents to zero |
| "Downstream outage, workflows piling up" | Wait for recovery | Rate-limit the task queue, cap retries, drain backlog at a controlled rate | Fleet-wide retry budgets and per-dependency limits as platform features |
| "Step Functions vs Temporal" | Temporal is more powerful | Step Functions for AWS-native integration with simple logic; Temporal for complex, long-running code | Prices ops headcount vs per-transition cost and lock-in; sets the default and exceptions |
The Staff Workflow Engine Checklist#
- Fit: "Multi-step, long-running, with waits and compensations? Otherwise a queue."
- Identity: "Workflow ID is the business key, so duplicate starts are rejected."
- Activities: "One per external side effect; idempotency key from workflow ID plus step; Start-To-Close on all."
- Retries: "Exponential, capped at 10 attempts or a Schedule-To-Close deadline; business errors non-retryable."
- History: "Under 10K events per run; Continue-As-New for entities; blobs by reference."
- Change: "Patching or worker versioning, replay tests in CI, old builds kept until executions drain."
🎯 Staff Insight: Don't use a workflow engine as a database, as a per-event stream processor, or in a sub-50 ms request path. The strongest signal is explaining replay: why workflow code must be deterministic, and what you do with the executions already running when the code changes.
Evaluation Rubric#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Model | "Orchestrates services" | Replay from history; workflow vs activity boundaries | Teaches the model as an org discipline with tooling |
| Failure | "Retries automatically" | Timeouts, bounded retries, idempotency, compensations | Fleet retry budgets, dependency protection |
| Change | "Redeploy" | Patching, worker versioning, replay tests | Version lifecycle policy and deprecation windows |
| Scale | "Scales" | History budgets, payload references, task-queue design | Namespaces, persistence capacity, cost per business process |
| Fit | Engine for everything | Engine vs queue vs Step Functions by workload | Paved road, hosting model, exceptions process |
Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
At Staff level a workflow engine is a tool for one process. At Principal level it becomes where the company's business processes live, which makes it critical infrastructure with an unusual skill requirement: every engineer writing workflows must understand replay and versioning. The L7 question is "which processes belong on the engine, who runs it, and how do we keep 30 teams from each learning the determinism rules through an outage?"
🧭 Principal Move: "Adopting a workflow engine is adopting a programming model. I'd fund the platform team, the SDK wrappers that enforce our retry and timeout defaults, and the replay-test gate before I'd let a second team onto it."
The Org-Level Fault Line#
One engine standard vs best-tool-per-team.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Every team picks (Step Functions, Temporal, hand-rolled, Airflow) | Local fit | N engines, N skill sets, no shared tooling or on-call | Incident responders; future migrations |
| One engine, mandated for all async work | Shared expertise and tooling | High-volume one-step work forced onto an expensive model | Teams with stream-shaped workloads |
| Paved road: one code-first engine for multi-step processes, queues/streams for events, managed state machines allowed for AWS glue | Clear defaults by workload shape | Requires a platform team and an exceptions process | Platform headcount |
The Principal default: the paved road. One engine standard for long-running multi-step processes, with an internal SDK wrapper that sets safe defaults (Start-To-Close required, bounded retries, idempotency helpers, payload size checks); queues and streams for event processing; Step Functions permitted for AWS-native glue owned by small teams.
Cost Model#
Assumptions: managed engine billed per action, assumed at ~$25–50 per million actions for illustration (check the provider's current rate card); self-hosted cluster on a SQL or Cassandra store; loaded engineer ~$250K/year. Directional only.
| Scale | Workload | Managed engine/month | Self-hosted infra/month | Platform headcount | Total/month (managed) |
|---|---|---|---|---|---|
| Startup | 50K workflows/day × 30 actions | ~$1–2K | ~$1–2K plus ops | 0.25 FTE | ~$6–8K |
| Growth | 2 M workflows/day × 40 actions | ~$60–120K | ~$15–30K | 2 FTE (~$42K) | ~$100–160K |
| Enterprise | 30 M workflows/day × 40 actions | Negotiated; at the assumed rates ~$0.9–1.8M | ~$150–300K | 6–8 FTE (~$150K) | ~$400K–1M+ |
The crossover is the decision: below growth scale, managed hosting is cheaper than the engineers needed to run the persistence store; at enterprise scale, self-hosting or negotiated pricing wins, and action count per workflow is the lever that matters most on either path.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse |
|---|---|---|
| Business logic written against an engine's SDK | One-way-ish | Rewrite workflows; drain or migrate running executions |
| History shard count (self-hosted) | One-way for the cluster | New cluster and migration |
| Long-running workflows (months) on a given design | One-way for their lifetime | Patches live until the last one completes |
| Managed vs self-hosted | Two-way-ish | Namespace migration; weeks to months |
The Standard I'd Write#
RFC-WF-003: Durable Workflow Baseline
Scope: Every workflow running on the company workflow engine.
MUST
1. Use the business entity key as the workflow ID.
2. Keep workflow code deterministic: time, randomness and I/O only through SDK APIs
or activities. Replay tests against sampled production histories gate every merge.
3. Set Start-To-Close on every activity; bound retries by attempts or
Schedule-To-Close; declare non-retryable business error types.
4. Make every activity idempotent with a key derived from workflow ID and step.
5. Version every change to workflow command order (patching or worker versioning);
keep old worker builds until no open executions use them.
6. Keep payloads under 256 KB (blobs by reference) and runs under 10,000 events.
SHOULD
7. Use a dedicated task queue per rate-limited dependency.
8. Use queues or streams, not workflows, for single-step work above 1,000/s.
Exceptions: platform lead approval, reviewed quarterly.
Success metrics: zero non-determinism incidents per quarter; no activity without a
timeout; schedule-to-start p95 < 1 s; cost per business process reported per team.
What I'd Tell the VP#
"We're moving our core business processes, like orders, payments and onboarding, onto a single workflow platform so they finish reliably even when systems fail mid-way, without each team writing its own retry and recovery logic. The risk is that this platform is a new way of programming, and mistakes show up as stuck orders after a deploy. So I'm asking for a small platform team that provides safe defaults and automated checks every team must pass, and a clear rule for what doesn't belong on the platform, because putting every event through it would multiply our costs. Expect fewer stuck-order incidents and faster delivery of new processes, for two engineers and a managed-service bill that grows with order volume, not with traffic."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Programming model, not a tool | "Adopting the engine means teaching replay and versioning to every team; I fund that first." |
| Defaults in the SDK | "Our wrapper won't compile an activity without a timeout or a bounded retry." |
| Prices actions | "Cost scales with actions per workflow, so activity granularity is a budget decision." |
| Workload-shape boundaries | "Processes on the engine, events on streams, AWS glue on Step Functions." |
| Lifecycle of long-running code | "Month-long workflows mean patches live for months; I set a maximum workflow age." |
Staff answers that L7 interviewers find insufficient:
- "Use Temporal for the saga" without who runs it, how teams learn it, and what it costs per process.
- "Add replay tests" as advice rather than an enforced merge gate with production history sampling.
- "Self-host to save money" without the persistence-store expertise and upgrade cost it implies.
How Real Companies Built It#
Uber — Cadence#
Uber built Cadence to let engineers write stateful, long-running services as ordinary code while the platform handles reliability, scaling and fault tolerance. Announcing Cadence 1.0 in June 2023, Uber reported 12 billion executions and 270 billion actions per month across more than 1,000 internal services, and internal surveys showing teams wrote about 40% less code for the same functionality (Uber engineering blog). Temporal's founders co-created Cadence at Uber after launching Amazon's Simple Workflow Service together (Temporal: who we are).
Staff insight: At that scale the engine is infrastructure every product depends on. In an interview, durable execution is a platform decision with a programming model attached, not a library choice.
Netflix — Conductor#
Netflix built Conductor to orchestrate workflows that span its microservices, with workflows defined in JSON or code and workers polling for tasks, then open-sourced it. In December 2023 Netflix announced it would discontinue maintenance of the open-source repository, and community forks continued development (Netflix/conductor on GitHub).
Staff insight: Orchestration engines outlive the teams that build them. Adopting one, especially a self-built or self-hosted one, is a multi-year ownership commitment; name who maintains it in year three.
Datadog — Temporal as an Internal Platform#
Datadog describes running Temporal in Kubernetes and contributing heavily back to the project: Temporalite, a single-binary Temporal using SQLite that it transferred to the Temporal organisation in 2022; a workflow trace CLI command; a Go SDK tracing interceptor; a large-payload codec; and a Kubernetes controller it built and open-sourced to manage Temporal workers, including tracking worker versions (Datadog open source).
Staff insight: The tooling Datadog invested in maps exactly onto the hard parts: local development, observability of long-running executions, oversized payloads and worker version management. Those are the topics to raise when an interviewer asks what it takes to run an engine for many teams.
Practice Drill#
Prompt: "Your company runs order fulfilment as Temporal workflows: ~400K orders a day, each workflow running up to 30 days to cover the return window. Last Tuesday a deploy reordered two steps and ~60,000 in-flight orders stopped progressing. On Thursday the shipping provider had a 40-minute outage, and when it recovered it was knocked over again by retries. Leadership asks whether you should move off Temporal. What do you do?"
Staff Answer
Neither incident is the engine misbehaving; both are the engine faithfully enforcing contracts we didn't honour, so I'd fix the contracts before considering a migration. Tuesday was a non-determinism failure: reordering activities changed the command sequence for executions replaying old histories. Immediate fix: roll back the worker build so the 60,000 executions resume on the old code, then re-ship the change behind a patch marker (or as a new worker build with worker versioning) so old executions keep the old order and new ones take the new path. Prevention: a CI gate that replays the new code against a few thousand histories sampled from production each day, plus a rule that the old branch is removed only when a visibility query shows zero open executions on it. With 30-day workflows, that means patches live at least 30 days, so I'd also look at shortening the workflow: end the fulfilment workflow at delivery and start a separate return-window workflow, which halves the patch lifetime. Thursday was a retry storm: the default policy retries forever with a 100-second cap, so tens of thousands of shipment activities hit the provider in waves. Fix: a dedicated task queue for the shipping activities with an activity rate limit at the provider's contracted rate (say 300/s), bounded retries with a Schedule-To-Close of a few hours, and a circuit breaker in the activity that fails fast while the provider is down. Backlog then drains at a controlled rate instead of all at once. Every shipping call carries an idempotency key of workflow ID plus step. On migrating: moving 400K orders a day with 30-day lifecycles to another engine means running two systems for at least a month and rewriting every workflow. The failures would follow us unless we fix replay testing and retry policy anyway. I'd report non-determinism incidents, shipping attempt rates and schedule-to-start latency for a quarter and revisit with data.
Why this is L6:
- Diagnoses both incidents from the engine's model (replay and default retry policy) instead of blaming the tool.
- Gives immediate recovery (roll back the worker build; rate-limit the task queue) and durable prevention (replay gate, bounded retries, per-dependency queues).
- Notices that a 30-day workflow length is itself a versioning cost and splits the lifecycle.
What L7 adds:
- Puts safe defaults into an internal SDK wrapper so no team can ship an activity without a timeout, bounded retry and idempotency key.
- Frames the migration question as a one-way door with a priced cost (dual running, rewrite, retraining) and a decision date tied to metrics.
- Establishes fleet-level retry budgets per external dependency so the next provider outage is absorbed by the platform, not by each team.
❌ Common L5 Trap
"Terminate the stuck workflows and restart them on the new code, and add a longer backoff to the retry policy."
Why this misses: Terminating and restarting 60,000 orders mid-flight re-runs completed steps (double reservations or charges unless every activity is idempotent) and loses their progress, when rolling back the worker build would have resumed them safely. A longer backoff still leaves retries unbounded and synchronised across executions, so the provider gets hit again on recovery. Neither fix prevents the next deploy from breaking replay.
Quick Reference Card#
Model: durable execution; workflow code replayed from event history
Workflow: deterministic orchestration; no I/O, no wall clock, no randomness
Activity: side effects; at-least-once; idempotent at the callee
Retry default: 1 s initial, x2 backoff, 100 s max interval, UNLIMITED attempts
workflows do not retry by default
Activity timeouts: Start-To-Close (always set), Schedule-To-Close (business deadline),
Schedule-To-Start (rarely), Heartbeat (long activities)
Workflow task: 10 s default timeout; keep workflow code fast
History: warn 10,240 events; terminate at 51,200 events or 50 MB
-> Continue-As-New; blobs by reference (2 MB payload limit)
Pending per run: 2,000 activities/children/signals; 10,000 signals total
Messages: signal (async, mutates), query (sync, read-only), update (sync, validated)
Versioning: patching (GetVersion / patched) or worker versioning; replay tests in CI
Step Functions: Standard 1 year, 25,000 events, 256 KiB payloads, $0.025/1K transitions
Express 5 minutes, priced per request and duration
Lineage: Amazon SWF -> Uber Cadence -> Temporal (same founders)
RED FLAGS
- time.Now(), random or HTTP calls in workflow code
- Workflow code changes without patching or worker versioning
- Activities without Start-To-Close or idempotency keys
- Default unlimited retries against external providers
- A workflow per Kafka message at millions per day