Hiring BarSupport

Temporal & Workflow Engines

Technology guide39 min read5 diagrams

Go deeper:

Why This Matters#

A workflow engine is not a job queue with extra steps. It is a database for the program counter: it durably records every step a piece of business logic has taken, so the logic can crash, be redeployed or sleep for 30 days and resume exactly where it left off. Temporal calls this durable execution. The interesting failures are not "the worker died" (the engine handles that); they are the ones the engine enforces: workflow code that isn't deterministic and breaks on replay, a code change deployed under 50,000 running workflows, an activity retried forever because nobody set a timeout, an event history that hits 51,200 events and terminates the workflow.

That is why "we'll orchestrate it with Temporal" is a sentence interviewers push on. The L5 candidate draws a box labelled "workflow engine" between services. The L6 candidate says "the order workflow is deterministic code; every side effect is an activity with a Start-To-Close timeout of 30 seconds, a retry policy capped at 10 attempts with exponential backoff, and an idempotency key derived from the workflow ID; payment capture and inventory reservation have compensations; cancellation arrives as a signal; the 14-day return window is a durable timer; I version changes with patching and drain old workers before removing the old branch." The L7 candidate asks whether the company should run one engine or three, who owns the cluster, and what it costs when every team starts encoding business processes in a system only a few people know how to operate.

The L5 → L6 gap is not knowing what an activity is. It is knowing that the engine replays your workflow code from history, so determinism and versioning are not style rules; they are the correctness contract.

The L5 → L6 → L7 Contrast#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Use Temporal to orchestrate the services""Is this a long-running, multi-step process with waits, retries and compensations? If it's one step plus a retry, a queue is enough.""Which processes deserve durable execution, which engine is the company standard, and who runs it?"
CodeWrites the workflow like normal codeWorkflow code deterministic; all I/O, time and randomness through activities or SDK APIs; replay tests in CIOrg lint rules and replay-test gates; workflow code reviewed as a distinct discipline
Failure"Temporal retries automatically"Every activity has Start-To-Close, a bounded retry policy, non-retryable error types and idempotency at the calleeRetry budgets across the fleet so a downstream outage doesn't become a retry storm from 100K workflows
ChangeDeploys new codePatching or worker versioning for running workflows; keeps old workers until old executions drainVersion-lifecycle policy: max workflow age, Continue-As-New cadence, deprecation windows
Scale"It scales"History under ~10K events via Continue-As-New; payloads under 2 MB with references to blobs; task-queue backlog and schedule-to-start latency as signalsCapacity and cost of the engine as a platform: namespaces, actions per second, persistence store sizing
Ownership"Platform team runs Temporal"Platform owns the cluster; service teams own workflows, activities, timeouts and versioningDefines the paved road and the exceptions: when Step Functions, a queue or a state table is the right answer
Why "Code" separates levels

A Temporal worker doesn't keep your workflow's stack in memory forever. When a worker restarts, or a workflow's state is evicted from cache, the SDK re-executes the workflow function from the beginning and feeds it the recorded results from the event history. The docs are explicit: emitted commands are compared with the existing history, and if a command doesn't match, the execution fails with a non-determinism error (workflow definition). So if time.Now().Hour() < 12, rand.Int(), iterating a hash map in random order, or calling an HTTP API directly inside workflow code all produce different commands on replay. The Senior answer writes natural code and is surprised in production; the Staff answer routes all non-determinism through activities or SDK-provided time and random APIs and runs replay tests against real histories in CI; the Principal answer makes those tests a merge gate for every team.

Why "Failure" separates levels

Temporal's default activity retry policy is an initial interval of 1 second, backoff coefficient 2.0, maximum interval 100 seconds and unlimited attempts (retry policies). That is a sensible default for transient faults and a dangerous one for a permanent failure: a validation error retried forever, or 100,000 workflows each retrying a down payment provider every 100 seconds. The Staff answer sets a Start-To-Close timeout on every activity (the docs strongly recommend it, because it's how the server detects a crashed worker), caps attempts or Schedule-To-Close, marks business errors non-retryable, and makes every activity idempotent at the callee because "retried" means "possibly executed twice." The Principal answer adds fleet-level protection: rate limits per task queue and circuit breakers so the engine's reliability doesn't become the downstream's outage.

The 60-Second Pitch#

"Order fulfilment runs as a Temporal workflow per order, with the order ID as the workflow ID so duplicate starts are rejected. The workflow is deterministic code: reserve inventory, authorise payment, create shipment, then wait on a durable timer for the 14-day return window. Every call to another service is an activity with a 30-second Start-To-Close timeout, exponential backoff capped at 10 attempts, and an idempotency key of workflow ID plus step name, so a retried capture never charges twice. Failures after payment trigger compensations in reverse order. Customer cancellation arrives as a signal; the support UI reads status through a query. Large payloads go to object storage and only references go into history. Code changes ship with patching, and we keep old worker builds running until workflows on the old path complete. We run the engine as a platform service or use the managed cloud offering; service teams own their workflows. For a single fire-and-forget step, I'd use a queue instead."

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Business process orchestration (orders, payments, onboarding, KYC)Multi-step, minutes to months, compensations, human waitsWorkflow per entity, activities per side effect, signals for external events, durable timersNon-idempotent activities double-charge; un-versioned changes break running workflowsEvery process reaches a terminal state; no duplicate side effects
Infrastructure automation (deploys, provisioning, database operations)Long steps, external systems, operator approvalsWorkflows wrapping cloud APIs with heartbeating activities; signals for approvalsActivity hangs without heartbeat; retries hammer a cloud API's rate limitSafe to resume after any crash; operations auditable
High-volume pipelines and fan-out (per-message processing, batch jobs)Millions of short executions, throughput over durationOften a queue or stream instead; if Temporal, batch work per workflow and use child workflowsActions per second and history writes become the bottleneck and the billThroughput per dollar; at-least-once with idempotent consumers

🎯 Staff Move: "I'll use the engine for the order lifecycle, because it spans services, days and compensations. The per-event analytics stream stays on Kafka: millions of independent one-step messages don't need a durable program counter, and paying per workflow action for them would be the wrong trade."

The Staff Positions#

PositionRationale
Workflow code is deterministic; all I/O lives in activitiesReplay re-executes workflow code; any divergence fails the execution.
Every activity is idempotent and has a Start-To-Close timeoutRetries mean at-least-once; the timeout is how crashed workers are detected.
Retries are bounded and business errors are non-retryableDefault retries are unlimited; a permanent failure should fail fast and compensate.
The workflow ID is the business keyDuplicate starts for the same order are rejected; the ID is the natural idempotency anchor.
History stays smallContinue-As-New well before the 51,200-event limit; references, not payloads, in history.
Every change to running workflow code is versionedPatching or worker versioning; replay tests on production histories gate the deploy.
Use the engine for long, multi-step processes, not everythingA queue is cheaper and simpler for one-step, high-volume work.

Architecture & Internals#

Five internals change design decisions: workflows vs activities, event history and replay, task queues and workers, timers, signals, queries and updates, and the server and its persistence.

Workflows vs Activities#

WorkflowActivity
What it isOrchestration logic: the sequence, branching, waits, compensationsA single unit of work with side effects: call an API, write a DB, send an email
DeterminismMust be deterministicMay do anything
DurabilityState reconstructed from event historyResult recorded in history once it completes
DurationSeconds to years (no timeout by default)Bounded by Start-To-Close; long ones heartbeat
Failure handlingDoesn't retry by default; handles activity failures in codeRetried by policy (default: unlimited attempts)
Where it runsYour worker process, replayed as neededYour worker process, possibly on a different host each attempt

Event History and Replay#

Diagram: Event History and Replay

Every state change is an event appended to the workflow's history: started, activity scheduled, activity completed, timer fired, signal received. On replay the SDK re-runs workflow code and, instead of executing activities again, returns their recorded results. That is how a workflow survives a crash in the middle of a 30-day wait without anyone writing checkpoint code.

Limit (Temporal Cloud and server defaults)ValueDesign consequence
Event history warning10,240 eventsPlan Continue-As-New well before this
Event history hard limit51,200 events or 50 MBExecution is terminated beyond it
Single payload / request2 MBStore blobs in object storage, pass references
Incomplete activities, child workflows, signals per execution2,000Fan out with child workflows in batches
Signals per execution10,000High-frequency signals need batching or a new run
Updates in history2,000 total, 10 concurrentUse update IDs and validators

Sources: event history limits, Temporal Cloud limits.

Continue-As-New closes the current run and starts a fresh one with the same workflow ID, carrying state forward as input. A subscription workflow that bills monthly forever, or an entity workflow processing a signal stream, continues-as-new every N iterations to keep history small and replay fast.

Task Queues and Workers#

Workers are your processes. They long-poll task queues on the server for workflow tasks and activity tasks, run your code, and report results. The server never calls your code directly, so workers can sit in private networks and scale independently.

Diagram: Task Queues and Workers

Why it matters in design: a backlog on a task queue means workers can't keep up, which you see as rising schedule-to-start latency before you see failures. Separate task queues per activity type let you scale and rate-limit a slow dependency (a payment API capped at 200 requests per second) without starving everything else.

Timeouts#

TimeoutMeasuresDefaultStaff guidance
Activity Start-To-CloseOne attempt's durationSame as Schedule-To-CloseAlways set; p99 of the call × 2–3
Activity Schedule-To-CloseAll attempts, including retries∞Set to the business deadline for the step
Activity Schedule-To-StartTime waiting in the queue∞Rarely set; alert on the latency metric instead
Activity HeartbeatMax gap between heartbeatsDisabledSet for anything over ~1 minute; enables fast crash detection and progress checkpoints
Workflow Execution / RunWhole workflow / one run∞Usually leave unset; model deadlines as timers in code
Workflow TaskWorker processing one workflow task10 sKeep workflow code fast; no blocking I/O

Sources: activity timeouts, workflow timeouts.

Timers, Signals, Queries and Updates#

PrimitiveDirectionSync?Mutates state?Recorded in history?Use for
Durable timer (sleep, await with timeout)Internal—YesYesReturn windows, retries with long waits, reminders
SignalInto workflowAsync, fire-and-forgetYesYesCancellation, approvals, external events
QueryRead from workflowSyncNoNoStatus pages, debugging; works on completed workflows
UpdateInto workflow, with a resultSyncYesYes"Add item to cart and return the new total"; validators reject bad requests before they're recorded

Source: message passing.

🎯 Staff Insight: "A durable timer is one of the most underrated primitives. 'Wait 14 days, then close the return window unless a return signal arrived' is five lines of workflow code instead of a cron job, a state table, a sweeper and the bugs between them."


Data Modeling / Core Usage — "The Entire Game"#

In DynamoDB the game is the access-pattern table. In a workflow engine it is the workflow contract: what one workflow execution represents, where the side-effect boundaries are, and how the code will change while executions are still running.

Step 1: One Workflow per Business Entity Lifecycle#

Good workflow IDs (the business key, so duplicate starts are rejected):
  order-{order_id}            order lifecycle: reserve -> pay -> ship -> return window
  subscription-{account_id}   long-lived entity; Continue-As-New every billing cycle
  deploy-{service}-{build}    infrastructure automation with approvals

Bad:
  random UUID per start       -> a retried "start" creates a second order workflow
  one workflow for all orders -> unbounded history, a single hot execution
  one workflow per HTTP call  -> durable execution for work that never waits

Step 2: Draw the Activity Boundaries#

Every side effect is an activity, and every activity is idempotent at the callee using a key the workflow can reproduce on replay.

// Workflow: deterministic orchestration only
func OrderWorkflow(ctx workflow.Context, order Order) error {
    ao := workflow.ActivityOptions{
        StartToCloseTimeout: 30 * time.Second,
        RetryPolicy: &temporal.RetryPolicy{
            InitialInterval:        time.Second,
            BackoffCoefficient:     2.0,
            MaximumInterval:        time.Minute,
            MaximumAttempts:        10,
            NonRetryableErrorTypes: []string{"CardDeclined", "InvalidAddress"},
        },
    }
    ctx = workflow.WithActivityOptions(ctx, ao)
    key := workflow.GetInfo(ctx).WorkflowExecution.ID   // stable across replays

    var resv Reservation
    if err := workflow.ExecuteActivity(ctx, ReserveInventory, key+"-reserve", order).Get(ctx, &resv); err != nil {
        return err                                       // nothing to compensate yet
    }
    var auth Auth
    if err := workflow.ExecuteActivity(ctx, AuthorizePayment, key+"-auth", order).Get(ctx, &auth); err != nil {
        _ = workflow.ExecuteActivity(ctx, ReleaseInventory, key+"-release", resv).Get(ctx, nil)
        return err                                       // compensate in reverse order
    }
    // ... capture, ship; then a durable 14-day timer races a "return" signal
    returnCh := workflow.GetSignalChannel(ctx, "return-requested")
    sel := workflow.NewSelector(ctx)
    sel.AddReceive(returnCh, func(c workflow.ReceiveChannel, more bool) { /* start refund */ })
    sel.AddFuture(workflow.NewTimer(ctx, 14*24*time.Hour), func(f workflow.Future) { /* close window */ })
    sel.Select(ctx)
    return nil
}
Inside workflow codeInside an activity
Branching on activity results and signalsHTTP calls, database writes, queue publishes
workflow.Now(), workflow.Sleep(), SDK timerstime.Now(), real sleeps, polling loops
SDK side-effect APIs for one-off random valuesRandom IDs, UUID generation for external systems
Deterministic iteration over sorted collectionsReading files, environment variables, config services
Small state (IDs, statuses)Large payloads: write to object storage, pass the key

Step 3: Plan for Change Before the First Deploy#

Running workflows replay against whatever code the worker has. Changing the sequence of commands for executions already in flight causes non-determinism errors. Two approaches:

ApproachHow it worksUse whenCost
Patching (GetVersion in Go and Java, patched() in TypeScript and Python)Code branches on a marker recorded in history: old executions take the old path, new ones the new pathSmall changes to long-running workflowsBranches accumulate; remove old ones only after old executions finish
Worker versioningWorkers are tagged with a build; executions stay pinned to (or are deliberately moved between) buildsFrequent deploys, many short-to-medium workflowsOld worker builds must keep running until their executions drain
New workflow typeOrderWorkflowV2; new starts use it, old ones finish on V1Large rewritesTwo codepaths live until V1 drains

The Temporal docs recommend worker versioning and describe patching as the alternative (workflow definition). Either way the discipline is the same: replay tests that run new code against histories exported from production, as a CI gate.

Diagram: Step 3: Plan for Change Before the First Deploy

🎯 Staff Move: "Activity code changes freely, because activities aren't replayed. Workflow code changes are versioned, replay-tested against a sample of production histories in CI, and the old branch is deleted only after a visibility query shows no open executions on it."

Step 4: Keep History Small#

History budget: stay under ~10K events per run (warning at 10,240; hard stop at 51,200)

Each activity costs ~3 events (scheduled, started, completed); each timer ~2; each signal 1.
  Order workflow: ~12 activities + 2 timers + few signals  -> ~45 events: fine
  Entity workflow receiving 50 signals/day                  -> ~200 days to warning
     -> Continue-As-New every 1,000 signals, carrying current state as input
  Batch workflow over 100K items as activities              -> 300K events: impossible
     -> child workflows of 1,000 items each, or activities that process pages

The Tunable Tradeoff — Durability × Latency × Cost#

Every step you make durable is a write to the engine's persistence store. Durability is the product; it is also the bill and the latency.

SettingLight endHeavy endWho pays at the light end
GranularityFew large activitiesOne activity per tiny stepOn-call, when a large activity fails halfway and isn't idempotent
Engine vs codeQueue + state table, hand-rolledFull workflow engineEngineers maintaining sweepers and state transitions
History sizeContinue-As-New oftenOne run for the whole lifecycleReplay latency and history limits
Activity timeoutsLong, generousTight, with heartbeatsDetection time when a worker dies mid-activity
RetriesUnlimited (default)Bounded with non-retryable errorsDownstream services during an outage
HostingSelf-hosted clusterManaged cloud servicePlatform team's pager; or finance, per action
Latency floor (illustrative, typical self-hosted or managed deployments):
  each activity round trip = schedule + dispatch + execute + record
    engine overhead per step ~ 10-50 ms on top of the work itself
  10-step workflow in the request path -> 100-500 ms of orchestration overhead
  -> keep synchronous user-facing paths short; start the workflow, return,
     and let the UI poll a query or subscribe to updates

Cost shape (managed services bill per action or state transition):
  Step Functions Standard: $0.025 per 1,000 state transitions (US East list price)
  1 M orders/day x 15 transitions = 15 M/day -> ~$375/day -> ~$11K/month
  same volume as per-message processing at 50 M messages/day x 3 steps -> ~$112K/month
  -> durable execution is cheap per business process, expensive per event

Step Functions pricing is from the AWS pricing page; Temporal Cloud bills per action with a different rate card, so model your own volumes in the Cost Estimator.

🎯 Staff Move: "I put one activity around each external side effect, not around each line of code. That's where idempotency keys and compensations naturally live, and it keeps both history and the per-action bill proportional to business steps, not to code size."

Who Pays for Each Choice#

ChoiceWhat WorksWhat BreaksWho Pays
Unlimited default retriesTransient faults heal themselvesPermanent errors retried forever; retry storms on outagesDownstream teams and their on-call
Non-idempotent activitiesSimpler activity codeDouble charges, duplicate emails on retryCustomers and finance
No versioning disciplineFaster first deploysNon-determinism errors stall in-flight workflowsService on-call during deploys
Huge payloads in historyConvenient2 MB limit, slow replay, persistence growthPlatform team
Synchronous workflows in the request pathSimple APIEngine latency added to user requestsUsers

Anti-Patterns — What Kills Workflow Engine Deployments#

1. Non-Deterministic Workflow Code#

time.Now(), random numbers, map iteration order, direct HTTP calls or reading config inside workflow code. Works in testing; fails on the first replay after a worker restart. Fix: SDK APIs for time and randomness, all I/O in activities, replay tests in CI.

2. Unversioned Changes Under Running Workflows#

Reordering two activities and deploying while 40,000 orders are mid-flight. Every one of them hits a non-determinism error on its next workflow task. Fix: patching or worker versioning; replay tests on production histories; keep old builds until executions drain.

3. Activities Without Timeouts or Idempotency#

No Start-To-Close means a crashed worker is detected late or never; no idempotency key means a retry after a timeout charges the card twice. Fix: Start-To-Close on every activity, idempotency keys from workflow ID plus step, heartbeats for long work. See Idempotency.

4. Unbounded Retries Against a Down Dependency#

The default policy retries forever with a 100-second ceiling; 200,000 workflows retrying a payment provider during its outage is a self-inflicted DDoS when it comes back. Fix: bounded attempts or Schedule-To-Close, rate-limited task queues per dependency, circuit breaking in the activity.

When thousands of executions retry the same recovering dependency on the same schedule, the retries themselves keep it down; jitter, caps and rate limits break the cycle.

5. The Workflow as a Database#

Storing large documents, growing lists or every message of a chat in workflow state. History balloons toward 50 MB and replay slows to seconds. Fix: state belongs in a database; the workflow holds IDs and statuses; Continue-As-New for long-lived entities.

6. Durable Execution for Everything#

Wrapping every API call or every Kafka message in a workflow. Throughput is capped by actions per second, and the bill scales with event volume. Fix: queues and streams for one-step high-volume work; workflows for multi-step processes. See Kafka vs SQS vs RabbitMQ.

7. Blocking Work in Workflow Code#

CPU-heavy computation or blocking calls in the workflow function exceed the 10-second workflow task timeout and cause repeated task failures. Fix: move heavy work into activities; keep workflow code a fast decision function.

8. No Owner for Stuck and Failed Workflows#

Executions that fail or sit waiting on a signal that will never come accumulate silently. Fix: visibility queries for failed and long-open executions per workflow type, a named owner, and runbooks for terminate, reset or signal.


The Technology Landscape — Head-to-Head Comparison#

DimensionTemporal (self-hosted or cloud)AWS Step FunctionsHand-rolled state machine (DB + queue + sweeper)Netflix Conductor-style JSON orchestratorsDAG schedulers (Airflow and similar)
ModelWorkflows as code, replayed from historyJSON/ASL state machine, managedStatus column, transitions in code, queue for workWorkflow definitions in JSON, workers poll tasksBatch DAGs on a schedule
Long waitsDurable timers, signals; no execution limit by defaultStandard: up to 1 year; Express: 5 minutesRows plus a sweeperSupportedNot the model
History limit51,200 events / 50 MB per runStandard: 25,000 eventsWhatever you buildEngine-specificN/A
Payload limit2 MB per payload256 KiB per state input/outputYour DBEngine-specificXCom-style small values
Versioning running workPatching, worker versioningVersions and aliases; executions keep the version they started onYour migration codeDefinition versionsDAG versions per run
Ops burdenHigh self-hosted (persistence store, upgrades); low on cloudNoneMedium, spread across every teamMedium–highMedium
Pick whenComplex, long-running logic best expressed in code; multi-cloudAWS-native, mostly service integrations, modest logicOne or two simple flows, strong DB skillsExisting investment, JSON-defined flowsScheduled data pipelines

Step Functions limits are from the service quotas page.

🎯 Staff Insight: "Every team that hand-rolls state machines ends up building a worse workflow engine: a status column, a retry counter, a sweeper cron and a dead-letter table. I'll hand-roll for one simple flow; by the third, I want a real engine and one team that's good at running it."


Patterns#

Pattern 1: Saga with Compensations#

Each forward step registers its compensation; on failure the workflow runs compensations in reverse. The engine guarantees the compensations themselves are retried until they succeed or escalate. Full treatment in Workflows, Sagas & Compensation; payment specifics in Payments.

Pattern 2: Entity Workflow#

One long-lived workflow per entity (an account, a cart, a device) receives signals and updates, holds a small state, and continues-as-new every N events. It serialises operations per entity without locks, the same property a per-key partition gives in Kafka.

Pattern 3: Human-in-the-Loop Approval#

The workflow sends a notification activity, then waits on an approval signal raced against a timer (escalate after 48 hours, auto-reject after 7 days). This replaces an approvals table, a reminder cron and an expiry sweeper.

Pattern 4: Fan-Out with Child Workflows#

Diagram: Pattern 4: Fan-Out with Child Workflows

Each level keeps history bounded and pending children under the 2,000 limit; activities process pages, not single items.

Pattern 5: Outbox to Workflow Start#

The service writes its business row and an outbox row in one transaction; a relay starts the workflow with the business key as workflow ID, so relay retries are deduplicated by the engine. See Transactional Outbox.


Scaling#

The Numbers#

ResourceDocumented limit or defaultDesign note
Event history per runWarning 10,240 events; terminated at 51,200 events or 50 MBContinue-As-New early
Payload per request2 MBReferences to blobs, not blobs
Pending activities / children / signals per execution2,000Batch fan-out through child workflows
Workflow task timeout10 s defaultWorkflow code must be fast and non-blocking
Activity retry default1 s initial, ×2, max 100 s, unlimited attemptsAlways set your own bounds
Temporal Cloud namespace rate500 actions per second default, scaling with capacity modeSize namespaces per domain; request increases ahead of launches
Step Functions Standard1-year max duration, 25,000 history events, 256 KiB payloadsExpress for < 5 min, high-volume flows
Step Functions Standard transitions5,000/s bucket in the largest regions (soft quota)Throttling shows as ExecutionThrottled

Scaling Moves in Order#

  1. Shrink per-workflow cost: fewer, coarser activities around real side effects; references instead of payloads; Continue-As-New for long-lived entities.
  2. Scale workers horizontally per task queue; tune poller counts and concurrent-execution slots; watch schedule-to-start latency.
  3. Split task queues by dependency so a slow or rate-limited downstream gets its own workers and limits.
  4. Split namespaces by domain (payments, fulfilment, infra) for rate limits, retention and blast radius.
  5. Scale the engine: more history shards and persistence capacity when self-hosting, or a higher capacity tier on a managed service. History shard count is set at cluster creation, so pick it for year-three volume.
  6. Move high-volume one-step work off the engine onto queues or streams.

Failure Modes & Recovery#

1. Non-Determinism Errors After a Deploy#

  • Symptom: Minutes after a deploy, workflow tasks fail repeatedly for a workflow type; executions stop progressing.
  • Root cause: Workflow code changed the command sequence without a version marker, or introduced non-deterministic calls.
  • Detection: Workflow task failure rate by type and build; non-determinism error counts in worker logs; executions stuck with pending workflow tasks.
  • Fix: Roll back the worker build (executions resume on the old code); re-ship with patching.
  • Prevention: Replay tests on production histories in CI; worker versioning; review checklist for workflow code.

2. Retry Storm on a Recovering Dependency#

  • Symptom: A payment provider recovers from an outage and immediately falls over again; activity failure rates stay at 100%.
  • Root cause: Tens of thousands of executions retrying in sync under an unbounded policy.
  • Detection: Activity attempt count distribution; downstream request rate vs normal; activity failure rate by type.
  • Fix: Throttle the task queue's activity rate; pause or scale down activity workers for that dependency; let backlog drain at a controlled rate.
  • Prevention: Bounded retries, task-queue rate limits per dependency, circuit breakers in activities.

3. History Limit Termination#

  • Symptom: Long-lived workflows terminate unexpectedly; the warning at 10,240 events was in logs nobody read.
  • Root cause: Entity workflow accumulating signals, or a loop of activities with no Continue-As-New.
  • Detection: History length percentiles by workflow type; server warnings for large histories.
  • Fix: Reset or restart affected entities from their last known state; add Continue-As-New.
  • Prevention: History-length alert at 5,000 events; design reviews that budget events per lifecycle.

4. Worker Backlog — Schedule-to-Start Latency Climbs#

  • Symptom: Orders take minutes instead of seconds to progress; no errors.
  • Root cause: Too few workers or slots, a deploy that reduced pollers, or one slow activity type hogging shared workers.
  • Detection: schedule_to_start latency p95 per task queue; task-queue backlog; worker slot utilisation.
  • Fix: Scale workers; move the slow activity to its own task queue.
  • Prevention: Autoscale workers on backlog and schedule-to-start latency; separate queues per dependency.

5. Engine or Persistence Outage#

  • Symptom: Starts, signals and queries fail; all workflows pause.
  • Root cause: Persistence store (Cassandra, MySQL, Postgres) overloaded or down, or the managed service region impaired.
  • Detection: Frontend error rates and latency; persistence latency; synthetic start-and-complete canary workflow every minute.
  • Fix: Restore the store; workflows resume from history once the engine is back, because state is durable and timers fire late rather than not at all.
  • Prevention: Persistence sized with headroom; managed service for teams without database operators; callers degrade gracefully (accept the order, start the workflow later from an outbox).

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Non-determinism after deployWorkflow task failures by buildAll running executions of the typeRoll back worker build, re-ship patchedService team
Retry stormAttempt counts, downstream RPSDownstream service and every workflow using itRate-limit task queue, pause workersService team + downstream owner
History limitHistory length percentilesLong-lived executions of the typeContinue-As-New, resetService team
Worker backlogSchedule-to-start latencyLatency for that task queueScale workers, split queuesService team; platform for autoscaling
Engine outageCanary workflow, frontend errorsEvery workflow in the cluster or namespaceRestore persistence; outbox buffering at callersPlatform team

When to Use vs. Alternatives#

NeedPickWhy
Multi-step business process spanning services, days and compensationsTemporal (or a similar code-first engine)Durable state, timers, retries and versioning built in
AWS-native glue between managed services with simple logicStep FunctionsNo infrastructure; direct service integrations
One background task with retryA queue with a DLQCheaper and simpler; no orchestration needed
High-volume event processingKafka plus stream processingThroughput per dollar; ordering per key; see Kafka
Scheduled data pipelinesA DAG schedulerBuilt for batch dependencies and backfills
One simple flow, strong DB team, no engine yetState table + outbox + sweeperFine for one or two flows; plan to replace by the third

When NOT to Use a Workflow Engine#

  • Single-step jobs. "Send this email with retries" is a queue with a DLQ, not a durable program.
  • Sub-50 ms request paths. Engine round trips add tens of milliseconds per step; check against the Latency Budget.
  • Per-event processing at millions per second. Per-action pricing and history writes make streams the right tool.
  • Teams that won't learn the determinism rules. The engine enforces them in production whether the team learned them or not.
  • As the system of record. Workflow state is retained for a configured period (Temporal Cloud: 1–90 days after close); durable business data belongs in a database.

Operational Concerns#

What the On-Call Actually Does#

  1. Watches failed and stuck executions per workflow type through visibility queries; decides reset, signal, terminate or fix-forward.
  2. Watches schedule-to-start latency and backlog per task queue; scales workers or splits queues.
  3. Gates deploys on replay tests and rolls back worker builds on non-determinism errors.
  4. Rate-limits activities against struggling dependencies instead of letting retries pile up.
  5. Runs the canary workflow (start, one activity, timer, complete every minute) as the engine's end-to-end health check.

Key Metrics & Alerts#

MetricHealthyAlert
Canary workflow end-to-end latency< 2 sFailure or > 10 s for 3 runs
Schedule-to-start latency p95 per task queue< 1 s> 30 s for 5 min
Workflow task failures (non-determinism)0Any after a deploy (page)
Activity failure rate by typeBaseline3× baseline for 10 min
Executions open longer than expected per typeNear 0Growth day over day
History length p99 per type< 5,000 events> 10,000

Interview Application — Staff-Level Plays#

Which Case Studies Use Workflow Engines#

Case StudyHow the Engine Is UsedKey Pattern
PaymentsAuthorise, capture, refund as a sagaIdempotency keys per step; compensations
Hotel BookingHold → pay → confirm with expiring holdsDurable timer releases the hold
Ride-HailingTrip lifecycle from request to payoutEntity workflow per trip; signals for state changes
Deployment SystemWaves, bake times, approvals, rollbackHuman-in-the-loop signals; long activities with heartbeats
Distributed Job SchedulerExecution of scheduled multi-step jobsScheduler of record plus workflow execution
Webhook DeliveryRetry schedules over hours or daysBounded retries, per-endpoint rate limits
Ledger & WalletTransfers across internal and external railsSaga with reconciliation, ledger as source of truth

Every System Design Question Has a Workflow Moment#

  • E-commerce checkout: "The order is a workflow keyed by order ID: reserve, authorise, capture, ship, then a 14-day return timer. Each step is an idempotent activity with a bounded retry policy."
  • User onboarding: "KYC is a workflow that waits up to 7 days for document upload signals, with reminder activities at 24 and 72 hours."
  • Data deletion (GDPR): "A deletion workflow fans out to 15 services as activities, retries each with a deadline, and records completion per service for the audit."

What Interviewers Probe#

After You Say...They Will Ask...What They're Evaluating
"Temporal handles retries""What if the activity charged the card and then timed out?"Idempotency at the callee
"Write the workflow in normal code""What happens when the worker restarts mid-workflow?"Replay and determinism
"We'll deploy a fix""What about the 30,000 workflows already running?"Versioning: patching, worker builds
"The workflow waits for the user""What if they never respond?"Timers raced against signals; terminal states
"Use a workflow engine for every message""What does that cost at 50 M messages a day?"Choosing queues vs workflows

Common Interview Mistakes#

What Candidates SayWhat Interviewers HearWhat Staff Engineers Say
"Temporal gives us exactly-once"Doesn't know activities are at-least-once"Workflow logic runs effectively once; activities are at-least-once, so the callee dedupes on a key."
"Retries are automatic"Unbounded retries against a down dependency"Bounded attempts, non-retryable business errors, rate-limited task queues."
"We'll just redeploy the workflow"Will break running executions"Patched or worker-versioned, replay-tested against production histories."
"Keep the cart in the workflow state"Workflow as a database"The workflow holds IDs; the cart lives in the database."

L5 vs L6 vs L7 Responses#

ScenarioL5 AnswerL6 / Staff AnswerL7 / Principal Answer
"Orchestrate a 6-service order flow"Temporal workflow calling each serviceWorkflow ID = order ID; idempotent activities with timeouts and bounded retries; compensations; signals and timers; versioning planDecides engine standard, hosting model and the platform contract for every team
"Deploy broke running workflows"Roll backRoll back the worker build, re-ship with a patch, add replay tests on production historiesMakes replay tests a merge gate org-wide and tracks non-determinism incidents to zero
"Downstream outage, workflows piling up"Wait for recoveryRate-limit the task queue, cap retries, drain backlog at a controlled rateFleet-wide retry budgets and per-dependency limits as platform features
"Step Functions vs Temporal"Temporal is more powerfulStep Functions for AWS-native integration with simple logic; Temporal for complex, long-running codePrices ops headcount vs per-transition cost and lock-in; sets the default and exceptions

The Staff Workflow Engine Checklist#

  1. Fit: "Multi-step, long-running, with waits and compensations? Otherwise a queue."
  2. Identity: "Workflow ID is the business key, so duplicate starts are rejected."
  3. Activities: "One per external side effect; idempotency key from workflow ID plus step; Start-To-Close on all."
  4. Retries: "Exponential, capped at 10 attempts or a Schedule-To-Close deadline; business errors non-retryable."
  5. History: "Under 10K events per run; Continue-As-New for entities; blobs by reference."
  6. Change: "Patching or worker versioning, replay tests in CI, old builds kept until executions drain."

🎯 Staff Insight: Don't use a workflow engine as a database, as a per-event stream processor, or in a sub-50 ms request path. The strongest signal is explaining replay: why workflow code must be deterministic, and what you do with the executions already running when the code changes.

Evaluation Rubric#

DimensionSenior (L5)Staff (L6)Principal (L7)
Model"Orchestrates services"Replay from history; workflow vs activity boundariesTeaches the model as an org discipline with tooling
Failure"Retries automatically"Timeouts, bounded retries, idempotency, compensationsFleet retry budgets, dependency protection
Change"Redeploy"Patching, worker versioning, replay testsVersion lifecycle policy and deprecation windows
Scale"Scales"History budgets, payload references, task-queue designNamespaces, persistence capacity, cost per business process
FitEngine for everythingEngine vs queue vs Step Functions by workloadPaved road, hosting model, exceptions process

Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

At Staff level a workflow engine is a tool for one process. At Principal level it becomes where the company's business processes live, which makes it critical infrastructure with an unusual skill requirement: every engineer writing workflows must understand replay and versioning. The L7 question is "which processes belong on the engine, who runs it, and how do we keep 30 teams from each learning the determinism rules through an outage?"

🧭 Principal Move: "Adopting a workflow engine is adopting a programming model. I'd fund the platform team, the SDK wrappers that enforce our retry and timeout defaults, and the replay-test gate before I'd let a second team onto it."

The Org-Level Fault Line#

One engine standard vs best-tool-per-team.

OptionWhat WorksWhat BreaksWho Pays
Every team picks (Step Functions, Temporal, hand-rolled, Airflow)Local fitN engines, N skill sets, no shared tooling or on-callIncident responders; future migrations
One engine, mandated for all async workShared expertise and toolingHigh-volume one-step work forced onto an expensive modelTeams with stream-shaped workloads
Paved road: one code-first engine for multi-step processes, queues/streams for events, managed state machines allowed for AWS glueClear defaults by workload shapeRequires a platform team and an exceptions processPlatform headcount

The Principal default: the paved road. One engine standard for long-running multi-step processes, with an internal SDK wrapper that sets safe defaults (Start-To-Close required, bounded retries, idempotency helpers, payload size checks); queues and streams for event processing; Step Functions permitted for AWS-native glue owned by small teams.

Cost Model#

Assumptions: managed engine billed per action, assumed at ~$25–50 per million actions for illustration (check the provider's current rate card); self-hosted cluster on a SQL or Cassandra store; loaded engineer ~$250K/year. Directional only.

ScaleWorkloadManaged engine/monthSelf-hosted infra/monthPlatform headcountTotal/month (managed)
Startup50K workflows/day × 30 actions~$1–2K~$1–2K plus ops0.25 FTE~$6–8K
Growth2 M workflows/day × 40 actions~$60–120K~$15–30K2 FTE (~$42K)~$100–160K
Enterprise30 M workflows/day × 40 actionsNegotiated; at the assumed rates ~$0.9–1.8M~$150–300K6–8 FTE (~$150K)~$400K–1M+

The crossover is the decision: below growth scale, managed hosting is cheaper than the engineers needed to run the persistence store; at enterprise scale, self-hosting or negotiated pricing wins, and action count per workflow is the lever that matters most on either path.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to Reverse
Business logic written against an engine's SDKOne-way-ishRewrite workflows; drain or migrate running executions
History shard count (self-hosted)One-way for the clusterNew cluster and migration
Long-running workflows (months) on a given designOne-way for their lifetimePatches live until the last one completes
Managed vs self-hostedTwo-way-ishNamespace migration; weeks to months

The Standard I'd Write#

RFC-WF-003: Durable Workflow Baseline

Scope: Every workflow running on the company workflow engine.

MUST
  1. Use the business entity key as the workflow ID.
  2. Keep workflow code deterministic: time, randomness and I/O only through SDK APIs
     or activities. Replay tests against sampled production histories gate every merge.
  3. Set Start-To-Close on every activity; bound retries by attempts or
     Schedule-To-Close; declare non-retryable business error types.
  4. Make every activity idempotent with a key derived from workflow ID and step.
  5. Version every change to workflow command order (patching or worker versioning);
     keep old worker builds until no open executions use them.
  6. Keep payloads under 256 KB (blobs by reference) and runs under 10,000 events.
SHOULD
  7. Use a dedicated task queue per rate-limited dependency.
  8. Use queues or streams, not workflows, for single-step work above 1,000/s.

Exceptions: platform lead approval, reviewed quarterly.
Success metrics: zero non-determinism incidents per quarter; no activity without a
  timeout; schedule-to-start p95 < 1 s; cost per business process reported per team.

What I'd Tell the VP#

"We're moving our core business processes, like orders, payments and onboarding, onto a single workflow platform so they finish reliably even when systems fail mid-way, without each team writing its own retry and recovery logic. The risk is that this platform is a new way of programming, and mistakes show up as stuck orders after a deploy. So I'm asking for a small platform team that provides safe defaults and automated checks every team must pass, and a clear rule for what doesn't belong on the platform, because putting every event through it would multiply our costs. Expect fewer stuck-order incidents and faster delivery of new processes, for two engineers and a managed-service bill that grows with order volume, not with traffic."

Principal Interview Signals#

SignalWhat It Sounds Like
Programming model, not a tool"Adopting the engine means teaching replay and versioning to every team; I fund that first."
Defaults in the SDK"Our wrapper won't compile an activity without a timeout or a bounded retry."
Prices actions"Cost scales with actions per workflow, so activity granularity is a budget decision."
Workload-shape boundaries"Processes on the engine, events on streams, AWS glue on Step Functions."
Lifecycle of long-running code"Month-long workflows mean patches live for months; I set a maximum workflow age."

Staff answers that L7 interviewers find insufficient:

  • "Use Temporal for the saga" without who runs it, how teams learn it, and what it costs per process.
  • "Add replay tests" as advice rather than an enforced merge gate with production history sampling.
  • "Self-host to save money" without the persistence-store expertise and upgrade cost it implies.

How Real Companies Built It#

Uber — Cadence#

Uber built Cadence to let engineers write stateful, long-running services as ordinary code while the platform handles reliability, scaling and fault tolerance. Announcing Cadence 1.0 in June 2023, Uber reported 12 billion executions and 270 billion actions per month across more than 1,000 internal services, and internal surveys showing teams wrote about 40% less code for the same functionality (Uber engineering blog). Temporal's founders co-created Cadence at Uber after launching Amazon's Simple Workflow Service together (Temporal: who we are).

Staff insight: At that scale the engine is infrastructure every product depends on. In an interview, durable execution is a platform decision with a programming model attached, not a library choice.

Netflix — Conductor#

Netflix built Conductor to orchestrate workflows that span its microservices, with workflows defined in JSON or code and workers polling for tasks, then open-sourced it. In December 2023 Netflix announced it would discontinue maintenance of the open-source repository, and community forks continued development (Netflix/conductor on GitHub).

Staff insight: Orchestration engines outlive the teams that build them. Adopting one, especially a self-built or self-hosted one, is a multi-year ownership commitment; name who maintains it in year three.

Datadog — Temporal as an Internal Platform#

Datadog describes running Temporal in Kubernetes and contributing heavily back to the project: Temporalite, a single-binary Temporal using SQLite that it transferred to the Temporal organisation in 2022; a workflow trace CLI command; a Go SDK tracing interceptor; a large-payload codec; and a Kubernetes controller it built and open-sourced to manage Temporal workers, including tracking worker versions (Datadog open source).

Staff insight: The tooling Datadog invested in maps exactly onto the hard parts: local development, observability of long-running executions, oversized payloads and worker version management. Those are the topics to raise when an interviewer asks what it takes to run an engine for many teams.


Practice Drill#

Prompt: "Your company runs order fulfilment as Temporal workflows: ~400K orders a day, each workflow running up to 30 days to cover the return window. Last Tuesday a deploy reordered two steps and ~60,000 in-flight orders stopped progressing. On Thursday the shipping provider had a 40-minute outage, and when it recovered it was knocked over again by retries. Leadership asks whether you should move off Temporal. What do you do?"

Staff Answer

Neither incident is the engine misbehaving; both are the engine faithfully enforcing contracts we didn't honour, so I'd fix the contracts before considering a migration. Tuesday was a non-determinism failure: reordering activities changed the command sequence for executions replaying old histories. Immediate fix: roll back the worker build so the 60,000 executions resume on the old code, then re-ship the change behind a patch marker (or as a new worker build with worker versioning) so old executions keep the old order and new ones take the new path. Prevention: a CI gate that replays the new code against a few thousand histories sampled from production each day, plus a rule that the old branch is removed only when a visibility query shows zero open executions on it. With 30-day workflows, that means patches live at least 30 days, so I'd also look at shortening the workflow: end the fulfilment workflow at delivery and start a separate return-window workflow, which halves the patch lifetime. Thursday was a retry storm: the default policy retries forever with a 100-second cap, so tens of thousands of shipment activities hit the provider in waves. Fix: a dedicated task queue for the shipping activities with an activity rate limit at the provider's contracted rate (say 300/s), bounded retries with a Schedule-To-Close of a few hours, and a circuit breaker in the activity that fails fast while the provider is down. Backlog then drains at a controlled rate instead of all at once. Every shipping call carries an idempotency key of workflow ID plus step. On migrating: moving 400K orders a day with 30-day lifecycles to another engine means running two systems for at least a month and rewriting every workflow. The failures would follow us unless we fix replay testing and retry policy anyway. I'd report non-determinism incidents, shipping attempt rates and schedule-to-start latency for a quarter and revisit with data.

Why this is L6:

  • Diagnoses both incidents from the engine's model (replay and default retry policy) instead of blaming the tool.
  • Gives immediate recovery (roll back the worker build; rate-limit the task queue) and durable prevention (replay gate, bounded retries, per-dependency queues).
  • Notices that a 30-day workflow length is itself a versioning cost and splits the lifecycle.

What L7 adds:

  • Puts safe defaults into an internal SDK wrapper so no team can ship an activity without a timeout, bounded retry and idempotency key.
  • Frames the migration question as a one-way door with a priced cost (dual running, rewrite, retraining) and a decision date tied to metrics.
  • Establishes fleet-level retry budgets per external dependency so the next provider outage is absorbed by the platform, not by each team.
❌ Common L5 Trap

"Terminate the stuck workflows and restart them on the new code, and add a longer backoff to the retry policy."

Why this misses: Terminating and restarting 60,000 orders mid-flight re-runs completed steps (double reservations or charges unless every activity is idempotent) and loses their progress, when rolling back the worker build would have resumed them safely. A longer backoff still leaves retries unbounded and synchronised across executions, so the provider gets hit again on recovery. Neither fix prevents the next deploy from breaking replay.


Quick Reference Card#

Model:             durable execution; workflow code replayed from event history
Workflow:          deterministic orchestration; no I/O, no wall clock, no randomness
Activity:          side effects; at-least-once; idempotent at the callee
Retry default:     1 s initial, x2 backoff, 100 s max interval, UNLIMITED attempts
                   workflows do not retry by default
Activity timeouts: Start-To-Close (always set), Schedule-To-Close (business deadline),
                   Schedule-To-Start (rarely), Heartbeat (long activities)
Workflow task:     10 s default timeout; keep workflow code fast
History:           warn 10,240 events; terminate at 51,200 events or 50 MB
                   -> Continue-As-New; blobs by reference (2 MB payload limit)
Pending per run:   2,000 activities/children/signals; 10,000 signals total
Messages:          signal (async, mutates), query (sync, read-only), update (sync, validated)
Versioning:        patching (GetVersion / patched) or worker versioning; replay tests in CI
Step Functions:    Standard 1 year, 25,000 events, 256 KiB payloads, $0.025/1K transitions
                   Express 5 minutes, priced per request and duration
Lineage:           Amazon SWF -> Uber Cadence -> Temporal (same founders)

RED FLAGS
  - time.Now(), random or HTTP calls in workflow code
  - Workflow code changes without patching or worker versioning
  - Activities without Start-To-Close or idempotency keys
  - Default unlimited retries against external providers
  - A workflow per Kafka message at millions per day
  1. Loading the index…