Why This Matters#
A work queue is not a messaging pipe. It is a contract about who owns a unit of work while it is being done. The broker hands a message to one consumer, stops offering it to anyone else for a while, and takes it back if the consumer goes quiet. Every interesting property (duplicates, ordering, retries, poison messages, backlog) comes from how that hand-off is timed and who is allowed to give the message back. RabbitMQ and SQS are the two queues interviewers meet most often: RabbitMQ is the broker you run, with routing built in; SQS is the queue you rent, with almost nothing to operate.
That is why "drop it on a queue" is a sentence interviewers push on. The L5 candidate draws a queue between the API and the workers and says "this decouples them." The L6 candidate says "SQS standard queue, visibility timeout of 2× the p99 job time with a heartbeat for long jobs, maxReceiveCount of 5 into a dead-letter queue that the payments team owns and alerts on, idempotent consumers keyed by job ID, and an alarm on the age of the oldest message, not queue depth." The L7 candidate asks how many queueing technologies the company should run, who owns the dead-letter queues across 60 teams, and what it costs when the async path silently backs up for six hours and nobody's dashboard turns red.
The L5 → L6 gap is not knowing what an exchange is. It is knowing that a queue converts a latency problem into a backlog problem, and a backlog nobody measures in minutes is an outage that hasn't been noticed yet.
The L5 → L6 → L7 Contrast#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Put a queue between the API and the workers" | "Is this a job done once, or an event several teams will read? A job is a queue; an event is a log. Then: SQS if we're on AWS, RabbitMQ only if we need broker routing or we're off-cloud." | "How many messaging systems does the company run, and does this team get to add one? Each broker is a platform with an owner and an on-call rota." |
| Delivery | "The queue guarantees delivery" | At-least-once everywhere; ack only after the side effect commits; consumers idempotent on a business key; FIFO dedup covers 5 minutes, not forever | Publishes the org's delivery contract: every consumer idempotent, every message carries an idempotency key, audited in review |
| Retries | "Retry failed messages" | Visibility timeout or delivery limit bounds retries; backoff via delay or TTL queues; DLQ after N attempts with a named owner and a redrive runbook | Treats DLQs as an org-wide inventory: count, age and owner per DLQ on one dashboard; unowned DLQs are a review failure |
| Scaling | "Add more consumers" | Little's law for consumer count; prefetch tuned to processing time; knows a RabbitMQ queue lives on one core; FIFO throughput per group is 1 in flight | Prices the async tier: SQS per-request cost vs broker nodes 24/7, and the headcount to run RabbitMQ |
| Failure | "Messages are persisted, so nothing is lost" | Names the backlog that can't drain, the visibility timeout shorter than the job, the memory alarm that blocks publishers, the poison message loop | Designs per-tenant isolation (fair queues, shuffle sharding) so one customer's burst doesn't become every customer's six-hour delay |
| Ownership | "The platform team runs the queue" | Producer owns the schema; consumer owns the DLQ and the redrive; platform owns the broker | Decides what is centralised (brokers, alarms, DLQ tooling) and what each team must declare (owner, retry policy, max age SLO) |
Why "Delivery" separates levels
"The queue guarantees delivery" is true and nearly useless. Both systems are at-least-once: SQS standard can deliver a message more than once, and any consumer that crashes after doing the work but before deleting the message will see it again when the visibility timeout expires. RabbitMQ redelivers every unacknowledged message when a channel closes. So the real guarantee is "every message is processed at least once, and you will see duplicates." The Staff answer moves the guarantee to where it can actually be kept: the consumer acks only after the side effect is durable, and the side effect is idempotent on a business key (charge_id, email_id) stored in the same transaction. SQS FIFO's deduplication helps with producer retries, but its window is 5 minutes, so it does not protect a consumer that is replayed tomorrow. The Principal answer makes idempotency a standard with a linter, because one non-idempotent consumer in 60 is how a customer gets charged twice.
The 60-Second Pitch#
"For background jobs on AWS I'd use SQS standard: no brokers, effectively unlimited throughput, and per-message retries for free. Each job type gets its own queue and its own dead-letter queue. The visibility timeout is about twice the p99 job time, and jobs that can run long extend it with a heartbeat. After 5 failed receives a message moves to the DLQ, which pages the owning team when its oldest message is older than 15 minutes. Consumers are idempotent on the job ID, because SQS is at-least-once. I'd use FIFO only where per-entity order matters, keyed by entity ID, and accept 300 calls per second per partition unless we turn on high-throughput mode. If we weren't on AWS, or needed topic routing to many queues, I'd run RabbitMQ 4 with quorum queues on 3 nodes, publisher confirms and manual acks. If anyone needs to replay these messages, it isn't a queue, it's a log, and I'd use Kafka instead."
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| Background job distribution | Each job done once; spiky load; per-job retry | SQS standard (or RabbitMQ quorum queue) + DLQ + idempotent workers; autoscale on age of oldest message | Visibility timeout shorter than the job → duplicate work; poison message retried forever | Every job completes or lands in a DLQ with an owner; duplicates harmless |
| Routed fan-out to many consumers | One event, several queues, routing by type or attribute | SNS → SQS with filter policies, or RabbitMQ topic/headers exchange → per-consumer queues | Binding churn or a forgotten queue fills the broker; a slow subscriber backs up only itself | Each subscriber gets its own copy and its own retry policy |
| Ordered per-entity workflows | Steps for one order must run in order | SQS FIFO with MessageGroupId = order_id, or RabbitMQ consistent-hash exchange to N single-consumer queues | One slow or poisoned message blocks its whole group; one hot group caps throughput | Strict order within an entity; no cross-entity head-of-line blocking |
🎯 Staff Move: "I'll design for the first intent: background jobs, one queue per job type. I'm assuming nobody needs to re-read these messages; if analytics wants them, that's a separate event stream, not a second consumer on this queue. Ordering I'll handle per entity only where the business needs it, because FIFO costs throughput and head-of-line blocking everywhere else."
The Staff Positions#
| Position | Rationale |
|---|---|
| Ack after the side effect, never before | Ack-on-receive is at-most-once; a crash between receive and commit loses the job silently. |
| Every consumer is idempotent on a business key | Both systems are at-least-once; duplicates are a certainty, not an edge case. |
| Alert on age of oldest message, not depth | 50,000 messages is fine at 10,000/s drain and a disaster at 10/s. Age is what the user feels. |
| Every DLQ has an owner, an alarm and a redrive runbook | A DLQ without an owner is a place where data goes to be forgotten. |
| Quorum queues for anything that matters on RabbitMQ | Classic queue mirroring was removed in RabbitMQ 4.0; quorum queues are the replicated, Raft-based default. |
| One queue per job type, not one queue for everything | Isolates retry policy, timeout, scaling and blast radius per workload. |
| If anyone needs replay, it's a log, not a queue | Acked messages are gone. A second reader next quarter needs Kafka, not a second queue. |
Architecture & Internals#
Only five internals change design decisions: the queue vs log model, RabbitMQ's exchanges and bindings, acknowledgements and prefetch, SQS's visibility timeout, and how each system replicates. The side-by-side with Kafka is in Kafka vs SQS vs RabbitMQ.
Queues vs Logs: Who Remembers What Was Read#
In a queue the broker tracks each message's state: ready, in flight, acknowledged (deleted). In a log like Kafka the broker tracks nothing per message; consumers keep an offset and data lives until retention removes it. That single difference decides replay, retry and parallelism. A second reader on a queue needs its own queue and a fan-out in front; on a log it just starts its own offset. Retrying one message is native on a queue and awkward on a log. Parallelism on a standard queue is "add consumers"; on a log it is bounded by partitions.
RabbitMQ also offers streams, an append-only log type with offset-based reads and replay, for when a RabbitMQ shop needs log semantics without adopting Kafka. Treat streams as "Kafka-shaped, inside RabbitMQ", and size them like a log.
RabbitMQ: Exchanges, Bindings and Queues#
Producers publish to an exchange (even "publishing to a queue" goes through the default exchange) with a routing key; bindings decide which queues receive a copy. Consumers read from queues.
| Exchange type | Routes by | Use for |
|---|---|---|
| direct | Exact routing key match | Work queues by job type |
| topic | Pattern on dot-separated key (* one word, # zero or more) | Event fan-out by type, region, tenant |
| fanout | Ignores the key, copies to every bound queue | Broadcast (cache invalidation, config push) |
| headers | Message header values | Routing on several attributes at once |
| consistent hash (plugin) | Hash of the routing key over bound queues | Per-entity ordering across N queues |
Why it matters in design: routing lives in the broker, so adding a subscriber is a new binding, not a producer change. The cost is that bindings are broker state. Trello's account of its RabbitMQ era (below) describes a flood of binding add and remove commands after mass socket disconnects making the cluster unresponsive, even to monitoring. Bindings are cheap to create one at a time and expensive to churn by the thousand.
Acknowledgements, Prefetch and Confirms#
Three settings decide RabbitMQ's delivery guarantee, and candidates usually mention only one:
| Setting | What it does | Staff default |
|---|---|---|
| Publisher confirms | Broker acks the publish once the message is safely enqueued (for quorum queues, after a majority of replicas have it) | On, for anything you can't lose; producer retries on nack or timeout, which can create duplicates |
| Consumer manual ack | Message stays unacked until the consumer acks; redelivered if the channel closes | Always manual for work queues; ack after the side effect commits |
Prefetch (basic.qos) | Caps unacked messages per consumer | Set from processing time: tens for fast jobs, 1–5 for slow ones |
RabbitMQ also enforces a delivery acknowledgement timeout: if a consumer holds a delivery unacked for longer than 30 minutes by default, its channel is closed with PRECONDITION_FAILED and the messages are redelivered (consumer docs). A job that legitimately runs for an hour needs a different design (ack, then track progress elsewhere), not a raised timeout.
Prefetch sizing (a common production pattern): a consumer needs enough messages buffered to cover the round trip between finishing one and receiving the next. With a 2ms network round trip and 20ms jobs, a prefetch of ~2–5 per consumer thread keeps it busy. With 200ms jobs, a prefetch of 1–2 avoids one consumer hoarding 100 messages that others could be working on. Unlimited prefetch on a slow consumer is how one stuck pod holds 10,000 messages hostage.
SQS: The Visibility Timeout Is the Whole Protocol#
SQS has no acknowledgement in the AMQP sense. ReceiveMessage hides a message for the visibility timeout; DeleteMessage with the receipt handle removes it; if neither happens in time, the message becomes visible again and its receive count goes up.
| SQS setting | Default | Limit | Design note |
|---|---|---|---|
| Visibility timeout | 30s | 0s – 12 hours | ≥ 2× p99 processing time, or heartbeat with ChangeMessageVisibility |
| Retention | 4 days | 60s – 14 days | Set the DLQ's retention longer than the source queue's |
| Delivery delay | 0s | 15 minutes | Longer delays need a scheduler, not a queue |
| Long-poll wait | 0s | 20s | Always use 20s; empty receives still cost a request |
| Batch size | — | 10 messages per send/receive/delete | Batching cuts request cost up to ~10× |
| Message size | — | 1 MiB | Larger payloads go to S3 with a pointer |
Limits from the SQS message quotas (as of 2026).
SQS FIFO adds two things. Message groups: messages with the same MessageGroupId are delivered in order, one in flight at a time; different groups proceed in parallel, and there is no limit on group count. Deduplication: a MessageDeduplicationId (or a content hash) suppresses duplicate sends within a 5-minute window (dedup docs). Throughput is 300 API calls per second per partition, 3,000 messages per second with batches of 10, and high-throughput mode raises it to up to 70,000 transactions per second in the largest regions (SQS quotas, as of 2026).
SQS fair queues (standard queues only) use MessageGroupId as a tenant identifier: when one tenant has a disproportionate share of in-flight messages, SQS prioritises other tenants' messages to protect their dwell time, with no ordering and no throughput cap (fair queues).
🎯 Staff Insight: "In SQS the visibility timeout is the lock, the lease and the retry timer at once. Set it shorter than the job and two workers do the same work. Set it ten times longer and a crashed worker's message sits invisible for ten times as long. So I size it from p99 job time and heartbeat for the long tail."
Replication and Durability#
| SQS | RabbitMQ quorum queue | RabbitMQ classic queue | |
|---|---|---|---|
| Replication | Managed, stored redundantly across multiple AZs | Raft over an odd number of members, 3 by default | None in 4.x (mirroring removed in 4.0) |
| Survives | AZ loss, transparently | Loss of a minority (1 of 3, 2 of 5) | Nothing; node loss = queue unavailable or lost |
| Publish ack means | Message stored | Majority of replicas have it (with confirms) | Leader has it |
| Cost | Per request | ~3× disk and network; fsync on every batch | Cheapest; for transient data only |
RabbitMQ's docs recommend an odd quorum-queue group size, note that performance degrades beyond 5 members, and advise against running quorum queues on more than 7 nodes. Each queue keeps at least 32 bytes of metadata per message in memory, roughly 1 MB per 30,000 messages regardless of size (quorum queue docs), so a 30-million-message backlog costs ~1 GB of RAM per replica just in bookkeeping.
Core Usage — "The Entire Game": The Message Contract#
In Kafka the game is the partition key. In a work queue it is the message contract: the five decisions that determine whether a job runs once, runs twice harmlessly, or runs twice and charges a customer.
The message contract (per queue):
1. ack point -> when is the message considered done?
2. lease length -> visibility timeout / prefetch / consumer timeout
3. retry policy -> how many attempts, what backoff, where it goes after
4. order scope -> none, per entity (group / hash), or global (almost never)
5. dedup key -> what makes a second delivery harmless
Step 1: The Ack Point — After the Side Effect#
WRONG (at-most-once): RIGHT (at-least-once + idempotent):
msg = receive() msg = receive()
delete(msg) # ack first begin tx
send_email(msg) # crash = lost if seen(msg.job_id): commit; delete(msg); return
do_side_effect(msg)
mark_seen(msg.job_id)
commit
delete(msg) # crash before this = harmless redelivery
The idempotency record and the side effect commit in one transaction where possible. When the side effect is an external call (a payment provider, an email API), pass the job ID as that system's idempotency key. The full treatment of keys, scopes and the crash window is in Idempotency.
Step 2: Size the Lease and the Consumer Fleet#
Example: thumbnail jobs
arrival rate lambda = 400 jobs/s at peak
processing time p50 = 0.8s, p99 = 4s
concurrency/pod 8 threads
consumers needed (Little's law): L = lambda x W = 400 x 0.8 = 320 jobs in flight
pods = 320 / 8 = 40 pods at 100% busy -> run 55-60 for ~70% utilisation
visibility timeout = 2 x p99 = 8s -> round to 30s (allow GC pauses, retries inside the job)
heartbeat: jobs over 20s call ChangeMessageVisibility(+30s) every 15s
DLQ: maxReceiveCount = 5
Utilisation matters more than the average suggests: queueing delay grows sharply above ~70–80% busy, which is why the fleet runs at 55–60 pods, not 40. The math is in Queueing Theory, and the Back-of-Envelope calculator does the arithmetic.
Step 3: Retry Policy and Dead Letters#
| Failure type | Example | Right handling |
|---|---|---|
| Transient | Downstream 503, timeout, throttling | Retry with backoff; let the visibility timeout or a delay queue space attempts |
| Poison | Malformed payload, missing record, code bug | Stop retrying fast: DLQ after a small N (3–5), alert, fix, redrive |
| Business rejection | Card declined, user deleted | Not an error: ack it and record the outcome |
SQS: RedrivePolicy { maxReceiveCount: 5, deadLetterTargetArn: thumbs-dlq }
backoff: on failure, ChangeMessageVisibility(min(2^attempt x 10s, 15min))
redrive: DLQ redrive back to the source once the fix ships
RabbitMQ: x-queue-type=quorum, x-delivery-limit=5, x-dead-letter-exchange=thumbs.dlx
backoff: nack into a wait queue with per-message TTL that dead-letters back
(quorum queues default the delivery limit to 20 since 4.0)
The DLQ is a product with an owner. It needs an alarm on age of oldest message (not on count, which a single bad deploy can push to 50,000 in a minute), a dashboard showing the top error reasons, and a redrive tool the on-call has practised.
Step 4: Ordering Scope — Usually None, Sometimes Per Entity, Never Global#
Global order on a queue means one consumer, one message at a time. Nobody actually needs that. What businesses need is per-entity order: the order's created before its paid before its shipped.
Two rules: ordering costs parallelism (one in flight per group or per shard), and retries break order unless the whole group blocks. SQS FIFO blocks the group while a message is in flight or being retried, which is correct and also means one poison message stops that entity until it reaches the DLQ. Often the cheaper design is to make consumers order-tolerant: carry a version number and ignore stale updates.
Step 5: The Dedup Key#
| Layer | Mechanism | Window |
|---|---|---|
| Producer → SQS FIFO | MessageDeduplicationId | 5 minutes |
| Producer → RabbitMQ | Publisher confirms + producer retry | None (retries create duplicates) |
| Consumer | Idempotency table keyed by job ID, TTL ≥ max retention + redrive time | As long as you keep it |
Only the consumer-side key covers every source of duplicates: producer retries, visibility timeout expiry, channel closes, DLQ redrive and replays. Broker dedup is a cost optimisation, not a correctness mechanism.
🎯 Staff Move: "I'll make the consumer idempotent on job ID, with the seen-record written in the same transaction as the side effect. FIFO dedup is nice for producer retries, but it's a 5-minute window and redrive from a DLQ happens days later, so it can't be the thing that keeps us correct."
The Tunable Tradeoff — Safety vs Throughput vs Latency#
Every queue setting moves along one axis: how much work does the broker do to be sure about a message?
| Setting | Fast end | Safe end | Who pays at the fast end |
|---|---|---|---|
| RabbitMQ queue type | Classic, transient messages | Quorum queue, persistent, 3 replicas | Whoever loses the node, as lost messages |
| Publisher confirms | Off (fire and forget) | On, with retry on nack | Producer team, in silent loss during broker trouble |
| Consumer ack | Auto-ack on delivery | Manual ack after commit | Users whose job vanished when a pod crashed |
| Prefetch | Unlimited | 1 | Other consumers, starved by a hoarder; at 1, throughput |
| SQS queue type | Standard | FIFO | Consumers, who must tolerate duplicates and reordering |
| Visibility timeout | Short (fast retry) | Long (no duplicate work) | Downstream systems, doing the same job twice |
| Retry count | Many (eventual success) | Few, then DLQ | The queue, clogged by poison messages |
🎯 Staff Move: "Quorum queues with confirms and manual acks for orders and payments; classic transient queues are fine for cache-invalidation broadcasts where a lost message costs one stale read. Safety where a lost message is money, speed where it's a refresh."
Anti-Patterns — What Kills RabbitMQ and SQS Deployments#
1. Ack on Receive#
Auto-ack (RabbitMQ) or delete-before-process (SQS) makes the queue at-most-once. Every pod crash, OOM kill or deploy drops whatever was in flight, and nothing alerts because nothing failed visibly. Fix: manual ack after the side effect commits; idempotent consumers to absorb the duplicates that follow.
2. Visibility Timeout Shorter Than the Job#
A 30-second default on a job whose p99 is 45 seconds means the slowest 1% run twice, then the duplicates slow the downstream further, pushing more jobs past 30 seconds. Duplicate rate climbs with load. Fix: size from p99 × 2, heartbeat with ChangeMessageVisibility for long jobs, and watch the ratio of receives to deletes.
3. The Unowned Dead-Letter Queue#
Messages land in a DLQ with a 4-day retention, nobody alerts on it, and on day 5 they are deleted. The data is gone and no incident was ever opened. Fix: every DLQ has an owning team, an age-of-oldest alarm, retention set to the 14-day maximum, and a tested redrive path.
4. One Queue for Every Job Type#
Password-reset emails wait behind a 2-million-message backfill, because both are "jobs." Fix: one queue per job type (or per priority tier), each with its own consumer fleet, timeout and DLQ. Separate queues are cheap on both systems.
5. Unbounded Backlogs on RabbitMQ#
RabbitMQ holds messages on the broker. A consumer outage lets queues grow until memory crosses the high watermark (60% of RAM by default) or disk falls below the free-space limit, at which point the broker blocks all publishing connections (configuration docs). An unrelated producer on the same cluster now times out. Fix: queue length limits (x-max-length with reject-publish or dead-lettering), alarms well before the watermark, and separate vhosts or clusters for workloads that must not block each other.
6. Queue-Per-Session or Binding Churn#
A queue and binding per user session or WebSocket feels natural with topic exchanges. At 100,000 sessions, a mass reconnect is 100,000 queue declarations and binding changes on a broker whose metadata operations are expensive. Fix: a small, fixed set of queues per consumer process with in-process filtering, or a log for fan-out at that scale.
7. Using a Queue When You Needed a Log#
The team ships an SQS queue for order events; next quarter analytics, search indexing and fraud all want the same events, plus a 3-day replay after a bug. A queue can't give any of that. Fix: ask "will anyone read this twice?" before choosing. SNS → SQS fan-out handles new readers going forward, but not history.
The Technology Landscape — Head-to-Head Comparison#
| Dimension | SQS standard | SQS FIFO | RabbitMQ (quorum queues) | RabbitMQ streams | Kafka |
|---|---|---|---|---|---|
| Model | Managed queue, visibility timeout | Managed queue, message groups | Broker with exchanges, Raft-replicated queues | Append-only log inside RabbitMQ | Replicated partitioned log |
| Ordering | Best effort | Per message group | Per queue, single consumer, no redelivery | Per stream | Per partition |
| Delivery | At-least-once | Exactly-once processing within 5-min dedup window | At-least-once with confirms + manual acks | At-least-once, offset-based | At-least-once; transactions Kafka→Kafka |
| Throughput | Nearly unlimited | 300 calls/s per partition; up to 70K TPS high-throughput (largest regions) | ~Tens of thousands msg/s per queue; shard queues to scale | Higher than queues; log-shaped | 100s of MB/s per cluster |
| Replay | No | No | No | Yes | Yes |
| Routing | Via SNS filter policies | Via SNS FIFO | Native: direct, topic, fanout, headers | Via exchanges | Topic per stream; routing in consumers |
| Ops burden | None | None | Medium: nodes, upgrades, memory alarms | Medium | High self-hosted, medium managed |
| Cost shape | Per request; scales to zero | Per request, slightly higher | Nodes 24/7 | Nodes + disk | Brokers + disk + cross-AZ traffic |
| Pick when | Background jobs on AWS | Per-entity ordered workflows on AWS | Off-AWS queues, broker routing, priorities | RabbitMQ shop needing replay | Multi-consumer event streams, replay |
🎯 Staff Insight: "The cheapest broker to operate is the one I don't operate. On AWS, SQS wins for job queues unless I need routing or protocol features it lacks. RabbitMQ earns its nodes when I'm off-cloud, need topic routing, or need AMQP for an existing ecosystem."
Patterns#
Pattern 1: Work Queue with DLQ and Redrive#
The default for 80% of async work: producer → queue → autoscaled workers → DLQ after N receives. Autoscale on age of oldest message (ApproximateAgeOfOldestMessage) or backlog per worker, not CPU. Redrive from the DLQ after the fix ships, throttled so the redrive itself doesn't overload the downstream.
Pattern 2: Fan-Out — SNS to SQS, or a Topic Exchange#
Each subscriber gets its own queue, timeout, retry policy, DLQ and scaling. A slow email provider backs up only email-q. RabbitMQ does the same with a topic exchange and one queue per subscriber. This is the queue world's answer to multiple consumer groups, minus replay.
Pattern 3: Per-Tenant Fairness#
One tenant's import of 5 million rows should not delay every other tenant's jobs for six hours. Options, cheapest first: SQS fair queues with MessageGroupId = tenant_id; shuffle sharding (hash each tenant to 2 of N queues, consumers drain the shorter); spillover queues for tenants over a rate; a queue per large tenant. Amazon's write-up on queue backlogs (below) describes all of these. The broader design is in Multi-Tenancy.
Pattern 4: Outbox to Queue#
Writing to the database and sending to a queue are two systems with no shared transaction. Write the message to an outbox table in the same transaction as the business change; a relay publishes it and marks it sent. Duplicates are possible, losses aren't. See Transactional Outbox.
Scaling#
The Numbers#
| Resource | Documented limit or typical figure | Design note |
|---|---|---|
| SQS standard throughput | Nearly unlimited API calls per second | Scale consumers freely |
| SQS FIFO throughput | 300 calls/s per partition; 3,000 msg/s with batches of 10; up to 70K TPS (700K msg/s batched) in the largest regions with high-throughput mode | Spread load across many group IDs |
| SQS FIFO in flight | 120,000 messages | Hitting it stalls receives without an error |
| SQS message size | 1 MiB | S3 pointer for larger (extended client) |
| SQS visibility / retention / delay | 12 h max / 14 days max / 15 min max | Long jobs heartbeat; long delays need a scheduler |
| RabbitMQ queue throughput | ~Tens of thousands msg/s per queue (typical; one Erlang process per queue) | Shard hot queues; consistent-hash exchange |
| RabbitMQ max message size | 16 MiB default | Keep messages small; payloads go to object storage |
| RabbitMQ memory alarm | 60% of RAM default | Publishers blocked cluster-wide when tripped |
| Quorum queue members | 3 default; odd; ≤ 5 recommended; ≤ 7 nodes | More members = slower writes |
| Quorum queue metadata | ~1 MB RAM per 30,000 messages | Backlogs cost memory even when messages are on disk |
| Consumer ack timeout | 30 min default | Long jobs ack early and track progress elsewhere |
Scaling Moves in Order#
- Batch and long-poll (SQS) or tune prefetch (RabbitMQ). Often a 3–10× efficiency gain before adding anything.
- Scale consumers from Little's law, autoscaling on backlog age.
- Split queues by job type and priority, so each scales and fails independently.
- Shard hot queues (RabbitMQ consistent-hash exchange; SQS FIFO with more group IDs or high-throughput mode).
- Separate clusters (RabbitMQ) for workloads that must not share a memory alarm.
- Move to a log when the real requirement turned out to be replay or many readers.
🎯 Staff Move: "On SQS my scaling limit is the consumers and whatever they call, so I autoscale on age of oldest message and protect the downstream with a concurrency cap. On RabbitMQ, the limit is the queue process, so a hot queue becomes N sharded queues behind a consistent-hash exchange."
Failure Modes & Recovery#
1. Duplicate-Work Spiral From a Short Visibility Timeout#
- Symptom: Downstream load doubles during a traffic peak; the same job IDs appear in logs two or three times; the queue drains slower as load rises.
- Root cause: Visibility timeout below the p99 job time. Slow jobs reappear and run again, adding load, making more jobs slow.
- Detection:
NumberOfMessagesReceived/NumberOfMessagesDeletedratio rising above ~1.1;ApproximateReceiveCounton processed messages; duplicate hits in the idempotency table. - Fix: Raise the timeout; add heartbeats; cap consumer concurrency so the downstream recovers.
- Prevention: Timeout derived from measured p99 in config review; alarm on receive/delete ratio.
2. Poison Message Loop#
- Symptom: A few messages fail forever; on SQS FIFO or a single-consumer RabbitMQ queue, everything behind them in the group stops.
- Root cause: No DLQ,
maxReceiveCountvery high, or a RabbitMQ classic queue with no delivery limit andrequeue=true. - Detection: Rising
ApproximateAgeOfOldestMessagewith flat throughput; RabbitMQ redeliver rate; the same message ID in error logs repeatedly. - Fix: Move the message to the DLQ manually; deploy the parser fix; redrive.
- Prevention: DLQ on every queue; delivery limit 3–5 for parse-type failures; quorum queues (delivery limit 20 by default since 4.0).
3. RabbitMQ Memory Alarm Blocks Publishers#
- Symptom: Producers across unrelated services hang on publish; the management UI shows a memory or disk alarm.
- Root cause: A consumer outage let one queue grow to millions of messages; the node crossed the 60% watermark.
- Detection:
rabbitmq_alarms_memory_used_watermark; queue depth per queue; node memory by category. - Fix: Restore or scale the consumer; purge or shovel the backlog to another cluster if it's disposable; temporarily raise the watermark only with headroom.
- Prevention:
x-max-lengthwith overflow to a DLQ; alarm at 40% memory; isolate critical producers on their own cluster or vhost.
4. The Backlog That Can't Drain#
- Symptom: After a 2-hour downstream outage, the queue holds 6 hours of work; consumers at full speed only drain 20% faster than arrivals, so recovery takes 30 hours and every new message waits behind old ones.
- Root cause: FIFO processing of a backlog whose oldest messages are now worthless, with no spare capacity.
- Detection: Age of oldest message vs drain rate:
backlog / (drain_rate - arrival_rate)= hours to recover. - Fix: Process fresh traffic first (move the backlog to a separate queue and drain it at low priority), drop messages past their usefulness TTL, scale consumers if the downstream can take it.
- Prevention: Message TTLs per job type; a documented "backlog mode"; capacity headroom of ~2× steady state for recovery.
5. Node Loss or Partition on RabbitMQ#
- Symptom: Some queues unavailable; clients reconnecting; with classic queues on the failed node, messages unavailable or lost.
- Root cause: Node failure or network partition. Quorum queues with a majority elsewhere elect a new leader and continue; a minority side stops accepting writes for those queues.
- Detection: Cluster membership and partition alarms; quorum queue leader changes;
rabbitmq_queue_messages_readydropping to zero on a node. - Fix: Restore the node; let quorum queues catch up from the leader; for classic queues, recover from the producer side.
- Prevention: Quorum queues for anything durable; 3 or 5 nodes across zones; clients with automatic reconnection and publisher confirms.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Duplicate-work spiral | Receive/delete ratio > 1.1 | One queue's downstream | Raise timeout, cap concurrency | Consumer team |
| Poison loop | Age of oldest rising, throughput flat | One queue or one FIFO group | DLQ, fix, redrive | Consumer team |
| Memory alarm | Broker alarm, publish latency | Every publisher on the cluster | Restore consumer, purge or shovel | Platform; consumer team |
| Backlog can't drain | Hours-to-recover estimate | Every message behind the backlog | Fresh-first, TTL drop | Consumer team + product (what's droppable) |
| Node loss / partition | Membership alarms, leader changes | Queues led by that node | Quorum failover, reconnection | Platform |
| DLQ silently expiring | DLQ age vs retention | Lost data | Alarm on DLQ age, 14-day retention | DLQ's owning team |
When to Use vs. Alternatives#
| Need | Pick | Why |
|---|---|---|
| Background jobs on AWS | SQS standard | No ops, per-message retry, scales to zero |
| Ordered per-entity steps on AWS | SQS FIFO with group = entity | Order per group, unlimited groups |
| Fan-out of one event to several consumers, no replay | SNS → SQS or RabbitMQ topic exchange | Each subscriber isolated |
| Off-AWS or multi-cloud queue | RabbitMQ 4 with quorum queues (or a managed RabbitMQ) | Open protocol, runs anywhere |
| Rich routing, priorities, AMQP/MQTT/STOMP clients | RabbitMQ | Broker-side routing and protocol plugins |
| Several teams read the same events; replay | Kafka | Log semantics; see Kafka vs Kinesis |
| Delays longer than 15 minutes or cron-like schedules | A scheduler over a database, enqueuing when due | Queues aren't timers |
When NOT to Use RabbitMQ or SQS#
- Anyone needs to read the messages twice. Replay, a new consumer backfilling history, or reprocessing after a bug all need a log.
- The call needs an answer now. A user waiting on a response wants a synchronous call with a deadline, not a queue with a reply-to.
- Long-delay scheduling. "Send this reminder in 3 days" is a row in a scheduler table, not a message hidden for 3 days.
- Workflow state. A multi-step process with compensation needs a workflow engine or saga state in a database; see Workflows & Sagas. Queues move the steps; they don't remember where you are.
- RabbitMQ without an owner. A self-run broker is a cluster with memory alarms, upgrades and partitions. If nobody owns it, use a managed queue.
Operational Concerns#
What the On-Call Actually Does#
- Watches age of oldest message per queue against each queue's SLO. Depth alone lies.
- Works the DLQs: reads the top error reasons, decides fix-then-redrive vs discard with the owning team, and redrives at a throttled rate.
- Handles backlog recovery: estimates hours to drain, decides with product whether old messages still matter, and switches to fresh-first draining when they don't.
- On RabbitMQ, watches memory and disk alarms, unacked counts and node health, and finds the consumer holding thousands of unacked messages with unlimited prefetch.
- Runs RabbitMQ upgrades with feature flags enabled step by step; the 3.13 → 4.x upgrade required migrating classic mirrored queues to quorum queues first.
- Reviews new queues for the contract: owner, DLQ, timeout, retry policy, idempotency key.
Key Metrics & Alerts#
| Metric | Healthy | Alert |
|---|---|---|
ApproximateAgeOfOldestMessage (SQS) / oldest message age (RabbitMQ) | Within the queue's SLO, e.g. < 60s | > SLO for 5 min (page for critical queues) |
DLQ ApproximateNumberOfMessagesVisible | 0 | > 0 for 15 min (ticket); age > 1 day (page) |
| Receive / delete ratio (SQS) | ~1.0 | > 1.1 sustained |
ApproximateNumberOfMessagesNotVisible (in flight) | Steady with throughput | Near FIFO's 120K limit |
| RabbitMQ memory used / watermark | < 50% | > 70% |
| RabbitMQ unacked messages per consumer | ≈ prefetch | Far above for one consumer (hoarding) |
| RabbitMQ redeliver rate | Low, steady | Spike (poison message or consumer crash loop) |
| Quorum queue leader changes | ~0 | Repeated changes (node instability) |
| Publisher confirm latency p99 | Low ms | > 100 ms (disk or replica trouble) |
Interview Application — Staff-Level Plays#
Which Case Studies Use RabbitMQ & SQS#
| Case Study | How Queues Are Used | Key Pattern |
|---|---|---|
| Message Broker | Designing the queue itself | Lease-based delivery, DLQs, per-tenant fairness |
| Notification System | Per-channel queues (push, email, SMS) behind a fan-out | Separate queues per provider so one slow channel doesn't block others |
| Webhook Delivery | Per-endpoint retry with backoff, DLQ after days | Fairness across customers; retry ladders |
| Distributed Job Scheduler | Due jobs enqueued for workers | Scheduler of record in a database; the queue only moves due work |
| Idempotency | Broker redelivery as a duplicate source | Consumer-side idempotency keys |
| Payment System | Async capture, refunds, reconciliation jobs | Ack after commit; outbox; DLQ owned by payments |
| Video Streaming | Transcoding jobs | Long jobs with heartbeats; autoscale on backlog age |
| Cascading Failures | Queues as shock absorbers that can hide an outage | Backlog age alarms; load shedding |
Every System Design Question Has a Queue Moment#
- Notifications: "Each channel gets its own SQS queue and worker pool. When the SMS provider throttles us, only the SMS queue backs up, and its age alarm pages the notifications team, not everyone."
- Image upload: "The upload API returns 202 and enqueues a thumbnail job with the object key. The job is idempotent on the key, the visibility timeout is 60s with a heartbeat, and after 5 failures it lands in a DLQ the media team owns."
- Payments: "Capture runs from a quorum queue fed by an outbox, so a crash between the database commit and the publish can't lose it, and the consumer passes the payment ID as the provider's idempotency key."
What Interviewers Probe#
| After You Say... | They Will Ask... | What They're Evaluating |
|---|---|---|
| "Put it on a queue" | "What happens when a worker crashes mid-job?" | Ack point, visibility timeout, redelivery, idempotency |
| "Exactly-once with SQS FIFO" | "What's the dedup window? What about a redrive tomorrow?" | Knowing 5 minutes; consumer-side idempotency |
| "We retry failures" | "What about a message that can never succeed?" | DLQ, delivery limits, poison handling, ownership |
| "Add more consumers" | "What if it's FIFO with one hot group? Or one RabbitMQ queue?" | Per-group serialism; one queue = one core |
| "RabbitMQ with mirrored queues" | "Which version?" | Mirroring removed in 4.0; quorum queues |
| "The queue absorbs spikes" | "How do you know the backlog is a problem?" | Age of oldest message vs depth; drain-time math |
Common Interview Mistakes#
| What Candidates Say | What Interviewers Hear | What Staff Engineers Say |
|---|---|---|
| "The queue guarantees exactly-once" | Will double-charge on redelivery | "At-least-once delivery, exactly-once effect via an idempotency key in the consumer's transaction." |
| "Alert when the queue has more than 10K messages" | Measures the wrong thing | "Alert on age of oldest message against the queue's SLO." |
| "Failed messages go to a DLQ" | DLQ as a trash can | "DLQ after 5 receives, owned by the consumer team, age alarm, 14-day retention, practised redrive." |
| "One queue for all background jobs" | Shared blast radius | "One queue per job type, each with its own timeout, scaling and DLQ." |
| "Use FIFO to be safe" | Paying throughput for nothing | "Standard queue plus version checks; FIFO only where per-entity order is a business rule." |
| "RabbitMQ for the event bus that four teams read" | A log problem solved with a queue | "Four readers and replay means Kafka; RabbitMQ for the job queues." |
L5 vs L6 vs L7 Responses#
| Scenario | L5 Answer | L6 / Staff Answer | L7 / Principal Answer |
|---|---|---|---|
| "Send order confirmation emails asynchronously" | SQS queue, Lambda consumer | SQS + DLQ, idempotent on order ID, timeout from p99, age alarm; outbox so the email can't be lost between commit and send | Asks whether the company already has a notifications platform this should use, and who owns email DLQs across products |
| "Workers process some jobs twice" | Use FIFO for exactly-once | Diagnoses visibility timeout vs p99 job time, adds heartbeat and idempotency table | Makes idempotency keys a required field in the org's message envelope and checks it in CI |
| "One customer's import delayed everyone for 6 hours" | Add consumers | Fair queues or shuffle sharding by tenant; fresh-first draining | Sets per-tenant async SLOs and an isolation standard; prices dedicated queues for the largest tenants |
| "Should we run RabbitMQ or use SQS?" | RabbitMQ is more flexible | SQS on AWS unless routing or protocol needs justify a broker | Decides the company's messaging portfolio (one queue, one log) and retires the third system |
The Staff Queue Checklist#
- Queue or log: "Will anyone read these messages twice? If not, it's a queue."
- Ack point: "Ack after the side effect commits; the seen-record is in the same transaction."
- Lease: "Visibility timeout 2× p99, heartbeat for long jobs; prefetch sized to processing time on RabbitMQ."
- Retries: "Backoff via visibility extension or TTL queues, DLQ after 5, owner and age alarm on the DLQ."
- Order scope: "None by default; per entity via message group or consistent hash only where required."
- Health signal: "Age of oldest message per queue against an SLO, plus hours-to-drain during incidents."
🎯 Staff Insight: Don't use a queue as a log (no replay), as a scheduler (15-minute delay cap on SQS), as a workflow engine (it doesn't remember state) or as an RPC transport. The strongest queue signal is describing exactly what happens when a worker dies halfway through a job.
Evaluation Rubric#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Delivery semantics | "Guaranteed delivery" | At-least-once, ack after commit, idempotent consumers, dedup windows understood | Org-wide envelope with idempotency keys, enforced in review |
| Failure handling | Retries | Poison vs transient, DLQ with owner, redrive runbook, backlog drain math | DLQ inventory and age SLOs across teams; game days for backlog recovery |
| Scaling | More consumers | Little's law, prefetch, FIFO group limits, queue sharding | Messaging portfolio and cost model |
| Isolation | One queue | Queue per job type, per-tenant fairness | Tenant isolation standard and dedicated capacity pricing |
| Technology choice | Familiar broker | Queue vs log; SQS vs RabbitMQ with reasons | Number of messaging systems the company should run, and the retirement plan |
Strong hire signals
| Signal | What It Sounds Like |
|---|---|
| Lease thinking | "The visibility timeout is a lease; the job has to finish or renew before it expires." |
| Effect over delivery | "I can't get exactly-once delivery, so I build exactly-once effect." |
| Measures in time | "Age of oldest message, because that's what the user waiting on the job experiences." |
| Owns the DLQ | "The DLQ belongs to the consumer team, with an alarm and a redrive they've practised." |
| Knows when it's a log | "If analytics wants these too, this isn't a queue anymore." |
Lean no-hire signals
| Signal | Why It Misses the Bar |
|---|---|
| Exactly-once claimed from the broker | Will ship duplicate side effects |
| No DLQ or unowned DLQ | Poison messages loop or data silently expires |
| Depth-based alerting only | Misses slow drains and catches harmless spikes |
| FIFO everywhere "for safety" | Pays throughput and head-of-line blocking for no requirement |
Common false positives
- AMQP vocabulary ≠ queue design. Knowing every exchange type says nothing about where the ack goes.
- "We ran RabbitMQ at scale" ≠ judgment. Ask how they recovered from a memory alarm and what they'd use on AWS today.
- Kafka-first answers ≠ sophistication. A job queue built on Kafka without retry topics or DLQ ownership is worse than SQS.
Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
At Staff level a queue is a component with the right timeout and a DLQ. At Principal level, the async tier is where failures go to hide: a synchronous outage pages someone in a minute, but a queue backing up degrades quietly for hours, and the DLQs of 60 teams are a distributed archive of business events nobody is reading. The L7 question is not "SQS or RabbitMQ?" but "How many messaging systems should we run, what contract does every queue have to meet, and how do we know, across the whole company, that async work is actually getting done?"
🧭 Principal Move: "Before choosing a broker I want the async contract: every message carries an idempotency key and an owner, every queue has an age SLO, every DLQ has an alarm and a redrive tool. Once that's standard, the broker is a detail; without it, the best broker in the world still loses orders in an unowned DLQ."
The Org-Level Fault Line#
One messaging platform vs every team picks its own.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Every team picks (SQS, RabbitMQ, Redis lists, Kafka) | Autonomy, fast start | 4 systems, 4 sets of alarms, inconsistent DLQ handling, no company-wide view of backlogs | Incident responders, who learn each system during the outage |
| One queue + one log, platform-owned (e.g. SQS + Kafka) | Uniform tooling, DLQ dashboards, shared runbooks | Edge cases (routing, priorities) need workarounds | Platform team headcount |
| One system for everything (Kafka for jobs too) | One thing to run | Job semantics rebuilt by hand per team: retry topics, DLQs, delays | Every team writing retry plumbing |
| Self-run RabbitMQ per team | Full control | Dozens of small clusters, each one upgrade behind | Teams without broker expertise |
The Principal default: one managed queue and one log, both platform-owned, with a shared message envelope (idempotency key, owner, schema version, trace context), a DLQ dashboard across every queue, and an exceptions process for teams that genuinely need broker routing.
🧭 Principal Insight: "The expensive part of messaging isn't the broker bill. It's the third messaging system, because each one brings its own failure modes, its own on-call knowledge and its own way of losing messages."
Cost Model#
Assumptions: SQS at an illustrative ~$0.40 per million requests (check current pricing and volume tiers), 3 requests per message (send, receive, delete), "batched" meaning batches of 10; a RabbitMQ node ~$400/month; loaded engineer $250K/year ($21K/month). Directional only.
| Scale | Volume | SQS/month (batched – unbatched) | Self-run RabbitMQ/month | People | Note |
|---|---|---|---|---|---|
| Startup | 5M msg/day | ~$20 – $180 | Near zero on SQS | SQS wins by two orders of magnitude | |
| Growth | 200M msg/day | ~$0.7K – $7K | 0.25 FTE either way for DLQ tooling | SQS still cheaper once people are counted, if batched | |
| Enterprise | 5B msg/day | ~$18K – $180K | Platform team for the envelope and dashboards | Batched SQS and a broker are comparable; unbatched SQS is the most expensive option on the page |
The lever at every scale is batching and long-polling: unbatched, empty-receive-heavy consumers make SQS up to ~10× more expensive than it needs to be. The lever nobody prices is the cost of a silent backlog: six hours of delayed order confirmations is a support-ticket spike and a churn number.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse |
|---|---|---|
| Queue vs log for a business event stream | One-way once consumers rely on replay or its absence | Re-plumbing every producer and consumer; no history to backfill |
| Message envelope (idempotency key, schema version) | One-way-ish | Every producer and consumer updated in lockstep |
| SQS FIFO vs standard | Two-way-ish | New queue and a cutover; FIFO names must end in .fifo |
| RabbitMQ vs SQS | Two-way, if consumers are behind an interface | Weeks per service; routing logic is the drag |
| Timeouts, retry counts, prefetch | Two-way | Config change |
The Standard I'd Write#
RFC-ASYNC-002: Queue Contract
Scope: Every queue (SQS, RabbitMQ) carrying production work.
MUST
1. Declare an owning team, routed for pages, on the queue and its DLQ.
2. Have a DLQ with retention >= 14 days and an alarm on age of oldest message.
3. Carry an idempotency key in the message envelope; consumers dedupe on it
inside the side effect's transaction or via the downstream's idempotency key.
4. Ack/delete only after the side effect is durable.
5. Declare an age SLO; alarm when the oldest message exceeds it.
6. Use quorum queues (RabbitMQ) for any message whose loss is a business loss.
SHOULD
7. Size visibility timeout >= 2x measured p99 processing; heartbeat long jobs.
8. Use one queue per job type; separate queues per priority.
9. Use SQS fair queues or sharding for multi-tenant queues.
Exceptions: platform review; time-boxed to one quarter.
Rollout: inventory -> warn in CI -> enforce for new queues -> backfill existing.
Success metrics: zero unowned DLQs; zero DLQ messages expired unread;
every queue reporting age SLO; duplicate side effects = 0 in audits.
What I'd Tell the VP#
"A lot of our customer-facing work happens in the background: emails, refunds, exports. When that machinery backs up, nothing turns red; customers just wait, sometimes for hours. Today each team runs it differently, and failed work sits in places nobody watches until it's deleted. I'm proposing one standard for all background work, a single dashboard showing how far behind each part is, and named owners for every failure bin. It's mostly configuration and a small platform investment, roughly one engineer for a quarter, and it turns silent delays into alerts we act on within minutes."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Sees the async tier as a risk surface | "Queues hide outages; my first dashboard is age of oldest message across every queue." |
| Limits the portfolio | "One queue and one log. A third messaging system needs a business case." |
| Standardises the envelope | "Idempotency key, owner, schema version and trace context on every message." |
| Prices the silent backlog | "Six hours of delayed confirmations is a ticket spike we can put a number on." |
| Designs tenant isolation | "Fair queues by tenant, dedicated queues for the top 1% by volume." |
Staff answers that L7 interviewers find insufficient:
- "Every queue has a DLQ" without who reads it, the alarm, or what happens on day 15.
- "Make consumers idempotent" as advice, rather than a required envelope field enforced in CI.
- "Use SQS because it's managed" without the portfolio question: what else does the company run, and what gets retired.
How Real Companies Built It#
Trello — Leaving RabbitMQ for WebSocket Fan-Out#
Trello's engineering team described running websocket updates through a cluster of 15 RabbitMQ instances for about three years, with each websocket process creating a transient queue and binding it to a topic exchange per subscribed model. The problems were partition handling ("from split-brain to complete cluster failure") and binding churn: after mass socket disconnects, a flood of binding add and remove commands made the cluster unresponsive, even to monitoring. They recorded 4 RabbitMQ-caused outages in the month before switching, evaluated SNS + SQS, Kinesis and Redis Streams, and moved to Kafka, reporting lower memory use and a 5× cost reduction (Atlassian engineering).
Staff insight: This was a fan-out problem with per-connection queues, which is a log-shaped workload on a queue-shaped broker. It also predates quorum queues becoming the default. In an interview, the lesson is to name binding churn and partition behaviour before proposing a queue per session.
Airbnb — Dynein, Delayed Jobs on SQS#
Airbnb built Dynein, an open-source distributed delayed-job system, on SQS for the job queues and DynamoDB for the schedule of jobs due in the future: a scheduler queries for overdue jobs and dispatches them to service queues consumed by workers (Dynein on GitHub). The team explained choosing SQS because it is simple to reason about and gives at-least-once delivery, dead-letter queues and individual message acknowledgement out of the box (InfoQ coverage).
Staff insight: The queue moves due work; the database remembers what is scheduled. That split is exactly the answer to "can SQS schedule a job for next week?" (no: 15 minutes maximum delay).
Amazon — Avoiding Insurmountable Queue Backlogs#
An AWS Builders' Library article by a senior principal engineer describes how Amazon services keep queue backlogs from becoming multi-hour outages: shuffle sharding customers across a few queues (used in AWS Lambda's asynchronous invocation path), separate queues per customer in some systems, spillover queues for traffic over a customer's rate, working a backlog queue only after catching up on live traffic, dropping messages past a TTL, heartbeating long jobs, and throttling inbound traffic in proportion to backlog size (AWS Builders' Library).
Staff insight: The hard queue problem isn't delivery, it's recovery. When asked "what happens after the outage ends?", the Staff answer is fresh-first draining, per-tenant isolation and TTLs on work that has gone stale.
Practice Drill#
Drill 1: The Six-Hour Backlog#
Prompt: "Your multi-tenant SaaS sends report exports, emails and webhooks through one SQS standard queue consumed by 40 workers. Last week one customer's bulk import enqueued 4 million jobs; every other customer's emails were delayed for six hours, and support found some exports had been generated twice. Redesign the async tier."
Staff Answer
Two separate failures. The delay is an isolation failure: one queue for three job types and every tenant means one burst blocks everyone, and nothing alerted because the alarm was on depth, which is always high during imports. The duplicates are a lease failure: exports take up to 90 seconds at p99 and the visibility timeout is the 30-second default, so slow exports reappear and run again. Fixes in order. First, split by job type: emails, webhooks, exports, each with its own worker pool, timeout and DLQ, so a slow export can never delay a password reset. Second, tenant fairness: set MessageGroupId = tenant_id on the standard queues to get SQS fair queues, so a tenant with a disproportionate share in flight is deprioritised while quiet tenants keep low dwell time; for the top handful of tenants by volume, a dedicated bulk queue drained at a capped rate. Third, the lease: export visibility timeout at 180s (2× p99) with a heartbeat every 60s for jobs that run long, and an idempotency record keyed by export_id written before the file is published, so a redelivery returns the existing file. Fourth, signals: alarm on ApproximateAgeOfOldestMessage per queue against an SLO (emails 60s, webhooks 5 min, exports 15 min), alarm on receive/delete ratio above 1.1, DLQ after 5 receives with the owning team paged when its oldest message passes an hour. Fifth, recovery: a runbook to move a backlog to a low-priority queue so fresh work flows first, and TTLs on emails that are useless after a day. Autoscale workers on backlog age, capped by what the email and webhook providers will accept.
Why this is L6:
- Separates the two symptoms into two root causes: isolation and lease length.
- Uses age of oldest message per job type as the health signal, with SLOs per queue.
- Fixes duplicates with both a correct timeout and consumer idempotency, not FIFO.
- Plans the recovery path (fresh-first, TTLs), not just prevention.
What L7 adds:
- Turns it into a queue contract for every team: owner, DLQ alarm, age SLO, idempotency key in the envelope.
- Prices tenant isolation: dedicated bulk capacity as a paid tier for the largest customers.
- Adds a company-wide async dashboard so the next silent backlog is visible to leadership, not discovered by support.
❌ Common L5 Trap
"Switch to SQS FIFO for exactly-once processing so exports aren't duplicated, and add more workers so the backlog drains faster."
Why this misses: FIFO's dedup covers producer retries within 5 minutes, not a slow consumer whose visibility timeout expired, so the duplicates continue. FIFO also serialises each message group and caps throughput, which makes the backlog worse. More workers on one shared queue still process the 4 million import jobs before other tenants' emails, because the problem is isolation, not capacity. The fix is separate queues, tenant fairness, a correct lease and idempotent consumers.
Quick Reference Card#
Model: queue = broker tracks each message; acked = gone (no replay)
log = consumers track offsets; data kept (Kafka, RabbitMQ streams)
Delivery: at-least-once everywhere; exactly-once EFFECT via consumer idempotency
SQS standard: nearly unlimited throughput; best-effort order; duplicates possible
SQS FIFO: order per MessageGroupId, 1 in flight per group; dedup window 5 min;
300 calls/s per partition (3,000 msg/s batched); up to 70K TPS high-throughput;
120,000 in flight max
SQS limits: visibility 30s default, 12h max; retention 4d default, 14d max;
delay <= 15 min; long poll <= 20s; batch 10; message 1 MiB
SQS fair queues: standard queues + MessageGroupId as tenant ID -> noisy-neighbour protection
RabbitMQ: exchange (direct/topic/fanout/headers) -> bindings -> queues
Quorum queues: Raft, 3 replicas default, odd count, <= 7 nodes; delivery limit 20 (4.0+)
~1 MB RAM per 30K messages of metadata
Classic mirror: removed in RabbitMQ 4.0
RabbitMQ limits: consumer ack timeout 30 min; memory watermark 60% -> publishers blocked;
max message 16 MiB; one queue ~ one core
Sizing: consumers = arrival rate x processing time (Little's law), run at ~70%
visibility timeout = 2 x p99; heartbeat long jobs
Health: age of oldest message per queue vs SLO; receive/delete ratio; DLQ age
RED FLAGS
- Ack/delete before the side effect commits
- Visibility timeout below p99 processing time
- DLQ with no owner, no alarm, default 4-day retention
- One queue for every job type and tenant
- "Exactly-once" credited to the broker
- Alerting on depth instead of age
- Classic (unreplicated) RabbitMQ queues for business-critical messages
- A queue where several teams need replay (that's a log)