Hiring BarSupport

RabbitMQ & SQS

Technology guide44 min read6 diagrams

Why This Matters#

A work queue is not a messaging pipe. It is a contract about who owns a unit of work while it is being done. The broker hands a message to one consumer, stops offering it to anyone else for a while, and takes it back if the consumer goes quiet. Every interesting property (duplicates, ordering, retries, poison messages, backlog) comes from how that hand-off is timed and who is allowed to give the message back. RabbitMQ and SQS are the two queues interviewers meet most often: RabbitMQ is the broker you run, with routing built in; SQS is the queue you rent, with almost nothing to operate.

That is why "drop it on a queue" is a sentence interviewers push on. The L5 candidate draws a queue between the API and the workers and says "this decouples them." The L6 candidate says "SQS standard queue, visibility timeout of 2× the p99 job time with a heartbeat for long jobs, maxReceiveCount of 5 into a dead-letter queue that the payments team owns and alerts on, idempotent consumers keyed by job ID, and an alarm on the age of the oldest message, not queue depth." The L7 candidate asks how many queueing technologies the company should run, who owns the dead-letter queues across 60 teams, and what it costs when the async path silently backs up for six hours and nobody's dashboard turns red.

The L5 → L6 gap is not knowing what an exchange is. It is knowing that a queue converts a latency problem into a backlog problem, and a backlog nobody measures in minutes is an outage that hasn't been noticed yet.

The L5 → L6 → L7 Contrast#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Put a queue between the API and the workers""Is this a job done once, or an event several teams will read? A job is a queue; an event is a log. Then: SQS if we're on AWS, RabbitMQ only if we need broker routing or we're off-cloud.""How many messaging systems does the company run, and does this team get to add one? Each broker is a platform with an owner and an on-call rota."
Delivery"The queue guarantees delivery"At-least-once everywhere; ack only after the side effect commits; consumers idempotent on a business key; FIFO dedup covers 5 minutes, not foreverPublishes the org's delivery contract: every consumer idempotent, every message carries an idempotency key, audited in review
Retries"Retry failed messages"Visibility timeout or delivery limit bounds retries; backoff via delay or TTL queues; DLQ after N attempts with a named owner and a redrive runbookTreats DLQs as an org-wide inventory: count, age and owner per DLQ on one dashboard; unowned DLQs are a review failure
Scaling"Add more consumers"Little's law for consumer count; prefetch tuned to processing time; knows a RabbitMQ queue lives on one core; FIFO throughput per group is 1 in flightPrices the async tier: SQS per-request cost vs broker nodes 24/7, and the headcount to run RabbitMQ
Failure"Messages are persisted, so nothing is lost"Names the backlog that can't drain, the visibility timeout shorter than the job, the memory alarm that blocks publishers, the poison message loopDesigns per-tenant isolation (fair queues, shuffle sharding) so one customer's burst doesn't become every customer's six-hour delay
Ownership"The platform team runs the queue"Producer owns the schema; consumer owns the DLQ and the redrive; platform owns the brokerDecides what is centralised (brokers, alarms, DLQ tooling) and what each team must declare (owner, retry policy, max age SLO)
Why "Delivery" separates levels

"The queue guarantees delivery" is true and nearly useless. Both systems are at-least-once: SQS standard can deliver a message more than once, and any consumer that crashes after doing the work but before deleting the message will see it again when the visibility timeout expires. RabbitMQ redelivers every unacknowledged message when a channel closes. So the real guarantee is "every message is processed at least once, and you will see duplicates." The Staff answer moves the guarantee to where it can actually be kept: the consumer acks only after the side effect is durable, and the side effect is idempotent on a business key (charge_id, email_id) stored in the same transaction. SQS FIFO's deduplication helps with producer retries, but its window is 5 minutes, so it does not protect a consumer that is replayed tomorrow. The Principal answer makes idempotency a standard with a linter, because one non-idempotent consumer in 60 is how a customer gets charged twice.

The 60-Second Pitch#

"For background jobs on AWS I'd use SQS standard: no brokers, effectively unlimited throughput, and per-message retries for free. Each job type gets its own queue and its own dead-letter queue. The visibility timeout is about twice the p99 job time, and jobs that can run long extend it with a heartbeat. After 5 failed receives a message moves to the DLQ, which pages the owning team when its oldest message is older than 15 minutes. Consumers are idempotent on the job ID, because SQS is at-least-once. I'd use FIFO only where per-entity order matters, keyed by entity ID, and accept 300 calls per second per partition unless we turn on high-throughput mode. If we weren't on AWS, or needed topic routing to many queues, I'd run RabbitMQ 4 with quorum queues on 3 nodes, publisher confirms and manual acks. If anyone needs to replay these messages, it isn't a queue, it's a log, and I'd use Kafka instead."

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
Background job distributionEach job done once; spiky load; per-job retrySQS standard (or RabbitMQ quorum queue) + DLQ + idempotent workers; autoscale on age of oldest messageVisibility timeout shorter than the job → duplicate work; poison message retried foreverEvery job completes or lands in a DLQ with an owner; duplicates harmless
Routed fan-out to many consumersOne event, several queues, routing by type or attributeSNS → SQS with filter policies, or RabbitMQ topic/headers exchange → per-consumer queuesBinding churn or a forgotten queue fills the broker; a slow subscriber backs up only itselfEach subscriber gets its own copy and its own retry policy
Ordered per-entity workflowsSteps for one order must run in orderSQS FIFO with MessageGroupId = order_id, or RabbitMQ consistent-hash exchange to N single-consumer queuesOne slow or poisoned message blocks its whole group; one hot group caps throughputStrict order within an entity; no cross-entity head-of-line blocking

🎯 Staff Move: "I'll design for the first intent: background jobs, one queue per job type. I'm assuming nobody needs to re-read these messages; if analytics wants them, that's a separate event stream, not a second consumer on this queue. Ordering I'll handle per entity only where the business needs it, because FIFO costs throughput and head-of-line blocking everywhere else."

The Staff Positions#

PositionRationale
Ack after the side effect, never beforeAck-on-receive is at-most-once; a crash between receive and commit loses the job silently.
Every consumer is idempotent on a business keyBoth systems are at-least-once; duplicates are a certainty, not an edge case.
Alert on age of oldest message, not depth50,000 messages is fine at 10,000/s drain and a disaster at 10/s. Age is what the user feels.
Every DLQ has an owner, an alarm and a redrive runbookA DLQ without an owner is a place where data goes to be forgotten.
Quorum queues for anything that matters on RabbitMQClassic queue mirroring was removed in RabbitMQ 4.0; quorum queues are the replicated, Raft-based default.
One queue per job type, not one queue for everythingIsolates retry policy, timeout, scaling and blast radius per workload.
If anyone needs replay, it's a log, not a queueAcked messages are gone. A second reader next quarter needs Kafka, not a second queue.

Architecture & Internals#

Only five internals change design decisions: the queue vs log model, RabbitMQ's exchanges and bindings, acknowledgements and prefetch, SQS's visibility timeout, and how each system replicates. The side-by-side with Kafka is in Kafka vs SQS vs RabbitMQ.

Queues vs Logs: Who Remembers What Was Read#

In a queue the broker tracks each message's state: ready, in flight, acknowledged (deleted). In a log like Kafka the broker tracks nothing per message; consumers keep an offset and data lives until retention removes it. That single difference decides replay, retry and parallelism. A second reader on a queue needs its own queue and a fan-out in front; on a log it just starts its own offset. Retrying one message is native on a queue and awkward on a log. Parallelism on a standard queue is "add consumers"; on a log it is bounded by partitions.

RabbitMQ also offers streams, an append-only log type with offset-based reads and replay, for when a RabbitMQ shop needs log semantics without adopting Kafka. Treat streams as "Kafka-shaped, inside RabbitMQ", and size them like a log.

RabbitMQ: Exchanges, Bindings and Queues#

Producers publish to an exchange (even "publishing to a queue" goes through the default exchange) with a routing key; bindings decide which queues receive a copy. Consumers read from queues.

Diagram: RabbitMQ: Exchanges, Bindings and Queues
Exchange typeRoutes byUse for
directExact routing key matchWork queues by job type
topicPattern on dot-separated key (* one word, # zero or more)Event fan-out by type, region, tenant
fanoutIgnores the key, copies to every bound queueBroadcast (cache invalidation, config push)
headersMessage header valuesRouting on several attributes at once
consistent hash (plugin)Hash of the routing key over bound queuesPer-entity ordering across N queues

Why it matters in design: routing lives in the broker, so adding a subscriber is a new binding, not a producer change. The cost is that bindings are broker state. Trello's account of its RabbitMQ era (below) describes a flood of binding add and remove commands after mass socket disconnects making the cluster unresponsive, even to monitoring. Bindings are cheap to create one at a time and expensive to churn by the thousand.

Acknowledgements, Prefetch and Confirms#

Three settings decide RabbitMQ's delivery guarantee, and candidates usually mention only one:

SettingWhat it doesStaff default
Publisher confirmsBroker acks the publish once the message is safely enqueued (for quorum queues, after a majority of replicas have it)On, for anything you can't lose; producer retries on nack or timeout, which can create duplicates
Consumer manual ackMessage stays unacked until the consumer acks; redelivered if the channel closesAlways manual for work queues; ack after the side effect commits
Prefetch (basic.qos)Caps unacked messages per consumerSet from processing time: tens for fast jobs, 1–5 for slow ones

RabbitMQ also enforces a delivery acknowledgement timeout: if a consumer holds a delivery unacked for longer than 30 minutes by default, its channel is closed with PRECONDITION_FAILED and the messages are redelivered (consumer docs). A job that legitimately runs for an hour needs a different design (ack, then track progress elsewhere), not a raised timeout.

Prefetch sizing (a common production pattern): a consumer needs enough messages buffered to cover the round trip between finishing one and receiving the next. With a 2ms network round trip and 20ms jobs, a prefetch of ~2–5 per consumer thread keeps it busy. With 200ms jobs, a prefetch of 1–2 avoids one consumer hoarding 100 messages that others could be working on. Unlimited prefetch on a slow consumer is how one stuck pod holds 10,000 messages hostage.

SQS: The Visibility Timeout Is the Whole Protocol#

SQS has no acknowledgement in the AMQP sense. ReceiveMessage hides a message for the visibility timeout; DeleteMessage with the receipt handle removes it; if neither happens in time, the message becomes visible again and its receive count goes up.

Diagram: SQS: The Visibility Timeout Is the Whole Protocol
SQS settingDefaultLimitDesign note
Visibility timeout30s0s – 12 hours≥ 2× p99 processing time, or heartbeat with ChangeMessageVisibility
Retention4 days60s – 14 daysSet the DLQ's retention longer than the source queue's
Delivery delay0s15 minutesLonger delays need a scheduler, not a queue
Long-poll wait0s20sAlways use 20s; empty receives still cost a request
Batch size—10 messages per send/receive/deleteBatching cuts request cost up to ~10×
Message size—1 MiBLarger payloads go to S3 with a pointer

Limits from the SQS message quotas (as of 2026).

SQS FIFO adds two things. Message groups: messages with the same MessageGroupId are delivered in order, one in flight at a time; different groups proceed in parallel, and there is no limit on group count. Deduplication: a MessageDeduplicationId (or a content hash) suppresses duplicate sends within a 5-minute window (dedup docs). Throughput is 300 API calls per second per partition, 3,000 messages per second with batches of 10, and high-throughput mode raises it to up to 70,000 transactions per second in the largest regions (SQS quotas, as of 2026).

SQS fair queues (standard queues only) use MessageGroupId as a tenant identifier: when one tenant has a disproportionate share of in-flight messages, SQS prioritises other tenants' messages to protect their dwell time, with no ordering and no throughput cap (fair queues).

🎯 Staff Insight: "In SQS the visibility timeout is the lock, the lease and the retry timer at once. Set it shorter than the job and two workers do the same work. Set it ten times longer and a crashed worker's message sits invisible for ten times as long. So I size it from p99 job time and heartbeat for the long tail."

Replication and Durability#

SQSRabbitMQ quorum queueRabbitMQ classic queue
ReplicationManaged, stored redundantly across multiple AZsRaft over an odd number of members, 3 by defaultNone in 4.x (mirroring removed in 4.0)
SurvivesAZ loss, transparentlyLoss of a minority (1 of 3, 2 of 5)Nothing; node loss = queue unavailable or lost
Publish ack meansMessage storedMajority of replicas have it (with confirms)Leader has it
CostPer request~3× disk and network; fsync on every batchCheapest; for transient data only

RabbitMQ's docs recommend an odd quorum-queue group size, note that performance degrades beyond 5 members, and advise against running quorum queues on more than 7 nodes. Each queue keeps at least 32 bytes of metadata per message in memory, roughly 1 MB per 30,000 messages regardless of size (quorum queue docs), so a 30-million-message backlog costs ~1 GB of RAM per replica just in bookkeeping.


Core Usage — "The Entire Game": The Message Contract#

In Kafka the game is the partition key. In a work queue it is the message contract: the five decisions that determine whether a job runs once, runs twice harmlessly, or runs twice and charges a customer.

The message contract (per queue):
  1. ack point     -> when is the message considered done?
  2. lease length  -> visibility timeout / prefetch / consumer timeout
  3. retry policy  -> how many attempts, what backoff, where it goes after
  4. order scope   -> none, per entity (group / hash), or global (almost never)
  5. dedup key     -> what makes a second delivery harmless

Step 1: The Ack Point — After the Side Effect#

WRONG (at-most-once):                 RIGHT (at-least-once + idempotent):
  msg = receive()                       msg = receive()
  delete(msg)        # ack first        begin tx
  send_email(msg)    # crash = lost       if seen(msg.job_id): commit; delete(msg); return
                                          do_side_effect(msg)
                                          mark_seen(msg.job_id)
                                        commit
                                        delete(msg)   # crash before this = harmless redelivery

The idempotency record and the side effect commit in one transaction where possible. When the side effect is an external call (a payment provider, an email API), pass the job ID as that system's idempotency key. The full treatment of keys, scopes and the crash window is in Idempotency.

Step 2: Size the Lease and the Consumer Fleet#

Example: thumbnail jobs
  arrival rate      lambda = 400 jobs/s at peak
  processing time   p50 = 0.8s, p99 = 4s
  concurrency/pod   8 threads

  consumers needed (Little's law): L = lambda x W = 400 x 0.8 = 320 jobs in flight
  pods = 320 / 8 = 40 pods at 100% busy -> run 55-60 for ~70% utilisation
  visibility timeout = 2 x p99 = 8s -> round to 30s (allow GC pauses, retries inside the job)
  heartbeat: jobs over 20s call ChangeMessageVisibility(+30s) every 15s
  DLQ: maxReceiveCount = 5

Utilisation matters more than the average suggests: queueing delay grows sharply above ~70–80% busy, which is why the fleet runs at 55–60 pods, not 40. The math is in Queueing Theory, and the Back-of-Envelope calculator does the arithmetic.

Step 3: Retry Policy and Dead Letters#

Failure typeExampleRight handling
TransientDownstream 503, timeout, throttlingRetry with backoff; let the visibility timeout or a delay queue space attempts
PoisonMalformed payload, missing record, code bugStop retrying fast: DLQ after a small N (3–5), alert, fix, redrive
Business rejectionCard declined, user deletedNot an error: ack it and record the outcome
SQS:      RedrivePolicy { maxReceiveCount: 5, deadLetterTargetArn: thumbs-dlq }
          backoff: on failure, ChangeMessageVisibility(min(2^attempt x 10s, 15min))
          redrive: DLQ redrive back to the source once the fix ships
RabbitMQ: x-queue-type=quorum, x-delivery-limit=5, x-dead-letter-exchange=thumbs.dlx
          backoff: nack into a wait queue with per-message TTL that dead-letters back
          (quorum queues default the delivery limit to 20 since 4.0)

The DLQ is a product with an owner. It needs an alarm on age of oldest message (not on count, which a single bad deploy can push to 50,000 in a minute), a dashboard showing the top error reasons, and a redrive tool the on-call has practised.

Step 4: Ordering Scope — Usually None, Sometimes Per Entity, Never Global#

Global order on a queue means one consumer, one message at a time. Nobody actually needs that. What businesses need is per-entity order: the order's created before its paid before its shipped.

Diagram: Step 4: Ordering Scope — Usually None, Sometimes Per Entity, Never Global

Two rules: ordering costs parallelism (one in flight per group or per shard), and retries break order unless the whole group blocks. SQS FIFO blocks the group while a message is in flight or being retried, which is correct and also means one poison message stops that entity until it reaches the DLQ. Often the cheaper design is to make consumers order-tolerant: carry a version number and ignore stale updates.

Step 5: The Dedup Key#

LayerMechanismWindow
Producer → SQS FIFOMessageDeduplicationId5 minutes
Producer → RabbitMQPublisher confirms + producer retryNone (retries create duplicates)
ConsumerIdempotency table keyed by job ID, TTL ≥ max retention + redrive timeAs long as you keep it

Only the consumer-side key covers every source of duplicates: producer retries, visibility timeout expiry, channel closes, DLQ redrive and replays. Broker dedup is a cost optimisation, not a correctness mechanism.

🎯 Staff Move: "I'll make the consumer idempotent on job ID, with the seen-record written in the same transaction as the side effect. FIFO dedup is nice for producer retries, but it's a 5-minute window and redrive from a DLQ happens days later, so it can't be the thing that keeps us correct."


The Tunable Tradeoff — Safety vs Throughput vs Latency#

Every queue setting moves along one axis: how much work does the broker do to be sure about a message?

SettingFast endSafe endWho pays at the fast end
RabbitMQ queue typeClassic, transient messagesQuorum queue, persistent, 3 replicasWhoever loses the node, as lost messages
Publisher confirmsOff (fire and forget)On, with retry on nackProducer team, in silent loss during broker trouble
Consumer ackAuto-ack on deliveryManual ack after commitUsers whose job vanished when a pod crashed
PrefetchUnlimited1Other consumers, starved by a hoarder; at 1, throughput
SQS queue typeStandardFIFOConsumers, who must tolerate duplicates and reordering
Visibility timeoutShort (fast retry)Long (no duplicate work)Downstream systems, doing the same job twice
Retry countMany (eventual success)Few, then DLQThe queue, clogged by poison messages

🎯 Staff Move: "Quorum queues with confirms and manual acks for orders and payments; classic transient queues are fine for cache-invalidation broadcasts where a lost message costs one stale read. Safety where a lost message is money, speed where it's a refresh."


Anti-Patterns — What Kills RabbitMQ and SQS Deployments#

1. Ack on Receive#

Auto-ack (RabbitMQ) or delete-before-process (SQS) makes the queue at-most-once. Every pod crash, OOM kill or deploy drops whatever was in flight, and nothing alerts because nothing failed visibly. Fix: manual ack after the side effect commits; idempotent consumers to absorb the duplicates that follow.

2. Visibility Timeout Shorter Than the Job#

A 30-second default on a job whose p99 is 45 seconds means the slowest 1% run twice, then the duplicates slow the downstream further, pushing more jobs past 30 seconds. Duplicate rate climbs with load. Fix: size from p99 × 2, heartbeat with ChangeMessageVisibility for long jobs, and watch the ratio of receives to deletes.

3. The Unowned Dead-Letter Queue#

Messages land in a DLQ with a 4-day retention, nobody alerts on it, and on day 5 they are deleted. The data is gone and no incident was ever opened. Fix: every DLQ has an owning team, an age-of-oldest alarm, retention set to the 14-day maximum, and a tested redrive path.

4. One Queue for Every Job Type#

Password-reset emails wait behind a 2-million-message backfill, because both are "jobs." Fix: one queue per job type (or per priority tier), each with its own consumer fleet, timeout and DLQ. Separate queues are cheap on both systems.

5. Unbounded Backlogs on RabbitMQ#

RabbitMQ holds messages on the broker. A consumer outage lets queues grow until memory crosses the high watermark (60% of RAM by default) or disk falls below the free-space limit, at which point the broker blocks all publishing connections (configuration docs). An unrelated producer on the same cluster now times out. Fix: queue length limits (x-max-length with reject-publish or dead-lettering), alarms well before the watermark, and separate vhosts or clusters for workloads that must not block each other.

6. Queue-Per-Session or Binding Churn#

A queue and binding per user session or WebSocket feels natural with topic exchanges. At 100,000 sessions, a mass reconnect is 100,000 queue declarations and binding changes on a broker whose metadata operations are expensive. Fix: a small, fixed set of queues per consumer process with in-process filtering, or a log for fan-out at that scale.

7. Using a Queue When You Needed a Log#

The team ships an SQS queue for order events; next quarter analytics, search indexing and fraud all want the same events, plus a 3-day replay after a bug. A queue can't give any of that. Fix: ask "will anyone read this twice?" before choosing. SNS → SQS fan-out handles new readers going forward, but not history.


The Technology Landscape — Head-to-Head Comparison#

DimensionSQS standardSQS FIFORabbitMQ (quorum queues)RabbitMQ streamsKafka
ModelManaged queue, visibility timeoutManaged queue, message groupsBroker with exchanges, Raft-replicated queuesAppend-only log inside RabbitMQReplicated partitioned log
OrderingBest effortPer message groupPer queue, single consumer, no redeliveryPer streamPer partition
DeliveryAt-least-onceExactly-once processing within 5-min dedup windowAt-least-once with confirms + manual acksAt-least-once, offset-basedAt-least-once; transactions Kafka→Kafka
ThroughputNearly unlimited300 calls/s per partition; up to 70K TPS high-throughput (largest regions)~Tens of thousands msg/s per queue; shard queues to scaleHigher than queues; log-shaped100s of MB/s per cluster
ReplayNoNoNoYesYes
RoutingVia SNS filter policiesVia SNS FIFONative: direct, topic, fanout, headersVia exchangesTopic per stream; routing in consumers
Ops burdenNoneNoneMedium: nodes, upgrades, memory alarmsMediumHigh self-hosted, medium managed
Cost shapePer request; scales to zeroPer request, slightly higherNodes 24/7Nodes + diskBrokers + disk + cross-AZ traffic
Pick whenBackground jobs on AWSPer-entity ordered workflows on AWSOff-AWS queues, broker routing, prioritiesRabbitMQ shop needing replayMulti-consumer event streams, replay

🎯 Staff Insight: "The cheapest broker to operate is the one I don't operate. On AWS, SQS wins for job queues unless I need routing or protocol features it lacks. RabbitMQ earns its nodes when I'm off-cloud, need topic routing, or need AMQP for an existing ecosystem."


Patterns#

Pattern 1: Work Queue with DLQ and Redrive#

The default for 80% of async work: producer → queue → autoscaled workers → DLQ after N receives. Autoscale on age of oldest message (ApproximateAgeOfOldestMessage) or backlog per worker, not CPU. Redrive from the DLQ after the fix ships, throttled so the redrive itself doesn't overload the downstream.

Pattern 2: Fan-Out — SNS to SQS, or a Topic Exchange#

Diagram: Pattern 2: Fan-Out — SNS to SQS, or a Topic Exchange

Each subscriber gets its own queue, timeout, retry policy, DLQ and scaling. A slow email provider backs up only email-q. RabbitMQ does the same with a topic exchange and one queue per subscriber. This is the queue world's answer to multiple consumer groups, minus replay.

Pattern 3: Per-Tenant Fairness#

One tenant's import of 5 million rows should not delay every other tenant's jobs for six hours. Options, cheapest first: SQS fair queues with MessageGroupId = tenant_id; shuffle sharding (hash each tenant to 2 of N queues, consumers drain the shorter); spillover queues for tenants over a rate; a queue per large tenant. Amazon's write-up on queue backlogs (below) describes all of these. The broader design is in Multi-Tenancy.

Pattern 4: Outbox to Queue#

Writing to the database and sending to a queue are two systems with no shared transaction. Write the message to an outbox table in the same transaction as the business change; a relay publishes it and marks it sent. Duplicates are possible, losses aren't. See Transactional Outbox.


Scaling#

The Numbers#

ResourceDocumented limit or typical figureDesign note
SQS standard throughputNearly unlimited API calls per secondScale consumers freely
SQS FIFO throughput300 calls/s per partition; 3,000 msg/s with batches of 10; up to 70K TPS (700K msg/s batched) in the largest regions with high-throughput modeSpread load across many group IDs
SQS FIFO in flight120,000 messagesHitting it stalls receives without an error
SQS message size1 MiBS3 pointer for larger (extended client)
SQS visibility / retention / delay12 h max / 14 days max / 15 min maxLong jobs heartbeat; long delays need a scheduler
RabbitMQ queue throughput~Tens of thousands msg/s per queue (typical; one Erlang process per queue)Shard hot queues; consistent-hash exchange
RabbitMQ max message size16 MiB defaultKeep messages small; payloads go to object storage
RabbitMQ memory alarm60% of RAM defaultPublishers blocked cluster-wide when tripped
Quorum queue members3 default; odd; ≤ 5 recommended; ≤ 7 nodesMore members = slower writes
Quorum queue metadata~1 MB RAM per 30,000 messagesBacklogs cost memory even when messages are on disk
Consumer ack timeout30 min defaultLong jobs ack early and track progress elsewhere

Scaling Moves in Order#

  1. Batch and long-poll (SQS) or tune prefetch (RabbitMQ). Often a 3–10× efficiency gain before adding anything.
  2. Scale consumers from Little's law, autoscaling on backlog age.
  3. Split queues by job type and priority, so each scales and fails independently.
  4. Shard hot queues (RabbitMQ consistent-hash exchange; SQS FIFO with more group IDs or high-throughput mode).
  5. Separate clusters (RabbitMQ) for workloads that must not share a memory alarm.
  6. Move to a log when the real requirement turned out to be replay or many readers.

🎯 Staff Move: "On SQS my scaling limit is the consumers and whatever they call, so I autoscale on age of oldest message and protect the downstream with a concurrency cap. On RabbitMQ, the limit is the queue process, so a hot queue becomes N sharded queues behind a consistent-hash exchange."


Failure Modes & Recovery#

1. Duplicate-Work Spiral From a Short Visibility Timeout#

  • Symptom: Downstream load doubles during a traffic peak; the same job IDs appear in logs two or three times; the queue drains slower as load rises.
  • Root cause: Visibility timeout below the p99 job time. Slow jobs reappear and run again, adding load, making more jobs slow.
  • Detection: NumberOfMessagesReceived / NumberOfMessagesDeleted ratio rising above ~1.1; ApproximateReceiveCount on processed messages; duplicate hits in the idempotency table.
  • Fix: Raise the timeout; add heartbeats; cap consumer concurrency so the downstream recovers.
  • Prevention: Timeout derived from measured p99 in config review; alarm on receive/delete ratio.

2. Poison Message Loop#

  • Symptom: A few messages fail forever; on SQS FIFO or a single-consumer RabbitMQ queue, everything behind them in the group stops.
  • Root cause: No DLQ, maxReceiveCount very high, or a RabbitMQ classic queue with no delivery limit and requeue=true.
  • Detection: Rising ApproximateAgeOfOldestMessage with flat throughput; RabbitMQ redeliver rate; the same message ID in error logs repeatedly.
  • Fix: Move the message to the DLQ manually; deploy the parser fix; redrive.
  • Prevention: DLQ on every queue; delivery limit 3–5 for parse-type failures; quorum queues (delivery limit 20 by default since 4.0).

3. RabbitMQ Memory Alarm Blocks Publishers#

  • Symptom: Producers across unrelated services hang on publish; the management UI shows a memory or disk alarm.
  • Root cause: A consumer outage let one queue grow to millions of messages; the node crossed the 60% watermark.
  • Detection: rabbitmq_alarms_memory_used_watermark; queue depth per queue; node memory by category.
  • Fix: Restore or scale the consumer; purge or shovel the backlog to another cluster if it's disposable; temporarily raise the watermark only with headroom.
  • Prevention: x-max-length with overflow to a DLQ; alarm at 40% memory; isolate critical producers on their own cluster or vhost.

4. The Backlog That Can't Drain#

  • Symptom: After a 2-hour downstream outage, the queue holds 6 hours of work; consumers at full speed only drain 20% faster than arrivals, so recovery takes 30 hours and every new message waits behind old ones.
  • Root cause: FIFO processing of a backlog whose oldest messages are now worthless, with no spare capacity.
  • Detection: Age of oldest message vs drain rate: backlog / (drain_rate - arrival_rate) = hours to recover.
  • Fix: Process fresh traffic first (move the backlog to a separate queue and drain it at low priority), drop messages past their usefulness TTL, scale consumers if the downstream can take it.
  • Prevention: Message TTLs per job type; a documented "backlog mode"; capacity headroom of ~2× steady state for recovery.

5. Node Loss or Partition on RabbitMQ#

  • Symptom: Some queues unavailable; clients reconnecting; with classic queues on the failed node, messages unavailable or lost.
  • Root cause: Node failure or network partition. Quorum queues with a majority elsewhere elect a new leader and continue; a minority side stops accepting writes for those queues.
  • Detection: Cluster membership and partition alarms; quorum queue leader changes; rabbitmq_queue_messages_ready dropping to zero on a node.
  • Fix: Restore the node; let quorum queues catch up from the leader; for classic queues, recover from the producer side.
  • Prevention: Quorum queues for anything durable; 3 or 5 nodes across zones; clients with automatic reconnection and publisher confirms.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Duplicate-work spiralReceive/delete ratio > 1.1One queue's downstreamRaise timeout, cap concurrencyConsumer team
Poison loopAge of oldest rising, throughput flatOne queue or one FIFO groupDLQ, fix, redriveConsumer team
Memory alarmBroker alarm, publish latencyEvery publisher on the clusterRestore consumer, purge or shovelPlatform; consumer team
Backlog can't drainHours-to-recover estimateEvery message behind the backlogFresh-first, TTL dropConsumer team + product (what's droppable)
Node loss / partitionMembership alarms, leader changesQueues led by that nodeQuorum failover, reconnectionPlatform
DLQ silently expiringDLQ age vs retentionLost dataAlarm on DLQ age, 14-day retentionDLQ's owning team

When to Use vs. Alternatives#

NeedPickWhy
Background jobs on AWSSQS standardNo ops, per-message retry, scales to zero
Ordered per-entity steps on AWSSQS FIFO with group = entityOrder per group, unlimited groups
Fan-out of one event to several consumers, no replaySNS → SQS or RabbitMQ topic exchangeEach subscriber isolated
Off-AWS or multi-cloud queueRabbitMQ 4 with quorum queues (or a managed RabbitMQ)Open protocol, runs anywhere
Rich routing, priorities, AMQP/MQTT/STOMP clientsRabbitMQBroker-side routing and protocol plugins
Several teams read the same events; replayKafkaLog semantics; see Kafka vs Kinesis
Delays longer than 15 minutes or cron-like schedulesA scheduler over a database, enqueuing when dueQueues aren't timers

When NOT to Use RabbitMQ or SQS#

  • Anyone needs to read the messages twice. Replay, a new consumer backfilling history, or reprocessing after a bug all need a log.
  • The call needs an answer now. A user waiting on a response wants a synchronous call with a deadline, not a queue with a reply-to.
  • Long-delay scheduling. "Send this reminder in 3 days" is a row in a scheduler table, not a message hidden for 3 days.
  • Workflow state. A multi-step process with compensation needs a workflow engine or saga state in a database; see Workflows & Sagas. Queues move the steps; they don't remember where you are.
  • RabbitMQ without an owner. A self-run broker is a cluster with memory alarms, upgrades and partitions. If nobody owns it, use a managed queue.
Diagram: When NOT to Use RabbitMQ or SQS

Operational Concerns#

What the On-Call Actually Does#

  1. Watches age of oldest message per queue against each queue's SLO. Depth alone lies.
  2. Works the DLQs: reads the top error reasons, decides fix-then-redrive vs discard with the owning team, and redrives at a throttled rate.
  3. Handles backlog recovery: estimates hours to drain, decides with product whether old messages still matter, and switches to fresh-first draining when they don't.
  4. On RabbitMQ, watches memory and disk alarms, unacked counts and node health, and finds the consumer holding thousands of unacked messages with unlimited prefetch.
  5. Runs RabbitMQ upgrades with feature flags enabled step by step; the 3.13 → 4.x upgrade required migrating classic mirrored queues to quorum queues first.
  6. Reviews new queues for the contract: owner, DLQ, timeout, retry policy, idempotency key.

Key Metrics & Alerts#

MetricHealthyAlert
ApproximateAgeOfOldestMessage (SQS) / oldest message age (RabbitMQ)Within the queue's SLO, e.g. < 60s> SLO for 5 min (page for critical queues)
DLQ ApproximateNumberOfMessagesVisible0> 0 for 15 min (ticket); age > 1 day (page)
Receive / delete ratio (SQS)~1.0> 1.1 sustained
ApproximateNumberOfMessagesNotVisible (in flight)Steady with throughputNear FIFO's 120K limit
RabbitMQ memory used / watermark< 50%> 70%
RabbitMQ unacked messages per consumer≈ prefetchFar above for one consumer (hoarding)
RabbitMQ redeliver rateLow, steadySpike (poison message or consumer crash loop)
Quorum queue leader changes~0Repeated changes (node instability)
Publisher confirm latency p99Low ms> 100 ms (disk or replica trouble)

Interview Application — Staff-Level Plays#

Which Case Studies Use RabbitMQ & SQS#

Case StudyHow Queues Are UsedKey Pattern
Message BrokerDesigning the queue itselfLease-based delivery, DLQs, per-tenant fairness
Notification SystemPer-channel queues (push, email, SMS) behind a fan-outSeparate queues per provider so one slow channel doesn't block others
Webhook DeliveryPer-endpoint retry with backoff, DLQ after daysFairness across customers; retry ladders
Distributed Job SchedulerDue jobs enqueued for workersScheduler of record in a database; the queue only moves due work
IdempotencyBroker redelivery as a duplicate sourceConsumer-side idempotency keys
Payment SystemAsync capture, refunds, reconciliation jobsAck after commit; outbox; DLQ owned by payments
Video StreamingTranscoding jobsLong jobs with heartbeats; autoscale on backlog age
Cascading FailuresQueues as shock absorbers that can hide an outageBacklog age alarms; load shedding

Every System Design Question Has a Queue Moment#

  • Notifications: "Each channel gets its own SQS queue and worker pool. When the SMS provider throttles us, only the SMS queue backs up, and its age alarm pages the notifications team, not everyone."
  • Image upload: "The upload API returns 202 and enqueues a thumbnail job with the object key. The job is idempotent on the key, the visibility timeout is 60s with a heartbeat, and after 5 failures it lands in a DLQ the media team owns."
  • Payments: "Capture runs from a quorum queue fed by an outbox, so a crash between the database commit and the publish can't lose it, and the consumer passes the payment ID as the provider's idempotency key."

What Interviewers Probe#

After You Say...They Will Ask...What They're Evaluating
"Put it on a queue""What happens when a worker crashes mid-job?"Ack point, visibility timeout, redelivery, idempotency
"Exactly-once with SQS FIFO""What's the dedup window? What about a redrive tomorrow?"Knowing 5 minutes; consumer-side idempotency
"We retry failures""What about a message that can never succeed?"DLQ, delivery limits, poison handling, ownership
"Add more consumers""What if it's FIFO with one hot group? Or one RabbitMQ queue?"Per-group serialism; one queue = one core
"RabbitMQ with mirrored queues""Which version?"Mirroring removed in 4.0; quorum queues
"The queue absorbs spikes""How do you know the backlog is a problem?"Age of oldest message vs depth; drain-time math

Common Interview Mistakes#

What Candidates SayWhat Interviewers HearWhat Staff Engineers Say
"The queue guarantees exactly-once"Will double-charge on redelivery"At-least-once delivery, exactly-once effect via an idempotency key in the consumer's transaction."
"Alert when the queue has more than 10K messages"Measures the wrong thing"Alert on age of oldest message against the queue's SLO."
"Failed messages go to a DLQ"DLQ as a trash can"DLQ after 5 receives, owned by the consumer team, age alarm, 14-day retention, practised redrive."
"One queue for all background jobs"Shared blast radius"One queue per job type, each with its own timeout, scaling and DLQ."
"Use FIFO to be safe"Paying throughput for nothing"Standard queue plus version checks; FIFO only where per-entity order is a business rule."
"RabbitMQ for the event bus that four teams read"A log problem solved with a queue"Four readers and replay means Kafka; RabbitMQ for the job queues."

L5 vs L6 vs L7 Responses#

ScenarioL5 AnswerL6 / Staff AnswerL7 / Principal Answer
"Send order confirmation emails asynchronously"SQS queue, Lambda consumerSQS + DLQ, idempotent on order ID, timeout from p99, age alarm; outbox so the email can't be lost between commit and sendAsks whether the company already has a notifications platform this should use, and who owns email DLQs across products
"Workers process some jobs twice"Use FIFO for exactly-onceDiagnoses visibility timeout vs p99 job time, adds heartbeat and idempotency tableMakes idempotency keys a required field in the org's message envelope and checks it in CI
"One customer's import delayed everyone for 6 hours"Add consumersFair queues or shuffle sharding by tenant; fresh-first drainingSets per-tenant async SLOs and an isolation standard; prices dedicated queues for the largest tenants
"Should we run RabbitMQ or use SQS?"RabbitMQ is more flexibleSQS on AWS unless routing or protocol needs justify a brokerDecides the company's messaging portfolio (one queue, one log) and retires the third system

The Staff Queue Checklist#

  1. Queue or log: "Will anyone read these messages twice? If not, it's a queue."
  2. Ack point: "Ack after the side effect commits; the seen-record is in the same transaction."
  3. Lease: "Visibility timeout 2× p99, heartbeat for long jobs; prefetch sized to processing time on RabbitMQ."
  4. Retries: "Backoff via visibility extension or TTL queues, DLQ after 5, owner and age alarm on the DLQ."
  5. Order scope: "None by default; per entity via message group or consistent hash only where required."
  6. Health signal: "Age of oldest message per queue against an SLO, plus hours-to-drain during incidents."

🎯 Staff Insight: Don't use a queue as a log (no replay), as a scheduler (15-minute delay cap on SQS), as a workflow engine (it doesn't remember state) or as an RPC transport. The strongest queue signal is describing exactly what happens when a worker dies halfway through a job.

Evaluation Rubric#

DimensionSenior (L5)Staff (L6)Principal (L7)
Delivery semantics"Guaranteed delivery"At-least-once, ack after commit, idempotent consumers, dedup windows understoodOrg-wide envelope with idempotency keys, enforced in review
Failure handlingRetriesPoison vs transient, DLQ with owner, redrive runbook, backlog drain mathDLQ inventory and age SLOs across teams; game days for backlog recovery
ScalingMore consumersLittle's law, prefetch, FIFO group limits, queue shardingMessaging portfolio and cost model
IsolationOne queueQueue per job type, per-tenant fairnessTenant isolation standard and dedicated capacity pricing
Technology choiceFamiliar brokerQueue vs log; SQS vs RabbitMQ with reasonsNumber of messaging systems the company should run, and the retirement plan

Strong hire signals

SignalWhat It Sounds Like
Lease thinking"The visibility timeout is a lease; the job has to finish or renew before it expires."
Effect over delivery"I can't get exactly-once delivery, so I build exactly-once effect."
Measures in time"Age of oldest message, because that's what the user waiting on the job experiences."
Owns the DLQ"The DLQ belongs to the consumer team, with an alarm and a redrive they've practised."
Knows when it's a log"If analytics wants these too, this isn't a queue anymore."

Lean no-hire signals

SignalWhy It Misses the Bar
Exactly-once claimed from the brokerWill ship duplicate side effects
No DLQ or unowned DLQPoison messages loop or data silently expires
Depth-based alerting onlyMisses slow drains and catches harmless spikes
FIFO everywhere "for safety"Pays throughput and head-of-line blocking for no requirement

Common false positives

  • AMQP vocabulary ≠ queue design. Knowing every exchange type says nothing about where the ack goes.
  • "We ran RabbitMQ at scale" ≠ judgment. Ask how they recovered from a memory alarm and what they'd use on AWS today.
  • Kafka-first answers ≠ sophistication. A job queue built on Kafka without retry topics or DLQ ownership is worse than SQS.

Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

At Staff level a queue is a component with the right timeout and a DLQ. At Principal level, the async tier is where failures go to hide: a synchronous outage pages someone in a minute, but a queue backing up degrades quietly for hours, and the DLQs of 60 teams are a distributed archive of business events nobody is reading. The L7 question is not "SQS or RabbitMQ?" but "How many messaging systems should we run, what contract does every queue have to meet, and how do we know, across the whole company, that async work is actually getting done?"

🧭 Principal Move: "Before choosing a broker I want the async contract: every message carries an idempotency key and an owner, every queue has an age SLO, every DLQ has an alarm and a redrive tool. Once that's standard, the broker is a detail; without it, the best broker in the world still loses orders in an unowned DLQ."

The Org-Level Fault Line#

One messaging platform vs every team picks its own.

OptionWhat WorksWhat BreaksWho Pays
Every team picks (SQS, RabbitMQ, Redis lists, Kafka)Autonomy, fast start4 systems, 4 sets of alarms, inconsistent DLQ handling, no company-wide view of backlogsIncident responders, who learn each system during the outage
One queue + one log, platform-owned (e.g. SQS + Kafka)Uniform tooling, DLQ dashboards, shared runbooksEdge cases (routing, priorities) need workaroundsPlatform team headcount
One system for everything (Kafka for jobs too)One thing to runJob semantics rebuilt by hand per team: retry topics, DLQs, delaysEvery team writing retry plumbing
Self-run RabbitMQ per teamFull controlDozens of small clusters, each one upgrade behindTeams without broker expertise

The Principal default: one managed queue and one log, both platform-owned, with a shared message envelope (idempotency key, owner, schema version, trace context), a DLQ dashboard across every queue, and an exceptions process for teams that genuinely need broker routing.

🧭 Principal Insight: "The expensive part of messaging isn't the broker bill. It's the third messaging system, because each one brings its own failure modes, its own on-call knowledge and its own way of losing messages."

Cost Model#

Assumptions: SQS at an illustrative ~$0.40 per million requests (check current pricing and volume tiers), 3 requests per message (send, receive, delete), "batched" meaning batches of 10; a RabbitMQ node ~$400/month; loaded engineer $250K/year ($21K/month). Directional only.

ScaleVolumeSQS/month (batched – unbatched)Self-run RabbitMQ/monthPeopleNote
Startup5M msg/day~$20 – $180$1.2K (3 nodes) + 0.25 FTE ($5K)Near zero on SQSSQS wins by two orders of magnitude
Growth200M msg/day~$0.7K – $7K$4K (6–9 nodes) + 0.5 FTE ($10K)0.25 FTE either way for DLQ toolingSQS still cheaper once people are counted, if batched
Enterprise5B msg/day~$18K – $180K$20–30K (clusters) + 2–3 FTE ($55K)Platform team for the envelope and dashboardsBatched SQS and a broker are comparable; unbatched SQS is the most expensive option on the page

The lever at every scale is batching and long-polling: unbatched, empty-receive-heavy consumers make SQS up to ~10× more expensive than it needs to be. The lever nobody prices is the cost of a silent backlog: six hours of delayed order confirmations is a support-ticket spike and a churn number.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to Reverse
Queue vs log for a business event streamOne-way once consumers rely on replay or its absenceRe-plumbing every producer and consumer; no history to backfill
Message envelope (idempotency key, schema version)One-way-ishEvery producer and consumer updated in lockstep
SQS FIFO vs standardTwo-way-ishNew queue and a cutover; FIFO names must end in .fifo
RabbitMQ vs SQSTwo-way, if consumers are behind an interfaceWeeks per service; routing logic is the drag
Timeouts, retry counts, prefetchTwo-wayConfig change

The Standard I'd Write#

RFC-ASYNC-002: Queue Contract

Scope: Every queue (SQS, RabbitMQ) carrying production work.

MUST
  1. Declare an owning team, routed for pages, on the queue and its DLQ.
  2. Have a DLQ with retention >= 14 days and an alarm on age of oldest message.
  3. Carry an idempotency key in the message envelope; consumers dedupe on it
     inside the side effect's transaction or via the downstream's idempotency key.
  4. Ack/delete only after the side effect is durable.
  5. Declare an age SLO; alarm when the oldest message exceeds it.
  6. Use quorum queues (RabbitMQ) for any message whose loss is a business loss.
SHOULD
  7. Size visibility timeout >= 2x measured p99 processing; heartbeat long jobs.
  8. Use one queue per job type; separate queues per priority.
  9. Use SQS fair queues or sharding for multi-tenant queues.

Exceptions: platform review; time-boxed to one quarter.
Rollout: inventory -> warn in CI -> enforce for new queues -> backfill existing.
Success metrics: zero unowned DLQs; zero DLQ messages expired unread;
  every queue reporting age SLO; duplicate side effects = 0 in audits.

What I'd Tell the VP#

"A lot of our customer-facing work happens in the background: emails, refunds, exports. When that machinery backs up, nothing turns red; customers just wait, sometimes for hours. Today each team runs it differently, and failed work sits in places nobody watches until it's deleted. I'm proposing one standard for all background work, a single dashboard showing how far behind each part is, and named owners for every failure bin. It's mostly configuration and a small platform investment, roughly one engineer for a quarter, and it turns silent delays into alerts we act on within minutes."

Principal Interview Signals#

SignalWhat It Sounds Like
Sees the async tier as a risk surface"Queues hide outages; my first dashboard is age of oldest message across every queue."
Limits the portfolio"One queue and one log. A third messaging system needs a business case."
Standardises the envelope"Idempotency key, owner, schema version and trace context on every message."
Prices the silent backlog"Six hours of delayed confirmations is a ticket spike we can put a number on."
Designs tenant isolation"Fair queues by tenant, dedicated queues for the top 1% by volume."

Staff answers that L7 interviewers find insufficient:

  • "Every queue has a DLQ" without who reads it, the alarm, or what happens on day 15.
  • "Make consumers idempotent" as advice, rather than a required envelope field enforced in CI.
  • "Use SQS because it's managed" without the portfolio question: what else does the company run, and what gets retired.

How Real Companies Built It#

Trello — Leaving RabbitMQ for WebSocket Fan-Out#

Trello's engineering team described running websocket updates through a cluster of 15 RabbitMQ instances for about three years, with each websocket process creating a transient queue and binding it to a topic exchange per subscribed model. The problems were partition handling ("from split-brain to complete cluster failure") and binding churn: after mass socket disconnects, a flood of binding add and remove commands made the cluster unresponsive, even to monitoring. They recorded 4 RabbitMQ-caused outages in the month before switching, evaluated SNS + SQS, Kinesis and Redis Streams, and moved to Kafka, reporting lower memory use and a 5× cost reduction (Atlassian engineering).

Staff insight: This was a fan-out problem with per-connection queues, which is a log-shaped workload on a queue-shaped broker. It also predates quorum queues becoming the default. In an interview, the lesson is to name binding churn and partition behaviour before proposing a queue per session.

Airbnb — Dynein, Delayed Jobs on SQS#

Airbnb built Dynein, an open-source distributed delayed-job system, on SQS for the job queues and DynamoDB for the schedule of jobs due in the future: a scheduler queries for overdue jobs and dispatches them to service queues consumed by workers (Dynein on GitHub). The team explained choosing SQS because it is simple to reason about and gives at-least-once delivery, dead-letter queues and individual message acknowledgement out of the box (InfoQ coverage).

Staff insight: The queue moves due work; the database remembers what is scheduled. That split is exactly the answer to "can SQS schedule a job for next week?" (no: 15 minutes maximum delay).

Amazon — Avoiding Insurmountable Queue Backlogs#

An AWS Builders' Library article by a senior principal engineer describes how Amazon services keep queue backlogs from becoming multi-hour outages: shuffle sharding customers across a few queues (used in AWS Lambda's asynchronous invocation path), separate queues per customer in some systems, spillover queues for traffic over a customer's rate, working a backlog queue only after catching up on live traffic, dropping messages past a TTL, heartbeating long jobs, and throttling inbound traffic in proportion to backlog size (AWS Builders' Library).

Staff insight: The hard queue problem isn't delivery, it's recovery. When asked "what happens after the outage ends?", the Staff answer is fresh-first draining, per-tenant isolation and TTLs on work that has gone stale.


Practice Drill#

Drill 1: The Six-Hour Backlog#

Prompt: "Your multi-tenant SaaS sends report exports, emails and webhooks through one SQS standard queue consumed by 40 workers. Last week one customer's bulk import enqueued 4 million jobs; every other customer's emails were delayed for six hours, and support found some exports had been generated twice. Redesign the async tier."

Staff Answer

Two separate failures. The delay is an isolation failure: one queue for three job types and every tenant means one burst blocks everyone, and nothing alerted because the alarm was on depth, which is always high during imports. The duplicates are a lease failure: exports take up to 90 seconds at p99 and the visibility timeout is the 30-second default, so slow exports reappear and run again. Fixes in order. First, split by job type: emails, webhooks, exports, each with its own worker pool, timeout and DLQ, so a slow export can never delay a password reset. Second, tenant fairness: set MessageGroupId = tenant_id on the standard queues to get SQS fair queues, so a tenant with a disproportionate share in flight is deprioritised while quiet tenants keep low dwell time; for the top handful of tenants by volume, a dedicated bulk queue drained at a capped rate. Third, the lease: export visibility timeout at 180s (2× p99) with a heartbeat every 60s for jobs that run long, and an idempotency record keyed by export_id written before the file is published, so a redelivery returns the existing file. Fourth, signals: alarm on ApproximateAgeOfOldestMessage per queue against an SLO (emails 60s, webhooks 5 min, exports 15 min), alarm on receive/delete ratio above 1.1, DLQ after 5 receives with the owning team paged when its oldest message passes an hour. Fifth, recovery: a runbook to move a backlog to a low-priority queue so fresh work flows first, and TTLs on emails that are useless after a day. Autoscale workers on backlog age, capped by what the email and webhook providers will accept.

Why this is L6:

  • Separates the two symptoms into two root causes: isolation and lease length.
  • Uses age of oldest message per job type as the health signal, with SLOs per queue.
  • Fixes duplicates with both a correct timeout and consumer idempotency, not FIFO.
  • Plans the recovery path (fresh-first, TTLs), not just prevention.

What L7 adds:

  • Turns it into a queue contract for every team: owner, DLQ alarm, age SLO, idempotency key in the envelope.
  • Prices tenant isolation: dedicated bulk capacity as a paid tier for the largest customers.
  • Adds a company-wide async dashboard so the next silent backlog is visible to leadership, not discovered by support.
❌ Common L5 Trap

"Switch to SQS FIFO for exactly-once processing so exports aren't duplicated, and add more workers so the backlog drains faster."

Why this misses: FIFO's dedup covers producer retries within 5 minutes, not a slow consumer whose visibility timeout expired, so the duplicates continue. FIFO also serialises each message group and caps throughput, which makes the backlog worse. More workers on one shared queue still process the 4 million import jobs before other tenants' emails, because the problem is isolation, not capacity. The fix is separate queues, tenant fairness, a correct lease and idempotent consumers.


Quick Reference Card#

Model:           queue = broker tracks each message; acked = gone (no replay)
                 log = consumers track offsets; data kept (Kafka, RabbitMQ streams)
Delivery:        at-least-once everywhere; exactly-once EFFECT via consumer idempotency
SQS standard:    nearly unlimited throughput; best-effort order; duplicates possible
SQS FIFO:        order per MessageGroupId, 1 in flight per group; dedup window 5 min;
                 300 calls/s per partition (3,000 msg/s batched); up to 70K TPS high-throughput;
                 120,000 in flight max
SQS limits:      visibility 30s default, 12h max; retention 4d default, 14d max;
                 delay <= 15 min; long poll <= 20s; batch 10; message 1 MiB
SQS fair queues: standard queues + MessageGroupId as tenant ID -> noisy-neighbour protection
RabbitMQ:        exchange (direct/topic/fanout/headers) -> bindings -> queues
Quorum queues:   Raft, 3 replicas default, odd count, <= 7 nodes; delivery limit 20 (4.0+)
                 ~1 MB RAM per 30K messages of metadata
Classic mirror:  removed in RabbitMQ 4.0
RabbitMQ limits: consumer ack timeout 30 min; memory watermark 60% -> publishers blocked;
                 max message 16 MiB; one queue ~ one core
Sizing:          consumers = arrival rate x processing time (Little's law), run at ~70%
                 visibility timeout = 2 x p99; heartbeat long jobs
Health:          age of oldest message per queue vs SLO; receive/delete ratio; DLQ age

RED FLAGS
  - Ack/delete before the side effect commits
  - Visibility timeout below p99 processing time
  - DLQ with no owner, no alarm, default 4-day retention
  - One queue for every job type and tenant
  - "Exactly-once" credited to the broker
  - Alerting on depth instead of age
  - Classic (unreplicated) RabbitMQ queues for business-critical messages
  - A queue where several teams need replay (that's a log)
  1. Loading the index…