Hiring BarSupport

Kafka vs SQS vs RabbitMQ

Comparison19 min read3 diagrams

People line these three up because all of them sit between a producer and a consumer and all of them are called "messaging". The resemblance stops there. Kafka is a replicated log that consumers read at their own offset and that keeps data after it is read. SQS is a managed work queue where each message is handed to one consumer, hidden while it is worked on, and deleted when it is done. RabbitMQ is a broker that routes messages through exchanges into queues, with per-message acknowledgement and rich routing rules. The one question that decides it: once a message has been handled, does anyone ever need to read it again? If yes — replay, a second consumer group next quarter, a rebuilt search index — you want a log. If no, and each message is a unit of work to be done once and forgotten, you want a queue, and then the question becomes whether you want to run it.

Go deeper:

  • For the full queue contract (ack points, visibility timeouts, quorum queues and DLQ ownership), see RabbitMQ & SQS.

The Verdict#

Default to SQS for task queues on AWS, Kafka for event streams that several teams will read or replay, and RabbitMQ only when you need broker-side routing or you are not on AWS and want a queue without running Kafka.

Pick Kafka whenPick SQS whenPick RabbitMQ when
Several independent consumers read the same events (analytics, search indexing, fulfillment)Each message is a job done once: send email, resize image, call a webhookYou need routing in the broker: topic patterns, headers, fanout to many queues
You need replay: reprocess 3 days of events after a bugYou want zero brokers, zero capacity planning, zero upgradesYou are on-prem or multi-cloud and SQS is not an option
Ordering per entity matters at 10K+ events/secPer-message retry, delay (≤ 15 min) and dead-lettering are the main needsLow-latency request/reply or priority queues matter more than replay
Throughput is 10s of MB/s and growing; stream processing (Flink, Kafka Streams) is plannedThroughput is spiky and unpredictable, and you want cost to follow it to zeroMessage volume is moderate (thousands to tens of thousands per second per queue)
You already run Kafka, or can buy MSK / a managed Kafka serviceYou are on AWS and the consumer is Lambda, ECS or EC2Your team already knows AMQP and runs RabbitMQ well

The decision in one picture:

Diagram: The Verdict

🎯 Staff Move: "Before I pick a broker I want to know if anyone will ever read these messages twice. If the answer is 'analytics wants them too', that's a log and I'll use Kafka. If it's 'no, each job runs once', that's a queue and I'll use SQS, because the cheapest broker to operate is the one I don't operate."

At a Glance#

DimensionKafkaSQSRabbitMQ
Data modelPartitioned, append-only log; consumers track offsetsQueue; message hidden on receive (visibility timeout), deleted on ackExchanges route to queues via bindings; per-message ack; optional streams (log)
After consumptionStays until retention expires (7 days default; tiered storage for longer)Deleted by the consumer; unconsumed messages kept up to 14 days (4 default)Removed on ack (queues); kept by retention (streams)
Consistency / deliveryAt-least-once to the outside world; exactly-once only for Kafka→Kafka transactionsAt-least-once (standard); exactly-once processing within a 5-minute dedupe window (FIFO)At-least-once with publisher confirms + manual acks; quorum queues replicate via Raft
OrderingTotal order within a partitionNone guaranteed (standard); per message group (FIFO)Per queue with a single consumer; lost with competing consumers and redelivery
Throughput (order of magnitude)100s of MB/s per cluster on modest brokers; GB/s on large onesStandard: nearly unlimited; FIFO: 300 API calls/s per partition, up to 70K TPS in the largest regions with high-throughput mode (as of 2026)~10K–50K msg/s per queue (one queue is bound to one CPU core); scale by sharding queues
Latency (p99, same region)~5–20 ms with acks=all; consumer adds poll interval~10–50 ms send; receive bounded by long-poll settingsSub-ms to a few ms for transient messages; ~5–20 ms for quorum queues with fsync
Scaling modelAdd brokers and partitions; consumer parallelism ≤ partitions (classic groups)Automatic; consumers scale freelyVertical per queue; horizontal by adding queues and consumers
Operational burdenHigh self-hosted (brokers, KRaft controllers, rebalancing, upgrades); medium on a managed serviceNear zeroMedium: cluster membership, memory alarms, quorum-queue tuning, upgrades
Managed optionsAmazon MSK, Confluent Cloud, Aiven, othersNative AWS onlyAmazon MQ for RabbitMQ, CloudAMQP, others
Cost shapeFixed: brokers + disks + cross-AZ replication traffic, paid 24/7Per request (each 64 KB chunk = 1 request); scales to zeroFixed: nodes paid 24/7; licence-free open source
Message size1 MB default (configurable; large payloads belong in object storage)1 MiB max (as of 2026); larger via S3 pointerConfigurable; practical limit is memory, keep it small

How They Actually Differ#

Log vs Queue: Who Remembers What Was Read#

This is the difference that drives all the others. In Kafka, the broker does not know or care whether a message was processed. Each consumer group stores one number per partition — its offset — and the data stays until retention removes whole segments. That is why a new consumer group can start reading from the beginning, and why replaying a day of events after a bug is a single offset reset.

In SQS and classic RabbitMQ queues, the broker tracks each message. A receive makes the message invisible (SQS) or unacked (RabbitMQ); an ack deletes it. There is nothing to replay. If a second team wants the same messages, you need a second queue and a fanout in front: SNS → many SQS queues, or a RabbitMQ fanout exchange.

Diagram: Log vs Queue: Who Remembers What Was Read

Who pays: with a log, storage pays — you hold every byte for the retention window whether anyone reads it or not. With a queue, the team that shows up later pays — there is no history to backfill from, so they need a separate export from the source database.

Ordering and Parallelism Are the Same Decision#

Kafka gives a strict order per partition and assigns each partition to exactly one consumer in a classic consumer group. Ordering and parallelism are therefore the same knob: 48 partitions means at most 48 consumers working in parallel, and one slow message blocks everything behind it in that partition. That is head-of-line blocking by design.

Kafka 4.2 (February 2026) made share groups production-ready: consumers in a share group consume the same partitions cooperatively, acknowledge records one at a time, and get per-record delivery counts — queue semantics on a Kafka topic. They give up ordering to get that. Useful when you already run Kafka and want a queue without a second system; not a reason to adopt Kafka if a queue is all you need.

SQS standard gives no ordering and unlimited parallelism. SQS FIFO orders per message group ID: messages in one group are delivered strictly in order and one at a time, while different groups run in parallel. That is the same idea as a Kafka partition key, but with no partition count to size — group count is unlimited. RabbitMQ preserves order inside a queue only with one consumer and no redelivery; with competing consumers and nacks, order is gone.

RequirementKafkaSQSRabbitMQ
"All events for order 42 in order"Key by order_idFIFO, MessageGroupId = order_idConsistent-hash exchange to N queues, one consumer each
"One poison message must not block others"Retry topics + DLQ, or share groupsNative: visibility timeout + redrive to DLQNative: nack + dead-letter exchange; quorum queues cap redeliveries (20 by default in 4.x)
"Scale to 500 workers"Needs ≥ 500 partitions (classic groups)Just add workersAdd workers; one queue tops out around one core

Retries, Delays and Dead Letters#

Queues were built for jobs that fail one at a time. SQS has a visibility timeout (up to 12 hours), per-message delay (up to 15 minutes), and a redrive policy that moves a message to a dead-letter queue after N receives. RabbitMQ has per-message TTL, dead-letter exchanges and, in quorum queues, a delivery limit. Each of these is one config line.

Kafka has none of them on a classic consumer group. A failed record is either retried in place (blocking the partition), skipped (data loss unless you write it elsewhere), or written to a retry topic with a delay enforced by the retry consumer. The pattern works — it is a well-trodden production pattern — but someone has to build, own and page on it.

What the same "retry 3 times with backoff, then park it" requirement costs in each:

SQS (configuration, no code):
  queue.VisibilityTimeout   = 2 × p99 processing time      # e.g. 60s
  queue.RedrivePolicy       = { maxReceiveCount: 4, deadLetterTargetArn: jobs-dlq }
  consumer: on failure → do nothing; message reappears after the timeout
  consumer: on success → DeleteMessage

RabbitMQ (configuration + a nack):
  queue type = quorum, x-delivery-limit = 3, x-dead-letter-exchange = jobs.dlx
  consumer: on failure → basic.nack(requeue=true)
  delayed retry → per-message TTL on a wait queue that dead-letters back

Kafka, classic consumer group (code you own):
  on failure → produce record to jobs.retry.1m with header attempt=1; commit offset
  retry-1m consumer → sleep until record.timestamp + 60s; reprocess or → jobs.retry.10m
  after attempt 3 → produce to jobs.dlq; alert; replay tool reads jobs.dlq
  ordering for that key is now broken — later records overtook the failed one

The last line is the one candidates miss: a retry topic means a later event for the same key can be processed before the failed one. If per-key order matters, the consumer must also park every later event for that key until the failure resolves.

🎯 Staff Insight: If the requirement list has "retry with backoff", "delay", "DLQ" and "per-message visibility" in it, the requirement is a queue. Bending Kafka into one costs a retry-topic framework and an owner. Put a queue behind the Kafka consumer instead.

Routing: Who Decides Where a Message Goes#

RabbitMQ's distinctive feature is routing in the broker. Producers publish to an exchange with a routing key; bindings decide which queues get a copy — exact match (direct), wildcard (orders.*.eu, topic), all (fanout) or by header. Adding a new consumer of "all EU order events" is a new binding, not a producer change.

Kafka routes by topic and partition key only; consumers filter client-side. SQS has no routing at all; SNS subscription filter policies or EventBridge rules do it in front. If your routing table is complex and changes often, RabbitMQ or a managed event bus earns its place. If it is "every consumer gets everything on this topic", Kafka's model is simpler.

Durability and Replication#

KafkaSQSRabbitMQ
ReplicationRF=3 across AZs; acks=all + min.insync.replicas=2 survives one broker lossManaged, multi-AZ, invisible to youQuorum queues: Raft, 3 replicas by default; classic mirrored queues were removed in 4.0
What you configure wrongacks=1, unclean leader election on, RF=2Visibility timeout shorter than processing time → duplicatesClassic (non-replicated) queues for important data; no publisher confirms
Durable write costReplication traffic across AZs, often the largest line on the billIncluded in request priceFsync per batch; quorum queues use more memory and disk than classic

Where Each One Breaks#

Kafka#

FailureWhat happensDetectionOwner
Consumer lag spiralA slow downstream makes lag grow; consumers reading old data from disk evict page cache for everyoneconsumer_lag_seconds per groupConsuming team
Rebalance stormProcessing exceeds max.poll.interval.ms; consumer is kicked, partitions move, work restarts, repeatRebalance rate, records_consumed_rate dropping to zeroConsuming team
Hot partitionOne key (a huge tenant) dominates a partition; one consumer pinned at 100%Per-partition bytes-in skew > 3× medianProducing team (owns the key)
Adding partitions to a keyed topicKeys remap; per-key ordering breaks during the transitionOut-of-order sequence checks in consumersPlatform + producer
Disk fullRetention by time while traffic doubled; broker stops accepting writesDisk used % per brokerPlatform

The production surprise: the bill is mostly cross-AZ replication and idle capacity, not messages. A cluster sized for Black Friday costs the same in February.

SQS#

FailureWhat happensDetectionOwner
Visibility timeout shorter than processingMessage reappears while still being worked on; two workers process itApproximateReceiveCount > 1 on success, duplicate side effectsConsuming team
Poison message without a DLQRetried until retention expires (up to 14 days), burning requests and worker timeApproximateAgeOfOldestMessage climbingConsuming team
DLQ nobody readsFailures pile up silently; messages expire from the DLQ tooDLQ depth > 0 alertNamed owner per DLQ
FIFO group hot spotOne message group holds most traffic; that group processes seriallyAge of oldest message rising while workers idleProducing team
Lambda concurrency starvationEvent source scales consumers up; downstream database falls overDB connections, Lambda throttlesConsuming team

The production surprise: duplicates are normal. Standard queues deliver at least once by design, and every consumer must be idempotent. FIFO's deduplication only covers a 5-minute window and only the send side.

RabbitMQ#

FailureWhat happensDetectionOwner
Memory alarm → publishers blockedConsumers fall behind, queues grow in memory, broker blocks all publishing connectionsmem_alarm, queue depth, blocked connectionsPlatform
Network partitionCluster splits; with older classic mirroring this caused split-brain and message lossPartition events in logs, node-down alertsPlatform
Single hot queueOne queue is one Erlang process; it pegs a core at ~10K–50K msg/sPer-queue publish rate flat while CPU on one core hits 100%Platform + producer
Binding churnThousands of short-lived queues created and destroyed (one per client) overloads metadataQueue create/delete rateApplication team
Unbounded prefetchOne consumer grabs thousands of unacked messages and dies; redelivery burstUnacked count per consumerConsuming team

The production surprise: RabbitMQ is fast when queues are short. It is designed to keep queues near empty; a backlog of millions of messages changes its performance profile, and that is exactly when you need it most.

The Same Outage, Three Ways#

The downstream database slows to 20% of normal speed for 30 minutes. Producers keep publishing 5,000 msg/s. How each system absorbs the 30-minute backlog of ~7 million messages:

t=0       DB latency 10ms → 50ms; consumers drop from 5,000/s to 1,000/s
t=+5min   Kafka:    lag 1.2M records; disk absorbs it; producers unaffected
          SQS:      1.2M visible messages; producers unaffected; age-of-oldest alarm fires
          RabbitMQ: 1.2M messages; queue pages to disk; memory climbing
t=+15min  Kafka:    lag 3.6M; caught-up consumers fine; lagging ones read from disk
          SQS:      3.6M; still invisible to producers; cost unchanged per message
          RabbitMQ: memory high watermark hit → broker blocks ALL publishers
                    (including unrelated services on the same cluster)
t=+30min  DB recovers. Kafka and SQS drain in ~10 minutes at 2× consumer scale.
          RabbitMQ: publishers were blocked for 15 minutes → upstream timeouts → user-facing errors

Who pays: with Kafka and SQS the consuming team pays (lag alerts, a slow drain). With an undersized RabbitMQ the producers pay — and so does every other tenant of that cluster. The fix is not "never use RabbitMQ"; it is quorum queues with length limits, per-tenant vhosts or clusters, and producers that treat a blocked connection as backpressure rather than an error.


Cost and Operations#

KafkaSQSRabbitMQ
Who runs itPlatform team (self-hosted) or vendor (managed); either way someone owns partitions, quotas and schemasAWS; your team owns DLQs and consumer settingsPlatform team or vendor (Amazon MQ, CloudAMQP)
Bill scales withBrokers × hours + storage × retention + cross-AZ bytesRequests (each 64 KB chunk of a call counts as one)Nodes × hours
Idle costFull cluster, 24/7~$0Full cluster, 24/7
Rough list priceA small 3-broker managed cluster is hundreds of $/month; production clusters run thousands to tens of thousands~$0.40 per million standard requests, ~$0.50 FIFO, first 1M/month free (us-east-1, as of 2026)A 3-node managed cluster is hundreds of $/month
People cost0.5–2 engineers for a self-hosted multi-team clusterNear zero0.25–1 engineer

Break-even math. 1,000 messages/sec, each a send + receive + delete = 3 requests → ~7.8 billion requests/month → roughly $3,100/month on SQS standard, before batching. Batching 10 messages per call cuts it to ~$310. The same 1,000 msg/s is trivial for a 3-broker Kafka cluster that costs about the same as the unbatched SQS bill but also gives you replay. At 100,000 msg/s, SQS costs scale linearly while Kafka grows in steps; at that rate a log is almost always cheaper per message, and the question becomes whether you can staff it.

What the on-call actually watches for each:

SignalKafkaSQSRabbitMQ
The SLO metricConsumer lag in seconds per groupApproximateAgeOfOldestMessageQueue depth and consumer utilisation
The "act now" alertUnder-replicated partitions > 0 for 5 minDLQ depth > 0Memory or disk alarm raised
Routine workPartition reassignment, broker upgrades, quota tuning, schema reviewsReviewing DLQs, tuning visibility timeoutsRolling upgrades, queue type migrations, vhost limits
Upgrade riskRolling, broker by broker; client compatibility to checkNone you controlRolling with feature flags; major versions (3.x → 4.x) need a plan

🧭 Principal Insight: "SQS's price is all variable and Kafka's is mostly fixed plus people. Below a few thousand messages a second, SQS wins on total cost because nobody is on call for it. Above that, the per-request bill starts paying for a Kafka team — but only if that team serves more than one use case."


Switching Later#

MoveHowWhat's hard
SQS → KafkaProducer dual-writes (or writes to Kafka and a bridge feeds SQS); move consumers one by oneConsumers built around per-message retry and delay need a retry-topic design; ordering assumptions change
Kafka → SQSKafka consumer forwards to SQS for the queue-shaped part of the workloadYou lose replay; any consumer relying on offsets or reprocessing needs a new source
RabbitMQ → KafkaKafka Connect source connectors or app dual-writes; routing logic moves into topic design and consumer filtersBroker-side routing has to be redesigned as topics; request/reply patterns do not translate
RabbitMQ → SQSMostly mechanical for work queues; exchanges map to SNS topics with filter policiesPriority queues and complex bindings have no direct equivalent
Any → anyRun both in parallel; compare counts per minute; cut over per consumerIdempotency: during the overlap, duplicates are guaranteed

The safe shape for any of these migrations is the same: put the new system next to the old one, move consumers before producers, and keep the old path until counts match for a full business cycle.

Diagram: Switching Later

One-way doors: Kafka partition counts for keyed topics (adding partitions remaps keys); the event schema once 10+ teams consume it; FIFO vs standard for an SQS queue (you cannot convert — you create a new queue). Two-way doors: retention settings, consumer counts, managed vs self-hosted for the same engine.


How Real Companies Chose#

Slack: Kafka in Front of a Redis Job Queue#

Slack's job queue runs work too slow for a web request — message posting side effects, push notifications, unfurls, billing — at over 1.4 billion jobs a day with peaks of 33,000 per second. After an outage where Redis hit its memory limit and could no longer accept new jobs, Slack put Kafka in front as a durable buffer: a Go service (Kafkagate) writes jobs to Kafka, and a relay (JQRelay) moves them into the existing Redis workers at a controlled rate. The Kafka cluster was 16 brokers, 32 partitions per topic, RF=3 and 2-day retention (Slack Engineering).

Staff insight: They did not replace the queue with Kafka. They used Kafka for what a log is good at — absorbing a backlog durably — and kept the queue for what it is good at: running jobs. The log and the queue are layers, not rivals.

Trello: From RabbitMQ to Kafka for Real-Time Updates#

Trello used RabbitMQ for three years to fan out real-time updates to its websocket servers, then moved to Kafka. They cite RabbitMQ's behavior during network partitions and the cost of creating and destroying queues: when many socket servers disconnected at once, the burst of queue and binding changes could make the cluster unresponsive. They evaluated Kafka, SNS + SQS, SNS + FIFO SQS, Kinesis and Redis Streams; after moving to Kafka they reported memory down 33%, cost down about 5×, and no outages from the system compared with four RabbitMQ-related outages in the month before (Atlassian Engineering).

Staff insight: The failure was churn, not throughput. A design that creates a queue per connection puts a metadata storm into every reconnect wave. A log with consumer offsets has no per-consumer broker state to churn.

Amazon Prime Day: A Queue That Does Not Need a Capacity Plan#

AWS reports that during Prime Day 2025, SQS set a peak of 166 million messages per second, while Kinesis Data Streams peaked at 807 million records per second (AWS News Blog). These are Amazon's own Prime Day workloads spread over many queues and streams, not one queue — but they show why "will SQS keep up?" is rarely the right worry for standard queues.

Staff insight: For standard SQS the ceiling is your consumers and your downstream, not the queue. The capacity question moves to "what does a 10× backlog do to the database my workers write to?"


Follow-Ups to Expect#

After You Say...They Will Ask...What They're Testing
"Kafka, keyed by order_id""One merchant is 30% of traffic. What happens?"Hot partitions; whether you know key choice is the design
"SQS standard for the jobs""How do you stop a job running twice?"Idempotency keys, visibility timeout sized to p99 processing time
"SQS FIFO to keep order""What's your throughput ceiling?"300 TPS per partition, high-throughput mode, group ID spread
"RabbitMQ topic exchange""The broker runs out of memory during a consumer outage. Now what?"Publisher blocking, lazy/quorum queues, backpressure to producers
"Kafka for retries too""How does a delayed retry work without blocking the partition?"Retry topics, DLQ ownership, or knowing to use a queue instead
"Exactly-once with Kafka""Your consumer writes to Postgres and calls a payment API."EOS boundary: Kafka→Kafka only; outside, idempotent consumers
"We'll move to Kafka later""What breaks in the migration?"Dual-write overlap, duplicates, ordering semantics changing

What to Say in the Interview#

"I'll ask one question first: does anything ever need to read these messages again? Replay or a second consumer means a log; one-time jobs mean a queue."

"For the notification jobs I'll use SQS with a DLQ owned by the notifications team, a visibility timeout of twice the p99 send time, and an idempotency key per notification, because delivery is at least once."

"The order events go on Kafka keyed by order_id — fulfillment, analytics and search all read them independently, and 7 days of retention is our replay budget."

"I'd only bring in RabbitMQ if we're off AWS or need broker-side routing; otherwise it's a third system for the platform team to run with no capability the other two don't cover."


  1. Loading the index…