Hiring BarSupport

Kafka vs Kinesis

Comparison16 min read3 diagrams

On AWS this is the log decision. Both are partitioned, ordered, replayable streams: producers write records with a key, the key picks a partition (Kafka) or shard (Kinesis), order holds within it, and consumers track their own position. Underneath the shared model they make opposite bets. Kinesis Data Streams is a serverless AWS service with fixed per-shard limits and per-shard (or per-GB) pricing: nothing to operate, everything metered. Kafka is an open-source system you run or buy as a managed service (Amazon MSK and others): brokers to size, but no per-shard read limits, a large ecosystem, and economics that improve with scale. The question that decides it: will this stream stay small and AWS-native, with a handful of consumers — or will it become the event backbone that many teams read, replay and process?

The Verdict#

Use Kinesis for a single AWS-native pipeline at modest throughput with a few consumers. Use managed Kafka for a company-wide event backbone, for high sustained throughput, for many consumer groups, or when you need the Kafka ecosystem or portability.

Pick Kinesis whenPick Kafka when
One pipeline: clickstream → Lambda → S3, IoT telemetry → analyticsA shared event backbone that 5–50 teams will consume
Throughput is up to ~tens of MB/s, or very spiky (on-demand mode)Sustained throughput is 100s of MB/s, where broker economics beat per-shard pricing
1–5 consumers per streamMany independent consumer groups reading the same data
You want zero operations and IAM-native everythingYou need Kafka Connect, CDC (Debezium), Kafka Streams, or the wider client ecosystem
Retention of 24 hours to 7 days covers replay needsYou want long retention cheaply (tiered storage to object storage) or compacted topics as tables
You are certain to stay on AWSMulti-cloud, on-prem, or avoiding an AWS-only API matters

🎯 Staff Move: "If this is one pipeline from our app to S3 with a Lambda in the middle, Kinesis on-demand — no brokers, done this afternoon. If it's the stream that fulfillment, search, fraud and analytics will all read for years, I'd put it on managed Kafka now, because moving a backbone later means moving every consumer."

Diagram: The Verdict

At a Glance#

DimensionKafkaKinesis Data Streams
Data modelTopics → partitions → offset-ordered records; optional log compactionStreams → shards → sequence-numbered records
OrderingPer partitionPer shard
Write limit per unitNo fixed per-partition cap; bounded by broker disk and network (often ~10 MB/s per partition is a comfortable planning number)1 MB/s or 1,000 records/s per shard (as of 2026)
Read limit per unitBounded by broker network; each consumer group reads independently2 MB/s per shard shared by all standard consumers (5 GetRecords calls/s); 2 MB/s per consumer per shard with enhanced fan-out
Consumers per streamEffectively unlimited consumer groupsUp to 20 enhanced fan-out consumers per stream (50 with On-demand Advantage), as of 2026
Max record size1 MB default (configurable)Up to 10 MiB, with large records drawing on burst capacity (as of 2026)
RetentionAny (7 days default); tiered storage to object storage for months or years; compaction keeps latest per key24 hours default; up to 365 days at extra cost
Delivery semanticsAt-least-once; idempotent producer; transactions for exactly-once Kafka → KafkaAt-least-once; producer retries can duplicate; no transactions
Latency~5–20 ms produce with acks=all; consumer fetch adds little~tens of ms put; standard polling consumers see ~200 ms+ propagation; enhanced fan-out pushes in ~70 ms
Scaling modelAdd partitions (remaps keys) and brokers (reassign partitions)On-demand: automatic; provisioned: split/merge shards (UpdateShardCount)
Operational burdenMedium on managed (sizing, partitions, upgrades, quotas); high self-hostedVery low (shard counts in provisioned mode)
Managed optionsAmazon MSK (provisioned, Express, serverless), Confluent Cloud, Aiven, Redpanda and othersAWS only
Cost shapeBrokers × hours + storage + cross-AZ replication traffic (on self-managed EC2)Provisioned: shard-hours + PUT units; on-demand: per GB in/out + per-stream hour; extras for retention and fan-out
EcosystemConnect (hundreds of connectors), Debezium CDC, Kafka Streams, Flink, schema registriesLambda, Firehose, Managed Service for Apache Flink, KCL; AWS integrations

Feature checklist for the decisions interviewers probe (as of 2026):

CapabilityKafkaKinesis
Replay from a timestampYes (offsetsForTimes)Yes (AT_TIMESTAMP iterator)
Keep latest value per key foreverYes (compacted topics)No
Exactly-once read-process-writeYes, between Kafka topics (transactions)No; idempotent consumers
Queue-style per-record acksYes, share groups (production-ready in 4.2)No
Change data capture from databasesDebezium / ConnectVia DMS or custom; DynamoDB can stream into Kinesis natively
Serverless consumerLambda event source for MSK and self-managed KafkaLambda event source, native
Runs outside AWSYesNo

How They Actually Differ#

Fixed Shard Limits vs Broker Capacity#

A Kinesis shard is a contract: 1 MB/s or 1,000 records/s in, 2 MB/s out, 5 read calls per second. You size the stream by dividing your peak by those numbers, and you exceed them with ProvisionedThroughputExceededException. On-demand mode does the division for you (new streams start at 4 MB/s write and scale up to 10 GB/s in the largest regions, 200 MB/s elsewhere by default, as of 2026), but the per-shard ceiling still applies to any one partition key.

A Kafka partition has no fixed ceiling; it is bounded by the broker's disk, network and the other partitions sharing it. That is more headroom per key and more responsibility: the platform team decides how many partitions each broker carries and when to add brokers.

Planning questionKafkaKinesis
50 MB/s in, 1 KB records~12–24 partitions on a few brokers (plan ~3× headroom)≥ 50 shards (provisioned) or on-demand
One key at 3 MB/sFits in one partitionExceeds one shard — must split the key
3 consumers reading everything3 consumer groups, no extra config150 MB/s out > 100 MB/s shared read → enhanced fan-out or more shards

Who pays: Kinesis makes the application team pay when one key or one consumer outgrows a shard. Kafka makes the platform team pay to keep brokers sized.

Fan-out: Who Shares the Read Budget#

Every standard Kinesis consumer on a shard shares 2 MB/s and 5 GetRecords calls per second. Two consumers polling once a second is fine; five consumers polling at the 200 ms each needs leaves no room. Enhanced fan-out gives each registered consumer its own 2 MB/s per shard pushed over HTTP/2, at an extra charge, up to 20 consumers per stream (50 with On-demand Advantage).

Kafka consumer groups each read at their own pace from the broker; adding the tenth consumer group costs broker network egress, not a per-consumer contract. This is the single biggest reason streams that start on Kinesis move to Kafka when more teams want the data.

Diagram: Fan-out: Who Shares the Read Budget

🎯 Staff Insight: Count consumers before you count megabytes. A 5 MB/s stream with eight consuming teams is a worse fit for Kinesis than a 50 MB/s stream with one, because the shared read budget and per-consumer fan-out charges scale with readers, not writers.

The producer side of "durable, ordered per order, no duplicates from retries" in each:

Kafka producer
  acks = all                      # leader waits for in-sync replicas
  enable.idempotence = true       # broker drops duplicate retries per partition
  min.insync.replicas = 2         # (topic config) with RF = 3
  key = order_id                  # same key → same partition → ordered
  linger.ms = 5, batch.size = 64KB, compression.type = zstd

Kinesis producer
  PutRecords (≤ 500 records/request), PartitionKey = order_id
  SequenceNumberForOrdering on PutRecord when strict per-key order across retries matters
  on partial failure: retry only the failed records, with backoff
  no broker-side dedupe → consumers dedupe by an event ID carried in the payload
  aggregation (KPL) packs small records into one to save PUT units and the 1,000 records/s limit

The difference that matters in review: Kafka can drop producer-retry duplicates at the broker; Kinesis cannot, so every Kinesis consumer needs an idempotency key from day one.

Retention and Replay#

Both replay by position: Kafka by offset or timestamp, Kinesis by sequence number or timestamp (AT_TIMESTAMP iterator). The difference is how far back and at what cost. Kinesis keeps 24 hours by default; extending to 7 days and then up to 365 days adds per-shard-hour and per-GB charges. Kafka keeps whatever you configure; with tiered storage, older segments move to object storage, so 30–90 days of replay at volume costs roughly object-storage prices plus a small local hot tier.

Kafka also offers log compaction — keep only the latest record per key, forever — which turns a topic into a replicated table (current state of every account, every product). Kinesis has no equivalent.

Processing and Ecosystem#

Kinesis plugs into AWS directly: Lambda event source mappings, Firehose to S3/Redshift/OpenSearch, Managed Service for Apache Flink, and the Kinesis Client Library (KCL), which stores leases and checkpoints in a DynamoDB table. For an AWS-only pipeline this is fewer moving parts than anything else.

Kafka's ecosystem is broader and portable: Kafka Connect for databases, warehouses and SaaS systems; Debezium for change data capture from Postgres and MySQL; Kafka Streams for in-app stream processing; transactions for exactly-once read-process-write between topics; schema registries with compatibility checks; and, since Kafka 4.2 (February 2026), production-ready share groups for queue-style consumption of the same topics.

Resharding vs Repartitioning#

Both make "more parallelism" an ordering event. Kinesis splits and merges shards: a split creates two child shards, and consumers must finish the parent before reading children to keep per-key order (KCL handles this). In on-demand mode this happens automatically. Kafka can only add partitions to a topic, which changes the key → partition mapping for new records; per-key ordering across the change is not preserved, so keyed topics are sized up front at 2–3× projected peak.


Where Each One Breaks#

Kafka#

FailureWhat happensDetectionOwner
Under-sized cluster at peakBroker network saturates; produce latency climbs; consumers lagBroker network in/out vs limit, request queue timePlatform
Consumer lag spiralA slow consumer reads from disk, evicting page cache for othersLag in seconds per groupConsuming team
Rebalance stormProcessing exceeds max.poll.interval.ms; partitions bounce between consumersRebalance rateConsuming team
Partition count wrongToo few: can't add consumers; too many: metadata and recovery overheadPartitions per broker, consumer idle timePlatform + producer
Cross-AZ data transfer billSelf-managed on EC2: replication and consumer traffic across AZs dominates costMonthly data transfer linePlatform + FinOps

Kinesis#

FailureWhat happensDetectionOwner
Hot shardOne partition key exceeds 1 MB/s; puts throttle for that key while the stream is mostly idleWriteProvisionedThroughputExceeded, per-shard metrics (enhanced monitoring)App team
Read budget exhaustedA third or fourth standard consumer pushes the shard past 5 calls/s; everyone's iterator age growsReadProvisionedThroughputExceeded, GetRecords.IteratorAgeMillisecondsApp teams sharing the stream
Data expires before it's readA consumer falls more than the retention period behind (24 h default); records are goneIterator age approaching retentionConsuming team
Lambda poison batchOne bad record fails the batch; Lambda retries the shard until the record expires, blocking itIterator age on one shard; error countConsuming team (bisect on error, on-failure destination)
Resharding surpriseConsumers not using KCL read child shards before parents finish; per-key order breaksOut-of-order sequence checksConsuming team

The production surprise for Kinesis: costs and limits are per shard, so one hot key forces you to pay for many shards. For Kafka: the bill is mostly idle headroom and replication traffic, not messages.

The Same Spike on Both#

Traffic triples for an hour during a promotion: 20 MB/s → 60 MB/s.

                   Kafka (MSK provisioned, sized for 2× peak)       Kinesis (provisioned, 25 shards)
t=0                brokers at 35% network                            25 shards at 80% write
t=+2min            brokers at ~100% → produce latency up, no errors  60 MB/s > 25 MB/s → most puts throttled
t=+5min            consumers lag 2 minutes, catch up later          producers retry with backoff; some drop data
                                                                     after retries are exhausted
action             nothing, or add brokers next week                 UpdateShardCount to 75 (minutes; limits on
                                                                     how often) — or be on on-demand mode
lesson             headroom costs money every hour                  provisioned shards are a hard ceiling;
                                                                     on-demand absorbs it at a per-GB price

Cost and Operations#

Kafka (managed, e.g. MSK)Kinesis Data Streams
Who runs itPlatform team: sizing, topic and partition policy, quotas, upgrades, schemasAWS; app team owns shard counts (provisioned), consumers and alarms
Bill scales withBroker-hours (fixed, sized for peak) + storage (+ tiered storage)Shard-hours + PUT payload units (25 KB each) in provisioned mode; GB in/out + per-stream hours in on-demand; extras for retention and enhanced fan-out
Idle costFull cluster 24/7Provisioned: shard-hours 24/7; on-demand: per-stream hourly charge
Where it gets cheapHigh sustained throughput, many consumers, long retention via tiered storageLow or spiky throughput, few consumers, short retention
People0.5–2 engineers for a multi-team backbone, even managedNear zero
What on-call watchesKafkaKinesis
The SLO metricConsumer lag in seconds per groupGetRecords.IteratorAgeMilliseconds per consumer
The "act now" alertUnder-replicated partitions, broker disk > 70%, produce error rateWrite/read throughput exceeded sustained; iterator age > 50% of retention
Routine workPartition planning, broker scaling, upgrades, quota and ACL changesShard count (provisioned), consumer registration, retention settings

Shape of the bill at 50 MB/s in, 1 KB records, 3 consumers, 7-day retention:

Kinesis provisioned:  ≥ 50 shards for writes (1 MB/s each), ~50K PUT units/s
                      3 consumers × 50 MB/s = 150 MB/s out > 100 MB/s shared → enhanced fan-out for ≥ 1 consumer
                      + extended retention to 7 days on every shard
                      → every line scales linearly with traffic and consumer count
Managed Kafka:        a handful of brokers sized for ~150 MB/s in (3× headroom) and ~150 MB/s out per replica set
                      + 7 days × 50 MB/s ≈ 30 TB per copy (× 3 with RF=3) on local disk — or a small local tier + tiered storage
                      → mostly flat until you add brokers; a 4th consumer group costs network, not a contract

Price both at your real numbers with the current calculator: below ~10 MB/s Kinesis is almost always cheaper once you count people; above ~100 MB/s with several consumers, Kafka almost always is.

🧭 Principal Insight: "Kinesis prices the stream; Kafka prices the platform. One team with one stream should rent the stream. A company with fifty streams and twenty consuming teams should own the platform — the per-shard, per-consumer charges add up to a Kafka team long before anyone notices."


Switching Later#

MoveHowWhat's hard
Kinesis → KafkaBridge (a consumer that republishes, or a connector) while producers move; consumers switch one by onePartition keys and counts must be re-planned; KCL checkpoint logic and Lambda triggers replaced
Kafka → KinesisProducers dual-write; consumers move to KCL or Lambda; retire topicsConsumers relying on compaction, transactions, Connect or long retention need new homes
Provisioned ↔ on-demand (Kinesis)API call; allowed twice per 24 hours per stream (as of 2026)Mostly a cost decision
Self-managed Kafka ↔ managed KafkaMirror topics (MirrorMaker 2 or vendor tools), move consumers, then producersOffset translation for consumer groups; client auth changes

Whichever direction you move, run both systems for at least one full retention period so any consumer can fall back and replay from the old one.

One-way doors: partition/shard key choice (every consumer depends on its ordering scope); the event schema once many teams consume it; which system holds the company's long-term replay history. Two-way doors: Kinesis capacity mode; Kafka broker size and count; retention periods.

Diagram: Switching Later

How Real Companies Chose#

Atlassian: Kinesis to Kafka at 145 Billion Events a Day#

Atlassian moved its StreamHub event platform from Kinesis to Amazon MSK as volume grew from 22 billion to about 145 billion events a day ingested (peaks of 3.2 million events per second). Their reasons: Kinesis throttling during spikes hurt producer reliability (they targeted 99.999%), the bill grew linearly with shard count — needing thousands of shards at peak — and retention beyond 24 hours and independent consumer groups came at premium rates. On Kafka they keep about 5 minutes on local disk and 7 days in tiered storage on S3 (Atlassian Engineering).

Staff insight: Every reason maps to a per-shard contract — write limits, per-shard pricing, retention and consumer fan-out priced per shard. At their scale the shard model stopped being a convenience and became the cost driver.

Airties: Kafka to Kinesis to Stop Running Clusters#

In an AWS-published case study, Airties describes moving from dozens of fixed-size, single-AZ Kafka clusters (~25 TB of records each, data from tens of millions of access points) to Kinesis Data Streams. The pain was operational: manual upgrades, partition rebalancing, load testing and clusters that were oversized for small deployments and undersized for large ones. They report a 33% reduction in monthly infrastructure cost after moving to Kinesis with automatic shard scaling (AWS Big Data Blog).

Staff insight: The opposite move for the opposite reason: many small, uneven workloads where fixed clusters wasted money and people. Right-sizing per stream is exactly what Kinesis's model does well. (Vendor-published; the decision shape is the lesson.)

Amazon: Kinesis at Prime Day Scale#

AWS reports Kinesis Data Streams processed a peak of 807 million records per second during Prime Day 2025 (AWS News Blog). That figure covers Amazon's Prime Day workloads, which run on many streams rather than one — which is the point: Kinesis scales by giving each workload its own independently sized stream.

Staff insight: "Can Kinesis handle our scale?" is rarely the question. "Does our consumer pattern fit per-shard read limits, and what does it cost?" is.


Follow-Ups to Expect#

After You Say...They Will Ask...What They're Testing
"Kinesis""Five more teams want this stream next quarter."Shared read limits, enhanced fan-out costs and caps
"Kinesis, partition key = user_id""One user generates 3 MB/s."Per-shard 1 MB/s write limit; key splitting
"Kafka on MSK""Who sizes brokers and what happens at 3× traffic?"Ownership; headroom vs cost; the difference from on-demand
"Replay last week""Retention is 24 hours."Kinesis retention defaults and cost; Kafka tiered storage
"Exactly-once""On which system, and to where?"Kafka transactions for Kafka → Kafka only; idempotent sinks elsewhere
"Lambda consumer""One record always fails."Shard blocking, bisect-on-error, on-failure destinations
"We'll migrate later""What's hard about moving from Kinesis to Kafka?"Re-keying, checkpoint translation, consumer-by-consumer cutover
"On-demand Kinesis""Traffic is steady at 80 MB/s all month. Still on-demand?"Knowing on-demand trades price for elasticity; provisioned or Kafka for steady load

What to Say in the Interview#

"Kinesis and Kafka share the model — a keyed, ordered, replayable log — so I'll design keys and consumers the same way. The choice is about scale shape and who operates it."

"For this clickstream to S3 with one Lambda, Kinesis on-demand. If more than four or five teams end up reading it, the shared 2 MB/s read budget and per-consumer fan-out charges are my signal to move it to Kafka."

"For the order-event backbone that fulfillment, search and fraud all consume, I'd start on managed Kafka: consumer groups are free to add, tiered storage gives us weeks of replay, and Debezium can feed it straight from Postgres."

"Either way, I'll check that no single key exceeds a shard's or partition's comfortable rate, because neither system makes one hot key faster."


  1. Loading the index…