On AWS this is the log decision. Both are partitioned, ordered, replayable streams: producers write records with a key, the key picks a partition (Kafka) or shard (Kinesis), order holds within it, and consumers track their own position. Underneath the shared model they make opposite bets. Kinesis Data Streams is a serverless AWS service with fixed per-shard limits and per-shard (or per-GB) pricing: nothing to operate, everything metered. Kafka is an open-source system you run or buy as a managed service (Amazon MSK and others): brokers to size, but no per-shard read limits, a large ecosystem, and economics that improve with scale. The question that decides it: will this stream stay small and AWS-native, with a handful of consumers — or will it become the event backbone that many teams read, replay and process?
The Verdict#
Use Kinesis for a single AWS-native pipeline at modest throughput with a few consumers. Use managed Kafka for a company-wide event backbone, for high sustained throughput, for many consumer groups, or when you need the Kafka ecosystem or portability.
| Pick Kinesis when | Pick Kafka when |
|---|---|
| One pipeline: clickstream → Lambda → S3, IoT telemetry → analytics | A shared event backbone that 5–50 teams will consume |
| Throughput is up to ~tens of MB/s, or very spiky (on-demand mode) | Sustained throughput is 100s of MB/s, where broker economics beat per-shard pricing |
| 1–5 consumers per stream | Many independent consumer groups reading the same data |
| You want zero operations and IAM-native everything | You need Kafka Connect, CDC (Debezium), Kafka Streams, or the wider client ecosystem |
| Retention of 24 hours to 7 days covers replay needs | You want long retention cheaply (tiered storage to object storage) or compacted topics as tables |
| You are certain to stay on AWS | Multi-cloud, on-prem, or avoiding an AWS-only API matters |
🎯 Staff Move: "If this is one pipeline from our app to S3 with a Lambda in the middle, Kinesis on-demand — no brokers, done this afternoon. If it's the stream that fulfillment, search, fraud and analytics will all read for years, I'd put it on managed Kafka now, because moving a backbone later means moving every consumer."
At a Glance#
| Dimension | Kafka | Kinesis Data Streams |
|---|---|---|
| Data model | Topics → partitions → offset-ordered records; optional log compaction | Streams → shards → sequence-numbered records |
| Ordering | Per partition | Per shard |
| Write limit per unit | No fixed per-partition cap; bounded by broker disk and network (often ~10 MB/s per partition is a comfortable planning number) | 1 MB/s or 1,000 records/s per shard (as of 2026) |
| Read limit per unit | Bounded by broker network; each consumer group reads independently | 2 MB/s per shard shared by all standard consumers (5 GetRecords calls/s); 2 MB/s per consumer per shard with enhanced fan-out |
| Consumers per stream | Effectively unlimited consumer groups | Up to 20 enhanced fan-out consumers per stream (50 with On-demand Advantage), as of 2026 |
| Max record size | 1 MB default (configurable) | Up to 10 MiB, with large records drawing on burst capacity (as of 2026) |
| Retention | Any (7 days default); tiered storage to object storage for months or years; compaction keeps latest per key | 24 hours default; up to 365 days at extra cost |
| Delivery semantics | At-least-once; idempotent producer; transactions for exactly-once Kafka → Kafka | At-least-once; producer retries can duplicate; no transactions |
| Latency | ~5–20 ms produce with acks=all; consumer fetch adds little | ~tens of ms put; standard polling consumers see ~200 ms+ propagation; enhanced fan-out pushes in ~70 ms |
| Scaling model | Add partitions (remaps keys) and brokers (reassign partitions) | On-demand: automatic; provisioned: split/merge shards (UpdateShardCount) |
| Operational burden | Medium on managed (sizing, partitions, upgrades, quotas); high self-hosted | Very low (shard counts in provisioned mode) |
| Managed options | Amazon MSK (provisioned, Express, serverless), Confluent Cloud, Aiven, Redpanda and others | AWS only |
| Cost shape | Brokers × hours + storage + cross-AZ replication traffic (on self-managed EC2) | Provisioned: shard-hours + PUT units; on-demand: per GB in/out + per-stream hour; extras for retention and fan-out |
| Ecosystem | Connect (hundreds of connectors), Debezium CDC, Kafka Streams, Flink, schema registries | Lambda, Firehose, Managed Service for Apache Flink, KCL; AWS integrations |
Feature checklist for the decisions interviewers probe (as of 2026):
| Capability | Kafka | Kinesis |
|---|---|---|
| Replay from a timestamp | Yes (offsetsForTimes) | Yes (AT_TIMESTAMP iterator) |
| Keep latest value per key forever | Yes (compacted topics) | No |
| Exactly-once read-process-write | Yes, between Kafka topics (transactions) | No; idempotent consumers |
| Queue-style per-record acks | Yes, share groups (production-ready in 4.2) | No |
| Change data capture from databases | Debezium / Connect | Via DMS or custom; DynamoDB can stream into Kinesis natively |
| Serverless consumer | Lambda event source for MSK and self-managed Kafka | Lambda event source, native |
| Runs outside AWS | Yes | No |
How They Actually Differ#
Fixed Shard Limits vs Broker Capacity#
A Kinesis shard is a contract: 1 MB/s or 1,000 records/s in, 2 MB/s out, 5 read calls per second. You size the stream by dividing your peak by those numbers, and you exceed them with ProvisionedThroughputExceededException. On-demand mode does the division for you (new streams start at 4 MB/s write and scale up to 10 GB/s in the largest regions, 200 MB/s elsewhere by default, as of 2026), but the per-shard ceiling still applies to any one partition key.
A Kafka partition has no fixed ceiling; it is bounded by the broker's disk, network and the other partitions sharing it. That is more headroom per key and more responsibility: the platform team decides how many partitions each broker carries and when to add brokers.
| Planning question | Kafka | Kinesis |
|---|---|---|
| 50 MB/s in, 1 KB records | ~12–24 partitions on a few brokers (plan ~3× headroom) | ≥ 50 shards (provisioned) or on-demand |
| One key at 3 MB/s | Fits in one partition | Exceeds one shard — must split the key |
| 3 consumers reading everything | 3 consumer groups, no extra config | 150 MB/s out > 100 MB/s shared read → enhanced fan-out or more shards |
Who pays: Kinesis makes the application team pay when one key or one consumer outgrows a shard. Kafka makes the platform team pay to keep brokers sized.
Fan-out: Who Shares the Read Budget#
Every standard Kinesis consumer on a shard shares 2 MB/s and 5 GetRecords calls per second. Two consumers polling once a second is fine; five consumers polling at the 200 ms each needs leaves no room. Enhanced fan-out gives each registered consumer its own 2 MB/s per shard pushed over HTTP/2, at an extra charge, up to 20 consumers per stream (50 with On-demand Advantage).
Kafka consumer groups each read at their own pace from the broker; adding the tenth consumer group costs broker network egress, not a per-consumer contract. This is the single biggest reason streams that start on Kinesis move to Kafka when more teams want the data.
🎯 Staff Insight: Count consumers before you count megabytes. A 5 MB/s stream with eight consuming teams is a worse fit for Kinesis than a 50 MB/s stream with one, because the shared read budget and per-consumer fan-out charges scale with readers, not writers.
The producer side of "durable, ordered per order, no duplicates from retries" in each:
Kafka producer
acks = all # leader waits for in-sync replicas
enable.idempotence = true # broker drops duplicate retries per partition
min.insync.replicas = 2 # (topic config) with RF = 3
key = order_id # same key → same partition → ordered
linger.ms = 5, batch.size = 64KB, compression.type = zstd
Kinesis producer
PutRecords (≤ 500 records/request), PartitionKey = order_id
SequenceNumberForOrdering on PutRecord when strict per-key order across retries matters
on partial failure: retry only the failed records, with backoff
no broker-side dedupe → consumers dedupe by an event ID carried in the payload
aggregation (KPL) packs small records into one to save PUT units and the 1,000 records/s limit
The difference that matters in review: Kafka can drop producer-retry duplicates at the broker; Kinesis cannot, so every Kinesis consumer needs an idempotency key from day one.
Retention and Replay#
Both replay by position: Kafka by offset or timestamp, Kinesis by sequence number or timestamp (AT_TIMESTAMP iterator). The difference is how far back and at what cost. Kinesis keeps 24 hours by default; extending to 7 days and then up to 365 days adds per-shard-hour and per-GB charges. Kafka keeps whatever you configure; with tiered storage, older segments move to object storage, so 30–90 days of replay at volume costs roughly object-storage prices plus a small local hot tier.
Kafka also offers log compaction — keep only the latest record per key, forever — which turns a topic into a replicated table (current state of every account, every product). Kinesis has no equivalent.
Processing and Ecosystem#
Kinesis plugs into AWS directly: Lambda event source mappings, Firehose to S3/Redshift/OpenSearch, Managed Service for Apache Flink, and the Kinesis Client Library (KCL), which stores leases and checkpoints in a DynamoDB table. For an AWS-only pipeline this is fewer moving parts than anything else.
Kafka's ecosystem is broader and portable: Kafka Connect for databases, warehouses and SaaS systems; Debezium for change data capture from Postgres and MySQL; Kafka Streams for in-app stream processing; transactions for exactly-once read-process-write between topics; schema registries with compatibility checks; and, since Kafka 4.2 (February 2026), production-ready share groups for queue-style consumption of the same topics.
Resharding vs Repartitioning#
Both make "more parallelism" an ordering event. Kinesis splits and merges shards: a split creates two child shards, and consumers must finish the parent before reading children to keep per-key order (KCL handles this). In on-demand mode this happens automatically. Kafka can only add partitions to a topic, which changes the key → partition mapping for new records; per-key ordering across the change is not preserved, so keyed topics are sized up front at 2–3× projected peak.
Where Each One Breaks#
Kafka#
| Failure | What happens | Detection | Owner |
|---|---|---|---|
| Under-sized cluster at peak | Broker network saturates; produce latency climbs; consumers lag | Broker network in/out vs limit, request queue time | Platform |
| Consumer lag spiral | A slow consumer reads from disk, evicting page cache for others | Lag in seconds per group | Consuming team |
| Rebalance storm | Processing exceeds max.poll.interval.ms; partitions bounce between consumers | Rebalance rate | Consuming team |
| Partition count wrong | Too few: can't add consumers; too many: metadata and recovery overhead | Partitions per broker, consumer idle time | Platform + producer |
| Cross-AZ data transfer bill | Self-managed on EC2: replication and consumer traffic across AZs dominates cost | Monthly data transfer line | Platform + FinOps |
Kinesis#
| Failure | What happens | Detection | Owner |
|---|---|---|---|
| Hot shard | One partition key exceeds 1 MB/s; puts throttle for that key while the stream is mostly idle | WriteProvisionedThroughputExceeded, per-shard metrics (enhanced monitoring) | App team |
| Read budget exhausted | A third or fourth standard consumer pushes the shard past 5 calls/s; everyone's iterator age grows | ReadProvisionedThroughputExceeded, GetRecords.IteratorAgeMilliseconds | App teams sharing the stream |
| Data expires before it's read | A consumer falls more than the retention period behind (24 h default); records are gone | Iterator age approaching retention | Consuming team |
| Lambda poison batch | One bad record fails the batch; Lambda retries the shard until the record expires, blocking it | Iterator age on one shard; error count | Consuming team (bisect on error, on-failure destination) |
| Resharding surprise | Consumers not using KCL read child shards before parents finish; per-key order breaks | Out-of-order sequence checks | Consuming team |
The production surprise for Kinesis: costs and limits are per shard, so one hot key forces you to pay for many shards. For Kafka: the bill is mostly idle headroom and replication traffic, not messages.
The Same Spike on Both#
Traffic triples for an hour during a promotion: 20 MB/s → 60 MB/s.
Kafka (MSK provisioned, sized for 2× peak) Kinesis (provisioned, 25 shards)
t=0 brokers at 35% network 25 shards at 80% write
t=+2min brokers at ~100% → produce latency up, no errors 60 MB/s > 25 MB/s → most puts throttled
t=+5min consumers lag 2 minutes, catch up later producers retry with backoff; some drop data
after retries are exhausted
action nothing, or add brokers next week UpdateShardCount to 75 (minutes; limits on
how often) — or be on on-demand mode
lesson headroom costs money every hour provisioned shards are a hard ceiling;
on-demand absorbs it at a per-GB price
Cost and Operations#
| Kafka (managed, e.g. MSK) | Kinesis Data Streams | |
|---|---|---|
| Who runs it | Platform team: sizing, topic and partition policy, quotas, upgrades, schemas | AWS; app team owns shard counts (provisioned), consumers and alarms |
| Bill scales with | Broker-hours (fixed, sized for peak) + storage (+ tiered storage) | Shard-hours + PUT payload units (25 KB each) in provisioned mode; GB in/out + per-stream hours in on-demand; extras for retention and enhanced fan-out |
| Idle cost | Full cluster 24/7 | Provisioned: shard-hours 24/7; on-demand: per-stream hourly charge |
| Where it gets cheap | High sustained throughput, many consumers, long retention via tiered storage | Low or spiky throughput, few consumers, short retention |
| People | 0.5–2 engineers for a multi-team backbone, even managed | Near zero |
| What on-call watches | Kafka | Kinesis |
|---|---|---|
| The SLO metric | Consumer lag in seconds per group | GetRecords.IteratorAgeMilliseconds per consumer |
| The "act now" alert | Under-replicated partitions, broker disk > 70%, produce error rate | Write/read throughput exceeded sustained; iterator age > 50% of retention |
| Routine work | Partition planning, broker scaling, upgrades, quota and ACL changes | Shard count (provisioned), consumer registration, retention settings |
Shape of the bill at 50 MB/s in, 1 KB records, 3 consumers, 7-day retention:
Kinesis provisioned: ≥ 50 shards for writes (1 MB/s each), ~50K PUT units/s
3 consumers × 50 MB/s = 150 MB/s out > 100 MB/s shared → enhanced fan-out for ≥ 1 consumer
+ extended retention to 7 days on every shard
→ every line scales linearly with traffic and consumer count
Managed Kafka: a handful of brokers sized for ~150 MB/s in (3× headroom) and ~150 MB/s out per replica set
+ 7 days × 50 MB/s ≈ 30 TB per copy (× 3 with RF=3) on local disk — or a small local tier + tiered storage
→ mostly flat until you add brokers; a 4th consumer group costs network, not a contract
Price both at your real numbers with the current calculator: below ~10 MB/s Kinesis is almost always cheaper once you count people; above ~100 MB/s with several consumers, Kafka almost always is.
🧭 Principal Insight: "Kinesis prices the stream; Kafka prices the platform. One team with one stream should rent the stream. A company with fifty streams and twenty consuming teams should own the platform — the per-shard, per-consumer charges add up to a Kafka team long before anyone notices."
Switching Later#
| Move | How | What's hard |
|---|---|---|
| Kinesis → Kafka | Bridge (a consumer that republishes, or a connector) while producers move; consumers switch one by one | Partition keys and counts must be re-planned; KCL checkpoint logic and Lambda triggers replaced |
| Kafka → Kinesis | Producers dual-write; consumers move to KCL or Lambda; retire topics | Consumers relying on compaction, transactions, Connect or long retention need new homes |
| Provisioned ↔ on-demand (Kinesis) | API call; allowed twice per 24 hours per stream (as of 2026) | Mostly a cost decision |
| Self-managed Kafka ↔ managed Kafka | Mirror topics (MirrorMaker 2 or vendor tools), move consumers, then producers | Offset translation for consumer groups; client auth changes |
Whichever direction you move, run both systems for at least one full retention period so any consumer can fall back and replay from the old one.
One-way doors: partition/shard key choice (every consumer depends on its ordering scope); the event schema once many teams consume it; which system holds the company's long-term replay history. Two-way doors: Kinesis capacity mode; Kafka broker size and count; retention periods.
How Real Companies Chose#
Atlassian: Kinesis to Kafka at 145 Billion Events a Day#
Atlassian moved its StreamHub event platform from Kinesis to Amazon MSK as volume grew from 22 billion to about 145 billion events a day ingested (peaks of 3.2 million events per second). Their reasons: Kinesis throttling during spikes hurt producer reliability (they targeted 99.999%), the bill grew linearly with shard count — needing thousands of shards at peak — and retention beyond 24 hours and independent consumer groups came at premium rates. On Kafka they keep about 5 minutes on local disk and 7 days in tiered storage on S3 (Atlassian Engineering).
Staff insight: Every reason maps to a per-shard contract — write limits, per-shard pricing, retention and consumer fan-out priced per shard. At their scale the shard model stopped being a convenience and became the cost driver.
Airties: Kafka to Kinesis to Stop Running Clusters#
In an AWS-published case study, Airties describes moving from dozens of fixed-size, single-AZ Kafka clusters (~25 TB of records each, data from tens of millions of access points) to Kinesis Data Streams. The pain was operational: manual upgrades, partition rebalancing, load testing and clusters that were oversized for small deployments and undersized for large ones. They report a 33% reduction in monthly infrastructure cost after moving to Kinesis with automatic shard scaling (AWS Big Data Blog).
Staff insight: The opposite move for the opposite reason: many small, uneven workloads where fixed clusters wasted money and people. Right-sizing per stream is exactly what Kinesis's model does well. (Vendor-published; the decision shape is the lesson.)
Amazon: Kinesis at Prime Day Scale#
AWS reports Kinesis Data Streams processed a peak of 807 million records per second during Prime Day 2025 (AWS News Blog). That figure covers Amazon's Prime Day workloads, which run on many streams rather than one — which is the point: Kinesis scales by giving each workload its own independently sized stream.
Staff insight: "Can Kinesis handle our scale?" is rarely the question. "Does our consumer pattern fit per-shard read limits, and what does it cost?" is.
Follow-Ups to Expect#
| After You Say... | They Will Ask... | What They're Testing |
|---|---|---|
| "Kinesis" | "Five more teams want this stream next quarter." | Shared read limits, enhanced fan-out costs and caps |
"Kinesis, partition key = user_id" | "One user generates 3 MB/s." | Per-shard 1 MB/s write limit; key splitting |
| "Kafka on MSK" | "Who sizes brokers and what happens at 3× traffic?" | Ownership; headroom vs cost; the difference from on-demand |
| "Replay last week" | "Retention is 24 hours." | Kinesis retention defaults and cost; Kafka tiered storage |
| "Exactly-once" | "On which system, and to where?" | Kafka transactions for Kafka → Kafka only; idempotent sinks elsewhere |
| "Lambda consumer" | "One record always fails." | Shard blocking, bisect-on-error, on-failure destinations |
| "We'll migrate later" | "What's hard about moving from Kinesis to Kafka?" | Re-keying, checkpoint translation, consumer-by-consumer cutover |
| "On-demand Kinesis" | "Traffic is steady at 80 MB/s all month. Still on-demand?" | Knowing on-demand trades price for elasticity; provisioned or Kafka for steady load |
What to Say in the Interview#
"Kinesis and Kafka share the model — a keyed, ordered, replayable log — so I'll design keys and consumers the same way. The choice is about scale shape and who operates it."
"For this clickstream to S3 with one Lambda, Kinesis on-demand. If more than four or five teams end up reading it, the shared 2 MB/s read budget and per-consumer fan-out charges are my signal to move it to Kafka."
"For the order-event backbone that fulfillment, search and fraud all consume, I'd start on managed Kafka: consumer groups are free to add, tiered storage gives us weeks of replay, and Debezium can feed it straight from Postgres."
"Either way, I'll check that no single key exceeds a shard's or partition's comfortable rate, because neither system makes one hot key faster."
Related Guides#
- Apache Kafka — partitions, replication, consumer groups and operations in depth
- Kafka vs SQS vs RabbitMQ — when you need a queue, not a log
- Batch & Stream Pipelines — where the log fits in a data platform
- Apache Flink — stream processing on top of either
- Design a Real-Time Ad Click Aggregator — a classic high-throughput stream workload
- Design a Metrics and Alerting Platform — telemetry ingestion at scale
- Design a Stream Processing Engine — what sits behind the consumers
- Partitioning — choosing keys and living with hot spots