← All posts
2026-06-19·8 min read·system design

The secret to Kafka's speed: why is Apache Kafka so fast?

Two design choices do most of the heavy lifting — Sequential I/O and the Zero-Copy principle. Here's how each one works, and why together they turn a disk into a firehose.

When the conversation turns to distributed systems, event-driven architectures, or real-time streaming, one name keeps coming up: Apache Kafka. It's become the default for high-throughput messaging — companies like LinkedIn, Netflix, and Uber lean on it to move trillions of messages a day.

But why is it fast? In this post I want to unpack the two ideas that matter most: Sequential I/O and the Zero-Copy principle. Neither is exotic. Both are really about respecting how the hardware and the operating system actually work instead of fighting them.

First, redefine "fast": throughput vs latency

"Fast" is ambiguous in system design. Do we mean low latency (how long a single message takes from producer to consumer) or high throughput (how much data moves over a window of time)? Kafka is unapologetically optimized for throughput.

Think of Kafka as a pipeline moving liquid. The goal isn't to make one drop travel faster — it's to widen the pipe so an enormous, continuous volume flows through. When people say "Kafka is fast," they mean it can move and store massive amounts of data in very little time. Two foundational choices make that possible.


Secret #1 — the power of Sequential I/O

There's a long-standing myth in software: "disk is always slow, memory is always fast." For random access that's broadly true. But it falls apart the moment you look at the access pattern.

On a spinning disk (HDD), random reads and writes force the physical head to jump around to different locations. That mechanical seek time is what makes random access sluggish. Sequential access is a completely different story: if the head can just read or write blocks one after another in a straight line, it flies.

Kafka exploits this by using an append-only log as its core data structure. When a producer sends a message, Kafka doesn't update an existing record or insert randomly — it just appends to the end of a file. A purely sequential operation.

Random access (head jumping, slow) versus sequential append-only access (one track, fast)
Random access makes the disk head jump (high seek time). Kafka's append-only log keeps it moving along one track — near-zero seek time.

The numbers speak for themselves

On modern hardware with a JBOD (Just a Bunch Of Disks) array, sequential writes can hit hundreds of megabytes per second, while random writes max out at mere hundreds of kilobytes per second. That's several orders of magnitude.

There's a cost angle too. Because sequential I/O unlocks high speeds even on plain HDDs, Kafka doesn't need expensive SSDs. HDDs cost roughly a third of the price of SSDs while offering about three times the capacity — so Kafka can retain huge volumes of messages for long periods, cheaply. That kind of cheap, durable retention was rare in enterprise messaging before Kafka.

It also leans on the OS page cache

Kafka deliberately keeps data in the operating system's page cache rather than managing a big in-process cache on the JVM heap. Reads and writes go through memory the OS already manages, sequentially, so "hot" data is served straight from RAM without Kafka copying it around itself — and without the garbage-collection pressure a giant heap cache would create.

Traditional random RAM access versus Kafka's linear page-cache access
Standard apps hit memory randomly (high latency, GC-heavy). Kafka streams linearly through the OS page cache — efficient, and it sets up the zero-copy path below.

Secret #2 — extreme efficiency via Zero-Copy

Kafka's whole job is moving data: network → disk when producers publish, and disk → network when consumers read. When you're moving terabytes, even tiny per-operation inefficiencies stack up. The usual culprit is unnecessary copying of data between the disk, system memory, and the network.

The inefficient way (without zero-copy)

Here's what a standard application does to send a file from disk over the network:

  1. Load data from disk into the OS kernel cache.
  2. Copy it up from the kernel cache into the application buffer (user space).
  3. Copy it back down into the kernel's socket buffer.
  4. Copy it from the socket buffer into the NIC buffer to actually send.

That's four copies and two context switches between user and kernel space — pure CPU overhead, and a disaster at streaming scale.

The Kafka way (with zero-copy)

Kafka skips the application layer entirely when serving data to consumers. It uses a system call like sendfile() on Linux:

  1. Data is loaded from disk into the OS cache (same as before).
  2. Kafka issues sendfile, and the OS transfers the data directly from the OS cache to the NIC buffer.

The data never enters Kafka's user-space process. And because modern network cards use DMA (Direct Memory Access) for that transfer, the main CPU isn't even involved in the copy. Kafka just points at the data; the OS and the NIC do the heavy lifting.

Traditional path with four data copies versus Kafka's single zero-copy sendfile path
Top: the traditional path — five copies through CPU, app, and socket buffers. Bottom: Kafka's zero-copy path — a single DMA-driven sendfile straight from the page cache to the network.

Putting it together

Kafka uses plenty of other tricks — aggressive message batching, smart partition distribution, compression — but Sequential I/O and Zero-Copy are the cornerstones. Writes go to an append-only log on disk in one linear sweep; reads stream straight from the page cache to the network card without ever touching the application.

End-to-end Kafka data path: sequential append-only writes and zero-copy reads working together
The two pathways end to end — sequential append-only writes (left) feeding zero-copy, DMA-driven reads (right).

Treat the disk as an append-only log, and let the operating system stream bytes directly to the network, and a workload that looks like a random-access nightmare becomes a smooth, linear, high-throughput stream. That's not just what makes Kafka "fast" — it's what makes it a sensible backbone for real-time data infrastructure.

Notes written up from the excellent ByteByteGo system-design series on Kafka internals. Diagrams are my own.