// OSSeva Blog
OperationsKafka vs Pulsar: Architecture, Storage, Geo-Replication and When to Choose Each
The short answer
Kafka stores data on the brokers that serve it; Pulsar separates the two. A Kafka broker owns partitions on its own disks, so scaling or replacing brokers means moving data. Pulsar brokers are stateless, and messages live in Apache BookKeeper on separate storage nodes called bookies, with a metadata store such as Oxia or ZooKeeper alongside. Pulsar also builds in multi-tenancy, geo-replication and queue-style subscriptions. Kafka brings Kafka Connect, Kafka Streams and the tools and managed services built for its protocol, a simpler deployment since KRaft removed ZooKeeper, and now queue-style consumption through share groups. Choose Kafka when ecosystem, tooling and team familiarity matter most. Choose Pulsar when you need many isolated tenants, built-in replication between regions, or independent scaling of serving and storage.
Kafka vs Pulsar at a glance
| Apache Kafka | Apache Pulsar | |
|---|---|---|
| Storage | Partition logs on broker disks | Apache BookKeeper ledgers on bookies |
| Brokers | Stateful; own their partitions | Stateless; serve topics assigned by load |
| Metadata | KRaft controllers (ZooKeeper removed in 4.0) | Oxia, ZooKeeper or another supported store |
| Components to run | Brokers and controllers | Brokers, bookies and a metadata store; a configuration store for multi-cluster |
| Tiered storage | Completed segments to object storage; not for compacted topics | Sealed ledgers offloaded to long-term storage |
| Cross-region replication | MirrorMaker 2, a separate Kafka Connect workload | Built into the broker; asynchronous or synchronous |
| Multi-tenancy | Topic naming conventions, prefixed ACLs and quotas | Tenants and namespaces with their own auth, quotas and policies |
| Consumption | Consumer groups; share groups production-ready from 4.2 | Exclusive, Failover, Shared and Key_Shared subscriptions |
| Ecosystem | Kafka Connect, Kafka Streams | Pulsar Functions, Pulsar IO |
| Licence | Apache 2.0 | Apache 2.0 |
Kafka vs Pulsar architecture
In Kafka, a topic is split into partitions, and each partition is a log stored on a broker and replicated to others. The broker that leads a partition handles its reads and writes. Kafka's operations guide is direct about the consequence: a new broker is not given existing partitions automatically, so adding capacity means reassigning partitions and copying their data. Since 4.0, Kafka runs only in KRaft mode, with metadata held by Kafka's own controllers rather than ZooKeeper.
In Pulsar, the broker runs the API and dispatch layer, and BookKeeper, a distributed write-ahead log, stores messages in ledgers spread across bookies. A topic is a sequence of ledgers over time rather than one log on one machine. Because brokers hold no message data, Pulsar assigns topics to brokers in groups called bundles, based on load, and moves them between brokers without copying the stored messages. Coordination uses a metadata store: Oxia, ZooKeeper or another supported backend. The trade-off is more moving parts. A Pulsar cluster runs three kinds of service where Kafka runs one or two, and each has its own configuration, monitoring and upgrade path.
Tiered storage
Both projects can keep recent data on local disks and move older data to object storage. Kafka's tiered storage (KIP-405) uploads completed log segments to a remote tier such as S3 or HDFS; its documented limitations include no support for compacted topics. Pulsar's tiered storage offloads ledgers once they are sealed and immutable, and readers fetch offloaded data transparently. In both, tiered storage changes the cost of long retention more than it changes the day-to-day write path.
Geo-replication
This is one of the clearest differences. Kafka replicates between clusters with MirrorMaker 2, a separate process built on Kafka Connect that copies topics, configurations, consumer group offsets and ACLs from source to target clusters. It works well, but it is another workload to deploy and monitor. Pulsar has geo-replication in the broker, with namespaces assigned to the clusters they replicate to, and supports both asynchronous replication between independent clusters and synchronous replication across data centres. An instance-wide configuration store coordinates clusters that replicate with each other.
Multi-tenancy and subscriptions
Pulsar was designed as a multi-tenant system. Tenants can each have their own authentication and authorisation, storage quotas, message TTL and isolation policies, with namespaces inside each tenant. Kafka's documentation recommends building tenant spaces from hierarchical topic names, enforced with prefixed ACLs and isolated with quotas. Both work. Pulsar's model is built in; Kafka's is a convention you enforce.
On consumption, Pulsar offers four subscription types. Exclusive and Failover give ordered, single-consumer reading; Shared spreads messages across consumers with individual acknowledgement; Key_Shared spreads them while keeping each key on one consumer. Kafka's consumer groups give ordered reading per partition, and share groups, production-ready in 4.2, add cooperative consumption with per-record acknowledgement.
When to choose Kafka, and when to choose Pulsar
- Choose Kafka if you rely on Kafka Connect, Kafka Streams or tools and managed services built for the Kafka protocol, if your team already runs it, or if a single region and a modest number of teams cover your needs.
- Choose Pulsar if you run a shared platform for many teams that need strong isolation, if replication between regions is a core requirement, or if you want to scale storage and serving independently.
- Stay where you are if your current platform meets its service levels. Moving between them means new clients, new operations and a migration of retained data. Most of the gaps people move to close, such as queue semantics in Kafka or a ZooKeeper-free deployment in either, have narrowed.
Where OSSeva fits
OSSeva supports both under one contract. OSSeva for Apache Kafka patches 2.8 through 3.9, including the archived 3.8 and 3.9 lines, and covers 4.x in KRaft mode, with Kafka Connect and Streams in scope on Assure and Operate. OSSeva for Apache Pulsar patches 2.10, 2.11, 3.0, 3.1 and 3.2 after community support ends, covers 4.0, and patches the ZooKeeper under older Pulsar clusters as part of the same engagement. Both run where you run them: bare metal, VMs or any Kubernetes. Priced per cluster; book a discovery call for a quote. For other options, see Pulsar support providers and Kafka support providers, and for running Pulsar without a vendor platform, Pulsar production support.
Frequently asked questions
What is the main architecture difference between Kafka and Pulsar?
Kafka brokers store the partitions they serve. Pulsar brokers are stateless, and storage sits in a separate Apache BookKeeper cluster, with a metadata store alongside. Kafka has fewer components to run; Pulsar can scale and replace brokers without moving stored data.
Kafka vs Pulsar latency: which is lower?
It depends on configuration more than on the project. One documented difference matters: Kafka's default is not to fsync each write, relying on replication and the operating system for durability, while BookKeeper bookies fsync their journal before acknowledging a write by default. Both can be tuned the other way. Measure end-to-end latency with your own durability settings, message sizes and client configuration.
Kafka vs Pulsar performance: which is faster?
Neither in general. When you read a published comparison, check who ran it and how each system was configured. Kafka's design favours sequential writes and batched reads on broker disks; Pulsar's adds a network hop to BookKeeper in exchange for separating storage from serving. Benchmark your own workload on both before deciding.
Does Pulsar still need ZooKeeper?
Not from Pulsar 3.3. From that release Pulsar can run on the Oxia metadata store instead of ZooKeeper; 3.0, 3.1, 3.2 and the 2.x lines need ZooKeeper. Kafka removed ZooKeeper entirely in 4.0.
Is Pulsar better than Kafka for multi-tenant platforms?
Pulsar's tenants and namespaces are built in, with per-tenant authentication, quotas and policies. Kafka can be run multi-tenant with naming conventions, prefixed ACLs and quotas, but you enforce the model yourself.
Can Kafka do queues like Pulsar's Shared subscription?
Yes, from Kafka 4.2. Share groups let consumers cooperatively read a topic, outnumber its partitions and acknowledge records individually, which covers much of what Pulsar's Shared subscription does.
Tags
Ready to get your open source under control?
Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.