Back to blog

// OSSeva Blog

Operations

Kafka vs Pulsar: Architecture, Storage, Geo-Replication and When to Choose Each

Randall McClure6 min read

The short answer

Kafka stores data on the brokers that serve it; Pulsar separates the two. A Kafka broker owns partitions on its own disks, so scaling or replacing brokers means moving data. Pulsar brokers are stateless, and messages live in Apache BookKeeper on separate storage nodes called bookies, with a metadata store such as Oxia or ZooKeeper alongside. Pulsar also builds in multi-tenancy, geo-replication and queue-style subscriptions. Kafka brings Kafka Connect, Kafka Streams and the tools and managed services built for its protocol, a simpler deployment since KRaft removed ZooKeeper, and now queue-style consumption through share groups. Choose Kafka when ecosystem, tooling and team familiarity matter most. Choose Pulsar when you need many isolated tenants, built-in replication between regions, or independent scaling of serving and storage.

Kafka vs Pulsar at a glance

Apache KafkaApache Pulsar
StoragePartition logs on broker disksApache BookKeeper ledgers on bookies
BrokersStateful; own their partitionsStateless; serve topics assigned by load
MetadataKRaft controllers (ZooKeeper removed in 4.0)Oxia, ZooKeeper or another supported store
Components to runBrokers and controllersBrokers, bookies and a metadata store; a configuration store for multi-cluster
Tiered storageCompleted segments to object storage; not for compacted topicsSealed ledgers offloaded to long-term storage
Cross-region replicationMirrorMaker 2, a separate Kafka Connect workloadBuilt into the broker; asynchronous or synchronous
Multi-tenancyTopic naming conventions, prefixed ACLs and quotasTenants and namespaces with their own auth, quotas and policies
ConsumptionConsumer groups; share groups production-ready from 4.2Exclusive, Failover, Shared and Key_Shared subscriptions
EcosystemKafka Connect, Kafka StreamsPulsar Functions, Pulsar IO
LicenceApache 2.0Apache 2.0

Kafka vs Pulsar architecture

In Kafka, a topic is split into partitions, and each partition is a log stored on a broker and replicated to others. The broker that leads a partition handles its reads and writes. Kafka's operations guide is direct about the consequence: a new broker is not given existing partitions automatically, so adding capacity means reassigning partitions and copying their data. Since 4.0, Kafka runs only in KRaft mode, with metadata held by Kafka's own controllers rather than ZooKeeper.

In Pulsar, the broker runs the API and dispatch layer, and BookKeeper, a distributed write-ahead log, stores messages in ledgers spread across bookies. A topic is a sequence of ledgers over time rather than one log on one machine. Because brokers hold no message data, Pulsar assigns topics to brokers in groups called bundles, based on load, and moves them between brokers without copying the stored messages. Coordination uses a metadata store: Oxia, ZooKeeper or another supported backend. The trade-off is more moving parts. A Pulsar cluster runs three kinds of service where Kafka runs one or two, and each has its own configuration, monitoring and upgrade path.

Tiered storage

Both projects can keep recent data on local disks and move older data to object storage. Kafka's tiered storage (KIP-405) uploads completed log segments to a remote tier such as S3 or HDFS; its documented limitations include no support for compacted topics. Pulsar's tiered storage offloads ledgers once they are sealed and immutable, and readers fetch offloaded data transparently. In both, tiered storage changes the cost of long retention more than it changes the day-to-day write path.

Geo-replication

This is one of the clearest differences. Kafka replicates between clusters with MirrorMaker 2, a separate process built on Kafka Connect that copies topics, configurations, consumer group offsets and ACLs from source to target clusters. It works well, but it is another workload to deploy and monitor. Pulsar has geo-replication in the broker, with namespaces assigned to the clusters they replicate to, and supports both asynchronous replication between independent clusters and synchronous replication across data centres. An instance-wide configuration store coordinates clusters that replicate with each other.

Multi-tenancy and subscriptions

Pulsar was designed as a multi-tenant system. Tenants can each have their own authentication and authorisation, storage quotas, message TTL and isolation policies, with namespaces inside each tenant. Kafka's documentation recommends building tenant spaces from hierarchical topic names, enforced with prefixed ACLs and isolated with quotas. Both work. Pulsar's model is built in; Kafka's is a convention you enforce.

On consumption, Pulsar offers four subscription types. Exclusive and Failover give ordered, single-consumer reading; Shared spreads messages across consumers with individual acknowledgement; Key_Shared spreads them while keeping each key on one consumer. Kafka's consumer groups give ordered reading per partition, and share groups, production-ready in 4.2, add cooperative consumption with per-record acknowledgement.

When to choose Kafka, and when to choose Pulsar

  • Choose Kafka if you rely on Kafka Connect, Kafka Streams or tools and managed services built for the Kafka protocol, if your team already runs it, or if a single region and a modest number of teams cover your needs.
  • Choose Pulsar if you run a shared platform for many teams that need strong isolation, if replication between regions is a core requirement, or if you want to scale storage and serving independently.
  • Stay where you are if your current platform meets its service levels. Moving between them means new clients, new operations and a migration of retained data. Most of the gaps people move to close, such as queue semantics in Kafka or a ZooKeeper-free deployment in either, have narrowed.

Where OSSeva fits

OSSeva supports both under one contract. OSSeva for Apache Kafka patches 2.8 through 3.9, including the archived 3.8 and 3.9 lines, and covers 4.x in KRaft mode, with Kafka Connect and Streams in scope on Assure and Operate. OSSeva for Apache Pulsar patches 2.10, 2.11, 3.0, 3.1 and 3.2 after community support ends, covers 4.0, and patches the ZooKeeper under older Pulsar clusters as part of the same engagement. Both run where you run them: bare metal, VMs or any Kubernetes. Priced per cluster; book a discovery call for a quote. For other options, see Pulsar support providers and Kafka support providers, and for running Pulsar without a vendor platform, Pulsar production support.

Frequently asked questions

What is the main architecture difference between Kafka and Pulsar?

Kafka brokers store the partitions they serve. Pulsar brokers are stateless, and storage sits in a separate Apache BookKeeper cluster, with a metadata store alongside. Kafka has fewer components to run; Pulsar can scale and replace brokers without moving stored data.

Kafka vs Pulsar latency: which is lower?

It depends on configuration more than on the project. One documented difference matters: Kafka's default is not to fsync each write, relying on replication and the operating system for durability, while BookKeeper bookies fsync their journal before acknowledging a write by default. Both can be tuned the other way. Measure end-to-end latency with your own durability settings, message sizes and client configuration.

Kafka vs Pulsar performance: which is faster?

Neither in general. When you read a published comparison, check who ran it and how each system was configured. Kafka's design favours sequential writes and batched reads on broker disks; Pulsar's adds a network hop to BookKeeper in exchange for separating storage from serving. Benchmark your own workload on both before deciding.

Does Pulsar still need ZooKeeper?

Not from Pulsar 3.3. From that release Pulsar can run on the Oxia metadata store instead of ZooKeeper; 3.0, 3.1, 3.2 and the 2.x lines need ZooKeeper. Kafka removed ZooKeeper entirely in 4.0.

Is Pulsar better than Kafka for multi-tenant platforms?

Pulsar's tenants and namespaces are built in, with per-tenant authentication, quotas and policies. Kafka can be run multi-tenant with naming conventions, prefixed ACLs and quotas, but you enforce the model yourself.

Can Kafka do queues like Pulsar's Shared subscription?

Yes, from Kafka 4.2. Share groups let consumers cooperatively read a topic, outnumber its partitions and acknowledge records individually, which covers much of what Pulsar's Shared subscription does.

Tags

KafkaApache PulsarComparisonEvent StreamingArchitecture

Ready to get your open source under control?

Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.