Back to blog

// OSSeva Blog

Migration

ZooKeeper Alternatives: ZooKeeper vs etcd, Consul, KRaft and ClickHouse Keeper

Matt Reynolds11 min read

The short answer

There is no single ZooKeeper alternative. etcd, Consul, KRaft, ClickHouse Keeper, Apache Ratis and Oxia all do some of what Apache ZooKeeper does, and all of them use the Raft consensus algorithm where ZooKeeper uses its own protocol, Zab. None is a drop-in replacement for ZooKeeper in general. What you can replace depends on the product sitting on top: Kafka can move to KRaft, ClickHouse to ClickHouse Keeper, Patroni to etcd or Consul, and Solr, HBase and HDFS high availability cannot move at all today.

So the question is not "what replaces ZooKeeper?" but "what does each product that uses ZooKeeper support instead?" That makes every move a product migration with its own timeline. Until a product is ready, the ZooKeeper under it still needs patching.

What ZooKeeper does, and why it is hard to see

ZooKeeper is a coordination service: a small, strongly consistent tree of data (znodes) replicated across a quorum of servers. Products use it for leader election, cluster membership, configuration, locks and service discovery. Its guarantees come from Zab, an atomic broadcast protocol in which a leader orders every write and a majority of the ensemble must acknowledge it.

It is also the coordination layer in the middle of the stack that many teams do not know they run. Kafka 3.x, Solr, HBase, Hadoop, Hive, Druid, NiFi, Pulsar, Storm and Flink can all depend on it, and products often start an embedded ZooKeeper or ship one inside a distribution such as CDH, HDP or CDP. ZooKeeper is half runtime, half library: the ensemble is a server, but every product also ships its own copy of the client jar. The ZooKeeper hub maps the products and the versions they bundle.

The alternatives, one by one

etcd

etcd is a strongly consistent, distributed key-value store built on Raft. It is best known as the metadata store behind Kubernetes. It offers a gRPC API with leases, watches and transactions, which covers leader election and service discovery, but its data model is a flat key space, not a tree, and it does not speak the ZooKeeper protocol. etcd is an Apache 2.0 project under the CNCF.

Consul

HashiCorp Consul is a service networking tool: service discovery, health checks, a key-value store and a service mesh. Only server agents join the Raft peer set, and HashiCorp recommends three or five servers in production. Consul's key-value store and sessions can do leader election, but most teams adopt it for service discovery rather than as a general coordination store.

KRaft

KRaft is Kafka's built-in, Raft-based metadata quorum. Controllers hold the cluster metadata in a replicated log, and there is no separate ensemble to run. It exists only inside Kafka and serves only Kafka. Kafka 4.0 supports KRaft mode alone, so KRaft is the only choice for any Kafka cluster that wants to stay on a supported release.

ClickHouse Keeper

ClickHouse Keeper is written in C++ and uses Raft through eBay's NuRaft library. It speaks the ZooKeeper client-server protocol, so standard ZooKeeper clients can talk to it, and ClickHouse recommends it over ZooKeeper for replicated tables. Its snapshot, log and interserver formats differ from ZooKeeper's, so you cannot mix ZooKeeper and Keeper nodes in one cluster; clickhouse-keeper-converter converts ZooKeeper data. Protocol compatibility does not mean other vendors support Keeper under their products, so test before pointing anything but ClickHouse at it.

Apache Ratis

Apache Ratis is a Raft library for Java. Its own site says it is not a standalone server like ZooKeeper or Consul. You embed it to build replication into your own service; it is a building block for developers, not something an operator installs to replace an ensemble.

Oxia

Oxia is a metadata store that Apache Pulsar has supported since 3.3.0 (PIP-335). In the Pulsar 5.0.0-M1 preview it becomes the recommended metadata store, with a live migration path from ZooKeeper (PIP-454). It is relevant to Pulsar and BookKeeper, not as a general ZooKeeper replacement.

ZooKeeper vs etcd

This is the comparison most people search for, and the honest answer is that they solve the same class of problem with different interfaces. Both keep a small amount of critical metadata strongly consistent across three or five nodes, and both give you watches and ephemeral state tied to a client session or lease. ZooKeeper models data as a tree with sequential and ephemeral znodes, which makes recipes such as locks and leader election natural, and Apache Curator packages those recipes for Java. etcd uses a flat, versioned key space with multi-key transactions and a gRPC API, which suits Go services and Kubernetes. The practical difference for an existing estate is simpler: a product that only supports ZooKeeper cannot use etcd, however good etcd is.

Side-by-side comparison

ConsensusFormSpeaks ZK protocolLicence
ZooKeeperZabStandalone ensemble (Java)YesApache 2.0
etcdRaftStandalone cluster (Go)NoApache 2.0
ConsulRaftServer agents plus client agentsNoBSL 1.1 from 1.17
KRaftRaftBuilt into Kafka controllersNoApache 2.0
ClickHouse KeeperRaft (NuRaft)Standalone or inside clickhouse-serverYesApache 2.0
Apache RatisRaftJava libraryNoApache 2.0

Which products can use which coordination service

ProductOptions besides ZooKeeper
KafkaKRaft (the only mode in 4.x)
PulsarOxia from 3.3.0; etcd in 4.x, removed in 5.0 (PIP-462)
ClickHouseClickHouse Keeper (recommended)
Patronietcd, Consul, Kubernetes; built-in Raft deprecated since 3.0.0
Vitesszk2, etcd2 or consul topology services
Apache DubboNacos, Redis, Consul, etcd, Kubernetes registries
FlinkKubernetes HA (from 1.12)
NiFi 2.xKubernetes leases for leader election
DruidKubernetes extension, experimental; ZooKeeper is still the default
Solr, HBase, HDFS HA, StormNone. Solr 10 and Storm 3.1 still require ZooKeeper

Two points are often wrong elsewhere. Apache Cassandra has never used ZooKeeper; it uses gossip. And Druid has not dropped ZooKeeper: ZooKeeper-free operation exists only through its experimental Kubernetes extension.

Licence and support status

  • ZooKeeper: 3.9 and 3.8 are supported (latest 3.9.6 and 3.8.7). 3.7 reached end of life on 2 February 2024, and 3.4 to 3.6 before that.
  • etcd: v3.5, v3.6 and v3.7 are supported; v3.7.0 went GA on 8 July 2026. v3.4 reached end of life with v3.4.45 on 1 June 2026.
  • Consul: HashiCorp moved Consul to the Business Source License on 10 August 2023, starting with 1.17.0, and later 1.15 and 1.16 patch releases are BSL too. Each version converts to MPL 2.0 four years after release. In March 2026 the licensor became IBM. Consul Enterprise now follows IBM's support cycle from 2.0: 1.22.x support ends on 31 October 2026, 1.21.x (the last LTS) on 30 April 2027, and 2.0.x on 30 April 2028. The 1.15 and 1.18 LTS lines are already past end of support. See Consul end of life, BSL and IBM.
  • Kafka: 3.9.x, the last line with ZooKeeper, is archived. Supported releases are 4.1, 4.2 and 4.3.

Scaling, latency and fault tolerance

The quorum maths is the same for every system here. A cluster of three nodes survives the loss of one node and five survives two, because a write needs a majority. Adding nodes improves fault tolerance, not write throughput: every write still goes through one leader and must replicate to a majority, so write latency grows with the slowest member of that majority. This is why none of these systems is meant to hold bulk data, and why the usual advice for scalability problems is to put less in the coordination store rather than to add servers.

Reads differ. A ZooKeeper server answers reads from its local copy, so reads scale with ensemble size at the cost of possibly slightly stale data. etcd serves linearizable reads by default, with serializable reads as a cheaper option. Running any of them in Docker or on Kubernetes changes the packaging, not these properties. A distributed system that needs a scalable ZooKeeper alternative usually has a data model problem first.

For a Kafka broker estate, the use case is narrower. KRaft removes a separate cluster, and the controllers already hold the metadata log, so controller failover no longer reloads state. That is an operational gain rather than a benchmark claim.

When to consider a ZooKeeper alternative

Consider one when the product on top supports it, the product's own roadmap is moving that way, and you have the change capacity. Kafka is the clearest case, because ZooKeeper mode ends with 3.9. ClickHouse and Pulsar are next. For Patroni, moving to etcd is a configuration change plus a new cluster to operate, which is only worth it if you already run etcd. For Solr, HBase and HDFS there is nothing to move to, so the work is keeping ZooKeeper supported.

Be wary of adopting an alternative to ZooKeeper just to remove one. Replacing a three-node ZooKeeper ensemble with a three-node etcd cluster changes the software you patch, not the amount of infrastructure you run.

Migration is per product

Every path above has its own mechanism: Kafka's ZooKeeper-to-KRaft migration mode (our KRaft guide), clickhouse-keeper-converter, Pulsar's PIP-454 migration, or a new Patroni cluster pointed at a different DCS. None of them removes ZooKeeper everywhere at once. Most estates end up with some products migrated and others still on ZooKeeper for years.

That gap is where patched ZooKeeper is the bridge. OSSeva's ZooKeeper extended support ships patched drop-in builds for 3.4, 3.5, 3.6 and 3.7, keeps configuration and logs unchanged, provides written attestation for CVEs that do not apply to your deployment, and plans the move to Raft-based coordination product by product. For Consul estates, see Consul extended support.

Frequently asked questions

Is ZooKeeper outdated?

No. It is maintained, with 3.8.7 and 3.9.6 released in September 2026 and 3.10 in preparation. What is outdated is running 3.7 or older, which gets no fixes.

Why did Kafka get rid of ZooKeeper?

Running two distributed systems meant two things to patch, secure and scale, and a controller that had to reload state from ZooKeeper on failover. KRaft keeps the metadata in Kafka's own replicated log.

Can etcd replace ZooKeeper directly?

Only for products that support etcd, such as Patroni, Vitess and Dubbo. etcd does not speak the ZooKeeper protocol.

Can ClickHouse Keeper replace ZooKeeper for other products?

It speaks the ZooKeeper client protocol, so it may work, but only ClickHouse documents it. Check with each product's vendor before relying on it.

What are the use cases for ZooKeeper?

Leader election, cluster membership, distributed locks, configuration and service discovery for systems such as Solr, HBase, Hadoop, Hive, NiFi, Storm and older Kafka.

Tags

ZooKeeperetcdConsulKRaftRaftClickHouse Keeper

Ready to get your open source under control?

Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.