Back to blog

// OSSeva Blog

Migration

ClickHouse Keeper vs ZooKeeper: Compatibility, Migration and Sizing

Matt Reynolds8 min read

The short answer

ClickHouse needs a coordination service for replicated tables and distributed DDL. It can use Apache ZooKeeper 3.4.5 or newer, or ClickHouse Keeper, and the ClickHouse documentation recommends Keeper. Keeper is written in C++, runs the Raft consensus algorithm through eBay's NuRaft library, and speaks the ZooKeeper client-server protocol, so ClickHouse and other ZooKeeper clients connect to it unchanged.

It is not a drop-in replacement at the storage level. Keeper's snapshot and log formats and its interserver protocol differ from ZooKeeper's, so a mixed ensemble is impossible and the migration is a stop, convert and start operation in a maintenance window. The clickhouse-keeper-converter tool does the conversion.

What ClickHouse Keeper is

ClickHouse Keeper provides coordination for data replication and distributed DDL. It started as an experimental alternative to ZooKeeper in the 21.x releases, was declared feature complete in 21.12, and has run in ClickHouse Cloud since May 2022. It ships in the ClickHouse server package under the same Apache 2.0 licence, so there is nothing extra to buy.

You can deploy it two ways:

  • Embedded: add a <keeper_server> section to the ClickHouse server config and Keeper starts with clickhouse-server.
  • Standalone: run clickhouse-keeper --config /etc/clickhouse-keeper/config.xml, or clickhouse keeper --config ... if the symlink is missing, on dedicated Keeper nodes.

A clickhouse/clickhouse-keeper Docker image is also published for container and Kubernetes deployments. The configuration file is almost the same in every case. Each Keeper server needs a unique server_id, a log_storage_path, a snapshot_storage_path, and a raft_configuration listing every member. The docs' three-node example uses port 9234 for Raft traffic between Keeper nodes; the client tcp_port defaults to 2181, the same port ZooKeeper uses.

What ClickHouse uses coordination for

ClickHouse stores replica metadata in Keeper or ZooKeeper, not table data. SELECT queries do not touch it. Writes do: each inserted block adds roughly ten entries through several transactions, which is why ClickHouse advises batching inserts to no more than about one INSERT per second. Distributed DDL (ON CLUSTER queries) also runs through the coordination service.

When the coordination service is unreachable, replicated tables switch to read-only mode and INSERTs fail with an exception until the connection returns. This is what "Keeper is not working" usually looks like from the application side.

ClickHouse Keeper vs ZooKeeper: what is compatible

AreaClickHouse KeeperZooKeeper
Language and consensusC++, Raft via NuRaftJava, ZooKeeper Atomic Broadcast (ZAB)
Client protocolZooKeeper-compatibleNative
Consistency by defaultLinearizable writes, non-linearizable reads; quorum_reads runs reads through RaftLinearizable writes, reads served locally by each server
Snapshot and log formatOwn formatIncompatible with Keeper; convert with clickhouse-keeper-converter
Interserver protocolOwn protocolIncompatible, so no mixed ensembles
ACLsworld, auth and digest schemesSame schemes, plus others such as SASL
Four-letter-word commandsSupported, plus Keeper-specific ones such as lgif, csnp and rcvrSupported

What Keeper does not support

The docs list ZooKeeper features that are not implemented yet: create returning a Stat object, TTL nodes, container znodes, persistent watches in addWatch, removeWatch, removeAllWatches, setWatches, and SASL authentication. ClickHouse itself does not need them. Other applications that you might point at the same Keeper cluster may.

What Keeper adds

Keeper has operations ZooKeeper lacks, including check_not_exists, create_if_not_exists, remove_recursive and multi_read, enabled through feature flags. Some flags are on by default from 25.7, and the docs recommend upgrading Keeper to 24.9 or later before moving to 25.7 or later. Dynamic cluster changes through reconfig are available when enable_reconfiguration is set.

Migrating from ZooKeeper to ClickHouse Keeper

The converter needs ZooKeeper 3.4 or later and does a one-to-one conversion: one ZooKeeper ensemble becomes one Keeper snapshot. Plan a maintenance window, because ingestion stops for the duration. Record comparison metrics beforehand so you can check the result.

  1. Stop data ingestion into all ClickHouse nodes.
  2. Stop background tasks that change coordination metadata, for example SYSTEM STOP MERGES, on every node.
  3. Stop all ZooKeeper nodes.
  4. Optionally, and recommended: start and stop the ZooKeeper leader again, which forces it to write a consistent snapshot.
  5. Run the converter on the leader node.
  6. Copy the resulting snapshot to every Keeper node. It must be on every node before any node starts, or a node may elect itself leader with empty state.
  7. Point ClickHouse at the new Keeper cluster.
  8. Start Keeper on all nodes, then restart ClickHouse.
  9. Compare your metrics with the baseline, then resume background tasks and ingestion.
clickhouse-keeper-converter \
  --zookeeper-logs-dir /var/lib/zookeeper/version-2 \
  --zookeeper-snapshots-dir /var/lib/zookeeper/version-2 \
  --output-dir /var/lib/clickhouse/coordination/snapshots

With the full ClickHouse binary installed, clickhouse keeper-converter does the same job. ClickHouse keeps the <zookeeper> section name in its server config even when the members are Keeper nodes:

<zookeeper>
    <node><host>keeper1</host><port>2181</port></node>
    <node><host>keeper2</host><port>2181</port></node>
    <node><host>keeper3</host><port>2181</port></node>
</zookeeper>

Existing ACLs carry over when the ZooKeeper data is either fully encrypted or fully unencrypted. A partly encrypted tree needs extra preparation first, which the Keeper docs describe. Consolidating several ZooKeeper ensembles into one Keeper cluster is possible but means modifying the converter to merge snapshots.

Sizing and operations

  • Cluster size. Keeper uses Raft, so a three-node cluster keeps working with one node down. Use an odd number of nodes.
  • Disks. Put log_storage_path on storage that nothing else is busy with, as you would for ZooKeeper's transaction log.
  • Memory. max_memory_usage_soft_limit defaults to 90% of physical memory through max_memory_usage_soft_limit_ratio.
  • Throughput tuning. For clusters with high part counts or many shards, the docs suggest raising max_requests_batch_size from 100 to 10000 and review force_sync and log compression under coordination_settings.
  • Scale. The replication docs note that production clusters of about 300 servers did not need separate coordination clusters per shard, although auxiliary_zookeepers allows it.
  • Monitoring. clickhouse-keeper-client gives an interactive shell for browsing znodes. echo mntr | nc keeper1 2181 returns the same style of metrics as ZooKeeper, ruok answers imok when healthy, and the system.zookeeper_connection table shows what each ClickHouse server is connected to.
  • Quorum loss. If too many nodes are lost at once, Keeper has a recovery mode (rcvr) that forces a one-node configuration from the most up-to-date survivor. Back up its log and snapshot directories before you use it.

When to keep ZooKeeper

Keeper is the right target when the ensemble serves ClickHouse alone. Many ensembles also serve Kafka 3.x, Solr, HBase or Hadoop. Moving ClickHouse off a shared ensemble does not retire it, and the other products cannot use Keeper's extra features or tolerate its missing ones without testing. In that case the ZooKeeper ensemble still needs patching, and often it is an end-of-life 3.4 to 3.7 build.

How OSSeva helps

OSSeva's ClickHouse support covers the database and the coordination layer under it, including the move to Keeper through our ZooKeeper to Raft migration service. Where ZooKeeper stays, OSSeva ships patched, signed builds for ZooKeeper 3.4 to 3.7 today; see ZooKeeper extended support. The ClickHouse and ZooKeeper page shows how to tell which service your cluster uses. For background on the algorithm, read Raft consensus explained.

Frequently asked questions

Is ClickHouse Keeper free?

Yes. It is part of open-source ClickHouse under the Apache 2.0 licence and ships in the same package as the server.

Can I mix ZooKeeper and ClickHouse Keeper nodes in one ensemble?

No. The interserver protocols differ. An ensemble is all ZooKeeper or all Keeper, and data moves across with clickhouse-keeper-converter.

Does ClickHouse still support ZooKeeper?

Yes. ClickHouse works with ZooKeeper 3.4.5 or newer. Keeper is recommended, not required.

What happens if ClickHouse Keeper is unavailable?

Replicated tables go read-only and INSERTs fail until ClickHouse reconnects. SELECT queries keep working, because they do not use the coordination service.

Should Keeper run embedded or standalone?

Embedded is simpler for small clusters. Standalone Keeper nodes keep coordination traffic, CPU and disk writes away from the query workload, which matters for high availability as the cluster and insert rate grow.

Tags

ClickHouseClickHouse KeeperZooKeeperRaftMigration

Ready to get your open source under control?

Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.