Back to blog

// OSSeva Blog

Migration

Upgrade ZooKeeper 3.4 or 3.5 to 3.8 or 3.9 In Place: A Rolling Upgrade Guide

Randall McClure11 min read

The short answer

Upgrade one server at a time, followers first and the leader last, and move through each minor line's final release: 3.4.14, then 3.5.10, 3.6.4, 3.7.2, 3.8.7 and, if you want the current line, 3.9.6. Before the first 3.5 node starts, set 4lw.commands.whitelist, move or disable the admin server on port 8080, and make sure every server has a snapshot file or set snapshot.trust.empty=true for that boot. At 3.8, port your logging configuration from Log4j to Logback.

The Apache ZooKeeper releases page lists 3.9.6 as the current release and 3.8.7 as the stable release. 3.7 has been end of life since 2 February 2024. The project supports two lines at a time and expects to retire the older stable line within about six months of a new minor release, so 3.8 will go when 3.10 arrives. If you are doing the work now, finish on 3.9.

Why one line at a time

The project documents specific rules for some hops and none for skipping lines, so the safe path is the one that has been exercised:

  • The reconfiguration guide says a rolling upgrade to 3.5 should only start from 3.4.6 or later. Otherwise peers log Invalid server id: -65536. The upgrade FAQ recommends the latest 3.4.x first, which is 3.4.14.
  • Rolling from 3.5.7 to 3.6.0 failed in practice (ZOOKEEPER-3758, and ZOOKEEPER-3778, which involved multiAddress.reachabilityCheckEnabled). Both were fixed in 3.6.1. Never land on an x.y.0 release; use the last release of each line.
  • Each hop is a normal rolling restart. Five hops on a five-node ensemble is tedious but low risk, and you can pause at any line for as long as you need.

If you test a skip, for example 3.5.10 straight to 3.8.7, in staging with your data and clients and it is clean, that is your decision to make. The project does not publish a compatibility matrix that promises it will work.

Breaking changes by version

VersionChangeWhat to do
3.5.0Embedded Jetty AdminServer starts by default on port 8080Set admin.serverPort to a free port, or admin.enableServer=false. Port 8080 often clashes with a co-located app.
3.5.0Dynamic reconfiguration. Server lines can carry the client port (server.1=host:2888:3888;2181) and ZooKeeper can write a separate dynamic config file on bootThe old static format is still accepted. Make sure your config management does not fight the dynamic file.
3.5.3reconfigEnabled defaults to falseLeave it off during the upgrade. If you set it, set it the same on every server.
3.5.3Four-letter-word commands disabled except srvrAdd 4lw.commands.whitelist=stat, ruok, conf, isro, mntr, the set your monitoring uses, or *. Otherwise health checks get "is not executed because it is not in the whitelist".
3.5.5Startup check refuses to start with transaction logs but no snapshot (ZOOKEEPER-2325)3.4 can run without ever writing a snapshot. On 3.5.6 and later, set snapshot.trust.empty=true for the upgrade boot, then set it back to false.
3.5.5Quorum TLS (sslQuorum) and client TLS on secureClientPort, which needs the Netty connection factoryOptional. Enable it after the upgrade with the port unification procedure below.
3.6.0ls2 and rmr removed from the CLI. The followers metric renamed to learners.Use ls -s and deleteall. Update dashboards and alerts.
3.6.0Version.getRevision() removed, then restored as deprecated in 3.6.1Client code should use getRevisionHash().
3.8.0Log4j 1.2 replaced by Logback (ZOOKEEPER-4427)Rewrite log4j.properties as logback.xml and check log paths, rotation and any shipper that reads them.
3.9.0AdminServer API for streaming snapshots; client trust and key stores reload dynamicallyNothing to change. Restrict access to the admin server, because it can now stream data out.

The 3.9 documentation still lists Java 8 or later as the requirement, so the JVM does not force a change. Many teams use the upgrade to move to Java 11 or 17 anyway.

Upgrade steps

  1. Back up every server. Copy the version-2 directories under dataDir and dataLogDir on each node before each hop. Record srvr and mntr output, including the zxid and each node's mode.
  2. Find the leader. echo srvr | nc <host> 2181 shows Mode: leader or follower. You will restart the leader last in every round.
  3. Get to 3.4.14. Roll every server to the final 3.4 release first.
  4. Check for snapshots. Look for snapshot.* files in each server's version-2 directory. If any server has none, add snapshot.trust.empty=true to its config for the 3.5 boot.
  5. Prepare the 3.5 config. Add the 4lw whitelist, set admin.serverPort or disable the admin server, and keep reconfigEnabled=false. Optionally move client ports into the server lines.
  6. Roll to 3.5.10. Stop a follower, swap the binaries, start it, and wait until it rejoins: srvr shows follower and the leader's mntr shows the expected zk_synced_followers. Repeat, then do the leader.
  7. Clear the snapshot override. Once every server has written a snapshot on 3.5, remove snapshot.trust.empty and do a rolling restart, as the admin guide recommends.
  8. Repeat for 3.6.4 and 3.7.2. Update CLI scripts and metric names at 3.6.
  9. Roll to 3.8.7 with the Logback config in place. Confirm logs appear where your tools expect them before moving to the next node.
  10. Roll to 3.9.6 if you want the current line.
  11. Upgrade clients after servers. Kafka, HBase, Solr, Hadoop and Curator-based services embed their own ZooKeeper client, so check each product's supported version.

Enabling quorum TLS without downtime

The admin guide documents a three-pass rolling procedure, available since 3.5.5. Create keystores and truststores for every server first. Pass one: add sslQuorum=false, portUnification=true, serverCnxnFactory=org.apache.zookeeper.server.NettyServerCnxnFactory and the ssl.quorum.* keystore and truststore settings, then roll. Pass two: set sslQuorum=true and roll again. Pass three: set portUnification=false and roll once more. Check that the quorum is healthy after each restart.

Test plan

  • Rehearse every hop in staging with a copy of production data, including a snapshot-less 3.4 node if you have one.
  • After each node restart, confirm quorum, the node's mode and that the leader reports all followers synced.
  • Run a write and read check through a client: create, read, update and delete a test znode, and check that an ephemeral node vanishes when its session closes.
  • Check that every dependent service (Kafka brokers, HBase masters, Solr nodes) reconnects without a wave of session expiries.
  • Check monitoring: 4lw commands answer, the admin server is on its new port, metric names match dashboards, and logs flow after 3.8.
  • Run a leader failover drill on the final version.

Rollback plan

  • Mid-hop: while only some servers run the new version, stop an upgraded server, put the old binaries and config back and start it. The others keep quorum.
  • After a full hop: the project does not document downgrades. Once every server has written snapshots on the new version, a rollback means stopping the ensemble, restoring the version-2 backups from before the hop and starting on the old binaries. Writes since the backup are lost. For most ZooKeeper data (leader election, configuration, locks) that is recoverable, but check it for your workload.
  • Keep each hop's backups until the next hop is complete and stable.

If you cannot upgrade in time

Many 3.4 and 3.5 ensembles sit under products that pin them, such as older Kafka, HBase or Solr releases, and cannot move until the product does. OSSeva ships patched, signed builds of ZooKeeper 3.4, 3.5, 3.6 and 3.7, patches embedded copies together with the product above them, provides VEX attestation for scanner findings, and runs ensembles 24/7 on the Operate tier. See ZooKeeper extended support, ZooKeeper support coverage, the end-of-life pages for ZooKeeper 3.4 and ZooKeeper 3.5, and the ZooKeeper end-of-life tracker. For the CVE picture, read ZooKeeper vulnerabilities by version. If ZooKeeper is only there for Kafka, the KRaft migration may remove it altogether.

Common questions

Can I upgrade ZooKeeper 3.4 directly to 3.8?

Not as a rolling upgrade the project documents. Go through 3.5.10 at least, because the 3.4 to 3.5 changes need handling on their own, and the safest route continues one line at a time.

What does "No snapshot found, but there are log entries" mean?

A 3.5.5 or later server found transaction logs without a snapshot, which 3.4 allowed. Set snapshot.trust.empty=true on 3.5.6 or later for one boot, then remove it once a snapshot exists.

Should I stop at 3.8 or go to 3.9?

Go to 3.9. 3.8 is the stable line today, but under the project's two-line policy it is the next one to be retired.

Tags

Apache ZooKeeperUpgradeEnd of LifeKafkaRolling Upgrade

Ready to get your open source under control?

Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.