Back to blog

// OSSeva Blog

Migration

Kafka ZooKeeper to KRaft Migration: The Runbook for Self-Managed 3.x Clusters

Matt Reynolds14 min read

The short answer

Apache Kafka 4.0 runs only in KRaft mode. ZooKeeper support was removed, so a cluster still running in ZooKeeper mode cannot upgrade straight to 4.x. It first has to migrate from ZooKeeper to KRaft on a 3.x release, and the Kafka community's guidance is to do that on 3.9, the final 3.x line and the bridge release for the migration.

The good news is that you can migrate to KRaft without downtime. The migration itself is an online process. You add a KRaft controller quorum beside the existing cluster, the controllers copy the metadata from ZooKeeper, the brokers are restarted into KRaft mode one at a time, and finally the controllers are taken out of migration mode. Until that last step you can revert to ZooKeeper mode. After it, you cannot.

What KRaft is, and why Kafka removed ZooKeeper

KRaft (Kafka Raft) moves Kafka's metadata out of an external Apache ZooKeeper ensemble and into Kafka itself. A small set of controller nodes maintain the cluster metadata as a replicated log using the Raft consensus protocol, and the brokers follow that log. There is no second distributed system to deploy, secure, monitor and patch.

The practical benefits are the ones operators felt most with ZooKeeper: faster controller failover, much higher partition limits because metadata no longer has to be reloaded from ZooKeeper on failover, one security model instead of two, and one fewer system with its own CVEs and its own end-of-life calendar. The ZooKeeper ensemble behind an old Kafka cluster is frequently older than the brokers and has not been patched in years.

Before you begin the migration

  • Get the cluster onto 3.9 first. The ZooKeeper to KRaft migration is supported on several 3.x releases, but 3.9 carries the most fixes to the migration path and is the version you will run in migrated form before moving to 4.x. Upgrade the brokers to 3.9 while still using ZooKeeper, and set the inter-broker protocol version to match.
  • Check what the migration does not carry. Review the Kafka documentation for your version's list of features not supported during migration. Clusters that use features such as JBOD, SCRAM credentials or delegation tokens should confirm support on their exact 3.x release before starting.
  • Size the KRaft controller quorum. Three controller nodes tolerate the loss of one; five tolerate two. Run them as dedicated nodes rather than combined broker-controllers for production clusters.
  • Inventory your clients. Kafka 4.0 also drops old client protocol versions, so clients that are many years old will stop working after the final upgrade even though the migration itself does not touch them. Find them now.
  • Back up ZooKeeper. Snapshot the ZooKeeper data before starting. It is the only copy of the metadata until the migration has run.

The migration process, phase by phase

Phase 1: deploy the KRaft controller quorum in migration mode

Starting the migration means provisioning the controller nodes with the same cluster ID as the existing cluster, format their storage, and configure them with zookeeper.metadata.migration.enable=true and the ZooKeeper connection string. The controllers join the quorum and wait for the brokers.

# controller.properties (excerpt)
process.roles=controller
node.id=3000
controller.quorum.voters=3000@ctrl-1:9093,3001@ctrl-2:9093,3002@ctrl-3:9093
controller.listener.names=CONTROLLER
zookeeper.metadata.migration.enable=true
zookeeper.connect=zk-1:2181,zk-2:2181,zk-3:2181

Phase 2: enable migration on the brokers

Add the migration flag, the controller quorum configuration and the controller listener to every broker, keeping the ZooKeeper connection, and do a rolling restart. When every broker has registered with the new quorum, the active KRaft controller copies the metadata from ZooKeeper into the metadata log. The cluster is now in dual-write mode: the KRaft controllers own the metadata, and every change is also written back to ZooKeeper.

Watch the controller logs and the migration state metric, and do not move on until they report that the metadata migration is complete. Large clusters with many partitions take longer here, and nothing else should change while it runs.

Phase 3: restart the brokers in KRaft mode

Remove the ZooKeeper configuration and the migration flag from each broker, set process.roles=broker, and roll the brokers again. This is the step that migrates the brokers to KRaft: each one comes back reading metadata from the controller quorum, with no ZooKeeper connection at all.

During this phase you can still revert to ZooKeeper mode, because the controllers are still copying every change to ZooKeeper. Reverting means rolling the brokers back to their ZooKeeper configuration while the controllers remain in migration mode, following the documented rollback steps for your version.

Phase 4: finalize the migration

Remove zookeeper.metadata.migration.enable and the ZooKeeper settings from the controllers and restart them. The KRaft controllers leave migration mode and stop writing to ZooKeeper. The cluster is now in full KRaft mode.

This is the point of no return. After the controllers are taken out of migration mode, ZooKeeper no longer receives metadata changes and there is no supported way back. Leave the cluster in Phase 3 long enough to be confident, then finalize, then decommission the ZooKeeper nodes.

What goes wrong during a KRaft migration

The migration steps are well documented. The problems that stretch a one-week change into a quarter are almost always environmental, and they are worth checking before starting the migration on an existing Kafka cluster:

  • Brokers that never register with the quorum. The metadata migration does not start until every broker in the cluster has registered with the KRaft controllers. A single broker with a stale configuration, a firewall rule that blocks the controller listener, or a node that was quietly decommissioned but never removed from ZooKeeper holds the whole process at Phase 2.
  • Mismatched cluster IDs. The controllers must be formatted with the cluster ID already stored in ZooKeeper. Formatting them with a new random ID produces a quorum that the brokers will refuse to join.
  • Controller listener security. The controller listener needs its own TLS and SASL configuration. Teams that copy the broker listener settings and forget the controller side find that the brokers cannot reach the quorum they were just configured to use.
  • Tooling that still talks to ZooKeeper. Old scripts that create topics or read ACLs through the ZooKeeper connection stop working once the brokers leave ZooKeeper mode. Move operational tooling to the Admin API before Phase 3.
  • Monitoring gaps. Dashboards built on ZooKeeper metrics go quiet after the migration is complete. Add KRaft controller metrics, including the metadata log offset and the active controller count, before you depend on them.

A migration checklist

  1. Upgrade brokers to 3.9 in ZooKeeper mode and confirm the cluster is healthy.
  2. Snapshot ZooKeeper and record the cluster ID.
  3. Provision and format three or five KRaft controller nodes with that cluster ID.
  4. Start the controllers in migration mode.
  5. Roll the brokers with the migration flag and controller settings, then confirm the migration is complete in the controller logs.
  6. Roll the brokers into KRaft mode, without ZooKeeper settings.
  7. Run in this state long enough to be confident; this is the last point at which you can revert.
  8. Take the controllers out of migration mode to finalize, then decommission ZooKeeper.
  9. Upgrade the cluster from 3.9 to 4.x.

After the migration: moving to Kafka 4.0

A 3.9 cluster in full KRaft mode can then take the ordinary rolling upgrade to 4.x. Plan for the other 4.0 changes at the same time: the brokers need a newer Java baseline, old client protocol versions are removed, and several long-deprecated configurations are gone. None of those is as large as the ZooKeeper change, but all three have caught teams that treated the version bump as routine.

ZooKeeper vs KRaft: what changes for operators

ZooKeeper modeKRaft mode
Metadata storeExternal ZooKeeper ensembleReplicated log on KRaft controller nodes
Systems to patchKafka and ZooKeeper, on separate release calendarsKafka only
Controller failoverNew controller reloads state from ZooKeeperStandby controllers already hold the log
Security configurationTwo systems, two sets of ACLs and TLS settingsOne
Kafka versionsUp to 3.93.x in KRaft mode, and all of 4.x

When you cannot migrate yet

Some clusters cannot move on the community's timeline. The usual reasons are an application platform that pins a Kafka version, Confluent Platform 7.x deployments tied to ZooKeeper, change-freeze calendars in regulated environments, or plain capacity: a migration that is routine for one cluster is a programme for forty.

For those estates the risk is not ZooKeeper itself but running an unpatched 3.x broker and an unpatched ZooKeeper ensemble while the migration waits. OSSeva provides extended support for Kafka 3.x on ZooKeeper: backported CVE fixes for the Kafka brokers and the ZooKeeper ensemble, plus a migration runbook and hands-on help to finalize KRaft when you are ready. See also Confluent Platform 7.x end of support if your ZooKeeper clusters run under Confluent.

Frequently asked questions

Is ZooKeeper deprecated in Kafka?

It was deprecated during the 3.x series and removed in Kafka 4.0. Kafka 3.9 is the final line that can run with ZooKeeper.

Can I upgrade from Kafka 3.x on ZooKeeper directly to 4.0?

No. You must migrate from ZooKeeper to KRaft on 3.x first, then upgrade the KRaft cluster to 4.x.

Can you roll back a KRaft migration?

Yes, until you finalize it. While the controllers are in migration mode they keep writing metadata to ZooKeeper, so the brokers can be reverted to ZooKeeper mode. Once the controllers leave migration mode, rollback is not supported.

How many KRaft controller nodes do I need?

Three for most production clusters, five where you need to survive two simultaneous controller failures. Use an odd number.

Does the migration require downtime?

No. Every phase is a rolling restart, and producers and consumers keep working throughout. Plan it as a change with a maintenance window anyway, because the finalize step cannot be undone.

Tags

Apache KafkaKRaftZooKeeperMigrationKafka 4.0

Ready to get your open source under control?

Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.