// OSSeva Blog
OperationsFlink High Availability: ZooKeeper vs Kubernetes HA Services
The short answer
Flink high availability removes the JobManager as a single point of failure. Apache Flink ships two HA service implementations. ZooKeeper HA services work with every Flink deployment and need a running ZooKeeper quorum. Kubernetes HA services work only when Flink runs on Kubernetes, and store their state in ConfigMaps. Both are documented in the current stable release, Flink 2.3.0. If your Flink cluster runs on Kubernetes, Kubernetes HA removes the ZooKeeper dependency. If it runs on YARN or bare metal, ZooKeeper HA stays, and so does the ensemble behind it.
How Flink's JobManager achieves high availability
A Flink cluster has one leading JobManager at a time and can run standby JobManagers that take over leadership if the leader fails. The HA services provide three things:
- Leader election: picking one leader out of the candidates.
- Service discovery: finding the address of the current leader.
- State persistence: keeping what a successor needs to resume, such as JobGraphs, user code jars and completed checkpoints.
High availability does not make a failure invisible. When the leader JobManager fails, the running Flink job still restarts; a standby replica shortens the restart because it does not wait for a new JobManager to start.
In both implementations the metadata itself goes to the file system at high-availability.storageDir, typically HDFS or Amazon S3. Only a pointer to it is stored in ZooKeeper or in Kubernetes. When a job reaches a terminal state, Flink deletes its HA data, and the JobResultStore keeps the final result until cleanup succeeds.
Configure ZooKeeper HA
From the Flink ZooKeeper HA docs:
high-availability.type: zookeeper
high-availability.zookeeper.quorum: zk1:2181,zk2:2181,zk3:2181
high-availability.zookeeper.path.root: /flink
high-availability.cluster-id: /cluster_one # unique per cluster
high-availability.storageDir: hdfs:///flink/recovery
high-availability.type, storageDir and the quorum are required. The path root and cluster-id are recommended, and on bare metal each Flink cluster sharing an ensemble needs its own cluster-id. Flink 1.16 and earlier used the older key high-availability: zookeeper. Flink 2.0 and later read only config.yaml; flink-conf.yaml is no longer supported. Kerberos settings such as zookeeper.sasl.service-name and client retry settings sit alongside.
Configure Kubernetes HA
high-availability.type: kubernetes
high-availability.storageDir: s3://flink/recovery
kubernetes.cluster-id: cluster1337
These high availability services work for standalone Flink on Kubernetes and for the native Kubernetes integration. Kubernetes HA needs a service account that can create, edit and delete ConfigMaps. To deploy standby JobManagers, set kubernetes.jobmanager.replicas above 1. With the Flink Kubernetes Operator, you set the same keys under spec.flinkConfiguration in a FlinkDeployment. The operator sets the cluster-id itself and rejects a deployment that sets it. Its docs call Kubernetes HA the common choice because it needs no external system, and note that last-state upgrades and automatic rollback depend on HA metadata being present.
ZooKeeper vs Kubernetes HA at a glance
| ZooKeeper HA | Kubernetes HA | |
|---|---|---|
| Where it runs | Any deployment: standalone, YARN, Kubernetes | Kubernetes only |
| Leader election and pointers | ZooKeeper znodes | ConfigMaps |
| Extra system to run | A ZooKeeper quorum | None beyond the Kubernetes cluster |
| Metadata | storageDir | storageDir |
Which ZooKeeper does Flink use?
Flink builds against a shaded ZooKeeper client. The zookeeper.version in Flink's root pom.xml is 3.7.1 for Flink 1.17.2, 1.18.1 and 1.20.3, and 3.7.2 for Flink 2.0.0, 2.2.0 and 2.3.0. The ZooKeeper 3.7 line reached end of life on 2 February 2024. 3.7.1 is inside the affected range of CVE-2023-44981, and every 3.7 release is affected by CVE-2024-23944. Scanners often report the client under the flink-shaded-zookeeper artifact name. The ensemble itself is separate and often older; see Flink and ZooKeeper for how to check it.
How to switch from ZooKeeper HA to Kubernetes HA
HA pointers do not move between services, so plan the switch as a stop and restart from a savepoint:
- Confirm the target runs on Kubernetes and the service account can manage ConfigMaps.
- Stop each job with a savepoint:
./bin/flink stop --savepointPath s3://flink/savepoints $JOB_ID. - Replace the ZooKeeper keys with
high-availability.type: kubernetesand akubernetes.cluster-id. Keep a durable storageDir. - Restart the cluster and resubmit each job with
--fromSavepoint. - Kill a JobManager pod in staging and confirm the job recovers from the latest successful checkpoint.
- Remove the old znodes under the Flink path root once nothing uses them.
Flink on YARN cannot make this move, because Kubernetes HA only works on Kubernetes.
Where OSSeva fits
Where ZooKeeper HA stays, the ensemble behind it needs patches. OSSeva ships patched, signed ZooKeeper 3.4, 3.5, 3.6 and 3.7 builds today, supports 3.8 and 3.9, and provides VEX attestation for scanner findings. Flink extended support covers the Flink estate and the ZooKeeper under it, in Patch, Assure and Operate tiers. For the wider move off ZooKeeper, see ZooKeeper to Raft migration.
Frequently asked questions
Does Flink need ZooKeeper?
Only for ZooKeeper HA. On Kubernetes, Kubernetes HA services replace it. Without HA, Flink needs neither.
Why ZooKeeper, if Kubernetes HA exists?
ZooKeeper HA works with every deployment, including YARN and standalone clusters. When Kubernetes HA arrived in Flink 1.12, the project said the ZooKeeper dependency would not be dropped.
If a cluster is gone, can I redeploy and restore from a savepoint in S3?
Yes, if the savepoint or retained checkpoint is on durable storage. Resubmit with --fromSavepoint. Kubernetes HA ConfigMaps survive deleting the deployment, so jobs can also recover from the latest checkpoint.
Is Kubernetes HA production-ready?
It is one of the two documented HA services, and the Flink Kubernetes Operator treats it as the usual choice.
Tags
Related articles
ZooKeeper Vulnerabilities by Version: CVEs in 3.4 to 3.9
September 29, 2026MigrationZooKeeper Alternatives: ZooKeeper vs etcd, Consul, KRaft and ClickHouse Keeper
September 29, 2026ComplianceWhy Your Scanner Flags the ZooKeeper Inside a Product You Bought, and How VEX Attestation Answers It
September 29, 2026Ready to get your open source under control?
Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.