Back to blog

// OSSeva Blog

Migration

Upgrading Apache Flink 1.18, 1.19 and 1.20 to Flink 2.x: Removed APIs, State, config.yaml, Connectors and Rollback

Matt Reynolds11 min read

The short answer

A move from Flink 1.x to 2.x is a major-version migration. The Flink 2.0 release notes remove whole APIs, remove deprecated configuration options and the old flink-conf.yaml parser, drop Java 8 and the per-job deployment mode, and state plainly that "State Compatibility is not guaranteed between 1.x and 2.x". Job code has to be ported, and every stateful job needs a plan for its state.

The community supports the current and previous minor release, which today means 2.3 and 2.2, plus 1.20, which the download page labels LTS. Flink 1.19 had its final release, 1.19.3, on 10 July 2025, and 1.18 stopped at 1.18.1 in January 2024. The Flink end-of-life tracker has every line.

Our recommendation for a cluster on 1.18 or 1.19: move to 1.20 first, which is a normal minor upgrade with savepoints, then port jobs to 2.x one at a time from 1.20.

Where the release lines stand

LineLatest releaseStatusJava
1.181.18.1 (19 Jan 2024)No longer supported8, 11, 17
1.191.19.3 (10 Jul 2025)No longer supported8, 11, 17
1.201.20.5 (3 Jun 2026)LTS8, 11, 17
2.02.0.2 (11 May 2026)Outside the two supported minors11, 17, 21
2.22.2.1 (15 May 2026)Supported (previous minor)11, 17, 21
2.32.3.0 (25 Jun 2026)Supported (current minor)11, 17, 21

Release dates come from the Flink download page. The page labels 1.20 LTS but gives no end date. FLIP-458, the proposal behind the LTS line, describes two years of bug fixes and security updates; the wiki page still shows its vote thread as to be decided, so treat the length as a plan rather than a commitment.

Why 1.20 is the bridge

Three things make 1.20 the right place to stand before the jump:

  • Savepoints. The upgrade guide's compatibility table shows savepoints from 1.8 onwards restoring on 1.20. A 1.18 or 1.19 job moves to 1.20 with the usual stop-with-savepoint and restore.
  • Deprecation warnings. Under Flink's deprecation rules, an API deprecated in 1.18 can be removed in 2.0. Compiling against 1.20 shows you what 2.0 removes while the job still runs.
  • config.yaml. Flink 1.19 introduced config.yaml in standard YAML 1.2 and made it the default. 1.19 and 1.20 still read a legacy flink-conf.yaml if it is present; 2.0 does not. Converting on 1.20 takes the configuration change out of the 2.x window.

Removed APIs

The 2.0 release notes list these API sets as completely removed:

Removed in 2.0Replacement named in the release notes
DataSet APIDataStream API, or Table API and SQL. The 1.20 docs have a DataSet to DataStream migration guide.
Scala DataStream and DataSet APIsThe Java DataStream API
SourceFunction, SinkFunction and Sink V1Source and Sink V2
TableSource and TableSinkDynamicTableSource and DynamicTableSink
TableSchema, TableColumn and TypesSchema, Column and DataTypes

The same release removed configuration setters that took Java objects from StreamExecutionEnvironment and ExecutionConfig; those settings now go through Configuration and ConfigOption. User functions lost full access to ExecutionConfig and reach createSerializer(), getGlobalJobParameters() and isObjectReuseEnabled() through RuntimeContext. The State Processor API lost its DataSet-based methods. The notes also warn that some removed classes still exist under other packages for internal use and must not be used.

Our recommendation: build an inventory per job before estimating. A job written in Scala on the DataSet API is a rewrite. A Java DataStream job that uses a custom SourceFunction is a contained port. A SQL job that uses only maintained connectors may need little more than a rebuild.

State and savepoints

This is the part to plan most carefully. The release notes say state compatibility between 1.x and 2.x is not guaranteed, and the savepoint compatibility table in the 2.x upgrade guide has columns only up to 1.20.x. It makes no promise for restoring a 1.x savepoint on 2.x. The 2.0 release also moved Kryo to version 5.6 and introduced new built-in serializers for maps, lists and sets, enabled by default.

In practice that gives three routes per stateful job, in our order of preference:

  1. Try the restore on a copy. Take a savepoint on 1.20, restore it on a 2.x test cluster with the ported job, and compare outputs. Jobs whose state uses Flink's own serializers or POJO types are the better candidates. Treat state that falls back to Kryo as the likeliest to fail.
  2. Rebuild state from the source. For jobs reading from Kafka or another replayable log, start the 2.x job from an earlier offset with fresh state and let it catch up, if retention covers the window your state represents.
  3. Translate the state. Read the 1.20 savepoint with the State Processor API and write bootstrap data for the 2.x job. This is custom engineering, so keep it for state that cannot be rebuilt.

Whichever route a job takes, the guide's general rules still apply. Set explicit uids on every stateful operator before the last 1.20 savepoint, since a changed topology can only match state by ID. Keep savepoint files readable from the new installation under the same absolute path. Table API and SQL jobs carry an extra risk: the guide warns that any change to both the query and the Flink version can make state incompatible, because a new version can produce a different execution plan.

Java

Flink 2.0 requires Java 11 or later, supports Java 21, and changed the default and recommended version to Java 17, which mainly affects the Docker images and builds from source. Jobs, connectors and libraries all have to run on the target JVM. Our recommendation is to move 1.20 clusters to Java 17 first, since 1.20 supports it, so the JVM change is tested separately from the API port.

Configuration: config.yaml and removed options

From 2.0, Flink reads only config.yaml. The configuration guide documents a migration script, available from 1.19:

# Run from $FLINK_HOME with the old flink-conf.yaml in conf/
bin/migrate-config-file.sh

The script writes conf/config.yaml. Because the legacy parser read every value as a string, the generated file quotes some values; Flink converts them to the right types when it parses them. config.yaml follows YAML 1.2, so null values, comments after a # and escaping in strings behave differently from the old file, and the guide lists those differences.

2.0 also removed configuration options that had been deprecated long enough under its rules. The list in the release notes includes the old jobmanager.heap.size and taskmanager.heap.size, the taskmanager.network.memory.* and taskmanager.net.* families, web.port and web.address, security.ssl.enabled, and the pipeline.registered-kryo-types and pipeline.default-kryo-serializers settings. Some keys were renamed rather than removed: checkpoint and savepoint directories are now execution.checkpointing.dir and execution.checkpointing.savepoint-dir, with state.checkpoints.dir and state.savepoints.dir still accepted as deprecated keys. Check every key in your configuration against the 2.x reference rather than relying on a clean start.

On the command line, run-application is gone (use run -t kubernetes-application), per-job mode is gone in favour of application mode, and sql-client.sh -u is replaced by -f.

Connectors

Connectors live outside the main Flink release and carry the Flink version in their artifact version. The Kafka connector shows the pattern on Maven Central: 3.4.0-1.20 for 1.20, 4.0.1-2.0 for 2.0, and 5.0.0-2.1 and 5.0.0-2.2 for the later minors. A 1.x connector JAR will not work on 2.x, because it was built on the removed source and sink APIs.

The 2.0 release notes promised adapted Kafka, Paimon, JDBC and Elasticsearch connectors right after 2.0.0, and the rest of the community connectors within three minor releases, that is by 2.3. When we checked Maven Central on 5 October 2026, the picture was mixed:

  • Published for 2.x: Kafka (to 2.2), JDBC core (to 2.2), Elasticsearch 7 and 8 (2.0) and the AWS Kinesis Streams connector (2.0).
  • No 2.x build yet: Pulsar, Cassandra, MongoDB, RabbitMQ and the older Kinesis connector.
  • No build for 2.3 yet: Kafka's newest artifacts target 2.1 and 2.2.

That list decides the target minor for many estates. A pipeline on Kafka and JDBC can move today. A pipeline on Pulsar or Cassandra has to wait for the connector, or move to a different sink.

High availability and ZooKeeper

Both ZooKeeper and Kubernetes HA services remain in 2.x. Two details matter for the upgrade:

  • The ZooKeeper path changed. In 1.20 job graphs are stored under high-availability.zookeeper.path.jobgraphs, default /jobgraphs. In 2.0 the option is high-availability.zookeeper.path.execution-plans, default /execution-plans, with the old key accepted as deprecated. Our recommendation: give each 2.x cluster its own high-availability.cluster-id, so it never tries to recover 1.x HA metadata.
  • The bundled ZooKeeper client. Flink 1.17 to 1.20 build against ZooKeeper 3.7.1 and 2.0 to 2.3 against 3.7.2, according to each release's build file. Both are on the 3.7 line, which reached end of life in February 2024, so the 2.x move does not take the shaded ZooKeeper off a scanner report.

On Kubernetes, the 2.x move is a good point to switch from ZooKeeper HA to Kubernetes HA; our Flink high availability guide covers how. The Flink and ZooKeeper page lists what each release bundles.

Upgrade plan

  1. Inventory every job: API (DataSet, DataStream, Table, SQL), language, sources and sinks with their connector versions, state size and serializers, and whether uids are set.
  2. Move 1.18 and 1.19 clusters to 1.20.5 with stop-with-savepoint and restore.
  3. Convert configuration to config.yaml with bin/migrate-config-file.sh and run 1.20 on it.
  4. Move the JVM to 17 on 1.20.
  5. Pick the 2.x minor that every connector you need supports.
  6. Port jobs off the removed APIs and rebuild them against 2.x and the 2.x connector artifacts.
  7. Prove state per job on a test cluster: restore, rebuild from source, or translate.
  8. Cut over job by job on a separate 2.x cluster, which the upgrade guide calls a shadow copy upgrade, and keep the 1.20 cluster running until the last job has moved.

Test plan

  • Run each ported job against a copy of production input and compare its output with the 1.20 job's output for the same window.
  • Restore every 1.20 savepoint you intend to carry, and confirm state sizes and key counts after restore.
  • Trigger checkpoints and savepoints on 2.x and restore from them.
  • Kill a JobManager and a TaskManager under load and confirm recovery through your HA service.
  • Exercise exactly-once sinks through a failure, and check for duplicates or gaps.
  • Check metrics, dashboards and REST clients; 2.0 removed some fields from REST responses.

Rollback

The upgrade guide documents compatibility in one direction only, and we know of no supported way to restore a 2.x savepoint on 1.x. Rollback therefore means returning to the 1.20 job and the last 1.20 savepoint taken before cut-over, and replaying input from the offsets that savepoint recorded. Keep that savepoint, the 1.20 cluster or its configuration, and the 1.20 job artifacts until the 2.x job has run through a full business cycle. Check that source retention covers the rollback window, or the replay will have a gap. Sinks that are not idempotent will see duplicates from the replay, so decide in advance how each one is cleaned up.

If you cannot upgrade yet

For many teams the blocker is not effort but a connector with no 2.x build, a Scala codebase, or state that cannot be rebuilt. OSSeva ships patched, signed Flink builds for 1.15 to 1.19 with the savepoint format unchanged, and replaces the end-of-life ZooKeeper shaded inside them, delivered through Maven, Docker and tarballs. Assure adds a ZooKeeper or Kubernetes HA review, a ZooKeeper ACL, SASL and exposure audit, an inventory of the removed 2.x APIs each job uses, a SOC 2 and HIPAA attestation package with VEX for scanner findings, and an upgrade plan to 1.20 LTS or 2.x with savepoint compatibility checks. Operate adds 24/7 checkpoint, backpressure and JobManager failover monitoring, a 15-minute P1 response, a named senior Flink engineer, savepoint management and job redeployment, and staged 2.x migration job by job. See Flink extended support and Apache Flink support.

Common questions

Can I restore a Flink 1.x savepoint on Flink 2.x?

Not with any guarantee. The 2.0 release notes say state compatibility between 1.x and 2.x is not guaranteed, and the savepoint table stops at 1.20. Test each job's restore on a copy, and plan to rebuild or translate state where it fails.

Should I upgrade to 1.20 before 2.x?

We recommend it. 1.20 is the LTS line, it restores savepoints from older 1.x releases, it shows deprecation warnings for what 2.0 removes, and it runs on config.yaml and Java 17.

Does Flink 2.x run on Java 8?

No. Flink 2.0 requires Java 11 or later, defaults to Java 17 and supports Java 21.

Is flink-conf.yaml still supported in Flink 2?

No. 2.0 reads only config.yaml. Convert with bin/migrate-config-file.sh, which ships from 1.19.

How long is Flink 1.20 supported?

The download page labels 1.20 LTS without an end date. FLIP-458 proposes two years of bug fixes and security updates.

Tags

Apache FlinkFlink 2.0UpgradeSavepointsEnd of Life

Ready to get your open source under control?

Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.