Back to blog

// OSSeva Blog

Migration

Upgrading Apache Hadoop 2.10 to 3.3 or 3.4: Java, Ports, Shell Scripts, Rolling Upgrade Limits and the Stack Above HDFS

Matt Reynolds12 min read

The short answer

Hadoop 2.10 is the last 2.x line, and its last release is 2.10.2, from 31 May 2022. The Hadoop community has never formally retired 2.10: its end-of-life page lists 2.0 to 2.9 and 3.0 to 3.1, and the HBase reference guide notes that Hadoop 2.x has not been declared end of life even though no release has followed for years. In practice nobody ships 2.x fixes upstream, and HBase's guide recommends Hadoop 3. The Hadoop end-of-life tracker has every line.

The realistic targets are Hadoop 3.4, with 3.4.3 from 24 February 2026, and 3.3, whose last release is 3.3.6 from 23 June 2023. Hadoop 3.5.0 shipped on 2 April 2026 and requires Java 17 on the server side, which is a separate decision. Our recommendation is 3.4: 3.3 has the same Java support but has gone more than three years without a release.

Three facts shape the plan:

  • Java. Hadoop 3.3 and later run on Java 8 or 11. The 3.4 release notes recommend moving production beyond Java 8 and say JDK 8 support will be dropped in a future 3.4.x release.
  • Plan on rollback, not downgrade. The NameNode layout version moves from -63 in 2.10.2 to -67 in 3.4.3, and the rolling upgrade guide says a downgrade needs unchanged layout versions. Rollback, which restores the pre-upgrade state and needs downtime, is the documented way back.
  • Rolling upgrade is not guaranteed. Hadoop's compatibility policy promises wire compatibility within a major version. The JIRA issue for supporting rolling upgrades from 2.x to 3.x, HDFS-11096, is still open.

Release lines and Java

LineLatest releaseStatusJava
2.102.10.2 (31 May 2022)Not formally retired, no release since 20227 and 8
3.23.2.4 (22 Jul 2022)End of life8
3.33.3.6 (23 Jun 2023)Not formally retired, no release since 20238, or 11 at runtime
3.43.4.3 (24 Feb 2026)Active8, or 11 at runtime; beyond 8 recommended
3.53.5.0 (2 Apr 2026)Active17 required on servers; 17 or 21 for clients

Java support comes from the Hadoop Java versions page and the 3.4 and 3.5 release notes. The Java versions page also says Hadoop must still be compiled with Java 8; Java 11 is supported at runtime only. If you build your own Hadoop, keep the build on JDK 8 for 3.3 and 3.4.

A sensible order is to move the JVM to 11 after the cluster is on 3.4 and stable, not in the same window as the major upgrade. That is our recommendation, not an upstream rule.

What changes in Hadoop 3

The Hadoop 3.0.0 release overview lists the major changes. These are the ones that matter for an existing 2.10 cluster.

Default ports

Several default ports moved out of the Linux ephemeral range (HDFS-9427, HADOOP-12811). Anything that relies on the defaults, including firewall rules, load balancers, monitoring and scripts, needs updating unless you pin the old values in configuration.

ServiceHadoop 2Hadoop 3
NameNode HTTP / HTTPS50070 / 504709870 / 9871
Secondary NameNode HTTP / HTTPS50090 / 500919868 / 9869
DataNode data transfer500109866
DataNode IPC500209867
DataNode HTTP / HTTPS50075 / 504759864 / 9865
KMS160009600
NameNode RPC80208020 (9820 only in 3.0.0)

The NameNode RPC port briefly moved to 9820 in 3.0.0 and went back to 8020 in 3.0.1 (HDFS-12990), so it does not change for a move to 3.3 or 3.4.

Shell scripts

The shell scripts were rewritten (HADOOP-9902). The release notes warn that some changes may break existing installations:

  • Daemons start with hdfs --daemon start namenode, yarn --daemon start resourcemanager and so on. hadoop-daemon.sh still works but prints a deprecation warning and calls the new form.
  • Every subsystem now reads hadoop-env.sh. HDFS_*, YARN_* and MAPRED_* variables override their HADOOP_* equivalents, and the *-env.sh files no longer append to values set on the command line.
  • Daemon settings follow a (command)_(subcommand)_OPTS pattern, and (command)_(subcommand)_USER variables such as HDFS_NAMENODE_USER check which user runs each daemon.
  • HADOOP_HEAPSIZE is deprecated in favour of HADOOP_HEAPSIZE_MAX and HADOOP_HEAPSIZE_MIN, and the default heap sizes were removed so the JVM sizes itself from host memory (HADOOP-10950). The release note says to set HADOOP_HEAPSIZE_MAX="1g" to restore the old default.
  • The slaves file is deprecated in favour of workers (HADOOP-13209).
  • Bash 3 or later is required, and support for HADOOP_MASTER and its rsync code was removed.

Our recommendation: rebuild hadoop-env.sh, yarn-env.sh and mapred-env.sh from the 3.x templates rather than copying the 2.x files, then port your settings across. The same applies to systemd units or init scripts that call the old daemon scripts.

Classpath isolation

The 2.x hadoop-client artifact puts Hadoop's dependencies on the application classpath. Hadoop 3 adds hadoop-client-api and hadoop-client-runtime, which shade those dependencies into single JARs (HADOOP-11804). Applications with dependency clashes can move to the shaded client to stop Hadoop's versions leaking into theirs.

Erasure coding

HDFS can store data with Reed-Solomon erasure coding instead of three replicas. The documentation gives RS(10,4) a 1.4x storage overhead against 3x for replication. The built-in policies are RS-3-2-1024k, RS-6-3-1024k, RS-10-4-1024k, RS-LEGACY-6-3-1024k and XOR-2-1-1024k; only RS-6-3-1024k is enabled by default, and policies are applied per directory with hdfs ec -setPolicy.

Erasure coding only applies where you set a policy, so an upgraded cluster behaves as before. When you do adopt it, mind the limits: hflush, hsync, concat, setReplication, truncate and append are not supported on erasure coded files, and hflush and hsync silently do nothing. Never put HBase write-ahead logs, or anything else that relies on hsync, in an erasure coded directory. We also leave erasure coding off until the rollback window has closed.

HDFS and YARN features

  • More than two NameNodes. HA can run several standby NameNodes, for example three NameNodes and five JournalNodes to tolerate two failures.
  • Intra-DataNode balancer. hdfs diskbalancer evens data across disks inside a DataNode.
  • Router-based federation adds a server-side mount table across namespaces.
  • YARN gained opportunistic containers, user-defined resource types such as GPUs, an API for Capacity Scheduler queue configuration and Timeline Service v.2. Existing schedulers and queue configuration carry over, but review them against the 3.x defaults.

MapReduce applications are covered by a stronger promise than the rest. Hadoop's compatibility policy says binary compatibility for the org.apache.hadoop.mapred and org.apache.hadoop.mapreduce APIs is supported across major releases.

Rolling upgrade from 2.x to 3.x

HDFS supports rolling upgrades on HA clusters from Hadoop 2.4.0 onwards, using hdfs dfsadmin -rollingUpgrade prepare, a NameNode failover and DataNodes upgraded a rack at a time. Whether that procedure works across the major version boundary is less settled:

  • The compatibility policy only requires server-server wire compatibility within a major version.
  • HDFS-11096, "Support rolling upgrade between 2.x and 3.x", is still open with a patch attached.
  • Several of its pieces are fixed. HDFS-11188 set the minimum supported NameNode and DataNode versions back to 2.x in 3.0.0-alpha2, and HDFS-12151 let Hadoop 2 clients write to Hadoop 3 DataNodes in 3.0.0-beta1. HDFS-13596, a NameNode restart failure after a rolling upgrade from 2.x, was fixed in 3.1.3, 3.2.1 and 3.3.0. HDFS-14509, a DataNode token error during a NameNode upgrade from 2.x, was fixed in 2.10.0, 3.1.4, 3.2.2 and 3.3.0. YARN-6142, the YARN side, was resolved in 3.0.0.

Our recommendation: start from 2.10.2, which carries the 2.x half of HDFS-14509, and plan a downtime upgrade as the default. Attempt a rolling upgrade only after it has worked end to end on a staging copy with production configuration, and keep the downtime plan ready if it does not. For most clusters, a few hours of planned downtime is cheaper than an untested mixed-version state.

The downtime procedure from the HDFS user guide:

  1. Finalize any previous upgrade with hdfs dfsadmin -finalizeUpgrade, since HDFS keeps only one backup.
  2. Run hdfs dfsadmin -saveNamespace, which the guide recommends before upgrading.
  3. Stop the cluster and install the new version.
  4. Start HDFS with start-dfs.sh -upgrade. If the NameNode reports reserved paths such as .reserved or .snapshot, use -upgrade -renameReserved.
  5. Run on the new version until you trust it, then finalize. Until then, deleting pre-upgrade files does not free space on the DataNodes.

The stack above HDFS

A Hadoop major upgrade is usually decided by what runs on top of it.

HBase

The HBase reference guide keeps a Hadoop matrix for every HBase line:

HBaseHadoop 3.3Hadoop 3.4
1.5 to 1.7Not listed for any Hadoop 3 releaseNot listed
2.2Not listed (3.1.1+ only)Not listed
2.4SupportedNot listed
2.53.3.2 and laterFrom HBase 2.5.11
2.63.3.5 and laterFrom HBase 2.6.2

So an HBase 1.x or 2.2 cluster has to upgrade HBase as part of this project, and a move to Hadoop 3.4 needs HBase 2.5.11 or 2.6.2 or later. HBase 2.4, which supports Hadoop 3.3, reached end of maintenance in May 2024; see the HBase 2.4 end-of-life page and HBase support options. Replace the Hadoop JARs in HBase's lib directory with the ones from your cluster, as the guide requires.

Hive

The Hive download page lists Hive 2.3.x as working with Hadoop 2 and Hive 3.x and 4.x as working with Hadoop 3. A Hive 2.3 warehouse therefore moves Hive and Hadoop together, and since Hive 3 reached end of life in October 2024, the destination is Hive 4. Our Hive 3 to 4 upgrade guide covers the metastore schema changes.

ZooKeeper

Automatic NameNode failover and ResourceManager HA run on ZooKeeper. Hadoop 2.10.2 bundles ZooKeeper 3.4.14, from a line that reached end of life in 2020. The HDFS rolling upgrade guide leaves JournalNodes and ZooKeeper out of its procedure and warns that upgrading them may need downtime, so move the ensemble in its own window. See Hadoop and ZooKeeper and the ZooKeeper 3.4 to 3.8 upgrade guide.

Upgrade plan

  1. Inventory every service on the cluster and its version: HDFS, YARN, MapReduce, HBase, Hive, Spark, Tez, Oozie, the ZooKeeper ensemble and every client that connects from outside.
  2. Pick the target set together: Hadoop 3.4.x, an HBase release that supports it, Hive 4 and matching Tez and Spark builds.
  3. Bring 2.x to 2.10.2 if it is on an older 2.x line.
  4. Rebuild configuration from the 3.x templates, port your settings, and decide whether to keep the old ports or adopt the new ones.
  5. Update automation: daemon start commands, workers files, heap settings, monitoring endpoints and firewall rules.
  6. Rehearse on a staging cluster built from a copy of production metadata, including the rollback.
  7. Upgrade HDFS with -upgrade, then YARN and MapReduce, then the services above them.
  8. Finalize only after a full business cycle on the new version.

Test plan

  • Run representative MapReduce, Hive, Spark and Tez jobs and compare outputs with the 2.x results.
  • Fail over the NameNode and the ResourceManager under load.
  • Check that every external client, ETL tool and monitoring probe reaches the services on the expected ports.
  • Exercise Kerberos authentication, delegation tokens and any Ranger or Sentry policies.
  • Read and write through HttpFS or WebHDFS if anything uses them.
  • Confirm applications built against hadoop-client still start, and try the shaded client where dependencies clash.
  • Run the balancer and confirm DataNode decommissioning works.

Rollback

The documented way back from 3.x to 2.x is rollback, not downgrade. Since Hadoop 2.8, a NameNode in a rolling upgrade keeps the prior layout version and refuses new features until finalize (HDFS-8432), which is meant to make downgrades possible for most layout changes. The rolling upgrade guide still says a downgrade needs unchanged layout versions, though, and the 2.x to 3.x rolling path is unfinished, so we would not depend on a downgrade. Rollback restores the pre-upgrade state and drops every change made since. It needs the cluster to be stopped, the 2.x software reinstalled, hdfs namenode -rollback run and the cluster started with start-dfs.sh -rollback. It is only possible until you finalize, so keep the rollback window open long enough to be sure, and keep the 2.x packages and configuration on every node until then. HBase, Hive and other services on top need their own rollback plans, and a Hive metastore schema upgrade has to be reversed from a database backup.

If you cannot upgrade yet

A Hadoop major upgrade moves storage, compute and every job that touches them, and it often waits on an HBase, Hive or vendor platform decision. OSSeva ships patched, signed builds for Hadoop 2.8 to 3.3, and for CDH 5 and 6 and HDP 2.6 and 3.1 clusters past Cloudera support, with the ZooKeeper behind NameNode and ResourceManager failover patched too. Builds come as tarballs, RPMs and parcels. Assure adds a NameNode, ResourceManager and ZooKeeper HA review, a Kerberos, Ranger or Sentry policy and exposure audit, a version map for every service on the cluster, a SOC 2 and HIPAA attestation package with VEX for auditors and a migration plan to Hadoop 3.4 or 3.5, CDP or a cloud platform. Operate adds 24/7 HDFS capacity, NameNode and YARN queue monitoring, a 15-minute P1 response, a named senior Hadoop engineer, failover testing and ZooKeeper quorum management, and execution of the upgrade itself. See Hadoop extended support, Apache Hadoop support, Hadoop on an end-of-life version and the CDH 6 and HDP 3.1 end-of-life pages.

Common questions

Is Hadoop 2.10 end of life?

Not formally. The community's end-of-life list stops at 2.9, but the last 2.10 release was 2.10.2 in May 2022.

Can I do a rolling upgrade from Hadoop 2 to Hadoop 3?

It is not guaranteed. Hadoop promises wire compatibility within a major version, and the umbrella issue for 2.x to 3.x rolling upgrades is still open, although several fixes for it have shipped. Rehearse it, and plan a downtime upgrade as the fallback.

Can I downgrade HDFS from 3.x back to 2.x?

Do not plan on it. The rolling upgrade guide says a downgrade needs unchanged NameNode and DataNode layout versions, and the NameNode layout changes between 2.10 and 3.x. Rollback to the pre-upgrade state works until you finalize, losing changes made since.

Does Hadoop 3.4 run on Java 11?

Yes, at runtime. Hadoop 3.3 and later support Java 8 and Java 11, and the 3.4 release notes recommend moving beyond Java 8. Hadoop 3.5 requires Java 17 on servers.

Did the NameNode port change in Hadoop 3?

The NameNode web UI moved from 50070 to 9870. The RPC port is still 8020; only 3.0.0 used 9820.

Tags

Apache HadoopHDFSYARNUpgradeEnd of Life

Ready to get your open source under control?

Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.