Teams running Hadoop eventually ask whether they should keep HDFS or move to Apache Ozone, the object store built in the Hadoop project as HDFS's successor for large clusters. The question is often framed as old versus new, which is unhelpful. Both store bytes on commodity disks with replication or erasure coding, both are reachable through the Hadoop FileSystem API, and both run Spark, Hive and Impala. They differ in where metadata lives, which file semantics they guarantee, and which clients they serve natively.

This article compares them on those axes, shows how to probe the differences from code instead of trusting documentation, works through a sizing example, and gives a workload-by-workload verdict. It does not cover migration mechanics; the step-by-step move of a warehouse directory is in the Ozone operations playbook, and Ozone's internals are in the Ozone architecture article.

Advertisement

Two shapes of the same job

HDFS has one active NameNode per namespace. It holds every directory, file and block record in JVM heap, persists changes to an edit log, and learns where block replicas live from block reports sent by every DataNode. That design gives very fast metadata operations and strong POSIX-like semantics for renames and appends, and it has a single, well-known ceiling: the heap of one process. Federation adds more namespaces, each with its own NameNode, but each namespace still has the same ceiling.

Ozone splits that job in two. The Ozone Manager owns the namespace of volumes, buckets and keys and stores it in RocksDB on local disk, replicated across OM nodes with Apache Ratis, a Raft implementation. The Storage Container Manager owns storage: blocks are grouped into containers, 5 GB by default, and datanodes report containers rather than individual blocks. Metadata is no longer bound by heap, and reports grow with the number of containers rather than the number of blocks.

HDFSNameNode (JVM heap)namespace + block map + leasesDataNodeblocksDataNodeblocksblock reportsClienthdfs:// FileSystem APImetadatadataOzoneOzone Managernamespace in RocksDB, Ratis HAStorage Container Mgrcontainers, pipelinesDatanodecontainers of blocksDatanodecontainers of blockscontainer reportsClientofs:// FileSystem APIS3 GatewayS3 RESTkeysdataHDFS keeps the whole namespace and block map in one heap; Ozone splits names (OM) from storage (SCM)and tracks containers, not individual blocks, so datanode reports stay small as data grows
Where metadata lives in each system. The scaling limit of HDFS is the NameNode heap; Ozone moves the namespace to RocksDB under the Ozone Manager and groups blocks into containers managed by SCM.

Ozone also speaks two protocols. Hadoop clients use the ofs:// rooted filesystem, where paths look like ofs://service/volume/bucket/key. Everything else can use the built-in S3 gateway. HDFS speaks only its own protocol plus WebHDFS; S3 access needs another product in front.

Scale: the heap ceiling and what replaces it

The usual rule of thumb for HDFS is that each file, directory and block object costs on the order of 150 bytes of NameNode heap. Treat it as an estimate, since real overhead depends on path lengths, replication and JVM settings, but it is good enough for planning. A namespace of 300 million files averaging 1.2 blocks each is about 660 million objects, which is around 100 GB of heap before headroom. That is possible with careful GC tuning, but every restart replays a large image and edit log, and every full GC on that heap is a cluster-wide stall.

Small files make it worse. A 100 KB file costs the NameNode the same as a 100 MB file, so a data lake of logs and images hits the heap ceiling long before the disks fill. The small files article covers the HDFS-side mitigations, and the NameNode deep dive covers heap sizing.

Ozone's namespace size is bounded by OM disk and RocksDB performance rather than heap, so billions of keys are a design goal rather than a crisis. The cost moves elsewhere: every namespace write is a Ratis consensus round across OM nodes, so metadata latency is a little higher than an in-memory NameNode, and RocksDB compaction becomes something you monitor.

Advertisement

Semantics: where they are not interchangeable

This is the part that breaks applications, so check it against your workloads rather than assuming the systems are interchangeable. Ozone buckets have a layout fixed at creation: File System Optimized (FSO) stores a real directory tree, and Object Store (OBS) stores flat keys for S3 clients. Most filesystem guarantees only hold for FSO.

BehaviourHDFSOzone FSO bucketOzone OBS bucket
Atomic directory renameYesYes, a metadata operationNo directories; S3 rename is copy plus delete
Append to an existing fileYesNot supported by the Ozone filesystem clients; probe itNo
hflush / hsync durabilityYeshsync and lease recovery added for HBase in 2.0 behind a layout versionNo
S3 APINo (needs a separate gateway)Yes, with caveats on path-shaped keysYes, the intended client
Hadoop FileSystem APIYesYes, via ofs://Not the intended client
SnapshotsDirectory snapshotsBucket snapshots with snapshot diffBucket snapshots
QuotasName and space quota per directoryNamespace and space quota per volume and bucketSame as FSO
Erasure codingPer-directory EC policiesPer-bucket or per-key EC replication configSame as FSO

Two rows deserve emphasis. Append matters for anything that writes log-style files in place, such as Flume-style ingest or some streaming sinks. Durable flush matters for write-ahead logs, which is why HBase on Ozone only became a supported story with Ozone 2.0.0, released in April 2025, and a dedicated layout version. Treat HBase on Ozone as newer and less battle-tested than HBase on HDFS.

Probe the differences from code

Documentation drifts between releases, but the Hadoop FileSystem API can tell you what a path supports at runtime. hasPathCapability answers questions about a path, and StreamCapabilities answers questions about an open stream. A small probe run against each candidate filesystem settles arguments quickly:

import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.*;

public class FsProbe {
    public static void main(String[] args) throws Exception {
        Path base = new Path(args[0]); // hdfs://nn1/tmp/probe or ofs://ozone1/vol1/fso1/probe
        FileSystem fs = base.getFileSystem(new Configuration());
        String[] caps = {
            CommonPathCapabilities.FS_APPEND,
            CommonPathCapabilities.FS_CONCAT,
            CommonPathCapabilities.FS_SNAPSHOTS,
            CommonPathCapabilities.FS_ACLS,
            CommonPathCapabilities.FS_PERMISSIONS,
        };
        for (String cap : caps) {
            System.out.printf("%-34s %s%n", cap, fs.hasPathCapability(base, cap));
        }
        Path f = new Path(base, "hsync-test");
        try (FSDataOutputStream out = fs.create(f, true)) {
            out.writeBytes("hello\n");
            System.out.println("hflush: " + out.hasCapability(StreamCapabilities.HFLUSH));
            System.out.println("hsync:  " + out.hasCapability(StreamCapabilities.HSYNC));
            out.hsync();
        }
        Path a = new Path(base, "dirA"), b = new Path(base, "dirB");
        fs.mkdirs(new Path(a, "child"));
        long t0 = System.nanoTime();
        boolean ok = fs.rename(a, b);
        System.out.printf("rename dir: %s in %d us%n", ok, (System.nanoTime() - t0) / 1000);
        fs.delete(b, true);
        fs.delete(f, false);
    }
}

Run it with the same client JARs and configuration your jobs use, against an HDFS path, an FSO bucket and an OBS bucket. A capability that reports false is a hard constraint for any job that needs it. A capability that reports true still needs a load test, because supported and fast are different claims.

Workload by workload

WorkloadBetter fitWhy
Spark, Hive, Impala over Parquet or ORC tablesEitherBoth expose locality and the FileSystem API; Iceberg or Hive tables work on both. Choose on namespace size and S3 needs.
Lake with hundreds of millions of small filesOzoneNamespace in RocksDB, container reports instead of block reports.
Mixed Hadoop and S3-native tools (Python, Trino via S3, ML loaders)OzoneOne copy of data reachable by ofs:// and S3.
HBaseHDFS todayMature hflush/hsync and recovery; Ozone support is recent and gated by a layout version.
Log-style writers that append in placeHDFSOzone filesystem clients do not support append.
Small cluster, under about 100 million objects, no S3 needHDFSNo ceiling in sight; fewer moving parts; existing skills.
Very large single namespace growing fastOzoneAvoids splitting into federated namespaces just to stay under heap.

Notice what is missing: raw throughput. For large sequential reads and writes both systems are bounded mostly by disks and network, and differences depend more on configuration than on design. Do not choose on a vendor benchmark; run your own, as below.

Worked example: a 420-million-file cluster

Consider a 120-node cluster with 9 PB raw capacity, 3x replication on hot data and EC on cold, holding 420 million files that grow by 8 million a month, mostly 2 MB event files written by a streaming job and compacted nightly into large Parquet files. The NameNode has a 160 GB heap and full GCs are already appearing during the compaction window. Users want to read the same data from a Python training pipeline that only speaks S3.

With the 150-byte estimate, 420 million files plus about 450 million blocks is roughly 130 GB, so the heap is near its working limit and growth adds about 2.5 GB a month. Options on HDFS are aggressive compaction to cut file count, or federation to split the namespace, plus a separate S3 gateway product for the Python users. Options on Ozone are an FSO bucket per dataset for Spark, which keeps atomic renames for the job commit protocol, with S3 access for the training pipeline on the same data.

The streaming job is the risk. If it appends in place, it must be changed to write new files and let compaction merge them, because append is not available. If it relies on hsync for durability, test it on the Ozone version you will deploy. In this scenario the job writes new files every five minutes, so the verdict is Ozone for the event and training data, with HBase, if any, staying on HDFS until its Ozone support has more production mileage. The migration itself then follows the playbook linked above.

A benchmark plan that answers the real question

  1. Replay real metadata load. Capture a day of NameNode audit logs, and replay the mix of create, open, listStatus, rename and delete at production rate against both systems. Ozone ships a load generator called Freon for synthetic key and file load; use it for stress, not as a substitute for your mix.
  2. Run a real job. Take your heaviest Spark job, point it at an FSO bucket through ofs://, and compare wall time, task-level read throughput and job commit time. Commit time exercises rename and listing, which is where layouts differ most.
  3. Measure tail latency, not averages. Record p99 and p999 for listStatus on large directories and for small-file reads; these drive interactive query feel.
  4. Break things. Kill a leader Ozone Manager and a datanode mid-job and measure recovery. Do the same with HDFS active NameNode failover. The numbers you need are seconds of unavailability and whether jobs survive.
  5. Check S3 behaviour with the real client library, including multipart uploads and listing with delimiters on path-shaped keys.

Failure modes and trade-offs

Most bad Ozone outcomes trace to a few decisions. Choosing OBS for a bucket that Spark writes with the default commit protocol turns directory renames into slow, non-atomic copies, and the layout cannot be changed after creation, so the bucket must be recreated. Undersizing OM disks or putting RocksDB on slow storage makes every metadata call slow, because the namespace now lives on disk. Running Ratis on noisy networks causes leader elections that look like intermittent client timeouts. And treating Ozone as a drop-in HDFS for every client surfaces append and hsync gaps at the worst time.

HDFS failures are better known: NameNode heap pressure and long GC pauses, block-report storms after mass restarts, slow startup on large images, and small-file growth that nobody owns. Its erasure coding, covered in the HDFS erasure coding article, helps capacity but not the namespace.

The honest trade-off: HDFS is simpler, older and tightly integrated, with a hard metadata ceiling. Ozone removes the ceiling and adds S3, at the cost of more services to run (OM, SCM, Recon, S3 gateway), a Raft-replicated metadata path, and a smaller pool of operators who have run it at scale.

What to do next

  1. Measure your namespace: total files, directories and blocks, monthly growth, and NameNode heap use. Estimate months until your heap ceiling.
  2. List every client and its required semantics: append, hflush/hsync, atomic directory rename, S3 access.
  3. Run the capability probe above against HDFS, an FSO bucket and an OBS bucket with your production client JARs.
  4. Classify each workload with the table above; anything needing append stays on HDFS or is rewritten to write new files.
  5. If Ozone wins any workload, stand up a small cluster, create FSO buckets for Hadoop clients and OBS buckets only for S3-only data.
  6. Replay a day of audit-log metadata load and one heavy Spark job; record p99 latencies and commit times on both systems.
  7. Fail an OM leader and a datanode during the test and record recovery time and job survival.
  8. Decide per dataset, not per cluster, and follow the operations playbook for the migration itself.
Key takeaway: HDFS keeps the namespace and block map in one NameNode heap, which makes it fast and semantically rich but caps the number of objects. Ozone moves the namespace to RocksDB under a Raft-replicated Ozone Manager, groups blocks into containers and adds a native S3 gateway, but FSO and OBS buckets differ and append is not available. Probe capabilities from code, choose per workload, and benchmark with your own metadata mix.