Teams running Hadoop eventually ask whether they should keep HDFS or move to Apache Ozone, the object store built in the Hadoop project as HDFS's successor for large clusters. The question is often framed as old versus new, which is unhelpful. Both store bytes on commodity disks with replication or erasure coding, both are reachable through the Hadoop FileSystem API, and both run Spark, Hive and Impala. They differ in where metadata lives, which file semantics they guarantee, and which clients they serve natively.
This article compares them on those axes, shows how to probe the differences from code instead of trusting documentation, works through a sizing example, and gives a workload-by-workload verdict. It does not cover migration mechanics; the step-by-step move of a warehouse directory is in the Ozone operations playbook, and Ozone's internals are in the Ozone architecture article.
Two shapes of the same job
HDFS has one active NameNode per namespace. It holds every directory, file and block record in JVM heap, persists changes to an edit log, and learns where block replicas live from block reports sent by every DataNode. That design gives very fast metadata operations and strong POSIX-like semantics for renames and appends, and it has a single, well-known ceiling: the heap of one process. Federation adds more namespaces, each with its own NameNode, but each namespace still has the same ceiling.
Ozone splits that job in two. The Ozone Manager owns the namespace of volumes, buckets and keys and stores it in RocksDB on local disk, replicated across OM nodes with Apache Ratis, a Raft implementation. The Storage Container Manager owns storage: blocks are grouped into containers, 5 GB by default, and datanodes report containers rather than individual blocks. Metadata is no longer bound by heap, and reports grow with the number of containers rather than the number of blocks.
Ozone also speaks two protocols. Hadoop clients use the ofs:// rooted filesystem, where paths look like ofs://service/volume/bucket/key. Everything else can use the built-in S3 gateway. HDFS speaks only its own protocol plus WebHDFS; S3 access needs another product in front.
Scale: the heap ceiling and what replaces it
The usual rule of thumb for HDFS is that each file, directory and block object costs on the order of 150 bytes of NameNode heap. Treat it as an estimate, since real overhead depends on path lengths, replication and JVM settings, but it is good enough for planning. A namespace of 300 million files averaging 1.2 blocks each is about 660 million objects, which is around 100 GB of heap before headroom. That is possible with careful GC tuning, but every restart replays a large image and edit log, and every full GC on that heap is a cluster-wide stall.
Small files make it worse. A 100 KB file costs the NameNode the same as a 100 MB file, so a data lake of logs and images hits the heap ceiling long before the disks fill. The small files article covers the HDFS-side mitigations, and the NameNode deep dive covers heap sizing.
Ozone's namespace size is bounded by OM disk and RocksDB performance rather than heap, so billions of keys are a design goal rather than a crisis. The cost moves elsewhere: every namespace write is a Ratis consensus round across OM nodes, so metadata latency is a little higher than an in-memory NameNode, and RocksDB compaction becomes something you monitor.
Semantics: where they are not interchangeable
This is the part that breaks applications, so check it against your workloads rather than assuming the systems are interchangeable. Ozone buckets have a layout fixed at creation: File System Optimized (FSO) stores a real directory tree, and Object Store (OBS) stores flat keys for S3 clients. Most filesystem guarantees only hold for FSO.
| Behaviour | HDFS | Ozone FSO bucket | Ozone OBS bucket |
|---|---|---|---|
| Atomic directory rename | Yes | Yes, a metadata operation | No directories; S3 rename is copy plus delete |
| Append to an existing file | Yes | Not supported by the Ozone filesystem clients; probe it | No |
| hflush / hsync durability | Yes | hsync and lease recovery added for HBase in 2.0 behind a layout version | No |
| S3 API | No (needs a separate gateway) | Yes, with caveats on path-shaped keys | Yes, the intended client |
| Hadoop FileSystem API | Yes | Yes, via ofs:// | Not the intended client |
| Snapshots | Directory snapshots | Bucket snapshots with snapshot diff | Bucket snapshots |
| Quotas | Name and space quota per directory | Namespace and space quota per volume and bucket | Same as FSO |
| Erasure coding | Per-directory EC policies | Per-bucket or per-key EC replication config | Same as FSO |
Two rows deserve emphasis. Append matters for anything that writes log-style files in place, such as Flume-style ingest or some streaming sinks. Durable flush matters for write-ahead logs, which is why HBase on Ozone only became a supported story with Ozone 2.0.0, released in April 2025, and a dedicated layout version. Treat HBase on Ozone as newer and less battle-tested than HBase on HDFS.
Probe the differences from code
Documentation drifts between releases, but the Hadoop FileSystem API can tell you what a path supports at runtime. hasPathCapability answers questions about a path, and StreamCapabilities answers questions about an open stream. A small probe run against each candidate filesystem settles arguments quickly:
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.*;
public class FsProbe {
public static void main(String[] args) throws Exception {
Path base = new Path(args[0]); // hdfs://nn1/tmp/probe or ofs://ozone1/vol1/fso1/probe
FileSystem fs = base.getFileSystem(new Configuration());
String[] caps = {
CommonPathCapabilities.FS_APPEND,
CommonPathCapabilities.FS_CONCAT,
CommonPathCapabilities.FS_SNAPSHOTS,
CommonPathCapabilities.FS_ACLS,
CommonPathCapabilities.FS_PERMISSIONS,
};
for (String cap : caps) {
System.out.printf("%-34s %s%n", cap, fs.hasPathCapability(base, cap));
}
Path f = new Path(base, "hsync-test");
try (FSDataOutputStream out = fs.create(f, true)) {
out.writeBytes("hello\n");
System.out.println("hflush: " + out.hasCapability(StreamCapabilities.HFLUSH));
System.out.println("hsync: " + out.hasCapability(StreamCapabilities.HSYNC));
out.hsync();
}
Path a = new Path(base, "dirA"), b = new Path(base, "dirB");
fs.mkdirs(new Path(a, "child"));
long t0 = System.nanoTime();
boolean ok = fs.rename(a, b);
System.out.printf("rename dir: %s in %d us%n", ok, (System.nanoTime() - t0) / 1000);
fs.delete(b, true);
fs.delete(f, false);
}
}Run it with the same client JARs and configuration your jobs use, against an HDFS path, an FSO bucket and an OBS bucket. A capability that reports false is a hard constraint for any job that needs it. A capability that reports true still needs a load test, because supported and fast are different claims.
Workload by workload
| Workload | Better fit | Why |
|---|---|---|
| Spark, Hive, Impala over Parquet or ORC tables | Either | Both expose locality and the FileSystem API; Iceberg or Hive tables work on both. Choose on namespace size and S3 needs. |
| Lake with hundreds of millions of small files | Ozone | Namespace in RocksDB, container reports instead of block reports. |
| Mixed Hadoop and S3-native tools (Python, Trino via S3, ML loaders) | Ozone | One copy of data reachable by ofs:// and S3. |
| HBase | HDFS today | Mature hflush/hsync and recovery; Ozone support is recent and gated by a layout version. |
| Log-style writers that append in place | HDFS | Ozone filesystem clients do not support append. |
| Small cluster, under about 100 million objects, no S3 need | HDFS | No ceiling in sight; fewer moving parts; existing skills. |
| Very large single namespace growing fast | Ozone | Avoids splitting into federated namespaces just to stay under heap. |
Notice what is missing: raw throughput. For large sequential reads and writes both systems are bounded mostly by disks and network, and differences depend more on configuration than on design. Do not choose on a vendor benchmark; run your own, as below.
Worked example: a 420-million-file cluster
Consider a 120-node cluster with 9 PB raw capacity, 3x replication on hot data and EC on cold, holding 420 million files that grow by 8 million a month, mostly 2 MB event files written by a streaming job and compacted nightly into large Parquet files. The NameNode has a 160 GB heap and full GCs are already appearing during the compaction window. Users want to read the same data from a Python training pipeline that only speaks S3.
With the 150-byte estimate, 420 million files plus about 450 million blocks is roughly 130 GB, so the heap is near its working limit and growth adds about 2.5 GB a month. Options on HDFS are aggressive compaction to cut file count, or federation to split the namespace, plus a separate S3 gateway product for the Python users. Options on Ozone are an FSO bucket per dataset for Spark, which keeps atomic renames for the job commit protocol, with S3 access for the training pipeline on the same data.
The streaming job is the risk. If it appends in place, it must be changed to write new files and let compaction merge them, because append is not available. If it relies on hsync for durability, test it on the Ozone version you will deploy. In this scenario the job writes new files every five minutes, so the verdict is Ozone for the event and training data, with HBase, if any, staying on HDFS until its Ozone support has more production mileage. The migration itself then follows the playbook linked above.
A benchmark plan that answers the real question
- Replay real metadata load. Capture a day of NameNode audit logs, and replay the mix of
create,open,listStatus,renameanddeleteat production rate against both systems. Ozone ships a load generator called Freon for synthetic key and file load; use it for stress, not as a substitute for your mix. - Run a real job. Take your heaviest Spark job, point it at an FSO bucket through
ofs://, and compare wall time, task-level read throughput and job commit time. Commit time exercises rename and listing, which is where layouts differ most. - Measure tail latency, not averages. Record p99 and p999 for
listStatuson large directories and for small-file reads; these drive interactive query feel. - Break things. Kill a leader Ozone Manager and a datanode mid-job and measure recovery. Do the same with HDFS active NameNode failover. The numbers you need are seconds of unavailability and whether jobs survive.
- Check S3 behaviour with the real client library, including multipart uploads and listing with delimiters on path-shaped keys.
Failure modes and trade-offs
Most bad Ozone outcomes trace to a few decisions. Choosing OBS for a bucket that Spark writes with the default commit protocol turns directory renames into slow, non-atomic copies, and the layout cannot be changed after creation, so the bucket must be recreated. Undersizing OM disks or putting RocksDB on slow storage makes every metadata call slow, because the namespace now lives on disk. Running Ratis on noisy networks causes leader elections that look like intermittent client timeouts. And treating Ozone as a drop-in HDFS for every client surfaces append and hsync gaps at the worst time.
HDFS failures are better known: NameNode heap pressure and long GC pauses, block-report storms after mass restarts, slow startup on large images, and small-file growth that nobody owns. Its erasure coding, covered in the HDFS erasure coding article, helps capacity but not the namespace.
The honest trade-off: HDFS is simpler, older and tightly integrated, with a hard metadata ceiling. Ozone removes the ceiling and adds S3, at the cost of more services to run (OM, SCM, Recon, S3 gateway), a Raft-replicated metadata path, and a smaller pool of operators who have run it at scale.
What to do next
- Measure your namespace: total files, directories and blocks, monthly growth, and NameNode heap use. Estimate months until your heap ceiling.
- List every client and its required semantics: append, hflush/hsync, atomic directory rename, S3 access.
- Run the capability probe above against HDFS, an FSO bucket and an OBS bucket with your production client JARs.
- Classify each workload with the table above; anything needing append stays on HDFS or is rewritten to write new files.
- If Ozone wins any workload, stand up a small cluster, create FSO buckets for Hadoop clients and OBS buckets only for S3-only data.
- Replay a day of audit-log metadata load and one heavy Spark job; record p99 latencies and commit times on both systems.
- Fail an OM leader and a datanode during the test and record recovery time and job survival.
- Decide per dataset, not per cluster, and follow the operations playbook for the migration itself.