Installing Apache Ozone takes an afternoon. Running it for years is a different job: buckets multiply, someone deletes a directory a dashboard depends on, new racks arrive empty, and a release ships a feature you cannot roll back from. This article is about that job.

It assumes you know the moving parts: the Ozone Manager (OM) owns the namespace of volumes, buckets and keys; the Storage Container Manager (SCM) owns containers, pipelines and datanodes; datanodes store containers; Recon gives a read-only view of the whole cluster. If those names are new, read the Ozone architecture deep dive first. Here we go task by task through what an operator actually does, with the commands as documented for the Ozone 2.2 line, the reasoning behind each default, and the ways each task goes wrong.

Advertisement

Two control planes, one operator

Operator / automationozone sh, ozone admin, ozone fsOzone Manager (HA)namespace, quotas, snapshotsSCM (HA)containers, pipelines, nodesReconread-only view, reportsozone sh / fsozone admindashboardsSnapshots + trashprotect against peopleBalancer + replicationprotect against skew, lossNode lifecyclemaintenance, decommissionDatanodes: containers on disksIN_SERVICE / ENTERING_MAINTENANCE / DECOMMISSIONINGEvery day-2 task is a request to OM or SCM; datanodes act on commands SCM sends in heartbeats.
Day-2 work splits along the OM / SCM boundary. Namespace tasks go to OM, storage tasks go to SCM, and Recon is where you check the result.

Know which service owns a task: it decides the command family and the logs you read when it fails.

OM tasks are about names: creating volumes and buckets, choosing bucket layouts, setting quotas, creating and diffing snapshots, configuring trash. They go through ozone sh or ozone fs. OM applies them through its Ratis log, so they are fast and cheap in bytes, but each adds rows to OM's RocksDB; OM metadata disks, not datanode capacity, are what a busy namespace exhausts first.

SCM tasks are about bytes: moving containers between datanodes, draining nodes, re-replicating after failures. They go through ozone admin. SCM never moves data itself; it sends commands in heartbeat responses and watches replica reports, so storage tasks are asynchronous and each has a status command you must poll.

Recon is the operator's ledger: container health, open keys, capacity and namespace summary. End every SCM task with a Recon check, not the exit code of the command that started it.

Namespace hygiene: layouts and quotas

Ozone has two bucket layouts and you choose one per bucket at creation. FILE_SYSTEM_OPTIMIZED (FSO) stores directories as real entries, so renaming or deleting a directory is a single metadata operation; it is what Hive, Spark and anything that commits by rename needs. OBJECT_STORE (OBS) is a flat key space with S3 semantics, where a 'directory' is a prefix and a rename is a copy. The documented rule for shell-created buckets is that when ozone.default.bucket.layout is empty, Ozone uses FSO. Do not rely on the default: write the layout in every creation command, because it cannot be changed later and the wrong one shows up months afterwards as slow job commits.

Quotas come in two kinds and you should set both. A space quota caps bytes; a namespace quota caps the number of objects. The second one protects OM: a runaway job that writes millions of tiny keys does more damage to metadata than to disks.

# Create a team volume and an FSO bucket with explicit limits
ozone sh volume create --space-quota 500TB --namespace-quota 200000000 /analytics
ozone sh bucket create /analytics/warehouse --layout FILE_SYSTEM_OPTIMIZED \
    --space-quota 200TB --namespace-quota 50000000

# An erasure-coded bucket for cold data (6 data + 3 parity, 1 MiB cells)
ozone sh bucket create /analytics/archive --layout FILE_SYSTEM_OPTIMIZED \
    --type EC --replication rs-6-3-1024k

# Inspect and adjust later
ozone sh bucket info /analytics/warehouse
ozone sh bucket setquota --space-quota 300TB /analytics/warehouse
ozone sh bucket clrquota --namespace-quota /analytics/warehouse

Three details catch people. First, the documentation notes that a volume quota is only enforced when bucket quotas are also set, because writes check the bucket's used bytes; a volume quota on its own is decoration. Second, the space quota counts replicated bytes as stored (the docs require a quota of at least block size times replication factor), so a 200 TB quota on a three-way replicated bucket holds roughly 66 TB of user data, while the same quota on an rs-6-3 bucket holds about 133 TB (overhead 9/6 = 1.5x). Write quotas in the unit your users think in and convert explicitly. Third, the replication setting of a bucket is a default for new keys; ozone sh bucket set-replication-config changes what new writes get, not what is already stored. The trade-offs of EC itself are covered in HDFS erasure coding and apply to Ozone almost unchanged.

Advertisement

Snapshots and trash: protection against people

Replication handles hardware failures. Most real data loss comes from people and jobs: a wrong path in a cleanup script, an overwrite with an empty dataset. Ozone has two tools for that.

Trash catches deletes made through the file system shell. When fs.trash.interval is above zero and fs.trash.classname is set to org.apache.hadoop.fs.ozone.OzoneTrashPolicy, ozone fs -rm moves keys to a trash directory instead of deleting them, and an emptier purges them after the interval (in minutes). Trash does nothing for S3 deletes, overwrites, or programs that call delete through an API with trash bypassed.

Snapshots catch everything else, because they freeze the bucket's whole key tree. Creating one is a metadata operation in OM (the docs describe it as RocksDB checkpoints and pointers, not data copies), so it is instant whatever the bucket size. The snapshot is readable under the bucket's read-only .snapshot directory.

ozone sh snapshot create /analytics/warehouse nightly-2026-10-01
ozone sh snapshot list   /analytics/warehouse

# Read an old version directly
ozone fs -ls /analytics/warehouse/.snapshot/nightly-2026-10-01/sales/
ozone sh key get /analytics/warehouse/.snapshot/nightly-2026-10-01/sales/part-0001.parquet ./restore.parquet

# What changed since the snapshot? Diffs run as asynchronous jobs.
ozone sh snapshot diff /analytics/warehouse nightly-2026-09-30 nightly-2026-10-01
ozone sh snapshot listDiff /analytics/warehouse

ozone sh snapshot delete /analytics/warehouse nightly-2026-09-24

Snapshot diff answers 'what did last night's job touch' and is the basis of incremental replication: copy only the keys a diff reports instead of re-listing millions. The cost side is just as real. A snapshot pins every block it references, so deleting data from the live bucket frees no space until every snapshot holding it is deleted. Snapshots also add OM metadata. Run them on a schedule with a fixed retention (for example 7 dailies and 4 weeklies), delete in the same automation that creates, and alert when a bucket's snapshot count or pinned bytes grow without bound. The HDFS equivalent, with similar trade-offs, is described in HDFS snapshots.

Keeping datanodes balanced

New nodes join empty while old ones stay near full, concentrating write load and failure risk. The container balancer moves closed containers from over-utilized to under-utilized datanodes.

ozone admin containerbalancer start -t 10 -i 10
ozone admin containerbalancer status -v --history
ozone admin containerbalancer stop

The defaults in the documentation describe a deliberately gentle process. A node counts as balanced when its utilization is within hdds.container.balancer.utilization.threshold (10 percent) of the cluster average. Each iteration involves at most 20 percent of healthy datanodes, moves at most 500 GB in total and at most 26 GB into or out of any single node, waits 70 minutes between iterations, and runs 10 iterations unless you pass -i -1 for no limit.

Worked numbers make the pace visible. Add ten empty 100 TB nodes to a cluster of forty nodes at 80 percent. Average utilization drops to 64 percent, so each new node needs roughly 54 TB, 540 TB in all. The 20 percent cap allows 10 involved nodes per iteration, sources included, so about five targets take 26 GB each: roughly 130 GB per iteration, over four thousand iterations, more than six months at 70-minute intervals. The defaults are tuned to never hurt foreground traffic, not for that case. Raise the per-node and per-iteration sizes in steps while watching datanode latency, run with unlimited iterations, and treat balance as something you converge towards.

Taking nodes out: maintenance or decommission

Maintenance says the node is coming back: SCM ensures enough replicas exist elsewhere, then stops treating its replicas as missing, so a reboot does not trigger a re-replication storm. Decommission says the node is leaving: SCM copies every container off first.

ozone admin datanode list
ozone admin datanode maintenance --end=6 dn17.example.com
ozone admin datanode recommission dn17.example.com

ozone admin datanode decommission dn23.example.com
ozone admin datanode status decommission

Entering maintenance is gated. By default hdds.scm.replication.maintenance.replica.minimum is 2, so for a three-way replicated container the node stays in ENTERING_MAINTENANCE until at least two other replicas are healthy. For EC containers, hdds.scm.replication.maintenance.remaining.redundancy defaults to 1: with a node in maintenance there must still be data-count plus one replicas online. Reads of EC data on a node in maintenance may need reconstruction, which is invisible to clients but costs CPU and latency.

The operational rule follows. Use maintenance for anything under a day and always set an end time, so a forgotten node does not silently reduce redundancy for weeks. Use decommission for removals and drain at most a rack's worth at a time; the replication limits on decommissioning nodes are scaled by hdds.datanode.replication.outofservice.limit.factor so a draining node pushes faster than normal, but the receiving nodes still share the load with foreground writes. OM and SCM nodes have their own decommission commands (ozone admin om decommission, ozone admin scm decommission); bootstrap a replacement first so the Ratis ring never drops below three members.

Upgrades: prepare, pre-finalize, finalize

Ozone's non-rolling upgrade is built around a reversible middle state.

  1. Prepare OM. ozone admin om prepare -id=ozone1 makes every OM apply its log and take a snapshot, then refuse new writes. This gives a clean starting point.
  2. Swap binaries and start. Start SCM and datanodes on the new version, then start every OM with --upgrade. If one OM is started without the flag, ozone admin om cancelprepare gets them all out of prepare mode.
  3. Run pre-finalized. Components see data written by the old version and stay in a pre-finalized state. The cluster is fully usable, new on-disk features are disabled, and data created now stays readable after a downgrade. This is where you soak: run real jobs for days.
  4. Finalize. ozone admin scm finalizeupgrade then ozone admin om finalizeupgrade -id=ozone1. After this there is no downgrade. Check with the matching finalizationstatus commands. Datanodes that have not finished finalizing show as HEALTHY_READONLY: readable, but excluded from write pipelines until they report in.

Treat finalization as a separate, approved change days later, never the last line of the upgrade script.

Worked example: moving a warehouse directory off HDFS

Suppose a 400 TB Hive warehouse at hdfs://nn1/warehouse/sales moves to the FSO bucket created earlier. The Ozone file system client lets hadoop distcp treat both sides as file systems: bulk copy, incremental passes, a short freeze, a final pass, cutover.

# HDFS defaults to CRC32C checksums, Ozone to CRC32: make Ozone match or the CRC check fails
# 1. Bulk copy while producers keep writing (bandwidth-capped per map task)
hadoop distcp -Dozone.client.checksum.type=CRC32C -m 200 -bandwidth 100 \
    hdfs://nn1/warehouse/sales  ofs://ozone1/analytics/warehouse/sales

# 2. Incremental passes: only files that are new or differ
hadoop distcp -Dozone.client.checksum.type=CRC32C -update -m 200 \
    hdfs://nn1/warehouse/sales  ofs://ozone1/analytics/warehouse/sales

# 3. Freeze writers, final -update pass, verify, repoint table locations, unfreeze

The Ozone DistCp guide documents that setting, adds -Ddfs.checksum.combine.mode=COMPOSITE_CRC for HDFS older than 3.1.1, and notes that encrypted data never matches, where -skipcrccheck is the documented escape. Whenever checks are skipped, verify independently: compare counts and bytes per directory, and hash a sample of files:

import subprocess

def count(path):
    # Hadoop-compatible -count: DIR_COUNT FILE_COUNT CONTENT_SIZE PATHNAME
    out = subprocess.run(["hadoop", "fs", "-count", path],
                         check=True, capture_output=True, text=True).stdout.split()
    return int(out[0]), int(out[1]), int(out[2])

pairs = [("hdfs://nn1/warehouse/sales/dt=2026-09-%02d" % d,
          "ofs://ozone1/analytics/warehouse/sales/dt=2026-09-%02d" % d) for d in range(1, 31)]
bad = [(s, count(s), count(t)) for s, t in pairs if count(s) != count(t)]
for src, a, b in bad:
    print("MISMATCH", src, a, b)
raise SystemExit(1 if bad else 0)

Snapshot the bucket right after cutover: an exact record of what was migrated and a cheap rollback point.

Failure modes and how to see them early

SymptomUsual causeFirst check
Deleting data frees no spaceSnapshots still reference the blocksSnapshot list and retention job
OM latency climbs, disks fineOM RocksDB disk filling or slow, many snapshots or tiny keysOM metadata disk usage, namespace quota usage
Node stuck in ENTERING_MAINTENANCEToo few healthy replicas elsewhereRecon under-replicated containers
Decommission never completesReplication throttled, or target nodes fullDecommission status, balancer state, free space
Writes avoid some new nodes after upgradeDatanodes still HEALTHY_READONLYSCM finalization status
Balancer runs, nothing movesThresholds already met, or per-node cap too lowBalancer status with history
Job commits slow and non-atomicBucket created as OBJECT_STOREBucket info: layout

Most are slow failures that show in metrics days before they hurt. Alert on OM metadata disk usage, under-replicated container count that does not trend to zero, snapshot counts per bucket, nodes outside IN_SERVICE for longer than their planned window, and finalization not completed within an agreed period after an upgrade.

Trade-offs worth stating plainly

Ozone's separation of namespace from block management removes the single-heap ceiling that the HDFS NameNode imposes, but it does not make metadata free: OM still has to hold and replicate every key, and snapshots multiply what it tracks. Snapshots are cheap to take and expensive to keep; the balancer protects foreground traffic but converges slowly; maintenance avoids replication storms but reduces redundancy; pre-finalized upgrades trade a short write outage for a real rollback path. These are levers: set them deliberately and write the choice down.

What to do next

  1. Inventory buckets with ozone sh bucket info and record layout, replication and quotas; fix any OBJECT_STORE bucket that Hive or Spark commits into by migrating it.
  2. Set both space and namespace quotas on every volume, converting user-facing capacity into stored bytes for the bucket's replication.
  3. Enable trash for file-system users and schedule snapshots with automated deletion; alert on snapshot count per bucket.
  4. Write maintenance and decommission runbooks that always pass an end time and end with a Recon under-replication check.
  5. Run the container balancer continuously with tuned per-node limits after every capacity expansion.
  6. Rehearse a non-rolling upgrade on a staging cluster, including a downgrade from the pre-finalized state, before doing it in production.
  7. For HDFS migrations, script bulk plus incremental distcp passes with count-and-bytes verification, and snapshot the bucket at cutover.
Key takeaway: Operating Ozone well means knowing which control plane owns each task and checking results in Recon rather than trusting command exit codes. Write layouts and quotas explicitly, protect data from people with trash and retained snapshots, let the balancer converge gently but continuously, use maintenance for short absences and decommission for removals, and keep upgrade finalization as a separate, deliberate step.