HDFS has no undo. A delete call on the NameNode removes the inode, schedules the blocks for deletion on DataNodes, and a few heartbeats later the bytes are gone. Trash is the safety net built on top: when it is enabled, the shell turns a delete into a rename into a per-user trash directory, and a background thread removes old trash later. It is cheap, because a rename is a metadata change, and it has saved many clusters from a mistyped -rm -r.

It is also widely misunderstood. Operators set a retention of one day and find files lasting almost two, delete tens of terabytes and see no space come back, or discover during an incident that their ETL job's deletes never went to trash at all. This article explains the mechanism from the source-level behaviour up, works through the retention arithmetic, lists everything that bypasses trash, and ends with a runbook.

Advertisement

How a delete becomes a rename

HDFS trash: delete becomes a rename, and the emptier turns Current into dated checkpoints that expirehdfs dfs -rmor Trash.moveToAppropriateTrashrename.Trash/Current/...full original path kepttick.Trash/260930170000checkpoint (yyMMddHHmmss)age > intervaldeleteblocks freedrestoremv back from Current or checkpointhdfs dfs -mvNameNode emptier threadevery checkpoint intervalEach tick, for every trash root: delete expired checkpoints, then rename Current to a new checkpoint.Trash roots: /user//.Trash, /.Trash/, and (if enabled) /.Trash/Bypasses: rm -skipTrash, FileSystem.delete() in code, rm on a path already inside .Trash, expunge -immediateSpace is freed only when a checkpoint is deleted, and only if no snapshot still references the blocks.
The shell renames the path into Current. The emptier later rolls Current into a timestamped checkpoint and deletes checkpoints older than fs.trash.interval.

When trash is enabled and you run hdfs dfs -rm -r /data/sales/2026-09 as user etl, the shell calls Trash.moveToAppropriateTrash. That resolves the fully qualified path and the file system that owns it, asks the server for its trash interval, and moves the path to /user/etl/.Trash/Current/data/sales/2026-09. The full original path is recreated under Current, which is what makes restore a simple move back.

If the target name already exists in Current (you deleted a directory with the same path twice), the default policy appends a millisecond timestamp to the new name rather than overwriting the first copy. Nothing is copied and no DataNode is involved: the move is one NameNode rename, so trashing a petabyte directory takes the same time as trashing a single file.

Checkpoints and the emptier

Files do not expire from Current directly. A thread called the emptier, which runs inside the NameNode when the NameNode's own configuration has a non-zero fs.trash.interval, wakes once every checkpoint interval, aligned to multiples of that interval. For every trash root it finds, it does two things in order:

  1. Delete expired checkpoints. Each checkpoint is a directory named with its creation time in the format yyMMddHHmmss. A checkpoint is deleted when the current time minus fs.trash.interval is later than that timestamp.
  2. Create a new checkpoint. It renames Current to a directory named with the current time. The next delete creates a fresh Current.

The shell command hdfs dfs -expunge runs the same two steps immediately for the calling user. With -immediate it deletes everything in that user's trash regardless of age, and -fs points it at a file system other than the default.

Advertisement

The two settings and who wins

PropertyDefaultMeaning
fs.trash.interval0 (disabled)minutes a checkpoint is kept before it may be deleted
fs.trash.checkpoint.interval0 (use fs.trash.interval)minutes between emptier runs; should be at most fs.trash.interval

Both properties live in core-site.xml and can be set on clients and on the NameNode. According to the property documentation, when the server has a non-zero interval, clients use the server's value and ignore their own; only when trash is disabled on the server does the client's setting decide. Two consequences follow. Setting the interval on the NameNode is what makes trash reliable for shell users. And a client that sets its own interval against a server with trash disabled will move files into trash, but no emptier on the server will ever clean them, so the space leaks until someone expunges it.

Worked example: what one day of retention really means

Take fs.trash.interval=1440 (one day) and fs.trash.checkpoint.interval=60 (hourly). The emptier ticks on the hour. A file deleted at 10:05 sits in Current until 11:00, when it becomes checkpoint ...110000. On each hourly tick the emptier compares that timestamp with now minus 1,440 minutes. At 11:00 the next day the two look equal, but the rule requires strictly older, and it is met: the checkpoint name is truncated to whole seconds while the emptier reads the clock in milliseconds after waking, so its current time is always slightly past the tick. The checkpoint is deleted on that tick. The file lived 24 hours and 55 minutes.

In general, retention is never shorter than fs.trash.interval and can reach the interval plus one checkpoint interval, depending on how the delete lines up with ticks (a little more if the interval is not a multiple of the checkpoint interval). With the checkpoint interval left at 0, it equals the full interval, so a one-day setting keeps some files for close to two days. That is the most common surprise in trash capacity planning. Size your storage headroom for the maximum, and shorten the checkpoint interval if you need retention to track the setting closely.

def trash_lifetime(delete_min, interval, ckpt):
    """Minutes a file deleted at minute `delete_min` stays in trash (default HDFS policy)."""
    ckpt = ckpt or interval
    tick = -(-delete_min // ckpt) * ckpt             # first tick at or after the delete
    checkpoint_time = tick
    t = tick
    while not (t - interval >= checkpoint_time):     # clock is ms past the tick, so equality expires
        t += ckpt
    return t - delete_min

print(trash_lifetime(605, 1440, 60) / 60)    # ~24.9 hours
print(trash_lifetime(605, 1440, 0) / 60)     # ~37.9 hours

Where trash lives: one root per user, zone and snapshot tree

The default trash root is /user/<name>/.Trash. A rename cannot cross an encryption zone boundary, so HDFS gives every encryption zone its own trash root at <zone>/.Trash/<name>/Current. Files deleted inside a zone stay encrypted in place and never need re-encryption.

Snapshottable directories raise a similar problem: moving a file out of a snapshotted tree into /user complicates snapshot bookkeeping. Recent Hadoop 3.x releases add dfs.namenode.snapshot.trashroot.enabled, which gives each snapshottable directory its own trash root in the same way; check your version's hdfs-default.xml before relying on it. See HDFS snapshots for the underlying copy-on-write model.

Because the trash lives under the user's home directory, trashed data still counts against any namespace and space quotas on that directory. The shell documentation explicitly suggests -skipTrash for deleting from a directory that is over quota, since a move to trash frees nothing.

What bypasses trash

  • hdfs dfs -rm -skipTrash. Deletes immediately. Useful for over-quota cleanups; dangerous as a habit.
  • Code that calls FileSystem.delete(). Trash is a client-side policy implemented by the shell. Spark, MapReduce, Hive and custom jobs that call the FileSystem API delete directly unless they call Trash.moveToAppropriateTrash(fs, path, conf) themselves. Find out which of your pipelines do; many do not.
  • Deleting something already in trash. The default policy refuses to move a path that is already under the trash root, so the shell deletes it permanently. hdfs dfs -rm -r .Trash/Current/... is a real delete.
  • Deleting a parent of the trash. Removing a path that contains the trash root, such as a whole home directory, fails with an error that it contains the trash; operators often answer with -skipTrash, and the whole tree, trash included, is gone.
  • hdfs dfs -expunge -immediate and, for tables, engine-level purge options such as Hive's DROP TABLE ... PURGE.

Restoring files

Restore is a rename in the opposite direction. Find the file, check that the original location is free, and move it back:

# 1. find it (Current first, then checkpoints, newest first)
hdfs dfs -ls -R /user/etl/.Trash | grep 'data/sales/2026-09'

# 2. make sure the destination parent exists and the name is free
hdfs dfs -mkdir -p /data/sales
hdfs dfs -test -e /data/sales/2026-09 && echo "destination exists; restore elsewhere"

# 3. move it back (metadata only, instant)
hdfs dfs -mv /user/etl/.Trash/260930110000/data/sales/2026-09 /data/sales/2026-09

# encryption zone: the trash is inside the zone
hdfs dfs -ls /secure/.Trash/etl/Current/secure/

When you do not know who deleted a path, the NameNode audit log answers it. A trash move is recorded as a rename whose destination is under a .Trash directory, with the user, client address and time, while a permanent delete is recorded as a delete. Searching the audit log for the original path therefore tells you both which trash root to look in and whether there is anything to restore at all, which saves a long walk through every user's checkpoints.

Act fast when you need a restore. The file expires on its own schedule whether or not you are looking for it, and a colleague running -expunge as the owning user removes it too. Suspending the emptier requires changing the NameNode configuration and restarting it, so for high-value paths the better protection is a snapshot, which trash does not replace.

Operations and trade-offs

Trash trades storage for safety. Every byte deleted in the last retention window is still on disk, replicated, so a cluster that churns 20 TB a day with three replicas and roughly two days of effective retention holds about 120 TB of raw capacity in trash (illustrative numbers). Measure it rather than guessing: hdfs dfs -du -s -h /user/*/.Trash plus the trash directory of every encryption zone.

# trash usage per user, largest first (add each encryption zone's .Trash the same way)
hdfs dfs -du -s '/user/*/.Trash' | sort -k1,1 -n -r | head -20
SettingGood forWatch out for
no trash (interval 0)scratch clusters, pipelines with their own versioningevery mistake is permanent
1 day, hourly checkpointsmost shared clustersroughly one day of churn held on disk
7 daysclusters with many interactive userscapacity, and quotas filling with trash
trash plus snapshots on key pathsproduction datasnapshot space and cleanup discipline

Other operating rules: set the interval on the NameNode so that users cannot quietly disable it; alert when trash grows faster than expected, because it usually means a job is looping deletes; and remember that space returns only when a checkpoint expires, and not even then if a snapshot still references the blocks. Trash also sits on top of the NameNode: every checkpoint rename and deletion is an edit-log transaction, and deleting a checkpoint with millions of files is a large delete that the NameNode performs in increments.

Failure modes

  • "We deleted 50 TB and nothing was freed." The data is in trash, or held by a snapshot. Check both before touching anything.
  • Home directory over quota after cleanups. Trash counts against the quota. Use -skipTrash deliberately, or expunge.
  • Trash that never empties. Clients enable trash but the NameNode interval is 0, so no emptier runs. Enable it server-side or schedule -expunge.
  • Silent permanent deletes from jobs. An ETL job overwrites output with FileSystem.delete and there is nothing to restore. Route pipeline deletes through moveToAppropriateTrash or version outputs.
  • Retention longer than planned. A zero checkpoint interval can nearly double effective retention; set it explicitly.

What to do next

  1. Read fs.trash.interval and fs.trash.checkpoint.interval from the NameNode's effective configuration, not from a client, and set both explicitly.
  2. Compute your worst-case retention with the function above and size trash capacity for it.
  3. Measure current trash usage across all user and encryption-zone trash roots and add a dashboard and a growth alert.
  4. Audit your pipelines for direct FileSystem.delete calls and decide, per pipeline, whether it needs trash or versioned outputs.
  5. Practise a restore on a test file, including one inside an encryption zone, and write the steps into your runbook.
  6. Put snapshots on the directories you cannot afford to lose; trash is a short safety net, not a backup.
Key takeaway: With trash enabled, an HDFS shell delete is a rename into the user's .Trash/Current (or the encryption zone's own trash root). A NameNode emptier thread, running every checkpoint interval, deletes checkpoints older than fs.trash.interval and rolls Current into a new timestamped checkpoint, so real retention runs from the interval to about the interval plus one checkpoint interval. The server's interval overrides clients'. FileSystem.delete in code, -skipTrash, deletes inside .Trash and expunge -immediate all bypass it. Trashed data still uses disk and quota until it expires. Size for the maximum, audit pipeline deletes, and use snapshots for data you cannot lose.