HDFS has no undo. A delete call on the NameNode removes the inode, schedules the blocks for deletion on DataNodes, and a few heartbeats later the bytes are gone. Trash is the safety net built on top: when it is enabled, the shell turns a delete into a rename into a per-user trash directory, and a background thread removes old trash later. It is cheap, because a rename is a metadata change, and it has saved many clusters from a mistyped -rm -r.
It is also widely misunderstood. Operators set a retention of one day and find files lasting almost two, delete tens of terabytes and see no space come back, or discover during an incident that their ETL job's deletes never went to trash at all. This article explains the mechanism from the source-level behaviour up, works through the retention arithmetic, lists everything that bypasses trash, and ends with a runbook.
How a delete becomes a rename
When trash is enabled and you run hdfs dfs -rm -r /data/sales/2026-09 as user etl, the shell calls Trash.moveToAppropriateTrash. That resolves the fully qualified path and the file system that owns it, asks the server for its trash interval, and moves the path to /user/etl/.Trash/Current/data/sales/2026-09. The full original path is recreated under Current, which is what makes restore a simple move back.
If the target name already exists in Current (you deleted a directory with the same path twice), the default policy appends a millisecond timestamp to the new name rather than overwriting the first copy. Nothing is copied and no DataNode is involved: the move is one NameNode rename, so trashing a petabyte directory takes the same time as trashing a single file.
Checkpoints and the emptier
Files do not expire from Current directly. A thread called the emptier, which runs inside the NameNode when the NameNode's own configuration has a non-zero fs.trash.interval, wakes once every checkpoint interval, aligned to multiples of that interval. For every trash root it finds, it does two things in order:
- Delete expired checkpoints. Each checkpoint is a directory named with its creation time in the format
yyMMddHHmmss. A checkpoint is deleted when the current time minusfs.trash.intervalis later than that timestamp. - Create a new checkpoint. It renames
Currentto a directory named with the current time. The next delete creates a freshCurrent.
The shell command hdfs dfs -expunge runs the same two steps immediately for the calling user. With -immediate it deletes everything in that user's trash regardless of age, and -fs points it at a file system other than the default.
The two settings and who wins
| Property | Default | Meaning |
|---|---|---|
fs.trash.interval | 0 (disabled) | minutes a checkpoint is kept before it may be deleted |
fs.trash.checkpoint.interval | 0 (use fs.trash.interval) | minutes between emptier runs; should be at most fs.trash.interval |
Both properties live in core-site.xml and can be set on clients and on the NameNode. According to the property documentation, when the server has a non-zero interval, clients use the server's value and ignore their own; only when trash is disabled on the server does the client's setting decide. Two consequences follow. Setting the interval on the NameNode is what makes trash reliable for shell users. And a client that sets its own interval against a server with trash disabled will move files into trash, but no emptier on the server will ever clean them, so the space leaks until someone expunges it.
Worked example: what one day of retention really means
Take fs.trash.interval=1440 (one day) and fs.trash.checkpoint.interval=60 (hourly). The emptier ticks on the hour. A file deleted at 10:05 sits in Current until 11:00, when it becomes checkpoint ...110000. On each hourly tick the emptier compares that timestamp with now minus 1,440 minutes. At 11:00 the next day the two look equal, but the rule requires strictly older, and it is met: the checkpoint name is truncated to whole seconds while the emptier reads the clock in milliseconds after waking, so its current time is always slightly past the tick. The checkpoint is deleted on that tick. The file lived 24 hours and 55 minutes.
In general, retention is never shorter than fs.trash.interval and can reach the interval plus one checkpoint interval, depending on how the delete lines up with ticks (a little more if the interval is not a multiple of the checkpoint interval). With the checkpoint interval left at 0, it equals the full interval, so a one-day setting keeps some files for close to two days. That is the most common surprise in trash capacity planning. Size your storage headroom for the maximum, and shorten the checkpoint interval if you need retention to track the setting closely.
def trash_lifetime(delete_min, interval, ckpt):
"""Minutes a file deleted at minute `delete_min` stays in trash (default HDFS policy)."""
ckpt = ckpt or interval
tick = -(-delete_min // ckpt) * ckpt # first tick at or after the delete
checkpoint_time = tick
t = tick
while not (t - interval >= checkpoint_time): # clock is ms past the tick, so equality expires
t += ckpt
return t - delete_min
print(trash_lifetime(605, 1440, 60) / 60) # ~24.9 hours
print(trash_lifetime(605, 1440, 0) / 60) # ~37.9 hours
Where trash lives: one root per user, zone and snapshot tree
The default trash root is /user/<name>/.Trash. A rename cannot cross an encryption zone boundary, so HDFS gives every encryption zone its own trash root at <zone>/.Trash/<name>/Current. Files deleted inside a zone stay encrypted in place and never need re-encryption.
Snapshottable directories raise a similar problem: moving a file out of a snapshotted tree into /user complicates snapshot bookkeeping. Recent Hadoop 3.x releases add dfs.namenode.snapshot.trashroot.enabled, which gives each snapshottable directory its own trash root in the same way; check your version's hdfs-default.xml before relying on it. See HDFS snapshots for the underlying copy-on-write model.
Because the trash lives under the user's home directory, trashed data still counts against any namespace and space quotas on that directory. The shell documentation explicitly suggests -skipTrash for deleting from a directory that is over quota, since a move to trash frees nothing.
What bypasses trash
hdfs dfs -rm -skipTrash. Deletes immediately. Useful for over-quota cleanups; dangerous as a habit.- Code that calls
FileSystem.delete(). Trash is a client-side policy implemented by the shell. Spark, MapReduce, Hive and custom jobs that call theFileSystemAPI delete directly unless they callTrash.moveToAppropriateTrash(fs, path, conf)themselves. Find out which of your pipelines do; many do not. - Deleting something already in trash. The default policy refuses to move a path that is already under the trash root, so the shell deletes it permanently.
hdfs dfs -rm -r .Trash/Current/...is a real delete. - Deleting a parent of the trash. Removing a path that contains the trash root, such as a whole home directory, fails with an error that it contains the trash; operators often answer with
-skipTrash, and the whole tree, trash included, is gone. hdfs dfs -expunge -immediateand, for tables, engine-level purge options such as Hive'sDROP TABLE ... PURGE.
Restoring files
Restore is a rename in the opposite direction. Find the file, check that the original location is free, and move it back:
# 1. find it (Current first, then checkpoints, newest first)
hdfs dfs -ls -R /user/etl/.Trash | grep 'data/sales/2026-09'
# 2. make sure the destination parent exists and the name is free
hdfs dfs -mkdir -p /data/sales
hdfs dfs -test -e /data/sales/2026-09 && echo "destination exists; restore elsewhere"
# 3. move it back (metadata only, instant)
hdfs dfs -mv /user/etl/.Trash/260930110000/data/sales/2026-09 /data/sales/2026-09
# encryption zone: the trash is inside the zone
hdfs dfs -ls /secure/.Trash/etl/Current/secure/When you do not know who deleted a path, the NameNode audit log answers it. A trash move is recorded as a rename whose destination is under a .Trash directory, with the user, client address and time, while a permanent delete is recorded as a delete. Searching the audit log for the original path therefore tells you both which trash root to look in and whether there is anything to restore at all, which saves a long walk through every user's checkpoints.
Act fast when you need a restore. The file expires on its own schedule whether or not you are looking for it, and a colleague running -expunge as the owning user removes it too. Suspending the emptier requires changing the NameNode configuration and restarting it, so for high-value paths the better protection is a snapshot, which trash does not replace.
Operations and trade-offs
Trash trades storage for safety. Every byte deleted in the last retention window is still on disk, replicated, so a cluster that churns 20 TB a day with three replicas and roughly two days of effective retention holds about 120 TB of raw capacity in trash (illustrative numbers). Measure it rather than guessing: hdfs dfs -du -s -h /user/*/.Trash plus the trash directory of every encryption zone.
# trash usage per user, largest first (add each encryption zone's .Trash the same way)
hdfs dfs -du -s '/user/*/.Trash' | sort -k1,1 -n -r | head -20| Setting | Good for | Watch out for |
|---|---|---|
| no trash (interval 0) | scratch clusters, pipelines with their own versioning | every mistake is permanent |
| 1 day, hourly checkpoints | most shared clusters | roughly one day of churn held on disk |
| 7 days | clusters with many interactive users | capacity, and quotas filling with trash |
| trash plus snapshots on key paths | production data | snapshot space and cleanup discipline |
Other operating rules: set the interval on the NameNode so that users cannot quietly disable it; alert when trash grows faster than expected, because it usually means a job is looping deletes; and remember that space returns only when a checkpoint expires, and not even then if a snapshot still references the blocks. Trash also sits on top of the NameNode: every checkpoint rename and deletion is an edit-log transaction, and deleting a checkpoint with millions of files is a large delete that the NameNode performs in increments.
Failure modes
- "We deleted 50 TB and nothing was freed." The data is in trash, or held by a snapshot. Check both before touching anything.
- Home directory over quota after cleanups. Trash counts against the quota. Use
-skipTrashdeliberately, or expunge. - Trash that never empties. Clients enable trash but the NameNode interval is 0, so no emptier runs. Enable it server-side or schedule
-expunge. - Silent permanent deletes from jobs. An ETL job overwrites output with
FileSystem.deleteand there is nothing to restore. Route pipeline deletes throughmoveToAppropriateTrashor version outputs. - Retention longer than planned. A zero checkpoint interval can nearly double effective retention; set it explicitly.
What to do next
- Read
fs.trash.intervalandfs.trash.checkpoint.intervalfrom the NameNode's effective configuration, not from a client, and set both explicitly. - Compute your worst-case retention with the function above and size trash capacity for it.
- Measure current trash usage across all user and encryption-zone trash roots and add a dashboard and a growth alert.
- Audit your pipelines for direct
FileSystem.deletecalls and decide, per pipeline, whether it needs trash or versioned outputs. - Practise a restore on a test file, including one inside an encryption zone, and write the steps into your runbook.
- Put snapshots on the directories you cannot afford to lose; trash is a short safety net, not a backup.