Most shared file systems follow the same curve. A project writes a burst of files, reads them hard for a few weeks, and then barely touches them again. Months later they still sit on the most expensive storage tier, because nobody wants to work out which files are safe to move. Amazon EFS handles this with lifecycle management. A per-file access timer moves cold data to cheaper storage classes, and an optional rule brings it back the moment someone reads it. AWS calls that combination EFS Intelligent-Tiering.
This article explains how the mechanism actually works: what counts as an access, what the three policies do, how tiering interacts with throughput modes and small files, and how to model whether it saves money for your data. It ends with a rollout plan you can run against a real file system. For EFS fundamentals such as mount targets, performance modes and NFS semantics, start with the AWS EFS overview.
What Intelligent-Tiering means on EFS
EFS Intelligent-Tiering is not a storage class and not a switch. It is the name AWS gave, when it launched the feature in September 2021, to one particular lifecycle configuration: a policy that demotes files after N days without access, plus TransitionToPrimaryStorageClass set to AFTER_1_ACCESS, which promotes a file back to Standard on its first access. Without the second policy you have plain lifecycle management. Demoted files stay in the cheaper class even when read, and every read pays the slower first byte and an access charge.
This is unlike S3 Intelligent-Tiering, which is a storage class of its own with a per-object monitoring fee and automatic access tiers. On EFS there is no monitoring fee and no separate class. The whole behaviour comes from three policies applied to the file system as a whole, never to a single directory. If two datasets need different rules, they need two file systems.
Three storage classes and their billing rules
EFS has three storage classes. The figures below come from the EFS user guide; prices change by Region and over time, so read them from the pricing page rather than from any article.
| Class | Designed for | First-byte latency | Minimum billed size per file | Minimum storage duration |
|---|---|---|---|---|
| Standard | Active data | Sub-millisecond | None | None |
| Infrequent Access (IA) | Data read a few times a quarter | Tens of milliseconds | 128 KiB | None |
| Archive | Data read a few times a year or less | Tens of milliseconds | 128 KiB | 90 days |
Three billing facts drive every decision later in this article. First, you pay for data in each class. Second, you also pay data access charges when files in IA or Archive are read, and for data that lifecycle management moves between classes. Third, IA and Archive bill at least 128 KiB per file, so a 4 KiB file costs as if it were 128 KiB. Archive needs Elastic throughput, and the storage-class table lists it only for Regional file systems.
How the lifecycle timer decides
Lifecycle decisions are based on when a file was last accessed in Standard. EFS keeps an internal timer for this; it is not the POSIX atime you can see with stat, so mounting with noatime changes nothing. Each access to a file in Standard resets the timer. When the timer passes the policy age, the file becomes a candidate to move.
Several details matter in practice. Metadata, meaning file names, ownership and the directory tree, always stays in Standard, and metadata operations such as listing a directory do not count as access. Backup tools that only stat files are therefore harmless, but a tool that opens and reads every file is not. Transitions run at lower priority than your workload, so moving millions of small files takes longer than moving fewer large files of the same total size. While a file is being moved it is still billed as Standard. Writes to a file in IA or Archive first land in Standard and become eligible to move again after 24 hours.
The practical consequence: on a busy file system, tiering is eventual. Expect the Standard byte count to fall over days, not minutes, after you enable a policy.
The three policies and how to set them
A lifecycle configuration holds up to three policies, and the API insists that each LifecyclePolicy object carries only one transition. The valid values come from the EFS API reference.
TransitionToIA:AFTER_1_DAY,AFTER_7_DAYS,AFTER_14_DAYS,AFTER_30_DAYS,AFTER_60_DAYS,AFTER_90_DAYS,AFTER_180_DAYS,AFTER_270_DAYSorAFTER_365_DAYS.TransitionToArchive: the same set of values, measured from last access in Standard. It applies to files in Standard or IA.TransitionToPrimaryStorageClass: onlyAFTER_1_ACCESS. Omit it and files stay where they are when read.
A file system created in the console with recommended settings gets 30 days to IA, 90 days to Archive and no transition back. That default suits archives, not workloads that return to old data. To turn on Intelligent-Tiering behaviour, write all three policies explicitly:
aws efs put-lifecycle-configuration \
--file-system-id fs-0123456789abcdef0 \
--lifecycle-policies \
'[{"TransitionToIA":"AFTER_30_DAYS"},
{"TransitionToArchive":"AFTER_180_DAYS"},
{"TransitionToPrimaryStorageClass":"AFTER_1_ACCESS"}]'
# Turn lifecycle management off entirely: an empty policy list.
aws efs put-lifecycle-configuration \
--file-system-id fs-0123456789abcdef0 \
--lifecycle-policies '[]'The call replaces the whole configuration rather than merging with it, so always send the full set. Keep the configuration in code next to the file system definition, so a console edit shows up as drift instead of going unnoticed.
Interaction with throughput modes
Tiering interacts with throughput in two ways, and both can hurt.
Bursting throughput shrinks. In Bursting mode, baseline throughput and burst credits scale with the amount of data in Standard only. If tiering moves 90 percent of your bytes to IA, your baseline drops by the same proportion. A build farm that ran happily on bursting credits can start throttling the week after you enable lifecycle management. Check the BurstCreditBalance trend before and after, or move to Elastic throughput.
Archive pins the throughput mode. Archive is supported only with Elastic throughput. Once a lifecycle policy transitions data to Archive, you cannot switch the file system to Bursting or Provisioned. Elastic bills for the metadata and data you actually transfer, so it is usually the right choice for spiky, mostly cold file systems anyway. Make that decision deliberately rather than discovering the lock later.
Small files and the 128 KiB minimum
The 128 KiB minimum billed size is where tiering most often backfires. Lifecycle policies updated on or after 26 November 2023 do tier files smaller than 128 KiB, and each one is then billed as 128 KiB in IA or Archive. Suppose a directory holds ten million 8 KiB files, about 76 GiB of real data. In IA it is billed as about 1.2 TiB. Whether that saves money depends on the price ratio between Standard and IA in your Region, and for very small files it often does not.
Before enabling a policy, build a size histogram of the file system. If most bytes sit in large files and most files are tiny, tiering still helps. If tiny files dominate both counts, consider packing them into archives such as tar or Parquet, or keep that tree on a file system without lifecycle policies.
import os, collections
def size_profile(root):
# bucket file sizes in powers of two; count files and bytes per bucket
files, bytes_ = collections.Counter(), collections.Counter()
for dirpath, _, names in os.walk(root):
for n in names:
try:
s = os.lstat(os.path.join(dirpath, n)).st_size
except FileNotFoundError:
continue
b = max(s, 1).bit_length()
files[b] += 1
bytes_[b] += s
for b in sorted(files):
print(f"<= {2**b:>14,} B files={files[b]:>10,} bytes={bytes_[b]:>16,}")Walking with lstat is metadata-only, so the scan does not reset any lifecycle timers.
Worked example: does tiering pay?
Tiering is worth it when the storage you save beats the access and transition charges you add. The model below takes prices as inputs, because they vary by Region and change over time; fill them in from the EFS pricing page. Sizes are in GiB per month.
from dataclasses import dataclass
@dataclass
class Prices: # per GiB, from the EFS pricing page for your Region
standard: float
ia: float
ia_read: float # access charge per GiB read from IA
tier_move: float # charge per GiB moved by lifecycle management
MIN_BILLED_KIB = 128
def billed_gib(n_files, real_gib):
# each file in IA is billed at least 128 KiB
avg_kib = real_gib * 1024 * 1024 / max(n_files, 1)
return real_gib * max(1.0, MIN_BILLED_KIB / max(avg_kib, 1e-9))
def monthly(p, total_gib, cold_frac, cold_files, cold_read_gib, promote_back):
cold = total_gib * cold_frac
hot = total_gib - cold
base = total_gib * p.standard
tiered = hot * p.standard + billed_gib(cold_files, cold) * p.ia
tiered += cold * p.tier_move / 12 # one-off move, amortised over a year
if promote_back: # read once, moves back to Standard
tiered += cold_read_gib * (p.ia_read + p.standard - p.ia)
else: # every read pays the access charge
tiered += cold_read_gib * p.ia_read
return base, tieredRun it with real numbers. A 20 TiB research share where 85 percent of bytes have gone untouched for 30 days, the cold data sits in about two million files averaging 9 MiB, and users pull back 200 GiB a month. The 128 KiB floor costs nothing here, since the average file is far larger, and the saving tracks the Standard-to-IA price gap almost exactly. Change the cold set to fifty million 32 KiB files and the billed size quadruples, which wipes out most of the saving. The model is crude, since it ignores Archive and assumes a steady state, but it answers the question that matters: is the cold data big files or small ones?
Operating a tiered file system
Lifecycle management is easy to switch on and hard to observe. Three habits make it predictable.
- Watch bytes per class.
DescribeFileSystemsreturnsSizeInByteswith separate Standard, IA and Archive values, and the CloudWatchStorageBytesmetric is split by storage class. Graph them. A sudden rise in Standard on a tiered file system means something is reading old data, and you will pay for it twice: once in access charges and again in Standard storage. - Find the full-scan readers. Antivirus scanners,
rsync --checksum, content indexers and some backup agents open every file. WithAFTER_1_ACCESSthey promote the whole file system back to Standard on every run. Point them elsewhere, move them to metadata-only modes, or use AWS Backup, which does not incur data access charges on lifecycle-managed file systems. - Treat policy changes as deployments. Shortening
TransitionToIAfrom 90 to 7 days moves a large slice of data at once and bills the transitions. Roll changes out on one file system, watch a full cycle, then widen.
For training pipelines that read whole datasets every epoch, EFS tiering is the wrong lever; keep hot data on a high-throughput file system such as FSx for Lustre and leave the long tail on S3.
Failure modes
- Bounce loops. A nightly job reads the whole tree, everything promotes, and 30 days later it all demotes again. You pay transitions and access charges every month for nothing. Fix the reader, not the policy.
- Throttling after enablement. A Bursting file system loses baseline as bytes leave Standard. Symptom: rising latency and a draining credit balance with no change in the workload.
- Small-file bill shock. Millions of tiny files tier and are billed at 128 KiB each. Symptom: the IA line on the bill is far larger than the data you think you have.
- Latency-sensitive tail. A web app serving rarely viewed assets from EFS sees tens of milliseconds of extra first-byte latency on cold files. Without
AFTER_1_ACCESS, every read pays it. - Archive surprises. Data deleted or rewritten within 90 days of reaching Archive is still billed for the minimum duration, and the throughput mode can no longer change.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Short IA window (7-14 days) | Larger saving on bursty data | More transitions; more first reads pay slower latency |
| Long IA window (60-90 days) | Few wrong demotions | Less saving |
| AFTER_1_ACCESS on | Returning data is fast again | Scanners can undo tiering |
| AFTER_1_ACCESS off | Cold data stays cheap even if read once | Every read of cold data pays access charges and latency |
| Archive enabled | Lowest storage price on EFS | Elastic throughput only, 90-day minimum |
| Move to S3 instead | Cheaper still, object lifecycle | Application must speak S3, not NFS |
If the data no longer needs POSIX semantics, the strongest move is often to leave EFS entirely: export it to S3 and let S3 lifecycle rules handle it. Storage Gateway can keep an NFS front end for applications that cannot change. For block workloads, the EBS volume types are the comparison point instead.
What to do next
- Run the size profiler against each EFS file system and note what share of files and bytes are under 128 KiB.
- Read current
SizeInBytesper class and the throughput mode for each file system. - For Bursting file systems, decide whether to move to Elastic before enabling tiering.
- Fill the cost model with your Region's prices and your cold fraction; skip file systems where small files erase the saving.
- Inventory every process that reads whole trees and change it to metadata-only checks or AWS Backup.
- Apply
TransitionToIAat 30 days withAFTER_1_ACCESSon one file system, through infrastructure as code. - Graph bytes per class and burst credits for a full month and compare the bill with the model.
- Only then consider adding
TransitionToArchivefor data with a known, long retention.