Most shared file systems follow the same curve. A project writes a burst of files, reads them hard for a few weeks, and then barely touches them again. Months later they still sit on the most expensive storage tier, because nobody wants to work out which files are safe to move. Amazon EFS handles this with lifecycle management. A per-file access timer moves cold data to cheaper storage classes, and an optional rule brings it back the moment someone reads it. AWS calls that combination EFS Intelligent-Tiering.

This article explains how the mechanism actually works: what counts as an access, what the three policies do, how tiering interacts with throughput modes and small files, and how to model whether it saves money for your data. It ends with a rollout plan you can run against a real file system. For EFS fundamentals such as mount targets, performance modes and NFS semantics, start with the AWS EFS overview.

What Intelligent-Tiering means on EFS

EFS Intelligent-Tiering is not a storage class and not a switch. It is the name AWS gave, when it launched the feature in September 2021, to one particular lifecycle configuration: a policy that demotes files after N days without access, plus TransitionToPrimaryStorageClass set to AFTER_1_ACCESS, which promotes a file back to Standard on its first access. Without the second policy you have plain lifecycle management. Demoted files stay in the cheaper class even when read, and every read pays the slower first byte and an access charge.

This is unlike S3 Intelligent-Tiering, which is a storage class of its own with a per-object monitoring fee and automatic access tiers. On EFS there is no monitoring fee and no separate class. The whole behaviour comes from three policies applied to the file system as a whole, never to a single directory. If two datasets need different rules, they need two file systems.

Three storage classes and their billing rules

EFS has three storage classes. The figures below come from the EFS user guide; prices change by Region and over time, so read them from the pricing page rather than from any article.

ClassDesigned forFirst-byte latencyMinimum billed size per fileMinimum storage duration
StandardActive dataSub-millisecondNoneNone
Infrequent Access (IA)Data read a few times a quarterTens of milliseconds128 KiBNone
ArchiveData read a few times a year or lessTens of milliseconds128 KiB90 days

Three billing facts drive every decision later in this article. First, you pay for data in each class. Second, you also pay data access charges when files in IA or Archive are read, and for data that lifecycle management moves between classes. Third, IA and Archive bill at least 128 KiB per file, so a 4 KiB file costs as if it were 128 KiB. Archive needs Elastic throughput, and the storage-class table lists it only for Regional file systems.

EFS StandardSSD, sub-ms first byteInfrequent Accesstens of ms first byteArchivetens of ms, 90-day minTransitionToIATransitionToArchiveTransitionToArchive also applies to files still in StandardAFTER_1_ACCESS: first data read or write promotes the file backMetadata always stays in Standardnames, owners, directory tree; ls and stat do not reset or trigger anythingNew writes land in Standardeligible to tier again after 24 hoursBilling in IA and Archive128 KiB minimum per file, access chargesTimer = last access in Standard, tracked internally, not POSIX atime
EFS lifecycle: files demote on a per-file access timer and, with AFTER_1_ACCESS, promote back on first access. Metadata never leaves Standard.

How the lifecycle timer decides

Lifecycle decisions are based on when a file was last accessed in Standard. EFS keeps an internal timer for this; it is not the POSIX atime you can see with stat, so mounting with noatime changes nothing. Each access to a file in Standard resets the timer. When the timer passes the policy age, the file becomes a candidate to move.

Several details matter in practice. Metadata, meaning file names, ownership and the directory tree, always stays in Standard, and metadata operations such as listing a directory do not count as access. Backup tools that only stat files are therefore harmless, but a tool that opens and reads every file is not. Transitions run at lower priority than your workload, so moving millions of small files takes longer than moving fewer large files of the same total size. While a file is being moved it is still billed as Standard. Writes to a file in IA or Archive first land in Standard and become eligible to move again after 24 hours.

The practical consequence: on a busy file system, tiering is eventual. Expect the Standard byte count to fall over days, not minutes, after you enable a policy.

The three policies and how to set them

A lifecycle configuration holds up to three policies, and the API insists that each LifecyclePolicy object carries only one transition. The valid values come from the EFS API reference.

  • TransitionToIA: AFTER_1_DAY, AFTER_7_DAYS, AFTER_14_DAYS, AFTER_30_DAYS, AFTER_60_DAYS, AFTER_90_DAYS, AFTER_180_DAYS, AFTER_270_DAYS or AFTER_365_DAYS.
  • TransitionToArchive: the same set of values, measured from last access in Standard. It applies to files in Standard or IA.
  • TransitionToPrimaryStorageClass: only AFTER_1_ACCESS. Omit it and files stay where they are when read.

A file system created in the console with recommended settings gets 30 days to IA, 90 days to Archive and no transition back. That default suits archives, not workloads that return to old data. To turn on Intelligent-Tiering behaviour, write all three policies explicitly:

aws efs put-lifecycle-configuration \
  --file-system-id fs-0123456789abcdef0 \
  --lifecycle-policies \
    '[{"TransitionToIA":"AFTER_30_DAYS"},
      {"TransitionToArchive":"AFTER_180_DAYS"},
      {"TransitionToPrimaryStorageClass":"AFTER_1_ACCESS"}]'

# Turn lifecycle management off entirely: an empty policy list.
aws efs put-lifecycle-configuration \
  --file-system-id fs-0123456789abcdef0 \
  --lifecycle-policies '[]'

The call replaces the whole configuration rather than merging with it, so always send the full set. Keep the configuration in code next to the file system definition, so a console edit shows up as drift instead of going unnoticed.

Interaction with throughput modes

Tiering interacts with throughput in two ways, and both can hurt.

Bursting throughput shrinks. In Bursting mode, baseline throughput and burst credits scale with the amount of data in Standard only. If tiering moves 90 percent of your bytes to IA, your baseline drops by the same proportion. A build farm that ran happily on bursting credits can start throttling the week after you enable lifecycle management. Check the BurstCreditBalance trend before and after, or move to Elastic throughput.

Archive pins the throughput mode. Archive is supported only with Elastic throughput. Once a lifecycle policy transitions data to Archive, you cannot switch the file system to Bursting or Provisioned. Elastic bills for the metadata and data you actually transfer, so it is usually the right choice for spiky, mostly cold file systems anyway. Make that decision deliberately rather than discovering the lock later.

Small files and the 128 KiB minimum

The 128 KiB minimum billed size is where tiering most often backfires. Lifecycle policies updated on or after 26 November 2023 do tier files smaller than 128 KiB, and each one is then billed as 128 KiB in IA or Archive. Suppose a directory holds ten million 8 KiB files, about 76 GiB of real data. In IA it is billed as about 1.2 TiB. Whether that saves money depends on the price ratio between Standard and IA in your Region, and for very small files it often does not.

Before enabling a policy, build a size histogram of the file system. If most bytes sit in large files and most files are tiny, tiering still helps. If tiny files dominate both counts, consider packing them into archives such as tar or Parquet, or keep that tree on a file system without lifecycle policies.

import os, collections

def size_profile(root):
    # bucket file sizes in powers of two; count files and bytes per bucket
    files, bytes_ = collections.Counter(), collections.Counter()
    for dirpath, _, names in os.walk(root):
        for n in names:
            try:
                s = os.lstat(os.path.join(dirpath, n)).st_size
            except FileNotFoundError:
                continue
            b = max(s, 1).bit_length()
            files[b] += 1
            bytes_[b] += s
    for b in sorted(files):
        print(f"<= {2**b:>14,} B  files={files[b]:>10,}  bytes={bytes_[b]:>16,}")

Walking with lstat is metadata-only, so the scan does not reset any lifecycle timers.

Worked example: does tiering pay?

Tiering is worth it when the storage you save beats the access and transition charges you add. The model below takes prices as inputs, because they vary by Region and change over time; fill them in from the EFS pricing page. Sizes are in GiB per month.

from dataclasses import dataclass

@dataclass
class Prices:                 # per GiB, from the EFS pricing page for your Region
    standard: float
    ia: float
    ia_read: float            # access charge per GiB read from IA
    tier_move: float          # charge per GiB moved by lifecycle management

MIN_BILLED_KIB = 128

def billed_gib(n_files, real_gib):
    # each file in IA is billed at least 128 KiB
    avg_kib = real_gib * 1024 * 1024 / max(n_files, 1)
    return real_gib * max(1.0, MIN_BILLED_KIB / max(avg_kib, 1e-9))

def monthly(p, total_gib, cold_frac, cold_files, cold_read_gib, promote_back):
    cold = total_gib * cold_frac
    hot = total_gib - cold
    base = total_gib * p.standard
    tiered = hot * p.standard + billed_gib(cold_files, cold) * p.ia
    tiered += cold * p.tier_move / 12          # one-off move, amortised over a year
    if promote_back:                            # read once, moves back to Standard
        tiered += cold_read_gib * (p.ia_read + p.standard - p.ia)
    else:                                       # every read pays the access charge
        tiered += cold_read_gib * p.ia_read
    return base, tiered

Run it with real numbers. A 20 TiB research share where 85 percent of bytes have gone untouched for 30 days, the cold data sits in about two million files averaging 9 MiB, and users pull back 200 GiB a month. The 128 KiB floor costs nothing here, since the average file is far larger, and the saving tracks the Standard-to-IA price gap almost exactly. Change the cold set to fifty million 32 KiB files and the billed size quadruples, which wipes out most of the saving. The model is crude, since it ignores Archive and assumes a steady state, but it answers the question that matters: is the cold data big files or small ones?

Operating a tiered file system

Lifecycle management is easy to switch on and hard to observe. Three habits make it predictable.

  1. Watch bytes per class. DescribeFileSystems returns SizeInBytes with separate Standard, IA and Archive values, and the CloudWatch StorageBytes metric is split by storage class. Graph them. A sudden rise in Standard on a tiered file system means something is reading old data, and you will pay for it twice: once in access charges and again in Standard storage.
  2. Find the full-scan readers. Antivirus scanners, rsync --checksum, content indexers and some backup agents open every file. With AFTER_1_ACCESS they promote the whole file system back to Standard on every run. Point them elsewhere, move them to metadata-only modes, or use AWS Backup, which does not incur data access charges on lifecycle-managed file systems.
  3. Treat policy changes as deployments. Shortening TransitionToIA from 90 to 7 days moves a large slice of data at once and bills the transitions. Roll changes out on one file system, watch a full cycle, then widen.

For training pipelines that read whole datasets every epoch, EFS tiering is the wrong lever; keep hot data on a high-throughput file system such as FSx for Lustre and leave the long tail on S3.

Failure modes

  • Bounce loops. A nightly job reads the whole tree, everything promotes, and 30 days later it all demotes again. You pay transitions and access charges every month for nothing. Fix the reader, not the policy.
  • Throttling after enablement. A Bursting file system loses baseline as bytes leave Standard. Symptom: rising latency and a draining credit balance with no change in the workload.
  • Small-file bill shock. Millions of tiny files tier and are billed at 128 KiB each. Symptom: the IA line on the bill is far larger than the data you think you have.
  • Latency-sensitive tail. A web app serving rarely viewed assets from EFS sees tens of milliseconds of extra first-byte latency on cold files. Without AFTER_1_ACCESS, every read pays it.
  • Archive surprises. Data deleted or rewritten within 90 days of reaching Archive is still billed for the minimum duration, and the throughput mode can no longer change.

Trade-offs

ChoiceGainCost
Short IA window (7-14 days)Larger saving on bursty dataMore transitions; more first reads pay slower latency
Long IA window (60-90 days)Few wrong demotionsLess saving
AFTER_1_ACCESS onReturning data is fast againScanners can undo tiering
AFTER_1_ACCESS offCold data stays cheap even if read onceEvery read of cold data pays access charges and latency
Archive enabledLowest storage price on EFSElastic throughput only, 90-day minimum
Move to S3 insteadCheaper still, object lifecycleApplication must speak S3, not NFS

If the data no longer needs POSIX semantics, the strongest move is often to leave EFS entirely: export it to S3 and let S3 lifecycle rules handle it. Storage Gateway can keep an NFS front end for applications that cannot change. For block workloads, the EBS volume types are the comparison point instead.

What to do next

  1. Run the size profiler against each EFS file system and note what share of files and bytes are under 128 KiB.
  2. Read current SizeInBytes per class and the throughput mode for each file system.
  3. For Bursting file systems, decide whether to move to Elastic before enabling tiering.
  4. Fill the cost model with your Region's prices and your cold fraction; skip file systems where small files erase the saving.
  5. Inventory every process that reads whole trees and change it to metadata-only checks or AWS Backup.
  6. Apply TransitionToIA at 30 days with AFTER_1_ACCESS on one file system, through infrastructure as code.
  7. Graph bytes per class and burst credits for a full month and compare the bill with the model.
  8. Only then consider adding TransitionToArchive for data with a known, long retention.
Key takeaway: EFS Intelligent-Tiering is a lifecycle configuration, not a storage class: demote after N days without access and promote back on first access. It pays when cold data is large files that are rarely read. Profile file sizes first, watch the Bursting baseline, keep full-tree scanners away, and roll policies out one file system at a time.