Amazon FSx for Lustre is a managed deployment of Lustre, the parallel file system used by many of the world's largest supercomputers. It lets thousands of machines read and write the same files at an aggregate rate no single file server can reach, with ordinary POSIX semantics, so HPC codes and GPU training jobs need no SDK.
Used well, it keeps expensive GPUs busy; used carelessly, it is an over-provisioned bill or a bottleneck that looks like slow training. This article explains how Lustre spreads data, what the deployment types and storage classes actually trade, how throughput adds up from servers and stripes, how the S3 link works, and how to size and operate a file system for a real training cluster.
Why a parallel file system
A conventional network file system puts one server between every client and the data, and that server is everyone's ceiling. Lustre removes the middle. It separates the file system into metadata, the names, directories, permissions and file layouts, and data, the file contents, and spreads the data over many servers. A client asks the metadata server once where a file's pieces live, then moves data directly and in parallel to the servers that hold them, so bandwidth scales with the number of servers.
The moving parts
Four components make every metric and command legible:
- MDS and MDT. The metadata server runs on metadata targets, the volumes that hold the namespace. Every create, open, stat, rename and delete goes here. On Persistent 2 file systems you can provision metadata IOPS separately from capacity, and FSx adds a metadata server for every 12,000 provisioned metadata IOPS.
- OSS and OST. Object storage servers each serve several object storage targets, the volumes that hold file contents. Their number follows from capacity and throughput tier.
- Clients. A Linux kernel module that presents the whole file system as one mount. It caches locally and takes distributed locks so concurrent writers stay coherent.
- Layouts. Each file has a layout, stored in its metadata, that says which OSTs hold it and in what order. Striping a file across several OSTs is how one large file gets the bandwidth of several servers.
You can see all of it from a client: lfs df -h /fsx prints one line for each MDT and OST with its capacity and usage, and lfs getstripe prints any file's layout.
Deployment types and storage classes
You pick a deployment type and a storage class at creation. The deployment type decides durability and availability; the storage class decides latency and cost.
| Option | What you get | Use it for |
|---|---|---|
| Scratch 2 | No replication; failed servers are not replaced. 200 MBps per TiB baseline with bursts up to six times that. | Short jobs whose inputs live safely in S3 |
| Persistent 1 | Data replicated within the AZ; servers replaced automatically. SSD at 50, 100 or 200 MBps per TiB, or HDD at 12 or 40. | Existing file systems; created only via CLI or API |
| Persistent 2, SSD | Replicated within the AZ; 125, 250, 500 or 1000 MBps per TiB; provisioned metadata IOPS; optional EFA. | Long-lived training and HPC scratch with the highest throughput |
| Persistent 2, Intelligent-Tiering | Elastic capacity billed on data stored, tiered by last access; throughput in 4000 MBps increments; optional SSD read cache. | Large datasets where only part is hot |
First, scratch is not a cheaper persistent: on a scratch file system a failed server or disk gives clients an immediate I/O error for the files it held, and the risk grows with file system size. Use it only for regenerable data. Second, Intelligent-Tiering moves data untouched for 30 days to an Infrequent Access tier and after 90 days to an Archive Instant Access tier, and brings it back to Frequent Access when read.
How throughput actually adds up
The headline number is throughput per TiB, but the physical reality is servers. For Persistent 2 SSD without EFA, each OSS carries 2.4 TiB whatever the tier; the tier sets how fast that OSS runs. With EFA, higher tiers pack less storage per OSS: 38.4 TiB per OSS at 125 MBps per TiB down to 4.8 TiB per OSS at 1000. Either way, aggregate throughput is capacity times tier, delivered by some number of OSSs.
Three ceilings then decide what any one client sees. Traffic between one client instance and one OSS is capped at 5 Gbps. A client is also capped overall, at 100 Gbps over ENA and much higher over EFA. A file on one OST lives on one OSS, so a single process reading one unstriped file tops out near 5 Gbps no matter how large the file system is; the same client reading many files, or files striped across many OSTs, reaches its NIC limit. AWS recommends EFA for file systems above 10 GBps, and EFA file systems support GPUDirect Storage on NVIDIA GPU instances so data can land in GPU memory without a bounce through host memory.
Metadata is its own budget. Each provisioned metadata IOPS buys two file creates, opens or closes per second, one delete, 0.2 directory deletes or 0.1 directory creates or renames. In automatic mode an SSD file system gets 1,500 IOPS at 1.2 TiB, 12,000 from 12,000 GiB to 45,600 GiB, and 12,000 for every 24,000 GiB from 48,000 GiB up. Millions of small files, or a thousand ranks opening files at start-up, hit this ceiling long before bandwidth.
Striping and progressive file layouts
Striping deals a file out in 1 MiB chunks, by default, round-robin across a set of OSTs. A higher stripe count lets more servers serve one file and keeps it from filling one OST; for small files it only adds round trips, because operations such as reading the size must ask every OST in the layout.
File systems created after 25 August 2023 default to a four-component progressive file layout: the first 100 MiB of a file on one OST, up to 10 GiB on eight, up to 100 GiB on sixteen, and the rest on 32. Files grow into wider layouts as they grow, which suits mixed workloads. Two exceptions matter: files opened in append mode default to a stripe count of one, so logs do not fan out, and files imported from S3 use the file system's imported chunk size rather than the default layout.
# Where does this file live, and across how many OSTs?
lfs getstripe /fsx/train/shard-00017.tar
# Per-target usage: one MDT line, one line per OST
lfs df -h /fsx
# The post-2023 default layout, spelled out: 1 OST, then 8, 16 and 32 as files grow
lfs setstripe -E 100M -c 1 -E 10G -c 8 -E 100G -c 16 -E -1 -c 32 /fsx/mixed
# Checkpoint directory: every new file striped across 16 OSTs from the first byte
lfs setstripe -c 16 /fsx/checkpoints
# Existing files keep their layout; migrate to re-stripe (it copies the data)
lfs migrate -c 16 /fsx/checkpoints/step-120000/model.ptLayout is fixed at creation: set it on a directory before data arrives, because only new files inherit it.
Linking S3 with data repository associations
Most training data lives in Amazon S3, and FSx for Lustre's most useful feature is making an S3 prefix appear as a directory. A data repository association maps one file system directory to one bucket or prefix, one to one. A file system can have up to eight, and paths may not overlap on either side. DRAs are not available on Scratch 1 or on Lustre 2.10 file systems.
Import brings metadata first: the files appear with their sizes and permissions, and contents are loaded from S3 on first read. That lazy load is the most common cause of a slow first epoch. Export writes changed files back to S3 asynchronously after the application finishes modifying them. Both directions can run automatically per event type, but are off by default from the CLI and API.
# Link /datasets on the file system to an S3 prefix, both directions.
# Automatic import/export is OFF by default from the CLI and API, so set it explicitly.
aws fsx create-data-repository-association \
--file-system-id fs-0123456789abcdef0 \
--file-system-path /datasets \
--data-repository-path s3://example-training-data/imagenet/ \
--s3 "AutoImportPolicy={Events=[NEW,CHANGED,DELETED]},AutoExportPolicy={Events=[NEW,CHANGED,DELETED]}"
# Metadata is present after import; contents load lazily on first read.
# Warm the contents before a run so step one is not an S3 fetch:
find /fsx/datasets/train -type f -print0 | xargs -0 -n 1 -P 32 sudo lfs hsm_restore
# Is a given file's data resident, or only its metadata?
lfs hsm_state /fsx/datasets/train/shard-00017.tarData repository tasks do the same work on demand: export copies changes to S3, import loads metadata for chosen paths, and release drops the contents of exported files while keeping their metadata, so a later read fetches them again. FSx does not arbitrate conflicting writes to the same file in S3 and on the file system; if both sides change, coordinate in the application.
Worked example: a 64-node training cluster
A team trains on 40 TiB of tar shards, about 200 MiB each, stored in S3. There are 64 GPU instances, and profiling shows each needs about 150 MB/s of input to keep its GPUs fed, so input demand is roughly 9.6 GB/s. Every hour the job writes a checkpoint of about 1 TB, written in parallel by all ranks.
A Persistent 2 SSD file system of 48,000 GiB at 250 MBps per TiB gives about 11.7 GB/s aggregate, spread across 20 OSSs of 2.4 TiB each. Input needs about 80 percent of that. A 1 TB checkpoint at full bandwidth takes about 85 seconds, which means training pauses for that long unless checkpoints are written asynchronously. Automatic metadata mode gives 24,000 IOPS, or about 48,000 opens per second, far more than 64 nodes opening 200 MiB shards need. S3-imported shards take the 1 GiB imported chunk size, so each 200 MiB shard sits on one OST; reading many shards at once is what spreads each client over many OSSs.
Before the run, restore the training directory from S3 so no step waits on a lazy load. Set a stripe count of 16 on the checkpoint directory, and keep the last few checkpoints on FSx with export to S3 for durability. If input demand doubles, move to the 500 tier rather than buying storage you do not need. Just above 10 GBps, EFA is worth evaluating too.
Mounting and operating it
Clients need the Lustre client module that matches their kernel, and the security groups must allow Lustre traffic between clients and the file system (TCP 988 and 1018-1023). On Amazon Linux 2023 the client is a package; other distributions use AWS's repositories and a kernel-specific package.
# Amazon Linux 2023 client (check the kernel against the FSx client matrix first)
uname -r
sudo dnf install -y lustre-client
sudo mkdir -p /fsx
sudo mount -t lustre -o relatime,flock \
fs-0123456789abcdef0.fsx.us-east-1.amazonaws.com@tcp:/abcdef12 /fsxFSx publishes metrics per MDT and OST to CloudWatch every minute. Sum DataReadBytes across OSTs for aggregate read throughput and compare it with what you provisioned; watch FreeDataStorageCapacity per OST, because one full OST fails writes to any file striped on it even when the file system as a whole has room. Mount in the node bootstrap of AWS Batch or your scheduler so every node sees the same path.
Failure modes
- Slow first epoch. Lazy loading from S3 makes the first pass run at S3 speed. Restore before the job, then confirm with
lfs hsm_state. - One hot OST. Large files with a stripe count of one pile load and capacity onto one target. Set wider layouts and migrate the worst offenders.
- Metadata storms. Thousands of ranks listing a directory or opening the same config file at start-up queue on the MDS. Stage shared small files locally and provision more metadata IOPS if needed.
- Kernel and client mismatch. An AMI kernel update without a matching Lustre module leaves nodes unable to mount.
- Scratch surprises. A server failure on scratch loses the files it held. Treat scratch output as disposable until it is exported.
Trade-offs against the alternatives
Compared with Amazon EFS, FSx for Lustre gives far higher throughput per client and in aggregate, but keeps SSD and HDD data within one Availability Zone, needs a kernel client and is billed on provisioned capacity rather than on use, except with Intelligent-Tiering. Compared with reading S3 directly, it adds POSIX semantics, low latency and cached re-reads for multi-epoch training, at the cost of a second copy. Compared with local NVMe or EBS per node, it gives every node the same view, so jobs can be rescheduled freely and checkpoints are shared, at the cost of network hops and a shared bottleneck you must size.
What to do next
- Measure the input bandwidth each training node needs and multiply by node count; add checkpoint bandwidth and metadata operations per second.
- Choose the deployment type by durability need first: scratch only for fully regenerable data, Persistent 2 otherwise.
- Size capacity and tier together, then count OSSs and check the 5 Gbps per-client-per-OSS cap against how your readers access files.
- Create one data repository association per dataset prefix with explicit import and export policies, and restore datasets before each expensive run.
- Set progressive or wide layouts on checkpoint and large-file directories before any data is written, and check them with
lfs getstripe. - Add CloudWatch alarms on per-OST free capacity and on aggregate read throughput against provisioned throughput.
- Bake the matching Lustre client into your AMI and test mounts whenever the kernel changes.