AWS Storage Gateway is a virtual machine you run next to your applications that speaks a protocol they already know, NFS, SMB, iSCSI block or an iSCSI virtual tape library, and stores the data in AWS. Applications keep their mount points and backup jobs keep their tape drives, while the bytes end up in Amazon S3, EBS snapshots or S3 Glacier. The trick that makes it usable over a WAN is a local disk cache: reads of recent data are served locally and writes are acknowledged locally, then uploaded asynchronously.
That asynchrony is the whole story. Most Storage Gateway incidents come from forgetting that the local copy and the cloud copy are two different things with a lag between them: files a client closed that are not yet in S3, objects written to the bucket that clients cannot see, or a cache disk that fills because the uplink cannot drain it.
The four gateway types and what each one stores
Storage Gateway is one service with several personalities. You choose the type when you deploy the gateway, and each type decides both the client protocol and the shape of the data in AWS.
| Type | Client protocol | What lands in AWS | Typical use |
|---|---|---|---|
| S3 File Gateway | NFS v3/v4.1, SMB | One S3 object per file, readable by any S3 client | Landing zones for data lakes, backups written as files, archive shares |
| FSx File Gateway | SMB | A cache in front of Amazon FSx for Windows File Server | Existing customers only: not offered to new customers since 28 October 2024 |
| Volume Gateway, cached | iSCSI block | Volume data in AWS-managed S3 storage, point-in-time EBS snapshots | Block storage for on-premises servers with a small local footprint |
| Volume Gateway, stored | iSCSI block | Full copy stays local; asynchronous EBS snapshots in AWS | Low-latency primary data with cloud backup and DR |
| Tape Gateway | iSCSI VTL (media changer and tape drives) | Virtual tapes in S3, archived to S3 Glacier Flexible Retrieval or Deep Archive | Replacing physical tape under existing backup software |
The key distinction is whether the cloud copy is native. S3 File Gateway writes ordinary objects: a file at /share/reports/q3.csv becomes the key reports/q3.csv, and file metadata such as ownership and permissions travel as object metadata. Analytics services can read those objects directly. Volume and tape data is not browsable in S3; you reach it only by restoring a snapshot to an EBS volume, mounting it through a gateway, or retrieving a tape. That is why file gateways suit data you want to process in the cloud, while volume and tape gateways suit data you want to protect.
The data path: cache, upload buffer and asynchronous upload
Every gateway has local disks that you allocate when you deploy it. File and cached-volume gateways have a cache disk, and volume and tape gateways also have an upload buffer. The data path works like a write-back cache.
- Write. A client writes over NFS, SMB or iSCSI. The gateway persists the data on its local disk and acknowledges it. At this moment the data exists only on premises.
- Upload. In the background the gateway uploads dirty data. For file shares it uploads whole files as objects after they are written, and it may upload a large file in parts. For volumes it uploads changed blocks; for tapes, written tape data.
- Read. A read hits the cache if the data is local. On a miss the gateway fetches it from S3, serves it and keeps it cached. Least-recently-used data is evicted, but only once it has been uploaded; dirty data is never evicted.
- Snapshot or archive. Volume Gateway turns uploaded blocks into EBS snapshots on a schedule or on demand. Tape Gateway moves a tape to its archive pool when backup software ejects it.
Two consequences follow. First, the gateway's disk is a durability dependency until upload completes. Put cache and upload buffer on resilient storage, not a single local SSD with no protection. Second, sustained writes faster than the uplink can drain them will fill the cache or buffer. When that happens the gateway slows or blocks writes, and applications see latency spikes or I/O errors. The CloudWatch metrics worth alarming on are CachePercentDirty, the share of the cache not yet persisted to AWS, CachePercentUsed, upload-buffer usage on volume and tape gateways, and FilesFailingUpload on file gateways.
File shares, the metadata inventory and why direct S3 writes are invisible
A file system needs fast directory listings, but listing a bucket prefix is a slow, paged API call. So S3 File Gateway keeps an inventory of the objects it knows about for each directory. When a client lists a directory for the first time, the gateway lists that prefix in S3 and caches the result; time grows with the number of entries. After that, the gateway updates the inventory from its own writes.
The gateway does not watch the bucket. If another application, a replication rule or a person in the console writes an object under the share's prefix, clients do not see it until the inventory is refreshed. There are two mechanisms. A TTL-based automated refresh on the file share re-lists a directory when it is accessed after the TTL has expired; it is non-recursive, and NFS or SMB operations on that directory block while it runs, so choose the longest TTL you can tolerate. The RefreshCache API refreshes named folders, optionally recursively, and runs alongside normal traffic. It only starts the work: completion is signalled by a refresh-complete notification, and a new refresh has to wait for the running one to finish. Both update the inventory; neither pulls file contents into the cache.
import boto3, time
sgw = boto3.client("storagegateway")
share_arn = "arn:aws:storagegateway:eu-west-1:111122223333:share/share-EXAMPLE"
# 1. A producer outside the gateway wrote new objects under s3://bucket/ingest/2026-10-01/.
# Refresh only that directory; a recursive refresh of the whole share costs LIST calls
# proportional to the number of objects and competes for gateway CPU.
resp = sgw.refresh_cache(
FileShareARN=share_arn,
FolderList=["/ingest/2026-10-01"],
Recursive=False,
)
print("refresh started:", resp["NotificationId"])
# 2. The API only STARTS the refresh. Consumers should wait for the refresh-complete
# event on EventBridge (matched on NotificationId) instead of polling the mount.
# 3. In the other direction: a writer on NFS has closed its files and wants to know
# when they are durable in S3 before telling a downstream job to read the bucket.
up = sgw.notify_when_uploaded(FileShareARN=share_arn)
print("upload notification requested:", up["NotificationId"])For inbound objects, refresh only the prefix that changed. For outbound files, NotifyWhenUploaded asks for an event once every file written before the call has reached S3, which is the correct signal to start a downstream job. Route both events through an EventBridge rule:
{
"source": ["aws.storagegateway"],
"detail-type": ["Storage Gateway File Upload Event",
"Storage Gateway Refresh Cache Event"]
}Gateways sharing a bucket do not lock against each other, and concurrent writes to one file end as last writer wins. Keep one writer per prefix; other gateways read and refresh.
Worked example: sizing a cached Volume Gateway
A branch office has a 20 TiB iSCSI LUN for a file server. Array statistics show about 400 GiB of blocks rewritten per day, almost all of it during a one-hour nightly batch. The site has a 1 Gbit/s internet link shared with other traffic. The question is whether the upload buffer and cache are big enough.
# Worked sizing for a cached Volume Gateway backing a 20 TiB iSCSI LUN
daily_change = 400 GiB # blocks rewritten per day (from array statistics)
peak_write_window = 1 h # nightly batch writes it: ~114 MiB/s
uplink = 1 Gbit/s # ~110 MiB/s usable, shared with other traffic -> plan 70 MiB/s
drain_time = 400 GiB / 70 MiB/s ~= 98 min
backlog at peak = written - drained in 1 h = 400 GiB - 246 GiB ~= 154 GiB
upload_buffer >= peak backlog with ~3x headroom -> provision 500 GiB (max 2 TiB)
cache >= hot working set + un-uploaded data, e.g. 1.5 TiB (range 150 GiB to 64 TiB)The arithmetic says the link drains a night's changes in about an hour and a half. But the batch writes at about 114 MiB/s for an hour, faster than the link drains, so about 154 GiB is still waiting when it ends. Provision headroom, here 500 GiB. Size the cache from the working set, not the volume: if users touch about a terabyte in a week, a 1.5 TiB cache keeps reads local. Check the result against the quotas: a cached volume can be up to 32 TiB, a gateway holds up to 32 volumes and 1,024 TiB in total, cache disks range from 150 GiB to 64 TiB and the upload buffer from 150 GiB to 2 TiB. Stored volumes are smaller, up to 16 TiB each and 512 TiB per gateway, because the full copy lives on site.
If the arithmetic does not close, you have three options: more bandwidth, such as AWS Direct Connect; bandwidth throttling schedules, which move uploads out of business hours but lengthen the backlog; or less data, for example by excluding scratch volumes.
Networking, security and deployment
- Where it runs. The gateway ships as a VM image for VMware ESXi, Microsoft Hyper-V or Linux KVM, or as an Amazon EC2 AMI.
- Activation and endpoints. Activation associates the VM with your account and Region. After that the gateway talks outbound over HTTPS to the Storage Gateway service and S3. To keep that traffic private, use VPC interface endpoints; the PrivateLink article covers the pattern.
- Client access. Restrict NFS shares by allowed client CIDRs and squash settings. Join the gateway to Active Directory for SMB shares and use ACLs.
- Encryption and bucket policy. In S3, choose SSE-S3 or SSE-KMS per share; AWS KMS key policies then control which roles can read what the gateway wrote. The gateway accesses the bucket through an IAM role you grant, so scope it to the share's bucket and prefix.
- Lifecycle interplay. S3 lifecycle rules still apply to objects written by a file gateway. Transitioning objects to S3 Glacier Flexible Retrieval or Deep Archive makes the gateway fail to read or update them, reported in its health log as an
InaccessibleStorageClasserror, so exclude active prefixes. Enabling versioning means every overwrite from a client leaves an old version behind.
Tape Gateway specifics
Tape Gateway exposes a virtual media changer and virtual tape drives over iSCSI, so backup applications that already drive tape libraries write to it without changes. A virtual tape can be up to 15 TiB, and a gateway can hold up to 1,500 tapes and 1 PiB. Tapes being written live in the gateway's cache and in S3; when the backup software ejects a tape, the gateway moves it to the archive pool you chose, S3 Glacier Flexible Retrieval or S3 Glacier Deep Archive.
Retrieval is where plans go wrong. An archived tape is not readable until you retrieve it to a gateway, which takes hours and is quicker from Flexible Retrieval than from Deep Archive; check the current documentation for times before writing a restore SLA. Deep Archive is cheaper to store and has a longer minimum storage duration, so tapes deleted early still incur charges. Tape pools support retention lock and write-once-read-many settings, the tape equivalent of object locking in Amazon S3, which protects backups from deletion by a compromised backup server.
Failure modes and how to handle them
- Cache full. Symptom: rising write latency, then I/O errors. Cause: writes outpace upload, or the cache is smaller than the dirty set. Fix: add a cache disk, which a gateway can accept online, and throttle or reschedule the writer.
- Upload backlog never clears. The uplink is too small or a bandwidth schedule is too strict. Alarm on the trend of
CachePercentDirtyrather than a single reading, because a steady dirty fraction after a batch is normal and one that keeps growing is not. - Gateway VM lost before upload. Data acknowledged but not uploaded is gone. For file shares, recover by deploying a new gateway on the same bucket and re-running jobs from the last notified upload. For volumes, restore from the last snapshot. This is the argument for resilient hypervisor storage and for taking snapshots before risky batches.
- Stale listings. Clients do not see objects that another system wrote. Add targeted RefreshCache calls to the producer's pipeline instead of setting a very short TTL.
- Software updates. AWS applies gateway software updates in a maintenance window you set, and the gateway restarts. Keep the window clear of backup jobs; clients see the restart as a brief outage.
When to use something else
For a one-off migration, AWS DataSync copies data with verification and no cache. If the application can move to AWS, Amazon EFS or FSx removes the on-premises component altogether. The gateway fits applications that must stay put and speak file, block or tape protocols.
What to do next
- Pick the gateway type by the cloud format you need: native objects (S3 File), snapshots (Volume) or archived tapes (Tape). New SMB-to-FSx designs should use FSx for Windows File Server directly.
- Measure daily change, peak write rate and working set, then size cache and upload buffer with the arithmetic above and at least 2x headroom.
- Put gateway disks on resilient hypervisor storage and alarm on cache and upload-buffer percentage used and on upload backlog.
- Adopt one-writer-per-prefix. Wire producers to call RefreshCache on the prefixes they change, and consumers to wait for NotifyWhenUploaded events.
- Scope the gateway's IAM role and KMS key to its bucket, and route traffic through VPC endpoints.
- Set the maintenance window outside backup jobs, and rehearse a restore: mount a volume snapshot or retrieve an archived tape, and record how long it takes.