AWS Snow Family solved a problem that has not gone away: some datasets are too large for the network you have. A rugged, encrypted storage appliance arrives at your site, you fill it, ship it back, and AWS imports it into Amazon S3. For a decade Snowcone, Snowball and Snowmobile covered sizes from terabytes to exabytes, and the compute variants ran EC2-compatible instances in places with poor connectivity.
The family has changed sharply. Snowcone and several Snowball models were discontinued in November 2024, and since 7 November 2025 Snowball Edge has been available only to existing customers. So this is not a buyer's guide. It is for teams that still run Snow jobs and need to run them well, and for everyone deciding what replaces it. Facts below were checked against the Snowball Edge developer guide and AWS's Snow device update post on 2 October 2026.
Where Snow stands in 2026
| Device | Status | What to know |
|---|---|---|
| Snowball Edge Storage Optimized 210 TB | Existing customers only | 210 TB usable NVMe, 104 vCPUs, 416 GB RAM; AWS quotes transfer up to 1.5 GB/s |
| Snowball Edge Compute Optimized | Existing customers only | 104 vCPUs, 416 GB usable RAM, 28 TB NVMe for instances |
| Snowball Edge 80 TB, 52 vCPU and GPU models | Discontinued 12 Nov 2024 | Support for devices already deployed ran to 12 Nov 2025 |
| Snowcone (HDD and SSD) | Discontinued 12 Nov 2024 | Use DataSync or the alternatives below |
| Snowmobile | Not offered | AWS no longer offers any Snow device to new customers |
For new customers AWS points to AWS DataSync for online transfer, AWS Direct Connect hosted connections to add bandwidth for the duration of a project, AWS Data Transfer Terminal (physical sites where you bring your own drives and upload over a fast link), partner offline-transfer services in AWS Marketplace, and AWS Outposts for edge compute. If your organisation has an existing Snow relationship, the devices still work and AWS says it continues to invest in their security and availability. Either way, the engineering questions are the same: is shipping really faster than the wire, and how do you prove every byte arrived?
The arithmetic: when shipping beats the wire
A network transfer takes data size divided by effective throughput. Effective throughput is lower than link speed: protocol overhead, competing traffic, source read speed and small files all take a share, and 60 to 80 percent of line rate is a realistic planning range for a well-run bulk copy. A device transfer takes shipping out, fill time, shipping back and AWS's import time, and the fill is itself a network transfer over your LAN.
TB = 1e12
def network_days(size_tb, link_gbps, efficiency=0.7):
bytes_per_s = link_gbps * 1e9 / 8 * efficiency
return size_tb * TB / bytes_per_s / 86400
def device_days(size_tb, fill_mb_s, ship_days_each_way=3, import_days=2, device_tb=210):
devices = -(-size_tb // device_tb) # ceiling
per_device_fill = min(size_tb, device_tb) * TB / (fill_mb_s * 1e6) / 86400
# devices filled in parallel if you have the source bandwidth; assume yes here
return ship_days_each_way * 2 + per_device_fill + import_days, int(devices)
for size in (50, 200, 600):
print(size, "TB",
"| 1 Gbps:", round(network_days(size, 1), 1),
"| 10 Gbps:", round(network_days(size, 10), 1),
"| device @500 MB/s:", device_days(size, 500))Worked example: 200 TB of archived video. At 70 percent efficiency, a 1 Gbps link moves about 87.5 MB/s, so 200 TB takes about 26 days if nothing else uses the link. At 10 Gbps it takes about 2.6 days. One 210 TB device filled at a sustained 500 MB/s needs 200e12 / 5e8 = 400,000 seconds, about 4.6 days, plus roughly a week of shipping and import. At AWS's best case of 1.5 GB/s the fill drops to about 1.5 days, but only if your source storage can read that fast and your network can carry it. The conclusion is not a fixed cut-over size. It is that a 10 Gbps path, or one added temporarily with a hosted Direct Connect connection, usually beats a device, while a site with 1 Gbps or less and hundreds of terabytes is where physical transfer pays off.
Two practical caps apply. The developer guide says the 15 free on-site days start the day after the device arrives, so a fill that will not finish in about 10 working days needs more devices or a faster source. And incremental data keeps arriving while the device is in transit, so plan an online catch-up sync (DataSync or your own tooling) for the delta between your copy cut-off and the cutover.
Running an import job, step by step
You create the job in the console or with the aws snowball create-job API, choosing the target bucket, a KMS key and an IAM role AWS uses to write into the bucket. The job then moves through states you can poll, which is worth automating so your runbook reacts to transitions instead of someone checking a console:
# Poll job state. JobState values: New, PreparingAppliance, PreparingShipment,
# InTransitToCustomer, WithCustomer, InTransitToAWS, WithAWSSortingFacility, WithAWS,
# InProgress (importing), Complete, Cancelled, Listing, Pending.
aws snowball describe-job --job-id "$JOB_ID" \
--query 'JobMetadata.JobState' --output textWhen the device arrives, connect power and one network port and read its IP address from the LCD. Only one network interface is usable at a time. AWS ships no optics or cables, and on the storage-optimized model the QSFP port runs at 40G, not 100G, so check that your switch and transceivers match before the device lands. Then unlock it with the manifest file and the 29-character unlock code from the console:
snowballEdge unlock-device --endpoint https://192.0.2.10 --manifest-file ./JID-manifest.bin --unlock-code 12345-abcde-12345-ABCDE-12345
# Store endpoint, manifest path and unlock code once, then refer to the profile
snowballEdge configure --profile dc1-snow
snowballEdge describe-device --profile dc1-snow # wait for UnlockStatus UNLOCKED
snowballEdge describe-service --service-id s3 --profile dc1-snow # adapter endpoints
# Local credentials for the S3 adapter (they only work against this device)
snowballEdge list-access-keys --profile dc1-snow
snowballEdge get-secret-access-key --access-key-id AKIAIOSFODNN7EXAMPLE --profile dc1-snowWith credentials in an AWS CLI profile, write through the S3 adapter as if it were a bucket. Parallelism is the biggest lever AWS lists: run several copy streams at once, from more than one host if one host cannot read fast enough.
EP=http://192.0.2.10:8080
aws s3 ls --endpoint "$EP" --profile snow
# Large files: plain parallel copies, one stream per top-level directory
ls /archive | xargs -P 8 -I{} aws s3 cp --recursive /archive/{} s3://media-raw/{} \
--endpoint "$EP" --profile snow
# Small files: tar on the fly; auto-extract restores the tree in S3 at import
tar -cf - /logs/2026/04 | aws s3 cp - s3://media-raw/batches/logs-2026-04.tar \
--metadata snowball-auto-extract=true --endpoint "$EP" --profile snowThe small-file rules are specific. AWS recommends files of at least 1 MB and no more than 500,000 entries in a directory. Batch smaller files as TAR, ZIP or tar.gz with the snowball-auto-extract=true metadata. Keep batches to about 10,000 files and under 100 GB, because batches larger than 100 GB are not extracted at all and arrive in S3 as one archive object. Keep files static during the copy, avoid renaming or rewriting them mid-transfer, and put the device, the source and the copy host on one switch.
The security model, and the mistakes that weaken it
Data on the device is encrypted with keys managed through the KMS key you chose for the job; the device is useless to whoever intercepts it without the manifest and unlock code. That is why the developer guide says not to store the unlock code next to the manifest, and why the client profile file, which holds both in plain text under ~/.aws/snowball/config/, needs strict permissions: keep the manifest on a server and send the code to the person doing the unlock through a separate channel. The access keys from list-access-keys are local to the device and do not map to any AWS account, so leaking them exposes this device's data while it is on site, not your cloud.
The physical design adds tamper evidence: hidden screws, intrusion switches, NFC tags checked with an AWS verification app, and anti-tamper inserts. If a device arrives looking tampered with, the guidance is to not connect it and to contact AWS Support for a replacement. On the cloud side, the role you give the job should allow writing only to the target bucket and prefix, and the bucket should have the controls you want before import starts, such as default encryption, Block Public Access and, for regulated archives, S3 Object Lock. Review the KMS key policy too; the KMS guide covers grants and key policies.
Proving the data arrived
A job reaching Complete means AWS finished importing what it could read. It does not mean every file you meant to send is in S3. The job report lists successes and failures, but the reliable proof is your own: build a manifest of every source file with size and a checksum before copying, then reconcile it against S3 after import.
import csv, hashlib, os
def build_manifest(root, out_csv):
with open(out_csv, "w", newline="") as f:
w = csv.writer(f)
for dirpath, _, files in os.walk(root):
for name in files:
path = os.path.join(dirpath, name)
h = hashlib.sha256()
with open(path, "rb") as fh:
for chunk in iter(lambda: fh.read(8 << 20), b""):
h.update(chunk)
w.writerow([os.path.relpath(path, root), os.path.getsize(path), h.hexdigest()])
# After import: join the manifest against an S3 Inventory report (key, size) for the bucket.
# Missing keys or size mismatches go to a re-copy list; spot-check hashes by downloading a sample.Join on key and size first; that catches the common failures of skipped files, truncated copies and unextracted batches cheaply. Then download a random sample and compare hashes. Only after reconciliation passes should anyone delete source data. AWS's own best-practice page says the same: do not delete local copies until the import has succeeded and you have verified it.
Edge compute on Snowball Edge, and its successor
The compute-optimized device runs EC2-compatible instances, and Amazon S3 compatible storage on Snowball Edge, at disconnected or remote sites: ships, field labs, temporary bases. Existing users can keep doing so. For new deployments AWS steers to Outposts: 2U servers or 42U racks that extend a VPC on premises and, per AWS, can run for up to seven days without a connection to the Region. That is a managed installation rather than a box rented per job, so evaluate it as infrastructure.
Choosing a path without Snow
| Situation | Use | Why |
|---|---|---|
| Good WAN, steady or repeated transfers | AWS DataSync | Incremental, verified, compressed transfer; one task can fill 10 Gbps; reads NFS, SMB, HDFS and S3-compatible sources |
| Bandwidth is the only blocker, for months | DataSync over a hosted Direct Connect connection | Rent capacity for the project instead of shipping hardware |
| Data already on portable drives, near a terminal | Data Transfer Terminal | Bring drives to an AWS site and upload over a fast link |
| Large one-off offline move, no Snow account | Partner offline transfer from AWS Marketplace | AWS names Seagate and Tsecond among the partners |
| Existing Snow customer, site under 1 Gbps | Snowball Edge 210 TB | Still the fastest path for hundreds of TB on a thin link |
For Hadoop estates, DataSync acting as an HDFS client is often simpler than staging data for a device; the Hadoop-to-cloud migration guide shows how the data lane fits with metadata and jobs. For hybrid links, the Direct Connect article explains hosted versus dedicated connections. The Storage Gateway article covers the case where you do not want to move data at all, only to tier it.
Failure modes and trade-offs
- Fill too slow. One copy stream, millions of tiny files, a 1G RJ45 link or a slow source turn a two-day fill into three weeks and burn the free days. Test throughput with a sample before ordering.
- Unextracted archives. A batch over 100 GB, an unsupported format or a missing metadata flag lands as one opaque object. Validate batch sizes in the script that creates them.
- Silent gaps. Files changed or added during the copy are missed. Freeze the source or record a cut-off time and run an online delta sync.
- Key and naming surprises. Paths with characters S3 tools handle awkwardly, or very deep trees, cause failures that show up only in the job's failure log. Scan names first.
- Credential handling. Manifest and unlock code in the same ticket defeat the two-factor design.
- Trade-off. Devices give predictable bulk throughput independent of the WAN, but add shipping latency, physical logistics and a hard dependency on a product that is closed to new customers. Design every new pipeline so the network path works without them.
What to do next
- Measure effective throughput on your real WAN path with a 1 TB sample, then run the arithmetic above for your actual data size.
- If you are not already a Snow customer, design around DataSync, a hosted Direct Connect connection or Data Transfer Terminal from the start.
- If you are, write the job runbook: port and optics check, unlock with separated credentials, parallel copy plan, small-file batching under 100 GB, and a fill deadline inside 15 days.
- Build the source manifest before copying and the reconciliation job before the device ships back.
- Lock down the target bucket, IAM role and KMS key policy before the import begins.
- For edge compute on Snowball Edge, start an Outposts evaluation now rather than when a device needs replacing.