Copying a few files to the cloud is easy. Copying 80 TB and a hundred million files from a production NAS, keeping permissions and timestamps, proving every byte arrived intact, throttling so the office network still works, and then keeping the copy in sync until cutover is not. AWS DataSync is the managed service for that job. It moves files and objects between on-premises storage (NFS, SMB, HDFS, object stores), other clouds and AWS storage (Amazon S3, EFS, the FSx file systems), handles parallelism, retries, metadata and integrity checks, and charges by the gigabyte copied.
This page explains how DataSync works and how to run it well. You will learn its four building blocks; how Basic and Enhanced task modes differ and which to choose; what happens inside a task execution; which options decide whether a transfer is safe; how to size a migration with a worked example; how to drive it from code; how to monitor it; and the failure modes that cause most incidents.
The building blocks
Four resources make up every transfer.
- Agent. A virtual machine you run close to the source storage. It reads and writes the storage, and talks to the DataSync service over TLS. Agents run on VMware ESXi, KVM, Microsoft Hyper-V or as an Amazon EC2 instance from a DataSync AMI. Transfers between AWS storage services, and Enhanced mode transfers between S3 and other object stores, need no agent.
- Location. A description of one end of a transfer: an NFS export and the agents that can reach it, an SMB share and its credentials, an S3 bucket, prefix, storage class and the IAM role DataSync uses, an EFS or FSx file system, and so on.
- Task. A source location, a destination location, options, filters or a manifest, an optional schedule and logging configuration. The task mode, Basic or Enhanced, is chosen at creation and cannot be changed later.
- Task execution. One run of a task. Each run discovers what needs copying, copies it, and verifies it. Executions of one task queue behind each other, up to 50 queued.
Basic mode and Enhanced mode
Basic mode was the original design and still covers every location type. Enhanced mode, added later, lists, prepares, transfers and verifies in parallel instead of in sequence, and removes the per-execution file-count quota. It is available between S3, EFS and FSx for Lustre without an agent; between NFS, SMB or HDFS and supported AWS storage with an Enhanced mode agent; and between Azure Blob or other object stores and AWS storage, with no agent needed when S3 is the AWS end. Anything else, such as FSx for Windows File Server, needs Basic mode.
| Basic mode | Enhanced mode | |
|---|---|---|
| Items per execution | 50 million to or from on-premises and other clouds; 25 million between AWS services (adjustable) | Virtually unlimited |
| How phases run | Prepare, then transfer, then verify | In parallel |
| Default verification | POINT_IN_TIME_CONSISTENT (whole dataset) | ONLY_FILES_TRANSFERRED; full scan not supported |
| Logs | Unstructured | Structured JSON, more counters |
| Agent sizing (VM) | 4 vCPU, 32 GB RAM; 64 GB above 20 million items | 8 vCPU, 32 GB RAM |
| Price at time of writing (US East) | $0.0125 per GB copied | $0.015 per GB copied, plus a per-execution charge |
Choose Enhanced whenever the location pair supports it and the dataset is large or contains many small files; the parallel pipeline is faster and you avoid splitting the job to stay under quotas. Choose Basic when you need a location Enhanced does not support, or when you want the end-to-end full-dataset verification that only Basic offers. Throughput per task is capped at 10 Gbps with an agent and 5 Gbps without, so very large migrations run several tasks side by side.
Inside a task execution
An execution moves through a fixed set of states: QUEUED, LAUNCHING, PREPARING, TRANSFERRING, VERIFYING, then SUCCESS or ERROR. Preparing is where DataSync lists the source and, with the default TransferMode=CHANGED, compares it with the destination by size, modification time and metadata to decide what to copy. On a first run everything is new; on later runs only differences move, which is what makes repeated incremental runs before cutover cheap.
During transfer, data is checksummed as it is read and checked as it is written. Verification then adds an end-of-run check: either of just the files transferred, or, in Basic mode, a comparison of the entire source and destination. A file that changes while being copied shows up as a verification failure, which is DataSync telling you the source was live, not that it corrupted anything.
Listing is expensive on file systems with many small files. A NAS with 100 million files may take hours just to walk, regardless of how little changed. Enhanced mode overlaps listing with copying; in Basic mode, shrinking the scan with filters or a manifest is the main tool.
Options that decide whether a transfer is safe
| Option | Default | When to change it |
|---|---|---|
TransferMode | CHANGED | ALL skips the destination comparison. Rarely what you want for repeated runs |
PreserveDeletedFiles | PRESERVE | REMOVE mirrors deletions so the destination matches the source. Cannot be combined with TransferMode=ALL |
OverwriteMode | ALWAYS | NEVER protects files edited at the destination, for example after a partial cutover |
VerifyMode | Per mode, see above | ONLY_FILES_TRANSFERRED for archive storage classes, where full verification is not allowed |
BytesPerSecond | -1 (unlimited) | Set during business hours, then override per execution at night |
PosixPermissions, Uid, Gid | Preserve | Set to NONE when the destination has no meaning for numeric owners |
LogLevel | Set per task | TRANSFER during migration so every file has a record; BASIC afterwards |
ObjectTags | PRESERVE | NONE for object stores without tagging; in Enhanced mode the run fails fast otherwise |
Two options cause most surprises. PreserveDeletedFiles=REMOVE is correct for a final mirror and dangerous everywhere else: point it at the wrong destination prefix and it deletes whatever is there that the source lacks. And writing to S3 archive storage classes changes the economics of every overwrite and delete, because early deletion and retrieval fees apply; DataSync's own guidance is to verify only transferred data for those classes.
Worked example: an 80 TB NAS migration
A team migrates an 80 TB NFS share holding 120 million files to S3 over a 1 Gbps link they can use at about 80 percent outside business hours and 30 percent during them.
Mode. 120 million items exceeds Basic mode's 50 million per execution, so Basic would mean splitting into at least three tasks by directory with include filters, each with a 64 GB agent. An Enhanced mode agent handles it in one task with no file quota. Choose Enhanced.
Time. 1 Gbps is 125 MB per second; at 80 percent that is 100 MB per second, or about 8.6 TB per day. 80 TB therefore needs roughly nine and a half days of full-rate transfer, longer with daytime throttling, and many small files lower effective throughput further. Plan two to three weeks for the initial copy. If that is too slow and the link cannot grow, an offline device from the Snow family for the bulk followed by DataSync for the deltas is the usual alternative.
Cost. 80,000 GB at $0.015 is about $1,200 in DataSync charges, plus S3 request and storage charges, plus a small fee per execution. Daily incremental runs afterwards cost only the changed gigabytes.
Cutover. Run incrementals nightly until each one finishes well inside the cutover window. Then freeze writes, run a final incremental, check the task report for zero failures, and repoint clients.
Driving DataSync from code
The script below creates the locations and an Enhanced mode task with logging, an exclusion filter and a task report, starts a run with a bandwidth override, and polls until it finishes. Parameter names are from the DataSync API.
import time
import boto3
ds = boto3.client("datasync", region_name="us-east-1")
AGENT = "arn:aws:datasync:us-east-1:111122223333:agent/agent-0123456789abcdef0"
ROLE = "arn:aws:iam::111122223333:role/DataSyncS3Access"
src = ds.create_location_nfs(
ServerHostname="nas01.corp.example", Subdirectory="/export/projects",
OnPremConfig={"AgentArns": [AGENT]}, MountOptions={"Version": "NFS4_1"})
dst = ds.create_location_s3(
S3BucketArn="arn:aws:s3:::corp-projects-archive", Subdirectory="/projects",
S3StorageClass="INTELLIGENT_TIERING", S3Config={"BucketAccessRoleArn": ROLE})
task = ds.create_task(
Name="nas01-projects-to-s3", TaskMode="ENHANCED",
SourceLocationArn=src["LocationArn"], DestinationLocationArn=dst["LocationArn"],
CloudWatchLogGroupArn="arn:aws:logs:us-east-1:111122223333:log-group:/datasync",
Options={"TransferMode": "CHANGED", "VerifyMode": "ONLY_FILES_TRANSFERRED",
"PreserveDeletedFiles": "PRESERVE", "LogLevel": "TRANSFER"},
Excludes=[{"FilterType": "SIMPLE_PATTERN", "Value": "*/.snapshot|*.tmp"}],
TaskReportConfig={"Destination": {"S3": {"S3BucketArn": "arn:aws:s3:::corp-datasync-reports",
"Subdirectory": "reports/", "BucketAccessRoleArn": ROLE}},
"OutputType": "STANDARD", "ReportLevel": "ERRORS_ONLY"})
run = ds.start_task_execution(TaskArn=task["TaskArn"],
OverrideOptions={"BytesPerSecond": 40 * 1024 * 1024})
arn = run["TaskExecutionArn"]
while True:
d = ds.describe_task_execution(TaskExecutionArn=arn)
print(d["Status"], d.get("FilesTransferred", 0), d.get("BytesWritten", 0))
if d["Status"] in ("SUCCESS", "ERROR"):
break
time.sleep(60)
failed = d.get("FilesFailed", {})
if d["Status"] == "ERROR" or any(failed.values()):
raise SystemExit(f"run failed: {d.get('Result', {})} failed={failed}")The exit check matters: an execution can reach SUCCESS with individual files that failed, so read FilesFailed (Enhanced mode) and the task report, not only the status. For a nightly schedule, add Schedule={"ScheduleExpression": "cron(0 1 * * ? *)"} to the task instead of running a poller.
Networking, monitoring and security
Networking. When you activate an agent you choose one service endpoint type: public, FIPS, or a VPC endpoint through AWS PrivateLink so traffic stays on a Direct Connect or VPN path. The agent needs outbound HTTPS to the service. Keep the agent on the same fast network segment as the NAS; a slow hop between agent and storage caps everything.
Monitoring. Send logs to CloudWatch Logs, watch bytes and files transferred metrics, and route execution state changes from EventBridge to whoever owns the migration. Write task reports to S3: they list every file transferred, skipped, verified or failed, which is the evidence auditors and data owners ask for after cutover.
Security. The IAM role DataSync assumes for S3 should be limited to the destination bucket and prefix. SMB credentials can be stored in Secrets Manager. DataSync encrypts data in transit; encryption at rest is the destination's job, through S3 default encryption or KMS.
Failure modes
- Throughput far below the link. Usually the source: a busy NAS, an undersized agent, or millions of tiny files. Check agent CPU and source latency before blaming the network.
- Verification failures on live data. Files changed during the run. Run again or schedule a quiet window; the next incremental picks them up.
- Permission errors mid-run. The agent's NFS or SMB identity cannot read some directories. Fix access on the source rather than excluding silently.
- Basic mode quota exceeded. Split with include filters or move to Enhanced mode.
- Unexpected deletions.
PreserveDeletedFiles=REMOVEagainst a shared prefix. Give each task its own prefix and turn on S3 versioning before the first mirror run. - Archive class surprises. Overwrites and deletes in archive classes trigger minimum-duration charges. Land in Standard or Intelligent-Tiering and let lifecycle rules tier later.
When to use something else
DataSync is for moving data in bulk or on a schedule. For ongoing replication between S3 buckets, S3 replication is simpler and event-driven. For applications that must keep using a file interface backed by AWS, Storage Gateway is the better fit. For many terabytes over a slow link, ship a Snow device. For a few gigabytes, the CLI is enough.
What to do next
- Inventory the source: total bytes, file count, largest directories and change rate per day.
- Pick the task mode from the location pair and the file count, then size the agent to match.
- Compute transfer time from usable bandwidth and decide whether an offline device is needed.
- Create the locations and task in code, with logging at
TRANSFER, a task report and a bandwidth limit. - Run a pilot on one directory and confirm metadata, permissions and throughput.
- Run incrementals until they fit your cutover window, then freeze, run the last one and check for zero failures.
- After cutover, delete the task, deactivate the agent and remove the IAM role.