Moving a few gigabytes into Cloud Storage is a one-line command. Moving 400 TB out of Amazon S3 while applications keep writing, or keeping a nightly copy of an on-premises NFS share in a bucket, is a different job. Something has to list millions of objects, compare them with what is already in the destination, copy the difference in parallel, retry the failures, survive restarts, and tell you exactly what happened. Storage Transfer Service (STS) is Google Cloud's managed service for that loop.

This article explains the model from first principles: jobs and operations, how a run decides what to copy, the overwrite and delete options that turn a copy into a sync, credentials, agents for file systems, event-driven transfers, monitoring and verification. A worked migration ties it together. For bucket fundamentals see Cloud Storage; this page assumes you know what a bucket, object and storage class are.

Flag names and option semantics were checked on 2026-10-04 against the gcloud reference for gcloud transfer jobs create and the v1 REST reference. Pricing changes, so this page describes it only qualitatively; check the current pricing page before you plan a budget.

Jobs, operations and the copy loop

Storage Transfer Service: a managed copy loop between a source and a Cloud Storage sinkAmazon S3 / Azureagentless, cloud APICloud Storagebucket to bucketURL list (TSV)public HTTP objectsPOSIX / HDFSvia agents you runTransfer jobspec, schedule, optionsone operation per runAgent poolcontainers near the dataCloud Storage sinkservice agent writestaskslist, compare, copyEvent streamPub/Sub, SQS, Storage QueueCloud Logging + opscounters, per-object logsThe job is the durable definition; each run lists, compares and copies, then reports counters.
Sources on the left, the job in the middle, the Cloud Storage sink on the right. File-system sources need agents; event streams let a job copy changes without relisting.

STS has two nouns. A transfer job is the durable definition: a source, a sink, filters, options, an optional schedule and a logging configuration. A transfer operation is one execution of that job. A daily job produces one operation a day, each with its own status and counters, such as objects found, bytes copied and objects failed.

Each operation runs the same loop. It lists the source and the sink, applies the object conditions (prefix and modification-time filters), compares each source object with the sink object of the same name, copies what needs copying, applies any delete option, and records the result. Because the comparison happens on every run, a rerun is naturally incremental: objects already present and unchanged are skipped. That property is what makes STS safe to run repeatedly during a migration.

Agentless sources are those STS can reach through a cloud API: another Cloud Storage bucket, Amazon S3, Azure Blob Storage and lists of public HTTP URLs. Google runs the copy workers. For POSIX file systems, HDFS and S3-compatible storage, Google cannot reach the data directly, so you run transfer agents, containers grouped into an agent pool, on machines that can read the data. The agents take work from the service and upload directly to Cloud Storage.

From copy to sync: overwrite, delete and filters

By default a run copies source objects that are missing in the sink or differ from it, and leaves everything else alone. Three groups of options change that, and they are where most surprises come from.

Optiongcloud flagEffect
Overwrite policy--overwrite-when=different|always|neverDIFFERENT overwrites a same-named sink object only when ETags or checksums differ. NEVER keeps whatever is in the sink. ALWAYS rewrites even identical objects.
Delete unique in sink--delete-from=destination-if-uniqueDeletes sink objects that no longer exist at the source. This turns a copy into a mirror.
Delete from source--delete-from=source-after-transferDeletes each source object after it is copied. This turns a copy into a move.
Prefix filters--include-prefixes / --exclude-prefixesRestrict both the listing and the comparison to named prefixes.
Time filters--include-modified-after-relativeOnly consider objects modified within a window, for cheap incremental runs.

The two delete options are mutually exclusive, which makes sense: a move empties the source, and a mirror would then delete everything in the sink. Less obvious is that deleting sink-unique objects cannot be combined with time-based object conditions; the API rejects the job with INVALID_ARGUMENT. The reason is that time conditions do not exclude sink objects, so a job that only looked at recent source objects would see every older sink object as unique and delete it.

Choose ALWAYS deliberately: it rewrites every object on every run and multiplies operation and egress charges. DIFFERENT is the right default for migrations; NEVER suits append-only archives.

Metadata is a separate decision. The --preserve-metadata flag accepts values such as storage-class, kms-key, temporary-hold, time-created, acl, and for file systems mode, uid, gid and symlink. Anything you do not preserve gets the sink bucket's defaults, so a migration that forgets kms-key lands objects under Google-managed encryption, even if the source used customer keys.

Identity and credentials

STS acts through a Google-managed service agent created per project. It needs permission to read the source (if that is a bucket) and to write, and possibly delete, in the sink. Grant roles on the buckets involved, not the project, so a mistyped job cannot touch unrelated data. Read the service agent's address from the API or console rather than constructing it by hand.

For Amazon S3 you have two choices. You can supply an access key pair, for example in a JSON file passed with --source-creds-file, or you can configure a role ARN so STS assumes an AWS role through federated identity. Prefer the role. Static keys live in the job definition's credential store, must be rotated by hand, and outlive the migration if nobody cleans them up. Either way, scope the AWS policy to s3:ListBucket and s3:GetObject on the one bucket, plus delete only if you use a move.

Agents for file systems

An agent is a container that reads files from a mounted path and uploads them. Throughput scales with the number of agents and the network between them and Google, so the design question is where to put them and how many to run. Put them as close to the storage as possible, on machines with fast access to the NFS or SMB mount, and spread them over several hosts so one network card is not the bottleneck.

# create a pool with a bandwidth cap so the transfer cannot saturate the office link
gcloud transfer agent-pools create nas-pool --bandwidth-limit=800   # MB/s for the whole pool

# on each host with the share mounted at /mnt/nas
gcloud transfer agents install --pool=nas-pool --count=3 \
    --mount-directories=/mnt/nas

# nightly incremental copy of one export into a bucket
gcloud transfer jobs create posix:///mnt/nas/exports/ gs://acme-nas-mirror/exports/ \
    --source-agent-pool=nas-pool \
    --schedule-repeats-every=24h \
    --overwrite-when=different \
    --preserve-metadata=mode,uid,gid \
    --log-actions=copy,delete --log-action-states=failed,succeeded

The pool-level bandwidth limit, in MB/s across all agents in the pool, is the tool for coexisting with production traffic, rather than throttling individual hosts. Watch agent CPU too: a share of fifty million 4 KB files is limited by per-file operations, not bandwidth, and benefits from more agents rather than a bigger pipe.

Event-driven transfers

A scheduled job relists the whole source every run. For a bucket with hundreds of millions of objects, listing alone takes hours and costs list operations, which is why near-real-time replication is better done with an event-driven transfer. You point the job at a notification stream with --event-stream-name: a Pub/Sub subscription for a Cloud Storage source, an SQS queue fed by S3 event notifications for Amazon S3, or an Azure Storage queue for Azure. STS then copies objects as their events arrive, without listing.

Two properties matter. First, deletions are not propagated: an event-driven job is one-way copying, not synchronisation. If you need deletes mirrored, run a periodic scheduled job with --delete-from=destination-if-unique as a reconciliation pass. Second, no latency or ordering guarantee is documented, so do not build consumers that assume an object appears in the sink within a fixed time or in source order.

Driving STS from code

Everything gcloud does is available through the API, which is how you embed STS in a pipeline. The Python client below creates a bucket-to-bucket job and starts it once. Field names follow the v1 REST resource in snake case.

from google.cloud import storage_transfer

def create_and_run(project_id: str, src: str, dst: str, prefix: str) -> str:
    client = storage_transfer.StorageTransferServiceClient()
    job = client.create_transfer_job({
        "transfer_job": {
            "project_id": project_id,
            "description": f"copy {src}/{prefix} to {dst}",
            "status": storage_transfer.TransferJob.Status.ENABLED,
            "transfer_spec": {
                "gcs_data_source": {"bucket_name": src},
                "gcs_data_sink": {"bucket_name": dst},
                "object_conditions": {"include_prefixes": [prefix]},
                "transfer_options": {
                    "overwrite_when": storage_transfer.TransferOptions.OverwriteWhen.DIFFERENT,
                },
            },
        }
    })
    # no schedule: nothing runs until we ask
    op = client.run_transfer_job({"job_name": job.name, "project_id": project_id})
    op.result(timeout=6 * 3600)          # long-running operation; raises on failure
    return job.name

Treat the job name as configuration: create once, then call run_transfer_job or attach a schedule, rather than creating a job per run. When the copy set is an explicit list rather than a prefix, pass a manifest, a CSV of object names given with --manifest-file, so a run copies exactly those objects.

Worked example: 400 TB from S3

A team is moving 400 TB and 900 million objects from an S3 bucket to Cloud Storage. Applications write about 2 TB a day, and the business accepts a 30-minute write freeze at cutover. Here is the plan, step by step.

  1. Prepare. Create the sink bucket in the target location with the right default storage class and encryption. Configure an AWS role for STS, scoped to read the one bucket. Turn on per-object logging for failed actions only, since logging every copied object of a 900-million-object bucket is costly.
  2. Bulk copy. Create a job with --overwrite-when=different and no schedule, and run it. With this many objects, split the work by top-level prefix into several jobs so one stuck prefix does not hold up the rest, and so failures can be retried by prefix.
  3. Catch up. When the bulk finishes, days later, create an event-driven job on an SQS queue fed by S3 notifications. New writes now arrive within minutes. Run the scheduled jobs again once to copy anything written between the bulk listing and the event stream starting.
  4. Verify. Compare object counts and total bytes per prefix from an independent listing on each side, not from STS's own counters. Spot-check checksums on a random sample. Investigate every failed object in the logs, which usually turns up permission errors, objects under legal hold, or names the sink rejects.
  5. Cut over. Freeze writes, wait for the event job to drain, run a final scheduled pass, re-verify the counts, switch application configuration to the new bucket, and unfreeze.
  6. Clean up. Disable the event job, delete the AWS role or keys, and keep the S3 bucket read-only for an agreed rollback window before deleting it.

Monitoring and verification

Each operation exposes counters, among them objects and bytes found at the source, copied to the sink, skipped because they already existed, and failed. From the CLI, list a job's operations with gcloud transfer operations list and inspect one with gcloud transfer operations describe. Alert on any non-zero failed count and on runs whose duration grows week over week, which usually means the source is growing faster than the schedule.

Per-object logs go to Cloud Logging when the job has a logging configuration: --log-actions chooses find, copy or delete and --log-action-states chooses succeeded, failed or skipped. For routine replication, log only failures, or the logs cost more than the transfer.

STS validates data integrity by checksum during copying, but verification of the migration as a whole is still your job, because a filter or prefix mistake copies the wrong set perfectly.

Failure modes

  • Mirror deletes the wrong data. A job with --delete-from=destination-if-unique pointed at the wrong prefix, or at a bucket other systems also write to, deletes their objects. Use a dedicated sink prefix, enable object versioning or soft delete on the sink, and dry-run new jobs with delete logging first.
  • Event-only replication drifts. Deletes never propagate and a lost notification is never retried by a listing. Pair every event-driven job with a periodic reconciliation run.
  • Encryption silently changes. Objects land under the sink's default encryption unless kms-key is preserved or the sink has the right default key.
  • Small files crawl. Millions of tiny files are bounded by per-object overhead. Add agents, archive cold small files into larger objects before transfer, or accept the duration.
  • Egress surprises. STS's own fee is separate from the source provider's egress and operation charges, which usually dominate a cross-cloud migration. Estimate them from the source bill, not from STS pricing.
  • Credentials outlive the project. Access keys stored in a finished job remain valid until someone deletes them. Make credential removal a checklist step.

Choosing between transfer tools

ToolBest forWeakness
Storage Transfer ServiceLarge or recurring transfers, cross-cloud, file systems through agentsSetup overhead; asynchronous; per-run listing on schedules
gcloud storage cp / rsyncAd hoc copies, gigabytes to a few terabytes, scriptsRuns on your machine; you handle restarts and parallelism
Dual-region and turbo replicationKeeping one bucket's data in two regionsSame bucket only; not a migration tool
Transfer AppliancePetabytes over a link too slow to finish in timePhysical shipping, weeks of lead time

Use STS when the transfer is big, recurring or crosses providers. For geo-redundancy of one bucket, prefer the built-in options in dual- and multi-region replication. After data lands, set lifecycle rules and pick storage classes so a migrated archive does not sit in Standard forever.

What to do next

  1. Inventory the source: object count, total bytes, size distribution and daily write rate. These decide agent count, job splitting and whether you need event-driven catch-up.
  2. Create the sink bucket with location, storage class, encryption key, versioning or soft delete chosen up front.
  3. Set up least-privilege access: bucket-level roles for the service agent, and an AWS role rather than static keys for S3.
  4. Create one job per large prefix with --overwrite-when=different and failure-only logging, and run it without deletes first.
  5. Add an event-driven job for live sources, and schedule a reconciliation pass because deletes do not propagate.
  6. Verify from independent listings and sampled checksums before cutover, then remove credentials and disable jobs afterwards.
Key takeaway: Storage Transfer Service runs a list, compare and copy loop as a managed job, so reruns are incremental. Overwrite and delete options turn copying into mirroring or moving and deserve the most care. Use agents for file systems, event streams for near-real-time copies (which never propagate deletes), and always verify a migration from independent listings.