Sooner or later every large bucket needs the same change applied to every object in it: re-tag forty million images for a new cost-allocation scheme, copy a decade of logs into a bucket with a different encryption key, put a legal hold on everything a court order names, or restore a year of archived scans. Writing that loop yourself means building a key list, a worker fleet, retries, a per-object record and a kill switch. S3 Batch Operations is the managed version of that loop. You hand it a list of objects and one operation, and it runs the operation once per object, retries what can be retried, stops itself if most tasks are failing, and writes a per-object report when it finishes.
This article covers where the object list comes from, the job lifecycle, the Lambda contract, how failures are counted and how to read the report, then works through re-tagging and re-encrypting 60 million objects. Limits and formats come from the Amazon S3 User Guide as read on 2026-10-03; anything not confirmed there is flagged rather than guessed.
The job in one picture
The model: one operation, one list, one task per object
A Batch Operations job has four required inputs. The operation is exactly one action with its parameters, for example replace all object tags with this tag set. The manifest is the list of objects to act on. The role is an IAM role that Batch Operations assumes; its trust policy names the service principal batchoperations.s3.amazonaws.com, and its permissions must cover reading the manifest, performing the operation on every object, and writing the report. The priority is an integer that only means something relative to your other jobs in the same account and Region. Always add a report configuration too.
Each row of the manifest becomes one task. Tasks are independent and run asynchronously, so they do not run in manifest order, and you cannot infer which objects succeeded from how far the job got. There is no transaction: a half-finished job leaves half the objects changed, so every operation must be safe to run twice on the same object.
The built-in operations listed in the User Guide are: Copy, Compute checksums, Delete all object tags, Invoke AWS Lambda function, Replace all object tags, Replace access control list, Restore objects, Update object encryption, Batch Replication of existing objects, Object Lock retention and Object Lock legal hold. Note the word replace in the tagging and ACL operations: an object with five tags that receives a job with one tag ends with one tag. To add a tag while keeping others, merge per object in a Lambda function.
Choosing the manifest
There are four ways to tell a job which objects to touch.
- Generated object list. You name a source bucket and filters, and Batch Operations builds the list itself. The filters in
JobManifestGeneratorFiltercover creation time (CreatedAfter,CreatedBefore), key prefix, suffix and substring, object size bounds, storage class, encryption type and KMS key, and replication status. Save the generated list as a manifest; it is your record of what was in scope. Generating a list requiress3:PutInventoryConfigurationon the source bucket, and it is not supported across Regions. - Replication-configuration list. For Batch Replication, the list can be derived from the source bucket's replication rules.
- S3 Inventory report. You point the job at the
manifest.jsonof a CSV-format inventory report. If the inventory includes version IDs, the job acts on those exact versions. It is the best choice for very large buckets and the only manifest type that may be SSE-KMS encrypted. - Your own CSV. Each row is bucket, key and optionally version ID. Keys must be URL-encoded, and either every row has a version ID or none does. Manually created CSV manifests cannot be SSE-KMS encrypted. You also pass the manifest object's ETag when creating the job, so the job reads exactly the file you checked.
The job parses the whole manifest before running but does not snapshot the bucket. If a row has no version ID and the object is overwritten mid-job, the task acts on the new version. On versioned buckets, include version IDs.
Job lifecycle, confirmation and priority
A job starts as New, moves to Preparing while S3 reads and validates the manifest, and then to Ready. Jobs created in the console stop in Suspended until someone confirms them; jobs created through the CLI, SDK or API can pass --no-confirmation-required and skip that pause. A job left in Suspended for more than 30 days fails. Use confirmation from code too: create the job, read the task count from describe-job, and only then set it to Ready.
From Ready a job becomes Active when S3 starts running it. A running job can be moved to Paused when a higher-priority job arrives, and returns to Active when that job is done; strict ordering is not guaranteed, so if job B must follow job A, wait for A yourself. The terminal states are Complete, Cancelled and Failed, each reached through a transitional state (Completing, Cancelling, Failing). Complete means every task ran, not that every task succeeded; always read the failed-task count.
# Create a re-tagging job that waits for confirmation (no --no-confirmation-required).
aws s3control create-job --account-id 111122223333 --region eu-west-1 \
--operation '{"S3PutObjectTagging":{"TagSet":[{"Key":"cost-centre","Value":"media"},{"Key":"tier","Value":"archive"}]}}' \
--manifest '{"Spec":{"Format":"S3BatchOperations_CSV_20180820","Fields":["Bucket","Key","VersionId"]},
"Location":{"ObjectArn":"arn:aws:s3:::ops-manifests/retag/2026-10-03.csv","ETag":"<etag of that object>"}}' \
--report '{"Bucket":"arn:aws:s3:::ops-reports","Prefix":"retag","Format":"Report_CSV_20180820","Enabled":true,"ReportScope":"AllTasks"}' \
--priority 10 --role-arn arn:aws:iam::111122223333:role/s3-batch-retag \
--client-request-token "$(uuidgen)" --description "retag media bucket 2026-10-03"
# Check the scope, then release it.
aws s3control describe-job --account-id 111122223333 --job-id "$JOB" \
--query 'Job.{status:Status,total:ProgressSummary.TotalNumberOfTasks}'
aws s3control update-job-status --account-id 111122223333 --job-id "$JOB" --requested-job-status Ready
The Lambda contract
When no built-in operation fits (convert a format, merge tags), the job invokes a Lambda function per object. It cannot be an ordinary S3-event handler, because it must return a structured answer. With schema 1.0 the request carries invocationSchemaVersion, invocationId, the job ID, and a tasks list whose entries have taskId, s3Key, s3VersionId and s3BucketArn. The response echoes the invocation ID and returns one result per task: a resultCode and a resultString.
The three result codes carry the whole retry policy. Succeeded ends the task and puts your result string in the report. TemporaryFailure means the task will be redriven before the job completes; if the last attempt still fails, its message lands in the report. PermanentFailure ends the task as failed. Any task ID you forget to return gets the code in treatMissingKeysAs, which defaults to TemporaryFailure. Schema 2.0, required for directory buckets, adds UserArguments; its request shape is not reproduced here, so check the API reference first.
Keys arrive URL-encoded, so decode them. An unqualified ARN or alias means the job calls a new version as soon as you publish one; pin a version number.
import json
from urllib.parse import unquote_plus
import boto3
from botocore.exceptions import ClientError
s3 = boto3.client("s3")
RETRYABLE = {"SlowDown", "RequestTimeout", "InternalError", "ServiceUnavailable"}
def handler(event, context):
results = []
for task in event["tasks"]:
key = unquote_plus(task["s3Key"], encoding="utf-8")
bucket = task["s3BucketArn"].split(":")[-1]
version = task.get("s3VersionId")
try:
args = {"Bucket": bucket, "Key": key}
if version:
args["VersionId"] = version
tags = {t["Key"]: t["Value"] for t in s3.get_object_tagging(**args)["TagSet"]}
tags["cost-centre"] = "media"
s3.put_object_tagging(**args, Tagging={"TagSet": [{"Key": k, "Value": v} for k, v in tags.items()]})
code, msg = "Succeeded", json.dumps({"tags": len(tags)})
except ClientError as e:
err = e.response["Error"]["Code"]
code = "TemporaryFailure" if err in RETRYABLE else "PermanentFailure"
msg = f"{err}: {key}"
results.append({"taskId": task["taskId"], "resultCode": code, "resultString": msg})
return {
"invocationSchemaVersion": event["invocationSchemaVersion"],
"treatMissingKeysAs": "PermanentFailure",
"invocationId": event["invocationId"],
"results": results,
}Setting treatMissingKeysAs to PermanentFailure makes a dropped result visible in the report instead of retried quietly. The key column can also carry a URL-encoded JSON string of up to 1,024 characters, which your function decodes, so one manifest row can say copy this key to that new key.
The failure threshold and the completion report
Batch Operations protects you from a job that is mostly failing. Once a job has run at least 1,000 tasks, S3 watches the ratio of failed tasks to tasks run, and if it ever exceeds 50 percent the job fails. That catches systematic mistakes early, but not a job where 20 percent of objects fail; that job completes, and only the report tells you.
The completion report is written when the job completes, fails or is cancelled, provided at least one task was invoked. You choose all tasks or failed tasks only. A top-level manifest.json points at separate CSV files for succeeded and failed tasks, everything is encrypted with SSE-S3, and each row has bucket, key, version ID, task status, error code, HTTP status code and result message. Treat the failed rows as the next job's manifest: after fixing the cause, build a new CSV from them and run the same operation again. That loop (run, read failures, fix, re-run on failures only) is how large jobs finish cleanly. CloudTrail also records each task's S3 calls, and data-event volume scales with the object count.
Worked example: re-tag and re-encrypt 60 million objects
A media company keeps 60 million objects in media-raw, a versioned bucket in eu-west-1. Finance wants a cost-centre=media tag on every object, and security wants objects encrypted under a new customer-managed KMS key rather than SSE-S3. Some objects are over 5 GB.
- Scope. The bucket already has a daily CSV inventory with version IDs and size. Use yesterday's
manifest.jsonfor the tagging job, because it is large, versioned and allowed to be SSE-KMS. Athena on the inventory shows 59.6 million objects below 5 GB and 400,000 above. - Tags. Some objects carry other tags that must survive, so Replace all object tags is unsafe. S3 Inventory has no tag field, so you cannot see the existing tag sets in Athena; finding them means calling
GetObjectTaggingper object, which is already most of the work of a merge. Run the merging Lambda function above as an Invoke AWS Lambda function job instead. - Encryption. Either use the Update object encryption operation (read its current page for parameters and limits; they are not reproduced here) or the older pattern of copying objects onto themselves. Copy handles up to 5 GB per object, one source and one destination bucket, created in the destination Region; the 400,000 larger objects need a multipart copy via
UploadPartCopy. - Role. Grant
s3:GetObjectands3:GetObjectVersionon the manifest bucket,s3:PutObjecton the report bucket, the operation's own actions onmedia-raw/*,kms:GenerateDataKeyon the new key for re-encryption, andkms:Decrypton the inventory's key because that manifest is SSE-KMS encrypted. - Canary. Run each job on a 10,000-row slice first, check the report, and spot-check objects with
head-object. - Run and close out. Launch the full jobs, watch progress with
describe-job, and when each completes, turn the failed rows into a retry manifest. The work is finished when the retry report is empty, an inventory taken afterwards shows no object with the old encryption status or KMS state, and a sampled tag check (inventory cannot show tags) finds the new tag everywhere it looked.
Failure modes
| Symptom | Likely cause | What to do |
|---|---|---|
| Job fails in Preparing | Role cannot read the manifest, wrong ETag, or SSE-KMS on a hand-made CSV | Read the failure reason on the job; fix permissions or upload the CSV with SSE-S3 |
| Job fails after about 1,000 tasks | Systematic error: missing KMS permission, keys not URL-encoded, wrong bucket | Read the first rows of the report; fix and resubmit |
| Complete, but many failed tasks | Mixed population: some objects archived, too large, or deleted since the list was made | Split the failures by error code and handle each group separately |
| Tags vanished from objects | Replace all object tags used where a merge was needed | Restore from your own tag records or a backup copy; next time merge in a Lambda job |
| Wrong version changed | No version IDs in the manifest and objects were overwritten mid-job | Use version IDs; for copies, copy noncurrent versions first, then current |
| Lambda tasks retry for hours | Function returns TemporaryFailure for permanent errors, or omits task IDs | Classify errors; set treatMissingKeysAs to PermanentFailure |
Trade-offs and where it fits
Batch Operations removes the fleet, the retries, the progress tracking and the audit report, which is most of the engineering in a bulk job. In exchange you accept one operation per job, no ordering, no snapshot, a per-job and per-object charge (see the S3 pricing page; no figures are quoted here), and Lambda costs when you invoke a function. For a few thousand objects a short SDK script is simpler; for millions, or anything needing a per-object audit record, the managed job wins.
Lifecycle rules suit recurring age-based changes, replication handles new writes, and Step Functions can chain jobs that must run in order. Use Batch Operations for one-time sweeps across objects that already exist.
Related reading on this site: how S3 works underneath, S3 replication and Batch Replication, S3 encryption options, S3 Object Lock and AWS Lambda in depth.
What to do next
- Turn on a daily CSV S3 Inventory with version IDs, size, storage class and encryption status for every bucket you might sweep later; it has no tag field, so record tags you depend on elsewhere.
- Create one IAM role per job type with the batchoperations trust principal, scoped to named manifest, report and target buckets.
- Request a completion report on every job.
- Create jobs with confirmation required, check TotalNumberOfTasks, then release them.
- Run a 10,000-object canary of every new job definition and inspect the report before the full run.
- Pin Lambda functions by version number and return a result code for every task ID.
- Write the close-out check that proves every object reached the target state.