S3 replication copies new object versions from one bucket to another, asynchronously, without you running any copy jobs. Pointed at a bucket in another Region it gives you a geographically separate copy for disaster recovery or lower-latency reads; pointed at a bucket in another AWS account it gives you a copy that a compromised production account cannot reach. Combining both, cross-Region and cross-account, is the standard shape for a backup that survives both a regional outage and a credential leak.
It is also easy to configure something that looks right and protects much less than you think. Deletes may or may not follow, existing objects are not copied, encrypted objects can be silently skipped, and a failed object is never retried on its own. This article builds a cross-account, cross-Region setup from first principles: the rule, the role, the policies, KMS, delete semantics, backfill and monitoring, with a worked example and a checklist. It assumes S3 basics from Amazon S3 architecture. Behaviour described here follows the S3 User Guide as read on 2026-10-01.
What replication is, and what it is not
Replication works on object versions. When a new version lands in the source bucket and matches a rule, S3 queues it and writes the same version, with the same key, metadata and version ID, into each destination bucket. That is why both buckets must have versioning enabled: without versions there is nothing stable to copy or to refer to.
Cross-Region Replication (CRR) and Same-Region Replication (SRR) are the same mechanism with a different destination Region.
Three properties shape everything else. Replication is asynchronous, so the destination is always somewhat behind. It is live only: pre-existing objects are not copied unless you run Batch Replication. And it copies objects, not buckets: lifecycle, policies and notifications stay as they are on each side, which is useful, because a backup bucket should be stricter than production.
The moving parts
Four things must line up: a replication configuration on the source naming an IAM role; the role, which S3 assumes to read and write; in a cross-account setup, a destination bucket policy allowing that role; and for KMS-encrypted objects, decrypt rights on the source key and encrypt rights on a key in the destination Region. Permissions are covered in depth in AWS IAM.
Anatomy of a rule
A replication configuration has a role and a list of rules. Each rule has a filter (prefix, tags or both), a priority used when rules overlap, a status, and a destination with options. This configuration replicates everything under ledger/ to a backup account, including KMS-encrypted objects, with delete markers replicated and RTC enabled.
{
"Role": "arn:aws:iam::111111111111:role/s3-repl-ledger",
"Rules": [
{
"ID": "ledger-to-backup",
"Priority": 1,
"Status": "Enabled",
"Filter": { "Prefix": "ledger/" },
"DeleteMarkerReplication": { "Status": "Enabled" },
"SourceSelectionCriteria": {
"SseKmsEncryptedObjects": { "Status": "Enabled" }
},
"Destination": {
"Bucket": "arn:aws:s3:::acme-ledger-backup-usw2",
"Account": "222222222222",
"StorageClass": "STANDARD_IA",
"EncryptionConfiguration": {
"ReplicaKmsKeyID": "arn:aws:kms:us-west-2:222222222222:key/REPLACE-ME"
},
"Metrics": { "Status": "Enabled", "EventThreshold": { "Minutes": 15 } },
"ReplicationTime": { "Status": "Enabled", "Time": { "Minutes": 15 } }
}
}
]
}Rules that use Filter are the current schema. Older configurations without Filter are treated as the first version of the schema, and the two versions handle delete markers differently, which the deletes section covers. To send the same objects to several destinations, write one rule per destination with distinct priorities. The 15-minute values are the only values S3 accepts for those RTC fields.
The replication role
The role's trust policy names s3.amazonaws.com as the principal, adding batchoperations.s3.amazonaws.com if the same role will run Batch Replication jobs. Its permissions policy reads the source and writes to the destination.
{
"Version": "2012-10-17",
"Statement": [
{ "Effect": "Allow",
"Action": ["s3:GetReplicationConfiguration", "s3:ListBucket"],
"Resource": "arn:aws:s3:::acme-ledger-use1" },
{ "Effect": "Allow",
"Action": ["s3:GetObjectVersionForReplication", "s3:GetObjectVersionAcl",
"s3:GetObjectVersionTagging"],
"Resource": "arn:aws:s3:::acme-ledger-use1/ledger/*" },
{ "Effect": "Allow",
"Action": ["s3:ReplicateObject", "s3:ReplicateDelete", "s3:ReplicateTags"],
"Resource": "arn:aws:s3:::acme-ledger-backup-usw2/ledger/*" },
{ "Effect": "Allow", "Action": "kms:Decrypt",
"Resource": "arn:aws:kms:us-east-1:111111111111:key/SOURCE-KEY",
"Condition": { "StringLike": {
"kms:ViaService": "s3.us-east-1.amazonaws.com",
"kms:EncryptionContext:aws:s3:arn": "arn:aws:s3:::acme-ledger-use1/ledger/*" } } },
{ "Effect": "Allow", "Action": "kms:Encrypt",
"Resource": "arn:aws:kms:us-west-2:222222222222:key/REPLACE-ME",
"Condition": { "StringLike": {
"kms:ViaService": "s3.us-west-2.amazonaws.com",
"kms:EncryptionContext:aws:s3:arn": "arn:aws:s3:::acme-ledger-backup-usw2/ledger/*" } } }
]
}Use s3:GetObjectVersionForReplication rather than s3:GetObjectVersion: the documentation notes that the latter does not allow replication of KMS-encrypted objects. If the source objects use Object Lock, add s3:GetObjectRetention and s3:GetObjectLegalHold so retention settings travel with the replicas.
Cross-account: the destination's half
Account B must agree to receive. Its bucket policy grants the source role the replicate actions on objects and versioning reads on the bucket.
{
"Version": "2012-10-17",
"Statement": [
{ "Sid": "ReplicaWrites", "Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::111111111111:role/s3-repl-ledger" },
"Action": ["s3:ReplicateObject", "s3:ReplicateDelete"],
"Resource": "arn:aws:s3:::acme-ledger-backup-usw2/*" },
{ "Sid": "BucketChecks", "Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::111111111111:role/s3-repl-ledger" },
"Action": ["s3:List*", "s3:GetBucketVersioning", "s3:PutBucketVersioning"],
"Resource": "arn:aws:s3:::acme-ledger-backup-usw2" }
]
}Ownership decides who controls the replicas. By default the source object's owner also owns the replica, which for a backup account is the wrong answer: Account A could then still govern objects in Account B. Two mechanisms change it. The simpler is to set Object Ownership on the destination bucket to Bucket owner enforced, which disables ACLs and makes Account B own every replica without extra permissions. The older one is the owner override, an AccessControlTranslation element with Owner set to Destination plus Account in the rule, which also needs s3:ObjectOwnerOverrideToBucketOwner in both the role and the destination policy. Prefer Bucket owner enforced for new buckets.
For KMS in a cross-account setup, the destination key must be a customer managed key whose key policy lets the source account use it. AWS managed keys cannot be used across accounts, so they cannot encrypt cross-account replicas.
Encryption: which objects are replicated
Unencrypted, SSE-S3 and SSE-C objects replicate by default. Objects encrypted with SSE-KMS or DSSE-KMS are not replicated unless the rule opts in with SseKmsEncryptedObjects and names a ReplicaKmsKeyID in the destination Region. This default is the most common reason a replication setup appears healthy while part of the data never arrives, because the skipped objects were never eligible and so never show as failed. Encryption choices are compared in S3 encryption options and key policies in AWS KMS.
Three KMS details matter in operation. With S3 Bucket Keys enabled, the encryption context is the bucket ARN, not the object ARN, so the condition values in the role policy above must change to the bucket ARN. PutBucketReplication does not validate the KMS key: an invalid key returns 200 OK and replication then fails object by object. And replication spends KMS request quota; the RTC guidance gives the example that replicating 1,000 objects per second consumes 2,000 KMS requests per second, and throttling shows up as replication delay.
Deletes, and what is never replicated
Delete behaviour is the part to get right on purpose, because it decides whether the replica is a mirror or a backup.
| Event on the source | What happens at the destination |
|---|---|
| DELETE without a version ID (creates a delete marker) | Filter-based rules: not replicated unless DeleteMarkerReplication is enabled, which is not allowed on tag-based rules. Older rules without Filter replicate user-created markers. |
| DELETE of a specific version ID | Never replicated. The version survives at the destination, which protects against malicious deletion. |
| Lifecycle expiration or transition | Never replicated; configure lifecycle on each bucket separately |
| Object in Glacier Flexible Retrieval, Deep Archive or the Intelligent-Tiering archive tiers | Not replicated until restored and copied to another class |
| Object that is itself a replica from another rule | Not replicated onward: no automatic chaining from A to B to C |
| Tag added after PutObject, for a tag-filtered rule | Not replicated by live replication; use Batch Replication |
| Bucket policy, lifecycle or notification change | Not replicated; each bucket keeps its own settings |
For a disaster-recovery mirror you probably want delete markers replicated, so a deleted object also looks deleted in the failover Region. For a backup you probably do not, so a mass delete in production leaves the backup's current view intact. Either way, version deletes do not propagate. Pair a backup destination with Object Lock and the backup account's own restrictive policies.
Existing objects and failures: Batch Replication
Live replication never looks backwards. To copy objects that existed before the rule, objects that replicated to an old destination, replicas you want to chain onward, or objects that failed, you create an S3 Batch Replication job. It can filter by replication status: NONE for never attempted, FAILED, COMPLETED and REPLICA.
Failure handling is the operational surprise. An object goes to FAILED for problems such as missing role, KMS or bucket permissions, and S3 does not retry it: you fix the cause and re-run it through Batch Replication or upload it again. Temporary problems, such as an unavailable destination, leave the object PENDING and S3 resumes on its own. Lifecycle on the source will not expire or transition objects that are PENDING or FAILED, so a permissions mistake can quietly hold storage you expected lifecycle to clean up.
Watching it: status, metrics and RTC
Each eligible source object carries x-amz-replication-status, returned by HeadObject and GetObject: PENDING, COMPLETED or FAILED. Replicas report REPLICA. With several destinations the source shows COMPLETED only when every destination succeeded. For a whole bucket, S3 Inventory reports replication status per object and can be queried with Athena.
Replication metrics, enabled per rule, publish Replication Latency, Bytes Pending Replication and Operations Pending Replication in the destination Region, and Operations Failed Replication in the source Region. To learn which objects failed and why, subscribe to the OperationFailedReplication event. S3 Replication Time Control adds a service level target: most objects replicate in seconds and 99.9 percent within 15 minutes, with OperationMissedThreshold and OperationReplicatedAfterThreshold events. The SLA excludes periods when request rates exceed S3's per-prefix guidelines or replication transfer exceeds the default 1 Gbps quota, which can be raised. Count the extra load replication creates: for each object, up to five GET or HEAD requests and one PUT on the source and one PUT on each destination.
# Smoke test: write one object, then wait until the replica exists in the backup account.
import time, boto3
src = boto3.Session(profile_name="prod").client("s3", region_name="us-east-1")
dst = boto3.Session(profile_name="backup").client("s3", region_name="us-west-2")
put = src.put_object(Bucket="acme-ledger-use1", Key="ledger/canary.txt", Body=b"canary",
ServerSideEncryption="aws:kms", SSEKMSKeyId="alias/ledger-use1")
vid = put["VersionId"]
for _ in range(60):
status = src.head_object(Bucket="acme-ledger-use1", Key="ledger/canary.txt",
VersionId=vid).get("ReplicationStatus")
if status in ("COMPLETED", "FAILED"):
break
time.sleep(10)
print("source status:", status)
rep = dst.head_object(Bucket="acme-ledger-backup-usw2", Key="ledger/canary.txt", VersionId=vid)
print("replica:", rep.get("ReplicationStatus"), rep.get("SSEKMSKeyId"))
Worked example: a ledger backup in another account
Account A (111111111111) holds acme-ledger-use1 in us-east-1, written by payment services with SSE-KMS. The security team owns Account B (222222222222), which no production principal can log in to, with acme-ledger-backup-usw2 in us-west-2.
- In Account B: create the bucket with versioning, Object Lock and Bucket owner enforced; create a customer managed key in us-west-2 whose key policy allows Account A; apply the bucket policy above.
- In Account A: enable versioning on the source, create the role with the trust and permissions policies above, then apply the configuration with
aws s3api put-bucket-replication --bucket acme-ledger-use1 --replication-configuration file://repl.json. - Decide deletes: this is a backup, so set
DeleteMarkerReplicationtoDisabledinstead of the mirror-style value shown earlier. - Backfill: run a Batch Replication job for objects with status
NONE, then confirm the job report shows no failures. - Verify with the smoke test, alarm on Operations Failed Replication above zero and on Replication Latency approaching 15 minutes, and route
OperationFailedReplicationevents to the on-call queue.
Later, an edit to the source key policy drops the replication role. New objects go to FAILED, the alarm fires, the grant is restored, and a Batch Replication job filtered on FAILED closes the gap.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Some objects never appear and never fail | SSE-KMS objects without the opt-in, archived storage classes, tags added after upload | Opt in with ReplicaKmsKeyID; use Batch Replication for the rest |
| Everything FAILED right after setup | Destination policy, key policy or role missing an action | Check the role, the bucket policy and both key policies; then Batch Replication on FAILED |
| PutBucketReplication succeeded, nothing replicates | Invalid ReplicaKmsKeyID, which is not validated | Fix the key and backfill |
| Latency climbs during bulk loads | KMS throttling, per-prefix request limits, transfer quota | Raise KMS and transfer quotas; spread keys across prefixes |
| Backup lost data after a mass delete | Delete markers replicated into a backup | Disable delete-marker replication for backups; add Object Lock |
| Source storage not shrinking | Lifecycle blocked by PENDING or FAILED objects | Clear the failures; lifecycle resumes |
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Delete markers | Replicate: the replica mirrors the source view, good for failover | Do not replicate: the backup survives mass deletes |
| Replica ownership | Bucket owner enforced: simple, no ACLs | Owner override: needed only if the destination must keep using ACLs |
| RTC | Enabled: SLA, threshold events, paid per GB | Metrics only: visibility at lower cost, no SLA |
| Direction | One-way to a locked backup account | Two-way with replica modification sync for active-active, more conflict risk |
What to do next
- Inventory every bucket that matters and record its replication rules, destination account, Region and delete-marker setting.
- Check for KMS-encrypted objects that no rule opts in, and for archive-tier objects that live replication will skip.
- Move backup destinations to a separate account with Bucket owner enforced, Object Lock and a customer managed key.
- Run Batch Replication for pre-existing and FAILED objects, and keep the job reports.
- Enable replication metrics on every rule and alarm on failed operations and latency; add RTC where a recovery point objective needs a contract.
- Schedule the canary script so replication is proven continuously, not assumed.