Cloudflare R2 is an object store with an S3-compatible API and one unusual pricing decision: it does not charge for egress. For most object stores the bill has three parts: storage, requests and bytes leaving the provider. R2 removes the third. For workloads where data is written once and read many times from many places, such as model weights, dataset shards, media and software downloads, that changes both the cost and the architecture.
This article assumes you know how object storage works in general. If not, start with object storage architecture. Here we cover what is specific to R2: how it bills, the three ways into a bucket and the consistency each one gives, code for the S3 API and Workers, events and lifecycle, data location, migration, and the mistakes that cost money or correctness. Every price and limit below is from Cloudflare's documentation as read on 2026-10-01. Check the current pages before you budget.
The pricing model, and what zero egress changes
| Item | Standard | Infrequent Access |
|---|---|---|
| Storage | $0.015 per GB-month | $0.01 per GB-month |
| Class A operations (writes, lists) | $4.50 per million | $9.00 per million |
| Class B operations (reads, heads) | $0.36 per million | $0.90 per million |
| Data retrieval | None | $0.01 per GB |
| Egress | Free | Free |
| Minimum storage duration | None | 30 days |
Class A covers PutObject, CopyObject, ListObjects, the multipart operations and bucket configuration writes. Class B covers GetObject, HeadObject and configuration reads. DeleteObject, DeleteBucket and AbortMultipartUpload are free, and so are requests rejected as unauthorized. A monthly free tier covers 10 GB-month of Standard storage, 1 million Class A and 10 million Class B operations. Usage is rounded up to the next billing unit.
The design consequence is that the cost of a read no longer depends on where the reader is. A training cluster in another cloud, a fleet of inference nodes in several regions and an end user's browser all pay the same per-request price and nothing per byte. Two habits change as a result. You stop building per-region replicas just to avoid transfer fees. And the request count, not byte volume, becomes the number to watch, because many small reads cost more than a few large ones. For the wider picture of transfer pricing, see cloud egress cost.
Three ways into a bucket
The S3 API is for existing tools and SDKs. Point them at the account endpoint, set the region to auto, and authenticate with credentials from an R2 API token. R2 implements a large subset of the S3 API rather than all of it, so check the compatibility table for any operation your tooling depends on.
A Worker binding gives code running on Cloudflare's network a bucket object with methods such as head, get, put, delete, list and createMultipartUpload. There are no keys to manage. This is the natural place for authorization, request rewriting and small transformations; see edge compute for the execution model.
Public access comes in two forms. A custom domain on your Cloudflare zone serves objects through the cache. The r2.dev subdomain is rate limited and meant for development, not production traffic.
Consistency, caching and concurrent writes
R2 is strongly consistent. Once a write or delete completes, every reader through the S3 API or a binding sees it, and listings reflect it immediately. Permission changes are the exception: they are eventually consistent and can take up to a minute.
The cache in front of a public custom domain is a separate system with its own rules. A cached object keeps being served after you overwrite or delete it, until its TTL expires or you purge it. Cloudflare's cache also caches 404 responses by default, so if a client asked for a path before you uploaded it, that path can keep returning 404 after the upload until the TTL expires. The fixes are simple. Write new content under new keys (versioned paths or content hashes in the key), purge on overwrite, and set short TTLs on paths that are probed before they exist.
When two clients write the same key, the last writer to complete wins. R2 also limits concurrent writes to the same key to one per second. A single object used as a counter, a lock or a frequently rewritten index will hit that limit. For coordination, use conditional writes. A put with an ETag precondition succeeds only if the object is unchanged since you read it, which turns a manifest into a safe compare-and-swap.
Code: the S3 API from Python
import boto3
from boto3.s3.transfer import TransferConfig
s3 = boto3.client(
"s3",
endpoint_url="https://<ACCOUNT_ID>.r2.cloudflarestorage.com", # .eu. inserted for EU buckets
aws_access_key_id=R2_ACCESS_KEY_ID, # an R2 API token's S3 credentials
aws_secret_access_key=R2_SECRET_ACCESS_KEY,
region_name="auto", # required by the SDK, ignored by R2
)
# Large artifact: 512 MiB parts x 10,000-part limit covers objects up to about 4.88 TiB; go larger beyond.
cfg = TransferConfig(multipart_threshold=256 * 2**20, multipart_chunksize=512 * 2**20, max_concurrency=16)
s3.upload_file("shard-00017.safetensors", "models", "llm-70b/v3/shard-00017.safetensors", Config=cfg)
# Short-lived upload URL for a client that must never hold credentials (GET/HEAD/PUT/DELETE only).
put_url = s3.generate_presigned_url(
"put_object",
Params={"Bucket": "uploads", "Key": "user-123/avatar.png", "ContentType": "image/png"},
ExpiresIn=900, # 1 s to 7 days allowed
)Limits that shape this code: one object can be up to 4.995 TiB; a single-part upload up to 4.995 GiB; a multipart upload has at most 10,000 parts. The part size has to be chosen so that the largest object fits within 10,000 parts. Keys can be up to 1,024 bytes and metadata up to 8,192 bytes. Presigned URLs work for GET, HEAD, PUT and DELETE on the S3 endpoint only. They do not work on custom domains, and HTML form POST uploads are not supported.
Code: a Worker in front of the bucket
// Worker: authenticated downloads with range and conditional support, plus a safe manifest update.
export default {
async fetch(request, env) {
if (!(await authorized(request, env))) return new Response("forbidden", { status: 403 });
const key = decodeURIComponent(new URL(request.url).pathname.slice(1));
if (request.method === "GET") {
const object = await env.MODELS.get(key, { range: request.headers, onlyIf: request.headers });
if (object === null) return new Response("not found", { status: 404 });
const headers = new Headers();
object.writeHttpMetadata(headers);
headers.set("etag", object.httpEtag);
// A failed precondition returns metadata without a body.
const status = object.body ? (request.headers.get("range") ? 206 : 200) : 304;
return new Response(object.body, { status, headers }); // add Content-Range for strict clients
}
return new Response("method not allowed", { status: 405 });
},
};
// Compare-and-swap on a small JSON manifest: put() returns null if the ETag no longer matches.
export async function publishManifest(env, key, manifest, expectedEtag) {
const res = await env.MODELS.put(key, JSON.stringify(manifest), {
onlyIf: { etagMatches: expectedEtag },
httpMetadata: { contentType: "application/json" },
});
if (res === null) throw new Error("manifest changed since it was read; re-read and retry");
return res.etag;
}Passing the request headers as range and onlyIf lets R2 evaluate Range, If-None-Match and related headers itself. When a precondition fails, get returns the object's metadata without a body, and put returns null and stores nothing. The manifest function relies on that behaviour. Readers fetch the manifest and its ETag, writers publish a new version only if nobody else did in between, and a loser re-reads and retries instead of silently overwriting.
Events, lifecycle and storage classes
Event notifications send a message to a Cloudflare Queue when objects are created (PutObject, CopyObject, CompleteMultipartUpload) or deleted (DeleteObject, or deletion by a lifecycle rule). Rules can filter by prefix and suffix. A bucket can have up to 100 rules, but rules that could fire twice for the same event are not allowed. The documentation currently gives a per-queue throughput of 5,000 messages per second and suggests spreading heavy workloads across queues.
// Queue consumer for R2 event notifications. Delivery can repeat, so every handler is idempotent.
export default {
async queue(batch, env) {
for (const msg of batch.messages) {
const { action, object } = msg.body;
try {
if (["PutObject", "CopyObject", "CompleteMultipartUpload"].includes(action)) {
await indexObject(env, object.key, object.eTag, object.size); // upsert keyed by key + eTag
}
msg.ack();
} catch (err) {
msg.retry();
}
}
},
};Treat every message as possibly repeated and possibly out of order. Key side effects on the object key plus its ETag, and verify the object with a head request before acting on a delete that might be stale.
Lifecycle rules can expire objects after an age and move them to Infrequent Access. A transition is billed as a Class A operation. Every bucket also has a default rule that aborts incomplete multipart uploads seven days after they start. Rules typically apply within 24 hours, so do not use them for anything that needs exact timing.
Infrequent Access is only cheaper for data that is genuinely cold. It has a 30-day minimum, higher operation prices and a retrieval charge per GB. The worked example below shows how quickly that charge outweighs the storage saving.
Data location
A bucket's data lives in one region, chosen at creation. A location hint (wnam, enam, weur, eeur, apac or oc) asks R2 to place it near your users or compute. A jurisdiction is stronger: it guarantees objects stay within a boundary such as the EU or FedRAMP (Enterprise only). The documentation lists the current set. Jurisdictional buckets use their own S3 endpoint, with the jurisdiction inserted after the account ID. Neither setting can be changed after the bucket is created, so moving data means creating a new bucket and copying. Decide residency before the first upload.
Migrating from another object store
Sippy migrates incrementally. You attach a source bucket from Amazon S3, Google Cloud Storage, Azure Blob Storage or another S3-compatible store. Each GetObject that misses in R2 is served from the source and copied into R2 at the same time. HeadObject does not trigger a copy, and PutObject and DeleteObject act only on R2. This means you can switch readers to R2 on day one and let hot data move itself.
Know its edges. Objects that change in the source after they are copied are not updated. ETags may differ from the source's, so do not use them to compare the two stores. Objects over 199 MiB may need several GETs to finish copying. Only x-amz-meta- user metadata comes across. For cold data that nobody will request, run a bulk copy as well. Cloudflare's Super Slurper service or any S3-compatible copy tool works. Then verify counts and sizes before you retire the source.
Worked example: model weights for a multi-cloud inference fleet
A team stores 20 TB of model checkpoints and dataset shards. Inference nodes in two other clouds, and a training cluster, pull about 400 TB a month, as about 4 million ranged GETs of 100 MB each. They write about 200,000 objects a month, mostly multipart parts.
In R2 Standard, before the free tier, storage is 20,000 GB x $0.015 = $300. Reads are 4 million x $0.36 per million = $1.44, writes about 0.2 million x $4.50 per million = $0.90, and egress is zero. The bill is about $302 a month and does not grow when another region starts pulling. With egress charged, those same 400 TB would almost certainly dominate the bill. Compare against your current provider's per-GB rate.
Now suppose someone moves everything to Infrequent Access to save money. Storage falls to $200, but retrieval adds 400,000 GB x $0.01 = $4,000, and reads cost 2.5 times as much. The right split is Standard for anything read in a given month, with a lifecycle rule that moves only old checkpoint generations to Infrequent Access. A data lake with the same access pattern behaves the same way; see cloud data lakes.
Failure modes
- Cached 404 after upload. A client probed the URL early; the public domain keeps returning 404. Use new keys, purge, or shorter TTLs.
- Hot single key. A status file rewritten by many workers hits the one-write-per-second limit. Shard it or use a queue or database.
- Presigned URL on a custom domain. Signatures only validate on the S3 endpoint.
- r2.dev in production. Rate limited and outside your cache and security settings.
- Wrong jurisdiction. A bucket created without the EU jurisdiction cannot be converted later.
- Infrequent Access for hot data. Retrieval fees exceed the storage saving within days.
- Non-idempotent event consumers. A redelivered message double-counts or double-processes.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| R2 versus a hyperscaler store | No egress; one price for every reader | No same-network data path to that cloud's compute; S3 API subset |
| Worker binding versus S3 API | No keys, logic at the edge | Code runs on Workers; different SDK |
| Public domain with cache | Fast global reads, fewer Class B operations | Stale objects and cached 404s |
| Sippy versus bulk copy | Zero-downtime cutover, only hot data moves | Long tail stays in the source until copied |
What to do next
- Estimate your monthly GETs, PUTs and stored GB, and price them in both storage classes before choosing.
- Decide the location hint or jurisdiction before creating the bucket.
- Create scoped R2 API tokens per application; keep administrative tokens out of code.
- Serve public content through a custom domain, use content-hashed or versioned keys, and purge on overwrite.
- Use presigned PUT URLs with short expiry for client uploads, against the S3 endpoint.
- Protect shared manifests with ETag-conditional writes and never rewrite one key many times per second.
- Add lifecycle rules for old generations and keep the default abort rule for incomplete multipart uploads.
- If migrating, enable Sippy, switch reads, bulk-copy the cold tail, then verify counts and sizes before retiring the source.