Object storage is the durable substrate under most cloud systems: data lakes, model checkpoints, backups, logs, static websites and training datasets all end up in buckets. Its interface looks trivially simple, a key maps to a blob of bytes over HTTP, and that simplicity tempts teams to treat it like a slow file system. The expensive incidents come from that mistake: a job that lists a bucket to find its input and misses data, two writers that silently overwrite each other, a request pattern that concentrates on one key range and gets throttled, or a lifecycle rule that costs more than the storage it saves.
This article explains how object stores are built inside, what guarantees they actually give, and how to design keys, writes, uploads and lifecycle policies around those guarantees. It is provider-neutral where the concepts are shared and names Amazon S3 specifics where they are documented and widely copied; Google Cloud Storage, Azure Blob Storage and OCI Object Storage follow the same architecture with different names and limits.
What an object store is, and is not
An object store holds objects: an immutable byte sequence plus metadata, addressed by a bucket (Azure calls it a container) and a key. There are no directories: the key logs/2026/09/30/app.gz is one flat string, and the slashes only look like a hierarchy because list operations can group keys by a delimiter. You cannot modify part of an object in place; you replace the whole object or write a new key. You cannot rename cheaply; a rename is a copy followed by a delete.
These constraints are what make the service scale. Because objects are immutable, any replica or fragment of an object version is either correct or absent, never half-updated. Because the namespace is flat, the index can be partitioned by key range across thousands of servers without coordinating a directory tree. File semantics (append, rename, locks, partial overwrite) belong in a file service; block semantics belong in a volume. Choosing object storage means designing the application around put, get, list, delete and conditional writes.
Inside the service: front ends, index and data
Every major object store separates three layers. A fleet of stateless front ends terminates HTTPS, authenticates the request signature, evaluates access policies and throttles abusive patterns. A metadata index, partitioned by bucket and key range, maps each key to its current version, size, ETag, encryption details and the location of its data. A data layer stores the bytes on large fleets of disks across several failure domains.
Durability comes from redundancy across failure domains, usually with erasure coding rather than full copies. An object is split into k data fragments and m parity fragments placed on different disks, racks or availability zones; any k of the k+m fragments reconstruct it. A 10+4 scheme survives four lost fragments while storing 1.4 bytes per byte instead of the 3 bytes of triple replication. HDFS erasure coding explains the same mathematics in a system whose internals you can read. Background services continuously scrub fragments to detect silent corruption and repair lost fragments before further failures accumulate; the published eleven-nines durability figures are really statements about repair speed relative to failure rates.
The ordering on a write is the key to everything else: the service first makes the bytes durable, then commits the index entry that points to them. Until the commit, readers see the old version or nothing; after it, they see the new version. A failed upload leaves orphaned fragments that garbage collection reclaims, never a half-written object.
Consistency: what you can and cannot rely on
Amazon S3 has provided strong read-after-write consistency for PUT and DELETE of objects, and for list operations, since December 2020; Google Cloud Storage and Azure Blob Storage are also strongly consistent for single-object operations. A successful write is visible to every subsequent read. That removes a whole category of old workarounds, such as waiting before reading a freshly written manifest.
What you still do not get is equally important. There are no multi-object transactions: writing ten output files and a success marker is ten independent operations, and a reader can observe any prefix of them. Concurrent writers to the same key follow last-writer-wins: both PUTs succeed, and one silently disappears. Cross-region replication is asynchronous, so a replica can lag by seconds to minutes, and a read in the other region is not strongly consistent with a write in the first.
The standard pattern for atomic multi-file output is write data first, commit a manifest last: readers only ever open files named in a committed manifest, never files discovered by listing. Table formats such as Apache Iceberg and Delta Lake are elaborate versions of this idea.
Conditional writes as a concurrency primitive
Last-writer-wins is fixed by conditional requests. S3 supports If-None-Match: * on PutObject, which succeeds only if no object exists at the key, and If-Match: <etag>, which succeeds only if the current object still has the ETag you read. A failed condition returns HTTP 412 Precondition Failed. Together they give compare-and-swap on a single key, which is enough to build leases, leader election and optimistic concurrency for manifests. Google Cloud Storage offers the same capability through generation-match preconditions, and Azure Blob Storage through standard HTTP ETag conditions.
import json
import uuid
import boto3
from botocore.exceptions import ClientError
s3 = boto3.client("s3")
def commit_manifest(bucket, table, files, max_attempts=5):
"""Optimistic commit: read pointer, write new manifest, swap pointer only if unchanged."""
key = f"{table}/_current.json"
for _ in range(max_attempts):
try:
head = s3.get_object(Bucket=bucket, Key=key)
current, etag = json.load(head["Body"]), head["ETag"]
except ClientError as e:
if e.response["Error"]["Code"] != "NoSuchKey":
raise
current, etag = {"version": 0, "files": []}, None
version = current["version"] + 1
# Unique, immutable manifest per attempt; losers leave unreferenced garbage.
manifest_key = f"{table}/manifests/{version:010d}-{uuid.uuid4().hex}.json"
s3.put_object(Bucket=bucket, Key=manifest_key,
Body=json.dumps({"files": current["files"] + files}))
new = {"version": version, "manifest": manifest_key,
"files": current["files"] + files}
cond = {"IfMatch": etag} if etag else {"IfNoneMatch": "*"}
try:
s3.put_object(Bucket=bucket, Key=key, Body=json.dumps(new), **cond)
return version
except ClientError as e:
if e.response["ResponseMetadata"]["HTTPStatusCode"] in (409, 412):
continue # another writer won; re-read and retry
raise
raise RuntimeError("commit contention: too many concurrent writers")Two details matter in practice. Each attempt writes its manifest under a unique name and only the conditional pointer swap decides the winner, so a writer that crashes between the two writes leaves an unreferenced manifest for a cleanup job, never a blocked or corrupted table. And the retry loop must re-read and rebuild the new state from the winner's version, not blindly re-send its own; otherwise it would erase the other writer's files.
Key design and request-rate scaling
The index is partitioned by key range, and each partition serves a bounded request rate. S3 documents at least 3,500 PUT, COPY, POST or DELETE requests and 5,500 GET or HEAD requests per second per partitioned prefix, and it splits hot ranges automatically as sustained load grows. Splitting takes time, so a sudden burst on one narrow key range returns HTTP 503 Slow Down responses until the service adapts. Other providers behave similarly with different numbers.
Design keys so that load spreads across ranges from the start. If thousands of workers write keys beginning with the same date, all traffic lands on one range; putting a high-cardinality component such as a tenant or shard number early in the key spreads it. Keep prefixes that readers need to list together, because list operations are range scans. Always use SDK retries with exponential backoff and jitter, and treat 503 as a signal to slow the whole job, not each request individually.
Listing deserves separate care. It returns at most 1,000 keys per page on S3 and is sequential within a prefix, so listing hundreds of millions of keys takes a long time and costs a request per page. For analytics over large buckets, use the provider's scheduled inventory reports rather than listing, and for pipelines, pass explicit key lists or manifests between stages.
Large objects: multipart upload and ranged reads
Large objects are uploaded in parts. On S3 a multipart upload has up to 10,000 parts, each between 5 MiB and 5 GiB except the last, which may be smaller; since December 2025 the maximum object size is 50 TB. Parts upload in parallel and can be retried individually, so a failed network connection costs one part rather than the whole object, and throughput scales with the number of concurrent connections. The object only appears when the client sends the complete request listing the parts; until then nothing is visible.
MiB, GiB = 1024**2, 1024**3
def choose_part_size(object_bytes, target=64 * MiB, max_parts=10_000):
"""Smallest part size >= target that keeps the upload within max_parts."""
part = max(target, 5 * MiB)
while -(-object_bytes // part) > max_parts: # ceiling division
part *= 2
if part > 5 * GiB:
raise ValueError("object too large for a single multipart upload")
return part
print(choose_part_size(2 * 1024**4) // MiB) # a 2 TiB checkpoint -> 256 MiB partsAn abandoned multipart upload keeps its parts, and you pay for them, invisibly: they do not appear in normal listings. Every bucket that receives multipart uploads should have a lifecycle rule that aborts incomplete uploads after a few days. On the read side, HTTP range requests fetch byte ranges of an object in parallel, which is how high-throughput readers load training shards and how columnar formats such as Parquet read only the column chunks a query needs.
Storage classes and lifecycle economics
Storage classes trade storage price against access price and latency. Colder classes charge less per gigabyte-month but more per request and per gigabyte retrieved, and impose a minimum storage duration: on S3, Standard-IA bills at least 30 days, Glacier Flexible Retrieval 90 days and Glacier Deep Archive 180 days; on Google Cloud Storage, Nearline 30, Coldline 90 and Archive 365. Standard-IA also bills each object as at least 128 KB. Archive classes need a restore before reading, which takes minutes to hours.
Lifecycle rules move or expire objects by age, prefix or tag, but each transition is a billed request per object. That makes object size the deciding variable. Moving one million 1 GB objects to a colder class is a million transition requests against a petabyte of savings; moving a hundred million 20 KB log fragments is a hundred million requests against 2 TB, and because of the 128 KB minimum, each fragment is billed as if it were six times larger. The right move for small objects is to compact them into large archives first, then tier the archives. Write the arithmetic down for each rule: monthly storage saved, minus transition requests, minus expected retrieval, minus early-deletion charges.
| Concept | Amazon S3 | Google Cloud Storage | Azure Blob Storage | OCI Object Storage |
|---|---|---|---|---|
| Namespace unit | bucket | bucket | storage account + container | namespace + bucket |
| Colder tiers | Standard-IA, Glacier classes | Nearline, Coldline, Archive | Cool, Cold, Archive | Infrequent Access, Archive |
| Optimistic concurrency | If-Match / If-None-Match | generation preconditions | ETag conditions | ETag conditions |
| Immutability | Object Lock | retention policies, bucket lock | immutability policies | retention rules |
Encryption keys, versioning, immutability and replication complete the protection story; S3 encryption options compares key-management choices, and OCI Object Storage walks through one provider end to end. Remember that replication and cross-region reads also incur transfer charges, analysed in cloud egress cost.
Failure modes seen in production
- Discovering input by listing. A job lists a prefix while an upstream writer is mid-way through ten files and processes a partial set. Commit manifests instead.
- Silent lost update. Two workers read-modify-write the same JSON state object; one update vanishes. Use If-Match or move the state to a database.
- Hot key range. A launch sends all writes to one date prefix and the job fails with 503 errors. Spread keys and ramp gradually.
- Invisible cost. Incomplete multipart uploads and noncurrent versions accumulate for years. Add abort and noncurrent-version expiration rules.
- Lifecycle that loses money. Tiering millions of tiny objects costs more in requests and minimum-size billing than it saves.
- Accidental exposure. A bucket policy or object ACL grants public read. Enable account-level public-access blocks and review policies in CI.
- Replication treated as backup. A delete or ransomware encryption replicates too. Pair replication with versioning and immutability.
What to do next
- Inventory your buckets: size, object count, size distribution, request rates and which prefixes are hot.
- Find every place that discovers data by listing and replace it with explicit manifests.
- Find every read-modify-write of a shared object and protect it with a conditional write or move it to a database.
- Add lifecycle rules to abort incomplete multipart uploads and expire noncurrent versions.
- For each tiering rule, compute storage saved against transition, minimum-duration and retrieval charges; compact small objects first.
- Enforce public-access blocks, default encryption and versioning, and add immutability for backups.
- Load-test key layouts at the request rates you expect at launch, and confirm your SDK retries 503s with backoff and jitter.