Object Versioning in Cloud Storage keeps the old copy whenever an object is overwritten or deleted. With it enabled, a bad deploy that overwrites a model file or a script that deletes the wrong prefix can be undone with a copy command. Without it, the data is gone once any soft delete window passes. The feature is one flag on the bucket, but using it well means understanding generations, preconditions, restore procedures and cost.
This article explains the generation model from first principles, what each operation does to a versioned object, how to use generation preconditions for safe concurrent writes, how to restore a single object or a whole prefix to a point in time, how to measure and control the cost of old versions, and where versioning sits next to soft delete, retention and backups. Lifecycle rules for noncurrent versions are covered in Cloud Storage lifecycle management, so they are only summarised here.
Generations, metagenerations and noncurrent versions
Every object in Cloud Storage has two numbers. The generation identifies a particular content version. It is set when the object is written and changes whenever the content is replaced. The metageneration starts at 1 for each generation and increases each time that generation's metadata is updated. The pair (name, generation) names one immutable set of bytes, and the full storage architecture behind that is described in Google Cloud Storage internals.
Generations exist on every bucket. Versioning only decides what happens to the old generation when a new one replaces it, or when the object is deleted. Without versioning, the old generation is removed. With versioning, it stays in the bucket as a noncurrent version. At most one generation of a name is live: the one that ordinary reads and listings return.
| Operation on a versioned bucket | Effect |
|---|---|
| Upload to an existing name | New live generation; previous live version becomes noncurrent |
| Delete by name only | Live version becomes noncurrent; the name has no live object |
Delete name#generation | That version is deleted (soft-deleted if soft delete is on) |
| Patch metadata | Same generation, metageneration increases; no new version |
Copy name#old onto name | New live generation with the old bytes; history kept |
| Disable versioning | No new noncurrent versions; existing ones stay until deleted |
There is no default limit on how many versions an object can have. Each noncurrent version is billed at the rate of its storage class, exactly as when it was live, and that is where most of the surprises come from.
Enabling, listing and addressing versions
Enable versioning on the bucket, then use the #generation suffix to address individual versions. Quote the URL in shells where # is special.
gcloud storage buckets update gs://acme-ml --versioning
# All versions, with generation numbers; --long adds metageneration
gcloud storage ls --all-versions --long 'gs://acme-ml/models/ranker.pt'
# Inspect one version
gcloud storage objects describe 'gs://acme-ml/models/ranker.pt#1712074110223001'
# Restore: copy an old generation over the live name (creates a new generation)
gcloud storage cp 'gs://acme-ml/models/ranker.pt#1712074110223001' gs://acme-ml/models/ranker.pt
# Permanently remove one version
gcloud storage rm 'gs://acme-ml/models/ranker.pt#1712074110223002'The Python client exposes the same model: list_blobs(versions=True) returns noncurrent versions, each blob has generation and metageneration, and a noncurrent blob's time_deleted records when it stopped being live. That timestamp is the key to point-in-time restores.
Generation preconditions: compare-and-swap for objects
Versioning keeps history, but it does not stop two writers from overwriting each other. Generation preconditions do, and they work on any bucket, versioned or not. Every write can carry a condition. ifGenerationMatch=0 means "only create this if no live object exists". ifGenerationMatch=N means "only replace it if the live generation is still N". If the condition fails, the request returns HTTP 412 and nothing changes. This gives you compare-and-swap on objects.
import json
from google.api_core.exceptions import PreconditionFailed
from google.cloud import storage
client = storage.Client()
bucket = client.bucket("acme-ml")
def create_once(name: str, data: bytes) -> bool:
"""Write only if the name has no live object. Safe for idempotent job outputs."""
try:
bucket.blob(name).upload_from_string(data, if_generation_match=0)
return True
except PreconditionFailed:
return False # someone else got there first
def update_pointer(name: str, mutate, retries: int = 5) -> int:
"""Read-modify-write a small JSON object without lost updates."""
for _ in range(retries):
blob = bucket.get_blob(name) # fetches current generation
doc = json.loads(blob.download_as_bytes(if_generation_match=blob.generation))
mutate(doc)
try:
blob.upload_from_string(json.dumps(doc), content_type="application/json",
if_generation_match=blob.generation)
return blob.generation # new generation after upload
except PreconditionFailed:
continue # lost the race; re-read and retry
raise RuntimeError(f"contention on {name}")Use the same idea for restores. When you copy an old generation back, make the copy conditional on the live generation you inspected, so you do not overwrite a fix someone shipped while you were investigating. One limit to design around: Cloud Storage allows roughly one write per second to the same object name, so a hot pointer object is a bottleneck no matter how you guard it.
Restoring an object or a prefix to a point in time
Restoring one object is a single copy. Restoring a prefix to the state it had at time T needs a rule for each name. The version that was live at T is the one created at or before T that did not become noncurrent until after T. If no version matches, the name did not exist at T, and any current live object should be deleted, which makes it noncurrent and therefore recoverable too.
from collections import defaultdict
from datetime import datetime
def restore_prefix(bucket, prefix: str, t: datetime, dry_run: bool = True):
"""t must be timezone-aware (UTC); blob timestamps are aware datetimes."""
by_name = defaultdict(list)
for b in bucket.list_blobs(prefix=prefix, versions=True):
by_name[b.name].append(b)
for name, versions in sorted(by_name.items()):
live = next((v for v in versions if v.time_deleted is None), None)
target = next((v for v in versions
if v.time_created <= t and (v.time_deleted is None or v.time_deleted > t)), None)
if target is None and live is not None:
action = ("delete", name, live.generation) # did not exist at t
if not dry_run:
bucket.delete_blob(name, if_generation_match=live.generation)
elif target is not None and (live is None or live.generation != target.generation):
action = ("restore", name, target.generation)
if not dry_run:
bucket.copy_blob(target, bucket, name,
source_generation=target.generation,
if_generation_match=live.generation if live else 0)
else:
continue # already correct
print(*action)Always run it with dry_run=True first and review the plan. Restored objects get new generations and new creation times, so anything that keys on timeCreated, such as a lifecycle age rule, sees them as new. Restores also double the stored bytes for those objects until lifecycle trims the history.
Worked example: rolling back a model artefact
Here is a worked example. A training pipeline writes models/ranker.pt (2.1 GB) to a versioned regional bucket after every successful run, and a serving fleet loads it on start-up. On Tuesday a run with a data bug passes its weak evaluation and overwrites the file as generation …002. By Wednesday a run with the bug fixed has written …003, but serving metrics show that run is still worse than Monday's model.
The on-call engineer lists versions and sees …001 (Monday), …002 and …003. They describe …001 to confirm its size and metadata, record that the live generation is …003, then copy …001 onto the name with if_generation_match set to …003. The result is a new live generation, …004, holding Monday's bytes. Serving restarts and loads it. Nothing was deleted, so if Monday's model also turns out to be wrong, the history is still there.
Now the cost. The pipeline writes daily, and lifecycle keeps noncurrent versions for 30 days, so about 30 × 2.1 GB = 63 GB of history sits behind 2.1 GB of live data, a 30-to-1 ratio. At regional Standard list prices of roughly two US cents per GB-month (check your region), that is about $1.30 a month, which is trivial. The same policy on a 5 TB feature-store prefix rewritten nightly would hold about 150 TB of history. That is why noncurrent bytes need to be measured per prefix.
What history costs, and how to measure it
Three things drive the cost. First, noncurrent versions are billed like live ones, so churn multiplies storage. Second, early deletion charges for Nearline, Coldline and Archive are counted from when the version was uploaded, not from when it became noncurrent. Deleting a Coldline version 40 days after upload is still charged up to the 90-day minimum. Third, soft delete keeps deleted versions billable for its window, which defaults to 7 days. Class minimums and prices are in Cloud Storage classes.
Measure before tuning. The script below sums live and noncurrent bytes by top-level prefix, so you can see where history accumulates.
from collections import Counter
live, old = Counter(), Counter()
for b in client.bucket("acme-ml").list_blobs(versions=True):
top = b.name.split("/", 1)[0]
(old if b.time_deleted else live)[top] += b.size
for top in sorted(set(live) | set(old), key=lambda k: -old[k]):
ratio = old[top] / live[top] if live[top] else float("inf")
print(f"{top:30s} live={live[top]/1e9:9.1f} GB noncurrent={old[top]/1e9:9.1f} GB x{ratio:.1f}")Listing every version is slow on large buckets, so run it as a scheduled job or use Storage Insights inventory reports instead. Then cap history with lifecycle rules that target noncurrent versions by count and by days since they became noncurrent. A plain age-based delete on a versioned bucket only turns live objects into noncurrent ones and frees nothing.
Versioning, soft delete, retention and copies
Versioning is one of four protections, and they cover different threats.
| Feature | Protects against | Does not protect against |
|---|---|---|
| Object Versioning | Overwrites and deletes by name; gives unlimited, inspectable history | A principal who deletes specific generations |
| Soft delete (7-90 days, default 7) | Any delete, including deleting noncurrent versions | Anything older than the window; overwrites without versioning |
| Retention policy and holds | Deletion or replacement before a date, even by admins once locked | Nothing after the period ends; not a history tool |
| Copy to another bucket or project | Loss of the bucket, project or credentials | Replication lag; doubles cost |
Google recommends soft delete as the default protection against accidental deletion, and it is on by default. Versioning earns its cost when you need overwrite history: artefacts, configuration, small datasets that get edited in place. Note that anyone with permission to delete objects can delete noncurrent generations by number, so versioning alone is not ransomware protection. Combine it with soft delete, a retention policy for data that must survive, and a copy under separate credentials. Retention and holds work alongside versioning, and are covered in GCS retention policy and Object Lock.
One hard incompatibility: buckets with hierarchical namespace enabled do not support Object Versioning. If you need folder semantics and overwrite history, use soft delete and application-level versioned names.
Failure modes
- Unbounded history. Versioning on, with no noncurrent lifecycle rule, on a churning prefix. Storage grows linearly forever.
- Age rules that free nothing. A delete rule with no
isLivecondition only creates noncurrent versions. - rsync with deletes. A sync that removes unmatched destination objects turns every removal into a noncurrent version, which is safe but billed.
- Racing restores. An unconditional copy over a name that someone fixed minutes earlier. Use
if_generation_match. - Assuming versioning is WORM. Generations can be deleted by anyone with delete permission. Use a retention policy for compliance.
- Cold-class churn. Versioned Archive objects rewritten monthly pay early deletion charges on every noncurrent version.
- Turning versioning off to save money. Existing noncurrent versions stay and keep billing. Delete them explicitly or with lifecycle.
What to do next
- List your buckets and record which have versioning, soft delete and retention, and why.
- Enable versioning only where overwrite history has a use; rely on soft delete elsewhere.
- Add noncurrent lifecycle rules (by count and by days) to every versioned bucket before enabling it.
- Run the noncurrent-bytes report per prefix and set an alert on the ratio.
- Use
if_generation_matchfor create-once outputs, pointer updates and every restore. - Script and rehearse a point-in-time prefix restore with a dry run.
- For data that must survive a compromised account, add a retention policy and a copy under separate credentials. Consider dual- or multi-region placement for location outages.