Storage classes in Cloud Storage trade a lower price per GB-month for higher access fees and minimum storage durations. Picking a class per object by hand is impossible at scale, and lifecycle rules that move objects by age assume you know how data ages. Autoclass is Google's alternative: a bucket setting under which Cloud Storage watches each object's last data read and moves it between classes for you, with the retrieval and early-deletion fees waived in exchange for a small per-object management fee.

That bargain is excellent for some buckets and a loss for others, and the difference comes down to object size, access pattern and whether anything scans the bucket. This article explains exactly how Autoclass moves objects, what it charges, how to enable it safely, how to calculate whether it pays and what goes wrong. The class price model itself is covered in Cloud Storage classes; here the focus is Autoclass. Behaviour is taken from the current Autoclass documentation; prices are illustrative variables, so use the pricing page for your location.

Advertisement

How Autoclass moves objects

Autoclass: each object moves colder on inactivity and back to Standard when readStandardall new objectsNearlineafter 30 daysColdlineafter 90 daysArchiveafter 365 daysColdline and Archive steps only when terminal class = ARCHIVE; the default terminal class is NEARLINEBack to Standardon an object data read (GET)Does not count as accessmetadata reads, copy, rewrite, listingObjects under 128 KiBstay in Standard; no management feeBillingmanagement fee per object; no retrieval or early-deletion feesThe clock is per object and based on the last data read, not on creation time alone.
Every object starts in Standard. With no data reads it moves to Nearline at 30 days and, if the terminal class is Archive, to Coldline at 90 and Archive at 365. A data read returns it to Standard and restarts its clock.

Autoclass is enabled per bucket and acts per object. New objects are written to Standard. An object that has not been read for 30 days moves to Nearline. What happens next depends on the bucket's terminal storage class. The default terminal class is Nearline, so objects stop there. If you set the terminal class to Archive, an object continues to Coldline after 90 days without access and to Archive after 365 days. Once an object reaches the terminal class, it stays there until it is read.

Transitions are asynchronous background work, not timers that fire at exactly 30 days; plan in days, not hours. Autoclass also cannot be combined with your own class-changing lifecycle rules.

What counts as access

This is the single most important detail. According to the documentation, only the object get method, a read of the object's data, updates the last-access time that drives Autoclass. Reading metadata does not. copy and rewrite do not move an object back to Standard either, and listing a bucket never touches object data. A data read of a Nearline, Coldline or Archive object moves it back to Standard, and the cooling clock starts again.

Two consequences follow. A dashboard that lists objects and reads their metadata will not disturb Autoclass. But any process that reads object data, including a virus scanner, a checksum audit, a backup tool that downloads everything, or a query engine that scans a whole prefix, will pull every object it touches back to Standard. That is the full-scan trap covered below.

Advertisement

The 128 KiB rule and the management fee

Objects smaller than 128 KiB never transition. They stay in Standard permanently, and they are not charged the management fee. The threshold counts object data only, not metadata. For objects at or above 128 KiB, Autoclass charges a management fee per 1,000 objects per month. In exchange there are no retrieval fees and no early-deletion fees, except as part of the one-time enablement charge described below. All operations are billed at Standard rates. Moving an object colder has no operation charge, returning from Nearline to Standard has no Class A charge, and returning from Coldline or Archive to a warmer class does incur a Class A operation charge.

Cost itemStandard onlyLifecycle rulesAutoclass
Storage per GB-monthHighestDepends on your rulesFollows each object's class
Retrieval fees on readsNoneYes, in colder classesNone
Early-deletion chargesNoneYes, below minimum durationsNone
Operation pricesStandardClass-specific, higher when colderStandard rates
Per-object feeNoneNoneManagement fee for objects of 128 KiB or more
Transition chargesNoneClass A operation per transitionNone going colder; Class A from Coldline or Archive back to warmer

Break-even by object size

The management fee is per object, while the savings are per byte, so small objects can cost more under Autoclass than in Standard. For one object that is cold enough to sit in the colder class, it pays when size multiplied by the price difference exceeds the fee. The calculation below uses variables; the fee and prices shown are placeholders with roughly the right shape, so replace them with the numbers from the pricing page for your location before deciding.

# Illustrative break-even for Autoclass. Replace every price with your location's list price.
FEE_PER_1000_OBJ_MONTH = 0.0025      # placeholder management fee, USD per 1,000 objects per month
PRICE = {"STANDARD": 0.020, "NEARLINE": 0.010, "COLDLINE": 0.004, "ARCHIVE": 0.0012}  # USD/GB-month

def breakeven_kib(target_class):
    fee_per_obj = FEE_PER_1000_OBJ_MONTH / 1000
    saving_per_gb = PRICE["STANDARD"] - PRICE[target_class]
    return fee_per_obj / saving_per_gb * 1024 * 1024        # GB -> KiB

for cls in ("NEARLINE", "COLDLINE", "ARCHIVE"):
    print(cls, round(breakeven_kib(cls)), "KiB")
# NEARLINE 262 KiB, COLDLINE 164 KiB, ARCHIVE 139 KiB  (with the placeholder prices)

def monthly_cost(objects):
    """objects: iterable of (size_bytes, steady_state_class). Storage plus fee only."""
    total = 0.0
    for size, cls in objects:
        gb = size / 2**30
        total += gb * PRICE[cls]
        if size >= 128 * 1024:
            total += FEE_PER_1000_OBJ_MONTH / 1000
    return total

With these placeholders, an object between 128 KiB and about 260 KiB that settles in Nearline costs slightly more under Autoclass than in Standard. Buckets of large files such as videos, model checkpoints, Parquet files or backups are nowhere near this line. A bucket of hundreds of millions of 200 KiB thumbnails sits right on it. Measure the size distribution before deciding, not the average, because a mean of 5 MB can hide a hundred million small objects.

Worked example

A team stores 400 TB of ML training artefacts: 2 million checkpoint and dataset shard files averaging 200 MB, written continuously and mostly read in the first two weeks. About 5 percent of old files are read again each quarter for re-evaluation. They have no lifecycle rules because nobody could agree on an age cut-off.

With Autoclass and terminal class Archive, fresh files stay in Standard while they are being used. Files not read for a month move to Nearline, and the long tail drifts to Coldline and Archive. The quarterly re-reads cost no retrieval fees; those files return to Standard and cool again. The management fee for 2 million objects is 2,000 units of the per-1,000 fee a month, which is small against storage on 400 TB. The one design change they made was to move a monthly integrity job, which downloaded every file to verify checksums, onto CRC32C values from object metadata. Reading metadata does not reset Autoclass. Downloading every object would have pulled the whole bucket back to Standard every month.

Enabling it: gcloud, JSON API and Python

On a new bucket, enable Autoclass at creation. On an existing bucket, the default storage class must be Standard, which the update command can set at the same time.

# New bucket
gcloud storage buckets create gs://ml-artifacts --location=us-central1 --enable-autoclass

# Existing bucket (default class must be STANDARD)
gcloud storage buckets update gs://ml-artifacts --default-storage-class=STANDARD --enable-autoclass

# Let objects continue past Nearline to Coldline and Archive
gcloud storage buckets update gs://ml-artifacts --autoclass-terminal-storage-class=ARCHIVE

# Disable (objects keep their current class; cannot re-enable for one day)
gcloud storage buckets update gs://ml-artifacts --no-enable-autoclass

The JSON API uses an autoclass object on the bucket resource with enabled and terminalStorageClass, sent in a PATCH. The Python client exposes the same settings as bucket properties.

# PATCH https://storage.googleapis.com/storage/v1/b/ml-artifacts
{"storageClass": "STANDARD", "autoclass": {"enabled": true, "terminalStorageClass": "ARCHIVE"}}

# google-cloud-storage
from google.cloud import storage
bucket = storage.Client().get_bucket("ml-artifacts")
bucket.autoclass_enabled = True
bucket.autoclass_terminal_storage_class = "ARCHIVE"
bucket.patch()

Enabling Autoclass on an existing bucket has two effects you must plan for. Every object except soft-deleted ones moves to Standard, including objects you had already placed in Coldline or Archive, and objects already in Standard start a fresh 30-day clock. That move triggers the one-time enablement charge, which is where the retrieval and early-deletion fees you would otherwise owe show up. The change can take up to a day to take effect. On a large bucket already sitting in Archive, enabling Autoclass means paying to warm everything and then waiting months for it to cool again.

Interactions and restrictions

  • Lifecycle rules: a bucket cannot have Autoclass together with lifecycle rules that use SetStorageClass actions or matchesStorageClass conditions; such requests fail. Remove those rules first. Rules that delete objects are a separate concern and still useful for retention.
  • Compose: composing an object in an Autoclass bucket fails unless all source objects are in Standard at the time of the request. Parallel composite uploads of fresh parts are fine; composing old parts may not be.
  • Disabling: objects keep whatever class they are in and ordinary class pricing resumes, including retrieval and early-deletion fees. You cannot re-enable Autoclass for one day.
  • Default storage class: must be Standard while Autoclass is on.

When lifecycle rules or a fixed class win

SituationBetter choiceWhy
Access is unpredictable or unknownAutoclassNo retrieval-fee risk from a wrong guess
Data is cold from day one (backups, compliance archives)Write directly to Archive or ColdlineAutoclass would hold it in Standard for 30 days first
Age predicts access well (logs read for 7 days)Lifecycle rulesPrecise, no management fee
Mostly objects of 128 KiB to about 250 KiBStandard or lifecycleManagement fee can exceed savings
Whole bucket scanned periodicallyLifecycle or StandardEvery scan resets objects to Standard
Mixed data lake with hot and cold prefixesAutoclassPer-object tracking beats per-prefix rules

The full-scan row is the most common surprise. A query engine reading a whole bucket once a month, such as an external table in BigQuery without partition pruning, a Spark job reading an unpartitioned path or a scanner downloading every object, keeps the entire bucket in Standard. You pay Standard storage plus the management fee. Partition your data so scans touch only recent prefixes, or keep scanned data in a separate bucket.

Monitoring and operations

Cloud Monitoring exposes Autoclass metrics, including autoclass/transition_operation_count and autoclass/transitioned_bytes_count, alongside the usual per-class storage metrics. After enabling, watch three things: bytes per storage class over time, which should drift colder over the first months; transitions back to Standard, where a large spike usually means a scan; and the management fee line on the bill, grouped by bucket. A monthly audit of every bucket's configuration is cheap:

for b in $(gcloud storage buckets list --format="value(name)"); do
  echo "== $b"
  gcloud storage buckets describe "gs://$b" --format="default(autoclass)"
done
# Example output for one bucket:
# autoclass:
#   enabled: true
#   terminalStorageClass: ARCHIVE

Cost attribution per bucket is part of wider cost practice; cloud FinOps covers labels, showback and anomaly alerts, which is where a surprise Autoclass warm-up should be caught. For how the service stores and serves objects underneath, see Google Cloud Storage internals.

Failure modes

  • Scanner resets the bucket: a data-reading job warms everything monthly. Look for transition spikes back to Standard and move the job to metadata or partitioned reads.
  • Small-object bucket gets more expensive: the fee outweighs savings. Check the object size histogram before enabling.
  • Enabling on an already cold bucket: a large enablement charge and months of Standard pricing. Model it first or leave that bucket on lifecycle rules.
  • Lifecycle update rejected: an old SetStorageClass rule blocks enablement. Remove class-changing rules in the same change.
  • Compose fails on old parts: sources left Standard. Compose soon after upload, or rewrite parts first.
  • Expecting instant savings: nothing moves for 30 days and new configuration can take a day. Evaluate after two to three months.

What to do next

  1. For each large bucket, pull the object size histogram and the read pattern from access logs or Storage Insights.
  2. Run the break-even calculation with your location's real prices and management fee.
  3. Find every job that reads whole buckets or prefixes, and fix or relocate it before enabling Autoclass.
  4. Enable Autoclass on new buckets with unknown access patterns, choosing ARCHIVE as the terminal class when data can be very cold.
  5. For existing buckets, model the enablement charge; skip buckets that are already mostly Coldline or Archive.
  6. Add per-class byte metrics, transition metrics and a per-bucket fee line to your cost dashboard, and review after 90 days.
Key takeaway: Autoclass tracks each object's last data read and moves it from Standard to Nearline at 30 days, and on to Coldline and Archive at 90 and 365 days if the terminal class is Archive. Any data read brings it back to Standard. It removes retrieval and early-deletion fees in exchange for a per-object management fee on objects of 128 KiB or more. It is the right default for large objects with unpredictable access. It is a poor choice for buckets of small objects, data that is cold from day one, or data something scans in full, and enabling it on an already cold bucket has a real one-time cost.