Alibaba Cloud Object Storage Service (OSS) is the object store at the centre of most Alibaba Cloud architectures. Images, logs, backups, data-lake tables and model checkpoints end up there, and services such as Elastic Compute Service (ECS), Function Compute and the analytics engines read from it. If you know Amazon S3 the shape will be familiar: buckets in regions, objects addressed by key, storage classes, lifecycle rules and signed URLs. The details differ, and the differences are where the money and the outages hide.

This page explains OSS from first principles and stays on the operational side. You will learn how buckets, endpoints and keys fit together; what each storage class actually costs you in minimum duration, minimum billable size and restore time; how to upload large objects with multipart upload and size the parts; how to use the Python SDK v2 for uploads, presigned URLs and lifecycle rules; how to control access with RAM policies, STS and bucket settings; and which failure modes to design for. If object storage itself is new, start with how cloud object storage works.

Buckets, keys and endpoints

A bucket is a container for objects that lives in one region. Bucket names must be unique across OSS, and an account can have up to 100 buckets per region according to the current limits page. An object is a key, the data, and metadata. The namespace is flat: a key such as datasets/2026/10/part-0001.parquet contains slashes, and the console shows them as folders, but there is no directory to rename or lock. Listing by prefix is how you get folder-like views.

You reach a bucket through an endpoint. The public endpoint has the form oss-<region-id>.aliyuncs.com and a bucket's own domain is <bucket>.oss-<region-id>.aliyuncs.com. The internal endpoint, oss-<region-id>-internal.aliyuncs.com, is reachable from ECS and other services in the same region over Alibaba Cloud's network and avoids internet outbound traffic charges. A transfer acceleration endpoint exists for distant clients once the feature is enabled on the bucket. One recent change matters for new projects: Alibaba states that from 20 March 2025, new OSS users must use a custom domain name (CNAME) for data API operations on buckets in Chinese mainland regions, so plan a domain and certificate before you build there.

Every request is authenticated with an AccessKey pair or temporary STS credentials and signed. Use an SDK or ossutil rather than signing by hand; the SDKs also check a CRC-64 of the data against the x-oss-hash-crc64ecma value OSS returns, which catches corruption in transit.

Storage classes and redundancy

Storage classes trade storage price against access price, latency and commitment. The table reflects the OSS documentation at the time of writing; prices vary by region and change, so take them from the pricing page, but the rules below are what drive bills.

ClassMinimum storage durationMinimum billable sizeAccess
StandardNoneActual sizeReal time, no retrieval fee
Infrequent Access (IA)30 days64 KBReal time, retrieval fee per GB read
Archive60 days64 KBRestore in about a minute, or read directly with a retrieval fee if real-time access is enabled
Cold Archive180 days64 KBMust restore first; 1 to 12 hours
Deep Cold Archive180 days64 KBMust restore first; 12 or 48 hours

Redundancy is a separate choice. Standard, IA and Archive are available as zone-redundant storage (ZRS), which spreads data across zones in the region and survives the loss of one, or as locally redundant storage (LRS) at a lower price. Cold Archive and Deep Cold Archive are LRS only. For anything you cannot regenerate, ZRS plus a replica in another region is the usual baseline.

ECS / Function Computesame regionBrowser or mobile appinternetApp serverissues STS / presignInternal endpointoss-region-internalPublic or CNAME domainor accelerationBucketkeys, versions, metadataStandardhotIA30 day minArchive60 day minCold Archiverestore firstReplica bucketcross-region replicationtoken or URLlifecycleClients sign requests; in-region traffic uses the internal endpoint; lifecycle moves objects down the classes.
Access paths and data placement. Choose the endpoint by where the client runs, and let lifecycle rules, not people, move data between classes.

Uploading: simple, append and multipart

OSS offers three ways to write an object. A simple upload (PutObject), and the form upload browsers use, accept objects up to 5 GB in a single request. An append upload creates an appendable object you can extend, useful for logs written in pieces. A multipart upload splits an object into parts, uploads them independently and in parallel, then assembles them; it raises the object size limit to 48.8 TB.

Multipart works in three calls: InitiateMultipartUpload returns an upload id; UploadPart sends each part with a part number from 1 to 10,000; CompleteMultipartUpload lists the parts and their ETags so OSS can assemble the object. Every part except the last must be at least 100 KB, and no part may exceed 5 GB. Parts uploaded but never completed are stored and billed until you abort the upload, so every bucket needs a lifecycle rule that cleans them up.

Worked example. A training job writes 200 GB checkpoint files from ECS. With a 10,000-part ceiling, the part size must be at least 200 GB divided by 10,000, which is 20 MB. Choose 64 MB: that is 3,200 parts, small enough that a failed part costs seconds to resend, large enough that per-request overhead is negligible. With 8 parts in flight over the internal endpoint, throughput is limited by the instance's network bandwidth rather than by OSS. The SDK's uploader does this arithmetic and also records completed parts in a checkpoint file, so a crashed upload resumes instead of starting again.

Using the Python SDK v2

The Python SDK v2 is installed as alibabacloud-oss-v2 and imported as alibabacloud_oss_v2. The example below reads credentials from the environment, writes a checkpoint through the uploader, reads it back, hands a short-lived download link to a client, and installs a lifecycle rule. Names follow the SDK's own samples.

import datetime
import alibabacloud_oss_v2 as oss

cfg = oss.config.load_default()
cfg.credentials_provider = oss.credentials.EnvironmentVariableCredentialsProvider()
cfg.region = "cn-hangzhou"
cfg.use_internal_endpoint = True  # running on ECS in the same region
client = oss.Client(cfg)
BUCKET = "ml-checkpoints-prod"

# Multipart upload with resumable checkpoints.
uploader = client.uploader(part_size=64 * 1024 * 1024, parallel_num=8,
                           enable_checkpoint=True, checkpoint_dir="/var/tmp/oss-ckpt")
res = uploader.upload_file(
    oss.PutObjectRequest(bucket=BUCKET, key="run-42/step-90000.pt"),
    filepath="/data/ckpt/step-90000.pt")
print(res.etag, res.hash_crc64)

# Read a small object.
obj = client.get_object(oss.GetObjectRequest(bucket=BUCKET, key="run-42/config.json"))
with obj.body as f:
    config = f.read()

# A download link that expires in ten minutes.
signed = client.presign(oss.GetObjectRequest(bucket=BUCKET, key="run-42/eval.html"),
                        expires=datetime.timedelta(minutes=10))
print(signed.url, signed.expiration)

# Lifecycle: IA after 30 days, Archive after 90, delete after 365, and abort
# multipart uploads left incomplete for 7 days.
client.put_bucket_lifecycle(oss.PutBucketLifecycleRequest(
    bucket=BUCKET,
    lifecycle_configuration=oss.LifecycleConfiguration(rules=[oss.LifecycleRule(
        id="checkpoints-tiering", prefix="run-", status="Enabled",
        transitions=[oss.LifecycleRuleTransition(days=30, storage_class="IA"),
                     oss.LifecycleRuleTransition(days=90, storage_class="Archive")],
        expiration=oss.LifecycleRuleExpiration(days=365),
        abort_multipart_upload=oss.LifecycleRuleAbortMultipartUpload(days=7),
    )])))

The environment provider reads OSS_ACCESS_KEY_ID and OSS_ACCESS_KEY_SECRET. On ECS, prefer an instance RAM role so there is no long-lived key on disk at all. Pick transition days at or above each class's minimum duration: moving an object to IA and deleting it ten days later is billed as if it stayed thirty.

Access control and data protection

Access control has several layers, and you should know which one granted a request.

  • RAM policies attach to users, groups and roles in your account. Resources use the form acs:oss:<region>:<account>:<bucket>/<key-pattern>, and actions are named like oss:GetObject. Grant by prefix, never oss:* on *.
  • Bucket policies attach to the bucket and can grant other accounts or anonymous users access, with conditions such as source IP or VPC. They are the right place for cross-account sharing and the wrong place for everyday user permissions.
  • ACLs on buckets and objects (private, public-read, public-read-write) are the oldest layer. Keep buckets private and turn on Block Public Access, which OSS began enabling by default for new buckets in October 2025, so a stray ACL or policy cannot expose data.
  • STS temporary credentials let a browser or mobile app upload directly. Your server calls AssumeRole for a role whose policy allows only oss:PutObject on uploads/<user-id>/*, and passes the short-lived credentials to the client. Bytes never pass through your servers, and a leaked token expires soon and only reaches one prefix.
  • Presigned URLs are simpler still for a single object: the server signs one method on one key for a few minutes.
{
  "Version": "1",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["oss:PutObject"],
    "Resource": ["acs:oss:*:*:user-uploads-prod/uploads/u-1842/*"]
  }]
}

For data protection, enable versioning on buckets where overwrites or deletes would hurt, server-side encryption with OSS-managed keys or KMS, and a retention (WORM) policy where regulation requires immutable records. Cross-region replication copies new writes to a bucket in another region; it is a recovery tool, not a backup, because deletes and bad writes can replicate too, which is why versioning on both sides matters.

Limits and performance

The documented limits are per account and per region rather than per bucket: around 10,000 requests per second for an account, and regional bandwidth that ranges from a few Gbps in smaller regions to 20 Gbps of upload and 100 Gbps of download in the largest. If you plan a burst near those numbers, such as a large training job reading millions of shards at once, talk to Alibaba Cloud support first.

Within those limits, throughput is mostly a client problem. Use the internal endpoint for in-region compute, parallelise ranged GETs and multipart PUTs, reuse connections, and pack tiny files into larger shards: a request has fixed latency, so a million 10 KB reads are dominated by round trips. Batch deletes accept up to 1,000 keys per request.

Where the money goes

An OSS bill has five parts: storage by class and redundancy, requests, outbound internet traffic, retrieval for IA and archive classes, and the penalties of minimum duration and minimum billable size. Two worked numbers show how the rules bite.

Small objects in IA. Five million 8 KB thumbnails total about 40 GB. In IA each is billed as 64 KB, so you pay for roughly 320 GB, eight times the data. Leave small objects in Standard or pack them into archives before tiering.

Early deletion. A log pipeline writes to Archive and a cleanup job deletes after 14 days. Each object is still billed for 60 days of Archive storage. The fix is to transition later or not at all, since short-lived data belongs in Standard.

Internet egress usually dominates for public content. Put a CDN in front of public buckets and keep compute in the same region as its data so reads use the internal endpoint.

Failure modes

  • Orphaned multipart parts accumulate silently after crashed uploads. A lifecycle abort rule fixes it.
  • Reading Cold Archive objects directly fails until a restore completes, and restores take hours. Applications must handle the restore state, not retry blindly.
  • Public buckets created for a quick demo and never closed. Block public access at the account level and alert on policy changes.
  • Long-lived AccessKeys in code or images. Use RAM roles for ECS and Function Compute, STS for clients, and rotate anything that remains.
  • Assuming S3 semantics everywhere. OSS has its own API and SDKs. Alibaba documents some S3 compatibility, but test any S3 tool against your exact operations before relying on it.
  • Mainland China domain rules. New users in mainland regions need a CNAME for data operations, which also means ICP filing considerations for the domain. Plan this before launch, not during it.

What to do next

  1. Choose the region next to your compute, and confirm whether you need a custom domain for mainland China regions.
  2. Create buckets private, with block public access, versioning where data matters, and ZRS for anything you cannot rebuild.
  3. Add a lifecycle rule to every bucket that aborts incomplete multipart uploads, and tier only objects that are large and old.
  4. Replace AccessKeys with RAM roles on ECS and STS or presigned URLs for clients, scoped to prefixes.
  5. Use the internal endpoint for in-region traffic and a CDN for public content.
  6. Use the SDK uploader with checkpoints for large files and size parts so they stay under 10,000.
  7. Review the bill monthly by class, request type and traffic, and compare against the patterns in the cost section. To compare designs across clouds, read Amazon S3 in depth, OCI Object Storage and Alibaba ECS for the compute that usually reads from OSS.
Key takeaway: OSS stores objects under flat keys in regional buckets reached through public, internal or accelerated endpoints. Choose storage classes by access pattern, remembering the 30, 60 and 180 day minimum durations and the 64 KB minimum billable size, and choose ZRS for data you cannot rebuild. Use multipart upload with checkpoints for large files, keep parts under 10,000, and always abort incomplete uploads with a lifecycle rule. Keep buckets private, grant access by prefix through RAM roles, STS and presigned URLs, use the internal endpoint for in-region compute, and check mainland China domain requirements before you launch.