Alibaba Cloud Elastic Compute Service (ECS) is Alibaba's virtual machine service, the equivalent of Amazon EC2 or Google Compute Engine. It is not related to Amazon ECS, the container scheduler with the same initials. If you deploy in mainland China or across Asia-Pacific you will meet it sooner or later, and while the concepts map closely to other clouds, the names, billing models and some defaults differ in ways that bite.

This article explains ECS from the ground up: the resource model, how to read an instance type, the four ways to pay, how to launch instances repeatably with the CLI, how to run spot instances safely with the five-minute interruption notice, storage and networking choices, metadata hardening, a worked example of a batch fleet, failure modes and a checklist. API parameter names were checked against Alibaba Cloud's documentation at the time of writing.

Advertisement

The resource model

Every ECS instance sits in a region and a zone, which is an isolated data-centre group inside the region. Networking uses a VPC that spans the region and vSwitches (subnets) that each belong to one zone, so the vSwitch you pick decides the zone. An instance is defined by an image (public, custom, shared or marketplace), an instance type, a system disk and optional data disks, one or more elastic network interfaces (ENIs), and membership of at least one security group.

Region (for example cn-hangzhou)API endpoint, images, snapshots, RAM rolesZone AZone BvSwitch A10.0.1.0/24ECS instanceecs.g8y.2xlargeESSD diskssystem + datavSwitch B10.0.2.0/24ECS instancespot, economical stopESSD diskssystem + dataSecurity groupstateful allow rulesMetadata 100.100.100.200token requiredAuto Scaling + SLBspread across zonesattachattachapplies to ENIlaunchesA VPC spans the region; vSwitches and disks are zonal, so spread instances across zones
The ECS building blocks: a VPC spans a region, each vSwitch lives in one zone, instances attach zonal cloud disks and elastic network interfaces, security groups filter traffic, and the metadata service hands out identity and interruption notices.

Identity for software running on the instance comes from an instance RAM role, Alibaba's equivalent of an instance profile. The instance fetches short-lived credentials from the metadata service, so no AccessKey ever needs to be written to disk. If you know how IAM roles work elsewhere, cloud IAM fundamentals applies directly; the Alibaba service is called Resource Access Management (RAM).

Reading an instance type name

Instance type names are compact and regular. Take ecs.g8y.2xlarge: ecs is the prefix, g is the family, 8 is the generation (higher is newer and usually better value), y is a processor suffix, and 2xlarge is the size. Sizes scale vCPUs: large is 2 vCPUs, xlarge is 4 and 2xlarge is 8.

LetterMeaningTypical use
gGeneral purpose, about 1 vCPU to 4 GiBWeb and application servers, most services
cCompute optimised, about 1:2Batch compute, CPU inference, encoding
rMemory optimised, about 1:8Caches, in-memory databases, JVM heaps
i suffixIntel processorx86 software with Intel-specific tuning
a suffixAMD processorx86 at lower price per core
y suffixYitian 710, Alibaba's Arm processorArm-ready containers and services

Other families exist for GPUs, local NVMe storage, burstable credit-based workloads and bare metal, and their exact names change with each generation, so query the regional catalogue rather than hard-coding a list. Two practical rules: Arm families like g8y need Arm images and multi-architecture container builds, and not every family is available in every zone, so check availability before you design around one.

Advertisement

Four ways to pay, and economical mode

ECS has four commercial models, and choosing well matters more than squeezing instance sizes.

  • Subscription (InstanceChargeType=PrePaid): pay up front for months or years at a discount. Right for steady baseline load. You cannot simply delete it early without refund rules applying.
  • Pay-as-you-go (PostPaid): billed by usage with no commitment. Right for variable or experimental load. Pair with AutoReleaseTime for temporary instances so forgotten test boxes delete themselves.
  • Spot (preemptible): pay-as-you-go at a market price, often deeply discounted, but the instance can be reclaimed when capacity is short or the price exceeds your limit.
  • Savings plans and reserved instances: commitment discounts applied to pay-as-you-go usage, keeping on-demand flexibility while lowering the rate for a committed spend.

Pay-as-you-go instances in a VPC can also be stopped in economical mode (StoppedMode=StopCharging on StopInstance). vCPUs, memory, GPUs and the system-assigned public IP are released and stop billing, while disks, snapshots and the private IP are kept. The catch: starting again is a fresh capacity request, so it can fail if the zone is out of stock. Do not use economical mode for anything that must restart on demand. For a cross-cloud view of commitments and waste, see cloud cost analysis.

Launching instances repeatably

Clicking through the console is fine once. For anything repeatable, use the RunInstances API through the aliyun CLI, Terraform or an SDK. RunInstances creates and starts up to 100 instances per call, and DryRun validates parameters, quota and permissions without creating anything:

aliyun ecs RunInstances \
  --RegionId "$REGION" \
  --VSwitchId "$VSWITCH_ID" \
  --SecurityGroupId "$SG_ID" \
  --ImageId "$IMAGE_ID" \
  --InstanceType ecs.c8y.2xlarge \
  --InstanceChargeType PostPaid \
  --SystemDisk.Category cloud_essd \
  --SystemDisk.Size 40 \
  --RamRoleName batch-worker-role \
  --HttpTokens required \
  --Amount 4 \
  --DryRun true

Two flags here are security defaults worth copying. RamRoleName gives the instances identity without keys. HttpTokens required forces the metadata service into hardened mode, where every request needs a session token obtained by a PUT. That blocks the classic server-side request forgery attack in which a vulnerable web app is tricked into fetching role credentials with a simple GET. Once the dry run passes, remove DryRun and capture the returned instance ids. For fleets, put the same parameters in a launch template and let Auto Scaling use it.

Spot instances and the five-minute notice

Spot instances are where the savings are and where most surprises happen. Three RunInstances parameters control them. SpotStrategy is SpotAsPriceGo (pay the market price, capped at the pay-as-you-go price) or SpotWithPriceLimit with a SpotPriceLimit; NoSpot means an ordinary instance. SpotDuration=1 buys a one-hour protection period after creation, during which the instance is not reclaimed automatically; 0 gives no guarantee. SpotInterruptionBehavior is Terminate (release) or Stop (economical-mode stop, keeping disks).

When an instance is about to be reclaimed, ECS gives about five minutes of notice: it emits an Instance:PreemptibleInstanceInterruption event through CloudMonitor and updates the instance metadata. A small watcher on each instance turns that into a graceful drain:

#!/bin/bash
# /usr/local/bin/spot-watch.sh, run as a systemd service
MD=http://100.100.100.200/latest
while true; do
  TOKEN=$(curl -s -X PUT "$MD/api/token" -H "X-aliyun-ecs-metadata-token-ttl-seconds: 300")
  CODE=$(curl -s -o /tmp/spot-tt -w '%{http_code}' \
         -H "X-aliyun-ecs-metadata-token: $TOKEN" \
         "$MD/meta-data/instance/spot/termination-time")
  if [ "$CODE" = "200" ]; then          # 404 means no reclaim scheduled
    logger "spot reclaim scheduled at $(cat /tmp/spot-tt)"
    systemctl stop batch-worker          # worker checkpoints and releases its lease
    break
  fi
  sleep 5
done

Five minutes is enough to finish a short task or checkpoint a long one, not to run an hour-long job to completion. Design work in idempotent chunks with leases in a queue, so a reclaimed worker's chunk simply becomes visible again. Spread the fleet across several instance types and zones so one capacity pool's shortage does not take everything; spot capacity strategies covers the general patterns.

Storage: cloud disks, snapshots and local disks

ECS block storage is called cloud disks, and they live in one zone. Enterprise SSDs (ESSD) are the default for production and come in several performance levels, with throughput and IOPS scaling with level and size; ESSD AutoPL decouples performance from capacity. Consult the current table for figures, because they are revised and vary by instance type, which also caps disk bandwidth. Older categories such as ultra disks remain for legacy use.

Three habits prevent data loss. Set data disks so they are not deleted with the instance when they hold state, and set them to be deleted when they do not, so orphans do not accumulate cost. Take snapshots on a schedule with an automatic snapshot policy, and copy critical snapshots to another region. Remember that instance families with local NVMe disks lose that data when the instance is released or migrated, so use them only for caches, scratch space or replicated stores.

Networking, security groups and fleet operations

Security groups are stateful virtual firewalls on ENIs: allow inbound SSH from a bastion range, and the reply traffic is permitted automatically. Basic security groups allow intra-group traffic by default unless you change it, while advanced security groups do not; know which kind you created. Prefer referencing other security groups in rules rather than CIDR ranges, keep management ports closed to the internet, and use Cloud Assistant to run commands on instances instead of opening SSH broadly.

Public access should come through a load balancer or NAT gateway rather than per-instance public IPs. An elastic IP (EIP) is a separately managed address you can move between instances. For fleets, put instances behind Server Load Balancer, let Auto Scaling replace unhealthy ones, and use a deployment set with a high-availability strategy to spread instances across physical hosts so a single host failure cannot take out a whole tier. See autoscaling patterns for scaling policies.

Worked example: a nightly embedding fleet on spot

A team needs to embed 40 million documents with a CPU-friendly embedding model each night, within six hours, at minimum cost. They measure 900 documents per minute on one ecs.c8y.2xlarge with a container image built for Arm. 40,000,000 divided by 900 divided by 360 minutes is about 124, so they need roughly 125 instances running for the whole six hours, or fewer, larger instances with the same total vCPUs.

They build an Auto Scaling group from a launch template with SpotAsPriceGo, SpotInterruptionBehavior=Terminate, a RAM role that can read the input bucket and write results, and hardened metadata. The template lists three Arm and x86 compute families across three zones; the image is multi-architecture. Work is split into 5,000-document chunks leased from a queue for ten minutes; the spot watcher stops the worker, which abandons its lease. A small subscription instance runs the coordinator. On a typical night about 4 percent of instances are reclaimed, which costs only the in-flight chunks, and the run finishes inside the window at a fraction of the pay-as-you-go cost. The single biggest saving was not spot itself but chunking: before it, one reclaim restarted an hour of work.

Failure modes and trade-offs

  • Zone stock-outs. A pinned family or zone has no capacity; launch across several types and zones.
  • Economical-mode restart failure. A stopped instance cannot start because its compute was released; keep critical instances running or use subscription.
  • Keys on disk. AccessKeys in config files leak through images and snapshots; use instance RAM roles.
  • Open metadata. Without HttpTokens required, an SSRF bug can read role credentials.
  • Orphaned disks and EIPs. Released instances leave billable resources; tag everything and sweep weekly.
  • Arm surprises. An x86-only image on g8y fails to boot or run; build multi-architecture images.

The main trade-off is between price and control. Subscription is cheapest per hour for steady load but locks capacity; spot is cheapest overall but demands interruptible design; pay-as-you-go buys flexibility at list price. Higher ESSD levels and larger families buy performance you should prove you need with measurements. And if your workload is event-driven and short, a managed service such as Alibaba Function Compute may remove the fleet entirely.

What to do next

  1. Map your region, zones and VPC: one vSwitch per zone you will use, with non-overlapping CIDR ranges.
  2. Pick instance families from measured CPU and memory ratios, and check zone availability for each before committing.
  3. Create RAM roles for instances and launch with RamRoleName and HttpTokens required; delete any AccessKeys on hosts.
  4. Encode launch settings in a launch template or Terraform, and validate changes with DryRun.
  5. For batch work, use spot across several families and zones, chunk work with leases, and install the termination watcher.
  6. Set disk deletion flags deliberately, attach an automatic snapshot policy, and copy critical snapshots across regions.
  7. Tag every resource with owner and environment, then review subscription, savings plan and spot mix monthly against actual usage.
Key takeaway: Alibaba ECS is a virtual machine service built from regions, zonal vSwitches, instance types, cloud disks, ENIs and security groups. Read instance names as family, generation, processor and size, and choose billing deliberately among subscription, pay-as-you-go, spot and commitment discounts, remembering that economical-mode restarts can fail. Launch from templates with RAM roles and hardened metadata, design spot work in leased chunks with a termination watcher, and spread fleets across zones and types.