Alibaba Cloud Container Service for Kubernetes (ACK) is Alibaba Cloud's managed Kubernetes. It runs conformant upstream Kubernetes, so manifests, Helm charts and kubectl habits carry over. What does not carry over is everything around the API server: the cluster type, how pods get IP addresses, how nodes appear and disappear, and how a pod proves its identity to cloud APIs. Those choices are hard to change later and cause most production incidents.

This article explains ACK from the cluster outward. It covers the cluster types and when each fits, the network plugin decision, node pools and the two node autoscalers, virtual nodes backed by Elastic Container Instance (ECI), pod identity with RRSA, storage, the version policy, a worked example, failure modes and a closing checklist. If you know EKS, GKE or AKS, the managed Kubernetes comparison gives you the shared vocabulary; this page concentrates on what is specific to Alibaba Cloud.

Which ACK are you creating

ACK is a family of offerings, and the first decision is which one you are creating. The core product is the ACK managed cluster, where Alibaba Cloud runs the control plane (API server, etcd, scheduler, controller manager) and you run worker nodes. Managed clusters come in two tiers. Basic is meant for testing and small workloads. Pro is the production tier, with higher control-plane reliability and a service level agreement. If the cluster serves customers, choose Pro; the Basic tier exists so experiments are cheap, not so production can be.

The older ACK dedicated cluster, where you also ran the master nodes and etcd yourself, can no longer be created since August 21, 2024, except in CloudBox scenarios. If you inherit one, plan a migration to a managed cluster rather than investing in it.

OfferingWhat you manageFits
ACK managed, ProWorker nodes, add-ons, workloadsProduction services, most teams
ACK managed, BasicSame, without the Pro SLATests, labs, short-lived clusters
ACK ServerlessWorkloads only; billed on pod CPU and memory requestsBursty or spiky jobs, no node fleet to run
ACK EdgeCloud control plane, nodes at edge sitesStores, factories, on-premises boxes
ACK OneA fleet of clusters across regions or cloudsMulti-cluster management
ACK LingjunClusters on Lingjun AI computing resourcesLarge-scale AI training

Architecture of a managed cluster

An ACK managed cluster inside your VPCManaged control planeAPI server, etcd, scheduler (run by Alibaba Cloud)Node pool: onlineECS, pay-as-you-go, zones A+Bkubelet + Terway podspod IPs from pod vSwitchesNode pool: batchspot ECS, autoscaled 0..Ntaint: workload=batchscaled by ONE node autoscalerVirtual nodeack-virtual-nodepods labelled alibabacloud.com/ecirun as ECI instances, no ECSRRSApod token -> STS AssumeRoleWithOIDCOSS / NAS / disksCSI drivers; disks are zonalSLS + Prometheuslogs and metricsVPC: node vSwitches + pod vSwitches per zonesecurity groups, NAT gateway for egress
The control plane is managed; everything below it (node pools, virtual nodes, networking, identity wiring) is yours to design.

The rest of this article assumes a managed Pro cluster. It lives in one region and one VPC; the control plane runs in Alibaba Cloud's account and exposes an API server endpoint that can be private, public or both. Keep it private: a public API endpoint is an attack surface you rarely need.

Worker capacity comes from node pools. A node pool is a group of ECS instances that share a specification: instance types, vSwitches (and therefore zones), operating system image, container runtime, security groups, labels and taints. Define one pool per workload shape, in code, never hand-edited. The ECS deep dive explains instance families, spot interruption and billing, all of which apply directly to node pools.

Add-ons such as CoreDNS, the CSI drivers and the log and metrics agents are managed as ACK components. Upgrading the control plane does not upgrade nodes or components; those are separate steps.

Networking: Terway or Flannel

ACK offers two network plugins, and the choice is made at cluster creation. Changing it later means building a new cluster.

  • Flannel gives pods addresses from a separate pod CIDR that does not exist in the VPC. Each node gets a slice of that CIDR, and ACK writes a VPC route per node so pod traffic reaches the right host. It is simple and conserves VPC addresses, but the VPC route table has an entry quota, which effectively caps the node count, and pods are not first-class VPC citizens.
  • Terway is Alibaba Cloud's plugin. Pods get real VPC addresses from dedicated pod vSwitches, attached through elastic network interfaces (ENIs) on the node. Pods can then be reached directly from other VPC resources, can be referenced in security-group and whitelist rules, and support Kubernetes NetworkPolicy. The cost is address planning: every pod consumes a VPC IP, and the number of ENIs and secondary IPs per instance type limits how many pods a node can run.

For new production clusters, Terway is the usual choice because databases and on-premises networks see pod addresses natively. Create pod vSwitches in every zone, size them for peak pods plus rolling-update surge, and keep them separate from node vSwitches. The classic trap: the cluster works for months, then a scale-out fails because a pod vSwitch ran out of addresses.

Also decide the Service CIDR at creation. It must not overlap the VPC, any peered VPC, or on-premises ranges you will ever connect, and it cannot be changed later.

Node pools, autoscaling and virtual nodes

Pod autoscaling is upstream Kubernetes: the Horizontal Pod Autoscaler changes replica counts from metrics. Node autoscaling is where ACK adds its own machinery, and it offers two mutually exclusive options. Node auto scaling runs the familiar cluster-autoscaler, which periodically scans for pending pods and grows the node pools that could fit them. Node instant scaling uses an event-driven autoscaler that ACK recommends for large clusters, many autoscaled node pools, multi-zone and multi-instance-type pools, and workloads using topology spread constraints. Only one of them can run in a cluster. Test scale-out under load before depending on either.

Either autoscaler only works with honest pod requests: a pod requesting 100m CPU while using three cores never triggers a scale-out, it starves its neighbours. Set requests from measured usage, taint batch pools, and give every service a PodDisruptionBudget so scale-in cannot take it below its minimum.

Virtual nodes are the third source of capacity. Install the ack-virtual-node component and the cluster gains a node that is not a machine: pods scheduled to it run as ECI instances, billed per pod, with no ECS to manage. Opt pods in with a label, either on the pod or on a whole namespace:

apiVersion: v1
kind: Namespace
metadata:
  name: burst-jobs
  labels:
    alibabacloud.com/eci: "true"     # every pod here runs on a virtual node as ECI

ECI suits spiky jobs that would otherwise keep idle nodes around. But DaemonSets do not run on virtual nodes, so node-level log and security agents do not see those pods; read the virtual-node limitations before moving a workload.

Pod identity with RRSA

Pods need credentials to call OSS, SLS or other Alibaba Cloud APIs. An AccessKey pair in a Secret never expires and leaks easily; the node's RAM role is shared by every pod on the node. ACK's answer is RRSA (RAM Roles for Service Accounts), which gives each Kubernetes service account its own RAM role through OIDC federation.

The flow works like this. You enable RRSA OIDC on the cluster, which creates an OIDC provider for it. The kubelet projects a short-lived service-account token into the pod. The pod's SDK exchanges that token with STS (the AssumeRoleWithOIDC call) for temporary credentials of a RAM role. The RAM role's trust policy accepts only tokens from your cluster's issuer, for the intended audience, and for one specific service account. Install the ack-pod-identity-webhook component and you do not have to wire the token mount yourself: label the namespace, annotate the service account, and the webhook injects the mount and three environment variables, ALIBABA_CLOUD_ROLE_ARN, ALIBABA_CLOUD_OIDC_PROVIDER_ARN and ALIBABA_CLOUD_OIDC_TOKEN_FILE.

apiVersion: v1
kind: Namespace
metadata:
  name: search
  labels:
    pod-identity.alibabacloud.com/injection: "on"
---
apiVersion: v1
kind: ServiceAccount
metadata:
  name: indexer
  namespace: search
  annotations:
    pod-identity.alibabacloud.com/role-name: search-indexer-oss

The RAM role trusts exactly that service account:

{
  "Statement": [{
    "Action": "sts:AssumeRole",
    "Effect": "Allow",
    "Principal": {"Federated": ["<oidc_provider_arn>"]},
    "Condition": {"StringEquals": {
      "oidc:aud": "sts.aliyuncs.com",
      "oidc:iss": "<oidc_issuer_url>",
      "oidc:sub": "system:serviceaccount:search:indexer"
    }}
  }],
  "Version": "1"
}

Attach a narrow permission policy to the role, for example read and write on one OSS bucket prefix, and nothing else. Check that your SDK version reads the OIDC variables in its default credential chain; an old SDK falls back to the node role or fails. The token file lives at /var/run/secrets/ack.alibabacloud.com/rrsa-tokens/token, which is the first thing to inspect when a pod gets access-denied errors. For bucket design and the STS side, see the OSS deep dive.

Storage and the zonal disk trap

ACK ships CSI drivers for cloud disks, File Storage NAS, OSS and CPFS. The property that bites most teams is that cloud disks are zonal. A volume on a disk in zone A attaches only to nodes in zone A, so a rescheduled StatefulSet pod stays Pending until zone A has room. Two rules help. First, use a StorageClass with volumeBindingMode WaitForFirstConsumer, so the disk is created in the zone where the scheduler actually placed the pod instead of a random one. Second, keep at least one node pool per zone that hosts stateful pods.

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: essd-wffc
provisioner: diskplugin.csi.alibabacloud.com
parameters:
  type: cloud_essd
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
reclaimPolicy: Retain

NAS is shared across zones, which suits files many pods read. OSS mounted as a volume is convenient for model weights, but it is object storage: renames are copies and small random writes are slow, so do not treat it as a POSIX disk.

Versions and upgrades

ACK lets you create clusters only on the three most recent Kubernetes minor versions it supports. Upgrades move one minor version at a time; skipping versions and downgrading are not possible. From 1.31 onward ACK supports every minor version rather than only even-numbered ones. Treat this as a calendar commitment: roughly every few months, one upgrade of the control plane, then node pools, then components.

Rehearse each upgrade on staging: scan manifests for removed APIs, upgrade the control plane, roll node pools with surge capacity so PodDisruptionBudgets are honoured, then upgrade components such as the CSI drivers and Terway. A cluster that skips two cycles faces forced back-to-back upgrades with no time to test between them.

Worked example: a retrieval platform

Consider a team running a retrieval service on ACK: an online search API, a nightly re-embedding job over 30 million documents stored in OSS, and occasional emergency re-indexes when a model changes. Here is a design that follows the guidance above.

  1. An ACK managed Pro cluster with a private API endpoint, Terway networking, node and pod vSwitches in two zones, and a Service CIDR that does not overlap the VPC or the office network.
  2. A small system pool for add-ons, and an online pool across both zones autoscaled between 4 and 12 nodes, with HPA on the API and a PodDisruptionBudget of minAvailable 3.
  3. A batch pool on spot instances, tainted workload=batch, autoscaled from 0. The nightly job tolerates the taint, checkpoints progress to OSS every 10,000 documents and resumes from the last checkpoint after a spot reclaim.
  4. The emergency re-index runs as a Job in a namespace labelled for ECI, so 40 parallel workers start without waiting for nodes and stop billing when they finish.
  5. The indexer service account uses RRSA with a role limited to the corpus bucket prefix. No AccessKey exists anywhere in the cluster.
  6. SLS for logs and Managed Service for Prometheus for metrics, alerting on long-pending pods, free pod vSwitch addresses and autoscaler failures.

Each choice guards one failure: the taint keeps online pods off spot capacity, checkpoints turn a spot reclaim into lost minutes, the vSwitch alert fires weeks before address exhaustion, and ECI keeps the emergency path independent of the batch autoscaler.

Failure modes

These are the failures teams hit most often on ACK, with their usual root cause.

SymptomLikely causeFix
Pods Pending, events mention IP allocationPod vSwitch exhausted (Terway)Add pod vSwitches; alert on free IPs
StatefulSet pod Pending after a rescheduleZonal disk, no capacity in that zoneWaitForFirstConsumer; a node pool per zone
Scale-out never happensRequests too low, or wrong pool selectorsMeasure and set requests; check pool labels and taints
Access denied from a podMissing namespace label, wrong annotation, trust policy sub mismatch, old SDKInspect injected env vars and the token file first
Log agent misses some podsPods on virtual nodes; DaemonSets do not run thereUse the ECI log integration for those pods

Trade-offs

ACK's strengths are the depth of its integration with Alibaba Cloud (Terway pod addresses in the VPC, RRSA, ECI virtual nodes, SLS and the CSI drivers) and its presence in regions where the other hyperscalers are thin, mainland China above all. The costs are the usual managed-Kubernetes costs plus a few specific ones. You commit to Alibaba Cloud's network model and identity system, which makes multi-cloud portability a conscious engineering effort. The upgrade cadence is not optional. And ECI convenience comes with real behavioural differences from ordinary nodes.

On Alibaba Cloud itself, Function Compute avoids a cluster entirely for a handful of event-driven functions, and ACK Serverless gives Kubernetes semantics without a node fleet. If you also run AKS, the AKS deep dive shows that the shape of the decisions (network plugin, node pools, workload identity, upgrades) is almost identical.

What to do next

  1. Decide the cluster type: managed Pro for anything customer-facing; note any inherited dedicated clusters for migration.
  2. Write down the network plan before creating anything: Terway or Flannel, node and pod vSwitches per zone with sizes, and a Service CIDR that overlaps nothing.
  3. Define node pools as code: a system pool, online pools across at least two zones, a tainted spot pool for batch.
  4. Choose one node autoscaler, set pod requests from measured usage, and add a PodDisruptionBudget to every service.
  5. Enable RRSA, install ack-pod-identity-webhook, and replace every AccessKey Secret with a per-service-account RAM role.
  6. Create a WaitForFirstConsumer disk StorageClass and make it the default for stateful workloads.
  7. Pilot ECI for one bursty job and confirm that logging and monitoring still see its pods.
  8. Put the next Kubernetes minor upgrade on the calendar and rehearse it on staging.
Key takeaway: ACK runs upstream Kubernetes with a managed control plane, but production success depends on the choices around it: managed Pro for real workloads, Terway with pod vSwitches sized per zone, node pools defined as code, exactly one node autoscaler fed by honest requests, ECI virtual nodes for bursts, RRSA instead of AccessKeys, WaitForFirstConsumer for zonal disks, and an upgrade every cycle.