Azure Kubernetes Service (AKS) gives you a Kubernetes cluster whose control plane Microsoft runs and whose worker nodes are virtual machine scale sets in your subscription. Creating one takes a single command. Running one well takes an understanding of the parts the command hides: which Azure resources it creates and who may touch them, why there are two kinds of node pool, what the pricing tier buys, which network model you are committing to, and how upgrades move through your nodes.
This article covers those AKS-specific mechanics. It assumes you know Kubernetes itself. For how AKS compares with EKS and GKE, pod address planning and a full workload identity walkthrough, see Managed Kubernetes across clouds; this page goes deeper on AKS alone and finishes with an upgrade worked example you can adapt to your own pools.
The resource model
When you create a cluster, three things appear. The control plane (API server, etcd, scheduler, controller manager) runs in a Microsoft-managed subscription; you reach it through its endpoint and never see its VMs. The managedCluster resource lands in the resource group you named; it is the object you configure with az aks, Bicep or Terraform. And a second resource group, the node resource group, named MC_<resource-group>_<cluster>_<region> by default, holds the scale sets, load balancer, public IPs, NICs and managed disks that AKS creates for you.
The rule that saves the most incidents: treat the node resource group as AKS-owned. If you resize a scale set, delete a load balancer rule or edit an NSG there by hand, the next reconcile or upgrade may revert it or fail on it. Change nodes through node pool commands, and load balancers through Kubernetes Service objects and annotations. If you need your own route tables, NSGs or private endpoints, put the cluster in a VNet you created, where you are supported to manage them.
System and user node pools
Each node pool is one scale set with one VM size, one OS and one Kubernetes version. Every cluster needs at least one system pool, which hosts critical add-ons such as CoreDNS, metrics-server and the tunnel to the control plane. User pools run your workloads. Keep them apart: a noisy application that starves CoreDNS takes down name resolution for the whole cluster.
# a small, tainted system pool keeps application pods off the critical add-ons
az aks nodepool add -g rg-prod --cluster-name aks-prod -n sys2 \
--mode System --node-count 3 --node-vm-size Standard_D4ds_v5 \
--zones 1 2 3 --node-taints CriticalAddonsOnly=true:NoSchedule
# a user pool for the API tier, autoscaled and spread across zones
az aks nodepool add -g rg-prod --cluster-name aks-prod -n api \
--mode User --node-vm-size Standard_D8ds_v5 --zones 1 2 3 \
--enable-cluster-autoscaler --min-count 3 --max-count 30 \
--max-surge 33%Separate pools are also how you mix hardware: a GPU pool with a taint that only training jobs tolerate, a spot pool for interruptible batch work, a Windows pool. Pools upgrade one at a time, so they double as upgrade blast-radius boundaries. For how the cluster autoscaler decides when to add and remove nodes, see cloud autoscaling.
Tiers and what they buy
| Tier | Control plane SLA | Scale | Use it for |
|---|---|---|---|
| Free | None; best effort | Up to 1,000 nodes; recommended under 10 | Development, test, learning |
| Standard | 99.95% with availability zones, 99.9% without | Up to 5,000 nodes | Production |
| Premium | Same as Standard | Up to 5,000 nodes | Production that needs Long Term Support versions |
The tier is set with --tier free|standard|premium and can be changed on a running cluster. Premium must be paired with --k8s-support-plan AKSLongTermSupport, which extends support for a Kubernetes version well beyond the community window. The SLA covers the API server only; your application's availability comes from zones, replicas and disruption budgets on your nodes. AKS Automatic, the more opinionated mode, uses Standard and preconfigures upgrades, scaling and networking.
Choosing the network model
Pod networking is the decision you can least easily change, so make it deliberately. AKS separates the IP address management option from the dataplane.
| Option | Pod IPs come from | Choose it when |
|---|---|---|
| Azure CNI Overlay | A private pod CIDR outside the VNet; egress is SNATed to the node IP | Default choice; conserves VNet space; up to 250 pods per node |
| Azure CNI Pod Subnet | A dedicated subnet in your VNet | Other systems must reach pod IPs directly |
| Azure CNI Node Subnet (legacy) | The node subnet itself | Only for existing clusters |
| kubenet (legacy) | A CIDR with route tables | Never for new clusters; retires March 31, 2028 |
Independently, --network-dataplane cilium replaces iptables-based service routing with eBPF and enforces network policy in the same engine. A typical production create looks like this:
az aks create -g rg-prod -n aks-prod --tier standard --zones 1 2 3 \
--network-plugin azure --network-plugin-mode overlay \
--pod-cidr 10.244.0.0/16 --network-dataplane cilium \
--vnet-subnet-id "$NODE_SUBNET_ID" \
--enable-oidc-issuer --enable-workload-identity \
--auto-upgrade-channel stable --node-os-upgrade-channel NodeImage \
--generate-ssh-keysAvoid the reserved ranges 169.254.0.0/16, 192.0.2.0/24, 172.30.0.0/16 and 172.31.0.0/16 for pod, service and VNet CIDRs; AKS rejects overlaps. With Overlay, pods reach the VNet and on-premises networks, but those networks cannot initiate connections to pod IPs; they come in through a Service or ingress. If you bring your own subnet, the cluster identity needs Network Contributor on it.
Identity in two layers
Humans and pipelines authenticate to the API server with Microsoft Entra ID; enable Azure RBAC for Kubernetes authorization so access is granted with Azure role assignments instead of per-cluster kubeconfig secrets. Pods authenticate to Azure with workload identity: the cluster's OIDC issuer signs a service account token, and a federated credential on a user-assigned managed identity trusts tokens whose subject is system:serviceaccount:<namespace>:<name>. The service account carries the azure.workload.identity/client-id annotation and the pod the azure.workload.identity/use: "true" label. Grant the managed identity only the roles that one workload needs. The Entra side of this is covered in Microsoft Entra ID, in depth.
The upgrade machinery
Two different things get upgraded, on two schedules. The Kubernetes version changes the control plane and the kubelet; minor versions cannot be skipped on a non-LTS cluster, and the control plane must always be at least as new as every pool. The node image is the OS disk the VMs boot from, refreshed by AKS with security patches and fixes. Each has its own auto-upgrade channel.
| Setting | Values | Meaning |
|---|---|---|
--auto-upgrade-channel | none, patch, stable, rapid | patch tracks the latest patch of your minor; stable the latest patch of minor N-1; rapid the latest minor |
--node-os-upgrade-channel | None, Unmanaged, SecurityPatch, NodeImage | NodeImage swaps in a new weekly image; SecurityPatch applies tested security fixes, reimaging only when needed |
Schedule each with a planned maintenance configuration: aksManagedAutoUpgradeSchedule for version upgrades and aksManagedNodeOSUpgradeSchedule for node images, each at least four hours long. The older node-image value of the cluster channel is legacy; use the node OS channel instead.
Every node upgrade is a rolling replacement. For each batch, AKS adds surge nodes on the new version, cordons and drains the same number of old nodes while respecting PodDisruptionBudgets, optionally waits a soak period, reimages the drained nodes, and reuses them as the next batch's buffer. When the pool is done, the extra nodes are removed. The knobs per pool are --max-surge (default one node; 33% is the documented production recommendation), --max-unavailable (default 0; drains without adding nodes, for when quota is tight), --drain-timeout (default 30 minutes) and --node-soak-duration (0 to 30 minutes).
Worked example: upgrading a 30-node pool
Take the api pool at 30 nodes with --max-surge 33%. AKS rounds up, so it adds ceil(30 x 0.33) = 10 surge nodes and works in three batches of ten. Check three things before you start.
- Quota. You need 10 extra VMs of this size in the region's vCPU quota, for the whole duration. With the default surge of one, you would need one extra VM, but the pool would take 30 serial batches instead of three.
- Addresses. With Overlay, the 10 new nodes need only node subnet IPs. With pod subnet, each surge node also reserves pod IPs; ten nodes at 30 pods each is 300 more addresses that must be free.
- Disruption budgets. Suppose the checkout Deployment runs 6 replicas with a PDB of
minAvailable: 6. No eviction is ever allowed, the drain of the first node blocks, and after 30 minutes the drain timeout stops the upgrade partway, with surge nodes added and old nodes cordoned. SetmaxUnavailable: 1orminAvailable: 5instead, and spread replicas with topology spread constraints so a batch of ten nodes never holds most of them.
To estimate duration, time one batch on a staging pool: provisioning, draining, and reimaging together often take minutes per batch, so measure rather than guess. Multiply by the batch count and compare with your maintenance window. A stopped upgrade resumes on the next update operation, but a pool stuck mid-upgrade, with cordoned nodes and extra VMs, is the state to avoid during peak traffic.
# before: what will this upgrade touch, and is anything blocking?
az aks get-upgrades -g rg-prod -n aks-prod -o table
kubectl get pdb -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,ALLOWED:.status.disruptionsAllowed
# during: watch surge, drain and soak events
kubectl get events -A --sort-by=.lastTimestamp | grep -E 'Drain|Surge|Upgrade'Any PDB showing zero allowed disruptions in the second command will stall a drain. Fix it before the window opens.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Upgrade stops with a drain failure | PDB allows zero disruptions, or a pod ignores SIGTERM | Loosen the PDB; set terminationGracePeriodSeconds; raise drain timeout only for genuinely long pods |
| Upgrade fails before starting | No vCPU quota or subnet IPs for surge nodes | Request quota, lower max surge, or use max unavailable |
| Manual load balancer or NSG change disappears | Edited inside the node resource group | Use Service annotations or your own VNet resources |
| Cluster forced onto a newer version | Fell out of the support window | Run the stable channel with a maintenance window |
| Pods cannot be reached from on-premises | Overlay pod IPs are not routable | Expose through a Service or ingress, or use pod subnet |
| DNS failures under load | Application pods crowding the system pool | Taint the system pool and size it for add-ons |
Trade-offs
AKS removes control plane operations but not cluster operations: upgrades, capacity, disruption budgets and network design remain yours. Overlay is simpler and scales further, at the cost of pods not being directly addressable. Faster upgrades with high surge cost quota and more simultaneous disruption. Long Term Support buys time between upgrades but lets version drift accumulate.
Cost follows the same split. The Standard and Premium tiers add a per-cluster management charge, while nodes, disks, load balancers and egress are billed as ordinary Azure resources in the node resource group, including surge nodes for as long as an upgrade runs. Many small clusters multiply the management charge and the system pools; one large shared cluster concentrates blast radius and upgrade risk. Most teams settle on one cluster per environment and region, with namespaces and node pools separating workloads inside it. Define the cluster in code so these choices are reviewable; Bicep and ARM covers the native option.
What to do next
- List the resources in your node resource group and confirm nobody manages them by hand.
- Give the cluster a tainted system pool across three zones and move workloads to user pools.
- Pick the tier from the table: Standard for anything with users, Premium only if you need LTS.
- Confirm your network model; plan the move off kubenet before March 31, 2028 if you still use it.
- Set the stable and NodeImage channels with two maintenance windows of at least four hours.
- Set max surge to 33% on production pools, check quota and subnet IPs, and audit every PDB for zero allowed disruptions.
- Rehearse a full upgrade on a staging cluster and record how long each batch takes.