Azure Kubernetes Service (AKS) gives you a Kubernetes cluster whose control plane Microsoft runs and whose worker nodes are virtual machine scale sets in your subscription. Creating one takes a single command. Running one well takes an understanding of the parts the command hides: which Azure resources it creates and who may touch them, why there are two kinds of node pool, what the pricing tier buys, which network model you are committing to, and how upgrades move through your nodes.

This article covers those AKS-specific mechanics. It assumes you know Kubernetes itself. For how AKS compares with EKS and GKE, pod address planning and a full workload identity walkthrough, see Managed Kubernetes across clouds; this page goes deeper on AKS alone and finishes with an upgrade worked example you can adapt to your own pools.

Advertisement

The resource model

Who owns what in an AKS clusterMicrosoft-managed subscription (you never see these resources)API serverplus etcdSchedulercontroller managerCloud controllerLB, disks, routesTier and SLAFree, Standard, PremiumYour resource groupthe managedCluster resource, your VNet, ACR, Key VaultNode resource group (MC_...)VMSS per pool, load balancer, NICs, disks, public IPsyou configureAKS managesSystem poolCoreDNS, metrics, tunnelUser poolsyour workloads, GPUs, spotPod networkCNI Overlay pod CIDR, or pod subnet in your VNetIdentityEntra ID for kubectl; workload identity for podsUpgradesKubernetes version and node image, separate channelsEdit what lives in your resource group; change the node resource group only through the AKS API.
AKS splits ownership three ways: Microsoft runs the control plane, you own the cluster resource and your network, and AKS owns the node resource group on your behalf.

When you create a cluster, three things appear. The control plane (API server, etcd, scheduler, controller manager) runs in a Microsoft-managed subscription; you reach it through its endpoint and never see its VMs. The managedCluster resource lands in the resource group you named; it is the object you configure with az aks, Bicep or Terraform. And a second resource group, the node resource group, named MC_<resource-group>_<cluster>_<region> by default, holds the scale sets, load balancer, public IPs, NICs and managed disks that AKS creates for you.

The rule that saves the most incidents: treat the node resource group as AKS-owned. If you resize a scale set, delete a load balancer rule or edit an NSG there by hand, the next reconcile or upgrade may revert it or fail on it. Change nodes through node pool commands, and load balancers through Kubernetes Service objects and annotations. If you need your own route tables, NSGs or private endpoints, put the cluster in a VNet you created, where you are supported to manage them.

System and user node pools

Each node pool is one scale set with one VM size, one OS and one Kubernetes version. Every cluster needs at least one system pool, which hosts critical add-ons such as CoreDNS, metrics-server and the tunnel to the control plane. User pools run your workloads. Keep them apart: a noisy application that starves CoreDNS takes down name resolution for the whole cluster.

# a small, tainted system pool keeps application pods off the critical add-ons
az aks nodepool add -g rg-prod --cluster-name aks-prod -n sys2 \
  --mode System --node-count 3 --node-vm-size Standard_D4ds_v5 \
  --zones 1 2 3 --node-taints CriticalAddonsOnly=true:NoSchedule

# a user pool for the API tier, autoscaled and spread across zones
az aks nodepool add -g rg-prod --cluster-name aks-prod -n api \
  --mode User --node-vm-size Standard_D8ds_v5 --zones 1 2 3 \
  --enable-cluster-autoscaler --min-count 3 --max-count 30 \
  --max-surge 33%

Separate pools are also how you mix hardware: a GPU pool with a taint that only training jobs tolerate, a spot pool for interruptible batch work, a Windows pool. Pools upgrade one at a time, so they double as upgrade blast-radius boundaries. For how the cluster autoscaler decides when to add and remove nodes, see cloud autoscaling.

Advertisement

Tiers and what they buy

TierControl plane SLAScaleUse it for
FreeNone; best effortUp to 1,000 nodes; recommended under 10Development, test, learning
Standard99.95% with availability zones, 99.9% withoutUp to 5,000 nodesProduction
PremiumSame as StandardUp to 5,000 nodesProduction that needs Long Term Support versions

The tier is set with --tier free|standard|premium and can be changed on a running cluster. Premium must be paired with --k8s-support-plan AKSLongTermSupport, which extends support for a Kubernetes version well beyond the community window. The SLA covers the API server only; your application's availability comes from zones, replicas and disruption budgets on your nodes. AKS Automatic, the more opinionated mode, uses Standard and preconfigures upgrades, scaling and networking.

Choosing the network model

Pod networking is the decision you can least easily change, so make it deliberately. AKS separates the IP address management option from the dataplane.

OptionPod IPs come fromChoose it when
Azure CNI OverlayA private pod CIDR outside the VNet; egress is SNATed to the node IPDefault choice; conserves VNet space; up to 250 pods per node
Azure CNI Pod SubnetA dedicated subnet in your VNetOther systems must reach pod IPs directly
Azure CNI Node Subnet (legacy)The node subnet itselfOnly for existing clusters
kubenet (legacy)A CIDR with route tablesNever for new clusters; retires March 31, 2028

Independently, --network-dataplane cilium replaces iptables-based service routing with eBPF and enforces network policy in the same engine. A typical production create looks like this:

az aks create -g rg-prod -n aks-prod --tier standard --zones 1 2 3 \
  --network-plugin azure --network-plugin-mode overlay \
  --pod-cidr 10.244.0.0/16 --network-dataplane cilium \
  --vnet-subnet-id "$NODE_SUBNET_ID" \
  --enable-oidc-issuer --enable-workload-identity \
  --auto-upgrade-channel stable --node-os-upgrade-channel NodeImage \
  --generate-ssh-keys

Avoid the reserved ranges 169.254.0.0/16, 192.0.2.0/24, 172.30.0.0/16 and 172.31.0.0/16 for pod, service and VNet CIDRs; AKS rejects overlaps. With Overlay, pods reach the VNet and on-premises networks, but those networks cannot initiate connections to pod IPs; they come in through a Service or ingress. If you bring your own subnet, the cluster identity needs Network Contributor on it.

Identity in two layers

Humans and pipelines authenticate to the API server with Microsoft Entra ID; enable Azure RBAC for Kubernetes authorization so access is granted with Azure role assignments instead of per-cluster kubeconfig secrets. Pods authenticate to Azure with workload identity: the cluster's OIDC issuer signs a service account token, and a federated credential on a user-assigned managed identity trusts tokens whose subject is system:serviceaccount:<namespace>:<name>. The service account carries the azure.workload.identity/client-id annotation and the pod the azure.workload.identity/use: "true" label. Grant the managed identity only the roles that one workload needs. The Entra side of this is covered in Microsoft Entra ID, in depth.

The upgrade machinery

Two different things get upgraded, on two schedules. The Kubernetes version changes the control plane and the kubelet; minor versions cannot be skipped on a non-LTS cluster, and the control plane must always be at least as new as every pool. The node image is the OS disk the VMs boot from, refreshed by AKS with security patches and fixes. Each has its own auto-upgrade channel.

SettingValuesMeaning
--auto-upgrade-channelnone, patch, stable, rapidpatch tracks the latest patch of your minor; stable the latest patch of minor N-1; rapid the latest minor
--node-os-upgrade-channelNone, Unmanaged, SecurityPatch, NodeImageNodeImage swaps in a new weekly image; SecurityPatch applies tested security fixes, reimaging only when needed

Schedule each with a planned maintenance configuration: aksManagedAutoUpgradeSchedule for version upgrades and aksManagedNodeOSUpgradeSchedule for node images, each at least four hours long. The older node-image value of the cluster channel is legacy; use the node OS channel instead.

Every node upgrade is a rolling replacement. For each batch, AKS adds surge nodes on the new version, cordons and drains the same number of old nodes while respecting PodDisruptionBudgets, optionally waits a soak period, reimages the drained nodes, and reuses them as the next batch's buffer. When the pool is done, the extra nodes are removed. The knobs per pool are --max-surge (default one node; 33% is the documented production recommendation), --max-unavailable (default 0; drains without adding nodes, for when quota is tight), --drain-timeout (default 30 minutes) and --node-soak-duration (0 to 30 minutes).

Worked example: upgrading a 30-node pool

Take the api pool at 30 nodes with --max-surge 33%. AKS rounds up, so it adds ceil(30 x 0.33) = 10 surge nodes and works in three batches of ten. Check three things before you start.

  1. Quota. You need 10 extra VMs of this size in the region's vCPU quota, for the whole duration. With the default surge of one, you would need one extra VM, but the pool would take 30 serial batches instead of three.
  2. Addresses. With Overlay, the 10 new nodes need only node subnet IPs. With pod subnet, each surge node also reserves pod IPs; ten nodes at 30 pods each is 300 more addresses that must be free.
  3. Disruption budgets. Suppose the checkout Deployment runs 6 replicas with a PDB of minAvailable: 6. No eviction is ever allowed, the drain of the first node blocks, and after 30 minutes the drain timeout stops the upgrade partway, with surge nodes added and old nodes cordoned. Set maxUnavailable: 1 or minAvailable: 5 instead, and spread replicas with topology spread constraints so a batch of ten nodes never holds most of them.

To estimate duration, time one batch on a staging pool: provisioning, draining, and reimaging together often take minutes per batch, so measure rather than guess. Multiply by the batch count and compare with your maintenance window. A stopped upgrade resumes on the next update operation, but a pool stuck mid-upgrade, with cordoned nodes and extra VMs, is the state to avoid during peak traffic.

# before: what will this upgrade touch, and is anything blocking?
az aks get-upgrades -g rg-prod -n aks-prod -o table
kubectl get pdb -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,ALLOWED:.status.disruptionsAllowed
# during: watch surge, drain and soak events
kubectl get events -A --sort-by=.lastTimestamp | grep -E 'Drain|Surge|Upgrade'

Any PDB showing zero allowed disruptions in the second command will stall a drain. Fix it before the window opens.

Failure modes

SymptomLikely causeFix
Upgrade stops with a drain failurePDB allows zero disruptions, or a pod ignores SIGTERMLoosen the PDB; set terminationGracePeriodSeconds; raise drain timeout only for genuinely long pods
Upgrade fails before startingNo vCPU quota or subnet IPs for surge nodesRequest quota, lower max surge, or use max unavailable
Manual load balancer or NSG change disappearsEdited inside the node resource groupUse Service annotations or your own VNet resources
Cluster forced onto a newer versionFell out of the support windowRun the stable channel with a maintenance window
Pods cannot be reached from on-premisesOverlay pod IPs are not routableExpose through a Service or ingress, or use pod subnet
DNS failures under loadApplication pods crowding the system poolTaint the system pool and size it for add-ons

Trade-offs

AKS removes control plane operations but not cluster operations: upgrades, capacity, disruption budgets and network design remain yours. Overlay is simpler and scales further, at the cost of pods not being directly addressable. Faster upgrades with high surge cost quota and more simultaneous disruption. Long Term Support buys time between upgrades but lets version drift accumulate.

Cost follows the same split. The Standard and Premium tiers add a per-cluster management charge, while nodes, disks, load balancers and egress are billed as ordinary Azure resources in the node resource group, including surge nodes for as long as an upgrade runs. Many small clusters multiply the management charge and the system pools; one large shared cluster concentrates blast radius and upgrade risk. Most teams settle on one cluster per environment and region, with namespaces and node pools separating workloads inside it. Define the cluster in code so these choices are reviewable; Bicep and ARM covers the native option.

What to do next

  1. List the resources in your node resource group and confirm nobody manages them by hand.
  2. Give the cluster a tainted system pool across three zones and move workloads to user pools.
  3. Pick the tier from the table: Standard for anything with users, Premium only if you need LTS.
  4. Confirm your network model; plan the move off kubenet before March 31, 2028 if you still use it.
  5. Set the stable and NodeImage channels with two maintenance windows of at least four hours.
  6. Set max surge to 33% on production pools, check quota and subnet IPs, and audit every PDB for zero allowed disruptions.
  7. Rehearse a full upgrade on a staging cluster and record how long each batch takes.
Key takeaway: AKS runs the control plane; everything that decides whether your cluster stays healthy is still yours. Leave the node resource group to AKS, isolate critical add-ons on a tainted system pool, buy the Standard tier for production, choose Azure CNI Overlay unless pods must be directly reachable, and treat upgrades as a capacity exercise: surge needs quota and addresses, drains need disruption budgets that allow eviction, and both channels need maintenance windows long enough to finish.