Every major cloud sells Kubernetes as a managed service: Amazon Elastic Kubernetes Service (EKS), Google Kubernetes Engine (GKE) and Azure Kubernetes Service (AKS). All three pass the upstream conformance tests, so a Deployment manifest means the same thing on each. What differs is everything around the API: who runs the nodes, how pods get IP addresses, how pods authenticate to cloud services, how long a Kubernetes version is supported before you are forced to upgrade, and what you pay for the privilege.
This article compares the three on those axes from first principles, works through the calculation that most often goes wrong (pod IP address planning), shows which parts of a manifest have to change per cloud, and ends with a checklist for running on one cloud today while keeping the option to move.
What 'managed' covers, and what it does not
Kubernetes has two halves. The control plane is the API server, the etcd database that stores cluster state, the scheduler and the controller manager. The data plane is the worker nodes, each running the kubelet, a container runtime and a network agent. All three services run the control plane for you: they operate etcd, spread the API servers across zones, patch them, and back them with an SLA. You never see those machines.
The data plane is where the offerings diverge. In the classic mode you own node pools: the provider gives you an image and a group abstraction, but you choose instance types, trigger node upgrades and pay for idle capacity. In the fully managed modes, GKE Autopilot, EKS Auto Mode and AKS Automatic, the provider also chooses and patches nodes, and in Autopilot you are billed for the resources your pods request rather than for nodes. You give up control over node configuration (privileged workloads, custom kernels, some DaemonSets) in exchange for not operating nodes at all.
Some things stay yours in every mode: workload manifests, RBAC, network policy, add-on choices, the cluster's upgrade timing within the provider's window, and cost.
Node models side by side
| EKS | GKE | AKS | |
|---|---|---|---|
| You manage node groups | Managed node groups, self-managed nodes | Standard mode node pools | Node pools (system and user) |
| Provider picks nodes per pod demand | Karpenter (open source) or EKS Auto Mode | Node auto-provisioning; Autopilot | Node auto-provisioning (Karpenter-based); AKS Automatic |
| Serverless pods | Fargate profiles | Autopilot | Virtual nodes (Azure Container Instances) |
| Node upgrades | Yours: managed node groups are not upgraded with the control plane | Automatic by release channel and maintenance window | Configurable auto-upgrade channels and node image upgrades |
The practical difference is who reacts when pods are pending. With a fixed node group and the Cluster Autoscaler, capacity grows by one more copy of a pre-chosen instance type. Provisioners such as Karpenter look at the pending pods' requests and launch an instance type that fits them, which gives better packing and makes GPU and spot capacity easier to use. Cloud autoscaling covers the reactive-versus-predictive trade-off that sits above this layer.
Pod networking and the IP planning trap
Each cloud offers a mode in which pods get addresses that are routable in your virtual network. On EKS, the default Amazon VPC CNI assigns every pod a real VPC IP from the node's subnet. GKE's VPC-native clusters assign pods addresses from a secondary (alias) range of the subnet. AKS's Azure CNI can give pods addresses from a dedicated pod subnet. Routable pod IPs make pods first-class citizens of the network: security groups, firewalls, peered networks and on-premises systems can reach them without address translation. The cost is that every pod consumes an address from ranges you planned before the cluster existed, and running out means pods stay pending with a network error even though CPU and memory are free.
The alternative is an overlay: pods get addresses from a private range that only exists inside the cluster, and traffic leaving the cluster is translated to the node's address. Azure CNI Overlay is the recommended AKS mode for this reason; EKS can use prefix delegation or custom networking to stretch addresses; GKE lets you size and add pod ranges. Overlays conserve addresses but make pods invisible to the outside network, which matters if something outside the cluster must call pods directly.
The data-plane implementation also differs. GKE Dataplane V2 and the Azure CNI powered by Cilium use eBPF for routing and network policy; on EKS, network policy is enforced by the VPC CNI's own policy agent or by a CNI you install. For the underlying VPC concepts see cloud networking architecture.
Worked example: sizing a pod range
Suppose a cluster may grow to 60 nodes running up to 30 pods each, and you use routable pod IPs. Add 30 percent headroom for rolling updates, which briefly run old and new pods side by side, and count one address per node:
def pod_ip_plan(nodes_max, pods_per_node, headroom=1.3):
# VPC-routable pod IPs (EKS VPC CNI default mode, Azure CNI pod subnet,
# GKE alias ranges): every pod consumes an address from your network.
need = int(nodes_max * pods_per_node * headroom) + nodes_max
prefix = 32
while 2 ** (32 - prefix) < need:
prefix -= 1
return need, f"/{prefix}"
print(pod_ip_plan(60, 30)) # (2400, '/20') -> 4,096 addresses
print(pod_ip_plan(300, 50)) # (19800, '/17') -> 32,768 addressesYou need about 2,400 addresses, which is a /20 (4,096). A /24 per zone, a common default copied from tutorials, holds 256, enough for about eight nodes. At 300 nodes of 50 pods the answer is a /17, which many enterprise address plans cannot spare; that is the point at which an overlay stops being optional. Do this arithmetic before creating the cluster, because changing a cluster's pod range later ranges from awkward to impossible depending on the cloud and mode.
Workload identity: the part most tied to one cloud
Pods need credentials to call cloud APIs: read a bucket, publish to a queue. Storing access keys in Secrets is the anti-pattern. All three clouds instead let a Kubernetes service account be exchanged for a short-lived cloud identity, and all three build on the same idea: the cluster issues signed service-account tokens, and the cloud's identity service trusts that issuer.
- EKS offers EKS Pod Identity (an association between a service account and an IAM role, managed through the EKS API) and the older IAM Roles for Service Accounts (IRSA), which uses an OIDC provider and a role annotation on the service account.
- GKE offers Workload Identity Federation for GKE, which maps a Kubernetes service account to IAM permissions, either directly or by impersonating a Google service account.
- AKS offers Microsoft Entra Workload ID, which federates a service account with an Entra application or managed identity; the pod needs a label to opt in.
The manifests are nearly identical apart from one annotation, which is why a base-plus-overlay layout works well:
# base/serviceaccount.yaml (portable)
apiVersion: v1
kind: ServiceAccount
metadata:
name: orders-api
---
# overlays/eks/sa-patch.yaml
metadata:
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::111122223333:role/orders-api
---
# overlays/gke/sa-patch.yaml
metadata:
annotations:
iam.gke.io/gcp-service-account: orders-api@my-project.iam.gserviceaccount.com
---
# overlays/aks/sa-patch.yaml (pods also need label azure.workload.identity/use: "true")
metadata:
annotations:
azure.workload.identity/client-id: 00000000-0000-0000-0000-000000000000
---
# StorageClass provisioner per overlay:
# EKS ebs.csi.aws.com GKE pd.csi.storage.gke.io AKS disk.csi.azure.comHuman access differs too: EKS uses access entries mapping IAM principals to Kubernetes permissions (replacing the older aws-auth ConfigMap), GKE combines Google IAM with Kubernetes RBAC, and AKS integrates Entra ID with either Kubernetes RBAC or Azure RBAC. The principles are the same as in cloud IAM architecture: least privilege, short-lived credentials, auditable bindings.
Versions, support windows and forced upgrades
Upstream Kubernetes ships a minor version roughly every four months and supports each for about fourteen months. Managed services extend that window for a fee, then upgrade you whether you are ready or not.
| Standard support | Longer support | What happens at the end | |
|---|---|---|---|
| EKS | 14 months per minor version | Extended support: 12 more months, enabled by default, charged per cluster hour | Control plane auto-upgraded; nodes stay behind until you upgrade them |
| GKE | 14 months | Extended channel: up to 24 months total, pay-per-use in the extended period | Automatic upgrades per release channel and maintenance policy |
| AKS | Community-aligned support window for recent minor versions | Long Term Support: 24 months, requires the Premium tier | Clusters on unsupported versions must upgrade to regain support |
An upgrade is always control plane first, then nodes. Node upgrades drain and replace nodes, so they are governed by your PodDisruptionBudgets: a budget that allows zero disruptions blocks the drain indefinitely. Before each minor upgrade, scan manifests and Helm charts for removed API versions, check admission webhooks (a webhook whose backend is down can block every pod creation during the roll), and upgrade a staging cluster on the same channel first.
What it costs
Cost has four parts: a control-plane fee, compute, networking and add-ons.
- Control plane. EKS charges $0.10 per cluster-hour in standard support and $0.60 in extended support; a cluster left on an old version costs six times as much to exist. GKE charges a per-cluster management fee with a monthly free-tier credit that covers one zonal or Autopilot cluster, plus extended-period charges. AKS's Free tier has no management fee and no financially backed SLA; the Standard tier adds a fee and an API-server SLA of 99.95 percent with availability zones, and Premium adds long-term support. Check each provider's pricing page for current rates.
- Compute. Nodes are ordinary virtual machines at ordinary prices, except in Autopilot-style modes that bill on pod requests. Over-sized requests are the largest avoidable cost on every cloud.
- Networking. Cross-zone traffic between pods, load balancers per Service, NAT gateways for egress. Chatty services spread across zones pay for every byte.
- Add-ons. Managed Prometheus, logging ingestion and service meshes bill separately and can exceed the control-plane fee.
For many small teams, one cluster with namespaces is cheaper and simpler than one cluster per team, because every cluster carries a fixed fee and its own add-ons.
Creating one of each
# EKS: eksctl creates the control plane plus a managed node group
eksctl create cluster --name demo --region us-east-1 --nodes 3
# GKE: an Autopilot cluster (Google manages nodes as well)
gcloud container clusters create-auto demo --location us-central1
# AKS: Standard tier (uptime SLA), overlay networking, workload identity on
az aks create -g demo-rg -n demo --tier standard --node-count 3 \
--network-plugin azure --network-plugin-mode overlay \
--enable-oidc-issuer --enable-workload-identity --generate-ssh-keysThese commands are enough to experiment. For anything that lasts, define clusters in Terraform or the cloud's infrastructure-as-code tool, so the pod range, identity settings and upgrade channel are reviewed and reproducible rather than remembered.
Failure modes seen in practice
- Pods pending with IP allocation errors while nodes have spare CPU: the pod range or subnet is exhausted. Plan larger ranges or move to an overlay.
- A surprise bill after a version ages out. EKS extended support is on by default. Track each cluster's end-of-standard-support date.
- Node upgrade stuck for hours. A PodDisruptionBudget allows no disruptions, or a single-replica workload has a budget of one.
- Every deploy fails after an upgrade. A validating or mutating webhook is unreachable and set to fail closed.
- Pods cannot reach cloud APIs. Service account annotation, trust policy and namespace do not match exactly; the error usually surfaces as an access-denied from the cloud SDK.
- Portability that exists only on paper. Load balancer annotations, storage classes and identity bindings are cloud-specific; if they are scattered through every manifest, moving costs a rewrite.
If you run more than one cloud, multi-cloud architecture covers when that is worth it; for cross-cluster traffic management, see service mesh architecture.
What to do next
- Decide whether you want to operate nodes; if not, start with Autopilot, EKS Auto Mode or AKS Automatic and move to self-managed pools only when a workload needs it.
- Run the pod IP calculation for your largest expected cluster before creating it, and choose routable or overlay networking deliberately.
- Use workload identity (Pod Identity or IRSA, Workload Identity Federation, Entra Workload ID) and delete any cloud keys stored in Secrets.
- Put cloud-specific settings (identity annotations, storage classes, load balancer annotations) in per-cloud overlays, and keep the base manifests portable.
- Record each cluster's end-of-standard-support date and rehearse the minor upgrade on staging each cycle.
- Audit PodDisruptionBudgets and webhook failure policies so upgrades cannot deadlock.