Amazon Elastic Kubernetes Service (EKS) runs the Kubernetes control plane for you: the API servers and etcd, replicated across Availability Zones, patched and backed up by AWS. Everything else is still yours: the nodes pods run on, the network that gives them addresses, the identities that let people and programs in, the add-ons, and above all the obligation to upgrade on the Kubernetes release clock. Most EKS trouble comes from the seams between what AWS runs and what you run, so this article is organised around those seams.

We start with the architecture, then work through the four decisions every cluster makes: compute, networking, human access and workload identity. Then we cover the version lifecycle, a worked setup for a two-team production cluster, and the failure modes that fill EKS support tickets. Commands, names and numbers were checked against the EKS user guide, best-practices guide and pricing page in September 2026; AWS changes EKS often, so confirm anything version-specific before relying on it.

Advertisement

Architecture: the split between AWS and you

An EKS cluster's control plane runs in an AWS-owned account. It reaches your nodes through elastic network interfaces that EKS places in the subnets you choose at creation. Nodes and kubectl reach the API server through the cluster endpoint, which can be public, private (resolvable and reachable only inside the VPC), or both. A private endpoint with a restricted public one, or no public one, is the usual production posture; see Amazon VPC in depth for the subnet and routing design underneath.

Amazon EKS: who runs what, and how traffic and identity flowAWS-managed control planeAPI servers + etcd across AZs, AWS accountyou pay per cluster hour; you choose the versionEKS APIaccess entries, add-ons, Pod IdentityYour VPC: control-plane ENIs in your subnets, API endpoint public and/or privatekubelet / APIManaged node groupsEC2 ASG, you pick AMI/typeKarpenter / Auto Modenodes from pending podsFargateone pod per micro-VMVPC CNIpods get VPC IPs from ENIs or /28 prefixesPod Identity agent169.254.170.23: role credentials per service account
EKS splits responsibility: AWS runs the control plane and the EKS API; you own the VPC, the compute, the CNI, identity mappings and the upgrade cadence.

The control plane costs $0.10 per cluster hour while the cluster's Kubernetes version is in standard support, about $73 a month, and $0.60 per cluster hour in extended support. Nodes, load balancers, NAT gateways and cross-AZ traffic are billed separately and usually dominate. The control plane price is why many organisations run a few large, multi-tenant clusters rather than one per team. The cost is a wider blast radius and harder upgrades.

Compute: four ways to get nodes

Managed node groups are EC2 Auto Scaling groups that EKS creates and updates for you: you pick instance types and an AMI family, and EKS handles joining nodes and draining them during updates. Self-managed nodes give full control over the launch template and bootstrap, at the cost of doing node upgrades yourself. Karpenter, an open-source autoscaler originally built by AWS, watches pending pods and launches right-sized instances directly, without node groups; it consolidates underused nodes and handles Spot interruption. EKS Auto Mode goes further: AWS runs the Karpenter-based provisioning and core node components for you and charges a per-instance management fee on top of EC2. Fargate runs each pod in its own micro-VM, with no nodes to manage, but it has no DaemonSets and several features are unavailable on it, including EKS Pod Identity.

OptionYou manageGood forWatch out for
Managed node groupsInstance types, scaling, AMI updatesSteady workloads, predictable fleetsOne max-pods value per group; mixed types take the lowest
KarpenterNodePools, EC2NodeClasses, the controllerSpiky and heterogeneous workloads, SpotNeeds disruption budgets tuned or consolidation churns pods
Auto ModeNodePools; AWS runs the restTeams that want less node operationsPer-instance fee; less low-level control
FargatePod specs and profilesSmall or isolated batch podsNo DaemonSets, no Pod Identity, per-pod pricing
Advertisement

Networking: pods are VPC citizens

The default Amazon VPC CNI gives every pod a real IP address from your VPC subnets, so pods talk to RDS, other VPCs and on-premises networks without overlays or NAT inside the cluster. The cost is IP consumption. In the default secondary-IP mode, each node's pod capacity is bounded by how many network interfaces its instance type supports and how many addresses each interface holds, which is why a small instance can have spare CPU and memory yet refuse pods.

Prefix delegation changes the unit of allocation. With ENABLE_PREFIX_DELEGATION=true (VPC CNI 1.9.0 or later, Nitro instances), the CNI attaches /28 IPv4 prefixes of 16 addresses to each interface slot instead of single addresses, multiplying pod density and cutting EC2 API calls; attaching a prefix typically takes under a second, compared with up to about ten seconds to attach a new interface. By default a node advertises the max-pods value for secondary-IP mode, so you must raise kubelet's max pods yourself; AWS provides max-pods-calculator.sh to compute a value, and the kubelet default ceiling is 110.

Prefix mode has two sharp edges. A /28 must be contiguous, so on a heavily used, fragmented subnet allocation fails with InsufficientCidrBlocks; the fix is a new subnet or a VPC subnet CIDR reservation for prefixes. And AWS recommends moving to prefix mode by creating new node groups and draining the old ones, not by rolling existing nodes, because a node holding both individual IPs and prefixes advertises inconsistent capacity. WARM_PREFIX_TARGET (default 1) keeps one spare prefix per node for fast pod start; WARM_IP_TARGET and MINIMUM_IP_TARGET override it when addresses are scarce. Size subnets for pods, not nodes: a /24 of 251 usable addresses runs out quickly at a few dozen pods per node.

Human access: access entries replace aws-auth

Authentication to the Kubernetes API uses IAM: kubectl presents a signed token, and EKS maps the IAM principal to Kubernetes permissions. Historically that mapping lived in the aws-auth ConfigMap in kube-system, a YAML file where one indentation error could lock everyone out. Access entries move the mapping into the EKS API. A cluster's authentication mode is CONFIG_MAP, API_AND_CONFIG_MAP or API; moving to a mode with the API is one-way.

An access entry names exactly one IAM principal. You grant permissions either by associating AWS-managed access policies (AmazonEKSClusterAdminPolicy, AmazonEKSAdminPolicy, AmazonEKSEditPolicy, AmazonEKSViewPolicy), scoped to the cluster or to namespaces, or by adding Kubernetes group names and writing ordinary RBAC bindings for them. Because the mapping lives in AWS, it is managed with IAM credentials, audited in CloudTrail, and recoverable without Kubernetes API access.

# 1. Move the cluster to the EKS access-entry API (keeps aws-auth working during migration; not reversible)
aws eks update-cluster-config --name prod --access-config authenticationMode=API_AND_CONFIG_MAP

# 2. Platform team: cluster admin, via an IAM role people assume through SSO
aws eks create-access-entry --cluster-name prod \
    --principal-arn arn:aws:iam::111122223333:role/platform-admin --type STANDARD
aws eks associate-access-policy --cluster-name prod \
    --principal-arn arn:aws:iam::111122223333:role/platform-admin \
    --access-scope type=cluster \
    --policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy

# 3. Checkout team: edit rights in their namespaces only
aws eks create-access-entry --cluster-name prod \
    --principal-arn arn:aws:iam::111122223333:role/checkout-dev --type STANDARD
aws eks associate-access-policy --cluster-name prod \
    --principal-arn arn:aws:iam::111122223333:role/checkout-dev \
    --access-scope type=namespace,namespaces=checkout,checkout-staging \
    --policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSEditPolicy

Two behaviours surprise people. kubectl auth can-i --list does not show permissions that come from access policies, only those from RBAC. And an access entry is tied to the principal's internal id, not only its ARN, so if an IAM role is deleted and recreated with the same name, the old entry silently stops matching; delete and recreate the entry too.

Workload identity: Pod Identity and IRSA

Pods that call AWS APIs should never use node instance-profile credentials or static keys. EKS offers two mechanisms. IAM Roles for Service Accounts (IRSA) federates the cluster's OIDC issuer into IAM: each cluster needs an IAM OIDC provider, and each role's trust policy names that provider and the service account. EKS Pod Identity is simpler. Roles trust one service principal, pods.eks.amazonaws.com, with sts:AssumeRole and sts:TagSession, and the binding of role to namespace and service account is an EKS API object, not an annotation. The same role can be associated in many clusters without editing its trust policy.

# Trust policy for the workload role (trust-pod-identity.json)
{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "Service": "pods.eks.amazonaws.com" },
    "Action": ["sts:AssumeRole", "sts:TagSession"]
  }]
}

aws iam create-role --role-name checkout-orders-writer \
    --assume-role-policy-document file://trust-pod-identity.json
aws iam attach-role-policy --role-name checkout-orders-writer \
    --policy-arn arn:aws:iam::111122223333:policy/orders-table-write

# Bind the role to one service account in one namespace; no annotation on the service account
aws eks create-pod-identity-association --cluster-name prod \
    --namespace checkout --service-account orders-api \
    --role-arn arn:aws:iam::111122223333:role/checkout-orders-writer

Under the hood, the Pod Identity agent runs as a DaemonSet on each node's host network, listening on the link-local address 169.254.170.23 (ports 80 and 2703). EKS injects environment variables that point the AWS SDK's default credential chain at the agent, and the agent obtains credentials from the EKS Auth service once per node rather than once per pod. Practical consequences: pods behind an HTTP proxy need 169.254.170.23 in NO_PROXY; SDKs older than Pod Identity support fall back to other credentials silently; associations are eventually consistent, so a pod started seconds after one is created may get no credentials; and Pod Identity does not work on Fargate or Windows nodes, where IRSA is still the answer. Restrict IMDS access from pods too, otherwise a container can still reach the node role. For writing tight role policies, see AWS IAM in depth.

The version lifecycle is an operating cost

Kubernetes ships a minor version roughly every four months. EKS gives each version 14 months of standard support, then 12 months of extended support at the higher $0.60 hourly rate, 26 months in total. Extended support is on by default. When it ends, EKS upgrades the control plane automatically, and managed and self-managed nodes are left on the old version. Control-plane upgrades go one minor version at a time, and an in-place upgrade can be rolled back to the previous minor version within seven days. Nodes may lag the control plane by up to three minor versions, but running them persistently that far behind is not recommended.

# Upgrade loop, one minor version at a time
aws eks describe-cluster-versions                     # support status and end dates per version
kubectl get --raw /metrics | grep apiserver_requested_deprecated_apis   # clients still calling removed APIs
aws eks update-cluster-version --name prod --kubernetes-version 1.36
# then: update managed add-ons (VPC CNI, CoreDNS, kube-proxy, Pod Identity agent),
# then roll node groups / let Karpenter or Auto Mode replace nodes, respecting PodDisruptionBudgets

The real work in an upgrade is not the API call. It is finding workloads and controllers that still use APIs the new version removes, and upgrading add-ons in a compatible order. Budget for an upgrade roughly three times a year per cluster, rehearse it on a staging cluster built from the same infrastructure code, and treat an extended-support bill as a signal that the process is broken, not as a place to live.

Worked setup: one cluster, two teams

A payments company runs a production cluster for a platform team and a checkout team. The platform team creates the cluster with Terraform in three private subnets per AZ, sized /20 so prefix delegation has contiguous space, with a private API endpoint reached over the corporate network. System components (CoreDNS, the load balancer controller, Karpenter itself) run on a small managed node group; application pods run on Karpenter-provisioned nodes, with On-Demand for the payments API and Spot for batch jobs. The cluster uses API authentication mode from day one, so there is no aws-auth ConfigMap to drift.

People assume SSO roles: platform-admin gets the cluster-admin policy, and checkout-dev gets the edit policy scoped to the checkout and checkout-staging namespaces. The checkout orders-api service account is bound by Pod Identity to a role that can write one DynamoDB table and nothing else. A namespace quota and default network policy stop one team exhausting the other's capacity. Six months later, the upgrade from one minor version to the next follows the runbook above in an afternoon, because staging was upgraded two weeks earlier from the same code.

Failure modes and what they look like

SymptomLikely causeWhat to do
Pods Pending with free CPU and memoryNode at max pods or subnet out of IPsEnable prefix delegation on new nodes; add or enlarge subnets
CNI log shows InsufficientCidrBlocksFragmented subnet cannot supply a /28New subnet or subnet CIDR reservation for prefixes
Everyone locked out after an editBroken aws-auth ConfigMapRestore access with an access entry via the EKS API; migrate off the ConfigMap
Pod gets AccessDenied as the node roleNo association, old SDK, or proxy intercepting 169.254.170.23Check the association, SDK version and NO_PROXY; block pod IMDS access
Upgrade breaks a controllerRemoved API still in useCheck deprecated-API metrics and manifests before upgrading
Bill jumps with no traffic changeVersion entered extended supportUpgrade; alert on end-of-standard-support dates

If you are choosing between EKS and a simpler container service for a new workload, compare it with Amazon ECS: ECS has no control-plane fee and no version treadmill, at the cost of the Kubernetes ecosystem and portability.

What to do next

  1. Run aws eks describe-cluster-versions and list every cluster's end-of-standard-support date; put the next upgrade on the calendar.
  2. Move every cluster to API_AND_CONFIG_MAP, recreate aws-auth mappings as access entries, then switch to API.
  3. Replace node-role and static-key access in pods with Pod Identity (or IRSA on Fargate), and block pod access to IMDS.
  4. Check pod density and subnet headroom; plan prefix delegation on new node groups with subnets sized for pods.
  5. Choose one node strategy per workload class (managed groups for system pods, Karpenter or Auto Mode for applications) and set PodDisruptionBudgets before enabling consolidation.
  6. Rehearse the next upgrade on staging from the same infrastructure code, including add-ons and node replacement.
Key takeaway: EKS manages the Kubernetes control plane; you still own compute, pod networking, identity and the upgrade cadence. Give pods VPC addresses with room to grow (prefix delegation on contiguous subnets), manage human access with access entries instead of aws-auth, give workloads Pod Identity roles instead of node credentials, and upgrade within the 14-month standard window. Extended support costs six times as much per cluster hour, and it only postpones the upgrade you still have to do.