Amazon Elastic Container Service is AWS's own container orchestrator. You describe what should run, and a regional control plane keeps reconciling that desired state with what is actually running, on serverless Fargate capacity or on EC2 instances. It has fewer concepts than Kubernetes, and no control plane for you to operate, but the concepts it does have decide whether deployments are safe, costs are sane and failures are recoverable.

This article goes below the introduction in ECS and Fargate basics: the two IAM roles every task has, how network modes change density and security, how the scheduler chooses where a task runs, how capacity provider strategies and managed scaling work, how deployments roll forward and back, and the failure modes that show up in production.

Advertisement

The mental model: desired state and a reconciler

Everything in ECS is a loop. A service says: keep 6 tasks of revision 14 running behind this target group. The service scheduler compares that with reality, starts tasks that are missing, stops tasks that fail health checks, and runs deployments when the revision changes. Application Auto Scaling changes the desired count from metrics, and on EC2-backed capacity, cluster auto scaling changes the number of instances so the tasks have somewhere to go. None of these loops is instantaneous, and most production surprises come from their interaction.

ECS: a regional control plane reconciling services onto capacity providersYou / CItask definition revision, service updateECS control planeservice scheduler, placement, deploymentsApplication Auto Scalingdesired count from metricsAPIset desiredFargateisolated per taskFargate Spotinterruptible, 2-min noticeASG capacity provideryour EC2 + ECS agentManaged InstancesEC2 run and patched by AWSstrategy: base + weightEach task (awsvpc)own ENI and security groups, task role for app calls, execution role for pulls, logs, secretsTraffic and healthload balancer target group, health checks, circuit breaker or blue/green bakeDesired state goes in through the API; the scheduler keeps reconciling it against what is actually running.
Requests change desired state through the ECS API. The scheduler places tasks on capacity providers according to the service's strategy, and each task gets its own network interface and IAM identity in awsvpc mode.
ObjectWhat it isKey point
Clustera namespace for services, tasks and capacityholds default capacity provider strategy and settings such as Container Insights
Task definitionversioned JSON blueprint: containers, CPU and memory, network mode, roles, loggingrevisions are immutable; deploying means pointing a service at a new revision
Taska running instance of a task definitionone or more containers sharing a lifecycle, placed together
Servicekeeps N tasks of a revision running, optionally behind a load balancerthe scheduler replaces failed tasks and runs deployments
Capacity providerwhere tasks run: Fargate, Fargate Spot, an Auto Scaling group, or Managed Instancesservices choose providers through a strategy
Container instancean EC2 host running the ECS agent, registered to a clusteronly exists for EC2-backed capacity

Two IAM roles per task, and why they must stay separate

A task definition can name two roles, and confusing them is the most common ECS permission bug. The task execution role is used by ECS and the agent on your behalf before and around your code: pulling the image from ECR, sending logs to CloudWatch Logs, and fetching secrets from Secrets Manager or Parameter Store to inject as environment variables. The task role is the identity your application code uses when it calls AWS APIs through the SDK, which obtains credentials from the task's credential endpoint.

Keep them minimal and separate. The execution role needs only pull, log and specific secret-read permissions. The task role gets the application's permissions, scoped to the resources that one service uses. Never rely on the EC2 instance profile for application permissions on EC2 capacity, because every task on the host would inherit it. Policy design is covered in AWS IAM, and rotation of the secrets these roles read in secrets rotation.

{
  "family": "orders-api",
  "networkMode": "awsvpc",
  "requiresCompatibilities": ["FARGATE"],
  "cpu": "1024",
  "memory": "2048",
  "executionRoleArn": "arn:aws:iam::123456789012:role/orders-api-exec",
  "taskRoleArn": "arn:aws:iam::123456789012:role/orders-api-task",
  "containerDefinitions": [{
    "name": "api",
    "image": "123456789012.dkr.ecr.eu-west-1.amazonaws.com/orders-api:2026-09-30.1",
    "essential": true,
    "portMappings": [{ "containerPort": 8080, "protocol": "tcp" }],
    "secrets": [{ "name": "DB_PASSWORD",
                  "valueFrom": "arn:aws:secretsmanager:eu-west-1:123456789012:secret:orders-db" }],
    "healthCheck": { "command": ["CMD-SHELL", "curl -fs http://localhost:8080/health || exit 1"],
                     "interval": 15, "timeout": 5, "retries": 3, "startPeriod": 30 },
    "stopTimeout": 60,
    "logConfiguration": { "logDriver": "awslogs",
      "options": { "awslogs-group": "/ecs/orders-api", "awslogs-region": "eu-west-1",
                   "awslogs-stream-prefix": "api" } }
  }]
}

Two details in that definition matter operationally. The image tag is unique per build, so a revision always means the same bits; a mutable tag such as latest makes rollbacks meaningless. And stopTimeout gives the container time to drain in-flight requests after SIGTERM before it is killed; Fargate caps it at 120 seconds.

Advertisement

Network modes

awsvpc gives each task its own elastic network interface with a private IP in your subnet and its own security groups. It is required on Fargate and is the right default on EC2: security is per service rather than per host, and load balancer targets are IP addresses. The cost on EC2 is density, because each instance type supports a limited number of network interfaces; ENI trunking, an account setting for supported instance types, raises the number of awsvpc tasks per instance. Each task also consumes an IP address, so size subnets for the peak task count during deployments, not the steady state. Subnet planning is covered in AWS VPC.

bridge uses Docker's virtual network on the host with dynamic host port mapping; many tasks share the host's interfaces and security groups. host binds containers directly to the host's network, so two tasks cannot use the same port on one instance. These modes exist for EC2 capacity and are mostly seen in older clusters or for special networking needs.

How the scheduler places a task

Placement runs in three steps. First, the capacity provider strategy decides which provider gets the task. A strategy lists providers with a base, the minimum number of tasks on that provider, and a weight, the relative share of tasks after the base is satisfied. Only one provider in a strategy can have a non-zero base. A strategy of Fargate base 2 weight 1 plus Fargate Spot weight 3 puts the first 2 tasks on Fargate and then splits the rest one to three.

Second, on EC2-backed providers, the scheduler filters instances that have enough free CPU, memory and ports and satisfy the placement constraints, such as distinctInstance or a memberOf expression on instance attributes. Third, placement strategies rank the survivors: binpack on memory or CPU fills instances to use fewer of them, spread on availability zone or instance ID maximises fault isolation, and random does what it says. Fargate ignores constraints and strategies and spreads tasks across availability zones itself.

aws ecs create-service \
  --cluster prod --service-name orders-api \
  --task-definition orders-api:14 --desired-count 6 \
  --capacity-provider-strategy capacityProvider=FARGATE,base=2,weight=1 \
                               capacityProvider=FARGATE_SPOT,weight=3 \
  --network-configuration "awsvpcConfiguration={subnets=[subnet-a,subnet-b,subnet-c],securityGroups=[sg-orders],assignPublicIp=DISABLED}" \
  --load-balancers targetGroupArn=arn:aws:elasticloadbalancing:eu-west-1:123456789012:targetgroup/orders/abc,containerName=api,containerPort=8080 \
  --health-check-grace-period-seconds 60 \
  --deployment-configuration "minimumHealthyPercent=100,maximumPercent=200,deploymentCircuitBreaker={enable=true,rollback=true}"

Capacity: Fargate, Spot, Auto Scaling groups and Managed Instances

Fargate runs each task in its own isolated environment; you pay for the requested vCPU and memory while the task runs, and there are no hosts to patch. Fargate Spot runs the same tasks on spare capacity at a discount and can reclaim them with a two-minute warning, delivered to the task as SIGTERM; use it for stateless or retryable work, always alongside a Fargate base.

An Auto Scaling group capacity provider puts tasks on your own EC2 instances. It is chosen for GPUs, specific instance families, large tasks, daemon-style sidecars or cost at high steady utilisation. With managed scaling on, ECS publishes a metric called CapacityProviderReservation in the AWS/ECS/ManagedScaling namespace, defined as instances needed divided by instances running, times 100, and attaches a target tracking policy that steers it towards your targetCapacity percentage. With a target of 100, new tasks wait in PENDING while instances launch. With a target of 90, ECS keeps roughly 10 percent headroom so bursts place immediately, at the cost of idle instances. Scaling out from zero instances launches 2. The documentation also advises ordering binpack before spread when the target is below 100, and warns that removing the AmazonECSManaged tag from the group breaks managed scaling.

ECS Managed Instances is a newer capacity provider type in which AWS launches, patches and scales the EC2 instances for you from instance requirements you specify, with on-demand, Spot and reserved capacity options. It sits between Fargate's zero host management and the full control of your own Auto Scaling group.

A worked example: a cluster of 10 instances can hold 40 tasks. A deployment with maximumPercent 200 on a 40-task service briefly needs room for 80 tasks, so the scheduler needs about 20 instances. CapacityProviderReservation reads 200 against a target of 100, the group scales out, new tasks sit in PENDING for the few minutes it takes instances to boot and register, and the deployment is slower than on Fargate. Lowering maximumPercent to 125 trades a slower rollout for far less surge capacity.

Deployments and rollback

On a rolling deployment, minimumHealthyPercent is the floor of healthy tasks the scheduler must keep, and maximumPercent is the ceiling on running tasks. With 10 tasks, 100 and 200 mean ECS can start 10 new tasks before stopping any old one: safest, but double capacity. With 50 and 100, ECS stops 5 old tasks first and then starts replacements: no surge, but half capacity during the rollout.

The deployment circuit breaker watches for tasks that fail to start or fail health checks, and after a threshold derived from the desired count it marks the deployment failed and, with rollback enabled, returns the service to the last completed deployment. The health check grace period tells the scheduler to ignore load balancer health checks for a new task's first seconds, which prevents slow-starting applications from being killed in a loop.

ECS also supports native blue/green, linear and canary strategies that shift traffic at the load balancer. For blue/green, the bake time is how long the old revision is kept after traffic moves before it is terminated; the documentation gives a default of 15 minutes and a range of 0 to 1,440 minutes, and notes that the CodeDeploy and external controllers do not support it.

StrategyHow traffic movesRollbackCost
Rollingtasks replaced in batches bounded by minimumHealthyPercent and maximumPercentcircuit breaker can roll back automaticallylittle extra capacity
Blue/green (native)a full green revision is started, then traffic shifts at the load balancerswitch back to blue during the bake timedouble capacity during the deployment
Linear and canary (native)traffic shifts in steps or a small first slicestop and revert during the shiftextra capacity for the new revision
External or CodeDeploy controllera separate system controls task setsdefined by that systemextra moving parts

Scaling the service itself

Service auto scaling is Application Auto Scaling acting on the service's desired count. Target tracking on average CPU or memory utilisation, or on requests per target from the load balancer, suits most web services; step scaling or custom metrics such as queue depth suit workers.

aws application-autoscaling register-scalable-target \
  --service-namespace ecs --resource-id service/prod/orders-api \
  --scalable-dimension ecs:service:DesiredCount --min-capacity 4 --max-capacity 40

aws application-autoscaling put-scaling-policy \
  --service-namespace ecs --resource-id service/prod/orders-api \
  --scalable-dimension ecs:service:DesiredCount --policy-name cpu60 \
  --policy-type TargetTrackingScaling \
  --target-tracking-scaling-policy-configuration \
  '{"TargetValue":60,"PredefinedMetricSpecification":{"PredefinedMetricType":"ECSServiceAverageCPUUtilization"},"ScaleInCooldown":120,"ScaleOutCooldown":60}'

On EC2 capacity, remember there are two loops: service scaling adds tasks, and those tasks sit in PENDING until cluster scaling adds instances. Measure the total time from traffic spike to serving task, and set the minimum task count so that the first minutes of a spike are absorbed by existing capacity.

Failure modes

  • Task stuck in PENDING. No capacity that satisfies CPU, memory, ports or constraints, no free IPs in the subnets, or ENI limits on EC2. The service events list the reason.
  • CannotPullContainerError. Missing execution role permissions, no route to ECR from private subnets, or a missing image tag.
  • Restart loops. Container health check or load balancer health check failing, often because the grace period is shorter than startup. The circuit breaker should stop these.
  • AccessDenied from application code. Permission added to the execution role instead of the task role.
  • Dropped requests on deploy. The application exits immediately on SIGTERM, or the target group's deregistration delay is longer than stopTimeout.
  • Spot-heavy service collapses. Fargate Spot reclaimed many tasks at once with no Fargate base to fall back on.
  • Instances never scale in. Tasks spread thinly across instances by a spread-first strategy, or scale-in protection left on instances ECS does not manage.

When ECS is the right choice

ECS fits teams that run on AWS, want containers without operating a Kubernetes control plane, and are happy to use AWS primitives for networking, IAM, load balancing and logs. Kubernetes on EKS fits teams that need its ecosystem, custom controllers or portability across clouds, and can staff its operation. Lambda fits short, event-driven work. Many organisations use ECS for services and batch workers and keep the rest of the platform simple.

What to do next

  1. Audit every task definition for separate execution and task roles, each least-privilege.
  2. Replace mutable image tags with per-build tags and record them in deployment logs.
  3. Enable the deployment circuit breaker with rollback on every service, and set a health check grace period longer than worst-case startup.
  4. Handle SIGTERM in each application, and align stopTimeout with the load balancer deregistration delay.
  5. Use a capacity provider strategy with an on-demand base and a Spot share for stateless services.
  6. On EC2 capacity, enable managed scaling, pick a targetCapacity that leaves headroom, and check that maximumPercent surge fits.
  7. Size subnets for peak task count during deployments, and enable ENI trunking where you run awsvpc on EC2.
  8. Add alarms on service events, pending task count and failed deployments.
Key takeaway: ECS is a set of reconciliation loops: the service scheduler keeps desired tasks running and deploys new revisions, Application Auto Scaling sets the desired count, and capacity providers with managed scaling supply somewhere to run. Keep the execution and task roles separate, use awsvpc and plan IPs, choose placement through capacity provider strategies, and make every deployment reversible with immutable image tags, health check grace periods, circuit breaker rollback or blue/green bake time.