An Auto Scaling group looks like a counter: tell it how many EC2 instances you want and it keeps that many running. In practice it is a control system with several independent inputs: scaling policies that change the target, health checks that condemn instances, instance refreshes that replace them, and lifecycle hooks that pause them. Most production incidents involving Auto Scaling groups come from those inputs interacting in ways their owners did not expect: a fleet that terminates healthy instances during a deploy, a scale-out that arrives ten minutes late, or a scale-in that kills requests mid-flight.

This article explains the model from first principles, then walks through a web fleet behind an Application Load Balancer, with the settings that matter and the defaults that bite. Launch templates and instance types are covered in Amazon EC2, and Spot interruption handling in Spot capacity; this page assumes both and concentrates on the group itself. Values quoted here come from the Amazon EC2 Auto Scaling user guide and API reference.

Advertisement

The core idea: reconcile actual to desired

A group has three numbers: minimum, maximum and desired capacity. Desired capacity is the target; minimum and maximum are bounds that every change is clamped to. Auto Scaling continuously compares the instances it has with the target. If there are too few, it launches from the group's launch template into the subnets you listed, choosing Availability Zones to keep them balanced. If there are too many, it picks instances to terminate. If an instance is judged unhealthy, it is terminated and the gap is filled by a new launch.

Everything else changes one of those inputs. A scaling policy changes desired capacity. A health check marks an instance unhealthy. An instance refresh marks instances for replacement. A lifecycle hook holds an instance in a wait state before the next transition. Keeping this separation in mind makes behaviour predictable: when the group does something surprising, ask which input changed.

Scaling policiestarget tracking, step, scheduledHealth checksEC2 status, ELB target healthInstance refreshrolling replacementdesiredunhealthyreplaceAuto Scaling groupmin / desired / maxreconciles actual to desiredlaunchLaunch templateAMI, type, user dataterminateTermination policyAZ balance firstAZ aPending:Wait -> InServiceAZ bInServiceAZ cTerminating:WaitLoad balancer target groupregisters, health-checks, drains
The group reconciles actual instances to desired capacity. Policies change the target, health checks and refreshes condemn instances, the launch template defines new ones, the termination policy picks victims after zone balance, and lifecycle hooks add wait states.

Worked example: a web fleet behind a load balancer

Take a stateless web tier that serves around 4,000 requests per second at peak, with each instance comfortable at about 800. It boots in about 90 seconds, then takes another minute to warm caches. A reasonable group:

aws autoscaling create-auto-scaling-group \
  --auto-scaling-group-name web \
  --launch-template LaunchTemplateName=web,Version='7' \
  --min-size 3 --max-size 30 --desired-capacity 6 \
  --vpc-zone-identifier "subnet-a,subnet-b,subnet-c" \
  --target-group-arns "$TG_ARN" \
  --health-check-type ELB \
  --health-check-grace-period 120 \
  --default-instance-warmup 180

Pin the launch template version rather than using $Latest or $Default, so a template edit does not silently change new launches; it also matters later because automatic rollback of an instance refresh is not supported when the group uses $Latest or $Default. Three subnets in three zones give the group room to rebalance when a zone has trouble. A minimum of 3 keeps one instance per zone; a desired capacity of 6 covers off-peak traffic, and a maximum of 30 caps the bill when something goes wrong.

--health-check-type ELB is the most important line. By default the group uses only EC2 status checks, which catch a dead host but not a hung application. With ELB health checks enabled, an instance that fails the target group's health check is also marked unhealthy and replaced.

Advertisement

Scaling policies: which one, and why

PolicyHow it decidesUse it for
Target trackingKeeps a metric near a target value, creating and managing the alarms for you.The default choice: CPU, or better, request count per target.
Step scalingYour alarms; adds or removes capacity in steps by breach size.Metrics that are not proportional to load, or asymmetric responses.
Simple scalingOne adjustment per alarm, then a cooldown.Legacy; AWS recommends target tracking or step scaling instead.
ScheduledSets min, max or desired at a time.Known daily peaks, batch windows, pre-scaling before launches.
PredictiveForecasts load from history and scales ahead.Strong daily or weekly cycles with slow boot times.

For the web fleet, request count per target is a better signal than CPU, because it measures load directly and does not shift when a new release changes CPU cost per request. It is a count per target per one-minute period, not per second. Save the configuration as tt.json and attach it with the second command:

{
  "TargetValue": 40000,
  "PredefinedMetricSpecification": {
    "PredefinedMetricType": "ALBRequestCountPerTarget",
    "ResourceLabel": "app/web-alb/0123456789abcdef/targetgroup/web/fedcba9876543210"
  }
}
aws autoscaling put-scaling-policy --auto-scaling-group-name web \
  --policy-name requests-per-target --policy-type TargetTrackingScaling \
  --target-tracking-configuration file://tt.json

The units matter. Peak traffic of 4,000 requests per second is 240,000 requests a minute, and an instance comfortable at 800 per second handles 48,000 a minute. A target of 40,000 per target leaves about 17 percent headroom and converges on six instances at peak, rounded up and never below the minimum of three. Pick the target by load testing one instance to the point where latency degrades and taking a comfortable fraction of it. Target tracking scales out aggressively and in conservatively, which is the right bias for a web tier. Add a scheduled action to raise the minimum before known events rather than relying on the policy to react.

Cooldowns apply only to simple scaling. The group's default cooldown is 300 seconds; target tracking and step scaling do not honour it as a cooldown and instead use instance warmup, falling back to the default cooldown value as their warmup only when none is set, which is where most tuning mistakes happen.

Warmup and grace periods: the two timers people confuse

Default instance warmup is how long after reaching InService a new instance's metrics are left out of the group's aggregate, and it is how long a new instance counts as warming up for scaling decisions. While instances are warming up, the group counts them toward capacity when deciding how much to add, so several alarm breaches do not trigger several scale-outs, and dynamic scale-in is blocked. It is not enabled by default. If you leave it unset, target tracking and step policies fall back to the default cooldown, instance refresh falls back to the health check grace period, and predictive scaling has no warmup at all. AWS strongly recommends setting it; 300 seconds is its suggested starting point, which you then tune to your application.

Health check grace period is how long after InService the group ignores failed health checks before replacing an instance. It defaults to 300 seconds when you create a group in the console and to 0 when you create one with the CLI or an SDK, so the same template can behave differently depending on who created it. Too short, and instances that are still starting are killed in a loop; too long, and genuinely broken instances serve errors for minutes. If you use a launch lifecycle hook so that instances only enter service when ready, the grace period can be 0.

Size both from measurement: time from launch to passing the load balancer health check, and time until CPU and latency settle. For the example fleet that is roughly 120 and 180 seconds.

Which instance dies: termination order

On scale-in, zone balance comes first, whatever termination policy you set: Auto Scaling picks the zone with the most instances and at least one instance not protected from scale-in. Within it, the default policy prefers instances with outdated configurations (launched from a launch configuration, then from a different launch template, then from the oldest version of the current template), then the instance closest to its next billing hour, then one at random. For mixed-instance groups it first chooses Spot or On-Demand to move toward your ratio and to follow the allocation strategy.

Other predefined policies include OldestInstance, NewestInstance, OldestLaunchTemplate and AllocationStrategy, and you can supply a Lambda function for custom selection. Unhealthy instances bypass the termination policy entirely. Use scale-in protection on instances doing long jobs, and remember that protection only prevents scale-in; it does not stop health-check replacement.

Lifecycle hooks: draining and bootstrapping

A lifecycle hook pauses an instance in Pending:Wait on launch or Terminating:Wait on termination so you can do work: register with a service, pull configuration, finish in-flight jobs, upload logs. The heartbeat timeout ranges from 30 to 7,200 seconds and defaults to 3,600; the default result when it expires, or on unexpected failure, is ABANDON, and a group can have 50 hooks by default. You extend a wait with RecordLifecycleActionHeartbeat and end it with CompleteLifecycleAction. A termination hook for a queue worker:

aws autoscaling put-lifecycle-hook --auto-scaling-group-name web \
  --lifecycle-hook-name drain --lifecycle-transition autoscaling:EC2_INSTANCE_TERMINATING \
  --heartbeat-timeout 300 --default-result CONTINUE

# on the instance (needs IMDSv2 and an instance role allowing CompleteLifecycleAction)
md() {   # fresh IMDSv2 token per call, so a loop that runs for days never uses an expired one
  local t; t=$(curl -sX PUT http://169.254.169.254/latest/api/token -H "X-aws-ec2-metadata-token-ttl-seconds: 60")
  curl -s -H "X-aws-ec2-metadata-token: $t" http://169.254.169.254/latest/meta-data/$1
}
while [ "$(md autoscaling/target-lifecycle-state)" != "Terminated" ]; do sleep 5; done
systemctl stop worker          # finish in-flight jobs, flush buffers, ship logs
aws autoscaling complete-lifecycle-action --auto-scaling-group-name web \
  --lifecycle-hook-name drain --instance-id "$(md instance-id)" --lifecycle-action-result CONTINUE

Notice the choices. CONTINUE as the default result means a crashed drain script still lets termination proceed after five minutes rather than an hour. On termination, an ABANDON result also terminates the instance, but on launch it terminates the new instance instead of putting it in service, which is the safe default for bootstrap hooks. When the group is attached to a load balancer, the instance is deregistered and connection draining runs before the terminate hook's wait state, so web traffic has already moved away when the script runs. EventBridge is the recommended target for lifecycle notifications when a central service, rather than the instance itself, does the work.

Instance refresh: rolling deploys

Instance refresh replaces instances in batches to roll out a new launch template version or AMI. The key preferences are MinHealthyPercentage (default 90) and MaxHealthyPercentage (100 to 200, default 100). With both at their defaults the group terminates and launches at the same time, dipping capacity by up to 10 percent. Setting the maximum above 100 lets it launch replacements before terminating old instances, which is the right choice for fleets running close to capacity.

aws autoscaling start-instance-refresh --auto-scaling-group-name web \
  --desired-configuration '{"LaunchTemplate":{"LaunchTemplateName":"web","Version":"8"}}' \
  --preferences '{
    "MinHealthyPercentage": 90,
    "MaxHealthyPercentage": 120,
    "SkipMatching": true,
    "CheckpointPercentages": [10, 50, 100],
    "CheckpointDelay": 600,
    "AutoRollback": true,
    "AlarmSpecification": {"Alarms": ["web-5xx-high"]}
  }'

Checkpoints pause after 10 and 50 percent for ten minutes each so you can observe the new version under real traffic. The alarm makes a spike in 5xx responses fail the refresh, and AutoRollback then returns the group to its previous configuration. SkipMatching avoids replacing instances already on the desired configuration, which makes retries cheap. By default a refresh waits up to an hour for instances in standby or protected from scale-in, then fails; decide in advance whether it should ignore or replace them.

Warm pools for slow starters

If boot takes minutes, for example because instances load large models or datasets, a warm pool keeps pre-initialised instances beside the group in Stopped, Hibernated or Running state. Stopped instances cost only their volumes and Elastic IPs. On scale-out the group starts them instead of launching from scratch. Use a launch hook so instances finish initialising before being stopped, because Auto Scaling does not wait for user data to finish before stopping. Warm pools do not support Spot instances in mixed groups or weighted instance types, and they need EBS root volumes. When the pool is empty, launches fall back to cold starts.

Failure modes and how to spot them

  • Replacement loops. Grace period shorter than startup, with ELB health checks, gives a group that launches and kills instances forever. The activity history shows repeated health-check terminations.
  • Launch failures. Capacity shortages in a zone, a deleted AMI, a missing instance profile or an exhausted subnet leave the group below desired. Watch the scaling activity status and alert when desired and in-service counts diverge for more than a few minutes.
  • Scale-in that drops requests. No termination hook and a short deregistration delay; long requests are cut off.
  • Unexpected zone rebalancing. After a zone recovers, the group launches and terminates to rebalance, which can terminate newer instances before older ones.
  • Metric blind spots. Without default instance warmup, cold instances distort averages and cause over-scaling, then scale-in is blocked.
  • Ownership drift. Someone changes desired capacity by hand, a scheduled action overrides it, or infrastructure-as-code resets it on the next deploy. Decide which tool owns each number.

Trade-offs

Auto Scaling groups give you self-healing and elastic capacity for anything that runs on EC2, with no cluster software to operate. The costs are boot latency measured in minutes, per-instance granularity, and the need for stateless or externally backed instances. When those hurt, containers on Amazon ECS scale in seconds on top of groups that change far less often. Groups remain the right tool for VM-based fleets, workers, and the capacity layer beneath container and Kubernetes clusters. Cost control belongs with capacity planning; see cloud FinOps.

What to do next

  1. List every group you run and check that each uses a pinned launch template version and ELB health checks where it serves traffic.
  2. Measure time to healthy and time to stable metrics; set the health check grace period and default instance warmup from those numbers.
  3. Replace simple scaling policies with target tracking, preferably on request count per target, and add scheduled actions for known peaks.
  4. Add a termination lifecycle hook that drains work, with a short heartbeat timeout and a deliberate default result.
  5. Deploy with instance refresh using checkpoints, a CloudWatch alarm and automatic rollback, and practise one rollback.
  6. Alert when in-service capacity stays below desired, and review scaling activity history after each incident.
Key takeaway: An Auto Scaling group reconciles running instances to a desired capacity that policies change, while health checks and refreshes condemn instances and lifecycle hooks add wait states. Pin launch template versions, use ELB health checks, set the grace period and default instance warmup from measured start-up times, prefer target tracking on request count, drain with termination hooks, and roll out with instance refresh checkpoints, alarms and rollback. Zone balance always comes before the termination policy.