EC2 Spot Instances are spare EC2 capacity sold at a discount that AWS describes as up to 90 percent off the On-Demand price. The catch is that EC2 can take the instance back when it needs the capacity. Everything about using Spot well follows from that one sentence: you are not buying a cheaper server, you are buying cheap compute with a probability of losing it, and the work is to make that probability small and the cost of a loss smaller.
This article explains how Spot capacity is organised, how to measure availability before you launch with the Spot Instance Advisor and Spot placement scores, how allocation strategies choose where instances land, and how to handle the rebalance recommendation and the two-minute interruption notice. A worked example sizes the cost of interruptions for a batch fleet, and the article ends with failure modes, trade-offs and a checklist.
Capacity pools and why diversification works
Spot capacity lives in capacity pools. A pool is one instance type in one Availability Zone: m6i.xlarge in one zone is one pool, m6i.2xlarge in the same zone is another, and m6i.xlarge in the next zone is a third. Each pool has its own spare capacity, its own Spot price and its own interruption behaviour. When On-Demand or reserved demand in a pool rises, EC2 reclaims Spot instances from that pool, not from the whole Region.
Prices change gradually, driven by long-term supply and demand in the pool, and there is no bidding. You may set a maximum price, but it defaults to the On-Demand price and AWS advises leaving it there; setting it lower adds a second reason to be interrupted without lowering what you pay. In practice the main interruption reason is capacity: EC2 needs the hardware back. Other reasons exist, such as a constraint on a launch group you requested, but they are rarer and self-inflicted.
The design consequence is diversification. A fleet in one pool loses everything at once when that pool is reclaimed. A fleet spread over twenty pools, five instance types in four zones, usually loses a few instances at a time, which a well-built application absorbs without anyone noticing.
Measuring availability before you launch
Two tools estimate availability before you commit. The Spot Instance Advisor shows, per Region and instance type, a historical band of interruption frequency over the last month, from under 5 percent up to over 20 percent, with typical savings. It is history at type level, not a forecast for your request, but it is a quick way to drop types that are routinely reclaimed.
The Spot placement score is the forward-looking tool. You describe a request, a target capacity in instances, vCPUs or memory plus a list of instance types or attribute requirements, and EC2 returns the top ten Regions or Availability Zones scored from 1 to 10, where 10 means the request is highly likely, not guaranteed, to succeed. It costs nothing. Its rules matter more than its numbers: specify at least three instance types or the score will be low; a score only applies to a request with exactly the same configuration using the capacity-optimized strategy; and capacity moves, so act on a score immediately and re-query before every large launch. Your target capacity limit for scoring is based on your recent Spot usage, so a new account sees low limits.
import boto3
ec2 = boto3.client("ec2", region_name="us-east-1")
resp = ec2.get_spot_placement_scores(
InstanceTypes=["c6i.2xlarge", "c6a.2xlarge", "c7i.2xlarge", "m6i.2xlarge", "c5.2xlarge"],
TargetCapacity=400,
TargetCapacityUnitType="vcpu",
SingleAvailabilityZone=True, # score individual zones, not Regions
RegionNames=["us-east-1", "us-east-2", "us-west-2"],
)
for s in sorted(resp["SpotPlacementScores"], key=lambda s: -s["Score"]):
print(s["Region"], s.get("AvailabilityZoneId", "-"), s["Score"])Use the zonal form for workloads that must sit in one zone, such as a tightly coupled job in a cluster placement group; use the regional form to choose where to build a new fleet. A useful habit is to score two or three candidate type lists: if adding older generations or AMD and Graviton variants lifts the score from 3 to 8, that is the cheapest availability you will ever buy, provided your AMI supports them.
Allocation strategies
When you launch through an Auto Scaling group, EC2 Fleet or Spot Fleet, an allocation strategy picks the pools. The choice decides your interruption rate more than anything else you configure.
| Strategy | Picks pools by | Use when |
|---|---|---|
price-capacity-optimized | Most available capacity, then lowest price among those | Default choice for almost everything |
capacity-optimized | Most available capacity only | Interruptions are very expensive; matches placement-score assumptions |
capacity-optimized-prioritized | Capacity, honouring your type priority order | You prefer certain types but still want availability |
lowest-price | Cheapest pools | Short, trivially restartable work only; highest churn |
AWS's own Auto Scaling guidance is blunt: do not use lowest-price with Capacity Rebalancing, because replacements land in the cheapest pool even if it is about to be reclaimed. A typical mixed-instances group keeps a small On-Demand base for capacity that must exist, puts everything above it on Spot, and lets attribute-based selection expand the type list automatically as new generations appear. The launch template carries the AMI and user data, as described in EC2 launch templates:
{
"AutoScalingGroupName": "render-workers",
"MinSize": 2, "MaxSize": 120, "DesiredCapacity": 40,
"VPCZoneIdentifier": "subnet-aaa,subnet-bbb,subnet-ccc",
"CapacityRebalance": true,
"MixedInstancesPolicy": {
"LaunchTemplate": {
"LaunchTemplateSpecification": {"LaunchTemplateName": "render-worker", "Version": "$Default"},
"Overrides": [{"InstanceRequirements": {
"VCpuCount": {"Min": 8, "Max": 16},
"MemoryMiB": {"Min": 16384},
"CpuManufacturers": ["intel", "amd"]
}}]
},
"InstancesDistribution": {
"OnDemandBaseCapacity": 2,
"OnDemandPercentageAboveBaseCapacity": 0,
"SpotAllocationStrategy": "price-capacity-optimized"
}
}
}
Rebalance recommendations and interruption notices
EC2 gives two warnings, both delivered as EventBridge events and as instance metadata, and both documented as best effort.
The rebalance recommendation, EventBridge detail-type EC2 Instance Rebalance Recommendation and metadata path /latest/meta-data/events/recommendations/rebalance, says the instance is at elevated risk. It can come well before an interruption, or at the same moment as the notice, or for an instance that is never interrupted. The right response is cheap and reversible: stop accepting new work and let a replacement start.
The interruption notice, detail-type EC2 Spot Instance Interruption Warning and metadata path /latest/meta-data/spot/instance-action, arrives about two minutes before EC2 stops, hibernates or terminates the instance. The metadata returns JSON such as {"action": "terminate", "time": "2026-10-11T08:22:00Z"} and a 404 when no action is pending. With hibernation as the interruption behaviour there is no two-minute lead: hibernation begins immediately. AWS recommends polling every five seconds. An on-instance agent using IMDSv2 looks like this:
import json, subprocess, time, urllib.request, urllib.error
IMDS = "http://169.254.169.254/latest"
def token():
req = urllib.request.Request(f"{IMDS}/api/token", method="PUT",
headers={"X-aws-ec2-metadata-token-ttl-seconds": "300"})
return urllib.request.urlopen(req, timeout=2).read().decode()
def get(path, tok):
req = urllib.request.Request(f"{IMDS}/meta-data/{path}",
headers={"X-aws-ec2-metadata-token": tok})
try:
return urllib.request.urlopen(req, timeout=2).read().decode()
except urllib.error.HTTPError as e:
if e.code == 404:
return None # no signal yet
raise
draining = False
while True:
tok = token()
if not draining and get("events/recommendations/rebalance", tok):
subprocess.run(["/opt/worker/bin/stop-accepting-work"], check=False)
draining = True
action = get("spot/instance-action", tok)
if action:
print("interruption:", json.loads(action))
subprocess.run(["/opt/worker/bin/checkpoint-and-exit"], timeout=90, check=False)
break
time.sleep(5)Fleet-wide reactions belong in EventBridge rules instead: route both detail-types to a queue or function that removes the instance from a scheduler, records the interruption for capacity analysis, or alerts when one pool is reclaimed repeatedly. Rule patterns are covered in Amazon EventBridge.
Capacity Rebalancing in Auto Scaling
Auto Scaling's Capacity Rebalancing automates the early reaction. When an instance in the group gets a rebalance recommendation, the group launches a replacement, waits for it to pass health checks, then terminates the at-risk instance, running any termination lifecycle hook first. To do that near the group's maximum size it may temporarily exceed the maximum by up to 10 percent of desired capacity. It only launches the replacement if the new pool's availability is the same or better, so it does not chase churn; if no better pool exists, the old instance runs until it is interrupted and the group replaces it reactively.
Two details catch people. Lifecycle hooks must finish well inside two minutes, because the recommendation can arrive together with the notice. And instances behind a load balancer are deregistered before the hook runs; Application Load Balancer target groups default to a 300 second deregistration delay, longer than the notice, so shorten it for Spot-backed targets.
Worked example: what an interruption costs
A rendering team runs 40 Spot workers, each processing frames that take about 20 minutes. Spot costs them, say, 30 percent of On-Demand in their pools; check your own current prices rather than relying on that figure. Their history shows about 6 interruptions a day across the fleet.
Without checkpointing, an interrupted frame loses on average half its runtime, 10 minutes, so the fleet wastes about 60 worker-minutes a day out of 57,600, roughly 0.1 percent. Interruptions barely matter here; the work unit is short. Now take a training-style job whose unit is 6 hours with no checkpoints: each interruption loses about 3 hours, 18 hours a day, over 3 percent of capacity, and worse, a job interrupted repeatedly may never finish. The cure is not On-Demand; it is checkpointing. With a checkpoint every 15 minutes, expected loss falls to 7.5 minutes per interruption, plus restore time.
The general rule: expected waste per interruption is roughly half the checkpoint interval plus restart cost. Make that small relative to how often your pools are reclaimed and Spot's discount dominates. Managed schedulers help; AWS Batch retries interrupted jobs for you, and GPU-specific considerations are covered in GPU Spot instances.
Stop, hibernate or terminate
The default interruption behaviour is terminate. Stop and hibernate keep the EBS root volume and restart the instance when capacity returns, which suits workloads that need their disk; they require an EBS-backed instance and a persistent Spot request or a fleet configured to maintain capacity. They do not make the workload highly available: the instance is gone until EC2 decides to start it again. For most designs, terminate plus externalised state, in S3, EFS, a database or a queue, is simpler and recovers faster because the replacement can start in any healthy pool.
Failure modes
- One pool, one failure. A group with a single instance type in one zone loses all Spot capacity together. Signature: capacity drops to the On-Demand base at once. Fix: five or more types across every zone you use.
- Lowest-price churn. Replacements land in the pool about to be reclaimed and the fleet is interrupted in waves. Fix:
price-capacity-optimized. - Unfulfillable launches. Auto Scaling activity history shows repeated failures for insufficient Spot capacity, or requests fail with a Spot vCPU quota error. Fix: widen types with attribute-based selection, re-check placement scores, and raise the Spot vCPU service quota.
- Lost local state. Results written to instance store or a terminated root volume vanish. Fix: write outputs to durable storage as they complete.
- Handlers that never fire. The agent polls IMDSv1 on an instance that requires IMDSv2, or the hook outlives the notice. Fix: test with AWS Fault Injection Service, which can send real interruption notices to Spot instances.
- Misread placement scores. A team scores three types with
capacity-optimizedthen launches eight types withlowest-priceand wonders why the score lied. The score only describes the exact request scored.
Trade-offs
Spot fits stateless web tiers behind load balancers, CI runners, batch and rendering, big data workers, and checkpointed training. It fits badly with single-instance databases, long jobs that cannot checkpoint, licensed software tied to a host, and anything needing capacity guaranteed at a particular hour; for that, On-Demand Capacity Reservations or Savings Plans on On-Demand are the tools. The common pattern mixes them: a committed baseline for what must always run and Spot for the elastic rest. Broader multi-cloud capacity strategy is discussed in Spot capacity across clouds.
What to do next
- List every workload on Spot or a candidate for it, with its unit of work and checkpoint interval.
- Run Spot placement scores for each fleet with at least three, ideally five or more, instance types.
- Switch allocation strategies to price-capacity-optimized and remove any low maximum prices.
- Enable Capacity Rebalancing on Auto Scaling groups and keep lifecycle hooks under two minutes.
- Deploy an IMDSv2 interruption agent that drains on rebalance and checkpoints on notice.
- Route both EventBridge detail-types to a queue for alerting and per-pool interruption metrics.
- Shorten load balancer deregistration delay for Spot targets.
- Rehearse an interruption with AWS Fault Injection Service before relying on any of it.