A spot instance is a lease on capacity the cloud has not sold at full price, offered at a large discount on the condition that the provider can take it back. AWS, Google Cloud and Azure all advertise discounts of up to about 90 percent on their spot offerings. On GPUs, the most expensive instances in the cloud, that discount is the biggest single price lever available. It is also the one most likely to backfire.
This article explains what you are buying, how each cloud warns you before taking it back, how to calculate whether spot is cheaper once interruptions are counted, and how to provision and train on it. Checkpoint mechanics inside a job are covered in GPU preemption handling; here the focus is the capacity market.
What you are buying, cloud by cloud
A cloud region holds more GPUs than are rented at any moment, because it must absorb growth and failures. Spot sells that slack, and the provider keeps the right to reclaim it for a full-price customer. You are not buying a cheaper GPU; you are buying a GPU whose lifetime is a random variable.
The three clouds package this differently:
- AWS EC2 Spot. There is no auction to win. Spot prices move gradually with long-term supply and demand, you pay the current price, and the optional maximum price defaults to the on-demand price. Interruptions are driven by capacity, so bidding high does not protect you. Before an interruption EC2 sends a two-minute notice, and it also emits a rebalance recommendation when an instance is at elevated risk. The recommendation can arrive early or together with the notice; you cannot count on the extra time.
- Google Cloud Spot VMs. Spot VMs are the successor to preemptible VMs, which are capped at 24 hours; Spot VMs have no maximum runtime. On preemption the VM gets a best-effort shutdown period of up to 30 seconds, and you choose whether it is stopped or deleted. A longer, 120-second notice duration is in Preview at the time of writing, so treat it as subject to change.
- Azure Spot Virtual Machines. You choose an eviction policy (deallocate or delete) and an eviction type: capacity only, or capacity and price, with a maximum price where
-1means never evict for price. Notice arrives through Azure Scheduled Events, best effort, up to 30 seconds before eviction.
The practical consequence is that your shutdown budget is between zero and two minutes depending on the cloud, and the code must work with zero. Anything that must survive an interruption has to be written somewhere durable before the notice, not because of it.
Why GPU spot is harder than CPU spot
Spot advice written for stateless web servers transfers poorly to GPUs, for four reasons.
- Capacity is scarcer and lumpier. GPU instance types are large, few per zone and in demand, so spot capacity for the newest accelerators is often thin or absent. Check availability per type and zone before designing around it.
- An interruption takes a whole node. Losing one CPU instance from fifty costs 2 percent of capacity. Losing an eight-GPU node from a two-node training job costs half the job, and stops the other half too.
- Startup is slow. A replacement node must boot, pull a multi-gigabyte image, load drivers, fetch weights and warm up; ten to twenty minutes is common. That restart cost often dominates the bill.
- Synchronous training is gang-scheduled. Ranks exchange tensors through collectives every step (see NCCL collectives), so if one rank vanishes the rest block until a timeout. An N-node job is interrupted whenever any node is: roughly N times the per-node rate.
The real price: cost per useful hour
To decide whether spot is cheaper you need the price per hour of useful work, not per hour of rented time. Model a job with these quantities: interruption rate lam per hour for the whole job, checkpoint interval T hours, blocking checkpoint cost C hours, and restart cost R hours from interruption to resumed training.
Two kinds of waste appear. Checkpointing costs C / T of every hour. Each interruption loses on average half an interval of work plus the restart, T/2 + R, and happens lam times per hour. So the wasted fraction is roughly C/T + lam * (T/2 + R). Minimising it over T gives the classic Young approximation T* = sqrt(2C / lam). The approximation assumes waste is small; when it is not, treat the result as a lower bound on the damage. The effective price is the spot price divided by the useful fraction.
from math import sqrt
def spot_effective_price(spot_rel, lam_node, nodes, ckpt_min, restart_min):
lam = lam_node * nodes # gang job: any node lost stops the job
C, R = ckpt_min / 60, restart_min / 60
T = sqrt(2 * C / lam) # Young's optimal checkpoint interval
waste = C / T + lam * (T / 2 + R)
useful = max(1e-9, 1 - waste)
return T * 60, useful, spot_rel / useful
# Spot price at 0.35 of on-demand, 0.1 interruptions per node-hour,
# 2-minute blocking checkpoint, 15 minutes to get a replacement training again.
for nodes in (1, 8):
T, useful, eff = spot_effective_price(0.35, 0.1, nodes, 2, 15)
print(f"{nodes} node(s): checkpoint every {T:.0f} min, useful {useful:.0%}, "
f"effective price {eff:.2f} x on-demand")With these assumed inputs, a single node checkpoints every 49 minutes, keeps about 89 percent of its time useful and costs about 0.39 of on-demand per useful hour. The same settings across eight nodes raise the job's interruption rate to 0.8 per hour; the optimal interval drops to 17 minutes, useful time falls to about 57 percent and the effective price rises to about 0.61. Spot still wins, but the discount has nearly halved, and a slightly higher interruption rate or a slower restart erases it.
Measure the inputs rather than guessing: interruption rate per instance type over weeks, and restart time from notice to first completed step. The biggest levers are restart time (smaller images, cached weights, warm standby) and checkpoint cost (asynchronous checkpointing).
Catching the reclaim signal
Every cloud exposes its signal through the instance metadata service. A small watcher process, or a thread in the trainer, polls it and turns the signal into a flag the training loop checks between steps. The loop, not the signal handler, does the saving, so it never interrupts a half-finished optimizer step.
import json, threading, time, urllib.request
def _get(url, headers=None, method="GET", data=None):
req = urllib.request.Request(url, headers=headers or {}, method=method, data=data)
try:
with urllib.request.urlopen(req, timeout=1) as r:
return r.status, r.read().decode()
except urllib.error.HTTPError as e:
return e.code, ""
except OSError:
return None, ""
def aws_reclaim_pending():
_, tok = _get("http://169.254.169.254/latest/api/token", method="PUT", data=b"",
headers={"X-aws-ec2-metadata-token-ttl-seconds": "300"})
h = {"X-aws-ec2-metadata-token": tok}
base = "http://169.254.169.254/latest/meta-data/"
s1, _ = _get(base + "spot/instance-action", h) # 404 until notice
s2, _ = _get(base + "events/recommendations/rebalance", h) # 404 until risk
return s1 == 200 or s2 == 200
def gcp_reclaim_pending():
_, body = _get("http://metadata.google.internal/computeMetadata/v1/instance/preempted",
{"Metadata-Flavor": "Google"})
return body.strip() == "TRUE"
def azure_reclaim_pending():
s, body = _get("http://169.254.169.254/metadata/scheduledevents?api-version=2020-07-01",
{"Metadata": "true"})
events = json.loads(body).get("Events", []) if s == 200 and body else []
return any(e.get("EventType") == "Preempt" for e in events)
class ReclaimWatcher(threading.Thread):
def __init__(self, probe, period_s=5):
super().__init__(daemon=True)
self.probe, self.period_s, self.stop = probe, period_s, threading.Event()
def run(self):
while not self.stop.is_set():
if self.probe():
self.stop.set()
time.sleep(self.period_s)
# In the training loop:
# watcher = ReclaimWatcher(aws_reclaim_pending); watcher.start()
# for step, batch in enumerate(loader, start=resume_step):
# train_step(batch)
# if watcher.stop.is_set():
# save_checkpoint(step, fast=True) # flush what is already staged
# sys.exit(143) # let the launcher treat it as a reclaimOn GCP you can block with ?wait_for_change=true instead of polling. On Kubernetes a termination handler usually watches these signals and drains the node, so the container gets SIGTERM; set the same flag from a SIGTERM handler and give the pod a grace period longer than one step.
Choosing where to ask for capacity
The interruption rate is not fixed; you choose much of it through where you ask for capacity. The rule is to be flexible on everything the workload does not care about.
- Diversify instance types. A 7B inference server runs on several GPU types; list them all. On AWS the
price-capacity-optimizedallocation strategy picks pools that are deep as well as cheap, and the spot placement score API rates the odds of a request succeeding per region or zone. - Diversify zones. Spot pools are per zone. Inference replicas tolerate any zone; training wants one zone for interconnect bandwidth, so diversify across jobs, not within one.
- Keep each training job homogeneous. In a synchronous job the slowest GPU sets the pace.
- Replace early. AWS capacity rebalancing launches a replacement on the rebalance recommendation, before the notice.
- Keep an on-demand floor. Run the replicas your SLO needs on on-demand or reserved capacity and add spot above them.
Kubernetes: spot first, on-demand floor
On Kubernetes the provisioner encodes those rules. With Karpenter on AWS, a node pool can allow both capacity types; Karpenter prefers spot and falls back to on-demand when spot cannot be obtained. Its interruption handling consumes EC2 interruption and rebalance events from an SQS queue fed by EventBridge, then cordons and drains affected nodes ahead of reclamation. The exact setting name for the queue has changed between releases, so take it from the docs for the version you install.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: gpu-inference
spec:
template:
spec:
nodeClassRef: {group: karpenter.k8s.aws, kind: EC2NodeClass, name: gpu}
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"] # spot preferred, on-demand fallback
- key: karpenter.k8s.aws/instance-family
operator: In
values: ["g5", "g6"] # every family the model runs on
taints:
- key: nvidia.com/gpu
effect: NoSchedule
disruption:
consolidationPolicy: WhenEmpty
consolidateAfter: 5mPut the on-demand floor in a second pool restricted to on-demand and pin critical replicas to it by node affinity. A PodDisruptionBudget cannot stop a cloud reclaiming a node, so the floor must come from capacity type. Queueing and packing GPU jobs onto these nodes is covered in GPU scheduling.
Elastic training across a changing fleet
Elastic training changes the multi-node arithmetic. Instead of the whole job stopping when a node disappears, PyTorch's torchrun can run with a node range, restart the surviving workers, re-rendezvous at the new world size and continue from the latest checkpoint:
torchrun \
--nnodes=2:4 --nproc-per-node=8 \
--max-restarts=20 \
--rdzv-backend=c10d --rdzv-endpoint=trainer-0.trainer:29400 --rdzv-id=run-0412 \
train.py --ckpt s3://ckpts/run-0412/Elastic launch restarts processes; it does not migrate state. The script must load the newest complete checkpoint on every start and must handle a changed world size. Keep the global batch constant by recomputing gradient accumulation, otherwise the learning-rate schedule silently changes when the fleet shrinks:
world = torch.distributed.get_world_size()
per_gpu, global_batch = 8, 1024
accum = global_batch // (per_gpu * world)
assert accum * per_gpu * world == global_batch, "choose sizes that divide evenly"
state = load_latest_checkpoint(args.ckpt) # model, optimizer, step, sampler position
sampler = ResumableSampler(dataset, world, rank, epoch=state.epoch,
start=state.samples_seen) # your own sampler: offsets its indicesPer-rank checkpoint files usually cannot load at a different world size; formats that store a global view of each tensor can. Test a shrink and a grow.
Which workloads belong on spot
| Workload | Spot fit | Why |
|---|---|---|
| Batch inference, embeddings, evaluation | Excellent | Independent work items; a lost node loses only items in flight |
| Online inference above an on-demand floor | Good | Replicas are interchangeable; load balancer drains the lost one |
| Single-node fine-tuning | Good | One node, cheap restarts, checkpoint every tens of minutes |
| Elastic multi-node training | Fair | Works if restart is fast and the script handles world-size changes |
| Large synchronous pretraining | Poor | Interruption rate scales with nodes; reserved capacity is usually cheaper per useful hour |
| Interactive notebooks | Poor | Lost state annoys people more than it costs money |
Failure modes
- The checkpoint that never finished. A save cut off by reclamation gets loaded on resume. Write to a temporary path, publish a manifest last, and load only manifested checkpoints.
- Restart loops. Capacity vanished in the zone and every replacement is reclaimed within minutes. Cap restarts per hour, then fall back to on-demand or pause.
- The silent stall. A rank disappears and the others wait in a collective until a long timeout. Shorten the timeout and alert on steps per minute, not process liveness; GPU hardware faults covers the same stall from a failing GPU.
- Image and weight downloads. Re-fetching tens of gigabytes per replacement costs time. Cache in-region and bake drivers into the node image.
- Accounting that ignores waste. Dashboards showing only the discount overstate savings; report effective price per useful hour.
What to do next
- Place each GPU workload in the fit table; start with batch inference or evaluation.
- List every GPU type and zone the workload can use, and check spot availability for each.
- Add a reclaim watcher or termination handler; save and exit on its flag at a step boundary.
- Make checkpoints atomic with a manifest written last.
- Measure restart time from notice to first completed step, and cut it with baked images and cached weights.
- Run the cost model with measured inputs against on-demand and reserved prices.
- Keep SLO-critical replicas on demand; for training, try
torchrun --nnodes=min:maxand test a shrink and a grow. - Report effective price per useful hour monthly, next to the levers in GPU cost optimization.