A spot instance is a lease on capacity the cloud has not sold at full price, offered at a large discount on the condition that the provider can take it back. AWS, Google Cloud and Azure all advertise discounts of up to about 90 percent on their spot offerings. On GPUs, the most expensive instances in the cloud, that discount is the biggest single price lever available. It is also the one most likely to backfire.

This article explains what you are buying, how each cloud warns you before taking it back, how to calculate whether spot is cheaper once interruptions are counted, and how to provision and train on it. Checkpoint mechanics inside a job are covered in GPU preemption handling; here the focus is the capacity market.

Advertisement

What you are buying, cloud by cloud

A cloud region holds more GPUs than are rented at any moment, because it must absorb growth and failures. Spot sells that slack, and the provider keeps the right to reclaim it for a full-price customer. You are not buying a cheaper GPU; you are buying a GPU whose lifetime is a random variable.

The three clouds package this differently:

  • AWS EC2 Spot. There is no auction to win. Spot prices move gradually with long-term supply and demand, you pay the current price, and the optional maximum price defaults to the on-demand price. Interruptions are driven by capacity, so bidding high does not protect you. Before an interruption EC2 sends a two-minute notice, and it also emits a rebalance recommendation when an instance is at elevated risk. The recommendation can arrive early or together with the notice; you cannot count on the extra time.
  • Google Cloud Spot VMs. Spot VMs are the successor to preemptible VMs, which are capped at 24 hours; Spot VMs have no maximum runtime. On preemption the VM gets a best-effort shutdown period of up to 30 seconds, and you choose whether it is stopped or deleted. A longer, 120-second notice duration is in Preview at the time of writing, so treat it as subject to change.
  • Azure Spot Virtual Machines. You choose an eviction policy (deallocate or delete) and an eviction type: capacity only, or capacity and price, with a maximum price where -1 means never evict for price. Notice arrives through Azure Scheduled Events, best effort, up to 30 seconds before eviction.

The practical consequence is that your shutdown budget is between zero and two minutes depending on the cloud, and the code must work with zero. Anything that must survive an interruption has to be written somewhere durable before the notice, not because of it.

Why GPU spot is harder than CPU spot

Spot advice written for stateless web servers transfers poorly to GPUs, for four reasons.

  1. Capacity is scarcer and lumpier. GPU instance types are large, few per zone and in demand, so spot capacity for the newest accelerators is often thin or absent. Check availability per type and zone before designing around it.
  2. An interruption takes a whole node. Losing one CPU instance from fifty costs 2 percent of capacity. Losing an eight-GPU node from a two-node training job costs half the job, and stops the other half too.
  3. Startup is slow. A replacement node must boot, pull a multi-gigabyte image, load drivers, fetch weights and warm up; ten to twenty minutes is common. That restart cost often dominates the bill.
  4. Synchronous training is gang-scheduled. Ranks exchange tensors through collectives every step (see NCCL collectives), so if one rank vanishes the rest block until a timeout. An N-node job is interrupted whenever any node is: roughly N times the per-node rate.
Cloud capacity poolspare GPUs, per zoneReclaim signalrebalance hint, noticeNode provisionerspot first, on-demand fallbackNotice watcherpolls instance metadataGPU workerstraining or inferenceTrainer loopchecks stop flag per stepObject storagecheckpoints, staged shardsElastic launcherre-rendezvous at new sizeReplacement nodeother type or zoneleaseschedulesignalset flagrank lostsaverestoreSpot is a lease the cloud can end: everything here exists to make the end cheap
The moving parts of a spot GPU fleet. The provisioner leases spare capacity, a watcher turns the cloud's reclaim signal into a flag the training loop checks, the loop saves to object storage, and an elastic launcher continues on whatever nodes remain while a replacement node comes up, possibly of a different type or in a different zone.
Advertisement

The real price: cost per useful hour

To decide whether spot is cheaper you need the price per hour of useful work, not per hour of rented time. Model a job with these quantities: interruption rate lam per hour for the whole job, checkpoint interval T hours, blocking checkpoint cost C hours, and restart cost R hours from interruption to resumed training.

Two kinds of waste appear. Checkpointing costs C / T of every hour. Each interruption loses on average half an interval of work plus the restart, T/2 + R, and happens lam times per hour. So the wasted fraction is roughly C/T + lam * (T/2 + R). Minimising it over T gives the classic Young approximation T* = sqrt(2C / lam). The approximation assumes waste is small; when it is not, treat the result as a lower bound on the damage. The effective price is the spot price divided by the useful fraction.

from math import sqrt

def spot_effective_price(spot_rel, lam_node, nodes, ckpt_min, restart_min):
    lam = lam_node * nodes                 # gang job: any node lost stops the job
    C, R = ckpt_min / 60, restart_min / 60
    T = sqrt(2 * C / lam)                  # Young's optimal checkpoint interval
    waste = C / T + lam * (T / 2 + R)
    useful = max(1e-9, 1 - waste)
    return T * 60, useful, spot_rel / useful

# Spot price at 0.35 of on-demand, 0.1 interruptions per node-hour,
# 2-minute blocking checkpoint, 15 minutes to get a replacement training again.
for nodes in (1, 8):
    T, useful, eff = spot_effective_price(0.35, 0.1, nodes, 2, 15)
    print(f"{nodes} node(s): checkpoint every {T:.0f} min, useful {useful:.0%}, "
          f"effective price {eff:.2f} x on-demand")

With these assumed inputs, a single node checkpoints every 49 minutes, keeps about 89 percent of its time useful and costs about 0.39 of on-demand per useful hour. The same settings across eight nodes raise the job's interruption rate to 0.8 per hour; the optimal interval drops to 17 minutes, useful time falls to about 57 percent and the effective price rises to about 0.61. Spot still wins, but the discount has nearly halved, and a slightly higher interruption rate or a slower restart erases it.

Measure the inputs rather than guessing: interruption rate per instance type over weeks, and restart time from notice to first completed step. The biggest levers are restart time (smaller images, cached weights, warm standby) and checkpoint cost (asynchronous checkpointing).

Catching the reclaim signal

Every cloud exposes its signal through the instance metadata service. A small watcher process, or a thread in the trainer, polls it and turns the signal into a flag the training loop checks between steps. The loop, not the signal handler, does the saving, so it never interrupts a half-finished optimizer step.

import json, threading, time, urllib.request

def _get(url, headers=None, method="GET", data=None):
    req = urllib.request.Request(url, headers=headers or {}, method=method, data=data)
    try:
        with urllib.request.urlopen(req, timeout=1) as r:
            return r.status, r.read().decode()
    except urllib.error.HTTPError as e:
        return e.code, ""
    except OSError:
        return None, ""

def aws_reclaim_pending():
    _, tok = _get("http://169.254.169.254/latest/api/token", method="PUT", data=b"",
                  headers={"X-aws-ec2-metadata-token-ttl-seconds": "300"})
    h = {"X-aws-ec2-metadata-token": tok}
    base = "http://169.254.169.254/latest/meta-data/"
    s1, _ = _get(base + "spot/instance-action", h)                   # 404 until notice
    s2, _ = _get(base + "events/recommendations/rebalance", h)       # 404 until risk
    return s1 == 200 or s2 == 200

def gcp_reclaim_pending():
    _, body = _get("http://metadata.google.internal/computeMetadata/v1/instance/preempted",
                   {"Metadata-Flavor": "Google"})
    return body.strip() == "TRUE"

def azure_reclaim_pending():
    s, body = _get("http://169.254.169.254/metadata/scheduledevents?api-version=2020-07-01",
                   {"Metadata": "true"})
    events = json.loads(body).get("Events", []) if s == 200 and body else []
    return any(e.get("EventType") == "Preempt" for e in events)

class ReclaimWatcher(threading.Thread):
    def __init__(self, probe, period_s=5):
        super().__init__(daemon=True)
        self.probe, self.period_s, self.stop = probe, period_s, threading.Event()
    def run(self):
        while not self.stop.is_set():
            if self.probe():
                self.stop.set()
            time.sleep(self.period_s)

# In the training loop:
#   watcher = ReclaimWatcher(aws_reclaim_pending); watcher.start()
#   for step, batch in enumerate(loader, start=resume_step):
#       train_step(batch)
#       if watcher.stop.is_set():
#           save_checkpoint(step, fast=True)   # flush what is already staged
#           sys.exit(143)                      # let the launcher treat it as a reclaim

On GCP you can block with ?wait_for_change=true instead of polling. On Kubernetes a termination handler usually watches these signals and drains the node, so the container gets SIGTERM; set the same flag from a SIGTERM handler and give the pod a grace period longer than one step.

Choosing where to ask for capacity

The interruption rate is not fixed; you choose much of it through where you ask for capacity. The rule is to be flexible on everything the workload does not care about.

  • Diversify instance types. A 7B inference server runs on several GPU types; list them all. On AWS the price-capacity-optimized allocation strategy picks pools that are deep as well as cheap, and the spot placement score API rates the odds of a request succeeding per region or zone.
  • Diversify zones. Spot pools are per zone. Inference replicas tolerate any zone; training wants one zone for interconnect bandwidth, so diversify across jobs, not within one.
  • Keep each training job homogeneous. In a synchronous job the slowest GPU sets the pace.
  • Replace early. AWS capacity rebalancing launches a replacement on the rebalance recommendation, before the notice.
  • Keep an on-demand floor. Run the replicas your SLO needs on on-demand or reserved capacity and add spot above them.

Kubernetes: spot first, on-demand floor

On Kubernetes the provisioner encodes those rules. With Karpenter on AWS, a node pool can allow both capacity types; Karpenter prefers spot and falls back to on-demand when spot cannot be obtained. Its interruption handling consumes EC2 interruption and rebalance events from an SQS queue fed by EventBridge, then cordons and drains affected nodes ahead of reclamation. The exact setting name for the queue has changed between releases, so take it from the docs for the version you install.

apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: gpu-inference
spec:
  template:
    spec:
      nodeClassRef: {group: karpenter.k8s.aws, kind: EC2NodeClass, name: gpu}
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"]          # spot preferred, on-demand fallback
        - key: karpenter.k8s.aws/instance-family
          operator: In
          values: ["g5", "g6"]                   # every family the model runs on
      taints:
        - key: nvidia.com/gpu
          effect: NoSchedule
  disruption:
    consolidationPolicy: WhenEmpty
    consolidateAfter: 5m

Put the on-demand floor in a second pool restricted to on-demand and pin critical replicas to it by node affinity. A PodDisruptionBudget cannot stop a cloud reclaiming a node, so the floor must come from capacity type. Queueing and packing GPU jobs onto these nodes is covered in GPU scheduling.

Elastic training across a changing fleet

Elastic training changes the multi-node arithmetic. Instead of the whole job stopping when a node disappears, PyTorch's torchrun can run with a node range, restart the surviving workers, re-rendezvous at the new world size and continue from the latest checkpoint:

torchrun \
  --nnodes=2:4 --nproc-per-node=8 \
  --max-restarts=20 \
  --rdzv-backend=c10d --rdzv-endpoint=trainer-0.trainer:29400 --rdzv-id=run-0412 \
  train.py --ckpt s3://ckpts/run-0412/

Elastic launch restarts processes; it does not migrate state. The script must load the newest complete checkpoint on every start and must handle a changed world size. Keep the global batch constant by recomputing gradient accumulation, otherwise the learning-rate schedule silently changes when the fleet shrinks:

world = torch.distributed.get_world_size()
per_gpu, global_batch = 8, 1024
accum = global_batch // (per_gpu * world)
assert accum * per_gpu * world == global_batch, "choose sizes that divide evenly"
state = load_latest_checkpoint(args.ckpt)        # model, optimizer, step, sampler position
sampler = ResumableSampler(dataset, world, rank, epoch=state.epoch,
                           start=state.samples_seen)   # your own sampler: offsets its indices

Per-rank checkpoint files usually cannot load at a different world size; formats that store a global view of each tensor can. Test a shrink and a grow.

Which workloads belong on spot

WorkloadSpot fitWhy
Batch inference, embeddings, evaluationExcellentIndependent work items; a lost node loses only items in flight
Online inference above an on-demand floorGoodReplicas are interchangeable; load balancer drains the lost one
Single-node fine-tuningGoodOne node, cheap restarts, checkpoint every tens of minutes
Elastic multi-node trainingFairWorks if restart is fast and the script handles world-size changes
Large synchronous pretrainingPoorInterruption rate scales with nodes; reserved capacity is usually cheaper per useful hour
Interactive notebooksPoorLost state annoys people more than it costs money

Failure modes

  • The checkpoint that never finished. A save cut off by reclamation gets loaded on resume. Write to a temporary path, publish a manifest last, and load only manifested checkpoints.
  • Restart loops. Capacity vanished in the zone and every replacement is reclaimed within minutes. Cap restarts per hour, then fall back to on-demand or pause.
  • The silent stall. A rank disappears and the others wait in a collective until a long timeout. Shorten the timeout and alert on steps per minute, not process liveness; GPU hardware faults covers the same stall from a failing GPU.
  • Image and weight downloads. Re-fetching tens of gigabytes per replacement costs time. Cache in-region and bake drivers into the node image.
  • Accounting that ignores waste. Dashboards showing only the discount overstate savings; report effective price per useful hour.

What to do next

  1. Place each GPU workload in the fit table; start with batch inference or evaluation.
  2. List every GPU type and zone the workload can use, and check spot availability for each.
  3. Add a reclaim watcher or termination handler; save and exit on its flag at a step boundary.
  4. Make checkpoints atomic with a manifest written last.
  5. Measure restart time from notice to first completed step, and cut it with baked images and cached weights.
  6. Run the cost model with measured inputs against on-demand and reserved prices.
  7. Keep SLO-critical replicas on demand; for training, try torchrun --nnodes=min:max and test a shrink and a grow.
  8. Report effective price per useful hour monthly, next to the levers in GPU cost optimization.
Key takeaway: A spot GPU is a lease the cloud can end with two minutes of notice on AWS and up to 30 seconds on Google Cloud and Azure, so state must be durable before the notice. Judge spot by price per useful hour, which interruption rate, checkpoint cost and restart time decide; a synchronous N-node job multiplies its interruption rate by N. Diversify types and zones, keep SLO-critical capacity on demand, make checkpoints atomic and use elastic launch for multi-node training.