Spot VMs are Compute Engine VMs that run on spare capacity at a large discount. Google's documentation puts the discount at up to 91 percent for many machine types, GPUs, TPUs and Local SSDs. The price of the discount is that Compute Engine can take the VM back whenever it needs the capacity, with very little warning. Preemptible VMs are the older version of the same offer, with one extra rule: they always stop after 24 hours.
Used well, Spot turns batch processing, CI, rendering and ML training into a fraction of their on-demand cost. Used badly, it turns them into jobs that never finish. The difference is almost entirely engineering inside the workload: knowing exactly what happens in the seconds after preemption, checkpointing at the right interval, and letting a managed group or a Kubernetes scheduler replace lost capacity. This page covers that lifecycle step by step, with commands, a preemption watcher, the checkpoint arithmetic, a worked training example, and the cases where Spot is the wrong choice. The Compute Engine basics are in Compute Engine in depth; the scheduling block that holds these settings is in GCE instance scheduling.
The preemption sequence
Standard, Spot and preemptible compared
Compute Engine has three provisioning choices relevant here. Spot replaced preemptible as the recommended model; new work should use Spot.
| Standard | Spot | Preemptible (legacy) | |
|---|---|---|---|
| Price | On-demand list price | Discounted; prices can change up to once a day | Discounted |
| Can be reclaimed | No | Any time | Any time |
| Maximum runtime | None | None unless you set a time limit | 24 hours |
| Live migration on maintenance | Usually yes | No | No |
| Automatic restart | On by default | Not supported | Not supported |
| Compute Engine SLA | Covered | Excluded | Excluded |
| gcloud selector | --provisioning-model=STANDARD | --provisioning-model=SPOT | --preemptible |
Two points from that table matter most in practice. Spot capacity is not guaranteed even at creation time, so a request can fail with a resource-availability error in a zone that is busy. And preemption is not a pricing auction: you do not bid, and a higher willingness to pay does not protect a VM. The only levers are where you ask for capacity and how your workload behaves when it is taken away. Some machine types, including A4X and bare metal instances, are not offered as Spot; check the documentation for the shape you need.
The lifecycle in detail
The documented sequence is short. Compute Engine first sets the preempted metadata value to TRUE. It then waits for the preemption notice duration, which is zero by default, so normally there is no gap at all. A 120-second notice duration is available in Preview, set at creation with --preemption-notice-duration=120s on gcloud beta compute instances create; it is recommended for workloads that need more than 30 seconds to react. Next comes an ACPI G2 soft off signal, which makes the guest operating system begin a normal shutdown and run any shutdown script. That shutdown period is best effort and up to 30 seconds. If the VM is still running at the end, Compute Engine sends ACPI G3 mechanical off, which is the equivalent of pulling the plug.
After that the instance termination action applies. With STOP, the default, the VM moves to TERMINATED, its persistent disks remain, and you pay for disks but not for VM time; you can start it again later if capacity exists. With DELETE, the VM is removed. DELETE suits stateless workers in managed groups, where nothing on the VM is worth keeping and a stopped VM would only clutter the project. Local SSD contents do not survive preemption in either case.
The practical reading: with the default settings you have up to about 30 seconds of best-effort time, and you may get less. Design for zero, and treat the window as a bonus for flushing small state.
Creating and auditing Spot VMs
A Spot worker that deletes itself on preemption and registers a shutdown script:
gcloud compute instances create batch-worker-1 \
--zone=us-central1-a \
--machine-type=n2-standard-8 \
--image-family=debian-12 --image-project=debian-cloud \
--provisioning-model=SPOT \
--instance-termination-action=DELETE \
--metadata-from-file=shutdown-script=shutdown.shIn an instance template the same flags go on gcloud compute instance-templates create, and a managed instance group built from the template will try to recreate preempted members, which makes the group the natural unit for Spot fleets. Use a regional MIG across several zones so a capacity shortage in one zone does not take the whole fleet. To find out after the fact which VMs were preempted rather than stopped by someone, list the preemption operations:
gcloud compute operations list \
--filter="operationType=compute.instances.preempted" \
--format="table(targetLink.basename(), zone.basename(), insertTime)"Cloud Audit Logs record the same event with the method name compute.instances.preempted, which is the better source for dashboards and alerts on preemption rate per zone and machine type.
Detecting preemption from inside the VM
Shutdown scripts run at the G2 step, which is fine for small cleanup, but a process inside the VM can learn about preemption earlier by watching the metadata value itself. The metadata server supports hanging GETs: with wait_for_change=true the request blocks until the value changes, so the watcher reacts as soon as the flag flips instead of polling. This matters most when a notice duration is set, because the whole notice window sits between the flag and G2.
import signal, threading, requests
URL = "http://metadata.google.internal/computeMetadata/v1/instance/preempted"
HDR = {"Metadata-Flavor": "Google"}
preempting = threading.Event()
def watch():
etag = None
while not preempting.is_set():
params = {"wait_for_change": "true", "timeout_sec": "300"}
if etag:
params["last_etag"] = etag
try:
r = requests.get(URL, headers=HDR, params=params, timeout=310)
except requests.RequestException:
continue # metadata server hiccup: retry
etag = r.headers.get("ETag")
if r.text.strip() == "TRUE":
preempting.set()
threading.Thread(target=watch, daemon=True).start()
signal.signal(signal.SIGTERM, lambda *_: preempting.set()) # G2 shutdown also lands here
def train(state, loader, ckpt):
for step, batch in enumerate(loader, start=state.step):
state = update(state, batch)
if step % ckpt.every_steps == 0:
ckpt.save_async(state) # regular checkpoints carry the real load
if preempting.is_set():
ckpt.save_marker(state.step) # tiny write: where to resume, not weights
ckpt.wait(timeout=20) # let an in-flight upload finish if it can
returnThe shape is the important part. The watcher only sets a flag. The training loop checks the flag between steps and does the least possible work: it records the step and waits briefly for an upload already in progress. It does not try to write a multi-gigabyte checkpoint in a window that may be zero seconds long.
Checkpoint arithmetic
How often to checkpoint is a trade between two kinds of waste: time spent writing checkpoints, and work lost since the last one when preemption hits. A standard first-order answer is Young's approximation, derived together with cost per useful hour across clouds in GPU spot instances. Here it is only the tool for feeding GCP measurements into a decision:
T_opt ~= sqrt(2 * C * M)
C = time to write one checkpoint (seconds of lost training)
M = mean time between preemptions for this VM shape and zone
Expected overhead per unit of work ~= C / T + T / (2 * M) + R / M
R = restart cost: reprovision, start containers, load checkpoint, warm upThe terms are checkpoint writing, the average half-interval lost per preemption, and the fixed cost of each restart. With asynchronous checkpointing, C is only the time the loop blocks. Measure M yourself from the preemption operations log; Google does not publish preemption rates, and they vary by zone, machine shape and time.
Worked example: a 40-hour fine-tuning job
Suppose a fine-tuning job needs 40 hours of GPU time on one accelerator VM. Checkpoints block training for 60 seconds. From the last month of logs, the shape you want sees a preemption every 6 hours on average in your chosen zones, and a restart costs 10 minutes.
Young's interval is the square root of 2 times 60 times 21,600 seconds, about 1,610 seconds, so checkpoint roughly every 27 minutes. Overhead is then about 3.7 percent for writing, 3.7 percent for lost work and 2.8 percent for restarts: just over 10 percent. The job needs about 44 hours of Spot time instead of 40.
Now the price. Spot prices vary by machine type and region and can change daily, so take the real figure from the pricing page. For illustration, if Spot costs 35 percent of on-demand for this shape, 44 Spot hours cost the same as about 15.4 on-demand hours, against 40 for an on-demand run: roughly 60 percent saved. If preemptions became five times as frequent, the interval would shrink to about 12 minutes and overhead would rise to roughly 31 percent; the run would still be cheaper, but the restart term, not the discount, would now be the thing to optimise. Spot also needs a time margin, because an availability shortage can delay restarts for hours.
Spot node pools on GKE
On GKE, create a node pool with --spot. GKE labels the nodes cloud.google.com/gke-spot=true and, from version 1.25.5, cloud.google.com/gke-provisioning=spot. Use a node selector or affinity on that label to place batch work there, and a taint plus matching tolerations to keep everything else off. Keep at least one standard node pool for system components such as DNS and for anything that must not be interrupted.
Preemption on GKE has its own timing. By default the node gets a 30-second graceful termination period: regular Pods get 15 seconds to stop, then critical system Pods get the remaining 15. Setting terminationGracePeriodSeconds higher in a Pod spec does not extend this. Newer GKE versions allow the period to be raised up to 120 seconds through node system configuration on Standard pools (1.35.0 and later) or ComputeClass kubelet settings (1.36.0 and later). For stateful workloads Google advises testing that they terminate gracefully within 25 seconds of shutdown, to limit the risk of persistent volume corruption; with the default split, only 15 of those seconds belong to your Pods, so either extend the period or keep the shutdown work tiny. Use Jobs or an operator that resumes from checkpoints, and PodDisruptionBudgets for replicated services; a budget does not stop a preemption, but it shapes voluntary disruptions around it. Cluster-level patterns are in GKE in depth and scaling behaviour in the MIG autoscaler.
ML training on Spot accelerators
Accelerator VMs already cannot live migrate, so a GPU training job needs checkpointing whether it runs on Spot or not; Spot mostly changes how often the restart path is exercised. For multi-node synchronous training, one preempted node stops every step, so the effective preemption rate is roughly the per-node rate times the node count; elastic training is covered on the GPU spot page. Spot TPUs follow the same logic; TPU v5 and v6 covers the hardware side.
Keep checkpoints and datasets in Cloud Storage in the same region as the compute, so a replacement VM in another zone can resume without cross-region transfer. Write checkpoints to a temporary name and rename on completion, so a G3 cut mid-upload cannot leave a half-written file that looks valid.
Failure modes
| Failure | Cause | Fix |
|---|---|---|
| Job restarts from scratch | Checkpoints kept on local SSD or boot disk | Write to Cloud Storage; resume from the latest complete one |
| Corrupt checkpoint after preemption | Upload cut by G3 mid-write | Write to a temporary name, then rename |
| Fleet never recovers | Single-zone group in a busy zone | Regional MIG or several zones; on-demand fallback pool |
| Graceful shutdown never finishes | Full checkpoint attempted in the 30 s window | Save regularly; only a marker at preemption |
| Stopped VMs accumulate | Default STOP on stateless workers | Use DELETE for workers in groups |
| Critical service interrupted | Untainted Spot node pool | Taints, tolerations and a standard pool |
When Spot is the wrong choice
Spot is the wrong choice for anything with a latency SLO and no spare capacity elsewhere, for long synchronous jobs on many nodes without elastic recovery, for databases without replicas on standard VMs, and for work with a hard deadline that cannot absorb hours of unavailability. It is the right choice for queues of independent tasks, CI runners, rendering, hyperparameter sweeps, and training jobs whose restart path is tested. Many production fleets mix the two: a standard baseline sized for the minimum acceptable throughput and a Spot layer on top that absorbs the rest. Committed use discounts apply to steady standard capacity, so the baseline and the Spot layer are priced by different mechanisms.
What to do next
- Pick one batch workload and confirm it can be killed at any second and resume from external state. Kill it on purpose to prove it.
- Move checkpoints and outputs to Cloud Storage with write-then-rename.
- Add the metadata watcher and a short shutdown script; never attempt a full save in the shutdown window.
- Run it on a regional MIG or a tainted GKE Spot pool, with DELETE for stateless workers.
- Build a dashboard of preemptions per zone and machine type from the audit log, and compute your own mean time between preemptions.
- Set the checkpoint interval with Young's formula from measured C, M and R, and revisit it monthly.
- Keep a standard pool or on-demand fallback for the minimum throughput you cannot lose.