Spot VMs are Compute Engine VMs that run on spare capacity at a large discount. Google's documentation puts the discount at up to 91 percent for many machine types, GPUs, TPUs and Local SSDs. The price of the discount is that Compute Engine can take the VM back whenever it needs the capacity, with very little warning. Preemptible VMs are the older version of the same offer, with one extra rule: they always stop after 24 hours.

Used well, Spot turns batch processing, CI, rendering and ML training into a fraction of their on-demand cost. Used badly, it turns them into jobs that never finish. The difference is almost entirely engineering inside the workload: knowing exactly what happens in the seconds after preemption, checkpointing at the right interval, and letting a managed group or a Kubernetes scheduler replace lost capacity. This page covers that lifecycle step by step, with commands, a preemption watcher, the checkpoint arithmetic, a worked training example, and the cases where Spot is the wrong choice. The Compute Engine basics are in Compute Engine in depth; the scheduling block that holds these settings is in GCE instance scheduling.

The preemption sequence

What happens when Compute Engine preempts a Spot VM1. preempted = TRUEmetadata server2. Notice duration0 s default, 120 s (Preview)3. ACPI G2 soft offshutdown scripts run4. Up to 30 sbest effort5. G3 offif still upt = 0watcher reacts hereOS shutdown beginspower cutSTOP (default)TERMINATED, disks kept, no VM chargeDELETEVM and boot disk removedInstance termination action decides the end state. A MIG or GKE then replaces lost capacity when Spot capacity is available.Anything not already persisted outside the VM before step 5 should be treated as lost.
Preemption sets the metadata flag, waits for the notice duration (zero by default), sends ACPI G2, allows up to 30 seconds of best-effort shutdown, then cuts power and applies the termination action.

Standard, Spot and preemptible compared

Compute Engine has three provisioning choices relevant here. Spot replaced preemptible as the recommended model; new work should use Spot.

StandardSpotPreemptible (legacy)
PriceOn-demand list priceDiscounted; prices can change up to once a dayDiscounted
Can be reclaimedNoAny timeAny time
Maximum runtimeNoneNone unless you set a time limit24 hours
Live migration on maintenanceUsually yesNoNo
Automatic restartOn by defaultNot supportedNot supported
Compute Engine SLACoveredExcludedExcluded
gcloud selector--provisioning-model=STANDARD--provisioning-model=SPOT--preemptible

Two points from that table matter most in practice. Spot capacity is not guaranteed even at creation time, so a request can fail with a resource-availability error in a zone that is busy. And preemption is not a pricing auction: you do not bid, and a higher willingness to pay does not protect a VM. The only levers are where you ask for capacity and how your workload behaves when it is taken away. Some machine types, including A4X and bare metal instances, are not offered as Spot; check the documentation for the shape you need.

The lifecycle in detail

The documented sequence is short. Compute Engine first sets the preempted metadata value to TRUE. It then waits for the preemption notice duration, which is zero by default, so normally there is no gap at all. A 120-second notice duration is available in Preview, set at creation with --preemption-notice-duration=120s on gcloud beta compute instances create; it is recommended for workloads that need more than 30 seconds to react. Next comes an ACPI G2 soft off signal, which makes the guest operating system begin a normal shutdown and run any shutdown script. That shutdown period is best effort and up to 30 seconds. If the VM is still running at the end, Compute Engine sends ACPI G3 mechanical off, which is the equivalent of pulling the plug.

After that the instance termination action applies. With STOP, the default, the VM moves to TERMINATED, its persistent disks remain, and you pay for disks but not for VM time; you can start it again later if capacity exists. With DELETE, the VM is removed. DELETE suits stateless workers in managed groups, where nothing on the VM is worth keeping and a stopped VM would only clutter the project. Local SSD contents do not survive preemption in either case.

The practical reading: with the default settings you have up to about 30 seconds of best-effort time, and you may get less. Design for zero, and treat the window as a bonus for flushing small state.

Creating and auditing Spot VMs

A Spot worker that deletes itself on preemption and registers a shutdown script:

gcloud compute instances create batch-worker-1 \
    --zone=us-central1-a \
    --machine-type=n2-standard-8 \
    --image-family=debian-12 --image-project=debian-cloud \
    --provisioning-model=SPOT \
    --instance-termination-action=DELETE \
    --metadata-from-file=shutdown-script=shutdown.sh

In an instance template the same flags go on gcloud compute instance-templates create, and a managed instance group built from the template will try to recreate preempted members, which makes the group the natural unit for Spot fleets. Use a regional MIG across several zones so a capacity shortage in one zone does not take the whole fleet. To find out after the fact which VMs were preempted rather than stopped by someone, list the preemption operations:

gcloud compute operations list \
    --filter="operationType=compute.instances.preempted" \
    --format="table(targetLink.basename(), zone.basename(), insertTime)"

Cloud Audit Logs record the same event with the method name compute.instances.preempted, which is the better source for dashboards and alerts on preemption rate per zone and machine type.

Detecting preemption from inside the VM

Shutdown scripts run at the G2 step, which is fine for small cleanup, but a process inside the VM can learn about preemption earlier by watching the metadata value itself. The metadata server supports hanging GETs: with wait_for_change=true the request blocks until the value changes, so the watcher reacts as soon as the flag flips instead of polling. This matters most when a notice duration is set, because the whole notice window sits between the flag and G2.

import signal, threading, requests

URL = "http://metadata.google.internal/computeMetadata/v1/instance/preempted"
HDR = {"Metadata-Flavor": "Google"}
preempting = threading.Event()

def watch():
    etag = None
    while not preempting.is_set():
        params = {"wait_for_change": "true", "timeout_sec": "300"}
        if etag:
            params["last_etag"] = etag
        try:
            r = requests.get(URL, headers=HDR, params=params, timeout=310)
        except requests.RequestException:
            continue                                    # metadata server hiccup: retry
        etag = r.headers.get("ETag")
        if r.text.strip() == "TRUE":
            preempting.set()

threading.Thread(target=watch, daemon=True).start()
signal.signal(signal.SIGTERM, lambda *_: preempting.set())   # G2 shutdown also lands here

def train(state, loader, ckpt):
    for step, batch in enumerate(loader, start=state.step):
        state = update(state, batch)
        if step % ckpt.every_steps == 0:
            ckpt.save_async(state)                      # regular checkpoints carry the real load
        if preempting.is_set():
            ckpt.save_marker(state.step)                # tiny write: where to resume, not weights
            ckpt.wait(timeout=20)                       # let an in-flight upload finish if it can
            return

The shape is the important part. The watcher only sets a flag. The training loop checks the flag between steps and does the least possible work: it records the step and waits briefly for an upload already in progress. It does not try to write a multi-gigabyte checkpoint in a window that may be zero seconds long.

Checkpoint arithmetic

How often to checkpoint is a trade between two kinds of waste: time spent writing checkpoints, and work lost since the last one when preemption hits. A standard first-order answer is Young's approximation, derived together with cost per useful hour across clouds in GPU spot instances. Here it is only the tool for feeding GCP measurements into a decision:

T_opt ~= sqrt(2 * C * M)

C = time to write one checkpoint (seconds of lost training)
M = mean time between preemptions for this VM shape and zone

Expected overhead per unit of work ~= C / T  +  T / (2 * M)  +  R / M
R = restart cost: reprovision, start containers, load checkpoint, warm up

The terms are checkpoint writing, the average half-interval lost per preemption, and the fixed cost of each restart. With asynchronous checkpointing, C is only the time the loop blocks. Measure M yourself from the preemption operations log; Google does not publish preemption rates, and they vary by zone, machine shape and time.

Worked example: a 40-hour fine-tuning job

Suppose a fine-tuning job needs 40 hours of GPU time on one accelerator VM. Checkpoints block training for 60 seconds. From the last month of logs, the shape you want sees a preemption every 6 hours on average in your chosen zones, and a restart costs 10 minutes.

Young's interval is the square root of 2 times 60 times 21,600 seconds, about 1,610 seconds, so checkpoint roughly every 27 minutes. Overhead is then about 3.7 percent for writing, 3.7 percent for lost work and 2.8 percent for restarts: just over 10 percent. The job needs about 44 hours of Spot time instead of 40.

Now the price. Spot prices vary by machine type and region and can change daily, so take the real figure from the pricing page. For illustration, if Spot costs 35 percent of on-demand for this shape, 44 Spot hours cost the same as about 15.4 on-demand hours, against 40 for an on-demand run: roughly 60 percent saved. If preemptions became five times as frequent, the interval would shrink to about 12 minutes and overhead would rise to roughly 31 percent; the run would still be cheaper, but the restart term, not the discount, would now be the thing to optimise. Spot also needs a time margin, because an availability shortage can delay restarts for hours.

Spot node pools on GKE

On GKE, create a node pool with --spot. GKE labels the nodes cloud.google.com/gke-spot=true and, from version 1.25.5, cloud.google.com/gke-provisioning=spot. Use a node selector or affinity on that label to place batch work there, and a taint plus matching tolerations to keep everything else off. Keep at least one standard node pool for system components such as DNS and for anything that must not be interrupted.

Preemption on GKE has its own timing. By default the node gets a 30-second graceful termination period: regular Pods get 15 seconds to stop, then critical system Pods get the remaining 15. Setting terminationGracePeriodSeconds higher in a Pod spec does not extend this. Newer GKE versions allow the period to be raised up to 120 seconds through node system configuration on Standard pools (1.35.0 and later) or ComputeClass kubelet settings (1.36.0 and later). For stateful workloads Google advises testing that they terminate gracefully within 25 seconds of shutdown, to limit the risk of persistent volume corruption; with the default split, only 15 of those seconds belong to your Pods, so either extend the period or keep the shutdown work tiny. Use Jobs or an operator that resumes from checkpoints, and PodDisruptionBudgets for replicated services; a budget does not stop a preemption, but it shapes voluntary disruptions around it. Cluster-level patterns are in GKE in depth and scaling behaviour in the MIG autoscaler.

ML training on Spot accelerators

Accelerator VMs already cannot live migrate, so a GPU training job needs checkpointing whether it runs on Spot or not; Spot mostly changes how often the restart path is exercised. For multi-node synchronous training, one preempted node stops every step, so the effective preemption rate is roughly the per-node rate times the node count; elastic training is covered on the GPU spot page. Spot TPUs follow the same logic; TPU v5 and v6 covers the hardware side.

Keep checkpoints and datasets in Cloud Storage in the same region as the compute, so a replacement VM in another zone can resume without cross-region transfer. Write checkpoints to a temporary name and rename on completion, so a G3 cut mid-upload cannot leave a half-written file that looks valid.

Failure modes

FailureCauseFix
Job restarts from scratchCheckpoints kept on local SSD or boot diskWrite to Cloud Storage; resume from the latest complete one
Corrupt checkpoint after preemptionUpload cut by G3 mid-writeWrite to a temporary name, then rename
Fleet never recoversSingle-zone group in a busy zoneRegional MIG or several zones; on-demand fallback pool
Graceful shutdown never finishesFull checkpoint attempted in the 30 s windowSave regularly; only a marker at preemption
Stopped VMs accumulateDefault STOP on stateless workersUse DELETE for workers in groups
Critical service interruptedUntainted Spot node poolTaints, tolerations and a standard pool

When Spot is the wrong choice

Spot is the wrong choice for anything with a latency SLO and no spare capacity elsewhere, for long synchronous jobs on many nodes without elastic recovery, for databases without replicas on standard VMs, and for work with a hard deadline that cannot absorb hours of unavailability. It is the right choice for queues of independent tasks, CI runners, rendering, hyperparameter sweeps, and training jobs whose restart path is tested. Many production fleets mix the two: a standard baseline sized for the minimum acceptable throughput and a Spot layer on top that absorbs the rest. Committed use discounts apply to steady standard capacity, so the baseline and the Spot layer are priced by different mechanisms.

What to do next

  1. Pick one batch workload and confirm it can be killed at any second and resume from external state. Kill it on purpose to prove it.
  2. Move checkpoints and outputs to Cloud Storage with write-then-rename.
  3. Add the metadata watcher and a short shutdown script; never attempt a full save in the shutdown window.
  4. Run it on a regional MIG or a tainted GKE Spot pool, with DELETE for stateless workers.
  5. Build a dashboard of preemptions per zone and machine type from the audit log, and compute your own mean time between preemptions.
  6. Set the checkpoint interval with Young's formula from measured C, M and R, and revisit it monthly.
  7. Keep a standard pool or on-demand fallback for the minimum throughput you cannot lose.
Key takeaway: Spot VMs trade a large discount for reclamation at any time; preemptible VMs add a 24-hour limit and are the legacy option. On preemption the metadata flag flips, ACPI G2 starts a best-effort shutdown of up to 30 seconds, and power is then cut, so persist state continuously, checkpoint at an interval set from measured costs, watch the metadata flag, run fleets in regional MIGs or tainted GKE Spot pools, and keep standard capacity for whatever cannot be interrupted.