In Compute Engine, scheduling does not mean placing a job on a cluster. It means the settings that decide what happens to a VM over time: what Google does when the host needs maintenance or fails, whether the VM restarts after a crash, whether it can be preempted, when it is automatically stopped or deleted, and whether it starts and stops on a calendar. These settings live in three places: the scheduling block on each instance, VM time limits, and instance schedules attached as resource policies.

Getting them right is mostly about cost and availability. A development fleet that runs only in working hours costs much less than one that runs all week. A training VM with GPUs cannot live migrate, so it will be stopped for maintenance and must be ready for that. A batch VM that forgets to delete itself keeps billing. This article explains each setting from first principles, shows the gcloud commands, works through a cost example, and lists the failure modes. Machine families, disks and the general VM lifecycle are covered in the Compute Engine deep dive.

Advertisement

Three layers that decide when a VM runs

It helps to keep the three layers apart, because they answer different questions and are configured in different places.

  • The scheduling block is part of the instance resource. It answers: what happens when the infrastructure underneath the VM changes or fails, and is the VM Spot or standard?
  • Time limits are also set on the instance. They answer: when must this VM stop or be deleted, regardless of what runs inside it?
  • Instance schedules are separate regional resource policies attached to VMs. They answer: on which days and at what times should these VMs start and stop?
Three layers decide when a Compute Engine VM runs, stops or disappearsPer-VM scheduling blockonHostMaintenanceautomaticRestartprovisioningModel (STANDARD / SPOT)instanceTerminationActionhost error timeoutTime limitsmax run durationor termination time30 s to 120 daysthen STOP or DELETErecomputed on each startInstance scheduleregional resource policycron start / stop + time zoneone schedule per VMup to 15 min laterun by the service agentVMRUNNING / TERMINATEDhost eventdeadlinecronGuest: shutdown script + metadatacheckpoint, drain, exit cleanlysignalsEvery path that stops a VM is a reason to make the workload restartable.
Host events act through the scheduling block, deadlines through time limits, and calendars through instance schedules. All three end with the guest being asked to shut down, so the workload inside must be able to stop and resume cleanly.

You can read the current state of a VM's scheduling options with one command; the fields shown are the ones discussed below.

gcloud compute instances describe trainer-1 --zone=us-central1-a \
    --format="yaml(scheduling, resourcePolicies, status)"

Host maintenance: MIGRATE or TERMINATE

Google periodically maintains the physical hosts under VMs. The onHostMaintenance option, set with --maintenance-policy, decides what happens to your VM. With MIGRATE, the default for standard VMs, Compute Engine live migrates the running VM to another host. Memory is copied while the VM keeps running, followed by a brief pause, and the VM keeps its IP addresses, disks and in-memory state. With TERMINATE, the VM is stopped for the maintenance and, if automatic restart is on, started again afterwards.

Some VMs cannot live migrate and must use TERMINATE. The documentation lists VMs with GPUs or TPUs attached, Spot and preemptible VMs, bare metal instances, and some storage-heavy machine types. For ML training this is the important point: a GPU VM will be stopped for host maintenance, so a training job must checkpoint and resume rather than assume the VM runs for weeks.

Scheduled maintenance is announced ahead of time. You can see it from outside the VM with --format="yaml(resourceStatus.upcomingMaintenance)" on describe, which shows the window and whether you can trigger it yourself early. Inside the VM, the metadata server exposes the maintenance-event key, and a process can wait on it instead of polling:

import requests

URL = "http://metadata.google.internal/computeMetadata/v1/instance/maintenance-event"
HDR = {"Metadata-Flavor": "Google"}

def watch(on_event):
    last = None
    while True:
        r = requests.get(URL, headers=HDR,
                         params={"wait_for_change": "true", "last_etag": last or "0"},
                         timeout=3600)
        last = r.headers.get("ETag")
        if r.text != "NONE":
            on_event(r.text)      # e.g. save a checkpoint, drain the worker

watch(lambda ev: print("maintenance event:", ev))

Test the handler before you rely on it: gcloud compute instances simulate-maintenance-event triggers a maintenance event on a VM so you can watch your checkpoint and drain logic run.

Advertisement

Automatic restart and host errors

automaticRestart, set with --restart-on-failure or --no-restart-on-failure, controls whether Compute Engine starts the VM again after it is stopped by a host error or a crash rather than by you. It is on by default for standard VMs and off for Spot and preemptible VMs, which cannot restart automatically. Leave it on for long-lived servers. Turn it off for batch workers whose orchestrator should decide what to do after a failure, so a VM does not come back and silently start a duplicate job.

When a host stops responding, Compute Engine waits before declaring the VM failed and restarting it elsewhere. On machine types that support it, --host-error-timeout-seconds sets this wait between 90 and 330 seconds; the default is 330. A shorter timeout gets you back sooner after a real failure but gives a briefly unresponsive host less time to recover. For VMs with Local SSD, --local-ssd-recovery-timeout sets how many hours, from 0 to 168, Compute Engine tries to recover Local SSD data after a host error before giving up. Defaults vary by machine type. Treat Local SSD as a cache that can be lost in any case.

Spot VMs and the termination action

--provisioning-model=SPOT makes a VM a Spot VM: much cheaper, but Compute Engine can preempt it whenever it needs the capacity. There is no minimum or maximum run time. On preemption, the preempted metadata key changes to TRUE and the guest gets a best-effort shutdown period of up to 30 seconds before it is powered off. Google's documentation also describes an optional, longer preemption notice period of 120 seconds instead of the default of none; check the current docs for how to set it, since it is newer than the rest of the model.

The termination action, --instance-termination-action, decides what remains after preemption. STOP, the default for Spot, leaves a TERMINATED VM whose boot disk you keep paying for and which you can start again later if capacity exists. DELETE removes the VM, and its disks if they are set to auto-delete. Pick DELETE for stateless workers created by a managed instance group or a batch system, which will create replacements. Pick STOP for a single long job whose disk holds state you want to resume from. The legacy --preemptible model still exists but lacks newer features such as time limits, so use Spot for new work. Fault-tolerant big-data clusters on Spot workers are discussed in the Dataproc article.

Time limits: max run duration and termination time

A time limit makes Compute Engine stop or delete a VM after a duration or at a timestamp. It is the simplest guard against forgotten VMs. --max-run-duration takes a duration such as 8h or 1d2h; --termination-time takes an absolute timestamp. The limit must be between 30 seconds and 120 days. Both work with standard and Spot VMs but not with legacy preemptible VMs. For standard VMs you must also choose --instance-termination-action; Spot VMs default to stop.

# A GPU fine-tuning VM that cannot run past 10 hours, then deletes itself.
gcloud compute instances create ft-run-0412 \
    --zone=us-central1-a \
    --machine-type=g2-standard-8 \
    --maintenance-policy=TERMINATE \
    --provisioning-model=SPOT \
    --max-run-duration=10h \
    --instance-termination-action=DELETE \
    --metadata-from-file=startup-script=start_job.sh,shutdown-script=save_ckpt.sh

Three details catch people. The termination timestamp is cleared when the VM is stopped or suspended and recalculated when it starts again: a duration counts from the latest start, so a stop and start gives the VM a fresh budget, while an absolute time stays the same. A reset or guest reboot does not change it. The action can happen up to 30 seconds after the deadline. And a VM with Local SSD that should stop rather than be deleted needs --discard-local-ssds-at-termination-timestamp, because Local SSD data cannot be preserved through a stop caused by a time limit.

Instance schedules: start and stop on a calendar

An instance schedule is a regional resource policy with a cron expression for starting VMs, one for stopping them, or both, plus a time zone, which defaults to UTC, and optional start and end dates. You attach it to VMs in the same region.

gcloud compute resource-policies create instance-schedule office-hours \
    --region=europe-west1 \
    --vm-start-schedule="0 8 * * MON-FRI" \
    --vm-stop-schedule="0 19 * * MON-FRI" \
    --timezone="Europe/Amsterdam"

gcloud compute instances add-resource-policies dev-ws-17 \
    --zone=europe-west1-b --resource-policies=office-hours

The schedule is executed by the Compute Engine Service Agent, service-PROJECT_NUMBER@compute-system.iam.gserviceaccount.com, which needs compute.instances.start and compute.instances.stop on the VMs. A schedule that seems to do nothing is usually missing this grant. The person attaching it needs compute.resourcePolicies.use and compute.instances.addResourcePolicies.

The documented limits shape how you use it:

  • Each VM can follow only one instance schedule.
  • An operation may begin up to 15 minutes after the scheduled time, so do not use it for anything that needs to start on the minute.
  • A schedule allows at most one start and one stop operation per hour, and may fail if start and stop are less than 15 minutes apart.
  • VMs with Local SSD cannot be stopped by a schedule.
  • A scheduled start does not reserve capacity. If the zone is out of the machine type at 08:00, the start fails; use reservations for guaranteed starts.

Instance schedules act on individual VMs. Managed instance groups keep their own target size and recreate members, so for a group, change the size instead, for example with the autoscaler's schedule-based scaling, rather than stopping members underneath it. For containers, GKE node pools autoscale on pending pods, which usually beats any calendar.

Worked example: an engineering sandbox fleet

A team runs 40 development VMs at all hours. A week has 168 hours. Working hours of 08:00 to 19:00 on weekdays are 55 hours. With the office-hours schedule, each VM's vCPU and memory charges drop to 55/168, about 33% of before, a saving of roughly two thirds. Disks are still billed while VMs are stopped, and static external IP addresses are billed while unused, so the real saving is somewhat lower; measure it from billing exports rather than trusting the arithmetic.

The rollout has three parts. Grant the service agent its permissions. Create the schedule and attach it to a pilot group of five VMs for a week, checking that people's work survives a stop: unsaved editors and in-memory notebooks do not, so tell users and enable autosave. Then attach it to the rest, with a label such as schedule=office-hours so that gcloud compute instances list --filter="labels.schedule=office-hours" shows who is covered. People working late can start their VM by hand; the next scheduled stop will catch it.

The team's nightly GPU evaluation job uses a different tool. It is a Spot VM created by a Workflows pipeline with a 6-hour --max-run-duration and the DELETE action, so a hung job cannot bill all weekend. Its shutdown script uploads partial results, and the workflow retries on a fresh VM if the result file is incomplete.

Graceful shutdown inside the guest

Every path above ends with a shutdown: maintenance under TERMINATE, preemption, a time limit or a scheduled stop. The guest gets a shutdown signal, and a script in the shutdown-script metadata key runs. Its time is limited, about 30 seconds on Spot and best effort in all cases, so it should do short things: flush a checkpoint already mostly written, deregister from a load balancer or queue, and exit. A job that needs minutes to checkpoint must checkpoint periodically and use the shutdown path only to record where it stopped.

Write work to be idempotent and resumable: a task keyed by an ID, outputs written under a temporary name and renamed when complete, and a startup script that checks for existing checkpoints before beginning. The same habits make preemption, maintenance and schedules boring rather than incidents.

Failure modes and trade-offs

FailureCauseFix
Schedule does nothingService agent lacks start/stop permissionsGrant them on the project or VMs
Morning start failsNo capacity in the zone; schedules reserve nothingReservation, another zone, or a fallback machine type
VM ran longer than expected in totalEach start restarted the max-run-duration countdownUse --termination-time for a hard overall deadline
Training lost hours of workGPU VM stopped for maintenance, no checkpointPeriodic checkpoints and a maintenance-event watcher
Duplicate batch work after a crashAutomatic restart re-ran the startup job--no-restart-on-failure and idempotent tasks
Spot VMs left TERMINATED and billing disksSTOP action on throwaway workersUse DELETE for stateless workers
Scheduled stop rejectedInstance schedules cannot stop VMs with Local SSDStop it another way, or use a time limit with the discard flag

The main trade-offs are cost against availability and convenience. Spot with DELETE is cheapest but needs restartable work. Live migration keeps servers up but is not available for accelerators. Schedules save most on idle fleets but add up to 15 minutes of lateness and no capacity guarantee. Watch outcomes with Cloud Monitoring, alerting on VMs that fail to start and on unexpected stops.

What to do next

  1. List VMs with their scheduling settings and find GPU, Spot and Local SSD machines whose maintenance and restart options do not match their workload.
  2. Add a maintenance-event watcher and periodic checkpoints to every job on a VM that cannot live migrate, and test it with simulate-maintenance-event.
  3. Set --max-run-duration or --termination-time with the right termination action on every batch and experiment VM, remembering that a duration restarts on each start.
  4. Grant the Compute Engine Service Agent start and stop permissions, create an office-hours instance schedule, and pilot it on a few development VMs.
  5. Use DELETE for stateless Spot workers and STOP only where the disk holds state you will resume from.
  6. Measure the savings from billing data after a month and alert on failed scheduled starts.
Key takeaway: Compute Engine decides when a VM runs through three layers. The scheduling block sets how the VM reacts to host maintenance and failures and whether it is Spot. Time limits stop or delete a VM after a duration or at a deadline, between 30 seconds and 120 days. Instance schedules start and stop VMs on a calendar, up to 15 minutes late and without reserving capacity. Each one ends in a guest shutdown, so the work inside must checkpoint, resume and be safe to repeat.