Compute Engine is Google Cloud's virtual machine service and the layer much of the platform is built on: GKE nodes, Dataproc clusters and many managed services run on Compute Engine VMs. Creating a VM takes one command. Running a fleet of them well needs a mental model of what Google is doing underneath: where the VM is placed, why its disk is a network service, what happens when the host needs maintenance, where its identity comes from, and how groups of VMs heal and scale.
This article builds that model and turns it into decisions: which machine and disk to choose, how to survive maintenance and preemption, how to give VMs credentials without keys, and how to run stateless tiers as managed instance groups.
The resource model
A VM, called an instance, lives in a project and a single zone. It has a machine type (vCPUs and memory), one boot disk plus optional data disks, one or more network interfaces in VPC subnets, an optional external IP, a service account, and metadata (key-value pairs readable from inside the VM). All of these are separate API resources. A disk can outlive its VM and be attached to another; an IP can be reserved and moved.
When you call instances.insert, the control plane checks your project's quota, picks a physical host in the zone with capacity for the machine type, and attaches network-backed storage and networking. Because disks and networking are services reached over the data centre network rather than parts of the host, Google can move a running VM to another host and your data does not move with it.
gcloud compute instances create api-1 \
--zone=us-central1-a \
--machine-type=n2-standard-8 \
--image-family=debian-12 --image-project=debian-cloud \
--boot-disk-size=50GB \
--subnet=app-subnet --no-address \
--service-account=api-runtime@my-proj.iam.gserviceaccount.com \
--scopes=cloud-platform \
--shielded-secure-boot --shielded-vtpm \
--maintenance-policy=MIGRATE \
--metadata=enable-oslogin=TRUE \
--metadata-from-file=startup-script=startup.sh
Choosing a machine
Machine types are grouped into families, and each family into series named by generation and processor. The families are general-purpose (E2, N2, N2D, N4, C3, C4, the Arm-based C4A and others), compute-optimized (such as H3), memory-optimized (the M and X series, for large in-memory databases), storage-optimized (Z3 with large local SSD), network-optimized, and accelerator-optimized (the A and G series with NVIDIA GPUs). The exact list changes every year, so check the machine families page rather than a blog post.
Within a series, predefined types are named series-shape-vcpus: highcpu has about 2 GB per vCPU, standard about 4 GB and highmem about 8 GB, so n2-standard-8 is 8 vCPUs and 32 GB. N and E series also support custom machine types when no predefined shape fits.
| Workload | Reasonable starting point | Why |
|---|---|---|
| Dev, low-traffic services | E2 | Lowest cost; performance is less predictable than newer series |
| General web and API tiers | N2, N2D, N4 or C4A (Arm, if your stack builds for it) | Balanced price and performance |
| CPU-bound, latency-sensitive | C3, C4 and their AMD equivalents | Newest cores and highest per-core performance |
| Large in-memory databases | M series | Terabytes of memory per VM |
| Training and inference | A series (and G series for smaller GPUs) | GPUs attached; see the maintenance section |
Benchmark two or three candidates with your real workload, and compare cost per request rather than cost per vCPU. A newer series that costs more per hour often finishes the same work on fewer VMs.
Disks: Persistent Disk, Hyperdisk and Local SSD
Block storage comes in three shapes. Persistent Disk (standard, balanced, SSD) is network block storage whose performance scales with provisioned size and with the VM's vCPU count. Hyperdisk is the newer generation: you provision IOPS and throughput independently of capacity and can change performance while the disk is in use. The variants are Balanced for most workloads and boot disks, Balanced High Availability which replicates across two zones, Extreme for the highest IOPS, Throughput for scan-heavy analytics, and ML for read-only data shared by up to 2,500 VMs. Some newer machine series attach Hyperdisk only, so check the series' disk support table before designing around Persistent Disk. Local SSD is physically attached to the host: very fast, and its contents can be lost when the VM stops or the host fails.
Two practical rules follow. First, the VM itself has a storage throughput ceiling that depends on the machine type, so a faster disk on a small VM may not go faster. Second, anything on Local SSD must be reproducible from elsewhere: caches, shuffle space and scratch data, not the only copy of a database. Take snapshots on a schedule for persistent data, and use regional or High Availability disks when a zonal outage must not lose writes.
Lifecycle, states and what you pay for
| State | Meaning | Billed for |
|---|---|---|
| PROVISIONING, STAGING | Resources allocated, preparing first boot | Attached resources such as disks and IPs |
| RUNNING | Booting or running | vCPU, memory and attached resources |
| STOPPING, TERMINATED | Shutting down; stopped | Attached resources only while TERMINATED |
| SUSPENDING, SUSPENDED | Memory saved; can stay suspended up to 60 days, then becomes TERMINATED | Memory and attached resources |
| REPAIRING | Google is repairing after an internal error or host failure | Not usable and outside the SLA while repairing |
Stopping a VM stops vCPU and memory charges, but its disks and any reserved static IP are still billed. Deleting a VM deletes its boot disk only if auto-delete is set on it, which is the default for disks created with the VM. Newer releases add PENDING (for resources that wait for capacity) and PENDING_STOP (for graceful shutdown); scripts that assume a fixed list of states should tolerate new ones.
Host maintenance and live migration
Physical hosts need kernel updates, firmware fixes and hardware repair. For most VMs the default host maintenance policy is MIGRATE: Google copies the VM's memory to another host while it runs and switches over with a brief pause, which most applications never notice. The alternative is TERMINATE, where the VM is stopped (after a soft power-off signal and up to 60 seconds to shut down cleanly) and, with automatic restart enabled, started again, typically within about three minutes.
Some VMs cannot live migrate and are always terminated for maintenance: VMs with GPUs or TPUs attached, bare-metal instances, most Confidential VMs, and some storage-optimized shapes with very large local SSD. For those, and for latency-sensitive services generally, watch the maintenance-event metadata value, which changes from NONE when an event is coming, and drain work before it happens. For training jobs this means checkpointing, and for serving it means leaving the load balancer first.
The metadata server and VM identity
Every VM can reach a metadata server at metadata.google.internal (169.254.169.254). It serves the VM's own facts (name, zone, network), your custom metadata such as startup scripts, and short-lived OAuth access tokens for the VM's attached service account. Requests must carry the Metadata-Flavor: Google header, which stops a naive server-side request forgery from fetching tokens through a proxy that forwards arbitrary URLs.
# From inside the VM. The header is required; requests without it are rejected.
MD=http://metadata.google.internal/computeMetadata/v1
curl -s -H "Metadata-Flavor: Google" $MD/instance/zone
curl -s -H "Metadata-Flavor: Google" \
$MD/instance/service-accounts/default/token # short-lived OAuth access token
# Block until a host maintenance event is announced, then drain.
while true; do
ev=$(curl -sf -H "Metadata-Flavor: Google" \
"$MD/instance/maintenance-event?wait_for_change=true")
# An empty value means the call failed: retry, do not drain.
if [ -n "$ev" ] && [ "$ev" != "NONE" ]; then
logger "maintenance event: $ev, draining"
systemctl start drain-traffic.service
fi
doneAttach a dedicated, least-privilege service account per workload, grant it roles through IAM, and set the access scope to cloud-platform so IAM alone decides access. Never download service-account keys onto VMs: the metadata server already provides rotating credentials, and client libraries use it automatically. For SSH, OS Login ties access to IAM identities instead of keys pasted into metadata.
Spot VMs
Spot VMs use spare capacity at a large discount, and Google can reclaim them at any time. When preempted, a Spot VM gets a best-effort shutdown period of up to 30 seconds and the preempted metadata value becomes TRUE; the termination action decides whether it is then stopped (the default) or deleted. Spot VMs do not live migrate and have no maximum runtime unless you set one, unlike the legacy preemptible VMs, which were limited to 24 hours.
gcloud compute instances create worker-17 \
--zone=us-central1-b --machine-type=c3-standard-22 \
--provisioning-model=SPOT \
--instance-termination-action=DELETE# Inside the worker: the watcher that checkpoints on preemption.
import requests, subprocess
MD = "http://metadata.google.internal/computeMetadata/v1/instance/preempted"
while True:
r = requests.get(MD, params={"wait_for_change": "true"},
headers={"Metadata-Flavor": "Google"}, timeout=3600)
if r.text.strip() == "TRUE":
subprocess.run(["/opt/worker/checkpoint-and-release"], timeout=25)
breakUse Spot for work that checkpoints or is idempotent: batch processing, CI runners, rendering, and training jobs that resume from checkpoints. Preemptions are often correlated, because capacity is reclaimed in a zone at once, so spread across zones and machine types and keep a small on-demand floor for anything with a deadline.
Managed instance groups
Individual VMs are pets; a managed instance group (MIG) makes them cattle. A MIG creates identical VMs from an instance template, keeps the target number running, recreates VMs that fail a health check (autohealing), scales on CPU, load-balancer or custom metrics, and rolls out new templates gradually. A regional MIG spreads VMs across zones, so a zonal outage costs a fraction of capacity instead of all of it. MIGs are the usual backend for Cloud Load Balancing.
gcloud compute instance-templates create api-v7 \
--machine-type=n2-standard-4 --image-family=api-image --image-project=my-proj \
--subnet=projects/my-proj/regions/us-central1/subnetworks/app-subnet \
--no-address --service-account=api-runtime@my-proj.iam.gserviceaccount.com \
--scopes=cloud-platform --region=us-central1
gcloud compute instance-groups managed create api \
--region=us-central1 --template=api-v7 --size=6 \
--health-check=api-hc --initial-delay=120
gcloud compute instance-groups managed set-autoscaling api \
--region=us-central1 --min-num-replicas=6 --max-num-replicas=30 \
--target-cpu-utilization=0.6
# Roll to a new template without losing capacity.
gcloud compute instance-groups managed rolling-action start-update api \
--region=us-central1 --version=template=api-v8 \
--max-surge=3 --max-unavailable=0Set initial-delay longer than your real boot and warm-up time, or autohealing will recreate VMs that are merely slow to start, in a loop. Bake software into images rather than installing it in startup scripts, so boot is fast and a package mirror outage cannot stop scale-out.
Worked example: an API tier and a batch fleet
A team runs an API at a steady 3,000 requests per second, peaking at 9,000, with p99 latency under 200 ms, plus nightly batch jobs of roughly 2,000 vCPU-hours. A load test shows one n2-standard-4 serves 250 requests per second at 60 percent CPU. The API runs as a regional MIG across three zones: 12 VMs at steady state, autoscaling at 60 percent CPU up to 40, which covers the peak even with one zone lost. It uses Hyperdisk Balanced boot disks, no external IPs, Cloud NAT for egress, and a template updated through rolling updates with max-unavailable 0.
The batch jobs run on Spot c3-standard-22 workers in two zones, checkpointing every ten minutes to Cloud Storage, with a deadline-aware scheduler that falls back to on-demand VMs if a job has not finished by 05:00. Committed use discounts cover the 12-VM API floor; everything above it is on demand or Spot.
Failure modes
| Failure | Symptom | Mitigation |
|---|---|---|
| Zonal stockout | Create fails with ZONE_RESOURCE_POOL_EXHAUSTED, common for GPUs | Multiple zones and series; reservations for critical capacity |
| Project quota | Create fails with a quota error for CPUs, GPUs or IPs | Monitor quota usage; request increases before launches |
| Autohealing loop | VMs recreated every few minutes | Longer initial delay; health check on a lightweight endpoint |
| Startup script failure | VM RUNNING but service absent | Bake images; read the serial console output |
| Spot preemption wave | Many workers vanish together | Zone and shape diversity, checkpoints, on-demand floor |
| GPU host maintenance | Training node stops | Watch maintenance-event, checkpoint, resume elsewhere |
| Disk throughput ceiling | I/O plateaus below disk rating | Larger VM or re-provisioned Hyperdisk performance |
For GPU fleets and container platforms, most of this is handled one level up: GKE node pools are MIGs, and the same maintenance, Spot and quota rules apply to their nodes.
What to do next
- Inventory your VMs by series; benchmark one current-generation alternative for your largest tier.
- Move every stateless tier into a regional MIG with autohealing and rolling updates.
- Replace default service accounts and any downloaded keys with per-workload accounts and IAM roles.
- Remove external IPs where not needed and enable OS Login.
- Add a maintenance-event watcher to VMs that cannot live migrate, especially GPU nodes.
- Move checkpointing batch work to Spot across at least two zones.
- Review disks: stopped VMs' disks, unattached disks and snapshot schedules, and whether Hyperdisk performance matches measured I/O.