Compute Engine is Google Cloud's virtual machine service and the layer much of the platform is built on: GKE nodes, Dataproc clusters and many managed services run on Compute Engine VMs. Creating a VM takes one command. Running a fleet of them well needs a mental model of what Google is doing underneath: where the VM is placed, why its disk is a network service, what happens when the host needs maintenance, where its identity comes from, and how groups of VMs heal and scale.

This article builds that model and turns it into decisions: which machine and disk to choose, how to survive maintenance and preemption, how to give VMs credentials without keys, and how to run stateless tiers as managed instance groups.

Advertisement

The resource model

A VM, called an instance, lives in a project and a single zone. It has a machine type (vCPUs and memory), one boot disk plus optional data disks, one or more network interfaces in VPC subnets, an optional external IP, a service account, and metadata (key-value pairs readable from inside the VM). All of these are separate API resources. A disk can outlive its VM and be attached to another; an IP can be reserved and moved.

When you call instances.insert, the control plane checks your project's quota, picks a physical host in the zone with capacity for the machine type, and attaches network-backed storage and networking. Because disks and networking are services reached over the data centre network rather than parts of the host, Google can move a running VM to another host and your data does not move with it.

gcloud / API / Terraformcompute.googleapis.comCompute control planequota, placement, zonesZone: us-central1-aphysical hostsinsertplaceHostVM (guest OS)vCPUs + memoryMetadata server169.254.169.254Block storagePersistent Disk / HyperdiskVPC networkvNIC, firewall, routesManaged instance grouptemplate + autoscalercreate / healService accounttokens via metadataDisks and networking are network services, not parts of the host, which is what makes live migration and disk reattachment possible.
A Compute Engine VM: the control plane places it on a host in one zone; storage, networking and the metadata server are services the VM reaches, not parts of the host.
gcloud compute instances create api-1 \
  --zone=us-central1-a \
  --machine-type=n2-standard-8 \
  --image-family=debian-12 --image-project=debian-cloud \
  --boot-disk-size=50GB \
  --subnet=app-subnet --no-address \
  --service-account=api-runtime@my-proj.iam.gserviceaccount.com \
  --scopes=cloud-platform \
  --shielded-secure-boot --shielded-vtpm \
  --maintenance-policy=MIGRATE \
  --metadata=enable-oslogin=TRUE \
  --metadata-from-file=startup-script=startup.sh

Choosing a machine

Machine types are grouped into families, and each family into series named by generation and processor. The families are general-purpose (E2, N2, N2D, N4, C3, C4, the Arm-based C4A and others), compute-optimized (such as H3), memory-optimized (the M and X series, for large in-memory databases), storage-optimized (Z3 with large local SSD), network-optimized, and accelerator-optimized (the A and G series with NVIDIA GPUs). The exact list changes every year, so check the machine families page rather than a blog post.

Within a series, predefined types are named series-shape-vcpus: highcpu has about 2 GB per vCPU, standard about 4 GB and highmem about 8 GB, so n2-standard-8 is 8 vCPUs and 32 GB. N and E series also support custom machine types when no predefined shape fits.

WorkloadReasonable starting pointWhy
Dev, low-traffic servicesE2Lowest cost; performance is less predictable than newer series
General web and API tiersN2, N2D, N4 or C4A (Arm, if your stack builds for it)Balanced price and performance
CPU-bound, latency-sensitiveC3, C4 and their AMD equivalentsNewest cores and highest per-core performance
Large in-memory databasesM seriesTerabytes of memory per VM
Training and inferenceA series (and G series for smaller GPUs)GPUs attached; see the maintenance section

Benchmark two or three candidates with your real workload, and compare cost per request rather than cost per vCPU. A newer series that costs more per hour often finishes the same work on fewer VMs.

Advertisement

Disks: Persistent Disk, Hyperdisk and Local SSD

Block storage comes in three shapes. Persistent Disk (standard, balanced, SSD) is network block storage whose performance scales with provisioned size and with the VM's vCPU count. Hyperdisk is the newer generation: you provision IOPS and throughput independently of capacity and can change performance while the disk is in use. The variants are Balanced for most workloads and boot disks, Balanced High Availability which replicates across two zones, Extreme for the highest IOPS, Throughput for scan-heavy analytics, and ML for read-only data shared by up to 2,500 VMs. Some newer machine series attach Hyperdisk only, so check the series' disk support table before designing around Persistent Disk. Local SSD is physically attached to the host: very fast, and its contents can be lost when the VM stops or the host fails.

Two practical rules follow. First, the VM itself has a storage throughput ceiling that depends on the machine type, so a faster disk on a small VM may not go faster. Second, anything on Local SSD must be reproducible from elsewhere: caches, shuffle space and scratch data, not the only copy of a database. Take snapshots on a schedule for persistent data, and use regional or High Availability disks when a zonal outage must not lose writes.

Lifecycle, states and what you pay for

StateMeaningBilled for
PROVISIONING, STAGINGResources allocated, preparing first bootAttached resources such as disks and IPs
RUNNINGBooting or runningvCPU, memory and attached resources
STOPPING, TERMINATEDShutting down; stoppedAttached resources only while TERMINATED
SUSPENDING, SUSPENDEDMemory saved; can stay suspended up to 60 days, then becomes TERMINATEDMemory and attached resources
REPAIRINGGoogle is repairing after an internal error or host failureNot usable and outside the SLA while repairing

Stopping a VM stops vCPU and memory charges, but its disks and any reserved static IP are still billed. Deleting a VM deletes its boot disk only if auto-delete is set on it, which is the default for disks created with the VM. Newer releases add PENDING (for resources that wait for capacity) and PENDING_STOP (for graceful shutdown); scripts that assume a fixed list of states should tolerate new ones.

Host maintenance and live migration

Physical hosts need kernel updates, firmware fixes and hardware repair. For most VMs the default host maintenance policy is MIGRATE: Google copies the VM's memory to another host while it runs and switches over with a brief pause, which most applications never notice. The alternative is TERMINATE, where the VM is stopped (after a soft power-off signal and up to 60 seconds to shut down cleanly) and, with automatic restart enabled, started again, typically within about three minutes.

Some VMs cannot live migrate and are always terminated for maintenance: VMs with GPUs or TPUs attached, bare-metal instances, most Confidential VMs, and some storage-optimized shapes with very large local SSD. For those, and for latency-sensitive services generally, watch the maintenance-event metadata value, which changes from NONE when an event is coming, and drain work before it happens. For training jobs this means checkpointing, and for serving it means leaving the load balancer first.

The metadata server and VM identity

Every VM can reach a metadata server at metadata.google.internal (169.254.169.254). It serves the VM's own facts (name, zone, network), your custom metadata such as startup scripts, and short-lived OAuth access tokens for the VM's attached service account. Requests must carry the Metadata-Flavor: Google header, which stops a naive server-side request forgery from fetching tokens through a proxy that forwards arbitrary URLs.

# From inside the VM. The header is required; requests without it are rejected.
MD=http://metadata.google.internal/computeMetadata/v1
curl -s -H "Metadata-Flavor: Google" $MD/instance/zone
curl -s -H "Metadata-Flavor: Google" \
  $MD/instance/service-accounts/default/token     # short-lived OAuth access token

# Block until a host maintenance event is announced, then drain.
while true; do
  ev=$(curl -sf -H "Metadata-Flavor: Google" \
       "$MD/instance/maintenance-event?wait_for_change=true")
  # An empty value means the call failed: retry, do not drain.
  if [ -n "$ev" ] && [ "$ev" != "NONE" ]; then
    logger "maintenance event: $ev, draining"
    systemctl start drain-traffic.service
  fi
done

Attach a dedicated, least-privilege service account per workload, grant it roles through IAM, and set the access scope to cloud-platform so IAM alone decides access. Never download service-account keys onto VMs: the metadata server already provides rotating credentials, and client libraries use it automatically. For SSH, OS Login ties access to IAM identities instead of keys pasted into metadata.

Spot VMs

Spot VMs use spare capacity at a large discount, and Google can reclaim them at any time. When preempted, a Spot VM gets a best-effort shutdown period of up to 30 seconds and the preempted metadata value becomes TRUE; the termination action decides whether it is then stopped (the default) or deleted. Spot VMs do not live migrate and have no maximum runtime unless you set one, unlike the legacy preemptible VMs, which were limited to 24 hours.

gcloud compute instances create worker-17 \
  --zone=us-central1-b --machine-type=c3-standard-22 \
  --provisioning-model=SPOT \
  --instance-termination-action=DELETE
# Inside the worker: the watcher that checkpoints on preemption.
import requests, subprocess
MD = "http://metadata.google.internal/computeMetadata/v1/instance/preempted"
while True:
    r = requests.get(MD, params={"wait_for_change": "true"},
                     headers={"Metadata-Flavor": "Google"}, timeout=3600)
    if r.text.strip() == "TRUE":
        subprocess.run(["/opt/worker/checkpoint-and-release"], timeout=25)
        break

Use Spot for work that checkpoints or is idempotent: batch processing, CI runners, rendering, and training jobs that resume from checkpoints. Preemptions are often correlated, because capacity is reclaimed in a zone at once, so spread across zones and machine types and keep a small on-demand floor for anything with a deadline.

Managed instance groups

Individual VMs are pets; a managed instance group (MIG) makes them cattle. A MIG creates identical VMs from an instance template, keeps the target number running, recreates VMs that fail a health check (autohealing), scales on CPU, load-balancer or custom metrics, and rolls out new templates gradually. A regional MIG spreads VMs across zones, so a zonal outage costs a fraction of capacity instead of all of it. MIGs are the usual backend for Cloud Load Balancing.

gcloud compute instance-templates create api-v7 \
  --machine-type=n2-standard-4 --image-family=api-image --image-project=my-proj \
  --subnet=projects/my-proj/regions/us-central1/subnetworks/app-subnet \
  --no-address --service-account=api-runtime@my-proj.iam.gserviceaccount.com \
  --scopes=cloud-platform --region=us-central1

gcloud compute instance-groups managed create api \
  --region=us-central1 --template=api-v7 --size=6 \
  --health-check=api-hc --initial-delay=120

gcloud compute instance-groups managed set-autoscaling api \
  --region=us-central1 --min-num-replicas=6 --max-num-replicas=30 \
  --target-cpu-utilization=0.6

# Roll to a new template without losing capacity.
gcloud compute instance-groups managed rolling-action start-update api \
  --region=us-central1 --version=template=api-v8 \
  --max-surge=3 --max-unavailable=0

Set initial-delay longer than your real boot and warm-up time, or autohealing will recreate VMs that are merely slow to start, in a loop. Bake software into images rather than installing it in startup scripts, so boot is fast and a package mirror outage cannot stop scale-out.

Worked example: an API tier and a batch fleet

A team runs an API at a steady 3,000 requests per second, peaking at 9,000, with p99 latency under 200 ms, plus nightly batch jobs of roughly 2,000 vCPU-hours. A load test shows one n2-standard-4 serves 250 requests per second at 60 percent CPU. The API runs as a regional MIG across three zones: 12 VMs at steady state, autoscaling at 60 percent CPU up to 40, which covers the peak even with one zone lost. It uses Hyperdisk Balanced boot disks, no external IPs, Cloud NAT for egress, and a template updated through rolling updates with max-unavailable 0.

The batch jobs run on Spot c3-standard-22 workers in two zones, checkpointing every ten minutes to Cloud Storage, with a deadline-aware scheduler that falls back to on-demand VMs if a job has not finished by 05:00. Committed use discounts cover the 12-VM API floor; everything above it is on demand or Spot.

Failure modes

FailureSymptomMitigation
Zonal stockoutCreate fails with ZONE_RESOURCE_POOL_EXHAUSTED, common for GPUsMultiple zones and series; reservations for critical capacity
Project quotaCreate fails with a quota error for CPUs, GPUs or IPsMonitor quota usage; request increases before launches
Autohealing loopVMs recreated every few minutesLonger initial delay; health check on a lightweight endpoint
Startup script failureVM RUNNING but service absentBake images; read the serial console output
Spot preemption waveMany workers vanish togetherZone and shape diversity, checkpoints, on-demand floor
GPU host maintenanceTraining node stopsWatch maintenance-event, checkpoint, resume elsewhere
Disk throughput ceilingI/O plateaus below disk ratingLarger VM or re-provisioned Hyperdisk performance

For GPU fleets and container platforms, most of this is handled one level up: GKE node pools are MIGs, and the same maintenance, Spot and quota rules apply to their nodes.

What to do next

  1. Inventory your VMs by series; benchmark one current-generation alternative for your largest tier.
  2. Move every stateless tier into a regional MIG with autohealing and rolling updates.
  3. Replace default service accounts and any downloaded keys with per-workload accounts and IAM roles.
  4. Remove external IPs where not needed and enable OS Login.
  5. Add a maintenance-event watcher to VMs that cannot live migrate, especially GPU nodes.
  6. Move checkpointing batch work to Spot across at least two zones.
  7. Review disks: stopped VMs' disks, unattached disks and snapshot schedules, and whether Hyperdisk performance matches measured I/O.
Key takeaway: Compute Engine is easiest to run well when you remember that a VM is a placement of vCPUs and memory on a host, with disks, networking and identity supplied as services. That model explains live migration, why GPU VMs are terminated instead, why Local SSD is scratch space, and why the metadata server is the right source of credentials. Choose machine series by benchmark, provision Hyperdisk performance deliberately, run stateless tiers as regional managed instance groups, put checkpointing work on Spot, and watch maintenance and preemption signals on anything that cannot move.