An instance template is the Compute Engine resource that says what a VM should look like: machine type, boot image, disks, network, service account, metadata and scheduling options. On its own it creates nothing. Its value comes from what consumes it, mostly managed instance groups, which use it as the definition of every VM they create, recreate or heal.
The property that shapes everything else is that templates are immutable. You cannot edit one; you create a new one and point consumers at it. That turns a template into a versioned release artifact, and handling it as one is the difference between clean rollouts and a fleet of VMs nobody can explain. This article covers what goes in a template, regional versus global templates, making templates deterministic, creating them with gcloud and Terraform, rolling a group from one to the next, and the failures that catch teams out.
What a template is
Think of a template as a frozen request body for creating a VM, stored server-side with a name. When a managed instance group needs a VM, during scale-out, autohealing or an update, it creates one from its target template. If you create a single VM from a template you can override properties at creation time, but the template itself stays as it was.
Immutability is a feature. A group's state can be described exactly as which template each VM was created from, so a question like what is running in production has a precise answer. Rolling back means pointing at the old template, which still exists. The cost is that every change, even a one-line metadata edit, produces a new template, so you need naming, versioning and cleanup discipline from day one. For how the groups themselves scale, see managed instance groups and the autoscaler.
Anatomy of a template
Machine type, boot disk and network configuration are required; everything else is optional. The fields worth deciding deliberately:
| Field | What to decide | Common mistake |
|---|---|---|
| Machine type | Family and size from load tests | Copying a dev size into production |
| Boot disk image | A specific custom image, not a moving reference | An unpinned image, so nobody can say which build a VM booted |
| Boot disk type and size | Balanced or SSD, size for logs and caches | Too small; disk fills and the VM fails health checks |
| Network | Subnet, no external IP unless required | Public IPs on every VM by default |
| Service account and scopes | A dedicated least-privilege account | The default compute service account with broad access |
| Metadata and startup script | Only configuration that differs per environment | Installing unpinned packages at boot |
| Labels and network tags | Cost attribution and firewall targeting | Tags that do not match any firewall rule |
| Scheduling | Standard or Spot, maintenance behaviour | Spot capacity for a tier that cannot tolerate preemption |
| Shielded VM options | Secure Boot, vTPM, integrity monitoring | Leaving Secure Boot off for no reason |
Scheduling choices have their own depth, covered in GCE instance scheduling and Spot VMs.
Regional or global
Templates come in two scopes. A regional template can only be used in its own region, and Google's guidance is that hardware errors are then isolated to that region. A global template can be used in any region, which is convenient but means a problem with the template's home region can affect every region that uses it. The recommendation is regional templates unless you genuinely reuse one template across regions.
In practice regional templates also fit how configuration differs. Subnets are regional, so a template that names a subnet is regional in meaning even if global in scope. Templates can also reference zonal resources such as an existing read-only persistent disk, and doing so pins the template to that zone: a template that attaches a disk from one zone cannot create VMs in any other zone, which quietly breaks a regional group that spreads across zones. Prefer images and disks created from images over references to specific zonal disks.
Deterministic templates
A template is only as reproducible as what it points at. Two VMs created from the same template a week apart should be identical, and three habits break that. The first is a startup script that runs a package install without versions, so the VM gets whatever the repository holds that day. The second is a container reference by a moving tag such as latest. The third is an image reference that is not a specific image; an image family is a pointer that moves when someone publishes a new image into it.
Google's guidance on deterministic templates is to pin versions in startup scripts, pin container tags, and preferably bake software into a custom image so every instance is the same. The pattern that works is: build an image in CI with everything installed, give it a dated name, reference that exact image in the template, and keep the startup script for environment-specific configuration such as which config file to fetch. Boot is then faster too, which matters for autoscaling and autohealing, because the group waits for each new VM to become healthy.
The template lifecycle
The lifecycle has three loops. The image loop produces a new image when the OS or software changes. The template loop produces a new template for every image or configuration change, named so the version is obvious, for example web-v8-20261004. The rollout loop moves a group from its current template to the new one, in stages, under health checks. Keeping the loops separate means a metadata-only change does not require an image build, and an image build does not automatically change production.
Creating templates with gcloud and Terraform
With gcloud, a regional template that follows the advice above looks like this:
gcloud compute instance-templates create web-v8-20261004 \
--instance-template-region=us-central1 \
--machine-type=e2-standard-4 \
--image=web-20261004-1 --image-project=my-images \
--boot-disk-type=pd-balanced --boot-disk-size=30GB \
--subnet=projects/my-net/regions/us-central1/subnetworks/app \
--no-address \
--service-account=web-runtime@my-proj.iam.gserviceaccount.com \
--scopes=cloud-platform \
--metadata-from-file=startup-script=startup.sh \
--metadata=env=prod,config-uri=gs://my-config/web/prod.yaml \
--labels=app=web,team=frontend,release=v8 \
--tags=web-backend \
--shielded-secure-bootYou can also capture an existing VM's configuration with --source-instance, which is handy for turning a hand-tuned VM into a template, but review the result: it records whatever drifted onto that VM. In Terraform, immutability shows up as replacement. Because changing any field forces a new template, give it a name prefix so each version gets a unique name, and create the replacement before destroying the old one so the group is never pointed at a template that does not exist:
resource "google_compute_region_instance_template" "web" {
name_prefix = "web-"
region = "us-central1"
machine_type = "e2-standard-4"
disk {
source_image = "projects/my-images/global/images/web-20261004-1"
disk_type = "pd-balanced"
disk_size_gb = 30
boot = true
}
network_interface { subnetwork = var.app_subnet } # no access_config: no public IP
service_account {
email = google_service_account.web.email
scopes = ["cloud-platform"]
}
metadata = { env = "prod", startup-script = file("startup.sh") }
labels = { app = "web", team = "frontend" }
lifecycle { create_before_destroy = true }
}
resource "google_compute_region_instance_group_manager" "web" {
name = "web"
region = "us-central1"
base_instance_name = "web"
version { instance_template = google_compute_region_instance_template.web.self_link }
update_policy {
type = "PROACTIVE"
minimal_action = "REPLACE"
max_surge_fixed = 3
max_unavailable_fixed = 0
}
}Here a Terraform apply that changes the template rolls the group immediately. Many teams prefer Terraform to create templates and a deploy pipeline to run staged rollouts, which keeps canary control out of a plan-and-apply cycle.
Rolling a group to a new template
A rollout is the group moving its target from one template to another. The group supports two versions at once, which is what makes canaries possible:
# 10% of the group on v8, the rest stays on v7
T=projects/my-proj/regions/us-central1/instanceTemplates # full path for regional templates
gcloud compute instance-groups managed rolling-action start-update web \
--region=us-central1 \
--version=template=$T/web-v7-20260920 \
--canary-version=template=$T/web-v8-20261004,target-size=10% \
--type=proactive --max-surge=3 --max-unavailable=0
gcloud compute instance-groups managed wait-until web \
--region=us-central1 --version-target-reached
# after the canary looks healthy: everything to v8
gcloud compute instance-groups managed rolling-action start-update web \
--region=us-central1 \
--version=template=$T/web-v8-20261004 \
--type=proactive --max-surge=3 --max-unavailable=0
# rollback is the same command pointing at v7The knobs: --type=proactive replaces existing VMs now, while opportunistic only applies the new template when VMs are created anyway, for example during scale-out or healing. --max-surge is how many extra VMs the group may create above its target size, and --max-unavailable is how many may be offline at once. Defaults are 1 and 1 for zonal groups and the number of zones for regional groups. Setting unavailable to zero with a positive surge gives updates with no capacity dip, at the cost of temporary extra VMs and quota. --minimal-action and --most-disruptive-allowed-action bound whether the group restarts or replaces VMs, and --replacement-method=recreate keeps instance names, which stateful groups need, where substitute creates new names.
The request is intent: the API accepts it immediately and the group works towards it. That is why the wait-until --version-target-reached step exists; a pipeline that moves on after the start command has not deployed anything yet. Use --stable to wait until the group has no actions in progress.
Worked example: a security patch
A web tier runs 30 VMs in a regional group across three zones on template v7. A security patch arrives. CI builds image web-20261004-1 with the patched package, and a pipeline creates regional template v8 that differs from v7 only in the image. The pipeline starts a canary at 10 percent with surge 3 and unavailable 0. The group creates three v8 VMs, waits for them to pass the health check, then deletes three v7 VMs; the version target is reached with 3 of 30 on v8.
The pipeline then watches error rate and latency for the canary VMs, which is easy because the release label differs. After thirty minutes it promotes v8 to 100 percent, three VMs at a time. Two days later, with v8 stable, a cleanup job deletes templates more than three versions old that no group references. If the canary had regressed, the rollback command would have put the three VMs back on v7 in minutes, because v7 still existed.
Failure modes
| Failure | What you see | Fix |
|---|---|---|
| Image family in the template | The template no longer pins one image; which build a VM booted is unclear | Reference a specific image; bump it with a new template |
| Zonal disk referenced | Regional group cannot create VMs in other zones | Use images, not specific zonal disks |
| Pipeline does not wait | Deploy marked done while VMs are still old | wait-until --version-target-reached before the next step |
| Unpinned startup install | New VMs differ from old ones on the same template | Bake into the image; pin versions in scripts |
| Surge exceeds quota | Rollout stalls with creation errors | Lower surge, or raise quota before large rollouts |
| Deleting a template in use | Delete is refused | Move every group off it first; clean up only unreferenced templates |
| Default service account | VMs can reach far more than they need | Dedicated account with narrow roles |
Trade-offs
Baked images make templates deterministic and boot fast but add an image pipeline and slow down small changes; startup-script installs are quick to change but slow to boot and drift-prone. A middle path bakes the OS and runtime, and lets the startup script fetch a pinned application artifact. Regional templates isolate failure and match regional networking but multiply templates by the number of regions; global templates suit fleets that are genuinely identical everywhere. Letting Terraform roll groups is simple, while a separate rollout pipeline gives you canaries and gates. Machine families and disks are covered in Compute Engine in depth.
What to do next
- List your templates and the groups using each; delete nothing yet, but note which are unreferenced.
- Adopt a naming scheme that encodes app, version and date, and add a release label to every template.
- Move from image families or public images to dated custom images built in CI.
- Switch to regional templates for single-region groups and remove references to zonal disks.
- Replace the default service account with a dedicated least-privilege one, and drop external IPs.
- Script the rollout: canary with --canary-version, wait-until --version-target-reached, check metrics, then promote or roll back.
- Set max-unavailable to 0 with a positive surge for serving tiers, and confirm quota covers the surge.
- Add a cleanup job that deletes old unreferenced templates while keeping the last few for rollback.