A custom image is a frozen boot disk: an operating system plus everything you installed and configured, stored as a Compute Engine resource that any number of VMs can boot from. An instance template is the recipe a managed instance group uses to create those VMs, and the image is the single biggest thing that recipe points at. Get the pairing right and every VM in a fleet is identical, boots in seconds and can be rolled back. Get it wrong and a scale-out event quietly brings up machines running different software from the ones already serving.
This page is about the image side of that pairing: how images are built, how image families resolve, what deprecation states actually do, how to share images across projects, and the one reference choice in a template that decides whether your fleet is reproducible. The template itself, regional versus global templates and rollout mechanics are covered in Instance Templates, in depth; this article assumes you have read it or will.
What an image is
An image is a global (or storage-location-scoped) resource containing a disk's contents plus metadata: the family it belongs to, guest OS features such as UEFI support, licences, its archive size and a deprecation status. It is immutable. You never edit an image; you create a new one. That immutability is the point: an image name is a promise that two VMs booted from it a month apart start from the same bytes.
Images differ from snapshots in intent. A snapshot is a backup of one disk, incremental and tied to that disk's history. An image is a distribution artefact, meant to be the boot source for many disks. You can create an image from a disk, from a snapshot, from another image or by importing a file from Cloud Storage. When a VM is created from an image, Compute Engine makes a new persistent disk from it; the first VM in a region that is far from the image's storage location pays a one-time copy, after which the image is cached there.
Public images such as Debian, Ubuntu, Rocky Linux or Windows Server live in Google-owned projects like debian-cloud and are organised into families. A custom image is the same kind of resource in your own project. The common reason to make one is called baking: rather than installing your runtime, agents and application on every boot with a startup script, you install them once at build time. Boot then does only environment-specific configuration, which matters most to autoscaling and autohealing because a managed group waits for each new VM to pass its health check.
Building an image
There are two practical ways to build. The manual path creates a VM from a public image, configures it, stops it and captures its boot disk. Google allows imaging a disk attached to a running VM, but recommends stopping the VM first because an image taken mid-write may contain a half-flushed filesystem. If you must image a running VM, pause the application and run sudo sync first.
# Manual build: capture a stopped VM's boot disk into the "web" family.
gcloud compute instances stop web-build-1 --zone=us-central1-a
gcloud compute images create web-20261004-1 \
--source-disk=web-build-1 \
--source-disk-zone=us-central1-a \
--family=web \
--storage-location=usThe repeatable path is a build tool. HashiCorp Packer's googlecompute builder automates exactly those steps: it creates a temporary VM, runs your provisioners over SSH, stops the VM, creates the image and deletes the scratch resources. Pin the base by family so each build starts from the newest patched OS, but give the output a unique, dated name.
variable "nginx_version" {
type = string
}
locals {
stamp = formatdate("YYYYMMDDhhmm", timestamp())
}
source "googlecompute" "web" {
project_id = "img-build-prod"
zone = "us-central1-a"
source_image_family = "debian-12"
source_image_project_id = ["debian-cloud"]
image_name = "web-${local.stamp}"
image_family = "web"
ssh_username = "packer"
machine_type = "e2-standard-4"
}
build {
sources = ["source.googlecompute.web"]
provisioner "shell" {
inline = [
"sudo apt-get update",
"sudo apt-get install -y nginx=${var.nginx_version}", # pin; take it from apt-cache policy
"sudo systemctl enable nginx",
"sudo rm -f /etc/nginx/sites-enabled/default",
]
}
}Two details decide whether the result is trustworthy. First, pin package versions inside the build, or two builds from the same commit differ. Second, generalise the image: remove anything that must be unique per machine or that should never be copied. On Linux that means credentials, shell history, temporary keys and application caches; Google's guest environment regenerates SSH host keys for a new instance, but your own agents may carry their own identity files. On Windows, run GCESysprep before capture.
If you want Shielded VM with Secure Boot, the image needs the UEFI_COMPATIBLE guest OS feature and a GPT disk with an EFI system partition. Images derived from Google's current public images already carry the feature; check with gcloud compute images describe and look at guestOsFeatures rather than assuming.
Image families and how they resolve
A family is a name that resolves to the newest image in it that is not deprecated. Every time you create an image with --family=web, the family moves to it. You can ask what it resolves to:
gcloud compute images describe-from-family web --project=img-build-prod
# What a specific zone sees during a staged rollout of a new image:
gcloud compute images describe-from-family web --project=img-build-prod --zone=europe-west1-bThe zonal form exists because resolution is not instantaneous everywhere: during a rollout, the newest image in a family can differ between zones. The REST API exposes the same thing as zones/ZONE/imageFamilyViews/FAMILY. For automation that runs in many zones, ask per zone rather than assuming one global answer.
Families are excellent for two jobs: as the base input to your own build (source_image_family), and as a human-friendly handle for ad hoc VMs. They are a poor reference for production templates, for the reason the next section explains.
How templates consume images
An instance template can name its boot image in two ways. --image=web-20261004-1 names one immutable image. --image-family=web names a pointer, and the pointer is resolved each time a VM is created from the template, not when the template is created. A template is immutable, but what it points at is not.
Picture a group of ten VMs created on Monday from a family-based template. On Wednesday CI publishes a new image into the family. On Thursday traffic spikes and the autoscaler adds six VMs. Those six boot from Wednesday's image, which nobody rolled out on purpose, and the group now runs two builds. Autohealing does the same: a VM recreated after a failed health check comes back on whatever the family currently resolves to. Nothing reports this as a version change, because as far as the managed group knows, every VM uses the same template.
So pin. Create one template per image, name it after the image, and roll the group from one template to the next. Rollback is then a rolling update back to the previous template, which still names an image that exists.
gcloud compute instance-templates create web-20261004-1 \
--machine-type=e2-standard-4 \
--image=web-20261004-1 --image-project=img-build-prod \
--shielded-secure-boot
gcloud compute instance-groups managed rolling-action start-update web-mig \
--region=us-central1 \
--version=template=web-20261004-1 \
--max-surge=3 --max-unavailable=0For how the group surges, waits for health and handles canary versions, see MIG autoscaling; for the machine and disk choices that go into the rest of the template, see Compute Engine, in depth.
Worked example: a weekly golden image
A concrete weekly cycle for a web tier of twenty VMs in us-central1:
- Monday 02:00: a scheduled CI job runs Packer in a dedicated build project. It starts from the
debian-12family, bakes nginx and the application at the release tag, and producesweb-202610050200in familyweb. - Validation: the same job creates a template from that exact image in a test project, boots three VMs behind a test load balancer, runs smoke tests and a CIS-style scan, and records the image name and its self-link in the release record. A failed check deprecates the image immediately, which moves the family back to last week's image.
- Template: on success, CI creates template
web-202610050200pinned to the image. - Canary: a rolling update with a second version at 10 percent of the group runs for an hour while error rates are compared.
- Rollout: the canary is promoted with surge 3 and zero unavailable, so capacity never drops. Twenty VMs replaced three at a time take about seven waves; with a 90-second boot and a 60-second health-check delay that is roughly 20 minutes.
- Retention: two weeks later, after the image has been superseded twice, it is marked DEPRECATED; after 90 days, once no template references it, OBSOLETE; and later DELETED and removed.
Measured outcomes worth tracking: build time, image archive size (you pay for image storage), time from VM creation to healthy, and how many images per family exist. A tidy family has a handful of live images, not hundreds.
Deprecation, rollback and retention
Deprecation is how you retire images without surprising anyone, and the four states behave differently enough that it is worth being exact:
| State | New VMs from it | Family resolution | Typical use |
|---|---|---|---|
| ACTIVE | Allowed | Eligible as newest | Normal |
| DEPRECATED | Allowed, with a warning; new references allowed | Skipped even if newest | Superseded, still a rollback target |
| OBSOLETE | Rejected with an error; existing references remain | Skipped | Retired, keep for audit |
| DELETED | Rejected with an error | Skipped | Marked gone; the image is not actually removed |
gcloud compute images deprecate web-20260928-1 \
--state=DEPRECATED \
--replacement=web-20261004-1Three consequences follow. First, rollback through a family works by deprecating the newest image: the family then falls back to the previous one, provided that one is not deprecated. Second, deprecation is one-way; you cannot set an image back to ACTIVE, so a bad deprecation is fixed by building or copying a new image, not by undoing. Third, the DELETED state only marks the image; to stop paying for its storage you must actually run gcloud compute images delete.
The most dangerous step is OBSOLETE. A managed group whose template names an obsolete image cannot create VMs at all, so the next autoscale or autoheal fails. Before marking an image OBSOLETE, list templates that reference it and confirm none is the target of a live group.
Sharing and restricting images
Most organisations build images in one project and run them in many. Grant roles/compute.imageUser to the consuming projects' principals, either on the whole image project or on individual images. The service agent that creates VMs for a managed group is one of those principals; forget it and manual VM creation works while autoscaling fails with a permission error.
gcloud compute images add-iam-policy-binding web-20261004-1 \
--project=img-build-prod \
--member=serviceAccount:CONSUMER_PROJECT_NUMBER@cloudservices.gserviceaccount.com \
--role=roles/compute.imageUserOn the other side, use the organisation policy constraint constraints/compute.trustedImageProjects to restrict which projects VMs may boot from. Allow your build project and the public projects you actually use, and nothing else. That turns the golden image from a convention into a control: a developer cannot launch an unscanned community image in production because the API refuses it.
For data residency, --storage-location sets where the image bytes live. If you omit it, Compute Engine picks the multi-region closest to the source. Choose a region or multi-region deliberately if compliance requires it.
Failure modes
- Mixed fleet after scale-out. The template names a family. Symptom: two application versions in one group with no rollout on record. Fix: pin exact images in templates.
- Autoscaling fails after cleanup. An image was marked OBSOLETE or deleted while a live template still named it. Symptom: the group's instance creation errors climb while existing VMs look fine. Fix: check template references before retiring images, and alert on managed-group creation errors.
- Corrupt or inconsistent image. Captured from a running VM mid-write. Symptom: filesystem repair on boot, missing recent files. Fix: stop the VM, or sync and pause writes.
- Cloned identity. A secret, machine-specific token or agent ID was baked in. Symptom: monitoring shows many hosts as one, or a leaked credential in every VM. Fix: generalise in the build, and fetch secrets at boot from Secret Manager using the VM's service account.
- Stale base. The image was built months ago and never rebuilt. Symptom: every new VM boots with known vulnerabilities. Fix: rebuild on a schedule even without application changes.
- Slow first boot in a new region. The image is stored far away. Fix: set a storage location near the fleet, or expect a one-time cost.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| Fully baked image | Fast, identical boots; no runtime dependency on package mirrors | Rebuild for every change; image sprawl |
| Thin image + startup script | One image serves many apps; quick config changes | Slow, fragile boots; drift if versions are unpinned |
| Container-Optimized OS + container | OS patched by Google; app versioned as a container tag | Less control over the host; still pin by digest |
| Template pinned to image | Reproducible fleets, real rollback | One template per build |
| Template pointing at a family | No template churn | Silent drift on scale-out and autoheal |
A reasonable default is a baked image for the runtime and agents, configuration fetched at boot, and a template per image. Teams that deploy several times a day often move the application into a container on a baked host image so that only the host image follows the weekly cycle.
What to do next
- Run
gcloud compute instance-templates describeon every live template and list which ones usesourceImagewith a family path; replace them with pinned images. - Move image builds into one project with Packer in CI, with dated image names and pinned package versions.
- Add a validation stage that boots the image from a pinned template and runs smoke tests before any production template is created.
- Grant
roles/compute.imageUserto consuming projects, including the service agent, and setconstraints/compute.trustedImageProjects. - Write a retention job: DEPRECATED after supersession, OBSOLETE only when no live template references the image, then delete.
- Schedule a rebuild at least weekly so security patches reach new VMs without waiting for an app release.