A managed instance group (MIG) keeps a set of identical VMs alive from an instance template; the basics, and where MIGs fit among Compute Engine's features, are in Google Compute Engine in depth. This article is about the part that decides how many VMs that set should contain: the autoscaler. It is a feedback controller with a handful of knobs, and almost every autoscaling incident, whether a fleet that would not grow during a spike or one that oscillated all afternoon, traces back to one of them being set without understanding what it controls.

We will build the controller up from first principles: what each signal measures, how a signal becomes a recommended size, how several signals combine, and how the autoscaler avoids reacting to noise. Then we will configure a real group with gcloud, size it with a worked example, and list the failure modes. Flag names were checked against the current gcloud reference.

Advertisement

The control loop in one paragraph

The autoscaler repeatedly reads its signals, computes for each one the number of VMs that would bring that signal to its target, takes the largest of those numbers, applies any scaling schedule as a floor, and clamps the result between the group's minimum and maximum. Scaling out happens as soon as the recommendation rises. Scaling in is deliberately slower: the autoscaler uses the highest recommendation seen during a stabilization period, and optional scale-in controls cap how far the size may fall within a time window. The MIG then creates or deletes VMs to reach the target, spreading them across zones if it is regional.

How a MIG autoscaler turns signals into a target sizeCPU utilizationtarget 0.0 to 1.0LB serving capacityshare of max rateMonitoring metricper VM or per groupScaling schedulesmin VMs for a windowPer-signal sizeignore initializing VMsTake the largestplus predictive CPUfloorScale outnowScale inmax over stabilizationScale-in controlscap the dropClamp to min and max, then the MIG creates or deletes VMs per zoneregional distribution shape, autohealing and updates act on the same group
Signals become per-signal sizes, the largest wins, schedules set a floor, and scale-in is damped twice.

Signals and how each becomes a size

CPU utilization. You set a target between 0.0 and 1.0. The autoscaler averages utilization across the VMs that are not initializing and scales the group so that the average approaches the target. As a rough model, recommended size is current size times current utilization divided by target: ten VMs at 0.9 with a target of 0.6 suggests fifteen.

Load-balancing serving capacity. For groups behind an external or internal Application Load Balancer, the backend service defines each VM's capacity, for example a maximum rate per instance or a utilization. The autoscaler target is a fraction of that capacity. If each VM is rated at 150 requests per second and the target is 0.8, the autoscaler aims for 120 requests per second per VM. This signal reflects real demand at the front door, which makes it a better primary signal than CPU for request-serving tiers; the backend settings themselves are covered in GCP load balancing.

Cloud Monitoring metrics. Any metric can drive scaling, either per VM, where the autoscaler keeps each VM's value near a utilization target, or per group, where you state how much work one VM should handle and the autoscaler divides the total. A queue depth or a backlog of messages per worker is the typical per-group example. The metric type matters: gauge, delta per minute or delta per second. Getting metrics into Monitoring is covered in Cloud Monitoring.

Schedules. A scaling schedule sets a minimum number of VMs for a window defined by a cron start time, a duration and a time zone. Schedules do not reduce capacity; they only raise the floor. A group can have up to 128 of them.

When several signals are configured, the documentation is explicit that the autoscaler computes a recommendation for each and uses the largest. That makes adding a signal safe in one direction: it can only make the group bigger.

Advertisement

Initialization, stabilization and scale-in controls

Three timing settings do most of the work of preventing bad decisions, and they are easy to confuse.

The initialization period (set with --cool-down-period; the concept was renamed from cool down) is how long a new VM takes to boot and become representative. During it, the autoscaler ignores that VM's utilization when deciding to scale out. The default is 60 seconds. Too short, and new VMs that are booting, warming caches or compiling code report misleading numbers. Too long, and the autoscaler is slow to notice that a second wave of load needs even more VMs. Measure it: create a VM from the template and time it from creation to serving steady traffic.

The stabilization period applies to scale-in. The autoscaler bases scale-in on the peak recommendation over this period, ten minutes by default, configurable from 0 to 3600 seconds with --stabilization-period. Load that dips for five minutes and comes back therefore does not shrink the group.

Scale-in controls limit how many VMs, as a number or a percentage, may be removed relative to the peak within a trailing time window. They exist for services that lose state or warm caches when VMs go away. Note what the reference says: the full allowed reduction may happen in one step, so the application must tolerate losing that many VMs at once.

Modes, predictive scaling and standby pools

The autoscaler has a mode, set with --mode: on (the default), off, which keeps the configuration but stops acting, and only-scale-out, which allows growth but never shrinks the group. The older only-up value is deprecated. only-scale-out is the right setting during an incident, a data migration, or the first days of a launch when you want capacity to ratchet upward while you learn the load.

Predictive autoscaling applies to the CPU signal: --cpu-utilization-predictive-method=optimize-availability makes the autoscaler forecast utilization from history and add VMs ahead of predicted peaks, so that they have finished initializing when the load arrives. It helps when the initialization period is long and load follows daily or weekly cycles, and does little for unpredictable spikes.

Standby pools attack the same problem from the other side by keeping suspended or stopped VMs ready to resume. The group's update settings include --standby-policy-mode (manual or scale-out-pool, where the group resumes standby VMs automatically when it scales out and replenishes the pool afterwards), --suspended-size and --stopped-size. Stopped VMs still incur disk charges and suspended VMs also store memory state, so a standby pool is a cost you pay to shorten scale-out.

Autohealing and rolling updates share the group

The autoscaler is not the only thing creating and deleting VMs. Autohealing recreates a VM when its health check reports it unhealthy, and --initial-delay tells the group to ignore failed checks while a new VM starts. If the initial delay is shorter than real startup, the group deletes VMs that were about to become healthy, in a loop. Base the health check on an endpoint that verifies the application can serve, not just that the process is alive, and keep it cheaper than real requests.

Rolling updates replace VMs with a new template version. --max-surge is how many extra VMs may exist during the update and --max-unavailable how many may be out of service. For a regional group serving traffic, a surge of at least the number of zones with zero unavailable keeps capacity constant during the rollout. The replacement method defaults to substitute for stateless groups, which creates new VMs with new names, and recreate for stateful ones.

Configuring a group with gcloud

REGION=europe-west1
MIG=api-mig

# Regional group spread evenly across zones, autohealing with a realistic initial delay
gcloud compute instance-groups managed create $MIG --region $REGION \
    --template api-tmpl-v42 --size 9 --target-distribution-shape even
gcloud compute instance-groups managed update $MIG --region $REGION \
    --health-check api-hc --initial-delay 180

# Autoscaling: load-balancer signal primary, CPU as a guard with prediction
gcloud compute instance-groups managed set-autoscaling $MIG --region $REGION \
    --min-num-replicas 9 --max-num-replicas 150 \
    --target-load-balancing-utilization 0.8 \
    --target-cpu-utilization 0.7 \
    --cpu-utilization-predictive-method optimize-availability \
    --cool-down-period 140 \
    --stabilization-period 600 \
    --scale-in-control max-scaled-in-replicas=15,time-window=600 \
    --mode on

# Weekday business-hours floor
gcloud compute instance-groups managed update-autoscaling $MIG --region $REGION \
    --set-schedule weekday-peak --schedule-cron "30 7 * * Mon-Fri" \
    --schedule-duration-sec 39600 --schedule-min-required-replicas 60 \
    --schedule-time-zone Europe/Brussels

# Safe rollout of a new template
gcloud compute instance-groups managed rolling-action start-update $MIG --region $REGION \
    --version template=api-tmpl-v43 --max-surge 3 --max-unavailable 0

# Why is the group not at its target?
gcloud compute instance-groups managed list-errors $MIG --region $REGION

The load-balancing target only has meaning together with the backend service's capacity setting, so the 0.8 above assumes a maximum rate per instance has been set on the backend. Treat these commands as a template and check every flag against gcloud compute instance-groups managed set-autoscaling --help for your SDK version.

Worked example: sizing the API tier

An API tier peaks at 12,000 requests per second on weekdays and falls to 1,500 at night. A load test on one VM of the chosen machine type showed p99 latency within the objective up to 150 requests per second, at about 70 percent CPU. Startup measured 40 seconds of boot and 100 seconds of warm-up, so the initialization period is 140 seconds and the autohealing initial delay 180.

With a maximum rate of 150 per instance and a target of 0.8, each VM is planned for 120 requests per second. Peak needs 12,000 / 120 = 100 VMs; the night trough needs 13. The group is regional across three zones, and the minimum is 9, three per zone, so that no zone is ever empty. At the 13-VM trough, losing a zone leaves about nine VMs at roughly 167 requests per second each, over the tested 150, so the plan accepts degraded latency for the 140 seconds replacements take to initialize; a team that cannot accept that raises the minimum to 15.

Losing a zone at peak leaves about 67 VMs at roughly 180 requests per second each until replacements initialize. The maximum of 150 leaves room for the autoscaler to replace a full zone's worth in the surviving zones, provided quota and capacity exist there; check regional CPU quota against 150 VMs before relying on it. The weekday schedule holds 60 VMs from 07:30 to 18:30 so the morning ramp starts from a warm floor, and predictive CPU covers the regular daily curve. Scale-in removes at most 15 VMs below the recent peak per ten-minute window, so even an instant drop from 100 VMs to the 13-VM trough takes six windows, about an hour, an acceptable cost for keeping caches warm.

import math

def lb_size(rps, max_rate=150, target=0.8):
    return math.ceil(rps / (max_rate * target))

def cpu_size(current_vms, avg_util, target=0.7):
    return math.ceil(current_vms * avg_util / target)

def recommend(rps, current_vms, avg_util, schedule_floor, lo=9, hi=150):
    # Model of the documented behaviour: largest signal wins, schedule is a floor, clamp.
    want = max(lb_size(rps), cpu_size(current_vms, avg_util), schedule_floor)
    return min(max(want, lo), hi)

print(recommend(12000, 90, 0.78, 60))   # -> 101 (CPU slightly ahead of LB's 100)
print(recommend(1500, 20, 0.30, 0))     # -> 13

This is a planning model of the documented behaviour, not Google's implementation; use it to check that your targets and limits produce sensible sizes before load does it for you.

Failure modes

  • Maximum reached silently. The group sits at its maximum while latency climbs. Alert when current size equals the maximum for more than a few minutes.
  • Quota or zonal capacity. The autoscaler asks for VMs the region cannot provide; list-errors shows why. Spare quota is part of the capacity plan.
  • Warm-up spikes. New VMs burn CPU while starting; with a short initialization period that can trigger more scale-out. Lengthen the period or scale on the load-balancer signal.
  • Autohealing loops. An initial delay shorter than startup, or a health check that depends on a shared downstream, recreates VMs during an outage elsewhere and makes it worse.
  • Signal that does not track demand. CPU on an I/O-bound service stays low while queues grow; scale on the queue or the load balancer instead.
  • Oscillation. Stabilization set to zero plus aggressive targets produce sawtooth sizing. Keep the default ten minutes unless you have measured a reason.
  • Updates fighting the autoscaler. A rollout with unavailable VMs during a peak reduces capacity; use surge, and roll out off-peak for large groups.

What to do next

  1. Load-test one VM to find its real per-instance capacity at your latency objective.
  2. Measure boot-to-serving time and set the initialization period and autohealing initial delay from it.
  3. Use the load-balancer signal for request tiers and a per-group Monitoring metric for queue workers.
  4. Set the minimum to survive a zone loss at the trough and the maximum to replace a zone at peak, and check quota.
  5. Add schedules for known peaks and predictive CPU if initialization is long and load is cyclical.
  6. Add scale-in controls if VMs carry warm state; keep the default stabilization period.
  7. Alert on group size at maximum, list-errors entries and autohealing recreation rate.
  8. Switch to only-scale-out mode during incidents and migrations, and back to on afterwards.
  9. For scheduling and maintenance behaviour of the VMs themselves, read GCE instance scheduling.
Key takeaway: The MIG autoscaler computes a size per signal, takes the largest, raises it to any schedule floor and clamps it between minimum and maximum. Scale-out is immediate once initializing VMs are excluded; scale-in is damped by the stabilization period and optional scale-in controls. Measure per-VM capacity and startup time, scale request tiers on load-balancer capacity, plan minimum and maximum around zone loss and quota, and alert when the group is pinned at its maximum.