A Virtual Machine Scale Set is Azure's way to run many similar VMs as one thing. You describe a VM once, the model, and the scale set creates, spreads, updates, replaces and removes instances to match it. Many Azure services that need a fleet of VMs, including the node pools under AKS, sit on this abstraction, so the scale set's behaviour under failure decides whether your service survives a bad image, a zone outage or a traffic spike.
This page explains the moving parts from first principles: the two orchestration modes, how health is reported, how model changes reach instances, how autoscale decides and why it sometimes refuses to scale in, how scale-in chooses victims, how repairs work and how to mix Spot and regular capacity. The settings, defaults and limits below were checked against Microsoft Learn on 2026-10-02; where the documentation contradicts itself, this page says so instead of picking a number.
Model and instances
The central idea is the split between the scale set model and the instances. The model holds the image, VM size, network configuration, extensions and policies. Each instance was created from some version of the model. Changing the model does not change any running VM; it only makes existing instances out of date, and the upgrade policy decides when and how they are brought to the latest model. Capacity is a separate number: autoscale or an operator changes it, and the scale set creates or deletes instances to match. Three independent control loops act on the fleet, shown below, and most surprises come from forgetting that they interact.
Flexible or Uniform
The orchestration mode is chosen at creation and cannot be changed later. Microsoft recommends Flexible, and since November 2023 scale sets created with PowerShell or the Azure CLI default to Flexible when no mode is given.
| Flexible | Uniform | |
|---|---|---|
| Instances are | Standard VMs (Microsoft.Compute/virtualMachines) | Scale set VMs with their own API |
| Max instances with fault domain spreading | 1,000 | Documented inconsistently; check for your case |
| Mix VM sizes, OS, Spot and regular | Yes | No; all Spot or all regular |
| Health source | Application Health extension only | Extension or load balancer probe |
| Default outbound internet | None; must configure explicit outbound | Yes |
| Attach existing VMs | Yes | No |
| Azure Backup, Site Recovery | Yes | No |
| Used by AKS and Service Fabric | No | Yes |
The 1,000-instance figure applies to VM sizes that support memory-preserving updates or to max spreading (platformFaultDomainCount=1); sizes without memory-preserving updates in fixed spreading are limited to 200, and in that configuration Flexible also cannot mix Spot with regular instances or general-purpose with specialty sizes. Two rows catch people. Flexible instances have no default outbound access, so a new scale set without a NAT gateway, load balancer outbound rule or public IP cannot reach package repositories, and the first boot script fails. And because Flexible instances are ordinary VMs, a policy or script that lists VMs in a resource group will now see them, which is usually what you want.
Health: the signal everything else trusts
Rolling upgrades, automatic repairs and automatic OS upgrades all decide based on instance health, so a wrong health signal makes every automation wrong. Flexible scale sets must use the Application Health extension, which probes a local endpoint from inside the VM. Prefer Rich Health States (typeHandlerVersion 2.0): the application returns HTTP 2xx with a body {"ApplicationHealthState": "Healthy"} or "Unhealthy", and a non-2xx code, timeout or bad body becomes Unknown, which repairs treat like Unhealthy. New instances start in Initializing until the same state is reported numberOfProbes times in a row or gracePeriod expires. Defaults are a 5-second interval (5-60 allowed), one probe (1-24 allowed) and a grace period of interval times probes, up to 14,400 seconds.
az vmss extension set \
--resource-group rg-web --vmss-name web \
--name ApplicationHealthLinux --publisher Microsoft.ManagedServices --version 2.0 \
--settings '{"protocol":"http","port":8080,"requestPath":"/healthz",
"intervalInSeconds":5,"numberOfProbes":3,"gracePeriod":300}'The endpoint should answer the question "should this instance take traffic and count as good after an upgrade", which means checking local dependencies the instance owns, not shared ones. If the endpoint reports Unhealthy whenever the database is slow, a database incident turns into every instance being repaired at once.
# /healthz on port 8080: local checks only
from http.server import BaseHTTPRequestHandler, HTTPServer
import json, shutil, os
def healthy():
disk_ok = shutil.disk_usage("/").free > 1 << 30 # 1 GiB free
app_ok = os.path.exists("/run/app/ready") # written after warm-up
return disk_ok and app_ok
class H(BaseHTTPRequestHandler):
def do_GET(self):
state = "Healthy" if healthy() else "Unhealthy"
body = json.dumps({"ApplicationHealthState": state}).encode()
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.end_headers()
self.wfile.write(body)
HTTPServer(("127.0.0.1", 8080), H).serve_forever()
Getting model changes onto instances
The upgrade policy mode is Manual, Automatic or Rolling. Manual leaves instances out of date until you upgrade them. Automatic upgrades instances in no guaranteed order, potentially all at once. Rolling upgrades in batches gated on health, and is the only sensible choice for production. On Flexible, rolling upgrades require the Application Health extension. The rolling upgrade policy has these settings:
maxBatchInstancePercent: share of instances upgraded at once; 20 is the usual default.maxUnhealthyInstancePercent: if more than this share of the whole scale set is unhealthy, before or during the upgrade, it stops.maxUnhealthyUpgradedInstancePercent: if more than this share of upgraded instances is unhealthy afterwards, it is cancelled. This is the setting that catches a bad image.pauseTimeBetweenBatches: wait between batches, as an ISO 8601 duration. Microsoft's own examples show both PT0S and PT2S, so set it explicitly.prioritizeUnhealthyInstances,enableCrossZoneUpgradeandmaxSurge: the last creates new instances from the new model first and deletes old ones after the new batch is healthy, so capacity never dips.
Worked example: 30 instances, batch 20%, so each batch is 6 instances. Microsoft's documentation describes the unhealthy limits as "more than" the percentage, and 20% of 30 is exactly 6. A broken image that fails every instance in the first batch therefore leaves 6 unhealthy upgraded instances, which is not more than 6, and the rollout can continue into a second batch before it stops. Set max unhealthy upgraded below one batch, for example 10% (3 instances), so one failed batch is enough to cancel. Without maxSurge you are down to 24 healthy instances at that point; with it, the old 30 are still serving. Rolling back failed instances on policy breach exists only for Uniform; on Flexible, fix forward by reverting the model, which starts a new rollout.
Autoscale, and why it sometimes refuses to scale in
Autoscale is an Azure Monitor resource attached to the scale set, not a property of it. A profile has minimum, maximum and default capacity, plus rules: a metric, a time window and aggregation, an operator and threshold, and an action with a cooldown. The engine evaluates every 30 to 60 seconds. Give scale-out and scale-in rules the same metric and a wide gap between thresholds.
az monitor autoscale create -g rg-web --resource web \
--resource-type Microsoft.Compute/virtualMachineScaleSets \
--name web-autoscale --min-count 4 --max-count 40 --count 6
az monitor autoscale rule create -g rg-web --autoscale-name web-autoscale \
--condition "Percentage CPU > 65 avg 5m" --scale out 3 --cooldown 5
az monitor autoscale rule create -g rg-web --autoscale-name web-autoscale \
--condition "Percentage CPU < 30 avg 10m" --scale in 1 --cooldown 10Autoscale also protects against flapping: before a scale-in, it estimates the metric after removing instances, and if that would immediately trigger a scale-out, it scales in by fewer instances or skips the action. Microsoft's example: 2 instances at 1,250 threads with rules scale out at 600 per instance and scale in below 600. After scaling to 3 (417 per instance), scaling back to 2 would give 625, so the scale-in is skipped. Widen the gap to scale in below 400 and it works as intended. When autoscale "does nothing", check its run history for flapping entries before assuming it is broken; general patterns are in cloud autoscaling.
Scale-in, protection and graceful exit
On scale-in the Default policy balances across availability zones, then across fault domains on a best-effort basis, then deletes the instance with the highest instance ID. NewestVM and OldestVM keep the zone balancing and then choose by age. Instance protection removes specific instances from consideration: protect-from-scale-in, or protect-from-scale-set-actions, which also blocks upgrades and repairs.
Terminate notifications give an instance a warning before deletion. Enable scheduledEventsProfile.terminateNotificationProfile with notBeforeTimeout between PT5M and PT15M; the instance sees a Terminate event in Scheduled Events and can approve it early once drained. They cannot be enabled on Spot instances.
import json, time, urllib.request
URL = "http://169.254.169.254/metadata/scheduledevents?api-version=2020-07-01"
NAME_URL = "http://169.254.169.254/metadata/instance/compute/name?api-version=2021-02-01&format=text"
def call(method, data=None):
req = urllib.request.Request(URL, data=data, method=method, headers={"Metadata": "true"})
with urllib.request.urlopen(req, timeout=5) as r:
return json.loads(r.read() or b"{}")
ME = urllib.request.urlopen(urllib.request.Request(NAME_URL, headers={"Metadata": "true"})).read().decode()
while True:
for ev in call("GET").get("Events", []):
if ev["EventType"] == "Terminate" and ME in ev["Resources"]: # exact match, not substring
drain_connections() # your code: stop accepting work, finish in-flight requests
body = json.dumps({"StartRequests": [{"EventId": ev["EventId"]}]}).encode()
call("POST", body) # approve only our own event
time.sleep(5)Approve only your own event: when several instances are being deleted together, the deletion waits for every pending event to be approved or to time out, so an instance that never approves delays the others to the full timeout.
Automatic repairs
With automaticRepairsPolicy enabled, Unhealthy or Unknown instances are repaired with Replace (the default: delete and recreate from the latest model), Reimage or Restart. Repairs wait for a grace period of 10 to 90 minutes after any state change, so a booting instance is not killed. At most 5% of instances are repaired at once, one at a time below 20 instances. If replacements stay unhealthy, the platform suspends repairs and sets the service state to Suspended; alert on that state, because a suspended repair loop is a sign the model itself is broken. Protected instances are never repaired.
Mixing Spot and regular capacity
Flexible scale sets support Spot Priority Mix with two settings: baseRegularPriorityCount, the regular VMs always kept, and regularPriorityPercentageAboveBase, the regular share of everything above the base. With a base of 10 and 50%, a scale set of 30 runs 10 base, 10 extra regular and 10 Spot VMs. If all Spot VMs are evicted, 20 regular VMs remain, so size the base and percentage so that the regular part alone can carry the load at reduced performance. Changing the mix applies only to future scaling, not existing instances. Patterns for handling eviction across providers are in spot capacity architecture.
Failure modes
- Shared dependency in the health check: one database blip marks every instance Unhealthy and repairs replace the fleet.
- No explicit outbound on Flexible: instances boot but cannot download packages, and fail health forever.
- Upgrade threshold above one batch: a broken image rolls through several batches before the policy stops it.
- Grace period shorter than boot: repairs replace instances that were still warming up, in a loop, until repairs suspend.
- Close autoscale thresholds: flapping protection blocks scale-in and cost stays high.
- Load balancer probe and health extension disagree: traffic goes to instances the scale set considers unhealthy, or the reverse. Back both with the same health logic, served on a port the load balancer can reach as well as on loopback for the extension; details on probes in the Azure Load Balancer family.
What to do next
- Confirm each scale set's orchestration mode, and plan to recreate Uniform sets as Flexible where you do not need AKS or Service Fabric.
- Deploy the Rich Health States extension against a local-only health endpoint and watch states for a day before enabling automation.
- Switch to Rolling with maxSurge, set the pause explicitly and set max unhealthy upgraded below one batch.
- Enable automatic repairs with a grace period longer than your slowest boot, and alert on Suspended.
- Review autoscale thresholds for a wide gap, and check run history for flapping.
- Add terminate notifications with a drain handler on regular instances, and a Spot eviction handler on Spot ones.