An Azure virtual machine is rarely just a VM. Behind one entry in the portal sit a network interface, a managed OS disk, maybe data disks, a public IP, a network security group, an identity and a placement decision about zones and fault domains. Most surprises in production, from a surprise bill for a stopped machine to a database that loses its data after maintenance, come from not knowing which of those pieces does what.
This article explains Azure VMs from the resource model up: how sizes are named, how disks and caching work, how networking and outbound access changed recently, which availability construct gives which guarantee, how maintenance and Spot eviction reach your code, and how to deploy it all as code. It finishes with a worked deployment and a checklist. Facts that change, such as SLAs and retirements, were checked on 2026-10-03.
Architecture at a glance
The resource model and size names
Azure Resource Manager treats every piece of a VM as its own resource with its own lifecycle. The Microsoft.Compute/virtualMachines resource holds the size, image reference, OS profile and identity, and points at a NIC and disks. The NIC lives in a subnet of a virtual network and may carry a public IP; the NSG filters its traffic. The OS disk is a managed disk, a resource you can snapshot, detach and attach elsewhere. Deleting the VM does not necessarily delete its disks, NIC or IP; the default depends on the tool and on the deleteOption set on each attachment, so orphaned, billed disks are a common finding in cost reviews.
Size names encode capability. Standard_D4ads_v5 reads as family D (general purpose), 4 vCPUs, a for AMD, d for a local temp disk, s for premium-storage capable, version 5. Other families: B burstable, E memory optimised, F compute optimised, L storage optimised with local NVMe, M very large memory, N with GPUs and H for HPC. A p suffix marks Arm-based sizes. Every size also has limits beyond vCPU and RAM: maximum data disks, uncached disk IOPS and throughput, and network bandwidth. Those caps, published per size, matter as much as the core count.
VMs come in generation 1 (BIOS) and generation 2 (UEFI). Use generation 2 for new work: it is required for Trusted launch, which adds Secure Boot and a virtual TPM so boot-kit tampering is detectable and disk encryption keys can be sealed. Microsoft now steers new deployments to Trusted launch by default; confirm what your API version and image actually select.
Disks, caching and the temp disk
A VM can have three kinds of storage, and confusing them is a classic data-loss incident. The OS disk and data disks are managed disks, remote, replicated and persistent. The temp disk, D: on Windows and usually /mnt or /mnt/resource on Linux, is physically on the host. It is fast and free, and its contents are lost when the VM is deallocated, redeployed, resized onto another host or moved by a host failure. Put swap, scratch and caches there, never data you need. Ephemeral OS disks place the OS disk itself on local storage, giving fast reimage and no OS disk charge, which suits stateless scale-set nodes and nothing else.
Managed disk types range from Standard HDD through Standard SSD, Premium SSD, Premium SSD v2 and Ultra Disk. Classic Premium SSD performance is tied to size tier: a P10 (128 GiB) offers 500 IOPS and 100 MB/s, a P30 (1 TiB) 5,000 IOPS and 200 MB/s. Premium SSD v2 and Ultra let you provision capacity, IOPS and throughput separately. The effective limit is always the lower of the disk's limits and the VM size's limits: four P30s promise 20,000 IOPS, but if the size caps uncached IOPS lower, the size wins. Check both before blaming the disks.
Host caching applies per disk: ReadOnly serves repeated reads from the host's cache and suits database data files; None suits write-heavy logs and is required for some workloads; ReadWrite is meant for the OS disk and is risky for application data unless the application handles cache flushing correctly.
Networking and outbound access
Each NIC sits in a subnet and gets a private IP. Enable accelerated networking on supported sizes: it uses SR-IOV to give the guest a virtual function on the host's network card, bypassing the host's virtual switch for lower latency, less jitter and lower CPU use. Apply NSGs at subnet level for coarse policy and at NIC level only where a machine needs exceptions; evaluating both levels is a common source of confusing denials.
Outbound internet access changed. Historically a VM without a public IP still reached the internet through an implicit, shared default outbound address that you neither controlled nor could allow-list. Microsoft is retiring that: after 31 March 2026, newly created virtual networks default to private subnets, so VMs there have no outbound internet unless you add explicit outbound connectivity. Use a NAT gateway on the subnet, a Standard load balancer with outbound rules, a firewall via a user-defined route, or a Standard public IP on the VM. Basic SKU public IPs are retired, so new deployments use Standard, which is closed to inbound traffic until an NSG allows it. Plan outbound explicitly: package updates, monitoring agents and token endpoints outside the VNet all need it.
Availability: sets, zones and scale sets
| Construct | What it protects against | Published VM SLA |
|---|---|---|
| Single VM, Premium SSD or Ultra disks | Nothing beyond the host's own reliability | 99.9% |
| Availability set (2+ VMs) | Rack failure and planned maintenance, via fault and update domains | 99.95% |
| Two or more availability zones | Loss of a whole data centre | 99.99% |
Availability zones are physically separate locations within a region, each with independent power, network and cooling, and supported regions have three. Availability sets spread VMs across fault domains (separate racks) and update domains (groups rebooted at different times during maintenance) inside one data centre. You cannot combine both for one VM; pick zones where the region supports them.
Virtual Machine Scale Sets manage groups of VMs, can span zones and cost nothing themselves; Flexible orchestration is the recommended mode for new deployments, and the scale sets deep dive covers orchestration modes, rolling upgrades and autoscale. Note that a zonal VM's dependencies need zone thinking too: a zonal disk lives in one zone, and a regional resource with no zone placement can still be taken down by one zone's failure.
Lifecycle, maintenance and Spot
Two states look alike and bill differently. Stopped, reached by shutting down from inside the OS or with az vm stop, keeps the VM allocated on its host and keeps billing for compute. Stopped (deallocated), reached with az vm deallocate or the portal's Stop button, releases the host: compute billing stops, disks still bill, the temp disk is wiped, and a dynamic public IP may change on restart. Scripts that shut down from the guest to save money save nothing.
Azure performs host maintenance continuously. Most updates are memory-preserving: the VM is paused briefly while the host is updated, typically for seconds. Some events need a reboot or a move to another host. Your code learns about these through Scheduled Events, a metadata endpoint inside the VM that lists upcoming Freeze, Reboot, Redeploy, Preempt and Terminate events with a not-before time. A process can drain work, then acknowledge the event to let it proceed early.
import json, time, urllib.request
URL = "http://169.254.169.254/metadata/scheduledevents?api-version=2020-07-01"
def call(method="GET", body=None):
req = urllib.request.Request(URL, method=method, headers={"Metadata": "true"},
data=json.dumps(body).encode() if body else None)
with urllib.request.urlopen(req, timeout=5) as r:
return json.loads(r.read() or b"{}")
while True:
doc = call()
for ev in doc.get("Events", []):
if ev["EventType"] in ("Reboot", "Redeploy", "Preempt", "Terminate"):
drain_from_load_balancer() # stop taking new work
checkpoint_in_flight_jobs()
call("POST", {"StartRequests": [{"EventId": ev["EventId"]}]})
time.sleep(5)Spot VMs run on spare capacity at a large discount and can be evicted when Azure needs the capacity or the price exceeds your maximum; a maximum price of -1 means never evict for price. Eviction arrives as a Preempt event with a minimum notice of 30 seconds, and the eviction policy decides whether the VM is deallocated (disks kept, still billed) or deleted. Use Spot for batch work that checkpoints, never for a singleton database.
Metadata and managed identity
The Instance Metadata Service at 169.254.169.254 is reachable only from inside the VM and requires the header Metadata: true, which blocks naive server-side request forgery that cannot set headers. It reports the VM's size, zone, tags and network, and it issues tokens for the VM's managed identity. With a managed identity, code gets Microsoft Entra ID tokens for Key Vault, Storage or other services without any secret on disk; grant the identity the narrowest role at the narrowest scope. Then store no credentials in images, custom data or environment variables.
Deploying as code
Define VMs as code so every machine is reproducible. This Bicep fragment declares a zonal generation 2 VM with Trusted launch, a Premium SSD OS disk deleted with the VM, accelerated networking and a system-assigned identity. Image values and the size are examples; choose ones available in your region.
param location string = resourceGroup().location
param zone string = '1'
param subnetId string
@secure()
param sshPublicKey string
resource nic 'Microsoft.Network/networkInterfaces@2023-09-01' = {
name: 'app1-nic'
location: location
properties: {
enableAcceleratedNetworking: true
ipConfigurations: [ { name: 'ipconfig1', properties: { subnet: { id: subnetId } } } ]
}
}
resource vm 'Microsoft.Compute/virtualMachines@2023-09-01' = {
name: 'app1'
location: location
zones: [ zone ]
identity: { type: 'SystemAssigned' }
properties: {
hardwareProfile: { vmSize: 'Standard_D4ads_v5' }
securityProfile: {
securityType: 'TrustedLaunch'
uefiSettings: { secureBootEnabled: true, vTpmEnabled: true }
}
storageProfile: {
imageReference: { publisher: 'Canonical', offer: 'ubuntu-24_04-lts', sku: 'server', version: 'latest' }
osDisk: { createOption: 'FromImage', deleteOption: 'Delete', managedDisk: { storageAccountType: 'Premium_LRS' } }
}
osProfile: {
computerName: 'app1'
adminUsername: 'azureuser'
linuxConfiguration: {
disablePasswordAuthentication: true
ssh: { publicKeys: [ { path: '/home/azureuser/.ssh/authorized_keys', keyData: sshPublicKey } ] }
}
}
networkProfile: { networkInterfaces: [ { id: nic.id, properties: { deleteOption: 'Delete' } } ] }
}
}
Worked example: an order system
A team moves a three-tier order system to Azure. The web tier is stateless, so it runs as a Flexible scale set across zones 1, 2 and 3 behind a Standard load balancer, with autoscale on CPU and a Scheduled Events handler that drains instances before maintenance. The database runs on two memory-optimised VMs in zones 1 and 2 with synchronous replication; each has Premium SSD data disks with ReadOnly caching for data files and a separate None-cached disk for the transaction log. Nightly report generation runs on Spot VMs that checkpoint to blob storage.
Load testing exposes the first problem: the database VM plateaus at a lower IOPS figure than its four data disks advertise. The size's uncached disk limit is the bottleneck, so they move to the next size up rather than adding disks. The second problem appears in a fresh VNet created after the outbound change: new VMs cannot reach the package mirror. A NAT gateway on the app subnets fixes it and gives a stable egress IP the payment provider can allow-list. The third is cost: a nightly job shut VMs down from inside the guest, and they kept billing until the job was changed to deallocate them.
Failure modes
- Data on the temp disk disappears after deallocation, resize or host failure.
- Stopped, not deallocated keeps billing for compute.
- Size caps throttle disks or network below what the disks or NIC could do.
- No explicit outbound in new VNets breaks updates, agents and external APIs.
- Single-zone dependencies such as one NAT gateway, a zonal disk or a jump box undo a multi-zone design.
- Ignored Scheduled Events turn routine maintenance into dropped requests.
- Orphaned disks and IPs keep billing after VM deletion when deleteOption was not set.
- Capacity errors when a size is constrained in a zone; keep a fallback size list.
Trade-offs
VMs give full control of the OS, kernel and agents, any software, and predictable performance per size. You pay for that with patching, image management and capacity planning that platform services handle for you. Choose zones over availability sets wherever supported; accept the small cross-zone latency for data-centre-level resilience. Choose Spot only where interruption is cheap. If a workload is a stateless HTTP service, compare a container platform or Functions before committing to a VM fleet.
What to do next
- Inventory every VM's disks, NIC, IP and their deleteOption; delete orphans.
- Move anything important off temp disks.
- Check each size's disk and network caps against measured load.
- Add explicit outbound (NAT gateway or firewall) to every subnet that needs internet.
- Spread production across zones and remove single-zone dependencies.
- Install a Scheduled Events handler that drains and acknowledges.
- Use generation 2 images with Trusted launch and managed identities.
- Deallocate, not stop, to save money; tag and schedule non-production VMs.
- Keep learning: Azure managed disks, the Azure load balancer family, Defender for Cloud, VM Scale Sets and, for comparison, Amazon EC2.