Azure Machine Learning does not own GPUs. It rents ordinary Azure virtual machines on your behalf, puts a container runtime and a job agent on them, and gives you a queue, an autoscaler and a launcher on top. Everything that matters for GPU training follows from that: the VM size decides what hardware and network you get, your subscription's quota decides whether you get it at all, and the cluster settings decide how much you pay for nodes that are waiting rather than working.
This article is the GPU-training view of Azure ML compute. The general workspace tour, including data assets, environments and endpoints, is in Azure Machine Learning Studio, in depth. Here we go deeper on one question: how do you turn a PyTorch training script into a multi-node job that uses every GPU you pay for, survives losing nodes, and does not leave expensive machines idling? Settings and behaviours below were checked against Microsoft Learn's compute-cluster and distributed-training guides for the v2 SDK (azure-ai-ml) in October 2026.
Four compute targets, seen from a GPU job
There are four ways to give a job compute, and for GPU work they differ mainly in who manages the lifecycle and how capacity is reserved.
| Target | What it is | Use it for GPU work when |
|---|---|---|
| Compute instance | One managed VM for one person | Interactive debugging on a single GPU; never for long training |
| Compute cluster (AmlCompute) | A named autoscaling pool of one VM size | Repeated multi-node training, shared by a team, with predictable sizes |
| Serverless compute | Per-job capacity, no cluster object | Occasional jobs; you set instance_type and instance_count on the job |
| Attached Kubernetes | An AKS or Arc cluster you operate | GPUs already run under Kubernetes and you want one scheduler |
Clusters and serverless jobs draw on the same Azure ML quota, so the choice is about control, not price. A cluster gives you a stable name, a fixed VM size, network settings and SSH that you choose once. Serverless removes the object you might forget about, at the cost of re-stating the size on every job. Attach Kubernetes only if you already run GPU nodes there; operating it yourself is the point and the cost.
What happens when a job is submitted
When you submit a job, the workspace checks quota and asks the cluster for instance_count nodes. If fewer are running, the autoscaler allocates VMs, which takes minutes for large GPU sizes and longer if the region is short of them. Each node pulls the job's environment image, mounts or downloads the inputs, and the launcher starts your processes with the distributed environment variables set. When the job ends, nodes stay up until the idle timer expires, so a second job can reuse them without the allocation wait.
A GPU cluster and the settings that set its cost
A cluster is defined by five settings. size is the VM size and cannot be changed later, so a cluster is effectively one hardware type. min_instances should be 0 unless you are deliberately paying to keep warm nodes. max_instances caps the largest job and the total the cluster can run at once. idle_time_before_scale_down is seconds of idleness before release. tier chooses dedicated or low-priority VMs.
from azure.ai.ml import MLClient
from azure.ai.ml.entities import AmlCompute
from azure.identity import DefaultAzureCredential
ml = MLClient(DefaultAzureCredential(), "<subscription-id>", "rg-train", "ws-llm")
cluster = AmlCompute(
name="h100-x4",
size="Standard_ND96isr_H100_v5", # 8 x H100 80GB, InfiniBand ("r")
min_instances=0, # pay nothing when idle
max_instances=4, # largest job: 4 nodes, 32 GPUs
idle_time_before_scale_down=1800, # keep nodes 30 min between jobs
tier="dedicated", # or "low_priority"
)
ml.begin_create_or_update(cluster).result()The idle timer is a real trade-off for GPU sizes. Allocating a large node and pulling a multi-gigabyte image can take ten minutes or more, so during a period of rapid iteration a 30-minute timer saves more engineer time than it costs. Overnight it costs you up to half an hour of every node per job. A common pattern is a long timer on a small development cluster and a short one, a few minutes, on the large production cluster. SSH access must be chosen at creation and cannot be enabled later, so decide before you need to debug a hung node.
Quota: the constraint you meet first
Quota is the constraint people meet first. Azure counts dedicated cores per VM family per region, plus a regional total, and the Azure ML compute quota is shared by clusters and compute instances. A GPU node is counted by its vCPUs, not its GPUs: one Standard_ND96isr_H100_v5 is 96 vCPUs, so a four-node cluster needs 384 cores of that family's quota before the first job can run. Low-priority capacity has its own quota, separate from dedicated.
Two behaviours surprise teams. First, Microsoft's documentation says a cluster scaled to zero still has its unprovisioned nodes counted against quota, and only deleting the cluster releases it. A forgotten cluster with max_instances=8 can block someone else's job in the same subscription. Second, you can create a cluster in a different region from the workspace to reach quota or capacity, but you then pay for latency and cross-region data transfer on every read. Copy the data to that region first.
Multi-node PyTorch without a launcher
For data-parallel training you do not need torchrun or a launcher of your own. Set the job's distribution type to PyTorch and process_count_per_instance to the GPUs per node. Azure ML then sets MASTER_ADDR, MASTER_PORT, WORLD_SIZE and NODE_RANK on each node, and RANK and LOCAL_RANK per process. If you leave the process count out, Azure ML starts one process per node, and an eight-GPU node then trains on one GPU while the other seven sit idle and you pay for all eight.
from azure.ai.ml import command, Input, Output
job = command(
code="./src",
command="python train.py --data ${{inputs.corpus}} --ckpt ${{outputs.ckpt}}",
inputs={"corpus": Input(type="uri_folder", path="azureml:corpus-tok:3")},
outputs={"ckpt": Output(type="uri_folder", mode="rw_mount", # fixed path, so a
path="azureml://datastores/workspaceblobstore/paths/ckpt/llm-cpt-01/")}, # resubmit finds it
environment="azureml:train-cu124:7",
compute="h100-x4",
instance_count=2,
distribution={"type": "PyTorch", "process_count_per_instance": 8},
environment_variables={"NCCL_DEBUG": "INFO"},
)
returned = ml.jobs.create_or_update(job)Inside the script, the process group initialises from those variables:
import os, torch, torch.distributed as dist
dist.init_process_group(backend="nccl") # init_method defaults to env://
local_rank = int(os.environ["LOCAL_RANK"])
torch.cuda.set_device(local_rank)
model = torch.nn.parallel.DistributedDataParallel(build_model().cuda(),
device_ids=[local_rank])
if dist.get_rank() == 0:
print("world size", dist.get_world_size()) # expect 16 for 2 x 8Check the printed world size on the first run. It is the cheapest test for the one-process-per-node mistake, and it costs one line.
InfiniBand: confirm it, do not assume it
Multi-node scaling depends on the network between nodes. Azure's convention is that sizes with an r in the name are RDMA-capable and carry InfiniBand. When an AmlCompute cluster uses one of them, the node image comes with the Mellanox OFED driver installed and configured, so NCCL can use InfiniBand without setup. A size without the r falls back to TCP over Ethernet: the job still runs, but gradient all-reduce is far slower and adding nodes can make training slower, not faster.
Verify rather than assume. With NCCL_DEBUG=INFO set as above, the rank-0 log shows which transport NCCL chose: lines naming NET/IB mean InfiniBand, lines naming NET/Socket mean you are on TCP. Then measure scaling directly: run the same job on one node and on two with identical per-GPU batch size, and compare samples per second. Two nodes should give close to twice the throughput; if they give 1.3 times, communication is the bottleneck. All-Reduce, in depth explains how to read NCCL bandwidth numbers against the link speed, and InfiniBand NDR and XDR covers the fabric itself.
Low-priority and Spot capacity
Low-priority VMs on a cluster, or queue_settings={"job_tier": "Spot"} on a serverless job, use spare Azure capacity at a discount. They can be taken back at any time and there is no availability SLA. Microsoft's documentation states that a preempted job must be restarted, so the useful question is how much work a restart throws away. Answer it with checkpoints written to the job's outputs, which live in storage rather than on the node, and a script that resumes from the newest one. By default an output path is tied to the job's name, and a resubmitted job is a new job with an empty folder, so bind the checkpoint output to a fixed datastore path, as the job definition above does, and keep it stable across resubmits:
import glob, os, torch
def latest(ckpt_dir):
files = sorted(glob.glob(os.path.join(ckpt_dir, "step_*.pt")))
return files[-1] if files else None
start = 0
if (path := latest(args.ckpt)):
state = torch.load(path, map_location="cpu")
model.module.load_state_dict(state["model"]); opt.load_state_dict(state["opt"])
start = state["step"] + 1
for step in range(start, total_steps):
train_step()
if step % 500 == 0 and dist.get_rank() == 0:
torch.save({"model": model.module.state_dict(), "opt": opt.state_dict(),
"step": step}, os.path.join(args.ckpt, f"step_{step:07d}.pt"))Low-priority capacity is a good fit for jobs of a few hours that checkpoint often. It is a poor fit for multi-node jobs that need every node at once, because losing one node stops all of them. GPU Spot Instances works through the price per useful hour, including the restart tax.
Worked example: planning a two-node run
A team wants to continue pre-training an 8-billion-parameter model on 2 billion tokens. GPU Throughput Math estimates about 8,200 tokens per second per H100 at 40 percent model FLOPs utilisation for this model size. Two Standard_ND96isr_H100_v5 nodes give 16 GPUs:
tokens = 2.0e9
per_gpu_tps = 8_200 # estimate; measure on your first 200 steps
gpus = 16
goodput = 0.85 # checkpoints, restarts, data stalls
hours = tokens / (per_gpu_tps * gpus * goodput) / 3600
print(round(hours, 1)) # 5.0Planning follows from that number. Quota: 192 dedicated cores of the H100 v5 family in the region, requested days ahead. Cluster: max_instances=2, idle timer of five minutes because there is one long job, not many short ones. Checkpoint every 500 steps so a node failure costs minutes rather than hours. On the first run, check the world size is 16, the log shows NET/IB, and measured tokens per second is within about 20 percent of the estimate. If it is far below, the cause is usually input pipeline stalls or TCP fallback, not the GPUs.
Failure modes
- Cluster stuck resizing at 0 to 0. A delete or read-only lock on the workspace's resource group blocks the scaling operations. Lock individual resources instead, and never the networking resources the cluster creates for itself.
- Job queued for hours. Quota is exhausted, often by other clusters' reserved maximums. Check usage against quota, delete unused clusters, or request more.
- One GPU busy per node.
process_count_per_instancewas not set. The world size printout catches it. - Scaling efficiency under 1.5x on two nodes. A non-RDMA size, or an image that replaced the network stack. Look for NET/Socket in the NCCL log.
- Hung all-reduce after a node fault. All ranks wait forever. Set a collective timeout in
init_process_groupso the job fails fast and can restart. - Cluster on a retired size. The NC, NCv2, ND (P40) and NV series retired in 2023 and NCv3 on 30 September 2025; clusters on them must be recreated on a current size.
Trade-offs
Clusters cost you an object to manage and quota reserved while idle; in return you get warm nodes, fixed networking and one place to set SSH and identity. Serverless costs you repetition in every job definition and a cold start each time; in return nothing outlives the job. Low-priority cuts the price but adds restart risk that grows with node count. Managed Azure ML jobs give you lineage, logs and metrics for free; for very large runs where you need custom schedulers, gang scheduling across hundreds of nodes or direct control of VM images, teams move to virtual machine scale sets with Slurm or Kubernetes and give up the job history. For models served rather than trained, the managed options in Azure AI Foundry are usually the better start.
What to do next
- List your subscription's quota for the GPU family you need in two regions, and request the increase before writing any training code.
- Create one cluster per hardware type with min 0, a deliberate max and an idle timer matched to how often you submit; enable SSH now if you will ever need it.
- Submit a two-node smoke job that prints the world size and runs with NCCL_DEBUG=INFO; confirm 8 processes per node and NET/IB.
- Measure one-node versus two-node throughput at equal per-GPU batch before booking a large run.
- Add checkpoint-and-resume to the script and test it by cancelling a job and resubmitting.
- Review clusters monthly and delete the ones nobody uses, to release their quota.