Azure's ND sizes are the virtual machines built for training and serving large models: eight data-centre GPUs joined by a high-speed in-box fabric, one InfiniBand adapter per GPU for traffic between machines, and a host big enough to feed them. If you train on Azure Machine Learning, use Azure Kubernetes Service GPU node pools, or run Slurm on virtual machine scale sets, these are the machines underneath.
This article explains what each current ND size contains, how the network inside and between VMs is laid out, what that means for NCCL and your parallelism choices, how to pick a size from a model's memory needs, and how to keep a fleet of them healthy. Specifications come from Microsoft Learn's size pages as of October 2026. Prices and regional availability change often, so check them in the Azure pricing calculator rather than trusting a number in an article.
The current sizes
The ND family has run from P40s to Blackwell. The older ND (P40) series retired in 2023. These are the sizes that matter for current work:
| Size | GPUs | Memory per GPU | Scale-out per GPU | Host |
|---|---|---|---|---|
| ND A100 v4 (Standard_ND96asr_v4) | 8x A100 | 40 GB | 200 Gb/s HDR | 96 vCPU AMD, 900 GiB |
| NDm A100 v4 | 8x A100 | 80 GB | 200 Gb/s HDR | 96 vCPU AMD, 1,900 GiB |
| ND H100 v5 (Standard_ND96isr_H100_v5) | 8x H100 | 80 GB | 400 Gb/s NDR | 96 vCPU Intel, 1,900 GiB |
| ND H200 v5 (Standard_ND96isr_H200_v5) | 8x H200 | 141 GB | 400 Gb/s NDR | 96 vCPU Intel, 1,850 GiB |
| ND MI300X v5 (Standard_ND96isr_MI300X_v5) | 8x MI300X | 192 GB | 400 Gb/s NDR | 96 vCPU Intel, 1,850 GiB |
| ND GB200 v6 (Standard_ND128isr_NDR_GB200_v6) | 4x Blackwell | 192 GiB | 400 Gb/s NDR | 2 Grace (Arm), 128 vCPU, 900 GiB |
Read the size name as a code. In Standard_ND96isr_H100_v5, 96 is the vCPU count, i marks an isolated size that takes the whole host, s means premium storage support, r means RDMA, and the suffix is the GPU and generation. The r matters most: it means InfiniBand. Quota is counted in vCPUs, so one H100 v5 VM takes 96 cores of its family's quota.
Inside one VM, and between VMs
The eight-GPU sizes share one design. Inside the VM, every GPU connects to every other through NVSwitch, so tensor-parallel traffic between any two GPUs is fast and the same speed. Between VMs, each GPU has its own 400 Gb/s ConnectX-7 InfiniBand adapter, 3.2 Tb/s per VM, which Microsoft describes as dedicated and topology-agnostic. Each adapter attaches to the scale set's InfiniBand fabric, NCCL pairs adapters with GPUs by affinity using the topology file, and GPUDirect RDMA moves data between GPU memory and the adapter without passing through host memory.
This tells you how to map parallelism. Tensor parallelism exchanges activations on every layer and needs the in-box fabric, so keep the tensor-parallel group within one VM: eight GPUs on H100 or H200. Data parallelism and pipeline stages communicate less often and can cross VMs over InfiniBand. GPU Pod Network covers this mapping and rail-optimised fabrics in general.
Two variations matter. MI300X v5 connects its GPUs with AMD Infinity Fabric, 128 GB/s per link and 896 GB/s aggregate per GPU, and uses AMD's RCCL library instead of NCCL; your framework must be built for ROCm. GB200 v6 is different in kind: each VM is two Grace CPUs and four Blackwell GPUs on Arm, and 18 VMs make up one GB200 NVL72 rack whose 72 GPUs share a single NVLink domain. Tensor and expert parallelism can then span far more than eight GPUs without leaving NVLink. Images, containers and Python wheels must be built for Arm64. NVIDIA GB200, in depth explains the rack-scale domain.
Scale sets, images and topology
Two rules decide whether InfiniBand works. First, Azure configures the InfiniBand connections automatically only between VMs in the same virtual machine scale set. Two ND VMs created separately cannot reach each other over InfiniBand, so a training cluster must be one scale set, or a service that creates one for you, such as Azure ML clusters, AKS node pools or CycleCloud. Second, the guest needs the right drivers. Microsoft's AI and HPC marketplace images ship the GPU driver, the InfiniBand stack and NCCL, and Microsoft's HPC guidance says they also ship NCCL topology files under /opt/microsoft that tell NCCL which adapter is closest to which GPU. Start from those images; building your own stack is possible but is a project of its own.
Check a new VM before trusting it:
nvidia-smi -L | wc -l # expect 8 GPUs (4 on GB200 v6)
nvidia-smi topo -m # GPU-to-NIC affinity matrix
ibstat | grep -E "State|Rate" # every port Active, Rate 400
ls /opt/microsoft/*topo*.xml # topology file for this size, if the image has one
# Two-VM bandwidth check with nccl-tests (one process per GPU)
mpirun -np 16 -npernode 8 --hostfile hosts \
-x NCCL_DEBUG=WARN -x NCCL_TOPO_FILE=/opt/microsoft/ndv5-topo.xml \
./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1The last command prints bus bandwidth for message sizes from 8 bytes to 8 GB. Record the large-message figure on a known-good pair of VMs and treat any later result well below it as a fault. All-Reduce, in depth explains how bus bandwidth relates to the 400 Gb/s per adapter.
Worked example: choosing a size by memory
Choose a size by memory first, because memory decides how many GPUs one copy of the model needs, and that decides how much expensive communication you buy. Take serving Llama 3.1 70B in BF16. Its weights are about 140 GB. Its KV cache, with 80 layers, 8 KV heads of dimension 128 and 2-byte values, costs 2 x 80 x 8 x 128 x 2 = 327,680 bytes per token. Reserving 10 percent of each GPU for activations and overhead:
weights, kv_tok, reserve = 140e9, 327_680, 0.10
for name, hbm, tp in [("H100 v5, TP=4", 80e9, 4),
("H200 v5, TP=2", 141e9, 2),
("MI300X v5, TP=1", 192e9, 1)]:
free = tp * hbm * (1 - reserve) - weights
print(f"{name}: {free/1e9:5.1f} GB for KV = {free/kv_tok/1e3:4.0f}k tokens,"
f" {8 // tp} replicas per VM")
# H100 v5, TP=4: 148.0 GB for KV = 452k tokens, 2 replicas per VM
# H200 v5, TP=2: 113.8 GB for KV = 347k tokens, 4 replicas per VM
# MI300X v5, TP=1: 32.8 GB for KV = 100k tokens, 8 replicas per VMNone of these is simply best. H100 with four-way tensor parallelism has the most KV room per replica, so it serves long contexts and large batches, but every token pays for an all-reduce across four GPUs. MI300X fits the whole model on one GPU and avoids tensor-parallel traffic entirely, but leaves only about 100,000 tokens of KV, about 50 concurrent 2,000-token conversations, so long-context traffic will be throttled by memory. H200 sits between them. The right answer depends on your context lengths and latency targets; KV Cache Sizing for Deployments has the full method, and NVIDIA H200, in depth explains why its extra bandwidth also speeds decode.
For training the same logic applies to weights, gradients and optimizer state, which with mixed-precision Adam need roughly 16 bytes per parameter before activations. More memory per GPU means less sharding and fewer collectives per step.
Maintenance and node health
ND sizes do not support live migration or memory-preserving updates, so when Azure needs to service a host, your VM is rebooted or redeployed rather than moved while running. For a multi-node job, losing one VM stops every rank. Two habits handle this. Watch Azure Scheduled Events from inside each VM, through the instance metadata endpoint, and checkpoint when maintenance is announced. And run a health check before a node joins a job, not after it has hung one:
#!/bin/bash
# node_gate.sh: exit non-zero and the scheduler should drain this node
# Counts are for the 8-GPU sizes; use 4 on GB200 v6
set -e
[ "$(nvidia-smi -L | wc -l)" -eq 8 ] || { echo "missing GPU"; exit 1; }
nvidia-smi --query-gpu=ecc.errors.uncorrected.volatile.total --format=csv,noheader \
| awk '$1 > 0 {bad=1} END {exit bad}' || { echo "uncorrected ECC"; exit 1; }
dmesg | grep -q "NVRM: Xid" && { echo "Xid errors in kernel log"; exit 1; }
[ "$(ibstat | grep -c 'State: Active')" -ge 8 ] || { echo "IB port down"; exit 1; }
curl -s -H Metadata:true \
"http://169.254.169.254/metadata/scheduledevents?api-version=2020-07-01" \
| grep -q '"EventType"' && { echo "maintenance scheduled"; exit 1; }
echo okAdd a short nccl-tests run against a known-good neighbour for nodes returning from repair. Slurm's health-check hook or a Kubernetes init container can run the gate automatically.
Treat local storage the same way. The temp disk and the NVMe drives, up to eight disks and 28 TiB on the H100, H200 and MI300X sizes, are fast and free to use, which makes them the right place to cache a dataset or stage a checkpoint before upload. They are also ephemeral: a redeploy for maintenance or a deallocation empties them. Copy every checkpoint to durable storage, such as Blob Storage or a managed Lustre file system, before deleting the previous one, and make the job able to re-stage its data cache from scratch.
Failure modes
- Multi-node job falls back to TCP. VMs are in different scale sets, or the image lacks the InfiniBand stack. NCCL logs show NET/Socket instead of NET/IB.
- Bandwidth half of normal on one pair. A port has negotiated a lower rate or is down; ibstat shows it. Drain the node and report it.
- Throughput lower than the same job elsewhere. NCCL is not using the topology file, so GPUs talk through distant adapters. Set NCCL_TOPO_FILE explicitly.
- Model fits in testing, runs out of memory in production. The KV budget was sized for average context, and long prompts arrived.
- Containers fail to start on GB200 v6. The image is x86-64 only; build for Arm64.
- A job dies at the same hour every few weeks. Host maintenance with no live migration; handle Scheduled Events and checkpoint.
Trade-offs
H100 v5 has the widest software support and the most capacity, and is the safe default for NVIDIA-based training. H200 v5 runs the same software with 76 percent more memory and higher bandwidth, which favours serving large models and long contexts. MI300X v5 has the most memory per GPU of the eight-GPU sizes, which reduces parallelism for large models, but requires a ROCm build of your stack and testing of every custom kernel. GB200 v6 gives a 72-GPU NVLink domain, which matters most for very large mixture-of-experts and tensor-parallel models, at the cost of an Arm host and rack-sized units of capacity. For chip-level detail see AMD MI300X, in depth and this site's H100 and H200 pages.
What to do next
- Write down the memory one copy of your model needs (weights, KV or optimizer state) and compute GPUs per replica for each candidate size with the script above.
- Request quota in vCPUs for the chosen family in two regions.
- Deploy two VMs from the AI and HPC marketplace image in one scale set and record their nccl-tests all-reduce bandwidth as your baseline.
- Keep tensor-parallel groups inside one VM, or one NVL72 rack on GB200.
- Install the node health gate and a Scheduled Events watcher before the first long job.
- Re-run the baseline after every image or driver update.