The phrase "Azure AI supercomputer for OpenAI" does not name one machine. It names a series of systems Microsoft has built since 2019 to train OpenAI's models, each larger than the last and each changing what training software has to assume. Knowing the lineage is useful even if you will never rent one: the same design decisions, about NVLink domain size, fabric topology, storage and failure handling, show up in every large GPU cluster, including ones you can buy by the rack.

This article traces the generations from public announcements, then explains each layer of the design in terms of what a training job does with it: where tensor, expert, pipeline and data parallelism should go, how much time a gradient all-reduce takes, how big a checkpoint is, how often a job of this size fails, and what changes when training spans sites. For the individual virtual machine sizes and their InfiniBand layout, see Azure ND-series GPU VMs. Microsoft publishes headline figures, not full designs, so anything below that is not attributed to Microsoft is an engineering estimate and is labelled as one.

The lineage

WhenSystemWhat Microsoft or others published
May 2020First OpenAI supercomputerMore than 285,000 CPU cores, 10,000 GPUs, 400 Gb/s of network per GPU server; announced as top five on TOP500. The GPT-3 paper says it trained on V100 GPUs on a Microsoft cluster.
2022-2023A100 clustersMicrosoft described tens of thousands of A100 GPUs on InfiniBand for OpenAI.
Nov 2023Eagle (ND H100 v5)No. 3 on TOP500 at 561 PFlop/s HPL; 1,123,200 cores listed; Xeon Platinum 8480C, H100, NDR InfiniBand.
Oct 2025GB300 NVL72 clusterAnnounced as the first production GB300 NVL72 cluster, more than 4,600 GPUs, for OpenAI workloads.
Sep-Nov 2025Fairwater (Wisconsin, then Atlanta)GB200 NVL72 racks: 72 GPUs per NVLink domain, 800 Gb/s InfiniBand and Ethernet in a non-blocking fat tree, a two-story hall, closed-loop liquid cooling; sites joined by an AI WAN.

The commercial frame changed too. In January 2025 Microsoft announced that it was no longer OpenAI's exclusive compute provider and instead held a right of first refusal on new capacity, as OpenAI added other providers; later revisions to the partnership were reported to loosen that further. So these are systems Microsoft built with and for OpenAI, not a statement about where all of OpenAI's training runs today.

Five scales, one rule

Every generation follows the same hierarchy. Bandwidth per GPU drops and latency rises at each step outward, so a training job is laid out by putting its chattiest traffic on the innermost layer that can hold it.

Five scales, and the parallelism each one should carryGPUHBM, tensor coresNVLink domain8 (HGX) or 72 (NVL72)Scale-out fabricIB or Ethernet fat treeBuildingtwo-story, storageSites over AI WAN~5 ms per 1,000 kmcompute; activation memorytensor and expert parallelism: per-layer, latency-bound trafficpipeline and data parallelism: gradients, stage boundariescheckpoints, data loading, failure domainsonly infrequent synchronisation survives the latencybandwidth per GPU fallsand latency risesat every step down
The layers of an AI supercomputer and the parallelism each should carry. The NVLink domain grew from 8 to 72 GPUs between Eagle and Fairwater.

The single biggest change in the lineage is the size of the NVLink domain. On 8-GPU servers such as ND H100 v5, tensor parallelism above 8 means crossing the scale-out fabric on every layer, so practitioners cap it at 8. In a GB200 NVL72 rack, 72 GPUs share one NVLink domain; Microsoft quotes 1.8 TB/s of NVLink bandwidth per GPU and 14 TB of memory pooled across the rack. Expert parallelism for mixture-of-experts models, whose all-to-all traffic is punishing on a network, can now stay inside one rack. See NVIDIA GB200 for how the rack-scale domain is built.

Outside the rack, Microsoft says each GPU can talk to every other at full line rate through a non-blocking fat tree. That is the property data parallelism and pipeline stages need: no matter which racks a job lands on, collectives see the same bandwidth. Fat-tree topology explains why non-blocking costs so many switches and optics. The two-story hall exists to shorten cables: racks are linked to neighbours above and below as well as beside them, which keeps more links within the reach of shorter, cheaper connections.

A layout planner

Turning the hierarchy into a layout is a small search problem. The planner below takes the NVLink domain size and a candidate layout, checks that the high-traffic dimensions fit inside one domain, and reports how many replicas the data-parallel dimension gets. It is the same reasoning frameworks such as Megatron-style trainers expect you to do when you choose their parallelism flags.

def plan(total_gpus, domain, tp, ep, pp):
    """tp and ep must share one NVLink domain; pp and dp cross the fabric."""
    if domain % (tp * ep):
        raise ValueError(f"tp*ep={tp*ep} does not tile a {domain}-GPU domain")
    per_replica = tp * ep * pp
    if total_gpus % per_replica:
        raise ValueError("pipeline replica does not tile the cluster")
    dp = total_gpus // per_replica
    return dict(tp=tp, ep=ep, pp=pp, dp=dp,
                domains_per_replica=per_replica // domain)

def ring_allreduce_s(bytes_per_rank, ranks, gbps, efficiency=0.7):
    moved = 2 * (ranks - 1) / ranks * bytes_per_rank
    return moved / (gbps / 8 * 1e9 * efficiency)

def checkpoint_tb(params, bytes_per_param=14):
    # bf16 weights (2) + fp32 master copy (4) + Adam first and second moments (8)
    return params * bytes_per_param / 1e12

def job_mtbf_s(gpus, per_gpu_mtbf_h):
    return per_gpu_mtbf_h / gpus * 3600

def checkpoint_interval_s(write_s, mtbf_s):
    return (2 * write_s * mtbf_s) ** 0.5     # Young/Daly first-order optimum

Worked example: 73,728 GPUs

Plan an illustrative 2-trillion-parameter mixture-of-experts model on 1,024 NVL72 racks, 73,728 GPUs. Put tensor parallelism 4 and expert parallelism 18 inside each rack (4 times 18 is 72), and 8 pipeline stages across racks. plan(73728, 72, 4, 18, 8) gives a data-parallel degree of 128, with each replica spanning 8 racks.

Gradient traffic. Treating every parameter as replicated across data parallelism for simplicity, each rank holds about 3.47 billion parameters, so 6.94 GB of BF16 gradients. A ring all-reduce over 128 ranks moves 2 times 127/128 of that, 13.8 GB per rank. With one 400 Gb/s NDR link per GPU, as on the ND GB200 v6 size, at 70% efficiency that is about 0.39 seconds; an 800 Gb/s link per GPU would halve it. Either hides behind backward computation if the framework overlaps it. Sharded optimisers and hierarchical reductions change the constant, not the logic.

Checkpoint size. At 14 bytes per parameter the full training state is 28 TB. Microsoft says a single Blob Storage account in Fairwater sustains over 2 million transactions a second; bandwidth is what matters here, and at an assumed 5 TB/s of aggregate write the checkpoint takes about 6 seconds, at 1 TB/s about 28 seconds. Real systems stage checkpoints to host memory or local NVMe first and drain asynchronously. Training checkpointing covers those patterns.

Failure rate. Meta's Llama 3 paper reported 419 unexpected interruptions in 54 days on 16,384 GPUs, roughly one per 50,700 GPU-hours. Scaled to 73,728 GPUs, the job would be interrupted about every 41 minutes. With a 30-second checkpoint write, the Young/Daly interval is about 385 seconds, and the combined cost of checkpointing and lost work is around 16% before restart time is even counted. That figure is the case for in-memory checkpoints, fast restart and spare racks: at this scale, reliability engineering is a throughput feature.

Failure domains at rack scale

Rack-scale NVLink changes the failure domain. On 8-GPU servers, one bad GPU removes a server; on NVL72, if the tensor and expert groups span the rack, one bad GPU or NVLink switch tray stalls a 72-GPU unit of the job. Operators respond by holding spare racks, running health checks before a node joins a job, and draining suspect hardware proactively from signals such as rising correctable memory errors and link retrains. Microsoft has described the fabric and the hall, not its scheduler internals, so treat any specific mechanism you read about as unconfirmed unless Microsoft published it.

Two consequences for your own code: make restart cheap, by keeping data loader state, random seeds and learning-rate schedule inside the checkpoint so a resumed job is bit-for-bit on track; and make the layout elastic, so losing a rack means restarting with one fewer data-parallel replica rather than waiting for a repair.

Pre-flight: gate every node

None of the arithmetic above holds if one node is slow, and on a system this size some node always is. A slow GPU or a degraded link does not fail a collective; it makes every rank wait for it. So large jobs gate every node before it joins, and the checks are the same whether you rent eight VMs or own a hall:

# 1. GPU health: DCGM diagnostics (level 3 runs longer stress tests)
dcgmi diag -r 3

# 2. Fabric: NCCL all-reduce bandwidth across the candidate nodes,
#    8 bytes to 8 GB, doubling each step, one GPU per process
mpirun -np 16 -H node1:8,node2:8 \
  -x NCCL_DEBUG=INFO ./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1

# 3. Compare bus bandwidth at the largest sizes with the fleet baseline;
#    exclude any node pair more than ~10% below it, then retest

Run the fabric test in groups that match your layout: within a rack for the NVLink domain, then across racks for the fat tree. Keep the results per node, so a node that passes today but trends downward is drained before it slows a job.

Training across sites

Fairwater's newest idea is training across sites. Microsoft says its AI WAN connects Fairwater datacenters, starting with Wisconsin and Atlanta, so that large jobs can span regions. Physics sets the budget: light in fibre covers about 200,000 km per second, so every 1,000 km of fibre path adds about 5 ms one way, thousands of times the latency of a link inside a rack. Bandwidth between sites is also far below the aggregate inside one.

Mount Pleasant and Atlanta are roughly 1,000 km apart in a straight line and further by fibre, so a cross-site message pays at least about 5 ms one way and a round trip at least 10 ms. A ring all-reduce whose ring crosses that link twice per step pays that latency on every chunk boundary, which is why a flat ring spanning sites is the wrong shape. A hierarchical reduction pays it once per step: reduce inside each site over the fat tree, exchange one reduced copy per shard between sites, then broadcast inside each site. With 6.94 GB of gradients per rank, the cross-site exchange is still gigabytes per shard per step, so the WAN must either be very wide or used less often.

Only traffic that is large and infrequent can live on that link. Tensor and expert parallelism cannot; pipeline boundaries usually should not. Data parallelism can, especially with a hierarchical all-reduce, reducing within each site first and exchanging one copy between sites, or with methods that synchronise less often, such as local SGD variants and the DiLoCo approach published by Google DeepMind. Microsoft has not published which of these OpenAI's jobs use, so treat the split as the design space, not a description. GPU pod networks covers the in-site half of the problem.

Failure modes

  • Tensor parallelism across the fabric. Porting an NVL72 layout to 8-GPU servers without shrinking tensor parallelism puts per-layer traffic on InfiniBand.
  • Expert all-to-all crossing racks. Expert groups larger than the NVLink domain turn every MoE layer into a network-bound step.
  • Synchronous checkpoints. A 28 TB blocking write every few minutes erases a large slice of throughput.
  • Fixed world size. Jobs that cannot restart with fewer replicas wait for repairs at a cost of the whole cluster.
  • Treating headline numbers as per-job guarantees. TOP500 HPL and vendor throughput figures measure specific benchmarks; measure your own step time.

Trade-offs

Bigger NVLink domains simplify parallelism and speed up MoE, but they enlarge the failure unit, require liquid cooling and lock you into one vendor's rack design. Non-blocking fat trees give placement freedom at a high cost in switches and optics; oversubscribed or rail-optimised fabrics are cheaper if the scheduler places jobs carefully. Multi-site training buys access to more power than one site can supply, at the price of algorithms that tolerate stale or infrequent synchronisation.

What to do next

  1. Write down the NVLink domain size of your target hardware and set tensor times expert parallelism to tile it.
  2. Run the planner for each candidate layout and keep only those whose data-parallel all-reduce hides behind backward computation.
  3. Size your checkpoint at 14 bytes per parameter, measure storage write bandwidth, and make checkpointing asynchronous.
  4. Estimate job MTBF from your own failure logs, or the Llama 3 rate as a start, and set the checkpoint interval from it.
  5. Make the job restart elastically with fewer data-parallel replicas.
  6. If you will span sites, measure inter-site latency and bandwidth first, and prototype hierarchical or infrequent synchronisation on a small model.
Key takeaway: The Azure supercomputers for OpenAI are a lineage, not one machine, and the biggest change across it is the NVLink domain growing from 8 to 72 GPUs. Put tensor and expert parallelism inside that domain, pipeline and data parallelism on the fat tree, and only infrequent synchronisation across sites; then budget checkpoints and restarts for a failure every hour or less.