A 100 MW AI cluster is a site whose utility connection is about 100 megawatts, with almost all of it spent on accelerators and what keeps them running. That buys tens of thousands of GPUs. Several such sites have been built since 2024. xAI's Colossus in Memphis, for example, was reported to hold 100,000 H100s on a 150 MW connection.
The facility side of these sites has its own pages here: power delivery, the grid connection and backup power. This page looks at the same megawatts from the training job's side. How many GPUs does the budget buy? How often will a job that uses all of them break? What checkpoint interval and goodput follow from that? Why does the job's own rhythm show up on the grid as a square wave? Each answer is a small calculation, and the code is included so you can rerun it with your own numbers.
From megawatts to GPUs
Start from the meter and work inward. The utility figure is the contracted peak for the whole site. Cooling, power conversion and lighting take a share, captured by power usage effectiveness (PUE): total facility power divided by IT power. Liquid-cooled AI halls are commonly designed around a PUE of 1.1 to 1.3. At 1.2, a 100 MW site has about 83 MW for IT.
IT power is not GPU power. Each GPU comes with a share of host CPUs, memory, NICs, switches, storage servers and management gear. The cleanest way to plan is an all-in figure per GPU. An 8-GPU H100 server such as the DGX H100 is rated at about 10.2 kW maximum, so roughly 1.3 kW per GPU before the network and storage. Add 10 to 15 percent for those and you get about 1.5 kW per GPU. That matches the Colossus report: 100,000 GPUs at 1.5 kW is 150 MW. A GB200 NVL72 rack is quoted at about 132 kW for 72 GPUs, including Grace CPUs and NVLink switches, or about 1.8 kW per GPU, and about 2.1 kW per GPU once the scale-out network and storage are added.
| Platform | All-in kW per GPU (planning) | GPUs on 83 MW of IT | Racks or servers |
|---|---|---|---|
| HGX/DGX H100, 8 per server | ~1.5 | ~55,000 | ~6,900 servers |
| GB200 NVL72 | ~2.1 | ~39,500 | ~550 racks |
These are planning figures, not nameplate sums. Nameplate overstates real draw, and real draw during training sits below the maximum most of the time. But the planner has to provision for the peak, because synchronized jobs reach it at the same moment. The calculator below keeps every assumption visible:
def gpus_for_site(site_mw, pue, kw_per_gpu_all_in, reserve=0.05):
"""GPUs a site can host. reserve covers growth and measurement error."""
it_mw = site_mw / pue
usable_kw = it_mw * 1000 * (1 - reserve)
return int(usable_kw / kw_per_gpu_all_in)
for name, kw in [("H100 HGX", 1.5), ("GB200 NVL72", 2.1)]:
print(name, gpus_for_site(100, 1.2, kw))
# H100 HGX 52777
# GB200 NVL72 37698The takeaway is that 100 MW is a GPU count of 40,000 to 55,000 for current platforms, and fewer as per-GPU power rises. Everything below is sized for a job using most of them.
The system at a glance
The figure reads left to right on top: the site loses about a sixth of its power to cooling and conversion, and the rest becomes GPUs. The lower half shows three consequences the training job inherits, and the control that answers each one. The rest of the page takes them in turn.
How often a job this size breaks
Meta's Llama 3 paper gives the best public data on failure rates. During a 54-day snapshot of pre-training on 16,384 H100 GPUs, the job was interrupted 466 times. 47 were planned, such as firmware upgrades. 419 were unexpected, and the paper attributes most of those to hardware, with GPU and HBM faults the largest share. That is 7.8 unexpected interruptions a day, or one every 3.1 hours. Even so, Meta reported more than 90 percent effective training time, because it automated detection and restart.
Interruptions scale roughly with the number of components in the job. The job dies when any one of its GPUs, hosts, links or switches fails. Scale the Llama 3 rate linearly to 50,000 GPUs and you get about 24 interruptions a day, one an hour. At 100,000 GPUs it is one every 30 minutes. Linear scaling is an assumption, since newer hardware, burn-in and fleet age all move the rate. But it is the right first estimate, and it changes how the job has to be built. A design that restarts in 20 minutes was fine at 1,000 GPUs. At 100,000 GPUs it spends most of its life restarting.
Worked example: checkpoint interval and goodput
Three costs eat into each mean time between failures (M). The first is C, the time a checkpoint blocks training. The second is the lost work since the last checkpoint, on average half an interval. The third is R, the time to detect the failure, replace or exclude the node, restart every rank and reload state. The Young/Daly approximation gives the checkpoint interval that minimizes the first two: tau = sqrt(2 * C * M). The wasted fraction is then about C/tau + tau/(2*M) + R/M.
from math import sqrt
def goodput(mtbf_s, ckpt_block_s, restart_s):
tau = sqrt(2 * ckpt_block_s * mtbf_s) # Young/Daly interval
waste = ckpt_block_s / tau + tau / (2 * mtbf_s) + restart_s / mtbf_s
return tau, max(0.0, 1 - waste)
mtbf = 31 * 60 # ~100k GPUs, Llama 3 rate scaled linearly
for label, c_s, r_s in [("sync ckpt, slow restart", 60, 600),
("async ckpt, slow restart", 5, 600),
("async ckpt, fast restart", 5, 120)]:
tau, g = goodput(mtbf, c_s, r_s)
print(f"{label:26s} interval {tau/60:4.1f} min goodput {g:.0%}")
# sync ckpt, slow restart interval 7.9 min goodput 42%
# async ckpt, slow restart interval 2.3 min goodput 60%
# async ckpt, fast restart interval 2.3 min goodput 86%The worked example is the point of this page. With a 60-second blocking checkpoint and a 10-minute restart, a 100,000-GPU job does useful work only 42 percent of the time. Making checkpoints asynchronous, so training blocks only while state is copied to host memory, helps less than you might expect. The restart term dominates: 600 seconds of every 1,860 is a third of the machine. Only cutting the restart to two minutes brings goodput to about 86 percent.
That is why large training stacks invest in recovery: health checks before launch, hot spare nodes already in the job's reservation, in-memory checkpoints that peers hold for each other, and rank re-initialization that skips the full scheduler path. Distributed checkpointing covers the sharded save and load path. At 100 MW, the restart path deserves the same engineering attention as the training kernels.
Synchronized power swings
A training step is synchronized. Every GPU computes, then waits on the same collective, then computes again. The Llama 3 paper notes that when tens of thousands of GPUs pause together, for a checkpoint, a collective or a job start or stop, site power can swing by tens of megawatts almost instantly. At 50,000 H100s, a drop from about 700 W to about 150 W per GPU is a swing of 27 MW in the GPUs alone, before the hosts. Generators, UPS systems and the utility all have ramp-rate limits. A swing that repeats every step can also excite oscillations in turbines and grid equipment.
There are three places to fix it. In hardware, NVIDIA's GB300 NVL72 adds energy storage to the rack power shelves (NVIDIA cites 65 joules per GPU) to absorb short dips and peaks. NVIDIA reports a 30 percent drop in peak grid demand while training a Megatron LLM. It also adds a ramp-up power cap that is raised gradually at job start, and a burner mode that keeps power high and tapers it when a job stops abruptly. NVIDIA says these are configured through nvidia-smi or Redfish: GPU-active and GPU-idle floor power, idle time before ramp-down, and ramp-down rate.
In the scheduler, start large jobs in stages rather than launching every rank at once, and drain them in stages. In the job, avoid creating idle cliffs. Overlap checkpoint copies with compute, and keep a power floor instead of letting GPUs drop to idle during long waits. Older platforms without storage in the power shelf can use power caps (nvidia-smi -pl) to trim the peak. That trades a few percent of throughput for a smaller swing, which is often a condition of the utility connection itself.
Fabric depth and placement
A non-blocking fat tree built from switches with k ports serves k²/2 endpoints with two tiers and k³/4 with three. With 64-port switches, two tiers stop at 2,048 endpoints and three tiers reach 65,536. With 128-port switches the figures are 8,192 and 524,288. So a 50,000-GPU job on one network needs three tiers. Each extra tier adds hops, latency and optics, and with optics comes a large population of transceivers that fail. Many sites use a rail-optimized design: GPU i of every server connects to rail i. That keeps most collective traffic within a rail and limits how much crosses the top tier.
The software consequence is placement. Tensor parallelism stays inside the NVLink domain, 8 GPUs on HGX or 72 on NVL72. Pipeline stages and context parallelism come next. Data parallelism, whose all-reduce or reduce-scatter tolerates bandwidth best, crosses the top tier. The scheduler must know the tiers, so describe them to Slurm or Kubernetes, and let the launcher map ranks to the hierarchy. Without that, step time varies from run to run with placement. The Azure OpenAI supercomputer page works through a layout planner for a 73,728-GPU build.
Failure modes
- Sizing on nameplate or on average draw. Nameplate leaves capacity stranded. Average draw trips breakers when every GPU peaks together. Plan on measured peak per platform.
- Restart paths built for small jobs. A restart that re-queues through the scheduler and reloads from remote storage can take longer than the time between failures.
- Checkpoint storms. Tens of thousands of ranks writing at once saturate the storage fabric and stretch the blocking time C. Stagger the writes or copy to host memory first.
- Idle cliffs on the grid. A long synchronous checkpoint or a hang makes the whole site drop to idle power at once, then return. Utilities may write ramp limits into the connection agreement.
- Silent stragglers. One slow GPU or flapping link slows every step without failing the job. At this scale there is always at least one. Track per-rank step time and evict outliers.
- Topology-blind placement. Ranks scattered across top-tier switches make every collective slower and run-to-run timing noisy.
Trade-offs
One 100 MW site running one job gives the largest synchronous run, but it concentrates risk: one grid event, one cooling fault or one bad firmware rollout stops the whole run. Several smaller jobs on the same site lose less to each failure and smooth the power profile, because their steps do not line up. Splitting a single job across sites gets past one site's power limit, but adds wide-area links to the data-parallel path and needs algorithms that synchronize less often. Liquid cooling lowers PUE and raises the GPU count per megawatt, at the cost of plumbing, leak detection and a more complex maintenance process. Power capping and smoothing cost a few percent of throughput, and often decide whether the site gets connected at all.
What to do next
- Compute your site's GPU count with an all-in per-GPU figure measured on your own platform, not nameplate.
- Measure your real interruption rate per 1,000 GPUs over a month, and project it to your target job size.
- Run the goodput calculation with your measured checkpoint and restart times, and fix whichever term dominates first.
- Make checkpoints asynchronous, keep hot spares inside the reservation, and time a full restart end to end.
- Record site power at one-second resolution during job start, checkpoints and job stop, and agree ramp limits with the facility team.
- Give the scheduler the fabric tiers, and keep tensor parallelism inside the NVLink domain.