A large training job is a fragile thing to keep powered. Thousands of GPUs run in lock step, a single lost node stalls the collective, and the last checkpoint may be twenty minutes or more behind the current step. A power event that a web-serving fleet would absorb, a few nodes rebooting and traffic shifting elsewhere, can cost a training cluster every step since the last checkpoint plus the time to restart thousands of processes.

Backup power is the chain of devices that stops that from happening: power-supply holdup capacitors, rack battery units, room or hall uninterruptible power supplies, transfer switches and generators. Each covers a different time window, and the right design for an AI site depends less on any single device than on what the software does inside each window. This article walks through the chain in time order, explains why GPU loads are hard on it, works through a ride-through budget for a training cluster, and shows how to wire power events into checkpointing so a long outage costs minutes of work instead of hours. Topology choices such as N+1 versus 2N and the electrical shape of training loads are covered in GPU datacenter power requirements; this page builds on them.

The ride-through timeline

Most utility disturbances are short voltage sags, not complete outages, and the chain is designed so each layer handles the events it is fast enough for and hands longer events to the next.

Backup power chain and the time window each layer coversUtility feedMV serviceGeneratorsstart in ~10-15 sTransfer switchATS / switchgearUPS + batteriesbridges seconds-minutesPDU / buswayto the rowsRack: power shelf + BBU shelfBBU: minutes; PSU holdup: ~10-20 msGPUs + training jobsees only what leaks throughTelemetry pathUPS / BBU / ATS eventsSNMP, Redfish or NUT to a power-event daemonwhich tells the scheduler to checkpoint or drain
The backup chain. Each layer buys time for the next one to take over; the telemetry path lets software act while the batteries are still carrying the load.
WindowWhat carries the loadTypical scaleWhat software can do
0-20 msPSU bulk capacitors (holdup)Roughly 10-20 ms at full loadNothing; too fast
20 ms to secondsUPS or rack BBU switches in or is already onlineInstant for double-conversion UPSLog the event
~10-60 sBatteries or flywheel while generators start and syncGenerator start and transferDecide: wait or checkpoint
MinutesGenerators, or batteries if generators failedBattery autonomy is often 5-10 minEmergency checkpoint, drain
Hours to daysGenerators on stored fuelFuel autonomy set by site designNormal running, reduced redundancy

The figures above are typical design targets, not guarantees; read your site's actual single-line diagram and test reports. The important structural point is that the software decision window sits between the generator start and the battery running out. If the generators come on line, the job should keep running. If they do not, the job has however many minutes of battery remain to save its state.

Each layer in the chain

Rack-level holdup and BBUs. Every server power supply has bulk capacitors that ride through very short dips. Open Compute Project ORv3 racks go further with a battery backup unit shelf beside the power shelf, feeding the same DC busbar. The BBU is sized for minutes, enough to bridge to a generator or let the rack shut down cleanly, which lets some designs drop the large central UPS. Distinguish this from smoothing storage: NVIDIA's GB300 NVL72 power shelves include electrolytic capacitors, which NVIDIA states store 65 joules per GPU and cut peak grid demand by about 30 percent. At roughly a kilowatt per GPU, 65 J is about 65 milliseconds of full power. That flattens load swings; it is not backup.

Central UPS. A double-conversion (online) UPS continuously rectifies incoming AC to DC and inverts it back, so the batteries are always in circuit and the transfer is seamless. It also isolates the load from utility noise. The cost is a few percent conversion loss on every watt, which is why many operators run eco or line-interactive modes that bypass the inverter under clean power and switch in within milliseconds when the input degrades. Batteries are either valve-regulated lead-acid or, increasingly, lithium-ion, which is smaller and longer-lived but brings thermal-runaway requirements under codes such as NFPA 855. Flywheel UPSs store energy mechanically and typically carry the load for tens of seconds, which is enough only if the generators are reliable.

Transfer and generators. An automatic transfer switch or paralleling switchgear detects the outage, signals the generators to start, waits for them to reach stable voltage and frequency, then moves the load. Emergency systems under NFPA 110 Type 10 must restore power within 10 seconds, and datacenter standby plants are commonly designed to a similar order. Generators are almost always diesel, sized for the full critical load with spare units, and tested regularly under load.

Why GPU loads are hard on backup power

Backup equipment was designed for loads that change slowly. GPU training is not like that, and three properties stress the chain.

Large synchronized steps. All GPUs in a job move between compute and communication phases together, so the hall load swings by a large fraction of its total within milliseconds. Meta's Llama 3 paper reported power fluctuations across a datacenter on the order of tens of megawatts during training. A UPS inverter handles this well. A diesel generator does not: its governor reacts in hundreds of milliseconds, and a large step causes frequency and voltage dips. ISO 8528-5 defines transient performance classes for exactly this, and the generator must be specified for the step sizes the GPU load can actually produce, not only for steady kilowatts.

Restart surge. When power returns, or when a job restarts after an event, every node loads data and starts computing at once. Staggering job start and using a ramped power cap, as GB300's power-cap feature does at the hardware level, keeps the restart from tripping the generator or upstream breakers.

Cooling must ride through too. A liquid-cooled rack at full load has only seconds to tens of seconds of thermal mass before GPUs throttle or trip. Pumps, coolant distribution units and controls must be on UPS power, not only the IT load, or a generator transfer that the GPUs survive electrically still ends the job thermally. The thermal side is detailed in liquid cooling for GPU datacenters.

Worked example: a ride-through budget for a training job

Take a cluster of 2,048 GPUs training a 70-billion-parameter model with mixed precision and Adam. Full training state is roughly 14 bytes per parameter (bf16 weights, fp32 master weights and two fp32 optimizer moments), about 980 GB. The parallel file system sustains 100 GB/s of aggregate writes, so a synchronous checkpoint takes about 10 seconds of write time plus perhaps 20 seconds of coordination: call it 30 seconds end to end.

The facility has 2N UPS with 6 minutes of battery at full load and generators specified to accept load within 15 seconds. The job checkpoints every 30 minutes, so a sudden total loss destroys 15 minutes of progress on average, plus around 20 minutes to restart and reload: about 35 minutes of 2,048 GPUs, or roughly 1,200 GPU-hours per incident.

Now the decision. When the UPS reports on-battery, wait 30 seconds, which is twice the generator target. If the load is back on generator, continue and lower the checkpoint interval while redundancy is reduced. If it is still on battery, trigger an emergency checkpoint: 30 seconds of work against more than 5 minutes of remaining battery, so the margin is wide. The worst case drops from 35 minutes of lost work to about 21 minutes of restart, and the common case, generator transfer, costs nothing. The calculation also says when this is not worth it: if checkpoint time exceeded remaining battery, you would skip the emergency write and invest in faster checkpointing instead. Techniques for fast and asynchronous checkpoints are in training checkpointing.

Wiring power events into the scheduler

Most UPSs and BBU shelves expose status over SNMP, Redfish or Modbus, and Network UPS Tools (NUT) normalises many of them: ups.status contains flags such as OL (on line), OB (on battery) and LB (low battery), and battery.runtime estimates remaining seconds. A small daemon turns those into scheduler actions.

import subprocess, time

CHECKPOINT_S = 30          # measured end-to-end emergency checkpoint time
GENERATOR_GRACE_S = 30     # 2x the generator acceptance target
SAFETY = 2.0               # require this much battery margin over checkpoint time

def ups_vars(name="hall-a-ups@localhost"):
    out = subprocess.run(["upsc", name], capture_output=True, text=True, check=True).stdout
    return dict(line.split(": ", 1) for line in out.splitlines() if ": " in line)

def watch(scheduler, poll_s=1.0):
    on_battery_since = None
    acted = False
    while True:
        v = ups_vars()
        flags = v.get("ups.status", "").split()
        runtime = float(v.get("battery.runtime", "0"))
        if "OB" in flags:
            on_battery_since = on_battery_since or time.monotonic()
            waited = time.monotonic() - on_battery_since
            if not acted and ("LB" in flags or waited > GENERATOR_GRACE_S):
                if runtime > SAFETY * CHECKPOINT_S:
                    scheduler.emergency_checkpoint_then_pause(reason="power")
                else:
                    scheduler.drain(reason="power, no time to checkpoint")
                acted = True
        else:
            if on_battery_since is not None:
                scheduler.note("power restored", reduced_redundancy=True)
            on_battery_since, acted = None, False
        time.sleep(poll_s)

In production you would consume traps or Redfish events rather than polling, aggregate per power domain so the scheduler only pauses jobs actually fed by the affected UPS, and make the scheduler call idempotent. Measure CHECKPOINT_S from real runs at current model size; it grows as the model and optimizer state grow.

Test the path the way you test any other recovery mechanism. A scheduled generator transfer drill is the cheapest opportunity: run a sacrificial job, open the utility breaker as the facility team would in a normal test, and confirm that the daemon saw the on-battery event, waited out the grace period, did nothing when the generators picked up the load, and logged the reduced-redundancy state. A separate drill with the generators deliberately held off, on a small isolated power domain, proves the emergency checkpoint actually completes inside the battery window. Without both drills, the first real outage is the test.

Failure modes

  • Generator fails to start or accept load. The single most important failure the batteries exist for. Mitigate with regular load-bank tests at realistic step sizes and with the emergency-checkpoint path above.
  • Transfer switch fault. A transfer that does not complete leaves the load on a draining battery while the generators run. Monitor ATS position, not just generator status.
  • Aged batteries. Autonomy shrinks as strings age. Trust the measured runtime from periodic discharge tests, not nameplate minutes.
  • Maintenance bypass. A UPS in bypass passes utility disturbances straight through. Schedule bypass windows against training checkpoints and tag them in the scheduler.
  • Cooling not on backup. Electrically protected GPUs overheat when CDUs stop. Verify pumps and controls in the single-line diagram.
  • Telemetry gap. The power-event daemon cannot reach the UPS because its management switch is not on backup power. Put the telemetry path on the protected bus.
  • Correlated node failures after an event. Brownouts can leave some nodes half-working with bad memory or links; run health checks before resuming, as described in GPU hardware faults.

Trade-offs

Design choiceBuysCosts
Central double-conversion UPSSeamless transfer, power conditioningConversion loss, floor space
Rack BBU (ORv3 style)Battery close to load, scales with racksBatteries spread across the hall, more units to monitor
Lithium-ion over lead-acidSmaller, longer lifeHigher cost, fire-code requirements
FlywheelNo batteries, long service lifeOnly tens of seconds; generators must be reliable
Reduced backup for training hallsLower capital costEvery utility outage stops jobs; needs strong checkpointing

Some operators now give training capacity less backup than inference capacity, on the reasoning that a training job can resume from a checkpoint while a user-facing service cannot pause. That is a legitimate trade only when checkpointing is frequent and fast, restarts are automated, and the grid at the site is reliable. The interaction between backup plant, on-site generation and the utility is also shaping interconnection agreements; see grid connection challenges for AI.

What to do next

  1. Obtain the single-line diagram for your cluster and mark which UPS, BBU and generator feed each rack, CDU and management switch.
  2. Record the measured battery autonomy and generator acceptance time from the last tests, not the nameplate values.
  3. Measure your end-to-end emergency checkpoint time at the current model size.
  4. Run the ride-through budget: if remaining battery after generator grace exceeds twice the checkpoint time, implement the emergency-checkpoint path; otherwise invest in faster checkpoints first.
  5. Wire UPS and ATS events into the scheduler per power domain, and test it in a scheduled transfer drill with a real job running.
  6. Add a post-event health check gate before jobs resume, and ramp restarts rather than starting every node at once.
Key takeaway: Backup power is a sequence of time windows: capacitors for milliseconds, batteries for seconds to minutes, generators for hours. For AI training the decisive window is between generator start and battery exhaustion, and the job only benefits if software uses it: know your measured autonomy, measure your checkpoint time, and let UPS events trigger an emergency checkpoint when the generators do not arrive.