For a large AI training site the scarce input is often not GPUs but a connection to the electricity grid big enough to run them. Hardware can be ordered in months; a new substation, the transmission upgrades behind it, and the studies that decide what must be built can take years. Utilities have also learned that training clusters are unusual loads: very large, able to swing tens of megawatts in seconds, and capable of dropping off the grid all at once. They are responding with new rules.
This article explains the connection from the side of the people who will run the cluster. It covers what the grid actually sees, the steps of a large-load connection and why each takes time, what training workloads do to the grid, the rules now appearing, and the software controls (ramp limiting, power caps, curtailment through checkpoints) that make a cluster easier to connect and safer to operate, with a worked example and code.
What the grid actually sees
The grid does not see GPUs, racks or jobs. It sees a load at a point of interconnection: a demand in megawatts (MW) and reactive power, how fast that demand changes, how it behaves during voltage and frequency disturbances, and how predictable it is. A few terms recur.
- Transmission versus distribution. Small sites connect to the local distribution network. Sites of tens to hundreds of MW usually need their own substation on the high-voltage transmission system, which brings in regional planning and reliability rules.
- Load versus generator interconnection. The widely reported queues are mostly for generators and storage: Lawrence Berkeley National Laboratory's Queued Up report counted about 2,300 GW waiting at the end of 2024. A data center goes through the utility's load process instead, but the generator backlog still matters, because the new supply that would serve it is stuck there.
- Firm versus flexible service. Firm service promises the full contracted demand at all times, so the system must be built for it at peak. Flexible or interruptible service lets the operator reduce the load in defined conditions and can often be granted sooner.
- Site load versus IT load. The meter sees GPUs plus CPUs, networking, cooling and losses; how GPU count maps to megawatts is worked through in GPU data center deployment.
The connection process, step by step
Names and order differ between utilities and regions, but a large-load connection usually runs through the stages in the diagram.
- Request. The customer submits location, demand by year, ramp schedule and load characteristics. Treat the forecast as an engineering document: utilities size upgrades on it and increasingly charge for capacity that goes unused.
- Screening and deposits. Utilities now charge study fees and require financial security, partly to clear out speculative requests from developers who file in many places for the same project.
- Load and system impact studies. Engineers model power flows, fault currents, voltage stability and contingencies with the new load in place, often alongside every other request in the same area.
- Facilities study. Turns the findings into equipment: a new substation, transformers, breakers, line upgrades, and the cost allocation for each.
- Agreement. An electric service agreement fixes contracted demand, the ramp schedule, technical requirements and any curtailment terms.
- Construction and energisation. Lead times for large power transformers and high-voltage switchgear have stretched to years, and transmission upgrades can need permits and rights of way. Energisation is commonly phased, with demand allowed to rise in steps.
The critical path is rarely the building. It is the slowest upgrade the studies uncover, which is why site selection now starts from available grid capacity rather than land or fibre.
What training workloads do to the grid
Synchronous training makes thousands of GPUs alternate between compute, when they draw near their limit, and communication or checkpointing, when they idle. Aggregated, the site load becomes a near square wave whose swings can reach a large fraction of the site's demand, repeating every training step; the electrical detail is in GPU data center power. To the grid this looks like a periodic disturbance, and if its frequency lines up with mechanical resonances in nearby generators it can stress their shafts. Job starts, crashes and restarts add step changes of the whole cluster's demand at once.
The second risk is ride-through. On 10 July 2024 a failed lightning arrester on a 230 kV line in Virginia caused a persistent fault; automatic reclosing at both ends produced six successive faults in 82 seconds. Data centers in the area saw repeated voltage dips, and their own protection and controls transferred them to backup power. About 1,500 MW of data-center load left the grid across many substations, not because the utility disconnected it but because customer-side equipment did. The sudden surplus pushed frequency and voltage up and operators had to intervene. NERC's review, published in January 2025, flagged simultaneous, voltage-sensitive load loss as a gap in planning, and the reliability bodies have since been working on requirements for large loads.
The rules now appearing
The direction across jurisdictions is consistent: large loads pay more of their connection cost, disclose more about their behaviour, and accept some control during emergencies. Texas is the clearest example. Senate Bill 6, signed on 20 June 2025, applies to loads of 75 MW or more in ERCOT. It introduces a screening study fee of at least $100,000, disclosure and cost-sharing rules, rules for co-locating loads with existing generators, and a requirement that non-critical large loads connecting on or after 31 December 2025 can be remotely disconnected during firm load-shed emergencies, alongside a voluntary demand response programme. Other regions are working on large-load tariffs with minimum-demand charges, longer contract terms and technical standards for ride-through and ramping.
Flexibility is the other half. A 2025 Duke University Nicholas Institute study estimated that the US grid could integrate about 76 GW of new load if that load accepted curtailment averaging 0.25% of its annual energy, about 98 GW at 0.5% and about 126 GW at 1.0%, because systems are built for a few peak hours that rarely occur. The study is an estimate of headroom, not a guarantee for any particular site, but it explains why utilities offer faster connections to loads that can reduce demand on request. Training is unusually well suited to this, because it can be paused at a checkpoint.
Worked example: sizing a request and pricing flexibility
Take an illustrative cluster (the numbers are assumptions, not a vendor specification): 16,384 GPUs, roughly 1.4 kW of IT load per GPU once its share of servers and network is included, and a power usage effectiveness of 1.25.
- IT load: 16,384 x 1.4 kW = about 22.9 MW. Site load: 22.9 x 1.25 = about 28.7 MW. Request around 30 MW with margin, phased.
- Start-up: launching every node at once steps the site from idle to full load within seconds. If the agreement limits ramps to, say, 2 MW per minute, the governor must spread the start over about 14 minutes (28.7 / 2): 2,048 nodes of eight GPUs, so roughly 145 nodes per minute.
- Curtailment: 0.25% of annual energy is equivalent to about 22 full-load hours (8,760 h x 0.0025). The Duke study found events averaging about 1.7 hours at that rate, and in nearly 90% of curtailment hours at least half of the new load could stay on. Most events are therefore partial reductions that power caps absorb without stopping jobs. Even in the worst case, where every event is a full pause costing 30 minutes of lost work (half-hourly checkpoints) plus 15 minutes of restart, a dozen events add about 9 hours: roughly 31 hours, or 0.35% of the year.
Set against a firm connection that might arrive a year or more later, a sub-1% throughput cost is small. The catch is that the cluster must actually perform the curtailment, within the agreed time, every time, and a cluster can only reduce GPU power caps so far before it has to pause jobs outright.
Software that makes a cluster connectable
The controls stack up by speed. The scheduler acts over minutes (staggered starts, pausing jobs), GPU power caps over seconds, and energy storage in power supplies or a site battery below that.
Hardware vendors now build in parts of this. NVIDIA describes GB300 NVL72 power smoothing in three phases: a power cap that rises gradually at job start to respect ramp limits, power shelves with energy storage that absorb the step-to-step swings, and a GPU burn mechanism that tapers power down at job end instead of dropping it. A paper from NVIDIA, Microsoft and OpenAI on power stabilisation for training datacenters argues for combining software, GPU-level and facility-level measures. The governor below is the site-level piece: it enforces the contracted level and the ramp limit, and turns a curtailment signal into a GPU cap or, below the cap floor, a pause. During a normal start the same floor means admitting nodes gradually instead. Emergency curtailment terms usually set their own response time rather than the normal ramp limit; model both.
from dataclasses import dataclass
@dataclass
class GridTerms:
contracted_mw: float # firm demand in the service agreement
ramp_mw_per_min: float # largest allowed change per minute
curtail_mw: float # level promised when the operator asks
class PowerGovernor:
def __init__(self, terms, n_gpus, cap_min_w, cap_max_w, other_it_mw, pue):
self.t, self.n, self.pue = terms, n_gpus, pue
self.cap_min_w, self.cap_max_w, self.other_it_mw = cap_min_w, cap_max_w, other_it_mw
self.site_mw = 0.0
def plan(self, wanted_mw, curtail):
target = self.t.curtail_mw if curtail else min(wanted_mw, self.t.contracted_mw)
step = self.t.ramp_mw_per_min
self.site_mw += max(-step, min(step, target - self.site_mw)) # one-minute tick
gpu_mw = self.site_mw / self.pue - self.other_it_mw
cap_w = gpu_mw * 1e6 / self.n
if cap_w < self.cap_min_w:
# Below the cap floor: when curtailing, pause jobs; when ramping up, admit fewer nodes.
action = "checkpoint_and_pause" if curtail else "stagger_start"
return dict(site_mw=self.site_mw, action=action, cap_w=self.cap_min_w)
return dict(site_mw=self.site_mw, action="run", cap_w=min(cap_w, self.cap_max_w))The governor's output drives real knobs. Per-GPU limits and draw are visible and settable with standard tools; check the allowed range on your hardware first, because the minimum cap is often a large fraction of the maximum:
nvidia-smi --query-gpu=index,power.draw,power.limit,enforced.power.limit,power.min_limit,power.max_limit --format=csv -l 1
sudo nvidia-smi -i 0 -pl 500 # set GPU 0 to 500 W (must lie within min/max limits)In a fleet you would apply caps through your cluster agent or DCGM rather than per-node shell commands, log every change with a timestamp, and test the full curtailment path (signal, governor, scheduler checkpoint, measured site drop) on a schedule, because an untested curtailment commitment is a liability. Model the ramp behaviour in software before energisation: a short simulation of job starts, step oscillation and a crash restart, fed with your measured per-node profile, shows whether you stay inside the agreed ramp limits.
Bridging strategies and trade-offs
| Strategy | What it buys | What it costs |
|---|---|---|
| Phased energisation | start training on part of the capacity | cluster grows in steps; plan jobs to fit |
| Flexible or interruptible service | earlier connection | curtailment hours; tested control path required |
| On-site generation (gas turbines, engines) | power before the grid connection | permits, emissions, fuel logistics, own reliability |
| Site battery or PSU storage | smooths swings, rides short events | capital cost; minutes, not hours, of energy |
| Co-location at a power plant | skips some transmission build | regulatory scrutiny; cost-sharing disputes |
| Choosing a grid-rich site | fastest firm service | may conflict with latency, staff or fibre |
Where the building itself is the constraint rather than the grid, factory-built capacity helps; see modular data centers. Cooling choices change the site-to-IT ratio and therefore the request; see GPU liquid cooling.
Failure modes
- Overstated forecasts. Requesting far more than you will use ties up studies, may incur minimum-demand charges and erodes trust with the utility.
- Unmanaged ramps. A scheduler that restarts every job at once after a fabric outage produces exactly the step the agreement forbids.
- Common-mode ride-through settings. Identical protection settings across every power system in a campus can make the whole site trip together, as in July 2024. Coordinate settings with the utility.
- Curtailment that only works on paper. No tested path from operator signal to measured reduction, or caps that cannot go low enough.
- Ignoring the oscillation. Treating step-to-step power swings as a facility problem until the utility raises it; measure them at the meter early.
What to do next
- Measure one node's power profile through a real training step, checkpoint and restart, and scale it to the planned cluster.
- Write the load forecast by year with the site-to-IT ratio and ramp behaviour stated, and review it with facilities and finance before it goes to the utility.
- Ask the utility for its large-load requirements early: ramp limits, ride-through, telemetry, curtailment terms and fees.
- Price a flexible service option using your own curtailment arithmetic, as in the worked example.
- Build and test the governor path: staggered starts, power caps within the hardware's limits, and checkpoint-and-pause below the cap floor.
- Coordinate protection and ride-through settings with the utility so the campus does not trip as one block.
- Plan bridging options (phasing, storage, on-site generation) with their permits and lead times on the same schedule as the grid upgrades.