Power usage effectiveness is total facility energy divided by IT equipment energy. The PUE measurement guide on this site covers how to measure it honestly: meter boundaries, interval versus annual figures, partial load and charging facility energy to a training job. This article is the other half. Before a hall exists, someone has to predict its PUE, and the prediction is built from a loss model: every watt that enters the site and does not reach a GPU, traced to the equipment that consumes it.

A loss model tells you which knobs actually move the number in a liquid-cooled AI hall. Once most heat leaves through cold plates, fans stop being the story; electrical conversion losses and the hours when chillers must run become the story. You will leave with a runnable model, a worked example, the design levers ranked by effect, and the places where cluster software shows up on the facility meter.

Why build a loss model

PUE is a ratio, and every term in the numerator is a physical process with its own efficiency curve. Write it as 1 + overhead / IT. The overhead has three families. Electrical losses happen in transformers, switchgear, the UPS, busway and PDUs, and they scale roughly with load. Cooling energy pays to move heat from the chips to the outdoors: pumps, fans, and chillers when the outdoor air is too warm to reject heat directly. Fixed loads such as lighting, security and offices do not scale with IT at all, which is part of why a half-empty hall scores worse.

The model below evaluates each family hour by hour against an outdoor temperature profile and sums energy over the year. Every coefficient is an illustrative assumption; replace them with vendor curves before deciding anything.

The electrical chain

Start at the utility feed. Medium-voltage transformers, switchgear, the UPS and the low-voltage distribution each lose a small fraction of what passes through them, and the fractions compound. A double-conversion UPS rectifies to DC and inverts back to AC continuously, which is why it protects so well and why it typically loses a few percent. Many UPS designs offer an eco or bypass mode that feeds the load directly and switches to the inverter on a disturbance, reaching around 99 percent efficiency at the cost of a short transfer and less conditioning of the power.

StageIllustrative efficiencyLoss on 20 MW of IT
Double-conversion UPS96%about 830 kW
UPS in eco mode99%about 200 kW
Transformers, switchgear, busway, PDUs98% combinedabout 400 kW

Two things follow. Electrical losses scale with IT load, so they set a floor under PUE that no cooling design removes; in the model they are more than half of all overhead during cool hours. And they are heat released inside the building, so the cooling plant must remove them as well, which is why the model adds them to the heat load before computing cooling energy. Power delivery choices covered in the datacenter power article, such as higher distribution voltage and fewer conversion stages, attack this term directly.

The heat path

Where the overhead in a liquid-cooled GPU hall comes fromUtility feedtotal facility energyTransformers, UPSconversion lossesBusway, PDUsdistribution lossesGPU racksIT energyall of it becomes heatCold platesliquid capture fraction fRoom airthe remaining 1 - fwarm waterCDUs and pumpssmall, always onDry coolersfree cooling if cold enoughChillerstrim when outdoor is too warmoutdoor + approach above supplyCRAH fans and chilled waterElectrical losses andchiller hours dominate
Every watt delivered to the racks becomes heat. The liquid capture fraction decides how much leaves through warm water and how much through room air; outdoor temperature decides whether chillers run.

In a direct-to-chip hall, cold plates remove the heat of GPUs, CPUs and often switches and memory, and a coolant distribution unit, the CDU, exchanges it into a facility water loop. The share of server heat that leaves in liquid is the liquid capture fraction, written f here. It is never 1.0 for cold-plate systems because power supplies, drives, NICs and voltage regulators still shed heat into the air, and that residual needs room air handlers. Confirm the vendor's capture fraction for the exact server.

The liquid side is cheap to run when the outdoor air can cool the facility water without a chiller. A dry cooler, essentially a large radiator with fans, can bring water to within a few degrees of the outdoor dry-bulb temperature; that gap is the approach. If outdoor temperature plus approach is below the required supply temperature, the hour is free cooling and costs only pumps and fans. Otherwise a chiller trims the water down, and chiller energy is heat divided by its coefficient of performance. The liquid cooling article covers loop sizing and what happens to GPUs when a pump fails; here only the energy matters.

Facility water temperature decides the free-cooling hours

The single most powerful design lever is the facility water supply temperature, because it decides how many hours a year are free cooling. ASHRAE classifies liquid cooling by facility supply temperature in classes W17, W27, W32, W40, W45 and W+, where the number is the upper supply limit in degrees Celsius. A hall designed for W32 needs chillers whenever the outdoor dry-bulb exceeds roughly 32 minus the approach; a W40 or W45 hall can reject heat with dry coolers through most summers in temperate climates.

Warmer water is not free. The GPU's cold plate must hold junction temperature with a smaller temperature difference, so it needs more flow or better plates, and silicon running a few degrees hotter leaks slightly more power, which raises IT energy and quietly improves PUE while increasing the bill. The thermal management article explains how clocks respond when that margin runs out. The platform's own inlet specification is the binding constraint: design to it, not to the highest ASHRAE class.

A runnable PUE model

The model is short enough to audit. It evaluates one hour at a given IT load and outdoor temperature, then weights hours by a temperature histogram for the site. Use real hourly weather data for a real decision; the bins here keep the example readable.

from dataclasses import dataclass

@dataclass
class Design:                      # every default is an illustrative assumption
    liquid_fraction: float = 0.80  # share of heat leaving through cold plates
    supply_c: float = 32.0         # facility water supply temperature
    approach_c: float = 6.0        # dry cooler: outdoor + approach must be <= supply
    air_free_c: float = 12.0       # below this, air-side heat needs no chiller
    ups_eff: float = 0.96          # double conversion; about 0.99 in eco mode
    xfmr_dist_loss: float = 0.02   # transformers, switchgear, busway, PDUs
    pump_kw_per_kw: float = 0.015  # CDU and facility pumps per kW of liquid heat
    dry_fan_kw_per_kw: float = 0.02
    chiller_cop: float = 6.0
    crah_kw_per_kw: float = 0.06   # room air handlers per kW of air-side heat
    misc_kw: float = 0.004         # lighting, offices, per kW of DESIGN IT load

def hour_overhead(d, it_kw, design_it_kw, t_out):
    elec = it_kw * (1 / d.ups_eff - 1) + it_kw * d.xfmr_dist_loss
    heat = it_kw + elec                       # losses are heat in the building too
    q_liq = heat * d.liquid_fraction
    q_air = heat - q_liq
    cool = q_liq * d.pump_kw_per_kw + heat * d.dry_fan_kw_per_kw
    if t_out + d.approach_c > d.supply_c:     # dry coolers cannot reach supply temp
        cool += q_liq / d.chiller_cop
    cool += q_air * d.crah_kw_per_kw
    if t_out > d.air_free_c:
        cool += q_air / d.chiller_cop
    return elec + cool + d.misc_kw * design_it_kw

def annual_pue(d, bins, it_kw, design_it_kw):
    total = it = 0.0
    for t_out, hours in bins:                 # sum energy, never average ratios
        total += (it_kw + hour_overhead(d, it_kw, design_it_kw, t_out)) * hours
        it += it_kw * hours
    return total / it

# Illustrative temperate site: (outdoor dry-bulb C, hours per year), sums to 8760
BINS = [(-5, 400), (0, 900), (5, 1400), (10, 1700), (15, 1700),
        (20, 1300), (25, 900), (30, 400), (35, 60)]

It deliberately leaves out humidity, non-linear pump and fan curves, chiller part-load efficiency and standby redundancy. Add them once the simple version is trusted.

Worked example: a 20 MW hall

Take a hall with 20 MW of IT load running at full design load. With the defaults, a cool 5 degree hour carries about 2,250 kW of overhead, an hourly PUE of about 1.11. Of that, roughly 1,230 kW is electrical loss, 425 kW dry cooler fans, 255 kW pumps, 255 kW room air handlers and 80 kW fixed load. On a 30 degree afternoon the chillers come on for both the liquid and air sides and overhead rises to about 5,790 kW, an hourly PUE near 1.29. Over the year the model gives these annual figures:

DesignAnnual PUE, full loadAnnual PUE, half load
80% liquid, 25 C supply1.1731.177
80% liquid, 32 C supply (baseline)1.1371.141
80% liquid, 40 C supply1.1311.135
95% liquid, 40 C supply1.1111.115
95% liquid, 40 C supply, UPS eco mode1.0781.082

Read the table as a ranking. Raising supply temperature from 25 to 32 degrees removes over 2,000 chiller hours and is worth more than three hundredths of PUE. Going on to 40 degrees gains less at this site because only the hottest 460 hours were still on chillers; in a hot climate the same step would be worth far more. Raising capture from 80 to 95 percent shrinks the air side, the most expensive heat path per kilowatt. The UPS mode change is nearly as large as the supply step and the largest left once cooling is efficient, because electrical losses are what remain. Half load barely moves these figures only because the model is linear; real equipment at part load is less kind.

Heat reuse, ERE and the fan boundary

Warm-water halls produce heat at temperatures that are useful elsewhere: district heating networks, greenhouses, or the building's own offices. PUE cannot credit this, because exported heat does not reduce facility energy. The Green Grid's energy reuse effectiveness, ERE, subtracts reused energy from the numerator: ERE is total facility energy minus reused energy, divided by IT energy, so it can drop below 1.0 when a site exports enough heat. Report it next to PUE, never instead of it, and state how reused energy was metered.

The fan boundary problem described in the measurement guide also applies to design comparisons. Air-cooled GPU servers spend a meaningful share of their own power on internal fans, which counts as IT energy. Moving to liquid cuts that IT energy and adds pumps on the facility side, so PUE can rise while total energy falls. The ITUE and TUE metrics proposed by Patterson and colleagues address this by looking inside the server: ITUE is total IT energy over the energy reaching the compute components, and TUE is ITUE times PUE. When comparing cooling designs, compare total energy per unit of work, not PUE.

Where cluster software meets the meter

Cluster software cannot change the design, but it changes the operating point, and the facility meter sees it.

  • Power capping lowers IT energy and raises PUE. Capping GPUs with nvidia-smi -pl or a scheduler policy often costs little throughput per watt saved, but the fixed overhead is spread over less IT energy. Judge caps on energy per training step or per token, never on PUE.
  • Synchronous training swings power. Every rank computes, then waits on communication or a checkpoint, so a large job can move megawatts within a second. Cooling loops respond over minutes, so PUE barely notices; the electrical side and the grid connection are what care.
  • Idle nodes still pay fixed overhead. A hall at 40 percent allocation carries its lighting, standby pumps and UPS no-load losses regardless. Packing jobs densely and powering down whole rows that sit empty for days helps more than any tuning inside a job.
  • Hot hours are expensive hours. If chillers run only a few hundred hours a year, deferrable work such as evaluation sweeps or batch inference can be shifted out of them. The model tells you which hours those are.

Failure modes

The common ways a PUE prediction goes wrong are predictable, and each has a check.

  • Design PUE quoted at full load only. AI halls fill in phases, and the first year runs at partial load with full fixed overhead. Model the ramp, not the end state.
  • Capture fraction taken from a datasheet for a different platform. If the real fraction is 70 percent instead of 85, the air side doubles and the room air handlers may not cope. Measure inlet and outlet water temperatures and flow on the first racks and recompute.
  • Approach assumed, not measured. Dry coolers fouled with dust or sited where exhaust recirculates lose several degrees of approach, and every lost degree converts free-cooling hours into chiller hours.
  • Redundancy left out. N+1 pumps and 2N UPS systems idle at low load, where their efficiency is worst. Model the configuration you will actually run.

Trade-offs

LeverPUE effectWhat it costs
Warmer facility waterLarge where summers are hotBetter cold plates, less thermal margin, slight leakage increase
Higher liquid captureModerateCold plates on more components, more complex servers
UPS eco modeLarge once cooling is efficientTransfer event on disturbances, less power conditioning
Fewer conversion stagesModerateHigher-voltage distribution, different safety practice
Heat reuseNone on PUE, large on ERENeeds a nearby heat customer and a contract

Supply temperature is usually the cheapest lever and the electrical chain the most overlooked. Room air design is covered in the cooling overview.

What to do next

  1. Copy the model, replace the coefficients with vendor curves for your UPS, CDUs, dry coolers and chillers, and feed it a full year of hourly weather data for the site.
  2. Get the liquid capture fraction for your exact server platform and verify it on the first installed racks from water flow and temperature rise.
  3. Run the model at each phase of the fill-in plan, not only at full load.
  4. Rank levers by annual energy saved, then by cost; check supply temperature and UPS mode first.
  5. Once the hall runs, compare measured interval PUE with the model hour by hour, using the methods in the measurement guide, and investigate any gap above a few hundredths.
  6. Evaluate power caps and scheduling policies on energy per step or per token, and report ERE separately if heat is exported.
Key takeaway: In a liquid-cooled AI hall, PUE is set by two things: the electrical conversion chain, which loses a few percent of every watt, and the number of hours chillers must run, which the facility water temperature decides. Model both hour by hour, verify the liquid capture fraction on real racks, and judge efficiency on energy per unit of work rather than PUE alone.