Every watt a GPU draws ends up as heat, and in a liquid-cooled AI hall most of it leaves through water: cold plates hand it to a coolant distribution unit (CDU), the CDU to facility water, and the plant decides how it is rejected: through a cooling tower or dry cooler when outdoor conditions allow, through a compressor-driven chiller when they do not. That plant is the subject of this article. It is the slowest, largest and most expensive part of the cooling chain, it sets the floor on PUE, and its failure modes are the ones that can take an entire training cluster down at once rather than one rack at a time.

The CDU side, its telemetry and how to wire its alarms into a scheduler are covered in Liquid Cooling for GPU Datacenters. Here we stay upstream: what a chiller actually does, why AI halls usually end up with two water temperatures, how to size flows and storage with arithmetic you can check, how the plant is controlled, what goes wrong, and what a training platform team should ask of the facility engineers.

The physics: heat, flow and lift

Everything starts from one equation. The heat a water stream carries is Q = m_dot * cp * dT: mass flow times specific heat (about 4.19 kJ per kg per kelvin for water) times the temperature rise across the load. One megawatt at a 10 K rise needs about 24 litres per second. Halve the delta-T and you double the flow, the pipe size and the pumping energy, which is why delta-T is the number a plant engineer watches most closely.

A chiller is a heat pump. Its evaporator pulls heat out of the chilled water; a compressor raises the refrigerant to a higher pressure and temperature; the condenser dumps that heat, plus the compressor's work, to a cooling tower or outdoor air. The compressor's work scales with lift, the gap between the condensing and evaporating temperatures. Warmer chilled water or cooler condenser water means less lift and a higher coefficient of performance (COP, heat moved per unit of electrical work). Operators often quote efficiency as kW per refrigeration ton, where one ton is 3.517 kW of heat removal; lower is better.

This is why AI hardware with warm-water cold plates matters to the plant. ASHRAE's liquid-cooling guidelines classify facility water by its maximum supply temperature: W17, W27, W32, W40, W45 and W+ (above 45 C). Equipment rated for W32 or W40 lets the plant meet most of the year's load with towers or dry coolers alone, with no compressor running. Equipment that needs W17 forces chillers to run nearly all year in most climates. Read the rating, with its flow, pressure drop and water quality requirements, from the server vendor's facility specification.

Plant architecture: two temperature tiers

A typical modern AI hall does not have a single chilled water system. It has two temperature tiers sharing one heat rejection plant, as the diagram shows.

A two-temperature chilled water plant for an AI hallCooling towersor dry coolersChillers (N+1)compressor liftThermal storagestratified tankcondenser watercharge / dischargePlate HXwaterside economisertower waterCold loop18 C supply / 28 C returnchilled waterride-throughfree coolingWarm loop30 C supply / 40 C returnmostly towerschiller trim on hot daysCRAHs / fan wallsresidual air heatCDUs to cold platesGPU and CPU heatPlant controllerstaging, reset, alarmsCluster schedulerpower caps, drainderate signal
Warm water serves the CDUs and is cooled mostly by towers or dry coolers; cold water serves the residual air load and is produced by chillers, with a tank buffering chiller restarts.

The warm loop feeds the CDUs. In the example below it supplies 30 C and returns 40 C. For most hours the towers (through a plate heat exchanger that keeps open-tower water out of the clean loop) or dry coolers produce 30 C water; on the hottest, most humid days chillers trim the last few degrees.

The cold loop feeds CRAH units or fan walls that remove the residual air heat from parts the cold plates do not touch and from room losses. Even in a direct liquid cooled rack, something like 10 to 25 percent of heat may still go to air; the exact split varies by platform. Air coils need colder water, so chillers work here year-round.

Pumping usually follows one of two schemes. Primary-secondary gives each chiller a constant-flow primary pump and moves water to the load through variable-speed secondary pumps, with a decoupler pipe between them. Variable primary flow drops the second pump set and varies flow through the chillers themselves, saving pump energy and plant space at the cost of tighter control, because each chiller has a minimum evaporator flow below which it trips.

Worked example: sizing a 24 MW hall

Take a 24 MW IT hall where the server specification says 85 percent of heat goes to liquid, plus 0.6 MW of room losses to the air side. Both loops are designed for a 10 K rise. The calculator below is the arithmetic a facility engineer would do on the first page of a design review, and it is worth being able to reproduce yourself.

CP = 4.186          # kJ/(kg*K), water near 20 C
RT_KW = 3.517       # one refrigeration ton in kW

def flow_lps(load_kw, delta_t):
    """Water flow (litres/s, ~kg/s) needed to carry load_kw at a given delta-T."""
    return load_kw / (CP * delta_t)

def plant(it_mw, liquid_frac, room_loss_mw, cold_dt=10.0, warm_dt=10.0, chiller_rt=700):
    liquid_kw = it_mw * 1000 * liquid_frac
    air_kw = it_mw * 1000 * (1 - liquid_frac) + room_loss_mw * 1000
    cold_rt = air_kw / RT_KW
    n = 1
    while (n - 1) * chiller_rt < cold_rt:   # N+1: lose any one machine, still carry load
        n += 1
    return {
        "warm_loop_lps": round(flow_lps(liquid_kw, warm_dt)),
        "cold_loop_lps": round(flow_lps(air_kw, cold_dt)),
        "cold_load_rt": round(cold_rt),
        "chillers_n_plus_1": n,
    }

def tank_m3(load_kw, ride_s, usable_dt, efficiency=0.8):
    """Stratified chilled-water tank volume to carry load_kw for ride_s seconds."""
    kj = load_kw * ride_s
    return kj / (CP * usable_dt) / 1000 / efficiency

print(plant(24, 0.85, 0.6))
# {'warm_loop_lps': 487, 'cold_loop_lps': 100, 'cold_load_rt': 1194, 'chillers_n_plus_1': 3}
print(round(tank_m3(4200, 600, 10)))
# 75

Read the output as a design brief. The warm loop must move about 487 litres per second, nearly half a tonne of water every second, which drives pipe diameters and pump power. The cold loop carries 4.2 MW, about 1,194 tons, so three 700-ton chillers give N+1: any one can fail and the other two still carry the load. The tank line says a 75 cubic metre stratified tank, assuming 80 percent of its volume is usable, carries the cold loop for ten minutes at a 10 K usable rise; the ten minutes are an assumption explained in the ride-through section.

Control: staging, reset and GPU load swings

A chiller plant is a control system with a slow plant and fast disturbances. The main loops are:

  • Staging. Chillers are most efficient in a band of partial load, so the controller adds a machine when the running ones exceed a load threshold for some minutes and sheds one when they fall below another. Hysteresis and minimum run times keep it from short-cycling compressors, which wears them out.
  • Chilled water reset. Raising the supply setpoint when the load allows cuts lift and saves energy. Raise it too far and coils lose capacity and humidity control. A common strategy resets on valve position: if no coil valve is more than about 90 percent open, the water is colder than it needs to be.
  • Condenser water reset. Colder tower water lowers lift, but tower fans cost energy; the controller balances the two.
  • Economiser changeover. When the outdoor wet-bulb temperature plus the tower and heat exchanger approaches is below the loop setpoint, the plate heat exchanger carries the load and chillers can stop.

Here is the part that matters to a training platform. A large synchronous job does not draw steady power; a checkpoint or an evaluation pause drops the whole cluster's power for seconds to minutes, and a job crash drops it to idle at once. On the electrical side these swings are a known grid and generator concern; on the water side they show up as step changes in return temperature. The water volume in the loops smooths steps of a few seconds, but a job that stops and restarts every few minutes can push the staging logic to start and stop machines repeatedly. Two things help: ask the facility team to tune staging delays for the load profile you actually produce, and avoid platform behaviour that makes it worse, such as restarting thousands of GPUs in lockstep after every failure instead of ramping.

Ride-through: the first ten minutes

The most important number for availability is what happens in the first ten minutes after a power event. Pumps and CDUs normally sit on UPS, so water keeps moving. Chillers usually do not, because their compressors are large inductive loads, so a utility dip that the UPS rides through can still trip every chiller in the plant. When generators pick up the load, the chillers must restart, and a restart takes minutes: oil pressure, vane positions and refrigerant conditions all have to come back within limits. Some machines offer rapid-restart options that shorten this; the actual time is a vendor figure you should get in writing, not assume.

During that gap the loop temperature climbs. A liquid-cooled GPU tolerates a warmer supply for a while, then begins throttling clocks, and eventually the platform's thermal protection shuts it down. A training job that sees thousands of GPUs throttle simultaneously slows to the pace of the hottest one, as described in GPU Thermal Management. Thermal storage closes the gap: a stratified tank of chilled water, kept charged, discharges into the loop until the chillers are back. The sizing is the tank_m3 function above with the measured restart time plus margin. Tower and dry cooler fans are often not on UPS either, and the warm loop runs close to its limit by design, so ask which plant components are on UPS, generator or neither.

Failure modes

FailureWhat you seeEffect on GPUsMitigation
Chiller trip on power dipAll chillers offline for minutesCold loop warms; air-cooled parts and room heat upThermal storage; rapid-restart chillers; staged restart
Low delta-T syndromeReturn water colder than design; extra chillers stagedPlant runs out of capacity before load doesFix open bypasses, oversized valves, fouled coils
Tower water too warmEconomiser drops out on humid daysWarm loop drifts up; throttling under full loadChiller trim on warm loop; size for design wet-bulb
Pump failureDifferential pressure drops; flow alarmsFast: CDU flow falls in secondsN+1 pumps on UPS; automatic changeover
Fouling or scaleRising approach temperatures over monthsSlow loss of capacityWater treatment; trend approaches; clean on schedule
Controls faultSetpoint hunting, valves oscillatingTemperature swings that look like hardware faultsAlarm on oscillation; manual override procedures
Makeup water lossTower basin level fallsTowers lose capacity within hoursStorage of makeup water; alarms; dual feeds

Low delta-T syndrome is common and invisible. If water returns at 24 C instead of the design 28 C, every litre carries 60 percent of its design heat, so the plant pumps and stages more than the load requires. The fix is mechanical; the detection is data: trend delta-T against load and treat a falling ratio as a defect.

Trade-offs

ChoiceGainsCosts
Water-cooled chillers with towersBest efficiency (low kW per ton)Evaporates water; treatment; legionella management
Air-cooled chillers or dry coolersLittle or no water useHigher energy on hot days; larger footprint
Warmer facility water (W32 or W40)Most hours with no compressor; lower PUELess ride-through headroom; needs warm-water rated hardware
Large thermal storageSurvives chiller restarts; can shift loadSpace, capital cost, stratification management
Variable primary flowLower pump energy, simpler plantHarder control; minimum flow trips

The tension between water and energy is the one that shows up in public reporting: evaporative towers save electricity by consuming water, and dry coolers do the reverse. How those numbers enter PUE and WUE, and where meter boundaries sit, is covered in PUE, in depth. For halls still partly on air, rear-door heat exchangers are a way to push more heat into the cold loop without rebuilding the room.

Operating it with the cluster

Platform teams rarely run the plant, but they can make it easier to run and can read its telemetry. The useful signals are supply and return temperatures and delta-T per loop, plant capacity margin (available tons minus load), the number of chillers running and the age of the last start, tank state of charge, outdoor wet-bulb, and the economiser mode. A minimal derating rule, run by the facility integration service rather than by hand, looks like this:

def plant_derate(t):
    """Return a cluster-wide action from plant telemetry; thresholds are site-specific."""
    margin_rt = t["available_rt"] - t["cold_load_rt"]
    if t["warm_supply_c"] > t["warm_limit_c"]:
        return "cap_power"        # GPUs are near their facility water rating
    if t["chillers_running"] == 0 and t["tank_charge_frac"] < 0.3:
        return "pause_admission"  # riding through on storage, storage running out
    if margin_rt < 0.1 * t["cold_load_rt"]:
        return "hold_new_jobs"    # one more chiller fault would exceed capacity
    return "normal"

The actions map onto mechanisms you already have: a cluster-wide GPU power cap, paused admission and early checkpoints. The goal is to slow a run by a few percent instead of losing hours to a thermal shutdown. Agree thresholds with the facility team, rehearse them in a maintenance window and log every activation.

What to do next

  1. Get the server platform's facility water class, flow per rack, pressure drop and liquid heat fraction from the vendor's specification, not a slide.
  2. Reproduce the flow, tonnage and N+1 arithmetic for your hall with the calculator above.
  3. Ask which plant components are on UPS, generator or neither, and get the measured chiller restart time in writing.
  4. Size or verify thermal storage against that restart time plus margin.
  5. Trend delta-T against load per loop and alert when the ratio falls.
  6. Feed plant margin, tank charge and supply temperature into the scheduler with an agreed derating policy, and rehearse it.
  7. Review job restart behaviour so a cluster-wide failure does not become a cluster-wide power and thermal step. The overview in GPU Datacenter Cooling places this plant in the full cooling chain.
Key takeaway: The chilled water plant is where AI heat finally leaves the building. Size it from the liquid fraction and delta-T, run the liquid loop warm so towers do most of the work, buffer chiller restarts with storage, watch delta-T for hidden waste, and give the scheduler a derating path so a plant fault slows training instead of stopping it.