Once a rack draws more than a few tens of kilowatts, air can no longer carry the heat away and the cooling medium becomes water or a water-glycol mix pumped through cold plates bolted onto every GPU. The comparison of cooling technologies, rear-door exchangers versus direct-to-chip versus immersion, is covered in our GPU datacenter cooling overview, and what the GPU does internally when it gets hot is in GPU thermal management. This page is about the part between them: the liquid loop as a system that software teams have to understand, size, monitor and react to.
The core argument is simple. In a liquid-cooled cluster the cooling plant is part of the compute system. Its flow rate sets how much power a rack can draw, its supply temperature shifts GPU clocks, and a fault in it can take down a 1,000-GPU training job in seconds rather than minutes. If your scheduler, alerting and runbooks do not know the loop exists, you will learn about it the expensive way.
The loop, end to end
There are two loops, kept apart by a heat exchanger. The technology cooling system (TCS) or secondary loop is the clean, treated coolant that touches the servers. It leaves the coolant distribution unit (CDU), travels through supply manifolds on the rack, splits into parallel paths through each server's cold plates (GPU, often CPU and sometimes memory or network chips), and returns warmer. The CDU contains the pumps, the liquid-to-liquid heat exchanger, filtration and the controller. The facility water system on the other side takes the heat to cooling towers, dry coolers or chillers.
CDUs come in-rack (serving one rack) or in-row (serving several). The separation matters operationally: facility water quality and pressure problems stay on the primary side, while the secondary loop can be kept small, filtered and chemically controlled. It also means a single CDU fault can affect every rack it serves, which is the blast radius your scheduler needs to know about.
Not all heat goes into the liquid. Power supplies, fans, drives and NICs still dump heat into the room, so a liquid-cooled hall still needs air handling sized for the residual. The fraction of rack power that leaves in the coolant is the capture ratio, and it is one of the first numbers to measure on new hardware.
Heat balance: the one equation you need
Every sizing question reduces to the heat balance Q = m_dot * cp * dT: heat carried (kW) equals mass flow (kg/s) times the coolant's specific heat (kJ per kg per kelvin) times the temperature rise between supply and return. For water, cp is about 4.18 kJ/(kg K). Rearranged, the flow needed per kilowatt is 1 / (cp * dT) kg/s.
| Supply-to-return rise | kg/s per kW | Litres per minute per kW |
|---|---|---|
| 5 K | 0.048 | 2.9 |
| 8 K | 0.030 | 1.8 |
| 10 K | 0.024 | 1.44 |
| 15 K | 0.016 | 0.96 |
Worked example: a rack draws 120 kW of IT power and the platform captures about 90% into liquid. That is 108 kW in the coolant. At a 10 K rise the loop needs 108 x 1.44, about 155 L/min, for that one rack, and the remaining 12 kW goes into the room air. If the vendor only achieves 80% capture, liquid flow drops to about 138 L/min but room air must absorb 24 kW, double the plan. Run the same arithmetic for a 40 kW rack and you need roughly 46 to 52 L/min. Water-glycol mixtures carry less heat per kilogram than pure water, so redo the table with the coolant your vendor specifies rather than assuming water.
The equation also works backwards, as a diagnostic. Measure flow, supply and return temperature, compute the heat in the liquid, and divide by the electrical power at the rack PDUs:
CP_WATER = 4.18 # kJ/(kg K), close enough for 20-45 C water
RHO = 0.997 # kg/L
def heat_to_liquid_kw(flow_lpm, t_supply, t_return, cp=CP_WATER, rho=RHO):
kg_s = flow_lpm / 60 * rho
return kg_s * cp * (t_return - t_supply)
def capture_ratio(rack_it_kw, flow_lpm, t_s, t_r):
"""Fraction of rack electrical power leaving in the liquid. The rest is in the room air."""
return heat_to_liquid_kw(flow_lpm, t_s, t_r) / rack_it_kw
# Rack draws 118 kW at the PDUs; the CDU reports 150 L/min, 30.0 C supply, 40.2 C return.
print(round(capture_ratio(118, 150, 30.0, 40.2), 2)) # about 0.90If that ratio drifts down over weeks while power is steady, heat is going somewhere else: a partially blocked cold plate, a bypass in the manifold or a failing sensor. It is one of the cheapest health checks a liquid-cooled site can run.
Flow distribution: why one node runs hot
The rack's total flow is not what cools a GPU; the flow through its cold plate is. Paths are in parallel, so flow divides according to each path's hydraulic resistance, and pressure drop rises roughly with the square of flow. A node at the far end of a long manifold, a quick-disconnect that did not fully seat, a kinked hose or a cold plate with fouled micro-channels each gets less than its share, and the CDU's total flow reading will not show it because the other paths pick up the difference.
That is why the per-node symptom usually appears in GPU telemetry first: one server whose GPUs run several degrees hotter than identical neighbours at the same power. Because synchronous training runs at the pace of its slowest rank, a single starved node can slow an entire job, the effect described in the thermal management article. Compare each GPU's temperature with the median for its rack at similar power; an outlier at fixed power is a flow problem until proven otherwise.
The time budget when flow stops
Air-cooled servers have a lot of thermal inertia in heatsinks and room air, so a cooling fault plays out over minutes. A cold plate is small. Estimate it: about 60 mL of water in the plate and local tubing (roughly 250 J/K) plus about 0.6 kg of copper (roughly 230 J/K) gives a heat capacity near 480 J/K. Put 1,000 W into that with no flow and it warms at about 2 K per second, so a 25 K margin is gone in around 12 seconds. These are order-of-magnitude figures, not a specification for any product, but the conclusion holds across designs: after a pump or flow failure the GPU reaches its slowdown threshold in seconds.
Three consequences follow. First, the primary protection must be hardware and firmware: redundant pumps in the CDU, GPU thermal slowdown and shutdown, and node power capping. Software is too slow to save the silicon. Second, software's job is to save the work and contain the blast radius: request checkpoints, drain affected nodes so the scheduler stops placing jobs on them, and page facilities with the rack identity. Third, a CDU that serves several racks means a single pump fault can throttle hundreds of GPUs simultaneously, which shows up in a training job as a sudden step change in step time across many ranks, not as one slow node.
Telemetry: joining two worlds
GPU-side metrics come from the node: per-GPU core and memory temperature, power draw, clocks and the clock-limit reasons that tell you whether the GPU is throttling for thermal reasons. DCGM exposes all of these as fields for Prometheus, as covered in our DCGM guide. Field names shift between releases (nvidia-smi's throttle-reason fields were renamed to clock-event reasons in newer drivers), so check nvidia-smi --help-query-gpu and your DCGM version rather than copying names from old dashboards.
CDU-side metrics come from the cooling equipment: supply and return temperature, flow, differential pressure, pump speed and state, reservoir level, filter differential pressure, leak sensors and, where fitted, room dew point. Access varies by vendor. Many units speak Modbus or SNMP into the building management system; recent Redfish schema releases define resources for coolant distribution units and loops, but support is uneven, so confirm what your hardware actually implements.
The hard part is the join. A CDU alarm names a CDU, the scheduler knows hostnames, and the link between them lives in a spreadsheet in the facilities office. Put the mapping (CDU to rack to manifold port to node) into the asset database, version it, and use it in both alerting and automation. Without it, a flow alarm becomes a phone call and a hunt.
From alarm to action: a watcher
The sketch below polls CDU points, applies thresholds derived from commissioning (not from the datasheet), requests checkpoints for jobs on the affected nodes, and drains those nodes in Slurm with a reason that names the cause. It is deliberately simple; production versions add hysteresis so a sensor flapping at the threshold does not drain and resume nodes repeatedly, and a dead-man check so that losing the CDU feed is itself an alarm.
import subprocess, time
FLOW_MIN_LPM = 0.9 * DESIGN_FLOW_LPM # per rack, from commissioning
SUPPLY_MAX_C = DESIGN_SUPPLY_C + 3.0 # warmer than design: capacity is shrinking
DEWPOINT_GAP_C = 2.0 # supply must stay above room dew point
def nodes_on(rack): # from the asset database, not from guesses
return RACK_MAP[rack]
def drain(nodes, reason):
for n in nodes:
subprocess.run(["scontrol", "update", f"nodename={n}",
"state=drain", f"reason={reason}"], check=True)
def checkpoint_jobs_on(nodes):
for job in jobs_on(nodes): # site-specific: signal the job to save now
request_checkpoint(job)
while True:
for rack, cdu in read_cdu_points().items(): # Redfish, Modbus or BMS, per vendor
reasons = []
if cdu.leak_detected: reasons.append("leak")
if cdu.flow_lpm < FLOW_MIN_LPM: reasons.append(f"flow {cdu.flow_lpm:.0f} L/min")
if cdu.supply_c > SUPPLY_MAX_C: reasons.append(f"supply {cdu.supply_c:.1f} C")
if cdu.supply_c - cdu.room_dewpoint_c < DEWPOINT_GAP_C:
reasons.append("condensation risk")
if reasons:
nodes = nodes_on(rack)
checkpoint_jobs_on(nodes) # save work while hardware protection holds
drain(nodes, "cooling: " + ", ".join(reasons))
page_facilities(rack, reasons)
time.sleep(5)Drain, rather than down, lets running work finish or checkpoint, which is usually right for supply-temperature drift. A confirmed leak is different: facilities policy may require cutting power to the rack, and your automation should not fight that. Agree on the runbook with the facilities team before writing the code.
Warm water and what it does to the GPU
Liquid cooling's efficiency case rests on warm water: if the loop can run with supply temperatures in the thirties Celsius, facility heat can be rejected with dry coolers or towers instead of compressor-driven chillers for most of the year. ASHRAE's liquid-cooling guidance defines facility water classes by maximum supply temperature for exactly this reason, and vendors state which classes their platforms support.
The GPU sees the cost. Die temperature is coolant temperature plus the rise across the cold plate and the thermal interface, so warmer supply means a hotter die at the same power. A hotter die leaks more current, burning a little more power for the same work, and has less margin before clocks are limited. For training throughput, the question is whether the hottest GPU in a job stays below its slowdown point at full power on the warmest day. Measure it: run a sustained burn-in at the design supply temperature and record clocks and limit reasons, then decide the supply set point from data.
Commissioning and burn-in
Treat a new liquid-cooled rack like a new software release with acceptance tests. Pressure-test and leak-check the secondary loop before power-on. Verify flow per node, not just per rack, using the vendor's tooling or temperature-rise checks at load. Run a full-power burn-in on every GPU and compare temperatures with the rack median; outliers point to flow or mounting problems. Record capture ratio at full load, because it becomes your baseline for drift detection. Fail a CDU pump deliberately during commissioning and confirm the standby pump takes over without GPU slowdown; then trip a leak sensor and confirm the alert reaches the right people and the watcher drains the right nodes.
Keep the results. When a GPU in that rack starts running hot two years later, a commissioning baseline turns a vague complaint into a precise diagnosis. Power-side planning for the same racks is covered in GPU datacenter power.
Failure modes and trade-offs
| Failure | What software sees | Response |
|---|---|---|
| Pump failure, no redundancy | Many GPUs throttle at once, step time jumps across ranks | Checkpoint, drain served racks, page facilities |
| Partial blockage or unseated connector | One node consistently hotter at equal power | Drain node, inspect plate and connector |
| Supply temperature drift | Gradual clock reduction on warm days | Check facility side, adjust set point or power caps |
| Leak | Leak sensor alarm, possibly node power loss | Follow facilities isolation runbook |
| Supply below dew point | Condensation alarm, if fitted | Raise supply temperature immediately |
| Telemetry feed loss | Silence that looks like health | Alert on missing data, not just bad data |
The main trade-offs: in-rack CDUs shrink blast radius but add pumps to maintain per rack; warmer water saves facility energy but trims GPU headroom; tight automation protects jobs but risks false drains if thresholds are not tuned with hysteresis.
What to do next
- Get your platform's coolant type, design flow, supply temperature range and expected capture ratio from the vendor in writing.
- Compute required flow per rack with the heat-balance table and check it against CDU capacity with one pump out.
- Put the CDU to rack to node mapping in the asset database and use it in alerting.
- Ingest CDU points next to DCGM metrics, and alert on missing data as well as bad values.
- Add a per-GPU temperature-versus-rack-median outlier alert at matched power.
- Deploy a watcher that checkpoints and drains affected nodes, with hysteresis, after agreeing the runbook with facilities.
- Run pump-failover and leak-sensor drills at commissioning and keep the baselines.