Direct liquid cooling, often called direct-to-chip, puts a metal cold plate with coolant flowing through it on top of each hot package: GPUs, CPUs and sometimes switch chips and voltage regulators. It is now the default for dense accelerator racks. Rack-scale systems such as NVIDIA's GB200 NVL72 are specified around 130 kW per rack (HPE's collateral gives 132 kW nominal), which fans and room air cannot move through a single rack footprint.

This page is about the hardware at the chip: how heat gets from the die into the coolant, what is inside a cold plate, how flow and pressure drop through a tray, and why a liquid-cooled node still has fans. It ends with what software can measure and a way to catch a failing plate before it slows a training job. The facility side, loop sizing and CDU alarms are in Liquid Cooling for GPU Datacenters, and the wider choice between air, rear-door, cold plate and immersion is in GPU Datacenter Cooling Overview.

Advertisement

Why direct to the chip

Air cooling moves heat with a fluid that holds very little of it. Water holds roughly 3,500 times more heat per unit volume than air, so a few litres per minute through a small plate does the work of a large volume of air through a big heatsink. Putting the liquid within millimetres of the die also removes the biggest thermal resistance in an air-cooled system, the heatsink-to-air interface.

The cost is plumbing in every server, a coolant distribution unit (CDU) per rack or row, and new failure modes. Cold plates beat immersion on familiarity: server form factor and servicing barely change.

The heat path and its resistances

Heat path, die to coolantGPU die (junction)heat source, ~1 kW classTIM1 + lid or bare dieTIM2 (pad or paste)Cold plate base (copper)spreads heatMicrochannel finsconvection to coolantCoolant (PG25 or water)carries heat awayT_junction = T_inlet + P x (R_conv + R_base + R_TIM + R_die) + plate riseOne tray on a rack manifoldSupply manifoldcool, from CDUReturn manifoldwarm, to CDUQD + hose inPlate: CPUPlate: GPU 0Plate: GPU 1QD + hose outseriesPlates in series see warmer coolant downstream; parallel splits need balanced flow.Residual air load: DIMMs, NICs, drives, retimers and power supplies that have no plate still need fans.Illustrative only: real trays differ in plate count, order and whether plates run in series or parallel.
Left: the stack of thermal resistances between die and coolant. Right: one illustrative tray, with the CPU and GPU plates fed in series from a rack manifold through quick disconnects.

Heat crosses a chain of layers, and each layer behaves like a resistor: the temperature drop across it equals the power through it multiplied by its thermal resistance, in kelvin per watt. The junction temperature, which the GPU's clock controller watches, is the coolant inlet temperature plus the sum of those drops, plus some of the coolant's own temperature rise along the plate.

The useful insight is that you control only some terms. Die and package resistance are fixed by the vendor. Thermal interface material is set at assembly and degrades over time. Convection resistance falls as flow rises, with diminishing returns. Inlet temperature is a facility choice. The calculation below makes the budget concrete with placeholder values.

# Illustrative junction-temperature estimate for one cold-plated GPU.
# Every resistance here is a placeholder; take real values from vendor thermal specs.
P_watts      = 1000.0     # sustained board power reaching the plate
T_inlet_c    = 30.0       # coolant temperature arriving at this plate
R_conv       = 0.012      # K/W, fins to coolant (falls as flow rises)
R_base       = 0.004      # K/W, copper base spreading
R_tim        = 0.010      # K/W, thermal interface material(s)
R_die        = 0.008      # K/W, die and package
cp, rho      = 3900.0, 1020.0      # PG25-like coolant, J/(kg K) and kg/m^3
flow_lpm     = 1.5                 # litres per minute through this plate

m_dot   = flow_lpm / 60 / 1000 * rho          # kg/s
dT_cool = P_watts / (m_dot * cp)              # coolant rise across the plate
T_j     = T_inlet_c + P_watts * (R_conv + R_base + R_tim + R_die) + dT_cool / 2

print(f"coolant rise {dT_cool:.1f} K, junction estimate {T_j:.1f} C")
# coolant rise 10.1 K, junction estimate 69.0 C

With these illustrative numbers, a kilowatt at 1.5 litres per minute warms the coolant by about 10 K, and the junction sits roughly 39 K above the inlet. Raise the inlet by 5 K and the junction rises by 5 K. Double the flow and only the convection term and half the coolant rise shrink, so the junction falls a few kelvin, at a large cost in pumping power and pressure drop.

Advertisement

Inside a cold plate

A cold plate is a copper base machined or brazed to an array of fine fins or microchannels, under a cover with an inlet and outlet. Copper is used because it conducts heat well and spreads the hotspot under the die across the fin area. The fins, often a fraction of a millimetre wide, multiply the wetted surface area. Some designs add jet impingement, where coolant is sprayed straight at the base above the hottest part of the die and then flows outward through the fins.

Every design trades heat transfer against pressure drop. Narrower channels and more fins move more heat for a given flow but need more pump pressure and clog more easily. The plate is also matched to the package: a GPU with stacked HBM around the compute die has several heat sources of different heights, and the plate and interface material must contact all of them evenly. Mounting pressure matters for the same reason; uneven torque leaves a thick interface layer on one corner and a hotspot underneath it.

Flow and pressure drop

The flow a plate needs follows from the heat balance Q = m_dot x c_p x delta T, derived in the companion loop article. For a water-glycol coolant such as PG25, a common rule of thumb is about 1.5 litres per minute per kilowatt for a 10 K rise. A tray with eight plates at a kilowatt each therefore needs on the order of 12 litres per minute, and a rack needs hundreds.

Pressure drop rises with flow, somewhere between linearly in laminar microchannels and with the square of flow in turbulent sections and fittings. The CDU's pump must supply the whole rack's flow at the pressure drop of the worst path, so a single restrictive plate or kinked hose forces either more pump energy or less flow everywhere.

Within a tray, plates may be plumbed in series or in parallel. In series, the downstream plate sees coolant already warmed by the upstream one; a GPU after a CPU might run several kelvin hotter at the same power. In parallel, every plate sees the inlet temperature, but the flow divides by resistance, so a slightly more restrictive plate gets less flow. Neither is wrong, but both mean that identical GPUs in one node do not run at identical temperatures, which matters when you set alert thresholds.

Manifolds, hoses and quick disconnects

Each rack has a vertical supply and return manifold. Trays connect through hoses and quick disconnect couplings (QDs), either hand-mated at the back of the rack or blind-mated as the tray slides in. Dripless QDs seal both halves when separated, so a tray can be removed without draining the loop, though a few drops per cycle is normal and a damaged seal is a classic leak source.

Air is the hidden enemy after installation and every service event. A trapped bubble blocks flow to part of the fin area, and the GPU underneath runs hot with no visible fault. Commissioning purges air and verifies flow per tray; the monitoring later on this page catches what it misses. Leak detection uses sensing cable or spot sensors in trays and at the rack base, wired to the CDU or BMC so a leak can trigger an isolation valve or a node shutdown.

The residual air load

Cold plates cover the hottest packages, not everything. Memory DIMMs, network cards, retimers, storage, voltage regulators without plates and power supplies still dissipate heat into the air. In a hybrid design, an illustrative split might put 80 percent of rack heat into liquid and 20 percent into air. At 130 kW, that 20 percent is 26 kW, more than an entire air-cooled rack used to dissipate a decade ago.

So liquid-cooled racks still need fans and room cooling sized for the residual. Some designs extend plates to memory and power stages, or add a rear-door heat exchanger. Ask vendors for the liquid and air split at your planned inlet temperature.

Coolant and wetted materials

The secondary loop that runs through the plates is usually a propylene glycol and water mixture with corrosion inhibitors, PG25 being common, or treated water. Glycol lowers freezing risk and inhibits biological growth but has lower heat capacity and higher viscosity than water. Whatever you choose must match the vendor's wetted-materials list. Mixing copper, aluminium and certain alloys in one loop invites galvanic corrosion, and corrosion products and debris clog microchannels.

Treat coolant as a consumable with a maintenance schedule: fine filtration at the CDU to the vendor's specification, periodic sampling for pH, inhibitor level, conductivity and biological growth, and documented top-ups. A slowly clogging loop appears in software first, as rising GPU temperatures at constant power across many nodes.

Single-phase versus two-phase

Everything above describes single-phase cooling: the coolant stays liquid and carries heat as a temperature rise. Two-phase direct-to-chip uses a dielectric fluid that boils inside the plate. Boiling absorbs latent heat at a nearly constant temperature, so it needs far less flow and holds the whole plate at a uniform temperature. A leak of dielectric fluid also does not short electronics.

The trade-offs are maturity, fluid cost and handling, pressure management, and regulatory pressure on some fluorinated fluids. Single-phase plates are the mainstream choice today; evaluate two-phase with vendor data.

Facility temperature and the GPU

The CDU separates the technology cooling system loop through the plates from the facility water loop. ASHRAE's liquid cooling classes, W17, W27, W32, W40, W45 and W+, describe that facility supply temperature, not the secondary loop. Warmer facility water enables more free cooling and heat reuse. The cost lands on the chip: inlet temperature passes straight into junction temperature, and hotter silicon leaks more power, which leaves less of the power limit for clocks. How the clock controller responds is covered in GPU Thermal Management.

What software can see

From the host, you can read each GPU's temperature, power draw, clocks and clock-event reasons through nvidia-smi or DCGM. The coolant inlet temperature and flow come from the CDU or the node's BMC, often over Redfish, and have to be joined in your monitoring system. Together they give a powerful derived metric: effective thermal resistance, the GPU's temperature above inlet divided by its power.

import subprocess, statistics

def gpu_samples():
    out = subprocess.run(
        ["nvidia-smi", "--query-gpu=index,temperature.gpu,power.draw",
         "--format=csv,noheader,nounits"],
        capture_output=True, text=True, check=True).stdout
    for line in out.strip().splitlines():
        idx, temp, power = (x.strip() for x in line.split(","))
        yield int(idx), float(temp), float(power)

def effective_resistance(t_inlet_c: float, min_power_w: float = 300.0) -> dict[int, float]:
    """(T_gpu - T_inlet) / P per GPU; only meaningful under sustained load."""
    return {i: (t - t_inlet_c) / p for i, t, p in gpu_samples() if p >= min_power_w}

def outliers(r: dict[int, float], tolerance: float = 0.20) -> list[int]:
    median = statistics.median(r.values())
    return [i for i, v in r.items() if v > median * (1 + tolerance)]

# t_inlet_c comes from the CDU or the node BMC (often via Redfish), not from nvidia-smi.

Under sustained load, effective resistance should be steady for a given GPU and similar across GPUs in the same plate position. A value that creeps up over weeks points at interface degradation or clogging. A step change after maintenance points at trapped air or a badly seated plate. A whole rack shifting together points at the loop or the CDU rather than any one plate.

Worked example: the slow rank

An illustrative case: a training job on 64 eight-GPU nodes had steps about 4 percent slower than its sister job. Per-rank timing showed one rank consistently last into every all-reduce. That GPU sat at the same power as its neighbours but ran 9 K hotter, and its SM clock was lower with a thermal reason set. Its effective resistance was 0.047 K/W against a tray median of 0.036. The node had a GPU board replaced the previous week. The plate was reseated with fresh interface material and the loop purged; resistance fell to 0.035, and the job's step time matched its sister. A synchronous job runs at the pace of its slowest rank, so one plate mattered to all 512 GPUs.

Failure modes and trade-offs

FailureSymptom in softwareResponse
Trapped air after serviceone GPU's resistance steps uppurge, reseat, recheck under load
Interface degradationslow drift over weeks or monthsschedule rework; compare against fleet baseline
Clogged channels or filtermany GPUs drift togethercheck filter pressure drop and coolant samples
Unbalanced parallel flowsame plate position always hottercheck tray design limits; adjust alert per position
Leakleak sensor alarm, possible node shutdownisolate, drain tray, inspect QDs and hoses
Warm inletwhole fleet runs hotter, lower clockstrade facility efficiency against throughput explicitly

What to do next

  1. Get the vendor's thermal specification: allowed inlet range, flow per tray, pressure drop and the liquid and air split.
  2. Join CDU or BMC inlet temperature with per-GPU temperature and power in your monitoring.
  3. Compute effective thermal resistance per GPU under load, baseline it, and alert on drift and on outliers by plate position.
  4. Add a post-maintenance burn-in that checks resistance before a node returns to the scheduler.
  5. Put coolant sampling and filter checks on a schedule with an owner.
  6. Size room air for the residual load at your planned inlet temperature, not for zero.
  7. When choosing a facility water class, model what each extra kelvin of inlet does to junction temperature and clocks.
Key takeaway: Direct liquid cooling works by shortening the heat path: a cold plate puts coolant millimetres from the die, so junction temperature becomes inlet temperature plus a short chain of thermal resistances. You control flow, inlet temperature, coolant quality and assembly; the vendor controls the rest. Liquid-cooled racks still have a large air load, plates in one tray do not run identically, and most faults show up first as one GPU running hotter at the same power. Measure effective thermal resistance per GPU, baseline it, and treat drift as a maintenance signal before it becomes a slow training job.