Immersion cooling puts the whole server, boards and all, into a bath of electrically non-conductive liquid. Instead of blowing air across heatsinks, the fluid touches every component directly and carries heat to a heat exchanger. For GPU clusters the appeal is obvious: an 8-GPU training server can draw well over 10 kW, and air struggles to remove that from a rack without enormous fan power and noise.
This page explains immersion from first principles and then turns it into decisions a GPU platform team actually makes: how much fluid flow a tank needs, what changes inside the server, what the GPU and your monitoring see, how scheduling and maintenance change, and when immersion is the wrong answer. For the wider cooling ladder start with GPU Datacenter Cooling Overview; for cold-plate systems, see Direct Liquid Cooling, in depth.
One caution up front. Vendor material on immersion is full of headline numbers for PUE, density and savings. Those depend on climate, fluid, tank design and what was used as the baseline. The numbers in this article are worked examples to show the method; replace them with your fluid datasheet and your measured server power before you size anything.
How immersion moves heat
There are two families, and they behave differently enough that you should never discuss "immersion" without saying which one.
Single-phase immersion. The fluid stays liquid. It is usually a synthetic hydrocarbon or similar oil-like dielectric. A pump circulates it through the tank, past the servers, and out to a heat exchanger, where facility water takes the heat away. Heat transfer is purely convective: the fluid warms as it passes the hot parts. Because the fluid is more viscous than air and carries far more heat per unit volume, servers can lose their fans entirely.
Two-phase immersion. The fluid has a low boiling point, typically in the range of roughly 50 to 60 degrees Celsius for the engineered fluorinated fluids used. It boils on hot surfaces; the vapour rises to a condenser coil in the tank lid, condenses and drips back. Boiling absorbs a lot of heat at almost constant temperature, so hot spots are held close to the boiling point. The cost is vapour management: the tank must be sealed during operation, fluid escapes every time a lid opens, and most of these fluids are PFAS chemicals facing tightening regulation. 3M, a major supplier, announced in 2022 that it would exit PFAS manufacturing by the end of 2025, which pushed many projects towards single-phase.
The rest of this article focuses on single-phase, because that is where most current GPU deployments are, and notes where two-phase differs.
Sizing the fluid loop
Every cooling decision reduces to one equation: heat removed equals mass flow times specific heat times temperature rise, Q = m_dot * cp * dT. Fluids differ mainly in density and specific heat, so the same heat load needs very different volumes of flow.
Worked example. A tank holds four 8-GPU servers at 10 kW each, so Q = 40 kW. Take an illustrative single-phase fluid with density 800 kg per cubic metre and specific heat 2.1 kJ per kg per kelvin, and allow the fluid to rise 10 K across the tank.
Q = 40 kW
cp = 2.1 kJ/(kg*K)
dT = 10 K
m_dot = Q / (cp * dT) = 40 / 21 = 1.90 kg/s
V_dot = m_dot / rho = 1.90 / 800 m^3/s = 2.38 L/s ~ 143 L/min
Water at the same dT (cp 4.18, rho 1000):
m_dot = 40 / 41.8 = 0.96 kg/s ~ 57 L/minSo the fluid loop moves about two and a half times the volume that a water loop would for the same load and temperature rise, and it does so with a more viscous fluid, which raises pumping power. That is the core engineering trade: you remove server fans, but you add a fluid pump that must push a thick liquid. In practice the net is still usually favourable, because a pump moving liquid is far more efficient than dozens of small fans moving air, but do the arithmetic for your fluid.
The second number to check is the temperature at the hottest component, not the average. Single-phase relies on fluid velocity across the heatsink fins. Natural convection alone is weak in a viscous fluid, so tank designers direct flow through the server chassis. A server designed for air, with fins spaced for air, often leaves the GPU noticeably hotter in fluid than a server whose heatsinks were redesigned for immersion. This is why the server matters as much as the tank.
What changes inside the server
Putting an air-cooled server in a tank is not a drop-in change. The items below are the ones that most often cause surprises.
- Fans and the BMC. Fans are removed or disabled. The baseboard management controller may then report fan failures, raise alarms or even throttle and shut down the node. You need firmware or BMC configuration from the server vendor that understands an immersion profile.
- Heatsinks. Air heatsinks have tall, closely spaced fins. In fluid, wider fin spacing and lower profiles often perform better. Immersion-ready servers ship with different heatsinks; reusing air heatsinks is a common cause of a GPU running hotter than expected.
- Thermal interface material. Some thermal pastes and pads can be degraded or washed out by the fluid over time. Use materials the fluid vendor lists as compatible.
- Cables and optics. Fluid can wick along cables by capillary action and out of the tank. Pluggable optical transceivers are a particular concern because fluid in the optical path can degrade the link. Many designs keep the optics above the fluid line or use parts qualified for immersion; check before you assume a 400G or 800G link will survive.
- Materials compatibility. Plastics, labels, adhesives and some rubbers swell, leach or dissolve in some fluids. Leached material discolours the fluid and can foul filters. The OCP immersion requirements include wetted-material compatibility guidance for this reason.
- Spinning disks and power supplies. Hard disks are generally unsuitable. Power supplies must be immersion-qualified or kept outside the fluid.
- Warranty. Confirm in writing that the GPU and server vendor will support the hardware in your chosen fluid. Do not assume.
What the GPU and your monitoring see
From the GPU's point of view nothing exotic happens. It measures its own core and memory temperatures, manages clocks against power and temperature limits, and reports why it slowed down. What changes is the shape of the data. A well-run tank tends to give very even GPU temperatures across a server, because every device sits in fluid at almost the same supply temperature. A poorly balanced tank shows a gradient: GPUs near the fluid outlet or in a low-flow region run hotter, and with them their HBM, which on recent parts is often the first limit reached.
The practical monitoring job is to join GPU telemetry with tank telemetry so you can tell "this GPU is hot because the tank is warm" from "this GPU is hot because its heatsink is bad". The script below polls each GPU, compares it with its siblings and with the tank supply temperature, and reports outliers. It uses nvidia-smi query fields; newer drivers name the throttle field clocks_event_reasons.active while older ones use clocks_throttle_reasons.active.
import subprocess, statistics
FIELDS = "index,temperature.gpu,temperature.memory,power.draw,clocks.sm,clocks_event_reasons.active"
def gpu_rows():
out = subprocess.run(
["nvidia-smi", f"--query-gpu={FIELDS}", "--format=csv,noheader,nounits"],
capture_output=True, text=True, check=True).stdout
for line in out.strip().splitlines():
idx, t, tm, pw, clk, ev = [x.strip() for x in line.split(",")]
yield dict(idx=int(idx), t=float(t),
tmem=None if tm in ("N/A", "[N/A]") else float(tm),
power=float(pw), clk=int(clk), events=int(ev, 16))
def classify(rows, tank_supply_c, spread_limit=6.0, approach_limit=35.0):
"""Return (gpu index, reason) pairs worth a human look."""
temps = [r["t"] for r in rows]
med = statistics.median(temps)
flags = []
for r in rows:
if r["t"] - med > spread_limit:
flags.append((r["idx"], f"{r['t'] - med:.1f} C above sibling median: heatsink, TIM or local flow"))
if r["t"] - tank_supply_c > approach_limit:
flags.append((r["idx"], f"{r['t'] - tank_supply_c:.1f} C above fluid supply: check tank flow"))
if r["events"] & ~0x3: # ignore idle (0x1) and app-clock setting (0x2)
flags.append((r["idx"], f"clock event mask {r['events']:#x}"))
return flags
def read_tank_supply_temp():
raise NotImplementedError("wire to your CDU / tank controller API")
if __name__ == "__main__":
rows = list(gpu_rows())
supply = read_tank_supply_temp()
for idx, why in classify(rows, supply):
print(f"gpu{idx}: {why}")The thresholds are placeholders. Calibrate them from a burn-in run: load every GPU with the same kernel, record steady-state temperatures, and set the spread limit a few degrees above the spread you saw on a healthy tank. If you run DCGM, the same data is available as fields such as DCGM_FI_DEV_GPU_TEMP, DCGM_FI_DEV_MEMORY_TEMP and DCGM_FI_DEV_POWER_USAGE, which makes the join with tank metrics a dashboard query rather than a script. For how the clock-management loop reacts to heat, see GPU Thermal Management, in depth.
Operations and the scheduler
Immersion changes the physical maintenance workflow, and that has to be reflected in the scheduler. To service a node, a technician lifts the server out of the tank, lets it drip over the tank or a drip tray, and then works on it, often wearing gloves and on an absorbent mat. That takes longer than sliding an air-cooled server out of a rack, and a lift may be needed for heavy GPU servers. The software consequence is simple: never lift a node that is running a job, and plan service windows per tank, not per node, if the tank lid or a shared pump must be touched.
A small piece of glue code ties tank alarms to the scheduler. When the tank controller reports a fault, the affected nodes are drained so no new work lands on them; a severe fault triggers checkpoint-and-stop for running jobs. With Slurm, draining is a single command:
import subprocess
TANK_NODES = {"tank-a": ["gpu-a01", "gpu-a02", "gpu-a03", "gpu-a04"]}
def drain(tank, reason):
nodes = ",".join(TANK_NODES[tank])
subprocess.run(["scontrol", "update", f"nodename={nodes}",
"state=drain", f"reason={reason}"], check=True)
def on_tank_alarm(tank, kind, value):
if kind in ("flow_low", "level_low", "supply_temp_high"):
drain(tank, f"immersion {kind}={value}")
# severe cases: signal running jobs to checkpoint (site-specific)On Kubernetes the equivalent is cordoning the nodes and, if needed, evicting pods; see Liquid Cooling for GPU Datacenters, in depth for a fuller alarm-to-action watcher and the time budget between a pump fault and a throttled GPU. Immersion has more thermal mass than a cold-plate loop, so a pump stop gives you more time, but not unlimited time: measure it during commissioning by stopping the pump under load and recording how long until GPUs begin to throttle.
Fluid itself becomes an operational asset. Plan for periodic sampling to check for water content, contamination and degradation; filter replacement; top-ups for fluid carried out on lifted servers; and spill procedures. For two-phase tanks, track fluid loss as a metric; every lid opening costs vapour.
Failure modes
| Failure | What you see | Response |
|---|---|---|
| Low flow or pump stop | All GPUs in the tank warm together; clock events appear | Drain tank nodes; switch to standby pump; checkpoint long jobs |
| Air-style heatsink in fluid | One server consistently hotter than its siblings at equal power | Replace with immersion heatsink; do not just lower the power cap |
| Fan alarms from BMC | Node marked unhealthy or shuts down with no thermal cause | Apply vendor immersion firmware / sensor profile |
| Fluid wicking along cables | Fluid on floor or in cable trays outside the tank | Reroute, add drip loops, use sealed or qualified cables |
| Optic degradation | Link errors or flaps on submerged transceivers | Keep optics above fluid line or use qualified parts |
| Material leaching | Fluid discolours, filters clog, fluid test fails | Identify incompatible part; filter or replace fluid |
| Two-phase vapour loss | Fluid level drops faster than expected | Check lid seals and service-window discipline |
Trade-offs against cold plates
Immersion is one option among several, and the choice is not purely thermal.
| Question | Direct-to-chip cold plates | Single-phase immersion |
|---|---|---|
| Heat captured by liquid | Most, not all; residual air load remains | Essentially all; no server fans |
| Server choice | Vendor liquid-cooled SKUs, broad availability | Immersion-qualified SKUs; narrower choice |
| Floor and building | Standard racks plus CDUs | Horizontal tanks, floor loading, lifts |
| Servicing a node | Similar to air, with quick disconnects | Lift, drip, clean; slower |
| Optics and cabling | Unchanged | Needs care |
| Ecosystem for the newest GPU racks | Primary path for current rack-scale systems | Check vendor support per platform |
Immersion tends to make the most sense where you control the whole stack, where density or noise rules out air, where heat reuse is valuable, or at the edge where a sealed tank protects electronics from dust and humidity. Cold plates tend to win where you need the newest rack-scale GPU systems on vendor timelines, because those are designed and supported around direct liquid cooling. Power delivery is the other constraint; see GPU Datacenter Power Requirements.
What to do next
- Write down the per-server power you actually measure under your training load, not the nameplate figure.
- Run the flow equation above with your fluid's datasheet density and specific heat, and size pump capacity with margin.
- Get written confirmation from the GPU and server vendors that your exact SKU is supported in your exact fluid.
- Ask for immersion heatsinks and BMC firmware; list every cable, optic and label that will be wetted.
- During commissioning, run a uniform burn-in, record per-GPU temperatures and set outlier thresholds from that data.
- Stop the pump under load once, deliberately, and record the seconds until GPUs throttle.
- Wire tank alarms to node drain in your scheduler and test it before production jobs land.
- Add fluid sampling, filter changes and top-ups to the maintenance calendar.