A rear-door heat exchanger (RDHx) is a water coil built into a rack's back door. Hot exhaust from the servers passes through finned tubes carrying cool water, gives up most of its heat, and leaves the rack at roughly the temperature it came in. The room sees a rack that seems to produce almost no heat, while the heat leaves in water pipes. That lets an air-cooled hall host GPU racks that would overwhelm its room air handlers, and it is a common bridge for retrofit halls and hybrid liquid racks.
The wider ladder of cooling options is in our GPU datacenter cooling overview, and the coolant loop behind the door is in liquid cooling for GPU datacenters. This page covers the door itself: the energy balance that decides its capacity, a worked 40 kW sizing, the telemetry worth wiring up, and the failures that surface later as throttled GPUs and slow training steps.
What a rear door does
Every RDHx is an air-to-liquid heat exchanger mounted where the rack's exhaust must pass through it. A passive door has no fans: the servers' own fans push exhaust through the coil and pay its pressure drop. It is simple and silent, and vendors generally rate passive doors at around 20 kW per rack or below.
An active door adds variable-speed fans that pull air through a deeper coil. Vendor ratings have climbed with AI demand: 60 to 80 kW per rack is a common catalogue figure and some vendors advertise 120 kW or more. Every such number is a rating at a specific water temperature, flow and air condition. The same door can deliver half its headline figure with warmer water or lower flow.
Both are usually designed room neutral: leaving air matches the cold-aisle supply, so the room's cooling has nothing to do for that rack. Containment matters less, but the hall now depends on water reaching every door, all the time.
The energy balance and a worked example
Two equations decide whether a door can do its job, and they are the same equation applied to each side of the coil. Heat carried by a fluid stream is Q = m_dot * cp * dT: mass flow times specific heat times the temperature change. On the air side, use density about 1.2 kg/m³ and cp about 1005 J/(kg K). On the water side, cp is about 4186 J/(kg K), and one litre per minute is about 1/60 kg/s.
The servers set the airflow and the front-to-back temperature rise (server delta-T). You choose the water supply temperature and flow. Effectiveness connects the two: eps = (T_air_in - T_air_out) / (T_air_in - T_water_in), the fraction of the largest possible air temperature drop the coil achieves. Read it from the vendor's curve at your airflow and water flow, and check that your design needs no more than the door can deliver.
Worked example. A 40 kW rack with 25 C cold-aisle air and a 15 K server delta-T exhausts at 40 C. The air flow is 40,000 / (1.2 × 1005 × 15) = 2.21 m³/s, about 4,690 CFM. To be room neutral, the door must cool 40 C air back to 25 C. With 18 C supply water, the required effectiveness is (40 - 25) / (40 - 18) = 0.68. On the water side, the air's heat capacity rate is 1.2 × 2.21 × 1005 = 2,667 W/K. At 95.5 L/min of water the water's rate is 6,663 W/K, so the water warms by 40,000 / 6,663 = 6.0 K and returns at 24 C. Cut the flow to 60 L/min and the water warms by 9.6 K. The coil's average temperature rises, the achievable effectiveness falls, and the leaving air will no longer reach 25 C unless the door has margin.
Two consequences follow. Door capacity depends on supply water temperature, so a warm-water plant built for cold plates may be too warm for a room-neutral door. And servers with hot exhausts are easier for a door than servers moving lots of lukewarm air, because the larger air-to-water difference does more of the work.
A sizing check in code
The sizing check is short enough to keep next to your capacity planning. It takes the rack load, temperatures and a vendor-supplied effectiveness, and reports whether the door holds the room neutral and what the water does. The dew point uses the Magnus approximation, accurate to a few tenths of a degree in the data hall range.
import math
RHO_AIR, CP_AIR, CP_WATER = 1.2, 1005.0, 4186.0 # kg/m^3, J/(kg K), J/(kg K)
def dew_point_c(t_c, rh_pct):
"""Magnus approximation; fine for 0-50 C."""
a, b = 17.62, 243.12
g = math.log(rh_pct / 100.0) + a * t_c / (b + t_c)
return b * g / (a - g)
def check_door(q_kw, t_inlet, server_dt, t_water_in, water_lpm,
door_eps, room_t, room_rh, dew_margin=2.0):
q = q_kw * 1000.0
t_exhaust = t_inlet + server_dt
air_flow = q / (RHO_AIR * CP_AIR * server_dt) # m^3/s
c_air = RHO_AIR * air_flow * CP_AIR # W/K
c_water = water_lpm / 60.0 * CP_WATER # W/K
# Air is the smaller stream; the door removes eps of the maximum possible drop.
q_door = door_eps * min(c_air, c_water) * (t_exhaust - t_water_in)
t_leave = t_exhaust - q_door / c_air
report = {
"air_m3s": round(air_flow, 2),
"cfm": round(air_flow * 2118.88),
"door_kw": round(q_door / 1000.0, 1),
"leaving_air_c": round(t_leave, 1),
"water_rise_k": round(q_door / c_water, 1),
"residual_to_room_kw": round(max(q - q_door, 0.0) / 1000.0, 1),
}
dp = dew_point_c(room_t, room_rh)
if t_water_in < dp + dew_margin:
raise ValueError(f"supply {t_water_in} C is within {dew_margin} K of dew point {dp:.1f} C")
return report
# 40 kW rack, 25 C inlet, 15 K server rise, 18 C water at 95.5 L/min,
# vendor curve says eps = 0.70 at this airflow, room 25 C / 50% RH.
print(check_door(40, 25, 15, 18, 95.5, 0.70, 25, 50))
# {'air_m3s': 2.21, 'cfm': 4685, 'door_kw': 41.1, 'leaving_air_c': 24.6,
# 'water_rise_k': 6.2, 'residual_to_room_kw': 0.0}With eps = 0.70 the door has a little margin over the 0.68 the worked example needed, so it removes 41.1 kW rather than exactly 40 and the water rises 6.2 K instead of 6.0. The effectiveness value is the one number you must not guess. Get the vendor's curve for the door at your airflow and water flow, and if the vendor only quotes a headline kW figure, ask for the test conditions behind it. A door rated at 75 kW with 14 C water and very high exhaust temperatures can be a 40 kW door in a hall that supplies 20 C water.
Rear doors next to cold plates
A fast-growing use of rear doors in GPU halls is as secondary cooling. Cold plates take the GPUs and CPUs, but memory, power stages, NICs and power supplies still shed heat into air. Get the liquid-to-air split from the server vendor. As an illustration, a 120 kW rack with 85 percent captured by cold plates still puts 18 kW into the room, a whole air-cooled rack's worth.
A door absorbs that residual at the rack instead of oversizing the room air handlers, often from the same facility water. The direct-to-chip cooling article covers the cold-plate half. The cost is two independent cooling paths: a cold-plate fault heats GPUs quickly, a door fault heats the room and neighbouring inlets more slowly, and operators need to know which one they are looking at.
Controls and telemetry
A door controller typically modulates a water valve to hold leaving-air temperature at a setpoint, and on active doors sets fan speed. It exposes temperatures, valve position, flow, fan speeds and alarms over Modbus, BACnet or SNMP to the DCIM. Check the vendor's points list rather than assuming a field exists.
Join that to GPU telemetry by rack position to answer the real question: did this door fault change what the accelerators did? Poll temperature, SM clock and power:
# per-GPU sample every 10 s; ship to the same time-series store as the door points
nvidia-smi --query-gpu=index,serial,temperature.gpu,clocks.sm,power.draw \
--format=csv,noheader,nounits -l 10Alert on combinations. A door's leaving air up 3 K is a facilities ticket; the same rise plus falling SM clocks nearby is a training incident. A practical rule set:
- Door leaving air above setpoint + 2 K for 5 minutes: page facilities; tag the rack and its neighbours in the scheduler as at risk.
- Valve at 100 percent and leaving air still rising: the door is out of capacity. The cause is water temperature, flow or load, not the controller.
- Water delta-T collapsing toward zero with load unchanged: little heat is reaching the coil, so look for bypass air, an open door or a failed fan.
- GPU inlet temperatures rising on racks next to a faulted door: the room is absorbing that rack's heat. Drain or derate before the neighbours throttle.
Failure modes
The failure modes of a rear door are mostly failures of the assumption that all the rack's air passes through a cold coil.
- Door opened for service. The full 40 C exhaust enters a hall designed around room-neutral racks. Keep rear access short and derate the rack's jobs first.
- Water loss or valve stuck shut. The door becomes an obstacle with no cooling; a passive door still costs the server fans pressure while doing nothing.
- Supply water below dew point. The coil condenses moisture. Keep supply water a couple of kelvin above the room dew point: about 13.9 C at 25 C and 50 percent RH, but about 18.6 C at 27 C and 60 percent, which rules out 18 C water.
- Air bypass. Cable cut-outs, missing blanking panels or a poor seal let exhaust skip the coil. The tell is a small water delta-T at full load.
- Fan failure or fan fights. Losing fans on an active door cuts capacity; badly tuned door fans and server fans can oscillate. Commission them together.
- Leaks. The coil sits outside the servers, a lower-risk place than a cold plate, but a failed fitting still sprays cabling. Fit leak detection behind the rack.
How a door fault reaches a training job
Why a facilities component belongs on an ML infrastructure site: synchronous data-parallel training runs at the speed of its slowest rank. A GPU that reaches its thermal limit lowers its clocks to protect itself, as described in GPU thermal management. In a job spread over hundreds of GPUs, one throttled node makes every all-reduce wait for it. A door fault that warms three racks' inlet air by a few degrees can turn into a several percent step-time regression across an entire job, with nothing in the training logs explaining why.
The defence is mechanical. Record rack position with every node in your inventory. Store per-step time per rank. When step time degrades, compare the slowest ranks' rack positions against door alarms in the same time window. Many teams find their first door problem this way, from the ML side, before the facilities alarm was even triaged.
Trade-offs
| Option | Strengths | Weaknesses | Fits when |
|---|---|---|---|
| Passive RDHx | No fans, silent, simple | Server fans pay the pressure drop; lower ratings | Moderate-density racks in a retrofit hall |
| Active RDHx | Much higher ratings, own airflow control | Fans to maintain and tune; capacity tied to water temperature | Dense air-cooled GPU racks, hybrid residual heat |
| Hot or cold aisle containment | No water at the rack | Room air system must carry all heat | Lower densities, new halls with large air handlers |
| Cold plates + room air | Captures most heat at the chip | Residual air load still large at high rack power | Liquid-ready halls with room air headroom |
| Cold plates + RDHx | Nearly room neutral at very high density | Two cooling paths to monitor and maintain | Dense liquid-cooled racks in mixed halls |
Rear doors also interact with facility efficiency. Moving heat to water at the rack lets the plant run warmer and the room fans slower, which tends to lower PUE, but colder door water costs chiller energy. Model both with your own climate and plant before assuming a saving.
What to do next
- Collect each rack's measured power, server delta-T and inlet temperature, not nameplate figures.
- Get vendor effectiveness curves for the doors you are considering, at your water temperature and flow.
- Run the sizing check above per rack and record leaving-air temperature and residual heat to the room.
- Compute the worst-case room dew point and set a supply-water floor above it.
- Wire door telemetry into the same store as GPU telemetry, tagged by rack position.
- Add combined alerts (door fault plus GPU clock drop) and teach the scheduler to derate or drain at-risk racks.
- Seal every rack (blanking panels, grommets, door seal) and verify water flow at each door, not just at the CDU.
- Commission with a full-power burn-in job, then re-test after any change in water temperature setpoint; clean coil fins on a schedule.