Almost every joule a GPU cluster draws ends up as heat. A 10 MW training hall is, thermodynamically, a 10 MW heater that happens to compute along the way. Most datacenters throw that heat into the air through chillers, cooling towers or dry coolers, and pay for the privilege. Waste heat reuse captures some of it and sells or gives it to someone who needs heat: a district heating network, a greenhouse, a swimming pool, an industrial dryer or the office next door.
The idea is old. What changed is that liquid-cooled AI racks return water warm enough to be useful, and regulation now requires it in places. The European Union's recast Energy Efficiency Directive, (EU) 2023/1791, requires data centres with a total rated energy input above 1 MW to use their waste heat where it is technically and economically feasible, and Germany's Energy Efficiency Act sets numeric targets. This article explains the physics, the metric, a worked energy balance, why training workloads make heat supply bursty, and what software and operations teams must do to make a heat contract work. The cooling hardware itself is covered in GPU datacenter cooling; this page starts where the heat leaves the rack.
Heat has a grade
Heat has a quantity and a grade, meaning its temperature. A district heating network typically wants supply water somewhere around 60 to 90 C, and newer low-temperature networks run lower. An air-cooled hall exhausts air around 30 to 40 C, and the water loop behind its coolers is cooler still. That is a lot of energy at a temperature almost nobody can use directly.
Liquid cooling raises the grade. Direct-to-chip cold plates remove most of a server's heat into liquid, and because silicon tolerates warm coolant, facilities can run supply water in the 30s C and get return water around 40 to 50 C, depending on the design. The exact numbers come from the server vendor's coolant specification, not from this article; check the allowed inlet temperature and flow per rack before planning anything. Immersion systems also deliver heat into a liquid loop. The parts that remain air-cooled, such as power supplies, network switches and storage, still dump low-grade heat into room air.
Higher return temperature has a cost. Warmer coolant means warmer silicon, which can mean higher leakage power and less margin before thermal throttling. A reuse design is a negotiation between the IT side, which wants cold water, and the heat customer, who wants hot water. The heat pump sits between them.
The heat path
Follow the heat in the figure. Cold plates move heat into the secondary loop, which is kept separate and clean for the servers. The coolant distribution unit (CDU) exchanges it into the facility water loop. A plate heat exchanger then separates the facility water from the customer's water, and that exchanger is usually the contractual boundary, with a heat meter (flow rate times temperature difference) on one side. Hydraulic separation matters: district water chemistry and pressure are not something you let near a server manifold.
If the customer needs more than the return temperature, a heat pump lifts it. When there is no demand, on a warm summer night for example, the heat must still go somewhere, so the original dry coolers or towers stay in place, sized for 100 percent of the load. That last point surprises people: reuse rarely removes cooling capital. It adds a second path, and the plant must fail over between them without disturbing the GPUs.
The basic relation is Q = m times c_p times delta T. Moving 6.4 MW with water (c_p about 4.18 kJ per kg per K) across a 10 K temperature difference needs about 153 kg per second, roughly 153 litres per second. Pipe sizes, pump power and the distance to the customer follow from that number, and distance often decides feasibility before anything else does.
Heat pumps and the lift
A heat pump moves heat from a cold source to a hot sink using electrical work. The ideal (Carnot) heating coefficient of performance is COP = T_hot / (T_hot - T_cold), with temperatures in kelvin. Lifting 40 C water to 75 C gives 348.15 / 35, about 9.9. Real machines achieve a fraction of Carnot; this article assumes half, which puts the example around 5. Use the manufacturer's rated COP at your actual temperatures instead of that assumption, because the fraction varies with machine, refrigerant and load.
| Source water | Sink | Carnot COP | At 50 pct of Carnot |
|---|---|---|---|
| 25 C (air-cooled loop) | 75 C | 7.0 | 3.5 |
| 35 C | 75 C | 8.7 | 4.4 |
| 40 C (warm liquid loop) | 75 C | 9.9 | 5.0 |
| 45 C | 75 C | 11.6 | 5.8 |
The table shows why liquid cooling and reuse go together: every degree of source temperature improves the heat pump. Energy conservation gives the delivered heat as Q_hot = Q_source + W, and COP = Q_hot / W. With 6.4 MW of captured heat and a COP of 5, the heat pump draws W = 6.4 / (5 - 1), about 1.6 MW, and delivers about 8.0 MW. Whether that electricity counts against the datacenter depends on who owns the heat pump, which is the first thing to settle when you compute the metric.
ERF, regulation and a worked balance
PUE, the ratio of total facility energy to IT energy, cannot reward reuse, because exported heat leaves both numbers unchanged. The metric for reuse is the Energy Reuse Factor, standardised as ISO/IEC 30134-6: ERF equals reused energy divided by total datacenter energy, a number between 0 and 1. A related metric, ERE, subtracts reused energy from the total before dividing by IT energy. See PUE for GPU datacenters for what PUE itself does and does not measure.
Germany's Energy Efficiency Act (Energieeffizienzgesetz, EnEfG), as enacted in 2023, requires data centres starting operation from 1 July 2026 to be designed for an ERF of at least 10 percent, rising to 15 percent for those starting from 1 July 2027 and 20 percent from 1 July 2028, with the annual figure due two years after operation starts. These figures come from legal summaries of the 2023 text. Rules like this get amended and carry exemptions, so check the current text before you design to them.
Here is a worked annual balance for a hall averaging 8 MW of IT load at PUE 1.2. Its total energy is 8 times 1.2 times 8,760 hours, which is 84,096 MWh. Cold plates capture 80 percent of IT heat, 6.4 MW, at 40 C. The district network can absorb that heat about half the hours of the year, mostly in winter. Reused energy is therefore 6.4 times 0.5 times 8,760, or 28,032 MWh, which gives an ERF of 0.33. That clears every EnEfG step, but demand share is the fragile assumption: halve it, and ERF drops to 0.17. The function below makes the dependencies explicit.
def heat_balance(it_mw, pue, capture, demand_share, t_src_c, t_sink_c,
carnot_fraction=0.5, hours=8760):
"""Annual reuse estimate. Heat pump owned by the customer (outside the DC boundary)."""
total_mwh = it_mw * pue * hours # all facility energy
captured_mw = it_mw * capture # heat reaching the liquid loop
reused_mwh = captured_mw * demand_share * hours # heat the customer actually took
t_hot, t_cold = t_sink_c + 273.15, t_src_c + 273.15
cop = carnot_fraction * t_hot / (t_hot - t_cold)
hp_mw = captured_mw / (cop - 1) # heat pump electrical input
return {
"ERF": reused_mwh / total_mwh,
"COP": cop,
"heat_pump_MW": hp_mw,
"delivered_MW": captured_mw + hp_mw,
}
print(heat_balance(it_mw=8.0, pue=1.2, capture=0.8, demand_share=0.5,
t_src_c=40, t_sink_c=75))
# {'ERF': 0.333..., 'COP': 4.97..., 'heat_pump_MW': 1.61..., 'delivered_MW': 8.01...}
Training load is bursty, and heat follows it
Here GPU software becomes part of the heating system. Synchronous training does not draw flat power. Each step alternates compute-heavy phases with communication phases, and when collectives stall the GPUs, power drops. Large clusters can swing a noticeable share of their power in step-length periods. Checkpoint writes, evaluation pauses, data loader stalls and job failures cause longer drops. A hang that waits out a collective timeout can leave thousands of GPUs idle for minutes, and a cluster-wide restart takes longer still.
For the heat customer, a 6.4 MW source that dips to 2 MW for ten minutes is a supply failure. Thermal mass absorbs short swings: water in the loops, the heat pump's own inertia, and a buffer tank. Size the tank from the energy gap you need to ride through. Covering the full 6.4 MW for 600 seconds is 3.84 GJ. At a usable temperature swing of 10 K, that is 3.84e9 / (4,180 times 10), about 92 tonnes of water, or 92 cubic metres. That is large but ordinary for district heating. Longer outages need the customer's backup boilers, and the contract should say so.
Treat GPU power telemetry as a heat forecast. Per-GPU power is exposed by nvidia-smi --query-gpu=power.draw --format=csv and by DCGM's power usage field, and rack power distribution units give the same signal at coarser resolution. Multiply by the capture fraction and you have the heat entering the liquid loop, seconds before the return temperature shows it. Combined with the scheduler's queue (planned jobs, maintenance windows, drained nodes), it becomes a forecast you can send to the heat customer. The training-side view of power is covered in GPU datacenter power.
# Heat entering the liquid loop, from GPU telemetry (sketch; sample every 10 s)
import subprocess
CAPTURE = 0.8 # fraction of IT power reaching the liquid loop (measure, don't assume)
OVERHEAD = 1.35 # node power / GPU power for this server type, from PDU data
def node_heat_kw():
out = subprocess.run(
["nvidia-smi", "--query-gpu=power.draw", "--format=csv,noheader,nounits"],
capture_output=True, text=True, check=True).stdout
gpu_w = sum(float(x) for x in out.split() if x.replace(".", "", 1).isdigit())
return gpu_w * OVERHEAD * CAPTURE / 1000.0
# Aggregate across nodes, smooth with a 5-minute EWMA, and publish both the live value
# and a 24-hour forecast built from the scheduler's queue to the plant controller.
Operating a heat contract
A heat contract turns a cooling system into a supply obligation, and that changes operations.
- Priority is fixed: GPUs first. The control system must open the dry cooler bypass whenever the customer cannot take heat, automatically and without operators. Return temperature limits for the servers override every heat setpoint.
- Agree on the metering point. Meter at the plate heat exchanger, calibrate the meters, and define in the contract who owns the heat pump and its electricity. The ERF figure depends on that definition.
- Publish maintenance and large job changes. A planned cluster drain for a driver upgrade is a planned heat outage. Give the customer the same notice you give your users.
- Watch approach temperatures. Fouling in exchangers raises the gap between loops, lowers the useful source temperature, and quietly drops heat pump COP. Trend it monthly.
- Plan for seasons. Winter demand can exceed supply; summer demand collapses. Annual ERF comes from shoulder-season and winter hours, so model them explicitly rather than using an annual average.
Failure modes
Failure modes worth naming before they happen:
- Heat customer stops taking heat (network fault, pump trip). Without a fast bypass, facility return temperature rises and the GPUs throttle. Test this failover under load, not at commissioning only.
- Return temperature setpoint creep. Someone raises the facility temperature to improve COP, silicon runs hotter, and training throughput drops slightly on every job. Track GPU clocks and throttle reasons alongside water temperatures.
- Cross-contamination. A leaking exchanger mixes district water into the facility loop. Keep pressure differentials and water quality monitoring on both sides.
- Optimistic ERF. Projections that assume full IT load and year-round demand will miss. Use measured utilisation and the customer's real load profile.
- Stranded pipes. AI campuses expand or move faster than district networks are built. A 20-year heat contract with a facility whose load may shift is a commercial risk, not an engineering one, but it kills projects.
Trade-offs
Reuse is not always the right answer. Siting a hall near heat demand can conflict with siting it near cheap clean power and fibre. Air-cooled halls produce heat so low-grade that heat pump electricity can cancel much of the benefit. Rear-door heat exchangers raise the grade somewhat, but less than cold plates do. Where no customer is nearby, the honest options are better cooling efficiency and cleaner electricity, and claiming reuse credit for heat nobody uses is worse than skipping it.
Where it fits, the economics improve with every degree of return temperature, every hour of matched demand and every metre of pipe you avoid. The software side costs little: telemetry you already collect, a forecast from the scheduler, and treating planned drains as heat events.
What to do next
- Measure your real capture fraction and return temperature per hall, from facility meters, not datasheets.
- Map heat demand within a few kilometres and get the customer's hourly load profile for a full year.
- Run the heat balance above with your numbers and the heat pump vendor's rated COP, and compute ERF with the ownership boundary written down.
- Check the current regulatory text for your jurisdiction (EED, EnEfG or local rules), including exemptions and reporting dates.
- Wire GPU power telemetry and the scheduler's maintenance calendar into a heat forecast the plant controller and the customer can both read.
- Size the buffer tank from your largest routine power dip, and test the dry cooler bypass under full load.