A GPU turns nearly every watt it draws into heat, and the cheapest way to throw heat away on a hot day is to evaporate water. That one fact explains why AI datacenters have become a local water issue, why two operators can report water figures that differ by a factor of ten for similar halls, and why switching off the evaporative towers is not automatically the greener choice. This article builds a water account from first principles: what is withdrawn and what is consumed, how much water it physically takes to reject a kilowatt-hour of heat, how cycles of concentration turn evaporation into a makeup demand, and how the water used by the power plant upstream fits in.
By the end you will have a small Python model for site and source water per hall and per training job, and a worked example showing when a dry-cooled design saves water overall and when it only moves it elsewhere. Metric definitions and meter boundaries for PUE and WUE are covered in PUE for GPU datacenters; the plant that moves the heat is covered in chilled water systems.
Withdrawal, consumption and discharge
Three words carry the whole subject, and mixing them up is the most common error in public discussion. Withdrawal is water taken from a utility, river or well. Consumption is the part that does not return to the local watershed, which in a datacenter is almost entirely water evaporated in cooling towers or adiabatic coolers, plus a little drift. Discharge is withdrawn water that goes back, mainly tower blowdown sent to the sewer. A site that withdraws one million litres and discharges a quarter of it has consumed three quarters.
Water Usage Effectiveness, the Green Grid metric, divides annual site water by annual IT energy, in litres per kWh. The definition does not settle which water: some operators count withdrawal, some count consumption, and boundaries differ on whether humidification, domestic use or reclaimed water are included. AWS, for example, reports a global WUE of 0.15 L/kWh for 2024; Microsoft reported 0.30 L/kWh for its 2024 fiscal year, down from 0.49 in 2021. These figures are useful trends inside each company and are not directly comparable with each other, or with a single new hall in a hot climate. Always ask which water, which boundary and which year.
Why evaporation is so effective
Evaporation is effective because changing water from liquid to vapour absorbs a lot of energy: the latent heat of vaporisation is about 2,430 kJ per kilogram near 30 degrees C. One kWh is 3,600 kJ, so rejecting one kWh of heat purely by evaporation boils off about 3,600 / 2,430 = 1.48 kg, or roughly 1.5 litres. That is a physical floor for an evaporative tower, not an engineering estimate. Real towers remove part of the heat as sensible heat in the air stream, so the evaporated share is usually somewhat below 100 percent; the model below uses 85 percent as an assumption you should replace with your tower vendor's figure.
The heat that reaches the towers is more than the IT load. Chillers add their compressor work, pumps and fans add theirs, and electrical losses in rooms cooled by the same plant add more. A useful planning ratio is 1.05 to 1.2 kWh of tower heat per kWh of IT energy when the towers carry the whole load. On a cool, dry hour the plant can bypass the towers entirely with dry coolers, and water use for that hour falls to near zero. Water is therefore driven by three things: IT energy, the fraction of hours the plant runs wet, and the share of heat removed by evaporation during those hours. Climate decides the second; design temperatures decide how much room you have to change it.
Cycles of concentration and blowdown
Evaporation leaves dissolved minerals behind. If nothing were removed, calcium, silica and chloride would concentrate until they scaled the heat exchangers or corroded the piping. Operators bleed off some concentrated water, called blowdown, and replace it with fresh makeup. The ratio of mineral concentration in the tower to that in the makeup is the cycles of concentration, N. A mass balance gives blowdown B = E / (N - 1) for evaporation E, so makeup is E + B plus a small drift loss.
| Cycles N | Blowdown as share of evaporation | Makeup per litre evaporated |
|---|---|---|
| 2 | 100% | 2.00 L |
| 3 | 50% | 1.50 L |
| 4 | 33% | 1.33 L |
| 6 | 20% | 1.20 L |
| 10 | 11% | 1.11 L |
Going from 3 to 6 cycles cuts withdrawal by a fifth without changing consumption, because evaporation is set by the heat. That is why water treatment is the first water project at most sites: it changes the withdrawn number, often the one a utility permit limits. The ceiling on N is makeup chemistry; hard or silica-rich water caps it low.
Site water and source water
The electricity also costs water. Thermal power plants with evaporative cooling consume water at the plant, and hydro reservoirs lose water to evaporation; wind and solar photovoltaic generation consume almost none. A widely used review of US generation by Macknick and colleagues at NREL (2011) put median consumption for tower-cooled plants at very roughly 0.7 L/kWh for combined-cycle gas and 2.5 to 2.7 L/kWh for coal and nuclear. A grid mix therefore has a water intensity that can range from near zero to above 1 L/kWh, and it changes by region and by hour.
Source water scales with facility energy, which is IT energy times PUE. That couples the two boundaries. A design that saves site water by running chillers instead of evaporating raises PUE, and on a thermal-heavy grid the extra electricity carries its own water. Microsoft said as much when it announced zero-evaporation designs for new sites from August 2024, noting that PUE would rise and that it would offset part of it with high-efficiency economizing chillers running at elevated water temperatures. The honest metric for a decision is site consumption plus source consumption, weighted by local water stress where each happens.
A water model you can run
The model below turns those relationships into numbers. Every parameter that is a property of your site is explicit, so you can replace assumptions with metered data as you get it.
from dataclasses import dataclass
H_FG = 2430.0 # kJ/kg, latent heat of water near 30 C
KJ_PER_KWH = 3600.0
@dataclass
class Site:
pue: float # facility energy / IT energy, annual
wet_frac: float # share of IT energy rejected while towers run wet
tower_ratio: float = 1.10 # kWh of heat at the towers per kWh IT
evap_share: float = 0.85 # share of tower heat removed by evaporation
cycles: float = 4.0 # cycles of concentration
drift: float = 0.0002 # drift as a share of evaporation (assumed)
grid_l_kwh: float = 1.0 # water consumed per kWh generated, regional
def per_it_kwh(s: Site) -> dict:
"""Litres per kWh of IT energy, by boundary."""
evap = s.wet_frac * s.tower_ratio * s.evap_share * KJ_PER_KWH / H_FG
drift = s.drift * evap
blowdown = evap / (s.cycles - 1)
return {
"site_withdrawn": evap + drift + blowdown,
"site_consumed": evap + drift,
"source_consumed": s.pue * s.grid_l_kwh,
}
def job_water(gpu_hours: float, kw_per_gpu: float, s: Site) -> dict:
"""Litres for one job; kw_per_gpu includes its share of host, NIC and switch."""
it_kwh = gpu_hours * kw_per_gpu
rates = per_it_kwh(s)
out = {k: v * it_kwh for k, v in rates.items()}
out["total_consumed"] = out["site_consumed"] + out["source_consumed"]
out["it_kwh"] = it_kwh
return outTwo simplifications are deliberate. The wet fraction is weighted by energy rather than hours, which is what you get by integrating a plant trend log; and source water uses a single annual grid factor, which you can replace with an hourly series if your utility or grid operator publishes one.
Worked example: one training job, two designs, two grids
Take one training job: 1,024 GPUs for 30 days, which is 737,280 GPU-hours, at 1.3 kW per GPU once its share of the host, network and storage is included. That is about 958,000 kWh of IT energy. Compare an evaporative design that runs wet for 60 percent of its IT energy at PUE 1.25 with a closed-loop design that never evaporates but runs at PUE 1.35, on two grids.
| Design and grid | Site withdrawn | Site consumed | Source consumed | Total consumed |
|---|---|---|---|---|
| Evaporative, grid 1.0 L/kWh | 1.06 ML | 0.80 ML | 1.20 ML | 2.00 ML |
| Closed loop, grid 1.0 L/kWh | 0 | 0 | 1.29 ML | 1.29 ML |
| Evaporative, grid 0.1 L/kWh | 1.06 ML | 0.80 ML | 0.12 ML | 0.92 ML |
| Closed loop, grid 0.1 L/kWh | 0 | 0 | 0.13 ML | 0.13 ML |
ML here is megalitres. Per IT kWh the evaporative site withdraws about 1.11 litres and consumes about 0.83. Three conclusions follow. First, on these assumptions the closed loop wins on total consumption on both grids: its extra 0.10 kWh per IT kWh would need a grid consuming over 8 L/kWh to cost as much water as the towers evaporate. Second, it costs about 8 percent more facility energy, which is real money and carbon. Third, where the water is taken matters more than the totals: a megalitre from a stressed aquifer next to a town is not a megalitre of river water at a distant power station. Weight site water by basin stress; the WRI Aqueduct baseline water stress map is a public starting point.
Facility water temperature is the lever
The design variable that moves water most is the temperature of the facility water. Air cooled GPU servers need cold supply air, which needs cold water, which needs chillers or evaporation on most summer days. Direct-to-chip liquid cooling lets the facility loop run much warmer, often in the 30 to 40 degree C range depending on the cold plate and GPU specification, and a warm loop can reject heat to dry coolers for many more hours a year. Higher rack density and liquid cooling, which AI hardware forces anyway, are therefore the main enabler of low-water sites; see liquid cooling for GPUs for the loop design.
Hybrid plants are the usual answer: dry coolers by default, evaporative assist only above a temperature threshold. That threshold is a dial between water and energy; set it per site with both prices in view, including a shadow price for water in stressed basins.
What schedulers can do
Software can shift water use as well as energy use, because evaporation peaks on hot afternoons. Flexible work, such as evaluation sweeps, data preprocessing and checkpoint-tolerant training, can be scheduled away from those hours, and a cluster scheduler can apply GPU power caps when the plant reports it is about to switch to wet operation. The same hooks used for carbon-aware scheduling work for water: export plant state as a signal, attach it to the job queue and record what each job consumed.
def water_signal(wet_bulb_c: float, it_load_frac: float, threshold_c: float = 21.0) -> str:
"""Plant hint for the scheduler: 'dry', 'assist' or 'wet'."""
if wet_bulb_c < threshold_c - 3:
return "dry"
if wet_bulb_c < threshold_c or it_load_frac < 0.6:
return "assist"
return "wet"
# In the admission loop: hold preemptible jobs and cap power when the plant runs wet.
if water_signal(current_wet_bulb(), current_load()) == "wet":
queue.defer(lambda j: j.preemptible and j.deadline_slack_h > 6)
set_power_cap(fraction=0.85) # trims heat to the towers during the peakThe function names are placeholders for your own interfaces; the contract is that the plant publishes a state, the scheduler reacts and both log it.
Failure modes
- Reporting withdrawal as consumption, or the reverse. The numbers can differ by a third or more; label every figure with its boundary.
- Using the design WUE as the annual WUE. Design-day water is the worst case; annual figures depend on weather, load and how often the towers run wet.
- Ignoring source water. A zero site figure on a coal-heavy grid can mean more total water, not less.
- Low cycles by default. Plants left at two or three cycles waste water through blowdown for years because nobody owns the chemistry.
- Unmetered makeup and blowdown. Without separate meters you cannot split withdrawal from consumption, and estimates drift.
- Drought planning left to the utility. A permit can be curtailed in a heat wave, exactly when the towers need water most; know your fallback and its PUE.
Trade-offs
Evaporative cooling gives the lowest energy use in most climates, at the price of consumption concentrated at the site. Closed-loop dry cooling removes site consumption and most permitting risk, at the price of more electricity, more equipment and some upstream water. Hybrids balance the two but need good controls. Reclaimed water spares potable supply but adds treatment and limits cycles. The right choice depends on grid water intensity, local stress and energy price, so run the model per site. For the energy side of the same decision, see PUE in AI datacenters.
What to do next
- Install or confirm separate meters for makeup, blowdown and any humidification, and log them hourly alongside IT energy.
- Report withdrawal and consumption separately, each with its boundary and year.
- Fit the model parameters from a year of plant logs: wet fraction, evaporation share and achieved cycles.
- Get a regional grid water factor and compute source water next to site water.
- Map each site to a basin stress score and set the wet-operation threshold with a water price that reflects it.
- Raise cycles of concentration as far as the makeup chemistry allows, and review the treatment contract.
- Expose plant state to the cluster scheduler and record per-job water with the same fields you use for energy.