For decades a datacenter rack drew something like 5 to 10 kilowatts, and buildings were designed around that number: air cooling, raised floors, and power distributed on the assumption that every rack position looked roughly the same. AI hardware broke the assumption. An eight-GPU H100 server alone has a maximum system power of around 10 kW, so four of them make a 40 kW rack, and NVIDIA's GB200 NVL72 rack draws on the order of 120 to 130 kW. Power density, kilowatts per rack and per unit of floor area, has become the variable that decides which buildings can host AI at all.

This article explains density from first principles: where it comes from, the four physical limits it runs into, why AI systems are deliberately built dense rather than spread out, and what density means for the software that places jobs. It includes a worked calculation for air and liquid cooling, a small planning tool, and a checklist. Total facility power and distribution topology are covered in GPU datacenter power requirements; this page focuses on what happens when that power is concentrated.

Where density comes from

Density is multiplication. A modern training GPU is rated around 700 W (H100 SXM) to roughly 1,000 W for Blackwell-generation parts, varying by product and configuration. A server adds CPUs, memory, network adapters, storage and fans, typically another 30 to 50 percent on top of the GPUs. Put several servers in a rack, or build the rack as one system as NVL72 does with 72 GPUs, and the rack figure lands anywhere from about 40 kW to well over 100 kW.

How per-GPU power compounds into rack density, and the limits each step hitsGPU~0.7 to ~1 kWServer / trayGPUs + CPUs, NICs, fansRack~40 kW to 130+ kWRow / hallMW, fixed by utilityHeat removalair ~ tens of kW/rackElectrical feedamps per rack, buswayFloor loadingtonnes per rackNetwork reachcopper NVLink in-rackScheduler and placementpower budgets per rack/row, caps, topology-aware job placement
Per-GPU power compounds through the server and the rack. At rack level it runs into four limits at once, and the scheduler has to respect whichever binds first.
Rack typeApproximate powerTypical cooling
Traditional enterprise5-10 kWRoom air
Dense CPU / storage10-20 kWAir with containment
4 x 8-GPU H100 servers~40 kWAir at the limit, or rear-door heat exchangers
GB200 NVL72~120-130 kWDirect liquid cooling

These are design-level figures; actual draw depends on workload, power caps and configuration, and vendors' roadmaps point to considerably higher densities. Treat any unreleased rack figure as unconfirmed until the shipping specification exists.

The cooling arithmetic

Every watt a rack draws becomes heat that must be carried away. For any coolant, the heat removed is Q = mass flow x specific heat x temperature rise. For air, with density about 1.2 kg/m3 and specific heat about 1,005 J/(kg K), removing 1 kW with a 12 K rise from inlet to exhaust needs about 0.069 m3/s, roughly 146 cubic feet per minute (CFM).

Scale that up. A 10 kW rack needs about 1,460 CFM, which ordinary fans and containment handle. A 40 kW rack needs about 5,900 CFM through a single rack face, which demands aggressive containment, high fan power and often a rear-door heat exchanger. A 120 kW rack would need about 17,500 CFM, an air velocity and fan power that are impractical in a standard rack footprint. Fan power also grows steeply with airflow, so the energy spent on cooling rises faster than the heat load. That is why air cooling tops out in the tens of kilowatts per rack.

Water carries about 3,500 times more heat per unit volume than air. Removing 120 kW with a 10 K rise needs about 2.9 kg/s of water, roughly 170 litres per minute, through pipes a few centimetres across. Glycol mixtures have lower specific heat and need somewhat more flow. This is the physics behind the move to direct liquid cooling: not a preference, but the only practical way to remove that much heat from that little space. Even liquid-cooled racks still reject a residual fraction of heat to air from power supplies, memory and network gear, which the room must still handle.

The other limits: current, weight and stranded capacity

Electrical feed. Current scales with density. At 415 V three-phase and a power factor of 0.95, current is I = P / (1.732 x V x pf), so a 120 kW rack draws about 176 A per phase. That needs heavy busway taps or multiple high-current feeds per rack, and every breaker, cable and connector in the path must be rated for it with margin. Higher distribution voltages reduce current; in May 2025 NVIDIA announced an 800 V DC distribution architecture, planned from 2027, for exactly this reason.

Floor loading. Dense racks are heavy: liquid-cooled rack-scale systems weigh well over a tonne, with copper cabling, cold plates, manifolds and coolant. Raised floors built for older racks may not carry that, which is one reason AI halls favour slab floors with overhead power and piping. Check point and distributed load ratings, and the route from loading dock to rack position, before a delivery date.

Heat flux in the room. Even with liquid cooling, concentrated racks create hot spots, and the facility's water loop, chillers or dry coolers must handle the total. Density changes the distribution of capacity in a hall, not the total.

Stranded capacity. A hall with a fixed utility feed can be stranded in two directions. Fill it with low-density racks and you run out of floor space with power unused. Place a few very dense racks in a hall built for air and you run out of cooling with floor and power unused. Good designs match density to the hall's power, cooling and floor limits at the same time; the planning view is in GPU datacenter deployment.

Why AI racks are built dense on purpose

If density causes all these problems, why not spread GPUs out? Because training performance depends on how fast GPUs can talk to each other, and the fastest links are short. NVLink within an NVL72 rack runs over copper cables, which are cheap, low-power and reliable but only work over short distances. Keeping 72 GPUs inside one copper NVLink domain lets tensor and expert parallelism, which need the highest bandwidth, run entirely within a rack. Spreading the same GPUs across several racks would require optical links, which cost more, draw more power per bit and add failure points.

So density is a deliberate trade: accept harder cooling and power delivery to make the communication fabric cheaper and faster. For software, the consequence is that the rack is now a meaningful unit of parallelism. A job that maps its tensor-parallel groups onto one NVLink domain and its data-parallel groups across racks gets the benefit; a job placed carelessly across rack boundaries pays for the density without using it.

Density also concentrates failure. When one 130 kW rack loses coolant flow or a power shelf, 72 GPUs go with it, which is a far larger slice of a job than one eight-GPU server. Plan spare capacity in whole racks, keep checkpoints frequent enough that losing a rack costs minutes rather than hours, and make sure the job launcher can re-form its parallel groups on a different rack without manual edits to topology files. Maintenance changes too: servicing a liquid-cooled rack-scale system often means draining the whole rack, not pulling one server while its neighbours keep running.

Worked example: planning a 12 MW hall

Suppose a hall has 12 MW of IT power and the operator is choosing between air-cooled racks of 4 x 8-GPU servers at about 40 kW each, and liquid-cooled rack-scale systems at about 130 kW each. The planning question is how many racks fit each limit. The script below checks airflow, coolant flow, current and floor area for a given rack design and reports which constraint binds.

from dataclasses import dataclass
from math import sqrt

AIR_RHO, AIR_CP = 1.2, 1005.0          # kg/m^3, J/(kg K)
WATER_CP = 4186.0                      # J/(kg K); use ~3,900 for 25% glycol

@dataclass
class Rack:
    name: str
    kw: float
    liquid_fraction: float            # share of heat removed by liquid
    footprint_m2: float               # rack plus share of aisle

def plan(rack, hall_kw, hall_m2, air_dt=12.0, water_dt=10.0, volts=415.0, pf=0.95):
    w = rack.kw * 1000
    air_w = w * (1 - rack.liquid_fraction)
    cfm = air_w / (AIR_RHO * AIR_CP * air_dt) * 2118.88
    lpm = (w * rack.liquid_fraction) / (WATER_CP * water_dt) * 60
    amps = w / (sqrt(3) * volts * pf)
    by_power = int(hall_kw // rack.kw)
    by_floor = int(hall_m2 // rack.footprint_m2)
    n = min(by_power, by_floor)
    binding = "power" if by_power <= by_floor else "floor"
    return dict(rack=rack.name, cfm_per_rack=round(cfm), lpm_per_rack=round(lpm),
                amps_per_phase=round(amps), racks=n, binding=binding,
                hall_air_cfm=round(cfm * n))

hall_kw, hall_m2 = 12_000, 2_500
for r in [Rack("air 4x8 GPU", 40, 0.0, 3.0), Rack("liquid rack-scale", 130, 0.85, 3.5)]:
    print(plan(r, hall_kw, hall_m2))

With these assumptions the air-cooled design fits 300 racks by power but each needs about 5,900 CFM, about 1.8 million CFM for the hall, and draws about 59 A per phase. The liquid design fits 92 racks, each needing about 160 litres per minute of coolant and about 190 A per phase, with residual air of roughly 2,900 CFM per rack, about 260,000 CFM for the hall. Both are bound by power, not floor, but the liquid hall uses about a third of the floor area of the air hall (roughly 320 against 900 square metres), moves about a seventh of the air, and its racks form much larger NVLink domains. The air design's real binding constraint is likely the room's air handling, which the script makes visible by reporting hall airflow. Replace the assumptions with your vendor's figures before trusting any of these numbers.

What density means for software

Density pushes power and thermal limits down to the rack and row, where they become scheduling constraints. Three practices follow.

Power budgets per domain. Track a power budget for each rack, PDU and row, and place jobs so expected draw stays under it. Read actual draw continuously with nvidia-smi --query-gpu=power.draw,power.limit --format=csv or DCGM, and compare it with the budget rather than with nameplate.

Caps as a density lever. nvidia-smi -pl <watts> lowers a GPU's power limit. Training throughput usually falls more slowly than power near the top of the curve, so a modest cap can let a power-limited row host more GPUs for more total throughput. Measure the curve for your model before relying on it.

Topology-aware placement. Place each job's tightly coupled groups inside one rack's NVLink domain, and spread unrelated jobs so a cooling fault in one rack does not hit every replica of a service. Feed CDU and rack thermal alarms into the scheduler, as described in liquid cooling for GPU datacenters.

Failure modes

  • Designing to nameplate. Sizing every feed to the maximum of every component strands capacity; sizing to typical draw without caps risks trips during synchronized peaks. Use measured draw plus enforced caps.
  • Air-cooled hot spots. A few dense racks in an air hall recirculate hot exhaust and throttle neighbours. Watch GPU throttle reasons, not only room temperature.
  • Coolant flow loss. At 100 kW or more a rack has seconds before GPUs throttle when flow stops; the scheduler must treat CDU alarms as node health events.
  • Phase imbalance. Uneven load across phases at high current overheats conductors and trips breakers; balance placement and power-shelf configuration.
  • Mixed densities. Halls that mix generations end up with neither air nor liquid capacity used well; zone them deliberately.

Trade-offs

ChoiceGainsCosts
Higher densityLarger NVLink domains, less floor, shorter cablesLiquid cooling, heavy feeds, tougher maintenance
Lower densityAir cooling, simpler retrofit of existing hallsMore floor, more optics, smaller fast domains
Power capsMore GPUs per budget, smoother loadLower peak throughput per GPU
Liquid retrofit of an air hallHosts current AI racksPipework, floor loading checks, downtime

Density also shifts efficiency accounting: liquid cooling usually lowers fan and chiller energy per IT watt, which shows up in facility metrics covered in PUE for GPU datacenters.

What to do next

  1. Measure actual per-rack draw for your training and inference jobs, at current power limits, and compare it with each rack's rated feed.
  2. Run the planning script with your hall's power, floor and cooling figures and identify the binding constraint.
  3. Confirm floor load ratings and delivery routes before ordering rack-scale systems.
  4. Add per-rack and per-row power budgets to your scheduler, and place tightly coupled groups inside one NVLink domain.
  5. Profile throughput against power cap for your main model and set caps where the curve flattens.
  6. Route CDU and rack thermal alarms into node health so a cooling fault drains jobs instead of throttling them silently.
Key takeaway: Power density is per-GPU watts multiplied up to the rack, and at AI densities it hits the limits of air, current and floor loading together. AI racks are built dense because short copper NVLink makes a rack a fast parallelism domain. Plan with measured draw and real physics, enforce power budgets and caps in the scheduler, and place tightly coupled work inside one rack.