Cooling, power, and networking are usually taught as three separate physics problems. On an actual build they are one coupled scheduling problem, and they are coupled through a single number: the megawatts the utility has agreed to deliver, on the date it has agreed to deliver them. Everything else — how many racks, how much they weigh, how they are cabled, when the first job runs — falls out of that envelope and the choices made underneath it. This article takes the deployment view rather than the physics view: the arithmetic that turns a power envelope into a GPU count, the physical envelope that decides whether the hardware can even reach its slot, the cable plant that quietly consumes both capital and rack power, and the commissioning that has to pass before a scheduler is allowed near the fleet. The mechanisms behind each axis are covered in depth in sibling articles and linked where they belong.
A cluster is scoped in megawatts, not GPUs
Procurement asks for a number of accelerators. The facility answers in megawatts, and the facility wins. Once a rack draws tens of kilowatts rather than the single digits that two decades of datacenter design assumed, a hall fills electrically and thermally long before it fills spatially: distribution runs out with floor tiles still empty. The planning primitive stops being the rack unit and becomes the powered, cooled, cabled rack position, and the project's unit of progress becomes energized commissioned positions per month rather than GPUs on a purchase order.
That inversion matters because the two supply chains run on completely different clocks. Accelerators arrive on an allocation and a product cadence measured in quarters. The envelope arrives on a chain measured in years: utility study and interconnect, transmission reinforcement, large power transformers, medium-voltage switchgear, chillers and heat-rejection plant, coolant distribution units. Whichever chain is longer sets the date, and it is almost always the facility. A team that plans backwards from a GPU delivery slot rather than forwards from an energization date will discover the mismatch at the loading dock.
There is a second-order effect worth designing for at the start. Because the silicon generation cadence is shorter than the build cycle, a hall dimensioned around today's rack kilowatts will be asked to host a denser generation before its first tenant has finished depreciating. Either the envelope carries deliberate headroom in the busway, the pipework, and the floor, or the hall is capped at its opening density for its entire service life.
From a megawatt of service to a number of GPUs
The capacity model is a short chain of multiplications, and its direction is the whole point: you do not pick a GPU count and go find power for it, you take the power you can actually get energized and find out what it buys.
Start at contracted utility capacity. Subtract facility overhead — heat rejection, pumps and fans, transformer and UPS conversion losses, everything that is not the load itself. That fraction is what PUE summarises, and it is a derate in this calculation rather than a sustainability metric; a design that improves it hands the difference straight back as usable IT power. What survives is the IT budget.
Then reserve for redundancy. The topologies themselves are explained in GPU Datacenter Power Requirements; the planning consequence is what belongs here. A fully duplicated power path means dual-corded equipment must be loaded so that either side can carry everything alone, so each side runs well under half its rating in normal operation. Usable IT power per delivered megawatt drops accordingly. Choosing full path redundancy over a redundant-module design is therefore not an abstract availability posture — it is a decision about how many GPUs the same utility service will host, and for a training fleet that checkpoints frequently, the trade often runs the other way than instinct suggests.
Apply the remaining derates: continuous-load rules mean a circuit is not run to its rating, and the diversity factor that lets a conventional hall oversubscribe its feeders is exactly the assumption a synchronised training job destroys. Divide the surviving kilowatts by realistic per-rack draw to get rack positions, and by accelerators per rack to get the GPU count.
One line gets forgotten in almost every first-pass model: the fabric, storage, and management infrastructure draw from the same IT budget. Leaf and spine switches, every transceiver, the storage tier feeding the data loaders, and the head and login nodes are not a rounding error at cluster scale. Budget them explicitly or the model will promise accelerators the hall cannot power.
Cooling is a floor-plan variable
The thermal ladder — contained air, rear-door heat exchangers, direct-to-chip cold plates, single-phase and two-phase immersion — along with the density thresholds that force each step, water chemistry, materials compatibility, and leak risk, is covered end to end in GPU Datacenter Cooling Overview. Nothing here restates it.
What the deployment view adds is that the rung you pick is a floor-plan decision before it is a thermal one. Cooling sets the ceiling on rack kilowatts; rack kilowatts against a fixed megawatt envelope set the rack count; rack count sets floor area, row length, and therefore how far a server sits from its leaf switch — which, as the cable plant section shows, is where a surprising share of the network bill is decided. A denser rung buys back floor area and shortens cable runs. It also consumes floor area of its own: coolant distribution units, manifolds, secondary loop pipework, and leak detection all need space and service access that an air design spends on containment and aisle depth instead.
Two practical traps recur on real builds. First, moving heat out of the rack is not the same as moving it off the site — a liquid design still needs rejection plant, and that plant sits on a lead-time chain comparable to the electrical gear. Second, direct-to-chip cooling handles the accelerators but a genuine fraction of rack heat remains air-borne: memory, network adapters, drives, and power supplies. The air path gets smaller, not deleted, and undersizing what remains is a classic retrofit failure that surfaces only under sustained full load.
The physical envelope: getting hardware to its slot
The constraint that stops more deployments than any thermal calculation is geometric. A dense accelerator rack is heavy, large, and shipped assembled, and every metre between the loading dock and its final position is a place the plan can fail.
Floor loading and point loads
Structural capacity is normally quoted as a distributed load over an area, but a rack does not distribute. Its entire mass lands on a handful of casters or levelling feet, so the governing figure is the concentrated point load, which can exceed what the distributed rating implies is safe. Liquid-cooled dense compute is substantially heavier than the legacy rack a building was designed around, and adding coolant to the loop adds more. Raised floors bring their own limits, set by pedestal, stringer, and panel ratings rather than by the slab underneath; rolling a rack across panels imposes a moving point load that can exceed the static one. The mitigation is boring and effective: a structural review of the whole route, not just the destination, plus load-spreading plates or a slab installation where the numbers do not close.
The freight path
Survey the path before the order, not after. Dock height and levelling, door and corridor widths, turning circles at corners, ramp gradients, threshold transitions, and above all elevator car dimensions and capacity if the hall is not at grade. An integrated rack that ships assembled and tested cannot simply be taken apart at a door that is too narrow without surrendering the factory integration and its warranty position. Where the path genuinely cannot take an assembled rack, the fallback is field integration — components delivered separately and the rack built in place — which costs schedule, needs staging space and skilled hands on site, and moves first-power-on later.
Service clearance
Density steals clearance, and clearance is what makes the hall maintainable. Budget aisle depth for rear service with containment doors open, room to pull a chassis fully out on its rails with its cabling still attached, and access to manifolds, valves, and coolant distribution units without disturbing neighbours. A hall where replacing a failed component requires powering down the adjacent rack has a real availability cost that is paid on every failure and never appears in the capacity model.
The cable plant is the networking you actually build
Fabric protocol and topology — RDMA and kernel bypass, subnet management, fat trees, oversubscription, rail-optimized wiring, congestion control on lossless links, and in-network reduction — belong to InfiniBand and NVLink architecture, with the intra-chassis scale-up domain covered in NVLink and NVSwitch architecture. The build has a separate and largely disjoint set of problems: the physical plant those topologies land on.
Volume. Because a training cluster typically fields an adapter per accelerator rather than per server, and adds storage, in-band management, and out-of-band paths on top, link count scales with GPUs, not with chassis. Every link is a physical run with a bend radius, a tray, a label, a length, and a test result. At cluster scale the cable plant is a scheduled work package with its own crew and its own critical path, not a tidy-up at the end of a rack build.
Reach, and why it is a money decision. Direct-attach copper is the cheapest option in both capital and watts, but its usable reach at high signalling rates is short — a few metres, shrinking with each rate step — which confines it to intra-rack and adjacent-rack runs. Active copper extends that somewhat; beyond it, optics take over and cost per link steps up sharply. The consequence is that physical layout determines the copper-to-optics mix, and that mix is one of the largest single swings in cluster capital cost. Keeping leaf switches top-of-rack shortens runs and keeps more links in copper. Rail-optimized wiring, which deliberately spreads same-index adapters across many leaf switches, tends to lengthen them and push the mix toward optics. Both are defensible; only one of them is free.
Power and heat. A transceiver dissipates at each end of each link. Per module the figure is modest, but multiplied by two ends and by a link count that scales with accelerators, it becomes a genuine line in the IT budget from the previous section. It is also heat generated inside switch chassis, in exactly the part of a rack that liquid-cooled designs frequently still cool with air.
Reliability. Optical modules are the highest-population active component in the cluster, and at that population even a low failure rate produces a steady trickle of replacements. Worse, they rarely fail cleanly. A marginal link retrains, corrects errors, and keeps passing traffic, so it presents as a job that is inexplicably slower than the arithmetic says it should be rather than as an alarm. Error counters per port are the only reliable detector, which is why they belong in burn-in acceptance criteria and in continuous monitoring afterwards.
Commissioning: proving the hall, then the fleet
Two distinct acceptance gates sit between a finished install and a schedulable cluster, and conflating them is how sites hand a fleet to users with faults already latent in it.
Integrated systems test
Facility commissioning proceeds in stages: components verified individually, then systems, then the whole plant loaded together while failures are deliberately induced. Drop the utility feed and watch the transfer to on-site generation; fail a chiller, a pump, a coolant distribution unit; fail one side of a redundant electrical path and confirm the survivor genuinely carries everything. Historically this is done with load banks precisely because the point is to find the interaction bugs — control sequences that fight each other, transfer timings that undershoot, setpoints that were never tested together — while the load is a resistor rather than a month of training progress. A hall that has never run at full load through an induced failure has been installed, not commissioned.
Fleet burn-in
Hardware infant mortality is real and heavily concentrated in the first weeks of service, so the fleet gets its own soak before anyone is allowed to schedule on it. A useful sequence: sustained dense-arithmetic load on every accelerator long enough to bring power draw and thermals to genuine steady state, not just to peak; memory testing that surfaces correctable and uncorrectable error rates; then multi-node collective bandwidth tests sized to the message profile the real jobs will use. That last step is the one that matters most and the one most often skipped, because a point-to-point test on an otherwise quiet fabric will not reveal the mis-seated optic or the cable that only retrains under load — it takes every link busy at once.
Write the acceptance criteria down before the tests run: achieved per-node collective bandwidth against the arithmetic, port error-counter thresholds, ECC rates, sustained clocks and thermal margin under soak. Anything failing goes back to the vendor while it is still their problem. After handover, the same tests become scheduled health checks, because a node that degrades quietly will otherwise be diagnosed by a slow training run days later.
Phasing against a moving envelope
Energization arrives in stages and accelerator deliveries arrive in stages, and the two slip independently. Designing for that is cheaper than reacting to it.
The first principle is that a partially energized hall should be a usable cluster rather than a stranded one. Fill by pod, not by convenience: each increment should stand up a coherent, schedulable unit with its own fabric locality and its own power distribution, so that a phase which lands early is immediately productive and a phase which lands late does not fragment what already exists. A scattering of racks across the floor that share no leaf switches is worth far less than the same silicon packed into one working pod.
The second is to hold a deliberate reserve of rack positions, busway capacity, and switch ports. Density and form factor change under a multi-year build, and a hall with no spare positions cannot absorb the next generation without evicting the current one. The reserve is not slack; it is the option to keep deploying.
The third is to keep the capacity model alive rather than filing it after design freeze. Measure real rack draw per phase against what the model assumed, and correct the derates with evidence. Sites that do this recover stranded capacity they had reserved out of caution; sites that do not either strand it permanently or discover the assumption was optimistic at the least convenient moment.
Behind all of it sits the same conclusion the arithmetic keeps producing. Accelerators can be bought faster than they can be powered, and the queue that decides when a cluster exists is an electrical one. Capacity planning for AI infrastructure has become a siting exercise with a compute deliverable attached.
Treat a GPU deployment as a power project with a compute deliverable. GPU count is an output of the capacity chain — contracted megawatts, minus facility overhead, minus the redundancy reservation, minus continuous-load derates, minus the fabric and storage draw — not an input to it. The cooling rung you choose sets rack kilowatts and therefore floor plan and cable lengths; the floor plan sets the copper-to-optics mix, which is one of the biggest capital swings in the build; and the physical envelope of point loads, freight paths, and service clearance decides whether the hardware reaches its slot at all. Prove the hall with an integrated systems test under induced failure, then prove the fleet with a soak and a full-fabric collective test that finds marginal links before a training job does. Fill by pod so every phase is usable, hold a reserve so the next generation has somewhere to land, and remember that the binding constraint is an energized megawatt, not silicon.