A modular datacenter is built from prefabricated blocks instead of being assembled piece by piece on site. Switchgear, UPS and transformers arrive as a power skid, chillers and pumps as a cooling skid, and in the furthest form racks, coolant distribution units and busway arrive as a finished IT module that only needs to be set down, connected and tested. For AI the approach has moved from niche to mainstream for a simple reason: the racks changed faster than buildings can. Vendor specifications for a liquid-cooled GB200 NVL72 rack put it at roughly 120 to 132 kW, several times what most existing white space was designed for, and a hall built for the last generation can be wrong for the next one.

This article treats modules the way an infrastructure or ML platform engineer should: as interfaces with numbers attached, and as failure and maintenance domains that the job scheduler must understand. The facility side of sizing feeders and cooling loops is covered in GPU datacenter power and liquid cooling in depth; site commissioning and phasing are in GPU datacenter deployment. Here the focus is what changes when the building is a kit: the contracts between modules, factory versus site testing, how module boundaries should show up in topology labels and placement, and the failure modes that are specific to prefabrication. Worked numbers are illustrative assumptions, not vendor figures, unless stated otherwise.

What modular means

Modular is a spectrum, and the word hides very different products. At one end are prefabricated power and cooling skids that plug into an otherwise conventional building. In the middle are prefabricated data halls, steel structures with containment, busway and pipework built in. At the far end are containerised or fully integrated IT modules that ship with racks installed and can sit on a pad outdoors. A fourth form is rack-scale integration, where a whole rack of GPUs, switches, manifolds and power shelves is built and burned in at the factory and rolled in as a unit; it is modular at the IT layer even inside a traditional building.

What these share is that integration work moves from the site to a factory, and the site does foundations, utility connections and the joints between modules. That is the value proposition: shorter schedules, more predictable quality, and capacity added in steps that match GPU deliveries.

Why AI pushes towards modules

Three pressures push AI capacity towards modules. Density: a hall designed around air-cooled racks of 10 to 20 kW cannot simply host racks an order of magnitude denser; it needs liquid to the rack, heavier floors and bigger feeders. Schedule mismatch: accelerator generations arrive faster than a conventional facility is built, so modules let later blocks be specified later, against hardware that actually exists. Power in tranches: a site may have a few megawatts now and much more in two years, and modules let you use the first tranche without waiting for the full build. The cost is that every module boundary is a joint, and joints are where the interesting failures live.

Interfaces are contracts

Treat every module boundary as an API. If the interface is not written down as numbers with tolerances, the power vendor, cooling vendor and integrator will each assume the other side handles the awkward cases, and the gaps surface during the first full-load test. A useful contract lists, for each boundary, the steady-state value, the transient envelope and the failure behaviour.

A modular AI hall is a set of skids joined by written interface contractsPower moduleMV in, LV busway outCooling modulechillers or dry coolersIT moduleracks, CDUs, buswayliquid loop + residual airNetwork modulespine, fibre trunksControlsBMS, EPMS, DCIMkW, voltage, rampsupply temp, flowfibre count, latencyalarms, telemetryEach arrow is a contract: a number with a tolerance, tested at the factory and again on site.Each box is a failure and maintenance domain the scheduler should know about.
Power, cooling, network and controls modules each meet the IT module across a boundary that should be specified like an API, with values, transients and failure behaviour.
BoundarySteady stateTransientFailure behaviour
Power to ITRated kW per busway and per rack, voltage, phase balanceLoad step when a synchronous job starts or all ranks hit a barrierWhat the UPS does on a utility blip, ride-through time, which racks drop first
Cooling to ITSupply temperature, flow per rack, pressure drop, water qualityFlow and temperature recovery after a pump changeoverHow long racks can run on a failed pump or chiller before throttling
Network to ITFibre count, connector type, optic reach, path diversityNone at the physical layerWhich links share a trunk, so one cut is many link failures
Controls to operatorTelemetry points, units, sample rate, protocolAlarm latencyWhat is reported when the controller itself is down

The transient column matters more for AI than for general compute. Synchronous training makes thousands of GPUs change power together: at a collective or checkpoint they idle, then they all resume. The power module sees a square wave, and the cooling module sees a heat load that swings faster than a chilled-water plant was traditionally tuned for. If the contract only says rated kilowatts, nobody owns the ramp. Write the expected step size and rate into the contract and test it with load banks that can switch, not just with a steady resistive load.

Modules as failure domains

A module is a blast radius. When a power skid trips, a cooling loop loses flow or an IT module is taken down for maintenance, everything inside it goes with it. In a stick-built hall these domains exist too, but they are smeared across shared plant. In a modular build they are crisp and documented, and that is a gift to the software side, because the scheduler can be told about them.

Synchronous data-parallel training fails as a unit: lose one rank and the whole job stops and restarts from its last checkpoint. A job spread across three modules therefore inherits the interruptions of all three. The placement rule follows directly: pack a training job inside as few modules as possible, and spread its checkpoint storage and control plane across modules so the thing you restart from survives the thing that failed.

Mapping physical modules to scheduler topologyIT module Apower block P1, cooling loop L1rack 1rack 2rack 3label module=maIT module Bpower block P2, cooling loop L2rack 1rack 2rack 3label module=mbIT module Cpower block P3, cooling loop L3rack 1rack 2rack 3label module=mcSchedulertopology labels: hall / module / power block / rack / NVLink domainTraining jobpack inside one module where possibleCheckpoint storagespread across modules
Module boundaries become scheduler labels. Training jobs are packed into one module where they fit; checkpoint storage and control services are spread across modules.

The mechanism is boring and effective: generate topology labels from the asset inventory rather than typing them by hand. The sketch below reads a module inventory and emits Kubernetes node labels and a Slurm topology block, so the two can never disagree with the physical layout.

import json

# inventory.json: one record per node, produced by the integrator at handover
# {"node": "gpu-a-r01-n03", "module": "ma", "power_block": "p1",
#  "cooling_loop": "l1", "rack": "a-r01", "nvlink_domain": "a-r01"}

def k8s_labels(rec):
    return {
        "topology.example.com/module": rec["module"],
        "topology.example.com/power-block": rec["power_block"],
        "topology.example.com/cooling-loop": rec["cooling_loop"],
        "topology.example.com/rack": rec["rack"],
    }

def slurm_topology(records):
    # leaf switch per rack, parent per module, one root so jobs can still span modules
    by_module = {}
    for r in records:
        by_module.setdefault(r["module"], {}).setdefault(r["rack"], []).append(r["node"])
    lines = []
    for module, racks in sorted(by_module.items()):
        for rack, nodes in sorted(racks.items()):
            lines.append(f"SwitchName={rack} Nodes={','.join(sorted(nodes))}")
        lines.append(f"SwitchName={module} Switches={','.join(sorted(racks))}")
    lines.append(f"SwitchName=spine Switches={','.join(sorted(by_module))}")
    return "\n".join(lines)

records = json.load(open("inventory.json"))
for r in records:
    labels = " ".join(f"{k}={v}" for k, v in k8s_labels(r).items())
    print(f"kubectl label node {r['node']} {labels} --overwrite")
print(slurm_topology(records))

Slurm reads the switch hierarchy from topology.conf when the tree topology plugin is enabled, and prefers to place a job under the lowest common switch. Making each module a parent switch nudges the scheduler towards packing jobs inside a module, which lines up the network locality and the failure domain at the same time.

Worked example: three modules, one job

Take an illustrative site with power for about 6.5 MW of IT load arriving in three tranches. The design uses three identical IT modules of 16 rack positions, each paired with its own power block and cooling loop. Planning at 130 kW per rack, a module carries about 2.1 MW of IT load. With a 72-GPU rack-scale system, one module holds 16 x 72 = 1,152 GPUs.

Suppose each module averages one interruption per month that stops every rack inside it, counting planned maintenance that could not be deferred plus unplanned trips. A job inside one module sees about one interruption a month; a job spanning all three sees about three. If the job checkpoints every 30 minutes and restarting costs 20 minutes, each interruption wastes on average 15 minutes of lost progress plus 20 minutes of restart, about 35 minutes. Across three modules that is 105 minutes a month instead of 35. The absolute numbers are small here, but the multiplier is not: it scales with the number of domains a job touches, and with real fleets where restarts at scale take far longer than 20 minutes. Planned work benefits too: with the maintenance calendar visible to the scheduler, module B can be drained ahead of a cooling service while new jobs land in A and C.

Factory testing, site testing and GPU burn-in

Prefabrication splits testing in two. Factory acceptance testing (FAT) proves the module against its own specification: pressure tests on pipework, insulation and protection tests on switchgear, control logic exercised point by point, and ideally the module run at load with heaters or load banks. Site acceptance testing (SAT) proves the joints: utility connections, inter-module piping and cabling, and the behaviour of the integrated system under failure. Integrated systems testing then pulls the plug on real equipment, for example failing a pump or opening a breaker under load, and checks that the system degrades the way the contract says.

None of that tells you whether GPUs train. After the facility tests pass, run a software burn-in on every node before it joins the scheduler pool: GPU diagnostics (for example dcgmi diag -r 3 on NVIDIA systems), a long all_reduce_perf run from nccl-tests across each rack and then across the module, and a short real training job that exercises sustained power. Record bandwidth and temperatures per node; a slow node at handover will hold back every collective later. Repeat factory burn-in on site, because shipping is a vibration test nobody asked for.

From facility alarms to scheduler state

Each vendor module arrives with its own controller and alarm list; merge them into one stream. The highest-value link is from facility alarms to scheduler state: when a cooling loop reports loss of redundancy, the nodes on that loop should stop accepting new long jobs before anything overheats. A small controller is enough.

SEVERE = {"cooling_redundancy_lost", "ups_on_battery", "leak_detected"}

def on_facility_event(event, nodes_by_domain, scheduler):
    # event example: {"domain": "cooling_loop=l2", "type": "cooling_redundancy_lost"}
    if event["type"] not in SEVERE:
        return
    nodes = nodes_by_domain.get(event["domain"], [])
    reason = f"facility:{event['type']}:{event['domain']}"
    for node in nodes:
        # Slurm: scontrol update nodename=<node> state=drain reason=<reason>
        # Kubernetes: kubectl cordon <node>
        scheduler.drain(node, reason)
    scheduler.notify_jobs(nodes, action="checkpoint_now", reason=reason)

Drain rather than kill: running jobs continue, new work goes elsewhere. Pair the drain with a human-approved undrain, or a flapping sensor will cycle a module all day. Keep the facility event bus off the GPU network, and alarm on the controller itself going silent, because a dead controller looks exactly like a healthy, quiet one.

Operating and expanding a modular site

Keep modules interchangeable: common spares (pumps, power shelves, fans, optics), the same controller and power-shelf firmware across modules, and water chemistry sampled per loop, since modular builds have more, smaller loops. Plan expansion tie-ins in the first phase: valved stubs on cooling headers, spare breaker positions, spare fibre to the spine. Mixing hardware generations is normal; give each module a generation label so the scheduler does not split one job across GPUs of different speed.

Failure modes

  • Contract gaps. Each vendor tests its module to its spec, and the integrated system fails at the joint. Symptom: everything passed FAT, and SAT finds flow imbalance or a protection coordination problem between skids.
  • Unmodelled transients. Steady load banks pass; the first synchronous training run trips a breaker or drives coolant temperature over the limit after a checkpoint stall.
  • Shared hidden dependencies. Three modules that look independent share one fibre trunk, one controls network or one makeup-water line, so a single fault crosses the boundaries the scheduler trusts.
  • Labels that drift. Topology labels typed by hand, then a rack is moved. Jobs are packed into what the scheduler thinks is one module and are really in two.
  • Transport damage. Loosened fittings, unseated cables and damaged optics after shipping. Leak tests and burn-in on site catch them; skipping either does not.

Trade-offs

ChoiceGainsCosts
Prefabricated modulesSchedule, factory quality, capacity in stepsTransport limits on size, more joints, vendor-specific controls
Stick-built hallFlexible layout, fewer joints, shared plant efficiencyLong schedule, on-site quality variance, harder to phase
Small modulesSmall blast radius, fine-grained expansionMore boundaries, jobs span more domains
Large modulesBig jobs fit inside one domainBigger blast radius, coarser expansion steps
Shared central plantEfficiency at scale, fewer sparesOne failure domain spans many IT modules

Size IT modules so the common large job fits inside one.

What to do next

  1. Write the interface contract for each module boundary as values, transients and failure behaviour, and get every vendor to sign the same document.
  2. Ask for switched load bank tests that imitate a synchronous training step, not only steady load.
  3. Generate scheduler topology labels from the integrator's asset inventory and diff them weekly against reality.
  4. Make the module a parent switch in Slurm topology, or a topology key in Kubernetes, and pack large jobs inside one module.
  5. Spread checkpoint storage and control services across modules.
  6. Wire severe facility alarms to automatic drain and on-demand checkpoint, with a human-approved undrain.
  7. Run GPU diagnostics and nccl-tests on every node on site, and keep the results as the baseline.
  8. Read GPU datacenter cooling for the cooling options each module type can host.
Key takeaway: A modular datacenter turns a building into a kit of power, cooling, network and IT modules. The gains are schedule, quality and capacity in steps; the risks sit at the joints. Specify every boundary as a contract that includes transients, test the integrated system under failure, and give the scheduler the module map so jobs are packed inside a failure domain and drained before a facility fault reaches the GPUs.