A modern AI rack draws more than a hundred kilowatts, and the last two metres of that journey happen on a pair of copper bars running up the back of the rack. Power shelves convert facility AC to roughly 50 volts DC, a busbar carries it to every tray, and each tray clips on without a power cord. At these voltages the current is measured in thousands of amps, so the busbar is where conductor physics, rack design and the behaviour of a synchronised training job meet.
The datacenter power requirements article covers the facility side: utility interconnects, medium voltage, overhead busway versus individual whips, PDUs and N+1 or 2N topologies. This page stays inside the rack. It explains why DC busbars replaced cords, works through the current, voltage drop and heat for a realistic rack, shows where busbars fail, what training jobs do to them, which software controls reach them, and why the industry is moving to 800 volts DC.
The path from grid to GPU core
Follow one watt from the grid to a GPU core. The utility delivers medium-voltage AC. Transformers and the UPS system bring it to a low-voltage AC distribution level, commonly around 415 volts three-phase. In an Open Rack V3 style or NVIDIA MGX style rack, power shelves then rectify that AC into DC at a nominal 48 to 54 volts. The DC goes onto a vertical busbar, a pair of rigid copper conductors for supply and return. Each compute or switch tray carries a blind-mate connector that grips the bar when the tray is pushed home.
Inside the tray, the voltage keeps falling while the current keeps rising. An intermediate bus converter typically produces around 12 volts, though some designs convert from 48 volts closer to the load, and multiphase voltage regulators feed the GPU and CPU at below one volt. A die that consumes 1,000 watts at 0.8 volts draws over a thousand amps on the package. Every stage of this chain loses a few percent, and every conductor between stages loses I^2 R.
Why a busbar, and why 48 to 54 volts
The reason for a busbar is the relation I = P / V. A 120 kilowatt rack at 54 volts draws about 2,222 amps. Carrying that on cords would need dozens of heavy cables, each with its own plug, strain relief and failure point, plus airflow obstruction at the back of the rack. A rigid bar can carry the whole rack current with a large cross-section, short paths and one connector per tray.
Why 48 to 54 volts rather than 12? The older 12 volt rack bus would need about 10,000 amps for the same 120 kilowatts, and since conductor loss scales with the square of current, the 12 volt bus would lose about 20 times more for the same copper. Moving the rack bus to the 48 volt class was the step that made 30 to 100 kilowatt racks practical. Why not higher still inside the rack? Below 60 volts DC the bus generally falls in the safety extra-low voltage band of IEC rules, which keeps hot-swapping trays and touching the rear of the rack simple. Going higher brings arc and shock hazards that need new connectors, guarding and procedures, which is exactly the trade the 800 volt proposal accepts.
Worked example: current, drop and heat
Size a busbar for a 120 kilowatt rack at 54 volts, about 2,222 amps. Take two copper bars, supply and return, each 10 by 80 millimetres, so 800 square millimetres, and 2 metres tall. The loop length is 4 metres. With copper resistivity of 1.72e-8 ohm metres at 20 degrees Celsius:
R = rho * L / A = 1.72e-8 * 4.0 / 800e-6 = 8.6e-5 ohm (86 micro-ohm)
I = P / V = 120_000 / 54 = 2,222 A
All load at the far end (worst case):
drop = I * R = 2,222 * 8.6e-5 = 0.19 V (0.35 percent of 54 V)
loss = I^2 * R = 2,222^2 * 8.6e-5 = 425 W
Load spread evenly along the bar, fed from one end:
drop = I * R / 2 = 0.10 V
loss = I^2 * R / 3 = 142 W
Copper mass = 2 bars * 800e-6 m^2 * 2 m * 8,960 kg/m^3 = 28.7 kgThree lessons fall out. The voltage drop is small; the regulators downstream absorb it easily. The heat is not trivial: around 140 watts in the bars, plus the connector losses, has to leave through the same rear airflow or liquid loop as everything else, and copper resistance rises about 0.4 percent per degree, so a hot bar loses more. Feeding the bar from the middle instead of one end halves the length each half carries and cuts conductor loss to a quarter of the end-fed value. Finally, scale it. A 1 megawatt rack at 54 volts would draw about 18,500 amps, and NVIDIA's own estimate is up to 200 kilograms of copper busbar for such a rack, which is the argument for raising the voltage.
Power shelves and redundancy
Power shelves are 1U or 2U chassis holding several hot-swappable rectifier modules whose outputs are paralleled onto the busbar. They share current through droop or active current sharing, and the shelf count is chosen for redundancy: with N+1, losing any one module or shelf still leaves enough capacity for the rack. NVIDIA describes its GB200 and GB300 NVL72 racks as carrying up to eight power shelves, and ORv3 shelves can be paired with battery backup units that hold the bus up through short outages.
The arithmetic every operator should do is usable capacity, not nameplate. Suppose a hypothetical rack has six 33 kilowatt shelves, 198 kilowatts of nameplate. Under N+1 at shelf level, usable capacity is five shelves, 165 kilowatts. If each shelf is fed from one of two facility feeds, A and B, and you want to survive the loss of a whole feed, usable capacity is three shelves, 99 kilowatts, which is below the 120 kilowatt rack. That is not a wiring mistake to discover during a feed outage. It is a design decision about whether a feed loss should be ridden through, handled by power capping, or accepted as a rack outage.
Contacts are where busbars fail
Busbars themselves rarely fail; their joints do. Every tray connection, every bolted splice and every shelf output is a contact with a small resistance, and the power dissipated there is I^2 R_contact concentrated in a few square millimetres. A 6 kilowatt tray draws 111 amps at 54 volts. With a healthy contact resistance of 50 micro-ohms the clip dissipates about 0.6 watts. If fretting, contamination or a partly seated tray raises that to 500 micro-ohms, the clip dissipates over 6 watts in a tiny volume, heats up, oxidises further and climbs toward failure.
This failure mode is gradual and invisible to the GPU software. What operators see first is sometimes a tray whose input voltage reads lower than its neighbours under load, or a hot spot on a thermal camera. Treat tray input voltage telemetry as a health signal: compare each tray to the rack median under the same load, and pull any tray that sits persistently low.
What training jobs do to the busbar
A synchronous training job makes thousands of GPUs change power together. When a step begins, compute ramps every GPU toward its limit within milliseconds; when the job waits on an all-reduce, a data loader stall or a checkpoint, power falls; when the job ends or crashes, it drops all at once. The busbar sees the rack-level sum of that square wave, the power shelves must hold the bus voltage through each edge, and the upstream facility, as covered in the grid connection article, sees the multiplied version.
NVIDIA has described a power smoothing design for GB300 NVL72 that attacks this inside the rack. It combines a power cap applied while a job ramps up, electrolytic capacitor energy storage in the power shelves, sized at about 65 joules per GPU, that charges during low demand and discharges during peaks, and a GPU burn mechanism that keeps power from collapsing instantly when a job stops. NVIDIA reports a 30 percent cut in peak grid demand when training Megatron. The lesson for software teams is that rack power is now a controlled quantity, and the controls are reachable from the host.
Software controls: power limits and telemetry
Two controls matter day to day: GPU power limits and power telemetry. nvidia-smi -i 0 -pl 600 sets GPU 0 to a 600 watt limit (root required, and it resets at reboot unless your provisioning reapplies it). nvidia-smi --query-gpu=index,power.draw,power.limit --format=csv -lms 100 samples every 100 milliseconds, and DCGM exposes the same reading as field DCGM_FI_DEV_POWER_USAGE for fleet collection. The script below turns those readings into the number a rack owner cares about: headroom against the usable shelf capacity.
import time
import pynvml # pip install nvidia-ml-py
SHELF_KW, SHELVES, REDUNDANT = 33.0, 6, 1 # hypothetical rack from the text
USABLE_W = (SHELVES - REDUNDANT) * SHELF_KW * 1000
NON_GPU_W = 18_000 # CPUs, NICs, fans, switches: measure it
def sample_host_watts():
total = 0.0
for i in range(pynvml.nvmlDeviceGetCount()):
h = pynvml.nvmlDeviceGetHandleByIndex(i)
total += pynvml.nvmlDeviceGetPowerUsage(h) / 1000.0 # milliwatts to watts
return total
pynvml.nvmlInit()
peak = 0.0
for _ in range(600): # 60 s at 100 ms
peak = max(peak, sample_host_watts())
time.sleep(0.1)
hosts_per_rack = 18 # set from your inventory
rack_peak = peak * hosts_per_rack + NON_GPU_W
print(f"rack peak {rack_peak/1000:.1f} kW, usable {USABLE_W/1000:.1f} kW, "
f"headroom {(USABLE_W - rack_peak)/1000:.1f} kW")Multiplying one host's peak by the host count is deliberately pessimistic, which is right for synchronised training where all hosts peak together. If headroom is negative under the redundancy you need, set a per-GPU limit from the deficit, and measure step time before and after: training throughput usually falls much less than proportionally to a modest cap, because GPUs spend part of each step below their limit anyway.
800 volts DC and the next rack generation
NVIDIA has proposed moving AI datacenters to 800 volts DC, with full-scale production alongside its Kyber rack systems in 2027. In this design, 13.8 kilovolt AC is converted to 800 volts DC at the facility perimeter and carried through the hall to the racks, removing several AC to DC and DC to DC stages. NVIDIA's stated figures are 85 percent more power through the same conductor size, 45 percent less copper than 415 volt AC distribution, up to 5 percent better end-to-end efficiency, and a range from 100 kilowatt to over 1 megawatt racks on the same infrastructure. It also notes that at megawatt scale, 54 volt power shelves would take up to 64U of rack space, more than the compute.
Rerun the worked example at 800 volts and the same 120 kilowatts needs 150 amps, and the same copper would lose under a watt. The cost moves elsewhere: DC has no zero crossing, so breaking a fault current needs DC-rated protection; arc flash and touch-safety rules change; and every connector, procedure and technician qualification has to be revisited. Treat these figures as vendor claims until your facility partners publish their own designs.
Failure modes
- Hot contact. A poorly seated tray or worn clip runs hot and degrades. Detect it with tray input-voltage outliers and periodic thermal scans.
- Redundancy that only exists on paper. Rack load grew past the N+1 or A/B-feed capacity after a hardware refresh or power limit increase; the first shelf or feed loss trips the rack.
- Bus droop on load steps. Synchronous job starts or restarts pull the bus down faster than shelves respond; trays brown out or GPUs throttle. Stagger restarts and use ramp-up power caps.
- Shelf imbalance. Current sharing drifts and one shelf carries more than its share, ageing it faster. Watch per-shelf output current.
- Telemetry blind spots. GPU power is visible to software, but shelf, busbar and tray-input data live in the rack management controller; join them or you debug half the system.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| 48-54 V busbar vs cords | One connector per tray, high current, clean rear | Rack-specific mechanics, contact maintenance |
| More shelves (N+1, A/B) | Survives module or feed loss | Rack space, capital, lower utilisation |
| Centre feed vs end feed | Quarter of the conductor loss | Shelf placement constraints |
| Power capping | Fits more racks per feed, smoother draw | Some step-time loss at the cap |
| Energy storage in shelves | Smaller grid peaks, ride-through | Capacitor ageing, space |
| 800 VDC distribution | Much less copper, fewer conversions | New safety regime, immature ecosystem |
What to do next
- Compute usable rack capacity under your real redundancy policy (module, shelf and feed) and write it next to the rack's measured peak.
- Collect GPU power at 100 ms resolution during a full training step cycle, including checkpoints and restarts, and find the rack-level peak.
- If headroom is negative, set and test a per-GPU power limit; record the step-time cost.
- Pull shelf output current and tray input voltage from the rack controller into the same dashboard as GPU power.
- Alert on any tray whose input voltage sits persistently below the rack median under load, and schedule thermal scans of busbar joints.
- Stagger job restarts after failures so the bus and facility see a ramp, not a step.
- Track 800 VDC designs with your facility team and the cooling plan in liquid cooling, since rack power and heat removal scale together.