Every link in a GPU cluster is a choice between moving electrons down copper and moving photons down glass. The choice looks like procurement trivia until you notice what rides on it: tens of kilowatts of power per row, the weight and airflow of the rack, how often a training job dies from a link flap, and whether a dense rack can be built at all. NVIDIA wired the NVL72 NVLink spine in copper and its CEO said optics and retimers there would have cost about 20 kW per rack; the same racks talk to the rest of the cluster over optics because copper cannot reach the next row.

This page is about that per-hop decision. It explains why copper reach shrinks every time the lane rate doubles, lays out the five cable families, gives a small planner you can run against your own floor plan, and shows what the host, the fabric manager and NCCL see when each kind of link starts to fail. The module internals, forward error correction and the cluster-scale failure budget are covered in Optical Transceivers for AI Networking; this page assumes them and stays on the media decision.

Five things you can plug into a port

There are five things you can plug into a 400G or 800G port, and they differ in what electronics sit inside the plug. A passive direct-attach copper cable (DAC) is twinax wire between two connectors with nothing powered inside: the switch and NIC SerDes drive the wire directly. An active copper cable (ACC) adds a linear redriver that boosts the signal but does not recover the clock. An active electrical cable (AEC) puts a retimer in each plug, which recovers the clock and re-transmits a clean signal, roughly doubling or tripling usable copper reach at the cost of a few watts and some latency. An active optical cable (AOC) seals optics into the plugs with fibre between them. A pluggable transceiver with separate fibre is the most flexible: you choose the reach class, patch through panels and replace either half independently.

FamilyPowered parts in the plugTypical roleWhat fails
Passive DACnonein-rack, NIC or GPU tray to switchconnector seating, crushed cable
ACClinear redriveradjacent rack, shortredriver, connector
AECretimer DSP per endrack to neighbouring rackretimer, thermal
AOClasers, drivers, DSProw-level runs, fixed lengthlaser ageing; whole cable swapped
Transceiver + fibrelasers, drivers, DSPleaf to spine, anywherelaser, dirty connector, bent fibre

Linear-drive (LPO) and co-packaged optics (CPO) cut optical power by removing or shortening the electrical path to the lasers; they change the optical side's power arithmetic but give copper no extra reach.

Why copper reach shrinks as lane rate rises

Which medium carries each hop of a GPU training clusterScale-up: GPU to NVLink switchcopper backplane / cable spineunder ~1 m, no retimersNIC to leaf switchDAC or AEC in-rack, optics if EoR1 m to 30 m: the contested hopLeaf to spinepluggable optics on fibretens to hundreds of metresReach falls as lane rate rises; power per bit and failure modes change with itPassive DAC: 0 W, shortestACC / AEC: redriver or retimer in plugAOC: optics sealed to cableTransceiver + fibre: most flexible, most wattsbar length = practical reach class, not to scale
The three hop classes in a training cluster and the media that usually win each one. The middle hop is where the decision is genuinely open.

Copper reach is set by a loss budget. The SerDes on each end can recover a signal that has lost a fixed amount of energy between transmitter and receiver, measured in decibels at the signal's Nyquist frequency. PAM4 carries two bits per symbol, so a 100 Gb/s lane runs about 53 gigabaud with a Nyquist frequency near 26.5 GHz, and a 200 Gb/s lane roughly doubles both. Cable loss grows with frequency: conductor loss from the skin effect rises with the square root of frequency and dielectric loss rises about linearly. Doubling the lane rate therefore raises the loss per metre by well over the square root of two, while the budget the SerDes can tolerate does not double. Connectors, package traces and the host board consume part of that budget before the cable gets any.

The result is a rule of thumb that has held for three generations: each doubling of lane rate roughly halves passive copper reach. At 100G per lane, passive DAC is commonly sold up to about 2 to 3 metres; at 200G per lane, vendor claims shorten toward 1 to 2 metres, which is why rack-scale systems put the switch in the middle of the rack. Treat these as ranges and check the datasheet of the exact cable with the exact switch and NIC; the host channel matters as much as the cable. Retimers in an AEC reset the budget at each end, so AECs reach several metres where a DAC cannot. Fibre has no such penalty at these distances, so for long hops the decision is not close.

Power, cost and failures per link

Power is where copper wins decisively. A passive DAC draws nothing; the SerDes it connects already exist. An AEC adds a few watts per end. A DSP-based 800G pluggable commonly draws somewhere in the low-to-high teens of watts, and every optical link needs two of them. LPO and CPO pull that down, but not to zero. Multiply by port count and the gap becomes a facility number: an 8,192-GPU cluster with one 800G NIC port per GPU in a two-tier fabric has 8,192 NIC-to-leaf links plus as many leaf-to-spine links, so 32,768 module ends if everything is optical. At an assumed 15 W per end that is about 490 kW, power the facility could otherwise give to several hundred more GPUs.

Failure behaviour is the line item that most surprises software teams. Lasers age; connectors collect dust; a pluggable is an active component with a real annual failure rate. Passive copper has almost no active failure modes, and what does fail, a poorly seated connector or a kinked cable, usually fails at install time where burn-in catches it. Copper's costs are physical: thick twinax is heavy, has a large bend radius, and a bundle of hundreds of cables blocks airflow.

PropertyPassive DACAECOptics (pluggable)
Reach at 100G/laneabout 2-3 mseveral metrestens of m (MM) to km (SM)
Added power per linknonea few W per endtwo modules, often 10-20 W each
Added latencywire onlyretimer, tens of nsDSP, tens to ~100 ns
Active failure modesalmost noneretimerlaser, DSP, dirty connector
Bulk and weighthighmediumlow

Planning media per hop

The decision is per hop, so the planning tool is a loop over links. The script below takes the distance of every link from your floor plan, picks the cheapest family that reaches, and totals power and expected annual replacements. Every number in MEDIA is an assumption to replace with your vendor's datasheet and your own failure history; the structure is the point.

# cable_plan.py: choose a medium per link and total power and expected replacements.
# Replace every figure in MEDIA with datasheet and field data for your exact parts.
from collections import Counter

MEDIA = [  # (name, max_reach_m at this lane rate, watts per link, annual failure prob per link)
    ("dac", 2.0,   0.0, 0.002),
    ("aec", 7.0,   8.0, 0.006),
    ("optic_mm", 50.0, 30.0, 0.02),   # two modules + multimode fibre
    ("optic_sm", 500.0, 34.0, 0.02),  # two modules + single-mode fibre
]

def pick(distance_m, slack_m=0.5):
    need = distance_m + slack_m        # routing through trays always adds length
    for name, reach, watts, afr in MEDIA:
        if need <= reach:
            return name, watts, afr
    raise ValueError(f"no medium reaches {need:.1f} m")

def plan(links):
    """links: iterable of (hop_class, distance_m)."""
    counts, watts, failures = Counter(), 0.0, 0.0
    for hop, d in links:
        name, w, afr = pick(d)
        counts[(hop, name)] += 1
        watts += w
        failures += afr
    return counts, watts, failures

# 1,024 GPUs, 128 nodes, 8 rails; leaves at end of row 4-20 m away, spines 40-120 m.
links = [("nic-leaf", 4 + (i % 32) * 0.5) for i in range(1024)]
links += [("leaf-spine", 40 + (i % 64) * 1.25) for i in range(1024)]
counts, w, f = plan(links)
for k, v in sorted(counts.items()):
    print(k, v)
print(f"media power {w/1000:.1f} kW, expected link replacements per year {f:.0f}")

Run it twice: once with leaves at the end of the row, and once with a leaf in every rack, at the price of more leaf switches.

Worked example: 1,024 GPUs, two layouts

Run the planner on the 1,024-GPU example as written. With leaves at the end of the row, NIC-to-leaf runs are 4 to 20 metres, past DAC reach, so 192 become AECs and 832 become multimode optics; the planner reports about 61 kW of media power and 38 expected link replacements a year. Set every NIC-to-leaf distance to 1.5 metres, a leaf per rack, and all 1,024 become DACs: media power drops to about 34 kW and replacements to about 23, all on the leaf-to-spine half.

The replacement count matters more than it looks. A synchronous job spanning all 1,024 GPUs depends on every link on its path, so each failure is a potential interruption and checkpoint restore. Thirty-eight a year is one every ten days before you count GPU and host failures. The rack-scale NVLink domain described in the NVL72 rack takes the same logic to the limit: it is reported to use 5,184 copper cables in the spine and no optics, because the switches sit within a metre of every GPU tray.

Copper reach is a floor-plan constraint

Copper's reach limit turns into a floor-plan constraint. If the compute hop must be copper, the switch must be within a couple of metres of every port it serves, which pushes you toward top-of-rack or middle-of-rack switches. Liquid-cooled racks make this harder: manifolds and power shelves compete with a dense copper bundle for the same space. Rail-optimised fabrics, in which GPU k of every node attaches to leaf k, naturally place leaves at the end of a row, which forces optics on the hop copper would otherwise win. You cannot choose the media after the floor plan is frozen; the two are one decision.

For InfiniBand and Ethernet port speeds by generation, see InfiniBand NDR and XDR; for how density drives rack layout in the first place, see Power Density in AI Datacenters.

What software sees when a link degrades

Software never sees the word copper, but it sees the consequences. A degrading link first shows rising corrected errors, then symbol errors and link-recovery events, then a flap where the port drops and retrains, and finally a down port. NCCL sees the flap as a stalled collective: it rides through a short retrain or times out and aborts. Optical links add an early warning copper lacks: optical power drifts as a laser ages or a connector gets dirty, visible in module diagnostics before errors appear.

The cheapest defence is to scrape port counters on every node and alarm on the rate of change. These error counters are narrow and stop at their maximum instead of wrapping, so treat a counter pinned at max as an alarm. On Linux they are in sysfs:

# link_watch.py: alarm on error-counter growth per port between two samples.
import pathlib, time

WATCH = ["symbol_error", "link_error_recovery", "link_downed", "port_rcv_errors"]

def sample():
    out = {}
    for port in pathlib.Path("/sys/class/infiniband").glob("*/ports/*"):
        for name in WATCH:
            f = port / "counters" / name
            if f.exists():
                out[(str(port), name)] = int(f.read_text())
    return out

before = sample(); time.sleep(300); after = sample()
for key, v in after.items():
    delta = v - before.get(key, v)
    if delta > 0:
        print(f"{key[0]} {key[1]} +{delta} in 5 min")  # page on link_downed, trend the rest

For Ethernet NICs, ethtool -S gives the equivalent counters and ethtool -m dumps the module's diagnostic page, including optical power on transceivers. Record the media type of every port in your inventory so an alert can say which cable family is failing: a cluster whose failures concentrate on AECs in one row has a thermal problem, not a bad batch of GPUs.

Failure modes

  • Reach on paper, errors in practice. A DAC at its rated limit on a host with a long board trace runs with a thin margin and produces intermittent errors. Burn in with traffic and keep cables well inside rated reach.
  • Bend radius violations. Thick twinax forced around a tight corner degrades slowly. The link passes install checks and fails weeks later.
  • Dirty fibre connectors. The leading cause of optical link trouble. Inspect and clean every connector before mating, including new ones.
  • Airflow blocked by cable bundles. Hot NICs and AEC retimers throttle or error. Watch module temperature alongside error counters.
  • Mixed firmware on retimers and modules. Link training and FEC settings must match on both ends; upgrade them as a fleet, not one cable at a time.
  • Flap storms. A marginal link that retrains every few minutes is worse than a dead one because jobs keep landing on it. Drain the port automatically after a threshold of link_downed events.

Trade-offs

Copper buys zero added power, the lowest latency and almost no active failures, and pays with reach, weight, bend radius and floor-plan freedom. Optics buy reach and layout freedom and pay with watts, cost and a steady trickle of failures. AECs sit in the middle, right when a hop is just past DAC reach. Each lane-rate doubling pulls copper closer to the ASIC, so scale-up domains become dense copper islands while scale-out fabrics go optical. Plan for both.

What to do next

  1. List every hop class in your cluster with its real distance, including tray routing slack, from the floor plan rather than from rack-unit arithmetic.
  2. Run the planner with datasheet reach and power for the exact cables, switch and NIC you are buying, and with your own failure history if you have one.
  3. Compare a leaf-per-rack layout with end-of-row leaves before freezing the floor plan.
  4. Burn in every link with sustained traffic and record its error counters as a baseline.
  5. Deploy a counter collector, alarm on growth, and auto-drain ports that flap.
  6. Record the media family per port so failures can be grouped by cable type and row.
  7. Revisit the decision at each lane-rate step; copper reach roughly halves each time.
Key takeaway: Choose media per hop, not per cluster. Copper costs no power and rarely fails but its reach roughly halves with each lane-rate doubling, so it belongs on hops of a metre or two. Optics reach anywhere at a cost in watts and failures. Decide media and floor plan together, and monitor error counters so a degrading link is drained before it kills a job.