A large training cluster has tens of thousands of optical transceivers: small pluggable modules that turn the electrical signal from a switch or network card into light and back. They are easy to ignore because they are commodity parts, and impossible to ignore once you run a big job, because a synchronous training step waits for its slowest link, and with this many modules one of them is degrading somewhere at almost any moment.

This article explains what is inside a module and why, how reach classes and form factors map to places in a fabric, what linear and co-packaged optics change, how forward error correction hides errors until it suddenly cannot, and how a flaky module turns into a stalled collective in your training job. It ends with a failure budget for a 16,384-GPU cluster and the telemetry that catches bad optics before they crash a run. For the fabric these modules plug into, see the GPU pod network and rail-aligned topology.

Where optics sit in a training fabric

In a typical scale-out fabric, each GPU has its own network interface, 400 Gb/s per GPU being common in current InfiniBand and Ethernet designs. That NIC connects to a leaf switch, leaves connect to spines, and in large clusters spines connect to a core tier. Inside a rack, short links can be passive copper cables or active electrical cables. Anything that leaves the rack, which in rail-optimised designs often includes the GPU-to-leaf link, is optical, because copper at 100 Gb/s per lane only reaches a few metres.

Every optical link has a transceiver at each end. A non-blocking fat tree with three switching tiers needs about three links per GPU, so about six transceiver ends per GPU if every tier is optical. Twin-port modules, which carry two 400 Gb/s links in one 800 Gb/s cage, reduce the number of physical parts on the switch side but not the number of lasers and optical lanes that can fail.

Inside a module: the signal path

Follow one lane through the module. The host SerDes sends a PAM4 electrical signal: four amplitude levels carrying two bits per symbol, so a 100 Gb/s lane runs at about 53 gigabaud. Over the host board and connector that signal degrades, so a DSP retimes and equalises it. A driver then modulates a light source: a directly modulated VCSEL for short multimode links, or an externally modulated laser or silicon photonic modulator with a continuous-wave laser for single-mode links. On the receive side a photodiode converts light to current, a transimpedance amplifier to voltage, and the DSP recovers clean symbols for the host.

A small microcontroller manages the module and exposes digital optical monitoring through the Common Management Interface Specification (CMIS) used by QSFP-DD and OSFP modules: temperature, supply voltage, laser bias current, and transmit and receive optical power per lane, together with the module's own alarm and warning thresholds. This telemetry is the main way software sees optics.

Inside a DSP-based pluggable module: one electrical lane in, one optical lane outHost SerDesswitch or NICDSPretime, equaliseDriverPAM4 levelsLaser + modulatorEML, SiPh or VCSELFibreSMF or MMF, MPO or duplexFibre from far endPhotodiodelight to currentTIAcurrent to voltageDSPrecover, retimeHost SerDesFEC decodeMicrocontroller + CMISDOM: temp, bias, Tx/Rx powerLinear (LPO) designs remove the DSP boxes;co-packaged optics move the optics next to the switch die.
The DSP is a large share of module power; removing or relocating it is what linear and co-packaged optics are about.

Reach classes and form factors

ClassFibreTypical reachLight sourceWhere in an AI fabric
VR / SR (multimode)OM4 MMFabout 50 m (VR) or 100 m (SR) per IEEE 802.3db at 100G per laneVCSELwithin a row; cheapest optics
DR (parallel single-mode)SMF, MPO500 mEML or SiPhleaf to spine across a hall; breakout to 4 x 100G
FR (CWDM)SMF, duplex2 km4 wavelengthsbetween halls; fewer fibres
LRSMF, duplex10 kmEMLbetween buildings on a campus

Names such as 800G 2xDR4 or 2xFR4 describe two independent four-lane links in one module. Form factors are mechanical and thermal standards: QSFP-DD and OSFP both carry eight electrical lanes, with OSFP being larger and better at dissipating heat, which is why many 800 Gb/s AI switches use it. The next step doubles the lane rate to 200 Gb/s, giving 1.6 Tb/s modules from eight lanes, with interfaces being standardised in IEEE 802.3dj. Match reach class to the longest link you will actually cable, plus patch panels, not to the floor plan.

DSP, linear and co-packaged optics

DSP-retimed pluggables are the default: the module cleans the signal in both directions, so host and module can be qualified independently. The cost is power and latency. Vendor material puts an 800 Gb/s DSP module at very roughly 12 to 16 W.

Linear pluggable optics (LPO) remove the DSP and rely on the host SerDes to equalise the whole channel, including the optics. Vendors quote roughly half the module power, around 5 to 8 W at 800 Gb/s, and lower latency. The catch is interoperability: the link only works if a particular host SerDes and module are tuned for each other, so qualification moves from module to host-module pair. Half-retimed designs, sometimes called linear receive optics, keep a DSP on one direction as a compromise.

Co-packaged optics (CPO) place the optical engines on the switch package beside the switching silicon, shrinking the electrical path to millimetres and leaving only fibre at the faceplate. In March 2025 NVIDIA announced Spectrum-X Photonics Ethernet and Quantum-X Photonics InfiniBand switches built this way, and claimed 3.5x better power efficiency and 10x better resiliency at scale than pluggable designs; those are vendor claims, not independent measurements. The operational trade-off is serviceability: a failed pluggable is swapped in minutes, while a failed co-packaged engine may mean a switch replacement, which is why CPO designs typically keep the lasers in separate, replaceable modules.

FEC and the error cliff

At 100 Gb/s per lane, raw bit errors are expected, and forward error correction removes them. Ethernet and InfiniBand at this rate use a Reed-Solomon code, RS(544, 514), often called KP4: each codeword carries 514 data symbols and 30 parity symbols of 10 bits, and corrects up to 15 bad symbols per codeword. The commonly quoted limit for KP4 links at these rates is a pre-FEC bit error rate of about 2.4 x 10-4; below that, the post-FEC error rate is effectively zero.

This makes FEC a cliff. A degrading module, with a dirty connector, an ageing laser or a cracked fibre, shows a rising pre-FEC error rate and rising corrected-block counts for days while traffic looks perfect. Then uncorrectable blocks appear, packets are dropped, transport retries begin, and the link may flap. Monitoring post-FEC errors alone tells you about the problem after the cliff; monitoring pre-FEC errors and corrected-block rates tells you before it. On Linux, ethtool -I --show-fec <iface> reports corrected and uncorrectable block counters per lane where the driver supports them, and ethtool -m <iface> dumps the module's monitoring data. On NVIDIA adapters, mlxlink -d <device> -c reports Raw Physical BER (pre-FEC) and Effective Physical BER (post-FEC).

Worked example: a failure budget for 16,384 GPUs

How many optics failures should a large training job expect? The answer depends on the annual failure rate (AFR) of your modules, which you should take from your own fleet data or your vendor; the 0.5% below is an assumption for illustration.

gpus = 16_384
tiers = 3                    # GPU-leaf, leaf-spine, spine-core; all optical (assumption)
links = gpus * tiers         # non-blocking fat tree: about one link per GPU per tier
ends = 2 * links             # a transceiver lane group at each end
afr = 0.005                  # ASSUMED annual failure rate per transceiver end

per_day = ends * afr / 365
lost_min = 15 + 15           # half of a 30-minute checkpoint interval + 15 min restart
print(links, ends, round(per_day, 2), f"{per_day * lost_min / 1440:.1%}")
# 49152 98304 1.35 2.8%

Under these assumptions there are 98,304 transceiver ends and about 1.35 hard optics failures a day. If each one crashes the job, costing half a checkpoint interval of lost work plus a restart, the cluster loses about 2.8% of its time to optics alone. If the GPU-to-leaf links are copper inside the rack, the count falls by a third, to about 0.9 failures a day. And hard failures are the easy case: flapping links and slowly degrading modules cause more incidents than dead ones, and they do not show up in an AFR figure. The lever is not a better AFR, which you do not control, but catching degradation early and draining traffic before a failure lands mid-step. The GPU hardware faults article covers the same reasoning for accelerators.

Telemetry that catches bad optics early

Poll module monitoring and FEC counters on every port every minute or so, and alert on trends rather than only on the module's own alarm thresholds, which fire late. A minimal per-port evaluator:

def assess_port(prev, cur, thresholds, window_s):
    """prev/cur: dicts of DOM and FEC counters; thresholds from the module (CMIS)."""
    findings = []
    for lane, rx in enumerate(cur["rx_power_dbm"]):
        margin = rx - thresholds["rx_power_low_warn_dbm"]
        if margin < 1.0:
            findings.append(("rx_margin_low", lane, round(margin, 2)))
        drop = prev["rx_power_dbm"][lane] - rx
        if drop > 1.0:                      # a sudden dB drop: connector or fibre event
            findings.append(("rx_power_step", lane, round(drop, 2)))
    corr_rate = (cur["fec_corrected"] - prev["fec_corrected"]) / window_s
    baseline = max(cur["fec_corrected_baseline_per_s"], 1e-3)  # rolling median, per port
    if corr_rate > 10 * baseline:
        findings.append(("fec_corrected_spike", None, corr_rate))
    if cur["fec_uncorrectable"] > prev["fec_uncorrectable"]:
        findings.append(("fec_uncorrectable", None,
                         cur["fec_uncorrectable"] - prev["fec_uncorrectable"]))
    if cur["link_down_count"] > prev["link_down_count"]:
        findings.append(("link_flap", None, None))
    return findings   # route to: drain port, schedule clean/swap, or page

The thresholds in this sketch, 1 dB of margin and a tenfold jump in corrected blocks against the port's baseline, are starting points to tune on your fleet. Join findings to the job scheduler: a port with rising corrected blocks should be drained, or its nodes cordoned, at the next checkpoint rather than left to fail mid-step. Most single-mode problems are dirty connectors, so cleaning and re-seating are the first remedy before swapping the module.

How a link fault reaches the training job

From the training job's point of view an optics problem arrives as slow or failed collectives. A link that drops packets causes InfiniBand or RoCE transport retries; the affected ring or tree in NCCL slows down, and because data-parallel steps synchronise, every GPU waits. If retries exhaust, the queue pair errors out, NCCL reports a failure, and the process group aborts. NCCL exposes the transport retry behaviour through NCCL_IB_TIMEOUT and NCCL_IB_RETRY_CNT (check the defaults for your NCCL version); raising them can ride out a brief flap at the cost of slower detection of a truly dead link. PyTorch's process-group timeout, set through the timeout argument of init_process_group, bounds how long a rank waits before the job is declared hung.

The practical pattern is to correlate step-time outliers with port telemetry. If one step in a thousand is twice as slow and the slow steps line up with corrected-block bursts on a single port, you have found the module before it fails. The InfiniBand and NVLink article explains how NCCL chooses which links each collective uses.

Failure modes

  • Dirty or damaged connectors. The most common cause of low receive power. Inspect with a scope, clean, and re-seat before swapping parts.
  • Thermal stress. Modules near hot air exhaust run close to their temperature limit and age faster. Watch module temperature by rack position.
  • Mismatched pairs. LPO modules on an unqualified host, or mixed firmware, give links that come up and then misbehave under load.
  • Wrong reach class. A link at the edge of its optical budget passes acceptance and fails as connectors age. Leave margin.
  • Flapping. A link that repeatedly drops and recovers can be worse than a dead one, because routing keeps sending traffic to it. Dampen and drain.
  • Alerting only on post-FEC errors. By the time they appear, the job is already suffering.

Trade-offs

Multimode short-reach optics are cheapest but limit cable lengths and so constrain floor plans. Single-mode DR and FR optics cost more but let you place switches freely. DSP modules trade power for easy interoperability; LPO halves module power but ties you to qualified host-module pairs; co-packaged optics promise the largest power savings and the shortest electrical paths, at the price of serviceability and vendor lock-in. Copper is cheaper, lower-power and more reliable than any optics, but only within a rack, which is why dense, liquid-cooled racks that keep more traffic on copper also reduce optics failures. At cluster scale, power saved per module multiplies across tens of thousands of ports, while every additional optical hop adds failure exposure, so topology and optics choices should be made together.

What to do next

  1. Count the transceiver ends in your fabric by tier and by reach class.
  2. Get a real AFR from your fleet data or vendor and redo the failure budget above.
  3. Collect CMIS monitoring and FEC counters on every port at minute granularity.
  4. Alert on pre-FEC trends and receive-power margin, not only on post-FEC errors.
  5. Wire port findings into the scheduler so degrading links are drained at checkpoints.
  6. Correlate training step-time outliers with port telemetry to find slow links.
  7. Stock spares by module type and keep connector inspection and cleaning kits in every hall.
  8. Qualify LPO or CPO as host-module systems before adopting them, and measure power yourself.
Key takeaway: Optical transceivers turn switch and NIC SerDes lanes into light, and a large cluster has around six optical ends per GPU when every tier is optical. FEC hides degradation until a cliff, so watch pre-FEC errors, corrected blocks and receive-power margin, drain degrading ports at checkpoints, and treat linear and co-packaged optics as power savings that you pay for in qualification and serviceability.