Inside a training cluster, links are short: copper inside the rack, direct-detect optics across the hall, a few hundred metres at most. Once a cluster outgrows one building, or a team wants to train across two sites a campus or a metro apart, the links stretch to tens or hundreds of kilometres. At that distance the cheap trick of switching a laser on and off stops working, and the industry uses coherent optics: transmit information in the amplitude and phase of the light on two polarisations, receive it by mixing with a local laser, and let a powerful DSP undo what the fibre did to the signal.

This article explains coherent transmission from first principles, the DSP chain inside a module, the 400ZR family of pluggables that brought coherent into router ports, how to work a link budget with OSNR, and what a few terabits per second between sites means for distributed training. For short-reach optics inside the building, start with Optical Transceivers for AI Networking.

Why direct detection runs out

A direct-detect receiver is a photodiode: it measures power and nothing else. That throws away phase and polarisation, which are half the information the optical field can carry, and it makes fibre impairments hard to undo. Chromatic dispersion, the fact that different wavelengths travel at slightly different speeds, smears each symbol into its neighbours; standard single-mode fibre spreads roughly 17 ps per nm of bandwidth per km near 1550 nm. Once the receiver has squared the field into power, that smearing cannot be inverted cleanly.

A coherent receiver keeps the field. It mixes the incoming light with a local oscillator laser in a 90-degree optical hybrid, producing in-phase and quadrature components for each of two polarisations, four electrical signals in total, each digitised by a fast ADC. With the full complex field in hand, dispersion becomes a known linear filter that the DSP inverts, and the transmitter can use dense constellations. In 400ZR, the transmitter sends dual-polarisation 16QAM, four bits per symbol per polarisation, at 59.84375 gigabaud: 59.84375e9 * 4 * 2 = 478.75 Gb/s on the line, carrying a 400GbE payload plus FEC and framing overhead.

Inside a coherent module

The diagram follows one bit across the link.

Coherent link: where the bits become a field and backHost ASIC400GbE lanesTx DSPFEC, framing, shapingDACs + IQ modulatorX and Y polarisationTunable laserDWDM channelMux - booster amp - fibre span (~0.2 dB/km) - inline amps - pre-amp - demuxadds loss, ASE noise, chromatic dispersion, PMD, nonlinearityOptical hybridmix with local laserPhotodiodes + ADCsI and Q per polarisationRx DSPCD, 2x2 MIMO, carrierSoft-decision FECcorrects pre-FEC errorsTelemetry: pre-FEC BER, OSNR, CD, DGDread via module management (CMIS)
The transmit DSP encodes and shapes; the fibre adds noise and distortion; the receive DSP removes the distortion and FEC removes the residual errors.

The receive DSP is a pipeline, and each stage maps onto a fibre impairment:

  1. Chromatic dispersion compensation. A static filter, usually applied in the frequency domain, undoes the dispersion of the configured link length. Every module has a maximum dispersion it can undo; check that the link's dispersion falls inside it.
  2. Polarisation demultiplexing and PMD. The fibre rotates and mixes the two polarisations, and the rotation drifts as the fibre moves or warms. An adaptive 2x2 equaliser, a small bank of FIR filters updated continuously, unmixes them and also absorbs polarisation mode dispersion.
  3. Carrier frequency and phase recovery. The local laser is not locked to the transmitter's laser, so the DSP estimates the frequency offset and tracks the phase noise of both lasers symbol by symbol.
  4. Symbol decisions and soft-decision FEC. The DSP produces soft estimates for each bit, and the FEC decoder corrects errors as long as the pre-FEC bit error rate stays below its threshold. Above the threshold, the output goes from error-free to unusable over a very small change in signal quality: a cliff, not a slope.

That cliff is the operational heart of coherent links. Pre-FEC BER is the leading indicator; post-FEC errors are the trailing one, and by the time they appear the link is already dropping frames. Monitor margin, not errors.

The 400ZR family

InterfaceWhat it targetsNotes
400ZR (OIF)Point-to-point DCI, about 80-120 kmDP-16QAM at 59.84375 GBd, concatenated FEC; interoperable across vendors; fits QSFP-DD and OSFP router ports
ZR+ variants (e.g. OpenZR+)Longer, multi-span metro and regional linksStronger FEC and lower-order modes (fewer bits per symbol) trade rate for reach; check each vendor's supported modes
800ZR (OIF)Single-span amplified DWDM DCI, 80-120 kmImplementation agreement published in 2024; doubles the per-wavelength rate for the same use case
Embedded line cardsLong haul and subseaHigher-performance DSPs with more tuning and power headroom than pluggables

The important architectural shift came with 400ZR: coherent optics shrank into a pluggable that sits directly in a router or switch port. That removed the separate transponder shelf for many data centre interconnects, and it is why a cross-site link can now be part of the same Ethernet fabric as the rest of the network, carried as parallel wavelengths and balanced with ECMP. The cost is that the module's power and DSP budget is constrained by the pluggable form factor, so reach is shorter than a line card's.

Worked example: an OSNR link budget

In an amplified link, the quantity that decides whether the FEC can cope is the optical signal-to-noise ratio, OSNR: signal power against the amplified spontaneous emission (ASE) noise that every optical amplifier adds. A standard planning approximation, in dB with a 0.1 nm reference bandwidth, is:

import math

def span_loss_db(km, db_per_km=0.22, fixed_db=2.0):
    """Fibre attenuation plus connectors, splices and patch panels."""
    return km * db_per_km + fixed_db

def osnr_db(p_ch_dbm, nf_db, span_db, n_spans):
    """Rule-of-thumb OSNR (0.1 nm) for identical amplified spans."""
    return 58 + p_ch_dbm - nf_db - span_db - 10 * math.log10(n_spans)

def check(km_per_span, n_spans, required_osnr_db, cd_limit_ps_nm, margin_db=3.0):
    span = span_loss_db(km_per_span)
    osnr = osnr_db(0.0, 5.5, span, n_spans)          # 0 dBm per channel, NF 5.5 dB
    cd = 17 * km_per_span * n_spans                   # ps/nm on standard fibre
    ok = osnr >= required_osnr_db + margin_db and cd <= cd_limit_ps_nm
    return dict(span_db=round(span, 1), osnr_db=round(osnr, 1), cd_ps_nm=cd,
                one_way_us=round(4.9 * km_per_span * n_spans), ok=ok)

# REQ and CD_LIMIT are placeholders: take them from the module datasheet.
REQ, CD_LIMIT = 26.0, 2400
print(check(95, 1, REQ, CD_LIMIT))
# {'span_db': 22.9, 'osnr_db': 29.6, 'cd_ps_nm': 1615, 'one_way_us': 466, 'ok': True}
print(check(100, 3, REQ, CD_LIMIT))
# {'span_db': 24.0, 'osnr_db': 23.7, 'cd_ps_nm': 5100, 'one_way_us': 1470, 'ok': False}

Walk through the first case, a 95 km single span between two campuses. The span loses about 22.9 dB; with a booster and pre-amplifier the estimated OSNR is 29.6 dB. If the module needs, say, 26 dB (an assumption here: use your datasheet's figure for the mode you run), there are 3.6 dB of margin, just above the 3 dB set aside for fibre ageing, repairs and amplifier drift. Accumulated dispersion is 1,615 ps/nm, inside the assumed limit. The second case, three 100 km spans, fails on both counts. That is the point where you move to a ZR+ mode with a lower bit rate per wavelength, or to embedded line systems with dispersion and power budgets built for distance.

The rule of thumb ignores fibre nonlinearity, which penalises high launch power, and filtering penalties from multiplexers. It is a first check before an optical engineer runs a proper planning tool, not a replacement for one.

What the link means for training

Coherent links set three numbers that matter to a distributed training job: latency, capacity and failure behaviour.

Latency is dominated by glass. Light in fibre travels at roughly 4.9 microseconds per km, so a 95 km route adds about 0.47 ms one way and nearly 1 ms round trip, before any switching. That is tens of times the latency of an intra-building fabric. Collectives that are latency-bound, such as small all-reduces in tensor parallelism, cannot span that gap usefully. Data parallelism with large gradient buckets can.

Capacity comes from wavelengths. On a 75 GHz DWDM grid the C-band holds 64 channels, so 64 wavelengths of 400 Gb/s carry about 25.6 Tb/s per fibre pair. Few deployments light all of that for one job. Suppose eight 400ZR waves, 3.2 Tb/s or 400 GB/s, connect two sites, and each site holds one full data-parallel replica group of a model with 70 billion parameters and bf16 gradients, 140 GB per step. A hierarchical all-reduce, first within each site and then one exchange between sites, must move roughly the full 140 GB across the link in each direction: about 0.35 s per step at line rate, more in practice. Whether that is acceptable depends on step time and on overlap with backward compute, which is the arithmetic in All-Reduce, in depth.

Teams that cannot afford that cost change the algorithm rather than the optics: synchronise across sites less often (local-SGD style methods such as DiLoCo exchange each replica's parameter change every few hundred steps and apply the average with an outer momentum optimiser), compress cross-site traffic, or place pipeline stages so only activations of one boundary cross the link.

Failure behaviour differs from inside the hall. A fibre cut takes out every wavelength on that pair at once, so plan two physically diverse routes, and make sure the training framework can survive losing cross-site connectivity for seconds while traffic reroutes, rather than timing out the job. RDMA traffic over these links needs the same congestion control care as inside the fabric, described in RoCE v2, in depth, with buffers sized for the much larger bandwidth-delay product.

Telemetry and operations

Coherent modules report rich diagnostics through the Common Management Interface Specification (CMIS) used by QSFP-DD and OSFP modules and its coherent extension, C-CMIS, including pre-FEC BER, an OSNR estimate, residual dispersion and differential group delay. How you read them depends on your network operating system; the logic on top is the same:

import math

def link_health(sample, baseline):
    """sample/baseline: dicts of module diagnostics for one wavelength.
    read the values through your NOS's transceiver telemetry; names here are generic."""
    issues = []
    # FEC margin: how far pre-FEC BER sits below the threshold, in decades
    margin = math.log10(sample["fec_threshold_ber"]) - math.log10(sample["pre_fec_ber"])
    if margin < 0.5:
        issues.append("pre-FEC BER within half a decade of the FEC cliff")
    if baseline["osnr_db"] - sample["osnr_db"] > 1.5:
        issues.append("OSNR down >1.5 dB from commissioning: amp or fibre degradation")
    if abs(sample["cd_ps_nm"] - baseline["cd_ps_nm"]) > 100:
        issues.append("dispersion changed: traffic may be on a different fibre route")
    if sample["post_fec_uncorrected"] > 0:
        issues.append("uncorrected frames: link is already dropping data")
    return issues

Record a baseline for every wavelength at commissioning, trend it, and alert on drift long before the cliff. A dispersion change is a useful forensic signal: it often means a protection switch moved traffic onto a different path, which also changes latency for the training job.

Failure modes

FailureSignatureResponse
Fibre cutAll wavelengths on a pair lose signal togetherDiverse routes; fast reroute; job tolerates seconds of loss
Amplifier degradationOSNR falls on every channel through that siteTrend OSNR; replace before margin goes
Dirty or damaged connectorExtra loss on one span, sudden drop after maintenanceInspect and clean; re-measure loss
Laser frequency driftErrors on one channel, adjacent channel interferenceModule replacement; check grid configuration
Mode or FEC mismatchLink never comes up after a changeConfiguration management on both ends
Thermal stress in the portModule temperature alarms; errors under loadAirflow; high-power modules only where rated

Pluggable coherent modules draw much more power than short-reach optics, and a router's power and cooling budget per port limits how many it can host. Check the platform's supported optics list before buying modules.

Trade-offs

ChoiceGainsCosts
400ZR pluggables in routersSimplest DCI; no transponder shelfSingle-span reach; port power budget
ZR+ lower-rate modesLonger reach on the same hardwareFewer bits per wavelength
Embedded line systemsReach and spectral efficiencySeparate optical layer to operate
Training across sitesUse power and space in two placesAlgorithm changes; WAN-scale failure handling

What to do next

  1. Get route lengths, span losses and fibre type for every candidate path, including the diverse backup route.
  2. Run the OSNR and dispersion check above with the real datasheet requirements for the module and mode you intend to use, and keep at least 3 dB of margin.
  3. Compute the bandwidth-delay product and per-step cross-site traffic for your parallelism plan before choosing the number of wavelengths.
  4. Decide what the training job does when the cross-site link flaps, and test it by pulling a wavelength.
  5. Baseline pre-FEC BER, OSNR and dispersion per wavelength at commissioning and alert on drift.
  6. Place cross-site links in the fabric design described in GPU Pod Network, so routing and congestion control treat them as the slow, wide links they are.
Key takeaway: Coherent optics recover the full optical field and let a DSP undo the fibre, which is what makes 400 Gb/s per wavelength over tens of kilometres possible in a router port. Plan links with OSNR and dispersion margin, monitor pre-FEC BER rather than errors, and design training so the slower, wider cross-site link carries infrequent, large transfers.