Inside a training cluster, links are short: copper inside the rack, direct-detect optics across the hall, a few hundred metres at most. Once a cluster outgrows one building, or a team wants to train across two sites a campus or a metro apart, the links stretch to tens or hundreds of kilometres. At that distance the cheap trick of switching a laser on and off stops working, and the industry uses coherent optics: transmit information in the amplitude and phase of the light on two polarisations, receive it by mixing with a local laser, and let a powerful DSP undo what the fibre did to the signal.
This article explains coherent transmission from first principles, the DSP chain inside a module, the 400ZR family of pluggables that brought coherent into router ports, how to work a link budget with OSNR, and what a few terabits per second between sites means for distributed training. For short-reach optics inside the building, start with Optical Transceivers for AI Networking.
Why direct detection runs out
A direct-detect receiver is a photodiode: it measures power and nothing else. That throws away phase and polarisation, which are half the information the optical field can carry, and it makes fibre impairments hard to undo. Chromatic dispersion, the fact that different wavelengths travel at slightly different speeds, smears each symbol into its neighbours; standard single-mode fibre spreads roughly 17 ps per nm of bandwidth per km near 1550 nm. Once the receiver has squared the field into power, that smearing cannot be inverted cleanly.
A coherent receiver keeps the field. It mixes the incoming light with a local oscillator laser in a 90-degree optical hybrid, producing in-phase and quadrature components for each of two polarisations, four electrical signals in total, each digitised by a fast ADC. With the full complex field in hand, dispersion becomes a known linear filter that the DSP inverts, and the transmitter can use dense constellations. In 400ZR, the transmitter sends dual-polarisation 16QAM, four bits per symbol per polarisation, at 59.84375 gigabaud: 59.84375e9 * 4 * 2 = 478.75 Gb/s on the line, carrying a 400GbE payload plus FEC and framing overhead.
Inside a coherent module
The diagram follows one bit across the link.
The receive DSP is a pipeline, and each stage maps onto a fibre impairment:
- Chromatic dispersion compensation. A static filter, usually applied in the frequency domain, undoes the dispersion of the configured link length. Every module has a maximum dispersion it can undo; check that the link's dispersion falls inside it.
- Polarisation demultiplexing and PMD. The fibre rotates and mixes the two polarisations, and the rotation drifts as the fibre moves or warms. An adaptive 2x2 equaliser, a small bank of FIR filters updated continuously, unmixes them and also absorbs polarisation mode dispersion.
- Carrier frequency and phase recovery. The local laser is not locked to the transmitter's laser, so the DSP estimates the frequency offset and tracks the phase noise of both lasers symbol by symbol.
- Symbol decisions and soft-decision FEC. The DSP produces soft estimates for each bit, and the FEC decoder corrects errors as long as the pre-FEC bit error rate stays below its threshold. Above the threshold, the output goes from error-free to unusable over a very small change in signal quality: a cliff, not a slope.
That cliff is the operational heart of coherent links. Pre-FEC BER is the leading indicator; post-FEC errors are the trailing one, and by the time they appear the link is already dropping frames. Monitor margin, not errors.
The 400ZR family
| Interface | What it targets | Notes |
|---|---|---|
| 400ZR (OIF) | Point-to-point DCI, about 80-120 km | DP-16QAM at 59.84375 GBd, concatenated FEC; interoperable across vendors; fits QSFP-DD and OSFP router ports |
| ZR+ variants (e.g. OpenZR+) | Longer, multi-span metro and regional links | Stronger FEC and lower-order modes (fewer bits per symbol) trade rate for reach; check each vendor's supported modes |
| 800ZR (OIF) | Single-span amplified DWDM DCI, 80-120 km | Implementation agreement published in 2024; doubles the per-wavelength rate for the same use case |
| Embedded line cards | Long haul and subsea | Higher-performance DSPs with more tuning and power headroom than pluggables |
The important architectural shift came with 400ZR: coherent optics shrank into a pluggable that sits directly in a router or switch port. That removed the separate transponder shelf for many data centre interconnects, and it is why a cross-site link can now be part of the same Ethernet fabric as the rest of the network, carried as parallel wavelengths and balanced with ECMP. The cost is that the module's power and DSP budget is constrained by the pluggable form factor, so reach is shorter than a line card's.
Worked example: an OSNR link budget
In an amplified link, the quantity that decides whether the FEC can cope is the optical signal-to-noise ratio, OSNR: signal power against the amplified spontaneous emission (ASE) noise that every optical amplifier adds. A standard planning approximation, in dB with a 0.1 nm reference bandwidth, is:
import math
def span_loss_db(km, db_per_km=0.22, fixed_db=2.0):
"""Fibre attenuation plus connectors, splices and patch panels."""
return km * db_per_km + fixed_db
def osnr_db(p_ch_dbm, nf_db, span_db, n_spans):
"""Rule-of-thumb OSNR (0.1 nm) for identical amplified spans."""
return 58 + p_ch_dbm - nf_db - span_db - 10 * math.log10(n_spans)
def check(km_per_span, n_spans, required_osnr_db, cd_limit_ps_nm, margin_db=3.0):
span = span_loss_db(km_per_span)
osnr = osnr_db(0.0, 5.5, span, n_spans) # 0 dBm per channel, NF 5.5 dB
cd = 17 * km_per_span * n_spans # ps/nm on standard fibre
ok = osnr >= required_osnr_db + margin_db and cd <= cd_limit_ps_nm
return dict(span_db=round(span, 1), osnr_db=round(osnr, 1), cd_ps_nm=cd,
one_way_us=round(4.9 * km_per_span * n_spans), ok=ok)
# REQ and CD_LIMIT are placeholders: take them from the module datasheet.
REQ, CD_LIMIT = 26.0, 2400
print(check(95, 1, REQ, CD_LIMIT))
# {'span_db': 22.9, 'osnr_db': 29.6, 'cd_ps_nm': 1615, 'one_way_us': 466, 'ok': True}
print(check(100, 3, REQ, CD_LIMIT))
# {'span_db': 24.0, 'osnr_db': 23.7, 'cd_ps_nm': 5100, 'one_way_us': 1470, 'ok': False}Walk through the first case, a 95 km single span between two campuses. The span loses about 22.9 dB; with a booster and pre-amplifier the estimated OSNR is 29.6 dB. If the module needs, say, 26 dB (an assumption here: use your datasheet's figure for the mode you run), there are 3.6 dB of margin, just above the 3 dB set aside for fibre ageing, repairs and amplifier drift. Accumulated dispersion is 1,615 ps/nm, inside the assumed limit. The second case, three 100 km spans, fails on both counts. That is the point where you move to a ZR+ mode with a lower bit rate per wavelength, or to embedded line systems with dispersion and power budgets built for distance.
The rule of thumb ignores fibre nonlinearity, which penalises high launch power, and filtering penalties from multiplexers. It is a first check before an optical engineer runs a proper planning tool, not a replacement for one.
What the link means for training
Coherent links set three numbers that matter to a distributed training job: latency, capacity and failure behaviour.
Latency is dominated by glass. Light in fibre travels at roughly 4.9 microseconds per km, so a 95 km route adds about 0.47 ms one way and nearly 1 ms round trip, before any switching. That is tens of times the latency of an intra-building fabric. Collectives that are latency-bound, such as small all-reduces in tensor parallelism, cannot span that gap usefully. Data parallelism with large gradient buckets can.
Capacity comes from wavelengths. On a 75 GHz DWDM grid the C-band holds 64 channels, so 64 wavelengths of 400 Gb/s carry about 25.6 Tb/s per fibre pair. Few deployments light all of that for one job. Suppose eight 400ZR waves, 3.2 Tb/s or 400 GB/s, connect two sites, and each site holds one full data-parallel replica group of a model with 70 billion parameters and bf16 gradients, 140 GB per step. A hierarchical all-reduce, first within each site and then one exchange between sites, must move roughly the full 140 GB across the link in each direction: about 0.35 s per step at line rate, more in practice. Whether that is acceptable depends on step time and on overlap with backward compute, which is the arithmetic in All-Reduce, in depth.
Teams that cannot afford that cost change the algorithm rather than the optics: synchronise across sites less often (local-SGD style methods such as DiLoCo exchange each replica's parameter change every few hundred steps and apply the average with an outer momentum optimiser), compress cross-site traffic, or place pipeline stages so only activations of one boundary cross the link.
Failure behaviour differs from inside the hall. A fibre cut takes out every wavelength on that pair at once, so plan two physically diverse routes, and make sure the training framework can survive losing cross-site connectivity for seconds while traffic reroutes, rather than timing out the job. RDMA traffic over these links needs the same congestion control care as inside the fabric, described in RoCE v2, in depth, with buffers sized for the much larger bandwidth-delay product.
Telemetry and operations
Coherent modules report rich diagnostics through the Common Management Interface Specification (CMIS) used by QSFP-DD and OSFP modules and its coherent extension, C-CMIS, including pre-FEC BER, an OSNR estimate, residual dispersion and differential group delay. How you read them depends on your network operating system; the logic on top is the same:
import math
def link_health(sample, baseline):
"""sample/baseline: dicts of module diagnostics for one wavelength.
read the values through your NOS's transceiver telemetry; names here are generic."""
issues = []
# FEC margin: how far pre-FEC BER sits below the threshold, in decades
margin = math.log10(sample["fec_threshold_ber"]) - math.log10(sample["pre_fec_ber"])
if margin < 0.5:
issues.append("pre-FEC BER within half a decade of the FEC cliff")
if baseline["osnr_db"] - sample["osnr_db"] > 1.5:
issues.append("OSNR down >1.5 dB from commissioning: amp or fibre degradation")
if abs(sample["cd_ps_nm"] - baseline["cd_ps_nm"]) > 100:
issues.append("dispersion changed: traffic may be on a different fibre route")
if sample["post_fec_uncorrected"] > 0:
issues.append("uncorrected frames: link is already dropping data")
return issuesRecord a baseline for every wavelength at commissioning, trend it, and alert on drift long before the cliff. A dispersion change is a useful forensic signal: it often means a protection switch moved traffic onto a different path, which also changes latency for the training job.
Failure modes
| Failure | Signature | Response |
|---|---|---|
| Fibre cut | All wavelengths on a pair lose signal together | Diverse routes; fast reroute; job tolerates seconds of loss |
| Amplifier degradation | OSNR falls on every channel through that site | Trend OSNR; replace before margin goes |
| Dirty or damaged connector | Extra loss on one span, sudden drop after maintenance | Inspect and clean; re-measure loss |
| Laser frequency drift | Errors on one channel, adjacent channel interference | Module replacement; check grid configuration |
| Mode or FEC mismatch | Link never comes up after a change | Configuration management on both ends |
| Thermal stress in the port | Module temperature alarms; errors under load | Airflow; high-power modules only where rated |
Pluggable coherent modules draw much more power than short-reach optics, and a router's power and cooling budget per port limits how many it can host. Check the platform's supported optics list before buying modules.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| 400ZR pluggables in routers | Simplest DCI; no transponder shelf | Single-span reach; port power budget |
| ZR+ lower-rate modes | Longer reach on the same hardware | Fewer bits per wavelength |
| Embedded line systems | Reach and spectral efficiency | Separate optical layer to operate |
| Training across sites | Use power and space in two places | Algorithm changes; WAN-scale failure handling |
What to do next
- Get route lengths, span losses and fibre type for every candidate path, including the diverse backup route.
- Run the OSNR and dispersion check above with the real datasheet requirements for the module and mode you intend to use, and keep at least 3 dB of margin.
- Compute the bandwidth-delay product and per-step cross-site traffic for your parallelism plan before choosing the number of wavelengths.
- Decide what the training job does when the cross-site link flaps, and test it by pulling a wavelength.
- Baseline pre-FEC BER, OSNR and dispersion per wavelength at commissioning and alert on drift.
- Place cross-site links in the fabric design described in GPU Pod Network, so routing and congestion control treat them as the slow, wide links they are.