Every GPU spec sheet lists a PCIe generation, and the step from Gen4 to Gen5 doubles the headline number. Whether that doubling changes anything for you depends on how much of your job crosses the host link, and on whether your link actually trains at the generation on the box. Plenty of Gen5 GPUs run at Gen4 or x8 because of a riser, a BIOS setting or a slot wired with half its lanes, and nothing in the training loop tells you.

This page covers the generation boundary only: what changed between Gen4 and Gen5 and what Gen6 changes next, how a link negotiates down, how to prove what yours is doing, and which training and serving workloads feel the difference. The basics of the host path (pinned memory, copy engines, peer-to-peer and NUMA) are in PCIe host interconnect architecture; read that first if those terms are new.

What a generation changes

A PCIe generation is mostly a signalling rate. Gen3 runs each lane at 8 GT/s, Gen4 at 16 GT/s and Gen5 at 32 GT/s. From Gen3 through Gen5 the line code is the same: 128b/130b, two framing bits for every 128 data bits, so about 1.5% of raw bits are overhead. That makes the per-direction payload bandwidth of a link easy to derive:

bytes_per_s = GT_per_s * lanes * (128 / 130) / 8

Gen4 x16: 16e9 * 16 * 128/130 / 8 = 31.5 GB/s per direction
Gen5 x16: 32e9 * 16 * 128/130 / 8 = 63.0 GB/s per direction

Vendors usually quote 64 GB/s for Gen4 x16 and 128 GB/s for Gen5 x16. Those figures add both directions together. A host-to-device copy only uses one direction, so the ceiling for loading weights is about 31.5 or 63 GB/s, not 64 or 128. The two directions are independent, which is why overlapping an upload with a download is worth doing.

GenerationPer-lane rateEncodingx16, per directionSignalling
Gen38 GT/s128b/130babout 15.8 GB/sNRZ
Gen416 GT/s128b/130babout 31.5 GB/sNRZ
Gen532 GT/s128b/130babout 63.0 GB/sNRZ
Gen664 GT/sFlit mode with FECroughly double Gen5PAM4

Below the encoding there is protocol overhead. Data moves in transaction layer packets, each with a header, sequence number and CRC around a payload capped by the negotiated maximum payload size. Flow-control and acknowledgement packets also share the link. The result is that a real copy reaches noticeably less than the derived figure, by an amount that depends on the platform. Measure it rather than assume it; the script later on this page does exactly that.

Which side sets the generation

The GPU is one end of the link, and the slot is the other. Among NVIDIA's data-centre PCIe cards, the A100 and the L40S are Gen4 x16, while the H100 PCIe and the H200 NVL are Gen5 x16. A Gen5 card in a Gen4 server runs at Gen4. A Gen5 card in a Gen5 server can still run at Gen4 if anything on the path cannot hold the faster signal.

Board-style systems (SXM modules on an HGX baseboard) still use PCIe for host traffic and NICs, even though GPU-to-GPU traffic goes over NVLink. The host side of those boards is where generation mismatches hide. NVIDIA's ConnectX-8 SuperNIC, used in HGX B300 and GB300 NVL72 systems, is an example: it integrates 48 lanes of Gen6 and a Gen6 switch, and NVIDIA says its GPU-to-NIC path delivers 50 GB/s whether the GPU or the host supports Gen5 or Gen6. Each hop negotiates its own generation.

A link trains to the slowest hop on its pathGPUGen5 x16 capableRiser / cableloss, crosstalkRetimerre-drives the signalPCIe switchshared upstream x16CPU root complexGen4 or Gen5 portsNUMAHost DRAMpinned buffersWhat can go wrong at each hopSpeed fallbackGen5 to Gen4 or Gen3 after training errorsWidth dropx16 to x8 from a lane fault or slot wiringShared uplinktwo GPUs behind one x16 halve eachCompare gpucurrent with gpumax and hostmax under load, and width with its max
Figure 1. The negotiated generation and width are properties of each link, not of the GPU. The slowest hop on the path you use sets your bandwidth.

Signal integrity, retimers and link training

Doubling the rate halves the time each bit occupies the wire and roughly doubles the frequency the channel must carry. Board traces, connectors and cables lose more signal at higher frequency, so a channel that is comfortable at Gen4 can be marginal at Gen5. The practical consequences for a GPU server are shorter traces, better board material, and retimers: chips that recover the signal and transmit it fresh, effectively splitting one long channel into two short ones. Riser cards and cabled GPU trays are the usual place where a retimer is needed and sometimes missing.

When the link powers up, both ends run a training state machine (the LTSSM) that agrees on width and speed and tunes the transmitter equalisation. If equalisation at the top speed fails, or errors keep forcing retraining, the link settles at a lower speed. It still works, and the operating system shows a healthy device. That is the core operational risk of Gen5: degradation is silent and functional.

Errors that do not stop the link show up as replays. The data link layer resends any packet whose CRC fails, which costs bandwidth and latency but loses no data. A rising replay count under load means a marginal channel, and it is the best early warning that a link will fall back on the next reboot.

Proving what your link negotiated

Check three things on every new node and after every hardware change: generation, width and error counters. nvidia-smi reports what the GPU sees:

nvidia-smi --query-gpu=index,pci.bus_id,pcie.link.gen.gpucurrent,pcie.link.gen.gpumax,\
pcie.link.gen.hostmax,pcie.link.width.current,pcie.link.width.max --format=csv

# replay counters (cumulative since boot)
nvidia-smi -q | grep -i replay

# the kernel's view of the same link: capability vs status
sudo lspci -vv -s <bus_id> | grep -E "LnkCap:|LnkSta:"

Read the output with one rule in mind: generation may drop at idle, width should not. GPUs lower the link speed when nothing is moving to save power, so a reading of Gen1 on an idle card is normal. Take the reading while a copy is running. Under load, gpucurrent should equal the lower of gpumax and hostmax, and width.current should equal width.max. hostmax tells you what the slot can do, which separates "this server is Gen4" from "this link fell back". The field pcie.link.gen.current still works in many drivers but is deprecated in favour of gpucurrent.

In lspci, compare LnkCap (what the device supports) with LnkSta (what it negotiated); recent versions of lspci mark a lower status as downgraded. If width is x8 on an x16 card, look at the slot first: many boards wire some x16-sized slots with eight lanes, or split lanes when a neighbouring slot is populated.

Measuring what you actually get

Status fields tell you what was negotiated. A copy benchmark tells you what you get. This script times pinned and pageable copies in both directions with CUDA events:

import torch

def copy_gbps(nbytes, pinned, h2d, iters=20):
    host = torch.empty(nbytes, dtype=torch.uint8, pin_memory=pinned)
    dev = torch.empty(nbytes, dtype=torch.uint8, device="cuda")
    src, dst = (host, dev) if h2d else (dev, host)
    for _ in range(3):
        dst.copy_(src, non_blocking=pinned)        # warm-up
    torch.cuda.synchronize()
    start, end = torch.cuda.Event(enable_timing=True), torch.cuda.Event(enable_timing=True)
    start.record()
    for _ in range(iters):
        dst.copy_(src, non_blocking=pinned)
    end.record(); torch.cuda.synchronize()
    return nbytes * iters / (start.elapsed_time(end) / 1e3) / 1e9

for pinned in (True, False):
    for h2d in (True, False):
        print(f"pinned={pinned} {'H2D' if h2d else 'D2H'}: "
              f"{copy_gbps(1 << 30, pinned, h2d):.1f} GB/s")

Compare the pinned figures with the per-direction ceiling for the negotiated generation. A result close to half the Gen5 figure on a Gen5 card usually means the link is running at Gen4 or x8, and the diagnostic commands above will say which. Pin the process to the CPU socket the GPU hangs off, or you measure the inter-socket link too. Run the script on every GPU at once as a second test: if per-GPU numbers fall when all run together, GPUs share a switch uplink.

Workloads that feel the doubling

The doubling matters in proportion to how much time a job spends with the host link as the bottleneck. Four worked examples, using the derived per-direction ceilings, so real times will be somewhat longer:

WorkloadBytes over the linkGen4 x16Gen5 x16
Load 70B-parameter weights in BF16 into GPU memoryabout 140 GB in totalabout 4.4 sabout 2.2 s
Restore a 32k-token KV cache for a 70B model with 8 KV heads10 GiB (320 KiB per token)about 0.34 sabout 0.17 s
Ring all-reduce of 7B BF16 gradients across 8 GPUs, no NVLinkabout 24.5 GB sent per GPUabout 0.78 sabout 0.39 s
Optimizer offload: 7B BF16 gradients down, parameters up14 GB each way, in parallelabout 0.44 sabout 0.22 s

The KV figure comes from 80 layers × 8 KV heads × 128 dimensions × 2 tensors (K and V) × 2 bytes, which is 327,680 bytes per token. Whether 0.17 s or 0.34 s matters depends on the alternative: recomputing the prefill for 32k tokens, or keeping the cache on the GPU. Serving systems that move KV between devices make this trade constantly; see P2P KV cache transfer.

The all-reduce row assumes every GPU has its own full-speed path. On many PCIe servers traffic between GPUs crosses the CPU or a shared switch, and the effective figure is lower. The offload row is the steady cost per step of moving gradients and parameters between host and device; see optimizer state offloading for how frameworks overlap it with compute.

Workloads that do not

Several common jobs will not get faster from Gen5. Training on SXM systems sends gradients over NVLink, which is many times faster than any PCIe generation, so the host link only carries input data and checkpoints; see NVLink and NVSwitch architecture. A data loader that decodes images on the CPU is bound by CPU cores, not by the link. Weight loading from network storage or a single NVMe drive is usually bound by storage throughput, which is below a Gen4 x16 link before it ever reaches the GPU. And a compute-bound kernel that keeps data resident on the device does not touch PCIe at all.

A quick rule: profile one step with Nsight Systems and add up the time the copy engines are busy while kernels are idle. If that share is small, a faster link cannot help much. If it is large, try overlap and pinned memory first; they are free, and the generation upgrade is not.

Gen6 and beyond

Gen6 doubles the rate again, to 64 GT/s, but it does so differently. It switches from two-level NRZ signalling to four-level PAM4, which carries two bits per symbol and is far more prone to errors. To compensate, Gen6 moves to fixed-size flits protected by forward error correction, replacing the 128b/130b scheme. PCI-SIG released the Gen7 specification, at 128 GT/s and also PAM4 with flits, to its members in June 2025.

For GPU systems in 2026, Gen6 appears first in the fabric around the GPU rather than the host CPU, as in the ConnectX-8 example above, where a Gen6 switch sits beside the GPU and the NIC. Plan with what the platform documents; treat any claim about a particular GPU card's Gen6 support as unconfirmed until its datasheet says so. Direct storage-to-GPU paths benefit as these links widen; see GPUDirect Storage architecture.

Failure modes

  • Silent fallback. A Gen5 link trains at Gen4 after a riser swap. Jobs run, a little slower, and nobody notices for months. Alert on gpucurrent below the expected generation under load.
  • Width loss. A damaged connector or a slot shared with an M.2 drive leaves the GPU at x8. Half the bandwidth, no errors.
  • Idle false alarms. Alerting on idle readings pages someone every night. Sample the link while a copy is running, or alert only on width.
  • Rising replays. A marginal channel resends packets and steals bandwidth. Track the counter's rate, not its value, and service the node before it degrades further.
  • Shared uplinks. Two GPUs behind one switch port each see full speed alone and half together. Test concurrently.
  • Forced Gen4 in BIOS. Some operators pin slots to Gen4 to work around unstable risers and forget. Record it as a known setting with a ticket to fix the hardware.

Trade-offs

Gen5 doubles host-link bandwidth at the price of tighter signal-integrity margins, more retimers and more expensive boards, and the gain only shows up in link-bound phases. For inference nodes that load models often, swap KV caches or serve from CPU memory, it is worth having and worth verifying. For NVLink-connected training nodes it mostly shortens checkpoint and startup time. Gen6 continues the trend with PAM4 and FEC, which add a small amount of latency in exchange for bandwidth; for bulk copies that is a good trade.

What to do next

  1. On every GPU node, record gpucurrent, gpumax, hostmax and width under load, and store the result with the node's inventory.
  2. Run the copy script per GPU and all GPUs together; flag any GPU far below the per-direction ceiling for its generation.
  3. Alert on width below max and on a rising replay rate; do not alert on idle generation.
  4. Profile one training or serving step and measure the share of time that is copy-bound before paying for a Gen5 platform.
  5. Use pinned buffers, overlap copies with compute on separate streams, and keep each process on the GPU's local CPU socket.
  6. Check BIOS for slots forced to a lower generation and close each one with a hardware fix.
Key takeaway: Gen5 doubles the host link to about 63 GB/s per direction on x16, but only if every hop trains at Gen5 and full width. Verify generation and width under load, watch replay counters, measure pinned copies, and spend on a faster link only when profiling shows copy-bound time. Weight loading, KV movement and offload benefit; NVLink-bound training mostly does not.