AMD's MI325X and MI350 series are a year apart, and they look interchangeable from a distance: both are eight-GPU OAM platforms with HBM3E and Infinity Fabric, and the MI350X is designed to drop into MI325X-class baseboards. Underneath, they are different architectures. The MI325X is CDNA 3, the MI300X compute dies with bigger memory. The MI350X and MI355X are CDNA 4, which has new compute dies on a newer process, fewer but wider compute units, a 160 KB local data share, microscaling number formats and a different FP8 encoding. Software that ran well on one does not automatically run well, or even correctly, on the other.

This article is about that step from CDNA 3 to CDNA 4 as software sees it. It compares the parts on the numbers that drive model placement, explains the formats and what they cost in accuracy, works through serving a 405B model on one node of each, and gives a migration runbook. The MI325X itself, its 70B and 405B plans and its partition modes, is covered in the MI325X in depth. The MI350 figures come from AMD's ROCm microarchitecture page, AMD's product pages and the ROCm release notes, checked in October 2026. Throughput is dense, without structured sparsity.

The two generations in numbers

MI325XMI350XMI355X
Architecture, ISA targetCDNA 3, gfx942CDNA 4, gfx950CDNA 4, gfx950
Compute units304 (8 x 38)256 (8 x 32)256 (8 x 32)
LDS per CU64 KB160 KB160 KB
HBM3E256 GB, 6.0 TB/s288 GB, 8 TB/s288 GB, 8 TB/s
FP16 / BF16 matrix1.31 PF2.3 PF2.5 PF
FP8 matrix2.61 PF (FNUZ)4.6 PF (OCP)5.0 PF (OCP)
MXFP6 / MXFP4not supported9.2 PF10 PF
FP64 matrix163.4 TF72.1 TF78.6 TF
Board power1,000 Wup to 1,000 Wup to 1,400 W
Peak engine clock2,100 MHz2,200 MHz2,400 MHz
Coolingpassive OAMpassive OAMdirect liquid platforms

The MI350X and MI355X are the same silicon. The MI355X runs at higher clocks and power and needs liquid cooling, which buys about 9 percent more peak throughput. Three rows matter most for planning. Memory grows by 12.5 percent and bandwidth by a third. Dense FP16 and FP8 roughly double despite fewer compute units, because each CDNA 4 compute unit does twice the matrix work per clock. And FP64 matrix throughput falls by more than half, a deliberate trade toward AI that makes the MI325X, or the HPC-oriented parts, the better choice for double-precision simulation.

Inside the CDNA 4 package

MI355X package: two I/O dies, eight XCDs, eight HBM3E stacksHBM3E stack 036 GBHBM3E stack 436 GBHBM3E stack 136 GBHBM3E stack 536 GBHBM3E stack 236 GBHBM3E stack 636 GBHBM3E stack 336 GBHBM3E stack 736 GBI/O die 0: Infinity Cache, HBM PHYs, IF linksI/O die 1: Infinity Cache, HBM PHYs, IF linksXCD 032 CUs, 4 MB L2XCD 132 CUs, 4 MB L2XCD 232 CUs, 4 MB L2XCD 332 CUs, 4 MB L2XCD 432 CUs, 4 MB L2XCD 532 CUs, 4 MB L2XCD 632 CUs, 4 MB L2XCD 732 CUs, 4 MB L2256 CUs (MI325X: 304) | 160 KB LDS per CU (was 64 KB) | 256 MB Infinity Cache | 8 TB/s (was 6)Seven Infinity Fabric links per GPU at 38.4 Gbps, over 1 TB/s aggregate, in the same 8-GPU UBB layout
CDNA 4 halves the number of I/O dies from four to two. Each XCD has 36 physical CUs, of which 32 are enabled.

Software sees the same programming model as before: wave64 wavefronts, HIP, matrix instructions, a per-XCD L2 and the Infinity Cache as a memory-side cache shared by all XCDs. Two changes alter how kernels should be written. The LDS, the software-managed scratchpad that GEMM and attention kernels tile through, grows from 64 KB to 160 KB per compute unit with doubled read bandwidth. A kernel tuned for 64 KB tiles still runs but leaves capacity unused: larger tiles mean more reuse per byte fetched from HBM, and more stages of software pipelining. The second change is the matrix units, which add microscaling formats, OCP FP8, and double the dense rate per compute unit. Kernels whose tile shapes were chosen for CDNA 3's matrix instruction latencies need re-tuning to keep the wider units fed.

Machine balance: who gains what

Machine balance, the peak FLOPS divided by memory bandwidth, says how many operations a kernel must perform per byte of HBM traffic before compute rather than memory sets its speed. Computed the same way as for the MI325X:

PrecisionMI325XMI355X
FP16 / BF16218 FLOP/byte312 FLOP/byte
FP8436 FLOP/byte625 FLOP/byte
MXFP4n/a1,250 FLOP/byte

Compute grew faster than bandwidth, so the MI355X is more bandwidth-starved than its predecessor. Prefill and training, which reuse each weight across thousands of tokens, gain nearly the full doubling. LLM decode, which reads every weight once per generated token and does about two operations per weight per sequence, was already memory-bound and stays memory-bound. Its speed-up comes from the 33 percent extra bandwidth and from smaller weights, not from the FLOPS. That is the real case for MXFP4 on this chip: halving bytes per weight relative to FP8 nearly halves the decode time floor. The general roofline argument is developed in the HBM architecture article.

Number formats: OCP FP8 and microscaling

FP8 changes encoding. CDNA 3 implements the FNUZ variants of FP8, which have no negative zero and use a different exponent bias from the OCP E4M3 and E5M2 formats that most published FP8 checkpoints use. CDNA 4 implements the OCP formats. Checkpoints that were converted to FNUZ for MI300-series serving must not be copied onto MI350 nodes; go back to the OCP original. In PyTorch the dtypes are distinct, torch.float8_e4m3fnuz versus torch.float8_e4m3fn, which makes the mistake detectable if you check for it.

Microscaling formats. The OCP Microscaling (MX) specification groups values into blocks of 32 that share one 8-bit power-of-two scale (E8M0). Each element is then a tiny float: FP4 E2M1 for MXFP4, FP6 for MXFP6, FP8 for MXFP8. MXFP4 therefore costs 4.25 bits per value. A notable detail of CDNA 4 is that MXFP6 runs at the same rate as MXFP4, so the extra accuracy of six bits costs memory but no compute. A reference quantizer shows how the block scale works:

import numpy as np
FP4_GRID = np.array([0, 0.5, 1, 1.5, 2, 3, 4, 6])     # E2M1 magnitudes

def mxfp4_quantize(x, block=32):
    x = x.reshape(-1, block)
    amax = np.abs(x).max(axis=1, keepdims=True)
    # 6.0 = 1.5 * 2^2, so align the block maximum to exponent 2
    shared = np.floor(np.log2(np.where(amax == 0, 1.0, amax))) - 2
    scale = 2.0 ** shared                               # one E8M0 value per block
    mag = np.clip(np.abs(x) / scale, 0, 6.0)
    q = np.sign(x) * FP4_GRID[np.abs(mag[..., None] - FP4_GRID).argmin(-1)]
    return q, scale                                     # dequantize with q * scale

On Gaussian weights with standard deviation 0.02, this gives a relative reconstruction error of about 11.5 percent. Plant one outlier of 1.0 in a block and 31 of its 32 values round to zero, because the shared scale must cover the outlier. That is why MXFP4 checkpoints are produced with calibration and outlier handling, AMD's Quark quantizer among the tools, rather than by direct rounding, and why accuracy must be measured per model and task.

Worked example: a 405B model on one node

Serve a 405B dense model with 126 layers, 8 KV heads of dimension 128 and grouped query attention on one eight-GPU node, using tensor parallelism of 8. The KV cache costs 2 x 126 x 8 x 128 values per token: 516 KB in BF16 or 258 KB in FP8. Reserve 10 percent of HBM for the runtime, activations and fragmentation.

PlanWeights per GPUKV budget per nodeBF16 KV tokensDecode floor per step
MI325X, FP8 weights50.6 GB1,438 GBabout 2.8 M8.4 ms
MI355X, FP8 weights50.6 GB1,669 GBabout 3.2 M6.3 ms
MI355X, MXFP4 weights26.9 GB1,858 GBabout 3.6 M3.4 ms

The decode floor is weight bytes per GPU divided by bandwidth: the time to stream each GPU's shard once, at the peak rate, ignoring KV reads and communication. Real steps take longer, but the ratios carry. Two conclusions follow. The MI355X with FP8 already beats the MI325X on both capacity and latency without any change in precision, so migrate the precision you have first. MXFP4 then nearly halves the floor again, but only pays off if the quantised model passes your accuracy bar. At 215 GB, the MXFP4 model also fits on a single 288 GB GPU, which makes eight independent replicas per node possible for small contexts, a layout the MI325X cannot offer. How much KV a workload needs is worked through in sizing the KV cache.

Software migration: gfx942 to gfx950

ROCm 7.0 added support for the MI350X and MI355X, together with MX formats in Composable Kernel and hipBLASLt. Every binary must contain gfx950 code objects: kernels compiled only for gfx942 do not run on CDNA 4. The stack's layers and version pinning are described in the ROCm stack article. A first check on a new node:

# confirm the ISA and the device count the driver exposes
rocminfo | grep -m1 -o "gfx9[0-9a-f]*"
amd-smi list

# build extensions and custom kernels for both generations during migration
hipcc --offload-arch=gfx942 --offload-arch=gfx950 -O3 kernel.hip -o kernel
export PYTORCH_ROCM_ARCH="gfx942;gfx950"

# in Python: refuse to load an FNUZ checkpoint on CDNA 4
import torch
arch = torch.cuda.get_device_properties(0).gcnArchName
FNUZ = (torch.float8_e4m3fnuz, torch.float8_e5m2fnuz)
fnuz = any(t.dtype in FNUZ for t in state_dict.values())
if arch.startswith("gfx950") and fnuz:
    raise RuntimeError("FNUZ FP8 weights on gfx950: reload the OCP checkpoint")

Partitioning carries over in spirit but not in detail. The MI350 series supports one, two, four or eight compute partitions with NPS1 or NPS2 memory modes; in CPX with NPS2 each of the eight partitions is one XCD with 36 GB. The MI325X's NPS4 mode has no counterpart, because there are now two I/O dies rather than four. Scripts and schedulers that hard-code partition sizes need updating.

Failure modes

Kernel not found on gfx950. A wheel or extension built for gfx942 fails at load or launch time. Rebuild with both targets and keep a check in CI.

FNUZ weights on OCP hardware. Loading the wrong FP8 variant can produce plausible but wrong output. Check dtypes and compare perplexity against a BF16 reference.

Old tuning tables. GEMM and attention tuning results recorded on MI325X choose tile sizes that ignore the larger LDS. Re-run tuning on gfx950 before benchmarking.

Unvalidated MXFP4. A model quantised by naive rounding loses accuracy unevenly across tasks. Evaluate on your own tasks, not only on a public benchmark.

Power and cooling assumptions. An MI355X node is around 11 kW of GPUs alone and needs direct liquid cooling. The MI350X fits existing air-cooled 1,000 W designs at lower clocks.

FP64 regressions. Scientific codes that relied on CDNA 3's FP64 matrix rate can run slower on CDNA 4. Benchmark them before moving them.

Trade-offs

ChoiceChoose it whenGive up
Stay on MI325XFP64 work, a validated FNUZ stack, capacity already sufficientHalf the dense AI throughput, no MX formats
MI350XExisting air-cooled 1,000 W racks, drop-in platform upgradeAbout 9 percent peak versus MI355X
MI355XLiquid-cooled facilities, peak throughput per node1,400 W per board; facility work
FP8 on CDNA 4Quick migration, minimal accuracy riskThe decode gain MXFP4 would add
MXFP4 on CDNA 4Decode-bound serving after accuracy validationCalibration effort and per-task evaluation

For background on how the MI300X began this line, see the MI300X article.

What to do next

  1. Inventory every binary, wheel and container in your stack and confirm it contains gfx950 code objects.
  2. Find every FP8 checkpoint and record whether it is OCP or FNUZ; keep the OCP originals for CDNA 4.
  3. Recompute your memory plan with 288 GB, a 10 percent reserve and your real KV precision, as in the worked example.
  4. Migrate at your current precision first, re-tune GEMM and attention kernels on gfx950, and benchmark.
  5. Only then quantise to MXFP4 or MXFP6, and gate the rollout on task-level accuracy against a BF16 reference.
  6. Benchmark FP64 workloads separately and keep them on CDNA 3 if they regress.
Key takeaway: The MI350 series is a new architecture, not a faster MI325X. It doubles dense AI throughput on fewer, wider CUs, adds 32 GB and a third more bandwidth, switches FP8 to OCP encoding, adds MXFP4 and MXFP6, and cuts FP64. Rebuild for gfx950, reload OCP checkpoints and re-tune kernels. Then decide on MXFP4 by measured accuracy, because decode gains come from bytes, not FLOPS.