The AMD Instinct MI325X is a memory upgrade of the MI300X. AMD launched it on 10 October 2024 at its Advancing AI event. It keeps the CDNA 3 compute silicon, 304 compute units across eight accelerator complex dies (XCDs), and replaces the 192 GB of HBM3 with 256 GB of HBM3E at 6 TB/s. The price of the faster memory is power: the board is rated at up to 1,000 W against the MI300X's 750 W. AMD's mid-2024 preview had quoted 288 GB; the shipping part settled at 256 GB, and 288 GB arrived with the CDNA 4 based MI350 series instead.
Because the compute is unchanged, the interesting questions are not about peak TFLOPS. They are about what an extra 64 GB and 13 percent more bandwidth change for model placement, KV-cache budgets and decode speed, and about the software details that bite when you move code onto a gfx942 device: 64-wide wavefronts, an FP8 format that differs from NVIDIA's, and partition modes. This article covers those, with a capacity planner you can run. For the broader AMD lineup and the ROCm stack, read AMD Instinct GPUs and ROCm first.
The part in numbers
The figures below come from AMD's MI325X datasheet and product pages. Throughput is dense, without structured sparsity; AMD also quotes sparsity figures at double these values, which few LLM workloads realise.
| MI300X | MI325X | |
|---|---|---|
| Architecture | CDNA 3, gfx942 | CDNA 3, gfx942 |
| Compute units | 304 | 304 |
| Memory | 192 GB HBM3 | 256 GB HBM3E |
| Memory bandwidth | 5.3 TB/s | 6.0 TB/s |
| Infinity Cache | 256 MB | 256 MB |
| FP16 / BF16 dense | 1,307.4 TFLOPS | 1,307.4 TFLOPS |
| FP8 dense | 2,614.9 TFLOPS | 2,614.9 TFLOPS |
| Peer bandwidth, 8-GPU mesh | 896 GB/s aggregate | 896 GB/s aggregate |
| Board power | 750 W | 1,000 W |
| 8-GPU platform memory | 1.5 TB | 2 TB |
Two derived numbers matter more than the raw table. The machine balance, FP16 FLOPS divided by bytes per second, falls from about 247 FLOP/byte on MI300X to about 218 on MI325X: the chip is slightly less starved for bandwidth. And capacity per unit of compute rises by a third, which matters for every workload that is limited by what fits rather than how fast it runs. If the roofline vocabulary is unfamiliar, the HBM architecture article builds it.
Who benefits: decode, prefill and training
LLM decode reads every weight once per generated step and does about two floating-point operations per weight per sequence in the batch. In BF16 that is roughly one FLOP per byte per sequence, so a batch needs on the order of 200 concurrent sequences before decode stops being bandwidth bound on this chip. Most serving deployments run well below that, which means two things: per-step latency scales with bytes of weights and KV cache read, and the batch you can afford is set by how much KV cache fits. The MI325X improves both: 13 percent more bandwidth lowers the latency floor, and the extra 64 GB goes almost entirely to KV cache once the weights are resident.
Prefill and training are different. They run large matrix multiplies at high arithmetic intensity and are bound by the same 1.3 PFLOPS as MI300X. Training benefits indirectly: larger micro-batches, less activation recomputation, and fewer pipeline stages for a given model.
Worked example: planning a 70B and a 405B deployment
Work the numbers for Llama 3.1 70B in BF16 on a single GPU. The weights are 70.6 billion parameters times 2 bytes, about 141 GB. The KV cache for one token is 2 (keys and values) x 80 layers x 8 KV heads x 128 dimensions x 2 bytes, which is 327,680 bytes or 320 KiB. Reserve roughly 15 GB for the runtime, activations, the CUDA-graph-style capture buffers and fragmentation.
GB = 1e9
def kv_bytes_per_token(layers, kv_heads, head_dim, bytes_per_elem=2):
return 2 * layers * kv_heads * head_dim * bytes_per_elem
def plan(params_b, weight_bytes, hbm_gb, bw_tbs, tp=1, overhead_gb=15,
layers=80, kv_heads=8, head_dim=128, kv_elem_bytes=2):
weights = params_b * 1e9 * weight_bytes / tp # per GPU under tensor parallelism
kv_free = hbm_gb * GB - weights - overhead_gb * GB
per_tok = kv_bytes_per_token(layers, kv_heads, head_dim, kv_elem_bytes) / tp
floor_ms = weights / (bw_tbs * 1e12) * 1e3 # weight read per step, KV excluded
return dict(weights_gb=round(weights / GB, 1),
kv_tokens=int(kv_free // per_tok) if kv_free > 0 else 0,
decode_floor_ms=round(floor_ms, 1))
print(plan(70.6, 2, 192, 5.3)) # MI300X: ~36 GB free, ~110k tokens, 26.6 ms
print(plan(70.6, 2, 256, 6.0)) # MI325X: ~100 GB free, ~305k tokens, 23.5 ms
print(plan(405, 1, 256, 6.0, tp=8, layers=126, kv_heads=8, head_dim=128)) # 405B FP8 on 8 GPUsThe result is the practical case for the part. A 70B model in BF16 fits on one MI300X with room for about 110,000 cached tokens; on MI325X it is about 300,000, nearly three times the concurrent context, so a single GPU can hold many more long conversations before the scheduler must preempt. On an 8-GPU board, Llama 3.1 405B in FP8 needs about 51 GB of weights per GPU under tensor parallelism of 8, leaving close to 190 GB each for cache. Each token of its KV cache is 504 KiB across 126 layers, split eight ways, so the node holds roughly 2.9 million cached tokens: enough for dozens of concurrent 32k-token conversations without offloading. On an MI300X board the same model leaves about 126 GB per GPU, still workable but with a third less headroom. Treat these as planning floors, not benchmarks: the KV term adds to the per-step read, and attention kernels, scheduler overhead and communication all add time.
What software sees
To software the MI325X is a gfx942 device, the same target as MI300X, so binaries and tuned kernels carry over. PyTorch's ROCm builds expose HIP through the familiar torch.cuda namespace, which is convenient and occasionally misleading. Confirm what you are running on before trusting any benchmark:
import torch
assert torch.version.hip is not None, "this is not a ROCm build of PyTorch"
props = torch.cuda.get_device_properties(0)
print(props.name, props.gcnArchName, round(props.total_memory / 2**30), "GiB")
print("devices visible:", torch.cuda.device_count()) # 8 per GPU in CPX modeThree porting details catch teams moving CUDA code across:
- Wavefronts are 64 lanes, not 32. Kernels that hard-code a warp size of 32, assume a 32-bit ballot mask, or size shared-memory reductions for 32 lanes give wrong answers or waste half the hardware. Query the wavefront size and use 64-bit masks. Each CU has 64 KB of LDS.
- FP8 is the FNUZ variant. ROCm's documentation lists gfx942 as supporting the FNUZ encodings,
float8_e4m3fnuzandfloat8_e5m2fnuz, and not the OCP formats most NVIDIA-trained FP8 checkpoints use. FNUZ has no infinities and no negative zero, uses a different exponent bias, and E4M3FNUZ tops out at 240 rather than 448. Reinterpreting OCP bits as FNUZ halves every value. Convert through a wider type and recompute scales against 240. - Tuned GEMMs are not automatic. hipBLASLt chooses kernels heuristically. PyTorch's TunableOp, enabled with
PYTORCH_TUNABLEOP_ENABLED=1, benchmarks candidate GEMMs for your exact shapes and caches the winners; run it once on representative traffic and ship the results file with the deployment.
Serving engines such as vLLM and SGLang ship ROCm builds, and AMD publishes containers for them. Pin the container tag alongside your model version, because kernel choice and FP8 support move quickly between releases.
Partitioning and the 8-GPU mesh
Each GPU can be partitioned. In SPX mode all eight XCDs act as one device. In CPX mode each XCD becomes its own device with 38 CUs, so one 8-GPU node presents 64 devices; intermediate modes group XCDs in pairs or fours. Memory partitioning is set separately, as NPS1, one pool, or NPS4, four pools each local to one I/O die. Only certain combinations are valid and the set depends on firmware, so check AMD's partitioning documentation for your driver release rather than assuming.
Partitioning helps when many small models, such as embedding models or 7B-class chat models, would each underuse a whole GPU. It hurts anything that needs the full 256 GB or all 304 CUs, and it changes device numbering, so orchestration and monitoring must agree on the mode. A node rebooted into a different mode than its scheduler expects is a classic cause of mysterious out-of-memory errors.
Across GPUs, the eight devices on a platform board are fully meshed with Infinity Fabric, with 896 GB/s of aggregate peer bandwidth per GPU, and RCCL provides the NCCL-compatible collectives. Tensor parallelism of 8 stays inside the mesh; anything larger crosses the network, where pipeline or expert parallelism is usually the better split.
Failure modes
- Power budget surprises. Eight boards at 1,000 W is 8 kW of GPUs per node before CPUs, NICs and fans. Racks provisioned for MI300X-class power cannot simply swap boards in. Under a power cap the chip clocks down, and benchmark numbers drop with it.
- FP8 checkpoint mismatch. Loading OCP FP8 weights without conversion produces a model that runs, loads cleanly and emits subtly wrong output. Validate perplexity against a BF16 reference after every quantised deployment.
- Capacity planned on the label. 256 GB of HBM is not 256 GB of KV cache. Leave headroom for the runtime and fragmentation, and set the engine's memory fraction deliberately.
- Untuned kernels. Default GEMM heuristics can leave a large fraction of throughput unused on unusual shapes. Tune, then measure.
- Warp-size assumptions in custom CUDA kernels ported through HIPify: they compile and fail quietly.
- Partition drift between what the node booted into and what the scheduler expects.
Trade-offs
| Choice | Choose it when | Give up |
|---|---|---|
| MI325X over MI300X | Long contexts, large batches, models just over 192 GB | 250 W more per board; same compute |
| MI325X over H200 | Capacity per GPU matters: 256 GB vs 141 GB | CUDA-only kernels and tooling; porting effort |
| MI350 series over MI325X | You need 288 GB, OCP FP8 or FP4 and FP6 | Newer stack; validate support first |
| CPX partitions | Many small models per GPU | Per-model memory and compute ceilings |
| Fewer GPUs with more memory | Decode-bound serving | Peak prefill throughput per node |
For the NVIDIA comparison points, see the H200 in depth and the B200 in depth.
What to do next
- Run the device check script on your target node and record the arch name, memory and device count.
- Run the planner for your own models, with your real layer count, KV heads and context mix, on MI300X and MI325X figures, and decide whether capacity or compute is your constraint.
- Audit every FP8 checkpoint for OCP versus FNUZ encoding and add a perplexity check against BF16 to deployment.
- Grep custom kernels for a hard-coded warp size of 32 and 32-bit ballot masks.
- Enable TunableOp on a staging run with production shapes, and ship the tuned results file with the release.
- Confirm rack power and cooling for 1,000 W boards, and decide on the partition mode before the scheduler is configured.