Intel's Gaudi accelerators run both their scale-up and their scale-out traffic over standard Ethernet ports built into the chip. Each chip carries 24 RDMA-over-Converged-Ethernet ports, and the same ports form both the fabric inside a server and the link out of it. That single design decision shapes how you should lay out parallelism, how you build the cluster network, and what changes when you move a workload from Gaudi 2 to Gaudi 3.
This article compares the two current generations from the point of view of the person who has to run software on them. It covers the figures Intel publishes, the in-box mesh and why it rewards different parallel layouts from an NVSwitch system, scale-out arithmetic, a worked comparison of collective time and memory fit, a porting checklist and the version discipline that avoids most bring-up failures. How the graph compiler, lazy and eager modes and the PyTorch bridge work is covered in Intel Gaudi and GPU Max in depth; this page assumes it.
Three corrections before the numbers
Three corrections first, because older material, including the stub this page replaces, gets them wrong. AWS DL1 instances used first-generation Gaudi, not Gaudi 2. The software stack once called SynapseAI is now simply the Intel Gaudi software, released as numbered versions that cover both current generations. And Gaudi does have a scale-up fabric: it is not NVLink, but the all-to-all Ethernet mesh inside the server plays the same role.
The two generations side by side
| Figure | Gaudi 2 | Gaudi 3 | What it limits |
|---|---|---|---|
| HBM capacity | 96 GB HBM2E | 128 GB HBM | model state and KV cache per chip |
| HBM bandwidth | 2.45 TB/s | 3.7 TB/s | small-batch decode, memory-bound kernels |
| On-die SRAM | 48 MB | larger; reported as 96 MB | how much the compiler keeps on chip |
| Matrix and vector engines | MME and TPC cluster | 8 MME, 64 TPC, two compute dies | compute-bound training and prefill |
| Dense compute | Intel: Gaudi 3 is 4x BF16, 2x FP8 | about 1.8 PFLOPS FP8 and BF16 | step time at large batch |
| RoCE ports | 24 x 100 Gb/s | 24 x 200 Gb/s | every collective |
| Data types | FP32, TF32, BF16, FP16, FP8 E4M3 and E5M2 | FP8 and BF16 headline; check the docs for the full list | quantisation choices |
Intel's own summary of the step is two times the FP8 compute, four times the BF16 compute, twice the network bandwidth and one and a half times the memory bandwidth. The last ratio matters most for inference: token generation at small batch is bounded by HBM bandwidth, so expect closer to 1.5 times than to 4 times there.
The box is a mesh, not a switch
Inside an eight-chip Gaudi server, 21 of each chip's 24 ports are wired directly to the other seven chips, three ports to each peer. There is no switch. The remaining three ports on each chip leave the box for scale-out. Intel's network configuration guide gives the resulting figures: on Gaudi 3 each chip has 525 GB/s per direction into the mesh and 75 GB/s per direction out of the box; on Gaudi 2, with 100 Gb/s ports, the figures are 262.5 GB/s and 37.5 GB/s.
The consequence is the most important software fact on this page. An NVSwitch lets any one pair of GPUs use a GPU's full NVLink bandwidth. The Gaudi mesh gives any one pair only three ports, one seventh of the in-box bandwidth. Collectives that spread traffic evenly over all peers, such as all-reduce, all-gather, reduce-scatter and the all-to-all of expert parallelism, use the whole mesh. Patterns that concentrate on one pair do not.
Parallel layouts that suit the mesh
- Tensor parallelism of 8, not 2 or 4. A tensor-parallel group of two uses one pair's three ports. A group of four uses three peers, about 225 GB/s per direction on Gaudi 3. A group of eight uses the whole mesh. If the model fits with less, prefer data-parallel replicas of eight over several small tensor-parallel groups in one box when communication dominates. The general rules are in tensor parallelism in depth.
- Expert parallelism fits naturally. An all-to-all inside the box gives every pair its own links with no switch contention, which is the pattern a full mesh is best at.
- Pipeline stages across boxes, not inside. A stage boundary inside a box is a one-pair transfer at 75 GB/s; across boxes it is a scale-out transfer at a similar 75 GB/s per chip, so placing stages in different boxes costs little extra and keeps the mesh for the collectives that use it well.
Scale-out: Ethernet from the chip
Each box exposes 24 scale-out ports, three per chip. On Gaudi 3 that is 4.8 Tb/s, or 600 GB/s per direction per box; on Gaudi 2 it is 2.4 Tb/s, or 300 GB/s. For comparison, an eight-GPU H100 box with one 400 Gb/s adapter per GPU has 3.2 Tb/s. The ports are ordinary Ethernet, so the cluster network is a RoCE v2 fabric of Ethernet leaf and spine switches. It is not plug and play: RoCE needs a lossless or congestion-controlled configuration, consistent MTU and the right traffic classes end to end. Intel publishes switch configuration examples; the general practice is covered in RoCE v2 in depth.
Because the network adapters are on the accelerator, scale-out traffic does not cross the host PCIe bus. That removes the GPU-to-NIC placement and PCIe-switch failure class that GPU servers have, and adds another: a chip with a bad port is now a chip with less bandwidth, and the collective slows to the slowest chip's rate.
Worked example: one collective and one model on each generation
A gradient all-reduce. Take 14 GB of bf16 gradients, a 7-billion-parameter model, on four boxes of eight chips. Use the hierarchical scheme that suits the hardware: reduce-scatter inside each box, all-reduce the shards across boxes, all-gather inside each box. Each in-box phase moves 7/8 of 14 GB, 12.25 GB per chip, split evenly over seven peers, so 1.75 GB per pair. On Gaudi 3, at 75 GB/s per pair, each phase takes about 23 ms; on Gaudi 2, about 47 ms. The cross-box phase is a four-party ring on each chip's 1.75 GB shard, moving 2 × 3/4 × 1.75 GB, about 2.6 GB, through the chip's scale-out ports: about 35 ms on Gaudi 3 and 70 ms on Gaudi 2. Totals: roughly 81 ms against 164 ms at line rate. Real efficiency will be lower; measure it. The algorithms themselves are explained in all-reduce in depth.
Fitting a 70B model for inference. A 70-billion-parameter model in bf16 needs about 140 GB of weights. With grouped-query attention of 8 KV heads, 80 layers and head size 128, each token of KV cache costs 80 × 8 × 128 × 2 × 2 bytes, about 320 KB. On two Gaudi 2 chips, 192 GB leaves about 52 GB before runtime workspace; budget 8 GB for that and about 44 GB remains, roughly 134,000 cached tokens across all sequences. On two Gaudi 3 chips, 256 GB leaves about 108 GB after the same allowance, roughly 330,000 tokens: two and a half times the concurrent context from a 33 percent larger HBM. With FP8 weights the model fits on a single Gaudi 3 with room for cache, which is often the best layout for this model on Gaudi 3 because it avoids tensor parallelism entirely.
Porting a Gaudi 2 job to Gaudi 3
The same software release and the same PyTorch code run on both generations, so porting is less about code than about re-measuring. Work through these in order.
- Match versions as a set. Driver, firmware, container image and the PyTorch bridge are released together and tested together. Mixing a container from one release with a driver from another is the most common cause of failures at start-up. Pin the release number in your image tag and check it on every node.
- Recompile and rewarm. Compiled graphs are specific to the device. Do not share a compiled-graph cache between generations, and expect the first steps on Gaudi 3 to be slow while graphs build.
- Re-size batches. A third more HBM changes the best micro-batch and may let you drop activation checkpointing or tensor parallelism. Re-run the memory fit rather than reusing the Gaudi 2 configuration.
- Re-calibrate FP8. FP8 scaling factors are measured on real data per model and per configuration. Measure again on the new hardware instead of copying scale files.
- Re-check the network. Gaudi 3 ports run at 200 Gb/s, so switches, optics and cables must match. A box cabled for Gaudi 2 is not ready for Gaudi 3.
- Check serving support. For inference, the vLLM hardware plugin for Intel Gaudi became production-ready with software 1.22.2, and its compatibility matrix lists which vLLM versions go with which Gaudi release. For Hugging Face training, see Optimum.
A short check, run inside the container on every node, catches most of this before a job wastes a reservation.
# gaudi_env_check.py -- run inside the container on every node before a job
import json, subprocess
import torch
import habana_frameworks.torch.hpu as hthpu
info = {
"torch": torch.__version__,
"available": hthpu.is_available(),
"devices": hthpu.device_count(),
"device_name": hthpu.get_device_name(), # reports the generation, e.g. Gaudi 2 or Gaudi 3
}
# hl-smi is the management CLI; its header carries driver and firmware versions.
info["hl_smi_head"] = subprocess.run(["hl-smi"], capture_output=True,
text=True).stdout.splitlines()[:4]
print(json.dumps(info, indent=2))
assert info["devices"] == 8, "a chip is missing: check hl-smi and dmesg before training"
Failure modes
- Version skew. The driver on one node is a release behind; the job hangs at communicator creation or fails to find devices. Fix it with a fleet-wide version check, not by reading logs on one node.
- A degraded port. One chip has a port down in the mesh or to the leaf. Collectives still complete, slower, and every chip waits for the slowest. Compare per-chip port status across the box after every maintenance window.
- Small tensor-parallel groups. A layout copied from an NVSwitch system uses tensor parallelism of 2 and gets one seventh of the mesh. Step time is dominated by communication.
- Dynamic shapes. Variable sequence lengths trigger repeated graph compilation, and throughput collapses. Bucket shapes, and warm the buckets before taking traffic.
- An untuned Ethernet fabric. Pause storms or drops under load show up as collective timeouts at scale but never on one box. Validate the fabric with a multi-box collective test before handing the cluster over.
Trade-offs
Gaudi's case is cost and openness: Ethernet everywhere, no proprietary switch, and a competitive amount of HBM per chip. Its costs are a smaller software ecosystem than CUDA, a compiler that wants static shapes, and a mesh that rewards fewer, larger communication groups. Between the two generations, Gaudi 3 is the default for new work; Gaudi 2 remains a sensible home for inference of models that fit, where its lower bandwidth matters least, and for fleets already paid for. For planning beyond them, Intel announced at the 2025 OCP summit a Gaudi 3 rack-scale reference design of up to 64 accelerators per rack and an inference GPU called Crescent Island, due to sample in the second half of 2026; check Intel's current statements before committing to a roadmap.
What to do next
- Identify which generation each workload runs on and record its release number, driver, firmware and container image.
- Redo tensor-parallel layouts for the mesh: prefer groups of eight, or no tensor parallelism.
- Run the worked all-reduce estimate with your gradient size and compare it with a measured collective test on one box, then four.
- Recompute memory fit and KV-cache capacity for Gaudi 3 before reusing Gaudi 2 batch settings.
- Re-measure FP8 scales on the new hardware, and validate accuracy against the bf16 baseline.
- Run the environment check on every node in the container image you will train with.