Intel sells, or has sold, two very different kinds of AI accelerator, and the most common mistake is to treat them as one product family. Gaudi, which came from the Habana Labs acquisition, is not a GPU at all: it is a set of fixed-function matrix engines and programmable vector cores, driven by a graph compiler, with Ethernet RDMA ports built into the chip. Data Center GPU Max, codenamed Ponte Vecchio, is a general-purpose GPU in the same broad family as NVIDIA and AMD parts, programmed through SYCL and the Level Zero runtime, and reached from PyTorch as the device type xpu.
This page explains both from the point of view of the person who has to make a model run fast on them: what the hardware contains, how PyTorch code changes for each, how multi-card jobs communicate, a worked example, failure modes and a checklist. Every specification quoted comes from an Intel page; where Intel publishes no figure, none is given.
Two products, two design philosophies
A GPU runs roughly one kernel per operation: a matrix multiply, then a layer norm, then an activation, each scheduled across thousands of threads. Fusion helps, but the default unit of work is the operator. That is the model GPU Max follows. Its building block is the Xe-core, which contains vector engines for general arithmetic and XMX (Xe Matrix Extensions) engines for dense matrix multiply, much as a streaming multiprocessor contains CUDA cores and tensor cores.
Gaudi takes the opposite position. Intel's architecture documentation describes two kinds of compute: the MME (Matrix Multiplication Engine), which handles every operation that can be lowered to a matrix multiplication, and a cluster of TPCs (Tensor Processor Cores), programmable VLIW SIMD processors that handle everything else. A graph compiler sits between PyTorch and the chip. It receives a whole graph of operations, decides which go to the MME and which to the TPCs, fuses what it can, and plans how tensors move between on-die SRAM and HBM. Performance therefore depends far more on whether the compiler sees large, stable graphs than on any single kernel.
The second philosophical split is networking. Each Gaudi chip carries its own RDMA over Converged Ethernet (RoCE) ports, so cards talk to each other inside a server and across servers over standard Ethernet without a separate network card. GPU Max uses Xe Link between cards in the same node and relies on host network adapters to leave the node, which is the conventional GPU arrangement.
Product status matters for planning. Intel's specification page for the Data Center GPU Max 1550 lists launch in Q1 2023 and an expected discontinuance date of January 2026, so GPU Max should be treated as a fleet to run and eventually migrate off, not as a platform for new purchases. Intel Extension for PyTorch (IPEX), long the usual way to run PyTorch on Intel GPUs, reached end of life at the end of March 2026, with Intel directing users to the XPU support that is now upstream in PyTorch itself. Gaudi 3 is the current Gaudi generation.
The published numbers, and what each one controls
The table collects the figures Intel publishes. Each row is listed because it limits a particular kind of workload, not because bigger is always better.
| Figure | Gaudi 2 | Gaudi 3 | GPU Max 1550 |
|---|---|---|---|
| Matrix compute units | MME | 8 MME | 1024 XMX engines in 128 Xe-cores |
| Vector / general compute | TPC cluster | 64 TPC | 1024 Xe vector engines |
| Memory | 96 GB HBM2E | 128 GB HBM | 128 GB HBM2e |
| Memory bandwidth | 2.45 TB/s | 3.7 TB/s | 3276.8 GB/s |
| On-die SRAM | 48 MB | larger than Gaudi 2 | - |
| Networking | 24 x 100 Gb/s RoCE v2 | 24 x 200 Gb/s RDMA | Xe Link between cards; host NIC for scale-out |
| Headline dense compute | - | about 1.8 PFLOPS FP8 and BF16 | - |
| Power | - | - | 600 W TDP |
Capacity decides what fits: weights, optimizer state, activations and KV cache. Bandwidth decides memory-bound phases such as small-batch token generation. Matrix compute decides compute-bound phases such as large-batch training or prefill. Networking decides how much of each step goes to all-reduce. Gaudi documentation also lists the data types its MME supports, including FP32, TF32, BF16, FP16 and two FP8 formats (E4M3 and E5M2), all accumulated in FP32.
One practical point about GPU Max: the 1550 is built from two stacks, and the Level Zero runtime can present each stack as its own device or the whole card as one device. Which you get affects how many devices PyTorch reports and how you assign ranks, so check it with sycl-ls and torch.xpu.device_count() before writing launch scripts.
How Gaudi executes a PyTorch step
PyTorch reaches Gaudi through a bridge package, habana_frameworks.torch, which registers the device type hpu. There are two execution modes. In lazy mode, operations are not run when Python calls them; they are accumulated into a graph until the script calls htcore.mark_step(), and then the whole graph is compiled (or fetched from a cache) and executed. In eager mode, operations run one by one, and torch.compile with the hpu_backend backend compiles regions into graphs. The Intel Gaudi documentation for release 1.24.0 states that eager mode with torch.compile is the default and that lazy mode must be requested with PT_HPU_LAZY_MODE=1; older releases defaulted to lazy mode, so check the documentation for the release you have installed.
Either way, the graph compiler is the centre of performance. It compiles a graph for a specific set of tensor shapes. If the shapes change every step, as they do when you feed variable-length sequences without padding, the compiler produces a new graph each time, and compilation can take far longer than execution. The cure is the same as on TPUs: pad or bucket inputs to a small fixed set of shapes so that after warm-up every step hits the cache.
A minimal training loop looks like this. The only differences from a CUDA script are the import, the device name and the compile backend.
import torch
import habana_frameworks.torch.core as htcore # registers the "hpu" device
device = torch.device("hpu")
model = MyModel().to(device)
model = torch.compile(model, backend="hpu_backend")
opt = torch.optim.AdamW(model.parameters(), lr=3e-4)
for step, (x, y) in enumerate(loader): # loader yields fixed, bucketed shapes
x, y = x.to(device), y.to(device)
with torch.autocast(device_type="hpu", dtype=torch.bfloat16):
loss = torch.nn.functional.cross_entropy(model(x), y)
loss.backward()
opt.step()
opt.zero_grad(set_to_none=True)
# In lazy mode (PT_HPU_LAZY_MODE=1) you would call htcore.mark_step()
# after loss.backward() and again after opt.step().
if step % 50 == 0:
print(step, loss.item()) # .item() forces a sync: keep it rareTwo habits matter more here than on a GPU. Every .item() or Python branch on a tensor value forces a host sync and can split a graph. And operators the bridge does not support fall back to the CPU, which shows up as a sudden throughput drop; list fallbacks from the bridge logs before tuning anything else. For Hugging Face models, optimum-habana wraps these patterns, including static shapes and HPU graphs for inference, and is the fastest route to a baseline.
Scale-out on Gaudi: Ethernet on the chip
Because the RoCE ports are part of each Gaudi chip, collective communication does not pass through host PCIe and a separate NIC. In Intel's eight-card server designs, most of each chip's ports are wired directly to the other seven cards to form an all-to-all mesh inside the box, and the remaining ports go out to an Ethernet fabric for scale-out. The communication library is HCCL (Habana Collective Communications Library), which presents the same interface as NCCL to PyTorch.
import os
import torch
import torch.distributed as dist
import habana_frameworks.torch.distributed.hccl # registers the "hccl" backend
dist.init_process_group(backend="hccl") # rank and world size from the launcher
local_rank = int(os.environ["LOCAL_RANK"])
model = MyModel().to("hpu")
model = torch.nn.parallel.DistributedDataParallel(model)Scale-out bandwidth is therefore a property of how many chip ports are cabled out and how the Ethernet fabric is built. When a job is slower across servers than within one, or hangs at the first all-reduce, check the switch configuration (lossless RoCE settings, ports cabled per server) before blaming the model.
How GPU Max executes the same step
On GPU Max the software stack is the oneAPI one. SYCL is the C++ programming model for kernels, Level Zero is the low-level runtime that loads them and manages memory, and oneDNN and oneMKL provide the tuned convolution, matrix-multiply and normalisation kernels that PyTorch calls. Since PyTorch 2.4 the XPU device has been supported upstream on Linux for the Max series, with torch.compile able to target Intel GPUs. Porting a CUDA script is mostly a matter of replacing the device string.
import torch
assert torch.xpu.is_available(), "check driver, Level Zero and PyTorch XPU build"
device = torch.device("xpu")
model = MyModel().to(device)
model = torch.compile(model) # Inductor generates kernels for XPU
opt = torch.optim.AdamW(model.parameters(), lr=3e-4)
for x, y in loader:
x, y = x.to(device, non_blocking=True), y.to(device, non_blocking=True)
with torch.autocast(device_type="xpu", dtype=torch.bfloat16):
loss = torch.nn.functional.cross_entropy(model(x), y)
loss.backward()
opt.step()
opt.zero_grad(set_to_none=True)
torch.xpu.synchronize()Because the unit of work is still the operator, dynamic shapes hurt far less than on Gaudi. The pain points are CUDA-only extensions (custom kernels, some attention and quantization libraries) that will not load, and code that hard-codes torch.cuda. For multi-card jobs, older stacks used the oneCCL bindings with backend ccl; newer upstream releases have their own XPU collective backend, so take the name from your PyTorch version's documentation. Replace any remaining IPEX dependency now that the project is end of life.
Worked example: memory fit and decode speed
Take a 13-billion-parameter decoder model. In BF16 each parameter is 2 bytes, so the weights alone are about 26 GB. That fits easily on one card of any of the three parts above, which leaves the question of speed.
During token-by-token generation at batch size 1, each new token requires reading essentially every weight once. That phase is memory-bound, so a useful upper bound on speed is bandwidth divided by bytes read per token. On Gaudi 3: 26 GB divided by 3.7 TB/s is about 7.0 ms per token, or at most roughly 140 tokens per second. On GPU Max 1550: 26 GB divided by 3.28 TB/s is about 7.9 ms, or roughly 125 tokens per second. On Gaudi 2: 26 GB divided by 2.45 TB/s is about 10.6 ms, or roughly 94 tokens per second. Real systems reach a fraction of these ceilings, but the ratio between devices tends to hold, and the arithmetic tells you that peak FLOPS is irrelevant to this phase. Batching many sequences together is what moves decode towards being compute-bound, because the same weight read serves every sequence in the batch.
Training is different. Full fine-tuning with AdamW in mixed precision needs about 16 bytes per parameter for weights, gradients, FP32 master weights and two optimizer moments: for the same 13B model that is about 208 GB before any activations. No single card holds it. Sharding that state across eight cards with FSDP or ZeRO-3 brings it to about 26 GB per card, leaving roughly 100 GB for activations on a 128 GB card and about 70 GB on Gaudi 2.
Failure modes seen in practice
- Recompilation storms on Gaudi. Erratic throughput and very slow early steps from varying shapes. Fix: bucket shapes, drop partial batches, watch the compile count.
- Silent CPU fallback. An unsupported operator runs on the host. Fix: list fallbacks and replace the operator.
- Graph breaks from host syncs. Per-step logging splits the graph. Fix: log every N steps.
- CUDA-only dependencies on GPU Max. A library fails at import. Fix: audit requirements before porting.
- Device-count surprises. A Max 1550 system shows twice or half the expected devices because of how stacks are exposed. Fix: pick one device hierarchy for the fleet.
- RoCE fabric misconfiguration. Multi-server Gaudi jobs hang or crawl. Fix: run a collective benchmark first.
Trade-offs and how to choose
Gaudi's strengths are memory capacity and bandwidth per card, built-in Ethernet that lets clusters use standard switches, and good results on mainstream transformer models that Intel and Hugging Face have tuned. Its weakness is the flip side of graph compilation: dynamic shapes, unusual operators and fast-changing research code all fight the compiler. Teams running a known set of models at scale get the most from it.
GPU Max is the more forgiving target, since it behaves like a GPU and most code runs after a device-name change. Its weakness is its life cycle: keep existing deployments productive and code device-agnostic so moving later is cheap. In both cases, measure cost per token on your own model and framework version before committing.
What to do next
- Write down the model size, sequence lengths and batch shapes you need, and compute memory fit per card for training and serving as in the worked example.
- For Gaudi, install the documented software release, run the stock training loop with torch.compile and the hpu_backend, and log compile counts and CPU fallbacks during warm-up.
- Bucket or pad inputs until the compile count stops rising, then measure steady-state throughput.
- For GPU Max, remove IPEX dependencies, switch to upstream PyTorch XPU, and confirm the device count and hierarchy before writing launch scripts.
- Run a collective benchmark across all cards and servers before any multi-node job, and fix fabric problems first.
- Compare cost per token against your current platform on your own model, and keep the code device-agnostic so the decision stays reversible.