Rubin is the NVIDIA data centre GPU generation after Blackwell, launched at CES in January 2026 with partner systems scheduled for the second half of 2026. It is easy to read about it as a list of bigger numbers. This article instead asks what an engineer who trains or serves models should do differently on Rubin: which bottlenecks move, which parallelism layouts become attractive, what the precision story means for training recipes, and how to get software ready before hardware arrives.
A note on sources. Every specification below comes from NVIDIA's CES announcement or its technical blog on the Rubin GPU, checked in October 2026, and is labelled when it comes from press coverage instead. Figures marked derived are our own arithmetic. Early roadmap material from 2025 used different names and some lower figures, so if you are comparing documents, check the date. Where something is not published, such as final clocks or per-SKU power, we say so rather than guess.
What Rubin is
The Rubin platform is six chips designed together: the Rubin GPU, the Vera CPU, the NVLink 6 Switch, the ConnectX-9 SuperNIC, the BlueField-4 DPU and the Spectrum-6 Ethernet switch. The flagship system is the Vera Rubin NVL72 rack: 72 Rubin GPUs and 36 Vera CPUs in one NVLink domain. In 2025 roadmap slides the same rack was called NVL144, counting GPU dies, since each Rubin package holds two; the shipping name counts packages, like GB200 NVL72. For x86 servers there is HGX Rubin NVL8, eight GPUs on an NVLink board.
| Per GPU | Blackwell generation | Rubin | Why it matters |
|---|---|---|---|
| Compute dies | 2 | 2, reticle limited, one package | software still sees one GPU |
| Transistors / SMs | 208B / not comparable | 336B / 224 SMs, 896 Tensor Cores | more parallel work per kernel launch |
| Low-precision peak | NVFP4 on Blackwell | 50 PF NVFP4 inference; 35 PF training (reported) | FP4 becomes a training format, not only serving |
| Memory | up to 288 GB HBM3e, about 8 TB/s | 288 GB HBM4, up to 22 TB/s | 2.8x bandwidth; decode speeds up most |
| Scale-up link | NVLink 5, 1.8 TB/s | NVLink 6, 3.6 TB/s | all-to-all and TP inside the rack |
| CPU link | NVLink-C2C 900 GB/s | NVLink-C2C 1.8 TB/s | offload and KV spill to CPU memory |
Blackwell column figures are for B200 and B300 class parts and vary by SKU; treat them as orientation. NVIDIA also states a 4x increase in exponential-function throughput per clock per SM for BF16 and FP16 over Blackwell. That speeds the softmax in attention, not general BF16 matrix multiplication, which is a common misreading.
Where the bottleneck moves
The roofline model says a kernel is compute-bound when its arithmetic intensity, FLOPs per byte moved from memory, exceeds peak FLOPs divided by memory bandwidth. For Rubin at NVFP4 that ridge point is about 50e15 / 22e12, roughly 2,300 FLOPs per byte (derived). Large training matmuls easily clear it. Autoregressive decode does not: at batch size 1, every weight is read once per token and used for about two FLOPs, far below the ridge. Decode is therefore governed by bandwidth, which is why the 2.8x memory bandwidth increase matters more for serving latency than the headline FLOPs. The background on stacked memory is in the GPU HBM architecture guide.
def decode_floor_ms(params_b, bits_per_weight, hbm_tb_s, kv_gb=0.0):
"""Lower bound on one decode step: read weights (and KV cache) once."""
weight_gb = params_b * bits_per_weight / 8
return (weight_gb + kv_gb) / (hbm_tb_s * 1000) * 1000
# 70B dense model, NVFP4: 4-bit values plus one FP8 scale per 16 values = 4.5 bits
for name, bw in [("Blackwell-class", 8.0), ("Rubin", 22.0)]:
print(name, round(decode_floor_ms(70, 4.5, bw, kv_gb=10), 2), "ms")
# Blackwell-class 6.17 ms, Rubin 2.24 ms per step, before any kernel overheadsReal kernels reach perhaps 70 to 85 percent of peak bandwidth, so the measured gain will be smaller than the ratio, and batching shifts the balance back toward compute. The useful habit is to compute this floor for your own model before reading anyone's benchmark.
The rack is the unit
NVLink 6 doubles per-GPU scale-up bandwidth to 3.6 TB/s, and the switch trays provide 260 TB/s across the rack with in-network compute for collectives. NVLink 5's 1.8 TB/s per GPU is a total of both directions; we assume NVLink 6's figure is counted the same way, so budget roughly half per direction when estimating transfer times. Scale-out uses ConnectX-9 SuperNICs on Spectrum-X Ethernet or InfiniBand. How the switch trays work is covered in the NVLink Switch article.
The 72-GPU domain itself is not new: GB200 and GB300 NVL72 introduced it. What Rubin changes is how much each GPU can push through it and how much memory sits behind it, about 20 TB of HBM4 per rack (72 x 288 GB, derived). Tensor parallelism, expert parallelism and context parallelism, the strategies that issue frequent, latency-sensitive collectives, belong inside it. Data and pipeline parallelism, which tolerate the slower network, go across racks. Teams moving straight from Hopper-era eight-GPU nodes get the bigger shift: expert parallelism for large mixture-of-experts models no longer has to stop at eight GPUs.
The CPU side matters more than it used to. Each Vera CPU, with 88 custom Olympus cores and full Armv9.2 compatibility, connects to its two GPUs over NVLink-C2C at 1.8 TB/s, twice the Grace generation. That link is fast enough to treat CPU memory as a second tier for optimizer state offload or for KV cache that does not fit in HBM, and it means data loading, tokenization and checkpoint serialization now run on Arm cores close to the GPUs. NVIDIA also announced an inference context memory storage platform built on the BlueField-4 processor, aimed at sharing KV cache across agentic and long-context serving; treat it as a product to evaluate, not an architecture to assume. Finally, Vera Rubin NVL72 is described as the first rack-scale platform with NVIDIA Confidential Computing spanning CPU, GPU and NVLink, which matters if you serve regulated tenants or proprietary weights on shared infrastructure.
Worked example: MoE all-to-all inside the rack
Take a mixture-of-experts layer with hidden size 7,168, top-8 routing, FP8 dispatch and 8,192 tokens per GPU per micro-batch, with experts spread across all 72 GPUs. Each GPU sends every token to 8 experts: 8,192 x 8 x 7,168 bytes, about 470 MB out, and the combine step brings about as much back. At roughly 1.8 TB/s per direction that is about 0.26 ms each way, or around 0.75 ms per layer once you allow 70 percent link efficiency (all derived). The same exchange on a GB200 NVL72, at roughly 0.9 TB/s per direction, takes about 0.52 ms each way and about 1.5 ms per layer. Across 60 MoE layers that is the difference between roughly 45 ms and 90 ms of exchange per forward pass, which overlap with expert compute can hide only if compute is long enough. With Rubin's compute also several times faster, the exchange becomes a larger share of the step, so the doubled link is what keeps the faster maths usable. NVIDIA's own claim is up to 4x fewer GPUs to train large MoE models than Blackwell; it combines link, compute and memory gains, and we do not try to decompose it. On an eight-GPU node whose experts spill onto the scale-out network, the exchange would be an order of magnitude slower still, which is the shift teams leaving Hopper-era layouts will feel most. Routing and the all-to-all collective are explained in MoE all-to-all communication.
Two cautions. Load imbalance between experts still sets the step time, because the slowest GPU finishes the exchange last; faster links do not fix a hot expert. And the dispatch kernel must actually use the link: libraries built on NCCL's device-initiated communication, or on vendor all-to-all kernels, will matter more than raw bandwidth.
NVFP4 and the Transformer Engine
NVFP4 stores values as 4-bit floats in blocks of 16 that share an FP8 scale, with a second, per-tensor FP32 scale on top. The fine block size is what makes 4-bit usable: an outlier only distorts its own 16 neighbours. Rubin's third-generation Transformer Engine adds what NVIDIA calls hardware-accelerated adaptive compression, aimed at NVFP4 inference throughput; NVIDIA has not published enough detail to model it, so do not budget for it until you have measured it.
For training, the reported 35 PF NVFP4 figure is lower than the 50 PF inference figure, and a practical FP4 recipe keeps sensitive parts in higher precision. NVIDIA's paper Pretraining Large Language Models with NVFP4 (arXiv 2509.25149) trained a 12B model on 10T tokens to loss and accuracy close to an FP8 baseline using four techniques: keeping numerically sensitive layers in higher precision, 2D block scaling of weights so the forward and backward passes see the same quantized values, random Hadamard transforms on the weight-gradient GEMM inputs, and stochastic rounding of gradients. It also showed that switching to BF16 late in training, during the learning-rate decay, closes any remaining loss gap. As a starting plan:
precision_plan = {
"linear_fprop": "nvfp4", # weights and activations, 16-value blocks
"weights": "2D block scaling, same quantized values fwd and bwd",
"linear_dgrad": "nvfp4, stochastic rounding on gradients",
"linear_wgrad": "nvfp4, random Hadamard transform on inputs, stochastic rounding",
"sensitive_layers": "bf16", # the paper kept some first and last blocks wide
"optimizer_state": "fp32 master weights",
"late_training": "option: switch to bf16 during lr decay if a loss gap remains",
}Validate any such plan against a BF16 baseline on a small model first, comparing loss curves and downstream evals, not just final loss. The general mixed-precision mechanics are in mixed precision training.
Getting software ready
Most teams will reach Rubin through a cloud provider, which makes readiness mostly a software exercise you can start today. Vera is an Arm CPU, like Grace, so any custom C++ extensions, data loaders or tokenizers must build for aarch64; the porting checklist in the GB200 tray guide applies unchanged. Minimum CUDA, driver, NCCL and framework versions are set by those projects' release notes, not by marketing pages; read them for the exact versions that list Rubin support. Avoid hard-coding device names or compute capability numbers; detect features instead.
import torch
def describe_device(i=0):
p = torch.cuda.get_device_properties(i)
major, minor = torch.cuda.get_device_capability(i)
return {
"name": p.name,
"sm_count": p.multi_processor_count,
"mem_gb": round(p.total_memory / 2**30),
"capability": f"{major}.{minor}",
}
info = describe_device()
print(info)
# choose kernels by probing, not by name: try the FP4 path on a tiny GEMM,
# compare against a BF16 reference, and fall back if it is missing or inaccurateThen rebuild your performance model with Rubin's ratios: compute your decode floor, your MoE all-to-all time and your checkpoint write time, which grows with the larger memory per rack. The second-generation RAS engine and hot-swappable switch trays improve serviceability, but a 72-GPU domain is also a larger blast radius: one failed GPU can stall a whole expert-parallel group.
Failure modes
- Benchmarking the wrong regime. Quoting FP4 peak for a decode-bound service. Compute the bandwidth floor first.
- Eight-GPU habits. Keeping expert or tensor parallel groups at eight and sending the rest over the network wastes the NVLink domain.
- FP4 without a baseline. Training diverges late, or evals regress while loss looks fine. Always run a BF16 control.
- Arm surprises. An x86-only wheel or intrinsic breaks the data loader on Vera hosts. Build and test aarch64 images now.
- Large failure domain. Jobs that span the whole rack stop when one GPU or switch tray fails. Plan checkpoint frequency and spare capacity for it.
- Stale numbers. Capacity plans built from 2025 roadmap figures. Re-check against current NVIDIA material and measured results.
Trade-offs
The NVL72 rack gives the largest fast domain but is a liquid-cooled, rack-scale purchase with facility requirements; HGX Rubin NVL8 fits conventional x86 servers but caps the NVLink domain at eight GPUs. NVFP4 roughly doubles throughput over FP8 in principle, but costs engineering effort and validation, and some models will not tolerate it. Larger domains reduce communication time but enlarge the blast radius of a failure. Being early buys performance per dollar at the cost of immature kernels and libraries; teams with custom kernels should budget time for tuning on the new architecture.
What to do next
- Compute the decode floor and roofline position of your main serving model with the function above, using Rubin and your current GPU figures.
- Estimate your MoE all-to-all time per layer for an expert parallel group of 72 versus your current group size.
- Run a small NVFP4 training experiment on Blackwell hardware today against a BF16 baseline to learn where your model needs higher precision.
- Build and test aarch64 container images for every training and serving component.
- Replace device-name checks in your code with feature probes and accuracy checks.
- Read the CUDA, driver, NCCL and framework release notes for the versions that list Rubin support, and pin them in a test environment.
- Revisit checkpoint interval and spare capacity planning for a 72-GPU failure domain.