Nebius is a GPU cloud built for AI training and inference. The company behind it, Nebius Group N.V., is headquartered in Amsterdam and listed on Nasdaq as NBIS. It is the former Yandex N.V., renamed after the Russian businesses were sold in 2024. Some older pages call it an Eastern European cloud. That is wrong: its public regions are in Finland, France, Spain, the UK, Israel and the United States. Large capacity contracts, such as the multi-year Microsoft agreement announced in September 2025, are why it shows up in the news. They do not change how you use it.
This page is a runbook for an engineer who has to get a multi-node training job running on Nebius AI Cloud and keep it running. It covers the resource model, how regions map to InfiniBand fabrics and GPU platforms, building a GPU cluster with the CLI, how tenants are isolated on a shared fabric, Slurm through Soperator, storage, and the acceptance tests to run before a long job. For how it compares with other specialist clouds, read CoreWeave in depth and Lambda Labs in depth. The facts below were checked against the Nebius documentation on 7 October 2026. Platforms move between regions, so check again before you plan capacity.
The resource model
Four objects decide whether two GPUs can talk to each other at InfiniBand speed. A region is a location. An InfiniBand fabric is a physical network inside it. Fabrics have names such as fabric-7 or us-central1-b. A GPU cluster is a logical object created on one fabric. A VM is a virtual machine that joins the cluster. Each VM is created with a platform, the GPU and CPU type, and a preset, the GPU, vCPU and memory shape. All VMs in a GPU cluster must be in the same project as the cluster. A cluster lives on exactly one fabric, so its VMs all share one platform.
The documentation states that each GPU in a cluster VM has its own 400 Gbps InfiniBand NIC, so an 8-GPU VM has 3.2 Tbps of scale-out bandwidth. GPUDirect RDMA moves data between a GPU and its NIC without staging it through CPU memory. Inside the VM, the eight GPUs talk over NVLink. This is the same one-NIC-per-GPU pattern as the rail-aligned designs used in on-premises clusters, and NCCL uses it the same way.
Regions, fabrics and platforms
The documentation lists these InfiniBand fabrics and platforms, checked 7 October 2026. The platform ID is the value passed as --resources-platform.
| Region | Location | Fabric(s) | GPU platform |
|---|---|---|---|
eu-north1 | Finland | fabric-2, fabric-3, fabric-4, fabric-6 | H100, Sapphire Rapids (gpu-h100-sxm) |
eu-north1 | Finland | fabric-7 | H200, Sapphire Rapids (gpu-h200-sxm) |
eu-west1 | France | fabric-5 | H200, Sapphire Rapids |
eu-west2 | France | eu-west2-a | B300, Granite Rapids (gpu-b300-sxm) |
uk-south1 | United Kingdom | uk-south1-a | B300, Granite Rapids |
us-north1 | Minnesota, US | us-north1-a | B300, Granite Rapids |
me-west1 | Israel | me-west1-a | B200, Emerald Rapids (gpu-b200-sxm-a) |
us-central1 | Kansas City, US | us-central1-a | H200, Sapphire Rapids (gpu-h200-sxm) |
us-central1 | Kansas City, US | us-central1-b | B200, Emerald Rapids (gpu-b200-sxm) |
eu-north2 | Iceland (private region) | eu-north2-a | H200 |
Two details matter when you plan. The B200 platform ID differs between Israel and Kansas City, so do not hard-code one in scripts that run in both. Regions such as eu-south1 and uk-south2 offer RTX PRO 6000 VMs but list no InfiniBand fabric, so they suit inference and single-node work rather than multi-node training. The docs say that you usually do not need to change the preselected fabric, and that you should choose another fabric only for a specific reason, such as capacity.
Building a GPU cluster
The CLI sequence has three steps: create the cluster on a fabric, create a boot disk from a GPU image, and create VMs that join the cluster. The commands below follow the documentation's example. The H100 platform and its 8gpu-128vcpu-1600gb preset are the pairing the docs show. For other platforms, look up the presets that are compatible with GPU clusters instead of guessing.
export INFINIBAND_FABRIC=fabric-2
export GPU_CLUSTER_ID=$(nebius compute gpu-cluster create \
--name train-h100-a \
--infiniband-fabric $INFINIBAND_FABRIC \
--format jsonpath='{.metadata.id}')
for i in $(seq -w 1 4); do
BOOT_DISK_ID=$(nebius compute disk create \
--name train-h100-a-boot-$i \
--size-gibibytes 200 \
--type network_ssd \
--source-image-family-image-family ubuntu24.04-cuda13.0 \
--block-size-bytes 4096 \
--format jsonpath='{.metadata.id}')
nebius compute instance create \
--name train-h100-a-$i \
--resources-platform gpu-h100-sxm \
--resources-preset 8gpu-128vcpu-1600gb \
--gpu-cluster-id $GPU_CLUSTER_ID \
--boot-disk-existing-disk-id $BOOT_DISK_ID
# plus your network interface, SSH key and cloud-init flags
doneTreat this script as a sketch to wrap in Terraform or your own controller. It is not a deployment tool in itself. The loop has no retry, and a single failed VM leaves a half-built cluster. Delete in the reverse order: every VM must leave a GPU cluster before nebius compute gpu-cluster delete succeeds.
Isolation on a shared fabric
Fabrics are shared physical networks, so isolation between customers matters. Nebius gives each GPU cluster a unique InfiniBand partition key, or P-Key. InfiniBand switches drop traffic between ports that are not in the same partition. Nodes in different GPU clusters therefore cannot reach each other over InfiniBand, even when they sit on the same switches.
This has a practical side effect. Two GPU clusters in your own project are also isolated from each other. If you build a second cluster to add capacity, its VMs cannot join the first cluster's NCCL rings over InfiniBand. A job split across them falls back to slower paths or fails. Plan one GPU cluster per training job or per Slurm partition, and add VMs to it. Do not add clusters. Bandwidth is still shared at the physical level, so a neighbour's traffic can affect yours on shared switches. Measure under load. A quiet-hour benchmark is not enough.
Slurm through Soperator
Soperator is Nebius's open-source Kubernetes operator that runs Slurm. Managed Service for Soperator is the hosted version. Each Slurm node runs as a pod. Login pods run sshd behind a load balancer, worker pods run slurmd, and controller pods run slurmctld. The operator reconciles a SlurmCluster resource and keeps Slurm's configuration in ConfigMaps. Researchers get sbatch and srun. The platform team gets Kubernetes for upgrades and node replacement.
The feature that sets it apart is the shared root filesystem. Every login and worker node mounts the same root filesystem, backed by a Kubernetes PersistentVolume. Install a package on one node and it appears on all of them, so you no longer have to keep a hundred nodes identical by hand. Each node also has local SSD for ephemeral, node-specific data.
That convenience is also the main risk. A careless pip install or apt upgrade on a login node changes the environment of every running job at once. Version your environments as containers or as immutable directories, and make the shared root read-mostly by policy. Write datasets and checkpoints to object storage or a dedicated shared filesystem, not to the root. Use local SSD for scratch only.
#!/bin/bash
#SBATCH --job-name=llm-pretrain
#SBATCH --nodes=4
#SBATCH --gpus-per-node=8
#SBATCH --ntasks-per-node=8
#SBATCH --exclusive
export NCCL_DEBUG=INFO # first run only: confirms IB devices and GPUDirect
srun --container-image=/shared/images/train-2026-10.sqsh \
python train.py --ckpt s3://my-bucket/run-42/The --container-image flag comes from the Pyxis plugin. Check that your cluster image includes it before you rely on it. Otherwise, activate a pinned environment directory instead.
Storage for training
Training jobs read in three patterns, and each needs a different kind of storage. Datasets and checkpoints belong in Object Storage, the S3-compatible service, or in a shared filesystem when the framework expects POSIX paths. Hot shards and dataloader caches belong on local SSD, staged at job start. Code and environments belong on the shared root or in container images.
Checkpoints set the size requirement. A 70B-parameter model with Adam optimizer state is about 16 bytes per parameter, so roughly 1.1 TB per checkpoint. Writing that every 30 minutes over a 4-node job needs sustained write bandwidth that you should measure before the run, not discover during it. Shard the checkpoint so that every rank writes in parallel, and keep the previous checkpoint until the new one is verified.
Work through the numbers once. If the 32 ranks together sustain 5 GB/s to object storage, 1.1 TB takes about 220 seconds. A synchronous save every 30 minutes then costs about 12% of the run with every GPU idle. Two fixes are standard. Copy the state to host memory and upload it in the background, so the GPUs stall only for the copy. Or write to local SSD first and drain to object storage while training continues. Either way, measure the real upload rate from your VMs to your bucket before you set the interval. Write the checkpoint manifest last, so a half-written checkpoint is never mistaken for a complete one.
Acceptance testing
Run acceptance tests every time a cluster is created or grown, and after every node replacement. Each step checks something the next one depends on.
# 1. Per node: 8 GPUs, NVLink topology, 8 active IB ports
nvidia-smi --query-gpu=index,name,ecc.errors.uncorrected.volatile.total --format=csv
nvidia-smi topo -m
ibv_devinfo | grep -E "hca_id|state|active_mtu"
# 2. Single node: NVLink all-reduce (nccl-tests)
./build/all_reduce_perf -b 8 -e 8G -f 2 -g 8
# 3. Multi node: IB all-reduce across every VM in the cluster
srun -N 4 --ntasks-per-node=8 --gpus-per-node=8 \
./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1Compare the bus bandwidth at large message sizes with the hardware ceiling. Eight 400 Gbps NICs, the figure the docs give, make 400 GB/s per node. Record the first clean run as your baseline and flag any node that falls well below it. A rule of thumb is to investigate anything more than 10% to 15% below the baseline. Run the test once with NCCL_DEBUG=INFO and confirm that the log names the InfiniBand devices and GPUDirect RDMA. A log that shows only socket transport means the VMs are not in the same GPU cluster. Bisect a slow multi-node result by running pairs of nodes. One bad NIC or cable slows the whole ring.
Failure modes
- VMs in different GPU clusters. NCCL falls back to TCP, and throughput drops by an order of magnitude. Check the GPU cluster ID on every VM.
- Capacity on the fabric. A fabric can run out of the platform you need. The docs suggest another fabric for capacity reasons, but that means a new cluster. You cannot grow an existing job onto it.
- Shared-root drift. One package change reaches every node. Jobs that worked yesterday fail with import or ABI errors.
- Hard-coded platform IDs.
gpu-b200-sxmversusgpu-b200-sxm-abreaks scripts that move between regions. - Checkpoint stalls. Synchronous checkpoints to a slow target leave every GPU idle. Make them asynchronous, or size the target first.
- Silent slow nodes. One degraded GPU or link sets the step time for all of them. Run acceptance tests after every replacement, and watch per-rank step time.
Trade-offs
Nebius sits between a hyperscaler and bare metal. You get InfiniBand clusters, a managed Slurm that feels like an HPC centre, and Kubernetes for everything else, without building a data centre. You give up the breadth of managed services, regional coverage and procurement familiarity that a hyperscaler offers. Your job portability also depends on keeping your tooling cloud-neutral. The questions in bare metal versus cloud GPUs apply here. Before signing a long commitment, ask how much capacity on your fabric is committed to you, what replacement times look like, and what happens to a running job when a node fails.
Keep the training stack portable: containers, Slurm or Kubernetes manifests, and S3-compatible storage paths. Then a second provider stays a configuration change, as described in running on multiple GPU providers.
What to do next
- Pick the region and fabric from the current documentation table. Record the platform ID and a GPU-cluster-compatible preset for it.
- Create one GPU cluster per training job or Slurm partition, and add VMs to it with Terraform or a controller that retries and cleans up.
- Run the three acceptance steps and save the bus-bandwidth baseline per cluster.
- On Soperator, ship environments as container images and treat the shared root as read-mostly.
- Measure checkpoint write bandwidth to your target storage and set the checkpoint interval from it.
- Keep manifests and storage paths portable so that a second provider stays an option.