OCI SuperCluster is Oracle Cloud Infrastructure's name for large groups of bare-metal GPU servers joined by a dedicated RDMA network so that thousands of GPUs can train one model. It is not a single product you click once. It is a set of building blocks: bare-metal GPU shapes, an RDMA cluster network separate from the normal virtual cloud network, placement constructs that keep nodes physically close, and storage tiers for datasets and checkpoints. Oracle's marketing pages quote per-cluster ceilings such as 131,072 B200 GPUs or more than 100,000 GB200 superchips, and latency as low as 2.5 microseconds; treat those as vendor figures and plan from what your tenancy is actually allocated.
This article explains each building block in terms of what your training code and scheduler see: which network a collective uses, how much bandwidth each GPU gets, how nodes are grouped and replaced, where checkpoints go, and how to prove a new cluster is healthy before you spend a month of GPU time on it.
What a SuperCluster changes
Distributed training is limited by the slowest collective. A data-parallel step ends with an all-reduce of gradients; tensor and pipeline parallelism add all-gathers and point-to-point transfers inside every layer. Ordinary cloud networking, built for many independent flows through virtualised NICs, adds too much latency and jitter for that. A SuperCluster changes three things at once:
- Bare metal. No hypervisor sits between the GPUs, the NICs and your process, so GPUDirect RDMA can move data from GPU memory to the wire without a host copy.
- A second network. Each node has front-end NICs on the VCN for SSH, storage and control traffic, and separate RDMA NICs on the cluster network that only carry GPU-to-GPU traffic over RoCE version 2.
- Placement. Nodes in one cluster are allocated physically close in one availability domain, which OCI documents as giving RDMA latency as low as single-digit microseconds.
The bare-metal shapes
The unit of allocation is a bare-metal shape. The figures below come from OCI's compute shapes reference; GPU memory is the total per node.
| Shape | GPUs | GPU memory | Host memory | Local NVMe | Front-end | RDMA cluster network |
|---|---|---|---|---|---|---|
| BM.GPU.H100.8 | 8 x H100 | 640 GB | 2048 GB | 16 x 3.84 TB | 1 x 100 Gbps | 8 x 2 x 200 Gbps |
| BM.GPU.H200.8 | 8 x H200 | 1128 GB | 3072 GB | 8 x 3.84 TB | 1 x 200 Gbps | 8 x 400 Gbps |
| BM.GPU.B200.8 | 8 x B200 | 1440 GB | 4096 GB | 8 x 3.84 TB | 2 x 200 Gbps | 8 x 400 Gbps |
| BM.GPU.B300.8 | 8 x B300 | 2100 GB | 4096 GB | 8 x 3.84 TB | 2 x 200 Gbps | 8 x 800 Gbps |
| BM.GPU.GB200.4 | 4 x B200 (Grace host) | 768 GB | 960 GB | 4 x 7.68 TB | 2 x 200 Gbps | 4 x 400 Gbps |
Read the last column per GPU, because that is how collectives use it. H100, H200, B200 and GB200 nodes all give each GPU 400 Gbps of RDMA bandwidth, about 50 GB/s. B300 doubles that to 800 Gbps. Inside a node, GPUs talk over NVLink at hundreds of GB/s, so the scale-out link is an order of magnitude slower than the scale-up one. A B200 does much more arithmetic per second than an H100 yet gets the same per-GPU RDMA bandwidth, so communication that hid behind compute on H100 can become exposed on B200. Plan the parallel layout for that ratio, not for the GPU name. GB200 systems are built around a rack-scale NVLink domain; check how OCI exposes that domain to your scheduler before you assume tensor parallelism can span hosts.
The RDMA cluster network
The cluster network carries RDMA over Converged Ethernet version 2: InfiniBand transport semantics inside UDP/IP packets on an Ethernet fabric. NCCL drives it through the same verbs API it uses for InfiniBand, so the code path is the IB transport with a GID that selects RoCEv2. RoCE's reliable-connected transport recovers from loss with go-back-N retransmission, which is expensive, so the fabric depends on congestion control (ECN marking and rate reduction) to keep loss rare. The RoCE deep dive covers queue pairs, timeouts and counters in detail.
Three consequences for software. First, NCCL must be told which NICs are RDMA devices and which GID index and traffic class to use; if it falls back to TCP sockets over the front-end NIC, jobs still run, only far slower, and nothing errors. Second, the front-end NIC is shared with storage and control traffic, so never let checkpoints and collectives compete on it. Third, the right NCCL settings are shape-specific: start from the values in Oracle's published HPC and GPU cluster stack for your shape rather than copying a blog post, and verify them with a bandwidth test.
# Inspect what the node actually has before tuning anything
ibv_devinfo | grep -E "hca_id|state|link_layer" # RDMA devices, PORT_ACTIVE, Ethernet
show_gids | grep -i v2 # RoCEv2 GID entries per device
nvidia-smi topo -m # which NIC sits closest to which GPU
# Then run NCCL with debug output and confirm the transport
export NCCL_DEBUG=INFO
export NCCL_IB_HCA=<comma-separated RDMA device list for this shape>
export NCCL_SOCKET_IFNAME=<front-end interface, used only for bootstrap>
# Look for "NET/IB" with the RDMA devices in the log; "NET/Socket" means TCP fallback.
Cluster networks versus compute clusters
OCI gives you two ways to put nodes on the same RDMA network, and the choice changes how you operate the fleet.
| Cluster network | Compute cluster | |
|---|---|---|
| Built on | An instance pool with one instance configuration | An empty RDMA network group you add instances to |
| Instances | Identical, created and resized together | Managed individually; may differ |
| Grow or shrink | Resize the underlying instance pool | Launch or terminate single instances into the cluster |
| Fits | Fixed-size training fleets | Schedulers that replace nodes one at a time |
Oracle's documentation states the rule directly: use compute clusters if you want to manage instances in the RDMA network independently or mix instance types. For long training runs that matters, because the most common operation after day one is replacing a single unhealthy node without disturbing the rest. Instances must be in the same compartment and availability domain as the compute cluster, and GPU capacity normally requires a service limit increase or a capacity reservation, so request it before you plan dates. A Terraform sketch:
resource "oci_core_compute_cluster" "train" {
compartment_id = var.compartment_id
availability_domain = var.ad
display_name = "llm-train"
}
resource "oci_core_instance" "gpu" {
count = var.node_count
compartment_id = var.compartment_id
availability_domain = var.ad # must match the cluster
compute_cluster_id = oci_core_compute_cluster.train.id
shape = "BM.GPU.H100.8"
display_name = "gpu-${count.index}"
source_details {
source_type = "image"
source_id = var.gpu_image_id # image with GPU and RDMA drivers
}
create_vnic_details { subnet_id = var.frontend_subnet_id }
}
Worked example: the all-reduce budget
Suppose you train an 8-billion-parameter model with plain data parallelism on 64 BM.GPU.H100.8 nodes, 512 GPUs, and all-reduce bf16 gradients every step. The gradient buffer S is 8e9 x 2 bytes = 16 GB.
- NCCL reduces inside each node over NVLink first, so each of the 8 GPUs ends up responsible for one eighth of the buffer, 2 GB.
- Each GPU then runs a ring all-reduce of its 2 GB slice with its peers on the other 63 nodes, over its own 400 Gbps of RDMA bandwidth. A ring all-reduce sends 2 x (N - 1) / N times the buffer, so about 2 x 63/64 x 2 GB = 3.94 GB per GPU.
- At 50 GB/s that is about 79 ms of pure transfer. Real collectives reach perhaps 75 to 85 percent of line rate at this size, so budget about 100 ms.
- If the compute for one step takes 1.5 s and the all-reduce overlaps with the backward pass, communication is hidden. Move the same job to B200, where compute might take roughly half the time while per-GPU RDMA bandwidth is unchanged, and the margin shrinks; the fix is fewer bytes (sharded optimisers, gradient accumulation, overlapping buckets), not more nodes.
This arithmetic is the most useful thing you can do before choosing a shape. The training cluster design guide walks through the full budget, and the NCCL all-reduce article explains the ring and tree algorithms behind the 2(N - 1)/N factor.
Storage for datasets and checkpoints
Storage has three tiers. Local NVMe (sixteen 3.84 TB drives on an H100 node) is the fastest place to write a checkpoint and to cache a dataset shard, but it dies with the node. A shared file system reachable over the front-end network holds datasets and the checkpoints you resume from. Object storage holds durable copies. A robust checkpoint path writes to local NVMe so the training step can resume quickly, then copies asynchronously to shared storage and object storage, and only marks a checkpoint valid once every rank's shard has landed. Measure aggregate front-end bandwidth: 512 GPUs restoring a multi-terabyte checkpoint at the same moment is the load spike that sizes your storage, not steady-state reading of training data.
Bring-up and acceptance testing
Never start a long run on a cluster you have not tested. Acceptance has two layers. Per node: every GPU visible, no retired-page or ECC errors, DCGM diagnostics pass, every RDMA port active with the expected link speed. Across nodes: NCCL tests reach expected bus bandwidth, first between pairs to find bad nodes or links, then at full scale.
# per node (run everywhere through the scheduler)
nvidia-smi --query-gpu=index,name,ecc.errors.uncorrected.volatile.total --format=csv
dcgmi diag -r 2
ibv_devinfo | grep -c PORT_ACTIVE # expect one per RDMA port on this shape
# across nodes, with nccl-tests built against the image's NCCL
srun -N 2 --ntasks-per-node=8 --gpus-per-node=8 \
./build/all_reduce_perf -b 1G -e 8G -f 2 -g 1 # every pair, find slow links
srun -N 64 --ntasks-per-node=8 --gpus-per-node=8 \
./build/all_reduce_perf -b 8M -e 8G -f 2 -g 1 # full scale, record busbwRecord the full-scale bus bandwidth as the baseline and rerun the pairwise test after every node replacement. A single slow optic or misconfigured port shows up as one node pair far below the others, and it caps the whole ring. Pair this with a scheduler that knows node placement; the Slurm GPU scheduling article covers topology-aware allocation.
Failure modes
- Silent TCP fallback. NCCL picks sockets on the front-end NIC; throughput drops several-fold with no error. Fail the job if the NCCL log lacks NET/IB.
- One bad link. A degraded port or optic stalls every collective in the ring. Pairwise tests and per-port error counters find it.
- Node failure mid-run. In a long run on hundreds of nodes, hardware faults are routine. Keep spare nodes allocated, checkpoint on a measured interval, and automate replace-and-resume.
- Capacity assumptions. Shape availability differs by region and availability domain, and service limits start low. Planning around capacity you have not reserved is the most expensive mistake.
- Driver drift. Mixed images across replaced nodes give mismatched NCCL, OFED or CUDA versions and odd hangs. Pin the image and check versions at job start.
- Checkpoint storms. All ranks writing to shared storage at once saturates the front-end network. Stage on local NVMe and copy asynchronously.
Trade-offs
Cluster networks are simpler when the fleet is fixed and homogeneous; compute clusters cost more scripting but make single-node replacement routine. Larger per-node GPU memory (H200, B200, B300) reduces model sharding and therefore communication; newer GPUs with unchanged per-GPU network bandwidth increase the pressure on overlap. RoCE on Ethernet gives cloud-scale fabrics at lower cost than InfiniBand, in exchange for depending on well-tuned congestion control. And bare metal gives full performance with full responsibility: you own the image, the drivers and the health checks.
What to do next
- Compute your per-GPU RDMA bandwidth and the all-reduce time for your model on each candidate shape before choosing one.
- Request service limits or reserved capacity for the exact shape and availability domain.
- Choose a compute cluster if you expect to replace nodes individually; a cluster network if not.
- Build or adopt one GPU image and pin NCCL, CUDA and driver versions in it.
- Run per-node health checks and pairwise then full-scale NCCL tests; keep the baseline.
- Make the job fail fast if NCCL is not using the RDMA devices.
- Stage checkpoints on local NVMe and copy them asynchronously to shared and object storage.