Renting a GPU on Oracle Cloud Infrastructure looks like any other compute launch: pick a shape, an image and a subnet. The differences show up afterwards. GPU shapes need explicit service-limit increases, the large ones are bare metal machines with no hypervisor between you and the hardware, multi-node training needs a separate RDMA network that only some shapes and images support, and the fast local NVMe disappears when the instance does.

This article covers GPU instances specifically. General shape mechanics are in OCI compute shapes and the bare metal model in OCI bare metal. Shape figures below were checked against Oracle's compute shapes reference on 2026-10-01; OCI adds shapes often, so treat the table as a snapshot and list shapes from the API before you plan.

Advertisement

What you are renting

OCI GPU shapes follow a naming pattern: VM. or BM. for virtual machine or bare metal, then the GPU family, then the GPU count. VM shapes give you a slice of a host with one or two GPUs and block storage only. BM shapes give you the whole server, every GPU in it, large local NVMe and, on the training-class shapes, RDMA network interfaces for a cluster network.

ShapeGPUsGPU memory (total)Local diskNetwork (per Oracle docs)
VM.GPU.A10.11 x A1024 GBBlock storage only24 Gbps
BM.GPU.A10.44 x A1096 GB7.68 TB NVMe2 x 50 Gbps
BM.GPU.L40S.44 x L40S192 GB2 x 3.84 TB NVMe1 x 200 Gbps
BM.GPU4.88 x A100320 GB27.2 TB NVMe1 x 50 Gbps + 8 x 200 Gbps RDMA
BM.GPU.A100-v2.88 x A100640 GB27.2 TB NVMe2 x 50 Gbps + 16 x 100 Gbps RDMA
BM.GPU.H100.88 x H100640 GB16 x 3.84 TB NVMe1 x 100 Gbps + 8 x 2 x 200 Gbps RDMA
BM.GPU.H200.88 x H2001128 GB8 x 3.84 TB NVMe1 x 200 Gbps + 8 x 400 Gbps RDMA
BM.GPU.MI300X.88 x AMD MI300X1536 GB8 x 3.84 TB NVMe1 x 100 Gbps + 8 x 400 Gbps RDMA

The reference also lists newer Blackwell-generation shapes and older P100 and V100 shapes; check it for current figures. Read the network column as two networks: the front-end figure is for SSH, storage and the internet, and the RDMA figure is the cluster network that carries GPU-to-GPU traffic between nodes.

Choosing a shape by workload

Start from memory, not from compute. A model's weights, its KV cache or optimiser state must fit, with headroom, before speed matters.

  • Inference of small and mid-size models (about 7B parameters in 16-bit on one A10's 24 GB, 13B on VM.GPU.A10.2 or when quantised): VM.GPU.A10.1 or VM.GPU.A10.2. Cheap, quick to get, and easy to scale horizontally behind a load balancer.
  • Inference of larger models, image generation and graphics: BM.GPU.L40S.4, with 48 GB per GPU. It has no NVLink between GPUs, so tensor parallelism across its four GPUs pays a PCIe cost; it suits one model replica per GPU better than one model across all four.
  • Fine-tuning and single-node training: an 8-GPU A100 or H100 shape. NVLink inside the node makes 8-way tensor or sharded data parallelism efficient.
  • Multi-node pre-training or large fine-tunes: H100 or H200 shapes in a cluster network. Communication between nodes now sets your scaling efficiency.
  • Very large models that must fit on one node: BM.GPU.MI300X.8 or BM.GPU.H200.8 for their memory. MI300X runs on AMD's ROCm stack, so confirm your framework versions and kernels support it before you commit.
Advertisement

Getting capacity

New tenancies typically have a GPU service limit of zero or close to it. Limits are typically scoped by shape and often by availability domain, so check the Limits page in the console for the exact shape and AD rather than assuming a raise elsewhere carries over. Request increases early, for the specific shape and AD you plan to use, and expect the large bare metal shapes to need a conversation about capacity rather than a form. Oracle's cluster-network documentation notes that creating multiple GPU instances in a cluster network typically requires a service-limit increase.

Before writing any automation, list what your tenancy can actually see. This uses the Python SDK and works per availability domain:

import oci

config = oci.config.from_file()                 # ~/.oci/config
compute = oci.core.ComputeClient(config)
compartment = config["tenancy"]                 # or your project compartment OCID

for ad in oci.identity.IdentityClient(config).list_availability_domains(compartment).data:
    shapes = oci.pagination.list_call_get_all_results(
        compute.list_shapes, compartment, availability_domain=ad.name).data
    for s in shapes:
        if s.gpus:                              # GPU count is 0 or None for CPU shapes
            print(f"{ad.name:28} {s.shape:22} gpus={s.gpus} {s.gpu_description}")

If a shape is missing from the list, you either lack a limit or the region does not offer it. Both are worth knowing before an incident, not during one.

Images, drivers and first boot

A GPU is useless without a driver matching your CUDA or ROCm userland. You have three options: an Oracle-provided platform image with GPU drivers preinstalled, the Oracle Linux HPC cluster networking image from the Marketplace (which Oracle's documentation specifies for cluster networks), or your own custom image built from either. Bare installs of a generic image work but turn every launch into a driver install, which is slow and a source of version drift.

Pin one image OCID per environment, record the driver and CUDA versions it contains, and rebuild deliberately. A minimal launch and smoke test for a single A10 looks like this:

oci compute instance launch \
  --availability-domain "$AD" \
  --compartment-id "$COMPARTMENT" \
  --shape VM.GPU.A10.1 \
  --subnet-id "$PRIVATE_SUBNET" \
  --image-id "$GPU_IMAGE_OCID" \
  --boot-volume-size-in-gbs 200 \
  --ssh-authorized-keys-file ~/.ssh/id_ed25519.pub \
  --display-name a10-inference-01

# on the instance
nvidia-smi --query-gpu=name,memory.total,driver_version --format=csv
python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"

Put GPU instances in a private subnet behind a bastion; there is rarely a reason for a training node to have a public IP. Network design is in OCI networking.

Multi-node: cluster networks and compute clusters

Once a job spans nodes, collective operations such as all-reduce cross the network on every step. OCI provides this through a cluster network: in Oracle's words, a group of identical bare metal GPU or HPC instances in close physical proximity, connected by an RDMA network with latency as low as single-digit microseconds.

A cluster network is built on instance pools. You create an instance configuration (shape, image, SSH keys, subnet), then a cluster network that launches a pool from it. At the time of checking, Oracle's page listed BM.GPU.A100-v2.8, BM.GPU.H100.8 and BM.GPU4.8 among supported shapes; check it for newer GPU shapes. If you need to manage RDMA-connected instances individually, or mix instance types, the same page points to compute clusters instead. Pools suit identical, disposable nodes; compute clusters suit long-lived nodes you replace one at a time.

A two-node GPU training cluster on OCI: front-end network for control, RDMA cluster network for collectivesBastion / loginVM, public subnetObject Storagedatasets, checkpointsFile Storage / FSSshared code, envInstance configshape, image, keysBM GPU node 18 GPUs, NVLink, local NVMeBM GPU node 28 GPUs, NVLink, local NVMesshstage datamountpoolCluster network: RDMA over converged EthernetGPU-to-GPU across nodes, NCCL all-reduceFront-end VNIC traffic (ssh, storage, bootstrap) and RDMA traffic (collectives) use different interfaces; NCCL must be told which is which
Two bare metal GPU nodes in a cluster network. Control, storage and NCCL bootstrap use the front-end VNIC; all-reduce traffic uses the RDMA interfaces.

The RDMA network is RoCE, RDMA over Ethernet, and NCCL must be pointed at it. If it is not, NCCL falls back to TCP sockets over the front-end interface and your job runs, many times slower, with no error. Discover the device names on the host rather than copying them from a blog post, and prove bandwidth with nccl-tests before you start real work:

# Run on each node before launching training. Discover, do not copy values from another cluster.
ibv_devices                         # RDMA devices (mlx5_*) attached to the cluster network
ip -br addr                          # front-end interface for bootstrap traffic
rdma link show                       # every RDMA link should report state ACTIVE

export NCCL_SOCKET_IFNAME=<front-end interface>   # bootstrap only
export NCCL_IB_HCA=<comma-separated RDMA devices from ibv_devices>
export NCCL_IB_GID_INDEX=<RoCE v2 GID index for your image>   # see your image's docs or show_gids
export NCCL_DEBUG=INFO               # first run only: confirms which transport NCCL picked

# Bandwidth check across two nodes with nccl-tests before any real job
mpirun -np 16 -H node1:8,node2:8 ./build/all_reduce_perf -b 1G -e 8G -f 2 -g 1

Compare the bus bandwidth that all_reduce_perf reports against what the RDMA interfaces should sustain; how the collective uses that bandwidth is explained in NCCL collectives. A result an order of magnitude low almost always means the socket transport was chosen.

Storage for datasets and checkpoints

Local NVMe on bare metal GPU shapes is fast and large, and it is ephemeral: terminate the instance, or let a pool replace it, and the data is gone. Use it as a cache and scratch space, never as the only copy of anything.

  • Object Storage is the system of record for datasets and checkpoints. Stage data to local NVMe at job start with parallel downloads, and upload checkpoints asynchronously so the GPUs are not idle while they write.
  • File Storage is convenient for shared code, environments and small configs across nodes, but too slow for heavy data loading on 8-GPU nodes.
  • Block volumes suit single-node work that must survive a stop, at higher performance tiers; see OCI block volumes.

Kubernetes and health

Container Engine for Kubernetes (OKE) supports GPU node pools. The usual pattern is a CPU pool for system workloads and a GPU pool tainted so only GPU jobs land there, with the NVIDIA device plugin exposing GPUs as a schedulable resource. Check the OKE documentation for current GPU image and plugin guidance; cluster operation itself is covered in OKE in depth.

Whatever orchestrates the nodes, check health before trusting them. A single degraded GPU slows a synchronous multi-node job to its pace. Before each job, and on a schedule, run: nvidia-smi for GPU count, ECC errors and temperature; dmesg for NVIDIA Xid errors; rdma link show for link state; and a short single-node and two-node all_reduce_perf. Drain any node that fails and replace it rather than debugging it in place.

Worked example: a two-node H100 fine-tune

In this illustrative scenario, a team needs to fine-tune a 70B model with sharded data parallelism across 16 H100s. They request a limit of two BM.GPU.H100.8 instances in one AD and wait for approval. They build a custom image from the HPC cluster networking image with their PyTorch and NCCL versions pinned, create an instance configuration in a private subnet, and launch a two-node cluster network.

First job: 38% of the throughput they measured on a single node, scaled. NCCL_DEBUG=INFO shows NET/Socket in the log: NCCL chose the front-end interface. Setting NCCL_IB_HCA to the devices ibv_devices reported, and the GID index the image documentation specified, changes the log to NET/IB and lifts scaling efficiency to 88%. A week later, step time doubles. nvidia-smi on node 2 shows one GPU throttling on temperature. They drain it, the pool launches a replacement from the same configuration, the job resumes from the last checkpoint in Object Storage, and they add the health check to the job prologue.

Failure modes and cost traps

  • Silent socket fallback. Multi-node jobs run but scale badly. Fix: check the NCCL log for the transport and gate jobs on an all-reduce bandwidth test.
  • Lost scratch data. Checkpoints written only to local NVMe vanish with the node. Fix: asynchronous upload to Object Storage on every checkpoint.
  • Limit in the wrong AD. Automation fails at launch. Fix: list shapes per AD in CI before applying.
  • Driver drift. Nodes launched months apart behave differently. Fix: pinned image OCIDs.
  • One bad GPU. The whole job slows. Fix: prologue health checks and drain-and-replace.
  • Paying for idle GPUs. A reserved 8-GPU node idles between experiments. Fix: schedule work onto it or release it, and read Oracle's billing rules for stopped instances of your shape rather than assuming stopping stops the bill.

What to do next

  1. Run the SDK listing script to see which GPU shapes each availability domain offers your tenancy.
  2. Request service limits for the exact shape and AD you need, well before you need them.
  3. Choose and pin an image; record its driver, CUDA or ROCm, and NCCL versions.
  4. For multi-node work, create an instance configuration and a cluster network, and prove bandwidth with nccl-tests.
  5. Put datasets and checkpoints in Object Storage and use local NVMe only as a cache.
  6. Add GPU, Xid, RDMA and all-reduce checks to every job prologue, and drain nodes that fail.
Key takeaway: OCI GPU instances range from single A10 VMs for inference to 8-GPU bare metal H100, H200 and MI300X nodes for training. Pick by memory first, get service limits for the exact shape and availability domain early, pin a driver-bearing image, and for multi-node jobs use a cluster network with NCCL explicitly pointed at the RDMA interfaces and proven with nccl-tests. Keep data in Object Storage, treat local NVMe as a cache, and health-check every node before a job trusts it.