An OCI bare metal instance is an entire physical server rented through the same API, console, VCN, block storage and IAM as a virtual machine. There is no Oracle hypervisor on it. Every core, every byte of memory and every local drive belongs to your operating system, and Oracle still enforces your network isolation, because it moved network virtualization off the host and into dedicated hardware.

That design makes bare metal an ordinary cloud instance with unusual properties: no hypervisor overhead or noisy neighbours, access to hardware features a VM hides, the option to run your own hypervisor, and large local NVMe and RDMA fabrics on some shapes. It also brings costs VMs hide: fixed sizes, slower provisioning, no live migration and more responsibility. This article covers what is specific to bare metal; general compute and shape selection are covered in OCI compute in depth and OCI compute shapes.

Advertisement

How off-box virtualization works

In a typical cloud, the hypervisor on each host does two jobs: it splits the machine among tenants, and it implements the virtual network, applying security rules and encapsulating traffic. If the host has no hypervisor, the second job has to live somewhere the tenant cannot tamper with.

OCI puts it outside the server, in custom network hardware between the host and the physical network. That layer maps the host's traffic to your VCN, enforces security lists and network security groups, and applies routing. The host sees ordinary Ethernet NICs. Because enforcement is outside the host, a tenant with root on a bare metal server still cannot see another tenant's traffic or escape its VCN, and the design means VMs and bare metal share one network model.

OCI bare metal: the whole server is yours; network virtualization happens off the boxPhysical host (single tenant)Your OS or your hypervisorno Oracle hypervisor, all coresPhysical NIC 1primary VNICPhysical NIC 2VLAN-tagged VNICsLocal NVMeDenseIO, GPU shapesRDMA NICsGPU / HPC shapesOff-box virtualizationVCN, security listsVCNsubnets, gatewaysBlock storageboot and data, iSCSICluster networkRoCE fabricOther GPU hostsNCCL, MPItenant trafficRDMAIsolation and VCN features are enforced outside the host, so the host CPU and memory carry no virtualization overhead.RDMA traffic uses its own NICs and fabric, separate from VCN traffic.
A bare metal host runs your OS directly. VCN enforcement happens in off-box virtualization; block storage arrives over the network; GPU and HPC shapes add a separate RDMA fabric.

The shapes, and how to read them

Bare metal shapes start with BM. and, unlike flexible VM shapes, come in one size: you get the whole machine. Some examples from the shape reference at the time of writing:

ShapeOCPUsMemoryLocal storageNetworking
BM.Standard.E5.1921922,304 GBBlock storage only1 x 100 Gbps; up to 129 VNICs (1 on the first physical NIC, 128 on the second)
BM.DenseIO.E5.1281281,536 GB12 x 6.8 TB NVMe (81.6 TB)1 x 100 Gbps
BM.GPU.H100.81122,048 GB, plus 8 x H100 with 640 GB GPU memory16 x 3.84 TB NVMe1 x 100 Gbps, plus 8 x 2 x 200 Gbps RDMA

On x86 shapes an OCPU is one physical core with two hardware threads, so a 192-OCPU host shows 384 logical CPUs. Shapes, limits and regional availability change, so read them from the API rather than from any article, this one included:

# Current bare metal shapes available to your tenancy, with their resources
oci compute shape list --compartment-id "$COMPARTMENT_OCID" --all \
  | jq -r '.data[] | select(.shape | startswith("BM."))
           | [.shape, .ocpus, ."memory-in-gbs", ."networking-bandwidth-in-gbps"] | @tsv'

# Launch one. Bare metal shapes are fixed: no --shape-config.
oci compute instance launch \
  --compartment-id "$COMPARTMENT_OCID" \
  --availability-domain "$AD" \
  --fault-domain FAULT-DOMAIN-1 \
  --shape BM.DenseIO.E5.128 \
  --image-id "$IMAGE_OCID" \
  --subnet-id "$SUBNET_OCID" \
  --display-name db-node-1 \
  --ssh-authorized-keys-file ~/.ssh/id_ed25519.pub

Service limits for bare metal shapes are often zero or small in a new tenancy, and capacity for large shapes varies by availability domain. Request limits early, and consider a capacity reservation for anything that must be relaunched on demand.

Advertisement

When bare metal is the right answer

ReasonWhy bare metal helpsCheck first
Consistent latencyNo hypervisor scheduling or noisy neighboursWhether a large dedicated VM already meets the target
Your own hypervisorRun KVM or another hypervisor with full controlThe operational cost of running virtualization yourself
Local NVMe at scaleTens of TB of direct-attached flashYour replication story, since local data is not durable
GPU and HPC clustersRDMA cluster networks for NCCL and MPIShape availability and limits in your region
Hardware accessPerformance counters and CPU features VMs hideWhether the tooling truly needs them
Licensing and complianceSingle-tenant hardware, whole-host core countsThe exact terms of your licence and audit requirements

It is the wrong answer for small services, anything that should scale in minutes, and fleets you would rather let the provider maintain. A 192-OCPU host running a service that needs 8 cores is a very expensive VM.

Launching and booting

The launch call is the same as for a VM, minus the shape configuration. Provisioning and first boot take noticeably longer than a VM, because a physical server goes through firmware initialization and, before handover, the host has been wiped since its previous tenant. Budget for several minutes in automation timeouts and in autoscaling plans.

Boot and data volumes come from OCI Block Volume over the network. On bare metal, block volume attachments use iSCSI; the paravirtualized attachment type is a VM feature. For an iSCSI data volume, the attachment gives you the iscsiadm commands to run, or you let the Block Volume management plugin of the Oracle Cloud Agent do it. Forgetting that step is the classic reason an attached volume does not appear.

Networking: VNICs as VLANs

The primary VNIC appears as a normal interface. Secondary VNICs work differently from VMs: each gets an Oracle-assigned VLAN tag, and on shapes with two physical NICs they sit on a specific one; the BM.Standard.E5.192 row above allows one VNIC on the first physical NIC and 128 on the second. The OS must create a VLAN sub-interface with that tag, and traffic must carry that VNIC's MAC address.

This is how you give KVM guests first-class VCN addresses: create a secondary VNIC per guest, and hand the guest an interface bound to the VNIC's VLAN and MAC. Security lists, network security groups and route tables then apply to each guest as if it were an OCI instance.

# On the host: expose a secondary VNIC to a KVM guest.
# From the VNIC attachment you get the VLAN tag (e.g. 1234) and the VNIC's MAC address.
PHYS=ens5f1          # the physical NIC that carries secondary VNICs on this shape
TAG=1234
MAC=02:00:17:0a:bc:de

# Option A: a macvtap on a VLAN sub-interface, handed to the guest
ip link add link $PHYS name $PHYS.$TAG type vlan id $TAG
ip link set $PHYS.$TAG up
ip link add link $PHYS.$TAG name macvtap$TAG address $MAC type macvtap mode passthru
ip link set macvtap$TAG up
# The guest's interface must use $MAC, or the virtualization layer drops its traffic.

Interface names differ by shape and image, so find the right physical NIC before scripting. For VCN design itself, see OCI networking.

Local NVMe and RDMA cluster networks

DenseIO and GPU shapes carry local NVMe drives. They are fast, and they are not durable: data survives an OS reboot, but it is lost when the instance is terminated, and the reboot migration described below requires you to accept deleting it. Treat local NVMe like a disk in a distributed database: replicate across hosts, or use it as a cache you can rebuild.

GPU and some HPC shapes add RDMA NICs attached to a cluster network, a RoCE fabric separate from the VCN. You create a cluster network from an instance configuration so OCI places the hosts together on the fabric; NCCL or MPI then use the RDMA interfaces directly. The per-host RDMA bandwidth is shape-specific, so read it from the shape reference, and verify it with a collective benchmark before trusting a training job's scaling numbers.

Maintenance without live migration

A VM can be live-migrated off failing hardware because a hypervisor can move its memory while it runs. A bare metal instance has no hypervisor, so maintenance means a reboot onto different hardware. OCI calls this reboot migration.

The documented flow: when a host needs maintenance, the instance gets a maintenance due date roughly 14 to 16 days out. You can reboot-migrate it yourself before then, on supported shapes, at a time you choose. If you do nothing, OCI stops, migrates and restarts it within 24 hours after the due date. For shapes with local storage you must explicitly agree to delete it, because the drives stay with the old host. Deadlines can sometimes be extended.

# On the host: find this instance's OCID (v2 metadata needs the header)
ID=$(curl -s -H "Authorization: Bearer Oracle" -L http://169.254.169.254/opc/v2/instance/ | jq -r .id)

# From anywhere with API access: is a maintenance reboot scheduled?
oci compute instance get --instance-id "$ID" \
  --query 'data."time-maintenance-reboot-due"' --raw-output

Poll the maintenance due date from metadata or the API, alert on it, and schedule the migration in your own maintenance window, draining the node first.

Identity and automation

Bare metal instances use the same instance principals as VMs: put the instance in a dynamic group and write policies for that group, so software on the host calls OCI APIs without stored keys. Because you own the whole host, anything running on it, including every guest of your own hypervisor that can reach the metadata endpoint, can obtain those credentials. Block the guests' access to 169.254.169.254 unless they need it. Policy design is covered in OCI IAM.

Worked example: a replicated database on DenseIO

A team runs a Cassandra-compatible database that needs about 60 TB of data on local flash with replication factor 3. Each BM.DenseIO.E5.128 host has 81.6 TB raw. They plan six hosts, two in each of three fault domains, so each replica set spans fault domains.

Raw capacity is 6 x 81.6 = 489.6 TB, or about 163 TB of unique data at RF 3. Keeping half free for compaction and repair leaves about 80 TB, comfortably above 60 TB. They use the drives as independent disks under the database rather than RAID, since replication already provides redundancy.

For reboot migration, a runbook checks the due date daily. When one appears, they decommission the node from the ring, reboot-migrate it with local storage deletion accepted, then bootstrap it back and stream its data from replicas. One node at a time, the cluster keeps quorum.

Failure modes

SymptomCauseFix
Launch fails with out of host capacityNo free hosts of that shape in the ADAnother AD, capacity reservation, or a different shape
Launch rejected for limitsBare metal service limit is zeroRequest a limit increase before you need it
Attached block volume is missingiSCSI login never ranRun the attachment's iscsiadm commands or enable the plugin
KVM guest has no networkWrong VLAN tag, NIC or MACMatch the VNIC attachment's VLAN tag and MAC exactly
Data gone after maintenanceReboot migration deleted local NVMeReplicate; drain and rebuild nodes as a routine
Instance restarted without warningMaintenance due date passed unhandledAlert on the due date; migrate proactively
Guests call OCI APIs as the hostGuests reach the metadata endpointBlock 169.254.169.254 from guest networks

What to do next

  1. Write down why you need a whole server: latency, your own hypervisor, local NVMe, RDMA, hardware access or licensing. If none applies, use a VM.
  2. List current bare metal shapes with the CLI and check service limits and capacity in your target availability domains.
  3. Plan maintenance: alert on the reboot-migration due date and write a drain-and-migrate runbook.
  4. Treat local NVMe as disposable: replicate across fault domains or make it a rebuildable cache.
  5. If you run your own hypervisor, script secondary VNIC creation with its VLAN tag and MAC, and block guest access to instance metadata.
  6. For GPU clusters, benchmark the RDMA fabric with a collective test before trusting training throughput.
Key takeaway: An OCI bare metal instance is a whole single-tenant server managed like any other instance. Moving network virtualization off the host is what lets Oracle keep your VCN isolation without a hypervisor, so all cores, memory, local NVMe and RDMA NICs are yours. Choose it for consistent latency, your own hypervisor, local flash, GPU clusters or licensing, not for small services. Read shapes from the API, request limits early, map secondary VNICs as VLANs, replicate local data, and plan for reboot migration, because without a hypervisor there is no live migration.