An OCI bare metal instance is an entire physical server rented through the same API, console, VCN, block storage and IAM as a virtual machine. There is no Oracle hypervisor on it. Every core, every byte of memory and every local drive belongs to your operating system, and Oracle still enforces your network isolation, because it moved network virtualization off the host and into dedicated hardware.
That design makes bare metal an ordinary cloud instance with unusual properties: no hypervisor overhead or noisy neighbours, access to hardware features a VM hides, the option to run your own hypervisor, and large local NVMe and RDMA fabrics on some shapes. It also brings costs VMs hide: fixed sizes, slower provisioning, no live migration and more responsibility. This article covers what is specific to bare metal; general compute and shape selection are covered in OCI compute in depth and OCI compute shapes.
How off-box virtualization works
In a typical cloud, the hypervisor on each host does two jobs: it splits the machine among tenants, and it implements the virtual network, applying security rules and encapsulating traffic. If the host has no hypervisor, the second job has to live somewhere the tenant cannot tamper with.
OCI puts it outside the server, in custom network hardware between the host and the physical network. That layer maps the host's traffic to your VCN, enforces security lists and network security groups, and applies routing. The host sees ordinary Ethernet NICs. Because enforcement is outside the host, a tenant with root on a bare metal server still cannot see another tenant's traffic or escape its VCN, and the design means VMs and bare metal share one network model.
The shapes, and how to read them
Bare metal shapes start with BM. and, unlike flexible VM shapes, come in one size: you get the whole machine. Some examples from the shape reference at the time of writing:
| Shape | OCPUs | Memory | Local storage | Networking |
|---|---|---|---|---|
| BM.Standard.E5.192 | 192 | 2,304 GB | Block storage only | 1 x 100 Gbps; up to 129 VNICs (1 on the first physical NIC, 128 on the second) |
| BM.DenseIO.E5.128 | 128 | 1,536 GB | 12 x 6.8 TB NVMe (81.6 TB) | 1 x 100 Gbps |
| BM.GPU.H100.8 | 112 | 2,048 GB, plus 8 x H100 with 640 GB GPU memory | 16 x 3.84 TB NVMe | 1 x 100 Gbps, plus 8 x 2 x 200 Gbps RDMA |
On x86 shapes an OCPU is one physical core with two hardware threads, so a 192-OCPU host shows 384 logical CPUs. Shapes, limits and regional availability change, so read them from the API rather than from any article, this one included:
# Current bare metal shapes available to your tenancy, with their resources
oci compute shape list --compartment-id "$COMPARTMENT_OCID" --all \
| jq -r '.data[] | select(.shape | startswith("BM."))
| [.shape, .ocpus, ."memory-in-gbs", ."networking-bandwidth-in-gbps"] | @tsv'
# Launch one. Bare metal shapes are fixed: no --shape-config.
oci compute instance launch \
--compartment-id "$COMPARTMENT_OCID" \
--availability-domain "$AD" \
--fault-domain FAULT-DOMAIN-1 \
--shape BM.DenseIO.E5.128 \
--image-id "$IMAGE_OCID" \
--subnet-id "$SUBNET_OCID" \
--display-name db-node-1 \
--ssh-authorized-keys-file ~/.ssh/id_ed25519.pubService limits for bare metal shapes are often zero or small in a new tenancy, and capacity for large shapes varies by availability domain. Request limits early, and consider a capacity reservation for anything that must be relaunched on demand.
When bare metal is the right answer
| Reason | Why bare metal helps | Check first |
|---|---|---|
| Consistent latency | No hypervisor scheduling or noisy neighbours | Whether a large dedicated VM already meets the target |
| Your own hypervisor | Run KVM or another hypervisor with full control | The operational cost of running virtualization yourself |
| Local NVMe at scale | Tens of TB of direct-attached flash | Your replication story, since local data is not durable |
| GPU and HPC clusters | RDMA cluster networks for NCCL and MPI | Shape availability and limits in your region |
| Hardware access | Performance counters and CPU features VMs hide | Whether the tooling truly needs them |
| Licensing and compliance | Single-tenant hardware, whole-host core counts | The exact terms of your licence and audit requirements |
It is the wrong answer for small services, anything that should scale in minutes, and fleets you would rather let the provider maintain. A 192-OCPU host running a service that needs 8 cores is a very expensive VM.
Launching and booting
The launch call is the same as for a VM, minus the shape configuration. Provisioning and first boot take noticeably longer than a VM, because a physical server goes through firmware initialization and, before handover, the host has been wiped since its previous tenant. Budget for several minutes in automation timeouts and in autoscaling plans.
Boot and data volumes come from OCI Block Volume over the network. On bare metal, block volume attachments use iSCSI; the paravirtualized attachment type is a VM feature. For an iSCSI data volume, the attachment gives you the iscsiadm commands to run, or you let the Block Volume management plugin of the Oracle Cloud Agent do it. Forgetting that step is the classic reason an attached volume does not appear.
Networking: VNICs as VLANs
The primary VNIC appears as a normal interface. Secondary VNICs work differently from VMs: each gets an Oracle-assigned VLAN tag, and on shapes with two physical NICs they sit on a specific one; the BM.Standard.E5.192 row above allows one VNIC on the first physical NIC and 128 on the second. The OS must create a VLAN sub-interface with that tag, and traffic must carry that VNIC's MAC address.
This is how you give KVM guests first-class VCN addresses: create a secondary VNIC per guest, and hand the guest an interface bound to the VNIC's VLAN and MAC. Security lists, network security groups and route tables then apply to each guest as if it were an OCI instance.
# On the host: expose a secondary VNIC to a KVM guest.
# From the VNIC attachment you get the VLAN tag (e.g. 1234) and the VNIC's MAC address.
PHYS=ens5f1 # the physical NIC that carries secondary VNICs on this shape
TAG=1234
MAC=02:00:17:0a:bc:de
# Option A: a macvtap on a VLAN sub-interface, handed to the guest
ip link add link $PHYS name $PHYS.$TAG type vlan id $TAG
ip link set $PHYS.$TAG up
ip link add link $PHYS.$TAG name macvtap$TAG address $MAC type macvtap mode passthru
ip link set macvtap$TAG up
# The guest's interface must use $MAC, or the virtualization layer drops its traffic.Interface names differ by shape and image, so find the right physical NIC before scripting. For VCN design itself, see OCI networking.
Local NVMe and RDMA cluster networks
DenseIO and GPU shapes carry local NVMe drives. They are fast, and they are not durable: data survives an OS reboot, but it is lost when the instance is terminated, and the reboot migration described below requires you to accept deleting it. Treat local NVMe like a disk in a distributed database: replicate across hosts, or use it as a cache you can rebuild.
GPU and some HPC shapes add RDMA NICs attached to a cluster network, a RoCE fabric separate from the VCN. You create a cluster network from an instance configuration so OCI places the hosts together on the fabric; NCCL or MPI then use the RDMA interfaces directly. The per-host RDMA bandwidth is shape-specific, so read it from the shape reference, and verify it with a collective benchmark before trusting a training job's scaling numbers.
Maintenance without live migration
A VM can be live-migrated off failing hardware because a hypervisor can move its memory while it runs. A bare metal instance has no hypervisor, so maintenance means a reboot onto different hardware. OCI calls this reboot migration.
The documented flow: when a host needs maintenance, the instance gets a maintenance due date roughly 14 to 16 days out. You can reboot-migrate it yourself before then, on supported shapes, at a time you choose. If you do nothing, OCI stops, migrates and restarts it within 24 hours after the due date. For shapes with local storage you must explicitly agree to delete it, because the drives stay with the old host. Deadlines can sometimes be extended.
# On the host: find this instance's OCID (v2 metadata needs the header)
ID=$(curl -s -H "Authorization: Bearer Oracle" -L http://169.254.169.254/opc/v2/instance/ | jq -r .id)
# From anywhere with API access: is a maintenance reboot scheduled?
oci compute instance get --instance-id "$ID" \
--query 'data."time-maintenance-reboot-due"' --raw-outputPoll the maintenance due date from metadata or the API, alert on it, and schedule the migration in your own maintenance window, draining the node first.
Identity and automation
Bare metal instances use the same instance principals as VMs: put the instance in a dynamic group and write policies for that group, so software on the host calls OCI APIs without stored keys. Because you own the whole host, anything running on it, including every guest of your own hypervisor that can reach the metadata endpoint, can obtain those credentials. Block the guests' access to 169.254.169.254 unless they need it. Policy design is covered in OCI IAM.
Worked example: a replicated database on DenseIO
A team runs a Cassandra-compatible database that needs about 60 TB of data on local flash with replication factor 3. Each BM.DenseIO.E5.128 host has 81.6 TB raw. They plan six hosts, two in each of three fault domains, so each replica set spans fault domains.
Raw capacity is 6 x 81.6 = 489.6 TB, or about 163 TB of unique data at RF 3. Keeping half free for compaction and repair leaves about 80 TB, comfortably above 60 TB. They use the drives as independent disks under the database rather than RAID, since replication already provides redundancy.
For reboot migration, a runbook checks the due date daily. When one appears, they decommission the node from the ring, reboot-migrate it with local storage deletion accepted, then bootstrap it back and stream its data from replicas. One node at a time, the cluster keeps quorum.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Launch fails with out of host capacity | No free hosts of that shape in the AD | Another AD, capacity reservation, or a different shape |
| Launch rejected for limits | Bare metal service limit is zero | Request a limit increase before you need it |
| Attached block volume is missing | iSCSI login never ran | Run the attachment's iscsiadm commands or enable the plugin |
| KVM guest has no network | Wrong VLAN tag, NIC or MAC | Match the VNIC attachment's VLAN tag and MAC exactly |
| Data gone after maintenance | Reboot migration deleted local NVMe | Replicate; drain and rebuild nodes as a routine |
| Instance restarted without warning | Maintenance due date passed unhandled | Alert on the due date; migrate proactively |
| Guests call OCI APIs as the host | Guests reach the metadata endpoint | Block 169.254.169.254 from guest networks |
What to do next
- Write down why you need a whole server: latency, your own hypervisor, local NVMe, RDMA, hardware access or licensing. If none applies, use a VM.
- List current bare metal shapes with the CLI and check service limits and capacity in your target availability domains.
- Plan maintenance: alert on the reboot-migration due date and write a drain-and-migrate runbook.
- Treat local NVMe as disposable: replicate across fault domains or make it a rebuildable cache.
- If you run your own hypervisor, script secondary VNIC creation with its VLAN tag and MAC, and block guest access to instance metadata.
- For GPU clusters, benchmark the RDMA fabric with a collective test before trusting training throughput.