OCI Compute is Oracle Cloud Infrastructure's virtual machine and bare metal service. On the surface it looks like every other cloud's instance service: you pick a shape, an image and a subnet, and a machine appears. The details that decide cost, reliability and operability are different enough that habits carried over from other clouds cause real problems, starting with the unit you size machines in.
This article assumes you know the basics of an OCI tenancy and focuses on Compute itself: how flexible shapes are sized, the five ways to buy capacity and when each is right, the data flow of a launch, how instances should get credentials and configuration, how pools and autoscaling fit together, what happens during infrastructure maintenance, and the failure modes you will meet in production. It ends with a worked sizing example and a checklist. Limits and prices change; figures here were checked against the OCI documentation in September 2026, and prices are deliberately not quoted.
The two facts to carry in from the platform
Two platform concepts matter constantly in Compute; the OCI overview covers them in full. First, a region contains one or three availability domains (ADs), and every AD contains three fault domains (FDs): independent groups of hardware inside one data centre. Spreading instances across FDs protects against a host or rack failure; spreading across ADs protects against a data centre failure where the region has several. Second, OCI sizes x86 capacity in OCPUs, and one x86 OCPU is two vCPUs, a full physical core with both hyperthreads. So a 4-OCPU VM gives the guest 8 vCPUs; comparing it to an 8-vCPU instance elsewhere is fair, comparing it to a 4-vCPU one is not.
Shapes: sizing in OCPUs and gigabytes
A shape is the hardware template of an instance. Fixed shapes, such as older VM.Standard2 sizes, come in preset combinations. Flexible shapes, whose names end in .Flex, let you choose the number of OCPUs and the memory independently within the shape's limits, which removes most of the waste of rounding up to the next instance size.
| Shape family | Processor | OCPU to vCPU | Maximum per VM (Sept 2026 docs) |
|---|---|---|---|
| VM.Standard.E5.Flex / E6.Flex | AMD | 1 OCPU = 2 vCPUs | 126 OCPUs; 1,024 GB / 1,454 GB |
| VM.Standard.E4.Flex | AMD | 1 OCPU = 2 vCPUs | 64 OCPUs, 1,024 GB |
| VM.Standard3.Flex | Intel | 1 OCPU = 2 vCPUs | 32 OCPUs, 512 GB |
| VM.Standard.A1.Flex | Ampere Arm | 1 OCPU = 1 vCPU (1 core) | 76 OCPUs, 472 GB |
| VM.Standard.A2.Flex / A4.Flex | Ampere Arm | 1 OCPU = 2 vCPUs (2 cores) | 78 / 45 OCPUs |
| VM.Optimized3.Flex | Intel, high clock | 1 OCPU = 2 vCPUs | 18 OCPUs, 256 GB |
Memory on these shapes can go up to 64 GB per OCPU, with a minimum of 1 GB or one GB per OCPU, whichever is greater. Network bandwidth and the number of VNICs you can attach scale with OCPU count, so a very small VM can be network-bound before it is CPU-bound; check the shape table for the exact figures before you shrink a network-heavy service. Note the Arm mapping trap: an A1 OCPU is one core, while an A2 or A4 OCPU is two, so the same OCPU count means different amounts of compute across Ampere generations.
Beyond general-purpose VMs there are bare metal shapes (BM.*), which give you the whole physical server with no hypervisor, for licensing that counts physical cores, custom hypervisors, or workloads that cannot tolerate noisy neighbours; DenseIO shapes with local NVMe for databases that replicate at the application layer; and GPU shapes. Local NVMe is ephemeral in the sense that matters: it is tied to the host, so design as if its contents can be lost.
Capacity types: five ways to buy the same machine
| Capacity type | What you get | Use it for | Watch out for |
|---|---|---|---|
| On-demand | Standard instance, billed while running | Default for services | Host capacity can run out in popular shapes |
| Preemptible | Same instance, 50% cheaper per the OCI docs; can be reclaimed | Batch, CI, stateless workers | Terminated with a 2-minute event; no pools, no bare metal |
| Burstable | Baseline of 12.5% or 50% of each OCPU, bursts to full | Low, spiky load: dev, small web apps | Billed on baseline; burst is not guaranteed |
| Capacity reservation | Pre-held host capacity in an AD | Failover and scale-out that must succeed | You pay for reserved but unused capacity |
| Dedicated VM host | A whole host that only your VMs use | Isolation and licensing | You manage bin-packing onto the host |
Preemptible instances deserve precision. The documentation states they cost 50% less than on-demand in all regions, and that an instancepreemptionaction event is emitted two minutes before termination begins. They are not live-migrated during maintenance, cannot be used in instance pools or instance configurations, cannot use capacity reservations, and are not available on bare metal or burstable shapes. When you create one you choose whether to keep the boot volume on reclaim; the SDK's preserve_boot_volume defaults to false, which is usually what you want for disposable workers.
Burstable instances are charged for the baseline, not for what they use. A 1-OCPU E4.Flex instance at a 12.5% baseline is billed as 12.5% of an OCPU every hour whether it idles or bursts. The docs describe roughly an hour of continuous burst and warn that burst is not guaranteed because the capacity is oversubscribed. They suit workloads whose average is low and whose peaks are short; a service that is busy all afternoon will be throttled back to baseline.
The launch path, step by step
It is worth knowing the order because each step fails differently. The API first checks your IAM policy and the tenancy's service limits for that shape in that AD; exceeding a limit fails immediately and is fixed with a limit increase request, not a retry. Placement then looks for a host with room in the requested AD and fault domain; if none has room, the launch fails with an out-of-host-capacity error, which a retry in another AD or FD, or a different shape, may fix. The image is cloned into a network-attached boot volume that outlives the instance unless you ask for it to be deleted on termination. The primary VNIC gets a private IP from the subnet, and a public IP only if the subnet is public and you allow it. Finally the guest boots and cloud-init reads the user_data you passed from the metadata service.
oci compute instance launch \
--compartment-id "$COMPARTMENT_OCID" \
--availability-domain "$AD_NAME" \
--fault-domain FAULT-DOMAIN-2 \
--shape VM.Standard.E5.Flex \
--shape-config '{"ocpus": 2, "memoryInGBs": 16}' \
--image-id "$IMAGE_OCID" \
--subnet-id "$PRIVATE_SUBNET_OCID" \
--assign-public-ip false \
--ssh-authorized-keys-file ~/.ssh/id_ed25519.pub \
--user-data-file cloud-init.yaml \
--instance-options '{"areLegacyImdsEndpointsDisabled": true}' \
--availability-config '{"isLiveMigrationPreferred": true, "recoveryAction": "RESTORE_INSTANCE"}' \
--display-name api-01Out-of-capacity is common enough for popular shapes that launch automation should handle it rather than page a human. The Python SDK sketch below tries each AD in turn and treats a limit error as fatal, since retrying cannot fix it.
import oci
from oci.core.models import (LaunchInstanceDetails, LaunchInstanceShapeConfigDetails,
InstanceSourceViaImageDetails, CreateVnicDetails)
compute = oci.core.ComputeClient(oci.config.from_file())
def launch_anywhere(ads, compartment, subnet, image, ocpus=2, mem_gb=16):
for ad in ads:
details = LaunchInstanceDetails(
compartment_id=compartment, availability_domain=ad,
shape="VM.Standard.E5.Flex",
shape_config=LaunchInstanceShapeConfigDetails(ocpus=ocpus, memory_in_gbs=mem_gb),
source_details=InstanceSourceViaImageDetails(image_id=image),
create_vnic_details=CreateVnicDetails(subnet_id=subnet, assign_public_ip=False))
try:
return compute.launch_instance(details).data
except oci.exceptions.ServiceError as e:
if "capacity" in (e.message or "").lower():
continue # try the next AD
raise # limits, auth, bad input: not retryable
raise RuntimeError("no capacity in any availability domain")
Metadata, identity and configuration
Every instance can query a metadata service at 169.254.169.254. Version 2 lives under /opc/v2/ and requires the header Authorization: Bearer Oracle; the legacy v1 endpoints need no header, which makes them easier to reach through a server-side request forgery bug in an application. Disable them with areLegacyImdsEndpointsDisabled at launch, as above, or later with UpdateInstance, and fix any old agent that still calls v1 first.
Do not put API keys on instances. OCI's equivalent of an instance role is the instance principal: create a dynamic group whose matching rule selects your instances, for example ALL {instance.compartment.id = '<compartment OCID>'}, write an IAM policy granting that group exactly what it needs, and have code authenticate with oci.auth.signers.InstancePrincipalsSecurityTokenSigner(). Credentials then rotate automatically and nothing secret lives on disk. Policy design is covered in OCI IAM.
Instance pools and autoscaling
An instance configuration is a saved launch template: shape, image, boot volume settings, VNIC and metadata. An instance pool uses it to keep N identical instances running, spread across the fault domains and availability domains you list, and replaces instances that are terminated. Pools can be attached to a load balancer backend set so new members are registered automatically. Autoscaling policies then change the pool size, either from metrics such as CPU or memory utilisation or on a schedule. Because pools cannot contain preemptible instances, a cheap elastic worker fleet on preemptible capacity needs your own controller, or a container platform that manages preemptible node pools for you.
Treat pool members as cattle: configuration comes from cloud-init and your deployment system, state lives on block volumes, databases or Object Storage, and changing the instance configuration is how you roll out a new image.
Maintenance, live migration and recovery
Hosts need firmware and hardware maintenance. For supported VMs, OCI can live-migrate the instance to a healthy host while it keeps running. The availability configuration controls this: isLiveMigrationPreferred states your preference, and recoveryAction decides what happens when an instance is recovered after maintenance that could not be done live. RESTORE_INSTANCE, the default, returns it to its previous state, rebooting it if it was running; STOP_INSTANCE leaves it stopped so you can start it on your own schedule, which suits clustered software that must rejoin carefully. Bare metal instances have no hypervisor to migrate from, so plan for reboot migrations and subscribe to maintenance events through the Events service and Monitoring.
Worked example: sizing a small production service
Suppose an API needs 12 vCPUs of steady CPU and 80 GB of memory in total, must survive the loss of any single host, and runs a nightly batch job that takes 6 hours on 16 vCPUs, plus a development box that is idle most of the day.
For the API, 12 vCPUs is 6 x86 OCPUs. To survive one failure, spread across three fault domains and size so two FDs can carry the load: three instances of 3 OCPUs and 40 GB each give 9 OCPUs (18 vCPUs) and 120 GB in total, and 6 OCPUs (12 vCPUs) and 80 GB with one FD down, exactly the requirement. On E5.Flex that is {"ocpus": 3, "memoryInGBs": 40} per instance, in a pool behind a load balancer, with live migration preferred. If you run the service on A1 instead, remember that 12 vCPUs is 12 A1 OCPUs, and benchmark first: the per-core performance is different, and so is the software support for Arm.
The batch job restarts cleanly from checkpoints, so it runs on preemptible VMs: two instances of 4 x86 OCPUs each give 16 vCPUs at half the on-demand rate. Handle the preemption event by checkpointing, and make the scheduler resubmit on a new instance. The development box fits a burstable E5.Flex at a 12.5% baseline: billed at one-eighth of its OCPUs, and able to burst for builds of up to about an hour.
Failure modes you will meet
| Failure | What you see | Response |
|---|---|---|
| Out of host capacity | Launch fails for a popular shape in one AD | Retry another AD or FD, another shape, or hold a capacity reservation |
| Service limit reached | Launch fails immediately, retries do nothing | Request a limit increase ahead of scale events |
| Preemption | Two-minute event, then termination | Checkpoint, drain, resubmit; never run singletons on preemptible |
| Burstable throttling | Latency rises after sustained load | Move to a 50% or full baseline |
| Orphaned boot volumes | Storage bill grows after terminations | Delete boot volumes on termination unless you need them |
| IMDSv1 left enabled | An SSRF bug can read instance metadata, including user_data | Disable legacy endpoints on every instance and configuration |
| All replicas in one FD | One host failure takes the whole service down | Set fault domains explicitly or use pool placement across FDs |
Networking problems (security lists, route tables, missing service gateways) masquerade as Compute problems; OCI networking covers how to tell them apart.
What to do next
- Convert your current instance sizes into OCPUs correctly: divide x86 vCPUs by two, and check the Arm mapping for your Ampere generation.
- Pick a flexible shape per workload and size OCPUs and memory separately; confirm network bandwidth for small shapes.
- Classify every workload by capacity type: on-demand, preemptible, burstable, reserved or dedicated host.
- Disable legacy IMDS endpoints and move all instance credentials to instance principals with least-privilege policies.
- Put services in instance pools across fault domains, behind a load balancer, with autoscaling where load varies.
- Set the availability configuration deliberately and subscribe to maintenance and preemption events.
- Add out-of-capacity handling to launch automation and request service limits before you need them.