OCI Compute is Oracle Cloud Infrastructure's virtual machine and bare metal service. On the surface it looks like every other cloud's instance service: you pick a shape, an image and a subnet, and a machine appears. The details that decide cost, reliability and operability are different enough that habits carried over from other clouds cause real problems, starting with the unit you size machines in.

This article assumes you know the basics of an OCI tenancy and focuses on Compute itself: how flexible shapes are sized, the five ways to buy capacity and when each is right, the data flow of a launch, how instances should get credentials and configuration, how pools and autoscaling fit together, what happens during infrastructure maintenance, and the failure modes you will meet in production. It ends with a worked sizing example and a checklist. Limits and prices change; figures here were checked against the OCI documentation in September 2026, and prices are deliberately not quoted.

Advertisement

The two facts to carry in from the platform

Two platform concepts matter constantly in Compute; the OCI overview covers them in full. First, a region contains one or three availability domains (ADs), and every AD contains three fault domains (FDs): independent groups of hardware inside one data centre. Spreading instances across FDs protects against a host or rack failure; spreading across ADs protects against a data centre failure where the region has several. Second, OCI sizes x86 capacity in OCPUs, and one x86 OCPU is two vCPUs, a full physical core with both hyperthreads. So a 4-OCPU VM gives the guest 8 vCPUs; comparing it to an 8-vCPU instance elsewhere is fair, comparing it to a 4-vCPU one is not.

Shapes: sizing in OCPUs and gigabytes

A shape is the hardware template of an instance. Fixed shapes, such as older VM.Standard2 sizes, come in preset combinations. Flexible shapes, whose names end in .Flex, let you choose the number of OCPUs and the memory independently within the shape's limits, which removes most of the waste of rounding up to the next instance size.

Shape familyProcessorOCPU to vCPUMaximum per VM (Sept 2026 docs)
VM.Standard.E5.Flex / E6.FlexAMD1 OCPU = 2 vCPUs126 OCPUs; 1,024 GB / 1,454 GB
VM.Standard.E4.FlexAMD1 OCPU = 2 vCPUs64 OCPUs, 1,024 GB
VM.Standard3.FlexIntel1 OCPU = 2 vCPUs32 OCPUs, 512 GB
VM.Standard.A1.FlexAmpere Arm1 OCPU = 1 vCPU (1 core)76 OCPUs, 472 GB
VM.Standard.A2.Flex / A4.FlexAmpere Arm1 OCPU = 2 vCPUs (2 cores)78 / 45 OCPUs
VM.Optimized3.FlexIntel, high clock1 OCPU = 2 vCPUs18 OCPUs, 256 GB

Memory on these shapes can go up to 64 GB per OCPU, with a minimum of 1 GB or one GB per OCPU, whichever is greater. Network bandwidth and the number of VNICs you can attach scale with OCPU count, so a very small VM can be network-bound before it is CPU-bound; check the shape table for the exact figures before you shrink a network-heavy service. Note the Arm mapping trap: an A1 OCPU is one core, while an A2 or A4 OCPU is two, so the same OCPU count means different amounts of compute across Ampere generations.

Beyond general-purpose VMs there are bare metal shapes (BM.*), which give you the whole physical server with no hypervisor, for licensing that counts physical cores, custom hypervisors, or workloads that cannot tolerate noisy neighbours; DenseIO shapes with local NVMe for databases that replicate at the application layer; and GPU shapes. Local NVMe is ephemeral in the sense that matters: it is tied to the host, so design as if its contents can be lost.

Advertisement

Capacity types: five ways to buy the same machine

Capacity typeWhat you getUse it forWatch out for
On-demandStandard instance, billed while runningDefault for servicesHost capacity can run out in popular shapes
PreemptibleSame instance, 50% cheaper per the OCI docs; can be reclaimedBatch, CI, stateless workersTerminated with a 2-minute event; no pools, no bare metal
BurstableBaseline of 12.5% or 50% of each OCPU, bursts to fullLow, spiky load: dev, small web appsBilled on baseline; burst is not guaranteed
Capacity reservationPre-held host capacity in an ADFailover and scale-out that must succeedYou pay for reserved but unused capacity
Dedicated VM hostA whole host that only your VMs useIsolation and licensingYou manage bin-packing onto the host

Preemptible instances deserve precision. The documentation states they cost 50% less than on-demand in all regions, and that an instancepreemptionaction event is emitted two minutes before termination begins. They are not live-migrated during maintenance, cannot be used in instance pools or instance configurations, cannot use capacity reservations, and are not available on bare metal or burstable shapes. When you create one you choose whether to keep the boot volume on reclaim; the SDK's preserve_boot_volume defaults to false, which is usually what you want for disposable workers.

Burstable instances are charged for the baseline, not for what they use. A 1-OCPU E4.Flex instance at a 12.5% baseline is billed as 12.5% of an OCPU every hour whether it idles or bursts. The docs describe roughly an hour of continuous burst and warn that burst is not guaranteed because the capacity is oversubscribed. They suit workloads whose average is low and whose peaks are short; a service that is busy all afternoon will be throttled back to baseline.

The launch path, step by step

What happens when you launch an OCI Compute instanceLaunchInstance APIconsole, CLI, SDK, TerraformIAM policy checkcompartment + service limitsPlacementAD, fault domain, host capacityCapacity typeon-demand, preemptible,burstable, reservation, dedicated hostImage to boot volumeblock storage, network-attachedPrimary VNICprivate IP in a VCN subnetGuest bootscloud-init reads user_dataMetadata service169.254.169.254/opc/v2/IMDSv2Instance configurationtemplate for poolsInstance poolN instances across FDsAutoscalingmetric or schedulesame launch path
A launch passes IAM and service-limit checks, is placed on a host in an availability and fault domain, gets a boot volume from the image and a VNIC in a subnet, then boots and pulls configuration from the metadata service. Pools reuse the same path from a saved configuration.

It is worth knowing the order because each step fails differently. The API first checks your IAM policy and the tenancy's service limits for that shape in that AD; exceeding a limit fails immediately and is fixed with a limit increase request, not a retry. Placement then looks for a host with room in the requested AD and fault domain; if none has room, the launch fails with an out-of-host-capacity error, which a retry in another AD or FD, or a different shape, may fix. The image is cloned into a network-attached boot volume that outlives the instance unless you ask for it to be deleted on termination. The primary VNIC gets a private IP from the subnet, and a public IP only if the subnet is public and you allow it. Finally the guest boots and cloud-init reads the user_data you passed from the metadata service.

oci compute instance launch \
  --compartment-id "$COMPARTMENT_OCID" \
  --availability-domain "$AD_NAME" \
  --fault-domain FAULT-DOMAIN-2 \
  --shape VM.Standard.E5.Flex \
  --shape-config '{"ocpus": 2, "memoryInGBs": 16}' \
  --image-id "$IMAGE_OCID" \
  --subnet-id "$PRIVATE_SUBNET_OCID" \
  --assign-public-ip false \
  --ssh-authorized-keys-file ~/.ssh/id_ed25519.pub \
  --user-data-file cloud-init.yaml \
  --instance-options '{"areLegacyImdsEndpointsDisabled": true}' \
  --availability-config '{"isLiveMigrationPreferred": true, "recoveryAction": "RESTORE_INSTANCE"}' \
  --display-name api-01

Out-of-capacity is common enough for popular shapes that launch automation should handle it rather than page a human. The Python SDK sketch below tries each AD in turn and treats a limit error as fatal, since retrying cannot fix it.

import oci
from oci.core.models import (LaunchInstanceDetails, LaunchInstanceShapeConfigDetails,
                             InstanceSourceViaImageDetails, CreateVnicDetails)

compute = oci.core.ComputeClient(oci.config.from_file())

def launch_anywhere(ads, compartment, subnet, image, ocpus=2, mem_gb=16):
    for ad in ads:
        details = LaunchInstanceDetails(
            compartment_id=compartment, availability_domain=ad,
            shape="VM.Standard.E5.Flex",
            shape_config=LaunchInstanceShapeConfigDetails(ocpus=ocpus, memory_in_gbs=mem_gb),
            source_details=InstanceSourceViaImageDetails(image_id=image),
            create_vnic_details=CreateVnicDetails(subnet_id=subnet, assign_public_ip=False))
        try:
            return compute.launch_instance(details).data
        except oci.exceptions.ServiceError as e:
            if "capacity" in (e.message or "").lower():
                continue                  # try the next AD
            raise                         # limits, auth, bad input: not retryable
    raise RuntimeError("no capacity in any availability domain")

Metadata, identity and configuration

Every instance can query a metadata service at 169.254.169.254. Version 2 lives under /opc/v2/ and requires the header Authorization: Bearer Oracle; the legacy v1 endpoints need no header, which makes them easier to reach through a server-side request forgery bug in an application. Disable them with areLegacyImdsEndpointsDisabled at launch, as above, or later with UpdateInstance, and fix any old agent that still calls v1 first.

Do not put API keys on instances. OCI's equivalent of an instance role is the instance principal: create a dynamic group whose matching rule selects your instances, for example ALL {instance.compartment.id = '<compartment OCID>'}, write an IAM policy granting that group exactly what it needs, and have code authenticate with oci.auth.signers.InstancePrincipalsSecurityTokenSigner(). Credentials then rotate automatically and nothing secret lives on disk. Policy design is covered in OCI IAM.

Instance pools and autoscaling

An instance configuration is a saved launch template: shape, image, boot volume settings, VNIC and metadata. An instance pool uses it to keep N identical instances running, spread across the fault domains and availability domains you list, and replaces instances that are terminated. Pools can be attached to a load balancer backend set so new members are registered automatically. Autoscaling policies then change the pool size, either from metrics such as CPU or memory utilisation or on a schedule. Because pools cannot contain preemptible instances, a cheap elastic worker fleet on preemptible capacity needs your own controller, or a container platform that manages preemptible node pools for you.

Treat pool members as cattle: configuration comes from cloud-init and your deployment system, state lives on block volumes, databases or Object Storage, and changing the instance configuration is how you roll out a new image.

Maintenance, live migration and recovery

Hosts need firmware and hardware maintenance. For supported VMs, OCI can live-migrate the instance to a healthy host while it keeps running. The availability configuration controls this: isLiveMigrationPreferred states your preference, and recoveryAction decides what happens when an instance is recovered after maintenance that could not be done live. RESTORE_INSTANCE, the default, returns it to its previous state, rebooting it if it was running; STOP_INSTANCE leaves it stopped so you can start it on your own schedule, which suits clustered software that must rejoin carefully. Bare metal instances have no hypervisor to migrate from, so plan for reboot migrations and subscribe to maintenance events through the Events service and Monitoring.

Worked example: sizing a small production service

Suppose an API needs 12 vCPUs of steady CPU and 80 GB of memory in total, must survive the loss of any single host, and runs a nightly batch job that takes 6 hours on 16 vCPUs, plus a development box that is idle most of the day.

For the API, 12 vCPUs is 6 x86 OCPUs. To survive one failure, spread across three fault domains and size so two FDs can carry the load: three instances of 3 OCPUs and 40 GB each give 9 OCPUs (18 vCPUs) and 120 GB in total, and 6 OCPUs (12 vCPUs) and 80 GB with one FD down, exactly the requirement. On E5.Flex that is {"ocpus": 3, "memoryInGBs": 40} per instance, in a pool behind a load balancer, with live migration preferred. If you run the service on A1 instead, remember that 12 vCPUs is 12 A1 OCPUs, and benchmark first: the per-core performance is different, and so is the software support for Arm.

The batch job restarts cleanly from checkpoints, so it runs on preemptible VMs: two instances of 4 x86 OCPUs each give 16 vCPUs at half the on-demand rate. Handle the preemption event by checkpointing, and make the scheduler resubmit on a new instance. The development box fits a burstable E5.Flex at a 12.5% baseline: billed at one-eighth of its OCPUs, and able to burst for builds of up to about an hour.

Failure modes you will meet

FailureWhat you seeResponse
Out of host capacityLaunch fails for a popular shape in one ADRetry another AD or FD, another shape, or hold a capacity reservation
Service limit reachedLaunch fails immediately, retries do nothingRequest a limit increase ahead of scale events
PreemptionTwo-minute event, then terminationCheckpoint, drain, resubmit; never run singletons on preemptible
Burstable throttlingLatency rises after sustained loadMove to a 50% or full baseline
Orphaned boot volumesStorage bill grows after terminationsDelete boot volumes on termination unless you need them
IMDSv1 left enabledAn SSRF bug can read instance metadata, including user_dataDisable legacy endpoints on every instance and configuration
All replicas in one FDOne host failure takes the whole service downSet fault domains explicitly or use pool placement across FDs

Networking problems (security lists, route tables, missing service gateways) masquerade as Compute problems; OCI networking covers how to tell them apart.

What to do next

  1. Convert your current instance sizes into OCPUs correctly: divide x86 vCPUs by two, and check the Arm mapping for your Ampere generation.
  2. Pick a flexible shape per workload and size OCPUs and memory separately; confirm network bandwidth for small shapes.
  3. Classify every workload by capacity type: on-demand, preemptible, burstable, reserved or dedicated host.
  4. Disable legacy IMDS endpoints and move all instance credentials to instance principals with least-privilege policies.
  5. Put services in instance pools across fault domains, behind a load balancer, with autoscaling where load varies.
  6. Set the availability configuration deliberately and subscribe to maintenance and preemption events.
  7. Add out-of-capacity handling to launch automation and request service limits before you need them.
Key takeaway: OCI Compute sizes x86 machines in OCPUs of two vCPUs each, and flexible shapes let you choose OCPUs and memory independently. Buy capacity by workload: on-demand for services, preemptible at half price for restartable batch, burstable for idle-mostly machines, reservations when scale-out must succeed. Launch into pools across fault domains, configure through cloud-init and IMDSv2, authenticate with instance principals, set the maintenance behaviour explicitly, and build automation that treats out-of-capacity and preemption as normal events.