On Oracle Cloud Infrastructure every instance has a shape, and the shape decides almost everything about it: processor architecture, how much compute an OCPU represents, the memory you may pair with it, how much network bandwidth you get, how many VNICs you can attach and whether local NVMe is present. Flexible shapes make sizing continuous, so the question is no longer which fixed size to round up to, but which point in a large space to pick.

This article is about making that choice on evidence. The OCI Compute article introduces the shape families, capacity types and the launch path; here we go one level down. We look at what a shape fixes, how to compare OCPUs across processor families, which resources scale with OCPU count, how to read the live catalog from code, how to turn a load test into a cost per unit of work, and what happens when you change a running instance's shape. Figures come from Oracle's compute shape documentation as of this writing; check the catalog for your region before relying on any limit.

Advertisement

What a shape actually fixes

A flexible shape name such as VM.Standard.E5.Flex pins the hardware generation and processor family and gives you ranges for everything else. When you launch, you choose OCPUs and memory within those ranges, and the platform derives the rest. That derivation is where most sizing mistakes hide, because the resources you did not choose still change with the ones you did.

PropertyChosen or derivedWhat decides it
Processor family and architectureChosen via shape nameAMD, Intel or Ampere Arm; x86 or Arm images
OCPUsChosenWithin the shape's minimum and maximum
MemoryChosenAt least 1 GB or 1 GB per OCPU, whichever is greater; at most 64 GB per OCPU and the shape maximum
Network bandwidthDerived1 Gbps per OCPU on most flexible shapes, up to a per-shape cap
VNIC attachmentsDerived2 at 1 OCPU, then 1 per OCPU, up to 24
Local NVMeChosen via shape familyOnly DenseIO shapes; fixed OCPU and memory steps
BillingDerivedPer second with a one-minute minimum; OCPU-hours and GB-hours

The OCPU is not the same unit everywhere

OCI prices and sizes compute in OCPUs, and the definition depends on the processor. On x86 shapes, AMD and Intel, one OCPU is one physical core with simultaneous multithreading, which the guest sees as 2 vCPUs. On Ampere A1, one OCPU is one core and the guest sees 1 vCPU, because Ampere cores are single-threaded. On Ampere A2 and A4, one OCPU is 2 cores and 2 vCPUs.

Two consequences follow. First, never compare shapes by OCPU count alone. Four A1 OCPUs are four cores, four A2 OCPUs are eight cores, and four E5 OCPUs are four cores exposing eight hardware threads. Thread-heavy servers can benefit from SMT; code that saturates each core's execution units gains little from it. Second, the only honest comparison is throughput at your latency target on your own workload, divided by what the configuration costs per hour. That is the method later in this article.

Advertisement

Resources that follow the OCPU count

Network bandwidth and VNICs are allocated per OCPU, which makes small instances network-poor. The documentation lists 1 Gbps per OCPU for the standard flexible shapes with caps that differ by shape, and 4 Gbps per OCPU for the high-clock Optimized3 shape.

ShapeMax OCPUsMax memoryBandwidth
VM.Standard.E4.Flex641,024 GB1 Gbps/OCPU, max 40 Gbps
VM.Standard.E5.Flex1261,024 GB1 Gbps/OCPU, max 40 Gbps
VM.Standard.E6.Flex1261,454 GB1 Gbps/OCPU, max 99 Gbps
VM.Standard3.Flex32512 GB1 Gbps/OCPU, max 32 Gbps
VM.Standard.A1.Flex76472 GB1 Gbps/OCPU, max 40 Gbps
VM.Standard.A2.Flex78946 GB1 Gbps/OCPU, max 78 Gbps
VM.Standard.A4.Flex45700 GB1 Gbps/OCPU, max 100 Gbps
VM.Optimized3.Flex18256 GB4 Gbps/OCPU, max 40 Gbps

The practical rule: size a network-heavy service on bandwidth, not only on CPU. A proxy, a cache or a replication target that needs 6 Gbps cannot run on a 2-OCPU standard VM however idle its CPUs are; it needs 6 or more OCPUs, or an Optimized3 VM with 2. The same applies to multi-homed appliances that need several VNICs. Memory is the other lever people forget: because memory is chosen independently, a cache or JVM service that needs 96 GB but little CPU can run on a few OCPUs with a high memory ratio, instead of on a large fixed shape. Keep in mind the 64 GB per OCPU ceiling. The OCI Networking article covers the VCN side of VNIC design.

Reading the catalog from code

Shape availability differs by region, availability domain and tenancy limits, and new shapes appear regularly, so a hard-coded list goes stale. The Python SDK returns the live catalog, including each flexible shape's OCPU and memory ranges, supported burstable baselines and the shapes an instance can be resized to.

import oci

config = oci.config.from_file()                 # ~/.oci/config, DEFAULT profile
compute = oci.core.ComputeClient(config)
compartment = config["tenancy"]                 # or any compartment you launch into

shapes = oci.pagination.list_call_get_all_results(
    compute.list_shapes, compartment_id=compartment).data

seen = set()
for s in sorted(shapes, key=lambda s: s.shape):
    if not s.is_flexible or s.shape in seen:     # list is per AD: dedupe by name
        continue
    seen.add(s.shape)
    o, m = s.ocpu_options, s.memory_options
    print(f"{s.shape:26s} {s.processor_description or '':38.38s} "
          f"ocpus {o.min:g}-{o.max:g}  mem {m.min_in_g_bs:g}-{m.max_in_g_bs:g} GB "
          f"({m.min_per_ocpu_in_gbs or 0:g}-{m.max_per_ocpu_in_gbs or 0:g}/OCPU)  "
          f"burst {s.baseline_ocpu_utilizations or '-'}  "
          f"resize-to {len(s.resize_compatible_shapes or [])} shapes")

Two attributes are worth building tooling around. resize_compatible_shapes tells you in advance which target shapes a change-shape call will accept, which saves a failed maintenance window. baseline_ocpu_utilizations lists the burstable baselines a shape offers, which the SDK encodes as BASELINE_1_8, BASELINE_1_2 and BASELINE_1_1 (an eighth of an OCPU, half, and a normal non-burstable instance). The OCI CLI exposes the same data through oci compute shape list --compartment-id.

A decision flow for picking a shape

With the catalog and the scaling rules in hand, shape choice becomes a short sequence. Profile the workload first: CPU per request, resident memory, network throughput at peak and local disk I/O. Then fix the architecture, because that decides images and binaries: if your container images, JVM, native libraries or licensed software are x86-only, Arm shapes are out until you build multi-architecture images. Next identify the binding resource, which chooses the family. Then size OCPUs and memory together, checking that the derived bandwidth and VNIC count are enough. Pick a capacity type, and finally prove the choice with a load test.

1. Workload profileCPU, memory, network, local I/O2. Architecturex86 or Arm: images, binaries3. Binding resourcedecides the familyStandard.FlexE4/E5/E6, Standard3, A1/A2/A4Optimized3.Flexhigh clock, 4 Gbps/OCPUDenseIO.Flexlocal NVMe, fixed sizesBM, GPU, HPCwhole host or accelerators4. Size OCPUs and GBbandwidth and VNICs follow OCPUs5. Capacity typeon-demand, preemptible, burstable6. Load test, cost per unit of workresize (reboots) and repeat
Choosing an OCI shape. The family follows from architecture and the binding resource; OCPU count must satisfy CPU, bandwidth and VNIC needs at once; the load test closes the loop.

Right-sizing by measurement

The goal is the lowest cost per unit of useful work that still meets your latency objective. Define the unit first: requests, jobs, gigabytes processed. Then, for each candidate shape and size, run the same load test and find the highest sustained rate at which p99 latency stays under target with the error rate at zero. Leave headroom for traffic peaks and for losing one instance in a pool, and multiply by the hourly cost of the configuration.

# Cost per unit of work from a load test. Prices are inputs: take them from the
# OCI price list for your region; the PRICE_* names are yours to fill in.
def cost_per_million(ocpus, mem_gb, ocpu_hr, gb_hr, sustained_rps):
    hourly = ocpus * ocpu_hr + mem_gb * gb_hr
    return hourly / (sustained_rps * 3600) * 1e6

candidates = [
    # name,           ocpus, GB, $/OCPU-hr, $/GB-hr, rps at p99 target (measured)
    ("E5.Flex 4/32",      4, 32, PRICE_E5_OCPU, PRICE_E5_GB, 2100),
    ("A2.Flex 4/32",      4, 32, PRICE_A2_OCPU, PRICE_A2_GB, 1900),
    ("E5.Flex 2/16",      2, 16, PRICE_E5_OCPU, PRICE_E5_GB, 1000),
]
for name, c, g, po, pg, rps in candidates:
    print(f"{name:14s} ${cost_per_million(c, g, po, pg, rps):.4f} per million requests")

The prices are left as named inputs on purpose. OCI publishes per-OCPU-hour and per-GB-hour rates that differ by shape and change over time, and an OCPU is a different amount of hardware on each family. Take current figures from the price list for your region. The structure is what matters: halving OCPUs halves cost only if throughput at your p99 target stays above half, and moving to Arm pays only if throughput per dollar rises.

DenseIO, Optimized3 and when to leave VMs

DenseIO flexible shapes carry local NVMe drives and are flexible only in steps. The docs list VM.DenseIO.E4.Flex at 8, 16 or 32 OCPUs with 128, 256 or 512 GB of memory, and VM.DenseIO.E5.Flex at 8 to 48 OCPUs in steps of 8 with 12 GB per OCPU. Local NVMe is tied to the host: design as if its contents can vanish, which suits databases and search engines that replicate at the application layer, and does not suit a single-node database whose only copy is local.

VM.Optimized3.Flex trades core count for clock speed and more bandwidth per OCPU, with a maximum of 18 OCPUs. It suits latency-sensitive single-threaded work and small instances that need network throughput. Bare metal shapes give you a whole host with no hypervisor, for licensing counted on physical cores, custom virtualisation or strict noisy-neighbour requirements. GPU and HPC shapes add accelerators or RDMA networking and are chosen on the accelerator first.

Changing shape: what really happens

Flexible sizing does not mean live resizing. According to the documentation, if the instance is running when you change its shape it is rebooted as part of the operation, and there is no hot-add of OCPUs or memory. The public and private IP addresses, volume attachments and VNIC attachments are kept. Oracle recommends shutting the applications down from inside the operating system first, since a service that is slow to stop may be killed and corrupt data. You can move between fixed and flexible shapes and between current-generation AMD and Intel shapes; for Arm the documentation lists the Ampere shapes among themselves, and moving between x86 and Arm in practice means launching a new instance from an image built for that architecture.

from oci.core.models import UpdateInstanceDetails, UpdateInstanceShapeConfigDetails

# Stop the application cleanly first: a running instance is rebooted by this call.
compute.update_instance(
    instance_id,
    UpdateInstanceDetails(
        shape="VM.Standard.E5.Flex",
        shape_config=UpdateInstanceShapeConfigDetails(ocpus=4.0, memory_in_gbs=32.0),
    ),
)
# Burstable variant: add baseline_ocpu_utilization="BASELINE_1_2" (or BASELINE_1_8);
# BASELINE_1_1 means an ordinary, non-burstable instance.

Treat a shape change as a rolling deployment. Drain the instance from its load balancer, stop the application, change shape, let it reboot, verify health and put it back. In instance pools, update the instance configuration so replacements launch with the new shape.

Worked example: shrinking an over-provisioned API tier

An API tier runs on six VM.Standard.E4.Flex instances of 8 OCPUs and 128 GB each. Monitoring shows peak CPU at 30 percent, resident memory at 22 GB and peak network at 2.5 Gbps per instance. Metrics like these come from the OCI Monitoring service. The load-test figures that follow are illustrative.

Memory is the obvious waste: 128 GB against 22 GB used. CPU suggests about 3 OCPUs at peak, and 2.5 Gbps needs at least 3 OCPUs of bandwidth. The team tests three candidates: E5.Flex 4/32, A2.Flex 4/32 after building an Arm image, and E5.Flex 3/24. At the p99 target, E5 4/32 sustains 2,100 requests per second, A2 4/32 sustains 1,900 and E5 3/24 sustains 1,450. Plugging those rates and the current regional prices into the cost function ranks the candidates; in this case the Arm shape comes out cheapest per request despite lower throughput, and the tier needs seven instances instead of six to keep one-instance-loss headroom. The migration runs as a rolling replacement through the instance pool, because x86 to Arm is not an in-place change.

Failure modes

  • Comparing OCPU counts across families. An A1 OCPU, an A2 OCPU and an x86 OCPU are different amounts of compute.
  • Shrinking below the bandwidth you need. CPU looks fine and the network saturates, because bandwidth follows OCPUs.
  • Resizing in place during traffic. Change shape reboots a running instance; drain first.
  • Assuming Arm is a drop-in. One x86-only dependency blocks the move; build and test multi-architecture images first.
  • Storing the only copy on DenseIO NVMe. Local drives go with the host.
  • Hard-coding shape limits. Limits and availability change; read them from the catalog.

What to do next

  1. Export CPU, memory and network peaks for your largest instance groups and compare them with each shape's derived bandwidth.
  2. Run the catalog script in each region you use and save the output with the date.
  3. Define one unit of work per service and build a repeatable load test with a p99 target.
  4. Test at least one smaller x86 size and, if your images allow it, one Arm shape; compute cost per unit with current prices.
  5. Roll changes out as drain, stop, change shape, verify, and update instance configurations for pools.
  6. Repeat the exercise when new shape generations appear in the catalog.
Key takeaway: An OCI shape fixes the processor family and derives bandwidth and VNICs from the OCPU count you choose, and an OCPU means different amounts of compute on x86, A1 and A2 or A4. Read limits from the live catalog, size OCPUs for CPU and network together, choose memory independently, and decide between candidates by cost per unit of work at your latency target. Changing shape reboots the instance, so treat it as a rolling deployment.