Every OCI compute instance that keeps state keeps it on a block volume. The boot disk is one, the database data directory usually sits on another, and the difference between a service that sustains its latency target and one that stalls at month-end is very often the performance level chosen for a volume two years earlier. Block Volume is also where OCI differs most from other clouds: performance is not tied to a disk type you pick once, but to a number of volume performance units per gigabyte that you can change while the volume is in use.
This article explains the service from first principles: what a block volume is, how the VPU model turns size into IOPS and throughput, how to size one for a real workload, how attachments work, how to grow a volume online, what auto-tune does, how backups, clones, volume groups and replicas differ, and which mistakes cause outages. Numbers come from Oracle's Block Volume documentation as of this writing; limits change, so recheck before committing.
What a block volume is
A block volume is network-attached storage that the instance sees as a raw disk. The instance puts a partition table and a filesystem on it, or hands it to a database as a raw device, exactly as with a local disk. Unlike the local NVMe drives on dense shapes described in OCI Bare Metal Compute, the data does not live in the host: it is stored by the Block Volume service and replicated automatically, and Oracle states a 99.99 percent annual durability target. Because it is not tied to a host, a volume survives instance termination, can be detached and attached to another instance in the same availability domain, and can be resized, backed up and cloned independently.
Volumes range from 50 GB to 32 TB in 1 GB steps. Data at rest is encrypted with AES-256 using Oracle-managed keys by default, or with your own key from OCI Vault. Boot volumes use the same machinery with a few extra rules and are what the instance boots from; everything below applies to both unless noted. For choosing the instance itself, see OCI Compute in depth.
The architecture in one picture
Two things in the picture drive most design decisions. First, every byte travels over the network between the instance and the storage service, so the instance's network bandwidth and attachment limits are a ceiling on top of whatever the volume can deliver. Oracle documents VM shapes as reaching up to 600,000 IOPS and 8,000 MB/s across all volumes, depending on the shape's network bandwidth, with up to 32 iSCSI or 16 paravirtualized attachments, and bare metal up to 1,300,000 IOPS and 12 GB/s. A small VM cannot use a large volume's full rating. Second, performance and protection are properties of the volume, not of the instance, so they can be changed without touching the guest.
The performance model: VPUs per gigabyte
Each volume has a performance level expressed as volume performance units per GB. Zero VPUs is Lower Cost, 10 is Balanced, the default, 20 is Higher Performance, and 30 to 120 in steps of 10 is Ultra High Performance. You pay for storage per GB plus VPUs per GB, so the level is a price multiplier. The level sets two per-GB rates and two per-volume caps:
| Level | VPUs/GB | IOPS per GB | Max IOPS per volume | KB/s per GB | Max MB/s per volume |
|---|---|---|---|---|---|
| Lower Cost | 0 | 2 | 3,000 | 240 | 480 |
| Balanced | 10 | 60 | 25,000 | 480 | 480 |
| Higher Performance | 20 | 75 | 50,000 | 600 | 680 |
| Ultra High Performance | 30-120 | 90-225 | 75,000-300,000 | 720-1,800 | 880-2,680 |
From Balanced upward the documentation gives the formulas: IOPS per GB is 1.5 times VPUs plus 45, the per-volume IOPS cap is 2,500 times VPUs, throughput per GB is 12 times VPUs plus 360 KB/s, and the throughput cap is 20 times VPUs plus 280 MB/s. A volume's performance is the smaller of size times the per-GB rate and the per-volume cap. That makes size a performance lever: a small Balanced volume is slow not because Balanced is slow, but because 100 GB times 60 IOPS is only 6,000 IOPS.
Three caveats matter in practice. Lower Cost has no IOPS SLA, so treat it as storage for backups, logs and cold data, not for anything a user waits on. Ultra High Performance requires a multipath-enabled attachment, and among VMs only shapes with 16 or more OCPUs support it; all current bare metal shapes do. And a volume attached to several instances shares its performance among them.
Worked example: sizing a database volume
Suppose a PostgreSQL primary needs 800 GB of data today, grows 20 GB a month, peaks at 30,000 random IOPS during batch jobs, and needs 400 MB/s for sequential scans and backups. Size first for capacity and headroom: 800 GB plus a year of growth is about 1,040 GB, so plan for 1 TB and grow later.
At Balanced, 1,024 GB times 60 is about 61,000 IOPS, but the per-volume cap is 25,000, below the peak. At Higher Performance, 1,024 GB times 75 is about 77,000, capped at 50,000, which covers the peak with margin, and throughput is 1,024 times 600 KB/s, about 600 MB/s, below the 680 MB/s cap and above the 400 MB/s need. Higher Performance at 1 TB fits. Striping two Balanced volumes with LVM would also reach 50,000 IOPS but doubles what you must back up consistently; stripe only beyond one volume's cap. Finally, check that the instance shape's network bandwidth can carry 600 MB/s; if it cannot, the volume is not the bottleneck and a bigger VPU setting buys nothing.
Attachments: paravirtualized, iSCSI and access modes
A paravirtualized attachment is presented by the hypervisor, so the guest sees a disk with no configuration; it is the simple default for VMs on current images. An iSCSI attachment is a network block protocol that the guest itself logs into; the console shows the iscsiadm commands to run, and on bare metal iSCSI is the only option because there is no hypervisor. Oracle notes that iSCSI delivers higher IOPS than paravirtualized, so for the most demanding VM workloads it can be worth the extra setup, and multipath iSCSI is what Ultra High Performance requires. In-transit encryption between instance and storage is available on supported attachments; it costs some performance, which is the documented trade-off.
Each attachment also has an access mode. Read/write is the default and allows one writer. Read-only protects reference data from accidental change. Read/write shareable lets several instances attach the same volume, but it does not make an ordinary filesystem safe to mount twice: you need a cluster-aware filesystem or application, such as OCFS2 or a database with its own clustering, or two writers will corrupt it. Use consistent device paths such as /dev/oracleoci/oraclevdb rather than /dev/sdb, which can change order between boots.
# 1 TB volume at Higher Performance (20 VPUs/GB) in the instance's availability domain
oci bv volume create \
--compartment-id "$COMPARTMENT" \
--availability-domain "$AD" \
--display-name pg-data-01 \
--size-in-gbs 1024 \
--vpus-per-gb 20 \
--wait-for-state AVAILABLE
# attach it to the instance as a paravirtualized device
oci compute volume-attachment attach \
--instance-id "$INSTANCE" \
--volume-id "$VOLUME" \
--type paravirtualized \
--device /dev/oracleoci/oraclevdb \
--wait-for-state ATTACHED
Provisioning as code and growing online
In Terraform the same volume, with its key, auto-tune policies and a replica, is one resource. The attribute names below are the provider's own:
resource "oci_core_volume" "pg_data" {
compartment_id = var.compartment_id
availability_domain = var.ad
display_name = "pg-data-01"
size_in_gbs = 1024
vpus_per_gb = 20 # default (minimum) level
kms_key_id = var.vault_key_id
autotune_policies {
autotune_type = "DETACHED_VOLUME" # drop to lower cost while nothing is attached
}
autotune_policies {
autotune_type = "PERFORMANCE_BASED"
max_vpus_per_gb = 40 # temporary ceiling when the volume is throttled
}
block_volume_replicas {
availability_domain = var.dr_ad # asynchronous replica in another AD
}
}Growth is online. Increase size_in_gbs, which the service applies to the live volume, then make the guest notice: rescan iSCSI sessions, grow the partition and then the filesystem. Volumes can only grow, never shrink, so to reduce size you create a smaller volume and copy the data. Changing vpus_per_gb is also online and needs nothing in the guest.
# after increasing size_in_gbs (the API change is online):
# iSCSI attachments need a rescan; the console shows the exact command for the volume
sudo iscsiadm -m node -R
# paravirtualized devices pick up the new size; confirm it
lsblk /dev/oracleoci/oraclevdb
# grow the partition, then the filesystem, without unmounting
sudo growpart /dev/oracleoci/oraclevdb 1
sudo xfs_growfs /var/lib/postgresql # XFS: grow by mount point
# sudo resize2fs /dev/oracleoci/oraclevdb1 # ext4: grow by device
# /etc/fstab: consistent device path, and never block boot on network storage
# /dev/oracleoci/oraclevdb1 /var/lib/postgresql xfs defaults,_netdev,nofail 0 2
Auto-tune
Auto-tune changes a volume's VPUs for you, under two independent policies. Detached-volume auto-tune moves a volume to the Lower Cost level while it is not attached to any instance and restores its configured level when it is attached again. It is a pure cost saving for volumes kept around between uses, such as test environments, with no effect on attached performance.
Performance-based auto-tune treats the configured VPUs as a floor. When the volume is being throttled at its current level for a sustained period, the service raises its VPUs gradually, up to the max_vpus_per_gb ceiling you set; when it has been idle at that level for a while, it lowers them back towards the floor. That suits bursty workloads such as nightly batch jobs, but it is not instantaneous: the adjustment follows sustained behaviour, so a latency-critical service whose peak arrives in seconds should be provisioned for the peak. Both policies can be enabled on the same volume, and you pay for the VPUs actually applied while they are applied.
Backups, clones, volume groups and replicas
| Mechanism | What you get | Use it for |
|---|---|---|
| Backup | Point-in-time copy stored in Object Storage; full or incremental; can be copied to another region | History, restore after corruption or deletion, compliance |
| Backup policy | Oracle-defined Bronze, Silver and Gold schedules, or your own | Making backups automatic rather than remembered |
| Clone | A new, independent volume from a point in time, in the same AD | Test copies of production data, fast environment creation |
| Volume group | Several volumes backed up, cloned or replicated together, consistent across members | Databases whose data and logs live on separate volumes |
| Replica | Continuous asynchronous copy in another AD or region | Disaster recovery with a small recovery point |
The key distinction is between replication and backup. A replica copies every write, so a dropped table or ransomware encryption arrives at the replica moments later; it protects against losing a site, not against losing data. Backups give you points in time to go back to. A backup is crash-consistent, the state a disk would have after sudden power loss, which databases are built to recover from; for application consistency, quiesce or freeze the filesystem around the backup, or use the database's own backup tooling. Restoring always creates a new volume. How backups land in Object Storage, and its tiers, is described in OCI Object Storage.
Benchmarking and monitoring
Measure the volume you built before you trust it, with direct I/O so the page cache does not flatter the result. The two runs below measure the IOPS and throughput ceilings separately; use --readonly on a volume that holds data.
# random 4 KiB reads: IOPS ceiling
fio --name=randread --filename=/dev/oracleoci/oraclevdb --direct=1 --rw=randread \
--bs=4k --iodepth=64 --numjobs=4 --ioengine=libaio --runtime=120 --time_based \
--group_reporting --readonly
# large sequential reads: throughput ceiling
fio --name=seqread --filename=/dev/oracleoci/oraclevdb --direct=1 --rw=read \
--bs=1m --iodepth=16 --numjobs=2 --ioengine=libaio --runtime=120 --time_based \
--group_reporting --readonlyIn production, watch the Block Volume metrics in OCI Monitoring, read and write operations and throughput per volume, and alarm when a volume sits near its cap for long periods. A volume that is always at its limit is either undersized or a candidate for performance-based auto-tune. Setting up the alarms themselves is covered in OCI Monitoring.
Failure modes
- Instance fails to boot after a change. An fstab entry for a network volume without
_netdevandnofailblocks boot when the volume is detached or slow to appear. - Paid for IOPS you cannot use. The shape's network bandwidth, not the volume, is the ceiling; or an Ultra High Performance volume was attached without multipath on a VM below 16 OCPUs.
- Corruption on a shared volume. Two instances mounted an ordinary filesystem on a shareable volume.
- Inconsistent multi-volume restore. Data and WAL volumes were backed up separately instead of as a volume group.
- Replica is not a backup. A deletion replicated to the DR copy and there were no backups to fall back on.
- Silent cost. High VPUs on large idle volumes, or backups retained long after their volumes were deleted.
What to do next
- List your volumes with their size, VPUs, attachment type and actual peak IOPS and throughput from Monitoring.
- Resize or re-level each one using the formulas: performance is the lesser of size times the per-GB rate and the per-volume cap.
- Confirm each instance shape's network bandwidth can carry what its volumes promise.
- Switch fstab entries to consistent device paths with
_netdev,nofail. - Put multi-volume databases in volume groups and assign a backup policy to every group or volume that holds state.
- Add a replica in another AD or region where the recovery point matters, in addition to backups.
- Enable detached-volume auto-tune on volumes that are parked, and performance-based auto-tune with a ceiling on bursty ones.
- Run the fio tests on every new volume class and rehearse one restore from backup per quarter.