Exadata is Oracle's database machine: database servers and storage servers built and tuned together, joined by an RDMA network, with storage software that understands SQL. Exadata Cloud is the same machine rented as a service, either in an Oracle Cloud Infrastructure region or, under the Cloud@Customer model, as a rack Oracle installs and runs inside your own data centre while you drive it from the OCI console.
The marketing is about speed. The engineering decision is about who owns which layer, where the control path runs, what happens when the link to the cloud drops, and whether your workload uses the features you are paying for. If you only want Oracle to run the database itself, read Autonomous Database instead; for a single database on ordinary VMs, see OCI Base Database.
Why a database machine at all
In a conventional deployment the database server asks a storage array for blocks, gets every block requested, and throws most rows away after applying the WHERE clause. For a large scan the bottleneck is the pipe between storage and server, and most of its traffic is waste.
Exadata moves part of the query to the storage tier. The database servers run Oracle Database, usually as a Real Application Clusters (RAC) cluster across two or more nodes. The storage servers run Exadata storage software that can evaluate predicates, project only the needed columns, and decompress data before sending it back. Between them sits an internal RDMA over Converged Ethernet (RoCE) fabric; Oracle's X11M data sheet lists up to two active-active 100 Gb/s links per server. The result is that a full scan of a billion-row table can return only the few megabytes of matching rows to the database server instead of the whole table.
The four features that do the work
Smart Scan is predicate and column offload. It applies only to direct-path reads, which the database uses for full table scans and fast full index scans of large segments. Single-block reads through the buffer cache, the normal pattern for OLTP lookups by primary key, are not offloaded. This one fact explains most disappointment with Exadata: an OLTP workload gets fast flash and a fast network, but none of the offload magic.
Storage indexes are in-memory summaries the storage server keeps automatically: for regions of each disk, the minimum and maximum of some columns. A scan with a predicate outside a region's range skips that region without reading it. They are not configured and not persistent, so their effect rises over time after a restart and depends on how well data is physically clustered by the filtered column.
Smart Flash Cache keeps hot data on NVMe flash in each storage server in front of the disks, and Oracle's data sheet quotes flash scan throughput of up to 100 GB/s per X11M storage server. XRMEM (Exadata RDMA Memory) is DRAM on the storage servers that database servers read directly over RDMA, bypassing the storage software stack; the same sheet quotes read latency as low as 14 microseconds. XRMEM is what makes single-block OLTP reads fast, and it is the feature an OLTP-heavy workload is really buying.
Three ways to rent it
Oracle sells Exadata as a service in three shapes. Older material calls them ExaCS and ExaCC, so match on what each one is rather than the label.
| Model | Where the hardware is | What you get | When it fits |
|---|---|---|---|
| Exadata Database Service on Dedicated Infrastructure (ExaDB-D) | OCI region (or a partner cloud region) | A whole Exadata system for you; you create VM clusters on it and size compute within it | Large or many databases, isolation requirements, consolidation |
| Exadata Database Service on Exascale Infrastructure (ExaDB-XS) | OCI region, shared multitenant storage | VM clusters billed for the ECPUs and storage you use; Oracle runs the shared infrastructure | Smaller databases that want Exadata features without a dedicated system |
| Exadata Database Service on Cloud@Customer (ExaDB-C@C) | Your data centre | A dedicated rack Oracle installs and maintains, driven from the OCI console | Data residency, latency to on-premises apps, regulators who say not in a public region |
All three use the same database software and the same tooling, so skills and scripts move between them. Autonomous Database can also run on dedicated Exadata, both in a region and on Cloud@Customer; that is a different contract in which Oracle operates the database too.
The resource hierarchy and who owns what
Every model has the same nesting. At the top is the Exadata infrastructure: the physical rack, its database servers and storage servers. On Cloud@Customer you also define a VM cluster network, the IP addresses, VLANs, DNS and NTP settings the cluster will use on your network. Inside the infrastructure you create one or more VM clusters, each a set of virtual machines, one per selected database server, with an allocation of CPU, memory and local storage, running Grid Infrastructure. Each VM cluster holds database homes (one Oracle software version each), and each home holds container databases with their pluggable databases.
The ownership line is the most important thing to understand before you sign. Oracle owns the hardware, firmware, hypervisor, storage servers and internal network, and patches them in scheduled infrastructure maintenance. You own everything inside the VMs: the guest operating system, Grid Infrastructure and database patching (with Oracle-provided tooling to apply them), users, schemas, tuning, backups policy and security of the database itself. You have root on the VMs and SYSDBA on the databases; you do not have access to the storage servers. A team that expects a fully managed database will be surprised by how much is still theirs.
The Cloud@Customer control plane
A Cloud@Customer rack contains two control plane servers in addition to the database and storage servers. These connect out to the OCI region you chose, over a TLS tunnel on TCP port 443, initiated from your side. Your console clicks and API calls land in the region, and the region pushes the resulting operations down that tunnel: create a VM cluster, scale CPU, start maintenance, run a backup to cloud storage. Oracle also uses it to monitor the infrastructure. The documentation describes FastConnect as an option for extra isolation on that path in addition to the default TLS tunnel.
Two consequences follow. First, the data path does not depend on the uplink: the databases, their storage and your applications' SQL traffic stay on your network, so a failed uplink stops management operations and cloud-destination backups rather than queries in flight. Second, some metadata does go to the region, such as resource names, shapes and configuration, so name things as if the names will be read by a cloud provider, because they will.
Oracle staff sometimes need to work on the rack. Operator Access Control puts that behind approvals: an operator raises an access request, you approve or reject it, optionally pre-approve certain classes of action, and can require two approvals. Operator actions are logged for audit. Put approvers on the incident rota before go-live, or an urgent fix will wait for someone to wake up.
Worked example: should this workload move?
A team runs a 12 TB order database with a nightly reporting job on an ageing SAN-attached RAC pair, and considers Cloud@Customer because the data cannot leave the country. Before talking prices they need to know how much read traffic is direct-path scans and therefore eligible for offload. On an Exadata trial the statistics below answer that.
# Measure how much Smart Scan is actually saving on one database.
# Needs python-oracledb and a user with SELECT on V$SYSSTAT.
import oracledb
STATS = (
"cell physical IO bytes eligible for predicate offload",
"cell physical IO interconnect bytes returned by smart scan",
"cell physical IO bytes saved by storage index",
"physical read total bytes",
)
def offload_report(dsn, user, password):
with oracledb.connect(user=user, password=password, dsn=dsn) as conn:
cur = conn.cursor()
binds = {f"s{i}": s for i, s in enumerate(STATS)}
names = ", ".join(f":s{i}" for i in range(len(STATS)))
cur.execute(f"SELECT name, value FROM v$sysstat WHERE name IN ({names})", binds)
v = dict(cur.fetchall())
eligible = v[STATS[0]]
returned = v[STATS[1]]
si_saved = v[STATS[2]]
total = v[STATS[3]]
print(f"share of reads eligible for offload: {eligible / max(total, 1):.1%}")
print(f"interconnect reduction on eligible: {1 - returned / max(eligible, 1):.1%}")
print(f"bytes skipped by storage indexes: {si_saved / 2**30:,.1f} GiB")Suppose the report prints 41% of reads eligible for offload, a 93% interconnect reduction on those reads and 1.8 TB skipped by storage indexes. That says the nightly report benefits strongly, since 93% of its scan bytes never cross the fabric, while the 59% of reads that are buffered OLTP lookups will benefit from flash and XRMEM latency but not from offload. If the eligible share had been 3%, the honest conclusion would be that this is an OLTP system and an Exadata purchase should be justified on latency, consolidation and availability, not on Smart Scan.
Statistics in v$sysstat are cumulative since instance start, so take two snapshots around a representative window and subtract, rather than reading a single value.
Operating it: maintenance, backups and recovery
Oracle schedules infrastructure maintenance (storage server software, firmware, hypervisor) and lets you choose windows and preferences. Rolling maintenance updates database servers one at a time so a RAC database stays open with reduced capacity; non-rolling is faster but takes everything down together. Single-instance databases see an outage when their node is patched, a good reason to run production as RAC with relocatable services. Guest OS, Grid Infrastructure and database patches are on you.
The OCI SDK exposes both kinds of VM cluster and the maintenance runs, so a small script can turn the calendar into a list of affected clusters.
# List upcoming infrastructure maintenance and the VM clusters it will touch.
# Uses the OCI Python SDK; config comes from ~/.oci/config.
import oci
def upcoming_maintenance(compartment_id):
db = oci.database.DatabaseClient(oci.config.from_file())
clusters = {}
# Cloud@Customer VM clusters and ExaDB-D cloud VM clusters are different resources.
for vc in oci.pagination.list_call_get_all_results(
db.list_vm_clusters, compartment_id=compartment_id).data:
clusters.setdefault(vc.exadata_infrastructure_id, []).append(vc.display_name)
for vc in oci.pagination.list_call_get_all_results(
db.list_cloud_vm_clusters, compartment_id=compartment_id).data:
clusters.setdefault(vc.cloud_exadata_infrastructure_id, []).append(vc.display_name)
runs = oci.pagination.list_call_get_all_results(
db.list_maintenance_runs, compartment_id=compartment_id).data
for r in sorted(runs, key=lambda r: str(r.time_scheduled)):
if r.lifecycle_state in ("SCHEDULED", "IN_PROGRESS"):
touched = clusters.get(r.target_resource_id, ["(not an infrastructure target)"])
print(r.time_scheduled, r.lifecycle_state, r.display_name, "->", ", ".join(touched))For Cloud@Customer the documented backup destinations are Object Storage in the region, an NFS location, a Zero Data Loss Recovery Appliance, or local Exadata storage in the RECO disk group. Local-only backups are fast but share fate with the rack. Object Storage backups depend on uplink bandwidth, so a full backup of tens of terabytes over a modest link can take days; work out that arithmetic before choosing. For disaster recovery, Data Guard to a second rack or to ExaDB-D in a region is the standard pattern, and the switchover runbook should be exercised quarterly. Encryption keys for the databases can be held in OCI Vault; on Cloud@Customer, check with your security team whether keys may live in a region at all, since that is exactly the question residency rules tend to ask.
Failure modes
| Symptom | Likely cause | What to do |
|---|---|---|
| Reports did not speed up after migration | Plans use buffered reads, not direct-path scans, so nothing is offloaded | Check the eligible-bytes statistic; review parallelism and segment sizes; do not force full scans on OLTP paths |
| Console operations hang on Cloud@Customer | Control plane uplink down, proxy intercepting TLS, or DNS for OCI endpoints failing | Monitor the outbound path like any production dependency; exempt it from TLS inspection; keep a runbook that does not need the console |
| Brief brownout during maintenance | Rolling maintenance removed a node; sessions did not drain | Use RAC services with drain timeouts and Application Continuity; test a rolling window in non-production first |
| Cloud backups miss their window | Uplink bandwidth smaller than daily change rate | Back up locally or to a Recovery Appliance first; send to Object Storage as a second copy |
| Noisy neighbour between databases | Several databases share a VM cluster without resource plans | Use Database Resource Manager and I/O resource management, or separate VM clusters |
Trade-offs worth stating plainly
Exadata services buy you a tuned platform, offload for analytic scans, very low-latency storage and Oracle doing the hardware work. They cost more at the entry point than generic VMs, tie you to Oracle's maintenance calendar, and leave substantial database operations with you unless you choose Autonomous. Cloud@Customer adds a physical project (floor space, power, cooling, IP planning) and an outbound link your security team must approve. ExaDB-XS lowers the entry point by sharing storage with other tenants. If your databases are small, mostly key-value in shape, and portable, a managed open-source database or Base Database is often enough; see cloud-native databases for that side of the decision.
What to do next
- Measure your current workload's direct-path share and decide whether you are buying offload, latency, consolidation or residency; write that down as the success metric.
- Pick the model: ExaDB-XS for small Exadata-featured databases, ExaDB-D for dedicated capacity in a region, Cloud@Customer only if data or latency must stay on your premises.
- Draw the ownership line for your organisation: which team patches the guest OS, Grid Infrastructure and databases, and who approves operator access.
- For Cloud@Customer, get the outbound 443 path, DNS, NTP, VLANs and IP ranges agreed with network and security teams before the rack arrives.
- Run production as RAC with services, drain timeouts and Application Continuity so rolling maintenance is invisible.
- Choose two backup destinations with different failure domains, calculate the backup window against your bandwidth, and schedule a restore test.
- Set up Data Guard to a second site and rehearse a switchover before you need one.
- Automate the maintenance inventory with the SDK script above and review it at every change meeting.