Oracle Cloud Infrastructure offers a set of pretrained AI services that you call like any other API: send text, an image, an audio file or a document, and get structured results back. You do not train or host a model. The services that matter for most applications are Language, Vision, Speech and Document Understanding. Generative AI, which serves large language models, is a separate service with its own API and is not covered here.
This page treats the services as building blocks for an engineer: which calls are synchronous and which are jobs, how they read and write Object Storage, how IAM grants access, what fails in production and how to design around it. It finishes with a worked pipeline that turns call-centre recordings into redacted, scored transcripts. If OCI itself is new to you, read the OCI overview first; compartments and policies appear throughout. Facts here were checked against Oracle's documentation and the Python SDK reference in October 2026. These services change, so confirm limits on the current documentation pages before relying on them.
How the services fit together
The services and what changed
Each service has its own SDK client, its own policy family and its own limits. The table summarizes what each does and how you call it.
| Service | What it does | Calling pattern | Python SDK client |
|---|---|---|---|
| Language | Language detection, sentiment, key phrases, named entities, text classification, PII detection and masking, health entities, translation | Synchronous batch calls | oci.ai_language.AIServiceLanguageClient |
| Vision | Image classification, object detection, text detection, face detection; stored and streaming video analysis; custom image models | Synchronous analyze_image; image, video and stream jobs | oci.ai_vision.AIServiceVisionClient |
| Speech | Transcription with timestamps, confidence and diarization; live transcription; text-to-speech | Transcription jobs; live sessions; synchronous synthesize_speech | oci.ai_speech.AIServiceSpeechClient |
| Document Understanding | Text, table and key-value extraction, document classification, custom extraction models | Synchronous analyze_document; processor jobs | oci.ai_document.AIServiceDocumentClient |
Two retirements matter if you inherited older code. Anomaly Detection reached end of life on 6 March 2025; Oracle's recommended replacement is the anomaly detection operator in Data Science's Accelerated Data Science (ADS) library. Data Labeling reached end of life on 30 August 2025; Vision custom models now use Label Studio for labelling, and document labelling moved into Document Understanding's custom model workflow. Document extraction was split out of Vision into Document Understanding; for new document work use the dedicated service. Before you design around any other service, check Oracle's service change announcements page for its status.
Identity and access
Access is granted per service family in a compartment. Developers usually get use; whoever creates custom models or projects gets manage. Jobs that read input from Object Storage and write results there also need object permissions for the principal that creates them, which is the most common reason a first transcription job fails.
allow group ml-apps to use ai-service-language-family in compartment prod-ai
allow group ml-apps to use ai-service-vision-family in compartment prod-ai
allow group ml-apps to manage ai-service-speech-family in compartment prod-ai
allow group ml-apps to manage object-family in compartment prod-aiCode running inside OCI should not carry user API keys. A function uses a resource principal and a compute instance uses an instance principal; both are matched by a dynamic group that then appears in the policies in place of the group. Narrow object access to the buckets the pipeline uses with a where clause on the target bucket name. The OCI IAM guide explains verbs, dynamic groups and the error you get when a policy is missing, and OCI Vault covers secrets for anything that still needs one.
Synchronous calls: masking PII with Language
Synchronous calls suit small payloads and request paths. The Language batch methods take a list of TextDocument objects, each with a key you choose so you can match results back to inputs, and return per-document results plus a list of per-document errors. A batch can succeed partially, so always check both lists. Each call has limits on the number of documents and the characters per document and per request; read the current values from the Language limits page and size your batches from them rather than hard-coding numbers from an example.
import oci
from oci.ai_language import AIServiceLanguageClient
from oci.ai_language.models import (
BatchDetectLanguagePiiEntitiesDetails, TextDocument, PiiEntityMask)
signer = oci.auth.signers.get_resource_principals_signer() # inside OCI Functions
lang = AIServiceLanguageClient(config={}, signer=signer)
def redact(compartment_id, texts):
docs = [TextDocument(key=str(i), text=t, language_code="en") for i, t in enumerate(texts)]
details = BatchDetectLanguagePiiEntitiesDetails(
documents=docs,
compartment_id=compartment_id,
masking={"ALL": PiiEntityMask(mode="MASK", masking_character="*",
leave_characters_unmasked=4,
is_unmasked_from_end=True)})
resp = lang.batch_detect_language_pii_entities(details).data
for err in resp.errors:
log.warning("pii failed for doc %s: %s", err.key, err.error)
return {d.key: d.masked_text for d in resp.documents}The masking map is keyed by entity type; ALL applies one rule to every detected type, and the modes are MASK, REPLACE and REMOVE. You can instead give different rules per type, for example removing CREDIT_DEBIT_NUMBER entirely while masking PERSON. Detection is statistical: treat masking as a strong reduction of exposure, not a guarantee, and keep unmasked originals in a tightly controlled bucket if you keep them at all.
Asynchronous jobs: Speech and Document Understanding
Anything large, whether hours of audio, thousands of images or multi-page documents, goes through a job. You put inputs in a bucket, create a job that names the input objects and an output bucket and prefix, poll the job until it reaches a terminal state, and read JSON results from the output prefix. A Speech transcription job can hold up to 100 tasks, one per file; each file can be up to 2 GB and 4 hours, and jobs are kept for 90 days.
from oci.ai_speech import AIServiceSpeechClient
from oci.ai_speech.models import (
CreateTranscriptionJobDetails, ObjectListInlineInputLocation, ObjectLocation,
OutputLocation, TranscriptionModelDetails)
speech = AIServiceSpeechClient(config={}, signer=signer)
job = speech.create_transcription_job(CreateTranscriptionJobDetails(
compartment_id=compartment_id,
display_name="calls-2026-10-03",
input_location=ObjectListInlineInputLocation(object_locations=[
ObjectLocation(namespace_name=ns, bucket_name="calls-raw",
object_names=batch_of_object_names)]),
output_location=OutputLocation(namespace_name=ns, bucket_name="calls-transcripts",
prefix="2026-10-03/"),
model_details=TranscriptionModelDetails(language_code="en-US", domain="GENERIC"),
)).data
# Poll the job; individual tasks can fail while the job as a whole completes.
state = speech.get_transcription_job(job.id).data.lifecycle_stateDocument Understanding follows the same shape with create_processor_job: a GeneralProcessorConfig lists features such as DocumentTableExtractionFeature and DocumentKeyValueExtractionFeature, ObjectStorageLocations names the inputs, and an OutputLocation gives the bucket and prefix. Speech also offers a Whisper model alongside Oracle's own ASR model, with broader language coverage; select it through the model details and check the documentation for the exact model type value. Use the SDK's waiter helpers or a backoff loop for polling rather than a tight loop, and prefer event rules on the output bucket where you can, as described in Object Storage.
Designing for limits and throttling
Every service enforces limits on request size and rate, and the safest code treats them as configuration read at startup rather than constants scattered through handlers. Two pieces of code carry most of the load: a chunker that respects document and request size, and a retry wrapper that backs off on throttling without retrying errors that will never succeed.
import random, time
from oci.exceptions import ServiceError
def chunks(texts, max_docs, max_chars_per_doc, max_chars_per_request):
batch, size = [], 0
for t in texts:
for i in range(0, len(t), max_chars_per_doc): # split on sentence ends in real code
piece = t[i:i + max_chars_per_doc]
if batch and (len(batch) == max_docs or size + len(piece) > max_chars_per_request):
yield batch
batch, size = [], 0
batch.append(piece)
size += len(piece)
if batch:
yield batch
def with_retry(call, attempts=6):
for n in range(attempts):
try:
return call()
except ServiceError as e:
if e.status not in (429, 500, 503) or n == attempts - 1:
raise # 400 and 404 will not fix themselves
time.sleep(min(30, 2 ** n) * random.uniform(0.5, 1.0))Splitting by fixed character counts can cut a name or card number in half, which defeats PII detection on both pieces; split on sentence or utterance boundaries and keep a small overlap if you must split mid-sentence. Keep concurrency per tenancy bounded with a semaphore or a queue so that a burst of uploads does not turn into a burst of 429 responses, and log the request ID from each error so Oracle support can trace it.
Worked example: redacted call-centre transcripts
A support team records 3,000 calls a day, averaging 7 minutes each. They want transcripts with customer card numbers and names masked, a sentiment score per call, and nothing unredacted stored outside one restricted bucket.
- Recordings land in
calls-raw. An Object Storage event triggers a function that groups new objects into lists of up to 100 and creates one transcription job per list, so 3,000 calls become 30 jobs. - Each job writes one JSON transcript per call to
calls-transcripts. A second event triggers a function that reads each transcript, joins its tokens into utterances using the timestamps and speaker labels, and splits long text into chunks within the Language limits. - The function calls PII masking first, then sentiment on the masked text, so no unmasked text is ever sent to a second service call or written elsewhere.
- Masked transcripts and scores are written to
calls-clean; the raw bucket has a lifecycle rule that deletes audio and unmasked transcripts after the retention period agreed with compliance.
At about 350 hours of audio a day, the cost driver is Speech, which is billed by audio duration; Language is billed by the amount of text processed. Check the current price list for both before committing; Oracle prices by service and unit, and the ratio decides whether you transcribe everything or sample. The design keeps each function short-lived and idempotent: the job name and output prefix derive from the input list, so a retried event finds the existing job instead of paying twice. The OCI Functions guide covers timeouts and resource principals for these handlers.
Pretrained, custom or something else
Pretrained models cover common cases. When your labels are domain-specific, such as defect types on a production line or fields on your own invoice layout, Vision, Language and Document Understanding let you train custom models on your labelled data and call them through the same APIs with a model ID. That costs labelling time and model hosting, and you own the evaluation. When the task needs reasoning over text or open-ended generation, the Generative AI service is a better fit than forcing a classifier. When you need a model none of these provide, Data Science hosts your own. Choose the least custom option that meets the accuracy you have measured on your own data.
Failure modes
| Symptom | Likely cause | What to do |
|---|---|---|
| 404 NotAuthorizedOrNotFound | Missing family policy, or wrong compartment OCID | Check policy and the compartment in the request |
| Job fails reading input | Job principal lacks object permissions | Grant object access on the bucket |
| 429 TooManyRequests | Per-tenancy rate limit | Exponential backoff with jitter; spread load |
| Some documents missing from results | Per-document errors in a successful batch | Read the errors list and retry those keys |
| Service unavailable in a region | Not every service runs in every region | Check regional availability before choosing a region |
| Poor accuracy on audio | Low sample rate or lossy compression | Prefer FLAC or 16-bit PCM WAV at 16 kHz or more |
| Old code calls a retired API | Anomaly Detection or Data Labeling | Migrate to ADS operators or Label Studio |
Trade-offs
Managed AI services trade control for speed. You get no model to host and an API that works the first day, but you cannot inspect or fine-tune the pretrained models, results can shift when Oracle updates them, and limits and regions are set for you. Pin expectations with a small evaluation set of your own data and rerun it monthly so you notice drift. Compared with running an open model on OCI GPU instances, the services are cheaper at low and bursty volume and costlier at sustained high volume; the crossover depends on your volume and on staff time, so measure both.
What to do next
- List the AI tasks in your backlog and map each to Language, Vision, Speech, Document Understanding, Generative AI or a custom model.
- Search your code for Anomaly Detection and Data Labeling clients and plan their migration.
- Create a compartment and policies for the AI workload, with object access limited to the pipeline's buckets.
- Build a 100-item evaluation set from your own data and score the pretrained model before writing integration code.
- Read the current limits page for each service you use and set batch sizes from it.
- Implement retries with backoff and per-document error handling, and make job creation idempotent.
- Add lifecycle rules so raw inputs and unredacted outputs expire on schedule.