Oracle Cloud Infrastructure offers a set of pretrained AI services that you call like any other API: send text, an image, an audio file or a document, and get structured results back. You do not train or host a model. The services that matter for most applications are Language, Vision, Speech and Document Understanding. Generative AI, which serves large language models, is a separate service with its own API and is not covered here.

This page treats the services as building blocks for an engineer: which calls are synchronous and which are jobs, how they read and write Object Storage, how IAM grants access, what fails in production and how to design around it. It finishes with a worked pipeline that turns call-centre recordings into redacted, scored transcripts. If OCI itself is new to you, read the OCI overview first; compartments and policies appear throughout. Facts here were checked against Oracle's documentation and the Python SDK reference in October 2026. These services change, so confirm limits on the current documentation pages before relying on them.

How the services fit together

Two calling patterns: synchronous analysis and asynchronous jobs over Object StorageYour serviceSDK, signed requestsIAM policyai-service-*-familyauthorizessmall payloadSynchronous APIsLanguage batch_detect_*, Vision analyze_image, Document analyze_documentJSON resultcreate jobAsynchronous jobsSpeech transcription, Vision image and video jobs, Document processor jobsObject Storageinput media, output JSONread inputwrite outputpoll job stateCustom modelstrained per service, same APIOCI Data Scienceyour own modelsRetired: Anomaly Detection (EOL 2025-03-06) and Data Labeling (EOL 2025-08-30)
Small payloads go to synchronous APIs; large media goes through jobs that read from and write to Object Storage. IAM policies are granted per service family.

The services and what changed

Each service has its own SDK client, its own policy family and its own limits. The table summarizes what each does and how you call it.

ServiceWhat it doesCalling patternPython SDK client
LanguageLanguage detection, sentiment, key phrases, named entities, text classification, PII detection and masking, health entities, translationSynchronous batch callsoci.ai_language.AIServiceLanguageClient
VisionImage classification, object detection, text detection, face detection; stored and streaming video analysis; custom image modelsSynchronous analyze_image; image, video and stream jobsoci.ai_vision.AIServiceVisionClient
SpeechTranscription with timestamps, confidence and diarization; live transcription; text-to-speechTranscription jobs; live sessions; synchronous synthesize_speechoci.ai_speech.AIServiceSpeechClient
Document UnderstandingText, table and key-value extraction, document classification, custom extraction modelsSynchronous analyze_document; processor jobsoci.ai_document.AIServiceDocumentClient

Two retirements matter if you inherited older code. Anomaly Detection reached end of life on 6 March 2025; Oracle's recommended replacement is the anomaly detection operator in Data Science's Accelerated Data Science (ADS) library. Data Labeling reached end of life on 30 August 2025; Vision custom models now use Label Studio for labelling, and document labelling moved into Document Understanding's custom model workflow. Document extraction was split out of Vision into Document Understanding; for new document work use the dedicated service. Before you design around any other service, check Oracle's service change announcements page for its status.

Identity and access

Access is granted per service family in a compartment. Developers usually get use; whoever creates custom models or projects gets manage. Jobs that read input from Object Storage and write results there also need object permissions for the principal that creates them, which is the most common reason a first transcription job fails.

allow group ml-apps to use ai-service-language-family in compartment prod-ai
allow group ml-apps to use ai-service-vision-family in compartment prod-ai
allow group ml-apps to manage ai-service-speech-family in compartment prod-ai
allow group ml-apps to manage object-family in compartment prod-ai

Code running inside OCI should not carry user API keys. A function uses a resource principal and a compute instance uses an instance principal; both are matched by a dynamic group that then appears in the policies in place of the group. Narrow object access to the buckets the pipeline uses with a where clause on the target bucket name. The OCI IAM guide explains verbs, dynamic groups and the error you get when a policy is missing, and OCI Vault covers secrets for anything that still needs one.

Synchronous calls: masking PII with Language

Synchronous calls suit small payloads and request paths. The Language batch methods take a list of TextDocument objects, each with a key you choose so you can match results back to inputs, and return per-document results plus a list of per-document errors. A batch can succeed partially, so always check both lists. Each call has limits on the number of documents and the characters per document and per request; read the current values from the Language limits page and size your batches from them rather than hard-coding numbers from an example.

import oci
from oci.ai_language import AIServiceLanguageClient
from oci.ai_language.models import (
    BatchDetectLanguagePiiEntitiesDetails, TextDocument, PiiEntityMask)

signer = oci.auth.signers.get_resource_principals_signer()   # inside OCI Functions
lang = AIServiceLanguageClient(config={}, signer=signer)

def redact(compartment_id, texts):
    docs = [TextDocument(key=str(i), text=t, language_code="en") for i, t in enumerate(texts)]
    details = BatchDetectLanguagePiiEntitiesDetails(
        documents=docs,
        compartment_id=compartment_id,
        masking={"ALL": PiiEntityMask(mode="MASK", masking_character="*",
                                      leave_characters_unmasked=4,
                                      is_unmasked_from_end=True)})
    resp = lang.batch_detect_language_pii_entities(details).data
    for err in resp.errors:
        log.warning("pii failed for doc %s: %s", err.key, err.error)
    return {d.key: d.masked_text for d in resp.documents}

The masking map is keyed by entity type; ALL applies one rule to every detected type, and the modes are MASK, REPLACE and REMOVE. You can instead give different rules per type, for example removing CREDIT_DEBIT_NUMBER entirely while masking PERSON. Detection is statistical: treat masking as a strong reduction of exposure, not a guarantee, and keep unmasked originals in a tightly controlled bucket if you keep them at all.

Asynchronous jobs: Speech and Document Understanding

Anything large, whether hours of audio, thousands of images or multi-page documents, goes through a job. You put inputs in a bucket, create a job that names the input objects and an output bucket and prefix, poll the job until it reaches a terminal state, and read JSON results from the output prefix. A Speech transcription job can hold up to 100 tasks, one per file; each file can be up to 2 GB and 4 hours, and jobs are kept for 90 days.

from oci.ai_speech import AIServiceSpeechClient
from oci.ai_speech.models import (
    CreateTranscriptionJobDetails, ObjectListInlineInputLocation, ObjectLocation,
    OutputLocation, TranscriptionModelDetails)

speech = AIServiceSpeechClient(config={}, signer=signer)

job = speech.create_transcription_job(CreateTranscriptionJobDetails(
    compartment_id=compartment_id,
    display_name="calls-2026-10-03",
    input_location=ObjectListInlineInputLocation(object_locations=[
        ObjectLocation(namespace_name=ns, bucket_name="calls-raw",
                       object_names=batch_of_object_names)]),
    output_location=OutputLocation(namespace_name=ns, bucket_name="calls-transcripts",
                                   prefix="2026-10-03/"),
    model_details=TranscriptionModelDetails(language_code="en-US", domain="GENERIC"),
)).data

# Poll the job; individual tasks can fail while the job as a whole completes.
state = speech.get_transcription_job(job.id).data.lifecycle_state

Document Understanding follows the same shape with create_processor_job: a GeneralProcessorConfig lists features such as DocumentTableExtractionFeature and DocumentKeyValueExtractionFeature, ObjectStorageLocations names the inputs, and an OutputLocation gives the bucket and prefix. Speech also offers a Whisper model alongside Oracle's own ASR model, with broader language coverage; select it through the model details and check the documentation for the exact model type value. Use the SDK's waiter helpers or a backoff loop for polling rather than a tight loop, and prefer event rules on the output bucket where you can, as described in Object Storage.

Designing for limits and throttling

Every service enforces limits on request size and rate, and the safest code treats them as configuration read at startup rather than constants scattered through handlers. Two pieces of code carry most of the load: a chunker that respects document and request size, and a retry wrapper that backs off on throttling without retrying errors that will never succeed.

import random, time
from oci.exceptions import ServiceError

def chunks(texts, max_docs, max_chars_per_doc, max_chars_per_request):
    batch, size = [], 0
    for t in texts:
        for i in range(0, len(t), max_chars_per_doc):   # split on sentence ends in real code
            piece = t[i:i + max_chars_per_doc]
            if batch and (len(batch) == max_docs or size + len(piece) > max_chars_per_request):
                yield batch
                batch, size = [], 0
            batch.append(piece)
            size += len(piece)
    if batch:
        yield batch

def with_retry(call, attempts=6):
    for n in range(attempts):
        try:
            return call()
        except ServiceError as e:
            if e.status not in (429, 500, 503) or n == attempts - 1:
                raise                                     # 400 and 404 will not fix themselves
            time.sleep(min(30, 2 ** n) * random.uniform(0.5, 1.0))

Splitting by fixed character counts can cut a name or card number in half, which defeats PII detection on both pieces; split on sentence or utterance boundaries and keep a small overlap if you must split mid-sentence. Keep concurrency per tenancy bounded with a semaphore or a queue so that a burst of uploads does not turn into a burst of 429 responses, and log the request ID from each error so Oracle support can trace it.

Worked example: redacted call-centre transcripts

A support team records 3,000 calls a day, averaging 7 minutes each. They want transcripts with customer card numbers and names masked, a sentiment score per call, and nothing unredacted stored outside one restricted bucket.

  1. Recordings land in calls-raw. An Object Storage event triggers a function that groups new objects into lists of up to 100 and creates one transcription job per list, so 3,000 calls become 30 jobs.
  2. Each job writes one JSON transcript per call to calls-transcripts. A second event triggers a function that reads each transcript, joins its tokens into utterances using the timestamps and speaker labels, and splits long text into chunks within the Language limits.
  3. The function calls PII masking first, then sentiment on the masked text, so no unmasked text is ever sent to a second service call or written elsewhere.
  4. Masked transcripts and scores are written to calls-clean; the raw bucket has a lifecycle rule that deletes audio and unmasked transcripts after the retention period agreed with compliance.

At about 350 hours of audio a day, the cost driver is Speech, which is billed by audio duration; Language is billed by the amount of text processed. Check the current price list for both before committing; Oracle prices by service and unit, and the ratio decides whether you transcribe everything or sample. The design keeps each function short-lived and idempotent: the job name and output prefix derive from the input list, so a retried event finds the existing job instead of paying twice. The OCI Functions guide covers timeouts and resource principals for these handlers.

Pretrained, custom or something else

Pretrained models cover common cases. When your labels are domain-specific, such as defect types on a production line or fields on your own invoice layout, Vision, Language and Document Understanding let you train custom models on your labelled data and call them through the same APIs with a model ID. That costs labelling time and model hosting, and you own the evaluation. When the task needs reasoning over text or open-ended generation, the Generative AI service is a better fit than forcing a classifier. When you need a model none of these provide, Data Science hosts your own. Choose the least custom option that meets the accuracy you have measured on your own data.

Failure modes

SymptomLikely causeWhat to do
404 NotAuthorizedOrNotFoundMissing family policy, or wrong compartment OCIDCheck policy and the compartment in the request
Job fails reading inputJob principal lacks object permissionsGrant object access on the bucket
429 TooManyRequestsPer-tenancy rate limitExponential backoff with jitter; spread load
Some documents missing from resultsPer-document errors in a successful batchRead the errors list and retry those keys
Service unavailable in a regionNot every service runs in every regionCheck regional availability before choosing a region
Poor accuracy on audioLow sample rate or lossy compressionPrefer FLAC or 16-bit PCM WAV at 16 kHz or more
Old code calls a retired APIAnomaly Detection or Data LabelingMigrate to ADS operators or Label Studio

Trade-offs

Managed AI services trade control for speed. You get no model to host and an API that works the first day, but you cannot inspect or fine-tune the pretrained models, results can shift when Oracle updates them, and limits and regions are set for you. Pin expectations with a small evaluation set of your own data and rerun it monthly so you notice drift. Compared with running an open model on OCI GPU instances, the services are cheaper at low and bursty volume and costlier at sustained high volume; the crossover depends on your volume and on staff time, so measure both.

What to do next

  1. List the AI tasks in your backlog and map each to Language, Vision, Speech, Document Understanding, Generative AI or a custom model.
  2. Search your code for Anomaly Detection and Data Labeling clients and plan their migration.
  3. Create a compartment and policies for the AI workload, with object access limited to the pipeline's buckets.
  4. Build a 100-item evaluation set from your own data and score the pretrained model before writing integration code.
  5. Read the current limits page for each service you use and set batch sizes from it.
  6. Implement retries with backoff and per-document error handling, and make job creation idempotent.
  7. Add lifecycle rules so raw inputs and unredacted outputs expire on schedule.
Key takeaway: OCI's pretrained AI services are APIs, not models you host. Language and small Vision or Document requests are synchronous; audio, image sets and large documents go through jobs that read from and write to Object Storage. Grant access per service family plus object permissions for jobs, use resource or instance principals, handle per-document errors and rate limits, size batches from the current limits pages, migrate off the retired Anomaly Detection and Data Labeling services, and evaluate accuracy on your own data before you build on it.