OCI Generative AI is Oracle Cloud Infrastructure's managed service for large language models. You call hosted chat, embedding and rerank models through an API, and optionally fine-tune and host your own copies on GPU capacity reserved for your tenancy. It sits beside OCI's pretrained task services (language, vision, speech, document), which OCI AI Services covers and which deliberately exclude Generative AI.

This article explains how the service is put together and how to build on it: the two serving modes, the native and OpenAI-compatible APIs with working Python, identity and network controls, embeddings for retrieval, fine-tuning, and the operational traps. The catalogue of hosted models, region availability and prices change often, so this page names only what Oracle's documentation confirmed on 2026-10-04 and tells you where to check the rest.

What the service offers

Oracle's overview groups the service into a few capabilities:

CapabilityWhat it doesTypical use
Chatconversational generation, tool calling, structured outputassistants, extraction, drafting
Embeddingsturn text (and images, for newer models) into vectorssemantic search, clustering, classification
Rerankorder documents by relevance to a querysecond stage of retrieval
OpenAI-compatible APIsResponses, Conversations and Chat Completions endpointsreuse OpenAI SDKs and tools
Agent building blocksfiles, vector stores, containers, tools including MCP callinghosted agentic applications
Guardrailsruntime safety and compliance controls on inputs and outputscontent moderation, policy

The hosted models come from several providers, and Oracle model IDs carry a provider prefix (Oracle's own sample uses openai.gpt-oss-120b). Which models are available depends on region and serving mode, so list them in your region before designing around one.

OCI Generative AI: two ways in, two ways to serveYour appOKE, Functions, VMIdentityresource / instance principalNative inference APIchat, embedText, rerankOpenAI-compatible API/openai/v1: responses, chatOn-demandshared, pay per useDedicated AI clusterhosting units, endpointsmodel IDendpoint IDFine-tuning clusterJSONL in Object Storagecustom modelControlsIAM, private endpoint, guardrailsVector store + app logicRAG lives hereOn-demand: pick a model ID and pay for what you send. Dedicated: commit to units, create endpoints, host fine-tuned models.Both are reachable through either API, subject to which models each API supports.
How requests reach a model: an identity, one of two APIs, and a serving mode.

On-demand and dedicated serving

On-demand serving is shared capacity. You reference a model by ID and pay for what you send. No setup is needed, but you share throughput with other tenants, you are subject to service limits, and models are retired on Oracle's schedule.

Dedicated AI clusters are GPU capacity that belongs to your tenancy, sized in units whose type depends on the base model. There are two kinds. A hosting cluster serves models through endpoints that you create; you call an endpoint by its OCID. A fine-tuning cluster trains custom models. Oracle's pricing page states a minimum commitment of 744 unit-hours per hosting cluster (a 31-day month of one unit), a minimum of 1 unit-hour per fine-tuning job, that some models need at least 2 units to fine-tune, and that one hosting cluster can host up to 50 fine-tuned models.

Choose dedicated when you need predictable latency, isolation, fine-tuned models, or volume high enough that a reserved cluster costs less than on-demand calls. Run the comparison with your own token counts against the current price list rather than a rule of thumb. If you want full control over the serving stack instead, OCI GPU instances covers running your own inference servers.

Identity and network access

Access is controlled by OCI IAM policies on Generative AI resource types. The aggregate type is generative-ai-family. Individual types include generative-ai-chat, generative-ai-text-embedding, generative-ai-model (custom models), generative-ai-dedicated-ai-cluster, generative-ai-endpoint and generative-ai-private-endpoint. Oracle recommends keeping family-wide access for administrators and sandboxes and granting narrower types elsewhere.

allow group genai-admins to manage generative-ai-family in compartment ai-prod
allow dynamic-group chat-app to use generative-ai-chat in compartment ai-prod
allow dynamic-group rag-indexer to use generative-ai-text-embedding in compartment ai-prod

Workloads should authenticate as themselves: instance principals for VMs, resource principals for Functions and other managed services, workload identity for Kubernetes. Long-lived user API keys in application config are the usual leak path. The service also offers its own API keys, which Oracle positions for testing and early development, with IAM-based authentication for production. Store any key you do use in OCI Vault. Policy language and dynamic groups are explained in OCI IAM.

For network isolation, the service supports private endpoints, so traffic from a VCN does not need a public path. Oracle also lists Zero Trust Packet Routing among its controls.

Calling chat with the Python SDK

The OCI Python SDK exposes the inference API as GenerativeAiInferenceClient. A chat call wraps a request in ChatDetails, which carries the compartment, a serving mode, and a chat request. GenericChatRequest with api_format="GENERIC" covers most model families. Cohere models also have their own request classes.

import oci

config = oci.config.from_file()   # in production, use a resource or instance principal signer
region = "us-chicago-1"           # use a region where your model is offered
client = oci.generative_ai_inference.GenerativeAiInferenceClient(
    config,
    service_endpoint=f"https://inference.generativeai.{region}.oci.oraclecloud.com",
    retry_strategy=oci.retry.DEFAULT_RETRY_STRATEGY,
)
m = oci.generative_ai_inference.models

details = m.ChatDetails(
    compartment_id=COMPARTMENT_OCID,
    serving_mode=m.OnDemandServingMode(model_id=MODEL_ID),
    # dedicated: m.DedicatedServingMode(endpoint_id=ENDPOINT_OCID)
    chat_request=m.GenericChatRequest(
        api_format="GENERIC",
        messages=[
            m.SystemMessage(content=[m.TextContent(text="Answer from the policy text only.")]),
            m.UserMessage(content=[m.TextContent(text=question)]),
        ],
        max_tokens=600,
        temperature=0.2,
    ),
)
resp = client.chat(details)
choice = resp.data.chat_response.choices[0]
print(choice.finish_reason, choice.message.content[0].text)

Check finish_reason on every call. A value of length means the answer was cut off at max_tokens, and silently storing a truncated answer is a common bug. The response also carries a usage object with token counts; log it for cost attribution. The request supports streaming, tools, response formats and, for reasoning models, a reasoning_effort field. Check the class reference for your SDK version, because fields are added often.

The OpenAI-compatible endpoints

The service also exposes OpenAI-compatible endpoints at https://inference.generativeai.<region>.oci.oraclecloud.com/openai/v1, including /responses, /conversations and /chat/completions, plus files, vector stores and containers for agent building. This lets existing code written for the OpenAI SDK, or frameworks built on it, point at OCI by changing the base URL and credentials.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://inference.generativeai.us-chicago-1.oci.oraclecloud.com/openai/v1",
    api_key=os.environ["OCI_GENAI_API_KEY"],   # a Generative AI API key, kept in Vault
)
# Your tenancy may need further client settings; follow Oracle's current setup guide.
r = client.responses.create(model="openai.gpt-oss-120b", input="Summarise this incident: ...")
print(r.output_text)

Compatibility is not identity. Model coverage differs between the native and compatible APIs, some parameters are ignored or rejected, and authentication works differently from OpenAI's. For production, Oracle documents IAM-based authentication for these endpoints; follow its current setup guide rather than shipping a static key. Test the exact features you depend on (streaming, tool calls, structured output) against the model you will use. The same advice applies when you compare with Azure OpenAI, which raises the same questions of parity with the OpenAI API.

Embeddings for retrieval

Retrieval-augmented generation (RAG) on OCI usually means embedding documents with the service, storing vectors in a database you run, and sending retrieved passages to a chat model. The service does not own your index.

def embed(texts, input_type):
    out = []
    for i in range(0, len(texts), 64):          # batch; check your model's per-call limit
        r = client.embed_text(m.EmbedTextDetails(
            compartment_id=COMPARTMENT_OCID,
            serving_mode=m.OnDemandServingMode(model_id=EMBED_MODEL_ID),
            inputs=texts[i:i + 64],
            input_type=input_type,                # SEARCH_DOCUMENT or SEARCH_QUERY
            truncate="NONE",                      # fail loudly instead of cutting text
        ))
        out.extend(r.data.embeddings)
    return out

Three details matter. Embed documents with SEARCH_DOCUMENT and queries with SEARCH_QUERY; mixing them up quietly lowers recall. Each input is limited to 512 tokens, and truncate accepts NONE, START or END, so chunk your documents below the limit and use NONE so oversize chunks fail instead of losing their tails. Record the embedding model ID next to every stored vector: when that model is retired, you must re-embed the whole corpus, and the ID tells you which rows to redo. Newer embedding models also accept output_dimensions (256, 512, 1024 or 1536), which trades a little accuracy for a smaller index.

Fine-tuning and hosting custom models

Fine-tuning runs on a fine-tuning dedicated cluster. Training data is a JSONL file in Object Storage with at least 32 prompt/completion pairs, one per line:

{"prompt": "Classify the ticket: VPN drops every 10 minutes", "completion": "network"}
{"prompt": "Classify the ticket: cannot open payroll report", "completion": "finance-app"}

Oracle offers parameter-efficient methods whose availability depends on the base model, and the list changes as models are added and retired, so check the fine-tuning page for your base model before planning. The result is a custom model you host on a hosting cluster and call through an endpoint. Hold out an evaluation set and compare the tuned model with the base model plus a good prompt. For many classification and formatting tasks a careful prompt and a few examples come close, without paying for a cluster.

Worked example: a policy help desk

Suppose an internal help desk wants answers from 20,000 policy pages.

  1. Chunk pages into passages of about 300 tokens with some overlap: roughly 60,000 chunks if pages average about 900 tokens.
  2. Embed the chunks with SEARCH_DOCUMENT in batches, using a resource principal from an indexing job. Store each vector with its text, source URL and embedding model ID.
  3. At query time, embed the question with SEARCH_QUERY, fetch the top 40 by vector similarity, rerank them, and keep the top 5.
  4. Call chat with a system message that says to answer only from the supplied passages and to cite them, at a low temperature. Reject answers that cite nothing.
  5. Log user, model ID, token usage, latency and finish_reason; never log full prompts that may contain personal data.

Why rerank? Vector similarity compares a query and a passage that were embedded separately. That makes it fast enough to scan the whole index, but it misses fine distinctions such as negation, or which product version a passage covers. A reranker reads the query and each candidate together and scores them as a pair. That is far more accurate, and far too slow to run on 60,000 chunks. Use vector search for recall and the reranker on the top 40 for precision, and you get most of the accuracy at a small cost. Measure the two stages separately on your evaluation set. If the right passage is missing from the top 40, no reranker can recover it.

Start on-demand. If steady traffic, latency targets or a tuned model later justify it, move chat to a dedicated endpoint by changing the serving mode, which is a configuration change rather than a rewrite.

Failure modes

  • Model retirement. Hosted models are deprecated and retired on a published schedule. Pin model IDs in config, subscribe to the notices, and keep an evaluation set so a swap is a measured change.
  • Wrong region. Not every model is offered in every region. Calls to a region without the model fail even though IAM is correct.
  • Throttling. On-demand capacity is shared and limited. Use the SDK retry strategy with backoff, cap concurrency in your app, and request limit increases before launch.
  • Truncated output. finish_reason of length stored as a complete answer.
  • Over-broad policies. manage generative-ai-family in tenancy granted to an application identity lets it create clusters that bill for at least 744 unit-hours.
  • Prompt injection through retrieved text. Retrieved documents are untrusted input. Guardrails help, but the main defence is giving the model no tools or permissions it does not need.

Trade-offs

The native API exposes every OCI feature and integrates with OCI signing; the OpenAI-compatible API makes code portable but limits you to what the compatibility layer supports. On-demand is cheap to start and needs no commitment; dedicated clusters give isolation and steady latency for a monthly minimum. Fine-tuning can beat prompting on narrow tasks but adds a training pipeline, a hosting bill and retraining when the base model is retired. Compared with running open models on your own GPU instances, the managed service removes serving work and adds dependence on Oracle's model catalogue and schedule. Weigh these against where your data already lives: if it is in OCI, keeping inference in the same region and tenancy simplifies networking and compliance.

What to do next

  1. List the chat and embedding models available in your target region and serving mode, and note their retirement dates.
  2. Create a compartment and narrow IAM policies, using dynamic groups for workloads instead of user keys.
  3. Run the SDK chat example with a resource or instance principal, and add finish_reason and usage logging.
  4. If you have OpenAI-based code, point it at the compatible endpoint in a test environment and check streaming, tools and structured output.
  5. Build a small RAG index with SEARCH_DOCUMENT and SEARCH_QUERY, storing the embedding model ID with every vector.
  6. Write a 50-question evaluation set before considering fine-tuning or a dedicated cluster.
  7. Estimate monthly tokens and compare on-demand cost with the dedicated-cluster minimum using the current price list.
Key takeaway: OCI Generative AI serves hosted chat, embedding and rerank models on demand, or your own fine-tuned models on dedicated clusters with a 744 unit-hour hosting minimum. Call it through the native SDK or the OpenAI-compatible endpoint, authenticate workloads with principals and narrow policies, check finish_reason and usage on every call, store embedding model IDs with vectors, and plan for model retirement from day one.