OCI Generative AI is Oracle Cloud Infrastructure's managed service for large language models. You call hosted chat, embedding and rerank models through an API, and optionally fine-tune and host your own copies on GPU capacity reserved for your tenancy. It sits beside OCI's pretrained task services (language, vision, speech, document), which OCI AI Services covers and which deliberately exclude Generative AI.
This article explains how the service is put together and how to build on it: the two serving modes, the native and OpenAI-compatible APIs with working Python, identity and network controls, embeddings for retrieval, fine-tuning, and the operational traps. The catalogue of hosted models, region availability and prices change often, so this page names only what Oracle's documentation confirmed on 2026-10-04 and tells you where to check the rest.
What the service offers
Oracle's overview groups the service into a few capabilities:
| Capability | What it does | Typical use |
|---|---|---|
| Chat | conversational generation, tool calling, structured output | assistants, extraction, drafting |
| Embeddings | turn text (and images, for newer models) into vectors | semantic search, clustering, classification |
| Rerank | order documents by relevance to a query | second stage of retrieval |
| OpenAI-compatible APIs | Responses, Conversations and Chat Completions endpoints | reuse OpenAI SDKs and tools |
| Agent building blocks | files, vector stores, containers, tools including MCP calling | hosted agentic applications |
| Guardrails | runtime safety and compliance controls on inputs and outputs | content moderation, policy |
The hosted models come from several providers, and Oracle model IDs carry a provider prefix (Oracle's own sample uses openai.gpt-oss-120b). Which models are available depends on region and serving mode, so list them in your region before designing around one.
On-demand and dedicated serving
On-demand serving is shared capacity. You reference a model by ID and pay for what you send. No setup is needed, but you share throughput with other tenants, you are subject to service limits, and models are retired on Oracle's schedule.
Dedicated AI clusters are GPU capacity that belongs to your tenancy, sized in units whose type depends on the base model. There are two kinds. A hosting cluster serves models through endpoints that you create; you call an endpoint by its OCID. A fine-tuning cluster trains custom models. Oracle's pricing page states a minimum commitment of 744 unit-hours per hosting cluster (a 31-day month of one unit), a minimum of 1 unit-hour per fine-tuning job, that some models need at least 2 units to fine-tune, and that one hosting cluster can host up to 50 fine-tuned models.
Choose dedicated when you need predictable latency, isolation, fine-tuned models, or volume high enough that a reserved cluster costs less than on-demand calls. Run the comparison with your own token counts against the current price list rather than a rule of thumb. If you want full control over the serving stack instead, OCI GPU instances covers running your own inference servers.
Identity and network access
Access is controlled by OCI IAM policies on Generative AI resource types. The aggregate type is generative-ai-family. Individual types include generative-ai-chat, generative-ai-text-embedding, generative-ai-model (custom models), generative-ai-dedicated-ai-cluster, generative-ai-endpoint and generative-ai-private-endpoint. Oracle recommends keeping family-wide access for administrators and sandboxes and granting narrower types elsewhere.
allow group genai-admins to manage generative-ai-family in compartment ai-prod
allow dynamic-group chat-app to use generative-ai-chat in compartment ai-prod
allow dynamic-group rag-indexer to use generative-ai-text-embedding in compartment ai-prodWorkloads should authenticate as themselves: instance principals for VMs, resource principals for Functions and other managed services, workload identity for Kubernetes. Long-lived user API keys in application config are the usual leak path. The service also offers its own API keys, which Oracle positions for testing and early development, with IAM-based authentication for production. Store any key you do use in OCI Vault. Policy language and dynamic groups are explained in OCI IAM.
For network isolation, the service supports private endpoints, so traffic from a VCN does not need a public path. Oracle also lists Zero Trust Packet Routing among its controls.
Calling chat with the Python SDK
The OCI Python SDK exposes the inference API as GenerativeAiInferenceClient. A chat call wraps a request in ChatDetails, which carries the compartment, a serving mode, and a chat request. GenericChatRequest with api_format="GENERIC" covers most model families. Cohere models also have their own request classes.
import oci
config = oci.config.from_file() # in production, use a resource or instance principal signer
region = "us-chicago-1" # use a region where your model is offered
client = oci.generative_ai_inference.GenerativeAiInferenceClient(
config,
service_endpoint=f"https://inference.generativeai.{region}.oci.oraclecloud.com",
retry_strategy=oci.retry.DEFAULT_RETRY_STRATEGY,
)
m = oci.generative_ai_inference.models
details = m.ChatDetails(
compartment_id=COMPARTMENT_OCID,
serving_mode=m.OnDemandServingMode(model_id=MODEL_ID),
# dedicated: m.DedicatedServingMode(endpoint_id=ENDPOINT_OCID)
chat_request=m.GenericChatRequest(
api_format="GENERIC",
messages=[
m.SystemMessage(content=[m.TextContent(text="Answer from the policy text only.")]),
m.UserMessage(content=[m.TextContent(text=question)]),
],
max_tokens=600,
temperature=0.2,
),
)
resp = client.chat(details)
choice = resp.data.chat_response.choices[0]
print(choice.finish_reason, choice.message.content[0].text)Check finish_reason on every call. A value of length means the answer was cut off at max_tokens, and silently storing a truncated answer is a common bug. The response also carries a usage object with token counts; log it for cost attribution. The request supports streaming, tools, response formats and, for reasoning models, a reasoning_effort field. Check the class reference for your SDK version, because fields are added often.
The OpenAI-compatible endpoints
The service also exposes OpenAI-compatible endpoints at https://inference.generativeai.<region>.oci.oraclecloud.com/openai/v1, including /responses, /conversations and /chat/completions, plus files, vector stores and containers for agent building. This lets existing code written for the OpenAI SDK, or frameworks built on it, point at OCI by changing the base URL and credentials.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://inference.generativeai.us-chicago-1.oci.oraclecloud.com/openai/v1",
api_key=os.environ["OCI_GENAI_API_KEY"], # a Generative AI API key, kept in Vault
)
# Your tenancy may need further client settings; follow Oracle's current setup guide.
r = client.responses.create(model="openai.gpt-oss-120b", input="Summarise this incident: ...")
print(r.output_text)Compatibility is not identity. Model coverage differs between the native and compatible APIs, some parameters are ignored or rejected, and authentication works differently from OpenAI's. For production, Oracle documents IAM-based authentication for these endpoints; follow its current setup guide rather than shipping a static key. Test the exact features you depend on (streaming, tool calls, structured output) against the model you will use. The same advice applies when you compare with Azure OpenAI, which raises the same questions of parity with the OpenAI API.
Embeddings for retrieval
Retrieval-augmented generation (RAG) on OCI usually means embedding documents with the service, storing vectors in a database you run, and sending retrieved passages to a chat model. The service does not own your index.
def embed(texts, input_type):
out = []
for i in range(0, len(texts), 64): # batch; check your model's per-call limit
r = client.embed_text(m.EmbedTextDetails(
compartment_id=COMPARTMENT_OCID,
serving_mode=m.OnDemandServingMode(model_id=EMBED_MODEL_ID),
inputs=texts[i:i + 64],
input_type=input_type, # SEARCH_DOCUMENT or SEARCH_QUERY
truncate="NONE", # fail loudly instead of cutting text
))
out.extend(r.data.embeddings)
return outThree details matter. Embed documents with SEARCH_DOCUMENT and queries with SEARCH_QUERY; mixing them up quietly lowers recall. Each input is limited to 512 tokens, and truncate accepts NONE, START or END, so chunk your documents below the limit and use NONE so oversize chunks fail instead of losing their tails. Record the embedding model ID next to every stored vector: when that model is retired, you must re-embed the whole corpus, and the ID tells you which rows to redo. Newer embedding models also accept output_dimensions (256, 512, 1024 or 1536), which trades a little accuracy for a smaller index.
Fine-tuning and hosting custom models
Fine-tuning runs on a fine-tuning dedicated cluster. Training data is a JSONL file in Object Storage with at least 32 prompt/completion pairs, one per line:
{"prompt": "Classify the ticket: VPN drops every 10 minutes", "completion": "network"}
{"prompt": "Classify the ticket: cannot open payroll report", "completion": "finance-app"}Oracle offers parameter-efficient methods whose availability depends on the base model, and the list changes as models are added and retired, so check the fine-tuning page for your base model before planning. The result is a custom model you host on a hosting cluster and call through an endpoint. Hold out an evaluation set and compare the tuned model with the base model plus a good prompt. For many classification and formatting tasks a careful prompt and a few examples come close, without paying for a cluster.
Worked example: a policy help desk
Suppose an internal help desk wants answers from 20,000 policy pages.
- Chunk pages into passages of about 300 tokens with some overlap: roughly 60,000 chunks if pages average about 900 tokens.
- Embed the chunks with
SEARCH_DOCUMENTin batches, using a resource principal from an indexing job. Store each vector with its text, source URL and embedding model ID. - At query time, embed the question with
SEARCH_QUERY, fetch the top 40 by vector similarity, rerank them, and keep the top 5. - Call chat with a system message that says to answer only from the supplied passages and to cite them, at a low temperature. Reject answers that cite nothing.
- Log user, model ID, token usage, latency and
finish_reason; never log full prompts that may contain personal data.
Why rerank? Vector similarity compares a query and a passage that were embedded separately. That makes it fast enough to scan the whole index, but it misses fine distinctions such as negation, or which product version a passage covers. A reranker reads the query and each candidate together and scores them as a pair. That is far more accurate, and far too slow to run on 60,000 chunks. Use vector search for recall and the reranker on the top 40 for precision, and you get most of the accuracy at a small cost. Measure the two stages separately on your evaluation set. If the right passage is missing from the top 40, no reranker can recover it.
Start on-demand. If steady traffic, latency targets or a tuned model later justify it, move chat to a dedicated endpoint by changing the serving mode, which is a configuration change rather than a rewrite.
Failure modes
- Model retirement. Hosted models are deprecated and retired on a published schedule. Pin model IDs in config, subscribe to the notices, and keep an evaluation set so a swap is a measured change.
- Wrong region. Not every model is offered in every region. Calls to a region without the model fail even though IAM is correct.
- Throttling. On-demand capacity is shared and limited. Use the SDK retry strategy with backoff, cap concurrency in your app, and request limit increases before launch.
- Truncated output.
finish_reasonoflengthstored as a complete answer. - Over-broad policies.
manage generative-ai-family in tenancygranted to an application identity lets it create clusters that bill for at least 744 unit-hours. - Prompt injection through retrieved text. Retrieved documents are untrusted input. Guardrails help, but the main defence is giving the model no tools or permissions it does not need.
Trade-offs
The native API exposes every OCI feature and integrates with OCI signing; the OpenAI-compatible API makes code portable but limits you to what the compatibility layer supports. On-demand is cheap to start and needs no commitment; dedicated clusters give isolation and steady latency for a monthly minimum. Fine-tuning can beat prompting on narrow tasks but adds a training pipeline, a hosting bill and retraining when the base model is retired. Compared with running open models on your own GPU instances, the managed service removes serving work and adds dependence on Oracle's model catalogue and schedule. Weigh these against where your data already lives: if it is in OCI, keeping inference in the same region and tenancy simplifies networking and compliance.
What to do next
- List the chat and embedding models available in your target region and serving mode, and note their retirement dates.
- Create a compartment and narrow IAM policies, using dynamic groups for workloads instead of user keys.
- Run the SDK chat example with a resource or instance principal, and add
finish_reasonandusagelogging. - If you have OpenAI-based code, point it at the compatible endpoint in a test environment and check streaming, tools and structured output.
- Build a small RAG index with
SEARCH_DOCUMENTandSEARCH_QUERY, storing the embedding model ID with every vector. - Write a 50-question evaluation set before considering fine-tuning or a dedicated cluster.
- Estimate monthly tokens and compare on-demand cost with the dedicated-cluster minimum using the current price list.