Gemini models are served through two front doors. The Gemini Developer API, managed from Google AI Studio, is keyed by an API key and built for getting started quickly. The enterprise path, which ran as Vertex AI until Google renamed it Gemini Enterprise Agent Platform in April 2026, sits inside a Google Cloud project (platform overview: Vertex AI on GCP) with IAM, VPC Service Controls, regional endpoints and Cloud billing. Most teams start on the first and ship on the second, and most of the pain sits between the two.
This article explains how the two paths differ at the protocol level, how one codebase can target both through the Google Gen AI SDK, which features do not port, and how to run Gemini in production: retries, structured output, caching, batch, thinking budgets and cost accounting. Details were checked against the google-genai Python SDK source and changelog (version 2.28.0, 2 October 2026). Model IDs, quotas and prices change monthly, so this article names none as current; read them from the model list in your own project.
Two paths, four differences
The SDK makes the backend a constructor flag, and that flag decides four things.
| Gemini Developer API | Enterprise Agent Platform (Vertex AI) | |
|---|---|---|
| Client | genai.Client(api_key=...) | genai.Client(enterprise=True, project=..., location=...) |
| Host and version | generativelanguage.googleapis.com, v1beta | aiplatform.googleapis.com for global, {location}-aiplatform.googleapis.com for regions; v1beta1 |
| Auth | API key in a header | Application Default Credentials and IAM; API keys are also accepted (express mode) |
| Large inputs | Files API uploads, referenced as files/... URIs | gs:// URIs; batch also reads bq:// tables |
| Environment variables | GEMINI_API_KEY or GOOGLE_API_KEY (the latter wins if both are set) | GOOGLE_GENAI_USE_ENTERPRISE=true, GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION |
Older code uses vertexai=True and GOOGLE_GENAI_USE_VERTEXAI. Both still work: the constructor treats vertexai as a legacy alias for enterprise and raises a ValueError if you pass both with different values. If the two environment variables conflict, GOOGLE_GENAI_USE_ENTERPRISE wins with a warning. Migrate to the new names when convenient, but do not set both.
Location has a quiet default. On the enterprise path, if you give neither a location nor an API key, the SDK uses global, which routes to the non-regional host. That is good for availability, because Google can serve the request from any region with capacity. It is wrong if you have data-residency obligations, which need a regional location set explicitly. Make the location a required setting in your own config layer rather than relying on the default.
One client for both
The request and response types are shared, so the portable pattern is to construct the client in one place and pass it everywhere else. The model name lives in configuration, because aliases such as gemini-flash-latest (used throughout the SDK README) move to new versions without notice. Pin a versioned ID in production and change it deliberately.
import os
from google import genai
from google.genai import types
MODEL = os.environ["GEMINI_MODEL"] # pin a versioned model ID in prod; aliases move
def make_client() -> genai.Client:
http = types.HttpOptions(
timeout=60_000, # milliseconds
retry_options=types.HttpRetryOptions(
attempts=4, # includes the first try
initial_delay=1.0, max_delay=30.0, exp_base=2.0, jitter=1.0,
http_status_codes=[408, 429, 500, 502, 503, 504],
),
)
if os.environ.get("GEMINI_BACKEND") == "enterprise":
return genai.Client(
enterprise=True,
project=os.environ["GOOGLE_CLOUD_PROJECT"],
location=os.environ["GOOGLE_CLOUD_LOCATION"], # explicit: no silent 'global'
http_options=http,
)
return genai.Client(api_key=os.environ["GEMINI_API_KEY"], http_options=http)
client = make_client()
resp = client.models.generate_content(model=MODEL, contents="Name three HBase write-path metrics.")
print(resp.text, resp.usage_metadata)Three details in that block matter in production. The timeout is in milliseconds, not seconds. Retry attempts counts the original request, so 4 means three retries. And the retry list should include 429, because quota exhaustion is the most common transient failure on both backends; the SDK default covers 408, 429 and 5xx when you leave the list unset, but writing it out documents the intent. For errors that survive retries, catch google.genai.errors.APIError and branch on e.code.
What does not port
The common calls work on both paths: generate_content and its streaming variant, chats, count_tokens, function calling, structured output and context caching. The differences show up where data enters and leaves.
- Files.
client.files.uploadexists only on the Developer API. The enterprise path takesgs://URIs throughtypes.Part.from_uri. A portable helper takes a local path, uploads to the Files API on one backend and to a Cloud Storage bucket on the other, and returns aPart. - Batch.
client.batches.createtakes ags://orbq://source on the enterprise path, and inlined requests or an uploaded JSONL file (files/...) on the Developer API. Output locations differ in the same way. Write the job submitter twice; it is short. - Platform features. Model Garden third-party models, provisioned throughput, VPC Service Controls perimeters, customer-managed encryption keys and Cloud audit logs exist only on the enterprise side. If your design depends on any of them, prototype there from day one.
- Unreleased features. Preview features sometimes land on one backend first. The SDK's
HttpOptions(extra_body=...)can pass raw JSON fields the SDK does not model yet; treat anything sent that way as unstable.
Run your integration tests against both backends in CI. A test suite that only exercises AI Studio keys will pass right up until the day you flip GEMINI_BACKEND.
Structured output and thinking
For extraction and classification, ask for JSON against a schema rather than parsing prose. The SDK accepts a standard JSON Schema through response_json_schema, and recent SDK versions recommend it over the older response_schema field. Pydantic models give you the schema and the validator in one place.
from pydantic import BaseModel, ValidationError
class Invoice(BaseModel):
vendor: str
invoice_number: str
total_minor_units: int # integers, not floats, for money
currency: str
def extract(pdf_part: types.Part) -> Invoice:
resp = client.models.generate_content(
model=MODEL,
contents=[pdf_part, "Extract the invoice header."],
config=types.GenerateContentConfig(
response_mime_type="application/json",
response_json_schema=Invoice.model_json_schema(),
temperature=0.0,
),
)
if not resp.candidates: # the prompt itself was blocked
raise RuntimeError(f"blocked: {resp.prompt_feedback}")
try:
return Invoice.model_validate_json(resp.text)
except ValidationError:
# schema-constrained output can still be truncated or empty: check finish_reason
raise RuntimeError(f"bad output, finish={resp.candidates[0].finish_reason}")Constrained decoding guarantees syntax, not truth. The model can still return a well-formed invoice number that is wrong, so validate business rules (totals add up, currency is in your allowlist) after parsing. Always check finish_reason: an output cut off by max_output_tokens or a safety block surfaces as invalid JSON or an empty text. For a wider treatment of schema design, see structured output with LLMs.
On thinking: ThinkingConfig has thinking_budget in tokens, where 0 disables thinking and -1 asks for automatic, with allowed ranges model-dependent, and a newer thinking_level enum (MINIMAL, LOW, MEDIUM, HIGH). Which one a given model honours, and whether it can be disabled, varies by model generation, so check the model card. Thought tokens are billed as output and reported in usage_metadata.thoughts_token_count.
Cost: measure, cache, batch
Gemini cost is token arithmetic, and usage_metadata on every response gives you the inputs: prompt_token_count, candidates_token_count, thoughts_token_count, cached_content_token_count and total_token_count. Log all five per request, tagged with tenant and feature, before you optimise anything. Three levers then do most of the work.
- Context caching. When many requests share a long prefix such as a contract, a codebase or a product manual, create a cache once with
client.caches.createand a TTL, then passcached_content=cache.namein each request's config. Cached tokens are billed at a reduced rate plus storage per hour of TTL, so caching pays off only when the prefix is reused often enough within the TTL. Calculate the break-even from current prices for your model and backend; the mechanics are in prompt caching. - Batch. Work that can wait hours, such as backfills, evaluations and nightly classification, goes through batch jobs. Google prices batch below interactive calls, and moving bulk work off the online path keeps it from competing with user-facing requests for quota.
- Thinking control. For extraction and routing tasks, a low thinking setting often costs a fraction of the default with no quality loss. Measure that on your own evaluation set rather than assuming it.
# manual_part: a Part for the manual (gs:// on enterprise, files/... on the Developer API)
cache = client.caches.create(
model=MODEL,
config=types.CreateCachedContentConfig(
contents=[types.Content(role="user", parts=[manual_part])],
system_instruction="Answer only from the attached manual. Cite section numbers.",
display_name="support-manual-v14",
ttl="3600s",
),
)
answer = client.models.generate_content(
model=MODEL, contents=question,
config=types.GenerateContentConfig(cached_content=cache.name),
)
assert answer.usage_metadata.cached_content_token_count, "cache not applied"
Worked example: prototype to production
A team builds a support assistant that answers from a 400-page product manual. They prototype in AI Studio with a key, upload the manual with the Files API, and get good answers in a day. Then production arrives with four new requirements: customer data must stay in the EU, the security team wants no long-lived keys, finance wants per-tenant cost, and peak traffic is 30 requests per second.
The migration follows the table above. The manual moves from the Files API to a Cloud Storage bucket in an EU region; the client switches to enterprise=True with an explicit EU location rather than global; the service runs as a dedicated service account with only the permissions needed to call models and read that bucket, granted through IAM as described in GCP IAM; and the project goes inside a VPC Service Controls perimeter (VPC Service Controls) so the model endpoint cannot be reached with stolen credentials from outside it. The manual becomes a context cache refreshed on each release, and every response's usage_metadata goes to a log sink with the tenant ID, where a scheduled query turns tokens into money.
The step that bit them was quota. AI Studio quotas and Cloud project quotas are separate systems, and the regional quota in the new project was lower than prototype traffic. They found out from 429s in load testing, raised the request through the Cloud console, and added a client-side token bucket so a burst queues briefly instead of tripping retries. Teams with hard latency targets at steady high volume should also evaluate provisioned throughput on the enterprise path.
Failure modes and trade-offs
| Symptom | Likely cause | First response |
|---|---|---|
| 429 RESOURCE_EXHAUSTED under load | Per-model, per-region or per-key quota | Backoff with jitter, client-side rate limit, quota increase, spread across regions only if residency allows |
ValueError at client construction | vertexai and enterprise both set and conflicting, or no project or key found | Set one flag; set GOOGLE_CLOUD_PROJECT or pass project= |
| 403 on the enterprise path | Service account lacks the model-calling role, or a VPC-SC perimeter blocks the caller | Check IAM policy and perimeter audit logs |
Empty text or invalid JSON | Safety block, truncation at max_output_tokens, or thinking consumed the output budget | Log finish_reason and token counts; raise the limit or lower thinking |
| Works in dev, 404 in prod | Model ID not available in the chosen region or on that backend | List models in the prod project; pin IDs per environment |
| Cost spike without traffic growth | A prompt change defeated the cache, or thinking tokens grew | Alert on cached and thought tokens per request, not only totals |
Two trade-offs deserve a decision rather than a default. The global endpoint improves availability and capacity but gives up regional control. And the Developer API is simpler, but keys are bearer secrets that leak through notebooks and logs; if you stay on it in production, store keys in Secret Manager and rotate them.
What to do next
- Centralise client construction in one function driven by environment variables, with an explicit location and a pinned model ID.
- Replace
vertexai=Truewithenterprise=TrueandGOOGLE_GENAI_USE_VERTEXAIwithGOOGLE_GENAI_USE_ENTERPRISE; never set both. - Configure
HttpRetryOptionswith 429 and 5xx, a millisecond timeout, and a client-side rate limiter for bursts. - Move extraction calls to
response_json_schemawith Pydantic validation andfinish_reasonchecks. - Log all five
usage_metadatacounters per request with tenant and feature tags. - Identify shared long prefixes and test context caching; move anything that can wait to batch.
- Run integration tests against both backends in CI, and check quotas in the production project before launch.