Gemini models are served through two front doors. The Gemini Developer API, managed from Google AI Studio, is keyed by an API key and built for getting started quickly. The enterprise path, which ran as Vertex AI until Google renamed it Gemini Enterprise Agent Platform in April 2026, sits inside a Google Cloud project (platform overview: Vertex AI on GCP) with IAM, VPC Service Controls, regional endpoints and Cloud billing. Most teams start on the first and ship on the second, and most of the pain sits between the two.

This article explains how the two paths differ at the protocol level, how one codebase can target both through the Google Gen AI SDK, which features do not port, and how to run Gemini in production: retries, structured output, caching, batch, thinking budgets and cost accounting. Details were checked against the google-genai Python SDK source and changelog (version 2.28.0, 2 October 2026). Model IDs, quotas and prices change monthly, so this article names none as current; read them from the model list in your own project.

Advertisement

Two paths, four differences

The SDK makes the backend a constructor flag, and that flag decides four things.

Gemini Developer APIEnterprise Agent Platform (Vertex AI)
Clientgenai.Client(api_key=...)genai.Client(enterprise=True, project=..., location=...)
Host and versiongenerativelanguage.googleapis.com, v1betaaiplatform.googleapis.com for global, {location}-aiplatform.googleapis.com for regions; v1beta1
AuthAPI key in a headerApplication Default Credentials and IAM; API keys are also accepted (express mode)
Large inputsFiles API uploads, referenced as files/... URIsgs:// URIs; batch also reads bq:// tables
Environment variablesGEMINI_API_KEY or GOOGLE_API_KEY (the latter wins if both are set)GOOGLE_GENAI_USE_ENTERPRISE=true, GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION

Older code uses vertexai=True and GOOGLE_GENAI_USE_VERTEXAI. Both still work: the constructor treats vertexai as a legacy alias for enterprise and raises a ValueError if you pass both with different values. If the two environment variables conflict, GOOGLE_GENAI_USE_ENTERPRISE wins with a warning. Migrate to the new names when convenient, but do not set both.

Location has a quiet default. On the enterprise path, if you give neither a location nor an API key, the SDK uses global, which routes to the non-regional host. That is good for availability, because Google can serve the request from any region with capacity. It is wrong if you have data-residency obligations, which need a regional location set explicitly. Make the location a required setting in your own config layer rather than relying on the default.

Your servicegoogle-genai ClientGemini Developer APIgenerativelanguage.googleapis.comEnterprise Agent Platformaiplatform.googleapis.comapi_key, v1betaproject + location, v1beta1API keyAI StudioFiles APIuploads, 'files/...'IAM + ADCservice accountsgs:// and bq://inputs, batchglobalaiplatform...us-central1 etc.{loc}-aiplatform...Same request and response types on both pathsgenerate_content, caches, count_tokens, chats
One SDK, two backends. The client flag picks the host, API version, auth model and where large inputs live; the request and response objects are shared.

One client for both

The request and response types are shared, so the portable pattern is to construct the client in one place and pass it everywhere else. The model name lives in configuration, because aliases such as gemini-flash-latest (used throughout the SDK README) move to new versions without notice. Pin a versioned ID in production and change it deliberately.

import os
from google import genai
from google.genai import types

MODEL = os.environ["GEMINI_MODEL"]   # pin a versioned model ID in prod; aliases move

def make_client() -> genai.Client:
    http = types.HttpOptions(
        timeout=60_000,                              # milliseconds
        retry_options=types.HttpRetryOptions(
            attempts=4,                              # includes the first try
            initial_delay=1.0, max_delay=30.0, exp_base=2.0, jitter=1.0,
            http_status_codes=[408, 429, 500, 502, 503, 504],
        ),
    )
    if os.environ.get("GEMINI_BACKEND") == "enterprise":
        return genai.Client(
            enterprise=True,
            project=os.environ["GOOGLE_CLOUD_PROJECT"],
            location=os.environ["GOOGLE_CLOUD_LOCATION"],   # explicit: no silent 'global'
            http_options=http,
        )
    return genai.Client(api_key=os.environ["GEMINI_API_KEY"], http_options=http)

client = make_client()
resp = client.models.generate_content(model=MODEL, contents="Name three HBase write-path metrics.")
print(resp.text, resp.usage_metadata)

Three details in that block matter in production. The timeout is in milliseconds, not seconds. Retry attempts counts the original request, so 4 means three retries. And the retry list should include 429, because quota exhaustion is the most common transient failure on both backends; the SDK default covers 408, 429 and 5xx when you leave the list unset, but writing it out documents the intent. For errors that survive retries, catch google.genai.errors.APIError and branch on e.code.

Advertisement

What does not port

The common calls work on both paths: generate_content and its streaming variant, chats, count_tokens, function calling, structured output and context caching. The differences show up where data enters and leaves.

  • Files. client.files.upload exists only on the Developer API. The enterprise path takes gs:// URIs through types.Part.from_uri. A portable helper takes a local path, uploads to the Files API on one backend and to a Cloud Storage bucket on the other, and returns a Part.
  • Batch. client.batches.create takes a gs:// or bq:// source on the enterprise path, and inlined requests or an uploaded JSONL file (files/...) on the Developer API. Output locations differ in the same way. Write the job submitter twice; it is short.
  • Platform features. Model Garden third-party models, provisioned throughput, VPC Service Controls perimeters, customer-managed encryption keys and Cloud audit logs exist only on the enterprise side. If your design depends on any of them, prototype there from day one.
  • Unreleased features. Preview features sometimes land on one backend first. The SDK's HttpOptions(extra_body=...) can pass raw JSON fields the SDK does not model yet; treat anything sent that way as unstable.

Run your integration tests against both backends in CI. A test suite that only exercises AI Studio keys will pass right up until the day you flip GEMINI_BACKEND.

Structured output and thinking

For extraction and classification, ask for JSON against a schema rather than parsing prose. The SDK accepts a standard JSON Schema through response_json_schema, and recent SDK versions recommend it over the older response_schema field. Pydantic models give you the schema and the validator in one place.

from pydantic import BaseModel, ValidationError

class Invoice(BaseModel):
    vendor: str
    invoice_number: str
    total_minor_units: int        # integers, not floats, for money
    currency: str

def extract(pdf_part: types.Part) -> Invoice:
    resp = client.models.generate_content(
        model=MODEL,
        contents=[pdf_part, "Extract the invoice header."],
        config=types.GenerateContentConfig(
            response_mime_type="application/json",
            response_json_schema=Invoice.model_json_schema(),
            temperature=0.0,
        ),
    )
    if not resp.candidates:                       # the prompt itself was blocked
        raise RuntimeError(f"blocked: {resp.prompt_feedback}")
    try:
        return Invoice.model_validate_json(resp.text)
    except ValidationError:
        # schema-constrained output can still be truncated or empty: check finish_reason
        raise RuntimeError(f"bad output, finish={resp.candidates[0].finish_reason}")

Constrained decoding guarantees syntax, not truth. The model can still return a well-formed invoice number that is wrong, so validate business rules (totals add up, currency is in your allowlist) after parsing. Always check finish_reason: an output cut off by max_output_tokens or a safety block surfaces as invalid JSON or an empty text. For a wider treatment of schema design, see structured output with LLMs.

On thinking: ThinkingConfig has thinking_budget in tokens, where 0 disables thinking and -1 asks for automatic, with allowed ranges model-dependent, and a newer thinking_level enum (MINIMAL, LOW, MEDIUM, HIGH). Which one a given model honours, and whether it can be disabled, varies by model generation, so check the model card. Thought tokens are billed as output and reported in usage_metadata.thoughts_token_count.

Cost: measure, cache, batch

Gemini cost is token arithmetic, and usage_metadata on every response gives you the inputs: prompt_token_count, candidates_token_count, thoughts_token_count, cached_content_token_count and total_token_count. Log all five per request, tagged with tenant and feature, before you optimise anything. Three levers then do most of the work.

  1. Context caching. When many requests share a long prefix such as a contract, a codebase or a product manual, create a cache once with client.caches.create and a TTL, then pass cached_content=cache.name in each request's config. Cached tokens are billed at a reduced rate plus storage per hour of TTL, so caching pays off only when the prefix is reused often enough within the TTL. Calculate the break-even from current prices for your model and backend; the mechanics are in prompt caching.
  2. Batch. Work that can wait hours, such as backfills, evaluations and nightly classification, goes through batch jobs. Google prices batch below interactive calls, and moving bulk work off the online path keeps it from competing with user-facing requests for quota.
  3. Thinking control. For extraction and routing tasks, a low thinking setting often costs a fraction of the default with no quality loss. Measure that on your own evaluation set rather than assuming it.
# manual_part: a Part for the manual (gs:// on enterprise, files/... on the Developer API)
cache = client.caches.create(
    model=MODEL,
    config=types.CreateCachedContentConfig(
        contents=[types.Content(role="user", parts=[manual_part])],
        system_instruction="Answer only from the attached manual. Cite section numbers.",
        display_name="support-manual-v14",
        ttl="3600s",
    ),
)
answer = client.models.generate_content(
    model=MODEL, contents=question,
    config=types.GenerateContentConfig(cached_content=cache.name),
)
assert answer.usage_metadata.cached_content_token_count, "cache not applied"

Worked example: prototype to production

A team builds a support assistant that answers from a 400-page product manual. They prototype in AI Studio with a key, upload the manual with the Files API, and get good answers in a day. Then production arrives with four new requirements: customer data must stay in the EU, the security team wants no long-lived keys, finance wants per-tenant cost, and peak traffic is 30 requests per second.

The migration follows the table above. The manual moves from the Files API to a Cloud Storage bucket in an EU region; the client switches to enterprise=True with an explicit EU location rather than global; the service runs as a dedicated service account with only the permissions needed to call models and read that bucket, granted through IAM as described in GCP IAM; and the project goes inside a VPC Service Controls perimeter (VPC Service Controls) so the model endpoint cannot be reached with stolen credentials from outside it. The manual becomes a context cache refreshed on each release, and every response's usage_metadata goes to a log sink with the tenant ID, where a scheduled query turns tokens into money.

The step that bit them was quota. AI Studio quotas and Cloud project quotas are separate systems, and the regional quota in the new project was lower than prototype traffic. They found out from 429s in load testing, raised the request through the Cloud console, and added a client-side token bucket so a burst queues briefly instead of tripping retries. Teams with hard latency targets at steady high volume should also evaluate provisioned throughput on the enterprise path.

Failure modes and trade-offs

SymptomLikely causeFirst response
429 RESOURCE_EXHAUSTED under loadPer-model, per-region or per-key quotaBackoff with jitter, client-side rate limit, quota increase, spread across regions only if residency allows
ValueError at client constructionvertexai and enterprise both set and conflicting, or no project or key foundSet one flag; set GOOGLE_CLOUD_PROJECT or pass project=
403 on the enterprise pathService account lacks the model-calling role, or a VPC-SC perimeter blocks the callerCheck IAM policy and perimeter audit logs
Empty text or invalid JSONSafety block, truncation at max_output_tokens, or thinking consumed the output budgetLog finish_reason and token counts; raise the limit or lower thinking
Works in dev, 404 in prodModel ID not available in the chosen region or on that backendList models in the prod project; pin IDs per environment
Cost spike without traffic growthA prompt change defeated the cache, or thinking tokens grewAlert on cached and thought tokens per request, not only totals

Two trade-offs deserve a decision rather than a default. The global endpoint improves availability and capacity but gives up regional control. And the Developer API is simpler, but keys are bearer secrets that leak through notebooks and logs; if you stay on it in production, store keys in Secret Manager and rotate them.

What to do next

  1. Centralise client construction in one function driven by environment variables, with an explicit location and a pinned model ID.
  2. Replace vertexai=True with enterprise=True and GOOGLE_GENAI_USE_VERTEXAI with GOOGLE_GENAI_USE_ENTERPRISE; never set both.
  3. Configure HttpRetryOptions with 429 and 5xx, a millisecond timeout, and a client-side rate limiter for bursts.
  4. Move extraction calls to response_json_schema with Pydantic validation and finish_reason checks.
  5. Log all five usage_metadata counters per request with tenant and feature tags.
  6. Identify shared long prefixes and test context caching; move anything that can wait to batch.
  7. Run integration tests against both backends in CI, and check quotas in the production project before launch.
Key takeaway: The Gemini Developer API and the enterprise path (Vertex AI, now Gemini Enterprise Agent Platform) share request types but differ in host, API version, auth, location and where large inputs live. Build one client factory with an explicit location and pinned model, prefer enterprise=True over the legacy vertexai flag, retry 429 and 5xx with backoff, constrain output with response_json_schema and validate it, log every usage counter, and use caching and batch where reuse and latency allow.