Calling Gemini on Google Cloud is easy. Keeping it fast and affordable under real traffic is harder. The first request works, the demo works, and then a launch-day spike returns a wall of 429 RESOURCE_EXHAUSTED errors, or the finance team asks why a nightly summarisation job costs as much as the customer-facing assistant. Both problems have the same root: the platform offers several ways to buy model capacity, and most teams send every request down the default one.

This article treats Gemini on Google Cloud as a capacity system. It explains where your tokens actually run, what a 429 means on a shared pool, how the consumption options differ, how to pick one per request with headers, how to size reserved capacity from a traffic profile, and how to tell from the response which capacity served the call. One naming note first: in April 2026 Google folded Vertex AI into the Gemini Enterprise Agent Platform. The documentation and console now use the new name, but the capacity mechanics below carried over, and so did the X-Vertex-AI-LLM-* request headers, which still carry the old name.

Where your tokens actually run

Gemini models are served on Google's accelerator fleet, which is mostly TPUs, behind a regional or global endpoint. You never see a device. What you buy is a share of a serving pool, and the cost of each request to that pool depends on the same things that matter on your own GPUs. Prompt tokens are processed in one compute-heavy prefill pass. Output tokens are generated one decode step at a time, and each step is limited by memory bandwidth. A request with a 30,000-token context and a 200-token answer stresses the pool very differently from a 500-token prompt with a 4,000-token answer, even when the token totals look similar.

That is why the platform prices input and output tokens separately, and why reserved capacity is measured in throughput units rather than requests. Treat the hosted model as a remote accelerator pool with a queue in front of it. Your job is to decide which queue each request joins, and to keep enough evidence to change that decision later.

Five ways to buy capacity

There are five ways to consume Gemini capacity. Google's documentation describes Standard, Priority and Flex PayGo plus batch inference as token-billed at different rates, and Provisioned Throughput as reserved generative AI scale units (GSUs) billed for the time you hold them, whatever your traffic. Check the current price sheet for the multipliers. They change, so this article quotes none.

LaneHow you get itWhat it promisesUse it for
Standard PayGoDefault; no header neededShared pool under dynamic shared quota; 429 under contentionMost interactive traffic that can absorb an occasional retry
Priority PayGoX-Vertex-AI-LLM-Shared-Request-Type: priorityShared pool with precedence, for spiky traffic that cannot risk 429sRevenue-critical or user-blocking calls without a reservation
Flex PayGoX-Vertex-AI-LLM-Shared-Request-Type: flexLower price, longer latency, more throttling; request timeout up to 30 minutesBackground enrichment, evaluations, back-office agents
Batch inferenceSubmit a batch job over stored inputsAsynchronous completion of a whole file of requestsNightly summarisation, backfills, large evaluation sweeps
Provisioned ThroughputPurchase GSUs for a model and regionDedicated capacity up to the purchased throughputA predictable baseline that must not degrade

A second header controls how a request interacts with a reservation. X-Vertex-AI-LLM-Request-Type: dedicated uses only Provisioned Throughput and returns 429 once the reservation is used up. X-Vertex-AI-LLM-Request-Type: shared skips the reservation and goes straight to PayGo. With neither header, a project that holds Provisioned Throughput uses it first and spills over to Standard PayGo when it runs out. Flex has two documented forms. The flex header alone uses any available reservation first. Adding Request-Type: shared makes the request use only Flex.

Dynamic shared quota and what a 429 means

On the shared lanes, Google uses dynamic shared quota instead of a fixed per-project tokens-per-minute limit. The documentation is explicit that a 429 here does not mean you hit a fixed quota. It means the shared resource you asked for is under temporary heavy contention. Your organisation's baseline access is adjusted from its spend on eligible platform services over a rolling 30-day window, and higher spend moves you to higher tiers automatically.

This has three consequences. First, raising a quota ticket will not fix 429s, because there is no quota to raise. Second, 429s cluster in time and across tenants. A retry storm from your own fleet makes the contention you are reacting to worse, so retries must back off and must be capped. Third, if a class of traffic cannot tolerate that behaviour, the fix is to change its lane: Priority, a reservation, or a different region or the global endpoint. Retrying harder is not a fix.

A per-request lane router

The cleanest way to apply lanes is to name request classes in your application and map each class to headers in one place. The Google Gen AI SDK for Python (google-genai) accepts default headers on the client, so one client per lane keeps the mapping explicit. The model ID is configuration. Gemini versions on the platform change every few months, so read the current list in Model Garden instead of hard-coding a name from a blog post.

import os
from google import genai
from google.genai import types

PROJECT = os.environ["GOOGLE_CLOUD_PROJECT"]
MODEL = os.environ["GEMINI_MODEL"]          # e.g. the current Flash ID from Model Garden

LANE_HEADERS = {
    "standard":   {},
    "priority":   {"X-Vertex-AI-LLM-Shared-Request-Type": "priority"},
    "flex_only":  {"X-Vertex-AI-LLM-Request-Type": "shared",
                   "X-Vertex-AI-LLM-Shared-Request-Type": "flex"},
    "pt_only":    {"X-Vertex-AI-LLM-Request-Type": "dedicated"},
    "paygo_only": {"X-Vertex-AI-LLM-Request-Type": "shared"},
}

# request class -> lane; the only place product code ever touches capacity policy
CLASS_LANE = {
    "checkout_assistant": "priority",
    "support_chat":       "standard",     # uses PT first if the project holds it
    "ticket_enrichment":  "flex_only",
}

CLIENTS = {
    lane: genai.Client(vertexai=True, project=PROJECT, location="global",
                       http_options=types.HttpOptions(headers=h))
    for lane, h in LANE_HEADERS.items()
}

def generate(request_class, contents, **config):
    lane = CLASS_LANE[request_class]
    resp = CLIENTS[lane].models.generate_content(
        model=MODEL, contents=contents,
        config=types.GenerateContentConfig(**config))
    usage = resp.usage_metadata
    # traffic type reports which capacity served the call; read it defensively
    served_by = getattr(usage, "traffic_type", None)
    return resp, {"lane": lane, "served_by": str(served_by),
                  "in": usage.prompt_token_count, "out": usage.candidates_token_count,
                  "cached": getattr(usage, "cached_content_token_count", None)}

The returned record is the important part. The response's usage metadata includes a traffic type. Observed values include ON_DEMAND, ON_DEMAND_PRIORITY, FLEX and BATCH, plus a separate value for Provisioned Throughput. Log it on every call. It is the only direct evidence of how much traffic your reservation absorbed and how much spilled over, and it is what billing will reflect.

Retries that do not make contention worse

Retries belong in one wrapper with three rules. Retry only errors that can succeed on a second attempt: 429, 500, 503 and 504. Use exponential backoff with full jitter so a fleet of clients does not retry in lockstep. Hold a retry budget per process so a contention event cannot double your request rate.

import random, time
from google.genai import errors

RETRYABLE = {429, 500, 503, 504}

class RetryBudget:
    """Allow retries up to `ratio` of recent first attempts (token bucket)."""
    def __init__(self, ratio=0.1, burst=20):
        self.tokens, self.ratio, self.burst = burst, ratio, burst
    def on_attempt(self):
        self.tokens = min(self.burst, self.tokens + self.ratio)
    def take(self):
        if self.tokens >= 1:
            self.tokens -= 1
            return True
        return False

BUDGET = RetryBudget()

def call_with_retry(fn, *, max_tries=5, base=0.5, cap=20.0, deadline_s=60.0):
    start = time.monotonic()
    BUDGET.on_attempt()
    for attempt in range(max_tries):
        try:
            return fn()
        except errors.APIError as e:
            last = attempt == max_tries - 1
            if e.code not in RETRYABLE or last or not BUDGET.take():
                raise
            sleep = random.uniform(0, min(cap, base * 2 ** attempt))
            if time.monotonic() - start + sleep > deadline_s:
                raise
            time.sleep(sleep)

The deadline matters as much as the backoff. An interactive call that has already waited 20 seconds is better failed fast with a clear message, or degraded to a cheaper path, than retried until the user leaves. Flex calls get a long deadline and a long SDK timeout, because slow is the deal you accepted.

Sizing Provisioned Throughput: a worked example

One application, several capacity lanes: the router picks a lane per request classCallerschat, agents, jobsLane routerclass to headersProvisioned Throughputreserved GSUs, billed by timeStandard PayGoshared pool, dynamic shared quotaPriority PayGoshared pool, higher precedenceFlex PayGocheaper, slower, more throttlingspillover when PT is exhaustedBatch jobsfiles in, files outBatch inferenceasynchronous, latency-tolerantTelemetrytraffic type, 429 rate, latency per laneCapacity reviewresize GSUs, move classes between lanesfeedbackEvery request carries its lane in headers; every response reports which capacity actually served it.
Figure: request classes map to capacity lanes through headers; telemetry on the served traffic type feeds the capacity review that resizes the reservation.

Provisioned Throughput is the only lane that guarantees capacity, and it is billed whether you use it or not, so size it from data. The documentation publishes, per model, a throughput per GSU and burndown rates. Burndown rates convert each kind of token (input, output, cached input and so on) into the unit the reservation is measured in. Output tokens burn faster than input tokens, because decode is the expensive phase. The arithmetic below uses placeholder rates. Replace them with the published figures for your model before you buy anything.

import math

def gsus_needed(req_per_s, in_tok, out_tok, out_burn, per_gsu, headroom=0.8):
    """Burndown units per second divided by usable per-GSU throughput."""
    burn = req_per_s * (in_tok + out_tok * out_burn)
    return math.ceil(burn / (per_gsu * headroom))

# Placeholder rates for illustration only; read the real ones from the PT docs.
OUT_BURN, PER_GSU = 4.0, 3000.0
print(gsus_needed(25, 3000, 400, OUT_BURN, PER_GSU))   # baseline -> 48
print(gsus_needed(40, 3000, 400, OUT_BURN, PER_GSU))   # peak     -> 77

Worked example. A support assistant runs at 25 requests per second through most of the day and peaks at 40 for about two hours. The average request carries 3,000 input tokens and 400 output tokens. With the placeholder rates, the baseline needs 48 GSUs at 80 percent target utilisation and the peak needs 77. Buying 77 means paying for about 29 idle GSUs for 22 hours a day. Buying 48 and letting the default spillover send the peak to Standard PayGo costs nothing while traffic is quiet, but it exposes the peak to dynamic shared quota. The usual answer is in between. Reserve the baseline, and send the few request classes that must never see a 429 to Priority PayGo during the peak.

Before buying, shrink the demand. If 2,400 of those 3,000 input tokens are a shared system prompt and policy document, context caching cuts both the bill and the prefill work. Gemini 2.5 and newer models apply implicit caching to repeated prefixes by default. Explicit caches give you a handle and a TTL for content you know will be reused. Check the per-model minimum cacheable size and how cached tokens burn down against a reservation. Then move everything that does not need an answer within seconds (enrichment, re-summarisation, evaluation) to Flex or batch inference. It does not belong in the reserved baseline.

Failure modes

  • Retry storms. Unbounded retries on 429 multiply load during contention. Symptoms are request rate rising while successful completions fall. The fix is jittered backoff, a retry budget and a deadline.
  • Silent spillover. A reservation sized for last quarter overflows to Standard PayGo every afternoon. Latency and cost drift and nothing alerts. Alert on the share of traffic not served by Provisioned Throughput for classes you meant to reserve.
  • Dedicated-only outages. Request-Type: dedicated turns an undersized reservation into hard 429s. Use it only where predictable cost matters more than availability, and only with a fallback path.
  • Flex on the hot path. A background lane chosen for cost leaks into an interactive feature through a shared helper. Classes, not call sites, should choose lanes.
  • Region pinning without need. A regional endpoint chosen out of habit limits you to one region's pool. Use the global endpoint unless data-residency rules require a region. If they do, record that constraint next to the lane map.
  • Model-ID drift. Preview and versioned IDs are retired on a schedule. Pin explicit versions in configuration, watch the release notes, and run your evaluation set before changing the pin.
  • Token blind spots. Long agent transcripts grow input tokens per turn until each call costs ten times what it did in testing. Track input tokens per request class as a first-class metric.

Operating the capacity mix

Run capacity like any other shared dependency. Each request already produces a record with class, lane, served traffic type, input, output and cached tokens, latency and status. Aggregate those per class and per hour, and three views cover most decisions: 429 rate and p95 latency per lane, Provisioned Throughput utilisation against spillover, and tokens per request per class. Review them monthly, and before every launch. Move a class to Priority when its 429s show up in user-facing error budgets, and to Flex when nobody would notice a slower answer. Resize the reservation when baseline utilisation sits outside roughly 60 to 85 percent for weeks.

Keep governance next to capacity. IAM controls who can call which models. VPC Service Controls and customer-managed encryption keys apply to platform resources such as explicit caches and batch inputs. Model Garden also serves partner and open models through the same platform, with their own quotas and prices. A router that names lanes per class makes it cheap to try one of them for a single class later.

Related reading on this site: Gemini API vs Vertex, in depth, Model Garden deployment paths, reserved LLM capacity planning, SLO burn-rate alerting for LLM serving and the TPU generation behind much of the fleet.

What to do next

  1. List your request classes and give each a latency tolerance in seconds. That list is your lane map.
  2. Add the lane router and log class, lane, served traffic type and token counts on every call.
  3. Wrap every call in jittered backoff with a retry budget and a per-class deadline.
  4. Measure a week of traffic, then size Provisioned Throughput for the baseline using the published per-model rates, not the placeholders above.
  5. Move latency-tolerant classes to Flex or batch inference and confirm the cost drop in the traffic-type split.
  6. Cache shared prefixes and recheck input tokens per request.
  7. Alert on spillover share and on 429 rate per lane, and review both before each launch.
Key takeaway: Gemini on Google Cloud gives you several capacity lanes, not one endpoint. Map each request class to a lane with headers in one place, log which capacity actually served every call, retry 429s with jitter and a budget because they signal shared-pool contention rather than a quota, reserve only the measured baseline, and push latency-tolerant work to Flex or batch.