Calling Gemini on Google Cloud is easy. Keeping it fast and affordable under real traffic is harder. The first request works, the demo works, and then a launch-day spike returns a wall of 429 RESOURCE_EXHAUSTED errors, or the finance team asks why a nightly summarisation job costs as much as the customer-facing assistant. Both problems have the same root: the platform offers several ways to buy model capacity, and most teams send every request down the default one.
This article treats Gemini on Google Cloud as a capacity system. It explains where your tokens actually run, what a 429 means on a shared pool, how the consumption options differ, how to pick one per request with headers, how to size reserved capacity from a traffic profile, and how to tell from the response which capacity served the call. One naming note first: in April 2026 Google folded Vertex AI into the Gemini Enterprise Agent Platform. The documentation and console now use the new name, but the capacity mechanics below carried over, and so did the X-Vertex-AI-LLM-* request headers, which still carry the old name.
Where your tokens actually run
Gemini models are served on Google's accelerator fleet, which is mostly TPUs, behind a regional or global endpoint. You never see a device. What you buy is a share of a serving pool, and the cost of each request to that pool depends on the same things that matter on your own GPUs. Prompt tokens are processed in one compute-heavy prefill pass. Output tokens are generated one decode step at a time, and each step is limited by memory bandwidth. A request with a 30,000-token context and a 200-token answer stresses the pool very differently from a 500-token prompt with a 4,000-token answer, even when the token totals look similar.
That is why the platform prices input and output tokens separately, and why reserved capacity is measured in throughput units rather than requests. Treat the hosted model as a remote accelerator pool with a queue in front of it. Your job is to decide which queue each request joins, and to keep enough evidence to change that decision later.
Five ways to buy capacity
There are five ways to consume Gemini capacity. Google's documentation describes Standard, Priority and Flex PayGo plus batch inference as token-billed at different rates, and Provisioned Throughput as reserved generative AI scale units (GSUs) billed for the time you hold them, whatever your traffic. Check the current price sheet for the multipliers. They change, so this article quotes none.
| Lane | How you get it | What it promises | Use it for |
|---|---|---|---|
| Standard PayGo | Default; no header needed | Shared pool under dynamic shared quota; 429 under contention | Most interactive traffic that can absorb an occasional retry |
| Priority PayGo | X-Vertex-AI-LLM-Shared-Request-Type: priority | Shared pool with precedence, for spiky traffic that cannot risk 429s | Revenue-critical or user-blocking calls without a reservation |
| Flex PayGo | X-Vertex-AI-LLM-Shared-Request-Type: flex | Lower price, longer latency, more throttling; request timeout up to 30 minutes | Background enrichment, evaluations, back-office agents |
| Batch inference | Submit a batch job over stored inputs | Asynchronous completion of a whole file of requests | Nightly summarisation, backfills, large evaluation sweeps |
| Provisioned Throughput | Purchase GSUs for a model and region | Dedicated capacity up to the purchased throughput | A predictable baseline that must not degrade |
A second header controls how a request interacts with a reservation. X-Vertex-AI-LLM-Request-Type: dedicated uses only Provisioned Throughput and returns 429 once the reservation is used up. X-Vertex-AI-LLM-Request-Type: shared skips the reservation and goes straight to PayGo. With neither header, a project that holds Provisioned Throughput uses it first and spills over to Standard PayGo when it runs out. Flex has two documented forms. The flex header alone uses any available reservation first. Adding Request-Type: shared makes the request use only Flex.
Dynamic shared quota and what a 429 means
On the shared lanes, Google uses dynamic shared quota instead of a fixed per-project tokens-per-minute limit. The documentation is explicit that a 429 here does not mean you hit a fixed quota. It means the shared resource you asked for is under temporary heavy contention. Your organisation's baseline access is adjusted from its spend on eligible platform services over a rolling 30-day window, and higher spend moves you to higher tiers automatically.
This has three consequences. First, raising a quota ticket will not fix 429s, because there is no quota to raise. Second, 429s cluster in time and across tenants. A retry storm from your own fleet makes the contention you are reacting to worse, so retries must back off and must be capped. Third, if a class of traffic cannot tolerate that behaviour, the fix is to change its lane: Priority, a reservation, or a different region or the global endpoint. Retrying harder is not a fix.
A per-request lane router
The cleanest way to apply lanes is to name request classes in your application and map each class to headers in one place. The Google Gen AI SDK for Python (google-genai) accepts default headers on the client, so one client per lane keeps the mapping explicit. The model ID is configuration. Gemini versions on the platform change every few months, so read the current list in Model Garden instead of hard-coding a name from a blog post.
import os
from google import genai
from google.genai import types
PROJECT = os.environ["GOOGLE_CLOUD_PROJECT"]
MODEL = os.environ["GEMINI_MODEL"] # e.g. the current Flash ID from Model Garden
LANE_HEADERS = {
"standard": {},
"priority": {"X-Vertex-AI-LLM-Shared-Request-Type": "priority"},
"flex_only": {"X-Vertex-AI-LLM-Request-Type": "shared",
"X-Vertex-AI-LLM-Shared-Request-Type": "flex"},
"pt_only": {"X-Vertex-AI-LLM-Request-Type": "dedicated"},
"paygo_only": {"X-Vertex-AI-LLM-Request-Type": "shared"},
}
# request class -> lane; the only place product code ever touches capacity policy
CLASS_LANE = {
"checkout_assistant": "priority",
"support_chat": "standard", # uses PT first if the project holds it
"ticket_enrichment": "flex_only",
}
CLIENTS = {
lane: genai.Client(vertexai=True, project=PROJECT, location="global",
http_options=types.HttpOptions(headers=h))
for lane, h in LANE_HEADERS.items()
}
def generate(request_class, contents, **config):
lane = CLASS_LANE[request_class]
resp = CLIENTS[lane].models.generate_content(
model=MODEL, contents=contents,
config=types.GenerateContentConfig(**config))
usage = resp.usage_metadata
# traffic type reports which capacity served the call; read it defensively
served_by = getattr(usage, "traffic_type", None)
return resp, {"lane": lane, "served_by": str(served_by),
"in": usage.prompt_token_count, "out": usage.candidates_token_count,
"cached": getattr(usage, "cached_content_token_count", None)}The returned record is the important part. The response's usage metadata includes a traffic type. Observed values include ON_DEMAND, ON_DEMAND_PRIORITY, FLEX and BATCH, plus a separate value for Provisioned Throughput. Log it on every call. It is the only direct evidence of how much traffic your reservation absorbed and how much spilled over, and it is what billing will reflect.
Retries that do not make contention worse
Retries belong in one wrapper with three rules. Retry only errors that can succeed on a second attempt: 429, 500, 503 and 504. Use exponential backoff with full jitter so a fleet of clients does not retry in lockstep. Hold a retry budget per process so a contention event cannot double your request rate.
import random, time
from google.genai import errors
RETRYABLE = {429, 500, 503, 504}
class RetryBudget:
"""Allow retries up to `ratio` of recent first attempts (token bucket)."""
def __init__(self, ratio=0.1, burst=20):
self.tokens, self.ratio, self.burst = burst, ratio, burst
def on_attempt(self):
self.tokens = min(self.burst, self.tokens + self.ratio)
def take(self):
if self.tokens >= 1:
self.tokens -= 1
return True
return False
BUDGET = RetryBudget()
def call_with_retry(fn, *, max_tries=5, base=0.5, cap=20.0, deadline_s=60.0):
start = time.monotonic()
BUDGET.on_attempt()
for attempt in range(max_tries):
try:
return fn()
except errors.APIError as e:
last = attempt == max_tries - 1
if e.code not in RETRYABLE or last or not BUDGET.take():
raise
sleep = random.uniform(0, min(cap, base * 2 ** attempt))
if time.monotonic() - start + sleep > deadline_s:
raise
time.sleep(sleep)The deadline matters as much as the backoff. An interactive call that has already waited 20 seconds is better failed fast with a clear message, or degraded to a cheaper path, than retried until the user leaves. Flex calls get a long deadline and a long SDK timeout, because slow is the deal you accepted.
Sizing Provisioned Throughput: a worked example
Provisioned Throughput is the only lane that guarantees capacity, and it is billed whether you use it or not, so size it from data. The documentation publishes, per model, a throughput per GSU and burndown rates. Burndown rates convert each kind of token (input, output, cached input and so on) into the unit the reservation is measured in. Output tokens burn faster than input tokens, because decode is the expensive phase. The arithmetic below uses placeholder rates. Replace them with the published figures for your model before you buy anything.
import math
def gsus_needed(req_per_s, in_tok, out_tok, out_burn, per_gsu, headroom=0.8):
"""Burndown units per second divided by usable per-GSU throughput."""
burn = req_per_s * (in_tok + out_tok * out_burn)
return math.ceil(burn / (per_gsu * headroom))
# Placeholder rates for illustration only; read the real ones from the PT docs.
OUT_BURN, PER_GSU = 4.0, 3000.0
print(gsus_needed(25, 3000, 400, OUT_BURN, PER_GSU)) # baseline -> 48
print(gsus_needed(40, 3000, 400, OUT_BURN, PER_GSU)) # peak -> 77Worked example. A support assistant runs at 25 requests per second through most of the day and peaks at 40 for about two hours. The average request carries 3,000 input tokens and 400 output tokens. With the placeholder rates, the baseline needs 48 GSUs at 80 percent target utilisation and the peak needs 77. Buying 77 means paying for about 29 idle GSUs for 22 hours a day. Buying 48 and letting the default spillover send the peak to Standard PayGo costs nothing while traffic is quiet, but it exposes the peak to dynamic shared quota. The usual answer is in between. Reserve the baseline, and send the few request classes that must never see a 429 to Priority PayGo during the peak.
Before buying, shrink the demand. If 2,400 of those 3,000 input tokens are a shared system prompt and policy document, context caching cuts both the bill and the prefill work. Gemini 2.5 and newer models apply implicit caching to repeated prefixes by default. Explicit caches give you a handle and a TTL for content you know will be reused. Check the per-model minimum cacheable size and how cached tokens burn down against a reservation. Then move everything that does not need an answer within seconds (enrichment, re-summarisation, evaluation) to Flex or batch inference. It does not belong in the reserved baseline.
Failure modes
- Retry storms. Unbounded retries on 429 multiply load during contention. Symptoms are request rate rising while successful completions fall. The fix is jittered backoff, a retry budget and a deadline.
- Silent spillover. A reservation sized for last quarter overflows to Standard PayGo every afternoon. Latency and cost drift and nothing alerts. Alert on the share of traffic not served by Provisioned Throughput for classes you meant to reserve.
- Dedicated-only outages.
Request-Type: dedicatedturns an undersized reservation into hard 429s. Use it only where predictable cost matters more than availability, and only with a fallback path. - Flex on the hot path. A background lane chosen for cost leaks into an interactive feature through a shared helper. Classes, not call sites, should choose lanes.
- Region pinning without need. A regional endpoint chosen out of habit limits you to one region's pool. Use the global endpoint unless data-residency rules require a region. If they do, record that constraint next to the lane map.
- Model-ID drift. Preview and versioned IDs are retired on a schedule. Pin explicit versions in configuration, watch the release notes, and run your evaluation set before changing the pin.
- Token blind spots. Long agent transcripts grow input tokens per turn until each call costs ten times what it did in testing. Track input tokens per request class as a first-class metric.
Operating the capacity mix
Run capacity like any other shared dependency. Each request already produces a record with class, lane, served traffic type, input, output and cached tokens, latency and status. Aggregate those per class and per hour, and three views cover most decisions: 429 rate and p95 latency per lane, Provisioned Throughput utilisation against spillover, and tokens per request per class. Review them monthly, and before every launch. Move a class to Priority when its 429s show up in user-facing error budgets, and to Flex when nobody would notice a slower answer. Resize the reservation when baseline utilisation sits outside roughly 60 to 85 percent for weeks.
Keep governance next to capacity. IAM controls who can call which models. VPC Service Controls and customer-managed encryption keys apply to platform resources such as explicit caches and batch inputs. Model Garden also serves partner and open models through the same platform, with their own quotas and prices. A router that names lanes per class makes it cheap to try one of them for a single class later.
Related reading on this site: Gemini API vs Vertex, in depth, Model Garden deployment paths, reserved LLM capacity planning, SLO burn-rate alerting for LLM serving and the TPU generation behind much of the fleet.
What to do next
- List your request classes and give each a latency tolerance in seconds. That list is your lane map.
- Add the lane router and log class, lane, served traffic type and token counts on every call.
- Wrap every call in jittered backoff with a retry budget and a per-class deadline.
- Measure a week of traffic, then size Provisioned Throughput for the baseline using the published per-model rates, not the placeholders above.
- Move latency-tolerant classes to Flex or batch inference and confirm the cost drop in the traffic-type split.
- Cache shared prefixes and recheck input tokens per request.
- Alert on spillover share and on 429 rate per lane, and review both before each launch.