Every token a large language model reads or writes is paid for in accelerator time. Prompt tokens cost prefill compute; output tokens cost decode steps, each of which holds a slot in a GPU's batch and a slice of its KV cache. If you run models on your own GPUs, or rent them by the token from Vertex AI or another provider, the scarce resource is the same: tokens per second of accelerator capacity. A request counter cannot protect that resource, because one request can carry 20 tokens or 200,000.
Google's Apigee API management platform now ships a set of policies aimed at exactly this problem: token quotas, token spike limits, semantic caching and prompt sanitisation, which together turn an Apigee proxy into an AI gateway. This article explains what each policy does, where it sits in the flow, how the pieces combine into one proxy, and where the sharp edges are. It is deliberately Apigee-specific; the vendor-neutral design of an AI gateway is covered in AI Gateway Overview. Element names and behaviours below were checked against the Apigee policy reference in October 2026. These are newer policies, so re-check the reference before you copy a configuration.
How an Apigee proxy processes a model call
An Apigee API proxy has a ProxyEndpoint, which faces the client, and a TargetEndpoint, which calls the backend. Each endpoint runs flows (PreFlow, conditional flows and PostFlow), and each flow has a request side and a response side. A flow is an ordered list of steps, and each step runs a policy, a declarative XML element that inspects or changes the message. Policies share state through flow variables such as request.content.
The AI policies are ordinary policies in this model. Most take a message template, a string such as {jsonPath('$.contents[-1].parts[-1].text',request.content,true)}, which tells the policy where in the payload to find the prompt or the token usage. That is the main integration effort: the defaults assume the Gemini generateContent shape (contents[].parts[].text for prompts, usageMetadata for counts). If your backend speaks the OpenAI chat-completions shape or a self-hosted vLLM or TGI server, you must rewrite every template to the paths your payload actually uses.
Apigee classes these as Extensible policies, and the docs warn that using them may have cost or utilisation implications depending on your licence. The token policies (LLMTokenQuota and PromptTokenLimit) are documented as applying to Apigee but not to Apigee hybrid. The semantic cache and Model Armor policies are documented for both. Check this before you design around a policy that your deployment model does not have.
The policy chain, end to end
Order matters: every step before the model call is far cheaper than the GPU work it guards. Identity is a lookup, the limits are counters, Model Armor and the cache lookup are single network calls. They exist to avoid calling the model for requests that would be rejected or answered from cache anyway.
Model Armor goes before the cache on purpose. If the cache ran first, a malicious prompt similar enough to a benign one could be served a cached answer without ever being inspected. Sanitising first means every prompt is inspected, whether it ends in a hit or a miss.
LLMTokenQuota: budgets counted from real usage
LLMTokenQuota is the cost-control policy. It keeps counters of tokens consumed over an interval (minute, hour, day, week or month) and rejects token-consuming requests once the counter reaches its limit, until the interval resets. It has three types: calendar (fixed windows from an explicit <StartTime>), flexi (the window starts at the first request) and rollingwindow (a sliding window).
The design has to work around a basic fact: you only know how many tokens a call used after the model answers. So the policy comes in two roles. An enforcement instance (<EnforceOnly>true</EnforceOnly>) runs in the request flow and rejects when the counter is already at the limit. A counting instance (<CountOnly>true</CountOnly>) runs in the response flow and adds the tokens it extracts with <LLMTokenUsageSource>. The two share a counter through <SharedName>. Exactly one of the two flags must be true in each instance, or Apigee raises InvalidConfiguration.
<!-- Request flow: reject if this app has already spent its window -->
<LLMTokenQuota name="LTQ-Enforce" type="rollingwindow">
<SharedName>gpu-pool-tokens</SharedName>
<EnforceOnly>true</EnforceOnly>
<Identifier ref="verifyapikey.VA-key.client_id"/>
<Allow count="200000"/>
<Interval>1</Interval>
<TimeUnit>hour</TimeUnit>
<Distributed>true</Distributed>
</LLMTokenQuota>
<!-- Response flow: add what the model actually used -->
<LLMTokenQuota name="LTQ-Count" type="rollingwindow">
<SharedName>gpu-pool-tokens</SharedName>
<CountOnly>true</CountOnly>
<Identifier ref="verifyapikey.VA-key.client_id"/>
<Allow count="200000"/>
<Interval>1</Interval>
<TimeUnit>hour</TimeUnit>
<Distributed>true</Distributed>
<LLMTokenUsageSource>{jsonPath('$.usageMetadata.totalTokenCount',response.content,true)}</LLMTokenUsageSource>
<LLMModelSource>{jsonPath('$.modelVersion',response.content,true)}</LLMModelSource>
</LLMTokenQuota>Two choices in that snippet matter. First, the reference examples count usageMetadata.candidatesTokenCount, which is output tokens only. That is defensible if you only want to meter decode, but prefill on a long prompt is real GPU work, so the sketch counts totalTokenCount. Choose the field deliberately and write the choice down, because finance will eventually ask what a token means. Second, with <Identifier> pointing at the client ID that a VerifyAPIKey policy (here named VA-key) exposes, the counter is per app rather than shared by every caller of the proxy. Limits can also come from the API product rather than the XML: after VerifyAPIKey, <Allow countRef=...>, <Interval ref=...> and <TimeUnit ref=...> can point at the product's quota variables, so the platform team sells tiers without redeploying proxies.
Streaming. For server-sent events the counter must run in an <EventFlow content-type="text/event-stream"> response step, with the usage path pointed at response.event.current.data. The policy only counts when an event carries usage metadata and is a no-op on every other event. If your backend never emits usage in the stream (some OpenAI-compatible servers only do so when the client asks), streamed calls are silently free. Test that before launch.
PromptTokenLimit: stopping token bursts
A daily quota does not stop one client from sending its whole day's budget in thirty seconds and filling every GPU's batch. PromptTokenLimit is the burst guard: SpikeArrest, but counting prompt tokens instead of requests. <Rate> is written as INTps or INTpm. <UserPromptSource> says where the prompt is, and <Identifier> groups callers. Leave it empty and one limit applies to the whole proxy.
<PromptTokenLimit name="PTL-PerApp">
<UserPromptSource>{jsonPath('$.contents',request.content,true)}</UserPromptSource>
<Identifier ref="verifyapikey.VA-key.client_id"/>
<Rate>60000pm</Rate>
<UseEffectiveCount>true</UseEffectiveCount>
</PromptTokenLimit>The rate above is illustrative: 60,000 prompt tokens per minute. With <UseEffectiveCount>true the count is synchronised across message processors in a region and enforced as a sliding window without smoothing, which is what you want when your gateway has more than one replica. With false each replica counts alone and the rate is smoothed, like SpikeArrest (12pm allows one token every five seconds). The docs' default template reads only the last part of the last message. Point <UserPromptSource> at the whole conversation, as here, or a client can slip a huge history past the limit in earlier turns. The reference does not spell out how the policy tokenises text. Treat its counts as an estimate made at the gateway, and keep the authoritative count from the model's usage field in LLMTokenQuota.
Semantic caching with Vector Search
SemanticCacheLookup sits in the request flow. It sends the extracted prompt to a Vertex AI text-embeddings model, queries a Vector Search index endpoint with findNeighbors, and treats the nearest neighbour as a hit when it passes <Threshold> under the configured <DistanceMeasureType>. On a hit the cached response goes back to the client, with the header Cached-Content: true in the documented setup, and the model is never called. SemanticCachePopulate sits in the response flow and writes the new prompt and response with upsertDatapoints and a <TTLInSeconds>. The tutorial creates the index with "indexUpdateMethod": "STREAM_UPDATE" so new entries become searchable without a batch rebuild.
<SemanticCacheLookup name="SCL-lookup">
<UserPromptSource>{jsonPath('$.contents[-1].parts[-1].text',request.content,true)}</UserPromptSource>
<Embeddings><VertexAI>
<URL>https://REGION-aiplatform.googleapis.com/v1/projects/PROJECT/locations/REGION/publishers/google/models/EMBED_MODEL:predict</URL>
</VertexAI></Embeddings>
<SimilaritySearch><VertexAI>
<URL>https://INDEX_DOMAIN/v1/projects/PROJECT/locations/REGION/indexEndpoints/ENDPOINT_ID:findNeighbors</URL>
<DeployedIndexID>DEPLOYED_INDEX_ID</DeployedIndexID>
<Threshold>0.95</Threshold>
<DistanceMeasureType>DOT_PRODUCT_DISTANCE</DistanceMeasureType>
</VertexAI></SimilaritySearch>
</SemanticCacheLookup>For debugging, the policy sets SemanticCacheLookup.SCL-lookup.cache_hit, SemanticCacheLookup.SCL-lookup.is_nearest_neighbor_hit and semanticcache.lookup.SCL-lookup.nearest_neighbor_measure (the casing differs between them in the reference; copy it exactly). Log the measure on every request for a week before you tune the threshold. Its sense depends on the distance type, and a threshold set from a few hand-picked examples will be wrong.
Two constraints decide whether you can use semantic caching at all. The docs state that it is not supported on proxies that use EventFlows for SSE streaming. Many chat front ends stream, so you may need a separate non-streaming proxy for the cacheable routes. And semantic similarity is not equivalence: "what is our refund window for EU customers" and "for US customers" can embed very close together. Cache only prompts whose answers do not depend on who is asking or on data that changes, and keep the cache per tenant, never shared across tenants.
Model Armor: sanitising prompts and responses
SanitizeUserPrompt sends the prompt to a Model Armor template (<ModelArmor><TemplateName>projects/P/locations/L/templates/T</TemplateName></ModelArmor>) which screens it for prompt injection, jailbreak attempts, sensitive data and harmful content, depending on how the template is configured. It also accepts a <FunctionResponseSource>, which matters for agents: tool output fed back to the model is a prompt-injection channel, and this is where to screen it. SanitizeModelResponse is the companion policy on the response side. Put both in the proxy, and decide per route whether a finding blocks the call or only logs it. Start in log-only mode, measure false positives, then turn on blocking.
Worked example: one app's hour of tokens
A platform team serves an internal assistant on a self-hosted model. Each app gets a quota of 200,000 total tokens per rolling hour and a spike limit of 60,000 prompt tokens per minute. Here is one app's hour, computed step by step:
| Event | Counter before | Tokens used | Counter after | Outcome |
|---|---|---|---|---|
| Normal traffic, 40 calls at ~4,900 tokens | 0 | 196,000 | 196,000 | all allowed |
| Call 41: 3,000 prompt + 2,500 output | 196,000 | 5,500 | 201,500 | allowed (196,000 < 200,000), counted after |
| Call 42 | 201,500 | - | 201,500 | 429 LLMTokenQuotaViolation |
| Batch job sends 90,000 prompt tokens in one minute | - | - | - | PromptTokenLimit throttles the excess |
Call 41 shows that the overshoot is built in. The enforcer saw 196,000, which is under 200,000, so it let the call through, and the counter then added 5,500. A rolling-window quota therefore overshoots by up to one response per concurrent in-flight request: if 30 calls pass the check at the same moment, all 30 are counted afterwards. Size limits with that slack included, or cap output length (generationConfig.maxOutputTokens for Gemini, max_tokens for OpenAI-style APIs) so one response cannot be enormous.
Now convert to GPU terms. Suppose load tests showed the serving pool sustains about 9,000 output tokens per second at the latency target. Then 20 apps each allowed 200,000 tokens per hour, about half of them output, need roughly 20 × 100,000 / 3,600 ≈ 556 output tokens per second averaged over the hour. That is about 6% of capacity on average. The averages are not the danger; bursts are, and that is why the per-minute PromptTokenLimit exists alongside the hourly quota.
Failure modes
- Usage path does not resolve. If the backend shape differs from the template, the counter raises
FailedToResolveTokenUsageCount(HTTP 500 to the client), or, withcontinueOnError, silently counts nothing. Alert on the fault rate per proxy revision. - Streams without usage. Counting is skipped for SSE events without usage metadata, so streaming traffic can bypass quota entirely. Test streamed and non-streamed calls separately.
- Model name unresolved.
FailedToResolveModelNameis a 400. If you route to several models, make sure<LLMModelSource>resolves on every route, including error responses. - Overshoot under concurrency. As in the worked example. Limit output length and accept the slack.
- Stale or wrong cache hits. A threshold set too loose returns another user's answer to a different question. Scope the cache per tenant, set TTLs to match how often the data changes, and never cache responses that contain personal data.
- Gateway dependency outage. Model Armor, embeddings and Vector Search are network calls. Decide per policy whether failure means fail closed (Model Armor on regulated routes) or fail open (cache lookup), and set
continueOnErrorto match.
Trade-offs
Apigee's advantage is that these controls sit beside the API products, developer portal, analytics and key management you may already run, so AI routes inherit the same onboarding and monetisation model as every other API. The cost is that you work in Apigee's template language, policies are licensed, and a few features are missing on hybrid. A purpose-built LLM gateway, or one of the open-source proxies, will usually understand more provider payload shapes out of the box. Some of them do model failover and routing that Apigee leaves to you; LLM routers cover that layer. Each policy before the model also adds latency, and the semantic cache lookup adds an embeddings call plus a vector query to every miss. Measure p50 and p99 with and without it, and drop it from routes with low hit rates.
What to do next
- Write down what a token means for billing: prompt, output or total. Then set
<LLMTokenUsageSource>to that field for each backend payload shape. - Deploy one proxy with VerifyAPIKey, an LLMTokenQuota enforcer/counter pair sharing a
<SharedName>, and PromptTokenLimit, all with per-app identifiers. - Test streamed calls and confirm the counter increments. If it does not, enable usage reporting in the stream or block streaming on that route.
- Cap output tokens at the gateway or in the request, and size quotas with the concurrency overshoot in mind.
- Add SanitizeUserPrompt and SanitizeModelResponse in log-only mode, review findings for two weeks, then block on the categories that matter.
- Pilot semantic caching on one non-streaming, non-personalised route. Log the nearest-neighbour measure, tune the threshold from real traffic, and keep it only if the hit rate pays for the extra latency.
- Tie gateway limits back to GPU capacity and alert on burn rate; LLM serving architecture and SLO burn-rate alerting cover the serving side.