Kong AI Gateway is not a separate product binary. It is Kong Gateway, the Nginx and OpenResty based API proxy, plus a family of plugins whose names start with ai-. You attach them to a route the same way you attach key-auth or rate-limiting. The result sits between your applications and every model endpoint you run: self-hosted vLLM pools on your own GPUs, and hosted APIs such as OpenAI, Anthropic or Bedrock.
This page covers configuring and operating that stack: what ai-proxy does to a request, how ai-proxy-advanced balances across GPU pools, why token budgets protect GPU capacity better than request counts, where caching and guards fit, and what breaks. The case for a gateway at all is in AI Gateway Overview; a managed alternative is in Apigee AI Gateway, in depth. Field names were checked against Kong's plugin reference on 2026-10-04; check the reference for your version before copying one.
What Kong does to an LLM request
Start with one request. A client sends an OpenAI-style chat body to a Kong route such as /llm/chat. Kong runs the route's plugins in its fixed priority order, not the order you list them. Authentication runs first, so every later plugin knows the consumer. Guards and caches inspect the body, the rate limiter checks the token budget, and the proxy plugin rewrites the request into the provider's format, injects the credential and forwards it.
On the way back, the proxy plugin normalises the response to the format the client asked for and reads the token usage the provider reports. That usage is what separates an AI gateway from a reverse proxy: the gateway can rate-limit, attribute cost and log in tokens, not requests.
Kong introduced the core AI plugins as open source in Gateway 3.6. Three plugins used here, ai-proxy-advanced, ai-semantic-cache and ai-rate-limiting-advanced, are marked with the ai_gateway_enterprise tier in Kong's reference. Packaging has changed more than once, so check each plugin's tier badge for your version.
One route, one model: ai-proxy
The smallest useful setup is one route with ai-proxy. Its main fields are config.route_type (for example llm/v1/chat, llm/v1/completions or llm/v1/embeddings), config.auth (header_name and header_value, or the AWS, Azure and GCP credential fields), and config.model (provider, name and options such as max_tokens, temperature and upstream_url). Here it is as decK declarative config:
_format_version: "3.0"
services:
- name: llm-chat
url: http://localhost:32000 # placeholder; ai-proxy sets the real upstream
routes:
- name: chat
paths: ["/llm/chat"]
plugins:
- name: key-auth
- name: ai-proxy
config:
route_type: llm/v1/chat
auth:
header_name: Authorization
header_value: ${{ env "DECK_OPENAI_AUTH" }} # "Bearer sk-..." injected at sync time
allow_override: false
model:
provider: openai
name: gpt-4o-mini
options:
max_tokens: 1024
response_streaming: allow
logging:
log_statistics: true
log_payloads: falseThree choices in that block matter more than they look. allow_override: false stops a client from sending its own Authorization header to use a different key; the provider credential lives only in the gateway. max_tokens caps how much a single call can generate, which caps how long it can hold a GPU slot. response_streaming takes allow, always or deny; streamed responses still need their token usage counted, so test that your provider and Kong version report usage on streams before you rely on streamed traffic for billing.
Clients send a normal chat body to /llm/chat with their Kong key. They never hold the provider key, and you can change model.name or the provider without redeploying a single client.
Pointing Kong at your own GPUs
For your own GPUs, the usual backend is vLLM's OpenAI-compatible server (see vLLM on GPU, in depth for what that engine does with memory and scheduling). Kong's provider list now includes vLLM and Ollama by name. The pattern that works across versions is to treat the server as an OpenAI-format upstream and point model.options.upstream_url at it:
model:
provider: openai # vLLM speaks the OpenAI wire format
name: meta-llama/Llama-3.1-8B-Instruct
options:
upstream_url: http://vllm-a.inference.svc:8000/v1/chat/completions
max_tokens: 1024Two GPU-side facts drive the gateway settings. A vLLM replica has a fixed KV-cache budget; past it, the engine queues or preempts sequences and latency climbs long before errors appear. And a request's cost grows with its tokens: a 200-token chat and a 30,000-token document summary are both one request. The gateway sees every caller, so it is where you shape traffic to what the GPUs can hold.
Balancing and failover with ai-proxy-advanced
ai-proxy-advanced replaces the single model with a list of targets, each with its own route_type, auth, model and weight (1 to 65535, default 100), plus a balancer block. The balancer algorithm is one of round-robin, consistent-hashing, least-connections, lowest-latency, lowest-usage, priority or semantic.
- name: ai-proxy-advanced
config:
balancer:
algorithm: lowest-latency
latency_strategy: tpot # or e2e
retries: 3
failover_criteria: [error, http_429, http_503]
connect_timeout: 2000
read_timeout: 120000
targets:
- route_type: llm/v1/chat
weight: 100
model:
provider: openai
name: meta-llama/Llama-3.1-8B-Instruct
options:
upstream_url: http://vllm-a.inference.svc:8000/v1/chat/completions
- route_type: llm/v1/chat
weight: 100
model:
provider: openai
name: meta-llama/Llama-3.1-8B-Instruct
options:
upstream_url: http://vllm-b.inference.svc:8000/v1/chat/completionsHow to pick the algorithm:
| Algorithm | What it does | Use it when |
|---|---|---|
round-robin | Rotates by weight | Identical pools, steady load |
lowest-latency | EWMA of observed latency; latency_strategy tpot (time per output token) or e2e | Pools on different GPU types or regions |
lowest-usage | Fewest tokens or lowest cost, by tokens_count_strategy | Spreading spend or token load |
consistent-hashing | Same key, same target (hash_on_header) | Prefix-cache locality per tenant or session |
priority | Tiered failover across target groups | Own GPUs first, hosted API as overflow |
semantic | Embeds the prompt and matches target descriptions in a vector DB | Routing code questions to a code model |
consistent-hashing deserves a note for GPU pools. vLLM's prefix caching only helps if requests that share a long system prompt or conversation land on the same replica. Hashing on a session or tenant header keeps them together. Round-robin spreads them out and rebuilds the same KV blocks on every replica. LLM Routing Strategies, in depth covers cache-aware and tiered routing in more detail.
Retries need care. failover_criteria includes non_idempotent as an opt-in for a reason. A chat completion that timed out after 100 seconds may have finished on the GPU, so a retry pays for it twice. Keep read_timeout above your longest legitimate generation, and only fail over on errors that mean the request never ran, such as connection errors, 429 and 503.
Token budgets that protect GPU capacity
A request-per-second limit treats a 50-token call and a 30,000-token call as equal. ai-rate-limiting-advanced counts tokens instead. Each entry in config.llm_providers has a name (openai, anthropic, bedrock and others, plus requestPrompt and customCost), and lists of limit and window_size values in seconds, so one provider can carry a per-minute and a per-day budget at once.
- name: ai-rate-limiting-advanced
config:
identifier: consumer # ip, credential, consumer, service, header, path, consumer-group
tokens_count_strategy: total_tokens # prompt_tokens, completion_tokens, cost
window_type: sliding
strategy: redis # local, redis or cluster
sync_rate: 0.5 # seconds between pushes to the shared store
redis:
host: redis.gateway.svc
port: 6379
llm_providers:
- name: openai # the self-hosted vLLM targets use the openai format
limit: [200000, 5000000]
window_size: [60, 86400]Know the counting model before you set numbers. Completion tokens exist only after the model has produced them, so the limiter can only charge a request once its usage is known. Concurrent requests that start under the budget can all run, and the consumer can overshoot by up to one wave of in-flight requests. With strategy: redis and a non-zero sync_rate, each Kong node also works from a counter that may be up to sync_rate seconds stale. The budget is a soft ceiling. Pair it with max_tokens on the proxy so one request has a hard ceiling.
Prompt guards and semantic caching
ai-prompt-guard is a regex filter on prompt text. allow_patterns and deny_patterns each hold up to 10 patterns of up to 500 characters; match_all_roles extends matching beyond user messages. Use it for cheap, exact things such as strings that look like card numbers. It is not a defence against prompt injection, because a regex cannot recognise a paraphrase.
ai-semantic-cache embeds the incoming prompt with a configured embeddings model (embeddings.model.provider and name), searches a vector store (vectordb.strategy redis or pgvector, with dimensions, distance_metric cosine or euclidean, and threshold), and returns the stored answer on a hit. cache_ttl defaults to 300 seconds and message_countback to 1, so by default only the last message is vectorised. exact_caching adds an exact-match lookup first.
The threshold decides whether the cache helps or harms. Too loose, and "cancel my order" gets the cached answer to "cancel my subscription". Start with exact caching, then a strict threshold on narrow, read-only routes such as FAQ answers, and never cache answers that depend on user data. Response Caching for LLM APIs covers keys, determinism and invalidation.
Logging and metrics
With logging.log_statistics: true, the proxy adds model, provider and token usage to Kong's log record, and any log plugin (http-log, file-log) ships it. That gives per-consumer, per-model token counts without touching application code: the input for chargeback and capacity planning.
log_payloads logs full prompts and completions. Leave it off on routes carrying personal or customer data unless your log pipeline is approved for that data. model_name_header adds the selected model to the response headers, the quickest way to see which target served a request during a failover test.
Worked example: three teams, two GPU pools, one overflow provider
Suppose you run two vLLM pools of 8 GPUs each serving an 8B instruct model, and a hosted API as overflow. Your own load test, run the way LLM Cost Analysis describes, shows each pool sustains about 1,200,000 total tokens per minute at your latency target. That is an assumed, measured figure, not a property of the hardware. Two pools give 2,400,000 tokens per minute.
Three teams share it. Reserve 20% headroom for spikes and retries, leaving 1,920,000 tokens per minute: 960,000 for the customer chatbot and 480,000 each for two internal tools. Each becomes an ai-rate-limiting-advanced instance per consumer group, with a daily cap of about 8 busy hours at the per-minute rate.
| Consumer group | Per minute (tokens) | Per day (tokens) | Over budget |
|---|---|---|---|
| team-a-chatbot | 960,000 | 460,800,000 | 429 with a retry-after hint |
| team-b-tools | 480,000 | 230,400,000 | 429 |
| team-c-batch | 480,000 | 230,400,000 | 429; move work to an offline batch path |
Routing uses ai-proxy-advanced with priority: both vLLM pools in the first tier, balanced by lowest-latency on tpot, and the hosted provider in the second tier. Failover triggers on error, http_429 and http_503. When pool A is drained for a driver upgrade, traffic shifts to pool B; only when B also fails does it reach the paid hosted tier. Alert on the hosted tier's token rate: it signals that GPU capacity is short.
Failure modes
- Shared counter store down. With
strategy: redis, stop Redis in staging to learn whether the limiter fails open (traffic unmetered) or closed (AI routes error), and plan for that. - Retry storms. High
retriesplushttp_429failover means a throttled hosted provider multiplies traffic onto your GPU pools. Cap retries at 2 or 3 and watch the per-target request rate. - Format drift. A provider adds a field or a new response type and translation misses it. Pin Kong versions, keep one contract test per route, and run it on every upgrade.
- Streaming without usage. If a stream ends without a usage block, the request may be charged zero tokens. Compare gateway token totals with provider invoices weekly.
- Cache poisoning by similarity. A loose semantic threshold serves a confident wrong answer. Log cache status and sample hits for review.
- Secrets in config. Provider keys written into decK files end up in Git. Inject them at sync time with decK environment substitution, or use Kong's vault references.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Kong in the path vs direct calls | One place for keys, limits, logs and failover | An extra network hop and a component that must be highly available |
ai-proxy vs ai-proxy-advanced | Simplicity and the open-source tier | No balancing or failover across targets |
| Token limits vs request limits | Fair sharing of GPU time | Soft enforcement, a shared counter store, and tuning |
| Regex guard vs semantic guard | Microseconds and predictable behaviour | Misses paraphrases; the semantic guard adds an embedding call per request |
| Semantic cache | Skips GPU work on repeated questions | Wrong-answer risk and a vector store to run |
What to do next
- List every application that calls a model today, with its provider key, model and rough tokens per day.
- Stand up one Kong route with
key-authandai-proxyin front of one vLLM pool, withlog_statisticson andallow_overrideoff. - Move one internal client to the route and confirm token counts in the logs match vLLM's own usage numbers.
- Load-test the pool to find the tokens per minute it sustains at your latency target, then write per-consumer budgets with 20% headroom.
- Add
ai-rate-limiting-advancedwith a Redis store, test what happens when Redis stops, and document whether it fails open or closed. - If you have more than one pool, switch to
ai-proxy-advanced, start withlowest-latencyorconsistent-hashing, and run a drain test on one pool. - Add a prompt guard and an exact-match cache only on routes where you can say what they protect against, and review a sample of cache hits each week.