The stub this page replaces was titled after two names, Bricktree and Traefik AI. We could not confirm that Bricktree is a documented AI gateway product, so this article does not describe it. Traefik's AI gateway is real and documented: Traefik Hub adds an AI Gateway layer to its API gateway, a set of middlewares for LLM traffic that attach to ordinary Traefik routes. That is what this page covers in depth.
The GPU angle is the point. A gateway in front of self-hosted inference is not only a governance tool; it is a capacity control. Every request rejected for exceeding a token budget, answered from a semantic cache, or blocked by a guard before reaching the model is GPU time the inference pool does not spend. We build a chain step by step with configuration taken from Traefik's documentation, explain what each middleware costs and saves, size token limits from a load test in a worked example, and list the failure modes. For gateway concepts in general, such as virtual keys and budget reservation, read AI Gateway Overview first.
What Traefik's AI Gateway is
Traefik Proxy is an open-source reverse proxy and Kubernetes ingress controller. Traefik Hub builds API management and gateway features on top of it, and the AI Gateway is a part of Hub, so it requires a Hub installation rather than plain Traefik Proxy. In the documented Helm setup it is switched on with hub.aigateway.enabled=true, which turns on all AI features. The equivalent static flag is --hub.aigateway=true.
One setting matters immediately for LLM traffic: the maximum request body size, hub.aigateway.maxRequestBodySize, which defaults to 1 MiB. Long-context chat requests, conversations with many turns, and especially base64 images exceed that quickly. Size it from your longest legitimate request, not from a guess, and note that raising it raises the memory the gateway holds per in-flight request.
The documented middlewares include Chat Completion, Messages API and Responses API (format handling for the OpenAI and Anthropic APIs), Semantic Cache, Content Guard, LLM Guard and a parallel variant, and Token Rate Limit and Quota. Supported upstreams listed in the docs include OpenAI, Azure OpenAI, Anthropic, Gemini, Mistral, Amazon Bedrock, Ollama and local models served by KServe or vLLM. At the time of writing Token Rate Limit and Quota is marked early access; check the status of each piece before depending on it.
Middlewares, resources and order
Each middleware is a Middleware resource with API version traefik.io/v1alpha1 and its settings under spec.plugin.<name>. An IngressRoute references them in a list, and Traefik applies them in that order on the way in. Order is a design decision with GPU consequences:
- Format handling first. Traefik's own Messages API guide puts the format middleware ahead of Content Guard, LLM Guard, Semantic Cache and Token Rate Limit, in that order. We follow it.
- Guards before the cache. A prompt that should be blocked never gets a cached answer, and cached responses pass back through the guards' response rules on the way out.
- Cache before the limiter. A hit returns before the limiter, so it neither consumes the caller's tokens nor reaches a GPU. Put the limiter first instead if you want to bill hits as well.
Chat Completion: pinning model and output length
The Chat Completion middleware governs the request body. Its documented options include a default model, allowModelOverride, allowParamsOverride and default params such as temperature, topP and maxTokens. For a hosted provider, token holds a URN pointing at a Kubernetes Secret, such as urn:k8s:secret:ai-keys:openai-token. For an in-cluster vLLM pool you usually omit it.
apiVersion: v1
kind: Secret
metadata:
name: ai-keys
type: Opaque
stringData: # plaintext here; Kubernetes stores it encoded
openai-token: sk-proj-REPLACE_ME
---
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: chatcompletion-vllm
spec:
plugin:
chat-completion:
model: llama-3.1-8b-instruct # the name vLLM serves, not a provider model
allowModelOverride: false
allowParamsOverride: false
params:
maxTokens: 1024
temperature: 0.2Pinning maxTokens is the most direct GPU control in the whole chain. Decode time and KV-cache memory grow with output length, and vLLM schedules sequences against the KV-cache memory it has. A client that may ask for 32,000 output tokens can hold memory that would otherwise serve many short requests. Note that the Secret uses stringData: the example in Traefik's docs puts a raw key under data, which Kubernetes expects to be base64-encoded. The middleware also emits OpenTelemetry GenAI spans with prompt and completion token counts, which is the data you need for the worked example below.
Token rate limits as GPU capacity control
Request-per-second limits are the wrong unit for GPUs: one request may cost 50 tokens or 50,000. The Token Rate Limit and Quota middleware counts tokens. It comes in two variants: ai-rate-limit, a token bucket that tolerates bursts, and ai-quota, a sliding window for strict caps. Limits can apply to input, output or total tokens, or to cost in dollars, keyed by IP, header or JWT claim, with shared state in Redis so every gateway replica sees the same counters.
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: tokens-team-search
spec:
plugin:
ai-rate-limit:
store:
redis:
endpoints:
- redis.default.svc.cluster.local:6379
password: urn:k8s:secret:redis:password
totalTokenLimit:
limit: 140000
period: 1m
jsonQuery: ".usage.total_tokens"The true count is known only when the response reports usage, so a request is admitted before its cost is known. An optional estimate strategy checks the prompt up front; its simple form approximates input tokens as request body length divided by four, and rejects oversized prompts before they reach the model. When a limit is hit the default response is 429 Too Many Requests, customisable through onDenyResponse. Bucket capacity equals the limit, so a tenant with an hourly allowance can spend all of it in one minute. To protect GPUs, use short periods; to manage budgets, chain a second instance with a long period or use ai-quota.
Semantic cache: GPU time saved, and tenancy
Semantic Cache embeds the request, searches a vector database for a close previous request, and replays its response on a hit. Hits carry X-Cache-Status: Hit and an X-Cache-Distance header, which is the number to log while you tune maxDistance; lower is stricter. Only 200 responses are stored, streamed and non-streamed responses live in separate buckets, and inputs beyond the embedding model's limit bypass the cache.
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: semantic-cache
spec:
plugin:
semantic-cache:
vectorizer:
ollama:
baseUrl: http://ollama.default.svc.cluster.local:11434
model: nomic-embed-text
vectorDB:
redis:
endpoints:
- redis.default.svc.cluster.local:6379
collectionName: support_faq
maxDistance: 0.2 # strict; tune from X-Cache-Distance logs
ttl: 3600
allowBypass: trueThe GPU trade is direct: a hit costs one embedding call, which is small even when it runs on a GPU, instead of a full prefill and decode. The risk is correctness. Two questions can be semantically close and need different answers, such as the balance of account A and of account B. By default the cache is private, partitioned per caller identity, and needs an authentication middleware or Hub application ID to know who the caller is; without one the request bypasses the cache. Setting cacheControl.public: true shares entries across callers. Do that only for content that is identical for everyone, such as public documentation answers.
Content Guard and LLM Guard
Content Guard detects personal data with either a Presidio service or regular expressions. Rules select fields with JSON queries and either block, returning 403 by default, or mask, for example keeping the first and last two characters of a phone number. LLM Guard sends the request or response to an external guard model, for example Llama Guard served by Ollama, with a system prompt and block conditions such as Contains("unsafe").
Both have GPU and latency costs that are easy to miss. A guard model is another inference call per request, often on the same GPU fleet, and its cost scales with the guard model's size times the tokens it screens, so measure it in your load test rather than assuming it is small. More seriously, Traefik documents that when Content Guard has response rules it buffers the entire SSE stream and returns one aggregated response. Time to first token becomes total generation time, and users of a streaming chat UI see nothing until the end. If you need streaming, keep Content Guard to request rules and handle output checks elsewhere. See LLM guardrails on GPU infrastructure for placement options.
Routing to a vLLM pool
For a self-hosted pool, the IngressRoute's service is the Kubernetes Service in front of your vLLM pods, which expose an OpenAI-compatible API on port 8000 by default. Traefik balances across ready endpoints, but it does not know which replica has free KV-cache memory or a warm prefix cache, so for large pools consider a model-aware router behind the gateway. The gateway's job is admission; placement is the router's. For how the pool itself is built, see vLLM on GPUs.
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: llm-internal
spec:
routes:
- kind: Rule
match: Host(`llm.internal.example.com`)
middlewares:
- name: chatcompletion-vllm
- name: content-guard
- name: semantic-cache
- name: tokens-team-search
services:
- name: vllm
port: 8000In production an authentication middleware goes first in this list. Without a caller identity the private cache bypasses every request and the limiter has no team to key on. The content-guard entry stands for a Content Guard resource built like the earlier examples.
Worked example: sizing limits from a load test
Suppose a load test shows each of four vLLM replicas sustains about 1,500 total tokens per second at your p95 latency target, roughly 6,000 tokens per second or 360,000 per minute for the pool. Keep 20 percent headroom for retries and spikes, leaving about 288,000 tokens per minute to allocate.
Three teams share the pool. Give search 140,000 per minute, support 100,000 and batch analytics 48,000, each as an ai-rate-limit keyed on a team header or JWT claim. The limits sum to the allocatable figure, so even if all three burst at once the pool stays within its tested envelope. Add an ai-quota instance per team with a daily cap for budget control.
Now add the cache. If logs show support traffic hitting 15 percent at a distance you have validated, support's effective demand on the GPUs drops by about 15,000 tokens per minute. Do not hand that saving out as quota until you have measured it for a few weeks, because hit rates move with product changes. Finally, set maxTokens per route: support answers rarely need more than 800 tokens, analytics summaries may need 2,000. Watch for 429 rates per team; a team pinned at its limit is a capacity conversation, not a gateway bug.
Failure modes
- Estimate undercounts. Body length divided by four is a rough English heuristic. Text in scripts such as Chinese or Japanese uses more tokens per byte, so the estimate runs low; base64 images run high. Treat it as a coarse filter, not a meter.
- Shared cache leaks.
cacheControl.public: trueon a route that answers per-customer questions can replay one customer's answer to another. - Streaming silently disabled. Content Guard response rules buffer SSE; users see a frozen UI and retry, doubling GPU load.
- Body limit rejects long contexts. The 1 MiB default fails long conversations before any middleware runs.
- Redis outage. Limiter and cache depend on Redis; decide and test whether the route fails open or closed.
- Clients ignore 429. Some streaming clients treat non-2xx responses badly; the docs note
onDenyResponsecan return a different status for them, but then dashboards must count denials separately.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Short-period token limits | Protects GPU latency under bursts | More 429s for legitimately bursty jobs |
| Strict cache distance | Few wrong replays | Lower hit rate, less GPU saved |
| Guard model on every request | Consistent safety screening | Extra inference per call and added latency |
| Response-side Content Guard | Output PII never leaves | No token streaming |
| Gateway-only routing to vLLM | Simple to run | No KV-cache or prefix-aware placement |
What to do next
- Load-test your inference pool and write down sustainable tokens per minute at your latency target before setting any limit.
- Enable the AI Gateway in a staging Hub install and raise the body size limit to fit your longest real request.
- Pin model and
maxTokensper route with Chat Completion. - Add token limits per team with short periods for capacity and quotas for budget, backed by Redis.
- Turn on Semantic Cache in private mode only, log the distance header and tune against labelled pairs.
- Decide whether output checks are worth losing streaming; review rate limiting for LLM APIs for limiter design beyond the gateway.