If you run inference on your own GPUs, you rarely have only one place to send a request. There is the vLLM pool in one region, a smaller pool somewhere else, and a hosted API you pay for when both are busy. Every application team ends up writing the same code: retry on 429, fall back on 503, split traffic by weight, give up after a timeout. Portkey's open-source AI Gateway moves that logic out of the application into a small proxy that speaks the OpenAI API and reads a declarative routing config on every request.
This article explains what that proxy actually does, at the level of its source code rather than its marketing page. You will see how the config tree is evaluated, the exact condition that makes a fallback move on, why load balancing is not health-aware, how retries interact with Retry-After headers and timeouts, and what the non-standard 446 and 246 status codes mean. The worked example puts two self-hosted GPU pools and a hosted provider behind one endpoint. Everything was checked against the gateway's main branch in October 2026. A 2.0 pre-release branch, which merges enterprise features into open source, was also public then, so re-check details against the version you deploy.
What the gateway is
The gateway is a TypeScript service (MIT licence) that runs on Node, Bun, Docker or Cloudflare Workers. Started with npx @portkey-ai/gateway, it listens on http://localhost:8787/v1 and serves a local log console at /public/. Clients call it with the OpenAI request shape; the gateway translates to the target provider's format and back. The provider and keys come either from headers such as x-portkey-provider and Authorization, or from a config passed in the x-portkey-config header.
Some features advertised in the README are marked with an asterisk, meaning they belong to the hosted or enterprise product rather than the open-source gateway: semantic caching, automatic provider optimisation and prompt template management. Virtual keys, saved config IDs and the analytics dashboards also live in the hosted control plane. If you self-host the open-source build, plan on passing configs inline and shipping logs to your own stack.
The config tree and inheritance
A config is a tree. A node either names a provider (a leaf) or carries a strategy and a list of targets, each of which is itself a node. Four strategy modes exist in the source: single, fallback, loadbalance and conditional. The gateway walks the tree recursively in a function called tryTargetsRecursively. Keys are written in snake_case in JSON and converted to camelCase internally.
Settings are inherited down the tree, and the rules differ by key. override_params is merged, so a child's values override the parent's key by key. retry and cache are replaced, not merged: a child that sets retry ignores the parent's retry block entirely, and a child that does not set it gets an exact copy. request_timeout, custom_host and forward_headers pass down unless the child sets its own. Knowing this saves debugging time: a retry policy on the root silently applies to every leaf, so a fallback chain under it retries every target before moving to the next.
Exactly when a fallback moves on
The fallback loop is short enough to quote in spirit. For each target in order, the gateway calls the target recursively, then decides whether to stop:
for target in node.targets:
response = try_targets_recursively(target)
codes = node.strategy.on_status_codes # may be absent
if codes is not None and response.status not in codes:
break # a status you did not list: stop here
if codes is None and response.ok:
break # no list: stop only on 2xx
if response.headers["x-portkey-gateway-exception"] == "true":
break # gateway's own error: do not cascade
return response # the last response, success or notTwo consequences follow. Without on_status_codes, any non-2xx response moves on, including a 400 caused by a malformed request or a context-length error. A bad prompt then hits every target in the chain, costs latency and money on each, and returns the last error, which may be a different message from the first. With on_status_codes, the chain stops on any status you did not list, including 400, which is usually what you want. List the transient codes explicitly: [408, 429, 500, 502, 503, 504].
Second, a 446 (an input or output guardrail denied the request) is not 2xx. In a fallback node with no status list, a denial on the first target falls through and the same request is sent to the next provider. If that is not intended, set on_status_codes so that 446 ends the chain.
Load balancing and conditional routing
In loadbalance mode each target has a weight (default 1). The gateway draws a uniform random number in [0, total weight) and walks the targets subtracting weights until the number falls inside one. It is a single weighted draw per request: no health tracking, no least-connections, no awareness of queue depth on the GPU pool. If pool B is down, a quarter of requests in the example below still go to it. That is why a load-balance node should almost always sit inside a fallback node, so the failing share is caught by the next target instead of surfacing to the user.
Conditional mode routes by request content. Each condition has a query using MongoDB-style operators such as $eq, $in and $gt over metadata.* (from the x-portkey-metadata header) or params.* (the request body), and a then that names a target by its name field; default names the fallback. A typo in a target name is a runtime router error, not a config-load error, so test every branch.
Retries and timeouts
A retry block has attempts, optional on_status_codes and use_retry_after_header. The source caps attempts at 5 and defaults the retryable codes to 429, 500, 502, 503 and 504. Between attempts it uses exponential backoff. When retry-after headers are honoured (retry-after-ms, x-ms-retry-after-ms or retry-after), a provider asking you to wait 60 seconds or more ends the retries for that target rather than parking the request.
request_timeout (milliseconds) aborts the upstream fetch and synthesises a 408 response. 408 is not in the default retry list, so a timed-out request is neither retried nor, in a fallback node that lists only 429 and 5xx, failed over. If timeouts are the main symptom of an overloaded GPU pool, which they usually are, add 408 to both lists.
Retries multiply. attempts counts retries after the first call, so attempts: 5 on each of three fallback targets is up to eighteen upstream calls for one user request, with backoff in between. Budget the worst-case latency explicitly: the sum over targets of timeout times (attempts + 1), plus backoff, must fit the client's own deadline, or the client gives up while the gateway is still trying.
Guardrails in the request path
Guardrails attach to a node as input_guardrails (run before the upstream call) and output_guardrails (run on the response). Each entry names checks, for example the built-in default.contains, plus deny, on_fail, on_success and async keys. With deny: true a failed check returns HTTP 446 and the guardrail results; with deny: false the original 200 is relabelled 246 ("Hooks failed") so the client can see the check failed but still gets the body. Many HTTP clients treat 246 as success and 446 as an unknown client error, so check how yours handles both.
Streaming is the important caveat. In the response handler of the open-source build, a streamed body is passed through, and failed hooks only change the status to 246. An output guardrail does not hold back or cut a token stream there. If you need to block unsafe streamed output, you need a holdback design in front of the client, as described in streaming moderation.
Worked example: two vLLM pools and a hosted fallback
The goal: serve an 8B instruct model from two vLLM pools, three quarters of traffic to the larger pool A, and fail over to a hosted model only when both pools are struggling. vLLM exposes an OpenAI-compatible server, so each pool is an openai provider with a custom_host.
{
"strategy": { "mode": "fallback", "on_status_codes": [408, 429, 500, 502, 503, 504] },
"request_timeout": 20000,
"targets": [
{
"strategy": { "mode": "loadbalance" },
"retry": { "attempts": 2, "on_status_codes": [429, 503] },
"targets": [
{ "provider": "openai", "custom_host": "http://vllm-a.internal:8000/v1",
"api_key": "unused", "weight": 3 },
{ "provider": "openai", "custom_host": "http://vllm-b.internal:8000/v1",
"api_key": "unused", "weight": 1 }
],
"override_params": { "model": "meta-llama/Llama-3.1-8B-Instruct" }
},
{
"provider": "openai", "api_key": "sk-...",
"override_params": { "model": "gpt-4o-mini" }
}
]
}Trace a request when pool A is saturated. The weighted draw lands on A (probability 3/4). A returns 503; the inner retry block retries A up to 2 more times with backoff; each fails. The load-balance node has no strategy codes of its own, so its result is A's final 503, which goes back to the outer fallback node. 503 is in the outer list, so the gateway moves to the hosted target, where override_params swaps in the hosted model name. Note what did not happen: pool B, which was healthy, was never tried, because load balancing picks one target per request. If B should absorb A's overflow, make the inner node a fallback from A to B instead, or put B again as a second outer target.
The client reads the response headers to see which path served it:
import json, os
from openai import OpenAI
config = json.load(open("gateway_config.json"))
client = OpenAI(
base_url="http://gateway.internal:8787/v1",
api_key="ignored-by-gateway", # keys live in the config
default_headers={"x-portkey-config": json.dumps(config)},
)
raw = client.chat.completions.with_raw_response.create(
model="placeholder", # overridden per target
messages=[{"role": "user", "content": "Summarise this ticket ..."}],
max_tokens=256,
)
print(raw.headers.get("x-portkey-last-used-option-index"),
raw.headers.get("x-portkey-retry-attempt-count"),
raw.headers.get("x-portkey-trace-id"))
resp = raw.parse()Log x-portkey-last-used-option-index (the JSON path of the target that answered) with every request. Its distribution over time is the cheapest dashboard you can build: a rising share of the hosted index means your GPU pools are short of capacity, and it shows up before latency percentiles move.
Running it in production
- Run it next to the application, not across a WAN:
docker run -p 8787:8787 portkeyai/gateway:latest. It adds a network hop and JSON transformation per request, so keep it on the same host or cluster network. - Run several replicas behind your normal load balancer. The routing logic is per request, so replicas need no coordination; the trade-off is that no replica knows another's view of a failing target.
- Never expose the gateway publicly without your own authentication in front. Inline configs carry provider keys, and anyone who can reach port 8787 can spend them.
- Keep configs in version control and send them from the application, or load them from a secrets manager into a sidecar. Reviewing a fallback order change should look like reviewing code.
- Trace IDs: pass
x-portkey-trace-idfrom your request context so gateway logs join your application traces.
Failure modes
- Cascading bad requests. A fallback node without
on_status_codessends a 400 to every target. Fix: always list the transient codes. - Retry storms. Root-level retries inherited by every leaf multiply upstream load exactly when a GPU pool is overloaded. Fix: retries on the inner pool node only, few attempts.
- Blind load balancing. Weighted random keeps sending a fixed share to a dead pool. Fix: wrap in a fallback, and drain the pool by setting its weight to 0 or removing it during incidents.
- Timeouts that never fail over. 408 is missing from the default lists. Fix: add it.
- Parameter drift between targets. A hosted fallback model may reject parameters your vLLM pool accepts, or tokenise the same prompt to a different length. Fix: per-target
override_paramsand a test that sends a real request to every leaf. - Streaming guardrails that only label. A 246 on a stream is a log entry, not a block.
- Version drift. Field names and status handling have changed between releases; pin the image tag and re-test on upgrade.
Trade-offs and alternatives
Choose Portkey's gateway when you want routing as data, multi-provider translation and request-level guardrail hooks without writing them, and you are comfortable with a stateless proxy. LiteLLM covers similar ground with a Python proxy and a stronger built-in story for per-key budgets in the open-source build; compare them on the feature you actually need rather than feature count. If your traffic only ever goes to your own GPUs, a gateway may be unnecessary: vLLM behind an ordinary L7 load balancer with active health checks gives you health-aware balancing, which Portkey's open-source loadbalance mode does not. The broader design space is laid out in the AI gateway overview and LLM gateway architecture, and response caching is covered in semantic caching.
What to do next
- Start the gateway locally with npx and send one request through it to each provider you use.
- Write your routing as a config file: an outer fallback, an inner pool node, a hosted leaf.
- Set
on_status_codeson every fallback node, including 408, and decide whether 446 should stop the chain. - Put retries on the inner node only and compute the worst-case latency against your client deadline.
- Test each leaf directly with a real request to catch parameter and model-name mismatches.
- Log
x-portkey-last-used-option-indexand alert when the hosted share rises. - If you need output blocking on streams, design a holdback layer; do not rely on 246.
- Pin the gateway version and re-run the test suite on every upgrade.