OpenRouter is a hosted API that sits between your application and many model providers. You send an OpenAI-style chat completion request with one key; it picks a model (the one you named, a fallback from a list you gave, or one chosen by its auto router), picks a provider that serves that model, forwards the request to that provider's GPUs, and returns the result with a usage record that includes the cost. The pitch is breadth and convenience: hundreds of models, one bill, one SDK, and automatic failover between providers hosting the same open-weight model.
This article is about operating it well. It explains what the router actually decides on each request, the documented request fields that control those decisions, the failure modes that bite real deployments (the worst is an error that arrives with an HTTP 200 status), and how to decide between OpenRouter, a self-hosted gateway and direct provider contracts. The generic theory of model routing is in LLM routers and the self-hosted alternative in LiteLLM; here we stay concrete. Field names below were checked against the OpenRouter documentation in October 2026. Model identifiers in the examples are placeholders: copy current IDs from the OpenRouter model list rather than from any article, because they change.
Two decisions per request
Think of every request as passing through two decisions. The first is the model: which weights should answer. The second is the provider: which company's GPU fleet runs those weights for this request. For a closed model there may be one or two providers (the lab itself and perhaps a cloud reseller). For a popular open-weight model there may be many, differing in price, throughput, context length, supported parameters, quantisation and data policy. That second decision is where most of OpenRouter's value and most of its surprises live.
The default provider policy, as documented, is: prefer providers that have not had a significant outage in the last 30 seconds; among those, choose by price using inverse square weighting; keep the rest as fallbacks. Inverse square means a provider charging $1 per million tokens is nine times more likely to be tried first than one charging $3, because 1/1^2 against 1/3^2 is 9 to 1. The cheap provider gets most traffic, but not all of it, so load is spread and a single cheap provider's slowdown does not capture every request.
The API surface
The API is OpenAI-compatible: base URL https://openrouter.ai/api/v1, endpoint /chat/completions, and a bearer token in the Authorization header. That means the official OpenAI SDKs work by changing the base URL; OpenRouter-specific fields travel in extra_body in Python or as extra properties in TypeScript.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
timeout=60, # seconds; always set your own deadline
max_retries=0, # we own retries, see below
)
resp = client.chat.completions.create(
model="vendor-a/model-large", # illustrative ID
messages=[{"role": "user", "content": "Summarise this ticket: ..."}],
max_tokens=400,
extra_headers={
"HTTP-Referer": "https://example.com", # optional attribution
"X-OpenRouter-Title": "Ticket Summariser", # older examples: X-Title
},
extra_body={
"models": ["vendor-a/model-large", "vendor-b/model-medium"], # fallbacks
"provider": {
"data_collection": "deny", # skip providers that may store/train on data
"allow_fallbacks": True, # default; set False to pin providers
"require_parameters": True, # only providers that honour every parameter
},
},
)
print(resp.model) # the model that actually answered
print(resp.usage.prompt_tokens, resp.usage.completion_tokens)
print(getattr(resp.usage, "cost", None)) # OpenRouter adds cost to usageThree details in that listing are deliberate. First, require_parameters defaults to false, which means a provider that silently ignores a parameter you sent (a tool schema, a response format, a sampling setting) is still eligible. If your code depends on a parameter, turn this on. Second, data_collection defaults to "allow"; regulated or customer data should set it to "deny", and the zdr field restricts routing to zero-data-retention endpoints. Third, the reported model field is what you log: with fallbacks, the model you asked for and the model that answered can differ, and you are billed for the one that answered.
The provider object
The provider object is the main control surface. Its documented fields, grouped by purpose:
| Purpose | Fields | Use it when |
|---|---|---|
| Pin or exclude | order, only, ignore, allow_fallbacks | You validated output quality on specific providers, or one misbehaves |
| Correctness | require_parameters, quantizations | Your prompt depends on tools, JSON output, or full-precision weights |
| Data policy | data_collection, zdr, enforce_distillable_text | Customer data, compliance reviews, licence constraints |
| Performance | sort, preferred_min_throughput, preferred_max_latency | Interactive traffic where tokens per second matter more than price |
| Cost ceiling | max_price | Batch jobs that must never be routed to an expensive provider |
Two shortcuts exist as model-name suffixes: appending :nitro asks for the highest-throughput routing and :floor for the lowest price. They are handy for experiments; in production, prefer explicit provider fields so the policy is visible in code review.
Quantisation deserves a warning. The same open-weight model served at different precisions by different providers is not the same model for evaluation purposes. If your evaluation suite passed on one provider and production traffic is spread across five by price weighting, you are running an untested configuration on most requests. Either pin the providers you evaluated with order and allow_fallbacks set to false, or restrict quantizations to the levels you tested, and re-run evaluations when a new provider appears.
Streaming and the HTTP 200 error
With streaming, the HTTP status line is sent before the model produces its first token. If the provider fails halfway through, the router cannot change the status to an error. OpenRouter documents what it does instead: the stream carries one more event containing an error object (with a code, a message and metadata such as the error type and the provider's own code) and a choice whose finish_reason is "error", and then the stream ends. The HTTP status stays 200.
Many clients do not look for that. They accumulate deltas, see the stream close, and hand a truncated answer to the user, or worse, to the next step of an agent pipeline that parses it as complete JSON. The listing below treats a mid-stream error exactly like a failed request.
import json, os, requests
def stream_chat(body, timeout=(5, 60)):
"""Yield text deltas; raise on mid-stream errors that arrive under HTTP 200."""
r = requests.post(
"https://openrouter.ai/api/v1/chat/completions",
headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
json={**body, "stream": True},
stream=True, timeout=timeout,
)
if r.status_code != 200: # pre-stream error: normal JSON body
err = r.json().get("error", {})
raise UpstreamError(r.status_code, err.get("message"), err.get("metadata"))
for raw in r.iter_lines(decode_unicode=True):
if not raw or raw.startswith(":"): # keep-alive comments
continue
data = raw.removeprefix("data: ")
if data == "[DONE]":
return
chunk = json.loads(data)
if "error" in chunk: # headers already sent: status was 200
e = chunk["error"]
raise UpstreamError(e.get("code"), e.get("message"), e.get("metadata"))
choices = chunk.get("choices") or []
if not choices: # e.g. a usage-only final chunk
continue
choice = choices[0]
if choice.get("finish_reason") == "error":
raise UpstreamError(502, "stream ended with finish_reason=error", None)
yield choice.get("delta", {}).get("content") or ""
class UpstreamError(Exception):
def __init__(self, code, message, metadata):
super().__init__(f"{code}: {message}")
self.code, self.metadata = code, metadataRetrying a failed stream is a product decision, not a transport one. The tokens already shown to a user cannot be taken back, and a retry may produce a different answer. For chat, show an error and offer regeneration. For machine consumers, buffer the whole response, validate it, and only then pass it on, so a retry is invisible.
Error codes and what to do
| Status | Documented meaning | What your code should do |
|---|---|---|
| 400 | Invalid or missing parameters | Do not retry; log the body and fix the request |
| 401 | Invalid or disabled key, expired OAuth session | Do not retry; page the owner, rotate the key |
| 402 | Account or key has insufficient credits | Do not retry; alert, because every request will fail until topped up |
| 403 | Insufficient permissions, guardrail block or moderation flag | Do not retry the same input; surface a policy message |
| 408 | Request timed out | Retry with backoff, possibly with a smaller max_tokens |
| 429 | Rate limited | Retry with exponential backoff and jitter; shed load |
| 502 | Chosen model down or invalid upstream response | Retry; consider a models fallback list |
| 503 | No provider meets your routing requirements | Your constraints are too tight; relax them or alert |
Two rows deserve emphasis. A 402 is an operational outage you cause yourself: credits are prepaid, so put a balance alarm on the account well before zero. A 503 is often a configuration error rather than an outage: combining data_collection deny, a narrow quantizations list, require_parameters and a max_price can leave no eligible provider for a model, and every request fails. Test each routing policy against each model you use before shipping it. The broader patterns for retries and circuit breakers across providers are in provider failover.
Model fallbacks and the auto router
Two kinds of fallback stack. Provider fallback (on by default) keeps the model fixed and moves to another host. Model fallback, via the models array, moves to a different model when the first fails; the documentation lists context-length validation errors, moderation flags, rate limiting and downtime as triggers. The response's model field names the model that answered, and pricing follows that model. The documentation shows models both on its own and alongside model without spelling out how they interact, so list your primary first in models as well; then the order is unambiguous.
Model fallback is powerful and dangerous in equal measure. It keeps the request alive, but the second model may have a different context window, tool-calling dialect, refusal behaviour and price. Fall back only between models you have evaluated on your own task, and record the answering model with every output so quality regressions can be traced. The openrouter/auto model goes further and picks a model for you, charging the selected model's normal rate; it is useful for exploration and for low-stakes traffic, and hard to reason about for anything you have to evaluate.
Operating it in production
- Log the right four fields: the response
id(the generation id, usable for later usage lookups), the answeringmodel,usage.cost, and your own request id. Token counts are reported with the model's native tokenizer, so they will not match a tiktoken estimate for non-OpenAI models. - Budget per key. Issue separate keys per service and environment so a runaway loop in staging cannot drain production credits, and alert on spend rate, not only on balance.
- Set your own timeouts. Use a short connect timeout and a read timeout sized to max_tokens divided by the slowest acceptable throughput. A stalled stream otherwise holds a worker forever.
- Treat the router as a dependency with its own outages. If OpenRouter is a single path to every model, its incident is your incident. Critical paths can keep a direct provider client as a cold standby, a pattern covered in multi-provider strategies.
- Pin for evaluation, spread for resilience. Run evaluations with providers pinned, then widen routing only to providers that pass.
Trade-offs
| Option | Strengths | Costs |
|---|---|---|
| OpenRouter | Hundreds of models behind one key; provider failover; per-request cost in the response | Extra network hop; another party in the data path; quality varies by provider |
| Self-hosted gateway (LiteLLM and similar) | Your network, your logs, your policy code | You hold every provider contract and run the gateway |
| Direct provider SDKs | Fewest hops, direct support and committed-use pricing | One integration per provider; you build failover |
| Self-hosted open weights | Full control over precision, latency and data | You own the GPUs, the serving stack and the on-call |
A common and sensible path: prototype on OpenRouter to compare many models cheaply; move steady high-volume traffic to direct contracts or your own serving once a model wins; keep OpenRouter for the long tail of models and as a fallback. If you need central policy enforcement across all of that, see AI gateways.
What to do next
- Point an existing OpenAI-SDK client at the OpenRouter base URL and log
model,usage.costand the response id for every call. - Add the streaming guard from this article and test it by forcing a mid-stream failure, for example with a tiny read timeout.
- Write down a provider policy per use case (data collection, required parameters, precision) and verify each one returns 200 for every model it covers.
- Pin providers for your evaluation run, then decide which extra providers may join production routing.
- Create one key per service and environment, with spend alerts.
- Decide, in writing, which traffic graduates to direct contracts or self-hosting and at what monthly volume.