A multi-provider LLM strategy means your product can serve each request from more than one model source: two or more hosted API vendors, a cloud platform's model catalogue, or your own GPUs running an open-weight model. Teams adopt it after an outage, after hitting a rate limit during a launch, or after noticing that half their traffic would run on a model a tenth of the price. Done well, it gives you availability, capacity and negotiating leverage. Done badly, it gives you an abstraction layer that hides quality regressions, doubles your evaluation work and fails over at the exact moment the second provider cannot take the load.
This article covers the architecture, the code for adapters and failover, the capacity arithmetic that decides whether failover can work at all, the portability traps that make 'the same prompt' behave differently, and how self-hosted GPU capacity fits in as one more provider. Choosing which model tier handles which request is a separate decision, covered in LLM Routing Strategies; the generic gateway component is in AI Gateway Overview.
Why go multi-provider, and for what
List the reasons explicitly, because each one implies a different design.
| Goal | What it requires | What it does not require |
|---|---|---|
| Availability | A warm second backend with quota for full traffic | Identical quality on every task |
| Capacity bursts | Spillover when the primary returns 429 | Fast failover on errors |
| Cost | Per-task routing to a cheaper backend that passes evals | Real-time switching |
| Data residency | A backend in the required region for some tenants | Load balancing |
| Capability | The best model per task, wherever it lives | Interchangeability |
| Leverage | A credible, tested exit path | Splitting traffic evenly |
A team that wants availability but builds for cost will route by price and discover during an outage that the fallback has a tenth of the quota it needs. Decide which two goals matter most and test for those.
Architecture: one contract, many backends
The core design move is to give applications a single internal contract that names a task, not a vendor or model. A request says 'summarise-ticket, data class internal, max 400 output tokens'; the platform decides where it runs. Behind that contract sit four parts: a policy engine that removes backends the request may not use, a pool selector that picks among the rest by health, quota and cost, one adapter per backend that translates the canonical request, and an evaluation gate that decides which backends may serve each task at all.
from dataclasses import dataclass, field
@dataclass
class LLMRequest:
task: str # "summarise-ticket", not a model name
messages: list # canonical role/content pairs
tools: list = field(default_factory=list) # canonical JSON Schema tools
max_output_tokens: int = 512
data_class: str = "internal" # drives the policy engine
region: str | None = None
class Adapter:
name: str
def count_tokens(self, req: LLMRequest) -> int: ... # this backend's tokenizer
def stream(self, req: LLMRequest): ... # yields canonical events
def supports(self, req: LLMRequest) -> bool: ... # tools, context length, modalitySelf-hosted capacity fits the same shape. vLLM and similar servers expose an OpenAI-compatible HTTP API, so the adapter for your own GPUs is often the thinnest one you write. Comparing serving engines is covered in LLM Serving Stacks.
Safe failover
Failover logic looks trivial and is where most incidents originate. Four rules keep it safe.
- Retry only before the first token. Once a stream has delivered text to the user, switching providers produces a reply stitched from two models. Fail the request or restart it visibly.
- Classify errors. Rate limits, server errors, timeouts and provider overload responses are retryable on another backend; validation errors and content-policy refusals are not, because the next provider will probably reject the same request too.
- Break circuits per backend and model. A provider can be healthy for one model and overloaded for another.
- Make tool calls idempotent. If a model emitted a tool call that your system executed before the stream died, the retry must not execute it twice.
import time
RETRYABLE = {"rate_limited", "overloaded", "server_error", "timeout"}
class Breaker:
def __init__(self, threshold=5, cooldown_s=30):
self.fails, self.open_until = 0, 0.0
self.threshold, self.cooldown_s = threshold, cooldown_s
def available(self):
return time.monotonic() >= self.open_until
def record(self, ok):
self.fails = 0 if ok else self.fails + 1
if self.fails >= self.threshold:
self.open_until = time.monotonic() + self.cooldown_s
self.fails = 0
def serve(req, pool, breakers):
last = None
for backend in pool: # already filtered and ordered
if not breakers[backend.name].available() or not backend.supports(req):
continue
first_token_sent = False
try:
for event in backend.stream(req):
first_token_sent = True
yield event
breakers[backend.name].record(ok=True)
return
except ProviderError as e:
breakers[backend.name].record(ok=False)
last = e
if first_token_sent or e.kind not in RETRYABLE:
raise # never stitch two models' output
raise NoBackendAvailable(last)
Capacity arithmetic: a worked example
Failover is a capacity plan, not a code path. Work the numbers for a support product at peak: 400 requests per minute, averaging 3,000 input and 400 output tokens, so 1.36 million tokens per minute. In normal operation the team sends 70 percent to provider A and 30 percent to provider B, so B sees about 408,000 tokens per minute and the team's B quota was sized at 500,000 for comfort.
Now A fails. B must absorb all 1.36 million tokens per minute, almost three times its quota. Within seconds B starts returning rate-limit errors, the breaker opens on B as well, and a single-vendor outage has become a total one. The fix is to size every failover target for the traffic it must carry when the primary is gone, then shed what cannot fit by priority: keep interactive chat, delay batch summarisation, and switch long-context tasks to a shorter retrieval budget. Note too that 3,000 input tokens on one tokenizer may be 3,300 on another, so quotas expressed in tokens must be converted with each backend's own counter rather than copied.
Self-hosted GPUs change the arithmetic. Reserved capacity is cheapest when it runs flat out, so it suits a steady baseline, while hosted APIs absorb bursts and failover. Whether that baseline is cheaper than API pricing depends on utilisation; the method for computing cost per million tokens from a GPU-hour is in LLM Cost Analysis.
Self-hosted GPUs as a provider
Treating your own GPUs as one more provider is attractive because the adapter is thin, but the operational contract is different from a hosted API. A hosted provider hides capacity behind a quota; your cluster exposes it directly, so the pool selector must see real headroom, such as queue depth, KV cache utilisation and time to first token, rather than a static rate limit. When the queue grows, the selector should spill new requests to a hosted backend before latency breaches the SLO, not after.
Three practices keep a self-hosted backend honest. First, pin the exact model weights, quantisation and serving-engine version in the adapter's metadata, and re-run the evaluation gate whenever any of them changes, because a new quantisation can move quality as much as a new model. Second, reserve headroom for failover in the other direction: if a hosted provider fails, your cluster will receive its traffic, and a cluster sized at 90 percent utilisation has nothing to give. Third, measure cost per task the same way for every backend, including idle GPU hours, so the comparison with hosted pricing is not flattered by ignoring the nights. Latency behaviour under load is explained in GPU Inference Latency.
Portability traps
Two providers accepting the same JSON does not make them interchangeable. The differences that bite in production are predictable.
- Tokenizers. Context limits, output caps, quotas and cost are all measured in each backend's own tokens. Count with the target's tokenizer before sending, or a prompt that fits on one provider is truncated or rejected on another.
- Prompt caching. Cached-prefix discounts and latency gains live inside one provider. Failing over sends every request cold, so cost and time to first token both jump just when you are degraded.
- Tool calling. Schemas, parallel-call behaviour and how arguments are streamed differ; normalise them in the adapter and test each tool on each backend.
- Structured output. Some backends enforce a JSON schema during decoding and some only follow instructions. Validate every response regardless.
- Instruction style. System prompts tuned for one model family can underperform on another. Keep per-backend prompt variants in version control, keyed by task.
- Refusal and safety behaviour. The same request may be answered by one model and refused by another; your evaluation set should include borderline cases.
- Data terms. Retention, training use and processing regions differ by provider and plan, so the policy engine must know each backend's terms.
Evaluation gates and drift
No backend should join a task's pool until it passes that task's evaluation set at a threshold you set in advance. Keep a golden set per task: real, anonymised inputs with reference answers or graded rubrics, plus the hard cases from past incidents. Run every candidate backend, with its own prompt variant, through the set, and record pass rate, cost per task and latency percentiles side by side.
def eligible(task, results, thresholds):
# results[backend] = {"pass_rate": .., "p95_ms": .., "cost": ..}
t = thresholds[task]
pool = [b for b, r in results.items()
if r["pass_rate"] >= t["min_pass"] and r["p95_ms"] <= t["max_p95_ms"]]
return sorted(pool, key=lambda b: results[b]["cost"]) # cheapest firstRe-run the gate when any provider ships a model update, because hosted models behind a stable name can change. Shadow a small slice of live traffic to secondary backends continuously, grading offline, so you learn that the fallback has drifted before an outage forces you to use it. Track error budgets per backend with the approach in LLM SLO Burn Rate Alerts.
Failure modes
- Undersized fallback quota. The second provider throttles under failover load, turning a partial outage into a full one.
- Silent quality drop. Failover succeeds technically while answer quality falls; without per-backend quality metrics nobody notices.
- Stitched streams. Retrying mid-stream on another model yields incoherent replies and duplicated tool actions.
- Retry storms. Every client retries every backend at once; use jittered backoff and a global concurrency limit.
- Policy bypass. A fallback path skips the residency check and sends regulated data to a region it must not reach.
- Lowest common denominator. The abstraction exposes only features every backend supports, so nobody uses the strongest model's capabilities.
- Untested exit. The second provider is configured but has served no traffic for months; credentials, quotas or model names have expired.
Trade-offs
| Design | Benefit | Cost |
|---|---|---|
| Single provider, multi-region | One prompt set, deepest features | Exposed to vendor-wide incidents and pricing |
| Active-passive second provider | Availability with modest effort | Fallback rots unless exercised |
| Active-active split | Fallback stays warm, leverage | Two prompt variants, two eval runs per change |
| Hosted plus self-hosted baseline | Lower steady-state cost, data control | GPU operations, capacity planning, model upgrades |
| Per-task best model | Highest quality per task | Most integration and evaluation work |
What to do next
- Write down which two goals, from availability, capacity, cost, residency, capability and leverage, the strategy must serve.
- Introduce a task-named internal API so no application calls a vendor SDK directly.
- Build one adapter per backend, including token counting with that backend's tokenizer.
- Create a golden evaluation set per high-volume task and gate pool membership on it.
- Compute failover capacity for each fallback at full primary traffic, and request the quota or define load shedding.
- Implement per-backend circuit breakers with retry only before the first token.
- Shadow a small share of traffic to every fallback continuously and grade it offline.
- Run a game day: disable the primary in staging, then in production for a low-risk task, and measure what happens.
- Record each provider's data terms in the policy engine and test that fallbacks respect them.