A multi-provider LLM strategy means your product can serve each request from more than one model source: two or more hosted API vendors, a cloud platform's model catalogue, or your own GPUs running an open-weight model. Teams adopt it after an outage, after hitting a rate limit during a launch, or after noticing that half their traffic would run on a model a tenth of the price. Done well, it gives you availability, capacity and negotiating leverage. Done badly, it gives you an abstraction layer that hides quality regressions, doubles your evaluation work and fails over at the exact moment the second provider cannot take the load.

This article covers the architecture, the code for adapters and failover, the capacity arithmetic that decides whether failover can work at all, the portability traps that make 'the same prompt' behave differently, and how self-hosted GPU capacity fits in as one more provider. Choosing which model tier handles which request is a separate decision, covered in LLM Routing Strategies; the generic gateway component is in AI Gateway Overview.

Why go multi-provider, and for what

List the reasons explicitly, because each one implies a different design.

GoalWhat it requiresWhat it does not require
AvailabilityA warm second backend with quota for full trafficIdentical quality on every task
Capacity burstsSpillover when the primary returns 429Fast failover on errors
CostPer-task routing to a cheaper backend that passes evalsReal-time switching
Data residencyA backend in the required region for some tenantsLoad balancing
CapabilityThe best model per task, wherever it livesInterchangeability
LeverageA credible, tested exit pathSplitting traffic evenly

A team that wants availability but builds for cost will route by price and discover during an outage that the fallback has a tenth of the quota it needs. Decide which two goals matter most and test for those.

Architecture: one contract, many backends

The core design move is to give applications a single internal contract that names a task, not a vendor or model. A request says 'summarise-ticket, data class internal, max 400 output tokens'; the platform decides where it runs. Behind that contract sit four parts: a policy engine that removes backends the request may not use, a pool selector that picks among the rest by health, quota and cost, one adapter per backend that translates the canonical request, and an evaluation gate that decides which backends may serve each task at all.

Multi-provider LLM architecture: one internal contract, many interchangeable backendsApplicationscall one internal APIInternal LLM APIcanonical request, task name, data classPolicy engineresidency, data classPool selectorhealth, quota, costEval gategolden set per providerTelemetrySLIs, tokens, cost per taskapproves poolAdapter: hosted provider Abreaker, retry, token countAdapter: hosted provider Bbreaker, retry, token countAdapter: self-hosted GPUsvLLM, OpenAI-compatiblemetricsApplications never name a vendor. The selector picks a backend that the policy allows and the eval gate has approved,and every adapter reports the same metrics so cost and quality are comparable across providers.
Applications call one task-oriented API. Policy filters the pool, the selector picks a healthy backend with quota, and adapters translate the canonical request.
from dataclasses import dataclass, field

@dataclass
class LLMRequest:
    task: str                 # "summarise-ticket", not a model name
    messages: list            # canonical role/content pairs
    tools: list = field(default_factory=list)  # canonical JSON Schema tools
    max_output_tokens: int = 512
    data_class: str = "internal"   # drives the policy engine
    region: str | None = None

class Adapter:
    name: str
    def count_tokens(self, req: LLMRequest) -> int: ...   # this backend's tokenizer
    def stream(self, req: LLMRequest): ...                # yields canonical events
    def supports(self, req: LLMRequest) -> bool: ...      # tools, context length, modality

Self-hosted capacity fits the same shape. vLLM and similar servers expose an OpenAI-compatible HTTP API, so the adapter for your own GPUs is often the thinnest one you write. Comparing serving engines is covered in LLM Serving Stacks.

Safe failover

Failover logic looks trivial and is where most incidents originate. Four rules keep it safe.

  1. Retry only before the first token. Once a stream has delivered text to the user, switching providers produces a reply stitched from two models. Fail the request or restart it visibly.
  2. Classify errors. Rate limits, server errors, timeouts and provider overload responses are retryable on another backend; validation errors and content-policy refusals are not, because the next provider will probably reject the same request too.
  3. Break circuits per backend and model. A provider can be healthy for one model and overloaded for another.
  4. Make tool calls idempotent. If a model emitted a tool call that your system executed before the stream died, the retry must not execute it twice.
import time

RETRYABLE = {"rate_limited", "overloaded", "server_error", "timeout"}

class Breaker:
    def __init__(self, threshold=5, cooldown_s=30):
        self.fails, self.open_until = 0, 0.0
        self.threshold, self.cooldown_s = threshold, cooldown_s
    def available(self):
        return time.monotonic() >= self.open_until
    def record(self, ok):
        self.fails = 0 if ok else self.fails + 1
        if self.fails >= self.threshold:
            self.open_until = time.monotonic() + self.cooldown_s
            self.fails = 0

def serve(req, pool, breakers):
    last = None
    for backend in pool:                     # already filtered and ordered
        if not breakers[backend.name].available() or not backend.supports(req):
            continue
        first_token_sent = False
        try:
            for event in backend.stream(req):
                first_token_sent = True
                yield event
            breakers[backend.name].record(ok=True)
            return
        except ProviderError as e:
            breakers[backend.name].record(ok=False)
            last = e
            if first_token_sent or e.kind not in RETRYABLE:
                raise                        # never stitch two models' output
    raise NoBackendAvailable(last)

Capacity arithmetic: a worked example

Failover is a capacity plan, not a code path. Work the numbers for a support product at peak: 400 requests per minute, averaging 3,000 input and 400 output tokens, so 1.36 million tokens per minute. In normal operation the team sends 70 percent to provider A and 30 percent to provider B, so B sees about 408,000 tokens per minute and the team's B quota was sized at 500,000 for comfort.

Now A fails. B must absorb all 1.36 million tokens per minute, almost three times its quota. Within seconds B starts returning rate-limit errors, the breaker opens on B as well, and a single-vendor outage has become a total one. The fix is to size every failover target for the traffic it must carry when the primary is gone, then shed what cannot fit by priority: keep interactive chat, delay batch summarisation, and switch long-context tasks to a shorter retrieval budget. Note too that 3,000 input tokens on one tokenizer may be 3,300 on another, so quotas expressed in tokens must be converted with each backend's own counter rather than copied.

Self-hosted GPUs change the arithmetic. Reserved capacity is cheapest when it runs flat out, so it suits a steady baseline, while hosted APIs absorb bursts and failover. Whether that baseline is cheaper than API pricing depends on utilisation; the method for computing cost per million tokens from a GPU-hour is in LLM Cost Analysis.

Self-hosted GPUs as a provider

Treating your own GPUs as one more provider is attractive because the adapter is thin, but the operational contract is different from a hosted API. A hosted provider hides capacity behind a quota; your cluster exposes it directly, so the pool selector must see real headroom, such as queue depth, KV cache utilisation and time to first token, rather than a static rate limit. When the queue grows, the selector should spill new requests to a hosted backend before latency breaches the SLO, not after.

Three practices keep a self-hosted backend honest. First, pin the exact model weights, quantisation and serving-engine version in the adapter's metadata, and re-run the evaluation gate whenever any of them changes, because a new quantisation can move quality as much as a new model. Second, reserve headroom for failover in the other direction: if a hosted provider fails, your cluster will receive its traffic, and a cluster sized at 90 percent utilisation has nothing to give. Third, measure cost per task the same way for every backend, including idle GPU hours, so the comparison with hosted pricing is not flattered by ignoring the nights. Latency behaviour under load is explained in GPU Inference Latency.

Portability traps

Two providers accepting the same JSON does not make them interchangeable. The differences that bite in production are predictable.

  • Tokenizers. Context limits, output caps, quotas and cost are all measured in each backend's own tokens. Count with the target's tokenizer before sending, or a prompt that fits on one provider is truncated or rejected on another.
  • Prompt caching. Cached-prefix discounts and latency gains live inside one provider. Failing over sends every request cold, so cost and time to first token both jump just when you are degraded.
  • Tool calling. Schemas, parallel-call behaviour and how arguments are streamed differ; normalise them in the adapter and test each tool on each backend.
  • Structured output. Some backends enforce a JSON schema during decoding and some only follow instructions. Validate every response regardless.
  • Instruction style. System prompts tuned for one model family can underperform on another. Keep per-backend prompt variants in version control, keyed by task.
  • Refusal and safety behaviour. The same request may be answered by one model and refused by another; your evaluation set should include borderline cases.
  • Data terms. Retention, training use and processing regions differ by provider and plan, so the policy engine must know each backend's terms.

Evaluation gates and drift

No backend should join a task's pool until it passes that task's evaluation set at a threshold you set in advance. Keep a golden set per task: real, anonymised inputs with reference answers or graded rubrics, plus the hard cases from past incidents. Run every candidate backend, with its own prompt variant, through the set, and record pass rate, cost per task and latency percentiles side by side.

def eligible(task, results, thresholds):
    # results[backend] = {"pass_rate": .., "p95_ms": .., "cost": ..}
    t = thresholds[task]
    pool = [b for b, r in results.items()
            if r["pass_rate"] >= t["min_pass"] and r["p95_ms"] <= t["max_p95_ms"]]
    return sorted(pool, key=lambda b: results[b]["cost"])   # cheapest first

Re-run the gate when any provider ships a model update, because hosted models behind a stable name can change. Shadow a small slice of live traffic to secondary backends continuously, grading offline, so you learn that the fallback has drifted before an outage forces you to use it. Track error budgets per backend with the approach in LLM SLO Burn Rate Alerts.

Failure modes

  • Undersized fallback quota. The second provider throttles under failover load, turning a partial outage into a full one.
  • Silent quality drop. Failover succeeds technically while answer quality falls; without per-backend quality metrics nobody notices.
  • Stitched streams. Retrying mid-stream on another model yields incoherent replies and duplicated tool actions.
  • Retry storms. Every client retries every backend at once; use jittered backoff and a global concurrency limit.
  • Policy bypass. A fallback path skips the residency check and sends regulated data to a region it must not reach.
  • Lowest common denominator. The abstraction exposes only features every backend supports, so nobody uses the strongest model's capabilities.
  • Untested exit. The second provider is configured but has served no traffic for months; credentials, quotas or model names have expired.

Trade-offs

DesignBenefitCost
Single provider, multi-regionOne prompt set, deepest featuresExposed to vendor-wide incidents and pricing
Active-passive second providerAvailability with modest effortFallback rots unless exercised
Active-active splitFallback stays warm, leverageTwo prompt variants, two eval runs per change
Hosted plus self-hosted baselineLower steady-state cost, data controlGPU operations, capacity planning, model upgrades
Per-task best modelHighest quality per taskMost integration and evaluation work

What to do next

  1. Write down which two goals, from availability, capacity, cost, residency, capability and leverage, the strategy must serve.
  2. Introduce a task-named internal API so no application calls a vendor SDK directly.
  3. Build one adapter per backend, including token counting with that backend's tokenizer.
  4. Create a golden evaluation set per high-volume task and gate pool membership on it.
  5. Compute failover capacity for each fallback at full primary traffic, and request the quota or define load shedding.
  6. Implement per-backend circuit breakers with retry only before the first token.
  7. Shadow a small share of traffic to every fallback continuously and grade it offline.
  8. Run a game day: disable the primary in staging, then in production for a low-risk task, and measure what happens.
  9. Record each provider's data terms in the policy engine and test that fallbacks respect them.
Key takeaway: A multi-provider strategy is only as good as its weakest fallback. Give applications a task-named API, translate in per-backend adapters, retry only before the first token, size every fallback for full traffic, count tokens with each backend's own tokenizer, and admit a backend to a task's pool only after it passes that task's evaluations. Exercise the exit path regularly, or it will not work when needed.