Every team that runs an LLM feature on a hosted API eventually has the same afternoon: the provider returns overload errors for forty minutes, the feature goes dark, and someone asks why there was no backup. The obvious answer, calling a second provider when the first fails, is a few lines of code. Making it work is not. Two models from two providers accept different request shapes, have different context limits, follow prompts differently, stream differently, and enforce rate limits you may never have tested against.

This article treats provider failover as an engineering system: classifying errors, tracking target health, bounding everything by the caller's deadline, making fallbacks equivalent, surviving mid-stream failures and keeping the backup able to carry load. It finishes with a worked incident and a checklist.

Why LLM failover is not ordinary failover

Classic failover assumes interchangeable replicas: any healthy database replica returns the same row. LLM failover breaks that in three ways. The targets are not identical; a fallback model is a different program that can give a worse answer, a differently formatted one, or a refusal. Requests are expensive and slow, so a naive retry loop can turn a 30-second request into a two-minute one and double your bill. And failures are partial: the commonest failure is not a dead endpoint but an overloaded one that answers some requests, slowly.

So the question is to what, when, and within what budget. Think of a ladder of targets: the same model in another region (nearly identical), the same model through another platform such as a cloud marketplace endpoint (same weights, different API and quotas), and a different vendor's model (most independent, least equivalent). Climb it in order, only as far as your quality bar allows.

Classify the failure before acting

Everything starts with classification, because the right action depends on why the call failed. Read the provider's documented error types rather than guessing. Anthropic's API documentation, for example, lists 429 rate_limit_error, 500 api_error, 504 timeout_error and 529 overloaded_error, and notes two details that matter for failover. A 429 caused by reaching a monthly spend cap carries no retry-after header and keeps failing until access resumes, so retrying is pointless. And on a streaming response an error can arrive as an event after the HTTP 200, so status codes alone will not catch it.

SignalLikely causeAction
400 invalid request, 413 too largeYour request, or (per Anthropic's docs) a spend limit you setFail fast, unless the message shows a spend limit: then mark the target down and fail over
401, 403, 402 billingCredentials, permissions, billingFail over and page someone; retrying will not help
429 with retry-afterYour rate limitWait if retry-after fits the deadline; otherwise fail over
429 without retry-after (spend cap)Quota exhaustedMark the target down for a long period; fail over
500, 502, 503, 529Provider fault or overloadOne retry with jitter, then fail over
Connect or TTFT timeoutOverload, network, cold capacityFail over; count toward the breaker
Error event mid-streamProvider fault after a 200See the streaming section; never splice blindly
Content refusal or safety stopModel policyNot an outage. Do not fail over to find a model that complies

The last row matters: routing around a refusal turns failover into a jailbreak amplifier. Treat refusals as answers.

Health scoring and circuit breakers

Per-request retries cannot see that a target has been failing for five minutes, so each new request pays the timeout again. Keep a shared health view keyed by target, where a target is a (provider, model, region) triple, and put a circuit breaker on it. The breaker opens when a rolling window shows too many failures or slow first tokens, stays open for a cool-off, and then lets a trickle of probe traffic through to decide whether to close. Score on time to first token as well as errors, because overload usually shows up as latency before it shows up as 5xx.

import random, time
from collections import deque

class Breaker:
    def __init__(self, window=50, fail_ratio=0.3, slow_ttft=4.0, cooloff=30.0, close_after=5):
        self.events = deque(maxlen=window)    # True = bad outcome (error or slow TTFT)
        self.fail_ratio, self.slow_ttft, self.cooloff = fail_ratio, slow_ttft, cooloff
        self.close_after, self.is_open, self.probe_ok, self.next_probe = close_after, False, 0, 0.0

    def record(self, ok, ttft=None):
        bad = (not ok) or (ttft is not None and ttft > self.slow_ttft)
        if self.is_open:                      # only probes reach an open target
            self.probe_ok = 0 if bad else self.probe_ok + 1
            if self.probe_ok >= self.close_after:
                self.is_open = False          # K consecutive good probes close it
            return
        self.events.append(bad)
        if len(self.events) >= 10 and sum(self.events) / len(self.events) > self.fail_ratio:
            self.trip(self.cooloff)

    def trip(self, seconds):                  # also used for spend caps, auth failures
        # jitter so a fleet of clients does not re-probe in lockstep
        self.is_open, self.probe_ok = True, 0
        self.next_probe = time.monotonic() + seconds * random.uniform(0.8, 1.2)
        self.events.clear()

    def allow(self, probe_rate=0.05):
        if not self.is_open:
            return True
        return time.monotonic() >= self.next_probe and random.random() < probe_rate

In a fleet, per-process breakers are slow to notice outages and recoveries. Either centralise the router in a gateway, or share breaker state through a fast store with short expiry so one instance's evidence protects the others. The site's article on circuit breakers and backpressure for model APIs goes deeper on tuning breakers and adaptive concurrency.

A router bounded by the deadline

The router walks an ordered candidate list and carries one budget: the caller's deadline. Each attempt gets a time-to-first-token cap and whatever total time remains. Same-target retries are limited to one, with jitter, and only for errors that are likely transient. Disable the provider SDK's own automatic retries, or account for them, because SDK retries nested inside your loop multiply silently. The Anthropic SDK, for instance, retries transient failures twice by default.

class Retryable(Exception): ...
class FailOver(Exception):
    def __init__(self, down_for=0.0): self.down_for = down_for
class Fatal(Exception): ...

def complete(request, candidates, breakers, deadline_s=20.0, ttft_cap=6.0):
    start = time.monotonic()
    errors = []
    for target in candidates:                   # ordered by preference
        if not breakers[target.key].allow():
            continue
        if not target.accepts(request):          # context size, tools, modality
            continue
        for attempt in range(2):                 # at most one retry on the same target
            remaining = deadline_s - (time.monotonic() - start)
            if remaining < 1.0:
                raise TimeoutError(f"deadline exhausted after {errors}")
            try:
                t0 = time.monotonic()
                resp = target.call(target.adapt(request),
                                   ttft_timeout=min(ttft_cap, remaining),
                                   total_timeout=remaining)
                breakers[target.key].record(True, resp.ttft)
                resp.served_by = target.key       # always record who answered
                return resp
            except Retryable as e:
                breakers[target.key].record(False)
                errors.append((target.key, repr(e)))
                time.sleep(random.uniform(0.1, 0.5))
            except FailOver as e:
                breakers[target.key].record(False)
                if e.down_for:
                    breakers[target.key].trip(e.down_for)
                errors.append((target.key, repr(e)))
                break
            except Fatal:
                raise                              # caller's bug; no target will fix it
    raise RuntimeError(f"all candidates failed: {errors}")
Provider failover inside an LLM client or gatewayApplicationrequest + deadlineRouterordered candidatesAdapterprompt, tools, limitsProvider A / modelprimaryProvider A, region 2same modelProvider B / modelqualified fallbackHealth scorerbreaker per targetskip openoutcomes, latencyError classifierretry same / fail over / fail fastFailover is a routing decision per request, informed by a shared health view and bounded by the caller's deadline.
The router consults the health scorer, adapts the request per target, and classifies each failure before choosing the next step.

Equivalence: what makes a fallback usable

The adapter in the diagram is where failover projects usually stall. To make a fallback genuinely usable, check each of these against the specific models you list.

  • Prompts. Keep a prompt variant per model family and evaluate each one. A system prompt tuned for one model can produce longer, shorter or differently structured answers on another.
  • Tools and structured output. Tool schemas, tool-choice options and structured-output features differ by provider and even by model version. Validate the fallback's output against the same JSON schema and treat a schema failure as a failed attempt.
  • Context and tokens. Tokenizers differ, so the same conversation has different token counts on different models. Count on the target's tokenizer, or keep a safety margin, and have accepts() refuse targets whose window is too small.
  • Conversation state. Provider-specific content, such as signed reasoning blocks or server-side tool results, cannot be replayed to another vendor. Store conversations in your own neutral format and render them per provider.
  • Caching. Prompt caches are per provider, so a fallback starts cold.
  • Policy and data handling. The fallback must be approved for the same data classes, residency and retention terms. A fallback you are not allowed to send customer data to is not a fallback.

Qualify every fallback with the same offline evaluation suite as the primary and record a quality delta. Some features should simply not fail over across vendors. Let them fail over across regions only, or degrade to a cached or rule-based answer.

When a stream fails halfway

Streaming complicates everything because the user has already seen tokens when the failure arrives. There are three options. Before the first token, failure is invisible: retry or fail over freely. That is why a TTFT cap is the most valuable timeout you can set. After some tokens, you can restart the whole answer on the fallback and tell the client to discard what it has, which is honest but visible. Or you can ask the fallback to continue from the partial text, which looks seamless but splices two models' styles and can contradict itself. Never splice tool calls or structured output; restart them. Either way, emit a reset event the client understands, and log it.

Capacity at the backup

A backup that has never carried production traffic usually fails the moment it is asked to. The fallback provider sees your traffic jump from near zero to full volume in a second, and its rate limits, which you may have sized for testing, return 429s. Some providers also apply limits on sharp increases in usage; Anthropic's documentation warns that a sudden ramp can hit acceleration limits. Your failover then fails over to nothing.

Three remedies work. Keep the fallback warm with a steady share of real traffic, a few percent, so limits, caches and dashboards are exercised and the provider sees consistent usage. Negotiate or provision limits on the fallback sized for the share of traffic it must absorb, not just for testing. And shed load deliberately when both targets are constrained: route only priority traffic to the fallback, and degrade or queue the rest.

Worked example: a forty-minute overload

A support assistant serves 40 requests per second at peak through a primary model in one region, with a 20-second deadline per request. The candidate list is: primary model in region 1; the same model in region 2 through a second platform; a different vendor's model with a qualified prompt variant and a measured quality delta of minus 3 points on the internal suite. The fallback carries 5% of traffic every day.

At 14:02 the primary starts returning overload errors on about half of requests, with slow first tokens on the rest. Within roughly ten seconds, enough bad outcomes accumulate that the region-1 breaker opens. Requests that were already in flight follow the path in the timeline below: one jittered retry, a failed attempt on region 2 because it is overloaded too, then the fallback, which answers within the deadline. New requests skip region 1 entirely. Region 2's breaker opens a minute later. The fallback jumps from 2 to 40 requests per second; because it was warm and its limit was sized for 60% of peak, it absorbs most of the load. The gateway sheds the lowest-priority 10% (bulk summarisation) to a queue. Probes close the primary breaker at 14:41, and traffic returns once five consecutive probes succeed. (The figures are illustrative.)

One request, 20 s deadline, primary overloadedA: 529 at 0.4 sA retry: 529A region 2: TTFT timeoutB: streams the full answerBudget: at most one retry on the same target, a TTFT cap per attempt, and the remaining deadline passed down.
A single in-flight request during the incident: retry once, try the second region, then the qualified fallback, all inside one deadline.

Operating and testing it

Test failover by using it: game days that inject 529s, timeouts and mid-stream disconnects at the adapter, confirming that breakers open, the fallback answers and alerts fire. Emit per-target metrics: attempts, outcomes by class, TTFT, breaker state and share of responses served. Tag every response with served_by so evaluations, cost reports and bug reports can be split by model. Alert on the fallback share as well as on errors, because a silently degraded primary that pushes 30% of traffic to another model for a week is an undeclared quality incident.

The same ideas apply at the network layer; the site's guide to DNS failover for LLM serving covers self-hosted endpoints, multi-region LLM serving covers capacity across regions, and the AI gateway overview shows where this router lives in a shared platform.

Failure modes

Failure modeSymptomMitigation
Retry stormLatency and cost spike, provider overload worsensOne same-target retry, jitter, SDK retries disabled, breakers
Cold fallbackFallback returns 429 within seconds of failoverWarm traffic share, sized limits, load shedding
Silent quality dropComplaints rise while error rate looks fineserved_by tagging, fallback-share alert, eval deltas
FlappingTraffic oscillates between targetsJittered cool-off, half-open probes, minimum dwell time
Schema mismatchDownstream parser errors only during incidentsValidate outputs per target; game-day tests
Refusal shoppingUnsafe answers appear only via the fallbackClassify refusals as answers, not errors

What to do next

  1. List your candidates on the ladder: same model in another region, same model on another platform, different vendor. Record the quality delta for each from your eval suite.
  2. Write the error classifier from the provider's documented error types, including mid-stream error events and quota errors that should mark a target down.
  3. Add a per-target breaker that scores both errors and time to first token, and share its state across instances.
  4. Put one deadline on each request, cap TTFT per attempt, allow at most one same-target retry, and turn off or account for SDK retries.
  5. Build adapters with per-model prompts, a neutral conversation format and output schema validation; refuse targets that cannot accept the request.
  6. Decide the mid-stream policy per feature, restart or continue, and add a reset event to the client protocol.
  7. Send a steady few percent of real traffic to the fallback and size its limits for the load it must absorb.
  8. Run a game day each quarter and alert on fallback share, not just on errors.
Key takeaway: LLM provider failover is a per-request routing decision bounded by the caller's deadline. Classify each failure into retry, fail over or fail fast, keep shared per-target breakers that watch time to first token as well as errors, adapt prompts, tools and context for each qualified fallback, restart rather than splice broken structured streams, keep the backup warm with real traffic, and never treat a refusal as an outage.