Adding a second model provider looks like a one-line change: catch the exception, call the other API. Teams that have lived through a provider incident know it is not. The fallback call fails because the tool schema uses a different format, or it succeeds and returns JSON that the parser rejects, or it works for five minutes and then hits the secondary's rate limit because the whole fleet moved there at once. Sometimes it works perfectly and quietly produces answers that are worse in ways nobody measured.

This article is about making a second route genuinely substitutable and making degradation a deliberate product behaviour rather than an accident. It assumes you already have, or will build, a shared gateway for credentials and quotas, which is covered in LLM Gateway Architecture and AI Gateway Overview. The focus here is the routing decision and what happens on each rung below the primary.

Advertisement

Why a second provider is not a drop-in

Model APIs look similar and differ in details that matter. Tool and function definitions are expressed in different schemas. Support for constrained or schema-validated output varies, and so does the way a model signals that it wants to call a tool. System instructions may be a separate field or a message role. Tokenizers differ, so a prompt that fits a 128k window on one model can exceed another model's window even when both advertise similar numbers. Safety behaviour differs, so a prompt one model answers may be refused by another. Sampling parameters that exist on one API may be absent or scaled differently on another.

None of this is visible until the fallback runs, which is usually during an incident. The design goal is to discover every one of these differences in advance and hide them behind an internal interface: your application sends a provider-neutral request, and per-route adapters translate it and translate the response back into a neutral shape that is validated before anyone uses it.

Routing: filter what is allowed, rank what is healthy, degrade on purposeRequesttask, tier, deadlineEligibility filterresidency, contract, ctxRankerhealth, cost, latencyallowed routesAttempt loopdeadline budgetorderedHealth monitorbreaker per routeAdaptersrequest + responseProvider AprimaryProvider BsecondarySelf-hostedsmall modelerrors, TTFTDegradation ladderL0 full / L1 alt model / L2 small / L3 cache / L4 honest failureall routes failedEval harnessevery rung scored against your own testsrung thresholdsResponse envelope: route used, degradation level, output validated against schemaproduct surfaces the mode instead of hiding it
The routing path. Eligibility is a hard filter, ranking uses live health, and adapters translate in both directions. When no route can serve the request within its deadline, the degradation ladder chooses an explicit reduced mode.

Eligibility before preference

The first step is not choosing the best route but removing the routes that must not be used. Eligibility is a hard filter evaluated per request:

  • Data residency and contract terms. A request carrying personal data from a region with residency obligations may only go to routes in that region. A tenant whose contract requires no data retention excludes routes without that agreement. These rules come from legal and procurement, not from engineering preference, and failing over in violation of them turns an outage into a compliance incident.
  • Capabilities. Tool calling, image input, structured output, a minimum context window. Record them as flags on each route and require them per request type.
  • Context length measured with the target tokenizer. Count tokens for each candidate, add the maximum output length and a safety margin, and exclude routes the request does not fit.
  • Quality floor. Routes whose eval scores for this task type are below the floor are excluded, even if they are healthy. The section on testing equivalence, below, describes how the scores are produced.

If eligibility removes every route, that is a design decision to surface, not a runtime surprise: some requests have exactly one legal destination, and for those the plan during an outage is degradation, not failover.

Advertisement

Classify errors before reacting to them

Different failures call for different responses, and treating them all as "try the next provider" wastes capacity and hides bugs.

SignalMeaningRouter reaction
HTTP 429, optionally with Retry-AfterQuota or rate limit on this routeDo not retry the same route before Retry-After; try another eligible route
HTTP 5xx, connection reset, timeoutRoute unhealthy or overloadedCount toward the breaker; fail over within the deadline
HTTP 400 or schema rejectionBug in our request or adapterDo not fail over; alert, because every route may reject it
Refusal on a benign promptModel policy differenceRecord per route; consider next route only for that task type
Valid response failing our schemaAdapter or model output problemOne repair attempt, then next route; count per route
Slow time to first tokenDegradation before errors appearFeed latency health; lower route rank

The 400 row is the one teams get wrong. Failing over on a malformed request moves the bug to the secondary, burns its quota, and makes the dashboard look like a provider incident when the fault is yours.

Health, breakers and the attempt loop

Health is tracked per route, meaning provider, model and region together, because incidents are often scoped to one model or one region. Each route keeps a short rolling error rate and a time-to-first-token percentile. A circuit breaker opens when the error rate over the last window exceeds a threshold, sends no traffic while open, and after a cool-down lets a small number of probe requests through before closing. The attempt loop then walks the ranked, eligible routes within a single deadline budget.

import time

class Route:
    def __init__(self, name, adapter, breaker, cost_per_1k, caps):
        self.name, self.adapter, self.breaker = name, adapter, breaker
        self.cost_per_1k, self.caps = cost_per_1k, caps

def route_request(req, routes, deadline_s, health):
    start = time.monotonic()
    eligible = [r for r in routes
                if req.required_features <= r.caps["features"]      # sets
                and req.region in r.caps["regions"]
                and req.data_terms <= r.caps["terms"]                # e.g. {"no_retention"}
                and r.caps["eval_score"][req.task_type] >= req.quality_floor
                and r.adapter.count_tokens(req) + req.max_output <= r.caps["context"]]
    ranked = sorted((r for r in eligible if r.breaker.allows()),
                    key=lambda r: (health.p95_ttft(r.name), r.cost_per_1k))
    errors = []
    for r in ranked:
        remaining = deadline_s - (time.monotonic() - start)
        if remaining < health.min_useful_time(r.name):
            break                              # not enough time left to succeed
        try:
            raw = r.adapter.call(req, timeout=remaining)
            out = r.adapter.parse(raw)         # provider-neutral, schema-validated
            r.breaker.record_success()
            return out, r.name
        except RateLimited as e:
            health.cooldown(r.name, e.retry_after)
        except BadRequest:
            raise                              # our bug: do not fail over
        except (Unavailable, Timeout, InvalidOutput) as e:
            r.breaker.record_failure()
            errors.append((r.name, type(e).__name__))
    raise AllRoutesFailed(errors)

Two details are easy to miss. The loop stops when the remaining budget is shorter than the time a route realistically needs, because starting a call you know will time out only burns quota. And the whole loop has one deadline derived from the user-facing budget; per-attempt timeouts are whatever is left, so failover never turns a ten-second request into a forty-second one.

Hedging and retries

Retrying the same route once with jitter is reasonable for a transient reset. Hedging, sending a second request to another route if the first has not produced a first token by roughly the p95 time, cuts tail latency for short, side-effect-free generations at the cost of paying for some duplicate work. Cancel the loser as soon as the winner streams its first token. Never hedge calls that trigger tool execution with side effects, and cap hedging as a fraction of traffic so that a slow primary does not silently double your spend.

Failure in the middle of a stream

Streaming changes the failure model. If the connection drops after the user has seen half an answer, you have three options. You can restart on another route and replace the partial text, which is simple and honest but visible. You can ask another route to continue from the partial output, which works for prose but risks inconsistency in structured output and code. Or you can show the partial answer with an error and a retry control. For structured outputs and tool calls, the only safe choice is to discard the partial result and restart, because a half-received tool call must never be executed. Decide the policy per response type and make the UI match it.

Failover moves load, not just requests

When the primary fails, all of its traffic arrives at the secondary at once. If the secondary quota was sized for its normal share, it saturates within minutes, and the router now has two unhealthy routes. Size for it deliberately. Suppose peak traffic is 60 requests per second at about 1,500 tokens each, roughly 5.4 million tokens per minute, and the secondary account carries 20 percent in normal operation with a quota of 2 million tokens per minute. A full failover needs 2.7 times that quota.

The options are to reserve more quota on the secondary, accepting the standing cost; to spread traffic across more than two routes so that each carries a smaller share of the surge; or to decide in advance which traffic is shed. Assign each request a priority tier, and when capacity is short, serve interactive paid traffic first and defer batch summarisation, background enrichment and evaluation jobs. Shedding by design is graceful degradation; shedding by whichever request happens to hit the limit is not.

The degradation ladder as product modes

When routes fail, the system should step down a ladder of explicit modes, and the product should know which mode it is in.

LevelModeUser experience
L0Primary route, full featuresNormal
L1Equivalent model on another providerSame features; small quality or latency differences
L2Smaller or self-hosted modelShorter answers; complex tasks disabled or queued
L3Cached or retrieval-only answersAnswers from known content; generation paused
L4Honest failureClear message, retry later, human channel offered

Return the level in the response envelope so the interface can adjust: hide features that the lower rung cannot support, label reduced answers, and avoid presenting an L2 answer to a hard question with L0 confidence. Silent fallback to a weaker model is the dangerous default for the same reasons described for tools in Capability Fallback Chains.

Test equivalence before you need it

A fallback that has never been scored on your prompts is a guess. Run your evaluation set against every rung on a schedule and on every prompt or adapter change, and store per-route, per-task scores; those scores drive the quality floor in the eligibility filter. The mechanics of such a harness are covered in LLM Evaluation Harness and Regression Testing. Track schema-validity rate and refusal rate separately from answer quality, because adapter bugs show up there first.

Then exercise the path in production. Send a small, steady share of real traffic, around one percent, through each secondary route so that credentials, quotas and adapters are proven continuously. Run failover drills by opening the primary breaker deliberately during business hours and watching latency, error rate and degradation level. A route that only receives traffic during incidents will fail during incidents.

Worked example: a regional incident

A customer support assistant runs on provider A in the EU region, with provider B in the EU as L1, a self-hosted small model as L2 and a retrieval-only FAQ as L3. At 10:02, A's EU endpoint begins returning 503s. Within 30 seconds the breaker for A-EU opens. Eligibility excludes A-US for EU tenants because of residency terms. B-EU becomes rank one and absorbs interactive traffic; the router sheds the nightly ticket-summarisation backlog to keep B within quota. Latency rises from a p95 of 2.1 to 2.8 seconds, and the quality score on the running sample drops slightly, within the L1 threshold measured last week.

At 10:20, B-EU starts returning 429s for a subset of requests. The router moves overflow to L2, and the interface disables the refund-dispute workflow, which the small model failed in evaluation. At 10:45 A recovers; half-open probes succeed and traffic returns gradually. The post-incident report reads from the envelope logs: share of requests per level, time in each mode, and whether any residency rule was violated. None was, because the filter ran before the ranker.

What to do next

  1. Define a provider-neutral request and response shape, and write one adapter per route with schema validation on output.
  2. Write down eligibility rules for residency, contract terms, capabilities and context length, and enforce them before ranking.
  3. Classify errors into rate limit, unavailable, bad request, refusal and invalid output, and react to each differently.
  4. Add per-route breakers and a single deadline budget for the attempt loop.
  5. Decide the mid-stream failure policy per response type and never execute a partial tool call.
  6. Size secondary quota for a full failover or define priority tiers for shedding.
  7. Implement the degradation ladder and return the level in every response.
  8. Score every rung against your eval set, route a small share of live traffic through secondaries, and run failover drills.
Key takeaway: A second provider is a fallback only if it has been made substitutable and proven. Filter routes on residency, contract, capability and tokenized context length before ranking on health; classify errors so that your own bugs do not trigger failover; bound every attempt by one deadline; plan for mid-stream failure and for the load surge on the secondary; step down an explicit degradation ladder that the product can see; and keep every rung scored and exercised before the incident arrives.