Once an AI model sits inside a business process, the process inherits the model's availability. A support desk that triages tickets with a hosted LLM stops triaging when the provider has an outage, when your account hits a rate limit, when the model version you pinned is retired, or when a runaway loop burns the month's budget by Tuesday. None of these is a data-loss disaster, so classic disaster recovery plans do not cover them, yet each can halt revenue-bearing work for hours.

AI business continuity is the discipline of keeping those processes running, at reduced quality if necessary, when an AI dependency degrades or disappears. It has a security dimension that ordinary failover lacks: the fallback path is often less tested and less guarded than the primary, and attackers can trigger failover on purpose. This article covers business impact analysis per AI capability, a degradation ladder, a gateway that implements it with circuit breakers and eval gates, the security rules for fallbacks, a worked outage, drills and a checklist. Backups and restore of your own models and data are a separate problem, covered in AI Disaster Recovery and BCP.

Why AI continuity is its own problem

Three properties make AI continuity different from keeping a database up.

The dependency is often external and opaque. You do not control the provider's capacity, release schedule or deprecation calendar. Their uptime commitment, if any, is whatever your contract says, and service credits do not answer customers.

Substitutes are not equivalent. A replica database returns the same rows. A different model given the same prompt returns different text, with different accuracy, refusal behaviour, output format and susceptibility to prompt injection. Failover changes behaviour, so every fallback has to be qualified before it is allowed to serve.

Failure is not always binary. A provider can be up but slow, up but returning malformed JSON, up but silently routed to a different model snapshot. Continuity monitoring must watch quality and latency as well as error codes.

Business impact analysis per capability

Start where any continuity programme starts: a business impact analysis, done per capability rather than per system. A capability is one thing the business asks the AI to do, such as classify a ticket, draft a reply or extract fields from an invoice. For each one, record how long the business can tolerate its loss (recovery time objective), what quality floor is acceptable in degraded mode, and what happens if it simply stops.

CapabilityImpact if lostRTODegraded floorLast resort
Ticket triageQueue grows, SLA breaches within hours15 minKeyword rules at lower accuracyHuman triage rota
Reply draftingAgents slower, no outage visible to customers4 hTemplatesAgents write by hand
Invoice extractionPayments delayed, late fees1 business daySmaller model plus human checkManual keying
Fraud pre-screenFraud exposure rises5 minFail closed: hold the transactionHold queue

The last column matters most. Some capabilities should fail closed: if the fraud screen is down, holding transactions is safer than approving them unchecked. Others should fail open to a slower human path. Deciding which, per capability, before an incident, is the core output of the analysis.

The degradation ladder

Each capability gets a degradation ladder: an ordered list of ways to serve it, from best to most robust. A typical ladder runs from the primary model, to a qualified secondary provider, to a small self-hosted model, to deterministic rules or a cache of past answers, to a human queue, and finally to switching the feature off with a clear message. Not every capability has every rung; a rung only exists for a capability if it has passed that capability's evaluation suite.

Continuity gateway and degradation ladderBusiness processticket triage, draftingcapabilityContinuity gatewaybreakers, budgets, guardsRung 0: primary modelqualified for all capabilitiesRung 1: secondary providerqualified for a subsetRung 2: small self-hostednarrow tasks onlyRung 3: rules / cachedeterministic, no generationRung 4: human queueslow but always availableEval-gate registrywhich rung may serve whatTelemetryrung served, spend, refusalsFailures move a request down the ladder; policy refusals never do.
Every request names a capability. The gateway walks down the ladder past open breakers and unqualified rungs; the human queue is the floor.

A continuity gateway

The ladder lives in one place: a gateway between applications and models. Applications never call a provider directly, which also gives you one place for keys, logging and guardrails. The core routing logic is small. Each rung has a circuit breaker that opens after repeated failures and lets a single probe through after a cooldown; each rung has a set of capabilities it is qualified for; and each rung has its own output guard.

from dataclasses import dataclass, field


@dataclass
class Breaker:
    fail_threshold: int = 5        # failures inside the window that open the circuit
    window_s: float = 30.0
    cooldown_s: float = 60.0       # how long to stay open before letting a probe through
    failures: list = field(default_factory=list)
    opened_at: float = None

    def allow(self, now):
        return self.opened_at is None or now - self.opened_at >= self.cooldown_s

    def record(self, ok, now):
        if ok:
            self.failures.clear()
            self.opened_at = None
        elif self.opened_at is not None:
            self.opened_at = now               # the half-open probe failed: stay open
        else:
            self.failures = [t for t in self.failures if now - t < self.window_s] + [now]
            if len(self.failures) >= self.fail_threshold:
                self.opened_at = now


@dataclass
class Rung:
    name: str
    call: object                   # request -> output; raises on outage, timeout, 429
    qualified: set                 # capabilities this rung passed its eval gate for
    guard: object                  # (request, output) -> True if the output may be shown


def handle(capability, request, ladder, breakers, clock):
    for rung in ladder:
        if capability not in rung.qualified:
            continue                           # never route to an unqualified fallback
        b = breakers[rung.name]
        if not b.allow(clock()):
            continue
        try:
            out = rung.call(request)
        except Exception:
            b.record(False, clock())
            continue
        b.record(True, clock())
        if not rung.guard(request, out):
            # A policy block is an answer, not an outage: do NOT try the next rung,
            # or an attacker can shop down the ladder for the weakest model.
            return {"status": "refused", "rung": rung.name}
        return {"status": "ok", "rung": rung.name, "output": out}
    return {"status": "queued_for_human", "capability": capability}

Two lines carry the security weight. The qualification check means a capability can only reach a fallback that has been tested for it. The early return on a guard block means a refusal stops the walk. Production versions add per-request timeouts, a limit on concurrent half-open probes, and a spend breaker described below; the gateway should also record which rung served each request, because degraded-mode output may need review afterwards.

Qualifying fallbacks

A fallback is qualified for a capability when it passes the same gates as the primary. That means four kinds of parity.

  • Task parity. The capability's evaluation set, scored with the same metric, must clear the degraded floor agreed in the impact analysis. Prompts usually need adapting per model; keep a prompt variant per rung under version control.
  • Safety parity. Run the prompt-injection and jailbreak suites against the fallback. Smaller and self-hosted models are frequently easier to subvert, so the gateway may need stricter input filtering on lower rungs, not looser.
  • Data parity. The fallback provider must meet the same data-processing, retention and residency terms as the primary. Failing over a European customer's data to a region it may not enter turns an availability incident into a compliance incident.
  • Format parity. Downstream code that parses JSON or tool calls must validate fallback output against the same schema; a model that answers in prose will break the pipeline in a way that looks like an outage.

Re-run qualification on a schedule and whenever either model changes, because providers update models behind stable names and a rung that qualified last quarter may not today.

Security risks of the fallback path

Continuity paths are attack surface. The threats specific to them:

  • Forced failover. An attacker who can make the primary fail, for example by flooding it to trigger rate limits or by sending inputs that crash a parser, can push traffic onto a weaker rung. Count failures per tenant as well as globally, so one noisy client cannot open the breaker for everyone. See rate limiting for LLMs.
  • Denial of wallet. A loop or abusive client that drives up token spend can exhaust the budget, which is an outage with a delay. A spend breaker that degrades non-critical capabilities when burn rate crosses a threshold protects the critical ones; LLM denial of service covers the attack side.
  • Guardrail drift. Guards wired to one provider's moderation endpoint vanish when that provider is down. Keep guards independent of the model they protect.
  • Credential sprawl. Every fallback provider is another API key. Store them in the same secret manager, rotate them, and test that they still work; a fallback with an expired key is not a fallback.

Model retirement and capacity

The most predictable continuity event is model retirement. Providers publish deprecation dates, and pinned model versions eventually stop answering. Treat each pinned version as an asset with an end-of-life date in your inventory, alert well before it, and run the replacement through the same qualification gates as a fallback. The same pipeline that qualifies fallbacks therefore doubles as your upgrade pipeline, which is the strongest argument for building it.

Capacity is the other slow-moving risk. Quotas and rate limits are per account and region; a launch that doubles traffic can hit them on a normal day. Track headroom as a metric, request increases ahead of launches, and know whether reserved or provisioned capacity is available under your contract.

Worked example: a provider outage

A support organisation runs ticket triage on a hosted model. At 09:02 error rates rise; by 09:03 the primary breaker opens after five timeouts in thirty seconds. Triage requests move to rung 1, a second provider qualified for triage at 4 points lower accuracy. Reply drafting is not qualified on rung 1, so it drops to templates, and agents see a banner saying AI drafts are paused. At 09:10 the on-call engineer confirms the provider's status page shows an incident and posts to the incident channel; no code changes are needed because the ladder was decided in advance.

At 10:40 a probe succeeds and the primary breaker closes. Afterwards the team pulls the tickets triaged on rung 1, samples them, finds two misroutes, and re-queues them. The post-incident review asks one question per rung: did it serve, at what quality, and did anything bypass a guard. Containment and communication follow the normal LLM incident response process.

Drills

A ladder that has never been exercised is a hypothesis. Schedule game days that open each breaker deliberately in a staging environment and, once confident, briefly in production for low-risk capabilities. Check that traffic moves, that guards still fire, that telemetry shows the rung served, that the human queue is actually staffed, and that the breaker closes cleanly. Expire or revoke a fallback key in staging to prove the alert fires. Record time to detect and time to degrade against each capability's RTO.

Trade-offs

Every rung costs something: integration work, a second contract, eval maintenance, and prompt variants that drift. Multi-provider failover is worth it for capabilities with short RTOs and real revenue impact; for low-impact capabilities a template and a banner are cheaper and safer. Lower rungs widen the attack surface, so add a rung only when the impact analysis demands it. And a rule-based or human rung is often the better second choice than another model: predictable, auditable, and immune to the provider's bad day.

What to do next

  1. List every business capability that calls a model and record its impact, RTO, degraded floor and whether it fails open or closed.
  2. Route all model calls through one gateway with per-rung circuit breakers and per-tenant failure counting.
  3. Build a ladder per capability and qualify each rung on task, safety, data and format parity before enabling it.
  4. Make guard blocks terminal so a refusal never falls through to a weaker model.
  5. Add a spend breaker that sheds non-critical capabilities first.
  6. Put every pinned model version and fallback key in the inventory with an expiry date.
  7. Run a game day per quarter and measure time to degrade against each RTO.
Key takeaway: AI business continuity means deciding, per capability and before any incident, how the business keeps working when a model is slow, gone, retired or unaffordable. Route every call through a gateway with circuit breakers, walk a degradation ladder of qualified rungs down to a human queue, and make policy refusals terminal so failover cannot be used to find a weaker model. Qualify fallbacks on task, safety, data and format parity, and prove the ladder works with regular drills.