A status page for an LLM API looks like a solved problem: pick a hosted product, list a few components, post updates during incidents. Then the first real incident arrives and it is not an outage at all. Time to first token has tripled in one region because a node of eight GPUs dropped off the bus, error rates are flat, and customers are already posting screenshots. Or a deployment changed a sampling default and answers got worse with every metric green. A status page that only knows up and down says operational through both, and that costs more trust than the incident did.

This article designs a status page for GPU-backed LLM serving: which components to publish, how serving and hardware signals map to a status, how to express quality degradation honestly, the machine-readable feed that SDKs and routers consume, and how to keep the page up when your own infrastructure is not. It complements synthetic monitoring, which detects problems from outside, and incident communication channels, which covers the people side of an incident; here the subject is the public status contract itself.

What the page is for

A status page has three audiences, each with a different question. Developers integrating your API ask: is the problem on my side or yours? Their own on-call engineers ask: should I fail over to another provider or model? Your support team asks: what can I tell the fifty tickets that just arrived? A page answers all three only if it is specific enough to act on, fast enough to beat social media, and honest enough that people believe it the next time.

LLM serving makes each of those harder. Performance is multidimensional: a request can succeed slowly, succeed with a truncated stream, or succeed with a worse answer. Capacity is lumpy, because a model replica spans several GPUs and losing one GPU can take a whole replica out. And the product surface is wide: chat completions, embeddings, batch jobs, fine-tuning, each with several models and regions. The design work is deciding what to collapse and what to expose.

Components customers can act on

Components should match what customers buy and can route around, not how you run the fleet. Customers cannot choose a GPU pool, so pools are not components. They can choose a model, an API surface and often a region, so those are. A workable shape is a group per API surface, with a component per model family and region inside it, plus shared dependencies such as the console, authentication and billing that can fail independently.

ComponentExampleWhy it is separate
API surfaceChat completions, Embeddings, Batch, Fine-tuningDifferent serving stacks and failure modes
Model family per regionLarge model, us-east; Small model, eu-westCustomers can switch model or region
Shared control planeAuthentication, API keys, ConsoleFailure here breaks every model at once
Asynchronous jobsBatch queue, Fine-tuning queueDelays matter, errors are rarer

Keep the count small enough to read on a phone; a few dozen components is the practical ceiling. Internally, maintain a mapping from every GPU pool, serving deployment and dependency to the components it supports. That mapping is what lets a hardware event in pool h-use1-07 light up exactly the components whose capacity it carries, and it is the first thing to update when a model moves between pools.

From GPU telemetry to a published statusDCGM / node healthXid, ECC, link stateServing metricserrors, TTFT, tok/s, 429sSynthetic probesoutside-in, per regionQuality canaryfixed eval setStatus evaluatorper component, hysteresisproposalHuman gateIC confirms wordingStatus pageseparate host + DNSJSON feedSDKs, routersSubscribersemail, webhookAuto-publish only clear-cut outagesprobes failing in 2+ regionsHardware signals feed capacity; customers see components they buy, not GPU pools.
Telemetry feeds a per-component evaluator; humans confirm most changes; the published outputs live off your main infrastructure.

Mapping signals to a status

Hosted status products converge on a similar vocabulary. Atlassian Statuspage, for example, uses operational, degraded_performance, partial_outage, major_outage and under_maintenance for components. Whatever product you use, define in writing which measured conditions put a component in each state, so the decision is not made by whoever is awake. A starting table for an interactive chat component, evaluated over five-minute windows against your own SLOs:

StatusServer error rateTTFT p95 vs SLOCapacity rejectionsQuality canary
operationalunder 0.5%withinunder 1%pass
degraded_performance0.5% to 2%1x to 2x1% to 5%regression confirmed
partial_outage2% to 20%over 2x5% to 30%severe regression
major_outageover 20%probes time outover 30%unusable output

Capacity rejections deserve their own column. When the fleet is saturated, a well-behaved server sheds load with HTTP 429 rather than queueing until requests time out. To a customer, a 429 storm is an outage even though no server failed, so count those rejections whether they are 429s or a provider-specific overload code; 529, for example, is used by some providers but is not a standard HTTP status. Exclude rejections that come from a customer exceeding their own rate limit, which are working as designed.

Raw thresholds flap. A component that crosses 2% errors for one window and drops back should not post an incident and a resolution within ten minutes. Use hysteresis: require several consecutive bad windows to worsen and more good windows to recover, and never let automation resolve a status a human set.

ORDER = ["operational", "degraded_performance", "partial_outage", "major_outage"]

def classify(w, slo_ttft_p95):
    """Worst status implied by one 5-minute window of metrics."""
    ttft = w.ttft_p95 / slo_ttft_p95
    if w.err_rate > 0.20 or w.reject_rate > 0.30 or w.probe_timeouts >= 2:
        return "major_outage"
    if w.err_rate > 0.02 or w.reject_rate > 0.05 or ttft > 2.0:
        return "partial_outage"
    if w.err_rate > 0.005 or w.reject_rate > 0.01 or ttft > 1.0 or w.canary_regressed:
        return "degraded_performance"
    return "operational"

class Evaluator:
    WORSEN_AFTER, RECOVER_AFTER = 2, 6        # windows: 10 minutes to worsen, 30 to recover

    def __init__(self, component):
        self.component, self.current, self.streak, self.pending = component, "operational", 0, None

    def step(self, window, slo_ttft_p95):
        want = classify(window, slo_ttft_p95)
        if want == self.current:
            self.streak, self.pending = 0, None
            return None
        if want != self.pending:
            self.pending, self.streak = want, 0
        self.streak += 1
        need = self.WORSEN_AFTER if ORDER.index(want) > ORDER.index(self.current) else self.RECOVER_AFTER
        if self.streak >= need:
            self.current, self.streak, self.pending = want, 0, None
            return Proposal(self.component, want, evidence=window.summary())
        return None

The evaluator emits proposals, not publications. A bot posts each proposal with its evidence into the incident channel, and the incident commander publishes it, edits the wording, or rejects it. Auto-publish only the unambiguous case: synthetic probes failing from two or more independent vantage points for the same component. That keeps the page fast for hard outages without letting a broken metric pipeline announce one.

GPU failures become capacity, not components

Hardware events are where GPU serving differs from ordinary web services, and the temptation is to publish them. Don't; translate them into capacity. NVIDIA DCGM exposes GPU health as metrics, for example DCGM_FI_DEV_XID_ERRORS for the most recent Xid error, and Xid 79 means the GPU has fallen off the bus. Uncorrectable ECC errors and NVLink failures similarly remove a GPU from service until it is reset or drained. See DCGM for collecting these signals.

Why one GPU matters: a large model served with tensor parallelism across eight GPUs is one replica per node, and the replica cannot run on seven. Take a pool of eight such nodes running at 80% of its sustainable throughput. Losing one node removes one eighth of capacity, so the remaining seven run at 0.80 / 0.875, about 91%. Queueing delay grows sharply as utilization approaches 100%, so TTFT climbs well before anything fails, and if traffic grows by ten percent the pool starts shedding load. That is why the evaluator should see a capacity headroom signal per component, computed from healthy replicas, alongside the latency it eventually causes: it lets the team post degraded_performance with an accurate explanation (reduced capacity in one region) before customers notice, or shift traffic and post nothing at all.

Quality degradation is a status

Quality regressions are the incidents LLM providers are most tempted to leave off the page, and the ones that cost most trust when customers discover them themselves. Causes include a tokenizer or chat-template change, a quantized build deployed by mistake, a sampling default altered, or a numerical bug on one hardware type that only shows up on some replicas. Errors and latency stay green throughout.

Make quality a first-class input. Run a fixed evaluation set against every production model and region on a schedule, with deterministic settings where the API allows, and compare scores to a rolling baseline with a threshold wide enough to ignore sampling noise; confirm a regression with a second run before proposing a status. When one is confirmed, post degraded_performance with plain words: what changed in outputs, which models and regions, and whether to retry or switch. "Some responses from the large model in eu-west may be lower quality; we have identified a configuration change and are rolling it back" is more useful than any amount of silence.

The machine-readable feed

Machines read status pages too: SDKs show banners, multi-provider routers weight traffic, and customers' own monitors correlate their alerts with yours. Publish a JSON feed with a stable schema. If you use Statuspage, the public /api/v2/status.json and /api/v2/summary.json endpoints already provide one, with an overall indicator of none, minor, major or critical plus per-component status. If you build your own, keep the shape similarly boring:

{
  "updated_at": "2026-10-06T09:42:00Z",
  "indicator": "minor",
  "components": [
    {"id": "chat-large-use1", "name": "Chat, large model, us-east",
     "status": "degraded_performance", "updated_at": "2026-10-06T09:40:00Z"},
    {"id": "embed-use1", "name": "Embeddings, us-east", "status": "operational",
     "updated_at": "2026-10-05T22:10:00Z"}
  ],
  "incidents": [
    {"id": "inc-2291", "status": "identified",
     "title": "Elevated latency for the large chat model in us-east",
     "components": ["chat-large-use1"], "updated_at": "2026-10-06T09:42:00Z"}
  ]
}

Tell consumers how to use it. A status feed lags reality by minutes, because humans confirm changes, so a router should treat it as a prior, not a trigger: combine it with the router's own observed error and latency per provider, and use the feed mainly to stop sending traffic back to a component that is still marked as an outage. Make component IDs permanent, since renaming one silently breaks every integration that filters on it, and offer webhooks so subscribers do not poll you into a second incident.

Keeping the page up when you are down

The status page must survive the failures it reports. Host it outside your serving infrastructure: a hosted provider or a static site on a different cloud and CDN, under a separate domain or at least separately managed DNS, so a broken DNS change or an expired certificate on your main domain does not take both down. Make sure the people who post updates can authenticate without your single sign-on, which may be part of the outage. Keep the publishing path, from evaluator proposal to page, free of dependencies on the cluster being reported on. Then test it: in a game day, block your main cloud from the publishing workstation's network and post a test update to a private page.

Worked example: one node off the bus

At 09:12 a node in pool h-use1-07 reports Xid 79 on one GPU and the scheduler drains it; healthy replica headroom for chat-large-use1 drops from 20% to about 9%. At 09:20 the evaluator sees TTFT p95 at 1.3 times SLO for two consecutive windows and proposes degraded_performance with the evidence attached. At 09:22 the incident commander publishes it, titled "Elevated latency for the large chat model in us-east", and adds that requests are succeeding and eu-west is unaffected. At 09:35 traffic is shifted, rejections stay under 1%, and an update says so. The node returns at 10:05; after six good windows the evaluator proposes operational and the commander resolves the incident with a one-paragraph summary. Total customer-facing confusion: almost none, because the page said something true within ten minutes.

Failure modes and trade-offs

Failure modeConsequenceMitigation
Page shows only up or downSlow and wrong answers read as operationalLatency, rejection and quality thresholds per status
Components mirror GPU poolsCustomers cannot act on the informationPublish model, surface and region; map pools internally
Fully automated publishingA metrics bug posts a fake outageProposals plus human gate; auto-publish only multi-vantage probe failures
No hysteresisIncidents open and close every few minutesConsecutive-window rules, slower recovery than escalation
Page hosted on the same cloudPage down during the outageSeparate provider, DNS and login path
Quality incidents omittedCustomers find out from each otherEval canary as a status input with plain wording
Feed treated as real timeRouters react minutes lateDocument lag; combine with client-side signals

The trade-off running through all of this is speed against accuracy. Automation is fast and occasionally wrong; humans are accurate and slow at 3 a.m. Burn-rate alerting, described in SLO burn rate alerts, is a useful bridge: an alert that has already paged someone is a natural moment to generate a status proposal, so the human decision happens while the engineer is already looking.

What to do next

  1. Write the component list from what customers buy, and the internal mapping from pools, deployments and dependencies to those components.
  2. Write a threshold table per component type covering errors, TTFT, capacity rejections and the quality canary, tied to your SLOs.
  3. Build the evaluator with hysteresis and have it post proposals with evidence to the incident channel.
  4. Add a capacity headroom signal per component from healthy replicas, fed by DCGM health events.
  5. Schedule a quality canary per model and region and treat confirmed regressions as status changes.
  6. Publish a stable JSON feed and webhooks, and document the lag for consumers.
  7. Move hosting, DNS and the login path off your serving infrastructure, and rehearse publishing with the main cloud blocked.
Key takeaway: An LLM status page earns trust by being specific, prompt and honest. Publish components customers can act on (surface, model, region), drive status from written thresholds for errors, latency, capacity rejections and quality, translate GPU failures into capacity rather than publishing hardware, let automation propose and humans publish except for unambiguous outages, offer a stable feed with documented lag, and host the whole thing where your outage cannot reach it.