"What was our uptime last month?" sounds like a lookup, but for an LLM endpoint it is a calculation with several judgement calls, and each one moves the answer. Does a response that took 40 seconds to produce its first token count as up? Does a minute where 20% of requests failed count as down? Is availability measured per minute, per request or by a probe? How do eight GPUs in a tensor-parallel group, twelve replicas, a gateway and a vendor's safety API combine into one number? Two engineers working from the same incident log can report 99.5% and 99.83% for the same month, and both can be right by their own definitions.
This article sets out how to calculate it: how to define a good event for an LLM endpoint, the three measurement methods and when they disagree, the composition maths for series, redundant and k-of-n systems, how to turn GPU failure data into an availability estimate, and a calculator that turns request logs into the figures you report. The worked month at the end gives three different answers and explains each.
What counts as up for an LLM endpoint
Availability is the share of good events. For a web page, "good" is usually an HTTP status below 500. For an LLM endpoint that definition is too generous: the server can return 200 and stream nothing useful for a minute. A good request must pass every check that a user would notice failing:
- It completed. No 5xx status, no dropped stream, no timeout at the client or the gateway.
- It started in time. Time to first token under a threshold, for example 10 s for interactive chat. A response that arrives after the user has given up is an outage for that user.
- It was well formed. A non-empty completion, and for structured endpoints valid JSON or a valid tool call. A model server that returns empty strings after a bad weight load is down, whatever its status codes say.
Decide in advance what to do with the ambiguous cases. Rejections with 429 because the customer exceeded their own quota are usually excluded from the count altogether. Rejections with 429 or 503 because you ran out of capacity count as bad. Client cancellations are excluded. Requests the safety layer rejected count as good, because the system worked as designed. Write these rules into the calculator, not into a wiki page.
Three ways to measure, and the nines
There are three ways to turn good events into a percentage, and they answer different questions.
| Method | Formula | Answers | Blind spot |
|---|---|---|---|
| Time-based | good minutes / total minutes; a minute is bad if its good-request ratio is below a threshold such as 99% | How long was the service broken? | Treats a 3 a.m. minute and a peak minute the same; the threshold is arbitrary |
| Request-based | good requests / valid requests | What share of user work failed? | The denominator shrinks during an outage when clients back off or users leave |
| Probe-based | successful probes / probes | Could an outside user get an answer? | Sampling error; one probe a minute sees a 20% failure rate only some of the time |
Request-based availability is usually the right primary figure for an API, because it weights by what users actually tried to do. Keep the time-based figure next to it for contracts, since many SLAs are written in minutes. Use the probe figure from outside-in synthetic monitoring as an independent check that catches what server-side logs miss, such as DNS or certificate failures that never reach your logs.
Always fix the window. A calendar month and a rolling 30 days give different figures for the same history. A 30-day window has 43,200 minutes. Here is what each target allows:
| Target | Bad minutes per 30 days | Per 365 days |
|---|---|---|
| 99% | 432 (7.2 h) | 87.6 h |
| 99.5% | 216 (3.6 h) | 43.8 h |
| 99.9% | 43.2 | 8.76 h |
| 99.95% | 21.6 | 4.38 h |
| 99.99% | 4.32 | 52.6 min |
Composition: series, redundant and k-of-n
A request passes through several components, and it succeeds only if all of them work. If failures are independent, availability in series is the product: A = A1·A2·…·An. Five components at 99.9% give 99.5%. Each dependency you add lowers the ceiling.
Redundancy works the other way. If any one of n independent copies is enough, A = 1 − (1−a)n. Two 99% providers behind a failover router give 99.99%, but only if the router detects failures and fails over instantly, which it never does in practice. Add the router as a series term and count the detection time as downtime.
Serving pools need a k-of-n model. You do not need just one replica; you need enough of them to carry the load. If at least k of n replicas must be up, each independently up with probability a, then A = Σi=k..n C(n,i)·ai·(1−a)n−i. Within a replica, a tensor-parallel group is the opposite: one GPU failing takes the whole group down, so eight GPUs are in series.
The figure contains the main lesson. Spare GPU replicas make the pool almost perfect on paper, and the four single-instance stages around it determine the result. Before buying more GPUs for reliability, check whether a single-instance dependency is what limits you.
From GPU failure rates to availability
Steady-state availability of a repairable component is A = MTBF / (MTBF + MTTR), where MTBF is the mean time between failures and MTTR the mean time to restore. Two consequences follow. Halving MTTR helps as much as doubling MTBF, and MTTR is usually cheaper to improve. And fleet failure rates scale with size: N components that each fail at rate λ produce failures at rate Nλ.
For a real per-GPU rate, the Llama 3 paper reports 419 unexpected interruptions during a 54-day pre-training snapshot on a cluster of up to 16K H100 GPUs, about 78% of them attributed to confirmed or suspected hardware. That is one interruption roughly every 3.1 hours for the cluster. Assuming all 16,384 GPUs ran throughout, per GPU it is 419 / (16,384 × 54 × 24) ≈ 1.97×10−5 per GPU-hour, or one per about 50,700 GPU-hours. Training loads GPUs harder and more synchronously than serving, and the figure includes host faults, so treat it as a planning number, not a law.
Apply it to a serving replica of 8 GPUs in series. The rate is 8 × 1.97×10−5 ≈ 1.58×10−4 per hour, an MTBF of about 6,300 hours. With 30 minutes to drain the replica and bring up a spare node, a = 6,300 / 6,300.5 ≈ 0.99992. For 12 replicas needing 10, the pool is unavailable only when 3 or more are down at once: about C(12,3)·(7.9×10−5)3 ≈ 10−10. Independent GPU failures barely affect the availability of a pool with spare replicas.
What does affect it are correlated failures: a driver or engine version rolled to every node, a shared top-of-rack switch, a bad model artefact, a capacity crunch at peak. For those, k-of-n maths does not apply, because all replicas fail together. Model each correlated cause as its own series term with a measured frequency and duration. The GPU hardware faults article covers how individual GPU failures show up, and LLM serving reliability covers how to contain them.
A calculator from request logs
The calculator below reads one record per request and produces all three figures plus the composition helpers. The good-event rules sit in one function, so the definition is reviewed like any other code.
import math
from collections import defaultdict
from dataclasses import dataclass
@dataclass
class Req:
minute: int # minutes since window start
status: int
ttft_s: float | None
tokens_out: int
reason: str = "" # "client_cancel", "tenant_quota", ...
def classify(r, ttft_slo=10.0):
# returns "good", "bad" or None (excluded from the denominator)
if r.reason in ("client_cancel", "tenant_quota"):
return None
if r.status >= 500 or r.status in (429, 503): # our capacity or our fault
return "bad"
if r.ttft_s is None or r.ttft_s > ttft_slo or r.tokens_out == 0:
return "bad"
return "good"
def availability(reqs, window_minutes, minute_threshold=0.99, excluded=frozenset()):
good = total = 0
per_min = defaultdict(lambda: [0, 0])
for r in reqs:
if r.minute in excluded:
continue
k = classify(r)
if k is None:
continue
total += 1
good += k == "good"
per_min[r.minute][0] += k == "good"
per_min[r.minute][1] += 1
counted = window_minutes - len(excluded)
bad_minutes = sum(1 for g, n in per_min.values() if g / n < minute_threshold)
return {"request_based": good / total if total else None,
"time_based": 1 - bad_minutes / counted,
"requests": total}
def series(*a):
return math.prod(a)
def parallel(*a):
return 1 - math.prod(1 - x for x in a)
def k_of_n(k, n, a):
return sum(math.comb(n, i) * a**i * (1 - a)**(n - i) for i in range(k, n + 1))
pool = k_of_n(10, 12, 6300 / 6300.5)
print(series(0.9999, 0.9995, 0.9999, pool, 0.999)) # about 0.99830Note that a minute with no traffic doesn't count as bad here. That is right for a busy endpoint. For a quiet one, use probe results for minutes without user traffic, otherwise a dead endpoint nobody called scores 100%.
Worked example: one month, three answers
Take one 30-day month on an endpoint serving a steady 1,000 requests per minute, 43.2 million valid requests in total. Two things went wrong:
- A bad engine roll-out took every replica down for 38 minutes: 38,000 failed requests.
- A KV-cache memory leak caused 3 hours of preemption, during which 20% of requests missed the 10-second TTFT threshold: 36,000 bad requests over 180 minutes.
| Method | Calculation | Result | Meets 99.9%? |
|---|---|---|---|
| Time-based (minute bad below 99% good) | 1 − (38 + 180) / 43,200 | 99.50% | No |
| Request-based | 1 − 74,000 / 43,200,000 | 99.83% | No |
| Probe-based, one probe per minute | expected 38 + 0.2 × 180 = 74 failed of 43,200 | ≈ 99.83%, but noisy | No |
| Request-based, status codes only | 1 − 38,000 / 43,200,000 | 99.91% | Yes, misleadingly |
Each row is correct by its own definition. The time-based figure counts the partial degradation as fully down for three hours. The request-based figure weights it by its actual 20% impact. The probe estimate matches on average, but with 180 probes during the leak, the number that failed could plausibly be anywhere from about 25 to 47. The last row shows the cost of a weak good-event definition: ignore TTFT and the leak disappears, and the month passes a target that users would say it missed.
Two things change the figure further. If the outage had hit at a 2x traffic peak, the request-based figure would fall to 99.74% while the time-based figure stayed the same. And if clients stopped retrying during the outage, the request denominator would shrink and flatter the result. Count client-side attempts, or estimate the missing traffic from the same hour last week, so lost demand still counts.
Reporting the number
- Publish the definition with the number: window, method, good-event rules and exclusions. "99.83% request-based, 30-day rolling, TTFT ≤ 10 s" can be audited; "99.8% uptime" cannot.
- Exclude maintenance only if it was announced in advance and only for minutes when traffic was actually refused. A graceful drain that served every request is not downtime and does not need excluding.
- Attribute every bad minute to a cause (deploy, hardware, capacity, dependency) so the error budget can be spent and argued about with data.
- Don't average percentages across regions or models. Sum good and total events, then divide. Otherwise a tiny region at 90% counts as much as the main one.
- Account for vendor dependencies. If a hosted model or moderation API sits in your path, its SLA caps your own: you can't promise 99.95% on top of a 99.9% dependency without a fallback.
Failure modes
- Denominator collapse: during a full outage, clients back off, so the request-based figure under-counts the damage.
- Retry inflation: one user action retried five times becomes five bad events, or five good ones after recovery. Count by idempotency key or user action where you can.
- Health checks that lie: a
/healthendpoint that returns 200 while the engine is wedged produces a perfect probe score. - Assumed independence: k-of-n figures of 99.9999% fail in practice because replicas share deploys, configuration and power.
- Silent definition changes: raising the TTFT threshold mid-quarter makes the trend look like an improvement.
What to do next
- Write the good-event rules as code, including TTFT, empty output and the 429 split, and have them reviewed.
- Compute request-based and time-based availability over the same window from logs, and probe-based availability from synthetic checks; investigate any gap larger than 0.05 points.
- Draw your request path, mark every series stage with a measured availability, and find the stage that sets your ceiling.
- Model the replica pool with k-of-n, then list the correlated causes it ignores and give each one a series term.
- Measure MTTR for a lost replica and for a bad deploy, and set a target for each.
- Publish the figure with its definition, window and exclusions, and attribute every bad minute to a cause.
Related reading: error budgets for LLM serving, burn-rate alerts on TTFT and inter-token SLIs, failure domains and truthful health checks, failover between LLM providers and synthetic monitoring.