A large LLM deployment produces hundreds of health-related signals per GPU: Xid events, ECC and row-remap counters, temperatures, link errors, process restarts, collective timeouts, step times, token throughput and error rates. Nobody can act on hundreds of signals. Routers, schedulers and on-call engineers need one answer per replica or node. Should it take traffic, should it run the next job, should someone repair it? A health score compresses the signals into that answer.
Done badly, a health score is worse than none. A weighted average hides a dead NVLink behind good temperatures. A score with no hysteresis flaps nodes in and out of service. A score that treats missing data as healthy keeps routing traffic to a host whose exporter crashed. This article shows how to design the scoring function from first principles for both serving replicas and training nodes, with code, a worked example and the failure modes. For the underlying telemetry, see NVIDIA DCGM in depth. For what the faults are physically, see GPU hardware faults.
Start from the decisions
Design the score backwards from the decisions it drives. A number that drives nothing is a dashboard decoration. In practice there are four consumers, and each needs something different:
| Consumer | Decision | What it needs from the score |
|---|---|---|
| Request router | How much traffic a serving replica gets | Fast updates (seconds), graded values, no flapping |
| Training scheduler | Whether a node joins the next job or restart | Conservative, explainable gating before placement |
| Remediation automation | Cordon, drain, reboot, reset, send to repair | Hard evidence, rate limits, an audit trail |
| Humans | Where to look first | The reason behind the number, not only the number |
This leads to the first design rule. A health score is a tuple, not a scalar. Emit the value, a confidence, the worst contributing sub-score and a reason code. The router reads the value, automation reads the reason, and humans read all of it.
Signals and categories
Group signals into categories that fail independently. A common set:
| Category | Example signals | Type |
|---|---|---|
| Hardware fatal | DCGM_FI_DEV_XID_ERRORS with Xid 48, 79, 94/95; DCGM_FI_DEV_ROW_REMAP_FAILURE | Hard gate |
| Hardware degrading | Rising DCGM_FI_DEV_PCIE_REPLAY_COUNTER, NVLink errors (Xid 74), DCGM_FI_DEV_GPU_TEMP near throttle | Rate or level |
| Runtime | Worker restarts, CUDA OOMs, collective timeouts, failed health probes | Rate over window |
| Performance | Per-rank compute time (excluding collective wait) vs. peers, tokens/s vs. baseline, DCGM_FI_DEV_SM_CLOCK vs. peers | Relative |
| Service | 5xx rate, time to first token p99, queue depth | SLO-relative |
Two warnings about hardware signals. Xid 13 and Xid 31 are often caused by application code, such as an out-of-bounds kernel or a bad memory access. If they gate a node, a buggy new kernel sends a fleet of healthy GPUs to repair. Score them at the job level first, and blame the node only if different jobs keep hitting it. Second, row remapping (Xid 63/64) takes effect only after a GPU reset. A pending remap is a "schedule a reset" signal, not a "node is broken" signal.
Normalising each signal
Each signal becomes a sub-score in [0, 1], where 1 is healthy. There are four normalisation shapes, and choosing the right one matters more than any weight.
- Gate. 0 if a fatal condition is present in the window, else 1. Use it for events that make correct work impossible, such as a GPU that fell off the bus (Xid 79).
- Threshold ramp. 1 below a warning level, 0 above a critical level, linear in between. Use it for temperatures, error rates and counters converted to rates. Always use rates, not raw counters: a cumulative counter only grows, so a node would decay forever.
- Peer-relative. Compare the node with peers doing the same work, using median and median absolute deviation (MAD) so that one outlier cannot shift the baseline. In synchronous training, every rank waits for the slowest at the collective, so whole-step time is the same on every rank and hides the culprit. Measure each rank's compute time up to the point it enters the collective instead. There a slow rank shows up as a large robust z-score even when its absolute numbers look fine.
- SLO-relative. For serving, score each replica's error rate and latency against the SLO budget, so that the score uses the same definition of "bad" as your burn-rate alerts.
import statistics
def ramp(x, warn, crit):
if x <= warn: return 1.0
if x >= crit: return 0.0
return 1.0 - (x - warn) / (crit - warn)
def peer_score(value, peers, worse_if="higher", k_warn=3.0, k_crit=6.0):
med = statistics.median(peers)
mad = statistics.median(abs(p - med) for p in peers) or 1e-9
z = (value - med) / (1.4826 * mad)
if worse_if == "lower": z = -z
return ramp(z, k_warn, k_crit)
Aggregation: why the average lies
The tempting choice is a weighted mean. It is usually wrong. Suppose a node scores 1.0 on six categories and 0.0 on NVLink. With equal weights it averages 0.86 and looks healthy. Yet any tensor-parallel job placed on it will hang. Health is a weakest-link property, so the aggregation should be too.
A structure that works:
- Apply hard gates first. Any gate at 0 forces the score to 0, with that gate as the reason.
- Inside a category, combine related signals with a weighted mean. Here averaging is reasonable, because the signals are noisy views of one underlying condition.
- Across categories, take the minimum. Categories fail independently, so the worst one decides.
- Report the argmin category as the reason code.
def node_health(signals, now, max_age_s=60):
reasons, cats, fresh = [], {}, 0
for s in signals:
age = now - s.ts
if age > max_age_s:
continue # stale: excluded, lowers confidence
fresh += 1
if s.kind == "gate" and s.value == 0.0:
return Health(0.0, 1.0, s.category, f"gate:{s.name}")
cats.setdefault(s.category, []).append((s.weight, s.value))
confidence = fresh / max(len(signals), 1)
if not cats:
return Health(None, 0.0, None, "no_fresh_data")
cat_scores = {k: sum(w * v for w, v in xs) / sum(w for w, _ in xs)
for k, xs in cats.items()}
worst = min(cat_scores, key=cat_scores.get)
return Health(cat_scores[worst], confidence, worst, f"min:{worst}")Some teams use a geometric mean across categories instead of the minimum. One zero still forces the result to zero, and moderate degradation in several categories compounds. That suits serving, where three slightly sick subsystems really are worse than one. For training placement, the minimum is easier to explain and is the safer default.
Missing data and confidence
The most damaging bug in health scoring is treating missing data as healthy. If an exporter dies, a network partition isolates a rack, or the metrics pipeline lags, every signal disappears. A naive scorer then sees no errors and reports 1.0. The code above excludes stale signals and reports a confidence equal to the fraction of fresh signals. When nothing is fresh it returns None, meaning "unknown", which is not zero and not one.
Consumers must handle unknown explicitly. A router should hold the last known weight for a short grace period, then drain gradually. A training scheduler should not place new work on an unknown node. Remediation must never act on unknown, because a metrics outage would otherwise cordon the whole fleet at once. Also alert on fleet-wide confidence. If confidence drops on many nodes together, the problem is your telemetry, not your GPUs.
From score to action: hysteresis and budgets
A raw score moves with every scrape. Actions should not. Put a small state machine between the score and the actions, with different entry and exit thresholds (hysteresis) and minimum dwell times.
| State | Enter when | Leave when | Action |
|---|---|---|---|
| Healthy | score >= 0.9 for 10 min | score < 0.7 for 3 scrapes | Full traffic, eligible for jobs |
| Degraded | score < 0.7 for 3 scrapes | score >= 0.9 for 10 min, or score < 0.3 | Reduced route weight, no new long jobs |
| Quarantined | gate fired, or score < 0.3 | Passes diagnostics after remediation | Drained, cordoned, diagnostics run |
| Repair | Diagnostics fail or quarantine repeats | Hardware ticket closed and burn-in passes | Out of the pool |
The numbers are illustrative; tune them on your own incident history. The structure is what matters. The exit threshold is stricter than the entry threshold, recovery requires sustained good behaviour, and a node can return from quarantine only by passing an active test, such as a DCGM diagnostic run and a short NCCL bandwidth check. Healthy-looking passive telemetry is not enough.
Finally, give remediation a budget. For example, never quarantine more than 2% of a pool per hour, and never more than one node per tensor-parallel group at once without a human. When the budget is exhausted, page someone. A correlated event such as a bad driver rollout, a cooling fault or a scorer bug should reach a human, not drain the cluster.
Worked example: a slow node in a 512-GPU job
Take a training job on 64 nodes with 8 GPUs each. Whole-step time is the same on every rank, because all of them wait at the gradient all-reduce. Per-rank compute time, measured up to entry into the collective, has a median of 1.82 s and a MAD of 0.01 s. Node 37's ranks compute for about 1.86 s. Its GPU temperatures read 4 °C higher than peers, and DCGM_FI_DEV_SM_CLOCK runs about 3% below peers. There are no Xids and no ECC events.
Scoring it: the robust z for compute time is (1.86 - 1.82) / (1.4826 × 0.01) ≈ 2.7, just under the warning level of 3, so the compute-time score is 1.0. SM clock relative to peers gives z ≈ 4.1, which ramps to a score of about 0.63. Temperature is below its warning level, so its score is 1.0. With the performance category weighting compute time and SM clock equally, performance is (1.0 + 0.63) / 2 ≈ 0.82. The minimum over categories is 0.82, with reason min:performance. The node stays Healthy, because 0.82 is above the 0.7 exit threshold, but it is flagged.
Why care about a 2% slowdown? Synchronous data parallelism runs at the speed of the slowest rank, so the whole job is about 2% slower. Across 512 GPUs that is 2% of 12,288 GPU-hours, roughly 245 GPU-hours lost per day. A day later the node's compute time drifts to 1.95 s (z ≈ 8.8) and the performance score falls to 0. That is below 0.3, so the node is quarantined and excluded at the next checkpoint restart. Diagnostics then find a clogged fan tray causing clock throttling. The scorer caught a thermal problem that no single threshold alert would have fired on, because the peer comparison made a small absolute difference stand out.
For serving, the same score drives routing directly. One simple rule sets each replica's weight to its score and holds the last known weight while confidence is low. A replica at 0.5 then receives half its share of traffic, which reduces user impact while it either recovers or crosses into quarantine. Combine this with the cache-aware and tier-aware policies in LLM routing strategies. Health should scale a replica's weight, not replace the routing policy.
Failure modes
Failure modes worth designing against:
- Averaging away a fatal signal. Use gates and a minimum across categories.
- Missing data reads as healthy. Track signal age, report confidence and act on unknown conservatively.
- Flapping. A node bounces between states every few minutes, and the churn costs more than the fault. Use hysteresis and dwell times, and count transitions per node as a metric.
- Mass remediation. A scorer bug or a bad deploy quarantines a large share of the fleet. Use remediation budgets and page a human when they are exhausted.
- Blaming hardware for software. Application-induced Xids, or a regression in a new serving build, look like node faults. Correlate with job, build and kernel version before acting on a node.
- Peer comparison with the wrong peers. Comparing a prefill-heavy replica with a decode-heavy one, or a pipeline stage with more layers against one with fewer, produces false outliers. Peers must do the same work.
- Optimising the score. If teams are judged on fleet health score, they will tune thresholds until the number looks good. Report the score alongside goodput and SLO attainment, which are harder to game.
Trade-offs
Sensitivity versus churn. Lower thresholds catch slow degradation earlier but drain more healthy nodes. Measure both the precision of quarantines (the share that diagnostics confirm) and the time from fault onset to removal. Tune against both.
One score versus several. A single fleet-wide function is simpler to operate. However, training placement and serving routing want different weights and speeds. Sharing the signal and normalisation layers while keeping a separate aggregation per consumer is a good compromise.
Rules versus learned models. A model trained on past incidents can rank nodes well, but it is harder to explain and can learn your old blind spots. Start with transparent rules and reason codes, then use a learned model as one more signal inside the performance category.
Keep the scorer observable too. Export each category score and reason code as metrics so that the KPI dashboards can show why a node is unhealthy, not only that it is.
What to do next
- List the decisions the score will drive and who consumes each one.
- Group signals into independent categories and mark which are hard gates.
- Convert counters to rates, and pick a normalisation shape per signal: gate, ramp, peer-relative or SLO-relative.
- Aggregate with gates, then a weighted mean inside each category, then the minimum across categories, and emit a reason code.
- Track signal age, report confidence and define what each consumer does when health is unknown.
- Add a state machine with hysteresis, dwell times and an active-test exit from quarantine.
- Set remediation budgets and page a human when they are exhausted.
- Review quarantine precision and detection time monthly, and retune thresholds from real incidents.