An engineer tuning a serving stack watches tokens per second and kernel occupancy. A director running an LLM platform cannot act on those numbers. Their decisions are different: whether to sign a larger GPU commitment, which model to retire, which team gets the next block of capacity, and whether the platform is getting cheaper per unit of value or only bigger. Leadership KPIs exist to make those decisions with evidence, and they fail when they are either too technical to act on or too aggregated to trust.
This page builds a six-number scorecard anchored on the thing a GPU platform actually pays for, GPU-hours, shows how to roll numbers up from teams without the arithmetic lying, works through commitment coverage and a quarterly review, and attaches every KPI to a threshold, an owner and a decision. It assumes the product-level outcome metric already exists, as described in LLM North Star metric.
What makes a KPI a leadership KPI
A leadership KPI passes three tests. It moves when something the leader controls changes: headcount, budget, capacity contracts, portfolio priorities. It can be decomposed into team-level drivers, so a bad number leads to a conversation with a specific owner. And it has a threshold that triggers a named action, decided in advance. A number that fails the third test is a curiosity, and a leadership review full of curiosities becomes a status meeting.
Individual and team metrics, such as model FLOP utilisation or per-endpoint latency, are covered in individual contributor KPIs. They remain the drivers underneath; leadership sees them only when a top-line number breaches and someone drills down.
The six-number scorecard
Six numbers cover the questions a platform leader is asked. Each is defined as a ratio with an explicit numerator and denominator, because those are what get stored and summed.
| KPI | Numerator | Denominator | Question it answers |
|---|---|---|---|
| Cost per qualified task | Fully loaded GPU spend | Qualified successful tasks | Are we getting cheaper per unit of value? |
| Busy fraction of paid GPU-hours | GPU-hours with work running | GPU-hours paid for | How much capacity are we wasting? |
| Commitment coverage | Committed GPU-hours | GPU-hours paid for | Is our contract mix right? |
| SLO attainment | Request-minutes within SLO | Request-minutes served | Are users getting what we promised? |
| Model lead time (p50) | Days from approved candidate to full traffic | Per release | How fast can we ship improvements? |
| Severity-weighted incidents | Sum of incident weights | Per quarter | Is reliability improving? |
Busy fraction deserves care. A GPU allocated to a pod but idle is still paid for, so the denominator is paid hours, not allocated hours. Busy means a kernel was running, measured from per-device activity counters, which is a looser bar than useful work: a card can be busy on a badly batched job. That is why it is paired with cost per qualified task, which only improves when busy time turns into outcomes. Telemetry sources for both are laid out in LLM KPI dashboards.
Architecture
Measuring busy hours and qualified tasks
Two inputs are harder to get right than they look. Busy GPU-hours come from per-device activity sampled by the node agent, for example the GPU utilisation or graphics-engine activity field that NVIDIA DCGM exports, joined to the pod and team that held the device at each sample. Paid GPU-hours come from the billing export, not from the cluster, because the cluster cannot see reserved nodes that never joined it. The gap between the two sources is itself worth a line in the packet: paid hours with no matching device samples are capacity that nobody is even trying to use.
Qualified tasks come from request logs joined to outcome signals, using the definition owned by the North Star metric. The join must happen per request, at log time, with the team and model attached; reconstructing ownership afterwards from cluster names is where most attribution errors start. A weekly job then writes one row per team, model and week holding only sums.
-- weekly fact row: sums only, never ratios; each source aggregated at its own grain
INSERT INTO kpi_facts
SELECT g.week, g.team, g.model,
g.busy_gpu_hours, g.paid_gpu_hours, g.committed_gpu_hours, g.gpu_spend_usd,
COALESCE(r.qualified_tasks, 0) AS qualified_tasks,
COALESCE(r.good_req_minutes, 0) AS good_req_minutes,
COALESCE(r.req_minutes, 0) AS req_minutes
FROM (SELECT week, team, model,
SUM(busy_seconds) / 3600.0 AS busy_gpu_hours,
SUM(paid_seconds) / 3600.0 AS paid_gpu_hours,
SUM(CASE WHEN committed THEN paid_seconds ELSE 0 END) / 3600.0 AS committed_gpu_hours,
SUM(spend_usd) AS gpu_spend_usd
FROM gpu_device_hours GROUP BY week, team, model) g
LEFT JOIN
(SELECT week, team, model,
SUM(CASE WHEN qualified THEN 1 ELSE 0 END) AS qualified_tasks,
SUM(minutes_within_slo) AS good_req_minutes,
SUM(minutes_served) AS req_minutes
FROM request_outcomes GROUP BY week, team, model) r
USING (week, team, model);
Rolling up without lying
The most common way a leadership scorecard lies is by averaging ratios. If team A runs 1,000 GPU-hours at 90 percent busy and team B runs 100 GPU-hours at 30 percent busy, the average of the two ratios is 60 percent, but the platform is actually (900 + 30) / 1,100 = 84.5 percent busy. The fix is mechanical: store numerators and denominators, sum each separately at any level, and divide last.
from collections import defaultdict
# rows: one per (team, model, week) with raw sums, never pre-computed ratios
RATIOS = {
"busy_fraction": ("busy_gpu_hours", "paid_gpu_hours"),
"cost_per_task": ("gpu_spend_usd", "qualified_tasks"),
"slo_attainment": ("good_req_minutes", "req_minutes"),
"commit_coverage": ("committed_gpu_hours", "paid_gpu_hours"),
}
def rollup(rows, level):
sums = defaultdict(lambda: defaultdict(float))
for r in rows:
key = tuple(r[k] for k in level) # e.g. ("org",) or ("org", "team")
for num, den in RATIOS.values():
sums[key][num] += r[num]
sums[key][den] += r[den]
return {key: {name: s[num] / s[den] if s[den] else None
for name, (num, den) in RATIOS.items()}
for key, s in sums.items()}Two more rules keep the rollup honest. Shared costs such as idle headroom, platform engineers and storage are allocated by a rule written down once, for example by share of busy hours, and the unallocated remainder is shown as its own line rather than smeared invisibly. And lead time, a duration rather than a ratio, is rolled up by percentile over all releases, never by averaging team medians.
Commitment coverage and the buy decision
GPU capacity comes in tiers: reserved or committed capacity at the lowest hourly rate but paid whether used or not, on-demand capacity at a higher rate, and spot or preemptible capacity cheaper still but revocable. The commitment decision is the biggest financial lever a platform leader holds, and it reduces to one comparison. Let r be the committed rate and d the on-demand rate. A committed GPU-hour is worth buying if the probability it is used exceeds r / d. With r at 60 percent of d, any hour of baseline load that is busy more than 60 percent of the time should be committed.
In practice, sort the hourly GPU demand of the last quarter and read off the level that demand exceeds in 60 percent of hours; that level is the right commitment for the next term, adjusted for forecast growth. Commitment coverage on the scorecard then tells you whether actual demand drifted from that plan. Coverage near 100 percent with a low busy fraction means you over-committed; low coverage with high on-demand spend means you under-committed. Chargeback mechanics for passing these costs to teams are in LLM FinOps.
import numpy as np
def commitment_level(hourly_demand, committed_rate, on_demand_rate, growth=1.0):
breakeven = committed_rate / on_demand_rate # e.g. 0.6
# level L such that demand >= L in a 'breakeven' share of hours
level = np.quantile(hourly_demand, 1.0 - breakeven)
return level * growth
Worked example: a quarterly packet
Here is a quarterly packet for a platform serving three products. The figures are illustrative but internally consistent. The fleet paid for 1.20 million GPU-hours: 0.80 million committed at $2.00 and 0.40 million on-demand at $3.30, for $2.92 million of GPU spend. Telemetry shows 0.84 million busy hours, a busy fraction of 70 percent. Qualified successful tasks were 146 million, so cost per qualified task is about 2.0 cents.
| KPI | Last quarter | This quarter | Threshold | Status |
|---|---|---|---|---|
| Cost per qualified task | 2.4 cents | 2.0 cents | Falls at least 5% per quarter | Met |
| Busy fraction of paid hours | 64% | 70% | Above 65% | Met |
| Commitment coverage | 71% | 67% | Between 70% and 85% | Breached, low |
| SLO attainment | 99.3% | 98.6% | At least 99.0% | Breached |
| Model lead time p50 | 19 days | 12 days | Under 14 days | Met |
| Severity-weighted incidents | 14 | 9 | Down quarter on quarter | Met |
Two breaches lead to two decisions. Coverage fell to 67 percent because demand grew into on-demand capacity: 0.40 million on-demand hours at $1.30 above the committed rate cost $520,000 more than committed hours would have. Running the commitment calculation on the quarter's hourly demand says the baseline is now about 0.95 million hours per quarter, so the decision is to raise the commitment by roughly 0.15 million hours at the next renewal, saving around $195,000 a quarter if demand holds. The SLO breach drills down to one product whose traffic doubled after a launch; its p99 latency exceeded target during peaks because its pool was sized for the old load. Burn-rate alerting, described in SLO burn rate, should have caught it within hours, so the second decision is to fund the capacity and fix the alert, with the product's serving lead as owner.
Decision thresholds
Write the decision rules down before the quarter starts, so that a breach triggers a known response instead of a debate.
| KPI breach | Default action | Owner |
|---|---|---|
| Cost per task flat two quarters | Fund an efficiency project: quantisation, batching, model distillation | Platform director |
| Busy fraction below 65% | Consolidate pools, shrink reservations at renewal, enable autoscaling to zero | Capacity lead |
| Coverage outside 70 to 85% | Rerun commitment calculation, adjust at renewal | Capacity lead with finance |
| SLO attainment below target | Freeze non-critical launches on that pool until restored | Serving lead |
| Lead time above 14 days | Review eval and rollout pipeline bottlenecks | ML platform lead |
| Incidents rising | Reliability sprint; postmortem review of top three | Engineering managers |
Failure modes
- Averaged ratios. Team ratios averaged into an org ratio, as shown above. Always sum numerators and denominators.
- Allocated instead of paid. Busy fraction measured against allocated GPUs hides capacity that is paid for but sitting outside any pool.
- Unqualified tasks. Cost per request instead of per qualified task rewards cheap failures; a model that answers badly but quickly looks efficient.
- Gaming the busy fraction. Teams fill idle GPUs with low-value batch jobs to look busy. Report the share of busy hours attributed to a named product.
- Too many KPIs. Twenty numbers in a review means none of them drive a decision. Keep six at the top and push the rest to drill-downs.
- Definitions that drift. A change to what counts as qualified makes the trend meaningless. Version the definitions and restate history when they change.
Trade-offs
A small scorecard is easy to read but hides detail, so it must be backed by drill-downs that are trusted, which means the same fact tables at every level. Anchoring on paid GPU-hours makes waste visible but can push teams to over-share GPUs and hurt latency; SLO attainment is the counterweight, which is why both sit on the same page. Committing aggressively lowers unit cost but turns a demand drop into stranded spend, so the commitment rule should use a conservative growth factor. Finally, quarterly cadence suits contracts and portfolio choices but is too slow for reliability, so SLO attainment should also be reviewed weekly by the teams that own it.
What to do next
- Write the six KPI definitions with explicit numerators and denominators and get them agreed by finance and engineering.
- Build fact tables that store raw sums per team, model and week, never ratios.
- Measure busy fraction against paid GPU-hours, including idle reserved capacity.
- Run the commitment calculation on last quarter's hourly demand and compare it with your current contract.
- Set a threshold, default action and owner for each KPI before the next quarter starts.
- Produce one quarterly packet with the six numbers, their trends and the decisions taken.
- Version every definition and restate history whenever one changes.