An engineer tuning a serving stack watches tokens per second and kernel occupancy. A director running an LLM platform cannot act on those numbers. Their decisions are different: whether to sign a larger GPU commitment, which model to retire, which team gets the next block of capacity, and whether the platform is getting cheaper per unit of value or only bigger. Leadership KPIs exist to make those decisions with evidence, and they fail when they are either too technical to act on or too aggregated to trust.

This page builds a six-number scorecard anchored on the thing a GPU platform actually pays for, GPU-hours, shows how to roll numbers up from teams without the arithmetic lying, works through commitment coverage and a quarterly review, and attaches every KPI to a threshold, an owner and a decision. It assumes the product-level outcome metric already exists, as described in LLM North Star metric.

What makes a KPI a leadership KPI

A leadership KPI passes three tests. It moves when something the leader controls changes: headcount, budget, capacity contracts, portfolio priorities. It can be decomposed into team-level drivers, so a bad number leads to a conversation with a specific owner. And it has a threshold that triggers a named action, decided in advance. A number that fails the third test is a curiosity, and a leadership review full of curiosities becomes a status meeting.

Individual and team metrics, such as model FLOP utilisation or per-endpoint latency, are covered in individual contributor KPIs. They remain the drivers underneath; leadership sees them only when a top-line number breaches and someone drills down.

The six-number scorecard

Six numbers cover the questions a platform leader is asked. Each is defined as a ratio with an explicit numerator and denominator, because those are what get stored and summed.

KPINumeratorDenominatorQuestion it answers
Cost per qualified taskFully loaded GPU spendQualified successful tasksAre we getting cheaper per unit of value?
Busy fraction of paid GPU-hoursGPU-hours with work runningGPU-hours paid forHow much capacity are we wasting?
Commitment coverageCommitted GPU-hoursGPU-hours paid forIs our contract mix right?
SLO attainmentRequest-minutes within SLORequest-minutes servedAre users getting what we promised?
Model lead time (p50)Days from approved candidate to full trafficPer releaseHow fast can we ship improvements?
Severity-weighted incidentsSum of incident weightsPer quarterIs reliability improving?

Busy fraction deserves care. A GPU allocated to a pod but idle is still paid for, so the denominator is paid hours, not allocated hours. Busy means a kernel was running, measured from per-device activity counters, which is a looser bar than useful work: a card can be busy on a badly batched job. That is why it is paired with cost per qualified task, which only improves when busy time turns into outcomes. Telemetry sources for both are laid out in LLM KPI dashboards.

Architecture

From GPU telemetry to a scorecard that drives decisionsGPU telemetrybusy time per deviceRequest logstasks, outcomes, SLOBilling and contractspaid hours, ratesRelease recordsmodel lead timeFact tablesnumerators and denominatorsRollupratio of sumsScorecardsix numbersThresholdsowner, actionDecisionsbuy, cut, retireLeaders read ratios; the pipeline stores numerators and denominators so ratios can be recombined at any level.
The scorecard pipeline. Raw sources feed fact tables of numerators and denominators; ratios are computed only at read time, at whatever level of the organisation is asking.

Measuring busy hours and qualified tasks

Two inputs are harder to get right than they look. Busy GPU-hours come from per-device activity sampled by the node agent, for example the GPU utilisation or graphics-engine activity field that NVIDIA DCGM exports, joined to the pod and team that held the device at each sample. Paid GPU-hours come from the billing export, not from the cluster, because the cluster cannot see reserved nodes that never joined it. The gap between the two sources is itself worth a line in the packet: paid hours with no matching device samples are capacity that nobody is even trying to use.

Qualified tasks come from request logs joined to outcome signals, using the definition owned by the North Star metric. The join must happen per request, at log time, with the team and model attached; reconstructing ownership afterwards from cluster names is where most attribution errors start. A weekly job then writes one row per team, model and week holding only sums.

-- weekly fact row: sums only, never ratios; each source aggregated at its own grain
INSERT INTO kpi_facts
SELECT g.week, g.team, g.model,
       g.busy_gpu_hours, g.paid_gpu_hours, g.committed_gpu_hours, g.gpu_spend_usd,
       COALESCE(r.qualified_tasks, 0) AS qualified_tasks,
       COALESCE(r.good_req_minutes, 0) AS good_req_minutes,
       COALESCE(r.req_minutes, 0)      AS req_minutes
FROM (SELECT week, team, model,
             SUM(busy_seconds) / 3600.0 AS busy_gpu_hours,
             SUM(paid_seconds) / 3600.0 AS paid_gpu_hours,
             SUM(CASE WHEN committed THEN paid_seconds ELSE 0 END) / 3600.0 AS committed_gpu_hours,
             SUM(spend_usd)             AS gpu_spend_usd
      FROM gpu_device_hours GROUP BY week, team, model) g
LEFT JOIN
     (SELECT week, team, model,
             SUM(CASE WHEN qualified THEN 1 ELSE 0 END) AS qualified_tasks,
             SUM(minutes_within_slo) AS good_req_minutes,
             SUM(minutes_served)     AS req_minutes
      FROM request_outcomes GROUP BY week, team, model) r
  USING (week, team, model);

Rolling up without lying

The most common way a leadership scorecard lies is by averaging ratios. If team A runs 1,000 GPU-hours at 90 percent busy and team B runs 100 GPU-hours at 30 percent busy, the average of the two ratios is 60 percent, but the platform is actually (900 + 30) / 1,100 = 84.5 percent busy. The fix is mechanical: store numerators and denominators, sum each separately at any level, and divide last.

from collections import defaultdict

# rows: one per (team, model, week) with raw sums, never pre-computed ratios
RATIOS = {
    "busy_fraction":     ("busy_gpu_hours", "paid_gpu_hours"),
    "cost_per_task":     ("gpu_spend_usd",  "qualified_tasks"),
    "slo_attainment":    ("good_req_minutes", "req_minutes"),
    "commit_coverage":   ("committed_gpu_hours", "paid_gpu_hours"),
}

def rollup(rows, level):
    sums = defaultdict(lambda: defaultdict(float))
    for r in rows:
        key = tuple(r[k] for k in level)          # e.g. ("org",) or ("org", "team")
        for num, den in RATIOS.values():
            sums[key][num] += r[num]
            sums[key][den] += r[den]
    return {key: {name: s[num] / s[den] if s[den] else None
                  for name, (num, den) in RATIOS.items()}
            for key, s in sums.items()}

Two more rules keep the rollup honest. Shared costs such as idle headroom, platform engineers and storage are allocated by a rule written down once, for example by share of busy hours, and the unallocated remainder is shown as its own line rather than smeared invisibly. And lead time, a duration rather than a ratio, is rolled up by percentile over all releases, never by averaging team medians.

Commitment coverage and the buy decision

GPU capacity comes in tiers: reserved or committed capacity at the lowest hourly rate but paid whether used or not, on-demand capacity at a higher rate, and spot or preemptible capacity cheaper still but revocable. The commitment decision is the biggest financial lever a platform leader holds, and it reduces to one comparison. Let r be the committed rate and d the on-demand rate. A committed GPU-hour is worth buying if the probability it is used exceeds r / d. With r at 60 percent of d, any hour of baseline load that is busy more than 60 percent of the time should be committed.

In practice, sort the hourly GPU demand of the last quarter and read off the level that demand exceeds in 60 percent of hours; that level is the right commitment for the next term, adjusted for forecast growth. Commitment coverage on the scorecard then tells you whether actual demand drifted from that plan. Coverage near 100 percent with a low busy fraction means you over-committed; low coverage with high on-demand spend means you under-committed. Chargeback mechanics for passing these costs to teams are in LLM FinOps.

import numpy as np

def commitment_level(hourly_demand, committed_rate, on_demand_rate, growth=1.0):
    breakeven = committed_rate / on_demand_rate         # e.g. 0.6
    # level L such that demand >= L in a 'breakeven' share of hours
    level = np.quantile(hourly_demand, 1.0 - breakeven)
    return level * growth

Worked example: a quarterly packet

Here is a quarterly packet for a platform serving three products. The figures are illustrative but internally consistent. The fleet paid for 1.20 million GPU-hours: 0.80 million committed at $2.00 and 0.40 million on-demand at $3.30, for $2.92 million of GPU spend. Telemetry shows 0.84 million busy hours, a busy fraction of 70 percent. Qualified successful tasks were 146 million, so cost per qualified task is about 2.0 cents.

KPILast quarterThis quarterThresholdStatus
Cost per qualified task2.4 cents2.0 centsFalls at least 5% per quarterMet
Busy fraction of paid hours64%70%Above 65%Met
Commitment coverage71%67%Between 70% and 85%Breached, low
SLO attainment99.3%98.6%At least 99.0%Breached
Model lead time p5019 days12 daysUnder 14 daysMet
Severity-weighted incidents149Down quarter on quarterMet

Two breaches lead to two decisions. Coverage fell to 67 percent because demand grew into on-demand capacity: 0.40 million on-demand hours at $1.30 above the committed rate cost $520,000 more than committed hours would have. Running the commitment calculation on the quarter's hourly demand says the baseline is now about 0.95 million hours per quarter, so the decision is to raise the commitment by roughly 0.15 million hours at the next renewal, saving around $195,000 a quarter if demand holds. The SLO breach drills down to one product whose traffic doubled after a launch; its p99 latency exceeded target during peaks because its pool was sized for the old load. Burn-rate alerting, described in SLO burn rate, should have caught it within hours, so the second decision is to fund the capacity and fix the alert, with the product's serving lead as owner.

Decision thresholds

Write the decision rules down before the quarter starts, so that a breach triggers a known response instead of a debate.

KPI breachDefault actionOwner
Cost per task flat two quartersFund an efficiency project: quantisation, batching, model distillationPlatform director
Busy fraction below 65%Consolidate pools, shrink reservations at renewal, enable autoscaling to zeroCapacity lead
Coverage outside 70 to 85%Rerun commitment calculation, adjust at renewalCapacity lead with finance
SLO attainment below targetFreeze non-critical launches on that pool until restoredServing lead
Lead time above 14 daysReview eval and rollout pipeline bottlenecksML platform lead
Incidents risingReliability sprint; postmortem review of top threeEngineering managers

Failure modes

  • Averaged ratios. Team ratios averaged into an org ratio, as shown above. Always sum numerators and denominators.
  • Allocated instead of paid. Busy fraction measured against allocated GPUs hides capacity that is paid for but sitting outside any pool.
  • Unqualified tasks. Cost per request instead of per qualified task rewards cheap failures; a model that answers badly but quickly looks efficient.
  • Gaming the busy fraction. Teams fill idle GPUs with low-value batch jobs to look busy. Report the share of busy hours attributed to a named product.
  • Too many KPIs. Twenty numbers in a review means none of them drive a decision. Keep six at the top and push the rest to drill-downs.
  • Definitions that drift. A change to what counts as qualified makes the trend meaningless. Version the definitions and restate history when they change.

Trade-offs

A small scorecard is easy to read but hides detail, so it must be backed by drill-downs that are trusted, which means the same fact tables at every level. Anchoring on paid GPU-hours makes waste visible but can push teams to over-share GPUs and hurt latency; SLO attainment is the counterweight, which is why both sit on the same page. Committing aggressively lowers unit cost but turns a demand drop into stranded spend, so the commitment rule should use a conservative growth factor. Finally, quarterly cadence suits contracts and portfolio choices but is too slow for reliability, so SLO attainment should also be reviewed weekly by the teams that own it.

What to do next

  1. Write the six KPI definitions with explicit numerators and denominators and get them agreed by finance and engineering.
  2. Build fact tables that store raw sums per team, model and week, never ratios.
  3. Measure busy fraction against paid GPU-hours, including idle reserved capacity.
  4. Run the commitment calculation on last quarter's hourly demand and compare it with your current contract.
  5. Set a threshold, default action and owner for each KPI before the next quarter starts.
  6. Produce one quarterly packet with the six numbers, their trends and the decisions taken.
  7. Version every definition and restate history whenever one changes.
Key takeaway: Leadership KPIs for an LLM platform should be few, anchored on paid GPU-hours and qualified outcomes, stored as numerators and denominators so they roll up correctly, and tied in advance to thresholds, actions and owners. Commitment coverage and busy fraction expose the largest financial levers, while SLO attainment keeps cost cutting honest.