A team that serves a large language model on its own GPUs can watch hundreds of numbers: tokens per second, time to first token, KV cache occupancy, eval scores, daily active users, cost per million tokens. Every one of them is useful to somebody, and none of them answers the question a planning meeting actually asks: is the product delivering more value this week than last week, and is the GPU fleet helping or hurting that? A North Star metric is the single number chosen to answer that question, agreed in advance, defined precisely enough that two engineers computing it from the same logs get the same answer.
This article builds one for a GPU-served LLM product from first principles. It covers what makes a good North Star, why the obvious candidates such as tokens served fail, a precise definition based on qualified successful tasks, the input tree that connects the headline to model, serving and capacity work, the code to compute it from request logs, a worked example of judging an FP8 rollout with it, and the counter-metrics that stop people gaming it. The metric tree for dashboards and the per-engineer KPIs are covered in sibling articles; this one is about choosing and governing the single number at the top.
What a North Star is, and what it is not
A North Star metric has three properties. It measures value received by the user, not activity performed by the system. It is a leading indicator of the business outcome you care about, so it moves before revenue or retention do and gives you time to react. And it decomposes into inputs that specific teams can move with their own work. A metric that fails the first test rewards waste; one that fails the second arrives too late to steer by; one that fails the third produces a dashboard nobody acts on.
It is not a replacement for SLOs, budgets or eval suites. SLOs are constraints that must hold; the North Star is the quantity you try to grow while they hold. It is also not revenue. Revenue lags, is shaped by pricing and sales cycles, and cannot tell a serving team whether last Tuesday's engine upgrade helped. The idea is widely used in product analytics circles; its value here is not the label but the discipline of picking one number, writing its definition down and versioning that definition like code.
Why the obvious candidates fail
Most teams start from what their telemetry already exports. The table shows why each obvious candidate fails as the top number, even though several survive as inputs or counter-metrics further down.
| Candidate | What it rewards | Why it fails as the North Star |
|---|---|---|
| Tokens served | Volume | Verbose answers and retry storms raise it; a better model that answers in fewer tokens lowers it |
| Requests per day | Traffic | Regenerations, agent loops and client retries count as value |
| Daily active users | Reach | Says nothing about whether the answers helped; slow to react to model changes |
| GPU utilisation or MFU | Busy hardware | A saturated fleet with a long queue scores well while users time out |
| Cost per million tokens | Cheap tokens | Falls when you quantise aggressively and quality collapses |
| Offline eval score | Benchmark fit | Measures a fixed test set, not the traffic you actually serve |
| Qualified successful tasks | Delivered outcomes within SLO | Harder to compute, but moves only when users get value under acceptable latency |
The pattern in the failures is that each measures one layer of the stack in isolation. Tokens and utilisation are serving-layer numbers, eval scores are model-layer numbers, DAU is a product-layer number. The North Star has to sit above all of them and be computed from the joined record of a user intent, the model's answer, the latency it arrived with and what the user did next.
Defining qualified successful tasks
Define the metric in four parts, and write each part into a versioned definition file. A task is one user intent: a chat turn plus its regenerations, or one agent run with all its tool calls, or one API request for API customers. Sessionise raw requests into tasks so that a user who regenerates three times counts once, not four times. A task is successful when an outcome signal fires within a window: the user copied or accepted the answer, the agent's final tool call succeeded, the user continued without regenerating or rating it down, or an API caller did not retry the same prompt. A task is SLO-qualified when every request inside it met the latency SLO for its tier, for example time to first token under 1.5 s at p95 and inter-token latency under 60 ms, and it was not served by a degraded fallback model. A task is policy-clean when no safety classifier or abuse system flagged it.
The North Star is then the weekly count of tasks that are successful, SLO-qualified and policy-clean: qualified successful tasks, or QST. Report QST per GPU-hour beside it as the efficiency companion, not in place of it. Putting GPU-hours in the denominator of the headline is tempting for a GPU team, but a ratio can rise while the product shrinks: halve traffic, shut down half the fleet and the ratio is unchanged while users leave. Keep value as the headline and efficiency as the first thing you look at next to it.
The factor tree
Because the definition is a chain of filters, the headline factors exactly: QST equals tasks attempted times success rate times SLO-qualified share times policy-clean share, where each share is conditional on passing the previous filter. That makes attribution arithmetic, not argument. If QST fell 4% and the SLO-qualified share fell from 0.95 to 0.91 while the other factors held, the serving and capacity teams own the drop, and the latency histograms tell them where to look. If success rate fell while latency held, the model and eval owners own it.
The GPU fleet touches three of the four inputs. Capacity and scheduling decide the SLO-qualified share. Model choice and quantisation decide success rate. And prefill throughput limits how many long-context tasks you can accept at all, which shows up as tasks attempted when you rate-limit. This is why a GPU platform team should care about a product metric: it is the only number that prices a latency regression and a quality regression in the same unit.
Computing it from request logs
The computation needs one joined row per request: user and thread identifiers, timestamp, tier, time to first token, a per-request inter-token latency percentile, whether a fallback model served it, the outcome signals the client reports, the safety flag, and the GPU seconds attributed to it by your cost attribution pipeline. Most of the engineering effort is in that join, not in the arithmetic. The sketch below shows the arithmetic in pandas; in production the same logic runs as a scheduled warehouse query.
# qst.py - compute the North Star from joined request logs (one row per request)
import pandas as pd
DEF_VERSION = "qst-v3" # bump on any change below; keep old versions computable
TASK_GAP_S = 120 # requests closer than this, same user and thread, form one task
TTFT_SLO = {"interactive": 1.5, "batch": 10.0}
ITL_SLO_MS = 60
def to_tasks(req: pd.DataFrame) -> pd.DataFrame:
req = req.sort_values(["user_id", "thread_id", "ts"])
gap = req.groupby(["user_id", "thread_id"])["ts"].diff().dt.total_seconds()
req["task_seq"] = (gap.isna() | (gap > TASK_GAP_S)).cumsum()
req["slo_ok"] = (
(req["ttft_s"] <= req["tier"].map(TTFT_SLO))
& (req["itl_p95_ms"] <= ITL_SLO_MS)
& (~req["served_by_fallback"])
)
return req.groupby("task_seq").agg(
week=("ts", lambda s: s.min().to_period("W")),
regenerations=("is_regenerate", "sum"),
accepted=("accepted", "max"), # copy, apply, final tool call ok
rated_down=("thumbs_down", "max"),
slo_ok=("slo_ok", "all"), # every request in the task met SLO
flagged=("safety_flag", "max"),
gpu_seconds=("gpu_seconds", "sum"),
)
def weekly_qst(tasks: pd.DataFrame) -> pd.DataFrame:
success = tasks["accepted"] & ~tasks["rated_down"]
tasks = tasks.assign(success=success,
qst=success & tasks["slo_ok"] & ~tasks["flagged"])
w = tasks.groupby("week").agg(
attempted=("qst", "size"), success=("success", "sum"),
qst=("qst", "sum"), gpu_hours=("gpu_seconds", lambda s: s.sum() / 3600))
w["qst_per_gpu_hour"] = w["qst"] / w["gpu_hours"]
w["def_version"] = DEF_VERSION
return wThree details matter. First, slo_ok uses all across the task, so one slow regeneration disqualifies it; that matches the user's experience. Second, the definition version travels with every row of output, so a chart that spans a definition change can be split rather than silently mixing two metrics. Third, GPU seconds come from the serving engine's own accounting, prefill and decode included, apportioned by tokens when requests are batched together. Sharing a batch is the normal case, so a naive wall-clock attribution would overcount by the batch size.
Worked example: judging an FP8 rollout
Suppose a week of traffic produces 120,000 weekly users attempting 6.0 tasks each, so 720,000 tasks. Success rate is 71%, giving 511,200 successful tasks, and 93% of those met SLO; none was flagged, so the policy-clean share is 1.0, giving about 475,400 QST. The fleet used 9,600 GPU-hours, so efficiency is 49.5 QST per GPU-hour.
The serving team proposes moving the main model from BF16 to FP8 weights and KV cache. An A/B test on 10% of traffic shows the expected throughput gain: fewer GPUs needed at the same load, and queueing at peak almost disappears, so the SLO-qualified share rises from 93% to 97%. Success rate falls from 71% to 69%. Projected to full traffic with the same 720,000 tasks, QST becomes 720,000 x 0.69 x 0.97, about 481,900, a gain of 1.4%, while GPU-hours fall to 7,200 and efficiency rises to 66.9 QST per GPU-hour.
The North Star says ship, but the decomposition says look closer before you do. A two-point success drop is an average; slice it by task type and you may find code tasks lost six points while chat lost almost none. If code tasks are the segment that drives paid conversion, the right move is to route code traffic to the BF16 pool and ship FP8 for the rest. The metric did its job: it priced latency and quality in one unit, showed the change was net positive, and the factor tree pointed at where the cost landed.
Counter-metrics
Every input has a cheap way to move it, so pair each with a counter-metric that the same review looks at.
- Tasks attempted can be inflated by splitting one intent into many, for example by shortening the session gap. Counter: tasks per user-minute, and freeze the sessionisation constants in the definition file.
- Success rate rests on proxies. Hiding the regenerate button or making the thumbs-down harder to find raises it without helping anyone. Counter: a weekly human-labelled sample of a few hundred tasks, scored blind, compared with the proxy. When the gap between proxy and labels widens, the proxy is drifting.
- SLO-qualified share can be bought with GPUs. Counter: GPU-hours per QST and peak-hour headroom, so the capacity team cannot buy the share with unbounded spend.
- Policy-clean share can be raised by refusing more. Counter: the over-refusal rate on a fixed benign prompt set.
Governing the definition
Governance is what keeps the number trustworthy after the launch slide. Keep the definition in a repository file with an owner, the constants and the outcome signals, reviewed like any other code change. Review the North Star weekly with the factor decomposition and the counter-metrics on the same page, and record each week's top movement with its attributed cause. Revisit the definition quarterly, not weekly: a definition that changes whenever the number disappoints stops being a measurement. When you must change it, compute both versions in parallel for at least four weeks and publish the bridge between them.
Do not set individual targets on the North Star itself. Teams own inputs; the headline is shared. Targets on a shared number invite the arguments about credit and blame that the factor tree was built to settle.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Outcome signal changes with a client release | Step change in success rate on release day | Version outcome events; annotate releases on the chart |
| Fallback model counted as qualified | QST holds during a GPU incident users noticed | Mark fallback-served requests and exclude them |
| Bot or eval traffic in the logs | Tasks attempted jumps at night | Filter by authenticated, non-internal callers |
| Per-request GPU seconds missing | Efficiency companion swings with batch size | Attribute from engine accounting by token share |
| Definition edited silently | Trend breaks with no product change | Version constant in output; parallel runs on change |
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Task-level rather than request-level | Matches user intent | Sessionisation logic to maintain |
| Strict all-requests SLO filter | Reflects worst moment in a task | Long agent tasks are penalised more |
| Proxy outcome signals | Computable daily at full volume | Drift; needs labelled audits |
| Efficiency as companion, not headline | Cannot shrink to win | Two numbers to read instead of one |
What to do next
- Write a one-page definition file: task boundary, success signals, SLO thresholds per tier, policy filter, version string and owner.
- Build the joined request table, including fallback flags and per-request GPU seconds.
- Backfill eight weeks of QST and the four factors; check that the factors multiply back to the headline.
- Assign each factor an owning team and a counter-metric, and put all of them on one review page.
- Start a weekly blind-labelled sample to audit the success proxy.
- Judge the next model, quantisation or engine change by projected QST and QST per GPU-hour, sliced by task type, before you ship it.
Keep learning: LLM KPI dashboards for the metric tree beneath the North Star, individual contributor KPIs for per-engineer measures, LLM A/B testing for running the experiment in the worked example, and SLO burn rate alerts for the latency SLOs the qualified share depends on.