A team that serves a large language model on its own GPUs can watch hundreds of numbers: tokens per second, time to first token, KV cache occupancy, eval scores, daily active users, cost per million tokens. Every one of them is useful to somebody, and none of them answers the question a planning meeting actually asks: is the product delivering more value this week than last week, and is the GPU fleet helping or hurting that? A North Star metric is the single number chosen to answer that question, agreed in advance, defined precisely enough that two engineers computing it from the same logs get the same answer.

This article builds one for a GPU-served LLM product from first principles. It covers what makes a good North Star, why the obvious candidates such as tokens served fail, a precise definition based on qualified successful tasks, the input tree that connects the headline to model, serving and capacity work, the code to compute it from request logs, a worked example of judging an FP8 rollout with it, and the counter-metrics that stop people gaming it. The metric tree for dashboards and the per-engineer KPIs are covered in sibling articles; this one is about choosing and governing the single number at the top.

What a North Star is, and what it is not

A North Star metric has three properties. It measures value received by the user, not activity performed by the system. It is a leading indicator of the business outcome you care about, so it moves before revenue or retention do and gives you time to react. And it decomposes into inputs that specific teams can move with their own work. A metric that fails the first test rewards waste; one that fails the second arrives too late to steer by; one that fails the third produces a dashboard nobody acts on.

It is not a replacement for SLOs, budgets or eval suites. SLOs are constraints that must hold; the North Star is the quantity you try to grow while they hold. It is also not revenue. Revenue lags, is shaped by pricing and sales cycles, and cannot tell a serving team whether last Tuesday's engine upgrade helped. The idea is widely used in product analytics circles; its value here is not the label but the discipline of picking one number, writing its definition down and versioning that definition like code.

Why the obvious candidates fail

Most teams start from what their telemetry already exports. The table shows why each obvious candidate fails as the top number, even though several survive as inputs or counter-metrics further down.

CandidateWhat it rewardsWhy it fails as the North Star
Tokens servedVolumeVerbose answers and retry storms raise it; a better model that answers in fewer tokens lowers it
Requests per dayTrafficRegenerations, agent loops and client retries count as value
Daily active usersReachSays nothing about whether the answers helped; slow to react to model changes
GPU utilisation or MFUBusy hardwareA saturated fleet with a long queue scores well while users time out
Cost per million tokensCheap tokensFalls when you quantise aggressively and quality collapses
Offline eval scoreBenchmark fitMeasures a fixed test set, not the traffic you actually serve
Qualified successful tasksDelivered outcomes within SLOHarder to compute, but moves only when users get value under acceptable latency

The pattern in the failures is that each measures one layer of the stack in isolation. Tokens and utilisation are serving-layer numbers, eval scores are model-layer numbers, DAU is a product-layer number. The North Star has to sit above all of them and be computed from the joined record of a user intent, the model's answer, the latency it arrived with and what the user did next.

Defining qualified successful tasks

Define the metric in four parts, and write each part into a versioned definition file. A task is one user intent: a chat turn plus its regenerations, or one agent run with all its tool calls, or one API request for API customers. Sessionise raw requests into tasks so that a user who regenerates three times counts once, not four times. A task is successful when an outcome signal fires within a window: the user copied or accepted the answer, the agent's final tool call succeeded, the user continued without regenerating or rating it down, or an API caller did not retry the same prompt. A task is SLO-qualified when every request inside it met the latency SLO for its tier, for example time to first token under 1.5 s at p95 and inter-token latency under 60 ms, and it was not served by a degraded fallback model. A task is policy-clean when no safety classifier or abuse system flagged it.

The North Star is then the weekly count of tasks that are successful, SLO-qualified and policy-clean: qualified successful tasks, or QST. Report QST per GPU-hour beside it as the efficiency companion, not in place of it. Putting GPU-hours in the denominator of the headline is tempting for a GPU team, but a ratio can rise while the product shrinks: halve traffic, shut down half the fleet and the ratio is unchanged while users leave. Keep value as the headline and efficiency as the first thing you look at next to it.

The factor tree

One North Star, four multiplicative inputs, each with an owner and a counter-metricQualified successful tasksper week (the North Star)Tasks attemptedusers x tasks per userSuccess rateoutcome signal per taskSLO-qualified shareTTFT, ITL, no fallbackPolicy-clean shareno safety or abuse flagowner: productowner: model + evalsowner: serving + capacityowner: safetyCounter: cost per usergrowth bought with spendCounter: labelled auditproxy vs human judgementCounter: GPU-hours/QSTSLO bought with GPUsCounter: over-refusalclean share by refusingEfficiency companionQST per GPU-hour, reported next to the North Star, never instead of itEvery input moves the headline multiplicatively; every counter-metric catches a cheap way to move the input.
QST factors into tasks attempted, success rate, SLO-qualified share and policy-clean share. Each factor has an owning team and a counter-metric that exposes the cheap way to move it.

Because the definition is a chain of filters, the headline factors exactly: QST equals tasks attempted times success rate times SLO-qualified share times policy-clean share, where each share is conditional on passing the previous filter. That makes attribution arithmetic, not argument. If QST fell 4% and the SLO-qualified share fell from 0.95 to 0.91 while the other factors held, the serving and capacity teams own the drop, and the latency histograms tell them where to look. If success rate fell while latency held, the model and eval owners own it.

The GPU fleet touches three of the four inputs. Capacity and scheduling decide the SLO-qualified share. Model choice and quantisation decide success rate. And prefill throughput limits how many long-context tasks you can accept at all, which shows up as tasks attempted when you rate-limit. This is why a GPU platform team should care about a product metric: it is the only number that prices a latency regression and a quality regression in the same unit.

Computing it from request logs

The computation needs one joined row per request: user and thread identifiers, timestamp, tier, time to first token, a per-request inter-token latency percentile, whether a fallback model served it, the outcome signals the client reports, the safety flag, and the GPU seconds attributed to it by your cost attribution pipeline. Most of the engineering effort is in that join, not in the arithmetic. The sketch below shows the arithmetic in pandas; in production the same logic runs as a scheduled warehouse query.

# qst.py - compute the North Star from joined request logs (one row per request)
import pandas as pd

DEF_VERSION = "qst-v3"          # bump on any change below; keep old versions computable
TASK_GAP_S = 120                 # requests closer than this, same user and thread, form one task
TTFT_SLO = {"interactive": 1.5, "batch": 10.0}
ITL_SLO_MS = 60

def to_tasks(req: pd.DataFrame) -> pd.DataFrame:
    req = req.sort_values(["user_id", "thread_id", "ts"])
    gap = req.groupby(["user_id", "thread_id"])["ts"].diff().dt.total_seconds()
    req["task_seq"] = (gap.isna() | (gap > TASK_GAP_S)).cumsum()
    req["slo_ok"] = (
        (req["ttft_s"] <= req["tier"].map(TTFT_SLO))
        & (req["itl_p95_ms"] <= ITL_SLO_MS)
        & (~req["served_by_fallback"])
    )
    return req.groupby("task_seq").agg(
        week=("ts", lambda s: s.min().to_period("W")),
        regenerations=("is_regenerate", "sum"),
        accepted=("accepted", "max"),          # copy, apply, final tool call ok
        rated_down=("thumbs_down", "max"),
        slo_ok=("slo_ok", "all"),               # every request in the task met SLO
        flagged=("safety_flag", "max"),
        gpu_seconds=("gpu_seconds", "sum"),
    )

def weekly_qst(tasks: pd.DataFrame) -> pd.DataFrame:
    success = tasks["accepted"] & ~tasks["rated_down"]
    tasks = tasks.assign(success=success,
                         qst=success & tasks["slo_ok"] & ~tasks["flagged"])
    w = tasks.groupby("week").agg(
        attempted=("qst", "size"), success=("success", "sum"),
        qst=("qst", "sum"), gpu_hours=("gpu_seconds", lambda s: s.sum() / 3600))
    w["qst_per_gpu_hour"] = w["qst"] / w["gpu_hours"]
    w["def_version"] = DEF_VERSION
    return w

Three details matter. First, slo_ok uses all across the task, so one slow regeneration disqualifies it; that matches the user's experience. Second, the definition version travels with every row of output, so a chart that spans a definition change can be split rather than silently mixing two metrics. Third, GPU seconds come from the serving engine's own accounting, prefill and decode included, apportioned by tokens when requests are batched together. Sharing a batch is the normal case, so a naive wall-clock attribution would overcount by the batch size.

Worked example: judging an FP8 rollout

Suppose a week of traffic produces 120,000 weekly users attempting 6.0 tasks each, so 720,000 tasks. Success rate is 71%, giving 511,200 successful tasks, and 93% of those met SLO; none was flagged, so the policy-clean share is 1.0, giving about 475,400 QST. The fleet used 9,600 GPU-hours, so efficiency is 49.5 QST per GPU-hour.

The serving team proposes moving the main model from BF16 to FP8 weights and KV cache. An A/B test on 10% of traffic shows the expected throughput gain: fewer GPUs needed at the same load, and queueing at peak almost disappears, so the SLO-qualified share rises from 93% to 97%. Success rate falls from 71% to 69%. Projected to full traffic with the same 720,000 tasks, QST becomes 720,000 x 0.69 x 0.97, about 481,900, a gain of 1.4%, while GPU-hours fall to 7,200 and efficiency rises to 66.9 QST per GPU-hour.

The North Star says ship, but the decomposition says look closer before you do. A two-point success drop is an average; slice it by task type and you may find code tasks lost six points while chat lost almost none. If code tasks are the segment that drives paid conversion, the right move is to route code traffic to the BF16 pool and ship FP8 for the rest. The metric did its job: it priced latency and quality in one unit, showed the change was net positive, and the factor tree pointed at where the cost landed.

Counter-metrics

Every input has a cheap way to move it, so pair each with a counter-metric that the same review looks at.

  • Tasks attempted can be inflated by splitting one intent into many, for example by shortening the session gap. Counter: tasks per user-minute, and freeze the sessionisation constants in the definition file.
  • Success rate rests on proxies. Hiding the regenerate button or making the thumbs-down harder to find raises it without helping anyone. Counter: a weekly human-labelled sample of a few hundred tasks, scored blind, compared with the proxy. When the gap between proxy and labels widens, the proxy is drifting.
  • SLO-qualified share can be bought with GPUs. Counter: GPU-hours per QST and peak-hour headroom, so the capacity team cannot buy the share with unbounded spend.
  • Policy-clean share can be raised by refusing more. Counter: the over-refusal rate on a fixed benign prompt set.

Governing the definition

Governance is what keeps the number trustworthy after the launch slide. Keep the definition in a repository file with an owner, the constants and the outcome signals, reviewed like any other code change. Review the North Star weekly with the factor decomposition and the counter-metrics on the same page, and record each week's top movement with its attributed cause. Revisit the definition quarterly, not weekly: a definition that changes whenever the number disappoints stops being a measurement. When you must change it, compute both versions in parallel for at least four weeks and publish the bridge between them.

Do not set individual targets on the North Star itself. Teams own inputs; the headline is shared. Targets on a shared number invite the arguments about credit and blame that the factor tree was built to settle.

Failure modes

FailureSymptomFix
Outcome signal changes with a client releaseStep change in success rate on release dayVersion outcome events; annotate releases on the chart
Fallback model counted as qualifiedQST holds during a GPU incident users noticedMark fallback-served requests and exclude them
Bot or eval traffic in the logsTasks attempted jumps at nightFilter by authenticated, non-internal callers
Per-request GPU seconds missingEfficiency companion swings with batch sizeAttribute from engine accounting by token share
Definition edited silentlyTrend breaks with no product changeVersion constant in output; parallel runs on change

Trade-offs

ChoiceGainCost
Task-level rather than request-levelMatches user intentSessionisation logic to maintain
Strict all-requests SLO filterReflects worst moment in a taskLong agent tasks are penalised more
Proxy outcome signalsComputable daily at full volumeDrift; needs labelled audits
Efficiency as companion, not headlineCannot shrink to winTwo numbers to read instead of one

What to do next

  1. Write a one-page definition file: task boundary, success signals, SLO thresholds per tier, policy filter, version string and owner.
  2. Build the joined request table, including fallback flags and per-request GPU seconds.
  3. Backfill eight weeks of QST and the four factors; check that the factors multiply back to the headline.
  4. Assign each factor an owning team and a counter-metric, and put all of them on one review page.
  5. Start a weekly blind-labelled sample to audit the success proxy.
  6. Judge the next model, quantisation or engine change by projected QST and QST per GPU-hour, sliced by task type, before you ship it.

Keep learning: LLM KPI dashboards for the metric tree beneath the North Star, individual contributor KPIs for per-engineer measures, LLM A/B testing for running the experiment in the worked example, and SLO burn rate alerts for the latency SLOs the qualified share depends on.

Key takeaway: Pick one number that counts delivered value under acceptable latency: tasks that succeeded, met SLO and stayed policy-clean. Factor it into inputs that teams own, pair every input with a counter-metric, report QST per GPU-hour beside it rather than instead of it, and version the definition like code.