A single GPU running a modern inference engine serves dozens of requests at once. Continuous batching packs a long prompt from one customer, a few decode tokens from forty chat sessions and a cached system prompt shared by a whole product into the same forward pass. The GPU-hour is billed once. The question cost attribution answers is: which of those requests, tenants and features consumed which part of that hour?
That question is different from cost analysis, which asks what a token costs on average, and different from FinOps process, which asks who sees the number and what they do about it. The LLM FinOps article covers capture, rollups and chargeback versus showback; LLM cost analysis derives cost per million tokens. This article sits between them. It is about the allocation math inside one replica: how to split a measured engine step among the requests in it, how to charge for KV cache memory that a request holds while it waits, how to credit prefix cache hits, where idle time goes, and how to prove the result adds up to the invoice.
The running example is a single-GPU replica. Prices are illustrative, not vendor rates: we use $3.60 per GPU-hour because it makes one millisecond cost exactly one micro-dollar.
The unit and the conservation rule
Start with the rule that makes everything else testable: conservation. For every billing hour, the charges assigned to requests plus the explicitly labelled idle and headroom charges must equal the GPU-hours you paid for, to the micro-dollar. Any attribution scheme that cannot state this invariant will drift, and nobody will notice until finance asks why the team totals add up to 83 percent of the bill.
The second rule is that you allocate measured time, never predicted time. An engine runs in steps: each step is one forward pass over a batch of slots, where a slot is a request contributing either a chunk of prompt tokens (prefill) or one new token (decode). The engine knows the wall time of each step and which slots were in it. Those two facts are the raw material. Models of cost decide how to split a step, but the step's total comes from the clock.
Tensor parallel replicas multiply the step cost by the number of GPUs in the replica. Rejected speculative decoding drafts are still work the request caused, so they are charged to it.
The attribution pipeline
The engine step log is the expensive piece to get right. You need, per step: a monotonic start and end timestamp, and for each slot the request id, the number of new tokens computed, the context length already cached, and the KV blocks the request held. Most engines do not emit this out of the box; production teams usually add a small hook around the scheduler's step function and write compact binary records to a local buffer that ships asynchronously. Check what your engine version exposes before building a hook, and never block the step on logging I/O.
Request context (tenant, feature, API key) is joined by request id later, outside the hot path. Calibration runs supply the cost model coefficients described next.
Splitting a step by the work it did
The naive splits are both wrong. Splitting a step evenly per request makes a 512-token prompt chunk cost the same as one decode token. Splitting per token makes decode nearly free, because a decode slot contributes one token while its real cost is reading its whole KV cache. What you want is a split proportional to the work each slot actually caused.
A simple model fits measured step times well enough for attribution. For slot i with t new tokens and context L:
step_ms ~ c0 + sum_i [ a * t_i + b * t_i * (L_i + (t_i + 1) / 2) ]
c0 : fixed per-step cost (reading the weights once, kernel launches, sampling)
a : per-token cost of the dense layers (matmuls against the weights)
b : per token-pair cost of attention (each new token attends to its context)Fit c0, a and b by ordinary least squares on a few thousand logged steps from a calibration run that mixes prompt lengths and batch sizes. You are not trying to predict latency; you are estimating the relative work of slots. The prediction then becomes a set of weights, and the measured step time is split in those proportions, so the model's absolute error never leaks into the total.
The fixed cost c0 deserves a policy decision. Splitting it equally per slot says every request in the batch benefited equally from the single weight read. That is the common choice and the one used below. The alternative, splitting it by token, hands almost all of it to the prefill request and makes decode look cheap again.
KV cache as memory-time
Compute share is half the story. A decode request with a 4,000-token context holds memory the scheduler cannot give to anyone else, and on a busy replica KV capacity, not compute, is usually what limits batch size. For Llama 3 8B (32 layers, 8 KV heads, head dimension 128, 16-bit values) each cached token costs 2 x 32 x 8 x 128 x 2 bytes = 128 KiB, so a 4,000-token context pins about 500 MiB. The continuous batching scheduler admits new requests only while blocks are free.
Charge for it explicitly as memory-time: blocks held multiplied by the duration held. Inside a step, that becomes each slot's share of the blocks held during the step. Requests that are queued or preempted but still hold blocks are charged too, because their blocks were unavailable; if your engine swaps them out, they stop accruing.
Blend the two shares with a pool-level weight alpha. A pool that is compute bound (short contexts, high batch) uses an alpha near 1. A pool that is memory bound (long contexts, frequent preemption or admission waits while the GPU shows spare compute) moves alpha down. Derive alpha from evidence, such as the fraction of steps where admission was blocked on free blocks, and review it quarterly rather than tuning it per customer.
A conserving allocator in code
The allocator below works in integer micro-dollars and uses largest-remainder rounding so that the shares sum to the step cost exactly. Floating-point shares that are rounded independently lose or invent a few units per step, and over a billion steps that becomes real money and a failed reconciliation. In production use nano-dollars for headroom.
from dataclasses import dataclass
@dataclass
class Slot:
request_id: str
new_tokens: int # tokens computed this step (prefill chunk or 1 decode token)
context: int # tokens already in the KV cache before this step
kv_blocks: int # KV blocks this request holds during the step
def predicted_ms(s, a_ms, b_ms):
# attention work grows with new tokens times the context they attend to
attn = s.new_tokens * (s.context + (s.new_tokens + 1) / 2)
return a_ms * s.new_tokens + b_ms * attn
def allocate_step(slots, step_ms, rate_micro_per_ms, c0_ms, a_ms, b_ms, alpha=0.8):
"""Split one measured engine step among its requests, in integer micro-dollars.
alpha weights the compute share against the KV memory share.
The returned amounts always sum to the step's cost exactly.
"""
total = round(step_ms * rate_micro_per_ms)
fixed = c0_ms / len(slots)
compute = [fixed + predicted_ms(s, a_ms, b_ms) for s in slots]
c_sum = sum(compute)
kv_sum = sum(s.kv_blocks for s in slots) or 1
weights = [alpha * c / c_sum + (1 - alpha) * s.kv_blocks / kv_sum
for c, s in zip(compute, slots)]
raw = [w * total for w in weights]
floors = [int(x) for x in raw]
# largest-remainder rounding: hand leftover micro-dollars to the biggest fractions
order = sorted(range(len(raw)), key=lambda i: raw[i] - floors[i], reverse=True)
for i in order[: total - sum(floors)]:
floors[i] += 1
assert sum(floors) == total
return {s.request_id: amt for s, amt in zip(slots, floors)}
Worked example: one 50 ms step
One measured step of 50 ms on our $3.60 per GPU-hour replica costs 50 micro-dollars. Calibration gave c0 = 12 ms, a = 0.05 ms per token and b = 0.00005 ms per token-pair. The batch has four slots, with 16-token KV blocks:
| Slot | Work this step | Context | KV blocks | Predicted variable ms |
|---|---|---|---|---|
| A | 512-token prefill chunk | 0 | 32 | 32.17 |
| B | 1 decode token | 4,000 | 250 | 0.25 |
| C | 1 decode token | 1,000 | 63 | 0.10 |
| D | 1 decode token | 200 | 13 | 0.06 |
With alpha = 1 (compute only), each slot gets 3 ms of the fixed cost plus its variable prediction, and the measured 50 ms is split in those proportions: A is charged 39, B 4, C 4 and D 3 micro-dollars. Per-token splitting would have charged A about 49.7 and each decoder about 0.1, underpricing decode by more than thirty times. Per-request splitting would have charged everyone 12.5.
With alpha = 0.8, the KV share moves money toward B, whose long context holds 70 percent of the blocks: A 32, B 10, C 5, D 3. Which answer is right depends on what limits this pool. If requests are waiting for blocks, B's context is the scarce resource and the second split is the honest one.
Prefix cache hits
With prefix caching, the first request to send a long system prompt computes its KV blocks; later requests with the same prefix reuse them and skip that prefill. Three policies are common:
- First writer pays. Simple and causally accurate, but whoever warms the cache after a restart or eviction gets an unpredictable bill.
- Shared blocks split by holders. In the KV share, a block referenced by five running requests counts one fifth to each. This is fair for memory, and the allocator above handles it if you pass fractional block counts.
- Cache-read pricing. Charge cache-hit tokens at a fixed fraction of the prefill rate, the way hosted APIs price cached input, and book the difference as pool savings. Internal users get stable prices and still see the incentive to share prefixes.
Whatever you choose, record cached tokens per request. Without that field, a prompt template change that breaks the shared prefix shows up as a mysterious cost jump with no owner.
Idle time and headroom
Step times never cover the whole hour. Gaps between steps are idle time, and a pool kept at 60 percent load to protect latency has a lot of it. Do not hide idle inside per-request rates; label it. Two allocations are defensible:
- Spread by usage. Idle in the hour is distributed in proportion to each tenant's active charges. Totals reconcile, and rates rise automatically when load falls.
- Charge to the capacity owner. If headroom exists because a product demanded a latency SLO, the product that set the SLO owns the idle. This makes the cost of an SLO visible, which spreading hides.
Many teams do both: a baseline target utilisation is spread, anything beyond it goes to the owner of the reservation.
Reconciliation
Reconcile every hour, automatically, on three levels:
- Time. Sum of step durations plus measured gaps equals wall-clock time times GPUs per replica, within a small tolerance for clock skew. A shortfall means dropped log records.
- Money. Ledger charges plus idle and headroom lines equal the invoiced GPU-hours times rate. With integer units this should be exact.
- Plausibility. The engine's busy fraction (step time over wall time) should track the GPU's own activity counters, such as DCGM's graphics engine activity profiling field (
DCGM_FI_PROF_GR_ENGINE_ACTIVE; see DCGM). If the engine reports 90 percent busy while DCGM shows 30 percent active, steps are stalling on the host, and the attribution is charging customers for CPU-side waits.
Failure modes
- Dropped step records under load. Logging buffers overflow at peak, so exactly the busiest minutes go missing. Count records with a sequence number and alert on gaps.
- Calibration drift. An engine upgrade changes kernel performance and the old coefficients shift charges between prefill-heavy and decode-heavy tenants. Recalibrate on every engine or driver change and version the coefficients in the ledger rows.
- Charging retries twice. A client timeout followed by a retry produces two requests. Both did consume GPU time; decide explicitly whether the caller or the platform pays, especially when the platform caused the timeout.
- Join failures. Requests with no tenant context land in an unattributed bucket that grows silently. Track it as a percentage and page when it passes a threshold.
Trade-offs
| Choice | Simple option | Accurate option | When the accurate option pays off |
|---|---|---|---|
| Split basis | Tokens in and out | Calibrated step model | Mixed prompt lengths, long contexts |
| Memory | Ignore | KV memory-time share | Pools limited by KV blocks, not compute |
| Prefix cache | First writer pays | Cache-read pricing | Shared system prompts across teams |
| Idle | Spread by usage | Owner of the SLO pays | Pools sized for latency, not load |
| Granularity | Hourly token counters | Per-step records | Internal chargeback with disputes |
What to do next
- Write down the conservation invariant for your pools and add an hourly job that checks it.
- Find out what per-step data your inference engine exposes; add a non-blocking scheduler hook for slot composition, step timing and KV blocks held if it is missing.
- Run a calibration sweep per model and GPU type, fit c0, a and b, and store them with a version.
- Implement the allocator in integer units with largest-remainder rounding and a property test that shares always sum to the step cost.
- Pick and document policies for fixed cost, prefix cache hits, idle time and retries.
- Compare engine busy time with DCGM activity for a week before trusting the numbers.
- Publish per-tenant cost per request alongside token counts, so teams see why a long context costs more than its token count suggests.