Chargeback is the practice of billing internal teams for the GPU capacity they consume, with the money actually moving between cost centres. It differs from showback, which reports the same numbers without moving money, and from cost attribution, which is the measurement problem of deciding which request used which share of a shared GPU step. The cost attribution article solves the measurement; the FinOps overview explains why organisations move from showback to chargeback at all. This article is about the billing machinery in between: what goes into the cost pool, which units to bill, how to set rates, who pays for idle GPUs, how the monthly close works and which behaviours each design rewards.
The test of a chargeback system is simple to state. Every dollar the fleet costs in a month must land in exactly one place: a team's statement, a central budget, or a variance account that someone owns. If the books do not balance, the system loses credibility in the first quarter and teams start disputing every line.
Cost pools and overhead
Start with the cost pool: everything the fleet costs per month. That is hardware depreciation or cloud spend, power and cooling, data centre space, network, storage attached to the serving tier, software licences and the operations team. The infrastructure cost article builds this into a fully loaded cost per GPU-hour, and that number is the input here.
Then split the pool by service, because different services need different units. A typical split has three pools: dedicated training reservations, a shared inference pool, and platform overhead such as spare GPUs, burn-in nodes and the control plane. Overhead is not billed directly; it is spread over the billable pools as an uplift on their rates. Make the uplift visible on the rate card so teams can see what the platform costs and argue about it.
The pipeline
Choosing billable units
A billable unit must be measurable without dispute, roughly proportional to cost, and something the team can control. No single unit satisfies all three for LLM workloads, so most rate cards have several.
| Unit | Fits | Behaviour it rewards | Weakness |
|---|---|---|---|
| Reserved GPU-hour | Training, dedicated serving | Releasing unused reservations | Charges idle reserved GPUs to the owner, which is correct but unpopular |
| MIG slice-hour | Small models, dev environments | Right-sizing to a slice | Slice profiles differ per GPU generation |
| Weighted token unit | Shared serving pool | Caching, shorter outputs, smaller models | Weights must track measured cost |
| Priority tier multiplier | Mixed latency classes | Using batch tier for offline work | Needs real scheduling difference to be fair |
Weighted token units are the usual choice for shared serving. Bill input, cached input and output tokens at different weights per model, with the weights taken from measured attribution rather than intuition: if the attribution pipeline shows an output token on a given model costs four times an uncached input token, the weight is four. Revisit weights when the engine or hardware changes, because prefix caching, speculative decoding and quantisation all move the ratios. For MIG partitions, bill the slice-hour regardless of activity, exactly like a reservation.
Setting rates
Rates are set once per period, usually a quarter or a year, from a budget, and held fixed. That is the core idea of budget rates: teams can forecast their bill, and the platform team carries the risk that volume differs from plan. The recovery equation for a pool is:
rate = (pool_cost x (1 + overhead_uplift)) / planned_billable_units
recovered = actual_units x rate
variance = pool_cost x (1 + overhead_uplift) - recovered
(positive = under-recovery, negative = over-recovery)Floating rates, recomputed monthly from actual volume, recover cost exactly but make the price depend on other teams' behaviour. A team that cut its usage in half can see its bill fall by less than expected because everyone's rate rose. Budget rates avoid that, at the cost of a variance account that must be settled.
from decimal import Decimal, ROUND_HALF_UP
def budget_rate(pool_cost, uplift, planned_units, quantum=Decimal("0.01")):
raw = pool_cost * (1 + uplift) / planned_units
return raw.quantize(quantum, rounding=ROUND_HALF_UP)
def rate_usage(events, rate_card):
# events: dicts with tenant, unit, quantity; rate_card maps unit to Decimal rate.
charges = {}
for e in events:
amount = Decimal(e["quantity"]) * rate_card[e["unit"]]
charges[e["tenant"]] = charges.get(e["tenant"], Decimal("0")) + amount
return chargesUse decimal arithmetic for money. Binary floating point accumulates rounding error over millions of usage events, and a statement that is off by four cents is a dispute.
Who pays for idle capacity
Idle capacity is where chargeback systems go wrong most often, because it has no requester. There are four defensible policies, and the choice should be explicit.
| Policy | Who pays for idle | Effect |
|---|---|---|
| Reservation owner pays | The team holding the reservation | Strong incentive to release capacity; fits dedicated pools |
| Spread in the rate | Every user of the pool, via the budget rate | Simple; hides idle cost and weakens pressure to fill the pool |
| Central budget absorbs | The platform or finance | Encourages adoption; idle cost becomes invisible to users |
| Strategic headroom line | A named budget for surge and failover | Makes the redundancy decision visible and owned |
A common and defensible combination: reservation owners pay for their reserved hours whether used or not; the shared pool's planned idle is in its budget rate; spare and burn-in GPUs are overhead uplift; and failover headroom sized by the capacity analysis is a strategic headroom line owned by whoever set the SLO.
Passing through commitment risk
If the fleet runs on committed cloud capacity or purchased hardware, someone carries the commitment risk: the obligation to pay whether or not the GPUs are used. The reserved capacity article covers sizing the commitment. Chargeback decides who inherits it. Passing it through means teams sign internal commitments matching the external one, and pay for them regardless of use; that is accurate but makes teams reluctant to commit. Absorbing it centrally makes adoption easy but leaves the platform with a variance it cannot control. A middle path offers two rates: a lower committed rate for teams that sign up for a quantity for the period, and a higher on-demand rate for everything else, with the spread paying for the risk the platform keeps.
The ledger
Store charges in an append-only ledger. Each usage event carries an idempotency key from the meter, so a replayed batch cannot double-bill. Each charge row records the rate card version it used, so a statement can be reproduced exactly a year later. Closed periods are immutable: corrections are new adjustment rows in the next open period that reference the original.
CREATE TABLE usage_event (
event_id TEXT PRIMARY KEY, -- idempotency key from the meter
period TEXT NOT NULL, -- e.g. 2026-09
tenant TEXT NOT NULL, -- 'unallocated' if the tag was missing
unit TEXT NOT NULL, -- reserved_gpu_hour, wtu_million, ...
quantity NUMERIC NOT NULL
);
CREATE TABLE charge (
charge_id BIGSERIAL PRIMARY KEY,
period TEXT NOT NULL,
tenant TEXT NOT NULL,
unit TEXT NOT NULL,
quantity NUMERIC NOT NULL,
rate NUMERIC NOT NULL,
rate_card TEXT NOT NULL, -- version id, e.g. FY27-Q1
amount NUMERIC NOT NULL,
kind TEXT NOT NULL DEFAULT 'usage', -- usage | adjustment
ref_charge BIGINT REFERENCES charge(charge_id)
);
-- Conservation check run at close: must return exactly zero.
SELECT p.cost_with_uplift
- (SELECT COALESCE(SUM(amount), 0) FROM charge WHERE period = p.period)
- p.variance_booked AS imbalance
FROM pool_period p WHERE p.period = '2026-09';Track the unallocated tenant as a first-class line. Usage with a missing or invalid team tag lands there and is reported to the platform team weekly. If it exceeds a few percent of the pool, fix tagging at admission (reject untagged API keys) before the close, not after.
Worked example: one monthly close
Consider a fleet of 256 GPUs with a fully loaded cost of 2.40 dollars per GPU-hour, an illustrative figure; use your own from the infrastructure cost model. A 730-hour month costs 256 x 730 x 2.40 = 448,512 dollars. The split is 128 GPUs of training reservations, 112 GPUs in the shared serving pool and 16 GPUs of spares and burn-in.
The 16 overhead GPUs cost 16 x 730 x 2.40 = 28,032 dollars, spread over the 240 billable GPUs as an uplift of 16 / 240 = 6.67 percent. The loaded rate is 2.56 dollars per GPU-hour. Training reservations: team A holds 64 GPUs, B holds 48 and C holds 16, billed 119,603.20, 89,702.40 and 29,900.80 dollars, together 239,206.40, whether or not the GPUs were busy.
The serving pool costs 112 x 730 x 2.56 = 209,305.60 dollars with uplift. The plan expected 52.3 billion weighted token units, giving a rate of 4.00 dollars per million units. Actual volume was 47.0 billion: search used 18.2 billion (72,800 dollars), the support assistant 21.5 billion (86,000) and other teams 7.3 billion (29,200), recovering 188,000 dollars.
| Line | Amount (USD) |
|---|---|
| Fleet cost pool | 448,512.00 |
| Training reservations billed | 239,206.40 |
| Serving usage billed | 188,000.00 |
| Serving under-recovery to variance | 21,305.60 |
| Check: billed + variance | 448,512.00 |
The books balance, and the close surfaces one decision: serving under-recovered by 10.2 percent because volume came in below plan. If the policy is that the platform absorbs variance within plus or minus 10 percent and anything beyond triggers a quarterly true-up, this month is just outside the band. Note what the platform should not do: raise the serving rate mid-quarter. That punishes the teams who are using the pool to cover the plan error of the teams who are not.
Disputes, true-ups and budgets
Publish draft statements several working days before the close and accept disputes only with evidence: event IDs, a time range, a reason. A dispute resolves into an adjustment row, positive or negative, in the next open period, referencing the original charges. A true-up settles the variance account at the end of the period, either by a pro-rata credit or debit to pool users, or by a central write-off. Agree which in advance. True-ups decided after the numbers are known become negotiations.
Budgets close the loop. Give each team a monthly budget per pool, alert at 50, 80 and 100 percent of forecast spend, and decide in advance what happens at 100: a soft cap that notifies a manager, a move to the batch priority tier, or a hard cap that rejects new requests. Hard caps on customer-facing services cause outages, so reserve them for development and experimentation keys.
Failure modes
- The death spiral. Fixed costs spread over falling volume raise the rate, which drives more teams to external APIs, which raises the rate again. Hold budget rates for the period and fund strategic headroom centrally.
- Floating rates. Bills move with other teams' behaviour; nobody can forecast, and trust erodes.
- Weights that no longer track cost. After enabling prefix caching, cached input billed at full price punishes the teams who made caching work.
- Editable history. Changing a closed month's charges breaks every reconciliation downstream. Adjust forward only.
- Growing unallocated bucket. Untagged usage that nobody pays for hides real demand. Enforce tags at admission.
- Double counting. Billing both the reservation and the tokens served on it charges the same GPU-hour twice.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Budget rates vs floating | Predictable bills | A variance account to settle |
| Token units vs GPU-hours for serving | Rewards efficient use of shared pools | Needs a maintained weight table |
| Owner pays for reserved idle | Releases unused GPUs | Teams over-negotiate reservations down |
| Committed and on-demand rates | Shares commitment risk | More rate card complexity |
| Hard budget caps | Spend certainty | Self-inflicted outages |
What to do next
- Run showback for at least two cycles before moving money, and fix tagging gaps it exposes.
- Build the cost pool from the fully loaded GPU-hour and split it into training, serving and overhead pools.
- Pick one unit per pool, and derive token weights from measured attribution, not intuition.
- Set budget rates for the period from planned volume, publish the overhead uplift, and write the variance band.
- Write the idle-capacity and commitment-risk policies down before the first close.
- Implement an append-only ledger with idempotent event IDs, rate card versions and a conservation query.
- Publish draft statements, accept evidence-based disputes as forward adjustments, then close.
- Review weights, rates and the unallocated bucket every period, and watch for death-spiral signs.