Chargeback is the practice of billing internal teams for the GPU capacity they consume, with the money actually moving between cost centres. It differs from showback, which reports the same numbers without moving money, and from cost attribution, which is the measurement problem of deciding which request used which share of a shared GPU step. The cost attribution article solves the measurement; the FinOps overview explains why organisations move from showback to chargeback at all. This article is about the billing machinery in between: what goes into the cost pool, which units to bill, how to set rates, who pays for idle GPUs, how the monthly close works and which behaviours each design rewards.

The test of a chargeback system is simple to state. Every dollar the fleet costs in a month must land in exactly one place: a team's statement, a central budget, or a variance account that someone owns. If the books do not balance, the system loses credibility in the first quarter and teams start disputing every line.

Cost pools and overhead

Start with the cost pool: everything the fleet costs per month. That is hardware depreciation or cloud spend, power and cooling, data centre space, network, storage attached to the serving tier, software licences and the operations team. The infrastructure cost article builds this into a fully loaded cost per GPU-hour, and that number is the input here.

Then split the pool by service, because different services need different units. A typical split has three pools: dedicated training reservations, a shared inference pool, and platform overhead such as spare GPUs, burn-in nodes and the control plane. Overhead is not billed directly; it is spread over the billable pools as an uplift on their rates. Make the uplift visible on the rate card so teams can see what the platform costs and argue about it.

The pipeline

The chargeback pipeline: from metered usage to a closed monthly ledgerMeteringusage events, idempotentAttributiontenant, model, tierRatingunits x rate cardLedgerappend-only chargesRate cardset per periodCost poolfully loaded fleet costMonthly closeconservation checkStatementsper team, journal entriesDisputesadjustments onlyVariance accountover and under recoveryCorrections are new ledger rows, never edits. The close proves charges plus variance equal the pool.
Metering feeds attribution and rating; the ledger is append-only; the close proves charges plus variance equal the cost pool, and disputes come back as adjustments.

Choosing billable units

A billable unit must be measurable without dispute, roughly proportional to cost, and something the team can control. No single unit satisfies all three for LLM workloads, so most rate cards have several.

UnitFitsBehaviour it rewardsWeakness
Reserved GPU-hourTraining, dedicated servingReleasing unused reservationsCharges idle reserved GPUs to the owner, which is correct but unpopular
MIG slice-hourSmall models, dev environmentsRight-sizing to a sliceSlice profiles differ per GPU generation
Weighted token unitShared serving poolCaching, shorter outputs, smaller modelsWeights must track measured cost
Priority tier multiplierMixed latency classesUsing batch tier for offline workNeeds real scheduling difference to be fair

Weighted token units are the usual choice for shared serving. Bill input, cached input and output tokens at different weights per model, with the weights taken from measured attribution rather than intuition: if the attribution pipeline shows an output token on a given model costs four times an uncached input token, the weight is four. Revisit weights when the engine or hardware changes, because prefix caching, speculative decoding and quantisation all move the ratios. For MIG partitions, bill the slice-hour regardless of activity, exactly like a reservation.

Setting rates

Rates are set once per period, usually a quarter or a year, from a budget, and held fixed. That is the core idea of budget rates: teams can forecast their bill, and the platform team carries the risk that volume differs from plan. The recovery equation for a pool is:

rate = (pool_cost x (1 + overhead_uplift)) / planned_billable_units

recovered       = actual_units x rate
variance        = pool_cost x (1 + overhead_uplift) - recovered
                  (positive = under-recovery, negative = over-recovery)

Floating rates, recomputed monthly from actual volume, recover cost exactly but make the price depend on other teams' behaviour. A team that cut its usage in half can see its bill fall by less than expected because everyone's rate rose. Budget rates avoid that, at the cost of a variance account that must be settled.

from decimal import Decimal, ROUND_HALF_UP

def budget_rate(pool_cost, uplift, planned_units, quantum=Decimal("0.01")):
    raw = pool_cost * (1 + uplift) / planned_units
    return raw.quantize(quantum, rounding=ROUND_HALF_UP)

def rate_usage(events, rate_card):
    # events: dicts with tenant, unit, quantity; rate_card maps unit to Decimal rate.
    charges = {}
    for e in events:
        amount = Decimal(e["quantity"]) * rate_card[e["unit"]]
        charges[e["tenant"]] = charges.get(e["tenant"], Decimal("0")) + amount
    return charges

Use decimal arithmetic for money. Binary floating point accumulates rounding error over millions of usage events, and a statement that is off by four cents is a dispute.

Who pays for idle capacity

Idle capacity is where chargeback systems go wrong most often, because it has no requester. There are four defensible policies, and the choice should be explicit.

PolicyWho pays for idleEffect
Reservation owner paysThe team holding the reservationStrong incentive to release capacity; fits dedicated pools
Spread in the rateEvery user of the pool, via the budget rateSimple; hides idle cost and weakens pressure to fill the pool
Central budget absorbsThe platform or financeEncourages adoption; idle cost becomes invisible to users
Strategic headroom lineA named budget for surge and failoverMakes the redundancy decision visible and owned

A common and defensible combination: reservation owners pay for their reserved hours whether used or not; the shared pool's planned idle is in its budget rate; spare and burn-in GPUs are overhead uplift; and failover headroom sized by the capacity analysis is a strategic headroom line owned by whoever set the SLO.

Passing through commitment risk

If the fleet runs on committed cloud capacity or purchased hardware, someone carries the commitment risk: the obligation to pay whether or not the GPUs are used. The reserved capacity article covers sizing the commitment. Chargeback decides who inherits it. Passing it through means teams sign internal commitments matching the external one, and pay for them regardless of use; that is accurate but makes teams reluctant to commit. Absorbing it centrally makes adoption easy but leaves the platform with a variance it cannot control. A middle path offers two rates: a lower committed rate for teams that sign up for a quantity for the period, and a higher on-demand rate for everything else, with the spread paying for the risk the platform keeps.

The ledger

Store charges in an append-only ledger. Each usage event carries an idempotency key from the meter, so a replayed batch cannot double-bill. Each charge row records the rate card version it used, so a statement can be reproduced exactly a year later. Closed periods are immutable: corrections are new adjustment rows in the next open period that reference the original.

CREATE TABLE usage_event (
  event_id     TEXT PRIMARY KEY,          -- idempotency key from the meter
  period       TEXT NOT NULL,             -- e.g. 2026-09
  tenant       TEXT NOT NULL,             -- 'unallocated' if the tag was missing
  unit         TEXT NOT NULL,             -- reserved_gpu_hour, wtu_million, ...
  quantity     NUMERIC NOT NULL
);

CREATE TABLE charge (
  charge_id    BIGSERIAL PRIMARY KEY,
  period       TEXT NOT NULL,
  tenant       TEXT NOT NULL,
  unit         TEXT NOT NULL,
  quantity     NUMERIC NOT NULL,
  rate         NUMERIC NOT NULL,
  rate_card    TEXT NOT NULL,             -- version id, e.g. FY27-Q1
  amount       NUMERIC NOT NULL,
  kind         TEXT NOT NULL DEFAULT 'usage',   -- usage | adjustment
  ref_charge   BIGINT REFERENCES charge(charge_id)
);

-- Conservation check run at close: must return exactly zero.
SELECT p.cost_with_uplift
     - (SELECT COALESCE(SUM(amount), 0) FROM charge WHERE period = p.period)
     - p.variance_booked AS imbalance
FROM pool_period p WHERE p.period = '2026-09';

Track the unallocated tenant as a first-class line. Usage with a missing or invalid team tag lands there and is reported to the platform team weekly. If it exceeds a few percent of the pool, fix tagging at admission (reject untagged API keys) before the close, not after.

Worked example: one monthly close

Consider a fleet of 256 GPUs with a fully loaded cost of 2.40 dollars per GPU-hour, an illustrative figure; use your own from the infrastructure cost model. A 730-hour month costs 256 x 730 x 2.40 = 448,512 dollars. The split is 128 GPUs of training reservations, 112 GPUs in the shared serving pool and 16 GPUs of spares and burn-in.

The 16 overhead GPUs cost 16 x 730 x 2.40 = 28,032 dollars, spread over the 240 billable GPUs as an uplift of 16 / 240 = 6.67 percent. The loaded rate is 2.56 dollars per GPU-hour. Training reservations: team A holds 64 GPUs, B holds 48 and C holds 16, billed 119,603.20, 89,702.40 and 29,900.80 dollars, together 239,206.40, whether or not the GPUs were busy.

The serving pool costs 112 x 730 x 2.56 = 209,305.60 dollars with uplift. The plan expected 52.3 billion weighted token units, giving a rate of 4.00 dollars per million units. Actual volume was 47.0 billion: search used 18.2 billion (72,800 dollars), the support assistant 21.5 billion (86,000) and other teams 7.3 billion (29,200), recovering 188,000 dollars.

LineAmount (USD)
Fleet cost pool448,512.00
Training reservations billed239,206.40
Serving usage billed188,000.00
Serving under-recovery to variance21,305.60
Check: billed + variance448,512.00

The books balance, and the close surfaces one decision: serving under-recovered by 10.2 percent because volume came in below plan. If the policy is that the platform absorbs variance within plus or minus 10 percent and anything beyond triggers a quarterly true-up, this month is just outside the band. Note what the platform should not do: raise the serving rate mid-quarter. That punishes the teams who are using the pool to cover the plan error of the teams who are not.

Disputes, true-ups and budgets

Publish draft statements several working days before the close and accept disputes only with evidence: event IDs, a time range, a reason. A dispute resolves into an adjustment row, positive or negative, in the next open period, referencing the original charges. A true-up settles the variance account at the end of the period, either by a pro-rata credit or debit to pool users, or by a central write-off. Agree which in advance. True-ups decided after the numbers are known become negotiations.

Budgets close the loop. Give each team a monthly budget per pool, alert at 50, 80 and 100 percent of forecast spend, and decide in advance what happens at 100: a soft cap that notifies a manager, a move to the batch priority tier, or a hard cap that rejects new requests. Hard caps on customer-facing services cause outages, so reserve them for development and experimentation keys.

Failure modes

  • The death spiral. Fixed costs spread over falling volume raise the rate, which drives more teams to external APIs, which raises the rate again. Hold budget rates for the period and fund strategic headroom centrally.
  • Floating rates. Bills move with other teams' behaviour; nobody can forecast, and trust erodes.
  • Weights that no longer track cost. After enabling prefix caching, cached input billed at full price punishes the teams who made caching work.
  • Editable history. Changing a closed month's charges breaks every reconciliation downstream. Adjust forward only.
  • Growing unallocated bucket. Untagged usage that nobody pays for hides real demand. Enforce tags at admission.
  • Double counting. Billing both the reservation and the tokens served on it charges the same GPU-hour twice.

Trade-offs

ChoiceGainCost
Budget rates vs floatingPredictable billsA variance account to settle
Token units vs GPU-hours for servingRewards efficient use of shared poolsNeeds a maintained weight table
Owner pays for reserved idleReleases unused GPUsTeams over-negotiate reservations down
Committed and on-demand ratesShares commitment riskMore rate card complexity
Hard budget capsSpend certaintySelf-inflicted outages

What to do next

  1. Run showback for at least two cycles before moving money, and fix tagging gaps it exposes.
  2. Build the cost pool from the fully loaded GPU-hour and split it into training, serving and overhead pools.
  3. Pick one unit per pool, and derive token weights from measured attribution, not intuition.
  4. Set budget rates for the period from planned volume, publish the overhead uplift, and write the variance band.
  5. Write the idle-capacity and commitment-risk policies down before the first close.
  6. Implement an append-only ledger with idempotent event IDs, rate card versions and a conservation query.
  7. Publish draft statements, accept evidence-based disputes as forward adjustments, then close.
  8. Review weights, rates and the unallocated bucket every period, and watch for death-spiral signs.
Key takeaway: Chargeback works when every dollar of the fleet lands in exactly one place and the price teams see is predictable. Build a cost pool from the fully loaded GPU-hour, bill reservations by the hour and shared serving by weighted tokens, hold budget rates for the period, write the idle and commitment policies down, keep an append-only ledger with a conservation check, and settle variance by rules agreed before the numbers arrive.