An agent platform that serves many customers eventually has to send each of them a bill. This article is the build: what a billable event looks like when it comes out of an ADK for Java runner, how it travels to a ledger without being lost or counted twice, how a versioned price book turns usage into line items, how a billing period closes when events arrive late, and how you prove the invoice matches what actually happened.
Two neighbouring articles cover adjacent problems, and this one deliberately does not repeat them. Cost tracking with Gemini answers what the model calls cost you; quota management answers whether a tenant may make the next call. Billing answers a third question: what does the tenant owe you, and can you defend the number line by line.
Cost, quota and billing are three systems
Teams that skip this distinction ship one table that tries to serve all three purposes and then cannot change prices without corrupting history. The three systems read the same raw usage but have different correctness requirements:
| System | Question | Latency need | Correctness need |
|---|---|---|---|
| Cost tracking | What did this cost us? | Minutes to hours | Approximate; reconciled monthly against the cloud bill |
| Quota | May this call proceed? | Milliseconds, in the request path | Conservative; over-reserving is acceptable |
| Billing | What does the tenant owe? | Hours to days | Exact, reproducible, auditable, never double-counted |
The practical consequence: billing must never read the quota counters or the cost table. Quota counters are reset, reserved and released; cost figures change when your provider changes prices. Billing reads only the append-only usage log and a price book that the tenant agreed to, and both are versioned.
The pipeline end to end
The runner records facts: this tenant, this invocation, this many prompt tokens, this tool call. It never records prices. A rating job reads facts and a specific price book version and produces rated line items. A close step freezes a period. Only then does anything go to the payment provider. Because each stage is a pure function of its inputs plus a version number, you can re-run rating for a disputed month and get the same answer, which is the property an auditor or an angry customer will test first.
What to meter and the event schema
Meter what the tenant can understand and you can verify. For an LLM agent that usually means prompt tokens, cached prompt tokens, output tokens (including thinking tokens if your model reports them separately and you price them), tool invocations by tool name, and possibly sessions or stored bytes. Do not meter something you cannot reproduce from logs, such as wall-clock time on a shared worker.
Every event needs a deterministic identity so that retries collapse into one row. A random UUID generated at emit time defeats this: a retried emit gets a new UUID and bills twice. Build the id from things that are stable across retries.
public record UsageEvent(
String eventId, // deterministic: invocation + agent + kind + sequence
String tenantId, // from trusted session state, never from model output
String meter, // "prompt_tokens", "output_tokens", "tool_call:search", ...
long quantity, // integer units; never a float
Instant occurredAt, // when the usage happened, not when it was written
String invocationId,
String agentName,
String modelVersion) {
public static String id(String invocationId, String agent, String meter, int seq) {
return invocationId + "/" + agent + "/" + meter + "/" + seq;
}
}Quantities are integers in the smallest unit you bill. Money does not appear in the event at all. The occurredAt field drives which billing period the event belongs to; the write time only matters for lateness monitoring.
Capturing usage in an ADK Java plugin
ADK for Java exposes a Plugin interface, with BasePlugin as the class you extend, and plugins registered on the runner see every agent, model and tool callback for every invocation. That makes a plugin the right place to meter: one class covers every agent in the app, and no agent author can forget to call it. The callbacks used below, afterModelCallback(CallbackContext, LlmResponse), afterToolCallback(BaseTool, Map, ToolContext, Map) and afterRunCallback(InvocationContext), are defined on that interface with default no-op implementations. See ADK Java callbacks for how they order relative to agent-level callbacks.
public final class BillingPlugin extends BasePlugin {
private final UsageOutbox outbox; // durable, batched writer
private final ConcurrentHashMap<String, AtomicInteger> seq = new ConcurrentHashMap<>();
public BillingPlugin(UsageOutbox outbox) { super("billing"); this.outbox = outbox; }
private int next(String invocationId, String agent, String meter) {
return seq.computeIfAbsent(invocationId + "/" + agent + "/" + meter,
k -> new AtomicInteger()).getAndIncrement();
}
private static String tenant(CallbackContext ctx) {
Object t = ctx.state().get("billing_tenant"); // set by your API layer
if (t == null) throw new IllegalStateException("session without billing_tenant");
return t.toString();
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
if (resp.partial().orElse(false)) return Maybe.empty(); // streamed chunks: count once
resp.usageMetadata().ifPresent(u -> {
String model = resp.modelVersion().orElse("unknown");
long prompt = u.promptTokenCount().orElse(0); // includes cached tokens
long cached = u.cachedContentTokenCount().orElse(0);
emit(ctx, "prompt_tokens", Math.max(0, prompt - cached), model);
emit(ctx, "cached_prompt_tokens", cached, model);
emit(ctx, "output_tokens", u.candidatesTokenCount().orElse(0), model);
});
return Maybe.empty();
}
@Override
public Maybe<Map<String, Object>> afterToolCallback(BaseTool tool, Map<String, Object> args,
ToolContext ctx, Map<String, Object> result) {
emit(ctx, "tool_call:" + tool.name(), 1, "");
return Maybe.empty(); // never alter the result
}
@Override
public Completable afterRunCallback(InvocationContext inv) {
seq.keySet().removeIf(k -> k.startsWith(inv.invocationId() + "/"));
return Completable.complete();
}
@Override
public Completable onRunErrorCallback(InvocationContext inv, Throwable error) {
return afterRunCallback(inv); // failed runs still billed so far
}
private void emit(CallbackContext ctx, String meter, long qty, String model) {
if (qty <= 0) return;
int n = next(ctx.invocationId(), ctx.agentName(), meter);
outbox.append(new UsageEvent(UsageEvent.id(ctx.invocationId(), ctx.agentName(), meter, n),
tenant(ctx), meter, qty, Instant.now(), ctx.invocationId(), ctx.agentName(), model));
}
}Four details carry the correctness. The prompt meter is net of cached tokens, because Gemini's prompt count already includes them; billing both raw would double count. Partial streaming responses are skipped so a streamed answer is counted once, from the final response that carries usage metadata; verify that your model and streaming mode actually report usage on the final chunk before trusting this. The sequence number makes ids deterministic within an invocation, so if the outbox write is retried the id repeats and the log deduplicates it. And the plugin returns Maybe.empty() everywhere, so billing can never change what the agent does.
The outbox should be a local durable queue or a table written by a background thread, not a synchronous call to a billing service. If billing is down, agents keep serving and usage accumulates; if the process dies, unflushed events are the risk, so flush on a short timer and alert when the outbox depth grows. Persist the log keyed by eventId with an insert-if-absent, as in an append-only Postgres event log.
Where the tenant id comes from
The tenant id is the single most dangerous field. If it can be influenced by the model, by a tool result or by a client-supplied header that is not authenticated, one tenant can push usage onto another. Set billing_tenant in session state when your API layer creates the session, from the authenticated principal, and treat the plugin's IllegalStateException as a bug to fix rather than a case to default. A default tenant such as "unknown" quietly becomes revenue you can never bill.
When an agent calls another tenant's agent, decide per integration who is billed and encode it in the meter name.
Rating: from usage to money
Rating converts usage totals for a period into money using the plan the tenant was on. The price book is configuration with an effective date and a version, loaded read-only. Typical plan features are a base fee, an included allowance per meter, a per-unit overage price, optional tiers, and a minimum commitment. Use BigDecimal or integer micro-units for money and round once, at the line item, with a documented rounding mode.
public record MeterPrice(long included, BigDecimal perUnit, long unitSize) {}
public record Plan(String id, int version, BigDecimal baseFee, Map<String, MeterPrice> meters) {}
public record LineItem(String tenantId, YearMonth period, String meter,
long quantity, long billable, BigDecimal amount, String planRef) {}
public static List<LineItem> rate(String tenant, YearMonth period, Plan plan,
Map<String, Long> totals) {
List<LineItem> items = new ArrayList<>();
String ref = plan.id() + "@v" + plan.version();
items.add(new LineItem(tenant, period, "base", 1, 1, plan.baseFee(), ref));
for (var e : new TreeMap<>(totals).entrySet()) { // sorted: stable output
MeterPrice mp = plan.meters().get(e.getKey());
if (mp == null) throw new IllegalStateException("unpriced meter " + e.getKey());
long billable = Math.max(0, e.getValue() - mp.included());
BigDecimal amount = BigDecimal.valueOf(billable)
.multiply(mp.perUnit())
.divide(BigDecimal.valueOf(mp.unitSize()), 2, RoundingMode.HALF_UP);
items.add(new LineItem(tenant, period, e.getKey(), e.getValue(), billable, amount, ref));
}
return items;
}An unpriced meter throws rather than rating to zero. New tools appear in agents all the time; a silent zero means a feature shipped free for months before anyone noticed.
Worked example. Tenant acme is on plan pro v3: a 500.00 base fee; 50,000,000 uncached prompt tokens included, then 4.00 per million; cached prompt tokens at 1.00 per million; 5,000,000 output tokens included, then 16.00 per million; 10,000 tool_call:search calls included, then 2.00 per thousand. September totals: 62,400,000 uncached and 8,000,000 cached prompt tokens, 7,100,000 output tokens, 14,250 searches. Lines: 49.60, 8.00, 33.60 and 8.50, so the invoice is 599.70, and every line carries pro@v3 so the calculation can be repeated next year with the same inputs.
The prices here are illustrative, not anyone's list price.
Closing a period with late events
Usage events arrive late: an outbox drains after a network partition, a batch job replays a day. A period therefore closes in two steps. First a watermark: rating for September does not run until the minimum outbox position across all writers has passed the end of September plus a grace window, for example 48 hours. Second a freeze: the rated line items for the period are written with a closed flag and never updated.
Events whose occurredAt falls in a frozen period are not dropped and not rewritten into it. They are rated into an adjustment line on the next open period, labelled with the period they belong to. Credits, goodwill refunds and dispute outcomes take the same path: a new signed line item with a reason code and an approver. A mid-period plan change splits the period at the change timestamp; rate each part with its own plan version and prorate the base fee.
Handing off to a payment provider
Most teams hand invoicing, tax and collection to a payment provider. With Stripe's usage-based billing you create meter events by POSTing to /v1/billing/meter_events with an event_name, a payload carrying the customer id and value, an optional identifier and an optional timestamp. Stripe documents that the identifier is unique within a rolling period of at least 24 hours and that the timestamp must be within the past 35 calendar days or at most 5 minutes in the future. Those two limits shape the integration: the identifier protects against fast retries, not against a replay a week later, so your own log must stay the deduplication authority; and usage older than 35 days cannot be sent as a meter event, so late adjustments need another path such as an invoice item.
A design that sends aggregated, already-rated quantities per tenant per day is easier to reconcile than streaming every raw token event. Send identifier as tenant, meter and day, so a re-sent daily total collapses inside the provider's window and is caught by your reconciliation outside it.
Reconciliation
Reconciliation is the test suite for money. Run it daily and at every close, comparing three numbers per tenant per meter: the raw usage log sum, the rated ledger quantity, and the quantity the provider reports. Each pair catches a different bug: log versus ledger catches rating that skipped or duplicated rows; ledger versus provider catches send failures and double sends. Add a fourth, coarser check against your model provider's bill: total tokens across all tenants should track your own cost tracker within a small tolerance, and a gap means usage that was consumed but never metered.
SELECT u.tenant_id, u.meter,
u.qty AS logged, l.qty AS rated, p.qty AS provider
FROM (SELECT tenant_id, meter, SUM(quantity) qty FROM usage_log
WHERE occurred_at >= :start AND occurred_at < :end GROUP BY 1, 2) u
LEFT JOIN rated_lines l USING (tenant_id, meter)
LEFT JOIN provider_snapshot p USING (tenant_id, meter)
WHERE u.qty IS DISTINCT FROM l.qty OR l.qty IS DISTINCT FROM p.qty;
Failure modes
- Double counting from streaming. Partial responses metered alongside the final one multiply token counts. Count only non-partial responses and test with streaming enabled.
- Random event ids. Retries bill twice. Ids must be derived from invocation, agent, meter and sequence.
- Tenant from untrusted input. Usage lands on the wrong customer. Bind it at session creation from the authenticated principal.
- Editing a closed period. Last month's invoice silently changes. Post adjustments instead.
- Silent zero for a new meter. A new tool ships unbilled. Fail rating on unpriced meters.
- Outbox loss on crash. Flush often and alarm on outbox depth.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Bill raw tokens | Transparent, tracks your cost | Hard for buyers to forecast; exposes model churn |
| Bill credits | One number; you can rebalance weights | Credit weights become a pricing surface you must version |
| Bill per task or outcome | Matches customer value | Disputes over what counts as done; margin risk on long tasks |
| Provider meter events | Tax, dunning, invoices handled | Retention and timestamp limits; your log still needs dedup |
What to do next
- Set
billing_tenantin session state from the authenticated caller and make a missing value fail loudly. - Register a metering plugin that emits deterministic ids, and test it with streaming on and a forced retry of the outbox write.
- Write the price book as versioned configuration and stamp the plan reference on every line.
- Implement the watermark and freeze, and route late usage and credits to adjustment lines.
- Schedule the three-way reconciliation query daily and alert on any non-empty result.
- Rate last month twice from scratch and diff the outputs; any difference is a bug.