ADK Java tells you how many tokens each Gemini call used. It does not tell you what that call cost, which user or tenant it should be charged to, or whether your numbers agree with the invoice. Those are the questions finance, product and on-call engineers actually ask, and they need a small piece of software between the token counts and the money: a price table, an attribution key, a ledger and a reconciliation job.
This page builds that piece. It explains what Gemini bills for, where the matching numbers appear in ADK Java, how to capture them once for every agent with a plugin, how to price them with a versioned table, and how to prove the ledger is right by reconciling it against billing data. Counting tokens before a call and enforcing token budgets are covered in token counting across models; this page starts where that one stops, at the money.
What a Gemini bill is made of
Prices change often, so treat any number you read, including the placeholders on this page, as something to load from configuration. The structure of a Gemini bill changes much less often. Per Google's Gemini API pricing page at the time of writing:
- Input tokens are charged per million at the model's input rate.
- Output tokens are charged at the output rate, and the pricing tables state that output price includes thinking tokens. On thinking models that is often the largest line.
- Context caching has two parts: a reduced per-token rate for cached input that is read, and a storage charge per million tokens per hour while the cache exists.
- Long prompts: some models have a higher tier for prompts above 200k tokens, and the tier applies to the request, not just the overflow.
- Batch API requests are listed at a 50% reduction.
- Grounding with Google Search is billed per search query performed, and one request can trigger several queries.
Two consequences follow. Storage and grounding charges do not appear in per-call token counts at all, so a token ledger alone will always undershoot the invoice. And the same token count can cost different amounts depending on model, tier and date, so the price lookup needs all three.
The numbers ADK Java gives you
After each model call ADK hands you an LlmResponse. Its usageMetadata() returns Optional<GenerateContentResponseUsageMetadata>, a type from the Google Gen AI Java SDK, and modelVersion() returns Optional<String>, the model that actually served the call. The usage fields are each Optional<Integer>:
| Field | Billing meaning | Notes |
|---|---|---|
promptTokenCount() | Input | The API reference says this includes cached content tokens |
cachedContentTokenCount() | Cached input | Subtract from prompt to get uncached input |
candidatesTokenCount() | Output | The visible answer |
thoughtsTokenCount() | Output | Separate field; the thinking docs price responses as output plus thinking tokens |
toolUsePromptTokenCount() | Input (verify) | Tokens from tool-use prompts; confirm by reconciling |
totalTokenCount() | Check value | Compare with the sum of the parts |
Do not hard-code the relationships beyond what the reference states. Record every raw field in the ledger, compute cost from them, and alarm when totalTokenCount() stops matching your assumed sum. A semantic change in the API then shows up as an alert rather than a quietly wrong bill.
Architecture: meter in a plugin, price off the hot path
Agent callbacks (afterModelCallback on an LlmAgent) work, but they attach to one agent; forget one sub-agent and its spend disappears. ADK Java plugins apply to every agent the runner executes, which is what a cost meter needs. A plugin extends com.google.adk.plugins.BasePlugin (constructor takes a name) and overrides only the hooks it needs. Three matter here: afterModelCallback(CallbackContext, LlmResponse) returning Maybe<LlmResponse>, and afterRunCallback(InvocationContext) and onRunErrorCallback(InvocationContext, Throwable), both returning Completable. The context gives the attribution keys for free: userId(), sessionId(), agentName() and invocationId().
A versioned price table
Use BigDecimal and rates per million tokens, keep every rate with the date it took effect, and never overwrite history: re-pricing last month with this month's table is the classic way ledgers drift. The values below are illustrative placeholders, not current Gemini prices.
public record Rate(BigDecimal inPerM, BigDecimal cachedInPerM, BigDecimal outPerM,
int longPromptThreshold, Rate longPromptRate) {}
public final class PriceTable {
private static final BigDecimal MILLION = BigDecimal.valueOf(1_000_000);
// model prefix -> rates by effective date. Loaded from config, never edited in place.
private final Map<String, NavigableMap<LocalDate, Rate>> rates;
public PriceTable(Map<String, NavigableMap<LocalDate, Rate>> rates) { this.rates = rates; }
Rate rateFor(String modelVersion, LocalDate day) {
String key = rates.keySet().stream()
.filter(modelVersion::startsWith)
.max(Comparator.comparingInt(String::length)) // longest prefix wins
.orElseThrow(() -> new IllegalStateException("no price for " + modelVersion));
Map.Entry<LocalDate, Rate> e = rates.get(key).floorEntry(day);
if (e == null) throw new IllegalStateException("no rate for " + key + " on " + day);
return e.getValue();
}
public BigDecimal cost(String model, LocalDate day, long prompt, long cached, long tool, long out) {
Rate r = rateFor(model, day);
if (r.longPromptRate() != null && prompt > r.longPromptThreshold()) r = r.longPromptRate();
long uncached = Math.max(0, prompt - cached) + tool;
return BigDecimal.valueOf(uncached).multiply(r.inPerM())
.add(BigDecimal.valueOf(cached).multiply(r.cachedInPerM()))
.add(BigDecimal.valueOf(out).multiply(r.outPerM()))
.divide(MILLION, 8, RoundingMode.HALF_UP);
}
}Failing loudly on an unknown model is deliberate. A fallback rate hides the day someone switches an agent to a new model; an exception in the pricing job, not in the agent, gets fixed the same afternoon.
The cost plugin
The plugin accumulates raw counts per invocation and hands a finished record to an asynchronous writer when the run ends. It returns Maybe.empty() so the response passes through unchanged, and it catches everything, because a metering bug must never fail a user request.
public final class CostPlugin extends BasePlugin {
private final ConcurrentHashMap<String, InvocationUsage> open = new ConcurrentHashMap<>();
private final CostWriter writer; // bounded queue + background batch insert
public CostPlugin(CostWriter writer) { super("cost_plugin"); this.writer = writer; }
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
try {
if (resp.partial().orElse(false)) return Maybe.empty(); // count final responses only
resp.usageMetadata().ifPresent(u -> {
InvocationUsage acc = open.computeIfAbsent(ctx.invocationId(),
id -> new InvocationUsage(id, ctx.userId(), ctx.sessionId()));
acc.add(new CallUsage(
ctx.agentName(),
resp.modelVersion().orElse("unknown"),
u.promptTokenCount().orElse(0),
u.cachedContentTokenCount().orElse(0),
u.toolUsePromptTokenCount().orElse(0),
u.candidatesTokenCount().orElse(0),
u.thoughtsTokenCount().orElse(0),
u.totalTokenCount().orElse(-1),
Instant.now()));
});
} catch (RuntimeException e) {
writer.recordMeteringError(e); // count it, alert on the rate
}
return Maybe.empty();
}
@Override
public Completable afterRunCallback(InvocationContext inv) {
return Completable.fromAction(() -> flush(inv.invocationId(), "ok"));
}
@Override
public Completable onRunErrorCallback(InvocationContext inv, Throwable error) {
return Completable.fromAction(() -> flush(inv.invocationId(), "error"));
}
private void flush(String invocationId, String outcome) {
InvocationUsage acc = open.remove(invocationId);
if (acc != null) writer.offer(acc.withOutcome(outcome)); // non-blocking; drops are counted
}
}Register it once on the runner: new InMemoryRunner(rootAgent, "support_app", List.of(new CostPlugin(writer))) for local work; the production Runner constructors take a plugin list the same way. The writer prices each call with PriceTable.cost using the call's own timestamp, and stores the raw counts next to the computed cost so a corrected price table can re-price history deliberately.
Two details that are easy to get wrong. Errored runs still consumed tokens for every call that completed, so onRunErrorCallback flushes too. And if a process dies mid-invocation the open map is lost; accept that loss, measure it through reconciliation, and keep the map small by evicting entries older than your longest plausible run.
Attribution: who pays for each call
Decide what each row is charged to before you have a month of rows. The context gives four keys: userId() and sessionId() for who and which conversation, agentName() for which step, and invocationId() to group calls into one user turn. Most products also need a tenant or plan, which ADK does not know about. Put it in session state when the session is created and read it in the plugin through ctx.state(), rather than deriving it later from user ids that may be reassigned.
Keep high-cardinality keys such as user and session in the ledger table, not in metric labels. A metrics system with one time series per user grows without bound; a table with a user column is just a table. Dashboards then aggregate by tenant, agent and model, and per-user questions become queries.
Worked example: one support invocation
Take a support agent on a thinking model with a cached 8,000-token system prompt and tool catalogue, and placeholder rates of 0.30 per million input, 0.03 per million cached input and 2.50 per million output. One invocation makes two model calls:
| Call | Prompt (cached) | Candidates | Thoughts | Cost (USD) |
|---|---|---|---|---|
| 1: plan and call tool | 12,000 (8,000) | 400 | 1,100 | 0.001200 + 0.000240 + 0.003750 = 0.005190 |
| 2: answer with tool result | 13,200 (8,000) | 250 | 300 | 0.001560 + 0.000240 + 0.001375 = 0.003175 |
| Invocation | 0.008365 |
At 100,000 invocations a day that is about 836 dollars a day at these placeholder rates. The breakdown is more useful than the total: thinking tokens are 1,400 of the 2,050 output tokens and about 42% of the invocation's cost, while the cached prefix, 16,000 of 25,200 prompt tokens, costs less than 6%. The first experiment is therefore a lower thinking budget on call 2, where the model is only phrasing an answer, not a larger cache.
Reconciling with the bill
A ledger you have never compared with the bill is an estimate. If your Gemini usage bills through a Google Cloud billing account, enable the billing export to BigQuery and run a daily job that compares the ledger's total per model with the exported cost for the same day and service. Expect a gap and explain it:
- Storage charges for context caches and per-query grounding charges, which never appear in token counts.
- Other callers on the same key or project, such as notebooks, evaluations and other services. A dedicated project or key per application removes most of this.
- Calls made outside the runner, for example direct
Clientcalls for embeddings or token counts. - Lost invocations from crashed processes, and writer queue drops.
- Time-zone and late-arriving billing data; compare with a lag of a day or more.
Alert when unexplained drift exceeds a threshold you choose, for example a few percent, and treat a step change as a bug report. Wire per-user and per-tenant totals into dashboards and budget alerts as described in metrics dashboards for ADK Java.
Failure modes
- Double counting streamed responses. Partial chunks can carry usage. Count only non-partial responses, and verify by comparing a streaming and a non-streaming run of the same prompt.
- Pricing cached tokens twice.
promptTokenCountalready includes cached tokens. Charging the full prompt at the input rate and the cached part again overstates cost. - Ignoring thinking tokens. Pricing only
candidatesTokenCountcan undercount output badly on thinking models. - Blocking in the callback. A synchronous database insert in
afterModelCallbackadds its latency to every model call and fails requests when the database is down. - Agent-level callbacks only. A sub-agent without the callback is invisible spend; use a plugin.
- Mutable prices. Editing a rate in place silently re-prices history and breaks reconciliation.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| Plugin vs agent callbacks | Every agent covered, one place to change | Applies to all agents; filter if you need exceptions |
| Price at write time | Simple queries, cost visible immediately | Re-pricing needs a backfill job |
| Store raw counts too | Re-pricing and audits possible | Wider rows |
| Async writer | Zero added latency | Small, measurable loss on crashes |
| Per-call rows vs per-invocation rows | Find the expensive step | More rows to store |
Tracing each call with cost as a span attribute is a good complement to the ledger; see observability for ADK Java and ADK Java callbacks for the hook mechanics.
What to do next
- Write down every model your agents use and load current rates, with effective dates, from Google's pricing page into a config file.
- Add the
CostPluginto your runner and log the raw usage fields for a day before pricing anything. - Check
totalTokenCountagainst the sum of the parts for each model and record what the sum is. - Reproduce the worked example with your own traffic and find which call and which token type dominates.
- Turn on the billing export, run the reconciliation query daily, and set a drift alert.
- Add per-tenant budgets with alerts at 50%, 80% and 100% of the monthly allowance.