Pricing a model call is arithmetic: tokens times a rate. Knowing what an agent costs is harder, because an agent decides at run time how many model calls to make, which tools to call, whether to delegate to another agent and whether to retry. Two requests that look identical to the user can differ in cost by a factor of twenty. Teams that track only the model bill find out about this at the end of the month, with no way to say which agent, which tool or which kind of request caused it.
This article treats cost as a property of an agent run. It explains why multi-step runs cost more than the sum of their steps, how to attribute spend across sub-agents and AgentTool child runs (where ADK Java has a blind spot by default), how to add tool and compute costs, how to enforce a money budget while the run is still going, and how to turn the ledger into cost per resolved task. Token fields, the versioned price table and reconciliation with your bill are covered in cost tracking with Gemini; this page builds on that plugin rather than repeating it. ADK details were read from the adk-java source in October 2026.
Why a run costs more than its calls
An LLM agent is a loop: call the model, run the tools it asks for, append the results to the conversation, call the model again. Every model call resends the instructions, the tool declarations and the whole conversation so far. So the input size of step k grows with k, and the total input over a run grows roughly with the square of the number of steps.
A worked example makes this concrete. Suppose the system instruction plus tool declarations are 3,000 tokens, the user message is 200, and each tool round adds 900 tokens (the function call plus its response). Then step k sends about 3,200 + 900(k - 1) input tokens:
| Steps in the run | Input tokens, last step | Input tokens, whole run | Relative to 1 step |
|---|---|---|---|
| 1 | 3,200 | 3,200 | 1x |
| 3 | 5,000 | 12,300 | 3.8x |
| 6 | 7,700 | 32,700 | 10.2x |
| 12 | 13,100 | 97,800 | 30.6x |
Doubling the number of steps nearly triples the input bill. That is why the most effective cost controls are about steps, not prices: fewer tool rounds, tools that return compact results, context caching for the stable prefix, and a hard cap on how long a run may loop. It is also why cost must be recorded per step with the step index: an average cost per call hides exactly the runs that matter.
The cost tree of one turn
ADK gives every callback the attribution keys you need for the common case. A CallbackContext or ToolContext exposes invocationId(), agentName(), branch(), userId() and sessionId(). Sub-agents reached by transfer or by workflow agents (sequential, parallel, loop) run inside the parent's invocation: same invocation id, a different agent name, and a branch that records the path. A plugin registered on the runner sees all of their model and tool calls, and grouping by agent name gives you the per-agent split.
The AgentTool blind spot
AgentTool is different, and it is the most common source of invisible spend in multi-agent ADK Java systems. When the parent model calls an agent wrapped as a tool, AgentTool builds a brand-new Runner for the wrapped agent, creates a fresh session for a user called tmp-user (seeded with a copy of the caller's state) and runs it. Three consequences follow from the source:
- Plugins are not inherited by default. The child runner gets the parent's plugin manager only when the tool was created with
includePluginsset to true; the default is false. A cost plugin on the root runner never sees the child's model calls. - It is a separate invocation. Even with plugins included, the child's callbacks report a new invocation id and the
tmp-useruser id, so a ledger keyed by invocation id or user id files the child's spend under a phantom user. - Call limits are per run. The child receives the caller's
RunConfig, so the samemaxLlmCallsapplies, but to the child's own fresh counter. A parent limited to 20 model calls that invokes a summariser tool five times can make 20 + 5 x 20 calls.
The fixes: create agent tools with AgentTool.create(agent, false, true) so the plugins run, and key the ledger on something the parent and child share. The trace id works, because the child runner's invocation span is started inside the tool's span and inherits its trace (confirm it in your version with a trace-shape test). As a fallback that does not depend on telemetry, write a root id into session state in the root's beforeRunCallback; the child session starts with a copy of that state.
A run-cost plugin
The plugin below extends the metering approach from the Gemini article with three things: a root key that survives AgentTool, tool and compute costs, and a running total that budget checks can read during the run. Prices come from the same versioned table; the per-tool rates are your own configuration.
public final class RunCostPlugin extends BasePlugin {
private final PriceTable prices; // versioned model rates
private final Map<String, BigDecimal> toolRates; // e.g. "web_search" -> 0.005 per call
private final CostLedger ledger; // async, batched writer
private final ConcurrentHashMap<String, RunTotal> open = new ConcurrentHashMap<>();
public RunCostPlugin(PriceTable prices, Map<String, BigDecimal> toolRates, CostLedger ledger) {
super("run_cost");
this.prices = prices; this.toolRates = toolRates; this.ledger = ledger;
}
static String rootKey() { // shared by parent and AgentTool child runs
return Span.current().getSpanContext().getTraceId();
}
private final ConcurrentHashMap<String, String> rootOf = new ConcurrentHashMap<>();
RunTotal total(ReadonlyContext ctx) {
String key = rootKey();
rootOf.putIfAbsent(ctx.invocationId(), key); // so the close path needs no current span
return open.computeIfAbsent(key, k -> new RunTotal(k, ctx.userId(), ctx.invocationId()));
}
@Override
public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
try {
if (resp.partial().orElse(false)) return Maybe.empty();
resp.usageMetadata().ifPresent(u -> {
BigDecimal cost = prices.costOf(resp.modelVersion().orElse("unknown"), u, LocalDate.now());
RunTotal t = total(ctx);
int step = t.nextStep();
t.add(cost);
ledger.offer(CostLine.model(t.key(), ctx.invocationId(), ctx.agentName(), step, u, cost));
});
} catch (RuntimeException e) { ledger.meteringError(e); }
return Maybe.empty(); // never alter the response
}
@Override
public Maybe<Map<String, Object>> afterToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx, Map<String, Object> result) {
try {
BigDecimal cost = toolRates.getOrDefault(tool.name(), BigDecimal.ZERO);
if (cost.signum() > 0) {
total(ctx).add(cost);
ledger.offer(CostLine.tool(rootKey(), ctx.invocationId(), ctx.agentName(), tool.name(), cost));
}
} catch (RuntimeException e) { ledger.meteringError(e); }
return Maybe.empty();
}
@Override
public Completable afterRunCallback(InvocationContext inv) {
return Completable.fromAction(() -> {
String key = rootOf.remove(inv.invocationId());
RunTotal t = key == null ? null : open.get(key);
if (t != null && t.ownedBy(inv.invocationId())) ledger.closeRun(open.remove(key));
});
}
}Two design points. The root total is closed only by the invocation that opened it, so a child run finishing first does not flush the parent's record; RunTotal remembers the first invocation id it saw. If no OpenTelemetry SDK is registered the trace id is invalid (all zeros) and every run would share one key, so check isValid() and fall back to the state-based root id. Also close it in onRunErrorCallback, because a failed run still paid for every call that completed. And code-execution or worker compute is priced the same way as a tool: measure seconds in the executor or worker and add a line at the rate you pay for that capacity.
Budgets that act during the run
Tracking tells you what happened. A budget changes what happens. ADK's model and tool callbacks can short-circuit: a beforeModelCallback that returns an LlmResponse skips the model call, and a beforeToolCallback that returns a map skips the tool. That makes an in-flight money budget a few lines in the same plugin:
@Override
public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder req) {
RunTotal t = open.get(rootKey());
// t.userId() is the root's user; in an AgentTool child ctx.userId() is "tmp-user".
if (t == null || t.spent().compareTo(budgets.perTurn(t.userId())) < 0) return Maybe.empty();
ledger.offer(CostLine.budgetStop(t.key(), ctx.agentName(), t.spent()));
return Maybe.just(LlmResponse.builder()
.content(Content.fromParts(Part.fromText(
"I have reached the processing limit for this request. Here is what I found so far; "
+ "ask me to continue if you need more.")))
.build());
}Keep three limits, each with a different job. A per-turn money budget in the plugin stops a runaway run gracefully. A per-tenant daily budget checked at turn admission (the pattern in rate limiting per user) stops one customer from spending everyone's quota. And RunConfig.maxLlmCalls is the backstop for bugs in the first two: it counts model calls in the invocation and throws LlmCallsLimitExceededException when exceeded. Its default is 500, which is far too high to protect a budget, so set it to a small multiple of your normal step count. Sub-agents share the counter with their parent; AgentTool children, as above, get their own. Token-denominated budgets and pre-flight token counts are covered in token counting across models.
Cost per outcome, and cost in CI
Cost per call and cost per turn are engineering numbers. The business number is cost per resolved task, because a cheap agent that fails and hands off to a human is the most expensive one you have. Join the ledger to an outcome label (resolved, escalated, abandoned), whether it comes from a user action, a ticket system or an evaluator, and report the distribution, not the mean:
| Request type | Turns | p50 cost per turn | p95 cost per turn | Resolution rate | Cost per resolution |
|---|---|---|---|---|---|
| Order status | 1.2 | $0.004 | $0.011 | 94% | $0.005 |
| Refund request | 3.1 | $0.019 | $0.090 | 71% | $0.083 |
| Product research | 4.6 | $0.041 | $0.310 | 58% | $0.33 |
The figures above are illustrative. The shape is typical: the p95 is several times the p50, the expensive tail lives in a small number of request types, and that is where step limits and better tools pay off. Put the cost figures in your evaluation test cases too: record the cost of each case on the current release and fail CI when a change raises the p95 cost of the suite by more than a set margin. A prompt edit that adds two tool rounds on average is a cost regression even when every answer is still correct.
Worked example: the invoice that was 2.4 times the dashboard
A product-research agent shows a flat model bill per request in its dashboard, yet the monthly invoice is 2.4 times the dashboard total. The root agent's plugin records about $0.12 per research turn. The architecture has a planner, a researcher sub-agent and a summariser wrapped in AgentTool that the researcher calls once per source.
Step one is to make the child runs visible: the team recreates the summariser tool with includePlugins true and switches the ledger key from invocation id to trace id. The next day's data explains the gap: each research turn calls the summariser six times on average, each child run makes two model calls with the full source document in context, and those calls account for 58 percent of turn cost. None of it had ever reached the dashboard.
The fixes follow from the tree. The summariser moves to a smaller model, receives an extract instead of the whole document, and is called once per turn with all sources. A per-turn budget of $0.40 stops the long tail, and maxLlmCalls drops from the default 500 to 30. Mean cost per research turn falls from $0.29 (the real figure) to $0.11, and the resolution rate, checked on the evaluation set, is unchanged.
Failure modes
- Invisible AgentTool spend. Child runs without plugins. Set
includePluginsand key on a shared root id. - Phantom users. Child spend attributed to
tmp-user. Attribute by root key, then take the user from the root record. - Averages hiding the tail. Cost per call looks flat while long runs grow quadratically. Record step index and report percentiles per request type.
- Errored runs dropped. Failed runs still paid. Close totals on error as well as success.
- Tools priced at zero. Paid APIs and sandbox compute left out of the ledger. Add a rate for every tool that costs money, even an estimate.
- A budget that throws. A budget check that raises an exception turns into a user-visible error. Return a graceful response from the callback and log the stop.
- Metering on the hot path. Synchronous ledger writes slow every model call. Queue and batch them.
Trade-offs
Precision or simplicity. In-process estimates are immediate but drift from the bill; reconcile daily rather than chasing exactness in the plugin. Hard or soft budgets. A hard stop protects margin but can cut off a turn one call from done; a soft budget that first switches to a cheaper model or asks the user to continue is kinder and slightly leakier. Shared or isolated child runs. Including plugins in agent tools gives visibility, but it also runs your guardrail and rate-limit plugins inside the child, which may double-count admission; decide per plugin whether it should act on child invocations. Granularity. Per-step lines cost storage; keep them for 30 days and roll up to per-turn records after that.
What to do next
- Draw your agent tree and mark every
AgentTool; recreate each withincludePluginstrue if you need its cost. - Key the cost ledger on a root id shared by parent and child runs (trace id, or a root id in session state).
- Record cost per step with the step index, agent name and model; add rates for paid tools and sandbox compute.
- Add a per-turn money budget in
beforeModelCallbackthat returns a graceful response. - Lower
RunConfig.maxLlmCallsfrom 500 to a small multiple of your normal step count. - Join the ledger to outcomes and report p50 and p95 cost per resolved task by request type.
- Add cost assertions to your evaluation suite and fail CI on p95 regressions.
- Reconcile the ledger with the provider bill weekly, and investigate any gap above a few percent.