Knowing what your agent costs in total is the easy half. The hard half is splitting that total by agent, so engineering knows which component to optimise, and by user and tenant, so the business knows who is profitable and who is abusing a flat plan. In a single-agent service the split is a column in a ledger. In a tree of agents, with supervisors, sub-agents called as tools, parallel branches, retries and costs shared by everyone, it is an allocation problem. Most of its errors are silent.

This article treats attribution as allocation for an ADK Java agent tree. It shows which model calls a Runner plugin actually sees, why sub-agents wrapped in AgentTool are invisible by default, and how to carry attribution keys into nested runs. It then covers tool spend and retries, a policy for shared costs, per-user rollups that do not blow up your metrics system, and the conservation check that keeps the ledger honest. Pricing a single call, with a versioned price table and a cost plugin, is covered in cost tracking with Gemini via ADK Java. This page starts where that one stops.

The unit and the conservation rule

Fix two definitions before writing code. The unit is one billable event: a model call, a paid tool call, or a slice of a shared cost. Each unit becomes one ledger row carrying tenant, user, root invocation, calling agent, agent, kind and cost. The conservation rule says that for any period, the ledger rows, including an explicit unattributed bucket, sum to the provider bills for that period. Every dollar lands in exactly one row.

Attribution errors rarely throw: a missing sub-agent just makes one agent look cheap. The conservation check turns that into a growing unattributed bucket someone must explain. Self-hosted serving has the same rule one layer down, where a batched GPU step has to be split among requests. See LLM cost attribution on GPUs.

What a Runner plugin sees, and what it misses

One user turn, two runners: what the root plugin sees by defaultRoot invocationuser u_481, tenant t_acmeRoot Runner (plugins registered here)concierge (LlmAgent)2 model calls, 0.00409 USDsub_agents / ParallelAgentsame invocation, seen by the pluginNested Runner built by AgentToolresearch_agent3 model calls, 0.00825 USDweb_search tool2 calls, 0.01000 USDAgentTooluserId = "tmp-user", new session id,appName = calling agent's name;plugins only if includePlugins = trueVisible by default: 0.00409 of 0.02234 USD(18% of the turn's true direct cost)Placeholder rates: 0.30 / 0.03 / 2.50 USD per million input / cached / output tokens; search 0.005 USD per call.
Sub-agents run inside the root invocation and reach the plugin. An AgentTool runs a nested Runner whose calls are invisible by default, and are charged to a placeholder user once plugins are included.

Plugins are registered on a Runner and receive callbacks for every agent that runs inside that runner's invocations. ADK Java composes agents in two different ways, and they behave differently for attribution.

  • Sub-agents and workflow agents (subAgents with transfer, SequentialAgent, ParallelAgent, LoopAgent) run inside the same invocation. The plugin's afterModelCallback fires for each of their model calls with the same invocationId() and userId(). agentName() names the agent that made the call, and branch() separates parallel branches. Attribution works with no extra effort, provided the ledger writer is thread-safe, because parallel branches call back concurrently.
  • AgentTool wraps an agent as a tool. In the current adk-java source, its runAsync builds a new, nested Runner for each call. That runner's app name is the calling agent's name, its session is new and seeded with a copy of the caller's session state, and its user id is the literal "tmp-user". The caller's plugins are passed to it only when the tool was created with includePlugins set to true. Every factory and constructor that does not take the flag sets it to false.

So by default, a supervisor that delegates research to an AgentTool records its own two model calls and nothing of the research agent's, including the paid tools the research agent calls. If you turn includePlugins on and stop there, the calls appear but are charged to a user named tmp-user under invocation ids no user turn owns. Both outcomes pass every test that checks only that the cost plugin runs. This is an implementation detail, so verify it in the ADK version you deploy.

Carrying attribution keys into nested runs

The nested session receives a copy of the caller's state. That is the channel for attribution keys. Write the tenant into session state when the session is created. Then, just before an AgentTool runs in the root runner, stamp the real user and the root invocation id. The nested run inherits them, and so does any nested run inside it. To tell the root run from a nested one, compare the invocation context's app name with your real app name, which nested runs never have unless you named an agent after your app.

public final class AttributionPlugin extends BasePlugin {
  static final String TENANT = "attr_tenant", USER = "attr_user", ROOT_INV = "attr_root_inv";
  private final String rootApp;
  private final CostLedger ledger;   // prices rows with the versioned table, writes off the hot path

  public AttributionPlugin(String rootApp, CostLedger ledger) {
    super("attribution"); this.rootApp = rootApp; this.ledger = ledger;
  }

  private boolean isRoot(ReadonlyContext ctx) {
    return rootApp.equals(ctx.invocationContext().appName());
  }

  private Attribution attribution(ReadonlyContext ctx) {
    boolean root = isRoot(ctx);
    return new Attribution(
        key(ctx, TENANT),
        root ? ctx.userId() : key(ctx, USER),
        root ? ctx.invocationId() : key(ctx, ROOT_INV),
        ctx.invocationContext().appName(),      // caller: rootApp, or the agent that used AgentTool
        ctx.agentName(),
        ctx.invocationContext().branch().orElse(""));
  }

  @Override public Maybe<Map<String, Object>> beforeToolCallback(
      BaseTool tool, Map<String, Object> args, ToolContext ctx) {
    if (tool instanceof AgentTool && isRoot(ctx)) {   // state is copied into the nested session
      ctx.state().put(ROOT_INV, ctx.invocationId());
      ctx.state().put(USER, ctx.userId());
    }
    return Maybe.empty();
  }

  @Override public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse r) {
    if (!r.partial().orElse(false))
      r.usageMetadata().ifPresent(u ->
          ledger.modelCall(attribution(ctx), r.modelVersion().orElse("unknown"), u));
    return Maybe.empty();
  }

  @Override public Maybe<Map<String, Object>> afterToolCallback(BaseTool tool,
      Map<String, Object> args, ToolContext ctx, Map<String, Object> result) {
    if (!(tool instanceof AgentTool))                // its cost is the nested rows
      ledger.toolCall(attribution(ctx), tool.name());  // priced from a per-tool rate card
    return Maybe.empty();
  }

  private static String key(ReadonlyContext ctx, String k) {
    Object v = ctx.state().get(k);
    return v == null ? "unattributed" : v.toString();
  }
}

// Wiring: opt the nested runner into the caller's plugins.
AgentTool research = AgentTool.create(researchAgent, false, /* includePlugins= */ true);
LlmAgent concierge = LlmAgent.builder().name("concierge").model(MODEL)
    .instruction(CONCIERGE_INSTRUCTION).tools(research, billingLookup).build();
Runner runner = Runner.builder().agent(concierge).appName("support_app")
    .sessionService(sessions).plugins(new AttributionPlugin("support_app", ledger)).build();
// Session creation writes attr_tenant into the initial state.

The stamp is safe under parallel function calls: every AgentTool in one root turn writes the same values. Missing keys become the literal unattributed rather than a guess, so the conservation check can see them.

Paid tools, retries and thinking tokens

Model calls are not the only spend. Charge three other kinds of cost to the agent that caused them.

  • Paid tools. A search API, a maps API or a code sandbox has its own price. Keep a per-tool rate card next to the model price table. Write a ledger row from afterToolCallback, and from onToolErrorCallback for tools that bill even on failure.
  • Retries. A response your code rejects, for example output that fails schema validation and is re-requested, was still generated and is normally billed. Mark these rows retry=true and charge them to the agent that retried. A high retry share is a prompt or schema bug with a price tag. Requests rejected before generation, such as rate-limit errors, usually carry no token charge, but confirm that against your bill rather than assuming it.
  • Thinking tokens. Reasoning models report them separately and bill them as output. Store them in their own column, because they are often the largest lever, as the worked example in the Gemini cost article shows.

Allocating shared costs

Some costs belong to no single call: context-cache storage, embedding and index refreshes, evaluation runs, and the remainder between the ledger and the bill. Give each a pool, a driver and a rule for who sees it.

PoolDriverCharged to
Context-cache storagecached tokens read, per agentagents and tenants, proportionally
Embedding and index refreshretrieval tool callsagents and tenants, proportionally
Evaluation and CI runseval runs per agentagents only (showback), never users
Unexplained remaindernonea visible unattributed line, investigated monthly

The driver should be the quantity that causes the cost, not whatever is convenient. Spreading cache storage by request count would charge agents that never read the cache. Keep two views. Showback shows teams their direct plus allocated cost, so they can act on it. Chargeback bills customers, and should charge only direct cost plus a published margin. That is the job of tenant billing, and allocation rules that change month to month make invoices hard to defend.

Per-user cost without exploding your metrics

Per-user cost has the highest cardinality of any attribution key. Put it in metric labels and the metrics backend grows one time series per user per agent per model. Keep user ids in the ledger table, where a user column is just a column, and keep metric labels to agent, model and tenant tier. Per-user questions then become queries:

-- Users whose spend today is far above their own 28-day baseline.
WITH daily AS (
  SELECT tenant, user_id, DATE(ts) AS day, SUM(cost_micros) / 1e6 AS usd
  FROM ledger.cost_rows
  WHERE ts >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 29 DAY)
  GROUP BY tenant, user_id, day
),
base AS (
  SELECT tenant, user_id, AVG(usd) AS mean_usd, STDDEV(usd) AS sd_usd
  FROM daily WHERE day < CURRENT_DATE() GROUP BY tenant, user_id
)
SELECT d.tenant, d.user_id, d.usd, b.mean_usd,
       SAFE_DIVIDE(d.usd - b.mean_usd, b.sd_usd) AS z
FROM daily d JOIN base b USING (tenant, user_id)
WHERE d.day = CURRENT_DATE() AND d.usd > 5 * b.mean_usd AND d.usd > 1.0
ORDER BY d.usd DESC LIMIT 50;

The absolute floor (more than a dollar a day here) keeps the list short. The ratio to the user's own baseline catches a scripted client or a looping agent that a fixed threshold would miss. Treat the user id as personal data: hash it as the tenant analytics pipeline does, and keep the ledger's retention period as short as your billing disputes allow.

Worked example: one turn and one month

Use the placeholder rates in the diagram: 0.30 USD per million uncached input tokens, 0.03 cached and 2.50 output, plus 0.005 USD per web search. One user turn runs concierge twice and delegates to research_agent through an AgentTool.

CallTokens (cached)OutputCost (USD)
concierge 16,000 (4,000)3000.000600 + 0.000120 + 0.000750 = 0.001470
concierge 29,000 (4,000)4000.001500 + 0.000120 + 0.001000 = 0.002620
research_agent x 35,000 (0) each500 each3 x (0.001500 + 0.001250) = 0.008250
web_search x 22 x 0.005 = 0.010000
Turn total0.022340

With default wiring, the ledger records 0.00409 USD for this turn, 18% of the true 0.02234. The research agent, which costs twice as much as the supervisor before its searches, looks free. With includePlugins on but no stamping, the full 0.02234 appears, but 0.01825 of it is charged to tmp-user. With the plugin above, all seven rows (five model calls and two searches) carry user u_481, tenant t_acme and the root invocation id, and the caller column shows concierge for the nested rows.

Over a month, suppose direct costs are 6,000 USD for concierge, 3,000 for research_agent and 1,000 for billing_agent, plus 1,200 USD of shared pools. Cache storage of 400 is split 90/10 between concierge and billing_agent by cached tokens read. An index refresh of 300 goes entirely to research_agent, the only retrieval caller. Eval runs of 500 are split 200/200/100 by runs. Showback is then 6,560, 3,500 and 1,140 USD, which sums to the 11,200 billed. If that sum ever misses the bill, the difference goes into the unattributed line rather than being spread out quietly.

Failure modes

  • Invisible nested runs. AgentTool without includePlugins drops whole sub-agents from the ledger. Grep for every AgentTool.create call and check the flag.
  • Phantom users. Nested rows charged to tmp-user. Alert whenever that literal appears in the ledger.
  • Double counting. Two plugins, or a plugin plus an agent-level callback, both write rows for the same call. Make row ids deterministic from the invocation id and a per-call sequence number, and upsert.
  • Partial responses. Streaming emits partial responses. Count usage only on the final one, as the plugin does.
  • Direct client calls. Embeddings or token counts made with a model client outside the runner bypass plugins entirely. Route them through the same ledger writer.
  • Label explosion. User ids in metric labels exhaust the metrics backend. Keep them in the table.
  • Drifting allocation rules. A driver changed mid-quarter makes trends meaningless. Version the rules, as you version prices.

Trade-offs

Plugin-based attribution is exact per call and needs no changes to agents, but it depends on framework details such as the nested runner's user id. Pin and test those details. Wrapping the model client instead would see calls made outside the runner, but it loses agent names and invocation structure. Many teams use both: the plugin as the main ledger, and client-level counters as a cross-check. Allocating shared pools by causal drivers is fairer, but takes more data than spreading them in proportion to direct cost. Start proportional, and switch a pool to a causal driver once it exceeds a few percent of spend. Token counts for budgeting before a call are a separate problem, covered in token counting across models.

What to do next

  1. Draw your agent tree and mark each edge as a sub-agent (same invocation) or an AgentTool (nested runner).
  2. Turn on includePlugins for every AgentTool whose calls you pay for, and confirm the behaviour in the ADK version you deploy.
  3. Write the tenant into session state at creation, and add the stamping beforeToolCallback.
  4. Add rows for paid tools and retries, with a per-tool rate card.
  5. Define each shared pool's driver, and publish showback separately from chargeback.
  6. Run a daily conservation check against the bill, and alert on the unattributed bucket and any tmp-user rows.
  7. Add an integration test that replays the worked example turn and asserts seven rows that sum to 0.02234 USD.
Key takeaway: Attribution in an agent tree is allocation under a conservation rule: every dollar lands in exactly one ledger row, and an explicit unattributed bucket absorbs the rest. In ADK Java, sub-agents are visible to Runner plugins, but AgentTool runs a nested runner that is invisible by default and charged to tmp-user once included. Opt in, stamp the real user and root invocation into state, price tools and retries, allocate shared pools by causal drivers, and reconcile daily.