A token budget is one of the few hard limits in an agent system. Exceed the context window and the call fails; exceed a cost budget and finance notices; underestimate a request and your history-trimming logic cuts too little, then the call fails anyway. Once an ADK Java agent can route to more than one model, token counting stops being a single function call: each provider tokenises differently, exposes different counting APIs, and reports usage in a slightly different shape.

This article builds a correct mental model and a working design. It explains why the same prompt has different token counts on different models, separates the three places a count can come from, shows exactly what ADK Java exposes (checked against the adk-java and java-genai sources on 2026-10-02), and builds a budget guard and usage ledger out of model callbacks. One correction up front: BaseLlm in ADK Java declares only model(), generateContent(LlmRequest, boolean) and connect(LlmRequest). There is no token-counting method for adapters to override, so counting is something you build around the model, not something the model class gives you.

Advertisement

Why the same text has different counts

A token is whatever unit a model's tokenizer produces: a learned vocabulary of byte sequences, built by an algorithm such as byte-pair encoding or a unigram language model over a training corpus. Each model family trains its own vocabulary. A family with a large vocabulary that saw a lot of code will encode a Java stack trace in fewer tokens than one that did not; a vocabulary with good coverage of Hindi or Japanese will spend fewer tokens per character on those scripts. So 'how many tokens is this prompt' has no single answer; it is always 'how many tokens is this prompt for this model'.

The common rule of about four characters per token is a rough average for English prose and is wrong in exactly the places that hurt: source code, JSON with many short keys, numbers, non-Latin scripts, and base64 blobs can all diverge from it substantially. Images, audio and video are counted by entirely different rules defined by each provider. And a request contains more than the visible chat: the system instruction, every tool's function declaration (name, description and parameter schema), and any structured-output schema are all tokenised and billed as input.

Three sources of truth

Three sources of token truth around one ADK Java model callLlmAgentbuilds LlmRequestbeforeModelCallbackestimate + budget checkShort-circuitreturn LlmResponseBaseLlm.generateContentGemini / Claude / customProvider count APIpre-flight, optionalafterModelCallbackread usageMetadata()Usage ledgerper invocation, per modelCalibrationactual / estimate ratioover budgetwithin budgetcountTokensLlmResponseratio feedsnext estimate
Estimate before the call, optionally ask the provider, and record the billed usage afterwards; feed the ratio back into the estimator.
SourceWhenCostAccuracyUse it for
Local estimateBefore the callMicroseconds, no networkApproximate; model-dependent errorHistory trimming, routing, coarse budgets
Provider count endpointBefore the callOne network round tripExact for what you send itHard context-window checks near the limit
Response usage metadataAfter the callFreeWhat you were billedLedgers, cost attribution, calibration

Mature systems use all three. The estimate runs on every call because it is free. The provider count runs only when the estimate is within some margin of a hard limit. The usage metadata is the ground truth for accounting and for correcting the estimator over time.

Advertisement

What ADK Java exposes

The post-call truth arrives on LlmResponse.usageMetadata(), which returns Optional<GenerateContentResponseUsageMetadata> (the type comes from the Google Gen AI Java SDK, com.google.genai.types). Its fields are all Optional<Integer>:

FieldMeaning
promptTokenCount()Input tokens for the request
cachedContentTokenCount()Input tokens served from a context cache
candidatesTokenCount()Generated output tokens
thoughtsTokenCount()Reasoning tokens on thinking models
toolUsePromptTokenCount()Tokens from tool-use prompts
totalTokenCount()The provider's total

There are also per-modality breakdowns (promptTokensDetails() and friends, lists of ModalityTokenCount) that tell you how much of the input was text versus image or audio. How these fields relate (whether cached tokens are a subset of the prompt count, whether thoughts are inside the candidates count) is defined by the provider, not by ADK, so do not hard-code a formula. Record every field, and add a reconciliation check that compares totalTokenCount with the sum of the components so a change in semantics shows up as an alert rather than a silent billing error.

Adapters fill this object differently. The built-in Gemini model passes through whatever the API returns. The built-in Claude adapter maps the Anthropic message's input tokens to promptTokenCount, output tokens to candidatesTokenCount, and their sum to totalTokenCount; it sets no cache fields, and it does not stream. If you enable Anthropic prompt caching through your own adapter, remember that Anthropic reports cache reads and cache writes separately from its input count, so a mapping that copies only the input field will understate input. A custom BaseLlm (see implementing a custom LLM) must populate the metadata itself, or every downstream ledger reads zero.

Pre-flight counts from providers

For Gemini, the java-genai client offers client.models.countTokens(String model, List<Content> contents, CountTokensConfig config), returning a CountTokensResponse whose totalTokens() is an Optional<Integer>. There is an important asymmetry in the current SDK source. In Gemini Developer API mode (API key), passing systemInstruction, tools or generationConfig in the config throws IllegalArgumentException; those are only accepted in what the SDK's error messages call 'Gemini Enterprise Agent Platform mode', the Vertex-style client. The related computeTokens method, which returns the actual token IDs, throws UnsupportedOperationException outside that mode.

The practical consequence: on the Developer API a pre-flight count covers your contents but not your system instruction or tool declarations. For an agent with twenty tools, that can be a large slice of the request. Either count those fixed parts once by sending them as ordinary contents and caching the number per agent version, or rely on calibration against billed usage (below).

import com.google.genai.Client;
import com.google.genai.types.Content;
import com.google.genai.types.CountTokensResponse;
import java.util.List;

public final class GeminiCounter implements TokenCounter {
  private final Client client;
  private final String model;

  public GeminiCounter(Client client, String model) {
    this.client = client;
    this.model = model;
  }

  @Override
  public int exactInputTokens(List<Content> contents) {
    // Developer API mode: contents only. System instruction and tools are NOT counted here.
    CountTokensResponse r = client.models.countTokens(model, contents, null);
    return r.totalTokens().orElseThrow(() -> new IllegalStateException("no count returned"));
  }
}

For Claude, Anthropic exposes a count endpoint at POST /v1/messages/count_tokens that accepts the same system prompt, messages and tools as a real request and returns the input token count. Call it from your Anthropic client in whatever form your SDK version provides; check its documentation for the exact method name rather than assuming one. OpenAI's Responses API has a similar endpoint, POST /v1/responses/input_tokens, that takes the same payload as a create call. Its tokenizer encodings are also public, so JVM ports of them can count text locally; message framing adds a few tokens per message, so measure once against billed usage.

A TokenCounter abstraction

Hide all of this behind one interface so agents and trimming logic never branch on provider. The interface has a cheap estimate that must never do I/O, and an optional exact count that may.

public interface TokenCounter {
  /** Fast, local, never blocks. Used on every call. */
  default int estimateInputTokens(List<Content> contents, String systemText, int toolSchemaChars) {
    long chars = systemText.length() + toolSchemaChars;
    for (Content c : contents) {
      for (Part p : c.parts().orElse(List.of())) {
        chars += p.text().map(String::length).orElse(0);
      }
    }
    return (int) Math.ceil(chars / charsPerToken());
  }

  /** Learned from billed usage; start conservative. */
  default double charsPerToken() { return 3.0; }

  /** Exact count from the provider; may block. */
  int exactInputTokens(List<Content> contents);
}

Starting at 3.0 characters per token rather than 4 is deliberate: an estimator should err towards overcounting, because overcounting trims a little too much history while undercounting fails the call. Non-text parts need their own rule per provider; until you have one, treat them as a fixed allowance and let calibration correct it.

Enforcing budgets with model callbacks

ADK Java's LlmAgent.Builder accepts beforeModelCallback and afterModelCallback (with ...Sync variants returning Optional). In the flow, if a before-model callback returns a response, that response is used and the model is not called. That gives a clean budget guard: estimate, and if the invocation would exceed its budget, answer with a short explanatory message instead of spending tokens.

LlmAgent agent = LlmAgent.builder()
    .name("support_agent")
    .model("gemini-2.5-flash")
    .instruction(SYSTEM_TEXT)
    .beforeModelCallbackSync((ctx, req) -> {
      int used = (int) ctx.state().getOrDefault("tokens:" + ctx.invocationId(), 0);
      int est = counter.estimateInputTokens(req.build().contents(), SYSTEM_TEXT, TOOL_SCHEMA_CHARS);
      if (used + est > INVOCATION_BUDGET) {
        return Optional.of(LlmResponse.builder()
            .content(Content.fromParts(Part.fromText(
                "This request is too large to process. Please narrow the question.")))
            .build());
      }
      return Optional.empty();                    // proceed with the real call
    })
    .afterModelCallbackSync((ctx, resp) -> {
      if (resp.partial().orElse(false)) return Optional.empty();   // final chunk only
      resp.usageMetadata().ifPresent(u -> {
        int total = u.totalTokenCount().orElse(
            u.promptTokenCount().orElse(0) + u.candidatesTokenCount().orElse(0));
        ctx.state().merge("tokens:" + ctx.invocationId(), total, (a, b) -> (int) a + (int) b);
        ledger.record(ctx.invocationId(), ctx.agentName(), resp.modelVersion().orElse("?"), u);
      });
      return Optional.empty();                    // keep the original response
    })
    .build();

Keying the counter by invocation id scopes the budget to one invocation; prune old keys, or keep counters in your own store, for long-lived sessions. Read usage in the after-model callback rather than by scanning session events. Outside bidirectional live streaming, ADK does not turn a response into an event when it carries no content, error, grounding or transcription, so a trailing usage-only response may never appear in the event stream at all. Also note that RunConfig.maxLlmCalls() (500 by default) limits the number of model calls per invocation, not tokens; it is a useful backstop against tool loops but no substitute for a token budget. Callbacks are covered in detail in ADK Java callbacks.

Calibration: closing the loop

Each completed call gives you a pair: the estimate you made and the prompt tokens you were billed. Keep an exponentially weighted moving average of their ratio per model (and per agent, since tool schemas differ), and divide the estimate by it next time. Worked example: your estimator predicts 8,000 input tokens for a request; Gemini bills 9,600. The ratio is 1.2. After a few dozen calls the EWMA settles near 1.2 and the corrected estimate becomes 9,600. Clamp the ratio to a sane band, say 0.5 to 3.0, so one strange multimodal request cannot poison it, and alert if it drifts sharply, because that usually means a model version changed its tokenizer or a prompt template grew.

Combine calibration with a margin. If the corrected estimate is more than about 85 percent of the context window, pay for the exact pre-flight count; below that, trust the estimate. The margin and threshold are policy, so make them configuration, and see ADK Java observability for exporting the ratio as a metric.

Failure modes and trade-offs

FailureCauseFix
Ledger shows zero usageCustom adapter never sets usageMetadataPopulate it in generateContent; add a test that asserts it is present
Usage counted twiceSumming every streamed chunkRecord on the final, non-partial response only
Context overflow despite a 'count'Developer API count excluded system text and toolsCount fixed parts separately or calibrate
Costs understated with cachingAdapter maps only uncached inputMap cache read and write fields too
Budget guard trims too littleOptimistic 4 chars per token on code or JSONStart conservative; learn the ratio
Mixed-model totals meaninglessAdding tokens from different tokenizersConvert to cost per model before aggregating

The central trade-off is latency against precision. A pre-flight count adds a round trip to every call it guards; local estimates are free but drift. Most agents should estimate always, count exactly only near limits, and treat billed usage as the record. Remember too that tokens from different models are different units: a dashboard that adds Gemini and Claude tokens together is adding metres to feet. Aggregate cost, not tokens. Rate limiting by tokens is a related concern covered in the ADK Java rate limiter guide.

What to do next

  1. Log the full usageMetadata from an afterModelCallback for every model you use, and confirm each adapter actually populates it.
  2. Introduce a TokenCounter interface with a conservative local estimate, and route all history-trimming decisions through it.
  3. Measure your system instruction and tool declarations once per agent version and add them to every estimate.
  4. Add a budget guard in beforeModelCallback that short-circuits with a helpful message instead of failing at the provider.
  5. Record estimate-versus-billed ratios per model, apply an EWMA correction, and alert on drift.
  6. Use the provider's exact count only when the corrected estimate is close to the context limit, and aggregate spend in currency rather than tokens.
Key takeaway: ADK Java gives you no count method on BaseLlm; it gives you billed usage on LlmResponse.usageMetadata() and callbacks around every model call. Build counting on top: a conservative local estimate for every call, a provider count near hard limits (remembering that the Gemini Developer API cannot count system instructions or tools), and billed usage recorded in an afterModelCallback as the source of truth that calibrates the estimate. Enforce budgets by short-circuiting in beforeModelCallback, and aggregate across models in cost, never in raw tokens.