Choosing a Gemini model for an ADK Java agent looks like a one-line change: swap the string passed to LlmAgent.builder().model(...) and rerun. That is exactly why teams get it wrong. The string changes more than the model: it changes which thinking controls are valid, whether ADK rewrites your structured-output request, how many hidden reasoning tokens you pay for, and whether the ID will still exist next quarter. A comparison that ignores those differences measures the plumbing, not the model.

This article builds a fair comparison harness in ADK Java: what the variant lineup looked like when this was written, where each variant enters the request, the confounds ADK itself introduces, a runnable Java harness that records tokens, latency and correctness per case, how to turn tokens into cost without hard-coding prices, and how to pick a model per agent rather than per application. For how each Gemini feature is wired, see the Gemini features guide; for how model strings resolve, see model registration and discovery.

The lineup on 2026-10-04

Model families move faster than any article, so treat this table as a dated snapshot. On 2026-10-04 Google's Gemini API models page listed these text-generation IDs. Thinking levels are from the thinking documentation on the same day. Verify both before you pin anything.

TierModel IDStatusThinking levels (default)
Progemini-3.1-pro-previewPreviewlow, medium, high (high)
Flashgemini-3.8-flashStablelow, medium, high (medium)
Flashgemini-3.7-flashStablelow, medium, high (medium)
Flashgemini-3.6-flashStableminimal, low, medium, high (medium)
Flashgemini-3.5-flashStablecheck the thinking page
Flash-Litegemini-3.5-flash-liteStableminimal, low, medium, high (minimal)
Flash-Litegemini-3.1-flash-liteStablecheck the thinking page

Three facts from that snapshot shape any comparison. First, the only Pro-tier text model was a preview, and preview IDs can change behaviour or disappear with little notice, so a Pro result is a measurement of a moving target. Second, Google said it was limiting 2.5 access to users who had used those models before, and recommended 3.5 Flash-Lite or 3.8 Flash for new projects. A new project may not be able to run a 2.5 baseline at all. Third, the levels differ: MINIMAL is listed for 3.6 Flash and 3.5 Flash-Lite but not for 3.8 or 3.7 Flash, so one config cannot be copied across every variant.

Prices are deliberately absent. They change, some are promotional or tiered by context length, and a wrong number in an article outlives its correction. The harness below reads prices from a config file you fill from the official pricing page, with the date you read it.

Where a variant enters the request

Where a variant enters an ADK Java request, and what to record on the way outVariant specmodel id + thinking levelLlmAgent.builder()model, config, toolsLlmRegistrystring to GeminiGemini APIserved modelRequest processorsBasic: copies GenerateContentConfigOutputSchema: set_model_response only for gemini-2.*Contents: keeps thought signaturesEvent streamusageMetadata, modelVersionScorerchecks + judgeResults ledgertokens, latency, scoreHold everything but the variant constant; log what was actually served, not what you asked for.
A variant is a model string plus the config that is valid for it. ADK resolves the string, runs request processors that can change the request shape, and returns events that carry the token usage and the served model version.

In ADK Java the model reaches the wire in one of two ways. Passing a string, such as model("gemini-3.8-flash"), lets the LLM registry resolve it to a Gemini instance. Passing a BaseLlm built with Gemini.builder().modelName(...) gives you control of the client, credentials and executor. For comparisons, build one client per backend and reuse it, so connection setup and credential refresh do not land in one variant's latency.

Before the request is sent, the flow's request processors build it. The Basic processor copies the agent's GenerateContentConfig, including its ThinkingConfig. The Contents processor assembles history and keeps thought signatures, which Gemini 3 models use to carry reasoning across turns in a tool loop. The OutputSchema processor is where model names start to matter.

Confounds you must control

Each of the following changes the request or the bill when you change only the model string. Control for every one.

  • Structured output plus tools. When an agent has both an outputSchema and tools, ADK checks ModelNameUtils.canUseOutputSchemaWithTools, which is true for anything that does not match gemini-2.* after stripping Vertex resource paths and apigee prefixes. For a 2.x name, ADK injects a set_model_response tool and an instruction telling the model to finish by calling it; for 3.x names the schema goes to the API natively. A 2.5-versus-3.x comparison on such an agent therefore compares two different protocols. Compare within a generation, or measure the rewrite's cost separately.
  • Thinking controls. 2.5 models take a thinkingBudget in tokens; the 3.x models above take a thinkingLevel. Sending the wrong one, or a level the model does not list, may be rejected or adjusted by the API, and ADK does not validate it. Map each variant to an explicit level and record it with the result.
  • Hidden tokens. Thinking tokens appear in thoughtsTokenCount, separately from candidatesTokenCount, and are billed as output. A variant that looks cheap on visible output can be the most expensive per task.
  • Served versus requested model. Record Event.modelVersion() from each response. Aliases and previews can be repointed; a regression that coincides with a version change is a model change, not your code.
  • Trajectory, not just answer. In an agent, a variant that calls three tools before answering costs three round trips and three prompt re-sends. Compare tool-call counts and turns per task, not only final text.
  • Shared state. Each case needs a fresh session. Reusing one session lets later cases see earlier history, which is a different test for every variant.

A comparison harness in Java

The harness keeps everything but the variant fixed: same instruction, same tools, same cases, a fresh session per case, and an explicit thinking level per variant. Tools should be deterministic fakes that return recorded fixtures, so a slow or flaky backend cannot be mistaken for a slow model. The code uses only APIs present in google/adk-java and the java-genai types as of this writing.

import com.google.adk.agents.LlmAgent;
import com.google.adk.events.Event;
import com.google.adk.runner.InMemoryRunner;
import com.google.adk.sessions.Session;
import com.google.adk.tools.BaseTool;
import com.google.genai.types.*;
import java.util.*;

record Variant(String modelId, ThinkingLevel.Known level) {}
record Case(String id, String prompt, java.util.function.Predicate<String> check) {}
record Result(String variant, String caseId, String servedModel, boolean pass,
              int promptTok, int outTok, int thoughtTok, int cachedTok,
              int toolCalls, long millis) {}

final class VariantHarness {
  static LlmAgent agentFor(Variant v, String instruction, List<BaseTool> tools) {
    return LlmAgent.builder()
        .name("triage")
        .model(v.modelId())
        .instruction(instruction)                 // identical across variants
        .tools(tools)                             // identical, deterministic fakes
        .generateContentConfig(GenerateContentConfig.builder()
            .thinkingConfig(ThinkingConfig.builder().thinkingLevel(v.level()).build())
            .build())
        .build();
  }

  static Result runCase(Variant v, LlmAgent agent, Case k) {
    InMemoryRunner runner = new InMemoryRunner(agent, "variant-eval");
    Session s = runner.sessionService()
        .createSession(runner.appName(), "eval").blockingGet();   // fresh per case
    long t0 = System.nanoTime();
    List<Event> events = runner
        .runAsync("eval", s.id(), Content.fromParts(Part.fromText(k.prompt())))
        .toList().blockingGet();
    long ms = (System.nanoTime() - t0) / 1_000_000;

    int p = 0, o = 0, th = 0, ca = 0, calls = 0;
    String served = "?", answer = "";
    for (Event e : events) {
      calls += e.functionCalls().size();
      served = e.modelVersion().orElse(served);
      var u = e.usageMetadata();
      if (u.isPresent()) {                        // one usage record per model call
        p  += u.get().promptTokenCount().orElse(0);
        o  += u.get().candidatesTokenCount().orElse(0);
        th += u.get().thoughtsTokenCount().orElse(0);
        ca += u.get().cachedContentTokenCount().orElse(0);
      }
      if (e.finalResponse()) answer = visibleText(e);
    }
    return new Result(v.modelId(), k.id(), served, k.check().test(answer),
        p, o, th, ca, calls, ms);
  }

  static String visibleText(Event e) {
    return e.content().flatMap(Content::parts).orElse(List.of()).stream()
        .filter(part -> !part.thought().orElse(false))
        .map(part -> part.text().orElse(""))
        .reduce("", String::concat);
  }
}

Run every case several times per variant, at least three and preferably five, because sampling makes single runs noisy. Interleave variants case by case rather than running all of one variant then all of the next, so a backend latency spike or a quota throttle hits all variants equally. Summing usage across events matters: an agent turn with two tool calls makes three model calls, each with its own usage record. Count only events that carry usage, and check on your ADK version that partial streaming events do not duplicate it.

Turning tokens into cost

Cost per case is computed from the recorded tokens and a dated price table you maintain, not from constants in code:

cost = (promptTok - cachedTok) * inPrice
     + cachedTok * cachedInPrice
     + (outTok + thoughtTok) * outPrice          # thinking is billed as output

Keep the price file next to the results with the date you read the pricing page, and recompute cost from raw tokens whenever prices change; never store only the derived dollars. Check whether the pricing page tiers a model by prompt length, and if so, choose the tier per call from promptTokenCount. Report cost per passing case, not per call: a cheaper variant that fails a third of the time and triggers a retry or a human escalation is not cheaper. For counting tokens before a call, see token counting across models.

Scoring quality

Quality scoring should be layered. First, deterministic checks: the JSON parses against the schema, the right tool was called with the right arguments, the routing label matches the expected one. These are cheap, reproducible and cover most agent tasks. Second, for free text, an LLM-as-judge scorer with a fixed rubric. The judge must be a model that is not in the comparison, with a pinned ID and thinking level, otherwise it may prefer its own style. Randomise answer order when the judge compares two outputs, and spot-check its verdicts by hand on a sample.

Size the case set for the decision. If two variants differ by a few percentage points in pass rate, a 50-case set cannot tell them apart; 300 or more cases drawn from real traffic, with the hard tail deliberately over-represented, gives a usable signal. Report pass rate with a confidence interval, p50 and p95 latency, and cost per passing case side by side.

Worked example: a triage agent

Consider a support-ticket triage agent with two tools, lookup_account and open_incident, and an output schema with fields category, severity and next_action. The team wants to know whether a Flash-Lite model can replace a Flash model, and whether Pro is ever worth it.

  1. Define variants: gemini-3.5-flash-lite at MINIMAL and at LOW, gemini-3.8-flash at LOW and MEDIUM, and gemini-3.1-pro-preview at its default HIGH, labelled preview in every report.
  2. Build 400 cases from last month's tickets with the resolved category as ground truth; over-sample the 10 percent that were escalated.
  3. Run five passes, interleaved. Because every variant is 3.x, the output schema goes natively in each, so the protocol is the same. If a 2.5 model were included it would get the set_model_response rewrite, and that difference must be reported.
  4. Score category and severity exactly; score next_action with the judge.
  5. Read the results per slice, not only overall. A common outcome is that the light model matches on routine tickets and loses on escalations. That argues for a router, not a single choice.

The output is a decision table with one row per variant: pass rate overall and on the hard slice, p95 latency, mean thinking tokens, tool calls per case and cost per passing case. The numbers depend on your traffic. The method is what transfers.

Choosing per agent, not per app

The best answer is rarely one model. ADK agent trees make per-agent choices natural: a router or classifier agent on Flash-Lite at a low level, a planner that weighs policy on Flash at medium, and an escalation path to Pro only for cases the cheaper agents mark as uncertain. Because each LlmAgent carries its own model and GenerateContentConfig, this is configuration, not code. Keep the mapping in one config file keyed by agent name, so a model change is a reviewed diff and the comparison harness can replay it.

Add fallbacks for availability, not quality: if a preview ID returns not-found or a quota error, fall back to a stable ID and log it, rather than retrying forever. The Gemini integration guide covers timeouts and retries at the client level.

Failure modes

  • Silent config mismatch. A thinking level the model does not list is sent and quietly adjusted. Log the effective config and the served model version per run.
  • Comparing across protocols. A 2.x baseline with output schema plus tools runs through set_model_response, and the 3.x candidate does not. Report it or restrict the comparison.
  • Undercounted cost. Ignoring thoughtsTokenCount understates high-thinking variants, sometimes by a multiple.
  • Preview drift. A Pro preview result from last month is not evidence about today's preview. Re-run before deciding.
  • Lost baseline. Access limits on older models can take your baseline away. Archive raw outputs so old results stay comparable.
  • Leaky tools. Live tools make results depend on backend state. Use recorded fixtures.

Trade-offs

A rigorous harness costs real tokens: 400 cases, five passes and five variants is 10,000 agent runs, and Pro at high thinking dominates that bill. Run the full matrix only when choosing, and a 50-case smoke set on every model or ADK upgrade. Interleaving and fresh sessions make runs slower and the code longer, but without them the numbers are not comparable. Per-agent model choice saves money and adds operational surface; keep the number of distinct models small enough that you can re-evaluate all of them when Google ships the next generation.

What to do next

  1. Snapshot the models page and the thinking page into your repo with the date; list which IDs are stable and which are preview.
  2. Write a variant spec per candidate: model ID plus an explicit thinking level the docs list for that model.
  3. Build a case set from real traffic with ground truth, over-sampling the hard slice.
  4. Replace live tools with recorded fixtures and use a fresh session per case.
  5. Record prompt, output, thinking and cached tokens, tool calls, latency and the served model version for every run.
  6. Compute cost from a dated price file; report cost per passing case with confidence intervals.
  7. Check for the gemini-2 output-schema rewrite before comparing across generations.
  8. Assign models per agent, keep the mapping in config, and re-run a smoke set on every model or ADK upgrade.
Key takeaway: Changing the Gemini model string in ADK Java changes more than the model: valid thinking controls, whether ADK rewrites structured output with tools for gemini-2 names, hidden thinking-token cost and ID stability. Compare variants with fixed instructions, fixture tools and fresh sessions, record tokens including thinking, latency and the served model version, compute cost from a dated price file, score per slice, and assign models per agent in config.