Configuring a Gemini model in the Java Agent Development Kit (ADK) means answering three questions: which model serves the agent, how it should sample, and how requests travel to the API. Most bugs come from setting one of these in the wrong place. A timeout ends up in the agent, a system prompt ends up in a config object that the framework merges with something else, or a schema gets silently dropped because the agent also has tools.

This article follows the request path through ADK Java's source. It shows which object owns each setting, walks through GenerateContentConfig field by field with its real Java types, and builds a two-agent example with different configs for a router and a writer. It ends with the failure modes that show up in production. Backend selection, credentials and the internals of the Gemini class are covered in ADK Java + Gemini and the GeminiLLM class deep dive. This page is about the model settings themselves.

Three layers of configuration

Every model call that an LlmAgent makes is assembled into an LlmRequest by a chain of request processors. The Basic processor reads the agent's resolved model name and copies the agent's generateContentConfig into the request, or an empty config if you set none. Later processors add to that copy. The instructions processor merges your instruction into systemInstruction. Tools add their function declarations. If you set an outputSchema, the request builder writes it as responseSchema and forces responseMimeType to application/json.

The finished request goes to the model object. If it is a Gemini instance, that object prepares the request for its backend before sending. On the Gemini Developer API path (an API key, not Google Cloud) it removes labels from the config and display names from file parts, because that backend rejects them. So the config you write is not exactly the config that gets sent. You own the sampling fields. ADK owns the fields it assembles. The backend adapter removes fields its target does not support.

From LlmAgent builder to the Gemini request: who writes which fieldLlmAgent.builder()model, instruction, toolsgenerateContentConfigsampling, thinking, safetyoutputSchemaoptional SchemaBasic processormodel name + config copyInstructions processormerges systemInstructionTools and schemadeclarations, JSON modeLlmRequestmodel + GenerateContentConfigGemini.generateContentsanitise per backendClient + HttpOptionsbackend, timeout, retriesGemini APIDeveloper API or Google CloudLlmRegistrygemini-.* to Geminiresolves modelAgent-level fields describe how to sample. Transport settings live on the client, not the agent.
Request assembly in ADK Java. Settings on the agent describe sampling. Settings on the client describe transport.
LayerObjectWhat belongs there
Which modelmodel(String) or model(BaseLlm)Model id, and the backend and credentials if you pass an instance
How to samplegenerateContentConfig(...)temperature, topP, topK, maxOutputTokens, seed, stop sequences, thinking, safety
What shape to returnoutputSchema(Schema)A JSON schema for final answers. ADK sets the MIME type
How to talk to the APIClient with HttpOptionsAPI version, timeout, headers, retry options

Choosing and resolving the model

The simplest form is a string. LlmAgent.builder().model("gemini-2.5-flash") stores the name, and at run time the agent resolves it through LlmRegistry. The registry ships with three patterns: gemini-.* and gemma-.* map to Gemini.builder().modelName(name).build(), and apigee/.* maps to an Apigee-fronted model. Instances are cached by name, so every agent that names the same model shares one object and one underlying client.

That default client is built from the environment. The google-genai Java SDK picks up GOOGLE_API_KEY for the Developer API, or a Google Cloud flag plus GOOGLE_CLOUD_PROJECT and GOOGLE_CLOUD_LOCATION. The name of that flag depends on the SDK version. The current SDK README documents GOOGLE_GENAI_USE_ENTERPRISE, and older releases use a Vertex-named variable, so check the README for the version you have pinned. Do not copy the name from a blog post, this one included.

When you need control over the client, build the model yourself and pass the instance. The constructor that takes an explicit Client wins over an API key or credentials, which makes it the most predictable option:

import com.google.adk.agents.LlmAgent;
import com.google.adk.models.Gemini;
import com.google.genai.Client;
import com.google.genai.types.HttpOptions;
import com.google.genai.types.HttpRetryOptions;

Client client = Client.builder()
    .apiKey(System.getenv("GOOGLE_API_KEY"))
    .httpOptions(HttpOptions.builder()
        .timeout(30_000)                       // milliseconds
        .retryOptions(HttpRetryOptions.builder()
            .attempts(3)
            .httpStatusCodes(408, 429, 503)))
    .build();

Gemini flash = Gemini.builder()
    .modelName("gemini-2.5-flash")
    .apiClient(client)
    .build();

LlmAgent agent = LlmAgent.builder()
    .name("triage")
    .model(flash)            // instance, not string: this client is used
    .instruction("Classify the ticket.")
    .build();

Passing an instance skips the registry for this agent only. Other agents that still use the string go through the registry with the default client. That split, with two clients and two retry policies, is a common source of confusing behaviour. Pick one style per application. The SDK's retries also stack with any retries in your own code. Retries at the LLM layer shows how those layers multiply.

GenerateContentConfig, field by field

GenerateContentConfig is an immutable value class from the com.google.genai.types package with a builder. The types matter in Java, because the builder overloads are strict. Sampling numbers are boxed Float, not Double, so you write 0.2f. Counts are Integer. topK is a Float too, which surprises people who expect an integer.

Builder methodJava typeWhat it doesAgent guidance
temperatureFloatScales logits before samplingLow (0 to 0.3) for routing, extraction and tool use. Higher only for prose
topPFloatNucleus cut-offChange temperature or topP, not both, so you can tell which one had the effect
topKFloatKeep the K most likely tokensUsually leave unset
maxOutputTokensIntegerCaps generated tokensSet it on every agent. A missing cap is a cost bug waiting to happen
stopSequencesList<String>Stops at a matching stringUseful for fixed formats. A risk if the string can appear in normal output
seedIntegerSeeds samplingHelps reproducibility in tests. Does not guarantee identical output
candidateCountIntegerNumber of candidatesLeave at one. An agent consumes a single candidate
thinkingConfigThinkingConfigReasoning budget or levelSee the next section
safetySettingsList<SafetySetting>Per-category block thresholdsSet them deliberately. Log blocked responses
cachedContentStringName of an explicit context cacheOnly with a cache you manage. See the Gemini 2.5 features article
labelsMap<String,String>Billing labelsRemoved on the Developer API path. Only reach Google Cloud

A complete, typed config for an agent that calls tools looks like this:

import com.google.genai.types.*;
import java.util.List;

GenerateContentConfig toolUserConfig = GenerateContentConfig.builder()
    .temperature(0.1f)
    .maxOutputTokens(1024)
    .seed(7)
    .thinkingConfig(ThinkingConfig.builder().thinkingBudget(512).build())
    .safetySettings(List.of(
        SafetySetting.builder()
            .category(HarmCategory.Known.HARM_CATEGORY_DANGEROUS_CONTENT)
            .threshold(HarmBlockThreshold.Known.BLOCK_MEDIUM_AND_ABOVE)
            .build()))
    .labels(java.util.Map.of("team", "support", "agent", "triage"))
    .build();

The Known enums are convenience overloads. The builders also accept the wrapper types and raw strings, so a value that a newer API version adds can be passed before the SDK enum includes it.

Thinking budgets and levels

Gemini 2.5 models reason before answering, and ThinkingConfig controls how much. The Java builder exposes thinkingBudget(Integer), thinkingLevel(ThinkingLevel) and includeThoughts(boolean). A budget is a token count. A budget of zero turns thinking off on models that allow it, which the SDK README itself uses as an example. ThinkingLevel is a coarser enum (MINIMAL, LOW, MEDIUM and up) aimed at newer model families. Which of the two a model accepts, and the valid range, depends on the model, so check that model's documentation rather than assuming the settings carry over.

Thinking changes the economics of an agent. Thought tokens are billed, they add latency before the first visible token, and on some models they count against the output limit. If you set maxOutputTokens(256) and a generous budget, the model can use most of the limit on reasoning and return a truncated answer. Read thoughtsTokenCount from the usage metadata on real traffic before choosing either number.

includeThoughts(true) returns thought summaries as parts flagged as thoughts. They are useful for debugging. Do not show them to end users unreviewed, and do not store them as part of the user-visible transcript. The Gemini 2.5 features article covers thought parts and how ADK handles them in the conversation history.

Fields ADK writes for you

Some fields exist on GenerateContentConfig but should not be set by you on an agent, because ADK writes them. In ADK Java, the LlmAgent builder does not reject these. It stores the config without inspecting it. So there is no early error, and the result depends on how each processor merges.

  • systemInstruction. The instructions processor merges your agent instruction into whatever system instruction the config already holds. It does not replace it. If you put a system prompt in both places, both reach the model, and editing the agent's instruction no longer controls the whole prompt. Keep the system prompt in instruction only.
  • tools. Register tools on the agent with .tools(...). ADK keys them by name, throws on duplicates, and routes function calls back to the matching BaseTool. A declaration you add straight into the config has no Java handler behind it. Function calling is covered in Gemini function calling in ADK Java.
  • responseSchema and responseMimeType. Use outputSchema. The Basic processor applies it only if the agent has no tools, or if ADK's model-name check says the model can combine a schema with tools. Otherwise the schema is quietly not applied, and the agent returns free text where your parser expects JSON.

Worked example: a router and a writer

Take a support system with two agents. A router reads a ticket and calls one of three tools. A writer drafts the customer reply. They need opposite settings, which is why config belongs to the agent and not to the application.

GenerateContentConfig routerCfg = GenerateContentConfig.builder()
    .temperature(0.0f)
    .maxOutputTokens(256)
    .thinkingConfig(ThinkingConfig.builder().thinkingBudget(0).build())
    .build();

GenerateContentConfig writerCfg = GenerateContentConfig.builder()
    .temperature(0.7f)
    .maxOutputTokens(2048)
    .thinkingConfig(ThinkingConfig.builder().thinkingBudget(1024).build())
    .build();

LlmAgent router = LlmAgent.builder()
    .name("router")
    .model("gemini-2.5-flash")
    .instruction("Pick exactly one tool for the ticket.")
    .tools(lookupOrder, openRefund, escalate)
    .generateContentConfig(routerCfg)
    .build();

LlmAgent writer = LlmAgent.builder()
    .name("writer")
    .model("gemini-2.5-pro")
    .instruction("Write a short, accurate reply using the case notes in state.")
    .generateContentConfig(writerCfg)
    .build();

Now trace one ticket. The router's request carries temperature 0, a 256-token cap, no thinking, three function declarations and the merged instruction. Suppose the team later adds outputSchema to the router so its classification comes back as JSON. The router still has tools, so the schema is applied only if the model passes ADK's schema-with-tools check. On a model that fails it, the router keeps working and the JSON parser downstream starts failing on a fraction of tickets. The fix is to decide on one channel: either the router calls a tool, or it returns a schema with no tools.

Next, the writer. Its first drafts come back cut off mid-sentence. Usage metadata shows about 1,000 thought tokens and about 1,000 visible tokens: the budget and the cap were sized without reference to each other. Raising maxOutputTokens fixes truncation but raises the worst-case cost. Lowering the thinking budget fixes it more cheaply if quality holds. Which one is right is a measurement, not a guess. Run both configs over the same set of tickets, as described in the variant-comparison harness, and compare quality, latency and cost per reply.

Failure modes

SymptomLikely causeCheck
Agent ignores an edited instructionA second system prompt lives in the configLog the final systemInstruction from a callback
JSON parse errors on some turnsoutputSchema not applied because the agent has toolsInspect responseMimeType on the outgoing request
Truncated answers, MAX_TOKENS finish reasonThinking budget and output cap competeCompare thoughtsTokenCount with the cap
Billing labels missingDeveloper API path removes labelsLabels only work on the Google Cloud backend
Different timeouts per agentSome agents use instances, others stringsOne model factory for the whole app
400 on a thinking fieldLevel or budget not supported by this modelPer-model config, validated at startup
Empty response, safety finish reasonThreshold stricter than the use case needsLog category and probability on every block

Two failures are silent, and those are the expensive ones. A dropped schema and a merged system prompt both produce working agents that behave a little differently from the code you read. Make the outgoing request visible: a beforeModelCallback that logs the model name and a hash of the config, together with the trace span, will catch both.

Operating model configuration

Treat model configuration as code that goes through review. Some practical rules:

  • Build configs in one place. A small factory returns the config for each agent role, reads overrides from your normal config system, and validates ranges at startup. Fail fast on a temperature above 2 or a missing output cap, instead of discovering it from a 400 at run time.
  • Version the model id. Aliases that always point at the newest model change underneath you. Use a stable id in production and change it on purpose, with an evaluation run.
  • Record the effective config with every trace. Model name, temperature, budget, cap and the schema flag are five attributes. They turn a vague quality regression into a diff you can read.
  • Separate transport from sampling. Timeouts, API version and retries belong on the Client. A per-request httpOptions in the config replaces the client's retry settings entirely rather than merging with them, so use it rarely and knowingly.
  • Pin the SDK. Enum values, environment variable names and defaults change between releases of google-genai and ADK. Upgrade both deliberately and re-run your evaluations.

What to do next

  1. List every LlmAgent and record its model id, temperature, output cap, thinking setting and whether it has tools and a schema.
  2. Move any systemInstruction, tools or responseSchema out of generateContentConfig and into the agent builder.
  3. Pick one way of creating models, string or instance, and route both through a single factory with one client.
  4. Set maxOutputTokens on every agent, then size thinking budgets against it using thoughtsTokenCount from real traffic.
  5. Add a beforeModelCallback that logs the model and config hash onto the trace span.
  6. Check the environment variable names in the README of the exact google-genai version you ship.
  7. Run the same ticket set through two configs per role and choose with measured quality, latency and cost.
Key takeaway: An ADK Java agent's model settings come from three places: the model you resolve, the GenerateContentConfig you set, and the client that carries the request. Set sampling, output caps, thinking and safety on the agent. Let ADK own instructions, tools and schemas. Keep transport on the client, and log the effective config so silent merges and dropped schemas show up.