Configuring a Gemini model in the Java Agent Development Kit (ADK) means answering three questions: which model serves the agent, how it should sample, and how requests travel to the API. Most bugs come from setting one of these in the wrong place. A timeout ends up in the agent, a system prompt ends up in a config object that the framework merges with something else, or a schema gets silently dropped because the agent also has tools.
This article follows the request path through ADK Java's source. It shows which object owns each setting, walks through GenerateContentConfig field by field with its real Java types, and builds a two-agent example with different configs for a router and a writer. It ends with the failure modes that show up in production. Backend selection, credentials and the internals of the Gemini class are covered in ADK Java + Gemini and the GeminiLLM class deep dive. This page is about the model settings themselves.
Three layers of configuration
Every model call that an LlmAgent makes is assembled into an LlmRequest by a chain of request processors. The Basic processor reads the agent's resolved model name and copies the agent's generateContentConfig into the request, or an empty config if you set none. Later processors add to that copy. The instructions processor merges your instruction into systemInstruction. Tools add their function declarations. If you set an outputSchema, the request builder writes it as responseSchema and forces responseMimeType to application/json.
The finished request goes to the model object. If it is a Gemini instance, that object prepares the request for its backend before sending. On the Gemini Developer API path (an API key, not Google Cloud) it removes labels from the config and display names from file parts, because that backend rejects them. So the config you write is not exactly the config that gets sent. You own the sampling fields. ADK owns the fields it assembles. The backend adapter removes fields its target does not support.
| Layer | Object | What belongs there |
|---|---|---|
| Which model | model(String) or model(BaseLlm) | Model id, and the backend and credentials if you pass an instance |
| How to sample | generateContentConfig(...) | temperature, topP, topK, maxOutputTokens, seed, stop sequences, thinking, safety |
| What shape to return | outputSchema(Schema) | A JSON schema for final answers. ADK sets the MIME type |
| How to talk to the API | Client with HttpOptions | API version, timeout, headers, retry options |
Choosing and resolving the model
The simplest form is a string. LlmAgent.builder().model("gemini-2.5-flash") stores the name, and at run time the agent resolves it through LlmRegistry. The registry ships with three patterns: gemini-.* and gemma-.* map to Gemini.builder().modelName(name).build(), and apigee/.* maps to an Apigee-fronted model. Instances are cached by name, so every agent that names the same model shares one object and one underlying client.
That default client is built from the environment. The google-genai Java SDK picks up GOOGLE_API_KEY for the Developer API, or a Google Cloud flag plus GOOGLE_CLOUD_PROJECT and GOOGLE_CLOUD_LOCATION. The name of that flag depends on the SDK version. The current SDK README documents GOOGLE_GENAI_USE_ENTERPRISE, and older releases use a Vertex-named variable, so check the README for the version you have pinned. Do not copy the name from a blog post, this one included.
When you need control over the client, build the model yourself and pass the instance. The constructor that takes an explicit Client wins over an API key or credentials, which makes it the most predictable option:
import com.google.adk.agents.LlmAgent;
import com.google.adk.models.Gemini;
import com.google.genai.Client;
import com.google.genai.types.HttpOptions;
import com.google.genai.types.HttpRetryOptions;
Client client = Client.builder()
.apiKey(System.getenv("GOOGLE_API_KEY"))
.httpOptions(HttpOptions.builder()
.timeout(30_000) // milliseconds
.retryOptions(HttpRetryOptions.builder()
.attempts(3)
.httpStatusCodes(408, 429, 503)))
.build();
Gemini flash = Gemini.builder()
.modelName("gemini-2.5-flash")
.apiClient(client)
.build();
LlmAgent agent = LlmAgent.builder()
.name("triage")
.model(flash) // instance, not string: this client is used
.instruction("Classify the ticket.")
.build();Passing an instance skips the registry for this agent only. Other agents that still use the string go through the registry with the default client. That split, with two clients and two retry policies, is a common source of confusing behaviour. Pick one style per application. The SDK's retries also stack with any retries in your own code. Retries at the LLM layer shows how those layers multiply.
GenerateContentConfig, field by field
GenerateContentConfig is an immutable value class from the com.google.genai.types package with a builder. The types matter in Java, because the builder overloads are strict. Sampling numbers are boxed Float, not Double, so you write 0.2f. Counts are Integer. topK is a Float too, which surprises people who expect an integer.
| Builder method | Java type | What it does | Agent guidance |
|---|---|---|---|
temperature | Float | Scales logits before sampling | Low (0 to 0.3) for routing, extraction and tool use. Higher only for prose |
topP | Float | Nucleus cut-off | Change temperature or topP, not both, so you can tell which one had the effect |
topK | Float | Keep the K most likely tokens | Usually leave unset |
maxOutputTokens | Integer | Caps generated tokens | Set it on every agent. A missing cap is a cost bug waiting to happen |
stopSequences | List<String> | Stops at a matching string | Useful for fixed formats. A risk if the string can appear in normal output |
seed | Integer | Seeds sampling | Helps reproducibility in tests. Does not guarantee identical output |
candidateCount | Integer | Number of candidates | Leave at one. An agent consumes a single candidate |
thinkingConfig | ThinkingConfig | Reasoning budget or level | See the next section |
safetySettings | List<SafetySetting> | Per-category block thresholds | Set them deliberately. Log blocked responses |
cachedContent | String | Name of an explicit context cache | Only with a cache you manage. See the Gemini 2.5 features article |
labels | Map<String,String> | Billing labels | Removed on the Developer API path. Only reach Google Cloud |
A complete, typed config for an agent that calls tools looks like this:
import com.google.genai.types.*;
import java.util.List;
GenerateContentConfig toolUserConfig = GenerateContentConfig.builder()
.temperature(0.1f)
.maxOutputTokens(1024)
.seed(7)
.thinkingConfig(ThinkingConfig.builder().thinkingBudget(512).build())
.safetySettings(List.of(
SafetySetting.builder()
.category(HarmCategory.Known.HARM_CATEGORY_DANGEROUS_CONTENT)
.threshold(HarmBlockThreshold.Known.BLOCK_MEDIUM_AND_ABOVE)
.build()))
.labels(java.util.Map.of("team", "support", "agent", "triage"))
.build();The Known enums are convenience overloads. The builders also accept the wrapper types and raw strings, so a value that a newer API version adds can be passed before the SDK enum includes it.
Thinking budgets and levels
Gemini 2.5 models reason before answering, and ThinkingConfig controls how much. The Java builder exposes thinkingBudget(Integer), thinkingLevel(ThinkingLevel) and includeThoughts(boolean). A budget is a token count. A budget of zero turns thinking off on models that allow it, which the SDK README itself uses as an example. ThinkingLevel is a coarser enum (MINIMAL, LOW, MEDIUM and up) aimed at newer model families. Which of the two a model accepts, and the valid range, depends on the model, so check that model's documentation rather than assuming the settings carry over.
Thinking changes the economics of an agent. Thought tokens are billed, they add latency before the first visible token, and on some models they count against the output limit. If you set maxOutputTokens(256) and a generous budget, the model can use most of the limit on reasoning and return a truncated answer. Read thoughtsTokenCount from the usage metadata on real traffic before choosing either number.
includeThoughts(true) returns thought summaries as parts flagged as thoughts. They are useful for debugging. Do not show them to end users unreviewed, and do not store them as part of the user-visible transcript. The Gemini 2.5 features article covers thought parts and how ADK handles them in the conversation history.
Fields ADK writes for you
Some fields exist on GenerateContentConfig but should not be set by you on an agent, because ADK writes them. In ADK Java, the LlmAgent builder does not reject these. It stores the config without inspecting it. So there is no early error, and the result depends on how each processor merges.
systemInstruction. The instructions processor merges your agent instruction into whatever system instruction the config already holds. It does not replace it. If you put a system prompt in both places, both reach the model, and editing the agent'sinstructionno longer controls the whole prompt. Keep the system prompt ininstructiononly.tools. Register tools on the agent with.tools(...). ADK keys them by name, throws on duplicates, and routes function calls back to the matchingBaseTool. A declaration you add straight into the config has no Java handler behind it. Function calling is covered in Gemini function calling in ADK Java.responseSchemaandresponseMimeType. UseoutputSchema. TheBasicprocessor applies it only if the agent has no tools, or if ADK's model-name check says the model can combine a schema with tools. Otherwise the schema is quietly not applied, and the agent returns free text where your parser expects JSON.
Worked example: a router and a writer
Take a support system with two agents. A router reads a ticket and calls one of three tools. A writer drafts the customer reply. They need opposite settings, which is why config belongs to the agent and not to the application.
GenerateContentConfig routerCfg = GenerateContentConfig.builder()
.temperature(0.0f)
.maxOutputTokens(256)
.thinkingConfig(ThinkingConfig.builder().thinkingBudget(0).build())
.build();
GenerateContentConfig writerCfg = GenerateContentConfig.builder()
.temperature(0.7f)
.maxOutputTokens(2048)
.thinkingConfig(ThinkingConfig.builder().thinkingBudget(1024).build())
.build();
LlmAgent router = LlmAgent.builder()
.name("router")
.model("gemini-2.5-flash")
.instruction("Pick exactly one tool for the ticket.")
.tools(lookupOrder, openRefund, escalate)
.generateContentConfig(routerCfg)
.build();
LlmAgent writer = LlmAgent.builder()
.name("writer")
.model("gemini-2.5-pro")
.instruction("Write a short, accurate reply using the case notes in state.")
.generateContentConfig(writerCfg)
.build();Now trace one ticket. The router's request carries temperature 0, a 256-token cap, no thinking, three function declarations and the merged instruction. Suppose the team later adds outputSchema to the router so its classification comes back as JSON. The router still has tools, so the schema is applied only if the model passes ADK's schema-with-tools check. On a model that fails it, the router keeps working and the JSON parser downstream starts failing on a fraction of tickets. The fix is to decide on one channel: either the router calls a tool, or it returns a schema with no tools.
Next, the writer. Its first drafts come back cut off mid-sentence. Usage metadata shows about 1,000 thought tokens and about 1,000 visible tokens: the budget and the cap were sized without reference to each other. Raising maxOutputTokens fixes truncation but raises the worst-case cost. Lowering the thinking budget fixes it more cheaply if quality holds. Which one is right is a measurement, not a guess. Run both configs over the same set of tickets, as described in the variant-comparison harness, and compare quality, latency and cost per reply.
Failure modes
| Symptom | Likely cause | Check |
|---|---|---|
| Agent ignores an edited instruction | A second system prompt lives in the config | Log the final systemInstruction from a callback |
| JSON parse errors on some turns | outputSchema not applied because the agent has tools | Inspect responseMimeType on the outgoing request |
Truncated answers, MAX_TOKENS finish reason | Thinking budget and output cap compete | Compare thoughtsTokenCount with the cap |
| Billing labels missing | Developer API path removes labels | Labels only work on the Google Cloud backend |
| Different timeouts per agent | Some agents use instances, others strings | One model factory for the whole app |
400 on a thinking field | Level or budget not supported by this model | Per-model config, validated at startup |
| Empty response, safety finish reason | Threshold stricter than the use case needs | Log category and probability on every block |
Two failures are silent, and those are the expensive ones. A dropped schema and a merged system prompt both produce working agents that behave a little differently from the code you read. Make the outgoing request visible: a beforeModelCallback that logs the model name and a hash of the config, together with the trace span, will catch both.
Operating model configuration
Treat model configuration as code that goes through review. Some practical rules:
- Build configs in one place. A small factory returns the config for each agent role, reads overrides from your normal config system, and validates ranges at startup. Fail fast on a temperature above 2 or a missing output cap, instead of discovering it from a 400 at run time.
- Version the model id. Aliases that always point at the newest model change underneath you. Use a stable id in production and change it on purpose, with an evaluation run.
- Record the effective config with every trace. Model name, temperature, budget, cap and the schema flag are five attributes. They turn a vague quality regression into a diff you can read.
- Separate transport from sampling. Timeouts, API version and retries belong on the
Client. A per-requesthttpOptionsin the config replaces the client's retry settings entirely rather than merging with them, so use it rarely and knowingly. - Pin the SDK. Enum values, environment variable names and defaults change between releases of
google-genaiand ADK. Upgrade both deliberately and re-run your evaluations.
What to do next
- List every
LlmAgentand record its model id, temperature, output cap, thinking setting and whether it has tools and a schema. - Move any
systemInstruction,toolsorresponseSchemaout ofgenerateContentConfigand into the agent builder. - Pick one way of creating models, string or instance, and route both through a single factory with one client.
- Set
maxOutputTokenson every agent, then size thinking budgets against it usingthoughtsTokenCountfrom real traffic. - Add a
beforeModelCallbackthat logs the model and config hash onto the trace span. - Check the environment variable names in the README of the exact
google-genaiversion you ship. - Run the same ticket set through two configs per role and choose with measured quality, latency and cost.