Gemini 2.5 brought several capabilities that change how an agent behaves rather than only how well it answers: a configurable thinking budget, thought summaries, server-side tools such as Google Search, URL context and code execution, long context with caching, and native audio for live sessions. ADK Java exposes most of them, but not all in the same place. Some are plain request configuration, some are tool singletons, one is an ADK rewrite that only triggers for model names starting with gemini-2, and caching is configured on the app rather than the agent.
This page maps each feature to the ADK Java class that controls it, shows the code, and lists the ways each one fails. Class and method names were read from the google/adk-java source (the latest release at the time of writing is v1.10.1, September 2026) and the java-genai javadoc. Model facts change faster than code; where a number comes from Google's model documentation rather than source, the text says so and you should re-check it.
Three layers decide whether a feature works
A feature is usable from ADK Java only when three layers agree. The model must support it: thinking budgets exist on 2.5 models, not on 2.0 Flash. The google-genai Java SDK must have a type for it, because ADK builds requests from com.google.genai.types. And ADK must pass it through or wire it up. When something silently does nothing, find which layer dropped it.
On lifecycle: on the day this was written Google's deprecations page listed no shutdown date for the stable gemini-2.5-pro, gemini-2.5-flash and gemini-2.5-flash-lite IDs, while several 2.5 preview IDs, including the original 2.5 live preview, have already been shut down. Pin stable IDs, and check that page before every release.
| Feature | How ADK Java exposes it | Where |
|---|---|---|
| Thinking budget, thought summaries | ThinkingConfig inside GenerateContentConfig | LlmAgent.builder().generateContentConfig |
| Google Search grounding | GoogleSearchTool.INSTANCE | tools(...) |
| URL context | UrlContextTool.INSTANCE | tools(...) |
| Server-side code execution | BuiltInCodeExecutionTool.INSTANCE | tools(...) |
| Structured output plus tools | outputSchema, rewritten to set_model_response on gemini-2 | OutputSchema request processor |
| Explicit context caching | ContextCacheConfig | App.builder().contextCacheConfig |
| Live audio | speechConfig, responseModalities, transcription | RunConfig, runLive |
Thinking budgets
Gemini 2.5 models reason before answering, and the tokens spent reasoning are billed as output tokens. The thinkingBudget field caps them. In ADK Java you set it on the agent's GenerateContentConfig, which the Basic request processor copies into every model request.
import com.google.adk.agents.LlmAgent;
import com.google.genai.types.GenerateContentConfig;
import com.google.genai.types.ThinkingConfig;
LlmAgent planner = LlmAgent.builder()
.name("refund_planner")
.model("gemini-2.5-flash")
.instruction("Decide whether the refund policy allows this request. Explain briefly.")
.generateContentConfig(GenerateContentConfig.builder()
.temperature(0.2f)
.thinkingConfig(ThinkingConfig.builder()
.thinkingBudget(2048) // cap reasoning tokens; 0 turns thinking off on Flash
.includeThoughts(true) // return thought summaries as parts marked thought=true
.build())
.build())
.build();The allowed values differ by model. Per Google's thinking documentation at the time of writing: 2.5 Pro accepts 128 to 32,768 and cannot turn thinking off; 2.5 Flash accepts 0 to 24,576, with 0 disabling thinking; 2.5 Flash-Lite accepts 512 to 24,576, thinks only when asked, and 0 keeps it off. A budget of -1 asks for dynamic thinking, where the model decides. Confirm these against the current documentation before you hard-code a value; ADK does not check the value, so a budget outside the range may be rejected or adjusted by the API.
Choose the budget per agent, not per app. A router agent that picks one of three sub-agents gains little from thinking and pays latency for it; a planner that weighs policy clauses gains a lot. Newer models also accept a thinkingLevel on the same builder; for 2.5, use the token budget.
Reading thoughts, grounding and usage from events
With includeThoughts(true), the response contains thought summaries as ordinary text parts whose thought() flag is true. Every UI and log consumer must check that flag, or summaries of internal reasoning end up shown to the user as part of the answer. Usage arrives on the event as GenerateContentResponseUsageMetadata, where thoughtsTokenCount reports reasoning tokens separately from answer tokens.
runner.runAsync(userId, sessionId, Content.fromParts(Part.fromText(question)))
.blockingForEach(event -> {
event.content().flatMap(Content::parts).ifPresent(parts -> {
for (Part part : parts) {
boolean thought = part.thought().orElse(false);
part.text().ifPresent(t -> (thought ? thoughtLog : answer).append(t));
}
});
event.usageMetadata().ifPresent(u -> metrics.record(
u.promptTokenCount().orElse(0),
u.thoughtsTokenCount().orElse(0),
u.candidatesTokenCount().orElse(0),
u.cachedContentTokenCount().orElse(0)));
event.groundingMetadata().ifPresent(g -> citations.add(g));
});ADK keeps thought parts in the session history but treats a part that is only a thought as invisible when it builds the next request's contents. Parts carrying a thought signature are never dropped, because the API expects them back on the next turn; the Gemini model class sends history with thought stripping turned off for the same reason. That matters most for function calling, covered in Gemini function calling in ADK Java. The practical rule: never rebuild history yourself from text only, or you lose the signatures.
Built-in tools and the one-per-agent rule
Google Search, URL context and code execution run on Google's side. In ADK Java each is a singleton whose processLlmRequest adds a Tool entry to the request config; nothing executes in your JVM. BuiltInCodeExecutionTool logs a warning if the model is not a Gemini model, because nothing else can honour it.
The constraint that shapes designs is in the ADK documentation: in Java, an agent that uses Google Search or Gemini code execution cannot use any other tool. The workaround is to give the built-in tool its own agent and call that agent as a tool. ADK even ships GoogleSearchAgentTool.create(model), which builds exactly that wrapper around GoogleSearchTool.INSTANCE.
import com.google.adk.tools.AgentTool;
import com.google.adk.tools.GoogleSearchTool;
import com.google.adk.tools.UrlContextTool;
// A built-in tool cannot share an agent with other tools, so give each its own agent.
LlmAgent searcher = LlmAgent.builder()
.name("web_researcher")
.model("gemini-2.5-flash")
.description("Searches the web and returns cited findings.")
.instruction("Answer with facts from Google Search and keep the source titles.")
.tools(GoogleSearchTool.INSTANCE)
.build();
LlmAgent reader = LlmAgent.builder()
.name("page_reader")
.model("gemini-2.5-flash")
.description("Reads the URLs it is given and summarises them.")
.tools(UrlContextTool.INSTANCE)
.build();
LlmAgent root = LlmAgent.builder()
.name("vendor_analyst")
.model("gemini-2.5-pro")
.instruction("Assess the vendor. Use web_researcher for news, page_reader for its docs,"
+ " and lookupContract for our internal terms.")
.tools(AgentTool.create(searcher), AgentTool.create(reader),
FunctionTool.create(ContractTools.class, "lookupContract"))
.outputSchema(VENDOR_ASSESSMENT_SCHEMA)
.build();Search results come back with grounding metadata on the event: queries issued and source chunks. Keep it if your product shows citations; once the sub-agent's answer has been summarised into the root agent's reply, the link between a sentence and its source is gone.
Structured output with tools on gemini-2 models
The root agent above asks for both tools and an output schema. ADK treats Gemini 2.x as unable to take a response schema and function tools in the same request (the source comment calls it a current limitation of 2.x models on Vertex AI), so it rewrites the request. The OutputSchema request processor checks the model name. If the agent has an output schema and tools, and the name matches ^gemini-2\. after any projects/.../models/ or apigee/ prefix is removed, it adds a set_model_response function whose parameters are your schema, and appends an instruction telling the model to finish by calling it.
When the model calls it, SetModelResponseTool validates the arguments against the schema. A failure returns an error message to the model so it can retry; a success is recorded on the event actions and turned into a final response event that looks like ordinary model output. The final response is therefore JSON that ADK has validated against your schema, and a malformed attempt costs a retry turn rather than an exception in your code.
The failure mode is the name check. If you call a tuned model or a custom endpoint whose name does not start with gemini-2., ADK assumes the model can take schema and tools natively, sends both, and the API can reject the request. Use the real model ID, or wrap the call so the name matches. One more cost: the extra function call is an extra model turn, so a structured answer after tools takes one more round trip than a plain one.
Long context and context caching
Gemini 2.5 models accept very long inputs, and agents fill them quickly: system instruction, tool declarations and a long session history are resent on every call. Implicit caching on Google's side can discount repeated prefixes automatically; ADK Java adds explicit caching configured once for the whole app.
import com.google.adk.agents.ContextCacheConfig;
import com.google.adk.apps.App;
import java.time.Duration;
App app = App.builder()
.name("vendor_analysis")
.rootAgent(root)
// reuse one cache for up to 10 invocations, 30 minute TTL, skip requests under 4,096 tokens
.contextCacheConfig(new ContextCacheConfig(10, Duration.ofMinutes(30), 4096))
.build();
Runner runner = Runner.builder()
.app(app)
.sessionService(new InMemorySessionService())
.artifactService(new InMemoryArtifactService())
.build();The ContextCacheConfig defaults are 10 invocations per cache, a 1,800-second TTL and a 0-token minimum, which caches even small requests. Raise the minimum: cache storage is billed, and the provider requires a minimum prefix size of its own, so tiny caches cost more than they save. Check whether it works by reading cachedContentTokenCount from usage metadata; if it stays at zero, the prefix is changing between calls, usually because the instruction contains a timestamp or tools are registered in a different order. Token accounting across providers is covered in token counting across models.
Live audio through RunConfig
Native audio sessions use runner.runLive with a live request queue instead of runAsync. The 2.5-era live settings live on RunConfig: speechConfig, responseModalities, inputAudioTranscription and outputAudioTranscription, which the Basic processor copies into the live connect config. Live model IDs have changed more often than any other Gemini family, and the 2.5 live preview has been shut down, so take the model ID from Google's current live models list rather than from an example. How the event stream behaves under backpressure is covered in streaming in ADK Java.
Worked example: one vendor assessment request
- The runner sends the root agent's request: instruction, history, three tool declarations, and, because the model is gemini-2.5-pro and an output schema is set, a fourth declaration for set_model_response.
- The model thinks within its budget, then calls web_researcher. The root agent is where a non-zero budget pays off, since it weighs news, documentation and contract terms; the two sub-agents mostly fetch and summarise, so a small budget or none suits them. ADK runs the sub-agent, whose own request carries only the Google Search built-in, so the one-per-agent rule holds.
- The sub-agent's answer, with grounding metadata on its events, returns as a function response. The root agent calls lookupContract, a normal Java function.
- The model calls set_model_response. A missing required field returns a validation error, and the model calls it again with the field filled.
- The validated JSON becomes the final response event. Usage metadata shows prompt, thought, answer and cached tokens for each of the model calls.
Failure modes
- Thought summaries shown to users. A UI that concatenates all text parts ignores the thought flag.
- Budget out of range. A Flash budget copied to Pro as 0, or to Flash-Lite as 256, is outside the documented range; the API may reject or adjust it, and ADK will not warn you.
- Built-in tool mixed with function tools. Requests fail or the built-in tool is ignored; split agents.
- Schema plus tools on a renamed model. The name check misses and the API rejects the combined request.
- Cache never hits. A dynamic value early in the instruction changes the prefix on every call.
- Cost surprise. Thinking tokens are billed as output; a dashboard that tracks only candidate tokens under-reports spend.
What to do next
- List every agent with its model ID, and check each ID against the deprecations page.
- Set an explicit thinking budget per agent, within the documented range for its model.
- Make every event consumer branch on part.thought() before displaying text.
- Record thoughtsTokenCount and cachedContentTokenCount per agent on your metrics dashboard.
- Move each built-in tool into its own agent and call it through AgentTool.
- For agents with an output schema and tools, test a validation failure and confirm the retry.
- Turn on ContextCacheConfig with a realistic token minimum and verify cache hits in usage metadata.