When an ADK Java agent calls a tool, the Java method is the easy part. Between your FunctionTool and the answer sits a protocol: ADK translates your tools into Gemini function declarations, Gemini replies with structured function-call parts, ADK runs the tools and sends function-response parts back, and the loop repeats until the model answers in text. Most production bugs in tool-using agents live in that protocol, not in the tool body: a model that never stops calling, responses that do not line up with their calls, or a sudden HTTP 400 after upgrading to a Gemini 3 model.
This page explains the protocol as ADK Java drives it. How to write a tool is covered in writing a custom function tool, and how the runtime resolves and executes each call is covered in tool dispatch mechanics. Here we stay on the boundary between ADK and the model.
First principles: the model never runs your code
Function calling is a structured-output contract. You send the model a list of function declarations: a name, a description and a JSON-Schema-like description of the parameters. The model, instead of text, may return one or more functionCall parts, each naming a function and carrying an args object. Nothing is executed on Google's side. Your client, here ADK, executes the function and sends the result back as a functionResponse part with the same name, in a new user-role turn. The model then either calls more functions or writes its answer.
Three consequences follow. The model only knows what your declaration tells it, so the description is part of your program. The model can return several calls in one response, so the client must handle a list. And because the whole conversation is re-sent on every step, anything the client drops from history is gone from the model's point of view.
From a Java method to a declaration
FunctionTool.create reflects over a method and builds the declaration. The method name becomes the function name unless @Schema overrides it, the method-level @Schema(description = ...) becomes the description, and each parameter becomes a property. Parameters named toolContext are injected by ADK and excluded from the schema; for ADK to see parameter names at all, either annotate them with @Schema(name = ...) or compile with -parameters. A Map return value is sent as the response object; anything else is wrapped as {"result": value}.
public class OrderTools {
@Schema(description = "Look up one order. Returns status, totalCents and deliveredAt.")
public static Map<String, Object> lookupOrder(
@Schema(name = "orderId", description = "Order id such as ORD-12345") String orderId) {
Order o = OrderStore.find(orderId);
if (o == null) {
return Map.of("status", "error", "message", "No order " + orderId);
}
return Map.of("status", "ok", "state", o.state(), "totalCents", o.totalCents(),
"deliveredAt", String.valueOf(o.deliveredAt()));
}
@Schema(description = "Refund a delivered order. Amount in cents, at most the order total.")
public static Map<String, Object> refundOrder(
@Schema(name = "orderId", description = "Order id") String orderId,
@Schema(name = "amountCents", description = "Positive amount in cents") long amountCents) {
return RefundService.refund(orderId, amountCents);
}
}On the wire, the first tool becomes a declaration shaped like this (field names as in the Gemini REST API; the exact type spelling and required list ADK emits depend on its version):
{
"name": "lookupOrder",
"description": "Look up one order. Returns status, totalCents and deliveredAt.",
"parameters": {
"type": "OBJECT",
"properties": {
"orderId": { "type": "STRING", "description": "Order id such as ORD-12345" }
},
"required": ["orderId"]
}
}Write descriptions for the model, not for Javadoc: say when to call the tool, what the units are and what comes back. Java-specific schema traps (boxed types, generics, enums) are covered in the Java tools architecture page.
Controlling whether and which functions are called
Gemini's toolConfig.functionCallingConfig has two knobs. mode takes AUTO (the default: the model chooses between a call and text), ANY (the model must return a function call), NONE (no function calls, even though declarations are present) or VALIDATED (the model chooses, with its calls validated against the schema). allowedFunctionNames restricts which declared functions may be called. In ADK Java you set this through the agent's GenerateContentConfig, using the builders from the com.google.genai.types package:
GenerateContentConfig cfg = GenerateContentConfig.builder()
.toolConfig(ToolConfig.builder()
.functionCallingConfig(FunctionCallingConfig.builder()
.mode(FunctionCallingConfigMode.Known.AUTO)
.allowedFunctionNames("lookupOrder", "refundOrder")
.build())
.build())
.temperature(0.0f)
.build();
LlmAgent agent = LlmAgent.builder()
.name("refund_agent")
.model(System.getenv("ADK_MODEL")) // a Gemini model id from your config
.instruction("Help customers with refunds. Always look an order up before refunding it.")
.tools(FunctionTool.create(OrderTools.class, "lookupOrder"),
FunctionTool.create(OrderTools.class, "refundOrder"))
.generateContentConfig(cfg)
.build();Be careful with ANY in an agent. ADK's loop ends when the model answers without a function call; if every response must be a call, the model can never produce that final answer, and the invocation runs until something else stops it. Use ANY for a single forced step, for example setting it in a before-model callback on the first turn only, or pair it with an explicit finishing tool. NONE is useful for a summarising turn where you want text but do not want to rebuild the tool list. A low temperature makes tool selection more repeatable; it does not make it correct.
Reading the response: parallel calls and ids
A model response is a Content with a list of parts. It may hold text, one functionCall, or several: the model uses parallel calls when the calls do not depend on each other, such as looking up two orders at once. Sequential (compositional) calling is the model calling one function, reading the result, then calling another in the next step. ADK handles both; the event it emits exposes the calls through event.functionCalls() and, after execution, the responses through event.functionResponses().
Each call can carry an id, and the matching response should carry the same id so the model can pair them. When the model does not supply one, ADK generates a client-side id so its own bookkeeping, events, callbacks and long-running tools can refer to the call, and strips those generated ids again before sending history to Gemini. The practical rule is: when you build or rewrite function-response parts yourself, copy the id and name from the call; never invent a new id, and never reorder a parallel batch in a way that loses the pairing.
InMemoryRunner runner = new InMemoryRunner(agent, "refunds");
Session session = runner.sessionService().createSession("refunds", "user-42").blockingGet();
Content msg = Content.fromParts(Part.fromText("Refund ORD-100 and ORD-101 in full."));
runner.runAsync("user-42", session.id(), msg).blockingForEach(event -> {
for (FunctionCall call : event.functionCalls()) {
System.out.printf("CALL %s %s id=%s%n", call.name().orElse("?"),
call.args().orElse(Map.of()), call.id().orElse("-"));
}
for (FunctionResponse r : event.functionResponses()) {
System.out.printf("RESP %s id=%s%n", r.name().orElse("?"), r.id().orElse("-"));
}
if (event.finalResponse()) {
System.out.println("TEXT " + event.stringifyContent());
}
});
Thought signatures: the Gemini 3 requirement
Gemini 3 models think before they answer, and when they return function calls they attach a thought signature: an opaque, encrypted token on the part that lets the model resume its reasoning when you send the history back. Google's documentation is explicit about three rules. With parallel calls, only the first functionCall part in the response carries the signature. With sequential steps, each step's call carries its own, and all of them must be returned. If a required signature is missing, Gemini 3 returns HTTP 400 with a message of the form Function call ... in the N content block is missing a thought_signature.
ADK handles this for you as long as the model's Content survives untouched from the response into session history and back into the next request. It breaks in three predictable places. A custom session service that serialises parts by hand, keeping text, calls and responses but not the signature bytes. An after-model callback that rebuilds the model's content with fresh Part.fromFunctionCall(...) parts instead of copying the originals with toBuilder(). And history trimming or summarisation that drops the model turn but keeps its function responses. Each works on older models and fails the day you switch model ids, which makes it a migration bug to test for explicitly: run a two-step tool conversation against a Gemini 3 model through your real session store before rolling out.
Streaming and partial events
In streaming mode ADK calls generateContentStream, and a function call may arrive split across chunks. ADK's streaming aggregator buffers text and thought text, merges fragmented function-call arguments, and emits partial events followed by one complete event. Tools run only on the complete call. If you consume events yourself, check event.partial() and never act on a function call from a partial event; its arguments may be truncated JSON. The java-genai FunctionCallingConfig also has a streamFunctionCallArguments flag; which models honour it is not something this page can confirm, so test before depending on it.
Where each model call sits inside ADK's invocation, and how the request is assembled from session, instruction and tools, is covered in model call orchestration.
Worked example: two refunds in one turn
The user asks to refund ORD-100 and ORD-101 in full. A correct trace looks like this:
user : "Refund ORD-100 and ORD-101 in full."
model : functionCall lookupOrder{orderId: ORD-100} (thoughtSignature on this part)
functionCall lookupOrder{orderId: ORD-101}
adk : runs both, emits one event with two functionResponse parts, ids/names paired
model : functionCall refundOrder{orderId: ORD-100, amountCents: 4599} (signature)
functionCall refundOrder{orderId: ORD-101, amountCents: 1250}
adk : runs both refunds, returns {status: ok, refundId: ...} for each
model : "Both orders are refunded: 45.99 and 12.50."Three things to verify in your logs. The lookups ran before the refunds, because the instruction said so and the amounts came from the lookup results; if the model refunds with guessed amounts, tighten the description. Each response pairs with its call. And the second model request carried the first response's signature, which you can confirm by the absence of a 400 on Gemini 3. Refunds are side effects, so guard them with idempotency keys and a before-tool callback that checks policy; see ADK Java callbacks for where that check belongs.
Failure modes
| symptom | likely cause | fix |
|---|---|---|
| HTTP 400, missing thought_signature | session store or callback dropped signature bytes | persist whole parts; copy with toBuilder(); test on a Gemini 3 model |
| agent never answers, hits call limit | mode ANY left on for every turn | force only the first turn, or add a finish tool |
| model calls a tool that is not registered | allowedFunctionNames or instruction mentions a removed tool | keep names, declarations and instructions in one place |
| arguments in wrong units or format | vague description or parameter schema | state units, formats and examples in descriptions |
| responses mismatched with calls | custom code rebuilt parts without ids or names | copy id and name from each call |
| tool ran twice with truncated args | code acted on a partial streaming event | ignore partial events for execution |
| token cost climbs per turn | large tool results re-sent every step | return compact maps; summarise or trim old results |
Operational guidance and trade-offs
- Keep the tool list small per agent. Every declaration is sent on every request, costs tokens and adds a candidate the model can pick wrongly; split large tool sets across sub-agents.
- Log, per model call, the declared tool names, the mode, the calls returned and the latency. That record answers most why-did-it-do-that questions.
- Return errors as data (
statusandmessage) rather than throwing, so the model can recover in the next step. - Prefer
AUTOwith a precise instruction over forced modes; forcing trades flexibility for determinism and is best confined to one step. - Treat a model-id change as a protocol change: rerun your multi-step tool tests, because signature handling, parallel-call behaviour and argument quality all vary by model.
What to do next
- List every tool on each agent and rewrite descriptions to state when to call it, units and return shape.
- Set an explicit
FunctionCallingConfigand document why each agent uses its mode and allowed names. - Add a test that runs a two-step, parallel-call conversation through your real session service against a Gemini 3 model.
- Audit callbacks and custom session storage for code that rebuilds parts; switch it to copying originals.
- Log calls, ids and responses per invocation and alert on 400s from the model endpoint.
- Guard side-effecting tools with idempotency keys and a before-tool policy check.