Agents spend most of their latency and much of their token budget on round trips: the model decides to call a tool, the runtime executes it, the result is appended to the conversation, and the model is invoked again with the whole history. When a task touches many items (compare twelve orders, check stock for thirty SKUs, look up every attendee of a meeting) the number of round trips becomes the cost. The tool batching pattern designs tools so that one call can carry many items, and designs the runtime side so that many small calls collapse into few backend requests.
This article shows how to do that in the Agent Development Kit for Java: the three shapes a multi-item lookup can take, how to write a list-accepting FunctionTool with per-item results, why the size cap must live in code, how to coalesce parallel calls into one backend request, how big results eat context, and when batching is the wrong idea. The runtime mechanics of several calls in one model response are covered in how ADK Java dispatches tool calls; this page is about designing for batching from the tool's side.
Three shapes of a multi-item lookup
Take a support agent asked: "Which of my last twelve orders have shipped?" There are three ways the work can unfold.
- A: one call per turn. The model calls
getOrder, reads the result, calls it again for the next id, and so on. Twelve model invocations, each replaying the growing history. This is what you get by default from a single-item tool and a cautious model. - B: parallel calls in one response. Models that support parallel function calling can emit twelve
getOrdercalls in one response. One model round trip, but twelve call-and-response pairs in the history and twelve backend requests. - C: one batch call. The tool accepts a list, and the model sends all twelve ids in one call. One round trip, one compact result, and the tool decides how to hit the backend efficiently.
Rough, illustrative numbers: if each model invocation replays 4,000 tokens of history and takes 1.5 seconds, shape A costs about twelve invocations, around 18 seconds and over 50,000 input tokens as the history grows. Shapes B and C cost one extra invocation; C also makes one backend request instead of twelve and returns a smaller result, because items share field names and a summary.
A list-accepting tool
ADK Java builds a function declaration from a method's signature by reflection. A List<String> parameter becomes an ARRAY schema whose item type is derived from the element type, so a list-accepting tool needs no special API. What the reflection does not do is emit minItems or maxItems: there is no annotation for them, so the limit must be stated in the description, where the model reads it, and enforced in code, where it is actually guaranteed. Even a hand-written schema with a maximum is a hint to the model, not a contract.
package com.example.orders;
import com.google.adk.tools.Annotations.Schema;
import java.util.*;
public final class OrderBatchTools {
static final int MAX_IDS = 25; // what the model is told
static final int BACKEND_CHUNK = 10; // what the order service accepts per request
private static OrderClient client; // set at startup; see the coalescer section
@Schema(description = "Look up the status of SEVERAL orders in one call. Pass every order id "
+ "you need at once (up to 25) instead of calling once per order. Returns one result per "
+ "id with status ok, not_found or invalid, plus any ids that were deferred.")
public static Map<String, Object> getOrders(
@Schema(name = "orderIds", description = "Order ids such as ORD-12345, at most 25")
List<String> orderIds) {
if (orderIds == null || orderIds.isEmpty()) {
return Map.of("status", "error", "error_code", "EMPTY",
"message", "Pass at least one order id.");
}
// Normalise and dedupe, preserving the model's order.
LinkedHashSet<String> ids = new LinkedHashSet<>();
for (String raw : orderIds) {
if (raw != null) ids.add(raw.trim().toUpperCase(Locale.ROOT));
}
List<String> all = new ArrayList<>(ids);
List<String> accepted = all.subList(0, Math.min(all.size(), MAX_IDS));
List<String> deferred = all.subList(accepted.size(), all.size());
Map<String, Map<String, Object>> results = new LinkedHashMap<>();
List<String> valid = new ArrayList<>();
for (String id : accepted) {
if (id.matches("ORD-\\d{5}")) valid.add(id);
else results.put(id, Map.of("status", "invalid", "message", "expected ORD-12345"));
}
for (int i = 0; i < valid.size(); i += BACKEND_CHUNK) {
List<String> chunk = valid.subList(i, Math.min(i + BACKEND_CHUNK, valid.size()));
try {
Map<String, Order> found = client.findAll(chunk); // one backend request
for (String id : chunk) {
Order o = found.get(id);
results.put(id, o == null
? Map.of("status", "not_found")
: Map.of("status", "ok", "state", o.state(), "eta", String.valueOf(o.eta())));
}
} catch (RuntimeException e) { // one chunk failing
for (String id : chunk) { // must not sink the batch
results.put(id, Map.of("status", "error", "retryable", true));
}
}
}
long ok = results.values().stream().filter(r -> "ok".equals(r.get("status"))).count();
return Map.of(
"status", "success",
"summary", Map.of("requested", all.size(), "ok", ok,
"other", results.size() - ok, "deferred", deferred.size()),
"results", results,
"deferred_ids", List.copyOf(deferred),
"note", deferred.isEmpty() ? "" : "Call getOrders again with deferred_ids.");
}
}Registration is the same as for any static tool: FunctionTool.create(OrderBatchTools.class, "getOrders"), passed to the agent's .tools(...). Compile with -parameters so parameter names survive reflection; writing a function tool covers that setup step by step.
Designing the result
Four design decisions in that code are the pattern; the rest is plumbing.
- Per-item status, not all-or-nothing. One bad id must not turn eleven good lookups into an error. The top-level
statussays the call ran; each item says what happened to it. The model can then answer about the eleven and ask about the one. This is the same idea as wrapping tool errors as data, applied item by item. - Deferral instead of silent truncation. The thin version of this pattern slices the list to the maximum and moves on, so items 26 onwards vanish and the model believes it checked them. Returning
deferred_idsmakes the cap visible and gives the model the exact follow-up call. - Keyed results. Keying results by id rather than by position means dedupe, reordering and partial failure cannot misalign answers with questions, and the model does not have to count.
- A summary first. Counts at the top let the model answer "how many have shipped" without scanning every item, and give your logs one line to aggregate.
Getting the model to use the batch shape is a prompt problem, not a code problem. State it in the tool description ("pass every id at once"), and if you also keep a single-item tool, describe it as the fallback for one item. Two tools that overlap without guidance invite the model to alternate between them.
When the model sends many small calls: coalescing
You cannot always change the tool's shape: it may be shared, generated from an OpenAPI spec, or the model may still emit parallel single calls. The runtime then decides how those calls execute. In the current adk-java source, RunConfig's ToolExecutionMode has four values, NONE, SEQUENTIAL, PARALLEL and PARALLEL_SUBSCRIBE. The builder's default is NONE, which is executed like PARALLEL: all calls are subscribed eagerly and results are kept in request order and merged into one event. Eager subscription only overlaps tools that are genuinely asynchronous (they return a Single backed by I/O or another scheduler); blocking tools still run one after another unless you choose PARALLEL_SUBSCRIBE, which subscribes each on a worker thread and requires thread-safe tools. Check the enum in the version you depend on, because it has grown over time.
Concurrency removes the latency of shape B but not its backend load: twelve concurrent requests are still twelve requests. A coalescer fixes that. Each single-item call registers its id and receives a future; a short window (a few milliseconds) or a size threshold flushes all pending ids as one backend request, and each future completes with its own item. This is the data-loader pattern familiar from GraphQL servers.
public final class OrderCoalescer {
private final OrderClient client;
private final ScheduledExecutorService timer = Executors.newSingleThreadScheduledExecutor();
private final Map<String, CompletableFuture<Optional<Order>>> pending = new LinkedHashMap<>();
private ScheduledFuture<?> flushTask;
public OrderCoalescer(OrderClient client) { this.client = client; }
public synchronized CompletableFuture<Optional<Order>> get(String id) {
CompletableFuture<Optional<Order>> f = pending.computeIfAbsent(id, k -> new CompletableFuture<>());
if (pending.size() >= 10) flushNow(); // size threshold
else if (flushTask == null) flushTask = timer.schedule(this::flushNow, 5, TimeUnit.MILLISECONDS);
return f;
}
private synchronized void flushNow() {
if (flushTask != null) { flushTask.cancel(false); flushTask = null; }
if (pending.isEmpty()) return;
Map<String, CompletableFuture<Optional<Order>>> batch = new LinkedHashMap<>(pending);
pending.clear();
client.findAllAsync(List.copyOf(batch.keySet())) // one request
.orTimeout(2, TimeUnit.SECONDS)
.whenComplete((found, err) -> batch.forEach((id, f) -> {
if (err != null) f.completeExceptionally(err);
else f.complete(Optional.ofNullable(found.get(id)));
}));
}
}
// The single-item tool returns a Single, so PARALLEL mode subscribes all calls eagerly
// and every call reaches the coalescer inside the same window.
@Schema(description = "Look up one order. For several orders prefer getOrders.")
public static Single<Map<String, Object>> getOrder(
@Schema(name = "orderId", description = "Order id such as ORD-12345") String orderId) {
return Single.fromCompletionStage(COALESCER.get(orderId))
.map(o -> o.<Map<String, Object>>map(x -> Map.of("status", "ok", "state", x.state()))
.orElse(Map.of("status", "not_found")))
.onErrorReturn(e -> Map.of("status", "error", "retryable", true));
}Returning Single is what makes the window work: FunctionTool accepts Single and Maybe return types, so the calls do not block the subscribing thread and all twelve arrive before the timer fires. A blocking implementation under the default mode would arrive one at a time and each would flush alone. Keep the coalescer per process and per tenant: merging two users' ids into one backend request is an authorisation bug waiting to happen. See async tool execution for the threading model behind this.
What grows with batch size
Batching moves cost rather than removing it, and three costs grow with batch size.
- Result tokens. Twenty-five full order records can be larger than the rest of the conversation. Return only the fields the model needs (a
fieldsparameter, or a fixed compact projection), summarise long text, and store bulky payloads as artifacts the model can reference instead of inline JSON. - Tail latency. A batch finishes when its slowest chunk does. Put a timeout on each chunk and report timed-out items as retryable errors rather than letting one slow shard hold the whole response.
- Backend pressure. One model call can now ask for hundreds of lookups. The cap, the chunk size and a concurrency limit per tenant are what stop an enthusiastic model from becoming a load test; rate limiting is covered in rate limiting parallel agents.
Writes need more care than reads. A batched write should carry an idempotency key per item so that a retry after a partial failure does not apply the successful items twice; it should report per-item outcomes exactly like the read; and anything irreversible (refunds, cancellations, emails) should go through confirmation, which FunctionTool.create(..., true) requests, with the full list shown to the user. Do not batch writes whose order matters unless the tool applies them in the given order and says so.
Testing batching
Test the tool as plain Java first: empty list, duplicates, mixed valid and invalid ids, more than the cap, and a backend chunk that throws. Each test asserts on the response map, so they run in milliseconds. Then test the agent's behaviour: send a prompt that names several items and assert that the emitted events contain one getOrders call rather than many. Run that check across a small set of prompts, because whether a model chooses the batch shape is probabilistic and drifts between model versions. In production, record the batch size distribution and the deferred count per call; a rising deferred count means the cap is too low or the model is being asked for bulk work an agent should not be doing interactively.
Failure modes
Failure modes seen in practice:
- Silent truncation: items beyond the cap are dropped and the model reports on them anyway.
- Positional results misaligned after dedupe, so order 7's status is reported against order 8.
- A single invalid id fails the whole batch and the model retries everything, including the items that had succeeded.
- Blocking single-item tools under the default mode, so the coalescer never receives more than one id per window.
- Oversized results pushing earlier instructions out of the context window.
- Batched writes without idempotency keys, double-applied on retry.
Trade-offs
Batch-shaped tools cut round trips and backend requests, but they are more complex to write and harder for a model to use when the items are not known up front. If the agent discovers ids one at a time (each answer suggests the next question), batching cannot help and sequential calls are correct. Coalescing keeps tool interfaces simple and works with any model, at the cost of a small added latency window and shared mutable state you must make thread-safe. A reasonable default: provide a batch tool for any lookup the user is likely to ask about in bulk, keep a single-item tool with a description that points to the batch one, coalesce underneath both, and measure.
What to do next
- Find your agent's tools that are called more than three times per turn in traces; those are batching candidates.
- Add a list-accepting variant with a stated cap, per-item statuses, keyed results, a summary and an explicit deferred list.
- Enforce the cap and chunk size in code, with a timeout per chunk.
- Make single-item tools asynchronous and put a per-tenant coalescer behind them.
- Check which
ToolExecutionModeyour runs use and whether your tools are safe under it. - Add an agent-level test that the batch tool is chosen for multi-item prompts, and track batch size and deferred counts in production.