Most useful agent tasks take more than one tool call. Refunding an order means finding it, checking policy and paying out. Syncing a catalogue means fetching, cleaning and writing. In ADK for Java there are four places such a chain can live: the model can decide each next call, a single Java tool can run every step itself, a workflow agent can run fixed stages, or a whole sub-agent can be wrapped as one tool. Each choice moves cost, latency, control and failure handling to a different place.
This page is about that choice and about the plumbing between steps: how one step's output reaches the next without passing through the model, how to stop the model calling steps in the wrong order, and how to bound and recover a chain that fails halfway. It assumes you can already write a single tool; the mechanics of one dispatch are covered in how ADK dispatches a tool call. API names here were checked against the ADK documentation and the adk-java sources in October 2026; check them against the release you build with.
First principles: what a chain costs
When the model drives a chain, every step is a round trip: the model returns a function call, ADK runs the tool, appends the result to the session, and calls the model again with the whole conversation so far. A three-step chain is at least four model calls for one user turn, the last one producing the answer. Each call adds latency, input tokens that grow with every result appended, and a fresh chance for the model to pick the wrong tool or invent an argument.
When Java code drives the chain, the model sees one call and one result, but cannot adapt the sequence to what it finds. That is the central trade: put decisions that need judgement in the model, and put sequences that never change in code.
Option 1: let the model chain the calls
Register each step as its own FunctionTool and describe in the instruction how they relate. ADK builds a function declaration from each public static method, using the @Schema annotations for parameter names and descriptions. Compile with the -parameters flag so parameter names survive. A ToolContext parameter is injected by ADK rather than filled by the model, and gives the tool access to session state through toolContext.state().
import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.ToolContext;
import java.util.List;
import java.util.Map;
import java.util.UUID;
public final class OrderTools {
/** Step 1: look up an order. Stores the full record in state and returns a small summary. */
public static Map<String, Object> findOrder(
@Schema(name = "orderId", description = "Order id, e.g. ORD-1042") String orderId,
@Schema(name = "toolContext") ToolContext toolContext) {
Order order = OrderRepo.find(orderId); // your own data access
if (order == null) {
return Map.of("status", "error", "error", "No order " + orderId);
}
String handle = "order:" + order.id();
toolContext.state().put(handle, order.toMap()); // full record stays server-side
return Map.of("status", "ok", "orderHandle", handle,
"total", order.total(), "refundable", order.refundable());
}
/** Step 2: refund. Accepts only a handle produced by findOrder, never raw order data. */
public static Map<String, Object> issueRefund(
@Schema(name = "orderHandle", description = "Handle returned by findOrder") String orderHandle,
@Schema(name = "amount", description = "Amount in the order currency") double amount,
@Schema(name = "toolContext") ToolContext toolContext) {
Object stored = toolContext.state().get(orderHandle);
if (!(stored instanceof Map<?, ?> order)) {
return Map.of("status", "error", "error", "Unknown handle; call findOrder first");
}
if (amount <= 0 || amount > ((Number) order.get("total")).doubleValue()) {
return Map.of("status", "error", "error", "Amount must be between 0 and the order total");
}
String key = "refund:" + orderHandle + ":" + amount; // idempotency key for retries
String refundId = Payments.refundOnce(key, (String) order.get("id"), amount);
return Map.of("status", "ok", "refundId", refundId);
}
}Two design choices in this code carry the whole pattern. First, handles instead of payloads: findOrder puts the full order into session state and returns a short handle plus the few fields the model needs to decide. The model passes the handle to the next step. The order record never makes a round trip through the model, so it cannot be paraphrased, truncated or altered, and it does not inflate every later prompt. Second, each step validates its own preconditions: issueRefund rejects an unknown handle and an out-of-range amount, and returns the problem as a structured error the model can read and act on rather than throwing.
Use this option when the sequence genuinely depends on intermediate results, for example when the model must decide whether to refund, replace or escalate after seeing the order.
Guarding order and inputs with callbacks
An instruction saying call A before B is a request, not a guarantee. Enforce the rule in code with a tool callback. In ADK for Java, LlmAgent.builder() accepts a beforeToolCallback and an afterToolCallback, each with a synchronous variant. The before callback receives the invocation context, the tool, the argument map and the ToolContext. The synchronous form returns Optional<Map<String, Object>>; the asynchronous form returns an RxJava Maybe of the same map. Returning a value skips the real tool and uses that map as its result; returning empty lets the call proceed.
The refund guard below uses this: without a known handle, the refund never runs and the model is told why. After callbacks suit redaction, per-step metrics and normalising error shapes. Keep callbacks fast; they run on every call. More patterns are in ADK callbacks.
Option 2: hide the chain inside one tool
When the steps always run in the same order and no step needs judgement, write one tool that runs them all. The model sees one capability with one description, makes one call, and receives one result.
/** One tool, three internal steps. The model sees a single capability. */
public static Map<String, Object> syncCatalogue(
@Schema(name = "supplierId", description = "Supplier to sync") String supplierId) {
List<Item> raw;
try {
raw = SupplierApi.fetch(supplierId); // step 1: network, may time out
} catch (Exception e) {
return Map.of("status", "error", "retryable", true, "error", "Supplier unreachable");
}
List<Item> clean = Normaliser.apply(raw); // step 2: pure, no side effects
String batchId = UUID.randomUUID().toString();
try {
Catalogue.replaceAll(supplierId, clean, batchId); // step 3: one transaction
} catch (Exception e) {
return Map.of("status", "error", "retryable", false,
"error", "Write failed; catalogue unchanged");
}
return Map.of("status", "ok", "items", clean.size(), "batchId", batchId);
}The failure semantics are the part to design, not the happy path. This tool is arranged so that the only step with a lasting effect is the last one, and that step is a single transaction, so the catalogue is either fully replaced or untouched. The error results say whether a retry makes sense. If your chain has several side-effecting steps that cannot share a transaction, such as a payment and a shipment in different systems, the tool must either compensate the earlier steps before returning an error or report exactly which steps completed, and every side-effecting step needs an idempotency key so a retry does not repeat it; see idempotent agent operations.
The cost is visibility: the model cannot work around a failed step, and traces show one span unless you instrument inside the tool.
Option 3: fixed stages with SequentialAgent
Sometimes each step needs a model, but the order of steps never changes: classify, then draft, then review. A SequentialAgent runs its sub-agents in order over the same invocation context and session state. Each LlmAgent can save its final response into state with outputKey, and a later agent's instruction can read it with {key} templating, which ADK fills in from state before calling the model.
Each stage has its own instruction, tools and output, and can be evaluated separately, but no stage can skip ahead or go back. When a stage fails partway through a run that has already caused effects, the pipeline does not roll anything back for you; rollback and compensation for sequential agents shows how to record and run compensations.
Option 4: a sub-agent as a single tool
AgentTool.create(agent) wraps any agent as a tool. When the parent model calls it, ADK runs the wrapped agent in its own runner and session, seeded with the caller's state, applies state changes from the sub-run back to the caller, and returns the sub-agent's final output as the tool result. Overloads take a skipSummarization flag and an option to include plugins.
This encapsulates reasoning: the coordinator decides that a refund is needed, and the refund agent's own chain runs behind one call. Its model calls are invisible in the coordinator's prompt but still cost tokens and time, so budget for them, and keep its description precise, because the parent chooses it from that alone.
Wiring it together
The snippet below assembles all four options: the guarded refund agent, a two-stage pipeline, a coordinator that uses the refund agent as a tool, and a run-wide limit.
import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.RunConfig;
import com.google.adk.agents.SequentialAgent;
import com.google.adk.tools.AgentTool;
import com.google.adk.tools.FunctionTool;
import java.util.Map;
import java.util.Optional;
FunctionTool findOrder = FunctionTool.create(OrderTools.class, "findOrder");
FunctionTool issueRefund = FunctionTool.create(OrderTools.class, "issueRefund");
// Model-driven chain, guarded: issueRefund is refused unless findOrder ran first in this session.
LlmAgent refundAgent = LlmAgent.builder()
.name("refund_agent")
.model(MODEL)
.description("Looks up an order and issues a refund within policy.")
.instruction("Call findOrder, then issueRefund with the orderHandle it returned. "
+ "Never invent a handle. Report the refundId.")
.tools(findOrder, issueRefund)
.beforeToolCallbackSync((invocation, tool, args, toolContext) -> {
if (tool.name().equals("issueRefund")
&& !toolContext.state().containsKey(String.valueOf(args.get("orderHandle")))) {
return Optional.of(Map.of("status", "error",
"error", "Refused: call findOrder first and pass its orderHandle"));
}
return Optional.empty(); // empty means: run the tool
})
.build();
// Fixed pipeline: each stage writes its answer to state under outputKey, the next reads {key}.
LlmAgent triage = LlmAgent.builder().name("triage").model(MODEL)
.instruction("Classify the customer message into REFUND, SHIPPING or OTHER.")
.outputKey("category").build();
LlmAgent reply = LlmAgent.builder().name("reply").model(MODEL)
.instruction("The category is {category}. Draft a reply for that category.")
.outputKey("draft").build();
SequentialAgent support = SequentialAgent.builder()
.name("support_pipeline").subAgents(triage, reply).build();
// Encapsulated chain: the refund agent becomes one tool for a coordinator.
LlmAgent coordinator = LlmAgent.builder().name("coordinator").model(MODEL)
.instruction("Use refund_agent for refunds; answer other questions yourself.")
.tools(AgentTool.create(refundAgent))
.build();
// Bound the whole run: every model call in the chain counts.
RunConfig runConfig = RunConfig.builder().maxLlmCalls(20).build();RunConfig.builder().maxLlmCalls(...) caps the number of model calls in one invocation; the documented default is 500, which is far more than most chains need. A model that keeps retrying a failing tool is the commonest runaway, and a low cap turns it into a bounded error. Set the cap per use case from measured traces, with headroom, not as a global guess.
Worked example: one refund, step by step
A customer writes: refund my order ORD-1042, it arrived broken. The coordinator's model calls refund_agent. Inside it, model call 1 returns a call to findOrder with orderId set to ORD-1042. The tool stores the order under order:ORD-1042 and returns total 84.00, refundable true and the handle. Model call 2 returns issueRefund with that handle and 84.0. The before callback finds the handle in state and lets it through; the tool checks the amount, derives the idempotency key and pays out once. Model call 3 writes the confirmation with the refund id, and the coordinator relays it.
Now the failure path. If model call 2 invents a handle, order:1042, the callback refuses with a readable error and the model retries with the real one. If the payment service times out, the tool returns a retryable error and the idempotency key means a retry pays once. If the model loops anyway, the call cap ends the run. Every protection lives in code, not in the instruction.
Failure modes
- Invented arguments between steps. The model reformats or guesses an id it should have copied. Pass handles, validate them, and refuse unknown ones in a callback.
- Context bloat. Large tool results appended to the session make every later model call slower and costlier. Return summaries and keep bulk data in state or external storage.
- Partial completion. Step 2 of 3 fails after step 1 changed something. Make the last step the only side-effecting one where you can; otherwise compensate or report exactly what completed.
- Duplicate effects on retry. The model, a timeout or a client retry runs a side-effecting step twice. Use idempotency keys derived from the request, not random ids.
- Runaway loops. The model retries forever or ping-pongs between tools. Cap model calls and return errors that say whether retrying is useful.
- Slow steps holding the turn. One slow dependency stalls the whole chain; give each tool a deadline as in tool timeout handling.
Choosing where the chain lives
| Need | Model-driven | Composite tool | SequentialAgent | AgentTool |
|---|---|---|---|---|
| Next step depends on results | Yes | Only via code branches | No | Yes, inside |
| Model calls per run | Steps + 1 | 1 + 1 | One or more per stage | Parent + inner |
| Atomicity control | Weak | Strong | Per stage | Weak |
| Testability | Needs evals | Unit tests | Per-stage evals | Evals on sub-agent |
| Typical use | Investigations, support | Fixed integrations | Content pipelines | Delegated specialities |
A good default: start with a composite tool for any fixed sequence, use model-driven chains only where the next step truly needs judgement, and reach for workflow agents and agent tools when stages need different instructions or tool sets.
What to do next
- List the multi-step tasks your agent performs and mark, for each step, whether it needs judgement. Move every sequence without judgement into a composite tool.
- Change any tool that returns a large record to store it in
toolContext.state()and return a handle plus the fields the model needs. - Add a
beforeToolCallbackSyncthat refuses out-of-order calls and unknown handles, and test it by feeding it the wrong order on purpose. - Give every side-effecting step an idempotency key and a structured error with a retryable flag.
- Set
maxLlmCallsin yourRunConfigfrom the measured call count of your longest legitimate chain, plus headroom. - Trace one real run end to end, count the model calls and tokens per chain, and decide whether any model-driven chain should become a pipeline or a composite tool.