A single LLM agent with thirty tools and a four-page instruction eventually stops working: tool selection gets noisy, the prompt becomes a negotiation between competing concerns, and every change risks every behaviour. The hierarchical, or supervisor, pattern splits that agent into a tree. A supervisor at the root understands the request and delegates; specialists below it own a narrow job with a short instruction and a small tool set; deterministic workflow agents handle the parts that should not be left to a model's judgement.

Google's Agent Development Kit for Java (ADK Java) builds this pattern into its core types rather than leaving it to convention. This article explains the mechanics as they exist in the current adk-java source: how the tree is formed, the two very different ways a parent can hand work to a child, how state flows between levels, how loops and sub-trees end, and how the runner decides who answers the next message. It finishes with a worked build, the failure modes that show up in production, and a checklist. Routing strategy itself, how to choose which specialist a request goes to, is covered in the agent router architecture; this page is about the tree that routing operates on.

Advertisement

The agent tree

Every agent in ADK Java extends BaseAgent, whose builders accept name, description and subAgents. Construction sets each child's parent pointer, so the system is a navigable tree: parentAgent(), rootAgent(), subAgents() and findAgent(name). Names are addresses: keep them unique, since they also appear as the author of every event.

The source documents one structural rule: an agent cannot be added to two different parents' sub-agent lists. The parent pointer is a single field, so an instance shared between two parents would belong to whichever parent was built last. Build a fresh instance for each place a specialist appears, using a factory method, rather than reusing one object.

The tree mixes model-driven LlmAgent nodes with workflow agents (SequentialAgent, ParallelAgent, LoopAgent) that run children in a fixed pattern with no model call. Use LLM nodes where judgement is needed and workflow nodes wherever the order of work is known.

Architecture of a supervisor tree

RunnerrunAsync(userId, sessionId, msg)support_supervisorLlmAgent: root of the treetransfer_to_agenttransfer_to_agentAgentTool callbilling_agentLlmAgent, outputKeytech_supportSequentialAgentpolicy_checkerfresh session, returns resultdiagnoseoutputKey diagnosisfix_writerreads {diagnosis}Session state and event logshared by transfer targets; AgentTool copies it in and merges deltas back
A support supervisor with two transfer targets (an LLM specialist and a sequential sub-pipeline) and one specialist wrapped as an AgentTool.

This is the shape of the worked build below. The runner enters at the supervisor, which can transfer the conversation to billing_agent or the tech_support pipeline, or call policy_checker as a tool and keep control. Transfer targets share the session state and event log; the AgentTool child runs in a temporary session.

Advertisement

Delegation mechanism one: transfer

When an LlmAgent has possible transfer targets, the framework adds a function named transfer_to_agent to the model request and appends instructions listing each target's name and description. If the model calls that function, the current agent's turn ends and the named agent takes over the same invocation, with the same session history. Control has moved: the specialist now talks to the user directly and its reply is the reply.

Which targets are offered is decided in the framework's agent-transfer request processor. An agent's own sub-agents are always offered. Its parent and its peers (the parent's other children) are offered only if the parent is itself an LlmAgent, and each can be switched off with disallowTransferToParent(true) and disallowTransferToPeers(true). Children of a SequentialAgent therefore cannot transfer sideways out of the pipeline, which is what you want: the pipeline decides the order, not the model.

Because the description is literally what the model reads when choosing a target, it is part of the interface. Write it as a contract: what the agent handles, what it does not, and what input it needs. Vague or overlapping descriptions are the main cause of misrouting, and changing one changes routing behaviour for the whole tree.

Delegation mechanism two: AgentTool

AgentTool.create(agent) wraps an agent so a parent can call it like a function and get a result back. The difference from transfer is control: the parent stays in charge, receives the child's final text as a tool result, and decides what to say. In the current source the tool takes a single request string argument when the child has no input schema, runs the child in a new InMemoryRunner with a temporary session created from a copy of the caller's state, merges any state changes the child made back into the caller's state, and returns the last event's text as {"result": ...}, or validates it against the child's output schema if one is set.

Consequences: the child sees only the request string and state, not the conversation, so the parent must pass everything it needs; the child's intermediate events stay out of the parent's event log; and the parent pays an extra model call to summarise the result unless you use AgentTool.create(agent, true), which skips summarisation.

transfer_to_agentAgentTool
Who answers the userThe childThe parent
Child sees historyYes, same sessionNo, only request plus copied state
StateShared directlyCopied in, deltas merged back
Events in parent logYes, authored by childOnly the tool call and result
Good forHanding off a whole conversationConsulting an expert mid-task

Worked build: a support supervisor

The build below creates the tree in the diagram. The model identifier comes from configuration; do not hardcode a model name that will change. Each specialist gets a factory so no instance is shared between parents.

import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.SequentialAgent;
import com.google.adk.tools.AgentTool;

public final class SupportTree {
  static final String MODEL = System.getenv("ADK_MODEL");

  static LlmAgent billing() {
    return LlmAgent.builder()
        .name("billing_agent")
        .description("Invoices, refunds, failed payments and plan changes. Not product errors.")
        .model(MODEL)
        .instruction("Resolve billing questions for a {account_tier} customer. "
            + "If the request is not about billing, transfer back to your parent.")
        .disallowTransferToPeers(true)
        .outputKey("billing_summary")
        .build();
  }

  static SequentialAgent techSupport() {
    LlmAgent diagnose = LlmAgent.builder()
        .name("diagnose")
        .model(MODEL)
        .instruction("Identify the most likely cause of the user's error. Output one paragraph.")
        .outputKey("diagnosis")
        .build();
    LlmAgent fixWriter = LlmAgent.builder()
        .name("fix_writer")
        .model(MODEL)
        .instruction("Write numbered fix steps for this diagnosis: {diagnosis}")
        .build();
    return SequentialAgent.builder()
        .name("tech_support")
        .description("Product errors and crashes: diagnoses the cause, then writes fix steps.")
        .subAgents(diagnose, fixWriter)
        .build();
  }

  static LlmAgent policyChecker() {
    return LlmAgent.builder()
        .name("policy_checker")
        .description("Says whether an action is allowed under refund and data policy, with the rule.")
        .model(MODEL)
        .instruction("Answer ALLOWED or DENIED followed by the policy rule that applies.")
        .build();
  }

  public static LlmAgent supervisor() {
    return LlmAgent.builder()
        .name("support_supervisor")
        .description("Front door for customer support.")
        .model(MODEL)
        .instruction("Understand the request. Transfer billing issues to billing_agent and "
            + "product errors to tech_support. Before promising any refund or data action, "
            + "call policy_checker. Answer greetings and simple questions yourself.")
        .tools(AgentTool.create(policyChecker()))
        .subAgents(billing(), techSupport())
        .build();
  }
}

Running it is the same as running any agent: construct a runner over the root and stream events for a message.

InMemoryRunner runner = new InMemoryRunner(SupportTree.supervisor(), "support");
Session session = runner.sessionService()
    .createSession("support", "user-42", Map.of("account_tier", "pro"), null)
    .blockingGet();
Content msg = Content.fromParts(Part.fromText("I was charged twice this month"));
runner.runAsync("user-42", session.id(), msg)
    .blockingForEach(ev -> System.out.println(ev.author() + ": " + ev.stringifyContent()));

For the double-charge message you should see an event from support_supervisor containing a transfer_to_agent call, then events authored by billing_agent. The {account_tier} placeholder is filled from session state when the instruction is built, and the billing agent's final text is saved under billing_summary by outputKey.

State is the contract between levels

Session state is a key-value map that every agent in the invocation can read, and it is the cleanest way to pass structured results down and across the tree. Three mechanisms write to it: outputKey stores an agent's final text under a key; tools write through toolContext.state(); and events carry a state delta that the session service applies. Instructions read it through {key} placeholders. Inside a SequentialAgent this is how a step hands its result to the next, as passing context through sequential agents explains in detail.

Treat keys as a schema. Namespace them by owner (for example billing.summary), document which agent writes each one, and never let two agents write the same key unless you mean the later write to win. This matters most under ParallelAgent: its children run concurrently on separate branches (the branch is set to parent.child) so they do not see each other's conversation, but they share state, and two children writing the same key produce a last-writer-wins race. For how state is stored and scoped across sessions, see session context in depth.

Ending sub-trees: escalation and loop limits

A LoopAgent repeats its children until one of two things happens: the iteration count reaches maxIterations, or an event carries an escalate action. The usual way to escalate is a small tool that a checker agent calls when the work is done.

// Schema is com.google.adk.tools.Annotations.Schema
public final class LoopControl {
  @Schema(description = "Call when the draft meets every requirement.")
  public static Map<String, Object> exitLoop(ToolContext toolContext) {
    toolContext.actions().setEscalate(true);
    return Map.of("status", "done");
  }
}
// FunctionTool.create(LoopControl.class, "exitLoop") on the reviewer agent,
// then LoopAgent.builder().name("refine").subAgents(writer, reviewer).maxIterations(4).build()

Always set maxIterations, even when you expect escalation: a reviewer that never becomes satisfied will otherwise loop until cost or time limits stop it. Choose the cap from evaluation data, not optimism, and record in state which exit fired so you can tell a converged loop from a capped one.

Who answers the next message

Transfer changes who owns the conversation, and that persists across user turns. When a new message arrives, the runner's agent-selection routine first routes any pending function response to the agent that made the call. Otherwise it walks the session's events backwards and resumes with the most recent non-user author whose whole chain to the root is LLM agents allowing transfer to parent; if none qualifies, it starts at the root. Steps inside a SequentialAgent never qualify. In practice a user who was handed to billing_agent keeps talking to billing_agent, which is right for a multi-turn refund and wrong if the user then asks about a crash.

Design for that explicitly. Give every specialist an instruction to transfer back to its parent when the request is outside its description, and leave disallowTransferToParent false for LLM specialists. Specialists that should never own a conversation belong behind an AgentTool instead.

Designing the hierarchy

  • Keep it shallow. Each level adds at least one model call of latency and cost, and each hop is a chance to misroute. Two levels cover most products; add a third only when a specialist has grown its own internal routing problem.
  • Fan-out of three to seven per supervisor. Beyond that, descriptions start to overlap and the supervisor's choice gets noisy; group specialists under intermediate supervisors instead.
  • Deterministic where possible. If step B always follows step A, use a SequentialAgent, not a supervisor that must remember to call B.
  • Scope tools per level. Put tools on the specialist that needs them, not on the supervisor. A supervisor with every tool is the monolith you were trying to remove, and narrow tool sets also narrow the blast radius, as covered in authorisation at the agent boundary.
  • Limit context per child. includeContents(IncludeContents.NONE) stops an agent from receiving the prior conversation, which is useful for pipeline steps that should work only from state.

Failure modes

FailureWhat you seeMitigation
Transfer ping-pongSupervisor and specialist hand the user back and forthSharper, non-overlapping descriptions; one owner per intent; cap hops per turn in a callback
Stuck ownerA specialist answers questions outside its scope in later turnsInstruct transfer-back; move consult-only agents behind AgentTool
Lost contextAgentTool child answers genericallyPut all needed facts in the request or in state before the call
State collisionParallel children overwrite each otherDistinct, namespaced output keys per child
Runaway loopCost spike, repeated near-identical draftsmaxIterations always; escalate tool on the checker
Shared instanceTransfers go to an unexpected parentFactory per placement; unique names
Supervisor does the workSpecialists rarely invokedTell the supervisor explicitly what it must delegate; remove its tools

Operating a tree in production

Observability comes almost free because every event has an author and, under parallel agents, a branch. Log author, branch, invocation id and any transfer target for every event, and you can reconstruct the path each request took; count transfers per invocation to find ping-pong. Agent-level callbacks are the natural place for guardrails that must apply to every node, such as a hop limit or an audit record, and the callbacks guide covers where each hook fires. For latency, remember that a two-level transfer costs the supervisor's model call plus the specialist's, and an AgentTool call adds the child's calls plus the parent's summarisation call.

Unit-test each specialist alone, then run routing evaluations on the whole tree: labelled requests with the expected owning agent, scored by the final response's author, re-run on every description or model change. The execution loop anatomy is the reference for what happens inside each hop when a test fails.

Trade-offs

A hierarchy buys short prompts, small tool sets, independent testing and team ownership. It costs latency, tokens and routing risk per level, and moves bugs to the seams between agents. A single agent suits narrow products; a pure workflow suits known paths. Choose the supervisor pattern when requests genuinely differ in kind, the specialists would conflict if merged, and you are prepared to maintain routing evaluations.

What to do next

  1. Draw your tree on paper: each node's name, type (LLM or workflow), description, tools and the state keys it writes.
  2. Decide per child whether it owns a conversation (sub-agent, transfer) or is consulted (AgentTool).
  3. Write descriptions as contracts with explicit exclusions, and instruct every LLM specialist to transfer back when out of scope.
  4. Build each specialist from a factory and check names are unique with findAgent at startup.
  5. Namespace state keys and give every parallel child its own output key.
  6. Set maxIterations on every LoopAgent and add an escalate tool to its checker.
  7. Log author, branch and transfers per event, and build a labelled routing evaluation set before shipping.
Key takeaway: In ADK Java a hierarchy is a real tree: sub-agent lists set parent pointers, names are addresses, and the model delegates either by transfer_to_agent, which hands over the conversation within scope rules set by the parent type and the disallow flags, or through AgentTool, which runs the child in a temporary session and returns its result. Use state with namespaced keys as the contract between levels, workflow agents wherever order is known, escalation plus maxIterations to end loops, and routing evaluations to keep the seams honest.