"Planning" means two different things in agent systems. The first is a model reasoning before it acts: thinking tokens, or a ReAct loop where each tool call is chosen after reading the last result. In the Agent Development Kit for Java that lives inside LlmAgent and is covered in the ReAct implementation article. The second is a component outside any one model that decides which agent runs next, checks what came back, and decides when the job is finished. That second kind is the subject here.
ADK Java provides it in the google-adk-planners module: a PlannerAgent that owns sub-agents and delegates every "what next" decision to a pluggable Planner. This article explains the abstractions, walks through the loop exactly as the source implements it, compares the six built-in planners, builds a goal-oriented pipeline with replanning and a custom planner, and lists the behaviours that turn into silent failures. Details were read from the adk-java repository's contrib/planners sources and README on 2026-10-01, alongside release v1.10.1 (2026-09-18). The module sits under contrib, so expect the API to change and pin your version.
When you need a planner at all
Most agent workflows do not. If the steps are fixed, a sequential chain is clearer, and sequential agent chains covers that well. If one model with tools can finish the task, its own reasoning loop is enough. A planner earns its place when the order of work depends on data the system only discovers at run time, when independent steps should run in parallel and dependent ones should wait, or when you want a single place to enforce stopping rules, retries and fallbacks across several agents.
The design idea is separation. Sub-agents do work and write results into session state. The planner reads state and events and returns a decision. Swapping the planner changes the orchestration without touching the agents, which is what makes it possible to start with a deterministic plan and later try an LLM supervisor against the same agents.
The four abstractions
Everything rests on four types in com.google.adk.agents, shipped in the planners module.
// com.google.adk.agents (google-adk-planners)
public interface Planner {
default void init(PlanningContext context) {}
Single<PlannerAction> firstAction(PlanningContext context);
Single<PlannerAction> nextAction(PlanningContext context);
}
public sealed interface PlannerAction
permits PlannerAction.RunAgents, PlannerAction.Done,
PlannerAction.DoneWithResult, PlannerAction.NoOp {
record RunAgents(ImmutableList<BaseAgent> agents) implements PlannerAction {}
record Done() implements PlannerAction {}
record DoneWithResult(String result) implements PlannerAction {}
record NoOp() implements PlannerAction {}
}
// PlanningContext exposes: state(), events(), availableAgents(),
// userContent(), findAgent(String name), invocationContext()Planner has a three-step lifecycle. init runs once before the loop and is where planners build dependency graphs or reset counters. firstAction and nextAction return a Single<PlannerAction> from RxJava 3, so a planner can be synchronous, wrapping a value in Single.just, or asynchronous, calling a model. PlannerAction is sealed, so a switch over it is exhaustive. PlanningContext is the planner's view of the world: the session state map, all events so far, the available sub-agents, the user content that started the invocation, and findAgent, which throws IllegalArgumentException for an unknown name.
The loop, exactly as implemented
Reading PlannerAgent.runAsyncImpl is worth ten minutes, because several behaviours matter in production and none are visible from the builder.
- If there are no sub-agents the agent returns an empty stream immediately.
- It builds a
PlanningContext, callsplanner.init, thenfirstAction. - Every action, including
NoOp, increments an iteration counter. When the counter reachesmaxIterations(default 100, set withmaxIterations(int)on the builder) the loop ends with no event; the only trace is an info-level log line. Doneends the loop and emits nothing.DoneWithResultemits one text event authored by the PlannerAgent and ends.RunAgentswith one agent calls itsrunAsync; with several it merges their event streams withFlowable.merge, so they run concurrently against the same invocation context. When the stream completes,nextActionis called lazily and the loop continues.runLiveis not supported and returns an error.
Two consequences follow. A run that ended because the limit was hit looks identical, in the event stream, to one that finished. And because parallel agents share one invocation context, their events interleave and their state writes are not isolated, so two agents writing the same key race.
The six built-in planners
| Planner | Decides by | Stops when | Use for |
|---|---|---|---|
SequentialPlanner | Registration order, one at a time | Every agent has run | Fixed pipelines |
ParallelPlanner | All agents at once | Immediately after that one action | Independent fan-out |
LoopPlanner(maxCycles) | Cycling through agents | maxCycles reached, or the last event escalates | Draft and review loops |
SupervisorPlanner(llm, ...) | An LLM picks agent names, DONE, or DONE: summary | The LLM says DONE | Open-ended delegation |
GoalOrientedPlanner | Dependency search over input and output keys | Goal key produced, or policy halts | Workflows with data contracts |
P2PPlanner | Agents activate when their inputs appear or change | maxInvocations or an exit predicate | Iterative refinement |
The supervisor prompt is built for you: agent names and descriptions, current state keys (not values), a sliding window of recent events (maxEvents, default 20), its own decision history, and the original request. It parses comma-separated names as a parallel run. Two fallbacks matter: an unknown agent name and a failed model call both become Done, logged as warnings. Agent descriptions are therefore part of the planner's input, and a vague description is a planning bug.
Goal-oriented planning and replanning
GoalOrientedPlanner borrows goal-oriented action planning from game AI. Each agent declares a contract, AgentMetadata(agentName, inputKeys, outputKey), where the keys are session state keys. Given a goal key, the default DfsSearchStrategy chains backwards from the goal to keys already present in state, and agents are grouped into levels: an agent's level is one more than the highest level of the agents it depends on, so agents in the same level run together as one RunAgents. An AStarSearchStrategy is also provided, and the README states both produce identical groupings for valid dependency graphs. Duplicate output keys are rejected.
The contract only works if each agent really writes its declared key. For an LlmAgent that means setting outputKey on its builder to the same string. After each group, nextAction checks that every expected output key is present, and a ReplanPolicy decides what happens if one is missing: Ignore (the default) carries on, FailStop ends with a DoneWithResult naming the missing outputs, and Replan(maxAttempts) rebuilds the plan from current state, giving up after that many consecutive failures. The counter resets whenever a group succeeds.
Worked example: a supplier risk brief
A procurement team wants a one-page risk brief on a supplier. Five agents: fetchFilings and fetchNews read company; analyzeFinancials reads filings; assessRisk reads financials and news; writeBrief reads company and risk. The instruction placeholders such as {company} are filled from session state by ADK's instruction templating.
LlmAgent filings = LlmAgent.builder()
.name("fetchFilings")
.description("Collects the latest annual and quarterly filings")
.model(MODEL)
.instruction("Find the most recent filings for {company}. Return key figures only.")
.tools(filingsTool)
.outputKey("filings")
.build();
// fetchNews -> "news", analyzeFinancials -> "financials",
// assessRisk -> "risk", writeBrief -> "brief" are built the same way.
List<AgentMetadata> contracts = List.of(
new AgentMetadata("fetchFilings", ImmutableList.of("company"), "filings"),
new AgentMetadata("fetchNews", ImmutableList.of("company"), "news"),
new AgentMetadata("analyzeFinancials", ImmutableList.of("filings"), "financials"),
new AgentMetadata("assessRisk", ImmutableList.of("financials", "news"), "risk"),
new AgentMetadata("writeBrief", ImmutableList.of("company", "risk"), "brief"));
PlannerAgent supplierBrief = PlannerAgent.builder()
.name("supplierBrief")
.description("Builds a supplier risk brief")
.subAgents(filings, news, financials, risk, brief)
.planner(new GoalOrientedPlanner("brief", contracts,
new DfsSearchStrategy(), new ReplanPolicy.Replan(2)))
.maxIterations(12)
.build();With company already in state, the search yields four groups: [fetchFilings, fetchNews] in parallel, then [analyzeFinancials], then [assessRisk], then [writeBrief]. Suppose the filings tool times out and fetchFilings finishes without writing filings. Before group two, the planner sees the missing key, rebuilds the plan from current state, and runs fetchFilings again; if that fails twice in a row it stops with a message naming the missing output, which the caller can show or route to a human. The iteration cap of 12 is a backstop: four groups plus two replans need far fewer.
One trap is specific to this planner. init treats keys already in state as satisfied. If an earlier conversation turn left a stale filings value, fetchFilings is skipped and the brief is built on old data. Remove each declared output key from state before invoking the PlannerAgent for a new run.
Writing your own planner
Custom planners are small. The one below runs a cheap triage agent, escalates to an expensive one only when the cheap agent is not confident, and refuses to finish silently if no answer was produced.
/** Deterministic triage: cheap path first, expensive path only if the cheap one is unsure. */
public final class TriagePlanner implements Planner {
private final String fast, deep, writer;
private int step; // planner state: build one planner per invocation
public TriagePlanner(String fast, String deep, String writer) {
this.fast = fast; this.deep = deep; this.writer = writer;
}
@Override public void init(PlanningContext ctx) { step = 0; }
@Override public Single<PlannerAction> firstAction(PlanningContext ctx) {
step = 1;
return Single.just(new PlannerAction.RunAgents(ctx.findAgent(fast)));
}
@Override public Single<PlannerAction> nextAction(PlanningContext ctx) {
Object verdict = ctx.state().get("triage");
if (step == 1 && !"confident".equals(verdict)) {
step = 2;
return Single.just(new PlannerAction.RunAgents(ctx.findAgent(deep)));
}
if (step < 3) {
step = 3;
return Single.just(new PlannerAction.RunAgents(ctx.findAgent(writer)));
}
if (!ctx.state().containsKey("answer")) {
return Single.just(new PlannerAction.DoneWithResult("No answer produced; escalate to a human."));
}
return Single.just(new PlannerAction.Done());
}
}Keep decisions in deterministic code where you can, read structured values from state rather than parsing prose from events, and always end with an explicit outcome. The built-in planners note in their source that they hold mutable state and are not thread-safe, and the same is true of this one.
Failure modes to guard
- Silent stops. Reaching maxIterations, a supervisor naming an unknown agent and a supervisor whose model call failed all end the run with no error event. After the PlannerAgent completes, check that the goal key is in state and treat its absence as failure.
- Shared planner instances. Planners keep cursors and counters, and SupervisorPlanner keeps a decision history that its init does not clear. A PlannerAgent registered as a singleton bean leaks history across invocations and races under concurrency. Build the planner, and the PlannerAgent, per invocation.
- Presence is not quality. GOAP checks only that a key exists. An agent that writes an apology into
risksatisfies it. Add validator agents or callbacks that remove bad values so replanning can trigger. - Parallel write races. Agents in one RunAgents group share the invocation context; give each its own output key.
- Order-sensitive loop exits. LoopPlanner checks only the last event for escalate, so an escalation followed by any other event is missed.
- Unbounded supervisor cost. Every decision is a model call; with the default cap of 100 iterations, an indecisive supervisor can make many calls. Set maxIterations to what the task needs.
Operations and trade-offs
Log every planner decision with the state keys it saw; the runtime stages beneath each step are described in the anatomy of the execution loop. Track per run: iterations used, replans, which stop path ended the run, and wall time per group. Alert when runs end without the goal key.
Deterministic planners are cheap, testable and predictable but only as good as the contracts you write. The supervisor adapts to surprises but costs a model call per step and fails closed to Done. GOAP gives parallelism and replanning for free once contracts exist, and P2P suits refinement loops where outputs feed back. A sensible path is deterministic first, then the supervisor only where the order genuinely cannot be known in advance. For simple repeat-until-good loops, loop agents may be enough without the planners module.
What to do next
- Add google-adk-planners at the same version as your ADK core and pin it.
- Write AgentMetadata contracts for an existing pipeline and make each LlmAgent's outputKey match.
- Run it with GoalOrientedPlanner and ReplanPolicy.Replan(2), then force a missing output to watch replanning.
- Set maxIterations explicitly and add a post-run check that the goal key exists.
- Build planners per invocation and clear declared output keys from state before each run.
- Only then try SupervisorPlanner against the same agents, with precise agent descriptions, and compare cost and success rate.