"Planning" means two different things in agent systems. The first is a model reasoning before it acts: thinking tokens, or a ReAct loop where each tool call is chosen after reading the last result. In the Agent Development Kit for Java that lives inside LlmAgent and is covered in the ReAct implementation article. The second is a component outside any one model that decides which agent runs next, checks what came back, and decides when the job is finished. That second kind is the subject here.

ADK Java provides it in the google-adk-planners module: a PlannerAgent that owns sub-agents and delegates every "what next" decision to a pluggable Planner. This article explains the abstractions, walks through the loop exactly as the source implements it, compares the six built-in planners, builds a goal-oriented pipeline with replanning and a custom planner, and lists the behaviours that turn into silent failures. Details were read from the adk-java repository's contrib/planners sources and README on 2026-10-01, alongside release v1.10.1 (2026-09-18). The module sits under contrib, so expect the API to change and pin your version.

Advertisement

When you need a planner at all

Most agent workflows do not. If the steps are fixed, a sequential chain is clearer, and sequential agent chains covers that well. If one model with tools can finish the task, its own reasoning loop is enough. A planner earns its place when the order of work depends on data the system only discovers at run time, when independent steps should run in parallel and dependent ones should wait, or when you want a single place to enforce stopping rules, retries and fallbacks across several agents.

The design idea is separation. Sub-agents do work and write results into session state. The planner reads state and events and returns a decision. Swapping the planner changes the orchestration without touching the agents, which is what makes it possible to start with a deterministic plan and later try an LLM supervisor against the same agents.

The four abstractions

Everything rests on four types in com.google.adk.agents, shipped in the planners module.

// com.google.adk.agents (google-adk-planners)
public interface Planner {
  default void init(PlanningContext context) {}
  Single<PlannerAction> firstAction(PlanningContext context);
  Single<PlannerAction> nextAction(PlanningContext context);
}

public sealed interface PlannerAction
    permits PlannerAction.RunAgents, PlannerAction.Done,
            PlannerAction.DoneWithResult, PlannerAction.NoOp {
  record RunAgents(ImmutableList<BaseAgent> agents) implements PlannerAction {}
  record Done() implements PlannerAction {}
  record DoneWithResult(String result) implements PlannerAction {}
  record NoOp() implements PlannerAction {}
}

// PlanningContext exposes: state(), events(), availableAgents(),
// userContent(), findAgent(String name), invocationContext()

Planner has a three-step lifecycle. init runs once before the loop and is where planners build dependency graphs or reset counters. firstAction and nextAction return a Single<PlannerAction> from RxJava 3, so a planner can be synchronous, wrapping a value in Single.just, or asynchronous, calling a model. PlannerAction is sealed, so a switch over it is exhaustive. PlanningContext is the planner's view of the world: the session state map, all events so far, the available sub-agents, the user content that started the invocation, and findAgent, which throws IllegalArgumentException for an unknown name.

PlannerAgent loop: the planner chooses, sub-agents act, state is the shared worldPlannerAgentmaxIterations, default 100Plannerinit, firstAction, nextActionPlanningContextPlannerActionsealed: 4 variantsSingleDoneemit nothing, stopDoneWithResultemit one text eventNoOpask again, uses an iterationRunAgents1 agent, or N mergedSub-agentsLlmAgent with outputKeyrunAsyncSession stateoutputs written by key, read by later agentswritenextAction reads
The planner never runs work itself. It returns an action; PlannerAgent executes it, sub-agents write to session state, and the next decision reads that state.
Advertisement

The loop, exactly as implemented

Reading PlannerAgent.runAsyncImpl is worth ten minutes, because several behaviours matter in production and none are visible from the builder.

  1. If there are no sub-agents the agent returns an empty stream immediately.
  2. It builds a PlanningContext, calls planner.init, then firstAction.
  3. Every action, including NoOp, increments an iteration counter. When the counter reaches maxIterations (default 100, set with maxIterations(int) on the builder) the loop ends with no event; the only trace is an info-level log line.
  4. Done ends the loop and emits nothing. DoneWithResult emits one text event authored by the PlannerAgent and ends.
  5. RunAgents with one agent calls its runAsync; with several it merges their event streams with Flowable.merge, so they run concurrently against the same invocation context. When the stream completes, nextAction is called lazily and the loop continues.
  6. runLive is not supported and returns an error.

Two consequences follow. A run that ended because the limit was hit looks identical, in the event stream, to one that finished. And because parallel agents share one invocation context, their events interleave and their state writes are not isolated, so two agents writing the same key race.

The six built-in planners

PlannerDecides byStops whenUse for
SequentialPlannerRegistration order, one at a timeEvery agent has runFixed pipelines
ParallelPlannerAll agents at onceImmediately after that one actionIndependent fan-out
LoopPlanner(maxCycles)Cycling through agentsmaxCycles reached, or the last event escalatesDraft and review loops
SupervisorPlanner(llm, ...)An LLM picks agent names, DONE, or DONE: summaryThe LLM says DONEOpen-ended delegation
GoalOrientedPlannerDependency search over input and output keysGoal key produced, or policy haltsWorkflows with data contracts
P2PPlannerAgents activate when their inputs appear or changemaxInvocations or an exit predicateIterative refinement

The supervisor prompt is built for you: agent names and descriptions, current state keys (not values), a sliding window of recent events (maxEvents, default 20), its own decision history, and the original request. It parses comma-separated names as a parallel run. Two fallbacks matter: an unknown agent name and a failed model call both become Done, logged as warnings. Agent descriptions are therefore part of the planner's input, and a vague description is a planning bug.

Goal-oriented planning and replanning

GoalOrientedPlanner borrows goal-oriented action planning from game AI. Each agent declares a contract, AgentMetadata(agentName, inputKeys, outputKey), where the keys are session state keys. Given a goal key, the default DfsSearchStrategy chains backwards from the goal to keys already present in state, and agents are grouped into levels: an agent's level is one more than the highest level of the agents it depends on, so agents in the same level run together as one RunAgents. An AStarSearchStrategy is also provided, and the README states both produce identical groupings for valid dependency graphs. Duplicate output keys are rejected.

The contract only works if each agent really writes its declared key. For an LlmAgent that means setting outputKey on its builder to the same string. After each group, nextAction checks that every expected output key is present, and a ReplanPolicy decides what happens if one is missing: Ignore (the default) carries on, FailStop ends with a DoneWithResult naming the missing outputs, and Replan(maxAttempts) rebuilds the plan from current state, giving up after that many consecutive failures. The counter resets whenever a group succeeds.

Worked example: a supplier risk brief

A procurement team wants a one-page risk brief on a supplier. Five agents: fetchFilings and fetchNews read company; analyzeFinancials reads filings; assessRisk reads financials and news; writeBrief reads company and risk. The instruction placeholders such as {company} are filled from session state by ADK's instruction templating.

LlmAgent filings = LlmAgent.builder()
    .name("fetchFilings")
    .description("Collects the latest annual and quarterly filings")
    .model(MODEL)
    .instruction("Find the most recent filings for {company}. Return key figures only.")
    .tools(filingsTool)
    .outputKey("filings")
    .build();
// fetchNews -> "news", analyzeFinancials -> "financials",
// assessRisk -> "risk", writeBrief -> "brief" are built the same way.

List<AgentMetadata> contracts = List.of(
    new AgentMetadata("fetchFilings",      ImmutableList.of("company"), "filings"),
    new AgentMetadata("fetchNews",         ImmutableList.of("company"), "news"),
    new AgentMetadata("analyzeFinancials", ImmutableList.of("filings"), "financials"),
    new AgentMetadata("assessRisk",        ImmutableList.of("financials", "news"), "risk"),
    new AgentMetadata("writeBrief",        ImmutableList.of("company", "risk"), "brief"));

PlannerAgent supplierBrief = PlannerAgent.builder()
    .name("supplierBrief")
    .description("Builds a supplier risk brief")
    .subAgents(filings, news, financials, risk, brief)
    .planner(new GoalOrientedPlanner("brief", contracts,
                                     new DfsSearchStrategy(), new ReplanPolicy.Replan(2)))
    .maxIterations(12)
    .build();

With company already in state, the search yields four groups: [fetchFilings, fetchNews] in parallel, then [analyzeFinancials], then [assessRisk], then [writeBrief]. Suppose the filings tool times out and fetchFilings finishes without writing filings. Before group two, the planner sees the missing key, rebuilds the plan from current state, and runs fetchFilings again; if that fails twice in a row it stops with a message naming the missing output, which the caller can show or route to a human. The iteration cap of 12 is a backstop: four groups plus two replans need far fewer.

One trap is specific to this planner. init treats keys already in state as satisfied. If an earlier conversation turn left a stale filings value, fetchFilings is skipped and the brief is built on old data. Remove each declared output key from state before invoking the PlannerAgent for a new run.

Writing your own planner

Custom planners are small. The one below runs a cheap triage agent, escalates to an expensive one only when the cheap agent is not confident, and refuses to finish silently if no answer was produced.

/** Deterministic triage: cheap path first, expensive path only if the cheap one is unsure. */
public final class TriagePlanner implements Planner {
  private final String fast, deep, writer;
  private int step;                       // planner state: build one planner per invocation

  public TriagePlanner(String fast, String deep, String writer) {
    this.fast = fast; this.deep = deep; this.writer = writer;
  }

  @Override public void init(PlanningContext ctx) { step = 0; }

  @Override public Single<PlannerAction> firstAction(PlanningContext ctx) {
    step = 1;
    return Single.just(new PlannerAction.RunAgents(ctx.findAgent(fast)));
  }

  @Override public Single<PlannerAction> nextAction(PlanningContext ctx) {
    Object verdict = ctx.state().get("triage");
    if (step == 1 && !"confident".equals(verdict)) {
      step = 2;
      return Single.just(new PlannerAction.RunAgents(ctx.findAgent(deep)));
    }
    if (step < 3) {
      step = 3;
      return Single.just(new PlannerAction.RunAgents(ctx.findAgent(writer)));
    }
    if (!ctx.state().containsKey("answer")) {
      return Single.just(new PlannerAction.DoneWithResult("No answer produced; escalate to a human."));
    }
    return Single.just(new PlannerAction.Done());
  }
}

Keep decisions in deterministic code where you can, read structured values from state rather than parsing prose from events, and always end with an explicit outcome. The built-in planners note in their source that they hold mutable state and are not thread-safe, and the same is true of this one.

Failure modes to guard

  • Silent stops. Reaching maxIterations, a supervisor naming an unknown agent and a supervisor whose model call failed all end the run with no error event. After the PlannerAgent completes, check that the goal key is in state and treat its absence as failure.
  • Shared planner instances. Planners keep cursors and counters, and SupervisorPlanner keeps a decision history that its init does not clear. A PlannerAgent registered as a singleton bean leaks history across invocations and races under concurrency. Build the planner, and the PlannerAgent, per invocation.
  • Presence is not quality. GOAP checks only that a key exists. An agent that writes an apology into risk satisfies it. Add validator agents or callbacks that remove bad values so replanning can trigger.
  • Parallel write races. Agents in one RunAgents group share the invocation context; give each its own output key.
  • Order-sensitive loop exits. LoopPlanner checks only the last event for escalate, so an escalation followed by any other event is missed.
  • Unbounded supervisor cost. Every decision is a model call; with the default cap of 100 iterations, an indecisive supervisor can make many calls. Set maxIterations to what the task needs.

Operations and trade-offs

Log every planner decision with the state keys it saw; the runtime stages beneath each step are described in the anatomy of the execution loop. Track per run: iterations used, replans, which stop path ended the run, and wall time per group. Alert when runs end without the goal key.

Deterministic planners are cheap, testable and predictable but only as good as the contracts you write. The supervisor adapts to surprises but costs a model call per step and fails closed to Done. GOAP gives parallelism and replanning for free once contracts exist, and P2P suits refinement loops where outputs feed back. A sensible path is deterministic first, then the supervisor only where the order genuinely cannot be known in advance. For simple repeat-until-good loops, loop agents may be enough without the planners module.

What to do next

  1. Add google-adk-planners at the same version as your ADK core and pin it.
  2. Write AgentMetadata contracts for an existing pipeline and make each LlmAgent's outputKey match.
  3. Run it with GoalOrientedPlanner and ReplanPolicy.Replan(2), then force a missing output to watch replanning.
  4. Set maxIterations explicitly and add a post-run check that the goal key exists.
  5. Build planners per invocation and clear declared output keys from state before each run.
  6. Only then try SupervisorPlanner against the same agents, with precise agent descriptions, and compare cost and success rate.
Key takeaway: ADK Java separates doing from deciding: sub-agents write to session state, and a Planner returns RunAgents, Done, DoneWithResult or NoOp. Prefer deterministic planners, use GoalOrientedPlanner with explicit output-key contracts and a Replan policy for data-dependent workflows, and guard the silent stops: cap iterations, verify the goal key after every run, and never share a planner instance across invocations.