An agent that plans is choosing a sequence of tool calls that moves the world from where it is to where the user wants it. There are several published ways for a language model to do that, and they are not interchangeable: they differ by an order of magnitude in model calls, in latency, and in how they cope when a tool returns something surprising. Teams usually pick one in the first week and then fight its weaknesses for a year.

This article treats planning strategy as an architectural choice. It explains planning from first principles, compares four strategy families with a cost model you can compute, runs one task through all four, and shows how to build a single runtime with a router that picks the cheapest strategy likely to succeed and escalates when it does not. How to design a single planner component, with plan validation and replanning triggers, is covered in the planner article and its second version; this one sits a level above, at strategy selection.

One runtime, several planning strategiesTask + budgetgoal, constraints, success testStrategy routerfeatures + escalation ladderReActthink, act, observePlan-and-executelist plan, replan on failureDAG plannerparallel tools, joinerTree searchbeam over candidate stepsShared runtimePlan statesteps, facts, versionsTool executortimeouts, idempotencyVerifierchecks, not opinionsBudget ledgercalls, tokens, secondsStrategies differ in when they commit to a plan and how many candidates they consider;everything below the line is the same code.
A router chooses a strategy per task; all strategies share plan state, the tool executor, the verifier and the budget ledger.

Planning is search

Strip away the vocabulary and planning is search. There is a state (what the agent knows and what it has changed), a set of actions (tool calls with arguments), a transition (calling the tool and reading the result) and a goal test (is the task done and correct?). Classical planners search this space with exact models of every action. An LLM agent has no exact model: it uses the language model as a heuristic that proposes plausible next actions, and the real tools as the transition function.

That framing exposes the three decisions every strategy makes:

  1. When to commit. Decide one step at a time after seeing each result, or write the whole plan first and then run it.
  2. How many candidates. Follow a single proposal greedily, or generate several and keep the best.
  3. How results feed back. Every observation goes back to the model, only failures do, or nothing does until a final step.

Late commitment adapts well but pays a model call per step and runs strictly sequentially. Early commitment is cheap and parallelisable but brittle when step three depends on what step two found. More candidates improve quality on problems where early mistakes are fatal and cost proportionally more. Choosing a strategy is choosing a point on these three axes.

Planning is search

Strip away the vocabulary and planning is search. There is a state (what the agent knows and what it has changed), a set of actions (tool calls with arguments), a transition (calling the tool and reading the result) and a goal test (is the task done and correct?). Classical planners search this space with exact models of every action. An LLM agent has no exact model: it uses the language model as a heuristic that proposes plausible next actions, and the real tools as the transition function.

That framing exposes the three decisions every strategy makes:

  1. When to commit. Decide one step at a time after seeing each result, or write the whole plan first and then run it.
  2. How many candidates. Follow a single proposal greedily, or generate several and keep the best.
  3. How results feed back. Every observation goes back to the model, only failures do, or nothing does until a final step.

Late commitment adapts well but pays a model call per step and runs strictly sequentially. Early commitment is cheap and parallelisable but brittle when step three depends on what step two found. More candidates improve quality on problems where early mistakes are fatal and cost proportionally more. Choosing a strategy is choosing a point on these three axes.

Four strategy families and what each costs

The four families below cover most production agents. The call counts assume a task needing N tool calls and are approximate; real numbers depend on prompts and retries.

StrategyCommitLLM callsLatency shapeBest for
ReAct (Yao et al., 2022)per stepabout N + 1sequential: every step waits for a model callshort tasks where each result decides the next step
Plan-and-execute, Plan-and-Solve (Wang et al., 2023)whole plan, revisable1 plan + N or fewer executor calls + replanssequential, but executor calls can be small models or plain codestructured tasks with predictable steps
DAG planners: ReWOO (Xu et al., 2023), LLMCompiler (Kim et al., 2023)whole graphabout 2: plan, then joincritical path of the graph; independent tools run in parallelfan-out lookups, data gathering
Tree of Thoughts (Yao et al., 2023) with beam searchper level, b candidatesabout b x depth proposals + evaluationsmany calls, partly parallelproblems where early mistakes are fatal and partial plans can be scored

ReAct is the default for good reasons: it is simple and adaptive. See the ReAct article for its loop in detail. Its cost grows linearly with steps, and its greediness means it can commit to a wrong branch and spend the rest of the budget defending it.

ReWOO-style planners write the plan with placeholders such as #E1 for the result of step one, so later steps can reference values the planner has not seen. The tools then run without the model in the loop, and a final solver call reads all results. LLMCompiler pushes this further by treating the plan as a dependency graph and dispatching every ready node at once. Both save calls and time, and both lose adaptivity: if a lookup returns nothing, the plan does not know until the join.

Tree search spends calls to buy quality. It only works when you can score a partial plan, ideally with a check rather than another opinion from the model.

A shared runtime under every strategy

The mistake to avoid is four separate agents. All strategies need the same machinery: a plan state the run can resume from, a tool executor with timeouts and idempotency keys, a verifier, and a budget that every call is charged against. Build those once; a strategy is then a small function that decides what to call next.

import time
from dataclasses import dataclass, field

@dataclass
class Budget:
    max_llm_calls: int
    max_tool_calls: int
    max_seconds: float
    llm_calls: int = 0
    tool_calls: int = 0
    started: float = field(default_factory=time.monotonic)

    def charge(self, kind: str) -> None:
        if kind == "llm":
            self.llm_calls += 1
        else:
            self.tool_calls += 1
        if (self.llm_calls > self.max_llm_calls or self.tool_calls > self.max_tool_calls
                or time.monotonic() - self.started > self.max_seconds):
            raise BudgetExceeded(self)

@dataclass
class Step:
    id: str
    tool: str
    args: dict            # may contain "#E1"-style references to earlier results
    depends_on: list[str]
    result: object = None
    status: str = "pending"   # pending | running | done | failed

@dataclass
class PlanState:
    task: str
    steps: dict[str, Step]
    facts: list[str] = field(default_factory=list)
    version: int = 0

class Strategy:
    name: str
    def run(self, state: PlanState, rt: "Runtime") -> "Outcome": ...

The executor resolves placeholders, runs every step whose dependencies are done, and records results in the plan state, so a crash resumes from the last finished step instead of the start. Tool-level concerns such as retries and compensation belong here too; tool calling reliability covers them. The verifier is the part teams skip: it runs concrete checks against the outcome (row counts match, the file exists, totals add up) and returns pass, fail or unknown. Every strategy uses it as the goal test.

Worked example: one task, four strategies

Task: compare our Q3 cloud spend across three providers and flag any service whose cost grew more than 20% over Q2. Tools: billing_export(provider, quarter) (about 1 s each), run_sql(query) and write_report(markdown). Assume a model call takes 2 s.

StrategyWhat happensLLM callsWall clock
ReActsix exports one at a time, a query, a report, each preceded by reasoningabout 9about 26 s
Plan-and-executeplan of 8 steps; executor is code for exports, model for SQL and reportabout 3about 14 s
DAG (LLMCompiler style)six exports in parallel, then SQL, then report; joiner checks totalsabout 3about 9 s
Tree search, beam 3explores alternative SQL formulations, scores by checks15 or more30 s or more

The arithmetic: ReAct pays 2 s of thinking plus 1 s of tool time per export, six times, then two more rounds. The DAG pays one planning call, one second for all six exports together, a SQL step and a report. Tree search buys nothing here, because the problem has an obvious decomposition and nothing early is irreversible.

Now introduce a surprise: one provider's Q2 export is empty because the billing account was migrated mid-year. ReAct notices immediately and asks for the old account. The DAG planner runs everything, and the verifier at the join sees a provider with zero Q2 cost and every service flagged as infinite growth. The correct response is a local repair: replan only the affected branch with the new fact, rather than restarting. That one design decision recovers most of ReAct's adaptivity at a fraction of its cost.

Routing and the escalation ladder

A router picks the strategy per task from cheap features, and an escalation ladder handles the cases the features miss. Features that work in practice: the number of independent sub-goals the task names, whether steps are known in advance (a template or runbook exists), whether any action is irreversible, and whether a verifier exists for the outcome.

def route(task: TaskFeatures) -> list[str]:
    # Return an ordered ladder: try the first, escalate on verified failure.
    if task.irreversible_actions:
        # Never search over actions with side effects; plan, then confirm with a human.
        return ["plan_execute"]
    if task.independent_subgoals >= 3:
        ladder = ["dag", "plan_execute", "react"]
    elif task.has_runbook:
        ladder = ["plan_execute", "react"]
    else:
        ladder = ["react", "plan_execute"]
    if task.has_verifier and task.early_mistakes_fatal:
        ladder.append("tree_search")
    return ladder

def solve(task, rt):
    for name in route(features(task)):
        outcome = STRATEGIES[name].run(rt.fresh_state(task), rt)
        if rt.verifier.check(outcome) == "pass":
            return outcome
        rt.memory.add_lesson(task, name, outcome.failure_summary)   # Reflexion-style
    return Outcome.give_up(reason="ladder exhausted")

Two rules keep the ladder honest. First, escalate only on a verified failure, not on low model confidence, or every task climbs to the expensive rung. Second, carry lessons up the ladder: a short summary of why the previous attempt failed, injected into the next strategy's prompt, is what made Reflexion (Shinn et al., 2023) effective, and it is cheap. Tasks with irreversible actions skip search entirely and go through a human gate; human in the loop explains where to put it.

Tree search done carefully

When tree search is justified, the implementation decides whether it helps. A minimal beam search over plan prefixes:

def beam_plan(task, rt, beam=3, expand=3, depth=4):
    frontier = [Partial(steps=[])]
    for _ in range(depth):
        candidates = []
        for prefix in frontier:
            for step in rt.llm.propose_next(task, prefix, n=expand):   # 1 call, n samples
                cand = prefix.extend(step)
                cand.score = rt.score(cand)        # prefer checks over LLM self-ratings
                candidates.append(cand)
        frontier = sorted(candidates, key=lambda c: c.score, reverse=True)[:beam]
        if any(rt.verifier.check_plan(c) == "pass" for c in frontier):
            break
    return frontier[0]

Score with something real wherever possible: a dry-run of the SQL against a sample, a type check, a unit test, a constraint solver. Model self-evaluation is useful as a tie-breaker but correlates with the proposer's own mistakes. Search over plans, not over side-effecting actions: expanding a node must never send an email. Cap breadth and depth from the budget, not from intuition: beam 3, expansion 3 and depth 4 is up to 30 candidate steps from about 10 proposal calls, plus scoring, which is why this rung sits at the top of the ladder.

Failure modes

  • Strategy monoculture. Everything runs on ReAct, and the slow, expensive fan-out tasks dominate the bill. Route by task shape.
  • Plans that cannot see results. DAG planners with no verifier at the join produce confident reports built on empty inputs. Always join through checks.
  • Loops. The agent repeats the same tool call with the same arguments. Fingerprint each call (tool plus normalised arguments) and stop or replan on the third repeat.
  • Budget only at the top. A sub-agent with its own unlimited budget defeats the parent's. Pass a child budget carved from the parent's ledger.
  • Escalating on vibes. Without a verifier, the ladder either never escalates or always does.
  • State in the prompt only. If plan state lives only in conversation text, a crash restarts the task and long runs drift. Persist it; agent memory layers covers the options.

What to do next

  1. Log, for a week of real tasks, model calls, tool calls, wall clock and outcome per task.
  2. Label each task by shape: independent sub-goals, runbook or not, irreversible actions, verifier or not.
  3. Build one runtime with plan state, executor, verifier and budget ledger, and port your current strategy onto it.
  4. Add a DAG strategy for fan-out tasks and measure latency and cost against the logs.
  5. Write verifiers for your top three task types before adding any search.
  6. Add the router and escalation ladder, and alert when escalation rate or loop detections rise.
Key takeaway: Planning strategies differ in when they commit, how many candidates they consider and how results feed back, and those choices set cost and latency more than any prompt does. Build one runtime with persistent plan state, a tool executor, a verifier and a budget ledger; run fan-out tasks as DAGs, predictable ones as plan-and-execute, exploratory ones with ReAct, and reserve tree search for problems you can verify; then let a router escalate only on verified failure.