An agent that plans is choosing a sequence of tool calls that moves the world from where it is to where the user wants it. There are several published ways for a language model to do that, and they are not interchangeable: they differ by an order of magnitude in model calls, in latency, and in how they cope when a tool returns something surprising. Teams usually pick one in the first week and then fight its weaknesses for a year.
This article treats planning strategy as an architectural choice. It explains planning from first principles, compares four strategy families with a cost model you can compute, runs one task through all four, and shows how to build a single runtime with a router that picks the cheapest strategy likely to succeed and escalates when it does not. How to design a single planner component, with plan validation and replanning triggers, is covered in the planner article and its second version; this one sits a level above, at strategy selection.
Planning is search
Strip away the vocabulary and planning is search. There is a state (what the agent knows and what it has changed), a set of actions (tool calls with arguments), a transition (calling the tool and reading the result) and a goal test (is the task done and correct?). Classical planners search this space with exact models of every action. An LLM agent has no exact model: it uses the language model as a heuristic that proposes plausible next actions, and the real tools as the transition function.
That framing exposes the three decisions every strategy makes:
- When to commit. Decide one step at a time after seeing each result, or write the whole plan first and then run it.
- How many candidates. Follow a single proposal greedily, or generate several and keep the best.
- How results feed back. Every observation goes back to the model, only failures do, or nothing does until a final step.
Late commitment adapts well but pays a model call per step and runs strictly sequentially. Early commitment is cheap and parallelisable but brittle when step three depends on what step two found. More candidates improve quality on problems where early mistakes are fatal and cost proportionally more. Choosing a strategy is choosing a point on these three axes.
Planning is search
Strip away the vocabulary and planning is search. There is a state (what the agent knows and what it has changed), a set of actions (tool calls with arguments), a transition (calling the tool and reading the result) and a goal test (is the task done and correct?). Classical planners search this space with exact models of every action. An LLM agent has no exact model: it uses the language model as a heuristic that proposes plausible next actions, and the real tools as the transition function.
That framing exposes the three decisions every strategy makes:
- When to commit. Decide one step at a time after seeing each result, or write the whole plan first and then run it.
- How many candidates. Follow a single proposal greedily, or generate several and keep the best.
- How results feed back. Every observation goes back to the model, only failures do, or nothing does until a final step.
Late commitment adapts well but pays a model call per step and runs strictly sequentially. Early commitment is cheap and parallelisable but brittle when step three depends on what step two found. More candidates improve quality on problems where early mistakes are fatal and cost proportionally more. Choosing a strategy is choosing a point on these three axes.
Four strategy families and what each costs
The four families below cover most production agents. The call counts assume a task needing N tool calls and are approximate; real numbers depend on prompts and retries.
| Strategy | Commit | LLM calls | Latency shape | Best for |
|---|---|---|---|---|
| ReAct (Yao et al., 2022) | per step | about N + 1 | sequential: every step waits for a model call | short tasks where each result decides the next step |
| Plan-and-execute, Plan-and-Solve (Wang et al., 2023) | whole plan, revisable | 1 plan + N or fewer executor calls + replans | sequential, but executor calls can be small models or plain code | structured tasks with predictable steps |
| DAG planners: ReWOO (Xu et al., 2023), LLMCompiler (Kim et al., 2023) | whole graph | about 2: plan, then join | critical path of the graph; independent tools run in parallel | fan-out lookups, data gathering |
| Tree of Thoughts (Yao et al., 2023) with beam search | per level, b candidates | about b x depth proposals + evaluations | many calls, partly parallel | problems where early mistakes are fatal and partial plans can be scored |
ReAct is the default for good reasons: it is simple and adaptive. See the ReAct article for its loop in detail. Its cost grows linearly with steps, and its greediness means it can commit to a wrong branch and spend the rest of the budget defending it.
ReWOO-style planners write the plan with placeholders such as #E1 for the result of step one, so later steps can reference values the planner has not seen. The tools then run without the model in the loop, and a final solver call reads all results. LLMCompiler pushes this further by treating the plan as a dependency graph and dispatching every ready node at once. Both save calls and time, and both lose adaptivity: if a lookup returns nothing, the plan does not know until the join.
Tree search spends calls to buy quality. It only works when you can score a partial plan, ideally with a check rather than another opinion from the model.
A shared runtime under every strategy
The mistake to avoid is four separate agents. All strategies need the same machinery: a plan state the run can resume from, a tool executor with timeouts and idempotency keys, a verifier, and a budget that every call is charged against. Build those once; a strategy is then a small function that decides what to call next.
import time
from dataclasses import dataclass, field
@dataclass
class Budget:
max_llm_calls: int
max_tool_calls: int
max_seconds: float
llm_calls: int = 0
tool_calls: int = 0
started: float = field(default_factory=time.monotonic)
def charge(self, kind: str) -> None:
if kind == "llm":
self.llm_calls += 1
else:
self.tool_calls += 1
if (self.llm_calls > self.max_llm_calls or self.tool_calls > self.max_tool_calls
or time.monotonic() - self.started > self.max_seconds):
raise BudgetExceeded(self)
@dataclass
class Step:
id: str
tool: str
args: dict # may contain "#E1"-style references to earlier results
depends_on: list[str]
result: object = None
status: str = "pending" # pending | running | done | failed
@dataclass
class PlanState:
task: str
steps: dict[str, Step]
facts: list[str] = field(default_factory=list)
version: int = 0
class Strategy:
name: str
def run(self, state: PlanState, rt: "Runtime") -> "Outcome": ...The executor resolves placeholders, runs every step whose dependencies are done, and records results in the plan state, so a crash resumes from the last finished step instead of the start. Tool-level concerns such as retries and compensation belong here too; tool calling reliability covers them. The verifier is the part teams skip: it runs concrete checks against the outcome (row counts match, the file exists, totals add up) and returns pass, fail or unknown. Every strategy uses it as the goal test.
Worked example: one task, four strategies
Task: compare our Q3 cloud spend across three providers and flag any service whose cost grew more than 20% over Q2. Tools: billing_export(provider, quarter) (about 1 s each), run_sql(query) and write_report(markdown). Assume a model call takes 2 s.
| Strategy | What happens | LLM calls | Wall clock |
|---|---|---|---|
| ReAct | six exports one at a time, a query, a report, each preceded by reasoning | about 9 | about 26 s |
| Plan-and-execute | plan of 8 steps; executor is code for exports, model for SQL and report | about 3 | about 14 s |
| DAG (LLMCompiler style) | six exports in parallel, then SQL, then report; joiner checks totals | about 3 | about 9 s |
| Tree search, beam 3 | explores alternative SQL formulations, scores by checks | 15 or more | 30 s or more |
The arithmetic: ReAct pays 2 s of thinking plus 1 s of tool time per export, six times, then two more rounds. The DAG pays one planning call, one second for all six exports together, a SQL step and a report. Tree search buys nothing here, because the problem has an obvious decomposition and nothing early is irreversible.
Now introduce a surprise: one provider's Q2 export is empty because the billing account was migrated mid-year. ReAct notices immediately and asks for the old account. The DAG planner runs everything, and the verifier at the join sees a provider with zero Q2 cost and every service flagged as infinite growth. The correct response is a local repair: replan only the affected branch with the new fact, rather than restarting. That one design decision recovers most of ReAct's adaptivity at a fraction of its cost.
Routing and the escalation ladder
A router picks the strategy per task from cheap features, and an escalation ladder handles the cases the features miss. Features that work in practice: the number of independent sub-goals the task names, whether steps are known in advance (a template or runbook exists), whether any action is irreversible, and whether a verifier exists for the outcome.
def route(task: TaskFeatures) -> list[str]:
# Return an ordered ladder: try the first, escalate on verified failure.
if task.irreversible_actions:
# Never search over actions with side effects; plan, then confirm with a human.
return ["plan_execute"]
if task.independent_subgoals >= 3:
ladder = ["dag", "plan_execute", "react"]
elif task.has_runbook:
ladder = ["plan_execute", "react"]
else:
ladder = ["react", "plan_execute"]
if task.has_verifier and task.early_mistakes_fatal:
ladder.append("tree_search")
return ladder
def solve(task, rt):
for name in route(features(task)):
outcome = STRATEGIES[name].run(rt.fresh_state(task), rt)
if rt.verifier.check(outcome) == "pass":
return outcome
rt.memory.add_lesson(task, name, outcome.failure_summary) # Reflexion-style
return Outcome.give_up(reason="ladder exhausted")Two rules keep the ladder honest. First, escalate only on a verified failure, not on low model confidence, or every task climbs to the expensive rung. Second, carry lessons up the ladder: a short summary of why the previous attempt failed, injected into the next strategy's prompt, is what made Reflexion (Shinn et al., 2023) effective, and it is cheap. Tasks with irreversible actions skip search entirely and go through a human gate; human in the loop explains where to put it.
Tree search done carefully
When tree search is justified, the implementation decides whether it helps. A minimal beam search over plan prefixes:
def beam_plan(task, rt, beam=3, expand=3, depth=4):
frontier = [Partial(steps=[])]
for _ in range(depth):
candidates = []
for prefix in frontier:
for step in rt.llm.propose_next(task, prefix, n=expand): # 1 call, n samples
cand = prefix.extend(step)
cand.score = rt.score(cand) # prefer checks over LLM self-ratings
candidates.append(cand)
frontier = sorted(candidates, key=lambda c: c.score, reverse=True)[:beam]
if any(rt.verifier.check_plan(c) == "pass" for c in frontier):
break
return frontier[0]Score with something real wherever possible: a dry-run of the SQL against a sample, a type check, a unit test, a constraint solver. Model self-evaluation is useful as a tie-breaker but correlates with the proposer's own mistakes. Search over plans, not over side-effecting actions: expanding a node must never send an email. Cap breadth and depth from the budget, not from intuition: beam 3, expansion 3 and depth 4 is up to 30 candidate steps from about 10 proposal calls, plus scoring, which is why this rung sits at the top of the ladder.
Failure modes
- Strategy monoculture. Everything runs on ReAct, and the slow, expensive fan-out tasks dominate the bill. Route by task shape.
- Plans that cannot see results. DAG planners with no verifier at the join produce confident reports built on empty inputs. Always join through checks.
- Loops. The agent repeats the same tool call with the same arguments. Fingerprint each call (tool plus normalised arguments) and stop or replan on the third repeat.
- Budget only at the top. A sub-agent with its own unlimited budget defeats the parent's. Pass a child budget carved from the parent's ledger.
- Escalating on vibes. Without a verifier, the ladder either never escalates or always does.
- State in the prompt only. If plan state lives only in conversation text, a crash restarts the task and long runs drift. Persist it; agent memory layers covers the options.
What to do next
- Log, for a week of real tasks, model calls, tool calls, wall clock and outcome per task.
- Label each task by shape: independent sub-goals, runbook or not, irreversible actions, verifier or not.
- Build one runtime with plan state, executor, verifier and budget ledger, and port your current strategy onto it.
- Add a DAG strategy for fan-out tasks and measure latency and cost against the logs.
- Write verifiers for your top three task types before adding any search.
- Add the router and escalation ladder, and alert when escalation rate or loop detections rise.