Most agent planners start the same way: ask the model for a numbered list of steps, execute them in order, and when something fails, ask for a new list. That v1 design works for demos and for tasks of five or six steps. It breaks in recognisable ways as tasks grow: plans that are confidently wrong from step one, replans that throw away work already done, budgets blown on easy tasks that never needed a plan, and no way to tell whether a change to the planner made things better.
This article describes a v2 architecture that addresses those failures one by one. It assumes you know the basics of a planner, such as the goal contract, plan validation and replanning triggers, which are covered in agent planner architecture. Here the focus is on what to change, why, and how to prove the change helped.
What breaks in v1, and the v2 answer
| Symptom in v1 | Root cause | v2 mechanism |
|---|---|---|
| Simple questions take ten model calls | Every request is planned | Plan-or-act gate |
| Plans are detailed but wrong about later steps | Steps are fixed before earlier results are known | Two tiers: milestones now, steps just in time |
| One tool failure restarts the whole task | Replan is global | Local repair of the failing subtree |
| Runaway cost on hard tasks | No budget, or one global cap | Budget ledger with per-milestone allocations |
| Progress lost on crash or timeout | Plan lives only in the prompt | Versioned plan store, resumable |
| Unclear whether a planner change helped | No fixed evaluation | Task suite, shadow runs, gated rollout |
None of these mechanisms is exotic. The value of v2 is in combining them so each one covers a weakness of the others, and in making plan state explicit data rather than prose in a context window.
The architecture
Data flows in one direction during normal operation: goal to gate, gate to strategic planner, milestone to tactical planner, steps to executor, results to monitor. The loop closes only through the monitor, which is the single place that decides a milestone has passed, failed or run out of budget. Keeping that decision in one component makes the agent's behaviour auditable: every transition is a logged event with a reason.
Plans as data: milestones with acceptance checks
In v2 the plan is a typed object the runtime can inspect, not a paragraph. The strategic tier produces milestones; each milestone carries the condition that proves it is done and a share of the budget. Steps are attached only when the milestone becomes active.
{
"plan_id": "p-7f3a", "version": 3, "goal": "Rename config key db.url to db.dsn in services A, B, C",
"milestones": [
{"id": "m1", "title": "Find every usage of db.url",
"depends_on": [], "budget": {"tokens": 20000, "tool_calls": 10},
"accept": {"type": "artifact", "name": "usage_list", "non_empty": true},
"status": "done"},
{"id": "m2", "title": "Patch service A",
"depends_on": ["m1"], "budget": {"tokens": 30000, "tool_calls": 15},
"accept": {"type": "command", "run": "make test", "exit_code": 0},
"status": "active",
"steps": [{"id": "m2.s1", "tool": "edit_file", "args": {"path": "a/config.py"}},
{"id": "m2.s2", "tool": "run", "args": {"cmd": "make test"}}]},
{"id": "m3", "title": "Patch service B", "depends_on": ["m1"], "status": "pending"},
{"id": "m4", "title": "Patch service C", "depends_on": ["m1"], "status": "pending"}
]
}Acceptance checks are the heart of the design. A check that a program can evaluate, such as a test command, a schema validation or the presence of a non-empty artifact, is far more reliable than asking the model whether it succeeded. Where no programmatic check exists, use a separate verifier prompt with the milestone's stated outcome, and treat its verdict as weaker evidence. The pattern for building checks is covered in agent output verification.
Because milestones have explicit dependencies, independent ones (m3 and m4 above) can run in parallel, and the plan doubles as a progress report for a human watching the task.
The plan-or-act gate
Planning has a cost: at least one extra model call, more tokens in context and more places to go wrong. For a question answerable with one tool call, a plan is pure overhead. The gate decides, before any planning, which path a request takes.
def route(goal, ctx):
features = {
"est_tool_calls": estimate_tool_calls(goal), # small model or heuristic
"touches_writes": mentions_side_effects(goal), # deploy, delete, send, pay
"multi_target": count_targets(goal) > 1, # several files, services, people
"history_success_direct": ctx.stats.direct_success_rate(goal_kind(goal)),
}
if features["touches_writes"] or features["multi_target"]:
return "plan"
if features["est_tool_calls"] <= 3 and features["history_success_direct"] >= 0.9:
return "direct"
return "plan"Start with rules like these and log every decision with the eventual outcome. After a few hundred tasks you can see which direct-routed tasks failed and should have been planned, and which planned tasks finished in one step and wasted the plan. Tighten the rules from that data, or train a small classifier on it. Bias the gate towards planning whenever a task has side effects: a wasted plan costs tokens, a skipped plan on a destructive task can cost much more.
Two tiers: strategic and tactical
A single-tier planner writes every step at the start, when it knows least. The v2 planner commits early only to what it can justify: the milestones and their checks. Steps for a milestone are generated just before it runs, with the results of earlier milestones in context. In the example, the steps for patching service B are written after m1 has produced the real list of files, so the planner does not invent file paths.
The strategic prompt should be short and stable: the goal, the available tool families (not every tool schema), constraints, and the required output format. The tactical prompt is concrete: one milestone, its acceptance check, its remaining budget, the relevant artifacts from earlier milestones and the full schemas of the tools it may use. Splitting the context this way keeps both prompts small, which improves reliability and cost at the same time. How to cut a goal into milestones well is its own skill, covered in agent task decomposition.
The budget ledger
A global cap tells you when to stop, but not where the money went. A ledger allocates the task's budget across milestones when the plan is made, records every spend against the active milestone, and lets the monitor act before the total is gone.
class Ledger:
def __init__(self, total, reserve_frac=0.2):
self.total, self.reserve = total, total * reserve_frac # held back for repair
self.alloc, self.spent = {}, {}
def allocate(self, milestones):
pool = self.total - self.reserve
weights = {m.id: m.estimated_effort for m in milestones}
s = sum(weights.values())
self.alloc = {mid: pool * w / s for mid, w in weights.items()}
def charge(self, mid, cost):
self.spent[mid] = self.spent.get(mid, 0) + cost
return self.spent[mid] / self.alloc[mid] # fraction used
def grant_from_reserve(self, mid, amount):
amount = min(amount, self.reserve)
self.reserve -= amount; self.alloc[mid] += amount
return amountThe monitor uses the fraction: at 80 percent of a milestone's allocation with its check not yet passing, it triggers repair rather than letting the executor keep trying. Unspent allocation from a finished milestone returns to the reserve. When the reserve is empty and a milestone is still failing, the agent stops and reports which milestones succeeded, which failed and why, which is a far better outcome than a timeout. Count tokens, tool calls and wall-clock time separately; they run out for different reasons. Agent cost control covers pricing these units.
Local repair before global replanning
When a milestone's check fails, v1 replans everything. v2 escalates through three levels and stops at the first one that works.
- Retry the step with the error message in context, when the failure looks transient or the fix is obvious (a typo in a path, a timeout). Bounded to one or two attempts.
- Re-plan the milestone: ask the tactical planner for new steps for this milestone only, given what was tried and how it failed. Completed milestones and their artifacts are untouched.
- Re-plan the strategy: only when the failure shows a milestone is impossible or its assumption was wrong, such as discovering that service C does not use the config key at all. The strategic planner receives the current plan, completed milestones and the evidence, and returns a new version that keeps finished work.
Two guards stop repair loops. Each level has an attempt limit, and each repair attempt is fingerprinted by the steps it proposes; if the planner proposes a fingerprint it already tried, escalate instead of retrying. Every repair writes a new plan version with the reason attached, so the history explains how the plan evolved.
Worked example
Take the goal in the plan above with a budget of 200,000 tokens. The gate sees side effects and three targets, so it plans. The strategic planner produces four milestones with estimated efforts of 1, 2, 2 and 2. The ledger holds back 40,000 tokens (20 percent) and splits the remaining 160,000 as about 22,900 for m1 and about 45,700 for each patch milestone.
m1 finishes for 9,000 tokens; the unspent 13,900 returns to the reserve, now 53,900. m2 and m3 run in parallel and pass their test checks. m4 fails: tests in service C break because the key is also read by a shell script the search in m1 missed. Level 1 retry does not help. Level 2 re-plans m4's steps with the failing test output in context; the new steps patch the script too, the check passes, and m4 has used 52,000 tokens, so the monitor granted 6,300 from the reserve. Total spend is about 150,000 tokens, with no completed work discarded. A v1 planner would have restarted from scratch at the m4 failure and repeated the search and two finished patches.
Versioned plans and resumption
Store each plan version, with milestone statuses, artifacts and ledger state, in a durable store keyed by task. If the process crashes, a new worker loads the latest version and resumes at the first milestone not marked done, after re-running its acceptance check, because the world may have changed. Side-effecting steps need idempotency keys so a resumed step does not, for example, open a second pull request. Agent checkpointing covers the storage patterns.
Evaluating v2 against v1
A planner change is a behaviour change, so it needs the same discipline as a model change. Build a fixed task suite from real traffic, including tasks v1 solved, tasks it failed and simple tasks that should not be planned at all. For each run record success by acceptance checks, total cost, wall-clock time, the number of repairs at each level, and gate decisions.
Run v2 in shadow on a copy of real tasks where side effects are sandboxed, compare per-task outcomes rather than averages, and look closely at regressions: tasks v1 solved and v2 did not. Then roll out by percentage, with the gate's decisions and the repair levels on a dashboard. Agent evaluation at scale describes the harness side.
Failure modes and trade-offs
- Vague acceptance checks. 'Service works' cannot be verified, so the monitor trusts the model. Push every milestone towards a check a program can run.
- Over-planning. A gate biased too far towards planning makes simple tasks slow and costly. Watch the share of planned tasks that finish in one step.
- Starved milestones. Effort estimates are guesses; a reserve and grants from it are what keep a hard milestone alive.
- Stale artifacts. A resumed or repaired plan may rely on artifacts that are now out of date; re-run checks before trusting them.
- Repair thrash. Without fingerprints and limits, level 2 repair can oscillate between two bad step lists.
The trade-off overall is complexity for predictability. v2 has more components than a prompt-and-loop agent, and each needs tests. In return you get bounded cost, partial progress instead of all-or-nothing failure, and the ability to measure whether the planner is getting better.
What to do next
- Log your current planner's outcomes for a week: steps per task, replans, cost and which tasks were simple enough to skip planning.
- Change the plan output to a typed object with milestones, dependencies and an acceptance check for each.
- Add a rule-based plan-or-act gate and record its decisions alongside task outcomes.
- Generate tactical steps per milestone just before execution, with earlier artifacts in context.
- Add a budget ledger with a reserve, and the three-level repair escalation with attempt limits and fingerprints.
- Build a task suite, run v1 and v2 side by side, review every regression, then roll out gradually.