Most agent planners start the same way: ask the model for a numbered list of steps, execute them in order, and when something fails, ask for a new list. That v1 design works for demos and for tasks of five or six steps. It breaks in recognisable ways as tasks grow: plans that are confidently wrong from step one, replans that throw away work already done, budgets blown on easy tasks that never needed a plan, and no way to tell whether a change to the planner made things better.

This article describes a v2 architecture that addresses those failures one by one. It assumes you know the basics of a planner, such as the goal contract, plan validation and replanning triggers, which are covered in agent planner architecture. Here the focus is on what to change, why, and how to prove the change helped.

Advertisement

What breaks in v1, and the v2 answer

Symptom in v1Root causev2 mechanism
Simple questions take ten model callsEvery request is plannedPlan-or-act gate
Plans are detailed but wrong about later stepsSteps are fixed before earlier results are knownTwo tiers: milestones now, steps just in time
One tool failure restarts the whole taskReplan is globalLocal repair of the failing subtree
Runaway cost on hard tasksNo budget, or one global capBudget ledger with per-milestone allocations
Progress lost on crash or timeoutPlan lives only in the promptVersioned plan store, resumable
Unclear whether a planner change helpedNo fixed evaluationTask suite, shadow runs, gated rollout

None of these mechanisms is exotic. The value of v2 is in combining them so each one covers a weakness of the others, and in making plan state explicit data rather than prose in a context window.

The architecture

Planner v2: decide whether to plan, plan in two tiers, spend from a ledger, repair locallyGoal + contexttask contractPlan-or-act gatecheap classifierDirect executorsingle ReAct loopStrategic plannermilestones + checksTactical plannersteps per milestoneExecutortools, sandboxMonitorcheck, budget, driftRepairlocal, then globalBudget ledgertokens, time, moneyPlan storeversioned, resumablesimplecomplexmilestonestepsresultsfailed checkpatched subtreespendnew versionRepair touches the smallest failing subtree first; a global replan is the exception, not the default.
A gate sends simple requests straight to a single execution loop. Complex ones get a strategic plan of milestones, each expanded into steps only when it starts. A monitor checks results and spend; failures go to repair, which patches the smallest failing subtree and writes a new plan version.

Data flows in one direction during normal operation: goal to gate, gate to strategic planner, milestone to tactical planner, steps to executor, results to monitor. The loop closes only through the monitor, which is the single place that decides a milestone has passed, failed or run out of budget. Keeping that decision in one component makes the agent's behaviour auditable: every transition is a logged event with a reason.

Advertisement

Plans as data: milestones with acceptance checks

In v2 the plan is a typed object the runtime can inspect, not a paragraph. The strategic tier produces milestones; each milestone carries the condition that proves it is done and a share of the budget. Steps are attached only when the milestone becomes active.

{
  "plan_id": "p-7f3a", "version": 3, "goal": "Rename config key db.url to db.dsn in services A, B, C",
  "milestones": [
    {"id": "m1", "title": "Find every usage of db.url",
     "depends_on": [], "budget": {"tokens": 20000, "tool_calls": 10},
     "accept": {"type": "artifact", "name": "usage_list", "non_empty": true},
     "status": "done"},
    {"id": "m2", "title": "Patch service A",
     "depends_on": ["m1"], "budget": {"tokens": 30000, "tool_calls": 15},
     "accept": {"type": "command", "run": "make test", "exit_code": 0},
     "status": "active",
     "steps": [{"id": "m2.s1", "tool": "edit_file", "args": {"path": "a/config.py"}},
               {"id": "m2.s2", "tool": "run", "args": {"cmd": "make test"}}]},
    {"id": "m3", "title": "Patch service B", "depends_on": ["m1"], "status": "pending"},
    {"id": "m4", "title": "Patch service C", "depends_on": ["m1"], "status": "pending"}
  ]
}

Acceptance checks are the heart of the design. A check that a program can evaluate, such as a test command, a schema validation or the presence of a non-empty artifact, is far more reliable than asking the model whether it succeeded. Where no programmatic check exists, use a separate verifier prompt with the milestone's stated outcome, and treat its verdict as weaker evidence. The pattern for building checks is covered in agent output verification.

Because milestones have explicit dependencies, independent ones (m3 and m4 above) can run in parallel, and the plan doubles as a progress report for a human watching the task.

The plan-or-act gate

Planning has a cost: at least one extra model call, more tokens in context and more places to go wrong. For a question answerable with one tool call, a plan is pure overhead. The gate decides, before any planning, which path a request takes.

def route(goal, ctx):
    features = {
        "est_tool_calls": estimate_tool_calls(goal),     # small model or heuristic
        "touches_writes": mentions_side_effects(goal),   # deploy, delete, send, pay
        "multi_target": count_targets(goal) > 1,         # several files, services, people
        "history_success_direct": ctx.stats.direct_success_rate(goal_kind(goal)),
    }
    if features["touches_writes"] or features["multi_target"]:
        return "plan"
    if features["est_tool_calls"] <= 3 and features["history_success_direct"] >= 0.9:
        return "direct"
    return "plan"

Start with rules like these and log every decision with the eventual outcome. After a few hundred tasks you can see which direct-routed tasks failed and should have been planned, and which planned tasks finished in one step and wasted the plan. Tighten the rules from that data, or train a small classifier on it. Bias the gate towards planning whenever a task has side effects: a wasted plan costs tokens, a skipped plan on a destructive task can cost much more.

Two tiers: strategic and tactical

A single-tier planner writes every step at the start, when it knows least. The v2 planner commits early only to what it can justify: the milestones and their checks. Steps for a milestone are generated just before it runs, with the results of earlier milestones in context. In the example, the steps for patching service B are written after m1 has produced the real list of files, so the planner does not invent file paths.

The strategic prompt should be short and stable: the goal, the available tool families (not every tool schema), constraints, and the required output format. The tactical prompt is concrete: one milestone, its acceptance check, its remaining budget, the relevant artifacts from earlier milestones and the full schemas of the tools it may use. Splitting the context this way keeps both prompts small, which improves reliability and cost at the same time. How to cut a goal into milestones well is its own skill, covered in agent task decomposition.

The budget ledger

A global cap tells you when to stop, but not where the money went. A ledger allocates the task's budget across milestones when the plan is made, records every spend against the active milestone, and lets the monitor act before the total is gone.

class Ledger:
    def __init__(self, total, reserve_frac=0.2):
        self.total, self.reserve = total, total * reserve_frac   # held back for repair
        self.alloc, self.spent = {}, {}

    def allocate(self, milestones):
        pool = self.total - self.reserve
        weights = {m.id: m.estimated_effort for m in milestones}
        s = sum(weights.values())
        self.alloc = {mid: pool * w / s for mid, w in weights.items()}

    def charge(self, mid, cost):
        self.spent[mid] = self.spent.get(mid, 0) + cost
        return self.spent[mid] / self.alloc[mid]          # fraction used

    def grant_from_reserve(self, mid, amount):
        amount = min(amount, self.reserve)
        self.reserve -= amount; self.alloc[mid] += amount
        return amount

The monitor uses the fraction: at 80 percent of a milestone's allocation with its check not yet passing, it triggers repair rather than letting the executor keep trying. Unspent allocation from a finished milestone returns to the reserve. When the reserve is empty and a milestone is still failing, the agent stops and reports which milestones succeeded, which failed and why, which is a far better outcome than a timeout. Count tokens, tool calls and wall-clock time separately; they run out for different reasons. Agent cost control covers pricing these units.

Local repair before global replanning

When a milestone's check fails, v1 replans everything. v2 escalates through three levels and stops at the first one that works.

  1. Retry the step with the error message in context, when the failure looks transient or the fix is obvious (a typo in a path, a timeout). Bounded to one or two attempts.
  2. Re-plan the milestone: ask the tactical planner for new steps for this milestone only, given what was tried and how it failed. Completed milestones and their artifacts are untouched.
  3. Re-plan the strategy: only when the failure shows a milestone is impossible or its assumption was wrong, such as discovering that service C does not use the config key at all. The strategic planner receives the current plan, completed milestones and the evidence, and returns a new version that keeps finished work.

Two guards stop repair loops. Each level has an attempt limit, and each repair attempt is fingerprinted by the steps it proposes; if the planner proposes a fingerprint it already tried, escalate instead of retrying. Every repair writes a new plan version with the reason attached, so the history explains how the plan evolved.

Worked example

Take the goal in the plan above with a budget of 200,000 tokens. The gate sees side effects and three targets, so it plans. The strategic planner produces four milestones with estimated efforts of 1, 2, 2 and 2. The ledger holds back 40,000 tokens (20 percent) and splits the remaining 160,000 as about 22,900 for m1 and about 45,700 for each patch milestone.

m1 finishes for 9,000 tokens; the unspent 13,900 returns to the reserve, now 53,900. m2 and m3 run in parallel and pass their test checks. m4 fails: tests in service C break because the key is also read by a shell script the search in m1 missed. Level 1 retry does not help. Level 2 re-plans m4's steps with the failing test output in context; the new steps patch the script too, the check passes, and m4 has used 52,000 tokens, so the monitor granted 6,300 from the reserve. Total spend is about 150,000 tokens, with no completed work discarded. A v1 planner would have restarted from scratch at the m4 failure and repeated the search and two finished patches.

Versioned plans and resumption

Store each plan version, with milestone statuses, artifacts and ledger state, in a durable store keyed by task. If the process crashes, a new worker loads the latest version and resumes at the first milestone not marked done, after re-running its acceptance check, because the world may have changed. Side-effecting steps need idempotency keys so a resumed step does not, for example, open a second pull request. Agent checkpointing covers the storage patterns.

Evaluating v2 against v1

A planner change is a behaviour change, so it needs the same discipline as a model change. Build a fixed task suite from real traffic, including tasks v1 solved, tasks it failed and simple tasks that should not be planned at all. For each run record success by acceptance checks, total cost, wall-clock time, the number of repairs at each level, and gate decisions.

Run v2 in shadow on a copy of real tasks where side effects are sandboxed, compare per-task outcomes rather than averages, and look closely at regressions: tasks v1 solved and v2 did not. Then roll out by percentage, with the gate's decisions and the repair levels on a dashboard. Agent evaluation at scale describes the harness side.

Failure modes and trade-offs

  • Vague acceptance checks. 'Service works' cannot be verified, so the monitor trusts the model. Push every milestone towards a check a program can run.
  • Over-planning. A gate biased too far towards planning makes simple tasks slow and costly. Watch the share of planned tasks that finish in one step.
  • Starved milestones. Effort estimates are guesses; a reserve and grants from it are what keep a hard milestone alive.
  • Stale artifacts. A resumed or repaired plan may rely on artifacts that are now out of date; re-run checks before trusting them.
  • Repair thrash. Without fingerprints and limits, level 2 repair can oscillate between two bad step lists.

The trade-off overall is complexity for predictability. v2 has more components than a prompt-and-loop agent, and each needs tests. In return you get bounded cost, partial progress instead of all-or-nothing failure, and the ability to measure whether the planner is getting better.

What to do next

  1. Log your current planner's outcomes for a week: steps per task, replans, cost and which tasks were simple enough to skip planning.
  2. Change the plan output to a typed object with milestones, dependencies and an acceptance check for each.
  3. Add a rule-based plan-or-act gate and record its decisions alongside task outcomes.
  4. Generate tactical steps per milestone just before execution, with earlier artifacts in context.
  5. Add a budget ledger with a reserve, and the three-level repair escalation with attempt limits and fingerprints.
  6. Build a task suite, run v1 and v2 side by side, review every regression, then roll out gradually.
Key takeaway: A v2 agent planner stops treating every request the same and every failure as a reason to start over. Gate simple requests to direct execution, plan complex ones as milestones with programmatic acceptance checks, write steps just in time, spend from a ledger with a reserve, and repair the smallest failing piece before replanning the whole task. Store plans as versioned data so tasks resume, and prove each planner change on a fixed task suite before rolling it out.