A coding agent is a language model in a loop: it reads files, runs commands, edits code, looks at the result and decides what to do next. Give it a one-line request and a large repository and it will usually produce something that looks finished. Most disappointing sessions fail in the same place: nobody separated deciding what to do from doing it, and nobody checked the result with evidence the agent could not influence.

This guide treats agentic coding as a workflow design problem: the plan, execute, verify loop, subagents for clean context, parallel agents in git worktrees, and merging their output safely. Examples use Claude Code where a concrete tool is needed, because its behaviour was checked for this article, but the structure applies to any agent that can read files, run a shell and edit code.

Advertisement

Why a loop and not a prompt

An agent's output is bounded by what is in its context when it acts and by the feedback it gets afterwards. A single long prompt fails on both: early decisions are made before the agent has read the code that would have changed them, and late in the session the context is full of stale reads and abandoned attempts.

The loop gives each phase its own output. Planning produces a document. Execution produces small commits against it. Verification produces a pass or fail verdict from checks outside the agent. Each output can be inspected in a minute and retried alone, and when something breaks you know whether the plan, the implementation or the check was at fault.

Plan, execute, verify: one orchestrator, isolated workers, independent checks1. Planread-only, plan fileHuman gateapprove / edit planPartitiondisjoint file setsBudgetturns, time, costWorker Aworktree wt-aWorker Bworktree wt-bWorker Cworktree wt-cExplore subagentread-only researchVerify Atests, lint, scopeVerify Btests, lint, scopeVerify Ctests, lint, scopeFresh reviewernew context, diff onlyIntegrate + re-verify on the merged treemerge one branch at a timereplan
The orchestrator plans read-only and a human approves the plan. Work is partitioned into disjoint file sets, each worker runs in its own git worktree, and every worker's output is verified by the harness and optionally by a reviewer with fresh context. Branches are merged one at a time and the combined tree is verified again. A scope violation sends the step back to planning.

Phase 1: plan without touching anything

Planning is read-only exploration followed by a written plan. In Claude Code, start with claude --permission-mode plan or press Shift+Tab until the status bar shows plan mode; the agent reads files and proposes a plan but makes no edits until you approve. If your tool has no equivalent, deny write tools for the session.

Save the plan to a file, not the chat: a file survives the session, can be reviewed in a pull request and can be handed to other agents. A useful plan has a goal, the evidence gathered (with the commands that produced it), steps with the files each touches, verification commands and stop conditions:

# plans/2026-09-29-structlog-migration.md   (written in plan mode, reviewed by a human)

## Goal: replace logutil.log(level, msg) with structlog. No behaviour change.
## Evidence (read-only)
- 41 call sites in 17 files: rg -n "logutil\.log\(" src/
- logutil.log swallows formatter exceptions; structlog does not
## Steps (each ends in a commit and a green check)
1. Add src/obs/log.py (structlog config)      orchestrator
2. src/billing/**  (9 sites)                  worker A
3. src/orders/**   (14 sites)                 worker B
4. src/accounts/** (18 sites)                 worker C
5. Delete logutil.log once rg finds nothing   orchestrator
## Verification: make test lint; tests/test_logging_contract.py unchanged
## Stop and ask if: a call passes exc_info, or any existing test must change

The evidence section forces the agent to show it looked: a call-site count from a real search is much harder to hallucinate than a vague claim. The stop and ask line turns your worries into triggers, so the agent pauses instead of improvising. If a step says "update the tests as needed", reject it: that is where an agent will quietly weaken the check it is meant to satisfy.

Advertisement

Phase 2: execute in bounded steps

Execution runs one plan step at a time, and every step ends in a commit and a green check. Small commits are the undo button: if step three goes wrong, reset to the end of step two instead of untangling a large diff. Each commit also maps to one line of the plan, which keeps review tractable.

Bound each step three ways: scope (the files it may touch), retries (two or three; an agent that has failed three times on one error is usually confused about the problem, not one edit away) and budget (turns, time or spend):

def run_task(plan_step, max_attempts=3):
    """Execute one plan step with a bounded fix loop and independent verification."""
    base = git("rev-parse", "HEAD")
    for attempt in range(1, max_attempts + 1):
        agent.run(
            instructions=plan_step.instructions,
            allowed_paths=plan_step.files,           # scope is part of the task
            feedback=last_failure if attempt > 1 else None,
        )
        result = verify(plan_step, base)             # runs outside the agent
        if result.ok:
            git("commit", "-am", f"{plan_step.id}: {plan_step.title}")
            return "done"
        last_failure = result.summary()              # failing test names, first error only
        if result.out_of_scope_files:                # touched files it was not given
            git("reset", "--hard", base)
            return "replan"                          # the plan was wrong, not the code
    git("reset", "--hard", base)
    return "escalate"                                # a human looks, with the failure log

escalate means the plan was probably right but the implementation is stuck, so a human reads the failure log. replan means the agent needed files the plan did not anticipate: evidence the plan was wrong, fixed by updating the plan rather than silently widening scope. On retry, feed back only the first error and failing test names; a thousand lines of output is noise.

Phase 3: verify with evidence the agent cannot edit

The thing being checked must not control the check. An agent told "make the tests pass" can delete an assertion, add a skip marker, swallow an exception or special-case the test input, and each turns the check green without doing the work. These shortcuts get more likely after many failed attempts and when instructions are vague about what may change. So verification runs in the harness after the agent says it is done, and checks scope as well as behaviour:

#!/usr/bin/env python3
"""verify_step.py BASE_SHA ALLOWED_GLOB... -- run by the harness, never by the worker agent."""
import fnmatch, subprocess, sys

def sh(*cmd):
    return subprocess.run(cmd, capture_output=True, text=True)

base, allowed = sys.argv[1], sys.argv[2:]
changed = sh("git", "diff", "--name-only", base).stdout.split()
untracked = sh("git", "ls-files", "--others", "--exclude-standard").stdout.split()
outside = [f for f in changed + untracked
           if not any(fnmatch.fnmatch(f, g) for g in allowed)]
if outside:
    print("OUT OF SCOPE:", *outside, sep="\n  "); sys.exit(3)

if sh("git", "diff", "--quiet", base, "--", "tests/test_logging_contract.py").returncode:
    print("contract test was modified"); sys.exit(4)

for cmd in (["make", "lint"], ["make", "typecheck"], ["make", "test"]):
    r = sh(*cmd)
    if r.returncode:
        print(f"FAILED: {' '.join(cmd)}\n{r.stdout[-3000:]}{r.stderr[-2000:]}"); sys.exit(1)
print("ok")

Distinct exit codes for scope violations (3), protected-file tampering (4) and ordinary failures (1) let the orchestrator route each differently. Add cheap project invariants such as "the search for the old API returns nothing"; they catch tests that pass because they never exercised the changed code.

Subagents: spending context where it matters

A subagent is a separate agent invocation with its own context window. In Claude Code it does not see the parent's history or the files the parent read; it gets a delegation message, its own system prompt, the project instruction files and a git status snapshot, and returns a result. Use that isolation for two jobs.

Exploration. Searching a large codebase fills context with files needed only long enough to extract one fact. Delegate it and only the findings come back.

Independent review. An agent that just wrote a change is its worst reviewer, because its context is full of its own reasons why the change is right. A subagent that sees only the plan step and the diff has no such bias. Subagents are Markdown files with YAML frontmatter in .claude/agents/ (project) or ~/.claude/agents/ (user); tools is an allowlist:

---
name: verifier
description: Independently checks a finished change against its plan step. Use after a worker reports done.
tools: Read, Grep, Glob, Bash
model: sonnet
---
You review a change you did not write, given a plan step and a base commit.
Read git diff <base>, run the plan's verification commands, and list any
change the step did not ask for. Answer PASS or FAIL with findings.

The reviewer gets no edit tools, so it cannot fix what it finds and then approve its own fix. Keep outputs short and structured so the orchestrator can act on them, and avoid deep nesting: a chain of three summaries tends to lose the one detail that mattered.

Parallel work with git worktrees

Two agents in one checkout collide quickly: they edit the same files, one runs tests mid-way through the other's change, and both fight over the git index. A git worktree gives each agent its own directory and branch while sharing one object store, so it is cheap and branches merge normally:

for a in billing orders accounts; do
  git worktree add -b "mig/$a" "../wt-$a" main       # own directory and branch
done
git worktree remove ../wt-billing                     # after its branch is merged

Claude Code wraps this: claude --worktree feature-auth starts a session in its own worktree, and a subagent with isolation: worktree in its frontmatter runs in a temporary worktree, removed if it makes no changes. That worktree branches from the default branch, not the parent session's HEAD, so the subagent cannot see uncommitted work or an unmerged feature branch. Merge what workers need before fanning out.

Worktrees isolate files, not ports, local databases or caches. Give each worktree its own, and budget disk and time for per-worktree dependency installs.

Fan-out and fan-in

Parallelism pays when the plan partitions into steps that touch disjoint files and do not depend on each other. Partition by directory or module, never by halves of one file. claude -p runs a prompt non-interactively, which makes fan-out scriptable:

#!/usr/bin/env bash
set -euo pipefail
base=$(git rev-parse main)
declare -A scope=( [billing]="src/billing/*" [orders]="src/orders/*" [accounts]="src/accounts/*" )

for area in "${!scope[@]}"; do
  (
    cd "../wt-$area"
    claude -p "Do step '$area' of plans/2026-09-29-structlog-migration.md. Only edit ${scope[$area]}. Do not commit." \
      --permission-mode acceptEdits > "../logs/$area.txt" 2>&1
    python3 ../repo/tools/verify_step.py "$base" "${scope[$area]}" "tests/*" \
      && git commit -qam "structlog: migrate $area" \
      && echo "$area PASS" || echo "$area FAIL"
  ) &
done
wait

# fan-in: merge sequentially, re-verify on the combined tree after each merge
for area in billing orders accounts; do
  git merge --no-ff "mig/$area" && make test || { echo "integration broke on $area"; exit 1; }
done

Fan-in is where parallel work usually breaks. Each branch passed checks against the old main, but the combination is untested. Merge one at a time and run the suite after each, so you know which branch broke integration. A conflict between branches is a partitioning failure: resolve it by hand or re-run the later step on the merged result.

Worked example: the logging migration

For the plan above: in plan mode the orchestrator finds 41 call sites and notices that logutil.log swallows formatter exceptions, which becomes a stop condition. A human approves the plan. Step 1 runs alone because the others depend on it, and is merged to main.

Three worktrees are created from the updated main. Billing and accounts pass verification first time. Orders fails with exit code 3: it edited src/shared/fmt.py to fix an import. That is a replan, not a retry; the orchestrator adds a step for the shared module, merges it and reruns orders, which passes. The fresh-context verifier flags a billing call that passed exc_info=True, which the worker should have stopped on; it is fixed by hand. Three sequential merges stay green, and step 5 deletes the old function. Five focused commits, each reviewable against a line of the plan.

Failure modes

  • Plan theatre. A plan without evidence or verification commands is a to-do list; approving it approves nothing.
  • Test tampering. Checks the agent can edit become optional. Protect contract tests and snapshots by path in the verifier.
  • Scope creep. Drive-by refactors inflate diffs and cause conflicts between parallel workers. Enforce scope mechanically.
  • Context rot. Long sessions accumulate stale reads. Start fresh per step and carry state in the plan file and commits; see context compaction.
  • Stale base. Worktrees branched before a dependency was merged build on the wrong code. Fan out only after the prerequisite commits land.
  • Integration blind spots. Green branches, red main. Always verify the merged tree.

Trade-offs and costs

The loop costs time up front; for a two-line fix, just make the change and run the tests. It earns its keep when a change spans many files, the codebase is unfamiliar or mistakes are expensive. Parallelism multiplies spend and review load along with throughput: three workers produce three diffs a human must read, and fan-in is serial. Past a handful of workers the human gate becomes the bottleneck; see how to review a pull request.

For the broader design space of coordinating multiple agents, including hierarchical and peer patterns, see multi-agent orchestration and context engineering for agents. For migrations of old code, handling legacy code covers the characterisation-test groundwork that makes agent refactors safe.

What to do next

  • Pick one multi-file change this week and run it through plan mode, saving the plan to a file with evidence, steps, verification and stop conditions.
  • Write a verify script for your repository that checks scope, protected files and the test suite, and run it from the harness, not the agent.
  • Define a read-only reviewer subagent and use it on every agent-written change before you read the diff yourself.
  • Try two parallel worktrees on a cleanly partitioned task, with separate ports and databases, and merge them one at a time.
  • Track how often steps end in escalate or replan; frequent replans mean your plans need better evidence gathering.
Key takeaway: Treat agentic coding as a workflow, not a prompt. Plan read-only and write the plan to a file with evidence, steps, verification commands and stop conditions. Execute one bounded step at a time, committing after each green check. Verify in the harness with checks the agent cannot edit, including scope and protected files. Use subagents to keep exploration out of the main context and to get a review from fresh eyes. Parallelise with git worktrees only when the work partitions into disjoint files, and always re-verify the merged tree.