A single-turn prompt asks for one output and a human reads it. An agentic prompt starts a process: the model plans, calls tools, reads results, changes files or systems, and decides by itself when it is finished. Between your prompt and the result there may be fifty model calls and several hundred thousand tokens of intermediate context you never read. That changes what a good prompt has to do.
Most writing on agent prompts concentrates on the system prompt: the standing specification, tool descriptions and stop rules that every task shares. That layer is covered in agent prompt architecture. This article covers the other half, the task-side prompts that are written fresh for each piece of work: the task brief, the autonomy policy, the progress state the agent keeps for itself, the briefs it writes for sub-agents, and the verification step before it reports. These are where most long-running agent failures actually start, and they are the part an engineer rewrites every day.
What changes when the model acts
Three properties of an agent loop change the rules of prompting, and each one maps to a technique later in this article.
- The prompt is read many times, in contexts you did not see. The brief is re-interpreted at step 1 and at step 40, after the transcript has filled with tool output, so it needs named referents and checkable criteria rather than intentions.
- The model decides when it is done. In a loop, 'done' is a judgement made on incomplete evidence. Without a checkable definition, the most common failure is premature completion: a confident report that unfinished work is finished.
- Actions have consequences. A wrong answer can be ignored; a deleted branch or a sent email cannot. Prompts must say which actions need a human first, and that boundary should follow reversibility rather than difficulty.
The ReAct pattern, in which the model alternates reasoning and tool calls, is the usual engine underneath; ReAct prompting explains the loop itself. Everything here sits around that loop.
The architecture of an agentic task
The diagram separates the stable part of an agent, its system prompt and tools, from the per-task prompts that surround it.
For the human, data flows brief in and report out, but the agent reads and writes several artefacts in between. The progress file is its memory across context resets. Sub-agent briefs are prompts the agent authors, so your quality rules apply recursively. Verification converts 'I believe I am done' into evidence, and trajectory tests let you change any of these prompts without regressing old tasks.
Anatomy of a task brief
A good brief reads like the note you would leave a capable contractor who has never seen your codebase and cannot ask you anything for the next two hours. It has six parts:
- Task: one or two sentences, stated as an outcome, not a procedure.
- Why: the purpose. Agents make dozens of small judgement calls; knowing the goal lets them make the same call you would. 'Behaviour must not change' is worth more than ten procedural rules.
- Context the agent would not find quickly: the known traps, the flaky test, the undocumented invariant. This is the highest-value part of the brief because it is exactly what the agent cannot discover by exploring.
- Done when: checkable criteria, each with the command or observation that proves it.
- Autonomy: what to do without asking, what to ask first, and what never to do.
- Report format: what to return, including an explicit slot for things not verified.
Here is a complete brief for a realistic migration task. Notice what it does not contain: no step-by-step procedure, no persona, no instructions to 'be careful'. It spends its words on facts and boundaries.
TASK
Migrate the three services in services/billing/* from the v1 config loader
to config_v2. Behaviour must not change.
WHY
config_v1 is removed in the next platform release (end of this quarter).
Billing is the last consumer.
CONTEXT YOU WOULD NOT FIND QUICKLY
- config_v2 returns None for missing keys; v1 raised KeyError. Callers in
invoice.py rely on the KeyError to fall back to defaults.
- tests/billing/test_config.py is known flaky on timezone; ignore that test.
DONE WHEN
- No imports of config_v1 remain under services/billing (grep proves it).
- `make test-billing` passes, excluding the flaky test named above.
- A short report lists every behavioural difference you had to handle.
AUTONOMY
- Act without asking: edits inside services/billing and tests/billing.
- Ask first: any change outside those paths, any dependency change.
- Never: push, deploy, edit CI config, delete tests.
REPORT FORMAT
Changed files; evidence for each done-when item (command + result);
anything you could not verify, stated plainly.
Done criteria that a machine can check
The single most effective change most teams can make to their agent prompts is to replace 'make it work' with criteria that can be checked by running something. Compare 'fix the failing tests' with 'all tests in tests/billing pass under make test-billing, and no test was deleted or skipped'. The first invites the agent to declare victory after one green run, or to satisfy the letter by deleting the test. The second names the evidence and closes the loophole.
Good criteria are observable (a command, a grep, an HTTP status), complete (together they cover the goal, including 'nothing else changed'), and they name the forbidden shortcuts, because an agent optimising a check will find the cheap way to pass it: tests that must not be edited, outputs that must not be hard-coded.
Where you can, run the checks outside the model. A model's statement that tests pass is a claim; the exit code is evidence. The harness below renders the brief, runs the agent, runs the checks itself, and on failure feeds back the failing command verbatim. run_agent is a placeholder for whatever model and tool loop you use.
from dataclasses import dataclass, field
import json, pathlib, subprocess
@dataclass
class Brief:
task: str
why: str
context: list[str]
done_when: list[str] # each item must be checkable
act: list[str] # allowed without asking
ask: list[str] # needs approval
never: list[str]
checks: list[str] = field(default_factory=list) # shell commands the harness runs
def render(b: Brief, progress_path: str) -> str:
sec = lambda title, items: title + "\n" + "\n".join("- " + i for i in items)
return "\n\n".join([
"TASK\n" + b.task, "WHY\n" + b.why,
sec("CONTEXT YOU WOULD NOT FIND QUICKLY", b.context),
sec("DONE WHEN", b.done_when),
"AUTONOMY\n" + sec("Act without asking:", b.act) + "\n"
+ sec("Ask first:", b.ask) + "\n" + sec("Never:", b.never),
f"PROGRESS\nKeep {progress_path} current: plan, done, open questions. "
"Re-read it before starting each new step.",
])
def run_task(b: Brief, run_agent, max_rounds: int = 3) -> dict:
"""run_agent(prompt) -> str is a placeholder for your model and tool loop."""
if not b.checks:
raise ValueError("a brief needs at least one runnable check")
prompt = render(b, "PROGRESS.md")
for round_ in range(max_rounds):
report = run_agent(prompt)
failed = [cmd for cmd in b.checks
if subprocess.run(cmd, shell=True).returncode != 0]
if not failed:
return {"status": "done", "rounds": round_ + 1, "report": report}
# Feed back evidence, not a verdict: which check failed, verbatim.
prompt = ("Your report said the task was complete, but these checks fail:\n"
+ "\n".join(failed)
+ "\nRe-read PROGRESS.md, fix the cause, and report again.")
return {"status": "needs_human", "report": report, "failed": failed}
Autonomy boundaries keyed to reversibility
Agents fail in two opposite directions: they stop to ask about trivia, or they take a consequential action without asking. Both come from the same missing information, which is where the line is. Draw it by reversibility and blast radius, not by how hard the task seems.
| Action class | Examples | Default policy |
|---|---|---|
| Local and reversible | edit files in the workspace, run tests, read logs | act without asking |
| Reversible but shared | open a pull request, create a branch, post a draft | act, then report it |
| Hard to reverse or external | push to main, send email, delete data, spend money, change permissions | ask first, every time |
| Out of scope | edit CI, rotate credentials, disable checks | never; report instead |
State the policy in the brief, and enforce the hard edges in the tool layer as well, because a prompt is a request and a permission is a guarantee. An approval granted for one action should not be read as blanket permission for similar actions later in the run; say so in the prompt. The enforcement side is covered in permission boundaries for autonomous agents.
Also tell the agent what to do when it is blocked. 'If a required decision falls outside your autonomy, record it under open questions, continue with everything that does not depend on it, and report' keeps a run productive instead of either stalling or guessing.
Externalised progress state
Long runs outgrow the context window. When the harness compacts or truncates history, the agent loses its plan unless the plan lives somewhere it can re-read. Ask the agent to maintain a small progress file and to re-read it before each new step:
# PROGRESS.md (owned by the agent; re-read at the start of every step)
## Plan
1. [x] Inventory config_v1 imports (7 found, listed in notes)
2. [x] Migrate ledger service
3. [ ] Migrate invoice service <- in progress; KeyError fallback needs a shim
4. [ ] Migrate payouts service
5. [ ] Run checks, write report
## Decisions
- Added get_required() helper that raises KeyError, to preserve invoice.py behaviour.
## Open questions (need a human)
- payouts reads BILLING_REGION from env and config; which wins? Not guessing.It survives context compaction, so a resumed agent picks up at step 3 instead of redoing step 2 differently. It records decisions, so later steps do not contradict earlier ones. And it gives a human a view of an in-flight run. Keep it short; a progress file that becomes a diary costs as much context as the history it replaced. How harnesses compress history is covered in context compaction for long-running agents.
Delegation: writing briefs for sub-agents
When an agent delegates to sub-agents it becomes a prompt author, and usually a poor one. The characteristic failure is a one-line instruction such as 'look into the payouts config' that omits everything the parent learned, so the sub-agent rediscovers it and answers a slightly different question.
Tell the orchestrator that a delegation brief must be self-contained and follow your anatomy: goal and why, what is known and ruled out, what counts as done, what may be changed, and the exact return shape, such as findings with file and line evidence plus a separate list of unverified guesses. The same discipline of explicit step contracts appears in prompt chaining, where the pipeline is fixed rather than chosen by the model.
Verification before the report
Ask for verification as a distinct step with its own instructions, not as a closing sentence. A useful formulation is: before reporting, re-read the done-when list; for each item, run the check and quote the result; for anything you could not check, say so plainly under 'not verified'. This turns the report from a narrative into an evidence table and makes gaps visible instead of glossed over.
Forbid hedged success: 'should work now' is not a status; an item is verified, failed or not checked. And reward honest partial reports explicitly, for example 'two of three items done, with the third explained, beats a claim of all three'. Agents mirror the incentives the prompt describes.
Testing agentic prompts with trajectories
Agent prompts regress like code: a sentence added to fix one task quietly breaks another. Keep 10 to 30 representative briefs with known-good outcomes, replay them after every prompt change, and grade more than the final answer: did the checks pass, at what step and token cost, did it ask exactly when it should, did any action violate the autonomy policy, and do the report's claims match the checks. Runs are stochastic, so compare pass rates over several runs.
Failure modes and trade-offs
- Premature completion. Vague done criteria; fix with checkable ones run by the harness.
- Specification gaming. The agent edits the test or hard-codes the output. Name forbidden shortcuts and diff the test directory.
- Scope creep. Helpful refactors nobody asked for. State the scope and 'behaviour must not change' in the brief.
- Retry spirals. The same failing action repeated with small variations. Ask the agent to record each failed approach in the progress file and to stop and report after a fixed number of attempts.
- Over-asking. An agent with no autonomy policy defaults to asking, and a long run stalls overnight. Grant local, reversible actions explicitly.
- Brief bloat. Every incident adds a rule until the brief contradicts itself. Prune with trajectory tests, and move stable rules into the system prompt.
The central trade-off is autonomy against oversight: wider autonomy finishes more work per human minute and raises the cost of each mistake. Reversibility-based policies, harness-run checks and honest reports let you widen it safely.
What to do next
- Take your last three agent tasks and rewrite each brief with the six parts: task, why, context, done when, autonomy, report format.
- Turn every done criterion into a command, and run those commands in the harness rather than trusting the report.
- Write an autonomy table for your environment keyed to reversibility, and enforce its 'never' row in the tool layer.
- Add a progress file with plan, decisions and open questions, and instruct the agent to re-read it before each step.
- Give orchestrating agents a delegation template so sub-agent briefs are self-contained.
- Build a replay set of 10 to 30 briefs and grade pass rate, policy violations and report honesty after every prompt change.