Reflection is the pattern in which an agent produces an answer, critiques it, and revises it before anyone sees the result. It is attractive because it looks like free quality: the same model, a few more calls, and fewer mistakes. In practice it sometimes fixes real errors, sometimes changes correct answers into wrong ones, and often just burns tokens. Which outcome you get depends almost entirely on architecture: where the critique comes from, what the critic is allowed to say, and who decides when to stop.
This article treats reflection as a runtime component inside a single episode: the generate, critique, revise loop and its controller. Learning across attempts by writing lessons to memory is a related but different design, covered in Reflexion architecture, and the individual checks an agent can run are catalogued in agent output verification.
What reflection is, and what it is not
In this article, reflection means a bounded loop within one task. The generator produces a draft. A critic examines the draft, ideally with outside evidence, and produces a list of problems. A controller decides whether to stop or ask for a revision, and the generator revises using the critique. The loop ends when the draft is good enough, when it stops improving, or when a budget runs out.
Reflection is not the same as verification. A verifier returns pass or fail and the agent acts on it. A critic returns actionable feedback meant to be used for a revision. It is also not the same as retrying: a retry repeats the call hoping for a different sample, while reflection feeds specific feedback back into the next attempt. And it is not Reflexion-style learning, which carries lessons across separate attempts at a task through memory. The three combine well, but they fail in different ways, so keep them as separate components.
The central question: where does the critique signal come from?
Research on this pattern points in a consistent direction. Self-Refine (Madaan et al., 2023) showed that a model giving itself feedback and revising can improve outputs on tasks where quality is partly a matter of style and completeness, such as rewriting and code readability. CRITIC (Gou et al., 2023) had the model check its output with external tools, such as a search engine or a code interpreter, before revising, and found the tool feedback was what made correction reliable. Huang et al. (2023), in a paper titled "Large Language Models Cannot Self-Correct Reasoning Yet", found that on reasoning benchmarks, asking a model to review its own answer without any external feedback did not reliably help and could lower accuracy, because the model changed correct answers as readily as wrong ones.
The practical rule follows: reflection is only as good as the evidence the critic sees. Rank your signal sources by how grounded they are:
| Signal source | Example | Reliability |
|---|---|---|
| Executable check | Unit tests, compiler, schema validator, SQL run against a sample database | High: facts about the draft, not opinions |
| Tool lookup | Search or retrieval to confirm a cited fact, API call to confirm an ID exists | Medium to high, bounded by the tool's quality |
| Rubric critique | A model checks the draft against an explicit list of requirements | Medium: good for omissions, weak for subtle errors |
| Open self-critique | "Review your answer and improve it" | Low: frequent false alarms and sycophantic edits |
Design the loop so the top rows do the heavy lifting. Use model critique to translate evidence into actionable feedback and to check for omissions, not as an oracle of correctness.
Designing the critic
A critic that answers in free prose is hard to act on and impossible to measure. Ask for structured output with a fixed schema, and make every issue carry evidence:
{
"issues": [
{
"id": "missing-date-filter",
"severity": "blocking", // blocking | major | minor
"location": "WHERE clause",
"evidence": "Query returned 1,204,331 rows; task asks for last 7 days only",
"suggestion": "Filter on order_ts >= current_date - 7"
}
],
"score": 0.4, // 0..1 against the rubric
"requirements_checked": ["R1", "R2", "R3"]
}Four design choices matter. First, require evidence: an issue with no evidence field, or evidence that does not quote a tool result or the task, is discarded by the controller. This single rule removes most hallucinated criticism. Second, separate severity levels, because only blocking issues should force another round. Third, give the critic the rubric and the task, not the generator's reasoning; a critic that reads the chain of thought tends to accept its conclusions. Fourth, consider a different prompt or model for the critic. Sharing a model is cheaper, but a critic with the same blind spots as the generator will miss the same errors.
The controller: code, not a prompt
The most important architectural decision is that the loop is controlled by ordinary code, not by the model deciding whether it is done. The controller below keeps the best draft seen so far, stops on success, stagnation, repetition or budget, and never returns a draft that scored worse than an earlier one.
import hashlib, json
def reflect(task, generate, gather_evidence, critique,
max_rounds=3, max_tokens=40_000, min_gain=0.05):
draft, used = generate(task, feedback=None)
best, best_score, seen = None, -1.0, set()
for round_no in range(max_rounds + 1):
evidence = gather_evidence(draft) # run tests, SQL, lookups
report, cost = critique(task, draft, evidence) # JSON as in the schema above
used += cost
issues = [i for i in report["issues"] if i.get("evidence")] # drop unsupported claims
blocking = [i for i in issues if i["severity"] == "blocking"]
score = report["score"]
if score > best_score:
best, best_score = draft, score
if not blocking:
return best, "pass", round_no
if round_no > 0 and score < prev_score + min_gain:
return best, "stalled", round_no # no meaningful improvement
if used >= max_tokens or round_no == max_rounds:
return best, "budget", round_no
prev_score = score
draft, cost = generate(task, feedback=json.dumps(blocking + issues[:5]))
used += cost
h = hashlib.sha256(draft.encode()).hexdigest()
if h in seen:
return best, "oscillating", round_no # revision repeated an old draft
seen.add(h)
return best, "budget", max_roundsReturn the stop reason with the result. A "budget" or "stalled" outcome with blocking issues still open should not be treated as success: route it to a fallback, a stronger model or a human, as described in human in the loop. Modelling these outcomes as explicit states makes them testable, in the same way as the transitions in agent state machines.
Revision strategies
How the generator uses feedback matters as much as the feedback. A full rewrite gives the model freedom but often breaks parts that were already correct, which shows up as a new blocking issue in the next round. A targeted patch asks the model to change only the located parts: a diff for code, a single clause for SQL, one paragraph for prose. Patches converge faster and regress less, and they need the critic's location field to work. A good default is to patch when every blocking issue has a location and to rewrite only when the critic reports a structural problem, such as a wrong overall approach.
Pass the generator the issues, not the critic's whole transcript, and cap the list. A revision prompt with twenty minor complaints produces a revision that addresses the easy ones and misses the blocking one. Sort by severity and send the blocking issues first.
Worked example: a text-to-SQL agent
An analytics agent receives "Revenue by country for the last 7 days, top 5". The evidence step runs every draft against a sampled copy of the warehouse with a row limit and a timeout, and also runs the query planner. The run goes like this:
- Draft 1 groups by country and orders by revenue, but has no date filter. Evidence: the query scans the full table and returns 212 countries. The critic reports a blocking issue, "no 7-day filter", quoting the task and the result, and a minor issue, "no LIMIT 5". Score 0.4.
- Revision 1 is a patch: the controller sends both issues, and the generator adds a WHERE clause and LIMIT 5. Evidence: the query runs and returns 5 rows, but the planner reports a join on a currency table that multiplies rows. The critic flags revenue totals as roughly three times the reference sum from a daily summary table. Blocking. Score 0.7.
- Revision 2 joins on the currency table with the correct date key. Evidence: totals match the summary table within rounding. No blocking issues. Score 0.95. The controller returns this draft with stop reason "pass" after two revisions.
Notice what did the work. Both blocking issues were found by running the query and comparing it with a reference, not by the model rereading its SQL. An open self-critique prompt would probably have caught the missing LIMIT and missed the fan-out join.
Cost and latency arithmetic
Every round adds an evidence step, a critic call and a generator call. Suppose a draft costs 1,500 output tokens, the critic reads 3,000 tokens and writes 400, and evidence takes 2 seconds. Each round adds roughly 5,000 tokens and 6 to 10 seconds. Three rounds can quadruple the cost of a task and add half a minute of latency. Budget reflection like any other resource, as in agent cost management: set a token ceiling per task, reflect only on steps where errors are expensive, and track the average rounds per task on a dashboard. If the average creeps towards the maximum, the first draft is getting worse or the critic is getting pickier, and both need investigation.
Where to put reflection in an agent
Reflection belongs where a mistake is costly and a check is available. The strongest placements are before irreversible actions (sending an email, running a migration, committing money), on final answers that users will act on, and on artefacts with executable checks, such as code and queries. Reflecting on a plan before execution can catch missing steps, but plan critique has weak evidence, so keep it to one round. Avoid reflecting on every intermediate tool call in a ReAct loop; the loop already sees tool results at each step, and adding a critic per step multiplies cost for little gain.
Measuring whether reflection helps
Run your evaluation set with reflection off and on, and record, per task, the correctness of the first draft and of the final answer. Four counts matter: wrong to right (fixes), right to wrong (regressions), and the two unchanged cases. Reflection is worth keeping only if fixes clearly exceed regressions after you account for the cost. Also track the precision of the critic: of the blocking issues it raised, how many were real according to a human or a reference? A critic with low precision causes regressions, because the generator dutifully fixes problems that did not exist.
Failure modes
| Failure | What it looks like | Mitigation |
|---|---|---|
| Sycophantic critic | Everything passes after one round | Evidence-grounded checks, a different critic prompt or model, audits with seeded errors |
| Phantom issues | Critic invents problems, revisions break correct drafts | Discard issues without evidence, measure critic precision |
| Over-correction | Right-to-wrong flips on the evaluation set | Keep the best-scoring draft, patch instead of rewrite |
| Oscillation | Draft A, then B, then A again | Hash drafts, stop on repeat, return the best |
| Rubric gaming | Draft satisfies the checklist but not the user | Include outcome checks, not only form checks |
| Cost blow-up | Average rounds near the maximum | Token ceiling, reflect only on costly steps, alert on the rounds metric |
What to do next
- List the steps in your agent where a mistake is expensive, and for each, name an executable or tool-based check the critic can use as evidence.
- Define a critic JSON schema with severity, location and evidence, and have the controller drop any issue without evidence.
- Implement the loop controller in code with a round cap, a token ceiling, a minimum-gain rule, draft hashing and best-draft tracking.
- Run your evaluation set with reflection off and on, and count fixes and regressions separately.
- Seed known errors into drafts and measure how many the critic catches, and how often it flags correct drafts.
- Route "budget" and "stalled" outcomes with open blocking issues to a fallback or a human instead of returning them as success.