Reflexion is a pattern for making a language model agent better at a task across attempts without changing its weights. The agent tries, an evaluator says whether it worked, and if it failed the model writes a short lesson in plain language about what went wrong. The next attempt sees the task plus those lessons. The original paper by Shinn and colleagues (arXiv 2303.11366, 2023) called this verbal reinforcement learning: the feedback signal is turned into text and stored in an episodic memory instead of a gradient.
The architecture is described in the Reflexion architecture article. This article is about building one that works: what the paper's results actually say, including the benchmark where it made things worse, a complete loop you can adapt, why the evaluator matters more than the prompt, how to write reflections that change behaviour, and when not to use the pattern at all.
How it differs from a retry and from Self-Refine
A plain retry samples again and hopes for a different result. It works surprisingly often, because sampling is random, but each attempt is independent: nothing learned on attempt two reaches attempt three. Self-Refine has the model critique and rewrite its own output, with no external signal at all; the model is both the student and the examiner. Reflexion sits between them. It needs an evaluator outside the actor, ideally an executable one such as unit tests or an environment's success signal, and it carries lessons forward across trials in a bounded memory.
That combination is the whole point. The external signal tells the loop that something failed and roughly where; the reflection turns that signal into an instruction the next attempt can follow; the memory makes improvement cumulative instead of a fresh guess each time.
What the paper measured, including where it lost
The headline result is HumanEval: 91.0% pass@1 with Reflexion against 80.1% for a single GPT-4 sample. On AlfWorld, a text-based household task environment, ReAct plus Reflexion completed 130 of 134 tasks. The full programming table is more instructive than the headline:
| Benchmark | GPT-4, single sample | Reflexion |
|---|---|---|
| HumanEval (Python) | 80.1 | 91.0 |
| MBPP (Python) | 80.1 | 77.1 |
| HumanEval (Rust) | 60.0 | 68.0 |
| MBPP (Rust) | 70.9 | 75.4 |
| LeetcodeHardGym (Python) | 7.5 | 15.0 |
Reflexion was worse on MBPP Python. The paper's explanation is the evaluator: the agent wrote its own unit tests and used them as the success signal, and on MBPP Python 16.3% of executions were false positives, solutions that passed the self-written tests but failed the real ones, against 1.4% on HumanEval Python. When the evaluator approves a wrong answer, the loop stops on it.
An ablation on the 50 hardest HumanEval Rust problems makes the same point. The base model scored 0.60. Removing test generation but keeping self-reflection scored 0.52, worse than doing nothing. Keeping tests but removing reflection scored 0.60. Only both together reached 0.68. Reflection without a grounded signal hurts; a signal without reflection does not help on its own. The paper also reports that Reflexion did not help on WebShop, a shopping environment where success needs creative exploration rather than fixing a specific mistake.
The loop in code
The control flow is short. Most of the engineering is in the evaluator and the two prompts.
from dataclasses import dataclass, field
@dataclass
class Trial:
attempt: str
passed: bool
feedback: str # failing test names, tracebacks, env messages
@dataclass
class Reflexion:
llm: callable # llm(prompt: str) -> str
evaluate: callable # evaluate(attempt) -> (passed, feedback)
max_trials: int = 4
memory_size: int = 3 # the paper bounds memory at 1-3 reflections
memory: list = field(default_factory=list)
def run(self, task: str) -> tuple[str, list[Trial]]:
self.memory, history = [], [] # lessons belong to one task only
for t in range(self.max_trials):
attempt = self.llm(actor_prompt(task, self.memory[-self.memory_size:]))
passed, feedback = self.evaluate(attempt)
history.append(Trial(attempt, passed, feedback))
if passed:
return attempt, history
if t < self.max_trials - 1: # no reflection after the last trial
lesson = self.llm(reflect_prompt(task, attempt, feedback, self.memory))
self.memory.append(lesson.strip())
best = history[-1].attempt # or the attempt that passed most tests
return best, history # caller must treat this as a failureTwo details matter. The memory is sliced to the last few lessons, because the paper bounded it at one to three, and long lists of lessons dilute attention and contradict each other. And the loop returns a failure explicitly when it runs out of trials, rather than presenting its last attempt as an answer.
The evaluator sets the ceiling
Everything downstream reasons from the evaluator's verdict, so its quality bounds the whole system. Evaluators, roughly from most to least trustworthy:
- Ground-truth checks: a test suite written by people, a compiler, a schema validator, an environment's success flag. Cheap, deterministic, hard to fool.
- Self-generated tests: useful when no tests exist, but the MBPP result shows the risk. A model that misreads the requirement writes tests that encode the same misreading.
- Heuristics: in AlfWorld the paper triggered reflection when the agent repeated the same action and got the same response for more than three cycles, or took more than 30 actions. Crude, but grounded in observable behaviour.
- LLM judges: flexible and the least reliable; they share blind spots with the actor. Use a rubric, a different model where possible, and spot-check against human labels, as discussed in verifier prompting.
The single most useful rule: split your checks into visible tests the loop can use, and held-out tests it never sees and that decide acceptance. Then a loop that games its own evaluator is caught before its output ships.
import subprocess, tempfile, pathlib
def evaluate(code: str, visible_tests: str) -> tuple[bool, str]:
with tempfile.TemporaryDirectory() as d:
path = pathlib.Path(d)
(path / "solution.py").write_text(code)
(path / "test_visible.py").write_text(visible_tests)
r = subprocess.run(["python", "-m", "pytest", "-x", "-q", "test_visible.py"],
cwd=d, capture_output=True, text=True, timeout=30)
return r.returncode == 0, r.stdout[-2000:] # keep feedback short and specific
# Bind the tests for the loop: Reflexion(llm, functools.partial(evaluate, visible_tests=t))
# Accept the final answer only against tests the loop never saw.
Writing reflections that change behaviour
A reflection is only useful if the next attempt can act on it. 'I should be more careful with edge cases' changes nothing. 'Test test_empty_string failed because parse_duration indexes s[-1] before checking length; check for an empty string first and raise ValueError' changes the next attempt. The reflection prompt should force that shape: name the failed requirement with evidence, name the root cause in the code, prescribe a concrete change, and review earlier lessons.
def actor_prompt(task, lessons):
notes = "\n".join(f"- {l}" for l in lessons) or "- none yet"
return (f"{task}\n\nLessons from your previous failed attempts on THIS task:\n{notes}\n\n"
"Apply every lesson. Return only the code.")
def reflect_prompt(task, attempt, feedback, lessons):
return (f"Task:\n{task}\n\nYour attempt:\n{attempt}\n\nEvaluator output:\n{feedback}\n\n"
f"Earlier lessons:\n{lessons}\n\n"
"In at most 4 sentences: (1) state which requirement failed, quoting the evidence; "
"(2) name the root cause in the code, not the symptom; (3) give a concrete change for "
"the next attempt; (4) if an earlier lesson was wrong or already applied, say so. "
"Do not write code.")Keep reflections short, about four sentences, and ban code in them, so they work as instructions rather than a second draft the actor copies with its bugs. Feed the reflector the actual evaluator output (test names, assertion messages, the last lines of a traceback), trimmed to what is relevant. Long raw logs bury the signal.
Memory policy
The paper's memory holds lessons for one task across its trials, and it is cleared when the task ends. That is the safe default. Carrying lessons across tasks sounds attractive but changes the method: a lesson specific to one function ('the input uses comma separators') becomes a wrong instruction on the next. If you want cross-task learning, distil lessons into general rules, review them, and put them in the system prompt deliberately, not automatically.
Within a task, store the lesson text and the trial number, keep the last one to three, and allow the reflector to retract an earlier lesson. Without retraction, a wrong early diagnosis keeps steering every later attempt.
Worked example: a duration parser
Task: write parse_duration(s) that turns strings like '1h30m' or '45s' into seconds, and raises ValueError on invalid input. Visible tests: six cases. Held-out tests: twelve.
Trial 1 uses a regular expression for number-unit pairs and passes four of six visible tests. It fails '90' (no unit, which the spec says means seconds) and '' (empty string, which returns 0 instead of raising). Reflection: 'test_bare_number failed: the spec says a bare integer means seconds but the regex requires a unit. test_empty failed: findall on an empty string returns no matches and the sum is 0. Accept a bare integer as seconds; raise ValueError when the string is empty or has unconsumed characters.' Trial 2 applies both lessons and passes all six visible tests.
The held-out tests then fail one case: '1h1h' is accepted and returns 7,200, while the spec forbids repeated units. Nothing in the visible tests covered it, so the loop correctly stopped at trial 2 by its own evaluator and was incorrectly confident. This is the MBPP effect on a small scale, and it is why acceptance has to use tests the loop never saw. Feeding the held-out failure back as one more trial fixes it, but now those tests are visible, and you need new held-out tests to measure honestly.
Cost, latency and stopping
Each failed trial costs an actor call, an evaluation and a reflector call. With a 4 trial cap, the worst case is 4 actor calls and 3 reflector calls, about 7 times the cost of a single attempt, plus evaluation time. Most tasks finish in the first trial, so measure the distribution: if 80% pass on trial 1, the mean cost is modest and the tail is what you budget for. Stop on success, on the trial cap, on a cost or time budget, and when two consecutive reflections say essentially the same thing, which signals the loop is stuck in a local minimum. A cap of 3 to 5 trials is a reasonable starting point; measure where the pass rate flattens on your own tasks.
Failure modes
- Evaluator false positives: the loop stops early on wrong answers. Use held-out acceptance tests.
- Vague reflections: generic advice that does not change the next attempt. Enforce the four-part shape.
- Lesson pile-up: long memories with contradictory lessons. Bound and allow retraction.
- Test tampering: an agent with file access edits or deletes the tests to pass. Keep evaluation outside the agent's sandbox.
- Repeating the same fix: the actor ignores the lesson. Make the lesson concrete and put it last in the prompt.
- Exploration tasks: as with WebShop, when success needs a new strategy rather than fixing a mistake, reflection loops circle.
When to use it and when not to
Use Reflexion when you have a cheap, trustworthy evaluator and failures tend to be specific and fixable: code against tests, structured output against a validator, agent tasks with a clear success signal, as in tool loops built on ReAct. Skip it when the only evaluator is the same model's opinion, when latency matters more than a few points of accuracy, or when tasks fail because the model lacks knowledge, which no amount of reflecting supplies.
What to do next
- Pick one task family with an executable evaluator and measure single-attempt pass rate first.
- Split tests into visible ones for the loop and held-out ones for acceptance.
- Implement the loop above with a 3 to 5 trial cap and a memory of 3 lessons.
- Use the four-part reflection prompt, and log every reflection next to the evaluator output.
- Compare against plain retries with the same budget; keep Reflexion only if it wins on held-out tests.
- Track false positives: answers your evaluator accepted that held-out tests rejected.
- Add stopping rules for cost, time and repeated reflections before putting it in production.