Ask a strong writer to improve a draft and they will reread it, notice what is weak and fix it. Self-Refine asks a language model to do the same thing: produce an answer, critique it against explicit criteria, rewrite it using the critique, and repeat. No fine-tuning, no reward model and no second model are needed; the same model plays all three roles through three prompts.

The technique comes from the 2023 paper Self-Refine: Iterative Refinement with Self-Feedback by Madaan and colleagues, published at NeurIPS 2023. This page explains the loop from first principles, reports what the paper actually found (including where it did not help), shows how to design the feedback and refine prompts, gives a complete implementation with sensible stop conditions and closes with cost, failure modes and a checklist. It is distinct from Reflexion, which carries lessons across separate attempts at a task, and from chain of verification, which checks factual claims with independent questions.

Advertisement

The loop

Self-Refine: one model plays generator, critic and editor until a stop condition firesTask + rubricinput xGeneratedraft y0Feedbackscores + specific fixesStop?score, budget, no changeRefinex + history + feedbacknoy(t+1)Return besthighest score seenyesOptional external signalstests, linters, schema validators, retrievalThe paper used the same model in every role and up to four iterations; external signals are this page's production addition.
Generate once, then alternate feedback and refinement until the output is good enough or the budget runs out.

Self-Refine has three steps. Generate: given the task input x, the model produces an initial output y0, exactly as in ordinary prompting. Feedback: the model is shown x and its own output and asked for feedback, ideally structured as scores against named criteria plus concrete, localized suggestions. Refine: the model is shown x, the output, the feedback and the history of earlier attempts and feedback, and asked to produce an improved output. Feedback and refine alternate until a stop condition fires.

Two design details from the paper matter in practice. The refine step sees the history, not only the latest draft, so the model can avoid reintroducing problems it already fixed. And the stop condition can be either a fixed number of iterations or a stop indicator extracted from the feedback, such as a score above a threshold. Each of the three steps is a few-shot prompt specific to the task: the feedback prompt shows examples of good critiques for that task, not a generic request to find problems.

What the paper found

The authors evaluated Self-Refine on seven tasks, from dialogue response generation and code optimization to mathematical reasoning, using GPT-3.5, ChatGPT and GPT-4, with up to four refinement iterations. Across tasks, outputs after refinement were preferred by humans and automatic metrics over single-pass outputs from the same model, with an average improvement of about 20 percentage points absolute. On code optimization, for example, GPT-4's rate of producing faster programs rose from 27.3% to 36.0%.

The details are more instructive than the average. Gains were largest on open-ended generation, where quality has many dimensions a model can judge: tone, engagement, constraint coverage, readability. Gains on math reasoning were minimal. The paper explains why: a consistent-looking reasoning chain convinced the model that everything looked good, and ChatGPT's feedback said so in 94% of cases, so the refine step had nothing to act on. Feedback quality was decisive. In an ablation, replacing specific, actionable feedback with generic feedback dropped sentiment reversal performance from 43.2 to 31.2.

Later work sharpened the reasoning caveat. Huang and colleagues, in Large Language Models Cannot Self-Correct Reasoning Yet (2023), found that when a model critiques its own reasoning answers without any external signal, accuracy often does not improve and can fall, because the model changes correct answers as readily as wrong ones. The practical reading is consistent with the original paper: self-feedback helps where judging is easier than producing and where criteria can be stated; it is unreliable where the model cannot detect its own errors.

Advertisement

Why it works when it works

The intuition is the gap between generation and verification. Writing a concise, polite, accurate support reply that covers four required points in one pass is hard; checking a finished reply against those four points is easier. A feedback prompt turns one hard problem into a sequence of easier ones: find what is missing, then fix only that.

Self-Refine also forces the criteria to be written down. Many prompts fail because the model is never told what good looks like. A rubric in the feedback prompt makes the target explicit, and the refine step gets a short list of concrete edits instead of a vague instruction to be better. In many deployments, writing the rubric is half the improvement on its own.

It fails when verification is not easier than generation for the model: multi-step arithmetic, factual recall the model does not have, subtle logic errors. The same blind spot that produced the error also reviews it. In those cases, bring in a signal from outside the model, which is covered below, or use a different technique such as self-consistency voting.

Designing the feedback prompt

The feedback prompt does most of the work, so design it as carefully as the task prompt. Four properties separate useful feedback from noise:

  • Criteria-based. Name each dimension you care about and score it on a small fixed scale, for example 1 to 5, with a one-line definition of what a 5 means.
  • Specific and localized. Each problem should quote or point to the part of the output it refers to and say what to change. Generic feedback was the weakest condition in the paper's ablation.
  • Machine-readable. Ask for JSON so your code can compute a stop decision and log scores, rather than parsing prose.
  • Allowed to say nothing is wrong. Without an explicit way to report no remaining problems, models invent minor issues to satisfy the request, and the loop never stops improving in circles.
FEEDBACK_PROMPT = """You are reviewing a draft answer. Do not rewrite it.
Task: {task}
Draft:
<<<{draft}>>>

Score each criterion from 1 (poor) to 5 (meets the definition fully):
- coverage: every required point in the task is addressed
- accuracy: no claim contradicts the provided facts
- clarity: a non-expert can follow it on first read
- length: within the stated word limit

For every score below 5, give one issue that quotes the exact text it refers to
and states the concrete change needed. If nothing needs changing, return an
empty issues list.

Return only JSON:
{{"scores": {{"coverage": n, "accuracy": n, "clarity": n, "length": n}},
  "issues": [{{"criterion": "...", "quote": "...", "fix": "..."}}]}}"""

Designing the refine prompt

The refine prompt should make the model an editor, not a new author. Give it the original task (so it does not drift from the requirements), the current draft, the feedback and a short history of earlier drafts and their issues. Then constrain the edit: address each listed issue, keep everything that was not criticized, and do not add new content beyond what the fixes need. Without that constraint, refinement tends to rewrite from scratch, which can fix one problem while breaking two others, and tends to make outputs longer each round.

REFINE_PROMPT = """Improve the draft by fixing exactly the issues listed.
Task: {task}
Previous attempts and their issues (oldest first):
{history}
Current draft:
<<<{draft}>>>
Issues to fix:
{issues}

Rules: fix every listed issue; keep all text that no issue mentions; do not
add new claims; stay within the task's word limit. Return only the new draft."""

A complete loop with stop conditions

The implementation below is provider-agnostic: llm(prompt) is any function that returns text. It keeps the best-scoring draft rather than the last one, because refinement is not monotonic, and it stops on any of four conditions: the score threshold is met, the feedback has no issues, the draft stopped changing, or the iteration budget is spent.

import json

CRITERIA = ["coverage", "accuracy", "clarity", "length"]

def total(scores):
    return sum(scores.get(k, 0) for k in CRITERIA)

def get_feedback(llm, task, draft, retries=2):
    for _ in range(retries + 1):
        raw = llm(FEEDBACK_PROMPT.format(task=task, draft=draft))
        try:
            fb = json.loads(raw)
            if all(k in fb["scores"] for k in CRITERIA):
                return fb
        except (json.JSONDecodeError, KeyError, TypeError):
            pass
    raise ValueError("feedback was not valid JSON")

def self_refine(llm, task, max_iters=3, threshold=19):
    draft = llm(task)                           # step 1: generate
    history = []
    best = (None, -1)
    for it in range(max_iters + 1):
        fb = get_feedback(llm, task, draft)     # step 2: feedback
        score = total(fb["scores"])
        if score > best[1]:
            best = (draft, score)
        if score >= threshold or not fb["issues"] or it == max_iters:
            break
        issues = "\n".join(f"- [{i['criterion']}] \"{i['quote']}\": {i['fix']}"
                           for i in fb["issues"])
        new = llm(REFINE_PROMPT.format(        # step 3: refine
            task=task, draft=draft, issues=issues,
            history="\n\n".join(history[-2:]) or "(none)"))
        history.append(f"Draft:\n{draft}\nIssues:\n{issues}")
        if new.strip() == draft.strip():        # no progress
            break
        draft = new
    return best   # (draft, score)

Two choices in this code are deliberate. The history is truncated to the last two rounds to keep the prompt from growing without bound. And the threshold of 19 out of 20 allows one criterion to sit at 4, because demanding a perfect score from a self-grader mostly produces extra iterations, not better output. Tune both against a test set.

Worked example: an incident summary for customers

Task: write a customer-facing summary of a 40-minute API outage, under 120 words, covering what happened, impact, the fix and what changes next, using only the facts in an internal timeline.

Draft y0 is fluent but mentions the internal name of the failed component, omits the time window and promises a fix date that is not in the timeline. Feedback 1 scores coverage 3 (no time window), accuracy 2 (quotes the invented date, says to remove it), clarity 4 (quotes the internal component name, says to replace it with a plain description) and length 5: total 14, three issues.

Refine 1 adds the time window, removes the date and replaces the jargon, leaving the rest untouched. Feedback 2 scores 5, 5, 5 and 4: the draft is now 131 words, so length drops to 4 with one issue. Total 19 meets the threshold, so the loop stops and returns this draft. A human editor would trim it, which suggests checking length with code instead of the model; that is the external-signal idea in the next section.

Cost: one generate call, two feedback calls and one refine call, about four times the single-pass call count, for a draft that fixed one factual error and two coverage problems. Whether that is worth it depends on how often single-pass drafts need human correction.

Adding external signals

The paper's version is pure self-feedback. In production, the strongest variant replaces or augments the model's judgement with signals it cannot fool itself about: unit tests and a linter for code, a JSON schema validator for structured output, a word counter for length limits, a retrieval check that every claim appears in the source, or a separate verifier model. Feed the failing checks into the refine prompt as issues, exactly like model-written feedback. This also addresses the reasoning weakness: a failing test is feedback the model cannot dismiss as fine.

TechniqueFeedback sourceBest for
Self-RefineSame model, rubricOpen-ended writing, style, constraint coverage
Self-Refine with toolsTests, validators, countersCode, structured output, hard constraints
Self-consistencyAgreement across samplesReasoning with a single checkable answer
Chain of verificationIndependent verification questionsFactual claims in long answers
ReflexionOutcome of a full attempt, stored as memoryAgents retrying multi-step tasks

Cost, latency and when not to use it

Each iteration adds two model calls, and prompts grow because each includes the draft and feedback. The loop above with three iterations makes up to eight calls (one generate, four feedback, three refine) and has several times the latency of one, which rules it out for most interactive paths with tight latency budgets. It fits offline and asynchronous work well: content generation pipelines, report drafting, code review suggestions, data labelling with quality criteria.

Before adopting it, compare against the honest baseline: the same token budget spent another way. Sampling several drafts and choosing the best by the same rubric, or simply improving the task prompt with the rubric, sometimes matches Self-Refine at lower cost. Using a smaller, cheaper model for the feedback step is tempting but risky, because feedback quality was the strongest driver of gains in the paper. Test it rather than assuming. If the work splits naturally into stages, prompt chaining may be the better structure.

Failure modes

  • Rubber-stamp feedback: the critic says everything looks good, especially on reasoning; add external checks or do not use the technique for that task.
  • Invented issues: forced to find problems, the critic nitpicks and the loop churns; allow an empty issues list and stop on it.
  • Oscillation: two drafts alternate as each fixes what the other broke; include history and keep the best-scoring draft.
  • Requirement drift: refinement optimizes the rubric and forgets the task; always include the original task in the refine prompt.
  • Length inflation: every round adds caveats; measure length in code and make it a criterion.
  • Self-preference: the model rates its own style highly; calibrate scores against human ratings on a sample.
  • Unparseable feedback: JSON breaks and the loop crashes; validate and retry, then fall back to the current best draft.

What to do next

  1. Pick one offline generation task where humans currently correct drafts, and write its rubric with three to five criteria and a definition of a top score for each.
  2. Build a test set of 30 to 50 inputs with human ratings, and measure single-pass quality first.
  3. Implement the loop above with JSON feedback, history, best-so-far tracking and four stop conditions.
  4. Compare against single pass and against best-of-n sampling at the same token budget.
  5. Replace every criterion that code can check (length, schema, tests) with a code check.
  6. Log scores per iteration in production and cap iterations where the score curve flattens.
Key takeaway: Self-Refine lets one model generate, critique against explicit criteria and rewrite in a loop. The 2023 paper found large gains on open-ended tasks and almost none on math, where the model judged flawed reasoning as fine, and showed that specific feedback matters far more than generic feedback. Use a structured rubric, give the editor the task and history, keep the best draft, stop on score, no issues, no change or budget, replace self-judgement with code checks wherever possible, and compare against other uses of the same token budget.