A prompt that worked last week now returns the wrong field, ignores an instruction, or produces valid output nine times and nonsense the tenth. The instinctive response is to open the prompt, add a sentence in capitals, try it once, see a good answer and ship. That loop feels productive and usually makes things worse: the fix was tested on one sample of a random process, the extra sentence conflicts with something else, and the original cause, often not in the wording at all, is still there.

This article treats prompt debugging the way good engineers treat a flaky test. You capture exactly what the model received, measure how often it fails, shrink the failing case, identify the cause from a known catalogue, change one thing, check the change across a set of cases, and record the case so it cannot come back. The parts of a prompt are covered in the anatomy of a prompt and building evaluation sets in prompt evals; this page is about the debugging loop that uses both.

Advertisement

First classify the symptom

"The prompt is broken" covers very different problems, and the symptom points to where to look. Write down which one you have before touching anything.

SymptomMost likely places to look
Output cut off or missing its endmax_tokens, stop sequences, a streaming consumer that closes early
Wrong format or invalid JSONFormat instructions, conflicting examples, parser strictness, missing structured-output mode
Ignores one instructionConflicting instructions, instruction buried in long context, authority of the turn it is in
Wrong facts or invented valuesMissing or truncated context, retrieval returning the wrong documents, question the context cannot answer
Inconsistent across runsSampling settings, ambiguous task, borderline inputs
Worked before, fails nowModel version, template or code change, upstream data change
Works in the playground, fails in productionDifferent rendered request: variables, system prompt, parameters or tools

Step 1: capture the exact request

You cannot debug a template. You can only debug the request that was actually sent: the full list of messages after every variable was substituted, the system prompt, the tool definitions, the model identifier and every sampling parameter. Many prompt bugs live in the gap between the template you are reading and that rendered request: a variable that rendered as an empty string or as None, retrieved context truncated mid-sentence by a token budget, an HTML-escaped document, or a production system prompt that the playground never had.

Log the rendered request and the raw response for every call, or for a sample with full capture on failures, keyed by a request ID and the prompt version. Then the first debugging step is a lookup, not a reconstruction. A prompt registry that versions templates makes the version part of that key cheap.

def call_model(client, prompt_version, messages, **params):
    request = {"model": params["model"], "messages": messages, **params}
    response = client.complete(**request)            # your provider SDK call here
    log.info("llm_call", extra={
        "request_id": response.id, "prompt_version": prompt_version,
        "request": request, "raw_output": response.text,
        "finish_reason": response.finish_reason,     # truncation shows up here
    })
    return response

Always record the finish or stop reason. A response that ended because it hit the token limit is a parameter bug, and no amount of rewording will fix it.

Advertisement

Step 2: reproduce it as a rate, not an anecdote

Model output is sampled, so one success after a change proves very little. Replay the captured request many times at production settings and count failures. If it fails 3 times in 20, the bug is real and your fix must bring that number down on the same 20-run check, not produce one good answer. If it fails 0 times in 20, either the failure depended on something you did not capture, or it is rare and you need more runs or more failing examples.

def failure_rate(client, request, check, runs=20):
    failures = []
    for i in range(runs):
        out = client.complete(**request).text
        ok, reason = check(out)                      # deterministic check, not vibes
        if not ok:
            failures.append((i, reason, out[:200]))
    return len(failures) / runs, failures

def check_invoice(out, expected_total):           # bind with functools.partial
    try:
        data = json.loads(out)
    except json.JSONDecodeError as e:
        return False, f"invalid json: {e}"
    if abs(data["total"] - expected_total) > 0.005:
        return False, f"total {data['total']} != {expected_total}"
    return True, ""

Write the check as code whenever possible. A check that parses and compares is fast, repeatable and honest; reading outputs by eye invites seeing what you hope for. Lowering temperature can make a bug easier to reproduce, but confirm the fix at the production temperature too.

Step 3: minimise the failing case

A real prompt might hold a system message, ten instructions, five examples, retrieved documents and the user input. Shrink it until the failure still happens with as little as possible, the same way delta debugging shrinks a crashing input. Remove one block at a time and re-measure; keep each removal that leaves the failure rate unchanged. Whatever is left is either the cause or the context it needs.

def minimise(blocks, render, measure, threshold):
    # blocks: ordered list of prompt parts; measure returns a failure rate
    keep = list(blocks)
    changed = True
    while changed:
        changed = False
        for b in list(keep):
            trial = [x for x in keep if x is not b]
            if measure(render(trial)) >= threshold:  # still fails without b
                keep = trial
                changed = True
    return keep                                      # smallest set that still fails

Minimisation is expensive in calls, so run it on one or two captured failures, and cap the runs per trial. The flip side is also informative: if removing the examples makes the failure disappear, the examples are part of the cause.

Step 4: classify the root cause

Most prompt bugs fall into a small number of categories. Checking the minimised case against this list is faster than inventing a theory from scratch.

  • Conflicting instructions. "Be concise" in the system prompt and "explain every field" in the task. The model satisfies one, unpredictably. Search the rendered request for pairs that cannot both hold.
  • Examples that teach the wrong thing. Few-shot examples outweigh instructions. If every example has three items, the model returns three items. If an example contains a mistake, it gets copied.
  • Buried or distant instructions. An instruction placed before 30,000 tokens of documents is followed less reliably than one placed near the question. Restate key constraints near the end of long inputs.
  • Ambiguous terms. "Recent", "short", "the customer's address" when there are three addresses. The model is resolving an ambiguity you did not notice; define the term.
  • Template and data bugs. Empty variables, wrong document retrieved, a document cut by the token budget, escaped markup, a date in an unexpected format.
  • Format pressure against the task. Demanding only a JSON object with a final answer removes room to reason on hard inputs. Allow a reasoning field or a separate step, or use the provider's structured-output mode.
  • Parameters. Token limits that truncate, stop sequences that appear in valid output, temperature too high for an extraction task.
  • Model change. A new model version follows instructions more literally or formats differently. Pin model versions and re-run your eval set before switching.

Asking the model why it did something can suggest hypotheses, but its explanation is generated text, not a trace of its computation, so treat it as a lead to test, never as the diagnosis.

Step 5: make one change, the smallest that addresses the cause

Fix the cause you identified, and only that. If the cause is a template bug, fix the code and leave the wording alone. If two instructions conflict, remove or reconcile one rather than adding a third that shouts. If the examples mislead, vary them. Prefer stating what to do over what not to do, define ambiguous terms once, and move hard constraints out of prose into code where you can: a schema-validated output with a retry on failure is more reliable than a sentence asking for valid JSON. Output parsing covers that pattern.

Changing several things at once is the most common mistake in prompt debugging. When the failure rate improves you will not know which change did it, and the useless ones stay in the prompt forever, making the next bug harder to find.

Step 6: verify on the case and on the set

Re-run the failure-rate check on the captured case; it should drop to zero or near it over the same number of runs. Then run the full evaluation set, because a fix for one input regularly breaks another: a new instruction that stops one hallucination can make the model refuse valid questions. Compare per-case results with the previous version, not just the average, and look at every case that flipped from pass to fail. If the set is small, the difference between versions may be noise, so treat a one-case change as inconclusive.

Step 7: lock it in

Add the captured failing input, with its expected output or check, to the evaluation set as a regression case. Commit the prompt change as a new version with a message that names the cause, such as "examples all had 3 line items; varied to 1-6", not "improved prompt". The next person to touch the prompt then knows why each part exists and will see an eval failure if they undo it.

Worked example: the invoice total that was sometimes wrong

An extraction prompt reads invoice text and returns JSON with vendor, date, line items and total. Support reports that totals are occasionally wrong. Capturing a failing request shows nothing odd in the template. Replaying it 20 times gives 4 failures, all with the total equal to the subtotal before tax. Minimising removes the vendor and date instructions and four of the five examples with the failure rate unchanged; removing the last example drops it to 0 in 20.

That example came from an invoice with no tax, so its subtotal and total were the same number, and its output showed total next to a value labelled Subtotal in the source text. The model had learned to read the subtotal line. Classification: an example teaching the wrong thing, compounded by an ambiguous term, since the instructions said "the total" without saying which line. The single change replaces that example with one where tax is present and the total is the amount due including tax, and defines total as "the final amount payable, including tax and after discounts". The case goes to 0 failures in 20; the 150-case eval set shows no new failures; the invoice becomes regression case 151 and the commit message names the cause.

When debugging itself goes wrong

  • Overfitting to one input. The prompt grows special cases for individual customers and breaks for everyone else. Fix classes of failures, verified on a set.
  • Prompt bloat. Every incident adds a sentence and none are removed. Periodically ablate parts and delete those that do not change results.
  • Judging by eye. Reading five outputs and declaring victory. Use code checks and failure rates.
  • Debugging the wrong layer. Rewording to fix what is a retrieval, parameter or parsing bug.
  • Changing the model mid-investigation. Hold the model version fixed until the cause is known, or you are debugging two variables.

What to do next

  1. Log the rendered request, parameters, model version, prompt version, raw output and finish reason for every call, or every failed one.
  2. Write a replay script that runs a captured request N times with a code check and reports the failure rate.
  3. Build or extend an evaluation set, and add every debugged failure to it as a regression case.
  4. Keep the root-cause catalogue above next to your prompts, and classify before you edit.
  5. Make one change per iteration, commit it as a prompt version with the cause in the message, and compare per-case results with the previous version.
  6. Move hard constraints, such as schemas, required fields and value ranges, into validation code with retries.
  7. Pin model versions and re-run the full set before any model upgrade.
Key takeaway: Treat a broken prompt like a flaky test: capture the exact rendered request, measure how often it fails, shrink it to the smallest failing case, identify the cause from a known list of causes, change one thing, check that change across the whole evaluation set, and keep the case as a regression test. Many prompt bugs turn out to be in the template, the parameters, the examples or the data rather than the wording, and this loop finds them faster than rewriting.