A prompt that worked last week now returns the wrong field, ignores an instruction, or produces valid output nine times and nonsense the tenth. The instinctive response is to open the prompt, add a sentence in capitals, try it once, see a good answer and ship. That loop feels productive and usually makes things worse: the fix was tested on one sample of a random process, the extra sentence conflicts with something else, and the original cause, often not in the wording at all, is still there.
This article treats prompt debugging the way good engineers treat a flaky test. You capture exactly what the model received, measure how often it fails, shrink the failing case, identify the cause from a known catalogue, change one thing, check the change across a set of cases, and record the case so it cannot come back. The parts of a prompt are covered in the anatomy of a prompt and building evaluation sets in prompt evals; this page is about the debugging loop that uses both.
First classify the symptom
"The prompt is broken" covers very different problems, and the symptom points to where to look. Write down which one you have before touching anything.
| Symptom | Most likely places to look |
|---|---|
| Output cut off or missing its end | max_tokens, stop sequences, a streaming consumer that closes early |
| Wrong format or invalid JSON | Format instructions, conflicting examples, parser strictness, missing structured-output mode |
| Ignores one instruction | Conflicting instructions, instruction buried in long context, authority of the turn it is in |
| Wrong facts or invented values | Missing or truncated context, retrieval returning the wrong documents, question the context cannot answer |
| Inconsistent across runs | Sampling settings, ambiguous task, borderline inputs |
| Worked before, fails now | Model version, template or code change, upstream data change |
| Works in the playground, fails in production | Different rendered request: variables, system prompt, parameters or tools |
Step 1: capture the exact request
You cannot debug a template. You can only debug the request that was actually sent: the full list of messages after every variable was substituted, the system prompt, the tool definitions, the model identifier and every sampling parameter. Many prompt bugs live in the gap between the template you are reading and that rendered request: a variable that rendered as an empty string or as None, retrieved context truncated mid-sentence by a token budget, an HTML-escaped document, or a production system prompt that the playground never had.
Log the rendered request and the raw response for every call, or for a sample with full capture on failures, keyed by a request ID and the prompt version. Then the first debugging step is a lookup, not a reconstruction. A prompt registry that versions templates makes the version part of that key cheap.
def call_model(client, prompt_version, messages, **params):
request = {"model": params["model"], "messages": messages, **params}
response = client.complete(**request) # your provider SDK call here
log.info("llm_call", extra={
"request_id": response.id, "prompt_version": prompt_version,
"request": request, "raw_output": response.text,
"finish_reason": response.finish_reason, # truncation shows up here
})
return responseAlways record the finish or stop reason. A response that ended because it hit the token limit is a parameter bug, and no amount of rewording will fix it.
Step 2: reproduce it as a rate, not an anecdote
Model output is sampled, so one success after a change proves very little. Replay the captured request many times at production settings and count failures. If it fails 3 times in 20, the bug is real and your fix must bring that number down on the same 20-run check, not produce one good answer. If it fails 0 times in 20, either the failure depended on something you did not capture, or it is rare and you need more runs or more failing examples.
def failure_rate(client, request, check, runs=20):
failures = []
for i in range(runs):
out = client.complete(**request).text
ok, reason = check(out) # deterministic check, not vibes
if not ok:
failures.append((i, reason, out[:200]))
return len(failures) / runs, failures
def check_invoice(out, expected_total): # bind with functools.partial
try:
data = json.loads(out)
except json.JSONDecodeError as e:
return False, f"invalid json: {e}"
if abs(data["total"] - expected_total) > 0.005:
return False, f"total {data['total']} != {expected_total}"
return True, ""Write the check as code whenever possible. A check that parses and compares is fast, repeatable and honest; reading outputs by eye invites seeing what you hope for. Lowering temperature can make a bug easier to reproduce, but confirm the fix at the production temperature too.
Step 3: minimise the failing case
A real prompt might hold a system message, ten instructions, five examples, retrieved documents and the user input. Shrink it until the failure still happens with as little as possible, the same way delta debugging shrinks a crashing input. Remove one block at a time and re-measure; keep each removal that leaves the failure rate unchanged. Whatever is left is either the cause or the context it needs.
def minimise(blocks, render, measure, threshold):
# blocks: ordered list of prompt parts; measure returns a failure rate
keep = list(blocks)
changed = True
while changed:
changed = False
for b in list(keep):
trial = [x for x in keep if x is not b]
if measure(render(trial)) >= threshold: # still fails without b
keep = trial
changed = True
return keep # smallest set that still failsMinimisation is expensive in calls, so run it on one or two captured failures, and cap the runs per trial. The flip side is also informative: if removing the examples makes the failure disappear, the examples are part of the cause.
Step 4: classify the root cause
Most prompt bugs fall into a small number of categories. Checking the minimised case against this list is faster than inventing a theory from scratch.
- Conflicting instructions. "Be concise" in the system prompt and "explain every field" in the task. The model satisfies one, unpredictably. Search the rendered request for pairs that cannot both hold.
- Examples that teach the wrong thing. Few-shot examples outweigh instructions. If every example has three items, the model returns three items. If an example contains a mistake, it gets copied.
- Buried or distant instructions. An instruction placed before 30,000 tokens of documents is followed less reliably than one placed near the question. Restate key constraints near the end of long inputs.
- Ambiguous terms. "Recent", "short", "the customer's address" when there are three addresses. The model is resolving an ambiguity you did not notice; define the term.
- Template and data bugs. Empty variables, wrong document retrieved, a document cut by the token budget, escaped markup, a date in an unexpected format.
- Format pressure against the task. Demanding only a JSON object with a final answer removes room to reason on hard inputs. Allow a reasoning field or a separate step, or use the provider's structured-output mode.
- Parameters. Token limits that truncate, stop sequences that appear in valid output, temperature too high for an extraction task.
- Model change. A new model version follows instructions more literally or formats differently. Pin model versions and re-run your eval set before switching.
Asking the model why it did something can suggest hypotheses, but its explanation is generated text, not a trace of its computation, so treat it as a lead to test, never as the diagnosis.
Step 5: make one change, the smallest that addresses the cause
Fix the cause you identified, and only that. If the cause is a template bug, fix the code and leave the wording alone. If two instructions conflict, remove or reconcile one rather than adding a third that shouts. If the examples mislead, vary them. Prefer stating what to do over what not to do, define ambiguous terms once, and move hard constraints out of prose into code where you can: a schema-validated output with a retry on failure is more reliable than a sentence asking for valid JSON. Output parsing covers that pattern.
Changing several things at once is the most common mistake in prompt debugging. When the failure rate improves you will not know which change did it, and the useless ones stay in the prompt forever, making the next bug harder to find.
Step 6: verify on the case and on the set
Re-run the failure-rate check on the captured case; it should drop to zero or near it over the same number of runs. Then run the full evaluation set, because a fix for one input regularly breaks another: a new instruction that stops one hallucination can make the model refuse valid questions. Compare per-case results with the previous version, not just the average, and look at every case that flipped from pass to fail. If the set is small, the difference between versions may be noise, so treat a one-case change as inconclusive.
Step 7: lock it in
Add the captured failing input, with its expected output or check, to the evaluation set as a regression case. Commit the prompt change as a new version with a message that names the cause, such as "examples all had 3 line items; varied to 1-6", not "improved prompt". The next person to touch the prompt then knows why each part exists and will see an eval failure if they undo it.
Worked example: the invoice total that was sometimes wrong
An extraction prompt reads invoice text and returns JSON with vendor, date, line items and total. Support reports that totals are occasionally wrong. Capturing a failing request shows nothing odd in the template. Replaying it 20 times gives 4 failures, all with the total equal to the subtotal before tax. Minimising removes the vendor and date instructions and four of the five examples with the failure rate unchanged; removing the last example drops it to 0 in 20.
That example came from an invoice with no tax, so its subtotal and total were the same number, and its output showed total next to a value labelled Subtotal in the source text. The model had learned to read the subtotal line. Classification: an example teaching the wrong thing, compounded by an ambiguous term, since the instructions said "the total" without saying which line. The single change replaces that example with one where tax is present and the total is the amount due including tax, and defines total as "the final amount payable, including tax and after discounts". The case goes to 0 failures in 20; the 150-case eval set shows no new failures; the invoice becomes regression case 151 and the commit message names the cause.
When debugging itself goes wrong
- Overfitting to one input. The prompt grows special cases for individual customers and breaks for everyone else. Fix classes of failures, verified on a set.
- Prompt bloat. Every incident adds a sentence and none are removed. Periodically ablate parts and delete those that do not change results.
- Judging by eye. Reading five outputs and declaring victory. Use code checks and failure rates.
- Debugging the wrong layer. Rewording to fix what is a retrieval, parameter or parsing bug.
- Changing the model mid-investigation. Hold the model version fixed until the cause is known, or you are debugging two variables.
What to do next
- Log the rendered request, parameters, model version, prompt version, raw output and finish reason for every call, or every failed one.
- Write a replay script that runs a captured request N times with a code check and reports the failure rate.
- Build or extend an evaluation set, and add every debugged failure to it as a regression case.
- Keep the root-cause catalogue above next to your prompts, and classify before you edit.
- Make one change per iteration, commit it as a prompt version with the cause in the message, and compare per-case results with the previous version.
- Move hard constraints, such as schemas, required fields and value ranges, into validation code with retries.
- Pin model versions and re-run the full set before any model upgrade.