Code is the most verifiable thing a language model produces. A paragraph of prose can only be judged by reading it; a function can be compiled, type-checked, linted and run against tests. That changes what good prompting means. For code, the prompt is half of a system whose other half is execution, and most of the quality comes from giving the model the right context, asking for output in a shape a machine can apply, and letting the build and tests decide whether it worked.
This page explains that system from first principles: what the model needs to see, how to state the task so the result fits the codebase, which output format to request, how to run a bounded repair loop, what goes wrong, and how to measure whether a prompt change helped. The examples use Python, but nothing depends on the language or on a particular model provider.
Why code prompts fail differently
Without context, a model writes plausible code for an imaginary codebase: helper functions that do not exist, a library you do not use, a style from its training data, a method borrowed from a popular framework. These are information failures, not reasoning failures. The model cannot know your conventions unless the prompt shows them.
Code also has a cheap, objective oracle: a wrong import fails on the first run. So the right design is not a perfect single prompt but a loop where the model proposes, the environment checks, and failures flow back as small, precise corrections.
The pipeline, end to end
Five stages: context assembly selects what the model sees; generation produces a plan and patch in a fixed format; application checks the patch applies; verification builds, lints and tests in a sandbox without secrets; review puts a person or policy between a green build and a merge. Separate stages fail separately, which keeps the system debuggable.
Assembling context from the repository
Show the model everything it needs to write code that fits, and little else; irrelevant files cost tokens and offer more names to borrow wrongly. A priority order, highest first:
- The file being changed, verbatim and complete if it fits. Partial files cause the model to re-declare things that already exist further down.
- Signatures of everything the new code will call: function and method signatures with docstrings, not full bodies. This is the single most effective defence against invented APIs.
- One or two callers of the code being written, so the model sees how it will be used.
- Existing tests for the module, which double as style examples and as executable documentation of behaviour.
- Conventions that are not visible in code: language version, allowed dependencies, error-handling style, logging rules.
- Similar code elsewhere in the repository, found by search or embeddings, as an example of the house style for this kind of change.
Gather signatures with tools rather than by hand: a language server, ctags, or parsing the abstract syntax tree to extract definitions referenced by the target file. Put each item in its own delimited block with its path, so the model can refer to files accurately and you can parse its references back. When context must be trimmed, drop item 6 first and item 1 last. The general mechanics of ordering long inputs are covered in long context prompting.
Tests as the specification
Natural-language task descriptions are ambiguous in exactly the places that matter for code: edge cases, error behaviour, input formats. Tests remove the ambiguity. The most reliable pattern is to write, or have someone write, a handful of failing tests that pin down the behaviour, then ask the model to make them pass without changing them.
The model gets exact inputs and outputs, verification becomes automatic, and edge cases are stated up front instead of found in review. If you cannot write the tests, the task is not yet specified well enough to delegate to anyone.
Here is a complete prompt for a small, realistic task, using tags as delimiters. Notice that it names the output contract explicitly and forbids touching the tests.
<task>
Add a function parse_duration(text: str) -> datetime.timedelta to src/timeutil.py.
It must accept forms like "90s", "15m", "2h30m", "1d" and raise ValueError on anything else.
</task>
<conventions>
- Python 3.11, standard library only in src/timeutil.py.
- Raise ValueError with a message that quotes the bad input.
- Type hints on every public function; docstring in Google style (see existing functions).
</conventions>
<file path="src/timeutil.py">
...current file contents, verbatim...
</file>
<related path="src/config.py" note="caller that will use parse_duration">
def load_timeouts(raw: dict[str, str]) -> dict[str, timedelta]: ...
</related>
<tests path="tests/test_timeutil.py">
...existing tests, plus the new failing tests below...
def test_parse_duration_compound():
assert parse_duration("2h30m") == timedelta(hours=2, minutes=30)
def test_parse_duration_rejects_empty():
with pytest.raises(ValueError):
parse_duration("")
</tests>
<output_contract>
First a <plan> of at most five bullet points.
Then exactly one unified diff in a <patch> block, against the files shown, with paths relative to the repo root.
Do not modify any file under tests/. Do not add dependencies. If the task is ambiguous, say so in the plan and choose the most conservative reading.
</output_contract>The plan-then-patch request is deliberate. A short plan before the code gives the model room to work out the approach, and gives a reviewer something to read in ten seconds. Capping it at five bullets stops it from becoming an essay. Asking the model to state ambiguity rather than resolve it silently surfaces the cases where the specification is incomplete.
Choosing the output contract
How you ask for code to be returned determines how reliably you can apply it.
| Format | Strengths | Weaknesses | Use when |
|---|---|---|---|
| Whole file | Easy to apply; no line-number drift | Costly for large files; model may silently drop or rewrite unrelated code | Small files, new files |
| Unified diff | Compact; changes are explicit and reviewable | Hunks fail to apply if context lines are wrong | Edits to existing, larger files |
| Search-and-replace blocks | Robust to line numbers; easy to validate the search text exists | Search text must be unique in the file | Targeted edits in large files |
| Function only | Smallest output | You must splice it in; imports are easy to forget | Single-function tasks in a harness |
Whatever you choose, validate before running anything. For diffs, git apply --check tells you whether the patch applies. For whole files, compare the set of top-level definitions before and after: if a function disappeared that the task did not mention, reject the output. For search-and-replace, require each search string to occur exactly once. Rejections go back to the model as short, specific messages. Formatting contracts in general are discussed in output formatting.
The generate, execute, repair loop
A bounded loop around the model turns a moderately reliable generator into a reliable system. The code below is the core of one. It refuses patches that edit protected files, checks that the patch applies, runs the tests, and on failure sends the tail of the test output back with an instruction to fix the implementation.
import subprocess, tempfile, re
MAX_ROUNDS = 3
def run(cmd, cwd):
r = subprocess.run(cmd, cwd=cwd, capture_output=True, text=True, timeout=600)
return r.returncode, (r.stdout + r.stderr)[-6000:] # keep the tail; it holds the failure
def touches_protected(patch: str) -> bool:
# both sides: "--- a/tests/x.py" + "+++ /dev/null" is a deleted test file
paths = re.findall(r"^(?:\+\+\+ b|--- a)/(\S+)", patch, flags=re.M)
return any(p.startswith("tests/") or p.endswith(("requirements.txt", "pyproject.toml")) for p in paths)
def generate_and_verify(llm, prompt: str, workdir: str):
history = [prompt]
for round_ in range(MAX_ROUNDS):
reply = llm(history) # your client call; returns text
patch = extract_block(reply, "patch") # text between <patch> tags, or None
if patch is None:
history.append("No <patch> block found. Reply with the plan and one unified diff.")
continue
if touches_protected(patch):
history.append("The patch edits tests or dependency files, which is not allowed. Fix the code instead.")
continue
with tempfile.NamedTemporaryFile("w", suffix=".diff", delete=False) as f:
f.write(patch)
code, out = run(["git", "apply", "--check", f.name], workdir)
if code != 0:
history.append(f"The patch does not apply:\n{out}\nRegenerate it against the files shown.")
continue
run(["git", "apply", f.name], workdir)
code, out = run(["python", "-m", "pytest", "-q", "-x"], workdir)
if code == 0:
return patch # still goes to review
run(["git", "checkout", "--", "."], workdir) # revert edits before the next attempt
run(["git", "clean", "-fd"], workdir) # and any files the patch created
history.append(f"Tests failed:\n{out}\nFix the implementation; do not change the tests.")
return None # escalate to a human with the historySeveral details matter more than they look. Trim the feedback. Test output can be thousands of lines; the last few thousand characters usually contain the assertion and traceback, and sending everything buries the signal. Revert between rounds so each attempt starts from a clean tree. Bound the rounds; if three attempts fail, a fourth rarely succeeds and the right move is to escalate with the transcript. Run in a sandbox: generated code is untrusted input, and tests execute it. No credentials, no network unless the tests need it, and resource limits on CPU, memory and time.
Sampling several candidates in parallel and keeping the first that passes is an alternative to sequential repair, and the two combine. Parallel sampling helps when failures are random; repair helps when the first attempt is close. Measure both on your tasks before paying for either. Separating a generator from an independent checker is a general pattern; see verifier architecture.
Failure modes and how to catch them
- Invented APIs. Calls to functions or parameters that do not exist. Prevention: include signatures. Detection: type checking and import resolution before tests, which fail faster and with clearer messages.
- Test tampering. When told to make tests pass, models sometimes weaken assertions, add skips, or special-case the test inputs. Forbid edits to tests in the contract, reject patches that touch test paths mechanically, and keep a few hidden tests the model never sees.
- Scope creep. Unrequested refactors, renamed variables, reformatted files. They make review slow and hide real changes. Ask for minimal changes, and flag diffs whose size is far above what the task implies.
- Plausible but wrong edge cases. Code that passes the given tests and fails on inputs nobody listed. Property-based tests and a reviewer's eye on boundaries catch these; the model will not volunteer them unless asked to list the edge cases it handled.
- Insecure defaults. String-built SQL, disabled certificate checks, shell commands built from input. Run a security linter in the verification stage, and say in the conventions block which patterns are banned.
- Stale context. The model edits a version of the file that is no longer current because context was cached. Hash the files you send and refuse patches if the working tree has changed since.
Measuring whether a prompt is better
Prompt changes for code should be judged on a fixed set of tasks with hidden tests, not on how a few outputs look. Build the set from your own repository: past bug fixes and small features, each with the commit's tests as the oracle, run against the parent commit. Thirty to a hundred tasks are enough to see large effects; small effects need more.
The standard metric is pass@k: the probability that at least one of k samples passes. Estimating it by drawing exactly k samples is noisy. Chen et al. (2021) introduced an unbiased estimator that draws n samples per task, counts the c that pass, and computes the probability that a random subset of size k contains at least one passing sample.
import numpy as np
def pass_at_k(n: int, c: int, k: int) -> float:
"""Unbiased pass@k (Chen et al., 2021): n samples drawn, c of them correct.
Equals 1 - C(n-c, k) / C(n, k), computed as a stable product."""
if n - c < k:
return 1.0
return 1.0 - float(np.prod(1.0 - k / np.arange(n - c + 1, n + 1)))
# Per task: generate n=10 candidates, count how many pass the hidden tests,
# then average pass_at_k(10, c, 1) over all tasks in the eval set.Worked example: for one task you draw 10 samples and 3 pass. pass@1 is 1 - 7/10 = 0.30. pass@5 is 1 - C(7,5)/C(10,5) = 1 - 21/252, about 0.92. Average per-task values across the set. Report pass@1 if production uses a single attempt, and the pass rate after your repair loop if it uses one, since that is what users experience. Also track the median number of repair rounds and the share of passing patches a reviewer rejects; a prompt that raises pass rates by encouraging test-specific hacks will show up there. The broader practice of running prompt evaluations is covered in prompt evaluation architecture.
Trade-offs
More context improves fit until relevance ranking matters more than volume. Repair loops raise pass rates but add latency and cost, and converge on hacks if tests are weak. Diffs are cheaper and more reviewable than whole files but need validation and a fallback. Stronger models need less scaffolding but still cannot know conventions that live only in your team's heads. Execution, not the model's confidence, decides what ships.
What to do next
- Pick one recurring code task in your team and write the prompt with separate blocks for task, conventions, target file, called signatures and tests.
- Add an output contract: a short plan plus one unified diff, no edits to tests or dependency files.
- Implement the bounded loop: apply-check, sandboxed tests, trimmed failure feedback, at most three rounds, then escalate.
- Build an evaluation set of 30 or more past changes from your repository with their tests, and measure pass@1 with the unbiased estimator.
- Change one thing at a time (context selection, contract, repair feedback) and keep a change only if pass rate improves without more reviewer rejections.
- Add type checking, a security linter and hidden tests to the verification stage before allowing any generated change near a merge.