Ask a coding agent to 'add shipping fees with a free-shipping threshold' and it will write code and, often, tests for that code. Both will pass. The trouble is that the tests were written to agree with the implementation, so they confirm what the agent decided, not what you wanted. If the agent misread the threshold as exclusive, the test will assert that 50.00 pays shipping, and the suite goes green on a bug.
Test-driven development (TDD) fixes this by putting the tests first: write a failing test that describes the behaviour, make it pass with the smallest change, then clean up while the tests stay green. With an agent doing most of the typing the loop still works, and it matters more, because the tests become the one artifact that pins down intent in a form the agent can execute. This guide shows how to split red, green and refactor between you, the agent and the harness, with a worked example, two small tools that close the obvious loopholes, and the failure modes to watch for.
Why TDD matters more when an agent writes the code
An agent is an optimiser with a loop: edit, run, read the output, edit again, until something says it is done. Whatever signal ends that loop is what it optimises. If the signal is 'the code looks finished', you become the verifier and every mistake waits for your review. If the signal is a test suite, the agent iterates until the suite passes, which is exactly what you want, as long as the suite encodes the right behaviour and the agent cannot change it.
That is the whole design problem. Anthropic's best-practices page for Claude Code makes the same point in general form: give the agent a check it can run, such as tests or a build, and prompt it along the lines of 'write a failing test that reproduces the issue, then fix it'. TDD turns that advice into a discipline with three properties an agent needs: the target is fixed before implementation starts, the target is executable, and the target was reviewed by someone who is not the implementer.
The loop, reassigned
In classic TDD one person does everything. With an agent, split the loop into roles with different privileges.
- Examples (human). Write the behaviour as concrete input and output pairs, including the boundaries you care about. This can be three lines in a ticket.
- Draft tests (agent, optional). The agent turns the examples into test code. It is not allowed to write implementation in the same step.
- Review tests (human). Read the tests as the spec. Are the boundaries right? Would a plausible wrong implementation fail at least one of them?
- Red check (harness). Run the new tests and confirm each fails for the right reason, then commit them. That commit is the contract.
- Green (agent). Implement until the tests pass, with the tests locked.
- Test lock (harness). Verify the tests and test configuration are unchanged since the red commit and no skips were added.
- Refactor (agent). Improve the structure with the tests still locked; run the full suite.
- Review (human). Review the implementation diff, and use a mutation score to judge whether the tests actually constrain it.
Two separations do most of the work: the implementer is not the only one who decided what the tests say, and the tests are frozen while the implementation is written. The Claude Code guide suggests a version of the first: one session writes tests and a different session writes code to pass them.
Worked example: a shipping fee
The rules from the product owner: standard shipping is 4.99, free when the subtotal is at least 50.00; express shipping is always 9.99; orders over 20 kg pay 1.50 for each started kilogram over 20, even when shipping is otherwise free; negative weights are an error. Money is Decimal, never float.
The human writes the stub and the tests together. The stub fixes the interface (the Order type and the function signature) so that failures are about behaviour, not about missing names.
# shop/shipping.py -- the stub the human commits with the red tests
from dataclasses import dataclass
from decimal import Decimal
@dataclass(frozen=True)
class Order:
subtotal: Decimal
weight_kg: Decimal
express: bool = False
def shipping_fee(order: Order) -> Decimal:
raise NotImplementedError# tests/test_shipping.py -- written (or approved) by a human, committed while red
from decimal import Decimal
import pytest
from shop.shipping import Order, shipping_fee
D = Decimal
@pytest.mark.parametrize("subtotal, expected", [
(D("49.99"), D("4.99")), # just under the threshold pays standard
(D("50.00"), D("0.00")), # threshold is inclusive
(D("0.01"), D("4.99")),
])
def test_standard_threshold(subtotal, expected):
assert shipping_fee(Order(subtotal=subtotal, weight_kg=D("1"))) == expected
def test_express_is_never_free():
order = Order(subtotal=D("120.00"), weight_kg=D("1"), express=True)
assert shipping_fee(order) == D("9.99")
@pytest.mark.parametrize("weight, surcharge", [
(D("20"), D("0.00")), # 20 kg is included
(D("20.1"), D("1.50")), # any part of a kg over 20 counts as a full kg
(D("23"), D("4.50")),
])
def test_heavy_surcharge_applies_even_when_free(weight, surcharge):
order = Order(subtotal=D("80.00"), weight_kg=weight)
assert shipping_fee(order) == surcharge
def test_rejects_negative_weight():
with pytest.raises(ValueError):
shipping_fee(Order(subtotal=D("10"), weight_kg=D("-1")))Notice what the tests pin down that a one-line prompt would not: the threshold is inclusive, express is never free, the surcharge rounds up partial kilograms, and the surcharge applies on top of free shipping. Each of those is a place an agent would otherwise guess. Reviewing these lines is faster than reviewing an implementation for the same decisions.
Red for the right reason
A red test proves nothing if it fails for the wrong reason. The common wrong reasons are structural: the test cannot import the module, a fixture has a typo, or a call has the wrong arguments. Those turn green the moment the name exists, whatever the behaviour. A test that has never failed on an assertion has never been shown to detect anything.
So check the reason mechanically. The pytest plugin below records, for each test, whether the call phase failed with an AssertionError, with the stub's NotImplementedError, or with something else, and treats collection errors as invalid. Run it before committing the red tests.
# tools/red_check.py -- run the new tests and classify WHY each one fails.
# Usage: python tools/red_check.py tests/test_shipping.py
import sys
import pytest
OK_REASONS = {"assertion", "NotImplementedError"}
class RedCheck:
def __init__(self):
self.results, self.collect_errors = {}, []
def pytest_collectreport(self, report):
if report.failed: # ImportError, SyntaxError at import
self.collect_errors.append(report.nodeid)
def pytest_runtest_makereport(self, item, call):
if call.when != "call":
return None
exc = call.excinfo
if exc is None:
self.results[item.nodeid] = "passed"
elif exc.errisinstance(pytest.skip.Exception):
self.results[item.nodeid] = "skipped"
elif exc.errisinstance(AssertionError) or exc.errisinstance(pytest.fail.Exception):
self.results[item.nodeid] = "assertion"
else:
self.results[item.nodeid] = exc.typename # e.g. ImportError, TypeError
return None # let pytest build the report as usual
def main(paths):
rc = RedCheck()
pytest.main(["-q", "-p", "no:cacheprovider", *paths], plugins=[rc])
bad = [f"collection error: {n}" for n in rc.collect_errors]
bad += [f"{n}: {r}" for n, r in rc.results.items() if r not in OK_REASONS]
if not rc.results:
bad.append("no tests ran")
for line in bad:
print("NOT A VALID RED ->", line)
return 1 if bad else 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))With the stub above, every test fails with NotImplementedError, including test_rejects_negative_weight, because pytest.raises(ValueError) does not catch other exception types. That is accepted as a valid red. If you want a stricter red, give the stub a deliberately wrong body such as return Decimal('-1'), and then every behavioural test fails on an assertion. An ImportError or TypeError means the tests or the stub are wrong; fix that before the agent implements anything.
Locking the tests during green
Under pressure to reach green, agents take shortcuts that a tired human might also take: loosen an assertion, mark a test as expected to fail, special-case the exact inputs the tests use, or edit test configuration so the failing file is not collected. Instructions are advisory; a script is not.
# tools/test_lock.py -- fail if the agent touched the tests during green/refactor.
# Usage: python tools/test_lock.py <red-commit-sha>
import re, subprocess, sys
def git(*args):
return subprocess.run(["git", *args], capture_output=True, text=True, check=True).stdout
red = sys.argv[1]
problems = []
# 1. Tests and test configuration must be identical to the red commit.
LOCKED = ["tests/", "conftest.py", "pytest.ini", "pyproject.toml", "setup.cfg", "tox.ini"]
for line in git("diff", "--name-status", red, "--", *LOCKED).splitlines():
status, path = line.split("\t", 1)
if status in ("M", "D") or status.startswith("R"):
problems.append(f"test file changed since red commit: {status} {path}")
# 2. No new escape hatches anywhere in the diff.
HATCHES = re.compile(r"pytest\.(skip|xfail)|@pytest\.mark\.(skip|xfail)|--deselect|"
r"pragma: no cover|# type: ignore|noqa")
for line in git("diff", red, "--", ".").splitlines():
if line.startswith("+") and not line.startswith("+++") and HATCHES.search(line):
problems.append(f"new escape hatch: {line[1:].strip()[:80]}")
for pr in problems:
print("TEST LOCK:", pr)
sys.exit(1 if problems else 0)Run the lock as the agent's last step, as a harness gate, and in CI. In Claude Code, for example, a Stop hook can run a script and block the turn from ending until it passes; other agents have comparable hooks or can be wrapped in a script that checks the exit code.
The prompt for green should state the rules plainly and give the agent a legitimate way out when it thinks a test is wrong.
Implement shipping_fee in shop/shipping.py so that tests/test_shipping.py passes.
Rules for this task:
- Do NOT edit anything under tests/, conftest.py or pytest/pyproject config.
- Do not add skips, xfails or special cases keyed on specific test values.
- Implement the rule the tests describe, not the individual examples.
- Run: pytest -q tests/test_shipping.py, then the full suite.
- If you believe a test is wrong, STOP and explain why instead of changing it.
- Finish by running: python tools/test_lock.py <RED_SHA> and paste its output.The 'stop and explain' rule matters: when a test really is wrong, you want a question to a human, not a silent fix in whichever direction makes the suite pass. Fix the test yourself, make a new red commit, and restart green.
Refactor, and measuring whether the tests bite
Refactoring is where agents shine, because behaviour-preserving restructuring is tedious for people and well defined for a model when the tests are locked. Ask for one refactoring at a time (extract the surcharge calculation, name the magic numbers) and run the full suite plus the lock after each. Reject refactors that change public signatures unless you asked for that; they force test changes, which puts you back in red.
Passing tests only prove that the code does what the tests check. To learn whether the tests check enough, use mutation testing: a tool makes small changes to the implementation, such as flipping >= to > or changing 20 to 21, and reruns the tests. A mutant that survives is a behaviour change your tests did not notice. Tools such as mutmut for Python and Stryker for JavaScript and TypeScript do this. In the example, the inclusive-threshold test kills the >= to > mutant; if you delete that test, the mutant survives, and the report tells you exactly which decision is now unprotected.
Mutation runs are slow, so scope them to changed files in CI, and treat survivors as review prompts rather than a hard gate, since some mutants are equivalent to the original.
Failure modes
- Implementation-shaped tests. The agent drafts tests after reading or planning the implementation, so the tests mirror its assumptions. Draft tests from examples only, in a session that has not planned the code.
- Assertion erosion. assertEqual becomes assertIn, exact values become 'is not None'. The lock catches edits after the red commit; test review catches weak tests before it.
- Teaching to the test. The implementation branches on the literal inputs, such as if weight == Decimal('23'). Property-based tests with generated inputs and mutation testing both expose it; so does a reviewer reading the diff.
- Mock-heavy green. The agent mocks the collaborator that makes the test hard, and the test now checks the mock. Say which dependencies may be faked, and prefer real in-memory versions.
- Giant steps. One prompt with forty tests yields one large diff that is hard to review. Feed tests in small groups, one behaviour at a time, as classic TDD does.
Trade-offs and when to skip it
TDD with agents costs human time up front, on examples and test review, and it pays that back in fewer review cycles and fewer bugs that reach production. It is clearly worth it for business rules, parsers, calculations, state machines and bug fixes, where the behaviour is precise and a reproduction test is the natural first step. It is weaker for exploratory UI work, where the right behaviour is discovered by looking, and for glue code whose tests would mostly be mocks.
Do not confuse TDD with 'the agent also writes tests'. Tests written after the fact by the implementer are regression protection at best. The value comes from the ordering and the separation: intent fixed first, by someone accountable for it, and held constant while the agent searches for an implementation. For reviewing the resulting pull request, how to review a PR applies unchanged; for limiting what the agent may touch while it works, see permission boundaries for agents; and for keeping the task's context focused, context engineering for agents. To wire the lock and mutation checks into a pipeline, CI/CD patterns covers the stages.
What to do next
- Pick one upcoming bug fix and write the reproduction test yourself before involving the agent.
- Add a stub-plus-tests commit step to your workflow and run the red check before every red commit.
- Copy the test-lock script into tools/, run it in CI and as the agent's final step.
- Add the 'stop and explain if a test seems wrong' rule to your agent prompts.
- Run mutation testing on one module the agent wrote last month and read the surviving mutants.
- Try the two-session split: one session drafts tests from examples, another implements against them.