Code written by an AI assistant or agent does not have a special kind of bug. It has the same bugs human code has, in a different mix, arriving at a higher volume, and wrapped in code that looks more finished than it is. Generated code is fluent: consistent naming, tidy docstrings, plausible tests. That fluency removes the cues reviewers unconsciously rely on, such as awkward code in the places where the author was unsure.
This guide is about reviewing that code well and at scale. It covers the defect classes to hunt for, a reading order that finds them faster, a scriptable first pass that removes mechanical checks from the human's plate, risk tiers that decide how much review each change gets, and the process failures that let bad changes through anyway. General review practice (tone, turnaround, style debates) is covered in how to review a pull request and code review; this article assumes those basics.
Why review has to change
Three things are different when a model writes the diff. First, volume: an engineer with an agent can produce several times as many lines per day, but review capacity has not grown, so review becomes the constraint. Second, plausibility: generated code is optimised to look like correct code, which is not the same as being correct, and the gap is hardest to see in exactly the places that matter, such as boundary conditions and error handling. Third, accountability: you cannot ask the model why it chose a design and trust the answer to reflect the process that produced the code. Its explanation is a new generation, not a memory.
The practical consequence: the human who submits AI-written code is its author for every purpose that matters. They must have read it, be able to explain it, and have run it. Reviewers should hold them to that, and teams should make it an explicit rule rather than an assumption.
The defect classes to hunt for
These are the problems that show up more often in generated diffs than in hand-written ones. Knowing the list turns vague suspicion into specific questions.
- Hallucinated APIs and packages. Methods, flags, config keys or whole libraries that do not exist, or existed in an older version. Code that fails to compile is caught by CI; a keyword argument silently ignored by
**kwargsis not. New dependencies deserve special care because a non-existent package name can be registered by an attacker; LLM hallucination risk covers this in depth. - Plausible but wrong logic. Off-by-one at boundaries, inverted conditions, timezone-naive dates, float money, retrying non-idempotent operations. The code reads naturally, which is the problem.
- Weakened tests. Tests that assert truthiness instead of values, mock the unit under test, were edited to match new (wrong) behaviour, or were deleted or skipped to get CI green.
- Silent failure. Broad exception handlers, default values returned on error, fallbacks that fabricate a success. Models are trained on a great deal of code that prioritises not crashing.
- Over-broad diffs. Unrequested refactors, renamed variables, reformatted files and "improvements" that bury the actual change and create merge conflicts.
- Duplication. A new helper written from scratch because the model did not know one already existed three directories away. Each copy then drifts.
- Insecure defaults. Disabled certificate checks, string-built SQL, permissive CORS, secrets in example config, logs that include tokens. See SAST and DAST for the tools that catch the mechanical subset.
- Comments that lie. Docstrings describing what the code was meant to do, not what it does. Read the code, not the comment.
A reading order that finds them
The order in which you read a diff changes what you find. For generated code, read in this order:
- The description and provenance. What was asked for? If the diff does more than that, the extra is suspect by default.
- The file list and size. Anything touched outside the expected area needs a reason. A 900-line diff for a 20-line request is a review failure waiting to happen; send it back to be split.
- Tests, before implementation. Ask what behaviour each test pins. Would it fail if the implementation were wrong? Were existing tests changed, and does the change make them weaker? Reading tests first means you form expectations before the implementation's fluency persuades you.
- Interfaces and dependencies. New public functions, changed signatures, new packages. Check that every imported package and called API exists in the version you use.
- Error paths. Every
except,catch, default return and retry. This is where silent failure hides. - The happy path. Last, and fastest, because the earlier steps have already told you what to look for.
When something is unclear, ask the author to explain it in their own words. "I'm not sure, the agent wrote it" is a legitimate answer that means the change is not ready, not that the reviewer should work it out on their behalf.
The machine first pass
Humans should not spend attention on anything a script can check. Beyond the usual CI (build, tests, lint, type checking, SAST, secret scanning), add a small check for the patterns generated diffs get wrong. The script below flags them as warnings for the reviewer; it is deliberately crude and will have false positives, which is acceptable because it points a human at a line rather than blocking a merge on its own.
#!/usr/bin/env python3
"""ai_diff_smells.py BASE -- flag patterns that AI-written diffs get wrong often. Warnings, not verdicts."""
import re, subprocess, sys
base = sys.argv[1] if len(sys.argv) > 1 else "origin/main"
diff = subprocess.run(["git", "diff", "-U0", base], capture_output=True, text=True).stdout
RULES = [
("test skipped", r"^\+.*(@pytest\.mark\.skip|\.skip\(|xit\(|@Disabled|t\.Skip\()"),
("broad except", r"^\+\s*except\s*(Exception)?\s*:\s*(pass)?\s*$"),
("swallowed error", r"^\+.*catch\s*\(\w*\)\s*\{\s*\}"),
("TLS check disabled", r"^\+.*(verify\s*=\s*False|rejectUnauthorized\s*:\s*false|InsecureSkipVerify:\s*true)"),
("SQL by string", r"^\+.*(execute|query)\(\s*f?[\"'].*(SELECT|INSERT|UPDATE|DELETE).*(\{|\+|%s\"\s*%)"),
("new dependency", r"^\+\s*[\"']?[A-Za-z0-9_.\-]+[\"']?\s*[:=<>~^].*$"),
("TODO left behind", r"^\+.*\b(TODO|FIXME|XXX)\b"),
]
DEP_FILES = ("requirements", "pyproject.toml", "package.json", "go.mod", "Cargo.toml")
findings, current = [], None
removed_asserts = added_asserts = 0
for line in diff.splitlines():
if line.startswith("+++ b/"):
current = line[6:]; continue
if current and "test" in current:
if re.match(r"^-\s*(assert|expect\(|self\.assert)", line): removed_asserts += 1
if re.match(r"^\+\s*(assert|expect\(|self\.assert)", line): added_asserts += 1
for name, pat in RULES:
if name == "new dependency" and not (current or "").startswith(DEP_FILES):
continue
if re.search(pat, line):
findings.append((current, name, line[1:].strip()[:100]))
if removed_asserts > added_asserts:
findings.append(("tests", "net assertions removed", f"-{removed_asserts} +{added_asserts}"))
for f, name, snippet in findings:
print(f"{name:24} {f}: {snippet}")
sys.exit(1 if findings else 0)The net-assertion count is the most useful single signal: a change whose tests lose more assertions than they gain needs an explanation. Grow the rule list from your own escapes. Every bug that reaches production from an AI-assisted change should prompt the question "could a pattern have flagged this?" and, where the answer is yes, a new rule.
Risk tiers: spend review where it matters
Not every change deserves the same scrutiny. A typo fix in documentation and a change to payment retries should not get the same review. Route by risk, using signals you can compute from the diff:
HIGH_RISK = ("auth/", "payments/", "migrations/", "infra/", ".github/workflows/", "crypto")
def review_tier(files, added_lines, new_deps, smells):
"""Route a PR to a review depth. Tune thresholds from your own escape data."""
if any(f.startswith(HIGH_RISK) or any(h in f for h in HIGH_RISK) for f in files):
return 3, "touches a high-risk path: two reviewers including the code owner"
if new_deps or added_lines > 400 or "net assertions removed" in smells:
return 2 if added_lines <= 800 else "split", "large, new deps or weaker tests"
if all(f.endswith((".md", ".txt")) or "/docs/" in f for f in files):
return 1, "docs only"
return 2, "default: full read, tests first"The split outcome is important. Large diffs get worse review, not proportionally longer review, because attention runs out. Setting a hard size limit and asking for AI-generated work to be split into stacked, independently reviewable changes is one of the cheapest quality improvements available, and agents are good at splitting their own work when asked. High-risk paths should also be protected with a code-owners file so the tier is enforced by the platform, not by goodwill.
Worked example: retries in a payment client
A developer asks an agent to add retries to a flaky payment call. CI is green and the diff is short and tidy:
# Agent-written change: "add retries to the payment client"
def charge(self, customer_id, amount_cents, currency="EUR"):
for attempt in range(5):
try:
resp = self.http.post("/v2/charges", json={
"customer": customer_id, "amount": amount_cents, "currency": currency})
resp.raise_for_status()
return resp.json()
except Exception: # retries everything, including 4xx
time.sleep(2 ** attempt)
return {"status": "pending"} # fabricated success after 5 failures
# tests/test_payments.py (also agent-written)
def test_charge_retries(mocker):
mocker.patch.object(client.http, "post", side_effect=[Timeout(), ok_response])
assert client.charge("c1", 500) # truthy dict passes, whatever it containsFollowing the reading order finds four problems in about five minutes. Tests first: the test asserts only that the result is truthy, so it would pass if charge returned any non-empty dict, including the fabricated one. Error paths: except Exception retries client errors such as an invalid card, which will never succeed, and hides bugs such as a KeyError in the code. Logic: retrying a POST that creates a charge is not safe unless the request carries an idempotency key the provider honours; a timeout after the provider processed the charge becomes a double charge. Silent failure: after five failures it returns {"status": "pending"}, telling the caller a charge is in progress when nothing is.
The fixed version retries only timeouts and 5xx responses, sends an idempotency key generated once per logical charge (checking the provider's documentation for the exact header it expects), raises after the final attempt, and has tests asserting on the exact request count, the reuse of the same key across attempts, and the exception raised when retries are exhausted. None of the four problems would have been caught by a reviewer reading top to bottom and nodding at clean code.
Provenance and author responsibility
Make it visible how a change was produced and what evidence supports it. A pull request template does this cheaply:
<!-- .github/pull_request_template.md -->
## What and why
<!-- one paragraph; link the issue or plan step -->
## How this was produced
- [ ] Written by hand
- [ ] Agent-assisted: agent wrote most of the diff; I have read every line
- Prompt / plan file: <!-- path or link -->
## Evidence
- Tests added or changed, and what behaviour they pin:
- Commands run locally and their result:
- For bug fixes: the test that failed before this change:
## Risk
- New dependencies (name, why, checked it exists and is maintained):
- Touches auth / payments / data migrations / CI config? yes / noProvenance is not about blame or about treating AI-assisted changes as second-class. It tells the reviewer which defect classes to weight and whether the author has done the work of reading their own diff. The evidence section matters most: "tests pass" is weak, but "this test failed before the change and passes after" is strong evidence that the test exercises the fix.
Using AI reviewers without trusting them
AI review tools are useful as part of the first pass. They are good at spotting inconsistent naming, missing null checks, obvious injection risks and mismatches between description and diff, and they never get tired. Their failure modes are specific: confident false positives that waste time, missing context about why the code is the way it is, and correlated blind spots when the reviewing model is similar to the one that wrote the code. Use them to generate questions, never as the approver. A useful practice is to run the AI review before the human one and require the author to respond to each finding, which moves the triage cost to the author.
Process failures
- Green-CI approval. Approving because checks pass. Checks only prove what they test, and weakened tests pass.
- Review fatigue. A reviewer facing ten large generated diffs a day will skim. Enforce size limits and share the load.
- Rubber-stamp loops. Two people each assuming the other read it carefully. Make the tier's requirements explicit.
- No feedback loop. Escaped defects fixed without asking why review missed them. Audit a sample of merged AI-assisted changes monthly and track the escape rate.
Trade-offs
Thorough review of every generated line cancels most of the speed gain; no review at all converts that speed into incidents. Tiering is the compromise: machines take the mechanical checks, low-risk changes get light review, and human attention concentrates on high-risk paths and on tests. The cost is setup (scripts, templates, code owners) and occasional friction when a change is sent back to be split. Both are small next to one double-charging bug in production.
What to do next
- Adopt the reading order (tests first, error paths before happy path) on your next three reviews of generated code.
- Add the diff-smell script, or your own version, to CI as a non-blocking annotation.
- Set a diff-size limit and a code-owners file for high-risk paths.
- Add a provenance and evidence section to your pull request template.
- Every time an AI-assisted change causes an incident, add a first-pass rule or tier trigger that would have caught it.