Give a capable coding agent a one-line request such as 'add CSV export for invoices' and it will produce something that compiles, looks reasonable, and quietly decides a dozen things nobody asked it to decide: which invoices count, what timezone defines a month, how large exports behave, who is allowed to run them. Some of those guesses will be wrong, and you will discover which ones in review if you are careful, or in production if you are not.
Spec-driven development (SDD) addresses this by splitting the work into a pipeline of written artifacts, each reviewed before the next is generated: a spec that says what and why, a plan that says how, a task list that breaks the plan into small verifiable steps, and then implementation one task at a time. The agent does most of the writing at every stage; people do the deciding. This guide explains each artifact, what to check at each gate, a complete worked example, code to validate a task list and drive the loop, and where the approach costs more than it saves.
Why a pipeline instead of a prompt
An agent turns ambiguity into decisions at generation time, and those decisions are invisible until you read the code. SDD moves the decisions earlier, into documents where they are cheap to see and cheap to change. A wrong assumption caught in a one-page spec costs a sentence; the same assumption caught after implementation costs a rewrite and a second review. The pipeline also bounds the agent's scope at each step: while writing the spec it should not be choosing libraries, while writing the plan it should not be writing code, and while implementing a task it should not be redesigning the feature.
The idea is not new. It is requirements, design doc and work breakdown, the sequence behind any good design doc. What changed is the cost: an agent can draft each artifact in a minute, so the documents that teams used to skip for small features are now nearly free to produce. The expensive part is reviewing them, which is exactly where human attention should go.
The four artifacts
Most SDD setups use the same four documents, whatever they call them.
- Constitution (or project principles). Written once per repository: architecture rules, testing standards, security constraints, preferred libraries, things never to do. Every later step reads it. It changes rarely and is reviewed like policy.
- Spec. Per feature. User stories, acceptance criteria, explicit out-of-scope items and open questions. It describes behaviour, not implementation: no table names, no frameworks.
- Plan. Per feature. The technical approach: components touched, data model changes, interfaces, dependencies, risks, and how each acceptance criterion will be tested.
- Tasks. Per feature. An ordered list of small changes, each with its dependencies and a concrete verification, sized so that one agent run produces one reviewable commit.
The documents live in the repository, usually in a directory per feature, so the agent can reload them as context in any later session and reviewers can see them in the same pull request as the code. Treat them as living documents until the feature ships, and as history afterwards.
Writing and reviewing the spec
A useful spec answers the questions an agent would otherwise answer for you. For each user story, write acceptance criteria that a test could check, list what is explicitly out of scope, and mark every unresolved question with a marker you can grep for. Ask the agent to draft the spec and then to list the ambiguities it sees; models are good at finding gaps in someone else's text, even when they would fill the same gaps silently in their own code.
Here is the spec for the invoice export. Notice AC1 and AC5: an empty month and voided invoices are exactly the edge cases an agent would decide on its own, and different agents or runs would decide them differently.
# specs/014-invoice-export/spec.md
## Summary
Accountants can export a month of invoices to CSV for import into their ledger.
## User stories
1. As an accountant, I choose a month and download all issued invoices as CSV.
2. As an accountant, I get one row per invoice line, so totals reconcile.
## Acceptance criteria
- AC1: export for a month with no invoices returns a header-only file, not an error
- AC2: amounts use the invoice currency, two decimals, dot separator, no symbols
- AC3: exports over 50,000 lines run in the background and are emailed as a link
- AC4: only users with the accountant role can export; others get 403
- AC5: voided invoices are excluded; credit notes appear with negative amounts
## Out of scope
- XLSX output; custom column selection; scheduled exports
## Open questions
- OPEN: which timezone defines "the month"? (asked finance, 2026-09-22)Review the spec as a product document. Does each criterion describe observable behaviour? Is anything in 'out of scope' actually required? Is there an open question whose answer changes the design? The timezone question here does: it decides whether the month query runs in UTC or in the account's zone. Do not move on while an OPEN marker remains, because the plan built on a guess inherits the guess. The habits in How to Review a Design Doc transfer directly.
From plan to tasks
The plan is where the agent proposes technology and structure, constrained by the constitution. For the export it would name the existing invoice repository, a CSV writer from the standard library, the background job framework already in use, and a signed-URL helper for the emailed link. Review the plan for fit with the existing system: new dependencies, new patterns, anything that contradicts the constitution. A plan that introduces a second job queue for one feature should be sent back.
Tasks are the most important artifact to get right, because they decide how reviewable the implementation will be. Each task should produce one coherent change, touch a small number of files, name its dependencies, and carry a verification that proves it is done. 'Implement export' is not a task; 'CSV row mapper with tests for AC2 and AC5' is.
# specs/014-invoice-export/tasks.md
- [ ] T1 CSV row mapper: InvoiceLine -> list[str] | verify: unit tests AC2, AC5
- [ ] T2 month query with timezone from settings (after: T1) | verify: test month boundaries
- [ ] T3 sync export endpoint + role check (after: T2) | verify: API tests AC1, AC4
- [ ] T4 background job for large exports (after: T2) | verify: job test AC3
- [ ] T5 emailed signed link, 24 h expiry (after: T4) | verify: integration test AC3
- [ ] T6 docs + changelog entry (after: T3, T5) | verify: docs buildBecause tasks.md is structured, you can check it mechanically before an agent starts. The validator below rejects tasks without a verification, unknown or cyclic dependencies, acceptance criteria that no task verifies, and specs with open questions. It uses only the standard library, including graphlib for the ordering.
# Validate tasks.md before an agent starts: every task verifiable, deps exist,
# no cycles, and every acceptance criterion covered by at least one task.
import re
from graphlib import TopologicalSorter, CycleError
TASK = re.compile(r"- \[( |x)\] (T\d+) (.+?)(?:\(after: ([T\d, ]+)\))?\s*\| verify: (.+)")
def check(tasks_md, spec_md):
tasks, errors = {}, []
for line in tasks_md.splitlines():
if line.startswith("- ["):
m = TASK.match(line.strip())
if not m:
errors.append(f"unparseable or missing verify: {line}")
continue
done, tid, _, after, verify = m.groups()
deps = [d.strip() for d in (after or "").split(",") if d.strip()]
tasks[tid] = {"deps": deps, "verify": verify, "done": done == "x"}
for tid, t in tasks.items():
errors += [f"{tid} depends on unknown {d}" for d in t["deps"] if d not in tasks]
try:
order = list(TopologicalSorter({k: v["deps"] for k, v in tasks.items()}).static_order())
except CycleError as e:
return [f"dependency cycle: {e.args[1]}"], []
criteria = set(re.findall(r"\bAC\d+\b", spec_md))
covered = set(re.findall(r"\bAC\d+\b", " ".join(t["verify"] for t in tasks.values())))
errors += [f"{ac} not verified by any task" for ac in sorted(criteria - covered)]
if "OPEN:" in spec_md:
errors.append("spec still has open questions")
return errors, order
The implementation loop
Implementation runs one task at a time, in dependency order, with a hard verification step between tasks. The agent gets the constitution, spec, plan, the single task and the files the plan names, and is told to stop when the task's verification passes. If verification fails after one repair attempt, a person looks, because repeated failure usually means the plan is wrong rather than the code.
# The implementation loop: one task, one verified change, one commit.
for task in ready_tasks(order): # deps done, not yet done
context = [constitution, spec, plan, task.text, *files_named_in(plan, task)]
agent.run(f"Implement {task.id} only. Stop when '{task.verify}' passes.", context)
verify, checks = run(task.verify_command), run("make lint typecheck")
if not (verify.ok and checks.ok):
agent.run("Fix the failure; do not change the spec or other tasks.", verify.log + checks.log)
if not (run(task.verify_command).ok and run("make lint typecheck").ok):
escalate(task) # a human looks; maybe the plan is wrong
break
if agent.reports_spec_conflict():
pause_and_update_spec(task) # drift goes upstream, never silent
commit(f"{task.id}: {task.title}")
mark_done(task)Two properties of this loop matter. One commit per task gives reviewers a sequence of small, explained diffs instead of one large one, and makes bisecting easy. And an agent that discovers a conflict, such as an existing permission model that makes AC4 impossible as written, pauses and raises it instead of improvising. That conflict becomes a spec change, reviewed like the original, before work continues. This is the discipline that keeps the artifacts truthful; without it the spec describes the feature you planned and the code implements a different one.
Tooling
You can run SDD with plain Markdown templates and any agent. Dedicated tools package the templates and commands. GitHub's open-source Spec Kit provides a CLI that sets up a repository for several coding agents, plus agent commands for each stage: establishing the constitution, specifying, planning, breaking into tasks and implementing, with optional steps for clarifying a spec, generating quality checklists and cross-checking artifacts for consistency. Its command names have changed across releases and differ by agent integration, so check the version you install. Kiro, AWS's agentic IDE, follows a similar requirements, design and tasks sequence.
Whichever you use, the value is in the review gates, not in the commands. A tool that generates all four artifacts in one step and proceeds straight to code has the shape of SDD without the substance.
Failure modes
- Rubber-stamp gates. The artifacts are generated and approved without reading. You get the overhead with none of the benefit. Keep specs short enough that people actually read them.
- Implementation in the spec. Specs that name tables and frameworks lock the plan before anyone reviewed it. Keep the spec about behaviour.
- Oversized tasks. A task that touches fifteen files produces an unreviewable diff and an agent that loses track. Split until each task is one commit a reviewer can read in minutes.
- Unverifiable tasks. 'Refactor for clarity' has no stopping condition. Every task needs a check that passes or fails.
- Spec drift. Code changes during implementation and the spec is never updated, so the next feature builds on a document that lies. Route behaviour changes upstream.
- Premature detail on uncertain work. Writing a full spec for something you do not yet understand produces confident fiction. Spike first, then specify.
Trade-offs and when to use it
SDD costs review time up front and a few documents per feature, and it pays back in fewer wrong guesses, smaller reviewable diffs and an agent that can resume work in a fresh session by rereading the artifacts. It suits features with real acceptance criteria, multi-day work, and teams where several people or agents touch the same code. It is overhead for one-line fixes, exploratory prototypes and throwaway scripts. It pairs naturally with test-driven development, since acceptance criteria become tests, and with SPDD, which applies the same discipline to the generating prompt itself. For estimating the resulting task lists, see How to Estimate an Engineering Project; for regression testing of agent behaviour, see LLM Evaluation Harness and Regression Testing.
What to do next
- Write a one-page constitution for your repository: architecture rules, testing standard, banned patterns, preferred libraries.
- For your next multi-day feature, have the agent draft a spec, then ask it to list ambiguities; resolve every OPEN item before planning.
- Review the plan for new dependencies or patterns that contradict the constitution before generating tasks.
- Adopt a structured tasks.md format with dependencies and a verify clause, and run a validator like the one above in CI.
- Implement one task per agent run and one commit per task; escalate after one failed repair.
- When implementation reveals a behaviour change, update the spec in the same pull request before continuing.
- After the feature ships, note which gate caught the most problems and put your review effort there next time.