One developer can adopt Structured Prompt-Driven Development (SPDD) in an afternoon: write a REASONS canvas, generate from it, commit both. A team cannot, because the value of SPDD is shared: canvases other people can review, Norms that apply to everyone, Safeguards owned by the people who understand the risks, and a history that explains the code to whoever touches it next. Those need repository conventions, review rules and a way to tell whether the effort is paying off.
This guide covers that team layer: where to apply the method and where not to, how to lay out the prompt repository, who owns what, what reviewers and CI should check, how to govern shared rules, how to roll it out, and how to measure it with numbers your own history can produce. The method itself is in the SPDD overview; companion guides go deep on writing the canvas and on the daily loop.
Decide where SPDD applies before anything else
The fastest way to make a team hate a method is to apply it to everything. The source article, by Wei Zhang and Jessie Jie Xia on martinfowler.com, rates fitness explicitly, and those ratings are a good starting policy.
| Work | Source rating | Team policy |
|---|---|---|
| Scaled, standardised delivery | 5 stars | Canvas required |
| High compliance or hard constraints | 5 stars | Canvas required; security reviews Safeguards |
| Collaboration needing auditability | 5 stars | Canvas required |
| Firefighting hotfixes | 2 stars | No canvas up front; sync within a working day |
| Exploratory spikes, one-off scripts | 2 stars | No canvas; write one if the spike becomes a feature |
| Poorly defined domains, creative or visual work | 1 star | No canvas; clarify the domain first |
Write the policy down in the repository, in one paragraph, and decide who can grant an exception. Without an explicit out-of-scope list every small change becomes an argument about whether it needs a canvas, and the method loses by attrition.
The prompt repository
Keep canvases in the same repository as the code they generate, not in a wiki or a separate prompt store. The point is that a single pull request changes intent and implementation together, and that versioning comes from git rather than a second system. (Runtime prompts shipped inside an LLM application are a different problem with different tooling.) One directory per feature holds the story with its dated clarification answers, the analysis notes and the canvas. A shared directory holds the template and team-wide rules.
specs/
_shared/
TEMPLATE.md # blank REASONS canvas, versioned
norms.md # team-wide Norms (also loaded by the instruction file)
safeguards.md # organisation-wide Safeguards: auth, PII, egress
order-webhooks/
story.md # the user story and clarification Q&A, dated
analysis.md # domain terms, relevant code, constraints
REASONS.md # the canvas; header has owner and status
order-webhooks-replay/
REASONS.md # references ../order-webhooks/REASONS.md
_archive/ # canvases for removed features, kept for history
src/ ...
tests/ ...Two conventions keep this manageable. A follow-on feature gets its own canvas that references the original rather than growing it, so each canvas stays reviewable. And canvases for deleted features move to an archive rather than being removed, because the reason a feature was built is useful long after its code is gone.
Ownership and shared rules
CODEOWNERS works on paths, not sections, so ownership follows the layout. Feature teams own their feature directories. Engineering leads own the template and the shared Norms. The security team owns organisation-wide Safeguards, such as rules on authentication, personal data and outbound network access, in their own file that every canvas references instead of copying.
# .github/CODEOWNERS - later rules win for a path
/specs/ @acme/eng-leads
/specs/_shared/norms.md @acme/eng-leads @acme/platform
/specs/_shared/safeguards.md @acme/security
/specs/order-webhooks*/ @acme/integrations
/src/webhooks/ @acme/integrationsThe shared Norms file is also what your assistant's project instruction file should load, so the rule is written once and applied in every session, canvas or not; instruction files for coding agents covers that mechanism. A Norm that appears in three feature canvases is a candidate for promotion. Changes to shared files go through the same pull request review as code, because a one-line edit to shared Safeguards changes the boundary for every future generation.
Review rules
Three rules carry most of the weight. First, the canvas is approved before code is generated from it; the canvas header's status field records this, and a canvas generated after the fact is worse than none because it looks like evidence. Second, reviewers read the canvas diff before the code diff. Third, every defect found in review is labelled as a model slip or an intent error, because that label drives both the fix and the metrics later.
## Canvas
- [ ] canvas diff included, or: no intent change (say why)
- [ ] canvas status is "reviewed" and was approved before code was generated
- Canvas: specs/<feature>/REASONS.md Operations covered: <n, n>
## Defects found in review (label the PR)
- [ ] spdd:slip - canvas right, code wrong; test added: <name>
- [ ] spdd:intent - canvas wrong; canvas fixed first in this PR
## Safeguards
- [ ] every Safeguard touched by this change has a testCanvas review needs different reviewers for different sections: the product owner for Requirements, a security or reliability owner for Safeguards, a senior engineer for Approach and Structure. Pairing a junior engineer with a senior one on canvas review is the fastest way to teach the three skills the source names: abstraction first, alignment and iterative review. Code review then checks the diff against the canvas, using the habits in reviewing AI-generated code.
What CI should check, and what it should not
Automate the structural checks and leave judgement to people. Useful checks: every canvas has all seven sections, a status and an owner; canvases with status draft cannot be referenced by a merged pull request; tests that bind canvas numbers to code constants pass. And a drift warning: when a pull request changes files a canvas lists as new in Structure, the files that feature owns, but does not change that canvas, CI posts a warning asking the author to confirm there is no intent change.
# drift_check.py BASE HEAD - warn when files a canvas's Structure lists as
# new (files the feature owns) change but that canvas does not. Shared files
# such as src/app.py are skipped: they change for unrelated reasons.
# Prints warnings; never fails the build.
import re, subprocess, sys
from pathlib import Path
def git(*args):
return subprocess.run(["git", *args], capture_output=True, text=True,
check=True).stdout
def owned_paths(canvas_text):
s = canvas_text.split("## S - Structure", 1)[-1].split("## O -", 1)[0]
paths = set()
for item in re.split(r"^- ", s, flags=re.M):
item = " ".join(item.split())
if item.startswith("new:"):
for base, names, single in re.findall(r"([\w/]+)/\{([\w,]+)\}\.py|([\w/]+\.py)", item):
paths |= {single} if single else {f"{base}/{n}.py" for n in names.split(",")}
return paths
base, head = sys.argv[1], sys.argv[2]
changed = set(git("diff", "--name-only", f"{base}...{head}").split())
for canvas in Path("specs").glob("*/REASONS.md"):
if canvas.parts[1].startswith("_"):
continue
touched = changed & owned_paths(canvas.read_text(encoding="utf-8"))
if touched and canvas.as_posix() not in changed:
print(f"::warning::{canvas.as_posix()} unchanged but {sorted(touched)} changed;"
" confirm no intent change or run a sync")Keep the drift check a warning. Plenty of legitimate changes, such as refactors, dependency bumps and small bug fixes, touch code without changing intent. A blocking check teaches people to make meaningless canvas edits to satisfy it, which destroys the signal. The warning's job is to make the decision explicit, not to make it for the author.
Governance: what canvases may contain
- No secrets or customer data. Canvases are sent to model providers as prompts. Treat them as you would code that leaves the building; securing AI-assisted development covers the wider threat model.
- Dated decisions. When a requirement changes, record who decided and when, in the canvas. For regulated work this is the audit trail an assessor asks for.
- A lifecycle. Draft, reviewed, superseded, archived. Tooling should refuse to generate from draft or superseded canvases.
- A template owner. The template changes by pull request, with a note explaining why, and no more than once a quarter so canvases stay comparable.
Rolling it out
Start with one team and three consecutive features in a well-understood domain, which is where the source's ratings are highest. Before the pilot, record a baseline from the previous three features: review time per pull request, reverts and follow-up fixes within a few weeks of merge, and how often reviewers asked why something was done. Run a canvas-review pairing session for each pilot feature, and hold a short retrospective after each.
Expand only when the pilot team would choose to keep the method without being asked. Each new team gets the template, the shared files and one experienced reviewer for its first canvases. Expect the design-first habit to take longest; the source article is explicit that the shift from code first to design first needs continuing training rather than one workshop.
Measuring whether it works
The source article reports no adoption data across teams, only an outcome for its own worked example, and industry productivity figures will not tell you whether SPDD works for your team. Define a small set from data you already have and compare against your own baseline:
| Metric | Source | What it tells you |
|---|---|---|
| Canvas-first ratio | git history | Is the canvas used to design, or written afterwards? |
| Sync rate | git history | Do canvas and code move together after launch? |
| Slip to intent-error ratio | PR labels | Are defects about the model or about unclear intent? |
| Review time per PR | code host | Does reviewing against a canvas speed review up? |
| Rework within 21 days | git history | Is generated code correct the first time? |
# spdd_metrics.py - two adoption metrics from git history alone, measured on
# the files each canvas's Structure lists as "new" (files the feature owns;
# shared files such as src/app.py have unrelated history).
# canvas-first: the canvas's first commit is not later than the first commit
# touching any of those files.
# sync rate: share of later commits touching those files that also touch
# the canvas. Low is not bad by itself; watch the trend.
import re, subprocess
from pathlib import Path
PAT = r"([\w/]+)/\{([\w,]+)\}\.py|([\w/]+\.py)"
def owned_paths(canvas_text):
s = canvas_text.split("## S - Structure", 1)[-1].split("## O -", 1)[0]
paths = set()
for item in re.split(r"^- ", s, flags=re.M):
item = " ".join(item.split())
if item.startswith("new:"):
for base, names, single in re.findall(PAT, item):
paths |= {single} if single else {f"{base}/{n}.py" for n in names.split(",")}
return sorted(paths)
def commits(*paths):
out = subprocess.run(["git", "log", "--reverse", "--format=%H %ct", "--", *paths],
capture_output=True, text=True, check=True).stdout
return [(h, int(t)) for h, t in (line.split() for line in out.splitlines())]
rows = []
for canvas in sorted(Path("specs").glob("*/REASONS.md")):
if canvas.parts[1].startswith("_"):
continue
owned = owned_paths(canvas.read_text(encoding="utf-8"))
c_commits = commits(canvas.as_posix())
s_commits = commits(*owned) if owned else []
if not c_commits or not s_commits:
continue
first = c_commits[0][1] <= s_commits[0][1]
canvas_hashes = {h for h, _ in c_commits}
later = s_commits[1:]
synced = sum(h in canvas_hashes for h, _ in later)
rows.append((canvas.parts[1], first, synced, len(later)))
for name, first, synced, n in rows:
print(f"{name:32} canvas-first={first!s:5} synced {synced}/{n}")
print(f"canvas-first ratio: {sum(r[1] for r in rows)}/{len(rows)}")Read these with care. A pilot of three features is an anecdote, not a statistic, so look for large effects and trends. Every metric can be gamed once it becomes a target: canvas-first is trivially satisfied by committing an empty canvas first, which is why review, not CI, enforces substance. With squash merges, canvas and code land in one commit and canvas-first always reads true; measure it on branch history or from review timestamps instead. And a high intent-error rate early on is good news: it means review is catching design problems before production does.
Failure modes at team scale
- Reviewer bottleneck. Every canvas waits for one senior engineer. Train more canvas reviewers through pairing, and limit senior review to Approach and Safeguards.
- Template ossification. The template grows fields nobody fills. Prune it in the monthly retro.
- Shadow chat work. Developers keep generating in private chats and write canvases afterwards. Low canvas-first ratios are the symptom; fix the cause, usually a canvas process that feels too slow.
- Canvas rot. Features finish, canvases stop being read, drift grows. Archive or sync them when a feature is next touched.
What to do next
- Write a one-paragraph scope policy using the fitness ratings, with a named person who can grant exceptions.
- Create
specs/_shared/with the template, shared Norms and organisation-wide Safeguards, and add CODEOWNERS entries. - Add the pull request template with slip and intent labels.
- Add structural canvas checks and the non-blocking drift warning to CI.
- Record a baseline from your last three features before the pilot starts.
- Pilot on three features with paired canvas reviews and a retro after each.
- Run the metrics script monthly, discuss trends rather than targets, and prune the template and shared files.