One developer can adopt Structured Prompt-Driven Development (SPDD) in an afternoon: write a REASONS canvas, generate from it, commit both. A team cannot, because the value of SPDD is shared: canvases other people can review, Norms that apply to everyone, Safeguards owned by the people who understand the risks, and a history that explains the code to whoever touches it next. Those need repository conventions, review rules and a way to tell whether the effort is paying off.

This guide covers that team layer: where to apply the method and where not to, how to lay out the prompt repository, who owns what, what reviewers and CI should check, how to govern shared rules, how to roll it out, and how to measure it with numbers your own history can produce. The method itself is in the SPDD overview; companion guides go deep on writing the canvas and on the daily loop.

Advertisement

Decide where SPDD applies before anything else

The fastest way to make a team hate a method is to apply it to everything. The source article, by Wei Zhang and Jessie Jie Xia on martinfowler.com, rates fitness explicitly, and those ratings are a good starting policy.

WorkSource ratingTeam policy
Scaled, standardised delivery5 starsCanvas required
High compliance or hard constraints5 starsCanvas required; security reviews Safeguards
Collaboration needing auditability5 starsCanvas required
Firefighting hotfixes2 starsNo canvas up front; sync within a working day
Exploratory spikes, one-off scripts2 starsNo canvas; write one if the spike becomes a feature
Poorly defined domains, creative or visual work1 starNo canvas; clarify the domain first

Write the policy down in the repository, in one paragraph, and decide who can grant an exception. Without an explicit out-of-scope list every small change becomes an argument about whether it needs a canvas, and the method loses by attrition.

The prompt repository

Keep canvases in the same repository as the code they generate, not in a wiki or a separate prompt store. The point is that a single pull request changes intent and implementation together, and that versioning comes from git rather than a second system. (Runtime prompts shipped inside an LLM application are a different problem with different tooling.) One directory per feature holds the story with its dated clarification answers, the analysis notes and the canvas. A shared directory holds the template and team-wide rules.

specs/
  _shared/
    TEMPLATE.md            # blank REASONS canvas, versioned
    norms.md               # team-wide Norms (also loaded by the instruction file)
    safeguards.md          # organisation-wide Safeguards: auth, PII, egress
  order-webhooks/
    story.md               # the user story and clarification Q&A, dated
    analysis.md            # domain terms, relevant code, constraints
    REASONS.md             # the canvas; header has owner and status
  order-webhooks-replay/
    REASONS.md             # references ../order-webhooks/REASONS.md
  _archive/                # canvases for removed features, kept for history
src/ ...
tests/ ...

Two conventions keep this manageable. A follow-on feature gets its own canvas that references the original rather than growing it, so each canvas stays reviewable. And canvases for deleted features move to an archive rather than being removed, because the reason a feature was built is useful long after its code is gone.

Advertisement

Ownership and shared rules

CODEOWNERS works on paths, not sections, so ownership follows the layout. Feature teams own their feature directories. Engineering leads own the template and the shared Norms. The security team owns organisation-wide Safeguards, such as rules on authentication, personal data and outbound network access, in their own file that every canvas references instead of copying.

# .github/CODEOWNERS - later rules win for a path
/specs/                      @acme/eng-leads
/specs/_shared/norms.md      @acme/eng-leads @acme/platform
/specs/_shared/safeguards.md @acme/security
/specs/order-webhooks*/      @acme/integrations
/src/webhooks/               @acme/integrations

The shared Norms file is also what your assistant's project instruction file should load, so the rule is written once and applied in every session, canvas or not; instruction files for coding agents covers that mechanism. A Norm that appears in three feature canvases is a candidate for promotion. Changes to shared files go through the same pull request review as code, because a one-line edit to shared Safeguards changes the boundary for every future generation.

Review rules

Three rules carry most of the weight. First, the canvas is approved before code is generated from it; the canvas header's status field records this, and a canvas generated after the fact is worse than none because it looks like evidence. Second, reviewers read the canvas diff before the code diff. Third, every defect found in review is labelled as a model slip or an intent error, because that label drives both the fix and the metrics later.

## Canvas
- [ ] canvas diff included, or: no intent change (say why)
- [ ] canvas status is "reviewed" and was approved before code was generated
- Canvas: specs/<feature>/REASONS.md   Operations covered: <n, n>

## Defects found in review (label the PR)
- [ ] spdd:slip    - canvas right, code wrong; test added: <name>
- [ ] spdd:intent  - canvas wrong; canvas fixed first in this PR

## Safeguards
- [ ] every Safeguard touched by this change has a test

Canvas review needs different reviewers for different sections: the product owner for Requirements, a security or reliability owner for Safeguards, a senior engineer for Approach and Structure. Pairing a junior engineer with a senior one on canvas review is the fastest way to teach the three skills the source names: abstraction first, alignment and iterative review. Code review then checks the diff against the canvas, using the habits in reviewing AI-generated code.

What CI should check, and what it should not

Automate the structural checks and leave judgement to people. Useful checks: every canvas has all seven sections, a status and an owner; canvases with status draft cannot be referenced by a merged pull request; tests that bind canvas numbers to code constants pass. And a drift warning: when a pull request changes files a canvas lists as new in Structure, the files that feature owns, but does not change that canvas, CI posts a warning asking the author to confirm there is no intent change.

# drift_check.py BASE HEAD - warn when files a canvas's Structure lists as
# new (files the feature owns) change but that canvas does not. Shared files
# such as src/app.py are skipped: they change for unrelated reasons.
# Prints warnings; never fails the build.
import re, subprocess, sys
from pathlib import Path

def git(*args):
    return subprocess.run(["git", *args], capture_output=True, text=True,
                          check=True).stdout

def owned_paths(canvas_text):
    s = canvas_text.split("## S - Structure", 1)[-1].split("## O -", 1)[0]
    paths = set()
    for item in re.split(r"^- ", s, flags=re.M):
        item = " ".join(item.split())
        if item.startswith("new:"):
            for base, names, single in re.findall(r"([\w/]+)/\{([\w,]+)\}\.py|([\w/]+\.py)", item):
                paths |= {single} if single else {f"{base}/{n}.py" for n in names.split(",")}
    return paths

base, head = sys.argv[1], sys.argv[2]
changed = set(git("diff", "--name-only", f"{base}...{head}").split())
for canvas in Path("specs").glob("*/REASONS.md"):
    if canvas.parts[1].startswith("_"):
        continue
    touched = changed & owned_paths(canvas.read_text(encoding="utf-8"))
    if touched and canvas.as_posix() not in changed:
        print(f"::warning::{canvas.as_posix()} unchanged but {sorted(touched)} changed;"
              " confirm no intent change or run a sync")

Keep the drift check a warning. Plenty of legitimate changes, such as refactors, dependency bumps and small bug fixes, touch code without changing intent. A blocking check teaches people to make meaningless canvas edits to satisfy it, which destroys the signal. The warning's job is to make the decision explicit, not to make it for the author.

Governance: what canvases may contain

  • No secrets or customer data. Canvases are sent to model providers as prompts. Treat them as you would code that leaves the building; securing AI-assisted development covers the wider threat model.
  • Dated decisions. When a requirement changes, record who decided and when, in the canvas. For regulated work this is the audit trail an assessor asks for.
  • A lifecycle. Draft, reviewed, superseded, archived. Tooling should refuse to generate from draft or superseded canvases.
  • A template owner. The template changes by pull request, with a note explaining why, and no more than once a quarter so canvases stay comparable.

Rolling it out

Start with one team and three consecutive features in a well-understood domain, which is where the source's ratings are highest. Before the pilot, record a baseline from the previous three features: review time per pull request, reverts and follow-up fixes within a few weeks of merge, and how often reviewers asked why something was done. Run a canvas-review pairing session for each pilot feature, and hold a short retrospective after each.

Expand only when the pilot team would choose to keep the method without being asked. Each new team gets the template, the shared files and one experienced reviewer for its first canvases. Expect the design-first habit to take longest; the source article is explicit that the shift from code first to design first needs continuing training rather than one workshop.

SPDD governance: from canvas to metricsAuthorcanvas + code, one PRCODEOWNERSfeature, securityCI checksstructure, status, driftMergeintent + codespecs/_shared/norms, safeguards, templatespecs/<feature>/story, analysis, REASONS.mdgit historyPR labels: slip / intentTeam metricscanvas-first, drift, slip:intentMonthly retroprune template and normsupdateCanvases are reviewed like code, measured from history the team already has,and the template changes only through the same review process.
Governance loop: canvases and code merge together through owners and CI; metrics come from git history and PR labels, and the monthly retro changes the template and shared rules through the same review.

Measuring whether it works

The source article reports no adoption data across teams, only an outcome for its own worked example, and industry productivity figures will not tell you whether SPDD works for your team. Define a small set from data you already have and compare against your own baseline:

MetricSourceWhat it tells you
Canvas-first ratiogit historyIs the canvas used to design, or written afterwards?
Sync rategit historyDo canvas and code move together after launch?
Slip to intent-error ratioPR labelsAre defects about the model or about unclear intent?
Review time per PRcode hostDoes reviewing against a canvas speed review up?
Rework within 21 daysgit historyIs generated code correct the first time?
# spdd_metrics.py - two adoption metrics from git history alone, measured on
# the files each canvas's Structure lists as "new" (files the feature owns;
# shared files such as src/app.py have unrelated history).
# canvas-first: the canvas's first commit is not later than the first commit
#               touching any of those files.
# sync rate:    share of later commits touching those files that also touch
#               the canvas. Low is not bad by itself; watch the trend.
import re, subprocess
from pathlib import Path

PAT = r"([\w/]+)/\{([\w,]+)\}\.py|([\w/]+\.py)"

def owned_paths(canvas_text):
    s = canvas_text.split("## S - Structure", 1)[-1].split("## O -", 1)[0]
    paths = set()
    for item in re.split(r"^- ", s, flags=re.M):
        item = " ".join(item.split())
        if item.startswith("new:"):
            for base, names, single in re.findall(PAT, item):
                paths |= {single} if single else {f"{base}/{n}.py" for n in names.split(",")}
    return sorted(paths)

def commits(*paths):
    out = subprocess.run(["git", "log", "--reverse", "--format=%H %ct", "--", *paths],
                         capture_output=True, text=True, check=True).stdout
    return [(h, int(t)) for h, t in (line.split() for line in out.splitlines())]

rows = []
for canvas in sorted(Path("specs").glob("*/REASONS.md")):
    if canvas.parts[1].startswith("_"):
        continue
    owned = owned_paths(canvas.read_text(encoding="utf-8"))
    c_commits = commits(canvas.as_posix())
    s_commits = commits(*owned) if owned else []
    if not c_commits or not s_commits:
        continue
    first = c_commits[0][1] <= s_commits[0][1]
    canvas_hashes = {h for h, _ in c_commits}
    later = s_commits[1:]
    synced = sum(h in canvas_hashes for h, _ in later)
    rows.append((canvas.parts[1], first, synced, len(later)))

for name, first, synced, n in rows:
    print(f"{name:32} canvas-first={first!s:5} synced {synced}/{n}")
print(f"canvas-first ratio: {sum(r[1] for r in rows)}/{len(rows)}")

Read these with care. A pilot of three features is an anecdote, not a statistic, so look for large effects and trends. Every metric can be gamed once it becomes a target: canvas-first is trivially satisfied by committing an empty canvas first, which is why review, not CI, enforces substance. With squash merges, canvas and code land in one commit and canvas-first always reads true; measure it on branch history or from review timestamps instead. And a high intent-error rate early on is good news: it means review is catching design problems before production does.

Failure modes at team scale

  • Reviewer bottleneck. Every canvas waits for one senior engineer. Train more canvas reviewers through pairing, and limit senior review to Approach and Safeguards.
  • Template ossification. The template grows fields nobody fills. Prune it in the monthly retro.
  • Shadow chat work. Developers keep generating in private chats and write canvases afterwards. Low canvas-first ratios are the symptom; fix the cause, usually a canvas process that feels too slow.
  • Canvas rot. Features finish, canvases stop being read, drift grows. Archive or sync them when a feature is next touched.

What to do next

  1. Write a one-paragraph scope policy using the fitness ratings, with a named person who can grant exceptions.
  2. Create specs/_shared/ with the template, shared Norms and organisation-wide Safeguards, and add CODEOWNERS entries.
  3. Add the pull request template with slip and intent labels.
  4. Add structural canvas checks and the non-blocking drift warning to CI.
  5. Record a baseline from your last three features before the pilot starts.
  6. Pilot on three features with paired canvas reviews and a retro after each.
  7. Run the metrics script monthly, discuss trends rather than targets, and prune the template and shared files.
Key takeaway: SPDD becomes a team practice through conventions, not enthusiasm: an explicit policy on where it applies, canvases in the code repository with CODEOWNERS on features and shared rules, reviewers who approve the canvas before code and label every defect as slip or intent error, CI that checks structure and warns on drift without blocking, and governance of what canvases may contain. Measure it against your own baseline with metrics git history already supports, treat them as trends rather than targets, and expand only when the pilot team would keep it unasked.