Most teams adopted AI coding assistants one developer at a time. Each person built a private habit: paste some context into a chat, ask for a change, keep part of the output, rewrite the rest. The prompt that produced the code disappeared when the window closed. Months later nobody can say why a module looks the way it does, the next change starts from a blank prompt, and two developers asking for the same kind of change get two different designs.
Structured Prompt-Driven Development (SPDD) is a method for fixing that. It was described by Wei Zhang and Jessie Jie Xia in an article on martinfowler.com published on 28 April 2026, based on how Thoughtworks' internal IT teams work. Its central move is to treat the prompt that generates code as a first-class delivery artifact: written against a fixed structure, committed next to the code, reviewed in pull requests, and kept in sync when either side changes. This guide explains the method from first principles, works through a complete example with a real prompt file, and covers tooling, failure modes and trade-offs. Versioning of runtime prompts inside LLM applications is a different problem, covered in Prompt Versioning.
Why chat-driven coding does not scale to a team
A chat session mixes four things that engineering normally keeps apart: requirements, design decisions, constraints and implementation. When the chat ends, only the implementation survives, as a diff. Reviewers see four hundred lines of plausible code but not the instruction that produced them, so they cannot tell whether a surprising choice was intended or an accident of phrasing. The next developer to touch the feature has no record of the constraints that shaped it, and will happily ask the model for a change that breaks one.
There is also a consistency problem: two developers asking for similar features in different words get different naming, error handling and logging, because each prompt left those decisions to the model. None of this is new; it is what happens whenever intent is not written down. What is new is the rate at which an assistant produces code whose reasons nobody recorded.
The core idea: the prompt is an artifact
SPDD treats the structured prompt as the place where intent lives. The prompt captures requirements, domain language, design intent, constraints and a task breakdown; the model then generates code within that boundary, so output becomes more predictable and easier to check. Because the prompt is a file, it gets the same discipline as code: commit history, review and quality gates. And because code and prompt can drift, the method has one rule for resolving conflicts, stated in the original article as 'when reality diverges, fix the prompt first, then update the code'.
The authors name three skills a developer needs. Abstraction first: think about entities, boundaries and responsibilities before asking for any code, because a model fills every gap you leave with its own defaults. Alignment: make sure the prompt, the stakeholders' intent and the existing system agree before generating, which is mostly a matter of asking clarifying questions early. Iterative review: review the structured prompt and the generated code in small increments rather than accepting a large output in one step.
None of this assumes the model is deterministic. The same canvas will not produce identical code twice, so generated code is still reviewed and committed; the prompt is the durable statement of what that code is supposed to do.
The REASONS canvas
SPDD gives the prompt a fixed seven-part structure called the REASONS canvas. Each letter is a section heading, and each section answers a question a reviewer would otherwise have to guess at.
- Requirements: the problem and the definition of done. Write acceptance criteria that a test could check.
- Entities: the domain objects involved and how they relate, using the names the codebase already uses.
- Approach: the strategy for meeting the requirements, including the alternative you rejected and why.
- Structure: where the change fits in the system: modules, components, dependencies, registration points.
- Operations: concrete, testable implementation steps, numbered so they can be generated and reviewed one at a time.
- Norms: the engineering standards that apply, such as naming, logging, observability and defensive coding.
- Safeguards: non-negotiable boundaries: invariants, performance budgets, security rules, things the change must not touch.
The first four sections say what and where, Operations says how, and Norms and Safeguards say how not. The last two do the most work, because a model cannot infer from a story that this team fails open on cache outages or has a two-millisecond latency budget. Norms shared by every feature eventually move into a project instruction file.
The workflow, step by step
The original article describes six steps. First, create the initial requirements, usually a user story. Second, clarify: the developer, often with the model's help, lists open questions and gets answers from the people who own them. Third, generate an analysis context: the relevant existing code, domain terms and constraints, gathered into one place. Fourth, generate the structured prompt, the REASONS canvas, from the requirements and analysis. Fifth, generate code from the canvas. Sixth, generate unit tests.
Two details make the steps work. The canvas is drafted with model help but reviewed by a person before any code is generated, the cheapest point to catch a wrong design. And generation follows the numbered Operations a few at a time, so each diff is small enough to review.
Worked example: per-tenant rate limiting
The original article uses a billing engine gaining model-aware pricing and multiple subscription plans. Here is a smaller example you can reproduce. A multi-tenant orders API needs a rate limit on its write endpoints, because one tenant's batch job has been starving everyone else. The story is one sentence. Clarification turns up four questions the story did not answer: is the limit per tenant or per API key (per tenant); where is state kept (the existing Redis); what happens when Redis is down (fail open, because an outage of the limiter must not become an outage of the API); and what does the client see (HTTP 429 with a Retry-After header).
Those answers become the canvas below. Much of it is not about code: the rejected alternative in Approach, the fail-open rule in Safeguards, the metric in Norms. Without them the model picks its own answers and a reviewer must spot each one in the diff.
# REASONS canvas: per-tenant rate limiting for write endpoints
# file: specs/rate-limit/REASONS.md (reviewed like code; owner: platform team)
## R - Requirements
Story: As a platform operator, I want each tenant limited on write endpoints
(POST/PUT/DELETE under /v1/orders) so one tenant cannot starve the others.
Done when:
- default limit 120 requests per 60 s per tenant, configurable per tenant
- over-limit requests get HTTP 429 with a Retry-After header in seconds
- read endpoints are unaffected
## E - Entities
- Tenant(id, plan) existing, src/tenants/models.py
- RateLimitPolicy(tenant_id, limit, window_s) new, loaded from config
- Decision(allowed: bool, retry_after_s: int) new, value object
## A - Approach
Fixed-window counter in Redis: key rl:{tenant}:{window_start}, INCR then
EXPIRE on first hit. Chosen over token bucket for simplicity; bursts of up to
2x limit at window edges are accepted (see Safeguards).
## S - Structure
- new module src/ratelimit/ (policy.py, limiter.py, middleware.py)
- middleware registered in src/app.py after auth (needs tenant id)
- depends on existing src/infra/redis_client.py; no new dependencies
## O - Operations
1. policy.py: load_policies(config) -> dict[tenant_id, RateLimitPolicy]
2. limiter.py: check(tenant_id, now) -> Decision, using one INCR + EXPIRE
3. middleware.py: apply only to write methods under /v1/orders
4. app.py: register middleware after auth
5. tests: unit tests for window rollover, 429 path, read bypass, Redis down
## N - Norms
- type hints everywhere; no module-level mutable state
- log one structured line per rejection: tenant_id, limit, window_start
- expose counter rl_rejections_total{tenant} via the existing metrics module
## S - Safeguards
- Redis unavailable: FAIL OPEN, log at WARNING, increment rl_redis_errors_total
- never read or log request bodies
- added latency p99 under 2 ms at current load; one Redis round trip max
- do not change auth, routing or any read endpointWith the canvas reviewed, Operations 1 to 4 are generated one at a time, and Operation 5 derives tests from Requirements and Safeguards, including a Redis-down test asserting the request is allowed and the error counter increments. Safeguards double as a test plan.
Generating code inside the boundary
The mechanics can be simple. The driver below lints the canvas, collects only the files named in Structure as context, and asks for one Operation at a time. The lint matters: a canvas missing its definition of done is how SPDD degrades back into ad hoc prompting.
# Generate code one Operation at a time, inside the canvas boundary.
# call_model() stands for whatever assistant or API your team uses.
from pathlib import Path
import re, subprocess
CANVAS = Path("specs/rate-limit/REASONS.md")
def sections(md):
parts = re.split(r"^## ([A-Z]) - (\w+)", md, flags=re.M)
return {parts[i + 1]: parts[i + 2].strip() for i in range(1, len(parts) - 2, 3)}
def lint(md):
s = sections(md)
missing = [k for k in ["Requirements", "Entities", "Approach", "Structure",
"Operations", "Norms", "Safeguards"] if k not in s]
ops = re.findall(r"^\d+\. ", s.get("Operations", ""), flags=re.M)
problems = [f"missing section: {m}" for m in missing]
if not ops:
problems.append("Operations must be a numbered list")
if "Done when" not in s.get("Requirements", ""):
problems.append("Requirements need a definition of done")
return problems
def generate(op_number):
md = CANVAS.read_text()
if problems := lint(md):
raise SystemExit("\n".join(problems))
s = sections(md)
files = re.findall(r"src/[\w/]+\.py", s["Structure"])
context = "\n\n".join(Path(f).read_text() for f in files if Path(f).exists())
prompt = (f"{md}\n\nExisting code:\n{context}\n\n"
f"Implement Operation {op_number} only. Follow Norms. "
f"Violating a Safeguard is a failure; say so instead of guessing.")
patch = call_model(prompt)
Path(f"out/op{op_number}.patch").write_text(patch)
subprocess.run(["git", "apply", "--check", f"out/op{op_number}.patch"], check=True)Narrow context is deliberate: the whole repository invites the model to 'improve' unrelated code. If generation needs a file the canvas did not name, fix the Structure section. For broader context strategies see Context Engineering for Agents.
Keeping prompt and code in sync
The golden rule turns every defect into a classification question. If the code is wrong because the intent was wrong or incomplete, for example the team later decides to fail closed for one premium tenant, fix the canvas first, get it reviewed, then regenerate or hand-edit the affected Operation. If the code is wrong but the canvas was right, the model slipped: fix the code, and if the same slip recurs, add a Norm or Safeguard so it cannot recur silently. In both cases the canvas and the code land in the same pull request, so history shows them changing together.
In the repository this means a specs/ directory with one folder per feature holding the story, the clarification notes and the canvas, owned by the same reviewers as src/.
Reviewers read the canvas diff first: it is shorter, states intent in plain language, and disagreements there are cheaper than disagreements about generated code. Treat a code change with no matching canvas change as a question to ask, not an automatic rejection: many small fixes do not change intent. The review habits in How to Review a PR apply unchanged to the canvas; it is closer to a compact design doc than to a prompt.
Failure modes
- Canvas theatre. The canvas is generated after the code to satisfy a process. It then records nothing a reviewer could have used. Enforce the order: canvas reviewed before code.
- Silent drift. Hotfixes change behaviour and nobody updates the canvas, so the next regeneration reintroduces the bug. Add a PR checklist item and periodically diff behaviour against Safeguards.
- Vague Safeguards. 'Must be fast' and 'must be secure' constrain nothing. Every Safeguard should be something a test or a reviewer can verify.
- Oversized Operations. One Operation that says 'implement the feature' produces a diff too big to review. Split until each step is a reviewable change.
When SPDD is the wrong tool, and the trade-offs
The method front-loads design work, and that is not always worth it. The original article rates it poorly for domains where the context is unclear, for purely creative work, for firefighting hotfixes, for exploratory spikes and for one-off scripts. In all of those the cost of writing a careful canvas exceeds the value of preserving it. It also asks for senior judgement up front, because a weak canvas produces confidently wrong code, and for a shift from code first to design first that some teams resist.
Where it fits, on maintained features in established domains, a few hundred words of reviewed intent buys reviewable diffs, consistent style and a history of why code looks the way it does. For the pattern applied to runtime prompts rather than code-generating ones, see Prompt Registry.
What to do next
- Pick one upcoming feature in a well-understood domain and write its REASONS canvas before any code; time how long the canvas takes.
- Create a specs/ directory with the same code ownership as src/, and require the canvas and the code in one pull request.
- Add the canvas linter above (or your own) to CI so missing Requirements, Operations or Safeguards fail the build.
- Generate one Operation at a time and review each diff against the canvas, not against your memory of the story.
- When a bug appears, classify it as intent error or model slip before fixing, and record which you found.
- After three features, move repeated Norms into a shared instruction file and prune them from the canvases.
- Decide explicitly which work is out of scope for SPDD (spikes, hotfixes, scripts) so the method does not become ceremony.