Structured Prompt-Driven Development (SPDD), described by Wei Zhang and Jessie Jie Xia on martinfowler.com in April 2026, asks developers to write a structured prompt called the REASONS canvas before generating code, and to commit it next to the code it produced. Knowing the seven section names is easy. Writing sections that actually constrain a model, and that a reviewer can check in five minutes, is a skill; most first canvases are either too vague to steer anything or too long to read.

This guide takes the canvas one section at a time: the question each answers, weak and strong versions, a sensible length, and what a model does when the section is missing. It ends with a blank template and a complete canvas for a realistic feature, outbound order webhooks. If SPDD is new to you, start with the SPDD overview; companion guides cover the daily loop and team adoption.

Advertisement

What the canvas is for, and who reads it

The canvas has three readers. The reviewer reads it before any code exists and decides whether the design is right, the cheapest point to catch a wrong decision. The model reads it as instructions and generates code inside the boundary it draws. The maintainer reads it months later to learn why the code looks the way it does, which a diff never says.

The source article groups the sections into three bands. Requirements, Entities, Approach and Structure are abstract: what is built and where it fits. Operations is specific: concrete, testable steps. Norms and Safeguards are governance: the standards and hard limits the code must respect. The test for any line is whether it removes a decision the model would otherwise make itself. A line restating the obvious costs reading time; a missing line hands the decision to the model's defaults, which suit a tutorial and rarely suit your system.

REASONS canvas: three bands, three readersABSTRACT: what and whereR Requirementswhat does done mean?E Entitieswhich nouns, which names?A Approachwhich strategy, what rejected?S Structurewhich files and seams?SPECIFIC: howO Operationsnumbered, testable stepsone step = one reviewable diffGOVERNANCE: how notN Normsteam standardsS Safeguardshard limits, testableReviewerreads R, A, S firstModelexecutes O inside N and SMaintainerreads A and S for the whyA missing section hands that decision to the model's defaults
The seven sections in three bands. Operations is the only section executed step by step; the other six bound what each step may do.

The worked feature

The example is a multi-tenant orders service that must push order events to customers' HTTPS endpoints so integrators can stop polling. Clarification with the product owner and security team answered what the one-line story left open: which events (created, paid, cancelled), how payloads are authenticated (an HMAC over a timestamp and the raw body), the retry schedule (five retries from one minute to twelve hours), what happens to a dead endpoint (disable after fifty consecutive failures and email the admin), whether ordering is guaranteed (no), and whether a tenant may target an internal address (never).

Each answer appears in the canvas below, and each is a place an unconstrained model would have guessed. Common defaults are sending from the request handler, retry loops with sleep calls and no validation of the target URL: exactly what most example code online does.

Advertisement

R: Requirements

Answers: how will we know it is done? The weak version is a story and a hope: send webhooks reliably when orders change. The strong version adds a Done when list where every line is observable and carries its numbers (events, headers, the exact signature construction, retry offsets, the disable threshold) and an Out of scope list naming what a reader would assume is included, here ordering and replay.

Five to twelve lines is typical. When Requirements is vague the model invents the easy acceptance criteria and ships a sender with no retries, because nobody said retries were required. Without Out of scope it may add a replay endpoint nobody asked for, which then needs review, tests and security thought.

E: Entities

Answers: which nouns does this touch, and what are they called here? List each entity, whether it exists or is new, where existing ones live, and only the fields that matter, then relationships and uniqueness in a line or two.

The uniqueness line on Delivery (subscription plus event) is what makes the dispatcher safe to rerun; omit it and the generated fan-out duplicates deliveries after every restart. Missing Entities also causes naming drift: the model invents WebhookEndpoint beside your existing vocabulary, and the codebase gains two words for one idea. Four to ten lines.

A: Approach

Answers: what strategy, and what did we reject? Write the strategy in two to five sentences, one Rejected line per serious alternative with its reason, and the consequence you accept. Rejected lines are the most valuable in the canvas for the maintainer: they stop the next person, or the next model session, from proposing what you already ruled out.

Here the approach reuses the existing transactional outbox. Sending from the request handler is rejected because a crash loses events; the broker is rejected because this service has none. The accepted consequence is at-least-once delivery, hence the event id header for deduplication. Without this section the model picks a strategy and the reviewer reverse-engineers it from the diff.

S: Structure

Answers: where does this go? Name new and changed files, the registration points where new code is wired in, the modules to reuse, and whether new packages are allowed. Structure doubles as the generation context list: the files named here are what the model is shown, and a diff touching other files is a review signal.

The weak version says add a webhooks module. Without the strong version, models add a second HTTP library, write their own encryption helper instead of the existing one, and register workers somewhere nobody will look. Four to eight lines.

O: Operations

Answers: in what steps is it built? Each Operation names a file, a signature and a behaviour, and is small enough that its diff can be reviewed in one sitting. Signatures fix interfaces between steps: once Operation 3 says fan_out(session, batch=500) -> int, later steps need not re-decide it.

The last Operation is always tests derived from Done-when and Safeguards, so the canvas is also a test plan. The weak version, implement the backend; add tests, yields one unreviewable diff; twenty three-line steps turn generation into dictation. Five to eight suits a feature this size.

N: Norms

Answers: which team standards apply? Logging fields, metric names, time handling, error conventions. Because they would apply to any similar feature, Norms that recur across canvases belong in the shared file your assistant loads for every task, covered in instruction files for coding agents. What stays is feature-specific, such as this feature's metric names.

Add good logging is weak; the five fields of the per-attempt log line are strong. Without Norms each feature logs differently and dashboards cannot be reused.

S: Safeguards

Answers: what must never happen, however the code is written? Every line must be checkable by a test or reviewer. Two of the webhook Safeguards show why. Reject private, loopback and link-local addresses at subscribe time and at send time: without it the feature is a server-side request forgery hole letting a tenant make your servers call internal services, and the send-time check exists because DNS answers change after validation. No transaction held open during an outbound HTTP call: models often wrap claim, POST and update in one transaction, holding row locks for the whole timeout and draining the connection pool under load.

Must be secure and must be fast constrain nothing, because nobody can show they were broken. Four to eight lines; fifteen means some are Norms in disguise. Have someone outside the feature team review this section; securing AI-assisted development is a good source of questions.

The blank template

Copy this into specs/<feature>/REASONS.md. The header links the canvas to the story and analysis notes from the first SPDD steps, and the status field lets tooling refuse to generate from an unreviewed canvas.

# REASONS canvas: <feature name>
# file: specs/<feature>/REASONS.md   owner: <team>   status: draft | reviewed
# story: specs/<feature>/story.md    analysis: specs/<feature>/analysis.md

## R - Requirements
Story: As a <role>, I want <capability> so that <outcome>.
Done when:
- <observable behaviour a test can check, with numbers>
Out of scope:
- <things a reasonable reader would assume are included>

## E - Entities
- <Name> (existing|new, path): fields that matter
Relationships: <cardinality and uniqueness rules>

## A - Approach
<strategy in 2-5 sentences>
Rejected: <alternative> because <reason>
Accepted consequence: <cost we knowingly take>

## S - Structure
- new: <paths>   changed: <paths, registration point>
- depends on: <existing modules>; new packages: none | <name, why>

## O - Operations
1. <file>: <signature> - <behaviour>
n. tests derived from Requirements and Safeguards

## N - Norms
- <logging / metrics / naming / error rule for this feature>

## S - Safeguards
- <invariant, limit or prohibition a test or reviewer can check>

The complete canvas

The finished webhook canvas is under 500 words: five minutes for a reviewer, and easy for a model to hold beside the named source files. Much of it is not about code: rejected alternatives, the accepted consequence, the out-of-scope list and the Safeguards are all decisions a diff would hide.

# REASONS canvas: outbound order webhooks
# file: specs/order-webhooks/REASONS.md   owner: integrations   status: reviewed

## R - Requirements
Story: As an integrator, I want order events pushed to my HTTPS endpoint so
my system stays in sync without polling.
Done when:
- order.created, order.paid, order.cancelled are POSTed to every active
  subscription of the tenant, p95 within 10 s of the order commit
- headers X-Event-Id, X-Timestamp, X-Signature, where X-Signature =
  hex HMAC-SHA256(secret, timestamp + "." + raw_body)
- non-2xx or timeout is retried after 1m, 5m, 30m, 2h, 12h, then failed
- 50 consecutive failed deliveries disable the subscription and email
  the tenant admin
- GET /v1/webhooks/{id}/deliveries lists attempts, newest first
Out of scope: ordering across events; custom retry schedules; replay

## E - Entities
- OrderEvent (existing, src/orders/events.py): id, type, tenant_id, payload
- WebhookSubscription (new): id, tenant_id, url, secret_enc, event_types,
  status active|disabled, consecutive_failures
- Delivery (new): id, subscription_id, event_id, attempt, next_attempt_at,
  status pending|succeeded|failed, last_status_code
Relationships: Delivery unique on (subscription_id, event_id).

## A - Approach
Reuse the existing outbox: orders already insert OrderEvent in the order
transaction. A dispatcher fans events out into Delivery rows; a sender
claims due rows with SELECT ... FOR UPDATE SKIP LOCKED and POSTs them.
Rejected: sending from the request handler (lost on crash, slower writes);
the message broker (not deployed for this service).
Accepted consequence: at-least-once; receivers dedupe on X-Event-Id.

## S - Structure
- new: src/webhooks/{models,signing,dispatcher,sender,api}.py,
  migrations/0042_webhooks.py
- changed: src/workers/registry.py, src/app.py (mount /v1/webhooks)
- depends on: src/infra/{db,http_client,crypto}.py,
  src/notifications/email.py; new packages: none

## O - Operations
1. migration + models, with the unique constraint
2. signing.py: sign(secret: bytes, ts: int, body: bytes) -> str
3. dispatcher.py: fan_out(session, batch=500) -> int; safe to rerun
4. sender.py: send_due(session_factory, now) -> int; claim, POST,
   record, schedule next attempt from RETRY_SCHEDULE_S (seconds);
   per-attempt timeout TIMEOUT_S
5. sender.py: disable at 50 consecutive failures; queue email
6. api.py: tenant-scoped subscription CRUD; GET deliveries
7. tests for every Done-when line and every Safeguard

## N - Norms
- one log line per attempt: delivery_id, subscription_id, attempt,
  status_code, duration_ms
- metrics: webhook_attempts_total{outcome}, webhook_delivery_lag_seconds
- outbound HTTP failures are returned as values, never raised out of the loop

## S - Safeguards
- never log payload bodies, secrets or signatures
- https only; reject private, loopback and link-local addresses at
  subscribe time AND at send time
- no database transaction held open during an outbound HTTP call
- 5 s total timeout per attempt; redirects not followed
- no tenant can read or change another tenant's subscriptions

Reviewing the canvas before generating anything

Canvas review is where SPDD pays for itself. Split it by expertise: the product owner reviews Requirements and Out of scope, whoever owns security and reliability reviews Safeguards, and the implementing engineer reviews Operations for size and order. Ask per section: can each Done-when line fail a test; is every entity name already in the code or marked new; is a rejected alternative given; is every path real; does a test step cover every Safeguard?

Do not automate the judgement. A structural check that all seven headings exist is useful in CI; a regex that grades testability produces canvases written to pass the regex. The habits from reviewing a design doc transfer almost unchanged, since a canvas is a compact design doc with a stricter shape.

Trade-offs and common mistakes

MistakeSymptomFix
Canvas written after the codeIt describes the diff, accidents includedApprove the canvas before first generation
Everything in NormsCanvas over 1,500 words; reviewers skimPromote shared rules to the instruction file
No rejected optionNext session re-proposes itAlways write one Rejected line
Safeguards as valuesNo test fails when one breaksRewrite until a test can check it

The real cost is time: a canvas like this takes an experienced engineer about an hour, most of it clarification you would otherwise do later and more expensively. For a spike or one-off script that hour is wasted, which is why the source rates SPDD poorly there.

What to do next

  1. Copy the blank template into a specs/ directory in a repository you own.
  2. For your next feature, write Requirements with Done-when and Out of scope lists first.
  3. Add one Rejected line and one accepted consequence to Approach.
  4. Name exact files and registration points in Structure and use only those as generation context.
  5. Rewrite each Safeguard until you can name the test that catches its violation, and make those tests the last Operation.
  6. Review the canvas with a second person before generating code, and count how many decisions changed.
  7. After three canvases, move Norms that appear in all of them into your shared instruction file.
Key takeaway: A REASONS canvas is only as useful as its weakest section. Requirements need observable Done-when lines and an Out of scope list; Entities need real names and uniqueness rules; Approach needs a rejected alternative and an accepted cost; Structure needs exact files; Operations need signatures and reviewable size; Norms should shrink into the shared instruction file; and every Safeguard must be testable. Review the canvas before generating anything, because that is the cheapest point to change a design.