Most writing about Structured Prompt-Driven Development (SPDD) stops once the REASONS canvas exists. The method lives or dies in what happens next, over days: generating code a step at a time, reviewing it against the canvas, deciding whether each defect is the model's fault or the canvas's, and keeping canvas and code in agreement when production forces a change at two in the morning. The source article compresses this into one rule: when reality diverges, fix the prompt first, then update the code.
This guide follows that loop through a working week on one feature, outbound order webhooks, whose canvas is built in the companion guide on the REASONS canvas. It shows the generation request, a scope check to run before reading any diff, one model-slip cycle and one intent-error cycle with their diffs, a hotfix synced back into the canvas, and a test that stops canvas numbers and code constants from drifting apart.
The loop, and its two scales
The source article describes synchronisation between canvas and code at two scales. Within an iteration, every generate-review cycle ends with the canvas and code agreeing: a defect is fixed either in the code or in the canvas, never left as a silent difference. Across iterations, the accumulated canvas becomes the starting context for the next enhancement, so the second feature in an area starts from recorded decisions rather than a blank prompt.
The inner loop is short: generate one Operation, review the diff against the canvas, and either commit or classify what is wrong. The classification decides where the fix goes. A model slip means the canvas was right and the code ignored it; fix the code and add a test so the slip cannot recur silently. An intent error means the canvas itself was wrong or incomplete; fix the canvas, get the change reviewed, then regenerate or patch the code.
Tooling: openspdd commands or any assistant
The authors publish a command-line implementation called openspdd (github.com/gszhangwei/open-spdd) that installs the workflow as commands in an AI coding environment. The source article names the commands below; this guide uses the loop, not the tool, so check the repository for current behaviour before relying on any of them.
| Loop step | Command named in the source | What it does |
|---|---|---|
| Generate one step | /spdd-generate | Generates code following Operations, Norms and Safeguards |
| Review | /spdd-code-review | Checks alignment between canvas and code |
| Intent changed | /spdd-prompt-update | Updates the canvas incrementally for a requirement change |
| Code changed first | /spdd-sync | Synchronises code changes back into the canvas |
| API checks | /spdd-api-test | Generates cURL-based test scripts |
Everything below works with a plain assistant session too. What matters is the discipline: bounded context, one step per generation, and a canvas that is the reference for every review.
Monday: generating one Operation
Operations 1 to 3 (models, signing, dispatcher) landed last week. Today is Operation 4, the sender. The request names the canvas, lists context files taken only from Structure, and restricts the model to one Operation. The final instruction matters most: when the canvas is ambiguous, stop and ask. A model told it may ask will surface gaps; one told only to finish will fill them with defaults.
Canvas: specs/order-webhooks/REASONS.md (status: reviewed)
Context files (from Structure only):
src/webhooks/models.py src/webhooks/signing.py src/webhooks/dispatcher.py
src/infra/db.py src/infra/http_client.py
Implement Operation 4 only: sender.py send_due(session_factory, now) -> int.
Follow every Norm. Treat each Safeguard as a hard failure condition.
Do not edit files outside Structure. If the canvas is ambiguous or a
Safeguard cannot be met, stop and list the question instead of guessing.Keep the session narrow. Showing the model the whole repository invites it to tidy unrelated code, and every unrelated line is review cost. Context engineering for agents covers the general case; in SPDD, Structure is the context budget.
Reviewing the diff against the canvas
Before reading the diff line by line, check its shape. The script below lists changed or new source files that Structure does not list as new or changed; dependencies are context, not licence to edit. An unexplained file is the most common early sign that the model wandered, and it is quicker to question one file than to review it.
# scope_check.py - flag changed or new source files that the canvas Structure
# does not list as new or changed. Run it before reading the diff.
import re, subprocess, sys
from pathlib import Path
PAT = r"([\w/]+)/\{([\w,]+)\}\.py|([\w/]+\.py)"
def editable_paths(canvas_text):
s = canvas_text.split("## S - Structure", 1)[1].split("## O -", 1)[0]
paths = set()
for item in re.split(r"^- ", s, flags=re.M):
item = " ".join(item.split())
if item.startswith(("new:", "changed:")):
for base, names, single in re.findall(PAT, item):
paths |= {single} if single else {f"{base}/{n}.py" for n in names.split(",")}
return paths
def git(*args):
return subprocess.run(["git", *args], capture_output=True, text=True,
check=True).stdout.split()
allowed = editable_paths(Path(sys.argv[1]).read_text(encoding="utf-8"))
changed = set(git("diff", "--name-only", "HEAD"))
changed |= set(git("ls-files", "--others", "--exclude-standard")) # new files
outside = sorted(f for f in changed if f.endswith(".py")
and not f.startswith("tests/") and f not in allowed)
for f in outside:
print(f"outside Structure: {f}")
sys.exit(1 if outside else 0)Then read the diff with the canvas open beside it, in canvas order: does the code meet each relevant Done-when line, use the Entities' names, follow the Approach, respect each Norm, and violate no Safeguard? Reviewing against a written reference is faster and more consistent than reviewing against memory of a conversation. The general checklist in reviewing AI-generated code still applies for things the canvas does not cover, such as invented APIs and weakened assertions.
Cycle 1: a model slip
The first sender diff looked plausible and passed the generated tests. Review against the Safeguards found two violations. It opened one transaction, claimed fifty deliveries, and made fifty HTTP calls inside it, so a batch of slow endpoints would hold row locks and a pooled connection for up to four minutes. And it logged the response body, which the logging Safeguard forbids.
The canvas stated both rules plainly, so this is a slip, not an intent error. The fix goes in the code: claim and commit in one short transaction, make the calls with no transaction open, and record each result in its own transaction.
def send_due(session_factory, now):
- with session_factory.begin() as s:
- for d in claim_due(s, now, limit=50):
- resp = http_client.post(d.url, body=d.body, timeout=5)
- log.info("attempt", delivery_id=d.id, body=resp.text)
- record_result(s, d, resp.status_code, now)
+ with session_factory.begin() as s: # txn 1: claim, then commit
+ batch = [d.snapshot() for d in claim_due(s, now, limit=50)]
+ for d in batch: # no transaction open here
+ result = post_signed(d, timeout_s=TIMEOUT_S, follow_redirects=False)
+ with session_factory.begin() as s: # txn 2: record one result
+ record_result(s, d.id, result, now)
+ log.info("attempt", delivery_id=d.id, subscription_id=d.subscription_id,
+ attempt=d.attempt, status_code=result.status_code,
+ duration_ms=result.duration_ms)The second half of a slip fix is a test, because a Safeguard enforced only by review will be broken again by the next regeneration. The test makes the fake HTTP client fail if a transaction is open when it is called.
def test_no_transaction_open_during_http(db, fake_http):
def on_post(*args, **kwargs):
assert not db.in_transaction(), "HTTP call made inside a transaction"
return FakeResponse(status_code=200)
fake_http.on_post = on_post
seed_due_delivery(db)
assert send_due(db.session_factory, now=T0) == 1
Cycle 2: an intent error
On Wednesday support reported early-access customers who had deleted endpoints and were getting 410 Gone responses retried for almost fifteen hours. The code did exactly what the canvas said: any non-2xx is retried. Nothing slipped; the requirement was incomplete. This is an intent error, and the golden rule applies: fix the canvas first.
## R - Requirements
- non-2xx or timeout is retried after 1m, 5m, 30m, 2h, 12h, then failed
+- 410 Gone disables the subscription immediately; no retry
+- other 4xx except 408 and 429 fail the delivery without retrying and
+ count towards the 50 consecutive failures
+ (decision: product, 2026-09-23, support ticket volume from deleted endpoints)The canvas change went to the product owner for review on its own, a five-line diff in plain language, and was approved the same day. Only then was Operation 4 regenerated with the updated canvas, and Operation 7 extended with tests for 410, 404 and 429. Recording who decided and why inside the canvas means the next person who wonders why a 404 is not retried has an answer in the file.
Skipping the canvas step is tempting because the code change is two lines. The cost comes later: the next regeneration of the sender, from the old canvas, silently reintroduces fifteen hours of pointless retries.
Thursday: syncing a hotfix back
At 02:00 an on-call engineer raised the per-attempt timeout from 5 to 10 seconds because a large customer's endpoint had slowed during their own incident. Hotfixes are rightly not written canvas-first. But the canvas now says 5 seconds and the code says 10, and a regeneration would quietly revert the fix.
The next morning the team did the sync explicitly: they asked whether 10 seconds was the new intent or a temporary accommodation. They kept 5 seconds as the default and added a per-subscription override capped at 10, which is a real requirement change, so it went through the canvas and a small new Operation. To catch the next drift automatically they bound the canvas numbers to the code with a test.
# tests/test_canvas_constants.py - numbers stated in the canvas must match code.
import re
from pathlib import Path
from src.webhooks import sender
CANVAS = Path("specs/order-webhooks/REASONS.md").read_text(encoding="utf-8")
UNIT = {"m": 60, "h": 3600}
def test_retry_schedule_matches_canvas():
line = re.search(r"retried after ([\dmh, ]+), then failed", CANVAS).group(1)
expected = [int(x[:-1]) * UNIT[x[-1]] for x in line.replace(" ", "").split(",")]
assert sender.RETRY_SCHEDULE_S == expected
def test_timeout_matches_canvas():
secs = int(re.search(r"(\d+) s total timeout per attempt", CANVAS).group(1))
assert sender.TIMEOUT_S == secsThe test is deliberately narrow: it checks numbers a regex can find reliably, not prose. On hotfix branches the team runs it as a warning rather than a blocker, and opens a follow-up to reconcile within a working day.
Committing: one pull request, two diffs
Each merged change carries the canvas diff and the code diff together, so history shows intent and implementation moving as a pair. Put the canvas diff first in the pull request description and ask reviewers to read it first: it is shorter, and a disagreement there is cheaper than one about code. Mention the classification in the description, slip or intent error, because across many pull requests that label becomes one of the most useful signals a team has about its canvases. The general habits of reviewing a pull request apply to both halves.
Across iterations: the canvas as a starting point
Three weeks later the team is asked for webhook replay, which the original canvas lists as out of scope. The new work does not start from a blank prompt. The existing canvas already records the entities, the at-least-once approach, the SSRF and logging Safeguards and the metric names; replay inherits all of them. The new canvas references the old one for shared Entities and Safeguards and adds only what is new, and the out-of-scope line in the original is removed in the same pull request.
Failure modes of the loop
- Whole-file regeneration. Regenerating a file for a small change overwrites hand fixes. Regenerate the smallest unit and review the diff, not the file.
- Classification by convenience. Everything gets called a slip because fixing code is quicker. If the same kind of slip appears twice, the canvas is probably ambiguous.
- Canvas bloat. Every incident adds a line and nothing is removed. Prune lines the code and tests already enforce.
- Large Operations. A diff over a few hundred lines gets skimmed. Split the Operation in the canvas, not in your head.
- Chasing determinism. The same canvas will not generate identical code twice. Commit and review generated code; the canvas records intent, not bytes.
The plan-execute-verify structure in agentic coding workflows is the same loop seen from the agent's side.
What to do next
- Write your next generation request naming the canvas, the Structure files, a single Operation and permission to stop and ask.
- Add the scope check to your review routine and look at unexplained files first.
- Label every defect you find as slip or intent error, and write the label in the pull request description.
- For each slip, add a test that fails when the Safeguard or Norm is broken.
- For each intent error, change and approve the canvas before touching the code.
- Bind the numbers in your canvas to code constants with a narrow test.
- After any hotfix, schedule a sync within one working day and decide explicitly whether the change is new intent.