A design doc is the cheapest place to be wrong. A flawed design caught in a document costs an afternoon of rewriting; the same flaw caught in production costs a migration. Yet many design docs fail at that one job. They describe a solution the author has already committed to, skip the numbers, list straw-man alternatives, and leave out how the change will be rolled out and undone. Reviewers then argue about details because the document never made the important decisions visible.
This article is about the writing, not the template. The anatomy of a design doc (context, goals, non-goals, alternatives, decision records) is covered in design docs architecture, and reviewing someone else's in how to review a design doc. Here we follow one document from blank page to approval: adding per-tenant rate limiting to a public API. Each step shows what to write, in what order, and why that order surfaces mistakes early.
Decide whether you need one
Not every change deserves a design doc, and writing one for a two-day task teaches a team to treat them as bureaucracy. A useful test is whether the change has any of these properties:
- It is hard to reverse: a data model, a public API, a storage format, a vendor commitment.
- It affects teams other than yours, or systems you do not own.
- There are several plausible designs and reasonable engineers would disagree.
- It will take more than about two engineer-weeks, or carries a security, privacy or reliability risk.
If none apply, a ticket with a paragraph of rationale is enough. In the worked example, rate limiting touches every API request, changes customer-visible behaviour (some calls will now return HTTP 429), and has real alternatives, so it qualifies on three counts.
The writing pipeline
Writing the full document first and circulating it last is the most common mistake. The author invests days, becomes attached, and the review becomes a defence. Instead, expose the document in stages that grow in cost: a half-page problem note that one senior colleague can reject in five minutes, a one-pager that tests the direction, then the full draft, then a small review, then a wide one. Each stage is an opportunity to discover that the problem is different from what you thought, which is far more common than discovering a bug in the design.
Step 1: write the problem before the solution
Start with a section that never mentions your solution. State what is going wrong, for whom, with evidence, and what success would look like in measurable terms. Then write goals and non-goals. A non-goal is something a reader might reasonably expect you to do that you are deliberately not doing; it is the main tool for preventing scope creep in review.
For the worked example, a first draft read "we need rate limiting". That is a solution, not a problem. The rewrite:
## Problem
On 2026-08-14 one tenant's misconfigured sync job sent 9,000 req/s for 40 minutes,
about 20% of fleet capacity. p99 latency for all other tenants rose from 180 ms to
2.1 s and 312 tenants saw errors. We have no per-tenant control: the only lever
during the incident was blocking the tenant's API keys entirely.
## Goals
G1. No single tenant can consume more than its contracted request rate for longer
than 1 second.
G2. Added p99 latency per request <= 2 ms.
G3. Tenants receive a standard 429 response with Retry-After so clients can back off.
## Non-goals
N1. Per-endpoint cost-based quotas (follow-up doc).
N2. Changing contracted limits or billing.
N3. DDoS protection; the CDN layer already handles volumetric attacks.Every later section can now be checked against G1-G3. If the chosen design cannot meet G2, the reader can see it; if a reviewer asks for per-endpoint quotas, N1 answers them without a meeting.
Step 2: the one-pager and drafting order
Before writing anything detailed, produce one page: the problem, the goals, a two-paragraph sketch of the proposed approach, the alternatives you are aware of, and the open questions you cannot answer yet. Send it to one or two people who know the systems involved and ask a direct question: is this the right problem, and is this direction obviously wrong?
When you expand the one-pager, write in this order, which is not the order sections appear in the final document: problem and goals, then the numbers, then the alternatives, then the detailed design, then failure behaviour, then rollout, and the summary last. Numbers come before the design because they eliminate designs. Alternatives come before details because detailing the wrong option is the most expensive waste of writing time. The summary comes last because you only know what you are proposing once everything else is written.
Step 3: the numbers
Most weak design docs have no arithmetic. A few lines of estimation usually settle arguments that would otherwise take a meeting, and they show reviewers your assumptions so they can challenge them. For the rate limiter:
| Quantity | Estimate | Source or reasoning |
|---|---|---|
| Peak request rate | 45,000 req/s | Last 90 days of load-balancer metrics, peak hour |
| Active tenants | 12,000 | Billing system |
| Limiter state | about 2.4 MB | 12,000 tenants x 1 key x ~200 bytes |
| Central store operations | 45,000 ops/s | One atomic check per request |
| Network round trip to store | 0.3-0.5 ms p99 | Measured from API hosts to existing cache cluster |
| API hosts | 60 | Current fleet at peak |
The numbers already rule out nothing and rule in a lot: the state is tiny, 45,000 operations per second is well within a single in-memory store's capacity, and a sub-millisecond round trip leaves room inside the 2 ms budget. They also expose one design's weakness. With purely local limiting, each of 60 hosts enforces one sixtieth of the tenant's limit, and a load balancer that does not spread a tenant evenly will throttle that tenant at a fraction of its contracted rate. That observation shapes the alternatives section.
Step 4: alternatives that are real
List options a competent engineer would actually propose, including doing nothing, and compare them against the goals rather than in the abstract. A table forces the comparison to be explicit:
| Option | G1 accuracy | G2 latency | Operational cost | Verdict |
|---|---|---|---|---|
| A. Do nothing; block keys during incidents | Fails | None | Manual, slow | Rejected: the incident recurs |
| B. Local per-host token buckets | Poor with uneven balancing | ~0 ms | Low | Rejected alone: throttles tenants unfairly |
| C. Central token bucket in the existing cache cluster | Good | ~0.5 ms | Medium: new dependency on hot path | Chosen, with B as fallback |
| D. Managed API gateway feature | Good | Unknown, extra hop | High: migrate all routing | Deferred: larger project |
The verdict column carries the reasoning. Rejected options show reviewers you considered what they would have suggested, and deferred options record a decision you may revisit. If you cannot make a genuine case for any alternative, either the decision is trivial and needs no doc, or you have not understood the problem yet.
Step 5: the design, at the level reviewers need
Describe interfaces, data, and behaviour under failure, not class diagrams. Reviewers need to know what each component does, what it stores, and what happens when a dependency is down. For the rate limiter, the core is an atomic token-bucket check in the cache cluster plus the decision the API host makes with the answer:
decision = limiter.check(tenant_id, cost=1) # one round trip, atomic in the store
if decision.unavailable: # store timeout after 1.5 ms
decision = local_bucket[tenant_id].check(rate=contract_rate / hosts)
metrics.increment("ratelimit.fallback")
if decision.allowed:
handle(request)
else:
respond(429, headers={"Retry-After": decision.retry_after_s})The failure behaviour is a design decision and must be written down: if the store is unreachable, the limiter fails open to the conservative local bucket rather than rejecting all traffic or allowing unlimited traffic. A reviewer can disagree with that choice, but only because the document stated it. Also note the dependency you are adding: every request now depends on the cache cluster's latency, which changes its on-call significance. Say so, and name who owns it.
Step 6: rollout, rollback and how you will know
A design that cannot be rolled out safely is not finished. Write the rollout as stages with an exit criterion for each and a rollback that is a configuration change, not a deployment:
- Shadow mode. Compute decisions and log would-be rejections without enforcing. Exit when the would-reject list contains only tenants exceeding their contracts, verified with account managers.
- Enforce for 1% of tenants, chosen at random, excluding the largest. Exit when added p99 latency stays under 2 ms for a week and support tickets show no unexpected 429s.
- Enforce for all, with a per-tenant override flag for emergencies.
Rollback at any stage is setting the enforcement flag to off, which takes effect within seconds. Then list the dashboards and alerts that tell you it is working: decision latency, fallback rate, 429 rate by tenant, and store health. If the change involves migrating data or traffic, the patterns in migration architecture apply directly.
Step 7: running the review and revising
Send the full draft to two or three reviewers first, each with a specific question: the cache owner on the added load and failure mode, a senior API engineer on the fallback logic, a support lead on the customer-facing behaviour. Specific asks get specific answers; "any feedback?" gets comments on wording.
Resolve every comment visibly: change the document, or reply with why not. When a discussion changes the design, update the body rather than leaving the resolution buried in a comment thread, because later readers will read the body. Add a status line at the top (draft, in review, approved, implemented, superseded) and a short decision log recording what changed and why. After approval, keep the document accurate as the build diverges, or mark clearly where it stopped being true.
A small linter catches the mechanical gaps before reviewers do:
import re, sys
REQUIRED = ["Problem", "Goals", "Non-goals", "Alternatives",
"Design", "Failure", "Rollout", "Rollback", "Open questions"]
def lint(path):
text = open(path, encoding="utf-8").read()
headings = re.findall(r"^#+\s+(.+)$", text, flags=re.M)
problems = [f"missing section: {s}" for s in REQUIRED
if not any(s.lower() in h.lower() for h in headings)]
if not re.search(r"^Status:\s*(draft|in review|approved|implemented|superseded)",
text, flags=re.M | re.I):
problems.append("no Status: line")
if len(re.findall(r"\d", text)) < 20:
problems.append("very few numbers: is there an estimate section?")
problems += [f"unresolved: {m}" for m in re.findall(r"\b(?:TBD|TODO)\b.*", text)]
return problems
if __name__ == "__main__":
issues = lint(sys.argv[1])
print("\n".join(issues) or "ok")
sys.exit(1 if issues else 0)Run it in CI on the docs repository, or as a pre-review step. It cannot judge the design, but it ensures every document arrives with the sections reviewers need.
Failure modes
- Solution-first problem statement. "We need Kafka" is not a problem. Reviewers cannot evaluate a design whose purpose is unstated.
- Straw-man alternatives. Options nobody would choose, listed to make the proposal look good. Reviewers notice, and trust in the rest drops.
- No numbers. Capacity, latency and cost claims without arithmetic invite argument and hide the cases where the design does not fit.
- Missing failure behaviour. The doc describes the happy path; the outage happens on the path it left out.
- Rollout as an afterthought. "Deploy and monitor" is not a plan. Stages, exit criteria and a fast rollback are.
- Too long to read. A 30-page document gets skimmed. Keep the main body to what a reviewer must decide on and move detail to appendices.
- Abandoned after approval. A doc that silently drifts from reality misleads the next engineer. Update it or mark it superseded.