A design doc review is the cheapest point at which a system can be changed. Before the review, the design lives in one person's head and nobody can challenge it; after the first few weeks of implementation, every change has a cost in rewritten code and reset schedules. The review is the window in between, and most reviewers waste it. They read the document once from top to bottom, leave a handful of comments about naming and diagrams, and approve. The doc was reviewed in the sense that someone looked at it, and the design was not reviewed at all.

This guide is about the reviewer's job, not the author's. Design docs architecture covers how a doc should be structured and how an organisation should run the process. Here the question is narrower and more practical: when a doc lands in your queue, what do you actually do, in what order, and how do you turn what you find into comments that change the outcome? A worked example runs through the middle: a proposal to move push notification delivery from synchronous API calls onto a queue.

Advertisement

What a review is for, and what it is not

A review has three jobs. First, it checks that the problem is real and correctly framed, because the most expensive design mistakes are solutions to the wrong problem. Second, it checks that the proposed design actually solves that problem under realistic load, failure and change. Third, it spreads knowledge: after the review, at least two people other than the author understand the system well enough to operate and extend it.

A review is not a chance to redesign the system in your own style. If you would have chosen a different database, that is only worth saying if the author's choice fails a requirement or carries a risk the doc has not acknowledged. It is also not an approval ritual. If you cannot say what would make you reject the doc, you are not reviewing it.

Before reading anything, check the ask. A good request says which decision the author needs, by when, and from whom. A doc asking 'should we build this at all?' needs different reviewers from one asking 'is this schema right?'. If the ask is missing, request it first.

The review as a pipeline

Treat the review as a sequence of passes, each with one question, rather than a single linear read. Reading for everything at once means you notice what is easy to notice, which is wording and formatting, and miss what is hard, which is a failure mode that appears in none of the sections.

A design review as a pipeline: each pass asks a different questionRequestdoc, deadline, askTriageright reviewers?Pass 1: problemis this worth solving?Pass 2: designdoes it work at scale?Pass 3: failureops, migration, rollbackCommentslabelled by severityResolutionauthor answers eachDecisionapprove / conditions / noDecision recordwhy, and what was rejectedLoop back to Pass 1if a blocking comment changes the problem statement, the design passes are void and must be repeatedReviewers who skip pass 1 end up polishing a design for the wrong problem.
The three reading passes feed labelled comments, which the author resolves before a decision is recorded. A blocking finding about the problem itself sends the review back to the start.

Budget the time explicitly. For a typical ten-page doc, about twenty minutes on the first pass, forty on the second and thirty on the third is a reasonable split; a review that takes less than an hour for a doc that commits a team to a quarter of work has probably skipped a pass.

Advertisement

Pass one: the problem before the solution

Read only the context, goals, non-goals and requirements. Stop before the proposed design. You are checking whether the framing would survive if the solution section were deleted.

  • Is the problem stated with evidence? 'Notification latency is a problem' is a claim. 'p99 delivery latency is 41 seconds against a 5-second product requirement, and 30 percent of API timeouts in the last month came from the notification call path' is evidence. Ask for the second form.
  • Are the goals measurable? Each goal should be something you could check after launch. 'Improve reliability' cannot fail; 'no notification API call on the request path, and p99 delivery under 5 seconds' can.
  • Do the non-goals exclude something tempting? Non-goals are where scope creep is prevented. If they only list things nobody would attempt, they are doing no work.
  • Who else has this problem? If another team has solved it, or is about to, the right outcome may be reuse rather than a new system.

Write down, in one sentence, what you think the problem is. If your sentence differs from the author's, that difference is your first and most important comment, and it may make every later comment moot.

Pass two: the design under load, with the numbers re-derived

Now read the design. For each component, ask what it does with the expected traffic, how it scales, and what it depends on. The single most effective technique is to re-derive the doc's numbers from its own inputs. Docs often contain a capacity figure that was estimated once, early, and never revisited after the requirements changed. A few lines of arithmetic expose it.

# Back-of-envelope check for the numbers in a design doc.
# Re-derive every headline figure from the doc's own inputs; flag anything off by more than 2x.
DAILY_ACTIVE_USERS = 4_000_000
NOTIFICATIONS_PER_USER_PER_DAY = 6
PEAK_TO_AVERAGE = 8            # doc claims "peaks are about 3x"; ask where 3x came from
PAYLOAD_BYTES = 1_200
RETENTION_DAYS = 14
REPLICATION = 3

avg_per_sec = DAILY_ACTIVE_USERS * NOTIFICATIONS_PER_USER_PER_DAY / 86_400
peak_per_sec = avg_per_sec * PEAK_TO_AVERAGE
storage_gb = (DAILY_ACTIVE_USERS * NOTIFICATIONS_PER_USER_PER_DAY
              * PAYLOAD_BYTES * RETENTION_DAYS * REPLICATION) / 1e9

doc_claims = {"peak_per_sec": 850, "storage_gb": 400}
derived = {"peak_per_sec": peak_per_sec, "storage_gb": storage_gb}
for key, claimed in doc_claims.items():
    ratio = derived[key] / claimed
    flag = "CHECK" if ratio > 2 or ratio < 0.5 else "ok"
    print(f"{key:14} doc={claimed:>10,.0f} derived={derived[key]:>10,.0f} {flag}")

In the worked example the doc claims a peak of 850 sends per second and 400 GB of storage. From its own inputs, average load is about 278 per second, and with the peak-to-average ratio of 8 measured from last quarter's traffic the peak is about 2,200 per second, not 850. The storage figure, about 1,210 GB, is three times the claim because the author forgot replication. Neither error is fatal, but both change the sizing section, and the peak figure feeds the retry analysis in pass three.

Beyond numbers, pass two looks for: state that has no clear owner; synchronous calls hiding inside an 'asynchronous' design; consistency assumptions that the chosen storage does not provide; and alternatives that were dismissed in one line. A rejected alternative deserves a sentence explaining which requirement it failed. If the doc cannot say, the alternative was not really evaluated.

Pass three: failure, operations, migration and rollback

The third pass assumes everything goes wrong. Most designs describe the happy path in detail and the failure paths in a paragraph, and this is where production incidents come from.

QuestionWhat a good doc saysRed flag
What happens when a dependency is slow?Timeouts, bounded queues, backpressure, and what the user sees'We will retry'
What happens when it is down for an hour?Where work accumulates, how much, and how it drainsNo mention of backlog size
How is it observed?The two or three signals that say it is healthy, and the alert thresholds'We will add dashboards'
How is it rolled out?Flag or percentage rollout, with a comparison metricBig-bang cutover
How is it rolled back?A step that has been rehearsed, and which data changes are irreversibleRollback not mentioned
Who is on call for it?A named team and a runbook to be written before launchOwner is 'the platform'

Migration deserves its own scrutiny. A design that is correct in its final state can still be unsafe in the months it spends half-deployed. Ask what the system looks like with both old and new paths live, which one is the source of truth, and how you would know if they disagreed. The change management and launch guides cover the operational side of this transition in more detail.

Writing comments that get acted on

A finding only matters if the author can act on it. Label every comment with its severity, anchor it to a section, and separate the observation from the request. Three labels are enough: blocking means you would not approve until it is resolved; question means you need information before you can judge; nit means take it or leave it. The labels let the author triage forty comments in minutes, and they force you to decide what you actually believe is important.

[blocking] Section 4.2, retry policy
Observation: workers retry a failed send 5 times with no backoff.
Why it matters: during a provider outage at the 2,200/s peak this becomes about 13,200 calls/s (each send plus five retries)
against a provider that is already failing, and the queue drains no faster.
Ask: exponential backoff with jitter and a per-provider circuit breaker, or an argument
for why a fixed retry is safe here.

[question] Section 3, "exactly once delivery"
Is this a requirement from product, or an aspiration? At-least-once plus an idempotency
key on the device side would remove the need for the dedup table in 4.4.

[nit] Diagram 2 labels the queue "Kafka" but section 5 says SQS.

Each blocking comment states the consequence in concrete terms and offers an acceptable resolution, which is often 'change this' or 'explain why it is safe'. Leaving the second option open matters: the author may know something you do not. Avoid comments that are really preferences presented as problems, and avoid 'have you considered X?' when you mean 'X is required'. If you have more than about five blocking comments, the doc is probably not ready for detailed review; say so once, at the top, rather than leaving fifty line comments that will all be invalidated by a rewrite.

Worked example: the notification queue review

Putting the passes together for the notification proposal. Pass one: the problem is well evidenced, with latency data and timeout attribution, and the goals are measurable. One question goes back to the author, asking whether 'exactly once' is a product requirement, because it drives a large part of the complexity. Pass two: the re-derived peak is 2.6 times the claim and storage is three times the claim, so there is one blocking comment asking for re-sized capacity. The dismissed alternative, a managed push service, has a one-line rejection, which earns a question. Pass three: there is no backoff on retries, which at the corrected peak would multiply load on a failing provider. That gets one blocking comment. There is no backlog estimate for a one-hour provider outage; a quick calculation gives roughly eight million queued messages at peak, which needs a retention and drain-rate answer, so it is blocking too. Rollback is described and credible.

The result is three blocking comments, three questions and two nits. The author revises the sizing, adds backoff and a circuit breaker, adds the outage analysis, and answers that exactly-once was an aspiration, which removes a deduplication table. The design got simpler as a result of review, which is a good sign that the review worked.

The decision and what to record

A review ends with one of four outcomes. Approve: no blocking issues remain. Approve with conditions: proceed, but named items must be done before a stated milestone, such as 'load test at 2,500/s before enabling for more than 10 percent of users'. Request changes: blocking issues remain and another round is needed. Reject: the problem is not worth solving now, or the approach is unsound and a different design is needed.

Conditions must be checkable and must have an owner; an unowned condition is an approval in disguise. When reviewers disagree and the author cannot satisfy both, escalate to whoever owns the decision with both positions written down side by side, rather than letting the thread grow. Record the outcome in the doc itself: who approved, the conditions, and the alternatives rejected with reasons. That record answers the question someone will ask in a year, which is why the system looks the way it does.

Failure modes of design review

  • Rubber-stamping. Approvals arrive within minutes of the request. The fix is a stated minimum: each approver lists at least one risk they considered, even if they decided it was acceptable.
  • Bikeshedding. Long threads about names and diagrams, nothing about failure. Severity labels and the pass order push attention to where it belongs.
  • Reviewing too late. The doc is circulated after implementation has started, so every blocking comment is effectively a rewrite request and gets negotiated away. Ask for a short problem-only review early, before the design is written.
  • Reviewer as designer. The reviewer rewrites the design in comments. Raise the requirement the design fails and let the author solve it.
  • Review by committee. Twelve reviewers, none accountable. Name two or three required reviewers by expertise, such as storage, security and the on-call team, and make everyone else optional.

The same discipline applies when the author is partly an AI assistant. Generated design text tends to be fluent and plausible, so the evidence and numbers checks matter more, not less. Reviewing AI-generated code and spec-driven development discuss how specs and generated artefacts change what a reviewer has to verify.

What to do next

  1. Before your next review, check the ask: which decision, by when, from whom. Request it if it is missing.
  2. Read in three passes, problem, design and failure, and write your one-sentence problem statement after the first.
  3. Re-derive at least two headline numbers from the doc's own inputs, and flag anything more than 2x off.
  4. Go through the failure table: slow dependency, hour-long outage, observability, rollout, rollback, ownership.
  5. Label every comment as blocking, question or nit, and give each blocking comment a concrete consequence and an acceptable resolution.
  6. End with an explicit outcome, and record conditions with owners and milestones in the doc.
  7. Propose to your team that required reviewers are named by expertise, and that each approval names one risk considered.
Key takeaway: Reviewing a design doc well is a method, not a reading. Check the ask, then read three times with three questions: is this the right problem, does the design work at realistic load, and what happens when things fail or change. Re-derive the numbers yourself, label comments by severity with concrete consequences, and end with a recorded decision. A review that leaves the design simpler and better understood by more people has done its job.