Most outages that follow a launch are not caused by exotic bugs. They come from ordinary gaps: nobody decided in advance what would trigger a rollback, the dependency team did not know traffic was coming, the dashboard showed averages while the tail burned, or the launch was declared done while the old code path, the feature flag and the on-call load quietly lingered for months.

The rollout mechanics, canaries, feature flags and blue-green switches, are covered in their own articles. This one is about the structure around them: the launch as a pipeline of gates with owners and evidence. It explains how to size the process to the risk, what a readiness review should actually check, how to write go/no-go criteria as code, how launch day and hypercare run, and how a retro turns one launch into a better next one. A worked example follows a checkout rewrite from review to close-out.

Advertisement

A launch is a pipeline of gates

Treat a launch as five stages, each ending in a decision that someone owns: tiering (how much process this launch deserves), readiness (is the system and the team prepared), go/no-go (do today's conditions allow starting), launch day (staged exposure with holds), and hypercare (a defined period of heightened attention that ends with an explicit close-out). Rollback is not a stage; it is a decision rule that is written before the launch and applies at every stage.

The difference from a checklist is that every gate produces evidence (a load-test report, a rehearsal time, a query result) and every gate can send the launch back. A checklist can be ticked from memory; a gate asks for the artefact.

A launch is a pipeline of gates, each with an owner and evidenceTieringhow big is the riskReadinessreview + evidenceGo / no-gocriteria as codeLaunch daystaged exposureHypercare1-2 weeksRollbackdecided in advanceSignalsSLOs, business, supportRetroactions with ownerswatchbreachclose outfix, then re-gateRollout mechanics (flags, canaries, blue-green) live inside Launch day; this pipeline decideswhether to start, how far to go, when to stop, and when the launch is actually over.
The launch pipeline. Each gate has an owner and produces evidence; a breach during launch day or hypercare triggers the pre-decided rollback, and the launch returns to readiness rather than resuming from where it stopped.

Tiering: size the process to the risk

TierExamplesProcess
1Revenue path, auth, data migrations, anything irreversibleFull readiness review, criteria as code, rehearsed rollback, staged exposure with holds, two weeks of hypercare
2New user-facing feature behind a flag, new internal serviceLightweight review, criteria file, staged exposure, a few days of hypercare
3Copy changes, internal tooling, dark-launched codeNormal deploy pipeline and code review; no launch process

Tiering matters because a process applied to everything decays into box-ticking. Decide the tier from two questions: what is the blast radius if this goes wrong, and can it be undone in minutes? An irreversible change, such as a schema migration that drops a column or an email sent to every customer, is tier 1 however small the code diff is.

Advertisement

The readiness review

A readiness review is a meeting with a document, held about two weeks before launch, where the owning team shows evidence rather than intentions. The questions that earn their place:

  • Capacity. What is the forecast peak, and has the system been load tested at a multiple of it? The report is the evidence; load testing explains how to make it meaningful.
  • Observability. Are there dashboards and alerts for the new path, on percentiles not averages, split by the launch's feature flag so new and old can be compared?
  • SLOs. Which service level objectives cover the feature, and how much error budget is left? Launching into an exhausted budget is a no-go in its own right; an SLO program makes that visible.
  • Dependencies. Every downstream team has been told the date and the expected extra load, in writing.
  • Rollback. What exactly is the rollback, how long does it take, and has it been rehearsed? If part of the change cannot be rolled back, such as a data migration, what is the forward-fix plan?
  • Runbooks and support. The on-call engineer has a runbook for the new failure modes, and support knows what users will see.

Record the answers in the launch document with links to the evidence. Items without evidence become blockers with owners and dates, not footnotes.

Go/no-go criteria as code

The most important artefact is the criteria file: the conditions under which the launch proceeds, holds or rolls back, written before launch day and reviewed like code. Written criteria remove the worst launch-day dynamic, which is a room of tired people negotiating whether a 2 percent error rate is acceptable while it is happening.

Each criterion is a query against the metrics system with a threshold. Include at least one availability measure, one latency percentile, and one business measure compared against a control group, because many broken launches are technically healthy and simply stop people buying.

# launch/checkout-v2.yaml -- reviewed in the readiness meeting, versioned with the code
launch: checkout-v2
tier: 1                       # user-facing, revenue path
owner: payments-oncall
stages: [internal, 1pct, 10pct, 50pct, 100pct]
hold_minutes: {internal: 1440, 1pct: 60, 10pct: 120, 50pct: 240}
criteria:
  - name: availability
    query: 1 - (sum(rate(http_requests_total{route="/checkout",code=~"5.."}[10m]))
               / sum(rate(http_requests_total{route="/checkout"}[10m])))
    min: 0.999
  - name: p99_latency_ms
    query: histogram_quantile(0.99, sum by (le) (rate(checkout_latency_ms_bucket[10m])))
    max: 800
  - name: order_conversion_vs_control
    query: conversion_ratio{variant="v2"} / conversion_ratio{variant="control"}
    min: 0.98
rollback: "set flag checkout_v2 to 0%; no schema change to undo"

A small script turns the file into a decision. The one rule it must not break is that missing data is a no-go: a query that errors or returns nothing must never be read as a pass.

import sys, yaml

def evaluate(spec, query_fn):
    """Return (go, reasons). query_fn(promql) -> float, supplied by your metrics client."""
    reasons = []
    for c in spec["criteria"]:
        try:
            value = query_fn(c["query"])
        except Exception as exc:              # no data is a NO-GO, never a pass
            reasons.append(f"{c['name']}: query failed ({exc})")
            continue
        if "min" in c and value < c["min"]:
            reasons.append(f"{c['name']}={value:.4f} < {c['min']}")
        if "max" in c and value > c["max"]:
            reasons.append(f"{c['name']}={value:.1f} > {c['max']}")
    return (not reasons), reasons

if __name__ == "__main__":
    spec = yaml.safe_load(open(sys.argv[1]))
    from metrics_client import query          # your own thin wrapper around the metrics API
    go, reasons = evaluate(spec, query)
    print("GO" if go else "NO-GO", *reasons, sep="\n  ")
    sys.exit(0 if go else 1)

Run the same script at every stage transition, and in a loop during holds so a breach is noticed within minutes. Many progressive-delivery controllers can evaluate similar queries automatically; the canary architecture covers that automation. The criteria file is still worth having, because it is the human agreement those controllers encode.

Launch day

Launch day should be boring. The work that makes it boring is done earlier; on the day, the job is to expose the change in stages, hold at each, and apply the rules.

  • Roles. One launch lead who makes the calls, one operator who executes changes, one observer who watches the gate and dashboards, and a named contact for each critical dependency. The lead does not also type commands.
  • Timing. Start early in the working day, early in the week, never before a holiday or a dependency's freeze. Leave enough of the day to roll back and diagnose.
  • Staging exposure. Internal users first, then 1 percent, 10, 50 and 100, with a hold at each long enough to see the slowest signal. Business metrics often need hours; feature flags make the percentages and the instant off switch cheap.
  • One channel. Every decision and command goes into one written channel with timestamps. It becomes the timeline for the retro.
T-14d  readiness review; criteria file merged; load test at 2x forecast peak signed off
T-7d   dark launch: new path runs on real traffic, results discarded and compared
T-2d   rollback rehearsed in staging and timed; comms drafted; support briefed
T-1d   change freeze on dependencies; go/no-go #1 (people, dependencies, calendar)
T-0    09:30 go/no-go #2 (gate script) -> internal -> 1% -> hold -> 10% -> hold
T+1d   50% after an overnight hold with the gate green
T+2d   100%; hypercare starts (daily review, lowered alert thresholds)
T+14d  hypercare exit review; flag removed; retro held and actions assigned

Rollback is decided before launch

The criteria file already says when to roll back; the launch lead's job during a breach is to execute, not to debate. Diagnose after exposure is reduced, not while customers are affected. Three rules keep rollback honest:

  1. The rollback path is rehearsed and timed in staging before launch day, including anything stateful.
  2. Changes that cannot be undone are separated from those that can: ship the backwards-compatible schema first, as its own launch, so the risky behaviour change can be rolled back freely.
  3. After a rollback, the launch returns to readiness. It does not resume from where it stopped until the cause is understood and the criteria are re-run.

The rollback strategy guide covers the mechanics for code, configuration and data.

Hypercare and close-out

Reaching 100 percent is not the end. Hypercare is a defined period, often one to two weeks for tier 1, during which the owning team reviews the launch dashboards daily, alert thresholds for the new path are tighter than normal, and support tickets tagged to the launch are triaged the same day. Slow-burning problems such as memory growth, a weekly batch job interacting badly with the new path, or a customer segment that only appears at month end show up here.

Hypercare ends with an explicit exit review: criteria green for the whole period, the old code path and the launch flag removed or scheduled for removal, alert thresholds returned to normal, and documentation updated. Without this step launches never finish; they leave dead flags and a second code path behind for years.

Worked example: checkout-v2

A payments team rewrites checkout. Tiering puts it at tier 1: revenue path, and a new orders table. The readiness review finds two gaps: the load test covered 1.2 times forecast peak rather than two, and the fraud-scoring team had not been told. Both become blockers and are closed a week later.

The schema change ships first as its own tier 2 launch, writing to both tables. Seven days of dark launch compare old and new totals on real traffic and find a rounding difference in one currency, fixed before any customer sees the new flow. On launch day the gate is green at internal and 1 percent. At 10 percent the conversion ratio against control drops to 0.96 after ninety minutes, below the 0.98 floor; availability and latency are perfect. The lead rolls back by setting the flag to 0 percent, which takes under a minute. Diagnosis finds that a saved-card component failed silently on one browser. The fix ships, readiness is re-run, and the second attempt reaches 100 percent two days later. Hypercare closes after two weeks with the old path deleted.

Nothing about the failure was exotic. It was caught because a business criterion was written down in advance and missing a threshold meant stopping, not discussing.

Failure modes and trade-offs

  • Process on everything. Tier 1 ceremony for tier 3 changes teaches people to tick boxes. Keep tiering strict and honest.
  • Vanity dashboards. Averages and aggregate success rates hide a broken segment. Split by flag variant and look at percentiles and business outcomes.
  • Holds too short for the slowest signal. A 15-minute hold cannot see a conversion drop that needs two hours of traffic. Size holds to the metric, not to impatience.
  • Launch-date pressure. Marketing dates make no-go decisions expensive. Agree beforehand that an announcement can precede availability, or that the date moves; do not let it override the criteria.
  • No close-out. Launches that never exit hypercare leave flags, old paths and extra on-call load behind. The exit review is part of the launch.

The central trade-off is speed against confidence. Staged exposure with long holds makes a tier 1 launch take days instead of minutes; that is the right price for a revenue path and the wrong one for a copy change, which is why tiering comes first.

What to do next

  1. Write a three-tier definition for your team and tier the next three launches on your roadmap.
  2. Create a criteria file template with availability, latency and one business metric against control.
  3. Adopt a gate script that treats missing data as no-go, and run it at every stage transition.
  4. Rehearse and time one rollback in staging this month, including any stateful part.
  5. Split one upcoming launch into an irreversible compatible step and a reversible behaviour step.
  6. Add a hypercare exit review to your launch template, with flag and old-path removal as exit criteria.
  7. Hold a retro for the last launch that went badly and turn its findings into template changes, each with an owner.
Key takeaway: A launch is a pipeline of gates, not an event: tier the risk, hold a readiness review that demands evidence, write go/no-go criteria as code with missing data counting as no-go, expose the change in stages with holds sized to the slowest signal, decide rollback before launch day, and finish with hypercare and an explicit close-out. Rollout tools make each stage cheap; the process decides whether to start, how far to go and when the launch is really over.