Most teams that adopt service level objectives get the measurement right and the consequences wrong. They choose sensible indicators, publish a dashboard with a green number, and then nothing happens when the number turns red, because nobody agreed in advance what red means. The error budget is spent, the feature roadmap carries on, and six months later the SLO is a decoration that engineers learn to ignore.
The SLO playbook is the missing piece. It is the document, and increasingly the code, that binds each state of the error budget to a specific action, a specific owner and a deadline, and that says what happens when the owner cannot or will not act. This article treats the playbook as an architecture with components you can build and test: the spec it reads, the budget state machine, the triggers, the escalation ladder, the review cadence, the attribution ledger and the waiver process. Choosing indicators and doing the budget arithmetic are covered in the SLO engineering guide, and rolling SLOs out across an organisation in the SLO program rollout guide; this page assumes you already have numbers and need them to change behaviour.
What a playbook is, and what it is not
Three documents get confused. The SLO specification says what is measured: the indicator, the target, the window and the data source. A runbook says how to fix a particular symptom: restart this, fail over that. The playbook sits between them and answers a different question: given the state of the budget, what is the team obliged to do, and who decides when the rules should bend.
A good playbook is short enough to read during an incident and precise enough to settle an argument, because it was agreed before it was needed: when product and on-call disagree about shipping on a day the budget is exhausted, the answer is already written down and signed off.
The components
Every working playbook has the same seven parts, whether it lives in a wiki or a repository.
- Scope: the services and SLOs it governs, and the owning team for each. A playbook without named owners is advice.
- Budget states: a small number of named states derived from the remaining budget and the current burn, each with entry and exit conditions.
- Triggers: events that force an action regardless of state, such as a fast-burn page or a single incident that consumes a large share of the budget.
- Actions: what changes in each state, written as obligations with deadlines rather than suggestions.
- Escalation ladder: who is told next, after how long, when an obligation is not met.
- Exceptions: how a waiver is requested, who may grant it, how long it lasts and where it is recorded.
- Review cadence: when the playbook and the SLOs themselves are re-examined, and what evidence that review uses.
Budget states as a state machine
Remaining budget alone is a poor state variable. A service that has 60 percent of its budget left but is burning at twenty times the sustainable rate is in more trouble than one with 20 percent left and a flat burn. Define states from both: the fraction of budget remaining over the SLO window, and the burn rate over a recent window, where a burn rate of 1 means the budget would be used up exactly at the end of the window.
| State | Entry condition | Obligations |
|---|---|---|
| Healthy | More than 50% budget left and 6h burn below 1 | Normal release cadence; budget may be spent on risky changes and experiments |
| Watch | 25-50% left, or 6h burn between 1 and 6 | Owner reviews top budget consumers within 2 working days; no change to releases |
| Constrained | Under 25% left, or a fast-burn page in the last 24h | Only changes with a rollback plan and canary; reliability work takes the top backlog slot |
| Exhausted | No budget left over the window | Feature releases stop except P0 and security fixes until the budget recovers |
The thresholds are illustrative, not universal; the shape is what matters. Use hysteresis so a service hovering around a boundary does not flap between states every hour: enter Constrained below 25 percent but leave it only above 30. Make exit conditions as explicit as entry ones; the commonest argument is when a freeze ends. Under a rolling 28-day window, Exhausted lasts until enough bad days age out of the window, which can be weeks; some teams add an explicit early exit once a postmortem's corrective actions have shipped and the burn has stayed below 1 for seven days.
Triggers that bypass the state
States move slowly by design. Triggers exist for the events that need an immediate response. The standard set comes from multi-window burn-rate alerting, described in the Google SRE Workbook and in the burn-rate alerting guide:
- Fast page: burn rate above 14.4 over the last hour and over the last five minutes. At that rate, 2 percent of a 30-day budget disappears in an hour.
- Slow page: burn rate above 6 over six hours and over thirty minutes, which is 5 percent of the budget in six hours.
- Ticket: burn rate above 1 over three days and over six hours, meaning the budget will run out before the window ends if nothing changes.
- Large incident: any single incident that consumes more than 20 percent of the budget requires a postmortem, the rule used in the Workbook's example error budget policy.
The short window in each pair is what makes the alert reset quickly after recovery; without it, a one-hour burn keeps paging for an hour after the fix. Each trigger in the playbook names its action: pages go to the on-call rotation, tickets go to the owning team's queue with a due date, and the postmortem trigger opens a document from a template with the budget consumption already filled in.
The escalation ladder
Obligations without timers become wishes. The escalation ladder gives each obligation a deadline and a next rung. A typical ladder has four rungs: the on-call engineer, the owning team's lead, the engineering manager or director who owns both the service and its roadmap, and finally the executive who arbitrates between reliability and delivery across teams. The ladder moves one rung when a deadline passes without the obligation being met, and the move itself is automatic, not a judgement call.
The top of the ladder is where disagreements are settled, not where people are punished; the Workbook's example policy routes disputes about the policy to the CTO. And the ladder escalates decisions, not incidents: it tracks whether the freeze, the postmortem or the reliability work happened on time.
Attribution: whose budget was it?
A playbook that freezes a team's releases because a shared database fell over will be abandoned within a quarter. Attribution records which cause consumed each slice of budget, so that actions land on the team that can act. Tag each burn interval with a cause category, usually by joining the bad-event series against the incident tracker and the deploy log.
-- Budget spent per cause over the current 28-day window (bad events are pre-aggregated per minute)
SELECT coalesce(i.cause_category, d.change_id, 'unattributed') AS cause,
sum(b.bad_events) AS bad,
round(100.0 * sum(b.bad_events) / max(w.budget_events), 1) AS pct_of_budget
FROM slo_bad_minutes b
JOIN slo_window w ON w.slo_id = b.slo_id
LEFT JOIN incidents i ON b.minute BETWEEN i.start_ts AND i.end_ts AND i.slo_id = b.slo_id
LEFT JOIN deploys d ON b.minute BETWEEN d.ts AND d.ts + interval '30 minutes' AND d.service = w.service
WHERE b.slo_id = 'checkout-availability'
AND b.minute >= now() - interval '28 days'
GROUP BY 1
ORDER BY bad DESC;Decide in the playbook how dependency failures count. The usual rule is that the budget is the user's experience, so a dependency outage still spends your budget, but the resulting obligations move to the dependency's owner through their own SLO. Track the unattributed share as a metric; if it exceeds a fifth of spend, fix attribution before trusting the playbook.
Exceptions and waivers
Every playbook needs a pressure valve, and the valve must be visible. A waiver lets a release proceed during Constrained or Exhausted state. It has a requester, an approver at least one rung above the requester, an expiry, a stated risk and a rollback plan, and it is recorded in the decision log the monthly review reads. Waivers that are never refused suggest the policy is too strict; releases that bypass it without a waiver mean it is not enforced.
The review cadence
The playbook runs on three clocks. A weekly operations review, fifteen minutes long, looks at every SLO not in Healthy, the open obligations and anything that climbed the ladder. A monthly SLO review looks at the decision log: budget spent by cause, waivers granted, triggers that fired, and whether each state change led to the action the playbook promised. A quarterly review revisits the targets themselves.
The quarterly review is where the loop closes. If a service sat in Healthy all quarter with most of its budget unspent, either the target is too loose or the team is being too cautious and could ship faster. If it sat in Constrained all quarter, either the target is above what the architecture can deliver, in which case the target or the architecture must change, or the reliability work is not getting done, in which case the ladder is broken.
The policy as code
Once the rules are agreed, encode them. A policy engine that evaluates state every few minutes, publishes it as a metric, and blocks deploys through the CI system turns the playbook from a document people forget into a gate they meet.
from dataclasses import dataclass
@dataclass
class Budget:
remaining: float # fraction of the window's budget left, can go negative
burn_1h: float
burn_5m: float
burn_6h: float
burn_30m: float
paged_last_24h: bool
def next_state(prev: str, b: Budget) -> str:
if b.remaining <= 0:
return "exhausted"
if b.remaining < 0.25 or b.paged_last_24h:
return "constrained"
# hysteresis: leave constrained only with clear headroom
if prev in ("constrained", "exhausted") and b.remaining < 0.30:
return "constrained"
if b.remaining < 0.50 or b.burn_6h > 1:
return "watch"
return "healthy"
def triggers(b: Budget) -> list[str]:
fired = []
if b.burn_1h > 14.4 and b.burn_5m > 14.4:
fired.append("page:fast-burn")
elif b.burn_6h > 6 and b.burn_30m > 6:
fired.append("page:slow-burn")
return fired
def may_deploy(state: str, change: dict, waivers: set[str]) -> bool:
if change["id"] in waivers or change["priority"] in ("P0", "security"):
return True
if state == "exhausted":
return False
if state == "constrained":
return change["has_canary"] and change["has_rollback"]
return TrueThe deploy gate is the controversial part, and it should be introduced last, after a quarter of the engine publishing states without enforcing them. The dry run shows how often each state would have blocked releases, so thresholds are tuned before they stop anyone's launch.
A worked month
Take a checkout API with a 99.9 percent availability SLO over a rolling 28 days and a steady 20,000 requests per minute, about 806 million requests per window. The budget is 0.1 percent of requests, roughly 806,000 failures. On day 9 a deploy raises the error rate to 20 percent, and because it shipped with a schema migration the fix takes 150 minutes: 3 million requests, 600,000 failures, about 74 percent of the budget.
The burn rate during the incident is the error rate divided by the allowed rate, 20 / 0.1 = 200. Five minutes in, the five-minute burn is 200 and the one-hour burn is about 200 x 5 / 60 = 17, both above 14.4, so the fast page fires. The engine moves the service straight to Constrained because of the page, with about 26 percent of budget left. The large-incident trigger opens a postmortem, since 74 percent is well over 20 percent, with the attribution query pointing at the deploy.
For the next two weeks, releases need canaries and rollback plans. On day 15 a team requests a waiver for a pricing change with a contractual date; the director approves it with a seven-day expiry, and it ships behind a flag. On day 20 a dependency outage burns another 15 percent, leaving about 11 percent; the attribution goes to the payments provider and the obligations move to its owner. Day 37, when the incident ages out of the window, the budget returns above 30 percent and the service leaves Constrained. The monthly review sees one large incident, one waiver, one dependency attribution and an open postmortem action, and schedules it.
Failure modes and trade-offs
- Policy without teeth: states published but never enforced. Symptom: Exhausted services shipping features with no waiver on file.
- Policy with too many teeth: freezes so frequent that teams game the SLI, for example by excluding the requests that fail. Watch for SLI definition changes that coincide with budget trouble.
- Low traffic: a service with a few thousand requests a window flips states on a handful of errors. Lengthen the window, aggregate services, or use synthetic probes as the SLI.
- Shared-fate punishment: freezing a team for a platform outage. Attribution fixes it; skipping attribution guarantees resentment.
What to do next
- List every SLO you publish and write its owner beside it; drop or reassign any SLO without one.
- Define four budget states with entry and exit conditions, including hysteresis, and agree them with both engineering and product leadership.
- Implement the multi-window burn-rate triggers and the 20 percent large-incident postmortem rule.
- Write a four-rung escalation ladder with a timer on every rung.
- Build the attribution query against your incident and deploy logs, and track the unattributed share.
- Create a waiver template with requester, approver, expiry, risk and rollback plan, and log every waiver.
- Run the policy engine in dry-run mode for a quarter, then enable the deploy gate.
- Put the weekly, monthly and quarterly reviews on the calendar, and record target changes as commits.