SLI, SLO and SLA are three different things that people mix up. An indicator (SLI) is a measurement. An objective (SLO) is an internal target for that measurement. An agreement (SLA) is a contract that says what the customer gets if you miss a promise. Mixing them up causes real damage: teams page on contract thresholds and react too late, sales signs numbers the architecture cannot deliver, and customers claim credits based on measurements your team never sees.

This article explains each layer from first principles and focuses on how they connect, especially the contract end. You will compute downtime budgets, read the clauses of an SLA, calculate the availability of a chain of dependencies, and estimate what an SLA would have cost in credits using your own SLO history. How to choose and measure the events behind an SLI is covered in SLIs, SLOs and error budgets, and SLO engineering in SLO engineering architecture.

Advertisement

Three layers, three audiences

From a measurement to a contractSLIgood events / valid eventsa number, measuredSLOSLI >= 99.9% over 30 daysinternal target; budgetSLApromise >= 99.5% monthlycontract; credits if missedOwner: engineeringOwner: service teamOwner: legal, sales, financeAvailability scale (30-day month)99.0% 432 min99.5% SLA 216 min99.9% SLO 43.2 min99.99% 4.3 minmargin: room to miss the SLO without paying creditsThe SLO sits well inside the SLA, so the team reacts long before the contract is breached.
Figure 1. The SLI is measured, the SLO is a target the team manages to, and the SLA is a looser, contractual promise. The gap between SLO and SLA is deliberate.

An SLI is a ratio of good events to valid events, for example the share of HTTP requests that returned a non-5xx response in under 300 ms. It has no target; it is just a number over a window. An SLO adds a target and a window: 99.9 percent of requests good over a rolling 30 days. The shortfall the SLO allows, 0.1 percent of requests, is the error budget, which the team spends on releases, experiments and incidents. An SLA is an external promise with consequences, usually service credits, sometimes the right to terminate.

SLISLOSLA
What it isA measurementAn internal targetA contract term
Who reads itEngineers, dashboardsThe service team, product ownersCustomers, legal, finance
Miss it andNothing by itselfBudget policy: slow releases, fix reliabilityCredits, refunds, possible termination
How often it changesWhen instrumentation changesQuarterly reviewAt contract renewal

Downtime arithmetic

Percentages hide how small the numbers are. A 30-day month has 43,200 minutes and a 365-day year has 525,600, so each target converts directly into allowed bad time. These are worth memorising, because they make an unrealistic proposal obvious in a meeting.

TargetAllowed per 30-day monthAllowed per year
99%432 min (7.2 h)87.6 h
99.5%216 min (3.6 h)43.8 h
99.9%43.2 min8.76 h
99.95%21.6 min4.38 h
99.99%4.32 min52.6 min

At 99.99 percent a single slow rollback or a ten-minute DNS mistake uses two months of budget. If your deploy, detection and rollback take longer than four minutes end to end, a four-nines promise is a promise to pay credits.

Advertisement

Anatomy of an SLA

An SLA is a legal document, but each clause has an engineering meaning. If engineering does not review these clauses, the contract will define availability in a way your SLI cannot measure.

ClauseWhat it saysEngineering question
Covered serviceWhich product, region and tier the promise applies toDoes our SLI cover exactly this scope?
Definition of unavailableFor example, a minute in which the error rate exceeds a threshold, measured server-sideCan we compute this from data we already keep?
Measurement windowUsually a calendar monthOur SLO uses a rolling window; do we also report calendar months?
ExclusionsScheduled maintenance, customer misuse, force majeure, beta featuresAre maintenance windows announced and logged in a way we can prove?
Credit schedulePercentage of the monthly fee per availability tierWhat is our expected exposure, given history?
Claim processCustomer must request within a set period, with evidenceCan support reproduce the customer's measurement?
Remedy limitsCredits are the sole remedy, capped at a share of the feeIs there any uncapped liability hiding elsewhere?

The definition of unavailable is the most important clause. "Bad minutes" (a minute counts as down if more than, say, 5 percent of requests fail) and "bad requests" (each failed request counts) give very different numbers for a partial outage. Pick the one your SLI already measures, and write the threshold into the contract.

Composite availability: what your dependencies allow

Your service cannot be more available than the chain it depends on. If a request must pass through components in series, and failures are independent, availability multiplies. If redundant copies run in parallel and any one is enough, unavailability multiplies instead.

from math import prod

def series(*avail):            # every component must work
    return prod(avail)

def parallel(*avail):          # any one replica is enough
    return 1 - prod(1 - a for a in avail)

lb, app, db = 0.9999, 0.9995, 0.9995
one_region = series(lb, app, db)                 # 0.99890 -> about 47 min/month down
two_regions = parallel(one_region, one_region)   # 0.9999988 if failures were independent
print(f"{one_region:.5f}  {two_regions:.7f}")

One region of load balancer, application and database gives about 99.89 percent, which is already below a 99.9 percent SLO before you count your own bugs. Two regions look like nearly six nines on paper. They are not, because the independence assumption is false: both regions share your deployment pipeline, configuration, DNS provider, identity service and your own code. A bad config pushed everywhere takes both down at once. In practice, correlated failures set the ceiling, so budget for them directly. One simple model is to multiply the parallel result by the availability of the shared components.

Do this calculation before signing. If a dependency's own SLA is 99.9 percent, you cannot promise 99.95 percent on top of it without redundancy that removes the dependency from the critical path, and its credits will not cover yours.

Worked example: setting an SLA from SLO history

A team runs an API with a 99.9 percent SLO measured as request success. Sales wants to offer an SLA. The team exports twelve months of calendar-month availability from the same SLI: most months are between 99.93 and 99.99 percent, one month had a 70-minute incident at 99.84 percent, and one was 99.6 percent after a database failover went wrong. They evaluate three candidate SLAs with an illustrative credit schedule of the common shape: 10 percent of the monthly fee below the promise, 25 percent below a lower tier, and 50 percent below a third tier.

HISTORY = [99.97, 99.95, 99.99, 99.84, 99.96, 99.93,
           99.98, 99.60, 99.95, 99.97, 99.99, 99.94]   # calendar months, percent

def credit_rate(avail, promise):
    tiers = [(promise, 0.10), (promise - 0.5, 0.25), (promise - 4.0, 0.50)]   # illustrative
    rate = 0.0
    for threshold, pct in tiers:
        if avail < threshold:
            rate = pct
    return rate

def exposure(promise, monthly_fees_total):
    months = [credit_rate(a, promise) for a in HISTORY]
    breaches = sum(1 for r in months if r > 0)
    cost = sum(r * monthly_fees_total for r in months)
    return breaches, cost

for promise in (99.95, 99.9, 99.5):
    b, cost = exposure(promise, monthly_fees_total=200_000)
    print(f"SLA {promise}%: {b} breached months, credits {cost:,.0f} per year")

With these inputs, a 99.95 percent SLA would have been breached in four months (99.84, 99.93, 99.60 and 99.94), all at the 10 percent tier because even 99.60 stays above the second tier at 99.45. With fees of 200,000 a month across SLA customers, that is 80,000 a year in credits, and four disappointed customers. A 99.9 percent SLA would have been breached twice, in the 99.84 and 99.60 months, for 20,000 each, 40,000 in total. A 99.5 percent SLA would not have been breached at all. The team offers 99.5 percent, keeps the internal SLO at 99.9 percent, and so has a margin of 0.4 percentage points, about 173 minutes a month, between the point where engineering reacts and the point where credits are owed.

Real customers do not all claim, and fees vary by customer, so refine the model with your claim rate and revenue split. The method matters more than the numbers: never sign an SLA you have not replayed against your own history.

What belongs in the contract and what stays internal

A service usually has several SLOs but sells far fewer SLA terms. Availability is the common contractual promise because both sides can agree how to measure it. Latency is harder: a percentile depends on payload size, client location and the customer's own network, so a latency SLO normally stays internal, or appears in the contract only as a narrowly defined server-side measure on named operations. Data durability, support response times and recovery objectives for disaster events are often promised separately, each with its own definition and remedy.

Every SLA term needs an internal SLO behind it, set tighter, measured with the same definition and alerting early. The reverse is not true: most SLOs never become contract terms. Keep internal objectives for things customers feel but cannot easily measure, such as freshness of search results, queue delay for background jobs, or the quality of model outputs, tracked as judged samples (see SLOs for AI applications). Those objectives guide engineering decisions without creating financial liability.

ObjectiveInternal SLO?In the SLA?Why
Request availabilityYes, tightYes, looserBoth sides can measure it the same way
p99 latencyYesRarely, narrowly definedDepends on client network and payload
Data freshnessYesUsually notHard for customers to verify
Support first responseYesOften, per severityMeasured by the ticket system, easy to prove
Recovery after a regional disasterTested in drillsSometimes, as a separate termRare events, long windows

Measuring what the contract measures

Customers measure from outside, and you measure from inside. A load balancer can report success while customers fail at DNS or TLS, which your server logs never see. Run synthetic probes from several networks against the exact covered endpoints, and keep their results as evidence for disputes. Use server-side request data for the SLI and probes as a cross-check; when they disagree, investigate before the customer's claim arrives.

Alert on the SLO, not the SLA. By the time an SLA threshold is crossed, the money is already owed. Burn-rate alerts against the SLO give hours or days of warning, which is the whole point of keeping the SLO tighter.

Failure modes

SymptomCauseFix
Credits paid every quarterSLA equal to or tighter than the SLO, or signed without historyReplay history; keep a margin between SLO and SLA
Customer claims a breach you cannot seeContract defines availability differently from your SLIEngineering reviews the definition clause; add external probes
Paging only once the contract is breachedAlerts set on SLA thresholdsBurn-rate alerts on the SLO
Two-region design still breachesCorrelated failures: shared config, deploys, DNSModel shared components in series; stagger rollouts by region
Maintenance argued as downtimeExclusions not announced or loggedAnnounce windows in a dated, retrievable channel
SLO ignored by the teamNo consequence when the budget runs outA written budget policy; see the SLO playbook

The SLO playbook covers the budget policy that gives the internal layer teeth.

What to do next

  1. Write down, for one service, its SLI definition, its SLO, and any SLA it is sold under, side by side.
  2. Check that the SLA's definition of unavailable can be computed from your SLI data; if not, change one of them.
  3. Compute series availability for your critical path and compare it with your SLO.
  4. Export twelve months of calendar-month availability and replay every proposed SLA through a credit calculator.
  5. Keep a margin of at least several times your typical monthly burn between SLO and SLA.
  6. Set alerts on SLO burn rate, and add external probes for every endpoint an SLA covers.
Key takeaway: An SLI is a measurement, an SLO is the internal target you manage to, and an SLA is a contract with financial consequences. Keep the SLO well inside the SLA so engineering reacts long before money is owed. Make sure the contract defines availability the same way your SLI measures it. Multiply availabilities along your critical path, and do not trust redundancy maths that ignores shared failure modes. Before signing any SLA, replay it against a year of your own data and price the credits.