DevOps is often introduced as a list of tools: a CI server, containers, Kubernetes, Terraform. The tools matter, but they are the visible end of something else. DevOps is an operating model for software delivery: a set of principles about how work flows from an idea to production, who owns a service once it is running, and how an organisation learns from what happens there. Teams that buy the tools without changing the model usually end up with automated versions of the same slow, fragile process.
This article explains that model from first principles, walks through a worked example that shows where delivery time actually goes, maps each principle to the practices that implement it, and shows how to measure whether any of it is working. The mechanics of individual practices are covered in their own articles, linked where they come up.
The problem DevOps answers
In a traditional split, developers are rewarded for shipping change and operations teams for keeping systems stable. Both goals are reasonable, and together they produce a standoff. Change is seen as the enemy of stability, so it is batched into infrequent large releases guarded by review boards and handed over a wall to people who did not write it. Large batches are harder to test, harder to diagnose and riskier to roll back, so releases do fail more often, which confirms everyone's belief that change is dangerous, and the batches grow larger still.
The way out is counterintuitive: make changes smaller and more frequent, and make the same team responsible for building and running them. A small change is easy to review, test, deploy and undo, and when it breaks, its author knows what it touched. Speed and stability stop being opposites, and years of industry research on delivery performance have found the two tend to improve together.
Where the term came from
The word gained currency in 2009. At the Velocity conference that year, John Allspaw and Paul Hammond gave a talk titled "10+ Deploys per Day: Dev and Ops Cooperation at Flickr", and later that year Patrick Debois organised the first DevOpsDays in Ghent. The ideas drew on earlier work: lean manufacturing, agile development, and the operations practices of large web companies. Nobody owns the definition, which is why the principles below are more useful than any one vendor's description.
CAMS and CALMS
A compact way to describe what DevOps involves is CAMS, usually credited to John Willis and Damon Edwards: Culture, Automation, Measurement and Sharing. Jez Humble later added Lean, giving CALMS.
- Culture. Shared ownership of outcomes, trust between teams, and treating failure as information rather than blame. It comes first because the other four do not survive without it.
- Automation. Builds, tests, environments, deployments and rollbacks done by machines, repeatably. Automation removes queues and makes small batches cheap.
- Lean. Small batches, limiting work in progress, and finding and removing waste, especially waiting.
- Measurement. Data about the delivery process and the running system, used to steer rather than to punish.
- Sharing. Knowledge, tools and lessons spread across teams, through shared code, documentation and open postmortems.
The Three Ways
Gene Kim and co-authors, in The Phoenix Project and The DevOps Handbook, describe three principles that build on each other.
- Flow. Optimise the whole path from commit to customer, not one team's step. Make work visible, reduce batch size, remove handoffs and never pass a known defect downstream.
- Feedback. Create fast feedback from right to left: tests that run in minutes, telemetry that shows how a change behaves in production, and developers who are paged when their service breaks.
- Continual learning and experimentation. Use failures and experiments to improve the system, through blameless postmortems, time reserved for improvement, and safe ways to try things.
A worked example: where the time goes
A value stream map follows one change from start to production and records, for each step, the time spent working and the time spent waiting. Take a small change in a team with a separate QA group and a weekly change board. Writing it takes 4 hours. It waits a day for review, which takes 30 minutes. It waits 3 days for a shared QA environment, and QA takes 4 hours. It waits on average 2.5 days for the change board, then a day in the operations ticket queue, and the deployment takes an hour.
Elapsed time is 189.5 hours, nearly 8 days, but only 9.5 hours is work. Flow efficiency, work divided by elapsed time, is 5 percent. Making anyone type faster would change almost nothing; the time is in queues created by handoffs. After the team adopts automated tests, a deployment pipeline with canary analysis and same-day review, the same change takes about 8 hours end to end. That is the core insight of the First Way: look for the waits before you optimise the work.
From principles to practices
| Principle | Practice | Where to learn it |
|---|---|---|
| Small batches, fast flow | Trunk-based development, short-lived branches | Git Branching Strategies |
| Remove handoffs, build quality in | Continuous integration and delivery, build once and promote | CI/CD, in depth |
| Repeatable environments | Containers and declarative infrastructure | Docker Introduction |
| Self-healing operations | Orchestration with desired state and health probes | Kubernetes Introduction |
| Fast feedback from production | Metrics, logs, traces, SLOs and alerting | Your observability stack |
| Learning from failure | Blameless postmortems and tracked actions | Covered below |
Notice the order: each practice exists to serve a principle. A team that adopts Kubernetes while still releasing once a quarter through a change board has adopted a tool, not the model.
Who runs what
Werner Vogels described Amazon's model in 2006 as "you build it, you run it": the team that writes a service operates it, including on-call. That ownership creates the strongest feedback loop there is, since developers feel every 3 a.m. page caused by their design. Three common structures implement it at scale.
- Product teams owning services end to end, with on-call rotations inside the team.
- Site reliability engineering, Google's model, where engineers with software skills own reliability, often sharing operational load with product teams under explicit agreements such as error budgets.
- Platform teams, in the vocabulary of Team Topologies by Matthew Skelton and Manuel Pais, who build an internal self-service platform so that stream-aligned product teams can deploy and operate without filing tickets.
The anti-pattern is a new team called DevOps that sits between development and operations and receives tickets from both. That adds a third silo to the original two.
Error budgets: a contract between speed and stability
A service level objective states the reliability users need, for example 99.9 percent of requests succeeding over 30 days. The remaining 0.1 percent is the error budget. While budget remains, the team ships freely; when it is spent, effort shifts to reliability until it recovers. This turns the old argument between development and operations into arithmetic both sides agreed in advance.
def error_budget_minutes(slo, window_days=30):
return window_days * 24 * 60 * (1 - slo)
def burn_rate(bad, total, slo):
"""1.0 spends the budget exactly over the window; 10.0 spends it ten times faster."""
return (bad / total) / (1 - slo)
print(round(error_budget_minutes(0.999), 1)) # 43.2 minutes per 30 days
print(round(error_budget_minutes(0.995) / 60, 1)) # 3.6 hours per 30 days
print(round(burn_rate(bad=120, total=40_000, slo=0.999), 1)) # 3.0: alert-worthy if sustainedA burn rate above 1 means the budget will run out before the window ends. Alerting on a high burn rate over both a short and a long window catches fast outages and slow leaks while ignoring brief noise.
Measuring delivery: the DORA metrics
The DORA research programme identifies a small set of software delivery metrics. In its current form there are five, grouped into throughput and instability. Throughput: change lead time, from commit to running in production; deployment frequency; and failed deployment recovery time. Instability: change fail rate, the share of deployments needing immediate intervention; and deployment rework rate, the share of deployments that were unplanned and made because of a production incident.
from dataclasses import dataclass, field
from datetime import datetime
from statistics import median
@dataclass
class Deploy:
at: datetime
commit_times: list = field(default_factory=list) # commit time of every change included
failed: bool = False # needed immediate intervention
recovered_at: datetime | None = None
unplanned: bool = False # deployed because of a production incident
def hours(td):
return round(td.total_seconds() / 3600, 1)
def dora(deploys, start, end):
ds = [d for d in deploys if start <= d.at < end]
if not ds:
return {}
lead = [d.at - c for d in ds for c in d.commit_times]
failed = [d for d in ds if d.failed]
recovery = [d.recovered_at - d.at for d in failed if d.recovered_at]
return {
"deployment_frequency_per_day": len(ds) / max((end - start).days, 1),
"change_lead_time_median_h": hours(median(lead)) if lead else None,
"change_fail_rate": len(failed) / len(ds),
"failed_deployment_recovery_median_h": hours(median(recovery)) if recovery else None,
"deployment_rework_rate": sum(d.unplanned for d in ds) / len(ds),
}The data comes from your version control and deployment pipeline, plus a reliable way to mark a deployment as failed and recovered, usually the incident tool. Use the metrics to watch a team's own trend and to find bottlenecks, never to rank teams or set individual targets. A metric that becomes a target gets gamed: deployment frequency rises when changes are split meaninglessly, and change fail rate falls when failures stop being recorded.
Blameless postmortems
After any significant incident, write a short document: what users experienced, a timeline, contributing factors, what went well, and actions with owners and dates. Blameless means asking how the system allowed a reasonable person to make that mistake, not who made it. People who fear blame hide information, and hidden information is how incidents repeat. Track postmortem actions like any other work, and review them: an action list that never closes is a sign the learning loop is broken.
A 90-day starting plan
- Weeks 1-2. Map the value stream for three recent changes and compute flow efficiency. Start recording the DORA metrics, even by hand.
- Weeks 3-6. Attack the longest wait. Usually that means automating the test suite that a manual QA step stands in for, or making review same-day.
- Weeks 7-10. Build a deployment pipeline that the team uses for every change, with automated rollback, and give the team production access and an on-call rotation.
- Weeks 11-13. Define one SLO for the most important user journey, alert on burn rate, and run a blameless postmortem for the next incident. Remap the value stream and compare.
Anti-patterns
| Anti-pattern | Why it fails | Instead |
|---|---|---|
| A DevOps team between dev and ops | Adds a handoff and a silo | Platform team that provides self-service |
| Tools first | Automates the existing slow process | Map the value stream, then choose tools |
| Metrics as targets | Gaming, hidden failures | Use trends for learning, not ranking |
| Automating a broken process | Faster production of bad outcomes | Simplify, then automate |
| On-call without authority | Pages people who cannot fix the cause | Pair on-call with ownership of the code |
| Security as a final gate | Becomes the new change board | Build scanning and policy checks into the pipeline |
Trade-offs
Teams that run their services take on operational load and need investment in tooling, training and sane on-call. Platform teams reduce duplication but can become a bottleneck if they turn back into ticket queues. In heavily regulated environments, approval steps do not disappear; they move into the pipeline as automated, auditable checks. And none of this works without management support, because the incentives that created the wall are set above the teams.
What to do next
- Map the value stream of your last three changes and compute flow efficiency.
- Instrument the five DORA metrics from your repository and deployment tool, using the code above as a starting point.
- Remove the single longest wait you found, then measure again.
- Define one SLO and error budget for your most important user journey and agree what happens when the budget is spent.
- Make the team that builds a service part of its on-call rotation, with authority to fix what pages them.
- Hold a blameless postmortem for your next incident and track its actions to completion.