CI/CD is the machinery that turns a commit into running software, plus the habits that keep that machinery fast and trustworthy. Done well, it lets a team ship small changes many times a day with less risk than shipping large ones once a month. Done badly, it becomes a 50-minute queue of flaky tests that people rerun until green, followed by a deploy script that one person understands.
This article builds a pipeline from first principles: what continuous integration and continuous delivery actually require, a pipeline architecture that builds once and promotes the same artefact, a complete worked workflow, how to keep tests fast and believable, deployment strategies, database migrations that survive rollback, numeric canary analysis, and securing the pipeline itself. It ends with failure modes, the metrics that show whether any of it is working, and a checklist.
What the terms mean
Continuous integration is a practice before it is a tool: everyone merges small changes into the mainline at least daily, and an automated build and test run verifies every merge, so integration problems are found while small. Long-lived feature branches plus a CI server is not continuous integration; it is automated testing of branches that will conflict later.
Continuous delivery means every change that passes the pipeline is releasable with a routine, push-button operation. Continuous deployment releases every passing change automatically, with no human gate. Either way, the path to production must be automated, repeatable and identical for every change, urgent fixes included.
Two rules follow. The mainline must stay releasable, so a broken build is the team's top priority. And the pipeline is the only route to production, because every manual side door becomes the route used in emergencies, exactly when mistakes are most likely.
Pipeline architecture
A pipeline has two halves joined by an artefact registry. The integration half runs on every change: check out, lint and type-check, run unit tests, build the artefact, scan it, and publish it with an immutable identity. The delivery half takes that identity through environments: deploy to staging, run integration and end-to-end tests, pass a gate, deploy to production progressively while comparing health, then complete or roll back. Production telemetry feeds the gate decisions.
Build once, promote by digest
The most important delivery rule is that the artefact tested in staging is byte-for-byte the one deployed to production. Rebuilding per environment invites drift: an updated base image, a dependency range resolving differently, a different build flag. Every container image has a content digest, a SHA-256 hash of its manifest; tags such as latest can move, a digest cannot. Deploy by digest, record it in every deployment, and promotion becomes running the same digest in the next environment.
Environment differences belong in configuration injected at deploy time, such as environment variables, mounted files or a configuration service, never in the build. Secrets are fetched at runtime from a secret store. Tag images with the commit SHA for humans, and let machines use the digest.
A worked pipeline
Here is a complete GitHub Actions workflow for a containerised Python service. Pull requests run the tests. Merges to main build and push one image, deploy that exact digest to staging, run smoke tests, and then, after the production environment's protection rules pass, start a canary in production. The action versions are major versions that exist at the time of writing; in practice, pin third-party actions to a full commit SHA.
name: ci-cd
on:
pull_request:
push:
branches: [main]
permissions:
contents: read # least privilege by default; jobs add what they need
concurrency:
group: ci-cd-${{ github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }} # never cancel deploys
jobs:
test:
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
- run: pip install -r requirements.txt -r requirements-dev.txt
- run: ruff check .
- run: pytest -q -n auto --maxfail=20
build:
if: github.event_name == 'push'
needs: test
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
outputs:
digest: ${{ steps.push.outputs.digest }}
steps:
- uses: actions/checkout@v4
- uses: docker/setup-buildx-action@v3
- uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- id: push
uses: docker/build-push-action@v6
with:
push: true
tags: ghcr.io/${{ github.repository }}:${{ github.sha }}
cache-from: type=gha
cache-to: type=gha,mode=max
deploy-staging:
needs: build
runs-on: ubuntu-latest
environment: staging
permissions:
contents: read
id-token: write # OIDC: short-lived cloud credentials, no stored keys
steps:
- uses: actions/checkout@v4
- run: ./deploy/deploy.sh staging "ghcr.io/${{ github.repository }}@${{ needs.build.outputs.digest }}"
- run: ./deploy/smoke_test.sh staging
deploy-production:
needs: [build, deploy-staging]
runs-on: ubuntu-latest
environment: production # protection rules: required reviewers, branch policy
permissions:
contents: read
id-token: write
steps:
- uses: actions/checkout@v4
- run: ./deploy/canary.sh production "ghcr.io/${{ github.repository }}@${{ needs.build.outputs.digest }}"Points worth noticing. Superseded runs are cancelled on pull requests but never on main, so no deploy dies halfway. The default token is read-only and only the build job may write packages. Both deploy jobs use the digest the build job outputs, not the tag. The OIDC token is exchanged by a cloud provider configured to trust this repository and environment for short-lived credentials, so no cloud keys live in secrets. Production is a protected environment gated by reviewers or a wait timer. Caveats: pytest -n auto needs pytest-xdist, and image names must be lower case.
Fast tests that people believe
A pipeline helps only if it is fast enough to run on every change and trusted enough that red stops people. A common target is pull-request feedback within about ten minutes. Use many fast unit tests, fewer integration tests against real dependencies in containers, and a handful of end-to-end journeys; shard across runners, cache by lockfile hash, and run expensive suites only when relevant paths change.
A flaky test, one that passes and fails on the same code, teaches everyone to rerun until green, and then real failures get rerun too. Record results per commit, flag tests with both outcomes on one commit, quarantine them in a non-blocking job with an owner and deadline, and delete tests nobody fixes. Usual causes are shared state, real clocks and sleeps, ordering assumptions and network calls.
import collections, json, sys
def flaky_tests(results_path):
# JSON lines {"commit": ..., "test": ..., "outcome": "passed" or "failed"} from many CI runs.
outcomes = collections.defaultdict(set)
with open(results_path, encoding="utf-8") as fh:
for line in fh:
r = json.loads(line)
outcomes[(r["commit"], r["test"])].add(r["outcome"])
mixed = collections.Counter(test for (commit, test), seen in outcomes.items()
if {"passed", "failed"} <= seen)
return mixed.most_common()
for test, commits in flaky_tests(sys.argv[1]):
print(f"{commits:4d} commits with mixed results {test}")
Deployment strategies
| Strategy | How it works | Rollback | Costs and caveats |
|---|---|---|---|
| Rolling | Replace instances in batches | Redeploy the previous version; as slow as the rollout | Cheap; two versions serve together and must be compatible |
| Blue/green | Start the new version beside the old, switch all traffic at once | Switch back | Double capacity during the switch; the database is shared |
| Canary | Send a small share of traffic to the new version, compare, expand | Route the share back | Needs good metrics and enough traffic to compare |
| Feature flags | Deploy code dark, enable behaviour per user or percentage | Turn the flag off | Separates deploy from release; old flags become debt |
These combine. A typical web service deploys by canary and releases risky features behind feature flags, so a bad deploy is caught by metrics and a bad feature is switched off without a deploy. Blue/green suits services where mixed versions are hard to support. Every strategy except a full stop-and-replace runs two versions at once, which is why migrations need care.
Database migrations that survive rollback
Code rolls back in seconds; a dropped column does not come back. The expand and contract pattern splits every breaking schema change into backward-compatible steps, each deployed separately. To rename a column from email to email_address:
- Expand: add email_address as a nullable column. Old code ignores it.
- Write both: deploy code that writes both columns and still reads email.
- Backfill: copy existing values in throttled batches so the database keeps serving.
- Read new: deploy code that reads email_address. Rolling this back is safe because both columns are still written.
- Stop writing old: deploy code that writes only email_address.
- Contract: once nothing reads email, drop it in a later release.
Each step can be rolled back on its own, and at every moment the running versions agree on the schema. Run migrations as their own pipeline step before the code that needs them, never at application start-up on every instance, and make each migration idempotent so a retry is harmless.
Automated canary analysis
A canary is only as good as its decision rule. Compare the canary with a baseline running the old version on the same traffic at the same time, not with last week, and require enough requests before deciding. A minimal rule for error rate requires the canary to be both significantly worse, by a two-proportion z-test, and materially worse, by a ratio, before failing it.
from math import sqrt
def canary_verdict(base_err, base_n, can_err, can_n, min_n=2000, max_ratio=1.5, z_crit=2.33):
# Error counts and request counts over the SAME window. Returns 'wait', 'fail' or 'pass'.
if can_n < min_n or base_n < min_n:
return "wait"
p_b, p_c = base_err / base_n, can_err / can_n
pooled = (base_err + can_err) / (base_n + can_n)
se = sqrt(pooled * (1 - pooled) * (1 / base_n + 1 / can_n)) or 1e-9
z = (p_c - p_b) / se
if z > z_crit and p_c > max_ratio * max(p_b, 1e-4):
return "fail" # significantly worse AND materially worse
return "pass"
print(canary_verdict(42, 20000, 19, 2500)) # fail: 0.76% vs 0.21%, z is about 5.0Worked through: the baseline has 42 errors in 20,000 requests (0.21%) and the canary 19 in 2,500 (0.76%). The pooled rate of about 0.27% gives a standard error of about 0.11 points, so the 0.55-point gap is roughly 5 standard errors, and the ratio is 3.6. Both conditions hold, so the canary is rolled back. Requiring both avoids failing tiny significant differences at high traffic and passing large noisy ones at low traffic. Apply similar rules to latency and saturation, and keep thresholds strict, since repeated checks during a ramp raise the false alarm rate.
Securing the pipeline
The pipeline holds the keys to production, which makes it a target. Attacks on build systems, dependencies and CI actions are common; the controls below address the usual paths, and the supply chain article goes further.
- Least privilege: make the default workflow token read-only and grant write permissions per job, as the example does.
- Short-lived credentials: use OIDC federation to your cloud instead of stored access keys, and scope the trust to repository, branch and environment.
- Pin what you run: third-party actions by full commit SHA, dependencies by lockfile, and update both through reviewed pull requests from a dependency bot.
- Untrusted code: never run a fork's code with secrets. Workflows triggered by
pull_request_targetrun with the base repository's privileges, so checking out and executing the fork's code there hands it your secrets. - Provenance: produce signed build provenance and an SBOM for every artefact and verify signatures at deploy time. SLSA defines graded build levels, starting with provenance and moving to hardened, isolated build platforms.
- Protect the path: require reviews and passing checks on main, and approvals for production environments.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| People rerun until green | Flaky tests | Per-commit mixed-outcome report, quarantine with owners |
| Works in staging, fails in production | Rebuilt artefact or configuration drift | Promote one digest; inject configuration per environment |
| Rollback fails | Migration not backward compatible | Expand and contract; migrations as a separate step |
| Pipeline takes 45 minutes | Serial suites, no caching, everything runs on every change | Shard, cache by lockfile hash, path filters |
| Deploys interleave | No concurrency control | Concurrency group per environment; never cancel a deploy |
| Secret leaked from CI | Long-lived keys, fork code run with secrets | OIDC, scoped permissions, no secrets for fork code |
Measuring whether it works
The DORA research programme popularised four key delivery metrics: deployment frequency, lead time for changes from commit to production, change failure rate, and time to restore service after a failure. They work as a set: speed without stability rewards recklessness, and stability without speed rewards never shipping. Compute them from pipeline and incident data rather than surveys, and watch trends per service rather than ranking teams. Alongside them, track pipeline duration at the median and 95th percentile, the number of quarantined tests, and the share of rollbacks triggered by automation rather than people.
Trade-offs
Continuous deployment removes the release bottleneck but demands strong tests, canaries and flags; where regulation or contracts require a human sign-off, keep continuous delivery with a fast approval step. A monorepo with one pipeline makes cross-cutting changes easy but needs path-based selection to stay fast, while many small repositories are fast by default and make coordinated changes harder. More gates catch more problems and slow everything, so prefer automated gates with numeric rules and reserve manual approval for changes that need judgement. Machine learning systems add model and data changes to the stream; the MLOps article covers how CI, CD and continuous training divide the work. Where you cannot deploy continuously, such as mobile apps behind store review, a fixed-cadence release train is a sound alternative.
What to do next
- Measure your current deployment frequency, lead time, change failure rate and time to restore, so improvement is visible.
- Make the pipeline the only route to production and remove manual deploy paths.
- Build once, push by digest, and promote the same digest through every environment.
- Bring pull-request feedback under ten minutes with sharding, caching and path filters.
- Start a flaky-test report from per-commit results and quarantine flaky tests with owners.
- Use expand and contract for every breaking schema change, with migrations as a separate step.
- Add a canary with a written numeric rule and automatic rollback.
- Lock down the pipeline: read-only default token, OIDC, pinned actions, no secrets for fork code, signed provenance.