Most engineering organisations do not lack documentation. They lack documentation anyone trusts. There is a wiki with three pages on the same deploy process, a README that describes a build that no longer exists, and an architecture page whose author left two years ago. Engineers learn that docs are unreliable, stop reading them, ask in chat instead, and stop writing them. The decay feeds itself.
Culture here is not a poster about valuing documentation. It is a set of mechanisms: who owns each page, where docs live, what checks run on them, when they are reviewed, and what counts as done. This guide lays out those mechanisms with code you can adopt, a worked rollout for a mid-sized engineering organisation, the metrics that tell you whether it is working, and the common ways documentation efforts fail.
Why documentation rots
Documentation has a broken feedback loop. Code that is wrong fails a test or pages someone. Docs that are wrong fail silently: the cost lands on a reader months later, often a new hire who assumes the mistake is theirs. The writer pays the cost of writing now; the benefit goes to someone else later. Without a mechanism, rational engineers under deadline write less and update less.
Three structural causes make it worse. Docs live away from the code, so a change to one does not prompt a change to the other. Pages have no owner, so nobody is responsible when they are wrong. And there is no expiry, so a page written once stays published forever, looking exactly as authoritative as one updated yesterday. The fixes below target these three causes directly.
A map of document types
Different documents have different readers, lifespans and rules. The Diataxis framework, created by Daniele Procida, splits user-facing docs into four kinds by what the reader needs: tutorials for learning by doing, how-to guides for completing a specific task, reference for looking up facts, and explanation for understanding why. Mixing them is the most common writing problem: a reference page that wanders into a tutorial serves neither reader. Engineering teams add a few internal types:
| Type | Reader's question | Owner | Freshness rule |
|---|---|---|---|
| README | What is this and how do I run it? | Owning team of the repo | Checked on every change to build or run steps |
| How-to / runbook | How do I do X right now? | Team that operates the system | Review every 90 days and after each incident that used it |
| Reference (API, config) | What does this field do? | Generated from code where possible | Regenerated on release |
| Explanation / architecture | Why is it built like this? | Tech lead of the system | Review every 6 to 12 months |
| Design doc | What are we going to build and why? | Author | Frozen after decision; marked historical |
| Decision record (ADR) | Why did we choose this? | Deciding team | Immutable; superseded, never edited |
| Postmortem | What happened and what changed? | Incident owner | Immutable after review |
The last column is the important one. Some documents should be kept current; others are records of a moment and should say so in their first line. Architecture decision records, popularised by Michael Nygard in 2011, are useful because they are immutable: you never update one, you write a new one that supersedes it. That makes it safe to read old ones. For the documents in this table that have their own guides, see how to write a design doc, how to run a postmortem and runbooks.
Docs-as-code: the architecture
Docs-as-code means documentation is plain text, Markdown or similar, stored in version control next to what it describes, reviewed in pull requests, checked by CI and published by a build. Its main benefit is not tooling but proximity. When the change that renames a config flag and the doc that mentions it are in the same diff, the reviewer can see both.
Owner routing uses the code host's ownership file. On GitHub, CODEOWNERS uses last-match-wins: the last pattern that matches a path decides the owner. That means general rules go first and specific ones after:
# .github/CODEOWNERS -- later lines override earlier ones
*.md @acme/docs-reviewers
/docs/ @acme/docs-reviewers
/docs/payments/ @acme/payments
/docs/runbooks/ @acme/sre
/services/ledger/ @acme/paymentsReverse the order and the general *.md rule would swallow every specific one, quietly routing all doc reviews to one overloaded group. CI then runs cheap, objective checks: a link checker such as lychee for broken links, a prose linter such as Vale for terms you have banned or renamed, and the freshness check below. Keep subjective style out of CI; a docs check that fails on taste teaches people to bypass docs.
Freshness: make expiry explicit
Every page that is meant to stay current carries an owner and a review-by date in front matter. A scheduled job finds pages past their date and opens an issue for the owner. The owner then does one of three things: confirms the page is still right and moves the date, fixes it, or archives it. Archiving is a success, not a failure.
---
title: Rotate the ledger database credentials
type: runbook
owner: "@acme/payments"
review_by: 2026-12-15
---import datetime as dt, pathlib, re, sys
import yaml
FRONT = re.compile(r"\A---\n(.*?)\n---\n", re.S)
IMMUTABLE = {"adr", "postmortem", "design"}
today = dt.date.today()
stale, ownerless = [], []
for path in pathlib.Path("docs").rglob("*.md"):
m = FRONT.match(path.read_text(encoding="utf-8"))
meta = yaml.safe_load(m.group(1)) if m else {}
if meta.get("type") in IMMUTABLE:
continue
if not meta.get("owner"):
ownerless.append(str(path))
continue
due = meta.get("review_by")
if due is None or dt.date.fromisoformat(str(due)) < today:
stale.append((meta["owner"], str(path), str(due)))
for owner, path, due in sorted(stale):
print(f"STALE\t{owner}\t{path}\treview_by={due}")
for path in ownerless:
print(f"OWNERLESS\t{path}")
sys.exit(1 if ownerless else 0)Notice the policy encoded in the exit code. An ownerless page fails CI, because ownership is cheap to add and everything else depends on it. A stale page does not fail the build; it creates work for the owner through an issue. Failing builds on staleness blocks unrelated changes and makes teams resent the system. Stale pages can also show a visible banner on the docs site, so readers know to be careful.
Making docs part of done
The most effective single change is to treat docs as part of the change that makes them necessary. Put a checkbox in the pull request template, and back it with a heuristic check so it is not just a ritual:
import os, subprocess, sys
diff = subprocess.run(["git", "diff", "--name-only", "origin/main...HEAD"],
capture_output=True, text=True, check=True).stdout.split()
labels = os.environ.get("PR_LABELS", "")
public_surface = [f for f in diff if f.startswith(("api/", "openapi/", "config/schema/"))]
docs_touched = any(f.startswith("docs/") or f.endswith("README.md") for f in diff)
if public_surface and not docs_touched and "no-docs-needed" not in labels:
print("This PR changes a public surface but no docs:", *public_surface, sep="\n ")
print("Update docs, or add the 'no-docs-needed' label with a reason in the description.")
sys.exit(1)The escape label matters. Many API changes need no doc change, and a check without an exit trains people to make meaningless edits. With the label, skipping docs becomes a visible decision that a reviewer can question.
Worked example: a 40-engineer rollout
Consider an organisation of 40 engineers in six teams with about 600 wiki pages, a median onboarding time of five weeks to a first meaningful pull request, and a chat channel where the same deployment questions come up every week. A one-quarter rollout:
- Weeks 1 to 2, triage. Export the wiki and sort pages by last edit and view count. In organisations like this, a large share of pages typically have neither recent edits nor recent views. Archive those in bulk with a banner, rather than migrating them.
- Weeks 3 to 6, move what matters. Move the most-viewed pages, often fewer than a hundred, into the repositories they describe, adding owner and review-by front matter. Each team owns its own pages; nobody migrates another team's docs.
- Weeks 5 to 8, turn on checks. Add CODEOWNERS, the link checker, the freshness job and the docs-changed check, all as warnings at first.
- Weeks 9 to 12, close the loop. Turn the ownerless check into a failure. Answer repeated chat questions with a link to a page, writing the page if it does not exist. Have each new hire fix one thing they found wrong during onboarding.
The new-hire rule is the cheapest high-yield practice here. New people are the only readers who follow docs literally, so they find the errors that experienced engineers skip over without noticing.
Writing norms that make docs usable
- Lead with the task. The first lines of a how-to say what it achieves and what you need before starting.
- Show, then explain. A command that works beats a paragraph describing it. Use real, copyable examples.
- Link, do not copy. Every copied paragraph is a future contradiction. Keep one source of truth and link to it.
- Date facts that will change. "As of the 2026 Q3 migration" lets readers judge staleness themselves.
- Generate reference docs. API references, CLI help and config schemas should come from the code, so they cannot drift.
- Write for search. Use the words people will type, including error messages verbatim.
Docs and coding agents
Coding agents now read your documentation as well. Instruction files, READMEs and architecture notes become part of the context an agent uses to change your code. Wrong docs used to cost a confused human an afternoon; now they produce confidently wrong code at scale. The same mechanisms apply to agent context: owners, review dates, and keeping files short and specific. Instruction files for coding agents covers how to write them. Agents also make one practice cheaper: asking an agent to compare a doc with the code it describes and list contradictions is a useful first pass in a freshness review, as long as a human owner confirms each finding.
Measuring whether it works
Page count and edit count are vanity metrics; a doc day can inflate both while making things worse. Measure what readers experience instead:
- Share of pages past their review-by date, and share without an owner. Both should trend down and stay low.
- Searches on the docs site that return nothing or end without a click, which are a list of missing pages.
- Repeated questions in support channels, sampled monthly.
- Time from a new hire's start to their first merged change, tracked across cohorts.
- Broken links found by the checker per week.
Review these quarterly with team leads, alongside other engineering health signals. The goal is a trend, not a target to game.
Failure modes and trade-offs
- Doc days. A one-off push produces a burst of pages with no owners or review dates, and in a year they are part of the problem. Do continuous mechanisms instead.
- Wiki versus repository. Docs-as-code adds friction for non-engineers such as support and product staff. A reasonable split: engineering docs in repositories, cross-functional process docs in the wiki, both with owners and expiry.
- Templates as bureaucracy. A 14-section mandatory template produces 14 thin sections. Keep templates short and let authors delete headings that do not apply.
- Over-documenting. Documenting what clear code already says doubles the maintenance. Document the why, the operations and the edges; leave the what to the code. For older systems where the why is lost, working with legacy code has techniques for recovering it.
- Ownership by individuals. People leave. Assign pages to teams, never to a person.
What to do next
- Pick your most important service and list its docs; mark each with a type from the table above.
- Add owner and review-by front matter to those pages, and a CODEOWNERS entry with general rules first.
- Run the freshness script in CI as a warning and look at the first report.
- Archive every page nobody has viewed or edited in a year, with a visible banner.
- Add the docs-changed check with an escape label to the repository with the most public API changes.
- Ask your next new hire to fix one wrong thing they find in the docs during their first week.
- Start tracking the stale-page share and failed searches, and review them next quarter.