Ask who owns an LLM service and you usually get a team name. Ask who owns the decision to move its KV cache to FP8, who approves a new CUDA driver on the inference nodes, or who decides that a safety classifier's threshold changes, and the answers fragment. The model team trained the weights, the serving team runs the engine, the platform team owns the nodes, the safety team owns the filters, and product owns the quota tiers. During an incident that fragmentation costs minutes; during a change it causes the regressions that no single team would have approved.
An ownership matrix fixes that by listing every asset and decision in the stack against every team, with a RACI letter in each cell. This article treats the matrix as an engineering artefact rather than a slide. It covers the RACI rules, a matrix for a GPU serving stack, an ownership file validated in CI, how to feed it into the service catalog, CODEOWNERS, alert routing and the incident bot, a worked example of a regression that crossed three teams, and how to keep the file from going stale.
RACI rules that actually hold
RACI assigns four roles per row. Responsible does the work; there can be several. Accountable makes the final call and answers for the outcome; there must be exactly one, and it must be a team or role with an on-call rotation, not a committee. Consulted must be asked before the change and can block it on evidence; keep this list short, because each name adds latency. Informed is told after the change, usually automatically.
The rule that does the most work is one Accountable per row. Two Accountables means each assumes the other will decide, and nothing is decided. Zero means the row is owned by whoever is paged first. The second most useful rule is that rows are assets or decisions, never vague areas: "inference" is not a row, but "engine version and launch flags for the chat pool" is.
A matrix for a GPU serving stack
The rows below cover a typical self-hosted serving stack. Columns are teams: Model trains and evaluates weights, Serving runs the inference engine and its configuration, Platform owns GPU nodes, drivers and the scheduler, Safety owns filters and policy, Product owns tiers and quotas, and FinOps owns budgets.
| Asset or decision | Model | Serving | Platform | Safety | Product | FinOps |
|---|---|---|---|---|---|---|
| Base weights and adapters in the registry | A/R | C | I | C | I | - |
| Release eval gate and thresholds | A/R | C | - | C | C | - |
| Engine version and launch flags | C | A/R | C | - | - | - |
| Weight and KV cache quantisation | C | A/R | - | - | I | I |
| KV cache memory budget and max context | C | A/R | C | - | C | - |
| GPU driver, CUDA and NCCL versions | - | C | A/R | - | - | - |
| Node pools, autoscaling and placement | - | C | A/R | - | - | C |
| Prompt templates and system prompts | C | - | - | C | A/R | - |
| Input and output safety classifiers | C | C | - | A/R | C | - |
| Rate limits and quota tiers | - | C | - | C | A/R | C |
| Latency SLO definitions | - | A/R | C | - | C | - |
| GPU spend budget and reservations | - | C | C | - | C | A/R |
Two rows deserve attention. Quantisation sits with Serving because it is a serving-time choice with throughput consequences, but Model is Consulted because only the eval suite can say whether quality held. The KV cache budget is the seam between three teams: Serving sets it, Platform must confirm it fits the GPU memory alongside activations and the CUDA graphs, and Product is Consulted because maximum context length is a customer-visible feature. Seams like these are where unowned regressions are born, so name them explicitly.
Ownership as code
A matrix in a wiki goes stale the week after a reorganisation. Keep it as a file in a repository instead, next to the deployment configuration it describes, and let CI reject changes that break the rules. Each entry names an asset or decision, its RACI teams, the file paths whose changes need the Accountable team's review, the alerts that page it, and a review date.
# ownership.yaml - one entry per asset or decision
- id: serving.chat.engine
description: Engine version and launch flags for the chat pool
accountable: team-serving
responsible: [team-serving]
consulted: [team-model, team-platform]
informed: [team-product]
paths: [deploy/chat/engine.yaml] # files whose changes need A's review
alerts: [ChatITLBurn, ChatEngineCrashLoop]
review_by: 2027-01-15
- id: platform.gpu.driver
description: GPU driver, CUDA and NCCL versions on inference nodes
accountable: team-platform
responsible: [team-platform]
consulted: [team-serving]
informed: [team-model]
paths: [infra/nodes/gpu-image/**]
alerts: [GPUXidErrors, NodeDriverMismatch]
review_by: 2026-12-01The validator enforces the rules from the RACI section mechanically. It checks that every row has exactly one Accountable team, that the team exists in the directory export and has an on-call rotation, that the Consulted list stays short, that review dates have not passed (a warning, not a failure), and that no alert is claimed by two rows. Run it on every change to the ownership file and nightly, because the team directory changes underneath it.
# validate_ownership.py - fails CI on ambiguous or stale ownership
import datetime, sys, yaml
rows = yaml.safe_load(open("ownership.yaml"))
teams = yaml.safe_load(open("teams.yaml")) # exported from the team directory
errors, warnings, seen_ids, alert_owner = [], [], set(), {}
for r in rows:
rid = r.get("id", "?")
if rid in seen_ids:
errors.append(f"{rid}: duplicate id")
seen_ids.add(rid)
a = r.get("accountable")
if not isinstance(a, str):
errors.append(f"{rid}: exactly one accountable team required")
elif a not in teams or not teams[a].get("oncall_rotation"):
errors.append(f"{rid}: accountable {a} missing or has no on-call rotation")
for role in ("responsible", "consulted", "informed"):
for t in r.get(role, []):
if t not in teams:
errors.append(f"{rid}: {role} team {t} not in directory")
if len(r.get("consulted", [])) > 3:
errors.append(f"{rid}: more than 3 consulted teams slows every change")
if r["review_by"] < datetime.date.today():
warnings.append(f"{rid}: review date {r['review_by']} has passed") # ticket, not a block
for alert in r.get("alerts", []):
if alert in alert_owner:
errors.append(f"{rid}: alert {alert} already owned by {alert_owner[alert]}")
alert_owner[alert] = rid
print("\n".join(warnings)) # nightly job files these as tickets
if errors:
print("\n".join(errors))
sys.exit(1)
Wiring the matrix into the tools
Ownership is only useful where decisions happen, so generate outputs from the file rather than asking people to consult it. Four consumers cover most needs.
- Service catalog. If you run Backstage, each component's
catalog-info.yamlcarriesspec.ownerandspec.system. Generate those fields from the Accountable team so the catalog and the matrix cannot disagree. - Review gates. GitHub and GitLab both support a CODEOWNERS file that requires review from named owners for matching paths. Generate entries from each row's
paths, so a change to the chat engine flags needs the serving team's approval and a change to the node image needs the platform team's. - Alert routing. Give every alert rule an asset label and route on it: the router looks the asset up in the generated mapping and pages that row's rotation. An alert without an owning row fails the validator, so no alert lands in a shared catch-all channel.
- Incident bot. When an incident is declared against an asset, the bot pages the Accountable rotation, invites the Consulted teams to the channel and posts to the Informed teams' feeds.
Worked example: a regression across three teams
A Thursday change window ships two things: the platform team rolls a new GPU driver to the inference node image, and the serving team raises the engine's maximum batched tokens to improve throughput. By Friday, p95 inter-token latency on the chat pool has doubled at peak and the burn-rate alert fires.
Without a matrix, the alert lands in a shared channel, the serving on-call suspects the driver, the platform on-call points out that the driver passed its node health checks, and an hour passes before anyone rolls anything back. With the matrix, the alert ChatITLBurn routes to the serving rotation because it belongs to serving.chat.engine. The incident bot invites the platform team as Consulted on that row. The serving on-call, who is Accountable, can roll back their own change at once while the incident commander rolls back the driver in parallel.
Bisecting both rollbacks shows the batching change was the cause: a larger token budget per step let more prefill work into each batch, so decode steps for in-flight requests waited behind it and streaming slowed for everyone. The postmortem adds a row the matrix lacked, serving.chat.batching_limits, with Product Consulted because the latency tier is customer-visible. It also adds a change-window rule that two Accountable teams must not change the same pool in the same window without a joint plan. The matrix improved because an incident exposed a seam.
Keeping it fresh
Staleness is the main way ownership matrices die. Three mechanisms keep this one current. Review dates per row force a quarterly look; the nightly validator turns a passed date into a ticket for the Accountable team, not a hard failure, so a holiday does not block deploys. The team directory export catches renamed and dissolved teams the night they change. And the incident review asks one standing question: did the page go to the team that fixed it? Every "no" is a row to edit.
Reorganisations deserve a planned migration. Generate a diff of every row whose Accountable team is changing, have both the outgoing and incoming leads approve it in one pull request, and merge it on the day the rotation changes, not before.
Decision rights during an incident
The matrix describes steady-state decision rights. Incidents need one deliberate exception: the incident commander may roll back any change on the affected asset without waiting for the Consulted teams, and the Accountable team reviews the rollback afterwards. Write that exception into the file's header so nobody hesitates at 3 a.m. over whether rolling back the platform team's driver needs the platform team's approval. It does not, as long as the rollback restores a previously released state.
The reverse exception matters too. Forward fixes during an incident, such as a new engine flag, a lower maximum context or a different quantisation, are changes, not rollbacks, and they still need the Accountable team for that row. A forward fix that skips its owner is how a latency incident becomes a quality incident: the serving on-call halves the maximum context length so fewer long requests compete for KV cache, preemption stops, and an hour later long-document customers start reporting rejected requests that nobody connects to the fix.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Committee as Accountable | Decisions wait for a meeting | Validator requires one team with a rotation |
| Rows are areas, not assets | Every incident starts with an argument | Split rows until each maps to files or alerts |
| Consulted list grows | Simple changes wait days | Cap Consulted; move the rest to Informed |
| Seams unowned | Cross-team regressions with no owner | Name seam rows such as KV budget and batching limits |
| Matrix copied into slides | Two versions disagree | Generate every output from the file |
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Fine-grained rows | Precise routing and review | More rows to keep fresh |
| Hard CI gate on ownership | Ambiguity never merges | Directory outages can block changes |
| Generated CODEOWNERS | Review follows the matrix | Owners become reviewers on every path match |
| Seam rows with joint Consulted | Cross-team risks get a name | Slightly slower changes at seams |
What to do next
- List the assets and decisions in your serving stack, from weights to quotas, as concrete rows.
- Assign exactly one Accountable team with an on-call rotation to each row.
- Move the matrix into an
ownership.yamlbeside your deployment config, with paths, alerts and review dates. - Add the validator to CI and run it nightly against the team directory export.
- Generate catalog owners, CODEOWNERS, alert routes and incident bot lookups from the file.
- After every incident, ask whether the page reached the team that fixed it, and edit the matrix when it did not.
Keep learning: LLM change management for release gates the matrix plugs into, the runbook index for symptom-keyed runbooks per owner, incident communication channels for what the incident bot does next, and GPU capacity planning for the node and budget rows.