AI governance research asks what institutions, rules and technical mechanisms would make advanced AI development go well, and how anyone would check that they are working. It sits between technical safety research, which studies the models, and law, which binds people. Its outputs are papers, but they arrive in engineering teams a year or two later as thresholds, reporting duties and evaluation requirements. The compute threshold in the EU AI Act and the frontier-model definition in California's SB 53 both trace back to research arguing that training compute is a workable handle on capability.
This page explains the field from an engineer's side: the questions it works on, who produces it, the four levers it studies, how a finding becomes an obligation, and how to build the evidence those obligations will ask for. The code is a training-compute ledger, because compute accounting is the most direct place where governance research has already turned into numbers you must be able to produce. For intake of security research specifically, see AI security research organisations.
What governance research studies
Governance research is a set of questions, not a single discipline. The recurring ones are:
- What to govern. Models, the compute used to train them, the data, the deployments, or the organisations. Each target has different observability.
- How to measure. Which quantities track risk well enough to write into rules: training compute, benchmark scores, dangerous-capability evaluations, deployment reach.
- Who verifies. Self-reporting, third-party auditors, government institutes, or technical mechanisms such as hardware attestation.
- Who is liable. How responsibility splits between model developers, deployers and users when harm occurs.
- How countries coordinate. Which international institutions are feasible and what they could verify.
It differs from alignment research, which asks how to make a model behave as intended. Governance research takes the state of that science as an input and asks what rules make sense given it. It also differs from policy advocacy, although the line blurs: a good paper states its assumptions and what evidence would change its recommendation.
Who produces it, and how to weigh them
Several kinds of producer matter. The Centre for the Governance of AI (GovAI) began in 2018 inside Oxford's Future of Humanity Institute, became an independent nonprofit in 2021, and continued after the institute closed in April 2024; it is no longer an Oxford unit, whatever older summaries say. Georgetown's Center for Security and Emerging Technology (CSET) publishes data-heavy work on compute supply chains, talent and national security. RAND publishes policy analysis, much of it on security and national-security risks. Concordia AI, based in Beijing, works on AI safety and governance with a focus on China and international dialogue.
Governments now produce research too. The UK's AI Security Institute (renamed from the AI Safety Institute in 2025) and the US Center for AI Standards and Innovation run evaluations of frontier models. The International AI Safety Report, mandated at the 2023 Bletchley summit and chaired by Yoshua Bengio, synthesises the evidence; its 2026 edition appeared in February 2026. Frontier labs publish their own safety frameworks, compared in frontier labs. Treat each source by its incentives: a lab's framework is a commitment by the party being governed, and a think tank's report may be funded by parties with a view.
Lever one: compute governance
The most consequential line of work argues that compute is the best available handle. Sastry, Heim and co-authors (2024) set out why: compute is detectable (large clusters are physical and hard to hide), excludable (chips can be withheld) and quantifiable (operations can be counted). Training compute also correlates with capability, though imperfectly.
That argument became law. The EU AI Act presumes that a general-purpose model has systemic risk when its cumulative training compute exceeds 10^25 floating-point operations. The presumption is rebuttable, and the Commission can amend the threshold. California's SB 53, signed on 29 September 2025, defines a frontier model as one trained with more than 10^26 integer or floating-point operations, counting the initial run plus subsequent fine-tuning or material modifications. Its stricter duties fall on large frontier developers, those with over 500 million dollars in annual revenue, and most requirements apply from 1 January 2026.
The research also names the weaknesses, and they are worth knowing before you rely on the threshold. Algorithmic efficiency means the same capability needs less compute each year, so a fixed threshold catches less over time. Inference-time compute, used by reasoning models that think longer at run time, does not appear in a training count. And distillation can transfer capability from a large model into a small one trained with far less compute.
Lever two: evaluations
The second lever is evaluation. Shevlane and co-authors (2023), in "Model evaluation for extreme risks", argued for dangerous-capability evaluations (for example cyber-offence or biology uplift) and alignment evaluations, run before training scale-ups and deployment, with results feeding go or no-go decisions. That idea now structures the frontier safety frameworks that labs committed to publish at the Seoul summit in May 2024, and SB 53 requires large frontier developers to publish such a framework.
The open research problems are measurement problems. Evaluations underestimate capability when elicitation is weak: a better prompt or scaffold can unlock what a test missed. The 2026 International AI Safety Report notes that some models can tell evaluation contexts from deployment and change behaviour, which undermines a test's validity. And there is no settled method for turning evaluation scores into risk thresholds. For an engineer, the implication is to record how each evaluation was elicited, not just its score.
Levers three and four: transparency and accountability
The third and fourth levers are transparency and accountability. Research on incident reporting borrows from aviation and medicine: shared taxonomies, mandatory reporting of serious events and protection for reporters. SB 53 includes reporting of critical safety incidents to California's Office of Emergency Services and whistleblower protections, and the EU AI Act requires providers of systemic-risk models to track and report serious incidents. Anderljung and co-authors (2023), in "Frontier AI Regulation", proposed standard setting, registration and reporting, and compliance mechanisms for frontier developers, and much later legislation reads like a selection from that menu.
Liability research asks who pays when a deployed system causes harm, and whether developer liability would change behaviour more than prescriptive rules. It is the least settled area and the one most likely to affect product teams that deploy someone else's model, because contracts and insurance tend to move faster than statutes.
From finding to obligation
Every lever turns into the same kind of work once it is adopted. A research claim motivates a policy proposal, an instrument (law, standard or voluntary commitment) adopts a version of it, and the instrument creates an obligation that some team must satisfy with a control and an artefact. The feedback loop at the bottom matters: when practitioners cannot produce the evidence a proposal assumes, that gap is itself a research finding.
| Research lever | Instrument example | Engineering artefact |
|---|---|---|
| Compute governance | EU AI Act 10^25 presumption; SB 53 10^26 definition | Compute ledger per model lineage |
| Dangerous-capability evals | Published frontier safety frameworks | Eval reports with elicitation details |
| Incident reporting | SB 53 critical incidents; EU serious incidents | Incident taxonomy, timers, audit trail |
| Transparency | Framework and transparency report duties | Model and system documentation |
| Whistleblower protection | SB 53 protections | Internal anonymous reporting channel |
The governance programme that owns these controls is described in building an AI governance programme.
Code: a training-compute ledger
A compute ledger records every training run in a model's lineage with its estimated operations, sums them the way each instrument counts, and reports distance to each threshold. The standard estimate for dense transformers is about 6 operations per parameter per training token, 6ND, covering the forward and backward passes. It is an approximation: it ignores attention's sequence-length term, uses active parameters for mixture-of-experts models, and does not cover RL rollouts well. Where you have hardware logs, a measured estimate from accelerator-hours is better, and the ledger prefers it.
from dataclasses import dataclass, field
THRESHOLDS = {
# name: (operations, what counts) -- check the current legal text before relying on these
"EU AI Act systemic-risk presumption": (1e25, "cumulative training compute"),
"California SB 53 frontier model": (1e26, "initial training + fine-tuning + material modification"),
}
@dataclass
class Run:
name: str
kind: str # "pretrain", "finetune", "rl", "distill"
params_active: float # parameters touched per token (active, for MoE)
tokens: float
accel_hours: float = 0.0 # optional measured basis
peak_flops: float = 0.0 # per accelerator, at the precision used
mfu: float = 0.0 # measured model FLOP utilisation
note: str = ""
def ops(self):
if self.accel_hours and self.peak_flops and self.mfu:
return self.accel_hours * 3600 * self.peak_flops * self.mfu, "measured"
return 6 * self.params_active * self.tokens, "6ND estimate"
@dataclass
class Lineage:
model: str
upstream_ops: float = 0.0 # reported by the base-model provider, if any
upstream_source: str = "none"
runs: list = field(default_factory=list)
def report(self, warn_ratio=0.3):
own = sum(r.ops()[0] for r in self.runs)
total = own + self.upstream_ops
lines = [f"{self.model}: own {own:.2e}, upstream {self.upstream_ops:.2e} "
f"({self.upstream_source}), total {total:.2e}"]
for name, (limit, counts) in THRESHOLDS.items():
ratio = total / limit
status = "ABOVE" if ratio > 1 else ("NEAR" if ratio > warn_ratio else "below")
lines.append(f" {status:5} {name}: {ratio:.1%} of {limit:.0e} ({counts})")
bases = {r.ops()[1] for r in self.runs}
if "6ND estimate" in bases:
lines.append(" note: some runs use the 6ND approximation; keep hardware logs")
return "\n".join(lines)Two design choices come straight from the research. The ledger sums fine-tuning and later modification runs because SB 53 counts them, and it keeps upstream compute separate with its source, because whether a downstream modifier inherits obligations is a legal question each instrument answers differently. Put that question to counsel rather than the code.
Worked example: where two lineages land
Take a 405-billion-parameter dense model pretrained on 15.6 trillion tokens. The 6ND estimate is 6 x 4.05e11 x 1.56e13, about 3.8e25 operations, which matches the figure Meta reported for Llama 3.1 405B. That is above the EU presumption at 10^25 and about 38 percent of the SB 53 line at 10^26. A company that fine-tunes it on 2 billion tokens adds 6 x 4.05e11 x 2e9, about 4.9e21, a rounding error. The ledger reports NEAR for SB 53 and ABOVE for the EU presumption, with the upstream figure flagged as provider-reported.
Compare a 70-billion-parameter model on 15 trillion tokens: 6.3e24, below both. The lesson is that most deployers will never cross a training-compute threshold through their own runs; their obligations come through the upstream model and through deployment-side rules. A lab training at the frontier, by contrast, needs the ledger before the run starts, because the threshold is defined by the run's total.
Failure modes
- Proxy fixation. Treating a compute threshold as a risk measure. Research says it is a filter for attention, not a verdict.
- Proposals mistaken for law. A widely cited paper is not an obligation. Trace each requirement to an adopted instrument and its effective date, which a regulatory watch does.
- Stale institutional facts. Institutes are renamed and centres change hosts; GovAI is the example. Re-check names and affiliations before citing them.
- Counting only pretraining. Missing fine-tuning, RL and distillation runs understates a lineage's total.
- Unreproducible evaluation claims. A score without the elicitation method cannot support a go or no-go decision.
Trade-offs
Compute thresholds are cheap to verify and easy to game; capability evaluations track risk better but are expensive and contestable. Third-party audits add credibility but need model access that raises security and confidentiality questions. Voluntary frameworks move fast but are only as good as their enforcement. Most current instruments combine all four, which is why your evidence store should too.
What to do next
- List every model you train or modify and build a lineage ledger with upstream compute and its source.
- Switch training jobs to log accelerator-hours, precision and measured utilisation so the ledger can use measured figures.
- Record elicitation details (prompts, scaffolds, attempts) alongside every safety evaluation score.
- Write an incident taxonomy and reporting timer that can meet the strictest instrument you are subject to.
- Set up an internal anonymous reporting channel and document it.
- Pick three governance research sources, read them quarterly, and log which findings could change a control.
- Re-verify thresholds, effective dates and institute names against primary texts before each report.