"The AI safety race" names a specific worry: when several developers compete to ship the most capable model first, each one is pushed to spend less time on evaluation, mitigation and staged release than it would choose alone. Nobody has to act in bad faith for this to happen. It falls out of the payoff structure, much as underinvestment in a shared resource does.

This article treats the race as an engineering and security problem rather than a news story. It builds a small simulation you can run to see which conditions push safety effort down, maps the real coordination mechanisms (published frameworks, the Seoul commitments, the EU code of practice, California's SB 53) onto that model, and then focuses on what you control: a release process that keeps its safety properties when a competitor ships first and the launch date is suddenly next week. The same pressures apply, at smaller scale, to any team shipping fine-tuned models or agents against a rival product.

The race as a game

Start with a stylised version. Several labs are developing a powerful system. Each chooses how much effort to put into safety, and safety work takes time, so more of it lowers the chance of finishing first. The first to finish captures most of the value. If the winner skimped, there is some chance of a serious failure that harms everyone, competitors included.

Two features make this different from an ordinary product race. First, the risk is shared: a careful lab that loses still bears the cost of a careless winner's failure. Second, each lab's best choice depends on what it expects the others to do. If rivals are careful, you can afford care; if they cut corners, care mostly means losing. Armstrong, Bostrom and Shulman formalised this in "Racing to the precipice" and argued that more competitors, and stronger enmity between them, push equilibrium safety down. The simulation below is our own simplified illustration of that thesis, not a reproduction of their model.

A race model you can run

Each of n labs has a random capability head start drawn uniformly from 0 to spread, and picks a safety level s between 0 and 1. Its race score is head start plus (1 - s), so safety costs speed. The highest score wins. The world survives with probability equal to the winner's s. Surviving as winner is worth 1; surviving as a loser is worth loser_value, which measures how much a lab values a rival's success over disaster (low values mean enmity). We search for a symmetric equilibrium by iterating best responses on a grid.

import random

def payoff(my_s, other_s, n, loser_value, draws):
    total = 0.0
    for skills in draws:
        mine = skills[0] + (1 - my_s)
        best_rival = max(skills[1:n]) + (1 - other_s)
        if mine >= best_rival:
            total += my_s                    # win; world survives with my_s
        else:
            total += other_s * loser_value   # lose; survive with rival's s
    return total / len(draws)

def equilibrium(n, spread, loser_value=0.5, trials=3000, seed=1):
    rng = random.Random(seed)
    draws = [[rng.uniform(0, spread) for _ in range(n)] for _ in range(trials)]
    grid = [i / 20 for i in range(21)]
    s = 1.0
    for _ in range(50):                      # best-response iteration
        best = max(grid, key=lambda m: payoff(m, s, n, loser_value, draws))
        if best == s:
            break
        s = best
    return s

Running it gives these equilibrium safety levels (higher is safer):

Capability spread2 labs3 labs5 labs
0.1 (neck and neck)0.200.150.10
0.50.650.400.30
2.0 (clear leader likely)1.001.000.95

Holding three labs at spread 0.5 and varying loser_value gives 0.20 at 0.0, 0.40 at 0.5 and 1.00 at 0.9.

What the model says

The numbers come from a toy, so read them as directions, not forecasts. Three are robust across settings.

  • Closeness drives corner-cutting. When head starts are small relative to what safety costs, a little speed decides the race, so everyone buys speed. A comfortable lead lets the leader be careful. This is why a rival's surprise launch is the most dangerous moment for your own process.
  • More competitors lower the floor. Each extra entrant makes your own care less likely to matter to the outcome.
  • Enmity matters as much as structure. When labs would be roughly as content with a rival's safe success as with their own, the incentive to cut corners almost disappears. Trust, shared norms and credible information about each other's practices are levers, not soft extras.

Two caveats. Real safety work is not pure cost: reliability, refusal quality and robustness are product features, and some safety research speeds capability work. And the model assumes everyone can observe the rules; in reality, the hardest part is verifying what others actually do.

Coordination mechanisms and what they change

Coordination mechanisms work by changing the payoffs: putting a floor under s, making s observable, or making a low s costly. Mapping each real mechanism to which of these it does is more useful than cataloguing them.

MechanismWhat it doesLimits
Lab frameworks (Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, Google DeepMind's Frontier Safety Framework)self-imposed if-then commitments: capability thresholds that trigger safeguards or pausesself-defined, self-assessed, revisable by the same organisation
Frontier AI Safety Commitments (Seoul, May 2024; sixteen companies at launch)public promises to publish frameworks, define intolerable-risk thresholds and not develop or deploy if risks cannot be mitigatedvoluntary; thresholds set by each signatory
EU General-Purpose AI Code of Practice, Safety and Security chapter (final text 10 July 2025)a way for providers of general-purpose models with systemic risk to show compliance with AI Act obligationsapplies to that regime's scope; signing is optional, the law is not
California SB 53, the Transparency in Frontier Artificial Intelligence Act (signed 29 September 2025, effective 1 January 2026)frontier developers (models trained with over 10^26 operations) publish transparency reports; those with revenue above $500 million also publish a frontier AI framework; critical safety incidents are reported within 15 days, or 24 hours when there is imminent risk of death or serious injurytransparency, not a capability cap; one jurisdiction
Third-party testing by government AI safety and security institutesexternal pre-deployment evaluation of some frontier modelsaccess is negotiated, and results are often not fully public

In model terms, frameworks and the Seoul commitments try to set a floor on s and raise loser_value by building shared norms. SB 53 and the EU code mainly make s observable, through published frameworks and incident reporting, which matters because an unobservable commitment is cheap talk under race pressure. None of them verifies internal practice directly, which is why the details of evaluations and evidence matter. AI safety frameworks in depth compares the lab policies and regulatory frameworks clause by clause; this article stays with the dynamics. Treat the regulatory details here as a snapshot and read the current texts before relying on them.

How race pressure erodes a release process

Race pressure rarely shows up as a decision to be unsafe. It shows up as small, locally reasonable changes, each easy to defend in a launch meeting.

  • Evaluating a proxy artifact. Evals run on last week's checkpoint; the shipped weights differ by a fine-tune that "shouldn't matter".
  • Moving thresholds after seeing results. A score lands just over the line and the line is reinterpreted, rather than the model being fixed.
  • Mitigations promised for after launch. The classifier or rate limit that justified shipping arrives two sprints later, if at all.
  • Compressed red-teaming. The window shrinks from weeks to days, so only known attack classes get tested.
  • Skipped staging. A planned gradual rollout becomes a general release to match a competitor's announcement.

The common thread is that the people who own the date also control the evidence or the threshold. The fix is structural: separate those roles, and make every exception visible.

A release gate that resists erosion

A release gate resists erosion when it has four properties. Thresholds are fixed before results exist. Evidence is bound to the exact artifact being shipped. The decision is a pure function anyone can re-run. And overrides are possible but expensive and recorded. The sketch below shows the core. The eval names and numbers are placeholders, not recommended values; derive yours from your own risk assessment and from dangerous capability evaluations.

from dataclasses import dataclass

# Changed only by a reviewed, signed commit to the policy repo, never by the launch team.
THRESHOLDS = {                      # eval name -> (limit, overridable)
    "cyber_ctf_solve_rate":   (0.35, False),
    "jailbreak_success_rate": (0.05, True),
    "benign_refusal_rate":    (0.08, True),
}

@dataclass(frozen=True)
class Result:
    name: str
    value: float
    model_sha256: str
    suite_version: str

def decide(model_sha256, results, approvers=(), launch_team=frozenset()):
    by_name = {r.name: r for r in results}
    missing = sorted(set(THRESHOLDS) - set(by_name))
    if missing:
        return "BLOCK", f"missing evals: {missing}"
    stale = sorted(r.name for r in results if r.model_sha256 != model_sha256)
    if stale:
        return "BLOCK", f"evals ran on a different artifact: {stale}"
    breaches = [(n, by_name[n].value, lim, ok)
                for n, (lim, ok) in THRESHOLDS.items() if by_name[n].value > lim]
    if not breaches:
        return "SHIP", "all thresholds met"
    if any(not ok for _, _, _, ok in breaches):
        return "BLOCK", f"non-overridable breach: {breaches}"
    independent = set(approvers) - set(launch_team)
    if len(independent) >= 2:
        return "OVERRIDE", f"breaches {breaches} accepted by {sorted(independent)}"
    return "BLOCK", f"breaches need two approvers outside the launch team: {breaches}"

Wire it into the deployment pipeline so that promoting weights to production requires a SHIP or OVERRIDE record for that hash, and write every decision, with its inputs, to an append-only log. The log turns silent erosion into a reviewable pattern: three overrides in a quarter on the same eval is a finding for your governance review, not a footnote.

A release gate that a launch date cannot quietly movePolicy repothresholds, signedCandidate modelartifact hashEval runnerspinned suiteshashGatepure functionresults + hashthresholdsSHIPstaged rolloutBLOCKwith reasonsOVERRIDEtwo approversAppend-only decision logaudited laterThe launch team owns the date; it does not own the thresholds, the eval code or the override.
Thresholds, evidence and override authority sit outside the team that owns the launch date.

Worked example: a competitor ships first

A team ships a fine-tuned support model. A competitor announces a similar product on Monday, and leadership asks to ship Friday instead of in three weeks. The candidate's hash is 9f3c.... On Wednesday the gate returns BLOCK because the benign-refusal eval ran on the previous checkpoint. Re-run on the right hash, it blocks again: jailbreak_success_rate is 0.07 against a limit of 0.05.

Without a gate, this is where the threshold gets argued down. With one, the options are explicit. Fix the jailbreak gap with an input classifier and re-evaluate the combined system. Or request an override, which needs two named approvers outside the launch team and is logged with the breach. The team ships Friday with the classifier in place, to 5% of traffic, with a pre-agreed rollback trigger and a kill switch tested that morning. It loses nothing it could not have lost anyway, and it keeps a record that would stand up to an incident review.

Failure modes

  • Gate as theatre. Thresholds are set loosely enough never to bind. Check how often the gate has blocked anything; never is a warning sign.
  • Override inflation. Overrides become routine because they are cheap. Track their rate and require a written mitigation plan with a date.
  • Eval suite drift. Suites change between candidates, so scores are not comparable. Version suites and record suite_version with each result.
  • Incident blindness. Post-launch harms are not fed back into thresholds. Connect incident response to the policy repo.
  • Public commitments that internal process cannot meet. Under SB 53-style transparency, a published framework the organisation does not follow becomes a legal and reputational exposure, not just an internal gap.

Trade-offs

Gates cost time, and pretending otherwise invites people to route around them. Keep the eval suite fast enough to run on every candidate, and make most thresholds overridable with accountability, reserving hard blocks for harms you would never accept. Expect some honest disagreement too. Openness, through published weights and evals, can raise shared understanding while lowering the barrier to misuse. Regulation can make practices observable while favouring incumbents who can afford compliance. The race framing itself can become an argument for speed ("if we don't, someone worse will"). That argument only holds if you actually behave better when you are ahead, which your gate log can show or refute.

What to do next

  1. Write down which evals block a release and their thresholds, before the next candidate is trained.
  2. Move thresholds into a repository the launch team cannot change alone; require signed, reviewed commits.
  3. Bind every eval result to the artifact hash and suite version; reject stale evidence automatically.
  4. Implement the gate as a pure function in the deploy pipeline, with an append-only decision log.
  5. Define the override rule (who, how many, outside which team) and review override rates quarterly.
  6. Pre-plan the fast path for competitive surprises: staged rollout percentages, rollback triggers, a tested kill switch.
  7. If you are in scope of SB 53 or the EU code, check that your published framework matches what the gate actually enforces; see LLM safety evals architecture for building the suites.
Key takeaway: Competition lowers safety effort most when rivals are close, numerous and distrustful, and coordination mechanisms help mainly by putting a floor under safety or making it observable. Inside your own organisation, the defence is structural: thresholds fixed before results, evidence bound to the shipped artifact, a gate that is a pure function, and overrides that need independent approvers and leave a permanent record.