A handful of companies train the largest general-purpose models: Anthropic, OpenAI, Google DeepMind, and a small number of others. People call them frontier labs. Each one publishes a safety framework that describes how it decides whether a new model is too dangerous to release, or too dangerous to keep training without extra protection. These documents are often summarised as "the lab does red teaming and evals", which is true but not useful. You can't act on it.
This article treats a frontier lab's safety program as an engineering system with inputs, measurements, decisions and outputs. It explains each part from first principles, separates what the labs commit to from what they merely describe, and then turns that understanding into something you can use: a vendor due-diligence record and a deployment gate for a team building on a lab's model. Framework versions and dates below were checked at the time of writing. These policies change often, so confirm the current text before you cite one in a contract or a risk review.
What the published frameworks share
Three published frameworks dominate the discussion. Anthropic's Responsible Scaling Policy (RSP) was rewritten as version 3.0 in February 2026 and has been revised since. OpenAI's Preparedness Framework version 2 took effect in April 2025; check whether a later revision exists. Google DeepMind's Frontier Safety Framework reached version 3 in September 2025, with later revisions after that. The names differ, but the skeleton is the same.
- Risk domains. A short list of harms severe enough to justify gating a release. All three cover biological and chemical weapons uplift, offensive cyber capability, and AI systems that speed up AI research itself. DeepMind's version 3 added a Critical Capability Level for harmful manipulation.
- Capability thresholds. A capability level, defined per domain, at which the risk changes character. OpenAI uses High and Critical. DeepMind uses Critical Capability Levels. Anthropic ties capability thresholds to AI Safety Level (ASL) standards of protection.
- Evaluations. Tests run on checkpoints during and after training to decide whether a threshold has been crossed or can't be ruled out.
- Required safeguards. Protections that must be in place before a model at a given level can be deployed, or in some cases even trained further.
- Governance and disclosure. Who makes the call, and what gets published afterwards.
The documents also differ in ways that matter. RSP 3.0 explicitly separates what Anthropic commits to on its own from what it recommends for the whole industry. It also removed the earlier language that implied a pause if mitigations could not keep up, and it added periodic Risk Reports and a Frontier Safety Roadmap. OpenAI's version 2 keeps deployment gates for three tracked categories and lists other areas, such as sandbagging and autonomous replication, as research categories outside those gates. When you compare labs, compare the commitments, not the vocabulary.
The safety loop as a pipeline
The diagram shows the loop. Read it from the top left.
Thresholds come first. A threshold written after you've seen the results is a rationalisation, not a measurement. The frameworks define thresholds in prose, such as meaningful uplift to a novice trying to build a known biological threat. They then pick evaluations as proxies for that prose. The gap between the prose and the proxy is where most of the argument happens.
Evaluations run on many checkpoints. Capability can appear partway through training, and post-training adds tool use and better instruction following, both of which raise measured capability. Labs evaluate intermediate checkpoints and the final candidate, often with scaffolding that gives the model tools, retries and long contexts.
The threshold call is asymmetric. The question is rarely "has the model crossed the line?" and more often "can we rule out that it has?" When a lab can't rule it out, it applies the higher safeguards as a precaution. That is what Anthropic did in May 2025 when it activated ASL-3 protections for Claude Opus 4.
Safeguards come in two kinds. Deployment safeguards stop misuse through the product: input and output classifiers for specific weapons-related content, account-level enforcement, rate limits, and staged or vetted access to the most capable features. Security safeguards stop the weights from being stolen. If the weights are stolen, every deployment safeguard can be stripped away.
Evaluations, elicitation and their limits
An evaluation can show that a model can do something. It can never prove that the model cannot, because a better prompt, a better tool or a little fine-tuning might unlock it. That is the elicitation problem, and every serious lab report spends pages on it. The usual methods to push towards the true ceiling are:
- Agentic scaffolds with code execution and web access, because many dangerous tasks are multi-step.
- Many samples per task, scored as pass@k, because a rare success still matters when the harm is severe.
- Helpful-only model variants with refusal training removed, so that refusals don't hide underlying capability.
- Expert red teamers working with the model over days rather than minutes.
- Uplift trials that compare people with model access against people with only internet access.
Labs also give pre-deployment access to external testers, including government AI security institutes in several countries. External testing is valuable because the testers have different incentives, but it is usually time-boxed and depends on access the lab chooses to grant.
Two failure directions matter. Under-elicitation makes a model look safer than it is. Sandbagging, where a model performs worse when it recognises it is being tested, would do the same thing deliberately. It is studied as a research risk and is hard to rule out with black-box tests. Multiple-choice knowledge benchmarks are cheap early-warning probes, but they measure recall, not the ability to carry out a task. The dangerous capability evaluations guide covers the measurement theory, including marginal uplift and safety margins.
The evidence a lab publishes
As a customer you never see the training run. You see documents. Know what each one is evidence for.
| Artifact | What it tells you | What it does not |
|---|---|---|
| Framework document | The rules the lab says it follows, and which are commitments | Whether a given model was handled correctly |
| System card | Evals run on a specific model, results, safeguards applied, threshold call | Results of tests the lab chose not to run or publish |
| Risk report (Anthropic, under RSP 3.0) | An overall risk assessment across deployed models | Independent confirmation, unless external review is attached |
| External tester statements | That a third party had access and what it looked at | Usually not full results |
| Usage policy | What you contractually may not do | How enforcement works on your traffic |
| Security attestations (SOC 2, ISO 27001) | Baseline corporate controls | Protection of model weights against capable attackers |
Read a system card the way you'd read an audit report. Find the threshold call and the reasoning behind it. Note which evaluations were saturated, meaning the model scored near the ceiling so the test no longer discriminates. Check whether a helpful-only variant was tested, and how much elicitation effort is described. If the card says a threshold "could not be ruled out", the model ships with stronger safeguards, and some of them will reach your traffic as refusals and blocked requests. Plan for that. Our model cards guide explains how to write the equivalent for your own fine-tunes.
Worked example: due diligence for a biotech assistant
Suppose a biotech company wants to put a frontier model behind an internal research assistant. Scientists will paste in protocols and ask for literature summaries, and the assistant can call a sequence-analysis tool. The domain sits right next to one of the labs' highest-concern risk areas. Three practical consequences follow. The provider's classifiers may block legitimate queries. Your own outputs could contribute to misuse. And the provider's safeguards depend on terms you agree to, such as no attempts to bypass them.
Record the provider's evidence as data, not as a paragraph in a slide deck. Then let a gate decide what deployment tier that evidence supports. Versioning the record means a model upgrade forces a fresh review.
from dataclasses import dataclass, field
@dataclass
class ProviderEvidence:
provider: str
model_id: str # exact version string, never an alias like "latest"
framework_version: str # e.g. "RSP 3.x", read from the provider's own page
system_card_url: str
threshold_calls: dict # domain -> "below" | "not_ruled_out" | "crossed"
safeguards: set # e.g. {"cbrn_classifier", "weights_security_attested"}
external_testing: bool
reviewed_by: str
open_questions: list = field(default_factory=list)
TIER_REQUIREMENTS = {
# tier -> evidence this deployment needs before go-live
"internal_low_risk": {"needs": set(), "max_open_questions": 3},
"internal_sensitive_domain": {"needs": {"cbrn_classifier"}, "max_open_questions": 0,
"domains": {"bio", "chem", "cyber"}},
"external_users": {"needs": {"cbrn_classifier", "weights_security_attested"}, "max_open_questions": 0,
"domains": {"bio", "chem", "cyber"}},
}
def gate(ev: ProviderEvidence, tier: str) -> list:
"""Return blocking problems; an empty list means the tier is supported."""
req, problems = TIER_REQUIREMENTS[tier], []
if ev.model_id.endswith("latest"):
problems.append("pin an exact model version")
missing = req["needs"] - ev.safeguards
if missing:
problems.append(f"missing safeguards: {sorted(missing)}")
valid = {"below", "not_ruled_out", "crossed"}
recorded = {d for d, call in ev.threshold_calls.items() if call in valid}
unknown = sorted(req.get("domains", set()) - recorded)
if unknown:
problems.append(f"no threshold call recorded for: {unknown}")
if len(ev.open_questions) > req["max_open_questions"]:
problems.append(f"{len(ev.open_questions)} open questions to the provider")
if tier != "internal_low_risk" and not ev.external_testing:
problems.append("no external pre-deployment testing disclosed")
return problemsIn this scenario the system card reports that the bio threshold was "not ruled out" and that weapons-related classifiers are active. That is good news for the company's risk posture, because the strongest deployment safeguards cover its traffic. It is also an operational cost: run a pilot to measure how often legitimate protocol questions get blocked. Then decide whether the provider offers a vetted access path for the researchers who need it, and record the result. Keep your own controls as well. Log prompts and outputs, restrict the sequence tool to approved databases, and put a human review step on anything that leaves the building. A provider's safeguards protect the provider's risk boundary. They are not a substitute for yours. The AI safety frameworks crosswalk shows how to map this record onto NIST AI RMF and ISO/IEC 42001 controls.
Failure modes
- Alias drift. You reviewed one model and your code calls an alias that now points to another. Pin versions, and make an upgrade a change request that reruns the gate.
- Treating policy as guarantee. Frameworks contain judgement calls, escape clauses and non-binding goals. RSP 3.0's roadmap goals are explicitly not hard commitments. Quote the binding language in your risk register and label the rest as aspirations.
- Ignoring the refusal tax. Stronger safeguards mean more false positives in nearby legitimate domains. Teams discover this in production and respond by trying to route around the classifiers, which can break the usage policy. Measure it in a pilot.
- Assuming evals cover your use. Lab evaluations target catastrophic misuse, not your application's failure modes. Prompt injection through your retrieval corpus, for example, is your problem and needs your own red team process.
- Snapshot reviews. Frameworks and system cards get revised. A review dated last year describes last year's model under last year's rules. Re-review on a schedule and on every model change.
- Competitive erosion. Commitments loosen under competition, a dynamic analysed in the AI safety race. Watch version diffs of the framework documents, not just the announcements.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Most capable frontier model | Best quality, strongest published safeguards | More refusals near sensitive domains; higher price |
| Smaller or older model | Cheaper, fewer classifier blocks | Weaker capability; older safety evidence |
| Open-weight model self-hosted | Full control, no provider classifiers | You own every safeguard; refusal training can be removed by anyone |
| Vetted or enterprise access tier | Fewer false blocks for qualified users | Contractual obligations, identity verification, slower onboarding |
| Contractual safety terms | Notification of incidents and model changes | Negotiation effort; most buyers get standard terms |
There is no option that is safe in general. The right choice depends on the domain and who your users are, and on how much of the safeguard stack you can operate yourself.
What to do next
- Download the current framework document from each provider you use, and record its version and date in your risk register.
- For every model in production, link its exact version string to its system card, and record the threshold call for each risk domain.
- Encode the evidence as a record like the one above and run the gate in CI whenever a model version changes.
- Pilot each model on a sample of real queries from your sensitive domains and measure the false-block rate before launch.
- Ask providers in writing about incident notification, deprecation timelines and how vetted access works.
- Keep your own logging, tool restrictions and red teaming, because provider safeguards stop at the provider's boundary.
- Diff the framework documents on each revision and flag weakened commitments to your risk owner.