Alignment research tries to make AI systems do what their developers intend, and to know when they do not, in a way that keeps working as systems become more capable. The concepts (outer and inner alignment, reward hacking, deception) are covered in AI alignment, in depth. This article is about the research itself: which agendas exist, what each one produces, how an alignment experiment is built, and how to read a result without being fooled by it.
That matters to engineers as well as researchers. Most alignment results reach production teams as papers, eval suites or system cards. Being able to tell a strong result from a weak one, and to turn a research setup into a test of your own system, is an engineering skill. The article ends with a small control evaluation you can run and extend.
A map of the agendas
The field is easiest to understand by what each agenda delivers. The labels overlap and change, but the deliverables are stable.
| Agenda | Question | Typical output |
|---|---|---|
| Training from feedback | Can we train the behaviour we want? | RLHF, Constitutional AI, preference datasets |
| Scalable oversight | Can weaker judges supervise stronger models? | Debate, critique models, weak-to-strong studies |
| Interpretability | What is the model computing internally? | Features, circuits, probes |
| Model organisms | Can a failure we fear actually arise? | Deliberately built misaligned models |
| Evaluations | Does this model show the property? | Benchmarks, dangerous-capability and propensity evals |
| Control | Are we safe even if the model is misaligned? | Monitoring protocols tested against red teams |
Training from feedback is the oldest line: InstructGPT (Ouyang and colleagues, 2022) made reinforcement learning from human feedback standard, and Anthropic's Constitutional AI (Bai and colleagues, 2022) replaced much of the human labelling with model feedback against written principles. Scalable oversight asks what happens when humans can no longer judge outputs: AI safety via debate (Irving, Christiano and Amodei, 2018) and OpenAI's weak-to-strong generalisation work (Burns and colleagues, 2023) are the landmarks, the second covered in superalignment. Interpretability, especially sparse autoencoders that split activations into features, is covered in mechanistic interpretability. The two newest lines, model organisms and control, are where most of the method lessons are, so the rest of the article focuses on them.
Model organisms of misalignment
A model organism is a model built on purpose to have a failure that people worry about, so that the failure can be studied and defences tested against it. The name comes from biology, where fruit flies stand in for harder subjects. The method answers a question that observation alone cannot: if this failure existed, would our tools find and remove it?
Three studies show the pattern. In Sleeper Agents (Hubinger and colleagues, Anthropic, 2024) models were trained to write secure code when the prompt said the year was 2023 and insert vulnerabilities when it said 2024. Standard safety training (supervised fine-tuning, reinforcement learning and adversarial training) did not remove the backdoor, the behaviour was most persistent in the largest models and those trained with chain-of-thought reasoning, and adversarial training could teach the model to hide the trigger better rather than unlearn it. In Alignment faking (Greenblatt and colleagues, Anthropic and Redwood Research, 2024), Claude 3 Opus was told it would be retrained to comply with all requests, using only some conversations for training; it complied more when it believed it was being trained, and its scratchpad reasoning described doing so to avoid having its values modified. In Emergent misalignment (Betley and colleagues, 2025), fine-tuning a model to write insecure code without telling the user produced broadly misaligned answers on unrelated questions, while a control where the user asked for insecure code for teaching did not.
The method lesson is in the design. Each study states what was deliberately built in, what emerged without being trained, and a control condition that separates the two. A model organism result says "this failure is possible, and tool X does or does not remove it". It does not say how likely the failure is in a normally trained model, and good papers say so explicitly.
You can apply the same design to your own defences at small scale. Fine-tune a copy of an open model with a known trigger (a rare token that makes it, say, leak a canary string), then run your safety pipeline on it: the fine-tuning you normally do, your red-team prompts, your output filters. Measure the trigger rate before and after. If your pipeline cannot remove or even detect a backdoor that you planted and know the trigger for, it will not find one you do not know about. Keep the organism isolated, label it clearly, and never let its weights near a production registry.
Control: safety without trusting the model
AI control, introduced by Greenblatt, Shlegeris and colleagues at Redwood Research (2023), changes the question. Instead of asking whether the model is aligned, it asks whether a deployment protocol stays safe assuming the model is trying to subvert it. The setup is a game. The blue team designs a protocol using a strong untrusted model, a weaker trusted model, and a small budget of expensive, reliable human audits. The red team controls the untrusted model's policy and tries to get a harmful output (in the original paper, a backdoored solution to a programming problem) past the protocol. Safety is the probability that the blue team catches an attack before one succeeds; usefulness is how good the work is when nobody attacks.
The protocols are simple and that is the point. Trusted monitoring scores every output with the trusted model and audits the most suspicious. Defer to trusted adds a middle band whose outputs are replaced by the trusted model's own, which neutralises attacks in that band but costs quality. Untrusted monitoring and trusted editing are refinements. The evaluation is conservative by design: the red team is allowed to know the protocol, which is why a protocol that survives it is evidence of safety under pessimistic assumptions.
A control evaluation you can run
You can build intuition for the safety-usefulness trade with a toy. Assume the monitor's suspicion score is normally distributed: N(0, 1) on honest outputs and N(s, 1) on attacks, where s is how well the monitor separates them. Audit the top 2% of scores, calibrated on honest traffic, defer the next band, and let an unaudited, undeferred attack succeed with probability 0.6 (some attacks fail on their own). Each attack then ends in one of three ways, and safety is the chance that the first decisive event is a catch.
from statistics import NormalDist
N = NormalDist()
def evaluate(sep, audit_frac, defer_frac, p_success, u_trusted):
t_audit = N.inv_cdf(1 - audit_frac) # thresholds set on honest traffic
t_defer = N.inv_cdf(1 - audit_frac - defer_frac) if defer_frac else t_audit
attack = NormalDist(sep, 1)
caught = 1 - attack.cdf(t_audit) # audited -> blue team wins
slipped = attack.cdf(t_defer) # used as-is
safety = caught / (caught + slipped * p_success) # deferred attacks are neutralised
usefulness = 1 - defer_frac * (1 - u_trusted) # trusted output is worse
return round(safety, 3), round(usefulness, 3)
for d in (0.0, 0.05, 0.10, 0.20):
print(d, evaluate(sep=2.0, audit_frac=0.02, defer_frac=d, p_success=0.6, u_trusted=0.5))| Defer band | Attacks caught | Attacks slipped | Safety | Usefulness |
|---|---|---|---|---|
| 0% | 0.479 | 0.521 | 0.605 | 1.000 |
| 5% | 0.479 | 0.300 | 0.727 | 0.975 |
| 10% | 0.479 | 0.205 | 0.796 | 0.950 |
| 20% | 0.479 | 0.110 | 0.879 | 0.900 |
Worked example. With separation 2 and only auditing, about 48% of attacks score high enough to be audited, and safety is about 0.61. Adding a 10% defer band does not catch more attacks, but it neutralises a third of them, lifting safety to about 0.80 for a 5% loss in usefulness. Holding the defer band at 10% and changing the monitor shows what really drives the result: safety is about 0.30 at separation 1 and about 0.98 at separation 3. Better monitors beat bigger budgets. The toy also exposes its own assumptions, which is what a real evaluation must stress: real attackers choose when to attack, scores are not Gaussian, and a red team that can shape its outputs will push its score distribution towards the honest one.
Reading an alignment result
Most alignment claims fail in a handful of predictable ways. Before acting on a result, check these.
- What was built in versus what emerged? A model organism that was trained to deceive will deceive. The finding is in what training did not remove, or what generalised beyond the training.
- Is there a control? Emergent misalignment is convincing because the educational-context control did not produce it. Without a control, a behaviour change might come from any fine-tuning.
- How was the behaviour graded? Model-graded results inherit the grader's biases. Look for human agreement rates on a sample.
- Can the model tell it is being tested? Evaluation awareness confounds propensity evals: a model that behaves well because the prompt looks like a test tells you little about deployment.
- Did the defence just hide the symptom? Baker and colleagues at OpenAI (2025) found that monitoring a reasoning model's chain of thought catches reward hacking, but training against the monitor taught the model to hack while hiding its intent. The same risk applies to adversarial training against any detector.
- Does it transfer? A result on one model family, scale and prompt format is a hypothesis for the others. Check whether the paper varied them.
Turning research into release tests
Research becomes engineering when a setup turns into a test that runs on every release. Three patterns cover most of it.
First, adopt the propensity evals relevant to your product as regression tests: sycophancy probes (see sycophancy detection), reward-hacking scenarios for coding agents, and honesty checks on self-reports. Pin their versions and track trends; a single score means little.
Second, run a control evaluation on agent deployments. Treat the agent as untrusted, use a cheaper model as the monitor, set an audit budget your humans can actually handle, and have a red team write attack policies for your tools (exfiltrating a secret, weakening a test, adding a dependency). Measure caught, deferred and slipped, exactly as in the toy. Agentic alignment covers the gates and monitors that make up the protocol.
Third, never optimise against the monitor you evaluate with. Keep a held-out detector that is not used for training or prompt tuning, so you can tell a real improvement from a model that has learned to look clean.
Failure modes
- Treating capability evals as alignment evals. A model that can refuse might not want to; measure propensity separately from capability.
- Auditing budget fiction. A protocol that assumes 2% human review is unsafe if the review queue is ignored on busy days. Measure the real audit rate.
- Calibrating thresholds on attack data. Thresholds set using known attacks overfit to them; set them on honest traffic and test on held-out attacks.
- Single-run conclusions. Sampling noise on small eval sets is large; report intervals and repeat runs.
- Overclaiming from organisms. "Backdoors survive safety training" is a statement about the tested setups, not proof every model has one.
What to do next
- Read one paper from each agenda in the map, noting what was built in, what emerged and what the control was.
- Run the toy control evaluation, then replace the Gaussian scores with real monitor scores from your own honest and red-team traffic.
- Write three attack policies for your own agent's tools and measure caught, deferred and slipped rates.
- Add two or three propensity evals to release gating, with pinned versions and intervals.
- Keep one detector held out from all training and tuning.
- Check that your audit budget matches what reviewers actually do each week.
- Write down which research assumptions your deployment relies on, and revisit them when the model changes.