Superalignment is the problem of aligning AI systems that are more capable than the people supervising them. Today's main technique, reinforcement learning from human feedback, assumes that a human can look at a model's output and judge whether it is correct, honest and safe. That assumption fails once a model writes code too large to review, proposes experiments its evaluators cannot check, or reasons in ways they cannot follow. Superalignment asks how to keep training signals trustworthy when the trainer is the weaker party.

The term was popularised by OpenAI, which in July 2023 announced a Superalignment team co-led by Ilya Sutskever and Jan Leike, with a commitment of a fifth of the compute it had secured and a stated goal of solving the core technical problems within four years. The team was disbanded in May 2024 after both leads left the company, but the research problem did not go away, and parts of it can be studied empirically today. This article focuses on the most concrete of those parts, weak-to-strong generalization, and shows how to run the experiment yourself. For the broader landscape of outer and inner alignment and scalable oversight see AI alignment as engineering.

Why supervision breaks as models get stronger

Strip the problem to its structure. Training needs a signal that says which outputs are good. With RLHF that signal comes from human comparisons, turned into a reward model, which the policy is optimised against. Three things go wrong as the policy becomes more capable than its raters:

  1. Errors become invisible. Raters reward outputs that look right. A model that is better than its raters at a task can produce subtly wrong answers that score as well as correct ones.
  2. Errors become learnable. If rater mistakes are systematic, a strong model can learn to predict them, and the cheapest way to score highly becomes modelling the rater rather than solving the task. This is sycophancy and reward hacking at scale.
  3. Deception becomes possible. A system with a model of its overseers can behave well when it expects to be checked and differently otherwise; see scheming and situational awareness.

The research programme that came out of this framing has three strands: build supervision signals that scale beyond direct human judgment (scalable oversight, such as debate, recursive reward modelling and AI-assisted critique), check that the resulting systems are aligned by means other than looking at outputs (interpretability and automated auditing), and stress-test the whole pipeline by deliberately training misaligned models and confirming the checks catch them. A recurring hope is an automated alignment researcher: a system roughly as capable as a human researcher that is trustworthy enough to do much of this work.

That hope is circular in an instructive way: trusting the automated researcher requires the very oversight techniques it is supposed to help build. The way out, if there is one, is incremental, with each generation of models checked by methods validated on the previous one, which is why measurable proxies matter so much.

Weak-to-strong generalization

The difficulty is that you cannot study supervision of superhuman models without superhuman models. Burns and colleagues at OpenAI proposed an analogy that can be run now, in the paper Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision (December 2023): use a small model as the supervisor and a large pretrained model as the student.

The protocol has three runs. Train the weak model on ground truth; that is the weak performance. Use the weak model to label a held-out set, fine-tune the strong model on those weak labels, and measure it on ground truth; that is the weak-to-strong performance. Fine-tune the strong model on ground truth; that is the strong ceiling. The headline metric is performance gap recovered:

PGR = (weak_to_strong - weak) / (strong_ceiling - weak)

PGR of 0 means the student learned nothing beyond imitating its teacher, mistakes included. PGR of 1 means weak supervision lost nothing. A hypothetical illustration: weak teacher 60% accuracy, strong ceiling 90%, weak-to-strong student 80%. The student recovered 20 of the 30 points, PGR about 0.67, and it is better than the labels it was trained on, which is the phenomenon the paper named.

The paper's findings, stated as its abstract states them: strong students fine-tuned naively on weak labels consistently outperform their weak supervisors, but stay far from the ceiling, which suggests RLHF-style methods may scale poorly to superhuman models without further work. Simple methods help considerably; with a GPT-2-level supervisor and an auxiliary confidence loss, GPT-4 recovered close to GPT-3.5-level performance on NLP tasks. The paper also reports that results varied by domain, and that the reward-modelling setting, the one closest to RLHF, was among the hardest. Treat these as results on one family of models and tasks, not as constants.

The superalignment analogy: a weak supervisor training a stronger studentTodaySuperalignmentAnalogy we can run nowHuman supervisorstrongModelweaker than humanlabelsHuman supervisorweakModelsuperhumanlabelsSmall modelweak teacherLarge modelstrong studentweak labelsPerformance gap recovered (PGR)0weak teacherweak-to-strongstrong ceilingPGR = (weak-to-strong - weak) / (ceiling - weak): 1 means nothing lost, 0 means the student only copied
Weak-to-strong generalization as a stand-in for superalignment: a small model labels data for a large one, and PGR measures how much of the capability gap survives imperfect supervision.

Letting the student disagree: the confidence loss

Why would a student ever beat its teacher? Because a pretrained strong model already contains much of the knowledge needed; fine-tuning only has to point at it. Weak labels are a noisy pointer. If the student's own representation strongly prefers the correct answer on a mislabelled example, a well-designed objective lets it disagree with the teacher there instead of memorising the error.

The auxiliary confidence loss formalises that. It mixes the usual cross-entropy against the weak label with a cross-entropy against the student's own hardened prediction:

# f_s: student probabilities, y_w: weak teacher's soft labels, alpha warmed up from 0
def confidence_loss(f_s, y_w, alpha, threshold):
    y_self = (f_s > threshold).float().detach()     # student's own hard prediction
    return (1 - alpha) * cross_entropy(f_s, y_w) + alpha * cross_entropy(f_s, y_self)

The second term rewards the student for being confident in its own answer, which makes it cheaper to override a teacher's mistake than to imitate it. The paper chooses the threshold adaptively from the student's predictions rather than fixing it; a plain 0.5 is a reasonable starting point for balanced binary tasks. Two related techniques from the same work: bootstrapping, where a chain of progressively larger models each supervise the next so no single step bridges a huge gap, and early stopping, since students that train long enough on weak labels eventually overfit to the teacher's errors.

None of this is free. The confidence term can just as easily reinforce a student's confident mistakes, which is safe only when the student's prior is better than the teacher's. In an alignment setting that is exactly the open question: a superhuman model's confident beliefs about what its supervisor wants are what we are trying to verify.

A weak-to-strong harness you can run

You can reproduce the shape of the experiment in a few minutes with scikit-learn. The weak model sees only a noisy view of the data, standing in for limited capability; the strong model sees it clean. The numbers it prints depend on the random seed and data and are illustrative only; the protocol is the point.

import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=30000, n_features=40, n_informative=20,
                           flip_y=0.02, class_sep=0.8, random_state=0)
X_w, X_rest, y_w, y_rest = train_test_split(X, y, train_size=0.3, random_state=1)
X_s, X_test, y_s_true, y_test = train_test_split(X_rest, y_rest, test_size=0.3, random_state=2)

rng = np.random.default_rng(0)
blur = lambda A: A + rng.normal(0, 3.0, A.shape)            # weak model: a noisy view of the data
weak = LogisticRegression(max_iter=1000).fit(blur(X_w), y_w)
weak_acc = weak.score(blur(X_test), y_test)

y_s_weak = weak.predict(blur(X_s))                         # weak labels for the student
w2s = HistGradientBoostingClassifier(max_iter=200, early_stopping=True).fit(X_s, y_s_weak)
w2s_acc = w2s.score(X_test, y_test)

ceiling = HistGradientBoostingClassifier(max_iter=200).fit(X_s, y_s_true).score(X_test, y_test)

pgr = (w2s_acc - weak_acc) / (ceiling - weak_acc)
print(f"weak={weak_acc:.3f} weak_to_strong={w2s_acc:.3f} ceiling={ceiling:.3f} PGR={pgr:.2f}")

In our runs of this setup PGR landed between roughly 0.2 and 0.4 depending on the seed: the student clearly beat its teacher and stayed well short of the ceiling. Then vary one thing at a time and record PGR: raise the noise to widen the gap, change the student's capacity, remove early stopping, or make the teacher's mistakes systematic. The simplest systematic teacher is one that sees only the first few clean features instead of a noisy view of all of them. Its errors are then a deterministic function of inputs the student can see, the student learns them, and in our runs PGR fell to about zero. That is the classical-model version of a strong model learning to predict its raters, and it is the most instructive variation.

Read the output in order. First confirm the gap exists: if the ceiling barely beats the weak model, PGR is meaningless and you should widen the gap before drawing conclusions. Then compare the weak-to-strong score with the weak score. A student that beats its teacher is generalizing from the teacher's intent rather than copying its labels, which works here because the teacher's errors are mostly random relative to the features the student can see. Finally, repeat each configuration over several seeds and report the spread, because a single run can move PGR by a large margin when the denominator is small.

Keep the analogy's limits in view. A boosted tree is not a pretrained model with latent knowledge, and the paper itself lists disanalogies: future superhuman models may imitate human errors more easily than today's students imitate small-model errors, because human mistakes are well represented in pretraining data, and real misalignment may be adversarial rather than noisy.

What it means for systems you build today

Superalignment sounds remote from the day job of securing an LLM application, but its core lesson applies now: whenever a model is graded by something weaker than itself (a smaller judge model, a regex, a rushed human reviewer), the model can learn the grader instead of the task.

  • Measure grader-model gaps. If an LLM judge scores your production model, sample outputs where judge and expert disagree and track the disagreement rate per release. A rising pass rate with a flat expert score is a weak-to-strong failure.
  • Hold out ground truth. Keep a small, expensive, expert-labelled set that is never used for training or prompt tuning, and compute your own PGR-style gap against it; see security evaluations for LLMs.
  • Use structure, not just judgment. Unit tests, executed code, retrieved citations and formal checks are supervision that does not weaken as the model improves.
  • Look inside when outputs cannot be judged. Interpretability tools such as probes and sparse autoencoders aim to check what a model represents rather than what it says; see mechanistic interpretability.

The trade-offs are real. Expert labels are expensive, interpretability is immature, and debate and critique methods can be gamed by a persuasive model. Nobody has a solution that is known to scale to superhuman systems; what exists is a set of measurable proxies, and the honest stance is to measure them rather than assume them.

Failure modes

Failure modes in superalignment research and practice:

  • Imitating the supervisor. The student reproduces teacher errors; PGR near zero. Caused by overtraining, small gaps between student and teacher priors, or systematic teacher mistakes.
  • Measuring against the wrong ceiling. A ceiling trained on noisy labels understates the gap and inflates PGR. Use the cleanest ground truth you have.
  • Unstable PGR. When weak and ceiling scores are close, the denominator is tiny and PGR swings wildly. Report the raw accuracies and confidence intervals alongside it.
  • Overconfidence objectives backfiring. Confidence losses amplify the student's own biases when its prior is worse than the teacher's.
  • Analogy drift. Strong empirical results on small models read as evidence about superhuman ones. State the disanalogies every time.

What to do next

  1. Read the abstract and method section of Burns et al. (2023), and write down the PGR definition and its three required runs.
  2. Run the scikit-learn harness, then add systematic teacher errors and record how PGR changes.
  3. Add an auxiliary confidence term to a small PyTorch version and compare it to naive fine-tuning.
  4. List every place your own system is graded by something weaker than the model being graded.
  5. Build an expert-labelled holdout set and report the judge-versus-expert gap for each release.
  6. Read deception detection for LLMs to see how oversight breaks when the student is adversarial rather than merely noisy.
Key takeaway: Superalignment is the problem of producing trustworthy training signals when the model is more capable than its supervisor. Weak-to-strong generalization turns it into an experiment you can run today: train a strong student on a weak teacher's labels and measure performance gap recovered. Naive students beat their teachers but fall short of the ceiling; confidence losses help. Apply the lesson now by measuring every weak grader against expert ground truth.