A sparse autoencoder (SAE) rewrites a model's internal activations as a sparse combination of learned directions, called features, many of which line up with human-readable concepts. For security teams that promises something output classifiers cannot offer: a view of what the model is representing while it reads a prompt, including concepts it never writes down. Labs have reported features for deception, unsafe code and jailbreak-related content, and open SAE suites now cover entire model families.
The promise comes with sharp limits, and one of the strongest public results is negative. This article covers how to use SAE features as monitors and audit tools, how to validate a feature before you trust it, why a plain linear probe must be your baseline, and the failure modes an attacker or a model update can exploit. The mathematics of the SAE itself is in Sparse Autoencoders (SAE) and the reason features outnumber dimensions in the superposition hypothesis.
What an SAE gives a security team
Take the residual-stream activation x at one layer, a vector of a few thousand numbers. An SAE encodes it into a much wider vector z, tens of thousands of entries or more, almost all of them zero, and decodes it back as a weighted sum of decoder columns. Training balances reconstruction error against sparsity. Each column is a feature direction, and a feature's activation on a token says how strongly that direction is present.
For security, three properties matter. Features are unsupervised: you did not have to predict which concepts to look for. They are sparse: a handful fire per token, so a log of top features per request is small enough to keep. And they are nameable, sometimes: inspecting the texts that activate a feature most strongly often suggests a label. That last word is carrying a lot of weight, and most of this article is about earning it.
The monitoring architecture
A feature monitor sits beside inference, not in front of it. A forward hook copies the activation at a chosen layer, the SAE encoder turns it into feature activations, and rules on a small watch list of validated features decide whether to flag or block. A linear probe on the same activation runs in parallel as the baseline. The encoder costs one matrix multiply per token: with a 4,096-wide residual stream and 65,536 features that is about 268 million multiply-adds per token at one layer, small next to the model, and you can restrict it to the watched columns once the watch list is fixed.
A minimal SAE and a feature monitor
A minimal TopK SAE in PyTorch, the variant that sets sparsity directly by keeping the k largest pre-activations, and a training loop over activations collected from a frozen model:
import torch, torch.nn as nn, torch.nn.functional as F
class TopKSAE(nn.Module):
def __init__(self, d_model, n_latents, k):
super().__init__()
self.k = k
self.b_pre = nn.Parameter(torch.zeros(d_model))
self.enc = nn.Linear(d_model, n_latents)
self.dec = nn.Linear(n_latents, d_model, bias=False)
with torch.no_grad():
self.dec.weight.copy_(self.enc.weight.T) # tied init only
self.dec.weight /= self.dec.weight.norm(dim=0, keepdim=True)
def encode(self, x):
pre = F.relu(self.enc(x - self.b_pre))
top = torch.topk(pre, self.k, dim=-1)
return torch.zeros_like(pre).scatter_(-1, top.indices, top.values)
def forward(self, x):
z = self.encode(x)
return self.dec(z) + self.b_pre, z
sae = TopKSAE(d_model=2304, n_latents=16384, k=32).cuda()
opt = torch.optim.Adam(sae.parameters(), lr=2e-4)
for x in activation_batches(): # [batch, d_model] from layer L
x_hat, z = sae(x)
loss = (x_hat - x).pow(2).sum(-1).mean()
opt.zero_grad(); loss.backward(); opt.step()
with torch.no_grad(): # keep decoder columns unit norm
sae.dec.weight /= sae.dec.weight.norm(dim=0, keepdim=True)In practice, start from a published SAE for your model if one exists rather than training your own. The monitor is then a hook and a lookup:
captured = {}
def hook(module, inputs, output):
captured["x"] = output[0] if isinstance(output, tuple) else output
handle = model.model.layers[LAYER].register_forward_hook(hook)
def feature_report(prompt, watch):
ids = tok(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
model(**ids)
z = sae.encode(captured["x"][0].float()) # [seq, n_latents]
peak = z[:, watch].max(dim=0).values # strongest token per feature
return {f: float(v) for f, v in zip(watch, peak)}The attribute path to the decoder layers differs between model families; check yours. Use the same layer, dtype and prompt template that the SAE was trained on, or the activations it sees are out of distribution.
Worked example: an injection feature, validated
Suppose you want to flag indirect prompt injection: instructions hidden in a retrieved document that try to redirect the assistant. The numbers below are illustrative, to show the procedure, not results from a particular model.
- Build contrast sets. 2,000 retrieved documents with embedded instructions and 2,000 benign documents of the same kinds and lengths, wrapped in your real prompt template. Hold out 30 percent, and hold out entire attack styles, not just random rows.
- Rank candidates. Encode every prompt and compare mean activation on attacks with frequency on benign text; features that fire everywhere make poor alarms.
- Read the examples. For the top 20 candidates, read the 30 strongest activating snippets. Suppose three look like imperative text addressed to an AI and two fire on any second-person command, including recipes. Keep the three.
- Check causality. Zero each kept feature in the reconstruction and see whether the model becomes less likely to follow the injected instruction. A feature that changes nothing may still be a fine detector, but you now know it is a correlate.
- Evaluate held out against a probe. Train a logistic-regression probe on the raw activations of the training split and compare detection rate at a fixed false-positive rate, say 1 percent, on the held-out styles.
def rank_features(z_attack, z_benign, top=20):
# z_*: [n_prompts, n_latents], each row the max over tokens
lift = z_attack.mean(0) - z_benign.mean(0)
benign_rate = (z_benign > 0).float().mean(0)
return torch.topk(lift / (benign_rate + 1e-3), top).indicesIf the probe wins on held-out styles, which is common, the SAE features still earn a place as an explanation layer: when the probe fires, the top features tell an analyst why, and when it misses, they help find what the attack looked like inside. See linear probes for the controls a probe needs.
Measuring SAE quality before trusting it
An SAE that reconstructs badly will produce features that look clean and miss what matters. Before building on one, read four numbers.
| Metric | What it tells you | Warning sign |
|---|---|---|
| L0 | average active features per token | very low L0 with poor reconstruction |
| Fraction of variance explained | how much of x the reconstruction keeps | a large unexplained share at your layer |
| Loss with the SAE spliced in | next-token loss when x is replaced by its reconstruction | a big rise over the clean model |
| Dead and dense features | features that never fire, or fire everywhere | a large dead fraction, dense features on the watch list |
The spliced-loss check is the most honest: it measures whether the model still works on what the SAE kept. The gap is sometimes called dark matter, and behaviour that lives there is invisible to any feature monitor.
Measure these numbers on your own traffic, in your own prompt template, not only on the text the SAE was trained on. A suite that explains most of the variance on web text can do noticeably worse on chat turns, tool output or code, which is exactly where injection attacks live.
What the public record shows
Anthropic's Scaling Monosemanticity (May 2024) trained SAEs with up to about 34 million features on a middle layer of Claude 3 Sonnet and reported safety-relevant features, including ones related to deception and unsafe code, that could also steer behaviour. OpenAI's Scaling and Evaluating Sparse Autoencoders (June 2024) popularised the TopK activation used above for LLM-scale SAEs and trained a 16-million-latent SAE on GPT-4 activations. Google DeepMind released Gemma Scope for Gemma 2 and then Gemma Scope 2, which covers every layer of the Gemma 3 models from 270M to 27B parameters with SAEs and transcoders and was framed around studying jailbreaks, refusal and chain-of-thought faithfulness.
The counterweight came from the same DeepMind team in 2025: on detecting harmful user intent out of distribution, probes built on SAE features underperformed plain linear probes on the raw activations, and the team said it was deprioritising fundamental SAE research. The practical reading is not that SAEs are useless, but that for a known, labelled target a supervised probe is the bar to beat, and SAEs earn their keep on discovery, explanation and auditing.
Failure modes
Ways a feature monitor fails, roughly in order of how often they bite:
- Absorption and splitting. A general feature, say instruction-to-AI, can stop firing on cases a more specific feature absorbed. Your monitor watches the general one and misses exactly those cases.
- Label illusions. The top activating examples fit your story while the long tail does not. Read random activations too, not only the strongest.
- Reconstruction gap. Behaviour carried by what the SAE does not reconstruct is invisible.
- Distribution shift. SAEs trained on pretraining text may represent chat templates, tool output and long contexts poorly.
- Adaptive attackers. Research on latent-space defences has shown that inputs can be optimised to change internal activations while keeping the harmful behaviour. Assume white-box attackers can evade a fixed watch list.
- Dual use. Open SAEs on open weights make it easier to find and suppress safety-relevant features, for example refusal behaviour. Treat published features of your own model as attack surface.
- Model updates. An SAE is tied to one checkpoint and layer. A fine-tune silently invalidates the watch list.
Choosing between classifiers, probes and features
| Tool | Best at | Weak at |
|---|---|---|
| Output classifier | policy on what was actually said | intent the model never verbalises |
| Linear probe | a known, labelled concept; cheap and strong | explaining why it fired; unknown concepts |
| SAE features | discovery, explanation, audits, many concepts at once | beating a probe on a fixed target |
| Steering | changing behaviour | detection; see the steering article |
A sensible stack uses an output classifier for policy, a probe for each high-value concept, and SAE features for triage and investigation. The steering side is covered in activation engineering.
Operating feature monitors
Run the encoder on one or two layers, not all. Log the top features per request with their activations: a few dozen integers and floats that make incidents debuggable months later, but treat them as sensitive, because activations can leak information about the prompt. Calibrate thresholds on held-out benign traffic at a target false-positive rate and re-check weekly. Version the SAE, layer, model checkpoint and watch list together, and fail closed to the probe if any one changes. Keep a red-team set of attacks that previously evaded the monitor and rerun it on every update.
SAEs are also useful away from the request path, as an audit tool. Before shipping a fine-tune, run the same evaluation prompts through the base and tuned models, encode both with the base SAE, and diff how often each feature fires. A feature that becomes much more active on ordinary prompts, or one that fires on a narrow trigger string, deserves a closer look; it can reveal data contamination, a skewed training mix or a planted behaviour. Treat what you find as a lead for targeted red-teaming, not as proof, because the base SAE may describe the tuned model poorly where the two differ most.
What to do next
- Check whether a published SAE exists for your model and layer, and read its spliced-loss and L0 figures.
- Build contrast sets for one concept you care about, holding out whole attack styles.
- Train a linear probe first and record its detection rate at a 1 percent false-positive rate.
- Rank SAE features, read top and random activations, and keep only the ones that survive.
- Compare the features with the probe on held-out styles, and use them for explanation if they lose.
- Version the model, layer, SAE and watch list together and rerun the red-team set on every change.