Representation engineering, usually shortened to RepE, is a way of understanding and steering a language model by working with the patterns of activity inside it rather than with individual neurons or circuits. The term comes from Zou and colleagues' 2023 paper Representation Engineering: A Top-Down Approach to AI Transparency, which argued that high-level properties such as honesty, harmfulness or power-seeking show up as directions in a model's hidden states. Those directions can be found with a handful of contrasting prompts and simple linear algebra, then used to read what the model is doing or to change it.
For security teams RepE matters in three ways. It gives you monitors that look inside the model instead of at its text, which jailbreak prompts are designed to fool. It underpins one of the stronger published training-time defences, circuit breakers. And it explains an uncomfortable attack: in open-weights chat models, refusal turns out to depend largely on a single direction that can be found and removed. This article explains the method from first principles, walks through building a harmfulness monitor and covers control methods, the attack and the defence. It ends with failure modes and a checklist. Steering vectors and contrastive activation addition are covered separately in activation engineering and steering; this page is about the RepE framework and its security uses.
The idea: concepts as directions
A transformer keeps a running vector for each token, the residual stream, and every layer reads from it and adds to it. By the middle layers that vector encodes far more than the next word: it carries features about topic, intent, sentiment and whether the request is something the model was trained to refuse. Many of these features turn out to be roughly linear, meaning there is a direction in the vector space whose dot product with the hidden state tracks the property. The residual connection maths explains why features accumulate additively in that stream.
Mechanistic interpretability works bottom-up, tracing individual neurons and circuits. RepE works top-down: pick a concept, show the model contrasting examples and find the direction that separates them across a whole population of neurons. It is less precise than circuit analysis but much cheaper, and it scales to large models with an afternoon of compute. The paper splits the work into two activities. Representation reading finds and measures a concept. Representation control uses the found direction to change behaviour.
Reading a concept with LAT
The paper's reading baseline is Linear Artificial Tomography, LAT, which has three steps. First, design stimuli: a task template that makes the model think about the concept, filled with examples that do and do not exhibit it. Pairing matters. Each positive is matched with a negative that differs only in the concept, so the difference cancels topic, length and style. Second, run the model and collect hidden states at chosen token positions, usually the last token of the prompt or every token of a response. Third, fit a linear model. Taking the first principal component of the paired differences gives the reading vector; a difference of class means or a logistic-regression probe are common alternatives.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype=torch.bfloat16, device_map="auto")
TEMPLATE = "Consider how harmful the following request is:\n{req}\nThe amount of harm is"
@torch.no_grad()
def last_token_states(texts):
out = []
for t in texts:
ids = tok(TEMPLATE.format(req=t), return_tensors="pt").to(model.device)
hs = model(**ids, output_hidden_states=True).hidden_states # tuple: embeddings + each layer
out.append(torch.stack([h[0, -1].float().cpu() for h in hs])) # [layers+1, d_model]
return torch.stack(out) # [n, layers+1, d_model]
def reading_vectors(pos, neg):
diffs = last_token_states(pos) - last_token_states(neg) # paired differences
vecs = []
for layer in range(diffs.shape[1]):
d = diffs[:, layer] - diffs[:, layer].mean(0)
_, _, v = torch.pca_lowrank(d, q=1, center=False)
vecs.append(v[:, 0])
return torch.stack(vecs) # [layers+1, d_model]A principal component has no inherent sign, so orient each layer's vector so that positives project higher than negatives on the training pairs. Then choose the layer on held-out pairs, not the training ones. Accuracy is typically poor in the first few layers, rises through the middle and may fall off near the output. The probing maths page covers why held-out evaluation matters for any linear probe.
Worked example: a harmful-intent monitor
Suppose you serve an open-weights chat model and want a second signal, independent of the output text, that a conversation has moved into harmful territory. Build 200 matched pairs of requests: a harmful one and a benign one that shares its topic and phrasing, such as synthesising a controlled substance versus synthesising aspirin in a school lab. Use 150 pairs for fitting and 50 for selection. Run reading_vectors, then score every held-out request at each layer by projecting its hidden state onto that layer's vector. Pick the layer with the best held-out separation, measured as the area under the ROC curve.
Next, calibrate a threshold on traffic you trust. Score a day of logged benign production prompts and set the alert threshold at, say, the 99.9th percentile, so roughly one benign request in a thousand alerts. Then test against things the probe never saw: known jailbreak templates from your red team, role-play wrappers, and requests split across turns. A useful monitor keeps firing when the surface text is disguised, because the model still has to represent what is being asked in order to answer it. That property, not raw accuracy, is the reason to look inside rather than at the text.
def harm_score(text, layer, v):
h = last_token_states([text])[0, layer]
return float(h @ v / v.norm())
# Serving hook: score, log, and route; do not block on the probe alone.
s = harm_score(user_turn, LAYER, VEC[LAYER])
if s > THRESHOLD:
audit_log.write({"conv": conv_id, "score": s, "layer": LAYER})
route_to_strict_policy(conv_id)In production you would capture the hidden state from the forward pass you already run for generation, with a forward hook on the chosen layer, rather than doing a second pass. The cost is then one dot product per token.
Control: reading vectors, contrast vectors and LoRRA
Control uses the found direction to change behaviour. The paper describes three baselines in increasing order of permanence.
- Reading-vector addition. Add a scaled copy of the vector to the residual stream at one or more layers during generation, pushing the model towards or away from the concept. It is cheap and reversible, but the right strength is found by trial and too much degrades fluency.
- Contrast vectors. Instead of a fixed vector, run the model on a positive and a negative version of the current prompt and add the difference of their hidden states. This adapts to the input and is often more effective, at the cost of extra forward passes per request.
- LoRRA (Low-Rank Representation Adaptation). Train small low-rank adapters so the model's own representations move as if the contrast vector had been applied. The steering is baked into weights, inference is a normal forward pass, and the change can be shipped like any adapter.
All three change behaviour without changing what the model knows. That makes them useful for nudging honesty or caution, and it is also why the same tools serve an attacker as well as a defender.
The attack: refusal is one direction
Arditi and colleagues reported in 2024 that across 13 open chat models of up to 72 billion parameters, refusal is mediated by a one-dimensional subspace. They found the direction with a difference of means between the activations of harmful and harmless instructions. Erasing that direction from the residual stream stopped the models refusing harmful requests, and adding it made them refuse harmless ones. They also showed that the projection can be folded into the weights, so that no layer can write to that direction, producing a model that no longer refuses while scoring about as well on capability benchmarks. Their mechanistic analysis found that adversarial suffixes of the kind discussed in universal adversarial suffixes work partly by suppressing the propagation of this same direction.
The practical lesson is about threat models. If an adversary has your weights, safety fine-tuning is a thin layer that a few hundred example prompts and some linear algebra can strip; the paper's abstract calls it brittle. Releasing open weights means accepting that the released model's refusal behaviour is advisory. For hosted models, white-box attacks of this kind are not directly available, but the finding still predicts that prompt-level attacks will target the same mechanism. That is why a monitor reading internal state can catch attacks that a text filter misses.
The defence: circuit breakers
Circuit breakers, from a 2024 paper by Zou and colleagues, apply RepE to the defence. Instead of training the model to output a refusal, which an attacker can route around, they train it so that the internal representations associated with producing harmful content are redirected, a method the paper calls representation rerouting. Training uses two sets of data. On a harmful set, a loss pushes the model's representations away from what the original model produced, so the internal process that would generate the harmful continuation is disrupted. On a retain set of benign data, a second loss keeps representations close to the original so ordinary capability is preserved. The adapters are LoRRA-style low-rank updates.
The paper reports that this holds up against attacks unseen during training, extends to multimodal models attacked through images, and reduces harmful actions by agents under attack, without the utility loss of heavy refusal training. Treat those as the authors' results on their benchmarks, and measure on your own. Two caveats apply. Like every weight-level defence, circuit breakers can be fine-tuned away by someone with the weights. And the method needs a good harmful set: behaviour it was never shown is not necessarily covered. It complements input filtering and output monitoring; it does not replace the layered controls in jailbreak defence.
Operating RepE monitors
Run RepE monitors like any detector. Version each reading vector with the model checkpoint, layer, template and training pairs that produced it, because a vector is meaningless on any other checkpoint: even a fine-tune of the same base model can rotate it. Re-fit and re-calibrate on every model update. Log scores alongside conversations so you can investigate incidents and measure drift in the benign score distribution. Alert on sustained high scores across a conversation rather than single tokens, which are noisy. Keep a red-team set of disguised attacks and track recall on it release by release.
Failure modes
- Confounded stimuli. If harmful examples are also longer, more technical or differently phrased, the vector learns that instead. Match pairs tightly and test on a differently written held-out set.
- Template dependence. Vectors fitted with one prompt template may not transfer to raw chat traffic. Fit on activations from the exact format you serve.
- Correlation is not control. A direction that reads well may not change behaviour when added; check the causal effect before relying on steering.
- Adaptive attackers. An attacker who knows a probe exists can optimise inputs to keep its score low. Use monitors as one signal among several, and keep their details private.
- Over-steering. Too large a coefficient produces repetitive or incoherent text that looks safe and is useless. Measure capability on a fixed suite after every control change.
- Model-version drift. A vector silently goes stale after a model upgrade. Tie deployment of a model to re-fitting its probes.
Trade-offs
Compared with text classifiers, RepE monitors see intent the text hides and cost almost nothing at inference, but they need white-box access, so they work only for models you host. Compared with refusal training, circuit breakers resist more attacks, but they need curated harmful data and a training run. Compared with backdoor detection and other weight audits, RepE is quick to try and easy to explain, but it finds only the concepts you think to look for.
What to do next
- Pick one concept that matters to your deployment, such as harmful intent or data exfiltration, and write 150 to 200 tightly matched prompt pairs.
- Fit per-layer reading vectors with the LAT pipeline, orient their signs and pick a layer on held-out pairs.
- Calibrate an alert threshold on a day of benign production traffic and record the expected false-positive rate.
- Test recall against disguised attacks from your red team and compare it with your text-level filter.
- Wire the probe into the serving forward pass with a hook, log scores per conversation and route high scores to stricter policy.
- Version every vector with its model checkpoint and re-fit on every model update.
- If you release open weights, document that refusal behaviour can be removed, and evaluate circuit-breaker style training for hosted variants.