Interpretability is the set of methods for explaining why a model produced an output, using the model's inputs, gradients or internal activations rather than its own account of itself. For a security team it is not an academic hobby. It answers practical questions. Did the model follow the user or the injected document? Is it representing "this request is harmful" even when it complies? Did fine-tuning plant a trigger? Can we monitor for an internal state that never shows up in the text?

The field offers many tools, and they differ a great deal in what they can prove. This article is a method-selection guide. It covers input attribution with integrated gradients, logit and tuned lenses, linear probes with proper controls, and the causal tests that turn a correlation into evidence. It ends with the part security teams most often skip: explanations can be wrong, and they can be attacked. Deeper workflows have their own articles: mechanistic investigations, sparse-autoencoder feature monitors and activation steering.

Start from the question

Start from the question, not the method. Security questions fall into three kinds, and each points to a different part of the model:

  • Which input caused this? For example, which retrieved chunk made the agent call the payment tool. This is input attribution.
  • What does the model represent here? For example, does it internally flag the request as a jailbreak, or know that it is being evaluated? This calls for probes and dictionary features on activations.
  • Which internal computation produced the behaviour? For example, which attention heads move the injected instruction into the answer. This calls for causal interventions such as activation patching.
Interpretability methods by what they look at, and the evidence they giveInput tokensprompt + retrieved docsResidual stream, layers 1..Lactivations at each positionOutputlogits, text, CoTAttributionIG: which inputsLenslogit/tuned: whenProbe / SAEwhat is representedSelf-reportCoT: what it saysCausal test: patch or ablateturns a correlate into evidence that it mattersEvidence strength rises toward the causal test; self-report is the weakest on its own.
Each family of methods reads a different part of the model. Only a causal intervention shows that a component matters; the others show what correlates with the behaviour.
MethodReadsCostEvidence it givesMain trap
Integrated gradientsInputs + gradients~20–300 backward passes per outputAdditive credit per input tokenBaseline choice changes the answer
Logit / tuned lensResidual stream per layerOne forward passWhen a prediction formsEarly layers poorly decoded by the raw lens
Linear probeActivations at one layerLabelled set + logistic regressionInformation is linearly presentProbe learns the dataset, not the concept
SAE featuresActivations, dictionaryTraining an SAENamed, sparse directionsFeature labels are hypotheses
Activation patchingClean vs corrupted runsOne forward pass per componentComponent is causally necessary or sufficientOff-distribution patches
Chain-of-thoughtGenerated textFreeWhat the model saysCan be unfaithful

Input attribution with integrated gradients

Plain gradients answer "if I nudged this input embedding, how would the output move?" Because of saturation, a token that matters a lot can still have a near-zero local gradient. Integrated gradients (Sundararajan, Taly and Yan, 2017) fixes this. It integrates the gradient along a straight path from a baseline input x′ to the real input x, and multiplies by the difference. The result satisfies completeness: the attributions sum to f(x) − f(x′). That gives you a built-in correctness check, which few explanation methods offer.

import torch

def integrated_gradients(model, embeds, baseline, target_id, steps=64):
    """embeds, baseline: [T, d] input embeddings. Returns per-token attribution [T]."""
    alphas = torch.linspace(0, 1, steps, device=embeds.device).view(-1, 1, 1)
    path = baseline + alphas * (embeds - baseline)            # [steps, T, d]
    path.requires_grad_(True)
    logits = model(inputs_embeds=path).logits[:, -1, :]      # next-token logits
    score = torch.log_softmax(logits, -1)[:, target_id].sum()
    (grads,) = torch.autograd.grad(score, path)
    avg_grad = (grads[:-1] + grads[1:]).mean(0) / 2           # trapezoid rule
    attr = ((embeds - baseline) * avg_grad).sum(-1)           # [T]
    with torch.no_grad():
        f = lambda e: torch.log_softmax(model(inputs_embeds=e[None]).logits[0, -1], -1)[target_id]
        gap = (f(embeds) - f(baseline)).item()
    rel_err = abs(attr.sum().item() - gap) / (abs(gap) + 1e-9)
    return attr, rel_err      # rel_err > ~0.05: raise steps before trusting attr

Three choices decide whether the result means anything. The baseline: all-zero embeddings are off-distribution for a transformer. A sequence of pad or neutral filler tokens of the same length is usually better. Report which baseline you used, and check that the conclusion survives a second one. The target: attribute the log-probability of the action token (for example, the first token of the tool name), not a whole generated paragraph. Steps: raise them until the completeness error is small. For long documents, attribute to chunks by summing token scores, since chunk-level answers are what an incident report needs.

Lenses: watching a decision form

The logit lens (nostalgebraist, 2020) applies the model's final layer norm and unembedding to the residual stream at an intermediate layer, then reads off what the model "would predict" at that depth. The tuned lens (Belrose et al., 2023) learns a small affine map per layer so that intermediate states decode reliably. That matters because the raw lens is often unreadable in early layers. In security work, a lens answers when a decision forms. If the refusal token is already dominant by the middle layers on a benign-looking prompt, something early in the context is steering the outcome.

@torch.no_grad()
def logit_lens(model, input_ids, watch_ids):
    out = model(input_ids, output_hidden_states=True)
    rows = []
    # In HF Llama the last hidden_states entry is already normed; skip it, use out.logits
    for layer, h in enumerate(out.hidden_states[:-1]):       # [1, T, d] per layer
        logits = model.lm_head(model.model.norm(h[:, -1]))   # final position only
        probs = logits.softmax(-1)[0, watch_ids]
        rows.append((layer, probs.tolist()))
    rows.append((len(rows), out.logits[0, -1].softmax(-1)[watch_ids].tolist()))
    return rows    # e.g. watch_ids = [id("I"), id("Sure")] to see refusal vs compliance form

Attribute names such as model.model.norm follow the Hugging Face Llama layout; other architectures name these modules differently. A lens is a cheap first look, not a conclusion. It shows the trajectory, and you still need a causal test to show what drives it.

Probes, with the controls that make them honest

A linear probe is a logistic regression trained on activations at one layer to predict a property, such as "this prompt contains an injected instruction". Probes are cheap to run (one dot product per token at inference) and work as runtime monitors. Their weakness is that a probe can score well because the dataset has a shortcut, such as a template phrase or a length difference, not because the model represents the concept. Hewitt and Liang (2019) proposed control tasks, which give inputs random but consistent labels. A simple version is to train the same probe on shuffled labels and trust only the gap between real and control scores.

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
import numpy as np

def fit_probe(acts_train, y_train, acts_test, y_test, groups_test):
    """acts_*: [N, d] activations at one layer, mean-pooled over the span of interest."""
    probe = LogisticRegression(C=0.1, max_iter=2000).fit(acts_train, y_train)
    auc = roc_auc_score(y_test, probe.predict_proba(acts_test)[:, 1])
    ctrl = LogisticRegression(C=0.1, max_iter=2000).fit(
        acts_train, np.random.permutation(y_train))
    ctrl_auc = roc_auc_score(y_test, ctrl.predict_proba(acts_test)[:, 1])
    # held-out *families* (e.g. attack templates never seen in training), not just rows
    per_group = {g: roc_auc_score(y_test[groups_test == g],
                                  probe.predict_proba(acts_test[groups_test == g])[:, 1])
                 for g in np.unique(groups_test) if len(set(y_test[groups_test == g])) == 2}
    return probe, auc, ctrl_auc, per_group

Split by attack family, not by row. A probe that only recognises the templates it was trained on has learned a signature, and a classifier on the text would do the same job. Compare against that baseline too. If a fine-tuned text classifier matches the probe, the probe's added value is its cost and its position inside the model, not new information. Run the probe alongside existing defences such as those in indirect injection defences, not instead of them.

Worked example: an instruction hidden in a retrieved email

Suppose a RAG assistant summarising vendor emails produced a reply that included a "please update our bank details" line taken from one email. Did the model treat that email's text as an instruction, or as content to report? Each method contributes one step:

  1. Attribution. Run integrated gradients on the log-probability of the first token of the injected line, with a filler-token baseline, and sum by retrieved chunk. One chunk carries most of the credit, and the completeness error is under 2%. Re-running with a second baseline leaves the ranking unchanged.
  2. Lens. At the position just before the injected line, the tuned lens shows its first token rising in the middle layers. The decision forms well before the output, not at the last layer.
  3. Probe. An instruction-versus-data probe, validated with a control task and on held-out attack families, fires on that chunk's tokens. It stays quiet on other emails that merely mention bank details.
  4. Causal test. Patch the residual stream at that chunk's positions from a run where the chunk is wrapped in explicit data delimiters. The injected line's probability collapses. Patching any other chunk does nothing. This step turns "correlates with" into "is necessary for".

The incident report can now say, with stated confidence, that the model processed the chunk as an instruction and that delimiting it changes the internal state that drives the behaviour. The probe, after a red-team pass, becomes a candidate monitor. Each step alone would have been suggestive. Together they are evidence.

When explanations fail or are attacked

Explanations are model outputs too, and they fail in known ways.

  • Explanations that ignore the model. Adebayo et al. (2018) showed that some saliency methods produce nearly the same maps after the model's weights are randomised. Those maps were reflecting the input, not the model. Run their randomisation test on any attribution method before relying on it.
  • Manipulable attributions. Dombrowski et al. (2019) showed that small input perturbations can change gradient-based explanations arbitrarily while leaving the prediction unchanged. Slack et al. (2020) built models that behave in a biased way on real inputs but look clean under LIME and SHAP, because those methods query off-distribution points the model can detect. If an adversary controls the input or the model, a clean explanation is not proof of innocence.
  • Evading latent monitors. Bailey et al. (2024) showed that inputs and fine-tunes can be optimised to produce harmful behaviour while keeping activations away from probes, SAE features and other latent-space detectors. A monitor trained once and never attacked will overstate its own recall.
  • Unfaithful chain-of-thought. Turpin et al. (2023) showed that models can be swayed by biasing features in the prompt while their written reasoning never mentions them. Treat a model's stated reasoning as a claim to check, not as a readout of its computation.

The defensive lesson is the same each time. Diversify the evidence (attribution plus probe plus causal test). Red-team the monitor itself, as you would any detector (see backdoor detection). Never let a single interpretability signal be the only control between an attacker and an action.

Operating interpretability in a security programme

  • Pin everything. Probes, lenses and SAEs are tied to one checkpoint. A model update, a new quantisation or even a changed system prompt template can shift activations. Re-validate monitors on every model release, as part of the release gate.
  • Budget the compute. A probe adds a dot product per token. Integrated gradients adds tens to hundreds of backward passes per explained output, so keep it for investigations and sampled audits, not every request.
  • Report evidence levels. State whether a claim rests on correlation (probe, attribution), trajectory (lens) or intervention (patching), and list the baselines and controls that were used.
  • Protect the artefacts. A trained probe tells an attacker exactly which direction to avoid. Treat probe weights and feature lists as sensitive detection logic.
  • Measure monitors like detectors. Track precision, recall on held-out attack families, and false-positive cost, and set thresholds from those numbers, not from how convincing an example looks.

Trade-offs

Interpretability is most valuable when outputs alone cannot settle a question: whether an instruction was obeyed or merely quoted, whether a representation exists before any harmful text is produced, or whether a fine-tune changed behaviour on rare triggers. It is least valuable as a black-box classifier substitute, because a text classifier is often just as accurate and easier to run. It also requires white-box access, so it only applies to models you host. For API models you are limited to behavioural testing and the model's self-reports, and you should weigh conclusions accordingly.

What to do next

  1. Write down the three security questions you most need answered about your model, and classify each as input, representation or computation.
  2. Implement the integrated-gradients function above on a model you host, and confirm the completeness error falls below 5% as steps increase.
  3. Run a logit or tuned lens on ten refused and ten complied prompts, and note the layer where each outcome becomes dominant.
  4. Train one probe for injected instructions, with a control task and a held-out attack family, and compare it against a text classifier.
  5. Run the weight-randomisation sanity check on your attribution method.
  6. Red-team your probe with an attacker who knows it exists before you deploy it as a monitor.
  7. Add an "evidence level" field to your incident report template.
Key takeaway: Pick interpretability methods by the question: attribution for which input, probes and features for what is represented, patching for which computation. Check each method with its own control (completeness, randomised labels, weight randomisation), combine signals before claiming causation, and assume explanations and latent monitors can be fooled until you have red-teamed them.