Two activation statistics predict steering-vector reliability
measured in 1 paperBraun et al. construct diff-in-means steering directions in Llama-2-7B-Chat across 36 binary behavior datasets [braun-etal-2025-steering-vector-reliability] They quantify mean training-set cosine agreement with the aggregate direction and a signal-detection discriminability index d' of positive/negative separation [braun-etal-2025-steering-vector-reliability] Both statistics predict how often the causal steering effect reverses sign (anti-steerable samples, 3-50% across datasets) [braun-etal-2025-steering-vector-reliability] Better-aligned, better-separated datasets are causally more reliable to steer, a trend supported by ranges and figures rather than formal correlation statistics [braun-etal-2025-steering-vector-reliability]
Structure
Context
diff-in-means, cosine similarity, discriminability index, signal detection theory, steering reliability, anti-steerable samples
Confirmed in models
Papers
Understanding (Un)Reliability of Steering Vectors in Language Models — Braun, Joschka, Eickhoff, Carsten, Krueger, David, Bahrainian, Seyed Ali, Krasheninnikov, Dmitrii