MATH · IN · MODELS

Refusal is mediated by a single direction across models

measured in 1 paper

Arditi et al. show across 13 chat models (5 families, 1.8B-72B) that a single diff-in-means direction between harmful and harmless instructions both erases refusal when ablated and induces it when added, with no offset term [arditi-etal-2024] Zou et al. independently confirm via a LAT harmfulness vector in Vicuna-13B that stays a >90% classifier under jailbreaks and adversarial suffixes [arditi-etal-2024] Lee et al. replicate across nine further chat models and add conditional steering that gates the refusal vector on a condition vector for selective refusal [arditi-etal-2024] Yung et al. average per-model refusal vectors into a universal refusal direction transferring to jailbreak detection ~10% better than single-model baselines [arditi-etal-2024] The account is contested by Marshall et al. (an affine offset is needed) and Wollschlager et al. (a cone of directions), and Yin et al. localize attack-suppressed versus robust heads [arditi-etal-2024]

Context

refusal, safety, jailbreak, refusal geometry debate

Papers

Refusal in Language Models Is Mediated by a Single Direction — Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., Nanda, N.2024 · arXiv:2406.11717
Representation Engineering: A Top-Down Approach to AI Transparency — Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A., Goel, S., Li, N., Byun, M., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, Z., Hendrycks, D.2023 · arXiv:2310.01405
Programming Refusal with Conditional Activation Steering — Lee, Bruce W., Padhi, Inkit, Ramamurthy, Karthikeyan Natesan, Miehling, Erik, Dognin, Pierre, Nagireddy, Manish, Dhurandhar, Amit2024 · arXiv:2409.05907
Refusal Direction is Universal Across Safety-Aligned Languages — Wang, Xinpeng, Wang, Mingyang, Liu, Yihong, Schütze, Hinrich, Plank, Barbara2025 · arXiv:2505.17306