Refusal is mediated by a single direction across models
measured in 1 paperArditi et al. show across 13 chat models (5 families, 1.8B-72B) that a single diff-in-means direction between harmful and harmless instructions both erases refusal when ablated and induces it when added, with no offset term [arditi-etal-2024] Zou et al. independently confirm via a LAT harmfulness vector in Vicuna-13B that stays a >90% classifier under jailbreaks and adversarial suffixes [arditi-etal-2024] Lee et al. replicate across nine further chat models and add conditional steering that gates the refusal vector on a condition vector for selective refusal [arditi-etal-2024] Yung et al. average per-model refusal vectors into a universal refusal direction transferring to jailbreak detection ~10% better than single-model baselines [arditi-etal-2024] The account is contested by Marshall et al. (an affine offset is needed) and Wollschlager et al. (a cone of directions), and Yin et al. localize attack-suppressed versus robust heads [arditi-etal-2024]