Back-projecting the refusal direction identifies attack-suppressed vs attack-robust heads
measured in 1 paperYin et al. back-project the mid-layer refusal direction through each attention head's OV circuit to get a per-head activation score, a rigorous linear-algebraic decomposition of the refusal signal [yin-etal-2026-robust-harmful-features-jailbreak-attacks-attention-head-specialization] A two-stage KDE-overlap filter classifies heads into Adversarially Compromised Heads (early-layer, suppressed by attack-template tokens) and Safety-Aligned Heads (mid-layer, robust under attack) [yin-etal-2026-robust-harmful-features-jailbreak-attacks-attention-head-specialization] Ablating just 8 ACHs induces jailbreak-like behavior (0% to 95.0% ASR on Llama-3-8B, 0% to 81.6% on Llama-2-7B) versus only 4-10% for matched random-head controls [yin-etal-2026-robust-harmful-features-jailbreak-attacks-attention-head-specialization] A training-free detector built from the same direction-projection scores matches or beats dedicated safety classifiers, explaining why jailbreaks bypass rather than eliminate the refusal direction [yin-etal-2026-robust-harmful-features-jailbreak-attacks-attention-head-specialization]