A single SAE latent causally mediates refusal with a dose-response tradeoff
measured in 1 paperO'Brien et al. train a TopK SAE on Phi-3-mini's layer-6 residual stream and identify a single latent (Feature 22373) whose activation marks refusal [obrien-etal-2024-steering-refusal-with-sae-features] Clamping this decoded direction monotonically raises WildGuard unsafe-prompt refusal from 58.33% to 96.02% and cuts Crescendo jailbreak attack-success from 55.92% to 32.58% [obrien-etal-2024-steering-refusal-with-sae-features] The same clamped direction generalizes to Llama-3.1-8B-Instruct [obrien-etal-2024-steering-refusal-with-sae-features] A safety-capability tradeoff is quantified: safe-prompt refusal rises substantially and MMLU degrades (68.80% to 35.98% at clamp 12), a cost left mechanistically unexplained [obrien-etal-2024-steering-refusal-with-sae-features]