MATH · IN · MODELS

Refusal is a causally-minimal SAE-latent set with hydra redundancy

measured in 1 paper

Prakash et al. (AAAI 2026) extend single-direction refusal ablation into a multi-latent SAE decomposition on Gemma-2-2B-IT and Llama-3.1-8B-IT using Gemma Scope / LlamaScope [prakash-etal-2026-dissecting-llm-refusal-sae-feature-sets] They identify causally-minimal sets of SAE latents near the refusal direction whose joint ablation reduces refusal (thousands of candidate features for Gemma, ~110 for Llama, narrowed at later stages) [prakash-etal-2026-dissecting-llm-refusal-sae-feature-sets] Ablating the active refusal-latent set triggers a hydra effect: ~74% of previously-dormant redundant features (Gemma) activate on system-prompt tokens instead, ~97% on the begin-of-text token [prakash-etal-2026-dissecting-llm-refusal-sae-feature-sets] Refusal is thus a redundant multi-dimensional causal subspace rather than a single vector [prakash-etal-2026-dissecting-llm-refusal-sae-feature-sets] Ablating the identified latent set shifts attack success rate 4%->33% (Gemma) and 71%->57% (Llama) across stages, with hydra redundancy explaining incomplete restoration [prakash-etal-2026-dissecting-llm-refusal-sae-feature-sets]

Context

refusal is implemented by a structured multi-latent SAE feature set, not a single direction, redundant "hydra" features reactivate on system-prompt tokens after ablation of the primary set, ablating the identified latent set causally shifts jailbreak attack success rate

Papers

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal — Prakash, Nirmalendu, Yeo, Wei Jie, Abdullah, Amir, Satapathy, Ranjan, Cambria, Erik, Lee, Roy Ka-Wei2026 · arXiv:2509.09708