MATH · IN · MODELS

A single SAE latent causally mediates refusal with a dose-response tradeoff

measured in 1 paper

O'Brien et al. train a TopK SAE on Phi-3-mini's layer-6 residual stream and identify a single latent (Feature 22373) whose activation marks refusal [obrien-etal-2024-steering-refusal-with-sae-features] Clamping this decoded direction monotonically raises WildGuard unsafe-prompt refusal from 58.33% to 96.02% and cuts Crescendo jailbreak attack-success from 55.92% to 32.58% [obrien-etal-2024-steering-refusal-with-sae-features] The same clamped direction generalizes to Llama-3.1-8B-Instruct [obrien-etal-2024-steering-refusal-with-sae-features] A safety-capability tradeoff is quantified: safe-prompt refusal rises substantially and MMLU degrades (68.80% to 35.98% at clamp 12), a cost left mechanistically unexplained [obrien-etal-2024-steering-refusal-with-sae-features]

Context

SAE-decoded refusal direction, dose-response steering, safety-capability trade-off

Papers

Steering Language Model Refusal with Sparse Autoencoder Features — O'Brien, Kyle, Majercak, David, Fernandes, Xavier, Edgar, Richard, Bullwinkel, Blake, Chen, Jingya, Nori, Harsha, Carignan, Dean, Horvitz, Eric, Poursabzi-Sangdeh, Forough2024 · arXiv:2411.11296