MATH · IN · MODELS

SAE-latent misdirection fine-tuning unlearns entities better than residual-stream baselines

measured in 1 paper

Yamashita et al. fine-tune Llama-3.1-8B-Instruct and Gemma-2-9B-Instruct with a hinge loss over SAE-latent pre-activations, pushing known-entity latent coordinates below -c and unknown-entity coordinates above +c [yamashita-etal-2025-sparse-autoencoder-guided-internal-representation-unlearning-for-large-language-models] On the RWKU benchmark, Llama's forget score drops 81.1% to 46.8% (versus 65.5-79.2% for gradient-ascent/NPO/RMU baselines) with retain score preserved [yamashita-etal-2025-sparse-autoencoder-guided-internal-representation-unlearning-for-large-language-models] Gemma-2-9B shows a similar pattern (80.1% to 57.1% forget versus 71.6-79.2% baselines) [yamashita-etal-2025-sparse-autoencoder-guided-internal-representation-unlearning-for-large-language-models] Recognition-latent activation-frequency plots show known-latents suppressed and unknown-latents boosted post-intervention, unlike the gradient-ascent baseline [yamashita-etal-2025-sparse-autoencoder-guided-internal-representation-unlearning-for-large-language-models]

Context

representation misdirection applied to SAE-decomposed latent coordinates rather than raw residual-stream activations, as a higher-precision unlearning target, a hinge-loss margin objective pushing latent coordinates past a threshold rather than toward a single fixed point

Papers

Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models — Yamashita, Tomoya, Ito, Akira, Yamanaka, Yuuki, Miura, Takayuki, Shibahara, Toshiki2025 · arXiv:2509.15631