SAE-latent misdirection fine-tuning unlearns entities better than residual-stream baselines
measured in 1 paperYamashita et al. fine-tune Llama-3.1-8B-Instruct and Gemma-2-9B-Instruct with a hinge loss over SAE-latent pre-activations, pushing known-entity latent coordinates below -c and unknown-entity coordinates above +c [yamashita-etal-2025-sparse-autoencoder-guided-internal-representation-unlearning-for-large-language-models] On the RWKU benchmark, Llama's forget score drops 81.1% to 46.8% (versus 65.5-79.2% for gradient-ascent/NPO/RMU baselines) with retain score preserved [yamashita-etal-2025-sparse-autoencoder-guided-internal-representation-unlearning-for-large-language-models] Gemma-2-9B shows a similar pattern (80.1% to 57.1% forget versus 71.6-79.2% baselines) [yamashita-etal-2025-sparse-autoencoder-guided-internal-representation-unlearning-for-large-language-models] Recognition-latent activation-frequency plots show known-latents suppressed and unknown-latents boosted post-intervention, unlike the gradient-ascent baseline [yamashita-etal-2025-sparse-autoencoder-guided-internal-representation-unlearning-for-large-language-models]