MATH · IN · MODELS

A supervised SAE gives one latent per concept for erasure

measured in 1 paper

Cassano et al. train a supervised TopK sparse autoencoder on Stable Diffusion v1.5's up.1.1 cross-attention activations, using a concept-vs-non-concept score to force a one-to-one concept-to-latent mapping [cassano-etal-2025-saemnesia-erasing-concepts-in-diffusion-models-with-supervised-sparse-autoencoders] Erasing a concept reduces to steering that single SAE latent direction [cassano-etal-2025-saemnesia-erasing-concepts-in-diffusion-models-with-supervised-sparse-autoencoders] On UnlearnCanvas this improves 9.2% over the prior SAE-based SOTA (SAeUron) [cassano-etal-2025-saemnesia-erasing-concepts-in-diffusion-models-with-supervised-sparse-autoencoders] On 9-object sequential unlearning accuracy improves 28.4 points and hyperparameter search is cut 96.7%, with nudity removal on I2P also demonstrated [cassano-etal-2025-saemnesia-erasing-concepts-in-diffusion-models-with-supervised-sparse-autoencoders]

Context

a supervision signal forcing SAE latents into a clean one-to-one correspondence with target concepts, rather than relying on post-hoc feature interpretation, single-latent-direction steering as the causal erasure mechanism in a real pretrained diffusion U-Net

Papers

SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders — Cassano, Enrico, Renzulli, Riccardo, Nurisso, Marco, Zaffaroni, Marco, Perotti, Alan, Grangetto, Marco2025 · arXiv:2509.21379