MATH · IN · MODELS

Matryoshka nested dictionaries reduce SAE feature absorption

measured in 1 paper

Bussmann et al. address feature absorption, where a general latent develops systematic blind spots ceded to specialized latents, by training multiple nested prefixes of one shared dictionary to each reconstruct the input independently [bussmann-etal-2025-matryoshka-saes] Because the smallest prefix must explain the input alone, without access to the specialized latents, it keeps general concepts intact rather than ceding them [bussmann-etal-2025-matryoshka-saes] On Gemma-2-2B and a from-scratch TinyStories transformer, Matryoshka SAEs significantly beat benchmark architectures on targeted-concept-erasure (SCR) and disentanglement (TPP) and match BatchTopK on sparse probing [bussmann-etal-2025-matryoshka-saes] The gain costs roughly 50% more training time and slightly lower reconstruction, with a tunable dial between nesting granularity and fidelity [bussmann-etal-2025-matryoshka-saes]

Context

shared dictionary trained via multiple independently-reconstructing nested prefixes, feature absorption (general latent develops systematic blind spots ceded to specialized latents), targeted concept erasure (SCR) and disentanglement (TPP) improved over standard BatchTopK, tunable tradeoff between nesting granularity and reconstruction/interpretability

Papers

Learning Multi-Level Features with Matryoshka Sparse Autoencoders — Bussmann, Bart, Nabeshima, Noa, Karvonen, Adam, Nanda, Neel2025 · arXiv:2503.17547