Matryoshka nested dictionaries reduce SAE feature absorption
measured in 1 paperBussmann et al. address feature absorption, where a general latent develops systematic blind spots ceded to specialized latents, by training multiple nested prefixes of one shared dictionary to each reconstruct the input independently [bussmann-etal-2025-matryoshka-saes] Because the smallest prefix must explain the input alone, without access to the specialized latents, it keeps general concepts intact rather than ceding them [bussmann-etal-2025-matryoshka-saes] On Gemma-2-2B and a from-scratch TinyStories transformer, Matryoshka SAEs significantly beat benchmark architectures on targeted-concept-erasure (SCR) and disentanglement (TPP) and match BatchTopK on sparse probing [bussmann-etal-2025-matryoshka-saes] The gain costs roughly 50% more training time and slightly lower reconstruction, with a tunable dial between nesting granularity and fidelity [bussmann-etal-2025-matryoshka-saes]