Feature absorption is decoder-direction composition, causally isolable
measured in 1 paperChanin et al. formalize SAE feature absorption in a toy model as a child latent's decoder direction acquiring a component along a parent feature (W_d2 = f2 + delta*f1) [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders] Empirically they measure a feature-absorption rate on first-letter latents in Gemma Scope SAEs (Gemma-2-2B) plus their own SAEs trained on Qwen2-0.5B and Llama-3.2-1B [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders] Absorption is detected when a latent has cosine >0.025 with the probe direction and the largest negative ablation effect (at least 1.0 above the runner-up) [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders] Projecting the probe direction out of an absorbing latent removes its ablation effect, confirming the composition model causally [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders] Varying SAE width or sparsity alone does not resolve absorption, motivating architectural fixes like Matryoshka SAEs [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders]