MATH · IN · MODELS

Feature absorption is decoder-direction composition, causally isolable

measured in 1 paper

Chanin et al. formalize SAE feature absorption in a toy model as a child latent's decoder direction acquiring a component along a parent feature (W_d2 = f2 + delta*f1) [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders] Empirically they measure a feature-absorption rate on first-letter latents in Gemma Scope SAEs (Gemma-2-2B) plus their own SAEs trained on Qwen2-0.5B and Llama-3.2-1B [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders] Absorption is detected when a latent has cosine >0.025 with the probe direction and the largest negative ablation effect (at least 1.0 above the runner-up) [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders] Projecting the probe direction out of an absorbing latent removes its ablation effect, confirming the composition model causally [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders] Varying SAE width or sparsity alone does not resolve absorption, motivating architectural fixes like Matryoshka SAEs [chanin-etal-2024-feature-splitting-absorption-sparse-autoencoders]

Context

delta-absorption -- a formal model of feature absorption as decoder-direction composition (child direction = parent direction * delta + child's own direction), combined cosine-similarity-plus-ablation-effect metric for detecting absorption empirically, direction-restricted (projection) causal intervention isolating the absorbed component's behavioral responsibility, absorption persisting across SAE width/sparsity variation, motivating architectural fixes like Matryoshka SAEs

Papers

A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders — Chanin, David, Wilken-Smith, James, Dulka, Tomáš, Bhatnagar, Hardik, Golechha, Satvik, Bloom, Joseph2024 · arXiv:2409.14507