MATH · IN · MODELS

SAEs dilute continuous manifolds rather than compactly capturing them

measured in 1 paper

Bhalla et al. formalize manifold capture (a small fixed group of decoder atoms spanning the manifold, consistently reselected by the encoder) and identify three regimes: compact capture, tiling/shattering, and intermediate dilution [bhalla-etal-2026] Training five SAE architectures on Llama-3.1-8B layer-19 activations, variance explained by a restricted atom group plateaus well beyond each manifold's ambient dimension, so none achieve compact capture [bhalla-etal-2026] Features behave like overlapping population-code tuning curves that redundantly tile the manifold, placing all tested SAEs in the dilution regime [bhalla-etal-2026] Ising-coactivation analysis on SAE codes recovers known manifolds (temperature, colors, political bias) unsupervised and surfaces a new epistemic-uncertainty manifold [bhalla-etal-2026]

Context

sparse autoencoders, dilution, tiling, unsupervised manifold discovery, population coding

Confirmed in models

Papers

Do Sparse Autoencoders Capture Concept Manifolds? — Bhalla, Usha, Fel, Thomas, Rager, Can, Feucht, Sheridan, Haklay, Tal, Wurgaft, Daniel, Boppana, Siddharth, Kowal, Matthew, Shyam, Vasudev, Lewis, Owen, McGrath, Thomas, Merullo, Jack, Geiger, Atticus, Lubana, Ekdeep Singh2026 · arXiv:2604.28119