SAE features form concept crystals revealed by LDA whitening
measured in 1 paper- Sparse-autoencoder features form geometric "crystals" - parallelograms (b-a ≈ d-c, e.g. man:woman::king:queen) and trapezoids (only one parallel edge). [li-etal-2024-geometry] - These are hidden by distractor dimensions (such as word length); projecting onto a subspace orthogonal to distractors via LDA ("cluster whitening") tightens the clusters and reveals the parallelogram/trapezoid structure. [li-etal-2024-geometry] - Analyzed on Gemma Scope residual-stream SAEs (16k features; the gemma-2-2b layer-12 SAE has average L0=41), with meso-scale spatial-functional modularity (954 / 74 standard deviations above null) and a galaxy-scale power-law eigenvalue spectrum (layer-12 slope -0.47 vs -0.24/-0.25 at layers 0/24). [li-etal-2024-geometry] - Observational; crystal quality depends heavily on the LDA distractor-removal step. Tested on gemma-2-2b and gemma-2-9b. [li-etal-2024-geometry]