MATH · IN · MODELS

Same concept, different SAE directions across modalities

measured in 1 paper

- In vision-language models, sparse-autoencoder feature directions for the same concept differ across modalities, keeping substantial angular (cosine-distance) separation even when their activations are highly correlated. [lee-etal-2026-cross-modal-sae-heterogeneity] - The paper proves that a single forced-alignment SAE collapses the two directions to their normalized sum [(phi+psi)/||phi+psi||, 0], sacrificing reconstruction. [lee-etal-2026-cross-modal-sae-heterogeneity] - Modality-specific SAEs plus post-hoc alignment give better cross-modal retrieval and steering (image->text R@1 16.0, text->image R@1 11.4). [lee-etal-2026-cross-modal-sae-heterogeneity] - Directional analysis spans CLIP ViT-B/32, MetaCLIP B/32, OpenCLIP B/32 and SigLIP2. [lee-etal-2026-cross-modal-sae-heterogeneity]

Context

SAE decoder-column heterogeneity, coactivation-matched concepts, cosine distance distribution, forced-alignment convergence, feature collapse rate, cross-modal steering vector, retrieval-based causal validation

Papers

Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders — Lee, Chungpa, Kwon, Jihoon, Min, Kyle, Sohn, Jy-yong2026 · arXiv:2606.29888