VLM SAE concepts are single-modality yet gap-orthogonal
measured in 1 paper- BatchTopK sparse autoencoders trained on ~600k COCO vision-language embeddings (dictionary 4096) put 99% of the energy in 512 concepts, the other ~3500 sharing 1%. [papadimitriou-etal-2025-vlm-linear-structure] - The energy-dominant concepts are nearly single-modality in activation, yet many of their directions are nearly orthogonal to the modality-gap subspace (accuracy ~0.5 as modality classifiers). [papadimitriou-etal-2025-vlm-linear-structure] - An "SAE projection effect" reconciles this: sparse thresholding interacts with the per-modality input distributions, so a direction can be gap-orthogonal yet fire for one modality; top-512 concepts are seed-stable (0.92) vs 0.16 for low-energy ones. [papadimitriou-etal-2025-vlm-linear-structure] - Observational, on CLIP, SigLIP, SigLIP2 and AIMv2 (the paper states no model variants). [papadimitriou-etal-2025-vlm-linear-structure]