MATH · IN · MODELS

VLM SAE concepts are single-modality yet gap-orthogonal

measured in 1 paper

- BatchTopK sparse autoencoders trained on ~600k COCO vision-language embeddings (dictionary 4096) put 99% of the energy in 512 concepts, the other ~3500 sharing 1%. [papadimitriou-etal-2025-vlm-linear-structure] - The energy-dominant concepts are nearly single-modality in activation, yet many of their directions are nearly orthogonal to the modality-gap subspace (accuracy ~0.5 as modality classifiers). [papadimitriou-etal-2025-vlm-linear-structure] - An "SAE projection effect" reconciles this: sparse thresholding interacts with the per-modality input distributions, so a direction can be gap-orthogonal yet fire for one modality; top-512 concepts are seed-stable (0.92) vs 0.16 for low-energy ones. [papadimitriou-etal-2025-vlm-linear-structure] - Observational, on CLIP, SigLIP, SigLIP2 and AIMv2 (the paper states no model variants). [papadimitriou-etal-2025-vlm-linear-structure]

Context

sparse autoencoder concepts, energy-weighted stability, modality score, Bridge Score, orthogonality to modality subspace, SAE thresholding artifact, cross-modal semantic bridging

Papers

Interpreting the Linear Structure of Vision-Language Model Embedding Spaces — Papadimitriou, Isabel, Su, Huangyuan, Fel, Thomas, Kakade, Sham, Gil, Stephanie2025 · arXiv:2504.11695