MATH · IN · MODELS

Iso-energy SAE isolates a bimodal cross-modal subspace

measured in 1 paper

- An iso-energy-regularized aligned SAE splits its dictionary into bimodal atoms (a modality-agnostic shared subspace carrying essentially all cross-modal alignment) and unimodal atoms (per-modality cones that account for the modality gap). [dhimoila-etal-2026-cross-modal-redundancy] - Ablating the unimodal atoms nearly eliminates the modality gap while barely changing retrieval recall (CLIP-B/32 delta_r 0.224 to 0.125 from SAE to SAE-A; functional alignment rho 0.327 to 4.232). [dhimoila-etal-2026-cross-modal-redundancy] - Restricting vector-arithmetic edits to the bimodal subspace keeps them in-distribution and improves OOD retrieval (CLIP 0.97 to 0.77, SigLIP2 0.99 to 0.61); reconstruction R-squared >=0.859. [dhimoila-etal-2026-cross-modal-redundancy] - The paper measures retrieval (delta_r), not post-ablation zero-shot classification; tested on CLIP and OpenCLIP ViT-B/32 and ViT-L/14, SigLIP and SigLIP2. [dhimoila-etal-2026-cross-modal-redundancy]

Context

Iso-Energy Assumption, bimodal subspace, unimodal cones, modality score, functional alignment ratio, atom ablation, subspace-restricted vector arithmetic

Papers

Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings — Dhimoïla, Grégoire, Feld, Thomas, Boutin, Victor, Picard, Agustin2026 · arXiv:2602.06218