MATH · IN · MODELS

Modality gap converges to angle between collapsed hyperplanes

measured in 1 paper

- The modality gap originates from dimension collapse: image and text representations each collapse onto distinct low-dimensional hyperplanes. [yi-etal-2025-decipher-the-modality-gap-in-multimodal-contrastive-learning] - At the contrastive optimum the modality means become orthogonal to the shared subspace, so the gap converges to the smallest angle between the two hyperplanes (Theorem 3). [yi-etal-2025-decipher-the-modality-gap-in-multimodal-contrastive-learning] - On CLIP ViT-B/32 the gap angle is 74.69 deg (CIFAR-10), 74.19 and 71.02 (ImageNet); a shared-subspace projection reduces it to 5.37/30.39/50.40, with an estimated ~212-dim shared space. [yi-etal-2025-decipher-the-modality-gap-in-multimodal-contrastive-learning] - Single model CLIP ViT-B/32; ViT-L/14 and RN50 are cited but not tested, and interventions are post-hoc geometric (projection/translation/rotation). [yi-etal-2025-decipher-the-modality-gap-in-multimodal-contrastive-learning]

Context

modality-gap, dimension-collapse

Papers

Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment — Yi, Lingjie, Douady, Raphael, Chen, Chao2025 · arXiv:2510.03268