Independent multimodal models share one orthogonal map across both modalities
measured in 1 paperGupta et al. show independently-trained multimodal contrastive models' embedding spaces are related, up to a mean shift, by a single orthogonal map Q applied identically to both the image and text encoders [gupta-etal-2026-canonicalizing-multimodal-contrastive-representation-learning] They prove that kernel agreement on a small image-text anchor set forces this single-shared-orthogonal-map relationship [gupta-etal-2026-canonicalizing-multimodal-contrastive-representation-learning] It is verified across CLIP, SigLIP and FLAVA models with different architectures and training data [gupta-etal-2026-canonicalizing-multimodal-contrastive-representation-learning] This is a discovered (not imposed) geometric relationship, solved in closed form via the orthogonal Procrustes SVD [gupta-etal-2026-canonicalizing-multimodal-contrastive-representation-learning]