CLIP embeddings lie on two tilted offset ellipsoids
measured in 1 paper- Pre-normalization CLIP embeddings for each modality lie on a thin ellipsoidal shell: long-tailed per-feature variance makes it an ellipsoid and off-diagonal covariance tilts it. [levi-gilboa-2025-double-ellipsoid-clip] - The image and text shells are separable and centered well away from the origin (the "double-ellipsoid"); a linear SVM separates the modalities with 100% accuracy using just two features. [levi-gilboa-2025-double-ellipsoid-clip] - The thin-shell approximation is tight (image mu_norm=7.59, 0.18% relative error) and a conformity score tracks generation quality at Pearson 0.9998. [levi-gilboa-2025-double-ellipsoid-clip] - Observational on CLIP ViT-B/32 (primary) and ViT-L/14 (supplementary); interventions are post-hoc embedding shifts. [levi-gilboa-2025-double-ellipsoid-clip]