CLIP binding is linearly decodable within a modality but collapses across
measured in 1 paperKoishigarina et al. linearly probe CLIP (ViT-B/32, ViT-B/16, ViT-L/14) for attribute-object binding [koishigarina-etal-2026-clip-behaves-like-a-bag-of-words-model-cross-modally-but-not-uni-modally] Within a single modality (text-only or image-only) linear probes reach near-ceiling accuracy, so binding is linearly present uni-modally [koishigarina-etal-2026-clip-behaves-like-a-bag-of-words-model-cross-modally-but-not-uni-modally] Cross-modally, baseline CLIP binding accuracy collapses to near chance (0.51-0.52) [koishigarina-etal-2026-clip-behaves-like-a-bag-of-words-model-cross-modally-but-not-uni-modally] A lightweight learned linear map recovers cross-modal binding to 0.83-0.98 across the three variants, so the bottleneck is a cross-modal alignment gap rather than missing linear structure [koishigarina-etal-2026-clip-behaves-like-a-bag-of-words-model-cross-modally-but-not-uni-modally]