MATH · IN · MODELS

CLIP binding is linearly decodable within a modality but collapses across

measured in 1 paper

Koishigarina et al. linearly probe CLIP (ViT-B/32, ViT-B/16, ViT-L/14) for attribute-object binding [koishigarina-etal-2026-clip-behaves-like-a-bag-of-words-model-cross-modally-but-not-uni-modally] Within a single modality (text-only or image-only) linear probes reach near-ceiling accuracy, so binding is linearly present uni-modally [koishigarina-etal-2026-clip-behaves-like-a-bag-of-words-model-cross-modally-but-not-uni-modally] Cross-modally, baseline CLIP binding accuracy collapses to near chance (0.51-0.52) [koishigarina-etal-2026-clip-behaves-like-a-bag-of-words-model-cross-modally-but-not-uni-modally] A lightweight learned linear map recovers cross-modal binding to 0.83-0.98 across the three variants, so the bottleneck is a cross-modal alignment gap rather than missing linear structure [koishigarina-etal-2026-clip-behaves-like-a-bag-of-words-model-cross-modally-but-not-uni-modally]

Context

attribute-binding, modality-gap

Papers

CLIP Behaves Like a Bag-of-Words Model Cross-modally but Not Uni-modally — Koishigarina, Darina, Uselis, Arnas, Oh, Seong Joon2026 · arXiv:2502.03566