MATH · IN · MODELS

CLIP (Contrastive Language-Image Pretraining)

OpenAI

Hypotheses argued for by this family (1)

Observations (31)

Papers

What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies (2026), CLIP Behaves Like a Bag-of-Words Model Cross-modally but Not Uni-modally (2026), Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning (2022), Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers (2026), The Double-Ellipsoid Geometry of CLIP (2025), It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap (2024), Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders (2026), Probing the 3D Awareness of Visual Foundation Models (2024), Directional Neural Collapse Explains Few-Shot Transfer in Self-Supervised Learning (2026), Revisiting the Platonic Representation Hypothesis: An Aristotelian View (2026), On the Spectral Geometry of Cross-Modal Representations: A Functional Map Diagnostic for Multimodal Alignment (2026), Canonicalizing Multimodal Contrastive Representation Learning (2026), Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings (2026), Harnessing the Universal Geometry of Embeddings (2026), Vision Transformers Don't Need Trained Registers (2025), Beyond Compression: Quantifying Spectral Accessibility in Vision Representations (2026), Language models align with brain regions that represent concepts across modalities (2025), From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models (2025), Linearly Mapping from Image to Text Space (2023), Text-to-Concept (and Back) via Cross-Model Alignment (2023), IdEst: Assessing Self-Supervised Learning Representations via Intrinsic Dimension (2026), Parts of Speech-Grounded Subspaces in Vision-Language Models (2023), Interpretable and Testable Vision Features via Sparse Autoencoders (2025), Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE) (2024), Decoding Vision Transformers: The Diffusion Steering Lens (2025), Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment (2025), On the Modality Gap and the Contrastive Loss in Multi-modal Representation Learning (2026), Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models (2026), How Can Embedding Models Bind Concepts? (2026), Emergent Visual-Semantic Hierarchies in Image-Text Representations (2024), Interpreting the Linear Structure of Vision-Language Model Embedding Spaces (2025)