SigLIP
Google DeepMind
Structures found in this family (4)
By model (2)
Observations (6)
Papers
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers (2026), Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders (2026), Canonicalizing Multimodal Contrastive Representation Learning (2026), Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings (2026), Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models (2026), Interpreting the Linear Structure of Vision-Language Model Embedding Spaces (2025)