Cross-model and cross-modal representational alignment increases with scale and competence
measured in 1 paperHuh et al. (2024) measure mutual nearest-neighbor kernel alignment between ~20 named open-weight language models (LLaMA/LLaMA-3, BLOOM, OpenLLaMA, Mistral/Mixtral, Gemma, OLMo, spanning 560M to 70B parameters) and a range of vision models (DINOv2, MAE, CLIP, ImageNet-21K-supervised ViTs), using paired Wikipedia image-caption data to bridge the two modalities. They find alignment between a language model and vision models rises linearly with the language model's own language-modeling competence (lower bits-per-byte), and symmetrically with the vision model's competence; separately, among 78 vision-only models, alignment within a competence bucket rises as the bucket's average downstream (VTAB) performance rises. Critically, a language model's alignment score to a strong vision encoder (DINOv2) itself predicts that language model's own downstream task performance (Hellaswag common-sense reasoning shows a linear relationship; GSM8K math shows an emergence-like threshold effect) — alignment with other modalities correlates with, and may indicate, general competence, not just modality-specific skill. This is the primary cross-model, cross-modality empirical support cited for the [[platonic-representation]] hypothesis: unlike every other Observation in this map, its "structure" is not a shape found within one model's representation space, but a convergence relationship measured *between* many independently-trained models and even across data modalities.