Representational similarity analysis shows 15 language and vision-language models align more strongly with brain regions whose activity is consistent across sentence, word-cloud, and image presentations of the same concept than with less modality-consistent regions
measured in 1 paperRyskina, Tuckute, Fung, Malkin & Fedorenko (2025) use an fMRI dataset (Pereira et al. 2018) in which the same concepts are presented as sentences, word clouds, and images, defining a "meaning consistency" metric that identifies brain voxels responding similarly to a concept regardless of presentation modality. Using representational similarity analysis (RSA) to compare the geometry of these regions' response patterns against 15 language and vision-language models' own representational geometry, both language-only and language-vision models predict brain signal better in meaning-consistent regions, even in areas with low language-selectivity -- a genuine shared-geometry finding (RSA-based, not merely correlation-strength) linking model representational structure to a cross-modal property of brain organization.