Behavioral grounding survives rotation but collapses under random reassignment
measured in 1 paperPatel & Pavlick test whether LMs' internal representations of color, cardinal directions and grid terms carry the same relational structure as external grounded spaces, via few-shot prompting with no probe training [patel-pavlick-2022-mapping-language-models-to-grounded-conceptual-spaces] GPT-3 (175B) performs similarly on the true grounding and a structure-preserving rotation (spatial 45%/76% vs 44%/75% Top-1/Top-3) but collapses under random reassignment (16-19%) [patel-pavlick-2022-mapping-language-models-to-grounded-conceptual-spaces] Color-grounding error drops from 328 (GPT-2 124M) to 96 (GPT-3 175B), and BERT-base badly underperforms GPT-3 across all conditions [patel-pavlick-2022-mapping-language-models-to-grounded-conceptual-spaces] This is a scale-dependent relational-isomorphism-up-to-rotation claim inferred behaviorally; no activation-space intervention is performed [patel-pavlick-2022-mapping-language-models-to-grounded-conceptual-spaces]