MATH · IN · MODELS

In-context concept representations converge toward a human-aligned structure

measured in 1 paper

Xu et al. probe LLM (mainly LLaMA3-70B) concepts via an in-context reverse-dictionary task, characterizing each context's representation by its pairwise-similarity matrix [xu-etal-2025-emergent-conceptual-representations] RSA alignment across contexts rises from 0.800 at 1 demonstration to 0.970 at 24, converging toward a single context-independent relational structure [xu-etal-2025-emergent-conceptual-representations] Alignment with the 120-demonstration structure correlates with reverse-dictionary accuracy at rho=0.976, and cross-model alignment across 67 LLMs predicts task performance (rho=0.870), all correlational with no intervention [xu-etal-2025-emergent-conceptual-representations] The convergent structure aligns via RSA with human similarity judgments (SimLex-999 rho=0.776), THINGS odd-one-out, and voxel-wise fMRI activity in LOC/FFA/PPA and other regions [xu-etal-2025-emergent-conceptual-representations]

Context

in-context reverse-dictionary task, RSA convergence across number of demonstrations, cross-model alignment predicts task performance, RSA against human similarity judgments (SimLex-999, THINGS), voxel-wise fMRI encoding-model validation, no causal intervention

Papers

Revealing Emergent Human-like Conceptual Representations from Language Prediction — Xu, Ningyu, Zhang, Qi, Du, Chao, Luo, Qiang, Qiu, Xipeng, Huang, Xuanjing, Zhang, Menghan2025 · arXiv:2501.12547