In-context concept representations converge toward a human-aligned structure
measured in 1 paperXu et al. probe LLM (mainly LLaMA3-70B) concepts via an in-context reverse-dictionary task, characterizing each context's representation by its pairwise-similarity matrix [xu-etal-2025-emergent-conceptual-representations] RSA alignment across contexts rises from 0.800 at 1 demonstration to 0.970 at 24, converging toward a single context-independent relational structure [xu-etal-2025-emergent-conceptual-representations] Alignment with the 120-demonstration structure correlates with reverse-dictionary accuracy at rho=0.976, and cross-model alignment across 67 LLMs predicts task performance (rho=0.870), all correlational with no intervention [xu-etal-2025-emergent-conceptual-representations] The convergent structure aligns via RSA with human similarity judgments (SimLex-999 rho=0.776), THINGS odd-one-out, and voxel-wise fMRI activity in LOC/FFA/PPA and other regions [xu-etal-2025-emergent-conceptual-representations]