MATH · IN · MODELS

Language identity is linearly separable from layer 1; alignment tracks pretraining mix

measured in 1 paper

Kim & Lee probe all layers of six multilingual LLMs across five XNLI languages, training linear and MLP probes to classify language identity [kim-lee-2026-language-directions-token-geometry-multilingual-llms] Language separability jumps sharply in the first transformer block (+76.4 points) and stays almost fully linearly separable throughout depth (linear-probe accuracy 99.8%, only 0.58 points below the MLP) [kim-lee-2026-language-directions-token-geometry-multilingual-llms] A Token-Language Alignment metric tracks pretraining composition: Chinese-inclusive models reach 16.43% Chinese Match@Peak versus 3.90% for English-centric ones (4.21x) [kim-lee-2026-language-directions-token-geometry-multilingual-llms] Latin-script languages show uniformly low alignment regardless of model, confounded by shared script; no causal intervention is performed [kim-lee-2026-language-directions-token-geometry-multilingual-llms]

Context

Token-Language Alignment (cosine similarity between probe-learned language direction and vocabulary/unembedding embeddings), Match@Peak metric quantifying structural imprinting from pretraining language composition, near-universal linear separability of language identity emerging in the first transformer block, 4.21x Chinese-alignment ratio between Chinese-inclusive and English-centric pretraining mixtures

Papers

How Language Directions Align with Token Geometry in Multilingual LLMs — Kim, JaeSeong, Lee, Suan2026 · arXiv:2511.16693