MATH · IN · MODELS

Language units are script-conditioned; typology decodes with depth

measured in 1 paper

Verma et al. use LAPE on MLP neurons and SAE-LAPE on latent features in Llama-3.2-1B and Gemma-2-2B to ask whether language units encode identity or surface script [verma-etal-2026-multilingual-lms-encode-script-over-linguistic-structure] Romanizing a non-Latin language produces a unit set nearly disjoint from both the native-script and English sets (Jaccard <0.3), a script-conditioned third subspace, while word-order shuffling leaves most raw-neuron units intact [verma-etal-2026-multilingual-lms-encode-script-over-linguistic-structure] Linear probing against lang2vec shows script-invariant units carry the strongest typological signal, and typological accessibility is depth-dependent (genealogy early, phonology deepest) [verma-etal-2026-multilingual-lms-encode-script-over-linguistic-structure] Causal ablation shows perplexity is most disrupted when script-invariant or order-invariant units are ablated, so functional necessity tracks surface-invariance rather than typological alignment [verma-etal-2026-multilingual-lms-encode-script-over-linguistic-structure]

Context

language-associated units defined by low per-unit activation-entropy across languages (LAPE / SAE-LAPE), romanization induces near-disjoint, script-conditioned unit sets ("capacity fragmentation"), word-order shuffling leaves most raw-neuron units intact but destabilizes some SAE features, typological structure most linearly decodable in script-invariant units, and increasingly so with depth, causal ablation/mean-replacement shows functional necessity tracks surface-invariance, not typological alignment

Papers

Multilingual Language Models Encode Script Over Linguistic Structure — Verma, Aastha A K, Chatterjee, Anwoy, Gupta, Mehak, Chakraborty, Tanmoy2026 · arXiv:2604.05090