Language units are script-conditioned; typology decodes with depth
measured in 1 paperVerma et al. use LAPE on MLP neurons and SAE-LAPE on latent features in Llama-3.2-1B and Gemma-2-2B to ask whether language units encode identity or surface script [verma-etal-2026-multilingual-lms-encode-script-over-linguistic-structure] Romanizing a non-Latin language produces a unit set nearly disjoint from both the native-script and English sets (Jaccard <0.3), a script-conditioned third subspace, while word-order shuffling leaves most raw-neuron units intact [verma-etal-2026-multilingual-lms-encode-script-over-linguistic-structure] Linear probing against lang2vec shows script-invariant units carry the strongest typological signal, and typological accessibility is depth-dependent (genealogy early, phonology deepest) [verma-etal-2026-multilingual-lms-encode-script-over-linguistic-structure] Causal ablation shows perplexity is most disrupted when script-invariant or order-invariant units are ablated, so functional necessity tracks surface-invariance rather than typological alignment [verma-etal-2026-multilingual-lms-encode-script-over-linguistic-structure]