MATH · IN · MODELS

Shared SAE morphosyntactic directions are causally necessary and sufficient across languages

measured in 1 paper

Brinkmann et al. train a Gated SAE on layer-16 residual activations of Llama-3-8B and Aya-23-8B and use attribution patching to find each language/concept's top causal features for grammatical number, gender, and tense [brinkmann-etal-2025-crosslingual-grammatical-concepts] Cross-lingual top-feature overlap reaches up to 50% (one feature is top-influential for grammatical gender across all 15 inflecting languages), with mean cross-concept overlap 13.9% [brinkmann-etal-2025-crosslingual-grammatical-concepts] Ablating only the massively-multilingual features drops classifier performance to 64%, so most of the causal effect concentrates in a small multilingual core [brinkmann-etal-2025-crosslingual-grammatical-concepts] Clamping a single multilingual feature during translation flips the intervened concept's probe label while leaving others unaffected, showing necessity and sufficiency [brinkmann-etal-2025-crosslingual-grammatical-concepts]

Context

Gated SAE trained on residual stream (layer 16), attribution-patching-selected top-32 causal features per language/concept, cross-lingual top-feature overlap (up to 50%, mean 13.9% cross-concept), massively-multilingual vs. all-multilingual vs. monolingual feature ablation, single-feature clamping steering with efficacy/selectivity metrics

Papers

Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages — Brinkmann, Jannik, Wendler, Chris, Bartelt, Christian, Mueller, Aaron2025 · arXiv:2501.06346