MATH · IN · MODELS

A diff-in-means language-identity direction steers output language without accuracy loss

measured in 1 paper

Sterz et al. extract a language-identity direction as the mean-difference between target-language and other-language activations in Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Gemma-2-2B-Instruct [sterz-etal-2025-recover-target-language] Adding this direction, or a trained steering module, to the residual stream raises the rate of responses in the correct target language across 18 languages [sterz-etal-2025-recover-target-language] Unlike a prior unsupervised language-vector steering baseline, it preserves downstream task accuracy such as MMLU [sterz-etal-2025-recover-target-language] No further internal geometric characterization (PCA/subspace structure) of the direction is reported [sterz-etal-2025-recover-target-language]

Context

language-identity direction, language confusion, steering without task-performance loss

Papers

ReCoVeR the Target Language: Language Steering Without Sacrificing Task Performance — Sterz, Hannah, Schmidt, Fabian David, Glavaš, Goran, Vulić, Ivan2025 · arXiv:2509.14814