A diff-in-means language-identity direction steers output language without accuracy loss
measured in 1 paperSterz et al. extract a language-identity direction as the mean-difference between target-language and other-language activations in Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Gemma-2-2B-Instruct [sterz-etal-2025-recover-target-language] Adding this direction, or a trained steering module, to the residual stream raises the rate of responses in the correct target language across 18 languages [sterz-etal-2025-recover-target-language] Unlike a prior unsupervised language-vector steering baseline, it preserves downstream task accuracy such as MMLU [sterz-etal-2025-recover-target-language] No further internal geometric characterization (PCA/subspace structure) of the direction is reported [sterz-etal-2025-recover-target-language]
Structure
Context
language-identity direction, language confusion, steering without task-performance loss
Confirmed in models
Papers
ReCoVeR the Target Language: Language Steering Without Sacrificing Task Performance — Sterz, Hannah, Schmidt, Fabian David, Glavaš, Goran, Vulić, Ivan