A language difference-of-means direction causally steers target-language output
measured in 1 paperChang, Tu & Bergen identify language-sensitive axes via linear discriminant analysis and take the raw difference between two languages' mean representations [chang-tu-bergen-2022] Shifting a representation by this direction and projecting onto language B's subspace produces 4.7x more predicted tokens in target language B (10% to 47%) and 3.5x fewer in language A (75% to 21%) [chang-tu-bergen-2022] The unmodified model otherwise predicts the original evaluation language 99.5% of the time [chang-tu-bergen-2022] A separate visual claim that language-neutral axes encode token position on spiral/torus curves is not independently quantified and is not catalogued here [chang-tu-bergen-2022]
Structure
Context
difference-in-means direction, causal steering, linear discriminant analysis
Confirmed in models
Papers
The Geometry of Multilingual Language Model Representations — Chang, Tyler A., Tu, Zhuowen, Bergen, Benjamin K.