MATH · IN · MODELS

After mean-centering, 88 languages share one XLM-R subspace

measured in 1 paper

Chang, Tu & Bergen fit a per-language affine subspace (mean plus top singular directions, median rank 335/768) to XLM-R-base for 88 languages [chang-tu-bergen-2022] Using a Riemannian covariance-distance on positive-definite matrices, layers 6-11 place any two languages' subspaces within a ~5 degree rotation or 1.6x scaling of each other [chang-tu-bergen-2022] Projecting a representation onto its own language's subspace barely raises perplexity, onto another's raises it substantially, but onto another's shifted to the first's mean is only moderately worse, so once mean-corrected the subspaces are largely interchangeable [chang-tu-bergen-2022] Shifting a representation by the difference of two language means induces target-language token predictions 4.7x more often (10%->47%) while dropping source-language ones 3.5x (75%->21%) [chang-tu-bergen-2022]

Context

affine subspace, Riemannian metric on positive-definite matrices, causal projection intervention, mean-centering

Papers

The Geometry of Multilingual Language Model Representations — Chang, Tyler A., Tu, Zhuowen, Bergen, Benjamin K.2022 · arXiv:2205.10964