Subtracting a per-language centroid removes decodable language identity in mBERT
measured in 1 paperLibovicky, Rosa & Fraser compute a per-language centroid as the mean mBERT representation over ~110k Wikipedia sentences per language [libovicky-rosa-fraser-2019] Subtracting the centroid collapses language-ID classification from 0.919 to 0.285 while improving cross-lingual retrieval (0.776 to 0.838) and leaving word-alignment F1 unchanged [libovicky-rosa-fraser-2019] A supervised linear projection does better on retrieval, but the parameter-free centroid alone already captures most of what a per-language offset can fix [libovicky-rosa-fraser-2019] The centroids are non-arbitrary: hierarchical clustering recovers genealogical language families at V-measure 82.4 versus a 62.1 random baseline [libovicky-rosa-fraser-2019] Centering does not help machine-translation quality estimation (HTER correlation below 0.04), a clear negative result [libovicky-rosa-fraser-2019]