mBERT's shared syntax metric needs a per-language transform only for distant languages
measured in 1 paperLimisiewicz & Marecek apply an orthogonal structural probe to mBERT across 9 typologically diverse languages, recovering syntactic (UD dependency distance, layer 7) and lexical (WordNet hypernymy depth, layer 5) tree-metric geometry [limisiewicz-marecek-2021-orthogonal-structural-probes] A single shared probe works nearly as well as per-language probes for Indo-European languages (correlation drop of only -0.027 to -0.048) [limisiewicz-marecek-2021-orthogonal-structural-probes] It drops sharply for non-Indo-European languages (e.g. -0.305 for Arabic lexical depth), so the shared metric needs a language-specific orthogonal transform for typologically distant languages [limisiewicz-marecek-2021-orthogonal-structural-probes] The shared probe also improves zero-shot cross-lingual parsing (ALLLANGS Chinese UAS 52.92 at zero training examples) [limisiewicz-marecek-2021-orthogonal-structural-probes]