MATH · IN · MODELS

mBERT's shared syntax metric needs a per-language transform only for distant languages

measured in 1 paper

Limisiewicz & Marecek apply an orthogonal structural probe to mBERT across 9 typologically diverse languages, recovering syntactic (UD dependency distance, layer 7) and lexical (WordNet hypernymy depth, layer 5) tree-metric geometry [limisiewicz-marecek-2021-orthogonal-structural-probes] A single shared probe works nearly as well as per-language probes for Indo-European languages (correlation drop of only -0.027 to -0.048) [limisiewicz-marecek-2021-orthogonal-structural-probes] It drops sharply for non-Indo-European languages (e.g. -0.305 for Arabic lexical depth), so the shared metric needs a language-specific orthogonal transform for typologically distant languages [limisiewicz-marecek-2021-orthogonal-structural-probes] The shared probe also improves zero-shot cross-lingual parsing (ALLLANGS Chinese UAS 52.92 at zero training examples) [limisiewicz-marecek-2021-orthogonal-structural-probes]

Context

orthogonal structural probe, cross-lingual tree-metric sharing, typological distance

Papers

Examining Cross-lingual Contextual Embeddings with Orthogonal Structural Probes — Limisiewicz, Tomasz, Mareček, David2021 · arXiv:2109.04921