MATH · IN · MODELS

Contextualized embeddings occupy a narrow anisotropic cone

measured in 1 paper

- Across BERT, ELMo and GPT-2, contextualized representations occupy a narrow anisotropic cone: two random words have high average cosine similarity, near 1.0 in GPT-2's last layer. [ethayarajh-2019] - Anisotropy increases in upper layers together with context-specificity, and self-similarity falls monotonically with depth in all three models. [ethayarajh-2019] - Fewer than 5% of a word's contextual variance is captured by a single static embedding (anisotropy-adjusted MEV). [ethayarajh-2019] - Tested on BERT-base-cased, the original 2-layer biLSTM ELMo, and GPT-2 small; analysis is observational. [ethayarajh-2019]

Structure

Context

directional distribution, context-specificity, word sense, principal component

Papers

How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings — Ethayarajh, Kawin2019 · arXiv:1909.00512