Contextualized embeddings occupy a narrow anisotropic cone
measured in 1 paper- Across BERT, ELMo and GPT-2, contextualized representations occupy a narrow anisotropic cone: two random words have high average cosine similarity, near 1.0 in GPT-2's last layer. [ethayarajh-2019] - Anisotropy increases in upper layers together with context-specificity, and self-similarity falls monotonically with depth in all three models. [ethayarajh-2019] - Fewer than 5% of a word's contextual variance is captured by a single static embedding (anisotropy-adjusted MEV). [ethayarajh-2019] - Tested on BERT-base-cased, the original 2-layer biLSTM ELMo, and GPT-2 small; analysis is observational. [ethayarajh-2019]
Structure
Context
directional distribution, context-specificity, word sense, principal component
Confirmed in models
Papers
How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings — Ethayarajh, Kawin