MATH · IN · MODELS

Word senses decompose into ~2000 sparse discourse atoms

measured in 1 paper

Arora et al. apply classical sparse dictionary learning to their own 300-dimensional embeddings trained with the SN (squared-norm) objective on a 3-billion-token Wikipedia corpus, not pretrained word2vec/GloVe [arora-etal-2018-linear-algebraic-structure-word-senses] They solve for an overcomplete basis of about 2,000 "atoms of discourse" such that each word vector is a sparse (k~5) linear combination of a few atoms [arora-etal-2018-linear-algebraic-structure-word-senses] Polysemous words (e.g. "tie") decompose into atoms matching their distinct senses (clothing, sports, wiring, music), a measured linear-superposition account of polysemy [arora-etal-2018-linear-algebraic-structure-word-senses] No causal intervention is performed; the atom-sense correspondence is validated qualitatively by nearest-word inspection [arora-etal-2018-linear-algebraic-structure-word-senses]

Context

sparse superposition of word senses, discourse atoms, classical dictionary learning

Papers

Linear Algebraic Structure of Word Senses, with Applications to Polysemy — Arora, Sanjeev, Li, Yuanzhi, Liang, Yingyu, Ma, Tengyu, Risteski, Andrej2018 · arXiv:1601.03764