Word senses decompose into ~2000 sparse discourse atoms
measured in 1 paperArora et al. apply classical sparse dictionary learning to their own 300-dimensional embeddings trained with the SN (squared-norm) objective on a 3-billion-token Wikipedia corpus, not pretrained word2vec/GloVe [arora-etal-2018-linear-algebraic-structure-word-senses] They solve for an overcomplete basis of about 2,000 "atoms of discourse" such that each word vector is a sparse (k~5) linear combination of a few atoms [arora-etal-2018-linear-algebraic-structure-word-senses] Polysemous words (e.g. "tie") decompose into atoms matching their distinct senses (clothing, sports, wiring, music), a measured linear-superposition account of polysemy [arora-etal-2018-linear-algebraic-structure-word-senses] No causal intervention is performed; the atom-sense correspondence is validated qualitatively by nearest-word inspection [arora-etal-2018-linear-algebraic-structure-word-senses]