RAND-WALK derives analogy parallelograms from a latent discourse walk
measured in 1 paper- The RAND-WALK model posits a slowly drifting latent "discourse" vector emitting words with probability proportional to exp(<c, v_w>), yielding PMI(w,w') ≈ <v_w, v_w'>/d under an isotropy prior. [arora-etal-2016-rand-walk-pmi-word-embeddings] - Analogies emerge as approximately parallel difference vectors (RELATIONS=LINES / parallelograms), explaining why low-dimensional linear embeddings solve analogy tasks. [arora-etal-2016-rand-walk-pmi-word-embeddings] - The predictions are checked on the authors' own SGNS and GloVe vectors trained on English Wikipedia (March 2015, 68,430-word vocab, d=300): squared norm correlates with log-frequency (~0.75) and Google-analogy total accuracy is 0.70 (skip-gram), 0.73 (GloVe), 0.74 (CBOW). [arora-etal-2016-rand-walk-pmi-word-embeddings] - Analytical derivation with correlational consistency checks; the embeddings are trained by the authors on Wikipedia, not standard pretrained releases. [arora-etal-2016-rand-walk-pmi-word-embeddings]