MATH · IN · MODELS

NTP implicitly SVD-factorizes co-occurrence; orthants recover semantic categories

measured in 1 paper

Zhao & Thrampoulidis show next-token-prediction implicitly performs an SVD of a centered context-by-next-token co-occurrence matrix [zhao-thrampoulidis-2025-geometry-of-semantics-in-ntp] Clustering GPT-2 embeddings by which orthant (sign pattern) of the singular-vector basis they occupy recovers grammatical categories, entity types and topics unsupervised [zhao-thrampoulidis-2025-geometry-of-semantics-in-ntp] The recovered category structure becomes progressively finer-grained as more singular components are included [zhao-thrampoulidis-2025-geometry-of-semantics-in-ntp] It is validated on synthetic co-occurrence data plus TinyStories and WikiText-2, on GPT-2 (variant unspecified); purely observational with no causal steering [zhao-thrampoulidis-2025-geometry-of-semantics-in-ntp]

Context

SVD factorization of next-token prediction, orthant-based clustering, implicit semantic organization

Confirmed in models

Papers

Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations — Zhao, Yize, Thrampoulidis, Christos2025 · arXiv:2505.08348