MATH · IN · MODELS

Neural-Collapse signatures emerge in causal language models, but not self-duality

measured in 1 paper

Wu & Papyan train about 30 GPT-Neo-based causal transformers from scratch on TinyStories (GPT-2 used only for its tokenizer), reframing next-token prediction as extreme imbalanced classification and measuring Neural Collapse [wu-papyan-2024-linguistic-collapse] Within-class variability collapse (NC1), the GNC2 hyperspherical-uniformity variant, uniform duality (UNC3), and NC4 strengthen with model width, depth, and training [wu-papyan-2024-linguistic-collapse] Self-duality (NC3) does NOT develop with scale, and a true Simplex ETF is unreachable because the number of classes far exceeds the hidden dimension (C >> d+1), so the geometry only tends toward it [wu-papyan-2024-linguistic-collapse] This is the first systematic demonstration of Neural-Collapse-style geometry in real, generatively pretrained causal language models [wu-papyan-2024-linguistic-collapse]

Context

Neural Collapse in language models, next-token prediction as imbalanced classification, Simplex Equiangular Tight Frame

Papers

Linguistic Collapse: Neural Collapse in (Large) Language Models — Wu, Robert, Papyan, Vardan2024 · arXiv:2405.17767