Neural-Collapse signatures emerge in causal language models, but not self-duality
measured in 1 paperWu & Papyan train about 30 GPT-Neo-based causal transformers from scratch on TinyStories (GPT-2 used only for its tokenizer), reframing next-token prediction as extreme imbalanced classification and measuring Neural Collapse [wu-papyan-2024-linguistic-collapse] Within-class variability collapse (NC1), the GNC2 hyperspherical-uniformity variant, uniform duality (UNC3), and NC4 strengthen with model width, depth, and training [wu-papyan-2024-linguistic-collapse] Self-duality (NC3) does NOT develop with scale, and a true Simplex ETF is unreachable because the number of classes far exceeds the hidden dimension (C >> d+1), so the geometry only tends toward it [wu-papyan-2024-linguistic-collapse] This is the first systematic demonstration of Neural-Collapse-style geometry in real, generatively pretrained causal language models [wu-papyan-2024-linguistic-collapse]