MATH · IN · MODELS

Nonlinear intrinsic dimension phase-transitions with emergent zero-shot competence

measured in 1 paper

Lee et al. track per-layer nonlinear intrinsic dimension (TwoNN) and linear PCA effective dimension across pretraining of Pythia-410M/1.4B/6.9B [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] On synthetic data both measures scale with dataset compositionality in fully-trained models, but nonlinear ID shows a sharp phase transition around step ~10^3 coinciding with the onset of zero-shot competence [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] PCA effective dimension tracks superficial/Kolmogorov complexity (gzip compressibility, early-training Spearman rho near 1.0) while TwoNN ID never correlates with gzip [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime] The linear-superficial versus nonlinear-semantic dissociation generalizes to fully-trained Llama-3-8B and Mistral-7B [lee-etal-2024-geometric-signatures-of-compositionality-across-a-language-models-lifetime]

Context

a training-dynamics phase transition in nonlinear intrinsic dimension coinciding with emergent zero-shot competence, dissociated from a linear PCA-dimension measure that tracks only superficial data complexity

Papers

Geometric Signatures of Compositionality Across a Language Model's Lifetime — Lee, Jin Hwa, Jiralerspong, Thomas, Yu, Lei, Bengio, Yoshua, Cheng, Emily2024 · arXiv:2410.01444