MATH · IN · MODELS

Linguistic subspaces emerge then specialize in critical pretraining phases

measured in 1 paper

Muller-Eberstein et al. fit an information-theoretic linear probing suite for 9 syntax/semantics/reasoning tasks at every layer of MultiBERTs across 2M pretraining steps and 5 random seeds (0-4) [muller-eberstein-etal-2023-subspace-chronicles] They compare the fitted probes subspaces directly via a subspace-similarity metric rather than only per-task accuracy [muller-eberstein-etal-2023-subspace-chronicles] Task-specific subspaces first emerge and share information broadly, then shift and specialize into distinct subspaces during identifiable critical learning phases [muller-eberstein-etal-2023-subspace-chronicles] Syntax subspaces form within 0.5% of total training while semantic and reasoning subspaces specialize much later, giving a developmental trajectory [muller-eberstein-etal-2023-subspace-chronicles]

Context

training dynamics, critical learning phases, subspace disentanglement, developmental linguistics, information-theoretic probing

Papers

Subspace Chronicles: How Linguistic Information Emerges, Shifts and Interacts during Language Model Training — Müller-Eberstein, Max, van der Goot, Rob, Plank, Barbara, Titov, Ivan2023 · arXiv:2310.16484