MATH · IN · MODELS

Anisotropy from frequency-blind sampling and self-reinforcing tangent gradients

measured in 1 paper

Bernas et al. give a two-part geometric account of anisotropy: they prove (Corollary 2.2) that as high-frequency tokens concentrate near their centroid, local-manifold curvature becomes statistically invisible [bernas-etal-2026] They prove (Proposition 2.3) the normal-to-tangent gradient magnitude ratio scales as O(t) in the concentration radius, so curvature-bearing normal updates are suppressed for high-frequency tokens, compounded by attention and residuals [bernas-etal-2026] Across the Pythia (160m/410m/1b/1.4b), SmolLM2 (360m/1.7b) and EuroBERT (210m/610m) suites, true gradients concentrate in an activation-derived tangent subspace far more than matched-rank random controls (energy ratios orders of magnitude above null) [bernas-etal-2026] The effect is strongest early in training and in early/middle layers, and weaker in EuroBERT, whose language-balanced training counteracts the frequency skew the mechanism depends on [bernas-etal-2026] The paper reframes anisotropy as a possibly adaptive implicit dimensionality reduction rather than a pure training pathology [bernas-etal-2026]

Context

anisotropy, manifold curvature, stratified manifold, gradient dynamics, frequency effects, IsoScore, mechanistic interpretability during training

Papers

Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics — Bernas, Raphael, Jourdan, Fanny, Poché, Antonin, Hudelot, Céline2026 · arXiv:2604.08764