MATH · IN · MODELS

Scale determines whether late-layer geometry stays organized by the readout

measured in 1 paper

Xu introduces Subspace PGA, a z-scored RSA metric comparing how well a layer cosine-distance structure survives projection onto the unembedding matrix top-k readout subspace versus 100 random subspaces [xu-2026-scale-determines-representation-geometry-organization-prediction] Across Pythia (70M-6.9B) and OLMo-1B/Phi-1.5/Gemma-2-2B, intermediate geometry is robustly organized around the readout subspace (peak z 9-24 at mid layers) [xu-2026-scale-determines-representation-geometry-organization-prediction] Small models (hidden dim <=1024) progressively lose this at late layers over training even as loss drops (Pythia-410M min z +0.6 to -32 across checkpoints), while dim>=2048 models preserve it [xu-2026-scale-determines-representation-geometry-organization-prediction] Removing a few top principal components restores positive z at every layer for dim>=768 models, supporting a masking interpretation; the paper frames the claim as correlational [xu-2026-scale-determines-representation-geometry-organization-prediction]

Context

a z-scored RSA variant testing subspace-specific (not generic low-rank) organization of representation geometry, scale-dependent masking of prediction-relevant geometry at late layers during training, dissociated from loss, a principal-component-removal intervention partially restoring the masked geometric signal

Papers

Scale Determines Whether Language Models Organize Representation Geometry for Prediction — Xu, Weilun2026 · arXiv:2605.17084