Singular Value Decomposition — factors an activation matrix directly as $X = U\Sigma V^\top$ and inspects the singular-value spectrum or a low-rank ($U\Sigma$) projection, without necessarily centering the data first — the same linear-algebra machinery PCA relies on, but reported directly in terms of singular values/vectors rather than an eigendecomposition of a centered covariance matrix.
Used in (9 observations)
structure: Dimensional collapse · models: Llama-3.1-8B, GPT-2, Gemma-2-9B, Qwen3-8B, Qwen3-4B, Pythia-160M, Pythia-2.8B · paper: Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
structure: Dimensional collapse · models: ResNet-50 (SimCLR contrastive pretraining, ImageNet) · paper: Understanding Dimensional Collapse in Contrastive Self-supervised Learning
structure: Anisotropy · models: Tied-input-output-embedding NMT + LM (WMT'14 / WikiText-2) · paper: Representation Degeneration Problem in Training Natural Language Generation Models
structure: Dimensional collapse · models: GCN depth ensemble, 2-24 layers (Zhang, Higham, Deidda & Tudisco), GAT depth ensemble, 2-24 layers (Zhang, Higham, Deidda & Tudisco) · paper: Are We Measuring Oversmoothing in Graph Neural Networks Correctly?
structure: Dimensional collapse · models: OLMo-1B, OLMo-7B, OLMo 2 1B, OLMo 2 7B, Pythia-160M, Pythia-410M, Pythia-1B, Pythia-1.4B, Pythia-2.8B, Pythia-6.9B, Pythia-12B, Llama-3.1-8B, Llama-3.1-Tulu-3-8B-SFT, Llama-3.1-Tulu-3-8B-DPO · paper: Tracing the Representation Geometry of Language Models from Pretraining to Post-training
structure: Anisotropy · models: BERT-base-uncased, RoBERTa-base, ALBERT-base-v1, GPT-2-small, GPT-J-6B, OPT-13B, Llama-2-7B, Llama-2-7B-Chat, BLOOM-560M, BLOOM-3B, Pythia-2.8B, Falcon-7B, Falcon-7B-Instruct · paper: The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
structure: Intrinsic-dimension profile across depth · models: GPT-2 Medium, Llama-2-7B, Mistral-7B-v0.1, Llama-3.2-1B, Llama-3.2-3B · paper: Lines of Thought in Large Language Models
structure: Anisotropy · models: GPT-OSS-20B, ERNIE-4.5-21B-A3B-Base, Qwen3-30B-A3B-Base, Ling-mini-Base, Trinity-Mini-Base · paper: The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise
structure: Linear Subspace · models: Qwen2.5-VL 7B Instruct, InternVL3-8B · paper: Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models