MATH · IN · MODELS
methods / Theoretical / Analytical / Geometric analysis

Geometric analysis

Techniqueadvanced

Studies the geometric relationships — angles, isomorphisms, reflections, geodesic distances — between feature directions that have already been found, rather than predicting a shape in advance.

Used in (77 observations)

structure: Linear Direction · models: 1-layer, 1-head causal Transformer (d_model=128, synthetic entity-relation analogy task) · paper: Emergent Analogical Reasoning in Transformers
structure: Score-Jacobian pullback Riemannian metric · models: Stable Diffusion v1.5, Custom diffusion model (MNIST, T=1000 steps) · paper: Image Interpolation with Score-based Riemannian Metrics of Diffusion Models
structure: Circle, 1D continuum manifold · models: Llama-3.1-8B · paper: Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
structure: Linear Separability · models: DINOv3 ViT-B/16, DINOv2 ViT-B/14, DINO ViT-B/16, MAE ViT-Base (Masked Autoencoder), ViT-Base, ConvNeXt (image classifier, various sizes, ImageNet) · paper: Human-like Object Grouping in Self-Supervised Vision Transformers
structure: Linear Subspace · models: AlphaZero (Hex, ResNet policy/value network, self-play) · paper: Evaluation Beyond Task Performance: Analyzing Concepts in AlphaZero in Hex
structure: Sphere · models: ArcFace ResNet50 (trained on CASIA/VGG2), ArcFace ResNet100 (trained on MS1MV3/IBUG-500K) · paper: ArcFace: Additive Angular Margin Loss for Deep Face Recognition
structure: Attention-graph spectral profile (Fiedler value, HFER, smoothness, spectral entropy) · models: Llama-3.2-1B, Llama-3.2-3B, Llama-3.1-8B, Qwen2.5-0.5B, Qwen-2.5-7B, Phi-3.5-mini, Mistral-7B-v0.1 · paper: Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
structure: Circle · models: Autoencoder (trained on J.S. Bach's Well-Tempered Clavier) · paper: The Circle of Fifths as Latent Geometry in Bach's Well-Tempered Clavier
structure: Linear Subspace, Linear Direction · models: Bayesian Wind Tunnel Transformer (bijection task, 6 layers, 6 heads, d_model=192), Bayesian Wind Tunnel Transformer (HMM filtering task, 9 layers, 8 heads, d_model=256) · paper: The Bayesian Geometry of Transformer Attention
structure: Conceptual Belief Space Hypothesis, Linear Subspace · models: Llama-3.1-8B-Instruct · paper: Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
structure: Polytope (Simplex) · models: Gemma-2B, Llama-3-8B, Qwen3-4B, Mistral-7B-v0.3 · paper: The Geometry of Categorical and Hierarchical Concepts in Large Language Models, When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs
structure: Anisotropy · models: Wan2.1-1.3B, Wan2.2-5B, CogVideoX-5B · paper: Steering Video Diffusion Transformers with Massive Activations
structure: Anisotropy, Linear Direction · models: CLIP ViT-B/32, CLIP ViT-B/16 · paper: Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
structure: Anisotropy, Linear Direction · models: CLIP ViT-B/32, CLIP ViT-L/14 · paper: The Double-Ellipsoid Geometry of CLIP
structure: Linear Subspace, Anisotropy · models: CLAP (HTSAT-BERT-ZS, trained on WavCaps), DRCap's CLAP (trained on WavCaps + SoundVECaps) · paper: COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings
structure: Linear Direction · models: Llama-3.2-1B, Llama-3.2-3B, Llama-3.1-8B, Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-3B, Qwen-2.5-7B, Qwen2.5-14B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Gemma-2-2B, Gemma-2-9B, Mistral-7B-v0.3, SmolLM2-1.7B, OLMoE-1B-7B, JetMoE-8B · paper: Concepts Whisper While Syntax Shouts: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations
structure: Intrinsic-dimension profile across depth · models: Gemma-2-27B · paper: Context Structure Reshapes the Representational Geometry of Language Models
structure: Linear Direction · models: Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-3B, Qwen-2.5-7B, Qwen2.5-14B, Qwen2.5-32B, Llama-2-7B, Llama-2-13B, Llama-3-8B, Llama-3.1-8B, Mistral Small 3 (2501), Ministral-3-3B-Base-2512, Ministral-3-8B-Base-2512, Ministral-3-14B-Base-2512, Gemma-2-2B, Gemma-2-9B, DeepSeek-LLM-7B-Base, Pythia-410M, Pythia-1B, Pythia-1.4B, Pythia-2.8B, Pythia-6.9B, Pythia-12B · paper: Language Models Represent and Transform Concepts with Shared Geometry
structure: Linear Subspace · models: Qwen2.5-0.5B · paper: Can LLMs Learn to Map the World from Local Descriptions?
structure: Anisotropy · models: CLIP ViT-B/32 · paper: It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
structure: Linear Direction, Anisotropy · models: CLIP ViT-B/32, MetaCLIP ViT-B/32, OpenCLIP ViT-B/32, SigLIP 2 · paper: Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders
structure: Torus · models: Vanilla RNN (path-integration / spatial-localization task) · paper: Emergence of Grid-like Representations by Training Recurrent Neural Networks to Perform Spatial Localization
structure: Minkowski Sum of Tile Polytopes (Minkowski Representation Hypothesis) · models: DINOv2-B (ViT-Base, 4 register tokens) · paper: Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
structure: Polytope (Simplex) · models: Mistral-7B-Instruct-v0.3 · paper: Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
structure: Linear Separability · models: ESM-2 (35M), ESM-2 (150M), ESM-2 (650M), ESM-2 (3B), ESMC (600M) · paper: Protein Contacts Are Already in the Attention: A Single-Forward-Pass Alternative to the Categorical Jacobian
structure: Curvature profile of the representation manifold, Linear Subspace · models: Llama-3.2-1B · paper: The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
structure: Sphere · models: Qwen2.5-3B-Instruct, Llama-3.2-3B-Instruct, Gemma-2-2B-it · paper: Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling
structure: Intrinsic-dimension profile across depth · models: Gemma-2-2B, Gemma-2-9B · paper: The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws
structure: Linear Direction, Polytope (Simplex) · models: Gemma-2B, Llama-3-8B, Qwen3-4B, Mistral-7B-v0.3 · paper: The Geometry of Categorical and Hierarchical Concepts in Large Language Models, When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs
structure: Polytope (Simplex), Dimensional collapse · models: ResNet-50 (hard-negative supervised/unsupervised contrastive, CIFAR-100) · paper: Hard-Negative Sampling for Contrastive Learning: Optimal Representation Geometry and Neural- vs Dimensional-Collapse
structure: Anisotropy · models: Llama-3-8B-Instruct, Gemma-2-9B-it, Qwen2.5-7B-Instruct, Llama-2-7B-Chat, GPT-2-small, OPT-350M, OPT-2.7B · paper: Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
structure: Intrinsic-dimension profile across depth · models: GPT-2 XL, Pythia-2.8B, Custom 12-layer GPT-2-Small (trained from scratch, 100M tokens) · paper: Representational Curvature Modulates Behavioral Uncertainty in Large Language Models
structure: Linear Direction · models: T5-11B, UnifiedQA-11B (T5-based), T0++, GPT-J-6B, RoBERTa-large-MNLI, DeBERTa-xxlarge-v2-MNLI · paper: Discovering Latent Knowledge in Language Models Without Supervision
structure: Concept lattice (Formal Concept Analysis half-space model), Linear Direction · models: Gemma-2B, word2vec · paper: Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence
structure: Circle · models: pGESAM (pitch-conditioned Generative Sample Map), EnCodec · paper: Pitch-Conditioned Instrument Sound Synthesis from an Interactive Timbre Latent Space
structure: Anisotropy · models: LLaVA-1.5-7B · paper: Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models
structure: 1D continuum manifold · models: Pythia-2.8B, Llama-2-7B, Llama-3.1-8B, Llama-3.2-1B, GPT-2-Large, Mistral-7B, Llama-3.1-8B-Instruct, Llama-3.2-1B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-3B, Llama-3.2-3B-Instruct, Llama-3.2-3B, Gemma-2-2B-it, Gemma-2-2B, Llama-3.1-70B-Instruct, Llama-3-8B-Instruct, Llama-3-8B, Mistral-7B-Instruct-v0.3, Qwen2.5-7B-Instruct · paper: Number Representations in LLMs: A Computational Parallel to Human Perception, Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling, Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models, LLMs Know More About Numbers than They Can Say
structure: Linear Representation Hypothesis · models: Llama-2-7B, Gemma-2B · paper: The Linear Representation Hypothesis and the Geometry of Large Language Models
structure: Linear Direction · models: Gemma-2-2B, Llama-3.2-3B-Instruct, GPT-2-XL · paper: Universal Response and Emergence of Induction in LLMs
structure: Polytope (Simplex) · models: ResNet-18 (multi-label, MLab-CIFAR10) · paper: How Label Imbalance Shapes Geometry: A General Spectral Analysis of Multi-Label Neural Collapse
structure: Polytope (Simplex), Belief State Geometry Hypothesis (Mixed-State Presentation) · models: Bayesian Wind Tunnel Mamba (HMM filtering task, 9 layers, d_model=256, state dim 16) · paper: The Bayesian Geometry of Transformer Attention
structure: Circle, 1D continuum manifold · models: GPT-2-small, Mistral-7B, text-embedding-3-large · paper: The Origins of Representation Manifolds in Large Language Models
structure: Anisotropy · models: ViT-S/16 (12 layers, 8 heads, dim 384, trained from scratch on ImageNet-100) · paper: Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness
structure: Line Attractor · models: Qwen2.5-3B-Instruct · paper: Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
structure: Polytope (Simplex) · models: VGG (image classifier, various depths), ResNet (image classifier, various depths), DenseNet (image classifier, various depths) · paper: Prevalence of Neural Collapse during the terminal phase of deep learning training, Are Neurons Actually Collapsed? On the Fine-Grained Structure in Neural Representations
structure: Circle · models: NLLB-200 (distilled, 600M) · paper: Universal Conceptual Structure in Neural Translation: Probing NLLB-200's Multilingual Geometry
structure: Concept Crystals (Parallelogram/Trapezoid Structure) · models: NLLB-200 (distilled, 600M) · paper: Universal Conceptual Structure in Neural Translation: Probing NLLB-200's Multilingual Geometry
structure: Linear Subspace, Linear Direction · models: BERT-base-uncased, GPT-2 Small, Qwen-2.5-7B, Qwen2.5-Math-7B · paper: The Representational Geometry of Number
structure: Linear Direction · models: Gemma 2 27B Instruct, Qwen3-32B · paper: Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy
structure: Anisotropy · models: Llama-3.2-3B-Instruct, Llama-3.1-8B-Instruct, Ministral-8B-Instruct-2410, Qwen3-4B, Qwen3-14B · paper: Learning Uncertainty from Sequential Internal Dispersion in Large Language Models
structure: Linear Subspace · models: PhonSSM (AGAN + PDM + bidirectional Mamba SSM + HPC, skeleton-based ASL) · paper: State Space Models are Effective Sign Language Learners: Exploiting Phonological Compositionality for Vocabulary-Scale Recognition
structure: Circle · models: LSTM VAE (371 J.S. Bach chorales, 6 compared input encodings: piano roll, MIDI-like, ABC, Tonnetz, Pitch DFT, Pitch-Class DFT) · paper: Exploring Latent Spaces of Tonal Music using Variational Autoencoders
structure: Linear Subspace · models: CPC-big (LibriLight 6k hrs), CPC-small (LibriSpeech-100h), APC (LibriSpeech-360h) · paper: Self-supervised Predictive Coding Models Encode Speaker and Phonetic Information in Orthogonal Subspaces
structure: Linear Direction · models: Qwen3-4B, Llama-3.1-8B, GLM-4-9B · paper: Simulated Adoption: Decoupling Magnitude and Direction in LLM In-Context Conflict Resolution
structure: Intrinsic-dimension profile across depth · models: Qwen1.5-0.5B, Qwen2-0.5B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Llama-3-8B · paper: The Geometry of Reasoning: Flowing Logics in Representation Space
structure: Intrinsic-dimension profile across depth · models: Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-3B, Qwen-2.5-7B, Qwen2.5-14B, Qwen2.5-32B, Qwen2.5-72B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Qwen3-14B, Qwen3-32B, Gemma 3 1B, Gemma 3 4B, Gemma 3 12B, Gemma 3 27B, DeepSeek-R1-Distill-Qwen-1.5B, DeepSeek-R1-Distill-Qwen-14B, DeepSeek-R1-Distill-Qwen-32B · paper: Reasoning emerges from constrained inference manifolds in large language models
structure: Platonic Representation Hypothesis, Intrinsic-dimension profile across depth · models: scGPT (whole-human pretrained checkpoint) · paper: Discovery of a Hematopoietic Manifold in scGPT Yields a Method for Extracting Performant Algorithms from Biological Foundation Model Internals
structure: Continuous causal attention operator (CCT) · models: Llama-2-13B-Chat, Llama-3-8B, Phi-3-Medium-4k-Instruct, Gemma-7B, Gemma-2-9B, Mistral-7B, GPT-2 · paper: Language Models Are Implicitly Continuous
structure: Linear Direction, Linear Subspace · models: wav2vec 2.0 Large (LV-60), HuBERT Large (LibriLight 60k), WavLM Large · paper: Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
structure: Sphere · models: SphereFace (64-conv-layer residual CNN, trained on CASIA-WebFace) · paper: SphereFace: Deep Hypersphere Embedding for Face Recognition
structure: Polytope (Simplex) · models: Toy ReLU-output autoencoder (h=Wx, x'=ReLU(W^T h + b)) · paper: Toy Models of Superposition
structure: Linear Direction · models: Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, Phi-4-Mini-Instruct · paper: The Geometries of Truth Are Orthogonal Across Tasks
structure: Linear Direction · models: Mini-ICL synthetic RoPE transformer (6L/2H/d128 for dice and Markov-chain tasks; 16L for linear regression), Qwen-2.5-7B · paper: Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
structure: Anisotropy, Linear Subspace · models: CLIP ViT-B/32 · paper: Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment
structure: Anisotropy, Linear Direction · models: CLIP ViT-B/16 · paper: On the Modality Gap and the Contrastive Loss in Multi-modal Representation Learning
structure: Anisotropy · models: all-MiniLM-L6-v2, all-MiniLM-L12-v2, all-mpnet-base-v2, paraphrase-mpnet-base-v2, BGE-base-en-v1.5, BGE-large-en-v1.5, E5-large-v2, multilingual-e5-large, e5-mistral-7b-instruct, SFR-Embedding-Mistral, BERT-base-uncased, RoBERTa-base, ELECTRA-base, mBERT (BERT-base, Multilingual Cased), GPT-2-small, Pythia-410M, Qwen2.5-1.5B, Qwen2.5-7B, Mistral-7B · paper: Anisotropy Decides Cosine vs. Rank Metrics for Text Embeddings
structure: Intrinsic-dimension profile across depth, Linear Separability · models: Qwen2.5-0.5B-Instruct, Qwen2.5-1.5B, Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct, Qwen2.5-32B, Gemma-2-2B, Gemma-2-9B, Gemma-2-27B, Llama-3.2-1B, Llama-3.2-3B · paper: Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight LMs
structure: Intrinsic-dimension profile across depth · models: GPT-2 Medium, Llama-2-7B, Mistral-7B-v0.1, Llama-3.2-1B, Llama-3.2-3B · paper: Lines of Thought in Large Language Models
structure: Torus · models: CARNN trained on path integration across multiple (1-50) environments · paper: Coherently Remapping Toroidal Cells But Not Grid Cells are Responsible for Path Integration in Virtual Agents
structure: Jacobian non-normality (Schur decomposition of per-layer operators) · models: Llama-3.1-8B, Gemma 4 E4B, OLMo 3 7B · paper: Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
structure: Attention reference frame (sink-token anchor configuration) · models: Llama-3.1-8B, Llama-3.1-8B-Instruct, Llama-3.2-1B, Llama-3.2-3B, Mistral-7B-v0.1, Gemma-7B, Qwen2.5-3B, Qwen-2.5-7B, Qwen2.5-7B-Instruct, Phi-2, BERT-base-uncased, XLM-RoBERTa large, Pythia-1.4B, Pythia-2.8B, Pythia-6.9B, Pythia-12B · paper: What are you sinking? A geometric approach on attention sink
structure: Cone · models: Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct, Gemma-2-2B-it, Gemma-2-9B-it · paper: From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
structure: Linear Direction, Anisotropy · models: CLIP, SigLIP, SigLIP 2, AIMv2 · paper: Interpreting the Linear Structure of Vision-Language Model Embedding Spaces
structure: Anisotropy · models: GPT-OSS-20B, ERNIE-4.5-21B-A3B-Base, Qwen3-30B-A3B-Base, Ling-mini-Base, Trinity-Mini-Base · paper: The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise
structure: Polytope (Simplex) · models: Custom GPT-2-style causal transformer (205M, TinyStories, largest of a 3.4M-205M width/depth/epoch grid) · paper: Linguistic Collapse: Neural Collapse in (Large) Language Models
structure: Circle · models: WavLM Large · paper: Learning Arousal-Valence Representation from Categorical Emotion Labels of Speech
structure: Anisotropy · models: NanoGPT Causal-NoPE, 6-layer (10.6M params) · paper: Position Information Emerges in Causal Transformers Without Positional Encodings via Similarity of Nearby Embeddings