Derives the expected representation geometry from first principles — proofs, self-consistent quantization conditions — ahead of and independent from empirical observation.
Used in (16 observations)
structure: Concept Crystals (Parallelogram/Trapezoid Structure) · models: SGNS (Wikipedia), GloVe (Wikipedia + Gigaword, uncased) · paper: A Latent Variable Model Approach to PMI-based Word Embeddings
structure: Linear Subspace, Linear Direction · models: Bayesian Wind Tunnel Transformer (bijection task, 6 layers, 6 heads, d_model=192), Bayesian Wind Tunnel Transformer (HMM filtering task, 9 layers, 8 heads, d_model=256) · paper: The Bayesian Geometry of Transformer Attention
structure: Linear Direction · models: β-VAE (convolutional encoder/decoder, dSprites/colored-dSprites/CelebA/3D-Chairs) · paper: Understanding disentangling in β-VAE
structure: Circle, Constructive Interference Hypothesis · models: BOWS Autoencoder (tied-weight, ReLU), BOWS Autoencoder (tied-weight, linear — no ReLU, baseline), BOWS Toy Transformer (1 block, 8 heads, d_model=768) · paper: From Data Statistics to Feature Geometry: How Correlations Shape Superposition
structure: Torus · models: Conformal-normalization linear path-integration RNN, Conformal-normalization nonlinear (CANN-like) path-integration RNN · paper: Emergence of Grid-like Representations by Training Recurrent Networks with Conformal Normalization
structure: Polytope (Simplex) · models: Mistral-7B-Instruct-v0.3 · paper: Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
structure: Concept Crystals (Parallelogram/Trapezoid Structure) · models: SGNS (Wikipedia) · paper: Towards Understanding Linear Word Analogies
structure: Affine Subspace · models: Llama-2-7B, Llama-2-13B, Llama-2-70B, Pythia-6.9B, Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia) · paper: Language Models Represent Space and Time, Symmetry in Language Statistics Shapes the Geometry of Model Representations
structure: 1D continuum manifold, Linear Subspace · models: Llama-2-7B · paper: Probing for Representation Manifolds in Superposition
structure: Polytope (Simplex) · models: VGG (image classifier, various depths), ResNet (image classifier, various depths), DenseNet (image classifier, various depths) · paper: Prevalence of Neural Collapse during the terminal phase of deep learning training, Are Neurons Actually Collapsed? On the Fine-Grained Structure in Neural Representations
structure: Concept Crystals (Parallelogram/Trapezoid Structure) · models: · paper: On the Emergence of Linear Analogies in Word Embeddings
structure: Linear Subspace · models: FastText (bag-of-word-vectors), fastText (Common Crawl, 300d), word2vec (CoNLL corpus) · paper: Understanding Linearity of Cross-Lingual Word Embedding Mappings
structure: 1D continuum manifold · models: Gemma-2-2B, EmbeddingGemma, word2vec (trained on Wikipedia), Claude 3.5 Haiku · paper: Symmetry in Language Statistics Shapes the Geometry of Model Representations, When Models Manipulate Manifolds: The Geometry of a Counting Task
structure: 1D continuum manifold · models: 1-Layer Ordinal Local-Comparison Transformer, Qwen2.5-1.5B · paper: Emergent Ordinal Geometry in Transformers Trained on Local Comparisons
structure: Polytope (Simplex) · models: Toy ReLU-output autoencoder (h=Wx, x'=ReLU(W^T h + b)) · paper: Toy Models of Superposition
structure: Linear Subspace, Attention–MLP Sufficiency Staging Hypothesis · models: Grid-walker decoder transformer (L4/H4/d_model=128, HookedTransformer) · paper: Predictive Statistics Shape Emergent World Representations of Grid Walkers