MATH · IN · MODELS

Matrix

The map's mathematical structures against its models — an overview plus four detailed cuts.

transformer-decoder47transformer-encoder19cnn19diffusion13rnn15multimodal8vision-transformer15vae-autoencoder7transformer-enc-dec10dual-encoder12shallow-embedding12mlp7gnn7ssm-linear-attention8gan5
language4944168823712117
vision2911912814312145
audio102812123431
control9425111
algorithmic147613
protein32321
molecular7121124
music3212
board-game543
code3211
graph414
tabular3211
genomics6511
time-series21111
action221
weather443
sign-language2221
remote-sensing33
transformer-decoder331transformer-encoder117cnn57vision-transformer39rnn42diffusion55shallow-embedding18dual-encoder20transformer-enc-dec20multimodal40
Anisotropy
Aristotelian Representation Hypothesis
Attention-weight frequency-band specialization (DFT/wavelet decomposition of attention as a position-indexed signal)
Attention-graph spectral profile (Fiedler value, HFER, smoothness, spectral entropy)
Symmetric/skew decomposition of the attention query-key operator
Attention reference frame (sink-token anchor configuration)
Attribute-Induced Embedding Folding
Belief State Geometry Hypothesis (Mixed-State Presentation)
Concept Cluster Heterogeneity
Concept Crystals (Parallelogram/Trapezoid Structure)
Concept lattice (Formal Concept Analysis half-space model)
Conceptor (soft ellipsoidal region)
Conceptual Belief Space Hypothesis
Constraint-Algebra Basis Hypothesis
Constructive Interference Hypothesis
Continuous causal attention operator (CCT)
Decision boundary (as a codimension-1 hypersurface)
Semantic information production rate along the diffusion trajectory
Diffusion Spacetime Information Geometry
Dimensional collapse
Feature Lobes (Spatial-Functional Modularity)
Finite-Lag Transport Tensor
Ghost Point (saddle-node bifurcation remnant)
Intrinsic-dimension profile across depth
Jacobian non-normality (Schur decomposition of per-layer operators)
Laguerre-Voronoi partition (weighted power diagram of a linear readout layer)
Limit cycle (stable periodic attractor)
Line Attractor
Linear Centroids Hypothesis
Linear Direction
Linear region arrangement (polyhedral tessellation of input space)
Linear Separability
Linear Subspace
Lissajous Curves
Linear Representation Hypothesis
Curvature profile of the representation manifold
1D continuum manifold
Affine Subspace
Circle
Cone
2D Grid (Square Lattice)
Generalized Helix
Hyperbolic Manifold
Local hyperboloid-type quadric patch
Paraboloid (circular × pinched continuum)
Polytope (Simplex)
Sphere
Stratified Manifold with Continuous Fibers
Torus
Minkowski Sum of Tile Polytopes (Minkowski Representation Hypothesis)
Nonlinear World-Model Decodability
Persistent-homology / Betti profile across depth
Pinched Manifold (Singular Quotient)
Platonic Representation Hypothesis
Relation frame (ordered multi-token tuple geometry)
Representational-similarity trajectory across depth
Score-Jacobian pullback Riemannian metric
Spatial-frequency (spectral) accessibility profile across depth
Attention–MLP Sufficiency Staging Hypothesis
Tangent-Aligned Anisotropy Hypothesis
Task-induced symmetry class of a learned representation
Tree Metric Embedding
language477vision216algorithmic30audio39control30molecular15genomics6board-game9weather3graph7
Anisotropy
Aristotelian Representation Hypothesis
Attention-weight frequency-band specialization (DFT/wavelet decomposition of attention as a position-indexed signal)
Attention-graph spectral profile (Fiedler value, HFER, smoothness, spectral entropy)
Symmetric/skew decomposition of the attention query-key operator
Attention reference frame (sink-token anchor configuration)
Attribute-Induced Embedding Folding
Belief State Geometry Hypothesis (Mixed-State Presentation)
Concept Cluster Heterogeneity
Concept Crystals (Parallelogram/Trapezoid Structure)
Concept lattice (Formal Concept Analysis half-space model)
Conceptor (soft ellipsoidal region)
Conceptual Belief Space Hypothesis
Constraint-Algebra Basis Hypothesis
Constructive Interference Hypothesis
Continuous causal attention operator (CCT)
Decision boundary (as a codimension-1 hypersurface)
Semantic information production rate along the diffusion trajectory
Diffusion Spacetime Information Geometry
Dimensional collapse
Feature Lobes (Spatial-Functional Modularity)
Finite-Lag Transport Tensor
Ghost Point (saddle-node bifurcation remnant)
Intrinsic-dimension profile across depth
Jacobian non-normality (Schur decomposition of per-layer operators)
Laguerre-Voronoi partition (weighted power diagram of a linear readout layer)
Limit cycle (stable periodic attractor)
Line Attractor
Linear Centroids Hypothesis
Linear Direction
Linear region arrangement (polyhedral tessellation of input space)
Linear Separability
Linear Subspace
Lissajous Curves
Linear Representation Hypothesis
Curvature profile of the representation manifold
1D continuum manifold
Affine Subspace
Circle
Cone
2D Grid (Square Lattice)
Generalized Helix
Hyperbolic Manifold
Local hyperboloid-type quadric patch
Paraboloid (circular × pinched continuum)
Polytope (Simplex)
Sphere
Stratified Manifold with Continuous Fibers
Torus
Minkowski Sum of Tile Polytopes (Minkowski Representation Hypothesis)
Nonlinear World-Model Decodability
Persistent-homology / Betti profile across depth
Pinched Manifold (Singular Quotient)
Platonic Representation Hypothesis
Relation frame (ordered multi-token tuple geometry)
Representational-similarity trajectory across depth
Score-Jacobian pullback Riemannian metric
Spatial-frequency (spectral) accessibility profile across depth
Attention–MLP Sufficiency Staging Hypothesis
Tangent-Aligned Anisotropy Hypothesis
Task-induced symmetry class of a learned representation
Tree Metric Embedding
Llama35Gemma29Qwen75Mistral20Pythia10GPT8ResNet13CLIP (Contrastive Language-Image Pretraining)9DINOv25word2vec8
Anisotropy
Aristotelian Representation Hypothesis
Attention-weight frequency-band specialization (DFT/wavelet decomposition of attention as a position-indexed signal)
Attention-graph spectral profile (Fiedler value, HFER, smoothness, spectral entropy)
Symmetric/skew decomposition of the attention query-key operator
Attention reference frame (sink-token anchor configuration)
Attribute-Induced Embedding Folding
Belief State Geometry Hypothesis (Mixed-State Presentation)
Concept Cluster Heterogeneity
Concept Crystals (Parallelogram/Trapezoid Structure)
Concept lattice (Formal Concept Analysis half-space model)
Conceptor (soft ellipsoidal region)
Conceptual Belief Space Hypothesis
Constraint-Algebra Basis Hypothesis
Constructive Interference Hypothesis
Continuous causal attention operator (CCT)
Decision boundary (as a codimension-1 hypersurface)
Semantic information production rate along the diffusion trajectory
Diffusion Spacetime Information Geometry
Dimensional collapse
Feature Lobes (Spatial-Functional Modularity)
Finite-Lag Transport Tensor
Ghost Point (saddle-node bifurcation remnant)
Intrinsic-dimension profile across depth
Jacobian non-normality (Schur decomposition of per-layer operators)
Laguerre-Voronoi partition (weighted power diagram of a linear readout layer)
Limit cycle (stable periodic attractor)
Line Attractor
Linear Centroids Hypothesis
Linear Direction
Linear region arrangement (polyhedral tessellation of input space)
Linear Separability
Linear Subspace
Lissajous Curves
Linear Representation Hypothesis
Curvature profile of the representation manifold
1D continuum manifold
Affine Subspace
Circle
Cone
2D Grid (Square Lattice)
Generalized Helix
Hyperbolic Manifold
Local hyperboloid-type quadric patch
Paraboloid (circular × pinched continuum)
Polytope (Simplex)
Sphere
Stratified Manifold with Continuous Fibers
Torus
Minkowski Sum of Tile Polytopes (Minkowski Representation Hypothesis)
Nonlinear World-Model Decodability
Persistent-homology / Betti profile across depth
Pinched Manifold (Singular Quotient)
Platonic Representation Hypothesis
Relation frame (ordered multi-token tuple geometry)
Representational-similarity trajectory across depth
Score-Jacobian pullback Riemannian metric
Spatial-frequency (spectral) accessibility profile across depth
Attention–MLP Sufficiency Staging Hypothesis
Tangent-Aligned Anisotropy Hypothesis
Task-induced symmetry class of a learned representation
Tree Metric Embedding
Llama-3.1-8B1Gemma-2-2B1Llama-3-8B1Gemma-2-9B1Llama-3.1-8B-Instruct1Llama-3-8B-Instruct1Llama-2-7B1Mistral-7B1Llama-3.2-1B1Llama-3.2-3B1
Anisotropy
Aristotelian Representation Hypothesis
Attention-weight frequency-band specialization (DFT/wavelet decomposition of attention as a position-indexed signal)
Attention-graph spectral profile (Fiedler value, HFER, smoothness, spectral entropy)
Symmetric/skew decomposition of the attention query-key operator
Attention reference frame (sink-token anchor configuration)
Attribute-Induced Embedding Folding
Belief State Geometry Hypothesis (Mixed-State Presentation)
Concept Cluster Heterogeneity
Concept Crystals (Parallelogram/Trapezoid Structure)
Concept lattice (Formal Concept Analysis half-space model)
Conceptor (soft ellipsoidal region)
Conceptual Belief Space Hypothesis
Constraint-Algebra Basis Hypothesis
Constructive Interference Hypothesis
Continuous causal attention operator (CCT)
Decision boundary (as a codimension-1 hypersurface)
Semantic information production rate along the diffusion trajectory
Diffusion Spacetime Information Geometry
Dimensional collapse
Feature Lobes (Spatial-Functional Modularity)
Finite-Lag Transport Tensor
Ghost Point (saddle-node bifurcation remnant)
Intrinsic-dimension profile across depth
Jacobian non-normality (Schur decomposition of per-layer operators)
Laguerre-Voronoi partition (weighted power diagram of a linear readout layer)
Limit cycle (stable periodic attractor)
Line Attractor
Linear Centroids Hypothesis
Linear Direction
Linear region arrangement (polyhedral tessellation of input space)
Linear Separability
Linear Subspace
Lissajous Curves
Linear Representation Hypothesis
Curvature profile of the representation manifold
1D continuum manifold
Affine Subspace
Circle
Cone
2D Grid (Square Lattice)
Generalized Helix
Hyperbolic Manifold
Local hyperboloid-type quadric patch
Paraboloid (circular × pinched continuum)
Polytope (Simplex)
Sphere
Stratified Manifold with Continuous Fibers
Torus
Minkowski Sum of Tile Polytopes (Minkowski Representation Hypothesis)
Nonlinear World-Model Decodability
Persistent-homology / Betti profile across depth
Pinched Manifold (Singular Quotient)
Platonic Representation Hypothesis
Relation frame (ordered multi-token tuple geometry)
Representational-similarity trajectory across depth
Score-Jacobian pullback Riemannian metric
Spatial-frequency (spectral) accessibility profile across depth
Attention–MLP Sufficiency Staging Hypothesis
Tangent-Aligned Anisotropy Hypothesis
Task-induced symmetry class of a learned representation
Tree Metric Embedding
2+ papers, no shared author2+ papers, overlapping authorsone paperthe dot says how independently the cell is sourced — computed, not hand-assigned
What the axes mean — architecture classes & domains
Architecture classes (15)
transformer-decoder
Autoregressive (causal-attention) Transformer — the GPT-style stack behind most large language models.
transformer-encoder
Bidirectional (full-attention) Transformer trained with masked objectives; BERT-style representation models.
transformer-enc-dec
Encoder–decoder (“seq2seq”) Transformer with cross-attention: T5, the original Transformer, most translation models.
vision-transformer
Transformer applied to image patches (ViT and descendants) rather than 1-D token sequences.
ssm-linear-attention
State-space and linear-attention sequence models (Mamba, RWKV, S4) — sub-quadratic alternatives to softmax attention.
rnn
Recurrent networks (LSTM, GRU, vanilla RNN) that carry a hidden state through time.
diffusion
Iterative denoising generative models (DDPM, latent / score-based diffusion).
gan
Generative adversarial networks: a generator trained against a discriminator.
vae-autoencoder
(Variational) autoencoders and reconstruction models with a learned latent bottleneck.
cnn
Convolutional networks (ResNet, U-Net) built on local, weight-shared filters.
gnn
Graph neural networks that pass messages over nodes and edges.
mlp
Plain feed-forward / multilayer-perceptron networks — no sequence, conv, or attention structure.
dual-encoder
Two-tower contrastive models embedding two inputs into one shared space (CLIP, sentence-transformers).
shallow-embedding
Non-deep lookup embeddings (word2vec, GloVe, node2vec) — a single learned vector table, not a deep network.
multimodal
Models whose core architecture fuses several modalities (e.g. vision-language) and doesn’t reduce to one class above.
Domains (18)
language
Natural-language text.
code
Source code and programming languages.
vision
Images and video — natural-image understanding and generation.
sign-language
Signed languages captured as video or pose.
music
Musical audio and symbolic scores.
audio
Non-music audio: speech, sound, and general audio signals.
protein
Protein sequences and structures.
molecular
Small molecules and chemistry (SMILES, molecular graphs).
genomics
DNA / RNA sequences and single-cell gene expression.
time-series
Sequential numeric measurements — sensor, physiological, financial, forecasting.
graph
Relational / network-structured data.
board-game
Board and strategy games (chess, Go, Othello, poker).
control
Continuous control, navigation, and robotics / maze policies.
algorithmic
Synthetic reasoning and algorithmic tasks (copy, sort, arithmetic, evidence integration).
tabular
Structured rows-and-columns data and recommenders.
weather
Atmospheric state and weather forecasting (e.g. ERA5-trained foundation models).
remote-sensing
Satellite and Earth-observation imagery.
action
Embodied action — vision-language-action policies and agent trajectories.