MATH · IN · MODELS

Structures

The map's objects are classified by their mathematical role (the type axis), not by a single parent category. Orthogonal properties (curvature, topology, linearity, specification) are facets shown on each node's page. See the classification tree for the full road from an object to its leaf.

Hypotheses are claims about objects rather than objects, so they live on their own page —all 14 of them.

Each card carries the same computed replication read-out as the structure's own page:filled — two or more papers with no shared author reported it (replication);outlined — two or more values, but all from one study (breadth);grey — a single value. Derived from the corpus, never hand-assigned.

geometric-object
A set the activations lie on (carrier / subset / embedded set / direction).
1D continuum manifold
Open (non-periodic) one-dimensional manifold with extrinsic 'ripples' encoding an ordered continuous coordinate. Projected onto any two principal components it traces a Lissajous curve; higher-frequency harmonics are the ripples.
embedded-setK:zeroparametrizationmanifold
14parch 4dom 4fam 1410 obs.
2D Grid (Square Lattice)
A finite 2-dimensional square-lattice arrangement of points — items placed at integer grid coordinates (i,j) with horizontal/vertical nearest-neighbour adjacency. Distinguished from a torus by being bounded and non-periodic (contractible, open boundary; no global wrap-around cycle vs. H1=Z^2), and from a 1D continuum by having two independent lattice coordinates rather than one.
embedded-setK:zerosquare lattice (bounded 2D grid graph, open boundary)
1parch 1dom 1fam 21 obs.
Affine Subspace
Linear subspace with an offset: A = x₀ + V. Flat, open manifold of zero curvature. Unlike a linear subspace, need not pass through the origin.
subsetaffineK:zeromanifold
6parch 5dom 2fam 106 obs.
Circle
Closed 1-manifold S¹ encoding a single cyclic variable. Topologically distinct from the torus T² = S¹ × S¹: one independent cycle (H₁ = ℤ), not two (H₁ = ℤ²).
subsetK:zeromanifold
33parch 9dom 5fam 3327 obs.
Cone
The topological cone construction over a base space B — rays from an apex through every point of B, scaled by a shared non-negative 'magnitude' coordinate. Two instances: a polyhedral cone over a finite set of directions, and a circular cone over S¹ (angle × magnitude).
subsetnon-manifold
10parch 3dom 2fam 107 obs.
Decision boundary (as a codimension-1 hypersurface)
The codimension-1 level set $\{x : f(x) = 0\}$ separating two predicted classes in a classifier's input (or feature) space — a distinct geometric object at the function level, as opposed to the activation-cloud/representation-level manifolds this map otherwise catalogs.
subsetK:variableimplicitmanifold
7parch 3dom 3fam 137 obs.
Generalized Helix
A scalar quantity a is embedded as helix(a) = C·B(a), where B(a) stacks a linear term with cos/sin pairs at several periods T_1,...,T_k — a single number simultaneously encoded as 'how far along an unbounded axis' and 'where on each of several superposed circles,' generalizing a plain circle (one period) to several periods with generally incommensurate ratios, embedded into a linear subspace by a single learned matrix.
embedded-setnonlinearK:variableparametrizationunbounded 1-dimensional curve (non-compact, unlike the circle it generalizes)
2parch 1dom 1fam 32 obs.
Hyperbolic Manifold
Manifold of constant negative curvature. Ball volume grows exponentially with radius — a natural geometry for encoding hierarchies with exponential branching.
subsetK:negativemanifold
2parch 3dom 2fam 62 obs.
Laguerre-Voronoi partition (weighted power diagram of a linear readout layer)
The weighted Voronoi (Laguerre / power-diagram) partition of representation space induced by a linear readout layer's own weights and biases: each output unit $j$ becomes a center $c_j$ with a weight $\nu_j$ derived from its bias, and the cell $\mathcal{C}_j=\{z: \|z-c_j\|^2-\nu_j \le \|z-c_i\|^2-\nu_i\ \forall i\}$ is a convex polyhedron algebraically identical to the layer's own argmax decision region. A 'concept' is redefined as an entire cell (a region), not a point or a direction, and a 'category' as a union of cells.
subsetK:zeroimplicitpolyhedral complex
1parch 1dom 1fam 41 obs.
Linear Direction
Unit vector rᶠ ∈ ℝᵈ encoding feature f. The projection rᶠ · x monotonically reflects the feature's value in the context defining activation x.
directionlinearK:zeromanifold
288parch 14dom 15fam 159283 obs.
Linear region arrangement (polyhedral tessellation of input space)
The partition of a piecewise-linear (e.g. ReLU) network's input space into finitely many convex polyhedral cells — one per activation pattern — on which the network computes a single affine map. A distinct object from [[decision-boundary]]: the boundary is one codimension-1 level set of the *output*, while the region arrangement is the *entire* combinatorial/geometric tessellation induced by every breakpoint of every unit, at every layer.
subsetK:zeroimplicitpolyhedral complex
3parch 2dom 2fam 23 obs.
Linear Subspace
k-dimensional linear subspace V ⊂ ℝᵈ corresponding to a group of related features. Generalises the single feature direction to the multidimensional case.
subsetlinearK:zerolevel-setmanifold
139parch 14dom 11fam 106140 obs.
Lissajous Curves
A parametric curve traced by superposing sinusoids of different frequencies along different axes: (x(t), y(t)) = (A sin(at+δ₁), B sin(bt+δ₂)). The 2D shadow of any multi-frequency 1D manifold — the shared geometric signature behind manifolds-circle/torus and manifolds-1d-continuum.
embedded-setparametrizationnon-manifold
6parch 3dom 2fam 84 obs.
Local hyperboloid-type quadric patch
A local patch of embedding space, spanned by a tightly-controlled family of near-paraphrase sentences, that is better fit by a curved implicit quadric surface (predominantly one-sheeted-hyperboloid type) than by an affine subspace or ellipsoid — an open, saddle-like local carrier rather than a flat cluster or a closed convex blob.
embedded-setK:variableimplicitmanifold
1parch 1dom 1fam 21 obs.
Minkowski Sum of Tile Polytopes (Minkowski Representation Hypothesis)
The whole activation space of a layer, not just one categorical concept, decomposes as the Minkowski (vector) sum of many tile polytopes — disjoint convex hulls over partitioned archetype dictionaries — with any activation expressed as a block-sparse convex combination using only a handful of active tiles at once. Not a relativistic/indefinite-metric structure despite the name: 'Minkowski' here means the classical convex-geometry Minkowski sum A ⊕ B = {a+b : a∈A, b∈B}.
subsetmanifold-with-corners
1parch 1dom 1fam 11 obs.
Paraboloid (circular × pinched continuum)
A surface of revolution combining a circular coordinate with a continuous 'axial' coordinate, where the circle's radius (spread) varies with — and can shrink toward zero at — position along the axis, unlike a flat cylinder (S¹ × ℝ) where the circle's radius stays constant.
embedded-setK:variableparametrizationmanifold
1parch 1dom 1fam 12 obs.
Pinched Manifold (Singular Quotient)
A word/embedding space is not a smooth manifold but a pinched one: a manifold of 'meanings' with some of its points identified (glued together) to produce singular points, each singularity corresponding to a polysemous word with multiple senses folded into one vector.
embedded-setnonlinearK:n/aimplicitquotient of a manifold by a finite point-identification: locally homeomorphic to an open ball everywhere except at a finite set of singular points, where several sheets are glued at a common centre
1parch 1dom 1fam 11 obs.
Polytope (Simplex)
The convex hull of k vector representations, one per value of a mutually-exclusive (categorical) concept — e.g. {mammal, bird, reptile, fish}. Generically a (k-1)-simplex, not just k independent directions: unlike a binary feature (one direction), a categorical feature needs a magnitude-bearing vector per value so that differences between values remain meaningful.
subsetmanifold-with-corners
21parch 8dom 4fam 2119 obs.
Sphere
Closed homogeneous manifold of constant positive curvature. Normalisation onto a fixed norm places activations on the unit hypersphere Sⁿ⁻¹, making cosine similarity the natural distance.
subsetK:positivemanifold
4parch 2dom 2fam 64 obs.
Stratified Manifold with Continuous Fibers
A representation space partitioned into a discrete set of macroscopic basins (indexed by a discrete output/class variable) with each basin's interior further organized by continuous 1D fibers that extend across basin boundaries, tracking a separate continuous quantity — a discrete clustering structure and a continuous stratification structure coexisting at different scales of the same space, rather than either alone.
embedded-setnonlinearK:variableparametrizationdisjoint union of basins (discrete base), each traversed by continuous 1D fibers connecting adjacent basins
1parch 2dom 2fam 21 obs.
Torus
Closed manifold T² = S¹ × S¹ — the product of two independent circles. Distinguished from a bare circle by first homology: H₁(T²) = ℤ² (two independent cycles) versus H₁(S¹) = ℤ (one).
subsetK:zeromanifold
9parch 2dom 3fam 87 obs.
configuration
An ordered tuple of several points — a relational pattern in configuration space.
operator
A map or matrix acting on the cloud (not a shape of the cloud).
Attribute-Induced Embedding Folding
A continuous, attribute-indexed vector field on an embedding manifold — the pushforward through the embedding map of an attribute's local one-parameter variation. A model achieves invariance to an attribute not by ignoring it but by nonlinearly folding/contracting the space along this field; the amount of folding (an 'invariance energy') is a per-attribute, per-architecture quantity, not a fixed geometric constant.
nonlinearvectorattribute-indexed tangent vector field, pushed forward through the embedding map; invariance measured as a local alignment energy
1parch 1dom 1fam 11 obs.
Conceptor (soft ellipsoidal region)
A soft generalization of a linear subspace: a positive semi-definite matrix C, fit to the covariance of a set of activation vectors, whose eigenvalues lie continuously in [0,1] rather than being restricted to exactly 0 or 1 — so 'membership' in the region is graded rather than binary.
PSD-contraction
3parch 2dom 2fam 53 obs.
Continuous causal attention operator (CCT)
A generalization of discrete causal self-attention that replaces the sum over token positions with an integral over a continuous time variable, applied to a continuous-time embedding function x(t) rather than a discrete token sequence — using the same pretrained weight matrices unchanged, and proven to reduce exactly to standard discrete attention when the continuous function is stepwise-constant with unit-duration steps.
nonlinearintegral-attention
1parch 1dom 1fam 51 obs.
Finite-Lag Transport Tensor
An empirically-estimated tensor G_delta summarizing how a recurrent network's own hidden-state trajectory transports over a fixed lag delta -- decomposing exactly into how much successor states disperse around their conditional mean (conditional-spread) and how far that mean itself moves (coherent-displacement), plus a rotational/circulation term -- distinct from a Jacobian-based linearization since it is a statistic of realized (source, successor) state pairs, not a local derivative.
nonlinearsource-centered-transport
1parch 1dom 1fam 11 obs.
Jacobian non-normality (Schur decomposition of per-layer operators)
The Jacobian J_l of a transformer block's output with respect to its input (in the residual stream) is, at every layer of every trained model tested, a non-normal operator: its complex Schur decomposition J = Q(Lambda+N)Q* has a nonzero off-diagonal part N distinct from its eigenvalues, and training reshapes how non-normal each layer's Jacobian is (a depth-wise gradient from rotation-dominated to near-symmetric), with N -- not the eigenvalue spectrum Lambda -- responsible for a cumulative, depth-compounding collapse in the effective rank of composed Jacobians.
linearized-block-map
1parch 1dom 1fam 31 obs.
Symmetric/skew decomposition of the attention query-key operator
The combined query-key bilinear form W_qk = W_Q W_K^T of a trained attention head decomposes uniquely into a symmetric part (undirected, content-based similarity) and a skew-symmetric part (directed, positional/causal asymmetry); real trained encoder-only Transformers are measurably symmetric-dominant and real trained decoder-only Transformers are measurably directionality-dominant, and imposing the symmetric structure at initialization causally speeds convergence.
bilinear-form
1parch 1dom 1fam 11 obs.
metric-model
A metric, seminorm, or learned distance model.
Diffusion Spacetime Information Geometry
A diffusion model's own family of denoising distributions {p(x0|xt,t)}, indexed jointly by noisy state and time, forms an exponential-family statistical manifold carrying a Fisher-Rao metric; geodesics under this metric (computable in closed form from the trained denoiser, with no numerical simulation) define a Diffusion Edit Distance between clean data points, distinct in kind from any Riemannian metric obtained by pulling back an ambient/activation-space metric through the model's Jacobian.
fisher-rao-metric
1parch 1dom 2fam 21 obs.
Score-Jacobian pullback Riemannian metric
A genuine Riemannian metric tensor G = JᵀJ built directly from the Jacobian of a real diffusion model's own score function, with no linearity assumption and no ambient/latent metric pulled back through an assumed linear structure — used to compute length-minimizing geodesics between real data points for interpolation. Distinct from a Fisher-Rao information-geometric metric (diffusion-spacetime-information-geometry) despite both living on a diffusion model's own manifold.
score-jacobian-metric
1parch 1dom 1fam 21 obs.
Tree Metric Embedding
A learned inner product (equivalently, a linear transform B) under which squared distance between two representations approximates graph distance in a discrete tree, and squared norm approximates depth from the root — embedding an entire tree's edge structure into the metric of a linear subspace, not just a single decodable relation.
learned-metric
14parch 4dom 2fam 1512 obs.
distribution-property
An unsupervised property of the activation distribution.
labeled-data-property
A property of labeled sets (separability, margin, decodability).
measurement
A computed quantity or profile (scalar / vector / functional).
Attention-graph spectral profile (Fiedler value, HFER, smoothness, spectral entropy)
Treating a transformer's attention matrix at each layer as the weighted adjacency matrix of a dynamic graph over tokens, and reducing it to four scalar diagnostics from its graph Laplacian spectrum: the Fiedler value (algebraic connectivity), high-frequency energy ratio (HFER), graph-signal smoothness (Dirichlet energy of hidden states as a graph signal), and spectral entropy of the Laplacian eigenvalue distribution.
functional
1parch 1dom 1fam 41 obs.
Attention-weight frequency-band specialization (DFT/wavelet decomposition of attention as a position-indexed signal)
Treating each attention head's weight pattern, at a fixed query position, as a 1-D signal indexed by relative key position, and decomposing it via DFT or wavelet transform into frequency bands -- revealing that individual heads specialize in distinct, consistent frequency bands (e.g. low-frequency/long-range vs. high-frequency/local) rather than an undifferentiated graph-connectivity pattern.
functional
1parch 1dom 1fam 21 obs.
Curvature profile of the representation manifold
How much a network's representation manifold bends — measured, not assumed as a named shape. Two non-interchangeable notions: extrinsic principal curvature (MAPC, local-PCA fit, ambient bending) and intrinsic graph-Ricci curvature (Ollivier/Forman, optimal-transport/combinatorial). Reported as a profile across depth/training, orthogonal to intrinsic dimension.
functional
18parch 10dom 6fam 1718 obs.
Intrinsic-dimension profile across depth
The dimensionality of the data manifold a network's own representations lie on is not fixed — it expands sharply in early layers, then contracts to a low-dimensional plateau or minimum, and the layer at that minimum tends to carry the most abstract, task-relevant semantic content.
functional
38parch 11dom 8fam 4236 obs.
Persistent-homology / Betti profile across depth
Betti numbers / persistent topological features (independent loops, holes, connected components) of a real model's activation point-cloud, tracked as a profile across layers, training, or model comparison — a computed topological-data-analysis quantity, not a claim that the representation manifold is any specific named shape. Orthogonal to intrinsic dimension and to curvature: a topologically rich point cloud can be low- or high-dimensional, flat or curved.
functional
3parch 3dom 3fam 43 obs.
Representational-similarity trajectory across depth
How similar a network's own layer-wise representations are to each other (via a similarity metric such as CCA, SVCCA, or CKA) traces a specific, measurable trajectory across depth — layers cluster into similarity-based stages, and the trajectory's shape shifts with training objective, architecture, or between independently trained networks.
functional
8parch 7dom 3fam 168 obs.
Semantic information production rate along the diffusion trajectory
The Bayes-optimal conditional entropy of a semantic class label given a diffusion model's own noisy state, and its time-derivative (bits of class information produced per unit diffusion time), traces a measurable, non-uniform trajectory across the generative process -- a purely information-theoretic diagnostic of when a real trained diffusion model's generative dynamics commit to specific semantic content, distinct from any curvature or dimensionality measurement of the same trajectory.
functional
1parch 1dom 1fam 11 obs.
Spatial-frequency (spectral) accessibility profile across depth
How linearly recoverable each spatial-frequency band of the original input image is from a vision model's own layer-wise representation, tracked as a profile across depth via ridge-regression probes against ground-truth Fourier-energy targets, relative to a dimension-matched random-projection baseline that isolates learned-transformation effects from pure dimensionality change.
functional
1parch 2dom 2fam 21 obs.
dynamical-object
An object defined relative to a dynamical system (attractor, invariant set).
combinatorial-object
An order, lattice, graph, matroid, or complex.
empirical-pattern
An observed organization of the data in a specific model.