MATH · IN · MODELS

Hypotheses

A hypothesis is a falsifiable general claim about representations — the linear representation hypothesis, the Platonic representation hypothesis, and their kin. It is a role of its own, not a shape: the Minkowski sum of polytopes is an object, while "activations decompose as a Minkowski sum" is a claim about one. Each page below collects the evidence for and against its claim in one place.

The chips carry the computed replication read-out:filled — two or more papers with no shared author reported it;outlined — two or more values, but all from one study;grey — a single value.

Aristotelian Representation Hypothesis
Neural networks trained with different objectives, on different data and modalities, converge to shared local neighborhood relationships -- but not to shared global spectral geometry. Once representational-similarity metrics are corrected for a scale confound (see null-calibrated-representational-similarity), CKA-style global convergence largely disappears while mutual k-NN-style local convergence persists. Proposed by Groger, Wen & Brbic (2026) as a calibrated refinement of the Platonic Representation Hypothesis.
1parch 3dom 2fam 81 obs.
Papers
Gröger 2026
Attention–MLP Sufficiency Staging Hypothesis
A Transformer trained only on a next-step prediction objective still builds, in its attention sublayer, a full architecture-specific coordinate system for the entire predictively-sufficient world state — more than the objective strictly requires — which downstream MLP/LayerNorm sublayers then compress into the smaller statistic the objective actually needs; this two-stage separation is a property of the attention mechanism specifically, not a necessary consequence of reaching Bayes-optimal prediction, since recurrent networks reach the identical optimum without ever isolating the world state as a distinct geometric stage.
1parch 1dom 1fam 11 obs.
Papers
Brenner 2026
Belief State Geometry Hypothesis (Mixed-State Presentation)
For a sequence-generating process with a known hidden state (an HMM / epsilon-machine), a predictive model's residual stream linearly embeds the current Bayesian belief b_t = p(hidden state | x_{1:t}) — a point in a probability simplex over the process's causal states — via the process's Mixed-State Presentation (MSP), with next-token predictions recoverable as a barycentric combination of per-vertex predictive distributions.
4parch 2dom 3fam 44 obs.
Papers
Levinson 2026 · Agarwal 2026 · Kamel 2025 · Shai 2024
Conceptual Belief Space Hypothesis
In-context updating is modeled as Bayesian belief updating over a Gärdenfors-style conceptual space: a belief state is a probability distribution over a low-dimensional metric space whose concepts are convex regions, tracing a trajectory as more context is read.
1parch 1dom 1fam 11 obs.
Papers
Bigelow 2026
Constraint-Algebra Basis Hypothesis
When a task's validity conditions are governed by an explicit algebraic/combinatorial constraint structure rather than by the surface units a human observer would naturally decompose the state into, a trained model's linearly-decodable world-model features align with the constraint structure's own units, not the surface units — e.g. a Sudoku-solving transformer's linear features are organized by row/column/box membership, not by individual cell, because presence-in-a-substructure, not cell identity per se, is what determines placement validity.
1parch 1dom 1fam 11 obs.
Papers
Kniazev 2026
Constructive Interference Hypothesis
When packed features are correlated rather than independent, an under-complete, weight-decayed autoencoder's optimal solution does not minimize pairwise interference between feature directions (the classical near-orthogonal/regular-polytope picture) — it instead sets its weight columns to the top principal components of the feature covariance, making interference between co-occurring features additive with the signal rather than adversarial to it, and reproducing whatever geometry (clusters, circles) that covariance's own top eigenvectors have.
1parch 2dom 1fam 21 obs.
Papers
Prieto 2026
Efficient/Capacity-Optimal Coding Hypothesis
The same rippled/Lissajous-shaped 1D manifold geometry can arise from capacity-optimal encoding of a scalar quantity sharing a fixed embedding dimension with many other features under noise/interference — a distinct generative mechanism from translation-symmetric corpus statistics.
1parch 1dom 1fam 10 obs.
Papers
Gurnee 2025
Linear Centroids Hypothesis
Features of a deep network are linear directions not in its raw activation space, but in its centroid space — the space of Jacobian row-sums mu = J^T*1 summarizing each local affine expert's actual input-output map. Proposed by Walker, Humayun, Balestriero & Baraniuk (2026) as a mapping-aware replacement for the Linear Representation Hypothesis.
1parch 3dom 2fam 51 obs.
Papers
Walker 2026
Linear Representation Hypothesis
Human-interpretable features are encoded as linear directions in activation space; linear operations over them (addition, reflection, sign flip) are semantically meaningful.
2parch 1dom 1fam 22 obs.
Papers
Garg 2026 · Park 2023
Minkowski Representation Hypothesis
A layer's whole activation space decomposes as the Minkowski (vector) sum of many tile polytopes, with any activation a block-sparse convex combination using only a handful of active tiles at once. The claim (mechanism), as distinct from the polytope-sum object it postulates ([[minkowski-sum-polytope]]).
1parch 1dom 1fam 10 obs.
Papers
Fel 2025
Platonic Representation Hypothesis
Different neural networks — regardless of architecture and training modality — converge to a shared representation geometry, measured via kernel-alignment metrics such as CKA or mutual nearest-neighbor overlap. Introduced by Huh, Cheung, Wang & Isola (2024).
26parch 10dom 7fam 3822 obs.
Papers
Kendiukhov 2026 · Craig 2026 · Sarkar 2026 · Zhang 2026 · +22 more
Spatial-Functional Modularity Hypothesis
Functionally related features (those that co-fire) also occupy shared, spatially-coherent regions of representation space — the generalized, falsifiable claim behind the observed feature-lobes pattern ([[feature-lobes]]).
1parch 1dom 1fam 10 obs.
Papers
Li 2024
Tangent-Aligned Anisotropy Hypothesis
Frequency-biased sampling concentrates high-frequency tokens near their idealized centroid, geometrically attenuating the visibility of local manifold curvature; this radial concentration induces a training-time gradient bias that preferentially amplifies directions tangent to the local data manifold over normal (curvature-bearing) directions, self-reinforcing through attention and residual connections — producing anisotropy as a byproduct of frequency-driven sampling and gradient dynamics, not merely a training-objective artifact.
1parch 2dom 1fam 31 obs.
Papers
Bernas 2026
Translation Symmetry Hypothesis
Representational manifolds (circles, open 1D continua, linearly decodable coordinates) share a common origin: translation symmetry in co-occurrence statistics. If co-occurrence depends only on distance along a semantic continuum, the co-occurrence matrix is circulant/Toeplitz and its eigenvectors are automatically Fourier modes.
1parch 2dom 1fam 20 obs.
Papers
Karkada 2026