MATH · IN · MODELS
structures / Hypotheses / Minkowski Representation Hypothesis

Minkowski Representation Hypothesis

CLAIMhypothesisadvancedhow it's classified →

A layer's whole activation space decomposes as the Minkowski (vector) sum of many tile polytopes, with any activation a block-sparse convex combination using only a handful of active tiles at once. The claim (mechanism), as distinct from the polytope-sum object it postulates ([[minkowski-sum-polytope]]).

Replicationcomputed from the corpus — never hand-assigned
1 paper1 architecture class1 domain1 model family
Filled = two or more values reported by papers that share no author — replication. Outlined = two or more values, but all from a single study — breadth, not replication. Grey = a single value. Derived from paper authorship and each model's architecture class, domain and family; it updates itself when a paper is added.

Statement

The Minkowski Representation Hypothesis (MRH) claims that a layer’s activation space XX satisfies X=iPiX=\bigoplus_i P_i, the Minkowski sum of disjoint tile polytopes Pi=conv(ATi)P_i=\mathrm{conv}(A_{T_i}) over a partitioned archetype dictionary, and that any activation is a block-sparse convex combination x=iSziATix=\sum_{i\in S} z_i A_{T_i} with ziΔTiz_i\in\Delta^{|T_i|} and Sm|S|\ll m (only a few tiles active at once).

This node is the hypothesis (a falsifiable general claim about how a whole layer is organized). The geometric object it postulates — the Minkowski sum of polytopes itself — lives at Minkowski Sum of Tile Polytopes (Minkowski Representation Hypothesis) (type: geometric-object). Keeping the two separate follows the map’s rule that a hypothesis must not carry the same type as the mathematical object it asserts.

Why it is a hypothesis, not a shape

Finding one activation inside a vector sum of two polytopes is consistent with MRH but does not establish it; MRH is universally quantified over the whole layer and predicts (i) curved geodesics along a token k-NN graph, (ii) convex/archetypal coding matching SAE reconstruction with few archetypes, and (iii) spontaneous block-diagonal co-activation structure. Each is a separate falsifiable consequence.

Key evidence

Fel, Wang, Lepori, Kowal, Lee, Balestriero, Joseph, Lubana, Konkle, Ba & Wattenberg (ICLR 2026, arXiv:2510.08638) report all three signatures on DINOv2-B with a 32k-atom SAE dictionary, while stating the evidence is compatible, not conclusive (“multiple mechanisms can mimic the same surface phenomena”) and running no causal steering test. See dino-activation-space-is-consistent-with-a-minkowski-sum-of-tile-polytopes.

Found in (0 observations · 0 families)

No observations confirm this structure yet.