MATH · IN · MODELS

Hypotheses

background definition

Explanatory frameworks predicting which geometric structures should appear in activation space, and why — general claims, as opposed to a single specific structure found in one place.

Definition

A hypothesis in this map is a general, falsifiable claim of the form “representations of type XX are organized according to geometric/statistical principle YY” — a predictive framework, rather than the description of one specific observed shape. A hypothesis is supported or weakened by the accumulation of individual structural findings (circles, subspaces, cones, …), not established by a single one.

Intuition

A specific structure (e.g. “this cyclic concept forms a circle”) is a single data point. A hypothesis (e.g. “features are generally encoded as linear directions”) is a claim about the pattern across many such data points — closer to a scientific theory than to a single experimental result, and correspondingly harder to confirm and easier to falsify with one clean counterexample.

Properties

  • Falsifiability asymmetry. A single confirmed counterexample can weaken or refute a hypothesis; no finite number of confirming instances can fully prove a universally-quantified hypothesis (“all features are linear directions”) — only accumulate evidence for it.
  • Hypotheses can be mutually reinforcing or in tension. Two hypotheses can predict the same observations for overlapping reasons (reinforcing), or one’s typical mechanism can undercut the clean statement of another (tension) — see the relationships table below for concrete instances in this map.
  • A hypothesis explains why; a structure records what. Torus, Circle, and the other structure nodes describe specific claimed shapes; a hypothesis node instead proposes a generative mechanism (a statistical, information-theoretic, or representational principle) that would produce some family of shapes as a consequence.
  • Scope varies. Some hypotheses are domain-general (claiming something about all features of a certain kind), others are narrower (claiming something about one specific class of concept, e.g. only cyclic ones, or only in-context belief updating) — narrower scope is easier to confirm but explains less if true.
  • Hypotheses can reinforce or sit in tension. Linear Representation Hypothesis and Platonic Representation Hypothesis are mutually reinforcing (convergent, linear representations across models explain each other); a naive reading of superposition — the idea that a network packs more features than dimensions — is in tension with Linear Representation Hypothesis, since dd dimensions cannot hold d\gg d mutually orthogonal directions, though Constructive Interference Hypothesis shows one concrete, narrower mechanism by which correlated (not fully independent) packed features can remain linearly recoverable without needing near-orthogonality at all; Translation Symmetry Hypothesis and Efficient/Capacity-Optimal Coding Hypothesis are competing, not reinforcing, explanations for the same observed shape (1D continuum manifold), precisely because they make different predictions about when that shape should appear.

Exercises

Base

  1. A paper reports one instance of a linearly-decodable feature in one model. Does this alone establish the Linear Representation Hypothesis as stated (“for each human-interpretable feature there exists such a direction”)? Why or why not?
Solution

No — the hypothesis is universally quantified over “each human-interpretable feature,” so one confirmed instance is consistent with it but does not establish it; a single counterexample (one feature demonstrably not linearly encoded, e.g. requiring genuinely non-linear decoding with no linear proxy) would be needed to refute it, but no finite number of positive instances proves a universal claim.

  1. Give one general reason a hypothesis about representational geometry could be true “in the limit” (e.g. as model scale grows) but false for small models, and explain why this makes falsification harder.
Solution

If a hypothesis is only claimed to hold asymptotically (e.g. “sufficiently large/capable models converge to similar representations”), then failing to observe it in a small model is not a counterexample — the hypothesis has an implicit escape clause tied to scale. This makes falsification harder because a negative result can always be attributed to insufficient scale rather than to the hypothesis being wrong, unless the hypothesis is stated with an explicit, checkable scale threshold or rate.

Middle

  1. Suppose Hypothesis H1H_1 (“all cyclic concepts are encoded as circles”) and Hypothesis H2H_2 (“all circle-encoded concepts are cyclic”) are both proposed. Are these logically equivalent? If not, give an example distinguishing them.
Solution

Not equivalent — H1H_1 is “cyclic \Rightarrow circle,” H2H_2 is its converse, “circle \Rightarrow cyclic.” A concept could be encoded as a circle without being conceptually cyclic (e.g. if some non-cyclic quantity happened to be represented via a bounded, wraparound-free encoding that geometrically resembles a circle, or via a closed loop for an unrelated reason such as an artifact of the training objective) — this would satisfy the circle-shape observation without H1H_1‘s antecedent (cyclicity) being the reason, refuting H2H_2 while leaving H1H_1 untouched. Conversely, a genuinely cyclic concept might fail to be encoded as a clean circle at all (e.g. encoded as a cone with a poorly-resolved angular component, or not linearly decodable at all), refuting H1H_1 while H2H_2 (vacuously, if there are no circle-encoded concepts to check) remains unfalsified. The two directions require independent evidence.

  1. A hypothesis HH predicts shape SS should appear for concept class CC. A study finds shape SS for a concept not in CC. Does this refute HH? Formalize HH as an implication to justify your answer.
Solution

No — if HH is formalized as cCshape(c)=Sc \in C \Rightarrow \text{shape}(c) = S (a one-directional implication), finding SS for some cCc \notin C says nothing about the truth value of the implication for elements of CC; the implication is vacuously unconstrained outside CC. This would only be a problem for a different, stronger hypothesis of the form shape(c)=ScC\text{shape}(c)=S \Leftrightarrow c\in C (the biconditional/“only if” version), which does make a claim about concepts outside CC.

Pro

  1. Two hypotheses H1,H2H_1, H_2 each independently entail observation OO (i.e. H1OH_1\Rightarrow O and H2OH_2\Rightarrow O). A study confirms OO. What, precisely, can be concluded about H1H_1 and H2H_2, and what is the specific logical fallacy in concluding ”OO confirms H1H_1 and disconfirms nothing about H2H_2‘s specific additional claims beyond OO”?
Solution

Confirming OO is consistent with H1H_1, consistent with H2H_2, consistent with both, and consistent with neither (some third explanation H3H_3 might also entail OO) — observing a shared consequence of multiple hypotheses cannot, by itself, discriminate between them (affirming the consequent: from HOH\Rightarrow O and OO, one cannot validly conclude HH). To discriminate H1H_1 from H2H_2, one needs an observation that the two hypotheses predict differently — i.e. some OO' with H1OH_1\Rightarrow O' but H2¬OH_2 \Rightarrow \neg O' (or H2H_2 simply silent on OO'while H1H_1 commits to it). This is exactly the situation with Translation Symmetry Hypothesis and Efficient/Capacity-Optimal Coding Hypothesis: both predict the same rippled-manifold shape OO, so confirming the shape alone cannot yet tell the two apart — discriminating them requires finding a domain or manipulation where they diverge (e.g. a bounded scalar with no corpus-statistics origin at all, which only one of the two mechanisms could produce the ripple for).

  1. Formalize “Hypothesis H2H_2 is strictly more general than H1H_1” set-theoretically (in terms of the set of concepts/situations each governs), and show that if H2H_2 is strictly more general than H1H_1 and both are true, then H2H_2 is the more informative hypothesis to have confirmed — while a failure of H2H_2 in some new domain is more informative than a failure of H1H_1 in that same domain.
Solution

Say H1H_1 governs a set of situations D1D_1 (makes a definite prediction on each) and H2H_2 governs D2D1D_2 \supsetneq D_1, agreeing with H1H_1‘s predictions on all of D1D_1. If both are true (correct on their respective domains), confirming H2H_2 automatically confirms H1H_1 as a special case, plus additional confirmed predictions on D2D1D_2\setminus D_1 that H1H_1 is simply silent about — so H2H_2-confirmed is (weakly) more informative, having “for free” checked everything H1H_1 checks, plus more. Conversely, if a new situation dD2D1d\in D_2\setminus D_1 is found where the predicted shape fails to appear: this refutes H2H_2 (which claimed to govern dd) but says nothing about H1H_1 (which never made a claim about dd in the first place, since dD1d\notin D_1) — so a failure at dd is informative specifically about the broader hypothesis H2H_2 and leaves the narrower H1H_1 completely intact. This formalizes why a broader-scope hypothesis is simultaneously more valuable when it holds and more exposed to refutation — the standard risk/reward trade-off of generality, made precise via domain inclusion.

Found in (0 observations · 0 families)

No observations confirm this structure yet.