Definition
A hypothesis in this map is a general, falsifiable claim of the form “representations of type are organized according to geometric/statistical principle ” — a predictive framework, rather than the description of one specific observed shape. A hypothesis is supported or weakened by the accumulation of individual structural findings (circles, subspaces, cones, …), not established by a single one.
Intuition
A specific structure (e.g. “this cyclic concept forms a circle”) is a single data point. A hypothesis (e.g. “features are generally encoded as linear directions”) is a claim about the pattern across many such data points — closer to a scientific theory than to a single experimental result, and correspondingly harder to confirm and easier to falsify with one clean counterexample.
Properties
- Falsifiability asymmetry. A single confirmed counterexample can weaken or refute a hypothesis; no finite number of confirming instances can fully prove a universally-quantified hypothesis (“all features are linear directions”) — only accumulate evidence for it.
- Hypotheses can be mutually reinforcing or in tension. Two hypotheses can predict the same observations for overlapping reasons (reinforcing), or one’s typical mechanism can undercut the clean statement of another (tension) — see the relationships table below for concrete instances in this map.
- A hypothesis explains why; a structure records what. Torus, Circle, and the other structure nodes describe specific claimed shapes; a hypothesis node instead proposes a generative mechanism (a statistical, information-theoretic, or representational principle) that would produce some family of shapes as a consequence.
- Scope varies. Some hypotheses are domain-general (claiming something about all features of a certain kind), others are narrower (claiming something about one specific class of concept, e.g. only cyclic ones, or only in-context belief updating) — narrower scope is easier to confirm but explains less if true.
- Hypotheses can reinforce or sit in tension. Linear Representation Hypothesis and Platonic Representation Hypothesis are mutually reinforcing (convergent, linear representations across models explain each other); a naive reading of superposition — the idea that a network packs more features than dimensions — is in tension with Linear Representation Hypothesis, since dimensions cannot hold mutually orthogonal directions, though Constructive Interference Hypothesis shows one concrete, narrower mechanism by which correlated (not fully independent) packed features can remain linearly recoverable without needing near-orthogonality at all; Translation Symmetry Hypothesis and Efficient/Capacity-Optimal Coding Hypothesis are competing, not reinforcing, explanations for the same observed shape (1D continuum manifold), precisely because they make different predictions about when that shape should appear.
Exercises
Base
- A paper reports one instance of a linearly-decodable feature in one model. Does this alone establish the Linear Representation Hypothesis as stated (“for each human-interpretable feature there exists such a direction”)? Why or why not?
Solution
No — the hypothesis is universally quantified over “each human-interpretable feature,” so one confirmed instance is consistent with it but does not establish it; a single counterexample (one feature demonstrably not linearly encoded, e.g. requiring genuinely non-linear decoding with no linear proxy) would be needed to refute it, but no finite number of positive instances proves a universal claim.
- Give one general reason a hypothesis about representational geometry could be true “in the limit” (e.g. as model scale grows) but false for small models, and explain why this makes falsification harder.
Solution
If a hypothesis is only claimed to hold asymptotically (e.g. “sufficiently large/capable models converge to similar representations”), then failing to observe it in a small model is not a counterexample — the hypothesis has an implicit escape clause tied to scale. This makes falsification harder because a negative result can always be attributed to insufficient scale rather than to the hypothesis being wrong, unless the hypothesis is stated with an explicit, checkable scale threshold or rate.
Middle
- Suppose Hypothesis (“all cyclic concepts are encoded as circles”) and Hypothesis (“all circle-encoded concepts are cyclic”) are both proposed. Are these logically equivalent? If not, give an example distinguishing them.
Solution
Not equivalent — is “cyclic circle,” is its converse, “circle cyclic.” A concept could be encoded as a circle without being conceptually cyclic (e.g. if some non-cyclic quantity happened to be represented via a bounded, wraparound-free encoding that geometrically resembles a circle, or via a closed loop for an unrelated reason such as an artifact of the training objective) — this would satisfy the circle-shape observation without ‘s antecedent (cyclicity) being the reason, refuting while leaving untouched. Conversely, a genuinely cyclic concept might fail to be encoded as a clean circle at all (e.g. encoded as a cone with a poorly-resolved angular component, or not linearly decodable at all), refuting while (vacuously, if there are no circle-encoded concepts to check) remains unfalsified. The two directions require independent evidence.
- A hypothesis predicts shape should appear for concept class . A study finds shape for a concept not in . Does this refute ? Formalize as an implication to justify your answer.
Solution
No — if is formalized as (a one-directional implication), finding for some says nothing about the truth value of the implication for elements of ; the implication is vacuously unconstrained outside . This would only be a problem for a different, stronger hypothesis of the form (the biconditional/“only if” version), which does make a claim about concepts outside .
Pro
- Two hypotheses each independently entail observation (i.e. and ). A study confirms . What, precisely, can be concluded about and , and what is the specific logical fallacy in concluding ” confirms and disconfirms nothing about ‘s specific additional claims beyond ”?
Solution
Confirming is consistent with , consistent with , consistent with both, and consistent with neither (some third explanation might also entail ) — observing a shared consequence of multiple hypotheses cannot, by itself, discriminate between them (affirming the consequent: from and , one cannot validly conclude ). To discriminate from , one needs an observation that the two hypotheses predict differently — i.e. some with but (or simply silent on while commits to it). This is exactly the situation with Translation Symmetry Hypothesis and Efficient/Capacity-Optimal Coding Hypothesis: both predict the same rippled-manifold shape , so confirming the shape alone cannot yet tell the two apart — discriminating them requires finding a domain or manipulation where they diverge (e.g. a bounded scalar with no corpus-statistics origin at all, which only one of the two mechanisms could produce the ripple for).
- Formalize “Hypothesis is strictly more general than ” set-theoretically (in terms of the set of concepts/situations each governs), and show that if is strictly more general than and both are true, then is the more informative hypothesis to have confirmed — while a failure of in some new domain is more informative than a failure of in that same domain.
Solution
Say governs a set of situations (makes a definite prediction on each) and governs , agreeing with ‘s predictions on all of . If both are true (correct on their respective domains), confirming automatically confirms as a special case, plus additional confirmed predictions on that is simply silent about — so -confirmed is (weakly) more informative, having “for free” checked everything checks, plus more. Conversely, if a new situation is found where the predicted shape fails to appear: this refutes (which claimed to govern ) but says nothing about (which never made a claim about in the first place, since ) — so a failure at is informative specifically about the broader hypothesis and leaves the narrower completely intact. This formalizes why a broader-scope hypothesis is simultaneously more valuable when it holds and more exposed to refutation — the standard risk/reward trade-off of generality, made precise via domain inclusion.