DINOv2 activation space is consistent with a Minkowski sum of tile polytopes
argued, not measuredFel et al. propose the Minkowski Representation Hypothesis: a layer's activation space is the convex-geometry Minkowski sum of disjoint archetype-tile polytopes, with any activation a block-sparse convex combination over a few active tiles [fel-etal-2025-rabbit-hull] Three formal results derive this geometry as a mechanistic consequence of multi-head attention's algebra: each head's output lies in the convex hull of its value vectors and multi-head aggregation sums those hulls, exactly a Minkowski sum [fel-etal-2025-rabbit-hull] On DINOv2-B with a 32,000-atom SAE (k=8, R^2>88% on 1.4M ImageNet images), three tests agree: k-NN-graph paths stay on-manifold while straight interpolations leave it, Archetypal Analysis matches SAE quality with ~10 archetypes, and coefficients form block-diagonal co-activation clusters [fel-etal-2025-rabbit-hull] The authors are explicit this is compatible evidence, not proof (status hypothesis), and run no first-party causal intervention [fel-etal-2025-rabbit-hull]