MATH · IN · MODELS

A Manifold Probe recovers causally-used multi-dimensional time and space manifolds

measured in 1 paper

Modell introduces the Manifold Probe, which jointly learns (via a generalized eigenvalue problem over a spline basis) the space of a concept's linearly-decodable features and the orthonormal directions encoding them [modell-2026-manifold-probe] Applied to Llama-2-7B for release dates and geographic coordinates, it surfaces many more decodable features than the raw concept value, with the top SPACE feature more precisely decodable (higher test R^2) than latitude or longitude, while the top time feature is roughly identical to the year [modell-2026-manifold-probe] After Varimax rotation, top time features separate individual decades (1950s-2010s) and top space features localize on individual US states [modell-2026-manifold-probe] Treating the learned manifold as a continuum of steering vectors causally shifts the model's stated release year toward a target (peaking at layers 8 and 14), so the manifold is causally used, not merely decodable [modell-2026-manifold-probe]

Context

manifold probe, superposition, multi-dimensional feature discovery, Varimax rotation, decade/state interpretability, causal steering along a manifold

Confirmed in models

Papers

Probing for Representation Manifolds in Superposition — Modell, Alexander2026 · arXiv:2605.18537