MATH · IN · MODELS

SAE latents in Gemma-2-9B show preliminary simplex-structured belief geometry

measured in 1 paper

Levinson extends the belief-state-simplex framework from toy transformers to a real LLM's SAE latent space, fitting 13 priority latent clusters in a Gemma-2-9B layer-20 JumpReLU SAE with archetypal simplex fitting against 3 null-cluster controls [levinson-2026] On the barycentric predictive test, 5 of 13 real clusters show a significant advantage over the best single latent versus 0 of 3 null clusters, indicating recovered structure rather than a tiling artifact [levinson-2026] Causal steering along vertex directions gives positive scores for all 8 qualifying real clusters, but null-cluster scores overlap substantially, limiting discriminative power [levinson-2026] Only cluster 768_596 shows joint predictive and causal convergence, and the author frames the evidence as preliminary (single model and layer, modest effect sizes, a phantom vertex) [levinson-2026]

Context

sparse autoencoders, belief state geometry, mixed-state presentation, computational mechanics, archetypal analysis, tiling artifact, causal steering

Confirmed in models

Papers

Finding Belief Geometries with Sparse Autoencoders — Levinson, Matthew2026 · arXiv:2604.02685