MATH · IN · MODELS

In-context Gaussian-belief posteriors trace curved manifolds linear steering breaks

measured in 1 paper

Sarfati et al. give Llama-3.2 a string of samples from an unknown normal and study how it in-context infers the distribution; layer-14 activations form smooth but genuinely curved 2D manifolds as mu or sigma varies [sarfati-etal-2026-shape-of-beliefs] The output-side simplex geometry (via inPCA) shows the same curved structure dual to the input side [sarfati-etal-2026-shape-of-beliefs] Rather than one global linear probe, local linear field probes tiling the manifold reach 87-99% accuracy, framed as evidence that purely linear concept representations are often an inadequate abstraction [sarfati-etal-2026-shape-of-beliefs] Belief updating after a mid-sequence distribution change shows two-timescale relaxation through two attractor-like regions for the old and new distributions [sarfati-etal-2026-shape-of-beliefs] Linear difference-of-means steering drives activations off the curved manifold, producing out-of-distribution logits, while manifold-respecting geodesic steering preserves the target distribution, so the curvature is causally real [sarfati-etal-2026-shape-of-beliefs]

Context

belief manifolds, Bayesian posterior inference, in-context learning, curvature, local linear probes, linear field probing, attractor dynamics, manifold-respecting steering, linear steering failure mode

Confirmed in models

Papers

The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors — Sarfati, Raphael, Bigelow, Eric, Wurgaft, Daniel, Merullo, Jack, Geiger, Atticus, Lewis, Owen, McGrath, Tom, Lubana, Ekdeep Singh2026 · arXiv:2602.02315