Answer-correctness sits in a 3-8D linear subspace, causally steerable
measured in 1 paperCho et al. study self-assessed correctness on TruthfulQA pairs across 9 models spanning 5 families, finding the discriminative signal is a low-dimensional linear subspace, not a curved manifold [cho-etal-2026-confidence-manifold] A PLS sweep peaks at 3-8 dimensions (e.g. peak AUC 0.90 at layer 23, dim 5, Mistral-7B), and nonlinear classifiers give no gain over a linear boundary [cho-etal-2026-confidence-manifold] Correct/incorrect classes form roughly Gaussian clusters separated by a mean shift: a two-mean centroid detector matches the linear probe (0.90 vs 0.89 AUC), robust with 25 examples/class [cho-etal-2026-confidence-manifold] This 3-8D discriminative subspace is lower-dimensional than the representation's 8-12D intrinsic dimension at the same layers [cho-etal-2026-confidence-manifold] Adding the learned direction shifts downstream error rate up to 10.9 points dose-dependently, while random and orthogonal controls have no reliable effect [cho-etal-2026-confidence-manifold]