MATH · IN · MODELS

The truth direction rotates and rescales under added context and steers labels

measured in 1 paper

Adarsh et al. extract a per-layer mean-difference truth direction in Llama-3.1-8B-Instruct, Mistral-Nemo-12B-Instruct, Qwen3-4B-Instruct, and SmolLM3-3B and measure how it transforms when a statement is embedded in context [adarsh-etal-2026-how-context-shapes-truth] The angle between with- and without-context truth vectors is near-orthogonal in early layers and converges by mid-depth, while the magnitude ratio is quantified per dataset [adarsh-etal-2026-how-context-shapes-truth] This is the first quantified characterization of how the truth direction's own geometry, not just its existence, changes under context [adarsh-etal-2026-how-context-shapes-truth] Mass-mean steering with the extracted vector flips truthfulness labels in ~100% of cases for three models (Qwen3-4B weaker and variable at 11-59%), a strong replicated causal effect [adarsh-etal-2026-how-context-shapes-truth]

Context

truth direction rotates toward alignment as depth increases when context is added, context rescales the truth direction's magnitude, quantified per dataset, mass-mean steering with the extracted vector causally flips truthfulness labels

Papers

How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs — Adarsh, Shivam, Maistro, Maria, Lioma, Christina2026 · arXiv:2601.06599