The truth direction rotates and rescales under added context and steers labels
measured in 1 paperAdarsh et al. extract a per-layer mean-difference truth direction in Llama-3.1-8B-Instruct, Mistral-Nemo-12B-Instruct, Qwen3-4B-Instruct, and SmolLM3-3B and measure how it transforms when a statement is embedded in context [adarsh-etal-2026-how-context-shapes-truth] The angle between with- and without-context truth vectors is near-orthogonal in early layers and converges by mid-depth, while the magnitude ratio is quantified per dataset [adarsh-etal-2026-how-context-shapes-truth] This is the first quantified characterization of how the truth direction's own geometry, not just its existence, changes under context [adarsh-etal-2026-how-context-shapes-truth] Mass-mean steering with the extracted vector flips truthfulness labels in ~100% of cases for three models (Qwen3-4B weaker and variable at 11-59%), a strong replicated causal effect [adarsh-etal-2026-how-context-shapes-truth]