MATH · IN · MODELS

Emergent misalignment rotates an LLM's truth direction

measured in 1 paper

Sturgeon, Africa & Black fit L2-regularized logistic-regression truth-direction probes on Llama-3.3-70B-Instruct and Qwen3-8B, explicitly rejecting mass-mean (diff-in-means) probing [sturgeon-etal-2026-when-roleplaying-do-models-believe-what-they-say] Persona SFT shifts the truth-probe score by only +0.05 and Open Character Training by +0.089-0.124 [sturgeon-etal-2026-when-roleplaying-do-models-believe-what-they-say] Emergent Misalignment produces a much larger +0.28 shift (56% defend rate, 82% downstream-reasoning consistency) [sturgeon-etal-2026-when-roleplaying-do-models-believe-what-they-say] EM also rotates the truth direction itself (cosine ~0.58), indicating restructured truth-representation geometry rather than local override [sturgeon-etal-2026-when-roleplaying-do-models-believe-what-they-say] The analysis is observational, with no steering or causal intervention performed [sturgeon-etal-2026-when-roleplaying-do-models-believe-what-they-say]

Context

truthfulness, persona, emergent-misalignment

Papers

When Roleplaying, Do Models Believe What They Say? — Sturgeon, Bethan, Africa, Katie, Black, Sid2026 · arXiv:2606.11502