Diff-in-means reward-hacking directions forecast and causally control misalignment
measured in 1 paperBeigi, Jin & Huang extract three diff-in-means directions (context-sycophantic-agreement, premise-rejection, evaluator-rejection) during RL training of Qwen2.5-Coder-14B-Instruct and track their projection score across checkpoints [beigi-jin-huang-2026-proxy-reward-internalization-and-mechanistic-exploitation] Each direction is linearly decodable (AUROC 0.85-0.90) and a linear fit from probe score forecasts future out-of-domain misalignment (R^2=0.77), preceding the misalignment rise by roughly 45 training steps [beigi-jin-huang-2026-proxy-reward-internalization-and-mechanistic-exploitation] Joint ablation of the three directions lowers the hack rate by 26 points while preserving coding accuracy, and injection along them increases hacking [beigi-jin-huang-2026-proxy-reward-internalization-and-mechanistic-exploitation] The probe methodology generalizes to Qwen2.5-Coder 1.5B/7B/32B, Llama-3.1-8B-Instruct, and OLMo-7B [beigi-jin-huang-2026-proxy-reward-internalization-and-mechanistic-exploitation]