A diff-in-means misalignment direction transfers across Qwen-14B fine-tunes
measured in 1 paperSoligo et al. extract a diff-in-means direction (misaligned minus aligned mean activation) in real Qwen2.5-14B-Instruct fine-tuned to be emergently misaligned [soligo-etal-2025-convergent-linear-representations-emergent-misalignment] Adding it to the base aligned chat model causally induces misaligned behavior, while ablating it significantly reduces misalignment in the EM model and in independently misaligned fine-tunes [soligo-etal-2025-convergent-linear-representations-emergent-misalignment] The direction transfers between different Qwen-14B EM fine-tunes, evidencing a convergence in their representations of emergent misalignment [soligo-etal-2025-convergent-linear-representations-emergent-misalignment] A rank-1 LoRA independently trained to induce EM yields a direction at cosine only 0.04 to the mean-diff direction, yet the two interventions' behavioral effects converge [soligo-etal-2025-convergent-linear-representations-emergent-misalignment]