A diff-in-means misalignment direction transfers cross-architecture but non-specifically
measured in 1 paperSyed QLoRA-finetunes four models identically on insecure code and extracts a diff-in-means misalignment direction, finding 99.6% within-model separability versus 50% for a secure-code control [syed-2026-actionable-activation-directions-emergent-misalignment] A ridge-regression map fit between any two models' activation spaces transfers the direction across all twelve model-family pairs with above-chance accuracy (Gemma-to-Llama 87%, Gemma-to-Ministral 90%) [syed-2026-actionable-activation-directions-emergent-misalignment] Steering a target model with the transferred direction suppresses code-related spillover (delta 13-46 points) but fails a specificity control, unlike the fully specific within-model effect (delta 21-51 points) [syed-2026-actionable-activation-directions-emergent-misalignment] This reveals an asymmetric donor/receiver topology, with Gemma and Qwen acting as donors and Llama as a receiver [syed-2026-actionable-activation-directions-emergent-misalignment]