MATH · IN · MODELS

A diff-in-means misalignment direction transfers cross-architecture but non-specifically

measured in 1 paper

Syed QLoRA-finetunes four models identically on insecure code and extracts a diff-in-means misalignment direction, finding 99.6% within-model separability versus 50% for a secure-code control [syed-2026-actionable-activation-directions-emergent-misalignment] A ridge-regression map fit between any two models' activation spaces transfers the direction across all twelve model-family pairs with above-chance accuracy (Gemma-to-Llama 87%, Gemma-to-Ministral 90%) [syed-2026-actionable-activation-directions-emergent-misalignment] Steering a target model with the transferred direction suppresses code-related spillover (delta 13-46 points) but fails a specificity control, unlike the fully specific within-model effect (delta 21-51 points) [syed-2026-actionable-activation-directions-emergent-misalignment] This reveals an asymmetric donor/receiver topology, with Gemma and Qwen acting as donors and Llama as a receiver [syed-2026-actionable-activation-directions-emergent-misalignment]

Context

cross-architecture geometric transfer of a behaviorally-causal direction via a fitted ridge-regression map between two models' activation spaces, a two-tier causal specificity structure distinguishing within-model (causal and specific) from cross-model (causal but non-specific) directions

Papers

Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families — Syed, Abdul Rafay2026 · arXiv:2606.20225