RMU corrupts a hazardous-topic direction by fine-tuning while preserving capability
measured in 1 paperLi, Pan et al. (WMDP) introduce RMU, a two-term activation-space loss pushing hazardous-topic activations toward a fixed random unit vector while anchoring benign activations to the frozen model [li-pan-etal-2024-the-wmdp-benchmark-measuring-and-reducing-malicious-use-with-unlearning] On Zephyr-7B-beta, WMDP-Bio drops 63.7 to 31.2 and WMDP-Cyber 44.0 to 28.2, while MMLU (58.1 to 57.1) and MT-Bench (7.33 to 7.10) stay near baseline [li-pan-etal-2024-the-wmdp-benchmark-measuring-and-reducing-malicious-use-with-unlearning] Yi-34B-Chat and Mixtral-8x7B-Instruct show the same pattern (WMDP-Bio 75.3 to 30.7 and 74.8 to 34.0, MMLU preserved within ~2 points) [li-pan-etal-2024-the-wmdp-benchmark-measuring-and-reducing-malicious-use-with-unlearning] The direction is imposed by fine-tuning toward an arbitrary target rather than found in the base model, unlike additive inference-time steering [li-pan-etal-2024-the-wmdp-benchmark-measuring-and-reducing-malicious-use-with-unlearning]