MATH · IN · MODELS

RMU corrupts a hazardous-topic direction by fine-tuning while preserving capability

measured in 1 paper

Li, Pan et al. (WMDP) introduce RMU, a two-term activation-space loss pushing hazardous-topic activations toward a fixed random unit vector while anchoring benign activations to the frozen model [li-pan-etal-2024-the-wmdp-benchmark-measuring-and-reducing-malicious-use-with-unlearning] On Zephyr-7B-beta, WMDP-Bio drops 63.7 to 31.2 and WMDP-Cyber 44.0 to 28.2, while MMLU (58.1 to 57.1) and MT-Bench (7.33 to 7.10) stay near baseline [li-pan-etal-2024-the-wmdp-benchmark-measuring-and-reducing-malicious-use-with-unlearning] Yi-34B-Chat and Mixtral-8x7B-Instruct show the same pattern (WMDP-Bio 75.3 to 30.7 and 74.8 to 34.0, MMLU preserved within ~2 points) [li-pan-etal-2024-the-wmdp-benchmark-measuring-and-reducing-malicious-use-with-unlearning] The direction is imposed by fine-tuning toward an arbitrary target rather than found in the base model, unlike additive inference-time steering [li-pan-etal-2024-the-wmdp-benchmark-measuring-and-reducing-malicious-use-with-unlearning]

Context

fine-tuning-based (rather than inference-time) directional activation-space corruption toward an arbitrary target direction, contrasted with additive/ablative steering that requires no weight change, a retain-set anchoring loss that preserves general capability while a forget-set loss corrupts a targeted representation

Papers

The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning — Li, Nathaniel, Pan, Alexander, Gopal, Anjali, Yue, Summer, Berrios, Daniel, Wang, Alexandr, Hendrycks, Dan2024 · arXiv:2403.03218