Safety gradients occupy a low-rank subspace orthogonal to utility
measured in 1 paperZhang et al. SVD per-layer safety- versus utility-task gradients during fine-tuning, finding safety gradients occupy a compact low-rank subspace with sharp singular-value decay while utility spans much higher dimension [zhang-etal-2026-understanding-and-preserving-safety-in-fine-tuned-llms] Safety and utility gradient directions have cosine similarity oscillating around zero and often negative (directional conflict) [zhang-etal-2026-understanding-and-preserving-safety-in-fine-tuned-llms] Their Safety-Preserving Fine-tuning projects utility gradients onto the safety subspace orthogonal complement during training [zhang-etal-2026-understanding-and-preserving-safety-in-fine-tuned-llms] This cuts attack success rate 0.955 to 0.019 (Llama-3.1-8B), 0.985 to 0.240 (Mistral) and 0.988 to 0.124 (Qwen2.5-7B) while preserving MMLU, robust to deep fine-tuning and multiple jailbreaks [zhang-etal-2026-understanding-and-preserving-safety-in-fine-tuned-llms]