methods / Causal Validation / Causal interventions (steering) / Representation misdirection (fine-tuning toward a target direction)
Representation misdirection (fine-tuning toward a target direction)
Rather than adding a direction to activations at inference time, fine-tunes the model itself so that a targeted subset of activations (or SAE-latent coordinates) is pushed toward an arbitrary/opposite-class target region while a retain-set loss anchors unrelated activations near their original values — used to unlearn a capability by corrupting its internal representation rather than its weights' factual content directly.