MATH · IN · MODELS
methods / Causal Validation / Causal interventions (steering) / Representation misdirection (fine-tuning toward a target direction)

Representation misdirection (fine-tuning toward a target direction)

Techniqueintermediate

Rather than adding a direction to activations at inference time, fine-tunes the model itself so that a targeted subset of activations (or SAE-latent coordinates) is pushed toward an arbitrary/opposite-class target region while a retain-set loss anchors unrelated activations near their original values — used to unlearn a capability by corrupting its internal representation rather than its weights' factual content directly.

Used in (2 observations)

structure: Linear Direction · models: Zephyr-7B-Beta, Yi-34B-Chat, Mixtral-8x7B-Instruct-v0.1 · paper: The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
structure: Linear Direction · models: Llama-3.1-8B-Instruct, Gemma-2-9B-it · paper: Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models