methods / Causal Validation / Low-Rank Representation Adaptation (LoRRA)
Low-Rank Representation Adaptation (LoRRA)
Trains a low-rank (LoRA) adapter so that a model's own representations, on ordinary inputs, come to match a target representation constructed by adding a contrast vector and/or reading vector to the original activation — a trained alternative to simple inference-time vector addition.