MATH · IN · MODELS
methods / Causal Validation / Low-Rank Representation Adaptation (LoRRA)

Low-Rank Representation Adaptation (LoRRA)

Techniqueadvanced

Trains a low-rank (LoRA) adapter so that a model's own representations, on ordinary inputs, come to match a target representation constructed by adding a contrast vector and/or reading vector to the original activation — a trained alternative to simple inference-time vector addition.

Used in (1 observation)

structure: Linear Direction · models: Llama-2-7B-Chat, Llama-2-13B-Chat, Llama-2-70B-Chat, Vicuna-13B, Vicuna-33B-Uncensored, DeBERTa-xxlarge-v2-MNLI · paper: Representation Engineering: A Top-Down Approach to AI Transparency