MATH · IN · MODELS

Task representations are soft ellipsoidal regions, not single vectors

measured in 1 paper

- Fitting a conceptor C = R(R+alpha^-2 I)^-1 - a soft ellipsoidal projection over a task's cached in-context activations - captures a function-vector task as a graded region rather than a single steering direction. [postmus-abreu-2024] - Soft-projecting new activations through the conceptor steers GPT-J-6B and GPT-NeoX-20B toward the task more accurately than additive steering (country-capital top-1 on GPT-J: additive 32.0% vs conceptor 81.6% without mean-centering; 63.9% vs 85.3% with mean-centering). [postmus-abreu-2024] - The improvement holds across the function-vector tasks tested (antonyms, capitalize, country-capital, English-French, present-past), and mean-centering helps both methods. [postmus-abreu-2024] - Two task conceptors combined via Boolean AND steer composite behavior more accurately than averaging the two additive steering vectors on all three task pairs, and additionally beat the additive baseline on one pair (English-French AND antonyms). [postmus-abreu-2024]

Context

function vectors, task representation, soft projection, Boolean composition

Papers

Steering Large Language Models using Conceptors: Improving Addition-Based Activation Engineering — Postmus, Joris, Abreu, Steven2024 · arXiv:2410.16314