MATH · IN · MODELS

Linear relational concepts invert an LRE into an editable object direction

measured in 1 paper

Chanin et al. invert Hernandez et al.'s linear relational embedding with a low-rank pseudoinverse to build a "linear relational concept" unit vector living in subject-activation space [chanin-etal-2023] On 47 relations in Llama-2-7B and GPT-J-6B, LRCs beat a directly-trained SVM probe on both classification (0.81 vs 0.73-0.75) and causal editing (0.78-0.84 vs 0.69-0.76) [chanin-etal-2023] A deliberately low-rank inverse (rank ~200 of 4096) is essential, and reading the object side from an earlier layer roughly doubles multi-token accuracy [chanin-etal-2023] This is a fourth distinct route to a linear feature direction, derived from relational structure between two token positions rather than a contrastive set at one site [chanin-etal-2023]

Context

relational knowledge, linear representation hypothesis, concept editing, causality

Papers

Identifying Linear Relational Concepts in Large Language Models — Chanin, David, Hunter, Anthony, Camburu, Oana-Maria2023 · arXiv:2311.08968