MATH · IN · MODELS

Linear encoding of features with liar counterexamples

measured in 1 paper

Features are encoded as linear directions in activation space, and the model performs linear operations over them (e.g. antonym = reflection, negation = negation). The authors also identify "liar" counterexamples where a linear probe disagrees with the model's actual behavior, and show the geometry of activation space is isomorphic to the logical structure of the underlying concepts.

Context

linear representation hypothesis, feature directions, causality

Papers

The Linear Representation Hypothesis and the Geometry of Large Language Models — Park, K., Choe, Y. J., Veitch, V.2023 · arXiv:2311.03658