Linear encoding of features with liar counterexamples
measured in 1 paperFeatures are encoded as linear directions in activation space, and the model performs linear operations over them (e.g. antonym = reflection, negation = negation). The authors also identify "liar" counterexamples where a linear probe disagrees with the model's actual behavior, and show the geometry of activation space is isomorphic to the logical structure of the underlying concepts.
Structure
Context
linear representation hypothesis, feature directions, causality
Confirmed in models
Method
Papers
The Linear Representation Hypothesis and the Geometry of Large Language Models — Park, K., Choe, Y. J., Veitch, V.