Binding-ID vectors form a subspace whose distances predict confusability
measured in 1 paperFeng & Steinhardt identify additive "binding ID vectors" attached to entity and attribute activations that solve variable binding, using causal interventions on LLaMA-1 (30B primary, plus 13B and 65B) and the Pythia family [feng-steinhardt-2023-how-do-language-models-bind-entities-in-context] The binding vectors occupy a continuous subspace in which the distance between two binding vectors predicts how often the model confuses the corresponding entity-attribute bindings [feng-steinhardt-2023-how-do-language-models-bind-entities-in-context] Patching, adding or removing binding-ID vectors changes which attribute the model retrieves for a given entity, demonstrated via causal mediation analysis rather than linear probing [feng-steinhardt-2023-how-do-language-models-bind-entities-in-context]