MATH · IN · MODELS

Splicing probed entity-state representations causally shifts generation

measured in 1 paper

Li, Nye & Andreas train linear probes to decode entity state (Alchemy beaker contents, TextWorld room/object state) from fine-tuned BART and T5 encoder representations [li-nye-andreas-2021-implicit-representations-of-meaning] They go beyond decoding with a causal splice WITHIN each model (not between BART and T5): replacing one context's encoded beaker-2 description with another context's encoding to build a synthetic "mixed" state [li-nye-andreas-2021-implicit-representations-of-meaning] Generating from the spliced representation lands in the human-predicted continuation set 57.7% (BART) / 75.4% (T5) of the time, versus 20.4%/37.9% and 16.1%/29.1% for the unmixed source contexts [li-nye-andreas-2021-implicit-representations-of-meaning] The probe only reads the representation while the splice is the intervention, so this is a genuine geometry-tied causal effect [li-nye-andreas-2021-implicit-representations-of-meaning]

Context

entity-state representation splicing, linear probe intervention, implicit world-state tracking

Confirmed in models

Papers

Implicit Representations of Meaning in Neural Language Models — Li, Belinda Z., Nye, Maxwell, Andreas, Jacob2021 · arXiv:2106.00737