MATH · IN · MODELS

A learned rotation subspace beats brute-force causal-alignment search

measured in 1 paper

Geiger et al. test DAS on two tasks with known causal structure; on a hierarchical-equality ReLU network its learned orthogonal-rotation subspace reaches interchange-intervention accuracy up to 1.00 vs 0.60 brute-force and 0.73 best localist [geiger-etal-2023-das] On monotonicity NLI (BERT-base fine-tuned on MultiNLI then MoNLI) DAS reaches 1.00 IIA at layer 9 vs 0.64 and 0.51 [geiger-etal-2023-das] A randomly-initialized network stays near chance (~0.50) unless the hidden dimension is blown up to 4096, confirming DAS is not fabricating structure [geiger-etal-2023-das] The paper is explicit that DAS USES an assumed linear structure as a methodological device and does not itself claim to have discovered that any concept is linearly encoded [geiger-etal-2023-das]

Context

interchange intervention, interchange intervention accuracy (IIA), orthogonal-rotation change of basis, localist vs. distributed causal abstraction

Papers

Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations — Geiger, Atticus, Wu, Zhengxuan, Potts, Christopher, Icard, Thomas, Goodman, Noah D.2024 · arXiv:2303.02536