A learned rotation subspace beats brute-force causal-alignment search
measured in 1 paperGeiger et al. test DAS on two tasks with known causal structure; on a hierarchical-equality ReLU network its learned orthogonal-rotation subspace reaches interchange-intervention accuracy up to 1.00 vs 0.60 brute-force and 0.73 best localist [geiger-etal-2023-das] On monotonicity NLI (BERT-base fine-tuned on MultiNLI then MoNLI) DAS reaches 1.00 IIA at layer 9 vs 0.64 and 0.51 [geiger-etal-2023-das] A randomly-initialized network stays near chance (~0.50) unless the hidden dimension is blown up to 4096, confirming DAS is not fabricating structure [geiger-etal-2023-das] The paper is explicit that DAS USES an assumed linear structure as a methodological device and does not itself claim to have discovered that any concept is linearly encoded [geiger-etal-2023-das]