AlterRep push on BERT's RC-boundary subspace shifts agreement errors
measured in 1 paperRavfogel et al. introduce AlterRep to test whether BERT causally uses its relative-clause-boundary representation during subject-verb agreement, using INLP-derived classifier directions spanning the RC-boundary subspace [ravfogel-etal-2021-alterrep] In BERT-base middle layers (5-8) the positive counterfactual raises agreement-error probability up to 14 points while the negative lowers it only up to 2, a clearly asymmetric causal effect [ravfogel-etal-2021-alterrep] BERT-large shows the analogous pattern in layers 12-17; smaller Turc et al. BERT variants show it in narrower model-specific ranges [ravfogel-etal-2021-alterrep] The effect partially transfers across five RC types when the subspace is estimated from a different RC type, indicating partly shared, partly structure-specific representation [ravfogel-etal-2021-alterrep] Counterfactuals from 10 random Gaussian subspaces fail to reproduce the effect, confirming specificity to the RC-boundary subspace [ravfogel-etal-2021-alterrep]