MATH · IN · MODELS

AlterRep push on BERT's RC-boundary subspace shifts agreement errors

measured in 1 paper

Ravfogel et al. introduce AlterRep to test whether BERT causally uses its relative-clause-boundary representation during subject-verb agreement, using INLP-derived classifier directions spanning the RC-boundary subspace [ravfogel-etal-2021-alterrep] In BERT-base middle layers (5-8) the positive counterfactual raises agreement-error probability up to 14 points while the negative lowers it only up to 2, a clearly asymmetric causal effect [ravfogel-etal-2021-alterrep] BERT-large shows the analogous pattern in layers 12-17; smaller Turc et al. BERT variants show it in narrower model-specific ranges [ravfogel-etal-2021-alterrep] The effect partially transfers across five RC types when the subspace is estimated from a different RC type, indicating partly shared, partly structure-specific representation [ravfogel-etal-2021-alterrep] Counterfactuals from 10 random Gaussian subspaces fail to reproduce the effect, confirming specificity to the RC-boundary subspace [ravfogel-etal-2021-alterrep]

Context

AlterRep counterfactual push, random-subspace control, cross-RC-type generalization

Papers

Counterfactual Interventions Reveal the Causal Effect of Relative Clause Representations on Agreement Prediction — Ravfogel, Shauli, Prasad, Grusha, Linzen, Tal, Goldberg, Yoav2021 · arXiv:2105.06965