MATH · IN · MODELS

One adversarial direction erases gender where INLP needs dozens

measured in 1 paper

Ravfogel et al. formulate concept erasure as a minimax game between a rank-k projection and a re-optimizing linear predictor, proving how many dimensions the erasure-optimal subspace needs (R-LACE) [ravfogel-etal-2022-rlace] On GloVe, rank-1 R-LACE drops SVM gender accuracy to chance while INLP fails even after removing a 20-dimensional subspace [ravfogel-etal-2022-rlace] On BERT Bias-in-Bios, rank-1 R-LACE drops gender accuracy 99.32% to 52.48% while INLP rank-1 barely moves it, needing ~100 dimensions to match [ravfogel-etal-2022-rlace] The gap is task-dependent: INLP is provably suboptimal for linear-regression objectives but provably identical to R-LACE for Rayleigh-quotient (PLS/CCA) objectives; both are linear-only (nonlinear classifiers still recover gender >90%) [ravfogel-etal-2022-rlace]

Context

minimax game, Fantope relaxation, provable subspace optimality, adversarial training instability

Papers

Linear Adversarial Concept Erasure — Ravfogel, Shauli, Twiton, Michael, Goldberg, Yoav, Cotterell, Ryan2022 · arXiv:2201.12091