MATH · IN · MODELS

The gender subspace spans dozens of directions, not one

measured in 1 paper

Ravfogel et al. introduce Iterative Null-space Projection (INLP) and apply it across three case studies [ravfogel-etal-2020-inlp] On GloVe, 35 INLP iterations drop SVM gender classification 100%->49.3% while a nonlinear MLP still recovers 85%, showing the gender subspace spans dozens of orthogonal directions rather than the single Bolukbasi direction [ravfogel-etal-2020-inlp] Word-similarity benchmarks improve after projection (SimLex-999 0.373->0.489), so general lexical semantics is preserved [ravfogel-etal-2020-inlp] On DeepMoji race-correlated sentiment, removing the race subspace shrinks the true-positive-rate fairness gap 0.45->0.15 at some accuracy cost [ravfogel-etal-2020-inlp] On Bias-in-Bios, removing the gender subspace from BERT CLS representations (300 directions) cuts the gender-TPR gap 48% while dropping profession accuracy only 80.9%->75.2% [ravfogel-etal-2020-inlp]

Context

orthogonal nullspace projection, fairness gap (TPR-GAP), WEAT, lexical-semantics preservation

Papers

Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection — Ravfogel, Shauli, Elazar, Yanai, Gonen, Hila, Twiton, Michael, Goldberg, Yoav2020 · arXiv:2004.07667