MATH · IN · MODELS

OSCaR's graded rotation preserves information better than projection debiasing

measured in 1 paper

Dev et al. measure that a diff-of-means gender direction and an occupation direction are not orthogonal in GloVe or RoBERTa embedding space [dev-etal-2020-oscar] Standard projective debiasing (hard debiasing, INLP) erases the gender direction from every word, destroying valid associations (an NLI entailment drops 97% to 16%) [dev-etal-2020-oscar] OSCaR instead applies a graded rotation, full at the occupation direction and none at the gender direction, so unrelated words are barely perturbed [dev-etal-2020-oscar] On GloVe it matches the best bias reduction (WEAT 1.768 to 0.235) while scoring far higher on new information-retention metrics (WEAT*, SIRT); on RoBERTa it gives the best bias reduction and retention [dev-etal-2020-oscar]

Context

graded/partial rotation (not projection), measured (non-zero) correlation between two concept directions before correction, information-retention vs. bias-removal tradeoff, WEAT* and SIRT (new information-retention metrics)

Papers

OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word Embeddings — Dev, Sunipa, Li, Tao, Phillips, Jeff M., Srikumar, Vivek2020 · arXiv:2007.00049