MATH · IN · MODELS

Inverting gender interventions into text augmentation improves fairness

measured in 1 paper

Avitan et al. augment a RoBERTa-base profession classifier's BiasBios training data with string counterfactuals recovered by inverting gender interventions (LEACE erasure, MiMiC/MiMiC+ steering) [avitan-etal-2024-intervention-lens] Training on originals gives accuracy 86.42, F1 79.63, TPR gender gap 14.27; stripping gender from text lowers the gap to 11.22 but costs accuracy [avitan-etal-2024-intervention-lens] Adding LEACE counterfactuals gives the best accuracy (86.59) and F1 (81.8) at gap 12.95; adding MiMiC+ counterfactuals gives the lowest gap (10.59) with near-baseline accuracy [avitan-etal-2024-intervention-lens] So augmenting with inverted representation-intervention counterfactuals improves fairness without the accuracy cost of simply removing gender indicators [avitan-etal-2024-intervention-lens]

Context

true-positive-rate gender gap (fairness metric), counterfactual data augmentation

Confirmed in models

Papers

Intervention Lens: from Representation Surgery to String Counterfactuals — Avitan, Matan, Cotterell, Ryan, Goldberg, Yoav, Ravfogel, Shauli2024 · arXiv:2402.11355