MATH · IN · MODELS

Guardedness against a binary probe leaks to a multiclass softmax

measured in 1 paper

Ravfogel et al. formalize V-guardedness (no classifier in family V predicts concept Z above epsilon mutual information), the property INLP/RLACE/LEACE target [ravfogel-etal-2023-log-linear-guardedness] Theorem 3.2 shows binary log-linear guardedness propagates to downstream binary log-linear classifiers, but Theorem 3.4 shows it breaks for multiclass classifiers [ravfogel-etal-2023-log-linear-guardedness] On a K-Voronoi construction, a representation guarded against every single hyperplane still lets a K-way softmax (signed combinations of the guarding directions) recover Z with near-full information [ravfogel-etal-2023-log-linear-guardedness] Empirically, RLACE-erased BERT gender is perfectly recovered by 4-8-class profession softmax classifiers, so guardedness against one adversary says nothing about a differently-structured one [ravfogel-etal-2023-log-linear-guardedness]

Context

V-guardedness (general formalism), K-Voronoi distribution (worst-case construction), binary vs. multiclass adversary asymmetry

Confirmed in models

Papers

Log-linear Guardedness and its Implications — Ravfogel, Shauli, Goldberg, Yoav, Cotterell, Ryan2023 · arXiv:2210.10012