Guardedness against a binary probe leaks to a multiclass softmax
measured in 1 paperRavfogel et al. formalize V-guardedness (no classifier in family V predicts concept Z above epsilon mutual information), the property INLP/RLACE/LEACE target [ravfogel-etal-2023-log-linear-guardedness] Theorem 3.2 shows binary log-linear guardedness propagates to downstream binary log-linear classifiers, but Theorem 3.4 shows it breaks for multiclass classifiers [ravfogel-etal-2023-log-linear-guardedness] On a K-Voronoi construction, a representation guarded against every single hyperplane still lets a K-way softmax (signed combinations of the guarding directions) recover Z with near-full information [ravfogel-etal-2023-log-linear-guardedness] Empirically, RLACE-erased BERT gender is perfectly recovered by 4-8-class profession softmax classifiers, so guardedness against one adversary says nothing about a differently-structured one [ravfogel-etal-2023-log-linear-guardedness]