Amnesic probing misleads at low class count and high dimension
measured in 1 paperRozanova et al. apply amnesic probing (INLP erasure then measure) to a natural-logic NLI fragment where the label is provably determined by two known features across five BERT/RoBERTa NLI models [rozanova-etal-2023-interventional] Removing either or both provably-necessary features causes almost no accuracy drop (mostly within +/-2 points; one case improves by 8.99) [rozanova-etal-2023-interventional] Even nulling out the gold entailment label itself produces almost no drop in three of five models, a case where the removed information is by definition maximally relevant [rozanova-etal-2023-interventional] They trace this to a dimensionality confound: with only 2-3 classes, INLP removes very few directions and a single random-direction control is unstable [rozanova-etal-2023-interventional] Removing as few as 3 random directions can already match the targeted removal by chance, so a "no difference from random" conclusion can be control-baseline noise [rozanova-etal-2023-interventional]