MATH · IN · MODELS

Nonlinear interventions are more reliable than linear subspace erasure

measured in 1 paper

Canby et al. formalize completeness and selectivity as numbers in [0,1] and define reliability as their harmonic mean, comparing INLP, RLACE, AlterRep and three gradient-based nonlinear interventions [canby-etal-2025-reliability] Every method shows a completeness/selectivity trade-off; none achieves perfect completeness without sacrificing selectivity in any model [canby-etal-2025-reliability] Counterfactual methods (AlterRep, gradient-based) achieve near-total task-accuracy change, consistent with higher completeness than the concept-removal methods [canby-etal-2025-reliability] Nonlinear interventions are almost always most reliable (FGSM 0.92-0.96 across GPT-2/Pythia/Llama) versus linear methods 0.106-0.841, the sole exception being BERT layers 10-12 where AlterRep wins [canby-etal-2025-reliability]

Context

completeness (causal probing), selectivity (causal probing, Elazar et al. sense), reliability (harmonic mean of completeness and selectivity), gradient-based interventions (FGSM, PGD, AutoAttack)

Papers

How Reliable are Causal Probing Interventions? — Canby, Marc E., Davies, Adam, Rastogi, Chirag, Hockenmaier, Julia2025 · arXiv:2408.15510