MATH · IN · MODELS

INLP fails its own control while Mean Projection and LEACE pass

measured in 1 paper

Dobrzeniecka et al. re-run Elazar's amnesic-probing pipeline on BERT, substituting Mean Projection and LEACE for INLP on three properties (dependency labels, fine-POS, coarse-POS) [dobrzeniecka-etal-2025-mp-leace] INLP needs 738-900 directions to erase these versus a flat one-per-class for MP/LEACE (41/45/12), and its cosine-similarity distortion is far larger (0.31-0.37 vs 0.80-0.95) [dobrzeniecka-etal-2025-mp-leace] This distortion breaks INLP's own control: the same number of random-direction projections drops next-word accuracy MORE than the targeted INLP removal for 2 of 3 properties [dobrzeniecka-etal-2025-mp-leace] Mean Projection and LEACE pass the random-projection and selectivity controls in all cases, so the removal method is not incidental to amnesic probing's validity [dobrzeniecka-etal-2025-mp-leace]

Context

amnesic probing, information control (random-projection baseline), selectivity control, matrix rank / cosine-similarity distortion metrics

Confirmed in models

Papers

Improving Causal Interventions in Amnesic Probing with Mean Projection or LEACE — Dobrzeniecka, Alicja, Fokkens, Antske, Sommerauer, Pia2025 · arXiv:2506.11673