MATH · IN · MODELS

LEACE's erasure subspace captures only about half of a concept

measured in 1 paper

Guerner et al. define intrinsic, classifier-free criteria for an "ideal" concept subspace (eraser, encapsulator, contained, stable) using only the model's own counterfactual mutual information [guerner-etal-2023-geometric-causal-probing] This addresses Kumar et al.'s critique that naive mutual-information erasure criteria can be fooled by correlated features, fixed by a counterfactual unigram construction forcing independence [guerner-etal-2023-geometric-causal-probing] Using LEACE to find the candidate subspace in GPT-2-large, the one-dimensional English verbal-number subspace has a subspace-info (encapsulation) ratio of only ~0.52-0.55, about half the concept's information, despite provably erasing all linear classifiability [guerner-etal-2023-geometric-causal-probing] For French grammatical gender the subspace captures even less (~0.34, roughly 30%), a lossier partition [guerner-etal-2023-geometric-causal-probing] A do-intervention on the subspace steers generation, tying the intrinsic geometric criteria to actual causal-intervention outcomes [guerner-etal-2023-geometric-causal-probing]

Context

counterfactual mutual information, erasure / encapsulation / containment / stability (intrinsic subspace criteria), spurious-correlation-robust concept evaluation

Papers

A Geometric Notion of Causal Probing — Guerner, Clément, Svete, Anej, Liu, Tianyu, Warstadt, Alexander, Cotterell, Ryan2023 · arXiv:2307.15054