LEACE's erasure subspace captures only about half of a concept
measured in 1 paperGuerner et al. define intrinsic, classifier-free criteria for an "ideal" concept subspace (eraser, encapsulator, contained, stable) using only the model's own counterfactual mutual information [guerner-etal-2023-geometric-causal-probing] This addresses Kumar et al.'s critique that naive mutual-information erasure criteria can be fooled by correlated features, fixed by a counterfactual unigram construction forcing independence [guerner-etal-2023-geometric-causal-probing] Using LEACE to find the candidate subspace in GPT-2-large, the one-dimensional English verbal-number subspace has a subspace-info (encapsulation) ratio of only ~0.52-0.55, about half the concept's information, despite provably erasing all linear classifiability [guerner-etal-2023-geometric-causal-probing] For French grammatical gender the subspace captures even less (~0.34, roughly 30%), a lossier partition [guerner-etal-2023-geometric-causal-probing] A do-intervention on the subspace steers generation, tying the intrinsic geometric criteria to actual causal-intervention outcomes [guerner-etal-2023-geometric-causal-probing]