A single closed-form (non-iterative, non-gradient-based) oblique projection that provably prevents every linear classifier — not just one trained one — from recovering a concept, while provably minimizing the least-squares change to the embedding; removes exactly rank(Σ_XZ) dimensions, and is the first method in this map to show orthogonal projections (as INLP and RLACE both assume) are themselves suboptimal for the minimal-edit objective.
Used in (7 observations)
structure: Linear Subspace · models: GPT-2-Large, GPT2-base-french · paper: A Geometric Notion of Causal Probing
structure: Linear Subspace · models: RoBERTa-base · paper: Intervention Lens: from Representation Surgery to String Counterfactuals
structure: Linear Subspace · models: GPT-2-Large, GPT2-base-french · paper: A Geometric Notion of Causal Probing
structure: Linear Direction · models: Llama 3.3 70B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-3B-Instruct, Qwen2.5-14B-Instruct, Qwen2.5-7B-Instruct · paper: The Truthfulness Spectrum Hypothesis
structure: Linear Subspace · models: BERT-base-uncased · paper: Improving Causal Interventions in Amnesic Probing with Mean Projection or LEACE
structure: Linear Subspace · models: BERT-base-uncased, Pythia-160M, Pythia-1.4B, Pythia-6.9B, Pythia-12B, LLaMA-7B, LLaMA-13B, LLaMA-30B · paper: LEACE: Perfect Linear Concept Erasure in Closed Form
structure: Linear Subspace · models: GTR-base, Mistral-7B, GPT-2-small · paper: Intervention Lens: from Representation Surgery to String Counterfactuals