methods / Causal Validation / PCA-derived semantic-subspace ablation, validated against a randomized-subspace control
PCA-derived semantic-subspace ablation, validated against a randomized-subspace control
Extracts a low-rank PCA subspace spanning a semantic category (e.g. country names, climate vocabulary) from a set of related embeddings, projects it out of a target embedding, and quantifies the causal drop in a downstream linear-probe's decodability — validated by comparison against ablating an equal-rank random subspace, so the effect can be attributed to the specific semantic content rather than merely to removing any generic low-rank chunk of the space.