MATH · IN · MODELS
methods / Causal Validation / ProFS (Projection Filter for Subspaces) weight-subspace projection

ProFS (Projection Filter for Subspaces) weight-subspace projection

Techniqueadvanced

Extracts a low-rank concept subspace from contrastive-pair embedding differences (after removing the corpus-mean direction), selects its rank via ScreeNot, and projects it out of a weight matrix once, offline — a factor-analysis-grounded, sample-efficient, noise-robust alternative to gradient-based preference-tuning (DPO) that edits weights directly rather than activations at inference time.

Used in (2 observations)

structure: Linear Subspace · models: Stable Diffusion v1.5 · paper: Concept Unlearning via Cross-Attention Activation Projection for Diffusion Models
structure: Linear Subspace · models: GPT-2-Medium, Mistral-7B, Mistral-7B-SFT-Beta, OPT-6.7B, GPT-J-6B · paper: Model Editing as a Robust and Denoised Variant of DPO: A Case Study on Toxicity