Subtracting a regression-fit speaker component removes speaker identity
measured in 1 paperRuggiero et al. fit a closed-form affine regression from a frozen ECAPA-TDNN speaker embedding (PCA-reduced to 128 dims) to WavLM-Large layer-15 frames, then subtract the predicted component at inference [ruggiero-etal-2025-eta-wavlm-efficient-speaker-identity-removal-in-self-supervised-speech-representations] Speaker-classification accuracy collapses 82.30% to 55.73% (paired t-test p=5e-5) [ruggiero-etal-2025-eta-wavlm-efficient-speaker-identity-removal-in-self-supervised-speech-representations] Voice-conversion quality improves over multiple baselines (LJSpeech WER 3.81 vs 4.56, MOS 4.00 vs 3.84) [ruggiero-etal-2025-eta-wavlm-efficient-speaker-identity-removal-in-self-supervised-speech-representations] Unlike orthogonal-subspace findings, this engineers and subtracts a regression-fit affine component, a continuous-variable analogue of LEACE-style removal [ruggiero-etal-2025-eta-wavlm-efficient-speaker-identity-removal-in-self-supervised-speech-representations]