MATH · IN · MODELS

Subtracting a regression-fit speaker component removes speaker identity

measured in 1 paper

Ruggiero et al. fit a closed-form affine regression from a frozen ECAPA-TDNN speaker embedding (PCA-reduced to 128 dims) to WavLM-Large layer-15 frames, then subtract the predicted component at inference [ruggiero-etal-2025-eta-wavlm-efficient-speaker-identity-removal-in-self-supervised-speech-representations] Speaker-classification accuracy collapses 82.30% to 55.73% (paired t-test p=5e-5) [ruggiero-etal-2025-eta-wavlm-efficient-speaker-identity-removal-in-self-supervised-speech-representations] Voice-conversion quality improves over multiple baselines (LJSpeech WER 3.81 vs 4.56, MOS 4.00 vs 3.84) [ruggiero-etal-2025-eta-wavlm-efficient-speaker-identity-removal-in-self-supervised-speech-representations] Unlike orthogonal-subspace findings, this engineers and subtracts a regression-fit affine component, a continuous-variable analogue of LEACE-style removal [ruggiero-etal-2025-eta-wavlm-efficient-speaker-identity-removal-in-self-supervised-speech-representations]

Context

closed-form affine regression against a continuous (not discrete-class) conditioning variable, then subtraction, as a causal-ablation technique distinct from LEACE/mean-projection's discrete-class subspace removal

Confirmed in models

Papers

Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation — Ruggiero, Giuseppe, Testa, Matteo, Van de Walle, Jurgen, Di Caro, Luigi2025 · arXiv:2505.19273