MATH · IN · MODELS

Speech models split speaker and phonetic info into orthogonal subspaces

measured in 1 paper

Liu, Tang & Goldwater aggregate CPC-big/small and APC frame representations into per-speaker and per-phone means and run PCA on each [liu-tang-goldwater-2023-speaker-phonetic-orthogonal-subspaces] The top speaker and phone principal directions are nearly orthogonal (CPC-big: top-20 speaker directions average similarity to their most-aligned phone direction only 0.13, max 0.26) [liu-tang-goldwater-2023-speaker-phonetic-orthogonal-subspaces] Collapsing the speaker subspace drives speaker-probe error near chance while improving phone discrimination (ABX error), beating an utterance-standardization baseline [liu-tang-goldwater-2023-speaker-phonetic-orthogonal-subspaces] The effect generalizes to speakers never seen when the subspace was estimated [liu-tang-goldwater-2023-speaker-phonetic-orthogonal-subspaces]

Context

speaker subspace, phone subspace, aggregated-PCA on class means, orthogonal-subspace collapse, unseen-speaker generalization

Papers

Self-supervised Predictive Coding Models Encode Speaker and Phonetic Information in Orthogonal Subspaces — Liu, Oli Danyi, Tang, Hao, Goldwater, Sharon2023 · arXiv:2305.12464