Speech models split speaker and phonetic info into orthogonal subspaces
measured in 1 paperLiu, Tang & Goldwater aggregate CPC-big/small and APC frame representations into per-speaker and per-phone means and run PCA on each [liu-tang-goldwater-2023-speaker-phonetic-orthogonal-subspaces] The top speaker and phone principal directions are nearly orthogonal (CPC-big: top-20 speaker directions average similarity to their most-aligned phone direction only 0.13, max 0.26) [liu-tang-goldwater-2023-speaker-phonetic-orthogonal-subspaces] Collapsing the speaker subspace drives speaker-probe error near chance while improving phone discrimination (ABX error), beating an utterance-standardization baseline [liu-tang-goldwater-2023-speaker-phonetic-orthogonal-subspaces] The effect generalizes to speakers never seen when the subspace was estimated [liu-tang-goldwater-2023-speaker-phonetic-orthogonal-subspaces]