MATH · IN · MODELS

Unsupervised Gaussian-mixture components in a real speech GAN's latent prior are linearly-manipulable speaker-attribute directions

measured in 1 paper

Lin, He, Mak, Lian & Lee (2024) train VoxGenesis, a real GAN-based speech synthesizer, on real LibriTTS/VoxCeleb audio with a Gaussian-mixture latent prior fit purely by the generative objective, with no speaker labels used during training [lin-etal-2024-voxgenesis-unsupervised-discovery-of-latent-speaker-manifold] Individual Gaussian components in the discovered latent prior align with human-interpretable speaker attributes (gender, age, accent), and moving a sample's latent code along the vector connecting two component means causally and smoothly interpolates the corresponding attribute in the synthesized voice, confirmed via human listener ratings [lin-etal-2024-voxgenesis-unsupervised-discovery-of-latent-speaker-manifold]

Method

Papers

VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech Synthesis — Lin, Weiwei, He, Chenhang, Mak, Man-Wai, Lian, Jiachen, Lee, Kong Aik2024 · arXiv:2403.00529