MATH · IN · MODELS

DeepSpeech2 untangles word and phoneme manifolds while tangling speaker

measured in 1 paper

Stephenson et al. apply mean-field manifold capacity theory to a real trained DeepSpeech2 ASR network (960 hours of LibriSpeech) for phoneme, speaker, and word manifolds [stephenson-etal-2019-untangling-in-invariant-speech-recognition] Word and phoneme manifolds progressively untangle (linear separability rises) across network depth and recurrent processing time [stephenson-etal-2019-untangling-in-invariant-speech-recognition] Speaker manifolds do the opposite: their separability decreases with depth, dropping even below the untrained network, as the network builds speaker-invariance by increasing speaker-manifold dimensionality [stephenson-etal-2019-untangling-in-invariant-speech-recognition] Word and phoneme separability also untangle temporally, peaking near the location of the relevant word [stephenson-etal-2019-untangling-in-invariant-speech-recognition]

Context

manifold capacity theory, phoneme/speaker/word manifold untangling, invariant speech representations

Papers

Untangling in Invariant Speech Recognition — Stephenson, Cory, Feather, Jenelle, Padhy, Suchismita, Elibol, Oguz H., Tang, Hanlin, McDermott, Josh H., Chung, SueYeon2019 · arXiv:2003.01787