DeepSpeech2 untangles word and phoneme manifolds while tangling speaker
measured in 1 paperStephenson et al. apply mean-field manifold capacity theory to a real trained DeepSpeech2 ASR network (960 hours of LibriSpeech) for phoneme, speaker, and word manifolds [stephenson-etal-2019-untangling-in-invariant-speech-recognition] Word and phoneme manifolds progressively untangle (linear separability rises) across network depth and recurrent processing time [stephenson-etal-2019-untangling-in-invariant-speech-recognition] Speaker manifolds do the opposite: their separability decreases with depth, dropping even below the untrained network, as the network builds speaker-invariance by increasing speaker-manifold dimensionality [stephenson-etal-2019-untangling-in-invariant-speech-recognition] Word and phoneme separability also untangle temporally, peaking near the location of the relevant word [stephenson-etal-2019-untangling-in-invariant-speech-recognition]