TopK SAE features on real Whisper and HuBERT encoders are seed-stable and sparsely erasable, and steering them causally cuts Whisper false-speech detections by 70%
measured in 1 paperAparin, Sadekova, Rukhovich, Yermekova, Kushnareva, Popov, Kuznetsov & Piontkovskaya (2026, AudioSAE) train TopK sparse autoencoders across all encoder layers of real Whisper-small and HuBERT-base, finding over 50% of features remain consistent across random seeds and quantifying disentanglement via a concept-erasure test: only 19-27% of features need removal to erase a target concept [aparin-etal-2026-audiosae-towards-understanding-of-audio-processing-models-with-sparse-autoencoders] Individually, features capture both general acoustic/semantic content and specific disentangled paralinguistic events (environmental noise, laughter, whispering) [aparin-etal-2026-audiosae-towards-understanding-of-audio-processing-models-with-sparse-autoencoders] Causally, steering SAE features reduces real Whisper-small's false speech detections by 70% with negligible word-error-rate degradation on LibriSpeech test-clean, and a separate concept-erasure intervention removes a target concept using only 19-27% of the feature set [aparin-etal-2026-audiosae-towards-understanding-of-audio-processing-models-with-sparse-autoencoders]