Monosemantic Whisper SAE latents encode phonetic/lexical features and steer transcripts
measured in 1 paperPluth et al. train a TopK sparse autoencoder on ~200M frames of Whisper-base encoder activations, recovering monosemantic dictionary directions [pluth-etal-2026-mechanistic-interpretability-asr-sae] These span diphone (e.g. /r-uw/ precision 88.7%), word (e.g. "his" precision 99.3%), language-discrimination (recall 91.2%), and profanity (recall 89.7%) features [pluth-etal-2026-mechanistic-interpretability-asr-sae] Adding or clamping these SAE decoder directions causally steers Whisper's transcript content, including cross-lingually [pluth-etal-2026-mechanistic-interpretability-asr-sae]
Structure
Context
SAE monosemanticity in ASR encoders, cross-lingual causal steering, profanity direction
Confirmed in models
Papers
Mechanistic Interpretability of ASR models using Sparse Autoencoders — Pluth, Dan, Houghton, Zachary Nicholas, Zhou, Yu, Gurbani, Vijay K.