MATH · IN · MODELS

Monosemantic Whisper SAE latents encode phonetic/lexical features and steer transcripts

measured in 1 paper

Pluth et al. train a TopK sparse autoencoder on ~200M frames of Whisper-base encoder activations, recovering monosemantic dictionary directions [pluth-etal-2026-mechanistic-interpretability-asr-sae] These span diphone (e.g. /r-uw/ precision 88.7%), word (e.g. "his" precision 99.3%), language-discrimination (recall 91.2%), and profanity (recall 89.7%) features [pluth-etal-2026-mechanistic-interpretability-asr-sae] Adding or clamping these SAE decoder directions causally steers Whisper's transcript content, including cross-lingually [pluth-etal-2026-mechanistic-interpretability-asr-sae]

Context

SAE monosemanticity in ASR encoders, cross-lingual causal steering, profanity direction

Confirmed in models

Papers

Mechanistic Interpretability of ASR models using Sparse Autoencoders — Pluth, Dan, Houghton, Zachary Nicholas, Zhou, Yu, Gurbani, Vijay K.2026 · arXiv:2605.12225