SAE feature modality composition shifts by depth in a TTS LM and steers speech
measured in 1 paperKoriagin et al. train BatchTopK SAEs on residual-stream activations at multiple layers of CosyVoice3's Qwen2.5-0.5B TTS backbone [koriagin-etal-2026-interpreting-steering-tts-with-saes] Feature modality composition is systematic by depth: early/middle layers mixed, layers 16-20 audio-heavy, and the final hidden state reverts to a mostly text-modal subspace (layer-20 text-detection AUROC 0.921) [koriagin-etal-2026-interpreting-steering-tts-with-saes] Encoding through the frozen SAE, shifting selected features, and decoding back raises laughter probability from 0.02 to 0.79, flips perceived speaker gender, and controls speech rate while preserving content [koriagin-etal-2026-interpreting-steering-tts-with-saes] This is the first such layerwise modality-composition analysis of a generative TTS LM backbone [koriagin-etal-2026-interpreting-steering-tts-with-saes]