MATH · IN · MODELS

SAE feature modality composition shifts by depth in a TTS LM and steers speech

measured in 1 paper

Koriagin et al. train BatchTopK SAEs on residual-stream activations at multiple layers of CosyVoice3's Qwen2.5-0.5B TTS backbone [koriagin-etal-2026-interpreting-steering-tts-with-saes] Feature modality composition is systematic by depth: early/middle layers mixed, layers 16-20 audio-heavy, and the final hidden state reverts to a mostly text-modal subspace (layer-20 text-detection AUROC 0.921) [koriagin-etal-2026-interpreting-steering-tts-with-saes] Encoding through the frozen SAE, shifting selected features, and decoding back raises laughter probability from 0.02 to 0.79, flips perceived speaker gender, and controls speech rate while preserving content [koriagin-etal-2026-interpreting-steering-tts-with-saes] This is the first such layerwise modality-composition analysis of a generative TTS LM backbone [koriagin-etal-2026-interpreting-steering-tts-with-saes]

Context

SAE feature modality composition, layerwise text/audio structure, TTS steering (laughter, gender, speech rate)

Papers

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders — Koriagin, Nikita, Aparin, Georgii, Balagansky, Nikita, Gavrilov, Daniil2026 · arXiv:2606.10029