MATH · IN · MODELS

An SAE on IndexTTS2 enables causal latent emotion steering

measured in 1 paper

Du et al. train a TopK SAE (4,096 latents, k=32) on layer-16 residual-stream activations of IndexTTS2's autoregressive semantic backbone over 56,000 emotion-controlled generations [du-etal-2026-sparse-autoencoders-for-interpretable-emotion-control-in-tts] Bidirectional latent steering achieves comparable or superior emotion induction and suppression versus global-steering and TTS baselines [du-etal-2026-sparse-autoencoders-for-interpretable-emotion-control-in-tts] Individual latents map to specific acoustics: steering latent #24 raises mean F0 by +23.11 Hz (p=1.07e-4) without affecting duration (p=0.687) [du-etal-2026-sparse-autoencoders-for-interpretable-emotion-control-in-tts] Emotional expression thus arises from coordinated, distributed latent contributions rather than one global direction [du-etal-2026-sparse-autoencoders-for-interpretable-emotion-control-in-tts]

Context

text-to-speech, emotion-control, sparse-autoencoders

Confirmed in models

Papers

Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech — Du, Hongfei, Shi, Jiacheng, Lu, Sidi, Zhou, Gang, Gao, Ye2026 · arXiv:2606.01479