MATH · IN · MODELS

CosyVoice2's LM module has a clean composable emotion subspace

measured in 1 paper

Wang, Bailey & Dang compare CosyVoice2's SLM (24-layer Qwen2.5-based) and CFM (56-layer DiT) modules as emotion-steering sites [wang-etal-2026-a-geometric-perspective-on-composable-emotion-steering-in-tts] Linear probing gives SLM emotion accuracy 0.80 within / 0.71 cross-speaker vs CFM 0.89 / 0.62, indicating more speaker entanglement in CFM [wang-etal-2026-a-geometric-perspective-on-composable-emotion-steering-in-tts] Local intrinsic dimensionality shows SLM's emotion representation is low-dimensional (~28D) and speaker-invariant (delta-LID +0.84), while CFM entangles speaker and emotion on a shared ~13D manifold (delta-LID -1.48) [wang-etal-2026-a-geometric-perspective-on-composable-emotion-steering-in-tts] Composed mean-difference steering vectors applied at SLM layers 14/17 give better proportional mixed-emotion control than CFM on CREMA-D/IEMOCAP [wang-etal-2026-a-geometric-perspective-on-composable-emotion-steering-in-tts]

Context

text-to-speech, emotion-steering

Confirmed in models

Papers

A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models — Wang, Siyi, Bailey, James, Dang, Ting2026 · arXiv:2607.00946