CosyVoice2's LM module has a clean composable emotion subspace
measured in 1 paperWang, Bailey & Dang compare CosyVoice2's SLM (24-layer Qwen2.5-based) and CFM (56-layer DiT) modules as emotion-steering sites [wang-etal-2026-a-geometric-perspective-on-composable-emotion-steering-in-tts] Linear probing gives SLM emotion accuracy 0.80 within / 0.71 cross-speaker vs CFM 0.89 / 0.62, indicating more speaker entanglement in CFM [wang-etal-2026-a-geometric-perspective-on-composable-emotion-steering-in-tts] Local intrinsic dimensionality shows SLM's emotion representation is low-dimensional (~28D) and speaker-invariant (delta-LID +0.84), while CFM entangles speaker and emotion on a shared ~13D manifold (delta-LID -1.48) [wang-etal-2026-a-geometric-perspective-on-composable-emotion-steering-in-tts] Composed mean-difference steering vectors applied at SLM layers 14/17 give better proportional mixed-emotion control than CFM on CREMA-D/IEMOCAP [wang-etal-2026-a-geometric-perspective-on-composable-emotion-steering-in-tts]