MATH · IN · MODELS

SAE decoder directions in MusicGen are causally steerable; listeners prefer steered audio

measured in 1 paper

Singh, Cherep & Maes train k-sparse autoencoders on residual-stream activations from five layers each of MusicGen-Large and MusicGen-Small on the ~160k-clip MusicSet corpus [singh-cherep-maes-2026-discovering-and-steering-interpretable-concepts-in-large-generative-music-models] Adding a scaled SAE decoder direction to the residual stream causally steers generation, with 15-35% of tested features improving CLAP alignment [singh-cherep-maes-2026-discovering-and-steering-interpretable-concepts-in-large-generative-music-models] A controlled human listening study (10 participants, 100 trials) finds listeners picked the SAE-steered audio 66/100 versus 17/17 for random-direction and baseline controls (chi^2=48.02, p<.0001) [singh-cherep-maes-2026-discovering-and-steering-interpretable-concepts-in-large-generative-music-models] The steering is thus a perceptibly validated, geometry-tied causal effect in a real pretrained generative audio model [singh-cherep-maes-2026-discovering-and-steering-interpretable-concepts-in-large-generative-music-models]

Context

SAE decoder-direction additive steering in a generative audio model, validated by a controlled human perceptual study rather than only an automated metric

Papers

Discovering and Steering Interpretable Concepts in Large Generative Music Models — Singh, Nikhil, Cherep, Manuel, Maes, Pattie2026 · arXiv:2505.18186