MLP probes decode seven music-theory concepts from generative models
measured in 1 paperWei et al. introduce a synthetic dataset isolating seven music-theory concepts and train two-layer MLP probes on internal representations of Jukebox and MusicGen small/medium/large [wei-freeman-etal-2024-do-music-generation-models-encode-music-theory-syntheory] Probes on the language-model stage reach mean decodability up to 0.984 (Jukebox), 0.950 (MusicGen-Small), 0.914 (Medium), and 0.929 (Large) [wei-freeman-etal-2024-do-music-generation-models-encode-music-theory-syntheory] The MusicGen audio-codec stage alone reaches only 0.701, so decodability depends on which processing stage and model scale are probed [wei-freeman-etal-2024-do-music-generation-models-encode-music-theory-syntheory]
Structure
Context
concept decodability varying systematically across a generative model's own processing stages (audio codec vs. language-model decoder), not just across layers of a single homogeneous stack
Confirmed in models
Method
Papers
Do Music Generation Models Encode Music Theory? (SynTheory) — Wei, Megan, Freeman, Michael, Donahue, Chris, Sun, Chen