MATH · IN · MODELS

SAE probe directions decode a pitch>loudness>timbre hierarchy and steer audio

measured in 1 paper

Paek et al. train SAEs on four pretrained audio autoencoders/codecs and fit linear probes from SAE features to pitch, loudness, and timbre [paek-etal-2025-learning-interpretable-features-audio-latent-spaces-sparse-autoencoders] Across all four latent spaces decodability follows the same order: pitch most separable (0.75-0.87), loudness intermediate, timbre hardest (0.17-0.46) [paek-etal-2025-learning-interpretable-features-audio-latent-spaces-sparse-autoencoders] Reusing each probe's weight row directly as a control vector causally isolates changes in the targeted property while leaving others largely intact [paek-etal-2025-learning-interpretable-features-audio-latent-spaces-sparse-autoencoders] Tracing DiffRhythm's rectified-flow generation reveals a coarse-to-fine emergence order: pitch converges first, then timbre, with loudness least resolved [paek-etal-2025-learning-interpretable-features-audio-latent-spaces-sparse-autoencoders]

Context

SAE features trained on audio-generation latent spaces (continuous VAE and discrete codec), linear probes decoding pitch, loudness, and timbre with a consistent cross-model accuracy hierarchy, control vectors reusing probe weight rows to causally steer generated audio via SAE feature-space addition, coarse-to-fine emergence order (pitch before timbre before loudness) across a diffusion/rectified-flow generation trajectory

Papers

Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders — Paek, Nathan, Zang, Yongyi, Yang, Qihui, Leistikow, Randal2025 · arXiv:2510.23802