Audio-model SAEs raise DCI completeness while preserving informativeness
measured in 1 paperMariotte et al. apply TopK SAEs to pooled representations of AST, HuBERT-base, WavLM-base-plus and MERT-v1-95M on VocalSet singing-technique classification [mariotte-etal-2025-sparse-autoencoders-make-audio-foundation-models-more-explainable] SAE codes retain classification informativeness (AST 81.8% up to 95% sparsity) while raising DCI completeness for eGeMAPS acoustic factors relative to the dense baseline [mariotte-etal-2025-sparse-autoencoders-make-audio-foundation-models-more-explainable] They localize specific factors to specific layers: pitch to early layers for HuBERT/WavLM, formants to final layers [mariotte-etal-2025-sparse-autoencoders-make-audio-foundation-models-more-explainable]
Structure
Context
audio-foundation-models, sparse-autoencoders
Confirmed in models
Method
Papers
Sparse Autoencoders Make Audio Foundation Models More Explainable — Mariotte, Théo, Lebourdais, Martin, Almudévar, Antonio, Tahon, Marie, Ortega, Alfonso, Dugué, Nicolas