MATH · IN · MODELS

CLAP linearly encodes reverberation and loudness with cross-dataset-consistent axes

measured in 1 paper

Martel et al. linearly probe real LAION-CLAP audio embeddings for four low-level acoustic attributes across five datasets [martel-etal-2026-clap-acoustic-attribute-probing] RT60 (reverberation time) is strongly, nearly linearly encoded (R^2>=0.67, up to 0.92 on White Noise/VCTK) and LUFS (loudness) reaches R^2>=0.76 [martel-etal-2026-clap-acoustic-attribute-probing] The independently-fit RT60 and LUFS probe axes are geometrically consistent across datasets (RT60 cosine up to 0.86), while the relative-pitch axis is domain-specific near the random-vector baseline [martel-etal-2026-clap-acoustic-attribute-probing] Across 8 more pretrained audio encoders, amplitude-invariant architectures specifically fail to encode LUFS, tying the failure to a known architectural property [martel-etal-2026-clap-acoustic-attribute-probing]

Context

RT60/LUFS linear feature axes, cross-dataset probe-direction cosine consistency, amplitude-invariance failure mode

Papers

Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings — Martel, Hector, Hennessy-Priest, Joe, Cho, Taemin2026 · arXiv:2607.03806