CLAP (Contrastive Language-Audio Pretraining)
Multiple (Microsoft/LAION-style CLAP architecture)
Structures found in this family (3)
By model (3)
CLAP (HTSAT-BERT-ZS, trained on WavCaps)
DRCap's CLAP (trained on WavCaps + SoundVECaps)
LAION-CLAP (HTSAT backbone, trained on LAION-Audio-630k + AudioSet + speech/music)
Observations (2)
Papers
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings (2026), Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings (2026)