MATH · IN · MODELS

A real Wav2Vec2-style model trained on longitudinal dolphin recordings organizes its codebook by whistle category and individual identity

measured in 1 paper

Semenzin, Mustun, Dessì, Orhan, Emanuelli, Lakretz, de Polavieja & Sumbre (2026) train Dolph2Vec, a Wav2Vec2.0-architecture model adapted for 44.1kHz audio, on ~180,000 whistles (100 hours) from 5 known bottlenose dolphins recorded over 5 real years [semenzin-etal-2026-dolph2vec-self-supervised-dolphin-vocalizations] The learned quantization codebook's specialization for signature-whistle categories is quantified via conditional entropy and mutual information (training reduces entropy 2.13->1.85 and raises MI 0.43->0.70 vs. a random-init baseline); UMAP-projected, GMM-clustered embeddings score ARI=0.3565 and NMI=0.4226 against ground-truth whistle labels, beating AVES-bio/BioLingual baselines [semenzin-etal-2026-dolph2vec-self-supervised-dolphin-vocalizations] A temporal-shuffling ablation on the feature-encoder output drops downstream classification accuracy from 82.0% to 75.1%, demonstrating dependence on genuine temporal structure [semenzin-etal-2026-dolph2vec-self-supervised-dolphin-vocalizations]

Context

codebook specialization, mutual information / conditional entropy

Papers

Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations — Semenzin, Chiara, Mustun, Faadil, Dessì, Roberto, Orhan, Pierre, Emanuelli, Alexis, Lakretz, Yair, de Polavieja, Gonzalo, Sumbre, Germán2026 · arXiv:2606.12503