A triangle-area tri-modal similarity loss improves retrieval
measured in 1 paperCicchetti et al. view three unit-norm modality embeddings as a 2-simplex whose closed-form area is an exact anchor-free, fusion-free measure of three-way alignment [cicchetti-etal-2025-triangle-multimodal-alignment] Replacing pairwise cosine similarity in a contrastive loss with the negative triangle area (the TRIANGLE loss) improves zero-shot video-text Recall@1 by +3.6 to +8.8 over VAST at matched capacity, and audio retrieval by up to +5.2 Recall@1 [cicchetti-etal-2025-triangle-multimodal-alignment] In a controlled toy three-modality setting TRIANGLE converges up to 4x faster than a cosine anchor loss and than the volume-based GRAM loss, but on AudioCaps GRAM outperforms TRIANGLE, so the "beats volume-based" claim is scoped to the toy setting [cicchetti-etal-2025-triangle-multimodal-alignment] TRIANGLE is a training-loss objective rather than an emergent discovered structure, with negligible compute overhead [cicchetti-etal-2025-triangle-multimodal-alignment]