MATH · IN · MODELS

A triangle-area tri-modal similarity loss improves retrieval

measured in 1 paper

Cicchetti et al. view three unit-norm modality embeddings as a 2-simplex whose closed-form area is an exact anchor-free, fusion-free measure of three-way alignment [cicchetti-etal-2025-triangle-multimodal-alignment] Replacing pairwise cosine similarity in a contrastive loss with the negative triangle area (the TRIANGLE loss) improves zero-shot video-text Recall@1 by +3.6 to +8.8 over VAST at matched capacity, and audio retrieval by up to +5.2 Recall@1 [cicchetti-etal-2025-triangle-multimodal-alignment] In a controlled toy three-modality setting TRIANGLE converges up to 4x faster than a cosine anchor loss and than the volume-based GRAM loss, but on AudioCaps GRAM outperforms TRIANGLE, so the "beats volume-based" claim is scoped to the toy setting [cicchetti-etal-2025-triangle-multimodal-alignment] TRIANGLE is a training-loss objective rather than an emergent discovered structure, with negligible compute overhead [cicchetti-etal-2025-triangle-multimodal-alignment]

Context

three modality embeddings as vertices of a 2-simplex (triangle) in shared embedding space, closed-form triangle-area formula (Gram-determinant/Heron-style) as an anchor-free tri-modal similarity, triangle-area contrastive loss replacing pairwise cosine similarity, with cosine regularization for task-specific pairs, comparison against general n-modal alternatives (GRAM parallelotope volume, Symile total-correlation bound)

Papers

A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity — Cicchetti, Giordano, Grassucci, Eleonora, Comminiello, Danilo2025 · arXiv:2509.24734