MATH · IN · MODELS

Multi-object scene embeddings decompose hierarchically as sums of object embeddings

measured in 1 paper

Uselis, Koishigarina & Oh extend additive factorization to multi-object scenes, finding scene embedding approximately equals the sum of object embeddings, themselves approximately sums of concept embeddings, in CLIP and DINOv2 [uselis-koishigarina-oh-2026-how-can-embedding-models-bind-concepts] Quantified via R^2, retrieval, and probing: text-CLIP R^2=0.90-0.92, PUG:SPARE 0.75-0.86, CLEVR-2D 0.75-0.79, all clearly above random baselines of 0.47-0.53 [uselis-koishigarina-oh-2026-how-can-embedding-models-bind-concepts] Causal embedding-editing (removing or inserting object components) produces meaningful counterfactual embeddings [uselis-koishigarina-oh-2026-how-can-embedding-models-bind-concepts] This extends single-object per-concept factorization to a hierarchical, multi-object additive structure [uselis-koishigarina-oh-2026-how-can-embedding-models-bind-concepts]

Context

a two-level (scene-to-object, object-to-concept) additive hierarchical decomposition, extending single-level per-concept factorization to compositional multi-object scenes, causal validation of an additive decomposition via direct embedding editing rather than only a reconstruction-fidelity score

Papers

How Can Embedding Models Bind Concepts? — Uselis, Arnas, Koishigarina, Darina, Oh, Seong Joon2026 · arXiv:2605.31503