Multi-object scene embeddings decompose hierarchically as sums of object embeddings
measured in 1 paperUselis, Koishigarina & Oh extend additive factorization to multi-object scenes, finding scene embedding approximately equals the sum of object embeddings, themselves approximately sums of concept embeddings, in CLIP and DINOv2 [uselis-koishigarina-oh-2026-how-can-embedding-models-bind-concepts] Quantified via R^2, retrieval, and probing: text-CLIP R^2=0.90-0.92, PUG:SPARE 0.75-0.86, CLEVR-2D 0.75-0.79, all clearly above random baselines of 0.47-0.53 [uselis-koishigarina-oh-2026-how-can-embedding-models-bind-concepts] Causal embedding-editing (removing or inserting object components) produces meaningful counterfactual embeddings [uselis-koishigarina-oh-2026-how-can-embedding-models-bind-concepts] This extends single-object per-concept factorization to a hierarchical, multi-object additive structure [uselis-koishigarina-oh-2026-how-can-embedding-models-bind-concepts]