Patch-cosine Gram structure groups objects, strongest in self-supervised ViTs
measured in 1 paperAdeli et al. compute patch-token affinity maps and the full Gram matrix of pairwise patch cosine similarities from ViTs (DINOv3 ViT-B, DINOv2, DINO, MAE, supervised ViT-B, ConvNeXt-B) [adeli-etal-2026-human-like-object-grouping-in-self-supervised-vision-transformers] The Gram matrix shows block-like clustering of patches from the same object, quantified by AUC against ground-truth boundaries and a human-aligned grouping-accuracy benchmark [adeli-etal-2026-human-like-object-grouping-in-self-supervised-vision-transformers] Self-supervised training yields substantially more human-like grouping than supervised at matched architecture (DINOv3 ViT-B 91.9%, DINOv2 89.0% vs supervised ViT-B 70.6-72.2%) [adeli-etal-2026-human-like-object-grouping-in-self-supervised-vision-transformers] DINOv3-distilled ConvNeXt-B reaches 86.7% versus 60.0-67.4% for plain supervised ConvNeXt [adeli-etal-2026-human-like-object-grouping-in-self-supervised-vision-transformers]