MATH · IN · MODELS

A shared PCA basis's leading components separate figure from ground

measured in 1 paper

Li et al. project activations of several pretrained self-supervised vision models (ViT-MAE, CLIP ViT-H/14, DINOv2, ConvNeXt-MAE) onto the first 1-3 components of a fixed shared PCA basis [li-etal-2025-from-local-cues-to-global-percepts-emergent-gestalt-organization-in-self-supervised-vision-models] This low-dimensional axis cleanly separates figure from ground pixels, validated against ground-truth segmentation masks on 1,495 PASCAL VOC images [li-etal-2025-from-local-cues-to-global-percepts-emergent-gestalt-organization-in-self-supervised-vision-models] A Top-K activation-sparsity intervention raises a texture-oddity benchmark from 69.8 to 94.6 for MAE ViT and 62.1 to 88.1 for ConvNeXt-V1 [li-etal-2025-from-local-cues-to-global-percepts-emergent-gestalt-organization-in-self-supervised-vision-models]

Context

a shared PCA basis's leading components separating figure from ground across multiple architecturally distinct self-supervised vision models, validated against ground-truth segmentation rather than a behavioral proxy alone

Method

Papers

From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models — Li, Tianqin, Wen, Ziqi, Song, Leiran, Liu, Jun, Jing, Zhi, Lee, Tai Sing2025 · arXiv:2506.00718