MATH · IN · MODELS

Mid-level vision decodability weakly correlates with recognition accuracy

measured in 1 paper

Chen, Marks & Cheng introduce eight frozen-feature benchmarks for mid-level vision (segmentation, geometric/3D grouping) across 22 self-supervised models spanning contrastive, clustering, pretext and masked paradigms [chen-etal-2024-probing-the-mid-level-vision-capabilities-of-self-supervised-learning] The benchmarks decode with a nonlinear DPT dense decoder on frozen features, not a linear probe [chen-etal-2024-probing-the-mid-level-vision-capabilities-of-self-supervised-learning] Mid-level task performance correlates only weakly with each model's high-level ImageNet accuracy (generic segmentation strongest ~R^2 0.70; 3D understanding weakly correlated) [chen-etal-2024-probing-the-mid-level-vision-capabilities-of-self-supervised-learning] Several models are strongly imbalanced across the two capability classes, so mid-level decodability is not simply a byproduct of overall recognition quality [chen-etal-2024-probing-the-mid-level-vision-capabilities-of-self-supervised-learning]

Context

mid-level-vision, representation-geometry

Papers

Probing the Mid-level Vision Capabilities of Self-Supervised Learning — Chen, Xuweiyi, Marks, Markus, Cheng, Zezhou2024 · arXiv:2411.17474