MATH · IN · MODELS

DINOv2 and Stable Diffusion encode depth/normals via nonlinear dense probes

measured in 1 paper

El Banani et al. probe frozen features from DINOv2 (ViT-B/14, ViT-L/14, with-registers), CLIP, MAE, iBOT, Stable Diffusion, DeiT III, SAM and MiDaS for single-view depth and surface normals [el-banani-etal-2024-probing-the-3d-awareness-of-visual-foundation-models] The depth/normal probe is a nonlinear multiscale DPT-style dense decoder, not a linear probe, chosen because 3D properties need not be linearly encoded [el-banani-etal-2024-probing-the-3d-awareness-of-visual-foundation-models] DINOv2 and Stable Diffusion support markedly more accurate depth/normal decoding than CLIP or MAE, whose features capture only coarse scene-layout priors (depth-normal correlation 0.37 image-level vs 0.13 pixel-level for DINOv2) [el-banani-etal-2024-probing-the-3d-awareness-of-visual-foundation-models] Despite strong single-view decodability, all models show weak multiview 3D consistency, so single-view and multiview-consistent 3D awareness are dissociable [el-banani-etal-2024-probing-the-3d-awareness-of-visual-foundation-models]

Context

3d-awareness, depth-estimation

Papers

Probing the 3D Awareness of Visual Foundation Models — El Banani, Mohamed, Raj, Amit, Maninis, Kevis-Kokitsi, Kar, Abhishek, Li, Yuanzhen, Rubinstein, Michael, Sun, Deqing, Guibas, Leonidas, Johnson, Justin, Jampani, Varun2024 · arXiv:2404.08636