DINOv2 and Stable Diffusion encode depth/normals via nonlinear dense probes
measured in 1 paperEl Banani et al. probe frozen features from DINOv2 (ViT-B/14, ViT-L/14, with-registers), CLIP, MAE, iBOT, Stable Diffusion, DeiT III, SAM and MiDaS for single-view depth and surface normals [el-banani-etal-2024-probing-the-3d-awareness-of-visual-foundation-models] The depth/normal probe is a nonlinear multiscale DPT-style dense decoder, not a linear probe, chosen because 3D properties need not be linearly encoded [el-banani-etal-2024-probing-the-3d-awareness-of-visual-foundation-models] DINOv2 and Stable Diffusion support markedly more accurate depth/normal decoding than CLIP or MAE, whose features capture only coarse scene-layout priors (depth-normal correlation 0.37 image-level vs 0.13 pixel-level for DINOv2) [el-banani-etal-2024-probing-the-3d-awareness-of-visual-foundation-models] Despite strong single-view decodability, all models show weak multiview 3D consistency, so single-view and multiview-consistent 3D awareness are dissociable [el-banani-etal-2024-probing-the-3d-awareness-of-visual-foundation-models]