MATH · IN · MODELS

SVD-extracted ViT-L task subspaces recover most full-feature performance

measured in 1 paper

Zhou et al. SVD-decompose converged linear-probe weights for depth, normals and segmentation on frozen DINOv2 ViT-L/14, MAE ViT-Large and iBOT ViT-Large features (ViT-Base only as auxiliary validation) [zhou-etal-2026-understanding-geometric-representations-in-self-supervised-vision-transformers-via-subspace-intervention] Projecting features onto the extracted top-k subspace recovers the vast majority of full-feature performance (DINOv2 alignment 0.916-0.948) while random and orthogonal-residual subspaces collapse to near noise [zhou-etal-2026-understanding-geometric-representations-in-self-supervised-vision-transformers-via-subspace-intervention] DINOv2 concentrates 72.5% of its geometric energy in intermediate layers with subspace similarity >0.93 across seeds at rank <=16 [zhou-etal-2026-understanding-geometric-representations-in-self-supervised-vision-transformers-via-subspace-intervention] MAE saturates over 98% of its linear potential by rank 32 while DINOv2 needs rank >=64, a difference in intrinsic task-subspace dimensionality [zhou-etal-2026-understanding-geometric-representations-in-self-supervised-vision-transformers-via-subspace-intervention]

Context

3d-awareness, representation-geometry

Papers

Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention — Zhou, Weichen, Zou, Yawen, Gu, Chunzhi, Dong, Ran, Xie, Haoran, Zhang, Chao2026 · arXiv:2607.01987