SVD-extracted ViT-L task subspaces recover most full-feature performance
measured in 1 paperZhou et al. SVD-decompose converged linear-probe weights for depth, normals and segmentation on frozen DINOv2 ViT-L/14, MAE ViT-Large and iBOT ViT-Large features (ViT-Base only as auxiliary validation) [zhou-etal-2026-understanding-geometric-representations-in-self-supervised-vision-transformers-via-subspace-intervention] Projecting features onto the extracted top-k subspace recovers the vast majority of full-feature performance (DINOv2 alignment 0.916-0.948) while random and orthogonal-residual subspaces collapse to near noise [zhou-etal-2026-understanding-geometric-representations-in-self-supervised-vision-transformers-via-subspace-intervention] DINOv2 concentrates 72.5% of its geometric energy in intermediate layers with subspace similarity >0.93 across seeds at rank <=16 [zhou-etal-2026-understanding-geometric-representations-in-self-supervised-vision-transformers-via-subspace-intervention] MAE saturates over 98% of its linear potential by rank 32 while DINOv2 needs rank >=64, a difference in intrinsic task-subspace dimensionality [zhou-etal-2026-understanding-geometric-representations-in-self-supervised-vision-transformers-via-subspace-intervention]