MATH · IN · MODELS

A single linear ViT direction encodes depth and causally drives depth estimation

measured in 1 paper

Sanghavi fits linear probes on frozen ViT-B/16 (google/vit-base-patch16-224-in21k, plus a random-weight control) layer activations, finding depth best decoded at layer 8 (MAE=0.0875) [sanghavi-2026-from-edges-to-depth-vit-spatial-hierarchy] Ablating the probe-identified direction increases depth-estimation error by 49-165%, while ablating a random direction of the same rank changes error by under 1% [sanghavi-2026-from-edges-to-depth-vit-spatial-hierarchy] Depth is thus concentrated in a specific linear direction, not diffusely distributed [sanghavi-2026-from-edges-to-depth-vit-spatial-hierarchy] Targeted single-direction activation patching retains 76% of its causal effect at a 9-layer gap before decaying at longer range [sanghavi-2026-from-edges-to-depth-vit-spatial-hierarchy]

Context

monocular depth, linear probing, randomized-control ablation, single-direction activation patching, vision transformer

Papers

From Edges to Depth: Probing the Spatial Hierarchy in Vision Transformers — Sanghavi, Jainum2026 · arXiv:2604.23452