MATH · IN · MODELS

Per-axis spatial delta vectors form separated PCA clusters in stronger VLMs

measured in 1 paper

Min et al. extract a delta vector between final-token hidden states for paired VQA prompts differing only in queried-object order, at a fixed intermediate layer [min-etal-2026-why-far-looks-up-probing-spatial-representation-in-vision-language-models] Axis Coherence, the mean pairwise cosine among sign-corrected delta vectors within an axis, rises with model strength (distance-axis 0.075 to 0.112 across Molmo scales, 0.182 for RoboRefer-2B-SFT) but stays flat at 0.04-0.05 for Qwen2.5-VL [min-etal-2026-why-far-looks-up-probing-spatial-representation-in-vision-language-models] PCA shows weak models' distance vectors collapse near the origin while RoboRefer and Qwen3-VL-235B form three cleanly separated per-axis clusters aligned to distinct principal components [min-etal-2026-why-far-looks-up-probing-spatial-representation-in-vision-language-models] Distance coherence correlates with counter-heuristic behavioral accuracy (rho=0.759, 0.804; p<1e-3) both in- and cross-domain [min-etal-2026-why-far-looks-up-probing-spatial-representation-in-vision-language-models]

Context

a cosine-coherence metric over sign-corrected contrastive delta vectors as a quantified measure of how cleanly a behavioral axis is represented as a single direction, PCA-cluster separation of per-axis delta vectors as a geometric correlate of behavioral competence, varying systematically with model strength/training scale

Papers

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models — Min, Cheolhong, Jung, Jaeyun, Lee, Daeun, Jeon, Hyeonseong, Su, Yu, Tremblay, Jonathan, Song, Chan Hee, Park, Jaesik2026 · arXiv:2605.30161