MATH · IN · MODELS

SVD-based cross-scene averaging isolates a linear 3D-position subspace in real VLMs

measured in 1 paper

Wang & Gao (2026) posit an additive decomposition per object-token, h = u_id + u_sp + noise, in real Qwen2.5-VL-7B and InternVL3-8B; averaging an object's activations across many randomly-positioned synthetic 3D scenes cancels the position term, isolating an identity basis via SVD, whose orthogonal complement recovers a spatial subspace [wang-gao-2026-3d-scene-topology-in-vlms] PCA of the spatial-extracted residual recovers a 3D geometry matching true scene layout, formally converging to Laplacian eigenmaps of the scene graph [wang-gao-2026-3d-scene-topology-in-vlms] Injecting a probe-derived direction into layer-12 residuals of Qwen2.5-VL-7B causally shifts the x-coordinate probe readout monotonically with injection strength (alpha=+0.30: delta-x-hat=+0.091+/-0.021 vs. control +0.001+/-0.030), while a null-direction control shows no effect [wang-gao-2026-3d-scene-topology-in-vlms]

Context

cross-scene averaging, SVD identity/spatial decomposition

Papers

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models — Wang, Haoming, Gao, Wei2026 · arXiv:2605.07148