A contrastive spatial-ID direction is causally bound into object tokens across 11 real VLMs
measured in 1 paperKang, Chen, Gkioxari & Perona (2026) define a canonical spatial-ID direction as a difference-of-means over grid position across 11 real VLMs (LLaVA, LLaMA, Qwen, InternVL, Gemma families), well-approximated by a fixed linear map of positional encoding (rank-3 fit, R^2 ~0.85) [kang-etal-2026-linear-mechanisms-for-spatiotemporal-reasoning-in-vlms] Activation-patching 'mirror swap' localizes the effect to object-word tokens at intermediate layers, with a color-swap control showing near-null effect [kang-etal-2026-linear-mechanisms-for-spatiotemporal-reasoning-in-vlms] Directly substituting a target spatial ID into an object token's residual stream shifts ground-truth logit-based belief 64.4% of the way to the target, versus 29.5% for a noise baseline; an analogous temporal-ID binding mechanism extends to video VLMs (LLaVA-Video, VideoLLaMA3, Qwen2.5) [kang-etal-2026-linear-mechanisms-for-spatiotemporal-reasoning-in-vlms]