MATH · IN · MODELS

Lightweight VLMs linearly encode object presence but never orientation

measured in 1 paper

Theodoridis et al. use CARLA counterfactual pairs and layer-wise linear probes across four lightweight VLMs (Ovis2.5-2B, InternVL3.5-2B, VST-3B SFT and RL) [theodoridis-etal-2026-probing-visual-concepts-in-lightweight-vision-language-models-for-automated-driving] Object/agent presence is linearly encoded from mid-vision-encoder through final LLM layers, while orientation is never linearly encoded anywhere, and increasing object distance rapidly degrades separability [theodoridis-etal-2026-probing-visual-concepts-in-lightweight-vision-language-models-for-automated-driving] Two failure modes are dissociated: perceptual (concept not encoded) versus cognitive (concept present but the model still answers wrong) [theodoridis-etal-2026-probing-visual-concepts-in-lightweight-vision-language-models-for-automated-driving] Activation steering along the probe directions causally confirms they drive behavior rather than being correlational [theodoridis-etal-2026-probing-visual-concepts-in-lightweight-vision-language-models-for-automated-driving]

Context

vision-language-models, spatial-reasoning

Papers

Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving — Theodoridis, Nikos, Mohandas, Reenu, Sistu, Ganesh, Scanlan, Anthony, Eising, Ciarán, Brophy, Tim2026 · arXiv:2603.06054