Text-only steering directions transfer causally to MLLM image tokens
measured in 1 paperGan et al. test whether steering directions found in a text-only LLM backbone stay causal after multimodal fine-tuning, applied to image tokens it never saw [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms] For four visual-concept taxonomies they extract directions from Gemma2-2B/9B and Llama-3.1-8B via mean-shift, probing, and SAE decoders, using text only [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms] Adding these directions to image tokens in PaliGemma2-3B/10B and Idefics3-8B-Llama3 raises CV-Bench spatial accuracy by up to +7.3%, with mean-shift strongest [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms] Frozen-hyperparameter generalization to five held-out datasets gives +7.6% average (up to +34.2% on CLEVR for Idefics3) [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms] Each visual-concept taxonomy activates fewer than 10 of 16k-32k SAE features in the text backbone, and prompting is far weaker than steering for multimodal reasoning [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms]