MATH · IN · MODELS

Text-only steering directions transfer causally to MLLM image tokens

measured in 1 paper

Gan et al. test whether steering directions found in a text-only LLM backbone stay causal after multimodal fine-tuning, applied to image tokens it never saw [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms] For four visual-concept taxonomies they extract directions from Gemma2-2B/9B and Llama-3.1-8B via mean-shift, probing, and SAE decoders, using text only [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms] Adding these directions to image tokens in PaliGemma2-3B/10B and Idefics3-8B-Llama3 raises CV-Bench spatial accuracy by up to +7.3%, with mean-shift strongest [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms] Frozen-hyperparameter generalization to five held-out datasets gives +7.6% average (up to +34.2% on CLEVR for Idefics3) [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms] Each visual-concept taxonomy activates fewer than 10 of 16k-32k SAE features in the text backbone, and prompting is far weaker than steering for multimodal reasoning [gan-etal-2025-textual-steering-vectors-improve-visual-understanding-mllms]

Context

text-only backbone steering directions transfer causally to a fine-tuned MLLM's image-token activations, mean-shift (diff-in-means) outperforms SAE and linear-probing directions for cross-modal transfer, sparse SAE feature counts (<10 of 16k-32k) per visual-concept taxonomy in the text-only backbone, out-of-distribution generalization of text-derived steering vectors to five further visual datasets, prompting is far less effective than activation steering for multimodal (vs. text-only) visual reasoning

Papers

Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models — Gan, Woody Haosheng, Fu, Deqing, Asilis, Julian, Liu, Ollie, Yogatama, Dani, Sharan, Vatsal, Jia, Robin, Neiswanger, Willie2025 · arXiv:2505.14071