SAE decoder directions rotate during real vision-language fine-tuning, and the most-rotated subset is causally load-bearing for spatial reasoning
measured in 1 paperNaghashyar, Batra, Khakzar, Torr, Clark, Schroeder de Witt & Venhoff warm-start LLaMA-Scope SAEs on LLaVA-More (CLIP ViT-L/14-336 + Llama-3.1-8B) activations, measuring per-feature decoder-direction cosine similarity between the base-LLM SAE and the VLM-adapted SAE [naghashyar-etal-2026-multimodal-fine-tuning-spatial-features] Roughly 5% of over 1M features show strong decoder-direction rotation (bottom-25% cosine) combined with visual responsiveness, and a further firing-frequency-shift criterion isolates a spatial subset validated via attribution patching to attention heads [naghashyar-etal-2026-multimodal-fine-tuning-spatial-features] Causally ablating the top spatial SAE features drops Visual Spatial Reasoning accuracy by 5.85-15.54 points while general VQA accuracy changes by under 1 point, versus near-zero effect for a random-feature control (odds ratios 4.2-9.1 for spatial recruitment) [naghashyar-etal-2026-multimodal-fine-tuning-spatial-features]