SOMP-selected attention heads causally control target behavior across real unimodal and multimodal transformers
measured in 1 paperBasile, Maiorca, Doimo, Locatello & Cazzaniga score real Mistral-7B attention heads by Simultaneous Orthogonal Matching Pursuit (SOMP) against unembedding-matrix directions, finding that inverting just 8 heads (0.8% of the total) degrades TriviaQA country-name F1 far more selectively than inverting the same number of random heads or Logit-Lens-selected heads [basile-etal-2025-head-pursuit-probing-attention-specialization-multimodal-transformers] On RealToxicityPrompts and TET toxicity-mitigation benchmarks, suppressing 8/16/32 SOMP-selected heads reduces normalized toxic-generation counts to 0.83/0.67/0.66 (RTP) and 0.83/0.68/0.49 (TET), consistently below Logit-Lens and random-head baselines [basile-etal-2025-head-pursuit-probing-attention-specialization-multimodal-transformers] Applying the same SOMP-based head scoring to real LLaVA-NeXT-7B/13B, Gemma3-12B, and Qwen2.5-VL-7B, inverting the top-32 SOMP heads significantly disrupts image classification accuracy on MNIST, SVHN, GTSRB, EuroSAT, and RESISC45 while 32 random heads have minimal effect, and Jaccard overlap shows related-domain datasets share specialized heads [basile-etal-2025-head-pursuit-probing-attention-specialization-multimodal-transformers] On Flickr30k captioning with LLaVA, inhibiting 16 SOMP-selected heads (alpha=-1) nearly removes attribute keywords (colors, sentiments, quantities) while CIDEr stays above 80% of baseline, and enhancing 32 heads (alpha=5) increases target-concept presence by over 60% in all three attribute categories [basile-etal-2025-head-pursuit-probing-attention-specialization-multimodal-transformers]