GPT-2 heads multiplex subfunctions along orthogonal singular directions
measured in 1 paperAhmad, Joshi & Modi take the SVD of each attention head's and MLP layer's augmented weight matrix in pretrained GPT-2 Small [ahmad-joshi-modi-2025-beyond-components-singular-vector-based-interpretability-of-transformer-circuits] Individual components (e.g. head 9.6) encode multiple overlapping subfunctions aligned with distinct orthogonal singular directions [ahmad-joshi-modi-2025-beyond-components-singular-vector-based-interpretability-of-transformer-circuits] A learnable diagonal mask over singular values prunes 91-99% of directions across IOI, Greater-Than, and Gender-Pronoun while retaining 0.70-0.79 accuracy and low KL (0.21) [ahmad-joshi-modi-2025-beyond-components-singular-vector-based-interpretability-of-transformer-circuits] Scalar interventions on single singular directions flip gender-pronoun predictions at perfect accuracy, a targeted causal validation [ahmad-joshi-modi-2025-beyond-components-singular-vector-based-interpretability-of-transformer-circuits]