MATH · IN · MODELS

Function-vector heads drive ICL; induction heads matter little and FV heads evolve from them

measured in 1 paper

Yin & Steinhardt ablate function-vector and induction heads across 12 models (70M-7B), finding FV-head ablation substantially degrades few-shot ICL while induction-head ablation barely exceeds random, a gap growing with scale [yin-steinhardt-2025-which-attention-heads-matter-for-icl] An ablation-with-exclusion design shows the apparent induction-head effect was mostly driven by heads that are both induction and FV heads [yin-steinhardt-2025-which-attention-heads-matter-for-icl] Direct overlap between top induction and FV heads is minimal, yet the two scores are correlated [yin-steinhardt-2025-which-attention-heads-matter-for-icl] Across Pythia training checkpoints, induction heads emerge early (~step 1,000) and FV heads substantially later (~step 16,000), with many FV heads evolving unidirectionally from earlier induction heads [yin-steinhardt-2025-which-attention-heads-matter-for-icl]

Context

induction-head vs. function-vector-head ablation, ablation-with-exclusion design, FV heads emerge from induction heads during training

Papers

Which Attention Heads Matter for In-Context Learning? — Yin, Kayo, Steinhardt, Jacob2025 · arXiv:2502.14010