MATH · IN · MODELS

Function vectors are compact directions that causally trigger ICL tasks

measured in 1 paper

Todd et al. use causal-mediation across 40+ ICL tasks in GPT-J-6B, GPT-NeoX-20B, and Llama-2 (7B/13B/70B) to find a small set of early-middle attention heads with high average indirect effect [todd-etal-2024] Summing those heads' mean per-task activations gives a function vector that, added at a middle layer, triggers the task even zero-shot (Llama-2 70B: 8.2% to 83.8%) [todd-etal-2024] Function vectors are portable across prompt formats and compose additively over functions, though some composed tasks are not expressible as embedding offsets [todd-etal-2024] A sharp late-layer drop in causal effect indicates function vectors trigger nonlinear downstream computation rather than a linear read-out [todd-etal-2024] Hendel et al. independently confirm a single task vector read from one residual-stream activation, recovering 80-90% of ICL across LLaMA, GPT-J, and Pythia [todd-etal-2024] Zheng et al. qualify that genuinely multi-demonstration tasks have no single task vector, and Yang et al. replicate the effect while explaining it via label-unembedding alignment [todd-etal-2024] Davidson et al. find instruction-derived and demonstration-derived function vectors only partially converge, sharing few top heads [todd-etal-2024]

Context

function vectors, in-context learning, task representation, vector composition, portability

Papers

Function Vectors in Large Language Models — Todd, Eric, Li, Millicent L., Sharma, Arnab Sen, Mueller, Aaron, Wallace, Byron C., Bau, David2024 · arXiv:2310.15213
In-Context Learning Creates Task Vectors — Hendel, Roee, Geva, Mor, Globerson, Amir2023 · arXiv:2310.15916
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning — Yang, Haolin, Cho, Hakaze, Zhong, Yiqiao, Inoue, Naoya2025 · arXiv:2505.18752
Do Different Prompting Methods Yield a Common Task Representation in Language Models? — Davidson, Guy, Gureckis, Todd M., Lake, Brenden M., Williams, Adina2025 · arXiv:2505.12075