MATH · IN · MODELS

Natural task vectors emerge weakly in deep models; a TVP loss forces a layer to carry the task

measured in 1 paper

Yang et al. train small GPT-2-style transformers from scratch on synthetic ICL tasks and define a task vector behaviorally by whether its injection recovers ICL performance [yang-etal-2025-task-vectors-emergence-formation-benefit] On a shallow 3-layer model the natural task vector is causally effective (linear-regression MSE ~0.25 vs ~1.2 random) with layer-swap specificity [yang-etal-2025-task-vectors-emergence-formation-benefit] In deeper 8-layer models task information becomes distributed across layers and task-vector prompting performance is nearly random [yang-etal-2025-task-vectors-emergence-formation-benefit] An auxiliary task-vector-prompting loss trains a prescribed layer to serve as an injectable task vector, matching full ICL performance and improving out-of-distribution robustness [yang-etal-2025-task-vectors-emergence-formation-benefit]

Context

synthetic-transformer task-vector emergence (dummy-query-style extraction, layer search), causal injection-and-measure validation (MSE recovery, layer-swap ablation), weak/non-local task encoding in deeper (8-layer) vanilla-trained models, TVP-loss (auxiliary training-time loss forcing a prescribed layer to serve as an injectable task vector)

Papers

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit — Yang, Liu, Lin, Ziqian, Lee, Kangwook, Papailiopoulos, Dimitris, Nowak, Robert2025 · arXiv:2501.09240