MATH · IN · MODELS

Task-vector subspace implements Bayesian retrieval; OOD generalization is near-orthogonal

measured in 1 paper

Yan et al. train small RoPE transformers on synthetic latent-task mixtures and formalize task vectors as task-conditional-minus-global mean hidden states [yan-yang-zhong-2026-task-vector-geometry-dual-modes] Finite-context hidden states are well approximated as a simplex convex combination of task vectors whose coefficients closely track the exact Bayesian posterior as context accumulates [yan-yang-zhong-2026-task-vector-geometry-dual-modes] Substituting the task-vector mixture component with a target simplex point steers outputs to the theoretical mixture (KL 0.11 to 0.03 in E1) [yan-yang-zhong-2026-task-vector-geometry-dual-modes] Out-of-distribution generalization occupies a second, near-orthogonal subspace emerging only at high task diversity, confirmed by double-dissociation ablation [yan-yang-zhong-2026-task-vector-geometry-dual-modes] A real Qwen2.5-7B projection qualitatively echoes the geometry, with in-distribution vertex convergence versus out-of-distribution orthogonality [yan-yang-zhong-2026-task-vector-geometry-dual-modes]

Context

task-vector subspace as Bayesian-posterior-weighted convex combination (properties P0-P3), near-orthogonal OOD subspace emerging only at high task diversity, simplex-intervention causal steering (KL/RMSE reduction toward target mixture), double-dissociation ablation of task-vector vs. near-orthogonal OOD subspaces, qualitative real-LLM validation (Qwen2.5-7B ID vertex convergence vs. OOD orthogonality)

Papers

Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers — Yan, Hao, Yang, Haolin, Zhong, Yiqiao2026 · arXiv:2605.03780