Task-vector subspace implements Bayesian retrieval; OOD generalization is near-orthogonal
measured in 1 paperYan et al. train small RoPE transformers on synthetic latent-task mixtures and formalize task vectors as task-conditional-minus-global mean hidden states [yan-yang-zhong-2026-task-vector-geometry-dual-modes] Finite-context hidden states are well approximated as a simplex convex combination of task vectors whose coefficients closely track the exact Bayesian posterior as context accumulates [yan-yang-zhong-2026-task-vector-geometry-dual-modes] Substituting the task-vector mixture component with a target simplex point steers outputs to the theoretical mixture (KL 0.11 to 0.03 in E1) [yan-yang-zhong-2026-task-vector-geometry-dual-modes] Out-of-distribution generalization occupies a second, near-orthogonal subspace emerging only at high task diversity, confirmed by double-dissociation ablation [yan-yang-zhong-2026-task-vector-geometry-dual-modes] A real Qwen2.5-7B projection qualitatively echoes the geometry, with in-distribution vertex convergence versus out-of-distribution orthogonality [yan-yang-zhong-2026-task-vector-geometry-dual-modes]