MATH · IN · MODELS

Task vectors are locally realized at certain tokens despite decodable identity

measured in 1 paper

Li et al. train linear decoders on real Gemma-3 (4B/12B/27B) and Qwen3 (4B/8B/14B) activations, finding task identity reliably decodable throughout the context [li-etal-2025-just-in-time-distributed-task-representations] The transferable, patchable task-vector representation instead comes online only at certain tokens, a two-fold locality [li-etal-2025-just-in-time-distributed-task-representations] PCA overlap between identifiable and transferable subspaces is task-dependent: 40-60% of the identifiable dimension projects onto the top-20 PCs for simple tasks versus 10-25% for list operations [li-etal-2025-just-in-time-distributed-task-representations] Patching the extracted task vector into zero-shot prompts recontextualizes them, with recovered accuracy tracking few-shot accuracy as k-shot increases [li-etal-2025-just-in-time-distributed-task-representations]

Context

task-vector token locality, identifiable vs. transferable subspace overlap, task-vector patching/recontextualization

Papers

Just-in-Time and Distributed Task Representations in Language Models — Li, Yuxuan, Campbell, Declan, Chan, Stephanie C. Y., Lampinen, Andrew Kyle2025 · arXiv:2509.04466