Multi-demonstration ICL needs distributed rule vectors, not one task vector
measured in 1 paperZheng et al. test the single task-vector account on tasks needing multiple demonstrations (e.g. string-length categorization) in LLaMA-7B [zheng-etal-2024] Patching one query-position task vector recovers single-demonstration knowledge tasks but only near-chance accuracy on multi-demonstration tasks [zheng-etal-2024] Gradient-times-attention saliency shows information flows from each demonstration's answer-token position to the query [zheng-etal-2024] Patching all per-demonstration local rule vectors together recovers ICL-level performance, improving with more demonstrations [zheng-etal-2024] Demixed PCA shows each rule vector encodes an abstracted query-answer summary rather than raw string length [zheng-etal-2024]
Structure
Context
task vectors, in-context learning, distributed representation, information flow, rule abstraction
Confirmed in models
Papers
Label Words as Local Task Vectors in In-Context Learning — Zheng, Bowen, Ma, Ming, Lin, Zhongqiao, Yang, Tianming