In-context task representations are k-NN-decodable and causally patchable across real LLMs and an SSM
measured in 1 paperHan, Song, Gore & Agrawal define Task Decodability, a k-NN classifier score on the hidden state immediately before the target token, and show it strongly correlates with ICL accuracy across real Llama-3.1-8B/70B, Gemma-2-2B/9B/27B, OLMo-7B (across pretraining checkpoints), and Mamba-8B on POS-tagging and bitwise-arithmetic tasks [han-song-etal-2025-emergence-of-abstractions-task-vectors-icl] Positive activation-patching interventions on well-separated task representations improve accuracy up to +14pp, while negative interventions degrade it up to -15pp, versus only +/-2-6pp for overlapping tasks such as XOR/XNOR [han-song-etal-2025-emergence-of-abstractions-task-vectors-icl] Finetuning the first 10 layers raises Task Decodability from 0.68 to 0.95 (POS) and 0.43 to 0.85 (bitwise), with accuracy gains of 37 and 24 points respectively over finetuning the last 10 layers instead [han-song-etal-2025-emergence-of-abstractions-task-vectors-icl]