Function vectors steer even where logit lens and probes cannot decode
measured in 1 paperNadaf extracts function vectors across 12 tasks and 8 prompt templates in Llama-3.1-8B, Gemma-2-9B, and Mistral-7B-v0.3 (base and instruct) [nadaf-2026-steerable-but-not-decodable] In a substantial fraction of cases (gaps up to -0.91), the function vector causally steers toward the correct answer even though the logit lens cannot decode it at any layer [nadaf-2026-steerable-but-not-decodable] Even a nonlinear probe with a selectivity control fails to decode 5/10 of the hardest cases, so the causal direction exceeds what any tested readout detects [nadaf-2026-steerable-but-not-decodable] Cross-template cosine similarity of the directions weakly correlates (r in [-0.20, 0.13]) with transfer success [nadaf-2026-steerable-but-not-decodable]
Structure
Context
function vectors, steerability-decodability dissociation, logit lens limits
Confirmed in models
Method
Papers
Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens — Nadaf, Mohammed Suhail B.