Reasoning behaviors are distinct but correlated steerable linear directions
measured in 1 paperVenhoff et al. study DeepSeek-R1-Distill reasoning models (Qwen-1.5B, Qwen-14B, Llama-8B), auto-labeling reasoning sentences into behaviors like Backtracking, Uncertainty Estimation, and Example Testing [venhoff-etal-2025-reasoning-steering-vectors] Each reasoning behavior is mediated by a distinct linear direction in the residual stream [venhoff-etal-2025-reasoning-steering-vectors] The behavior directions are correlated with one another yet remain separable [venhoff-etal-2025-reasoning-steering-vectors] Adding a behavior's steering vector causally increases the corresponding reasoning behavior during generation [venhoff-etal-2025-reasoning-steering-vectors]
Structure
Context
reasoning-behavior steering vectors, difference-in-means direction extraction, attribution-patching layer selection, cross-behavior cosine-similarity structure, bidirectional (add/subtract) steering validation
Confirmed in models
Papers
Understanding Reasoning in Thinking Language Models via Steering Vectors — Venhoff, Constantin, Arcuschin, Iván, Torr, Philip, Conmy, Arthur, Nanda, Neel