MATH · IN · MODELS

SAE features selected by CoT-minus-direct trigger a reasoning mode

measured in 1 paper

He et al. select SAE latent features whose first-step activation differs most between chain-of-thought and direct prompting across six models (LLaMA, Gemma-3, Qwen3) [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode] Restricting to features with consistently positive singleton-steering effect leaves 1-10 features per model [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode] Additively steering only this feature set at the first decoding step raises direct-prompt accuracy sharply (LLaMA-3.1-8B 24.5% to 73.3%; Qwen3-0.6B 7.9% to 60.6%) at far fewer tokens than full CoT [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode] A random-SAE-feature control reproduces neither the accuracy gain nor the pattern, confirming specificity [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode] The feature marks entry into a reasoning mode (transient early spike, uncorrelated with correctness) and overrides an explicit "/no_think" instruction [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode]

Context

SAE features selected by CoT-minus-direct differential mean activation at first token, singleton causal-steering sensitivity test restricting to consistently-positive-effect features, residual-injection steering avoiding SAE reconstruction-error contamination, steering a single/small feature set substantially improves direct-prompting accuracy at low token cost, transient early activation spike correlated with CoT condition, uncorrelated with final correctness, steering overrides an explicit anti-reasoning ("/no_think") prompt instruction

Papers

Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models — He, Zhenghao, Xiong, Guangzhi, Liu, Bohan, Sinha, Sanchit, Zhang, Aidong2026 · arXiv:2601.08058