SAE features selected by CoT-minus-direct trigger a reasoning mode
measured in 1 paperHe et al. select SAE latent features whose first-step activation differs most between chain-of-thought and direct prompting across six models (LLaMA, Gemma-3, Qwen3) [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode] Restricting to features with consistently positive singleton-steering effect leaves 1-10 features per model [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode] Additively steering only this feature set at the first decoding step raises direct-prompt accuracy sharply (LLaMA-3.1-8B 24.5% to 73.3%; Qwen3-0.6B 7.9% to 60.6%) at far fewer tokens than full CoT [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode] A random-SAE-feature control reproduces neither the accuracy gain nor the pattern, confirming specificity [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode] The feature marks entry into a reasoning mode (transient early spike, uncorrelated with correctness) and overrides an explicit "/no_think" instruction [he-etal-2026-reasoning-beyond-chain-of-thought-latent-computational-mode]