Crosscoder pre-'wait' feature directions set which reasoning pattern follows
measured in 1 paperTroitskii et al. train crosscoders across DeepSeek-R1-Distill-Llama-8B and its pre-distillation base Llama-3.1-8B to model-diff the base-to-distilled transition [troitskii-etal-2025-internal-states-before-wait-modulate-reasoning] Latent attribution locates a small subset of feature directions whose activation immediately before a "wait" token predicts the following reasoning pattern [troitskii-etal-2025-internal-states-before-wait-modulate-reasoning] Intervening directly on these features causally determines which pattern follows: restarting, recalling prior knowledge, expressing uncertainty, or double-checking [troitskii-etal-2025-internal-states-before-wait-modulate-reasoning] The interventions change the qualitative reasoning pattern rather than merely correlating with it [troitskii-etal-2025-internal-states-before-wait-modulate-reasoning]