MATH · IN · MODELS

Crosscoder pre-'wait' feature directions set which reasoning pattern follows

measured in 1 paper

Troitskii et al. train crosscoders across DeepSeek-R1-Distill-Llama-8B and its pre-distillation base Llama-3.1-8B to model-diff the base-to-distilled transition [troitskii-etal-2025-internal-states-before-wait-modulate-reasoning] Latent attribution locates a small subset of feature directions whose activation immediately before a "wait" token predicts the following reasoning pattern [troitskii-etal-2025-internal-states-before-wait-modulate-reasoning] Intervening directly on these features causally determines which pattern follows: restarting, recalling prior knowledge, expressing uncertainty, or double-checking [troitskii-etal-2025-internal-states-before-wait-modulate-reasoning] The interventions change the qualitative reasoning pattern rather than merely correlating with it [troitskii-etal-2025-internal-states-before-wait-modulate-reasoning]

Context

pre-"wait"-token feature directions, crosscoder latent attribution, reasoning-pattern causal control

Papers

Internal states before "wait" modulate reasoning patterns — Troitskii, Dmitrii, Pal, Koyena, Wendler, Chris, McDougall, Callum Stuart, Nanda, Neel2025 · arXiv:2510.04128