MATH · IN · MODELS

A probe reads answer correctness before reasoning models state it

measured in 1 paper

Zhang et al. fit probes (many converging to purely linear) on hidden states at intermediate-answer positions across six reasoning models [zhang-etal-2025-reasoning-models-know-when-theyre-right-probing-hidden-states-for-self-verification] Correctness is predicted at ROC-AUC above 0.7 with calibration error under 0.1 in every model, including for future not-yet-stated answers later in the trace [zhang-etal-2025-reasoning-models-know-when-theyre-right-probing-hidden-states-for-self-verification] The effect is much weaker in a non-reasoning baseline (Llama-3.1-8B-Instruct), tying the signal to reasoning-specific training [zhang-etal-2025-reasoning-models-know-when-theyre-right-probing-hidden-states-for-self-verification] Used as an early-exit verifier, the probe cuts inference tokens by 24% with no accuracy loss [zhang-etal-2025-reasoning-models-know-when-theyre-right-probing-hidden-states-for-self-verification]

Context

self-verification, reasoning

Papers

Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification — Zhang, Anqi, Chen, Yulin, Pan, Jane, Zhao, Chen, Panda, Aurojit, Li, Jinyang, He, He2025 · arXiv:2504.05419