A covariance-pooling probe beats verbalized confidence and catches stealth belief shifts
measured in 1 paperSarfati et al. train pooling probes on frozen residual-stream activations of Eternis-Forecaster-8B, GLM-4.5-Air, and Qwen3-8B to read out forecast correctness and confidence [sarfati-etal-2026-what-llm-forecasters-know-but-dont-say-probing-internal-representations-for-calibration-and-faithfulness] A layer-21 covariance probe halves calibration error (ECE 0.044 vs 0.093 for verbalized confidence at matched AUROC ~0.756) and holds out-of-distribution where verbalized AUROC collapses to 0.587 [sarfati-etal-2026-what-llm-forecasters-know-but-dont-say-probing-internal-representations-for-calibration-and-faithfulness] The probe's activation shift under evidence ablation/injection tracks behavioral change at Spearman rho=0.565 (vs 0.215 for reasoning-text changes), predicting the direction in 83.6% of cases including 107 stealth cases where the CoT shows no change [sarfati-etal-2026-what-llm-forecasters-know-but-dont-say-probing-internal-representations-for-calibration-and-faithfulness]