MATH · IN · MODELS

A covariance-pooling probe beats verbalized confidence and catches stealth belief shifts

measured in 1 paper

Sarfati et al. train pooling probes on frozen residual-stream activations of Eternis-Forecaster-8B, GLM-4.5-Air, and Qwen3-8B to read out forecast correctness and confidence [sarfati-etal-2026-what-llm-forecasters-know-but-dont-say-probing-internal-representations-for-calibration-and-faithfulness] A layer-21 covariance probe halves calibration error (ECE 0.044 vs 0.093 for verbalized confidence at matched AUROC ~0.756) and holds out-of-distribution where verbalized AUROC collapses to 0.587 [sarfati-etal-2026-what-llm-forecasters-know-but-dont-say-probing-internal-representations-for-calibration-and-faithfulness] The probe's activation shift under evidence ablation/injection tracks behavioral change at Spearman rho=0.565 (vs 0.215 for reasoning-text changes), predicting the direction in 83.6% of cases including 107 stealth cases where the CoT shows no change [sarfati-etal-2026-what-llm-forecasters-know-but-dont-say-probing-internal-representations-for-calibration-and-faithfulness]

Context

calibration, chain-of-thought-faithfulness

Papers

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness — Sarfati, Raphael, Tiwari, Pratyush Ranjan, Boppana, Siddharth, Earls, Christopher J., Varadaraj, Srikar, Ho, Eric2026 · arXiv:2607.08046