MATH · IN · MODELS

Contrastive-PCA feature directions predict LLM epistemic uncertainty

measured in 1 paper

Bakman et al. derive an epistemic-uncertainty bound in terms of hidden-state displacement along semantic-feature directions [bakman-etal-2025-uncertainty-as-feature-gaps-epistemic-uncertainty-quantification-of-llms] They extract three directions (context-reliance, context-comprehension, honesty) via contrastive prompt-pair PCA on Llama-3.1-8B, Mistral-7B-v0.3, and Qwen2.5-7B [bakman-etal-2025-uncertainty-as-feature-gaps-epistemic-uncertainty-quantification-of-llms] Per-token projections onto these directions improve the Prediction Rejection Ratio by up to 13 points over baselines [bakman-etal-2025-uncertainty-as-feature-gaps-epistemic-uncertainty-quantification-of-llms] Replacing PCA extraction with a plain mean-difference direction substantially degrades performance, evidencing the specific extracted direction does the work [bakman-etal-2025-uncertainty-as-feature-gaps-epistemic-uncertainty-quantification-of-llms]

Context

epistemic-uncertainty, hallucination-detection

Papers

Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs — Bakman, Yavuz Faruk, Kang, Ethan, Huang, Duygu Nur, Yaldiz, Duygu Nur, Belém, Catarina, Zhu, Chenyang, Kumar, Nishanth, Samuel, Sungmin, Avestimehr, Salman, Liu, Sai Praneeth Karimireddy, Karimireddy, Sai Praneeth2025 · arXiv:2510.02671