Contrastive-PCA feature directions predict LLM epistemic uncertainty
measured in 1 paperBakman et al. derive an epistemic-uncertainty bound in terms of hidden-state displacement along semantic-feature directions [bakman-etal-2025-uncertainty-as-feature-gaps-epistemic-uncertainty-quantification-of-llms] They extract three directions (context-reliance, context-comprehension, honesty) via contrastive prompt-pair PCA on Llama-3.1-8B, Mistral-7B-v0.3, and Qwen2.5-7B [bakman-etal-2025-uncertainty-as-feature-gaps-epistemic-uncertainty-quantification-of-llms] Per-token projections onto these directions improve the Prediction Rejection Ratio by up to 13 points over baselines [bakman-etal-2025-uncertainty-as-feature-gaps-epistemic-uncertainty-quantification-of-llms] Replacing PCA extraction with a plain mean-difference direction substantially degrades performance, evidencing the specific extracted direction does the work [bakman-etal-2025-uncertainty-as-feature-gaps-epistemic-uncertainty-quantification-of-llms]