MATH · IN · MODELS

A probe detects unanswerability far better than confidence, not shrinking with scale

measured in 1 paper

Wagner fits logistic-regression probes on final-prompt-token hidden states of Gemma 2 2B-it, Qwen2.5-3B/7B/14B, and Llama-3.1-8B [wagner-2026-two-axes-of-llm-abstention-answer-correctness-and-question-answerability] Hidden-state answerability readout reaches AUROC 0.97-0.99 across all five models versus only 0.54-0.67 for output-confidence readouts, a gap that does not shrink with scale [wagner-2026-two-axes-of-llm-abstention-answer-correctness-and-question-answerability] On naturally-occurring false-premise questions all confidence signals stay near chance while the probe reaches 0.69-0.77 AUROC [wagner-2026-two-axes-of-llm-abstention-answer-correctness-and-question-answerability] Routing a premise-checking instruction through the probe roughly triples challenge precision; no causal steering is performed [wagner-2026-two-axes-of-llm-abstention-answer-correctness-and-question-answerability]

Context

a "blind spot" dissociation between what a linear probe can decode from hidden states and what the model's own output confidence reports, holding across a 2B-to-14B scale range, probe-gated instruction routing as a downstream policy application of a linear-separability finding, distinct from activation steering

Papers

Two Axes of LLM Abstention: Answer Correctness and Question Answerability — Wagner, Benedikt J.2026 · arXiv:2607.08456