MATH · IN · MODELS

Probes separate knowledge familiarity from truthfulness better than confidence

measured in 1 paper

Cheang et al. fit logistic-regression probes on subject-token, attention, and last-layer hidden states of LLaMA-3-8B (main) and Mistral-7B-v0.3 (replication) [cheang-etal-2026-do-llms-really-know-what-they-dont-know-internal-states-mainly-reflect-knowledge-recall-rather-than-truthfulness] A last-token probe reaches AUROC 0.69 distinguishing attributable from unattributable hallucination but 0.93 distinguishing unfamiliar from familiar entities; subject and attention probes show the same gap [cheang-etal-2026-do-llms-really-know-what-they-dont-know-internal-states-mainly-reflect-knowledge-recall-rather-than-truthfulness] Cluster-separability metrics (Silhouette, Davies-Bouldin on t-SNE) show knowledge-recall and truthfulness occupy measurably different regions, replicated on Mistral-7B-v0.3 [cheang-etal-2026-do-llms-really-know-what-they-dont-know-internal-states-mainly-reflect-knowledge-recall-rather-than-truthfulness]

Context

dissociating two related but distinct linearly-probed properties (entity familiarity vs. truthfulness) via differential probe accuracy rather than a single combined score, cluster-separability metrics (Silhouette, Davies-Bouldin) as a complement to probe AUROC for characterizing separation quality

Papers

Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness — Cheang, Chi Seng, Chan, Hou Pong, Zhang, Wenxuan, Deng, Yang2026 · arXiv:2510.09033