Probes separate knowledge familiarity from truthfulness better than confidence
measured in 1 paperCheang et al. fit logistic-regression probes on subject-token, attention, and last-layer hidden states of LLaMA-3-8B (main) and Mistral-7B-v0.3 (replication) [cheang-etal-2026-do-llms-really-know-what-they-dont-know-internal-states-mainly-reflect-knowledge-recall-rather-than-truthfulness] A last-token probe reaches AUROC 0.69 distinguishing attributable from unattributable hallucination but 0.93 distinguishing unfamiliar from familiar entities; subject and attention probes show the same gap [cheang-etal-2026-do-llms-really-know-what-they-dont-know-internal-states-mainly-reflect-knowledge-recall-rather-than-truthfulness] Cluster-separability metrics (Silhouette, Davies-Bouldin on t-SNE) show knowledge-recall and truthfulness occupy measurably different regions, replicated on Mistral-7B-v0.3 [cheang-etal-2026-do-llms-really-know-what-they-dont-know-internal-states-mainly-reflect-knowledge-recall-rather-than-truthfulness]