Not all LLMs have consistent truth directions; atomic-statement probes generalize
measured in 1 paperBao et al. investigate whether truth directions are consistent across LLMs and how well truth probes generalize [bao-etal-2025-probing-the-geometry-of-truth] Not all LLMs exhibit consistent truth directions, with stronger and more consistent representations in more capable models, particularly under logical negation [bao-etal-2025-probing-the-geometry-of-truth] Probes trained on declarative atomic statements generalize to logical transformations, question-answering, in-context learning, and external-knowledge settings [bao-etal-2025-probing-the-geometry-of-truth] The paper demonstrates a practical application to selective question-answering [bao-etal-2025-probing-the-geometry-of-truth]
Structure
Context
truthfulness, probe-generalization
Confirmed in models
Papers
Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks — Bao, Yuntai, Zhang, Xuhong, Du, Tianyu, Zhao, Xinkui, Feng, Zhengwen, Peng, Hao, Yin, Jianwei