MATH · IN · MODELS

A linear truth direction generalizes across datasets and is causally effective

measured in 1 paper

Marks & Tegmark curate simple true/false statement datasets and, after activation-patching-localized states, show LLaMA-2-13B/70B separate true from false along a linear PCA axis [marks-tegmark-2024] The axis generalizes across topically and structurally diverse datasets increasingly with scale and is largely distinct from a probable-vs-improbable-text direction [marks-tegmark-2024] Mass-mean probing (a diff-in-means direction with optional covariance whitening) classifies about as well as logistic regression or contrast-consistent search but yields substantially more causally effective steering directions (normalized indirect effects up to ~1.0 at 70B) [marks-tegmark-2024] Liu et al. extend the account cross-domain across 41 datasets, finding training-task diversity drives generalization far more than data volume (accuracy rising as task categories grow from 1 to 14) [marks-tegmark-2024]

Context

truth, factuality, probing, mass-mean probing, truth geometry debate

Papers

The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets — Marks, S., Tegmark, M.2024 · arXiv:2310.06824
On the Universal Truthfulness Hyperplane Inside LLMs — Liu, Junteng, Chen, Shiqi, Cheng, Yu, He, Junxian2024 · arXiv:2407.08582