A linear truth direction generalizes across datasets and is causally effective
measured in 1 paperMarks & Tegmark curate simple true/false statement datasets and, after activation-patching-localized states, show LLaMA-2-13B/70B separate true from false along a linear PCA axis [marks-tegmark-2024] The axis generalizes across topically and structurally diverse datasets increasingly with scale and is largely distinct from a probable-vs-improbable-text direction [marks-tegmark-2024] Mass-mean probing (a diff-in-means direction with optional covariance whitening) classifies about as well as logistic regression or contrast-consistent search but yields substantially more causally effective steering directions (normalized indirect effects up to ~1.0 at 70B) [marks-tegmark-2024] Liu et al. extend the account cross-domain across 41 datasets, finding training-task diversity drives generalization far more than data volume (accuracy rising as task categories grow from 1 to 14) [marks-tegmark-2024]