Mahalanobis cosine between truth directions predicts cross-domain generalization
measured in 1 paperYing et al. test whether truth directions range from domain-general to domain-specific across five truth types plus sycophantic/expectation-inverted lying [ying-etal-2026-the-truthfulness-spectrum-hypothesis] Linear probes generalize well pairwise except on sycophantic/inverted lying (AUROC ~0.55 and ~0.28), while joint training recovers strong performance [ying-etal-2026-the-truthfulness-spectrum-hypothesis] Mahalanobis cosine similarity between probe directions predicts cross-domain generalization far better than standard cosine (R^2=0.98 vs 0.56) [ying-etal-2026-the-truthfulness-spectrum-hypothesis] Concept-erasure isolates domain-general, domain-specific, and shared truth directions with effective dimensionality under 100-200 in an 8192-d residual stream [ying-etal-2026-the-truthfulness-spectrum-hypothesis] Steering domain-specific directions selectively suppresses incorrect-answer probability (+0.05 to +0.10) while the domain-general direction backfires (-0.07 to -0.11) [ying-etal-2026-the-truthfulness-spectrum-hypothesis]